Skip to main content
Deploy Qwen3-ASR from the Model Library and transcribe a short recording. Choose Batch to upload the file and receive a complete transcript, or Real-time to send audio in chunks and receive partial and final transcripts over WebSocket.

Before you begin

You need a Baseten account and a Baseten API key. The real-time Python client also requires uv. Each model deployment runs on a dedicated GPU; review the instance cost before deploying. The shell commands on this page use macOS or Linux syntax. Download the supplied recording:
This creates jfk.wav, an 11-second excerpt of John F. Kennedy’s inaugural address from the Whisper test assets. The file contains 16 kHz mono, signed 16-bit pulse-code modulation (PCM16) audio. The batch request uploads the WAV file; the real-time client reads the WAV container and sends only the audio samples.

Transcribe a batch

  1. Open Qwen3-ASR in the Model Library and select Deploy Batch. Sign in if prompted.
  2. Enter a Model name, review Hardware and pricing and Autoscaling, then select Deploy.
  3. Wait for the deployment to become active. Open the model’s overview and select Copy ID next to its name.
Set your API key and the batch model ID, then upload the supplied file:
The JSON response contains the complete transcript in its text field. For the supplied recording, it transcribes Kennedy’s remarks beginning “And so, my fellow Americans.” Punctuation can vary between runs. This request sends the whole file without real-time pacing. You don’t need a WebSocket connection or the Python client below. When you’re finished, deactivate the tutorial deployment unless you plan to keep using it.

Deploy a real-time model

  1. Open Qwen3-ASR in the Model Library and select Deploy Streaming, the library’s label for the real-time deployment. Sign in if prompted.
  2. Enter a Model name, review Hardware and pricing and Autoscaling, then select Deploy.
  3. Wait for the deployment to become active. Open the model’s overview and select Copy ID next to its name.
The batch and real-time deployments use different endpoints. Use the real-time model ID for the WebSocket client below.

Connect to the endpoint

Set your API key and the model ID you copied:
These commands set environment variables without printing them. Keep the API key on your server; don’t embed it in browser code. Save the following as transcribe.py, or download the complete client. The client connects to your model’s production WebSocket endpoint, sends audio while receiving transcripts, and closes after the server finishes processing the recording. See the Qwen3-ASR API reference for the message schema and audio requirements. The client declares its Python requirement and pinned WebSocket dependency in inline script metadata. The uv run command below selects Python 3.11 and installs websockets==15.0.1 in an environment for the script. uv downloads Python if needed.

Stream the audio

Run the client with the supplied recording:
The client sends 100 ms of audio per chunk at real-time pace. It wraps each chunk in an input_audio_buffer.append JSON message with base64 audio. After the last chunk, it sends input_audio_buffer.commit to flush the remaining speech. Sending and receiving run concurrently so you can receive partial transcripts while audio is still arriving. The client waits up to 60 seconds for each response and also bounds the total session time. If the connection fails before any transcript appears, check that the real-time deployment is active and the model ID and API key belong to the same workspace. An HTTP 500 during model startup can mean the replica hasn’t finished loading; inspect the deployment’s Logs and retry after it becomes ready. A format error means your WAV differs from the supplied 16 kHz mono PCM16 recording.

Read the transcript

The client prints updates labeled with a transcription number. A partial is a revisable hypothesis for that number. A final completes that segment. Store finals once per number instead of concatenating every partial update. One run with the supplied recording produced these final lines (partial updates omitted):
A recording can produce multiple final segments. The client joins them into the last Transcript: line. Partial wording and the number of updates can vary between runs.

Finish the session

The client waits for a final message with is_end_of_audio_flush: true before closing the WebSocket. Closing immediately after sending the audio can lose the last transcript. If you interrupt the client with Ctrl+C, run the file again to get a complete result. Your real-time transcription is complete when the client prints Transcript: and exits. Closing the connection releases the session. Both batch and real-time deployments remain active according to their autoscaling settings after requests finish. For a disposable tutorial deployment, open Dedicated Inference, select your model, open its deployment, and choose Deactivate deployment. Confirm the action to stop its replicas. See Deployment lifecycle for reactivation steps.

Next steps