> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Transcribe speech

> Deploy Qwen3-ASR for batch transcription over HTTP or real-time transcription over WebSocket.

Deploy Qwen3-ASR from the Model Library and transcribe a short recording. Choose **Batch** to upload the file and receive a complete transcript, or **Real-time** to send audio in chunks and receive partial and final transcripts over WebSocket.

## Before you begin

You need a Baseten account and a [Baseten API key](/organization/api-keys). The real-time Python client also requires [uv](https://docs.astral.sh/uv/getting-started/installation/). Each model deployment runs on a dedicated GPU; review the instance cost before deploying.

The shell commands on this page use macOS or Linux syntax.

Download the supplied recording:

```bash theme={"system"}
curl -fL https://docs.baseten.co/assets/audio/jfk.wav -o jfk.wav
```

This creates `jfk.wav`, an 11-second excerpt of John F. Kennedy's inaugural address from the [Whisper test assets](https://github.com/openai/whisper/blob/main/tests/jfk.flac). The file contains 16 kHz mono, signed 16-bit pulse-code modulation (PCM16) audio. The batch request uploads the WAV file; the real-time client reads the WAV container and sends only the audio samples.

## Transcribe a batch

1. Open [Qwen3-ASR in the Model Library](https://www.baseten.co/library/qwen-3-asr-1-7b/) and select **Deploy Batch**. Sign in if prompted.
2. Enter a **Model name**, review **Hardware and pricing** and **Autoscaling**, then select **Deploy**.
3. Wait for the deployment to become active. Open the model's overview and select **Copy ID** next to its name.

Set your API key and the batch model ID, then upload the supplied file:

```bash theme={"system"}
export BASETEN_API_KEY="YOUR_API_KEY"
export BASETEN_BATCH_MODEL_ID="YOUR_BATCH_MODEL_ID"

curl --fail-with-body --show-error \
  "https://model-${BASETEN_BATCH_MODEL_ID}.api.baseten.co/environments/production/sync/v1/audio/transcriptions" \
  -H "Authorization: Bearer ${BASETEN_API_KEY}" \
  -F "model=Qwen/Qwen3-ASR-1.7B" \
  -F "file=@jfk.wav"
```

The JSON response contains the complete transcript in its `text` field. For the supplied recording, it transcribes Kennedy's remarks beginning "And so, my fellow Americans." Punctuation can vary between runs.

This request sends the whole file without real-time pacing. You don't need a WebSocket connection or the Python client below. When you're finished, [deactivate the tutorial deployment](#finish-the-session) unless you plan to keep using it.

## Deploy a real-time model

1. Open [Qwen3-ASR in the Model Library](https://www.baseten.co/library/qwen-3-asr-1-7b/) and select **Deploy Streaming**, the library's label for the real-time deployment. Sign in if prompted.
2. Enter a **Model name**, review **Hardware and pricing** and **Autoscaling**, then select **Deploy**.
3. Wait for the deployment to become active. Open the model's overview and select **Copy ID** next to its name.

The batch and real-time deployments use different endpoints. Use the real-time model ID for the WebSocket client below.

## Connect to the endpoint

Set your API key and the model ID you copied:

```bash theme={"system"}
export BASETEN_API_KEY="YOUR_API_KEY"
export BASETEN_MODEL_ID="YOUR_MODEL_ID"
```

These commands set environment variables without printing them. Keep the API key on your server; don't embed it in browser code.

Save the following as `transcribe.py`, or [download the complete client](/assets/audio/transcribe.py). The client connects to your model's production WebSocket endpoint, sends audio while receiving transcripts, and closes after the server finishes processing the recording. See the [Qwen3-ASR API reference](/reference/inference-api/predict-endpoints/streaming-transcription-api#qwen3-asr-streaming) for the message schema and audio requirements.

The client declares its Python requirement and pinned WebSocket dependency in inline script metadata. The `uv run` command below selects Python 3.11 and installs `websockets==15.0.1` in an environment for the script. uv downloads Python if needed.

<Accordion title="Complete transcribe.py client">
  ```python theme={"system"}
  # /// script
  # requires-python = ">=3.11"
  # dependencies = ["websockets==15.0.1"]
  # ///

  """Stream a 16 kHz mono PCM16 WAV to Qwen3-ASR on Baseten.

  Requires Python 3.11+ and websockets==15.0.1.
  Usage: uv run --python 3.11 transcribe.py jfk.wav
  """

  import asyncio
  import base64
  import json
  import os
  import sys
  import wave

  from websockets.asyncio.client import connect
  from websockets.exceptions import WebSocketException

  SAMPLE_RATE = 16_000
  CHUNK_FRAMES = 1_600  # 100 ms of audio.


  def load_audio(path):
      with wave.open(path, "rb") as audio:
          if (audio.getframerate(), audio.getnchannels(), audio.getsampwidth()) != (
              SAMPLE_RATE, 1, 2
          ) or audio.getcomptype() != "NONE":
              raise ValueError("Use an uncompressed 16 kHz mono PCM16 WAV file.")
          pcm = audio.readframes(audio.getnframes())
      if not pcm:
          raise ValueError("The WAV file contains no audio.")
      return pcm


  async def transcribe(path):
      pcm = load_audio(path)
      key = os.environ["BASETEN_API_KEY"]
      model_id = os.environ["BASETEN_MODEL_ID"]
      url = f"wss://model-{model_id}.api.baseten.co/environments/production/websocket"
      headers = {"Authorization": f"Bearer {key}"}
      finals = {}

      async with connect(url, additional_headers=headers, open_timeout=60) as ws:
          await ws.send(json.dumps({
              "whisper_params": {"audio_language": "auto"},
              "streaming_params": {
                  "enable_partial_transcripts": True,
                  "partial_transcript_interval_s": 0.5,
              },
          }))

          async def send_audio():
              start = asyncio.get_running_loop().time()
              for offset in range(0, len(pcm), CHUNK_FRAMES * 2):
                  chunk = pcm[offset:offset + CHUNK_FRAMES * 2]
                  await ws.send(json.dumps({
                      "type": "input_audio_buffer.append",
                      "audio": base64.b64encode(chunk).decode("ascii"),
                  }))
                  # Pace by audio duration without accumulating network delays.
                  next_send = start + (offset + len(chunk)) / (SAMPLE_RATE * 2)
                  await asyncio.sleep(max(0, next_send - asyncio.get_running_loop().time()))
              await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))

          async def receive_transcripts():
              while True:
                  message = json.loads(await asyncio.wait_for(ws.recv(), timeout=60))
                  if message.get("type") == "error" or message.get("error"):
                      raise RuntimeError(f"Server error: {message}")
                  if message.get("type") != "transcription":
                      continue
                  number = message["transcription_num"]
                  text = " ".join(segment["text"] for segment in message.get("segments", []))
                  if message.get("is_final"):
                      finals[number] = text
                      print(f"final {number}: {text}", flush=True)
                      if message.get("is_end_of_audio_flush"):
                          return
                  else:
                      # Each update replaces the hypothesis for this number.
                      print(f"partial {number} (revisable): {text}", flush=True)

          sender = asyncio.create_task(send_audio())
          receiver = asyncio.create_task(receive_transcripts())
          try:
              # Bound the whole session even if unrelated events keep arriving.
              async with asyncio.timeout(len(pcm) / (SAMPLE_RATE * 2) + 120):
                  await asyncio.gather(sender, receiver)
          finally:
              for task in (sender, receiver):
                  task.cancel()
              await asyncio.gather(sender, receiver, return_exceptions=True)

      return " ".join(finals[number] for number in sorted(finals) if finals[number])


  if __name__ == "__main__":
      if sys.version_info < (3, 11):
          sys.exit("This client requires Python 3.11 or later.")
      if len(sys.argv) != 2:
          sys.exit("Usage: uv run --python 3.11 transcribe.py <16-kHz-mono-PCM16.wav>")
      try:
          print("Transcript:", asyncio.run(transcribe(sys.argv[1])))
      except KeyError as error:
          sys.exit(f"Missing environment variable or response field: {error}")
      except TimeoutError:
          sys.exit("Transcription timed out. Check deployment readiness and model logs.")
      except (OSError, ValueError, EOFError, wave.Error, RuntimeError, WebSocketException) as error:
          sys.exit(f"Transcription failed: {error}")
      except KeyboardInterrupt:
          sys.exit("Stopped. Run the file again to receive a complete transcript.")
  ```
</Accordion>

## Stream the audio

Run the client with the supplied recording:

```bash theme={"system"}
uv run --python 3.11 transcribe.py jfk.wav
```

The client sends 100 ms of audio per chunk at real-time pace. It wraps each chunk in an `input_audio_buffer.append` JSON message with base64 audio. After the last chunk, it sends `input_audio_buffer.commit` to flush the remaining speech.

Sending and receiving run concurrently so you can receive partial transcripts while audio is still arriving. The client waits up to 60 seconds for each response and also bounds the total session time.

If the connection fails before any transcript appears, check that the real-time deployment is active and the model ID and API key belong to the same workspace. An HTTP `500` during model startup can mean the replica hasn't finished loading; inspect the deployment's **Logs** and retry after it becomes ready. A format error means your WAV differs from the supplied 16 kHz mono PCM16 recording.

## Read the transcript

The client prints updates labeled with a transcription number. A partial is a revisable hypothesis for that number. A final completes that segment. Store finals once per number instead of concatenating every partial update.

One run with the supplied recording produced these final lines (partial updates omitted):

```text theme={"system"}
final 0: And so, my fellow Americans.
final 1: Ask not.
final 2: What your country can do for you.
final 3: Ask what you can do for your country.
Transcript: And so, my fellow Americans. Ask not. What your country can do for you. Ask what you can do for your country.
```

A recording can produce multiple final segments. The client joins them into the last `Transcript:` line. Partial wording and the number of updates can vary between runs.

## Finish the session

The client waits for a final message with `is_end_of_audio_flush: true` before closing the WebSocket. Closing immediately after sending the audio can lose the last transcript. If you interrupt the client with **Ctrl+C**, run the file again to get a complete result.

Your real-time transcription is complete when the client prints `Transcript:` and exits. Closing the connection releases the session. Both batch and real-time deployments remain active according to their autoscaling settings after requests finish.

For a disposable tutorial deployment, open **Dedicated Inference**, select your model, open its deployment, and choose **Deactivate deployment**. Confirm the action to stop its replicas. See [Deployment lifecycle](/deployment/manage/lifecycle#deactivate-a-deployment) for reactivation steps.

## Next steps

* [Measure voice inference performance](/inference/audio/performance) under your expected traffic.
* [Look up the Qwen3-ASR streaming protocol](/reference/inference-api/predict-endpoints/streaming-transcription-api#qwen3-asr-streaming) to integrate the model into your application.
* [Configure scale to zero](/deployment/manage/scaling#scale-to-zero) to preserve an endpoint that can wake for later requests.
