> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Speech and audio

> Choose a path for transcription, speech generation, speaker diarization, and audio understanding on Baseten.

Deploy speech-to-text and text-to-speech models on dedicated infrastructure, or send audio to audio-capable Model APIs. Choose batch or real-time transcription, speech generation, or audio understanding for your application.

Start with an optimized deployment from the [Model Library](https://www.baseten.co/library/). Each model card describes its input format, supported languages, hardware, and benchmarks. For custom preprocessing or inference code, define a [Truss model package](/development/model/overview) and deploy it with the [Baseten CLI](/reference/cli/baseten/model#push).

## Transcription

Convert speech to text with automatic speech recognition (ASR). Choose a client workflow that matches how audio reaches your application:

* **Batch transcription:** Send a complete audio file to [Whisper](https://www.baseten.co/library/whisper/) or [Qwen3-ASR](https://www.baseten.co/library/qwen-3-asr-1-7b/). Start with the [batch quickstart](/inference/audio/quickstart#transcribe-a-batch).
* **Real-time transcription:** Keep a WebSocket connection open and send audio as it arrives. Start with the [real-time quickstart](/inference/audio/quickstart#deploy-a-real-time-model) to deploy Qwen3-ASR and receive partial and final transcripts.

Real-time models have model-specific protocols. Qwen3-ASR uses JSON messages containing base64 audio; the real-time Whisper deployment accepts binary audio frames. Follow the [transcription WebSocket reference](/reference/inference-api/predict-endpoints/streaming-transcription-api) for the model you deploy.

## Speech generation

Convert text to speech (TTS) and return audio to your application. [Qwen3-TTS Streaming](https://www.baseten.co/library/qwen3-tts-12hz/) accepts text over WebSocket and returns audio incrementally. Its model card also describes voice cloning from reference audio.

Listen to speech samples in the [upstream Qwen3-TTS demo](https://huggingface.co/spaces/Qwen/Qwen3-TTS). Use the Baseten model card to select a deployment with the voice controls your application needs.

To write your own Python inference code, follow [Generate speech with Kokoro](/examples/text-to-speech). This example deploys a speech model with the Baseten CLI and returns generated audio.

## Speaker diarization

Associate time intervals in a recording with speaker labels. Use diarization when you need to separate speakers in a meeting, call, or interview. Labels such as `SPEAKER_00` distinguish speakers within the recording; they don't establish a person's identity.

For batch processing of complete recordings, use [MOSS Transcribe Diarize](/examples/models/transcription/moss-transcribe-diarize) to generate transcripts with speaker labels and timestamps. For standalone diarization, browse the [Model Library](https://www.baseten.co/library/) and review the [pyannote performance article](https://www.baseten.co/blog/how-baseten-makes-pyannotes-diarization-models-96x-faster/).

## Audio understanding

Ask a model to summarize, reason about, or answer questions about a recording. [Audio input with Model APIs](/inference/model-apis/audio) sends audio and text in a request to a hosted model. Use this path when you need a model's response to the audio, with no dedicated deployment to manage.

## Voice agents

Connect real-time transcription to an LLM, then send the LLM's response to a speech-generation model. Use [Model APIs](/inference/model-apis/overview) for the LLM stage, or deploy an LLM on [Dedicated Inference](/deployment/concepts).

Review the selected model's [reasoning controls](/inference/model-apis/reasoning#control-reasoning-depth) when evaluating response latency. Measure the full turn from the end of the user's speech to the start of response playback, alongside the [speech latency measurements](/inference/audio/performance#latency-and-quality-measurements). For framework support, see the LiveKit entry in [Integrations](/inference/integrations).

## Custom speech models

Fine-tune a speech model with Training Jobs using the ML Cookbook:

* **[Qwen3-ASR](https://github.com/basetenlabs/ml-cookbook/tree/main/examples/qwen3-asr-transformers):** Prepare audio/transcript pairs, train a checkpoint, and deploy it for transcription.
* **[Whisper](https://github.com/basetenlabs/ml-cookbook/tree/main/examples/whisper-transformers):** Fine-tune Whisper with Transformers and serve the checkpoint with the recipe's inference example.
* **[Qwen3-TTS](https://github.com/basetenlabs/ml-cookbook/tree/main/examples/qwen3-tts-transformers):** Prepare speech data, fine-tune voice generation, and serve the checkpoint with the recipe's Truss implementation.

For checkpoint deployment paths, see [Serve your trained model](/training/deployment#speech-checkpoints).

## Next steps

* [Transcribe speech](/inference/audio/quickstart) with the batch or real-time client and a supplied recording.
* [Measure voice inference performance](/inference/audio/performance) to choose concurrency and evaluate deployment cost.
