Transcription
Convert speech to text with automatic speech recognition (ASR). Choose a client workflow that matches how audio reaches your application:- Batch transcription: Send a complete audio file to Whisper or Qwen3-ASR. Start with the batch quickstart.
- Real-time transcription: Keep a WebSocket connection open and send audio as it arrives. Start with the real-time quickstart to deploy Qwen3-ASR and receive partial and final transcripts.
Speech generation
Convert text to speech (TTS) and return audio to your application. Qwen3-TTS Streaming accepts text over WebSocket and returns audio incrementally. Its model card also describes voice cloning from reference audio. Listen to speech samples in the upstream Qwen3-TTS demo. Use the Baseten model card to select a deployment with the voice controls your application needs. To write your own Python inference code, follow Generate speech with Kokoro. This example deploys a speech model with the Baseten CLI and returns generated audio.Speaker diarization
Associate time intervals in a recording with speaker labels. Use diarization when you need to separate speakers in a meeting, call, or interview. Labels such asSPEAKER_00 distinguish speakers within the recording; they don’t establish a person’s identity.
For batch processing of complete recordings, use MOSS Transcribe Diarize to generate transcripts with speaker labels and timestamps. For standalone diarization, browse the Model Library and review the pyannote performance article.
Audio understanding
Ask a model to summarize, reason about, or answer questions about a recording. Audio input with Model APIs sends audio and text in a request to a hosted model. Use this path when you need a model’s response to the audio, with no dedicated deployment to manage.Voice agents
Connect real-time transcription to an LLM, then send the LLM’s response to a speech-generation model. Use Model APIs for the LLM stage, or deploy an LLM on Dedicated Inference. Review the selected model’s reasoning controls when evaluating response latency. Measure the full turn from the end of the user’s speech to the start of response playback, alongside the speech latency measurements. For framework support, see the LiveKit entry in Integrations.Custom speech models
Fine-tune a speech model with Training Jobs using the ML Cookbook:- Qwen3-ASR: Prepare audio/transcript pairs, train a checkpoint, and deploy it for transcription.
- Whisper: Fine-tune Whisper with Transformers and serve the checkpoint with the recipe’s inference example.
- Qwen3-TTS: Prepare speech data, fine-tune voice generation, and serve the checkpoint with the recipe’s Truss implementation.
Next steps
- Transcribe speech with the batch or real-time client and a supplied recording.
- Measure voice inference performance to choose concurrency and evaluate deployment cost.