Skip to main content
Deploy speech-to-text and text-to-speech models on dedicated infrastructure, or send audio to audio-capable Model APIs. Choose batch or real-time transcription, speech generation, or audio understanding for your application. Start with an optimized deployment from the Model Library. Each model card describes its input format, supported languages, hardware, and benchmarks. For custom preprocessing or inference code, define a Truss model package and deploy it with the Baseten CLI.

Transcription

Convert speech to text with automatic speech recognition (ASR). Choose a client workflow that matches how audio reaches your application:
  • Batch transcription: Send a complete audio file to Whisper or Qwen3-ASR. Start with the batch quickstart.
  • Real-time transcription: Keep a WebSocket connection open and send audio as it arrives. Start with the real-time quickstart to deploy Qwen3-ASR and receive partial and final transcripts.
Real-time models have model-specific protocols. Qwen3-ASR uses JSON messages containing base64 audio; the real-time Whisper deployment accepts binary audio frames. Follow the transcription WebSocket reference for the model you deploy.

Speech generation

Convert text to speech (TTS) and return audio to your application. Qwen3-TTS Streaming accepts text over WebSocket and returns audio incrementally. Its model card also describes voice cloning from reference audio. Listen to speech samples in the upstream Qwen3-TTS demo. Use the Baseten model card to select a deployment with the voice controls your application needs. To write your own Python inference code, follow Generate speech with Kokoro. This example deploys a speech model with the Baseten CLI and returns generated audio.

Speaker diarization

Associate time intervals in a recording with speaker labels. Use diarization when you need to separate speakers in a meeting, call, or interview. Labels such as SPEAKER_00 distinguish speakers within the recording; they don’t establish a person’s identity. For batch processing of complete recordings, use MOSS Transcribe Diarize to generate transcripts with speaker labels and timestamps. For standalone diarization, browse the Model Library and review the pyannote performance article.

Audio understanding

Ask a model to summarize, reason about, or answer questions about a recording. Audio input with Model APIs sends audio and text in a request to a hosted model. Use this path when you need a model’s response to the audio, with no dedicated deployment to manage.

Voice agents

Connect real-time transcription to an LLM, then send the LLM’s response to a speech-generation model. Use Model APIs for the LLM stage, or deploy an LLM on Dedicated Inference. Review the selected model’s reasoning controls when evaluating response latency. Measure the full turn from the end of the user’s speech to the start of response playback, alongside the speech latency measurements. For framework support, see the LiveKit entry in Integrations.

Custom speech models

Fine-tune a speech model with Training Jobs using the ML Cookbook:
  • Qwen3-ASR: Prepare audio/transcript pairs, train a checkpoint, and deploy it for transcription.
  • Whisper: Fine-tune Whisper with Transformers and serve the checkpoint with the recipe’s inference example.
  • Qwen3-TTS: Prepare speech data, fine-tune voice generation, and serve the checkpoint with the recipe’s Truss implementation.
For checkpoint deployment paths, see Serve your trained model.

Next steps