Skip to main content

Setup

Sign in to Baseten with Truss, then install the websockets library.
Sign in to Baseten
Install websockets
Pick the model you want to deploy. Each tab is a self-contained recipe.
mistralai/Voxtral-Mini-4B-Realtime-2602 is a 4B-parameter encoder-decoder model.This preset serves Voxtral Mini Realtime on an RTX PRO 6000, tuned for low-latency streaming transcription.

Hardware

RTX_PRO_6000

Engine

vLLM (0.22.0 custom build)

Context

10K

Concurrency

48

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten:
You should see output similar to:
truss push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

This preset exposes a WebSocket streaming endpoint at /v1/realtime for low-latency, incremental transcription. See the streaming transcription API reference for the message protocol, Python client example, and supported audio formats.

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero