Setup
Install the Baseten CLI and sign in, then install the OpenAI SDK.Install and sign in to BasetenFor other platforms or a specific version, see the Baseten CLI install reference.
- macOS or Linux
- Windows
Terminal
Install the OpenAI SDK
Hardware
H100
Engine
vLLM 0.14.1
Concurrency
32
Write the config
Create and move into the project directory:config.yaml and paste the following:
config.yaml
vllm/vllm-openai:v0.14.1 image with Microsoft’s VibeVoice plugin patches applied at startup, serving weights pre-mounted at /models/vibevoice-asr so cold starts skip the 9.2 GB Hugging Face download. The server runs in eager mode with a 32k context and up to 16 concurrent sequences, exposing the model as vibevoice on the chat completions endpoint.
Deploy
Deploy the model with the Baseten CLI:Baseten CLI
baseten model push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.
Call the model
Your deployment serves an OpenAI-compatible chat completions API at/v1/chat/completions that accepts audio inputs.
Send audio as an audio_url content item on a chat message. The model returns the transcription as the assistant message content.
- Python
- cURL
main.py
Next steps
Call your model
Endpoint anatomy, authentication, and sync versus async inference
Autoscaling
Scale replicas with traffic, including scale to zero