Setup
Install the Baseten CLI and sign in, then install the OpenAI SDK.Install and sign in to BasetenFor other platforms or a specific version, see the Baseten CLI install reference.
- macOS or Linux
- Windows
Terminal
Install the OpenAI SDK
uvx truss login --browser and deploy with uvx truss push.
This variant ships in 2 presets tuned for different goals: Latency for lowest time-to-first-token, and Throughput for highest tokens per second. Pick the tab that matches your workload.
- Latency
- Throughput
This preset serves Llama 3.3 70B Instruct on H100:4 under vLLM with FP8 weights and tensor parallelism. It targets low time-to-first-token on the 70B chat model.Then create a file named This config runs Llama 3.3 70B Instruct on four H100 GPUs with vLLM, loading FP8 weights from You should see output similar to:
Hardware
H100 × 4
Engine
vLLM 0.26.0
Context
128K
Concurrency
128
Write the config
Create and move into the project directory:config.yaml and paste the following:config.yaml
nvidia/Llama-3.3-70B-Instruct-FP8 and sharding them across the four ranks. The server holds the batch to 128 sequences and 8192 batched tokens per scheduler step, and chunked prefill keeps long prompts from stalling the requests already decoding.Flags
Thestart_command passes these flags to the engine. Each one controls a runtime or serving behavior:Deploy
Push the config to Baseten with the Baseten CLI, or with the Truss CLI if you prefer it:baseten model push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.Call the model
Your deployment serves an OpenAI-compatible API.Now call your deployment to run inference:- Python
- cURL
main.py
Next steps
Call your model
Endpoint anatomy, authentication, and sync versus async inference
Autoscaling
Scale replicas with traffic, including scale to zero