Skip to main content
Meta’s Llama 3.3 70B instruction-tuned model. Both presets run on H100:4 under vLLM from NVIDIA’s FP8 checkpoint: one tuned for low time-to-first-token, one for total token throughput.

Setup

Install the Baseten CLI and sign in, then install the OpenAI SDK.
Install and sign in to Baseten
Terminal
For other platforms or a specific version, see the Baseten CLI install reference.
Install the OpenAI SDK
Prefer not to install? Sign in with uvx truss login --browser and deploy with uvx truss push. This variant ships in 2 presets tuned for different goals: Latency for lowest time-to-first-token, and Throughput for highest tokens per second. Pick the tab that matches your workload.
This preset serves Llama 3.3 70B Instruct on H100:4 under vLLM with FP8 weights and tensor parallelism. It targets low time-to-first-token on the 70B chat model.

Hardware

H100 × 4

Engine

vLLM 0.26.0

Context

128K

Concurrency

128

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml
This config runs Llama 3.3 70B Instruct on four H100 GPUs with vLLM, loading FP8 weights from nvidia/Llama-3.3-70B-Instruct-FP8 and sharding them across the four ranks. The server holds the batch to 128 sequences and 8192 batched tokens per scheduler step, and chunked prefill keeps long prompts from stalling the requests already decoding.

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten with the Baseten CLI, or with the Truss CLI if you prefer it:
You should see output similar to:
baseten model push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible API.Now call your deployment to run inference:
main.py

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero