Skip to main content
Deploy models from Hugging Face directly on Baseten using vLLM and Truss. You write a config.yaml, push with the Truss CLI, and get an OpenAI-compatible API endpoint, without custom Python or a Dockerfile. This guide walks through deploying Gemma 4 26B Instruct on two H100 GPUs with vLLM, using EAGLE3 speculative decoding and prefix caching. You’ll add a Hugging Face token, write a config, deploy to Baseten, and call the model’s OpenAI-compatible endpoint. Weights mirror once through the Baseten Delivery Network (BDN), so replicas scale up without re-downloading from Hugging Face.

Install and sign in

Before you begin, sign up or sign in to Baseten, then install uv, a fast Python package manager. Install the Truss CLI and connect it to your Baseten account. Browser login opens a tab to approve this device, so there’s no API key to copy and paste.
Install Truss
Sign in
Prefer not to install? Run uvx truss login --browser to use the same flow without a permanent install, and use uvx truss … for the rest of this guide.

Add a Hugging Face access token

Gemma requires a license click-through:
  1. Accept Google’s license terms on the Gemma model page. The weights in this example come from RedHatAI’s FP8 fork; your Hugging Face token grants access to both repos.
  2. Create a read-only user access token.
  3. Save the token as a secret named hf_access_token in your Baseten workspace.

Create a Truss project

Create a directory for your project:
vLLM server deployments only need a config.yaml. Skip custom Python and the model/ directory, which you use for custom preprocessing or postprocessing.

Write the config

Create a config.yaml with:
config.yaml
Here’s what each setting does:
  • weights tells BDN which Hugging Face checkpoint to mirror and where to mount it inside the container. auth_secret_name uses your hf_access_token secret for the gated download.
  • base_image and docker_server run vLLM as the serving process: start_command launches the server, and the endpoint fields tell Baseten which routes to forward for predictions and health checks.
  • --enable-prefix-caching reuses the KV cache when requests share a prompt prefix, such as a system prompt, RAG context, or multi-turn history.
  • The --speculative-config.* flags enable EAGLE3 speculative decoding, which runs a small draft model alongside the main model and accepts matching token predictions to cut decode latency.
  • resources provisions two H100 GPUs; start_command reads the GPU count with nvidia-smi and sets vLLM’s tensor parallelism to match.
  • runtime.health_checks gives vLLM time to load weights before Baseten routes traffic or restarts the replica.
  • model_metadata supplies the example request for the dashboard Try panel, and secrets declares which workspace secrets the container can read.

Deploy

Push the model to Baseten:
You should see:
The first deploy takes 5-10 minutes while Baseten pulls the vLLM base image and BDN mirrors the FP8 weights and the EAGLE3 speculator from Hugging Face. Subsequent scale-ups reuse the cached image and weights. You can watch progress in the logs linked above.

Call the model

Once the deployment shows Active in the dashboard, call it with a Baseten API key. The endpoint follows this shape: Anatomy of the model API endpoint. In https://model-abc123.api.baseten.co/environments/production/sync/v1, abc123 is the model ID and production is the environment that serves the request. Export your key before sending the request:
Replace {model_id} in the examples below with your model ID from the deploy output.
Send a streaming chat completion with the OpenAI SDK. Save the following as call_model.py:
call_model.py
Run the script with uv, which pulls the OpenAI SDK on the fly:
You should see the response stream back token by token:
The model argument in your request must match the --served-model-name flag in start_command, or the API returns a 400.
Any code that works with the OpenAI SDK works with your deployment: point base_url at your model’s endpoint. To route traffic through a third-party OpenAI-compatible gateway, see External LLM gateways.

Adapt to another model

The same pattern works across model families: BDN handles weight delivery, vLLM serves the model, and Baseten handles replicas, routing, and monitoring. Port the template incrementally, changing and validating one layer before moving to the next.
  • Weights: Point weights[].source at the new repo and update the path in start_command. Keep auth_secret_name for gated models, and pin a revision (for example, @main or a commit hash) for reproducibility.
  • Served model name: Set --served-model-name to the public model ID your clients will send, and update the model field in example_model_input to match.
  • Model-specific vLLM flags: Swap or drop reasoning and tool-call parsers (the gemma4 parsers only apply to Gemma 4). Remove the --speculative-config.* flags if your target has no published EAGLE3 speculator.
  • Hardware: Resize resources.accelerator for the new checkpoint’s memory footprint. Confirm utilization in the deployment logs and nvidia-smi.
  • Runtime tuning: Tune runtime.predict_concurrency alongside --max-num-seqs once you know your traffic pattern.
  • Rollback: Promote a working config to a separate environment and roll forward only after smoke tests pass.

Next steps

Other weight sources

Mirror weights from S3, GCS, and other BDN sources instead of Hugging Face.

SGLang

Serve the same class of model with SGLang instead of vLLM.

Custom Docker server

Run vLLM, SGLang, Triton, or any containerized inference server on Baseten.

Autoscaling

Configure replicas, concurrency targets, and scale-to-zero for production traffic.

Customize a model

Add custom Python when you need preprocessing, postprocessing, or unsupported architectures.