Skip to main content
Meta’s Llama 4 Scout is a 17B-active MoE with native multimodal support and a 10M token context window.

Setup

Sign in to Baseten with Truss, then install the OpenAI SDK.
Sign in to Baseten
Install the OpenAI SDK
This preset serves Llama 4 Scout on H100:4 with a 128K serving context and native multimodal support.

Hardware

H100 × 4

Engine

vLLM (0.22.0-cu129 build)

Context

128K

Concurrency

256

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten:
You should see output similar to:
truss push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible API. Now call your deployment to run inference:
main.py

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero