Skip to main content

Setup

Sign in to Baseten with Truss, then install the OpenAI SDK.
Sign in to Baseten
Install the OpenAI SDK
Pick the model you want to deploy. Each tab is a self-contained recipe.
openai/gpt-oss-20b is a 20B-parameter dense model with up to 128K context.This preset serves GPT-OSS 20B on a single H100 using the Harmony response format, tuned for low time-to-first-token.

Hardware

H100

Engine

TRT-LLM v2

Context

128K

Concurrency

64

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml

Key parameters

Baseten Inference Stack (BIS) reads these fields from the trt_llm block. Each one shapes how the engine is built and served:

Deploy

Push the config to Baseten:
You should see output similar to:
truss push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible API.Now call your deployment to run inference:
main.py

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero