Setup
Sign in to Baseten with Truss, then install the OpenAI SDK.Sign in to Baseten
Install the OpenAI SDK
Hardware
H100
Engine
vLLM (nightly-2c9c07c8… build)
Context
8K
Concurrency
8
Write the config
Create and move into the project directory:config.yaml and paste the following:
config.yaml
RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic checkpoint on a single H100 with vLLM built from the DiffusionGemma pull-request branch, because diffusion support has not yet landed in a vLLM release. Setting max-num-seqs to 8 and gpu-memory-utilization to 0.85 leaves the headroom that diffusion warmup needs for its logits buffers, and the deployment exposes an OpenAI-compatible chat completions endpoint under the served name google/diffusiongemma-26B-A4B-it.
Flags
Thestart_command passes these flags to the engine. Each one controls a runtime or serving behavior:
Deploy
Push the config to Baseten:truss push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.
Call the model
Your deployment serves an OpenAI-compatible API. Now call your deployment to run inference:- Python
- cURL
main.py
Next steps
Call your model
Endpoint anatomy, authentication, and sync versus async inference
Autoscaling
Scale replicas with traffic, including scale to zero