Skip to main content
Google’s DiffusionGemma diffusion language model (26B total, 4B active), served from an FP8 quantized checkpoint.

Setup

Sign in to Baseten with Truss, then install the OpenAI SDK.
Sign in to Baseten
Install the OpenAI SDK
This preset serves DiffusionGemma 26B A4B on a single H100 with FP8 weights and dynamic activation quantization, optimized for low-latency diffusion decoding.

Hardware

H100

Engine

vLLM (nightly-2c9c07c8… build)

Context

8K

Concurrency

8

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml
This config serves the RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic checkpoint on a single H100 with vLLM built from the DiffusionGemma pull-request branch, because diffusion support has not yet landed in a vLLM release. Setting max-num-seqs to 8 and gpu-memory-utilization to 0.85 leaves the headroom that diffusion warmup needs for its logits buffers, and the deployment exposes an OpenAI-compatible chat completions endpoint under the served name google/diffusiongemma-26B-A4B-it.

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten:
You should see output similar to:
truss push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible API. Now call your deployment to run inference:
main.py

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero