Skip to main content
Z.ai GLM-5.3 Flash is a multimodal mixture-of-experts model with 321B total and 18B active parameters. This preset serves the native FP8 checkpoint on eight H100 GPUs through vLLM with a 1M-token context window, image and video inputs, always-on reasoning, automatic tool calling, and multi-token prediction speculative decoding.

Setup

Sign in to Baseten with Truss, then install the OpenAI SDK.
Sign in to Baseten
Install the OpenAI SDK
This preset serves GLM-5.3 Flash on H100:8 from its native FP8 checkpoint, with a multi-token prediction head proposing five speculative tokens per step to speed up decoding.

Hardware

H100 × 8

Engine

vLLM (glm53-flash build)

Context

1M

Concurrency

16

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml
The container loads the FP8 checkpoint to /app/checkpoint/model and serves the OpenAI-compatible API on port 8000 with tensor parallel size 8 across the eight H100 GPUs. The engine caps each request at 1,048,576 tokens and runs 16 concurrent sequences, keeping weights in FP8 while the KV cache stays in BF16, because the current implementation does not support an FP8 KV cache on Hopper. Requests that omit reasoning_effort inherit the server default of high, and each prompt accepts one image and one video alongside its text.

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten:
You should see output similar to:
truss push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible API. Now call your deployment to run inference:
main.py
The server parses the model’s chain of thought into a separate reasoning_content field on the response. Read it alongside the final answer:
To let the model call tools, pass a tools array. The server returns structured tool_calls on the response:

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero