Skip to main content
Ornith 1.5 35B-A3B is an FP8 mixture-of-experts reasoning model for coding and agentic tasks. It activates about 3B parameters per token and serves OpenAI-compatible multimodal chat, reasoning content, and tool calls with vLLM.

Setup

Sign in to Baseten with Truss, then install the OpenAI SDK.
Sign in to Baseten
Install the OpenAI SDK
This preset serves Ornith 1.5 35B-A3B on H100:2 with FP8 weights and tensor parallel size 2, optimized for low-latency interactive chat, reasoning, and agent workflows.

Hardware

Engine

vLLM 0.19.1

Context

256K

Concurrency

32

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml
This preset deploys the FP8 checkpoint of Ornith 1.5 35B-A3B through vLLM on two H100 GPUs with tensor parallelism. It enables prefix caching and preserves the checkpoint’s native 256K context window. The deployment exposes an OpenAI-compatible chat completions endpoint with reasoning, tool calling, and one image per prompt.

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten:
You should see output similar to:
truss push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible API. Now call your deployment to run inference:
main.py
To access the model’s chain of thought, enable thinking mode. The server parses the reasoning output into a separate reasoning_content field on the response:
To let the model call tools, pass a tools array. The server returns structured tool_calls on the response:

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero