Skip to main content
JetBrains’ Mellum2 open MoE code model (12B total, 2.5B active) with a 131k-token context window and tool calling.

Setup

Sign in to Baseten with Truss, then install the OpenAI SDK.
Sign in to Baseten
Install the OpenAI SDK
This preset serves Mellum2 12B A2.5B Instruct on a single H100 through vLLM, optimized for low-latency code generation.

Hardware

H100

Engine

vLLM 0.23.0

Context

128K

Concurrency

128

Write the config

Create and move into the project directory:
Then create a file named config.yaml and paste the following:
config.yaml
This config runs the official vllm/vllm-openai:v0.23.0 image, the release that adds MellumForCausalLM support, and streams weights from JetBrains/Mellum2-12B-A2.5B-Instruct with the Run:ai streamer. The Hermes tool-call parser enables OpenAI-compatible function calling, and a concurrency ceiling of 128 keeps the deployment throughput-friendly for coding assistants.

Flags

The start_command passes these flags to the engine. Each one controls a runtime or serving behavior:

Deploy

Push the config to Baseten:
You should see output similar to:
truss push prints your model ID (abc1d2ef in the example). The examples below use it wherever you see {model_id}, and read your API key from the BASETEN_API_KEY environment variable.

Call the model

Your deployment serves an OpenAI-compatible API. Now call your deployment to run inference:
main.py
To let the model call tools, pass a tools array. The server returns structured tool_calls on the response:

Next steps

Call your model

Endpoint anatomy, authentication, and sync versus async inference

Autoscaling

Scale replicas with traffic, including scale to zero