Inference endpoints
- Call an OpenAI-compatible model you deployed with the OpenAI SDK and your deployment’s base URL.
- Send JSON to your Truss model or custom server with prediction endpoints.
- Call a Chain with Chain endpoints. See Invoke a Chain for examples.
- Call a model through Model APIs with Chat Completions or Messages.
OpenAI-compatible deployments
To call an OpenAI-compatible model you deployed, set your OpenAI client’s base URL to your deployment’s endpoint:/chat/completions or /embeddings. Use the model value shown in your engine’s documentation. See OpenAI SDK setup for an example.
Supported routes and request fields depend on your engine and its version. See the documentation for BIS-LLM (Baseten Inference Stack v2), Engine-Builder-LLM, or Baseten Embeddings Inference (BEI). For a custom container, see custom server endpoint mapping.
Model and Chain endpoints
Baseten assigns each deployed model or Chain a dedicated subdomain. The hostname identifies the model or Chain; the path selects an environment or deployment and an operation. Replace{endpoint} with an operation such as predict for models or run_remote for Chains.
For models:
string
required
Model ID from your model dashboard. Used in model endpoint hostnames.
string
required
Chain ID from your Chain dashboard. Used in Chain endpoint hostnames.
string
required
Environment name, such as
staging. Used when calling a named environment.string
required
Deployment ID. Used when calling a specific deployment.
Predict endpoints
Prediction endpoints accept a JSON request body and forward it to your model’spredict function, the configured custom server route, or the Chain entrypoint.
- Models
- Chains
- Regional
Status endpoints
Wake endpoints
Timeouts
Baseten applies server-side timeouts to inference and wake operations. A prediction timeout begins after routing forwards the request for inference. A synchronous request that exceeds this timeout returns a504.
The async prediction timeout applies to each inference attempt, not to the submission request. Baseten returns
201 after placing the request in the queue, and the async retry policy can start another attempt after a retryable timeout. A synchronous request can also wait for capacity for up to the separate 1200-second default parking timeout before inference begins. For the complete sequence, see Request lifecycle.
Timeouts aren’t user-configurable. Set client timeouts based on your model’s expected response time and any expected wait for capacity. For more information, see Configure HTTP clients. For how to interpret and respond to a 504, see Inference errors.
Request size
The ingress proxy limits each request body to 100 MB. Larger requests return413 Request Entity Too Large before reaching the model or Chain. The limit includes the JSON envelope and any base64-encoded media and cannot be changed.
For large audio, video, or batched media, upload the file to S3, GCS, Azure Blob, or other object storage. Send a presigned URL in the request instead of embedding the file. The model can fetch the file directly from storage.
For how to interpret and respond to a 413, including the separate async payload limit, see Inference errors.
Model APIs
Model APIs run supported models such as DeepSeek, GLM, and Kimi on infrastructure managed by Baseten. You choose a model slug and pay per million tokens. All models support tool calling, structured outputs, and JSON mode; some also support reasoning, vision, or audio. Model APIs implement the OpenAI Chat Completions format. To update an OpenAI client, change its base URL, API key, and model slug. Requests use this endpoint:https://inference.baseten.co/v1/messages and is in beta.
Endpoint references
- WebSockets: Environment, development, and deployment endpoints.
- Audio: Transcription and streaming transcription.
- Chat Completions: OpenAI-compatible requests.
- Messages: Anthropic-compatible requests.
- Server-side tool execution: Web search and fetch tools.