- Request rate limits: Maximum API requests per minute.
- Token rate limits: Maximum tokens processed per minute (input + output combined). Cached input tokens count at full weight: a token served from the KV cache costs less but counts the same toward the token rate limit as an uncached one.
To raise your limits, contact us to request email verification for higher Basic limits, or to move to Pro or Enterprise.
Pricing
Model APIs bill per million tokens. For current per-model rates, see the Model APIs pricing page. Fast tier variants have their own per-model pricing and rate limits; see Supported models for details.Cached input tokens
Cached input tokens are prompt tokens served from the KV cache, billed at a discounted rate. Every request participates in caching automatically, with no flags or opt-in steps. To raise cache hits across related requests, send thex-session-affinity header with a consistent value to pin them to the same region or replica. Scope one ID to a single conversation or full agentic task, including all subagents; reusing it across unrelated agents concentrates load on one replica and defeats the load-balancing benefit. A good ID is stable per task but uncorrelated across tasks. For example, use a SHA-256 hash of your internal conversation or session ID, truncated to 32 hex characters:
x-session.py
x-session-affinity header on every request in the conversation:
Budgets
You can set a monthly budget to receive email notifications as your workspace’s total spend approaches a configured amount. If you choose to enforce a budget, Baseten rejects Model API requests once your workspace reaches the budget. Budget enforcement doesn’t affect dedicated inference or training.Usage
Track Model APIs usage withGET /v1/model_apis/usage. The endpoint returns input, cached input, uncached input, and output token counts, plus request counts, in 1m, 1h, or 1d buckets. Group results by API key, model, and service tier, or filter the response to specific values.
Usage data is available from August 5, 2026 at 20:45 UTC (
2026-08-05T20:45:00Z). Baseten did not backfill earlier Model APIs usage.Next steps
Inference errors
Handle
429 Too Many Requests and other status codesModel APIs overview
Supported models and feature support