How to read an inference error
A failed response has two parts worth reading:- The status code: the broad category, returned in the HTTP response (
502,503, and so on). - The error body:
{"error": "<message>"}. The message is either a short string Baseten generates (such asError making prediction) or, when your model itself returned the error, the raw response your model produced, passed through unchanged.
Quick reference
Is it my model or Baseten?
When a request reaches your model and your model server responds with an error, Baseten passes that response through to you. When the request never gets a real answer from your model, Baseten returns its own error instead. Telling these apart is the fastest way to know where to debug. Your model returned the error: the status code and body come from your model server. Your handler raised an exception, returned a non-2xx status, or timed out internally. Debug it like any application bug, starting from the error body and your model logs. Your model was unreachable: the request failed in front of your container, so you get a message likeError making prediction with a 502 or 503. The container is down, restarting, was killed (for example, out of memory), or is still cold-starting.
The most common version of the second case is memory pressure. When you increase your payload or batch sizes, each request uses more memory, and the container can be killed mid-request (OOMKilled). Confirm it by checking your model logs for OOMKilled or repeated restarts, and your metrics for memory pressure and replica restarts. To fix it, reduce per-request memory with smaller batches or payloads, or move to a larger instance type.
Errors by status code
Each section below covers what the status code means, its common causes, and what to check first.HTTP 400, 401, 403, and 404: request and authentication errors
These point to the request itself, not your model.401: the request has no credentials; the body readsUnauthorized. Send a Baseten API key in theAuthorizationheader.403: either the key is invalid or revoked, or the target requires a regional endpoint. An invalid key returnsAuthentication failed; check the key’s value and that it belongs to the right workspace. If the key is valid, use the endpoint for the regional environment or regional deployment.404: the model or deployment ID doesn’t exist, or the model was deleted. Confirm the ID and that you’re calling the right predict endpoint.400: the request or URL is malformed. Check the request body is valid JSON and the path is correct.
404 when the request body’s model field doesn’t match a served model:
404 body is flat: {"error": "<message>"}. For vLLM, set model to the --served-model-name value. To list served models, call /v1/models through the sync endpoint:
~/.trussrc), see Troubleshooting inference.
HTTP 402: payment required
The account has an unresolved billing or payment issue, such as exhausted credits with no payment method on file. Check your billing and usage settings, or contact your account owner.HTTP 413: payload too large
The request body exceeded a size limit. Two limits can return a413:
- Edge limit: the ingress proxy caps every inbound request body at 100 MB and rejects larger requests before they reach your model or chain. It covers the full HTTP body, including the JSON envelope and any base64-encoded media, and isn’t configurable.
- Async limit: async requests have a smaller per-organization cap, 256 KiB (262,144 bytes) by default. The message reports both sizes:
payload size of X bytes exceeds maximum size of Y bytes.
The async payload limit is set per organization. Contact support to raise it for your organization. The 100 MB edge limit is fixed.
502, not a 413. See Is it my model or Baseten?.
HTTP 429: too many requests
The request exceeded a rate limit, or arrived when no capacity was available. Where the limit lives depends on the surface you’re calling:- Model APIs: the request exceeded the requests-per-minute limit or one of the model’s token rate limits. Check
x-ratelimit-remaining-requestsand the returned token limit headers to see which limit reached zero. Models with a combined token rate limit returnx-ratelimit-remaining-tokens. Models with split token rate limits returnx-ratelimit-remaining-uncached-input-tokensandx-ratelimit-remaining-output-tokens. See Rate limit response headers for the full header list. - Async endpoints:
/async_predictallows 12,000 requests per minute per organization, and the status and cancel endpoints allow 100 requests per second. See Async inference. - Dedicated deployments: a response with the
CAPACITY_EXCEEDEDerror code means every replica slot was full. Load shedding can also reject a request because queued payloads create routing memory pressure or the queue crosses its soft limit. See Request queuing and load shedding.
429 responses mean that backoff alone won’t restore enough capacity. For CAPACITY_EXCEEDED responses on dedicated deployments, increase max replicas or tune the concurrency target. Dedicated load shedding returns 429 for routing memory pressure or the soft queue limit and 529 for the hard queue limit.
HTTP 499: client closed request
A499 means the client disconnected before the response was written: a user closed a browser tab, a client-side timeout fired, or the calling service cancelled the request. Because the client has already gone away, a 499 shows up in your deployment’s request metrics and logs rather than as a response your code handles. The metrics count the request under status 499, and the logs record Client closed connection with the request ID.
When the client disconnects, Baseten cancels the in-flight work and aborts the running request on the inference engine, so you don’t pay for GPU time generating output nobody reads. See Request cancellation.
Occasional 499s are normal and need no server-side action. A steady stream of them usually means a client-side timeout shorter than your model’s response time, cancelling requests that would have succeeded. Set the client timeout to match your model’s expected latency. See Configure HTTP clients.
HTTP 500: internal server error
A500 has two meanings, distinguished by the error message:
Internal Server Error (in model/chainlet).: your model code raised an exception during the request. The full traceback is in your model logs; debug it like any application bug.Failed to route request: Baseten waited for a replica, but none became available before the parking timeout expired. The parking timeout is 1200 seconds by default. A cold start that takes longer than the timeout can cause this error, such as when a model loads weights after scaling from zero. Retry with exponential backoff. If it happens regularly, keep minimum replicas greater than zero to avoid scale-from-zero waits, or see Startup times for ways to make startup faster.
HTTP 502: bad gateway
A502 has two meanings, and they need different responses:
- Your container was unreachable: the most common case. The container crashed, restarted, or was killed (for example,
OOMKilled) mid-request. The body is a short message likeError making prediction. Check your model logs for a crash or restart. See Is it my model or Baseten?. - A Baseten-side error: a transient problem in the gateway or routing layer. If your model logs are clean and show no crash, retry with exponential backoff.
HTTP 503: service unavailable
The container isn’t available to take the request yet. This is usually transient. Retry with exponential backoff. Common causes:- Draining: an instance is shutting down during a deploy or scale-down. The message asks you to retry on another instance.
- Routing: the request couldn’t be routed to a workload plane, or a circuit breaker is open to protect an unhealthy upstream.
503. Baseten holds it at the routing layer until a replica is ready, then forwards it. See Request lifecycle.
Async requests return
503 rather than 502 when the async service isn’t set up on the workload plane yet. See Async inference.HTTP 504: gateway timeout
The prediction ran longer than the request timeout (1200 seconds for sync predict). Common causes are a model that’s too slow for the payload, an under-provisioned instance, or a hung request. Profile the model, raise its resources, or move long-running work to the async API, which allows up to 3600 seconds per inference attempt and can retry a timed-out attempt. If you see a timeout sooner than 1200 seconds, check your client’s own timeout. Set it to match your model’s expected response time. See Configure HTTP clients.HTTP 529: overloaded
A529 means the routing layer refused the request because the deployment couldn’t accept more work. One cause is a queue crossing its hard load-shedding limit. A 529 doesn’t indicate an account rate limit.
Retry with exponential backoff and jitter. If the response carries a Retry-After header, wait at least that long instead of using your own interval. If your client already retries 429, treat 529 the same way. Persistent 529 responses mean that backoff alone won’t restore enough capacity. Raise max replicas or tune the concurrency target.
529 isn’t a standard HTTP status code. Some clients treat unrecognized 5xx codes as fatal, so check that yours retries 529 instead of failing over.
Streaming and async errors
Not every failure arrives as a status code with a JSON body.Streaming responses
Baseten can send the200 response headers before the first output chunk. After sending those headers, it can’t change the HTTP status to report a later failure. A timeout or an unavailable model can end the stream early without an error event.
Check whether the stream completed
For Model API requests using Chat Completions or Messages, check both the stream-end marker and the reason generation stopped. The following checks cover these two formats, not the Responses API used for web search.
A token limit ends the stream with its usual marker, but the answer can be incomplete. Inspect the stop reason before accepting the result. A
200 status or a closed connection alone doesn’t establish that the stream completed.
These excerpts show raw server-sent events and omit earlier events. The OpenAI parsed iterator consumes [DONE] internally, and Anthropic’s text_stream yields only text. Use the streaming examples to read raw OpenAI events or iterate Anthropic events and check for premature termination.
- Chat Completions
- Messages
A Chat Completions stream ends with this marker:An interrupted stream might end after a content chunk like this, without the final
Stream-end marker
data: [DONE] marker:Partial output
[DONE] marks the end of the stream. Also inspect finish_reason in the preceding chunks to distinguish a normal stop from an output limit or a tool call.Recover from an interrupted stream
For Chat Completions and Messages, treat an explicit error event or a missing stream-end marker as a failed request. The stream can end without an error object, even if your SDK doesn’t raise an exception. Retrying starts a new request; it doesn’t resume the interrupted stream. Keep partial output separate from the retry’s output. Use bounded retries, and save completed results so a failed request doesn’t require rerunning an entire background job.- For dedicated deployments, check your model logs.
- For Model APIs, save the model name, timestamp, and any error details. Include the
x-baseten-request-idresponse header when available. Share these details when contacting support.
Async requests
Baseten returns submission errors immediately, including an invalid request, authentication failure, payload limit, rate limit, or unavailable async service. After Baseten accepts an async request, a later inference failure appears in the result payload’serrors array, which is empty on success. See Async inference.
Where to look next
When an error isn’t self-explanatory, these are the fastest places to confirm a cause:- Model logs: crashes, restarts,
OOMKilled, and your model’s own error output. - Metrics: memory pressure, replica count and restarts, and queue depth.
- Request lifecycle: how queuing, cold starts, and concurrency affect request handling.
- Async inference: for payloads or runtimes that shouldn’t go through the synchronous path.