Skip to main content
Use deployment logs to investigate startup failures and errors during inference. Capture a prediction’s request ID to find the log entries associated with that call.
For deployments built with Truss before 0.15.5, upgrade with pip install --upgrade truss and redeploy to enable per-request log filtering.

Scope by environment or deployment

The Logs tab can show entries from a single deployment or from every deployment in an environment.
  1. Sign in to app.baseten.co, choose Dedicated Inference, and select your model.
  2. Open the Logs tab.
  3. Use the scope dropdowns to select an environment or deployment, then choose the time range you want to inspect.
Environment scope aggregates logs across every deployment in that environment, including past deployments still serving traffic during a rollout. Use it to follow a request across deployment boundaries or to watch a promotion in progress. Deployment scope restricts logs to a single deployment ID. Use it to isolate behavior to one version, such as a development deployment. The same scope applies to live tail and historical search.

Events

The volume chart in the Logs tab displays platform events, such as scaling and deployment changes. Hover over a marker for details.

Get the request ID

Capture the request ID from the prediction response. HTTP responses expose it in X-Baseten-Request-Id; gRPC responses expose it in metadata. Accepted async requests also return it in the JSON body.
When you make a predict call, include the -sD- flag to print response headers alongside the body:
The request ID appears as a response header:

Filter logs by request ID

Once you have a request ID:
  1. Open your model’s Logs tab and select the environment or deployment that handled the request.
  2. Set a time range that includes the request.
  3. In Search and filter logs, enter the ID with the requestId: prefix:
  4. Open Column settings and turn on Show request ID if the column is hidden.
The view narrows to show log entries from that request. Entries with request context display the request ID alongside the replica ID when both columns are enabled.

Logging with request context

The Truss logging formatter attaches request context to Python logging records emitted during a predict call. Use a logger with the default Truss logging configuration:

Custom servers

For a custom server, extract the x-baseten-request-id header from incoming requests and include it as a top-level request_id key in your JSON log output. See the setup guides for custom HTTP servers and custom gRPC servers.

Download logs

Download logs as a file from the Logs tab. Set a historical time range to export matching entries across that range. In live view, the export starts at the oldest visible log entry. To download logs:
  1. Sign in to your workspace at app.baseten.co and choose Dedicated Inference in the sidebar, then select your model.
  2. Choose the Logs tab and set the deployment or environment scope, time range, and any filters (level, request ID, replica, or search).
  3. Open the download menu and choose Download CSV or Download JSON.
The file downloads automatically once it finishes preparing. A single download covers up to 7 days and 100,000 log lines. If you reach either limit, shorten the time range or add filters and export again.

Fetch logs from the CLI

For terminal and scripting workflows, see Fetch and stream logs. That guide covers the Baseten CLI and Management API, including filters and machine-readable output.

Export logs to an OTLP endpoint

Forward new logs to an observability backend that accepts OTLP over HTTP with Protocol Buffers encoding. Use the connection for external log search, alerting, and retention.
Log export is rolling out gradually. If the OTEL connection card isn’t visible in your settings, contact Baseten support to enable it for your organization.

What gets exported

Log export includes:
  • Build logs: image builds for new deployments.
  • Deploy and promotion logs: lifecycle events emitted as a deployment activates, scales, or is promoted to an environment.
  • Serving logs: stdout and stderr from your model replicas, including anything you write through Python’s logging module.
Records use the OTLP LogRecord format. Common attributes include the following fields. Their presence depends on the source of the log:
string
The log line.
string
ID of the model that emitted the log.
string
ID of the deployment that emitted the log.
string
Environment name when the deployment belongs to an environment.
string
Replica ID for serving logs.
string
Inference request ID when the log includes request context. Matches X-Baseten-Request-Id in the response.
string
Training job ID for training logs.
string
Chainlet ID for Chains.
string
Formatted Python traceback when the log contains an exception.
Baseten maps log levels to OTLP SeverityNumber and SeverityText and removes internal infrastructure labels before export. Exports start from the moment the connection is enabled. Historical logs aren’t backfilled. Delivery is best effort: Baseten retries transient failures, but records can be dropped when retries are exhausted or the export buffer fills.

Configure a connection

Each Baseten organization can have one OTLP destination at a time. The setting is organization-wide, so every team’s logs go to the same endpoint. To configure a connection:
  1. Sign in to your workspace at app.baseten.co and choose General settings under Organization settings in the sidebar.
  2. In the OTEL connection card, choose Add connection.
  3. For Endpoint URL, type the full HTTPS URL of your OTLP/HTTP logs receiver, including the path (/v1/logs for most receivers). For per-vendor values, see the integration notes below.
  4. For Header name, type the HTTP header your backend uses to authenticate.
  5. For Header value, type the credential for that header. Baseten stores the value encrypted and never displays it again.
  6. (Optional) Choose Add header to send up to three additional headers with every export request.
  7. Choose Save. New log records are forwarded once the connection is active.
To rotate credentials or change destinations, use the edit icon on the saved connection. Remove the connection to disable log export.

Verify the connection

Test on a saved connection sends a probe log record and reports whether your endpoint accepted it. It covers endpoints on these domains: grafana.net, datadoghq.com, datadoghq.eu, newrelic.com, honeycomb.io, dynatrace.com, elastic.co, axiom.co, coralogix.com, logz.io, signalfx.com, chronosphere.io, signoz.cloud, and sentry.io. Endpoints on any other domain receive logs normally, but Test returns an error instead of sending a probe. The probe uses OTLP/HTTP JSON. Exported logs use Protocol Buffers, so configure your receiver to accept that encoding even if the probe succeeds. Once enabled, the OTEL connection card reports an unhealthy connection when an export failed in the past 30 minutes and no export succeeded in the past 5 minutes. A connection with no delivery attempts appears healthy, so verify delivery in your backend. To verify delivery:
  1. Send a prediction that writes a log entry with request context.
  2. Capture X-Baseten-Request-Id from the response.
  3. Open your observability backend’s log search and search for that ID. Confirm that the matching log entry appears.
A prediction that emits no log entries won’t produce a matching record.

Integration notes

The endpoint and header values below come from each vendor’s OTLP/HTTP documentation. Check those docs for the most current values for your account and region.
Honeycomb accepts OTLP/HTTP at https://api.honeycomb.io/v1/logs (or a region-specific host such as https://api.eu1.honeycomb.io/v1/logs). Authenticate with an ingest API key:
string
required
https://api.honeycomb.io/v1/logs
string
required
x-honeycomb-team
string
required
Your Honeycomb ingest API key.
Check Honeycomb’s OpenTelemetry documentation for log ingestion and dataset routing. If your account requires an x-honeycomb-dataset header, add it as an additional header in the connection.
For another backend, use its HTTPS OTLP/HTTP logs endpoint and authentication header. The receiver must accept Protocol Buffers payloads.