Skip to main content
Export OpenTelemetry traces from Dedicated Inference deployments that use the Truss server. Built-in spans measure request handling, and custom spans measure work inside your model. Send them to an OpenTelemetry Protocol (OTLP) collector or an observability provider that accepts OTLP/HTTP. Built-in trace export is disabled by default because it adds request overhead. Custom instrumentation has its own exporter configuration. For traces from Python or Node.js inference callers, configure Performance Client tracing. For a custom server, configure instrumentation in your server or framework.

Export built-in traces with OTLP

Send Truss’s built-in spans to an OpenTelemetry Collector or another OTLP/HTTP endpoint:
Don’t place authentication credentials in environment_variables. If your provider requires credentials, send traces through a collector that stores them securely, or use the secret-backed Honeycomb configuration below.
  1. Configure an OTLP/HTTP receiver that the deployment can reach. For example, use an OpenTelemetry Collector that exports data to your observability provider. To send traces to Datadog, configure the collector with the Datadog exporter.
  2. Set the receiver’s full trace endpoint in config.yaml and enable tracing:
    config.yaml
  3. Deploy the updated configuration and send inference requests. In your collector or observability provider, look for spans with service.name set to truss-server.
boolean
default:"false"
Enables built-in Truss trace export when you configure an exporter. This setting doesn’t control a custom tracer that your model creates.
string
Full OTLP/HTTP trace endpoint, set in environment_variables. Truss passes this URL directly to the exporter, so include /v1/traces when your collector uses the standard path. Unset by default.

Export built-in traces to Honeycomb

For debugging, Truss can send built-in traces directly to Honeycomb’s endpoint in the United States, https://api.honeycomb.io/v1/traces, using a key from Baseten secrets. For other regions, configure a collector using the OTLP setup.
  1. Add your Honeycomb Ingest Key to Baseten secrets as HONEYCOMB_API_KEY.
  2. Update config.yaml:
    config.yaml
  3. Deploy the updated configuration and send inference requests. In Honeycomb, look for spans with service.name set to truss-server.
string
Dataset header sent to Honeycomb, set in environment_variables. Set this value and declare the API key secret to enable the dedicated Honeycomb exporter. Unset by default.
string | null
Honeycomb credential. Set the configuration value to null to resolve the key from Baseten secrets instead of placing it in config.yaml.
If you also set OTEL_EXPORTER_OTLP_ENDPOINT, Truss exports built-in spans to both destinations.

Continue a distributed trace

Use your client’s OpenTelemetry context propagation to inject trace headers into inference requests. Truss uses the incoming context as the parent of its built-in request span. Without valid incoming context, it starts a new trace.
string
W3C trace context identifying the caller’s trace and parent span. Optional.
string
Optional vendor-specific trace state accompanying traceparent.
The default OpenTelemetry sampler respects the incoming parent’s sampling decision. If that parent is unsampled, Truss doesn’t export its built-in spans. For Performance Client requests, use the TraceContext parameter in Python or Node.js. Configure client span export separately from the deployment’s exporter.

Add custom spans

Use the OpenTelemetry packages included with the Truss server to create custom spans and events. This example uses the collector endpoint from the OTLP configuration and a separate provider for model spans:
model.py
Send {"text": "hello"} to receive {"text": "HELLO"}. The custom provider exports a predict span and its child transform span under service.name user-model. For additional instrumentation, see the OpenTelemetry Python guide. Truss detaches its built-in context before calling your preprocess, predict, and postprocess methods. To connect custom spans to an incoming trace, access the request headers and explicitly extract their context. For streaming methods, manage span lifetime and context inside the generator: the generator runs after the method returns, and Truss doesn’t isolate every iteration. For Chains, ChainletOptions.enable_b10_tracing enables Baseten’s internal performance tracing independently of custom instrumentation. To find log records for an inference request, use request ID filtering.