Skip to main content
The Performance Client sends concurrent inference requests from Python, Node.js, or Rust. It batches embedding, reranking, and classification inputs and reuses HTTP connections across calls. Python and Node.js use the same Rust implementation. Use the client with a Baseten deployment or a compatible third-party endpoint, such as OpenAI. Set the base URL for the API you want to call and provide its API key. The client appends the request path, such as /v1/embeddings. To create and manage Baseten resources, use the management SDKs.

Libraries

Python

Synchronous and asynchronous inference methods.
Python reference · GitHub · PyPI

Node.js

Promise-based inference methods with TypeScript declarations.
Node.js reference · GitHub · npm

Rust

Asynchronous inference methods for Tokio applications.
Rust reference · GitHub · crates.io

Inference methods

Use the embedding, reranking, and classification methods with compatible Baseten Embeddings Inference deployments. For other JSON endpoints, such as Engine-Builder-LLM completions, use the generic request methods. They expect non-streaming responses. Set stream to false when the endpoint supports streaming.

Batching and concurrency

Embedding, reranking, and classification methods split text inputs into batches based on the text count and character threshold. A single text that exceeds the character threshold stays intact. Generic request methods send each payload as a separate request. Batch size and character thresholds don’t combine those payloads. The client returns generic responses in input order.
Configure RequestProcessingPreference with keyword arguments.
int
default:256
Maximum concurrent requests. Must be 1 to 1,024, or 1 to 512 when batch_size is below 16.
int
default:8
Maximum texts per embedding, reranking, or classification request. Accepts 1 to 1,024.

Timeouts and retries

Timeouts

float
default:3600
Timeout for each HTTP request, from 0.1 to 3,600 seconds.
float | None
Timeout for the entire method call, in seconds. Defaults to None. Must be at least timeout_s when set.
See CancellationToken for cancellation behavior and limitations.

Request headers

The client sets these headers from the per-request timeout. Servers that support them can stop work when a request expires.
string
Relative timeout in milliseconds, rounded up. A timeout of 30.5 seconds produces 30500.
string
Absolute deadline as a Unix timestamp in milliseconds.

Retries

By default, the client retries HTTP statuses 408, 409, 429, and 500 through 599. HTTP status retries don’t consume the retry budget for network and timeout failures.
int
default:5
Maximum retries per request, from 0 to 6. Set to 0 to turn off retries.
set[int]
HTTP status codes to exclude from retries. Defaults to an empty set.
int
default:125
Initial delay between retries, from 50 to 45,000 milliseconds.
The backoff delay multiplies by four after each retry and caps at 45,000 milliseconds. The client then adds a random delay of up to 99 milliseconds.

Hedging

Hedging starts a duplicate request while waiting for the original request to finish. Python and Node.js enforce a shared hedge budget for these extra attempts. Enabling hedging can make your model process the same input more than once.
float | None
Delay before starting a duplicate request, in seconds. Defaults to None, which disables hedging. When set, must be at least 0.045 seconds and less than timeout_s - 0.045.

Connections and response formats

Keep the same client for successive calls so it can reuse HTTP connections. To share connections between clients, pass the same HttpClientWrapper to their constructors. Use HTTP/1.1 for high concurrency workloads. See HTTP client configuration for connection and timeout tuning. The client requests MessagePack responses and zstd compression through its default Accept and Accept-Encoding headers, and decodes MessagePack responses. Set either header explicitly to override its default.
int
default:1
HTTP version passed to the client constructor. Use 1 for HTTP/1.1 or 2 for HTTP/2.
dict[str, str] | None
Additional HTTP request headers, set in request preferences. Defaults to None.

Client tracing

Python and Node.js can export client spans to an OpenTelemetry collector over OTLP/HTTP JSON with gzip compression. Set these environment variables before the first inference call. The client reads them once per process and ignores the application’s OTEL_* variables.
string
HTTP or HTTPS collector URL, such as http://localhost:4318. The client appends /v1/traces if the URL path doesn’t end with it. Unset by default. An unset or empty value disables span export.
string
Comma-separated name=value headers for collector requests, such as api-key=<COLLECTOR_KEY>,tenant=<TENANT>. Values are literal and aren’t URL-decoded. Unset by default, which adds no custom headers.
Use TraceContext to attach client spans to a parent trace.
An invalid collector URL or header causes inference calls to fail. Restart the process after correcting these settings. Background export failures don’t fail inference calls. Export is best effort, and the client doesn’t flush pending spans on process exit.

Environment variables

string
Default API key when you don’t pass one to the client.
string
Fallback API key when BASETEN_API_KEY is absent.
string
default:"perfclient"
Prefix for request IDs.
string
default:"warn"
Log level. Takes precedence over RUST_LOG. Accepts trace, debug, info, warn, or error.

Benchmarks

Baseten benchmarks measured more than 1,200 requests per second from one client. Performance Client throughput benchmark comparison