> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Performance clients

> Send concurrent inference requests with the Baseten Performance Client for Python, Node.js, or Rust.

The Performance Client sends concurrent inference requests from Python, Node.js, or Rust. It batches embedding, reranking, and classification inputs and reuses HTTP connections across calls. Python and Node.js use the same Rust implementation.

Use the client with a Baseten deployment or a compatible third-party endpoint, such as OpenAI. Set the base URL for the API you want to call and provide its API key. The client appends the request path, such as `/v1/embeddings`. To create and manage Baseten resources, use the [management SDKs](/reference/sdk/overview).

## Libraries

<CardGroup cols={2}>
  <Card title="Python" icon="python" iconType="brands">
    Synchronous and asynchronous inference methods.

    ```sh theme={"system"}
    pip install baseten_performance_client
    ```

    [Python reference](/reference/sdk/performance-client/python) · [GitHub](https://github.com/basetenlabs/truss/tree/main/baseten-performance-client/python_bindings) · [PyPI](https://pypi.org/project/baseten-performance-client/)
  </Card>

  <Card title="Node.js" icon="js" iconType="brands">
    Promise-based inference methods with TypeScript declarations.

    ```sh theme={"system"}
    npm install @basetenlabs/performance-client
    ```

    [Node.js reference](/reference/sdk/performance-client/node) · [GitHub](https://github.com/basetenlabs/truss/tree/main/baseten-performance-client/node_bindings) · [npm](https://www.npmjs.com/package/@basetenlabs/performance-client)
  </Card>

  <Card title="Rust" icon="rust" iconType="brands">
    Asynchronous inference methods for Tokio applications.

    ```sh theme={"system"}
    cargo add baseten_performance_client_core
    cargo add tokio --features macros,rt-multi-thread
    ```

    [Rust reference](/reference/sdk/performance-client/rust) · [GitHub](https://github.com/basetenlabs/truss/tree/main/baseten-performance-client/core) · [crates.io](https://crates.io/crates/baseten_performance_client_core)
  </Card>
</CardGroup>

## Inference methods

| Task | Python | Node.js | Rust |
| - | - | - | - |
| Generate embeddings | [`embed` / `async_embed`](/reference/sdk/performance-client/python#embed) | [`embed`](/reference/sdk/performance-client/node#embed) | [`process_embeddings_requests`](/reference/sdk/performance-client/rust#generate-embeddings) |
| Rerank texts | [`rerank` / `async_rerank`](/reference/sdk/performance-client/python#rerank) | [`rerank`](/reference/sdk/performance-client/node#rerank) | [`process_rerank_requests`](/reference/sdk/performance-client/rust#rerank-texts) |
| Classify texts | [`classify` / `async_classify`](/reference/sdk/performance-client/python#classify) | [`classify`](/reference/sdk/performance-client/node#classify) | [`process_classify_requests`](/reference/sdk/performance-client/rust#classify-texts) |
| Send generic HTTP requests | [`batch_post` / `async_batch_post`](/reference/sdk/performance-client/python#batch_post) | [`batchPost`](/reference/sdk/performance-client/node#batchpost) | [`process_batch_post_requests`](/reference/sdk/performance-client/rust#send-generic-requests) |

Use the embedding, reranking, and classification methods with compatible [Baseten Embeddings Inference](/engines/bei/overview) deployments. For other JSON endpoints, such as [Engine-Builder-LLM](/engines/engine-builder-llm/overview) completions, use the generic request methods. They expect non-streaming responses. Set `stream` to `false` when the endpoint supports streaming.

## Batching and concurrency

Embedding, reranking, and classification methods split text inputs into batches based on the text count and character threshold. A single text that exceeds the character threshold stays intact.

Generic request methods send each payload as a separate request. Batch size and character thresholds don't combine those payloads. The client returns generic responses in input order.

<Tabs>
  <Tab title="Python">
    Configure [`RequestProcessingPreference`](/reference/sdk/performance-client/python#request-preferences) with keyword arguments.

    <ParamField body="max_concurrent_requests" type="int" default={256}>
      Maximum concurrent requests. Must be 1 to 1,024, or 1 to 512 when `batch_size` is below 16.
    </ParamField>

    <ParamField body="batch_size" type="int" default={8}>
      Maximum texts per embedding, reranking, or classification request. Accepts 1 to 1,024.
    </ParamField>
  </Tab>

  <Tab title="Node.js">
    Configure [`RequestProcessingPreference`](/reference/sdk/performance-client/node#request-preferences) with positional arguments.

    <ParamField body={"maxConcurrentRequests"} id="param-node-max-concurrent-requests" type="number | null" default={256}>
      Maximum concurrent requests. Must be 1 to 1,024, or 1 to 512 when `batchSize` is below 16.
    </ParamField>

    <ParamField body={"batchSize"} id="param-node-batch-size" type="number | null" default={8}>
      Maximum texts per embedding, reranking, or classification request. Accepts 1 to 1,024.
    </ParamField>
  </Tab>

  <Tab title="Rust">
    [`RequestProcessingPreference`](/reference/sdk/performance-client/rust#request-preferences) fields start as `None`. The client applies effective defaults when a request starts.

    <ParamField body="max_concurrent_requests" type="Option<usize>">
      Maximum concurrent requests. Effective default: `128`. Must be 1 to 1,024, or 1 to 512 when `batch_size` is below 16.
    </ParamField>

    <ParamField body="batch_size" type="Option<usize>">
      Maximum texts per embedding, reranking, or classification request. Effective default: `128`. Accepts 1 to 1,024.
    </ParamField>
  </Tab>
</Tabs>

## Timeouts and retries

### Timeouts

<Tabs>
  <Tab title="Python">
    <ParamField body="timeout_s" type="float" default={3600}>
      Timeout for each HTTP request, from 0.1 to 3,600 seconds.
    </ParamField>

    <ParamField body="total_timeout_s" type="float | None">
      Timeout for the entire method call, in seconds. Defaults to `None`. Must be at least `timeout_s` when set.
    </ParamField>

    See [`CancellationToken`](/reference/sdk/performance-client/python#cancellationtoken) for cancellation behavior and limitations.
  </Tab>

  <Tab title="Node.js">
    <ParamField body={"timeoutS"} id="param-node-timeout-s" type="number | null" default={3600}>
      Timeout for each HTTP request, from 0.1 to 3,600 seconds.
    </ParamField>

    <ParamField body={"totalTimeoutS"} id="param-node-total-timeout-s" type="number | null">
      Timeout for the entire method call, in seconds. Defaults to `null`. Must be at least `timeoutS` when set.
    </ParamField>

    See [`CancellationToken`](/reference/sdk/performance-client/node#cancellationtoken) for cancellation behavior and limitations.
  </Tab>

  <Tab title="Rust">
    <ParamField body="timeout_s" type="Option<f64>">
      Timeout for each HTTP request, from 0.1 to 3,600 seconds. Effective default: `3600.0`.
    </ParamField>

    <ParamField body="total_timeout_s" type="Option<f64>">
      Timeout for the entire method call, in seconds. Unset by default. Must be at least `timeout_s` when set.
    </ParamField>

    See [cancellation and errors](/reference/sdk/performance-client/rust#cancellation-and-errors) for how to cancel in-flight requests.
  </Tab>
</Tabs>

### Request headers

The client sets these headers from the per-request timeout. Servers that support them can stop work when a request expires.

<ParamField header="Request-Timeout-Ms" type="string">
  Relative timeout in milliseconds, rounded up. A timeout of `30.5` seconds produces `30500`.
</ParamField>

<ParamField header="Request-Deadline-Ms" type="string">
  Absolute deadline as a Unix timestamp in milliseconds.
</ParamField>

### Retries

By default, the client retries HTTP statuses `408`, `409`, `429`, and `500` through `599`. HTTP status retries don't consume the retry budget for network and timeout failures.

<Tabs>
  <Tab title="Python">
    <ParamField body="max_retries" type="int" default={5}>
      Maximum retries per request, from 0 to 6. Set to `0` to turn off retries.
    </ParamField>

    <ParamField body="non_retryable_status_codes" type="set[int]">
      HTTP status codes to exclude from retries. Defaults to an empty set.
    </ParamField>

    <ParamField body="initial_backoff_ms" type="int" default={125}>
      Initial delay between retries, from 50 to 45,000 milliseconds.
    </ParamField>
  </Tab>

  <Tab title="Node.js">
    <ParamField body={"maxRetries"} id="param-node-max-retries" type="number | null" default={5}>
      Maximum retries per request, from 0 to 6. Set to `0` to turn off retries.
    </ParamField>

    <ParamField body={"nonRetryableStatusCodes"} id="param-node-non-retryable-status-codes" type="number[] | null">
      HTTP status codes to exclude from retries. Defaults to an empty array.
    </ParamField>

    <ParamField body={"initialBackoffMs"} id="param-node-initial-backoff-ms" type="number | null" default={125}>
      Initial delay between retries, from 50 to 45,000 milliseconds.
    </ParamField>
  </Tab>

  <Tab title="Rust">
    <ParamField body="max_retries" type="Option<u32>">
      Maximum retries per request, from 0 to 6. Effective default: `5`. Set to `0` to turn off retries.
    </ParamField>

    <ParamField body="non_retryable_status_codes" type="Option<HashSet<u16>>">
      HTTP status codes to exclude from retries. Unset by default.
    </ParamField>

    <ParamField body="initial_backoff_ms" type="Option<u64>">
      Initial delay between retries, from 50 to 45,000 milliseconds. Effective default: `125`.
    </ParamField>
  </Tab>
</Tabs>

The backoff delay multiplies by four after each retry and caps at 45,000 milliseconds. The client then adds a random delay of up to 99 milliseconds.

### Hedging

Hedging starts a duplicate request while waiting for the original request to finish. Python and Node.js enforce a shared hedge budget for these extra attempts. Enabling hedging can make your model process the same input more than once.

<Tabs>
  <Tab title="Python">
    <ParamField body="hedge_delay" type="float | None">
      Delay before starting a duplicate request, in seconds. Defaults to `None`, which disables hedging. When set, must be at least 0.045 seconds and less than `timeout_s - 0.045`.
    </ParamField>
  </Tab>

  <Tab title="Node.js">
    <ParamField body={"hedgeDelay"} id="param-node-hedge-delay" type="number | null">
      Delay before starting a duplicate request, in seconds. Defaults to `null`, which disables hedging. When set, must be at least 0.045 seconds and less than `timeoutS - 0.045`.
    </ParamField>
  </Tab>

  <Tab title="Rust">
    <ParamField body="hedge_delay" type="Option<f64>">
      Delay before starting a duplicate request, in seconds. Unset by default, which disables hedging. When set, must be at least 0.045 seconds and less than `timeout_s - 0.045`.
    </ParamField>

    Retry and hedge budgets can allow attempts after exhaustion. The per-request `max_retries` limit still applies. See [Rust request preferences](/reference/sdk/performance-client/rust#request-preferences).
  </Tab>
</Tabs>

## Connections and response formats

Keep the same client for successive calls so it can reuse HTTP connections. To share connections between clients, pass the same `HttpClientWrapper` to their constructors.

Use HTTP/1.1 for high concurrency workloads. See [HTTP client configuration](/inference/http-client-configuration) for connection and timeout tuning.

The client requests MessagePack responses and zstd compression through its default `Accept` and `Accept-Encoding` headers, and decodes MessagePack responses. Set either header explicitly to override its default.

<Tabs>
  <Tab title="Python">
    <ParamField body="http_version" type="int" default={1}>
      HTTP version passed to the client constructor. Use `1` for HTTP/1.1 or `2` for HTTP/2.
    </ParamField>

    <ParamField body="extra_headers" type="dict[str, str] | None">
      Additional HTTP request headers, set in request preferences. Defaults to `None`.
    </ParamField>
  </Tab>

  <Tab title="Node.js">
    <ParamField body={"httpVersion"} id="param-node-http-version" type="number | null" default={1}>
      HTTP version passed to the client constructor. Use `1` for HTTP/1.1 or `2` for HTTP/2.
    </ParamField>

    <ParamField body={"extraHeaders"} id="param-node-extra-headers" type="Record<string, string> | null">
      Additional HTTP request headers, set in request preferences. Defaults to `null`.
    </ParamField>
  </Tab>

  <Tab title="Rust">
    <ParamField body="http_version" type="u8" required>
      HTTP version passed to the client constructor. Use `1` for HTTP/1.1 or `2` for HTTP/2.
    </ParamField>

    <ParamField body="extra_headers" type="Option<HashMap<String, String>>">
      Additional HTTP request headers, set in request preferences. Unset by default.
    </ParamField>
  </Tab>
</Tabs>

## Client tracing

Python and Node.js can export client spans to an OpenTelemetry collector over OTLP/HTTP JSON with gzip compression. Set these environment variables before the first inference call. The client reads them once per process and ignores the application's `OTEL_*` variables.

<Tabs>
  <Tab title="Python">
    <ParamField body={"BASETEN_PERFORMANCE_CLIENT_OTLP_ENDPOINT"} id="param-python-otlp-endpoint" type="string">
      HTTP or HTTPS collector URL, such as `http://localhost:4318`. The client appends `/v1/traces` if the URL path doesn't end with it. Unset by default. An unset or empty value disables span export.
    </ParamField>

    <ParamField body={"BASETEN_PERFORMANCE_CLIENT_OTLP_HEADERS"} id="param-python-otlp-headers" type="string">
      Comma-separated `name=value` headers for collector requests, such as `api-key=<COLLECTOR_KEY>,tenant=<TENANT>`. Values are literal and aren't URL-decoded. Unset by default, which adds no custom headers.
    </ParamField>

    Use [`TraceContext`](/reference/sdk/performance-client/python#tracecontext) to attach client spans to a parent trace.
  </Tab>

  <Tab title="Node.js">
    <ParamField body={"BASETEN_PERFORMANCE_CLIENT_OTLP_ENDPOINT"} id="param-node-otlp-endpoint" type="string">
      HTTP or HTTPS collector URL, such as `http://localhost:4318`. The client appends `/v1/traces` if the URL path doesn't end with it. Unset by default. An unset or empty value disables span export.
    </ParamField>

    <ParamField body={"BASETEN_PERFORMANCE_CLIENT_OTLP_HEADERS"} id="param-node-otlp-headers" type="string">
      Comma-separated `name=value` headers for collector requests, such as `api-key=<COLLECTOR_KEY>,tenant=<TENANT>`. Values are literal and aren't URL-decoded. Unset by default, which adds no custom headers.
    </ParamField>

    Use [`TraceContext`](/reference/sdk/performance-client/node#tracecontext) to attach client spans to a parent trace.
  </Tab>
</Tabs>

An invalid collector URL or header causes inference calls to fail. Restart the process after correcting these settings. Background export failures don't fail inference calls. Export is best effort, and the client doesn't flush pending spans on process exit.

## Environment variables

<ParamField body="BASETEN_API_KEY" type="string">
  Default API key when you don't pass one to the client.
</ParamField>

<ParamField body="OPENAI_API_KEY" type="string">
  Fallback API key when `BASETEN_API_KEY` is absent.
</ParamField>

<ParamField body="PERFORMANCE_CLIENT_REQUEST_ID_PREFIX" type="string" default="perfclient">
  Prefix for request IDs.
</ParamField>

<Tabs>
  <Tab title="Python">
    <ParamField body="PERFORMANCE_CLIENT_LOG_LEVEL" type="string" default="warn">
      Log level. Takes precedence over `RUST_LOG`. Accepts `trace`, `debug`, `info`, `warn`, or `error`.
    </ParamField>

    ```sh theme={"system"}
    PERFORMANCE_CLIENT_LOG_LEVEL=info python script.py
    ```
  </Tab>

  <Tab title="Node.js">
    The Node.js client doesn't initialize a logging subscriber automatically.
  </Tab>

  <Tab title="Rust">
    <ParamField body="PERFORMANCE_CLIENT_LOG_LEVEL" type="string" default="warn">
      Log level. Takes precedence over `RUST_LOG`. Accepts `trace`, `debug`, `info`, `warn`, or `error`.
    </ParamField>

    ```sh theme={"system"}
    PERFORMANCE_CLIENT_LOG_LEVEL=debug cargo run
    ```
  </Tab>
</Tabs>

## Benchmarks

[Baseten benchmarks](https://www.baseten.co/blog/your-client-code-matters-10x-higher-embedding-throughput-with-python-and-rust/) measured more than 1,200 requests per second from one client.

<img src="https://mintcdn.com/baseten-preview/W3NbEem9OZkF5rdB/images/performance-client-diagram.png?fit=max&auto=format&n=W3NbEem9OZkF5rdB&q=85&s=8b1df45b528b98df2154736919de105c" alt="Performance Client throughput benchmark comparison" width="1200" height="565" data-path="images/performance-client-diagram.png" />
