> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Python performance client

> Send concurrent embedding, reranking, classification, and custom HTTP requests from Python.

Call inference methods synchronously or with `asyncio`. Synchronous methods release the Python global interpreter lock while requests run.

See the [Performance Client overview](/reference/sdk/performance-client/overview) for shared batching, retry, and connection settings.

## Installation

```bash theme={"system"}
pip install baseten_performance_client==0.1.15
```

Use Python 3.8 or later. Install NumPy separately to use `OpenAIEmbeddingsResponse.numpy()`.

## First request

Set your API key, an embeddings deployment URL, and the `model` value expected by that deployment:

```bash theme={"system"}
export BASETEN_API_KEY="<YOUR_API_KEY>"
export BASETEN_BASE_URL="https://model-YOUR_MODEL_ID.api.baseten.co/environments/production/sync"
export BASETEN_MODEL="<MODEL_NAME>"
```

Save this as `embeddings.py`:

```python embeddings.py theme={"system"}
import os

from baseten_performance_client import PerformanceClient, RequestProcessingPreference

client = PerformanceClient(
    base_url=os.environ["BASETEN_BASE_URL"],
    api_key=os.environ["BASETEN_API_KEY"],
)
preference = RequestProcessingPreference(
    max_concurrent_requests=32,
    batch_size=8,
    timeout_s=30.0,
)
response = client.embed(
    input=["Hello world", "Example text"],
    model=os.environ["BASETEN_MODEL"],
    encoding_format="float",
    preference=preference,
)
print(f"{len(response.data)} embeddings")
print(f"Total tokens: {response.usage.total_tokens}")
print(f"Elapsed seconds: {response.total_time:.3f}")
```

Run `python embeddings.py`. A successful response prints `2 embeddings`, followed by token usage and elapsed time. For reranking, classification, or generic requests, set the base URL to a deployment that supports the corresponding operation.

## `PerformanceClient`

```python theme={"system"}
PerformanceClient(
    base_url,
    api_key=None,
    http_version=1,
    client_wrapper=None,
    proxy=None,
    endpoint_pool=None,
)
```

<ParamField body="base_url" type="str" required>
  Base URL for requests. The client appends the request path. For an embeddings deployment, use the URL ending in `/sync`, without `/v1/embeddings`.
</ParamField>

<ParamField body="api_key" type="str | None">
  Inference API key. Defaults to `None`, so the client checks `BASETEN_API_KEY`, then `OPENAI_API_KEY`. Construction fails if no key is available.
</ParamField>

<ParamField body="http_version" type="int" default={1}>
  `1` for HTTP/1.1 or `2` for HTTP/2.
</ParamField>

<ParamField body="client_wrapper" type="HttpClientWrapper | None">
  Shared HTTP connection pool. When supplied, its HTTP version and proxy configuration take precedence. Defaults to `None`.
</ParamField>

<ParamField body="proxy" type="str | None">
  Proxy URL for a newly created HTTP client. Defaults to `None`.
</ParamField>

<ParamField body="endpoint_pool" type="EndpointPool | None">
  Set of endpoints the client can send requests to. When supplied, these URLs replace `base_url`. Defaults to `None`.
</ParamField>

### `get_client_wrapper`

`client.get_client_wrapper()` returns the `HttpClientWrapper` used by the client. Pass it to another client to share connections.

### `api_key`

<ResponseField name="api_key" type="str">
  The API key resolved when constructing the client.
</ResponseField>

## Inference methods

Each method has a synchronous and an asynchronous version with the same arguments and response type. Asynchronous methods return awaitables for `asyncio`; they don't submit jobs to the [asynchronous inference API](/inference/async).

| Synchronous method | Asynchronous method | Response | Path appended to the base URL |
| - | - | - | - |
| `embed` | `async_embed` | `OpenAIEmbeddingsResponse` | `/v1/embeddings` |
| `rerank` | `async_rerank` | `RerankResponse` | `/rerank` |
| `classify` | `async_classify` | `ClassificationResponse` | `/predict` |
| `batch_post` | `async_batch_post` | `BatchPostResponse` | Supplied `url_path`. |

### `embed`

Generates embeddings for the input texts.

```python theme={"system"}
client.embed(
    input,
    model,
    encoding_format=None,
    dimensions=None,
    user=None,
    preference=None,
)
```

<ParamField body="input" type="list[str]" required>
  Nonempty list of texts to embed.
</ParamField>

<ParamField body="model" type="str" required>
  Nonempty value expected by the model server.
</ParamField>

<ParamField body="encoding_format" type="str | None">
  `"float"` or `"base64"`. Defaults to `None`, which lets the endpoint select the format.
</ParamField>

<ParamField body="dimensions" type="int | None">
  Requested embedding dimensions. The Python binding accepts 1 through 1,000,000; the model determines which values it supports. Defaults to `None`.
</ParamField>

<ParamField body="user" type="str | None">
  User identifier sent to the endpoint. Defaults to `None`.
</ParamField>

<ParamField body="preference" type="RequestProcessingPreference | None">
  Batching, concurrency, timeout, and retry settings. Defaults to `None`.
</ParamField>

Returns [`OpenAIEmbeddingsResponse`](#openaiembeddingsresponse), with embedding results, token usage, and timing fields.

### `async_embed`

`await client.async_embed(...)` accepts the same arguments as `embed` and returns `OpenAIEmbeddingsResponse`.

```python async_embeddings.py theme={"system"}
import asyncio
import os

from baseten_performance_client import PerformanceClient, RequestProcessingPreference


async def main():
    client = PerformanceClient(
        base_url=os.environ["BASETEN_BASE_URL"],
        api_key=os.environ["BASETEN_API_KEY"],
    )
    response = await client.async_embed(
        input=["Hello world", "Example text"],
        model=os.environ["BASETEN_MODEL"],
        encoding_format="float",
        preference=RequestProcessingPreference(timeout_s=30.0),
    )
    print(f"{len(response.data)} embeddings")


asyncio.run(main())
```

Run this script with `python async_embeddings.py`. In an application that already runs an event loop, await `async_embed` from your existing async function.

### `rerank`

Scores texts against a query.

```python theme={"system"}
client.rerank(
    query,
    texts,
    raw_scores=False,
    model=None,
    return_text=False,
    truncate=False,
    truncation_direction="Right",
    preference=None,
)
```

<ParamField body="query" type="str" required>
  Query to score the texts against.
</ParamField>

<ParamField body="texts" type="list[str]" required>
  Nonempty list of texts to rerank.
</ParamField>

<ParamField body="raw_scores" type="bool" default={false}>
  Request raw scores from the server.
</ParamField>

<ParamField body="model" type="str | None">
  Select a model when the server supports it. Defaults to `None`.
</ParamField>

<ParamField body="return_text" type="bool" default={false}>
  Include the original text in each result.
</ParamField>

<ParamField body="truncate" type="bool" default={false}>
  Let the server truncate inputs.
</ParamField>

<ParamField body="truncation_direction" type="str" default="Right">
  Direction for the server to truncate inputs.
</ParamField>

<ParamField body="preference" type="RequestProcessingPreference | None">
  Batching, concurrency, timeout, and retry settings. Defaults to `None`.
</ParamField>

The model server determines support for `model`, `raw_scores`, `return_text`, `truncate`, and `truncation_direction`.

Returns [`RerankResponse`](#rerankresponse), with an original input index and score for each result.

Configure `client` for a reranking deployment, then score the texts:

```python theme={"system"}
response = client.rerank(
    query="How do I deploy a model?",
    texts=["Deploy a model with Truss.", "Create an API key in settings."],
    return_text=True,
    preference=RequestProcessingPreference(timeout_s=30.0),
)
for result in response.data:
    print(result.index, result.score, result.text)
```

### `async_rerank`

`await client.async_rerank(...)` accepts the same arguments as `rerank` and returns `RerankResponse`.

### `classify`

Assigns labels and scores to the input texts.

```python theme={"system"}
client.classify(
    inputs,
    model=None,
    raw_scores=False,
    truncate=False,
    truncation_direction="Right",
    preference=None,
)
```

<ParamField body="inputs" type="list[str]" required>
  Nonempty list of texts to classify. The client wraps each string in a one-element list in the request's `inputs` field.
</ParamField>

<ParamField body="model" type="str | None">
  Model value sent to the server. Defaults to `None`.
</ParamField>

<ParamField body="raw_scores" type="bool" default={false}>
  Request raw scores from the server.
</ParamField>

<ParamField body="truncate" type="bool" default={false}>
  Let the server truncate inputs.
</ParamField>

<ParamField body="truncation_direction" type="str" default="Right">
  Direction for the server to truncate inputs.
</ParamField>

<ParamField body="preference" type="RequestProcessingPreference | None">
  Batching, concurrency, timeout, and retry settings. Defaults to `None`.
</ParamField>

The server interprets `model`, `raw_scores`, `truncate`, and `truncation_direction`.

Returns [`ClassificationResponse`](#classificationresponse), whose `data` contains a list of label-score results for each input.

Configure `client` for a classification deployment, then classify the texts:

```python theme={"system"}
response = client.classify(
    inputs=["The setup worked well.", "The request failed."],
    preference=RequestProcessingPreference(timeout_s=30.0),
)
for group in response.data:
    for result in group:
        print(result.label, result.score)
```

### `async_classify`

`await client.async_classify(...)` accepts the same arguments as `classify` and returns `ClassificationResponse`.

### `batch_post`

Sends one HTTP request per payload and returns responses in input order.

```python theme={"system"}
client.batch_post(url_path, payloads, preference=None, method=None)
```

<ParamField body="url_path" type="str" required>
  Path appended to the selected base URL, such as `/predict`.
</ParamField>

<ParamField body="payloads" type="list" required>
  Nonempty list of JSON-compatible payloads.
</ParamField>

<ParamField body="preference" type="RequestProcessingPreference | None">
  Concurrency, timeout, retry, and header settings. Defaults to `None`.
</ParamField>

<ParamField body="method" type="str | None">
  Defaults to `None`, which uses `"POST"`. Also accepts `"GET"`, `"PUT"`, `"PATCH"`, `"DELETE"`, `"HEAD"`, and `"OPTIONS"`. Use uppercase names.
</ParamField>

The client sends request bodies for POST, PUT, and PATCH. `batch_size` and `max_chars_per_request` don't combine the payloads into a single request. For DELETE, HEAD, and OPTIONS, the response data contains empty dictionaries rather than decoded response bodies.

<Note>
  In `0.1.15`, `batch_post` and `async_batch_post` don't accept `custom_headers`, despite the argument appearing in the package's type hints. Set headers with `RequestProcessingPreference(extra_headers={"x-request-source": "batch-job"})`.
</Note>

Configure `client` for a deployment that serves `/v1/completions`, then send non-streaming completion requests:

```python theme={"system"}
payloads = [
    {"model": os.environ["BASETEN_MODEL"], "prompt": prompt, "stream": False}
    for prompt in ["Explain embeddings.", "Explain reranking."]
]
response = client.batch_post(
    url_path="/v1/completions",
    payloads=payloads,
    preference=RequestProcessingPreference(
        max_concurrent_requests=32,
        timeout_s=60.0,
        extra_headers={"x-request-source": "batch-job"},
    ),
)
for body, headers in zip(response.data, response.response_headers):
    print(body, headers)
```

### `async_batch_post`

`await client.async_batch_post(url_path, payloads, preference=None, method=None)` returns `BatchPostResponse`.

## Request preferences

Pass a `RequestProcessingPreference` to an inference method to set batching, concurrency, timeouts, and retries for that call. Use keyword arguments when creating preferences.

```python theme={"system"}
from baseten_performance_client import RequestProcessingPreference

preference = RequestProcessingPreference(
    max_concurrent_requests=64,
    batch_size=8,
    timeout_s=30.0,
)
```

`RequestProcessingPreference()` and `RequestProcessingPreference.default()` apply the defaults below. Fields are readable and writable. Passing `None` to an optional constructor argument uses its default.

<ParamField body="max_concurrent_requests" type="int" default={256}>
  Maximum concurrent requests. Must be 1 to 1,024, or 1 to 512 when `batch_size` is below 16.
</ParamField>

<ParamField body="batch_size" type="int" default={8}>
  Maximum inputs per embedding, reranking, or classification request. Accepts 1 to 1,024.
</ParamField>

<ParamField body="timeout_s" type="float" default={3600}>
  Timeout for each request, from 0.1 to 3,600 seconds.
</ParamField>

<ParamField body="max_chars_per_request" type="int" default={8000}>
  Character threshold for splitting text batches. Accepts 50 to 1,048,576. The client sends any text that exceeds the threshold intact in its own batch.
</ParamField>

<ParamField body="pin_initial_endpoint_once" type="bool" default={false}>
  Send all initial requests in this call to one endpoint from the pool.
</ParamField>

<ParamField body="hedge_delay" type="float | None">
  Delay in seconds before sending a duplicate request. Defaults to `None`, which disables hedging. When set, must be at least 0.045 seconds and less than `timeout_s - 0.045`.
</ParamField>

<ParamField body="total_timeout_s" type="float | None">
  Timeout for the entire operation. Must be at least `timeout_s` when supplied. Defaults to `None`.
</ParamField>

<ParamField body="hedge_budget_pct" type="float">
  Fraction used to calculate the operation's hedge budget. Defaults to `0.10`, or 10%. Accepts 0 to 3.
</ParamField>

<ParamField body="retry_budget_pct" type="float">
  Fraction used to calculate the operation's budget for timeout and network retries. Defaults to `0.05`, or 5%. HTTP-status retries don't consume this budget. Accepts 0 to 3.
</ParamField>

<ParamField body="max_retries" type="int" default={5}>
  Maximum retries per request. Accepts 0 to 6. Set to `0` to turn off retries.
</ParamField>

<ParamField body="initial_backoff_ms" type="int" default={125}>
  Initial delay between retries, from 50 to 45,000 milliseconds.
</ParamField>

<ParamField body="cancel_token" type="CancellationToken | None">
  Cancellation state checked by `async_embed` before starting. Doesn't interrupt requests in flight. Defaults to `None`. See [`CancellationToken`](#cancellationtoken) for method-specific behavior.
</ParamField>

<ParamField body="primary_api_key_override" type="str | None">
  Accepts and stores a key but doesn't change request authentication in `0.1.15`. Set the key with `PerformanceClient(api_key=...)`. Defaults to `None`.
</ParamField>

<ParamField body="extra_headers" type="dict[str, str] | None">
  Additional HTTP request headers. Defaults to `None`.
</ParamField>

<ParamField body="non_retryable_status_codes" type="set[int]">
  HTTP statuses to exclude from automatic retries. Defaults to an empty set.
</ParamField>

<ParamField body="trace_context" type="TraceContext | None">
  W3C parent trace context. Defaults to `None`.
</ParamField>

The client validates request preferences when you call an operation.

## Response types

All response types include these timing and header fields. Times are in seconds.

<ResponseField name="total_time" type="float">
  Time for the overall operation.
</ResponseField>

<ResponseField name="individual_request_times" type="list[float]">
  Time for each batch request.
</ResponseField>

<ResponseField name="response_headers" type="list[dict[str, str]]">
  Response headers for each batch request.
</ResponseField>

### `OpenAIEmbeddingsResponse`

<ResponseField name="object" type="str">
  Response object type.
</ResponseField>

<ResponseField name="model" type="str">
  Model identifier returned by the server.
</ResponseField>

<ResponseField name="data" type="list[OpenAIEmbeddingData]">
  Embedding results. See [`OpenAIEmbeddingData`](#openaiembeddingdata) for each result's fields.
</ResponseField>

<ResponseField name="usage" type="OpenAIUsage">
  Token counts. See [`OpenAIUsage`](#openaiusage).
</ResponseField>

`response.numpy()` converts float embeddings to a two-dimensional NumPy `float32` array. It raises `ValueError` for empty data, base64 embeddings, or inconsistent dimensions.

Install NumPy with `pip install numpy`, then convert the response from the first request:

```python theme={"system"}
vectors = response.numpy()
print(vectors.shape)
```

### `OpenAIEmbeddingData`

<ResponseField name="object" type="str">
  Embedding object type.
</ResponseField>

<ResponseField name="index" type="int">
  Position of the embedding in the response.
</ResponseField>

<ResponseField name="embedding" type="list[float] | str">
  Embedding values as floats or a base64 string, depending on the response format.
</ResponseField>

### `OpenAIUsage`

<ResponseField name="prompt_tokens" type="int">
  Number of prompt tokens.
</ResponseField>

<ResponseField name="total_tokens" type="int">
  Total token count.
</ResponseField>

### `RerankResponse`

<ResponseField name="object" type="str">
  Response object type.
</ResponseField>

<ResponseField name="data" type="list[RerankResult]">
  Reranking results. See [`RerankResult`](#rerankresult) for each result's fields.
</ResponseField>

### `RerankResult`

<ResponseField name="index" type="int">
  Original input index.
</ResponseField>

<ResponseField name="score" type="float">
  Reranking score.
</ResponseField>

<ResponseField name="text" type="str | None">
  Original input text, when included by the server.
</ResponseField>

### `ClassificationResponse`

<ResponseField name="object" type="str">
  Response object type.
</ResponseField>

<ResponseField name="data" type="list[list[ClassificationResult]]">
  A list of label-score results for each input. See [`ClassificationResult`](#classificationresult) for each result's fields.
</ResponseField>

### `ClassificationResult`

<ResponseField name="label" type="str">
  Classification label.
</ResponseField>

<ResponseField name="score" type="float">
  Classification score.
</ResponseField>

### `BatchPostResponse`

<ResponseField name="data" type="list">
  Decoded response payloads. Each element corresponds to one input payload.
</ResponseField>

## Helper types

Import these types from `baseten_performance_client`.

### `HttpClientWrapper`

`HttpClientWrapper(http_version=1, proxy=None)` creates a reusable HTTP connection pool. Pass it through `PerformanceClient(client_wrapper=...)` or to an `Endpoint`.

<ParamField body="http_version" type="int" default={1}>
  `1` for HTTP/1.1 or `2` for HTTP/2.
</ParamField>

<ParamField body="proxy" type="str | None">
  Proxy URL for the connection pool. Defaults to `None`.
</ParamField>

Set `BASETEN_SECOND_BASE_URL` to another deployment URL, then share the wrapper between clients:

```python theme={"system"}
import os

from baseten_performance_client import HttpClientWrapper, PerformanceClient

wrapper = HttpClientWrapper(http_version=1)
first = PerformanceClient(
    base_url=os.environ["BASETEN_BASE_URL"],
    api_key=os.environ["BASETEN_API_KEY"],
    client_wrapper=wrapper,
)
second = PerformanceClient(
    base_url=os.environ["BASETEN_SECOND_BASE_URL"],
    api_key=os.environ["BASETEN_API_KEY"],
    client_wrapper=wrapper,
)
```

### `CancellationToken`

`CancellationToken()` creates a token. Pass it as `RequestProcessingPreference(cancel_token=token)`. `token.cancel()` sets its cancelled state, and `token.is_cancelled()` returns that state. A cancelled token can't be reset.

In `0.1.15`, `async_embed` checks the token before starting and raises `ValueError` if it's already cancelled. The other Python inference methods don't check the token. No method polls it while requests run, so calling `cancel()` doesn't interrupt requests in flight. Use `timeout_s` and `total_timeout_s` to bound request duration.

### `TraceContext`

Use `TraceContext(traceparent, tracestate=None)` to attach requests to a parent trace. Pass it as `RequestProcessingPreference(trace_context=...)`.

<ParamField body="traceparent" type="str" required>
  W3C parent trace context.
</ParamField>

<ParamField body="tracestate" type="str | None">
  Additional W3C trace state. Defaults to `None`.
</ParamField>

The constructed object exposes these read-only properties:

<ResponseField name="traceparent" type="str">
  The supplied parent trace context.
</ResponseField>

<ResponseField name="tracestate" type="str | None">
  The supplied trace state, or `None`.
</ResponseField>

Don't also set `traceparent` or `tracestate` in `extra_headers`; the client rejects that combination.

To export client spans to an OpenTelemetry collector, configure [client tracing](/reference/sdk/performance-client/overview#client-tracing).

### `Endpoint`

`Endpoint` configures a base URL and the health checks used by an `EndpointPool`.

```python theme={"system"}
Endpoint(
    base_url,
    api_key,
    client_wrapper,
    deep_health_url=None,
    deployment_health_path=None,
    health_check_interval_s=None,
    health_check_timeout_s=None,
    health_check_retries=None,
    health_fail_on_first=False,
    deployment_timeout_is_no_vote=None,
    deep_timeout_is_no_vote=None,
)
```

<ParamField body="base_url" type="str" required>
  Endpoint base URL.
</ParamField>

<ParamField body="api_key" type="str" required>
  Key for health-check authentication. Inference requests use the client's API key.
</ParamField>

<ParamField body="client_wrapper" type="HttpClientWrapper" required>
  HTTP connection pool for health checks.
</ParamField>

<ParamField body="deep_health_url" type="str | None">
  Absolute URL for an additional health check. Defaults to `None`.
</ParamField>

<ParamField body="deployment_health_path" type="str | None">
  Relative health-check path. Defaults to `None`, which uses `/health`.
</ParamField>

<ParamField body="health_check_interval_s" type="float | None">
  Seconds between health checks. Defaults to `None`, which uses 10 seconds.
</ParamField>

<ParamField body="health_check_timeout_s" type="float | None">
  Timeout in seconds per health-check attempt. Defaults to `None`, which uses 6 seconds.
</ParamField>

<ParamField body="health_check_retries" type="int | None">
  Retries per health check. Defaults to `None`, which uses 2 retries.
</ParamField>

<ParamField body="health_fail_on_first" type="bool" default={false}>
  Stop evaluating checks after the first failure.
</ParamField>

<ParamField body="deployment_timeout_is_no_vote" type="bool | None">
  Ignore a timeout from the deployment check when deciding whether the endpoint is healthy. Defaults to `None`, which uses `True`.
</ParamField>

<ParamField body="deep_timeout_is_no_vote" type="bool | None">
  Ignore a timeout from the deep health check when deciding whether the endpoint is healthy. Defaults to `None`, which uses `True`.
</ParamField>

### `EndpointPool`

`EndpointPool(endpoints, endpoint_weights=None)` groups endpoints for request routing.

<ParamField body="endpoints" type="list[Endpoint]" required>
  Nonempty list of endpoints with distinct base URLs.
</ParamField>

<ParamField body="endpoint_weights" type="list[float] | None">
  Routing weights for the endpoints. Defaults to `None`, which gives each endpoint equal weight. Custom weights must match the endpoint count, be finite and nonnegative, and include at least one positive value.
</ParamField>

Pass the pool as `PerformanceClient(endpoint_pool=pool, ...)`. An endpoint's health-check key doesn't replace the client's inference API key.

## Errors

| Exception | Meaning |
| - | - |
| `requests.exceptions.Timeout` | Local or remote timeout. Its `args` contain `408`, the message, and `"local"` or `"remote"`. |
| `requests.exceptions.HTTPError` | HTTP failure. Its `args` contain the status code and message. |
| `ValueError` | Invalid parameters, serialization failures, network or connection errors, or cancellation. |

HTTP errors don't attach a `requests.Response` object. Read the exception's `args` instead of assuming `error.response.status_code` is available.

Handle errors from the embedding client:

```python theme={"system"}
import requests

try:
    response = client.embed(
        input=["Hello world"],
        model=os.environ["BASETEN_MODEL"],
        encoding_format="float",
        preference=RequestProcessingPreference(timeout_s=30.0),
    )
except requests.exceptions.HTTPError as error:
    print(f"HTTP {error.args[0]}: {error.args[1]}")
except requests.exceptions.Timeout as error:
    print(f"Timeout: {error}")
except ValueError as error:
    print(f"Client error: {error}")
```

## Package source

Check `baseten_performance_client.__version__` for the installed version. The [PyPI source distribution](https://pypi.org/project/baseten-performance-client/0.1.15/) includes the type hints and Rust implementation. See the [Python binding source](https://github.com/basetenlabs/truss/tree/main/baseten-performance-client/python_bindings) on GitHub for current development.
