Skip to main content
Call inference methods synchronously or with asyncio. Synchronous methods release the Python global interpreter lock while requests run. See the Performance Client overview for shared batching, retry, and connection settings.

Installation

Use Python 3.8 or later. Install NumPy separately to use OpenAIEmbeddingsResponse.numpy().

First request

Set your API key, an embeddings deployment URL, and the model value expected by that deployment:
Save this as embeddings.py:
embeddings.py
Run python embeddings.py. A successful response prints 2 embeddings, followed by token usage and elapsed time. For reranking, classification, or generic requests, set the base URL to a deployment that supports the corresponding operation.

PerformanceClient

str
required
Base URL for requests. The client appends the request path. For an embeddings deployment, use the URL ending in /sync, without /v1/embeddings.
str | None
Inference API key. Defaults to None, so the client checks BASETEN_API_KEY, then OPENAI_API_KEY. Construction fails if no key is available.
int
default:1
1 for HTTP/1.1 or 2 for HTTP/2.
HttpClientWrapper | None
Shared HTTP connection pool. When supplied, its HTTP version and proxy configuration take precedence. Defaults to None.
str | None
Proxy URL for a newly created HTTP client. Defaults to None.
EndpointPool | None
Set of endpoints the client can send requests to. When supplied, these URLs replace base_url. Defaults to None.

get_client_wrapper

client.get_client_wrapper() returns the HttpClientWrapper used by the client. Pass it to another client to share connections.

api_key

str
The API key resolved when constructing the client.

Inference methods

Each method has a synchronous and an asynchronous version with the same arguments and response type. Asynchronous methods return awaitables for asyncio; they don’t submit jobs to the asynchronous inference API.

embed

Generates embeddings for the input texts.
list[str]
required
Nonempty list of texts to embed.
str
required
Nonempty value expected by the model server.
str | None
"float" or "base64". Defaults to None, which lets the endpoint select the format.
int | None
Requested embedding dimensions. The Python binding accepts 1 through 1,000,000; the model determines which values it supports. Defaults to None.
str | None
User identifier sent to the endpoint. Defaults to None.
RequestProcessingPreference | None
Batching, concurrency, timeout, and retry settings. Defaults to None.
Returns OpenAIEmbeddingsResponse, with embedding results, token usage, and timing fields.

async_embed

await client.async_embed(...) accepts the same arguments as embed and returns OpenAIEmbeddingsResponse.
async_embeddings.py
Run this script with python async_embeddings.py. In an application that already runs an event loop, await async_embed from your existing async function.

rerank

Scores texts against a query.
str
required
Query to score the texts against.
list[str]
required
Nonempty list of texts to rerank.
bool
default:false
Request raw scores from the server.
str | None
Select a model when the server supports it. Defaults to None.
bool
default:false
Include the original text in each result.
bool
default:false
Let the server truncate inputs.
str
default:"Right"
Direction for the server to truncate inputs.
RequestProcessingPreference | None
Batching, concurrency, timeout, and retry settings. Defaults to None.
The model server determines support for model, raw_scores, return_text, truncate, and truncation_direction. Returns RerankResponse, with an original input index and score for each result. Configure client for a reranking deployment, then score the texts:

async_rerank

await client.async_rerank(...) accepts the same arguments as rerank and returns RerankResponse.

classify

Assigns labels and scores to the input texts.
list[str]
required
Nonempty list of texts to classify. The client wraps each string in a one-element list in the request’s inputs field.
str | None
Model value sent to the server. Defaults to None.
bool
default:false
Request raw scores from the server.
bool
default:false
Let the server truncate inputs.
str
default:"Right"
Direction for the server to truncate inputs.
RequestProcessingPreference | None
Batching, concurrency, timeout, and retry settings. Defaults to None.
The server interprets model, raw_scores, truncate, and truncation_direction. Returns ClassificationResponse, whose data contains a list of label-score results for each input. Configure client for a classification deployment, then classify the texts:

async_classify

await client.async_classify(...) accepts the same arguments as classify and returns ClassificationResponse.

batch_post

Sends one HTTP request per payload and returns responses in input order.
str
required
Path appended to the selected base URL, such as /predict.
list
required
Nonempty list of JSON-compatible payloads.
RequestProcessingPreference | None
Concurrency, timeout, retry, and header settings. Defaults to None.
str | None
Defaults to None, which uses "POST". Also accepts "GET", "PUT", "PATCH", "DELETE", "HEAD", and "OPTIONS". Use uppercase names.
The client sends request bodies for POST, PUT, and PATCH. batch_size and max_chars_per_request don’t combine the payloads into a single request. For DELETE, HEAD, and OPTIONS, the response data contains empty dictionaries rather than decoded response bodies.
In 0.1.15, batch_post and async_batch_post don’t accept custom_headers, despite the argument appearing in the package’s type hints. Set headers with RequestProcessingPreference(extra_headers={"x-request-source": "batch-job"}).
Configure client for a deployment that serves /v1/completions, then send non-streaming completion requests:

async_batch_post

await client.async_batch_post(url_path, payloads, preference=None, method=None) returns BatchPostResponse.

Request preferences

Pass a RequestProcessingPreference to an inference method to set batching, concurrency, timeouts, and retries for that call. Use keyword arguments when creating preferences.
RequestProcessingPreference() and RequestProcessingPreference.default() apply the defaults below. Fields are readable and writable. Passing None to an optional constructor argument uses its default.
int
default:256
Maximum concurrent requests. Must be 1 to 1,024, or 1 to 512 when batch_size is below 16.
int
default:8
Maximum inputs per embedding, reranking, or classification request. Accepts 1 to 1,024.
float
default:3600
Timeout for each request, from 0.1 to 3,600 seconds.
int
default:8000
Character threshold for splitting text batches. Accepts 50 to 1,048,576. The client sends any text that exceeds the threshold intact in its own batch.
bool
default:false
Send all initial requests in this call to one endpoint from the pool.
float | None
Delay in seconds before sending a duplicate request. Defaults to None, which disables hedging. When set, must be at least 0.045 seconds and less than timeout_s - 0.045.
float | None
Timeout for the entire operation. Must be at least timeout_s when supplied. Defaults to None.
float
Fraction used to calculate the operation’s hedge budget. Defaults to 0.10, or 10%. Accepts 0 to 3.
float
Fraction used to calculate the operation’s budget for timeout and network retries. Defaults to 0.05, or 5%. HTTP-status retries don’t consume this budget. Accepts 0 to 3.
int
default:5
Maximum retries per request. Accepts 0 to 6. Set to 0 to turn off retries.
int
default:125
Initial delay between retries, from 50 to 45,000 milliseconds.
CancellationToken | None
Cancellation state checked by async_embed before starting. Doesn’t interrupt requests in flight. Defaults to None. See CancellationToken for method-specific behavior.
str | None
Accepts and stores a key but doesn’t change request authentication in 0.1.15. Set the key with PerformanceClient(api_key=...). Defaults to None.
dict[str, str] | None
Additional HTTP request headers. Defaults to None.
set[int]
HTTP statuses to exclude from automatic retries. Defaults to an empty set.
TraceContext | None
W3C parent trace context. Defaults to None.
The client validates request preferences when you call an operation.

Response types

All response types include these timing and header fields. Times are in seconds.
float
Time for the overall operation.
list[float]
Time for each batch request.
list[dict[str, str]]
Response headers for each batch request.

OpenAIEmbeddingsResponse

str
Response object type.
str
Model identifier returned by the server.
list[OpenAIEmbeddingData]
Embedding results. See OpenAIEmbeddingData for each result’s fields.
OpenAIUsage
Token counts. See OpenAIUsage.
response.numpy() converts float embeddings to a two-dimensional NumPy float32 array. It raises ValueError for empty data, base64 embeddings, or inconsistent dimensions. Install NumPy with pip install numpy, then convert the response from the first request:

OpenAIEmbeddingData

str
Embedding object type.
int
Position of the embedding in the response.
list[float] | str
Embedding values as floats or a base64 string, depending on the response format.

OpenAIUsage

int
Number of prompt tokens.
int
Total token count.

RerankResponse

str
Response object type.
list[RerankResult]
Reranking results. See RerankResult for each result’s fields.

RerankResult

int
Original input index.
float
Reranking score.
str | None
Original input text, when included by the server.

ClassificationResponse

str
Response object type.
list[list[ClassificationResult]]
A list of label-score results for each input. See ClassificationResult for each result’s fields.

ClassificationResult

str
Classification label.
float
Classification score.

BatchPostResponse

list
Decoded response payloads. Each element corresponds to one input payload.

Helper types

Import these types from baseten_performance_client.

HttpClientWrapper

HttpClientWrapper(http_version=1, proxy=None) creates a reusable HTTP connection pool. Pass it through PerformanceClient(client_wrapper=...) or to an Endpoint.
int
default:1
1 for HTTP/1.1 or 2 for HTTP/2.
str | None
Proxy URL for the connection pool. Defaults to None.
Set BASETEN_SECOND_BASE_URL to another deployment URL, then share the wrapper between clients:

CancellationToken

CancellationToken() creates a token. Pass it as RequestProcessingPreference(cancel_token=token). token.cancel() sets its cancelled state, and token.is_cancelled() returns that state. A cancelled token can’t be reset. In 0.1.15, async_embed checks the token before starting and raises ValueError if it’s already cancelled. The other Python inference methods don’t check the token. No method polls it while requests run, so calling cancel() doesn’t interrupt requests in flight. Use timeout_s and total_timeout_s to bound request duration.

TraceContext

Use TraceContext(traceparent, tracestate=None) to attach requests to a parent trace. Pass it as RequestProcessingPreference(trace_context=...).
str
required
W3C parent trace context.
str | None
Additional W3C trace state. Defaults to None.
The constructed object exposes these read-only properties:
str
The supplied parent trace context.
str | None
The supplied trace state, or None.
Don’t also set traceparent or tracestate in extra_headers; the client rejects that combination. To export client spans to an OpenTelemetry collector, configure client tracing.

Endpoint

Endpoint configures a base URL and the health checks used by an EndpointPool.
str
required
Endpoint base URL.
str
required
Key for health-check authentication. Inference requests use the client’s API key.
HttpClientWrapper
required
HTTP connection pool for health checks.
str | None
Absolute URL for an additional health check. Defaults to None.
str | None
Relative health-check path. Defaults to None, which uses /health.
float | None
Seconds between health checks. Defaults to None, which uses 10 seconds.
float | None
Timeout in seconds per health-check attempt. Defaults to None, which uses 6 seconds.
int | None
Retries per health check. Defaults to None, which uses 2 retries.
bool
default:false
Stop evaluating checks after the first failure.
bool | None
Ignore a timeout from the deployment check when deciding whether the endpoint is healthy. Defaults to None, which uses True.
bool | None
Ignore a timeout from the deep health check when deciding whether the endpoint is healthy. Defaults to None, which uses True.

EndpointPool

EndpointPool(endpoints, endpoint_weights=None) groups endpoints for request routing.
list[Endpoint]
required
Nonempty list of endpoints with distinct base URLs.
list[float] | None
Routing weights for the endpoints. Defaults to None, which gives each endpoint equal weight. Custom weights must match the endpoint count, be finite and nonnegative, and include at least one positive value.
Pass the pool as PerformanceClient(endpoint_pool=pool, ...). An endpoint’s health-check key doesn’t replace the client’s inference API key.

Errors

HTTP errors don’t attach a requests.Response object. Read the exception’s args instead of assuming error.response.status_code is available. Handle errors from the embedding client:

Package source

Check baseten_performance_client.__version__ for the installed version. The PyPI source distribution includes the type hints and Rust implementation. See the Python binding source on GitHub for current development.