Skip to main content
Use baseten_performance_client_core to send concurrent inference requests from a Tokio application. See the Performance Client overview for shared batching, retry, and connection settings.

Installation

First request

Set your API key, an embeddings deployment URL, and the model value expected by that deployment:
Save this as src/main.rs in your Cargo project:
src/main.rs
Run cargo run. A successful response prints 2 embeddings and the elapsed time. For reranking, classification, or generic requests, set the base URL to a deployment that supports the corresponding operation.

PerformanceClientCore

String
required
Base URL for requests. Inference methods append their request path. With an endpoint pool, the first pool endpoint becomes the primary URL.
Option<String>
Explicit key, or None to read BASETEN_API_KEY, then OPENAI_API_KEY. Missing all three returns InvalidParameter.
u8
required
1 for HTTP/1.1 or 2 for HTTP/2.
Option<Arc<HttpClientWrapper>>
Existing connection pool, or None to create one.
Option<String>
Proxy URL for a newly created connection pool. If you supply a wrapper, the client uses the wrapper’s proxy configuration.
Option<Arc<EndpointPool>>
Pool of endpoints for request routing. Pass None to use base_url.
Use HttpClientWrapper::new(http_version, proxy)? to create an Arc<HttpClientWrapper>. Clone the Arc into multiple clients to share connections. client.get_client_wrapper() returns a clone of the client’s wrapper.

Inference methods

All methods are asynchronous and return Result<_, ClientError>. Configure each call with a RequestProcessingPreference.

Generate embeddings

Sends batches to /v1/embeddings. Returns (CoreOpenAIEmbeddingsResponse, Vec<Duration>, Vec<HeaderMap>, Duration).
Vec<String>
required
Texts to embed.
String
required
Model value expected by the server.
Option<String>
Embedding encoding sent to the server. Pass None to let the server select the format.
Option<u32>
Embedding dimensions sent to the server.
Option<String>
User identifier sent to the server.
&RequestProcessingPreference
required
Batching, concurrency, timeout, and retry settings for the call.
CoreOpenAIEmbeddingsResponse combines results from all batches:
Vec<CoreOpenAIEmbeddingData>
Embedding results.
CoreOpenAIUsage
Token counts reported by the server.

CoreOpenAIEmbeddingData

String
Object type returned by the server.
usize
Index of the input text.
CoreEmbeddingVariant
Match as CoreEmbeddingVariant::FloatVector(Vec<f32>) or CoreEmbeddingVariant::Base64(String).

CoreOpenAIUsage

u32
Number of prompt tokens reported by the server.
u32
Total number of tokens reported by the server.

Rerank texts

Sends batches to /rerank. Returns (CoreRerankResponse, Vec<Duration>, Vec<HeaderMap>, Duration).
String
required
Query to compare against the texts.
Vec<String>
required
Texts to rerank.
bool
required
Request raw scores.
Option<String>
Model value sent to the server.
bool
required
Include text in each result.
bool
required
Truncation setting interpreted by the server.
String
required
Truncation direction interpreted by the server, such as "Right".
&RequestProcessingPreference
required
Batching, concurrency, timeout, and retry settings for the call.
CoreRerankResponse combines results from all batches:
Vec<CoreRerankResult>
Reranking results.

CoreRerankResult

usize
Index of the input text.
f64
Reranking score.
Option<String>
Text returned when requested.

Classify texts

Sends batches to /predict, wrapping each input string in a single-element list. Returns (CoreClassificationResponse, Vec<Duration>, Vec<HeaderMap>, Duration).
Vec<String>
required
Texts to classify.
Option<String>
Model value sent to the server.
bool
required
Request raw scores.
bool
required
Truncation setting interpreted by the server.
String
required
Truncation direction interpreted by the server.
&RequestProcessingPreference
required
Batching, concurrency, timeout, and retry settings for the call.
Vec<Vec<CoreClassificationResult>>
One list of CoreClassificationResult values per input text.

CoreClassificationResult

String
Classification label.
f64
Score for the label.

Send generic requests

Sends one request per payload and expects non-streaming responses. Add serde_json = "1" to your dependencies to construct payloads.
String
required
Request path appended to the base URL.
Vec<serde_json::Value>
required
Payloads to send as separate requests.
&RequestProcessingPreference
required
Concurrency, timeout, and retry settings for the call.
HttpMethod
required
POST, GET, PUT, PATCH, DELETE, HEAD, or OPTIONS. Only POST, PUT, and PATCH send request bodies. DELETE, HEAD, and OPTIONS return empty maps instead of decoded response bodies.
Returns (Vec<(rmpv::Value, HeaderMap, Duration)>, Duration):
Vec<(rmpv::Value, HeaderMap, Duration)>
Response body, headers, and duration for each payload in input order.
Duration
Total elapsed time for the operation.

Response timing

Embedding, reranking, and classification methods return a four-element tuple. Tuple positions are zero-based:
CoreOpenAIEmbeddingsResponse | CoreRerankResponse | CoreClassificationResponse
Merged response. The response type and data fields depend on the method. All three response types include the timing and header fields below.
Vec<Duration>
Duration of each batch request.
Vec<HeaderMap>
Headers of each batch response.
Duration
Total elapsed time for the operation.
The merged response includes these fields:
f64
Total elapsed time in seconds.
Vec<f64>
Duration of each batch request in seconds.
Vec<HeaderMap>
Headers of each batch response.
Per-request timing and header entries correspond to batches, so their count can differ from the number of input texts.

Request preferences

RequestProcessingPreference::new() and Default::default() leave every field as None. At request start, the client applies the defaults below to any unset fields. Set values with with_<field>(value) builders or assign Some(value) to public fields.
Option<usize>
Maximum concurrent requests. Effective default: 128. Must be 1 to 1,024, or 1 to 512 when batch_size is below 16.
Option<usize>
Maximum texts per batch, from 1 to 1,024. Effective default: 128.
Option<usize>
Character threshold for splitting batches, from 50 to 256,000. Unset by default. A single text that exceeds the threshold stays intact.
Option<bool>
Send all initial requests in this call to one endpoint from the pool. Effective default: false.
Option<f64>
Per-request timeout, from 0.1 to 3,600 seconds. Effective default: 3600.0.
Option<f64>
Delay in seconds before a duplicate attempt. Unset by default. Must be at least 0.045 seconds and less than timeout_s - 0.045. Enabling hedging can make your model process the same input more than once.
Option<f64>
Deadline in seconds for the whole operation. Unset by default; must be at least timeout_s when set.
Option<f64>
Hedge budget fraction, from 0 to 3. Effective default: 0.10.
Option<f64>
Retry budget fraction, from 0 to 3. Effective default: 0.05.
Retry and hedge budgets don’t reliably cap extra attempts across a call. The per-request max_retries limit still applies. Leave hedge_delay unset to turn off hedging.
Option<u32>
Maximum retries per request, from 0 to 6. Effective default: 5. By default, the client retries HTTP statuses 408, 409, 429, and 500–599. Set to 0 to turn off retries.
Option<u64>
Initial backoff, from 50 to 45,000 milliseconds. Effective default: 125.
Option<CancellationToken>
Shared cancellation flag. The client creates a new token when unset. Calling cancel() alone doesn’t stop requests. See cancellation behavior.
Option<String>
Accepted but doesn’t change inference authentication. Set the key on the client instead.
Option<HashMap<String, String>>
Additional HTTP headers. Unset by default.
Option<HashSet<u16>>
Exclude status codes from automatic retries. Unset by default.

Cancellation and errors

Dropping an inference future aborts its in-flight request tasks. For example, use a timeout around the future or abort the Tokio task that owns it. CancellationToken::new(false) creates a shared flag. token.cancel() sets it, and token.is_cancelled() reads it. Inference requests don’t poll this flag, so calling token.cancel() alone doesn’t stop requests.
bool
required
When true, dropping a token clone also sets the cancellation flag.
Match ClientError to handle failures:

Endpoint pools

Create endpoints inside a Tokio runtime, because Endpoint::new starts background health checks.
  1. Create a shared HttpClientWrapper.
  2. Build EndpointConfig::new(base_url, health_api_key, wrapper) for each endpoint, then call Endpoint::new(config)?.
  3. Pass the endpoints to EndpointPoolConfig::new(endpoints). Optionally set .with_weights(weights).
  4. Call EndpointPool::new(config)? and pass its Arc to the client constructor.

Pool configuration

EndpointPoolConfig accepts these fields:
Vec<Endpoint>
required
At least one endpoint. Endpoint URLs must be unique.
Option<Vec<f64>>
Routing weights, initially None. When set, weights must match the endpoint count, be finite and nonnegative, and include at least one positive value.

Health snapshot

pool.health_snapshot() returns an EndpointPoolHealthSnapshot:
Vec<EndpointHealthStatus>
Health status of each endpoint.

EndpointHealthStatus

String
Endpoint base URL.
bool
Whether the endpoint is healthy.

Endpoint configuration

EndpointConfig provides builders for these fields:
Duration
Interval between health checks. Defaults to 10 seconds. Set with .with_health_check_interval(duration).
Duration
Health check timeout. Defaults to 6 seconds. Set with .with_health_check_timeout(duration).
u32
default:2
Number of health check retries. Set with .with_health_check_retries(retries).
usize
default:64
Retry attempt concurrency limit. Minimum: 4. Set with .with_retry_attempt_concurrency_limit(limit).
Option<EndpointHealthConfig>
Custom health checks, initially None. When unset, the endpoint uses a relative /health check. Set with .with_endpoint_health(config).
Use .with_standard_health_checks(deep_health_url, fail_on_first, deployment_health_path, deployment_timeout_is_no_vote, deep_timeout_is_no_vote) to configure deployment and optional deep health checks together. For custom checks, create an EndpointHealthConfig with these fields:
Vec<EndpointHealthCheckConfig>
required
Health checks created with EndpointHealthCheckConfig::relative(path) or ::absolute(url).
bool
required
Mark the endpoint unhealthy as soon as a health check fails.
EndpointHealthCheckConfig supports this option:
bool
default:true
Ignore a timeout from this check when deciding whether the endpoint is healthy. Set with .with_timeout_is_no_vote(value).

Environment and crate features

String
Default API key when the client constructor receives None.
String
Fallback key if BASETEN_API_KEY is absent.
String
default:"warn"
Log filter. Takes precedence over RUST_LOG.
String
default:"perfclient"
Request ID prefix.
The crate enables the rustls TLS feature by default. native-tls is also available. The crate initializes a tracing subscriber automatically.

Package source

Find releases on crates.io and the Rust source on GitHub for current development. For other types, utility functions, and constants, see the crate API reference.