baseten_performance_client_core to send concurrent inference requests from a Tokio application.
See the Performance Client overview for shared batching, retry, and connection settings.
Installation
First request
Set your API key, an embeddings deployment URL, and themodel value expected by that deployment:
src/main.rs in your Cargo project:
src/main.rs
cargo run. A successful response prints 2 embeddings and the elapsed time. For reranking, classification, or generic requests, set the base URL to a deployment that supports the corresponding operation.
PerformanceClientCore
String
required
Base URL for requests. Inference methods append their request path. With an endpoint pool, the first pool endpoint becomes the primary URL.
Option<String>
Explicit key, or
None to read BASETEN_API_KEY, then OPENAI_API_KEY. Missing all three returns InvalidParameter.u8
required
1 for HTTP/1.1 or 2 for HTTP/2.Option<Arc<HttpClientWrapper>>
Existing connection pool, or
None to create one.Option<String>
Proxy URL for a newly created connection pool. If you supply a wrapper, the client uses the wrapper’s proxy configuration.
Option<Arc<EndpointPool>>
Pool of endpoints for request routing. Pass
None to use base_url.HttpClientWrapper::new(http_version, proxy)? to create an Arc<HttpClientWrapper>. Clone the Arc into multiple clients to share connections. client.get_client_wrapper() returns a clone of the client’s wrapper.
Inference methods
All methods are asynchronous and returnResult<_, ClientError>. Configure each call with a RequestProcessingPreference.
Generate embeddings
/v1/embeddings. Returns (CoreOpenAIEmbeddingsResponse, Vec<Duration>, Vec<HeaderMap>, Duration).
Vec<String>
required
Texts to embed.
String
required
Model value expected by the server.
Option<String>
Embedding encoding sent to the server. Pass
None to let the server select the format.Option<u32>
Embedding dimensions sent to the server.
Option<String>
User identifier sent to the server.
&RequestProcessingPreference
required
Batching, concurrency, timeout, and retry settings for the call.
CoreOpenAIEmbeddingsResponse combines results from all batches:
Vec<CoreOpenAIEmbeddingData>
Embedding results.
CoreOpenAIUsage
Token counts reported by the server.
CoreOpenAIEmbeddingData
String
Object type returned by the server.
usize
Index of the input text.
CoreEmbeddingVariant
Match as
CoreEmbeddingVariant::FloatVector(Vec<f32>) or CoreEmbeddingVariant::Base64(String).CoreOpenAIUsage
u32
Number of prompt tokens reported by the server.
u32
Total number of tokens reported by the server.
Rerank texts
/rerank. Returns (CoreRerankResponse, Vec<Duration>, Vec<HeaderMap>, Duration).
String
required
Query to compare against the texts.
Vec<String>
required
Texts to rerank.
bool
required
Request raw scores.
Option<String>
Model value sent to the server.
bool
required
Include text in each result.
bool
required
Truncation setting interpreted by the server.
String
required
Truncation direction interpreted by the server, such as
"Right".&RequestProcessingPreference
required
Batching, concurrency, timeout, and retry settings for the call.
CoreRerankResponse combines results from all batches:
Vec<CoreRerankResult>
Reranking results.
CoreRerankResult
usize
Index of the input text.
f64
Reranking score.
Option<String>
Text returned when requested.
Classify texts
/predict, wrapping each input string in a single-element list. Returns (CoreClassificationResponse, Vec<Duration>, Vec<HeaderMap>, Duration).
Vec<String>
required
Texts to classify.
Option<String>
Model value sent to the server.
bool
required
Request raw scores.
bool
required
Truncation setting interpreted by the server.
String
required
Truncation direction interpreted by the server.
&RequestProcessingPreference
required
Batching, concurrency, timeout, and retry settings for the call.
Vec<Vec<CoreClassificationResult>>
One list of
CoreClassificationResult values per input text.CoreClassificationResult
String
Classification label.
f64
Score for the label.
Send generic requests
serde_json = "1" to your dependencies to construct payloads.
String
required
Request path appended to the base URL.
Vec<serde_json::Value>
required
Payloads to send as separate requests.
&RequestProcessingPreference
required
Concurrency, timeout, and retry settings for the call.
HttpMethod
required
POST, GET, PUT, PATCH, DELETE, HEAD, or OPTIONS. Only POST, PUT, and PATCH send request bodies. DELETE, HEAD, and OPTIONS return empty maps instead of decoded response bodies.(Vec<(rmpv::Value, HeaderMap, Duration)>, Duration):
Vec<(rmpv::Value, HeaderMap, Duration)>
Response body, headers, and duration for each payload in input order.
Duration
Total elapsed time for the operation.
Response timing
Embedding, reranking, and classification methods return a four-element tuple. Tuple positions are zero-based:CoreOpenAIEmbeddingsResponse | CoreRerankResponse | CoreClassificationResponse
Merged response. The response type and
data fields depend on the method. All three response types include the timing and header fields below.Vec<Duration>
Duration of each batch request.
Vec<HeaderMap>
Headers of each batch response.
Duration
Total elapsed time for the operation.
f64
Total elapsed time in seconds.
Vec<f64>
Duration of each batch request in seconds.
Vec<HeaderMap>
Headers of each batch response.
Request preferences
RequestProcessingPreference::new() and Default::default() leave every field as None. At request start, the client applies the defaults below to any unset fields. Set values with with_<field>(value) builders or assign Some(value) to public fields.
Option<usize>
Maximum concurrent requests. Effective default:
128. Must be 1 to 1,024, or 1 to 512 when batch_size is below 16.Option<usize>
Maximum texts per batch, from 1 to 1,024. Effective default:
128.Option<usize>
Character threshold for splitting batches, from 50 to 256,000. Unset by default. A single text that exceeds the threshold stays intact.
Option<bool>
Send all initial requests in this call to one endpoint from the pool. Effective default:
false.Option<f64>
Per-request timeout, from 0.1 to 3,600 seconds. Effective default:
3600.0.Option<f64>
Delay in seconds before a duplicate attempt. Unset by default. Must be at least 0.045 seconds and less than
timeout_s - 0.045. Enabling hedging can make your model process the same input more than once.Option<f64>
Deadline in seconds for the whole operation. Unset by default; must be at least
timeout_s when set.Option<f64>
Hedge budget fraction, from 0 to 3. Effective default:
0.10.Option<f64>
Retry budget fraction, from 0 to 3. Effective default:
0.05.max_retries limit still applies. Leave hedge_delay unset to turn off hedging.
Option<u32>
Maximum retries per request, from 0 to 6. Effective default:
5. By default, the client retries HTTP statuses 408, 409, 429, and 500–599. Set to 0 to turn off retries.Option<u64>
Initial backoff, from 50 to 45,000 milliseconds. Effective default:
125.Option<CancellationToken>
Shared cancellation flag. The client creates a new token when unset. Calling
cancel() alone doesn’t stop requests. See cancellation behavior.Option<String>
Accepted but doesn’t change inference authentication. Set the key on the client instead.
Option<HashMap<String, String>>
Additional HTTP headers. Unset by default.
Option<HashSet<u16>>
Exclude status codes from automatic retries. Unset by default.
Cancellation and errors
Dropping an inference future aborts its in-flight request tasks. For example, use a timeout around the future or abort the Tokio task that owns it.CancellationToken::new(false) creates a shared flag. token.cancel() sets it, and token.is_cancelled() reads it. Inference requests don’t poll this flag, so calling token.cancel() alone doesn’t stop requests.
bool
required
When
true, dropping a token clone also sets the cancellation flag.ClientError to handle failures:
Endpoint pools
Create endpoints inside a Tokio runtime, becauseEndpoint::new starts background health checks.
- Create a shared
HttpClientWrapper. - Build
EndpointConfig::new(base_url, health_api_key, wrapper)for each endpoint, then callEndpoint::new(config)?. - Pass the endpoints to
EndpointPoolConfig::new(endpoints). Optionally set.with_weights(weights). - Call
EndpointPool::new(config)?and pass itsArcto the client constructor.
Pool configuration
EndpointPoolConfig accepts these fields:
Vec<Endpoint>
required
At least one endpoint. Endpoint URLs must be unique.
Option<Vec<f64>>
Routing weights, initially
None. When set, weights must match the endpoint count, be finite and nonnegative, and include at least one positive value.Health snapshot
pool.health_snapshot() returns an EndpointPoolHealthSnapshot:
Vec<EndpointHealthStatus>
Health status of each endpoint.
EndpointHealthStatus
String
Endpoint base URL.
bool
Whether the endpoint is healthy.
Endpoint configuration
EndpointConfig provides builders for these fields:
Duration
Interval between health checks. Defaults to 10 seconds. Set with
.with_health_check_interval(duration).Duration
Health check timeout. Defaults to 6 seconds. Set with
.with_health_check_timeout(duration).u32
default:2
Number of health check retries. Set with
.with_health_check_retries(retries).usize
default:64
Retry attempt concurrency limit. Minimum:
4. Set with .with_retry_attempt_concurrency_limit(limit).Option<EndpointHealthConfig>
Custom health checks, initially
None. When unset, the endpoint uses a relative /health check. Set with .with_endpoint_health(config)..with_standard_health_checks(deep_health_url, fail_on_first, deployment_health_path, deployment_timeout_is_no_vote, deep_timeout_is_no_vote) to configure deployment and optional deep health checks together.
For custom checks, create an EndpointHealthConfig with these fields:
Vec<EndpointHealthCheckConfig>
required
Health checks created with
EndpointHealthCheckConfig::relative(path) or ::absolute(url).bool
required
Mark the endpoint unhealthy as soon as a health check fails.
EndpointHealthCheckConfig supports this option:
bool
default:true
Ignore a timeout from this check when deciding whether the endpoint is healthy. Set with
.with_timeout_is_no_vote(value).Environment and crate features
String
Default API key when the client constructor receives
None.String
Fallback key if
BASETEN_API_KEY is absent.String
default:"warn"
Log filter. Takes precedence over
RUST_LOG.String
default:"perfclient"
Request ID prefix.
rustls TLS feature by default. native-tls is also available. The crate initializes a tracing subscriber automatically.