> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Rust performance client

> Send concurrent inference requests from Rust with the Baseten Performance Client.

Use `baseten_performance_client_core` to send concurrent inference requests from a Tokio application.

See the [Performance Client overview](/reference/sdk/performance-client/overview) for shared batching, retry, and connection settings.

## Installation

```sh theme={"system"}
cargo add baseten_performance_client_core
cargo add tokio --features macros,rt-multi-thread
```

## First request

Set your API key, an embeddings deployment URL, and the `model` value expected by that deployment:

```bash theme={"system"}
export BASETEN_API_KEY="<YOUR_API_KEY>"
export BASETEN_BASE_URL="https://model-YOUR_MODEL_ID.api.baseten.co/environments/production/sync"
export BASETEN_MODEL="<MODEL_NAME>"
```

Save this as `src/main.rs` in your Cargo project:

```rust src/main.rs theme={"system"}
use baseten_performance_client_core::{PerformanceClientCore, RequestProcessingPreference};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let client = PerformanceClientCore::new(
        std::env::var("BASETEN_BASE_URL")?,
        Some(std::env::var("BASETEN_API_KEY")?),
        1,
        None,
        None,
        None,
    )?;
    let preference = RequestProcessingPreference::new()
        .with_batch_size(16)
        .with_max_concurrent_requests(32);

    let (response, _, _, elapsed) = client.process_embeddings_requests(
        vec!["Hello world".into(), "Example text".into()],
        std::env::var("BASETEN_MODEL")?,
        Some("float".into()),
        None,
        None,
        &preference,
    ).await?;

    println!("{} embeddings in {:.3}s", response.data.len(), elapsed.as_secs_f64());
    Ok(())
}
```

Run `cargo run`. A successful response prints `2 embeddings` and the elapsed time. For reranking, classification, or generic requests, set the base URL to a deployment that supports the corresponding operation.

## `PerformanceClientCore`

```rust theme={"system"}
PerformanceClientCore::new(
    base_url: String,
    api_key: Option<String>,
    http_version: u8,
    client_wrapper: Option<Arc<HttpClientWrapper>>,
    proxy: Option<String>,
    endpoint_pool: Option<Arc<EndpointPool>>,
) -> Result<PerformanceClientCore, ClientError>
```

<ParamField body="base_url" type="String" required>
  Base URL for requests. Inference methods append their request path. With an endpoint pool, the first pool endpoint becomes the primary URL.
</ParamField>

<ParamField body="api_key" type="Option<String>">
  Explicit key, or `None` to read `BASETEN_API_KEY`, then `OPENAI_API_KEY`. Missing all three returns `InvalidParameter`.
</ParamField>

<ParamField body="http_version" type="u8" required>
  `1` for HTTP/1.1 or `2` for HTTP/2.
</ParamField>

<ParamField body="client_wrapper" type="Option<Arc<HttpClientWrapper>>">
  Existing connection pool, or `None` to create one.
</ParamField>

<ParamField body="proxy" type="Option<String>">
  Proxy URL for a newly created connection pool. If you supply a wrapper, the client uses the wrapper's proxy configuration.
</ParamField>

<ParamField body="endpoint_pool" type="Option<Arc<EndpointPool>>">
  Pool of endpoints for request routing. Pass `None` to use `base_url`.
</ParamField>

Use `HttpClientWrapper::new(http_version, proxy)?` to create an `Arc<HttpClientWrapper>`. Clone the `Arc` into multiple clients to share connections. `client.get_client_wrapper()` returns a clone of the client's wrapper.

## Inference methods

All methods are asynchronous and return `Result<_, ClientError>`. Configure each call with a [`RequestProcessingPreference`](#request-preferences).

### Generate embeddings

```rust theme={"system"}
client.process_embeddings_requests(
    texts: Vec<String>,
    model: String,
    encoding_format: Option<String>,
    dimensions: Option<u32>,
    user: Option<String>,
    preference: &RequestProcessingPreference,
).await
```

Sends batches to `/v1/embeddings`. Returns `(CoreOpenAIEmbeddingsResponse, Vec<Duration>, Vec<HeaderMap>, Duration)`.

<ParamField body="texts" type="Vec<String>" required>
  Texts to embed.
</ParamField>

<ParamField body="model" type="String" required>
  Model value expected by the server.
</ParamField>

<ParamField body="encoding_format" type="Option<String>">
  Embedding encoding sent to the server. Pass `None` to let the server select the format.
</ParamField>

<ParamField body="dimensions" type="Option<u32>">
  Embedding dimensions sent to the server.
</ParamField>

<ParamField body="user" type="Option<String>">
  User identifier sent to the server.
</ParamField>

<ParamField body="preference" type="&RequestProcessingPreference" required>
  Batching, concurrency, timeout, and retry settings for the call.
</ParamField>

`CoreOpenAIEmbeddingsResponse` combines results from all batches:

<ResponseField name="data" type="Vec<CoreOpenAIEmbeddingData>">
  Embedding results.
</ResponseField>

<ResponseField name="usage" type="CoreOpenAIUsage">
  Token counts reported by the server.
</ResponseField>

#### `CoreOpenAIEmbeddingData`

<ResponseField name="object" type="String">
  Object type returned by the server.
</ResponseField>

<ResponseField name="index" type="usize">
  Index of the input text.
</ResponseField>

<ResponseField name="embedding_internal" type="CoreEmbeddingVariant">
  Match as `CoreEmbeddingVariant::FloatVector(Vec<f32>)` or `CoreEmbeddingVariant::Base64(String)`.
</ResponseField>

#### `CoreOpenAIUsage`

<ResponseField name="prompt_tokens" type="u32">
  Number of prompt tokens reported by the server.
</ResponseField>

<ResponseField name="total_tokens" type="u32">
  Total number of tokens reported by the server.
</ResponseField>

### Rerank texts

```rust theme={"system"}
client.process_rerank_requests(
    query: String,
    texts: Vec<String>,
    raw_scores: bool,
    model: Option<String>,
    return_text: bool,
    truncate: bool,
    truncation_direction: String,
    preference: &RequestProcessingPreference,
).await
```

Sends batches to `/rerank`. Returns `(CoreRerankResponse, Vec<Duration>, Vec<HeaderMap>, Duration)`.

<ParamField body="query" type="String" required>
  Query to compare against the texts.
</ParamField>

<ParamField body="texts" type="Vec<String>" required>
  Texts to rerank.
</ParamField>

<ParamField body="raw_scores" type="bool" required>
  Request raw scores.
</ParamField>

<ParamField body="model" type="Option<String>">
  Model value sent to the server.
</ParamField>

<ParamField body="return_text" type="bool" required>
  Include text in each result.
</ParamField>

<ParamField body="truncate" type="bool" required>
  Truncation setting interpreted by the server.
</ParamField>

<ParamField body="truncation_direction" type="String" required>
  Truncation direction interpreted by the server, such as `"Right"`.
</ParamField>

<ParamField body="preference" type="&RequestProcessingPreference" required>
  Batching, concurrency, timeout, and retry settings for the call.
</ParamField>

`CoreRerankResponse` combines results from all batches:

<ResponseField name="data" type="Vec<CoreRerankResult>">
  Reranking results.
</ResponseField>

#### `CoreRerankResult`

<ResponseField name="index" type="usize">
  Index of the input text.
</ResponseField>

<ResponseField name="score" type="f64">
  Reranking score.
</ResponseField>

<ResponseField name="text" type="Option<String>">
  Text returned when requested.
</ResponseField>

### Classify texts

```rust theme={"system"}
client.process_classify_requests(
    inputs: Vec<String>,
    model: Option<String>,
    raw_scores: bool,
    truncate: bool,
    truncation_direction: String,
    preference: &RequestProcessingPreference,
).await
```

Sends batches to `/predict`, wrapping each input string in a single-element list. Returns `(CoreClassificationResponse, Vec<Duration>, Vec<HeaderMap>, Duration)`.

<ParamField body="inputs" type="Vec<String>" required>
  Texts to classify.
</ParamField>

<ParamField body="model" type="Option<String>">
  Model value sent to the server.
</ParamField>

<ParamField body="raw_scores" type="bool" required>
  Request raw scores.
</ParamField>

<ParamField body="truncate" type="bool" required>
  Truncation setting interpreted by the server.
</ParamField>

<ParamField body="truncation_direction" type="String" required>
  Truncation direction interpreted by the server.
</ParamField>

<ParamField body="preference" type="&RequestProcessingPreference" required>
  Batching, concurrency, timeout, and retry settings for the call.
</ParamField>

<ResponseField name="data" type="Vec<Vec<CoreClassificationResult>>">
  One list of `CoreClassificationResult` values per input text.
</ResponseField>

#### `CoreClassificationResult`

<ResponseField name="label" type="String">
  Classification label.
</ResponseField>

<ResponseField name="score" type="f64">
  Score for the label.
</ResponseField>

### Send generic requests

```rust theme={"system"}
client.process_batch_post_requests(
    url_path: String,
    payloads_json: Vec<serde_json::Value>,
    preference: &RequestProcessingPreference,
    method: HttpMethod,
).await
```

Sends one request per payload and expects non-streaming responses. Add `serde_json = "1"` to your dependencies to construct payloads.

<ParamField body="url_path" type="String" required>
  Request path appended to the base URL.
</ParamField>

<ParamField body="payloads_json" type="Vec<serde_json::Value>" required>
  Payloads to send as separate requests.
</ParamField>

<ParamField body="preference" type="&RequestProcessingPreference" required>
  Concurrency, timeout, and retry settings for the call.
</ParamField>

<ParamField body="method" type="HttpMethod" required>
  `POST`, `GET`, `PUT`, `PATCH`, `DELETE`, `HEAD`, or `OPTIONS`. Only `POST`, `PUT`, and `PATCH` send request bodies. `DELETE`, `HEAD`, and `OPTIONS` return empty maps instead of decoded response bodies.
</ParamField>

Returns `(Vec<(rmpv::Value, HeaderMap, Duration)>, Duration)`:

<ResponseField name="0" type="Vec<(rmpv::Value, HeaderMap, Duration)>">
  Response body, headers, and duration for each payload in input order.
</ResponseField>

<ResponseField name="1" type="Duration">
  Total elapsed time for the operation.
</ResponseField>

### Response timing

Embedding, reranking, and classification methods return a four-element tuple. Tuple positions are zero-based:

<ResponseField name="0" type="CoreOpenAIEmbeddingsResponse | CoreRerankResponse | CoreClassificationResponse">
  Merged response. The response type and `data` fields depend on the method. All three response types include the timing and header fields below.
</ResponseField>

<ResponseField name="1" type="Vec<Duration>">
  Duration of each batch request.
</ResponseField>

<ResponseField name="2" type="Vec<HeaderMap>">
  Headers of each batch response.
</ResponseField>

<ResponseField name="3" type="Duration">
  Total elapsed time for the operation.
</ResponseField>

The merged response includes these fields:

<ResponseField name="total_time" type="f64">
  Total elapsed time in seconds.
</ResponseField>

<ResponseField name="individual_request_times" type="Vec<f64>">
  Duration of each batch request in seconds.
</ResponseField>

<ResponseField name="response_headers" type="Vec<HeaderMap>">
  Headers of each batch response.
</ResponseField>

Per-request timing and header entries correspond to batches, so their count can differ from the number of input texts.

## Request preferences

`RequestProcessingPreference::new()` and `Default::default()` leave every field as `None`. At request start, the client applies the defaults below to any unset fields. Set values with `with_<field>(value)` builders or assign `Some(value)` to public fields.

<ParamField body="max_concurrent_requests" type="Option<usize>">
  Maximum concurrent requests. Effective default: `128`. Must be 1 to 1,024, or 1 to 512 when `batch_size` is below 16.
</ParamField>

<ParamField body="batch_size" type="Option<usize>">
  Maximum texts per batch, from 1 to 1,024. Effective default: `128`.
</ParamField>

<ParamField body="max_chars_per_request" type="Option<usize>">
  Character threshold for splitting batches, from 50 to 256,000. Unset by default. A single text that exceeds the threshold stays intact.
</ParamField>

<ParamField body="pin_initial_endpoint_once" type="Option<bool>">
  Send all initial requests in this call to one endpoint from the pool. Effective default: `false`.
</ParamField>

<ParamField body="timeout_s" type="Option<f64>">
  Per-request timeout, from 0.1 to 3,600 seconds. Effective default: `3600.0`.
</ParamField>

<ParamField body="hedge_delay" type="Option<f64>">
  Delay in seconds before a duplicate attempt. Unset by default. Must be at least 0.045 seconds and less than `timeout_s - 0.045`. Enabling hedging can make your model process the same input more than once.
</ParamField>

<ParamField body="total_timeout_s" type="Option<f64>">
  Deadline in seconds for the whole operation. Unset by default; must be at least `timeout_s` when set.
</ParamField>

<ParamField body="hedge_budget_pct" type="Option<f64>">
  Hedge budget fraction, from 0 to 3. Effective default: `0.10`.
</ParamField>

<ParamField body="retry_budget_pct" type="Option<f64>">
  Retry budget fraction, from 0 to 3. Effective default: `0.05`.
</ParamField>

Retry and hedge budgets don't reliably cap extra attempts across a call. The per-request `max_retries` limit still applies. Leave `hedge_delay` unset to turn off hedging.

<ParamField body="max_retries" type="Option<u32>">
  Maximum retries per request, from 0 to 6. Effective default: `5`. By default, the client retries HTTP statuses `408`, `409`, `429`, and `500`–`599`. Set to `0` to turn off retries.
</ParamField>

<ParamField body="initial_backoff_ms" type="Option<u64>">
  Initial backoff, from 50 to 45,000 milliseconds. Effective default: `125`.
</ParamField>

<ParamField body="cancel_token" type="Option<CancellationToken>">
  Shared cancellation flag. The client creates a new token when unset. Calling `cancel()` alone doesn't stop requests. See [cancellation behavior](#cancellation-and-errors).
</ParamField>

<ParamField body="primary_api_key_override" type="Option<String>">
  Accepted but doesn't change inference authentication. Set the key on the client instead.
</ParamField>

<ParamField body="extra_headers" type="Option<HashMap<String, String>>">
  Additional HTTP headers. Unset by default.
</ParamField>

<ParamField body="non_retryable_status_codes" type="Option<HashSet<u16>>">
  Exclude status codes from automatic retries. Unset by default.
</ParamField>

## Cancellation and errors

Dropping an inference future aborts its in-flight request tasks. For example, use a timeout around the future or abort the Tokio task that owns it.

`CancellationToken::new(false)` creates a shared flag. `token.cancel()` sets it, and `token.is_cancelled()` reads it. Inference requests don't poll this flag, so calling `token.cancel()` alone doesn't stop requests.

<ParamField body="cancel_on_drop" type="bool" required>
  When `true`, dropping a token clone also sets the cancellation flag.
</ParamField>

Match `ClientError` to handle failures:

| Variant | Fields | Meaning |
| - | - | - |
| `InvalidParameter` | `String` | Missing key or invalid request configuration. |
| `Http` | `status: u16`, `message: String`, `customer_request_id: Option<String>` | HTTP failure after applicable retries. |
| `LocalTimeout` | `String`, `Option<String>` | Client timeout, with an optional request ID. |
| `RemoteTimeout` | `String`, `Option<String>` | Server timeout, with an optional request ID. |
| `Network` | `String` | Network failure. |
| `Connect` | `String` | Connection failure. |
| `Serialization` | `String` | Request or response conversion failure. |
| `Cancellation` | `String` | Operation cancelled. |

## Endpoint pools

Create endpoints inside a Tokio runtime, because `Endpoint::new` starts background health checks.

1. Create a shared `HttpClientWrapper`.
2. Build `EndpointConfig::new(base_url, health_api_key, wrapper)` for each endpoint, then call `Endpoint::new(config)?`.
3. Pass the endpoints to `EndpointPoolConfig::new(endpoints)`. Optionally set `.with_weights(weights)`.
4. Call `EndpointPool::new(config)?` and pass its `Arc` to the client constructor.

### Pool configuration

`EndpointPoolConfig` accepts these fields:

<ParamField body="endpoints" type="Vec<Endpoint>" required>
  At least one endpoint. Endpoint URLs must be unique.
</ParamField>

<ParamField body="weights" type="Option<Vec<f64>>">
  Routing weights, initially `None`. When set, weights must match the endpoint count, be finite and nonnegative, and include at least one positive value.
</ParamField>

### Health snapshot

`pool.health_snapshot()` returns an `EndpointPoolHealthSnapshot`:

<ResponseField name="endpoints" type="Vec<EndpointHealthStatus>">
  Health status of each endpoint.
</ResponseField>

#### `EndpointHealthStatus`

<ResponseField name="base_url" type="String">
  Endpoint base URL.
</ResponseField>

<ResponseField name="healthy" type="bool">
  Whether the endpoint is healthy.
</ResponseField>

### Endpoint configuration

`EndpointConfig` provides builders for these fields:

<ParamField body="health_check_interval" type="Duration">
  Interval between health checks. Defaults to 10 seconds. Set with `.with_health_check_interval(duration)`.
</ParamField>

<ParamField body="health_check_timeout" type="Duration">
  Health check timeout. Defaults to 6 seconds. Set with `.with_health_check_timeout(duration)`.
</ParamField>

<ParamField body="health_check_retries" type="u32" default={2}>
  Number of health check retries. Set with `.with_health_check_retries(retries)`.
</ParamField>

<ParamField body="retry_attempt_concurrency_limit" type="usize" default={64}>
  Retry attempt concurrency limit. Minimum: `4`. Set with `.with_retry_attempt_concurrency_limit(limit)`.
</ParamField>

<ParamField body="endpoint_health" type="Option<EndpointHealthConfig>">
  Custom health checks, initially `None`. When unset, the endpoint uses a relative `/health` check. Set with `.with_endpoint_health(config)`.
</ParamField>

Use `.with_standard_health_checks(deep_health_url, fail_on_first, deployment_health_path, deployment_timeout_is_no_vote, deep_timeout_is_no_vote)` to configure deployment and optional deep health checks together.

For custom checks, create an `EndpointHealthConfig` with these fields:

<ParamField body="checks" type="Vec<EndpointHealthCheckConfig>" required>
  Health checks created with `EndpointHealthCheckConfig::relative(path)` or `::absolute(url)`.
</ParamField>

<ParamField body="fail_on_first" type="bool" required>
  Mark the endpoint unhealthy as soon as a health check fails.
</ParamField>

`EndpointHealthCheckConfig` supports this option:

<ParamField body="timeout_is_no_vote" type="bool" default={true}>
  Ignore a timeout from this check when deciding whether the endpoint is healthy. Set with `.with_timeout_is_no_vote(value)`.
</ParamField>

## Environment and crate features

<ParamField body="BASETEN_API_KEY" type="String">
  Default API key when the client constructor receives `None`.
</ParamField>

<ParamField body="OPENAI_API_KEY" type="String">
  Fallback key if `BASETEN_API_KEY` is absent.
</ParamField>

<ParamField body="PERFORMANCE_CLIENT_LOG_LEVEL" type="String" default="warn">
  Log filter. Takes precedence over `RUST_LOG`.
</ParamField>

<ParamField body="PERFORMANCE_CLIENT_REQUEST_ID_PREFIX" type="String" default="perfclient">
  Request ID prefix.
</ParamField>

The crate enables the `rustls` TLS feature by default. `native-tls` is also available. The crate initializes a tracing subscriber automatically.

## Package source

Find releases on [crates.io](https://crates.io/crates/baseten_performance_client_core) and the [Rust source](https://github.com/basetenlabs/truss/tree/main/baseten-performance-client/core) on GitHub for current development.

For other types, utility functions, and constants, see the [crate API reference](https://docs.rs/baseten_performance_client_core/latest/baseten_performance_client_core/).
