/v1/embeddings. To create and manage Baseten resources, use the management SDKs.
Libraries
Python
Node.js
Rust
Inference methods
Use the embedding, reranking, and classification methods with compatible Baseten Embeddings Inference deployments. For other JSON endpoints, such as Engine-Builder-LLM completions, use the generic request methods. They expect non-streaming responses. Set
stream to false when the endpoint supports streaming.
Batching and concurrency
Embedding, reranking, and classification methods split text inputs into batches based on the text count and character threshold. A single text that exceeds the character threshold stays intact. Generic request methods send each payload as a separate request. Batch size and character thresholds don’t combine those payloads. The client returns generic responses in input order.- Python
- Node.js
- Rust
Configure
RequestProcessingPreference with keyword arguments.int
default:256
Maximum concurrent requests. Must be 1 to 1,024, or 1 to 512 when
batch_size is below 16.int
default:8
Maximum texts per embedding, reranking, or classification request. Accepts 1 to 1,024.
Timeouts and retries
Timeouts
- Python
- Node.js
- Rust
float
default:3600
Timeout for each HTTP request, from 0.1 to 3,600 seconds.
float | None
Timeout for the entire method call, in seconds. Defaults to
None. Must be at least timeout_s when set.CancellationToken for cancellation behavior and limitations.Request headers
The client sets these headers from the per-request timeout. Servers that support them can stop work when a request expires.string
Relative timeout in milliseconds, rounded up. A timeout of
30.5 seconds produces 30500.string
Absolute deadline as a Unix timestamp in milliseconds.
Retries
By default, the client retries HTTP statuses408, 409, 429, and 500 through 599. HTTP status retries don’t consume the retry budget for network and timeout failures.
- Python
- Node.js
- Rust
Hedging
Hedging starts a duplicate request while waiting for the original request to finish. Python and Node.js enforce a shared hedge budget for these extra attempts. Enabling hedging can make your model process the same input more than once.- Python
- Node.js
- Rust
float | None
Delay before starting a duplicate request, in seconds. Defaults to
None, which disables hedging. When set, must be at least 0.045 seconds and less than timeout_s - 0.045.Connections and response formats
Keep the same client for successive calls so it can reuse HTTP connections. To share connections between clients, pass the sameHttpClientWrapper to their constructors.
Use HTTP/1.1 for high concurrency workloads. See HTTP client configuration for connection and timeout tuning.
The client requests MessagePack responses and zstd compression through its default Accept and Accept-Encoding headers, and decodes MessagePack responses. Set either header explicitly to override its default.
- Python
- Node.js
- Rust
Client tracing
Python and Node.js can export client spans to an OpenTelemetry collector over OTLP/HTTP JSON with gzip compression. Set these environment variables before the first inference call. The client reads them once per process and ignores the application’sOTEL_* variables.
- Python
- Node.js
string
HTTP or HTTPS collector URL, such as
http://localhost:4318. The client appends /v1/traces if the URL path doesn’t end with it. Unset by default. An unset or empty value disables span export.string
Comma-separated
name=value headers for collector requests, such as api-key=<COLLECTOR_KEY>,tenant=<TENANT>. Values are literal and aren’t URL-decoded. Unset by default, which adds no custom headers.TraceContext to attach client spans to a parent trace.Environment variables
string
Default API key when you don’t pass one to the client.
string
Fallback API key when
BASETEN_API_KEY is absent.string
default:"perfclient"
Prefix for request IDs.
- Python
- Node.js
- Rust
string
default:"warn"
Log level. Takes precedence over
RUST_LOG. Accepts trace, debug, info, warn, or error.Benchmarks
Baseten benchmarks measured more than 1,200 requests per second from one client.