Skip to main content
Inference methods return promises for use with await. The package includes TypeScript declarations. See the Performance Client overview for shared batching, retry, and connection settings.

Installation

First request

Set your API key, an embeddings deployment URL, and the model value expected by that deployment:
Save this as embeddings.mjs. The .mjs extension enables ES modules and top-level await.
embeddings.mjs
Run node embeddings.mjs. A successful response prints 2 embeddings, followed by token usage and elapsed time. The operation examples below use the same imports and a client configured for the corresponding deployment.

PerformanceClient

string
required
Base URL for requests. The client appends the request path. For an embeddings deployment, use the URL ending in /sync, without /v1/embeddings.
string | null
Inference API key. When omitted, the client checks BASETEN_API_KEY, then OPENAI_API_KEY. Construction fails if no key is available.
number | null
default:1
1 for HTTP/1.1 or 2 for HTTP/2.
HttpClientWrapper | null
Shared HTTP connection pool. When supplied, its HTTP version and proxy configuration take precedence.
string | null
Proxy URL for a newly created HTTP client.
EndpointPool | null
Set of endpoints the client can send requests to. Version 0.1.15 doesn’t export Endpoint or EndpointPool, so you can’t create a pool through the package import.

getClientWrapper

client.getClientWrapper() returns the HttpClientWrapper used by the client. Pass it to another client to share connections.

Inference methods

Optional arguments are positional. Pass undefined or null to skip an argument. There are no separate synchronous or asyncEmbed methods.

embed

Generates embeddings for the input texts.
string[]
required
Nonempty list of strings to embed.
string
required
Model value expected by the server.
string | null
Use "float". In version 0.1.15, base64 embeddings become empty arrays in the returned response. If omitted, the server selects the format.
number | null
Embedding size supported by the model.
string | null
User identifier sent to the endpoint.
RequestProcessingPreference | null
Batching, concurrency, timeout, and retry settings for this call. See Request preferences.
Returns an embedding response with results, usage, and timing fields.

rerank

Scores texts against a query.
string
required
Query to score the texts against.
string[]
required
Nonempty list of texts to rerank.
boolean | null
default:false
Request raw scores from the server.
string | null
Model value sent to the server.
boolean | null
default:false
Include the input text in each result.
boolean | null
default:false
Let the server truncate inputs.
string | null
default:"Right"
Direction for the server to truncate inputs.
RequestProcessingPreference | null
Batching, concurrency, timeout, and retry settings for this call. See Request preferences.
The model server determines support for model, rawScores, returnText, truncate, and truncationDirection. Returns a reranking response. Configure client for a reranking deployment, then score the texts:

classify

Assigns labels and scores to the input texts.
string[]
required
Nonempty list of strings. The client sends each string as a one-element list in the request’s inputs field.
string | null
Model value sent to the server.
boolean | null
default:false
Request raw scores from the server.
boolean | null
default:false
Let the server truncate inputs.
string | null
default:"Right"
Direction for the server to truncate inputs.
RequestProcessingPreference | null
Batching, concurrency, timeout, and retry settings for this call. See Request preferences.
Returns a classification response. Configure client for a classification deployment, then classify the texts:

batchPost

Sends one HTTP request per payload and returns responses in input order.
string
required
Path appended to the base URL, such as /predict.
JsonValue[]
required
Nonempty array of JSON-compatible values. batchSize and maxCharsPerRequest don’t combine them into a single request.
RequestProcessingPreference | null
Concurrency, timeout, retry, and header settings. Set custom headers with extraHeaders; batchPost has no separate headers argument. See Request preferences.
string | null
default:"POST"
HTTP method. Also accepts "GET", "PUT", "PATCH", "DELETE", "HEAD", and "OPTIONS". Use uppercase names. The client sends request bodies for POST, PUT, and PATCH. For DELETE, HEAD, and OPTIONS, the response data contains empty objects rather than decoded response bodies.
Configure client for a deployment that serves /v1/completions, then send non-streaming completion requests:

Request preferences

Pass a RequestProcessingPreference to an inference method to set batching, concurrency, timeouts, and retries for that call. Its constructor accepts positional arguments in this order:
Pass undefined to retain an argument’s default when setting a later argument.
number | null
default:256
Maximum concurrent requests. Must be 1 to 1,024, or 1 to 512 when batchSize is below 16.
number | null
default:8
Maximum inputs per embedding, reranking, or classification request. Accepts 1 to 1,024.
number | null
default:3600
Timeout for each request, from 0.1 to 3,600 seconds.
number | null
default:8000
Character threshold for splitting text batches. Accepts 50 to 1,048,576. The client sends any text that exceeds the threshold intact in its own batch.
boolean | null
default:false
Send all initial requests in this call to one endpoint from the pool.
number | null
Delay in seconds before sending a duplicate request. Defaults to null, which disables hedging. When set, must be at least 0.045 seconds and less than timeoutS - 0.045.
number | null
Timeout for the entire operation. Defaults to null. Must be at least timeoutS when supplied.
number | null
Fraction used to calculate the operation’s hedge budget. Defaults to 0.10, or 10%. Accepts 0 to 3.
number | null
Fraction used to calculate the operation’s budget for timeout and network retries. Defaults to 0.05, or 5%. HTTP-status retries don’t consume this budget. Accepts 0 to 3.
number | null
default:5
Maximum retries per request. Accepts 0 to 6. Set to 0 to turn off retries.
number | null
default:125
Initial delay between retries, from 50 to 45,000 milliseconds.
CancellationToken | null
Cancellation state. Unset when omitted. Doesn’t prevent or interrupt inference requests in 0.1.15. See CancellationToken.
string | null
Accepts and stores a key but doesn’t change request authentication in 0.1.15. Defaults to null. Set apiKey on PerformanceClient instead.
Record<string, string> | null
Additional HTTP request headers. Defaults to null.
number[] | null
HTTP statuses to exclude from automatic retries. Defaults to an empty array.
TraceContext | null
W3C parent trace context. Defaults to null.
Read preferences through properties with the same names as the constructor arguments, except cancelToken, which has no public getter. Only traceContext has a setter. To change other preferences, construct a new instance. The client validates request preferences when you call an operation.

Response objects

Responses are plain JavaScript objects with snake_case property names, such as total_time. The TypeScript return type is Promise<any>; the package doesn’t export response classes or a numpy() helper. All responses include these fields:
number
Time in seconds for the overall operation.
number[]
Time in seconds for each batch request.
Record<string, string>[]
Response headers for each batch request.

Embedding response

string
Response object type.
string
Model name returned by the server.
object[]
Embedding results.
object
Token usage for the operation.

Reranking response

string
Response object type.
object[]
Reranking results.

Classification response

string
Response object type.
object[][]
One array of label-score results per input text.

Batch response

JsonValue[]
Decoded response payloads in input order. See batchPost for HTTP methods that return empty objects.

Helper types

The package entry point exports PerformanceClient, RequestProcessingPreference, HttpClientWrapper, and CancellationToken. Import these classes from @basetenlabs/performance-client. CommonJS applications can also use require("@basetenlabs/performance-client").

HttpClientWrapper

new HttpClientWrapper(httpVersion?, proxy?) creates a reusable HTTP connection pool. Pass the wrapper to the PerformanceClient constructor.
number | null
default:1
1 for HTTP/1.1 or 2 for HTTP/2.
string | null
Proxy URL for the connection pool.
Set BASETEN_SECOND_BASE_URL to another deployment URL, then share the wrapper between clients:

CancellationToken

new CancellationToken() creates a token. Pass it as the twelfth RequestProcessingPreference argument. token.cancel() sets its cancelled state. A cancelled token can’t be reset.
boolean
Whether the token is in its cancelled state. Read this property as token.isCancelled.
In 0.1.15, the Node.js inference methods don’t check or poll the token. Calling cancel() doesn’t prevent requests from starting or interrupt requests in flight. Use timeoutS and totalTimeoutS to bound request duration.

TraceContext

TraceContext is a TypeScript interface, not a runtime constructor:
string
required
W3C parent trace context.
string
Additional W3C trace state.
To attach requests to a parent trace, pass an object with these fields to the preference constructor or assign preference.traceContext. Don’t also set traceparent or tracestate in extraHeaders; the client rejects that combination. To export client spans to an OpenTelemetry collector, configure client tracing.

Endpoint

Endpoint configures a base URL and its health checks. In 0.1.15, it appears only in the TypeScript declarations, so you can’t construct it through the package import:
string
required
Endpoint base URL.
string
required
Key for health-check authentication. Inference requests use the client’s API key.
HttpClientWrapper
required
HTTP connection pool for health checks.
string | null
Absolute URL for an additional health check.
string | null
default:"/health"
Relative health-check path.
number | null
default:10
Seconds between health checks.
number | null
default:6
Timeout in seconds per health-check attempt.
number | null
default:2
Retries per health check.
boolean | null
default:false
Stop evaluating checks after the first failure.
boolean | null
default:true
Ignore a timeout from the deployment check when deciding whether the endpoint is healthy.
boolean | null
default:true
Ignore a timeout from the deep health check when deciding whether the endpoint is healthy.

EndpointPool

The TypeScript declaration new EndpointPool(endpoints, endpointWeights?) groups endpoints for a client.
Endpoint[]
required
Nonempty array of Endpoint instances with distinct base URLs.
number[] | null
Endpoint weights. Defaults to equal values. Custom weights must match the endpoint count, be finite and nonnegative, and include at least one positive value.
Like Endpoint, EndpointPool appears only in the type declarations. You can’t construct it through the package import or pass a new pool to PerformanceClient.

Errors

Inference methods reject their promises when a request or validation fails.
string
InvalidArg for invalid parameters. HTTP, network, connection, serialization, timeout, and cancellation failures have the code GenericFailure.
string
Failure description. Timeout messages identify Local timeout or Remote timeout.
HTTP errors don’t include a separate status code field. Handle errors from the embedding client:

Package source

The npm package includes the JavaScript exports, TypeScript declarations, and native bindings. See the Node.js binding source on GitHub for current development.