await. The package includes TypeScript declarations.
See the Performance Client overview for shared batching, retry, and connection settings.
Installation
First request
Set your API key, an embeddings deployment URL, and themodel value expected by that deployment:
embeddings.mjs. The .mjs extension enables ES modules and top-level await.
embeddings.mjs
node embeddings.mjs. A successful response prints 2 embeddings, followed by token usage and elapsed time. The operation examples below use the same imports and a client configured for the corresponding deployment.
PerformanceClient
string
required
Base URL for requests. The client appends the request path. For an embeddings deployment, use the URL ending in
/sync, without /v1/embeddings.string | null
Inference API key. When omitted, the client checks
BASETEN_API_KEY, then OPENAI_API_KEY. Construction fails if no key is available.number | null
default:1
1 for HTTP/1.1 or 2 for HTTP/2.HttpClientWrapper | null
Shared HTTP connection pool. When supplied, its HTTP version and proxy configuration take precedence.
string | null
Proxy URL for a newly created HTTP client.
EndpointPool | null
Set of endpoints the client can send requests to. Version
0.1.15 doesn’t export Endpoint or EndpointPool, so you can’t create a pool through the package import.getClientWrapper
client.getClientWrapper() returns the HttpClientWrapper used by the client. Pass it to another client to share connections.
Inference methods
Optional arguments are positional. Passundefined or null to skip an argument. There are no separate synchronous or asyncEmbed methods.
embed
Generates embeddings for the input texts.
string[]
required
Nonempty list of strings to embed.
string
required
Model value expected by the server.
string | null
Use
"float". In version 0.1.15, base64 embeddings become empty arrays in the returned response. If omitted, the server selects the format.number | null
Embedding size supported by the model.
string | null
User identifier sent to the endpoint.
RequestProcessingPreference | null
Batching, concurrency, timeout, and retry settings for this call. See Request preferences.
rerank
Scores texts against a query.
string
required
Query to score the texts against.
string[]
required
Nonempty list of texts to rerank.
boolean | null
default:false
Request raw scores from the server.
string | null
Model value sent to the server.
boolean | null
default:false
Include the input text in each result.
boolean | null
default:false
Let the server truncate inputs.
string | null
default:"Right"
Direction for the server to truncate inputs.
RequestProcessingPreference | null
Batching, concurrency, timeout, and retry settings for this call. See Request preferences.
model, rawScores, returnText, truncate, and truncationDirection.
Returns a reranking response.
Configure client for a reranking deployment, then score the texts:
classify
Assigns labels and scores to the input texts.
string[]
required
Nonempty list of strings. The client sends each string as a one-element list in the request’s
inputs field.string | null
Model value sent to the server.
boolean | null
default:false
Request raw scores from the server.
boolean | null
default:false
Let the server truncate inputs.
string | null
default:"Right"
Direction for the server to truncate inputs.
RequestProcessingPreference | null
Batching, concurrency, timeout, and retry settings for this call. See Request preferences.
client for a classification deployment, then classify the texts:
batchPost
Sends one HTTP request per payload and returns responses in input order.
string
required
Path appended to the base URL, such as
/predict.JsonValue[]
required
Nonempty array of JSON-compatible values.
batchSize and maxCharsPerRequest don’t combine them into a single request.RequestProcessingPreference | null
Concurrency, timeout, retry, and header settings. Set custom headers with
extraHeaders; batchPost has no separate headers argument. See Request preferences.string | null
default:"POST"
HTTP method. Also accepts
"GET", "PUT", "PATCH", "DELETE", "HEAD", and "OPTIONS". Use uppercase names. The client sends request bodies for POST, PUT, and PATCH. For DELETE, HEAD, and OPTIONS, the response data contains empty objects rather than decoded response bodies.client for a deployment that serves /v1/completions, then send non-streaming completion requests:
Request preferences
Pass aRequestProcessingPreference to an inference method to set batching, concurrency, timeouts, and retries for that call. Its constructor accepts positional arguments in this order:
undefined to retain an argument’s default when setting a later argument.
number | null
default:256
Maximum concurrent requests. Must be 1 to 1,024, or 1 to 512 when
batchSize is below 16.number | null
default:8
Maximum inputs per embedding, reranking, or classification request. Accepts 1 to 1,024.
number | null
default:3600
Timeout for each request, from 0.1 to 3,600 seconds.
number | null
default:8000
Character threshold for splitting text batches. Accepts 50 to 1,048,576. The client sends any text that exceeds the threshold intact in its own batch.
boolean | null
default:false
Send all initial requests in this call to one endpoint from the pool.
number | null
Delay in seconds before sending a duplicate request. Defaults to
null, which disables hedging. When set, must be at least 0.045 seconds and less than timeoutS - 0.045.number | null
Timeout for the entire operation. Defaults to
null. Must be at least timeoutS when supplied.number | null
Fraction used to calculate the operation’s hedge budget. Defaults to
0.10, or 10%. Accepts 0 to 3.number | null
Fraction used to calculate the operation’s budget for timeout and network retries. Defaults to
0.05, or 5%. HTTP-status retries don’t consume this budget. Accepts 0 to 3.number | null
default:5
Maximum retries per request. Accepts 0 to 6. Set to
0 to turn off retries.number | null
default:125
Initial delay between retries, from 50 to 45,000 milliseconds.
CancellationToken | null
Cancellation state. Unset when omitted. Doesn’t prevent or interrupt inference requests in
0.1.15. See CancellationToken.string | null
Accepts and stores a key but doesn’t change request authentication in
0.1.15. Defaults to null. Set apiKey on PerformanceClient instead.Record<string, string> | null
Additional HTTP request headers. Defaults to
null.number[] | null
HTTP statuses to exclude from automatic retries. Defaults to an empty array.
TraceContext | null
W3C parent trace context. Defaults to
null.cancelToken, which has no public getter. Only traceContext has a setter. To change other preferences, construct a new instance.
The client validates request preferences when you call an operation.
Response objects
Responses are plain JavaScript objects with snake_case property names, such astotal_time. The TypeScript return type is Promise<any>; the package doesn’t export response classes or a numpy() helper.
All responses include these fields:
number
Time in seconds for the overall operation.
number[]
Time in seconds for each batch request.
Record<string, string>[]
Response headers for each batch request.
Embedding response
string
Response object type.
string
Model name returned by the server.
object[]
Embedding results.
object
Token usage for the operation.
Reranking response
string
Response object type.
object[]
Reranking results.
Classification response
string
Response object type.
object[][]
One array of label-score results per input text.
Batch response
JsonValue[]
Decoded response payloads in input order. See
batchPost for HTTP methods that return empty objects.Helper types
The package entry point exportsPerformanceClient, RequestProcessingPreference, HttpClientWrapper, and CancellationToken. Import these classes from @basetenlabs/performance-client. CommonJS applications can also use require("@basetenlabs/performance-client").
HttpClientWrapper
new HttpClientWrapper(httpVersion?, proxy?) creates a reusable HTTP connection pool. Pass the wrapper to the PerformanceClient constructor.
number | null
default:1
1 for HTTP/1.1 or 2 for HTTP/2.string | null
Proxy URL for the connection pool.
BASETEN_SECOND_BASE_URL to another deployment URL, then share the wrapper between clients:
CancellationToken
new CancellationToken() creates a token. Pass it as the twelfth RequestProcessingPreference argument. token.cancel() sets its cancelled state. A cancelled token can’t be reset.
boolean
Whether the token is in its cancelled state. Read this property as
token.isCancelled.0.1.15, the Node.js inference methods don’t check or poll the token. Calling cancel() doesn’t prevent requests from starting or interrupt requests in flight. Use timeoutS and totalTimeoutS to bound request duration.
TraceContext
TraceContext is a TypeScript interface, not a runtime constructor:
string
required
W3C parent trace context.
string
Additional W3C trace state.
preference.traceContext. Don’t also set traceparent or tracestate in extraHeaders; the client rejects that combination.
To export client spans to an OpenTelemetry collector, configure client tracing.
Endpoint
Endpoint configures a base URL and its health checks. In 0.1.15, it appears only in the TypeScript declarations, so you can’t construct it through the package import:
string
required
Endpoint base URL.
string
required
Key for health-check authentication. Inference requests use the client’s API key.
HttpClientWrapper
required
HTTP connection pool for health checks.
string | null
Absolute URL for an additional health check.
string | null
default:"/health"
Relative health-check path.
number | null
default:10
Seconds between health checks.
number | null
default:6
Timeout in seconds per health-check attempt.
number | null
default:2
Retries per health check.
boolean | null
default:false
Stop evaluating checks after the first failure.
boolean | null
default:true
Ignore a timeout from the deployment check when deciding whether the endpoint is healthy.
boolean | null
default:true
Ignore a timeout from the deep health check when deciding whether the endpoint is healthy.
EndpointPool
The TypeScript declaration new EndpointPool(endpoints, endpointWeights?) groups endpoints for a client.
Endpoint[]
required
Nonempty array of
Endpoint instances with distinct base URLs.number[] | null
Endpoint weights. Defaults to equal values. Custom weights must match the endpoint count, be finite and nonnegative, and include at least one positive value.
Endpoint, EndpointPool appears only in the type declarations. You can’t construct it through the package import or pass a new pool to PerformanceClient.
Errors
Inference methods reject their promises when a request or validation fails.string
InvalidArg for invalid parameters. HTTP, network, connection, serialization, timeout, and cancellation failures have the code GenericFailure.string
Failure description. Timeout messages identify
Local timeout or Remote timeout.