> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Node.js performance client

> Send concurrent embedding, reranking, classification, and custom HTTP requests from Node.js.

Inference methods return promises for use with `await`. The package includes TypeScript declarations.

See the [Performance Client overview](/reference/sdk/performance-client/overview) for shared batching, retry, and connection settings.

## Installation

```bash theme={"system"}
npm install @basetenlabs/performance-client@0.1.15
```

## First request

Set your API key, an embeddings deployment URL, and the `model` value expected by that deployment:

```bash theme={"system"}
export BASETEN_API_KEY="<YOUR_API_KEY>"
export BASETEN_BASE_URL="https://model-YOUR_MODEL_ID.api.baseten.co/environments/production/sync"
export BASETEN_MODEL="<MODEL_NAME>"
```

Save this as `embeddings.mjs`. The `.mjs` extension enables ES modules and top-level `await`.

```javascript embeddings.mjs theme={"system"}
import {
  PerformanceClient,
  RequestProcessingPreference,
} from "@basetenlabs/performance-client";

const client = new PerformanceClient(
  process.env.BASETEN_BASE_URL,
  process.env.BASETEN_API_KEY,
);
const preference = new RequestProcessingPreference(
  32, // maxConcurrentRequests
  8,  // batchSize
  30, // timeoutS
);
const response = await client.embed(
  ["Hello world", "Example text"],
  process.env.BASETEN_MODEL,
  "float",
  undefined,
  undefined,
  preference,
);
console.log(`${response.data.length} embeddings`);
console.log(`Total tokens: ${response.usage.total_tokens}`);
console.log(`Elapsed seconds: ${response.total_time.toFixed(3)}`);
```

Run `node embeddings.mjs`. A successful response prints `2 embeddings`, followed by token usage and elapsed time. The operation examples below use the same imports and a `client` configured for the corresponding deployment.

## `PerformanceClient`

```typescript theme={"system"}
new PerformanceClient(
  baseUrl: string,
  apiKey?: string | null,
  httpVersion?: number | null,
  clientWrapper?: HttpClientWrapper | null,
  proxy?: string | null,
  endpointPool?: EndpointPool | null,
)
```

<ParamField body="baseUrl" type="string" required>
  Base URL for requests. The client appends the request path. For an embeddings deployment, use the URL ending in `/sync`, without `/v1/embeddings`.
</ParamField>

<ParamField body="apiKey" type="string | null">
  Inference API key. When omitted, the client checks `BASETEN_API_KEY`, then `OPENAI_API_KEY`. Construction fails if no key is available.
</ParamField>

<ParamField body="httpVersion" type="number | null" default={1}>
  `1` for HTTP/1.1 or `2` for HTTP/2.
</ParamField>

<ParamField body="clientWrapper" type="HttpClientWrapper | null">
  Shared HTTP connection pool. When supplied, its HTTP version and proxy configuration take precedence.
</ParamField>

<ParamField body="proxy" type="string | null">
  Proxy URL for a newly created HTTP client.
</ParamField>

<ParamField body="endpointPool" type="EndpointPool | null">
  Set of endpoints the client can send requests to. Version `0.1.15` doesn't export [`Endpoint`](#endpoint) or [`EndpointPool`](#endpointpool), so you can't create a pool through the package import.
</ParamField>

### `getClientWrapper`

`client.getClientWrapper()` returns the `HttpClientWrapper` used by the client. Pass it to another client to share connections.

## Inference methods

Optional arguments are positional. Pass `undefined` or `null` to skip an argument. There are no separate synchronous or `asyncEmbed` methods.

| Method | Path appended to the base URL |
| - | - |
| `embed` | `/v1/embeddings` |
| `rerank` | `/rerank` |
| `classify` | `/predict` |
| `batchPost` | Supplied `urlPath`. |

### `embed`

Generates embeddings for the input texts.

```typescript theme={"system"}
client.embed(
  input: string[],
  model: string,
  encodingFormat?: string | null,
  dimensions?: number | null,
  user?: string | null,
  preference?: RequestProcessingPreference | null,
): Promise<any>
```

<ParamField body="input" type="string[]" required>
  Nonempty list of strings to embed.
</ParamField>

<ParamField body="model" type="string" required>
  Model value expected by the server.
</ParamField>

<ParamField body="encodingFormat" type="string | null">
  Use `"float"`. In version `0.1.15`, base64 embeddings become empty arrays in the returned response. If omitted, the server selects the format.
</ParamField>

<ParamField body="dimensions" type="number | null">
  Embedding size supported by the model.
</ParamField>

<ParamField body="user" type="string | null">
  User identifier sent to the endpoint.
</ParamField>

<ParamField body="preference" type="RequestProcessingPreference | null">
  Batching, concurrency, timeout, and retry settings for this call. See [Request preferences](#request-preferences).
</ParamField>

Returns an [embedding response](#embedding-response) with results, usage, and timing fields.

### `rerank`

Scores texts against a query.

```typescript theme={"system"}
client.rerank(
  query: string,
  texts: string[],
  rawScores?: boolean | null,
  model?: string | null,
  returnText?: boolean | null,
  truncate?: boolean | null,
  truncationDirection?: string | null,
  preference?: RequestProcessingPreference | null,
): Promise<any>
```

<ParamField body="query" type="string" required>
  Query to score the texts against.
</ParamField>

<ParamField body="texts" type="string[]" required>
  Nonempty list of texts to rerank.
</ParamField>

<ParamField body="rawScores" type="boolean | null" default={false}>
  Request raw scores from the server.
</ParamField>

<ParamField body="model" type="string | null">
  Model value sent to the server.
</ParamField>

<ParamField body="returnText" type="boolean | null" default={false}>
  Include the input text in each result.
</ParamField>

<ParamField body="truncate" type="boolean | null" default={false}>
  Let the server truncate inputs.
</ParamField>

<ParamField body="truncationDirection" type="string | null" default="Right">
  Direction for the server to truncate inputs.
</ParamField>

<ParamField body="preference" type="RequestProcessingPreference | null">
  Batching, concurrency, timeout, and retry settings for this call. See [Request preferences](#request-preferences).
</ParamField>

The model server determines support for `model`, `rawScores`, `returnText`, `truncate`, and `truncationDirection`.

Returns a [reranking response](#reranking-response).

Configure `client` for a reranking deployment, then score the texts:

```javascript theme={"system"}
const response = await client.rerank(
  "How do I deploy a model?",
  ["Deploy a model with Truss.", "Create an API key in settings."],
  false,     // rawScores
  undefined, // model
  true,      // returnText
  false,     // truncate
  "Right",   // truncationDirection
  new RequestProcessingPreference(32, 8, 30),
);
for (const result of response.data) {
  console.log(result.index, result.score, result.text);
}
```

### `classify`

Assigns labels and scores to the input texts.

```typescript theme={"system"}
client.classify(
  inputs: string[],
  model?: string | null,
  rawScores?: boolean | null,
  truncate?: boolean | null,
  truncationDirection?: string | null,
  preference?: RequestProcessingPreference | null,
): Promise<any>
```

<ParamField body="inputs" type="string[]" required>
  Nonempty list of strings. The client sends each string as a one-element list in the request's `inputs` field.
</ParamField>

<ParamField body="model" type="string | null">
  Model value sent to the server.
</ParamField>

<ParamField body="rawScores" type="boolean | null" default={false}>
  Request raw scores from the server.
</ParamField>

<ParamField body="truncate" type="boolean | null" default={false}>
  Let the server truncate inputs.
</ParamField>

<ParamField body="truncationDirection" type="string | null" default="Right">
  Direction for the server to truncate inputs.
</ParamField>

<ParamField body="preference" type="RequestProcessingPreference | null">
  Batching, concurrency, timeout, and retry settings for this call. See [Request preferences](#request-preferences).
</ParamField>

Returns a [classification response](#classification-response).

Configure `client` for a classification deployment, then classify the texts:

```javascript theme={"system"}
const response = await client.classify(
  ["The setup worked well.", "The request failed."],
  undefined, // model
  false,     // rawScores
  false,     // truncate
  "Right",   // truncationDirection
  new RequestProcessingPreference(32, 8, 30),
);
for (const group of response.data) {
  for (const result of group) {
    console.log(result.label, result.score);
  }
}
```

### `batchPost`

Sends one HTTP request per payload and returns responses in input order.

```typescript theme={"system"}
client.batchPost(
  urlPath: string,
  payloads: JsonValue[],
  preference?: RequestProcessingPreference | null,
  method?: string | null,
): Promise<any>
```

<ParamField body="urlPath" type="string" required>
  Path appended to the base URL, such as `/predict`.
</ParamField>

<ParamField body="payloads" type="JsonValue[]" required>
  Nonempty array of JSON-compatible values. `batchSize` and `maxCharsPerRequest` don't combine them into a single request.
</ParamField>

<ParamField body="preference" type="RequestProcessingPreference | null">
  Concurrency, timeout, retry, and header settings. Set custom headers with `extraHeaders`; `batchPost` has no separate headers argument. See [Request preferences](#request-preferences).
</ParamField>

<ParamField body="method" type="string | null" default="POST">
  HTTP method. Also accepts `"GET"`, `"PUT"`, `"PATCH"`, `"DELETE"`, `"HEAD"`, and `"OPTIONS"`. Use uppercase names. The client sends request bodies for POST, PUT, and PATCH. For DELETE, HEAD, and OPTIONS, the response data contains empty objects rather than decoded response bodies.
</ParamField>

Configure `client` for a deployment that serves `/v1/completions`, then send non-streaming completion requests:

```javascript theme={"system"}
const payloads = ["Explain embeddings.", "Explain reranking."].map(prompt => ({
  model: process.env.BASETEN_MODEL,
  prompt,
  stream: false,
}));
const response = await client.batchPost(
  "/v1/completions",
  payloads,
  new RequestProcessingPreference(32, undefined, 60),
);
response.data.forEach((body, index) => {
  console.log(body, response.response_headers[index]);
});
```

## Request preferences

Pass a `RequestProcessingPreference` to an inference method to set batching, concurrency, timeouts, and retries for that call. Its constructor accepts positional arguments in this order:

```typescript theme={"system"}
new RequestProcessingPreference(
  maxConcurrentRequests?: number | null,
  batchSize?: number | null,
  timeoutS?: number | null,
  maxCharsPerRequest?: number | null,
  pinInitialEndpointOnce?: boolean | null,
  hedgeDelay?: number | null,
  totalTimeoutS?: number | null,
  hedgeBudgetPct?: number | null,
  retryBudgetPct?: number | null,
  maxRetries?: number | null,
  initialBackoffMs?: number | null,
  cancelToken?: CancellationToken | null,
  primaryApiKeyOverride?: string | null,
  extraHeaders?: Record<string, string> | null,
  nonRetryableStatusCodes?: number[] | null,
  traceContext?: TraceContext | null,
)
```

```javascript theme={"system"}
import { RequestProcessingPreference } from "@basetenlabs/performance-client";

const preference = new RequestProcessingPreference(
  64, // maxConcurrentRequests
  8,  // batchSize
  30, // timeoutS
);
```

Pass `undefined` to retain an argument's default when setting a later argument.

<ParamField body="maxConcurrentRequests" type="number | null" default={256}>
  Maximum concurrent requests. Must be 1 to 1,024, or 1 to 512 when `batchSize` is below 16.
</ParamField>

<ParamField body="batchSize" type="number | null" default={8}>
  Maximum inputs per embedding, reranking, or classification request. Accepts 1 to 1,024.
</ParamField>

<ParamField body="timeoutS" type="number | null" default={3600}>
  Timeout for each request, from 0.1 to 3,600 seconds.
</ParamField>

<ParamField body="maxCharsPerRequest" type="number | null" default={8000}>
  Character threshold for splitting text batches. Accepts 50 to 1,048,576. The client sends any text that exceeds the threshold intact in its own batch.
</ParamField>

<ParamField body="pinInitialEndpointOnce" type="boolean | null" default={false}>
  Send all initial requests in this call to one endpoint from the pool.
</ParamField>

<ParamField body="hedgeDelay" type="number | null">
  Delay in seconds before sending a duplicate request. Defaults to `null`, which disables hedging. When set, must be at least 0.045 seconds and less than `timeoutS - 0.045`.
</ParamField>

<ParamField body="totalTimeoutS" type="number | null">
  Timeout for the entire operation. Defaults to `null`. Must be at least `timeoutS` when supplied.
</ParamField>

<ParamField body="hedgeBudgetPct" type="number | null">
  Fraction used to calculate the operation's hedge budget. Defaults to `0.10`, or 10%. Accepts 0 to 3.
</ParamField>

<ParamField body="retryBudgetPct" type="number | null">
  Fraction used to calculate the operation's budget for timeout and network retries. Defaults to `0.05`, or 5%. HTTP-status retries don't consume this budget. Accepts 0 to 3.
</ParamField>

<ParamField body="maxRetries" type="number | null" default={5}>
  Maximum retries per request. Accepts 0 to 6. Set to `0` to turn off retries.
</ParamField>

<ParamField body="initialBackoffMs" type="number | null" default={125}>
  Initial delay between retries, from 50 to 45,000 milliseconds.
</ParamField>

<ParamField body="cancelToken" type="CancellationToken | null">
  Cancellation state. Unset when omitted. Doesn't prevent or interrupt inference requests in `0.1.15`. See [`CancellationToken`](#cancellationtoken).
</ParamField>

<ParamField body="primaryApiKeyOverride" type="string | null">
  Accepts and stores a key but doesn't change request authentication in `0.1.15`. Defaults to `null`. Set `apiKey` on `PerformanceClient` instead.
</ParamField>

<ParamField body="extraHeaders" type="Record<string, string> | null">
  Additional HTTP request headers. Defaults to `null`.
</ParamField>

<ParamField body="nonRetryableStatusCodes" type="number[] | null">
  HTTP statuses to exclude from automatic retries. Defaults to an empty array.
</ParamField>

<ParamField body="traceContext" type="TraceContext | null">
  W3C parent trace context. Defaults to `null`.
</ParamField>

Read preferences through properties with the same names as the constructor arguments, except `cancelToken`, which has no public getter. Only `traceContext` has a setter. To change other preferences, construct a new instance.

The client validates request preferences when you call an operation.

## Response objects

Responses are plain JavaScript objects with snake\_case property names, such as `total_time`. The TypeScript return type is `Promise<any>`; the package doesn't export response classes or a `numpy()` helper.

All responses include these fields:

<ResponseField name="total_time" type="number">
  Time in seconds for the overall operation.
</ResponseField>

<ResponseField name="individual_request_times" type="number[]">
  Time in seconds for each batch request.
</ResponseField>

<ResponseField name="response_headers" type="Record<string, string>[]">
  Response headers for each batch request.
</ResponseField>

### Embedding response

<ResponseField name="object" type="string">
  Response object type.
</ResponseField>

<ResponseField name="model" type="string">
  Model name returned by the server.
</ResponseField>

<ResponseField name="data" type="object[]">
  Embedding results.

  <Expandable title="Embedding properties">
    <ResponseField name="object" type="string">
      Embedding object type.
    </ResponseField>

    <ResponseField name="index" type="number">
      Original input index.
    </ResponseField>

    <ResponseField name="embedding" type="number[]">
      Embedding vector. Base64 embeddings become empty arrays in `0.1.15`. Call [`embed`](#embed) with `encodingFormat` set to `"float"` to receive the vector values.
    </ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="usage" type="object">
  Token usage for the operation.

  <Expandable title="Properties">
    <ResponseField name="prompt_tokens" type="number">
      Prompt token count.
    </ResponseField>

    <ResponseField name="total_tokens" type="number">
      Total token count.
    </ResponseField>
  </Expandable>
</ResponseField>

### Reranking response

<ResponseField name="object" type="string">
  Response object type.
</ResponseField>

<ResponseField name="data" type="object[]">
  Reranking results.

  <Expandable title="Result properties">
    <ResponseField name="index" type="number">
      Original input index.
    </ResponseField>

    <ResponseField name="score" type="number">
      Score for the input text.
    </ResponseField>

    <ResponseField name="text" type="string | null">
      Input text when returned by the server.
    </ResponseField>
  </Expandable>
</ResponseField>

### Classification response

<ResponseField name="object" type="string">
  Response object type.
</ResponseField>

<ResponseField name="data" type="object[][]">
  One array of label-score results per input text.

  <Expandable title="Result properties">
    <ResponseField name="label" type="string">
      Classification label.
    </ResponseField>

    <ResponseField name="score" type="number">
      Score for the label.
    </ResponseField>
  </Expandable>
</ResponseField>

### Batch response

<ResponseField name="data" type="JsonValue[]">
  Decoded response payloads in input order. See [`batchPost`](#batchpost) for HTTP methods that return empty objects.
</ResponseField>

## Helper types

The package entry point exports `PerformanceClient`, `RequestProcessingPreference`, `HttpClientWrapper`, and `CancellationToken`. Import these classes from `@basetenlabs/performance-client`. CommonJS applications can also use `require("@basetenlabs/performance-client")`.

### `HttpClientWrapper`

`new HttpClientWrapper(httpVersion?, proxy?)` creates a reusable HTTP connection pool. Pass the wrapper to the `PerformanceClient` constructor.

<ParamField body="httpVersion" type="number | null" default={1}>
  `1` for HTTP/1.1 or `2` for HTTP/2.
</ParamField>

<ParamField body="proxy" type="string | null">
  Proxy URL for the connection pool.
</ParamField>

Set `BASETEN_SECOND_BASE_URL` to another deployment URL, then share the wrapper between clients:

```javascript theme={"system"}
import { HttpClientWrapper, PerformanceClient } from "@basetenlabs/performance-client";

const wrapper = new HttpClientWrapper(1);
const first = new PerformanceClient(
  process.env.BASETEN_BASE_URL,
  process.env.BASETEN_API_KEY,
  1,
  wrapper,
);
const second = new PerformanceClient(
  process.env.BASETEN_SECOND_BASE_URL,
  process.env.BASETEN_API_KEY,
  1,
  wrapper,
);
```

### `CancellationToken`

`new CancellationToken()` creates a token. Pass it as the twelfth `RequestProcessingPreference` argument. `token.cancel()` sets its cancelled state. A cancelled token can't be reset.

<ResponseField name="isCancelled" type="boolean">
  Whether the token is in its cancelled state. Read this property as `token.isCancelled`.
</ResponseField>

In `0.1.15`, the Node.js inference methods don't check or poll the token. Calling `cancel()` doesn't prevent requests from starting or interrupt requests in flight. Use `timeoutS` and `totalTimeoutS` to bound request duration.

### `TraceContext`

`TraceContext` is a TypeScript interface, not a runtime constructor:

```typescript theme={"system"}
interface TraceContext {
  traceparent: string;
  tracestate?: string;
}
```

<ParamField body="traceparent" type="string" required>
  W3C parent trace context.
</ParamField>

<ParamField body="tracestate" type="string">
  Additional W3C trace state.
</ParamField>

To attach requests to a parent trace, pass an object with these fields to the preference constructor or assign `preference.traceContext`. Don't also set `traceparent` or `tracestate` in `extraHeaders`; the client rejects that combination.

To export client spans to an OpenTelemetry collector, configure [client tracing](/reference/sdk/performance-client/overview#client-tracing).

### `Endpoint`

`Endpoint` configures a base URL and its health checks. In `0.1.15`, it appears only in the TypeScript declarations, so you can't construct it through the package import:

```typescript theme={"system"}
new Endpoint(
  baseUrl: string,
  apiKey: string,
  clientWrapper: HttpClientWrapper,
  deepHealthUrl?: string | null,
  deploymentHealthPath?: string | null,
  healthCheckIntervalS?: number | null,
  healthCheckTimeoutS?: number | null,
  healthCheckRetries?: number | null,
  healthFailOnFirst?: boolean | null,
  deploymentTimeoutIsNoVote?: boolean | null,
  deepTimeoutIsNoVote?: boolean | null,
)
```

<ParamField body="baseUrl" type="string" required>
  Endpoint base URL.
</ParamField>

<ParamField body="apiKey" type="string" required>
  Key for health-check authentication. Inference requests use the client's API key.
</ParamField>

<ParamField body="clientWrapper" type="HttpClientWrapper" required>
  HTTP connection pool for health checks.
</ParamField>

<ParamField body="deepHealthUrl" type="string | null">
  Absolute URL for an additional health check.
</ParamField>

<ParamField body="deploymentHealthPath" type="string | null" default="/health">
  Relative health-check path.
</ParamField>

<ParamField body="healthCheckIntervalS" type="number | null" default={10}>
  Seconds between health checks.
</ParamField>

<ParamField body="healthCheckTimeoutS" type="number | null" default={6}>
  Timeout in seconds per health-check attempt.
</ParamField>

<ParamField body="healthCheckRetries" type="number | null" default={2}>
  Retries per health check.
</ParamField>

<ParamField body="healthFailOnFirst" type="boolean | null" default={false}>
  Stop evaluating checks after the first failure.
</ParamField>

<ParamField body="deploymentTimeoutIsNoVote" type="boolean | null" default={true}>
  Ignore a timeout from the deployment check when deciding whether the endpoint is healthy.
</ParamField>

<ParamField body="deepTimeoutIsNoVote" type="boolean | null" default={true}>
  Ignore a timeout from the deep health check when deciding whether the endpoint is healthy.
</ParamField>

### `EndpointPool`

The TypeScript declaration `new EndpointPool(endpoints, endpointWeights?)` groups endpoints for a client.

<ParamField body="endpoints" type="Endpoint[]" required>
  Nonempty array of `Endpoint` instances with distinct base URLs.
</ParamField>

<ParamField body="endpointWeights" type="number[] | null">
  Endpoint weights. Defaults to equal values. Custom weights must match the endpoint count, be finite and nonnegative, and include at least one positive value.
</ParamField>

Like [`Endpoint`](#endpoint), `EndpointPool` appears only in the type declarations. You can't construct it through the package import or pass a new pool to `PerformanceClient`.

## Errors

Inference methods reject their promises when a request or validation fails.

<ResponseField name="code" type="string">
  `InvalidArg` for invalid parameters. HTTP, network, connection, serialization, timeout, and cancellation failures have the code `GenericFailure`.
</ResponseField>

<ResponseField name="message" type="string">
  Failure description. Timeout messages identify `Local timeout` or `Remote timeout`.
</ResponseField>

HTTP errors don't include a separate status code field.

Handle errors from the embedding client:

```javascript theme={"system"}
try {
  const response = await client.embed(
    ["Hello world"],
    process.env.BASETEN_MODEL,
    "float",
    undefined,
    undefined,
    new RequestProcessingPreference(32, 8, 30),
  );
  console.log(`${response.data.length} embeddings`);
} catch (error) {
  console.error(error.code, error.message);
}
```

## Package source

The [npm package](https://www.npmjs.com/package/@basetenlabs/performance-client/v/0.1.15) includes the JavaScript exports, TypeScript declarations, and native bindings. See the [Node.js binding source](https://github.com/basetenlabs/truss/tree/main/baseten-performance-client/node_bindings) on GitHub for current development.
