Skip to main content
Choose per-replica concurrency by measuring speech latency under the traffic you expect to serve. Compare configurations that meet your quality and latency targets, then calculate cost using achieved throughput and utilization. Open your model’s Metrics view and select the deployment and time range for your test.
Metrics view showing completed WebSocket connections: green for normal closures and red for an internal error from an invalid-metadata test.

WebSocket metrics during verification. The green 1000s series groups normal closures. The red 1011s series records an invalid-metadata test. Points appear after sessions close.

Start with Ongoing websocket connections and Replicas to understand load. Use Inference volume to inspect how connections ended, and End-to-end connection duration to measure their lifetimes. A normal close code confirms transport cleanup; use the model’s final-flush signal to confirm transcription completion. A long connection can contain many short turns. Measure each turn in your client to find the delay your application experiences.

Workload evaluation

Choose measurements that match how your application sends and receives audio: Real-time factor (RTF) is processing time divided by audio duration. An RTF below 1 means processing is faster than the recording’s duration. RTFx is the reciprocal, audio duration divided by processing time. Calculate both from the durations recorded in your test. A real-time-paced transcription session takes approximately the recording’s duration even when the model has spare capacity. Measure batch throughput with complete recordings sent without real-time pacing.

Latency and quality measurements

Record event timestamps with a monotonic clock. Use the same start and end events when comparing configurations. Finalization includes the model’s voice activity detection (VAD) policy when silence triggers the final. A commit-to-flush measurement starts at an explicit client event and has a different boundary. Label them separately. For transcription quality, measure word error rate (WER) against reference transcripts, with the dataset and text normalization recorded. Semantic WER uses a different scoring method, so compare it only with results using the same evaluator. For speaker attribution, record diarization error rate (DER) and the evaluation’s overlap and boundary rules.

Deployment metrics

Deployment metrics describe the transport and replicas. Pair them with client events to diagnose individual turns. Time to first byte (TTFB) measures the first response byte. A control message can precede a transcript or audio payload. Instrument the first relevant payload in your client instead of treating TTFB as first audio or first transcript. An HTTP failure before the WebSocket upgrade and an abnormal WebSocket close occur at different stages. Use the WebSocket status-code reference to distinguish connection setup from an established session ending.

Concurrency and cost

Start with a warm single-session baseline, then increase simultaneous sessions while holding the audio and model settings constant. Include silence if your application keeps connections open between turns. Record unsuccessful sessions as failures, rather than excluding them from the latency results. On September 18, 2026, a verification run of the quickstart client produced the following results. Each batch used the same 11-second recording at real-time pace, with 100 ms chunks and partial updates requested every 0.5 seconds. The deployment used Qwen3-ASR-1.7B on one RTX Pro 6000 replica with vLLM 0.22.0 and a platform concurrency target of 96. The client ran on macOS against a deployment with Global region selection. The test didn’t record physical region placement. All 21 sessions completed. This small repeated-fixture test verifies the measurement procedure; it doesn’t establish production capacity or tail latency. The audio includes silence before commit, so these values don’t measure end-of-speech finalization. Use a production-sized dataset to choose capacity. For example, if your application requires p95 finalization below 800 ms, test candidate concurrency values against that exact boundary. Reject settings that exceed the target or increase the error rate. The 800 ms threshold is an example application target, not a Baseten guarantee. Read model-card concurrency recommendations alongside their GPU, dataset, and timing definition. A deployment’s initial autoscaling settings can differ from a benchmark recommendation. Confirm the values applied to your deployment before testing. Calculate unit cost over the same interval as the measured workload:
  • Transcription cost per audio hour: Total deployment cost divided by the total audio hours processed.
  • Speech generation cost per million input characters: Total deployment cost divided by input characters processed, multiplied by 1,000,000.
Include idle replicas and startup time in deployment cost. Count audio across all sessions, including overlapping sessions, when calculating processed audio hours. Use achieved throughput over the interval rather than a model card’s peak throughput. A lightly used deployment can have low latency and high cost per audio hour.

WebSocket capacity

Each WebSocket stays on one replica for its lifetime and counts toward that replica’s concurrency target while idle. Closing unused sessions releases capacity. During a promotion, existing sessions continue on the previous deployment while new connections route to the promoted deployment. The platform concurrency target controls replica scaling. The model server also has its own scheduling and memory limits. Increasing the platform target doesn’t increase the model’s processing capacity. Keep the replica range and model configuration fixed while comparing concurrency, then test the autoscaling behavior your production traffic needs. To change the target in the dashboard or through the Management API, follow Update autoscaling settings. Confirm the applied value before you run another test. Warm replicas avoid model-loading delay for new sessions. Choose a minimum replica count that meets your startup target, and a maximum that covers your tested peak load. WebSocket scheduling describes how Baseten places connections and scales down idle replicas.

Benchmark interpretation

Keep a record for each test so another engineer can reproduce it: Retest the selected setting with varied recordings and enough sessions to estimate the tail. Compare cold starts separately from warm inference. For text-to-speech, include playback buffering and check whether audio generation stalls after the first chunk.

Common symptoms

Use the measurement closest to the symptom before changing deployment settings.

Next steps