
WebSocket metrics during verification. The green 1000s series groups normal closures. The red 1011s series records an invalid-metadata test. Points appear after sessions close.
Workload evaluation
Choose measurements that match how your application sends and receives audio:
Real-time factor (RTF) is processing time divided by audio duration. An RTF below 1 means processing is faster than the recording’s duration. RTFx is the reciprocal, audio duration divided by processing time. Calculate both from the durations recorded in your test.
A real-time-paced transcription session takes approximately the recording’s duration even when the model has spare capacity. Measure batch throughput with complete recordings sent without real-time pacing.
Latency and quality measurements
Record event timestamps with a monotonic clock. Use the same start and end events when comparing configurations.
Finalization includes the model’s voice activity detection (VAD) policy when silence triggers the final. A commit-to-flush measurement starts at an explicit client event and has a different boundary. Label them separately.
For transcription quality, measure word error rate (WER) against reference transcripts, with the dataset and text normalization recorded. Semantic WER uses a different scoring method, so compare it only with results using the same evaluator. For speaker attribution, record diarization error rate (DER) and the evaluation’s overlap and boundary rules.
Deployment metrics
Deployment metrics describe the transport and replicas. Pair them with client events to diagnose individual turns.
Time to first byte (TTFB) measures the first response byte. A control message can precede a transcript or audio payload. Instrument the first relevant payload in your client instead of treating TTFB as first audio or first transcript.
An HTTP failure before the WebSocket upgrade and an abnormal WebSocket close occur at different stages. Use the WebSocket status-code reference to distinguish connection setup from an established session ending.
Concurrency and cost
Start with a warm single-session baseline, then increase simultaneous sessions while holding the audio and model settings constant. Include silence if your application keeps connections open between turns. Record unsuccessful sessions as failures, rather than excluding them from the latency results. On September 18, 2026, a verification run of the quickstart client produced the following results. Each batch used the same 11-second recording at real-time pace, with 100 ms chunks and partial updates requested every 0.5 seconds. The deployment used Qwen3-ASR-1.7B on one RTX Pro 6000 replica with vLLM 0.22.0 and a platform concurrency target of 96. The client ran on macOS against a deployment with Global region selection. The test didn’t record physical region placement.
All 21 sessions completed. This small repeated-fixture test verifies the measurement procedure; it doesn’t establish production capacity or tail latency. The audio includes silence before commit, so these values don’t measure end-of-speech finalization.
Use a production-sized dataset to choose capacity. For example, if your application requires p95 finalization below 800 ms, test candidate concurrency values against that exact boundary. Reject settings that exceed the target or increase the error rate. The 800 ms threshold is an example application target, not a Baseten guarantee.
Read model-card concurrency recommendations alongside their GPU, dataset, and timing definition. A deployment’s initial autoscaling settings can differ from a benchmark recommendation. Confirm the values applied to your deployment before testing.
Calculate unit cost over the same interval as the measured workload:
- Transcription cost per audio hour: Total deployment cost divided by the total audio hours processed.
- Speech generation cost per million input characters: Total deployment cost divided by input characters processed, multiplied by 1,000,000.
WebSocket capacity
Each WebSocket stays on one replica for its lifetime and counts toward that replica’s concurrency target while idle. Closing unused sessions releases capacity. During a promotion, existing sessions continue on the previous deployment while new connections route to the promoted deployment. The platform concurrency target controls replica scaling. The model server also has its own scheduling and memory limits. Increasing the platform target doesn’t increase the model’s processing capacity. Keep the replica range and model configuration fixed while comparing concurrency, then test the autoscaling behavior your production traffic needs. To change the target in the dashboard or through the Management API, follow Update autoscaling settings. Confirm the applied value before you run another test. Warm replicas avoid model-loading delay for new sessions. Choose a minimum replica count that meets your startup target, and a maximum that covers your tested peak load. WebSocket scheduling describes how Baseten places connections and scales down idle replicas.Benchmark interpretation
Keep a record for each test so another engineer can reproduce it:
Retest the selected setting with varied recordings and enough sessions to estimate the tail. Compare cold starts separately from warm inference. For text-to-speech, include playback buffering and check whether audio generation stalls after the first chunk.
Common symptoms
Use the measurement closest to the symptom before changing deployment settings.Next steps
- Change deployment capacity after selecting a tested setting.
- Export metrics into your monitoring system.
- Choose a deployment region near your clients.
- Fine-tune a speech model when quality evaluation identifies a task-specific gap.