Skip to main content
The Startup time card described on this page is rolling out to organizations and might not appear on your environment and deployment pages yet.
Startup time is the time a new replica takes from creation until it passes its readiness check and can accept traffic. Every new replica goes through startup, whether it’s the first replica of a deployment scaled to zero or one the autoscaler adds under load. A synchronous request that arrives while a deployment has no ready replicas waits for startup to finish, so startup time becomes part of that request’s latency.

Startup phases

A new replica runs these phases in order. The Startup time card uses the same phase names. Total startup time runs from replica creation to ready, so it also includes short gaps between phases, such as container sandbox setup. Pulling weights appears only when BDN delivers the weights. Weights that your code downloads in load(), or through the legacy model_cache key, count toward Loading model instead. For Truss-built images, Baseten can stream the container image: the node starts your container once the files it needs at startup arrive, and the rest of the image downloads during Loading model. Deployment logs show “streaming-enabled image” when streaming is active.

Measure startup time

The Startup time card on the environment and deployment pages shows how long your replicas take to start. The environment page aggregates startups across every deployment that served the environment during the selected range. The deployment page shows startups for that deployment only. Choose a range of the last 24 hours or 7 days. The last 30 days range appears once 30 days of startup history are available. The card shows:
  • Startup duration: the mean total duration, from replica creation to ready, with a bar for each of p50, p90, p95, and p99, broken down by phase. The percentiles describe your startups better than the mean. Select a phase in the legend to highlight it in the chart and filter the history table.
  • Startup history: one row per replica startup, with its outcome, a timeline of its phases, and its start and ready times. Filter by outcome or phase, search by replica or deployment, and sort by outcome, duration, or start time.
Each startup has one of these outcomes:
  • Succeeded: the replica became ready.
  • Failed: the replica stopped before it became ready, for example from an exception in load(), running out of memory, or an image pull error.
  • Stalled: the replica didn’t become ready within 30 minutes.
For a failed or stalled startup, the timeline shows the last phase the replica reached. Start there, then open the logs for that time window for detail. A failure in Pulling image points to the image or registry. A failure in Loading model points to your load() code or to memory: an out-of-memory error during load needs a larger instance or smaller weights. The card counts only a replica’s first startup. A container that crashes and restarts appears in the Restarts chart on the Metrics tab instead.

Diagnose slow startups

Find the phase that dominates your startups, then apply the matching fix: When p50 and p99 are both high and Loading model dominates, image and weight caching can’t help: the time goes to work that runs on every startup. Optimize load(), and if your model uses torch.compile, enable compilation caching.

Reduce startup time

Startup time drops most when you reduce the duration of the phase that dominates it.

Smaller container images

A smaller image pulls faster on a node that doesn’t have it cached:
  • Remove unused entries from requirements and system_packages in your config.
  • Deliver weights through BDN instead of building them into the image.

Faster weight loading

BDN runs automatically on engine-builder deployments. On any other deployment, turn it on by adding a weights block to your config:
config.yaml
  • Use allow_patterns and ignore_patterns to skip files your model doesn’t serve, such as .bin checkpoints when .safetensors files exist.
  • Pin a revision (@<commit-sha>) so cached weights stay valid across deployments.
  • If you use model_cache, run truss migrate to move to BDN. See Cached weights.
Quantized weights (FP8 or FP4) reduce both download and load time.

Faster model loading

Loading model covers everything in load(), so you control most of it:
  • Don’t download files in load(). Move them to a weights block so BDN caches them and the time moves to Pulling weights.
  • Load weights from safetensors files instead of pickled checkpoints.
  • Move work that doesn’t need to happen before the first request out of load(), or into build time.
  • Keep custom health checks from blocking readiness on long warm-up work.

Compilation caching

torch.compile creates artifacts while a replica starts. Torch compile caching, built on b10cache, persists those artifacts so a new replica can reuse them instead of compiling from scratch. Benchmark your deployment to measure the effect on startup time.

Keep startup off the request path

To keep requests from waiting on startup, set min_replica to 1 or higher so a replica is always running, or raise it ahead of known traffic spikes, using your p99 startup time as the lead time. See Traffic patterns for pre-warming.

Next steps

  • Metrics: Response time, replica, and restart charts that complement the Startup time card.
  • Request lifecycle: What happens to requests during replica startup, including queuing and timeout behavior.
  • Autoscaling: Control how often replicas start and how many stay running.
  • Billing and usage: How startup time is metered.
  • Troubleshooting: Diagnose slow or failed startups.