The Startup time card described on this page is rolling out to organizations and might not appear on your environment and deployment pages yet.
Startup phases
A new replica runs these phases in order. The Startup time card uses the same phase names. Total startup time runs from replica creation to ready, so it also includes short gaps between phases, such as container sandbox setup.
Pulling weights appears only when BDN delivers the weights. Weights that your code downloads in
load(), or through the legacy model_cache key, count toward Loading model instead.
For Truss-built images, Baseten can stream the container image: the node starts your container once the files it needs at startup arrive, and the rest of the image downloads during Loading model. Deployment logs show “streaming-enabled image” when streaming is active.
Measure startup time
The Startup time card on the environment and deployment pages shows how long your replicas take to start. The environment page aggregates startups across every deployment that served the environment during the selected range. The deployment page shows startups for that deployment only. Choose a range of the last 24 hours or 7 days. The last 30 days range appears once 30 days of startup history are available. The card shows:- Startup duration: the mean total duration, from replica creation to ready, with a bar for each of p50, p90, p95, and p99, broken down by phase. The percentiles describe your startups better than the mean. Select a phase in the legend to highlight it in the chart and filter the history table.
- Startup history: one row per replica startup, with its outcome, a timeline of its phases, and its start and ready times. Filter by outcome or phase, search by replica or deployment, and sort by outcome, duration, or start time.
- Succeeded: the replica became ready.
- Failed: the replica stopped before it became ready, for example from an exception in
load(), running out of memory, or an image pull error. - Stalled: the replica didn’t become ready within 30 minutes.
load() code or to memory: an out-of-memory error during load needs a larger instance or smaller weights.
The card counts only a replica’s first startup. A container that crashes and restarts appears in the Restarts chart on the Metrics tab instead.
Diagnose slow startups
Find the phase that dominates your startups, then apply the matching fix:
When p50 and p99 are both high and Loading model dominates, image and weight caching can’t help: the time goes to work that runs on every startup. Optimize
load(), and if your model uses torch.compile, enable compilation caching.
Reduce startup time
Startup time drops most when you reduce the duration of the phase that dominates it.Smaller container images
A smaller image pulls faster on a node that doesn’t have it cached:- Remove unused entries from
requirementsandsystem_packagesin your config. - Deliver weights through BDN instead of building them into the image.
Faster weight loading
BDN runs automatically on engine-builder deployments. On any other deployment, turn it on by adding aweights block to your config:
config.yaml
- Use
allow_patternsandignore_patternsto skip files your model doesn’t serve, such as.bincheckpoints when.safetensorsfiles exist. - Pin a revision (
@<commit-sha>) so cached weights stay valid across deployments. - If you use
model_cache, runtruss migrateto move to BDN. See Cached weights.
Faster model loading
Loading model covers everything inload(), so you control most of it:
- Don’t download files in
load(). Move them to aweightsblock so BDN caches them and the time moves to Pulling weights. - Load weights from
safetensorsfiles instead of pickled checkpoints. - Move work that doesn’t need to happen before the first request out of
load(), or into build time. - Keep custom health checks from blocking readiness on long warm-up work.
Compilation caching
torch.compile creates artifacts while a replica starts. Torch compile caching, built on b10cache, persists those artifacts so a new replica can reuse them instead of compiling from scratch. Benchmark your deployment to measure the effect on startup time.
Keep startup off the request path
To keep requests from waiting on startup, setmin_replica to 1 or higher so a replica is always running, or raise it ahead of known traffic spikes, using your p99 startup time as the lead time. See Traffic patterns for pre-warming.
Next steps
- Metrics: Response time, replica, and restart charts that complement the Startup time card.
- Request lifecycle: What happens to requests during replica startup, including queuing and timeout behavior.
- Autoscaling: Control how often replicas start and how many stay running.
- Billing and usage: How startup time is metered.
- Troubleshooting: Diagnose slow or failed startups.