> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Startup times

> The phases of replica startup, how to measure startup time, and the model and config changes that reduce it.

export const MiniColdStart = () => {
  const ref = React.useRef(null);
  const init = React.useRef(false);
  React.useEffect(() => {
    if (!ref.current || init.current) return;
    init.current = true;
    const PER = 11000;
    const STAGES = [["Scaled to zero", 0], ["Waking up", 0.5], ["Active", 1]];
    const STAGE_PH = {
      "Scaled to zero": 0.01,
      "Waking up": 0.20,
      "Active": 0.65
    };
    const SEGS = [["Acquiring resources", 0.06, 0.14], ["Pulling image", 0.14, 0.21], ["Pulling weights", 0.21, 0.33], ["Loading model", 0.33, 0.46]];
    const START = SEGS[0][1], READY = SEGS[SEGS.length - 1][2];
    const LX = 40, TX0 = 180, TX1 = 540, TY = 62, ROW = 22, BH = 10;
    const SPAN_Y = TY + SEGS.length * ROW + 8;
    const tx = v => TX0 + (v - START) / (READY - START) * (TX1 - TX0);
    let dotRadii = STAGES.map(() => 2.5);
    let frozen = null;
    const opts = {
      title: "Replica startup",
      desc: "Click a status or a startup phase to freeze the loop there. Click anywhere else to resume.",
      formula: "startup time = replica created → readiness check passes",
      W: 580,
      H: 205,
      onClick: function (x, y) {
        let target = null;
        if (y > 10 && y < 48) {
          let nearest = null, nearestDist = Infinity;
          STAGES.forEach(s => {
            const d = Math.abs(x - (40 + s[1] * 500));
            if (d < nearestDist) {
              nearestDist = d;
              nearest = s[0];
            }
          });
          if (nearestDist < 90) target = STAGE_PH[nearest];
        } else if (y >= TY && y < TY + SEGS.length * ROW) {
          const seg = SEGS[Math.floor((y - TY) / ROW)];
          target = (seg[1] + seg[2]) / 2;
        }
        frozen = target != null && frozen !== target ? target : null;
      },
      draw: function (ctx, t, p) {
        const h = window._mdHelpers;
        const ph = frozen != null ? frozen : t % PER / PER;
        let status = "Scaled to zero";
        if (ph < START) status = "Scaled to zero"; else if (ph < 0.94) status = ph < READY ? "Waking up" : "Active";
        ctx.strokeStyle = p.brd;
        ctx.lineWidth = 0.8;
        ctx.beginPath();
        ctx.moveTo(40, 36);
        ctx.lineTo(540, 36);
        ctx.stroke();
        STAGES.forEach((s, i) => {
          const sx = h.lerp(40, 540, s[1]);
          const cur = status === s[0];
          dotRadii[i] += ((cur ? 4.2 : 2.5) - dotRadii[i]) * 0.18;
          ctx.fillStyle = cur ? p.p : p.brd;
          ctx.beginPath();
          ctx.arc(sx, 36, dotRadii[i], 0, Math.PI * 2);
          ctx.fill();
          ctx.font = (cur ? "600 " : "500 ") + "9.5px ui-monospace,Menlo,monospace";
          ctx.fillStyle = cur ? p.txt : p.sub;
          ctx.textAlign = i === 0 ? "left" : i === STAGES.length - 1 ? "right" : "center";
          ctx.textBaseline = "middle";
          ctx.fillText(s[0], sx, 22);
        });
        const barA = h.lerp(1, 0.55, h.fade(ph, READY, 0.52)) * (1 - h.fade(ph, 0.84, 0.92) * 0.7);
        SEGS.forEach((seg, i) => {
          const cy = TY + i * ROW + ROW / 2;
          const active = ph >= seg[1] && ph < seg[2];
          const done = ph >= seg[2] && ph < 0.92;
          ctx.font = (active ? "600 " : "500 ") + "9.5px ui-monospace,Menlo,monospace";
          ctx.fillStyle = active ? p.w : done ? p.txt : p.sub;
          ctx.textAlign = "left";
          ctx.textBaseline = "middle";
          ctx.fillText(seg[0], LX, cy);
          ctx.strokeStyle = p.brdM;
          ctx.lineWidth = 1;
          ctx.beginPath();
          ctx.moveTo(TX0, cy);
          ctx.lineTo(TX1, cy);
          ctx.stroke();
          const x0 = tx(seg[1]) + 1, x1 = tx(seg[2]) - 1;
          ctx.globalAlpha = barA;
          const frac = ph < 0.92 ? h.fade(ph, seg[1], seg[2]) : 0;
          if (frac < 1) {
            ctx.setLineDash([3, 3]);
            ctx.strokeStyle = p.brd;
            ctx.beginPath();
            ctx.roundRect(x0, cy - BH / 2, x1 - x0, BH, 2);
            ctx.stroke();
            ctx.setLineDash([]);
          }
          if (frac > 0) {
            ctx.fillStyle = frac >= 1 ? p.p : p.w;
            ctx.beginPath();
            ctx.roundRect(x0, cy - BH / 2, (x1 - x0) * frac, BH, 2);
            ctx.fill();
          }
          ctx.globalAlpha = 1;
        });
        if (ph >= START && ph < READY) {
          const px = tx(ph);
          ctx.strokeStyle = p.w;
          ctx.lineWidth = 1;
          ctx.beginPath();
          ctx.moveTo(px, TY - 4);
          ctx.lineTo(px, SPAN_Y + 4);
          ctx.stroke();
        }
        const spanA = ph >= START && ph < 0.92 ? h.fade(ph, START, START + 0.02) * (1 - h.fade(ph, 0.84, 0.92)) : 0;
        if (spanA > 0) {
          const sx1 = tx(Math.min(ph, READY));
          const ready = ph >= READY;
          ctx.globalAlpha = spanA;
          ctx.strokeStyle = ready ? p.p : p.w;
          ctx.lineWidth = 1;
          ctx.beginPath();
          ctx.moveTo(TX0, SPAN_Y - 4);
          ctx.lineTo(TX0, SPAN_Y);
          ctx.lineTo(sx1, SPAN_Y);
          if (ready) ctx.lineTo(sx1, SPAN_Y - 4);
          ctx.stroke();
          ctx.font = "500 9.5px ui-monospace,Menlo,monospace";
          ctx.fillStyle = ready ? p.p : p.w;
          ctx.textAlign = "left";
          ctx.textBaseline = "middle";
          ctx.fillText(ready ? "Startup time" : "Startup time so far", LX, SPAN_Y);
          ctx.globalAlpha = 1;
        }
        ctx.font = "500 10px ui-monospace,Menlo,monospace";
        ctx.fillStyle = p.sub;
        ctx.textAlign = "left";
        ctx.textBaseline = "middle";
        let cap = "";
        if (ph < 0.03) cap = "No replicas running. The deployment is Scaled to zero."; else if (ph < START) cap = "A request arrives and triggers a scale-up"; else if (ph < SEGS[1][1]) cap = "Acquiring resources: the replica is scheduled onto a GPU node"; else if (ph < SEGS[2][1]) cap = "Pulling image: the node pulls the image (skipped if cached)"; else if (ph < SEGS[3][1]) cap = "Pulling weights: BDN downloads weights to the node and mounts them"; else if (ph < READY) cap = "Loading model: container starts, load() runs, readiness passes"; else if (ph < 0.84) cap = "The replica is ready and serves traffic"; else if (ph < 0.94) cap = "Traffic stops and the replica scales down"; else cap = "No replicas running. The deployment is Scaled to zero.";
        ctx.fillText(cap, 40, 186);
      }
    };
    let cleanup = null, destroyed = false, timer = null, retries = 0;
    const tryMount = () => {
      if (destroyed || !ref.current) return;
      if (window._mdMount) {
        cleanup = window._mdMount(ref.current, opts);
      } else if (retries++ < 60) {
        timer = setTimeout(tryMount, 30);
      }
    };
    tryMount();
    return () => {
      destroyed = true;
      if (timer) clearTimeout(timer);
      if (cleanup) cleanup();
      init.current = false;
    };
  }, []);
  return <div ref={ref} />;
};

export const MiniDiagramEngine = () => {
  React.useEffect(() => {
    if (window._mdMount) return;
    const isDark = () => document.documentElement.classList.contains("dark");
    const lerp = (a, b, t) => a + (b - a) * Math.min(1, Math.max(0, t));
    const fade = (p, a, b) => Math.min(1, Math.max(0, (p - a) / (b - a)));
    function P() {
      const d = isDark();
      return {
        bg: d ? "#021309" : "#fff",
        sub: "#869089",
        brd: d ? "#344339" : "#dee4de",
        brdM: d ? "#203026" : "#f4f9f3",
        q: d ? "#4a90ff" : "#2176ff",
        qFill: d ? "rgba(74,144,255,0.18)" : "rgba(199,220,255,0.7)",
        qDark: "#114aa6",
        w: d ? "#f7c42f" : "#9c7400",
        p: d ? "#19E76E" : "#0e863f",
        rb: "#005934",
        rbf: d ? "rgba(25,231,110,0.22)" : "rgba(178,247,207,0.55)",
        rsf: d ? "rgba(247,196,47,0.18)" : "rgba(253,237,188,0.55)",
        txt: d ? "#dee4de" : "#0c1d13"
      };
    }
    function repBox(ctx, x, y, w, h, state, label, prog, p, displayState) {
      let fl = p.bg, st = p.p, tc = p.p, ds = [];
      if (state === "stopped") {
        fl = "transparent";
        st = p.brd;
        tc = p.sub;
        ds = [3, 3];
      } else if (state === "starting") {
        fl = p.rsf;
        st = p.w;
        tc = p.w;
      } else if (state === "busy") {
        fl = p.rbf;
        st = p.rb;
        tc = p.rb;
      } else if (state === "stopping") {
        st = p.sub;
        tc = p.sub;
      }
      ctx.globalAlpha = state === "stopped" || state === "stopping" ? 0.55 : 1;
      ctx.setLineDash(ds);
      ctx.beginPath();
      ctx.roundRect(x, y, w, h, 6);
      ctx.fillStyle = fl;
      ctx.fill();
      ctx.strokeStyle = st;
      ctx.lineWidth = 1.3;
      ctx.stroke();
      ctx.setLineDash([]);
      ctx.globalAlpha = 1;
      ctx.font = "500 11px ui-monospace,Menlo,monospace";
      ctx.fillStyle = tc;
      ctx.textAlign = "left";
      ctx.textBaseline = "middle";
      ctx.fillText(label, x + 10, y + h / 2);
      ctx.font = "500 9.5px ui-monospace,Menlo,monospace";
      ctx.textAlign = "right";
      ctx.fillText(displayState || state, x + w - 10, y + h / 2);
      if (state === "starting" && prog > 0) {
        ctx.fillStyle = p.w;
        ctx.fillRect(x, y + h - 2, w * prog, 2);
      }
    }
    function setRich(el, s) {
      el.replaceChildren();
      const parts = s.split("`");
      for (let i = 0; i < parts.length; i++) {
        if (i % 2 === 0) {
          if (parts[i]) el.appendChild(document.createTextNode(parts[i]));
        } else {
          const code = document.createElement("code");
          code.textContent = parts[i];
          code.style.cssText = "font-family:ui-monospace,Menlo,monospace;font-size:0.92em;background:" + (isDark() ? "rgba(255,255,255,0.08)" : "rgba(0,0,0,0.05)") + ";padding:1px 4px;border-radius:3px";
          el.appendChild(code);
        }
      }
    }
    window._mdHelpers = {
      lerp,
      fade,
      repBox
    };
    window._mdMount = function (root, opts) {
      const W = opts.W || 580, H = opts.H || 200;
      const card = document.createElement("div");
      const tit = document.createElement("div");
      tit.style.cssText = "font:500 12px ui-monospace,Menlo,monospace;letter-spacing:-0.28px;margin:0 0 4px";
      const desc = document.createElement("p");
      desc.style.cssText = "margin:0 0 10px;font-size:12px;line-height:1.45;font-family:system-ui,-apple-system,sans-serif";
      const cv = document.createElement("canvas");
      cv.style.cssText = "display:block;width:100%;max-width:" + W + "px;touch-action:pan-y";
      const formula = document.createElement("div");
      formula.style.cssText = "margin:8px 0 0;font:500 11px ui-monospace,Menlo,monospace;letter-spacing:-0.28px;border-radius:4px;padding:4px 8px;display:inline-block";
      setRich(tit, opts.title);
      setRich(desc, opts.desc);
      if (opts.formula) setRich(formula, opts.formula);
      card.appendChild(tit);
      card.appendChild(desc);
      card.appendChild(cv);
      if (opts.formula) card.appendChild(formula);
      root.appendChild(card);
      const ctx = cv.getContext("2d");
      const dpr = window.devicePixelRatio || 1;
      cv.width = W * dpr;
      cv.height = H * dpr;
      cv.style.height = H + "px";
      ctx.scale(dpr, dpr);
      if (opts.onClick) {
        cv.style.cursor = "pointer";
        cv.addEventListener("click", e => {
          const r = cv.getBoundingClientRect();
          opts.onClick((e.clientX - r.left) / r.width * W, (e.clientY - r.top) / r.height * H);
        });
      }
      function applyTheme() {
        const d = isDark();
        card.style.cssText = "border:1px solid " + (d ? "#344339" : "#f4f9f3") + ";border-radius:8px;padding:16px 18px;margin:12px 0;background:" + (d ? "#021309" : "#fff") + ";max-width:" + W + "px";
        tit.style.color = d ? "#dee4de" : "#0c1d13";
        desc.style.color = d ? "#9CA59E" : "#5a675e";
        formula.style.background = d ? "#0C1D13" : "#f4f9f3";
        formula.style.borderColor = d ? "#203026" : "#dee4de";
        formula.style.border = "1px solid " + (d ? "#203026" : "#dee4de");
        formula.style.color = d ? "#19E76E" : "#0e863f";
        if (opts.title) setRich(tit, opts.title);
        if (opts.desc) setRich(desc, opts.desc);
        if (opts.formula) setRich(formula, opts.formula);
      }
      applyTheme();
      let visible = true, raf = 0, t0 = performance.now(), dirty = true;
      const obs = new IntersectionObserver(e => visible = e[0].isIntersecting, {
        threshold: 0.15
      });
      obs.observe(cv);
      const themeObs = new MutationObserver(() => {
        dirty = true;
        applyTheme();
      });
      themeObs.observe(document.documentElement, {
        attributes: true,
        attributeFilter: ["class"]
      });
      function loop(ts) {
        raf = requestAnimationFrame(loop);
        if (!visible) {
          dirty = true;
          return;
        }
        const t = ts - t0;
        ctx.clearRect(0, 0, W, H);
        opts.draw(ctx, t, P());
        dirty = false;
      }
      raf = requestAnimationFrame(loop);
      return () => {
        cancelAnimationFrame(raf);
        obs.disconnect();
        themeObs.disconnect();
        card.remove();
      };
    };
    return () => {
      delete window._mdMount;
      delete window._mdHelpers;
    };
  }, []);
  return <span />;
};

<MiniDiagramEngine />

<Note>
  The **Startup time** card described on this page is rolling out to organizations and might not appear on your environment and deployment pages yet.
</Note>

*Startup time* is the time a new replica takes from creation until it passes its readiness check and can accept traffic. Every new replica goes through startup, whether it's the first replica of a deployment scaled to zero or one the autoscaler adds under load. A synchronous request that arrives while a deployment has no ready replicas waits for startup to finish, so startup time becomes part of that request's latency.

<MiniColdStart />

## Startup phases

A new replica runs these phases in order. The **Startup time** card uses the same phase names. Total startup time runs from replica creation to ready, so it also includes short gaps between phases, such as container sandbox setup.

| Phase | Starts | Ends | What drives its duration |
| - | - | - | - |
| Acquiring resources | Baseten creates the replica. | The replica is scheduled onto a node. | Availability of the requested instance type. |
| Pulling image | The node starts pulling your container image. | The image is available on the node. | Image size. A replica on a node that already has your image skips this phase. |
| Pulling weights | The node starts mounting the [Baseten Delivery Network (BDN)](/development/model/bdn) weights volume. | All weight files are on the node and mounted. | Weight size and which BDN cache tier serves them. |
| Loading model | The node creates your model container. | The replica passes its readiness check. | Your setup code. The container starts, the model server starts, and your `load()` method runs. For inference engines like vLLM and SGLang, this includes loading weights into GPU memory, capturing CUDA graphs, compiling kernels with `torch.compile`, and profiling the KV cache. |

**Pulling weights** appears only when BDN delivers the weights. Weights that your code downloads in `load()`, or through the legacy `model_cache` key, count toward **Loading model** instead.

For Truss-built images, Baseten can stream the container image: the node starts your container once the files it needs at startup arrive, and the rest of the image downloads during **Loading model**. Deployment logs show "streaming-enabled image" when streaming is active.

## Measure startup time

The **Startup time** card on the environment and deployment pages shows how long your replicas take to start. The environment page aggregates startups across every deployment that served the environment during the selected range. The deployment page shows startups for that deployment only. Choose a range of the last 24 hours or 7 days. The last 30 days range appears once 30 days of startup history are available.

The card shows:

* **Startup duration:** the mean total duration, from replica creation to ready, with a bar for each of p50, p90, p95, and p99, broken down by phase. The percentiles describe your startups better than the mean. Select a phase in the legend to highlight it in the chart and filter the history table.
* **Startup history:** one row per replica startup, with its outcome, a timeline of its phases, and its start and ready times. Filter by outcome or phase, search by replica or deployment, and sort by outcome, duration, or start time.

Each startup has one of these outcomes:

* **Succeeded:** the replica became ready.
* **Failed:** the replica stopped before it became ready, for example from an exception in `load()`, running out of memory, or an image pull error.
* **Stalled:** the replica didn't become ready within 30 minutes.

For a failed or stalled startup, the timeline shows the last phase the replica reached. Start there, then open the [logs](/observability/logs) for that time window for detail. A failure in **Pulling image** points to the image or registry. A failure in **Loading model** points to your `load()` code or to memory: an out-of-memory error during load needs a larger instance or smaller weights.

The card counts only a replica's first startup. A container that crashes and restarts appears in the [Restarts](/observability/metrics#restarts) chart on the Metrics tab instead.

## Diagnose slow startups

Find the phase that dominates your startups, then apply the matching fix:

| Signal | Likely cause | What to do |
| - | - | - |
| Long **Acquiring resources** | Waiting for capacity | Choose a more available instance type. If this phase stays long, [contact support](mailto:support@baseten.co). |
| Long or frequent **Pulling image** | Large image, or replicas landing on nodes without your image | Build a [smaller image](#smaller-container-images). |
| Long **Pulling weights** | Large weights or cache misses | [Filter and pin your weights](#faster-weight-loading), or quantize the model. |
| Long **Loading model** | Heavy setup code, or weights downloaded in `load()` | [Speed up model loading](#faster-model-loading). |
| Many rows in **Startup history** | Autoscaling churn or scale-to-zero cycling | Tune your [autoscaling settings](/deployment/autoscaling/overview) so replicas start less often. |
| Wide gap between p50 and p99 | Some startups hit cold caches | Your p50 startups reuse cached images and weights while your p99 startups don't. Enable BDN and reduce your image size before you optimize `load()`. |

When p50 and p99 are both high and **Loading model** dominates, image and weight caching can't help: the time goes to work that runs on every startup. Optimize `load()`, and if your model uses `torch.compile`, enable [compilation caching](#compilation-caching).

## Reduce startup time

Startup time drops most when you reduce the duration of the phase that dominates it.

### Smaller container images

A smaller image pulls faster on a node that doesn't have it cached:

* Remove unused entries from `requirements` and `system_packages` in your config.
* Deliver weights through BDN instead of building them into the image.

### Faster weight loading

BDN runs automatically on engine-builder deployments. On any other deployment, turn it on by adding a [`weights`](/development/model/bdn) block to your config:

```yaml config.yaml theme={"system"}
weights:
  - source: "hf://Qwen/Qwen3-8B@<commit-sha>"
    mount_location: "/models/qwen"
    allow_patterns: ["*.safetensors", "*.json"]
```

* Use `allow_patterns` and `ignore_patterns` to skip files your model doesn't serve, such as `.bin` checkpoints when `.safetensors` files exist.
* Pin a revision (`@<commit-sha>`) so cached weights stay valid across deployments.
* If you use `model_cache`, run `truss migrate` to move to BDN. See [Cached weights](/development/model/model-cache).

Quantized weights (FP8 or FP4) reduce both download and load time.

### Faster model loading

**Loading model** covers everything in `load()`, so you control most of it:

* Don't download files in `load()`. Move them to a `weights` block so BDN caches them and the time moves to **Pulling weights**.
* Load weights from `safetensors` files instead of pickled checkpoints.
* Move work that doesn't need to happen before the first request out of `load()`, or into build time.
* Keep [custom health checks](/development/model/health-checks#custom-health-check-logic) from blocking readiness on long warm-up work.

### Compilation caching

`torch.compile` creates artifacts while a replica starts. [Torch compile caching](/development/model/runtime-caching#torch-compile-caching), built on [b10cache](/development/model/runtime-caching), persists those artifacts so a new replica can reuse them instead of compiling from scratch. Benchmark your deployment to measure the effect on startup time.

### Keep startup off the request path

To keep requests from waiting on startup, set [`min_replica`](/deployment/autoscaling/overview#param-min-replica) to 1 or higher so a replica is always running, or raise it ahead of known traffic spikes, using your p99 startup time as the lead time. See [Traffic patterns](/deployment/autoscaling/traffic-patterns#pre-warming-for-predictable-bursts) for pre-warming.

## Next steps

* [Metrics](/observability/metrics): Response time, replica, and restart charts that complement the **Startup time** card.
* [Request lifecycle](/deployment/autoscaling/request-lifecycle): What happens to requests during replica startup, including queuing and timeout behavior.
* [Autoscaling](/deployment/autoscaling/overview): Control how often replicas start and how many stay running.
* [Billing and usage](/organization/billing): How startup time is metered.
* [Troubleshooting](/troubleshooting/deployments#autoscaling-issues): Diagnose slow or failed startups.
