Skip to main content
Development deployments let you iterate on your model without redeploying from scratch each time you make a change. When you save a file, the Baseten CLI detects the change, calculates a patch, and applies it to the running deployment in seconds.
If baseten model push --watch isn’t a good fit, SSH access lets you connect to a running deployment with standard SSH. SSH works well for custom servers that watch doesn’t patch and for longer-lived interactive sessions. Pair it with a non-zero min_replicas and a code-syncing tool to iterate on a live container.

Start a development deployment

Create a development deployment and start watching for changes:
Terminal
The Baseten CLI creates a development deployment, waits for it to become ready, and begins watching your project directory for file changes.
To apply model code changes without restarting the inference server, add the --watch-hot-reload flag:
Terminal
See What gets live-patched for details and caveats about hot reload.

Re-attach to a development deployment

If you stop the watch session (Ctrl+C), re-attach to the existing development deployment with:
Terminal
You should see:
baseten model watch syncs any changes made while disconnected, then resumes watching. It requires an existing development deployment. If you don’t have one, use baseten model push --watch to create it. To apply model code changes without restarting, add the --hot-reload flag:
Terminal

What gets live-patched

Truss monitors your project directory (respecting .trussignore patterns) and applies patches for the following changes without a full rebuild: With the --watch-hot-reload or --hot-reload flags, the Baseten CLI hot-reloads model code changes by swapping the model class in-process without restarting the inference server. This preserves in-memory state like loaded weights and caches. If a patch includes non-model changes (such as requirements or config), it falls back to a standard restart.
Hot reload re-imports your module and updates __class__ on the existing model instance. It does not re-run __init__() or load(). If you add new instance state in those methods that predict() depends on, predict() calls will fail to see it. When your changes involve new instance state, stop the watch session and do a full reload with baseten model push --watch.

What requires a full redeploy

The patch system doesn’t support some changes. When you make these changes, stop the watch session and run baseten model push (or baseten model push --watch to start a new development deployment):
If a patch fails, the watcher prints an error and continues watching. Fix the issue in your source files and save again. For persistent failures, run baseten model push --watch to start fresh.

Limitations

Development deployments optimize for iteration, not production traffic:
  • Single replica: Fixed at 0 minimum, 1 maximum. No autoscaling beyond one replica.
  • No gRPC: Trusses with gRPC transport require a published deployment.
  • No TRT-LLM engine builds: TRT-LLM build flow requires a published deployment.
See Development deployments for the full autoscaling constraints.

Deploy to production

When you’re done iterating, deploy a published version:
Terminal
By default, baseten model push creates a published deployment with full autoscaling support. Published deployments can scale to multiple replicas and are suitable for production traffic. To deploy and promote directly to the production environment:

CLI reference: model push

Full list of options for the push command.

CLI reference: model watch

Full list of options for the watch command.

Autoscaling

Configure replicas, concurrency targets, and scale-to-zero for production.

Environments

Manage staging, production, and custom environments.