Skip to main content
Use POST /v1/models to deploy a Truss archive. To deploy a model published through Baseten for Model Labs, see Deploy a Model Labs listing. To create the model in a specific team, use POST /v1/teams/{team_id}/models with the same request body. Use GET /v1/teams to find the team ID. The organization-scoped endpoint uses the organization’s default team. Deploying an archive follows these steps:
  1. Prepare: POST /v1/prepare_model_upload validates the payload and returns temporary credentials scoped to an S3 location.
  2. Upload: push your Truss archive to that location.
  3. Create: POST /v1/models commits the upload as a new model.
For a team-scoped archive deployment, also pass team_id when preparing the upload.

Prepare the upload

Send a Truss config as a JSON object with a model name. Add a weights block to load weights through the Baseten Delivery Network. Set dry_run to true to validate without issuing credentials. The response carries the upload credentials and the S3 location to upload to:

Upload the archive

Package your Truss as a gzipped tar archive, then upload it to the returned s3_bucket and s3_key using the temporary credentials:
upload.py
A successful upload returns nothing. boto3 raises an exception if the temporary credentials have expired or the s3_key doesn’t match the one from the prepare step.

Create the model

Commit the upload with source.kind set to model_archive, the same deployment payload you validated, and the s3_key from the prepare step. The response returns the created model and its first deployment:

Check the deployment

Creation returns before the deployment is ready. Use the model and deployment IDs from the response to poll GET /v1/models/{model_id}/deployments/{deployment_id} until its status is ACTIVE:

Call the model

If the deployment fails, check its deployment logs. Once the deployment is ACTIVE, send inference requests to the model’s predict endpoint, using the model id from the create response and your API key. The request and response shapes match whatever your model’s predict method accepts and returns:

Next steps

Call your model

Stream responses, send async requests, and use the other inference transports.

Add a deployment

Push a new deployment to the model you created.