resources and the model in build_commands to match your model’s requirements.
Set up your environment
This guide uses the Baseten CLI. Sign in and installrequests to call the deployed model from Python. Browser login opens a tab to approve this device, so there’s no API key to copy and paste.
Sign in to Baseten
Install requests
Configure the model
Create a directory with aconfig.yaml file:
config.yaml:
config.yaml
base_image is a lightweight Python image. The build_commands install the system packages that the Ollama install script requires (curl, ca-certificates, and zstd), install Ollama, and then pull TinyLlama while the image builds. The slim base image doesn’t include these packages by default.
The build step also creates tinyllama:4t, a copy of the model with num_thread pinned to 4. Ollama sizes its thread pool from the host CPU count, not the instance’s vCPU count, so the unmodified model spreads across 192 threads on the host and a short reply takes minutes instead of seconds. Baking the thread count into the image means every request and every replica runs with it, and no request has to pass it. Raise num_thread if you move the deployment to a larger instance.
The start_command runs the Ollama server. The readiness_endpoint and liveness_endpoint both point to /api/tags, which returns successfully when Ollama is running. The predict_endpoint maps Baseten’s /predict route to Ollama’s /api/generate endpoint.
This example only needs 4 CPUs and 8 GB of memory. For a complete list of resource options, see the Resources page.
Deploy the model
Push the model to Baseten to start the deployment:Call the model
Ollama’s/api/generate is mapped to Baseten’s /predict route, so you can call the deployed model with any HTTP client:
- Baseten CLI
- cURL
- Python
To run inference with the Baseten CLI, use the
predict command:MODEL_ID with the model ID from your deployment output.
You should see:
Next steps
For higher-throughput serving on GPUs with OpenAI-compatible endpoints, see the vLLM and SGLang examples.Deploy LLMs with vLLM
Serve open-source LLMs on vLLM with prefix caching and the OpenAI-compatible API.
Deploy LLMs with SGLang
Serve open-source LLMs on SGLang’s high-performance runtime with the OpenAI-compatible API.