Skip to main content
Ollama is a popular lightweight LLM inference server, similar to vLLM or SGLang. This guide deploys an Ollama model as a custom Docker server on Baseten. This configuration serves TinyLlama with Ollama on a CPU instance. The deployment process is the same for larger Ollama models. Adjust the resources and the model in build_commands to match your model’s requirements.

Set up your environment

This guide uses the Baseten CLI. Sign in and install requests to call the deployed model from Python. Browser login opens a tab to approve this device, so there’s no API key to copy and paste.
Sign in to Baseten
Install requests

Configure the model

Create a directory with a config.yaml file:
Copy the following configuration into config.yaml:
config.yaml
The base_image is a lightweight Python image. The build_commands install the system packages that the Ollama install script requires (curl, ca-certificates, and zstd), install Ollama, and then pull TinyLlama while the image builds. The slim base image doesn’t include these packages by default. The build step also creates tinyllama:4t, a copy of the model with num_thread pinned to 4. Ollama sizes its thread pool from the host CPU count, not the instance’s vCPU count, so the unmodified model spreads across 192 threads on the host and a short reply takes minutes instead of seconds. Baking the thread count into the image means every request and every replica runs with it, and no request has to pass it. Raise num_thread if you move the deployment to a larger instance. The start_command runs the Ollama server. The readiness_endpoint and liveness_endpoint both point to /api/tags, which returns successfully when Ollama is running. The predict_endpoint maps Baseten’s /predict route to Ollama’s /api/generate endpoint. This example only needs 4 CPUs and 8 GB of memory. For a complete list of resource options, see the Resources page.

Deploy the model

Push the model to Baseten to start the deployment:
You should see output like:
Copy the model ID from the output for the next step. The first deploy takes several minutes while Baseten builds the image, and Ollama downloads TinyLlama into it. Later pushes of the same configuration reuse the cached image, so scaling up starts a server that already has the model on disk.

Call the model

Ollama’s /api/generate is mapped to Baseten’s /predict route, so you can call the deployed model with any HTTP client:
To run inference with the Baseten CLI, use the predict command:
Replace MODEL_ID with the model ID from your deployment output. You should see:

Next steps

For higher-throughput serving on GPUs with OpenAI-compatible endpoints, see the vLLM and SGLang examples.

Deploy LLMs with vLLM

Serve open-source LLMs on vLLM with prefix caching and the OpenAI-compatible API.

Deploy LLMs with SGLang

Serve open-source LLMs on SGLang’s high-performance runtime with the OpenAI-compatible API.