> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Vision

> Fine-tune a vision-language model with image and text pairs, deploy the checkpoint, and send an image prompt.

Fine-tune a vision-language model on examples that pair an image with a text prompt and an answer. The trainer renders the image as visual tokens, concatenates them with the prompt and answer tokens, and computes loss over the combined sequence in the same [`forward_backward()`](/reference/sdk/loops/training-client) and [`optim_step()`](/reference/sdk/loops/training-client) round trip as [text-only supervised fine-tuning (SFT)](/loops/train-on-your-data). The training loop, checkpointing, and deployment stay the same; only the input construction changes.

This guide fine-tunes with LoRA on [`hiyouga/geometry3k`](https://huggingface.co/datasets/hiyouga/geometry3k), a dataset of geometry diagrams with text problems and numeric answers, then deploys the checkpoint and queries it with image prompts. The data-building and loss-masking techniques work with any image-and-text dataset. Vision training supports a subset of the [supported base models](/loops/supported-models): Qwen VL (Qwen3.5, Qwen3.6, Qwen3.8) and Kimi (K2.6, K2.7 Code).

## Prerequisites

* A [Baseten account](https://app.baseten.co/signup) and a [workspace API key](/organization/api-keys) with access to Loops
* Python 3.12 or newer and [uv](https://docs.astral.sh/uv/) to install dependencies and run the scripts

<Steps>
  <Step title="Export your API key">
    Export the key as `BASETEN_API_KEY`:

    <CodeGroup>
      ```bash macOS/Linux theme={"system"}
      export BASETEN_API_KEY="paste-your-api-key-here"
      ```

      ```powershell Windows theme={"system"}
      setx BASETEN_API_KEY "paste-your-api-key-here"
      ```
    </CodeGroup>
  </Step>

  <Step title="Install the Baseten CLI">
    The `baseten loops` and `baseten model` commands later in this guide need the CLI:

    <Tabs>
      <Tab title="macOS or Linux">
        ```bash Terminal theme={"system"}
        brew tap basetenlabs/baseten
        brew trust --formula basetenlabs/baseten/baseten
        brew install baseten
        ```
      </Tab>

      <Tab title="Windows">
        Download and extract the binary, then move `baseten.exe` to a directory on your `PATH`:

        ```powershell Terminal theme={"system"}
        Invoke-WebRequest `
          https://github.com/basetenlabs/baseten-cli/releases/download/v1.0.0/baseten_1.0.0_windows_amd64.zip `
          -OutFile baseten.zip; Expand-Archive -Force baseten.zip .
        ```
      </Tab>
    </Tabs>

    For other platforms or a specific version, see the [Baseten CLI install reference](/reference/cli/baseten/overview#install).
  </Step>
</Steps>

<Note>
  Loops is in early access. To enable it for your workspace, [fill out the signup form](https://www.baseten.co/talk-to-us/loops-signup/).
</Note>

## Install

Create a uv project and install the dependencies:

```bash theme={"system"} theme={"system"}
uv init loops-vision
cd loops-vision
uv add baseten-loops datasets pillow torch torchvision transformers
```

## Create the training script

The script opens a training run on a base model, converts dataset examples into training data, and trains a LoRA adapter before saving a checkpoint.

Create a file named `train_vision.py` in the project directory:

```python train_vision.py theme={"system"} theme={"system"}
import io

from datasets import load_dataset
from PIL import Image
from transformers import AutoImageProcessor

from baseten.loops import (
    AdamParams,
    Datum,
    ImageChunk,
    ModelInput,
    ServiceClient,
    TensorData,
)
```

`ServiceClient` opens a session with Baseten. `create_lora_training_client` provisions a trainer for the base model and attaches a LoRA adapter with the rank you set. `get_tokenizer` returns the base model's tokenizer, and the image processor parses image dimensions later in the script:

```python train_vision.py theme={"system"} theme={"system"}
BASE_MODEL = "Qwen/Qwen3.5-0.8B"

service_client = ServiceClient()
training_client = service_client.create_lora_training_client(
    base_model=BASE_MODEL,
    rank=16,
)
tokenizer = training_client.get_tokenizer()
image_processor = AutoImageProcessor.from_pretrained(BASE_MODEL, trust_remote_code=True)
```

## Turn examples into training data

Each example in the `geometry3k` dataset contains an image, a problem, and an answer. The trainer uses each example as a [`Datum`](/reference/sdk/loops/types):

* `model_input` contains one [`ImageChunk`](/reference/sdk/loops/types) with the image bytes and one text chunk with the problem and answer tokens.
* `loss_fn_inputs` supplies the loss targets through a [`TensorData`](/reference/sdk/loops/types) named `target_tokens`, with one target token per input position.

`ImageChunk` takes the image bytes inline, as JPEG or PNG. One chunk holds one image; add more chunks for more images. Pass the image already resized: the trainer doesn't resize images and rejects a decoded image larger than 18 MB or 50 megapixels.

<Warning>
  Loops doesn't fetch external image URLs. Pass the bytes inline in `ImageChunk`. The `ImageAssetPointerChunk` type exists for Tinker compatibility, but its constructor raises `ValueError`.
</Warning>

`ImageChunk.expected_tokens` tells the trainer how many input tokens the image becomes. The model replaces the image with exactly that many image tokens, so the number must match what the model actually produces. For Qwen VL models, compute it from the image processor as `prod(image_grid_thw) // merge_size**2`.

Add these functions to build a `Datum` from a dataset example:

```python train_vision.py theme={"system"} theme={"system"}
def compute_expected_tokens(image: Image.Image) -> int:
    result = image_processor(images=[image])
    grid = result["image_grid_thw"][0]
    merge = image_processor.merge_size
    return int(grid[0] * grid[1] * grid[2]) // (merge * merge)


def to_image_chunk(image: Image.Image, expected_tokens: int) -> ImageChunk:
    buf = io.BytesIO()
    image.convert("RGB").save(buf, format="JPEG")
    return ImageChunk(
        data=buf.getvalue(),
        format="jpeg",
        expected_tokens=expected_tokens,
    )


def to_datum(example, tokenizer):
    image = example["images"][0]
    problem = example["problem"]
    answer = example["answer"]

    expected_tokens = compute_expected_tokens(image)
    img_chunk = to_image_chunk(image, expected_tokens)

    # geometry3k wraps the prompt in an <image> placeholder; strip it.
    prompt = problem.replace("<image>", "")
    p = tokenizer.encode(prompt, add_special_tokens=False)
    a = tokenizer.encode(answer, add_special_tokens=False)

    # The trainer doesn't shift labels: drop the last input token and make
    # each target the token that follows its position.
    text_chunk = ModelInput.from_ints(p + a[:-1]).chunks[0]
    targets = [-100] * (expected_tokens + len(p) - 1) + list(a)

    return Datum(
        model_input=ModelInput(chunks=[img_chunk, text_chunk]),
        loss_fn_inputs={
            "target_tokens": TensorData(
                data=targets, dtype="int64", shape=[len(targets)]
            )
        },
    )
```

The trainer expands each `ImageChunk` into `expected_tokens` copies of the model's image-placeholder token ID and appends the text chunk. Because `to_datum` omits the final answer token from the input, the input sequence is `[image: expected_tokens] [prompt: len(p)] [answer: len(a) - 1]`.

### Set the loss targets

`forward_backward` compares each input position against the token that should follow it. The trainer doesn't shift labels, so `to_datum` drops the final answer token from the input and assigns next-token targets to the remaining positions. The resulting input and `target_tokens` have the same length: `expected_tokens + len(p) + len(a) - 1`.

Mask the image and prompt positions with `-100`, which cross-entropy ignores, except for the last prompt position. That position targets the first answer token. The remaining positions target the subsequent answer tokens, including the final answer token:

```python theme={"system"}
targets = [-100] * (expected_tokens + len(p) - 1) + list(a)
```

## Add the training loop

At each training step, [`forward_backward()`](/reference/sdk/loops/training-client) runs the model on a batch and accumulates gradients. Then [`optim_step()`](/reference/sdk/loops/training-client) applies them with [`AdamParams`](/reference/sdk/loops/types). Add the loop to `train_vision.py`:

```python train_vision.py theme={"system"} theme={"system"}
BATCH_SIZE = 4

dataset = load_dataset("hiyouga/geometry3k", split="train[:4]")
data = [to_datum(ex, tokenizer) for ex in dataset]
print(f"prepared {len(data)} examples")

for step, start in enumerate(range(0, len(data), BATCH_SIZE), 1):
    batch = data[start : start + BATCH_SIZE]
    fb = training_client.forward_backward(data=batch).result(timeout=600.0)
    training_client.optim_step(
        AdamParams(learning_rate=4e-5)
    ).result(timeout=600.0)
    print(f"step {step} loss={fb.loss:.4f}")
```

The `train[:4]` slice limits this tutorial to four examples. For a complete fine-tune, use the full split. Image examples use more memory than text-only examples because the model processes image tensors and text together. Start with a small batch size, and increase it if GPU memory allows.

## Save a checkpoint

Save the tuned adapter as a named checkpoint with [`save_weights_for_sampler()`](/reference/sdk/loops/training-client). The checkpoint stays available after the training session ends. Add the save call to `train_vision.py`:

```python train_vision.py theme={"system"} theme={"system"}
save_resp = training_client.save_weights_for_sampler(name="vision-step-1").result(timeout=600.0)
print(f"saved checkpoint at {save_resp.path}")
```

## Run the training script

Run the script:

```bash theme={"system"} theme={"system"}
uv run python train_vision.py
```

The first run provisions the trainer before the first step starts, which takes a few minutes. Subsequent runs against the same base model can reuse the trainer by setting `LOOPS_REUSE_FROM_RUN_ID`.

The script prints the number of prepared examples, the training loss, and the checkpoint path:

```output theme={"system"} theme={"system"}
prepared 4 examples
step 1 loss=5.4734
saved checkpoint at bt://loops:e3m1863/sampler_weights/vision-step-1
```

## Stop the training resources

The trainer continues to bill after the script exits. The checkpoint path includes the run ID after `bt://loops:`. Replace `<run_id>` with your run ID, not the example ID `e3m1863`. List your active Loops runs and deactivate the run you created:

```bash theme={"system"} theme={"system"}
baseten loops run list
baseten loops run deactivate --run-id "<run_id>" --yes
```

## Deploy the checkpoint

Deploy the checkpoint to [Dedicated Inference](/development/model/overview). The deployment loads your LoRA adapter on top of the base weights and serves an OpenAI-compatible `/v1/chat/completions` endpoint for image inputs. Before you deploy, add an `hf_access_token` to [workspace secrets](/organization/secrets). The deployment uses this token to download the base weights from Hugging Face.

List the checkpoints for your run, replacing `<run_id>` with the same run ID:

```bash theme={"system"} theme={"system"}
baseten loops checkpoint list --run-id "<run_id>"
```

```output theme={"system"} theme={"system"}
                          Checkpoints for run: e3m1863
╭─────────┬──────────┬─────────┬─────────┬──────┬─────────┬─────────╮
│ ID      │ Name     │ Run     │ Target  │ Type │ Size    │ Created │
├─────────┼──────────┼─────────┼─────────┼──────┼─────────┼─────────┤
│ MPXxYJP │ vision…  │ e3m1863 │ sampler │ lora │ 20.6 MB  │ 2026-… │
╰─────────┴──────────┴─────────┴─────────┴──────┴─────────┴─────────╯
```

Replace `<checkpoint_id>` with the ID of your `vision-step-1` checkpoint from the list output. The deploy command asks for a model name, GPU type, GPU count, and Hugging Face secret name. The default secret name is `hf_access_token`.

```bash theme={"system"} theme={"system"}
baseten loops checkpoint deploy --checkpoint-id "<checkpoint_id>"
```

```output theme={"system"} theme={"system"}
Successfully created deployment: deployment-1
Model ID: q86o6283
Deployment ID: qvxp6gg
Deployment succeeded.
```

Use the **Model ID** and **Deployment ID** from your deploy output for `<model_id>` and `<deployment_id>` in all remaining commands. Check the deployment status:

```bash theme={"system"} theme={"system"}
baseten model deployment describe --model-id "<model_id>" --deployment-id "<deployment_id>" --jq '.status'
```

```output theme={"system"} theme={"system"}
"SCALED_TO_ZERO"
```

The deployment is ready when its status is `ACTIVE` or `SCALED_TO_ZERO`. If the status is `SCALED_TO_ZERO`, the first request starts a replica and takes longer than later requests.

## Send an image prompt

Send the image as base64-encoded data in the OpenAI `image_url` format.

The example sends this geometry diagram with the text prompt "Find x.":

<img src="https://mintcdn.com/baseten-preview/tKZwC5srcS2lPFRS/images/geometry3k-example.png?fit=max&auto=format&n=tKZwC5srcS2lPFRS&q=85&s=280ac7c8cfda87eedf13e2b4062a7621" alt="Geometry diagram from the geometry3k dataset showing a circle with two intersecting chords labeled 4, 8, 6, and x" width="250" data-path="images/geometry3k-example.png" />

Create `build_request.py` to encode the image and build the request:

```python build_request.py theme={"system"} theme={"system"}
import base64
import io
import json

from datasets import load_dataset
from PIL import Image

ds = load_dataset("hiyouga/geometry3k", split="train[:1]")
img = ds[0]["images"][0]
buf = io.BytesIO()
img.convert("RGB").save(buf, format="JPEG")
b64 = base64.b64encode(buf.getvalue()).decode()
problem = ds[0]["problem"].replace("<image>", "")

payload = {
    "model": "vision-step-1",
    "messages": [
        {
            "role": "user",
            "content": [
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/jpeg;base64,{b64}"},
                },
                {"type": "text", "text": problem},
            ],
        }
    ],
    "max_tokens": 64,
    "chat_template_kwargs": {"enable_thinking": False},
}

with open("request.json", "w") as f:
    json.dump(payload, f)
```

Run the script. It writes `request.json` with the base64-encoded image and the prompt:

```bash theme={"system"} theme={"system"}
uv run python build_request.py
```

In the payload, `model` is the checkpoint name. Send the request:

<Tabs>
  <Tab title="cURL">
    ```bash theme={"system"} theme={"system"}
    curl -X POST "https://model-<model_id>.api.baseten.co/deployment/<deployment_id>/sync/v1/chat/completions" \
      -H "Authorization: Bearer $BASETEN_API_KEY" \
      -H "Content-Type: application/json" \
      -d @request.json
    ```
  </Tab>

  <Tab title="Baseten CLI">
    ```bash theme={"system"} theme={"system"}
    baseten model predict --model-id "<model_id>" --deployment-id "<deployment_id>" --file request.json
    ```
  </Tab>
</Tabs>

```json theme={"system"} theme={"system"}
{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "We are given a circle with two chords intersecting inside it, labeled as circles with radii from the center to the endpoints.\n\nThe diagram shows:\n\n- A central vertical chord.\n- Two intersecting chords on the left side.\n- The labels are:\n  - 4 and 6 are"
      },
      "finish_reason": "length"
    }
  ],
  "usage": {"prompt_tokens": 81, "completion_tokens": 64, "total_tokens": 145}
}
```

One training step on four examples has little effect on the base model. Use more examples and training steps to improve accuracy on the geometry task.

## Send a text-only prompt

The same deployment accepts text-only requests. Omit the image content block:

<Tabs>
  <Tab title="cURL">
    ```bash theme={"system"} theme={"system"}
    curl -X POST "https://model-<model_id>.api.baseten.co/deployment/<deployment_id>/sync/v1/chat/completions" \
      -H "Authorization: Bearer $BASETEN_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model":"vision-step-1","messages":[{"role":"user","content":"What is the capital of France?"}],"max_tokens":32}'
    ```
  </Tab>

  <Tab title="Baseten CLI">
    ```bash theme={"system"} theme={"system"}
    baseten model predict --model-id "<model_id>" --deployment-id "<deployment_id>" \
      --data '{"model":"vision-step-1","messages":[{"role":"user","content":"What is the capital of France?"}],"max_tokens":32}'
    ```
  </Tab>
</Tabs>

## Deactivate the deployment

The deployment bills for its GPU while it's active. Deactivate it when you finish testing:

```bash theme={"system"} theme={"system"}
baseten model deployment deactivate --model-id "<model_id>" --deployment-id "<deployment_id>"
```

For scripts, pass `--yes` to skip the confirmation prompt. Deactivation doesn't delete the checkpoint, so you can deploy it again.

## Next steps

* **[Deploy a checkpoint](/loops/deploy-checkpoints)**: The full deploy flow, including the dashboard option and the `--config` flag for non-interactive deploys.
* **[Train on your data](/loops/train-on-your-data)**: The text-only SFT guide. Vision training follows the same round trip.
* **[Loops concepts](/loops/concepts)**: Sessions, trainers, samplers, checkpoints, and how weight sync works.
* **[Migrate from Tinker](/loops/tinker-compatibility)**: If you're porting a Tinker vision recipe, the `ImageChunk` type is call-compatible. Loops doesn't support `ImageAssetPointerChunk`.
