Skip to main content
Fine-tune a vision-language model on examples that pair an image with a text prompt and an answer. The trainer renders the image as visual tokens, concatenates them with the prompt and answer tokens, and computes loss over the combined sequence in the same forward_backward() and optim_step() round trip as text-only supervised fine-tuning (SFT). The training loop, checkpointing, and deployment stay the same; only the input construction changes. This guide fine-tunes with LoRA on hiyouga/geometry3k, a dataset of geometry diagrams with text problems and numeric answers, then deploys the checkpoint and queries it with image prompts. The data-building and loss-masking techniques work with any image-and-text dataset. Vision training supports a subset of the supported base models: Qwen VL (Qwen3.5, Qwen3.6, Qwen3.8) and Kimi (K2.6, K2.7 Code).

Prerequisites

1

Export your API key

Export the key as BASETEN_API_KEY:
2

Install the Baseten CLI

The baseten loops and baseten model commands later in this guide need the CLI:
Terminal
For other platforms or a specific version, see the Baseten CLI install reference.
Loops is in early access. To enable it for your workspace, fill out the signup form.

Install

Create a uv project and install the dependencies:

Create the training script

The script opens a training run on a base model, converts dataset examples into training data, and trains a LoRA adapter before saving a checkpoint. Create a file named train_vision.py in the project directory:
train_vision.py
ServiceClient opens a session with Baseten. create_lora_training_client provisions a trainer for the base model and attaches a LoRA adapter with the rank you set. get_tokenizer returns the base model’s tokenizer, and the image processor parses image dimensions later in the script:
train_vision.py

Turn examples into training data

Each example in the geometry3k dataset contains an image, a problem, and an answer. The trainer uses each example as a Datum:
  • model_input contains one ImageChunk with the image bytes and one text chunk with the problem and answer tokens.
  • loss_fn_inputs supplies the loss targets through a TensorData named target_tokens, with one target token per input position.
ImageChunk takes the image bytes inline, as JPEG or PNG. One chunk holds one image; add more chunks for more images. Pass the image already resized: the trainer doesn’t resize images and rejects a decoded image larger than 18 MB or 50 megapixels.
Loops doesn’t fetch external image URLs. Pass the bytes inline in ImageChunk. The ImageAssetPointerChunk type exists for Tinker compatibility, but its constructor raises ValueError.
ImageChunk.expected_tokens tells the trainer how many input tokens the image becomes. The model replaces the image with exactly that many image tokens, so the number must match what the model actually produces. For Qwen VL models, compute it from the image processor as prod(image_grid_thw) // merge_size**2. Add these functions to build a Datum from a dataset example:
train_vision.py
The trainer expands each ImageChunk into expected_tokens copies of the model’s image-placeholder token ID and appends the text chunk. Because to_datum omits the final answer token from the input, the input sequence is [image: expected_tokens] [prompt: len(p)] [answer: len(a) - 1].

Set the loss targets

forward_backward compares each input position against the token that should follow it. The trainer doesn’t shift labels, so to_datum drops the final answer token from the input and assigns next-token targets to the remaining positions. The resulting input and target_tokens have the same length: expected_tokens + len(p) + len(a) - 1. Mask the image and prompt positions with -100, which cross-entropy ignores, except for the last prompt position. That position targets the first answer token. The remaining positions target the subsequent answer tokens, including the final answer token:

Add the training loop

At each training step, forward_backward() runs the model on a batch and accumulates gradients. Then optim_step() applies them with AdamParams. Add the loop to train_vision.py:
train_vision.py
The train[:4] slice limits this tutorial to four examples. For a complete fine-tune, use the full split. Image examples use more memory than text-only examples because the model processes image tensors and text together. Start with a small batch size, and increase it if GPU memory allows.

Save a checkpoint

Save the tuned adapter as a named checkpoint with save_weights_for_sampler(). The checkpoint stays available after the training session ends. Add the save call to train_vision.py:
train_vision.py

Run the training script

Run the script:
The first run provisions the trainer before the first step starts, which takes a few minutes. Subsequent runs against the same base model can reuse the trainer by setting LOOPS_REUSE_FROM_RUN_ID. The script prints the number of prepared examples, the training loss, and the checkpoint path:

Stop the training resources

The trainer continues to bill after the script exits. The checkpoint path includes the run ID after bt://loops:. Replace <run_id> with your run ID, not the example ID e3m1863. List your active Loops runs and deactivate the run you created:

Deploy the checkpoint

Deploy the checkpoint to Dedicated Inference. The deployment loads your LoRA adapter on top of the base weights and serves an OpenAI-compatible /v1/chat/completions endpoint for image inputs. Before you deploy, add an hf_access_token to workspace secrets. The deployment uses this token to download the base weights from Hugging Face. List the checkpoints for your run, replacing <run_id> with the same run ID:
Replace <checkpoint_id> with the ID of your vision-step-1 checkpoint from the list output. The deploy command asks for a model name, GPU type, GPU count, and Hugging Face secret name. The default secret name is hf_access_token.
Use the Model ID and Deployment ID from your deploy output for <model_id> and <deployment_id> in all remaining commands. Check the deployment status:
The deployment is ready when its status is ACTIVE or SCALED_TO_ZERO. If the status is SCALED_TO_ZERO, the first request starts a replica and takes longer than later requests.

Send an image prompt

Send the image as base64-encoded data in the OpenAI image_url format. The example sends this geometry diagram with the text prompt “Find x.”: Geometry diagram from the geometry3k dataset showing a circle with two intersecting chords labeled 4, 8, 6, and x Create build_request.py to encode the image and build the request:
build_request.py
Run the script. It writes request.json with the base64-encoded image and the prompt:
In the payload, model is the checkpoint name. Send the request:
One training step on four examples has little effect on the base model. Use more examples and training steps to improve accuracy on the geometry task.

Send a text-only prompt

The same deployment accepts text-only requests. Omit the image content block:

Deactivate the deployment

The deployment bills for its GPU while it’s active. Deactivate it when you finish testing:
For scripts, pass --yes to skip the confirmation prompt. Deactivation doesn’t delete the checkpoint, so you can deploy it again.

Next steps

  • Deploy a checkpoint: The full deploy flow, including the dashboard option and the --config flag for non-interactive deploys.
  • Train on your data: The text-only SFT guide. Vision training follows the same round trip.
  • Loops concepts: Sessions, trainers, samplers, checkpoints, and how weight sync works.
  • Migrate from Tinker: If you’re porting a Tinker vision recipe, the ImageChunk type is call-compatible. Loops doesn’t support ImageAssetPointerChunk.