Qwen/Qwen3.5-2B to label support tickets as access or billing issues. It saves a report of the model’s answers so your application can find mistakes and choose data for the next round.
Before you begin
You need Loops access for your workspace, a Baseten API key, and uv. The script uses Python 3.12.uv installs the dependencies listed in the script, including baseten-loops==0.23.2.
Use the exact model ID from the supported models list. Check the list for the Base or Instruct version you need. The script also checks which models your workspace supports before creating a trainer.
You pay for the GPUs used by the trainer and sampler. The script requests deactivation after each round, including if training or evaluation fails, and records whether cleanup succeeded. Start with one round on a small model before running multiple rounds at once.
Create a working directory and set your API key. These commands use macOS or Linux shell syntax:
LOOPS_REUSE_FROM_RUN_ID, LOOPS_REUSE_FROM_SESSION_ID, LOOPS_BASE_URL, and LOOPS_TEAM. Unset these variables if they’re present in your environment.
Prepare training and evaluation data
Createtrain.jsonl with prompts and expected answers for training:
train.jsonl
eval.jsonl with different prompts to evaluate the model:
eval.jsonl
Build a training round
These snippets show the training and evaluation code insiderun_round(). The complete script after the steps includes the imports, data-loading helpers, and cleanup needed to run it.
1
Create a trainer
Create a session with
ServiceClient, then call create_lora_training_client() to start training from the selected base_model. The script reads key from BASETEN_API_KEY:rank controls the size of the LoRA adapter; this example uses 16. max_seq_len configures a 1,024-token sequence limit, and the script checks each example against that limit. ready_timeout sets how long the client waits for the trainer or sampler to become ready, in seconds. sampler_timeout limits how long it waits for a sampling request.2
Measure evaluation loss
Convert the training and evaluation rows into token inputs with the script’s
to_datum() helper. Call forward() on the evaluation data to measure loss before training. This call doesn’t update the model:to_datum() pairs each input token with the next token to predict. It uses -100 to exclude prompt tokens from the loss, so training focuses on the answer. Each Datum contains the input tokens as a ModelInput and the targets as TensorData..result(timeout=600) waits up to 600 seconds for the operation to finish. The report stores the result as baseline_loss for comparison after training.3
Train on batches and measure loss again
Call
forward_backward() to compute gradients for a batch, then optim_step() to update the model. AdamParams sets the optimizer’s learning_rate. Train on batches of up to eight examples, then measure loss on the same evaluation data:epochs controls how many times the model trains on the full dataset. The script defaults to two. Compare trained_loss with baseline_loss to check how training changed performance on these examples.4
Save a training checkpoint
Call The report stores the checkpoint path in
save_state() to save the model weights and optimizer state so you can resume training. name identifies the checkpoint within the run:trainer_checkpoint.5
Prepare the trained weights for sampling
Call The first
save_weights_for_sampler() to save weights for generating answers. Pass the returned path to create_sampling_client() to select those weights:create_sampling_client() call creates the run’s sampler. The sampler_checkpoint path also lets you deploy the saved weights for inference.6
Generate and score answers
Call The report records each generated answer and whether it matches the expected label. Your application can use incorrect answers to choose training data for the next round.
sample() for each evaluation prompt. num_samples=1 requests one answer. SamplingParams limits the answer to 16 tokens and stops at a newline:Save the complete script
Expand the code block, copy the script, and save it ascampaign.py. It combines the preceding steps with input validation, report saving, and run deactivation.
Complete campaign.py script
Complete campaign.py script
campaign.py
Run the script and check results
Run the script with the example data:round-1.json to check the results:
If
cleanup_status isn’t completed, use List runs to find runs with the saved session_id, then deactivate each run. Network failures or a terminated process can prevent cleanup from finishing. Check this before retrying so you don’t leave GPUs running.
The example data tests that the code works; it’s too small to measure model quality. The script scores an answer as correct when it matches the expected label, ignoring case and surrounding whitespace. Use evaluation criteria that fit your task.
Train with updated data
Your application can importrun_round() from campaign.py and use its returned report to find incorrect answers. Use those mistakes to guide your data selection, then start another round. Keep the evaluation file the same so you can compare results.
Save the updated training data as train-v2.jsonl and choose a new report path:
create_training_client_from_state_with_optimizer() with trainer_checkpoint. Record which starting point you use when comparing rounds.
Set limits in your application for spending, total rounds, and rounds running at once. Stop when the model meets your evaluation target or reaches a limit. Save each round’s results before starting the next one. Before retrying a failed worker, check its saved session and run IDs to avoid starting the same round twice.
Next steps
- Train on your data covers chat-format data and training in batches.
- Loops SDK reference documents the classes and parameters used by the script.
- Deploy a checkpoint serves saved sampler weights for inference.