Skip to main content
Use the Loops SDK to train a model from Python, evaluate its answers, and save checkpoints. Your application chooses the training data and decides when to run another round. This example trains Qwen/Qwen3.5-2B to label support tickets as access or billing issues. It saves a report of the model’s answers so your application can find mistakes and choose data for the next round.

Before you begin

You need Loops access for your workspace, a Baseten API key, and uv. The script uses Python 3.12. uv installs the dependencies listed in the script, including baseten-loops==0.23.2. Use the exact model ID from the supported models list. Check the list for the Base or Instruct version you need. The script also checks which models your workspace supports before creating a trainer. You pay for the GPUs used by the trainer and sampler. The script requests deactivation after each round, including if training or evaluation fails, and records whether cleanup succeeded. Start with one round on a small model before running multiple rounds at once. Create a working directory and set your API key. These commands use macOS or Linux shell syntax:
This example creates independent runs using the default API endpoint and your API key’s workspace. It rejects the SDK overrides LOOPS_REUSE_FROM_RUN_ID, LOOPS_REUSE_FROM_SESSION_ID, LOOPS_BASE_URL, and LOOPS_TEAM. Unset these variables if they’re present in your environment.

Prepare training and evaluation data

Create train.jsonl with prompts and expected answers for training:
train.jsonl
Create eval.jsonl with different prompts to evaluate the model:
eval.jsonl
Use the same evaluation examples in each round so you can compare results. Keep a separate test set for the final assessment; don’t use it for training or choosing new training data. The script rejects identical prompts in the training and evaluation files. Also check for examples that repeat the same question in different words.

Build a training round

These snippets show the training and evaluation code inside run_round(). The complete script after the steps includes the imports, data-loading helpers, and cleanup needed to run it.
1

Create a trainer

Create a session with ServiceClient, then call create_lora_training_client() to start training from the selected base_model. The script reads key from BASETEN_API_KEY:
rank controls the size of the LoRA adapter; this example uses 16. max_seq_len configures a 1,024-token sequence limit, and the script checks each example against that limit. ready_timeout sets how long the client waits for the trainer or sampler to become ready, in seconds. sampler_timeout limits how long it waits for a sampling request.
2

Measure evaluation loss

Convert the training and evaluation rows into token inputs with the script’s to_datum() helper. Call forward() on the evaluation data to measure loss before training. This call doesn’t update the model:
to_datum() pairs each input token with the next token to predict. It uses -100 to exclude prompt tokens from the loss, so training focuses on the answer. Each Datum contains the input tokens as a ModelInput and the targets as TensorData..result(timeout=600) waits up to 600 seconds for the operation to finish. The report stores the result as baseline_loss for comparison after training.
3

Train on batches and measure loss again

Call forward_backward() to compute gradients for a batch, then optim_step() to update the model. AdamParams sets the optimizer’s learning_rate. Train on batches of up to eight examples, then measure loss on the same evaluation data:
epochs controls how many times the model trains on the full dataset. The script defaults to two. Compare trained_loss with baseline_loss to check how training changed performance on these examples.
4

Save a training checkpoint

Call save_state() to save the model weights and optimizer state so you can resume training. name identifies the checkpoint within the run:
The report stores the checkpoint path in trainer_checkpoint.
5

Prepare the trained weights for sampling

Call save_weights_for_sampler() to save weights for generating answers. Pass the returned path to create_sampling_client() to select those weights:
The first create_sampling_client() call creates the run’s sampler. The sampler_checkpoint path also lets you deploy the saved weights for inference.
6

Generate and score answers

Call sample() for each evaluation prompt. num_samples=1 requests one answer. SamplingParams limits the answer to 16 tokens and stops at a newline:
The report records each generated answer and whether it matches the expected label. Your application can use incorrect answers to choose training data for the next round.

Save the complete script

Expand the code block, copy the script, and save it as campaign.py. It combines the preceding steps with input validation, report saving, and run deactivation.
campaign.py
The script saves the session ID before creating the trainer and the run ID when the trainer is ready. During cleanup, it closes the SDK clients and calls Deactivate a run for runs in that session.

Run the script and check results

Run the script with the example data:
A successful run prints its session ID, run ID, and report path. The IDs vary by run:
Open round-1.json to check the results: If cleanup_status isn’t completed, use List runs to find runs with the saved session_id, then deactivate each run. Network failures or a terminated process can prevent cleanup from finishing. Check this before retrying so you don’t leave GPUs running. The example data tests that the code works; it’s too small to measure model quality. The script scores an answer as correct when it matches the expected label, ignoring case and surrounding whitespace. Use evaluation criteria that fit your task.

Train with updated data

Your application can import run_round() from campaign.py and use its returned report to find incorrect answers. Use those mistakes to guide your data selection, then start another round. Keep the evaluation file the same so you can compare results. Save the updated training data as train-v2.jsonl and choose a new report path:
Each call starts a new LoRA training run from the selected base model. To continue training from a saved checkpoint, use create_training_client_from_state_with_optimizer() with trainer_checkpoint. Record which starting point you use when comparing rounds. Set limits in your application for spending, total rounds, and rounds running at once. Stop when the model meets your evaluation target or reaches a limit. Save each round’s results before starting the next one. Before retrying a failed worker, check its saved session and run IDs to avoid starting the same round twice.

Next steps