> ## Documentation Index
> Fetch the complete documentation index at: https://docs.baseten.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Programmatic training

> Use the Loops Python SDK to train a model, evaluate its answers, and save results for the next training round.

Use the [Loops SDK](/reference/sdk/loops/overview) to train a model from Python, evaluate its answers, and save checkpoints. Your application chooses the training data and decides when to run another round.

This example trains `Qwen/Qwen3.5-2B` to label support tickets as access or billing issues. It saves a report of the model's answers so your application can find mistakes and choose data for the next round.

## Before you begin

You need [Loops access](https://www.baseten.co/talk-to-us/loops-signup/) for your workspace, a [Baseten API key](/organization/api-keys), and [uv](https://docs.astral.sh/uv/getting-started/installation/). The script uses Python 3.12. `uv` installs the dependencies listed in the script, including `baseten-loops==0.23.2`.

Use the exact model ID from the [supported models](/loops/supported-models) list. Check the list for the Base or Instruct version you need. The script also checks which models your workspace supports before creating a trainer.

You pay for the GPUs used by the trainer and sampler. The script requests deactivation after each round, including if training or evaluation fails, and records whether cleanup succeeded. Start with one round on a small model before running multiple rounds at once.

Create a working directory and set your API key. These commands use macOS or Linux shell syntax:

```bash theme={"system"}
mkdir loops-campaign
cd loops-campaign
export BASETEN_API_KEY="YOUR_API_KEY"
```

This example creates independent runs using the default API endpoint and your API key's workspace. It rejects the SDK overrides `LOOPS_REUSE_FROM_RUN_ID`, `LOOPS_REUSE_FROM_SESSION_ID`, `LOOPS_BASE_URL`, and `LOOPS_TEAM`. Unset these variables if they're present in your environment.

## Prepare training and evaluation data

Create `train.jsonl` with prompts and expected answers for training:

```jsonl train.jsonl theme={"system"}
{"prompt":"Classify the support ticket as access or billing. Respond with only the label.\nTicket: I cannot sign in to my account.\nLabel:","answer":" access"}
{"prompt":"Classify the support ticket as access or billing. Respond with only the label.\nTicket: My card was charged twice.\nLabel:","answer":" billing"}
{"prompt":"Classify the support ticket as access or billing. Respond with only the label.\nTicket: Please reset my account password.\nLabel:","answer":" access"}
{"prompt":"Classify the support ticket as access or billing. Respond with only the label.\nTicket: The invoice shows the wrong price.\nLabel:","answer":" billing"}
```

Create `eval.jsonl` with different prompts to evaluate the model:

```jsonl eval.jsonl theme={"system"}
{"prompt":"Classify the support ticket as access or billing. Respond with only the label.\nTicket: I forgot my login credentials.\nLabel:","answer":" access"}
{"prompt":"Classify the support ticket as access or billing. Respond with only the label.\nTicket: I need a refund for a duplicate payment.\nLabel:","answer":" billing"}
```

Use the same evaluation examples in each round so you can compare results. Keep a separate test set for the final assessment; don't use it for training or choosing new training data. The script rejects identical prompts in the training and evaluation files. Also check for examples that repeat the same question in different words.

## Build a training round

These snippets show the training and evaluation code inside `run_round()`. The complete script after the steps includes the imports, data-loading helpers, and cleanup needed to run it.

<Steps>
  <Step title="Create a trainer">
    Create a session with [`ServiceClient`](/reference/sdk/loops/service-client), then call [`create_lora_training_client()`](/reference/sdk/loops/service-client) to start training from the selected `base_model`. The script reads `key` from `BASETEN_API_KEY`:

    ```python theme={"system"}
    service = ServiceClient(api_key=key)
    trainer = service.create_lora_training_client(
        base_model=base_model, rank=16, seed=42, max_seq_len=MAX_TOKENS,
        ready_timeout=1200, sampler_timeout=120, name=output_path.stem,
    )
    ```

    `rank` controls the size of the LoRA adapter; this example uses 16. `max_seq_len` configures a 1,024-token sequence limit, and the script checks each example against that limit. `ready_timeout` sets how long the client waits for the trainer or sampler to become ready, in seconds. `sampler_timeout` limits how long it waits for a sampling request.
  </Step>

  <Step title="Measure evaluation loss">
    Convert the training and evaluation rows into token inputs with the script's `to_datum()` helper. Call [`forward()`](/reference/sdk/loops/training-client) on the evaluation data to measure loss before training. This call doesn't update the model:

    ```python theme={"system"}
    tokenizer = trainer.get_tokenizer()
    train_data = [to_datum(tokenizer, row) for row in train_rows]
    eval_data = [to_datum(tokenizer, row) for row in eval_rows]
    record["baseline_loss"] = trainer.forward(data=eval_data).result(timeout=600).loss
    ```

    `to_datum()` pairs each input token with the next token to predict. It uses `-100` to exclude prompt tokens from the loss, so training focuses on the answer. Each [`Datum`](/reference/sdk/loops/types) contains the input tokens as a [`ModelInput`](/reference/sdk/loops/types) and the targets as [`TensorData`](/reference/sdk/loops/types).

    `.result(timeout=600)` waits up to 600 seconds for the operation to finish. The report stores the result as `baseline_loss` for comparison after training.
  </Step>

  <Step title="Train on batches and measure loss again">
    Call [`forward_backward()`](/reference/sdk/loops/training-client) to compute gradients for a batch, then [`optim_step()`](/reference/sdk/loops/training-client) to update the model. [`AdamParams`](/reference/sdk/loops/types) sets the optimizer's `learning_rate`. Train on batches of up to eight examples, then measure loss on the same evaluation data:

    ```python theme={"system"}
    for _ in range(epochs):
        for start in range(0, len(train_data), 8):
            trainer.forward_backward(data=train_data[start:start + 8]).result(timeout=600)
            trainer.optim_step(AdamParams(learning_rate=4e-5)).result(timeout=600)
    record["trained_loss"] = trainer.forward(data=eval_data).result(timeout=600).loss
    ```

    `epochs` controls how many times the model trains on the full dataset. The script defaults to two. Compare `trained_loss` with `baseline_loss` to check how training changed performance on these examples.
  </Step>

  <Step title="Save a training checkpoint">
    Call [`save_state()`](/reference/sdk/loops/training-client) to save the model weights and optimizer state so you can resume training. `name` identifies the checkpoint within the run:

    ```python theme={"system"}
    record["trainer_checkpoint"] = trainer.save_state(name="final").result(timeout=600).path
    ```

    The report stores the checkpoint path in `trainer_checkpoint`.
  </Step>

  <Step title="Prepare the trained weights for sampling">
    Call [`save_weights_for_sampler()`](/reference/sdk/loops/training-client) to save weights for generating answers. Pass the returned path to [`create_sampling_client()`](/reference/sdk/loops/training-client) to select those weights:

    ```python theme={"system"}
    weights = trainer.save_weights_for_sampler(name="final").result(timeout=600)
    record["sampler_checkpoint"] = weights.path
    sampler = trainer.create_sampling_client(model_path=weights.path)
    ```

    The first `create_sampling_client()` call creates the run's sampler. The `sampler_checkpoint` path also lets you deploy the saved weights for inference.
  </Step>

  <Step title="Generate and score answers">
    Call [`sample()`](/reference/sdk/loops/sampling-client) for each evaluation prompt. `num_samples=1` requests one answer. [`SamplingParams`](/reference/sdk/loops/types) limits the answer to 16 tokens and stops at a newline:

    ```python theme={"system"}
    record["evaluations"] = []
    for row in eval_rows:
        sample = sampler.sample(
            prompt=ModelInput.from_ints(tokenizer.encode(row["prompt"], add_special_tokens=False)),
            num_samples=1,
            sampling_params=SamplingParams(max_tokens=16, temperature=0, stop=["\n"]),
        )
        completion = tokenizer.decode(sample.sequences[0].tokens).strip()
        record["evaluations"].append({
            "prompt": row["prompt"], "expected": row["answer"].strip(),
            "completion": completion,
            "match": completion.casefold() == row["answer"].strip().casefold(),
        })
    ```

    The report records each generated answer and whether it matches the expected label. Your application can use incorrect answers to choose training data for the next round.
  </Step>
</Steps>

## Save the complete script

Expand the code block, copy the script, and save it as `campaign.py`. It combines the preceding steps with input validation, report saving, and run deactivation.

<Accordion title="Complete campaign.py script">
  ```python campaign.py theme={"system"}
  # /// script
  # requires-python = ">=3.12"
  # dependencies = ["baseten-loops==0.23.2", "requests==2.34.2"]
  # ///

  """Runs one train/evaluate round and releases its Loops compute."""

  import argparse
  import hashlib
  import json
  import os
  from contextlib import ExitStack
  from pathlib import Path

  import requests
  from baseten.loops import (
      AdamParams, Datum, ModelInput, SamplingParams, ServiceClient, TensorData,
  )

  BASE_MODEL = "Qwen/Qwen3.5-2B"
  API_URL = "https://api.baseten.co/v1/loops"
  MAX_TOKENS = 1024


  def read_examples(path):
      rows = [json.loads(line) for line in Path(path).read_text().splitlines() if line.strip()]
      if not rows:
          raise ValueError(f"{path} contains no examples")
      for row in rows:
          if not isinstance(row, dict) or any(
              not isinstance(row.get(key), str) or not row[key].strip()
              for key in ("prompt", "answer")
          ):
              raise ValueError("Each example needs a nonempty prompt and answer")
      return rows


  def to_datum(tokenizer, row):
      prompt = tokenizer.encode(row["prompt"], add_special_tokens=False)
      answer = tokenizer.encode(row["answer"], add_special_tokens=False)
      if not prompt or not answer or len(prompt) + len(answer) > MAX_TOKENS:
          raise ValueError("Use nonempty prompts and answers totaling at most 1024 tokens")
      tokens = (prompt + answer)[:-1]
      targets = [-100] * (len(prompt) - 1) + answer
      return Datum(
          model_input=ModelInput.from_ints(tokens),
          loss_fn_inputs={
              "target_tokens": TensorData(data=targets, dtype="int64", shape=[len(targets)])
          },
      )


  def api(method, path, key, **kwargs):
      response = requests.request(
          method, f"{API_URL}{path}",
          headers={"Authorization": f"Bearer {key}"}, timeout=30, **kwargs,
      )
      response.raise_for_status()
      return response.json()


  def release_session(session_id, key, record):
      # Also find runs created before a provisioning timeout returned a client.
      record["deactivated_run_ids"] = []
      try:
          runs = api("GET", "/runs", key)["runs"]
          owned = [run["id"] for run in runs if run["session_id"] == session_id]
          for run_id in owned:
              api("POST", f"/runs/{run_id}/deactivate", key)
              record["deactivated_run_ids"].append(run_id)
          record["cleanup_status"] = "completed"
      except BaseException as error:
          record["cleanup_status"] = "failed"
          record["cleanup_error_type"] = type(error).__name__
          raise


  def write_report(path, record):
      temporary = path.with_suffix(".tmp")
      temporary.write_text(json.dumps(record, indent=2))
      temporary.replace(path)


  def run_round(train_path, eval_path, output_path, *, epochs=2, base_model=BASE_MODEL):
      overrides = [name for name in (
          "LOOPS_REUSE_FROM_RUN_ID", "LOOPS_REUSE_FROM_SESSION_ID",
          "LOOPS_BASE_URL", "LOOPS_TEAM",
      ) if name in os.environ]
      if overrides:
          raise ValueError(f"Unset SDK overrides before running this example: {', '.join(overrides)}")
      train_rows, eval_rows = read_examples(train_path), read_examples(eval_path)
      if {row["prompt"] for row in train_rows} & {row["prompt"] for row in eval_rows}:
          raise ValueError("Keep evaluation prompts out of the training data")
      if epochs < 1:
          raise ValueError("epochs must be positive")
      output_path = Path(output_path)
      if output_path.exists():
          raise FileExistsError(f"Choose a new report path: {output_path}")
      output_path.parent.mkdir(parents=True, exist_ok=True)
      key = os.environ["BASETEN_API_KEY"]
      record = {
          "base_model": base_model, "epochs": epochs, "rank": 16, "seed": 42,
          "learning_rate": 4e-5, "status": "starting", "cleanup_status": "pending",
          "train_sha256": hashlib.sha256(Path(train_path).read_bytes()).hexdigest(),
          "eval_sha256": hashlib.sha256(Path(eval_path).read_bytes()).hexdigest(),
      }
      service = ServiceClient(api_key=key)
      record["session_id"] = service.session_id
      write_report(output_path, record)
      print(f"session_id={service.session_id}", flush=True)

      with ExitStack() as cleanup:
          cleanup.callback(write_report, output_path, record)
          cleanup.callback(release_session, service.session_id, key, record)
          try:
              supported = {model.model_name for model in service.get_server_capabilities().supported_models}
              if base_model not in supported:
                  raise ValueError(f"Model isn't available through Loops: {base_model}")
              trainer = service.create_lora_training_client(
                  base_model=base_model, rank=16, seed=42, max_seq_len=MAX_TOKENS,
                  ready_timeout=1200, sampler_timeout=120, name=output_path.stem,
              )
              cleanup.callback(trainer.close)
              record["run_id"] = trainer.run_id
              record["status"] = "training"
              write_report(output_path, record)
              print(f"run_id={trainer.run_id}", flush=True)
              tokenizer = trainer.get_tokenizer()
              train_data = [to_datum(tokenizer, row) for row in train_rows]
              eval_data = [to_datum(tokenizer, row) for row in eval_rows]
              record["baseline_loss"] = trainer.forward(data=eval_data).result(timeout=600).loss
              for _ in range(epochs):
                  for start in range(0, len(train_data), 8):
                      trainer.forward_backward(data=train_data[start:start + 8]).result(timeout=600)
                      trainer.optim_step(AdamParams(learning_rate=4e-5)).result(timeout=600)
              record["trained_loss"] = trainer.forward(data=eval_data).result(timeout=600).loss
              record["trainer_checkpoint"] = trainer.save_state(name="final").result(timeout=600).path
              weights = trainer.save_weights_for_sampler(name="final").result(timeout=600)
              record["sampler_checkpoint"] = weights.path
              write_report(output_path, record)

              sampler = trainer.create_sampling_client(model_path=weights.path)
              cleanup.callback(sampler.close)
              record["evaluations"] = []
              for row in eval_rows:
                  sample = sampler.sample(
                      prompt=ModelInput.from_ints(tokenizer.encode(row["prompt"], add_special_tokens=False)),
                      num_samples=1,
                      sampling_params=SamplingParams(max_tokens=16, temperature=0, stop=["\n"]),
                  )
                  completion = tokenizer.decode(sample.sequences[0].tokens).strip()
                  record["evaluations"].append({
                      "prompt": row["prompt"], "expected": row["answer"].strip(),
                      "completion": completion,
                      "match": completion.casefold() == row["answer"].strip().casefold(),
                  })
              record["status"] = "completed"
          except BaseException as error:
              record["status"] = "failed"
              record["error_type"] = type(error).__name__
              raise
      print(f"saved report to {output_path}", flush=True)
      return record


  if __name__ == "__main__":
      parser = argparse.ArgumentParser(description=__doc__)
      parser.add_argument("--train", default="train.jsonl")
      parser.add_argument("--eval", default="eval.jsonl")
      parser.add_argument("--output", default="round-1.json")
      args = parser.parse_args()
      run_round(args.train, args.eval, args.output)
  ```
</Accordion>

The script saves the session ID before creating the trainer and the run ID when the trainer is ready. During cleanup, it closes the SDK clients and calls [Deactivate a run](/reference/loops-api/runs/deactivate-a-run) for runs in that session.

## Run the script and check results

Run the script with the example data:

```bash theme={"system"}
uv run --python 3.12 campaign.py
```

A successful run prints its session ID, run ID, and report path. The IDs vary by run:

```text theme={"system"}
session_id=<SESSION_ID>
run_id=<RUN_ID>
saved report to round-1.json
```

Open `round-1.json` to check the results:

| Report field | How to use it |
| - | - |
| `session_id`, `run_id` | Find the run in the dashboard or [Loops API](/reference/loops-api/runs/get-a-run). |
| `train_sha256`, `eval_sha256` | Identify the exact training and evaluation files used. |
| `baseline_loss`, `trained_loss` | Compare the same evaluation examples before and after training. |
| `evaluations` | Compare each generated answer with the expected answer. |
| `trainer_checkpoint` | Resume training with weights and optimizer state. |
| `sampler_checkpoint` | Generate outputs from the saved weights or deploy them for inference. |
| `status`, `error_type` | Check whether training and evaluation completed or which error stopped them. |
| `cleanup_status`, `cleanup_error_type` | Check whether cleanup completed or which error stopped it. |
| `deactivated_run_ids` | Find the runs whose deactivation requests succeeded. |

If `cleanup_status` isn't `completed`, use [List runs](/reference/loops-api/runs/list-runs) to find runs with the saved `session_id`, then [deactivate each run](/reference/loops-api/runs/deactivate-a-run). Network failures or a terminated process can prevent cleanup from finishing. Check this before retrying so you don't leave GPUs running.

The example data tests that the code works; it's too small to measure model quality. The script scores an answer as correct when it matches the expected label, ignoring case and surrounding whitespace. Use evaluation criteria that fit your task.

## Train with updated data

Your application can import `run_round()` from `campaign.py` and use its returned report to find incorrect answers. Use those mistakes to guide your data selection, then start another round. Keep the evaluation file the same so you can compare results.

Save the updated training data as `train-v2.jsonl` and choose a new report path:

```bash theme={"system"}
uv run --python 3.12 campaign.py \
  --train train-v2.jsonl \
  --eval eval.jsonl \
  --output round-2.json
```

Each call starts a new LoRA training run from the selected base model. To continue training from a saved checkpoint, use [`create_training_client_from_state_with_optimizer()`](/reference/sdk/loops/service-client) with `trainer_checkpoint`. Record which starting point you use when comparing rounds.

Set limits in your application for spending, total rounds, and rounds running at once. Stop when the model meets your evaluation target or reaches a limit. Save each round's results before starting the next one. Before retrying a failed worker, check its saved session and run IDs to avoid starting the same round twice.

## Next steps

* [Train on your data](/loops/train-on-your-data) covers chat-format data and training in batches.
* [Loops SDK reference](/reference/sdk/loops/overview) documents the classes and parameters used by the script.
* [Deploy a checkpoint](/loops/deploy-checkpoints) serves saved sampler weights for inference.
