Skip to content

Project: controlled training comparison

Run two learning rates fairly and inspect the validation loss, accuracy, macro F1 and per-class recall after each epoch.

What to remember

A comparison is useful only when the split, starting model and training budget agree. This script creates one initial state and restores it for each trial. Each trial also gets the same seeded batch order.

It connects Course 2 metrics, schedules, search discipline and accumulation. The comparison is deliberately a two-value search using ordinary PyTorch; Optuna is useful for larger searches, but is not needed to understand the objective.

Important code

model.load_state_dict(initial)
loader = DataLoader(train, batch_size=16, shuffle=True,
                    generator=torch.Generator().manual_seed(23))
optimizer = torch.optim.SGD(model.parameters(), lr=lr, weight_decay=1e-3)
scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(
    optimizer, mode="min", factor=0.5, patience=1)

# Repeat for each epoch:
train_accumulated(model, loader, optimizer, "cpu", microbatches=3)
measured = metrics(model, validation)
scheduler.step(measured["loss"])

The loss controls the plateau scheduler and selects the final comparison winner. Accuracy and macro F1 describe different aspects of the same predictions; they are not averaged together into an invented score.

Setting This example Why it matters
Learning rate 0.01 and 0.1 Changes the size of SGD updates; compare the full histories.
Microbatch / accumulation 16 × 3 Up to 48 samples per optimizer update on one device.
Training samples 133 The last update has 37 samples; normalization uses that actual count.
Plateau patience / factor 1 / 0.5 After the allowed non-improving checks, reduce LR by half. Eight short epochs need not trigger a reduction.
Metric aggregation One confusion matrix for validation Avoids giving a short batch the same weight as a full one.

The model has no BatchNorm or Dropout, which makes the accumulation mechanism easier to inspect. The helper averages gradients by the actual group sample count.

Check the effective batch

Samples per optimizer update
Single-device arithmetic; a final incomplete group is smaller.

Run and inspect

python examples/training_comparison.py
python examples/training_comparison.py --epochs 20 --output artifacts/my-comparison

The offline CPU run saves comparison.json under the selected output directory. For each trial, inspect lr_used versus lr_next, loss, macro F1, per-class recall and the confusion matrix. Rows are true classes; columns are predictions.

This seeded synthetic problem demonstrates controlled comparison. Its scores are not a benchmark for a real application, and the selected LR may change with a different dataset or budget. Keep a separate test split for final assessment in a real experiment.

Try changing: compare a wider LR range, then change only accumulation steps. Does a different update count alter learning even though the number of epochs stayed fixed?

Review metrics and tuning · Review accumulation

Retained CPU execution

These are measured outputs of the small synthetic run on 2026-10-01: PyTorch 2.14.0+cu130, seed 17, two CPU threads, 133 training and 48 validation samples, eight epochs. They demonstrate the mechanism, not real-world classifier quality.

Validation loss for the two controlled learning rates

Validation macro F1 for the same trials

Actual validation confusion matrix for the selected LR

The lower final validation loss selects initial LR 0.1: loss 0.9382, accuracy 66.7%, macro F1 0.5486. The matrix includes all 48 validation predictions; rows are truth and columns predictions. Both trials keep their initial LR during this short run: the plateau condition is not reached.

Class 2 is recognized in only 1 of 7 validation examples (14.3% recall), despite 66.7% overall accuracy. This is why the per-class matrix and recall belong beside the winning scalar score.

The retained comparison report contains every epoch and count.

Optional physical batch versus accumulation benchmark

python examples/training_comparison.py --benchmark --device cpu
python examples/training_comparison.py --benchmark --device auto
# Optional on compatible CUDA hardware; also measures accumulated AMP FP16:
python examples/training_comparison.py --benchmark --device cuda

CPU is the default. Every variant starts from the same parameters and sample order: effective batch 48, three updates, final group 37. One warmup epoch is discarded; the next five complete epochs are timed. The timed region includes iteration, transfers, forward/backward and updates; model construction is outside it.

Measured CPU variant Median epoch Throughput CUDA memory
physical FP32 1.06 ms 125,059 samples/s N/A on CPU
accumulated FP32 1.90 ms 69,930 samples/s N/A on CPU

Maximum FP32 parameter difference after the final incomplete group: 1.49e-08. This tiny CPU experiment is sensitive to overhead; it does not establish a speed advantage on other models or devices. CUDA/AMP was unavailable in this session. On CUDA, benchmark.json also records peak allocated CUDA memory (including resident model state), free/total CUDA memory before trials, GPU name, optimizer/model configuration, precision, timing samples and throughput. Retained benchmark configuration and measurements.

Complete source

Open the maintained runnable script
"""Compare two learning rates fairly; original Course 2 recall experiment."""
from __future__ import annotations

import argparse
import copy
import json
import time
import statistics
from pathlib import Path

import torch
from torch import nn
from torch.utils.data import DataLoader, TensorDataset

from recall_patterns import train_accumulated


def device_for(name):
    if name == "auto":
        name = "cuda" if torch.cuda.is_available() else "cpu"
    if name == "cuda" and not torch.cuda.is_available():
        raise ValueError("CUDA unavailable; choose --device cpu")
    return torch.device(name)


def metrics(model, loader):
    model.eval()
    matrix = torch.zeros(3, 3, dtype=torch.long)
    total_loss = 0.0
    with torch.no_grad():
        for x, y in loader:
            device = next(model.parameters(), torch.empty(0)).device
            logits = model(x.to(device))
            total_loss += nn.functional.cross_entropy(logits, y.to(device), reduction="sum").item()
            predictions = logits.argmax(1).cpu()
            matrix += torch.bincount(y * 3 + predictions, minlength=9).reshape(3, 3)
    tp = matrix.diag().float()
    precision = tp / matrix.sum(0).clamp_min(1)
    recall = tp / matrix.sum(1).clamp_min(1)
    f1 = 2 * precision * recall / (precision + recall).clamp_min(1e-12)
    count = matrix.sum().item()
    return {
        "loss": total_loss / count if count else None,
        "accuracy": tp.sum().item() / count if count else None,
        "macro_f1": f1.mean().item() if count else None,
        "recall_by_class": recall.tolist() if count else [None] * 3,
        "confusion_matrix": matrix.tolist(),
    }


def benchmark(template, train, device, repetitions=5):
    """Tiny end-to-end training epoch; fixed order, effective batch 48, three updates."""
    if repetitions < 1 or not len(train):
        raise ValueError("Benchmark requires samples and positive repetitions")
    variants = [("physical FP32", 48, 1, False), ("accumulated FP32", 16, 3, False)]
    if device.type == "cuda":
        variants.append(("accumulated AMP FP16", 16, 3, True))
    cuda_memory = None
    if device.type == "cuda":
        free, total = torch.cuda.mem_get_info(device)
        cuda_memory = {"free_mib_before_trials": free / 1024**2, "total_mib": total / 1024**2,
                       "device_name": torch.cuda.get_device_name(device)}
    results = []
    for name, size, accumulation, amp in variants:
        elapsed, peaks, states = [], [], []
        for repetition in range(repetitions + 1):
            model = copy.deepcopy(template).to(device).train()
            optimizer = torch.optim.SGD(model.parameters(), lr=0.1)
            scaler = torch.amp.GradScaler("cuda", enabled=amp)
            loader = DataLoader(train, batch_size=size, shuffle=False)
            if device.type == "cuda":
                torch.cuda.synchronize(device)
                torch.cuda.reset_peak_memory_stats(device)
            start = time.perf_counter()
            optimizer.zero_grad(set_to_none=True)
            group = 0
            for step, (x, y) in enumerate(loader, 1):
                x, y = x.to(device), y.to(device)
                with torch.autocast(device_type=device.type, dtype=torch.float16, enabled=amp):
                    loss = nn.functional.cross_entropy(model(x), y, reduction="sum")
                scaler.scale(loss).backward()
                group += len(y)
                if step % accumulation == 0 or step == len(loader):
                    scaler.unscale_(optimizer)
                    for p in model.parameters():
                        p.grad.div_(group)
                    scaler.step(optimizer)
                    scaler.update()
                    optimizer.zero_grad(set_to_none=True)
                    group = 0
            if device.type == "cuda":
                torch.cuda.synchronize(device)
            seconds = time.perf_counter() - start
            if repetition:  # discard the first complete epoch as warmup
                elapsed.append(seconds)
                if device.type == "cuda":
                    peaks.append(torch.cuda.max_memory_allocated(device) / 1024**2)
            states.append({k: v.detach().cpu() for k, v in model.state_dict().items()})
        median = statistics.median(elapsed)
        results.append({"name": name, "microbatch": size, "accumulation": accumulation,
                        "effective_batch": 48, "samples": len(train), "updates": (len(train) + 47) // 48,
                        "median_seconds": median, "seconds": elapsed,
                        "samples_per_second": len(train) / median,
                        "peak_cuda_allocated_mib": max(peaks) if peaks else None})
        if len(results) == 1:
            reference = states[-1]
        results[-1]["max_parameter_difference_from_physical_fp32"] = max(
            (states[-1][k] - reference[k]).abs().max().item() for k in reference)
    return {"device": str(device), "torch": str(torch.__version__), "threads": torch.get_num_threads(),
            "scope": "tiny training epoch including loading and transfer; not a hardware ranking",
            "warmup_epochs": 1, "repetitions": repetitions, "variants": results,
            "cuda_memory": cuda_memory,
            "optimizer": {"name": "SGD", "learning_rate": 0.1}, "model": str(template),
            "amp": "measured" if device.type == "cuda" else "skipped: CUDA required"}


def run(output: Path, epochs=8, device="cpu", do_benchmark=False):
    torch.set_num_threads(2)
    torch.manual_seed(17)
    x = torch.randn(181, 6)
    # Imbalanced synthetic labels; class metrics show more than accuracy alone.
    y = torch.where(x[:, 0] > 0.8, 2, torch.where(x[:, 1] > 0.0, 1, 0))
    train = TensorDataset(x[:133], y[:133])
    validation = DataLoader(TensorDataset(x[133:], y[133:]), batch_size=17)
    device = device_for(device)
    template = nn.Sequential(nn.Linear(6, 12), nn.ReLU(), nn.Linear(12, 3)).to(device)
    initial = copy.deepcopy(template.state_dict())
    trials = []
    for lr in (0.01, 0.1):
        model = copy.deepcopy(template)
        model.load_state_dict(initial)
        # Recreate the generator: each trial sees the same shuffled batch order.
        loader = DataLoader(train, batch_size=16, shuffle=True,
                            generator=torch.Generator().manual_seed(23))
        optimizer = torch.optim.SGD(model.parameters(), lr=lr, weight_decay=1e-3)
        scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(
            optimizer, mode="min", factor=0.5, patience=1)
        history = []
        for epoch in range(epochs):
            used_lr = optimizer.param_groups[0]["lr"]
            # 3 x 16 = 48 samples/full update; final group uses its actual count.
            train_accumulated(model, loader, optimizer, device, microbatches=3)
            measured = metrics(model, validation)
            scheduler.step(measured["loss"])
            history.append({"epoch": epoch + 1, "lr_used": used_lr,
                            "lr_next": optimizer.param_groups[0]["lr"], **measured})
        trials.append({"initial_lr": lr, "history": history})
    output.mkdir(parents=True, exist_ok=True)
    report = {
        "data": "seeded synthetic classification; not a real-world benchmark",
        "selection": "lowest final validation loss; no test set used for tuning",
        "train_samples": len(train), "validation_samples": 48,
        "microbatch": 16, "accumulation_steps": 3, "seed": 17,
        "device": str(device), "torch": str(torch.__version__),
        "best_initial_lr": min(trials, key=lambda t: t["history"][-1]["loss"])["initial_lr"],
        "trials": trials,
    }
    (output / "comparison.json").write_text(json.dumps(report, indent=2), encoding="utf-8")
    if do_benchmark:
        (output / "benchmark.json").write_text(json.dumps(benchmark(template, train, device), indent=2), encoding="utf-8")
    for trial in trials:
        last = trial["history"][-1]
        print(f"initial_lr={trial['initial_lr']}: loss={last['loss']:.3f}, "
              f"accuracy={last['accuracy']:.3f}, macro_f1={last['macro_f1']:.3f}")
    return report


if __name__ == "__main__":
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("--output", type=Path, default=Path("artifacts/training-comparison"))
    parser.add_argument("--epochs", type=int, default=8)
    parser.add_argument("--device", choices=("cpu", "auto", "cuda"), default="cpu")
    parser.add_argument("--benchmark", action="store_true", help="Measure a warmed-up physical/accumulated comparison")
    args = parser.parse_args()
    if args.epochs < 1:
        parser.error("--epochs must be positive")
    run(args.output, args.epochs, args.device, args.benchmark)

The shared train_accumulated helper comes from examples/recall_patterns.py:

Open the accumulation helper
def train_accumulated(model, loader, optimizer, device, microbatches=4):
    """Single device; unweighted classification; average by actual group size."""
    if microbatches < 1:
        raise ValueError("microbatches must be positive")
    model.train()
    optimizer.zero_grad(set_to_none=True)
    loss_fn = torch.nn.CrossEntropyLoss(reduction="sum")
    group_samples = 0
    for step, (x, y) in enumerate(loader, start=1):
        x, y = x.to(device), y.to(device)
        loss_fn(model(x), y).backward()  # add gradients; no update yet
        group_samples += y.numel()
        if step % microbatches == 0 or step == len(loader):
            for parameter in model.parameters():
                if parameter.grad is not None:
                    parameter.grad.div_(group_samples)
            optimizer.step()
            optimizer.zero_grad(set_to_none=True)
            group_samples = 0

API reference: ReduceLROnPlateau.