The Training Loop: Gradient Descent in Practice
Forward, zero the grads, backward, step, repeat. Four lines that train every network in this course, and the two settings that decide whether they work: how big a step to take, and how much data to look at per step.
Inspect Architecture: Linear regression1. The idea
Chapter 4 took one step downhill. Chapter 5 built an engine that finds the gradient for any formula, and Chapter 6 made it fast on whole arrays. Training is what you get when you put that step in a loop: compute the loss, reset the gradients, run backward(), move every knob a little against its gradient, and repeat. Those four lines are the inner loop of every neural network, from the ride model to the capstone's language model.
Start the ride model at w = 0, b = 0, where its loss is 777.5, and let it run. With a learning rate of η = 0.01 it gets within 1% of the best possible line, w* = 3.041 and b* = 4.973, after 721 steps. It finds the answer without ever being told it. But the learning rate matters enormously. With η = 0.001 the loss is still 1.44 after 2,000 steps. With η = 0.016 it takes 450 steps. With η = 0.017, only slightly larger, every step overshoots further than the last, and after 2,000 steps the knobs are around −10³⁵.
That cliff edge isn't luck. Near its minimum, the loss is a bowl, and the bowl's steepness in each direction (its curvature) sets the largest step that doesn't overshoot: η must stay below 2 divided by the largest curvature, here 2 / 120.04 = 0.0167. The ride bowl is also lopsided, about 263 times steeper in one direction than the other, because km values run from 2 to 12 while b's input is always 1. A step small enough for the steep direction crawls along the flat one. Standardizing the input, subtracting its mean and dividing by its spread, makes the bowl perfectly round, and then one step with η = 0.5 lands exactly on the minimum.
The last setting is how much data to use per step. Real datasets don't fit into one gradient, so the loop takes mini-batches: a few examples at a time, in shuffled order. On a year of rides (240 of them), ten passes with the full batch leave the loss at 124, while batches of 16 reach the minimum within four passes. Batches of one get close even faster, but never settle down, and seeing why tells you a lot about how networks are trained.
2. The math
repeat: L = loss(θ) (forward)
θ.grad = 0 (zero the grads)
L.backward() (backward)
θ ← θ − η · θ.grad (step)w* = cov(x, y) / var(x) b* = ȳ − w*·x̄
rides: w* = 3.0411 b* = 4.9726 L* = 0.6644 from (0, 0): L = 777.5η = 0.001: L = 1.441, still far from the best
η = 0.01: within 1% of L* after 721 steps
η = 0.016: within 1% after 450 steps
η = 0.017: diverges (knobs around −10³⁵)L(θ) ≈ L* + ½ (θ − θ*)ᵀ H (θ − θ*) H = matrix of second derivatives
rides: H = (2/N) [ Σx² Σx ; Σx N ] = [ 118.5 13.5 ; 13.5 2 ]
eigenvalues of H: 120.04 and 0.456along an axis with curvature λ, each step multiplies
the distance to θ* by (1 − η·λ)
stable ⇔ |1 − η·λ| < 1 for every λ ⇔ η < 2 / λ_max = 2 / 120.04 = 0.0167κ = λ_max / λ_min = 120.04 / 0.456 = 263
with η = 0.016: steep axis × (1 − 0.016·120.04) = × (−0.92)
flat axis × (1 − 0.016·0.456) = × 0.9927x′ = (x − x̄) / s rides: x̄ = 6.75, s = 3.70
H′ = (2/N) [ Σx′² Σx′ ; Σx′ N ] = [ 2 0 ; 0 2 ] ⇒ κ = 1
η = 0.5: exact minimum in 1 step η = 0.1 or 0.9: within 1% in 27 steps∇L_B = (1/|B|) Σ_(i ∈ B) ∇ℓᵢ (on average, equal to the full gradient)
one epoch = one shuffled pass = N / |B| updates
240 rides, |B| = 16: 15 updates per epochL* = 1.996, starting loss 1008.3
|B| = 240: 10 updates, loss 124.3 after 10 epochs
|B| = 16: 150 updates, 1.999 after 4 epochs, then within 0.003 of L*
|B| = 1: 2,400 updates, 2.11 after 1 epoch, then bouncing between 2.05 and 2.52at θ*: full gradient ∇L = 0, but single-ride gradients ∇ℓᵢ ≠ 0
spread of ∇L_B ∝ 1 / √|B| (for batches drawn at random)• Take the bowl exactly: L(θ) = L* + ½ (θ − θ*)ᵀ H (θ − θ*). Its gradient is ∇L = H (θ − θ*), which is zero at θ* as it should be.
• One step of gradient descent gives θ_new − θ* = (θ − θ*) − η·H (θ − θ*) = (I − ηH)(θ − θ*). So the distance to the minimum is multiplied by the matrix (I − ηH) every step.
• H is symmetric, so it has perpendicular eigenvectors with eigenvalues λ₁, λ₂, … Along eigenvector k, (I − ηH) just multiplies by (1 − η·λₖ). The directions don't mix: each one shrinks or grows on its own.
• After t steps the component along direction k is (1 − η·λₖ)ᵗ times its starting value. That goes to 0 exactly when |1 − η·λₖ| < 1, meaning 0 < η < 2/λₖ. Every direction must shrink, so η < 2/λ_max.
• For the rides, 2/λ_max = 2/120.04 = 0.01666. At η = 0.016 the steep direction's factor is 1 − 1.92 = −0.92: it overshoots each time but shrinks. At η = 0.017 it's 1 − 2.04 = −1.04, and the error grows by 4% per step, which over 2,000 steps is a factor of about 10³⁴. ∎
• For the line ŷ = w·x + b with MSE, the second derivatives are ∂²L/∂w² = (2/N) Σ xᵢ², ∂²L/∂w∂b = (2/N) Σ xᵢ and ∂²L/∂b² = (2/N)·N = 2. Each is a constant, which is why the bowl is exact.
• For the rides, Σx² = 4 + 25 + 64 + 144 = 237 and Σx = 27, so H = [118.5, 13.5; 13.5, 2]. The off-diagonal 13.5 tilts the bowl, and the 118.5 against 2 stretches it.
• After standardizing, x′ has mean 0 and variance 1: Σ x′ᵢ = 0 and Σ x′ᵢ² = N. Then H′ = (2/N)[N, 0; 0, N] = [2, 0; 0, 2] = 2I. Both curvatures equal 2, so κ = 1: a perfectly round bowl.
• With H = 2I the step is θ_new − θ* = (I − η·2I)(θ − θ*) = (1 − 2η)(θ − θ*). Choosing η = 0.5 makes the factor 0, so one step lands exactly on θ*. η = 0.1 and η = 0.9 give factors 0.8 and −0.8, the same size, which is why both need 27 steps.
• Real networks aren't quadratic and their H changes from place to place, so no single step size is perfect. But the lesson carries over: inputs on similar scales give rounder bowls and faster training. Chapter 12 does the same thing inside the network. ∎
3. How it works
4. The code (python)
# Chapter 7: the training loop. Forward, zero the grads, backward, step, and repeat.
# Uses the Tensor engine from Chapter 6: save that chapter's code down to (not including) the line
# '# === 1.' as tensor.py next to this file. That's unbroadcast(), the Tensor class and count().
import numpy as np
from tensor import Tensor
km = np.array([2.0, 5.0, 8.0, 12.0])
minutes = np.array([11.0, 21.0, 28.0, 42.0])
def train(x, y, lr, steps, batch_size=None, seed=0):
"""Gradient descent on the line w*x + b. Returns the final knobs and the full-data loss after every update."""
w, b = Tensor(0.0), Tensor(0.0) # start knowing nothing
params = [w, b]
rng = np.random.default_rng(seed)
n = len(x)
batch_size = batch_size or n # default: the whole dataset every step
history = [float(np.mean((w.data * x + b.data - y) ** 2))]
updates = 0
while updates < steps:
order = rng.permutation(n) # one epoch = one pass through the data, shuffled
for start in range(0, n, batch_size):
rows = order[start:start + batch_size]
loss = ((w * x[rows] + b - y[rows]) ** 2).mean() # 1. forward
for p in params: # 2. zero the grads
p.grad = np.zeros_like(p.data)
loss.backward() # 3. backward
for p in params: # 4. step downhill
p.data = p.data - lr * p.grad
history.append(float(np.mean((w.data * x + b.data - y) ** 2)))
updates += 1
if updates == steps or not np.isfinite(history[-1]):
return float(w.data), float(b.data), history
return float(w.data), float(b.data), history
# === 1. The best line, in closed form, to know what we're aiming for ==============
w_best = np.mean((km - km.mean()) * (minutes - minutes.mean())) / np.var(km)
b_best = minutes.mean() - w_best * km.mean()
loss_best = np.mean((w_best * km + b_best - minutes) ** 2)
print(f"best line: w* = {w_best:.4f}, b* = {b_best:.4f}, loss = {loss_best:.4f}")
# === 2. Four learning rates, 2,000 steps each ======================================
print("\n lr w b loss after 2,000 first step within 1% of best")
for lr in [0.001, 0.01, 0.016, 0.017]:
w, b, hist = train(km, minutes, lr, 2000)
near = next((i for i, v in enumerate(hist) if v < 1.01 * loss_best), None)
print(f"{lr:6} {w:8.4g} {b:8.4g} {hist[-1]:12.4g} {near}")
# === 3. Why 0.017 explodes: the curvature of the bowl ===============================
X = np.column_stack([km, np.ones(4)])
H = 2 / len(km) * X.T @ X # second derivatives of the MSE in (w, b)
lam = np.linalg.eigvalsh(H)
print(f"\ncurvatures (eigenvalues of H): {lam.round(3)}, condition number {lam.max() / lam.min():.0f}")
print(f"largest stable learning rate: 2 / {lam.max():.2f} = {2 / lam.max():.4f}")
# === 4. Standardize the input: the bowl becomes round ================================
mu, sd = km.mean(), km.std()
km_std = (km - mu) / sd
for lr in [0.1, 0.5, 0.9]:
w_s, b_s, hist = train(km_std, minutes, lr, 30)
near = next((i for i, v in enumerate(hist) if v < 1.01 * loss_best), None)
print(f"standardized, lr {lr}: within 1% after {near} steps;"
f" back in km: w = {w_s / sd:.4f}, b = {b_s - w_s * mu / sd:.4f}")
# === 5. More data: a year of rides, and mini-batches ==================================
rng = np.random.default_rng(7)
year_km = np.round(rng.uniform(1, 15, 240), 1)
year_min = np.round(3 * year_km + 5 + rng.normal(0, 1.5, 240), 1) # 3 min/km + 5, plus noise
year_std = (year_km - year_km.mean()) / year_km.std()
Xy = np.column_stack([year_std, np.ones(240)])
best = np.linalg.lstsq(Xy, year_min, rcond=None)[0]
print(f"\nyear of rides: best loss {np.mean((Xy @ best - year_min) ** 2):.3f}, starting loss {np.mean(year_min ** 2):.3f}")
print("batch updates loss after each of 10 epochs")
for B in [240, 16, 1]:
steps = 10 * (240 // B) # 10 epochs, whatever the batch size
w_s, b_s, hist = train(year_std, year_min, 0.05, steps, batch_size=B)
per_epoch = hist[240 // B::240 // B]
print(f"{B:5} {steps:7} {[round(v, 3) for v in per_epoch]}")
w_s, b_s, _ = train(year_std, year_min, 0.05, 150, batch_size=16)
print(f"B = 16, back in km: w = {w_s / year_km.std():.3f}, b = {b_s - w_s * year_km.mean() / year_km.std():.3f}")
5. Practice
Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.
6. Go further
- Find the cliff edge yourself: bisect between η = 0.016 (converges) and η = 0.017 (diverges), running 2,000 steps each time, until the interval is narrower than 0.0001. How close do you get to 2 / λ_max = 0.01666?
- Calm down batch size 1: rerun the year-of-rides experiment with B = 1 and a shrinking learning rate, η_t = 0.05 / (1 + t / 500) where t counts updates. Does the loss now settle near 1.996? What happens if it shrinks too fast, say η_t = 0.05 / (1 + t)?
- Train the picnic neuron from Chapter 6 with this loop: full batch, η = 1.0, 500 steps, printing the loss every 100 steps. The picnic inputs already lie between 0 and 1. Does standardizing them still change how fast it trains?
7. Check yourself
In the catalog
Finished this chapter?