Learn machine learning from scratch

Start at chapter 1 and build up: what a model is, how it measures its own mistakes, and how it learns. Everything is worked out in plain Python, with the maths shown rather than assumed. Two shorter tracks hold reference notes on the maths and on modern architectures.

Track Progress: 0 Completed
0%
Core CourseLEVEL 2 · INTERMEDIATEChapter 7

The Training Loop: Gradient Descent in Practice

Forward, zero the grads, backward, step, repeat. Four lines that train every network in this course, and the two settings that decide whether they work: how big a step to take, and how much data to look at per step.

38 min Prerequisites: Chapters 4–6 (the Tensor class from Chapter 6, saved as tensor.py)
Inspect Architecture: Linear regression

1. The idea

Chapter 4 took one step downhill. Chapter 5 built an engine that finds the gradient for any formula, and Chapter 6 made it fast on whole arrays. Training is what you get when you put that step in a loop: compute the loss, reset the gradients, run backward(), move every knob a little against its gradient, and repeat. Those four lines are the inner loop of every neural network, from the ride model to the capstone's language model.

Start the ride model at w = 0, b = 0, where its loss is 777.5, and let it run. With a learning rate of η = 0.01 it gets within 1% of the best possible line, w* = 3.041 and b* = 4.973, after 721 steps. It finds the answer without ever being told it. But the learning rate matters enormously. With η = 0.001 the loss is still 1.44 after 2,000 steps. With η = 0.016 it takes 450 steps. With η = 0.017, only slightly larger, every step overshoots further than the last, and after 2,000 steps the knobs are around −10³⁵.

That cliff edge isn't luck. Near its minimum, the loss is a bowl, and the bowl's steepness in each direction (its curvature) sets the largest step that doesn't overshoot: η must stay below 2 divided by the largest curvature, here 2 / 120.04 = 0.0167. The ride bowl is also lopsided, about 263 times steeper in one direction than the other, because km values run from 2 to 12 while b's input is always 1. A step small enough for the steep direction crawls along the flat one. Standardizing the input, subtracting its mean and dividing by its spread, makes the bowl perfectly round, and then one step with η = 0.5 lands exactly on the minimum.

The last setting is how much data to use per step. Real datasets don't fit into one gradient, so the loop takes mini-batches: a few examples at a time, in shuffled order. On a year of rides (240 of them), ten passes with the full batch leave the loss at 124, while batches of 16 reach the minimum within four passes. Batches of one get close even faster, but never settle down, and seeing why tells you a lot about how networks are trained.

Two charts with logarithmic loss axes. Left, learning rate: the four-ride loss over 1,000 steps from w = b = 0. With η = 0.001 the loss falls quickly to about 5 and then creeps down slowly. With η = 0.01 and η = 0.016 it reaches the minimum, 0.664, within a few hundred steps, faster for 0.016. With η = 0.017 it shoots off the top of the chart. Right, batch size: the loss on 240 rides after each of 10 epochs with η = 0.05. The full batch (B = 240) falls slowly, from about 1,000 to 124. B = 16 reaches the minimum, about 2, by epoch 4. B = 1 drops to about 2 after one epoch and then keeps bouncing between 2.05 and 2.52.Two charts with logarithmic loss axes. Left, learning rate: the four-ride loss over 1,000 steps from w = b = 0. With η = 0.001 the loss falls quickly to about 5 and then creeps down slowly. With η = 0.01 and η = 0.016 it reaches the minimum, 0.664, within a few hundred steps, faster for 0.016. With η = 0.017 it shoots off the top of the chart. Right, batch size: the loss on 240 rides after each of 10 epochs with η = 0.05. The full batch (B = 240) falls slowly, from about 1,000 to 124. B = 16 reaches the minimum, about 2, by epoch 4. B = 1 drops to about 2 after one epoch and then keeps bouncing between 2.05 and 2.52.
The two settings of the training loop. Too small a step crawls and too large a step blows up; smaller batches make more progress per pass through the data, but the smallest never quite settles.

2. The math

The training loop
repeat: L = loss(θ) (forward) θ.grad = 0 (zero the grads) L.backward() (backward) θ ← θ − η · θ.grad (step)
θ stands for all the knobs and η is the learning rate. Zeroing matters because grads accumulate, as Chapter 5 showed. Forget it and each step uses the sum of all past gradients.
Where it should end up: the best line in closed form
w* = cov(x, y) / var(x) b* = ȳ − w*·x̄ rides: w* = 3.0411 b* = 4.9726 L* = 0.6644 from (0, 0): L = 777.5
For a straight line and squared error there's a formula for the answer, so we can check the loop against it. For neural networks there isn't, and the loop is the only way.
Four learning rates, 2,000 steps from (0, 0)
η = 0.001: L = 1.441, still far from the best η = 0.01: within 1% of L* after 721 steps η = 0.016: within 1% after 450 steps η = 0.017: diverges (knobs around −10³⁵)
A bigger step is faster, right up to the point where it stops working altogether. The next two boxes explain where that point is.
Near the minimum, every loss is a bowl
L(θ) ≈ L* + ½ (θ − θ*)ᵀ H (θ − θ*) H = matrix of second derivatives rides: H = (2/N) [ Σx² Σx ; Σx N ] = [ 118.5 13.5 ; 13.5 2 ] eigenvalues of H: 120.04 and 0.456
The eigenvalues of H are the bowl's curvatures along its two main axes. For the line with MSE the bowl is exact, not an approximation, because the loss is a quadratic in (w, b).
The largest stable learning rate
along an axis with curvature λ, each step multiplies the distance to θ* by (1 − η·λ) stable ⇔ |1 − η·λ| < 1 for every λ ⇔ η < 2 / λ_max = 2 / 120.04 = 0.0167
Between 1/λ and 2/λ the factor is negative: each step overshoots to the other side, but by less than before. Past 2/λ it overshoots by more each time, and the loss grows without limit. Derived in the first derivation.
The condition number: why the safe step crawls
κ = λ_max / λ_min = 120.04 / 0.456 = 263 with η = 0.016: steep axis × (1 − 0.016·120.04) = × (−0.92) flat axis × (1 − 0.016·0.456) = × 0.9927
The step has to be small enough for the steepest direction, which leaves the flattest one shrinking by under 1% per step. A large κ means a slow, ill-conditioned problem.
Standardizing makes the bowl round
x′ = (x − x̄) / s rides: x̄ = 6.75, s = 3.70 H′ = (2/N) [ Σx′² Σx′ ; Σx′ N ] = [ 2 0 ; 0 2 ] ⇒ κ = 1 η = 0.5: exact minimum in 1 step η = 0.1 or 0.9: within 1% in 27 steps
With mean 0 and variance 1, Σx′ = 0 and Σx′² = N, so both curvatures are 2. Convert back with w = w′/s and b = b′ − w′·x̄/s, which gives w* = 3.041 and b* = 4.973 again.
Mini-batches and epochs
∇L_B = (1/|B|) Σ_(i ∈ B) ∇ℓᵢ (on average, equal to the full gradient) one epoch = one shuffled pass = N / |B| updates 240 rides, |B| = 16: 15 updates per epoch
A random batch's gradient is a noisy but unbiased estimate of the full one. Smaller batches mean cheaper, noisier steps, and many more of them per pass through the data.
A year of rides: 10 epochs at η = 0.05
L* = 1.996, starting loss 1008.3 |B| = 240: 10 updates, loss 124.3 after 10 epochs |B| = 16: 150 updates, 1.999 after 4 epochs, then within 0.003 of L* |B| = 1: 2,400 updates, 2.11 after 1 epoch, then bouncing between 2.05 and 2.52
Every row reads each ride exactly ten times, so the work is the same. The data was generated as 3 min/km + 5 plus noise, and B = 16 recovers w = 2.975 and b = 5.031.
Why batch size 1 never settles
at θ*: full gradient ∇L = 0, but single-ride gradients ∇ℓᵢ ≠ 0 spread of ∇L_B ∝ 1 / √|B| (for batches drawn at random)
At the minimum the rides' pulls cancel on average, but each single ride still pulls its own way, so the knobs keep jiggling. Bigger batches average out more of the noise; a shrinking learning rate is the other cure (second exercise).

• Take the bowl exactly: L(θ) = L* + ½ (θ − θ*)ᵀ H (θ − θ*). Its gradient is ∇L = H (θ − θ*), which is zero at θ* as it should be.

• One step of gradient descent gives θ_new − θ* = (θ − θ*) − η·H (θ − θ*) = (I − ηH)(θ − θ*). So the distance to the minimum is multiplied by the matrix (I − ηH) every step.

• H is symmetric, so it has perpendicular eigenvectors with eigenvalues λ₁, λ₂, … Along eigenvector k, (I − ηH) just multiplies by (1 − η·λₖ). The directions don't mix: each one shrinks or grows on its own.

• After t steps the component along direction k is (1 − η·λₖ)ᵗ times its starting value. That goes to 0 exactly when |1 − η·λₖ| < 1, meaning 0 < η < 2/λₖ. Every direction must shrink, so η < 2/λ_max.

• For the rides, 2/λ_max = 2/120.04 = 0.01666. At η = 0.016 the steep direction's factor is 1 − 1.92 = −0.92: it overshoots each time but shrinks. At η = 0.017 it's 1 − 2.04 = −1.04, and the error grows by 4% per step, which over 2,000 steps is a factor of about 10³⁴. ∎

• For the line ŷ = w·x + b with MSE, the second derivatives are ∂²L/∂w² = (2/N) Σ xᵢ², ∂²L/∂w∂b = (2/N) Σ xᵢ and ∂²L/∂b² = (2/N)·N = 2. Each is a constant, which is why the bowl is exact.

• For the rides, Σx² = 4 + 25 + 64 + 144 = 237 and Σx = 27, so H = [118.5, 13.5; 13.5, 2]. The off-diagonal 13.5 tilts the bowl, and the 118.5 against 2 stretches it.

• After standardizing, x′ has mean 0 and variance 1: Σ x′ᵢ = 0 and Σ x′ᵢ² = N. Then H′ = (2/N)[N, 0; 0, N] = [2, 0; 0, 2] = 2I. Both curvatures equal 2, so κ = 1: a perfectly round bowl.

• With H = 2I the step is θ_new − θ* = (I − η·2I)(θ − θ*) = (1 − 2η)(θ − θ*). Choosing η = 0.5 makes the factor 0, so one step lands exactly on θ*. η = 0.1 and η = 0.9 give factors 0.8 and −0.8, the same size, which is why both need 27 steps.

• Real networks aren't quadratic and their H changes from place to place, so no single step size is perfect. But the lesson carries over: inputs on similar scales give rounder bowls and faster training. Chapter 12 does the same thing inside the network. ∎

3. How it works

1
Prepare the data
Put examples in rows and standardize each input column: subtract its mean, divide by its standard deviation. Remember both numbers, because new inputs need the same treatment.
2
Initialize the knobs
Make every parameter a Tensor. For the line, zeros are fine. Chapter 12 explains why networks need random starting values instead.
3
Loop over epochs and batches
Each epoch, shuffle the rows and cut them into batches. For each batch: forward, zero the grads, backward, step.
4
Watch the loss
Print or plot the loss every so often. Falling steadily: fine. Falling very slowly: try a larger η or standardize. Growing or NaN: η is too large.
5
Decide when to stop
Stop after a fixed budget of epochs, or when the loss has stopped improving for a while. Chapter 11 adds the most important rule: stop when the loss on data you didn't train on stops improving.

4. The code (python)

core_ch7.py
# Chapter 7: the training loop. Forward, zero the grads, backward, step, and repeat.
# Uses the Tensor engine from Chapter 6: save that chapter's code down to (not including) the line
# '# === 1.' as tensor.py next to this file. That's unbroadcast(), the Tensor class and count().
import numpy as np
from tensor import Tensor

km = np.array([2.0, 5.0, 8.0, 12.0])
minutes = np.array([11.0, 21.0, 28.0, 42.0])


def train(x, y, lr, steps, batch_size=None, seed=0):
    """Gradient descent on the line w*x + b. Returns the final knobs and the full-data loss after every update."""
    w, b = Tensor(0.0), Tensor(0.0)                     # start knowing nothing
    params = [w, b]
    rng = np.random.default_rng(seed)
    n = len(x)
    batch_size = batch_size or n                        # default: the whole dataset every step
    history = [float(np.mean((w.data * x + b.data - y) ** 2))]
    updates = 0
    while updates < steps:
        order = rng.permutation(n)                      # one epoch = one pass through the data, shuffled
        for start in range(0, n, batch_size):
            rows = order[start:start + batch_size]
            loss = ((w * x[rows] + b - y[rows]) ** 2).mean()        # 1. forward
            for p in params:                                        # 2. zero the grads
                p.grad = np.zeros_like(p.data)
            loss.backward()                                         # 3. backward
            for p in params:                                        # 4. step downhill
                p.data = p.data - lr * p.grad
            history.append(float(np.mean((w.data * x + b.data - y) ** 2)))
            updates += 1
            if updates == steps or not np.isfinite(history[-1]):
                return float(w.data), float(b.data), history
    return float(w.data), float(b.data), history


# === 1. The best line, in closed form, to know what we're aiming for ==============
w_best = np.mean((km - km.mean()) * (minutes - minutes.mean())) / np.var(km)
b_best = minutes.mean() - w_best * km.mean()
loss_best = np.mean((w_best * km + b_best - minutes) ** 2)
print(f"best line: w* = {w_best:.4f}, b* = {b_best:.4f}, loss = {loss_best:.4f}")

# === 2. Four learning rates, 2,000 steps each ======================================
print("\n   lr      w        b       loss after 2,000   first step within 1% of best")
for lr in [0.001, 0.01, 0.016, 0.017]:
    w, b, hist = train(km, minutes, lr, 2000)
    near = next((i for i, v in enumerate(hist) if v < 1.01 * loss_best), None)
    print(f"{lr:6}  {w:8.4g} {b:8.4g}   {hist[-1]:12.4g}        {near}")

# === 3. Why 0.017 explodes: the curvature of the bowl ===============================
X = np.column_stack([km, np.ones(4)])
H = 2 / len(km) * X.T @ X                          # second derivatives of the MSE in (w, b)
lam = np.linalg.eigvalsh(H)
print(f"\ncurvatures (eigenvalues of H): {lam.round(3)}, condition number {lam.max() / lam.min():.0f}")
print(f"largest stable learning rate: 2 / {lam.max():.2f} = {2 / lam.max():.4f}")

# === 4. Standardize the input: the bowl becomes round ================================
mu, sd = km.mean(), km.std()
km_std = (km - mu) / sd
for lr in [0.1, 0.5, 0.9]:
    w_s, b_s, hist = train(km_std, minutes, lr, 30)
    near = next((i for i, v in enumerate(hist) if v < 1.01 * loss_best), None)
    print(f"standardized, lr {lr}: within 1% after {near} steps;"
          f" back in km: w = {w_s / sd:.4f}, b = {b_s - w_s * mu / sd:.4f}")

# === 5. More data: a year of rides, and mini-batches ==================================
rng = np.random.default_rng(7)
year_km = np.round(rng.uniform(1, 15, 240), 1)
year_min = np.round(3 * year_km + 5 + rng.normal(0, 1.5, 240), 1)   # 3 min/km + 5, plus noise
year_std = (year_km - year_km.mean()) / year_km.std()
Xy = np.column_stack([year_std, np.ones(240)])
best = np.linalg.lstsq(Xy, year_min, rcond=None)[0]
print(f"\nyear of rides: best loss {np.mean((Xy @ best - year_min) ** 2):.3f}, starting loss {np.mean(year_min ** 2):.3f}")

print("batch  updates   loss after each of 10 epochs")
for B in [240, 16, 1]:
    steps = 10 * (240 // B)                       # 10 epochs, whatever the batch size
    w_s, b_s, hist = train(year_std, year_min, 0.05, steps, batch_size=B)
    per_epoch = hist[240 // B::240 // B]
    print(f"{B:5}  {steps:7}   {[round(v, 3) for v in per_epoch]}")
w_s, b_s, _ = train(year_std, year_min, 0.05, 150, batch_size=16)
print(f"B = 16, back in km: w = {w_s / year_km.std():.3f}, b = {b_s - w_s * year_km.mean() / year_km.std():.3f}")

5. Practice

Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.

P1 Start the ride model at w = 0, b = 0, where ∂L/∂w = −427.5 and ∂L/∂b = −51. After one step with η = 0.01, what is w?
w_new = w − η · ∂L/∂w.
Solution. 0 − 0.01 · (−427.5) = 4.275. The first step overshoots w* = 3.041, because the steep direction is being corrected almost entirely in one go.
P2 A one-knob loss is L(w) = 5(w − 2)². What is the largest learning rate that still converges?
The curvature is the second derivative, L″(w) = 10. Stability needs η < 2/λ.
Solution. λ = 10, so η < 2/10 = 0.2.
P3 Same loss, starting at w = 0, with η = 0.15. Each step multiplies the distance to the minimum by 1 − η·λ. Where is w after 3 steps?
1 − 0.15 · 10 = −0.5, and the starting distance is 0 − 2 = −2.
Solution. The distance becomes −2 · (−0.5)³ = −2 · (−0.125) = 0.25, so w = 2 + 0.25 = 2.25. It jumps from side to side: 3, 1.5, 2.25.
P4 A bowl has curvatures 120.04 and 0.456. What is its condition number? (Nearest whole number.)
κ = λ_max / λ_min.
Solution. 120.04 / 0.456 ≈ 263.
P5 Standardize the 12 km ride, with mean 6.75 km and standard deviation 3.70 km. (3 decimal places.)
x′ = (x − x̄) / s.
Solution. (12 − 6.75) / 3.70 = 5.25 / 3.70 ≈ 1.419: the 12 km ride is about 1.4 standard deviations above average.
P6 240 rides, batch size 16. How many updates does one epoch take?
One epoch reads every example once.
Solution. 240 / 16 = 15 updates per epoch.
P7 Training on standardized km gives w′ = 11.251. With s = 3.70, what is w in minutes per km? (3 decimal places.)
w′·x′ = w′·(x − x̄)/s, so w = w′ / s.
Solution. 11.251 / 3.70 ≈ 3.041, the closed-form w*.
P8 A batch of four rides gives per-ride slopes for b of [2, −1, 3, 0]. What is the batch gradient for b?
The batch loss is a mean, so its gradient is the mean of the per-example gradients.
Solution. (2 − 1 + 3 + 0) / 4 = 1.

6. Go further

  1. Find the cliff edge yourself: bisect between η = 0.016 (converges) and η = 0.017 (diverges), running 2,000 steps each time, until the interval is narrower than 0.0001. How close do you get to 2 / λ_max = 0.01666?
  2. Calm down batch size 1: rerun the year-of-rides experiment with B = 1 and a shrinking learning rate, η_t = 0.05 / (1 + t / 500) where t counts updates. Does the loss now settle near 1.996? What happens if it shrinks too fast, say η_t = 0.05 / (1 + t)?
  3. Train the picnic neuron from Chapter 6 with this loop: full batch, η = 1.0, 500 steps, printing the loss every 100 steps. The picnic inputs already lie between 0 and 1. Does standardizing them still change how fast it trains?

7. Check yourself

Answer all 5 questions correctly to complete the chapter · 0 / 5 done
Q1/5 During training the loss grows every step and soon becomes NaN. What's the most likely fix?
Q2/5 Why does the loop zero the gradients before calling backward()?
Q3/5 Why does standardizing the input speed up training of the ride model so much?
Q4/5 On the year of rides, why do batches of 16 beat the full batch after the same 10 epochs?
Q5/5 With batch size 1, the loss gets near the minimum and then keeps bouncing. Why?

In the catalog

Finished this chapter?