Learn machine learning from scratch

Start at chapter 1 and build up: what a model is, how it measures its own mistakes, and how it learns. Everything is worked out in plain Python, with the maths shown rather than assumed. Two shorter tracks hold reference notes on the maths and on modern architectures.

Track Progress: 0 Completed
0%
Core CourseLEVEL 2 · INTERMEDIATEChapter 9

Choosing Among Many: Classification with Softmax

Chapter 2's layer gave three scores for picnic, kite flying and reading. Softmax turns scores into probabilities that add up to 1, cross-entropy scores them, and the gradient is once again just 'predicted minus actual'.

38 min Prerequisites: Chapters 2, 3 and 6–8 (the Tensor class saved as tensor.py)
Inspect Architecture: Logistic regression

1. The idea

Chapter 2 built a layer of three neurons that each scored tomorrow's weather for one activity: a picnic, flying a kite, or staying in to read. For sun 0.7, rain 0.2 and wind 0.5 the scores were 0.8, 0.3 and 0.5. Picnic wins, but by how much? Scores can be any size and any sign, so they don't answer that. What we want is a probability for each activity, all positive and adding up to 1, like the single probability Chapter 3 got from the sigmoid.

Softmax does it in two moves. First, make every score positive with the exponential: e^0.8 = 2.23, e^0.3 = 1.35, e^0.5 = 1.65. Then divide each by their sum, 5.22. The result is 0.43 for the picnic, 0.26 for the kite and 0.32 for reading. The order of the scores never changes, but the gaps between them now mean something. With only two classes, softmax turns out to be exactly the sigmoid of the difference between the two scores, so this chapter generalizes Chapter 3 rather than replacing it.

The loss is Chapter 3's cross-entropy, unchanged: minus the log of the probability given to what actually happened. If they flew a kite tomorrow, the model gave that 0.26 and pays −ln 0.26 = 1.35. The gradient with respect to the scores is the same pleasant surprise as before: p − y, the predicted probabilities minus the one-hot vector of what happened, [0.43, −0.74, 0.32]. Push up the score of what happened, push down the others in proportion to how much probability they took.

Then we train one. A season of 150 days is labelled with what people actually did: read when it rained, fly a kite when it was windy, otherwise picnic, except that one day in ten they did something else anyway. A single linear layer with softmax gets 93.3% of the days right, and its weights are easy to read: the 'read' score puts +11.7 on rain. A hidden layer gets to 95.3%. The rule that made the labels only scores 96%, because of those days off-script, so anything higher would mean the model was memorizing noise. That's the subject of Chapter 11.

Two bar charts. Left, the scores of Chapter 2's layer for tomorrow's weather: picnic 0.8, kite 0.3, read 0.5. Right, the probabilities softmax gives at three temperatures. At T = 1: picnic 0.43, kite 0.26, read 0.32. At T = 0.5 the bars are more uneven, with picnic about 0.52 and kite about 0.19. At T = 2 they are flatter, all between about 0.30 and 0.38. In every case picnic is highest and kite lowest.Two bar charts. Left, the scores of Chapter 2's layer for tomorrow's weather: picnic 0.8, kite 0.3, read 0.5. Right, the probabilities softmax gives at three temperatures. At T = 1: picnic 0.43, kite 0.26, read 0.32. At T = 0.5 the bars are more uneven, with picnic about 0.52 and kite about 0.19. At T = 2 they are flatter, all between about 0.30 and 0.38. In every case picnic is highest and kite lowest.
Softmax keeps the order of the scores but turns them into probabilities. The temperature T decides how decisive the result is.

2. The math

Softmax: from scores to probabilities
pₖ = e^(zₖ) / Σⱼ e^(zⱼ) tomorrow: z = [0.8, 0.3, 0.5] e^z = [2.2255, 1.3499, 1.6487], sum 5.2241 p = [0.426, 0.258, 0.316]
z are the scores of Chapter 2's three neurons, also called logits. The exponential makes every score positive, and dividing by the sum makes them add up to 1.
What softmax keeps, and what it adds
every pₖ > 0, Σₖ pₖ = 1, bigger score ⇒ bigger probability prediction: argmaxₖ pₖ = argmaxₖ zₖ = picnic
To pick a single answer you don't need softmax at all: the largest score wins. Softmax is for training, and for knowing how sure the model is.
Adding a constant changes nothing, so subtract the maximum
softmax(z + c) = softmax(z) for any number c z = [1000, 999, 998]: e^1000 overflows to ∞ and ∞/∞ = NaN z − max = [0, −1, −2]: p = [0.665, 0.245, 0.090]
The common factor e^c cancels between top and bottom. Every real implementation shifts by the largest score first, so the biggest exponential is e⁰ = 1.
Two classes: softmax is the sigmoid
softmax([z₁, z₂])₁ = e^(z₁) / (e^(z₁) + e^(z₂)) = 1 / (1 + e^(−(z₁ − z₂))) = σ(z₁ − z₂) picnic vs kite only: σ(0.8 − 0.3) = σ(0.5) = 0.6225
Only the difference between the two scores matters. Chapter 3's single sigmoid neuron is a two-class softmax with the second score fixed at 0.
Temperature: how decisive
softmax(z / T) T = 0.5: [0.522, 0.192, 0.286] T = 1: [0.426, 0.258, 0.316] T = 2: [0.379, 0.295, 0.326]
Dividing by T < 1 stretches the gaps between scores and sharpens the probabilities; T > 1 shrinks them. Language models use exactly this knob when they sample their next word (Chapter 16).
Cross-entropy with one-hot labels
y = one-hot of what happened, e.g. kite = [0, 1, 0] ℓ = −Σₖ yₖ ln pₖ = −ln p_true kite: ℓ = −ln 0.2584 = 1.3533 a uniform guess, p = 1/3 each: ℓ = ln 3 = 1.0986
The same loss as Chapter 3: the surprise at what actually happened. With K classes, the 'know nothing' baseline is ln K, so a model scoring above ln 3 here is worse than guessing.
The gradient is p − y again
∂ℓ/∂zₖ = pₖ − yₖ kite: p − y = [0.426, −0.742, 0.316] the entries always sum to 0
The true class's score is pushed up by 1 − p_true; every other score is pushed down by the probability it took. Derived in the first derivation, and checked against the engine in the code.
A batch, in Chapter 6's shapes
Z = X Wᵀ + b (N × 3) P = softmax of each row L = mean of −ln P[i, labelᵢ] dZ = (P − Y) / N dW = dZᵀ X (3 × 3) db = column sums of dZ
Nothing new: the p − y rule per row, the 1/N from the mean, and Chapter 6's matrix-product and broadcasting rules. The code builds softmax from exp, a row sum (a product with a column of ones) and division, so the engine needs no changes.
A season of 150 days
labels: rainy (rain > 0.5) → read; else windy (wind > 0.6) → kite; else picnic then 1 day in 10 relabelled at random the rule itself matches 96% of the labels softmax regression (12 knobs), from zero, η = 1.0, 3,000 steps: loss 1.0986 → 0.3052, accuracy 93.3% with 8 hidden units: 94.7% with 16 hidden units: 95.3%
Starting from zero weights gives p = 1/3 for every day, so the loss starts at exactly ln 3. A hidden layer helps because the real boundaries are 'rain above 0.5' and 'wind above 0.6', which a single linear layer can only approximate.
Reading the trained weights
picnic: (sun, rain, wind) = (−0.61, −7.05, −5.77), bias +6.85 kite: (1.20, −4.62, 5.69), bias −1.94 read: (−0.58, 11.68, 0.08), bias −4.91
Read is a rain detector. Kite likes wind and dislikes rain. Picnic dislikes both but starts with a big head start from its bias. Sun matters least, because the rule that made the labels never looked at it.

• Let c be the true class, so ℓ = −ln p_c. Write ln p_c with the softmax formula: ln p_c = z_c − ln Σⱼ e^(zⱼ). So ℓ = −z_c + ln Σⱼ e^(zⱼ).

• Differentiate the second term with respect to any score zₖ, using the chain rule with (ln u)′ = 1/u: ∂/∂zₖ ln Σⱼ e^(zⱼ) = e^(zₖ) / Σⱼ e^(zⱼ) = pₖ.

• The first term, −z_c, has slope −1 with respect to z_c and 0 with respect to every other score. That's −yₖ, where y is the one-hot vector.

• Add the two parts: ∂ℓ/∂zₖ = pₖ − yₖ. ∎ The entries sum to Σ pₖ − Σ yₖ = 1 − 1 = 0, which makes sense: adding the same amount to every score changes nothing, so the gradient can't point in that direction.

• Writing the loss as −z_c + ln Σ e^(zⱼ) avoids ever computing the softmax in a way that can overflow or take ln 0. Frameworks provide a fused 'cross-entropy from logits' for exactly this reason; the first exercise builds one.

• Add the same number c to every score: e^(zₖ + c) / Σⱼ e^(zⱼ + c) = e^c·e^(zₖ) / (e^c·Σⱼ e^(zⱼ)). The factor e^c cancels, leaving pₖ. ∎

• Choose c = −max(z). The largest shifted score is 0 and all others are negative, so every exponential lies between 0 and 1. Nothing can overflow, and the denominator is at least 1, so nothing divides by zero.

• For two classes, shift by −z₂: p₁ = e^(z₁ − z₂) / (e^(z₁ − z₂) + e⁰) = 1 / (1 + e^(−(z₁ − z₂))) = σ(z₁ − z₂). ∎

• So a two-class softmax has one redundant score: only z₁ − z₂ matters. Chapter 3's sigmoid neuron is that case with the second score fixed at zero, and its gradient p − y from Chapter 4 is the two-class version of this chapter's rule.

3. How it works

1
One score per class
Give the last layer one neuron per class: Z = X Wᵀ + b has shape (examples × classes). These raw scores are the logits.
2
Softmax each row
Subtract the row's largest score, exponentiate, divide by the row's sum. Each row is now a probability distribution over the classes.
3
Encode the labels
Turn each label into a one-hot row: a 1 in its class's column, zeros elsewhere. Stacked, they form Y with the same shape as P.
4
Cross-entropy and backward
Average −ln of the probability given to each true class. The gradient reaching the scores is (P − Y) / N, and the engine carries it back into W and b.
5
Predict and measure
The predicted class is the largest score in each row. Report accuracy, but also the loss: two models with the same accuracy can differ a lot in how confident they are.

4. The code (python)

core_ch9.py
# Chapter 9: more than two classes. Softmax turns a layer's scores into probabilities.
# Uses the Tensor class from Chapter 6, saved as tensor.py (see Chapter 7).
import numpy as np
from tensor import Tensor

ACTIVITIES = ["picnic", "kite", "read"]


def softmax(z):
    """Plain NumPy, one row of scores per example. Subtracting the row's largest score changes nothing but avoids overflow."""
    e = np.exp(z - z.max(axis=-1, keepdims=True))
    return e / e.sum(axis=-1, keepdims=True)


# === 1. Chapter 2's layer, scored on tomorrow's weather ===========================
W = np.array([[2.0, -3.0, -1.0],     # picnic
              [-1.0, 0.0, 3.0],      # kite flying
              [0.0, 2.5, 0.0]])      # stay in and read
b = np.array([0.5, -0.5, 0.0])
tomorrow = np.array([0.7, 0.2, 0.5])                 # (sun, rain, wind)
z = W @ tomorrow + b
print("scores z:", z, " e^z:", np.exp(z).round(4), " sum:", np.exp(z).sum().round(4))
print("softmax: ", softmax(z).round(4))
for T in [0.5, 2.0]:
    print(f"T = {T}:   ", softmax(z / T).round(4))

# Two classes is the sigmoid in disguise: softmax([z1, z2])[0] = sigmoid(z1 - z2)
print("picnic vs kite only:", softmax(z[:2]).round(4), " sigmoid(0.8 - 0.3) =", round(1 / (1 + np.exp(-0.5)), 4))

# Why we subtract the maximum first
big = np.array([1000.0, 999.0, 998.0])
with np.errstate(over="ignore", invalid="ignore"):
    naive = np.exp(big) / np.exp(big).sum()
print("naive softmax of [1000, 999, 998]:", naive, "  shifted:", softmax(big).round(4))

# === 2. The loss, and its gradient p − y ============================================
p = softmax(z)
y = np.array([0.0, 1.0, 0.0])                        # tomorrow they went kite flying
print(f"\ncross-entropy if they flew a kite: -ln {p[1]:.4f} = {-np.log(p[1]):.4f}   (uniform guess: ln 3 = {np.log(3):.4f})")
print("gradient with respect to z, p − y:", (p - y).round(4), " sums to", round(float((p - y).sum()), 12))


def softmax_cross_entropy(Z, Y):
    """Mean cross-entropy of softmax(Z) against one-hot rows Y, built from Chapter 6's operations only."""
    shifted = Z + (-Z.data.max(axis=1, keepdims=True))            # a constant shift: no gradient needed
    E = shifted.exp()
    row_sums = E @ Tensor(np.ones((Z.shape[1], 1)))               # (N, 3)(3, 1): each row's total
    P = E / row_sums                                              # broadcast over the 3 columns
    return -(Tensor(Y) * P.log()).sum() / Z.shape[0], P

# Check the engine against p − y: make the scores a leaf and ask for their gradient
Z = Tensor(z[None, :])
loss, P = softmax_cross_entropy(Z, y[None, :])
loss.backward()
print("engine's dL/dz:", Z.grad.round(4), " loss", round(float(loss.data), 4))

# === 3. A season of days, and a classifier trained on them ===========================
rng = np.random.default_rng(9)
X = np.round(rng.random((150, 3)), 2)                # (sun, rain, wind) for 150 days
rule = np.where(X[:, 1] > 0.5, 2, np.where(X[:, 2] > 0.6, 1, 0))   # rainy: read; else windy: kite; else picnic
labels = rule.copy()
noisy = rng.random(150) < 0.1                        # 1 day in 10, people did something else
labels[noisy] = rng.integers(0, 3, noisy.sum())
Y = np.eye(3)[labels]                                # one-hot rows, (150, 3)
print(f"\n150 days: {np.bincount(labels).tolist()} picnic/kite/read;"
      f" the rule itself is right on {np.mean(rule == labels):.0%} of them")


def train(params, scores, lr, steps):
    for _ in range(steps):
        loss, _ = softmax_cross_entropy(scores(), Y)
        for prm in params:
            prm.grad = np.zeros_like(prm.data)
        loss.backward()
        for prm in params:
            prm.data = prm.data - lr * prm.grad
    loss, P = softmax_cross_entropy(scores(), Y)
    return float(loss.data), np.mean(P.data.argmax(axis=1) == labels)

# Softmax regression: one linear layer, 3 inputs -> 3 scores, starting from all zeros
W = Tensor(np.zeros((3, 3)))
b = Tensor(np.zeros(3))
start, _ = softmax_cross_entropy(Tensor(X) @ W.T + b, Y)
loss, acc = train([W, b], lambda: Tensor(X) @ W.T + b, lr=1.0, steps=3000)
print(f"softmax regression: loss {float(start.data):.4f} -> {loss:.4f}, accuracy {acc:.1%}")
for name, row, bias in zip(ACTIVITIES, W.data, b.data):
    print(f"  {name:6} weights (sun, rain, wind) = {row.round(2)}, bias {bias:+.2f}")

# With a hidden layer of 8 or 16 ReLU units (Chapter 8)
for hidden in [8, 16]:
    r = np.random.default_rng(0)
    W1, b1 = Tensor(r.normal(size=(hidden, 3))), Tensor(np.zeros(hidden))
    W2, b2 = Tensor(r.normal(size=(3, hidden))), Tensor(np.zeros(3))
    loss, acc = train([W1, b1, W2, b2], lambda: (Tensor(X) @ W1.T + b1).relu() @ W2.T + b2, lr=0.5, steps=5000)
    print(f"{hidden:2} hidden units:     loss {loss:.4f}, accuracy {acc:.1%}")

5. Practice

Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.

P1 What does softmax give for the scores [0, 0, 0]? Enter the first probability. (3 decimal places.)
e⁰ = 1 for every class.
Solution. 1 / (1 + 1 + 1) = 1/3 ≈ 0.333 for each class. Equal scores give equal probabilities.
P2 Two classes with scores [2, 0]. What probability does softmax give the first? (4 decimal places.)
Two-class softmax is σ(z₁ − z₂).
Solution. σ(2) = 1 / (1 + e⁻²) ≈ 0.8808.
P3 Tomorrow's probabilities are [0.426, 0.258, 0.316] for picnic, kite, read. If they stayed in and read, what is the cross-entropy? (3 decimal places.)
ℓ = −ln p_true.
Solution. −ln 0.316 ≈ 1.153. Lower than the kite's 1.353, because the model gave reading more probability.
P4 Same probabilities, and they flew a kite. What is the gradient entry for the picnic score, ∂ℓ/∂z_picnic?
pₖ − yₖ, and y_picnic = 0.
Solution. 0.426 − 0 = 0.426: the picnic score gets pushed down in proportion to the probability it took.
P5 Add up the three entries of p − y for any example. What do you get?
Both p and y sum to 1.
Solution. Σ pₖ − Σ yₖ = 1 − 1 = 0. Raising all scores together changes nothing, so the gradient has no component in that direction.
P6 Softmax of [5, 3, 4]: what is the first probability? Shift by the maximum first. (4 decimal places.)
z − 5 = [0, −2, −1].
Solution. 1 / (1 + e⁻² + e⁻¹) = 1 / (1 + 0.1353 + 0.3679) ≈ 0.6652, the same as for [1000, 999, 998].
P7 A classifier gets 140 of the 150 days right. What is its accuracy, as a fraction? (3 decimal places.)
Correct divided by total.
Solution. 140 / 150 ≈ 0.933, the softmax regression's 93.3%.
P8 A model with 10 classes predicts 0.1 for every class. What is its cross-entropy? (3 decimal places.)
−ln of the probability given to the true class, whatever it is.
Solution. −ln 0.1 = ln 10 ≈ 2.303. That's the 'know nothing' baseline for 10 classes, which Chapter 13's digit classifier has to beat.

6. Go further

  1. Build a fused SoftmaxCrossEntropy operation in tensor.py: its forward pass computes the mean of −z_true + ln Σ e^z per row (with the max subtracted), and its _backward adds (P − Y) / N to the scores' grad directly. Check it against softmax_cross_entropy() on the 150 days, then compare how many Tensors each version puts in the graph.
  2. Which days does the trained softmax regression get wrong? Print a 3 × 3 table of true activity against predicted activity (a confusion matrix), and look at the inputs of the mistakes. How many of them are the noisy days, where people ignored the rule?
  3. Sample tomorrow's activity 1,000 times from softmax(z / T) for T = 0.5, 1 and 2 (np.random.default_rng().choice with p=...). Count how often each activity comes up. How does T change the mix, and what happens as T approaches 0?

7. Check yourself

Answer all 5 questions correctly to complete the chapter · 0 / 5 done
Q1/5 What does softmax do to a row of scores?
Q2/5 Why do implementations subtract the largest score before exponentiating?
Q3/5 A model gives the true class probability 0.26. What gradient reaches that class's score?
Q4/5 With two classes, softmax is the same as which earlier tool?
Q5/5 The labelling rule matches 96% of the 150 days. Why shouldn't you aim for 100% training accuracy?

Finished this chapter?