Choosing Among Many: Classification with Softmax
Chapter 2's layer gave three scores for picnic, kite flying and reading. Softmax turns scores into probabilities that add up to 1, cross-entropy scores them, and the gradient is once again just 'predicted minus actual'.
Inspect Architecture: Logistic regression1. The idea
Chapter 2 built a layer of three neurons that each scored tomorrow's weather for one activity: a picnic, flying a kite, or staying in to read. For sun 0.7, rain 0.2 and wind 0.5 the scores were 0.8, 0.3 and 0.5. Picnic wins, but by how much? Scores can be any size and any sign, so they don't answer that. What we want is a probability for each activity, all positive and adding up to 1, like the single probability Chapter 3 got from the sigmoid.
Softmax does it in two moves. First, make every score positive with the exponential: e^0.8 = 2.23, e^0.3 = 1.35, e^0.5 = 1.65. Then divide each by their sum, 5.22. The result is 0.43 for the picnic, 0.26 for the kite and 0.32 for reading. The order of the scores never changes, but the gaps between them now mean something. With only two classes, softmax turns out to be exactly the sigmoid of the difference between the two scores, so this chapter generalizes Chapter 3 rather than replacing it.
The loss is Chapter 3's cross-entropy, unchanged: minus the log of the probability given to what actually happened. If they flew a kite tomorrow, the model gave that 0.26 and pays −ln 0.26 = 1.35. The gradient with respect to the scores is the same pleasant surprise as before: p − y, the predicted probabilities minus the one-hot vector of what happened, [0.43, −0.74, 0.32]. Push up the score of what happened, push down the others in proportion to how much probability they took.
Then we train one. A season of 150 days is labelled with what people actually did: read when it rained, fly a kite when it was windy, otherwise picnic, except that one day in ten they did something else anyway. A single linear layer with softmax gets 93.3% of the days right, and its weights are easy to read: the 'read' score puts +11.7 on rain. A hidden layer gets to 95.3%. The rule that made the labels only scores 96%, because of those days off-script, so anything higher would mean the model was memorizing noise. That's the subject of Chapter 11.
2. The math
pₖ = e^(zₖ) / Σⱼ e^(zⱼ)
tomorrow: z = [0.8, 0.3, 0.5] e^z = [2.2255, 1.3499, 1.6487], sum 5.2241
p = [0.426, 0.258, 0.316]every pₖ > 0, Σₖ pₖ = 1, bigger score ⇒ bigger probability
prediction: argmaxₖ pₖ = argmaxₖ zₖ = picnicsoftmax(z + c) = softmax(z) for any number c
z = [1000, 999, 998]: e^1000 overflows to ∞ and ∞/∞ = NaN
z − max = [0, −1, −2]: p = [0.665, 0.245, 0.090]softmax([z₁, z₂])₁ = e^(z₁) / (e^(z₁) + e^(z₂))
= 1 / (1 + e^(−(z₁ − z₂))) = σ(z₁ − z₂)
picnic vs kite only: σ(0.8 − 0.3) = σ(0.5) = 0.6225softmax(z / T)
T = 0.5: [0.522, 0.192, 0.286]
T = 1: [0.426, 0.258, 0.316]
T = 2: [0.379, 0.295, 0.326]y = one-hot of what happened, e.g. kite = [0, 1, 0]
ℓ = −Σₖ yₖ ln pₖ = −ln p_true
kite: ℓ = −ln 0.2584 = 1.3533
a uniform guess, p = 1/3 each: ℓ = ln 3 = 1.0986∂ℓ/∂zₖ = pₖ − yₖ kite: p − y = [0.426, −0.742, 0.316]
the entries always sum to 0Z = X Wᵀ + b (N × 3) P = softmax of each row
L = mean of −ln P[i, labelᵢ]
dZ = (P − Y) / N dW = dZᵀ X (3 × 3) db = column sums of dZlabels: rainy (rain > 0.5) → read; else windy (wind > 0.6) → kite; else picnic
then 1 day in 10 relabelled at random
the rule itself matches 96% of the labels
softmax regression (12 knobs), from zero, η = 1.0, 3,000 steps:
loss 1.0986 → 0.3052, accuracy 93.3%
with 8 hidden units: 94.7% with 16 hidden units: 95.3%picnic: (sun, rain, wind) = (−0.61, −7.05, −5.77), bias +6.85
kite: (1.20, −4.62, 5.69), bias −1.94
read: (−0.58, 11.68, 0.08), bias −4.91• Let c be the true class, so ℓ = −ln p_c. Write ln p_c with the softmax formula: ln p_c = z_c − ln Σⱼ e^(zⱼ). So ℓ = −z_c + ln Σⱼ e^(zⱼ).
• Differentiate the second term with respect to any score zₖ, using the chain rule with (ln u)′ = 1/u: ∂/∂zₖ ln Σⱼ e^(zⱼ) = e^(zₖ) / Σⱼ e^(zⱼ) = pₖ.
• The first term, −z_c, has slope −1 with respect to z_c and 0 with respect to every other score. That's −yₖ, where y is the one-hot vector.
• Add the two parts: ∂ℓ/∂zₖ = pₖ − yₖ. ∎ The entries sum to Σ pₖ − Σ yₖ = 1 − 1 = 0, which makes sense: adding the same amount to every score changes nothing, so the gradient can't point in that direction.
• Writing the loss as −z_c + ln Σ e^(zⱼ) avoids ever computing the softmax in a way that can overflow or take ln 0. Frameworks provide a fused 'cross-entropy from logits' for exactly this reason; the first exercise builds one.
• Add the same number c to every score: e^(zₖ + c) / Σⱼ e^(zⱼ + c) = e^c·e^(zₖ) / (e^c·Σⱼ e^(zⱼ)). The factor e^c cancels, leaving pₖ. ∎
• Choose c = −max(z). The largest shifted score is 0 and all others are negative, so every exponential lies between 0 and 1. Nothing can overflow, and the denominator is at least 1, so nothing divides by zero.
• For two classes, shift by −z₂: p₁ = e^(z₁ − z₂) / (e^(z₁ − z₂) + e⁰) = 1 / (1 + e^(−(z₁ − z₂))) = σ(z₁ − z₂). ∎
• So a two-class softmax has one redundant score: only z₁ − z₂ matters. Chapter 3's sigmoid neuron is that case with the second score fixed at zero, and its gradient p − y from Chapter 4 is the two-class version of this chapter's rule.
3. How it works
4. The code (python)
# Chapter 9: more than two classes. Softmax turns a layer's scores into probabilities.
# Uses the Tensor class from Chapter 6, saved as tensor.py (see Chapter 7).
import numpy as np
from tensor import Tensor
ACTIVITIES = ["picnic", "kite", "read"]
def softmax(z):
"""Plain NumPy, one row of scores per example. Subtracting the row's largest score changes nothing but avoids overflow."""
e = np.exp(z - z.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
# === 1. Chapter 2's layer, scored on tomorrow's weather ===========================
W = np.array([[2.0, -3.0, -1.0], # picnic
[-1.0, 0.0, 3.0], # kite flying
[0.0, 2.5, 0.0]]) # stay in and read
b = np.array([0.5, -0.5, 0.0])
tomorrow = np.array([0.7, 0.2, 0.5]) # (sun, rain, wind)
z = W @ tomorrow + b
print("scores z:", z, " e^z:", np.exp(z).round(4), " sum:", np.exp(z).sum().round(4))
print("softmax: ", softmax(z).round(4))
for T in [0.5, 2.0]:
print(f"T = {T}: ", softmax(z / T).round(4))
# Two classes is the sigmoid in disguise: softmax([z1, z2])[0] = sigmoid(z1 - z2)
print("picnic vs kite only:", softmax(z[:2]).round(4), " sigmoid(0.8 - 0.3) =", round(1 / (1 + np.exp(-0.5)), 4))
# Why we subtract the maximum first
big = np.array([1000.0, 999.0, 998.0])
with np.errstate(over="ignore", invalid="ignore"):
naive = np.exp(big) / np.exp(big).sum()
print("naive softmax of [1000, 999, 998]:", naive, " shifted:", softmax(big).round(4))
# === 2. The loss, and its gradient p − y ============================================
p = softmax(z)
y = np.array([0.0, 1.0, 0.0]) # tomorrow they went kite flying
print(f"\ncross-entropy if they flew a kite: -ln {p[1]:.4f} = {-np.log(p[1]):.4f} (uniform guess: ln 3 = {np.log(3):.4f})")
print("gradient with respect to z, p − y:", (p - y).round(4), " sums to", round(float((p - y).sum()), 12))
def softmax_cross_entropy(Z, Y):
"""Mean cross-entropy of softmax(Z) against one-hot rows Y, built from Chapter 6's operations only."""
shifted = Z + (-Z.data.max(axis=1, keepdims=True)) # a constant shift: no gradient needed
E = shifted.exp()
row_sums = E @ Tensor(np.ones((Z.shape[1], 1))) # (N, 3)(3, 1): each row's total
P = E / row_sums # broadcast over the 3 columns
return -(Tensor(Y) * P.log()).sum() / Z.shape[0], P
# Check the engine against p − y: make the scores a leaf and ask for their gradient
Z = Tensor(z[None, :])
loss, P = softmax_cross_entropy(Z, y[None, :])
loss.backward()
print("engine's dL/dz:", Z.grad.round(4), " loss", round(float(loss.data), 4))
# === 3. A season of days, and a classifier trained on them ===========================
rng = np.random.default_rng(9)
X = np.round(rng.random((150, 3)), 2) # (sun, rain, wind) for 150 days
rule = np.where(X[:, 1] > 0.5, 2, np.where(X[:, 2] > 0.6, 1, 0)) # rainy: read; else windy: kite; else picnic
labels = rule.copy()
noisy = rng.random(150) < 0.1 # 1 day in 10, people did something else
labels[noisy] = rng.integers(0, 3, noisy.sum())
Y = np.eye(3)[labels] # one-hot rows, (150, 3)
print(f"\n150 days: {np.bincount(labels).tolist()} picnic/kite/read;"
f" the rule itself is right on {np.mean(rule == labels):.0%} of them")
def train(params, scores, lr, steps):
for _ in range(steps):
loss, _ = softmax_cross_entropy(scores(), Y)
for prm in params:
prm.grad = np.zeros_like(prm.data)
loss.backward()
for prm in params:
prm.data = prm.data - lr * prm.grad
loss, P = softmax_cross_entropy(scores(), Y)
return float(loss.data), np.mean(P.data.argmax(axis=1) == labels)
# Softmax regression: one linear layer, 3 inputs -> 3 scores, starting from all zeros
W = Tensor(np.zeros((3, 3)))
b = Tensor(np.zeros(3))
start, _ = softmax_cross_entropy(Tensor(X) @ W.T + b, Y)
loss, acc = train([W, b], lambda: Tensor(X) @ W.T + b, lr=1.0, steps=3000)
print(f"softmax regression: loss {float(start.data):.4f} -> {loss:.4f}, accuracy {acc:.1%}")
for name, row, bias in zip(ACTIVITIES, W.data, b.data):
print(f" {name:6} weights (sun, rain, wind) = {row.round(2)}, bias {bias:+.2f}")
# With a hidden layer of 8 or 16 ReLU units (Chapter 8)
for hidden in [8, 16]:
r = np.random.default_rng(0)
W1, b1 = Tensor(r.normal(size=(hidden, 3))), Tensor(np.zeros(hidden))
W2, b2 = Tensor(r.normal(size=(3, hidden))), Tensor(np.zeros(3))
loss, acc = train([W1, b1, W2, b2], lambda: (Tensor(X) @ W1.T + b1).relu() @ W2.T + b2, lr=0.5, steps=5000)
print(f"{hidden:2} hidden units: loss {loss:.4f}, accuracy {acc:.1%}")
5. Practice
Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.
6. Go further
- Build a fused SoftmaxCrossEntropy operation in tensor.py: its forward pass computes the mean of −z_true + ln Σ e^z per row (with the max subtracted), and its _backward adds (P − Y) / N to the scores' grad directly. Check it against softmax_cross_entropy() on the 150 days, then compare how many Tensors each version puts in the graph.
- Which days does the trained softmax regression get wrong? Print a 3 × 3 table of true activity against predicted activity (a confusion matrix), and look at the inputs of the mistakes. How many of them are the noisy days, where people ignored the rule?
- Sample tomorrow's activity 1,000 times from softmax(z / T) for T = 0.5, 1 and 2 (np.random.default_rng().choice with p=...). Count how often each activity comes up. How does T change the mix, and what happens as T approaches 0?
7. Check yourself
In the catalog
Finished this chapter?