Measuring Error: Choosing a Loss Function
Training needs one number that says how wrong the model is. Which number you pick decides what 'good' means, from squared error for ride times to cross-entropy for probabilities.
🔍 Inspect Architecture: Logistic regression💡 1. Core Intuition & Concepts
In Chapter 1 we scored knob settings with 'average minutes off'. That was fine for a first try, but training needs a precise score, called the loss, and choosing it is a real design decision. The optimiser will do whatever it takes to push that number down, so the loss is your definition of what a good model is. Get it wrong and you get a model that is excellent at the wrong thing.
Start with the ride log and add the day you got a puncture: 20 km in 75 minutes. With knobs w = 3, b = 5 the misses are 0, −1, 1, −1 and −10 minutes. Mean absolute error goes from 0.75 to 2.6. Mean squared error jumps from 0.75 to 20.6, and that single ride is 97% of it. Squaring makes the model obsess over big misses. That's right if big misses really are costly, like a late ambulance, and wrong if they're flukes, like a puncture. The same split shows up in the best constant guess: squared error wants the mean (dragged to 35.4 by the 75), absolute error wants the median (28).
Now the picnic neuron. It outputs a score z, and we want a probability that tomorrow will be good, so we squash z with the sigmoid function into the range 0 to 1. How do we score a probability? Accuracy ('how many days did it call right?') is what we ultimately care about, but it's a staircase. In the code below, moving the bias from 0.3 to 0.6 leaves accuracy stuck at 100% while the model's confidence clearly changes. A staircase has no slope, and the next chapter needs slopes.
Cross-entropy is the standard answer. For each example, take the probability the model gave to what actually happened, and score −ln of it: your 'surprise'. A confident right answer costs almost nothing. A confident wrong answer costs a lot: saying 1% when the picnic was great costs 4.6, while squared error would charge only 0.98. It isn't an arbitrary choice either. The derivation below shows that minimising cross-entropy is the same as making the observed data as probable as possible.
📐 2. Mathematical Formulations & Derivations
L(θ) = (1/N) · Σᵢ ℓ(ŷᵢ, yᵢ)eᵢ = ŷᵢ − yᵢ
MAE = (1/N) Σ |eᵢ| MSE = (1/N) Σ eᵢ² RMSE = √MSEe = [0, −1, 1, −1, −10]
MAE = 13 / 5 = 2.6 MSE = (0 + 1 + 1 + 1 + 100) / 5 = 20.6 RMSE ≈ 4.54
share of MSE from the flat-tyre ride = 100 / 103 ≈ 97%argmin_c Σ (yᵢ − c)² = mean(y) argmin_c Σ |yᵢ − c| = median(y)
times [11, 21, 28, 42, 75]: mean = 35.4, median = 28σ(z) = 1 / (1 + e^(−z)) σ(0) = 0.5 σ(−z) = 1 − σ(z)
Monday: z = 0.8 ⇒ p = 1 / (1 + e^(−0.8)) = 0.690ℓ(p, y) = −[ y·ln p + (1 − y)·ln(1 − p) ] = −ln(probability given to what happened)
Monday (y = 1, p = 0.690): −ln 0.690 = 0.371 Tuesday (y = 0, p = 0.091): −ln 0.909 = 0.096
mean over the six recorded days = 0.3256p given to the true outcome: 0.9 0.5 0.1 0.01
squared error (1 − p)²: 0.010 0.250 0.810 0.980
cross-entropy −ln p: 0.105 0.693 2.303 4.605ℓ = −ln p_(correct class) guessing uniformly over K classes gives ℓ = ln K
K = 2: ln 2 ≈ 0.693 K = 10: ln 10 ≈ 2.303• For day i the model gives probability pᵢ that it'll be good. The probability it assigns to what actually happened is P(yᵢ) = pᵢ^yᵢ · (1 − pᵢ)^(1 − yᵢ). That's pᵢ when yᵢ = 1 and 1 − pᵢ when yᵢ = 0.
• Treating the days as independent, the probability the model assigns to the whole history is the product Λ(θ) = Πᵢ P(yᵢ). Seen as a function of the knobs θ, this is called the likelihood. Good knobs make what really happened look probable.
• A product of many numbers below 1 quickly underflows to 0 on a computer. Take the logarithm. It's increasing, so it doesn't move the maximum, and it turns the product into a sum: ln Λ(θ) = Σᵢ [ yᵢ ln pᵢ + (1 − yᵢ) ln(1 − pᵢ) ].
• Maximising ln Λ is the same as minimising −ln Λ, and dividing by N doesn't move the optimum either. What's left is (1/N) Σᵢ −[ yᵢ ln pᵢ + (1 − yᵢ) ln(1 − pᵢ) ], the mean binary cross-entropy.
• So cross-entropy isn't a convention someone picked. The knobs that minimise it are exactly the maximum-likelihood knobs. (The same argument with bell-curve noise on the ride times gives mean squared error.) ∎
⚙️ 3. Step-by-Step Computational Mechanism
💻 4. Code from Scratch (python)
# Chapter 3: one number for "how wrong". Regression losses first, then probabilities and cross-entropy.
import math
# --- regression: the ride log from Chapter 1, plus one flat-tyre day ------------
rides_km = [2, 5, 8, 12, 20]
rides_min = [11, 21, 28, 42, 75] # the 20 km ride had a puncture
def predict(km, w=3, b=5):
return w * km + b
def mae(errors): return sum(abs(e) for e in errors) / len(errors)
def mse(errors): return sum(e * e for e in errors) / len(errors)
def rmse(errors): return math.sqrt(mse(errors))
errors = [predict(x) - y for x, y in zip(rides_km, rides_min)]
print("errors:", errors)
for name, n in [("first 4 rides", 4), ("with flat tyre", 5)]:
e = errors[:n]
print(f"{name:>15}: MAE={mae(e):.3f} MSE={mse(e):.3f} RMSE={rmse(e):.3f}")
share = errors[4] ** 2 / sum(e * e for e in errors)
print(f"share of squared error from the flat-tyre ride: {share:.0%}")
# The best constant guess: mean for MSE, median for MAE
times = sorted(rides_min)
mean = sum(times) / len(times)
median = times[len(times) // 2]
print(f"\nbest constant under MSE = mean = {mean}, under MAE = median = {median}")
# --- classification: will the picnic go well? ------------------------------------
def sigmoid(z):
return 1 / (1 + math.exp(-z))
def bce(p, y):
"""Binary cross-entropy for one example: the 'surprise' at what actually happened."""
return -(y * math.log(p) + (1 - y) * math.log(1 - p))
w, b = [2.0, -3.0, -1.0], 0.5 # the picnic neuron from Chapter 2
history = [ # (sun, rain, wind) -> 1 = the picnic went well
([0.7, 0.2, 0.5], 1),
([0.1, 0.9, 0.3], 0),
([0.9, 0.0, 0.1], 1),
([0.4, 0.3, 0.9], 0),
([0.6, 0.1, 0.2], 1),
([0.5, 0.4, 0.4], 0),
]
print("\n z p(good) label loss")
total = 0.0
for x, y in history:
z = sum(wi * xi for wi, xi in zip(w, x)) + b
p = sigmoid(z)
total += bce(p, y)
print(f"{z:+.2f} {p:.3f} {y} {bce(p, y):.3f}")
print(f"mean cross-entropy: {total / len(history):.4f}")
# --- confident and wrong: cross-entropy vs squared error ---------------------------
print("\np(good) when it WAS good | squared error | cross-entropy")
for p in [0.9, 0.5, 0.1, 0.01]:
print(f"{p:>24} | {(1 - p) ** 2:>13.3f} | {bce(p, 1):>13.3f}")
# --- why not train on accuracy? Nudge the bias and watch both ----------------------
def scores(bias):
acc, ce = 0, 0.0
for x, y in history:
p = sigmoid(sum(wi * xi for wi, xi in zip(w, x)) + bias)
acc += int((p > 0.5) == bool(y))
ce += bce(p, y)
return acc / len(history), ce / len(history)
print("\nbias accuracy cross-entropy")
for bias in [0.3, 0.4, 0.5, 0.6, 0.7]:
acc, ce = scores(bias)
print(f"{bias:.1f} {acc:.3f} {ce:.4f}")
✏️ 5. Practice Problems
Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.
🛠️ 6. Try It Yourself
- Write a Huber loss: ½e² when |e| ≤ 2, otherwise 2·(|e| − 1). Compute it on the five-ride errors. How does the flat-tyre ride's share compare with its share of MSE and of MAE?
- Scan the picnic bias from −1 to 2 in steps of 0.1 and print mean cross-entropy and accuracy. Where is the cross-entropy lowest? Is accuracy still 100% there?
- Flip one label in `history` (say Wednesday becomes 0) and rerun. Which day now dominates the loss, and why that one?
🧠 7. Comprehension Checkpoint
In the catalog
Finished this chapter?