Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
🧠 Core CourseLEVEL 1 · BEGINNERChapter 3

Measuring Error: Choosing a Loss Function

Training needs one number that says how wrong the model is. Which number you pick decides what 'good' means, from squared error for ride times to cross-entropy for probabilities.

⏱ 16 min read🎯 Prerequisites: Chapters 1–2; logarithms (ln) at a 'what does it do' level
🔍 Inspect Architecture: Logistic regression

💡 1. Core Intuition & Concepts

In Chapter 1 we scored knob settings with 'average minutes off'. That was fine for a first try, but training needs a precise score, called the loss, and choosing it is a real design decision. The optimiser will do whatever it takes to push that number down, so the loss is your definition of what a good model is. Get it wrong and you get a model that is excellent at the wrong thing.

Start with the ride log and add the day you got a puncture: 20 km in 75 minutes. With knobs w = 3, b = 5 the misses are 0, −1, 1, −1 and −10 minutes. Mean absolute error goes from 0.75 to 2.6. Mean squared error jumps from 0.75 to 20.6, and that single ride is 97% of it. Squaring makes the model obsess over big misses. That's right if big misses really are costly, like a late ambulance, and wrong if they're flukes, like a puncture. The same split shows up in the best constant guess: squared error wants the mean (dragged to 35.4 by the 75), absolute error wants the median (28).

Now the picnic neuron. It outputs a score z, and we want a probability that tomorrow will be good, so we squash z with the sigmoid function into the range 0 to 1. How do we score a probability? Accuracy ('how many days did it call right?') is what we ultimately care about, but it's a staircase. In the code below, moving the bias from 0.3 to 0.6 leaves accuracy stuck at 100% while the model's confidence clearly changes. A staircase has no slope, and the next chapter needs slopes.

Cross-entropy is the standard answer. For each example, take the probability the model gave to what actually happened, and score −ln of it: your 'surprise'. A confident right answer costs almost nothing. A confident wrong answer costs a lot: saying 1% when the picnic was great costs 4.6, while squared error would charge only 0.98. It isn't an arbitrary choice either. The derivation below shows that minimising cross-entropy is the same as making the observed data as probable as possible.

📐 2. Mathematical Formulations & Derivations

A loss is an average of per-example scores
L(θ) = (1/N) · Σᵢ ℓ(ŷᵢ, yᵢ)
ℓ ('little L') scores one prediction against its true answer. L averages over the data set and is what training minimises. The metric you report to people (accuracy, minutes late) can be different from the loss you train on.
Three regression losses
eᵢ = ŷᵢ − yᵢ MAE = (1/N) Σ |eᵢ| MSE = (1/N) Σ eᵢ² RMSE = √MSE
RMSE is back in the original units (minutes), which makes it easier to talk about. MSE is what you optimise, because its square has a clean slope everywhere.
Worked: one bad ride
e = [0, −1, 1, −1, −10] MAE = 13 / 5 = 2.6 MSE = (0 + 1 + 1 + 1 + 100) / 5 = 20.6 RMSE ≈ 4.54 share of MSE from the flat-tyre ride = 100 / 103 ≈ 97%
Before the puncture, both MAE and MSE were 0.75. One outlier multiplies MSE by 27 but MAE by only 3.5.
The best constant guess
argmin_c Σ (yᵢ − c)² = mean(y) argmin_c Σ |yᵢ − c| = median(y) times [11, 21, 28, 42, 75]: mean = 35.4, median = 28
'argmin_c' means 'the c that makes this smallest'. The mean result is Chapter 1's derivation. For the median, moving c past a data point changes which side is bigger, so the balance point splits the data in half.
Sigmoid: from score to probability
σ(z) = 1 / (1 + e^(−z)) σ(0) = 0.5 σ(−z) = 1 − σ(z) Monday: z = 0.8 ⇒ p = 1 / (1 + e^(−0.8)) = 0.690
Large positive z gives p close to 1, large negative z gives p close to 0. The neuron's decision boundary (z = 0) is exactly where the model is 50/50.
Binary cross-entropy (log loss)
ℓ(p, y) = −[ y·ln p + (1 − y)·ln(1 − p) ] = −ln(probability given to what happened) Monday (y = 1, p = 0.690): −ln 0.690 = 0.371 Tuesday (y = 0, p = 0.091): −ln 0.909 = 0.096 mean over the six recorded days = 0.3256
Only one of the two terms is ever active: y·ln p when the picnic went well, (1 − y)·ln(1 − p) when it didn't.
Confident and wrong
p given to the true outcome: 0.9 0.5 0.1 0.01 squared error (1 − p)²: 0.010 0.250 0.810 0.980 cross-entropy −ln p: 0.105 0.693 2.303 4.605
Squared error can never exceed 1 on a probability, so it barely distinguishes 'wrong' from 'confidently wrong'. Cross-entropy grows without limit as p → 0, which forces the model to stay humble when it isn't sure.
More than two classes, and a sanity baseline
ℓ = −ln p_(correct class) guessing uniformly over K classes gives ℓ = ln K K = 2: ln 2 ≈ 0.693 K = 10: ln 10 ≈ 2.303
A yes/no model whose mean cross-entropy is above 0.693 on balanced data is worse than a coin flip. Always compare your loss with a trivial baseline. Chapter 9 builds the multi-class version with softmax.

• For day i the model gives probability pᵢ that it'll be good. The probability it assigns to what actually happened is P(yᵢ) = pᵢ^yᵢ · (1 − pᵢ)^(1 − yᵢ). That's pᵢ when yᵢ = 1 and 1 − pᵢ when yᵢ = 0.

• Treating the days as independent, the probability the model assigns to the whole history is the product Λ(θ) = Πᵢ P(yᵢ). Seen as a function of the knobs θ, this is called the likelihood. Good knobs make what really happened look probable.

• A product of many numbers below 1 quickly underflows to 0 on a computer. Take the logarithm. It's increasing, so it doesn't move the maximum, and it turns the product into a sum: ln Λ(θ) = Σᵢ [ yᵢ ln pᵢ + (1 − yᵢ) ln(1 − pᵢ) ].

• Maximising ln Λ is the same as minimising −ln Λ, and dividing by N doesn't move the optimum either. What's left is (1/N) Σᵢ −[ yᵢ ln pᵢ + (1 − yᵢ) ln(1 − pᵢ) ], the mean binary cross-entropy.

• So cross-entropy isn't a convention someone picked. The knobs that minimise it are exactly the maximum-likelihood knobs. (The same argument with bell-curve noise on the ride times gives mean squared error.) ∎

⚙️ 3. Step-by-Step Computational Mechanism

1
Decide what a mistake costs
Is being 10 minutes late 10× as bad as 1 minute, or 100×? Is a confident wrong call worse than an unsure one? The loss should encode your answer, because the optimiser will take it literally.
2
Regression: start with MSE
It's smooth and well understood, and its best constant is the mean. If your data has flukes you don't want to chase, switch to MAE or a hybrid such as Huber loss (squared for small errors, absolute for big ones).
3
Classification: probabilities + cross-entropy
Squash the score into a probability (sigmoid for yes/no), then score −ln of the probability given to the true answer. Report accuracy to people, but train on cross-entropy.
4
Sanity-check the number
Compare with a trivial model: always predicting the mean ride time, or 50/50 for a balanced yes/no (loss ln 2 ≈ 0.693). If you can't beat the baseline, something upstream is wrong.

💻 4. Code from Scratch (python)

core_ch3.py
# Chapter 3: one number for "how wrong". Regression losses first, then probabilities and cross-entropy.
import math

# --- regression: the ride log from Chapter 1, plus one flat-tyre day ------------
rides_km  = [2, 5, 8, 12, 20]
rides_min = [11, 21, 28, 42, 75]            # the 20 km ride had a puncture


def predict(km, w=3, b=5):
    return w * km + b


def mae(errors):  return sum(abs(e) for e in errors) / len(errors)
def mse(errors):  return sum(e * e for e in errors) / len(errors)
def rmse(errors): return math.sqrt(mse(errors))


errors = [predict(x) - y for x, y in zip(rides_km, rides_min)]
print("errors:", errors)
for name, n in [("first 4 rides", 4), ("with flat tyre", 5)]:
    e = errors[:n]
    print(f"{name:>15}: MAE={mae(e):.3f}  MSE={mse(e):.3f}  RMSE={rmse(e):.3f}")
share = errors[4] ** 2 / sum(e * e for e in errors)
print(f"share of squared error from the flat-tyre ride: {share:.0%}")

# The best constant guess: mean for MSE, median for MAE
times = sorted(rides_min)
mean = sum(times) / len(times)
median = times[len(times) // 2]
print(f"\nbest constant under MSE = mean = {mean}, under MAE = median = {median}")

# --- classification: will the picnic go well? ------------------------------------
def sigmoid(z):
    return 1 / (1 + math.exp(-z))


def bce(p, y):
    """Binary cross-entropy for one example: the 'surprise' at what actually happened."""
    return -(y * math.log(p) + (1 - y) * math.log(1 - p))


w, b = [2.0, -3.0, -1.0], 0.5               # the picnic neuron from Chapter 2
history = [                                 # (sun, rain, wind) -> 1 = the picnic went well
    ([0.7, 0.2, 0.5], 1),
    ([0.1, 0.9, 0.3], 0),
    ([0.9, 0.0, 0.1], 1),
    ([0.4, 0.3, 0.9], 0),
    ([0.6, 0.1, 0.2], 1),
    ([0.5, 0.4, 0.4], 0),
]
print("\n  z       p(good)  label  loss")
total = 0.0
for x, y in history:
    z = sum(wi * xi for wi, xi in zip(w, x)) + b
    p = sigmoid(z)
    total += bce(p, y)
    print(f"{z:+.2f}    {p:.3f}     {y}     {bce(p, y):.3f}")
print(f"mean cross-entropy: {total / len(history):.4f}")

# --- confident and wrong: cross-entropy vs squared error ---------------------------
print("\np(good) when it WAS good | squared error | cross-entropy")
for p in [0.9, 0.5, 0.1, 0.01]:
    print(f"{p:>24} | {(1 - p) ** 2:>13.3f} | {bce(p, 1):>13.3f}")

# --- why not train on accuracy? Nudge the bias and watch both ----------------------
def scores(bias):
    acc, ce = 0, 0.0
    for x, y in history:
        p = sigmoid(sum(wi * xi for wi, xi in zip(w, x)) + bias)
        acc += int((p > 0.5) == bool(y))
        ce += bce(p, y)
    return acc / len(history), ce / len(history)

print("\nbias   accuracy   cross-entropy")
for bias in [0.3, 0.4, 0.5, 0.6, 0.7]:
    acc, ce = scores(bias)
    print(f"{bias:.1f}    {acc:.3f}      {ce:.4f}")

✏️ 5. Practice Problems

Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.

P1 Knobs w = 3, b = 5 predict [11, 20, 29, 41] for the rides that took [11, 21, 28, 42] minutes. What is the MSE?
💡 Errors are 0, −1, 1, −1. Square them and average.
Solution. (0 + 1 + 1 + 1) / 4 = 0.75. It equals the MAE here only because every error is 0 or ±1.
P2 What is the RMSE for those same four rides? (3 decimal places.)
💡 RMSE = √MSE.
Solution. √0.75 ≈ 0.866 minutes.
P3 Add the flat-tyre ride (20 km, 75 min; the model predicts 65). What is the MAE over all five rides?
💡 Errors: 0, −1, 1, −1, −10.
Solution. (0 + 1 + 1 + 1 + 10) / 5 = 13 / 5 = 2.6.
P4 And the MSE over all five rides?
💡 Square each error first. The 10-minute miss becomes 100.
Solution. (0 + 1 + 1 + 1 + 100) / 5 = 103 / 5 = 20.6. The one bad ride contributes 100 of the 103.
P5 What is σ(0.8), the picnic probability for Monday? (3 decimal places.)
💡 σ(z) = 1 / (1 + e^(−z)), and e^(−0.8) ≈ 0.4493.
Solution. 1 / (1 + 0.4493) = 1 / 1.4493 ≈ 0.690.
P6 The model says p = 0.8 and the picnic went well (y = 1). What is the cross-entropy loss? (Use ln; 3 decimal places.)
💡 With y = 1 only the −ln p term is active.
Solution. −ln 0.8 ≈ 0.223.
P7 Same prediction p = 0.8, but the picnic was rained off (y = 0). What is the loss now? (3 decimal places.)
💡 With y = 0 the model gave probability 1 − p to what happened.
Solution. −ln(1 − 0.8) = −ln 0.2 ≈ 1.609, about 7× the loss of the previous problem for the same prediction.
P8 A lazy model always says p = 0.5 on a yes/no problem. What is its mean cross-entropy? (3 decimal places.)
💡 Every example costs the same, whichever label it has.
Solution. −ln 0.5 = ln 2 ≈ 0.693 on every example, so the mean is 0.693. That's the coin-flip baseline to beat.

🛠️ 6. Try It Yourself

  1. Write a Huber loss: ½e² when |e| ≤ 2, otherwise 2·(|e| − 1). Compute it on the five-ride errors. How does the flat-tyre ride's share compare with its share of MSE and of MAE?
  2. Scan the picnic bias from −1 to 2 in steps of 0.1 and print mean cross-entropy and accuracy. Where is the cross-entropy lowest? Is accuracy still 100% there?
  3. Flip one label in `history` (say Wednesday becomes 0) and rerun. Which day now dominates the loss, and why that one?

🧠 7. Comprehension Checkpoint

Answer all 5 questions correctly to complete the chapter · 0 / 5 done
Q1/5 Why does one outlier hurt mean squared error much more than mean absolute error?
Q2/5 You must predict every ride with the same constant c and you're scored with MAE. Which c is best?
Q3/5 On a day that turned out great (y = 1), the model said p = 0.01. How do the two losses react?
Q4/5 Why don't we train the picnic model directly on accuracy?
Q5/5 Minimising mean binary cross-entropy is equivalent to…

Finished this chapter?