Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
🧠 Core CourseLEVEL 1 · BEGINNERChapter 1

What a Model Is: A Function with Knobs

Forget brains and neurons for a moment. A model is an ordinary function whose behaviour is set by a few adjustable numbers, and learning means turning them.

⏱ 12 min read🎯 Prerequisites: Basic Python: variables, lists, for-loops
🔍 Inspect Architecture: Linear regression

💡 1. Core Intuition & Concepts

Say you cycle to work and want to guess how long a ride will take before you set off. You've written down four past rides: 2 km took 11 minutes, 5 km took 21, 8 km took 28, and 12 km took 42. A reasonable guess is that each kilometre costs some fixed number of minutes, plus a bit of overhead for unlocking the bike and waiting at the first light.

That guess is already a model: minutes = w × km + b. The distance is the input. The two numbers w (minutes per km) and b (overhead) are the knobs, which the literature calls parameters. Different knob settings give different models from the same family: w = 3, b = 5 predicts 11, 20, 29 and 41 minutes, which is close to what happened. w = 1, b = 0 predicts 2, 5, 8 and 12, which is badly wrong.

Every model in this course, up to the attention-based language model in the capstone, has this same shape: a function f(input; knobs) → prediction. What changes is how many knobs there are (2 here, billions in a large language model) and how they're wired together. Training is the job of finding knob values that make the predictions match the data.

In this chapter we turn the knobs by brute force: try lots of settings and keep the best. That works for 2 knobs and is hopeless for 20, and the reason it's hopeless is the reason the rest of the course exists.

📐 2. Mathematical Formulations & Derivations

The model: a function of the input AND the knobs
ŷ = f(x; w, b) = w·x + b
x is the input (km), ŷ ('y-hat') is the prediction (minutes). w and b are the knobs. The semicolon separates data, which you're given, from parameters, which you choose.
Parameter space: every point is a whole model
θ = (w, b) ∈ ℝ² e.g. θ₁ = (3, 5), θ₂ = (1, 0)
Collect the knobs into one vector θ ('theta'). Our model family is a 2-D plane of possible models, and training means searching that plane for a good point. A large language model does the same search in billions of dimensions.
Units keep you honest
[min/km] × [km] + [min] = [min]
w must be minutes per km and b must be minutes, or the sum is meaningless. Checking units is the fastest way to catch a formula written the wrong way round.
Scoring a setting: mean absolute error, worked out
L₁(w, b) = (1/N) · Σᵢ | w·xᵢ + b − yᵢ | L₁(3, 5) = ( |11−11| + |20−21| + |29−28| + |41−42| ) / 4 = (0 + 1 + 1 + 1) / 4 = 0.75
The score is itself a function, but of the knobs rather than of the input. Every point θ in parameter space gets a height L(θ), which gives a 'loss landscape' that training tries to walk down.
A second score: mean squared error
L₂(w, b) = (1/N) · Σᵢ ( w·xᵢ + b − yᵢ )²
Squaring makes big misses count much more than small ones (a 4-minute miss costs 16, not 4) and gives a smooth curve with no sharp corners. That smoothness is why most training uses it. Chapter 3 compares the two properly.
Residuals, and the best b for a fixed w
rᵢ = yᵢ − w·xᵢ b* = r̄ = (1/N) Σᵢ rᵢ w = 3: r = [11−6, 21−15, 28−24, 42−36] = [5, 6, 4, 6] ⇒ b* = 21/4 = 5.25
The residual is what's left of each answer after the slope has done its part. Under squared error, the best overhead is simply their average. The derivation below shows why.
Why brute force can't scale
settings to try = (values per knob) ^ (number of knobs)
With 61 × 101 values for our 2 knobs that's 6,161 tries, which takes milliseconds. Trying just 10 values each for a model with 100 knobs would take 10¹⁰⁰ tries, more than there are atoms in the observable universe.

• Fix w and write the error in terms of the residuals rᵢ = yᵢ − w·xᵢ: L(b) = (1/N) Σ (w·xᵢ + b − yᵢ)² = (1/N) Σ (b − rᵢ)².

• Expand the square: (b − rᵢ)² = b² − 2b·rᵢ + rᵢ². Averaging over i gives L(b) = b² − 2b·r̄ + mean(r²), where r̄ is the mean residual.

• Complete the square: b² − 2b·r̄ = (b − r̄)² − r̄². So L(b) = (b − r̄)² + [ mean(r²) − r̄² ].

• The bracket doesn't involve b at all (it's the variance of the residuals). The first term is a square, so it's never negative, and it's zero exactly when b = r̄.

• Therefore L(b) is smallest at b* = r̄, and the smallest error left over is the variance of the residuals. For w = 3: b* = 5.25 and L = 28.25 − 5.25² = 0.6875. ∎ (No calculus needed. Chapter 4 reaches the same answer with derivatives.)

⚙️ 3. Step-by-Step Computational Mechanism

1
Pick a family of functions
Decide on the model's shape before you look for knob values. Here it's a straight line, w·x + b. The family limits what the model can ever express: no setting of w and b will make a straight line curve.
2
Score a knob setting
Run the model on every example you have, compare with the true answers, and reduce the result to one number. A single score lets you rank two settings against each other.
3
Search for better knobs
Try many settings and keep the one with the lowest score. Here that means a grid over w and b. Later chapters replace this blind search with gradients, which point towards 'better' directly.
4
Use the trained model
Once the knobs are fixed, the model is an ordinary function again. Ask it about a ride you haven't taken yet (15 km, say) and it answers in microseconds.

💻 4. Code from Scratch (python)

core_ch1.py
# Chapter 1: a model is a function with knobs.
# Pure Python only: no NumPy yet. Every line should be something you could do by hand.

rides_km  = [2, 5, 8, 12]      # inputs
rides_min = [11, 21, 28, 42]   # what actually happened


def model(km, w, b):
    """Predict ride time: w minutes per km plus b minutes of overhead."""
    return w * km + b


def mean_abs_error(w, b):
    """Average number of minutes our guesses are off by."""
    total = 0.0
    for km, actual in zip(rides_km, rides_min):
        total += abs(model(km, w, b) - actual)
    return total / len(rides_km)


# 1) Turn the knobs by hand and compare
for w, b in [(1, 0), (4, 0), (3, 5)]:
    guesses = [model(km, w, b) for km in rides_km]
    print(f"w={w}, b={b}: guesses {guesses}  ->  off by {mean_abs_error(w, b):.2f} min")

# 2) Brute force: try every w in 0.0..6.0 and every b in 0.0..10.0 (steps of 0.1)
best = None
tries = 0
for i in range(61):
    for j in range(101):
        w, b = i / 10, j / 10
        err = mean_abs_error(w, b)
        tries += 1
        if best is None or err < best[0]:
            best = (err, w, b)

err, w, b = best
print(f"\nTried {tries} settings. Best: w={w}, b={b}, off by {err:.2f} min on average")

# 3) Use the trained model on a ride we have never taken
print(f"Predicted time for 15 km: {model(15, w, b):.1f} min")

# 4) Why this can't scale: count the tries for bigger models
for knobs in [2, 10, 100]:
    print(f"{knobs:>3} knobs x 10 values each -> {10 ** knobs:.1e} settings to try")

# 5) The algebra result: for a fixed w, the best b (squared error) is the mean residual
w = 3
residuals = [y - w * x for x, y in zip(rides_km, rides_min)]
b_star = sum(residuals) / len(residuals)
mse = sum((w * x + b_star - y) ** 2 for x, y in zip(rides_km, rides_min)) / len(rides_km)
print(f"\nw=3: residuals {residuals}, best b = {b_star}, mean squared error = {mse}")

✏️ 5. Practice Problems

Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.

P1 With knobs w = 2.5 and b = 6, how many minutes does the model predict for a 10 km ride?
💡 Plug into ŷ = w·x + b.
Solution. ŷ = 2.5 × 10 + 6 = 25 + 6 = 31 minutes.
P2 What is the mean absolute error of w = 4, b = 0 on the four recorded rides (2→11, 5→21, 8→28, 12→42)?
💡 Predictions are 8, 20, 32, 48. Take |prediction − actual| for each ride, then average.
Solution. Misses: |8−11| = 3, |20−21| = 1, |32−28| = 4, |48−42| = 6. Sum 14, divided by 4 rides = 3.5 minutes.
P3 Fix w = 3. Which b gives the smallest mean squared error on the rides?
💡 Compute the residuals yᵢ − 3·xᵢ and use the result from the derivation.
Solution. Residuals: 11−6 = 5, 21−15 = 6, 28−24 = 4, 42−36 = 6. Their mean is 21/4 = 5.25, so b* = 5.25.
P4 What is the mean squared error at w = 3, b = 5.25? (Give 4 decimal places.)
💡 Errors are b − rᵢ for residuals [5, 6, 4, 6]. Square them and average.
Solution. Errors: 0.25, −0.75, 1.25, −0.75. Squares: 0.0625 + 0.5625 + 1.5625 + 0.5625 = 2.75. Divided by 4 = 0.6875. That equals the variance of the residuals, as the derivation predicts.
P5 The model says w = 3 minutes per km. What is that in minutes per mile? (1 mile = 1.609 km; 2 decimal places.)
💡 A mile is longer than a km, so the number of minutes per mile should be bigger.
Solution. 3 min/km × 1.609 km/mile = 4.827 ≈ 4.83 min/mile. The knob's value depends on the units of the input, so rescaling the data changes the knobs.
P6 A model has 3 knobs and you try 11 values for each on a grid. How many settings must you score?
💡 Every value of knob 1 pairs with every value of knob 2 and of knob 3.
Solution. 11 × 11 × 11 = 11³ = 1331.

🛠️ 6. Try It Yourself

  1. Add a fifth ride, 20 km in 75 minutes (you got a flat tyre). Rerun the search. How much do the best knobs move? Why does one bad ride pull the line so far?
  2. Change the model to minutes = w × km with no b. What's the best error now? What does that tell you about the value of the overhead knob?
  3. Make the grid 10× finer (steps of 0.01). How many tries does that take, and does the error improve much?

🧠 7. Comprehension Checkpoint

Answer all 5 questions correctly to complete the chapter · 0 / 5 done
Q1/5 A model has 20 knobs and you try 10 candidate values for each one on a grid. How many settings must you score?
Q2/5 In ŷ = w·x + b trained on your ride log, which quantities does training change?
Q3/5 A knob setting has a mean absolute error of 0.75 on the rides. What does that tell you?
Q4/5 With w held fixed, which value of b minimises the mean squared error?
Q5/5 Your rides show a U-shape: very short and very long rides are both slow per km. What happens if you fit ŷ = w·x + b?

In the catalog

Finished this chapter?