What a Model Is: A Function with Knobs
Forget brains and neurons for a moment. A model is an ordinary function whose behaviour is set by a few adjustable numbers, and learning means turning them.
🔍 Inspect Architecture: Linear regression💡 1. Core Intuition & Concepts
Say you cycle to work and want to guess how long a ride will take before you set off. You've written down four past rides: 2 km took 11 minutes, 5 km took 21, 8 km took 28, and 12 km took 42. A reasonable guess is that each kilometre costs some fixed number of minutes, plus a bit of overhead for unlocking the bike and waiting at the first light.
That guess is already a model: minutes = w × km + b. The distance is the input. The two numbers w (minutes per km) and b (overhead) are the knobs, which the literature calls parameters. Different knob settings give different models from the same family: w = 3, b = 5 predicts 11, 20, 29 and 41 minutes, which is close to what happened. w = 1, b = 0 predicts 2, 5, 8 and 12, which is badly wrong.
Every model in this course, up to the attention-based language model in the capstone, has this same shape: a function f(input; knobs) → prediction. What changes is how many knobs there are (2 here, billions in a large language model) and how they're wired together. Training is the job of finding knob values that make the predictions match the data.
In this chapter we turn the knobs by brute force: try lots of settings and keep the best. That works for 2 knobs and is hopeless for 20, and the reason it's hopeless is the reason the rest of the course exists.
📐 2. Mathematical Formulations & Derivations
ŷ = f(x; w, b) = w·x + bθ = (w, b) ∈ ℝ² e.g. θ₁ = (3, 5), θ₂ = (1, 0)[min/km] × [km] + [min] = [min]L₁(w, b) = (1/N) · Σᵢ | w·xᵢ + b − yᵢ |
L₁(3, 5) = ( |11−11| + |20−21| + |29−28| + |41−42| ) / 4 = (0 + 1 + 1 + 1) / 4 = 0.75L₂(w, b) = (1/N) · Σᵢ ( w·xᵢ + b − yᵢ )²rᵢ = yᵢ − w·xᵢ b* = r̄ = (1/N) Σᵢ rᵢ
w = 3: r = [11−6, 21−15, 28−24, 42−36] = [5, 6, 4, 6] ⇒ b* = 21/4 = 5.25settings to try = (values per knob) ^ (number of knobs)• Fix w and write the error in terms of the residuals rᵢ = yᵢ − w·xᵢ: L(b) = (1/N) Σ (w·xᵢ + b − yᵢ)² = (1/N) Σ (b − rᵢ)².
• Expand the square: (b − rᵢ)² = b² − 2b·rᵢ + rᵢ². Averaging over i gives L(b) = b² − 2b·r̄ + mean(r²), where r̄ is the mean residual.
• Complete the square: b² − 2b·r̄ = (b − r̄)² − r̄². So L(b) = (b − r̄)² + [ mean(r²) − r̄² ].
• The bracket doesn't involve b at all (it's the variance of the residuals). The first term is a square, so it's never negative, and it's zero exactly when b = r̄.
• Therefore L(b) is smallest at b* = r̄, and the smallest error left over is the variance of the residuals. For w = 3: b* = 5.25 and L = 28.25 − 5.25² = 0.6875. ∎ (No calculus needed. Chapter 4 reaches the same answer with derivatives.)
⚙️ 3. Step-by-Step Computational Mechanism
💻 4. Code from Scratch (python)
# Chapter 1: a model is a function with knobs.
# Pure Python only: no NumPy yet. Every line should be something you could do by hand.
rides_km = [2, 5, 8, 12] # inputs
rides_min = [11, 21, 28, 42] # what actually happened
def model(km, w, b):
"""Predict ride time: w minutes per km plus b minutes of overhead."""
return w * km + b
def mean_abs_error(w, b):
"""Average number of minutes our guesses are off by."""
total = 0.0
for km, actual in zip(rides_km, rides_min):
total += abs(model(km, w, b) - actual)
return total / len(rides_km)
# 1) Turn the knobs by hand and compare
for w, b in [(1, 0), (4, 0), (3, 5)]:
guesses = [model(km, w, b) for km in rides_km]
print(f"w={w}, b={b}: guesses {guesses} -> off by {mean_abs_error(w, b):.2f} min")
# 2) Brute force: try every w in 0.0..6.0 and every b in 0.0..10.0 (steps of 0.1)
best = None
tries = 0
for i in range(61):
for j in range(101):
w, b = i / 10, j / 10
err = mean_abs_error(w, b)
tries += 1
if best is None or err < best[0]:
best = (err, w, b)
err, w, b = best
print(f"\nTried {tries} settings. Best: w={w}, b={b}, off by {err:.2f} min on average")
# 3) Use the trained model on a ride we have never taken
print(f"Predicted time for 15 km: {model(15, w, b):.1f} min")
# 4) Why this can't scale: count the tries for bigger models
for knobs in [2, 10, 100]:
print(f"{knobs:>3} knobs x 10 values each -> {10 ** knobs:.1e} settings to try")
# 5) The algebra result: for a fixed w, the best b (squared error) is the mean residual
w = 3
residuals = [y - w * x for x, y in zip(rides_km, rides_min)]
b_star = sum(residuals) / len(residuals)
mse = sum((w * x + b_star - y) ** 2 for x, y in zip(rides_km, rides_min)) / len(rides_km)
print(f"\nw=3: residuals {residuals}, best b = {b_star}, mean squared error = {mse}")
✏️ 5. Practice Problems
Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.
🛠️ 6. Try It Yourself
- Add a fifth ride, 20 km in 75 minutes (you got a flat tyre). Rerun the search. How much do the best knobs move? Why does one bad ride pull the line so far?
- Change the model to minutes = w × km with no b. What's the best error now? What does that tell you about the value of the overhead knob?
- Make the grid 10× finer (steps of 0.01). How many tries does that take, and does the error improve much?
🧠 7. Comprehension Checkpoint
In the catalog
Finished this chapter?