Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
🧠 Core CourseLEVEL 2 · INTERMEDIATEChapter 4

Which Way Is Downhill? Derivatives by Hand and by Nudging

The derivative answers the question brute force couldn't: if I turn this knob a little, does the error go up or down, and how fast? Derive it with the chain rule, then check it by nudging.

⏱ 20 min read🎯 Prerequisites: Chapters 1–3; high-school algebra (no calculus assumed)
🔍 Inspect Architecture: Linear regression

💡 1. Core Intuition & Concepts

Chapter 1 ended with a wall: searching every knob setting needs 10¹⁰⁰ tries for a modest model. The way around it is to stop searching and start asking. Stand at the current setting and ask each knob: if I nudge you up a tiny bit, does the loss go up or down, and by how much? That question has a precise answer, the derivative. Put every knob's answer together and you get the gradient, which points in the steepest uphill direction. Walk the opposite way and the loss drops, without trying any other settings.

Take the ride model at w = 3, b = 5 with mean squared error 0.75. Its derivatives are ∂L/∂w = −4.5 and ∂L/∂b = −0.5. Both are negative, so raising either knob lowers the loss, and right now w matters 9× more than b. One small step against the gradient, of size 0.01, takes the knobs to (3.045, 5.005) and the loss from 0.75 to 0.668. That's one step and zero search.

There are two ways to get a derivative. By hand, you apply a few rules (exact and fast, but easy to get wrong). Numerically, you literally nudge the knob and measure the change (always works, but it's approximate and needs two loss evaluations per knob, far too slow for millions of knobs). Practitioners use both: derive by hand, or let an autograd engine do it, then check against the numerical version. That check is called a gradient check and it catches most bugs.

The star of the chapter is the chain rule. Every neural network is functions nested inside functions: a loss of a sigmoid of a dot product. The chain rule says the slope of the whole chain is the product of each link's local slope. Chapter 5 turns exactly that sentence into about 50 lines of code.

📐 2. Mathematical Formulations & Derivations

The derivative: slope of a tiny step
f′(x) = lim_(h→0) ( f(x + h) − f(x) ) / h for small h: f(x + h) ≈ f(x) + f′(x) · h
Read the second line as 'nudge the input by h and the output moves by about f′(x) times h'. The derivative is a local exchange rate between input and output.
The handful of rules you need
(c)′ = 0 (xⁿ)′ = n·xⁿ⁻¹ (c·f)′ = c·f′ (f + g)′ = f′ + g′ (eˣ)′ = eˣ (ln x)′ = 1/x example: f(x) = 3x² + 2x ⇒ f′(x) = 6x + 2, f′(2) = 14
With these plus the chain rule you can differentiate every loss in this course.
The chain rule
d/dx f(g(x)) = f′(g(x)) · g′(x) in slope notation: dL/dw = (dL/de) · (de/dw) example: d/dx (2x + 1)³ = 3(2x + 1)² · 2 at x = 1: 3·9·2 = 54
Outer slope (evaluated at the inner value) times inner slope. For longer chains, keep multiplying one factor per link.
Partial derivatives and the gradient
∂L/∂w: slope in w with b held fixed ∇L = ( ∂L/∂w , ∂L/∂b ) −∇L is the direction of steepest descent
With many knobs, differentiate with respect to one while treating the others as constants. The vector of all those slopes is the gradient ('nabla L').
Gradient of MSE for the line (derived below)
eᵢ = w·xᵢ + b − yᵢ ∂L/∂w = (2/N) Σ eᵢ·xᵢ ∂L/∂b = (2/N) Σ eᵢ at (3, 5): e = [0, −1, 1, −1] ∂L/∂b = ½ · (−1) = −0.5 ∂L/∂w = ½ · (0·2 − 1·5 + 1·8 − 1·12) = ½ · (−9) = −4.5
Each ride pushes the slope w in proportion to its error times its distance, so long rides have more leverage on the slope than short ones.
One step downhill
θ_new = θ − η · ∇L η = 0.01: (3, 5) → (3 + 0.045, 5 + 0.005) = (3.045, 5.005) loss: 0.750 → 0.668
η ('eta') is the step size, also called the learning rate. Repeating this step is gradient descent, the subject of Chapter 7.
A minimum has zero gradient, and that recovers Chapter 1
∂L/∂b = (2/N) Σ (w·xᵢ + b − yᵢ) = 0 ⇒ b = (1/N) Σ (yᵢ − w·xᵢ) = mean residual w = 3: b* = 5.25
At the bottom of a smooth valley, no small nudge helps, so every slope is zero. Setting ∂L/∂b = 0 gives the completing-the-square result in one line.
Numerical derivatives: forward vs central
forward: ( f(x + h) − f(x) ) / h error ≈ ½·|f″(x)|·h central: ( f(x + h) − f(x − h) ) / (2h) error ≈ (1/6)·|f‴(x)|·h² sin at x = 1, h = 0.01: forward error 4.2e−3, central error 9.0e−6
Shrinking h by 10× cuts the forward error by 10× but the central error by 100×. See the second derivation for why.
Gradient check
relative error = | g_hand − g_numeric | / ( |g_hand| + |g_numeric| ) < 1e−7: fine ~ 1e−4: suspicious > 1e−2: a bug
Our hand-derived line gradient agrees with central differences to about 1e−10. Forgetting the factor 2 would give a relative error of 0.33.
Two derivatives you'll reuse constantly
σ′(z) = σ(z) · (1 − σ(z)) at z = 0.8: 0.690 · 0.310 = 0.2139 for ℓ = cross-entropy(σ(z), y): ∂ℓ/∂z = (−y/p + (1 − y)/(1 − p)) · p(1 − p) = p − y
The sigmoid's slope is largest at z = 0 (0.25) and vanishes for large |z|. Combined with cross-entropy, all the messy terms cancel: the gradient is just 'predicted minus actual'.

• Write the loss as nested functions: L = (1/N) Σᵢ eᵢ², with eᵢ = w·xᵢ + b − yᵢ.

• Sum rule and constant rule: the derivative of an average is the average of the derivatives, so ∂L/∂w = (1/N) Σᵢ ∂(eᵢ²)/∂w.

• Chain rule on one term: the outer function is u ↦ u², with slope 2u, and the inner function is eᵢ. So ∂(eᵢ²)/∂w = 2eᵢ · ∂eᵢ/∂w.

• Inner slopes: in eᵢ = w·xᵢ + b − yᵢ, the numbers xᵢ, b and yᵢ are constants when we vary w, so ∂eᵢ/∂w = xᵢ. Varying b instead gives ∂eᵢ/∂b = 1.

• Put it together: ∂L/∂w = (2/N) Σᵢ eᵢ·xᵢ and ∂L/∂b = (2/N) Σᵢ eᵢ. ∎ Every gradient in this course is built from these same moves: split the loss into links, take each link's local slope, and multiply along the chain.

• Near x, any smooth function can be expanded as f(x + h) = f(x) + h·f′(x) + (h²/2)·f″(x) + (h³/6)·f‴(x) + …

• Forward difference: subtract f(x) and divide by h to get f′(x) + (h/2)·f″(x) + … The leftover error is proportional to h. For sin at x = 1, h = 0.01: (0.01/2)·sin(1) ≈ 4.2e−3, matching the table.

• Expand the other side too: f(x − h) = f(x) − h·f′(x) + (h²/2)·f″(x) − (h³/6)·f‴(x) + … Subtracting it from f(x + h), the f(x) and f″ terms cancel: f(x + h) − f(x − h) = 2h·f′(x) + (h³/3)·f‴(x) + …

• Divide by 2h: f′(x) + (h²/6)·f‴(x) + … The error is now proportional to h². For sin at 1, h = 0.01: (0.0001/6)·cos(1) ≈ 9.0e−6, matching the table.

• Why not use h = 1e−15? A computer stores about 16 significant digits. f(x + h) − f(x) subtracts two nearly equal numbers and loses most of them, giving a rounding error of roughly 1e−16 / h that grows as h shrinks. The table shows it taking over below h ≈ 1e−8. For central differences, h ≈ 1e−5 is a good default. ∎

⚙️ 3. Step-by-Step Computational Mechanism

1
Write the loss as a chain
Name every intermediate quantity: error eᵢ = w·xᵢ + b − yᵢ, then squared error, then the average. Each link should be a function simple enough to differentiate on sight.
2
Take local slopes
Differentiate each link with respect to its own input: d(u²)/du = 2u, ∂e/∂w = x, ∂e/∂b = 1. No link needs to know about the others.
3
Multiply along the chain
The slope of the loss with respect to a knob is the product of the local slopes on the path from knob to loss. Where several paths exist (one per example), add their contributions.
4
Check numerically
Nudge each knob up and down by h ≈ 1e−5, compute the central difference, and compare with your formula using the relative error. Do this once whenever you write a new gradient.
5
Step against the gradient
Subtract a small multiple of the gradient from the knobs. The loss should go down. If it goes up, the step is too big or the sign is wrong.

💻 4. Code from Scratch (python)

core_ch4.py
# Chapter 4: which way is downhill? Derivatives by hand, checked numerically.
import math

rides_km  = [2, 5, 8, 12]
rides_min = [11, 21, 28, 42]
N = len(rides_km)


def loss(w, b):
    """Mean squared error of the line w*km + b on the ride log."""
    return sum((w * x + b - y) ** 2 for x, y in zip(rides_km, rides_min)) / N


def grad_by_hand(w, b):
    """dL/dw and dL/db, derived with the chain rule (see the maths section)."""
    errors = [w * x + b - y for x, y in zip(rides_km, rides_min)]
    dw = 2 / N * sum(e * x for e, x in zip(errors, rides_km))
    db = 2 / N * sum(errors)
    return dw, db


def grad_numeric(w, b, h=1e-5):
    """Central differences: nudge each knob up and down, measure the change."""
    dw = (loss(w + h, b) - loss(w - h, b)) / (2 * h)
    db = (loss(w, b + h) - loss(w, b - h)) / (2 * h)
    return dw, db


w, b = 3.0, 5.0
print(f"loss at w={w}, b={b}: {loss(w, b)}")
print("by hand :", [round(g, 6) for g in grad_by_hand(w, b)])
print("numeric :", [round(g, 6) for g in grad_numeric(w, b)])

# --- a gradient check, the way you'll use it in every later chapter -------------
def rel_error(a, b):
    return abs(a - b) / max(1e-12, abs(a) + abs(b))

for point in [(3.0, 5.0), (1.0, 0.0), (3.3, 4.1)]:
    exact, approx = grad_by_hand(*point), grad_numeric(*point)
    worst = max(rel_error(e, a) for e, a in zip(exact, approx))
    print(f"gradient check at {point}: worst relative error {worst:.1e}")

# --- forward vs central difference, and what happens when h gets too small ------
f = math.sin            # true derivative at x = 1 is cos(1)
x, true = 1.0, math.cos(1.0)
print("\n      h        forward error   central error")
for h in [1e-1, 1e-2, 1e-4, 1e-6, 1e-8, 1e-10, 1e-12]:
    fwd = (f(x + h) - f(x)) / h
    cen = (f(x + h) - f(x - h)) / (2 * h)
    print(f"{h:>8.0e}     {abs(fwd - true):.2e}        {abs(cen - true):.2e}")

# --- the gradient points uphill, so one small step against it lowers the loss -----
lr = 0.01
dw, db = grad_by_hand(w, b)
w2, b2 = w - lr * dw, b - lr * db
print(f"\none step downhill: ({w}, {b}) -> ({w2:.4f}, {b2:.4f}), loss {loss(w, b):.4f} -> {loss(w2, b2):.4f}")

# --- setting dL/db = 0 recovers Chapter 1's answer -------------------------------
w = 3.0
b_star = sum(y - w * x for x, y in zip(rides_km, rides_min)) / N
print(f"at w=3, dL/db = 0 when b = {b_star}; check: dL/db there = {grad_by_hand(w, b_star)[1]:.1e}")

# --- sigmoid: its derivative is s(1 - s) -------------------------------------------
def sigmoid(z): return 1 / (1 + math.exp(-z))
z = 0.8
s = sigmoid(z)
numeric = (sigmoid(z + 1e-5) - sigmoid(z - 1e-5)) / 2e-5
print(f"\nsigmoid'(0.8): formula s(1-s) = {s * (1 - s):.6f}, numeric = {numeric:.6f}")

✏️ 5. Practice Problems

Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.

P1 f(x) = 3x² + 2x. What is f′(2)?
💡 Use the power rule on each term: f′(x) = 6x + 2.
Solution. f′(x) = 6x + 2, so f′(2) = 12 + 2 = 14.
P2 What is the derivative of (2x + 1)³ at x = 1?
💡 Chain rule: outer u³ has slope 3u², inner 2x + 1 has slope 2.
Solution. 3(2x + 1)² · 2. At x = 1: 3 · 3² · 2 = 3 · 9 · 2 = 54.
P3 For the ride model at w = 3, b = 5 (errors e = [0, −1, 1, −1]), what is ∂L/∂b for MSE?
💡 ∂L/∂b = (2/N) Σ eᵢ with N = 4.
Solution. Σ eᵢ = −1, so ∂L/∂b = (2/4)·(−1) = −0.5.
P4 Same point: what is ∂L/∂w? (Distances x = [2, 5, 8, 12].)
💡 ∂L/∂w = (2/N) Σ eᵢ·xᵢ.
Solution. Σ eᵢxᵢ = 0·2 − 1·5 + 1·8 − 1·12 = −9, so ∂L/∂w = (2/4)·(−9) = −4.5.
P5 Estimate the derivative of f(x) = x² at x = 3 with a forward difference and h = 0.1.
💡 (f(3.1) − f(3)) / 0.1
Solution. (9.61 − 9) / 0.1 = 6.1. The true value is 6, so the error is 0.1 = (h/2)·f″ = 0.05 · 2, as the Taylor argument predicts.
P6 Now use a central difference with h = 0.1 for the same f(x) = x² at x = 3.
💡 (f(3.1) − f(2.9)) / 0.2
Solution. (9.61 − 8.41) / 0.2 = 6.0, exactly right. For a quadratic f‴ = 0, so the central difference has no error term at all.
P7 What is σ′(0), the slope of the sigmoid at z = 0?
💡 σ′(z) = σ(z)(1 − σ(z)) and σ(0) = 0.5.
Solution. 0.5 · 0.5 = 0.25. That's the sigmoid's steepest point; its slope only shrinks as |z| grows.
P8 Starting at (w, b) = (3, 5) with gradient (−4.5, −0.5), take one step with η = 0.01. What is the new w?
💡 w_new = w − η · ∂L/∂w.
Solution. 3 − 0.01 · (−4.5) = 3 + 0.045 = 3.045. Minus a negative slope means w goes up.

🛠️ 6. Try It Yourself

  1. Derive the gradient of mean absolute error for the line (the slope of |e| is +1 for e > 0 and −1 for e < 0) and check it numerically at (3.3, 4.1). Then try (3, 5), where one ride has error exactly 0. What does the central difference report there?
  2. Break the gradient on purpose: drop the factor 2 in grad_by_hand and rerun the gradient check. What relative error do you get, and would you have noticed without the check?
  3. Put the downhill step in a loop: 2,000 steps with η = 0.01, printing the loss every 200. Where do w and b end up? Compare with Chapter 1's grid search (w = 3.1, b = 4.8 under MAE). Why might they differ?

🧠 7. Comprehension Checkpoint

Answer all 5 questions correctly to complete the chapter · 0 / 5 done
Q1/5 At w = 3, b = 5 you compute ∂L/∂w = −4.5. What does that tell you?
Q2/5 Why is the central difference usually preferred over the forward difference?
Q3/5 Why not make h as small as 1e−15 for a very accurate numerical derivative?
Q4/5 What does the chain rule give for d/dx f(g(x))?
Q5/5 At the minimum of a smooth loss, what is the gradient?

Finished this chapter?