Which Way Is Downhill? Derivatives by Hand and by Nudging
The derivative answers the question brute force couldn't: if I turn this knob a little, does the error go up or down, and how fast? Derive it with the chain rule, then check it by nudging.
🔍 Inspect Architecture: Linear regression💡 1. Core Intuition & Concepts
Chapter 1 ended with a wall: searching every knob setting needs 10¹⁰⁰ tries for a modest model. The way around it is to stop searching and start asking. Stand at the current setting and ask each knob: if I nudge you up a tiny bit, does the loss go up or down, and by how much? That question has a precise answer, the derivative. Put every knob's answer together and you get the gradient, which points in the steepest uphill direction. Walk the opposite way and the loss drops, without trying any other settings.
Take the ride model at w = 3, b = 5 with mean squared error 0.75. Its derivatives are ∂L/∂w = −4.5 and ∂L/∂b = −0.5. Both are negative, so raising either knob lowers the loss, and right now w matters 9× more than b. One small step against the gradient, of size 0.01, takes the knobs to (3.045, 5.005) and the loss from 0.75 to 0.668. That's one step and zero search.
There are two ways to get a derivative. By hand, you apply a few rules (exact and fast, but easy to get wrong). Numerically, you literally nudge the knob and measure the change (always works, but it's approximate and needs two loss evaluations per knob, far too slow for millions of knobs). Practitioners use both: derive by hand, or let an autograd engine do it, then check against the numerical version. That check is called a gradient check and it catches most bugs.
The star of the chapter is the chain rule. Every neural network is functions nested inside functions: a loss of a sigmoid of a dot product. The chain rule says the slope of the whole chain is the product of each link's local slope. Chapter 5 turns exactly that sentence into about 50 lines of code.
📐 2. Mathematical Formulations & Derivations
f′(x) = lim_(h→0) ( f(x + h) − f(x) ) / h
for small h: f(x + h) ≈ f(x) + f′(x) · h(c)′ = 0 (xⁿ)′ = n·xⁿ⁻¹ (c·f)′ = c·f′ (f + g)′ = f′ + g′ (eˣ)′ = eˣ (ln x)′ = 1/x
example: f(x) = 3x² + 2x ⇒ f′(x) = 6x + 2, f′(2) = 14d/dx f(g(x)) = f′(g(x)) · g′(x) in slope notation: dL/dw = (dL/de) · (de/dw)
example: d/dx (2x + 1)³ = 3(2x + 1)² · 2 at x = 1: 3·9·2 = 54∂L/∂w: slope in w with b held fixed ∇L = ( ∂L/∂w , ∂L/∂b )
−∇L is the direction of steepest descenteᵢ = w·xᵢ + b − yᵢ ∂L/∂w = (2/N) Σ eᵢ·xᵢ ∂L/∂b = (2/N) Σ eᵢ
at (3, 5): e = [0, −1, 1, −1]
∂L/∂b = ½ · (−1) = −0.5 ∂L/∂w = ½ · (0·2 − 1·5 + 1·8 − 1·12) = ½ · (−9) = −4.5θ_new = θ − η · ∇L η = 0.01: (3, 5) → (3 + 0.045, 5 + 0.005) = (3.045, 5.005)
loss: 0.750 → 0.668∂L/∂b = (2/N) Σ (w·xᵢ + b − yᵢ) = 0 ⇒ b = (1/N) Σ (yᵢ − w·xᵢ) = mean residual
w = 3: b* = 5.25forward: ( f(x + h) − f(x) ) / h error ≈ ½·|f″(x)|·h
central: ( f(x + h) − f(x − h) ) / (2h) error ≈ (1/6)·|f‴(x)|·h²
sin at x = 1, h = 0.01: forward error 4.2e−3, central error 9.0e−6relative error = | g_hand − g_numeric | / ( |g_hand| + |g_numeric| )
< 1e−7: fine ~ 1e−4: suspicious > 1e−2: a bugσ′(z) = σ(z) · (1 − σ(z)) at z = 0.8: 0.690 · 0.310 = 0.2139
for ℓ = cross-entropy(σ(z), y): ∂ℓ/∂z = (−y/p + (1 − y)/(1 − p)) · p(1 − p) = p − y• Write the loss as nested functions: L = (1/N) Σᵢ eᵢ², with eᵢ = w·xᵢ + b − yᵢ.
• Sum rule and constant rule: the derivative of an average is the average of the derivatives, so ∂L/∂w = (1/N) Σᵢ ∂(eᵢ²)/∂w.
• Chain rule on one term: the outer function is u ↦ u², with slope 2u, and the inner function is eᵢ. So ∂(eᵢ²)/∂w = 2eᵢ · ∂eᵢ/∂w.
• Inner slopes: in eᵢ = w·xᵢ + b − yᵢ, the numbers xᵢ, b and yᵢ are constants when we vary w, so ∂eᵢ/∂w = xᵢ. Varying b instead gives ∂eᵢ/∂b = 1.
• Put it together: ∂L/∂w = (2/N) Σᵢ eᵢ·xᵢ and ∂L/∂b = (2/N) Σᵢ eᵢ. ∎ Every gradient in this course is built from these same moves: split the loss into links, take each link's local slope, and multiply along the chain.
• Near x, any smooth function can be expanded as f(x + h) = f(x) + h·f′(x) + (h²/2)·f″(x) + (h³/6)·f‴(x) + …
• Forward difference: subtract f(x) and divide by h to get f′(x) + (h/2)·f″(x) + … The leftover error is proportional to h. For sin at x = 1, h = 0.01: (0.01/2)·sin(1) ≈ 4.2e−3, matching the table.
• Expand the other side too: f(x − h) = f(x) − h·f′(x) + (h²/2)·f″(x) − (h³/6)·f‴(x) + … Subtracting it from f(x + h), the f(x) and f″ terms cancel: f(x + h) − f(x − h) = 2h·f′(x) + (h³/3)·f‴(x) + …
• Divide by 2h: f′(x) + (h²/6)·f‴(x) + … The error is now proportional to h². For sin at 1, h = 0.01: (0.0001/6)·cos(1) ≈ 9.0e−6, matching the table.
• Why not use h = 1e−15? A computer stores about 16 significant digits. f(x + h) − f(x) subtracts two nearly equal numbers and loses most of them, giving a rounding error of roughly 1e−16 / h that grows as h shrinks. The table shows it taking over below h ≈ 1e−8. For central differences, h ≈ 1e−5 is a good default. ∎
⚙️ 3. Step-by-Step Computational Mechanism
💻 4. Code from Scratch (python)
# Chapter 4: which way is downhill? Derivatives by hand, checked numerically.
import math
rides_km = [2, 5, 8, 12]
rides_min = [11, 21, 28, 42]
N = len(rides_km)
def loss(w, b):
"""Mean squared error of the line w*km + b on the ride log."""
return sum((w * x + b - y) ** 2 for x, y in zip(rides_km, rides_min)) / N
def grad_by_hand(w, b):
"""dL/dw and dL/db, derived with the chain rule (see the maths section)."""
errors = [w * x + b - y for x, y in zip(rides_km, rides_min)]
dw = 2 / N * sum(e * x for e, x in zip(errors, rides_km))
db = 2 / N * sum(errors)
return dw, db
def grad_numeric(w, b, h=1e-5):
"""Central differences: nudge each knob up and down, measure the change."""
dw = (loss(w + h, b) - loss(w - h, b)) / (2 * h)
db = (loss(w, b + h) - loss(w, b - h)) / (2 * h)
return dw, db
w, b = 3.0, 5.0
print(f"loss at w={w}, b={b}: {loss(w, b)}")
print("by hand :", [round(g, 6) for g in grad_by_hand(w, b)])
print("numeric :", [round(g, 6) for g in grad_numeric(w, b)])
# --- a gradient check, the way you'll use it in every later chapter -------------
def rel_error(a, b):
return abs(a - b) / max(1e-12, abs(a) + abs(b))
for point in [(3.0, 5.0), (1.0, 0.0), (3.3, 4.1)]:
exact, approx = grad_by_hand(*point), grad_numeric(*point)
worst = max(rel_error(e, a) for e, a in zip(exact, approx))
print(f"gradient check at {point}: worst relative error {worst:.1e}")
# --- forward vs central difference, and what happens when h gets too small ------
f = math.sin # true derivative at x = 1 is cos(1)
x, true = 1.0, math.cos(1.0)
print("\n h forward error central error")
for h in [1e-1, 1e-2, 1e-4, 1e-6, 1e-8, 1e-10, 1e-12]:
fwd = (f(x + h) - f(x)) / h
cen = (f(x + h) - f(x - h)) / (2 * h)
print(f"{h:>8.0e} {abs(fwd - true):.2e} {abs(cen - true):.2e}")
# --- the gradient points uphill, so one small step against it lowers the loss -----
lr = 0.01
dw, db = grad_by_hand(w, b)
w2, b2 = w - lr * dw, b - lr * db
print(f"\none step downhill: ({w}, {b}) -> ({w2:.4f}, {b2:.4f}), loss {loss(w, b):.4f} -> {loss(w2, b2):.4f}")
# --- setting dL/db = 0 recovers Chapter 1's answer -------------------------------
w = 3.0
b_star = sum(y - w * x for x, y in zip(rides_km, rides_min)) / N
print(f"at w=3, dL/db = 0 when b = {b_star}; check: dL/db there = {grad_by_hand(w, b_star)[1]:.1e}")
# --- sigmoid: its derivative is s(1 - s) -------------------------------------------
def sigmoid(z): return 1 / (1 + math.exp(-z))
z = 0.8
s = sigmoid(z)
numeric = (sigmoid(z + 1e-5) - sigmoid(z - 1e-5)) / 2e-5
print(f"\nsigmoid'(0.8): formula s(1-s) = {s * (1 - s):.6f}, numeric = {numeric:.6f}")
✏️ 5. Practice Problems
Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.
🛠️ 6. Try It Yourself
- Derive the gradient of mean absolute error for the line (the slope of |e| is +1 for e > 0 and −1 for e < 0) and check it numerically at (3.3, 4.1). Then try (3, 5), where one ride has error exactly 0. What does the central difference report there?
- Break the gradient on purpose: drop the factor 2 in grad_by_hand and rerun the gradient check. What relative error do you get, and would you have noticed without the check?
- Put the downhill step in a loop: 2,000 steps with η = 0.01, printing the loss every 200. Where do w and b end up? Compare with Chapter 1's grid search (w = 3.1, b = 4.8 under MAE). Why might they differ?
🧠 7. Comprehension Checkpoint
In the catalog
Finished this chapter?