Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

πŸŽ“ Track Progress: 0 Completed
0%
πŸ’» Modern ArchitecturesLEVEL 2 Β· INTERMEDIATEChapter 4

Multi-Layer Perceptrons & Analytical Backpropagation

Hidden representations, non-linear activations (ReLU/GELU), and matrix chain rule.

⏱ 18 min read🎯 Prerequisites: Matrix Multiplication & Partial Derivatives
πŸ” Inspect Architecture: Multilayer perceptron (MLP)

πŸ’‘ 1. Core Intuition & Concepts

Multi-Layer Perceptrons (MLPs) overcome the linear separability limitations of single perceptrons by stacking linear transformations with non-linear activation functions. Analytical backpropagation uses the multivariate calculus chain rule to propagate error gradients backwards from the scalar loss through intermediate hidden layers.

πŸ“ 2. Mathematical Formulations & Derivations

Forward Pass Equation
z[l]=a[lβˆ’1]W[l]+b[l]a[l]=g[l](z[l])z^{[l]} = a^{[l-1]} W^{[l]} + b^{[l]} \qquad a^{[l]} = g^{[l]}(z^{[l]})
l: layer index, W^[l]: weight matrix (D_{in}Γ—D_{out}), g: activation function.
Output Error Delta
Ξ΄[L]=βˆ‡a[L]LβŠ™gβ€²(z[L])=(y^βˆ’y)βŠ™gβ€²(z[L])\delta^{[L]} = \nabla_{a^{[L]}}\mathcal{L} \odot g'(z^{[L]}) = (\hat{y} - y) \odot g'(z^{[L]})
Hadamard element-wise product between loss derivative and activation slope.
Hidden Layer Backprop Chain Rule
Ξ΄[l]=(Ξ΄[l+1](W[l+1])T)βŠ™gβ€²(z[l])\delta^{[l]} = \left( \delta^{[l+1]} (W^{[l+1]})^T \right) \odot g'(z^{[l]})
Transposed weight projection routes downstream errors back through the network.

βš™οΈ 3. Step-by-Step Computational Mechanism

1
Forward Activation
Compute intermediate layer activations z^[1] = X W^[1] + b^[1], a^[1] = ReLU(z^[1]).
2
Loss Evaluation
Compute scalar objective L = BCE(y_hat, y) or MSE(y_hat, y).
3
Backward Gradient Delta
Compute output error delta^[2] = y_hat - y.
4
Hidden Gradients
Calculate dW^[2] = (a^[1])^T delta^[2], delta^[1] = (delta^[2] (W^[2])^T) * ReLU'(z^[1]), dW^[1] = X^T delta^[1].

πŸ’» 4. Code from Scratch (python)

mlp.py
import numpy as np

class MLPScratch:
    """2-Layer Multi-Layer Perceptron with Vectorized Backpropagation in NumPy."""
    def __init__(self, input_dim: int, hidden_dim: int, output_dim: int, lr: float = 0.01):
        self.lr = lr
        self.W1 = np.random.randn(input_dim, hidden_dim) * np.sqrt(2.0 / input_dim)
        self.b1 = np.zeros((1, hidden_dim))
        self.W2 = np.random.randn(hidden_dim, output_dim) * np.sqrt(2.0 / hidden_dim)
        self.b2 = np.zeros((1, output_dim))

    def _relu(self, z): return np.maximum(0, z)
    def _relu_grad(self, z): return (z > 0).astype(float)
    def _sigmoid(self, z): return 1.0 / (1.0 + np.exp(-np.clip(z, -20, 20)))

    def forward(self, X: np.ndarray) -> np.ndarray:
        self.X = X
        self.z1 = np.dot(X, self.W1) + self.b1
        self.a1 = self._relu(self.z1)
        self.z2 = np.dot(self.a1, self.W2) + self.b2
        self.a2 = self._sigmoid(self.z2)
        return self.a2

    def backward(self, y: np.ndarray):
        N = self.X.shape[0]
        delta2 = self.a2 - y
        dW2 = (1.0 / N) * np.dot(self.a1.T, delta2)
        db2 = (1.0 / N) * np.sum(delta2, axis=0, keepdims=True)

        delta1 = np.dot(delta2, self.W2.T) * self._relu_grad(self.z1)
        dW1 = (1.0 / N) * np.dot(self.X.T, delta1)
        db1 = (1.0 / N) * np.sum(delta1, axis=0, keepdims=True)

        self.W1 -= self.lr * dW1; self.b1 -= self.lr * db1
        self.W2 -= self.lr * dW2; self.b2 -= self.lr * db2

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter Β· 0 / 1 done
Q1/1 Why is an activation function like ReLU or GELU mandatory between linear layers in an MLP?

Finished this chapter?