π» Modern ArchitecturesLEVEL 2 Β· INTERMEDIATEChapter 4
Multi-Layer Perceptrons & Analytical Backpropagation
Hidden representations, non-linear activations (ReLU/GELU), and matrix chain rule.
π Inspect Architecture: Multilayer perceptron (MLP)π‘ 1. Core Intuition & Concepts
Multi-Layer Perceptrons (MLPs) overcome the linear separability limitations of single perceptrons by stacking linear transformations with non-linear activation functions. Analytical backpropagation uses the multivariate calculus chain rule to propagate error gradients backwards from the scalar loss through intermediate hidden layers.
π 2. Mathematical Formulations & Derivations
Forward Pass Equation
l: layer index, W^[l]: weight matrix (D_{in}ΓD_{out}), g: activation function.
Output Error Delta
Hadamard element-wise product between loss derivative and activation slope.
Hidden Layer Backprop Chain Rule
Transposed weight projection routes downstream errors back through the network.
βοΈ 3. Step-by-Step Computational Mechanism
1
Forward Activation
Compute intermediate layer activations z^[1] = X W^[1] + b^[1], a^[1] = ReLU(z^[1]).
2
Loss Evaluation
Compute scalar objective L = BCE(y_hat, y) or MSE(y_hat, y).
3
Backward Gradient Delta
Compute output error delta^[2] = y_hat - y.
4
Hidden Gradients
Calculate dW^[2] = (a^[1])^T delta^[2], delta^[1] = (delta^[2] (W^[2])^T) * ReLU'(z^[1]), dW^[1] = X^T delta^[1].
π» 4. Code from Scratch (python)
mlp.py
import numpy as np
class MLPScratch:
"""2-Layer Multi-Layer Perceptron with Vectorized Backpropagation in NumPy."""
def __init__(self, input_dim: int, hidden_dim: int, output_dim: int, lr: float = 0.01):
self.lr = lr
self.W1 = np.random.randn(input_dim, hidden_dim) * np.sqrt(2.0 / input_dim)
self.b1 = np.zeros((1, hidden_dim))
self.W2 = np.random.randn(hidden_dim, output_dim) * np.sqrt(2.0 / hidden_dim)
self.b2 = np.zeros((1, output_dim))
def _relu(self, z): return np.maximum(0, z)
def _relu_grad(self, z): return (z > 0).astype(float)
def _sigmoid(self, z): return 1.0 / (1.0 + np.exp(-np.clip(z, -20, 20)))
def forward(self, X: np.ndarray) -> np.ndarray:
self.X = X
self.z1 = np.dot(X, self.W1) + self.b1
self.a1 = self._relu(self.z1)
self.z2 = np.dot(self.a1, self.W2) + self.b2
self.a2 = self._sigmoid(self.z2)
return self.a2
def backward(self, y: np.ndarray):
N = self.X.shape[0]
delta2 = self.a2 - y
dW2 = (1.0 / N) * np.dot(self.a1.T, delta2)
db2 = (1.0 / N) * np.sum(delta2, axis=0, keepdims=True)
delta1 = np.dot(delta2, self.W2.T) * self._relu_grad(self.z1)
dW1 = (1.0 / N) * np.dot(self.X.T, delta1)
db1 = (1.0 / N) * np.sum(delta1, axis=0, keepdims=True)
self.W1 -= self.lr * dW1; self.b1 -= self.lr * db1
self.W2 -= self.lr * dW2; self.b2 -= self.lr * db2π§ 5. Comprehension Checkpoint
Answer all 1 questions correctly to complete the chapter Β· 0 / 1 done
Q1/1 Why is an activation function like ReLU or GELU mandatory between linear layers in an MLP?
In the catalog
Finished this chapter?