Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
💻 Modern ArchitecturesLEVEL 1 · BEGINNERChapter 1

Linear Regression & Gradient Descent

Deriving ordinary least squares, vectorization, and batch gradient descent in pure NumPy.

⏱ 10 min read🎯 Prerequisites: Basic Python, Vectors & Matrices
🔍 Inspect Architecture: Linear regression

💡 1. Core Intuition & Concepts

Linear regression models a scalar or vector target as a linear combination of explanatory input features with trainable weights and bias. Parameter optimization via gradient descent iteratively minimizes the Mean Squared Error (MSE) loss along the negative gradient direction.

📐 2. Mathematical Formulations & Derivations

Model & Mean Squared Error Loss
y^=Xw+bL(w,b)=1N∑i=1N(y^i−yi)2=1N∥Xw+b−y∥22\hat{y} = Xw + b \qquad \mathcal{L}(w,b) = \frac{1}{N}\sum_{i=1}^N (\hat{y}_i - y_i)^2 = \frac{1}{N} \|Xw + b - y\|^2_2
N: sample count, X: feature matrix (N×D), w: weight vector (D×1), b: scalar bias.
Analytical Matrix Gradients
∇wL=2NXT(Xw+b−y)∂L∂b=2N∑i=1N(y^i−yi)\nabla_w \mathcal{L} = \frac{2}{N} X^T (Xw + b - y) \qquad \frac{\partial \mathcal{L}}{\partial b} = \frac{2}{N} \sum_{i=1}^N (\hat{y}_i - y_i)
Vectorized partial derivatives with respect to weight vector w and scalar bias b.
Parameter Update Rule
w←w−α∇wLb←b−α∂L∂bw \leftarrow w - \alpha \nabla_w \mathcal{L} \qquad b \leftarrow b - \alpha \frac{\partial \mathcal{L}}{\partial b}
α: learning rate hyperparameter controlling the gradient step magnitude.

⚙️ 3. Step-by-Step Computational Mechanism

1
Initialize Parameters
Set weight vector w to zeros or small Gaussian noise N(0, 0.01) and bias b = 0.
2
Vectorized Forward Pass
Compute predictions across the full batch simultaneously using matrix multiplication y_hat = Xw + b.
3
Compute Error & Gradient
Calculate residual error e = y_hat - y and matrix-vector product ∇_w = (2/N) X^T e.
4
Step & Iterate
Subtract scaled gradients from parameters for E epochs until MSE loss stabilizes.

💻 4. Code from Scratch (python)

linreg.py
import numpy as np

class LinearRegressionScratch:
    """Linear Regression with Batch Gradient Descent implemented from scratch in NumPy."""
    def __init__(self, lr: float = 0.01, n_iters: int = 1000):
        self.lr = lr
        self.n_iters = n_iters
        self.weights = None
        self.bias = 0.0
        self.loss_history = []

    def fit(self, X: np.ndarray, y: np.ndarray):
        n_samples, n_features = X.shape
        self.weights = np.zeros(n_features)
        self.bias = 0.0

        for epoch in range(self.n_iters):
            # 1. Forward Pass: y_hat = X * w + b
            y_pred = np.dot(X, self.weights) + self.bias

            # 2. Compute Mean Squared Error Loss
            loss = np.mean((y_pred - y) ** 2)
            self.loss_history.append(loss)

            # 3. Vectorized Gradients
            error = y_pred - y
            dw = (2 / n_samples) * np.dot(X.T, error)
            db = (2 / n_samples) * np.sum(error)

            # 4. Update Parameters
            self.weights -= self.lr * dw
            self.bias -= self.lr * db

    def predict(self, X: np.ndarray) -> np.ndarray:
        return np.dot(X, self.weights) + self.bias

if __name__ == "__main__":
    np.random.seed(42)
    X = 2 * np.random.rand(100, 1)
    y = 4.0 + 3.0 * X.squeeze() + np.random.randn(100) * 0.05
    model = LinearRegressionScratch(lr=0.1, n_iters=300)
    model.fit(X, y)
    print(f"Fitted w: {model.weights[0]:.3f} (True: 3.000), b: {model.bias:.3f} (True: 4.000)")

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 What occurs if the learning rate α in gradient descent is set too large?

In the catalog

Finished this chapter?