Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
📐 The Math Behind MLLEVEL 1 · BEGINNERChapter 1

Linear Algebra, Vector Spaces & Matrix Calculus

Vector norms, geometric projections, trace derivative identities, and multivariate matrix calculus.

⏱ 15 min read🎯 Prerequisites: High School Algebra & Functions
🔍 Inspect Architecture: Linear regression

💡 1. Core Intuition & Concepts

Every machine learning dataset lives inside a vector space R^D, and every neural network layer applies a parameterized affine transformation W x + b followed by coordinate-wise non-linear curvature. Matrix calculus provides the exact formal mechanics for computing how scalar loss functions change with respect to multidimensional weight matrices.

📐 2. Mathematical Formulations & Derivations

Vector & Matrix Norms
∥x∥1=∑i=1D∣xi∣∥x∥2=xTx∥A∥F=Tr(ATA)=∑i,jAij2\|x\|_1 = \sum_{i=1}^D |x_i| \qquad \|x\|_2 = \sqrt{x^T x} \qquad \|A\|_F = \sqrt{\text{Tr}(A^T A)} = \sqrt{\sum_{i,j} A_{ij}^2}
L1 norm (Manhattan/sparsity), L2 norm (Euclidean distance), and Frobenius norm (matrix magnitude).
Inner Product, Cosine Similarity & Orthogonal Projection
⟨u,v⟩=uTv=∥u∥2∥v∥2cos⁡θproju(v)=uTv∥u∥22u\langle u, v \rangle = u^T v = \|u\|_2 \|v\|_2 \cos \theta \qquad \text{proj}_u(v) = \frac{u^T v}{\|u\|_2^2} u
Measures alignment angle θ and projects vector v onto the 1D subspace spanned by u.
Essential Matrix Calculus Identities
∇x(aTx)=a∇x(xTAx)=(A+AT)x∇WTr(AWB)=ATBT\nabla_x (a^T x) = a \qquad \nabla_x (x^T A x) = (A + A^T)x \qquad \nabla_W \text{Tr}(A W B) = A^T B^T
Analytical vector and matrix derivatives commonly used in loss minimization.
Matrix Form Least Squares Gradient
L(W)=1N∥XW−Y∥F2  ⟹  ∇WL=2NXT(XW−Y)\mathcal{L}(W) = \frac{1}{N} \|X W - Y\|_F^2 \implies \nabla_W \mathcal{L} = \frac{2}{N} X^T (X W - Y)
Vectorized gradient of matrix linear regression with N samples and multi-output dimensions.

• Let the residual error matrix be E = XW - Y in R^(N x K).

• Express the Frobenius norm using the trace operator: ||E||_F^2 = Tr(E^T E) = Tr((XW - Y)^T (XW - Y)).

• Expand the terms inside the trace: Tr(W^T X^T X W - W^T X^T Y - Y^T X W + Y^T Y).

• Apply the trace derivative identity ∇_W Tr(W^T A W) = (A + A^T) W = 2 A W (since A = X^T X is symmetric), and ∇_W Tr(A W) = A^T:

• ∇_W ||XW - Y||_F^2 = 2 X^T X W - 2 X^T Y = 2 X^T (X W - Y). Q.E.D.

⚙️ 3. Step-by-Step Computational Mechanism

1
Coordinate Transformations
Matrix-vector multiplication y = Ax rotates, shears, and scales the coordinate basis vectors e_i.
2
Jacobians & Multi-Output Sensitivity
For vector function f: R^n -> R^m, the Jacobian matrix J_ij = ∂f_i / ∂x_j maps input velocity to output perturbation.
3
Hessians & Local Curvature
The symmetric Hessian matrix H_ij = ∂²f / (∂x_i ∂x_j) defines the second-order quadratic landscape around critical points.

💻 4. Code from Scratch (python)

math_linalg.py
import numpy as np

# Numerical vs Analytical Matrix Gradient Verification
def check_matrix_gradient(N=50, D=5, K=2, eps=1e-5):
    np.random.seed(42)
    X = np.random.randn(N, D)
    W = np.random.randn(D, K)
    Y = np.random.randn(N, K)

    # Analytical gradient: (2/N) * X^T (X W - Y)
    grad_analytical = (2.0 / N) * np.dot(X.T, (np.dot(X, W) - Y))

    # Numerical gradient via central finite difference
    grad_numerical = np.zeros_like(W)
    for i in range(D):
        for j in range(K):
            W_pos = W.copy(); W_pos[i, j] += eps
            W_neg = W.copy(); W_neg[i, j] -= eps
            loss_pos = np.mean((np.dot(X, W_pos) - Y)**2)
            loss_neg = np.mean((np.dot(X, W_neg) - Y)**2)
            grad_numerical[i, j] = (loss_pos - loss_neg) / (2.0 * eps)

    max_diff = np.max(np.abs(grad_analytical - grad_numerical))
    print(f"Analytical vs Numerical Gradient Max Difference: {max_diff:.2e}")
    assert max_diff < 1e-6, "Gradient check failed!"
    print("✓ Matrix Calculus Gradient Identity Verified Successfully.")

if __name__ == "__main__":
    check_matrix_gradient()

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 Why is the gradient of a general quadratic form f(x) = x^T A x equal to (A + A^T)x instead of 2Ax?

In the catalog

Finished this chapter?