Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
📐 The Math Behind MLLEVEL 1 · BEGINNERChapter 2

Probability Distributions, Maximum Likelihood & MAP

Random variables, Bayes' Theorem, and deriving loss functions from probabilistic priors.

⏱ 15 min read🎯 Prerequisites: Basic Calculus, Integrals & Exponentials
🔍 Inspect Architecture: Linear regression

💡 1. Core Intuition & Concepts

Machine learning loss functions are not arbitrary heuristics — they are negative log-likelihoods derived from fundamental probability distributions. Maximum Likelihood Estimation (MLE) finds parameters maximizing data likelihood, while Maximum A Posteriori (MAP) incorporates Bayesian parameter priors that directly yield L1 (Lasso) and L2 (Ridge) regularization.

📐 2. Mathematical Formulations & Derivations

Multivariate Gaussian Density
p(x;μ,Σ)=1(2π)D/2∣Σ∣1/2exp⁡(−12(x−μ)TΣ−1(x−μ))p(x; \mu, \Sigma) = \frac{1}{(2\pi)^{D/2} |\Sigma|^{1/2}} \exp\left( -\frac{1}{2}(x - \mu)^T \Sigma^{-1} (x - \mu) \right)
μ: mean vector, Σ: symmetric positive semi-definite covariance matrix.
Bayes' Theorem & Posterior Formulation
P(θ∣D)=P(D∣θ)P(θ)P(D)∝L(θ;D)⋅P(θ)P(\theta | \mathcal{D}) = \frac{P(\mathcal{D} | \theta) P(\theta)}{P(\mathcal{D})} \propto \mathcal{L}(\theta; \mathcal{D}) \cdot P(\theta)
Posterior ∝ Likelihood × Prior. Taking -log yields Loss = NLL + Regularizer.
MLE Equivalence to Mean Squared Error (MSE)
arg⁡max⁡θ∑i=1Nlog⁡p(yi∣xi;θ,σ2)  ⟺  arg⁡min⁡θ12σ2∑i=1N(yi−fθ(xi))2\arg\max_\theta \sum_{i=1}^N \log p(y_i | x_i; \theta, \sigma^2) \iff \arg\min_\theta \frac{1}{2\sigma^2} \sum_{i=1}^N (y_i - f_\theta(x_i))^2
Gaussian observational noise y = f(x) + ε with ε ~ N(0, σ²) yields MSE loss.
MAP Priors & Regularization Equivalence
θ∼N(0,τ2I)  ⟹  Ridge (L2):λ∥w∥22θ∼Laplace(0,b)  ⟹  Lasso (L1):λ∥w∥1\theta \sim \mathcal{N}(0, \tau^2 I) \implies \text{Ridge } (L_2): \lambda \|w\|_2^2 \qquad \theta \sim \text{Laplace}(0, b) \implies \text{Lasso } (L_1): \lambda \|w\|_1
Gaussian priors penalize large weights (L2); Laplace priors enforce exact sparsity (L1).

• Assume observational data y_i = f_theta(x_i) + eps_i, where noise eps_i ~ N(0, sigma^2) is i.i.d.

• The conditional likelihood for sample i is p(y_i | x_i, theta) = (1 / sqrt(2 pi sigma^2)) * exp(-(y_i - f_theta(x_i))^2 / (2 sigma^2)).

• The joint log-likelihood over the dataset D is log L(theta) = sum_{i=1}^N [-0.5 * log(2 pi sigma^2) - (y_i - f_theta(x_i))^2 / (2 sigma^2)].

• Discard constants independent of theta and negate to minimize: argmin_theta (1 / (2 sigma^2)) sum_{i=1}^N (y_i - f_theta(x_i))^2 = argmin_theta MSE(theta). Q.E.D.

⚙️ 3. Step-by-Step Computational Mechanism

1
Likelihood Function
Treat observed data as fixed and compute the joint probability as a function of model parameters L(theta) = prod_{i=1}^N p(x_i | theta).
2
Log-Likelihood Transformation
Taking the natural logarithm transforms products of probabilities into stable sums and prevents numerical underflow.
3
Prior Distribution (MAP)
Multiplying by prior P(theta) introduces inductive bias, constraining parameter space and preventing overfitting.

💻 4. Code from Scratch (python)

math_prob_mle.py
import numpy as np

# Comparing MLE (Linear Regression), MAP with Gaussian Prior (Ridge), and MAP with Laplace Prior (Lasso)
def demonstrate_mle_vs_map():
    np.random.seed(42)
    N, D = 20, 8  # Small sample size where regularization matters
    X = np.random.randn(N, D)
    true_w = np.array([3.0, -2.0, 0.0, 0.0, 1.5, 0.0, 0.0, -1.0])
    y = np.dot(X, true_w) + np.random.randn(N) * 0.5

    # 1. MLE: (X^T X)^(-1) X^T y
    w_mle = np.linalg.solve(np.dot(X.T, X) + 1e-6 * np.eye(D), np.dot(X.T, y))

    # 2. MAP with Gaussian Prior (Ridge / L2): (X^T X + lambda * I)^(-1) X^T y
    lambda_ridge = 2.0
    w_ridge = np.linalg.solve(np.dot(X.T, X) + lambda_ridge * np.eye(D), np.dot(X.T, y))

    print("True Weights:   ", np.round(true_w, 2))
    print("MLE (Overfit):  ", np.round(w_mle, 2))
    print("MAP Ridge (L2): ", np.round(w_ridge, 2))

if __name__ == "__main__":
    demonstrate_mle_vs_map()

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 Why does a Laplace prior on weights induce exact parameter sparsity (zeros) while a Gaussian prior only shrinks weights smoothly?

In the catalog

Finished this chapter?