💡 1. Core Intuition & ConceptsMachine learning loss functions are not arbitrary heuristics — they are negative log-likelihoods derived from fundamental probability distributions. Maximum Likelihood Estimation (MLE) finds parameters maximizing data likelihood, while Maximum A Posteriori (MAP) incorporates Bayesian parameter priors that directly yield L1 (Lasso) and L2 (Ridge) regularization.
📐 2. Mathematical Formulations & DerivationsMultivariate Gaussian Density
p ( x ; μ , Σ ) = 1 ( 2 π ) D / 2 ∣ Σ ∣ 1 / 2 exp ( − 1 2 ( x − μ ) T Σ − 1 ( x − μ ) ) p(x; \mu, \Sigma) = \frac{1}{(2\pi)^{D/2} |\Sigma|^{1/2}} \exp\left( -\frac{1}{2}(x - \mu)^T \Sigma^{-1} (x - \mu) \right) p ( x ; μ , Σ ) = ( 2 π ) D /2 ∣Σ ∣ 1/2 1 exp ( − 2 1 ( x − μ ) T Σ − 1 ( x − μ ) ) μ: mean vector, Σ: symmetric positive semi-definite covariance matrix.
Bayes' Theorem & Posterior Formulation
P ( θ ∣ D ) = P ( D ∣ θ ) P ( θ ) P ( D ) ∝ L ( θ ; D ) ⋅ P ( θ ) P(\theta | \mathcal{D}) = \frac{P(\mathcal{D} | \theta) P(\theta)}{P(\mathcal{D})} \propto \mathcal{L}(\theta; \mathcal{D}) \cdot P(\theta) P ( θ ∣ D ) = P ( D ) P ( D ∣ θ ) P ( θ ) ∝ L ( θ ; D ) ⋅ P ( θ ) Posterior ∝ Likelihood × Prior. Taking -log yields Loss = NLL + Regularizer.
MLE Equivalence to Mean Squared Error (MSE)
arg max θ ∑ i = 1 N log p ( y i ∣ x i ; θ , σ 2 ) ⟺ arg min θ 1 2 σ 2 ∑ i = 1 N ( y i − f θ ( x i ) ) 2 \arg\max_\theta \sum_{i=1}^N \log p(y_i | x_i; \theta, \sigma^2) \iff \arg\min_\theta \frac{1}{2\sigma^2} \sum_{i=1}^N (y_i - f_\theta(x_i))^2 arg θ max i = 1 ∑ N log p ( y i ∣ x i ; θ , σ 2 ) ⟺ arg θ min 2 σ 2 1 i = 1 ∑ N ( y i − f θ ( x i ) ) 2 Gaussian observational noise y = f(x) + ε with ε ~ N(0, σ²) yields MSE loss.
MAP Priors & Regularization Equivalence
θ ∼ N ( 0 , τ 2 I ) ⟹ Ridge ( L 2 ) : λ ∥ w ∥ 2 2 θ ∼ Laplace ( 0 , b ) ⟹ Lasso ( L 1 ) : λ ∥ w ∥ 1 \theta \sim \mathcal{N}(0, \tau^2 I) \implies \text{Ridge } (L_2): \lambda \|w\|_2^2 \qquad \theta \sim \text{Laplace}(0, b) \implies \text{Lasso } (L_1): \lambda \|w\|_1 θ ∼ N ( 0 , τ 2 I ) ⟹ Ridge ( L 2 ) : λ ∥ w ∥ 2 2 θ ∼ Laplace ( 0 , b ) ⟹ Lasso ( L 1 ) : λ ∥ w ∥ 1 Gaussian priors penalize large weights (L2); Laplace priors enforce exact sparsity (L1).
▶ Proof Breakdown & Step-by-Step Derivation: Derivation: Gaussian Noise Assumption Yields MSE Loss [Toggle] • Assume observational data y_i = f_theta(x_i) + eps_i, where noise eps_i ~ N(0, sigma^2) is i.i.d.
• The conditional likelihood for sample i is p(y_i | x_i, theta) = (1 / sqrt(2 pi sigma^2)) * exp(-(y_i - f_theta(x_i))^2 / (2 sigma^2)).
• The joint log-likelihood over the dataset D is log L(theta) = sum_{i=1}^N [-0.5 * log(2 pi sigma^2) - (y_i - f_theta(x_i))^2 / (2 sigma^2)].
• Discard constants independent of theta and negate to minimize: argmin_theta (1 / (2 sigma^2)) sum_{i=1}^N (y_i - f_theta(x_i))^2 = argmin_theta MSE(theta). Q.E.D.
⚙️ 3. Step-by-Step Computational Mechanism1
Likelihood Function
Treat observed data as fixed and compute the joint probability as a function of model parameters L(theta) = prod_{i=1}^N p(x_i | theta).
2
Log-Likelihood Transformation
Taking the natural logarithm transforms products of probabilities into stable sums and prevents numerical underflow.
3
Prior Distribution (MAP)
Multiplying by prior P(theta) introduces inductive bias, constraining parameter space and preventing overfitting.
💻 4. Code from Scratch (python)import numpy as np
# Comparing MLE (Linear Regression), MAP with Gaussian Prior (Ridge), and MAP with Laplace Prior (Lasso)
def demonstrate_mle_vs_map():
np.random.seed(42)
N, D = 20, 8 # Small sample size where regularization matters
X = np.random.randn(N, D)
true_w = np.array([3.0, -2.0, 0.0, 0.0, 1.5, 0.0, 0.0, -1.0])
y = np.dot(X, true_w) + np.random.randn(N) * 0.5
# 1. MLE: (X^T X)^(-1) X^T y
w_mle = np.linalg.solve(np.dot(X.T, X) + 1e-6 * np.eye(D), np.dot(X.T, y))
# 2. MAP with Gaussian Prior (Ridge / L2): (X^T X + lambda * I)^(-1) X^T y
lambda_ridge = 2.0
w_ridge = np.linalg.solve(np.dot(X.T, X) + lambda_ridge * np.eye(D), np.dot(X.T, y))
print("True Weights: ", np.round(true_w, 2))
print("MLE (Overfit): ", np.round(w_mle, 2))
print("MAP Ridge (L2): ", np.round(w_ridge, 2))
if __name__ == "__main__":
demonstrate_mle_vs_map()🧠 5. Comprehension CheckpointAnswer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 Why does a Laplace prior on weights induce exact parameter sparsity (zeros) while a Gaussian prior only shrinks weights smoothly?
A. The Laplace distribution has a sharp, non-differentiable peak at w=0 with constant derivative sgn(w), pulling small weights directly to zero. B. The Gaussian distribution has heavier tails than the Cauchy distribution. C. The Laplace distribution cannot be integrated in closed form. D. Because the Gaussian Hessian is singular at the origin.
Finished this chapter?
Mark as Complete