Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

πŸŽ“ Track Progress: 0 Completed
0%
πŸ“ The Math Behind MLLEVEL 2 Β· INTERMEDIATEChapter 3

Information Theory, Entropy & Divergence Measures

Shannon entropy, cross-entropy, KL-divergence, Gibbs' inequality, and mutual information.

⏱ 18 min read🎯 Prerequisites: Probability & Logarithms

πŸ’‘ 1. Core Intuition & Concepts

Information theory formalizes uncertainty, surprise, and distance between probability distributions. It provides the foundation for decision tree splitting (Information Gain), categorical cross-entropy loss in neural networks, and modern alignment objectives like DPO and RLHF.

πŸ“ 2. Mathematical Formulations & Derivations

Shannon Entropy (Information Content)
H(P)=βˆ’βˆ‘x∈XP(x)log⁑2P(x)=Ex∼P[βˆ’log⁑2P(x)]H(P) = -\sum_{x \in \mathcal{X}} P(x) \log_2 P(x) = \mathbb{E}_{x \sim P}[-\log_2 P(x)]
Average number of bits required to encode an event drawn from distribution P.
Cross-Entropy Formula
H(P,Q)=βˆ’βˆ‘x∈XP(x)log⁑Q(x)=H(P)+DKL(Pβˆ₯Q)H(P, Q) = -\sum_{x \in \mathcal{X}} P(x) \log Q(x) = H(P) + D_{KL}(P \parallel Q)
Coding cost using model distribution Q when true distribution is P.
Kullback-Leibler (KL) Divergence
DKL(Pβˆ₯Q)=βˆ‘x∈XP(x)log⁑P(x)Q(x)β‰₯0(Gibbs’ Inequality)D_{KL}(P \parallel Q) = \sum_{x \in \mathcal{X}} P(x) \log \frac{P(x)}{Q(x)} \ge 0 \qquad \text{(Gibbs' Inequality)}
Relative entropy quantifying information lost when approximating P with Q.
Mutual Information & Information Bottleneck
I(X;Y)=H(X)βˆ’H(X∣Y)=DKL(P(X,Y)βˆ₯P(X)P(Y))I(X; Y) = H(X) - H(X|Y) = D_{KL}\big(P(X,Y) \parallel P(X)P(Y)\big)
Information shared between random variables X and Y; measures statistical dependence.

β€’ Recall Jensen's inequality: For a strictly concave function phi(t), E[phi(t)] <= phi(E[t]). The logarithm ln(t) is strictly concave.

β€’ Express -D_KL(P || Q) as an expectation under P: -D_KL(P || Q) = sum_x P(x) ln(Q(x)/P(x)) = E_{x ~ P}[ln(Q(x)/P(x))].

β€’ Apply Jensen's inequality: E_{x ~ P}[ln(Q(x)/P(x))] <= ln(E_{x ~ P}[Q(x)/P(x)]) = ln(sum_x P(x) (Q(x)/P(x))).

β€’ Simplify the summation: sum_x Q(x) = 1. Thus ln(1) = 0.

β€’ Therefore, -D_KL(P || Q) <= 0 => D_KL(P || Q) >= 0, with equality if and only if P(x) = Q(x) for all x. Q.E.D.

βš™οΈ 3. Step-by-Step Computational Mechanism

1
Self-Information
Surprise I(x) = -log P(x) is inversely proportional to probability: rare events carry high information.
2
Forward vs Reverse KL
Forward D_KL(P_true || Q) is mean-seeking (zero-avoiding); Reverse D_KL(Q || P_true) is mode-seeking (zero-forcing, used in RL/DPO).
3
Jensen-Shannon Symmetrization
D_JS(P || Q) = 0.5 * D_KL(P || M) + 0.5 * D_KL(Q || M) where M = 0.5*(P+Q) bounds distance in [0, 1].

πŸ’» 4. Code from Scratch (python)

math_info_theory.py
import numpy as np

# Numerical computation of Entropy, Cross-Entropy, and KL-Divergence
def compute_divergences(p, q, eps=1e-12):
    p = np.clip(p, eps, 1.0); p /= np.sum(p)
    q = np.clip(q, eps, 1.0); q /= np.sum(q)

    # Shannon Entropy H(P)
    H_p = -np.sum(p * np.log2(p))
    # Cross-Entropy H(P, Q)
    H_pq = -np.sum(p * np.log2(q))
    # KL Divergence D_KL(P || Q)
    D_kl = np.sum(p * np.log2(p / q))
    # Jensen-Shannon Divergence
    m = 0.5 * (p + q)
    D_js = 0.5 * np.sum(p * np.log2(p / m)) + 0.5 * np.sum(q * np.log2(q / m))

    return {"H(P)": H_p, "H(P,Q)": H_pq, "D_KL(P||Q)": D_kl, "D_JS": D_js}

if __name__ == "__main__":
    P = np.array([0.7, 0.2, 0.1])  # True distribution
    Q = np.array([0.5, 0.3, 0.2])  # Model prediction
    res = compute_divergences(P, Q)
    for k, v in res.items():
        print(f"{k:12s}: {v:.4f} bits")
    assert abs(res["H(P,Q)"] - (res["H(P)"] + res["D_KL(P||Q)"])) < 1e-6
    print("βœ“ Cross-Entropy Identity H(P, Q) = H(P) + D_KL(P || Q) Verified.")

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter Β· 0 / 1 done
Q1/1 Why is minimizing Cross-Entropy H(P_true, Q_model) mathematically equivalent to minimizing KL-Divergence D_KL(P_true || Q_model)?

Finished this chapter?