Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

πŸŽ“ Track Progress: 0 Completed
0%
πŸ“ The Math Behind MLLEVEL 4 Β· FRONTIERChapter 8

Continuous Flow Matching ODEs & Post-Training RL Mathematics

Rectified flow velocity fields, Bellman optimality, DPO closed-form substitution, and DeepSeek-R1 GRPO advantage.

⏱ 25 min read🎯 Prerequisites: Differential Equations & Probability
πŸ” Inspect Architecture: DeepSeek-V3 and R1 reasoning MoE

πŸ’‘ 1. Core Intuition & Concepts

Modern frontier AI models synthesize media via straight-line Ordinary Differential Equations (Rectified Flow Matching) and learn complex reasoning behaviors through group-normalized relative advantage reinforcement learning (DeepSeek-R1 GRPO). These mathematical breakthroughs replace complex diffusion SDEs with deterministic velocity vectors and eliminate the need for separate Value critic networks.

πŸ“ 2. Mathematical Formulations & Derivations

Rectified Flow Matching Straight-Line Trajectory
xt=(1βˆ’t)x0+tx1β€…β€ŠβŸΉβ€…β€ŠTargetΒ VelocityΒ vt(xt)=dxtdt=x1βˆ’x0x_t = (1-t) x_0 + t x_1 \implies \text{Target Velocity } v_t(x_t) = \frac{d x_t}{dt} = x_1 - x_0
Straight probability path connecting pure Gaussian noise x_0 to clean data point x_1.
Flow Matching Least-Squares Objective
LRFM(ΞΈ)=Et∼U(0,1),x0∼p0,x1∼p1[βˆ₯vΞΈ(xt,t)βˆ’(x1βˆ’x0)βˆ₯22]\mathcal{L}_{\text{RFM}}(\theta) = \mathbb{E}_{t \sim \mathcal{U}(0,1), x_0 \sim p_0, x_1 \sim p_1} \left[ \|v_\theta(x_t, t) - (x_1 - x_0)\|_2^2 \right]
Neural network v_ΞΈ directly learns the constant velocity vector field along the ODE trajectory.
Direct Preference Optimization (DPO) Closed-Form Objective
LDPO(ΞΈ)=βˆ’E(x,yw,yl)[log⁑σ(Ξ²log⁑πθ(yw∣x)Ο€ref(yw∣x)βˆ’Ξ²log⁑πθ(yl∣x)Ο€ref(yl∣x))]\mathcal{L}_{\text{DPO}}(\theta) = -\mathbb{E}_{(x, y_w, y_l)} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right) \right]
Directly solves the constrained RL problem without training a separate reward model or critic.
DeepSeek-R1 Group Relative Policy Optimization (GRPO) Advantage
Ai=riβˆ’mean({r1,…,rG})std({r1,…,rG})+Ο΅JGRPO(ΞΈ)=1Gβˆ‘i=1Gmin⁑(πθ(oi∣q)Ο€old(oi∣q)Ai,clip(… )Ai)A_i = \frac{r_i - \text{mean}(\{r_1, \dots, r_G\})}{\text{std}(\{r_1, \dots, r_G\}) + \epsilon} \qquad \mathcal{J}_{\text{GRPO}}(\theta) = \frac{1}{G}\sum_{i=1}^G \min\left( \frac{\pi_\theta(o_i|q)}{\pi_{\text{old}}(o_i|q)} A_i, \text{clip}(\dots) A_i \right)
Computes relative baseline advantages across group of G sampled rollouts, completely eliminating the Critic model.

β€’ In RLHF, the objective with KL penalty is max_pi E_{x, y ~ pi}[r(x, y)] - beta * D_KL(pi(y|x) || pi_ref(y|x)).

β€’ The closed-form analytical solution to this constrained optimization is pi*(y|x) = (1 / Z(x)) * pi_ref(y|x) * exp((1 / beta) * r(x, y)).

β€’ Rearranging yields the exact implicit reward: r(x, y) = beta * log(pi*(y|x) / pi_ref(y|x)) + beta * log Z(x).

β€’ Substitute this implicit reward into the Bradley-Terry preference probability P(y_w > y_l | x) = sigma(r(x, y_w) - r(x, y_l)).

β€’ The partition function terms beta * log Z(x) cancel out exactly, yielding the DPO loss without any reward model. Q.E.D.

βš™οΈ 3. Step-by-Step Computational Mechanism

1
Deterministic Flow Generation
Euler ODE integration dx/dt = v_theta(x, t) generates high-resolution samples in as few as 4-8 steps.
2
Group-Normalized Baseline
GRPO samples G responses for question q and normalizes rewards to create zero-mean advantages.
3
Emergent Long Chain-of-Thought
Pure rule-based correctness rewards (math/code tests) combined with GRPO drive models to self-reflect and verify.

πŸ’» 4. Code from Scratch (python)

math_flow_rl.py
import numpy as np

# GRPO Group-Normalized Advantage Calculation from Scratch (DeepSeek-R1)
def compute_grpo_advantages(rewards: np.ndarray, eps: float = 1e-8):
    mean_r = np.mean(rewards, axis=-1, keepdims=True)
    std_r = np.std(rewards, axis=-1, keepdims=True)
    advantages = (rewards - mean_r) / (std_r + eps)
    return advantages

if __name__ == "__main__":
    sample_rewards = np.array([
        [1.0, 0.0, 1.0, 0.0],  # Prompt 1: 2 correct, 2 incorrect
        [0.8, 0.9, 0.2, 0.8]   # Prompt 2: partial credits
    ])
    adv = compute_grpo_advantages(sample_rewards)
    print("Computed GRPO Relative Advantages:")
    print(np.round(adv, 3))
    assert np.allclose(np.mean(adv, axis=-1), 0.0), "Group mean must be zero!"
    print("βœ“ Zero-mean baseline normalization verified.")

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter Β· 0 / 1 done
Q1/1 How does DeepSeek-R1's GRPO eliminate the need for a separate Value (Critic) network during RL fine-tuning?

Finished this chapter?