π‘ 1. Core Intuition & ConceptsModern frontier AI models synthesize media via straight-line Ordinary Differential Equations (Rectified Flow Matching) and learn complex reasoning behaviors through group-normalized relative advantage reinforcement learning (DeepSeek-R1 GRPO). These mathematical breakthroughs replace complex diffusion SDEs with deterministic velocity vectors and eliminate the need for separate Value critic networks.
π 2. Mathematical Formulations & DerivationsRectified Flow Matching Straight-Line Trajectory
x t = ( 1 β t ) x 0 + t x 1 β
β βΉ β
β TargetΒ VelocityΒ v t ( x t ) = d x t d t = x 1 β x 0 x_t = (1-t) x_0 + t x_1 \implies \text{Target Velocity } v_t(x_t) = \frac{d x_t}{dt} = x_1 - x_0 x t β = ( 1 β t ) x 0 β + t x 1 β βΉ TargetΒ VelocityΒ v t β ( x t β ) = d t d x t β β = x 1 β β x 0 β Straight probability path connecting pure Gaussian noise x_0 to clean data point x_1.
Flow Matching Least-Squares Objective
L RFM ( ΞΈ ) = E t βΌ U ( 0 , 1 ) , x 0 βΌ p 0 , x 1 βΌ p 1 [ β₯ v ΞΈ ( x t , t ) β ( x 1 β x 0 ) β₯ 2 2 ] \mathcal{L}_{\text{RFM}}(\theta) = \mathbb{E}_{t \sim \mathcal{U}(0,1), x_0 \sim p_0, x_1 \sim p_1} \left[ \|v_\theta(x_t, t) - (x_1 - x_0)\|_2^2 \right] L RFM β ( ΞΈ ) = E t βΌ U ( 0 , 1 ) , x 0 β βΌ p 0 β , x 1 β βΌ p 1 β β [ β₯ v ΞΈ β ( x t β , t ) β ( x 1 β β x 0 β ) β₯ 2 2 β ] Neural network v_ΞΈ directly learns the constant velocity vector field along the ODE trajectory.
Direct Preference Optimization (DPO) Closed-Form Objective
L DPO ( ΞΈ ) = β E ( x , y w , y l ) [ log β‘ Ο ( Ξ² log β‘ Ο ΞΈ ( y w β£ x ) Ο ref ( y w β£ x ) β Ξ² log β‘ Ο ΞΈ ( y l β£ x ) Ο ref ( y l β£ x ) ) ] \mathcal{L}_{\text{DPO}}(\theta) = -\mathbb{E}_{(x, y_w, y_l)} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right) \right] L DPO β ( ΞΈ ) = β E ( x , y w β , y l β ) β [ log Ο ( Ξ² log Ο ref β ( y w β β£ x ) Ο ΞΈ β ( y w β β£ x ) β β Ξ² log Ο ref β ( y l β β£ x ) Ο ΞΈ β ( y l β β£ x ) β ) ] Directly solves the constrained RL problem without training a separate reward model or critic.
DeepSeek-R1 Group Relative Policy Optimization (GRPO) Advantage
A i = r i β mean ( { r 1 , β¦ , r G } ) std ( { r 1 , β¦ , r G } ) + Ο΅ J GRPO ( ΞΈ ) = 1 G β i = 1 G min β‘ ( Ο ΞΈ ( o i β£ q ) Ο old ( o i β£ q ) A i , clip ( β¦ β ) A i ) A_i = \frac{r_i - \text{mean}(\{r_1, \dots, r_G\})}{\text{std}(\{r_1, \dots, r_G\}) + \epsilon} \qquad \mathcal{J}_{\text{GRPO}}(\theta) = \frac{1}{G}\sum_{i=1}^G \min\left( \frac{\pi_\theta(o_i|q)}{\pi_{\text{old}}(o_i|q)} A_i, \text{clip}(\dots) A_i \right) A i β = std ({ r 1 β , β¦ , r G β }) + Ο΅ r i β β mean ({ r 1 β , β¦ , r G β }) β J GRPO β ( ΞΈ ) = G 1 β i = 1 β G β min ( Ο old β ( o i β β£ q ) Ο ΞΈ β ( o i β β£ q ) β A i β , clip ( β¦ ) A i β ) Computes relative baseline advantages across group of G sampled rollouts, completely eliminating the Critic model.
βΆ Proof Breakdown & Step-by-Step Derivation: Derivation: DPO Closed-Form Policy Representation [Toggle] β’ In RLHF, the objective with KL penalty is max_pi E_{x, y ~ pi}[r(x, y)] - beta * D_KL(pi(y|x) || pi_ref(y|x)).
β’ The closed-form analytical solution to this constrained optimization is pi*(y|x) = (1 / Z(x)) * pi_ref(y|x) * exp((1 / beta) * r(x, y)).
β’ Rearranging yields the exact implicit reward: r(x, y) = beta * log(pi*(y|x) / pi_ref(y|x)) + beta * log Z(x).
β’ Substitute this implicit reward into the Bradley-Terry preference probability P(y_w > y_l | x) = sigma(r(x, y_w) - r(x, y_l)).
β’ The partition function terms beta * log Z(x) cancel out exactly, yielding the DPO loss without any reward model. Q.E.D.
βοΈ 3. Step-by-Step Computational Mechanism1
Deterministic Flow Generation
Euler ODE integration dx/dt = v_theta(x, t) generates high-resolution samples in as few as 4-8 steps.
2
Group-Normalized Baseline
GRPO samples G responses for question q and normalizes rewards to create zero-mean advantages.
3
Emergent Long Chain-of-Thought
Pure rule-based correctness rewards (math/code tests) combined with GRPO drive models to self-reflect and verify.
π» 4. Code from Scratch (python)import numpy as np
# GRPO Group-Normalized Advantage Calculation from Scratch (DeepSeek-R1)
def compute_grpo_advantages(rewards: np.ndarray, eps: float = 1e-8):
mean_r = np.mean(rewards, axis=-1, keepdims=True)
std_r = np.std(rewards, axis=-1, keepdims=True)
advantages = (rewards - mean_r) / (std_r + eps)
return advantages
if __name__ == "__main__":
sample_rewards = np.array([
[1.0, 0.0, 1.0, 0.0], # Prompt 1: 2 correct, 2 incorrect
[0.8, 0.9, 0.2, 0.8] # Prompt 2: partial credits
])
adv = compute_grpo_advantages(sample_rewards)
print("Computed GRPO Relative Advantages:")
print(np.round(adv, 3))
assert np.allclose(np.mean(adv, axis=-1), 0.0), "Group mean must be zero!"
print("β Zero-mean baseline normalization verified.")π§ 5. Comprehension CheckpointAnswer all 1 questions correctly to complete the chapter Β· 0 / 1 done
Q1/1 How does DeepSeek-R1's GRPO eliminate the need for a separate Value (Critic) network during RL fine-tuning?
A. It samples a group of G rollouts per prompt and uses their empirical group mean and standard deviation as the baseline to compute relative advantage. B. It replaces the policy gradient with a supervised cross-entropy loss. C. It assumes all actions have a constant reward of 1.0. D. It runs Monte Carlo Tree Search at every token generation step.
Finished this chapter?
Mark as Complete