Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
💻 Modern ArchitecturesLEVEL 4 · FRONTIERChapter 11

Post-Training RL for Reasoning LLMs (DPO & DeepSeek-R1 GRPO)

Direct Preference Optimization (DPO), Group Relative Policy Optimization (GRPO), and self-reflection.

⏱ 25 min read🎯 Prerequisites: Reinforcement Learning, Policy Gradients & Cross-Entropy

💡 1. Core Intuition & Concepts

Modern reasoning models like DeepSeek-R1, OpenAI o1, and Claude 3.7 Sonnet achieve emergent chain-of-thought verification using post-training reinforcement learning. DeepSeek-R1 introduced Group Relative Policy Optimization (GRPO), computing relative advantages across a group of sampled responses to eliminate the value critic network.

📐 2. Mathematical Formulations & Derivations

Direct Preference Optimization (DPO) Loss
LDPO(θ)=−E(x,yw,yl)[log⁡σ(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_{\text{DPO}}(\theta) = -\mathbb{E}_{(x, y_w, y_l)} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right) \right]
Optimizes preferences directly on policy logits without requiring a separate reward model.
DeepSeek-R1 GRPO Group Advantage
Ai=ri−mean({r1,…,rG})std({r1,…,rG})+ϵA_i = \frac{r_i - \text{mean}(\{r_1, \dots, r_G\})}{\text{std}(\{r_1, \dots, r_G\}) + \epsilon}
Normalizes rewards across a group of G sampled response rollouts for the same prompt.

⚙️ 3. Step-by-Step Computational Mechanism

1
Sample Candidate Group
For prompt q, sample a group of G response trajectories {o_1, o_2, ..., o_G} from current policy π_θ.
2
Evaluate Rule-Based Rewards
Score each response with binary mathematical correctness, format adherence, or compiler tests r_i.
3
Compute Group-Relative Advantage
Calculate standardized advantage A_i = (r_i - r_bar) / (sigma_r + eps).
4
PPO-Clipped Policy Update with KL Penalty
Update policy parameters θ while constraining drift from reference model π_ref.

💻 4. Code from Scratch (python)

rl_grpo.py
import torch
import torch.nn.functional as F

def compute_dpo_loss(
    pi_logps_win: torch.Tensor,
    pi_logps_lose: torch.Tensor,
    ref_logps_win: torch.Tensor,
    ref_logps_lose: torch.Tensor,
    beta: float = 0.1
) -> torch.Tensor:
    log_ratio_win = pi_logps_win - ref_logps_win
    log_ratio_lose = pi_logps_lose - ref_logps_lose
    logits = beta * (log_ratio_win - log_ratio_lose)
    loss = -F.logsigmoid(logits).mean()
    return loss

def compute_grpo_advantage(rewards: torch.Tensor, eps: float = 1e-8) -> torch.Tensor:
    mean = rewards.mean(dim=-1, keepdim=True)
    std = rewards.std(dim=-1, keepdim=True)
    return (rewards - mean) / (std + eps)

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 How does DeepSeek-R1's GRPO eliminate the need for a separate Critic (Value) network during RL fine-tuning?

Finished this chapter?