💻 Modern ArchitecturesLEVEL 4 · FRONTIERChapter 11
Post-Training RL for Reasoning LLMs (DPO & DeepSeek-R1 GRPO)
Direct Preference Optimization (DPO), Group Relative Policy Optimization (GRPO), and self-reflection.
💡 1. Core Intuition & Concepts
Modern reasoning models like DeepSeek-R1, OpenAI o1, and Claude 3.7 Sonnet achieve emergent chain-of-thought verification using post-training reinforcement learning. DeepSeek-R1 introduced Group Relative Policy Optimization (GRPO), computing relative advantages across a group of sampled responses to eliminate the value critic network.
📐 2. Mathematical Formulations & Derivations
Direct Preference Optimization (DPO) Loss
Optimizes preferences directly on policy logits without requiring a separate reward model.
DeepSeek-R1 GRPO Group Advantage
Normalizes rewards across a group of G sampled response rollouts for the same prompt.
⚙️ 3. Step-by-Step Computational Mechanism
1
Sample Candidate Group
For prompt q, sample a group of G response trajectories {o_1, o_2, ..., o_G} from current policy π_θ.
2
Evaluate Rule-Based Rewards
Score each response with binary mathematical correctness, format adherence, or compiler tests r_i.
3
Compute Group-Relative Advantage
Calculate standardized advantage A_i = (r_i - r_bar) / (sigma_r + eps).
4
PPO-Clipped Policy Update with KL Penalty
Update policy parameters θ while constraining drift from reference model π_ref.
💻 4. Code from Scratch (python)
rl_grpo.py
import torch
import torch.nn.functional as F
def compute_dpo_loss(
pi_logps_win: torch.Tensor,
pi_logps_lose: torch.Tensor,
ref_logps_win: torch.Tensor,
ref_logps_lose: torch.Tensor,
beta: float = 0.1
) -> torch.Tensor:
log_ratio_win = pi_logps_win - ref_logps_win
log_ratio_lose = pi_logps_lose - ref_logps_lose
logits = beta * (log_ratio_win - log_ratio_lose)
loss = -F.logsigmoid(logits).mean()
return loss
def compute_grpo_advantage(rewards: torch.Tensor, eps: float = 1e-8) -> torch.Tensor:
mean = rewards.mean(dim=-1, keepdim=True)
std = rewards.std(dim=-1, keepdim=True)
return (rewards - mean) / (std + eps)🧠 5. Comprehension Checkpoint
Answer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 How does DeepSeek-R1's GRPO eliminate the need for a separate Critic (Value) network during RL fine-tuning?
In the catalog
Finished this chapter?