Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
💻 Modern ArchitecturesLEVEL 4 · FRONTIERChapter 10

Diffusion Models & Rectified Flow Matching

Continuous-time ODE trajectories, straight paths, and FLUX.1 velocity matching.

⏱ 22 min read🎯 Prerequisites: Probability, ODEs & Transformer blocks
🔍 Inspect Architecture: Flow matching / continuous normalizing flows

💡 1. Core Intuition & Concepts

Modern image and video generation models (FLUX.1, Stable Diffusion 3, Sora) have transitioned from stochastic DDPM diffusion to Rectified Flow Matching, learning straight ODE velocity trajectories connecting noise to data.

📐 2. Mathematical Formulations & Derivations

Rectified Flow Linear Interpolation
xt=(1−t)x0+tx1t∈[0,1]x_t = (1 - t) x_0 + t x_1 \qquad t \in [0, 1]
x_0: pure Gaussian noise, x_1: target training image, x_t: intermediate sample.
Flow Matching Velocity Objective
LFM(θ)=Et,x0,x1[∥vθ(xt,t)−(x1−x0)∥2]\mathcal{L}_{\text{FM}}(\theta) = \mathbb{E}_{t, x_0, x_1}\left[ \| v_\theta(x_t, t) - (x_1 - x_0) \|^2 \right]
Neural network v_θ learns the direct velocity vector pointing from noise to data.

⚙️ 3. Step-by-Step Computational Mechanism

1
Sample Noise & Data Pair
Draw clean data x_1 ~ p_data and noise x_0 ~ N(0, I).
2
Interpolate at Random Time t
Sample timestep t ~ U(0, 1) and compute straight trajectory point x_t = (1-t)x_0 + t x_1.
3
Predict Velocity Field
Pass noisy image x_t and timestep t through transformer backbone to predict velocity v_θ(x_t, t).
4
Euler ODE Generation
Sample by integrating dx/dt = v_θ(x, t) from t=0 to t=1 in 4 to 20 steps.

💻 4. Code from Scratch (python)

diffusion.py
import torch
import torch.nn as nn

class RectifiedFlowTrainer:
    """Rectified Flow Matching (FLUX.1 / SD3 style) Training Step in PyTorch."""
    def __init__(self, model: nn.Module):
        self.model = model

    def training_step(self, x_1: torch.Tensor) -> torch.Tensor:
        B = x_1.shape[0]
        x_0 = torch.randn_like(x_1)
        t = torch.rand((B, 1, 1, 1), device=x_1.device)
        x_t = (1.0 - t) * x_0 + t * x_1
        target_v = x_1 - x_0
        pred_v = self.model(x_t, t.squeeze())
        loss = torch.mean((pred_v - target_v) ** 2)
        return loss

    @torch.no_grad()
    def sample_euler(self, noise: torch.Tensor, steps: int = 10) -> torch.Tensor:
        dt = 1.0 / steps
        x_t = noise
        for step in range(steps):
            t = torch.full((noise.shape[0],), step * dt, device=noise.device)
            v = self.model(x_t, t)
            x_t = x_t + v * dt
        return x_t

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 Why does Rectified Flow Matching enable high-quality generation in fewer sampling steps (4-10 steps) than older DDPM diffusion models (50-1000 steps)?

Finished this chapter?