💻 Modern ArchitecturesLEVEL 4 · FRONTIERChapter 10
Diffusion Models & Rectified Flow Matching
Continuous-time ODE trajectories, straight paths, and FLUX.1 velocity matching.
🔍 Inspect Architecture: Flow matching / continuous normalizing flows💡 1. Core Intuition & Concepts
Modern image and video generation models (FLUX.1, Stable Diffusion 3, Sora) have transitioned from stochastic DDPM diffusion to Rectified Flow Matching, learning straight ODE velocity trajectories connecting noise to data.
📐 2. Mathematical Formulations & Derivations
Rectified Flow Linear Interpolation
x_0: pure Gaussian noise, x_1: target training image, x_t: intermediate sample.
Flow Matching Velocity Objective
Neural network v_θ learns the direct velocity vector pointing from noise to data.
⚙️ 3. Step-by-Step Computational Mechanism
1
Sample Noise & Data Pair
Draw clean data x_1 ~ p_data and noise x_0 ~ N(0, I).
2
Interpolate at Random Time t
Sample timestep t ~ U(0, 1) and compute straight trajectory point x_t = (1-t)x_0 + t x_1.
3
Predict Velocity Field
Pass noisy image x_t and timestep t through transformer backbone to predict velocity v_θ(x_t, t).
4
Euler ODE Generation
Sample by integrating dx/dt = v_θ(x, t) from t=0 to t=1 in 4 to 20 steps.
💻 4. Code from Scratch (python)
diffusion.py
import torch
import torch.nn as nn
class RectifiedFlowTrainer:
"""Rectified Flow Matching (FLUX.1 / SD3 style) Training Step in PyTorch."""
def __init__(self, model: nn.Module):
self.model = model
def training_step(self, x_1: torch.Tensor) -> torch.Tensor:
B = x_1.shape[0]
x_0 = torch.randn_like(x_1)
t = torch.rand((B, 1, 1, 1), device=x_1.device)
x_t = (1.0 - t) * x_0 + t * x_1
target_v = x_1 - x_0
pred_v = self.model(x_t, t.squeeze())
loss = torch.mean((pred_v - target_v) ** 2)
return loss
@torch.no_grad()
def sample_euler(self, noise: torch.Tensor, steps: int = 10) -> torch.Tensor:
dt = 1.0 / steps
x_t = noise
for step in range(steps):
t = torch.full((noise.shape[0],), step * dt, device=noise.device)
v = self.model(x_t, t)
x_t = x_t + v * dt
return x_t🧠 5. Comprehension Checkpoint
Answer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 Why does Rectified Flow Matching enable high-quality generation in fewer sampling steps (4-10 steps) than older DDPM diffusion models (50-1000 steps)?
In the catalog
Finished this chapter?