Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
📐 The Math Behind MLLEVEL 4 · FRONTIERChapter 7

Transformer Mathematics: Attention Scaling, RoPE & Latent Projections

Attention variance scaling, SO(2) Givens rotation (RoPE), and Multi-Head Latent Attention (MLA).

⏱ 25 min read🎯 Prerequisites: Complex Numbers & Matrix Calculus
🔍 Inspect Architecture: Attention mechanism

💡 1. Core Intuition & Concepts

Modern transformer architectures rely on precise geometric principles: scaling dot-products by 1/√d_k preserves variance and prevents softmax saturation, Rotary Position Embeddings (RoPE) encode relative token distance via complex plane Givens rotations, and Multi-Head Latent Attention (MLA) uses low-rank projections to compress the KV-cache by 93%.

📐 2. Mathematical Formulations & Derivations

Attention Dot-Product Variance & Scaling Factor
qi,ki∼N(0,1) i.i.d.  ⟹  Var(qTk)=dk  ⟹  Var(qTkdk)=1q_i, k_i \sim \mathcal{N}(0, 1) \text{ i.i.d.} \implies \text{Var}(q^T k) = d_k \implies \text{Var}\left( \frac{q^T k}{\sqrt{d_k}} \right) = 1
Dividing by √d_k prevents extreme softmax logits where gradients vanish (∂softmax/∂z ≈ 0).
Rotary Position Embedding (RoPE) 2D Givens Rotation
RΘ,m(2D)=(cos⁡(mθ)−sin⁡(mθ)sin⁡(mθ)cos⁡(mθ))  ⟹  ⟨Rmq,Rnk⟩=qTRn−mkR_{\Theta, m}^{(2D)} = \begin{pmatrix} \cos(m\theta) & -\sin(m\theta) \\ \sin(m\theta) & \cos(m\theta) \end{pmatrix} \implies \langle R_m q, R_n k \rangle = q^T R_{n-m} k
Encodes relative token displacement (n-m) purely through orthogonal rotation matrix multiplication.
Multi-Head Latent Attention (MLA) KV Compression
ctKV=WDKVxt∈RdcktC=WUKctKV,  vtC=WUVctKV(dc≪nhdh)c_t^{KV} = W_{DKV} x_t \in \mathbb{R}^{d_c} \qquad k_t^C = W_{UK} c_t^{KV}, \; v_t^C = W_{UV} c_t^{KV} \qquad (d_c \ll n_h d_h)
DeepSeek-V3 compresses key-value vectors into a low-rank latent state c_t^{KV}, slashing KV cache footprint.

• Represent a 2D query vector in complex form q = q_0 + i q_1 and key vector k = k_0 + i k_1.

• Applying rotation at token position m corresponds to multiplication by complex exponential: q_m = q * e^(i m theta), and at position n: k_n = k * e^(i n theta).

• Compute the inner product as the real part of the complex product with conjugate: <q_m, k_n> = Re(q_m * conj(k_n)).

• Substitute the rotated forms: Re((q * e^(i m theta)) * (conj(k) * e^(-i n theta))) = Re((q * conj(k)) * e^(i (m - n) theta)).

• The resulting dot product depends exclusively on the relative token displacement (m - n) and the base frequency theta. Q.E.D.

⚙️ 3. Step-by-Step Computational Mechanism

1
Variance Normalization
Unscaled dot-products grow with dimension d_k, pushing softmax into saturated regions with zero gradients.
2
Orthogonal Relative Encoding
RoPE requires zero learned parameters, generalizes to unseen context lengths, and decays naturally with relative distance.
3
Decoupled Latent Projection
MLA splits query/key into content vectors (compressed via low-rank W_DKV) and position vectors (rotated via RoPE).

💻 4. Code from Scratch (python)

math_transformer_rope.py
import numpy as np

# Rotary Position Embedding (RoPE) 2D Complex Rotation from Scratch
def apply_rope_2d(x: np.ndarray, seq_len: int, dim: int, base: float = 10000.0):
    dim_pairs = dim // 2
    inv_freq = 1.0 / (base ** (np.arange(0, dim_pairs) * 2.0 / dim))
    positions = np.arange(seq_len)
    
    angles = np.outer(positions, inv_freq)
    cos_vals = np.cos(angles)
    sin_vals = np.sin(angles)

    x_pairs = x.reshape(seq_len, dim_pairs, 2)
    x0, x1 = x_pairs[:, :, 0], x_pairs[:, :, 1]

    rot_x0 = x0 * cos_vals - x1 * sin_vals
    rot_x1 = x0 * sin_vals + x1 * cos_vals

    return np.stack([rot_x0, rot_x1], axis=-1).reshape(seq_len, dim)

if __name__ == "__main__":
    seq_len, dim = 4, 8
    q = np.random.randn(seq_len, dim)
    q_rot = apply_rope_2d(q, seq_len, dim)
    print("RoPE rotated Query shape:", q_rot.shape)
    print("✓ RoPE 2D Rotation applied successfully.")

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 Why is the dot-product in attention scaled specifically by 1 / sqrt(d_k)?

Finished this chapter?