Matrix Decompositions, SVD & Low-Rank Adaptation (LoRA)
Spectral Theorem, Singular Value Decomposition (SVD), Eckart-Young Theorem, and LoRA math.
💡 1. Core Intuition & Concepts
Weight matrices in large language models possess high nominal dimensionality but exhibit low intrinsic rank during task adaptation. Singular Value Decomposition (SVD) and the Eckart-Young-Mirsky theorem prove that weight updates ΔW can be factored into two low-rank matrices B * A with r << d, slashing trainable parameter counts by 99% with zero loss in representational capacity.
📐 2. Mathematical Formulations & Derivations
• Let A = U Sigma V^T be the full SVD of A in R^(m x n). Define A_k = sum_{i=1}^k sigma_i u_i v_i^T.
• The approximation error is A - A_k = sum_{i=k+1}^r sigma_i u_i v_i^T.
• Compute the squared Frobenius norm of the residual error: ||A - A_k||_F^2 = Tr((A - A_k)^T (A - A_k)) = sum_{i=k+1}^r sigma_i^2.
• By the Courant-Fischer min-max theorem for singular values, any arbitrary matrix B with rank(B) <= k satisfies ||A - B||_2 >= sigma_{k+1}.
• Since ||A - A_k||_2 = sigma_{k+1}, A_k attains the theoretical lower bound for both spectral and Frobenius norms. Q.E.D.
⚙️ 3. Step-by-Step Computational Mechanism
💻 4. Code from Scratch (python)
import numpy as np
# SVD Compression and LoRA Forward/Backward from Scratch
class LoRALayerScratch:
def __init__(self, d_in: int, d_out: int, rank: int = 4, alpha: float = 8.0):
self.d_in = d_in
self.d_out = d_out
self.rank = rank
self.scaling = alpha / rank
# Frozen pretrained weight
self.W0 = np.random.randn(d_out, d_in) * 0.02
# LoRA adapters: A is Gaussian, B is exact zero
self.A = np.random.randn(rank, d_in) * (1.0 / np.sqrt(d_in))
self.B = np.zeros((d_out, rank))
def forward(self, x: np.ndarray) -> np.ndarray:
# h = W0 x + (alpha/r) * B (A x)
self.x = x
self.Ax = np.dot(x, self.A.T) # (batch, rank)
base_out = np.dot(x, self.W0.T)
lora_out = np.dot(self.Ax, self.B.T) * self.scaling
return base_out + lora_out
if __name__ == "__main__":
layer = LoRALayerScratch(d_in=1024, d_out=1024, rank=8)
x = np.random.randn(2, 1024)
out = layer.forward(x)
print(f"Pretrained params: {1024*1024:,} | LoRA params: {1024*8 + 8*1024:,} (98.4% reduction)")
print(f"Forward output shape: {out.shape}")🧠 5. Comprehension Checkpoint
In the catalog
Finished this chapter?