About Epoch Zero
Epoch Zero is a free site for learning machine learning from scratch. It has three parts:
lessons that start from first principles, a field guide to model
architectures you can search, filter and compare, and project write-ups.
You can use all of it without an account. Signing in only keeps your progress and bookmarks in sync between
devices. See what we store.
How to read any of this
“Architecture” is almost never a single brand name — it's a stack of independent choices. Two models
with different names usually differ in exactly one row of this grid, which is why every 2025–26 LLM
release looks like a small edit inside one transformer block.
The choice grid
| Axis | Options you'll see |
| Token mixer | convolution, recurrence, attention, state space, long convolution, or a mix |
| Residual stream | plain, residual, dense, hyper-connections / mHC |
| Normalization | BatchNorm, LayerNorm, RMSNorm, QK-norm, GroupNorm, DeepNorm |
| Activation | ReLU, GELU, SiLU, GLU family, SwiGLU/GEGLU |
| Positional scheme | sinusoidal, learned, relative, RoPE (+ scaling), ALiBi, NoPE |
| Attention shape | MHA → MQA → GQA → MLA, sliding window, sparse, compressed |
| Sparsity | dense, mixture-of-experts (top-k, fine-grained, shared experts) |
| Objective | supervised, self-supervised, contrastive, masked, autoregressive, diffusion, RL |
How to pick, in practice
- Tabular data: gradient-boosted trees win most of the time. Reach for neural tabular models
only when you have many columns, mixed modality, or lots of data.
- Images: a modern CNN backbone or a ViT, pretrained. Detect/segment with the task-specific heads.
- Text: a pretrained transformer you fine-tune or prompt. Don't train from scratch.
- Long sequences / streaming at scale: the state-space and hybrid family exists exactly for this.
- Few hundred rows: linear models, k-NN and trees. A deep net will overfit and you'll blame the data.
- Anything generative: pick the family by output type — diffusion for images/audio, autoregressive
for text/code, flows and VAEs when you need a likelihood or a smooth latent space.
Where to keep reading
LLM Architecture Gallery — Sebastian Raschka
Hand-drawn diagrams of every major open-weight release, updated as they ship.
The Big LLM Architecture Comparison
Side-by-side of DeepSeek, Llama, Qwen, Gemma, Grok and friends: attention, MoE, normalization.
A Visual Guide to Attention Variants
MHA, MQA, GQA, MLA, sliding-window, sparse and hybrid attention, drawn out.
Recent Developments: KV Sharing, mHC, Compressed Attention
Gemma 4, ZAYA1, Laguna XS.2 and DeepSeek V4 — the 2026 long-context tricks.
scikit-learn — Supervised learning
The canonical taxonomy and implementation notes for every classical family here.
State-space LLMs: Do we need attention?
Mamba, StripedHyena, Based and the case for attention-free sequence models.
A Survey on Mixture of Experts in LLMs
Constant-updated paper list on routing, expert design and MoE training.
arXiv cs.LG — recent
Where the next entry on this page will come from.
About the links on each card
Architecture names are stable; paper URLs are not. Every entry stores its canonical paper
title, and each title is resolved against public bibliographic indexes — preferring an arXiv
landing page, then a DOI, then the publisher's page. Titles that cannot be matched confidently link to a
targeted search instead, rather than sending you to a guessed identifier that may have rotted. The verifier
also fetches a sample of the resolved URLs on every build to confirm they still resolve.