Linear regression
Fits a straight-line (hyperplane) relationship by least squares. The reference point every other model is measured against.
Two things get called βML architectureβ: the classical statistical model families, and the neural network wirings. Both are here, tagged by what they are and what data they eat.
Fits a straight-line (hyperplane) relationship by least squares. The reference point every other model is measured against.
Least squares plus an L2 penalty on the weights, which shrinks them and tames collinearity.
Least squares with an L1 penalty, which drives some coefficients exactly to zero.
L1 and L2 penalties combined, giving sparse selection without Lasso's instability on correlated groups.
Linear score squashed through a sigmoid into a probability; trained by maximum likelihood.
Linear predictor plus a link function and a noise family, so counts, rates and skewed targets all fit one framework.
Fits under a loss that shrugs off outliers (Huber, Tukey bisquare) or by trimming and resampling (RANSAC, Theil-Sen).
The original single-layer learning rule: nudge a weight vector whenever a point is misclassified.
Finds the maximum-margin separating hyperplane; with the kernel trick it becomes a nonlinear boundary without ever building the features.
Reformulates the SVM as a linear system, trading the margin-based quadratic program for a much simpler solve.
Learns a boundary around normal data only, flagging everything outside as anomalous.
Models each class as a Gaussian and classifies by posterior; LDA shares one covariance, QDA gives each class its own.
Recursively splits the feature space on whichever question most reduces impurity, producing a readable rule set.
Hundreds of de-correlated trees (bagged rows plus random feature subsets) averaged into one robust predictor.
Like a random forest but with random split thresholds as well as random feature subsets β more variance, less bias, faster.
Fits shallow trees sequentially, each one correcting the residual error of the running model.
A regularized, second-order, sparsity-aware implementation of gradient boosting that dominated tabular competitions.
Histogram-based leaf-wise boosting with GOSS and EFB β much faster training on large, sparse tabular sets.
Oblivious trees with ordered boosting, purpose-built to handle categorical features without leaking target statistics.
Reweights misclassified samples each round so a sequence of weak learners concentrates on hard cases.
Gradient boosting that predicts a full probability distribution rather than a point, using natural gradients.
Trains a meta-model on the out-of-fold predictions of several base models.
Anomalies get isolated by random splits in fewer steps, so average path length becomes the anomaly score.
Applies Bayes' rule while pretending features are independent given the class β wrong, but surprisingly effective.
Models data as a weighted sum of Gaussians, fit by Expectation-Maximization; gives soft cluster memberships.
A latent discrete state that evolves as a Markov chain and emits observations; decoding with Viterbi, fitting with Baum-Welch.
A discriminative undirected graphical model over a whole sequence, so labels can depend on each other globally.
Directed acyclic graphs encoding conditional independence between variables, with exact or approximate inference.
Undirected models where the joint distribution factors over cliques; inference by belief propagation or graph cuts.
A distribution over functions defined by a kernel; prediction is exact Bayesian inference with calibrated uncertainty.
Gaussian-process or tree-ensemble surrogates plus an acquisition function to pick the next experiment.
Documents as mixtures over topics, topics as distributions over words, inferred by variational or Gibbs sampling.
Energy-based stochastic networks; restricted forms train by contrastive divergence and stack into deep belief nets.
Turns intractable posterior integration into optimization by fitting a tractable approximation.
Alternates between assigning points to the nearest centroid and recomputing centroids until it settles.
Like k-means but with actual data points as centres, so any distance metric works and it resists outliers.
Builds a tree of merges (or splits); cut the dendrogram anywhere to get a chosen number of clusters.
Grows clusters from dense neighbourhoods and labels sparse points as noise; finds any shape, needs no k.
Hierarchical DBSCAN that extracts the most stable clusters across density levels, automatically.
Orders points by reachability distance, producing a plot you can read to choose epsilon and find nested structure.
Each point climbs the gradient of the local density estimate until it lands on a mode; modes are the clusters.
Uses the eigenvectors of a similarity graph's Laplacian to embed points, then clusters the embedding.
A grid of prototype vectors that compete and pull neighbours with them, producing a topology-preserving map.
Messages passed between points elect exemplars directly, so k emerges from the data.
Compares a point's local density to its neighbours'; much lower density means an outlier.
Rotates the data onto orthogonal directions of maximum variance, ranked so you can truncate.
Runs PCA in the feature space implied by a kernel, capturing nonlinear structure without ever forming it.
Low-rank factorization that works directly on a sparse term-document matrix, giving latent semantic axes.
Factors a non-negative matrix into two non-negative ones, producing parts-based additive representations.
Separates a mixture into statistically independent sources by maximizing non-Gaussianity.
Explains observed variables as linear functions of fewer latent factors plus per-variable noise.
Minimizes divergence between neighbour distributions in high and low dimension, laying clusters out for the eye.
Builds a fuzzy topological graph of the data, then optimizes a low-dimensional layout preserving it.
Neighbourhood-preserving embeddings designed to keep both local and global structure better than t-SNE or UMAP.
MDS preserves pairwise distances; Isomap preserves geodesic distances on a neighbourhood graph; LLE preserves local linear reconstruction.
Instead of shrinking the raw features, learn an embedding by solving a pretext task β contrastive pairs, masked modelling, augmentation invariance.
Stacked fully-connected layers with nonlinear activations, trained by backpropagation.
Hidden units that fire based on distance to a centre, making the network a learned kernel interpolator.
Each block learns a perturbation of its input through an identity skip, so gradient flows cleanly through hundreds of layers.
Gated skip connections that let a layer decide what fraction of the input to pass through unchanged.
Every layer receives the concatenated outputs of all preceding layers, maximizing feature reuse.
SELU activations plus the right initialization keep activations normalized, removing the need for batch norm.
A gating network routes each input to a few specialist sub-networks, so total parameters grow without compute growing with them.
One network generates the weights of another, letting a single model instantiate many specialized variants.
Learnable spline functions live on the edges, not the nodes, replacing fixed activations with trainable univariate functions.
Replaces discrete layers with a continuous vector field integrated by a solver, giving an ODE as the model.
Continuous-time recurrent units whose dynamics adapt to the input, with far fewer neurons than conventional nets.
Associative memory stored as an energy landscape where patterns are attractors of the dynamics.
Groups neurons into capsules that encode pose and instantiation, routed dynamically between layers.
Neurons communicate in discrete spikes with temporal dynamics, trainable by surrogate gradients or STDP.
A controller network coupled to differentiable external memory it learns to read and write.
Stores facts in an explicit memory and attends over it to answer queries β knowledge outside the weights.
Compress the input through a bottleneck and reconstruct it, learning a compact code in the middle.
Learn small weight-sharing filters that slide across the input, exploiting locality and translation invariance.
Two convolutional layers with pooling feeding a small MLP β the template every CNN still follows.
A deep CNN with ReLU, dropout and GPU training that halved ImageNet error and kicked off the deep learning era.
Depth built from nothing but 3x3 convolutions stacked in uniform blocks.
Parallel convolutions of different kernel sizes in one block, concatenated β multi-scale features at one depth.
Residual blocks at scale β ResNet-50/101/152 are the workhorse backbones that follow the skip-connection idea to depth.
Grouped convolutions (ResNeXt) or wider, shallower blocks (WideResNet) to trade depth for parallelism and width.
Depthwise-separable convolutions and channel shuffling cut multiply-adds by an order of magnitude.
Compound scaling: grow depth, width and resolution together in the ratio a small search found best.
A ResNet modernized with transformer-era choices β larger kernels, LayerNorm, GELU, fewer normalization stages.
Split the image into patches, treat each patch as a token, and run a plain transformer encoder over them.
ViT made trainable on ImageNet alone through distillation from a CNN teacher and heavy augmentation.
Windowed attention with shifted windows, so cost stays linear in image size while information crosses windows.
Mask most patches and train the encoder to reconstruct them β a self-supervised pretraining objective for ViTs.
Jointly embeds images and captions into one space with a contrastive objective, enabling zero-shot classification.
Replaces CLIP's softmax contrastive loss with a pairwise sigmoid loss, training better at small batch sizes.
Propose candidate regions, then classify each one with a CNN: R-CNN, Fast R-CNN, Faster R-CNN with an RPN.
Single-shot detection that predicts boxes and classes in one forward pass β real-time by design.
Multi-scale anchor-based single-shot heads; RetinaNet adds focal loss to fix the foreground/background imbalance.
Detection as set prediction: a transformer decoder with learned object queries, matched by Hungarian loss. No anchors, no NMS.
Fully-convolutional encoder-decoder with skip connections, producing a per-pixel output map.
Dilated convolutions and spatial pyramid pooling keep resolution while enlarging the receptive field.
A detection head plus a parallel branch predicting a per-instance binary mask.
A huge image encoder with a prompt-conditioned mask decoder that segments anything you point at or box.
A promptable unified segmentation model extending visual mask prediction to real-time streaming video through a spatio-temporal memory bank and memory attention.
Operate on point sets or voxels: PointNet's symmetric pooling, PointNet++, sparse 3D convolutions, 3D U-Net.
Extend vision models along time with 3D convolutions, (2+1)D factorisation, or spatio-temporal attention.
A shared hidden state carried through time, so the model sees a sequence as it unfolds.
Gated cell state with input, forget and output gates that let information persist across many steps.
LSTM simplified to two gates and one merged state β fewer parameters, similar accuracy.
Two RNNs, one reading forward and one backward, so every position sees full context.
LSTM modernized with exponential gating, a matrix memory and parallelizable training (sLSTM and mLSTM blocks).
An encoder RNN compresses the input, a decoder generates the output, and attention lets the decoder look back at every encoder state.
Dilated causal 1D convolutions stacked to cover long receptive fields, trained in parallel unlike an RNN.
A fixed random recurrent reservoir supplies rich dynamics; only the readout layer is trained.
A soft, content-based lookup: query a set of key-value pairs and read out a weighted average.
Stacked blocks of multi-head self-attention plus a position-wise feed-forward network, with residual connections and normalization.
Ways to dodge the quadratic attention matrix: sparse patterns, low-rank projections, kernels, clustering, or memory.
Most layers attend only locally while a few attend globally, cutting KV-cache cost at long context.
Masked-language pretraining over a bidirectional encoder; produces contextual embeddings, not generated text.
Causal self-attention trained to predict the next token; scaling this recipe produced the modern LLM.
An encoder reads the source and a cross-attending decoder writes the target, for any text-to-text task.
Transformers trained so a single pooled vector represents a sentence β bi-encoders with contrastive objectives.
Transformer blocks whose feed-forward layers are replaced by routed expert banks, activating a small slice per token.
Same transformer backbone, changed by post-training: instruction tuning, RLHF/DPO, then long chain-of-thought RL.
671B parameter Mixture-of-Experts (37B active) pairing Multi-Head Latent Attention (MLA) with fine-grained MoE routing, multi-token prediction, and large-scale reinforcement learning (GRPO).
A vision encoder, a projector, and a language model trained to talk about images (Flamingo, BLIP-2, LLaVA, Qwen-VL, InternVL).
The diffusion U-Net replaced by a transformer operating on latent patches, with adaptive normalization for the timestep.
Linear time-invariant recurrences with structured state matrices β RNN-like linear-time inference, trainable in parallel as a convolution.
The first generation of deep SSMs: diagonal or low-rank parameterizations of the state matrix with a parallel scan.
Makes the SSM parameters input-dependent (selective), so the model can choose what to remember or forget per token.
Restricts the state transition to a scalar-times-identity, exposing the recurrence as a matrix multiply (the SSD layer).
Bidirectional or hybrid SSM backbones for images and video, offering linear cost instead of quadratic attention.
An RNN whose recurrence is expressed as a linear-attention style kernel, so training parallelizes and inference is constant-memory.
Retention instead of attention β a decayed weighted sum that can be computed three ways: parallel, recurrent or chunkwise.
Real-gated linear recurrences mixed with local attention, packaged as a production open model.
Implicit long convolutions interleaved with gating, sub-quadratic in sequence length.
Alternates SSM and convolution blocks with a small amount of attention, tuned for throughput at long context.
Interleaves transformer and Mamba layers in one MoE model, so most layers are linear and a few do exact attention.
The empirical consensus: a mostly-linear stack with a minority of attention layers beats either pure design.
An autoencoder trained by maximizing a variational lower bound, giving a smooth continuous latent space you can sample.
A discrete codebook bottleneck turns continuous data into tokens that a prior model (often a transformer) can generate.
A generator and a discriminator trained against each other, the discriminator supplying the loss signal.
DCGAN made GANs trainable with convolutions; WGAN replaced the Jensen-Shannon objective with Wasserstein distance plus a Lipschitz constraint.
A style-based generator where a learned latent is injected at every resolution level, disentangling coarse from fine attributes.
Image-to-image translation, with CycleGAN's cycle-consistency letting it learn unpaired mappings.
Factor a signal into a factorized sequential distribution and predict it element by element: PixelCNN, PixelRNN, WaveNet.
Compose invertible transforms so the exact likelihood is computable and sampling is just an inverse pass.
Destroy data with noise over many steps and learn to reverse the process; sampling anneals noise back into a sample.
Generalizes diffusion as a stochastic differential equation over continuous time, trained on score matching.
Run the diffusion process inside a VQ-VAE latent space instead of pixel space, cutting cost by orders of magnitude.
Train one model on conditional and unconditional objectives, then interpolate between them at sampling for stronger conditioning.
Instead of reversing noise, learn a velocity field that moves samples along straight-ish paths from noise to data.
12B-parameter rectified flow transformer combining dual multimodal diffusion transformer (MMDiT) blocks with rotary position encodings for extreme visual fidelity and typography.
Trained or distilled so a sample is produced in one or a few steps instead of hundreds.
Learn an unnormalized energy function; low energy means likely data, and sampling means descending it.
Learns a stochastic policy whose sampling probability is proportional to a reward, giving diverse rather than single-best outputs.
Message passing over edges: each node aggregates its neighbours, repeated to propagate information across hops.
A spectral filter approximated by a first-order polynomial, which reduces to a normalized neighbourhood average.
Samples a fixed number of neighbours and concatenates self with aggregated neighbours β inductive, so new nodes work.
Attention weights over neighbours instead of fixed normalization, so important edges dominate.
Sum-aggregation message passing with a provably maximal expressive power for GNNs (bounded by 1-WL).
Attention over all nodes with structural encodings (shortest paths, Laplacians, random walks) injected as bias.
Architectures that respect 3D symmetry by construction β rotations and translations transform the output the same way they transform the input.
Embeds trees and hierarchies in negatively-curved space where volume grows exponentially with radius.
Learns the value of each action in each state; DQN approximates the Q-function with a CNN stabilized by replay and a target network.
Double DQN de-biases the max operator; dueling splits state value from advantage; prioritized replay samples surprising transitions more.
A policy (actor) proposes actions while a value function (critic) estimates how good they were, reducing gradient variance.
Many parallel workers each collect experience and update a shared policy asynchronously or in lockstep.
Trust-region policy optimization simplified into a clipped ratio objective that is safe enough to run for many epochs.
Deterministic policy gradients for continuous actions; TD3 adds twin critics and target smoothing to tame Q-value overestimation.
Maximum-entropy RL: maximize reward while staying as stochastic as possible, with twin critics and automatic temperature tuning.
Learn a model of the environment's dynamics and plan or train inside the learned model.
Learn a compressed latent model of the environment that can be rolled forward to imagine trajectories.
Self-play plus Monte Carlo tree search guided by a single network that predicts both policy and value.
Several learning agents share an environment; centralised critics with decentralised actors are the usual answer.
Train a reward model on human preferences, then optimize the policy against it, usually with a KL penalty.
Skips the reward model and the RL loop by turning the preference objective into a simple classification loss on the policy.
Drops the critic and normalizes advantages within a group of sampled answers; verifiable rewards (math, code tests) carry the signal.
Sequential decision-making with immediate reward only: explore to learn, exploit to earn.
Several attention heads run in parallel on projected subspaces, then concatenate β different heads learn different relations.
Share one key/value projection across groups of query heads, shrinking the KV cache by 4β8x at almost no quality cost.
Compresses keys and values into a low-rank latent vector before caching, cutting the KV cache far beyond GQA.
The 2026 frontier of cache reduction: reuse KV tensors across layers, budget attention per layer, compress via convolution.
Inject position by rotating query/key pairs by an angle proportional to position (RoPE), or by adding learned relative biases (ALiBi, T5).
BatchNorm normalizes across the batch, LayerNorm and RMSNorm across features β stabilizing optimization and enabling higher learning rates.
The nonlinearity between linear layers: ReLU, LeakyReLU, ELU, GELU, SiLU/Swish, Mish, and the gated GLU family.
The running sum that every block reads from and writes to β and new plumbing (hyper-connections, mHC) that widens it.
The per-token MLP that holds most parameters, usually widened 4x and gated in modern models.
Empirical power laws linking loss to parameters, data and compute β plus the finding that data matters more than previously assumed.
Byte-pair encoding, WordPiece or Unigram turn text into subword tokens β the model's actual alphabet.
Neural architectures designed for tables rather than repurposed from images: feature-tokenizing transformers and attentive flows.
Forecasting architectures beyond ARIMA: N-BEATS, N-HiTS, TFT, DeepAR, and patching transformers like PatchTST and TimesNet.
Large pretrained forecasters trained across many domains that zero-shot a new series without fitting.
From matrix factorization and BPR through factorization machines (DeepFM, DLRM) to two-tower retrieval and sequential SASRec/BERT4Rec.
Encode each item as a semantic ID and let a sequence model generate the next item's ID directly.
CTC and RNN-Transducer for alignment-free training, LAS for attention-based decoding, Conformer for the current encoder.
Encoder-decoder Transformer trained on 680,000 hours of weakly-supervised multilingual audio for robust zero-shot speech recognition, translation, and voice activity detection.
Text to waveform through neural vocoders (WaveNet, HiFi-GAN), end-to-end TTS (Tacotron, FastSpeech), and neural audio codecs (SoundStream, EnCodec).
Equivariant graph networks and attention over residues (AlphaFold2's Evoformer, ESM, RoseTTAFold, AlphaFold3's diffusion module).
Add a residual term for the governing differential equation, so the network is penalised for violating known physics.
Learn mappings between function spaces rather than between finite vectors, so one model generalizes across resolutions.
Large graph or transformer networks trained on reanalysis data that now beat numerical weather prediction at medium range.
Pair a retriever with a generator so knowledge lives in an index rather than only in weights (RAG, REALM, FiD, kNN-LM).
The unifying framework: architectures are defined by which symmetries they respect over which domain.