The architecture landscape, in one place.

Two things get called β€œML architecture”: the classical statistical model families, and the neural network wirings. Both are here, tagged by what they are and what data they eat.

Start from your problem

Architectures

Linear models7

Linear regression

1805Linear modelsclassicaltabularsignalsany domain

Fits a straight-line (hyperplane) relationship by least squares. The reference point every other model is measured against.

Ridge regression

1970Linear modelsclassicaltabularsignals

Least squares plus an L2 penalty on the weights, which shrinks them and tames collinearity.

Lasso regression

1996Linear modelsclassicaltabularsignals

Least squares with an L1 penalty, which drives some coefficients exactly to zero.

Elastic Net

2005Linear modelsclassicaltabularsignals

L1 and L2 penalties combined, giving sparse selection without Lasso's instability on correlated groups.

Logistic regression

1958Linear modelsclassicaltabulartextany domain

Linear score squashed through a sigmoid into a probability; trained by maximum likelihood.

Generalized linear models

1972Linear modelsclassicaltabular

Linear predictor plus a link function and a noise family, so counts, rates and skewed targets all fit one framework.

Robust regression

1964Linear modelsclassicaltabularsignals

Fits under a loss that shrugs off outliers (Huber, Tukey bisquare) or by trimming and resampling (RANSAC, Theil-Sen).

Kernel & margin5

Perceptron

1958Kernel & marginneuraltabulartext

The original single-layer learning rule: nudge a weight vector whenever a point is misclassified.

Support vector machines

1995Kernel & marginclassicaltabulartextimages

Finds the maximum-margin separating hyperplane; with the kernel trick it becomes a nonlinear boundary without ever building the features.

Least-squares SVM / LS-SVM

1999Kernel & marginclassicaltabular

Reformulates the SVM as a linear system, trading the margin-based quadratic program for a much simpler solve.

One-class SVM

1999Kernel & marginclassicaltabularsignals

Learns a boundary around normal data only, flagging everything outside as anomalous.

Trees & ensembles11

Decision tree (CART, ID3, C4.5)

1984Trees & ensemblesclassicaltabularany domain

Recursively splits the feature space on whichever question most reduces impurity, producing a readable rule set.

Random Forest

2001Trees & ensemblesclassicaltabularsignalsimages

Hundreds of de-correlated trees (bagged rows plus random feature subsets) averaged into one robust predictor.

Extra Trees

2006Trees & ensemblesclassicaltabularsignals

Like a random forest but with random split thresholds as well as random feature subsets β€” more variance, less bias, faster.

Gradient boosting machine

1999Trees & ensemblesclassicaltabularrecommendersany domain

Fits shallow trees sequentially, each one correcting the residual error of the running model.

XGBoost

2016Trees & ensemblesclassicaltabularrecommenders

A regularized, second-order, sparsity-aware implementation of gradient boosting that dominated tabular competitions.

LightGBM

2017Trees & ensemblesclassicaltabularrecommenders

Histogram-based leaf-wise boosting with GOSS and EFB β€” much faster training on large, sparse tabular sets.

CatBoost

2017Trees & ensemblesclassicaltabular

Oblivious trees with ordered boosting, purpose-built to handle categorical features without leaking target statistics.

AdaBoost

1995Trees & ensemblesclassicaltabularimages

Reweights misclassified samples each round so a sequence of weak learners concentrates on hard cases.

NGBoost

2019Trees & ensemblesclassicaltabular

Gradient boosting that predicts a full probability distribution rather than a point, using natural gradients.

Stacking and blending

1992Trees & ensemblesclassicaltabularany domain

Trains a meta-model on the out-of-fold predictions of several base models.

Isolation Forest

2008Trees & ensemblesclassicaltabularsignals

Anomalies get isolated by random splits in fewer steps, so average path length becomes the anomaly score.

Probabilistic & graphical11

Naive Bayes

1958Probabilistic & graphicalclassicaltexttabular

Applies Bayes' rule while pretending features are independent given the class β€” wrong, but surprisingly effective.

Gaussian mixture models (EM)

1977Probabilistic & graphicalclassicaltabularsignalsspeech

Models data as a weighted sum of Gaussians, fit by Expectation-Maximization; gives soft cluster memberships.

Hidden Markov models

1966Probabilistic & graphicalclassicalspeechtime seriestext

A latent discrete state that evolves as a Markov chain and emits observations; decoding with Viterbi, fitting with Baum-Welch.

Conditional random fields

2001Probabilistic & graphicalclassicaltexttime series

A discriminative undirected graphical model over a whole sequence, so labels can depend on each other globally.

Bayesian networks

1985Probabilistic & graphicalclassicaltabularany domain

Directed acyclic graphs encoding conditional independence between variables, with exact or approximate inference.

Markov and factor graphs

1974Probabilistic & graphicalclassicalimagessignals

Undirected models where the joint distribution factors over cliques; inference by belief propagation or graph cuts.

Gaussian processes

1973Probabilistic & graphicalclassicaltabularmoleculessignals

A distribution over functions defined by a kernel; prediction is exact Bayesian inference with calibrated uncertainty.

Bayesian optimization surrogates

1978Probabilistic & graphicalclassicalcontrol / agentsany domain

Gaussian-process or tree-ensemble surrogates plus an acquisition function to pick the next experiment.

Latent Dirichlet Allocation

2003Probabilistic & graphicalclassicaltext

Documents as mixtures over topics, topics as distributions over words, inferred by variational or Gibbs sampling.

Boltzmann machines / RBM / DBN

1986Probabilistic & graphicalclassicaltabularimages

Energy-based stochastic networks; restricted forms train by contrastive divergence and stack into deep belief nets.

Variational inference

1999Probabilistic & graphicalclassicaltabularany domain

Turns intractable posterior integration into optimization by fitting a tractable approximation.

Clustering11

k-means

1967Clusteringclassicaltabularsignalsimages

Alternates between assigning points to the nearest centroid and recomputing centroids until it settles.

k-medoids / PAM

1987Clusteringclassicaltabular

Like k-means but with actual data points as centres, so any distance metric works and it resists outliers.

Hierarchical clustering

1963Clusteringclassicaltabulartextmolecules

Builds a tree of merges (or splits); cut the dendrogram anywhere to get a chosen number of clusters.

DBSCAN

1996Clusteringclassicaltabularimages

Grows clusters from dense neighbourhoods and labels sparse points as noise; finds any shape, needs no k.

HDBSCAN

2013Clusteringclassicaltabulartext

Hierarchical DBSCAN that extracts the most stable clusters across density levels, automatically.

OPTICS

1999Clusteringclassicaltabular

Orders points by reachability distance, producing a plot you can read to choose epsilon and find nested structure.

Mean shift

1975Clusteringclassicalimagestabular

Each point climbs the gradient of the local density estimate until it lands on a mode; modes are the clusters.

Spectral clustering

1973Clusteringclassicalimagesgraphs

Uses the eigenvectors of a similarity graph's Laplacian to embed points, then clusters the embedding.

Self-organizing maps

1982Clusteringclassicaltabularsignals

A grid of prototype vectors that compete and pull neighbours with them, producing a topology-preserving map.

Affinity propagation

2007Clusteringclassicaltabulartext

Messages passed between points elect exemplars directly, so k emerges from the data.

Local Outlier Factor

2000Clusteringclassicaltabularsignals

Compares a point's local density to its neighbours'; much lower density means an outlier.

Dimensionality reduction11

Principal component analysis

1901Dimensionality reductionclassicaltabularimagessignals

Rotates the data onto orthogonal directions of maximum variance, ranked so you can truncate.

Kernel PCA

1998Dimensionality reductionclassicalimagestabular

Runs PCA in the feature space implied by a kernel, capturing nonlinear structure without ever forming it.

Truncated SVD / LSA

1988Dimensionality reductionclassicaltext

Low-rank factorization that works directly on a sparse term-document matrix, giving latent semantic axes.

Non-negative matrix factorization

1999Dimensionality reductionclassicaltextaudioimages

Factors a non-negative matrix into two non-negative ones, producing parts-based additive representations.

Independent component analysis

1994Dimensionality reductionclassicalsignalsaudio

Separates a mixture into statistically independent sources by maximizing non-Gaussianity.

Factor analysis

1904Dimensionality reductionclassicaltabularany domain

Explains observed variables as linear functions of fewer latent factors plus per-variable noise.

t-SNE

2008Dimensionality reductionclassicalany domainimagestext

Minimizes divergence between neighbour distributions in high and low dimension, laying clusters out for the eye.

UMAP

2018Dimensionality reductionclassicalany domaintextimages

Builds a fuzzy topological graph of the data, then optimizes a low-dimensional layout preserving it.

PaCMAP / TriMap

2021Dimensionality reductionclassicalany domain

Neighbourhood-preserving embeddings designed to keep both local and global structure better than t-SNE or UMAP.

Classical MDS and Isomap / LLE

1952Dimensionality reductionclassicalany domain

MDS preserves pairwise distances; Isomap preserves geodesic distances on a neighbourhood graph; LLE preserves local linear reconstruction.

Self-supervised representation learning

2018Dimensionality reductionhybridany domainimagestext

Instead of shrinking the raw features, learn an embedding by solving a pretext task β€” contrastive pairs, masked modelling, augmentation invariance.

Neural foundations17

Multi-layer perceptron

1986Neural foundationsneuraltabularany domain

Stacked fully-connected layers with nonlinear activations, trained by backpropagation.

Radial basis function network

1988Neural foundationsneuraltabularcontrol / agents

Hidden units that fire based on distance to a centre, making the network a learned kernel interpolator.

Residual networks

2015Neural foundationsneuralimagesany domain

Each block learns a perturbation of its input through an identity skip, so gradient flows cleanly through hundreds of layers.

Highway networks

2015Neural foundationsneuralany domain

Gated skip connections that let a layer decide what fraction of the input to pass through unchanged.

DenseNet

2016Neural foundationsneuralimages

Every layer receives the concatenated outputs of all preceding layers, maximizing feature reuse.

Self-normalizing networks

2017Neural foundationsneuraltabular

SELU activations plus the right initialization keep activations normalized, removing the need for batch norm.

Mixture of experts

1991Neural foundationshybridtextimagesany domain

A gating network routes each input to a few specialist sub-networks, so total parameters grow without compute growing with them.

Hypernetworks

2016Neural foundationsneuralany domain

One network generates the weights of another, letting a single model instantiate many specialized variants.

Kolmogorov-Arnold networks

2024Neural foundationsneuraltabularmolecules

Learnable spline functions live on the edges, not the nodes, replacing fixed activations with trainable univariate functions.

Neural ODEs

2018Neural foundationsneuraltime seriescontrol / agentsany domain

Replaces discrete layers with a continuous vector field integrated by a solver, giving an ODE as the model.

Liquid neural networks

2020Neural foundationsneuraltime seriescontrol / agents

Continuous-time recurrent units whose dynamics adapt to the input, with far fewer neurons than conventional nets.

Hopfield networks

1982Neural foundationsneuralany domain

Associative memory stored as an energy landscape where patterns are attractors of the dynamics.

Capsule networks

2017Neural foundationsneuralimages

Groups neurons into capsules that encode pose and instantiation, routed dynamically between layers.

Spiking neural networks

1997Neural foundationsneuralimagessignals

Neurons communicate in discrete spikes with temporal dynamics, trainable by surrogate gradients or STDP.

Neural Turing Machine / DNC

2014Neural foundationsneuraltextany domain

A controller network coupled to differentiable external memory it learns to read and write.

Memory networks

2014Neural foundationsneuraltext

Stores facts in an explicit memory and attends over it to answer queries β€” knowledge outside the weights.

Autoencoders

1986Neural foundationsneuralany domainimagestabular

Compress the input through a bottleneck and reconstruct it, learning a compact code in the middle.

Convolutional & vision10

Convolutional neural networks

1989Convolutional & visionneuralimagesvideosignals

Learn small weight-sharing filters that slide across the input, exploiting locality and translation invariance.

LeNet-5

1998Convolutional & visionneuralimages

Two convolutional layers with pooling feeding a small MLP β€” the template every CNN still follows.

AlexNet

2012Convolutional & visionneuralimages

A deep CNN with ReLU, dropout and GPU training that halved ImageNet error and kicked off the deep learning era.

VGG

2014Convolutional & visionneuralimages

Depth built from nothing but 3x3 convolutions stacked in uniform blocks.

Inception / GoogLeNet

2014Convolutional & visionneuralimages

Parallel convolutions of different kernel sizes in one block, concatenated β€” multi-scale features at one depth.

ResNet family

2015Convolutional & visionneuralimagesvideo

Residual blocks at scale β€” ResNet-50/101/152 are the workhorse backbones that follow the skip-connection idea to depth.

ResNeXt / WideResNet

2016Convolutional & visionneuralimages

Grouped convolutions (ResNeXt) or wider, shallower blocks (WideResNet) to trade depth for parallelism and width.

MobileNet / ShuffleNet

2017Convolutional & visionneuralimagesvideo

Depthwise-separable convolutions and channel shuffling cut multiply-adds by an order of magnitude.

EfficientNet

2019Convolutional & visionneuralimages

Compound scaling: grow depth, width and resolution together in the ratio a small search found best.

ConvNeXt

2022Convolutional & visionneuralimages

A ResNet modernized with transformer-era choices β€” larger kernels, LayerNorm, GELU, fewer normalization stages.

Vision transformers6

Vision transformers

2020Vision transformersneuralimagesvideo

Split the image into patches, treat each patch as a token, and run a plain transformer encoder over them.

DeiT

2020Vision transformersneuralimages

ViT made trainable on ImageNet alone through distillation from a CNN teacher and heavy augmentation.

Swin Transformer

2021Vision transformersneuralimagesvideo

Windowed attention with shifted windows, so cost stays linear in image size while information crosses windows.

Masked autoencoders

2021Vision transformersneuralimagesvideo

Mask most patches and train the encoder to reconstruct them β€” a self-supervised pretraining objective for ViTs.

CLIP / contrastive vision-language

2021Vision transformersneuralmultimodalimagestext

Jointly embeds images and captions into one space with a contrastive objective, enabling zero-shot classification.

SigLIP

2023Vision transformersneuralmultimodalimagestext

Replaces CLIP's softmax contrastive loss with a pairwise sigmoid loss, training better at small batch sizes.

Detection & segmentation11

YOLO family

2015Detection & segmentationneuralimagesvideo

Single-shot detection that predicts boxes and classes in one forward pass β€” real-time by design.

SSD / RetinaNet

2016Detection & segmentationneuralimages

Multi-scale anchor-based single-shot heads; RetinaNet adds focal loss to fix the foreground/background imbalance.

DETR and DINO-DETR

2020Detection & segmentationneuralimages

Detection as set prediction: a transformer decoder with learned object queries, matched by Hungarian loss. No anchors, no NMS.

FCN / U-Net

2015Detection & segmentationneuralimagesmolecules

Fully-convolutional encoder-decoder with skip connections, producing a per-pixel output map.

DeepLab / atrous segmentation

2016Detection & segmentationneuralimages

Dilated convolutions and spatial pyramid pooling keep resolution while enlarging the receptive field.

Promptable segmentation (SAM-style)

2023Detection & segmentationneuralimagesvideo

A huge image encoder with a prompt-conditioned mask decoder that segments anything you point at or box.

SAM 2 (Segment Anything in Images and Videos)

2024Detection & segmentationneuralimagesvideo

A promptable unified segmentation model extending visual mask prediction to real-time streaming video through a spatio-temporal memory bank and memory attention.

3D and point-cloud networks

2017Detection & segmentationneuralmoleculesimages

Operate on point sets or voxels: PointNet's symmetric pooling, PointNet++, sparse 3D convolutions, 3D U-Net.

Recurrent & sequence8

Recurrent neural networks

1986Recurrent & sequenceneuraltexttime seriesspeech

A shared hidden state carried through time, so the model sees a sequence as it unfolds.

LSTM

1997Recurrent & sequenceneuraltexttime seriesspeech

Gated cell state with input, forget and output gates that let information persist across many steps.

GRU

2014Recurrent & sequenceneuraltexttime seriesspeech

LSTM simplified to two gates and one merged state β€” fewer parameters, similar accuracy.

Bidirectional RNNs

1997Recurrent & sequenceneuraltextspeech

Two RNNs, one reading forward and one backward, so every position sees full context.

xLSTM

2024Recurrent & sequenceneuraltexttime series

LSTM modernized with exponential gating, a matrix memory and parallelizable training (sLSTM and mLSTM blocks).

Sequence-to-sequence with attention

2014Recurrent & sequenceneuraltextspeech

An encoder RNN compresses the input, a decoder generates the output, and attention lets the decoder look back at every encoder state.

Temporal convolutional networks

2016Recurrent & sequenceneuraltime seriesaudiosignals

Dilated causal 1D convolutions stacked to cover long receptive fields, trained in parallel unlike an RNN.

Attention & transformers4

Attention mechanism

2014Attention & transformerscomponenttextimagesmultimodal

A soft, content-based lookup: query a set of key-value pairs and read out a weighted average.

Transformer

2017Attention & transformersneuraltextimagesspeech

Stacked blocks of multi-head self-attention plus a position-wise feed-forward network, with residual connections and normalization.

Efficient attention variants

2020Attention & transformerscomponenttextany domain

Ways to dodge the quadratic attention matrix: sparse patterns, low-rank projections, kernels, clustering, or memory.

Pretrained language models9

Encoder-only models (BERT family)

2018Pretrained language modelsneuraltextcode

Masked-language pretraining over a bidirectional encoder; produces contextual embeddings, not generated text.

Sentence embedding models

2019Pretrained language modelsneuraltext

Transformers trained so a single pooled vector represents a sentence β€” bi-encoders with contrastive objectives.

Mixture-of-experts language models

2021Pretrained language modelshybridtextcode

Transformer blocks whose feed-forward layers are replaced by routed expert banks, activating a small slice per token.

Instruction-tuned and reasoning models

2022Pretrained language modelshybridtextcode

Same transformer backbone, changed by post-training: instruction tuning, RLHF/DPO, then long chain-of-thought RL.

DeepSeek-V3 and R1 reasoning MoE

2024Pretrained language modelsneuraltextcode

671B parameter Mixture-of-Experts (37B active) pairing Multi-Head Latent Attention (MLA) with fine-grained MoE routing, multi-token prediction, and large-scale reinforcement learning (GRPO).

Vision-language models

2021Pretrained language modelshybridmultimodalimagestext

A vision encoder, a projector, and a language model trained to talk about images (Flamingo, BLIP-2, LLaVA, Qwen-VL, InternVL).

Diffusion transformers

2022Pretrained language modelsgenerativeimagesvideoaudio

The diffusion U-Net replaced by a transformer operating on latent patches, with adaptive normalization for the timestep.

State-space & hybrids12

State-space models

2021State-space & hybridsneuraltime seriestextsignals

Linear time-invariant recurrences with structured state matrices β€” RNN-like linear-time inference, trainable in parallel as a convolution.

S4 / S4D / DSS / S5

2021State-space & hybridsneuralsignalstime series

The first generation of deep SSMs: diagonal or low-rank parameterizations of the state matrix with a parallel scan.

Mamba / selective SSM

2023State-space & hybridsneuraltexttime seriessignals

Makes the SSM parameters input-dependent (selective), so the model can choose what to remember or forget per token.

Mamba-2 / SSD

2024State-space & hybridsneuraltext

Restricts the state transition to a scalar-times-identity, exposing the recurrence as a matrix multiply (the SSD layer).

Vision Mamba / MambaVision

2024State-space & hybridsneuralimages

Bidirectional or hybrid SSM backbones for images and video, offering linear cost instead of quadratic attention.

RWKV

2021State-space & hybridsneuraltext

An RNN whose recurrence is expressed as a linear-attention style kernel, so training parallelizes and inference is constant-memory.

RetNet

2023State-space & hybridsneuraltext

Retention instead of attention β€” a decayed weighted sum that can be computed three ways: parallel, recurrent or chunkwise.

Griffin / Hawk / RecurrentGemma

2024State-space & hybridshybridtext

Real-gated linear recurrences mixed with local attention, packaged as a production open model.

Hyena and long convolutions

2022State-space & hybridsneuraltextmolecules

Implicit long convolutions interleaved with gating, sub-quadratic in sequence length.

StripedHyena

2023State-space & hybridshybridtextmolecules

Alternates SSM and convolution blocks with a small amount of attention, tuned for throughput at long context.

Jamba / Zamba

2024State-space & hybridshybridtext

Interleaves transformer and Mamba layers in one MoE model, so most layers are linear and a few do exact attention.

Hybrid attention and SSM models

2023State-space & hybridshybridtextany domain

The empirical consensus: a mostly-linear stack with a minority of attention layers beats either pure design.

Generative models17

Variational autoencoders

2013Generative modelsgenerativeimagesany domainmolecules

An autoencoder trained by maximizing a variational lower bound, giving a smooth continuous latent space you can sample.

VQ-VAE / VQ-GAN

2017Generative modelsgenerativeimagesaudio

A discrete codebook bottleneck turns continuous data into tokens that a prior model (often a transformer) can generate.

Generative adversarial networks

2014Generative modelsgenerativeimagesvideo

A generator and a discriminator trained against each other, the discriminator supplying the loss signal.

DCGAN / WGAN

2015Generative modelsgenerativeimages

DCGAN made GANs trainable with convolutions; WGAN replaced the Jensen-Shannon objective with Wasserstein distance plus a Lipschitz constraint.

StyleGAN family

2018Generative modelsgenerativeimages

A style-based generator where a learned latent is injected at every resolution level, disentangling coarse from fine attributes.

CycleGAN / Pix2Pix

2017Generative modelsgenerativeimages

Image-to-image translation, with CycleGAN's cycle-consistency letting it learn unpaired mappings.

Autoregressive image and audio models

2016Generative modelsgenerativeimagesaudio

Factor a signal into a factorized sequential distribution and predict it element by element: PixelCNN, PixelRNN, WaveNet.

Normalizing flows

2015Generative modelsgenerativetabularimagessignals

Compose invertible transforms so the exact likelihood is computable and sampling is just an inverse pass.

Denoising diffusion probabilistic models

2020Generative modelsgenerativeimagesaudiovideo

Destroy data with noise over many steps and learn to reverse the process; sampling anneals noise back into a sample.

Score-based and SDE models

2021Generative modelsgenerativeimagesaudio

Generalizes diffusion as a stochastic differential equation over continuous time, trained on score matching.

Latent diffusion models

2022Generative modelsgenerativeimagesvideo

Run the diffusion process inside a VQ-VAE latent space instead of pixel space, cutting cost by orders of magnitude.

Classifier-free guidance

2022Generative modelscomponentimagesaudio

Train one model on conditional and unconditional objectives, then interpolate between them at sampling for stronger conditioning.

Flow matching / rectified flow

2022Generative modelsgenerativeimagesvideoaudio

Instead of reversing noise, learn a velocity field that moves samples along straight-ish paths from noise to data.

FLUX.1 (Flow Matching Transformer)

2024Generative modelsgenerativeimages

12B-parameter rectified flow transformer combining dual multimodal diffusion transformer (MMDiT) blocks with rotary position encodings for extreme visual fidelity and typography.

Energy-based models

2006Generative modelsgenerativeany domain

Learn an unnormalized energy function; low energy means likely data, and sampling means descending it.

GFlowNets

2021Generative modelsgenerativemoleculesgraphs

Learns a stochastic policy whose sampling probability is proportional to a reward, giving diverse rather than single-best outputs.

Graph & geometric8

Graph neural networks

2005Graph & geometricneuralgraphsmolecules

Message passing over edges: each node aggregates its neighbours, repeated to propagate information across hops.

Graph convolutional networks

2016Graph & geometricneuralgraphs

A spectral filter approximated by a first-order polynomial, which reduces to a normalized neighbourhood average.

GraphSAGE

2017Graph & geometricneuralgraphsrecommenders

Samples a fixed number of neighbours and concatenates self with aggregated neighbours β€” inductive, so new nodes work.

Graph attention networks

2017Graph & geometricneuralgraphs

Attention weights over neighbours instead of fixed normalization, so important edges dominate.

Graph isomorphism network / MPNN

2018Graph & geometricneuralmoleculesgraphs

Sum-aggregation message passing with a provably maximal expressive power for GNNs (bounded by 1-WL).

Graph transformers

2021Graph & geometricneuralgraphsmolecules

Attention over all nodes with structural encodings (shortest paths, Laplacians, random walks) injected as bias.

Equivariant and geometric networks

2017Graph & geometricneuralmoleculesgraphs

Architectures that respect 3D symmetry by construction β€” rotations and translations transform the output the same way they transform the input.

Reinforcement learning15

Q-learning and DQN

2013Reinforcement learningreinforcement learningcontrol / agents

Learns the value of each action in each state; DQN approximates the Q-function with a CNN stabilized by replay and a target network.

Double / dueling / prioritized DQN

2015Reinforcement learningreinforcement learningcontrol / agents

Double DQN de-biases the max operator; dueling splits state value from advantage; prioritized replay samples surprising transitions more.

Actor-critic methods

1999Reinforcement learningreinforcement learningcontrol / agents

A policy (actor) proposes actions while a value function (critic) estimates how good they were, reducing gradient variance.

A3C / A2C

2016Reinforcement learningreinforcement learningcontrol / agents

Many parallel workers each collect experience and update a shared policy asynchronously or in lockstep.

PPO

2017Reinforcement learningreinforcement learningcontrol / agentstext

Trust-region policy optimization simplified into a clipped ratio objective that is safe enough to run for many epochs.

DDPG / TD3

2015Reinforcement learningreinforcement learningcontrol / agents

Deterministic policy gradients for continuous actions; TD3 adds twin critics and target smoothing to tame Q-value overestimation.

Soft actor-critic

2018Reinforcement learningreinforcement learningcontrol / agents

Maximum-entropy RL: maximize reward while staying as stochastic as possible, with twin critics and automatic temperature tuning.

Model-based RL (Dreamer, MuZero)

2019Reinforcement learningreinforcement learningcontrol / agents

Learn a model of the environment's dynamics and plan or train inside the learned model.

World models

2018Reinforcement learningreinforcement learningcontrol / agentsvideo

Learn a compressed latent model of the environment that can be rolled forward to imagine trajectories.

AlphaZero

2017Reinforcement learningreinforcement learningcontrol / agents

Self-play plus Monte Carlo tree search guided by a single network that predicts both policy and value.

Multi-agent RL

2017Reinforcement learningreinforcement learningcontrol / agents

Several learning agents share an environment; centralised critics with decentralised actors are the usual answer.

Direct preference optimization

2023Reinforcement learninghybridtext

Skips the reward model and the RL loop by turning the preference objective into a simple classification loss on the policy.

GRPO and verifier-based RL

2024Reinforcement learninghybridtextcode

Drops the critic and normalizes advantages within a group of sampled answers; verifiable rewards (math, code tests) carry the signal.

Bandits (UCB, Thompson sampling)

1933Reinforcement learningclassicalrecommendersany domain

Sequential decision-making with immediate reward only: explore to learn, exploit to earn.

Transformer components11

Multi-head attention

2017Transformer componentscomponenttextimages

Several attention heads run in parallel on projected subspaces, then concatenate β€” different heads learn different relations.

Grouped-query attention / MQA

2019Transformer componentscomponenttext

Share one key/value projection across groups of query heads, shrinking the KV cache by 4–8x at almost no quality cost.

Multi-head latent attention

2024Transformer componentscomponenttext

Compresses keys and values into a low-rank latent vector before caching, cutting the KV cache far beyond GQA.

KV sharing and compressed attention

2026Transformer componentscomponenttext

The 2026 frontier of cache reduction: reuse KV tensors across layers, budget attention per layer, compress via convolution.

Rotary and relative position encodings

2021Transformer componentscomponenttext

Inject position by rotating query/key pairs by an angle proportional to position (RoPE), or by adding learned relative biases (ALiBi, T5).

Normalization layers

2015Transformer componentscomponentany domain

BatchNorm normalizes across the batch, LayerNorm and RMSNorm across features β€” stabilizing optimization and enabling higher learning rates.

Activation functions

2010Transformer componentscomponentany domain

The nonlinearity between linear layers: ReLU, LeakyReLU, ELU, GELU, SiLU/Swish, Mish, and the gated GLU family.

Residual stream and hyper-connections

2015Transformer componentscomponentany domain

The running sum that every block reads from and writes to β€” and new plumbing (hyper-connections, mHC) that widens it.

Feed-forward block and gated FFN

2020Transformer componentscomponentany domain

The per-token MLP that holds most parameters, usually widened 4x and gated in modern models.

Scaling laws

2020Transformer componentscomponentany domain

Empirical power laws linking loss to parameters, data and compute β€” plus the finding that data matters more than previously assumed.

Tokenization

2016Transformer componentscomponenttext

Byte-pair encoding, WordPiece or Unigram turn text into subword tokens β€” the model's actual alphabet.

Domain-specific14

Tabular deep learning

2019Domain-specificneuraltabular

Neural architectures designed for tables rather than repurposed from images: feature-tokenizing transformers and attentive flows.

Time-series models

2019Domain-specifichybridtime series

Forecasting architectures beyond ARIMA: N-BEATS, N-HiTS, TFT, DeepAR, and patching transformers like PatchTST and TimesNet.

Time-series foundation models

2023Domain-specificneuraltime series

Large pretrained forecasters trained across many domains that zero-shot a new series without fitting.

Recommender architectures

2016Domain-specifichybridrecommenders

From matrix factorization and BPR through factorization machines (DeepFM, DLRM) to two-tower retrieval and sequential SASRec/BERT4Rec.

Generative retrieval

2022Domain-specificneuralrecommenderstext

Encode each item as a semantic ID and let a sequence model generate the next item's ID directly.

Speech recognition architectures

2006Domain-specifichybridspeech

CTC and RNN-Transducer for alignment-free training, LAS for attention-based decoding, Conformer for the current encoder.

Whisper (Speech Transformer)

2022Domain-specificneuralspeechaudio

Encoder-decoder Transformer trained on 680,000 hours of weakly-supervised multilingual audio for robust zero-shot speech recognition, translation, and voice activity detection.

Speech synthesis and audio codecs

2016Domain-specificgenerativespeechaudio

Text to waveform through neural vocoders (WaveNet, HiFi-GAN), end-to-end TTS (Tacotron, FastSpeech), and neural audio codecs (SoundStream, EnCodec).

Protein and molecule architectures

2020Domain-specificneuralmolecules

Equivariant graph networks and attention over residues (AlphaFold2's Evoformer, ESM, RoseTTAFold, AlphaFold3's diffusion module).

Physics-informed neural networks

2019Domain-specificneuralcontrol / agentssignals

Add a residual term for the governing differential equation, so the network is penalised for violating known physics.

Neural operators

2020Domain-specificneuralcontrol / agentssignals

Learn mappings between function spaces rather than between finite vectors, so one model generalizes across resolutions.

Weather and climate models

2023Domain-specificneuralimagestime series

Large graph or transformer networks trained on reanalysis data that now beat numerical weather prediction at medium range.

Retrieval-augmented architectures

2020Domain-specifichybridtext

Pair a retriever with a generator so knowledge lives in an index rather than only in weights (RAG, REALM, FiD, kNN-LM).

Geometric deep learning

2017Domain-specificneuralgraphsmoleculesimages

The unifying framework: architectures are defined by which symmetries they respect over which domain.

0 shortlisted