Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

πŸŽ“ Track Progress: 0 Completed
0%
πŸ’» Modern ArchitecturesLEVEL 2 Β· INTERMEDIATEChapter 5

Convolutional Networks & 2D Spatial Feature Maps

2D spatial convolutions, sliding windows, padding/stride geometry, and spatial pooling.

⏱ 16 min read🎯 Prerequisites: MLP & 2D Array Slicing

πŸ’‘ 1. Core Intuition & Concepts

Convolutional Neural Networks (CNNs) process grid-structured data like images by enforcing translation equivariance and local spatial connectivity through parameter-shared convolution kernels.

πŸ“ 2. Mathematical Formulations & Derivations

2D Discrete Cross-Correlation (Convolution)
S(i,j)=(Iβˆ—K)(i,j)=βˆ‘m=βˆ’kkβˆ‘n=βˆ’kkI(i+m,j+n)K(m,n)S(i,j) = (I * K)(i,j) = \sum_{m=-k}^{k} \sum_{n=-k}^{k} I(i+m, j+n) K(m, n)
Sliding kernel K computes dot products over local receptive field patches in input image I.
Output Feature Map Spatial Dimension
Hout=⌊Hin+2Pβˆ’KSβŒ‹+1H_{\text{out}} = \left\lfloor \frac{H_{\text{in}} + 2P - K}{S} \right\rfloor + 1
H: height, P: padding pixels, K: kernel size, S: stride step.

βš™οΈ 3. Step-by-Step Computational Mechanism

1
Pad Input
Apply zero padding P to image boundaries to preserve spatial dimensions.
2
Sliding Kernel Dot Products
Slide K x K filter across H x W grid, multiplying local receptive fields and summing with bias.
3
Max-Pooling Downsampling
Extract max activation in non-overlapping 2 x 2 blocks to reduce spatial dimensions by half.

πŸ’» 4. Code from Scratch (python)

cnn.py
import numpy as np

def conv2d_forward(X: np.ndarray, W: np.ndarray, b: np.ndarray, stride: int = 1, pad: int = 0):
    N, C_in, H, W = X.shape
    C_out, _, K_h, K_w = W.shape
    
    if pad > 0:
        X_pad = np.pad(X, ((0,0), (0,0), (pad, pad), (pad, pad)), mode='constant')
    else:
        X_pad = X

    H_out = int((H + 2 * pad - K_h) / stride) + 1
    W_out = int((W + 2 * pad - K_w) / stride) + 1
    out = np.zeros((N, C_out, H_out, W_out))

    for n in range(N):
        for c_o in range(C_out):
            for h_o in range(H_out):
                h_start = h_o * stride
                h_end = h_start + K_h
                for w_o in range(W_out):
                    w_start = w_o * stride
                    w_end = w_start + K_w
                    receptive_field = X_pad[n, :, h_start:h_end, w_start:w_end]
                    out[n, c_o, h_o, w_o] = np.sum(receptive_field * W[c_o]) + b[c_o]

    return out

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter Β· 0 / 1 done
Q1/1 If an input image is 32x32, kernel is 5x5, padding is 0, and stride is 1, what is the output spatial resolution?

Finished this chapter?