Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
🧠 Core CourseLEVEL 1 · BEGINNERChapter 2

One Neuron, Many Inputs: The Dot Product

Grow the one-input model to many inputs, meet the dot product, then stack neurons into layers and batches, and see why a layer needs a bend.

⏱ 15 min read🎯 Prerequisites: Chapter 1 (a model is a function with knobs)
🔍 Inspect Architecture: Perceptron

💡 1. Core Intuition & Concepts

Chapter 1's model had one input. Real decisions rarely do. Suppose you want to score how good tomorrow is for a picnic from three forecast numbers, each scaled to 0–1: sunshine, chance of rain, and wind. The natural extension of w·x + b is to give every input its own knob and add everything up: score = w₁·sun + w₂·rain + w₃·wind + b.

That weighted sum plus a bias is a neuron. Nothing more mysterious is going on. The weights encode how much each input matters and in which direction. A positive weight on sunshine means more sun raises the score, and a negative weight on rain means rain lowers it. The bias sets the baseline score when every input is zero.

Multiplying two lists element by element and summing the results is common enough to have its own name: the dot product. Once you see a neuron as a dot product, the rest follows. A layer is several neurons that read the same inputs, so it's several dot products, which is a matrix times a vector. Scoring many days at once is a matrix times a matrix. That's the entire reason GPUs, which are very fast at matrix multiplication, took over deep learning.

There's one catch, which the code below demonstrates: stacking layers of pure weighted sums gains you nothing, because two linear layers collapse into one. A small bend after each neuron (we use ReLU, max(0, z)) breaks the collapse. That bend is what lets deep networks represent curved, complicated functions.

📐 2. Mathematical Formulations & Derivations

One neuron: dot product plus bias
z = w · x + b = w₁x₁ + w₂x₂ + … + wₙxₙ + b = Σⱼ wⱼxⱼ + b
x holds the n inputs, w holds one weight per input, and b is the bias. With n = 1 this is exactly Chapter 1's model.
Worked example: tomorrow's picnic score
x = [0.7, 0.2, 0.5] w = [2.0, −3.0, −1.0] b = 0.5 z = 2.0·0.7 + (−3.0)·0.2 + (−1.0)·0.5 + 0.5 = 1.4 − 0.6 − 0.5 + 0.5 = 0.8
Mostly sunny, low rain chance, moderate wind: a positive score. Check each product by hand once.
The geometry of a dot product
w · x = ‖w‖ · ‖x‖ · cos θ ‖w‖ = √(w · w) = √(w₁² + … + wₙ²)
θ is the angle between the two vectors. The score is largest when x points the same way as w (cos θ = 1), zero when they're at right angles, and negative when they point apart. So a neuron's weight vector is the pattern it's looking for, and the dot product measures how closely an input matches it.
The decision boundary is a flat surface
boundary = { x : w · x + b = 0 } distance from origin = |b| / ‖w‖ signed distance of any x = (w · x + b) / ‖w‖
Every input on one side gets z > 0, and on the other side z < 0. In 2-D the boundary is a line, in 3-D a plane, in n-D a 'hyperplane'. The weights set its tilt and the bias shifts it away from the origin.
A layer: several neurons reading the same inputs
z = W x + b W: (n_out × n_in), x: (n_in), b and z: (n_out) parameters = n_in · n_out + n_out
Row k of W holds the weights of neuron k, so the layer output is one dot product per row. 3 inputs → 3 neurons means 9 weights plus 3 biases, 12 knobs in all.
The shape rule for matrix products
(m × n) · (n × p) → (m × p) the inner sizes (n) must match
Entry (i, k) of the result is the dot product of row i of the left matrix with column k of the right one. Most bugs in neural-network code are shape bugs, and this rule is how you check them on paper.
A batch: many examples at once
Z = X Wᵀ + b X: (examples × inputs), Wᵀ: (inputs × neurons), Z: (examples × neurons)
Each row of X is one day's forecast and each row of Z is that day's layer output. The bias row is added to every row, which is called broadcasting.
Why we need a bend (non-linearity)
W₂(W₁x + b₁) + b₂ = (W₂W₁)x + (W₂b₁ + b₂) → still one linear layer relu(z) = max(0, z)
Two linear layers are equivalent to one linear layer with different knobs. Putting ReLU between them stops that simplification.

• Take any two points x₁ and x₂ that both lie on the boundary, so w·x₁ + b = 0 and w·x₂ + b = 0.

• Subtract the two equations. The bias cancels: w·x₁ − w·x₂ = 0, which is w·(x₁ − x₂) = 0 because the dot product distributes over subtraction.

• x₁ − x₂ is an arrow lying inside the boundary, pointing from one boundary point to another. Its dot product with w is zero, so cos θ = 0 and the angle between them is 90°.

• This holds for every pair of boundary points, so w is perpendicular to the whole boundary. It points towards the side where z > 0.

• Consequence: moving an input along w changes the score fastest, and moving it along the boundary doesn't change the score at all. Chapter 4 generalises this idea as the gradient. ∎

⚙️ 3. Step-by-Step Computational Mechanism

1
Multiply matching entries
Pair input 1 with weight 1, input 2 with weight 2, and so on. The two lists must be the same length. When they aren't, it's the most common bug in neural-network code, and NumPy reports it as a shape error.
2
Sum, then add the bias
Adding up the products gives one number per neuron. The bias shifts it, so the neuron can still output something useful when all its inputs are zero.
3
Repeat per neuron: that's a layer
Give each neuron its own row of weights and its own bias. The inputs stay the same and only the knobs differ, so each neuron learns to look for a different pattern.
4
Repeat per example: that's a batch
Stack several inputs as rows of a matrix and one matrix multiplication scores them all. The maths is the same; it just runs in parallel.
5
Bend the output
Pass each neuron's z through ReLU. Negative values become 0 and positive values pass through unchanged. Without this step, any number of layers is no more expressive than one.

💻 4. Code from Scratch (python)

core_ch2.py
# Chapter 2: from one neuron to a layer to a batch, first in pure Python, then checked with NumPy.
import numpy as np


def dot(a, b):
    """Multiply matching entries and add them up."""
    assert len(a) == len(b), f"length mismatch: {len(a)} vs {len(b)}"
    return sum(ai * bi for ai, bi in zip(a, b))


def neuron(x, w, b):
    return dot(w, x) + b


def layer(x, W, b):
    """One neuron per row of W: same inputs, different knobs."""
    return [neuron(x, w_row, b_k) for w_row, b_k in zip(W, b)]


def relu(values):
    return [max(0.0, v) for v in values]


# --- one neuron: tomorrow's picnic score -------------------------------
tomorrow = [0.7, 0.2, 0.5]            # sunshine, rain chance, wind (0..1)
picnic_w = [2.0, -3.0, -1.0]
picnic_b = 0.5
print("picnic score:", round(neuron(tomorrow, picnic_w, picnic_b), 4))   # 0.8

# --- a layer: three neurons, three different opinions -------------------
W = [[ 2.0, -3.0, -1.0],   # picnic
     [-1.0,  0.0,  3.0],   # kite flying: loves wind
     [ 0.0,  2.5,  0.0]]   # stay-in-and-read: loves rain
b = [0.5, -0.5, 0.0]
print("layer output:", [round(z, 4) for z in layer(tomorrow, W, b)])

# --- a batch: four days at once ------------------------------------------
week = [[0.7, 0.2, 0.5],
        [0.1, 0.9, 0.3],
        [0.9, 0.0, 0.1],
        [0.4, 0.3, 0.9]]
batch_py = [layer(day, W, b) for day in week]

# The same batch with NumPy: one matrix multiplication, bias broadcast over rows
batch_np = np.array(week) @ np.array(W).T + np.array(b)
print("pure Python == NumPy:", np.allclose(batch_py, batch_np))
print("after ReLU:")
for day, z in zip(["Mon", "Tue", "Wed", "Thu"], batch_py):
    print(f"  {day}: {[round(v, 2) for v in relu(z)]}")

# --- why the bend matters: two linear layers collapse into one -----------
rng = np.random.default_rng(0)
W1, b1 = rng.normal(size=(4, 3)), rng.normal(size=4)
W2, b2 = rng.normal(size=(2, 4)), rng.normal(size=2)
x = np.array(tomorrow)

two_layers = W2 @ (W1 @ x + b1) + b2
one_layer  = (W2 @ W1) @ x + (W2 @ b1 + b2)       # a single layer with merged knobs
print("\nlinear stack == single layer:", np.allclose(two_layers, one_layer))

with_relu = W2 @ np.maximum(0, W1 @ x + b1) + b2
print("with ReLU in between, still equal?", np.allclose(with_relu, one_layer))

# --- geometry: the neuron as a pattern detector --------------------------
w = np.array(picnic_w)
for name, day in zip(["Mon", "Tue", "Wed", "Thu"], week):
    d = np.array(day)
    cos = w @ d / (np.linalg.norm(w) * np.linalg.norm(d))
    signed_dist = (w @ d + picnic_b) / np.linalg.norm(w)
    side = "picnic side" if signed_dist > 0 else "stay-home side"
    print(f"{name}: cos(angle to w) = {cos:+.2f}, distance to boundary = {signed_dist:+.2f} ({side})")

✏️ 5. Practice Problems

Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.

P1 Compute the dot product [1, 2, 3] · [4, −5, 6].
💡 Multiply matching entries, then add: 1·4 + 2·(−5) + 3·6.
Solution. 4 − 10 + 18 = 12.
P2 A neuron has weights w = [4, −2] and bias b = 1. What is its output z for the input x = [0.5, 1.0]?
💡 z = w·x + b.
Solution. z = 4·0.5 + (−2)·1.0 + 1 = 2 − 2 + 1 = 1.
P3 A layer reads 64 inputs and has 10 neurons. How many parameters (weights + biases) does it have?
💡 Use parameters = n_in · n_out + n_out.
Solution. 64 × 10 = 640 weights, plus 10 biases = 650.
P4 What is relu(−2.3) + relu(1.7)?
💡 relu(z) = max(0, z).
Solution. relu(−2.3) = 0 and relu(1.7) = 1.7, so the sum is 1.7.
P5 A neuron has w = [3, 4] and b = −10. How far is its decision boundary from the origin?
💡 Distance = |b| / ‖w‖, and ‖w‖ = √(3² + 4²).
Solution. ‖w‖ = √(9 + 16) = 5, so the distance is |−10| / 5 = 2.
P6 What is cos θ between w = [1, 0] and x = [1, 1]? (3 decimal places.)
💡 cos θ = (w·x) / (‖w‖ ‖x‖).
Solution. w·x = 1, ‖w‖ = 1, ‖x‖ = √2, so cos θ = 1/√2 ≈ 0.707, which is an angle of 45°.
P7 You pass a batch of 32 days (3 features each) through a layer of 5 neurons. How many numbers are in the output Z?
💡 Work out the shape of Z = X Wᵀ + b first.
Solution. X is 32 × 3 and Wᵀ is 3 × 5, so Z is 32 × 5 = 160 numbers: one score per day per neuron.

🛠️ 6. Try It Yourself

  1. Invent a fourth neuron for the layer (a 'go swimming' score, say). Add its row to W and its bias to b. What shape is W now, and how many knobs does the layer have?
  2. Swap the order of the inputs in `tomorrow` without changing W. What happens to the scores? What does that tell you about how a neuron 'knows' which input is which?
  3. Swap ReLU for the identity function (return the values unchanged) and rerun the collapse test. Then try the absolute value. Which ones break the collapse, and why?

🧠 7. Comprehension Checkpoint

Answer all 5 questions correctly to complete the chapter · 0 / 5 done
Q1/5 A layer has 4 neurons and each one reads the same 3 inputs. How many knobs (parameters) does the layer have?
Q2/5 For some input x, w·x is large and positive. What does that tell you about x?
Q3/5 Why do we put a non-linearity such as ReLU after each layer?
Q4/5 X has shape (32 × 3) and the layer's W has shape (5 × 3). What is the shape of X Wᵀ?
Q5/5 How is a neuron's weight vector w oriented relative to its decision boundary w·x + b = 0?

Finished this chapter?