One Neuron, Many Inputs: The Dot Product
Grow the one-input model to many inputs, meet the dot product, then stack neurons into layers and batches, and see why a layer needs a bend.
🔍 Inspect Architecture: Perceptron💡 1. Core Intuition & Concepts
Chapter 1's model had one input. Real decisions rarely do. Suppose you want to score how good tomorrow is for a picnic from three forecast numbers, each scaled to 0–1: sunshine, chance of rain, and wind. The natural extension of w·x + b is to give every input its own knob and add everything up: score = w₁·sun + w₂·rain + w₃·wind + b.
That weighted sum plus a bias is a neuron. Nothing more mysterious is going on. The weights encode how much each input matters and in which direction. A positive weight on sunshine means more sun raises the score, and a negative weight on rain means rain lowers it. The bias sets the baseline score when every input is zero.
Multiplying two lists element by element and summing the results is common enough to have its own name: the dot product. Once you see a neuron as a dot product, the rest follows. A layer is several neurons that read the same inputs, so it's several dot products, which is a matrix times a vector. Scoring many days at once is a matrix times a matrix. That's the entire reason GPUs, which are very fast at matrix multiplication, took over deep learning.
There's one catch, which the code below demonstrates: stacking layers of pure weighted sums gains you nothing, because two linear layers collapse into one. A small bend after each neuron (we use ReLU, max(0, z)) breaks the collapse. That bend is what lets deep networks represent curved, complicated functions.
📐 2. Mathematical Formulations & Derivations
z = w · x + b = w₁x₁ + w₂x₂ + … + wₙxₙ + b = Σⱼ wⱼxⱼ + bx = [0.7, 0.2, 0.5] w = [2.0, −3.0, −1.0] b = 0.5
z = 2.0·0.7 + (−3.0)·0.2 + (−1.0)·0.5 + 0.5 = 1.4 − 0.6 − 0.5 + 0.5 = 0.8w · x = ‖w‖ · ‖x‖ · cos θ ‖w‖ = √(w · w) = √(w₁² + … + wₙ²)boundary = { x : w · x + b = 0 } distance from origin = |b| / ‖w‖
signed distance of any x = (w · x + b) / ‖w‖z = W x + b W: (n_out × n_in), x: (n_in), b and z: (n_out)
parameters = n_in · n_out + n_out(m × n) · (n × p) → (m × p) the inner sizes (n) must matchZ = X Wᵀ + b X: (examples × inputs), Wᵀ: (inputs × neurons), Z: (examples × neurons)W₂(W₁x + b₁) + b₂ = (W₂W₁)x + (W₂b₁ + b₂) → still one linear layer
relu(z) = max(0, z)• Take any two points x₁ and x₂ that both lie on the boundary, so w·x₁ + b = 0 and w·x₂ + b = 0.
• Subtract the two equations. The bias cancels: w·x₁ − w·x₂ = 0, which is w·(x₁ − x₂) = 0 because the dot product distributes over subtraction.
• x₁ − x₂ is an arrow lying inside the boundary, pointing from one boundary point to another. Its dot product with w is zero, so cos θ = 0 and the angle between them is 90°.
• This holds for every pair of boundary points, so w is perpendicular to the whole boundary. It points towards the side where z > 0.
• Consequence: moving an input along w changes the score fastest, and moving it along the boundary doesn't change the score at all. Chapter 4 generalises this idea as the gradient. ∎
⚙️ 3. Step-by-Step Computational Mechanism
💻 4. Code from Scratch (python)
# Chapter 2: from one neuron to a layer to a batch, first in pure Python, then checked with NumPy.
import numpy as np
def dot(a, b):
"""Multiply matching entries and add them up."""
assert len(a) == len(b), f"length mismatch: {len(a)} vs {len(b)}"
return sum(ai * bi for ai, bi in zip(a, b))
def neuron(x, w, b):
return dot(w, x) + b
def layer(x, W, b):
"""One neuron per row of W: same inputs, different knobs."""
return [neuron(x, w_row, b_k) for w_row, b_k in zip(W, b)]
def relu(values):
return [max(0.0, v) for v in values]
# --- one neuron: tomorrow's picnic score -------------------------------
tomorrow = [0.7, 0.2, 0.5] # sunshine, rain chance, wind (0..1)
picnic_w = [2.0, -3.0, -1.0]
picnic_b = 0.5
print("picnic score:", round(neuron(tomorrow, picnic_w, picnic_b), 4)) # 0.8
# --- a layer: three neurons, three different opinions -------------------
W = [[ 2.0, -3.0, -1.0], # picnic
[-1.0, 0.0, 3.0], # kite flying: loves wind
[ 0.0, 2.5, 0.0]] # stay-in-and-read: loves rain
b = [0.5, -0.5, 0.0]
print("layer output:", [round(z, 4) for z in layer(tomorrow, W, b)])
# --- a batch: four days at once ------------------------------------------
week = [[0.7, 0.2, 0.5],
[0.1, 0.9, 0.3],
[0.9, 0.0, 0.1],
[0.4, 0.3, 0.9]]
batch_py = [layer(day, W, b) for day in week]
# The same batch with NumPy: one matrix multiplication, bias broadcast over rows
batch_np = np.array(week) @ np.array(W).T + np.array(b)
print("pure Python == NumPy:", np.allclose(batch_py, batch_np))
print("after ReLU:")
for day, z in zip(["Mon", "Tue", "Wed", "Thu"], batch_py):
print(f" {day}: {[round(v, 2) for v in relu(z)]}")
# --- why the bend matters: two linear layers collapse into one -----------
rng = np.random.default_rng(0)
W1, b1 = rng.normal(size=(4, 3)), rng.normal(size=4)
W2, b2 = rng.normal(size=(2, 4)), rng.normal(size=2)
x = np.array(tomorrow)
two_layers = W2 @ (W1 @ x + b1) + b2
one_layer = (W2 @ W1) @ x + (W2 @ b1 + b2) # a single layer with merged knobs
print("\nlinear stack == single layer:", np.allclose(two_layers, one_layer))
with_relu = W2 @ np.maximum(0, W1 @ x + b1) + b2
print("with ReLU in between, still equal?", np.allclose(with_relu, one_layer))
# --- geometry: the neuron as a pattern detector --------------------------
w = np.array(picnic_w)
for name, day in zip(["Mon", "Tue", "Wed", "Thu"], week):
d = np.array(day)
cos = w @ d / (np.linalg.norm(w) * np.linalg.norm(d))
signed_dist = (w @ d + picnic_b) / np.linalg.norm(w)
side = "picnic side" if signed_dist > 0 else "stay-home side"
print(f"{name}: cos(angle to w) = {cos:+.2f}, distance to boundary = {signed_dist:+.2f} ({side})")
✏️ 5. Practice Problems
Work these out on paper (or in Python) and type the number. Answers are checked with a small tolerance for rounding.
🛠️ 6. Try It Yourself
- Invent a fourth neuron for the layer (a 'go swimming' score, say). Add its row to W and its bias to b. What shape is W now, and how many knobs does the layer have?
- Swap the order of the inputs in `tomorrow` without changing W. What happens to the scores? What does that tell you about how a neuron 'knows' which input is which?
- Swap ReLU for the identity function (return the values unchanged) and rerun the collapse test. Then try the absolute value. Which ones break the collapse, and why?
🧠 7. Comprehension Checkpoint
In the catalog
Finished this chapter?