Learn ML from Scratch: Beginner to Advanced

A rigorous, step-by-step interactive curriculum. Master foundational mathematical derivations, vector calculus, pure NumPy/PyTorch implementations from scratch, and modern 2026 reasoning LLM architectures.

🎓 Track Progress: 0 Completed
0%
💻 Modern ArchitecturesLEVEL 1 · BEGINNERChapter 2

Logistic Regression & Cross-Entropy Classification

Sigmoid activation, maximum likelihood estimation, and binary decision boundaries.

⏱ 12 min read🎯 Prerequisites: Linear Regression & Exponentials
🔍 Inspect Architecture: Logistic regression

💡 1. Core Intuition & Concepts

Logistic regression extends linear models to binary classification by passing linear logits through the non-linear sigmoid activation function σ(z) = 1 / (1 + e^-z), mapping real numbers to probability interval (0, 1).

📐 2. Mathematical Formulations & Derivations

Sigmoid Activation Function
σ(z)=11+e−zσ′(z)=σ(z)(1−σ(z))\sigma(z) = \frac{1}{1 + e^{-z}} \qquad \sigma'(z) = \sigma(z)(1 - \sigma(z))
Maps real log-odds to probability interval [0, 1] with symmetric derivative.
Binary Cross-Entropy Loss (BCE)
LBCE(w,b)=−1N∑i=1N[yilog⁡(y^i)+(1−yi)log⁡(1−y^i)]\mathcal{L}_{\text{BCE}}(w,b) = -\frac{1}{N} \sum_{i=1}^N \left[ y_i \log(\hat{y}_i) + (1 - y_i) \log(1 - \hat{y}_i) \right]
Negative log-likelihood of Bernoulli distribution under maximum likelihood estimation.
BCE Loss Gradients
∇wL=1NXT(y^−y)∂L∂b=1N∑i=1N(y^i−yi)\nabla_w \mathcal{L} = \frac{1}{N} X^T (\hat{y} - y) \qquad \frac{\partial \mathcal{L}}{\partial b} = \frac{1}{N} \sum_{i=1}^N (\hat{y}_i - y_i)
Identical functional form to linear regression gradients, but with non-linear predictions.

⚙️ 3. Step-by-Step Computational Mechanism

1
Compute Linear Logits
Calculate z = Xw + b as the raw unconstrained score.
2
Apply Sigmoid Activation
Transform logits into predicted probabilities y_hat = σ(z) in (0, 1).
3
Evaluate Cross-Entropy Loss
Compute negative log-likelihood L = -[y log y_hat + (1-y) log(1-y_hat)].
4
Vectorized Gradient Descent
Update w <- w - alpha * (1/N) X^T(y_hat - y).

💻 4. Code from Scratch (python)

logreg.py
import numpy as np

class LogisticRegressionScratch:
    """Logistic Regression classifier implemented from scratch in NumPy."""
    def __init__(self, lr: float = 0.05, n_iters: int = 1000):
        self.lr = lr
        self.n_iters = n_iters
        self.weights = None
        self.bias = 0.0

    def _sigmoid(self, z: np.ndarray) -> np.ndarray:
        return np.where(z >= 0, 1 / (1 + np.exp(-z)), np.exp(z) / (1 + np.exp(z)))

    def fit(self, X: np.ndarray, y: np.ndarray):
        n_samples, n_features = X.shape
        self.weights = np.zeros(n_features)
        self.bias = 0.0

        for _ in range(self.n_iters):
            linear_model = np.dot(X, self.weights) + self.bias
            y_pred = self._sigmoid(linear_model)

            dw = (1 / n_samples) * np.dot(X.T, (y_pred - y))
            db = (1 / n_samples) * np.sum(y_pred - y)

            self.weights -= self.lr * dw
            self.bias -= self.lr * db

    def predict_proba(self, X: np.ndarray) -> np.ndarray:
        return self._sigmoid(np.dot(X, self.weights) + self.bias)

    def predict(self, X: np.ndarray, threshold: float = 0.5) -> np.ndarray:
        return (self.predict_proba(X) >= threshold).astype(int)

🧠 5. Comprehension Checkpoint

Answer all 1 questions correctly to complete the chapter · 0 / 1 done
Q1/1 What is the derivative of the Sigmoid function σ(z) expressed in terms of σ(z)?

Finished this chapter?