Statistics Behind Machine Learning: Understanding Logistic Regression

Blog post by Anand: "Statistics Behind Machine Learning: Understanding Logistic Regression" (published April 3, 2025; categories: Tech). From binary classification with sigmoid to multi-class with softmax — a step-by-step walkthrough of logistic regression, cross-entropy loss, and gradient descent with Python code.

In my previous blogs, I covered Statistics Behind Machine Learning: Understanding Simple and Multiple Linear Regression and Statistics Behind Machine Learning: Understanding Polynomial Regression. Those blogs focused on predicting continuous numerical values — like house prices or exam scores.

But many real-world problems are not about predicting numbers. They are about predicting categories. Will this customer buy or not? Is this email spam or not? Will this student pass or fail? That is classification, and that is what we cover today.

What Is Classification?

Classification is fundamentally different from regression. Regression outputs a continuous value. Classification outputs a discrete category.

You might think — why not just use linear regression and set a threshold? The problem is linear regression can produce values like 1.7, −0.4, or 93.2. Those do not make sense as probabilities.

You need something that stays between 0 and 1. That is where logistic regression comes in.

What Is Logistic Regression?

Despite its name, logistic regression is a classification algorithm — not a regression one. It predicts the probability that a data point belongs to a certain class.

Probability >= 0.5 → Predict Class 1. Probability < 0.5 → Predict Class 0

Part 1: Binary Classification Using the Sigmoid Function

Problem: Predict whether a student will pass (1) or fail (0) based on hours studied.

Step 1: The Linear Equation

x = input feature (hours studied). w = weight. b = bias. z = raw output

Example: w=0.8, b=−1.2, x=3 → z = 0.8×3 + (−1.2) = 1.2

Step 2: The Sigmoid Function

Example: σ(1.2) = 1 / (1 + e^−1.2) ≈ 0.769 → "76.9% chance of passing"

Step 3: Converting Probability to Class Label

0.5 is the default threshold — but real applications tune it. Cancer screening may use 0.2 (catch more cases). Spam filtering may use 0.7 (avoid false positives).

Step 4: The Loss Function — Binary Cross-Entropy (Log Loss)

y = actual label (0 or 1). ŷ = predicted probability

Step 5: Training via Gradient Descent

Initialize: w = 0, b = 0. Forward pass: compute z → sigmoid → probability. Calculate loss. Compute gradients: ∂L/∂w = (ŷ − y) · x and ∂L/∂b = (ŷ − y). Update weights: w = w − α · (ŷ − y) · x and b = b − α · (ŷ − y) (α = learning rate, typically 0.01). Repeat thousands of iterations until convergence

Step 6: Prediction

Binary Logistic Regression — Python Code

Output:

Part 2: Multi-Class Classification Using Softmax

What if there are more than two classes? For example: Fail / Average / Excellent. Sigmoid is not enough — we need probabilities that sum to 1 across all classes.

Step 1: One Linear Equation Per Class

Step 2: The Softmax Function

Numerical Example:

Step 3: Loss Function — Categorical Cross-Entropy

y is a one-hot vector. Example for true label "Average":

Step 4: Training

Same gradient descent loop as binary classification, but updates multiple weight sets — one per class.

Step 5: Prediction

No threshold — just pick the class with the highest probability.

Softmax Classification — Python Code

Output:

Key Takeaways

Threshold selection is a business decision, not a mathematical law.. Probability outputs are more valuable than hard class labels — confidence scores enable smarter downstream decisions.. Internally, the model treats labels as numbers ("Pass" = 1, "Fail" = 0).. Logistic regression is still industry-standard in banking, medicine, and tech for its speed and interpretability.. Softmax underpins modern AI — LLM output layers use softmax, making this knowledge foundational.

That's all for today. As usual, when you feel this content is valuable, follow me for more upcoming blogs.

Connect with Me:

Email: sanand03072005@gmail.com. Instagram: @anandsundaramoorthysa. LinkedIn: Anand Sundaramoorthy

Read it: https://www.anandsundaramoorthy.com/blog/statistics-behind-machine-learning-understanding-logistic-regression


Static rendering for crawlers. The full interactive site is at https://www.anandsundaramoorthy.com/blog/statistics-behind-machine-learning-understanding-logistic-regression.