From numbers to yes/no
Linear regression predicts a number. But many questions have yes/no answers: Will this student pass? Is this email spam? Is this tumour malignant? We want a probability between 0 and 1.
Logistic regression does this in two steps:
- Compute a score exactly like linear regression:
z = w₁x₁ + w₂x₂ + b - Squash it into (0, 1) with the sigmoid function:
σ(z) = 1 / (1 + e^(−z))
The sigmoid is an S-shaped curve: very negative z → almost 0, very positive z → almost 1, and z = 0 → exactly 0.5.
Seeing it in 3D
In the model, the floor holds two features (x₁, x₂) and the height of the coloured sheet is the predicted probability of class 1. Class-1 points float at the top (probability 1), class-0 points sit at the bottom. As training runs, the flat sheet tilts and bends into an S-shaped “cliff” that rises right between the two groups.
The decision boundary
We predict class 1 when p ≥ 0.5, which happens exactly when z ≥ 0. So the boundary is the set of points where
w₁x₁ + w₂x₂ + b = 0
— a straight line (the white line in the model). Logistic regression is therefore a linear classifier.
Training: cross-entropy and gradient descent
For a point with label y (0 or 1) and predicted probability p, the cross-entropy loss is:
loss = −[ y · log(p) + (1 − y) · log(1 − p) ]
It’s near 0 when the prediction is confident and right, and huge when it’s confident and wrong. Its gradient turns out to be beautifully simple:
∂loss/∂w = (p − y) · x ∂loss/∂b = (p − y)
So each step nudges the weights in proportion to the error p − y, just like gradient descent for linear regression.
Code
import numpy as np
def sigmoid(z):
return 1 / (1 + np.exp(-z))
# features: [hours studied, attendance %] (scaled); label: passed (1) or not (0)
X = np.array([[0.5, 0.4], [1.0, 0.5], [1.5, 0.3], [3.0, 0.9],
[3.5, 0.7], [4.0, 0.95], [2.0, 0.6], [4.5, 0.8]])
y = np.array([0, 0, 0, 1, 1, 1, 0, 1])
w, b, lr = np.zeros(2), 0.0, 0.5
for step in range(2000):
p = sigmoid(X @ w + b)
w -= lr * X.T @ (p - y) / len(y)
b -= lr * np.mean(p - y)
new_student = np.array([2.8, 0.85])
print("P(pass) =", sigmoid(new_student @ w + b))
With scikit-learn:
from sklearn.linear_model import LogisticRegression
model = LogisticRegression().fit(X, y)
print(model.predict_proba([[2.8, 0.85]])[0, 1])
Beyond two classes and straight lines
- Multi-class: use one-vs-rest, or the softmax generalisation (the last layer of most neural networks).
- Curved boundaries: add features like x₁², x₁x₂ — or use a neural network.
- Regularisation (L1/L2) keeps weights small to avoid overfitting.
Where is it used?
Spam filters, credit scoring, medical diagnosis, click-through prediction, and as the final layer of neural network classifiers. It’s fast, simple and its weights are easy to interpret.
Common mistakes
- Calling it “regression” and expecting it to predict numbers — it’s a classification method.
- Using mean squared error as the loss — it trains poorly with a sigmoid.
- Forgetting that the decision threshold doesn’t have to be 0.5 (e.g. medical tests may use a lower one).
Complexity at a glance
| Case / operation | Time | Why |
|---|---|---|
| Prediction | O(d) | One weighted sum and a sigmoid. |
| One training step | O(n · d) | |
| Extra space | O(d) |
Quick check
Test yourself — pick an answer to see if you got it.
1. What is the output range of the sigmoid function σ(z)?
That's why its output can be read as a probability.
2. σ(0) equals…
1 / (1 + e⁰) = 1/2. The decision boundary is exactly where z = 0.
3. What shape is the decision boundary of logistic regression (with the raw features)?
The boundary is w·x + b = 0, which is linear.
4. Which loss function is used to train logistic regression?
Cross-entropy heavily penalises confident wrong predictions and gives a smooth, convex loss.