1. Home
  2. AI & Machine Learning
  3. Logistic Regression

Logistic Regression

Predict yes/no probabilities with an S-shaped curve. Watch a 3D probability sheet bend to separate two classes as it trains.

Interactive 3DIntermediate13 min readAI/MLUpdated

Drag to rotate · Right-drag to pan · Click, then scroll to zoom · Space play · ←→ step

What's happening

Pseudocode

    Try this in the 3D model

    • Rotate the scene to look along the boundary. Can you see the S-shaped sigmoid?
    • Train on Overlapping data. Why does accuracy stay below 100%?
    • Try learning rate 0.1 and then 1.5. Which one finishes faster?
    • Watch the glowing (misclassified) points disappear as training goes on.

    From numbers to yes/no

    Linear regression predicts a number. But many questions have yes/no answers: Will this student pass? Is this email spam? Is this tumour malignant? We want a probability between 0 and 1.

    Logistic regression does this in two steps:

    1. Compute a score exactly like linear regression: z = w₁x₁ + w₂x₂ + b
    2. Squash it into (0, 1) with the sigmoid function:
    σ(z) = 1 / (1 + e^(−z))

    The sigmoid is an S-shaped curve: very negative z → almost 0, very positive z → almost 1, and z = 0 → exactly 0.5.

    Seeing it in 3D

    In the model, the floor holds two features (x₁, x₂) and the height of the coloured sheet is the predicted probability of class 1. Class-1 points float at the top (probability 1), class-0 points sit at the bottom. As training runs, the flat sheet tilts and bends into an S-shaped “cliff” that rises right between the two groups.

    The decision boundary

    We predict class 1 when p ≥ 0.5, which happens exactly when z ≥ 0. So the boundary is the set of points where

    w₁x₁ + w₂x₂ + b = 0

    — a straight line (the white line in the model). Logistic regression is therefore a linear classifier.

    Training: cross-entropy and gradient descent

    For a point with label y (0 or 1) and predicted probability p, the cross-entropy loss is:

    loss = −[ y · log(p) + (1 − y) · log(1 − p) ]

    It’s near 0 when the prediction is confident and right, and huge when it’s confident and wrong. Its gradient turns out to be beautifully simple:

    ∂loss/∂w = (p − y) · x        ∂loss/∂b = (p − y)

    So each step nudges the weights in proportion to the error p − y, just like gradient descent for linear regression.

    Code

    import numpy as np
    
    def sigmoid(z):
        return 1 / (1 + np.exp(-z))
    
    # features: [hours studied, attendance %] (scaled); label: passed (1) or not (0)
    X = np.array([[0.5, 0.4], [1.0, 0.5], [1.5, 0.3], [3.0, 0.9],
                  [3.5, 0.7], [4.0, 0.95], [2.0, 0.6], [4.5, 0.8]])
    y = np.array([0, 0, 0, 1, 1, 1, 0, 1])
    
    w, b, lr = np.zeros(2), 0.0, 0.5
    for step in range(2000):
        p = sigmoid(X @ w + b)
        w -= lr * X.T @ (p - y) / len(y)
        b -= lr * np.mean(p - y)
    
    new_student = np.array([2.8, 0.85])
    print("P(pass) =", sigmoid(new_student @ w + b))

    With scikit-learn:

    from sklearn.linear_model import LogisticRegression
    model = LogisticRegression().fit(X, y)
    print(model.predict_proba([[2.8, 0.85]])[0, 1])

    Beyond two classes and straight lines

    • Multi-class: use one-vs-rest, or the softmax generalisation (the last layer of most neural networks).
    • Curved boundaries: add features like x₁², x₁x₂ — or use a neural network.
    • Regularisation (L1/L2) keeps weights small to avoid overfitting.

    Where is it used?

    Spam filters, credit scoring, medical diagnosis, click-through prediction, and as the final layer of neural network classifiers. It’s fast, simple and its weights are easy to interpret.

    Common mistakes

    • Calling it “regression” and expecting it to predict numbers — it’s a classification method.
    • Using mean squared error as the loss — it trains poorly with a sigmoid.
    • Forgetting that the decision threshold doesn’t have to be 0.5 (e.g. medical tests may use a lower one).

    Complexity at a glance

    Case / operationTimeWhy
    PredictionO(d)One weighted sum and a sigmoid.
    One training stepO(n · d)
    Extra spaceO(d)

    Quick check

    Test yourself — pick an answer to see if you got it.

    1. What is the output range of the sigmoid function σ(z)?

    2. σ(0) equals…

    3. What shape is the decision boundary of logistic regression (with the raw features)?

    4. Which loss function is used to train logistic regression?

    Saved only in this browser — no account needed.
    Spotted a mistake or a bug in the 3D model?

    Report a mistake

    in Logistic Regression. Thank you — every report makes the lesson better for the next reader.

    We'll also include a link to the step of the 3D model you're on and your browser type, so we can reproduce it.