1. Home
  2. AI & Machine Learning
  3. Linear Regression

Linear Regression

The "hello world" of machine learning — fit the best straight line through data. See every error as a 3D square that shrinks as the model learns.

Interactive 3DBeginner12 min readAI/MLUpdated

Drag to rotate · Right-drag to pan · Click, then scroll to zoom · Space play · ←→ step

What's happening

Pseudocode

    Try this in the 3D model

    • Rotate the camera sideways to see the red squared-error squares. Press Train and watch them shrink.
    • Compare Train with gradient descent with Exact answer. Do they end at the same line?
    • Set the learning rate to 0.01. How many steps does it need now?
    • Press New data a few times. Can you guess w and b before training?

    What is it?

    Linear regression predicts a number from one or more inputs by drawing the best straight line through the data:

    ŷ = w · x + b
    • x — the input (e.g. hours studied),
    • ŷ (“y-hat”) — the prediction (e.g. marks),
    • w — the weight or slope: how much ŷ changes when x grows by 1,
    • b — the bias or intercept: the prediction when x = 0.

    Training means finding the w and b that make the line fit the data as well as possible.

    What does “best” mean? — Mean squared error

    For each data point, the residual (error) is ŷᵢ − yᵢ — the red vertical line in the 3D model. We square each error and take the average:

    MSE = (1/n) · Σ (ŷᵢ − yᵢ)²

    Squaring makes all errors positive (so +3 and −3 don’t cancel) and punishes big mistakes much more than small ones. In the model, each squared error is drawn as an actual square sticking out of the wall — rotate the camera to see them. “Least squares” literally means: make the total area of these squares as small as possible.

    Training with gradient descent

    We start with a bad line (w = 0, b = 0) and improve it using gradient descent. The gradients of the MSE are:

    ∂MSE/∂w = (2/n) Σ (ŷᵢ − yᵢ) · xᵢ
    ∂MSE/∂b = (2/n) Σ (ŷᵢ − yᵢ)

    Then step against them: w ← w − η·∂MSE/∂w, b ← b − η·∂MSE/∂b. Press Train to watch the line rotate and lift into place while the squares shrink.

    The exact answer — normal equation

    For a straight line there is also a formula that gives the best w and b directly:

    w = Σ(xᵢ − x̄)(yᵢ − ȳ) / Σ(xᵢ − x̄)²
    b = ȳ − w · x̄

    With many features it becomes w = (XᵀX)⁻¹ Xᵀy. It’s exact but needs a matrix inversion, which is slow for thousands of features — that’s why big models use gradient descent.

    Code

    import numpy as np
    
    x = np.array([1, 2, 3, 4, 5, 6], dtype=float)        # hours studied
    y = np.array([35, 45, 50, 62, 70, 78], dtype=float)  # marks
    
    # Gradient descent
    w, b, lr = 0.0, 0.0, 0.02
    for step in range(5000):
        y_hat = w * x + b
        dw = 2 * np.mean((y_hat - y) * x)
        db = 2 * np.mean(y_hat - y)
        w -= lr * dw
        b -= lr * db
    print(f"w = {w:.2f}, b = {b:.2f}")          # w ≈ 8.63, b ≈ 26.47
    
    # Exact least squares (normal equation)
    w_exact = np.sum((x - x.mean()) * (y - y.mean())) / np.sum((x - x.mean()) ** 2)
    b_exact = y.mean() - w_exact * x.mean()
    print(round(w_exact, 2), round(b_exact, 2))
    
    print("predicted marks for 7 hours:", w * 7 + b)

    With scikit-learn:

    from sklearn.linear_model import LinearRegression
    model = LinearRegression().fit(x.reshape(-1, 1), y)
    print(model.coef_[0], model.intercept_)

    Measuring how good the fit is

    • MSE / RMSE: average squared error (RMSE = √MSE, in the same units as y).
    • R² score: 1 means a perfect fit, 0 means “no better than always predicting the average”.

    Assumptions and limits

    • The relationship should be roughly linear — use polynomial features or other models for curves.
    • Outliers pull the line strongly, because errors are squared.
    • It predicts numbers, not categories — for yes/no questions use logistic regression.

    Where is it used?

    Predicting house prices, sales, temperatures, exam marks; finding trends in data; and as the building block of neural networks — every neuron starts with w·x + b!

    Common mistakes

    • Extrapolating far outside the data range (“100 hours of study → 900 marks”).
    • Forgetting to scale features with very different ranges before gradient descent.
    • Confusing correlation with causation.

    Complexity at a glance

    Case / operationTimeWhy
    Prediction for one inputO(d)d = number of features.
    One gradient descent stepO(n · d)Uses every training example.
    Normal equation (exact)O(n · d² + d³)Matrix inversion — slow for many features.
    Extra spaceO(d)

    Quick check

    Test yourself — pick an answer to see if you got it.

    1. What does the model ŷ = w·x + b learn during training?

    2. Why do we square the errors instead of just adding them up?

    3. A residual is…

    4. Linear regression predicts…

    Saved only in this browser — no account needed.
    Spotted a mistake or a bug in the 3D model?

    Report a mistake

    in Linear Regression. Thank you — every report makes the lesson better for the next reader.

    We'll also include a link to the step of the 3D model you're on and your browser type, so we can reproduce it.