What is it?
Linear regression predicts a number from one or more inputs by drawing the best straight line through the data:
ŷ = w · x + b
x— the input (e.g. hours studied),ŷ(“y-hat”) — the prediction (e.g. marks),w— the weight or slope: how much ŷ changes when x grows by 1,b— the bias or intercept: the prediction when x = 0.
Training means finding the w and b that make the line fit the data as well as possible.
What does “best” mean? — Mean squared error
For each data point, the residual (error) is ŷᵢ − yᵢ — the red vertical line in the 3D model. We square each error and take the average:
MSE = (1/n) · Σ (ŷᵢ − yᵢ)²
Squaring makes all errors positive (so +3 and −3 don’t cancel) and punishes big mistakes much more than small ones. In the model, each squared error is drawn as an actual square sticking out of the wall — rotate the camera to see them. “Least squares” literally means: make the total area of these squares as small as possible.
Training with gradient descent
We start with a bad line (w = 0, b = 0) and improve it using gradient descent. The gradients of the MSE are:
∂MSE/∂w = (2/n) Σ (ŷᵢ − yᵢ) · xᵢ
∂MSE/∂b = (2/n) Σ (ŷᵢ − yᵢ)
Then step against them: w ← w − η·∂MSE/∂w, b ← b − η·∂MSE/∂b. Press Train to watch the line rotate and lift into place while the squares shrink.
The exact answer — normal equation
For a straight line there is also a formula that gives the best w and b directly:
w = Σ(xᵢ − x̄)(yᵢ − ȳ) / Σ(xᵢ − x̄)²
b = ȳ − w · x̄
With many features it becomes w = (XᵀX)⁻¹ Xᵀy. It’s exact but needs a matrix inversion, which is slow for thousands of features — that’s why big models use gradient descent.
Code
import numpy as np
x = np.array([1, 2, 3, 4, 5, 6], dtype=float) # hours studied
y = np.array([35, 45, 50, 62, 70, 78], dtype=float) # marks
# Gradient descent
w, b, lr = 0.0, 0.0, 0.02
for step in range(5000):
y_hat = w * x + b
dw = 2 * np.mean((y_hat - y) * x)
db = 2 * np.mean(y_hat - y)
w -= lr * dw
b -= lr * db
print(f"w = {w:.2f}, b = {b:.2f}") # w ≈ 8.63, b ≈ 26.47
# Exact least squares (normal equation)
w_exact = np.sum((x - x.mean()) * (y - y.mean())) / np.sum((x - x.mean()) ** 2)
b_exact = y.mean() - w_exact * x.mean()
print(round(w_exact, 2), round(b_exact, 2))
print("predicted marks for 7 hours:", w * 7 + b)
With scikit-learn:
from sklearn.linear_model import LinearRegression
model = LinearRegression().fit(x.reshape(-1, 1), y)
print(model.coef_[0], model.intercept_)
Measuring how good the fit is
- MSE / RMSE: average squared error (RMSE = √MSE, in the same units as y).
- R² score: 1 means a perfect fit, 0 means “no better than always predicting the average”.
Assumptions and limits
- The relationship should be roughly linear — use polynomial features or other models for curves.
- Outliers pull the line strongly, because errors are squared.
- It predicts numbers, not categories — for yes/no questions use logistic regression.
Where is it used?
Predicting house prices, sales, temperatures, exam marks; finding trends in data; and as the building block of neural networks — every neuron starts with w·x + b!
Common mistakes
- Extrapolating far outside the data range (“100 hours of study → 900 marks”).
- Forgetting to scale features with very different ranges before gradient descent.
- Confusing correlation with causation.
Complexity at a glance
| Case / operation | Time | Why |
|---|---|---|
| Prediction for one input | O(d) | d = number of features. |
| One gradient descent step | O(n · d) | Uses every training example. |
| Normal equation (exact) | O(n · d² + d³) | Matrix inversion — slow for many features. |
| Extra space | O(d) |
Quick check
Test yourself — pick an answer to see if you got it.
1. What does the model ŷ = w·x + b learn during training?
Training adjusts the slope w and intercept b to make predictions close to the real values.
2. Why do we square the errors instead of just adding them up?
An error of +3 and −3 would sum to 0. Squaring makes every error positive and punishes large mistakes strongly.
3. A residual is…
The red vertical lines in the model are the residuals.
4. Linear regression predicts…
Regression predicts numbers; classification (e.g. logistic regression) predicts categories.