The problem: too many features
Real datasets can have dozens or thousands of features. Many are correlated — height and weight, or neighbouring pixels in an image — so they carry overlapping information. Too many features make models slow, hard to visualise, and prone to overfitting (the “curse of dimensionality”).
Principal Component Analysis finds a smaller set of new features — combinations of the old ones — that keep as much of the information as possible.
The intuition
Look at the 3D cloud in the model. It’s long in one direction, thinner in another and almost flat in the third. If you had to describe each point with just two numbers, you’d measure along the two directions where the points spread out most and ignore the thin direction. That’s exactly PCA.
The steps
- Centre the data — subtract the mean of each feature.
- Compute the covariance matrix
C(d × d). Its entries say how pairs of features vary together. - Find the eigenvectors and eigenvalues of
C.- The eigenvectors are the principal components (PC1, PC2, …) — perpendicular directions.
- Each eigenvalue is the variance along its component.
- Sort components by eigenvalue (largest first).
- Project the data onto the top
kcomponents. Now each point haskfeatures instead ofd.
Explained variance ratio = eigenvalue ÷ sum of eigenvalues. If PC1 + PC2 explain 97%, two numbers describe each point almost perfectly.
In the 3D model: the cloud is centred, the three PC arrows appear (lengths ∝ √variance), the cloud is rotated so PC1/PC2/PC3 line up with the axes, and finally PC3 is dropped — the cloud flattens onto a plane.
Code
From scratch with NumPy:
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(200, 3)) @ np.array([[3, 1, 0.5], [0, 1, 0.3], [0, 0, 0.2]])
Xc = X - X.mean(axis=0) # 1. centre
C = np.cov(Xc, rowvar=False) # 2. covariance (3 × 3)
values, vectors = np.linalg.eigh(C) # 3. eigen-decomposition
order = np.argsort(values)[::-1] # 4. sort, biggest first
values, vectors = values[order], vectors[:, order]
k = 2
X2 = Xc @ vectors[:, :k] # 5. project → 200 × 2
print("explained variance:", np.round(values / values.sum(), 3))
With scikit-learn:
from sklearn.decomposition import PCA
pca = PCA(n_components=2).fit(X)
X2 = pca.transform(X)
print(pca.explained_variance_ratio_)
Choosing k
Plot the cumulative explained variance and keep enough components to reach, say, 95%. Or pick 2–3 when the goal is visualisation.
Important details
- Scale features (standardise) first if they use different units — otherwise the feature with the biggest numbers dominates.
- PCA finds linear structure. For curved structure, try t-SNE, UMAP or kernel PCA (for visualisation).
- The new features are mixtures of the old ones, so they’re harder to interpret.
Where is PCA used?
Visualising high-dimensional data in 2D, compressing images, speeding up models, removing noise, face recognition (“eigenfaces”), and finding the main patterns in genetics and finance data.
Common mistakes
- Forgetting to centre (or scale) the data.
- Using PCA as a classifier — it’s a preprocessing step.
- Assuming the first component always matters most for your prediction task — it only maximises variance, not accuracy.
Complexity at a glance
| Case / operation | Time | Why |
|---|---|---|
| Covariance matrix (n points, d features) | O(n · d²) | |
| Eigen-decomposition | O(d³) | |
| Projecting the data | O(n · d · k) | |
| Extra space | O(d²) |
Quick check
Test yourself — pick an answer to see if you got it.
1. What is the first principal component?
PC1 captures the most variance; PC2 the most of what is left, and so on.
2. Why do we centre (subtract the mean) before PCA?
Covariance and the principal directions are defined around the mean.
3. The eigenvalues of the covariance matrix tell us…
Dividing each eigenvalue by their sum gives the "explained variance ratio".
4. PCA is an example of…
It needs no labels — it only looks at how the features vary together.