1. Home
  2. AI & Machine Learning
  3. Principal Component Analysis (PCA)

Principal Component Analysis (PCA)

Reduce the number of features while keeping most of the information. Watch a 3D cloud get centred, rotated onto its principal axes and flattened to 2D.

Interactive 3DIntermediate12 min readAI/MLUpdated

Drag to rotate · Right-drag to pan · Click, then scroll to zoom · Space play · ←→ step

What's happening

Pseudocode

    Try this in the 3D model

    • Before running, rotate the cloud. Can you see its long, medium and short directions?
    • After step 2, compare the arrow lengths with the percentages.
    • At the last step, how much variance is kept? Press New data and compare.

    The problem: too many features

    Real datasets can have dozens or thousands of features. Many are correlated — height and weight, or neighbouring pixels in an image — so they carry overlapping information. Too many features make models slow, hard to visualise, and prone to overfitting (the “curse of dimensionality”).

    Principal Component Analysis finds a smaller set of new features — combinations of the old ones — that keep as much of the information as possible.

    The intuition

    Look at the 3D cloud in the model. It’s long in one direction, thinner in another and almost flat in the third. If you had to describe each point with just two numbers, you’d measure along the two directions where the points spread out most and ignore the thin direction. That’s exactly PCA.

    The steps

    1. Centre the data — subtract the mean of each feature.
    2. Compute the covariance matrix C (d × d). Its entries say how pairs of features vary together.
    3. Find the eigenvectors and eigenvalues of C.
      • The eigenvectors are the principal components (PC1, PC2, …) — perpendicular directions.
      • Each eigenvalue is the variance along its component.
    4. Sort components by eigenvalue (largest first).
    5. Project the data onto the top k components. Now each point has k features instead of d.

    Explained variance ratio = eigenvalue ÷ sum of eigenvalues. If PC1 + PC2 explain 97%, two numbers describe each point almost perfectly.

    In the 3D model: the cloud is centred, the three PC arrows appear (lengths ∝ √variance), the cloud is rotated so PC1/PC2/PC3 line up with the axes, and finally PC3 is dropped — the cloud flattens onto a plane.

    Code

    From scratch with NumPy:

    import numpy as np
    
    rng = np.random.default_rng(0)
    X = rng.normal(size=(200, 3)) @ np.array([[3, 1, 0.5], [0, 1, 0.3], [0, 0, 0.2]])
    
    Xc = X - X.mean(axis=0)                      # 1. centre
    C = np.cov(Xc, rowvar=False)                 # 2. covariance (3 × 3)
    values, vectors = np.linalg.eigh(C)          # 3. eigen-decomposition
    order = np.argsort(values)[::-1]             # 4. sort, biggest first
    values, vectors = values[order], vectors[:, order]
    
    k = 2
    X2 = Xc @ vectors[:, :k]                     # 5. project → 200 × 2
    print("explained variance:", np.round(values / values.sum(), 3))

    With scikit-learn:

    from sklearn.decomposition import PCA
    pca = PCA(n_components=2).fit(X)
    X2 = pca.transform(X)
    print(pca.explained_variance_ratio_)

    Choosing k

    Plot the cumulative explained variance and keep enough components to reach, say, 95%. Or pick 2–3 when the goal is visualisation.

    Important details

    • Scale features (standardise) first if they use different units — otherwise the feature with the biggest numbers dominates.
    • PCA finds linear structure. For curved structure, try t-SNE, UMAP or kernel PCA (for visualisation).
    • The new features are mixtures of the old ones, so they’re harder to interpret.

    Where is PCA used?

    Visualising high-dimensional data in 2D, compressing images, speeding up models, removing noise, face recognition (“eigenfaces”), and finding the main patterns in genetics and finance data.

    Common mistakes

    • Forgetting to centre (or scale) the data.
    • Using PCA as a classifier — it’s a preprocessing step.
    • Assuming the first component always matters most for your prediction task — it only maximises variance, not accuracy.

    Complexity at a glance

    Case / operationTimeWhy
    Covariance matrix (n points, d features)O(n · d²)
    Eigen-decompositionO(d³)
    Projecting the dataO(n · d · k)
    Extra spaceO(d²)

    Quick check

    Test yourself — pick an answer to see if you got it.

    1. What is the first principal component?

    2. Why do we centre (subtract the mean) before PCA?

    3. The eigenvalues of the covariance matrix tell us…

    4. PCA is an example of…

    Saved only in this browser — no account needed.
    Spotted a mistake or a bug in the 3D model?

    Report a mistake

    in Principal Component Analysis (PCA). Thank you — every report makes the lesson better for the next reader.

    We'll also include a link to the step of the 3D model you're on and your browser type, so we can reproduce it.