1. Home
  2. AI & Machine Learning
  3. Gradient Descent

Gradient Descent

The algorithm that trains almost every AI model — walking downhill on a loss surface. Roll a ball across real 3D surfaces and tune the learning rate.

Interactive 3DBeginner12 min readAI/MLUpdated

Drag to rotate · Right-drag to pan · Click, then scroll to zoom · Space play · ←→ step

What's happening

Pseudocode

    Try this in the 3D model

    • Run it on the Bowl with η = 0.1, then 0.01. How many steps does each take?
    • Set η = 2.1 on the Bowl. What happens to the ball?
    • Choose Narrow valley with η = 0.7 and watch the zig-zag.
    • Choose Two valleys and click different starting points on the surface. Do you always reach the deepest valley?

    What problem does it solve?

    Training a machine-learning model means finding the parameters (weights) that make its predictions as good as possible. We measure “how bad” with a loss function: a big loss means many mistakes, a small loss means good predictions.

    If a model had only two parameters, x and y, we could draw the loss as a surface — a landscape where the height is the loss. Training = finding the lowest point of that landscape. That’s exactly the colourful surface in the 3D model.

    Real models have millions or billions of parameters, so we can’t draw the landscape or check every point. We need a smarter way.

    The idea: walk downhill in fog

    Imagine standing on a hillside in thick fog. You can’t see the valley, but you can feel the slope under your feet. So you take a step in the steepest downhill direction, feel again, take another step… and eventually you reach the bottom.

    That slope is the gradient, written ∇loss. It is a vector of partial derivatives — one per parameter — and it points in the direction where the loss increases fastest. So we step the opposite way:

    x ← x − η · ∂loss/∂x
    y ← y − η · ∂loss/∂y

    η (eta) is the learning rate — how big each step is. In the model, the green arrow shows the downhill direction at the ball’s position.

    The learning rate matters a lot

    Learning rate What happens Try it
    Too small (0.01) Tiny, safe steps — training takes forever Bowl, η = 0.01
    Just right (0.1–0.3) Smooth, quick convergence Bowl, η = 0.1
    A bit large (0.7) Overshoots and zig-zags across narrow valleys Narrow valley, η = 0.7
    Too large (2.1) Every step lands higher than before — diverges Bowl, η = 2.1

    Picking the learning rate is one of the most important decisions when training a neural network.

    Local minima

    On the Two valleys surface there is a deep valley (the global minimum) and a shallower one (a local minimum). At the bottom of either one the slope is zero, so gradient descent stops — it has no way to know a deeper valley exists. Where you start decides where you end up. Click different points on the surface to see this.

    In practice, techniques such as random restarts, momentum (keep some speed from previous steps so you can roll through small bumps) and the Adam optimiser help a lot. Interestingly, in very high-dimensional neural networks, true bad local minima turn out to be rare — the bigger problem is flat regions and saddle points.

    A worked example

    Take the bowl loss(x, y) = 0.5·(x² + y²). Its gradient is simply (x, y). Starting at (2, 1) with η = 0.1:

    Step x y gradient loss
    0 2.00 1.00 (2.00, 1.00) 2.50
    1 1.80 0.90 (1.80, 0.90) 2.03
    2 1.62 0.81 (1.62, 0.81) 1.64
    3 1.46 0.73 (1.46, 0.73) 1.33

    Each step multiplies the position by 0.9, so it slides smoothly towards (0, 0), where the loss is 0.

    Code

    def loss(x, y):
        return 0.5 * (x**2 + y**2)
    
    def gradient(x, y):
        return x, y                      # ∂loss/∂x = x, ∂loss/∂y = y
    
    x, y = 2.0, 1.0                      # starting guess
    lr = 0.1                             # learning rate η
    
    for step in range(100):
        gx, gy = gradient(x, y)
        x -= lr * gx                     # step downhill
        y -= lr * gy
        if (gx**2 + gy**2) ** 0.5 < 1e-4:
            break
    
    print(f"stopped after {step} steps at ({x:.4f}, {y:.4f}), loss = {loss(x, y):.6f}")

    Gradient descent for linear regression

    The same loop trains a real model. Here we fit a line ŷ = w·x + b to data by minimising the mean squared error:

    xs = [1, 2, 3, 4, 5]
    ys = [3, 5, 7, 9, 11]                # the true line is y = 2x + 1
    w, b, lr = 0.0, 0.0, 0.02
    
    for epoch in range(2000):
        n = len(xs)
        dw = sum(2 * (w * x + b - y) * x for x, y in zip(xs, ys)) / n
        db = sum(2 * (w * x + b - y)     for x, y in zip(xs, ys)) / n
        w -= lr * dw
        b -= lr * db
    
    print(round(w, 3), round(b, 3))      # ≈ 2.0 1.0

    Variants you will hear about

    • Batch gradient descent — uses the whole dataset for every step (accurate but slow).
    • Stochastic gradient descent (SGD) — uses one random example per step (noisy but fast).
    • Mini-batch — uses small batches like 32 or 256 examples: the standard in deep learning.
    • Momentum, RMSProp, Adam — smarter step rules that adapt speed and direction.

    Common mistakes

    • Adding the gradient instead of subtracting it (walking uphill!).
    • Not scaling input features — one steep direction forces a tiny learning rate (the narrow-valley problem).
    • Judging convergence too early: check that the loss has really stopped decreasing.

    Complexity at a glance

    Case / operationTimeWhy
    One step (n parameters)O(n)Compute the gradient and update every parameter.
    One step on a dataset of m examples (batch)O(m · n)The gradient sums over every training example.
    Stochastic / mini-batch stepO(b · n)Uses only b examples per step — much cheaper.
    Extra spaceO(n)

    Quick check

    Test yourself — pick an answer to see if you got it.

    1. In gradient descent, which direction do we move the parameters?

    2. What usually happens if the learning rate is too large?

    3. The update rule is x ← x − η·g. If x = 3, η = 0.1 and g = 4, what is the new x?

    4. Why might gradient descent stop in a "local minimum"?

    Saved only in this browser — no account needed.
    Spotted a mistake or a bug in the 3D model?

    Report a mistake

    in Gradient Descent. Thank you — every report makes the lesson better for the next reader.

    We'll also include a link to the step of the 3D model you're on and your browser type, so we can reproduce it.