Iterative Minimisation by Steepest Descent
Setting and solving works when the equations are tractable. With a hundred variables and no closed form, it is hopeless.
Steepest descent gives up on solving exactly and walks downhill instead.
The idea
Module 3 established that points in the direction of steepest ascent. To go down, go the other way.
and for several variables
The number is the step size — the learning rate in machine learning. Start somewhere, repeat, stop when the steps become negligible.
A worked run
Minimise , whose answer we already know is . Take , , starting at :
| 0 | 1 | 2 | 0.8 |
| 1 | 0.8 | 1.6 | 0.64 |
| 2 | 0.64 | 1.28 | 0.512 |
| 3 | 0.512 | 1.024 | 0.4096 |
Each step multiplies by , so . It approaches zero without ever arriving — which is normal. Iterative methods converge; they do not terminate.
The step size decides everything
| Behaviour on | |
|---|---|
| Too small (0.001) | Correct direction, painfully slow |
| Well chosen (0.1) | Steady geometric convergence |
| in one step — exact, by luck | |
| Too large (1.5) | : overshoots, oscillates, diverges |
With starting from : . The method runs away from the minimum it is meant to find. Too large a step is not merely slow — it is wrong.
Stopping
There is no exact arrival, so choose a criterion:
- below a tolerance — the ground is flat enough
- below a tolerance — progress has stalled
- a maximum iteration count, so a failing run still ends
The honest limitation
Steepest descent goes downhill from where it starts. Drop it into a valley and it finds the bottom of that valley. If a deeper one exists elsewhere, it will never learn of it — nothing in the method looks beyond the local slope.
This is the practical face of the local-versus-global distinction from Module 3, and it is why training a neural network from two different random initialisations can produce two different models.
Why it dominates in practice
Each step needs only the gradient — no second derivatives, no matrix inversion, no solving of systems. For a function of a million variables that is the difference between feasible and not. Nearly all modern machine learning is a refinement of this loop.