Unconstrained Optimisation and the Second Derivative Test
Optimisation is where calculus earns its keep: least cost, least error, most throughput. In one variable you set and check the second derivative. Two variables need slightly more machinery, for a reason worth understanding.
Critical points
At a smooth maximum or minimum the surface is level in every direction, so every directional derivative is zero. Since , that forces the gradient itself to vanish:
Solve both simultaneously to find the critical points.
The new possibility
In one variable a critical point is a max, a min, or an inflection. In two, a fourth case appears — and it is the interesting one.
At the origin both partials vanish. But along the -axis the surface is , a minimum; along the -axis it is , a maximum. It is a minimum in one direction and a maximum in another: a saddle point, shaped like a mountain pass or a Pringle.
Testing a couple of directions cannot detect this. Hence the test below.
The second derivative test
Compute the second partials and form the discriminant
evaluated at the critical point. Then:
| Condition | Conclusion |
|---|---|
| and | local minimum |
| and | local maximum |
| saddle point | |
| test fails — investigate directly |
is the determinant of the Hessian matrix
and by Clairaut's theorem , so is symmetric.
Why ? The cross term measures how the two directions interact. If it is large enough, the surface curves down along some diagonal even when it curves up along both axes — that is the saddle, and the subtraction is what detects it.
Worked classification
Critical point . Second partials:
So is a local minimum, with value
In fact completing the square gives , manifestly never negative — so this local minimum is also the global minimum.
Local versus global
The test is local: it examines curvature at a single point. A function can have several local minima of different depths, and the test cannot tell you which is lowest — it never looks beyond the immediate neighbourhood.
To find a global extremum you must compare all critical values, and also check the boundary of the region if one exists.
This distinction is not academic. Gradient descent finds a local minimum; whether it is the global one is, in general, unknown — which is why training a neural network twice from different starting points can give different results.