KTU S1

Unconstrained Optimisation and the Second Derivative Test

By the end you should be able to: Locate critical points of a two-variable function, classify them using the discriminant, and distinguish local from global extrema.

Optimisation is where calculus earns its keep: least cost, least error, most throughput. In one variable you set f′(x)=0f'(x)=0 and check the second derivative. Two variables need slightly more machinery, for a reason worth understanding.

Critical points

At a smooth maximum or minimum the surface is level in every direction, so every directional derivative is zero. Since Duf=∇f⋅uD_{\mathbf u}f = \nabla f\cdot \mathbf u, that forces the gradient itself to vanish:

∇f=0⟺fx=0   and   fy=0\nabla f = \mathbf 0 \qquad\Longleftrightarrow\qquad f_x = 0 \;\text{ and }\; f_y = 0

Solve both simultaneously to find the critical points.

The new possibility

In one variable a critical point is a max, a min, or an inflection. In two, a fourth case appears — and it is the interesting one.

f(x,y)=x2−y2f(x,y) = x^2 - y^2

At the origin both partials vanish. But along the xx-axis the surface is x2x^2, a minimum; along the yy-axis it is −y2-y^2, a maximum. It is a minimum in one direction and a maximum in another: a saddle point, shaped like a mountain pass or a Pringle.

Testing a couple of directions cannot detect this. Hence the test below.

The second derivative test

Compute the second partials and form the discriminant

D=fxxfyy−(fxy)2D = f_{xx}f_{yy} - (f_{xy})^2

evaluated at the critical point. Then:

ConditionConclusion
D>0D > 0 and fxx>0f_{xx} > 0local minimum
D>0D > 0 and fxx<0f_{xx} < 0local maximum
D<0D < 0saddle point
D=0D = 0test fails — investigate directly

DD is the determinant of the Hessian matrix

H=(fxxfxyfyxfyy)H = \begin{pmatrix} f_{xx} & f_{xy} \\ f_{yx} & f_{yy}\end{pmatrix}

and by Clairaut's theorem fxy=fyxf_{xy}=f_{yx}, so HH is symmetric.

Why −(fxy)2-(f_{xy})^2? The cross term measures how the two directions interact. If it is large enough, the surface curves down along some diagonal even when it curves up along both axes — that is the saddle, and the subtraction is what detects it.

Worked classification

f(x,y)=x2+y2−4x−6y+13f(x,y) = x^2+y^2-4x-6y+13 fx=2x−4=0  ⟹  x=2,fy=2y−6=0  ⟹  y=3f_x = 2x-4 = 0 \implies x = 2, \qquad f_y = 2y-6 = 0 \implies y = 3

Critical point (2,3)(2,3). Second partials:

fxx=2,fyy=2,fxy=0f_{xx} = 2, \quad f_{yy} = 2, \quad f_{xy} = 0 D=(2)(2)−02=4>0,fxx=2>0D = (2)(2) - 0^2 = 4 > 0, \qquad f_{xx} = 2 > 0

So (2,3)(2,3) is a local minimum, with value

f(2,3)=4+9−8−18+13=0f(2,3) = 4+9-8-18+13 = 0

In fact completing the square gives f=(x−2)2+(y−3)2f = (x-2)^2+(y-3)^2, manifestly never negative — so this local minimum is also the global minimum.

Local versus global

The test is local: it examines curvature at a single point. A function can have several local minima of different depths, and the test cannot tell you which is lowest — it never looks beyond the immediate neighbourhood.

To find a global extremum you must compare all critical values, and also check the boundary of the region if one exists.

This distinction is not academic. Gradient descent finds a local minimum; whether it is the global one is, in general, unknown — which is why training a neural network twice from different starting points can give different results.