Composition
Given f(x)=x2 and g(x)=sinx, feeding one into the other gives
(f∘g)(x)=f(g(x))=sin2x(g∘f)(x)=g(f(x))=sin(x2)
Order matters. These are different functions; try x=1 and compare.
The chain rule, one variable
dxdf(g(x))=f′(g(x))⋅g′(x)
Or in Leibniz form, which is easier to remember and to generalise:
dxdy=dudy⋅dxdu
What it says. A small nudge in x changes u by g′(x) times as much;
that change in u changes y by f′(u) times as much. Multiply the two
amplification factors.
For y=sin(x2): put u=x2, so y=sinu.
dxdy=cos(u)⋅2x=2xcos(x2)
The multivariable chain rule
Now let z=f(x,y) where both x and y depend on a parameter t. As t
moves, both inputs move, and each drags z with it:
dtdz=∂x∂zdtdx+∂y∂zdtdy
Sum the contributions, one per input. Each term is "how much z responds
to that input" times "how fast that input is moving".
Note the mixture of ∂ and d: partial for z against its own inputs,
ordinary for anything depending on t alone.
A worked instance with a check
z=x2+y2, with x=cost and y=sint.
∂x∂z=2x,∂y∂z=2y,dtdx=−sint,dtdy=cost
dtdz=2x(−sint)+2y(cost)=−2costsint+2sintcost=0
Check it directly. Substituting first,
z=cos2t+sin2t=1, a constant — so dz/dt=0. The two routes
agree.
The geometry: the point (cost,sint) travels around the unit circle, which
is itself a level curve of x2+y2. Moving along a contour, height never
changes. The chain rule and the picture say the same thing.
More inputs
The pattern extends. For w=f(x,y,z) with all three depending on t:
dtdw=∂x∂wdtdx+∂y∂wdtdy+∂z∂wdtdz
One term per input, always.
Why it matters later
Backpropagation in a neural network is the chain rule applied through many
layers. The gradient of a loss with respect to an early weight is a product of
the sensitivities at every layer in between — exactly this rule, at scale.