Optimization · visual field guide · N°02

Four ways to find the bottom
Gradient Descent Optimizers

Each optimizer has a different strategy for navigating a loss landscape. Same starting point, same goal — but they take very different paths. Pick a landscape and watch them race.

Landscape
Learning rate
×1.0
Speed
×3
step 0 / 300
Loss landscape · Ravine SGD oscillates; Adam glides straight to the bottom
SGD Vanilla
θ ← θ − α·∇L
Loss
Steps
Momentum β = 0.9
v ← β·v + ∇L
θ ← θ − α·v
Loss
Steps
RMSProp β = 0.9
s ← β·s + (1−β)·∇L²
θ ← θ − α·∇L / √s
Loss
Steps
Adam β₁=0.9 β₂=0.999
m ← β₁·m + (1−β₁)·∇L
v ← β₂·v + (1−β₂)·∇L²
θ ← θ − α·m̂/√v̂
Loss
Steps
Loss over iterations (log scale)
SGD
Momentum
RMSProp
Adam

In a narrow ravine, SGD bounces off the walls while Adam adapts its step size per dimension — gliding straight to the minimum. Switch to Rosenbrock to see why even Adam struggles with a curved banana-shaped valley.

Tweaks live

Trail length
60
Point size
8
Show gradients
Contour bands
10