LayerNorm vs BatchNorm vs GroupNorm

Same (Batch, Sequence, Dim) tensor, three ways to choose which numbers get averaged together. Drag the cube to orbit it, click a cell to pick it, and switch the norm type.

← → nudge the last slider · space replay the last sweep · Esc leave tour

Same tensor feeds every panel below.

Which cells get pooled together?

Every cube is one number. Drag to orbit, scroll to zoom, click a cell to select it. Full-size coloured cubes are the cells pooled into one μ and σ; the outlined column is the token plotted below.

Batch axis Sequence axis Dim axis
−1+1 colour & bar height = value

Faded selectors don't move the highlighted window — but they still choose which token and channel the charts below follow.

Normalize, one channel per bar

One bar per channel of the outlined token. The dark tick is the μ that channel subtracts, the shaded band is μ ± σ. Watch who shares a tick: LayerNorm shares one μ across the token, BatchNorm gives every channel its own, GroupNorm shares inside a group. Click a bar to focus it.

Both statistics are taken over the highlighted window — that choice is the only difference between the three norms.

channel value μ used by that channel μ ± σ raw value

The learnable affine step, one channel per bar

After normalizing, each channel gets its own learned scale γ and shift β. Drag a bar's top while on ×γ or +β to set that channel directly, or use the sliders.

x
→
(x−μ)/σ
→
× γd
→
+ βd
→
y
channel value normalized x̂ (before affine) ×γscale tag β shift