1One box moves
A block has a straight highway (the residual connection) and side branches that do the work.
Watch the gold LayerNorm box.
Attention
FFN
LayerNorm = rescale to size 1
⊕ = add branch back
Post-LN out = LN( x + F(x) )
Pre-LN out = x + F( LN(x) )
2Going up: how big does the signal get?
Stack many blocks. Send a signal from the bottom and watch its width — that's its size.
3Going down: how hard does each layer learn?
During training a learning signal (the gradient) flows back down. Brighter layer = bigger push on its weights.
4Make it deeper
Size of the push on the top layer, as we add more and more layers.
computing…
5So which one?
Each placement protects something different.
Post-LN
Original Transformer (2017), BERT
- +Every layer keeps a strong voice
- +Often slightly better results — when it trains
- −Big pushes everywhere → unstable, needs a slow “warm-up” start
- −Hard to make very deep
Pre-LN
GPT-2 onward, most modern LLMs
- +Clean highway → stable training
- +Easy to stack very deep, little or no warm-up
- −Signal keeps growing → top layers get quieter
- −Needs one extra LayerNorm at the very end
Post-LN protects each layer's voice. Pre-LN protects the highway.
6Your turn: build the block
Fill the 6 slots from bottom (input) to top (output). Tap a part, then tap a slot — or drag it.
Bonus: after stacking many of these blocks, what goes before the output head?