How a Transformer learns which word came first
A Transformer reads the whole sequence at once, so order is lost. Position encoding puts a unique, smooth signal into every slot.
← → step position · space play/pause · Esc leave tour
"dog" appears twice. The embedding knows the word, not the slot — so both copies start identical.
Rows = positions, columns = dimensions, colour = value. Hover to read a cell; click a row to jump there.
Even dims use sine, the paired odd dim uses cosine — same frequency, quarter-turn apart.
Even dims (0, 2, 4, …) use sine.
Odd dims (1, 3, 5, …) use cosine — same frequency.
A large base spreads the frequencies from fast to nearly constant, so every position stays unique — even in long sequences. Drag the base slider.
Each dimension pair is a clock hand with its own speed. Click one to spin the sequence; drag to set it yourself. The pale arc = angle swept so far.
How similar the current position is to every other one — neighbours look alike, distant positions drift to zero. Axis: Δ = distance from the reference (grey = absolute position). Hover to read, click to move.
Token embedding + position encoding = what the model actually sees.
Bar height = value (above/below the line = positive/negative), bar colour = dimension (same scale as the clocks). The heatmap uses colour for value instead.