Absolute Position Encoding

How a Transformer learns which word came first

A Transformer reads the whole sequence at once, so order is lost. Position encoding puts a unique, smooth signal into every slot.

← → step position · space play/pause · Esc leave tour

The problem

"dog" appears twice. The embedding knows the word, not the slot — so both copies start identical.

"dog" · slot 0

"dog" · slot 6

Position × dimension

Rows = positions, columns = dimensions, colour = value. Hover to read a cell; click a row to jump there.

−10+1
colour = value
i=0 · fasti max · slow
dimension colour (clocks & bars)
Hover a cell for its value.

The formula

Even dims use sine, the paired odd dim uses cosine — same frequency, quarter-turn apart.

Even dims (0, 2, 4, …) use sine.

Odd dims (1, 3, 5, …) use cosine — same frequency.

Why base = 10,000?

A large base spreads the frequencies from fast to nearly constant, so every position stays unique — even in long sequences. Drag the base slider.

Frequency clocks

Each dimension pair is a clock hand with its own speed. Click one to spin the sequence; drag to set it yourself. The pale arc = angle swept so far.

pos=0

Position similarity

How similar the current position is to every other one — neighbours look alike, distant positions drift to zero. Axis: Δ = distance from the reference (grey = absolute position). Hover to read, click to move.

Reference: pos=0

Putting it together

Token embedding + position encoding = what the model actually sees.

pos=0

Token embedding

+

Position encoding

=

Final input vector

Bar height = value (above/below the line = positive/negative), bar colour = dimension (same scale as the clocks). The heatmap uses colour for value instead.