Multi-Head Attention
Key concept: Attention heads can read the same tokens with different weights and contribute separate outputs.
Add heads one at a time, select a Reader, and inspect a head’s weight matrix. The final panel combines active outputs; this toy uses their mean in place of a learned output projection Wᴼ.
What to look for: each head follows a different pattern, so adding a head changes the combined result.
Open the demo in a new tab for more room.