← DSAI4207 demos

Lecture 4 · Demo 7

Multi-Head Attention

Key concept: Attention heads can read the same tokens with different weights and contribute separate outputs.

Add heads one at a time, select a Reader, and inspect a head’s weight matrix. The final panel combines active outputs; this toy uses their mean in place of a learned output projection Wᴼ.

What to look for: each head follows a different pattern, so adding a head changes the combined result.

Open the demo in a new tab for more room.