← Glossary
Glossary
What is Self-Attention?
Each token computes query, key, and value vectors. Attention weight between two tokens = dot product of their query and key, scaled and softmaxed. Output = weighted sum of value vectors. Lets every token see every other token.
What people say
How the model decides what to focus on