TAI BUI
← Glossary
Glossary

What is Self-Attention?

Each token computes query, key, and value vectors. Attention weight between two tokens = dot product of their query and key, scaled and softmaxed. Output = weighted sum of value vectors. Lets every token see every other token.

What people say

How the model decides what to focus on

Why it's called that