← Glossary
Glossary
What is KV Cache?
During autoregressive generation, caching the key and value matrices from previous tokens so you don't recompute them at each step. Trades memory for speed. Essential for fast LLM inference.
What people say
Makes inference faster