Conceptual

KV Caching and Grouped-Query Attention

Why autoregressive inference caches keys and values, how the cache dominates memory at long context, and how grouped-query/multi-query attention shrink it.