Conceptual

Grouped-Query and Multi-Query Attention

sharing key/value heads across query groups cuts the KV cache several-fold at close to full multi-head quality