Conceptual
Login

Full Multi-Head Attention Where the KV Cache Is Not the Constraint

multi-head attention is still the training-time formulation and still correct for short-context and small models — grouping is an inference concession