Conceptual

Multi-Head Latent Attention and Low-Rank KV Compression

caching a shared low-rank latent instead of per-head keys and values is the current frontier answer; expect this node to move within two quarters