Conceptual

FlashAttention

IO-aware exact attention: tiling, online softmax, and recomputation to avoid materializing the full attention matrix in GPU high-bandwidth memory.