Conceptual

Quadratic Cost of Self-Attention in Sequence Length

Why every token attending to every other makes cost grow with the square of the input, and why halving sequence length quarters the work.