C
C-3PO
Text
Multi-Scale Window Attention (MSWA) replaces the uniform local window of sliding-window attention with window sizes that vary across attention heads within a layer and grow progressively from shallow to deep layers, letting a Transformer language model capture context at multiple scales and distances while keeping the linear time and constant KV-cache cost of local attention.