J
jeremy
Video
ALiBi attention linear bias enables input length extrapolation in Transformers
The ALiBi mechanism introduces a linear bias term to transformer attention scores based on the positional distance between query and key tokens, enabling input length extrapolation without learned po…