Conceptual

FAST: Lightweight Audio Spectrogram Transformer with Lipschitz-Continuous Attention

FAST hybridizes convolutional layers (local time-frequency features) with transformer blocks (global context) in a MobileViT-inspired design, and stabilizes the compact model by replacing standard transformer components with Lipschitz-continuous variants (CenterNorm, Scaled Cosine Similarity Attention, Weighted Residual Shortcuts), matching or beating prior audio spectrogram transformers on ADIMA and AudioSet with up to 150x fewer parameters.