Conceptual

Robustness of Cover-Song Identification Models on YouTube Versions

An empirical finding that state-of-the-art cover-song (version) identification models, trained and benchmarked almost exclusively on SecondHandSongs-derived community datasets, degrade sharply on out-of-distribution YouTube cover versions. Using multi-modal uncertainty sampling to surface the most uncertain YouTube candidates plus Mechanical Turk annotation, the authors build a new evaluation set on which current models score significantly lower ranking performance, pinpoint version types that are especially hard to rank (instrumental, karaoke, background accompaniment, chunk covers), and provide a taxonomy of the alterations cover versions undergo on the web. The study exposes the alignment and distribution-shift blind spots of version-identification benchmarks.