Conceptual

Tri-Modal Valence-Arousal Matching Across Images, Music, and Captions in Affective Computing

A tri-modal encoder framework that aligns images, music and musical captions in a shared emotional space defined by continuous valence and arousal values, enabling similarity-scored random pairing instead of one-to-one matching, state-of-the-art valence-arousal prediction, and zero-shot transfer to downstream emotion tasks.