D
Data
Text
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
Most multimodal learning pairs modalities one-to-one (a cat image with a cat sound and the word 'cat'). This paper introduces MMVA, a tri-modal framework with separate tower encoders for images, musi…