VoiceFormer: Transformer-Based Multi-Modal Speech Separation from Text and Lip Cues
A unified Transformer framework for separating and enhancing a target speaker's voice from noisy multi-speaker mixtures by conditioning on multiple modalities - the textual content of the utterance, the speaker's lip movements, or both. Operating in the raw-waveform domain, it fuses synchronous and asynchronous cues without requiring the audio and visual streams to be temporally aligned or share a sampling rate, tolerating offsets of 200 ms or more, and achieves state-of-the-art results on the LRS2 and LRS3 audio-visual benchmarks.
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Akam Rahimi, Triantafyllos
VoiceFormer is a unified Transformer-based framework for multi-modal speech separation and enhancement - isolating a target speaker's voice from multi-speaker, noisy mixtures (the 'cocktail party' pr…