Conceptual

Multimodal Enrolment Vectors for Target-Speaker Separation

A transformer-based target-speaker separation system (VoiceVector) whose enrolment network produces speaker embeddings from audio-only, audio-visual (lip movements), or silent-video visual data, and whose separation network can be conditioned on multiple positive and negative enrolment vectors, removing the reliance on clean reference audio.