Conceptual

Heterogeneous Feature Fusion for Multi-Turn Spoken Dialogue Systems

A design for spoken dialogue systems that must respond to speech content plus paralinguistic and environmental cues - emotion, audio events, and background music - across multi-turn conversations. Students learn a heterogeneous feature fusion module that adaptively selects and combines different audio and text feature streams according to the dialogue context, and the accompanying recipe for training such a system on a mix of synthetic and real dialogue data, including how to find the synthetic-to-real ratio that maximizes real-world performance while covering scenarios too rare to collect naturally.