Conceptual

Multimodal RAG Framework for Long-Dialogue Emotional Causal Reasoning

CauseMotion: a framework that lets an LLM infer emotional cause-and-effect chains spanning dozens of conversational turns. It fuses audio-derived affective features (vocal emotion, emotional intensity, speech rate) with text, and uses Retrieval-Augmented Generation with a sliding-window mechanism to pull in the contextually relevant earlier segments of a long dialogue. The design targets the failure of plain LLMs to track intricate emotional causality across long-form conversations, yielding measurable gains in causal-inference accuracy over text-only baselines.