Multimodal Temporal Modeling for Continuous Group Emotion Recognition in Multi-party Dialogues
This study forms continuous recognition of the Group Emotion at a one-second resolution, introduces the Mixed state, which captures the emotional divergence among participants in the group, and proposes a multimodal temporal framework that integrates audio and video information using a sliding-window context.