These results establish fine-grained body language as a critical debiasing signal, recovering accuracy on the subtle expressions that text and speech alone fail to capture.
Abstract
Despite the rapid advancement of Emotion Recognition in Conversation (ERC), prevailing systems that primarily integrate language and speech exhibit substantial performance disparities on underrepresented emotion classes (e.g., Fear, Disgust). This study investigates whether fine-grained non-verbal visual modalities (facial action units, hand gestures, and body pose) can effectively mitigate these biases. We propose a multi-stream fusion architecture combining language, speech, and engineered pose-aware visual features, trained with class-imbalance-aware objectives. Experiments on MELD demonstrate that hybrid pose augmentation improves F1 on the least frequent classes: Fear +9.19%, Disgust +6.07%, Sadness +6.46%. We achieve an overall weighted F1 of 68.79%, competitive with recent state-of-the-art systems while uniquely targeting minority-class debiasing. These results establish fine-grained body language as a critical debiasing signal, recovering accuracy on the subtle expressions that text and speech alone fail to capture.
Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.
Arman Sajjadi, M. Nekou, Sayna Sarvar et al.· Signal, Image and Video Proc...· 0 citations
Findings indicate that reliability-aware fusion improves benchmark emotion recognition, while the 21.50 ms per-sample RA-Gated latency on an NVIDIA GeForce RTX 3060 supports real-time-oriented applicability.
Ritu Tyagi· Journal of Intelligent Decis...· 0 citations
LG-GER is proposed, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence that achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.
Ahmed-Shehab Khan, Zhiyuan Li, Yan Tong· 0 citations
The FER20E dataset provides a comprehensive benchmark for advancing emotion recognition in unconstrained and real-world scenarios, and a data annotation tool (DL-DAT) that follows a semi-automated, human-in-the-loop pipeline to enable scalable and reliable annotation.
The architecture, while forgoing the precision of a studio-lit, well-audio-recorded environment, is able to sacrifice some accuracy for the robustness required to function in a real-world human-computer interaction environment.