Skip to content
Open access

Fine-Grained Pose-Aware Visual Fusion for Emotion Recognition in Conversational Video Streams

Jul 2026 · Mathematics · Vol 14, pp. 2639 · 0 citations · 36 references

TL;DR

These results establish fine-grained body language as a critical debiasing signal, recovering accuracy on the subtle expressions that text and speech alone fail to capture.

Abstract

Despite the rapid advancement of Emotion Recognition in Conversation (ERC), prevailing systems that primarily integrate language and speech exhibit substantial performance disparities on underrepresented emotion classes (e.g., Fear, Disgust). This study investigates whether fine-grained non-verbal visual modalities (facial action units, hand gestures, and body pose) can effectively mitigate these biases. We propose a multi-stream fusion architecture combining language, speech, and engineered pose-aware visual features, trained with class-imbalance-aware objectives. Experiments on MELD demonstrate that hybrid pose augmentation improves F1 on the least frequent classes: Fear +9.19%, Disgust +6.07%, Sadness +6.46%. We achieve an overall weighted F1 of 68.79%, competitive with recent state-of-the-art systems while uniquely targeting minority-class debiasing. These results establish fine-grained body language as a critical debiasing signal, recovering accuracy on the subtle expressions that text and speech alone fail to capture.

Read PDF

Similar papers

Aug 2026

Bidirectional joint cross-attention framework for transformer based audio–visual emotion recognition

Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.

Arman Sajjadi, M. Nekou, Sayna Sarvar et al. · 0 citations
Preprint Aug 2026

LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

LG-GER is proposed, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence that achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.

Ahmed-Shehab Khan, Zhiyuan Li, Yan Tong · 0 citations
Open access Aug 2026

FER20E: An Extended Facial Expression Recognition Dataset With 20 Discrete Emotions

The FER20E dataset provides a comprehensive benchmark for advancing emotion recognition in unconstrained and real-world scenarios, and a data annotation tool (DL-DAT) that follows a semi-automated, human-in-the-loop pipeline to enable scalable and reliable annotation.

Kuldeep Singh Yadav, Lalan Kumar · 2 citations
Open access Aug 2026

Progressive Evolution of Emotion Detection: From Unimodal Baselines to a Quad-Modal Dynamic Fusion Architecture

The architecture, while forgoing the precision of a studio-lit, well-audio-recorded environment, is able to sacrifice some accuracy for the robustness required to function in a real-world human-computer interaction environment.

Atharv Shukla, Rashi Agarwal · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.