Skip to content
Open access

Emotion recognition from body movement through interpretable motion-aware sequential modeling

Aug 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 58 references
Medicine

TL;DR

A Window Transformer architecture grounded in the Multiple Instance Learning (MIL) paradigm, which decomposes the full sequence as a single temporal stream into overlapping windows and learns to assign greater relevance to those segments containing stronger emotional content, providing a more interpretable framework.

Abstract

Emotion recognition from bodily movement remains a challenging problem, particularly when only pose-based motion sequences are available and emotionally informative content is not uniformly distributed across time. In this work, we propose a Window Transformer architecture grounded in the Multiple Instance Learning (MIL) paradigm to address this challenge. Rather than processing the full sequence as a single temporal stream, the model decomposes it into overlapping windows and learns to assign greater relevance to those segments containing stronger emotional content. This formulation provides a more interpretable framework, since the learned relevance scores reveal which temporal regions drive the final prediction, while also yielding richer, context-aware representations of each segment. We evaluate the proposed approach on two publicly available datasets, MEED and DIEM-A, and compare it against a standalone Transformer baseline under different batch size and window configuration settings. The Window Transformer consistently outperforms the baseline and exhibits a more stable behavior across training configurations, achieving best accuracies of 56.35 ± 2.67% on MEED and 23.84 ± 0.83% on DIEM-A in a subject-independent scenario. Beyond performance gains, this work also establishes new benchmark values on both datasets, providing reference results that future research on bodily emotion recognition can build on.

Read PDF

Similar papers

Preprint Aug 2026

LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

LG-GER is proposed, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence that achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.

Ahmed-Shehab Khan, Zhiyuan Li, Yan Tong · 0 citations
Review Open access Aug 2026

A Systematic Review of Emotion Recognition: From Unimodal Signals to Multimodal Integration

These findings reveal that multimodal systems, which fuse visual, acoustic, linguistic, linguistic, and physiological signals, consistently outperform unimodal counterparts, achieving accuracy levels above 85% on benchmark datasets.

B. Bashir, Zayyanu Yunusa · 0 citations
#artificial intelligence Preprint Sep 2026

LongEmo: Towards Emotion Understanding and Reasoning in Long Videos

While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shap...

Shuo Zhang, Yifan Zhou, Han-Yu Wang et al. · 0 citations
#human-computer interacti... Preprint Aug 2026

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

This paper introduces OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction, and proposes Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation.

Jiahao Huang, Zheng Lian, Jingyi Zhang et al. · 0 citations
Book Open access Oct 2026

HAAS: Holistic Attention-free Animation from Speech using Mamba

Synthesizing holistic co-speech gestures that integrate facial expressions and full-body motion is essential for embodied conversational agents in fields such as virtual reality and film. Existing state-of-the-art approaches predominantly rely on attention-based architectures, which suffer from quadratic computational...

Laxmi Narayen Nagarajan Venkatesan, Harsh Vardhan Singh, Vansh Sinha et al. · 0 citations
Review Open access Aug 2026

Audio-Visual-Textual Fusion for Emotion Recognition: A Neurologically Inspired Attention Mechanism

Systems that recognise emotion from speech, facial behaviour and language combine the three streams with mechanisms borrowed from machine translation rather than from any account of how the brain performs the same task. This article sets out a fusion mechanism derived from four established findings in multisensory neur...

Navin Chandran · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.