Skip to content

CLMER: A Framework for Contrastive Learning-Based Multimodal Emotion Recognition.

Sep 2026 · IEEE Transactions on Neural Networks and Learning Systems · Vol PP · 0 citations
Medicine

TL;DR

Experimental evaluations on two public datasets DEAP, AMIGOS and a private dataset MAN-II demonstrate that CLMER significantly outperforms unimodal and traditional fusion approaches, achieving state-of-the-art performance in emotion classification tasks.

Abstract

Emotion recognition plays a crucial role in human-computer interaction and affective computing, yet its effectiveness is limited by the difficulty of integrating heterogeneous modalities with fundamentally different structures, such as physiological signals and visual data. In this article, we propose CLMER, a contrastive learning-based multimodal cross-attention framework designed to address the challenges of complex emotion recognition. The framework introduces a serialization strategy that converts pixel-level image data into time-series data, aligning it with the temporal characteristics of physiological signals. CLMER consists of three core components that work together to enable effective multimodal emotion recognition. The multimodal data preparation module preprocesses physiological and visual data, ensuring consistency across modalities. Building on this foundation, the contrastive learning (CL)-based feature extraction module generates temporal representations that capture the essential patterns embedded in the data through self-supervised (SS) learning. Finally, the multimodal fusion module employs cross-modal attention to integrate features with improved modality alignment. Experimental evaluations on two public datasets DEAP, AMIGOS and a private dataset MAN-II demonstrate that CLMER significantly outperforms unimodal and traditional fusion approaches, achieving state-of-the-art performance in emotion classification tasks. These findings highlight the framework's robust generalization, computational efficiency, and strong performance in multimodal emotion recognition, suggesting its potential for real-world deployment. Our code is available at https://github.com/liangyubuaa/CLMER.

View source

Similar papers

Aug 2026

MS²FL: Modality-shared and Modality-specific Feature Learning for Multimodal Emotion Recognition

Multimodal emotion recognition based on complementary physiological signals such as electroencephalogram (EEG) and eye movements can effectively reflect human emotional states, demonstrating significant potential in fields such as rehabilitation monitoring and driving safety. However, existing multimodal emotion recogn...

Xin-Hui Li, Hao-Yuan Chen, Minchao Wu et al. · 0 citations
Review Open access Aug 2026

A Systematic Review of Emotion Recognition: From Unimodal Signals to Multimodal Integration

These findings reveal that multimodal systems, which fuse visual, acoustic, linguistic, linguistic, and physiological signals, consistently outperform unimodal counterparts, achieving accuracy levels above 85% on benchmark datasets.

B. Bashir, Zayyanu Yunusa · 0 citations
Book Open access Oct 2026

Light-ED: Lightweight Multimodal Emotion Detection using Enhanced EfficientNet

Emotion recognition plays a key role in affective computing and human–computer interaction, where understanding emotions from multimodal signals such as facial expressions and speech remains challenging. Most existing methods treat data fusion and classification as separate stages, limiting performance and efficiency....

Wamika Jha, Mea Wang, U. Alim et al. · 0 citations
Sep 2026

Mamba-CrossMod: a multimodal affective analysis framework based on selective state space model

The proposed Mamba-CrossMod is a novel multimodal feature fusion framework that introduces Mamba-ATT, an enhanced attention mechanism based on a selective state-space model for capturing long-range dependencies with theoretically linear complexity.

Yi-Wen Tong, Jing Mu, Wen-Xin Chang et al. · 0 citations
Conference Aug 2026

Multimodal Emotion Recognition with Emotion-Specific Cross-Modal Attention Blocks

Multimodal emotion recognition is increasingly important for healthcare, education, and human-computer interaction. However, many existing systems learn a single shared representation for all emotions, which can blur subtle class-specific cues. This paper proposes an emotion-specific multimodal architecture that combin...

Gnanaseelan Dharshika, A. Ramanan · 0 citations
Open access Aug 2026

An Attention-Guided Framework for Feature-Level and Decision-Level Fusion in Multimodal Emotion Recognition

A comparative analysis of four multimodal configurations of early Fusion without attention, early Fusion with attention, late Fusion without attention, and late Fusion complemented by attention provides a methodological framework that can be used to develop more effective and understandable MER systems using a systemat...

Chintan Chatterjee, Brijesh Bhatt · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.