Skip to content
Open access

Parameter-Efficient Audio-Visual Dynamic Facial Expression Recognition with Mamba Fusion Adapters and Frame-Level Feature Arrangement

Aug 2026 · Italian National Conference on Sensors · Vol 26 · 0 citations · 63 references
Medicine

Abstract

Dynamic Facial Expression Recognition (DFER) has recently attracted significant interest due to its vital role in enabling empathetic and human-compatible technologies. Developing models that remain robust under in-the-wild variability is a key motivation for DFER research and its practical applications. Improving models through multimodal learning, which leverages audio and video data for richer, complementary representations, is one promising direction. However, existing methods still rely heavily on modality-specific encoders and coarse-grained content-level alignment, which hinders their ability to capture fine-grained emotional semantics and dynamic cross-modal interactions. To address this, we adopt parameter-efficient fine-tuning (PEFT) to facilitate audio-visual interaction. This strategy offers key advantages: (1) freezing parameters preserves upstream pretrained knowledge, ensuring that the model focuses solely on learning modules for audio-visual interaction and modal fusion; (2) a Mamba Fusion Adapter (MFAdapter) is inserted at each encoder layer to perform causal, audio-conditioned fusion over a frame-aligned token sequence, enabling efficient multi-level cross-modal injection with linear complexity; and (3) a Frame-level Feature Arrangement (FFA) strategy is introduced as a deterministic index-based arrangement scheme that arranges audio tokens at the video-frame rate, providing a frame-indexed temporal prior that supports the causal scan in MFAdapter; FFA reduces, but does not eliminate, coarse segment-level mismatch and is not claimed as verified frame-level synchronization. Notably, our method achieves competitive performance on DFEW and MAFW while updating 2.7% (4.7 million) of the model parameters, with the best WAR of 58.70% on MAFW among all compared methods and 76.62%/65.25% (WAR/UAR) on DFEW under the official five-fold cross-validation protocols.

Read PDF

Similar papers

Open access Aug 2026

Lightweight and robust audio-visual emotion recognition via multi-scale mamba temporal modeling and quality-aware expert fusion

Experiments demonstrate that the proposed lightweight audio-visual emotion recognition framework achieves competitive recognition accuracy with significantly fewer parameters and lower computational cost.

Tianxing Zhang, Hadi Affendy Bin Dahlan, Fadhilah Rosdi et al. · 0 citations
Aug 2026

Bidirectional joint cross-attention framework for transformer based audio–visual emotion recognition

Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.

Arman Sajjadi, M. Nekou, Sayna Sarvar et al. · 0 citations

A modified mobilenetV2 using selective kernels and coordinate attention for facial expression recognition

Facial Expression Recognition (FER) plays an important role in human-computer interaction (HCI). However, most highly accurate FER models rely on complex and computationally heavy architectures that makes them unsuitable for low resource devices. This study proposes a modified MobileNetV2-based lightweight network that...

Shahzad Ali · 0 citations
Open access 2026

DCHF: Dual-Stream Cooperative Perception with Hierarchical Fusion Network for Micro-Expression Recognition

: Micro-expression recognition (MER) is a challenging task because micro-expressions are extremely short in duration, weak in intensity, and often distributed over subtle local facial regions. Existing methods either rely on handcrafted descriptors with limited representation capacity or focus on single-stream deep mod...

Zishi Li, Xiao-Dong Huang · 0 citations
Book Open access Oct 2026

Light-ED: Lightweight Multimodal Emotion Detection using Enhanced EfficientNet

Emotion recognition plays a key role in affective computing and human–computer interaction, where understanding emotions from multimodal signals such as facial expressions and speech remains challenging. Most existing methods treat data fusion and classification as separate stages, limiting performance and efficiency....

Wamika Jha, Mea Wang, U. Alim et al. · 0 citations
Open access Aug 2026

An Attention-Enhanced ConvNeXtTiny Model for Robust Facial Expression Recognition

Facial Expression Recognition (FER) plays an important role in affective computing and human–computer interaction by enabling automated interpretation of human emotional states from facial images. Despite recent advances in deep learning, reliable FER remains challenging because of variations in facial appearance, illu...

Manisha B. Thombare, S. Gumaste · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.