Aug 2026· Italian National Conference on Sensors· Vol 26· 0 citations· 63 references
Medicine
Abstract
Dynamic Facial Expression Recognition (DFER) has recently attracted significant interest due to its vital role in enabling empathetic and human-compatible technologies. Developing models that remain robust under in-the-wild variability is a key motivation for DFER research and its practical applications. Improving models through multimodal learning, which leverages audio and video data for richer, complementary representations, is one promising direction. However, existing methods still rely heavily on modality-specific encoders and coarse-grained content-level alignment, which hinders their ability to capture fine-grained emotional semantics and dynamic cross-modal interactions. To address this, we adopt parameter-efficient fine-tuning (PEFT) to facilitate audio-visual interaction. This strategy offers key advantages: (1) freezing parameters preserves upstream pretrained knowledge, ensuring that the model focuses solely on learning modules for audio-visual interaction and modal fusion; (2) a Mamba Fusion Adapter (MFAdapter) is inserted at each encoder layer to perform causal, audio-conditioned fusion over a frame-aligned token sequence, enabling efficient multi-level cross-modal injection with linear complexity; and (3) a Frame-level Feature Arrangement (FFA) strategy is introduced as a deterministic index-based arrangement scheme that arranges audio tokens at the video-frame rate, providing a frame-indexed temporal prior that supports the causal scan in MFAdapter; FFA reduces, but does not eliminate, coarse segment-level mismatch and is not claimed as verified frame-level synchronization. Notably, our method achieves competitive performance on DFEW and MAFW while updating 2.7% (4.7 million) of the model parameters, with the best WAR of 58.70% on MAFW among all compared methods and 76.62%/65.25% (WAR/UAR) on DFEW under the official five-fold cross-validation protocols.
Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.
Arman Sajjadi, M. Nekou, Sayna Sarvar et al.· Signal, Image and Video Proc...· 0 citations
Facial Expression Recognition (FER) plays an important role in human-computer interaction (HCI). However, most highly accurate FER models rely on complex and computationally heavy architectures that makes them unsuitable for low resource devices. This study proposes a modified MobileNetV2-based lightweight network that...
: Micro-expression recognition (MER) is a challenging task because micro-expressions are extremely short in duration, weak in intensity, and often distributed over subtle local facial regions. Existing methods either rely on handcrafted descriptors with limited representation capacity or focus on single-stream deep mod...
Zishi Li, Xiao-Dong Huang· Computers, Materials & C...· 0 citations
Emotion recognition plays a key role in affective computing and human–computer interaction, where understanding emotions from multimodal signals such as facial expressions and speech remains challenging. Most existing methods treat data fusion and classification as separate stages, limiting performance and efficiency....
Facial Expression Recognition (FER) plays an important role in affective computing and human–computer interaction by enabling automated interpretation of human emotional states from facial images. Despite recent advances in deep learning, reliable FER remains challenging because of variations in facial appearance, illu...
Manisha B. Thombare, S. Gumaste· European Journal of Prosthod...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.