Skip to content
Open access

SE-POSTER: Channel-Enhanced Landmark Guided Transformer for Facial Emotion Recognition

Jul 2026 · Informatics · Vol 13, pp. 123 · 0 citations · 39 references

TL;DR

A ResNet-18–Transformer landmark-guided module called SE-POSTER that fuses lightweight Squeeze-and-Excitation attention modules into the multi-scale feature pyramid of the baseline POSTER architecture that is capable of steadily boosting recognition accuracy relative to the baseline POSTER and several state-of-the-art FER methods.

Abstract

Recognizing facial emotions automatically from images/videos (FER) still represents a difficult problem for emotion computing, mainly due to variations in the face pose, lighting, occlusion, facial features, and expression intensity in the wild. Recent CNN–Transformer-based hybrid models like POSTER have leveraged local feature learning, landmark guidance, and global dependency modeling to achieve strong performance. Yet these methods give the main focus to spatial and contextual representations while not really going deep into adaptive channel-wise feature importance over multi-scale representations. As different feature channels represent emotions in varying degrees, it is likely that by treating all feature channels equally, one would limit the ability of the learned features to discriminate effectively. To overcome this weakness, this article presents a ResNet-18–Transformer landmark-guided module called SE-POSTER that fuses lightweight Squeeze-and-Excitation (SE) attention modules into the multi-scale feature pyramid of the baseline POSTER architecture. The proposed method carries out feature channel recalibration adaptively at the level of features before Transformer-based global attention modeling, thus allowing the network to focus on emotionally informative feature channels and suppress less relevant responses. The inclusion of SE attention in the network enhances fine, mid, and global levels of feature representations at a very low cost in terms of computation. On the basis of the RAF-DB, FERPlus, and AffectNet datasets, enormous experiments prove that the SE-POSTER framework proposed is capable of steadily boosting recognition accuracy relative to the baseline POSTER and several state-of-the-art FER methods. Especially, the proposed model delivers 92.78% accuracy on RAF-DB while it also shows better robustness and generalization capability under difficult real-world conditions. Moreover, additional ablation studies reveal that multi-level channel recalibration is effective in improving discriminative emotional feature learning.

Read PDF

Similar papers

Open access Aug 2026

An Attention-Enhanced ConvNeXtTiny Model for Robust Facial Expression Recognition

Facial Expression Recognition (FER) plays an important role in affective computing and human–computer interaction by enabling automated interpretation of human emotional states from facial images. Despite recent advances in deep learning, reliable FER remains challenging because of variations in facial appearance, illu...

Manisha B. Thombare, S. Gumaste · 0 citations
Open access Aug 2026

Transformer With Decoupled Self-Attention Regularization for Age-Unbiased Facial Expression Recognition

A novel bias-mitigation method that decouples and regularizes age- and emotion-related components within the self-attention mechanism of Transformer to reduce age-related bias and enhance age-invariant emotion separation is proposed.

Jaeil Park, Sung-Bae Cho · 0 citations
Sep 2026

GCA-ODN: A global context-aware dropout network for joint facial landmark detection and emotion recognition under occlusion.

We propose the Global Context-Aware Dropout Network (GCA-ODN), a CNN-based, computationally practical neural architecture for joint facial landmark detection (FLD) and facial expression recognition (FER) under partial facial occlusion. GCA-ODN learns a shared embedding that encodes facial geometry and affective cues, i...

Muhammad Sadiq · 0 citations
Open access 2026

Dual-Stream Facial Emotion Recognition with Self-Supervised Pre-Training and Evidential Uncertainty

: Facial emotion recognition (FER) remains difficult in real-world settings. Inter-subject variability, lighting changes, occlusion, and class imbalance all limit performance. Most FER systems rely on one convolutional or transformer backbone. This narrows the features available for classification. This paper presents...

Rashid Jahangir, Nazik Alturki, M. Alreshoodi · 0 citations
Open access Sep 2026

Implementing Facial Emotion Recognition with Attention-Enhanced Feature Learning and Grad-CAM Explainability

Facial Emotion Recognition (FER) is a major study area in computer vision and affective computing because to its many applications in human-computer interaction, healthcare monitoring, intelligent surveillance, education, and behavioural analysis. Existing FER systems struggle with feature discrimination, model interpr...

Amruta Netaji Taur, Vijayshri A. Injamuri · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.