Skip to content

PLGSA-Transformer: Periocular Landmark-Guided Attention with Occlusion-Adaptive Cosine Thresholding for Cross-Modal Masked and Unmasked Face Recognition

Jul 2026 · arXiv.org · Vol abs/2607.03581 · 0 citations · 31 references
Computer Science

TL;DR

Results confirm that encoding periocular geometry into attention, with Transformer modelling and occlusion-adaptive thresholds, yields a robust, scalable solution for cross-modal masked face recognition.

Abstract

The widespread adoption of facial masks, accelerated by COVID-19 and mandated in security-sensitive settings, has exposed limitations of conventional face recognition systems. Existing approaches relying on fixed cosine thresholds, non-adaptive CNNs, and purely data-driven features fail to generalize when facial regions are occluded, creating a gap between lab performance and real-world deployability. This paper proposes PLGSA-Transformer, a cross-modal face matching framework with three contributions. First, Periocular Landmark-Guided Spatial Attention (PLGSA) uses MediaPipe landmarks to compute Gaussian heatmaps over the eye, brow, and forehead regions, fusing them with EfficientNetB3 features via a learnable residual gate to direct attention toward discriminative visible regions. Second, a Hybrid CNN-Transformer Branch reshapes feature maps into tokens processed by a two-layer Multi-Head Self-Attention encoder, enabling cross-regional dependency modelling. Third, the Occlusion-Adaptive Cosine Threshold (OACT) is a jointly trained head that raises the matching threshold in proportion to predicted occlusion severity. The model is evaluated on 858 images from Zenodo MDMFR (60%), Kaggle CelebA-HQ masked collection (25%), and author-collected images (15%), spanning both genders, ages 21-75, with varied mask types, trained via a unified loss combining contrastive verification, identity classification, and occlusion cross-entropy. PLGSA-Transformer achieves 97.22% pair verification accuracy with ROC AUC 1.0000, surpassing VGG-16-based MUFM (Abdullah et al., 2025; 95.0%), HOG classifiers (Adnan et al., 2020; 85.0%), and Feature-based Structural Measure (Shnain et al., 2017; 86.61%). These results confirm that encoding periocular geometry into attention, with Transformer modelling and occlusion-adaptive thresholds, yields a robust, scalable solution for cross-modal masked face recognition.

View source

Similar papers

Open access Sep 2026

Pyramid-guided multi-scale self-attention and channel–spatial refinement for occlusion-robust face recognition

A pyramid-guided multi-scale attention framework based on scale alignment and reliability-aware feature refinement improves occlusion robustness without sacrificing clean-face recognition performance, indicating its practical potential for identity verification and access-control applications involving masks, glasses,...

Qi-Nan Zhu · 0 citations
Conference Jul 2026

RCF-Net: Degradation-Aware Hybrid CNN–Transformer for Child Face Identification in Surveillance

Child face identification from surveillance video remains difficult because facial crops are frequently low-resolution, blurred, partially occluded, and captured under unstable illumination. Age-related facial variation further increases the difficulty of maintaining discriminative identity embeddings for children. Thi...

R. Arora, Akash Pandey, Navjeet Kaur · 0 citations
Open access Jul 2026

SE-POSTER: Channel-Enhanced Landmark Guided Transformer for Facial Emotion Recognition

Recognizing facial emotions automatically from images/videos (FER) still represents a difficult problem for emotion computing, mainly due to variations in the face pose, lighting, occlusion, facial features, and expression intensity in the wild. Recent CNN–Transformer-based hybrid models like POSTER have leveraged loca...

Alpamis Kutlimuratov, K. Sharipov, Piratdin Allayarov et al. · 0 citations
Open access Aug 2026

Geometry and mask aware vision transformer for masked face recognition in unconstrained scenarios

Facial identity identification in unrestricted real-world environments may benefit from this model, which performs well in identifying and verifying low-quality and cross-pose masked faces and outperforming the various state-of-the-art methods and previously proposed methods.

P. Kaur, Taqdir Kaur, Sahezpreet Singh · 0 citations
Conference Jul 2026

A Hybrid CNN–Transformer Network for Robust Masked and Occluded Face Recognition in Smart Surveillance Systems

Face recognition systems applied to smart surveillance settings often experience poor performance when the faces are partially occluded by a mask or other objects. Occlusions eliminate critical facial information, which makes face identification much more difficult for traditional deep learning models. To solve this is...

R. R, Anbalagan E · 0 citations
Open access Aug 2026

A Hybrid Vision Transformer and EfficientNet-B3 Framework for Facial Expression Recognition

A hybrid architecture that combines Vision Transformers (ViTs) to capture global context with EfficientNet-B3 for multi-scale feature extraction and highlights the promise of hybrid deep learning architectures in tackling real-world facial expression recognition challenges.

Sasan Karamizadeh, Saman Shojae Chaeikar, Mazdak Zamani · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.