Skip to content
Open access

A Hybrid Vision Transformer and EfficientNet-B3 Framework for Facial Expression Recognition

Aug 2026 · Journal of Imaging · Vol 12, pp. 360 · 0 citations · 43 references
Medicine

TL;DR

A hybrid architecture that combines Vision Transformers (ViTs) to capture global context with EfficientNet-B3 for multi-scale feature extraction and highlights the promise of hybrid deep learning architectures in tackling real-world facial expression recognition challenges.

Abstract

Facial expression recognition technology is vital for security, verification, and personalization, but it faces challenges due to variations in scale, illumination, occlusion, and facial expressions. This paper presents a hybrid architecture that combines Vision Transformers (ViTs) to capture global context with EfficientNet-B3 for multi-scale feature extraction. Unlike simple concatenation, our approach projects the ViT’s [CLS] token and the EfficientNet’s global pooling features into a shared 512-dimensional space before merging, enabling better alignment of global and local features. When tested on the FERPlus dataset, it reaches an accuracy of 94.4 ± 0.3%, surpassing several recent methods, notably existing transformer- and CNN-based methods. Ablation studies show each component’s contribution, with the full model outperforming the no-fusion version by 2.6%. With around 98 million parameters and an inference time of ~23 ms per image, it balances efficiency and high performance, suitable for real-time use on suitable hardware. Evaluation via confusion matrix, t-SNE visualization, and comparisons with recent techniques such as HLA-ViT (90.13%), AU-ViT (90.15%), and CCFER (91.24%) demonstrates its robustness and discriminative feature learning. This work highlights the promise of hybrid deep learning architectures in tackling real-world facial expression recognition challenges.

Read PDF

Similar papers

Open access Sep 2026

Cloud-Based Facial Expression Recognition Using Squeeze Vision Transformer

Purpose: This paper presents and evaluates a cloud-oriented facial expression recognition (FER) system built on a compact, squeeze-style Vision Transformer, and assesses its suitability as an accurate, resource-efficient modelling layer for cloud and edge-cloud deployment. Design/Methodology/Approach: A pretrained ViT-...

S. Al-darraji, A. J. Jalil · 0 citations

HighTech and Innovation

Dalya Alkhafaji, W. Teahan · 0 citations
Aug 2026

Hybrid Deep Learning Model for Fake Image Detection with Advanced Face Object Segmentation Method

The fast development of deepfake technologies for image generation produces more and more realistic manipulated facial images that are harder to distinguish from the real content. In this paper, we propose a novel Hybrid CNN–Vision Transformer (HybCNNViT) framework for robust deepfake image detection by combining discr...

Saurabh Kumar Jain, Mohd. Akbar · 0 citations
Open access Sep 2026

Hybrid Deep Learning System for Robust and Scalable Face Authentication

A biometric face identification system plays a crucial role in authentication system, requiring accuracy and computational efficiency, particularly in resource-constrained environments. The objective of this study is to develop a lightweight deep learning framework for facial authentication that achieves high performan...

Dalya Alkhafaji, W. Teahan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.