Cloud-Based Facial Expression Recognition Using Squeeze Vision Transformer
Abstract
Purpose: This paper presents and evaluates a cloud-oriented facial expression recognition (FER) system built on a compact, squeeze-style Vision Transformer, and assesses its suitability as an accurate, resource-efficient modelling layer for cloud and edge-cloud deployment. Design/Methodology/Approach: A pretrained ViT-small-patch16-224 backbone is fine-tuned for seven-class facial expression classification (angry, disgust, fear, happy, neutral, sad, and surprise). Images from three benchmark datasets (CK+, JAFFE, and KDEF) are preprocessed by resizing, converting to grayscale, converting to tensors, normalising, and expanding channels. The model is trained with cross-entropy loss and the Adam optimiser over ten epochs. Performance is assessed using accuracy, loss, per-class precision, recall, and F1-score, macro-averaged metrics, and confusion matrices. Research Limitation: Evaluation relies on controlled laboratory datasets and measures classification performance offline. Generalisation to in-the-wild data and end-to-end cloud-serving latency, throughput, and cost have not yet been tested. Findings: The model attains 97.5% accuracy on CK+ (macro F1 of 0.956), 95.0% on JAFFE (macro F1 of 0.946), and 86.1% on KDEF (macro F1 of 0.857). Accuracy is highest on the tightly controlled CK+ and JAFFE datasets and lower on the more heterogeneous KDEF, with Fear and Disgust the most frequently confused expressions across all three datasets. Practical Implication: The results indicate that compact, squeeze-style vision transformers can act as an accurate and computationally efficient modelling layer for cloud and edge-cloud FER services. They reduce the computational complexity and memory required per inference request, simplifying real-world deployment and scaling. Social Implication: Efficient cloud-based facial expression recognition can support socially beneficial applications such as remote healthcare and mental-health monitoring, education, and assistive human-computer interaction, provided demographic fairness and facial-data privacy receive due attention. Originality/value: The paper contributes a cloud-oriented facial expression recognition system built on a compact, squeeze-style vision transformer, showing how a parameter-efficient transformer backbone can be integrated into a cloud-and-edge pipeline for emotion recognition.