Skip to content

Emotionnet: A Deep Learning Framework for Real-Time Video-Based Facial Emotion Recognition

Sep 2026 · Natural Resources for Human Health · 0 citations · 20 references

Abstract

EmotionNet is a neural network that is used to recognize facial emotions in real-time and under dynamic and unconstrained conditions using videos. The suggested system combines both spatial features extraction and time modeling to ensure the capture of facial appearance and motion-based emotion features. A convoluted neural network (CNN) backbone that is lightweight is used to extract discriminative spatial characteristics of successive video frames and a bidirectional long short-term memory (Bi-LSTM) network is used to learn the time related characteristics of frame sequences. In order to increase resistance to changes in illumination, face alignment, and pose, the framework will include data augmentation, face alignment, and attention-based feature refinement. EmotionNet is designed to be used in the real-time surveillance systems and with edge devices, which is optimized to react to the low-latency inference. The model is trained and tested with benchmark facial emotion datasets through a multi-class classification model of the emotions of happiness, sadness, anger, fear, surprise, disgust, and neutrality. It is proven by experimental success that it performs better than standard CNN and single recurrent models with better accuracy, macro F1-score and temporal stability with less computational cost. Moreover, the role of temporal modeling and attention mechanisms in the improvement of recognition is justified by ablation experiments.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.