Multimodal Deep Learning for Context-Aware Engagement Estimation in Online Learning
Abstract
Student engagement estimation is a key challenge in online education, where instructors have limited access to the behavioral cues available in traditional classrooms. This paper presents our approach to the Context-Aware Student Engagement Detection (CASED) Challenge, which aims to estimate continuous engagement from multimodal recordings of online lectures. We propose a novel multimodal architecture that integrates student and instructor behavioral cues with audio streams, lecture-slide context, and personality metadata. Our approach demonstrates strong predictive capabilities, achieving a Mean Squared Error (MSE) of 0.03 and a Mean Absolute Error (MAE) of 0.14. Ablation experiments show that contextual information improves engagement estimation compared with student-centered behavioral cues alone. Despite these architectural optimizations, evaluating the model on unseen data underscores the challenge of predicting engagement using only 10-second temporal windows. Additionally, the absence of data on the student’s workspace, including screen positioning and user perspective, limits the ability to predict complex behavioral dynamics, which is fundamental for engagement prediction.