UAPE-CLIP: Uncertainty-Aware prompt learning with vision–language alignment for student engagement recognition
Abstract
Student engagement recognition underpins the understanding of learning behaviors, early warning of learning risks, and the design of instructional interventions. Most existing approaches primarily model visual spatiotemporal cues and make limited use of semantic priors encoded in label text, which can lead to brittle performance when cues are weak, states co-occur, or the deployment domain shifts. To address these challenges, we propose UAPE-CLIP (Uncertainty-Aware Prompt Learning for Student Engagement Prediction via Vision-Language Alignment). UAPE-CLIP leverages cross-modal semantic supervision during training while adopting a decoupled video-only inference pathway to avoid reliance on label text at deployment. We introduce an uncertainty-aware task-adaptive prompt learning mechanism that scales soft prompts according to prototype uncertainty, stabilizing state semantics and improving video–text alignment. To mitigate pseudo-negative interference caused by repeated semantic prototypes in multi-label settings, we employ a multi-positive contrastive objective. Multi-grained cross-modal matching is further combined with temporal modeling to capture both global behavioral patterns and key-frame evidence. Experiments on DAiSEE and AFEW show that UAPE-CLIP achieves 70.63% accuracy on DAiSEE (a +3.22-point gain over the strongest baseline) and 64.75% on AFEW (a +2.04-point gain over the best video-only contrastive method). Overall, UAPE-CLIP offers a practical balance among accuracy, computational cost, and deployment consistency for engagement analytics.