Multimodal DynaST: A Multimodal Framework With Dynamic Sample Selection and Temporal Context Aggregation for Weakly Supervised Video Anomaly Detection
Weakly supervised video anomaly detection (WSVAD) reduces annotation costs by utilizing only video-level labels during training. However, existing methods often suffer from inadequate temporal modeling and noisy supervision, which limit anomaly localization performance. To address these issues, this paper proposes mult...