Multimodal DynaST: A Multimodal Framework With Dynamic Sample Selection and Temporal Context Aggregation for Weakly Supervised Video Anomaly Detection
Abstract
Weakly supervised video anomaly detection (WSVAD) reduces annotation costs by utilizing only video-level labels during training. However, existing methods often suffer from inadequate temporal modeling and noisy supervision, which limit anomaly localization performance. To address these issues, this paper proposes multimodal DynaST, a multimodal framework for weakly supervised video anomaly detection. A Temporal Context Aggregation Module (TCAM) is designed to capture both global temporal dependencies and local contextual consistency through a global–local attention mechanism. Meanwhile, a Dynamic Sample Selection (DSS) strategy progressively identifies reliable training segments according to prediction confidence and temporal stability, reducing the influence of noisy supervision. In addition, a Feature Alignment Mechanism (FAM) aligns visual representations with semantic prototypes to enhance feature discriminability. An auxiliary audio branch is further introduced to provide complementary acoustic cues through decision-level fusion. Experimental results on UCF-Crime, ShanghaiTech, and XD-Violence demonstrate the effectiveness of the proposed method. multimodal DynaST achieves frame-level AUCs of 86.97% and 98.25% on UCF-Crime and ShanghaiTech, respectively, and obtains an AP of 79.29% on XD-Violence. Ablation studies further confirm the effectiveness of each component and the benefit of multimodal learning for robust anomaly detection.