Skip to content
Open access

Multimodal DynaST: A Multimodal Framework With Dynamic Sample Selection and Temporal Context Aggregation for Weakly Supervised Video Anomaly Detection

2026 · IEEE Access · Vol 14, pp. 144202-144217 · 0 citations · 37 references

Abstract

Weakly supervised video anomaly detection (WSVAD) reduces annotation costs by utilizing only video-level labels during training. However, existing methods often suffer from inadequate temporal modeling and noisy supervision, which limit anomaly localization performance. To address these issues, this paper proposes multimodal DynaST, a multimodal framework for weakly supervised video anomaly detection. A Temporal Context Aggregation Module (TCAM) is designed to capture both global temporal dependencies and local contextual consistency through a global–local attention mechanism. Meanwhile, a Dynamic Sample Selection (DSS) strategy progressively identifies reliable training segments according to prediction confidence and temporal stability, reducing the influence of noisy supervision. In addition, a Feature Alignment Mechanism (FAM) aligns visual representations with semantic prototypes to enhance feature discriminability. An auxiliary audio branch is further introduced to provide complementary acoustic cues through decision-level fusion. Experimental results on UCF-Crime, ShanghaiTech, and XD-Violence demonstrate the effectiveness of the proposed method. multimodal DynaST achieves frame-level AUCs of 86.97% and 98.25% on UCF-Crime and ShanghaiTech, respectively, and obtains an AP of 79.29% on XD-Violence. Ablation studies further confirm the effectiveness of each component and the benefit of multimodal learning for robust anomaly detection.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.