Skip to content

TASG-VAD: Weakly Supervised Video Anomaly Detection via Temporal Variation Attention and Adaptive Saliency Guidance

Aug 2026 · International journal of pattern recognition and artificial intelligence · 0 citations

TL;DR

To alleviate the model’s reliance on dominant anomalous segments, TASG-VAD introduces an Adaptive Saliency Guidance (ASG) strategy, which performs intra-video saliency ranking and dynamic masking to guide the model toward overlooked subtle anomalies.

Abstract

Weakly supervised video anomaly detection (WSVAD) is important in intelligent surveillance. Existing methods often overemphasize salient abnormal segments, overlook subtle clues, and model temporal dependencies ineffectively. To address these issues, we propose TASG-VAD, an efficient anomaly detection framework. The proposed method is developed along two main directions: dual-scale temporal modeling and subtle anomaly discovery. Specifically, Temporal Variation Attention (TVA) amplifies anomaly related dynamic changes through second-order temporal differences while suppressing interference from static backgrounds. In addition, the Dual-Scale Temporal Encoder (DSTE) combines a dual-branch structure, a parameter-free attention mechanism, and dual-scale temporal convolutions to simultaneously capture local fine-grained fluctuations and long range global dependencies. Furthermore, to alleviate the model’s reliance on dominant anomalous segments, TASG-VAD introduces an Adaptive Saliency Guidance (ASG) strategy, which performs intra-video saliency ranking and dynamic masking to guide the model toward overlooked subtle anomalies. Experimental results show that TASG-VAD achieves AUCs of 88.21% and 98.38% on UCF-Crime and ShanghaiTech, respectively, and an AP of 84.60% on XD-Violence. With only about 1.7M parameters, it maintains high accuracy and excellent inference efficiency, and significantly outperforms existing methods.

View source

Similar papers

Sep 2026

Spatio-temporal video anomaly detection via CNN-ViT autoencoder and farneback optical flow

This research introduces a novel memory-augmented CNN-ConvViT autoencoder framework for unsupervised video anomaly detection and introduces a Temporal-Aware Prototype Memory Module (TAPMM) that explicitly learns normal spatio-temporal behavior patterns.

Vandana Pathak, Manoj Diwakar, Neeraj Kumar Pandey et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection

An adaptive temporal modeling framework for WSVAD that explicitly accounts for variations in video dynamics across multiple temporal granularities is proposed and an adaptive similarity-based fusion strategy that dynamically integrates anomaly scores into video-level predictions is proposed, replacing fixed top-k aggre...

Chang-Yi Li, Yu Xiao · 0 citations
Open access 2026

Multimodal DynaST: A Multimodal Framework With Dynamic Sample Selection and Temporal Context Aggregation for Weakly Supervised Video Anomaly Detection

Weakly supervised video anomaly detection (WSVAD) reduces annotation costs by utilizing only video-level labels during training. However, existing methods often suffer from inadequate temporal modeling and noisy supervision, which limit anomaly localization performance. To address these issues, this paper proposes mult...

Wen-Jian Zhang, Xin-Xin Fan, Qi-Wei Pan et al. · 0 citations
2026

Reconstructive Visual Tuning for Weakly Supervised Video Anomaly Detection

Weakly supervised video anomaly detection (WS-VAD) presents a significant challenge in security video surveillance, as it aims to accurately identify anomaly frames in untrimmed videos with only video-level supervision. Several recent studies exploit vision-language pre-training models, e.g., CLIP, to take advantage of...

Shuang-Qing Zhang, Wei Xu, Yu-Qi Fang et al. · 2 citations
Open access Sep 2026

Learn the interactions: Weakly supervised video anomaly detection with human-object interactions

Weakly supervised video anomaly detection remains a challenging problem, primarily due to the scarcity of abnormal training samples and the lack of diverse feature representations, which hamper the learning of discriminative models. To address these issues, we introduce a novel weakly supervised cross-domain framework...

Mao-Wen Zhou, Erma Rahayu Mohd Faizal Abdullah, Aznul Qalid Md Sabri et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.