Aug 2026· International Journal of Multimedia Information Retrieval· Vol 15· 0 citations· 99 references
Computer Science
TL;DR
This work proposes a novel taxonomy categorizing 34 existing language-driven VAD methods based on learning paradigm, core architecture, adaptation strategy, and functional output, and serves as a foundational resource for the VAD community, advancing the understanding and application of language-driven models in addressing complex VAD tasks.
This work proposes a novel Text-Driven Video Anomaly Detection (TD-VAD) approach, which utilizes video-like text descriptions with temporal characteristics generated by LLM to train a VAD model, without any reliance on target-domain anomaly data.
Shuang-Qing Zhang, Lei-Lei Ma, Zhao Wang et al.· 1 citation
In this paper, we propose MuST-VAD, a mutual structured learning framework for weakly supervised video anomaly detection (VAD) in which an anomaly detector and a large vision-language model (LVLM) exchange their acquired knowledge. Detectors in weakly supervised VAD learn anomaly scores from features extracted by a fixed, task-agnostic backbone. These fixed features bound the achievable detection accuracy. Recent methods therefore transfer LVLM semantics into the detector as richer features. However, this transfer is one-way: what the detector learns about the target videos never returns to the LVLM. MuST-VAD extends the one-way transfer into a bidirectional learning loop. In this loop, the latest detector predictions supervise the LVLM adaptation, and the adapted LVLM returns updated representations that retrain the detector; the two models alternate these updates over small video groups. Both models train on detector-selected key clips, while confidence weighting and annotation-anchored question answering keep the exchanged supervision reliable. On UCF-Crime, our mutual learning improves the one-pass transfer baseline from 88.15% to 88.63% AUROC and from 37.25% to 42.46% average precision (AP), outperforming the state-of-the-art method in AP by 4.13 points.
Satoshi Hashimoto, Hitoshi Nishimura, Mori Kurokawa· 0 citations
Video anomaly detection is important for safety-critical monitoring, yet deployed systems must recognize anomaly categories absent from training. Existing weakly supervised and vision-language methods emphasize video- or segment-level scores or text–video similarity, limiting object-centric temporal reasoning and direct known-versus-unknown separation. To address these limitations, this paper proposes OW-YW-VAD, an object-centric open-world video anomaly detection framework. YOLO-World provides open-vocabulary object evidence, ByteTrack forms trajectories, and a temporal convolutional encoder and interaction module model motion and context. Known-class confidence, normal-memory distance, regional dynamics, and semantic uncertainty are fused to produce localized normal, known-anomaly, and unknown-anomaly decisions. Experimental results show that OW-YW-VAD obtains 76.10%, 98.60%, and 90.12% on UBnormal, ShanghaiTech, and UCF-Crime, respectively, in independent benchmark-specific Protocol-A real-video experiments. Experimental results show that the five-seed held-out-category real-video experiment yields a known-versus-unknown AUROC of 96.36±0.25%, OSCR of 88.67±0.19%, and unknown-class F1 of 90.50±0.16%. The prompt-role ablation further shows that object prompts dominate the fused decision, while the auxiliary prompt groups do not yet contribute positively under the held-out-category protocol. Experimental results show that the evaluation quantifies both benchmark anomaly detection and direct known-versus-unknown separation under the declared protocols.
You-Xi Li, Xiang-Jun Chen, Li-Ming Wang et al.· Scientific Reports· 0 citations
An in-depth survey of fifteen state-of-art methodologies including classical CNN models, temporal-spatial video recognition, transformer-based networks, explainable AI (XAI) models, and models that combine multimodal large language model (LLM) products are provided.
Shavnam Shavnam, Neha Dhiman· International Journal of Inn...· 0 citations
Video anomaly detection (VAD) is critical for automation systems and security surveillance. Recently, multimodal vision–language models (MLLMs) have attracted increasing attention due to their rich pre-trained knowledge and strong explainability. However, existing MLLM-based approaches struggle to adapt to real-world settings where anomaly definitions are complex: they either rely on the model’s built-in knowledge and mainly capture only generic anomalies, or require anomalous samples for supervised fine-tuning—which are often rare and may raise legal or privacy concerns. To address this challenge, we propose a Sparsity-Controllable Vision-Language Model (SCVLM) for scenario-related anomaly detection. SCVLM learns normality from unlabeled normal data by jointly reconstructing multimodal representations and summarizing textual descriptions, thus enabling anomalies to be detected as deviations from the learned normal patterns. We introduce a Sparsity-Controllable Memory Block (SCMB) to improve memory addressing mechanism for pretrained multimodal representations. During inference, anomalies are detected by fusing two modality-specific reconstruction-errors with an LLM-based textual anomaly scoring mechanism. Meanwhile, we fuse the interpretable cues from each detection branch to derive anomaly reasoning consistent with human commonsense. Extensive experiments on challenging benchmarks demonstrate that SCVLM achieves state-of-the-art detection performance. Our code is available at https://github.com/SCVLM/SCVLM
Jiangyun Chen, Yuanjie Dang, Peng Chen et al.· IEEE Transactions on Informa...· 0 citations
Video Anomaly Detection (VAD) aims to identify abnormal frames from discrete events within video sequences. Existing VAD methods suffer from heavy annotation burdens in fully-supervised paradigm, insensitivity to subtle anomalies in semi-supervised paradigm, and vulnerability to noise in weakly-supervised paradigm. To address these limitations, we propose a novel paradigm: Single-Frame supervised VAD (SF-VAD), which uses a single annotated abnormal frame per abnormal video. SF-VAD ensures annotation efficiency while offering precise anomaly reference, facilitating robust anomaly modeling, and enhancing the detection of subtle anomalies in complex visual contexts. To validate its effectiveness, we construct three SF-VAD benchmarks by manually re-annotating the ShanghaiTech, UCF-Crime, and XD-Violence datasets in a practical procedure. Further, we devise Frame-guided Progressive Learning (FPL), to generalize sparse frame supervision to event-level anomaly understanding. FPL first leverages evidential learning to estimate anomaly relevance guided by annotated frames. Then it extends anomaly supervision by mining discrete abnormal events based on anomaly relevance and feature similarity. Meanwhile, FPL decouples normal patterns by isolating distinct normal frames outside abnormal events, reducing false alarms. Extensive experiments show SF-VAD achieves state-of-the-art detection results while offering a favorable trade-off between performance and annotation cost. The benchmarks and code are available at https://github.com/Junxi-Chen/SF-VAD .
Junxi Chen, Liang Li, Yunbin Tu et al.· Neural Information Processin...· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.