Skip to content
Open access

YOLO world guides open world video anomaly detection

You-Xi Li Xiang-Jun Chen Li-Ming Wang Xiao-Chen Huang Qiang Liu
Sep 2026 · Scientific Reports · 0 citations

Abstract

Video anomaly detection is important for safety-critical monitoring, yet deployed systems must recognize anomaly categories absent from training. Existing weakly supervised and vision-language methods emphasize video- or segment-level scores or text–video similarity, limiting object-centric temporal reasoning and direct known-versus-unknown separation. To address these limitations, this paper proposes OW-YW-VAD, an object-centric open-world video anomaly detection framework. YOLO-World provides open-vocabulary object evidence, ByteTrack forms trajectories, and a temporal convolutional encoder and interaction module model motion and context. Known-class confidence, normal-memory distance, regional dynamics, and semantic uncertainty are fused to produce localized normal, known-anomaly, and unknown-anomaly decisions. Experimental results show that OW-YW-VAD obtains 76.10%, 98.60%, and 90.12% on UBnormal, ShanghaiTech, and UCF-Crime, respectively, in independent benchmark-specific Protocol-A real-video experiments. Experimental results show that the five-seed held-out-category real-video experiment yields a known-versus-unknown AUROC of 96.36±0.25%, OSCR of 88.67±0.19%, and unknown-class F1 of 90.50±0.16%. The prompt-role ablation further shows that object prompts dominate the fused decision, while the auxiliary prompt groups do not yet contribute positively under the held-out-category protocol. Experimental results show that the evaluation quantifies both benchmark anomaly detection and direct known-versus-unknown separation under the declared protocols.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.