Skip to content
Open access

Real-Time Anomaly Detection on Edge Devices via VLM Prompt Optimization

Jul 2026 · Electronics · 0 citations · 12 references

Abstract

Real-time video anomaly detection (VAD) under realistic edge constraints—sub-second latency, ≤25 W power, no cloud dependency, and human-interpretable output—remains an open problem. Existing lightweight video convolutional neural networks (X3D, MoViNets) are bound to closed-set training distributions, while recent vision–language-model-based VAD methods (LAVAD, VERA, Holmes-VAD) achieve 80–89% area under the curve (AUC) but rely on datacenter-grade GPUs and Chain-of-Thought (CoT) reasoning that pushes per-segment latency well above one second. This paper reframes the design target from peak accuracy to practical edge deployability and contributes two tightly coupled designs: (i) an edge-optimized inference stack that compresses Qwen3-VL-2B with 4-bit Activation-aware Weight Quantization (INT4 AWQ) and serves it through a TensorRT-LLM C++ runtime on NVIDIA Jetson Orin NX (16 GB, 25 W); and (ii) a fully automatic, CoT-free verbalized prompt optimization in which an 8B optimizer iteratively refines a natural-language definition block Dt using class-balanced (stratified) development batches on a disjoint development subset, with no human editing and no runtime cost on the edge device. Three findings support this framing: (a) the inference stack reduces per-segment latency to 0.25 s, a 7.4× speed-up and 55% memory reduction over a Python/PyTorch baseline; (b) verbalized prompt optimization improves zero-shot AUC from 71.82% (manual prompt) to 76.39%, outperforming GPT-4- and Gemini-Pro-generated prompts (74.12% and 74.35%) under the same edge backbone; and (c) single-frame input attains the highest mean AUC among one-, five-, and eight-frame windows—statistically comparable to the five-frame setting—while offering the lowest latency, making it the preferred operating point under the edge budget. While the absolute AUC (76.39%) is below recent server-side methods (CLIP-TSA 87.58%, VadCLIP 88.02%, Holmes-VAD 89.51%), our framework is the only one in this comparison that operates entirely on a ≤25 W edge device, providing a deployment-oriented operating point on the accuracy–feasibility frontier of VLM-based VAD.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.