Skip to content

Detect Early, Escalate Rarely: Anytime Detection of AI-Generated Video from the Compressed Bitstream

Jul 2026 · arXiv.org · Vol abs/2607.19476 · 0 citations · 39 references
Computer Science

TL;DR

Detectors for AI-generated video are evaluated offline and recast the task as streaming perception and score the motion field the codec already wrote into the bitstream as a parse, not a pixel-domain forward pass.

Abstract

Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasingly by a large vision-language model. Detection, however, is deployed online. We recast the task as streaming perception and score the motion field the codec already wrote into the bitstream. Reading that field is a parse, not a pixel-domain forward pass. Because the running aggregate is monotone, one end-calibrated threshold is anytime-valid at the data-dependent decision time. Recalibrating at each prefix is not. Escalation is priced in closed form. A compute budget maps to a deferral window, on a frontier monotone exactly where the deferral condition holds. On matched GenVidBench the codec stage reaches full-length AUC 0.64 at five orders of magnitude less compute than a pixel CNN, on CPU. Its gate holds the stopping-time false-positive rate at target while the real data match its calibration, and drifts above it under distribution shift. Deferring 15% of clips lifts accuracy from 0.75 to 0.78 at $7\times$ less compute (paired: McNemar $p<10^{-6}$). The stage-1 ordering replicates on AIGVDBench. We introduce no new detector. The contribution is the reframing, two guarantees, and the measured frontiers. Code, configurations, and evaluation splits: https://github.com/KurbanIntelligenceLab/streamdet.

View source

Similar papers

Open access Sep 2026

Trained Still Wins: Narrowing the Gap to Zero-Shot Video Anomaly Detection

Video anomaly detection has two very different answers to the question of how a detector should acquire its notion of “normal” for a given camera: adapt its parameters to hours of normal footage recorded by that exact camera, or perform no target-scene adaptation at all and rely on frozen, heavily pretrained backbones—...

P. Kanwal, Shylaja S. S., Prasad B. Honnavalli · 0 citations
Preprint Jul 2026

AI or Real: Detecting Partially Altered Videos Under Resource-Constrained Environments

The proliferation of generative video models has shifted the practical detection threat from fully fabricated clips to partially manipulated footages. Although modern detectors achieve strong accuracy using foundation backbones of 400M+ parameters, their resource footprint precludes edge deployment. In this paper, we p...

Tamoghna Chakraborty, Md. Nurul Absur, Sourya Saha et al. · 0 citations
#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a co...

Killian Steunou, Yannis Tevissen, M. E. El Yacoubi · 0 citations
#computer vision Preprint Sep 2026

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

ShallowStream is a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building and achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x an...

Jitai Hao, Ke-Shuai Yang, Qiang Huang et al. · 0 citations
Open access Jul 2026

Real-Time Anomaly Detection on Edge Devices via VLM Prompt Optimization

Real-time video anomaly detection (VAD) under realistic edge constraints—sub-second latency, ≤25 W power, no cloud dependency, and human-interpretable output—remains an open problem. Existing lightweight video convolutional neural networks (X3D, MoViNets) are bound to closed-set training distributions, while recent vis...

Sungmin Yu, Jongwon Moon, Hosub Yoon · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.