Skip to content
Preprint

Mind the Rift: Cross-Scale Coupling Mismatch for AI-Generated Video Detection

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

RIFT (Representation Inconsistency Forensics on Trajectories), an orthogonal forensic framework that addresses cross-scale coupling mismatch through three interlocking components: a macro stream that builds a dynamic baseline of expected temporal evolution via differential geometry and persistent homology on learned manifold trajectories, a micro stream that acts as a sensitive forensic probe via steganalytic filtering and temporal modeling, and a coupling divergence module that measures the conditional dependency between the two streams.

Abstract

As AI video generators achieve cinematic realism, reliable detection becomes essential for safeguarding digital trust. We identify cross-scale coupling mismatch as a new forensic signal, where scale refers to the level of abstraction (semantic dynamics vs. pixel-level residuals): in natural videos, macro-level temporal dynamics and micro-level residual patterns are intrinsically coupled by the unified imaging physics pipeline, whereas AI generators, whose training objectives do not explicitly preserve this joint distribution, systematically violate this coupling. Detecting such mismatch is challenging because it requires independently extracting information at both scales while simultaneously quantifying their cross-scale relationship. We propose RIFT (Representation Inconsistency Forensics on Trajectories), an orthogonal forensic framework that addresses this through three interlocking components: a macro stream that builds a dynamic baseline of expected temporal evolution via differential geometry and persistent homology on learned manifold trajectories, a micro stream that acts as a sensitive forensic probe via steganalytic filtering and temporal modeling, and a coupling divergence module that measures the conditional dependency between the two streams. Gram-Schmidt orthogonality guarantees the information-theoretic validity of this measurement. Experiments on two benchmarks (VidProM, 120K videos, 7 generators; GenVidBench, 68K videos, 4 generators) demonstrate that RIFT achieves 99.33% and 99.72% F1-score respectively, with 97.87% unseen-generator detection rate in leave-one-out evaluation, while exhibiting encoder agnosticism: scaling from ViT-S/14 (22M) to ViT-L/14 (300M) changes F1 by less than 0.1%, and switching to a different encoder family (DINOv1) reduces F1 by only 0.73 pp. Code is available at https://github.com/Litsay/RIFT

View source

Similar papers

Preprint Sep 2026

Delving into Asymmetric Information Dynamics for High-Fidelity Virtual Try-On

Virtual try-on (VTON) requires precise pixel-level fidelity, yet mainstream Diffusion Transformers (DiTs) often suffer from texture degradation and structural drift. We identify symmetric interactions in standard joint-attention mechanisms as a source of these failures. Although such interactions support semantic flexi...

Zi-Shu Qin, Zhi-Yu Jin, Pi-Pei Huang et al. · 0 citations
Preprint Aug 2026

VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

This work introduces meta-detection into AI-generated video detection, enabling reliable forgery detection by jointly optimizing predicted labels and supporting evidence within reinforcement learning, and introduces evidence-aware credit assignment, which preserves reliable label supervision while encouraging detectors...

Bo-Wei Liu, Zheng Lu, Yuhan Bian et al. · 1 citation
#artificial intelligence Preprint Sep 2026

PIVOT: Physics-Grounded Verification for AI-Generated Audio-Video Detection

As generative models continue to advance, AI-generated content (AIGC) is becoming increasingly realistic, weakening the artifact cues commonly exploited by existing detectors. Nevertheless, faithfully reproducing the physical behavior of real-world events remains challenging for current generators. We therefore explore...

Bo-Yuan Zheng, Kang-Ran Zhao, Xiao-Yu Zhang et al. · 0 citations
Preprint Sep 2026

ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos

This work introduces ManiVid, a unified forensic analysis task covering forgery detection, artifact grounding, and anomaly explanation for manipulated videos, and proposes ManiVidLens, a unified framework for explainable video forgery analysis.

Heng-Rui Kang, Zhong-Hao Yan, Yuxuan Yang et al. · 0 citations
Open access Aug 2026

DBINDS: detection based on initial noise difference sequence from diffusion model inversion for AI-generated videos

DBINDS, a diffusion-model-inversion-based detection framework that extends the analysis from the pixel domain to a diffusion-inversion-derived latent-noise space, is proposed and a composite of spatiotemporal correlation and spatiotemporal texture features is identified as the Best Dual Combination.

Yan-Lin Wu, Xiaogang Yuan, Dezhi An · 0 citations
Conference Open access Sep 2026

Transferable Attacks on Open-Vocabulary Video Instance Segmentation via Dual-Objective Triggers

The Dual-Objective Triggers (DOT) is presented, the first transferable attack on OV-VIS that simultaneously exploits the vision–language coupling and temporal coherence, and Phase-Guided Ad-versarial Training is introduced, which injects perturbations primarily in the phase spectrum while blending amplitudes with clean...

Ming-Hao Shou, Ke-Sen Wang, Tong Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.