Skip to content

G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement

Jul 2026 · arXiv.org · Vol abs/2607.04607 · 0 citations · 63 references
Computer Science

TL;DR

G2VD introduces a counterfactual intervention pipeline (CFIPipeline) that constructs counterfactual samples through VAE-based reconstruction and subsequent frequency-domain and pixel-domain alignment, thereby weakening spurious correlations between domain-specific bias and authenticity labels and design a causal disentanglement classifier.

Abstract

Rapid advances in AI video generation pose increasing security risks and call for reliable detectors with strong cross-domain generalization. Although existing methods perform well under in-domain evaluation, their performance degrades substantially on unseen generators. A key reason is shortcut learning, where detectors rely on domain-specific bias rather than intrinsic forensic cues. To address this issue, we propose G2VD, a generalizable AI-generated video detection framework based on counterfactual intervention and causal disentanglement. First, G2VD introduces a counterfactual intervention pipeline (CFIPipeline) that constructs counterfactual samples through VAE-based reconstruction and subsequent frequency-domain and pixel-domain alignment, thereby weakening spurious correlations between domain-specific bias and authenticity labels. Building on this intervention, we further design a causal disentanglement classifier that combines two domain-anchored branches with complementary objectives and a constraint based on the Hilbert-Schmidt Independence Criterion (HSIC), encouraging the causal and non-causal representations to capture intrinsic forensic cues and domain-specific bias, respectively. Experiments across four public datasets demonstrate strong cross-domain performance and consistent gains over baseline methods. In the challenging GenVidBench setting, G2VD achieves over 90\% overall ACC, with improvements of 0.194 in F1 and 0.104 in AUC over comparable state-of-the-art methods, while using only 10\% of the available training data. Code is available at https://github.com/DMOSCAR-98/G2VD.

View source

Similar papers

Preprint Aug 2026

VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

This work introduces meta-detection into AI-generated video detection, enabling reliable forgery detection by jointly optimizing predicted labels and supporting evidence within reinforcement learning, and introduces evidence-aware credit assignment, which preserves reliable label supervision while encouraging detectors...

Bo-Wei Liu, Zheng Lu, Yuhan Bian et al. · 1 citation
Preprint Sep 2026

Mind the Rift: Cross-Scale Coupling Mismatch for AI-Generated Video Detection

RIFT (Representation Inconsistency Forensics on Trajectories), an orthogonal forensic framework that addresses cross-scale coupling mismatch through three interlocking components: a macro stream that builds a dynamic baseline of expected temporal evolution via differential geometry and persistent homology on learned ma...

Siyu Li, Jin Yang, Weiheng Liang · 0 citations
Preprint Aug 2026

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

Feature-robust Augmentation is introduced, which comprises diversified degradation-aware augmentation strategies, and a supervised contrastive learning pattern paired with a mean-teacher architecture that stabilizes features against augmentations through consistency constraints that wins the first place in ACM Multimed...

Zhu Xu, Jia-Qi Tang, Po-Kai Chen et al. · 0 citations
Conference Open access Sep 2026

PURE: Purging Unrelated Representations for Content-Agnostic Forgery Detection

This work proposes PURE (Purging Unrelated Representations for Content-Agnostic Forgery Detection), which achieves content-agnostic detection through two complementary components: a Causal Semantic Generative (CSG) mechanism that disentangles semantic representations from forgery-irrelevant nuisance factors, and a Gaus...

Xin-Yu Wu, Dong Li, Minglai Shao et al. · 0 citations
Preprint Aug 2026

Environment-Invariant Subspace Learning for Generalizable Deepfake Detection

This work proposes an innovative Environment-Invariant Subspace Learning (EISL) framework, which aims to disentangle features into orthogonal forgery-relevant invariant factors and environment-related residual factors via a learnable low-rank projection and designs an Environmental Intervention module that generates di...

Sheng-Hao Chen, Hao Jia, Chen Li et al. · 0 citations
Preprint Aug 2026

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

This paper presents FADE, an effective training framework for counterfactual discovery and explanation that is built on an evidence-first, two-stage training paradigm, and demonstrates remarkable robustness when transitioning from constrained MCQs to unconstrained OQA and captioning.

Fufangchen Zhao, Jin-Hu Fu, Jiachen Lei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.