Skip to content
Preprint

Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection

Aug 2026 · 0 citations · 58 references
Computer Science

TL;DR

Defake-o3 is presented, an explainable AIGI detector that moves from speculative rationales to verifiable evidence, and Experiments show that Defake-o3 improves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence.

Abstract

The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often generate speculative rationales: they rely on vague or hallucinated artifacts, miss subtle localized flaws from the latest generators, and fail to provide evidence that can be visually verified. We present Defake-o3, an explainable AIGI detector that moves from speculative rationales to verifiable evidence. It combines interactive visual search with verifier-guided evidence alignment: the model iteratively zooms into suspicious regions to inspect fine-grained details, while an Evidence Verifier, trained from human verification annotations, provides reinforcement learning rewards that favor grounded evidence and penalize baseless claims. To support this objective, we construct GroundFake, a dataset designed for grounded explainable detection, with localized bounding-box evidence, human verification based on visual grounding and artifact specificity, corrected reasoning trajectories, and valid/invalid evidence supervision. We further introduce FakeFrontier, an out-of-distribution benchmark built from real images and outputs of 10 recent generators, together with an MLLM-based protocol for evaluating evidence quality and persuasiveness. Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 improves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence.

View source

Similar papers

Preprint Aug 2026

Grounding and Explaining Visual Evidence for AI-Generated Image Detection in Human-Centric Scenes

Rapid advances in image generation models call for interpretable AI-generated image detection methods that not only determine authenticity but also provide supporting visual evidence. Existing approaches may produce inconsistencies between generated explanations and localized evidence regions, undermining the reliabili...

Kun Guo, Yu-Zhou Yang, Haoyue Wang et al. · 0 citations
Preprint Sep 2026

Look Before You Judge: Training-Free Region Mining for Grounded and Explainable Deepfake Detection

Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evidence. Existing groun...

Chia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Evidence-RL: Towards Evidence-intensive Visual Reasoning

This work proposes Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding, which outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.

Haojie Huang, Xin-Lei Yu, Cheng-Ming Xu et al. · 3 citations
Preprint Aug 2026

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

Feature-robust Augmentation is introduced, which comprises diversified degradation-aware augmentation strategies, and a supervised contrastive learning pattern paired with a mean-teacher architecture that stabilizes features against augmentations through consistency constraints that wins the first place in ACM Multimed...

Zhu Xu, Jia-Qi Tang, Po-Kai Chen et al. · 0 citations
#natural language process... Preprint Sep 2026

MIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal Misinformation

Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. However, such methods do not verify what human fact-checkers...

Rui-Hong Zeng, Jonathan Tonglet, Preslav Nakov et al. · 0 citations
Preprint Aug 2026

VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

This work introduces meta-detection into AI-generated video detection, enabling reliable forgery detection by jointly optimizing predicted labels and supporting evidence within reinforcement learning, and introduces evidence-aware credit assignment, which preserves reliable label supervision while encouraging detectors...

Bo-Wei Liu, Zheng Lu, Yuhan Bian et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.