Skip to content
Preprint

PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMs

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

A Perception-as-Tool paradigm is introduced and instantiated as PATE-Forensics, which architecturally decouples detection and localization from explanation generation while coupling detection and localization as tightly as possible within a forensic perception tool.

Abstract

Existing explainable deepfake forensic methods typically rely on task-adapted MLLM to jointly address detection, localization, and explanation. Inspired by agent-style tool use, we instead introduce a Perception-as-Tool paradigm and instantiate it as PATE-Forensics, which architecturally decouples detection and localization from explanation generation while coupling detection and localization as tightly as possible within a forensic perception tool. The DINOv3-based tool couples a multi-granularity detection module that integrates global, patch-level, and segment-level evidence with a cue-guided localization module by spatializing the patch-level and segment-level evidence into forgery score maps that guide dense mask prediction. The original image and forensic perception outputs produced by the tool form structured forensic context for a general-purpose MLLM, which is guided by prompt constraints to generate explanations without task-specific fine-tuning. On DDL-X Track 3, PATE-Forensics achieves the best official score of 0.89, outperforming the second-ranked team by 0.19 points. Our code is available at https://github.com/yqli00000/PATE-Forensics.

View source

Similar papers

Preprint Sep 2026

Evidence-Guided Detection, Localization and Explanation for Text-Centric Image Forensics

The rapid progress of AIGC has made text-centric image manipulation increasingly accessible, creating new forensic challenges that require not only authenticity detection but also spatial grounding and evidence-based explanation. This paper presents our solution to the GenText-Forensics Challenge at ACM Multimedia 2026...

Pei-Feng Liu, Bin Li, Qing-Song Zhang et al. · 0 citations
Preprint Sep 2026

Agentic Tool-Augmented Reasoning for Explainable Image Forgery Detection

Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large language model (MLLM)-based approaches generate post-hoc explanations of predetermined classification results rather than reasoning from evidence. Inspired by the forensi...

Zhi-Ya Tan, Jing Huang, Chang-Tao Miao et al. · 0 citations
Open access Sep 2026

XAI-DRIVEN VIDEO FORGERY FORENSICS: A UNIFIED EXPLAINABLE AI FRAMEWORK FOR MULTI-TYPE VIDEO MANIPULATION DETECTION

XAI-VFF is presented, a unified Explainable AI framework that combines Gradient-weighted Class Activation Mapping (Grad-CAM), SHapley Additive exPlanations (SHAP), and Local Interpretable Model-Agnostic Explanations (LIME) for video forgery explanations.

S. P, R. Priya · 0 citations
Preprint Sep 2026

Team MSU GenText-Forensics Challenge 2026 Technical Report

Document text forgery has evolved beyond simple pixel-level manipulation: modern attacks alter not only the appearance of a document but also its meaning, and increasingly target the OCR&LLM pipelines that consume such documents. The ACM MM 2026 GenText-Forensics challenge therefore requires systems that not only decid...

Kirill Koltsov, A. Gushchin, D. Vatolin et al. · 0 citations
Preprint Sep 2026

From Detection to Localization: A Unified Forensics Framework for Fully Synthetic and Tampered Images

The rapid advancement of generative models has significantly worsened the problem of manipulated image detection, as these methods are capable of producing highly realistic forgeries, reinforcing the importance of multimedia forensics. Conventional approaches typically frame image manipulation detection as a binary cla...

Annalisa Gallina, M. Fiorucci, Marco Brigo et al. · 0 citations
Preprint Aug 2026

Primitive-Driven Compositional Forensic Visual Prompting for Open-World Face Anti-Spoofing

This work proposes a compositional forensic visual prompt learning framework that operates entirely in the visual feature space and employs patch-aware attention to refine a shared set of learnable micro-forensic primitives into localized forensic evidence units derived from image patches.

Fangling Jiang, Qi Li, Bing Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.