Skip to content

LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

Jul 2026 · arXiv.org · Vol abs/2607.25962 · 1 citation · 49 references
Computer Science

TL;DR

LaP-Forensics is presented, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence that supports the utility of the residual stream under the evaluated settings, while free-form textual faithfulness and reliability under post-processing remain open limitations.

Abstract

Recent generative models can produce images with few obvious visual artifacts, weakening detectors and explanations that rely only on surface appearance. We present LaP-Forensics, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence. A frozen Stable Diffusion DDIM inversion-reconstruction model provides a fixed reconstruction reference, and its residual map measures local compatibility with that reference. Independent projectors encode the RGB image and residual map before a structured Where-What-Why model predicts a textual analysis and an artifact mask.Supervised fine-tuning is followed by Group Relative Policy Optimization (GRPO), whose reward combines mask overlap with output-structure and evidence-reference terms. These text-side terms encourage the model to refer to the consistency map but do not constitute a verifier of free-form textual truth. A separate image-level head fuses RGB and DDIM-residual class features. Experiments show cross-generator detection on UniversalFakeDetect and competitive artifact localization on the official SynthScars benchmark. Controlled cue-construction, inversion-horizon, component, reward-term, and counterfactual analyses support the utility of the residual stream under the evaluated settings, while free-form textual faithfulness and reliability under post-processing remain open limitations.

View source

Similar papers

Conference Aug 2026

Semantic-Conditioned Forensic Consistency Model for Open-World Image Authenticity Verification

The recent upsurge in the development of sophisticated generative models has significantly improved the visual realism and semantic coherence of synthetic images, thus presenting a major challenge to the field of multimedia forensics. The conventional approaches often rely on either artifacts or high level semantic cues, limiting their robustness when handling images/videos generated by models that were previously unseen. The proposed work addresses this problem by developing a novel forensic consistency learning framework that is conditioned on semantic content. Specifically, the model leverages pretrained DINO (self-DIstillation with NO labels) encoder for visual content, forensic features capturing latent acquisition characteristics and image-level CLIP (Contrastive Language-Image Pretraining) features for global semantic content. A forensic predictor module estimates the expected forensic features conditioned on semantic information, enabling capture of inconsistencies between visual content and underlying artifacts. Additionally, a patch-level anomaly score module enables robust image-level prediction. The method was evaluated under cross-generator setting, by training on GenImage dataset augmented with ProGAN, and evaluated on UniversalFakeDetect (UFD) benchmark. The model achieves 91.5% Area Under ROC Curve, 92.2% Average Precision and 84.1% classification accuracy, significantly outperforming existing baselines like UFD and ResNet (upto 6-10% improvement). Extensive ablations further validate the effectiveness of proposed semantic-conditioned forensic modeling for open-world image authentication.

Venkata Satya Renuka Devi Bhamidipati, Srinivasa Rao Chanamallu, Sudheer Gopinathan · 0 citations
Preprint Aug 2026

FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and Localization

FUSED combines low-level forensic cues with high-level semantic features using a sparsely-gated Mixture-of-Experts architecture, enabling the model to adaptively prioritize the most relevant signal for each token.

Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska · 0 citations
Preprint Aug 2026

Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow

A generative detector that localizes tampering by estimating the local restoration cost required to align a query image with authentic visual-text statistics, rather than by learning forgery-specific decision boundaries is proposed, and Sparse-Constraint Rectified Flow is introduced, a detector-oriented adaptation of Flow Matching for spatially sparse anomaly localization.

Jiangling Zhang, Shuxuan Gao, Zeyu Chen et al. · 0 citations
Jul 2026

Uncertainty-Aware Deepfake Detection via Multi-View Structural Learning

An uncertainty-aware deepfake detection framework that identifies manipulations through inconsistencies across complementary evidence sources by introducing Inter-Branch Disagreement Calibration (IBDC), a disagreement-aware uncertainty modeling mechanism that links predictive uncertainty to conflicts among evidence streams.

Muhammad Umar Farooq, Kutub Uddin, Awais Khan et al. · 1 citation
Preprint Aug 2026

Invisible Shortcuts: Why Vision Encoders Know Your Camera

Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition. We hypothesize that large-scale semantic supervision, whether through categorical labels (ImageNet) or billion-scale captions (LAION), naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features. By introducing controlled metadata-semantics correlations, we show that stronger ones produce systematically higher sensitivity to metadata traces and larger performance degradation under metadata distribution shifts. We further explore mitigation strategies applied during and after pretraining that reduce sensitivity not only to targeted metadata but also to unseen ones, without sacrificing performance on downstream tasks. Metadata sensitivity also has a positive side: it partly explains the strong generated-image detection ability of some encoders, while its mitigation can improve out-of-distribution generalization. Code: https://github.com/ryan-caesar-ramos/visual-encoder-traces

Vladan Stojni'c, Ryan Ramos, Giorgos Kordopatis-Zilos et al. · 0 citations
Aug 2026

A confidence-guided hybrid network for image restoration

The confidence-guided hybrid network (CGHNet) is proposed, a parallel three-branch framework that jointly performs frequency-decoupled local restoration, global context modeling, and pixel-wise degradation prior estimation and its key component is a confidence-guided feature purification mechanism.

Xiaohui Kou, Yang Yan, Qiuyan Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.