FUSED combines low-level forensic cues with high-level semantic features using a sparsely-gated Mixture-of-Experts architecture, enabling the model to adaptively prioritize the most relevant signal for each token.
Abstract
Diffusion-based inpainting models modify only a localized part of an image, while many AI-image detectors rely on global artifacts and do not localize. These artifacts vary across generators, limiting detector transfer under distribution shifts. Recent work shows that restoring the authentic pixels outside the inpainted region removes these cues and can degrade pretrained detectors. To address this, we present FUSED, a unified framework for the joint detection and localization of AI-generated inpainting. FUSED combines low-level forensic cues with high-level semantic features using a sparsely-gated Mixture-of-Experts architecture, enabling the model to adaptively prioritize the most relevant signal for each token. For each input, FUSED predicts both an image-level manipulation score and a pixel-level mask of the inpainted area. On the OpenSDID cross-generator benchmark, FUSED achieves the best average detection and localization, with the largest gains on unseen generators. The same model transfers directly to the held-out AutoSplice and CocoGlide benchmarks, more than doubling localization performance. Evaluating each held-out benchmark with and without the global generator artifact further shows that all evaluated methods, ours included, partly read the artifact as evidence of manipulation, and FUSED remains the strongest under both conditions. Code and pretrained models are available at https://github.com/AntonNuzhdin/FUSED.
LaP-Forensics is presented, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence that supports the utility of the residual stream under the evaluated settings, while free-form textual faithfulness and reliability under post-processing remain open limitations.
Can Wang, Yuhao Wang, Yu-She Cao et al.· arXiv.org· 1 citation
The proposed Attention-Based Deep Learning Pipeline of AI-Created Image Recognition incorporates three integrated branches, including low-level statistical feature extraction, high-level semantic representation learning, and attention-based feature refinement mechanism, which support the robustness and generalization ability of the proposed model in detecting AI-generated images in a variety of generators and conditions.
Nadia Ali· Al-Noor Journal of Engineeri...· 0 citations
The rapid advancement of generative models has significantly worsened the problem of manipulated image detection, as these methods are capable of producing highly realistic forgeries, reinforcing the importance of multimedia forensics. Conventional approaches typically frame image manipulation detection as a binary classification task (real vs. generated), which limits the capability to distinguish and localize different forms of manipulation. To address these constraints, this work extends an existing detector by introducing a unified multiclass framework (real vs. fully generated vs. tampered). In addition to classifying image authenticity, the framework incorporates a segmentation branch to enable pixel-level localization of tampered regions. The proposed approach outperforms selected recent benchmarks, offering an efficient solution with improved classification accuracy and higher IoU scores for the localization task. Find the code at https://github.com/anngal01/From-Detection-to-Localization-A-Unified-Forensics-Framework-for-Fully-Synthetic-and-Tampered-Images.
Annalisa Gallina, M. Fiorucci, Marco Brigo et al.· 0 citations
This work explores an approach that integrates wavelet-based frequency analysis with deep learning to enhance deepfake detection, and suggests that wavelet sub-bands expose manipulation cues that are useful for detecting unseen fake classes, but they should not be interpreted as a uniform robustness improvement.
In this paper, we propose a unified multimodal forensic framework for joint image-authenticity detection and manipulation localization. The proposed method integrates frequency-domain features cues with visual representations to improve manipulation-sensitive feature learning. In addition, an object-context mask-refinement module is introduced to enhance pixel-level localization through contextual attention and boundary-aware residual refinement. We compare the proposed method against ten existing state-of-the-art approaches on benchmark datasets for image-level detection and pixel-level localization. Experimental results show that the proposed method achieves the best image-level detection performance, with an overall classification accuracy and Macro F1-score of 99.1%, demonstrating robust discrimination among real, fully synthetic, and tampered images. Furthermore, it achieves category-specific accuracies of 98.8%, 99.9%, and 98.5% for real, fully synthetic, and tampered images, respectively, highlighting its effectiveness across diverse generation and manipulation scenarios.
M.R. Aliabadi, Naidan Zhang, Frank Y. Shih· International journal of pat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.