Evidence Fusion for Analyzing Multimodal Image Manipulation
Abstract
In this paper, we propose a unified multimodal forensic framework for joint image-authenticity detection and manipulation localization. The proposed method integrates frequency-domain features cues with visual representations to improve manipulation-sensitive feature learning. In addition, an object-context mask-refinement module is introduced to enhance pixel-level localization through contextual attention and boundary-aware residual refinement. We compare the proposed method against ten existing state-of-the-art approaches on benchmark datasets for image-level detection and pixel-level localization. Experimental results show that the proposed method achieves the best image-level detection performance, with an overall classification accuracy and Macro F1-score of 99.1%, demonstrating robust discrimination among real, fully synthetic, and tampered images. Furthermore, it achieves category-specific accuracies of 98.8%, 99.9%, and 98.5% for real, fully synthetic, and tampered images, respectively, highlighting its effectiveness across diverse generation and manipulation scenarios.