This paper investigates synthetic digital image detection across different generative models. The task is relevant due to rapid development of generative artificial intelligence, which enables creation of visual content that is difficult to distinguish from real imagery and increases risks related to misinformation, visual forgery, and declining trust in digital media. The study compares several detection approaches: CLIP-based semantic representations, FFT-based frequency-domain features, fusion models combining semantic and frequencydomain information, and ensembles of detectors. Experiments were conducted on tinyGenImage, a subset of GenImage, using BigGAN and Stable Diffusion v1.5 images for training and validation. Performance across different sources was evaluated on images generated by Midjourney, Wukong, and GLIDE, which were not used during training. Results show that CLIP-based models provide a strong baseline, while FFT-only models perform weaker as standalone detectors. Fusion models did not consistently improve over CLIP baselines, whereas ensembling achieved the best overall performance and improved classification quality across different generative sources.
Vadim Borzov, E. Rybakov, M. Moseva et al.· 2026 Systems of Signal Synch...· 0 citations
This model leverages different layers of the large language model (LLM) decoder, treating them as multiple detectors that perform synchronous distortion detection using multi-level features, which helps mitigate the ambiguous foreground-background separation commonly encountered in the S2A problem.
Ziheng Jia, Yingji Liang, Jiaying Qian et al.· 0 citations
This work establishes a new paradigm for generated image detection by recasting the detection task as a problem of machine unlearning, and introduces two detection methods: data-free detection, which prunes model parameters to induce unlearning without data access, and data-driven detection, which optimizes LVMs to unlearn knowledge tied to generated images.
Jun Nie, Yonggang Zhang, Tongliang Liu et al.· 0 citations
GurAI is proposed, a transparent logistic late-fusion method that combines Rich384 and DeMamba logits that suggests that transparent late fusion can exploit complementary detector strengths more effectively than architectural redesign alone when facing generator diversity.
The findings demonstrate that AI-generated images serve as an effective diagnostic tool for robustness evaluation and failure mode analysis, and supports reliable deployment of product image quality assessment models and highlights the value of synthetic data as a structured robustness testing resource in multimedia and e-commerce systems.
Imad Tbaileh, Huthaifa I. Ashqar· Multimedia Systems· 0 citations
SafeIMG is introduced, a safety-oriented benchmark spanning 12 public- and individual-safety scenarios generated using GPT Image 2.0 that provides human annotations that localise suspicious regions and explain local artefacts and higher-level commonsense or physical inconsistencies.
Yizhi Wang, Yichen Xiao, Linan Yue et al.· 0 citations