Skip to content
Preprint

On the Separation of Human and AI-Generated Images in CLIP Embedding Space

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

A significant difference is revealed between the visual evidence reflected in CLIP representations and that readily accessible to human perception, raising broader questions about the relationship between artificial and human vision and, ultimately, between artificial and human aesthetic judgment.

Abstract

We identify a previously unreported phenomenon in CLIP representations: human and AI-generated paintings spontaneously separate along the dominant principal directions of their joint embedding distribution, without any supervised objective designed to distinguish the two classes. Rather than exploiting this phenomenon for detection, our objective is to interpret it: we seek to identify the visual information underlying the separation and to trace it back from the embedding space to the image domain. We pursue this objective through a progressive investigation combining interpretable image representations with gradient-based inversion, used systematically as an experimental probe of the relationships identified in feature space. Robustness experiments and increasingly expressive statistical descriptors progressively rule out several intuitive explanations based on global image properties and simple local statistics, and point instead to distributed multiscale image structure. Multiscale scattering provides the most informative interpretable representation considered, but offers only a partial account of the phenomenon. Direct inversion provides a complementary and striking observation: substantial displacements along the dominant CLIP directions can be induced by image perturbations that remain nearly imperceptible to human observers, showing that the directions involved in the separation are highly sensitive to image variations with very low perceptual salience for humans. Taken together, these results reveal a significant difference between the visual evidence reflected in CLIP representations and that readily accessible to human perception, raising broader questions about the relationship between artificial and human vision and, ultimately, between artificial and human aesthetic judgment.

View source

Similar papers

Preprint Aug 2026

Structured Local Differential Modeling for AI-Generated Image Detection

RippleNet is proposed, an AI-generated image detection framework based on local differential signals that adaptively identifies forgery-sensitive regions and constructs multi-directional, multi-scale differential representations within local neighborhoods, explicitly characterizing anomalous patterns in neighborhood st...

Jia-Zhen Yang, Ruijin Jin, Junjun Zheng et al. · 0 citations
#computer vision Open access Sep 2026

ASAP: Visual analytics for identifying and analyzing image patterns in AI-generated images

Generative image models can produce highly realistic images, raising concerns about potential misuse in creating deceptive content. Current deepfake approaches face several challenges, including limited generalizability, lack of interpretability, and poor actionability. To help address these, we present ASAP, an intera...

Jinbin Huang, Yuki Ueno, Chen Chen et al. · 0 citations
Preprint Aug 2026

Learning visual representations for compositional analysis of artworks and photographs

This work compares two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets.

F. Behrad, T. Tuytelaars, Johan Wagemans · 0 citations
Preprint Aug 2026

Understanding Why Foundation Models Work for Diffusion-Generated Image Detection

This work investigates what cues are exploited by foundation-model-based detectors to distinguish real images from diffusion-generated ones and suggests that foundation-model-based detectors succeed by capturing non-semantic low-to-mid frequency distributional discrepancies between real and diffusion-generated images.

D. Cozzolino, G. Poggi, L. Verdoliva · 0 citations
Preprint Sep 2026

Bottom-up Modeling of Repeated Elements via Single Image Analysis-by-Synthesis

Qualitative results reveal superior reconstructions and interpretable decompositions compared to classical decomposition, joint alignment, and 3D object modeling methods, while maintaining a simple 2D formulation, suggesting that meaningful object discovery can emerge from single image learning alone.

Syrine Kalleli, Alexei A. Efros, Mathieu Aubry · 0 citations
#machine learning Preprint Sep 2026

SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such a...

Shuang Liang, Le-Jun Liao, Shi-Yuan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.