Skip to content
Preprint

RA-ClipScore: Making Generative Model Evaluation More Interpretable

Aug 2026 · 0 citations · 58 references
Computer Science

TL;DR

RA-CLIPScore provides more robust and interpretable evaluations than prior methods, particularly under distribution misalignment or partially irrelevant textual attributes, and aligns more closely with human perception of visual diversity than existing semantic metrics.

Abstract

Generative models can produce images nearly indistinguishable from real data, yet rigorous and interpretable evaluation remains challenging. Conventional metrics such as FID provide only scalar scores with limited diagnostic insight. Widely adopted CLIP-based metrics enable semantic evaluation beyond simple training class labels, but inherit limitations from CLIP's training paradigm that restrict attribute-wise analysis. We propose RA-CLIPScore, a novel metric that mitigates these issues and extends CLIP-based evaluation to spatial distribution alignment, measuring whether generated objects adhere to the positional priors found in the training data. RA-CLIPScore introduces dual prompts to decouple competing attributes and leverages local patch tokens to capture fine-grained regional semantics. We evaluate image generative models on their ability to match both attribute and spatial distributions of the training data. Extensive experiments show that RA-CLIPScore provides more robust and interpretable evaluations than prior methods, particularly under distribution misalignment or partially irrelevant textual attributes. We further demonstrate how it reveals spatial biases in generative models. User evaluations confirm that Regional Single Attribute Divergence based on our RA-CLIPScore aligns more closely with human perception of visual diversity than existing semantic metrics.

View source

Similar papers

Preprint Aug 2026

Where Does Generative Difficulty Reside? An Empirical Study of Target Representations

This work studies raw pixels, SD-VAE latents and DINOv2 as well as MAE representation-autoencoder features within a unified masked autoregressive rectified-flow model and shows that compression, reconstruction fidelity, token dimensionality, and visible semantic clustering do not individually predict generative behavio...

Marcel Plocher, B. Schölkopf, Andreas Geiger et al. · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
Preprint Aug 2026

VGI-Bench: Probing Visual Intelligence in Video Generation Models

VGI-bench is introduced, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models, and it is hoped VGI-bench will help stimulate the development of next-generation video generation mode...

Xuan He, Cong Wei, Yuhao Cheng et al. · 1 citation
Preprint Aug 2026

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

The Evaluation Agent framework is proposed, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses and is efficient, promptable, explainable, and scalable across models and tools.

Shu-Lin Tian, Zi-Qi Huang, Fan Zhang et al. · 2 citations
Preprint Aug 2026

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.

Zehua Hao, Fang Liu, Qinliang Wang et al. · 0 citations

Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling

An Implicit Cultural Alignment Reward Model built upon a lightweight 4.2-billion-parameter Multimodal Large Language Model (MLLM) integrated with a Skip-connection Cross-Attention mechanism, enabling late-stage semantic features to directly attend to early-stage visual representations and better preserve culturally sal...

Boxin Chang, Yu-Chih Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.