Skip to content

Multi-Modal Language Models as Text-to-Image Model Evaluators

May 2025 · arXiv.org · Vol abs/2505.00759 · 4 citations · ⚡ 1 influential · 76 references
Computer Science

TL;DR

MT2IE is presented, an evaluation framework in which a single multimodal large language model (MLLM) acts as an evaluator agent, iteratively generating the evaluation prompts and scoring the resulting images, showing that MT2IE's image-text consistency scores have higher correlation with human judgment than metrics previously introduced in the literature.

Abstract

The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate T2I progress. We present Multimodal Text-to-Image Eval (MT2IE), an evaluation framework in which a single multimodal large language model (MLLM) acts as an evaluator agent, iteratively generating the evaluation prompts and scoring the resulting images. We show that MT2IE's image-text consistency scores have higher correlation with human judgment than metrics previously introduced in the literature. MT2IE generates prompts that are efficient at probing T2I model performance: closely recovering the official T2I model rankings of three structurally distinct benchmarks from just 20 generated evaluation prompts, 28-105x fewer than the benchmarks'own prompt sets. When compared to existing evaluation metrics such as CLIPScore, VIEScore, and VQAScore, MT2IE's T2I model rankings are more faithful and far more consistent across multiple evaluation seeds when using the same number of prompts. MT2IE can also adapt evaluation to the model being tested: rewriting each prompt based on the model's own measured performance to produce a bespoke per-model benchmark that still recovers the official rankings and keeps the evaluated model in an informative scoring range. We hope that these results will encourage the development of dynamic and interactive evaluation frameworks, and mitigate the deprecation of automatic evaluation benchmarks.

View source

Similar papers

for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models

A novel evaluation framework called Image2Text2Image is proposed, which leverages diffusion models, such as Stable Diffusion or DALL-E, for text-to-image generation, and does not rely on human-annotated reference cap-tions, making it a valuable tool for assessing image captioning models.

Jia-Hong Huang, Hongyi Zhu, Yixian Shen et al. · 0 citations
Preprint Sep 2026

Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment

With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study seman...

Yu Zhao, Jia-Rui Wang, Hui-Yu Duan et al. · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
Review Aug 2026

Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

Failure analysis shows that the leading systems primarily lose points through object miscounting and geometric artifacts, whereas the trailing systems more frequently produce garbled text, and Ideogram 3.0 also frequently omits requested elements.

Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmed Rashad · 1 citation
Book Open access Aug 2026

AutoDavis: Automatic and Dynamic Evaluation Protocol of Large Vision-Language Models on Visual Question-Answering

AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.

Han Bao, Yue Huang, Yan-Bo Wang et al. · 0 citations
Preprint Aug 2026

Simile Understanding in Text-to-Image Models: An Evaluation Framework

A scalable evaluation framework for simile understanding is proposed that includes a controlled simile dataset in which metaphorical vehicles are drawn from a predefined set of object-detectable categories and combined with diverse templates.

Luecheng Wang, Shintaro Ozaki, Hidetaka Kamigaito et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.