MT2IE is presented, an evaluation framework in which a single multimodal large language model (MLLM) acts as an evaluator agent, iteratively generating the evaluation prompts and scoring the resulting images, showing that MT2IE's image-text consistency scores have higher correlation with human judgment than metrics previously introduced in the literature.
Abstract
The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate T2I progress. We present Multimodal Text-to-Image Eval (MT2IE), an evaluation framework in which a single multimodal large language model (MLLM) acts as an evaluator agent, iteratively generating the evaluation prompts and scoring the resulting images. We show that MT2IE's image-text consistency scores have higher correlation with human judgment than metrics previously introduced in the literature. MT2IE generates prompts that are efficient at probing T2I model performance: closely recovering the official T2I model rankings of three structurally distinct benchmarks from just 20 generated evaluation prompts, 28-105x fewer than the benchmarks'own prompt sets. When compared to existing evaluation metrics such as CLIPScore, VIEScore, and VQAScore, MT2IE's T2I model rankings are more faithful and far more consistent across multiple evaluation seeds when using the same number of prompts. MT2IE can also adapt evaluation to the model being tested: rewriting each prompt based on the model's own measured performance to produce a bespoke per-model benchmark that still recovers the official rankings and keeps the evaluated model in an informative scoring range. We hope that these results will encourage the development of dynamic and interactive evaluation frameworks, and mitigate the deprecation of automatic evaluation benchmarks.
A novel evaluation framework called Image2Text2Image is proposed, which leverages diffusion models, such as Stable Diffusion or DALL-E, for text-to-image generation, and does not rely on human-annotated reference cap-tions, making it a valuable tool for assessing image captioning models.
Jia-Hong Huang, Hongyi Zhu, Yixian Shen et al.· 0 citations
With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study seman...
Yu Zhao, Jia-Rui Wang, Hui-Yu Duan et al.· 0 citations
In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...
Failure analysis shows that the leading systems primarily lose points through object miscounting and geometric artifacts, whereas the trailing systems more frequently produce garbled text, and Ideogram 3.0 also frequently omits requested elements.
Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmed Rashad· 1 citation
AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.
Han Bao, Yue Huang, Yan-Bo Wang et al.· Proceedings of the 32nd ACM...· 0 citations
A scalable evaluation framework for simile understanding is proposed that includes a controlled simile dataset in which metaphorical vehicles are drawn from a predefined set of object-detectable categories and combined with diverse templates.
Luecheng Wang, Shintaro Ozaki, Hidetaka Kamigaito et al.· 0 citations