Skip to content
Review

Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

Aug 2026 · 1 citation · 10 references
Computer Science

TL;DR

Failure analysis shows that the leading systems primarily lose points through object miscounting and geometric artifacts, whereas the trailing systems more frequently produce garbled text, and Ideogram 3.0 also frequently omits requested elements.

Abstract

Text-to-image models are typically reported on average-case prompts, which understates the gap between systems on compositionally demanding requests involving precise object counts, multi-object attribute binding, legible embedded text, and explicit spatial constraints. We evaluate four production text-to-image systems: Hunyuan 3.0, Gemini 3 Pro Image ("Nano Banana Pro"), Black Forest Labs FLUX.2, and Ideogram 3.0. The evaluation uses the 48 hardest prompts drawn from the DataSeeds.AI Sample Dataset (DSD), selected through an automated complexity-scoring pass over the full corpus. Every generated image is graded using an independent-judge rubric. GPT-5.4-Pro authors an atomic, weighted, mutually exclusive and collectively exhaustive (MECE) evaluation rubric, while Gemini 3.1 Pro Preview independently determines whether each criterion is satisfied. Gemini 3 Pro Image ranks first with a score of 84.8/100, narrowly ahead of FLUX.2 at 82.3/100. Ideogram 3.0 and Hunyuan 3.0 score 65.7/100 and 63.3/100, respectively. Failure analysis shows that the leading systems primarily lose points through object miscounting and geometric artifacts, whereas the trailing systems more frequently produce garbled text. Ideogram 3.0 also frequently omits requested elements. Full per-sample rubrics, scores, and failure annotations are available from the authors upon request.

View source

Similar papers

Preprint Sep 2026

T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation

Recent text-to-image models have become increasingly capable of rendering explicit text, but reliable localized text control requires more than generating the correct string. In applications such as product labeling, signage, and interface design, target text should be rendered within a designated text-bearing region w...

Yan Wang, Xin-Yi Hou, Wei-Guo Lin et al. · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
#small language model Preprint Aug 2026

NumBench: Diagnosing Counting Failures in Text-to-Image Models

The Confidence-Weighted Numeric Precision Score (\cwnps), which aggregates three calibrated detectors and discounts uncertain proposals, is proposed for scalable evaluation and developed a process model in which requested instances compete for a finite set of resolvable image regions.

Sandeep Wadhwa, M. Vatsa, Richa Singh et al. · 0 citations
Preprint Sep 2026

GenPuzzle: Benchmarking Visual Reasoning in Image Generation Models

Recent image generation systems increasingly combine multimodal understanding, reasoning, and synthesis, suggesting that they may do more than render plausible scenes. Yet existing evaluations emphasize aesthetics, prompt alignment, compositionality, or text-based answers, leaving unclear whether these systems can solv...

Chang-Peng Zhao, Yi-Ren Song, Jin-Peng Wang · 0 citations
Preprint Aug 2026

Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing

Modern text-to-image (T2I) models often have similar total scores but different strengths, making practical selection difficult. Fine-grained benchmarks decompose prompts into questions, yet often return them to prompt scores and fixed categories, weakening attribution and ignoring complexity. Related requirements are...

Shao-An Zhao, Fang Zhao, Xueqiang Guo et al. · 0 citations
Review Aug 2026

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

Poplar is presented, a reproducible Specify--Render--Inspect pipeline for human-centric image dataset synthesis and is released together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human-centric collections.

Zhishan Zou · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.