Assessing the Reliability of LLM-Based Architectural Design Image Generation: A Comprehensive Evaluation Framework and Benchmark (AGEB)
Abstract
Text-to-image systems have progressed from research prototypes to widely deployed tools, but high-fidelity imagery alone does not satisfy the requirements of professional architectural design. The Architecture Generation and Evaluation Benchmark (AGEB) is introduced as an end-to-end automated benchmark that assesses the reliability of architectural image generation along four axes: semantic correspondence with the design brief, a prompt-derived circulation proxy, perspective geometry, and no-reference technical quality. AGEB consists of 300 tasks organized into six cognitive levels, a unified generation protocol applied to five representative systems, five independent image generations per prompt, and a four-module evaluation pipeline combining a dual-expert chain-of-thought (COT) reasoning procedure, graph-theoretic circulation analysis, classical geometric consistency checks, and no-reference quality metrics (NIQE, BRISQUE, PIQE, Inception Score). Within this benchmark setting, the large language model is used only as an automated reasoning evaluator in the COT module, not as one of the scored image generators. Under one normalized prompt per task and five generated images per prompt for each system, the five systems exhibit differentiated strengths: GPT-Image-1 obtains the highest COT and circulation-proxy scores, Sora leads on BRISQUE, Inception Score, and PIQE, DALL-E 3 attains the highest perspective consistency, and Midjourney reaches the lowest NIQE. All five systems, however, score below 0.35 on the circulation proxy, indicating limited functional-programme compatibility within the controlled AGEB setting rather than verified pixel-level spatial functionality. Data and code are released at https://github.com/torfqy/Architecture-Generation-and-Evaluation-Benchmark-AGEB-