Skip to content
Open access

Retrieval-Augmented Floor Plan Generation with Pre-Trained Text-to-Image Models: A Saudi Building Code Study

Aug 2026 · Electronics · 0 citations · 27 references

TL;DR

The study sets out plainly what off-the-shelf models and prompt-level guidance deliver without fine-tuning, and where they still fall short of professional, code-verified design.

Abstract

Floor plan design normally depends on architectural training or CAD software, a barrier for homeowners, students, and small practices alike. The author asks a narrower question: can a general-purpose, pre-trained text-to-image model draw a usable floor plan straight from a written brief, and can a building code be folded into that process? To find out, four models, Gemini, DALL-E, DeepAI, and Stable Diffusion, were put through tests using ten descriptions of villas and apartment buildings. Relevant Saudi Building Code (SBC) clauses, covering minimum room sizes, corridor widths, accessibility, and fire-safety provisions, were retrieved and written into each prompt before generation, and the outputs were assessed quantitatively by accuracy and latency, with SBC compliance and realism recorded only as qualitative observations. Gemini was the fastest by a wide margin, averaging 8.3 s per plan against 40.5 for Stable Diffusion, the slowest, and it also scored highest for accuracy; that ordering is not established here, however, because the models were not scored by a common judge, and each description was generated only once. DALL-E drew the most realistic images but was slower and looser on detail. One limitation cut across all four: none reported room dimensions reliably, so compliance can be verified only in part from the image. A case is presented in which Gemini printed area labels directly on the image, yet a pixel-level measurement shows the room with the smallest printed area drawn as the largest of the three, an internal inconsistency that needs no external ground truth to demonstrate and that exposes the core limitation of current text-to-image decoders. On the strength of these results, the best model, Gemini, was built into a Django web application that turns a typed description into a viewable plan. The system is offered as an early-stage drafting assistant, not a code-verified architectural design tool. The study sets out plainly what off-the-shelf models and prompt-level guidance deliver without fine-tuning, and where they still fall short of professional, code-verified design.

Read PDF

Similar papers

Review Sep 2026

Evidence-gated multimodal parsing and vectorization of architectural floor plans

SALI-FP is introduced, an evidence-gated multimodal pipeline that converts a plan into reviewable semantic maps, objects, vectors, and relation records while constraining local revisions by image evidence.

Hong-Xuan Chen, Wenda Wang, Jia-Chen Lu et al. · 0 citations
Book Open access Aug 2026

CodeCAD: A Parametric CAD Dataset for Programmatic 3D Model Generation

This work presents a dataset of over 95,000 native OpenSCAD models, primarily containing mechanical parts and engineering components, designed specifically for the programmatic generation of 3D models, and aims to open new opportunities for programmatic 3D model generation.

D. Fresacher, Klaus Diepold · 0 citations
Preprint Aug 2026

Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning

The proposed framework leverages the reasoning and language-understanding capabilities of large language models (LLMs), while grounding the reasoning with geometry evidence and structured material/printer knowledge to generate reliable pre-print recommendations.

Zhao-Da Du, Qiaojie Zheng, Xiaoli Zhang · 0 citations
Open access Aug 2026

Implementation U-Net for Semantic Segmentation and Perspective Transformation in Floor Tile Virtual Try-On Systems

Conventional floor tile VTO systems rely on printed catalogs, physical samples, or marker-based AR, making it difficult for consumers to visualize ceramic products in real rooms. This study proposes a markerless VTO framework integrating U-Net floor segmentation and homography-based perspective transformation. The mode...

Fahrul Rozi, A. Juwita, Yusuf Eka Wicaksana et al. · 0 citations
Open access 2026

Assessing the Reliability of LLM-Based Architectural Design Image Generation: A Comprehensive Evaluation Framework and Benchmark (AGEB)

Text-to-image systems have progressed from research prototypes to widely deployed tools, but high-fidelity imagery alone does not satisfy the requirements of professional architectural design. The Architecture Generation and Evaluation Benchmark (AGEB) is introduced as an end-to-end automated benchmark that assesses th...

Xiao-Yang Fan, Hao-En Xin, Zi-Teng Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.