Skip to content
Open access

Evaluating Vision-language Models for Zero-shot Room Area Estimation in Floor Plan Images

Aug 2026 · Periodica polytechnica. Civil engineering · 0 citations · 38 references

TL;DR

It is suggested that while the evaluated open-source VLMs can approximate spatial proportions in simplified color-coded layouts, their quantitative accuracy remains significantly below that of deterministic image analysis methods.

Abstract

Vision-language models (VLMs) have shown strong performance on multimodal reasoning tasks, yet their ability to perform quantitative analysis of technical drawings remains largely unexplored. This study evaluates the instruction-guided zero-shot performance of three open-source VLMs, on the task of estimating room areas from color-coded residential floor plan images. Of the three models tested, two produced quantitatively evaluable structure outputs. A dataset of 100 floor plans, comprising 472 plan-level room-type evaluation comparisons derived from 894 ground-truth room instances, was used to compare predicted and ground truth areas. Results show moderate predictive correlation (R2 ≈ 0.71), but substantial estimation errors (MAPE ≈ 43%), with noticeably larger errors for smaller rooms. A classical pixel counting baseline, using color segmentation, achieved MAPE = 5.68% at the room-instance level, and MAPE = 1.58% when evaluated at the floor plan level, highlighting the limitations of the tested VLMs for precise geometric estimation under this controlled setup. These findings suggest that while the evaluated open-source VLMs can approximate spatial proportions in simplified color-coded layouts, their quantitative accuracy remains significantly below that of deterministic image analysis methods.

Read PDF

Similar papers

Preprint Aug 2026

StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference, and that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.

Michelle Lin · 0 citations
Preprint Oct 2026

From Reasoning Failures to Composable Video Spatial Intelligence

Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparin...

Peng-Zhan Sun, Jun-Bin Xiao, Ramanathan Rajaraman et al. · 0 citations
Preprint Sep 2026

Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation

Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to achieve state-of-the-art performance across multiple visual perception tasks. Motivated by this perspective, we ask how far image generators can go on public visual perception benchmarks in a zero-shot setting...

Shangzhe Di, Zhaokai Wang, Wei-Di Xie · 1 citation
Open access Aug 2026

Semantic Specificity Degradation in Zero-Shot Generative Vision-Language Models

Vision-Language Models (VLMs) have shown strong zero-shot performance in generating free-form image descriptions. However, most evaluations focus on hallucinated content, while less attention is given to whether models preserve specific object identities. This study investigates whether zero-shot generative VLMs retain...

Nahumi Nugrahaningsih, F. Sylviana · 0 citations
#small language model Preprint Sep 2026

Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models

A modular, predictor agnostic, tool-augmented framework that equips a small VLM with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing, matching a scripted pipeline on three of four tasks.

Kai Glantz, Clemens Grange · 0 citations
Conference Open access 2026

Enhancing Zero-Shot Visual Reasoning with Qwen2-VL on CLEVR

Overall, the results show that the baseline is fairly stable under lightweight prompt and decoding changes, however, these low-cost adjustments do not solve the model's main weakness, which still appears in spatial and compositional reasoning.

Zhong-Tian Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.