Aug 2026· Periodica polytechnica. Civil engineering· 0 citations· 38 references
TL;DR
It is suggested that while the evaluated open-source VLMs can approximate spatial proportions in simplified color-coded layouts, their quantitative accuracy remains significantly below that of deterministic image analysis methods.
Abstract
Vision-language models (VLMs) have shown strong performance on multimodal reasoning tasks, yet their ability to perform quantitative analysis of technical drawings remains largely unexplored. This study evaluates the instruction-guided zero-shot performance of three open-source VLMs, on the task of estimating room areas from color-coded residential floor plan images. Of the three models tested, two produced quantitatively evaluable structure outputs. A dataset of 100 floor plans, comprising 472 plan-level room-type evaluation comparisons derived from 894 ground-truth room instances, was used to compare predicted and ground truth areas. Results show moderate predictive correlation (R2 ≈ 0.71), but substantial estimation errors (MAPE ≈ 43%), with noticeably larger errors for smaller rooms. A classical pixel counting baseline, using color segmentation, achieved MAPE = 5.68% at the room-instance level, and MAPE = 1.58% when evaluated at the floor plan level, highlighting the limitations of the tested VLMs for precise geometric estimation under this controlled setup. These findings suggest that while the evaluated open-source VLMs can approximate spatial proportions in simplified color-coded layouts, their quantitative accuracy remains significantly below that of deterministic image analysis methods.
The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference, and that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.
Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparin...
Peng-Zhan Sun, Jun-Bin Xiao, Ramanathan Rajaraman et al.· 0 citations
Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to achieve state-of-the-art performance across multiple visual perception tasks. Motivated by this perspective, we ask how far image generators can go on public visual perception benchmarks in a zero-shot setting...
Vision-Language Models (VLMs) have shown strong zero-shot performance in generating free-form image descriptions. However, most evaluations focus on hallucinated content, while less attention is given to whether models preserve specific object identities. This study investigates whether zero-shot generative VLMs retain...
Nahumi Nugrahaningsih, F. Sylviana· Edu Komputika Journal· 0 citations
A modular, predictor agnostic, tool-augmented framework that equips a small VLM with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing, matching a scripted pipeline on three of four tasks.
Overall, the results show that the baseline is fairly stable under lightweight prompt and decoding changes, however, these low-cost adjustments do not solve the model's main weakness, which still appears in spatial and compositional reasoning.
Zhong-Tian Liu· ITM Web of Conferences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.