Skip to content
Review

Evidence-gated multimodal parsing and vectorization of architectural floor plans

Sep 2026 · 0 citations
Computer Science

TL;DR

SALI-FP is introduced, an evidence-gated multimodal pipeline that converts a plan into reviewable semantic maps, objects, vectors, and relation records while constraining local revisions by image evidence.

Abstract

Architectural floor plans remain a high-friction barrier to archive digitization and early design-model preparation because heterogeneous graphics encode spatial semantics and editable geometry together. We introduce SALI-FP, an evidence-gated multimodal pipeline that converts a plan into reviewable semantic maps, objects, vectors, and relation records while constraining local revisions by image evidence. In a full production audit of 11,534 heterogeneous plans, SALI-FP produced structured outputs for every plan, including 752,510 valid polygon-bearing objects. The same output form has supported initial drawing digitization and design-model preparation in practical design work. Public-benchmark calibration is paired with a 30-case matched visual evidence set in Appendix F, where room-scale coverage, openings, oblique boundaries, and circulation continuity can be inspected directly. SALI-FP offers an engineering-oriented interpretation-to-geometry workflow for reviewed CAD/BIM preparation and existing-building information recovery.

View source

Similar papers

Preprint Aug 2026

PolarSym: Polar Geometry-aware Attention for CAD Floorplan Parsing

PolarSym is proposed, a polar-coordinate geometry-aware attention framework for CAD plan parsing that improves the geometric awareness of Transformers at low computational cost, offering an effective geometric modeling paradigm for CAD plan parsing.

Kerui Chen, Yi-Qing Wang, Kangzhou Xin et al. · 0 citations
Preprint Aug 2026

Groundbench: Multi-Resolution Polygon Grounding Exposes the Geometry Gap in Vision-Language Models

An operational output-geometry gap spanning localisation, boundary construction, serialisation, and topology, rather than latent boundary perception alone is Measures an operational output-geometry gap spanning localisation, boundary construction, serialisation, and topology rather than latent boundary perception alone...

Zhong-Han Bian, Zhen-Ran Wang, Jin-Song Li et al. · 0 citations
Preprint Aug 2026

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

OmniMech is introduced, the first million-scale benchmark for evaluating VLMs on executable CAD generation from industrial manufacturing data, and experiments show that current VLMs and CAD-specialized models still struggle with executable program synthesis, fine-grained 3D reconstruction, and reliable enforcement of d...

Tai-Ting Lu, Run-Ze Liu, Zi-Wei Dong et al. · 0 citations
Preprint Aug 2026

MIRAGE-CAD: Construction-Mediated Multimodal Generation of Executable CAD Programs

MIRAGE-CAD maps each input to a shared construction representation and mediates program generation through an explicit construction-plan interface and shows that executable validity, geometric fidelity, and parametric responsiveness can diverge substantially and should therefore be evaluated separately.

Jizong Zhan · 0 citations
Open access Sep 2026

Spatially Constrained 3D Scene Graphs for Scene-Grounded Subtask Generation

3D scene graphs connect spatial perception with high-level language reasoning, but unverified false-positive groundings can impair downstream subtask generation. We propose a feed-forward, spatially constrained framework that grounds task-relevant objects and generates position-aware subtasks without an iterative task–...

Junsang Ryu, Si-Woo Lee, Inseon Choi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.