Open access
Jul 2026
From geometric labels to semantic understanding of indoor building components using multimodal large language models
Building-MLLM is proposed, a point cloud-centered multimodal large language model (MLLM) for indoor components, which models point clouds and instructions to generate responses across Simple Recognition, Complex Captioning, and Multi-Engineering Question Answering tasks, demonstrating superior indoor component language understanding and providing initial generalizability in transfer inference on other real-world datasets.
Shuju Jing, Chao Yin
· Automation in Construction · 0 citations