SceneReasoner: Decoupled Spatial Tokenization for Indoor Large-Scene Understanding with LLMs
Abstract. Most existing 3D vision-language models focus on object-level or single-room understanding and perform poorly in large-scale, multi-room indoor environments where task-relevant objects constitute only a small fraction of the total point cloud. When multi-room point clouds are fed directly into an LLM, critica...