Integrating VLM Perception and Geometric-Semantic SLAM for Indoor Environment Recognition
This paper presents a semantic mapping system that integrates Vision-Language Models (VLMs) with Simultaneous Localization and Mapping (SLAM) to enhance the understanding of indoor environments. While traditional SLAM systems primarily focus on geometric occupancy, they often lack the semantic context needed for high-level robotic tasks. The proposed framework bifurcates the mapping process into two synchronized subsystems: Geometric SLAM and VLM Recognition. The first system uses LiDAR and odometry data to perform real-time localization and mapping, providing a stable and accurate environment for semantic data integration. Simultaneously, the second system leverages a VLM with a ZED2 depth camera to identify furniture and perform precise object size estimation, mapping these semantic entities into the global coordinate frame. A data fusion layer then synthesizes these inputs into a comprehensive Semantic Map that includes both geometric structures and physical object dimensions. Experimental results demonstrate that our integrated approach achieves precise spatial anchoring of semantic entities and improves overall mapping stability, providing a robust foundation for robot environmental understanding in complex indoor scenarios.