PointLIBERO: Unlocking Spatial Awareness in VLAs With a Novel 3-D Dataset and a Lightweight Framework
Abstract
Vision-Language-Action models such as OpenVLA and DexVLA have demonstrated impressive generalization by leveraging large-scale 2D robotic datasets. However, their reliance on 2D RGB imagery significantly limits their 3D spatial reasoning, leading to spatial naivety in depth-sensitive tasks. Bridging this gap is hindered by two primary challenges: the data scarcity of high-quality 3D robotic demonstrations and the risk of catastrophic forgetting when fine-tuning pretrained models. To overcome these obstacles, we first introduce PointLIBERO, a new large-scale robotic manipulation benchmark that augments the widely-used LIBERO suite with trajectory-aligned point cloud data generated via a rigorous reconstruction pipeline. Concurrently, to effectively utilize this 3D modality, we propose a lightweight framework underpinned by a three-stage collaborative fine-tuning strategy. Notably, to address the modality misalignment, we incorporate a cross-modal contrastive alignment that synchronizes 3D geometric representations with the pre-trained 2D semantic space. These aligned features are then integrated via a zero-initialized point prompting mechanism, which treats 3D cues as learnable prompts to progressively infuse spatial awareness into the VLA’s attention layers. This framework ensures precise geometric grounding while preserving the priors of the pre-trained 2D baseline. Extensive experiments on both the PointLIBERO benchmark and real-world robotic settings demonstrate that our method achieves substantial performance gains over state-of-the-art 2D VLAs, particularly in spatially demanding tasks. Our code and dataset are available at https://github.com/YhpycLM/PointLIBERO Note to Practitioners—This paper addresses the practical challenge that modern robotic AI models often lack the 3D spatial awareness required for precise physical execution. Because they rely primarily on 2D images, these models struggle with depth-sensitive tasks and visual occlusion. To overcome this, we introduce a method to equip existing 2D robotic models with explicit 3D perception using point cloud data. We provide a new 3D-enhanced dataset and a lightweight framework that efficiently integrates spatial awareness into pre-trained models without the need to retrain them from scratch. This approach significantly improves the reliability and success rate of robots in complex spatial and long-horizon manipulation tasks. Beyond tabletop manipulation, these techniques can be applied to industrial assembly, automated quality inspection, and logistics handling, where 3D spatial reasoning is critical.