PointLIBERO: Unlocking Spatial Awareness in VLAs With a Novel 3-D Dataset and a Lightweight Framework
Vision-Language-Action models such as OpenVLA and DexVLA have demonstrated impressive generalization by leveraging large-scale 2D robotic datasets. However, their reliance on 2D RGB imagery significantly limits their 3D spatial reasoning, leading to spatial naivety in depth-sensitive tasks. Bridging this gap is hindere...