Skip to content
#robotics Preprint

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

Jul 2026 · 1 citation · 57 references
Computer Science

Abstract

Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstrations across diverse camera setups and the policy must generalize this mapping across viewpoints. We address this mismatch with robot-centric pointmaps, images whose pixels store the 3D coordinates of scene points in the robot frame. Pointmaps provide robot-frame 3D geometry while preserving the dense H x W grid expected by pretrained 2D VLAs, so they integrate into existing VLAs with minimal architectural change. On RoboCasa, pointmaps improve both pi0.5 and SmolVLA and outperform representative camera-viewpoint and 3D-aware baselines. In real-robot experiments, their advantage over an RGB-only policy widens when the camera is moved to a placement unseen during training.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Show-Harness: Just a VLM Agent Can Play Robots

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to"play"robots through a compact semantic interface linking intent to action. Show-Harness exposes...

Yan-Zhe Chen, Ze-Chen Bai, Zhi-Jun Cao et al. · 23 citations · ⚡2
#artificial intelligence Review Apr 2023

Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey

This survey reviews representative Transformer-based autonomous driving models and organizes them by task role, sensing configuration, and architectural design and analyzes how efficiency constraints reshaping model design choices in practice affects deployability, robustness, and safety.

J. Zhong, Zheng Liu, Xiangshan Chen · 21 citations
#artificial intelligence Open access Mar 2025

SurgRAW: Multi-Agent Workflow With Chain of Thought Reasoning for Robotic Surgical Video Analysis

This work introduces SurgCoTBench, the first reasoning-focused benchmark in RAS, and proposes SurgRAW, a clinically aligned Chain-of-Thought (CoT) driven agentic workflow for zero-shot multi-task reasoning in surgery, which surpasses mainstream VLMs and agentic systems and outperforms a supervised model.

Chang Han Low, Ziyue Wang, Tianyi Zhang et al. · 20 citations · ⚡2
#robotics Mar 2025

CoinFT: A Coin-Sized, Capacitive 6-Axis Force Torque Sensor for Robotic Applications

We introduce CoinFT, a capacitive 6-axis force/torque (F/T) sensor that is compact, light, low-cost, and robust with an average root-mean-squared error of 0.16 N for force and 1.08 mN m for moment when the input ranges from 0-14 N and 0-5 N in normal and shear directions, respectively. CoinFT is a stack of two rigid PC...

Hojung Choi, Jun En Low, Tae Myung Huh et al. · 20 citations · ⚡1

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

This work proposes AtomicVLA, a unified planning-and-execution framework that jointly generates task-level plans, atomic skill abstractions, and fine-grained actions, and introduces a flexible routing encoder that automatically assigns dedicated atomic experts to new skills, enabling continual learning.

Likui Zhang, Tao Tang, Zhihao Zhan et al. · 19 citations · ⚡3
#artificial intelligence Preprint Dec 2024

Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-Based Robots

This paper introduces POEF (POlicy EFfective Jailbreak), an automated red-teaming framework that takes into account the robot-specific constraints during both the optimization and evaluation processes and proposes two defense strategies that mitigate the behavior jailbreak risks.

Xuancun Lu, Zhen Huang, Xin-Feng Li et al. · 17 citations · ⚡5

Related blog posts

Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.