Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstrations across diverse camera setups and the policy must generalize this mapping across viewpoints. We address this mismatch with robot-centric pointmaps, images whose pixels store the 3D coordinates of scene points in the robot frame. Pointmaps provide robot-frame 3D geometry while preserving the dense H x W grid expected by pretrained 2D VLAs, so they integrate into existing VLAs with minimal architectural change. On RoboCasa, pointmaps improve both pi0.5 and SmolVLA and outperform representative camera-viewpoint and 3D-aware baselines. In real-robot experiments, their advantage over an RGB-only policy widens when the camera is moved to a placement unseen during training.
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to"play"robots through a compact semantic interface linking intent to action. Show-Harness exposes...
Yan-Zhe Chen, Ze-Chen Bai, Zhi-Jun Cao et al.· 23 citations· ⚡2
This survey reviews representative Transformer-based autonomous driving models and organizes them by task role, sensing configuration, and architectural design and analyzes how efficiency constraints reshaping model design choices in practice affects deployability, robustness, and safety.
This work introduces SurgCoTBench, the first reasoning-focused benchmark in RAS, and proposes SurgRAW, a clinically aligned Chain-of-Thought (CoT) driven agentic workflow for zero-shot multi-task reasoning in surgery, which surpasses mainstream VLMs and agentic systems and outperforms a supervised model.
Chang Han Low, Ziyue Wang, Tianyi Zhang et al.· IEEE Robotics and Automation...· 20 citations· ⚡2
We introduce CoinFT, a capacitive 6-axis force/torque (F/T) sensor that is compact, light, low-cost, and robust with an average root-mean-squared error of 0.16 N for force and 1.08 mN m for moment when the input ranges from 0-14 N and 0-5 N in normal and shear directions, respectively. CoinFT is a stack of two rigid PC...
Hojung Choi, Jun En Low, Tae Myung Huh et al.· arXiv.org· 20 citations· ⚡1
This work proposes AtomicVLA, a unified planning-and-execution framework that jointly generates task-level plans, atomic skill abstractions, and fine-grained actions, and introduces a flexible routing encoder that automatically assigns dedicated atomic experts to new skills, enabling continual learning.
Likui Zhang, Tao Tang, Zhihao Zhan et al.· arXiv.org· 19 citations· ⚡3
This paper introduces POEF (POlicy EFfective Jailbreak), an automated red-teaming framework that takes into account the robot-specific constraints during both the optimization and evaluation processes and proposes two defense strategies that mitigate the behavior jailbreak risks.
Xuancun Lu, Zhen Huang, Xin-Feng Li et al.· 17 citations· ⚡5
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.