Skip to content
Conference

Physical Reasoning VLA Models for Enhancing Robotic Manipulation

Jul 2026 · 2026 23rd International Conference on Ubiquitous Robots (UR) · pp. 363-368 · 0 citations · 51 references
Computer Science

Abstract

Vision-Language-Action (VLA) models have demonstrated strong performance in robot manipulation by leveraging pre-trained vision-language models to map observations directly to actions. However, existing approaches reason primarily at the visual or semantic level, lacking explicit understanding of the physical interactions that fundamentally govern manipulation tasks. In this paper, we propose Physics Reasoning VLA, a method that enables VLA models to explicitly reason about physical interactions prior to acting, grounded in two fundamental quantities: contact points, which specify where the target object interacts with the robot or surrounding environment, and contact forces, which describe the magnitude and direction of force applied at those locations. Rather than directly mapping observations to actions, our model first predicts contact points and forces at the pixel level via learnable physics queries, then incorporates the resulting physics-aware features alongside visual and language inputs to generate actions. To prevent physics reasoning from disrupting pre-trained visual and linguistic representations, we further introduce a hybrid attention mechanism that applies full attention over image, language, and proprioceptive tokens, while applying causal attention over physics query and action tokens. We evaluate our method on the RoboCasa simulation benchmark, demonstrating that physics reasoning consistently improves performance over the vanilla π0 baseline, with an average success rate improvement from 36.2% to 43.6%.

View source

Similar papers

Preprint Sep 2026

ForeTac-VLA: A Forecasting-Based Tactile-Vision-Language-Action Model for Contact-Rich Robotic Manipulation

Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, yet their reliance on visual perception limits robustness in contact-rich environments, where critical physical interaction states may not be visually observable. Existing tactile-enhanced VLA methods improve physical gro...

Zheng Tao, Xin Li, Xin Wang · 0 citations
#artificial intelligence Preprint Aug 2026

PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models

PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations, demonstrating that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA polic...

Davood Soleymanzadeh, Kai-Di Zhang, Zhi-Yuan Zhang et al. · 2 citations
Preprint Sep 2026

HINT: Human-Intent Inception for Long-Horizon Robot Manipulation

Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual...

Ming-Yu Mei, Haojie Xu, Shi-Hao Jin et al. · 0 citations
Preprint Sep 2026

H-VLA: Hierarchical Vision-Language-Action Model with Key-Action Reasoning and Motion Planning in a Unified Action Space

Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but many existing methods still rely on direct mappings from language and visual observations to dense actions. This formulation can weaken the semantic reasoning capability inherited from pre-trained Vision-Language Models (VLMs)...

Xiong-Feng Peng, Lu Xu, Yan-Dong Wang et al. · 0 citations
Preprint Sep 2026

PSR: Predictive Sensorimotor Representation Learning for Contact-Rich Manipulation

Contact-rich manipulation requires policies to generate precise actions by reasoning over contact forces, robot configurations, and interaction histories beyond visual observations. Existing methods passively condition on force feedback rather than actively predicting future contact dynamics, limiting their ability to...

Sheng-Bao Li, Peng Xu, Chao Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.