Skip to content
Preprint

PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning

Aug 2026 · 1 citation · 41 references
Computer Science

TL;DR

PhysMind is introduced, a training-free agentic framework that constructs one reusable, question-agnostic executable world per video that fits analytic continuous-time dynamics and latent physical parameters without unrolling a time-stepped simulator.

Abstract

Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to interpret these dynamics and reason reliably about future and counterfactual outcomes. We introduce PhysMind, a training-free agentic framework that constructs one reusable, question-agnostic executable world per video. PhysMind recovers a temporally consistent dynamic scene through object segmentation, mesh reconstruction, and 6D pose tracking, then fits analytic continuous-time dynamics and latent physical parameters without unrolling a time-stepped simulator. Given a question, it inspects, continues, or edits the world and answers from the resulting trajectories and interactions. Relative to direct chain-of-thought (CoT) reasoning with the same VLM, PhysMind improves accuracy by 38.23 points on CLEVRER and 8.08 points on Physion++. On counterfactual questions, it exceeds the strongest evaluated VLM baseline, GPT-5.5, by 19.25 points.

View source

Similar papers

Preprint Aug 2026

TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment, and shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation.

Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al. · 1 citation
Preprint Sep 2026

Principia: Relational Physics Tests for Video Models

Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law,...

Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan et al. · 0 citations
Jul 2026

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

VisualPatchWorld is introduced, which represents world dynamics as code and first selects a qualitative dynamical form with short active probes, then fits that form's free parameters from recorded state-action traces by minimizing multi-step prediction error.

Jiaxin Bai, Jia–Jie Xiong · 0 citations
#artificial intelligence Preprint Sep 2026

From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often con...

Meng Luo, Yi-Chen Liu, Jia-Hao Wang et al. · 0 citations
Jul 2026

ContactFlow: A video action conditioning that transfers across embodiments

This work proposes Contact Flow, an embodiment-agnostic action representation that encodes manipulation through the trajectory of 3D contact points between an actor and a target object, and integrates this model into a propose-imagine-verify-act pipeline, where generated rollouts are assessed by a vision-language model...

Sami Azirar, Enrico Pallotta, Jan Nogga et al. · 1 citation
Preprint Aug 2026

CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding

Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with robot action domains. Egocentric videos are used only as auxiliary pre-training. It remains unclear whether improved VLM physical understandi...

Hanwen Wan, Dafeng Chi, Lin-Bo Zhai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.