Skip to content

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning

Jul 2026 · arXiv.org · Vol abs/2607.11044 · 0 citations · 41 references
Computer Science

Abstract

Humans can infer hidden physical processes from sparse observations, yet current evaluation protocols for Vision Language Models fail to assess whether such physical reasoning is genuinely captured. To address this gap, we introduce Retrospective Physical Process Reasoning, a new evaluation paradigm to reason backward from outcomes under explicit physical constraints. Building on the paradigm, we present RetroHolmes, the first real-world benchmark for Retrospective Physical Process Reasoning, comprising object-centric image pairs annotated with reachability labels and causal step sequences across diverse physical transitions. Using RetroHolmes, we analyze state of the art Vision Language Models and uncover systematic failure modes, including judgment bias in reachability assessment and belief dominance over physical evidence, mirroring sycophancy behavior observed in large language models. We further demonstrate a simple analysis-by-synthesis instantiation with visual simulation as an intermediate step, validating the diagnostic value of RetroHolmes and highlighting the importance of physically grounded intermediate representations for physical reasoning.

View source

Similar papers

Preprint Sep 2026

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed, is presented.

Wenzhuo Xu, Yu-Chen Zhu, Chongjian Ge et al. · 0 citations
Preprint Aug 2026

SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs

SymboUQ is a symbolic uncertainty quantification framework that estimates final-answer reliability from reasoning traces by distinguishing symbolizability, whether a claim can be represented in the verifier's formal language, from semantic determinacy, whether its execution yields an entailed or contradicted verdict ra...

Da-Hai Yu, Lin Jiang, Rong-Chao Xu et al. · 0 citations
Open access 2026

Teaching LLMs to Infer Real-World Consequences through Embodied Semantic Grounding

: Large Language Models (LLMs) have recently advanced in real-world commonsense reasoning, including understanding everyday object behaviors and inferring their attributes from text. However, they remain limited in reasoning about the real-world consequences of events, such as how object failures, obstructions, or stru...

Manaswi Kulahara, Khadija Parwez, Faisal Alhwikem et al. · 0 citations
Preprint Aug 2026

The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics

Results show that capability failures can manifest as distributed, task-dependent changes in the structure of visible reasoning, and that CoT dynamics agnostic to whether the verbalized trace reflects the model's internal computations can help diagnose and correct failures.

Shashwat Sourav, Aishwarya H. Balwani · 1 citation
Preprint Aug 2026

Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning

Experiments show that Code-as-World-VL achieves state-of-the-art performance on QuantiPhy and surpasses leading proprietary models, highlighting the potential of executable world representations as a scalable foundation for physical intelligence.

Han-Yang Wang, Yiyang Cai, Weiliang Chen et al. · 4 citations · ⚡1
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.