This work introduces Causal-Plan-Bench, a high-fidelity diagnostic suite spanning four causal dimensions, curated via multi-stage verification, and initiate the first effort to turn agents from superficial token predictors into physically grounded causal reasoners, bridging language modeling and world modeling.
Zi-Rong Song, Zheng Lu, Ming-Qi Gao et al.· 3 citations
This work introduces Competitive Memory Readout, which explicitly incorporates same-class competitor evidence when retrieving target information from memory when retrieving target information from memory, and applies a lightweight adaptive restoration rule after competition.
The model achieves the best performance among similarly sized models on 19 of the 38 benchmarks and substantially outperforms strong competitors, including Qwen3.6-A3B and Cosmos 3.5 MoT-2B, and demonstrates strong performance on embodied agentic tasks requiring multi-turn interaction and long-horizon reasoning.