World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrate...
Zi-Jian Jin, Yun-Bei Zhang, Yuan-Zhe Liu et al.· 0 citations
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introdu...
Yun-Bei Zhang, Zi-Jian Jin, Yuan-Zhe Liu et al.· 0 citations
ChainSWE is introduced, the first benchmark for evaluating agents on sequential, dependent bug fixes within a shared codebase, and reveals a consistent performance drop by up to 70% as the chain length increases.
Qirui Jin, Lingching Tung, Kenan Li et al.· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.