World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution,...
Jie Wu, Yu-Zhi Huang, Jun-Qi Liu et al.· 0 citations
Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor envi...
Deyi Zhu, Hao-Yu Fan, Yinan Zhu et al.· 2 citations
This work proposes LayoutDSL, a novel LLM-based framework for learning an interior layout policy in a domain-specific language (DSL) action space that substantially improves spatial plausibility and design logicality over strong baselines and existing methods.
Yuhao Lu, Weichen Zhang, Wenyi Xiao et al.· 0 citations
The first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception is introduced, and ActiveFly, a closed-loop UAV agent that integrates visual-language reasoning with fine-grained control, is developed and deployed on a physical UAV platform.
This work designs a structured Chain-of-Thought (CoT) framework that explicitly models 3D environmental perception to ensure robust spatial understanding and reasoning and introduces a novel RL algorithm featuring multi-objective process rewards and a tailored advantage estimation method, facilitating fine-grained cred...
Zile Zhou, Huining Yuan, Weichen Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.