World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure e...
Hao-Yi Jiang, Liu Liu, Xin-Jiang Wang et al.· 0 citations
AeroAct is the first WAM instantiated and demonstrated for real-world aerial flight, and closed-loop simulation and real-world experiments show that temporal visual context improves target tracking and object-search performance, and that WAM-based policies can be executed on a physical quadrotor.
SkillHEX is introduced, a closed-loop framework coupling hypothesis-driven self-verification with evidence-guided tree search that translates falsifiable failure hypotheses into executable tests, producing diagnostic evidence as dense reward without additional environment attempts.
Yuru Feng, Yaoqi Chen, Bei-Di Zhao et al.· 0 citations
This work proposes MESA (a Multi-structure Evidence Selection framework for long-horizon Agent), which builds five complementary structure views of each trajectory and learns from end-to-end answer-level feedback to select and fuse a query-specific subset for a frozen answer model.
Beidi Zhao, Yaoqi Chen, Yuru Feng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.