General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physic...
Zhiqin Yang, Chen-Xin Li, Xiao-Meng Hu et al.· 0 citations
Robot foundation policies predict action chunks, but how many actions to execute before replanning depends on the current task phase. We introduce ChunkTrust, which treats the execution horizon as a latent variable inferred from action-expert evidence rather than a fixed hyperparameter. Its training-free Action-aware H...
Fan-Ding Huang, Jing-Yan Jiang, Shi-Feng Bao et al.· 0 citations
This report presents HyMobileAgent, a mobile GUI agent built on Hy3.0-VL-A3B, a vision-native foundation model featuring native any-resolution input, an A3B-scale deployment budget, and a 32K context window to model extended interaction histories.
H. Team, Hua-Wen Shen, Zheng-Yang Tang et al.· arXiv.org· 2 citations· ⚡1
JarvisHub is introduced, a canvas-native creative agent harness for long-horizon multimodal creation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.