CHOREO, a framework for training-free composition of heterogeneous humanoid skills, demonstrates that executable trajectories provide a scalable interface for accumulating and composing pretrained humanoid capabilities.
Abstract
Recent advances in humanoid robotics have produced diverse skills through reinforcement learning, motion imitation, and generative modeling. Yet these capabilities remain siloed because they are built around incompatible representations, interfaces, and controllers. We present CHOREO, a framework for training-free composition of heterogeneous humanoid skills. Our key observation is that, regardless of how a skill is learned, it can ultimately be expressed as an executable motion trajectory. Based on this observation, CHOREO converts each capability into SkillMotion, a unified representation that combines motion states, contacts, semantics, and boundary conditions. Skills are composed through direct continuation, cubic Hermite blending, or validated bridge motions, without retraining source models or updating models at test time. On Unitree G1 in MuJoCo, CHOREO organizes 2,950 admitted SkillMotion assets derived from heterogeneous sources and achieves 95.4\% sequence success across 130 multi-action tasks, including 93.8\% success on eight-action sequences. These results demonstrate that executable trajectories provide a scalable interface for accumulating and composing pretrained humanoid capabilities.
Robotic foundation models offer a promising path toward general-purpose humanoid robot control, often through hierarchical architectures. However, their effectiveness depends on the command interface between the planner and the controller, which must support accurate execution while remaining easy to predict, and ideal...
Fei-Yang Wu, Chen-Xiao Gao, Chen Yang et al.· 0 citations
Humanoid robots can acquire complex skills by imitating kinematic humanoid motion references, yet reliable references for contact-rich interactions remain difficult to obtain: motion capture deteriorates under occlusion and close physical contact, while retargeting introduces additional contact and geometric inconsiste...
Lalit Jayanti, Kashu Yamazaki, Yuto Shibata et al.· 0 citations
WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token is introduced and its data-scaling and temporal-context behavior under the tested recipes are characterized.
Chunkai Yang, An-Dong Yang, Di Huang et al.· 0 citations
SkillX is presented, a unified reinforcement learning framework that learns and composes multiple atomic soccer skills through a single command-conditioned policy, enabling the robot to execute atomic skills and transition among them such as dribbling, trapping, and shooting.
Zhang-Chen Ye, En-Xuan Ruan, Yi-Fei Bao et al.· 2 citations
Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and...
Shenghe Zheng, Wen-Bo Li, Ji-Yao Zhang et al.· 0 citations
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configur...
Hongyu Li, Bo-Wen Wen, Xing-Hao Zhu et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.