Behavior Foundation Models (BFMs) give humanoids a promptable policy over a latent behavior space, enabling one single vector to represent a motion to imitate, a pose to reach, or a reward to maximize. Forward-Backward representations successfully produce such spaces, but at the cost of hundreds of GPU-hours for a sing...
Tan-Dzung Do, Tuan Dat Phuong, Nico Bohlinger et al.· 0 citations
Whole-body compliant control is essential for deploying heavy humanoids under high payload in human-centric environments. Most prior force-aware learning-based pipelines focus on end-effector resistance, per-link upper-body springs, or end-effector stiffness modulation, leaving arbitrary-site perturbations on heavy pla...
Tan-Dzung Do, Cuc T.Trinh, Tuan Dat Phuong et al.· 1 citation
Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In particular, future obs...
Khang Nguyen, Hoang Pham Quang Nguyen, Ha Phuong Nguyen et al.· 0 citations
Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between policy queries. We present vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization. We relate query lat...
Khanh Duy Nguyen, Hoang M. Truong, A. T. Le· 0 citations
Global Tensor Motion Planning (GTMP) solves motion planning with batched tensor operations over a layered multipartite graph. We generalize GTMP so that adjacent-layer edges are realized by any black-box local planner (e.g., linear interpolation, splines, sampling-based planning, trajectory optimization, or generative...
Sai Coumar, A. T. Le, Zachary Kingston· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.