Skip to content

Author

Weiming Zhi

We have 9 of 37 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

SAKI: Skill Assembly and Kinematic Imitation from Human Videos for Long-Horizon Mobile Manipulation

Learning from human videos offers a promising route to acquiring diverse manipulation skills. Extending this capability beyond tabletop settings to long-horizon mobile manipulation requires adapting and composing demonstrated interactions across changing scenes and robot configurations. We present Skill Assembly and Ki...

Yi-Jie Lu, James Zhao, Wei-Ming Zhi · 0 citations
Preprint Sep 2026

StereoPatch: Patch-Aligned RGB-Depth Fusion for Spatial Perception in Robot Manipulation

StereoPatch is introduced, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction and suggests that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction.

Ya-Nan Zhou, Zhao-Yan Qian, James Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation

Successful mobile manipulation requires coordinated base and arm motion while maintaining accurate spatial positioning. However, demonstration-trained policies can struggle to realise the intended base motion reliably, leading to spatial misalignment and subsequent manipulation failures. We present MAVP (Map-Aware Visu...

Jin-He Tang, Rui Dai, Wei-Ming Zhi · 0 citations
Preprint Sep 2026

Learning In-Hand Object Reaching to General 6D Poses

POISE (Palm-relative Object reaching In SE(3), a sim-to-real reinforcement learning framework for in-hand 6D object pose reaching, combines diverse stable-grasp initialization, goal- and geometry-conditioned control, an adaptive 6D goal curriculum, and a compact reward scheme for pose reaching and grasp preservation.

Jun-Xiao Lin, Tian-Yue Wu, Jie Yin et al. · 0 citations
Preprint Aug 2026

NestDex: Nested Policy Learning with Copilot Assisted Teleoperation for Dexterous Manipulation

Dexterous manipulation promises substantially richer robot interaction with the physical world, but learning these behaviours remains constrained by the difficulty of collecting consistent, complete-task demonstrations. Unlike parallel-jaw manipulation, dexterous tasks require the operator to coordinate arm motion with...

James Zhao, Jin-He Tang, Mingyuan Ba et al. · 1 citation
Preprint Aug 2026

AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies

Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single-step commands. Yet perception errors and execution drift can move the robot outside the demonstration distribution, while the policy continues to produce smooth action c...

Jin-He Tang, Weiming Zhi · 1 citation

TRACE: Trajectory-Routed Causal Memory for Delayed-Evidence Visuomotor Imitation

This work introduces TRAjectory-routed Causal Evidence (TRACE), a memory framework for visuomotor imitation policies that stores task-relevant visual and robot-state evidence in a fixed-size latent memory that remains bounded over long episodes.

Zihao Li, Ran-Peng Qiu, Yin-Cong Chen et al. · 1 citation
Review Jul 2026

Tri-Manual Visuomotor Imitation Learning of Robot Policies

TriManPolicy is presented, a tri-manual imitation learning system that allows one operator to demonstrate behaviours for three arms while reconsidering when they occur, and policies trained on demonstrations retimed by DATS exhibit more efficient coordination while maintaining comparable observed task success.

James Zhao, Mingyuan Ba, Weiming Zhi · 1 citation
2025

Building 3D Representations and Generating Motions From a Single Image via Video-Generation

This work proposes a framework known as Video-Generation Environment Representation (VGER), which leverages the advances of large-scale video generation models to generate a moving camera video conditioned on the input image, and demonstrates its ability to produce smooth motions that account for the captured geometry...

Weiming Zhi, Ziyong Ma, Tianyi Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.