Skip to content

MOP: A Multimodal Object-Aware Policy for Robotic Manipulation via Geometry-Guided Fusion and Trajectory Prediction

Sep 2026 · IEEE Robotics and Automation Letters · Vol 11, pp. 10114-10121 · 0 citations · 49 references

Abstract

3D imitation learning has demonstrated capability in diverse visuomotor tasks but often struggles with small objects or high-precision manipulation due to geometric sparsity. To address this, we propose the Multimodal Object-aware Policy (MOP), a novel framework that incorporates a geometry-guided fusion module to adaptively integrate 2D semantic features with 3D geometry for precise control. Additionally, we introduce a lightweight Object Position Prediction (OPP) module that serves as an auxiliary supervision signal. The training labels for this module are generated by using a cost-effective vision-based software tracker, effectively replacing expensive hardware-based motion capture systems. We evaluate our policy across 56 tasks on 3 simulation benchmarks. Experimental results demonstrate that MOP significantly outperforms baselines, achieving higher success rates with lower variance and efficient inference. Real-world experiments on 4 tasks further validate the robustness and transferability of our approach.

View source