Skip to content

Author

Kai-Wen Zheng

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

SPD-MAPPO: Reinforcement Learning With Stochastic Policy Distillation for Multi-Vehicle Coordination in Open-pit Mine

Coordination of multiple autonomous trucks is crucial for enhancing the efficiency and safety of modern mining, yet it is challenged by dynamic vehicle-to-vehicle interactions and the complexity of mining transportation. Conventional rule-based methods struggle to balance efficiency with success rates and lack flexibility in diverse scenarios. While multi-agent reinforcement learning (MARL) shows great promise for cooperative tasks, its application in real world is often hampered by challenges in convergence. To address these challenges, we propose SPD-MAPPO, a novel multi-stage learning framework that transfers expertise from imitation learning(IL) to cooperative ability in MARL. The framework first employs IL to pre-train a policy with basic single-agent driving ability, which is subsequently refined for cooperative behaviors through MARL. Specifically, we design a Stochastic Policy Distillation (SPD) mechanism to bridge the gap between single-agent expertise and multi-agent coordination, and a multi-head critic network to achieve more precise credit assignment. We validate our method in a high-fidelity simulator with a real-world map of mine and a truck dynamics model. Our method outperforms typical rule-based and MARL methods in success rate, efficiency, and operational accuracy.

Kai-Wen Zheng, Yafei Wang, Yichen Zhang et al. · 0 citations