Skip to content
Conference

SPD-MAPPO: Reinforcement Learning With Stochastic Policy Distillation for Multi-Vehicle Coordination in Open-pit Mine

Jul 2026 · 2026 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM) · pp. 1-6 · 0 citations · 17 references

Abstract

Coordination of multiple autonomous trucks is crucial for enhancing the efficiency and safety of modern mining, yet it is challenged by dynamic vehicle-to-vehicle interactions and the complexity of mining transportation. Conventional rule-based methods struggle to balance efficiency with success rates and lack flexibility in diverse scenarios. While multi-agent reinforcement learning (MARL) shows great promise for cooperative tasks, its application in real world is often hampered by challenges in convergence. To address these challenges, we propose SPD-MAPPO, a novel multi-stage learning framework that transfers expertise from imitation learning(IL) to cooperative ability in MARL. The framework first employs IL to pre-train a policy with basic single-agent driving ability, which is subsequently refined for cooperative behaviors through MARL. Specifically, we design a Stochastic Policy Distillation (SPD) mechanism to bridge the gap between single-agent expertise and multi-agent coordination, and a multi-head critic network to achieve more precise credit assignment. We validate our method in a high-fidelity simulator with a real-world map of mine and a truck dynamics model. Our method outperforms typical rule-based and MARL methods in success rate, efficiency, and operational accuracy.

View source