Reliable Task Allocation for Multirobot System Using Safe Deep Reinforcement Learning
Abstract
This article addresses the critical reliability and safety challenges in multistation multirobot welding task allocation for automated manufacturing, where conventional task assignment approaches often overlook physical collision risks, leading to system failures and production downtime. Unlike existing methods that either ignore safety constraints or incorporate them heuristically as penalty terms in reward functions, we formulate the multirobot task allocation problem as a constrained Markov decision process (CMDP), enabling explicit and rigorous enforcement of operational safety constraints. Within this framework, we develop a safe deep reinforcement learning architecture that integrates a graph-based encoder for spatial feature extraction with a Lagrangian-based constrained policy optimization mechanism. This approach systematically balances the dual objectives of minimizing production cycle time and ensuring collision-free robot operations through adaptive safety-weight adjustment during training. Extensive experiments on both real-world welding cases and large-scale synthetic scenarios demonstrate that our method achieves zero safety violations while maintaining superior task efficiency, outperforming classical heuristics and conventional reinforcement learning baselines that exhibit notable collision rates. The results underscore the effectiveness of embedding safety as a formal constraint rather than an afterthought, providing a reliable and practical solution for industrial multirobot coordination.