A deep analysis of MARL applied to industrial multi-robot systems based on a systematic review is presented, with particular focus on cooperative manipulation tasks.
Abstract
Modern manufacturing faces increasing demands for flexibility, customization, and productivity under dynamic conditions. Multi-robot systems offer a promising solution by enabling cooperative execution of complex tasks, such as assembly and cooperative manipulation. In this context, Multi-Agent Reinforcement Learning (MARL) has emerged as a promising paradigm to enhance coordination and adaptability in industrial settings. MARL enables multiple agents to learn and interact in shared environments to achieve common goals within complex and dynamic industrial processes. In this paper, a deep analysis of MARL applied to industrial multi-robot systems based on a systematic review is presented, with particular focus on cooperative manipulation tasks. Following PRISMA guidelines, we analyze a total of 30 articles published between 2016 and 2026, selected independently by two of the authors from an initial pool of 102 records retrieved from Scopus and Web of Science. These articles were used to address five key questions regarding MARL algorithms, control architectures, industrial applications and validation practices. These research questions seek to examine gaps and trends at the research level which are important for the development of multi-agent control technologies. This review shows a clear prevalence of model-free algorithms under Centralized Training with Decentralized Execution (CTDE) architectures, with validation mainly performed in simulation. Despite promising results and high potential for impact, critical gaps remain in scalability, reproducibility, and sim-to-real transfer, limiting real deployment in manufacturing environments. To address these challenges and fill current gaps, we outline actionable research directions, such as hybrid MARL approaches, standardized industrial benchmarks, digital twin pipelines, and safety-aware deployment strategies, to accelerate MARL adoption in industrial environments.
Reinforcement learning (RL) has shown considerable promise for robotic decision-making, yet deploying multi-agent RL (MARL) on physical multi-robot systems in industrial environments remains challenging. This paper investigates the real-world applicability of decentralized MARL for multi-robot multi-machine tending. We propose Feature-fusion Multi-Agent Proximal Policy Optimization (FMAPPO), which fuses 2D LiDAR measurements with task-specific state information to enable safe decentralized multi-robot task assignment and navigation. A complete simulation-to-reality pipeline was developed using high-fidelity robotic simulation and ROS2 and deployed on physical mobile-manipulator platforms operating under realistic real-world conditions, with the robotic arms disabled during the experiments. We further investigate the sensitivity of the learned policy to command update frequency, an important consideration for real-world deployment. Comparative evaluation in simulation demonstrated that FMAPPO significantly outperformed state-of-the-art baselines with a large effect size, achieving improvements of 106\% and 21\% in parts delivery and 48\% and 11\% in parts collection over MAPPO and SMAPPO, respectively. FMAPPO also increased machine utilization by 31 and 10 percentage points, respectively, while reducing collisions by 18\% and 15\% and increasing the safety score by 14 and 6 percentage points compared with MAPPO and SMAPPO, respectively. Furthermore, real-world experiments demonstrated that the learned decentralized policies can coordinate multiple robots to service multiple machines while maintaining safe operation under real-world sensing and control constraints. Videos of the real-world experiment are available online https://anonymouspapers123.github.io/FMAPPO/.
A. Abdalwhab, Giovanni Beltrame, David St-Onge· 0 citations
Flexible manufacturing, characterized by high-mix, low-volume, and highly variable production, demands robotic systems with strong adaptability, dexterity, and intelligence that conventional offline-programmed industrial robots cannot provide. This paper presents a systematic review of key technologies for robot embodied intelligence oriented toward flexible manufacturing, organized around the closed loop of perception, decision-making, and execution. The purpose is to clarify the current research landscape, identify core technical bottlenecks, and outline future directions for embodied-intelligent manufacturing. Adopting a literature-analysis and comparative-review method, the study examines representative advances at three levels: multimodal environmental perception and real-time modeling, flexible adaptive precision manipulation, and intelligent decision-making for process planning and scheduling. The review finds that multimodal fusion and semantic SLAM are overcoming perception bottlenecks, that deep learning and force/position hybrid control are balancing flexible adaptability with high-precision operation, and that deep reinforcement learning and large models are advancing intelligent process planning. It concludes that data scarcity, model reliability, software-hardware integration, and ethical-legal standards remain the principal challenges to large-scale industrial deployment.
Zheng-Yang Chen· Advances in Engineering Inno...· 0 citations
A data-driven scoping review of 130 studies published between 2020 and 2026, following PRISMA-ScR guidelines, to systematically map the landscape of long-horizon RL for robotic manipulation and presents a gap atlas that identifies underexplored research directions across methodological and experimental dimensions.
Matthew Acs, Xiangnan Zhong· Discover Robotics· 0 citations
It is concluded that MARL is a promising solution for future intelligent collaborative robotics and highlights future research directions including federated reinforcement learning, explainable AI, edge-based robotic intelligence, and adaptive swarm robotics for Industry 4.0 applications.
Suresh Babu Reddy, Anita Verma· International Journal of Int...· 0 citations
In complex settings like smart manufacturing and human-robot teamwork, robots face internal disturbances and external uncertainties. Traditional control methods depend on accurate models and manual parameter tuning, leading to complex adjustment, weak dist urbance rejection, and poor generalization. Deep reinforcement learning (DRL) combines deep learning's feature extraction with reinforcement learning's sequential decision-making. Through end-to-end learning, it removes the need for exact system models and has become a key approach for robot adaptive control. This paper systematically reviews DRL-driven robot adaptive control. It first outlines core concepts and theory, building the technical framework that brings together DRL and adaptive control. It then analyzes mainstream DRL algorithm improvements and hybrid methods for adaptive control, explores main issues in Sim-to-Real transfer, and discusses safe DRL control modeling under constraints. The paper also introduces typical robot applications, examines current challenges and bottlenecks, and points out future trends. This review aims to clarify the development path of DRL-driven robot adaptive control, offering reference for theory, algorithm design, and engineering practice. It is hoped that this survey will help new researchers quickly grasp the landscape of the field.
Bipedal robots have gained a lot of attention in robotics because of their versatility in numerous application areas such as in rescue missions, military tasks, therapy, personal assistance and care for the aged. Their ability to move in tough and mixed-up spaces provides essential advantages over wheeled robots. However, they present challenges, such as complicated movement, balancing and control issues, and energy management problems. This paper utilized PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) to examine peer-reviewed publications in bipedal robots, covering both their modeling dynamics and strategic control, application areas and the challenges in development and management. After reviewing 165 of 15,214 publications that satisfied the inclusion criteria, the assessments were developed around the history of bipedal robots, marking important developments like dynamic control, passive-dynamic walking, and using sensors in real-time. Developments in artificial intelligence, machine learning, and reinforcement learning have made bipedal robots more stable, adaptable, and efficient. The novelty of this review lies in integrating rigid-body and reduced-order dynamic models, classical and learning-based control strategies, hardware limitations, and future research priorities within a unified comparative framework for bipedal robotics.
B. Kommey, E. Tamakloe, Safianu Umar et al.· AVITEC· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.