BeMapper is a novel evolutionary-augmented reinforcement learning framework that integrates a multi-agent bidirectionally-coordinated network (BicNet) with a distributed actor-critic architecture that incorporates success rate variance to penalize unstable behaviors and resolve credit assignment ambiguity.
Abstract
Despite recent advances in multi-agent path finding, achieving robust coordination in dynamic and crowded warehouse environments remains a bottleneck due to training instability and inefficient credit assignment. To address these challenges, we propose BeMapper, a novel evolutionary-augmented reinforcement learning framework that integrates a multi-agent bidirectionally-coordinated network (BicNet) with a distributed actor-critic architecture. Technically, our core novelty lies in three aspects: (1) A bidirectional feature fusion mechanism that enables agents to perceive collective spatial states beyond local observations; (2) An evolutionary-driven critic selection strategy that iteratively propagates high-performing models to accelerate convergence; (3) A multi-metric scoring system that incorporates success rate variance to penalize unstable behaviors and resolve credit assignment ambiguity. Extensive experiments demonstrate the superiority of BeMapper: it achieves a 98.66% mean success rate, outperforming state-of-the-art baselines Mapper (95.51%) and BicNet (93.78%) by 3.15%and 4.88%, respectively. Crucially, BeMapper yields a significantly higher average reward of 18.81, representing a relative improvement of 1.65 over Mapper and a substantial leap over BicNet's near-zero performance (0.04). Furthermore, in more crowded scenarios, BeMapper reduces the average travel steps to 36, being 5-9 steps shorter than competing methods, effectively enhancing operational throughput while ensuring robustness for large-scale industrial automation.
We present PRIMAL3, an ultra-large-scale learning-based framework for multi-agent pathfinding (MAPF) that integrates reinforcement learning, topology-aware communication, LaCAM3-guided training, and PIBT-based action refinement. PRIMAL3 targets failures at topologically critical states, where agents must coordinate dec...
Cheng-Yang He, T. Duhan, Gadiel Sznaier Camps et al.· 1 citation
Focusing on multi-agent path finding as an exemplary problem, this paper proposes to simplify two popular approaches to MAPF, namely multi-agent reinforcement learning and adaptive search, to enable seamless combination and transferability of methods without substantial engineering effort.
Thomy Phan· Proceedings of the Thirty-Fi...· 0 citations
The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.
X.-H. Fang, K. Chen, Cheng-Hao Ren et al.· Advanced Electromagnetics· 0 citations
This work introduces a unified RL formulation that jointly optimizes agent and environment policies, where the environment policy learns graph edge costs to provide global movement guidance via backward Dijkstra search and achieves significant improvements over the strong search-based planner, Causal-PIBT, across multi...
He Jiang, Jingtian Yan, Yulun Zhang et al.· 0 citations
M3W is a novel approach that applies mixture-of-experts (MoE) to world model instead of policy, enabling both learning and planning, and demonstrates superior performance, sample efficiency, and multi-task adaptability.
Zi-Jie Zhao, Zhong Zhao, Kaixuan Xu et al.· Neural Information Processin...· 9 citations
Cooperative multi-agent reinforcement learning (MARL) enables autonomous agents to coordinate in complex spatial environments. This study proposes a MARL framework for goal-directed navigation that integrates entangled state embeddings, copula-based joint action transformations, and a shared reward mechanism. Entangled...