Skip to content
Open access

Decentralized Model-Based ACKTR for Large-Scale Multi-Agent Path Planning Under Partial Observability

Aug 2026 · Electronics · 0 citations · 27 references

TL;DR

This work formulate large-scale MAPP as a partially observable networked Markov decision process as a decentralized model-based Actor-Critic using the Kronecker-factored trust region (DM-ACKTR) algorithm, which consistently obtains the highest TCR and lowest CR.

Abstract

Multi-agent path planning (MAPP) under partial observability requires agents to coordinate their movements and complete tasks efficiently without access to global information. The planning space and coordination complexity grow rapidly with increasing numbers of agents, targets, and obstacles. We formulate large-scale MAPP as a partially observable networked Markov decision process. Based on this formulation, we propose a decentralized model-based Actor-Critic using the Kronecker-factored trust region (DM-ACKTR) algorithm. The algorithm integrates local model learning with ACKTR-based policy optimization in an independent learning architecture. Each agent learns a local model to predict the next observation and reward. These predictions are used to construct additional transitions for Actor and Critic updates. A neighborhood-based communication mechanism incorporates information from nearby agents into value estimation. Region partitioning reduces each agent’s effective planning space. These improvements enable DM-ACKTR to continue outperforming the baseline algorithms as the scale of the MAPP problem increases. Experiments across three training and five evaluation scenarios show that DM-ACKTR achieves the best overall performance. Among the five evaluated algorithms, it consistently obtains the highest TCR and lowest CR, improving TCR by 2.06–4.35% and reducing CR by 11.26–25.95% relative to the respective best baselines.

Read PDF

Similar papers

Preprint Aug 2026

MDGAM-Based Cooperative Task Scheduling for Communication-Constrained Distributed Multi-Agent Systems

A neural scheduling framework for distributed multi-robot task allocation, consisting of a multi-decoder graph attention model (MDGAM) policy model and a critic-free group relative multi-agent policy gradient (GRMAPG) training algorithm, which improves task-completion performance over existing heuristic and learning-ba...

Licheng Wang, Ming-Tao Huang, Yuan Shen · 0 citations
Preprint Aug 2026

Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations

This work introduces a unified RL formulation that jointly optimizes agent and environment policies, where the environment policy learns graph edge costs to provide global movement guidance via backward Dijkstra search and achieves significant improvements over the strong search-based planner, Causal-PIBT, across multi...

He Jiang, Jingtian Yan, Yulun Zhang et al. · 0 citations
Preprint Aug 2026

Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration

A Planner-Conditioned Diffusion Policy (PCDP) is proposed, trained on demonstrations from multiple planner styles with planner identity as an explicit conditioning input, enabling a single shared model to learn a multimodal trajectory distribution and generate diverse, controllable trajectory candidates from the same o...

M. Teo, Jeric Lew, T. Duhan et al. · 0 citations
Open access Jul 2026

A Belief-Driven Hybrid Reinforcement Learning Framework for Decentralized Multi-Robot Navigation Under Partial Observability

Decentralized multi-robot navigation is difficult when robots must act from local observations without centralized coordination or explicit inter-robot communication. A belief-driven hybrid reinforcement learning framework is evaluated for planar multi-robot navigation under partial observability. Each robot builds a c...

V. Malathi, Pramod Sreedharan, Rthuraj Puthiyaveedu Rajesh et al. · 0 citations
Open access Aug 2026

Collision-aware cooperative multi-UAV path planning with hierarchical PPO-LSTM

The results indicate that separating waypoint-level strategy from recurrent local execution improves mission reliability and collision avoidance in the tested grid environments, while larger random-map benchmarks, fully controlled MAPPO/QMIX comparisons, and continuous 3-D simulation remain important future work.

Alparslan Güzey · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.