This study evaluates how architectural heterogeneity and reward design influence coordination in decentralized IPPO-based Multi-Agent Reinforcement Learning (MARL) systems operating in Search-and-Rescue environments with strict sequential task dependencies and provides empirically grounded design recommendations for cooperative multi-agent systems.
Abstract
This study evaluates how architectural heterogeneity and reward design influence coordination in decentralized IPPO-based Multi-Agent Reinforcement Learning (MARL) systems operating in Search-and-Rescue (SAR) environments with strict sequential task dependencies. Controlled experiments compare homogeneous and heterogeneous teams across multiple reward structures, observation ranges, task complexities, and asymmetric role distributions in a partially observable, communication-denied, decentralized environment. The results show that homogeneous teams consistently achieve faster convergence, higher coordination efficiency, and more stable cooperative behavior than balanced heterogeneous configurations. Under hybrid global–local reward schemes, homogeneous teams achieve a higher success rate than local-only reward schemes. A dedicated test confirms this advantage stems from team composition itself, not parameter sharing: it persists even under independent, non-shared policy networks. Additional experiments suggest that asymmetric role allocation may influence coordination dynamics in heterogeneous teams. The study provides empirically grounded design recommendations for cooperative multi-agent systems, derived from controlled simulation experiments under communication-limited sequential coordination constraints.
Cooperation emergence is a central problem in multi-agent systems because decentralized agents must coordinate while adapting to the changing behavior of others. Evolutionary game theory identifies strategically stable outcomes, but stability under a population adjustment dynamic need not imply that finite-sample learn...
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory t...
Does multi-agent LLM interaction help or hurt? Some work reports gains from debate (Du et al., 2024), critique loops (Chen et al., 2025), and mixture-of-agents synthesis (Wang et al., 2025), while other work finds that interaction adds cost without improving quality under equal budgets (Tran&Kiela, 2026; Xu et al., 202...
Summer Eunhyung Ann, Haokun Liu, Chen-Hao Tan· 3 citations
X-CODE is an explainable offline MARL that operates offline without environmental interaction, nor inter-agent communication, nor inter-agent communication, and exploits explainability-aware reward shaping to modify the relative preference among joint offline transitions during centralized training to improve decentral...
In complex, partially observable environments such as real-time strategy games, relying on decentralized multi-agent systems to learn efficient cooperative strategies is a challenging task. This paper proposes a Role Cognition Decomposition algorithm based on Multi-Agent Reinforcement Learning (RCD-MARL), which combine...
Cooperative multi-agent reinforcement learning (MARL) enables autonomous agents to coordinate in complex spatial environments. This study proposes a MARL framework for goal-directed navigation that integrates entangled state embeddings, copula-based joint action transformations, and a shared reward mechanism. Entangled...