Skip to content

H3C-BEACON: hierarchical hybrid heterogeneous control with Bayesian-elite adaptive coalition network for multi-agent reinforcement learning

Sep 2026 · Complex & Intelligent Systems · 0 citations
Reinforcement Learning in Robotics

TL;DR

H3C-BEACON is presented, a unified hierarchical framework for coordination and control in cooperative MARL that jointly integrates communication, probabilistic belief inference, adaptive coalition formation, and policy stabilisation and significantly improves reproducibility across random initialisations.

Abstract

Multi-agent reinforcement learning (MARL) in partially observable and non-stationary environments requires agents to simultaneously infer hidden states, coordinate under limited communication, and learn stable cooperative policies. While existing approaches typically address these challenges through independent mechanisms, their interactions remain insufficiently exploited. We present H3C-BEACON ( Hierarchical Hybrid Heterogeneous Control with Bayesian-Elite Adaptive Coalition Network ), a unified hierarchical framework for coordination and control in cooperative MARL that jointly integrates communication, probabilistic belief inference, adaptive coalition formation, and policy stabilisation. The framework combines six complementary components: (i) Dynamic Graph Attention Networks (DGAT) for distance-aware communication, (ii) Bayesian belief fusion for hidden-state estimation under partial observability, (iii) spectral coalition formation for adaptive role specialisation, (iv) a dual-critic architecture that separates global coordination from local decision making, (v) RTD++ elite-trajectory anchoring to improve policy optimisation stability, and (vi) bounded entropy control to balance exploration and exploitation. We evaluate H3C-BEACON on three widely used cooperative MARL benchmarks spanning communication-intensive, coordination, and imperfect-information settings. On the Multi-Agent Particle Environments, H3C-BEACON consistently improves coordination quality over MAPPO, achieving a perfect win rate across all five independent random seeds on while improving the best episode reward from $$-6.06\pm 0.70$$ - 6.06 ± 0.70 to $$-2.35\pm 0.62$$ - 2.35 ± 0.62 . On , the framework substantially reduces performance variability, producing a $$95\%$$ 95 % confidence interval approximately $$28\times $$ 28 × narrower than MAPPO ( $$\pm 0.57$$ ± 0.57 versus $$\pm 15.90$$ ± 15.90 ), indicating significantly improved reproducibility across random initialisations. Ablation experiments indicate that each architectural component contributes to overall performance, with win-rate reductions ranging from $$28\%$$ 28 % (−DGAT) to $$70\%$$ 70 % (−RTD++ or −Coalitions) when individual modules are removed. Under partial observability in Hanabi-full, H3C-BEACON increases the mean score from $$2.29\pm 0.23$$ 2.29 ± 0.23 to $$3.96\pm 0.82$$ 3.96 ± 0.82 ( $$+73\%$$ + 73 % ) while avoiding policy collapse across all runs, suggesting that RTD++ provides effective stabilisation during cooperative learning. Although MAPPO remains superior on StarCraft combat scenarios, a result consistent with the structural characteristics of that environment (homogeneous units, dense global state, and no explicit communication channel that would benefit from DGAT or coalition formation), the overall results indicate that jointly modelling communication, belief estimation, adaptive coalition formation, and stable optimisation provides a robust and effective framework for cooperative MARL in environments characterised by partial observability and decentralised coordination. All primary results use 5 independent random seeds with $$95\%$$ 95 % confidence intervals ( $$t_{0.975,4}{=}2.776$$ t 0.975 , 4 = 2.776 ).

Read PDF

Similar papers

SCRAMBLE: Safe Cooperative Risk-Aware Multi-Agent Belief Learning for Embodied Reinforcement Learning in Partially Observable Dynamic Environments

Safe decision-making in partially observable dynamic environments remains a fundamental challenge for multi-agent embodied reinforcement learning, where agents must cooperate with limited local observations while responding to evolving environmental changes and safety-critical interactions. Existing methods often empha...

Bo-Zhi Zhang, Jiang-Bo Wang, Tian Jing · 0 citations
Open access Aug 2026

Directional Pheromone Gradient Observations for Decentralized Multi-Agent Reinforcement Learning in Swarm Drone Search and Rescue

Findings indicate that directional pheromone-gradient observations provide an effective and communication-efficient mechanism for decentralized swarm coordination, improving search effectiveness and operational robustness in post-disaster SAR scenarios.

Peter Yacoub, Mohamed Malek Kaouach, Esraa Khatab et al. · 0 citations
Open access Aug 2026

Decentralized Model-Based ACKTR for Large-Scale Multi-Agent Path Planning Under Partial Observability

This work formulate large-scale MAPP as a partially observable networked Markov decision process as a decentralized model-based Actor-Critic using the Kronecker-factored trust region (DM-ACKTR) algorithm, which consistently obtains the highest TCR and lowest CR.

Ye-Min Liu, Jin-Hao Yang, Xiang-Yu Ma et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Attention-based Hierarchical Variational Information Bottleneck for Robust Multi-Agent Communication under Variable Bandwidth

Learning-based multi-agent communication under limited bandwidth does not only require deciding what to communicate, but also structuring messages so that partial transmissions remain useful. We study this problem under prefix truncation, where only the first part of each message is received. To address it, we propose...

Lukas Koch Vindbjerg, Qi Zhang, Yury Brodskiy et al. · 0 citations
Conference Aug 2026

Heterogeneous Multi-Agent Autonomous Learning and Safe Cooperative Decision-Making

Unmanned surface and underwater vehicles face challenges in autonomously learning cooperative encirclement for high-value targets under partial observability, intermittent communication, and collision risks. This paper proposes a heterogeneous multi-agent reinforcement learning framework with safe decisionmaking. The f...

Jiang-Li Cao, Chao Liu, Guo-Ping Zhang · 0 citations
Preprint Sep 2026

Fully Decentralized and Safety-Aware Multi-Agent Reinforcement Learning for Control on Networks

This paper develops a safe and fully decentralized multi-agent reinforcement learning (MARL) algorithm to solve a class of discrete-time control problems on networks, including the persistent monitoring problem. Fully decentralized control of agents, while offering numerous benefits, faces issues such as exponentially...

T. Rogalski, Shirantha Welikala · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.