Skip to content

MAB-PGC: Multiagent Mutual Awareness Belief via Progressive Graph Construction.

Sep 2026 · IEEE Transactions on Neural Networks and Learning Systems · Vol PP · 0 citations
Medicine

TL;DR

This work proposes a multiagent belief state computation method termed mutual awareness belief (MAB), which extends belief modeling to decentralized partially observable settings by incorporating agents' behavioral intentions and collaborative relationships and proves effective in the nonmonotonic Pursuit environment.

Abstract

In multiagent reinforcement learning (MARL), agents engage in collaborative or competitive interactions, adapting their policies to maximize rewards. However, in environments with only team rewards, agents cannot accurately perceive the rewards obtained from their individual actions. To address this limitation, we propose a multiagent belief state computation method termed mutual awareness belief (MAB). MAB extends belief modeling to decentralized partially observable settings by incorporating agents' behavioral intentions and collaborative relationships. This allows each agent to infer the contribution of individual behaviors to team rewards during decision-making and supports team-advantageous joint decisions. In addition, to address the issue of information lag, we develop progressive graph construction (PGC), a novel module for computing MAB. PGC constructs a dynamic graph with progressively updated node features and cooperation-informed edge weights. Through interaction with PGC, agents iteratively update and retrieve graph information to derive the MAB. Empirical results show that our method achieves competitive performance against baseline approaches on the StarCraft unit micromanagement benchmark and proves effective in the nonmonotonic Pursuit environment.

View source

Similar papers

#machine learning Preprint Sep 2026

HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning

Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the...

Xing-Long Luo, Yu-Ding Zhang, Yu-Heng Kuang et al. · 0 citations
Open access Sep 2026

Exploration by Multi-Agent Relationship Graph Reconstruction

Efficient exploration remains a key challenge in cooperative multi-agent reinforcement learning (MARL), where the novelty of a joint situation may arise not only from unfamiliar individual observations but also from previously underrepresented interaction patterns among agents. Existing intrinsic-reward approaches comm...

Ya-Xin Xu, Yin-Xiang He, Tian-Yi Liu et al. · 0 citations
#reinforcement learning Open access Sep 2026

Cooperative multi-agent reinforcement learning with entangled state representations and copula-based action coordination for urban navigation

Cooperative multi-agent reinforcement learning (MARL) enables autonomous agents to coordinate in complex spatial environments. This study proposes a MARL framework for goal-directed navigation that integrates entangled state embeddings, copula-based joint action transformations, and a shared reward mechanism. Entangled...

Jong-Min Kim · 0 citations
#artificial intelligence Preprint Sep 2026

Improving LLM Collaboration via Multi-Agent Preference Learning

Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comparative human or AI...

Shuo Liu, Xinzichen Li, Tian-Le Chen et al. · 0 citations
Open access Aug 2026

AGTA: Topology-Aware Sequential Decision-Making in Multi-Agent Reinforcement Learning

Action Generation with Topology Awareness (AGTA), a topology-aware sequential decision-making framework in MARL that integrates inter-agent correlation modeling with topology-guided decision-order optimization, and outperforms the state-of-the-art counterparts.

Kun Hu, Shanghua Wen, Wen-Di Wu et al. · 0 citations
#machine learning Preprint Sep 2026

MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning

Collaboration topology shapes both the performance and execution cost of LLM-based multi-agent systems. Because tasks differ in complexity and required capabilities, recent approaches generate task-specific collaboration graphs that specify agent participation and information flow. However, representative topology gene...

Kai-Rui Yang, Zi-Heng Yi, Xun-Kai Li et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.