Sep 2026· IEEE Transactions on Neural Networks and Learning Systems· Vol PP· 0 citations
Medicine
TL;DR
This work proposes a multiagent belief state computation method termed mutual awareness belief (MAB), which extends belief modeling to decentralized partially observable settings by incorporating agents' behavioral intentions and collaborative relationships and proves effective in the nonmonotonic Pursuit environment.
Abstract
In multiagent reinforcement learning (MARL), agents engage in collaborative or competitive interactions, adapting their policies to maximize rewards. However, in environments with only team rewards, agents cannot accurately perceive the rewards obtained from their individual actions. To address this limitation, we propose a multiagent belief state computation method termed mutual awareness belief (MAB). MAB extends belief modeling to decentralized partially observable settings by incorporating agents' behavioral intentions and collaborative relationships. This allows each agent to infer the contribution of individual behaviors to team rewards during decision-making and supports team-advantageous joint decisions. In addition, to address the issue of information lag, we develop progressive graph construction (PGC), a novel module for computing MAB. PGC constructs a dynamic graph with progressively updated node features and cooperation-informed edge weights. Through interaction with PGC, agents iteratively update and retrieve graph information to derive the MAB. Empirical results show that our method achieves competitive performance against baseline approaches on the StarCraft unit micromanagement benchmark and proves effective in the nonmonotonic Pursuit environment.
Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the...
Xing-Long Luo, Yu-Ding Zhang, Yu-Heng Kuang et al.· 0 citations
Efficient exploration remains a key challenge in cooperative multi-agent reinforcement learning (MARL), where the novelty of a joint situation may arise not only from unfamiliar individual observations but also from previously underrepresented interaction patterns among agents. Existing intrinsic-reward approaches comm...
Ya-Xin Xu, Yin-Xiang He, Tian-Yi Liu et al.· Applied Informatics· 0 citations
Cooperative multi-agent reinforcement learning (MARL) enables autonomous agents to coordinate in complex spatial environments. This study proposes a MARL framework for goal-directed navigation that integrates entangled state embeddings, copula-based joint action transformations, and a shared reward mechanism. Entangled...
Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comparative human or AI...
Shuo Liu, Xinzichen Li, Tian-Le Chen et al.· 0 citations
Action Generation with Topology Awareness (AGTA), a topology-aware sequential decision-making framework in MARL that integrates inter-agent correlation modeling with topology-guided decision-order optimization, and outperforms the state-of-the-art counterparts.
Kun Hu, Shanghua Wen, Wen-Di Wu et al.· Mathematics· 0 citations
Collaboration topology shapes both the performance and execution cost of LLM-based multi-agent systems. Because tasks differ in complexity and required capabilities, recent approaches generate task-specific collaboration graphs that specify agent participation and information flow. However, representative topology gene...
Kai-Rui Yang, Zi-Heng Yi, Xun-Kai Li et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…