Experiments on the StarCraft Multi-Agent Challenge benchmark demonstrate that CLRE achieves superior win rates and sample efficiency compared with state-of-the-art MARL methods, showing that the proposed approach enables diverse and efficient cooperation among agents in complex environments.
Fair Multi-Level Preference Optimization (Fair-MPO or $\Phi$-MPO), a new preference optimization framework for agentic learning, is proposed and a Fair Multi-Level Objective that addresses imbalance in agentic learning is introduced.
Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill et al.· 0 citations
In complex, partially observable environments such as real-time strategy games, relying on decentralized multi-agent systems to learn efficient cooperative strategies is a challenging task. This paper proposes a Role Cognition Decomposition algorithm based on Multi-Agent Reinforcement Learning (RCD-MARL), which combine...
Multi-agent reinforcement learning (MARL) is impeded by the combinatorial explosion of joint action spaces and the neglect of complex inter-agent dependencies, which severely hinder effective coordination and accurate credit as signment. To address these challenges, this paper proposes a novel framework termed Lightwei...
Jun Wang, Wen-Jie Ke, Lu Liu et al.· Neural Networks· 0 citations
Many cooperative multi‐agent tasks are naturally defined by graph‐structured objectives, where agents must collectively reach, for example, a desired relational configuration or satisfy a set of constraints. These objectives often encode spatial arrangements, inter‐agent relations, or constraints that can be formal...
Alessandro Amato, Raffaele Galliera, K. Venable et al.· The AI Magazine· 0 citations
River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization is proposed, which achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks.
Yi-Fan Yao, Bo Pang, Xuan-Phi Nguyen et al.· 2 citations
Efficient cooperation is challenging due to the usual partial observability of each agent in multi-agent reinforcement learning. Recurrent networks encode local interaction histories, but their hidden representations provide limited insight into the information underlying individual decisions. To address these challeng...
Yi-Jie Sun, Sanquan Sun, Yan-Da Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.