Skip to content
#reinforcement learning Review Open access

Learning Together and Against Each Other: How Multiple Reinforcement-Learning Agents Coordinate, Compete and Where the Field is Heading

Sep 2026 · International Journal of Innovative Science and Research Technology · 0 citations · 26 references
Reinforcement Learning in Robotics

TL;DR

The obstacles that distinguish MARL from its singleagent counterpart are discussed, namely a moving-target learning problem, the difficulty of dividing a shared reward among team members, limited local views, and growth of the joint action space, and why training with global information but acting on local information has become the standard design.

Abstract

When several self-directed learners share one environment, each must improve its own behavior while the others are also changing theirs. This is the setting studied by multi-agent reinforcement learning (MARL), and the agents involved may be working toward one goal, pulling in opposite directions, or some combination of the two. In this review we first lay out the mathematical objects used to describe such interaction, namely Markov games and their decentralized, partially observed variants, and then group the algorithmic literature into three lineages: agents that learn in isolation, methods that factor a team value into per-agent pieces (VDN, QMIX and their descendants), and actor-critic schemes that train against a critic with global knowledge (MADDPG, COMA, MAPPO). We discuss the obstacles that distinguish MARL from its singleagent counterpart, namely a moving-target learning problem, the difficulty of dividing a shared reward among team members, limited local views, and growth of the joint action space, and explain why training with global information but acting on local information has become the standard design. Uses of MARL in real-time strategy games, robot teams, automated vehicles, wireless resource sharing, and the orchestration of language-model agents are reviewed alongside the test suites the community relies on. We close by pointing to unresolved questions concerning scale, cooperation with unfamiliar partners, and deployment safety.

Read PDF

Similar papers

#large language models Review Open access Sep 2026

A Survey on Reinforcement Learning Optimization Methods for Multi-Agent Collaboration of Large Language Models

The integration of reinforcement learning (RL) into the optimization of multi-agent collaboration for Large Language Models (LLMs) is an important combination of two advanced areas, Multi-Agent Systems (MAS) and LLMs. This paper thoroughly examines the main approaches, evaluation standards, recent progress and existing...

Qian-Ling Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

Improving LLM Collaboration via Multi-Agent Preference Learning

Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comparative human or AI...

Shuo Liu, Xinzichen Li, Tian-Le Chen et al. · 0 citations
Preprint Aug 2026

IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents

Isolated Bilateral Reinforcement Learning (IB-RL), in which the two roles coevolve through joint rollouts while each role optimizes its own reward through fully independent advantages, action masks, and update paths, produces policies that generalize more effectively to unseen counterparts.

Senhao Wang, Chenghao Cai, Hai-Tao Hu et al. · 0 citations
Preprint Aug 2026

Training Small LLMs as Spatial Multi-Agent Policies

Training LLM-based multi-agent systems with multi-agent reinforcement learning with multi-agent reinforcement learning is rapidly gaining traction, and a parallel line of work argues that such systems should be judged by their behavior, not only their reward.

Yi Mao, Andrew Perrault · 0 citations
Preprint Sep 2026

Certifying cooperation: a novel approach to cooperative multi-agent task generation

A shared reward gives agents a common objective, but leaves open when, how and even whether they must cooperate to succeed. We address these questions in the Laser Learning Environment, a multi-agent path-finding environment where cooperation materializes as one agent blocking a laser to let a teammate pass safely. We...

Yannick Molinghen, Hugo Charels, Tom Lenaerts · 0 citations
Open access Aug 2026

Multi-Agent Reinforcement Learning via Agent-Specific Preference

This paper introduces Multi-AGent Preference-Integrated lEarning (MAGPIE), a framework that leverages agent-specific preference signals in the multi-agent learning process and can derive Nash equilibrium solutions.

Ni Mu, Yao Luan, Yiqin Yang et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.