Skip to content
Book Open access

Combining Policy Gradients with Quality-Diversity in Cooperative Multi-Agent Reinforcement Learning

Jul 2026 · GECCO Companion · pp. 137-140 · 0 citations · 17 references
Computer Science

TL;DR

MAPGA-ME is proposed, a multi-agent extension of PGA-MAP-Elites that integrates policy gradient updates into MAP-Elites for cooperative control and identifies key factors affecting the effectiveness of policy gradient-based QD in multi-agent learning.

Abstract

Quality-Diversity (QD) methods combined with policy gradients have shown strong performance in single-agent reinforcement learning, but extending them to multi-agent settings introduces challenges from partial observability and agent interactions. We propose MAPGA-ME, a multi-agent extension of PGA-MAP-Elites that integrates policy gradient updates into MAP-Elites for cooperative control. Our results show that directly transferring policy gradient mechanisms from single-agent QD does not consistently improve performance in multi-agent environments. In particular, a design choice effective in single-agent settings becomes less suitable under decentralized, partially observable conditions. Across multiple configurations, we identify key factors affecting the effectiveness of policy gradient-based QD in multi-agent learning, providing practical guidance for adapting these methods.

Read PDF

Similar papers

Open access Aug 2026

Multi-Agent Reinforcement Learning via Agent-Specific Preference

This paper introduces Multi-AGent Preference-Integrated lEarning (MAGPIE), a framework that leverages agent-specific preference signals in the multi-agent learning process and can derive Nash equilibrium solutions.

Ni Mu, Yao Luan, Yiqin Yang et al. · 0 citations
Jul 2026

Efficient Heterogeneous Exploration with Mutual Policy Divergence Maximization for Multiagent Reinforcement Learning.

This work introduces a novel MARL framework, Multi-Agent Divergence Policy Optimization (MADPO) with Mutual Policy Divergence Maximization (Mutual PDM), and proposes a new extension of CCS divergence for measuring policy divergence of more than two agents, the Generalized Conditional Cauchy-Schwarz (GCCS) divergence.

Haowen Dou, Lujuan Dang, Mingfei Lu et al. · 0 citations
Jul 2026

Learning from success: efficient selective learning methods for multi-agent sparse-reward tasks

This paper proposes effective multi-agent selective learning methods to boost sample-efficient training by learning from successful experiences, and adopts a retrogression-based selection method to identify successful agent trajectories from the team rewards.

Xin-Ning Chen, Xuan Liu, Yanwen Ba et al. · 0 citations
Conference Aug 2026

Adaptive Exploration and Curriculum Learning for Multi-Agent Reinforcement Learning

Multi-agent Reinforcement learning has gained significant attention for solving decision-making problems involving multiple autonomous agents. However, effective learning in MARL is still difficult due to environments, dependencies between agents, and poor exploration strategies. Although adaptive exploration and curri...

B. Adwaith, Kevin Francis, Remya Nair T · 0 citations
Conference Open access Sep 2026

Towards Streamlined Learning and Search for Multi-Agent Optimization

Focusing on multi-agent path finding as an exemplary problem, this paper proposes to simplify two popular approaches to MAPF, namely multi-agent reinforcement learning and adaptive search, to enable seamless combination and transferability of methods without substantial engineering effort.

Thomy Phan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.