Skip to content
Review Open access

From Multi-Agent Reinforcement Learning to Agentic AI: A Comprehensive Literature Review of Algorithmic Advances and Decision-Analytic Implications (2020-2025)

Aug 2026 · Applied Decision Analytics · 0 citations · 16 references

Abstract

Agentic artificial intelligence has evolved from a research aspiration to a deployable technology between 2020 and 2025. This evolution rests on two intertwined research trajectories: the maturation of multi-agent reinforcement learning (MARL) for coordinated sequential decision-making, and the emergence of large language model (LLM)-based agent architectures integrating symbolic reasoning, tool use, and natural-language communication into cooperative multi-agent workflows. This literature review synthesizes 57 peer-reviewed and openly archived contributions published since 2019 across journals and reputable venues, organized into a thematic taxonomy spanning value-decomposition algorithms (QMIX, QPLEX, Weighted QMIX, FACMAC), trust-region and sequence-model policy methods (MAPPO, HAPPO, MAT, HARL, UPDeT), communication and role learning (NDQ, I2C, ROMA, RODE), credit assignment (LICA, Difference Rewards Policy Gradients, DOP), game-theoretic equilibrium solvers (Pipeline PSRO, JPSRO, Online Double Oracle), open-ended and mixed-motive learning (Open-Ended Learning Team, CICERO, alliance dilemmas), and LLM-based agentic frameworks (AutoGen, MetaGPT, CAMEL, AgentVerse, ChatDev, Generative Agents, Voyager, ReAct, Reflexion, Tree of Thoughts). We compare benchmark and reproducibility infrastructure (PettingZoo, EPyMARL benchmarking, SMAC variants), examine application domains (autonomous driving, multi-agent pathfinding, software engineering, scientific discovery), and discuss implications for applied decision analytics, including human-in-the-loop arbitration, risk-bounded coordination, and verifiable autonomy. We close with an agenda of open problems including non-stationarity, credit assignment under partial observability, alignment and safety in deceptive agents, evaluation under distribution shift, and integrating symbolic reasoning with reinforcement-learned policies to guide the next phase of agentic AI research.

Read PDF