Skip to content

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review

Feb 2026 · arXiv.org · Vol abs/2604.16321 · 2 citations · 164 references
Computer Science

TL;DR

A Multi-Vocal Literature Review is conducted, combining insights from both academia and industry, including peer-reviewed studies and grey literature to systematically synthesize and analyze existing knowledge on LLM-based multi-agent systems for code generation.

Abstract

Large Language Models (LLMs) have enabled multi-agent systems to perform autonomous code generation for complex tasks. Despite the recent growth in research and industrial applications in this area, there is little work on synthesizing evidence from both academic and industrial sources to capture the current state of research on LLM-based multi-agent systems for code generation. To this end, we conducted a Multi-Vocal Literature Review (MLR), combining insights from both academia and industry, including peer-reviewed studies and grey literature. The aim of this study is to systematically synthesize and analyze existing knowledge on LLM-based multi-agent systems for code generation. Specifically, the review examines the motivations for their use, employed benchmarks and models, key challenges, proposed solutions, and potential directions for future research. We selected and reviewed 114 studies, and the key findings are: 1) the identified reasons for adopting multi-agent systems for code generation were classified into nine categories; 2) the models and evaluation benchmarks utilized across the studies were systematically analyzed to provide a structured overview of commonly adopted LLM configurations and assessment practices; 3) the reported challenges and corresponding solutions were synthesized into six main categories and 26 subcategories; and 4) future research directions were identified and organized into six main categories and 18 subcategories. The results of this MLR will assist researchers and practitioners in pursuing further studies and supporting the real-world adoption of multi-agent systems in industrial settings.

View source

Similar papers

Review Aug 2026

Developing LLM-based Multi-Agent Systems in Software Engineering: A Mixed-Method Experience Report

A comprehensive overview of the existing tools and frameworks for implementing MAS in software engineering and a set of lessons learned and challenges that can help researchers and practitioners to select a suitable MAS framework according to their needs are provided.

Maria Sâmyla Serafim de Oliveira, M. Ibiyo, Marco Gianrusso et al. · 1 citation
Review Open access Aug 2026

A Systematic Survey of LLM-Based Agentic AI Frameworks for Multi-Agent Coordination and Interoperability

The analysis indicates that the promise of LLM-based agents for scalable automation, collaborative reasoning, and complex workflow execution comes with significant challenges in long-horizon reliability, evaluation standardization, communication security, cost-efficient orchestration, governance, and the interpretabili...

N. Mohamed, Prasun Chakrabarti, S. Gupta · 0 citations
Jul 2026

A Multi-Agent Benchmarking Framework for Evaluating the Performance of Large Language Models in Logic Programming

A configurable multi-agent framework for benchmarking LLMs in Prolog code generation that combines a Code Generator Agent, a deterministic execution layer using SWI-Prolog, and an evaluator based on the LLM-as-a-Judge paradigm that supports model-agnostic experimentation and evaluates outputs across functional correctn...

Nikolaos Karamousalidis, P. Kefalas · 0 citations
Review Open access Aug 2026

Multi-Agent Collaborative Code Generator

The work presented here provides a five-agent collaborative architecture to provide continuous and verifiable code optimization by controlled specialization and iterative refinement and provides a scalable platform for intelligent, self-adjusting development environments with a 92.

Jesalkumari Varolia · 0 citations
Review Aug 2026

Integration of Multi-agent Reinforcement Learning (MARL) and Large Language Models (LLMs) in Agentic AI Systems: A Narrative Review

The paper delivers a comprehensive integrative review which examines current LLM–MARL research through various aspects of agent architectures and coordination mechanisms and memory augmentation and evaluation methodologies used in different fields of study.

Harsh Shingavi, Yash Ranbhare, Esha Shasri et al. · 0 citations
Review Open access Jul 2026

Multi-Agent LLM Pipeline for Code Writing: An Experimental Study of Writer-Reviewer Architecture

A multi-agent pipeline-based approach for solving competitive programming problems via Large Language Models (LLMs) is developed and the Writer-Reviewer pipeline where the Writer Agent produces Python solutions and the Reviewer Agent gives static natural language feedback is analyzed.

Temel Kaan Ekiz, M. Z. Konyar · 0 citations

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.