Skip to content
Open access

Multi-Agent debate system based on large language models: structured deliberation and validation in satellite communications

Susana Gómez Álvarez Alejandro Mozo Quesada Tomás Navarro Sergio Gálvez Rojas Francisco López Valverde
Aug 2026 · Journal of Intelligence and Information Systems · 0 citations · 33 references

TL;DR

This study proposes a moderated, domain-adaptive multi-agent debate framework applied to the high-stakes domain of satellite communications (SatCom), and assesses the efficacy of structured deliberation against single-agent baselines, and the impact of model heterogeneity versus homogeneity.

Abstract

Structured multi-agent debates among Large Language Models (LLMs) have emerged as a powerful paradigm for enhancing reasoning reliability and argumentative coherence. Motivated by the European Space Agency’s (ESA) interest in trustworthy AI for space operations, this study proposes a moderated, domain-adaptive multi-agent debate framework applied to the high-stakes domain of satellite communications (SatCom). Specifically, it assesses (i) the efficacy of structured deliberation against single-agent baselines, and (ii) the impact of model heterogeneity versus homogeneity. A single-agent baseline is compared against a multi-agent framework deploying a moderator and two domain-specialized experts. These systems utilize local 70B-parameter LLMs in homogeneous (Llama-3.3) and heterogeneous (Llama-3.3 + DeepSeek-R1 + Qwen-2.5) configurations, all augmented with a shared, curated Retrieval-Augmented Generation (RAG) corpus combining academic institutional sources and ESA material from the Nebula portal (SatNex V programme). Outputs from 213 technical queries are evaluated via LLM-as-a-judge across three phases: baseline proficiency, strategic reasoning, and executive readiness. Single-agent systems lead in encyclopedic tasks, where retrieval suffices over deliberation. However, both multi-agent configurations outperform in strategic reasoning, with heterogeneous debates achieving superior performance in executive scenarios by victory margins of up to 2.75 points on a 10-point scale. These results validate architectural diversity as a decisive factor in resolving high-complexity technical conflicts. Ultimately, this work delivers a generalizable, fully traceable deliberation framework suitable for real-world, mission-critical environments. Code, prompts, and evaluation data are publicly available at https://github.com/amozo-es/multi-agent-debate/ .

Read PDF

Similar papers

Review Jul 2026

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

This work presents a systematic literature review characterizing 141 primary studies on Multi-Agent Debate, and derives a three-dimensional taxonomy covering debate participants, the interaction mechanisms structuring the exchange, and the agreement protocols governing debate resolution, supported by formal notations to render MAD configurations.

Quim Motger, M. Oriol, J. Marco et al. · 1 citation
Open access Jul 2026

Language Model Council: A Multi-Agent Framework using Explainable AI

This work introduces the Language Model Council (LMC), a collaborative framework that combines the expertise of multiple specialized AI agents to evaluate a user query from different perspectives and outperforms traditional single-model systems by improving response quality, reducing hallucinations, and increasing user trust through enhanced explainability.

D. M, Shwetha Kr, G. Divya et al. · 0 citations
Review

ICMTEST-2026 - an Adaptive Multi-Agent Framework for Argumentation With a Legal Case Study

Performing complex, debate-oriented reasoning tasks does not only involve searching for information, but it also involves organizing the data into clear arguments and counterarguments. Traditional AI systems usually face problems in having a structured flow of arguments, and they fall apart when faced with cold-start conditions in the absence of precedents. To help meet these challenges, we propose an adaptive multi-agent framework that combines the principle of Retrieval-Augmented Generation (RAG) with structured debate and knowledge enrichment over time. This framework enables long-term learning and better performance over time, in contrast to the static systems that do not update their own knowledge base following a case. To illustrate the effectiveness of this framework, we present a case study of the problem involving a part of the Indian Penal Code. The system manages to mimic a trial between opposing parties, assess evidence, and give a thorough final report, including an overview of facts and a probable court ruling.

Iffat Patel, Abhishek Tiwari, Deepak Yadav et al. · 0 citations
Conference Aug 2026

Quality Score: A Behavioural Metric for Deliberation Quality in LLM-MAS Systems

Autonomous systems increasingly employ Large Language Models (LLMs) as a deliberative layer in situations where classical machine learning methods encounter out-of-distribution scenarios. The use of multiple models in a multi-agent configuration (LLM-MAS) enables mutual validation of responses and reduces the risk of errors arising from the inherent biases of a single model. A fundamental question arises, however: what is the quality of the deliberation process in such a dialogue? This paper introduces Quality Score (Q)-a behavioural, post-hoc metric grounded in observable dialogue interaction patterns, assessing the quality of the deliberation process in LLM-MAS systems. Q is computed from four components: Position Evolution (PE), Argument Diversity (AD), Consensus Timing (CT), and Mutual Acknowledgment (MA). Validation was conducted on 303 dialogues across three experimental conditions differing in the presence of Dialogue Service Level Objectives (DSLO). Results show that Q captures meaningful differences in deliberation dynamics: unconstrained dialogues achieve an average $\mathbf{Q}=\mathbf{0. 7 9 5}$, while DSLO conditions yield $\mathbf{Q}=\mathbf{0. 7 0 2}$ and $\mathbf{Q}=\mathbf{0. 7 4 7}$, respectively. Given a sufficiently large statistical sample, Q enables identification of which system configurations systematically produce higher-quality deliberation-knowledge potentially useful when designing LLM-MAS systems for applications in automation and robotics.

Ján Kapusta, Waldemar Bauer, Jerzy Baranowski · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.