Skip to content
Conference

Quality Score: A Behavioural Metric for Deliberation Quality in LLM-MAS Systems

Aug 2026 · International Conference on Methods & Models in Automation & Robotics · pp. 221-225 · 0 citations · 17 references

Abstract

Autonomous systems increasingly employ Large Language Models (LLMs) as a deliberative layer in situations where classical machine learning methods encounter out-of-distribution scenarios. The use of multiple models in a multi-agent configuration (LLM-MAS) enables mutual validation of responses and reduces the risk of errors arising from the inherent biases of a single model. A fundamental question arises, however: what is the quality of the deliberation process in such a dialogue? This paper introduces Quality Score (Q)-a behavioural, post-hoc metric grounded in observable dialogue interaction patterns, assessing the quality of the deliberation process in LLM-MAS systems. Q is computed from four components: Position Evolution (PE), Argument Diversity (AD), Consensus Timing (CT), and Mutual Acknowledgment (MA). Validation was conducted on 303 dialogues across three experimental conditions differing in the presence of Dialogue Service Level Objectives (DSLO). Results show that Q captures meaningful differences in deliberation dynamics: unconstrained dialogues achieve an average $\mathbf{Q}=\mathbf{0. 7 9 5}$, while DSLO conditions yield $\mathbf{Q}=\mathbf{0. 7 0 2}$ and $\mathbf{Q}=\mathbf{0. 7 4 7}$, respectively. Given a sufficiently large statistical sample, Q enables identification of which system configurations systematically produce higher-quality deliberation-knowledge potentially useful when designing LLM-MAS systems for applications in automation and robotics.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.