Quality Score: A Behavioural Metric for Deliberation Quality in LLM-MAS Systems
Autonomous systems increasingly employ Large Language Models (LLMs) as a deliberative layer in situations where classical machine learning methods encounter out-of-distribution scenarios. The use of multiple models in a multi-agent configuration (LLM-MAS) enables mutual validation of responses and reduces the risk of errors arising from the inherent biases of a single model. A fundamental question arises, however: what is the quality of the deliberation process in such a dialogue? This paper introduces Quality Score (Q)-a behavioural, post-hoc metric grounded in observable dialogue interaction patterns, assessing the quality of the deliberation process in LLM-MAS systems. Q is computed from four components: Position Evolution (PE), Argument Diversity (AD), Consensus Timing (CT), and Mutual Acknowledgment (MA). Validation was conducted on 303 dialogues across three experimental conditions differing in the presence of Dialogue Service Level Objectives (DSLO). Results show that Q captures meaningful differences in deliberation dynamics: unconstrained dialogues achieve an average $\mathbf{Q}=\mathbf{0. 7 9 5}$, while DSLO conditions yield $\mathbf{Q}=\mathbf{0. 7 0 2}$ and $\mathbf{Q}=\mathbf{0. 7 4 7}$, respectively. Given a sufficiently large statistical sample, Q enables identification of which system configurations systematically produce higher-quality deliberation-knowledge potentially useful when designing LLM-MAS systems for applications in automation and robotics.