Skip to content
Conference

Quorum-Inspired Multi-Agent Consensus for Detecting Reasoning Hallucinations in Retrieval-Augmented Generation: A Cautionary Study

Aug 2026 · International Conference Computational Vision and Bio Inspired Computing · pp. 21-28 · 0 citations · 23 references

Abstract

Collective decision-making in nature, such as bacterial quorum sensing and swarm consensus, has inspired the intuition that a panel of large language model (LLM) verifiers can outperform a single model through aggregated voting. In this work, we rigorously evaluate this hypothesis for detecting reasoning hallucinations defined as fabricated, ungrounded, or self-contradictory steps in multi-hop reasoning chains within retrieval-augmented generation (RAG) systems. We conduct experiments on HotpotQA using a fixed openweight generator and evaluate a heterogeneous three-agent quorum (Logic, Evidence, Critic) against single-judge and homogeneous ensemble baselines across two verifier-capability regimes: strong (24-32B parameters) and weak (7-8B parameters). Evaluation is performed on both a controlled injection benchmark with precise labels and a human-validated set of naturally generated responses. Our results show a consistent negative finding: the heterogeneous quorum does not outperform a single LLM judge in either regime (McNemar p $=\mathbf{0. 4 5}$ in the strong setting and $\mathbf{p}=\mathbf{1. 0 0}$ in the weak setting). Moreover, increasing model diversity does not improve performance over homogeneous ensembles, and scaling the number of agents does not materially change outcomes. These findings also replicate on two additional multi-hop datasets, 2WikiMultiHopQA and MuSiQue. Even on the more challenging MuSiQue benchmark $(\mathbf{F} \mathbf{1}=\mathbf{0. 8 8})$, neither larger quorums nor a learned aggregator outperform the single judge. We observe that the quorum provides only limited benefits, including improved ranking under soft agreement scoring in the weak regime (AUROC 0.92 vs. 0.86) and a tunable precision-abstention trade-off. We further identify a capability threshold at which sub-4B model fails to produce reliable verdicts, indicating a hard lower bound for effective verification. Additionally, we demonstrate that answer correctness is an imperfect proxy for reasoning faithfulness, showing only moderate agreement with human judgment $(\kappa$ = 0.54, 80% raw agreement on HotpotQA; $\kappa=0.59$ on 2WikiMultiHopQA and $\kappa=0.61$ on MuSiQue), with approximately 19% of sampled chains judged genuinely unfaithful. This highlights important limitations in current evaluation practices for hallucination detection. We release all code, prompts, and datasets to support reproducibility. Overall, our findings caution that increasing the number of agents does not inherently improve LLM-based verification and clarify the conditions under which multi-agent quorum systems provide benefits.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.