Jul 2026· Annual International ACM SIGIR Conference on Research and Development in Information Retrieval· pp. 4627-4632· 0 citations· 22 references
Computer Science
TL;DR
Contradictory Statement Multi-Agent Debate (CSMAD), a multi-agent framework that creates structured disagreement by generating a contradictory claim for each input claim, is proposed and consistently outperforms the strongest baseline for both large and medium-sized language models.
Abstract
Large Language Models (LLMs) are prone to hallucinations, producing fluent but factually incorrect statements. Recent multi-agent debate methods improve hallucination detection by jointly improving reasoning and decision-making. However, existing approaches either collaborate which amplifies shared overconfidence, or adopt adversarial preset stances, that can inject incorrect information complicating decision making. To address this, we propose Contradictory Statement Multi-Agent Debate (CSMAD), a multi-agent framework that creates structured disagreement by generating a contradictory claim for each input claim. CSMAD asks independent agents to evaluate the claim and the contradictory claim, which encourages different lines of reasoning without assigning preset stances. When the outcome is non-discriminative; both the contradictory statements are either accepted or rejected; the agents exchange rationales and update their judgments after considering opposing evidence. A final judge then decides the truth of the original claim, using both arguments as context. To make contradictory statement generation reliable, we add a Natural Language Inference (NLI) based verifier that checks whether the generated statement actually contradicts the original claim; if it does not, the system falls back to an explicit negation-based contradiction. Across public benchmarks for question answering and scientific claim verification, as well as a proprietary e-commerce claims dataset, we show that CSMAD consistently outperforms the strongest baseline for both large (Claude-3.5 Sonnet) and medium-sized (Qwen3-8B) language models, improving F1 by +2.3 and +4.1 points, respectively, while reducing LLM token cost by 28%.
Large language models (LLMs) can generate fluent and confident responses that are factually incorrect, unsupported by evidence, or inconsistent with the source material. These hallucinations reduce the reliability of LLM-based question answering, summariza tion, dialogue, and information retrieval systems, especially w...
ADAM-Bench (Auditing Dialogue Assertions with Multimodal Evidence), a benchmark for paper-grounded hallucinations in scientific dialogue, is introduced and two tasks are defined: hallucination detection and minimal evidence set localization.
Ze-Xing Zhang, Tian-Yang Lei, Ke-Wei Yang et al.· Proceedings of the 32nd ACM...· 0 citations
HallDetect, a lightweight, reference-free, and black-box framework for hallucination detection, is presented, a lightweight, reference-free, and black-box framework for hallucination detection that is evaluated not only on summarization but across a broader range of source-grounded generation settings.
Achir Oukelmoun, N. Semmar, Gaël de Chalendar· 0 citations
Hallucinations—fluent outputs containing incorrect or unsupported factual claims—remain an important obstacle to reliable use of large language models (LLMs). This study evaluates the scope and limits of HalluDetector, a reference-based detector that identifies contradiction-like evidence using lexical, numeric, unavai...
Seung-Ho Lee· Journal of high school scien...· 0 citations
Collective decision-making in nature, such as bacterial quorum sensing and swarm consensus, has inspired the intuition that a panel of large language model (LLM) verifiers can outperform a single model through aggregated voting. In this work, we rigorously evaluate this hypothesis for detecting reasoning hallucinations...
Anish Chaulagain, Dipika Acharya, Lucky Shrestha et al.· International Conference Com...· 0 citations
This paper studies hallucinated CoT as a discourse-structural phenomenon, not only a factual one, and suggests that discourse structure provides an interpretable signal for detecting reasoning hallucinations and complements existing factuality and uncertainty-based hallucination detectors.
Boris Galitsky· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.