Skip to content
Book Open access

CSMAD: Hallucination Detection via Multi-Agent Debate with NLI-Verified Contradictory Statements

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 4627-4632 · 0 citations · 22 references
Computer Science

TL;DR

Contradictory Statement Multi-Agent Debate (CSMAD), a multi-agent framework that creates structured disagreement by generating a contradictory claim for each input claim, is proposed and consistently outperforms the strongest baseline for both large and medium-sized language models.

Abstract

Large Language Models (LLMs) are prone to hallucinations, producing fluent but factually incorrect statements. Recent multi-agent debate methods improve hallucination detection by jointly improving reasoning and decision-making. However, existing approaches either collaborate which amplifies shared overconfidence, or adopt adversarial preset stances, that can inject incorrect information complicating decision making. To address this, we propose Contradictory Statement Multi-Agent Debate (CSMAD), a multi-agent framework that creates structured disagreement by generating a contradictory claim for each input claim. CSMAD asks independent agents to evaluate the claim and the contradictory claim, which encourages different lines of reasoning without assigning preset stances. When the outcome is non-discriminative; both the contradictory statements are either accepted or rejected; the agents exchange rationales and update their judgments after considering opposing evidence. A final judge then decides the truth of the original claim, using both arguments as context. To make contradictory statement generation reliable, we add a Natural Language Inference (NLI) based verifier that checks whether the generated statement actually contradicts the original claim; if it does not, the system falls back to an explicit negation-based contradiction. Across public benchmarks for question answering and scientific claim verification, as well as a proprietary e-commerce claims dataset, we show that CSMAD consistently outperforms the strongest baseline for both large (Claude-3.5 Sonnet) and medium-sized (Qwen3-8B) language models, improving F1 by +2.3 and +4.1 points, respectively, while reducing LLM token cost by 28%.

Read PDF

Similar papers

Conference Open access Sep 2026

Mitigating Hallucinations in Natural Language Generation through Prompt Engineering: A Mechanism- Oriented Narrative Review

Large language models (LLMs) can generate fluent and confident responses that are factually incorrect, unsupported by evidence, or inconsistent with the source material. These hallucinations reduce the reliability of LLM-based question answering, summariza tion, dialogue, and information retrieval systems, especially w...

Jia-Nian Lin · 0 citations
Book Open access Aug 2026

Who's Adam? Benchmarking Hallucinations in Scientific Dialogue

ADAM-Bench (Auditing Dialogue Assertions with Multimodal Evidence), a benchmark for paper-grounded hallucinations in scientific dialogue, is introduced and two tasks are defined: hallucination detection and minimal evidence set localization.

Ze-Xing Zhang, Tian-Yang Lei, Ke-Wei Yang et al. · 0 citations
Preprint Aug 2026

Decomposed Entailment for Factuality Checking and Hallucination Detection

HallDetect, a lightweight, reference-free, and black-box framework for hallucination detection, is presented, a lightweight, reference-free, and black-box framework for hallucination detection that is evaluated not only on summarization but across a broader range of source-grounded generation settings.

Achir Oukelmoun, N. Semmar, Gaël de Chalendar · 0 citations
Open access Sep 2026

A conservative benchmark of contradiction detection and its limits in LLM hallucination evaluation

Hallucinations—fluent outputs containing incorrect or unsupported factual claims—remain an important obstacle to reliable use of large language models (LLMs). This study evaluates the scope and limits of HalluDetector, a reference-based detector that identifies contradiction-like evidence using lexical, numeric, unavai...

Seung-Ho Lee · 0 citations
Conference Aug 2026

Quorum-Inspired Multi-Agent Consensus for Detecting Reasoning Hallucinations in Retrieval-Augmented Generation: A Cautionary Study

Collective decision-making in nature, such as bacterial quorum sensing and swarm consensus, has inspired the intuition that a panel of large language model (LLM) verifiers can outperform a single model through aggregated voting. In this work, we rigorously evaluate this hypothesis for detecting reasoning hallucinations...

Anish Chaulagain, Dipika Acharya, Lucky Shrestha et al. · 0 citations

Discourse Structure as an Interpretable Signal for Detecting Hallucinated Chain-of-Thought Reasoning in Large Language Models

This paper studies hallucinated CoT as a discourse-structural phenomenon, not only a factual one, and suggests that discourse structure provides an interpretable signal for detecting reasoning hallucinations and complements existing factuality and uncertainty-based hallucination detectors.

Boris Galitsky · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.