Skip to content
Conference

Two Minds are Safer Than One: Argumentative Llm Agents for Clinical Diagnosis

Jul 2026 · Annual International Computer Software and Applications Conference · pp. 130-139 · 0 citations · 26 references
Computer Science

Abstract

Large Language Models (LLMs) show strong potential for clinical reasoning, yet their deployment in medical decision support is hindered by hallucinations, overconfidence, and limited transparency. We propose Dialectic Diagnosis, an agentic framework in which two heterogeneous LLM agents engage in structured argumentative interaction inspired by clinical second-opinion workflows. A Clinical Reasoner proposes candidate diagnoses, while a Skeptical Critic challenges these hypotheses by identifying omissions, cognitive biases, and unsupported reasoning. Their interaction is governed by a formal finite-state machine (FSM) enforcing a disciplined proposecritique-resolve protocol, with final decisions produced by an Arbiter agent providing calibrated confidence estimates. To ensure transparency, we introduce a Diagnostic Argument Graph that explicitly represents supporting evidence, contradictions, and missing diagnoses. Evaluations on real-world clinical datasets (MIMIC-IV and eICU) show clear gains over single-agent LLM baselines, with improved diagnostic accuracy, lower calibration error, and fewer critical diagnostic omissions. These results indicate that structured argumentative interaction between LLM agents provides a principled path toward safer and more explainable clinical AI systems.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning

This work identifies a systematic failure mode in RL-trained retrieval agents: outcome-only rewards improve accuracy while degrading faithfulness, a phenomenon the authors term confident hallucination, and addresses this with a faithfulness-gated reward design.

Jiangyuan Chen, Cheng-Hao Zhang, Hengxing Cai · 0 citations

Right for the Wrong Reasons: A Benchmark for Hallucination and Clinical Safety in AI Health Triage

An empirical benchmark for evaluating clinical triage systems that assesses explanation quality alongside decision outcomes, and provides a reproducible, checkpoint-based evaluation pipeline and outline a roadmap for bias stress-testing, hallucination mitigation, and open benchmark release.

S. Marimuthu, Patricia L. Mabry, HealthPartners.Com · 0 citations
Preprint Aug 2026

EVADE: Evidence-Verified Agentic Diagnosis with Escape

This work introduces EVADE (Evidence-Verified Agentic Diagnosis with Escape), an inferential, non-training method that enhances the safety of deploying a single frozen VLM by verifying gate consistency across different image views rather than re-reading the model's own text.

Mohaimenul Azam Khan Raiaan, Nur Mohammad Fahad · 0 citations
Book Open access Jul 2026

CSMAD: Hallucination Detection via Multi-Agent Debate with NLI-Verified Contradictory Statements

Contradictory Statement Multi-Agent Debate (CSMAD), a multi-agent framework that creates structured disagreement by generating a contradictory claim for each input claim, is proposed and consistently outperforms the strongest baseline for both large and medium-sized language models.

Swapnil Gupta, Akshay Verma, Khush Gupta et al. · 0 citations
#natural language process... Preprint Aug 2026

Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols

The Evidence Package Benchmark is introduced, integrating 1,870 packages across six heterogeneous sources with explicit modality masks and evidence permissions, and EviBound, a protocol-aware evidence control framework is proposed, a protocol-aware evidence control framework for safer clinical NLP research.

Cheng-Yuan Gao, Jiang Wu, Tao Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.