Jul 2026· Annual International Computer Software and Applications Conference· pp. 130-139· 0 citations· 26 references
Computer Science
Abstract
Large Language Models (LLMs) show strong potential for clinical reasoning, yet their deployment in medical decision support is hindered by hallucinations, overconfidence, and limited transparency. We propose Dialectic Diagnosis, an agentic framework in which two heterogeneous LLM agents engage in structured argumentative interaction inspired by clinical second-opinion workflows. A Clinical Reasoner proposes candidate diagnoses, while a Skeptical Critic challenges these hypotheses by identifying omissions, cognitive biases, and unsupported reasoning. Their interaction is governed by a formal finite-state machine (FSM) enforcing a disciplined proposecritique-resolve protocol, with final decisions produced by an Arbiter agent providing calibrated confidence estimates. To ensure transparency, we introduce a Diagnostic Argument Graph that explicitly represents supporting evidence, contradictions, and missing diagnoses. Evaluations on real-world clinical datasets (MIMIC-IV and eICU) show clear gains over single-agent LLM baselines, with improved diagnostic accuracy, lower calibration error, and fewer critical diagnostic omissions. These results indicate that structured argumentative interaction between LLM agents provides a principled path toward safer and more explainable clinical AI systems.
This work identifies a systematic failure mode in RL-trained retrieval agents: outcome-only rewards improve accuracy while degrading faithfulness, a phenomenon the authors term confident hallucination, and addresses this with a faithfulness-gated reward design.
An empirical benchmark for evaluating clinical triage systems that assesses explanation quality alongside decision outcomes, and provides a reproducible, checkpoint-based evaluation pipeline and outline a roadmap for bias stress-testing, hallucination mitigation, and open benchmark release.
S. Marimuthu, Patricia L. Mabry, HealthPartners.Com· 0 citations
This work introduces EVADE (Evidence-Verified Agentic Diagnosis with Escape), an inferential, non-training method that enhances the safety of deploying a single frozen VLM by verifying gate consistency across different image views rather than re-reading the model's own text.
Mohaimenul Azam Khan Raiaan, Nur Mohammad Fahad· 0 citations
Contradictory Statement Multi-Agent Debate (CSMAD), a multi-agent framework that creates structured disagreement by generating a contradictory claim for each input claim, is proposed and consistently outperforms the strongest baseline for both large and medium-sized language models.
Swapnil Gupta, Akshay Verma, Khush Gupta et al.· Annual International ACM SIG...· 0 citations
A self-reflective framework in which an LLM generates an answer, identifies claims that may be uncertain, performs an internal verification stage, and revises the response before delivery is proposed.
Priti Sharma, Sachin Sharma· Iconic research and engineer...· 0 citations
The Evidence Package Benchmark is introduced, integrating 1,870 packages across six heterogeneous sources with explicit modality masks and evidence permissions, and EviBound, a protocol-aware evidence control framework is proposed, a protocol-aware evidence control framework for safer clinical NLP research.
Cheng-Yuan Gao, Jiang Wu, Tao Lu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.