Skip to content
Preprint

EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

EviDx is introduced, an evidence-aware active diagnosis framework that pairs patient-specific diagnostic environments with a clinical diagnostic scaffold and an observer-guided runtime harness that improves diagnostic performance and process stability while revealing model-dependent capability boundaries.

Abstract

Clinical diagnosis is an active evidence-seeking process in which clinicians acquire evidence, update competing hypotheses, and decide when the available evidence is sufficient for diagnosis. Yet many medical diagnosis systems built around large language models (LLMs) still formulate diagnosis as static case-to-answer prediction, with limited support for evidence acquisition. Agentic LLMs offer a dynamic alternative through tool use and intermediate diagnostic trajectories, but existing systems often under-specify how patient evidence should be exposed, scaffolded, and controlled at runtime. We introduce EviDx, an evidence-aware active diagnosis framework that pairs patient-specific diagnostic environments with a clinical diagnostic scaffold and an observer-guided runtime harness. In EviDx, $\mathcal{E}$-Synthesis constructs interactive environments from raw clinical cases; the scaffold organizes role-specialized agents, evidence tools, and evolving evidence states; and the harness regulates diagnostic termination by tracking uncertainty and evidence coverage. A 3-level evaluation pyramid assesses execution robustness, reasoning dynamics, and diagnostic outcomes. Experiments show that EviDx improves diagnostic performance and process stability while revealing model-dependent capability boundaries.

View source

Similar papers

#natural language process... Preprint Sep 2026

MedTRACE: Tool-Augmented Multimodal Clinical Reasoning Agents for Evidence-Grounded Decision-Making

Multimodal clinical decision-making requires reliable reasoning over heterogeneous evidence from electronic health records, medical images, and physiological signals. Existing models typically map these inputs directly to diagnoses without explicitly assessing evidence sufficiency, tool-use requirements, or diagnostic...

Ji Lu, Li-Fei Liu, Hao-Ran Yu et al. · 0 citations
Preprint Aug 2026

CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents

CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories, is introduced, demonstrating that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisi...

Xi-Wei Dai, Zi-Jie Meng, Zhiting Fan et al. · 0 citations
Preprint Sep 2026

From Given to Gathered Evidence: Agentic Learning for Longitudinal Medical Reasoning

Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected evidence rather than the ability to seek it across clinical records and longitudinal imaging. We propose CASE: a series of role-specific Clinical Agents for Seeking Evide...

Min-Ye Shao, Chao-Hui Yu, Yi-Xuan Wu et al. · 0 citations
Preprint Aug 2026

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

MediSkill-Evo is introduced, which self-evolves governed process knowledge without fine-tuning the backbone, and updates clinical, process, symbolic, and visual knowledge in four typed banks under type-specific validation and scope rules.

Ruoyu Wu, Shen-Fu Xie, Yin-Qian Sun et al. · 0 citations
#software testing Review Sep 2026

Agentic AI in Medicine: Challenges for Responsible Development and the Case for Clinical Testing Harnesses.

It is argued that each of the four challenges facing responsible development is addressed by a clinical testing harness: a structured evaluation environment comprising scenario libraries built from clinical edge cases, full-trajectory observability, explicit escalation testing, and staged evidence thresholds tied to sc...

E. Waisberg, Joseph W. Guarnieri · 0 citations
#artificial intelligence Review Oct 2026

TrustMed-RL: Long-Horizon Reinforcement Learning for Evidence-Grounded Clinical Diagnosis

Medical language models can produce correct diagnoses despite incomplete investigations and unsupported reasoning. To support long-horizon, evidence-grounded diagnosis, we introduce \textbf{TrustMed-RL}. Built from PubMed rare-disease cases and over 24,000 manually annotated image panels, it integrates interviews, exam...

Wen-Xin Zhan, Yi-Zheng Jiao, Hai-Feng Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.