EviDx is introduced, an evidence-aware active diagnosis framework that pairs patient-specific diagnostic environments with a clinical diagnostic scaffold and an observer-guided runtime harness that improves diagnostic performance and process stability while revealing model-dependent capability boundaries.
Abstract
Clinical diagnosis is an active evidence-seeking process in which clinicians acquire evidence, update competing hypotheses, and decide when the available evidence is sufficient for diagnosis. Yet many medical diagnosis systems built around large language models (LLMs) still formulate diagnosis as static case-to-answer prediction, with limited support for evidence acquisition. Agentic LLMs offer a dynamic alternative through tool use and intermediate diagnostic trajectories, but existing systems often under-specify how patient evidence should be exposed, scaffolded, and controlled at runtime. We introduce EviDx, an evidence-aware active diagnosis framework that pairs patient-specific diagnostic environments with a clinical diagnostic scaffold and an observer-guided runtime harness. In EviDx, $\mathcal{E}$-Synthesis constructs interactive environments from raw clinical cases; the scaffold organizes role-specialized agents, evidence tools, and evolving evidence states; and the harness regulates diagnostic termination by tracking uncertainty and evidence coverage. A 3-level evaluation pyramid assesses execution robustness, reasoning dynamics, and diagnostic outcomes. Experiments show that EviDx improves diagnostic performance and process stability while revealing model-dependent capability boundaries.
Multimodal clinical decision-making requires reliable reasoning over heterogeneous evidence from electronic health records, medical images, and physiological signals. Existing models typically map these inputs directly to diagnoses without explicitly assessing evidence sufficiency, tool-use requirements, or diagnostic...
CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories, is introduced, demonstrating that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisi...
Xi-Wei Dai, Zi-Jie Meng, Zhiting Fan et al.· 0 citations
Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected evidence rather than the ability to seek it across clinical records and longitudinal imaging. We propose CASE: a series of role-specific Clinical Agents for Seeking Evide...
Min-Ye Shao, Chao-Hui Yu, Yi-Xuan Wu et al.· 0 citations
MediSkill-Evo is introduced, which self-evolves governed process knowledge without fine-tuning the backbone, and updates clinical, process, symbolic, and visual knowledge in four typed banks under type-specific validation and scope rules.
Ruoyu Wu, Shen-Fu Xie, Yin-Qian Sun et al.· 0 citations
It is argued that each of the four challenges facing responsible development is addressed by a clinical testing harness: a structured evaluation environment comprising scenario libraries built from clinical edge cases, full-trajectory observability, explicit escalation testing, and staged evidence thresholds tied to sc...
E. Waisberg, Joseph W. Guarnieri· Annals of Biomedical Enginee...· 0 citations
Medical language models can produce correct diagnoses despite incomplete investigations and unsupported reasoning. To support long-horizon, evidence-grounded diagnosis, we introduce \textbf{TrustMed-RL}. Built from PubMed rare-disease cases and over 24,000 manually annotated image panels, it integrates interviews, exam...
Wen-Xin Zhan, Yi-Zheng Jiao, Hai-Feng Song et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.