Skip to content
Preprint

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

MediSkill-Evo is introduced, which self-evolves governed process knowledge without fine-tuning the backbone, and updates clinical, process, symbolic, and visual knowledge in four typed banks under type-specific validation and scope rules.

Abstract

Interactive clinical agents operate under partial observability, so reliable care depends on reaching the correct diagnosis through evidence-grounded, safe interactions. Yet existing agents struggle to convert experience into reusable process knowledge with explicit provenance and authority. To address this gap, we introduce MediSkill-Evo, which self-evolves governed process knowledge without fine-tuning the backbone. It realizes this self-evolution by updating clinical, process, symbolic, and visual knowledge in four typed banks under type-specific validation and scope rules. The Process-Constrained Preference Harness then turns validated knowledge into action by grounding candidates in evidence and prioritizing safer decisions. We evaluate on 300 MIMIC-IV-derived FullChain encounters, 180 hard-isolation conditions covering six process obligations, and 100 multimodal NEJM image-diagnosis cases. On Qwen FullChain, MediSkill-Evo improves diagnosis accuracy by 7.81% and treatment-intent coverage by 70.67% over the best-performing prior agent, while reducing critical failures by 43.04%. Under stress, it improves the stress-process composite by 7.77% and required-action completion by 12.41% over the best-performing agent for each metric, with stronger patient-fact, temporal-evidence, and triage-red-flag recovery and no controller-scored errors in unavailable-evidence, treatment, and triage safety checks. On multimodal NEJM diagnosis, MediSkill-Evo with optional MedSAM localization improves diagnosis accuracy by 2.56% and core score by 18.96% over the best-performing memory agent. Code is available at https://anonymous.4open.science/r/mediskill-evo_anonymous-68E7.

View source

Similar papers

Preprint Aug 2026

EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents

EviDx is introduced, an evidence-aware active diagnosis framework that pairs patient-specific diagnostic environments with a clinical diagnostic scaffold and an observer-guided runtime harness that improves diagnostic performance and process stability while revealing model-dependent capability boundaries.

Lihang Zeng, Shao-Ting Zhang, Xiaofan Zhang · 0 citations

EviGen: Predictive Evidence Scaffolding for Verifiable Clinical Rationale Generation

Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review of these records is impractical, and LLM-based processing is costly and often unreliable...

Feng-Nan Li, Heman Burre, Li-Wen Sun et al. · 0 citations
Review Open access Sep 2026

Transforming large language models into medical specialists via knowledge injection.

While general-purpose large language models (LLMs) demonstrate remarkable capabilities, their clinical application demands rigorous adaptation to ensure safety and accuracy. This review presents a comprehensive framework for transforming LLMs into trustworthy medical specialists. We detail three core knowledge-injectio...

Kiduk Kim, Jeong Min Song, Dong Yeong Kim et al. · 0 citations
#natural language process... Preprint Aug 2026

GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space

GPAgentBench-2K is introduced, the first Constrained MDP (CMDP) LLM-agent benchmark for primary-care clinical decision-making, constructed from expert-validated records of real-world GP encounters, and uncovers a clinical quality-safety gap.

Bo-Qi Chen, Xudong Liu, Y. Ao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making

The application of large language models (LLMs) to personalized medical assistants has garnered growing interest. However, existing medical benchmarks largely rely on static question answering with pre-selected evidence, leaving unclear whether LLMs can make reliable clinical decisions from real longitudinal electronic...

Jun Xiang, Zhi-Jie Bao, Rong Hu et al. · 1 citation
Review Open access 2026

Beyond Static Agents: A Six-Dimensional Taxonomy and Survey of Self-Evolving LLM Agents for Healthcare

Large language model (LLM)-based agents are increasingly being explored for healthcare tasks such as clinical decision support, care coordination and autonomous workflow execution. Beyond static pipelines, recent systems claim to self-evolve by adapting their tools, memory, reasoning, policy, context and coordination s...

Shubham Vatsal, Harsh Dubey, Ahsaas Bajaj · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.