2026· Journal of Intelligent Medicine and Healthcare· Vol 4, pp. 109-124· 0 citations· 31 references
TL;DR
The contribution is not that KRD is universally optimal but that layer-wise auditability is a design discipline whose cost in this setting was lower than its critics would have predicted.
Abstract
: High-stakes clinical decision support (CDS) demands a property that aggregate accuracy cannot capture: a trace that a clinician who was not in the room can inspect layer by layer when the system is wrong. We argue that the way to obtain this property is to refuse to entangle the large language model (LLM) with the rest of the pipeline. We propose KRD (Knowledge–Rule–Decision) , a four-component architecture that separates fact extraction, a compile-time clinical knowledge layer in the spirit of the LLM Wiki pattern of Karpathy, a rule layer of hand-written contraindications and heuristics, and a decision interface whose compose method short-circuits to a rule-cited blocking response whenever any hard violation fires. We evaluate KRD against a pure language model, a retrieval-augmented language model, a rule-only system, and a light hybrid on a benchmark of 32 type-1 diabetes scenarios. A strict version of the unsafe-suggestion rate stratifies the five systems monotonically into four distinct tiers from 0.867 down to zero, with S4 and S5 tied at the floor; the full KRD stack and the light hybrid reach the hard-safety ceiling together; KRD leads the light hybrid on evidence trace completeness by 25% relative and on reviewer correction burden by 12% relative, both directionally clear and borderline significant under bootstrap intervals; and KRD issues 17 language model calls per benchmark pass against the light hybrid’s 32, a 47% reduction that is a direct consequence of the architectural choice to evaluate the rule layer before invoking the model. We also report honestly that the evidence gate is inert on this benchmark because every compiled concept is graded A or B, and we trace five fact-extraction failures to a single field and a single linguistic pattern. The contribution is not that KRD is universally optimal but that layer-wise auditability is a design discipline whose cost in this setting was lower than its critics would have predicted.
Embedding domain-specific knowledge into LLMs may improve performance on specialized exam-style thoracic-surgery questions on this text-only benchmark, however, the present 56-item evaluation does not establish clinical equivalence, diagnostic accuracy in practice, multimodal competence, or readiness for real-world cli...
Qian Li, Yong-Xin Li, Chao Ye et al.· Frontiers in Artificial Inte...· 0 citations
A properly configured LLM can function as an effective decision-support tool in hospital compounding pharmacy, improving efficiency while maintaining high standards of completeness, accuracy, and regulatory compliance.
E. Castellana, MR Chiappetta· Hospital Pharmacy· 0 citations
DDx-Finder is presented, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demand...
H. Lim, H. Yi, J. Yoon et al.· medRxiv· 0 citations
Preparing training data for domain-specific medical Named Entity Recognition (NER) involves a trade-off between annotation quality, expert effort, and data privacy: manual annotation is costly, whereas cloud-based Large Language Models (LLMs) raise concerns about the control of sensitive clinical text. This article int...
Florian Freund, Philippe Tamla, Bao Tran et al.· Electronics· 0 citations
Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (R...
Jia-Yu Huang, Zi-Chen Tang, Qian-Hui Ling et al.· 0 citations
Lung cancer is among the deadliest cancers worldwide, largely because it is usually caught too late. Building AI tools for earlier, stage-aware diagnosis is hard in practice: patient scans are scattered across hospitals that cannot share them freely, and clinicians are reluctant to trust models that cannot explain thei...
Devyani Rawat, Shuchi Bhadula, Sachin Sharma· International Journal of Adv...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.