It is suggested that automated abstraction could enable census-level quality measurement, reduce reporting burden, and expand the range of clinically meaningful measures that can be operationalized from narrative clinical documentation.
Abstract
Background. Clinical quality measurement often relies on manual abstraction of medical records, an approach that is costly, burdensome, and often infeasible for measures requiring interpretation of narrative text; these constraints have shaped measure development itself, filtering out clinically important measures that are too difficult to operationalize. We evaluated whether neuro-symbolic artificial intelligence (NSAI), which combines large language model extraction with symbolic reasoning, could reliably abstract complex quality measures from narrative pathology reports. Methods. The NSAI system decomposes each measure into atomic questions and is aligned to real-world reports through case-based refinement, an iterative human-in-the-loop process. Using 2,000 independently double-abstracted reports, we compared NSAI-based abstraction against trained human abstractors across four pathology quality measures established by the College of American Pathologists. Results. The NSAI system's agreement with the adjudicated gold standard (Cohen's kappa = 0.95) matched or modestly exceeded that of the trained human abstractors measured against the same standard (kappa = 0.92), with particularly strong performance on Gastrointestinal Metaplasia (CAP 43). In component analyses, case-based refinement drove the largest accuracy gains (up to 0.25 in kappa), whereas architectural decomposition primarily reduced performance variance across language-model backends more than tenfold, a property essential for clinical deployment. Conclusions. These findings suggest that automated abstraction could enable census-level quality measurement, reduce reporting burden, and expand the range of clinically meaningful measures that can be operationalized from narrative clinical documentation.
LLMs represent a viable and flexible approach to diagnosis code extraction from unstructured clinical notes that can augment structured diagnoses and provide contextualizing metadata.
H. Razzaghi, Nhat Nguyen, M. Pargi et al.· JAMIA Open· 0 citations
Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific ques...
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur et al.· 0 citations
With the rise of artificial intelligence (AI), an increasing number of AI-based diagnostic tools are being developed. Before clinical implementation, these tools must be validated against existing gold standards. This requires trials that quantify the agreement between AI predictions and reference measurements. However...
Abstract Medical artificial intelligence (AI) benchmarks are increasingly used to assess the readiness of large language models for health-related tasks, but aggregate performance scores can obscure clinically meaningful variation across domains. HealthBench, an open benchmark of 5000 multi-turn health conversations, r...
R. McBain, Ellice Huang, Caroline A. Figueroa et al.· BMJ mental health· 0 citations
Current evidence supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
Yun-Jia Wu, Qi Yan, Dingcheng Tian· AI Medicine· 0 citations
This work proposes Knowledge-Guided Reasoning over Clinical Evidence with LLMs (KREL), a framework that leverages LLMs for clinical text understanding and reasoning while integrating external ICD coding guidelines as structured knowledge, and enables tight coupling between domain knowledge and LLM reasoning.
Xubin Chen, Yi-Peng Zhou, Wenxin Sun et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.