Skip to content
Open access

Neuro-Symbolic AI for Automated Pathology Quality Measurement

Jul 2026 · medRxiv · 0 citations
Medicine

TL;DR

It is suggested that automated abstraction could enable census-level quality measurement, reduce reporting burden, and expand the range of clinically meaningful measures that can be operationalized from narrative clinical documentation.

Abstract

Background. Clinical quality measurement often relies on manual abstraction of medical records, an approach that is costly, burdensome, and often infeasible for measures requiring interpretation of narrative text; these constraints have shaped measure development itself, filtering out clinically important measures that are too difficult to operationalize. We evaluated whether neuro-symbolic artificial intelligence (NSAI), which combines large language model extraction with symbolic reasoning, could reliably abstract complex quality measures from narrative pathology reports. Methods. The NSAI system decomposes each measure into atomic questions and is aligned to real-world reports through case-based refinement, an iterative human-in-the-loop process. Using 2,000 independently double-abstracted reports, we compared NSAI-based abstraction against trained human abstractors across four pathology quality measures established by the College of American Pathologists. Results. The NSAI system's agreement with the adjudicated gold standard (Cohen's kappa = 0.95) matched or modestly exceeded that of the trained human abstractors measured against the same standard (kappa = 0.92), with particularly strong performance on Gastrointestinal Metaplasia (CAP 43). In component analyses, case-based refinement drove the largest accuracy gains (up to 0.25 in kappa), whereas architectural decomposition primarily reduced performance variance across language-model backends more than tenfold, a property essential for clinical deployment. Conclusions. These findings suggest that automated abstraction could enable census-level quality measurement, reduce reporting burden, and expand the range of clinically meaningful measures that can be operationalized from narrative clinical documentation.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Scaling Clinical Judgment to Evaluate Medical AI

Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific ques...

Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur et al. · 0 citations
Open access Sep 2026

Statistical Methods for Assessing Diagnostic Agreement

With the rise of artificial intelligence (AI), an increasing number of AI-based diagnostic tools are being developed. Before clinical implementation, these tools must be validated against existing gold standards. This requires trials that quantify the agreement between AI predictions and reference measurements. However...

Maximilian Pilz · 0 citations
Review Open access Sep 2026

Making medical AI benchmarks clinically interpretable: the case of mental health

Abstract Medical artificial intelligence (AI) benchmarks are increasingly used to assess the readiness of large language models for health-related tasks, but aggregate performance scores can obscure clinically meaningful variation across domains. HealthBench, an open benchmark of 5000 multi-turn health conversations, r...

R. McBain, Ellice Huang, Caroline A. Figueroa et al. · 0 citations
#natural language process... Preprint Aug 2026

KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

This work proposes Knowledge-Guided Reasoning over Clinical Evidence with LLMs (KREL), a framework that leverages LLMs for clinical text understanding and reasoning while integrating external ICD coding guidelines as structured knowledge, and enables tight coupling between domain knowledge and LLM reasoning.

Xubin Chen, Yi-Peng Zhou, Wenxin Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.