Skip to content
Review

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

Jul 2026 · arXiv.org · Vol abs/2607.22954 · 0 citations · 34 references
Computer Science Mathematics

TL;DR

This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis and proposes a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis.

Abstract

Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale. Materials and Methods: We applied a two-stage LLM pipeline---open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)---to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts. Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess. Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis. Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis.

View source

Similar papers

Preprint Aug 2026

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

The Nimblemind Multi-Agent System is developed, an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, and evaluation was limited to a single-institution cohort and external validation is needed, demonstrating the feasibility of automated, auditable feature engineering for complex...

Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi et al. · 0 citations
Review Sep 2026

Automated Classification of Radiation Oncology Safety Events Using Large Language Models: A Novel Approach to Streamline Reporting and Enable Retrospective Analysis.

This is the first study to apply an LLM for automated classification of radiation oncology safety events and to outline a framework for future reporting systems that leverage artificial intelligence (AI)-assisted workflows, demonstrating strong potential for automating retrospective safety event tagging and streamlinin...

Qiong-Ge Li, Jian Liu, Xing Li et al. · 0 citations
Open access Sep 2026

A framework for developing, validating, and utilizing automated measures of order errors using the retract-and-reorder methodology

The methodology described can be applied to develop automated measures to detect a range of order error types, examine the epidemiology, and investigate the root causes of order errors in near-real-time, as well as rigorously evaluate the impact of preventive interventions.

P. Spector, A. Grauer, Jerard Z. Kneifati-Hayek et al. · 0 citations
Conference Open access 2026

Intelligent Narrative Summaries and Risk Scoring of Laboratory Panels with Large Language Models

A prompt-driven pipeline that converts FHIR R4 laboratory panels into structured, paragraph-length clinical narratives paired with a calibrated 0-1 risk score, using GPT-4o-mini as the generation engine is described, demonstrating feasibility for abnormality flagging and narrative generation.

Rahul Reddy Hanumanthgari · 0 citations
Review Aug 2026

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework.

Veronica Chatrath, Bryan Zhu, George Pu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.