This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis and proposes a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis.
Abstract
Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale. Materials and Methods: We applied a two-stage LLM pipeline---open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)---to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts. Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess. Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis. Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis.
LLMs represent a viable and flexible approach to diagnosis code extraction from unstructured clinical notes that can augment structured diagnoses and provide contextualizing metadata.
H. Razzaghi, Nhat Nguyen, M. Pargi et al.· JAMIA Open· 0 citations
The Nimblemind Multi-Agent System is developed, an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, and evaluation was limited to a single-institution cohort and external validation is needed, demonstrating the feasibility of automated, auditable feature engineering for complex...
This is the first study to apply an LLM for automated classification of radiation oncology safety events and to outline a framework for future reporting systems that leverage artificial intelligence (AI)-assisted workflows, demonstrating strong potential for automating retrospective safety event tagging and streamlinin...
Qiong-Ge Li, Jian Liu, Xing Li et al.· Practical Radiation Oncology· 0 citations
The methodology described can be applied to develop automated measures to detect a range of order error types, examine the epidemiology, and investigate the root causes of order errors in near-real-time, as well as rigorously evaluate the impact of preventive interventions.
P. Spector, A. Grauer, Jerard Z. Kneifati-Hayek et al.· JAMIA Open· 0 citations
A prompt-driven pipeline that converts FHIR R4 laboratory panels into structured, paragraph-length clinical narratives paired with a calibrated 0-1 risk score, using GPT-4o-mini as the generation engine is described, demonstrating feasibility for abnormality flagging and narrative generation.
Rahul Reddy Hanumanthgari· EPJ Web of Conferences· 0 citations
CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework.
Veronica Chatrath, Bryan Zhu, George Pu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.