Skip to content
Open access

The potential of LLMs in generating questions and answers with EHRs

Jul 2026 · Frontiers in Digital Health · Vol 8 · 1 citation · 32 references
Medicine

TL;DR

Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming, this study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians.

Abstract

Background This study aimed to generate medical qualification exam questions and their corresponding answers from real-world electronic health records (EHRs) with large language models (LLMs), and to compare their output to that of human medical experts. Methods Utilizing a multicenter bidirectional anonymized database China Elderly Comorbidity Medical Database (CECMed), a total of 8 LLMs: ERNIE 4, ChatGLM 4, Doubao, Hunyuan, Spark 4, Qwen, Llama 3, and Mistral were tasked with generating open-ended questions and answers based on a subset of sampled admission reports. LLMs generated the medical question and answer through few-shot prompting. An independent expert panel scored the AI-generated outputs based on multiple criteria, including coherence, sufficiency of key information, information correctness, factual consistency, evidence of statement, and professionalism, using 5-point Likert scales. Results For question generation, ERNIE 4 achieved the highest cumulative score (16.47). Human experts surpassed LLMs in sufficiency of key information (3.67) but lagged in information correctness (3.63 vs. LLMs' 4.03–4.57). The information correctness of ERNIE was significantly higher than the human's [0.93 (0.62, 1.24), p < 0.01]. For answer generation, humans led overall (14.49), while Doubao outperformed the other LLMs in coherence (3.57), factual consistency (3.60), and professionalism (3.53). The coherence of human's was significantly better than that of 8 LLMs, especially outperformed Llama [0.8 (0.37, 1.23), p < 0.01] and Mistral [0.87 (0.45, 1.28), p < 0.01]. Conclusions Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming. This study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians. Although current LLMs performed dissatisfactorily in some aspects, medical students and interns may find LLMs a useful auxiliary tool to support their learning. Clinical Trial Registration https://clinicaltrials.gov/study/NCT06316544, identifier: NCT06316544.

Read PDF

Similar papers

Open access Jul 2026

Large language models for interpretation of health checkup results

Large language models (LLMs) show strong generalization, yet their ability to interpret structured medical data remains insufficiently studied. This work evaluated four LLMs—Claude Sonnet 4, Gemini 2.5 Pro, GPT-4o, and LLaMA 3.1-70B—using comprehensive health checkup data from the Korean National Health Insurance Service. Multiple prompting strategies (few-shot, role-based, constraint-based, and Chain-of-Thought) were tested. Zero-shot accuracy averaged 0.69 (SD 0.06), increasing to 0.92 (0.06) with combined strategies and to 0.95 (0.07) with Chain-of-Thought. Claude Sonnet 4, Gemini 2.5 Pro, and GPT-4o achieved the highest accuracies (≥ 0.98), while LLaMA 3.1-70B showed lower but improvable performance. Item-level analysis of 10,000 cases demonstrated near-perfect accuracy (0.99–1.00) for most biochemical markers, including glucose, cholesterol, triglycerides, and liver enzymes. In contrast, blood pressure showed lower accuracy (0.61–0.91), with age-related decline, likely due to the complexity of multi-categorical thresholds requiring integration of systolic and diastolic values. Subgroup analyses revealed model-specific biases: sex-related biases were observed in body mass index (Claude Sonnet 4) and urine protein, serum creatinine, and gamma-glutamyl transferase (LLaMA 3.1-70B), while age-related biases were identified in blood pressure (Claude Sonnet 4, Gemini 2.5 Pro) and low-density lipoprotein cholesterol and hemoglobin (LLaMA 3.1-70B). Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.

Jiwon You, Hangsik Shin · 0 citations
Preprint Aug 2026

KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

Automatic Medical Coding (AMC), which assigns standardized International Classification of Diseases (ICD) codes to clinical notes, is essential for medical reimbursement, quality reporting, and clinical research. Existing pre-trained language model (PLM)-based methods typically formulate AMC as an extreme multi-label classification problem over a predefined code set, while recent large language model (LLM)-based approaches instead frame it as generation or multi-step reasoning. However, key challenges remain, including the extreme length of clinical notes that hinders effective interpretation, the vast ICD label space, and complex coding rules that are not explicitly captured by LLMs. In this work, we propose Knowledge-Guided Reasoning over Clinical Evidence with LLMs (KREL), a framework that leverages LLMs for clinical text understanding and reasoning while integrating external ICD coding guidelines as structured knowledge. This design enables tight coupling between domain knowledge and LLM reasoning, reducing hallucinations and improving compliance with coding standards. Experiments on benchmark datasets show that KREL consistently outperforms strong PLM-based and state-of-the-art LLM-based baselines.

Xubin Chen, Yipeng Zhou, Wenxin Sun et al. · 0 citations
Aug 2026

Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists.

BACKGROUND Large language models (LLMs) are increasingly explored for drug information support, yet their reliability and clinical applicability remain uncertain. This study evaluated multiple LLMs in responding to real-world drug information questions retrieved from a university hospital in Thailand, focusing on clarity in Thai, concordance with pharmacist responses, relevance, context awareness, and citation credibility. METHODS A total of 102 drug information questions were submitted to seven LLMs, generating 714 responses. In the inter-rater reliability phase, 10 pilot questions were submitted to all seven LLMs, and the resulting 70 responses were evaluated by three assessors, yielding 210 rating instances. Agreement was measured using intra-class correlation coefficient and Fleiss' kappa. Following satisfactory agreement, the 102 questions were divided into three subsets, each assessed by one assessor using a predefined rubric. Binary outcomes were coded to calculate sensitivity and domain fulfillment rates, with pharmacist-provided answers serving as the reference standard for concordance. Model performance was compared using Cochran's Q test, and citation-related issues were identified from qualitative comments. RESULTS Inter-rater reliability demonstrated substantial to almost perfect agreement, with coefficients ranging from 0.79-0.86. Overall, most LLMs performed well in clarity, relevance, and context awareness. Clarity scores ranged from 85.25%-92.75%, while fulfillment rates ranged from 0.97-1.00 for relevance and 0.90-0.99 for context awareness. Concordance was more variable, ranging from 0.70-0.86, whereas citation credibility was consistently weak, ranging from 0.03-0.36. A significant difference between LLMs was observed only in the concordance, with post hoc analyses identifying a significant difference between ChatGPT-5.2 Thinking (OpenAI, San Francisco, CA, USA) and Copilot Think Deeper (Microsoft, Redmond, WA, USA). (McNemar test, adjusted p = 0.0178). CONCLUSIONS LLMs show potential as tools for preliminary drug information retrieval and rapid responses generation in drug information services. However, variable concordance and persistent limitations in citation credibility indicate the need for continued pharmacist oversight.

Nuntapong Boonrit, Najwa Bin-Useng, Aphichaya Sirijariyawat et al. · 0 citations
Open access Jul 2026

Clinician expertise and prompt engineering enhance cancer information extraction in electronic health records by small language models.

BACKGROUND Real-world data (RWD) in unstructured electronic health records (EHRs) is crucial for understanding complex diseases like cancer, but extracting structured information is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four small, locally deployable language models for information extraction from Italian EHRs. METHODS We examine three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess the impact of prompt design on accuracy. We evaluate the performance of four open-source small language models (SLMs) for clinical information extraction from Italian electronic health records (EHRs) in the APOLLO 11 trial on non-small cell lung cancer (NSCLC). The extraction protocol involves four steps: problem definition, data preprocessing, Large Language Model (LLM)-based information extraction, and output evaluation. RESULTS We show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM (Tumor, Node, Metastasis) staging, PD-L1 (Programmed death-ligand 1), and ECOG-PS (Eastern Cooperative Oncology Group-Performance Status) are more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improve performance and reduced hallucinations. Clinical expertise enhances consistency in the extraction, particularly among students using annotated examples. CONCLUSIONS The study confirms that privacy-preserving SLMs can be deployed locally for efficient and secure cancer data extraction. Findings highlight the need for hybrid systems combining SLMs with expert input and underline the importance of aligning clinical documentation practices with SLM capabilities. This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.

Federica Corso, V. Peppoloni, L. Mazzeo et al. · 0 citations
Preprint Aug 2026

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

Praveen Reddy, C. Mandke, Suvrankar Datta et al. · 0 citations