Skip to content

A multi-site benchmarking framework for scalable extraction of geriatric care constructs from electronic health records

Sep 2026 · npj Health Systems · Vol 3 · 0 citations · 49 references
Medicine

TL;DR

Findings highlight key trade-offs among symbolic, generative, and instruction-tuned approaches for CGA and 4Ms phenotyping, suggesting that hybrid heuristic-LLM methods offer interpretability and stability, whereas large language models provide greater adaptability for complex clinical narratives.

Abstract

Current Natural Language Processing (NLP) algorithms for detecting geriatric conditions are largely limited to domain-specific models that fail to capture the interdependent, multidimensional nature of comprehensive geriatric assessment. This study aimed to develop and evaluate a comprehensive, scalable, and robust information extraction framework to identify Comprehensive Geriatric Assessment (CGA) and Age-Friendly Health Systems (AFHS) 4Ms-related data elements from unstructured electronic health record (EHR) text across multiple health systems. Using a team science approach grounded in the TRUST framework, we annotated pooled clinical notes from four health systems to produce a gold-standard dataset of 41 CGA- and 4Ms-related geriatric care data elements. Three information extraction approaches were implemented and evaluated: an in-context learning generative large language model (GPT-4o), a hybrid heuristic-LLM model (MedAgingIE), and an instruction-tuned open-source lightweight model (Qwen2-7B-Instruct). Performance was assessed on a blinded test set using macro- and micro-averaged metrics. GPT-4o achieved a macro F1-score of 0.56 and micro F1-score of 0.87; MedAgingIE achieved 0.55 and 0.92; and Qwen2-7B-Instruct achieved 0.30 and 0.81, respectively. MedAgingIE demonstrated the strongest consistency between precision and recall, while GPT-4o showed superior sensitivity for diverse, context-rich geriatric concepts. These findings highlight key trade-offs among symbolic, generative, and instruction-tuned approaches for CGA and 4Ms phenotyping, suggesting that hybrid heuristic-LLM methods offer interpretability and stability, whereas large language models provide greater adaptability for complex clinical narratives.

Read PDF

Similar papers

Open access Sep 2026

Evaluating RAG Configurations for Clinical Information Extraction from EHR Notes: Aged Care Case Study

Clinical information extraction from unstructured electronic health records is important for supporting clinical decision making and healthcare research. However, large language models can struggle to accurately extract domain-specific information without effective adaptation. Retrieval-augmented generation offers a...

Dinithi S. Vithanage, Quang Vinh Duong, Chao Deng et al. · 0 citations
Conference Open access 2026

Natural Language Processing for Prediction of Chronic Diseases from Electronic Health Records

This approach combines semantic understanding of clinical narratives with structural modeling of patient-disease-treatment relationships and successfully validates synthetic EHR data utility for privacy-preserving healthcare AI development while addressing critical requirements necessary for clinical decision support s...

U. Luke, P. Asuquo, Victor Anaga et al. · 0 citations
Open access Sep 2026

Information retrieval in pre-hospital care with visualization-oriented natural-language interface via LLMs

A novel framework that leverages the language understanding and code generation ability of Large Language Models (LLMs) to build an information retrieval system with Visualization-oriented Natural-language-based Inter-faces (V-NLI).

Xin Gao, Zheng-Ye Zhu, Xin-Yu Ma et al. · 0 citations
Open access Sep 2026

Less Can Be Better: Decomposing Clinical Data Modalities in Large Language Model-based Healthcare Applications

The benefits of multimodal data integration are task-dependent and healthcare LLMs should examine clinical data modalities according to specific tasks for efficient integration, and provide practical guidance for designing efficient clinical decision support systems.

Cheng Peng, Mengxian Lyu, Ziyi Chen et al. · 0 citations
Open access Aug 2026

Conceptualising heart disease prediction through a unified framework combining clinical theory and machine learning models

The results show that it is feasible to make better predictions and gain valuable insights by merging these two types of data and that integrating unstructured data allows for a more holistic view of patient health, leading to earlier detection, personalized interventions, and improved decision-making in clinical setti...

Diana Olivia, Yarakam Shiva Chaitanya Reddy, Vibha Prabhu et al. · 0 citations
Review Open access Sep 2026

Practical Guide to Large Language Models for Information Extraction in Behavioral Health Notes: Tutorial

Abstract Background Mental health clinical notes contain decision-critical information often absent from structured electronic health record fields. Large language models (LLMs) can extract clinically relevant signals from narrative text; however, variability in output format, limited reproducibility, and inconsistent...

Diya Saha, J. Edgcomb · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.