Skip to content
Open access

Intelligent Framework for Adverse Drug Event Identification Using Large Language Models and Retrieval-Augmented Generation: Development and Evaluation Study

Aug 2026 · Journal of Medical Internet Research · Vol 28 · 0 citations · 33 references
Medicine

TL;DR

Synergizing a curated domain-specific knowledge base with LLMs via a RAG architecture is an effective strategy for accurately identifying ADEs in unstructured Chinese clinical notes, providing a foundational open-source benchmark and a robust technical framework to advance pharmacovigilance, drug safety research, and clinical decision support.

Abstract

Abstract Background Adverse drug events (ADEs) pose significant public health challenges and economic burdens. While substantial ADE information is documented in unstructured clinical notes, its extraction remains difficult due to semantic complexity. Large language models (LLMs) offer promising text comprehension capabilities but are often hindered by domain-specific hallucinations. Objective This study aims to evaluate the effectiveness of retrieval-augmented generation (RAG) in improving the identification of ADEs using LLMs from Chinese clinical narratives and to establish a paradigm for this task. Methods We collected and preprocessed 19,983 Chinese clinical notes, retaining 18,432 high-quality records. Following a rigorous annotation and deduplication process, we established a gold-standard reference dataset (n=2510) and an ADE knowledge base (n=5144) using a standardized JSON schema. We evaluated 3 state-of-the-art LLMs (DeepSeek-V3 [DeepSeek], ERNIE 3.5-8K [Baidu], and GPT-4o [OpenAI]) under 3 prompt strategies: nonaugmented generation (NAG), static-augmented generation (SAG), and RAG. Performance was comprehensively assessed using precision, recall, and F1-score across 3 recognition matching levels (L1 exact, L2 sentence, and L3 overlap) via 1000 bootstrap resamples. Model robustness was further validated from real-world clinical progress notes, reflecting real-world ADE prevalence. Results We successfully constructed and publicly released the first Chinese ADE corpus derived from clinical notes. Across the tested LLMs, RAG yielded higher F1-scores than NAG and SAG at the L3 level. The optimal configuration, DeepSeek-V3 with RAG, achieved an overall L3-level F1-score of 0.9638 (95% CI 0.9541‐0.9727). Notably, the RAG approach increased the recall of GPT-4o from 0.6419 under NAG to 0.9241 under RAG (FDR P=.003). Evaluation on real-world datasets demonstrated clinical utility, with the RAG prompt maintaining high discriminatory capability (specificity: 0.9821; F2-score: 0.8885). Error analysis revealed that RAG successfully resolved common identification errors, both omissions and commissions, that were intractable for nonaugmented models. Conclusions Synergizing a curated domain-specific knowledge base with LLMs via a RAG architecture is an effective strategy for accurately identifying ADEs in unstructured Chinese clinical notes. This approach can mitigate hallucinations in LLMs, providing a foundational open-source benchmark and a robust technical framework to advance pharmacovigilance, drug safety research, and clinical decision support.

Read PDF

Similar papers

Open access Aug 2026

Using Natural Language Processing to Identify Adverse Drug Events Characterized by Medication Replacement in Primary Care Electronic Medical Records: Algorithm and Validation Study

Effective detection of ADEs in clinical notes may benefit from NLP models that approximate the clinical reasoning of health care providers as models evolve.

Alan Katz, Abhishek Dhankar, Gillian Fransoo et al. · 0 citations
Review Open access Aug 2026

Retrieval-Augmented Large Language Models for Clinically Aligned Adverse Event Coding in Acute Myeloid Leukemia Clinical Trials

Evaluation across three AML clinical trials showed that strict LLT-level string agreement underestimated clinical ap-propriateness, highlighting the importance of combining hierarchical evaluation metrics with clini-cal expert validation for AI-assisted MedDRA coding in hematology trials.

N. Dashti, M. Schneider, J. Eckardt et al. · 0 citations
Review Open access Nov 2025

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide...

Elena A. Mourelatou, Ioannis Katakis · 0 citations
Open access Aug 2026

Retrieval-augmented generation for medical question answering: a multi-metric performance evaluation

The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.

Yunus Kökver · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.