Skip to content
Conference Open access

Medical Text Knowledge Discovery and Clinical Decision Support Based on Natural Language Processing

2026 · ITM Web of Conferences · 0 citations

TL;DR

This paper summarizes the main challenges currently facing, including medical data privacy and labeling problems, interpretability and clinical credibility barriers of the model, and systemic barriers to multimodal fusion.

Abstract

Natural language processing (NLP) technology is the key driving force to unlock the value of massive, unstructured medical text data and promote the development of smart medicine. This paper aims to systematically review the application of NLP in the field of health care, focusing on how “medical text mining” drives “knowledge discovery” and ultimately serves “clinical decision support”. Firstly, this paper reviews the evolution of technology from the early rule method to the current pre training language model. Then, the core technologies such as medical information extraction, knowledge map construction, text classification and generation and their application in typical scenarios such as electronic medical record analysis, auxiliary diagnosis, prognosis prediction, and patient management are reviewed. Through the induction and comparison of existing studies, this paper summarizes the main challenges currently facing, including medical data privacy and labeling problems, interpretability and clinical credibility barriers of the model, and systemic barriers to multimodal fusion. Finally, this paper looks forward to the future research directions, such as the development of interpretable AI and the construction of NLP system for real-world evidence, in order to promote the transformation of this technology from research to safe, reliable and efficient clinical landing.

Read PDF

Similar papers

Review Open access Aug 2026

NATURAL LANGUAGE PROCESSING USING DEEP LEARNING: A COMPREHENSIVE REVIEW OF MODELS, TECHNIQUES, AND FUTURE DIRECTIONS

Natural language processing (NLP) has emerged as a key focus of AI research for the analysis, interpretation, extraction, summarisation, and generation of human language. The vast amount of unstructured textual data in scientific research, electronic health records, clinical notes, radiology reports, public health documents, and digital health platforms has driven the demand for sophisticated computational tools and techniques capable of extracting structured and actionable knowledge from language. NLP has been greatly advanced by deep learning, which allows for automatic representation learning, understanding context, modeling sequences, and generating large amounts of language by means of structures like CNN, RNN, LSTM, GRU, attention mechanisms, transformers, and large language models. This review aims to present a detailed overview of deep learning-based NLP models, methods, applications, challenges, and future directions, focusing on biomedical informatics, clinical text mining, digital health and biomathematical relevance. It has numerous applications such as biomedical literature mining, named entity recognition, relation extraction, clinical decision support, pharmacovigilance, radiology report generation, public health surveillance, and construction of knowledge graph. The specific focus lies in the application of NLP to identify biological entities, clinical variables and quantitative evidence that can be used to support biomathematical modeling. There are several current challenges such as domain shift, privacy, hallucination, bias, interpretability, and reproducibility. The success of future progress relies on reliable, comprehensible, domain specific and clinically verified NLP systems.

Dr. Pradeep Kumar Atulker, Dr. Rahul Kumar Hindustani, Ravi Shankar Nanduri et al. · 0 citations
Conference Open access 2026

Natural Language Processing for Prediction of Chronic Diseases from Electronic Health Records

This paper presents a comprehensive multi-modal artificial intelligence framework for the prediction of disease from electronic health records that integrates ClinicalBERT natural language processing with graph neural networks, temporal modeling and explainability analysis. Using Synthea synthetic EHR dat with SNOMED CT codes from 1,171 patients, our approach combines semantic understanding of clinical narratives with structural modeling of patient-disease-treatment relationships. The system achieves predictive performance with macro-averaged F1 score of 0.4512 and AUC of 0.9071 across six chronic conditions, demonstrating outstanding results for diabetes (F1=0.900) and hypertension (F1=0.949). Novel contributions include temporal progression forecasting over 12-month periods using LSTM-Transformer hybrid architecture and comprehensive explainability framework providing gradient-based feature importance analysis and automated clinical reasoning generation. The frameworks successfully validates synthetic EHR data utility for privacy-preserving healthcare AI development while addressing critical requirements necessary for clinical decision support system.

U. Luke, P. Asuquo, Victor Anaga et al. · 0 citations
Conference Jul 2026

Machine Learning-Augmented AI System for Medical Report Analysis, Clinical Summarization, and Diagnostic Decision Support

The extensive adoption of electronic health records has necessitated the development of automated systems that are capable of understanding unstructured clinical documents. Medical records, such as lab results, radiology findings, and discharge summaries, thus make manual analysis a slow and error-prone process. The paper introduces an AI-driven medical report analysis framework that employs natural language processing and deep learning to automatically locate and interpret the clinically significant information. The system proposed in this paper first preprocesses the medical text to identify the major entities such as diseases, symptoms, and drugs, and then translates them into structured clinical data. An attention-based neural model is used to produce brief analytical summaries, which help clinical decision-making. Experimentally, it was found that the proposed system not only outperformed the manual process in accuracy but also reduced the time. The framework, therefore, increases the efficiency of healthcare and opens up the potential for better utilization of electronic medical records.

Simranjit Singh Bedi, S. Kaswan, Sandeep Singh Kang · 0 citations
Open access Aug 2026

Using Natural Language Processing to Identify Adverse Drug Events Characterized by Medication Replacement in Primary Care Electronic Medical Records: Algorithm and Validation Study

Abstract Background Health care systems generate vast amounts of unstructured text, such as clinical notes, which capture nuanced patient experiences, clinical reasoning, and subtle indicators of health status. While health system research has traditionally relied upon structured data, natural language processing (NLP) enables the extraction of this rich textual information. Leveraging NLP could improve the identification and characterization of underreported adverse drug events (ADEs). Objective The primary objective of this study was to train and evaluate multiple NLP models, including both previously published architectures and a novel model, for the identification of ADEs from clinical notes. Methods Electronic medical records from the Manitoba Primary Care Research Network (MaPCReN) were used in this study. Clinical notes were annotated to indicate the presence of a possible ADE, the associated words or phrases, and the corresponding drug. A subselection algorithm was applied to ensure sufficient representation of notes containing ADEs for training a robust classifier. The cohort was restricted to patients aged 55 years and older and was annotated in 2 waves: Wave 1 comprised primary care encounter notes selected for a temporally linked emergency department (ED) visit, enriching it for acute presentations, while Wave 2 relaxed this requirement, and its acuity composition was uncharacterized. The annotated data were split into training and test sets. NLP models—including BioBERT, BlueBERT, a large language model (LLM) classifier, and an LLM embeddings–based classifier—were trained on both original clinical notes and notes reformatted into Subjective-Objective-Assessment-Plan (SOAP) structure. To approximate a clinician-inspired reasoning workflow, Mistral-7b-Instruct was used to extract presenting symptoms and generate a ranked list of potential etiologies. Model performance was evaluated using precision, recall, and F1-score. Results Of the 1085 annotated encounter notes, 355 included ADEs. Across 9 modeling approaches evaluated over 5 random seeds, models trained on SOAP-rewritten notes generally outperformed those trained on original notes. The SOAP Rewrite+ LLM Embeddings Classifier (GritLM-SOAP) achieved the highest mean F1-score (72.53%; 95% CI 62.40%‐81.82%), while BioBERT-SOAP weighted achieved the highest mean recall (76.34%; 95% CI 62.69%‐89.66%). End-to-end span-level extraction (named entity recognition+relation extraction) on ADE-positive test notes achieved a mean relaxed F1-score of 0.455, with relation extraction identified as the bottleneck. Conclusions Effective detection of ADEs in clinical notes may benefit from NLP models that approximate the clinical reasoning of health care providers. While the SOAP Rewrite+ LLM Embeddings Classifier demonstrated a reasonable balance of precision and recall, there is room for improvement as models evolve.

Alan Katz, Abhishek Dhankar, Gillian Fransoo et al. · 0 citations
Review Jul 2026

From Information Extraction to Clinical Reasoning: A Systematic Scoping Review of Large Language Models in Cancer Pathology Reports.

Pathology reports anchor cancer diagnosis and staging, yet their narrative structure limits reliable translation into structured, machine-actionable knowledge, creating a bottleneck between expert interpretation and scalable clinical intelligence. Despite decades of clinical natural language processing (NLP) research, pathology text remains among the most complex and consequential sources of medical data to operationalize at scale. Large language models (LLMs) offer new approaches for reading, extracting, and interpreting these reports. We synthesize current LLM work in cancer pathology using a four-level capability framework across the pathology report data lifecycle: (level 1) text preparation and quality checks, (level 2) information extraction, (level 3) guideline-based clinical reasoning, such as TNM staging and registry coding, and (level 4) interpretive synthesis, such as explanations, summarization, or decision support. Rather than grouping studies by NLP task labels, this framework tracks how LLM applications progress from preprocessing and extraction toward higher-level interpretation and synthesis. We followed PRISMA-ScR guidelines and searched four databases through September 2, 2025, identifying 41 eligible studies. Most studies focus on level 2 tasks, with fewer addressing level 3 and level 4 tasks. Encoder-based models, including domain-specific variants such as BioBERT, were commonly used for structured extraction tasks, whereas generative models, including GPT, LLaMA, and Mistral-family models, were increasingly evaluated for prompting-based extraction, staging, and summarization. Reported performance was often high for well-defined extraction tasks, but external validation was uncommon, and metrics varied across studies, limiting direct comparison. Overall, the evidence suggests that success in lower capability levels does not consistently translate to higher-level reasoning, especially when reports are inconsistent, required staging inputs are missing, or clinical assumptions must be inferred, which helps explain gaps between benchmark results and practical adoption. Future work should prioritize robust multi-site validation, clinically meaningful error analysis, transparent evaluation, and privacy-preserving implementation strategies to support safe integration in oncology.

Maryam Seifaddini, Mohammad Beheshti, Steven Richberg et al. · 0 citations
Open access Jul 2026

Clinician expertise and prompt engineering enhance cancer information extraction in electronic health records by small language models.

BACKGROUND Real-world data (RWD) in unstructured electronic health records (EHRs) is crucial for understanding complex diseases like cancer, but extracting structured information is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four small, locally deployable language models for information extraction from Italian EHRs. METHODS We examine three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess the impact of prompt design on accuracy. We evaluate the performance of four open-source small language models (SLMs) for clinical information extraction from Italian electronic health records (EHRs) in the APOLLO 11 trial on non-small cell lung cancer (NSCLC). The extraction protocol involves four steps: problem definition, data preprocessing, Large Language Model (LLM)-based information extraction, and output evaluation. RESULTS We show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM (Tumor, Node, Metastasis) staging, PD-L1 (Programmed death-ligand 1), and ECOG-PS (Eastern Cooperative Oncology Group-Performance Status) are more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improve performance and reduced hallucinations. Clinical expertise enhances consistency in the extraction, particularly among students using annotated examples. CONCLUSIONS The study confirms that privacy-preserving SLMs can be deployed locally for efficient and secure cancer data extraction. Findings highlight the need for hybrid systems combining SLMs with expert input and underline the importance of aligning clinical documentation practices with SLM capabilities. This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.

Federica Corso, V. Peppoloni, L. Mazzeo et al. · 0 citations