Skip to content
Open access

A text mining and ontology-based approach using phenotypes to obtain relevant literature for rare diseases

Aug 2026 · iScience · Vol 29 · 0 citations · 78 references
Medicine

TL;DR

This work demonstrates how AI-driven informatics can enhance rare disease diagnosis and exemplifies the role of digital tools in transforming precision medicine and healthcare delivery.

Abstract

Summary Diagnosing rare diseases remains a major challenge due to limited clinical knowledge and the frequent absence of diagnostic criteria. We present a digital framework that leverages large language models and biomedical text embeddings to bridge this gap. By mapping Human Phenotype Ontology terms to a shared vector space with millions of PubMed abstracts and full-text articles, our method enables phenotype-driven semantic search and ranks literature relevant to patient symptoms, even without explicit disease mentions. Validated on OMIM-derived benchmarks and applied to RASopathies, including NF1, Noonan, and Costello syndromes, our approach retrieved expected findings, supporting differential diagnosis and research. The framework is implemented in an open-source Python package, py-semtools, and it can be integrated into clinical decision support systems or adapted to other ontologies and corpora. This work demonstrates how AI-driven informatics can enhance rare disease diagnosis and exemplifies the role of digital tools in transforming precision medicine and healthcare delivery.

Read PDF

Similar papers

A novel approach for Rare Disease symptom extraction from Clinical Texts Using Knowledge Graph Embeddings and Large Language Models

This study demonstrates how the new-age technologies, such as GenAI and Natura Language Processing (NLP) can aid in collecting valuable clinical data, improving patient safety and rare disease identification, as well as integrating natural language processing and graph-based reasoning to enhance disease recognition.

A. Dhull, Kartik Pal, Manoj Kumar · 0 citations
Open access Aug 2026

megaMine: a scalable, rule-based framework for mining gene-cancer-drug evidence from biomedical literature

The rapid expansion of the oncology literature has outpaced manual curation of clinically relevant gene-cancer-drug associations and oncogenic driver evidence. Existing automated approaches often lack transparency or are difficult to scale across heterogeneous data sources. To address this gap, we developed megaMine, a...

Muhammad Junaid, K. Prazanowska, Ha-Eun Jeong et al. · 0 citations
Open access Aug 2026

Systematic detection of predictive gene sets by semantics-based selection

The findings demonstrate that Gene Ontology can be effectively leveraged for semantics-aware gene selection in clinical machine learning and offers a scalable strategy to uncover new, testable biological hypotheses–revealing gene functions that might otherwise remain hidden when examining only broadly selected gene com...

Peter Eckhardt-Bellmann, N. Taha, Silke D. Werle et al. · 0 citations
Conference Open access Sep 2026

Structure-Aware Contrastive Learning for Biomedical Embeddings: Bridging the Gap Between HPO and Clinical Literature

A new embedding adaptation procedure is defined whose fine-tuning approach is guided by a novel "Disease-Overlap" similarity measure, which prioritizes clinical co-occurrence of phenotypes over taxonomic distance, and optimizes the embedding space using AnglE Loss to mitigate gradient saturation.

Jose L. Mellina Andreu, Alejandro Cisterna García, Juan A. Botía · 0 citations
Preprint Jul 2026

LLMBDC: Language Model for Biological Domains Oriented Clustering of Gene Ontology

LLMBDC (Large Language Model for Biological Domains Oriented Clustering of Gene Ontology) provides a scalable, reproducible, and interpretable route to context-aware, system-level interpretation of GO enrichment results while preserving biological specificity.

Xi-Ming Ran, Jie Xu, Peng Jin et al. · 0 citations
Open access Sep 2026

Ontology-aware knowledge graph retrieval-augmented generation for clinical decision support

Effectively retrieving and interpreting the vast, diverse, and largely unstructured data contained within electronic health records (EHRs) present significant challenges for clinical decision support systems. Large language models (LLMs), when applied to complex healthcare datasets, frequently exhibit hallucinations, l...

Deepak Panneerselvam, Sasikala E · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.