Skip to content
Conference Open access

Structure-Aware Contrastive Learning for Biomedical Embeddings: Bridging the Gap Between HPO and Clinical Literature

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · pp. 6832-6840 · 0 citations · 33 references

TL;DR

A new embedding adaptation procedure is defined whose fine-tuning approach is guided by a novel "Disease-Overlap" similarity measure, which prioritizes clinical co-occurrence of phenotypes over taxonomic distance, and optimizes the embedding space using AnglE Loss to mitigate gradient saturation.

Abstract

Large Language Models (LLMs) are extensively used at biomedical text processing but often fail to capture the complex, functional relationships encoded in expert knowledge graphs like the Human Phenotype Ontology (HPO). This "semantic gap'" limits their utility in precision medicine tasks such as rare disease diagnosis, where distinguishing overlapping clinical presentations requires understanding underlying pathophysiological connections rather than just surface-level textual similarity. In this work, we propose a Neuro-Symbolic Alignment Framework that bridges this separation by integrating literature-mined specialized phenotypical descriptions with the ontological structure used as reference. Specifically, we augment phenotype representations with automatically selected text fragments from massive corpus of descriptions mined from scientific literature (PubMed), overcoming the typical data scarcity of standard ontology definitions. We define a new embedding adaptation procedure whose fine-tuning approach is guided by a novel "Disease-Overlap" similarity measure, which prioritizes clinical co-occurrence of phenotypes over taxonomic distance, and optimizes the embedding space using AnglE Loss to mitigate gradient saturation. Extensive evaluations show that our approach significantly outperforms state-of-the-art baselines, including SapBERT, on both intrinsic semantic correlation and practical downstream tasks, including synthetic patient disease ranking and solving real cases stored in Phenopacket, where our model achieves x4 top-1 accuracy than the previous best model.

Read PDF

Similar papers

Open access Sep 2026

LLM-H2G: biomedical semantic-enhanced hypergraph contrastive learning for herb–disease association prediction

Herb–disease association prediction is central to computational traditional medicine, but existing graph and hypergraph methods mainly rely on observed topology and underuse biomedical textual semantics, especially in heterogeneous or sparse association networks. We propose LLM-H2G, a biomedical semantic...

Jun Zhang, Hengchuang Yin, Chao Wu et al. · 0 citations
Open access 2026

Knowledge Distillation for Biomedical Text Classification: A Systematic Comparative Analysis of Multiple Teacher–Student Architectures

Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.

Amine Gonca Toprak, Aytuğ Onan · 0 citations
Preprint Aug 2026

Neighborhood-Aware Dual Biomedical Entity Linking

A three-stage framework made up of neighborhood-aware retrieval, dual reranking, and score fusion that achieves the state of the art on average across five widely-used benchmarks and remains efficient at inference.

Yicheng Tao, Jie Liu · 0 citations
Open access Aug 2026

Benchmarking MeSH-augmented embeddings for biomedical document similarity

The extensive volume of biomedical scientific literature requires efficient methods for retrieving relevant documents based on semantic technologies and biomedical concepts. While embedding-based methods have shown improvements over traditional keyword-based methods, the integration of domain-specific terminologies lik...

Rohitha Ravinder, Lukas Geist, Nelson Quiñones et al. · 0 citations
#artificial intelligence Preprint Aug 2026

MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts

The proposed methodology selects and preprocesses a large corpus of scientific articles on malaria, and then annotates them with entities of clinical significance, and leverages BioBERT, a state-of-the-art pre-trained language model, to encode the textual data into context-aware representations.

V. Anoop, N. Devika · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.