Skip to content

HypoKG: Evidence-Disciplined Biomedical Hypothesis Generation Beyond Endpoint Knowledge

Sep 2026 · 0 citations · 38 references
Computer Science Biology

TL;DR

These findings show that knowledge graphs support hypothesis generation in two ways: they identify biological endpoint pairs absent from the literature, and their mechanistic paths guide how LLMs reason between them.

Abstract

Large language models (LLMs) can generate biomedical hypotheses, but it remains unclear whether they truly reason from scientific evidence or simply produce convincing-sounding ideas. To study this, we combine three major biological databases: the Kyoto Encyclopedia of Genes and Genomes (KEGG), Rhea, and UniProt, into a unified biochemical knowledge graph and construct a benchmark of 550 paths connecting enzyme sources to rare disease endpoints, yielding 13,200 hypotheses from six LLMs under four conditions varying the biological information each model receives: source enzyme only, full biological path, or source and disease endpoint only. Hypotheses are scored using an expert-derived five-criterion rubric on a 1-5 scale per criterion. We find that models given both the source and disease endpoint often produce the highest-scoring hypotheses, showing that LLMs can generate compelling ideas from minimal information. However, these hypotheses are less grounded in the evidence. In contrast, models given the full biological path generate hypotheses more consistent with known mechanistic relationships. We call this evidence-disciplined reasoning. To confirm this effect, we shuffled intermediate path steps while keeping endpoints fixed. Evidence grounding dropped significantly (delta = -0.793, p<0.001), confirming models genuinely used path structure during reasoning. Our findings show that knowledge graphs support hypothesis generation in two ways: they identify biological endpoint pairs absent from the literature, and their mechanistic paths guide how LLMs reason between them.

View source

Similar papers

#artificial intelligence Review Aug 2026

Medical Causal Hypothesis Verification with Large Language Models

It is shown that while LLMs exhibit strong recall, they often perform poorly at providing valid scientific articles and evidence for support and at rejecting unsupported hypotheses, highlighting the need for rigorous evaluation before using LLMs for search and retrieval in healthcare settings.

Safiyyah Ahmed, Abrar Ansari, M. Islam et al. · 0 citations
Conference Open access Sep 2026

THGAgents: Traceable Biomedical Hypothesis Generation via Dynamic Causal Reasoning

THGAgents utilizes collaborative and dynamically updating agents to build a Traceable Causal Knowledge Graph, which serves as the foundation for the evidence-based knowledge structure and employs an LLM-driven heuristic search algorithm to traverse the complex network, balancing both novelty and rigor to deduce strict,...

Ming-Jia Yang, Kun-Hua Dong, K. Lim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Semantic Model of Genetic Evidence: A Step Toward Bridging the Basic-Science-Clinic Gap

A semantic model for scientific evidence with three core classes is introduced, specialize it for genetics, align it structurally to FHIR Evidence with a SEPIO-anchored credibility decomposition, and attach a compact dimensional vocabulary whose conditional-activation rules are validated by a SHACL schema for the imple...

M. Bouzinier, D. Etin · 0 citations

Optimizing large language model prompts for biomedical knowledge discovery

This work presents a scalable, reproducible framework for evaluating, optimizing, and interpreting LLMs for biomedical knowledge extraction, with a focus on gene–gene regulatory relation prediction, pathway component recognition, multimodal pathway figure understanding, and automated prompt optimization.

Muhammad Azam · 0 citations
#artificial intelligence Preprint Sep 2026

DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery

We introduce DoAtlas-2, a foundation for self-evolving causal biomedical discovery that organizes knowledge around causal mechanisms and advances through external evidence from human populations. DoAtlas-2 integrates 771 research resources covering more than 720,000 participants in 48 countries, from longitudinal clini...

Yu-Long Li, Rong Xia, Yuxuan Zhang et al. · 0 citations
#natural language process... Preprint Aug 2026

Quantitative Evidence Mining for Plausibility-Aware Biomedical AI

This work outlines a framework for plausibility-aware AI that treats extracted claims not as final answers but as auditable evidence objects, making clear what was measured, how much it changed, in which setting, with what uncertainty, and from which source.

N. S. Babaiha, Stefan Geißler, Marie-Christine Simon et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.