Biology

HypoKG: Evidence-Disciplined Biomedical Hypothesis Generation Beyond Endpoint Knowledge

How the science connects

Natural language p…Knowledge graphBiomedical informa…

AI Insight

This study investigates whether large language models (LLMs) generate biomedical hypotheses through genuine scientific reasoning or pattern matching. Researchers created a knowledge graph linking 550 enzyme-disease pathways and tested six LLMs under different information conditions, finding that models given only endpoints produced compelling but less evidence-grounded hypotheses, while those given full biological pathways generated hypotheses more consistent with known mechanisms—termed "evidence-disciplined reasoning." Shuffling pathway steps significantly reduced evidence grounding, confirming models actually use structural biological information during reasoning.


This research demonstrates that knowledge graphs can enhance AI-driven scientific discovery by both identifying novel biological connections and guiding mechanistically sound hypothesis generation. The findings have important implications for developing AI tools that support biomedical research while maintaining scientific rigor and reducing hallucination risks.


Understand the Science

Natural language processing 67 articles Explore Concept → Knowledge graph Concept coming soon Biomedical informatics Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Large language models (LLMs) can generate biomedical hypotheses, but it remains unclear whether they truly reason from scientific evidence or simply produce convincing-sounding ideas. To study this, we combine three major biological databases: the Kyoto Encyclopedia of Genes and Genomes (KEGG), Rhea, and UniProt, into a unified biochemical knowledge graph and construct a benchmark of 550 paths connecting enzyme sources to rare disease endpoints, yielding 13,200 hypotheses from six LLMs under four conditions varying the biological information each model receives: source enzyme only, full biological path, or source and disease endpoint only. Hypotheses are scored using an expert-derived five-criterion rubric on a 1-5 scale per criterion. We find that models given both the source and disease endpoint often produce the highest-scoring hypotheses, showing that LLMs can generate compelling ideas from minimal information. However, these hypotheses are less grounded in the evidence. In contrast, models given the full biological path generate hypotheses more consistent with known mechanistic relationships. We call this evidence-disciplined reasoning. To confirm this effect, we shuffled intermediate path steps while keeping endpoints fixed. Evidence grounding dropped significantly (delta = -0.793, p < 0.001), confirming models genuinely used path structure during reasoning. Our findings show that knowledge graphs support hypothesis generation in two ways: they identify biological endpoint pairs absent from the literature, and their mechanistic paths guide how LLMs reason between them.

Source: HypoKG: Evidence-Disciplined Biomedical Hypothesis Generation Beyond Endpoint Knowledge