Benchmarking commercial large language models for gene-disease-phenotype extraction from full-text human genetics literature
Abstract
Manual curation of gene-disease-phenotype relationships from the human genetics literature is a persistent bottleneck for maintaining its bioinformatics databases. Whereas large language models (LLMs) offer a promising alternative, there is currently no systematic benchmark that evaluates whether state-of-the-art commercial LLMs can perform this task reliably on the full-text articles. To address this gap, we introduce a standardized benchmark comprising 406 full-text articles covering 180 congenital heart disease-associated genes, and a multi-dimensional evaluation framework that incorporates fuzzy matching to account for synonyms and partial matches. We benchmarked seven state-of-the-art LLMs, GPT-4o, Claude-Opus-4, DeepSeek-R1, Grok-4, Qwen-3.5, Gemini-2.5 (Pro), and GPT-5 on the extraction of structured gene, disease, and phenotype fields. The top-performing model, Grok-4, achieved 97.6% overall accuracy, whereas the lowest-performing model reached approximately 88%, still surpassing many prior benchmarks employing zero-shot or n-shot prompting in biomedical relation extraction (RE) tasks. Our results provide a rigorous characterization of current LLMs capabilities and limitations. This paper contains two components. First, we conducted a human genetics field benchmark study on LLMs against a curated database. Second we developed the evaluation framework for this task. The benchmark dataset, evaluation framework, and model benchmarking outputs are made available online to support future studies in a reproducible manner.