Benchmarking commercial large language models for gene-disease-phenotype extraction from full-text human genetics literature
Manual curation of gene-disease-phenotype relationships from the human genetics literature is a persistent bottleneck for maintaining its bioinformatics databases. Whereas large language models (LLMs) offer a promising alternative, there is currently no systematic benchmark that evaluates whether state-of-the-art comme...