Skip to content
Book Open access

Benchmarking LLM Agents on Real-World Biological Database Curation for Data-Driven Scientific Discovery

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 47 references

TL;DR

BioDataLab evaluates the capability of autonomous agents to transform raw, heterogeneous biological resources into structured, analysis-ready databases, and underscores that while LLMs are proficient in downstream reasoning, autonomous upstream curation remains a formidable frontier.

Abstract

High-quality biological databases are the bedrock of data-driven scientific discovery. However, the construction of these resources remains a labor-intensive bottleneck, particularly for emerging research frontiers where structured data is non-existent. While LLM-based agents have catalyzed progress in downstream scientific modeling, their potential to automate the critical upstream challenge of database curation remains largely untapped. To bridge this gap, we introduce BioDataLab, a rigorous benchmark comprising 100 tasks meticulously derived from 57 high-impact database publications. BioDataLab evaluates the capability of autonomous agents to transform raw, heterogeneous biological resources into structured, analysis-ready databases. Unlike static evaluations, BioDataLab provides a fully interactive environment encompassing data retrieval, extraction, annotation, and integration, featuring process-oriented curation targets and contamination-control checks. We benchmark 11 state-of-the-art LLMs (including Gemini-3.0, GPT-5.2, and Claude-4.5) under different agent frameworks, revealing a substantial capability gap: the top-performing model achieves only a 40% success rate. Further error analysis identifies significant bottlenecks in multi-step tool orchestration and adherence to complex biological data formats. These findings underscore that while LLMs are proficient in downstream reasoning, autonomous upstream curation remains a formidable frontier. All data and codes are available at GitHub.

Read PDF

Similar papers

Review Open access Aug 2026

Detailed curation of biological samples and experimental designs for genomics using LLM-supported agentic workflows

An automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource, with performance near that of human curators, at approximately 1/20th the cost and at least 100 times the speed.

P. Pavlidis, B. O. Mancarci, A. Mãximo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AutoKD: Autonomous Knowledge Discovery

Scientific discovery in data-rich domains is currently constrained by human bandwidth: the growth in the volume and complexity of real-world data far outpaces the rate at which researchers can read, reason, and synthesize. Recent LLM-based multi-agent systems have begun to automate portions of the research cycle, but t...

Qin-Wen Ge, B. Ni, Hao Fu et al. · 0 citations
Preprint Aug 2026

ChemReporter: A Framework for Curating and Exporting Large-Scale Chemical Datasets for MLIP Training

Training set quality and diversity are key determinants of the reliability of machine learning interatomic potentials (MLIPs), yet using massive datasets in full is often impractical and redundant, making intelligent data selection essential. A major bottleneck, however, is the lack of infrastructure for uniformly acce...

M. Bluntzer, Jules Tilly, Christoph Brunken · 0 citations
Book Open access Aug 2026

Automating End-to-End Hybrid Query Processing: Benchmark, Solution, and Insights

Hybrid queries—natural language questions over structured data that require both database capabilities and LLM reasoning—have recently emerged as a prominent research topic. However, existing solutions remain overly dependent on manual workflows, and current benchmarks are limited in scale and diversity. To bridge this...

Bo Li, Chenzhan Wang, Long-Kang Lin et al. · 0 citations
Preprint Aug 2026

CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications

The system architecture, core workflows, and benchmarks that reach literature performance for quantum mechanical, physicochemical and bioactivity property prediction, and use cases involving time series datasets demonstrating applications beyond molecular chemistry datasets are described.

Brendan Smith, S. López-Moreno, E. Dolores-Cuenca et al. · 0 citations
Open access Aug 2026

Automating scientific annotations for open transcriptomic profiles via multi-stage agents

GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.

Xiaodan Zhang, S. Paithankar, Jing Pu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.