Skip to content
Book Open access

BioFlowBench: A Comprehensive Benchmark for Evaluating Bioinformatics Tool-use Capabilities of LLMs and Agents

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 9059-9070 · 1 citation · 22 references

Abstract

The rapid advancement of high-throughput technologies has led to an explosion of biological data and a subsequent surge in bioinformatics analysis tools, thereby creating an urgent demand for automated bioinformatics workflows. Recently Large Language Models (LLMs) and LLM-based agents show great potential in this area. However, existing benchmarks primarily focus on static question-answering (QA) tasks, failing to capture the knowledge-action gap between understanding tool usage and executing complex bioinformatics workflows. Furthermore, current evaluation paradigms often prioritize algorithmic success rates, while neglecting the biological validity. Moreover, the construction of execution benchmarks is challenging due to complex environmental dependencies and the high cost of manual annotation, leading to poor scalability. In this study, we propose BioFlowBench, a comprehensive benchmark designed to shift from static knowledge assessment to dynamic execution evaluation in bioinformatics tool utilization. First, we construct a multi-layered dataset consisting of 5,071 test samples, including Syntax Understanding, Contextual Application and Real-world Execution. Second, we introduce BioGen, an agent-based pipeline designed for the automated generation of executable benchmarks. By creating compact, low-overhead synthetic data, BioGen facilitates low-cost and large-scale testing. Third, we propose a multi-dimensional evaluation framework comprising static knowledge, structural integrity, functional validity, and efficiency metrics. Our experiments reveal that: (1) A significant gap exists between static QA and dynamic execution tasks, with top LLMs perform well on static QA but falter in real-world execution scenario; (2) specialized agents outperform general models in real-world execution through environmental interaction and iterative refinement; and (3) domain knowledge remains the primary bottleneck, often leading to executable but biologically inaccurate outputs. The code is available at: https://github.com/YufeiHouAnne/BioFlowBench and the dataset can be accessed at: https://www.scidb.cn/detail?dataSetId=aee284681d674f53bfc6dae44635e773.

Read PDF

Similar papers

Book Open access Aug 2026

Benchmarking LLM Agents on Real-World Biological Database Curation for Data-Driven Scientific Discovery

BioDataLab evaluates the capability of autonomous agents to transform raw, heterogeneous biological resources into structured, analysis-ready databases, and underscores that while LLMs are proficient in downstream reasoning, autonomous upstream curation remains a formidable frontier.

Jiaxian Yan, Xi Fang, Jintao Zhu et al. · 0 citations
Open access Sep 2026

PromptBio: An Agentic Platform for End-to-End Computational Biomedical Research

Modern biomedical research increasingly depends on complex computational analyses, yet translating a scientific question into a reliable workflow still requires substantial technical expertise and manual coordination. PromptBio is a multi-agent AI platform available through a web portal at https://promptbio.ai that add...

Min-Zhe Zhang, Wen-Hao Gu, Bo-Wei Han et al. · 0 citations
Book Open access Aug 2026

MicrobeQuest: A Context-Aware Multimodal Benchmark for Microbiology Information Extraction

AI for Science (AI4S) is rapidly advancing scientific discovery, yet its progress critically depends on the availability of large-scale, high-quality structured scientific data and reliable evaluation benchmarks. In microbiology, abundant multimodal information—spanning text, images, tables, and charts—is distributed a...

Ou Zheng, Xue Ren, Xuexia Su et al. · 0 citations
Review Open access Aug 2026

TaxoFlow: a step-by-step tutorial to build a nextflow pipeline for metagenomics taxonomic classification

An open, interactive and web-based tutorial that guides scholars with basic command-line skills through the detailed development of a validated and reproducible Nextflow metagenomics classification pipeline, which aims to lower technical barriers in microbiome bioinformatics and promote best practices in metagenomics d...

Jeferyd Yepes-García, Laurent Falquet · 0 citations
Open access Sep 2026

BioTester: an AI-driven automated testing framework for identifying potential quality risks in bioinformatics software

Abstract Bioinformatics software plays a critical role in clinical applications such as cancer screening and genetic disease diagnosis, where comprehensive quality management is essential for ensuring the accuracy and reliability of downstream analysis. However, current validation practices rely heavily on manually des...

Xin Lian, Jia-Yin Wang, Xiao-Yi Zhu et al. · 0 citations
Review Open access Aug 2026

Detailed curation of biological samples and experimental designs for genomics using LLM-supported agentic workflows

An automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource, with performance near that of human curators, at approximately 1/20th the cost and at least 100 times the speed.

P. Pavlidis, B. O. Mancarci, A. Mãximo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.