BioFlowBench: A Comprehensive Benchmark for Evaluating Bioinformatics Tool-use Capabilities of LLMs and Agents
A significant gap exists between static QA and dynamic execution tasks, with top LLMs perform well on static QA but falter in real-world execution scenario; specialized agents outperform general models in real-world execution through environmental interaction and iterative refinement; and domain knowledge remains the p...