Skip to content
Open access

Evaluating Large‐Language Models in Bioinformatics Applications

Sep 2026 · Advanced Computing · 0 citations · 31 references

Abstract

Large language models (LLMs) have significantly revolutionized natural language processing through their strong capabilities in text generation and reasoning. Yet, their applicability to bioinformatics applications remains largely unexplored. Here, we systematically evaluate state‐of‐the‐art LLMs across six representative task domains: drug–drug interaction prediction, antimicrobial and anticancer peptide identification, molecular optimization, gene and protein named entity recognition, single‐cell type annotation, and bioinformatics question answering. Our results show that general‐purpose LLMs can deliver competitive performance across most tasks, demonstrating their versatility for biological data analysis under minimal task‐specific adaptation. Meanwhile, our findings underscore critical limitations for the usage of LLMs in bioinformatics research. First, LLMs exhibit limited capability in chemical structure reasoning for molecular optimization, highlighting the need for integrating physics‐based structure constraints into generative modeling. Second, their sensitivity to ambiguity in gene/protein recognition and single‐cell annotation emphasizes the importance of domain‐specific knowledge in prompt design. Finally, their difficulty in answering computational biology questions suggests that future systems may benefit from combining LLMs with structured reasoning frameworks to support more biologically rigorous inference and decision‐making. These findings reveal both the promise and the boundaries of current LLMs in bioinformatics, offering a roadmap for advancing next‐generation LLM models tailored to the bioinformatics applications.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.