Skip to content

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

Jul 2026 · arXiv.org · Vol abs/2607.08803 · 0 citations · 96 references
Computer Science Biology

TL;DR

This work presents TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways, and introduces new instruction tasks for capabilities that current corpora barely cover.

Abstract

The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.

View source

Similar papers

Open access Aug 2026

Protein language models and the long tail of functional diversity

It is found that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations, and it is shown that singletons share mutual information with clustered sequences, making them learnable by PLMs and...

R. Vinod, Samir Char, Ava A. Amini et al. · 0 citations
Review Jul 2026

Genomic language models (gLMs): Emerging applications, challenges, and future directions in computational genomics.

Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining...

Mahinaz A. Mashhour, M. A. Wahed, Mai S. Mabrouk · 0 citations
#protein folding Review Open access Aug 2026

Large language models in bioinformatics: a comprehensive survey

This survey reviews the basic principles of LLMs and summarizes representative applications in gene and genome sequence analysis, protein structure and function prediction, and drug design, including virtual screening and personalized medicine.

Zhi-Gang Meng, Zhi-Kai Yang, Mingming Zhu et al. · 0 citations
Open access Jul 2026

∑0–EvoCell: An AI-Native Ontology that Unifies Evolutionary and Cell Biology in Latent Space

Foundation models for biology achieve impressive pattern recognition on molecular sequences and single-cell transcriptomics, yet they fail to outperform simple linear baselines for predicting genetic perturbation effects, exposing a gap between statistical correlation and mechanistic understanding. This gap is compound...

Lurong Pan · 0 citations
Open access Sep 2026

LAMBDA: a prophage detection benchmark for genomic language models.

This work introduces LAMBDA, a benchmark designed to rigorously evaluate genome language model embeddings through phage-bacteria sequence discrimination across four categories of increasing complexity: probing tasks, fine-tuning assessments, diagnostic tests, and genome-wide prophage detection.

LeAnn M. Lindsey, Nicole L. Pershing, K. Dufault-Thompson et al. · 0 citations
#machine learning Preprint Aug 2026

Task- and dataset-specific information in protein language models

Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs)....

R. Joeres, Ilya S. Senatorov, A. Kolchina et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.