This work presents TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways, and introduces new instruction tasks for capabilities that current corpora barely cover.
Abstract
The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.
It is found that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations, and it is shown that singletons share mutual information with clustered sequences, making them learnable by PLMs and...
R. Vinod, Samir Char, Ava A. Amini et al.· bioRxiv· 0 citations
Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining...
Mahinaz A. Mashhour, M. A. Wahed, Mai S. Mabrouk· Biochemical and Biophysical...· 0 citations
This survey reviews the basic principles of LLMs and summarizes representative applications in gene and genome sequence analysis, protein structure and function prediction, and drug design, including virtual screening and personalized medicine.
Zhi-Gang Meng, Zhi-Kai Yang, Mingming Zhu et al.· Frontiers in Genetics· 0 citations
Foundation models for biology achieve impressive pattern recognition on molecular sequences and single-cell transcriptomics, yet they fail to outperform simple linear baselines for predicting genetic perturbation effects, exposing a gap between statistical correlation and mechanistic understanding. This gap is compound...
This work introduces LAMBDA, a benchmark designed to rigorously evaluate genome language model embeddings through phage-bacteria sequence discrimination across four categories of increasing complexity: probing tasks, fine-tuning assessments, diagnostic tests, and genome-wide prophage detection.
LeAnn M. Lindsey, Nicole L. Pershing, K. Dufault-Thompson et al.· NAR Genomics and Bioinformat...· 0 citations
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs)....
R. Joeres, Ilya S. Senatorov, A. Kolchina et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.