Skip to content
Review

Genomic language models (gLMs): Emerging applications, challenges, and future directions in computational genomics.

Jul 2026 · Biochemical and Biophysical Research Communications - BBRC · Vol 831, pp. 154321 · 0 citations · 69 references
Medicine

TL;DR

Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining.

Abstract

Genomic language models (gLMs) are rapidly becoming important tools for learning biological information directly from sequence data. By adapting concepts from natural language processing, these models aim to capture contextual dependencies, regulatory grammar, evolutionary constraint, and sequence-level functional patterns that may be difficult to detect using alignment-based, motif-based, or conventional supervised methods alone. This systematic review evaluates recent model-development studies of genomic, RNA, nucleotide, codon-level, and regulatory DNA language models, with emphasis on model architecture, tokenization, training objective, biological task, benchmarking strategy, and reported limitations. A structured search of PubMed, Scopus, and Web of Science identified 469 records. After duplicate removal, screening, and full-text eligibility assessment, 58 studies met the strict inclusion criteria for primary model development or substantial model adaptation. The included studies covered diverse applications, including regulatory sequence prediction, variant-effect modeling, genome annotation, microbial and viral genome analysis, RNA splicing and regulation, codon optimization, mRNA design, and generative design of regulatory or RNA sequences. Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining. However, the evidence was heterogeneous and did not support a general claim of superiority over established bioinformatics tools or specialized supervised models. In several regulatory genomics tasks, specialized supervised models, k-mer-based approaches, or conventional deep-learning baselines remained competitive or superior to pretrained language-model representations. Generative models showed growing promise for RNA, codon, viral genome, and cis-regulatory element design, although many were evaluated mainly in silico. Overall, the field is advancing quickly, but broader impact will require standardized benchmarks, clearer reporting, stronger external validation, improved interpretability, and experimental confirmation of predicted or generated biological functions.

View source

Similar papers

Review Open access Jul 2026

Decoding viral protein sequences by large language models

This mini-review summarizes recent developments in devising and applying protein language models for biological sequences, emphasizing viral protein analysis, and outlines a road map for the potential application of LLMs in empowering virology research and pathogen surveillance.

Tianyi Fei, Siqi Li, Ziyue Yang et al. · 0 citations
Review Open access Jan 2024

Advancing bioinformatics with language models: components, applications, and perspectives

This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, drug discovery, and single-cell analysis, and highlights major challenges that remain insufficiently addressed in prior reviews.

Jiajia Liu, Mengyuan Yang, Yankai Yu et al. · 41 citations · ⚡4
Review Open access Jul 2026

SNP Detection Strategies in Genomic Research: A Comparative Review of Major Tools, Algorithms, Challenges and Applications

This review compares SNP detection programs such as GATK, BCFtools, FreeBayes, SAMtools, SAMtools, and DeepVariant and their algorithmic structures, namely pileup- based, haplotype-based, and machine-learning approaches and suggests that no single tool is the best.

Shikhi Baruri, Sunita Khanal · 0 citations
Open access Jul 2026

An embedding-based framework enables statistical testing of gene-set function hypotheses inferred by large language models

An embedding-based statistical framework is developed that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs.

Yanhao Tan, Li-Ju Wang, Tianyuzhou Liang et al. · 0 citations
Open access Aug 2026

Protein language models and the long tail of functional diversity

It is found that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations, and it is shown that singletons share mutual information with clustered sequences, making them learnable by PLMs and useful for training.

R. Vinod, Samir Char, Ava A. Amini et al. · 0 citations
Preprint Aug 2026

Frozen but Not Always Accessible: A Representation Analysis of Genomic Language Models

A representation-accessibility analysis of frozen genomic language models across regulatory, epigenetic, promoter, splice-site, and variant-effect prediction tasks shows that local biological signal is partially present in frozen representations, but is not always accessible through final pooled embeddings.

Nirjhor Datta, Swakkhar Shatabda, M. S. Rahman · 0 citations