Skip to content
Open access

CENO: A Genome-Scale World Model for Evolutionary Sequence Interpretation and Programmable Regulatory Design

Jul 2026 · bioRxiv · 0 citations
Biology

TL;DR

CENO is introduced, a family of long-context generative genomic world models designed to preserve local DNA grammar while extending usable context to regulatory and chromatin scales and provides a genome-scale sequence world-model framework for sequence interpretation, evolutionary reasoning, gene-scale reconstruction and programmable regulatory sequence generation.

Abstract

DNA encodes biological function across a continuum of sequence scales, from single-nucleotide and motif-level grammar to regulatory neighborhoods, chromatin-scale organization and evolutionary constraint. A useful model of genomes should therefore do more than classify short sequence windows: it should maintain nucleotide-resolution state over long contexts, score counterfactual mutations, condition on homologous sequence evidence and generate candidates that can be evaluated against structural or functional objectives. We define such a system operationally as a genomic world model: a general-purpose generative model of genome sequence space that unifies sequence understanding and sequence design through a shared state and likelihood interface. Here we introduce CENO, a family of long-context generative genomic world models designed to preserve local DNA grammar while extending usable context to regulatory and chromatin scales. CENO combines Mamba sequence-mixing layers, sparse attention layers and mixture-of-experts capacity in a single autoregressive backbone, and is trained at 300M, 600M and 1B parameter scales with a staged curriculum that progresses from 8k-token cross-domain genomic pretraining to 131k- and 1M-token whole-genome long-context continuation. We evaluate CENO under a unified world-model benchmark paradigm spanning retrieval, representation, counterfactual perturbation, reconstruction, evolutionary conditioning and design. CENO retains practical long-context inference and retrieves distal sequence in synthetic assays. In zero-shot long-context analyses, without task-specific fine-tuning, long-context continuation yields annotation- and chromatin-boundary-associated attention patterns and frozen-state representations that generalize across human cell types and mouse cell or tissue settings. To incorporate evolutionary information, we further post-train CENO on packed real multiple-sequence-alignment contexts and score variants by reference–mutant likelihood deltas, improving matched variant-effect prediction and producing evolutionary enrichment signals across species. Complementing these perturbation-based variant tests, we evaluate zero-shot long-sequence generation by partial-gene continuation, asking whether the model can recover withheld gene-scale sequence structure across eukaryotic, bacterial and archaeal species; recovery improves with model scale and later whole-genome long-context training. Finally, we use CENO as the backbone for a cell-type-specific enhancer design workflow in mouse cortex, coupling a CENO-based accessibility oracle with conditional supervised fine-tuning and oracle-guided reinforcement learning. Together, CENO provides a genome-scale sequence world-model framework for sequence interpretation, evolutionary reasoning, gene-scale reconstruction and programmable regulatory sequence generation.

Read PDF

Similar papers

Open access Jun 2026

Unlocking cis-regulatory landscapes across 500 million years of evolution and disease mechanisms

Abstract Genomic DNA encodes regulatory information that determines where, when, and to what extent genes are expressed. Theoretically, we should be able to identify these transcriptional “instructions” by examining genomic DNA sequence alone, yet this has remained challenging. Here we present the Vertebrate Regulatory MOdule Detector (VRMOD), a method that accurately predicts gene regulatory sequences using only the query genomic sequences. We applied VRMOD to 309 Ensembl genomes, generating a compendium of high-resolution, genome-position-fixed cis-regulatory modules without parameter tuning. We performed extensive computational evaluation and experimental validation of VRMOD predictions. Notably, VRMOD predicted three sub-enhancers within the human hs52 enhancer at the FTO locus from the VISTA database, including one missed by existing methods. Using a chicken embryo system and 3D tissue imaging, we showed that each sub-enhancer exhibits restricted spatiotemporal activity within specific subsets of tissues where the full enhancer is active. We further demonstrated VRMOD’s utility for identifying evolutionarily non-conserved enhancers, annotating regulatory sequences in non-model organisms, and identifying candidate disease-causal variants. Collectively, VRMOD provides a universal coordinate reference system for regulatory sequences across 309 vertebrate genomes and enables genome-wide annotation of non-coding regulatory elements in any vertebrate species using genomic sequence alone.

Tássia Mangetti Gonçalves, Casey L. Stewart, Samantha D. Baxley et al. · 0 citations
Review Open access Aug 2026

Deep Learning for Deciphering the Plant Cis-Regulatory Code

This review compares convolutional, Transformer-based and graph architectures used to represent local sequence features, chromatin state and three-dimensional genome organisation to their applications to transcription-factor binding, chromatin accessibility, gene expression, non-coding variant prioritisation and regulatory-sequence design.

Zhi-Meng Zhao, Si-Xuan Huang, Shi-Long Zhang et al. · 0 citations
Open access Sep 2026

Predicting genome-wide functional constraints with GPN-Star.

Genomic language models have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences1. However, standard genomic language models adapted from natural language processing often require large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks2-4. Here we introduce a genomic pretrained network with species tree and alignment representations (GPN-Star), which is a biologically grounded genomic language model featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammal and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales show task-dependent advantages of modelling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms previous methods in prioritizing pathogenic and fine-mapped genome-wide association study variants, yields strong enrichments of complex trait heritability and improves power in rare variant association testing5. Extending beyond humans, we train GPN-Star for five model organisms-Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana-demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful and flexible tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Cheng-Zhong Ye, Gonzalo Benegas, Carlos Albors et al. · 0 citations
Open access Sep 2026

NucleicBERT interprets RNA sequence space through self-supervised language modelling

Much of the human genome’s non-protein-coding fraction acts directly through RNA, yet the structural and functional roles encoded in these sequences remain poorly understood. Applying deep learning is hindered by scarce RNA structural data and it remains unclear what biological constraints such models can recover directly from the abundant RNA sequences alone. Here, to address these challenges, we developed NucleicBERT, a self-supervised masked-language model that learns contextual representations from single sequences without evolutionary information. Explainable artificial intelligence analyses show that the model organizes RNA sequences in latent space and encodes structural properties indicating that biologically meaningful constraints are learned from sequence correlations alone. When fine-tuned for downstream structural and functional tasks, NucleicBERT requires only single sequences while matching or exceeding current RNA prediction models. This alignment-free framework addresses the scarcity of annotated 3D RNA data while providing a rapid, computational complement to experimental techniques. By bridging abundant unlabelled sequence data with scarce structural annotations, NucleicBERT advances RNA structure prediction and informs how large language models encode biological information. RNA structure and function are hard to infer because annotations are scarce, despite abundant sequence data. Upadhyay et al. trained a self-supervised model on large-scale RNA data that derives biologically meaningful patterns from sequence correlations.

Utkarsh Upadhyay, Julian Herold, Markus Götz et al. · 0 citations
#small language model Open access Sep 2026

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5′-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5′-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5′-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism’s own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

A. Duggan, M. Newman, David R. McMillen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.