ContextAware Significance of Cross-gene Attribution for Discovering Elements (CASCADE) is introduced, a positionspecific statistical framework for identifying model-derived candidate regulatory elements from in silico saturation mutagenesis and shifts motif recovery from downstream of the transcription start site toward promoter sequence.
Abstract
Gene expression is governed by regulatory DNA and their associated trans factors acting in specific cell types, yet the sequences underlying this control remain poorly mapped in plants. Genome-pretrained DNA language models provide a route to interrogate regulatory sequence directly, but their attributions have largely been interpreted using bulk or whole-tissue data, and standard attribution pipelines can preferentially highlight sequences downstream of the transcription start (TSS) site rather than promoter-associated signals. Here, we train a celltype-resolved sequence-to-expression model from a single-cell soybean (Glycine max) atlas by coupling a soybean-adapted Genomic Pre-trained Network (GPN) to a shared sequence encoder with 66 cell-type-specific output heads. Across 38,339 protein-coding genes, the model achieves a mean per-cell-type, across-gene Pearson correlation of 0.683 and, recast as a highversus-low expression classification, reaches an area under the ROC curve of 0.92 to 0.97 across tissues, at or above dedicated plant sequence models. We then introduce ContextAware Significance of Cross-gene Attribution for Discovering Elements (CASCADE), a positionspecific statistical framework for identifying model-derived candidate regulatory elements from in silico saturation mutagenesis. Relative to the pooled null used by TF-MoDISco, CASCADE shifts motif recovery from downstream of the transcription start site toward promoter sequence, with 77% of CASCADE-exclusive motifs, compared with 12% of TF-MoDISco-exclusive motifs, falling within the promoter. Applied across the atlas, CASCADE identifies approximately 1.39 million candidate elements spanning broadly active, tissue-restricted and cell-type-restricted classes. Together, these analyses establish a position-aware approach for extracting promoterassociated regulatory hypotheses from sequence models and generate a cell-type-resolved map of candidate cis-regulatory elements.
This review compares convolutional, Transformer-based and graph architectures used to represent local sequence features, chromatin state and three-dimensional genome organisation to their applications to transcription-factor binding, chromatin accessibility, gene expression, non-coding variant prioritisation and regulatory-sequence design.
Results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints.
Single-cell foundation models have transformed transcriptomic analysis, yet most rely on fixed gene identifiers that limit transfer across species and data types. Here we present LucaCell, a sequence-centric foundation model that represents genes through pre-trained mRNA sequence embeddings rather than static gene annotations. Gene expression is discretized into bins and modeled with a Transformer encoder, enabling sequence-informed cell representation without a fixed gene-ID vocabulary. Pre-training on 85 million human and mouse single cells, LucaCell is evaluated on human, mouse and lemur gene expression profiles, human chromatin accessibility data, unaligned reads from more than 50 prokaryotic taxa, and five influenza A virus genomes. LucaCell enables manual-mapping-free cross-species cell type annotation and an alignment-free microbial embedding framework that simultaneously distinguishes bacterial species identity and intra-species physiological states. It also improves gene expression reconstruction by incorporating donor-specific exonic SNP information into mRNA sequence embeddings, and predicts cellular viral load across influenza A virus strains while highlighting infection-like transcriptional states in mock-infected cells. These results show that sequence-informed gene representation can improve the generalization of single-cell foundation models across species, data types, and predictive tasks.
Yan Sun, Yong He, Min-Si Ren et al.· bioRxiv· 0 citations
Abstract Motivation Predicting and deciphering the regulatory logic of enhancers remains a significant challenge due to their complex sequence features and the absence of consistent genetic or epigenetic signatures that distinguish them from other genomic regions. Existing machine learning methods capture nucleotide composition but often fail to model sequence context effectively. Results We present DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. Using ENCODE registry of candidate cis-regulatory elements (cCREs), we curated a benchmark dataset, consisting of 21 926 enhancers of 201 bp length and 46 159 enhancers of 350 bp length, as positive instances. The best fine-tuned model achieved 88.05% accuracy and a Matthews correlation coefficient of 76% on an independent dataset. Genome-wide application identified 1 684 595 enhancer regions covering 26.65% of the human genome. By performing integrative analyses with DNABERT-based transcription factor models, we identify 2681 statistically significant loss-of-function and 1917 gain-of-function enhancer variants, which respectively alter the function of 1623 and 1247 ENCODE-cCRE enhancers. Similarly, we identify 4057 candidate de novo enhancers, created by 5464 gain-of-function variants. These genome-wide enhancer annotations and candidate genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies. Availability and implementation DNABERT-Enhancer is freely available at https://github.com/DavuluriLab/DNABERT-Enhancer; Trained model predictions can be explored interactively via the web application at https://dnabert-enhancer-datarepo.streamlit.app/. The fine-tuned models are archived and citable through Zenodo (https://doi.org/10.5281/zenodo.19157566).
Rekha Sathian, P. Dutta, Ferhat Ay et al.· Bioinformatics· 0 citations
By integrating genome-wide association studies loci from Alzheimer's disease, multiple sclerosis, and schizophrenia, scReGAT identifies disease-associated cell types and uncovers candidate regulatory mechanisms underlying complex trait associations, positioning scReGAT as a robust and generalizable framework for decoding long-range gene regulation at single-cell resolution.
Abstract Motivation Genomic sequence-to-activity models can decipher gene regulatory mechanisms and predict the functional impact of regulatory variants. However, current models struggle to integrate information from sequences outside promoters, especially information from cell type specific regulatory elements. Results Here, we propose incorporating base-pair resolution evolutionary conservation data into genomic sequence-to-expression predictors. We explore two training strategies—training from scratch or fine-tuning an existing sequence-only model with additional conservation input. We find that in both cases, base-pair resolution conservation data improves cell type specific sequence-to-expression prediction, with training from scratch yielding the greatest benefit. The improvement in cell type specific expression prediction can be attributed in part to the fact that models trained on sequence and conservation data learn to better recognize cell type specific regulatory elements than models trained on sequence alone. Availability Code is available at https://github.com/ni-lab/basenji-phyloP.
Pooja Kathail, Forest L. Yang, Gabriel B. Loeb et al.· Bioinformatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.