Skip to content
Open access

A token-pruning framework enables efficient representation of the human genome for RNA modification analysis

Sep 2026 · Bioinformatics · Vol 42 · 0 citations · 51 references
Medicine

TL;DR

Experimental results show that ATSFormer consistently outperforms existing state-of-the-art methods while achieving substantial computational savings and structural analysis using AlphaFold3 supports the biological relevance of the motifs identified by ATSFormer.

Abstract

Abstract Motivation Modelling long genomic sequences remains challenging due to extreme sequence length, high redundancy, and the need for biological interpretability. Although Transformer-based architectures have achieved strong performance across genomic tasks, their high computational cost and reliance on fixed tokenization strategies limit their scalability and ability to focus on biologically informative regions. Results We propose ATSFormer, a token-pruning Transformer framework for efficient and biologically informed genomic sequence modelling. ATSFormer incorporates an attention-guided and parameter-free Adaptive Token Sampling (ATS) module into Transformer layers. Guided by attention-derived importance scores, ATS dynamically retains informative tokens while probabilistically discarding redundant ones, thereby reducing sequence length, FLOPs, and memory usage without introducing additional learnable parameters or extra training procedures. Importantly, the retained tokens correspond to key contributors to model predictions, enabling ATSFormer to highlight biologically meaningful sites and sequence motifs. We evaluated ATSFormer on four benchmark RNA modification datasets derived from RMVar 2.0, covering A-to-I, m1A, m5C, and m7G. Experimental results show that ATSFormer consistently outperforms existing state-of-the-art methods while achieving substantial computational savings. Furthermore, structural analysis using AlphaFold3 supports the biological relevance of the motifs identified by ATSFormer. Availability and implementation The source data and code are freely available at GitHub (https://github.com/1gao2/ATSFormer) and Zenodo (https://doi.org/10.5281/zenodo.21813541).

Read PDF

Similar papers

Preprint Aug 2026

Frozen but Not Always Accessible: A Representation Analysis of Genomic Language Models

A representation-accessibility analysis of frozen genomic language models across regulatory, epigenetic, promoter, splice-site, and variant-effect prediction tasks shows that local biological signal is partially present in frozen representations, but is not always accessible through final pooled embeddings.

Nirjhor Datta, Swakkhar Shatabda, M. Rahman · 0 citations
Book Open access Aug 2026

MotRNA: Encoding RNA Motifs via Explicit N-gram Memory

RNA motifs and Low-Complexity Repeats (LCRs)—recurrent, conserved sequences of nucleotides—serve as the fundamental vocabulary of RNA structure and biological function. While current RNA foundation models have revolutionized sequence modeling through Transformer architectures, they predominantly prioritize modeling glo...

Xiang-Yu Ji, Xin Wang, Yang Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transfo...

Ming-Rui Li, Si-Xian Shen, Min-Zhang Li et al. · 0 citations
Open access Jul 2026

Accelerating inference in genomic and proteomic foundation models via speculative decoding

Abstract Motivation Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Trans...

K. Provatas, Aris Karatzikos, Charalampos Koilakos et al. · 0 citations
Open access Sep 2026

LAMBDA: a prophage detection benchmark for genomic language models.

Transformer-based genomic sequence models represent an emerging frontier in computational biology. Yet, their embeddings have not yet shown the same level of predictive power as natural and protein language models, highlighting a gap between current implementations and theoretical promise. Existing benchmarks for DNA l...

LeAnn M. Lindsey, Nicole L. Pershing, K. Dufault-Thompson et al. · 0 citations
Open access Sep 2026

Predicting genome-wide functional constraints with GPN-Star.

Genomic language models have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences1. However, standard genomic language models adapted from natural language processing often require large model sizes and computational resources, yet still fall short of classical evol...

Cheng-Zhong Ye, Gonzalo Benegas, Carlos Albors et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.