It is demonstrated that nucleotide position within the input sequence alters the nature of SegmentNT’s raw prediction probabilities, which can be standardized to improve prediction consistency and identify potential approaches to account for these biases.
Abstract
Abstract Recent advances in large language models have extended to genomic applications, yet model robustness relative to context is unclear. Here, we demonstrate two intrinsic biases (input sequence length and nucleotide position) affecting SegmentNT results, a model included with the Nucleotide Transformer that provides nucleotide-level predictions of biological features. We demonstrate that nucleotide position within the input sequence (beginning, middle, or end) alters the nature of SegmentNT’s raw prediction probabilities, which can be standardized to improve prediction consistency. While longer input sequence length improves model performance, diminishing returns suggest a surprisingly small input length of ∼3072 nucleotides might be sufficient for many applications. We further identify a 24-nucleotide periodic oscillation in SegmentNT’s prediction probabilities, revealing an intrinsic bias potentially linked to the model’s training tokenization (6-mers) and architecture. We identify potential approaches to account for these biases and provide generalizable insights for utilizing nucleotide-resolution functional prediction models.
Experimental results show that ATSFormer consistently outperforms existing state-of-the-art methods while achieving substantial computational savings and structural analysis using AlphaFold3 supports the biological relevance of the motifs identified by ATSFormer.
Wen-Jia Gao, Jun-Lei Yu, Jun-Ru Jin et al.· Bioinformatics· 0 citations
A simple generative approach for creating synthetic 5′-UTR libraries based solely on the genomic sequence statistics of any desired organism and demonstrates that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mec...
A. Duggan, M. Newman, David R. McMillen· PLoS ONE· 0 citations
RNA motifs and Low-Complexity Repeats (LCRs)—recurrent, conserved sequences of nucleotides—serve as the fundamental vocabulary of RNA structure and biological function. While current RNA foundation models have revolutionized sequence modeling through Transformer architectures, they predominantly prioritize modeling glo...
Xiang-Yu Ji, Xin Wang, Yang Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
SPIRAL is presented, a layer-wise SAE analysis of BiRNA-BERT, a layer-wise SAE analysis of RNA language models where byte-pair tokenization breaks the one-token-one-nucleotide correspondence that nucleotide-level attribution assumes.
M. Hossain, MD. Roqunuzzaman Sojib, Md Toki Tahmid et al.· bioRxiv· 0 citations
Protein language models (pLMs) such as ESM-2 achieve strong zero-shot mutation-effect prediction, yet the internal computations supporting these predictions remain poorly understood. We introduce a sparse feature circuit framework that combines sparse autoencoders, integrated-gradients attribution, and activation patch...
Saishradha Mohanty, Manya Phutela, A. G. Green· bioRxiv· 0 citations
Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transfo...
Ming-Rui Li, Si-Xian Shen, Min-Zhang Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.