Aug 2026· Journal of Chemical Information and Modeling· Vol 66 16, pp.
10396-10411
· 0 citations· 46 references
Medicine
TL;DR
Deep3MVPF, a multiview deep learning framework for 3'UTR stability prediction and m6A site identification, integrates a multiscale convolutional neural network, a k-mer de Bruijn graph neural network, and a secondary-structure graph neural network to jointly model sequence, topological, and structural representations.
Abstract
Accurate prediction of mRNA stability and identification of N6-methyladenosine (m6A) sites are central to understanding post-transcriptional regulation. Because the 3' untranslated region (3'UTR) contains both stability-associated cis-elements and many m6A sites, it provides a suitable context for modeling RNA regulatory effects. However, most existing methods rely primarily on linear sequence information and do not adequately capture higher-order topology or RNA structural context. Here, we present Deep3MVPF, a multiview deep learning framework for 3'UTR stability prediction and m6A site identification. Deep3MVPF integrates a multiscale convolutional neural network, a k-mer de Bruijn graph neural network, and a secondary-structure graph neural network to jointly model sequence, topological, and structural representations. For 3'UTR stability prediction, the model was trained and evaluated on a zebrafish (Danio rerio) mRNA degradation data set and achieved an MSE of 0.0049. For m6A site identification, it was evaluated on nine human cell line data sets and achieved an average AUC of 0.970. Attribution analysis further showed that Deep3MVPF recovered regulatory features consistent with known biology, including the destabilizing GCACUU motif and stabilizing G-rich/G-quadruplex-associated signals. These results demonstrate that integrating heterogeneous RNA representations can improve predictive modeling and facilitate interpretation of post-transcriptional regulatory grammar.
Identifying transcription factor binding sites (TFBSs) is fundamental to understanding complex gene regulatory mechanisms and the functions of non-coding regions. Although existing methods have achieved substantial strides, capturing both local structural features and long-range spatial dependencies within DNA sequences remains a major challenge for improving prediction accuracy. In this study, we propose DNCLA, a deep learning model that synergizes multisize convolutional fusion, Bidirectional Long ShortTerm Memory (Bi-LSTM) networks, and a multi-head self-attention
mechanism. At the feature extraction level, DNCLA breaks through the limitations of traditional single-sequence encoding by fusing Nucleotide Chemical Properties (NCP) with Dinucleotide Physicochemical Properties (DPCP). NCP provides a refined characterization of chemical differences between bases based on ring structures, hydrogen bond sites, and functional group properties, while DPCP introduces parameters such as local structural stability and geometric flexibility of the DNA. Subsequently, the model extracts spatial evolution from these high-dimensional features through a multi-size convolutional module; captures long-range spatial dependencies using Bi-LSTM layers; and employs a multi-head self-attention mechanism to achieve adaptive weight distribution of global features, thereby enhancing the perception of key regulatory motifs. Results from training and testing the proposed model on 165 ChIPseq datasets demonstrate that DNCLA possesses robust generalization capabilities and high predictive performance in TFBSs identification. This suggests that the incorporation of physicochemical features better elucidates the essence of interactions between transcription factors and DNA.
Jingjue Wei, Jie Feng· Match-communications in Math...· 0 citations
It is demonstrated that sequence and structure provide complementary predictive information, and that sites with greater prediction sensitivity to structural perturbation exhibit distinct local structural profiles between cell lines, and provides a multifeature deep learning framework for accurate and interpretable structure-aware epitranscriptomic prediction.
Ming-Ze Sun, Di Zhang, Zhi-Yuan Li et al.· PLoS Computational Biology· 0 citations
PLM-ArgMe is presented that is based on a symmetry-sensitive Transformer framework using context-aware ESM-2 residue embeddings, which is mapped through a novel Bio-Symmetric Mirrored Sinusoidal Encoding strategy to address the biological symmetry hypothesis of arginine methylation.
Nitika Bhatt, Kartik Joshi, R. Rout et al.· Biochemical and Biophysical...· 0 citations
Abstract RNA secondary structure is essential for understanding the functions of non-coding RNAs, ribosomal RNAs, and viral genomes. However, accurate prediction of long RNA structures remains challenging due to complex long-range interactions and the limited availability of long-RNA training data. We present UFold-X, a dual-branch deep learning framework that combines a convolutional encoder for local structure modeling with a Mamba-based Visual State Space Module for capturing long-range dependencies. A dynamic gating mechanism adaptively integrates the two branches according to sequence length. UFold-X was evaluated on multiple benchmark datasets containing RNAs up to 5000 nucleotides. To rigorously assess generalization, we introduced a cross-clan benchmark for long RNAs. Under this stringent setting, UFold-X achieved performance comparable to state-of-the-art classical approaches while achieving the best performance among deep learning-based methods. Additional cross-family and within-family evaluations further demonstrated robust transferability and competitive predictive performance. UFold-X also maintained excellent computational efficiency, requiring only 0.08 s per sequence on average. To assess biological consistency, we developed a SHAPE-based reactivity prediction variant (UFold-X-R) and an integrated metric, the Hybrid Reactivity-Pairing Score (HRPS). UFold-X-R showed strong agreement with experimental icSHAPE data and achieved the highest HRPS among all evaluated methods. A user-friendly web server is available at https://ufold-x.ai4bread.com.
Lai-Yi Fu, Jia-Chun Li, Rui-Qi Wang et al.· Nucleic Acids Research· 0 citations
Protein-RNA interactions regulate diverse biological processes and are increasingly exploited in therapeutic RNA discovery, but accurate inferences of nucleotide preferences and reliable structure prediction remain challenging. Here, we present PRIS, a unified structure-based deep-learning framework that combines two complementary components: PRISeq for nucleotide probability estimation at each RNA position and PRIScore for residue-nucleotide distance prediction to discriminate native-like from incorrect poses. Both share a feature extractor that integrates an Anti-Symmetric Graph Attention Network (A-GAT) with sparse k-Maximum Inner Product (k-MPI) attention to capture long-range interactions across large graphs. PRIScore improves the selection of native-like protein-RNA predictions generated by AlphaFold3, achieving a top-1 success rate of 81.91% on a docking benchmark, compared to 79.26% for AlphaFold3. The selected structures are then fed into PRISeq, which infers position-specific binding preferences and screens RNA libraries. On a PWM benchmark, PRISeq achieved a mean absolute error (MAE) of 0.75, outperforming FoldX, Rosetta-based scoring functions, and NA-MPNN. In virtual screening against MS2 protein, PRISeq screens 129,248 RNA hairpins within 11.95 seconds, achieving the highest EF0.5% of 14.40, approximately double the best baseline. PRIS also effectively enriches active aptamers against NELF-E and GFP while preserving sequence diversity. By integrating structure selection with binding-preference inference, PRIS provides an efficient framework for large-scale RNA library screening and aptamer design.
Yi-Hao Zhao, Jing Han, Ji-Ke Wang et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.