Recent progress in predicting RNA modification site-disease associations is summarized, with a focus on common modifications such as m6A, m1A, and m7G, and key challenges in this field are highlighted.
Abstract
Abstract With the development of epitranscriptomics, studies have shown that RNA modification sites are closely related to many diseases. Because experimental validation is time-consuming and labor-intensive, an increasing number of studies have developed computational methods to predict potential associations between RNA modification sites and diseases. In this review, we summarize recent progress in predicting RNA modification site-disease associations, with a focus on common modifications such as m6A, m1A, and m7G. First, we systematically summarize commonly used databases and data sources and outline approaches for constructing similarity information for modification sites and diseases. We then review existing prediction methods—including network-based strategies, matrix completion, and machine learning—and discuss their typical advantages and limitations. Finally, we highlight key challenges in this field, including limited known associations, data imbalance, unclear definitions of negative samples, inconsistent evaluation standards across studies, and limited interpretability and experimental validation. We also suggest future directions, such as expanding high-quality datasets, integrating multi-omics data, establishing unified evaluation pipelines, and strengthening experimental validation. We hope this review provides a clear overview and practical guidance for studies on the associations between modification sites and diseases and supports the development of more reliable prediction methods.
Understanding how epigenome variation contributes to gene expression in disease and development is a fundamental challenge. Regulatory regions show cell type-specific epigenome activity and differ in their location, size, and distance to their target genes, complicating discovery and analysis. Recent machine learning models have been proposed to address these problems by learning functions for the prediction of gene expression from epigenomic data. Here, we use the large IHEC EpiATLAS dataset to benchmark state-of-the-art linear and nonlinear approaches. We optimize each approach for over 28,000 human genes, providing an inferred regulatory catalog of gene models. In-depth comparison reveals that gene characteristics and the epigenomic complexity of the locus influence the difficulty of predicting the epigenome-to-transcriptome association. The model performance is further evaluated using CRISPRi and eQTL validation data. Based on these models, we conduct histone-acetylation association studies in a systematic way to investigate how epigenetic variation impacts gene expression. The model-based analysis revealed genes and regulatory regions linked to B-cell leukemia in patient data with known disease-related functions. Our work provides a foundation for applications that link epigenome variation to gene expression in human cells, by benchmarking methods on a per-gene basis, illustrating their use in a disease context and making trained models available to the community.
Fatemeh Behjati Ardakani, Shamim Ashrafiyan, Laura Rumpf et al.· Genome Biology· 0 citations
Post-translational modifications (PTMs) are pivotal in modulating protein function and cellular processes. However, experimental identification of PTM sites remains costly and labor-intensive. Recent advances in artificial intelligence (AI) have enabled accurate and scalable in silico PTM site prediction from large-scale proteomic data. In this review, we provide a comprehensive and up-to-date overview of AI-driven PTM site prediction across more than ten PTM classes, covering single-PTM site prediction, multiple-PTM site prediction, inter-site crosstalk prediction, and functional prediction of modification sites. We systematically analyze and compare key AI frameworks, from conventional machine learning to deep learning, and summarize representative tools. We also identify key challenges and propose future directions for improvement. To facilitate application and ongoing progress, we provide practical guidelines for method selection and have established a dedicated website, which serves as a community benchmarking resource for the development of PTM site prediction tools. This website will be regularly updated with emerging prediction tools. By integrating comprehensive literature analysis with a dynamic online resource, we aim to provide a reliable foundation for understanding current capabilities and guiding the future development of PTM site prediction tools, thereby promoting the integration of AI into practical biomedical research applications.
Jia-Yi Ran, Xiaohan Zhang, Yunze Wang et al.· Genomics, Proteomics & Bioin...· 0 citations
Post-translational modifications (PTMs) are modifications of proteins that occur after translation. They exert their influence on human health by altering the properties of proteins. Recent progress in biomedical research has yielded extensive heterogeneous datasets detailing protein PTMs and their disease relevance. These datasets provide a substantial foundation for the development of methods predicting PTM-disease associations. Thus we develop BAMPDA, a matrix refactoring framework with heterogeneous inference designed to enhance the inference of PTM-disease associations. Our approach integrates multi source data encompassing both protein and disease characteristics, leveraging two successively applied matrix based algorithms for association prediction. Furthermore, rigorous 5-fold cross-validation demonstrates that BAM PDA significantly outperforms five recent baseline methods. Finally, we applied BAMPDA to four representative diseases to predict associated proteins and validate the role of PTMs in mediating these disease relationships. The results underscore BAMPDA's robust predictive capability in practical scenarios.
Xinting Zhang, Jie Ni, Zihao Song et al.· IEEE transactions on computa...· 0 citations