This review describes the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS and presents 30 methods designed to leverage AI in GWAS.
Abstract
The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.
ABSTRACT Genome‐wide association studies (GWAS) have shown that pleiotropy, whereby a single genetic variant or gene influences multiple traits, is common in complex human diseases. Detecting cross‐phenotype associations from GWAS summary statistics remains challenging because of small effect sizes, extensive multiple testing, heterogeneous effects, and possible differences in effect direction across traits. Methods that jointly analyze multiple traits can improve the ability to detect pleiotropic signals while retaining the practical advantages of summary statistic‐based analyses. Although a range of statistical approaches has been developed for this purpose, practical guidance on their application, assumptions, and interpretation remains limited. This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets. We also highlight the importance of accounting for effect heterogeneity, correlation, and biological group structure at the gene and pathway levels in the detection and interpretation of pleiotropic association signals.
Christina Y. Feng, P. Sugier, Nan Zou et al.· Statistics in Medicine· 0 citations
Genomic selection (GS) has revolutionized animal breeding by accelerating genetic gain through genome-wide marker data. As genotyping technologies advance and data dimensionality grows, the statistical foundations of GS are shifting from classical linear frameworks, which assume additive genetic effects, toward advanced computational models that capture complex nonlinear relationships in genomic data. The commercialization of genotyping arrays for livestock and poultry, coupled with steadily declining sequencing costs, has led to an exponential increase in the availability of high-density genomic data. However, challenges persist, including scenarios where the number of genetic markers far exceeds the number of samples with phenotypic data, and the growing complexity of relationships within genomic data. These issues significantly limit the applicability of traditional evaluation models. In parallel, computational power has increased significantly over the last few decades, providing the capacity necessary for highly complex analyses. While traditional linear mixed models provide a robust framework for incorporating biological priors and modeling additive genetic effects, they often rely on simplified assumptions. In contrast, machine learning (ML) and deep learning (DL) algorithms, which do not rely on predefined parametric models, are well-suited to capturing complex nonlinear relationships and offer effective solutions to the aforementioned challenges. This review provides a comprehensive overview of GS methodologies. We first cover the statistical foundations of linear mixed and Bayesian models, and then survey modern ML and DL approaches. We discuss the assumptions, advantages, and limitations of each method and, by comparing the computational efficiency and predictive accuracy of these diverse approaches, aim to provide practical guidance for optimizing genomic evaluation strategies in the era of big data breeding.
Lifei Zhang, Mingzhu Zhang, Xinle Wang et al.· Journal of Animal Science an...· 0 citations
Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work presents an empirical analysis of UQ in deep learning models, focusing on genomics applications. In a series of experiments, we contrast Deep Ensembles, Bayesian Neural Networks, and Monte Carlo-dropout methods. We assess their ability to quantify uncertainty in different scenarios, accounting for common dataset characteristics in two genomic application areas and modalities: sequence-to-activity models, and single-cell expression analysis. Our systematic comparison framework provides guidelines for the applicability and reliability of UQ methods in genomics, highlighting their strengths and limitations in different scenarios. We show that Bayesian Neural Networks are better at capturing uncertainty caused by strong class imbalance and out-of-distribution data in genomics, despite their computational disadvantages. Moreover, we show how uncertainty scores can be used to select high-quality predictions in protein-RNA interactions.
A clinically oriented, pipeline-based synthesis of contemporary AI applications in genomic medicine, focusing on factors that determine model robustness and clinical utility, and common sources of failure in real-world genomic AI systems.
Alexandra-Maria Blaga, Răzvan-Octavian Mihuț, A. Treteanu et al.· International Journal of Mol...· 0 citations