Skip to content
Open access

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease

Aug 2026 · Frontiers in Genetics · Vol 17 · 0 citations · 44 references
Medicine

TL;DR

This study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort, and support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

Abstract

Introduction Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. Methods In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. Results The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%–53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. Discussion Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

Read PDF

Similar papers

Open access Jan 2026

From Germline Variants to Tumor Outcome: GWAS‐Based Functional Genomics Prioritizes Colorectal Cancer Susceptibility Genes and Links SMAD9 to Prognosis

Integrating statistical genetics, regulatory prediction, and locus‐directed experiments prioritized SMAD9, MAP3K2, FADS1, and ACTR1B as CRC susceptibility genes, and placed experimental bounds on the proposed MAP3K2 and ACTR1B mechanisms.

Chengguang Hu, Guang Yang, Han Xiong et al. · 0 citations
Review Open access Aug 2026

Suggestive genome-wide associations with inflammatory biomarkers in an admixed population, including a missense variant in the OR6K6 olfactory receptor gene associated with MCP-1

New SNPs linked to inflammatory biomarkers in a highly admixed Brazilian population are identified, including a missense variant in an olfactory receptor gene linked to MCP-1, which may be biologically important for inflammation and could affect the risk of cardiometabolic diseases.

J. Leite, G. F. L. Pascoal, G. B. S. Duarte et al. · 0 citations
Open access Sep 2026

Integrative Transcriptomic and Genetic Analysis Nominally Prioritizes SLC51A and TPMT as Candidate Diagnostic Genes for Thoracic Aortic Aneurysm

This study identifies SLC51A and TPMT as candidate diagnostic genes for TAA and proposes a simple two‑gene nomogram, which indicates that variation in SLC51A and TPMT expression is associated with coordinated changes in RNA metabolism, macromolecular catabolism, energy utilization, and cell‑cycle–related processes.

Jun-Chao Huang, Jin-Shan Zhou, Ya-Kun Liu et al. · 0 citations
Open access Sep 2026

Enhanced power and transferability for genetics-driven metabolomic biomarker discovery in admixed American cohorts

Despite metabolomics transforming our understanding of risk factors and aetiology of metabolic diseases, profiling is rarely performed for people of non-European ancestries, on whom much of metabolic disease burden falls. Metabolome-wide association studies (MWAS) can be performed using genetic scores to predict metabo...

T. Oreskovic, D. Jin, E. Trichia et al. · 0 citations
Open access Aug 2026

Locus-specific stratification and prioritization unveil genetic risk mechanism underlying complex diseases

An approach comprising locus-specific stratification (LSS) and gene regulatory prioritisation score (GRPS), which uniquely considers multi-signals during fine-mapping and target gene identification, to address issues arising from multi-signals in complex diseases.

Jing Zhang, Qiao-Qiao Liu, Ye Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.