Skip to content
Review Open access

SNPannotator: automated functional annotation of genetic variants and linked proxies

Aug 2026 · Bioinformatics · Vol 42 · 1 citation · 31 references
Medicine

TL;DR

The SNPannotator package is introduced, an automated post-GWAS analysis software package designed to streamline the interpretation of GWAS findings and provide a practical framework for efficiently deriving biologically meaningful insights from GWAS data.

Abstract

Abstract Summary Genome-wide association studies (GWASs) have identified thousands of genetic variants associated with complex traits and diseases. However, explaining the mechanisms underlying phenotypic variation remains challenging. Here, we introduce SNPannotator, an automated post-GWAS analysis software package designed to streamline the interpretation of GWAS findings. Our pipeline implements a multi-step process that identifies proxy variants in high linkage disequilibrium (LD) with associated lead variants, then queries comprehensive resources (including Ensembl, the GTEx Portal, the eQTL Catalog, and STRING DB) for genomic position, deleteriousness, regulatory annotations, clinical significance, trait associations, expression (eQTLs) and splicing quantitative trait loci (sQTLs), and functional enrichment analyses and compiles the results into user-friendly reports. This package is implemented in the R programming language and includes auxiliary functions for variant lookup and LD exploration. SNPannotator provides a practical framework for efficiently deriving biologically meaningful insights from GWAS data and for assisting researchers in prioritizing candidate variants for functional validation. Availability and implementation The SNPannotator package is available from the Comprehensive R Archive Network (CRAN) at https://cran.r-project.org/web/packages/SNPannotator. The development version and tutorial is available on GitHub (https://github.com/omicslaboratory/SNPannotator). The online version of the package is available at https://omicslab.org/snpannotator.

Read PDF

Similar papers

Open access Sep 2026

Modeling pathway overlap increases accuracy of GWAS gene set enrichment

Genome-wide association studies (GWAS) have identified thousands of loci associated with complex traits and diseases, and extensive efforts are underway to translate these variant-level signals into biological mechanism. A widely applied approach is pathway enrichment analysis, which tests whether genetic associations concentrate within biological pathways beyond background polygenic expectations. However, pathway databases contain extensive sharing of genes across pathways (pathway overlap), an underappreciated source of bias that creates structural dependencies in enrichment statistics and obscures pathway-specific genetic signal. Moreover, the degree of pathway overlap is increasing as pathway resources expand. Here, we introduce Gene Swap Randomization (GSR), an empirical framework that preserves pathway size and multi- pathway gene membership in the null model, enabling explicit adjustment for pathway overlap. Applying GSR to enrichment results from the Molecular Signatures Database (MSigDB) across twelve complex traits and four pathway analysis approaches (MAGMA, PascalX, GSA-MiXeR, and PRSet), we show that pathway overlap can produce enrichment under polygenicity even in the absence of pathway-specific biology. GSR improves prioritization of biologically relevant pathways supported by independent gene-disease associations (Open Targets, Malacards), regulatory interactions (DoRothEA), and tissue-specific expression patterns (GTEx). GSR improves concordance with external benchmarks in 60.8% of comparisons overall and 79.3% disease- association benchmarks, corresponding to improvement in 10 of 16 aggregated method-validation framework comparisons. We demonstrate that pathway overlap is a key source of bias in GWAS pathway enrichment, that pathway-specific disease enrichment persists after conditioning on overlap, and that GSR improves biological insight by distinguishing pathway-specific genetic signal from enrichment driven by pathway overlap.

A. Cote, W. R. Kesting, J. García-González et al. · 0 citations
Review Open access Aug 2026

A Guide for Exploring Pleiotropic Associations in Genome‐Wide Association Studies Using Summary Statistics

This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets.

Christina Y. Feng, P. Sugier, Nan Zou et al. · 0 citations
Open access Aug 2026

Robust Inference With Ghostknockoffs in Genome‐Wide Association Studies With Sample Relatedness

Genome‐wide association studies (GWASs) have been extensively adopted to depict the underlying genetic architecture of complex traits. Recent studies show that knockoff‐based methods can identify variants with unique, potentially causal effects on phenotypes. However, their statistical validity and effectiveness in studies with related individuals, such as the UK Biobank, remain unexplored. In this paper, we extensively evaluate a simple and effective analytical strategy that integrates GhostKnockoffs and state‐of‐the‐art marginal association tests. We show that this approach is robust to arbitrary relatedness structure as long as the input Z‐scores are derived from valid generalized linear mixed models. This robustness also extends GhostKnockoffs to other GWASs settings, including meta‐analysis of studies with sample overlap when the input score test Z‐scores are properly calibrated, and association test statistics beyond score tests in independent sample settings. We demonstrate the method's validity and practical utility using simulation studies and a meta‐analysis of nine European ancestral genome‐wide association studies and whole exome/genome sequencing studies for the Alzheimer's disease.

Xinran Qi, M. Belloy, Jiaqi Gu et al. · 0 citations
Open access Jan 2026

From Germline Variants to Tumor Outcome: GWAS‐Based Functional Genomics Prioritizes Colorectal Cancer Susceptibility Genes and Links SMAD9 to Prognosis

Introduction Genome‐wide association studies (GWAS) have identified over 200 germline risk loci for colorectal cancer (CRC), yet the causal variants and genes behind most GWAS signals remain unknown and the link between inherited risk and tumor outcome is largely unexplored. Connecting germline single nucleotide polymorphisms (SNPs) to gene expression through expression quantitative trait loci (eQTL) and to clinical outcome is needed to interpret this inherited risk. Methods We combined CRC GWAS summary data (73,149 cases and 112,467 controls of European ancestry) with GTEx v8 eQTL from colon sigmoid, colon transverse, small intestine terminal ileum, and whole blood, integrating causal transcriptome‐wide association study (cTWAS) with SuSiE fine‐mapping, Bayesian colocalization, MAGMA SNP‐to‐gene analysis, and AlphaGenome variant‐effect prediction. Prioritized proteins were assessed by western blot. Four experimentally selected regulatory variants were genotyped in HCT116, SW480, RKO, and NCM460 cells; genotype‐protein associations were tested across cell‐line means, and cis‐regulatory effects were assessed by allele‐specific expression (ASE) and reference‐versus‐alternate dual‐luciferase assays. Prognostic relevance was evaluated in TCGA colorectal tumors. Results cTWAS identified seven genes with posterior inclusion probability (PIP) > 0.50, and colocalization across 78 gene‐tissue pairs revealed 20 associations with PP.H4 > 0.80. Four genes showed convergent evidence: SMAD9 (PIP = 0.916, PP.H4 = 0.981), MAP3K2 (PIP = 0.762, PP.H4 = 0.827), FADS1 (PIP = 0.632, PP.H4 = 0.942), and ACTR1B (PIP = 0.566, PP.H4 = 0.994). All four proteins were reduced in CRC cells. Alt‐allele dosage was inversely associated with SMAD9 (Pearson r = −0.985, BH‐adjusted p = 0.029) and FADS1 protein abundance (r = −0.995, BH‐adjusted p = 0.020), but not with MAP3K2 or ACTR1B. Reporter and ASE assays detected the clearest allele‐specific effects at SMAD9 and FADS1, whereas MAP3K2 and ACTR1B were null in the tested systems. Lower SMAD9 expression was nominally associated with poorer survival (log − rank p = 0.022 to 0.050), but these associations did not survive Benjamini–Hochberg correction across 12 tests. Conclusions Integrating statistical genetics, regulatory prediction, and locus‐directed experiments prioritized SMAD9, MAP3K2, FADS1, and ACTR1B as CRC susceptibility genes. Functional evidence was strongest and most directionally coherent for SMAD9, demonstrated allele‐specific but context‐dependent regulation at FADS1, and placed experimental bounds on the proposed MAP3K2 and ACTR1B mechanisms.

Chengguang Hu, Guang Yang, Han Xiong et al. · 0 citations
Review Open access Jul 2026

SNP Detection Strategies in Genomic Research: A Comparative Review of Major Tools, Algorithms, Challenges and Applications

This review compares SNP detection programs such as GATK, BCFtools, FreeBayes, SAMtools, SAMtools, and DeepVariant and their algorithmic structures, namely pileup- based, haplotype-based, and machine-learning approaches and suggests that no single tool is the best.

Shikhi Baruri, S. Khanal · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.