Skip to content
Open access

WGS2IBI: a cloud-based workflow for individualized Bayesian inference from whole genome sequencing data

Jul 2026 · BMC Bioinformatics · 0 citations

TL;DR

WGS2IBI provides a scalable, reproducible, and accessible workflow resource for WGS analysis, enabling efficient population- and individual-level genetic studies without local installation and lowering practical computational barriers for large-scale WGS studies.

Abstract

Scalable and reproducible genomic workflows that support both individual- and population-level analyses are critically needed in precision medicine. We present WGS2IBI, a cloud-based, modular workflow that integrates whole-genome sequencing (WGS) preprocessing, population-level variant screening, and the previously developed Individualized Bayesian Inference (IBI) framework within a reproducible analysis pipeline. Implemented using the Common Workflow Language (CWL) and Docker, WGS2IBI ensures portability and reproducibility across computational environments. Benchmarking on the Jackson Heart Study (JHS) TOPMed Freeze 9 cohort demonstrated efficient large-scale genomic processing and analysis. Preprocessing reduced 102 million variants to 18 million variants in approximately two hours at a total cost of $20.83. Population-level analyses using Global Search and Fisher’s exact test (FET) were completed for $0.44 and $11.44, respectively. Individual-level analysis using IBIwas completed for $3.66. Analyses across additional TOPMed cohorts showed predictable scaling across cohort and chromosome sizes. Deployment on both BioData Catalyst (BDC) and the Gabriella Miller Kids First (KF) Data Resource Center via CAVATICA further demonstrated portability across two major NIH cloud ecosystems. To provide a biologically meaningful use case beyond workflow benchmarking, we also applied WGS2IBI to real hypertension data from 1,821 unrelated Framingham Heart Study (FHS) participants, where IBI-prioritized variants were enriched for lower minor allele frequency and included variants mapping to genes with prior blood-pressure relevance. WGS2IBI provides a scalable, reproducible, and accessible workflow resource for WGS analysis, enabling efficient population- and individual-level genetic studies without local installation. Dual deployment on BDC and KF expands usability across diverse NIH genomic ecosystems, supporting both adult and pediatric research communities. This workflow lowers practical computational barriers for large-scale WGS studies and enables integrated evaluation of individual- and population-level genetic analyses.

Read PDF

Similar papers

Review Open access Aug 2026

Variant calling in nonmodel organisms with snpArcher

A step-by-step protocol for snpArcher, a Snakemake-based workflow that takes raw sequencing reads and a reference genome as input and produces a filtered, joint-called VCF suitable for downstream population genomic analysis, is presented.

Cade Mirchandani, Abdelmajid Omarjee, Guillaume Achaz et al. · 0 citations
Open access Aug 2026

A Deterministic Framework for Integrated Genome Variant Interpretation - The ‘GenomeVAP’

High-throughput genomic sequencing generates vast amounts of data, yet the interpretation of individual-genetic variants remains hindered by the dispersion of relevant evidence across various databases. We present a modular, web-based framework (GenomeVAP) designed for deterministic evidence integration in genomic rese...

Anusha Sunder · 0 citations
Open access Sep 2026

Automated GenePy gene-burden computation via a reproducible Nextflow workflow integrated with the Genomics England (GEL) Lifebit platform

Abstract Interpretation of rare-disease genomes remains constrained by variant-centric analytical frameworks that insufficiently capture the cumulative impact of multiple variants within a gene. GenePy provides an individual-level, gene-based burden metric that integrates variant consequence, allele frequency, and zygo...

Iman Nazari, Guo Cheng, J. Ashton et al. · 0 citations
Review Open access Sep 2026

Geomosaic: a flexible bioinformatics platform integrating complementary metagenomic analyses from sequencing reads to genomes

Metagenomic analyses can be performed at multiple analytical levels, including read-based, assembly-based, and genome-resolved approaches, each capturing complementary biological information while introducing distinct analytical biases and trade-offs. However, existing workflows are commonly optimized for a single anal...

D. Corso, Edoardo Taccaliti, B. Barosa et al. · 0 citations
Preprint Aug 2026

SVPLEX: A Nextflow Pipeline for Cohort-level Structural Variant Calling

SVPLEX is a Nextflow pipeline for cohort-level structural variant detection from short-read whole-genome sequencing data and generates a merged consensus callset across the analysis cohort, which can be used to assess cohort-specific variation, remove technical artefacts, and serve as input for rare disease variant pri...

J. Munro, M. F. Bennett, M. Bahlo · 0 citations
Open access Sep 2026

Upscaling Genotyping by Amplicon Sequencing With GBAS‐GUI

ABSTRACT Genotyping by amplicon sequencing (GBAS) is a relatively low‐cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large‐scale genetic monitoring projects. However, most existing analytical pipelines are either marker‐spec...

Sebastian Sonnenberg, Thapasya Vijayan, C. Rupprecht et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.