Skip to content
Preprint

SVPLEX: A Nextflow Pipeline for Cohort-level Structural Variant Calling

Aug 2026 · 0 citations
Biology

TL;DR

SVPLEX is a Nextflow pipeline for cohort-level structural variant detection from short-read whole-genome sequencing data and generates a merged consensus callset across the analysis cohort, which can be used to assess cohort-specific variation, remove technical artefacts, and serve as input for rare disease variant prioritisation workflows.

Abstract

SVPLEX is a Nextflow pipeline for cohort-level structural variant detection from short-read whole-genome sequencing data. The pipeline implements six different structural variant callers with different strengths and weaknesses, integrating different levels of evidence for SVs, and generates a merged consensus callset across the analysis cohort. Callset filtering is achieved by leveraging consensus among multiple individual callers and by ensuring that deletion and duplication calls are supported by observable changes in read depth. The output merged cohort SV callset can then be used to assess cohort-specific variation, remove technical artefacts, and serve as input for rare disease variant prioritisation workflows. SVPLEX is user-friendly, reproducible, scalable, and can be executed flexibly on either a local workstation, a high-performance compute (HPC) cluster, or deployed on cloud infrastructure. The required inputs are alignment files for the cohort of interest, and the output is a single merged cohort structural variant VCF. SVPLEX is available on GitHub (bahlolab/SVPLEX) and is licensed under the MIT open-source licence.

View source

Similar papers

Review Open access Aug 2026

Variant calling in nonmodel organisms with snpArcher

A step-by-step protocol for snpArcher, a Snakemake-based workflow that takes raw sequencing reads and a reference genome as input and produces a filtered, joint-called VCF suitable for downstream population genomic analysis, is presented.

Cade Mirchandani, Abdelmajid Omarjee, Guillaume Achaz et al. · 0 citations
Open access Sep 2026

gVCF2CNV: a scalable pipeline for CNV detection from whole-genome sequencing data

Motivation Copy-number variants (CNVs) contribute to human disease and population trait variation. CNV detection from large whole-genome sequencing cohorts remains computationally demanding, as most methods require BAM or CRAM files. Genomic VCF (gVCF) files are smaller, routinely generated by standard variant-calling...

Mame Seynabou Diop, Florian Bénitìere, Kuldeep Kumar et al. · 0 citations
Review Open access Aug 2026

nf-cavalier: A Nextflow Pipeline for Rare Disease Variant Prioritization and Reporting

nf-cavalier is a Nextflow pipeline that automates genomic variant annotation, filtering, and reporting for individuals with rare Mendelian diseases. The pipeline takes as input variant callsets for an individual, family, or rare disease cohort, together with a target gene panel or a phenotype of interest. Variants are...

J. Munro, Joshua Reid, M. Bahlo et al. · 0 citations
Open access Sep 2026

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller

ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinicall...

J. E. Perdomo, Mian Umair Ahsan, Jasmine Akoto et al. · 0 citations
Open access Aug 2026

SVlog: a logic programming framework for understanding structural variation in genomic disease

Structural variants (SVs) are a diverse group of genetic variants defined by a minimum size of 50 base pairs. SVs account for the majority of all variant bases in a person’s genome and are commonly implicated in inherited disease and cancer. However, SV analysis is complex due to their wide variation in type and size,...

M. Gudkov, Andre L. M. Reis, M. Kumaheri et al. · 0 citations
Open access Sep 2026

VariantFlow: a selective-execution engine for efficient population genomic computation on large variant datasets

Population-scale sequencing now produces variant call sets with thousands of samples and millions of sites, making post-calling analysis a recurring bottleneck. Because the Variant Call Format stores every field of every record together, a tool answering a field-limited question still parses the unused annotations, FOR...

E. Estaji, Jian-Feng Mao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.