SVPLEX is a Nextflow pipeline for cohort-level structural variant detection from short-read whole-genome sequencing data and generates a merged consensus callset across the analysis cohort, which can be used to assess cohort-specific variation, remove technical artefacts, and serve as input for rare disease variant prioritisation workflows.
Abstract
SVPLEX is a Nextflow pipeline for cohort-level structural variant detection from short-read whole-genome sequencing data. The pipeline implements six different structural variant callers with different strengths and weaknesses, integrating different levels of evidence for SVs, and generates a merged consensus callset across the analysis cohort. Callset filtering is achieved by leveraging consensus among multiple individual callers and by ensuring that deletion and duplication calls are supported by observable changes in read depth. The output merged cohort SV callset can then be used to assess cohort-specific variation, remove technical artefacts, and serve as input for rare disease variant prioritisation workflows. SVPLEX is user-friendly, reproducible, scalable, and can be executed flexibly on either a local workstation, a high-performance compute (HPC) cluster, or deployed on cloud infrastructure. The required inputs are alignment files for the cohort of interest, and the output is a single merged cohort structural variant VCF. SVPLEX is available on GitHub (bahlolab/SVPLEX) and is licensed under the MIT open-source licence.
A step-by-step protocol for snpArcher, a Snakemake-based workflow that takes raw sequencing reads and a reference genome as input and produces a filtered, joint-called VCF suitable for downstream population genomic analysis, is presented.
Cade Mirchandani, Abdelmajid Omarjee, Guillaume Achaz et al.· Molecular biology and evolut...· 0 citations
Motivation Copy-number variants (CNVs) contribute to human disease and population trait variation. CNV detection from large whole-genome sequencing cohorts remains computationally demanding, as most methods require BAM or CRAM files. Genomic VCF (gVCF) files are smaller, routinely generated by standard variant-calling...
nf-cavalier is a Nextflow pipeline that automates genomic variant annotation, filtering, and reporting for individuals with rare Mendelian diseases. The pipeline takes as input variant callsets for an individual, family, or rare disease cohort, together with a target gene panel or a phenotype of interest. Variants are...
J. Munro, Joshua Reid, M. Bahlo et al.· bioRxiv· 0 citations
ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinicall...
J. E. Perdomo, Mian Umair Ahsan, Jasmine Akoto et al.· NAR Genomics and Bioinformat...· 0 citations
Structural variants (SVs) are a diverse group of genetic variants defined by a minimum size of 50 base pairs. SVs account for the majority of all variant bases in a person’s genome and are commonly implicated in inherited disease and cancer. However, SV analysis is complex due to their wide variation in type and size,...
M. Gudkov, Andre L. M. Reis, M. Kumaheri et al.· bioRxiv· 0 citations
Population-scale sequencing now produces variant call sets with thousands of samples and millions of sites, making post-calling analysis a recurring bottleneck. Because the Variant Call Format stores every field of every record together, a tool answering a field-limited question still parses the unused annotations, FOR...
E. Estaji, Jian-Feng Mao· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.