Aug 2026· Molecular biology and evolution· Vol 43· 0 citations· 25 references
Medicine
TL;DR
A step-by-step protocol for snpArcher, a Snakemake-based workflow that takes raw sequencing reads and a reference genome as input and produces a filtered, joint-called VCF suitable for downstream population genomic analysis, is presented.
Abstract
Abstract Population genomic studies in nonmodel organisms increasingly depend on whole-genome resequencing, yet translating raw reads into reliable variant callsets remains a practical challenge due to the complexity of multistep bioinformatics pipelines and the absence of species-specific best practices. Here, we present a step-by-step protocol for snpArcher, a Snakemake-based workflow that takes raw sequencing reads and a reference genome as input and produces a filtered, joint-called Variant Call Format (VCF) file suitable for downstream population genomic analysis. We guide users through six phases: installation and environment setup, sample sheet creation, run configuration, execution on local or high-performance computing systems, quality control review using an interactive HTML dashboard, and downstream analysis, focusing on postprocessing and filtering. The quality control (QC) dashboard aggregates individual-level metrics including principal component analysis, relatedness estimation, depth-missingness diagnostics, and admixture analysis to help identify batch effects, contamination, cryptic relatedness, and outlier samples before downstream analysis. We demonstrate the impact of sequential filtering steps on the site frequency spectrum and demographic inference using a dataset of 137 burrowing owl (Athene cunicularia) genomes, showing how removal of low-coverage individuals, sex-linked scaffolds, and regions of excess heterozygosity eliminates artifacts that would otherwise bias inference of population size history. This protocol is intended as a practical companion to the original snpArcher publication, enabling researchers working with nonmodel organisms to produce and evaluate analysis-ready variant callsets in a reproducible manner.
WGS2IBI provides a scalable, reproducible, and accessible workflow resource for WGS analysis, enabling efficient population- and individual-level genetic studies without local installation and lowering practical computational barriers for large-scale WGS studies.
Yasaman J. Soofi, Md Asad Rahman, Jinxing Ren et al.· BMC Bioinformatics· 0 citations
SVPLEX is a Nextflow pipeline for cohort-level structural variant detection from short-read whole-genome sequencing data and generates a merged consensus callset across the analysis cohort, which can be used to assess cohort-specific variation, remove technical artefacts, and serve as input for rare disease variant pri...
BLink-seq is presented, a novel Haplotagging method that is compatible with standard short-read next-generation sequencing platforms, is locally reproducible with low-cost reagents, and is scalable for high-throughput sample processing.
Azwad R Iqbal, Pavel V. Dimens, J. Rick et al.· bioRxiv· 0 citations
MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer–based statistics to detect biological variation without requiring alignment, is presented.
L. Mboning, Maciej Dlugosz, Marek Kokot et al.· bioRxiv· 0 citations
A workflow to minimize the effect of WES capture inconsistencies in single-nucleotide variation (SNV) data is proposed, which leads to a considerable decrease in the batch effect signal, potentially increasing the likelihood of finding true biological signals.
Laura Jarosz, M. Ochocki, Julia Merta et al.· Methods and Protocols· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.