Skip to content
Open access

A Sample to Results Workflow for Compositional Analysis of Multiplexed Amplicon Sequencing Experiments

Jul 2026 · bioRxiv · 0 citations · 45 references
Medicine Biology

TL;DR

A standardized workflow for multiplexed amplicon sequencing from sample collection through data analysis for diverse sample types that produces data and publication-ready figures for multiple taxonomic and functional genes for carbon, nitrogen, phosphorus, sulfur, and arsenic cycling for each sample analyzed is developed.

Abstract

Microbial communities play key roles in the transformation and cycling of elements ranging from required macronutrients to toxic metalloids. Next-generation sequencing has been applied across multiple ecosystems to probe the interplay of microbial community structure and functional potential with respect to elemental cycling. Shotgun metagenomics collects marker gene sequences without amplification and is costly for large numbers of samples and deep coverage. Conversely, amplicon sequencing of taxonomic marker genes, e.g. 16S and 18S rRNA, is cost-effective for large numbers of samples, but provides limited functional insight. A middle ground between the two approaches is needed to analyze community structure and functional potential within a sample while remaining cost-effective with high throughput. To address this need, we developed a standardized workflow for multiplexed amplicon sequencing from sample collection through data analysis for diverse sample types, including freshwater, sediments, and soils, that produces data and publication-ready figures for multiple taxonomic and functional genes for carbon, nitrogen, phosphorus, sulfur, and arsenic cycling for each sample analyzed. The workflow’s utility was shown by analyzing 11 taxonomic and functional gene amplicons sequenced from 25 samples with high technical replicate similarity. The workflow is named CAMASE for Compositional Analysis of Multiplex Amplicon Sequencing Experiments. This proof-of-concept shows that CAMASE economically produces standard amplicon sequencing outputs (ASV/OTU counts and taxonomy, PCA, and relative abundance plots) for hundreds of amplicon by sample combinations and provides specific recommendations for implementation. GRAPHICAL ABSTRACT Samples are collected in a preservative and material collected on filters prior to DNA extraction. Target gene amplicons are produced in parallel with internal barcodes enabling sequencing in a single run followed by compositional data analysis. All wet lab protocols, code markdowns, and templates for required metadata files are available at https://hansonlabgit.dbi.udel.edu/aprange/CAMASE. Created in BioRender. Bennett, A. (2026) https://BioRender.com/ymnojt0

Read PDF

Similar papers

Review Open access Aug 2026

Bioinformatic tools for microbiome analysis: from raw sequences to biological insights

This review presents a practical, workflow-oriented guide to microbiome data analysis, from raw DNA sequence processing to statistical interpretation and biological insight, and highlights emerging technologies, including machine learning methods that are beginning to reshape the field.

Jenna Poelzer, D. Wishart · 0 citations
Review Open access Jul 2026

Metagenomics: Tools To Unpack the Total Genomes of Microorganisms: Methods, Applications, and Emerging Frontiers- A Narrative Review

Applications across healthcare, environmental science, agriculture, biotechnology, and industry are reviewed with particular emphasis on clinical metagenomic next-generation sequencing (mNGS) for infectious disease diagnostics, antimicrobial resistance (AMR) surveillance, gut microbiome research, and precision medicine.

A. Alsharksi · 0 citations
Open access Aug 2026

PUDU (pipeline for universal diversity unveiling): an accessible end-to-end workflow for taxonomic profiling and ecological visualization of environmental microbiomes across amplicon, shotgun, and long-read sequencing

Background Environmental microbiome research has advanced through three complementary sequencing modalities, targeted 16S rRNA amplicon sequencing, whole-genome shotgun (WGS) metagenomics, and long-read full-length 16S rRNA profiling, each supported by distinct toolsets with heterogeneous outputs, variable configurations, and different levels of reproducibility documentation. Existing pipelines are typically modality-specific, require substantial configuration expertise, or produce outputs that need further custom scripting before standard ecological analyses can begin. This analytical fragmentation introduces avoidable technical variability and complicates cross-study reproducibility and comparability. PUDU addresses this by integrating all three modalities into a single reproducible workflow with simplified configuration, harmonized outputs across classifiers, and direct compatibility with downstream ecological analysis frameworks. Results We present PUDU (Pipeline for Universal Diversity Unveiling), a modular Snakemake workflow that supports amplicon (short-read 16S), shotgun metagenomics (WGS), and long-read 16S analyses from raw reads to standardized outputs for downstream microbial ecology. PUDU performs technology-aware preprocessing and centralized quality control, and integrates established taxonomic approaches, including DADA2 for amplicons, Emu for full-length 16S long reads, and Kraken2/Bracken and Centrifuger for WGS. Across methods, PUDU produces harmonized count and relative-abundance tables at user-defined taxonomic ranks, Krona files, and a standardized Phyloseq-compatible R object to streamline diversity analyses and statistical workflows. PUDU also provides an integrated Shiny interface for metadata-aware alpha/beta diversity, ordination, community composition, and shared-taxa exploration with exportable figures and taxa tables. We demonstrate PUDU on two publicly available environmental datasets spanning rhizosphere WGS and long-read marine sediment 16S, yielding broadly consistent community-level patterns across classifiers (Spearman ρ = 0.936 at phylum level; PERMANOVA R2 = 0.87–0.95) with peak memory below 45 GB on a standard Linux workstation. Conclusion PUDU is an end-to-end, reproducible, and extensible framework that enables standardized taxonomic profiling and ecology-oriented analysis across sequencing modalities. By combining harmonized outputs, Phyloseq interoperability, and an integrated visualization layer, PUDU facilitates reproducible, standardized, and comparable environmental microbiome analysis from raw reads to interpretable ecological insights.

Alejandro Medaglia-Mata, Pablo Rojas-Rodríguez, V. Bystrý et al. · 0 citations
Review Oct 2026

Insights into the analysis of microbial communities in fermented foods from the perspective of DNA-based techniques.

The quality, flavor, and stability of fermented foods depend on the microbial community. However, microbial dynamics are difficult to observe directly, leading to limited control over fermentation. High-throughput sequencing is a revolutionary tool for microbial characterization, among which DNA-based amplicon and metagenomic sequencing are core techniques. Nevertheless, the related data processing workflows in the context of fermented foods have not yet been systematically summarized, hindering the translation of research findings into fermentation practices. This review clarifies the applications of amplicon and metagenomic sequencing in fermented foods. For amplicon sequencing, the impacts of target regions, data preprocessing, and reference databases are addressed. For metagenomic sequencing, sequencing strategies, read-based and binning-based analytical methods, functional annotation, and species-specific databases are discussed. In addition, major strategies for downstream analysis of community data are summarized, including microbial diversity, co-occurrence networks, niche and community assembly, key environmental drivers, and machine learning-based prediction. Amplicon sequencing efficiently reveals microbial succession during fermentation but has limitations in functional annotation. Metagenomic sequencing is notable for functional annotation, enabling the linkage between microbial communities and metabolic potential alongside community characterization. Standardized data preprocessing and specific databases are critical for improving characterization. For community data, integrated analysis allows uncovering the driving factors of microbial succession, thereby helping to regulate fermentation. Notably, the compositional nature of the data must be considered and validated to avoid spurious associations. In summary, the exponential growth of sequencing data will propel the era of precision fermentation.

Hao Zhou, Lijun Yan, Ling Zhang et al. · 0 citations
Open access Aug 2026

A streamlined workflow for high throughput metaproteomic analysis of the rumen microbiome.

This study addresses current limitations in the application of metaproteomics to rumen microbiome research by developing a streamlined and scalable sample preparation workflow that enables more efficient processing of larger sample sets.

A. L. Pedersen, L. Dayon, Michael Affolter et al. · 0 citations
Open access Aug 2026

Novel Insights into Metagenomic-Assembled Genomes from Layer Chicken Housing Environment

Culture-independent techniques are playing a major role in exploring unique and novel microbial communities from complex ecosystems, leading to an outstanding impact on our basic understanding of the tree of life. Microbial communities are not extensively studied in layer chicken housing environments, particularly from the point of view of taxa carrying antimicrobial resistance genes, virulence genes and their functional potential. This study aimed to extract metagenomic-assembled genomes (MAGs) from the Illumina short-reads shotgun metagenomics sequenced data that originated from an Alberta poultry barn environment and then to study host tracking of antimicrobial resistance genes (ARGs) and the roles of genes involved in functions related to ammonia production, short-chain fatty acid (SCFA)-related pathways, sulfur metabolism, methane emission, stress and disinfectant-related pathways. A total of 251 high-quality MAGs were extracted, including 249 bacterial and two archaeal genomes from sequencing data of 30 metagenomic sequencing samples comprising 15 air and 15 manure samples collected from 15-layer farms. Interestingly 22 bacterial MAGs were not classified to species levels using GTDB-based classification. ARGs were mainly harbored by the genera Staphylococcus, Alistepes, Romboutsia, and Enterococcus. Bacteroides is a main taxon carrying ARGs in air samples. Ammonia production-related genes were mainly tracked in Staphylococcus, Ruminococcus and Corynebacterium genera. The assimilatory sulfate reduction genes responsible for sulfur metabolism and hydrogenase-related genes responsible for hydrogen cycling were traced from Staphylococcus originated from both air and manure. The current study provides characterizations of MAGs from a poultry housing environment by linking microbial taxa with virulence, resistance, and metabolic functions. The findings emphasize the role of microbiota in shaping gas emissions and AMR, with implications for poultry health and worker’s safety and the ultimate aim of sustainable poultry production.

Awais Ghaffar, M. F. Abdul-Careem · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.