Skip to content

Author

F. Sedlazeck

5 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Isocall enables scalable transcript identification from long-read RNA-sequencing data

Long-read RNA sequencing directly resolves the full structures of RNA transcripts. Advances in throughput now enable the generation of deeply sequenced cohorts of hundreds of samples, making joint transcript discovery across large datasets possible. However, existing transcript identification methods were designed for small datasets, which limits their applicability at this scale. Here, we present Isocall, a scalable and deterministic computational method for jointly calling transcripts from multiple PacBio long-read RNA sequencing samples. Isocall converts aligned full-length non-concatemer reads into compact per-sample transcript profiles, merges these profiles across samples, and jointly identifies known and novel transcripts supported by reads in the analyzed dataset. Filtering is tunable as presets provide coarse control and individual parameters, including relative abundance and internal priming thresholds, provide fine control. Isocall demonstrated high precision in our accuracy benchmarks, including the 69 transcript WTC11 SIRV spike-in controls, for which Isocall reported 0-2 false-positive transcripts per sample across the three SIRV mixes at default settings. To demonstrate scalability, we applied Isocall to 206 samples, totalling 3.5 billion raw reads, from the Human Pangenome Reference Consortium. After parallelized pbmm2 alignment and Isocall profile, the call step performed joint calling across the entire dataset in 25 minutes, using 1.3 GB of peak memory and 8 threads. Finally, in Genome in a Bottle samples with matched SNP genotypes, splice-site polymorphisms provide an additional measure of call accuracy: Isocall recovered 337 polymorphic splice sites, including a de novo donor site in BTN3A1 that corresponds to a complete isoform switch on the mutant allele.

E. Dolzhenko, Megan D. Schertzer, Ryan Gossart et al. · 0 citations
Open access Sep 2026

Integrated map of somatic mosaicism across human tissues in 25 individuals

Although all cells in the body descend from one genome, they accumulate distinct genetic and epigenetic changes over a lifetime, producing a mosaic of somatic variation that can shape development, aging, and disease. This mosaicism is often studied in isolation, leaving unclear how these forms relate within and between individuals. Here, we present the first integrated analysis of the Somatic Mosaicism across Human Tissues (SMaHT) Network’s production resource, profiling up to 20 tissues from 25 donors using short- and long-read, duplex, single-cell, transcriptomic, and epigenomic sequencing, alongside donor-specific near-telomere-to-telomere assemblies. Somatic mutation burden cannot be captured by a single data type or metric, as tissues accumulate distinct variant classes largely independently of one another. Long-read and single-cell data resolved cell-type-specific mutational processes, traced mobile element insertions to source loci, and revealed the developmental timing and functional consequences of individual mutations. Donor-specific assemblies uncovered elevated mutation rates within centromeres and segmental duplications inaccessible to standard reference genomes, while haplotype-resolved chromatin and methylation data showed that nongenetically-deterministic epigenetic states are pervasive across tissues. Together, these findings provide an integrated, multi-scale portrait of somatic mosaicism across the human body, establishing a baseline against which its contributions to aging and disease can be measured.

F. Sedlazeck, Tim H. H. Coorens, Peter J. Park et al. · 0 citations
Open access Aug 2026

SALRR: Scalable Analysis of Long-Read RNA-Seq Enables Comprehensive Transcriptome Profiling in Human Brain

Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.

C. Kouam, Jackson Mingle, Pilar Álvarez Jerez et al. · 0 citations
Open access Aug 2026

A complete diploid human genome benchmark for personalized genomics

SUMMARY Human genome sequencing typically relies on mapping reads to a reference genome to call variants, but this approach introduces technical biases, excluding duplicated and structurally polymorphic regions of the genome. To overcome this, we present a telomere-to-telomere genome benchmark with near-perfect accuracy across 99.4% of the diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), which were absent from prior benchmarks. We annotated genes and repeats on both haplotypes, including 19,956 protein-coding genes on the maternal haplotype and 19,190 on the paternal haplotype, and developed new methods to measure the accuracy of reads, phased variant call sets, and assemblies against a diploid reference. Genome-wide analyses show that de novo assembly resolves 2%–7% more sequence and outperforms variant calling accuracy by an order of magnitude, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Nancy F. Hansen, Nathan Dwarshuis, Hyun Joo Ji et al. · 7 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.