This work establishes SBX and the AXELIOS 1 as a transformative platform for high-scale single-cell isoform sequencing and demonstrates that this consensus approach successfully captures the vast isoform diversity of single-cell libraries and enables the accurate measure of differential isoform expression across distinct cell types in peripheral blood mononuclear cells.
Abstract
Single-cell RNA sequencing has transformed our understanding of cellular systems, yet the reliance on short-read sequencing restricts analysis to gene-level quantification and obscures the immense biological diversity generated by alternative splicing. While long-read sequencing technologies can capture full-length RNA and resolve transcript isoforms, current platforms remain constrained by throughput and high per-base costs, rendering them impractical for modern million-cell applications. To address this critical limitation, we developed and optimized sequencing-by-expansion (SBX) chemistry for high-throughput single-cell RNA isoform profiling. Integrated within the AXELIOS 1 sequencing platform, SBX employs a unique biochemical conversion process that transforms complementary DNA into expanded surrogate high signal-to-noise polymers called Xpandomers which are sequenced via translocation through a dense nanopore array yielding over 9.5 billion reads in a two-hour run. To leverage this unique data type for long-read single-cell RNA isoform sequencing, we developed the Consensus UMI Deduplication using Longest Length (CUDLL) algorithm, which computationally consolidates variable-length raw SBX reads into single, high-fidelity consensus reads, elevating sequence accuracy to 99.83% and maximizing per transcript read length. We demonstrate that this consensus approach successfully captures the vast isoform diversity of single-cell libraries and enables the accurate measure of differential isoform expression across distinct cell types in peripheral blood mononuclear cells. Furthermore, SBX coupled with CUDLL efficiently resolves T-cell and B-cell receptor clonotypes directly from whole-transcriptome libraries without the need for VDJ-specific target enrichment. Ultimately, this work establishes SBX and the AXELIOS 1 as a transformative platform for high-scale single-cell isoform sequencing.
Short-read sequencing-based single-cell transcriptomics represents the current gold standard for studying cellular transcriptomes but remains limited in its ability to resolve full-length transcript isoforms and splicing patterns. Long-read single-cell and single-nucleus RNA sequencing (LR sc/snRNA-seq) enables the transcriptome-wide characterization of full-length isoforms at cellular resolution, yet the relative performance of commercially available workflows remains insufficiently explored. Here, using nuclei extracted from a standardized multi-species benchmark sample and Oxford Nanopore Technologies long-read sequencing, we systematically benchmarked four LR snRNA-seq strategies: 10x Genomics 3’, 10x Genomics 5’, ArgenTag, and Parse Biosciences. Comparing transcriptome features qualitatively and quantitatively, as well as the concordance with matched short-read data and the ability to resolve cellular heterogeneity, we identified substantial method-specific differences in read length and yield, transcript coverage, isoform detection, and recovery of sample-specific biological information, with the 10x Genomics 3’ and 5’ assays emerging as the most balanced approaches for comprehensive isoform-resolved single-nucleus transcriptomics. Altogether, our study provides a systematic assessment of four commercially available workflows for performing LR snRNA-seq and highlights key methodological trade-offs related to distinct library preparation strategies, thus providing practical guidance for future isoform-resolved transcriptome studies at the single-nucleus level.
F. Köhler, Anna Delgado-Tejedor, Maik Zehnsdorf et al.· bioRxiv· 0 citations
Abstract Transcription—the process by which genomic DNA is converted into RNA—is a highly dynamic and tightly controlled process across all domains of life. Although bacteria were once regarded as relatively simple organisms, their transcriptomes are now recognized to be remarkably complex, heterogeneous, and subject to multilayered regulation. Despite the availability of an abundance of sequenced bacterial genomes, a comprehensive understanding of how bacteria tune their transcriptional output to adapt to changing environments remains lacking. To this end, SEnd-seq (simultaneous 5′ and 3′ end sequencing) was developed as a high-throughput approach uniquely capable of simultaneously capturing both 5′ and 3′ ends of individual RNA molecules, enabling the reconstruction of full-length transcripts. By capturing each RNA molecule as a distinct molecular entity with single-nucleotide resolution, SEnd-seq has uncovered previously unrecognized transcriptional features across diverse bacterial species, including even the well-studied Escherichia coli. This method performs robustly across a wide range of RNA species and organisms, including hard-to-lyse pathogens such as Mycobacterium tuberculosis. Moreover, SEnd-seq exhibits high sensitivity for detecting low-abundance RNA and is compatible with various target RNA enrichment strategies, as well as genetic, chemical, and functional perturbations, enabling context-specific transcriptomic analyses. As a versatile and broadly adaptable technology, SEnd-seq provides comprehensive insights into transcriptional regulation, RNA processing, and the coordination between RNA-based processes, thereby uncovering potential targets for antibiotic development. In the present review, we summarize the features and applications of SEnd-seq and discuss its future methodological development and expansion into broader biological and biomedical research contexts.
Fusion transcripts contribute to cancer, inherited diseases, developmental disorders, and evolution. Long-read RNA sequencing enables direct sequencing of full-length transcripts, creating new opportunities to detect complex fusion architectures, including previously inaccessible multi-segmented fusion transcripts. However, accurate transcriptome-wide fusion detection remains challenging because existing methods struggle to distinguish genuine fusion events from technical artefacts. Here, we present a comprehensive benchmark of transcriptome-wide fusion detection using simulated datasets and transcriptomes from three cancer cell lines across Oxford Nanopore Technologies (ONT) cDNA, PCR-cDNA, and direct RNA sequencing, Pacific Biosciences (PacBio) Kinnex sequencing, Illumina short-read RNA sequencing, six long-read fusion callers, and multiple analysis strategies. False-positive fusion calls remained the dominant limitation across sequencing platforms and algorithms. Increasing sequencing depth improved recall but also amplified spurious fusion calls, whereas higher read-support thresholds improved precision at the expense of sensitivity. ONT PCR-cDNA sequencing combined with CTAT-LR-Fusion achieved the best overall balance between precision and recall, whereas JAFFAL was the only caller to reliably identify simulated tri-gene fusions. Consensus calling reduced false positives but markedly reduced sensitivity, with only one of 400 simulated fusions detected by all six callers. Breakpoint localisation emerged as a major limitation across all methods. Long-read sequencing consistently recovered more validated fusion transcripts than short-read sequencing, enabled detection of complex tri-gene fusions, and produced more biologically plausible fusion landscapes with fewer promiscuous gene partners. Collectively, our results establish the first comprehensive benchmarking framework for transcriptome-wide fusion detection, using long-read RNA sequencing, and provide practical guidance for selecting sequencing workflows and computational strategies, while identifying key priorities for future algorithm development.
Ryley Dorney, Siyuan Wu, J. Y. Hung et al.· bioRxiv· 0 citations
Cell-to-cell transcriptional heterogeneity, or noise, is an intrinsic property of the transcriptome with implications for development, disease progression, and aging. Bulk RNA-seq masks this variability by averaging gene expression across cells, whereas single-cell RNA sequencing (scRNA-seq) resolves it. Nevertheless, separating biological noise from technical variance remains challenging, particularly across platforms with different chemistries. We benchmarked two widely adopted technologies, Evercode WT (SPLiT-seq, Parse Biosciences) and Chromium (10x Genomics), on human lymphoblastoid nuclei. Evercode WT achieved targeted sequencing depth and nuclei number far more reliably, and its random-hexamer priming yielded more intronic reads and non-coding RNA genes; Chromium recovered more cells and detected polyadenylated transcripts and cell-line markers more sensitively. Despite these opposing biases, the platforms showed comparable gene detection and strongly correlated expression profiles. Using datasets from both platforms, we defined a noise metric detrended from mean expression and showed that per-gene estimates were reproducible across chemistries. Noise was lower in G2M than in G1 and was most strongly associated with gene length rather than exonic length. Expression of genes with CpG-island promoters was less variable than that of those without. This study establishes a platform-independent basis for quantifying transcriptional noise and a framework for selecting an appropriate scRNA-seq platform. Highlights Parse Evercode WT demonstrates superior predictability in targeted cell recovery and sequencing depth estimations compared to 10x Chromium. Distinct biases: Parse captures intronic sequences; 10x targets polyadenylated mRNA. Detrended transcriptional noise is reproducible across both barcoding chemistries. Noise scales with gene length, not exonic length, and is lower at CGI promoters.
Rafal Czapiewski, M. Chiang, James Ding et al.· bioRxiv· 0 citations
Long-read RNA sequencing directly resolves the full structures of RNA transcripts. Advances in throughput now enable the generation of deeply sequenced cohorts of hundreds of samples, making joint transcript discovery across large datasets possible. However, existing transcript identification methods were designed for small datasets, which limits their applicability at this scale. Here, we present Isocall, a scalable and deterministic computational method for jointly calling transcripts from multiple PacBio long-read RNA sequencing samples. Isocall converts aligned full-length non-concatemer reads into compact per-sample transcript profiles, merges these profiles across samples, and jointly identifies known and novel transcripts supported by reads in the analyzed dataset. Filtering is tunable as presets provide coarse control and individual parameters, including relative abundance and internal priming thresholds, provide fine control. Isocall demonstrated high precision in our accuracy benchmarks, including the 69 transcript WTC11 SIRV spike-in controls, for which Isocall reported 0-2 false-positive transcripts per sample across the three SIRV mixes at default settings. To demonstrate scalability, we applied Isocall to 206 samples, totalling 3.5 billion raw reads, from the Human Pangenome Reference Consortium. After parallelized pbmm2 alignment and Isocall profile, the call step performed joint calling across the entire dataset in 25 minutes, using 1.3 GB of peak memory and 8 threads. Finally, in Genome in a Bottle samples with matched SNP genotypes, splice-site polymorphisms provide an additional measure of call accuracy: Isocall recovered 337 polymorphic splice sites, including a de novo donor site in BTN3A1 that corresponds to a complete isoform switch on the mutant allele.
E. Dolzhenko, Megan D. Schertzer, Ryan Gossart et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.