Reanalyzed whole-genome sequencing data from 1,578 unsolved probands and identified pathogenic variants in multiple snRNA genes, including RNU4-2, RNU2-2, RNU5B-1, and RNU4ATAC, and developed an snRNA-extended WES approach by incorporating capture probes targeting 50 snRNA genes into a standard exome design.
Abstract
Summary Pathogenic variants in small nuclear RNA (snRNA) genes have recently emerged as a major cause of Mendelian disorders, particularly neurodevelopmental disorders, yet they remain difficult to detect in routine diagnostics because conventional whole-exome sequencing (WES) does not capture snRNA loci. Here, we reanalyzed whole-genome sequencing (WGS) data from 1,578 unsolved probands and identified pathogenic variants in multiple snRNA genes, including RNU4-2, RNU2-2, RNU5B-1, and RNU4ATAC, accounting for 1.2% (19 patients) of previously unsolved cases. We then developed an snRNA-extended WES approach by incorporating capture probes targeting 50 snRNA genes into a standard exome design. Benchmarking demonstrated robust, uniform coverage across all targeted snRNA loci without increasing sequencing depth. Applying this approach to patient samples reliably detected disease-causing snRNA variants previously identified by WGS. Our results establish snRNA-extended WES as a cost-effective and scalable strategy to improve diagnostic yield and bridge the gap between recent gene discoveries and clinical genomic practice.
The results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.
F. Mazzarotto, Özem Kalay, E. Arslan et al.· Genome Medicine· 0 citations
This work describes the full spectrum of genetic variation and shows that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions.
J. Lin, J. Gustafson, J. Wertz et al.· medRxiv· 0 citations
A single genomic assay that delivers complete information across variant classes remains an aspirational goal. Currently, researchers and clinicians rely on an inefficient, expensive combination of short-read sequencing for single-nucleotide variants (SNVs) and small indels, comparative genomic hybridization (CGH) arrays for copy number variants (CNVs), and optical mapping and long-read sequencing for complex rearrangements, limiting the full potential of genomic discovery. To address these issues, TruPath Genome provides a one-test-for-all solution. By combining PCR-free whole-genome sequencing (WGS) with proximity-mapped read technology, it achieves high-resolution detection of SNVs and indels alongside long-range phasing for CNVs and structural variant (SV) refinement. We applied TruPath Genome on six clinical samples that were previously resolved by conventional methods. Across the cohort, TruPath Genome delivered coverage and variant-calling performance comparable to conventional WGS while achieving superior long-range phasing and enabling precise breakpoint resolution for clinically relevant structural events. This highlights TruPath Genome’s potential to consolidate genomic testing pipelines, accelerate diagnosis, and expand access to advanced genomic insights. Furthermore, its ultra-long-range data facilitates telomere-to-telomere assemblies and pangenome development, advancing our understanding of genome biology at an unprecedented scale.
Background: Small nuclear RNAs (snRNAs) are RNA components of the major and minor spliceosomes that play a core role in splice-site recognition and control of the splicing process. Variants in genes that produce snRNAs are increasingly recognised as major contributors to rare disorders, including neurodevelopmental disorders (NDD) and retinal dystrophies (collectively termed RNUopathies, a subset of spliceosomopathies). Clinical interpretation of variants in snRNAs is, however, challenging and existing guidance to support clinical variant classification does not adequately capture the unique features of snRNAs that necessitate a bespoke approach. Methods: We quantified the elevated background mutation rate in snRNA genes using de novo variants from 12,007 trios and assessed mutation density in 76,215 genome sequenced individuals in gnomAD. We convened a panel of clinical, research, and industry scientists with wide-ranging expertise in clinical variant interpretation and classification and expert knowledge in snRNA genes to draft and refine a guidance document. Results: We detail important considerations for variant classification in snRNA genes. These include: the difficulties of variant identification which requires genome or targeted sequencing approaches, the large number of gene paralogs with high sequence identity that complicate read mapping and variant calling, and historical inaccuracies in snRNA gene annotation. Further we show a ~50-fold increase in de novo mutation rate in snRNA genes compared to intergenic sequence and discuss the implications of this for variant classification. We provide a set of specific recommendations for classifying variants in snRNA genes. Finally, we introduce RNUdb, an interactive web-based tool to support snRNA variant annotation and classification. Conclusions: We provide the first guidance for clinical variant classification in snRNA genes and anticipate that this will support routine screening and analysis of snRNA genes in clinical genetic testing.
Elston N. D'souza, A. Blakes, Robin Paluch et al.· medRxiv· 0 citations
Splicing is a complex molecular mechanism in eukaryotic cells essential to gene expression and regulation, involving more than 300 protein-coding genes (PCGs) and 43 small nuclear RNA (snRNA) genes. However, fewer than 30 gene-disease relationships have been described as spliceosomopathies to date. This discrepancy suggests the splicing machinery as an underexplored area for human disease gene discovery. For snRNA currently classified as pseudogenes, we prioritized candidates with similar epigenomic, genomic, and hypermutability features as functional snRNA genes. Population-variant-depletion analysis was performed to identify regions under negative selection. We analyzed rare variants in PCGs and snRNA genes and prioritized snRNA pseudogenes across a large heterogeneous rare disease cohort. There was high concordance for prioritizing genes annotated as pseudogenes by the variant-depleted region analysis (9) and by random forest models of hypermutation, genomic and epigenomic features (6). We identified 26 variants of interest across six PCGs with established gene-disease relationships (GDRs) and 14 genes not yet disease-associated, including one pseudogene across 30 individuals. For snRNAs genes, we identified 49 variants of interest located in seven genes with established GDR and 11 genes not yet disease-associated, including two pseudogenes across 80 individuals. This study highlights the importance of splicing-related PCG and snRNA in the genetic etiology of rare diseases. By leveraging specialized approaches for prioritizing pseudogenes, combined with the PCG and snRNA analysis, the genes and variants expand the variant pathogenicity spectrum of spliceosomopathies and suggest variants for follow-up case series and future functional validation.
O. Messaoud, S. DiTroia, R. Tarawneh et al.· medRxiv· 0 citations
Whole-exome sequencing (WES) enables the identification of rare germline variants contributing to pediatric diseases. Trio-based sequencing, comparing affected children with their parents, is particularly effective for rare disease genetics. However, WES data analysis requires bioinformatics expertise, varies across institutions, and is often incompatible with clinical workflows. We developed T-Rex (Trio Rare variant analysis of EXomes), a cross-platform desktop application that enables the standardized and local analysis of WES germline Trio data without the need for programming knowledge. T-Rex integrates state-of-the-art tools for alignment, dual-variant calling (GATK HaplotypeCaller + VarScan2), annotation (SNPEff/SNPSift), rare-variant filtering based on population frequencies (gnomAD), and family-based statistical testing, including the Transmission Disequilibrium Test with multiple-testing correction. Benchmarking of the dual-caller strategy on the Genome in a Bottle Ashkenazim Trio demonstrates high precision (99.2%) while maintaining robust sensitivity (91.1%). User testing (n = 13) confirmed quick learning across clinicians and researchers. Application to a cohort of n = 121 pediatric cancer Trio datasets, filtering for rare protein-coding variants (MAF ≤ 0.1% in gnomAD v4.1), validated all assessable previously reported pathogenic variants. Overall, T-Rex enables clinicians to robustly analyze WES Trio data in compliance with data protection regulations without requiring additional software licenses. As one of the first platforms for comprehensive WES Trio analysis that requires no programming expertise while providing reproducible, end-to-end workflows for clinical genomics, T-Rex facilitates collaborative research between clinics and reduces reliance on external providers.
Sara-Luisa Reh, C. Walter, J. Lohse et al.· Scientific Reports· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.