Aug 2026· Nature Biomedical Engineering· 0 citations· 43 references
Medicine
TL;DR
This work presents CRISPRLungo, a computational pipeline specifically designed for long-read amplicon sequencing of gene edited samples that incorporates unique molecular identifier-based error correction and statistical filtering to distinguish true editing events from background noise, enabling robust detection of small indels and structural variants.
Abstract
Long-read sequencing can characterize complex genome editing-induced DNA sequence changes such as large deletions, insertions and inversions that are difficult to detect using short-read sequencing. However, PCR amplification and sequencing errors complicate accurate variant detection, and existing analysis tools are not optimized for gene editing specific allelic outcomes. Here we present CRISPRLungo, a computational pipeline specifically designed for long-read amplicon sequencing of gene edited samples. CRISPRLungo incorporates unique molecular identifier-based error correction and statistical filtering to distinguish true editing events from background noise, enabling robust detection of small indels and structural variants. Through systematic benchmarking using simulated datasets, we demonstrate that CRISPRLungo outperforms existing approaches in both accuracy and read recovery. CRISPRLungo supports both Oxford Nanopore and PacBio platforms and identifies previously undetected structural variant edits such as inversions in published CRISPR datasets. To demonstrate allele-specific edit quantification, we applied CRISPRLungo to analyse edited primary cells from a patient harbouring compound heterozygous SBDS mutations, accurately quantifying SBDS editing outcomes despite contaminating reads from the homologous SBDSP1 pseudogene. To maximize accessibility, we developed a fully client-side web application requiring no installation, making advanced long-read analysis accessible to researchers regardless of computational expertise. CRISPRLungo is freely available at https://github.com/pinellolab/CRISPRLungo with a user-friendly web interface available at https://pinellolab.github.io/CRISPRLungo .
CRISPR interference (CRISPRi) is a powerful technology for studying loss-of-function phenotypes, enabling transient and reversible control of gene expression without the introduction of double-stranded DNA breaks. The cost of conducting large-scale CRISPR screens necessitates the selection of effective and specific single-guide RNAs for the design of compact libraries. While several genome-wide CRISPRi-Cas9 libraries have been created, updates to transcript annotations, the generation of higher-resolution chromatin accessibility datasets, and the development of newer on-target prediction models motivate an updated CRISPRi library design approach. Here, we generate large CRISPRi datasets tiling essential and nonessential genes. We compare the performance of multiple KRAB domain systems, develop an updated CRISPRi-specific on-target scoring scheme, and quantitatively characterize off-target effects associated with seed-sequence patterns. We leverage these findings to design an optimized CRISPRi-Cas9 library, Katsano, and validate its performance with genome-wide viability screens.
Smriti Srikanth, Fengyi Zheng, Laura M Drepanos et al.· Cell Genomics· 0 citations
While long-read sequencing technologies (e.g., PacBio Revio, ONT) have revolutionized high-quality genome assembly for the human pangenome, mitochondrial genome (mtDNA) analysis still largely relies on short-read and Sanger sequencing. However, short-read sequencing often lacks the resolution required to resolve complex variations due to the unique features of mtDNA, such as high mutation rates and repetitive homopolymeric regions, which frequently lead to alignment artifacts and mapping ambiguities. To address this, we evaluated whether applying long-read sequencing to mtDNA improves analytical quality in empirical data. Through comprehensive bioinformatics analyses, we compared the performance of long-read sequencing against short-read sequencing and microarrays. Our results revealed that long-read sequencing detected the highest number of variants (n = 533), significantly outperforming both short-read sequencing (n = 525) and microarrays (n = 49). Notably, both sequencing methods provided significantly higher resolution in haplogroup assignment compared to microarrays in terms of phylogenetic depth (p < 0.05). Long-read sequencing demonstrated superior detection power, particularly for InDels. We identified two novel non-synonymous variants, including a unique InDel detected exclusively by long-read sequencing. Protein modeling and stability analysis validated that this InDel causes structural instability (RMSD > 2.0 Å, ΔΔG = -45.21 kcal/mol). Furthermore, we confirmed that this novel InDel is shared among haplogroup A samples in both the 1000 Genomes Project ONT dataset and the Korean population, highlighting the practical implications of long-read sequencing for molecular biology and population genetics.
Hyung Jun Kim, Sung Min Kim, Kyungheon Yoon et al.· bioRxiv· 0 citations
Allele-specific CRISPR/Cas editing is a powerful tool with great potential for treating genetic diseases and for uncovering the effects of allelic diversity. By targeting commonly inherited single nucleotide polymorphisms (SNPs), a small number of gRNAs can treat many more individuals than targeting rare disease mutations. However, current tools for identifying common targetable variants and generating CRISPR guide RNAs (gRNA) have fundamental conceptual and technical limitations. Here, we introduce EXCAVATE-HT (EXtracting Common Allelic VAriants for Targeted Editing in High-Throughput) a bioinformatic tool that mines population variant data to generate CRISPR libraries targeting genomic loci for allele-specific editing. Users define their loci of interest, Cas species, and SNP frequency, then EXCAVATE-HT outputs an annotated list of allele-specific gRNAs. EXCAVATE-HT can also generate libraries of gRNA pairs to enable excision. We illustrate the use of EXCAVATE-HT to design and characterize multiple gRNA libraries for allele-specific targeting of the disease gene, Cone-Rod Homeobox (CRX). EXCAVATE-HT revealed multiple excisions that could treat >30-fold more patients than targeting a single CRX disease mutation.
Akshita G Saxena, G. D. Ramey, John A. Capra et al.· bioRxiv· 0 citations
CRISPR knockout (CRISPRko) and CRISPR interference (CRISPRi) are two workhorse technologies for loss-of-function studies, yet direct comparisons between the two are scant relative to their widespread adoption. Here, we establish benchmarking libraries for Cas9-based CRISPRko and CRISPRi screens using Perturb-seq as the read-out. For both modalities, we observe consistent transcriptional signatures among cells with the same genes perturbed, strong evidence of on-target signal. We also examine tradeoffs between modalities: while CRISPRi guides demonstrate heightened rates of off-target activity, we also observe artifacts stemming from the cellular response to double-stranded breaks with the use of CRISPRko. The libraries and analyses presented here will be a useful benchmarking and de-risking resource for any group preparing for a large-scale Perturb-seq screen.
Laura M Drepanos, Berta Escude Velasco, Abigail J Chase et al.· bioRxiv· 1 citation
Rare diseases, most of which have a genetic basis, remain a major challenge due to diagnostic delays and limited therapeutic options, particularly within the Middle Eastern regions. These countries exhibit a heightened prevalence of genetic disorders attributable to their distinctive genetic architecture. Advances in long-read sequencing (LRS) technologies have significantly improved our ability to detect complex genetic variations, including structural variants (SVs), repeat expansions, and mutations in previously inaccessible genomic regions, thereby increasing the diagnostic yield in rare disease cohorts. In parallel, the rapid evolution of gene-editing platforms such as CRISPR/Cas9, base editors, and prime editors has opened new possibilities for addressing the biological pathways of the disease and achieving precise therapeutic correction of pathogenic variants causing the disease. Importantly, the integration of LRS with gene-editing approaches establishes a continuum from accurate variant discovery and functional characterization to the development of personalized therapies. This review highlights recent progress in both fields, discusses their complementary roles in rare disease research, and explores the translational opportunities and ethical challenges of combining these technologies to advance precision medicine. In addition, the review addresses emerging ethical and regulatory considerations associated with the clinical translation of long-read sequencing and next-generation gene-editing technologies, particularly in the context of rare disease precision medicine.
Anshida Konamveettil Abdul Latheef, Mohammad K Ali, O. Farahat et al.· Frontiers in Medicine· 0 citations
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.