Skip to content
Open access

Coordinate- and Sequence-Based Features for a new Combined Annotation-Dependent Depletion Framework of Structural Variants (CADD-SV v2.0)

Jul 2026 · bioRxiv · 0 citations · 6 references
Biology

TL;DR

This version introduces a unified Random Forest model trained on an expanded set of proxy-neutral and proxy-deleterious variants drawn from human and non-human primate genomes, and substantially improves the computational workflow, increasing predictive power for genome-wide SV interpretation.

Abstract

Structural variants are a major source of genomic variation and contribute to human disease and evolution through diverse mechanisms, yet their functional interpretation remains challenging. We present CADD-SV v2.0, an improved machine learning framework for scoring SV deleteriousness that expands on the original CADD-SV implementation. This version introduces a unified Random Forest model trained on an expanded set of proxy-neutral and proxy-deleterious variants drawn from human and non-human primate genomes. The model integrates updated genomic annotations, including constraint metrics, regulatory elements, and chromatin architecture features. It scores Deletions, Insertions, Duplications and Inversions based on a single scoring framework that uses both the variant and its flanking regions. To complement this framework, we also explore sequence-based annotations derived from SegmentNT, a deep learning model that provides functional predictions from DNA sequence at nucleotide resolution. Our analysis evaluated whether sequence-derived functional signals can provide additional information for SV prioritization and whether additional models with these features alone or in combination with previous coordinate-based annotations can be used.\ CADD-SV v2.0 outperforms its previous version and other tools in prioritizing deleterious variants across major SV types, including some previously unsupported, and substantially improves the computational workflow, increasing predictive power for genome-wide SV interpretation.

Read PDF

Similar papers

Open access Sep 2026

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller

Abstract Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning–based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

J. E. Perdomo, Mian Umair Ahsan, Jasmine Akoto et al. · 0 citations
Open access Aug 2026

A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models

It is demonstrated that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and it is proposed that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.

João Sartori, Ana Carolina Ramos Guimarães, Lucas de Almeida Machado · 0 citations
Open access Aug 2026

A high-resolution human pangenome structural variant resource for improved disease association

This work describes the full spectrum of genetic variation and shows that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions.

J. Lin, J. Gustafson, J. Wertz et al. · 0 citations
Open access Aug 2026

DNCLA: A Deep Learning Model for TFBS Identification Based on Structural and Conformational Properties of Nucleotides and Dinucleotides

Identifying transcription factor binding sites (TFBSs) is fundamental to understanding complex gene regulatory mechanisms and the functions of non-coding regions. Although existing methods have achieved substantial strides, capturing both local structural features and long-range spatial dependencies within DNA sequences remains a major challenge for improving prediction accuracy. In this study, we propose DNCLA, a deep learning model that synergizes multisize convolutional fusion, Bidirectional Long ShortTerm Memory (Bi-LSTM) networks, and a multi-head self-attention mechanism. At the feature extraction level, DNCLA breaks through the limitations of traditional single-sequence encoding by fusing Nucleotide Chemical Properties (NCP) with Dinucleotide Physicochemical Properties (DPCP). NCP provides a refined characterization of chemical differences between bases based on ring structures, hydrogen bond sites, and functional group properties, while DPCP introduces parameters such as local structural stability and geometric flexibility of the DNA. Subsequently, the model extracts spatial evolution from these high-dimensional features through a multi-size convolutional module; captures long-range spatial dependencies using Bi-LSTM layers; and employs a multi-head self-attention mechanism to achieve adaptive weight distribution of global features, thereby enhancing the perception of key regulatory motifs. Results from training and testing the proposed model on 165 ChIPseq datasets demonstrate that DNCLA possesses robust generalization capabilities and high predictive performance in TFBSs identification. This suggests that the incorporation of physicochemical features better elucidates the essence of interactions between transcription factors and DNA.

Jingjue Wei, Jie Feng · 0 citations
Feb 2025

Accurate de novo transcription unit annotation from run-on and sequencing data

Functional element annotations are critical tools used to provide insight into the molecular processes governing cell development, differentiation, and disease. Run-on and sequencing assays measure the production of nascent RNAs and can provide an effective data source for discovering functional elements. However, the accurate inference of functional elements from run-on sequencing data remains an open problem because the signal is noisy and challenging to model. Here we investigated computational approaches that convert run-on and sequencing data into annotations representing transcription units, including genes and non-coding RNAs. We developed a convolutional neural network, called convolutional discovery of gene anatomy using PRO-seq (CGAP), trained to identify different anatomical features of a transcription unit, which were then stitched together into transcript annotations using a hidden Markov model (HMM). Comparison with existing methods showed a significant performance improvement using our novel CGAP-HMM approach. We developed a voting system that ensembles the top three annotation strategies, resulting in large and significant improvements in transcription unit annotation accuracy over the best performing individual method. Finally, we also report a conditional generative adversarial network (cGAN) as a generative approach to transcription unit annotation that shows promise for further development. Collectively our work provides novel tools for de novo transcription unit annotation from run-on and sequencing data that are accurate enough to be useful in many applications.

Paul R. Munn, Jay Chia, Charles G. Danko · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.