Skip to content
Open access

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting

Aug 2026 · Frontiers in Digital Health · Vol 8 · 0 citations · 41 references
Medicine

TL;DR

A probabilistic gradient boosting model on variant pathogenicity prediction that applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making is presented.

Abstract

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

Read PDF

Similar papers

Review Open access Aug 2026

A machine learning framework for predictive interpretation of variants of uncertain significance in hereditary cancer

This reproducible pipeline provides a clinically grounded computational approach to VUS triaging in precision oncology, with external validation supporting its generalizability to independent hereditary cancer gene datasets.

Nayeema Nizamuddin, Soham Biswas, Akshaykumar Zawar et al. · 0 citations

Explainable AI for analyzing cancer outcomes using large-scale genome sequencing data

A multi-tier, explainable AI framework designed to risk-stratify patients and predict overall survival using clinical and genomic covariates is developed and demonstrates that explainable machine learning models can robustly predict survivability and highlight actionable features for oncology dashboards.

P. Nalela · 0 citations
#machine learning Preprint Sep 2026

Interpretable Multi-Instance Learning Enables Early Prediction of Key Molecular Alterations from Routine Flow Cytometry in Acute Myeloid Leukemia

Background: Molecular testing for NPM1 and FLT3-ITD mutations guides critical early treatment decisions in acute myeloid leukemia (AML), but results can take weeks, long after these decisions must be made. Flow cytometry, already performed within hours of admission as part of routine care, may carry enough signal to pr...

Jonathan Legrand, Aguirre Mimoun, B. D. de Senneville et al. · 0 citations
Open access Aug 2026

Machine Learning approaches for the detection of disease-causing variants in whole-genome data need to address the expression of functional genes

A benchmark is constructed that includes real data and synthetic data with known generating mechanisms and various dataset sizes and levels of noise, and shows that it is necessary to take into account the expression of functional genes in order to successfully predict disease.

Camilla Mapstone, Julia Handl, David Talavera · 0 citations
#explainable ai Open access Aug 2026

aiDIVA – hybrid AI for rare disease diagnostics using evidence-based, machine learning and language models

aiDIVA is presented, an ensemble-AI combining statistical and machine learning models trained on genomic and phenotypic data to identify causal variants among tens of thousands per patient, and applies a random forest model to classify pathogenicity and generates evidence-based scores for dominant and recessive disease...

D. Boceck, L. Laugwitz, M. Sturm et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.