Skip to content
Review Open access

A machine learning framework for predictive interpretation of variants of uncertain significance in hereditary cancer

Aug 2026 · Frontiers in Systems Biology · Vol 6 · 0 citations · 34 references
Medicine

TL;DR

This reproducible pipeline provides a clinically grounded computational approach to VUS triaging in precision oncology, with external validation supporting its generalizability to independent hereditary cancer gene datasets.

Abstract

Introduction Variant interpretation remains a major bottleneck in clinical genomics, with variants of uncertain significance (VUS) representing a critical unresolved challenge due to insufficient evidence for definitive classification. Existing in silico tools exhibit variable and often inconsistent performance complicating clinical decision-making, particularly in the context of hereditary cancer genomics. Methods In this study, we developed a machine learning framework trained on 1,04,646 high-confidence ClinVar germline variants (3-star+ review status) annotated with Ensembl VEP (v114, GRCh38) and CADD v1.6 pathogenicity scores to classify variants as Pathogenic or Benign, subsequently applying the trained model to reclassify 40894 ClinVar VUS. Train/test partitioning was performed at the variant level (80/20 split) to prevent data leakage, with hyperparameter optimization via GridSearchCV and performance assessed by 10-fold cross-validation. Four classifiers were evaluated viz. Logistic Regression, Support Vector Machine, Random Forest and XGBoost, with Random Forest achieving the highest performance (AUC-ROC = 0.9995, 95% CI: 0.9993–0.9997; 10-fold CV AUC = 0.9992 ± 0.0004). Probability thresholds of P ≥ 0.80 (Pathogenic) and P <= 0.20 (Benign) were derived from Precision-Recall curve analysis, achieving empirically validated precision of 99.63% and 99.77% respectively on held-out test variants. Results and Discussion Applied to 40,894 ClinVar VUS, the model reclassified 19393 (47.4%) as Likely Pathogenic and 8,957 (21.9%) as Likely Benign, while 12,544 (30.7%) were conservatively retained as uncertain. External validation on 7,462 ENIGMA-classified BRCA1/BRCA2 variants from the BRCA Exchange database, completely independent of the ClinVar training data demonstrated an overall concordance of 98.83% (AUC = 1.0000). Further validation of VUS reclassification against 671 variants classified as VUS in ClinVar but definitively classified by ENIGMA yielded an overall concordance of 89.57% (Pathogenic: 96.4%, Benign: 87.4%). SHAP-based explainability analysis confirmed that predictions were predominantly driven by biologically interpretable features, including CADD Phred score, VEP functional impact tier, variant consequence class and population allele frequency, consistent with ACMG/AMP evidence criteria. This reproducible pipeline provides a clinically grounded computational approach to VUS triaging in precision oncology, with external validation supporting its generalizability to independent hereditary cancer gene datasets.

Read PDF

Similar papers

Open access Aug 2026

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting

A probabilistic gradient boosting model on variant pathogenicity prediction that applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making is presented.

Karthik V, S. Prejesh, Sumedh Deepak Kudale et al. · 0 citations
Open access Aug 2026

Machine Learning Triage of LDLR Variants of Uncertain Significance Using Predictor Concordance and ACMG-Aligned Evidence Mapping

Background: Variants of uncertain significance (VUS) in the LDLR gene remain a major barrier to the molecular diagnosis of familial hypercholesterolemia. Although computational approaches offer scalable prioritization, their clinical utility is limited by predictor discordance, incomplete annotations, and inflated perf...

BalaSubramani Gattu Linga, Faisal E. Ibrahim, Nader I. Al-Dewik · 0 citations
#machine learning Preprint Sep 2026

Interpretable Multi-Instance Learning Enables Early Prediction of Key Molecular Alterations from Routine Flow Cytometry in Acute Myeloid Leukemia

Background: Molecular testing for NPM1 and FLT3-ITD mutations guides critical early treatment decisions in acute myeloid leukemia (AML), but results can take weeks, long after these decisions must be made. Flow cytometry, already performed within hours of admission as part of routine care, may carry enough signal to pr...

Jonathan Legrand, Aguirre Mimoun, B. D. de Senneville et al. · 0 citations
Open access Sep 2026

Explainable machine learning for breast cancer prediction in resource-constrained settings: A multi-algorithmic framework integrating shap-based transparency with clinical decision support

Breast cancer remains the most commonly diagnosed malignancy among women globally, with disproportionately higher mortality rates in low- and middle-income countries (LMICs) where diagnostic delays and limited specialist pathology capacity are widespread. While machine learning (ML) approaches achieve strong predictive...

Oluwaseun Adebayo Bamodu, Sumaiya Nezam, C. Chung · 0 citations
Conference Aug 2026

Gradient Boosting Techniques in a Risk-Aware and Explainable Machine Learning Framework for Heart Disease Prediction

The complicated connection between medical risk factors and the serious ramification of misdiagnosis highlights the essential challenge of detecting cardiovascular disease in its early stages. Although most examinations to date have focused on accuracy-centric evaluation, which may not absolutely account for clinical s...

H. Suresh, P. R. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.