Skip to content
Open access

Integration of Advanced Machine Learning and Statistical Methods for High-Dimensional Genomic Data Analysis

Jul 2026 · Precision Journal of Applied Mathematics and Statistics · 0 citations

Abstract

Introduction: The growing volume of high-throughput genomic data has enabled opportunities for cancer prognosis, patient stratification and biomarker identification. But gene expression data typically consist of thousands of molecular variables and relatively few patient observations, making the data analytically challenging in terms of dimensionality, redundancy, noise, overfitting, and interpretability. To overcome these problems this work proposes a hybrid statistical and machine-learning framework to analyze breast cancer gene-expression profiles. Methodology: The proposed workflow was tested on the dataset of Molecular Taxonomy of Breast Cancer International Consortium (METABRIC), that combines genomic measurements with related clinical data. The data preparation stages included missing-value treatment, standardisation of features, variance-based filtering, hypothesis-driven statistical screening based on t-tests and analysis of variance (ANOVA), correlation analysis, and dimensionality reduction using principal component analysis. After feature engineering, several predictive and exploratory algorithms were designed, such as: Random Forest, Multilayer Perceptron (MLP), Extreme Gradient Boosting (XGBoost), K-Means clustering, and Ensemble Learning. The accuracy, precision, recall, F1 score, Receiver Operating Characteristic Area Under the Curve (ROC-AUC), cross validation performance and silhouette coefficient measures were used to assess model effectiveness. Results: Experimental results showed that MLP classifier outperformed other classifiers in terms of prediction accuracy while ensemble model resulted in the best ROC-AUC value. Conclusion: The findings suggest that statistical feature selection combined with machine learning-based feature importance analysis can boost predictive power while simultaneously providing greater importance for biologically relevant features. The overall proposed framework offers a comprehensive strategy for genomic classification, prioritisation of candidates as biomarkers, and the creation of data-driven decision-support tools in the context of breast cancer research and precision oncology applications.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.