A reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings is proposed and several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model.
Abstract
Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings. The method is evaluated on three biomedical datasets, including scenarios with tens of thousands of attributes and substantially fewer samples. Across these datasets, the proposed approach achieved clear separation between control and disease groups while selecting a compact set of features. Several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model, while additional selected features suggest potential novel candidates for further investigation. These results indicate that the proposed framework may provide an interpretable and computationally efficient alternative for feature selection in high-dimensional computational biology applications.
Overall, GeDiRep provides a structured and interpretable framework for high-dimensional gene expression analysis by selecting reduced yet discriminative and biologically meaningful gene subsets.
Cihan Kuzudisli, B. Qaqish, Burcu Bakir-Gungor et al.· Mathematics· 0 citations
Three widely used variable selection approaches are compared under a range of data-generating conditions and to illustrate their performance using an ovarian cancer miRNA dataset, finding Boruta offers a favorable balance between sensitivity and specificity in highly correlated settings and LASSO provides stricter fals...
Reza Arabi Belaghi, Hulya Yurekli, Farzaneh Hamidi et al.· BMC Medical Research Methodo...· 0 citations
The Bootstrap-Enhanced Regularization Method (BERM) is a robust approach for variable selection and coefficient estimation in complex biomedical datasets, achieving the highest overall balanced accuracy while maintaining competitive coefficient estimation performance across a range of simulated sparsity, noise, and dim...
INTRODUCTION
Gene feature selection is essential in bioinformatics and medical research, as it identifies gene subsets closely associated with specific diseases or biological traits from high-dimensional gene datasets. High-dimensional gene data causes the curse of dimensionality, leading to sparsity, complex inter-fea...
Zhilin Wangy, Wei-Ping Ding, Jin-Quan Zhang et al.· Journal of Advanced Research· 0 citations
The findings suggest that statistical feature selection combined with machine learning-based feature importance analysis can boost predictive power while simultaneously providing greater importance for biologically relevant features.
Z. Khan, M. Sohail, Habiba Mehak· Precision Journal of Applied...· 0 citations
The newly introduced RBAs were among the strongest performing, and by robustly retaining both main effects and 2-way epistatic interactions, these algorithms preserve predictive signals for downstream modeling.
Kia Kazemi-Nia, H. Bandhey, P. Freda et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.