Skip to content
Open access

Bootstrap-enhanced regularization addressing multicollinearity and skewness in high-dimensional immunophenotyping data

Aug 2026 · BMC Bioinformatics · 0 citations

TL;DR

The Bootstrap-Enhanced Regularization Method (BERM) is a robust approach for variable selection and coefficient estimation in complex biomedical datasets, achieving the highest overall balanced accuracy while maintaining competitive coefficient estimation performance across a range of simulated sparsity, noise, and dimensionality scenarios.

Abstract

Accurate identification and estimation of variables associated with outcomes or disease states are critical for advancing diagnosis, prognosis, and precision medicine in biomedical research. Regularized regression techniques, such as lasso, are widely employed to enhance interpretability by reducing model complexity and identifying significant variables. However, these methods face two major challenges: (1) the exclusion of important variables due to high correlation with included predictors, and (2) the presence of skewness in human biomedical datasets, which violates key statistical assumptions. Current approaches that fail to address these issues simultaneously may lead to biased interpretations and unreliable coefficient estimates. To overcome these limitations, we propose an enhanced two-step approach, the Bootstrap-Enhanced Regularization Method (BERM). BERM outperformed existing regularization methods in variable selection, achieving the highest overall balanced accuracy while maintaining competitive coefficient estimation performance across a range of simulated sparsity, noise, and dimensionality scenarios. We further demonstrated the effectiveness of BERM by applying it to a human immunophenotyping dataset to identify important immune parameters in the autoimmune disease, type 1 diabetes. BERM is a robust approach for variable selection and coefficient estimation in complex biomedical datasets. Its consistent performance across a wide range of data conditions supports more reliable identification of important variables. An open-source implementation of BERM is available as an R package on GitHub ( https://github.com/xiaorudong/berm ).

Read PDF

Similar papers

Open access Jul 2026

Revisiting Logistic Regression for High-Dimensional Gene Expression Data

A reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings is proposed and several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model.

Rossana O. Souza, W. Rodrigues, Bráulio Couto et al. · 0 citations
Open access Aug 2026

A simulation-based comparison of Boruta, LASSO, and Elastic Net for variable selection in logistic regression, with an ovarian cancer miRNA application

Three widely used variable selection approaches are compared under a range of data-generating conditions and to illustrate their performance using an ovarian cancer miRNA dataset, finding Boruta offers a favorable balance between sensitivity and specificity in highly correlated settings and LASSO provides stricter fals...

Reza Arabi Belaghi, Hulya Yurekli, Farzaneh Hamidi et al. · 0 citations
Aug 2026

Beyond predictive performance: Interpretability challenges and feature importance bias in XGBoost-based readmission models.

Song et al. report a machine-learning framework based on the eXtreme Gradient Boosting (XGBoost) algorithm for predicting 1-year unplanned readmissions among elderly patients with coronary heart disease (CHD). This commentary examines critical limitations in the interpretability and methodological robustness of such mo...

S. Oka, Maito Suzuki, Y. Takefuji · 0 citations
Jul 2026

Assessing penalized approaches for estimating causal treatment effects under extremely limited overlap in oncology.

BACKGROUND Overlap weighting (OW) is increasingly used to estimate treatment effects in observational cancer studies. OW has attractive features: it targets the clinical equipoise population and mitigates the influence of extreme propensity score (PS) weights. Additionally, under regularity conditions, when the PS mode...

Sangwon Lee, H. Chae, Dong-woo Choi et al. · 0 citations
Open access Aug 2026

metadeconfoundR: Covariate analysis of high-dimensional cross-sectional omics data

This work benchmarked metadeconfoundR against state-of-the-art methods for identifying biomarkers using simulated ground truth derived from microbiome data, and demonstrated its ability to disentangle confounding effects while preserving statistical power.

T. Birkner, Chia-Yu Chen, Morgan Essex et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.