Knowledge-Driven Feature Selection with the Grouping–Scoring–Modeling Framework for Biomarker Discovery in High-Dimensional Transcriptomic Data
Biomarker discovery from high-dimensional transcriptomic data is frequently hindered by the “curse of dimensionality” and model selection bias. To address this, we propose the Grouping–Scoring–Modeling (G-S-M) framework, a knowledge-driven pipeline that anchors feature selection in established disease–gene associations. G-S-M operates within a 100-iteration Monte Carlo ensemble architecture utilizing internal cross-validation to ensure unbiased evaluation. We evaluated this framework on seven cancer datasets, where it demonstrated robust discrimination with an overall mean F1-score of 0.84 across all datasets. The framework achieved the strongest performance on Acute Myeloid Leukemia (mean F1 = 0.99, AUC-ROC = 1.00) and maintained competitive accuracy even on challenging cohorts, while producing biologically interpretable gene panels traceable to named disease associations. Permutation tests (10,000 iterations) confirmed statistically significant disease–gene enrichment (p < 0.0001) in five of seven datasets, and independent protein interaction network analyses demonstrated significant enrichment of the selected features. Released as an open-source software suite with interactive interfaces, G-S-M provides a reproducible computational framework for candidate biomarker discovery.