Gene expression datasets provide valuable information for disease classification and biomarker discovery; however, their high dimensionality and limited sample size may limit classification performance and reduce biological interpretability. This study proposes GeDiRep, a prior knowledge-guided framework for identifying compact and informative gene subsets. The method first organizes filtered genes into disease-associated groups using curated gene–disease associations from DisGeNET. Each group is then scored according to its predictive contribution, and representative genes are selected using Random Forest-based feature importance. Representative genes from the top-ranked groups are progressively accumulated, and the resulting gene subsets are used to assess classification performance on the test set. Experiments on eight microarray datasets showed that GeDiRep reduced the average number of selected features from 40.3 to 8.3 compared with G-S-M while improving the average AUC from 0.83 to 0.87. In comparison with traditional feature selection methods using the same number of genes, GeDiRep also achieved competitive AUC values. Biological analyses, including term–gene network, hub gene, and heatmap analyses, supported the functional relevance and stability of several selected genes. Overall, GeDiRep provides a structured and interpretable framework for high-dimensional gene expression analysis by selecting reduced yet discriminative and biologically meaningful gene subsets.
Cihan Kuzudisli, B. Qaqish, Burcu Bakir-Gungor et al.· Mathematics· 0 citations
Biomarker discovery from high-dimensional transcriptomic data is frequently hindered by the “curse of dimensionality” and model selection bias. To address this, we propose the Grouping–Scoring–Modeling (G-S-M) framework, a knowledge-driven pipeline that anchors feature selection in established disease–gene associations. G-S-M operates within a 100-iteration Monte Carlo ensemble architecture utilizing internal cross-validation to ensure unbiased evaluation. We evaluated this framework on seven cancer datasets, where it demonstrated robust discrimination with an overall mean F1-score of 0.84 across all datasets. The framework achieved the strongest performance on Acute Myeloid Leukemia (mean F1 = 0.99, AUC-ROC = 1.00) and maintained competitive accuracy even on challenging cohorts, while producing biologically interpretable gene panels traceable to named disease associations. Permutation tests (10,000 iterations) confirmed statistically significant disease–gene enrichment (p < 0.0001) in five of seven datasets, and independent protein interaction network analyses demonstrated significant enrichment of the selected features. Released as an open-source software suite with interactive interfaces, G-S-M provides a reproducible computational framework for candidate biomarker discovery.
Malik Yousef, Jens Allmer, Yasin Inal et al.· Applied Sciences· 0 citations
Portfolio construction aims to balance expected return and risk through effective asset allocation. This study proposes a portfolio formation framework that integrates machine learning-based return prediction with Markowitz mean–variance portfolio optimization. Random Forest, XGBoost, Multilayer Perceptron, and Support Vector Regression models are employed to predict the cross-sectional excess returns of stocks using financial indicators derived from technical and macroeconomic variables. These predictions are incorporated into the portfolio optimization process to determine portfolio weights. The resulting strategies are evaluated against benchmark portfolios including an equal-weighted portfolio and the BIST 100 index. Empirical results show that machine learning-based stock selection improves portfolio performance. In particular, the XGBoost-based portfolio achieves the best results with an annual return of 75.30% and a Sharpe ratio of 1.80. SHAP analysis further indicates that momentum and price-based technical indicators play a dominant role in model predictions.
Mustafa Etcil, Hüseyin Akkaş, Burak Kolukısa et al.· Signal Processing and Commun...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.