Advanced Machine Learning Approaches for Predictive Analytics: Comparative Evaluation of Classification Performance in High-Dimensional Data
Abstract
This study examined advanced machine learning approaches for predictive analytics by comparatively evaluating the classification performance of Support Vector Machine (SVM), Random Forest, and Gradient Boosting in high-dimensional data. A high-dimensional dataset containing multiple observations and a large number of features was used for the analysis. Data were collected from a structured secondary dataset and prepared through systematic data preprocessing, including data cleaning, handling of missing values, removal of duplicate records, feature preparation, and appropriate transformation of variables. Feature selection and dimensionality reduction were applied to address irrelevant and redundant features and reduce the complexity of the high-dimensional feature space. The dataset was divided into training and testing sets, and SVM, Random Forest, and Gradient Boosting models were trained and optimized using cross-validation and hyperparameter tuning. Classification performance was evaluated using accuracy, precision, recall, F1-score, and ROC-AUC, while comparative analysis was conducted to examine model performance before and after feature selection and dimensionality reduction. The data analysis was performed using Python with Pandas, NumPy, and Scikit-learn, while Matplotlib and Seaborn were used for data visualization. The study provides a systematic framework for evaluating different machine learning approaches under high-dimensional conditions and examining the contribution of feature management to classification performance. The significance of the study lies in supporting researchers and data analysts in selecting appropriate machine learning models and preprocessing strategies for complex high-dimensional datasets, contributing to more efficient, reliable, and evidence-based predictive analytics.