Comparative Evaluation of Feature Selection Strategies and Machine Learning Classifiers for IoT Botnet Detection
Abstract
The rapid expansion of Internet of Things (IoT) devices has increased the need for accurate botnet detection methods that can operate with a compact set of network-traffic features. This study presents a controlled comparative evaluation of eight feature selection strategies and six machine learning classifiers using the N-BaIoT dataset. The feature selection strategies cover statistical, projection-based, recursive, hybrid, and ensemble approaches. Each method was evaluated at five feature-set sizes comprising 2, 3, 5, 10, and 15 features. Six classifiers—Decision Tree, Random Forest, Gradient Boosting, support vector machine with a radial basis function kernel, Logistic Regression, and K-Nearest Neighbours—were tested using fixed settings across all feature configurations. The complete design produced 240 model configurations that were assessed on the same evaluation set using accuracy, recall, precision, F1-score, training time, and inference latency. The results show that the RFE-Hybrid Lasso method produced the strongest overall feature subsets, with an average accuracy of 98.94% and an average F1-score of 98.09%. The best individual configuration combined RFE-Hybrid Lasso, 15 selected features, and Decision Tree, achieving 99.63% accuracy, 98.87% recall, 99.41% precision, and a 99.14% F1-score. Decision Tree also recorded the highest average classifier performance, reaching 99.18% accuracy with a training time of 2.1 minutes and an inference latency of 0.18 ms. The findings indicate that RFE-based hybrid selection improves detection performance over single statistical methods and that Decision Tree offers the most favourable performance–cost balance among the evaluated classifiers.