It is demonstrated that accuracy alone can be misleading for model evaluation and that integrating energy, runtime, and memory metrics enables more sustainable and resource-aware machine learning model selection.
Abstract
Machine learning should be judged by how well it predicts, and computational resources are not accounted for in predictive accuracy. Given the growing emphasis on energy consumption and resource efficiency, decision-supporting frameworks should go beyond accuracy. This study presents an energy-based benchmarking approach for supervised learning models. Ten classical algorithms were evaluated on three textual and tabular datasets. The energy consumption of preprocessing, training, and inference was monitored with Intel RAPL via pyRAPL along with the runtime, peak memory usage, and predictive performance statistics (accuracy, precision, recall, F1-score, and AUC). Experiments were conducted in a controlled CPU-based environment to ensure comparability. The computational role of this feature is found to be appreciably diverse. Results show that Random Forest achieved the highest overall balance between predictive performance and efficiency (CI = 0.950, PPI = 0.907), while Logistic Regression provided a competitive trade-off (CI = 0.905, EI = 0.998). Gaussian Naïve Bayes was the most energy-efficient model with a mean energy consumption of 127 J, whereas Support Vector Classifier (SVC) incurred the highest computational cost, consuming 45,758 J and requiring 3925 s on average. The Pareto analysis identified Random Forest, Logistic Regression, Passive Aggressive, and Decision Tree as non-dominated solutions. These findings demonstrate that accuracy alone can be misleading for model evaluation and that integrating energy, runtime, and memory metrics enables more sustainable and resource-aware machine learning model selection. The proposed framework provides practical guidance for Green AI, Tiny Machine Learning (TinyML), edge computing, and other resource-constrained deployment environments.
This study presents a comprehensive comparative evaluation of three widely used machine learning algorithms: Logistic Regression, Decision Tree, and Random Forest, for classification tasks, and demonstrates that Random Forest achieves superior generalization performance compared to the other models.
Badhur Ammulya, G. Nagalakshmi· International Journal of Lat...· 0 citations
This study proposes a hybrid XDSVM framework for multi-class student performance prediction by integrating a Support Vector Machine (SVM) and a compact Deep Neural Network (DNN). Boruta feature selection identifies 20 key predictors from engagement and interaction data, capturing over 70% normalized feature gain. The D...
Helal Masud, Billal Hossain, Asaduzzaman Asaduzzaman et al.· IEEE Jordan Conference on Ap...· 0 citations
Overall, logistic regression demonstrates strong and good generalization ability, while ensemble methods constitute robust alternatives for binary classification on structured data.
Zahra Benider, H. Bouzahir, Jaafar Idrais· EPJ Web of Conferences· 0 citations
The main conclusions obtained demonstrate that it is possible to build reliable models capable of maximizing fault detection and minimizing false alarms, which is vital for predictive maintenance systems in industries where accuracy is critical.
Juan Carlos Gonzalez Islas, Jesus Vargas-Ortiz, Aldo Márquez-Grajales et al.· ITEGAM- Journal of Engineeri...· 0 citations
Logistic Regression, Random Forest, soft voting, and weighted voting using a publicly available educational dataset comprising 4,424 student records, 36 predictors, and three outcome classes provided accurate, interpretable, and decision-oriented predictions, although external validation and prospective evaluation are...
Haleeful Jud· Journal of Intelligent Decis...· 0 citations
The widespread deployment of machine learning (ML) systems in critical domains has exposed the limitations of accuracy-centric evaluation, particularly under conditions involving class imbalance, noise, and distributional shifts. Existing studies frequently employ alternative metrics in isolation and lack a unified fra...
Ade Putra, E. Noche, Diksha D. Gabhane· Journal of Data Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.