XAI-AutoDL: An Explainable Automated System for Deep Learning-Based Classification Using Meta-Features
Abstract
The increasing use of machine learning (AutoML) and deep learning means hyperparameter optimization can save time. Nevertheless, it is not clear how they compare on structured tabular classification under tight computational constraints. Here we present a resource-aware benchmark that compares the performance of Random Forests, FLAML AutoML, and an Optuna-tuned deep learning pipeline based on PyTorch Multi-Layer Perceptron (MLP). This benchmark uses stratified 60/20/20 train-validation-test splits and a 60-second optimization budget for the automated methods. All methods share the same preprocessing pipeline. Accuracy, F1-score, Precision, and Recall are evaluated using Friedman and pairwise Wilcoxon signed-rank tests. With mean Accuracy, F1-score, Precision, and Recall of 0.8587, 0.8521, 0.8541, and 0.8587 respectively, FLAML achieved the best average performance over all four metrics. However, Friedman testing indicates no significant differences across these approaches at 0.05. The SHAP and LIME methods were applied on the Breast Cancer Wisconsin dataset as part of the aggregate benchmarking. The results indicate that the structure of the training exercise had an effect on model performance.