Multi-Method Explainable Deep Learning for Acute Stress Analysis using Physiological Signals
Abstract
Wearable devices measure body signals such as skin conductance, pulse, skin temperature and movement, and machine-learning models can use these signals to detect acute stress. Recent work has added explainability so that a user can see which signal influenced a decision. However, most studies use only a single explanation method, judge that explanation visually, and rarely report how confident the model is. This raises a practical question: can these explanations be trusted? This paper studies that question on the WESAD dataset under strict subject-independent (Leave-One-Subject-Out) evaluation. Five deep models (LSTM, BiLSTM, GRU, CNN-LSTM, Transformer) and five classical models are trained on identical folds. Explanations are generated with three methods (Integrated Gradients, SHAP, LIME) and their agreement is measured across six held-out participants; explanations are further examined with faithfulness tests, cross-architecture comparison, counterfactual and temporal analysis; and predictive uncertainty is estimated with deep ensembles and assessed for calibration. Three findings are reported. First, the five deep architectures perform almost identically (0.846–0.863 accuracy) and are statistically indistinguishable (all Holm-corrected p = 1.00, small effect sizes), while a classical Random Forest is more accurate (0.911); after Holm–Bonferroni correction no individual classical-vs-deep pair is significant at 15 subjects, but the direction is consistent across every comparison, with medium-to-large effect sizes for nine of the ten comparisons (rank-biserial 0.50–0.68) and a smaller effect for the tenth (Logistic Regression vs Transformer, +0.25), so at this data scale increasing architectural complexity is not the productive direction. Second, across six subjects the three explanation methods agree only partially, heart rate is the single most important channel in two to four of six subjects depending on the method, and pairwise method agreement is low-to-moderate and highly variable (Spearman ρ from 0.10 ± 0.45 to 0.70 ± 0.20), which shows that a single explanation method is insufficient on its own and that reliability must be measured rather than assumed; faithfulness is in the expected direction (mean insertion AUC 0.794 > deletion AUC 0.756) but modest and subject-dependent. Third, the ensemble is well calibrated (Expected Calibration Error 0.065) and, under a coverage–risk (deferral) analysis, answering only the most-confident 50% of cases raises selective accuracy from 0.884 to 0.960. The study concludes that trustworthiness, measured explanation reliability together with calibration, rather than accuracy alone, is the appropriate objective for wearable stress detection.