Explainable Domain-Adaptive CNN–Transformer for Bidirectional Cross-Domain Bearing Fault Diagnosis
Abstract
Industry 5.0 requires resilient, adaptive, and trustworthy manufacturing systems capable of maintaining reliable diagnostic performance across heterogeneous industrial environments. However, data-driven fault diagnosis models often experience substantial performance degradation when transferred across machines, operating conditions, and data acquisition platforms because of domain distribution shifts. This study proposes an explainable domain-adaptive CNN–Transformer framework for unsupervised bidirectional cross-domain bearing fault diagnosis using the Case Western Reserve University (CWRU) and Paderborn University (PU) datasets. The framework integrates one-dimensional convolutional layers for extracting local high-frequency vibration patterns, Transformer encoders for modelling long-range temporal dependencies, and Maximum Mean Discrepancy (MMD)-based feature-distribution alignment for learning transferable domain-invariant representations. Under the Unsupervised Domain Adaptation (UDA) protocol, the source domain supplies labelled samples for classification learning, whereas the target domain contributes unlabelled features only for MMD-based alignment; target labels are withheld from training and model selection and are used only for final evaluation. Conventional 1D-CNN and bidirectional long short-term memory baselines achieve over 90% accuracy in-domain but fall to 82.14% and 79.88%, respectively, for CWRU→PU, corresponding to domain-drop magnitudes of 12.07 and 12.99 percentage points. The proposed framework achieves 98.63% and 96.82% in-domain accuracy on CWRU and PU, respectively, and 92.46% for CWRU→PU and 94.18% for PU→CWRU, with domain-drop magnitudes of 6.17 and 2.64 percentage points. Ablation results confirm the complementary contributions of convolutional feature extraction, Transformer-based temporal modelling, and domain alignment. Furthermore, attention, saliency, and feature-importance analyses show that the model focuses on fault-relevant vibration regions and informative diagnostic characteristics, including kurtosis, root-mean-square (RMS), and crest factor, improving prediction transparency. These findings support accurate, transferable, and interpretable vibration-based condition monitoring across heterogeneous bearing datasets.