Jul 2026· Jambura Journal of Electrical and Electronics Engineering· 0 citations
TL;DR
An unsupervised hybrid deep learning framework for unknown bearing fault diagnosis and severity assessment using vibration signals that combines Continuous Wavelet Transform, Convolutional Neural Networks, and Long Short-Term Memory autoencoders is presented.
Abstract
This paper presents an unsupervised hybrid deep learning framework for unknown bearing fault diagnosis and severity assessment using vibration signals. The proposed framework combines Continuous Wavelet Transform (CWT), Convolutional Neural Networks (CNN), and Long Short-Term Memory (LSTM) autoencoders. Trained exclusively on 3,779 healthy data segments, the model detects anomalies via reconstruction error analysis evaluated across a total of 7,585 test segments. Experiments on the CWRU dataset show that the proposed hybrid model achieves competitive performance (AUC 0.990, accuracy 90.42%, and F1-score 89.57%) compared to spectral baselines, while uniquely preserving temporal dynamics—a critical advantage for non-stationary industrial environments. However, outer race faults were not reliably detected under the global threshold, which we report as a key limitation. A severity assessment and Health Index are also introduced for interpretable predictive maintenance.
Reliable bearing fault detection is essential for predictive maintenance in industrial systems; however, obtaining labelled fault data is often expensive, time-consuming, and impractical in real-world deployments. To address this challenge, this study proposes a healthy-only self-supervised anomaly detection framework for bearing health monitoring using vibration measurements. The proposed approach combines convolutional neural networks and Transformer-based temporal modelling to learn informative representations from healthy vibration signals without requiring fault labels during representation learning. Three self-supervised learning strategies—reconstruction-based, contrastive, and a unified contrastive–reconstruction objective—are investigated to evaluate the effectiveness of different representation learning approaches. The learned latent representations are subsequently analysed using Isolation Forest and Mahalanobis-distance anomaly scoring methods. To provide a realistic assessment of generalisation, a strict grouped cross-validation protocol is employed, where data are partitioned at the sample level to prevent information leakage between training and testing sets. Furthermore, prevalence-aware experiments are conducted under 5% and 10% fault prevalence scenarios to assess deployment robustness. Experimental results on the Paderborn bearing dataset demonstrate that the proposed CNN + Transformer model trained with combined contrastive and reconstruction objectives and evaluated using Isolation Forest achieves the best overall performance, obtaining a ROC-AUC of 0.878±0.015, a PR-AUC of 0.958±0.005, and an F1-score of 0.590±0.040. The results consistently outperform classical feature-based approaches, One-Class SVM, and autoencoder baselines. Ablation analysis further shows that combining contrastive and reconstruction objectives produces more informative representations than either objective alone. The findings demonstrate that the proposed healthy-only self-supervised framework provides an effective and label-efficient approach for rolling bearing anomaly detection and shows promise for predictive maintenance applications where labelled fault data are limited or unavailable.
Syed Sajjad Haider Zaidi, Alex Shenfield, Hongwei Zhang et al.· Processes· 0 citations
In this study, a hybrid deep learning architecture is proposed for robust vibration-based fault diagnosis in industrial machinery by jointly modeling time-domain, frequency-domain, and temporal dynamics. In the time domain, convolutional layers combined with an Enhanced Gated Attention (EGA) mechanism emphasize informative signal components while suppressing noise. Temporal evolution is modeled using Neural Ordinary Differential Equations (Neural ODEs), enabling smooth and stable continuous-time feature representations. In parallel, a Fourier Neural Operator (FNO) extracts frequency-domain characteristics, augmented with gated attention to focus on fault-related spectral patterns. Long Short-Term Memory (LSTM) layers capture long-range dependencies, while Squeeze-and-Excitation (SE) blocks adaptively recalibrate channel-wise feature responses. A multi-scale attention-based fusion module integrates domain-specific representations and auxiliary features to enhance discrimination under varying operating conditions. The proposed model is evaluated on the SUBFv1.0 dataset through extensive ablation studies and experiments under multiple noise levels, achieving 98.41% accuracy in noise-free conditions and maintaining performance above 91% even at 5 dB SNR. Unlike existing multi-path approaches that combine heterogeneous features in a loosely coupled or discrete manner, the proposed architecture uniquely integrates multi-domain feature learning with bidirectional attention mechanisms and continuous-time temporal dynamics, enabling coherent cross-domain interaction and robust fault characterization.
Canan Taştimur· Information Technology and C...· 0 citations
To address the degradation of cross-condition diagnostic performance caused by feature-scale drift in rolling bearing vibration signals under variable operating conditions, this paper proposes a spectral-guided adaptive multi-scale convolutional neural network (SAMACNN). First, PSD sequences and time-frequency features are introduced as dual-stream inputs. While the time-frequency main branch extracts local information, the Spectral Transformer spectral bypass branch captures long-range dependencies in harmonic structures. Second, dynamic gating weights are generated for the multi-scale convolutional branches, enabling sample-conditioned multi-scale feature selection and fusion and alleviating scale mismatch caused by fixed receptive fields and static fusion. Finally, data collected from two bearing fault simulation test rigs are used to verify the effectiveness and superiority of the proposed algorithm. The experimental results show that the proposed SAMACNN method achieves average accuracies of 97.16% and 97.56% on the two datasets, respectively, outperforming the ablation variants and demonstrating strong robustness and generalization capability in complex variable-condition measurement environments.
DaXin Li, Wang Hong, Hai Xue et al.· Engineering Research Express· 0 citations
Early detection of bearing faults in rotating machinery is essential for predictive maintenance. Although deep learning-based methods have achieved strong results in fault diagnosis, they usually require large amounts of labeled data. In industrial settings, however, faulty samples are limited, which restricts the applicability of fully supervised approaches. In this study, a bearing fault diagnosis framework based on unsupervised representation learning is proposed for limited-label scenarios. Firstly, a convolutional autoencoder is trained on raw vibration signals without using labels to learn informative latent representations. After that, these learned representations are classified using only a limited number of labeled samples. The proposed method is evaluated on the CWRU bearing dataset under same-load and cross-load settings with both single and dual-channel inputs. Experimental results show that the proposed framework achieves strong performance under low-label conditions and that the dual-channel setup further improves classification performance.
Ahmet Kaplan, Kürşat İnce, Murat Beken· Signal Processing and Commun...· 0 citations
Deep learning models for bearing fault diagnosis on the Case Western Reserve University (CWRU) dataset routinely achieve near-perfect accuracy. Yet their decisions remain largely opaque at the signal level, so engineers cannot determine where in the raw vibration waveform the model focuses. This study aims to bridge the interpretability gap in deep learning-based bearing fault diagnosis by developing a signal-level explainability framework for the Wide Deep Convolutional Neural Network (WDCNN) on the Case Western Reserve University (CWRU) dataset. SHAP DeepExplainer is applied directly to raw 2,048-point vibration segments to produce per-sample-point attribution maps; Fault Signature Maps (FSMs) are formalized as class-averaged SHAP fingerprints in three variants (signed, absolute, and variance) and validated via discriminability index, severity monotonicity, and split-half stability, complemented by a three-variant ablation study examining how architectural decisions affect both accuracy and explainability. On the 10-class CWRU dataset under a rigorous temporal split (1,240/310/750 samples), WDCNN achieves 99.87% test accuracy with a macro F1-score of 0.997. FSMs demonstrate high reproducibility (split-half stability = 0.940); severity monotonicity ranges from 8.6% to 17.6%; and ablation reveals that removing Batch Normalization increases FSM discriminability by 2.3× at a 3.6% accuracy cost. As the first work applying SHAP DeepExplainer at full 2,048-point raw signal resolution for WDCNN, this study establishes that fault discrimination relies on transient impulse morphology rather than bearing characteristic frequencies, a finding invisible to feature-level XAI, and introduces a multi-resolution diagnostic framework bridging deep learning accuracy with physically interpretable vibration analysis for trustworthy deployment in safety-critical industrial environments.
T. Suharto, Kadarsah Suryadi, B. Iskandar et al.· Emerging Science Journal· 0 citations