A vision–language-guided multimodal framework with dynamic prototype memory for small-sample fault diagnosis
Abstract
Rotating machinery fault diagnosis under limited training samples and varying operating conditions remains challenging, as data scarcity and domain shift often weaken the generalisation capability of deep learning models. Recent advances in large-scale pretrained foundation models have provided transferable priors for data-limited scenarios. Meanwhile, time-series signals transformed into visual representations preserve structured patterns that are well suited to visual modelling. Motivated by these advantages, this study proposes a vision–language (V–L)-guided multimodal framework with dynamic prototype memory for small-sample fault diagnosis of rotating machinery. The proposed framework combines semantic priors from a pretrained V–L model with temporal representation learning, cross-modal alignment and fusion, and an adaptive classification strategy to improve fault representation under limited-data conditions. Experiments on the Case Western Reserve University and Paderborn University bearing datasets under multiple small-sample training ratios and cross-condition transfer tasks demonstrate that the proposed framework achieves improved diagnostic performance over several representative deep learning baselines. The results indicate that integrating V–L semantic priors with stable temporal representation learning and cross-modal alignment provides an effective solution for fault diagnosis under limited-sample and varying-condition settings.