Skip to content

Cross-Modal Guidance Learning for Zero-Shot Industrial Anomaly Detection

2026 · IEEE Transactions on Automation Science and Engineering · Vol 23, pp. 14941-14954 · 0 citations · 44 references

Abstract

Zero-shot industrial anomaly detection (ZIAD) aims to develop a unified model capable of directly identifying unseen anomaly categories in images without requiring reference samples. Recently, large-scale Vision-Language Models (VLMs) such as CLIP have shown great potential for solving this task. However, existing methods typically rely on manual text prompts to guide VLMs in anomaly detection, which often fail to capture fine-grained semantic cues, leading to limited accuracy. To address the challenge, this paper proposes a novel Cross-Modal Guidance Learning (CMGL) framework for ZIAD. Instead of handcrafted textual prompts, CMGL introduces learnable prompting mechanism to fully exploit the collaborative guidance between visual and textual modalities for efficient unseen anomaly detection. Leveraging the frozen image encoder of pre-trained CLIP, the CMGL extracts multi-scale patch tokens and global tokens of the input image as visual representations. Then, informed by the cross-modal information, adaptive prompt vectors are constructed to obtain textual representations. In the process, a Learnable Context Block (LCBlock) and a Multi-Layer Perceptron (MLP) are introduced to extract holistic semantics and fine-grained details, and an Adaptive State Vector Module (ASVM) is designed to learn generalized normal and abnormal state vectors from extensive text descriptions. By aggregating the outputs of these components, textual representations of the image are acquired through the frozen text encoder. Finally, a Local-Global Token Integrator (LGTI) and an Uncertainty-Aware Anomaly Fusion Module (UAFM) are proposed to achieve anomaly recognition and localization through visual–textual alignment. Extensive experiments on multiple industrial datasets demonstrate the superiority of our method. Note to Practitioners—This paper presents a Cross-Modal Guidance Learning (CMGL) framework to address anomaly detection under the zero-shot setting. Unlike previous approaches that rely on manually crafted text prompts, the proposed CMGL derives task-relevant prompt cues from cross-modal data by the designed learnable prompting mechanism, guiding the model to automatically recognize and localize unseen anomaly categories without requiring any reference samples. Extensive experiments demonstrate the effectiveness and strong generalization capability of the proposed approach. Benefiting from these properties, our method provides a novel and effective ZIAD solution for identifying potential anomalies in real-world industrial scenarios where data distributions are uncertain or anomaly-related information cannot be clearly specified. Our project page is publicly available at https://aicoder12.github.io/CMGL/

View source

Similar papers

Conference Aug 2026

DPRF-CLIP:Dual-Path Residual Fusion for Zero-Shot Industrial Anomaly Detection

Zero-Shot Anomaly Detection (ZSAD) aims to accurately identify anomalous samples from unseen categories without relying on target class training data. In industrial quality inspection scenarios, collecting training samples for target defect categories is often impractical due to production constraints and data scarcity, and ZSAD methods can effectively address the challenge of reliable anomaly detection under limited data conditions. Recently, vision-language models have shown strong generalization and inherent zero-shot capabilities, greatly facilitating their wide application in zero-shot industrial anomaly detection tasks with competitive and reliable detection performance. However, they have critical practical limitations: insufficient attention to fine-grained image details and poor adaptability to the specific requirements of industrial anomaly detection tasks. To address these limitations, we propose DPRF-CLIP, a CLIP-based ZSAD framework. It uses a pre-trained ResNet network to extract fine-grained local image features, which are fused into CLIP’s visual encoder via a specially designed bidirectional cross-attention module. A feature enhancement module is also integrated to further strengthen the model’s ability in capturing fine-grained local visual patterns. Comprehensive experiments on real-world industrial benchmarks (MVTec AD, VisA) show DPRF-CLIP achieves competitive performances, validating its effectiveness and strong generalization in industrial anomaly detection.

Shuai Liu, Tian-Jiao Ma · 0 citations
Jul 2026

VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection

Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although recent CLIP-based methods have demonstrated promising generalization through vision-language alignment, they remain limited in capturing diverse anomaly semantics and subtle local variations. To address these limitations, we propose VFAD, a unified framework that combines variational semantic prompting with frequency-adaptive representation learning. Specifically, we introduce a Variational Semantic Prompt Extractor (VSPE), which adaptively aggregates anomaly-relevant local semantics from dense patch tokens and regularizes them through a variational information bottleneck, thereby incorporating fine-grained visual cues and enabling more precise cross-modal alignment. Furthermore, we develop a Frequency-Adaptive Representation Aggregation (FARA) module that leverages wavelet-based frequency decomposition and frequency-specific expert aggregation to enhance anomaly-discriminative visual representations. By jointly strengthening semantic guidance and visual representation learning, VFAD improves both anomaly discrimination and fine-grained localization. Extensive experiments on 13 industrial and medical benchmarks demonstrate that VFAD consistently outperforms existing state-of-the-art ZSAD methods across diverse anomaly scenarios. The code will be publicly available upon publication.

Peng Chen, Kaige Li, Wei Wang et al. · 0 citations
Open access Aug 2026

Multi-level visual-language models feature learning for generalizable anomaly detection

Zero-shot anomaly detection (ZSAD) aims to identify anomalies in target datasets without accessing their samples. Although CLIP and other large-scale vision language models show strong generalization, their potential for multi-level feature extraction in ZSAD remains underexplored. To address this, we propose a Multi-Level Feature Learning (MLFL) framework to enhance the zero-shot capability of CLIP via hierarchical alignment. MLFL adopts a two-stage training paradigm: Multi-Level Text Prompt Tuning (MLTP) and Multi-Level Text-Image Feature Alignment (MLFA). MLTP learns object-agnostic and object-aware prompts tailored to different encoder blocks. MLFA aligns textual and visual features using linear layers for shallow blocks and a Deep Feature Alignment (DFA) module for deep blocks. To compress parameters and preserve semantics, we introduce a Generalized Prompt Distillation (GPD) module that distills object-aware prompts into a unified representation. Experiments on seven industrial datasets achieve state-of-the-art performance, and deployment tests on edge devices demonstrate the potential applicability of the framework in practical industrial scenarios.

Jianfeng Qiu, Junfa Li, Juan Xie et al. · 0 citations
Preprint Aug 2026

Co-Evolutionary Prompt Optimization with Cross-Category Transfer for Zero-Shot Anomaly Detection

Zero-shot anomaly detection (ZSAD) has gained significant attention for its practical value in industrial inspection. Recently, CLIP-based approaches have been widely adopted in ZSAD due to their strong vision-language generalization capabilities. However, existing methods commonly employ continuous prompt embeddings for prompt optimization and encode semantics in latent vectors, which lack interpretability and scalability. To this end, we propose CoEvoAD, a co-evolutionary framework for discrete prompt selection. CoEvoAD performs prompt search in the discrete natural-language space using an evolutionary algorithm. Candidate prompts are iteratively generated, evaluated, and selected throughout population evolution, thus preserving the interpretability and composability of natural language. Furthermore, we introduce a Cross-Category Transfer Objective (CCTO), which treats held-out source categories as proxies for unseen categories and scores prompt rules based on their estimated cross-category transferability, effectively improving cross-category generalization. Extensive experiments are conducted to validate the effectiveness of CoEvoAD, and the results show that it achieves state-of-the-art performance across multiple anomaly detection datasets. The code is available at https://github.com/rstao-bjtu/CoEvoAD.

Si-Si Zhu, Chang-Wei Yu, Renshuai Tao et al. · 0 citations
Open access Aug 2026

Myriad: a large multimodal model applying vision experts for industrial anomaly detection

A novel large multimodal model applying vision experts for industrial anomaly detection (abbreviated as Myriad), which treats conventional IAD models as VEs and converts their anomaly maps into lightweight prompts that steer a frozen Q-Former toward suspicious regions, while a compact low-rank adapter shapes features for IAD.

Yuanze Li, Haolin Wang, Shihao Yuan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.