Skip to content

VCF-CLIP: Visual Context-Driven Fine-Grained Prompt Learning for Zero-Shot Anomaly Detection.

Jul 2026 · IEEE Transactions on Neural Networks and Learning Systems · Vol PP · 0 citations
Medicine

Abstract

Benefiting from recent advances in vision-language models (VLMs), numerous CLIP-based zero-shot anomaly detection (ZSAD) methods have been proposed to address the cold-start problem. Despite their impressive performance, these methods still depend on manual prompt engineering, and their coarse-grained text prompts struggle to capture the diverse patterns of anomalies, resulting in suboptimal visual-text alignment. To overcome these limitations, we propose VCF-CLIP, a visual context-driven fine-grained prompt learning framework built upon CLIP. The novelties of VCF-CLIP lie in two main aspects. First, we propose the prompt prototype learning (PPL) strategy, which learns a pair of unified prompt prototypes representing general normal and anomalous states in a loss-guided manner, thereby eliminating the need for manual prompt design. Second, we propose a lightweight prompt refinement adapter that dynamically aggregates multiscale and multilevel visual features to iteratively refine the prompt prototypes, enabling the generation of instance-specific prompts enriched with fine-grained information. We conduct extensive experiments on 14 benchmarks across industrial and medical domains, and show that VCF-CLIP outperforms existing state-of-the-art ZSAD methods.

View source