The potential of vision-language models (VLMs) such as CLIP for zero-shot anomaly detection (ZSAD) is constrained by an inherent semantic-localization dichotomy. While CLIP’s global features excel at image-level classification, they lack the spatial sensitivity required for pixel-level segmentation. Existing approaches attempt to alleviate this issue through prompt optimization, which introduces a trade-off between global semantic discrimination and local localization accuracy, thereby limiting cross-domain generalization capability. To resolve this, we introduce HD-CLIP, a framework that decouples these competing objectives. HD-CLIP employs dedicated pathways for classification and localization, guided by a hierarchical and dynamic prompting mechanism that provides multilevel, content-adaptive cues. A novel localization-distilled supervision (LDS) loss then unifies these pathways by creating a probabilistic bridge between spatial evidence and semantic judgment. Extensive experiments on 14 challenging datasets achieve state-of-the-art performance, confirming that HD-CLIP effectively bridges the semantic-localization divide and advances the potential of VLMs for robust ZSAD applications.
Jielin Jiang, Yunxiang Chang, Hao Yin et al.· IEEE Transactions on Instrum...· 0 citations
This work proposes Hint-Conditioned Generative Recommendation (HCGRec), a semantic-ID generative recommendation framework that recovers learning signal for such hard training instances and introduces hint-aware credit decomposition, using supervised learning to preserve item-semantic and prefix-structure alignment for hinted tokens and GRPO to optimize the sampled suffix.
Kangning Zhang, Hao Fang, Xukun Luo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.