HD-CLIP: Hierarchical Dynamic Prompting and Decoupled Learning for Zero-Shot Anomaly Detection
The potential of vision-language models (VLMs) such as CLIP for zero-shot anomaly detection (ZSAD) is constrained by an inherent semantic-localization dichotomy. While CLIP’s global features excel at image-level classification, they lack the spatial sensitivity required for pixel-level segmentation. Existing approaches...