Training-Free and Self-Improving Industrial Anomaly Detection with Vision-Language Models
Abstract
Although industrial anomaly detection has attained highly accurate pixel-level anomaly detection on standard benchmarks, current methods merely produce heatmaps and anomaly scores. These outputs remain insufficient to address the core concerns of inspectors-the type, severity, and root cause of a defect. We present a framework that couples a training-free anomaly localizer with a vision-language model verifier. The localizer exploits DINOv3 ViT-L features and PCA reconstruction error to achieve a pixel-level AUROC of 0.979 on MVTec AD, while the verifier, which operates exclusively at inference, translates detected anomalies into natural language inspection reports. The entire pipeline is training-free: PCA fitting completes in roughly 30 seconds per product category, with no model training required. Using a frozen VLM, selectively inspecting only high-score regions cuts VLM calls by 30.1% without sacrificing recall, and incremental learning from VLM-verified samples autonomously improves detection accuracy without any human annotation. This training-free, self-improving pipeline bridges the gap between pixel-accurate anomaly detection and semantic information needed for real-world industrial inspection.