InspectorGPT, a VLM framework centered on comparative reasoning, is proposed, which demonstrates superior multi-dimensional performance and generalization to unseen benchmarks, validating comparative reasoning for comprehensive industrial inspection.
Abstract
Industrial anomaly detection is a critical component of modern manufacturing. Most traditional unsupervised methods rely on modelling normal feature distributions, inherently limiting generalization to unknown categories. To improve generalizability, some recent methods incorporate vision-language models (VLMs) for zero-shot detection via text prompts. However, we observe that reasoning-oriented post-training can cause anomaly discrimination to collapse, with some fine-tuned models performing worse than their base VLMs. Existing methods also provide only textual decisions or coarse boxes, without pixel-level segmentation. A more explicit detection principle comes from human inspection: anomalies are identified by comparing a query image with a defect-free reference. Inspired by this, we propose InspectorGPT, a VLM framework centered on comparative reasoning. Given a normal reference and a query image, InspectorGPT compares them to identify discrepancies and perform multiple inspection tasks with detailed reasoning. We internalize this capability through Chain-of-Thought (CoT) fine-tuning and Group Relative Policy Optimization (GRPO) with tailored, verifiable rewards. We further introduce InspectorGPT-Seg for pixel-level anomaly masks. Segmentation supervision improves anomaly discrimination but weakens semantic reasoning, while joint training fails to balance them. We therefore train the two branches separately and combine them through task-vector fusion. Extensive experiments demonstrate superior multi-dimensional performance and generalization to unseen benchmarks, validating comparative reasoning for comprehensive industrial inspection.
Although industrial anomaly detection has attained highly accurate pixel-level anomaly detection on standard benchmarks, current methods merely produce heatmaps and anomaly scores. These outputs remain insufficient to address the core concerns of inspectors-the type, severity, and root cause of a defect. We present a f...
Industrial anomaly detection (IAD) requires reliable identification and precise localization of subtle defects, yet most existing methods depend on manually tuned decision thresholds and large collections of defect-free samples, limiting scalability in real-world production. To address these constraints, we present gro...
Industrial anomaly detection (IAD) is evolving beyond conventional detection and localization toward multimodal inspection systems that can describe, explain, and reason about fine-grained defects. Although recent multimodal large language model (MLLM)-based methods improve anomaly understanding through textual reasoni...
Jaron Yeh, Yen-Wei Chang, Jiang Liu et al.· 0 citations
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions ma...
Meng-Yang Zhao, Zhuo-Lin He, Hai-Yang Yu et al.· 0 citations
A novel large multimodal model applying vision experts for industrial anomaly detection (abbreviated as Myriad), which treats conventional IAD models as VEs and converts their anomaly maps into lightweight prompts that steer a frozen Q-Former toward suspicious regions, while a compact low-rank adapter shapes features f...
Yuanze Li, Haolin Wang, Shihao Yuan et al.· Science China Information Sc...· 0 citations
Industrial visual inspection is a key task in intelligent manufacturing and quality control. However, defective samples in real production lines are usually scarce, diverse in appearance, and expensive to annotate, which makes supervised models that rely on large numbers of defective samples difficult to adapt to new p...
Yan Wang, Guan Zhang, Chun-Xiao Wu et al.· Artificial Intelligence and...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.