Skip to content
Conference Open access

GRASP: Enhancing Zero-Shot Multi-Modal Anomaly Detection via Geometric Refinement and Semantic Prompting

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · 0 citations · 38 references

TL;DR

A unified zero-shot multimodal anomaly detection framework GRASP is proposed, with substantial improvements of 3.0 points in I-AUROC and 2.7 points in AUPRO compared to the SOTA method.

Abstract

Zero-shot multimodal anomaly detection is critical for identifying structural defects that are often invisible to traditional RGB images, particularly in scenarios lacking target domain samples. However, existing methodologies face two significant impediments: the prohibitive computational overhead caused by relying on multi-view rendering for depth processing, and the inability of static text prompts to adapt to fine-grained local anomalies. To address these challenges, a unified zero-shot multimodal anomaly detection framework GRASP is proposed. Firstly, a Frequency Domain Enhancement module is introduced to replace costly rendering with spectral transformations, directly synthesizing high-fidelity depth images for efficient utilization. Secondly, a Prompt-Conditioned Variational module is designed to bridge the semantic gap by grounding global textual descriptions into local visual nuances. Finally, a Dual Cross-Injection Alignment module is proposed to enable robust feature fusion for enhancing anomaly classification performance, while a Pyramid Anomaly Map Recalibration module further refines anomaly localization across multiple scales. Extensive experiments on MVTec 3D-AD and Eyecandies demonstrate that GRASP establishes a new state-of-the-art, yielding substantial improvements of 3.0 points in I-AUROC and 2.7 points in AUPRO compared to the SOTA method.

Read PDF

Similar papers

Preprint Aug 2026

Rethinking Auxiliary Modalities in Multimodal Zero-shot Anomaly Detection: From Semantic Fusion to Conditional Modulation

Recent foundation model-based methods have endowed RGB images with strong zero-shot anomaly detection (ZSAD) through vision-language pretraining. However, RGB observations alone remain limited in perceiving anomalies dominated by geometric deformation, depth variation, or subtle surface changes. Auxiliary modalities ca...

Peng Wu, Xin Ge, Yujia Sun et al. · 0 citations
2026

Cross-Modal Guidance Learning for Zero-Shot Industrial Anomaly Detection

Zero-shot industrial anomaly detection (ZIAD) aims to develop a unified model capable of directly identifying unseen anomaly categories in images without requiring reference samples. Recently, large-scale Vision-Language Models (VLMs) such as CLIP have shown great potential for solving this task. However, existing meth...

Tiyu Fang, Lin Zhang, Ran Song et al. · 0 citations
Preprint Aug 2026

Dual Anchors, Do It Better: Hierarchical Group Merging for Zero-Shot Anomaly Detection

Zero-shot anomaly detection (ZSAD) aims to identify anomalies in unseen domains, a setting that is particularly critical for industrial and medical applications where domain shifts are prevalent. However, most CLIP-based ZSAD methods anchor semantics solely on the text modality, making performance highly sensitive to p...

Jimin Roh, Dongkyu Kim, Suk-Ju Kang · 0 citations
Open access Sep 2026

Boundary-Guided Dual-Perspective Cross-Modal Fusion Network for RGB-IR Object Detection

A Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) is proposed to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives.

Hu Lin, Zhi-Wei Fu, Xiu-Mei Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.