Skip to content
Open access

H-SCGF: Multi-Modal 3-D Anomaly Detection Combining Structural Gating With Spatially Consistent Feature Matching

2026 · IEEE Access · Vol 14, pp. 146659-146671 · 0 citations · 31 references

Abstract

Multi-modal 3D anomaly detection is a key technology in industrial automated quality control, crucial for identifying complex manufacturing defects. However, existing methods still face significant challenges in handling structural anomalies that disrupt higher-order statistical relationships or spatial organization. This paper introduces H-SCGF (Hierarchical Structurally-Consistent Guided Fusion), a detection framework designed to improve structural anomaly localization by combining guided multi-modal fusion with spatially consistent feature matching. First, at the feature extraction level, we employ strong self-supervised models, namely DINO (ViT-Base/8) for appearance features and Point-MAE for geometric features, ensuring high-quality representations from the outset. Second, at the feature fusion level, we design a trainable Structural Gating Fusion (SGF) module and optimize it together with lightweight adapter layers using a Hierarchical Structure Consistency Loss ( $L_{HSC}$ ), while keeping the DINO and Point-MAE backbones frozen. This module guides the fusion of feature streams by generating dynamic gating signals from the computed second-order statistical (Gram matrix) differences between RGB and geometric features across multiple scales, thereby producing the fused descriptor, the output of SGF, that is sensitive to structural consistency. Finally, at the anomaly scoring level, we propose a Multi-bank Spatially-Consistent Scoring (MSCS) mechanism. This mechanism utilizes three parallel feature banks (for appearance, geometry, and fused descriptors) and adds spatial context verification to traditional nearest-neighbor matching: it checks whether a test patch neighborhood is consistent with the neighborhood of its matched normal prototype. Here, neighborhood structure means the local spatial arrangement stored with that prototype. Experiments on the MVTec 3D-AD and Eyecandies benchmark datasets demonstrate that H-SCGF achieves an average pixel-level AUPRO of 97.2% on MVTec 3D-AD and 89.1% on Eyecandies (with a peak of 98.5% on Confetto), indicating its practical value for industrial anomaly localization.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.