A Vision Transformer for Bearing Fault Diagnosis Based on Wavelet Transform and Frequency-Domain Circulant Attention
Abstract
Deep learning models often struggle with complex noise and variable loads in industrial fault diagnosis. This paper proposes a highly efficient framework, the Frequency-domain Circulant Attention Vision Transformer (FC-ViT), for robust rotating machinery monitoring. Raw vibration signals are first converted into 2D time-frequency representations via Continuous Wavelet Transform (CWT). FC-ViT then utilizes an improved Frequency-domain Circulant Attention mechanism, which achieves a log-linear computational complexity of $O(N \log N)$ , to isolate fault-related impulses from heavy stochastic and non-Gaussian impulsive noise. Validated on the Case Western Reserve University (CWRU) and Paderborn University (PU) datasets, FC-ViT achieves 100% accuracy under noise-free conditions. At an extreme −5 dB SNR, it maintains 89.4%–98.3% accuracy, significantly outperforming both state-of-the-art 1D diagnostic networks and generic 2D vision models (e.g., Swin-Transformer, ConvNeXt, and CA-DeiT). These results demonstrate superior noise immunity and cross-load generalization. Furthermore, comprehensive hardware deployment evaluations confirm its ultra-low inference latency and minimal memory footprint, providing a practical and real-time solution for real-world industrial condition monitoring.