The proposed MSFormer incorporates a parallel multi-scale Convolutional Neural Network architecture and hierarchical Transformer modules to comprehensively process 1D vibration signals to provide a powerful and precise intelligent solution for mechanical fault diagnosis.
Abstract
In rotating machinery monitoring, obtaining highly discriminative fault features from complex vibration signals remains a significant challenge for deep learning-based diagnostic models. In this paper, a novel intelligent fault diagnosis method named MSFormer is proposed. The MSFormer incorporates a parallel multi-scale Convolutional Neural Network (CNN) architecture and hierarchical Transformer modules to comprehensively process 1D vibration signals. By utilizing varying kernel sizes, the multi-scale CNN extracts both high-frequency local transient impulses and low-frequency global degradation trends. Subsequently, the Transformer modules are employed to model the long-range dependencies within the extracted feature sequences, effectively mitigating the interference of environmental noise. Extensive experiments are conducted on bearing fault experimental data to evaluate the proposed method. Four state-of-the-art models are compared under the same experimental settings. Quantitative metrics and qualitative tools are utilized for comprehensive evaluation. Experimental results indicate that MSFormer achieves a 92.67% accuracy, 92.54% F1-score, and 93.22% precision, demonstrating significant superiority. MSFormer provides a powerful and precise intelligent solution for mechanical fault diagnosis.
To address the degradation of cross-condition diagnostic performance caused by feature-scale drift in rolling bearing vibration signals under variable operating conditions, this paper proposes a spectral-guided adaptive multi-scale convolutional neural network (SAMACNN). First, PSD sequences and time-frequency features are introduced as dual-stream inputs. While the time-frequency main branch extracts local information, the Spectral Transformer spectral bypass branch captures long-range dependencies in harmonic structures. Second, dynamic gating weights are generated for the multi-scale convolutional branches, enabling sample-conditioned multi-scale feature selection and fusion and alleviating scale mismatch caused by fixed receptive fields and static fusion. Finally, data collected from two bearing fault simulation test rigs are used to verify the effectiveness and superiority of the proposed algorithm. The experimental results show that the proposed SAMACNN method achieves average accuracies of 97.16% and 97.56% on the two datasets, respectively, outperforming the ablation variants and demonstrating strong robustness and generalization capability in complex variable-condition measurement environments.
DaXin Li, Wang Hong, Hai Xue et al.· Engineering Research Express· 0 citations
A robust and noise-resilient bearing fault diagnosis framework that integrates advanced signal processing with hybrid deep learning techniques is presented, demonstrating strong robustness and generalization capability.
Sujit Kumar, Manish Kumar, Bam Bahadur Sinha· International Journal of Dyn...· 0 citations
In this study, a hybrid deep learning architecture is proposed for robust vibration-based fault diagnosis in industrial machinery by jointly modeling time-domain, frequency-domain, and temporal dynamics. In the time domain, convolutional layers combined with an Enhanced Gated Attention (EGA) mechanism emphasize informative signal components while suppressing noise. Temporal evolution is modeled using Neural Ordinary Differential Equations (Neural ODEs), enabling smooth and stable continuous-time feature representations. In parallel, a Fourier Neural Operator (FNO) extracts frequency-domain characteristics, augmented with gated attention to focus on fault-related spectral patterns. Long Short-Term Memory (LSTM) layers capture long-range dependencies, while Squeeze-and-Excitation (SE) blocks adaptively recalibrate channel-wise feature responses. A multi-scale attention-based fusion module integrates domain-specific representations and auxiliary features to enhance discrimination under varying operating conditions. The proposed model is evaluated on the SUBFv1.0 dataset through extensive ablation studies and experiments under multiple noise levels, achieving 98.41% accuracy in noise-free conditions and maintaining performance above 91% even at 5 dB SNR. Unlike existing multi-path approaches that combine heterogeneous features in a loosely coupled or discrete manner, the proposed architecture uniquely integrates multi-domain feature learning with bidirectional attention mechanisms and continuous-time temporal dynamics, enabling coherent cross-domain interaction and robust fault characterization.
Canan Taştimur· Information Technology and C...· 0 citations
A rolling bearing fault diagnosis method based on multi-scale depthwise separable convolution (MDSC) and a convolutional neural network–Transformer hybrid model (CNN-Transformer) is proposed to address the non-stationarity of fault signals and the difficulty of jointly capturing local and global features. First, continuous wavelet transform (CWT) converts one-dimensional vibration signals into two-dimensional time-frequency images to enhance fault representation. Then, multi-scale convolution (MSC) and depthwise separable convolution (DSC) are introduced to extract local impulsive features and fault patterns at different scales with fewer parameters. A CNN-Transformer architecture is further developed, where convolutional neural network (CNN) captures local details and Transformer models global dependencies. In addition, pretraining-finetuning, data augmentation, label smoothing, and normal sample optimization are adopted to improve training stability and diagnostic performance. Experimental results show accuracies of 98.80% on the Xi’an Jiaotong University bearing dataset (XJTU-SY) and 100.00% on the Case Western Reserve University bearing dataset (CWRU), demonstrating strong discriminative ability, stability, and robustness.
Shuai Yang, Yanchao Chen, Yang Yu· Engineering Research Express· 0 citations
Fault signals in rotating machinery typically manifest as long time-series data embedded with local high-frequency impulses. Traditional deep learning methods often struggle to simultaneously capture these transient local impacts and model long-term global degradation features. To address this challenge, this paper proposes a novel GCU-SAM enhanced Transformer for intelligent fault diagnosis. The proposed network integrates a Gated Convolutional Unit (GCU) with a Self-Attention Mechanism (SAM). By introducing the GCU as a local inductive bias prior to the global attention module, the model dynamically captures and purifies local impulse responses via reset and update gates. Subsequently, a cascaded multi-head self-attention mechanism models long-sequence global evolution trends, forming an integrated framework for local fine-grained perception and global correlation modeling. Validated on the CWRU bearing and SEU gearbox datasets, the proposed architecture achieves superior diagnostic performance with an average F1-score of 99.01%. Compared to representative baselines, including 1D-CNN, BiLSTM, and the standard Transformer, the GCU-SAM significantly boosts diagnostic accuracy and effectively overcomes the early-stage optimization oscillations inherent in pure attention mechanisms. By achieving a deep multi-scale fusion of local abrupt changes and global degradation trends, the model exhibits exceptional feature-clustering discriminative capability and convergence stability, providing a robust and highly accurate solution for the intelligent fault diagnosis of complex rotating machinery.
Jing Li, Lei Hu, Peng Luo· Italian National Conference...· 0 citations
Deep learning models often struggle with complex noise and variable loads in industrial fault diagnosis. This paper proposes a highly efficient framework, the Frequency-domain Circulant Attention Vision Transformer (FC-ViT), for robust rotating machinery monitoring. Raw vibration signals are first converted into 2D time-frequency representations via Continuous Wavelet Transform (CWT). FC-ViT then utilizes an improved Frequency-domain Circulant Attention mechanism, which achieves a log-linear computational complexity of $O(N \log N)$ , to isolate fault-related impulses from heavy stochastic and non-Gaussian impulsive noise. Validated on the Case Western Reserve University (CWRU) and Paderborn University (PU) datasets, FC-ViT achieves 100% accuracy under noise-free conditions. At an extreme −5 dB SNR, it maintains 89.4%–98.3% accuracy, significantly outperforming both state-of-the-art 1D diagnostic networks and generic 2D vision models (e.g., Swin-Transformer, ConvNeXt, and CA-DeiT). These results demonstrate superior noise immunity and cross-load generalization. Furthermore, comprehensive hardware deployment evaluations confirm its ultra-low inference latency and minimal memory footprint, providing a practical and real-time solution for real-world industrial condition monitoring.
Zhijun Teng, Li Jiqi, Mingyang Sun et al.· IEEE Access· 0 citations