Multimodal and multiscale adaptive feature fusion for fault diagnosis of rotating machinery
Abstract
To overcome insufficient feature extraction, poor generalization, and high computational costs in rotating machinery fault diagnosis, this paper proposes Vision Transformer with multi-channel and multiscale adaptive feature fusion (MCMSAF-ViT), a lightweight acoustic-vibration bimodal ViT. First, 1D time-series signals are transformed into 2D images via data encoding and JET mapping. Next, a parallel dual-channel architecture extracts multiscale features using varying dilated convolutions. Spatial and channel attention mechanisms dynamically weight these features to enhance discriminative representation before fusing them for classification. Validated on datasets from the University of Ottawa and Huazhong University of Science and Technology, MCMSAF-ViT achieves consistently high accuracy, outperforming baselines like ResNet and EfficientNet under noise and complex working conditions. Moreover, the parameter count of MCMSAF-ViT is reduced to the 105 level, demonstrating a favorable balance between diagnostic accuracy and model compactness. These results provide a compact and effective framework for rotating machinery fault diagnosis.