Skip to content

Machine learning-based fault detection in low-speed bearings using a multi-environmental dataset

Jul 2026 · Engineering Research Express · Vol 8 · 0 citations · 28 references
Physics

TL;DR

The near-perfect linear separability indicates that the dataset’s binary, controlled-laboratory labelling rather than intrinsic bearing-degradation physics drives the clean classification, and validation on 500–1000 or more samples with progressive-degradation labelling is essential before any operational claim can be supported.

Abstract

This study provides a systematic robustness evaluation of classical machine learning for vibration-based bearing fault detection in low-RPM internal combustion engines (1000–2000 RPM) across a controlled temperature × humidity grid, a regime underrepresented in benchmark datasets that emphasise high-speed applications. A publicly available dataset from a 658cc engine (–10 °C–45 °C, 0%–100% humidity) was analysed; vibration features were derived from the non-zero channels of a tri-axial acquisition, with the bearing-housing vibration carried primarily by channel Ch3. To prevent temporal leakage, 390 263 continuous measurements were aggregated into 89 steady-state units, each spanning 90 s, yielding a deliberately independence-preserving but low sample-to-feature ratio (89:92). Four algorithms Random Forest, Support Vector Machine, Logistic Regression, and Neural Network were evaluated using stratified 5-fold cross-validation. All models achieved apparent accuracy exceeding 95%, with Random Forest performing best (97.8% ± 2.7%), but no statistically significant differences were found (Friedman test, p = 0.732). Vibration features, particularly crest factor and root mean square, provided the greatest discriminative power, while environmental factors accounted for less than 17% combined importance. The near-perfect linear separability (98.2% with Logistic Regression) indicates that the dataset’s binary, controlled-laboratory labelling rather than intrinsic bearing-degradation physics drives the clean classification. Accordingly, the reported accuracies are apparent upper-bound estimates from an exploratory study, not expected field performance; validation on 500–1000 or more samples with progressive-degradation labelling is essential before any operational claim can be supported.

View source

Similar papers

Conference Jul 2026

Machine learning-based bearing failure research

It is confirmed that traditional machine learning models, optimized through manual feature engineering, can provide a ‘high-precision, low-risk’ solution for bearing fault diagnosis and offers significant reference value for the intelligent operation of aero-engines and other industrial equipment.

Qianxi Ye, Pengfang Gao · 0 citations
Aug 2026

Machine learning-based feature-driven model generation and evaluation for multi-fault bearing diagnosis using XGBoost and time-domain statistical features of vibration data

The findings affirm the efficacy of XGBoost in bearing fault classification and emphasise the diagnostic value of carefully selected time-domain features, as well as suggesting strong potential for deploying such models in real-time condition monitoring and predictive maintenance systems.

A. Bhende · 0 citations
Aug 2026

Multi-Class Fault Detection and Diagnosis of Rolling Bearings: a Machine Learning Approach

Results show that using statistical vibration features with ensemble classifiers is a good way to diagnose multi-class bearing faults and establishes a comprehensive benchmark for ML- and DL-based rolling bearing FDD.

M. I. Quamar, Abdulrazaq Nafiu Abubakar, Ali Nasir · 0 citations
Open access 2026

Prediction and Classification of Residual Service Life in Wind Turbine Bearings Under Variable Speed Conditions Using Hybrid Machine Learning Models

The increasing demand for reliability in wind turbine systems makes early bearing fault detection under variable-speed conditions a persistent challenge. This paper proposes a hybrid methodology for classifying degradation stages and estimating a relative RUL-related degradation indicator for bearings by integrating synthetic data modeling, feature selection, and a combined unsupervised–supervised learning approach. Synthetic vibration signals are generated through logistic-curve interpolation with pink noise, enabling controlled degradation simulation. Features from time, frequency, and time–frequency domains were ranked using Mutual Information, and Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) was employed to identify progressive wear stages. Cluster centers serve as anchors for mapping degradation into RUL percentages, while classification ensures stage consistency. Experimental results demonstrate six well-defined clusters for inner-race faults (Silhouette 0.5190), three moderate clusters for outer-race faults (0.2339), and overlapping patterns for rolling element faults (–0.1527), with zero RUL deviation in the best case. The proposed framework combines real and synthetic data to enhance generalization while reducing computational cost, offering a reliable and scalable solution for predictive maintenance and assessment of degradation progression in wind turbine bearings.

Gustavo Gomes Do Valle, Benjamin Soudhan, Meisam Mahdavi et al. · 0 citations
Open access Jul 2026

Artificial Intelligence-Driven Quality Control in Mechanical Manufacturing: Vibration-Based Multiclass Gear Fault Detection Using LightGBM

The results show that vibration-based machine learning can support robust, near-real-time fault identification in mechanical manufacturing environments and highlights the importance of chronological validation, feature engineering over multiple time windows, and the trade-off between predictive performance and deployment efficiency.

P. Malega, J. Kováč, Róbert Munkáči et al. · 0 citations