Skip to content

Component Failure Prediction for the SCANIA Component X Dataset

Sep 2026 · International Symposium on Networks, Computers and Communications · pp. 1-6 · 0 citations · 8 references

Abstract

Predictive maintenance systems use operational data to identify equipment that is approaching failure before the failure causes downtime or unnecessary repair cost. This paper presents a machine learning workflow for the SCANIA Component X predictive maintenance dataset, a real-world heavy-duty vehicle dataset containing anonymized operational readouts, vehicle specifications, repair outcomes, and multi-class failure-window labels. The goal is to classify each vehicle into one of five risk classes, where class 0 indicates no imminent failure and classes 1 through 4 represent increasingly urgent failure windows. Because the data is highly imbalanced, raw accuracy is not enough to choose a useful model. The official SCANIA cost matrix penalizes missed failures much more heavily than unnecessary inspections, so model selection is driven by total validation cost while accuracy and macro F1 are reported for context. We compare all-class-0, nearest-centroid, Gaussian Naive Bayes, Random Forest, distance-weighted k-nearest neighbors, logistic regression, LightGBM classifiers. The best complete-validation operating point is a cost-sensitive LightGBM classifier with 54.4% validation accuracy and validation cost 35,599, compared with the all-class-0 baseline cost of 57,400. The model is also exported to an interactive dashboard that replays telemetry and scores inserted records using the same feature construction pipeline.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.