Skip to content
Conference Open access

A CNN and BiLSTM Network for Predicting Job Failures in Dynamic Cloud Workloads

2026 · E3S Web of Conferences · 0 citations · 11 references

Abstract

Cloud service providers face significant challenges in preventing hardware and software failures due to the large-scale and heterogeneous nature of cloud computing. Although many studies have focused on characterising failed jobs, fewer have explored proactive failure prediction. This paper presents a deep learning-based failure prediction model that integrates Convolutional Neural Networks (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) networks to identify job failures before they occur. The proposed model improves the performance of cloud computing applications by reducing job failures and optimising resource utilisation. Using the Google Cluster Traces dataset, we analyse failure patterns and evaluate the effectiveness of the model across multiple performance metrics. The results demonstrate the robustness of the proposed scheme, achieving an accuracy of 99.96%, along with a high F1-score of 99.92% when compared to existing models. These findings highlight the potential of deep learning in proactive failure mitigation, providing a foundation for future advances in cloud workload reliability.

Read PDF