Machine Learning for Electric Submersible Pump Fault Diagnosis: Physical Interpretability, Cross-Unit Generalization, and Evaluation Pitfalls
Abstract
Machine learning (ML) techniques have been widely applied to fault detection and diagnosis in Electric Submersible Pumps (ESPs), often reporting high predictive accuracy. However, high performance does not necessarily imply that learned decision boundaries reflect physically meaningful fault mechanisms. This study distinguishes epistemic interpretability, associated with model transparency, from physical interpretability, defined here as the stability of diagnostic decisions across distinct physical units, with operating-regime stability discussed as a broader requirement that cannot be directly isolated from the reference feature file used in this study. Current ML practices in ESP fault diagnosis are examined through a structured literature review and an empirical analysis of a public multi-pump vibration dataset. Representative supervised models are evaluated under sample-wise and cross-unit (pump-wise) validation. Performance is assessed using accuracy, macro-F1, confusion matrices, Principal Component Analysis-based class-space geometry, centroid distances, stability metrics, and bootstrap-based confidence intervals and empirical significance tests for cross-unit degradation. Results show that sample-wise validation can overestimate robustness, whereas cross-unit evaluation reveals performance degradation in most nonlinear and ensemble configurations, with statistical support in several high-capacity models. Class imbalance mitigation reduces majority-class bias but does not eliminate structural misclassification patterns. Representation-level analysis shows that fault condition is the stronger organizing factor in the feature space, while pump identity contributes a measurable but non-dominant fraction of feature-space variance. These findings indicate that predictive accuracy alone is insufficient to characterize diagnostic quality in ESP applications. Explicit cross-unit validation and stability-oriented evaluation are required to assess diagnostic robustness under conditions closer to deployment-relevant generalization.