Litho-fluid facies prediction by integrating well logs and seismic attributes: a SMOTE-enhanced machine learning approach addressing class imbalance
Abstract
As exploration efforts intensify, deep lithologic hydrocarbon reservoirs have emerged as principal objectives in hydrocarbon prospecting. Various types of hydrocarbon reservoirs, including shale, tight sandstone, and carbonate rock, are closely related to lithology. However, lithologic predictions alone provide incomplete reservoir characterization, fluid properties constitute an essential discriminator. Leveraging field dataset characteristics, we defined litho-fluid facies through integrated petrophysical parameters (porosity and permeability), lithology (sandstone or mudstone), and fluid type (oil, gas, water or dry). Commencing with well logs and 3D seismic attributes, we optimized litho-fluid facies sensitivity via cross-plot analysis and correlation analysis. To address sample imbalance in litho-fluid facies labeling, we implemented Synthetic Minority Over-sampling Technique (SMOTE) for sample balancing. Ultimately, litho-fluid facies classification based on multiple machine learning (ML) methods is achieved. The ML models are trained using two fully well logs (incorporating attributes of seismic trace nearby borehole), with blind-test validation conducted on a third well. Subsequently, we estimate 3D litho-fluid facies data through the result of seismic inversion and signal analysis. Results demonstrate that Categorical Boosting (CatBoost) classifier delivers superior reliability and robustness in facies discrimination. This integrated workflow presents significant potential for deployment in underexplored areas with sparse well logs, enabling predictive reservoir in data-constrained exploration settings.