Skip to content
Open access

Using ensemble learning and Gaussian mixture model to predict petrophysical properties and hydraulic flow units in carbonate reservoirs.

Jul 2026 · Scientific Reports · 0 citations
Medicine

Abstract

Accurate prediction of porosity and permeability in complex carbonate reservoirs is very important for understanding reservoirs, but remains challenging due to inherent heterogeneity. This study develops a robust, machine learning-driven workflow to enhance the prediction of these critical petrophysical properties and the identification of Hydraulic Flow Units. The methodology integrates conventional core data and geophysical well logs, employing advanced data preprocessing, including depth matching, which significantly improved the log-core porosity correlation. A key innovation involves using a Gaussian Mixture Model for unsupervised Hydraulic Flow Unit identification, which outperformed traditional empirical methods and K-Means clustering by yielding five distinct Hydraulic Flow Units with high intra-unit porosity-permeability correlations (R2 up to 0.93) validated by Mercury Injection Capillary Pressure data. For predictive modeling, a comprehensive comparison of algorithms revealed that a Voting ensemble meta-algorithm with a Multi-Layer Perceptron base learner delivered superior performance for both porosity (on integrated data from three wells) and permeability (modeled per Hydraulic Flow Unit). The final models successfully estimated properties in non-cored intervals and a blind well, demonstrating high accuracy and generalizability. This integrated approach provides a reliable and theory-grounded framework for characterizing heterogeneous carbonate reservoirs, reducing dependency on extensive coring operations.

Read PDF

Similar papers

Aug 2026

Ensemble Machine-Learning-Based Horizontal Permeability Prediction and Hydraulic Flow Unit Characterization of the Hugin Sandstone

Understanding permeability is essential for evaluating reservoir quality and field development planning. Reliable permeability estimation can reduce the uncertainty in reservoir characterization, particularly in intervals where core data are limited. As the industry relies on log-based interpretations and empirical correlations, the limitations of these approaches become apparent. Data-driven approaches offer a promising alternative to conventional empirical methods. The data set in this study comprises 252 samples with seven features derived from conventional well logs. Data preprocessing includes handling missing values, smoothing logs, feature engineering to add an extra input, and transformation with the Yeo-Johnson technique. A center moving average filter was used to reduce variance and improve data consistency. Ensemble machine-learning (ML) and baseline models were developed and evaluated using a 75-25 train-test split, four-fold cross validation, and model complexity assessment. Ensemble methods outperformed baseline models, with extremely randomized trees (ET) and random forest (RF) emerging as the most stable, achieving a mean R² of 0.93 and 0.89 and a low R² standard deviation (0.3). Multilinear regression (MLR) and artificial neural networks (ANNs) show limited accuracy, while gradient boosting (GB) and extreme gradient boosting (XGBoost) methods exhibit overfitting despite perfect training scores. Predicted kh values were compared with core data. Both linear and nonlinear empirical equations were derived using MLR, polynomial regression, and a power-law model. The power-law model (empirical equation) achieved an R² value of 0.79 and can therefore be used to estimate permeability. Additionally, a Gaussian mixture model (GMM) was used for unsupervised classification of hydraulic flow units (HFU) using the flow zone indicator (FZI), computed from the continuous permeability curve obtained from the best ML model. Thus, ML-based permeability prediction is an indispensable component of HFU modeling. The model identified three distinct flow zones, enabling HFU clustering and defining their corresponding petrophysical properties and depositional environments.

Vikram Kumar, Sayantan Ghosh, S. Maiti · 0 citations
Open access Aug 2026

Ensemble Machine Learning for Seismic Attribute-Based Porosity Prediction and Lithofacies Classification in Clastic Hydrocarbon Reservoirs: A Comparative Workflow with Uncertainty Assessment

Accurate prediction of reservoir porosity and reliable lithofacies classification are fundamental to hydrocarbon exploration and reservoir development because they directly influence reserve estimation, well placement, and production optimization. Conventional seismic interpretation methods often struggle to capture the complex nonlinear relationships between seismic attributes and reservoir properties, particularly in heterogeneous clastic formations. This study presents an integrated machine learning workflow for simultaneous porosity prediction and lithofacies classification using post-stack seismic attributes calibrated with well-log observations. Twenty seismic attributes representing amplitude, frequency, phase, geometric, and textural characteristics were extracted from a three-dimensional seismic volume and screened using a systematic feature-selection strategy. Four supervised machine learning algorithms, namely Random Forest (RF), Support Vector Regression (SVR), Gradient Boosting Regression (GBR), and Artificial Neural Networks (ANN), were developed and compared for porosity estimation, while corresponding classification models were evaluated for lithofacies prediction. Model performance was assessed using k-fold cross-validation and blind-well validation to ensure robust generalization. Predictive uncertainty was quantified through ensemble-based confidence estimation and incorporated into the final reservoir property volumes. Illustrative placeholder results indicate that ensemble learning algorithms consistently outperform conventional regression approaches by effectively capturing nonlinear relationships among seismic attributes while providing improved porosity prediction accuracy and more reliable lithofacies discrimination. The proposed workflow integrates feature selection, comparative machine learning evaluation, blind-well validation, and uncertainty assessment into a unified framework that can be readily adapted to other clastic hydrocarbon reservoirs. This study demonstrates the potential of modern machine learning techniques for quantitative seismic reservoir characterization while providing confidence-aware predictions for exploration and field-development decision making.

Pronab Chowdhury · 0 citations
Open access Jul 2026

Machine learning-based seismic attribute analysis for porosity prediction and lithofacies classification in clastic hydrocarbon reservoirs

Accurate prediction of reservoir properties such as porosity, permeability, and lithofacies distribution is essential for reliable hydrocarbon reserve estimation, well placement, and field development planning. Conventional deterministic methods for relating seismic response to reservoir properties are limited by the nonlinear and multivariate nature of the seismic–petrophysical relationship, and single-attribute correlations frequently fail to capture the complexity of clastic reservoir systems. This study presents a machine learning-assisted workflow that integrates multi-attribute seismic analysis with wireline log data to predict porosity and classify lithofacies within a clastic hydrocarbon reservoir. A suite of seismic attributes including root-mean-square (RMS) amplitude, sweetness, instantaneous frequency, coherence, envelope amplitude, and acoustic impedance derived from post-stack inversion was extracted and calibrated against log measurements at control wells. Feature selection based on correlation ranking and recursive feature elimination identified the most informative attributes, and three supervised learning models random forest (RF), support vector regression (SVR), and a feed-forward artificial neural network (ANN) were trained for porosity prediction. A separate classification stage using random forest and gradient boosting was applied for lithofacies discrimination. The results demonstrate that the ANN and RF models substantially outperform single- and multi-attribute linear regression, achieving a coefficient of determination (R²) of 0.86–0.88 for porosity prediction at blind-test wells, while the ensemble classifier reached an overall lithofacies classification accuracy of 88%. Integrating machine learning with multi-attribute seismic analysis reduces prediction uncertainty and produces spatially continuous reservoir property volumes that support more reliable reservoir characterization and lower-risk exploration decision-making.

Rodwan A. Elbarouni · 4 citations
Conference Open access Aug 2026

Prediction of Porosity and Permeability Using Well Log and Core Data: A Data-Driven Approach

Accurate prediction of porosity and permeability is very important for reservoir characterization and hydrocarbon extraction. Traditional workflow in the form of empirical correlations is usually difficult, time-consuming, spatially limiting, and entirely dependent on formation geology. The current study examines a different approach, which uses machine learning (ML) regression models based on well-log and core data. Three regression architectures including Random Forest, CatBoost, and K-Nearest Neighbors (KNN) were trained and validated based on a dataset consisting of 340 samples of shaly sand gas reservoirs. The gamma ray (GR), resistivity (RLLD), spontaneous potential (SP), bulk density (RHOB), neutron porosity (NPHI), and depth were used as the input variables with the core-derived porosity (CPHI) and permeability (CKHG) being used as the targets. The quantitative measures of performance of the models included R2, Mean Absolute Error (MAE), and Root Mean Squared Error (RMSE). The findings indicated that KNN regression was better than its counterparts as it achieved R2 = 0.8933 and R2 = 0.9340 in terms of porosity and permeability prediction, respectively, and more acceptable metrics of errors showed. Comparatively, the traditional empirical methods showed a significantly lower accuracy rate. The findings highlight that machine learning has the potential to provide precise, scalable, and low-cost predictions of the reservoir properties which could lead to better choices for exploration and production activities.

Md Alamin Islam, Bintun Zaman, Shahria Nayem Ahmed · 0 citations