Skip to content

Ensemble Machine-Learning-Based Horizontal Permeability Prediction and Hydraulic Flow Unit Characterization of the Hugin Sandstone

Aug 2026 · Petrophysics · 0 citations

Abstract

Understanding permeability is essential for evaluating reservoir quality and field development planning. Reliable permeability estimation can reduce the uncertainty in reservoir characterization, particularly in intervals where core data are limited. As the industry relies on log-based interpretations and empirical correlations, the limitations of these approaches become apparent. Data-driven approaches offer a promising alternative to conventional empirical methods. The data set in this study comprises 252 samples with seven features derived from conventional well logs. Data preprocessing includes handling missing values, smoothing logs, feature engineering to add an extra input, and transformation with the Yeo-Johnson technique. A center moving average filter was used to reduce variance and improve data consistency. Ensemble machine-learning (ML) and baseline models were developed and evaluated using a 75-25 train-test split, four-fold cross validation, and model complexity assessment. Ensemble methods outperformed baseline models, with extremely randomized trees (ET) and random forest (RF) emerging as the most stable, achieving a mean R² of 0.93 and 0.89 and a low R² standard deviation (0.3). Multilinear regression (MLR) and artificial neural networks (ANNs) show limited accuracy, while gradient boosting (GB) and extreme gradient boosting (XGBoost) methods exhibit overfitting despite perfect training scores. Predicted kh values were compared with core data. Both linear and nonlinear empirical equations were derived using MLR, polynomial regression, and a power-law model. The power-law model (empirical equation) achieved an R² value of 0.79 and can therefore be used to estimate permeability. Additionally, a Gaussian mixture model (GMM) was used for unsupervised classification of hydraulic flow units (HFU) using the flow zone indicator (FZI), computed from the continuous permeability curve obtained from the best ML model. Thus, ML-based permeability prediction is an indispensable component of HFU modeling. The model identified three distinct flow zones, enabling HFU clustering and defining their corresponding petrophysical properties and depositional environments.

View source

Similar papers

Conference Aug 2026

Core-Calibrated Machine Learning–Based Permeability Estimation: A Western Niger Delta Case Study

Accurate permeability estimation is essential for optimum reservoir characterization, simulation, and development planning; however, its prediction remains challenging due to the limited availability of core-derived permeability measurements, often constrained by acquisition cost and data coverage. This study presents an integrated machine learning workflow for permeability prediction calibrated against core permeability data using wireline log information from wells located in western Niger Delta. The workflow involved systematic data preprocessing, depth-based alignment of core and well log data, and feature engineering to derive additional petrophysical attributes relevant to fluid flow behavior. Key input variables used in the model include Gamma Ray (GR), Bulk Density (RHOB), Resistivity, Neutron Porosity, and several derived parameters such as shale volume, resistivity index, effective porosity, neutron–density separation, and bulk volume water. Permeability values were transformed into logarithmic space to improve modeling stability and capture the wide range of permeability values typical of heterogeneous clastic reservoirs. Five (5) supervised machine learning algorithms; Random Forest (RF), Extreme Gradient Boosting (XGB), Extra Trees Model (ETM), AdaBoost (ADB) and Decision Tree (DT) were developed and evaluated using an 80–20% train–test split, with blind test well excluded for validation against measured core permeability. The Extra Trees model demonstrated the highest predictive performance, achieving an R² value of ~ 0.9164, Mean Absolute Error (MAE) of ~ 1.5392, and Root Mean Squared Error (RMSE) of ~ 4.1011, indicative of high correlation between predicted and core-measured permeability. The result visualized predicted versus actual permeability along the depth axis, providing a vertical reservoir-scale understanding of permeability distribution. The results indicate that machine learning models can effectively capture complex petrophysical relationships and provide reliable permeability estimates in intervals lacking core measurements.

Nwakanma Alexander, E. Ekpenyong, Maduabuchi Ogu · 0 citations
Conference Aug 2026

Machine Learning–Based Rate of Penetration Prediction Using Multi-Well Field Data in Deviated Wells

Reliable forecasting of drilling rate of penetration (ROP) remains a technical challenge due to its complex dependence on several operational, geological, and directional conditions that are highly nonlinear and well-specific. These difficulties are amplified in deviated wells, where traditional empirical relations often fail to capture the combined effects of lithological variability and directional changes. This study develops a machine learning (ML) model for ROP prediction using multi-well field data and evaluates candidate ML algorithms/models not only on numerical accuracy but also on their ability to reproduce reasonable trend responses to changes in key drilling parameters. The objective is to establish a validation approach that integrates statistical performance, cross-well generalization, and sensitivity behavior consistent with drilling mechanics. Field data from 18 wells were used to develop the model, while two additional wells were reserved for independent validation. A total of 91,140 drilling datasets, comprising 24 parameters/features, were collected. After preprocessing and selective outlier screening, 85,695 records were retained for model training and testing. Feature selection resulted in nine highly influential parameters: True Vertical Depth (TVD), Weight on Bit (WOB), rotational speed (Rotation), mud flow rates (FLOW), mud density, shale content (Shale), Sandstone content (Sandstone), inclination (Inc), and dogleg severity (DLS). Eleven supervised ML algorithms were evaluated, representing instance-based, tree-based, boosting-based, neural-network, and stacking-ensemble model families. Model performance was assessed using feature-importance ratings and statistical metrics, including train/test R2, MSE, RMSE, and MAE, which were used to select the best-performing model. Finally, sensitivity analysis was conducted to evaluate the selected model's ability to reproduce physically meaningful ROP responses. Model comparison across the training and testing splits demonstrated strong overall predictive performance, with training R2 values ranging from 0.91 to 0.99 and testing R2 values ranging from 0.88 to 0.94. Among the evaluated models, the Gradient Boosting (GB) model provided the best overall performance, with training R2 = 0.9940, testing R2 = 0.9399, RMSE = 36.88, and MAE = 21.52. The GB model also showed strong agreement with the averaged feature-importance ranking and reproduced physically meaningful ROP sensitivity trends for the dominantly influential drilling parameters. SHAP and feature-importance analyses confirmed that TVD, FLOW, WOB, shale content, and rotation were the most influential variables controlling ROP. External validation on two unseen wells further demonstrated strong generalization, with R2 = 0.9694 for Well-19 and R2 = 0.96 for Well-20. This study presents a physically informed framework for evaluating and selecting data-driven ML models for deviated wells. The approach combines conventional accuracy metrics with feature-importance consistency, reasonable trend prediction in sensitivity analysis, and strong agreement between model predictions and unseen-well data. Through this integrated evaluation, a model is identified that not only reproduces historical data accurately but also yields correct responses under a diverse range of varying input parameters. The proposed methodology establishes a reproducible basis for developing more reliable ROP forecasting tools for complex well trajectories.

Nayem Ahmed, Ramadan Ahmed, V. Soriano et al. · 0 citations
Conference Open access Aug 2026

Prediction of Porosity and Permeability Using Well Log and Core Data: A Data-Driven Approach

Accurate prediction of porosity and permeability is very important for reservoir characterization and hydrocarbon extraction. Traditional workflow in the form of empirical correlations is usually difficult, time-consuming, spatially limiting, and entirely dependent on formation geology. The current study examines a different approach, which uses machine learning (ML) regression models based on well-log and core data. Three regression architectures including Random Forest, CatBoost, and K-Nearest Neighbors (KNN) were trained and validated based on a dataset consisting of 340 samples of shaly sand gas reservoirs. The gamma ray (GR), resistivity (RLLD), spontaneous potential (SP), bulk density (RHOB), neutron porosity (NPHI), and depth were used as the input variables with the core-derived porosity (CPHI) and permeability (CKHG) being used as the targets. The quantitative measures of performance of the models included R2, Mean Absolute Error (MAE), and Root Mean Squared Error (RMSE). The findings indicated that KNN regression was better than its counterparts as it achieved R2 = 0.8933 and R2 = 0.9340 in terms of porosity and permeability prediction, respectively, and more acceptable metrics of errors showed. Comparatively, the traditional empirical methods showed a significantly lower accuracy rate. The findings highlight that machine learning has the potential to provide precise, scalable, and low-cost predictions of the reservoir properties which could lead to better choices for exploration and production activities.

Md Alamin Islam, Bintun Zaman, Shahria Nayem Ahmed · 0 citations
Jul 2026

A Tuned-Filtration and Supervised Machine-Learning Framework for Robust ROP Prediction and Drilling Optimization

Improving drilling efficiency remains a central objective in petroleum operations due to its direct impact on time and cost. This study develops a predictive framework for estimating the Rate of Penetration (ROP) using supervised machine learning techniques combined with systematic data conditioning. The analysis is based on more than 11,000 measurements obtained from four horizontal wells. A rigorous preprocessing strategy was implemented to enhance data reliability, including removal of invalid entries and statistical outliers using the interquartile range method. This procedure reduced the dataset to 7,297 high-quality observations. In addition, target stabilization was introduced through Exponential Moving Average smoothing (spans of 5 and 10), which reduced short-term fluctuations and improved the learnability of the ROP signal. Three tree-based regression models—Decision Tree, Random Forest, and Gradient Boosting—were evaluated under both default configurations and optimized settings. Results show that model performance is strongly influenced by data conditioning. The Random Forest model achieved the highest accuracy, with a coefficient of determination (R2) of 0.96 and a mean squared error (MSE) of 26 when trained on the EMA-10 dataset. Gradient Boosting exhibited the largest improvement from hyperparameter tuning, with R2 increasing from 0.86 to 0.95. To bridge the gap between model development and practical use, the trained models were implemented in interactive applications for real-time prediction and parameter optimization. The outcomes demonstrate that careful preprocessing and noise-aware modeling significantly enhance predictive capability.

B. Elahifar, Thomas Philip Fagerli · 0 citations