Skip to content
Conference Open access

A probabilistic data-driven framework for modelling clay consolidation and settlement behavior

2026 · E3S Web of Conferences · Vol 730, pp. 03004 · 0 citations · 6 references

Abstract

Predicting consolidation settlement in large-scale reclamation projects remains difficult because laboratory consolidation test data are sparse in space and natural marine clays show strongly non-linear structured behavior. This study develops a probabilistic data-driven framework to reconstruct spatially continuous e–log p′ and log k–e relationships from sparse consolidation test data. A deep neural network is combined with repeated K-fold cross-validation to estimate the ensemble mean response and its 95% confidence interval. The framework was trained using consolidation test data from 49 boreholes at Kobe Airport and was evaluated using two independent blind-test boreholes that were not used in model development. The predicted mean curves reproduced the main features of the observed compression and permeability responses, including depth-dependent yield behavior and post-yield changes in compressibility. The engineering applicability of the framework was examined through settlement analysis at monitoring point KC-1, where no site-specific borehole data were available. In this analysis, soil deformation was treated as one-dimensional, whereas pore-water flow was modeled as two-dimensional. The predicted material relationships were used as input. The calculated settlement history reproduced the main observed trend, while the late-stage difference from the measurements showed the importance of deeper strata outside the present modeling scope. The proposed framework provides a practical way to interpolate consolidation behavior in space with quantified ML-related uncertainty for large reclamation projects with similar geological and data conditions.

Read PDF

Similar papers

Aug 2026

Ensemble Machine-Learning-Based Horizontal Permeability Prediction and Hydraulic Flow Unit Characterization of the Hugin Sandstone

Understanding permeability is essential for evaluating reservoir quality and field development planning. Reliable permeability estimation can reduce the uncertainty in reservoir characterization, particularly in intervals where core data are limited. As the industry relies on log-based interpretations and empirical correlations, the limitations of these approaches become apparent. Data-driven approaches offer a promising alternative to conventional empirical methods. The data set in this study comprises 252 samples with seven features derived from conventional well logs. Data preprocessing includes handling missing values, smoothing logs, feature engineering to add an extra input, and transformation with the Yeo-Johnson technique. A center moving average filter was used to reduce variance and improve data consistency. Ensemble machine-learning (ML) and baseline models were developed and evaluated using a 75-25 train-test split, four-fold cross validation, and model complexity assessment. Ensemble methods outperformed baseline models, with extremely randomized trees (ET) and random forest (RF) emerging as the most stable, achieving a mean R² of 0.93 and 0.89 and a low R² standard deviation (0.3). Multilinear regression (MLR) and artificial neural networks (ANNs) show limited accuracy, while gradient boosting (GB) and extreme gradient boosting (XGBoost) methods exhibit overfitting despite perfect training scores. Predicted kh values were compared with core data. Both linear and nonlinear empirical equations were derived using MLR, polynomial regression, and a power-law model. The power-law model (empirical equation) achieved an R² value of 0.79 and can therefore be used to estimate permeability. Additionally, a Gaussian mixture model (GMM) was used for unsupervised classification of hydraulic flow units (HFU) using the flow zone indicator (FZI), computed from the continuous permeability curve obtained from the best ML model. Thus, ML-based permeability prediction is an indispensable component of HFU modeling. The model identified three distinct flow zones, enabling HFU clustering and defining their corresponding petrophysical properties and depositional environments.

Vikram Kumar, Sayantan Ghosh, S. Maiti · 0 citations
Open access Jul 2026

Physically Constrained and Location-Aware Machine Learning for Joint Prediction of Clay Compression and Recompression Indices

Compression index (Cc) and recompression index (Cur) are essential parameters in one-dimensional consolidation and settlement analysis, yet their direct determination from oedometer testing is time-consuming, costly, and often limited by sparse recompression data. This study develops an interpretable and physically constrained machine-learning framework for the joint prediction of Cc and Cur from four routinely measured index properties: liquid limit (LL), plasticity index (PI), initial void ratio (e), and natural water content (w). A curated subset of 459 natural clay records from the global CLAY/Cc/6/6203 database was used to benchmark single-output and multi-output Random Forest, gradient-boosted tree, and deep neural network models. In addition to conventional random train–test and cross-validation protocols, a leave-one-location-out validation was introduced to evaluate transferability across 81 Country–Location groups. Under the random-split setting, Cc was predicted with moderate-to-good accuracy, with baseline models achieving test R2 values of approximately 0.61–0.70 and a geotechnically enriched Random Forest model increasing the test R2 to 0.777. Cur was more difficult to predict. Although feature enrichment improved its test R2 to 0.507, location-aware validation reduced Cur performance substantially, confirming its stronger dependence on site-specific stress history, fabric, and geological structure. SHAP interpretation identified e and w as the dominant controls on Cc, while Cur exhibited weaker and more diffuse dependence on the available index properties. A physically constrained target transformation based on the bounded ratio of Cur/Cc guaranteed mechanically admissible predictions with Cur < Cc, but did not fully recover the missing information needed for accurate Cur estimation. The proposed constraint is not a governing-equation-based physics-informed model. Rather, it is a mechanically constrained target transformation that preserves the admissible relationship Cur < Cc. The results show that routine index properties can support the useful preliminary prediction of Cc, whereas Cur should be treated as a screening-level estimate unless explicit stress history descriptors are available.

Zeroual Abdelatif, A. Baghbani, Aissa Lahlouhi et al. · 0 citations
Aug 2026

In Situ Stress Prediction Using Multiple Seismic Attributes Based on the Random Forest Algorithm

In situ stress characterisation is crucial for the exploration and development of coal resources. Conventional methods based on well‐log data offer high‐accuracy point measurements but are limited to wellbore locations, whereas traditional seismic‐based predictions provide extensive spatial coverage but have lower accuracy and often fail to capture the complex, non‐linear relationships between seismic attributes and the stress field. To bridge this gap, we introduce a data‐driven approach that integrates the strengths of both data types using a random forest (RF) model. In our approach, geomechanical attributes (Young's modulus and Poisson's ratio) and geometric attributes (curvature), derived from pre‐stack inversion, serve as the model's input features. High‐resolution stress values calculated from well logs using the combined spring model serve as the training labels. The RF model, optimised via grid search and cross‐validation, demonstrates high predictive accuracy. We apply the proposed RF‐based method to a field dataset from the Daji area in Shanxi Province, North China, successfully generating a continuous three‐dimensional (3D) in situ stress volume. The resulting stress field exhibits strong spatial consistency with regional tectonic features, validating the model's accuracy and geological applicability. This study demonstrates that a machine learning framework can effectively link seismic data with well‐log‐derived reference stress labels, extending sparse reference stress information into a continuous 3D volume that reliably characterises the inter‐well stress heterogeneity. This provides a practical and effective framework for in situ stress analysis in complex geological settings.

Shiqi Peng, Suping Peng, Chuangjian Li et al. · 0 citations
Conference Aug 2026

Machine Learning–Based Rate of Penetration Prediction Using Multi-Well Field Data in Deviated Wells

Reliable forecasting of drilling rate of penetration (ROP) remains a technical challenge due to its complex dependence on several operational, geological, and directional conditions that are highly nonlinear and well-specific. These difficulties are amplified in deviated wells, where traditional empirical relations often fail to capture the combined effects of lithological variability and directional changes. This study develops a machine learning (ML) model for ROP prediction using multi-well field data and evaluates candidate ML algorithms/models not only on numerical accuracy but also on their ability to reproduce reasonable trend responses to changes in key drilling parameters. The objective is to establish a validation approach that integrates statistical performance, cross-well generalization, and sensitivity behavior consistent with drilling mechanics. Field data from 18 wells were used to develop the model, while two additional wells were reserved for independent validation. A total of 91,140 drilling datasets, comprising 24 parameters/features, were collected. After preprocessing and selective outlier screening, 85,695 records were retained for model training and testing. Feature selection resulted in nine highly influential parameters: True Vertical Depth (TVD), Weight on Bit (WOB), rotational speed (Rotation), mud flow rates (FLOW), mud density, shale content (Shale), Sandstone content (Sandstone), inclination (Inc), and dogleg severity (DLS). Eleven supervised ML algorithms were evaluated, representing instance-based, tree-based, boosting-based, neural-network, and stacking-ensemble model families. Model performance was assessed using feature-importance ratings and statistical metrics, including train/test R2, MSE, RMSE, and MAE, which were used to select the best-performing model. Finally, sensitivity analysis was conducted to evaluate the selected model's ability to reproduce physically meaningful ROP responses. Model comparison across the training and testing splits demonstrated strong overall predictive performance, with training R2 values ranging from 0.91 to 0.99 and testing R2 values ranging from 0.88 to 0.94. Among the evaluated models, the Gradient Boosting (GB) model provided the best overall performance, with training R2 = 0.9940, testing R2 = 0.9399, RMSE = 36.88, and MAE = 21.52. The GB model also showed strong agreement with the averaged feature-importance ranking and reproduced physically meaningful ROP sensitivity trends for the dominantly influential drilling parameters. SHAP and feature-importance analyses confirmed that TVD, FLOW, WOB, shale content, and rotation were the most influential variables controlling ROP. External validation on two unseen wells further demonstrated strong generalization, with R2 = 0.9694 for Well-19 and R2 = 0.96 for Well-20. This study presents a physically informed framework for evaluating and selecting data-driven ML models for deviated wells. The approach combines conventional accuracy metrics with feature-importance consistency, reasonable trend prediction in sensitivity analysis, and strong agreement between model predictions and unseen-well data. Through this integrated evaluation, a model is identified that not only reproduces historical data accurately but also yields correct responses under a diverse range of varying input parameters. The proposed methodology establishes a reproducible basis for developing more reliable ROP forecasting tools for complex well trajectories.

Nayem Ahmed, Ramadan Ahmed, V. Soriano et al. · 0 citations