Skip to content
Open access

Tunnel Water Inflow Prediction Using CatBoost and Comparative Hyperparameter Optimization Strategies

Jul 2026 · Applied Sciences · Vol 16, pp. 6882 · 1 citation · 30 references

Abstract

Accurate prediction of tunnel water inflow in water-rich fault zones is important for groundwater control design and construction risk prevention. In this study, a per-linear-meter tunnel water inflow database containing 425 valid samples was established through orthogonal numerical simulations based on a three-dimensional steady-state seepage model with a grouting ring. The input variables included four hydraulic and grouting parameters and two excavation-position descriptors, namely the excavation-position distance and excavation-position category, thereby reflecting both the water-blocking effect of grouting reinforcement and the spatial variation in water inflow as the excavation face approached the fault zone. Considering that the samples were generated from 25 orthogonal simulation cases at different excavation positions, grouped validation was adopted to reduce information leakage at the simulation-case level. Four baseline machine learning models, including SVM, RF, XGBoost, and CatBoost, were evaluated using ten repeated grouped hold-out validations. CatBoost achieved the best overall baseline generalization performance, with an average test R2 of 0.6209 ± 0.0405, MAE of 0.1084 ± 0.0079, and RMSE of 0.1555 ± 0.0085. CatBoost was therefore selected for further hyperparameter optimization. Subsequently, random search, Bayesian optimization, the Osprey Optimization Algorithm, and the Grey Wolf Optimizer were compared under the same search space and computational budget. Hyperparameter optimization was conducted only within the training set using grouped cross-validation, and the independent grouped test set was used only for final evaluation. The results showed that the unoptimized CatBoost model achieved the best overall balance between prediction accuracy, stability, and computational efficiency. Although RS-CatBoost slightly improved MAE and MAPE among the optimized models, none of the optimization strategies consistently outperformed the unoptimized CatBoost baseline, indicating that the choice of hyperparameter optimization algorithm played a secondary role under the current dataset and grouped-validation framework. The proposed framework is intended as a preliminary modeling reference under controlled numerical simulation conditions, and its practical engineering reliability requires further validation using field monitoring data or independent benchmark cases.

Read PDF

Similar papers

Aug 2026

Ensemble Machine-Learning-Based Horizontal Permeability Prediction and Hydraulic Flow Unit Characterization of the Hugin Sandstone

Understanding permeability is essential for evaluating reservoir quality and field development planning. Reliable permeability estimation can reduce the uncertainty in reservoir characterization, particularly in intervals where core data are limited. As the industry relies on log-based interpretations and empirical correlations, the limitations of these approaches become apparent. Data-driven approaches offer a promising alternative to conventional empirical methods. The data set in this study comprises 252 samples with seven features derived from conventional well logs. Data preprocessing includes handling missing values, smoothing logs, feature engineering to add an extra input, and transformation with the Yeo-Johnson technique. A center moving average filter was used to reduce variance and improve data consistency. Ensemble machine-learning (ML) and baseline models were developed and evaluated using a 75-25 train-test split, four-fold cross validation, and model complexity assessment. Ensemble methods outperformed baseline models, with extremely randomized trees (ET) and random forest (RF) emerging as the most stable, achieving a mean R² of 0.93 and 0.89 and a low R² standard deviation (0.3). Multilinear regression (MLR) and artificial neural networks (ANNs) show limited accuracy, while gradient boosting (GB) and extreme gradient boosting (XGBoost) methods exhibit overfitting despite perfect training scores. Predicted kh values were compared with core data. Both linear and nonlinear empirical equations were derived using MLR, polynomial regression, and a power-law model. The power-law model (empirical equation) achieved an R² value of 0.79 and can therefore be used to estimate permeability. Additionally, a Gaussian mixture model (GMM) was used for unsupervised classification of hydraulic flow units (HFU) using the flow zone indicator (FZI), computed from the continuous permeability curve obtained from the best ML model. Thus, ML-based permeability prediction is an indispensable component of HFU modeling. The model identified three distinct flow zones, enabling HFU clustering and defining their corresponding petrophysical properties and depositional environments.

Vikram Kumar, Sayantan Ghosh, S. Maiti · 0 citations
Open access Jul 2026

Prediction of In-Situ Properties of Coastal Soils Using GWO-XGBoost Model

In coastal port infrastructure, accurate prediction of soil profiles and Standard Penetration Test (SPT) N-values at intermediate borehole locations is critical for safe, economical and resilient foundation design, as across all three dimensions subsoil conditions can vary significantly over short distances since soil is heterogeneous, anisotropic, and unpredictable material. In order to predict in-situ properties at intermediate points, the application of Machine Learning models is necessary which would save time and cost required during the preliminary design and detailed planning phases. In order to maximize predictive accuracy and to navigate complex hyperparameter search spaces, a metaheuristic - optimized approach like the Grey Wolf Optimizer (GWO) - eXtreme Gradient Boosting (XGBoost) hybrid model was used, which is superior to classical interpolation techniques and conventional Machine Learning models such as Random Forest and XGBoost which suffer from limited generalization on sparse geotechnical datasets and suboptimal hyperparameter selection. A dataset of 385 borehole records from 72 geotechnically investigated locations, spanning depths of 0.5 m to 87 m of three major deep-sea port development sites, namely Machilipatnam, Ramayapatnam, and Durgarajpatnam, Andhra Pradesh, India, was compiled and processed using systematic data cleaning, geotechnical imputation, spatial feature engineering, and three normalization strategies Z-Score Standardization, Min-Max Scaling, and Robust Scaling in order to predict continuous SPT N-values and categorical soil profiles simultaneously at unsampled locations. The GWO algorithm optimized XGBoost hyperparameters including learning rate, maximum depth, estimator count, and L1/L2 regularization coefficients. The classical interpolation methods failed critically, with Inverse Distance Weighting (IDW) yielding R² = 0.1005 and Radial Basis Function (RBF) producing negative R² values. Optimized XGBoost improved performance to RMSE = 3.7072 and R² = 0.9315 and Random Forest achieved 87.32% soil classification accuracy and R² = 0.7601. The GWO–XGBoost model with Min-Max scaling confirmed strong generalization, attaining 100% primary soil type classification accuracy and R² = 0.9384, with a robust cross-validation score of 0.9095 ± 0.0195. The GWO–XGBoost framework offers a cost-effective tool with high-accuracy, for detailed subsurface characterization, which align with UN Sustainable Development Goals SDG 9 (Industry, Innovation and Infrastructure) focussing on target 9.1 (Resilient Infrastructure) as prediction would help in preventing failures due to complex environmental conditions and for designing resilient coastal infrastructure. Under SDG 9, target 9.4 (stainable Industrialization/Innovation) is addressed as Grey Wolf Optimization + XGBoost is a data-driven, innovative approach that improves sustainability and engineering efficiency compared to traditional field testing which is carbon-intensive.

Sail S · 0 citations
Jul 2026

Machine Learning Prediction of the Ground Reaction Curve in Sand with the MATLAB GUI Platform

The proposed framework combined a curated database, neural network-based curve prediction, and hyperparameter optimization, providing a robust approach for evaluating the soil arching effect, providing a robust approach for evaluating the soil arching effect.

Cheng-shuang Yin, Liu-mei Wei, Han-lin Wang et al. · 0 citations
Open access Jul 2026

Deformation Prediction Model for Soft Rock Tunnels Based on NWOA-LSTM Model

Surrounding rock deformation in soft rock tunnels is controlled by complex nonlinear interactions among geological conditions, construction parameters, and support measures, making accurate prediction challenging. In this study, a project-scale deformation prediction dataset was constructed using field monitoring data from a sandy shale tunnel project. Eight engineering factors were selected as input variables, including excavation method, initial support strength, closure time, tunnel burial depth, lithology, rock integrity, groundwater condition, and the relative orientation between major structural planes and the tunnel. A hybrid prediction framework integrating a novel whale optimization algorithm (NWOA) and a long short-term memory (LSTM) network was developed. The proposed NWOA improves the standard whale optimization algorithm by introducing a nonlinear convergence strategy, an adaptive weight coefficient, and a dynamic spiral position updating mechanism to enhance the hyperparameter search process of the LSTM model. Model performance and stability were further assessed using repeated and nested cross-validation. The corresponding RMSEs were 0.2294 ± 0.0734 and 0.2219 ± 0.0751 percentage points, the MAEs were 0.1427 ± 0.0350 and 0.1505 ± 0.0585 percentage points, and the R2 values were 0.8660 ± 0.0548 and 0.8732 ± 0.0599, respectively. These comparable results support project-specific predictive performance for the investigated tunnel sections.

Fanmeng Kong, Bo Wang, Xin Li et al. · 0 citations
Open access Jul 2026

Tunnel Water Inflow Prediction and Uncertainty Quantification Using Vine Copula-Coupled Sparse Polynomial Chaos Expansion

Accurate prediction and uncertainty quantification of tunnel water inflow are critical for construction safety, risk mitigation, and groundwater-control planning. However, conventional analytical and numerical methods are often limited by simplified assumptions and high computational cost, while many machine learning models lack reliable uncertainty quantification. To address these limitations, this study proposes a framework by coupling vine copula dependence modeling with sparse polynomial chaos expansion (SPCE). The framework utilizes a vine copula to characterize the asymmetric dependence among input parameters. Two distinct approaches are used to construct the SPCE models: the arbitrary polynomial chaos expansion (aPCE) method, assuming independence in the original space, and the Rosenblatt transform-based polynomial chaos expansion (Rt-PCE) method, which maps correlated inputs into an independent space via the Rosenblatt transform to establish rigorous orthogonal polynomials. Validation using a database of 600 cases shows that SPCE models achieve point accuracy comparable to artificial neural network (ANN) and Gaussian process regression (GPR) with superior numerical stability. Notably, Rt-PCE yields the best predictive robustness and outperforms both benchmarks in probability density fitting, particularly in capturing extreme tail behavior. Furthermore, the study confirms that neglecting input dependence biases probability estimations, whereas vine copula-based modeling effectively captures both the central tendency and tail features of inflow distributions. The proposed framework provides decision support for resource-efficient intervention planning under uncertain hydrogeological conditions.

Jing Qian, Zhihao Zhao, You Dong et al. · 0 citations
Conference Aug 2026

Development and Validation of an Empirical Correlation for Water Production Forecasting in Petroleum Reservoirs

Accurate water production prediction is critical for field development optimization, surface facility design, and production management in mature oil reservoirs. This study develops empirical correlations for forecasting Water-Oil Ratio (WOR) using 9,161 production records from seven Volve Field wells (Norwegian North Sea, 2007–2016). Four approaches were evaluated: multiple linear regression, power law correlation, polynomial regression, and an exponential model, benchmarked against established methods including the X-Plot, Ershaghi-Omoregie, Buckley-Leverett, Arps decline curve, and Chan diagnostic techniques. Feature engineering generated derived variables including cumulative oil production, pressure ratio, production time, gas-oil ratio, and productivity index. After removing non-physical values and treating extreme WOR observations, data were split 80/20 for training and validation. The power law correlation achieved the strongest test-set performance (R2 = 0.845, RMSE = 2.374, MAE = 1.065), expressing WOR as a function of cumulative oil production, pressure ratio, and production time. It outperformed all conventional benchmarks, with the Arps decline-based method representing the best traditional comparator but at substantially lower accuracy. These results demonstrate that simple empirical correlations, when derived from high-quality datasets, can reliably forecast water production behavior. The proposed correlation provides a practical, easily implemented tool for production forecasting, water handling capacity planning, and operational decision-making within standard reservoir engineering workflows.

E. Echikwau, M. Mba, M. M. Ekereke et al. · 0 citations