Skip to content
Open access

Novel Interpretable Machine Learning Models for Predicting Compressive Strength of Nano-Silica Concrete

Aug 2026 · Engineering Research Express · 0 citations

TL;DR

This study provides a robust, interpretable, and generalizable ML framework for optimizing nano-silica concrete mix design and highlights the strong potential of ML, particularly ensemble models combined with explainable AI techniques, to improve prediction reliability, reduce trial-and-error experimentation, and support more cost-efficient and sustainable concrete design.

Abstract

This study presents a comprehensive comparative analysis of several machine learning (ML) models for predicting the compressive strength (CS) of nano-silica (NS)-enhanced concrete. A large dataset comprising 724 experimental mix designs was compiled from the literature, including various variables such as cement content, water-to-binder ratio, fine and coarse aggregates, nano-silica content, superplasticizer content, and curing time. Six ML algorithms were developed and evaluated: Interaction Model, Full Quadratic (FQ), Artificial Neural Network (ANN), M5P-Tree, Gradient Boosting (GB), and Random Forest (RF). Model performance was assessed using R², RMSE, MAE, scatter index (SI), and objective value (OBJ). Among all models, the RF model achieved the highest predictive accuracy, followed by ANN and GB models. Sensitivity analysis revealed curing time as the most influential factor, while partial dependence plots exhibited the nonlinear effect of nano-silica quantity, with an optimal strength response about 15 kg/m3. In addition, SHAP (SHapley Additive exPlanations) analysis was employed to enhance model interpretability, confirming the dominant influence of curing age and water-to-cement ratio on compressive strength prediction. Compared with many previous studies relying on limited datasets or single-model approaches, this study provides a robust, interpretable, and generalizable ML framework for optimizing nano-silica concrete mix design. The findings highlight the strong potential of ML, particularly ensemble models combined with explainable AI techniques, to improve prediction reliability, reduce trial-and-error experimentation, and support more cost-efficient and sustainable concrete design.

Read PDF

Similar papers

Open access Jul 2026

Optimization and predictive measurement of compressive strength of iron ore slag modified concrete using data-driven supervised machine learning algorithms.

This study develops an integrated machine learning-experimental framework to predict the compressive strength (CS) of concrete incorporating ternary industrial wastes glass powder, marble powder, and iron ore slag. For this purpose, a dataset comprising 366 mix ratios and corresponding CS values was compiled from various sources for analysis. Advanced machine learning (ML) algorithms, including extreme gradient boosting (XGB), gradient boosting, and random forest (RF), were employed alongside hybrid techniques such as XGB-GBR and XGB-RF to evaluate the influence of these materials on strength. Based on the outcomes of the analysis, the hybrid XGB-GBR model demonstrates the highest balanced performance for both training (R2 = 0.911) and testing (R2 = 0.869) data sets. For validating the ML modeling and developing an interactive graphical user interface (GUI), experimental evaluation of CS and scanning electron microscopy was conducted. Additionally, feature importance modeling and optimization identified curing age and coarse aggregate as the most influential factors that would impact the model prediction. The contribution of this research lies in the combined modeling and experimental evaluation of a ternary waste concrete system, along with the development of a GUI. This deployable GUI will enhance the industrial applicability of ML-based concrete optimization by reducing material costs, minimizing trial batching, and supporting sustainable mix design practices.

Md. Samsuzzaman Sobuz, Md. Kawsarul Islam Kabbo, Abdullah Alzlfawi et al. · 0 citations
Open access Aug 2026

Innovative sustainable concrete with waste glass materials: an explainable machine learning for compressive strength prediction

Sustainable concrete incorporating waste glass materials has emerged as a promising solution to reduce environmental impacts associated with cement production and natural aggregate depletion. Accurate prediction of compressive strength (CS) is essential for optimizing such mixtures and ensuring structural reliability. In this study, five machine learning models: Random Forest (RF), K-Nearest-Neighbors (KNN), Adaptive-Boosting (AdaBoost), Light Gradient Boosting Machine (LightGBM), and Extreme-Gradient-Boosting (XGBoost) were developed and optimized using Grid Search to predict the CS of concrete containing glass powder (GP) and glass sand (GS). A dataset of 270 experimental samples was utilized, incorporating eight input parameters, including curing duration, cement content, GP, GS, water, density, sand, and basalt. Among the models, LightGBM demonstrated superior predictive performance during testing, achieving a determination coefficient (R² = 0.964) and Root-Mean-Square-Error (RMSE = 2.05 MPa), followed by XGBoost (R² = 0.955) and RF (R² = 0.953). In contrast, KNN and AdaBoost exhibited comparatively lower performance. SHAP and Partial Dependence Plot (PDP) analyses identified curing duration and cement content as the most influential parameters, while water and GP exhibited negative effects on CS. To enhance practical applicability, the LightGBM model was deployed through a user-friendly GUI, enabling rapid and reliable prediction of CS and providing an accurate, interpretable, and practical decision-support tool for sustainable concrete mixture design.

Abdelrahman Shams, S. R. Wani, Eman Mousa et al. · 0 citations
Open access Jul 2026

Data-driven modeling of compressive strength in sustainable self-compacting concrete incorporating recycled aggregates using ensemble learning techniques

Abstract This study develops a robust framework for estimating the compressive strength of self-compacting concrete (SCC) incorporating recycled aggregates using supervised machine learning (ML) techniques. A comprehensive experimental database comprising 582 concrete mix designs was used, encompassing diverse input variables including binder content, water, coarse and fine aggregates, recycled aggregate proportion, superplasticizer dosage, and curing time. Seven ML algorithms—XGBoost, CatBoost, AdaBoost, Extra Trees, Bagging Regressor, K-Nearest Neighbors, and Radius Neighbors—were systematically trained using a stratified 70/15/15 data split and optimized via grid search with five-fold cross-validation. Model performance was evaluated using coefficient of determination (R 2), root mean squared error, and MAE across training, validation, and testing datasets. Among all models, XGBoost demonstrated the highest accuracy, achieving an average R 2 of 0.9799, RMSE of 2.87 MPa, and mean absolute error of 1.97 MPa. The Permutation Feature Importance analysis revealed that binder content, water, and coarse aggregate were the most influential predictors of strength. This study confirms that ensemble ML models, particularly XGBoost, can reliably predict the compressive strength of SCC with recycled aggregates, while offering transparent insights into material behavior. The results provide a valuable tool for sustainable mix design optimization and practical implementation in eco-efficient concrete construction.

A. Khan, M. D. Rasheed, Muhammad Huzaifa Naveed et al. · 0 citations
Open access Jul 2026

Comparative Analysis of AI and Statistical Models for Predicting Mechanical and Durability-Related Properties of Alkali-Activated Recycled Aggregate Concrete

Alkali-activated recycled aggregate concrete (AARAC) offers a sustainable alternative to traditional concrete but suffers from complex, non-linear mechanical behavior that challenges conventional prediction methods. This study develops and compares five machine learning models, linear regression (LR), M5P, Random Forest (RF), K-Nearest Neighbors (KNN) and XGBoost, for predicting the compressive strength (Cs), flexural strength (Fs), splitting tensile strength (Ss), pull-out bond strength (PT), and water absorption (Wa%) of AARAC. A dataset of 360 experimental samples, incorporating natural aggregate, recycled concrete aggregate (RCA), cement block aggregate (CBA), water-to-cement ratio (W/C), alkaline treatment status, and slump, was used. Models were evaluated via train/test split (80/20) and 10-fold cross-validation using R2, MAE, RMSE, and MAPE. Random Forest achieved the highest test R2 (0.8736) and lowest test MAPE (1.418%) and XGBoost (R2 = 0.8605, MAPE = 1.557%). KNN and M5P performed moderately, while LR was the weakest (R2 = 0.6958, MAPE = 2.147%). All tree-based models exhibited overfitting, with training R2 up to 0.98. Scatter plot analysis revealed systematic underprediction by RF for Cs (constant offset of ~2 MPa) and increasing bias for PT, Ss, and Wa% at higher values. XGBoost gave perfect predictions for PT and Wa% but underpredicted Cs and Fs. K-fold cross-validation confirmed XGBoost as the most robust (mean R2 = 0.9844). Correlation analysis showed W/C strongly increases Wa% (r = 0.80) and decreases PT (r = −0.73); RCA negatively affects mechanical properties, while CBA and alkaline treatment improve them. The study concludes that ensemble tree models, particularly Random Forest, are superior for AARAC prediction, but systematic bias requires post hoc calibration.

Ahmed D. Almutairi, Abd Al-Kader A. Al Sayed · 0 citations
Open access 2026

Machine Learning Approaches for Predicting Compressive Strength of Concrete: A Comparative Performance Analysis

Accurate prediction of concrete compressive strength is essential for effective mix design, quality control, and structural performance assessment. Conventional empirical models often exhibit limited accuracy due to the complex and nonlinear interactions among concrete constituents. This study investigates the applicability of several machine learning models for predicting the compressive strength of concrete using a publicly available experimental dataset comprising 1030 concrete mixtures. Linear regression was adopted as a baseline model and compared with support vector regression, random forest regression, and artificial neural networks. The performance of machine learning models was meticulously assessed using the coefficient of determination, root mean square error, and mean absolute error. Additionally, the models underwent five-fold cross-validation to evaluate their robustness and generalization capabilities. The results unambiguously demonstrate that machine learning models significantly outperform linear regression models. Cross-validation results confirm the stability and reliability of the developed models. Feature importance analysis reveals that curing age and cement content are the most influential parameters affecting compressive strength, followed by water content, which is consistent with established concrete material behavior. The findings demonstrate that machine learning models, particularly random forest regression, can serve as effective supporting tools for preliminary concrete mix design and performance evaluation.

S. Rouabah · 0 citations
Open access Jul 2026

Machine Learning-Based Compressive Strength Prediction and Multi-Objective Optimization of Ultra-High Performance Concrete

The compressive strength of ultra-high-performance concrete (UHPC) is jointly influenced by multiple factors, including material composition, mixture proportion parameters, and curing regime. Conventional empirical methods are therefore insufficient to accurately characterize the highly nonlinear relationships involved. To improve the prediction accuracy of UHPC compressive strength and to achieve mixture proportion optimization that simultaneously considers mechanical performance, economic efficiency, and environmental impact, this study developed random forest (RF), artificial neural network (ANN), gradient boosting decision tree (GBDT), and extreme gradient boosting (XGBoost) models based on 810 publicly available UHPC experimental datasets. Model performance was evaluated using R2, RMSE, MAE, and MAPE. To enhance the robustness of model validation, repeated K-fold cross-validation, sensitivity analysis with different random seed splits, and benchmark model comparisons were further introduced. The results indicate that the XGBoost model achieved superior predictive performance on both the test set and robustness validation, with test-set R2, RMSE, MAE, and MAPE values of 0.9604, 7.77, 5.58, and 4.80, respectively. The model was further interpreted using SHAP, PDP, and ICE methods, and the results revealed that curing age, fiber content, silica fume content, and water-to-binder ratio were important variables affecting the compressive strength of UHPC. Furthermore, XGBoost was used as a surrogate model and coupled with NSGA-II and TOPSIS methods for multi-objective optimization. Under the constraints of compressive strength, water-to-binder ratio, superplasticizer-to-binder ratio, and absolute volume, a computationally recommended UHPC mixture proportion balancing strength, cost, and carbon emissions was obtained. This study provides a reproducible machine-learning-assisted approach for UHPC compressive strength prediction and low-carbon, cost-effective mixture proportion design.

Rong Li, Teng Zhou, Siyu Lu et al. · 0 citations