Biochar is a porous, carbon-rich soil amendment that can enhance soil water retention capacity by modifying pore structure and physicochemical properties. Understanding the soil−water characteristic curve (SWCC) of biochar-amended soils is essential for evaluating their hydrological behavior and promoting the application of biochar in engineering practice. Given the demonstrated feasibility and accuracy of machine learning methods for predicting soil parameters, this study employed six machine learning models, namely, decision tree, random forest, XGBoost, LightGBM, CatBoost, and artificial neural network, to predict the SWCC of biochar-amended soils based on a constructed dataset. Feature importance analysis and partial dependence analysis were further conducted to reveal the influence patterns of key variables. The results indicate that all six models exhibit good predictive capability, with gradient boosting models (XGBoost, CatBoost, and LightGBM) performing best. Suction is the dominant factor controlling the volumetric water content variation, while soil particle-size distribution and dry density provide the physical basis for water retention. Biochar content, pyrolysis temperature, and feedstock type further modulate the water retention capacity of amended soils. Overall, the findings demonstrate that machine learning approaches can effectively predict the SWCC of biochar-amended soils and provide insights into the controlling mechanisms of soil water retention.
Despite the potential of biochar to enhance crop nutrient uptake, the complex, feedstock-dependent interactions among biochar properties, soil biological processes, and plant root traits remain poorly understood. Hence, this study employed machine learning algorithms (Linear regression and Random Forest) with advanced feature selection methods (Correlation-based Feature Subset Selection Evaluator, Correlation-Based Attribute Evaluator, Principal Components Attribute Transformer, and Stepwise Regression) to predict potassium (K+) and magnesium (Mg2+) uptake by wheat (Triticum aestivum L.) roots. A greenhouse experiment was conducted for 42 days, utilizing soils amended with four biochar treatments derived from wheat stubble, wood residues, rice husk, and corn residues (25 g biochar kg⁻¹ soil). The models were trained using a comprehensive dataset of over 40 parameters from the experiment (including variables such as soil basal respiration, root cation exchange capacity, and biochar specific surface area). The results indicated that Random Forest models achieved excellent predictive performance and outperformed Linear Regression, demonstrating superior ability to capture nonlinear relationships and feature interactions. Feature selection revealed distinct mechanisms: K+ uptake was primarily driven by soil basal respiration, root adenosine triphosphate (ATP) content, and biochar surface properties (zeta potential and Brunauer-Emmett-Teller (BET) surface area), highlighting microbial mobilization and energy-dependent active transport. In contrast, Mg2+ uptake was driven predominately by biochar oxygenation, biochar Mg2+ content, root cation exchange capacity, and carboxyl groups on root cell walls (yielding a predictive CC of 0.89 and MAE of 0.71). The significance of biochar's intrinsic Mg2+ content suggests a direct nutritional contribution from the amendment itself, which is a fundamentally different pathway than the indirect mechanisms highlighted for K+. Biochar efficacy was highly feedstock-dependent, with wheat stubble and wood residue biochars outperforming rice husk and corn residue biochars. These quantitative results provide specific mechanistic insights into biochar-mediated variation of cationic nutrition and demonstrate the power of machine learning for unraveling complex rhizosphere dynamics.
K. Ghassemi-Golezani, Salar Farhangi-Abriz, S. Rahimzadeh· Scientific Reports· 0 citations
Soil fertility is a crucial aspect of agricultural productivity and sustainability, determines the soil's capacity to provide essential nutrients necessary for plant growth and development. This study focuses on the analysis and prediction of soil fertility using ensemble learning techniques in the South Gondar Zone. By examining various soil parameters, including macro and micronutrients, soil structure, pH, and organic matter content, the research aims to develop predictive models that accurately assess soil fertility levels. The objective of this study is to analyze and predict soil fertility using machine learning techniques, specifically targeting the south Gondar Zone, in the Amhara region of Ethiopia. The dataset for this study comprises 20,168 instances, including both fertile and non-fertile samples, with 17 selected attributes. Several machine learning models were evaluated on both the original and SMOTE-balanced datasets. The models included Random Forest, AdaBoost, and XGBoost classifiers. I applied Random Forest classifier consistently demonstrated the highest performance, with testing accuracies of 94.55% on the original dataset and 94.37% on the SMOTE-balanced dataset. AdaBoost also showed strong performance, achieving testing accuracies of 94.6% and 94.31% on the original and SMOTE-balanced datasets, respectively. XGBoost performed well but was slightly less accurate compared to the ensemble methods. However, XGBoost showed robust performance on both datasets, with testing accuracies of 94.04% and 94.27%. Feature importance analysis using the Random Forest classifier has shown that factors such as Clay, CEC, CaCO
3
, Sand, and Mn significantly impact soil fertility prediction. as a conclusion Random Forest classifier, an ensemble-based learning technique, was the most reliable and accurate model for predicting soil fertility. These results highlight the importance of specific soil properties in determining fertility and can guide targeted soil management practices to improve agricultural productivity.
Tigist Tewabe· American Journal of Robotics...· 0 citations
Accurate prediction of shale methane adsorption capacity is crucial for reservoir evaluation. This study integrates 486 experimental datasets to develop a multivariate machine learning prediction model. Six key geological parameters, including depth, total organic carbon (TOC), moisture, porosity, vitrinite reflectance (
Ro
), and clay minerals, were selected as features. Correlation analysis methods were used to reveal the nonlinear relationships between various geological parameters and methane adsorption capacity, as well as the intrinsic coupling structures among these parameters. The predictive performance of three machine learning models, namely, Support Vector Machine (SVM), Random Forest (RF), and Extreme Gradient Boosting (XGBoost), was systematically compared. The results show that TOC and
Ro
are the significant factors controlling methane adsorption. The XGBoost model achieved superior performance on both training and test sets, demonstrating higher prediction accuracy and stronger generalization capability compared to SVM and RF. This study provides novel insights and methodologies for assessing shale gas adsorption potential and optimizing development strategies.
Hongjian Zhu, Ning Zhang, Zongquan Hu et al.· Frontiers in Earth Science· 0 citations
Soil organic carbon (SOC) is a key indicator of cropland quality, soil fertility, and the carbon sequestration potential of agroecosystems. Accurate characterization of its spatial distribution is essential for black soil conservation and regional soil carbon management. However, regional-scale SOC prediction is often constrained by limited field observations, which can reduce model generalizability and predictive reliability. In this study, we developed a limited-sample SOC prediction framework for the Liaohe Plain using 310 surface (0–20 cm) soil samples collected in 2025 and multi-source environmental covariates, including climate, vegetation, soil spectral, and terrain variables. The framework used the Tabular Prior-Data Fitted Network (TabPFN), whose performance was compared with that of Random Forest, Support Vector Machine, CatBoost, K-Nearest Neighbors, and XGBoost. Model performance was evaluated using 100 repetitions of random 80:20 holdout validation and repeated five-fold spatial cross-validation based on spatially constrained clustering, while sampling-induced relative uncertainty was quantified using 100 repeated random sampling and model-fitting runs. Under random holdout validation, TabPFN showed competitive predictive performance, with mean R2 and RMSE values of 0.608 ± 0.038 and 4.026 ± 0.184 g kg−1, respectively. Repeated spatial cross-validation yielded more conservative performance estimates, with mean R2 and RMSE values of 0.535 ± 0.072 and 4.33 ± 0.31 g kg−1, respectively, indicating that random splitting may overestimate model performance when sampling sites are spatially clustered. Spatial prediction showed that cropland SOC ranged from 4.63 to 27.04 g kg−1, with generally lower values in the west and higher values in the northeast. Areas with high sampling-induced relative uncertainty were mainly concentrated in the northern, northeastern, and marginal regions. These findings provide a methodological basis for SOC mapping, supplementary sampling optimization, and regional soil carbon management under limited-sample conditions, although the temporal robustness of the results requires confirmation using independent data from additional years.
Yongqiang Yang, Yanzhi Zhao, Shuang Gang et al.· Agronomy· 0 citations
Understanding the environmental fate of microplastics (MPs) in agricultural soils remains a major challenge, particularly under field conditions where soil structure and hydraulic processes jointly regulate particle transport and retention. This study investigated whether hydro-physical soil functioning can explain the distribution and accumulation of MPs in pistachio orchard soils from a semi-arid region of southeastern Türkiye. A total of 42 soil samples were analyzed for MP abundance, size distribution, and morphology, together with key hydro-physical properties including texture, porosity, bulk density, aggregate stability, organic matter content, and soil water retention characteristics. To identify the dominant controls on MP occurrence, explainable machine learning approaches combining Random Forest (RF), Gradient Boosting Decision Trees (GBDT), and SHAP (SHapley Additive exPlanations) analysis were employed. Microplastic abundance differed among management systems. Former landfill or construction sites represented the largest proportion of the total recorded microplastic abundance (40.9%), followed by conventionally managed (25.2%), manure-amended (24.5%), and sewage-sludge-amended orchards (9.4%). Median microplastic abundances were 1433, 667, 4633, and 633 particles kg−1 soil, respectively. Fine-sized MPs constituted the dominant particle fraction and exhibited strong associations with pore-system characteristics, indicating that pore-size compatibility governs their retention and mobility within the soil matrix. Morphology-specific analyses further revealed contrasting relationships between soil hydro-physical properties and individual MP forms, suggesting distinct retention pathways for granules, films, fragments, and fibers. Explainable AI analysis identified organic matter, silt content, bulk density, and water retention characteristics as the most influential predictors of MP occurrence. Among the tested models, RF demonstrated superior predictive robustness and generalization capacity. The findings demonstrate that hydro-physical soil functioning plays a central role in determining microplastic fate in agricultural soils and highlight the value of interpretable machine learning frameworks for uncovering the mechanisms underlying contaminant retention and redistribution. Integrating soil structural indicators with explainable artificial intelligence offers a promising pathway for improving microplastic risk assessment in agroecosystems.
Kubra Polat, H. Günal, Murat Birol et al.· Land· 0 citations