Aug 2026· Applied Sciences· Vol 16, pp. 7859· 0 citations· 34 references
Abstract
Accurate prediction of input–output relationships in natural gas exploration is essential for improving exploration efficiency and optimizing investment allocation. However, this task is severely hindered by data sparsity and strong nonlinear characteristics inherent in oil and gas exploration systems, rendering conventional statistical methods and single machine learning models ineffective. This study develops a novel integrated framework combining data augmentation, nonlinear feature engineering, and ensemble learning to achieve accurate and interpretable prediction of exploration input–output matching under limited data constraints. Taking four core exploration indicators—including the number of exploration wells, total drilling depth, reserve abundance, and proven reserves—as input variables, adaptive prediction models were constructed for seven typical hydrocarbon basin exploration systems. To ensure comprehensive algorithmic exploration, nine advanced algorithms, including mainstream ensemble methods (RandomForest), few-shot neural networks (FewShot_NN), and kernel-based regressions, were systematically benchmarked. Furthermore, SHapley Additive exPlanations (SHAPs) was adopted to enhance model interpretability, and non-parametric Wilcoxon signed-rank tests were introduced to rigorously validate statistical significance. The results demonstrate that the optimal predictive pathway varies across different geological systems. Specifically, RandomForest and GBDT exhibit superior performance in systems with moderate heterogeneity (e.g., Jialingjiang and Changxing–Feixianguan Formations), whereas FewShot_NN and Kernel Ridge achieve the highest accuracy under extreme data sparsity and volatility (e.g., Xujiahe Formation and Lower Permian). The established framework yields a coefficient of determination (R2) greater than 0.96 for the majority of study cases, with overall absolute percentage errors heavily minimized. SHAP analysis further verifies that drilling depth and reserve abundance are the dominant controlling factors. This data-driven framework provides a robust and interpretable technical tool for the intelligent management of energy resources.
These findings demonstrate that energy efficient model cascades require evaluation beyond clean accuracy, with explicit attention to routing reliability under distribution shift, and select a model cascade at the pareto-optimum of accuracy, routing quality, and energy consumption that achieves competitive predictive pe...
Pallavi Mitra, J. Kushwaha, F. Biessmann· 0 citations
Concerns about the dependability and credibility of prediction outputs have grown as a result
of the expanding use of machine learning (ML) systems in high-stakes industries like
healthcare, finance, autonomous systems, and public governance. Uncertainty estimate is still
somewhat underemphasized, despite its crucia...
Precious Chidum Amadi· International Journal of Com...· 0 citations
Adaptive representation selection can improve a constrained-observation robustness-efficiency frontier in matched benchmark settings, but does not universally dominate trace baselines.
Active feature acquisition learns policies that sequentially acquire features to maximize information about a target variable. We study how to learn and evaluate such policies from finite offline data using prior-data fitted networks (PFNs), which are off-the-shelf models that output posterior predictive distributions...
Y. Kobayashi, Divyam Madaan, Shalmali Joshi· 0 citations
A structured narrative review using predefined searches of Web of Science Core Collection and Scopus to synthesize empirical XAI applications across six water-research domains, with emphasis on method selection, model and data compatibility, explanation reliability, and operational implementation.