Skip to content

Author

D. Rajapakshe

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Assessing Features of PFAS Groundwater Occurrence Using SHAP-Enhanced Machine Learning.

Per- and polyfluoroalkyl substances (PFAS) are persistent groundwater contaminants that pose long-term risks to water sources and public health. Predicting PFAS occurrence remains challenging due to high-dimensional environmental data and limited model interpretability. In this study, we benchmarked explainable machine learning models to predict PFAS occurrence in groundwater and to identify the features strongly associated with estimated PFAS occurrence. A comprehensive dataset of 12,406 groundwater characterization records collected between 2001 and 2019 was compiled with 172 explanatory features to describe PFAS source proximity, land use, hydrogeology, soil properties, meteorology, and sampling sites characteristics. Four tree-based ensemble classifiers, including Random Forest, XGBoost, LightGBM, and CatBoost, were evaluated under multiple classification schemes. Binary classification, which split total PFAS concentrations into 'low' and 'high' categories based on the median value, achieved the most robust and generalizable performance, with testing accuracy above 94% and well-established precision-recall curves. Increasing the number of classification bins degraded performance, particularly for intermediate bins, highlighting intrinsic separability limits in PFAS occurrence data rather than model deficiencies. Model interpretability was addressed using SHapley Additive exPlanations (SHAP), which revealed that sampling year and proximity to major PFAS sources were the dominant predictors across all models. Additional contributions were attributed to proximity to other PFAS sources, soil texture, hydrologic features, land use, and precipitation patterns. SHAP interaction analyses further revealed model-learned temporal variation in the attribution of source-proximity predictors. These patterns are interpreted as hypothesis generating model associations that may be related to regulatory changes, evolving monitoring strategies, analytical-era differences, and possible secondary-source influences, rather than as confirmation of specific environmental processes. Collectively, this study achieved interpretable prediction for PFAS occurrence in groundwater, and offered a transparent, data-driven framework to inform PFAS risk assessment.

Lin Wang, Yun Ma, D. Rajapakshe et al. · 0 citations