Skip to content
#explainable ai Open access

Integrasi Explainable AI (SHAP) Pada Model Machine Learning Untuk Analisis Faktor Penentu Kualitas Fasilitas Kesehatan Berbasis Web

Aug 2026 · BETRIK · 0 citations · 2 references

TL;DR

This study aims to integrate an Explainable Artificial Intelligence (XAI) approach using SHapley Additive exPlanations (SHAP) into an XGBoost model developed in Google Colab and deployed as an interactive web dashboard via Streamlit, providing intuitive clinical and managerial transparency for public health planning.

Abstract

The quality and equitable distribution of healthcare facilities are vital indicators of regional public service development. Although modern Machine Learning models such as Extreme Gradient Boosting (XGBoost) achieve exceptionally high predictive performance in classifying medical facility quality levels, their black-box nature often limits decision-making transparency for policymakers. This study aims to integrate an Explainable Artificial Intelligence (XAI) approach using SHapley Additive exPlanations (SHAP) into an XGBoost model developed in Google Colab and deployed as an interactive web dashboard via Streamlit. The dataset comprises aggregated BPJS Healthcare Facility data across Indonesian regencies/cities supplemented with operational indicators. The experimental results demonstrate superior modeling performance, achieving an accuracy of 95.15% and an F1-Score of 97.30%. Through SHAP Summary Plot analysis, Skor_Fasilitas_UGD and Total_Faskes were identified as the primary dominant factors driving high-quality facility classifications. The Streamlit web application successfully visualizes individual feature contributions (SHAP Waterfall Plot) in real-time, providing intuitive clinical and managerial transparency for public health planning

Read PDF

Similar papers

Open access Aug 2026

Pembuatan Model Machine Learning untuk Klasifikasi Risiko Pinjaman Perbankan Menggunakan Random Forest dan XGboost

Credit risk is one of the major challenges in managing loan portfolios in the banking sector, highlighting the need for approaches capable of identifying loan risk patterns and factors associated with problematic loans based on historical data. This study aims to develop and compare the performance of Random Forest and Extreme Gradient Boosting (XGBoost) algorithms in classifying loan risk patterns and to identify the features that contribute to the classification results. The study uses a historical loan dataset processed through exploratory data analysis, categorical data transformation, an 80:20 training and testing data split, and class imbalance handling using Random Under Sampling and SMOTE. Model optimization was performed through hyperparameter tuning using GridSearchCV, while model performance was evaluated based on accuracy, precision, recall, and F1-score. Model interpretation was conducted using Feature Importance, SHAP (SHapley Additive Explanations), and decision tree visualization. The results show that XGBoost with preprocessing and the second hyperparameter tuning achieved the best performance, with an accuracy of 98%, precision of 97%, recall of 85%, and F1-score of 90%, outperforming Random Forest on the dataset used. Interpretability analysis indicates that recoveries, total_rec_prncp, out_prncp, last_pymnt_amnt, and funded_amnt are among the features that contribute substantially to the classification results. These findings indicate that preprocessing, class imbalance handling, and parameter optimization can improve classification performance while providing insights into the factors contributing to loan risk patterns based on historical data.

Zacki Ferdinansyah, Rahmat Budiarsa · 0 citations
Open access Jul 2026

Pendekatan Interpretable Machine Learning untuk Analisis Keberhasilan Kampanye Pemasaran Menggunakan CatBoost dan SHAP

Predicting the success of digital marketing campaigns remains a significant challenge due to the complex interactions among various variables, such as budget allocation and advertising channel selection. This study aims to develop a marketing analytics model that achieves high predictive accuracy while also providing clear interpretability of the prediction results. The study uses the SalesMind Marketing Campaigns 2026 dataset, which simulates 3,478 digital marketing campaign records from 2026. The dataset consists of categorical variables such as ad_channel and campaign_type, as well as numerical variables including marketing_spend, impressions, and conversion_rate as the prediction target. The proposed approach applies Interpretable Machine Learning by combining the CatBoost algorithm to predict conversion rates and SHAP (SHapley Additive exPlanations) to analyze the contribution of each variable. Model optimization was performed using GridSearchCV, resulting in excellent performance with an RMSE of 0.0012, an MAE of 0.005, and a coefficient of determination (R²) of 99.12%. The analysis results indicate that budget allocation is the most dominant factor in improving conversion rates without showing indications of diminishing marginal effectiveness. In addition, the use of interactive platforms such as Meta and TikTok significantly contributes to campaign effectiveness. These findings contribute to providing an accurate and informative predictive model that can support strategic decision-making in digital marketing management more effectively.

Aprilisa Arum Sari, Oktalia Kumala Sari · 0 citations
Aug 2026

Explainable artificial intelligence–driven SHAP-based feature selection for interpreting black-box fuzzy modeling: an autonomous decision-making framework

Integrating explainable artificial intelligence (XAI) for water quality assessment (WQA) is necessary to protect human health, ensure the availability of clean water and sustainable conservation of the environment. Black-box machine learning (ML) models perform well but lack cognitive insight, limiting their use in decision-support systems. This research addresses the need by putting forward a concise, explainable technique for binary classification of water quality. The research presents a comprehensive autonomous decision-making framework for assessing water quality, incorporating XAI, feature selection and fuzzy IF-THEN reasoning. To ensure the integrity of the statistical analysis, missing data has been addressed through the implementation of multiple imputation by chained equations (MICE). Effective use of highly performing classifier facilitated the predictions. Shapley additive explanations (SHAP) were adopted to identify and normalize significant characteristics. To boost interpretation, fuzzy linguistic terms are generated using arcsinh-based quartile partitioning, facilitating the formation of SHAP-driven fuzzy IF–THEN rules. The validity of the rules has been monitored by activation strength analysis to confirm the consistency between fuzzy inference and predictions of the models. Outperforming all other models, the Random Forest algorithm scored highest accuracy of almost 78%. Through the integration of SHAP computation, the most influential criteria of water quality were determined. The fuzzy modeling process was made simpler by rule creation based on the most essential variables with time complexity reduction of 46.9%. The established rules offered clear conclusions about the evaluation of water potability and successfully translate complicated ML results into indicators that humans can grasp. Selection of model with high accuracy, activation strength analysis of rules, inter-fold consistency in SHAP ranks demonstrates that the suggested framework attains both predictive and interpretative stability. The proposed XAI–SHAP black-box fuzzy model facilitates the development of autonomous decision-making framework for public health, water quality and environmental sustainability. Despite the scientific merits of soundness and understanding of the described SHAP-fuzzy architecture, numerous limitations have to be admitted. The RF model had experienced medium predictive accuracy with an accuracy of about 78% throughout cross-validation. The standard deviation is relatively small indicating the model is stable and is always generalizing correctly. On the other hand, it may be possible to enhance the classification effectiveness with better advanced data pre-processing, dealing overlapping of criteria and class imbalance and use of integrated models, which have a higher predictive strength. The framework can support water management authorities in making faster, more transparent decisions about water quality. By combining explainable AI with fuzzy reasoning, it helps non-experts understand why certain assessments are made, improving trust and accountability. It can be applied to real-time monitoring systems to detect contamination risks early and prioritize interventions. The approach also enables efficient resource allocation by focusing on the most influential parameters. The framework can improve public health by enabling earlier detection of water contamination and more reliable quality assessments. Its transparency helps build trust among communities, regulators and stakeholders by clearly explaining decisions. Better water management can support equitable access to safe water, particularly in vulnerable regions. However, disparities in data availability and technical infrastructure may widen gaps between well-resourced and underserved areas. In contrast to traditional methods that utilize SHAP exclusively as an interpretation instrument, the suggested framework developed a trustworthy methodology by integrating a transparent connection between ML explanations and linguistic fuzzy reasoning. The proposed XAI–SHAP fuzzy model facilitates the development of decision-making framework for public health, water quality and environmental sustainability. It uniquely combines SHAP-based feature normalization, arcsinh-based quartile partitioning, understandable IF-THEN fuzzy rules and activation strength analysis. This guarantees both interpretability and stability, offering a clear and elucidative method for binary classification of water quality.

Freeha Qamar, Muhammad Riaz · 0 citations
Open access Jul 2026

Optimasi Hiperparameter XGBoost Regression untuk Prediksi Harga Saham BBNI Berbasis Transaksi Historis

Stock investment in the banking sector, such as PT Bank Negara Indonesia (Persero) Tbk (BBNI), carries high risks due to dynamic market volatility, necessitating accurate prediction methods to support investment decision-making. This study aims to optimize the performance of the XGBoost Regression algorithm in predicting BBNI stock prices based on historical transaction data from the last five years. The methodology applied includes data preprocessing for price format validation, feature engineering using technical indicators (SMA, EMA, MACD, RSI), and hyperparameter optimization using the Grid Search Cross-Validation technique. The experimental results demonstrate that hyperparameter optimization effectively refines the model's predictive stability. While maintaining a highly precise Mean Absolute Percentage Error (MAPE) of 2.00%, the Grid Search technique successfully reduced the nominal error (RMSE) and improved the model's goodness-of-fit (R-Squared). These findings confirm that the optimized model offers superior generalization capabilities in capturing price volatility compared to the baseline model. Consequently, this optimized predictive model provides a robust analytical tool for investors and financial analysts to mitigate risks and formulate effective trading strategies in the highly fluctuating banking stock market.

Muhammad Irfan · 0 citations
Jul 2026

Explainable Machine Learning for Type 2 Diabetes Screening Using Shap Feature Attribution on NHANES 2017-2018 Data

Type 2 Diabetes Mellitus (T2DM) presents a critical public health challenge, particularly in Southeast Asian lowand middle-income countries where healthcare resources are constrained. This paper evaluates and compares two machine learning classifiers - Random Forest (RF) and XGBoost - for T2DM risk classification using the NHANES 2017-2018 dataset (5.393 adult participants). SHAP (SHapley Additive exPlanations) is applied to both models to provide clinically interpretable feature attribution. XGBoost achieved the highest overall performance with accuracy of 91.84%, precision of 0.8372, F1-score of 0.7105, and AUC-ROC of 0.929. SHAP analysis consistently identified HbA1c, age, and waist circumference as dominant predictors across both models. This work constitutes the ML classification and explainability phase of a broader programme toward an Explainable AI-Driven Digital Twin Framework for T2DM management in Southeast Asian health information systems; Digital Twin architecture and HL7 FHIR integration are reserved for subsequent phases.

Helen Sastypratiwi, T. Wah, Saadial Razalli Bin Azzuhri · 0 citations
Open access Jul 2026

Pengembangan model prediksi drug related problems menggunakan machine learning pada pasien geriatri dengan polifarmasi

Background: Drug Related Problems (DRPs) are a major cause of reduced therapeutic quality, prolonged hospital stays, and high healthcare costs among geriatric patients. Polypharmacy and multimorbidity increase therapeutic complexity, necessitating predictive methods capable of early identification of high-risk patients. Purpose: To develop a machine learning-based prediction model for drug related problems (DRPs) in geriatric patients with polypharmacy. Method: A retrospective, analytical observational study design was employed, utilizing electronic medical record data from hospitalized geriatric patients. A total of 2,458 patients meeting the inclusion criteria were analyzed. Prediction models were developed using Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, and Extreme Gradient Boosting (XGBoost). Model evaluation was conducted using Accuracy, Precision, Recall, F1 score, and the Area Under the Receiver Operating Characteristic Curve (ROC-AUC). Model interpretation was performed using Shapley Additive Explanations (SHAP). Results: A total of 70.6% of patients experienced DRPs, with the most common category being treatment effectiveness. Factors significantly associated with the occurrence of DRPs included advanced age, the number of medications, chronic kidney disease, the use of high-alert medications, and major drug interactions (p<0.05). Conclusion: The incidence of Drug Related Problems (DRPs) among geriatric patients undergoing polypharmacy is high, predominantly within the Treatment Effectiveness category. The XGBoost algorithm proved to be the most effective at predicting DRPs, with key predictors including the number of medications, drug interactions, renal function, age, and High Alert Medications. Suggestion: Future research should conduct external validation using multicenter data from various hospitals. Additionally, the integration of the model into hospital information systems needs to be evaluated through prospective studies to assess its impact on reducing the incidence of DRPs.   Keywords: Clinical Decision Support System; Drug Related Problems; Machine Learning; Older Adults; Polypharmacy; XGBoost.

Ika Sutra Perwirahayu Aji Saputri · 0 citations

Related blog posts