Skip to content
Open access

An Explainable Machine Learning Framework for Credit Risk Assessment and Optimization in the Indian Banking Sector

Jul 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

An Explainable Machine Learning (XML) framework for credit risk assessment that combines an ensemble classifier, integrating XGBoost, Random Forest, and LightGBM, with an integrated SHAP-and-LIME explainability layer is proposed and evaluated using a large-scale retail and priority-sector loan dataset drawn from public sector, private sector, regional rural, and small finance bank segments operating in India.

Abstract

Credit risk assessment remains a foundational function of commercial banking, yet the Indian banking sector's persistent non-performing asset (NPA) burden, heterogeneous borrower base, and expanding priority-sector and microfinance lending create distinct challenges that generic, globally trained credit scoring models do not adequately address. This paper proposes an Explainable Machine Learning (XML) framework for credit risk assessment that combines an ensemble classifier, integrating XGBoost, Random Forest, and LightGBM, with an integrated SHAP-and-LIME explainability layer, and evaluates the framework using a large-scale retail and priority-sector loan dataset drawn from public sector, private sector, regional rural, and small finance bank segments operating in India. The framework incorporates SMOTE-ENN based class-imbalance handling to address the low base default rate typical of retail lending portfolios, and an explainability-guided feature refinement step that uses SHAP attributions to iteratively prune low-value features and retain regulator-interpretable risk factors. The proposed model was evaluated on a dataset of 42,860 loan accounts and benchmarked against five baseline models. Results show that the proposed explainable ensemble achieves 93.2% accuracy, 91.0% precision, 89.4% recall, and a 90.2% F1-score, exceeding the strongest baseline (LightGBM) by 5.6 percentage points in F1-score, with an area under the ROC curve (AUC) of 0.967. Segment-wise analysis reveals materially higher default risk concentration in regional rural banks and small finance institutions relative to public and private sector banks, with debt-to-income ratio, credit bureau (CIBIL) score, and repayment delinquency history emerging as the most influential predictors across segments.

Read PDF

Similar papers

Open access Jul 2026

An Interpretability Analysis of Credit Default Prediction Using Random Forest with SHAP and LIME

This study explores the use of Explainable Artificial intelligence techniques to improve the interpretability of credit default prediction and highlights the practical value of explainable machine learning in developing more understandable, trustworthy, and accountable credit risk assessment systems for real-world financial decision-making.

Muskan, B. Sidhu · 0 citations
Review Open access Aug 2026

Explainable Machine Learning for Credit Risk Management and Intelligent Lending Decisions in Nepalese Cooperative Banks: A Mathematical Review

It is argued that predictive accuracy and regulatory transparency are not competing objectives but complementary necessities for institutional survival in Nepal’s cooperative sector.

S. K. Sahani, Tsair-Fwu Lee, Digvijay Pandey et al. · 0 citations
Open access Jul 2026

An Integrated Explainable Deep Learning Framework for Loan Default Prediction and Credit Risk Decision Support System

An Integrated Data-Driven Loan Management Framework that comprises an Artificial Neural Network (ANN) for credit default prediction, Structured Query Language (SQL) to systematically extract and transform data, and Interactive Visual Analytics Dashboards to aid in providing transparency within the decision-making process is discussed.

Ponsak .S. Bande, Chimuzuoroke E. Ugwuja, Blessing .C. Uzo et al. · 0 citations
Open access Aug 2026

Towards Sustainable Financial Inclusion: A Comparative Study of Ensemble Architectures and SHAP-Based Explainability in Bank Loan Prediction

By empirically proving that high-performance algorithms can be mathematically blind to demographic biases, this framework directly advances SDG 10 (Reduced Inequalities) and provides the accountable, feature-level justifications required for secure and sustainable financial inclusion (SDG 8).

Htet Nge Nge Ko, Aung Htoo Khine, Shadab Kalhoro et al. · 0 citations
Review Open access Aug 2026

Machine Learning for Individual Credit Risk Assessment: A Systematic Literature Review of State-of-the-Art Methods, Challenges and Perspectives

A systematic literature review of ML applications in credit risk assessment (CRA), covering publications from January 2016 to May 2026, synthesise prevailing methodologies into a unified end-to-end credit risk modelling framework spanning data preprocessing, feature engineering, model training, evaluation, and operational deployment.

Bolun Zhang, Jun Luo, Ruobing Wu et al. · 0 citations
Open access Jul 2026

Optimizing Credit Risk Assessment in Ghanaian Micro-Lending Institutions: A Comparative Analysis of Random Forest, Extra Tree Classifier, and Ensemble Machine Learning Models

Credit risk assessment is pivotal to the sustainability of micro-lending institutions, particularly in emerging economies such as Ghana, where conventional evaluation methods remain predominantly manual and subjective. Traditional approaches, which rely on face-to-face interviews, personal judgments, and simple background checks, are vulnerable to human biases, inconsistencies, and inefficiencies that contribute to elevated default rates and broader financial instability. This study investigates the application of machine learning (ML) techniques, specifically Random Forest (RF), Extra Tree Classifier (ETC), and a probability-averaged Ensemble Classifier, to enhance credit risk assessment in Ghanaian micro-lending institutions. Using a quantitative experimental research design, the study analysed 32,581 loan records drawn from Tepa Man Microfinance Institution. Data preprocessing included missing-value imputation, one-hot encoding, and class balancing via random oversampling, applied exclusively to the training set. Model performance was evaluated through 10-fold stratified cross-validation using accuracy, precision, recall, F1-score, AUC-ROC, Cohen's Kappa, and Matthews Correlation Coefficient (MCC). Hyperparameters were set to scikit-learn defaults (n_estimators = 100, random_state = 42) to ensure reproducibility. The Random Forest and Extra Tree Classifiers each achieved a mean accuracy of 99.33% and an AUC-ROC of 0.9997, results that are consistent with the high-quality, real-world dataset and are critically interpreted in the context of potential overfitting risks. Feature importance analysis identified the loan-to-income ratio and interest rate as the dominant predictors of default. The Ensemble Method, which averages class probabilities across both base models, achieved 84.25% accuracy and an AUC of 0.9231, demonstrating stronger generalization than the individual classifiers. The study concludes that integrating ML models can substantially improve the accuracy, consistency, and reliability of credit risk evaluations, thereby reducing default rates and supporting financial inclusion in Ghana's microfinance sector.

P. Addo, Samuel Kofi Akpatsa, Emmanuel Mensah et al. · 0 citations