Jul 2026· International Journal of Latest Technology in Engineering Management & Applied Science· 0 citations· 9 references
TL;DR
An Integrated Data-Driven Loan Management Framework that comprises an Artificial Neural Network (ANN) for credit default prediction, Structured Query Language (SQL) to systematically extract and transform data, and Interactive Visual Analytics Dashboards to aid in providing transparency within the decision-making process is discussed.
Abstract
The credit risk assessment remains a critical challenge for financial institutions as manual and semi-automated loan evaluation processes often produce inconsistent decisions, high default rates, and operational inefficiencies. This paper discusses an Integrated Data-Driven Loan Management Framework that comprises an Artificial Neural Network (ANN) for credit default prediction, Structured Query Language (SQL) to systematically extract and transform data, and Interactive Visual Analytics Dashboards to aid in providing transparency within the decision-making process. For the study, a public loan dataset was used to pre-process the target dataset containing a total of 38,577 records and 24 attributes with pre-processing techniques such as feature engineering, one-hot encoding and class re-balancing with the use of Synthetic Minority Oversampling Technique (SMOTE) yielding 43 model-ready inputs.
The final ANN architecture was developed with 2 hidden layers (43 and 21 neurons with ReLU activation) and a sigmoid output neuron for binary classification. SHapley Additive exPlanations (SHAP) were computed to provide interpretability for each prediction made by the model. The overall accuracy of the model is 87% and the weighted precision, recall, and F1-score are 0.87 for the held-out test set (n = 12,858). The framework is implemented as a web-based decision support tool using Flask that enables users to receive risk scores in real-time along with explainable outputs and visual dashboards. The results from this experimentation indicate that an integrated pipeline (including data querying, predictive modeling, interpretability, and visualization) provides better decision-making and stakeholder transparency than using a model in isolation; therefore, it serves as a proof-of-concept prototype for intelligent loan management in banks and other financial institutions.
Accurate loan default prediction is crucial for managing credit risk. Traditional models struggle with class imbalance, where defaults are rare. This study creates a hybrid system combining machine learning with Deep Neural Networks, using Adaptive Synthetic Sampling (ADASYN) to effectively address this imbalance. The research utilized a comprehensive dataset of 255,347 loan records with 18 predictive features, including borrower demographics, financial attributes, and loan characteristics. The study employed rigorous exploratory data analysis, four models were systematically evaluated: Logistic Regression, Random Forest, XGBoost, and the proposed Hybrid ML-DNN system with ADASYN integration. The study's methodology involved preprocessing data, applying ADASYN to handle class imbalance, and creating a hybrid model. This combines interpretable traditional ML with powerful deep neural networks to better predict complex, non-linear patterns in loan defaults. Traditional baseline models (Logistic Regression, Random Forest, XGBoost) achieved misleadingly high overall accuracies (88%+) but failed catastrophically at default detection, with recall rates between 0% and 8% for the minority class. These models essentially learned to predict all loans as "good," rendering them operationally ineffective for risk assessment. In stark contrast, the Hybrid ML-DNN model with ADASYN achieved a 52% recall rate for defaults a more than six-fold improvement over the best baseline model while maintaining a 79.83% overall accuracy and 0.7571 ROC AUC score. The model successfully identified 3,074 actual defaults out of 5,931 total defaults in the test set, compared to baseline models that identified fewer than 500 defaults. However, this improvement came with a precision trade-off of 29% for the default class, resulting in 7,442 false positives, highlighting the inherent business decision between risk mitigation and customer acquisition. The study proves ADASYN is essential for effective loan default prediction, not just an enhancement. It provides a robust, transparent AI framework for financial institutions, enabling improved risk assessment and promoting more stable and responsible lending practices
J. Godson· International journal of re...· 0 citations
Large volumes of loan applications motivate automated decision-support systems that can reduce processing delays and improve consistency while controlling credit risk. This study presents a unified supervised-learning framework comparing XGBoost, Gradient Boosting, and CatBoost for loan approval prediction. The experiments use the Dream Housing Finance dataset containing 614 applications and 12 predictive variables after removing Loan_ID. The pipeline includes missing-value treatment, feature engineering, scaling, SMOTE-based class balancing applied only to training data, and evaluation on a held-out test set of 169 samples. Perfect training performance is treated as a diagnostic warning rather than evidence of generalization. CatBoost achieved the best held-out accuracy (88.17%), precision (88.37%), recall (88.37%), and F1-score (88.37%), with 10 false approvals and 10 false rejections. Confusion-matrix analysis, false-positive and false-negative rates, balanced accuracy, and Wilson confidence intervals indicate the most balanced performance among the evaluated models. The framework is intended as a prototype decision-support approach; larger multi-institutional validation, probability-based discrimination analysis, explainability, calibration, and fairness assessment are required before deployment in real lending environments.
N. Dandotiya, Kirti Jain, Prashant Kumar Shrivastava· 2026 International Conferenc...· 0 citations
Credit risk assessment is a core component of financial decision-making. This study develops an explainable machine learning framework for modeling loan approval decisions on heterogeneous tabular data, centered on a Cross-Attentional Tabular Transformer that applies bidirectional cross-attention between numerical and categorical feature groups. The prediction target is historical loan-approval status, treated as a proxy for, not a direct measure of, borrower default risk; a supplementary validation on a dataset with an authentic default label is also reported. Class imbalance is addressed through focal loss, and post hoc interpretability is provided through SHAP analysis. Three classifiers, Random Forest, Gradient Boosting, and the proposed transformer, are evaluated on a 5000-sample credit dataset using accuracy, precision, recall, F1-score, ROC-AUC, and average precision. Gradient Boosting achieves the best performance (accuracy 0.9640, F1-score 0.9189), with Random Forest comparable; the proposed transformer reaches 0.9530 accuracy and 0.8949 F1, without surpassing the ensembles and at substantially higher computational cost. A five-split robustness comparison additionally evaluates XGBoost, LightGBM, CatBoost, and calibrated logistic regression: all three Gradient-Boosting variants and both classical ensembles exceed the transformer’s performance on every metric, while calibrated logistic regression does not. The evaluated baseline set excludes deep tabular architectures such as TabNet, FT-Transformer, SAINT, and TabPFN-style methods. Across the three primary classifiers, SHAP identifies credit score, employment status, and income as the dominant features, consistent with domain expectations. The results characterize the observed performance–efficiency trade-off between ensemble methods and attention-based tabular learning under the evaluated data conditions.
Bowen Dong, Xinyu Zhang, Ziwei Hong et al.· Entropy· 0 citations
An Explainable Machine Learning (XML) framework for credit risk assessment that combines an ensemble classifier, integrating XGBoost, Random Forest, and LightGBM, with an integrated SHAP-and-LIME explainability layer is proposed and evaluated using a large-scale retail and priority-sector loan dataset drawn from public sector, private sector, regional rural, and small finance bank segments operating in India.
A. Agrawal, Vaibhav C. Gandhi· International journal of com...· 0 citations
This study explores the use of Explainable Artificial intelligence techniques to improve the interpretability of credit default prediction and highlights the practical value of explainable machine learning in developing more understandable, trustworthy, and accountable credit risk assessment systems for real-world financial decision-making.
Muskan, B. Sidhu· International Journal of Com...· 0 citations
Standard credit-risk scorecards rely on linear ratio thresholds that break down when feature interactions are nonlinear and observations carry temporal dependencies. Qualitative signals embedded in corporate disclosures—tone shifts, forward-looking hedges, and sector-specific terminology—remain largely ignored by numeric-only models, even though such signals often precede ratio deterioration. This paper introduces a tri-modal deep learning framework that jointly trains three complementary branches: a Convolutional Neural Network (CNN) for cross-sectional ratio-pattern detection, a Long Short-Term Memory (LSTM) network for multi-quarter trend modelling, and a Natural Language Processing (NLP) branch for disclosure-text encoding. Prior to deep-model training, LASSO regularisation removes collinear financial indicators and SMOTE oversampling corrects the severe class imbalance characteristic of distress datasets. A feature-concatenation fusion layer integrates all three branch outputs; the resulting vector feeds a sigmoid classifier that produces a calibrated distress probability. Benchmarked against five baselines on four financial datasets, the model reaches 94.8% accuracy and 91.3% minority-class recall, with a 4.1-point F1 advantage over the strongest single-modality competitor.
Paiinti Meenakshi, Muthaluru Bhuvaneshwari, Jalla Ganesh et al.· International Conference Com...· 0 citations