Skip to content
Open access

A Resampling Ensemble Model for Multi-Window Corporate Default Prediction Under Class Imbalance

Jul 2026 · Systems · Vol 14, pp. 776 · 0 citations · 61 references

TL;DR

This study provides an effective early-warning tool for financial institutions and relevant stakeholders to identify high-risk firms, and enriches empirical evidence on the time-varying drivers of corporate default risk.

Abstract

Effective identification of corporate default risk is crucial for maintaining financial stability and safeguarding investors’ interests. Existing models remain limited in addressing class imbalance and the dynamic evolution of default-related features over time. To overcome these challenges, we propose an adaptive spherical neighborhood resampling and class-specific reliability evidential reasoning model (ASNR-crER). By combining feature-weighted minority sample reconstruction with reliability-guided recursive evidence fusion, the proposed model aims to improve the prediction accuracy of both default and non-default firms under class imbalance. This study uses Chinese listed small enterprises from 2000 to 2023 as the research sample, comprising 10,449 firm-year observations from 2182 firms. By matching default status in year t with firm indicators from t-0 to t-5, six rolling prediction windows are constructed. The empirical results show that: (1) Compared with mainstream benchmark methods, ASNR-crER achieves the best overall performance in terms of accuracy, AUC, and F1 across all prediction windows, indicating that it can more reliably identify high-risk default firms while maintaining strong recognition of non-default firms. (2) SHAP analysis indicates that financial, non-financial, and macroeconomic indicators exert time-varying effects on corporate default risk. Financial indicators, including “Retained earnings/total assets”, “Other receivables/current assets”, and “Annualized return on assets”, reflect internal capital accumulation and profitability, serving as key predictors of default risk. Non-financial indicators, such as “Top 10 Tradable Shares H-index” and “Top 10 shareholders H-index”, can provide supplementary signals for medium-term risk identification. Macroeconomic indicators, including “M2 YoY growth rate”, “Urban HH per capita income”, and “Benchmark short-term loan rate”, show stronger explanatory power in longer prediction windows. Therefore, this study provides an effective early-warning tool for financial institutions and relevant stakeholders to identify high-risk firms, and enriches empirical evidence on the time-varying drivers of corporate default risk.

Read PDF

Similar papers

Conference Jul 2026

Predictive Financial Health Evaluation of Companies Using Random Forest and Logistic Regression Techniques

Nearly 50,000 corporate insolvencies were recorded across European markets in 2023, yet most deployed prediction models for corporate insolvency were calibrated on financial data that preceded the significant changes in supply chains and interest rates that followed the end of the pandemic. Identifying deteriorating financial health before default is framed here as a supervised binary classification task over 64 liquidity, leverage, and profitability ratio features drawn from the Polish Companies Bankruptcy Dataset: 10,503 firm-year observations from the third annual reporting period with a 20:1 class imbalance. SMOTE-based oversampling restores minority representation before training, and SHAP-guided feature ranking subsequently retains 22 high-impact ratios while reducing inter-feature collinearity without discarding predictive signal. This feature count is sufficient because the 22 ratios together account for 85% of the model’s total attribution mass in the preliminary Random Forest model. A soft-voting ensemble of Random Forest and Logistic Regression achieves 91.2% accuracy and AUC 0.943, outperforming standalone Random Forest (AUC 0.927), XGBoost (AUC 0.916), and Logistic Regression alone (AUC 0.891) across five stratified cross-validation folds with score stability within ±0.8%. Ensemble tree methods capture nonlinear ratio interactions that linear discriminant approaches structurally cannot represent, particularly across leverage and interest-coverage features where distress signals emerge two reporting periods before formal default.

K.Vengatesan, P. J, Sayyad Samee et al. · 0 citations
Aug 2026

From linear to machine learning models: an empirical study on real earnings management detection in Indian listed firms

Earnings Management practices degrade the reporting quality and potentially deceive stakeholders. This paper addresses the evaluation of multiple models to identify the most effective model for detecting Real Earnings Management (R.E.M.). Financial data of non-financial BSE 500 listed companies from April 1, 2014, to March 31, 2024, was utilized for analysis. The study has compared the prediction and classification rate of Linear Regression (LR), logistic regression, support vector machines, random forests and Deep Belief Neural Networks (DBNNs). Further, the empirical analysis has been re-conducted using a dataset of S&P 500 firms for the period 2020–2024, to assess the robustness of the results. Empirical results have revealed the fact that DBNNs outperform other models in both classification and prediction of R.E.M. in Indian listed firms and the results have remained robust across a dataset of US listed firms. The present research offers empirical support to the relatively scarce literature by employing deep neural networks for the prediction and categorization of R.E.M.

Radhika, Meena Sharma, Anu Gupta · 0 citations
Conference Jul 2026

Financial Risk Prediction using LASSO-GBDT Hybrid Modeling with TOPSIS-based Multi-Criteria Evaluation

Standard credit-risk scorecards rely on linear ratio thresholds that break down when feature interactions are nonlinear and observations carry temporal dependencies. Qualitative signals embedded in corporate disclosures—tone shifts, forward-looking hedges, and sector-specific terminology—remain largely ignored by numeric-only models, even though such signals often precede ratio deterioration. This paper introduces a tri-modal deep learning framework that jointly trains three complementary branches: a Convolutional Neural Network (CNN) for cross-sectional ratio-pattern detection, a Long Short-Term Memory (LSTM) network for multi-quarter trend modelling, and a Natural Language Processing (NLP) branch for disclosure-text encoding. Prior to deep-model training, LASSO regularisation removes collinear financial indicators and SMOTE oversampling corrects the severe class imbalance characteristic of distress datasets. A feature-concatenation fusion layer integrates all three branch outputs; the resulting vector feeds a sigmoid classifier that produces a calibrated distress probability. Benchmarked against five baselines on four financial datasets, the model reaches 94.8% accuracy and 91.3% minority-class recall, with a 4.1-point F1 advantage over the strongest single-modality competitor.

Paiinti Meenakshi, Muthaluru Bhuvaneshwari, Jalla Ganesh et al. · 0 citations
Open access Aug 2026

SMOTETomek-DNN: A Machine Learning Framework for Credit Risk Prediction with an Imbalanced Dataset

Accurate credit risk prediction plays a critical role in strengthening the financial stability of microfinance institutions, especially in developing economies, where increasing loan defaults and imbalanced borrower records create significant challenges for reliable decision-making. Although machine learning approaches have improved credit assessment practices, existing models often favor majority-class borrowers and fail to detect high-risk default cases effectively because of severe class imbalance. This limitation highlights the need for more robust and imbalanced-sensitive predictive frameworks. This study aims to develop an effective machine learning-based credit risk prediction framework by integrating data balancing strategies with ensemble and deep learning models. This study systematically investigates the impact of baseline learning and multiple resampling techniques, including oversampling, undersampling, and hybrid methods, when applied to Random Forest, XGBoost, LightGBM, CatBoost, and Deep Neural Network classifiers. The effectiveness of the proposed models was assessed using imbalance-aware evaluation measures, particularly ROC-AUC and Geometric Mean, along with conventional classification metrics. The experimental findings demonstrate that incorporating resampling techniques substantially improved the default detection performance. The DNN model combined with SMOTETomek achieved the best results, obtaining 94.9% F1-score, ROC-AUC 98%, and 97.2% of G-Mean. CatBoost also exhibited consistent competitiveness across different sampling configurations. These findings suggest that hybrid sampling integrated with advanced learning architectures can provide a reliable and practical solution for managing credit risk in imbalanced microfinance datasets, supporting improved lending decisions and sustainable financial operations in the future.

Tiruneh Kebede Dubale, Siraj Sebhatu Seyoum · 0 citations
Conference Open access Jul 2026

Predicting Bond Default Risk Based on the XGBoost Model

At present, China faces a large scale of bond defaults. To address the problems of low prediction accuracy and limited indicator systems in traditional models, this paper conducts research on bond default risk prediction based on the XGBoost model. This study selects 616 valid samples of credit bonds from entity industries in the Wind database from 2019 to 2024, and constructs a multi-dimensional indicator system covering bond characteristics, financial leverage, and other factors. After data cleaning and feature selection, the dataset is divided into training and test sets at a 7:3 ratio, and the XGBoost model is trained using 3-fold cross-validation. The results show that the sustainability of corporate profitability is the core influencing factor of default. The model achieves an AUC of 0.997 on the test set, with no severe overfitting, demonstrating excellent generalization ability and prediction accuracy. This provides methodological support for the prevention and control of bond default risk.

Shurui Xie · 0 citations