Explainable Earnings-Quality Risk Screening Using Hybrid Machine Learning and Governance Signals: An Auditor-Oriented Framework for Indian Listed Firms
Aug 2026· International Journal For Multidisciplinary Research· 0 citations· 38 references
TL;DR
The study contributes an auditor-oriented architecture that separates predictive screening from the professional conclusion, embeds explanation quality and calibration into model evaluation, and maps model outputs to review procedures, suitable for future validation on verified Indian firm-year enforcement, restatement, and qualified-report outcomes.
Abstract
Artificial intelligence is increasingly used to prioritize audit attention, yet high predictive accuracy alone is insufficient in assurance settings where reviewers must understand why a firm-year has been classified as risky. This study develops an Explainable Earnings-Quality Risk Screening framework (EQR-XAI) for Indian listed-firm auditing. The framework is intentionally distinct from conventional financial-misstatement classifiers: it predicts an earnings-quality risk state rather than asserting fraud, combines accounting-ratio, cash-flow, governance, related-party, and auditor-transition signals, and produces auditor-readable explanations through SHAP attribution and rule-based reason codes. The research design specifies a firm-year panel assembled from publicly available annual reports and exchange disclosures, with a reproducible synthetic evaluation panel used in this paper to demonstrate the complete analytical workflow without representing simulated values as observed corporate facts. Comparative models include logistic regression, random forest, and gradient boosting. The illustrative experiment shows the proposed hybrid model reaching ROC-AUC 0.91 and F1 0.84, with calibration error below the benchmark models. Accrual intensity, the cash-flow-to-profit gap, receivable growth, related-party intensity, and auditor change emerge as the leading explanation drivers. The study contributes an auditor-oriented architecture that separates predictive screening from the professional conclusion, embeds explanation quality and calibration into model evaluation, and maps model outputs to review procedures. The framework is suitable for future validation on verified Indian firm-year enforcement, restatement, and qualified-report outcomes.
The evaluation of accounting transactions is increasingly challenging due to the growing volume of financial records, severe class imbalance, and the limited transparency of existing audit support systems. Many current machine learning approaches emphasize prediction accuracy while providing insufficient interpretability and weak support for risk-oriented audit decisions. To address these issues, this paper proposes an Intelligent Accounting framework based on explainable machine learning for risk-oriented transaction outcome prediction. The proposed framework integrates accounting-driven feature engineering, supervised learning, SHAP based explainable artificial intelligence, and probability-based risk scoring into a unified decision-support pipeline. Logistic Regression is adopted as the core predictive model due to its robustness, interpretability, and model parsimony under highly imbalanced transaction data. Experimental results on accounting dataset consisting of 1,000 transaction records show that Logistic Regression achieved the highest PR-AUC of 0.9737 and ROC-AUC of 0.6458 compared with Random Forest and XGBoost. The risk scoring mechanism also ranked problematic transactions within the highest-risk group, supporting audit prioritization. In addition, graphical SHAP analysis provides qualitative insights by identifying Operating Expenses, log_Operating Expenses, transaction timing, Transaction Volume, Profit Margin, Revenue, Expenditure, Cash Flow, Gross Profit, and Accuracy Score as influential factors affecting transaction outcomes. These findings show that the proposed framework not only predicts transaction outcomes but also explains the accounting factors behind each decision. Overall, this study transforms conventional transaction classification into an interpretable, risk-oriented, and audit-driven intelligent accounting system for transparent financial decision support.
J. K. Siregar, Astari Dianty, Antonius Bimo Rentor et al.· International Seminar on Int...· 0 citations
An Explainable Machine Learning (XML) framework for credit risk assessment that combines an ensemble classifier, integrating XGBoost, Random Forest, and LightGBM, with an integrated SHAP-and-LIME explainability layer is proposed and evaluated using a large-scale retail and priority-sector loan dataset drawn from public sector, private sector, regional rural, and small finance bank segments operating in India.
A. Agrawal, Vaibhav C. Gandhi· International journal of com...· 0 citations
Profitability deterioration often appears before bankruptcy, default, or formal financial distress. This paper develops an explainable machine learning framework to identify listed companies with high next-year profitability downside risk. Using a public firm-year panel constructed from SEC Financial Statement Data Sets and Stooq historical stock price data, this study examines U.S. non-financial listed companies from 2015 to 2024. High-risk observations are defined as firms whose next-year ROA change falls in the bottom 30 percent within the same industry-year group. Logistic Regression, Random Forest, XGBoost, and LightGBM are evaluated under a chronological validation design. XGBoost performs best in the out-of-sample test set, with an AUC of 0.836 and a PR-AUC of 0.653. SHAP results indicate that revenue growth, operating margin, leverage, ROA, operating cash flow growth, and stock volatility are the main risk drivers.
Xiaoqian Lin· Applied and Computational En...· 0 citations
This study explores the use of Explainable Artificial intelligence techniques to improve the interpretability of credit default prediction and highlights the practical value of explainable machine learning in developing more understandable, trustworthy, and accountable credit risk assessment systems for real-world financial decision-making.
Muskan, B. Sidhu· International Journal of Com...· 0 citations
Corporate financial distress imposes sizable and persistent costs on shareholders, creditors, employees, and the broader economy. However, practical risk governance in emerging markets requires not only accurate risk ranking but also reliable probabilities that can be translated into monitoring thresholds and escalation actions. This study develops an early-warning framework for Vietnamese listed non-financial firms that targets decision-useful one-year-ahead distress probabilities. It evaluates whether calibrated and explainable machine-learning models remain operationally usable under temporal change. The analysis uses a firm-year panel of companies listed on the Ho Chi Minh City Stock Exchange and the Hanoi Stock Exchange over 2014–2024, comprising 7305 observations. A strict chronological design is implemented to emulate forward deployment: model estimation uses 2015–2019, tuning and probability calibration use 2020–2021, and final evaluation is conducted once on an out-of-time hold-out period of 2022–2024. During the test period, distress prevalence increases to 15.42%, compared with approximately 12% in earlier windows. A conventional probabilistic benchmark is compared with multiple machine-learning classifiers under an identical feature space and temporal protocol. An explanation layer is also applied to support governance-oriented interpretation. Out-of-time evidence indicates that tree-ensemble methods provide the strongest combination of ranking performance and probability accuracy, with the random forest providing the best out-of-time performance among the evaluated models, with reasonable discrimination and the lowest probability error, although the magnitude of the AUC improvement should be interpreted as moderate rather than exceptional. The random forest’s calibrated probabilities support monotonic risk stratification and transparent, capacity-constrained watchlists. Selecting the top 10% of firm-years by predicted risk captures 48.3% of distress events with 44.3% precision, corresponding to a 2.87-fold lift over the base rate. Incremental gains from the structural, market-implied distance-to-default proxy are limited once standard accounting and market measures are included. This suggests that routinely available accounting and market variables already capture most of the relevant distress information in this setting. Overall, the results support an out-of-time calibrated and explainable pipeline as a practical foundation for auditable monitoring and tiered escalation in Vietnam’s listed corporate sector.
Tuyen Le Nam, Tam Phan Huy· Journal of International Com...· 0 citations