Skip to content
Open access

The Role of Machine Learning in Accounting: A Study on Automated Financial Statement Analysis and Prediction

Jul 2026 · Journal of Administrative and Economic Sciences · Vol 19, pp. 59-67 · 0 citations · 19 references

TL;DR

Whether appending lowdimensional text features to strong predictive models enhances accuracy is examined by examining whether appending lowdimensional text features to strong predictive models enhances accuracy, and a "Signal Dilution Effect" is revealed.

Abstract

The integration of machine learning into accounting has advanced predictive analytics, yet the optimal fusion of structured financial data (“hard data”) and textual sentiment (“soft data”) remains ambiguous. This study addresses this gap by examining whether appending lowdimensional text features to strong predictive models enhances accuracy. Using an XGBoost model trained on financial ratios from 10-K reports, we compare its performance against a naive hybrid model incorporating Loughran-McDonald (2011) dictionary-based sentiment features from the Management Discussion & Analysis (MD&A) section. Empirical results reveal a “Signal Dilution Effect”: the hybrid model statistically underperforms the pure structured-data model. SHAP value analysis confirms that while structured ratios dominate predictive contributions, naive text concatenation introduces noise that degrades the performance of advanced non-linear classifiers. This research contributes to accounting analytics by cautioning against indiscriminate feature stacking and advocating for separated processing strategies in financial prediction systems.

Read PDF

Similar papers

Open access Jul 2026

Machine Learning-Based Detection of Financial Statement Fraud: Integrating Beneish Ratios, Linguistics, and Stock Volatility on the Indonesia Stock Exchange

This research aimed to develop and evaluate a financial statement fraud (FSF) detection model for companies listed on the Indonesia Stock Exchange (IDX) during the 2024–2025 period using a machine learning approach. An ablation study design was employed to test 27 model combinations across three feature scenarios: Beneish M-Score financial ratios, linguistic features extracted from Management Discussion and Analysis (MD&A) texts using the InSet lexicon, and 30-day stock price volatility. Nine classification algorithms were evaluated using precision, recall, and F1-score metrics. The best-performing model combined financial ratios and linguistic features using Gradient Boosting, achieving an F1-score of 0.545. In contrast, the addition of stock price volatility as a feature did not improve model performance and instead reduced the classification ability of all tested algorithms. These findings indicate that the Indonesian capital market may have limited ability to anticipate indications of financial statement fraud before related information is publicly disclosed. This study concludes that integrating accounting and linguistic information is more effective than relying solely on financial ratios or incorporating market data. These findings contribute to the development of machine learning-based FSF detection literature in Indonesia and provide an alternative approach for auditors, investors, and regulators to enhance the effectiveness of early financial statement fraud detection.

Abdurrochman Halomoan Hasibuan, Vidyarto Nugroho · 0 citations
Review Jul 2026

A Review of Machine Learning Applications for Credit Default Risk Prediction and Early Warning Systems

The paper shows that ensemble learning models have superior predictive power and argues that behavioral data complement traditional datasets for underbanked populations, such as "credit invisibles," and makes a strong case for XAI being essential for model transparency, combating bias, and meeting regulatory requirements.

Bojun Chen · 0 citations
Aug 2026

From linear to machine learning models: an empirical study on real earnings management detection in Indian listed firms

Earnings Management practices degrade the reporting quality and potentially deceive stakeholders. This paper addresses the evaluation of multiple models to identify the most effective model for detecting Real Earnings Management (R.E.M.). Financial data of non-financial BSE 500 listed companies from April 1, 2014, to March 31, 2024, was utilized for analysis. The study has compared the prediction and classification rate of Linear Regression (LR), logistic regression, support vector machines, random forests and Deep Belief Neural Networks (DBNNs). Further, the empirical analysis has been re-conducted using a dataset of S&P 500 firms for the period 2020–2024, to assess the robustness of the results. Empirical results have revealed the fact that DBNNs outperform other models in both classification and prediction of R.E.M. in Indian listed firms and the results have remained robust across a dataset of US listed firms. The present research offers empirical support to the relatively scarce literature by employing deep neural networks for the prediction and categorization of R.E.M.

Radhika, Meena Sharma, Anu Gupta · 0 citations
Review Open access Jul 2026

Machine Learning for Stock Market Prediction: A Review of Sentiment, Technical, Macroeconomic, and Fundamental Factors

Predicting the stock market has become an important research topic in recent years, because accurate result can support investment decision and reduce financial risks.Traditional statistical models are hard to analyze in the nonlinear market, leading researchers to develop machine learning models. This paper reviews recent studies about stock prediction based on machine-learning methods from four perspectives: traditional machine learning models, sentiment analysis, technical indicators, macroeconomic and fundamental factors. The reviewed literature shows that machine learning methods can have better result than traditional models in handling complex financial data. Furthermore, sentiment information, technical indicators, and macroeconomic or fundamental factors can significantly improve prediction performance. However, each approach has limitations. The review finds that using multiple factors may relate to accurate predictions. Therefore, combining different factors through hybrid models is considered the most effective strategy. Finally, this review reveals a clear transition from single-factor prediction models to multi-factor prediction systems and provides a comprehensive understanding of current developments in stock prediction and highlights potential directions for further research.

Tianyuan Shi · 0 citations
Open access Aug 2026

Towards Sustainable Financial Inclusion: A Comparative Study of Ensemble Architectures and SHAP-Based Explainability in Bank Loan Prediction

By empirically proving that high-performance algorithms can be mathematically blind to demographic biases, this framework directly advances SDG 10 (Reduced Inequalities) and provides the accountable, feature-level justifications required for secure and sustainable financial inclusion (SDG 8).

Htet Nge Nge Ko, Aung Htoo Khine, Shadab Kalhoro et al. · 0 citations
Jul 2026

Toward Intelligent Accounting: An Explainable Machine Learning Framework for Risk-Oriented Transaction Outcome Prediction

The evaluation of accounting transactions is increasingly challenging due to the growing volume of financial records, severe class imbalance, and the limited transparency of existing audit support systems. Many current machine learning approaches emphasize prediction accuracy while providing insufficient interpretability and weak support for risk-oriented audit decisions. To address these issues, this paper proposes an Intelligent Accounting framework based on explainable machine learning for risk-oriented transaction outcome prediction. The proposed framework integrates accounting-driven feature engineering, supervised learning, SHAP based explainable artificial intelligence, and probability-based risk scoring into a unified decision-support pipeline. Logistic Regression is adopted as the core predictive model due to its robustness, interpretability, and model parsimony under highly imbalanced transaction data. Experimental results on accounting dataset consisting of 1,000 transaction records show that Logistic Regression achieved the highest PR-AUC of 0.9737 and ROC-AUC of 0.6458 compared with Random Forest and XGBoost. The risk scoring mechanism also ranked problematic transactions within the highest-risk group, supporting audit prioritization. In addition, graphical SHAP analysis provides qualitative insights by identifying Operating Expenses, log_Operating Expenses, transaction timing, Transaction Volume, Profit Margin, Revenue, Expenditure, Cash Flow, Gross Profit, and Accuracy Score as influential factors affecting transaction outcomes. These findings show that the proposed framework not only predicts transaction outcomes but also explains the accounting factors behind each decision. Overall, this study transforms conventional transaction classification into an interpretable, risk-oriented, and audit-driven intelligent accounting system for transparent financial decision support.

J. K. Siregar, Astari Dianty, Antonius Bimo Rentor et al. · 0 citations