Adaptive Process Mining and Selective Monitoring for Algorithmic Auditing: A Survey of Representations, Learning Policies, Decision Strategies, and Open Problems
Aug 2026· Machine Learning and Knowledge Extraction· Vol 8, pp. 234· 0 citations· 69 references
TL;DR
A four-layer framework that connects process representation, learning, inspection allocation, and governance within a single budgeted sequential decision problem over event streams for selective algorithmic auditing is introduced.
Abstract
Selective algorithmic auditing requires deciding which process evidence should receive attention when exhaustive review is infeasible. This Review introduces a four-layer framework that connects process representation, learning, inspection allocation, and governance within a single budgeted sequential decision problem over event streams. Unlike prior reviews centered on predictive process monitoring, explainability, cost analysis, or bibliometric structure, the proposed framework examines how these functions interact when human review, computation, latency, and documentation capacity are constrained. A structured and targeted survey of 89 unique publication families is used to illustrate and critically examine event-log, Petri-net, graph, object-centric, neural, uncertainty-aware, sequential, bandit, reinforcement learning, and audit architecture approaches. The reviewed evidence indicates that substantial bodies of work address the individual layers, but cross-layer evaluation remains fragmented and uses heterogeneous datasets, objectives, and validation protocols. The synthesis identifies five priorities: audit-ready benchmarks, explicit inspection budget protocols, calibrated uncertainty, transfer across organizational contexts, and reproducible governance interfaces. The main contribution is a computational framework and a corpus-bounded research agenda that connects representation, learning, inspection allocation, and governance for selective algorithmic auditing.
Financial budget preparation remains vulnerable to errors arising from data extraction and mapping, spreadsheet manipulation, forecast assumptions, version control, judgemental adjustment and weak reconciliation. Artificial intelligence (AI) is increasingly proposed as a means of reducing these errors, yet the evidence base is fragmented across management accounting, accounting information systems, forecasting, robotic process automation, machine learning, anomaly detection and human-AI decision research. This critical narrative review evaluates how far current evidence supports AI-enabled error reduction in organisational budgeting and where new sources of error emerge. Literature published principally from 2015 to 16 June 2026 was identified through multidisciplinary and business-focused scholarly indexes, supplemented by citation searching and verification against authoritative bibliographic records. The synthesis distinguishes deterministic automation from predictive machine learning, anomaly detection and generative AI because these technologies address different failure modes and carry different assurance requirements. Evidence is strongest for reducing repetitive transfer and processing errors, improving selected accounting estimates and forecasts, and widening exception screening. Direct causal evidence that AI improves end-to-end corporate budget accuracy remains limited, while field evidence increasingly shows that benefits depend on data quality, process standardisation, confidence-aware human intervention and effective internal control. Important countervailing risks include data leakage, model drift, false precision, brittle automation, automation bias and unreliable numerical reasoning by large language models. The review therefore argues that AI should be treated as a layered control and decision-support architecture rather than an autonomous budget preparer. The most defensible design combines governed source data, deterministic calculations, validated forecasting models, exception detection, logged human overrides and continuous performance monitoring. Future research should test these arrangements in real budgeting cycles using common error taxonomies and outcome measures that capture accuracy, rework, reconciliation failures, uncertainty and control effectiveness.
F. Abed· Asian Journal of Economics B...· 0 citations
Freshness-constrained audit capacity (FCAC) is developed, a decision-support framework that treats automation as an authorization decision constrained by action risk, evidence freshness, and shared review capacity.
Technology change management in large financial institutions depends on risk assessments that are accurate, consistent, and auditable. In practice, many institutions still rely on self-reported questionnaires. Those questionnaires are subjective, easy to game, and poor at separating routine changes from the ones that later trigger major incidents. This paper presents SENTRY, a risk assessment platform that replaces questionnaire-based scoring with a deterministic machine learning pipeline built from gradient-boosted decision trees (XGBoost) and hybrid retrieval-augmented generation (RAG). The system combines structured operational metadata, application dependency graphs, and historical incident records with a hybrid semantic and lexical search over historical change requests. The retrieval step captures the risk signal in unstructured change request text, then compresses that signal into a single scalar feature before model inference. That design keeps the model deterministic and preserves per-prediction explainability via SHAP values. Evaluated on enterprise-scale change data, SENTRY achieves a ROC AUC of 0.87 and 85% overall accuracy, and it detects high-risk changes at roughly 3.25 times the rate of the existing process. We close by examining the architectural trade-offs behind this design and what they imply for the use of machine learning in regulated change management.
Daniel Arulpragasam, Christer Henrysson, E. Ly et al.· 0 citations
This study examines artificial intelligence (AI) in external auditing by synthesizing existing evidence, clarifying key concepts, identifying theoretical and methodological gaps, and outlining future research directions. A systematic literature review and bibliometric analysis were conducted on 130 peer-reviewed articles retrieved from Scopus and Web of Science databases. The review followed the PRISMA 2020 guidelines, while VOSviewer was used to map research trends, and thematic clusters. Research on AI in auditing has grown substantially, with the United States, China, and the United Kingdom leading scholarly contributions. The analysis identified three dominant research streams: machine learning and fraud detection, audit analytics and big data, and AI adoption and governance. Commonly applied AI techniques include machine learning, neural networks, natural language processing, robotic process automation, and expert systems. The study suggested that AI enhances fraud detection, risk assessment, and audit quality, while raising concerns regarding algorithmic bias, transparency, and professional skepticism. The study develops an integrated framework linking AI applications to audit quality and provides a research agenda to guide future inquiry. The findings offer practical insights for auditors, regulators, and organizations seeking to implement AI responsibly and effectively in audit processes.
Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empirical audit protocol for structured intermediate outputs: first audit dataset shortcuts, then isolate bundled prompt changes, check whether intermediate labels are answer-associated, test decomposed semantic evidence, and audit provider-level execution failures. A 480-example synthetic development set initially suggested large gains from a state-structured prompt bundle, but TF-IDF diagnostics showed lexical separability and no positive standalone Ignore cases. We therefore construct a frozen 160-example controlled counterfactual set with 40 matched four-way families and rule-derived reference policies. On this set, exposing the four state definitions improves accuracy, but an isolated explicit state-output field does not significantly improve policy accuracy for Llama-3.3-70B and gives only a marginal, non-significant gain for GPT-OSS-120B. Supplying benchmark-associated state labels shifts policy predictions, but because those labels deterministically map to policies, this is a label-conditioning diagnostic rather than evidence of a faithful internal mechanism. Family-level and seed-stability analyses further show that example-level accuracy overstates counterfactual consistency: complete four-way family success is rare. An exploratory follow-up that elicits decomposed semantic evidence also fails to improve routing for the cleanly evaluated endpoint; the corresponding GPT-OSS condition was unavailable because of provider-side request validation. We evaluate policy classification only, not downstream responses, tool actions, or memory-store mutation.
Yihang Chen, Pinyan Qian, Su Wang et al.· 0 citations
Artificial intelligence has diffused rapidly through the recruitment function, from automated résumé screening and chatbot-led candidate engagement to algorithmic assessment and predictive analytics, and vendors and adopters advance strong claims about efficiency and quality-of-hire gains. Simultaneously, high-profile failures have made algorithmic hiring a focal case in debates about automated discrimination and its regulation. This paper reviews the multidisciplinary literature on AI in recruitment — spanning human resource management, information systems, computer science research on algorithmic fairness, and the emerging regulatory scholarship — to assess what is credibly known about its benefits, its risks, and the conditions that separate the two. The review finds robust evidence for process-efficiency gains but thin and mixed evidence for quality-of-hire improvement; a well-established taxonomy of bias mechanisms (training data bias, proxy discrimination, and feedback loops) with documented instances in deployed systems; and an emerging governance literature converging on auditability, human oversight, and outcome monitoring as the practices that condition whether adoption helps or harms diversity outcomes. The paper develops a governance-centred framework for HR practice, maps it against incoming regulation including the EU AI Act's classification of employment AI as high-risk, and sets out a research agenda focused on the gap between vendor claims and independently verifiable outcomes.
Mehmoona Akram, Syeda Fatima Hussain, R. Anwar· Journal of Management Resear...· 0 citations