Sep 2026· Zenodo (CERN European Organization for Nuclear Research)
Human-Automation Interaction and Safety
Abstract
Human-in-the-loop (HITL) architectures retain human verification in artificial intelligence (AI)-assisted decisions to preserve safety. Human presence does not guarantee AI-independent judgment. Multiple substantive checks can share one AI-dependent information pathway, and a Reciprocal Validation Spiral can reinforce that dependency when human approval becomes evidence for further AI trust and delegation. This paper treats independence as graded and defines a Hole in the Loop at a decision or verification point when formal HITL remains intact and zero substantive verification pathways meet a domain-specific minimum criterion for effective independence. Before such a threshold is operationalized, a limiting case is identifiable: every substantive check is downstream of the same AI output or AI-conditioned judgment, and no checker forms a pre-AI judgment or consults AI-independent evidence. The Spiral can generate, deepen, or stabilize a Hole, while a Hole can also exist without it. When the common AI-dependent pathway is systematically wrong, no effectively independent route remains to register the discrepancy before downstream action. The error can pass with multiple approvals and a completed audit trail while agreement rises, override rates fall, and procedural compliance appears complete. Correlated agreement can therefore be recorded as independent confirmation. The paper argues that calling oversight of a high-risk system meaningful carries an ethical obligation to preserve effectively independent verifiability at decision points that matter and to make its dependency structure auditable. The relevant unit of AI safety therefore moves from human presence to the preservation and observability of effectively independent verification pathways.
GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations of such models.
Xiaotian Zhang, Chun-yan Li, Yi Zong et al.· arXiv.org· 216 citations· ⚡17
Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.
Jinhe Bi, Yifan Wang, Danqi Yan et al.· arXiv.org· 73 citations· ⚡4
The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.
Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson· EUROMICRO Conference on Soft...· 64 citations· ⚡6
This paper designs Markov decision processes (MDPs) for different combinatorial problems and proposes to train conditional GFlowNets to sample from the solution space and demonstrates that GFlowNet policies can efficiently find high-quality solutions.
Dinghuai Zhang, H. Dai, Esmeralda S. Whitammer et al.· Advances in Neural Informati...· 59 citations· ⚡8
An empirical study on the current state of practice in artificial intelligence ethics is conducted by means of a multiple case study of five case companies, which indicates a gap between research and practice in the area.
Ville Vakkuri, Kai-Kristian Kemell, Joni Kultanen et al.· arXiv.org· 56 citations· ⚡6