Sep 2026· Social science computer review· 0 citations· 66 references
TL;DR
A real-world AI evaluation framework focused on AI-in-use: how people actually appropriate, adapt, and work around AI systems in context, and what consequences follow over time is proposed.
Abstract
Organizations are adopting generative AI faster than the evidence base needed to govern it. Existing evaluation tools such as benchmarks, alignment scores, and safety tests were built for model development, not for judging whether systems will create value, introduce friction, or shift risk in specific real-world settings. As a result, there is little systematic evidence about how AI behaves once it is embedded in everyday work. This paper proposes a real-world AI evaluation framework focused on AI-in-use: how people actually appropriate, adapt, and work around AI systems in context, and what consequences follow over time. Instead of treating variability across users, tasks, and settings as noise to be controlled away, the framework treats that variation as the central source of deployment-relevant evidence. It sets out four design principles for producing decision-ready evidence at scale and proposes a shared evaluation architecture combining a structured observation environment, a metrics hub, and reusable consortium models that summarize system behavior across contexts. Rather than replacing traditional benchmarks, this framework adds a sociotechnical evidence layer that connects model capabilities to the organizational and practitioner level outcomes where deployment decisions are actually made.
It is argued that treating the human and the model as a single joint cognitive system is the central design principle for the next generation of decision systems.
Ashore-Onisemo Funmilayo· INTERNATIONAL JOURNAL OF SOC...· 0 citations
A three-arm experiment that varies the manner of AI use against a no-AI anchor and measures performance 15-20 min later, once the tool has been removed, results in an immediate near-transfer decrement rather than demonstrated lasting de-skilling.
Cheng-Hai Liu, Jia-Yao Guo, Rong-Xia Gao et al.· Frontiers in Psychology· 0 citations
A common assumption in human-AI collaboration is that better outcomes require more context. In knowledge work, however, some of the information needed to complete a task is difficult to access or express. In this paper, we introduce a typology developed through a Research-through-Design study in corporate goal-setting...
Aeneas Stankowski, Lisa Wiese, Malini B. Leveque· Adjunct Proceedings of the 1...· 0 citations
This work examines how six decision-support mechanisms affect engagement, trust, and collaborative task performance in a diabetes meal-planning scenario and argues for a contextual, balanced pairing of CFF and XAI design that accounts for interactivity, decision frequency, and task complexity.
Oliver Henderson· International Journal of Com...· 0 citations
A two-dimensional design space is introduced in which both dimensions are organised into five operational levels, making the coupling explicit and navigable, and six architectural tactics for adjusting a deployment’s position within it are proposed, offering a shared vocabulary for compliance-aware agentic AI design.
D. Safin, Dian Baltaa, Timon Sengewaldb et al.· EGOV-CeDEM-ePart 2026· 0 citations
Lifecycle assurance is synthesized through lifecycle assurance: a conceptual framing focused on producing evidence that data can support a specific AI claim when its influence may be embedded in model behavior, model-based judgments, or agent actions.
Hariharan Gopinath, Jan Bosch, H. Olsson· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.