Skip to content

Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use

Sep 2026 · 0 citations · 43 references
Computer Science

TL;DR

A structured, seeded review of 24 focal empirical publications, finding no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure.

Abstract

Researchers assessing competent generative-AI use at work must choose among self-reports, objective tests, and measures of oversight and reliance. We conducted a structured, seeded review of 24 focal empirical publications, starting from the 2024 COSMIN-based review and adding a targeted update through 17 August 2026. We grouped the measures into four domains: knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents. In an exploratory meta-analysis, we pooled three direct subjective-objective correlations from one research program (REML r = .055; Hartung-Knapp 95% CI [-.047, .156]; combined reported N = 2,765). We could not resolve a discrepancy between the largest study's reported correlation and p-value, leaving its weight uncertain. Adding a synthetic mean of 12 cross-factor correlations from a fourth study gave r = .079 (95% CI [-.025, .181]). This sensitivity analysis concerns a broader comparison. From this small evidence base, we cannot establish a population correlation, validate workplace cutoffs, or justify substituting self-ratings for performance scores. We identified tests of foundation knowledge (AICOS-S and GLAT) and measures of verification, reliance, trust, and dependency. We found no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure; some cover subsets. We propose a four-layer workplace battery with non-compensatory decision rules, but have not tested its thresholds or whether it improves on other assessment approaches.

View source

Similar papers

Review Open access Sep 2026

Self-Regulated Learning in Generative AI-Assisted Academic Writing: A Systematic Review of ESL/EFL University Students

The growing use of generative AI writing assistants, such as ChatGPT, in ESL/EFL academic contexts has sparked debate about their influence on self-regulated learning (SRL). This systematic review examines 41 empirical studies published between 2021 and March 2026 to evaluate how university students engage with AI tool...

J. Indongo, S. Ithindi · 0 citations
Review Open access Oct 2026

Metacognition in AI-Supported Second Language Learning: A Systematic Review of Constructs, Measures, and Evidence (2016–2026)

Metacognition has informed research on second and foreign language (L2) learning for three decades, while AI tools have supported such learning for nearly as long. Yet, no review has examined how the field conceptualises and measures metacognition under AI mediation. This systematic review identified 60 eligible studie...

Hong Yi, Qiang Chen, Zhuo Wang · 0 citations
Review Open access Sep 2026

Generative AI and adaptive systems for customising help-seeking scaffolds: A systematic review

Help-seeking is crucial in self-regulated learning (SRL), but generic scaffolds often do not meet diverse learner needs. This review examines AI, including large language models (LLMs), in customising help-seeking scaffolds, their effects, and methodological and ethical constraints. Using SRL theory, the review analyse...

Jecha S. Jecha, Chimaobi Charles Igweani, Anum Banaras et al. · 0 citations
Open access Sep 2026

A Hierarchy of Inferability: The DIRECT Framework for Defining Human involvement in AI-Assisted Identification and Rating of Perceptual Indicators in Open Teacher Discourse

This study examines the conditions under which generative AI (GenAI) can support the identification and rating of perceptual indicators related to technology adoption in open-ended teacher discourse. Grounded in the Technology Acceptance Model (TAM), the study conceptualizes three levels of inferential demand: direct-e...

Adi Yaakov-Azaria, Merav Rotary-Saban, Anat Cohen et al. · 0 citations
Review Open access Sep 2026

Effectiveness of Argument-Driven Inquiry Combined with Peer Assessment in Developing Epistemic Cognition and Laboratory Report Writing Skills in General Chemistry: A Randomized Cluster Trial

We ran a cluster-randomized controlled trial with 228 undergraduate general chemistry students, separating what Argument-Driven Inquiry (ADI) and structured peer assessment each contribute alone, and what they contribute together, to epistemic cognition and laboratory report writing quality. Randomization was at the se...

Sarah Rasheed · 0 citations
#small language model Review Open access Sep 2026

Conditional validity in LLM-mediated L2 assessment: an argument-based systematic review and meta-analysis

Introduction Large language models (LLMs) are increasingly used for scoring and feedback in second-language (L2) assessment, yet the validity of the resulting interpretations remains contested. This review evaluated when LLM-mediated assessment is psychometrically and educationally defensible using an argument-based va...

L. Alghamdi, T. Alghizzi · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.