A shared vocabulary, a design space, and evidence that measured monitoring and task performance are separable design targets are contributed, suggesting that measured monitoring and task performance are separable design targets.
Abstract
AI assistance places a metacognitive demand on users, who must judge their own competence and the system's. Yet designers lack comparative evidence on which interventions to choose, where to place them, and how to tell whether they worked. We elicited 30 interventions from 11 experts and, with prior work, organized them into a design space of time (when an intervention acts), level (whose competence is judged), and source (who supplies the monitoring cue). A between-subjects experiment (N = 917; 12 planning-and-organizing problems) compared a per-task reliability card, contrasting replies, pause points, and post-problem reflection against a baseline LLM assistant. Reliability cards and contrasting replies reduced estimation error and overconfidence and increased aggregate confidence discrimination. No task-performance improvement or average within-item discrimination gain was established. We contribute a shared vocabulary, a design space, and evidence that measured monitoring and task performance are separable design targets.
A generative AI teaching partner should support reasoning over supplying conclusions; however, this has not been tested against learning in an authentic course. Drawing on design-based research, we specify the position as a conjecture map and report a first design cycle in two graduate-level research methods courses. S...
F. Zahra, Wei Wang, Frances K. Harper et al.· 0 citations
It is suggested that while explainable ChatGPT effectively strengthens the understanding of specific, explained content, it does not confer immediate, generalized gains in metacognitive abilities or holistic writing outcomes.
Mohammad Mousazadeh· International Journal of Eng...· 0 citations
Introduction Generative AI is increasingly used by college students for explanation, evaluation, and decision support, raising the question of whether perceived reliance accurately tracks the extent to which AI advice shapes final judgments. Methods In a three-condition between-subjects experiment, 342 undergraduate st...
Shao-Wei Ren· Frontiers in Psychology· 0 citations
People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-L...
Robin Welsch, Michelle Rausch, Pascal Knierim et al.· 0 citations
Coding assistants raise task performance, but learners plan and monitor less. Giving less away, the usual fix, conflates two things: how much work a system carries (cognitive load) and what the learner must decide before help arrives (metacognitive demand). Our principle, preserved metacognitive demand, holds the secon...
Xin-Meng Hou, Yuxuan Weng, Chin-Hsien Yeh et al.· 0 citations
Collaborative academic reading depends on coordination that existing tools do not provide. Readers need to know when peers are confused, who can help, and when AI assistance would add value rather than interrupt. Two interaction problems remain unresolved: when AI should intervene (grounded in sustained behavioural sig...
Sakil Sarker, Yasaman A. Basti, Heidar Davoudi et al.· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.
MIT News · Artificial Intelligence· news.mit.eduJul 7, 2026
The professor of physics and inaugural director of the NSF AI Institute for Artificial Intelligence and Fundamental Interactions will lead LNS and continue his research in particle physics.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.