Skip to content
Preprint

Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems

Aug 2026 · 0 citations · 74 references
Computer Science

TL;DR

The Multi-Dimensional Assessment for AI Cognition (MAAC) is introduced, a theoretically grounded framework for shifting evaluation from what text-based AI systems produce to how they think, complementing existing outcome-based benchmarks with cognitively grounded, multi-dimensional evaluation.

Abstract

Evaluating artificial intelligence systems has historically relied on outcome-based benchmarks that measure task accuracy, robustness, or fairness. While indispensable, these benchmarks provide limited diagnostic insight into the underlying cognitive processes that generate performance-leaving critical questions unanswered about how AI systems reason, integrate memory, manage complexity, or avoid generating false information. This paper introduces the Multi-Dimensional Assessment for AI Cognition (MAAC), a theoretically grounded framework for shifting evaluation from what text-based AI systems produce to how they think. MAAC defines nine cognitively motivated dimensions: Cognitive Load, Tool Execution, Content Quality, Memory Integration, Complexity Handling, Hallucination Control, Knowledge Transfer, Processing Efficiency, and Process-Outcome Alignment. Each dimension is grounded in established cognitive science theory-drawing on Marr's tri-level hypothesis, Baddeley's working memory model, Sweller's cognitive load theory, and unified theories of cognition. Five theoretical analyses provide initial support for the framework's coherence and empirical testability: dimension-to-theory mapping; a coverage matrix assessing breadth and non-redundancy; a formal gap analysis relative to current evaluation practice; a worked diagnostic illustration; and a set of a priori interdependency predictions for future empirical testing. MAAC provides a theoretical and operational framework for principled process-level cognitive assessment of text-based AI systems, complementing existing outcome-based benchmarks with cognitively grounded, multi-dimensional evaluation.

View source

Similar papers

Open access Aug 2026

The critical-thinking paradox in generative AI-integrated learning: distinguishing efficiency from cognitive depth—a differentiated framework and testable propositions

Generative artificial intelligence (GenAI) has rapidly entered educational settings, yet a fundamental question remains unresolved: does GenAI enhance learning or does it improve immediate performance while reducing the cognitive activity on which durable learning depends? This Hypothesis and Theory article offers a th...

Jia-Yin Lin, N. M. Al-Hada · 1 citation
#artificial intelligence Review Open access Sep 2026

Thinking with AI or Thinking Less? A Narrative Review of Generative AI Dependence, Cognitive Offloading, and Human Cognitive Functioning

The new generation of digital assistance programs, including generative artificial intelligence (GenAI), is increasingly shifting from information retrieval to the generation of explanations, arguments, code, abstracts, and solutions. The extent to which GenAI support of human cognition can extend human capacity versus...

Deva Adithiya L.M.D, Samarawickrema N. S., Peiris J. A. H. L. et al. · 0 citations
Review Open access Sep 2026

Evaluation of Large Language Models as Tools, Models, and Partners in Creative Thinking Research: A Selective Narrative Review with the GCA Framework

Creativity research faces three persistent bottlenecks: divergent-thinking scoring is labour-intensive, cognitive models of creativity remain underspecified, and laboratory tasks fall short of real-world creative achievement. Large language models (LLMs) offer potential solutions, but the field lacks a structured frame...

Ke-Xin Huang, Chun-Lei Liu, Jia-Qin Yang · 0 citations
Open access Sep 2026

Adaptive judgment in the cognitive reflection test: A computational analysis.

The model generates the intuitive errors elicited by the bat-and-ball problem and explains why these errors can be seen as byproducts of adaptive cognition, and shows how such judgments can be understood as rational adaptations to the learning environment.

Yi-Jin Hu, Sudeep Bhatia · 0 citations
Open access Sep 2026

Process-Oriented, Behaviorally Anchored Assessment of Clinical Reasoning in Large Language Models and the Effect of Extended Thinking: Protocol for a Prospective, Multigroup, Comparative Study.

BACKGROUND Most clinical reasoning evaluations in large language models (LLMs) score only the final answer, usually multiple-choice accuracy, which explains little about model reasoning. Two developments stress this gap: reasoning-optimized models are now common, and several expose an explicit extended thinking control...

Adnan Agha, M. Jalil, Eram Anwar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.