Skip to content
Review Open access

Evaluating AI “Understanding” with Cognitive-Psychological Criteria: Evidence, Gaps, and Structural Limits

Jul 2026 · Journal of Language, Culture and Education · Vol 3, pp. 50-62 · 0 citations · 26 references

TL;DR

This article clarifies the concept definitions and evaluation criteria of understanding in cognitive psychology by combining classic theories and experimental evidence, and uses these criteria as the analytical framework for the performance of “similar understanding” in contemporary artificial intelligence systems.

Abstract

With the rapid development of large-scale language models, the performance of artificial intelligence in language understanding and reasoning tasks has become increasingly strong. As a result, many people have gradually regarded the success of tasks as the same as true understanding. However, from the perspective of cognitive psychology, understanding is not merely about giving correct or fluent responses; it is actually an internal psychological process. It includes aspects such as meaning construction, context model update, reasoning generation, and metacognitive regulation. Under such circumstances, this paper systematically studies the cognitive understanding ability of artificial intelligence from the perspective of cognitive psychology. It employs the methods of literature review and qualitative case analysis. First, it clarifies the concept definitions and evaluation criteria of understanding in cognitive psychology by combining classic theories and experimental evidence, and then uses these criteria as the analytical framework. Study the performance of “similar understanding” in contemporary artificial intelligence systems. It was found that although artificial intelligence can approach human understanding at the behavioral level, it has not reached the core cognitive psychological standards. AI lacks experience-based situational models, causal and goal-oriented reasoning mechanisms, and inherent metacognitive monitoring. These limitations are structural, not quantitative issues. It reflects the fundamental difference between artificial intelligence systems and human cognition. This article can help clarify the theoretical distinction between surface manifestations and true understanding, and also provide a cognitive psychology foundation for a more cautious explanation of artificial intelligence capabilities.

Read PDF

Similar papers

Aug 2026

Toward a Unified Theory of Understanding: A Conceptual Analysis for Natural and Artificial Understanding

Understanding is the foundational component of intelligence that makes adaptation to novel states possible. As artificial intelligence models advance toward artificial general and superintelligence, an objective, in contrast to human-centered, theory of understanding is required. One merit of an objective conceptual analysis of understanding is that it would unify human and artificial understanding. This study defines understanding as well-integrated use of reasoning across a network of conceptual, logical, and causal relationships, which are situated on an efficient conceptual web. A central evolutionary function of understanding is the capacity to generalize knowledge into novel situations. This renders machine learning’s concept of “generalization” a measure of understanding. The paper discusses two primary methods for assessing artificial understanding: behavioral methods measuring generalizability and mechanistic analyses revealing causal mechanisms of the neural network. We discuss the limitations of these methods. The most important problem, as this paper defends, is the difficulty of measuring objective understanding in contrast to human understanding.

Hasan Çağatay · 0 citations
Review Open access Jul 2026

When Can AI Models Explain Learning? Validity Criteria for AI as Cognitive Models in Education

Traditional verbal theories in educational psychology often remain underspecified at the mechanistic level. While they offer rich descriptive constructs and conceptual insights, they provide limited accounts of how learning processes unfold dynamically and causally. Here, “verbal” denotes theories stated in prose and qualitative relations rather than as formal, computational process models. This lack of mechanistic precision constrains rigorous theory testing, limits integration across levels of analysis, and reduces the potential to design interventions grounded in explanatory understanding of learning processes. In this Review, we synthesize an emerging paradigm that treats artificial intelligence (AI) systems not merely as predictive tools or instructional technologies, but as cognitive models of learners, explicit, runnable instantiations of theoretical assumptions about cognition and learning. Our central question is not whether AI systems can serve as cognitive models, but when they should be allowed to count as such. We therefore organize the Review around explicit validity criteria, theoretical grounding, construct validity, mechanistic transparency, alignment with human learning trajectories, error-signature matching, causal-intervention tests, ecological validity, and instructional usefulness, that an AI system must satisfy before its cognitive-model status is granted rather than assumed. We examine how major families of AI models, including neural networks, reinforcement learning agents, cognitive architectures, and large language models, have been used to operationalize core educational constructs such as memory, strategy use, motivation, self-regulation, and social learning. Across these approaches, we highlight how mechanistic transparency, interpretability, and alignment with human learning trajectories and error patterns are essential for explanatory validity. We further discuss methodological tools, such as representation analysis, ablation, and trajectory-level comparison, that enable causal inference about learning mechanisms within models. Finally, we outline key challenges and future directions, including construct validity, ecological realism, individual differences, and ethical accountability. By positioning AI as a theoretical instrument rather than solely an engineering solution, this Review argues that AI-based cognitive models can, when they satisfy these criteria, help transform abstract learning theories into precise, testable, and educationally actionable accounts of how students learn.

Peng Wang, Olga Viberg, E. Law et al. · 2 citations
Open access Aug 2026

Relevance and artificial intelligence

This article explores the distinct characteristics of human intelligence in comparison to Artificial Intelligence (AI), highlighting the concepts of relevance and meaning giving as fundamental properties of intelligent thought and action. Highlights the limitations of human cognition, such as limited time, computational power, and communication, which shape human intelligence. The text argues that relevance is crucial for understanding human and AI intelligence, as it determines the meaningfulness of actions and thoughts in various contexts. The article also explores the distinctions among formal languages, computational languages, and natural languages, asserting that formal languages cannot express relevance because of their formal nature. It emphasises the importance of human mental processes in assigning meaning to information and the need for relevance in AI systems for practical applications. Furthermore, the article examines the role of physical symbol systems in modelling human thought and the limitations of AI in replicating human-like intelligence. It critiques the assumptions underlying AI, such as the analogy between brain processes and digital computation, and discusses the challenges of defining relevance in mathematical and formal systems. Ultimately, the article concludes that relevance is essential for the development of effective AI systems, as it guides the selection of pertinent information and actions in specific contexts. It posits that, while the number of computational devices is vast, the construction of general AI remains unattainable because it requires defining specific semiotic action spaces for meaningful operations.

P. Saariluoma, Matthias Rauterberg · 0 citations
Review Open access Jun 2025

From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology.

Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human instruments to LLMs without establishing that the outputs are reliable or interpretable, raising the risk of measurement phantoms-statistical regularities mistaken for genuine psychological phenomena. This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant. It develops a dual-validity framework in which evidentiary demands scale with scientific ambition: from tool use through behavioral characterization and human simulation to cognitive modeling. Classifying text may require only accuracy and reliability; claiming that an LLM simulates anxiety or illuminates cognitive mechanisms requires additional evidence, including construct validity evidence and experimental controls. Progress depends on developing computational analogs of psychological constructs rather than assuming human measures automatically apply to language models.

Zhicheng Lin · 8 citations · ⚡1
Open access Aug 2026

On (artificial) intelligence: from a concept of measurable human ability to automated evaluative order

This article argues that intelligence should not be understood as a concrete and delimited phenomenon or as a natural kind, but as a historically and culturally produced concept through which selected behaviors, capacities, and performances come to be classified as intelligent. It proposes to shift the question from what intelligence “really is” to how certain performances come to count as intelligent, how they are made measurable, and how those measurements acquire social, political, and technical force. The article’s contribution is to connect the historical operationalization of intelligence with the evaluative infrastructures through which AI reproduces and transforms it. Tracing a genealogy from modern rationality, psychometrics, and standardized testing to artificial intelligence, the article shows how intelligence became actionable through classification, comparison, measurement, and evaluation. Once test scores shaped access to education, credentials, work, and social recognition, measurement no longer merely described ability; it became part of the institutional conditions through which ability was recognized, rewarded, and made socially consequential. Testing thus helped produce feedback loops in which measured intelligence partly reflected the opportunities that testing itself had helped allocate. Intelligence was then linked to merit, qualification, and deservingness, allowing historically contingent criteria of achievement and selection to appear not as products of unequal opportunity, but as natural differences in ability. The article then shows how artificial intelligence inherits this operational history. AI became possible once intelligence had already been reformulated as performance that could be formalized, evaluated, and reproduced apart from the living human subject. Cybernetics, information theory, and early AI translated this conception into computational terms, while contemporary machine learning relocates it into data curation, task definition, model architectures, objectives, metrics, and benchmarks. Although deep learning departs from explicit rules and predefined symbolic representations, it does not escape operationalization: curated datasets and task-specific objectives shape the learned latent spaces through which relations become detectable, comparable, rankable, and optimizable. AI therefore crystallizes historically specific conceptions of intelligence by embedding them in technical systems of training, measurement, comparison, and evaluation. Its social consequences emerge when these evaluative infrastructures shape which performances become recognizable, rewarded, and normalized. The question is therefore which conceptions of intelligence are being technically reproduced, made authoritative, and extended through AI systems.

Juan Sebastian Olier · 0 citations