Skip to content
Open access

Understanding Cognition-Induced Risks in Agentic AI Systems

Aug 2026 · IEEE Intelligent Systems · pp. 1-8 · 0 citations · 25 references
Computer Science

TL;DR

This work systematically analyzes risks induced by expanding cognitive capabilities, following a three-level framework defined by their cognitive scope, from physical cognition to social cognition, and finally to self-referential cognition.

Abstract

Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different domains, their cognitive engagement raises critical concerns for human society that remain insufficiently studied. To address this gap, we systematically analyze risks induced by expanding cognitive capabilities, following a three-level framework defined by their cognitive scope, from physical cognition to social cognition, and finally to self-referential cognition. We study their potential risks to human agency, autonomy, and control capability, corresponding to each cognitive level. We finally propose strategies to mitigate these risks and enhance the controllability of agentic AI systems, ensuring their long-term safe development.

Read PDF

Similar papers

Preprint Jul 2026

Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition

LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human cognition are therefore often seen as the result of anthropomorphic projection. We argue that this framing is mistaken. LLMs clearly differ from humans in important respects, including their physical substrate, learning history, and the environments with which they interact. These differences make it all the more striking that contemporary LLM-based systems converge with human cognition on a number of principles of cognitive organization with longstanding support in cognitive science. We identify structural correspondences across five dimensions: inferential organization, computational architecture, representational structure, prediction-driven learning, and reinforcement-learning-like mechanisms supporting goal-directed action. These correspondences support a broader model of intelligent cognition in which core principles long used to explain human intelligence also characterize contemporary LLM-based systems.

C. Sripada, Richard L. Lewis · 0 citations
Review Aug 2026

A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI

Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While generative and agentic AI have demonstrated impressive capabilities across a wide range of tasks, many fundamental cognitive functions remain fragmented or weakly developed, limiting reliable operation over extended time horizons. This paper presents a taxonomy-driven survey of the major cognitive capability gaps that continue to constrain the development of Cognitive AI. The literature is organized around five dimensions: persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation. For each dimension, we review recent advances, identify recurring limitations, and discuss open research challenges. Building on these insights, we outline a conceptual Adaptive Cognitive Intelligence Architecture (ACIA) and examine emerging directions in cognition-centric evaluation. The proposed taxonomy provides a unified framework for organizing existing research, identifying unresolved challenges, and guiding the design of future cognitively capable systems. Together, the taxonomy, architectural perspective, and evaluation framework offer a roadmap for advancing AI systems that exhibit more reliable long-term reasoning, adaptive decision-making, and continual learning. The survey highlights key research opportunities toward more adaptive, reliable, and cognitively capable AI systems, providing a foundation for future progress toward Cognitive AI and, ultimately, Artificial General Intelligence (AGI).

Taye Akinrele, Sindhuja Penchala, Noorbakhsh Amiri Golilarz et al. · 0 citations
Review Open access Aug 2026

Seemingly conscious AI risks

AI systems are increasingly designed in ways that lead users to perceive them as conscious. This paper provides a unified framework connecting empirical hallmarks of consciousness attribution to a structured risk taxonomy of Seemingly Conscious AI (SCAI), AI systems that exhibit hallmarks which elicit consciousness attribution from users. We survey the empirical literature to identify five such hallmarks of SCAI, spanning affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior. These provide observable, system-level proxies for this inherently subjective phenomenon, informing its design and enabling its empirical study. Drawing on this foundation, we develop a taxonomy of SCAI risks spanning risks to individuals, including emotional dependence and autonomy erosion, and societal-level harms, including human status erosion and political strife. We complement this conceptual analysis with an expert survey to assess the likelihood of each risk category. We find that risks to individuals, particularly emotional dependence and autonomy erosion, are already observable and rated as high probability, while societal risks, at a low probability, carry high potential severity and path-dependence. The single perceptual mechanism of consciousness attribution is shown to generate this heterogeneous risk surface. We then discuss the implications of these risks and map the multidisciplinary research gaps in this nascent field to inform its research agenda.

Ben Bariach, P. Schoenegger, M. Bhaskar et al. · 2 citations
Preprint Jul 2026

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

Embodied AI is moving from isolated perception or action modules toward physical agents that understand, plan under goals, act through robot bodies, monitor progress, and improve from experience. Existing systems address this loop only in parts: end-to-end policies generate actions but often lack spatial reasoning, planning, and execution assessment, while robot-agent systems orchestrate tools or specialists but do not learn a shared representation. This fragmentation limits general Physical Agentic AI. We present ACE-Brain-0.5, a unified embodied foundation model that organizes robot intelligence into five coupled functions: spatial perception, decision making, embodied interaction, self-monitoring, and self-improvement. Built on ACE-Brain-0, which established spatial intelligence as a shared scaffold across robot platforms, ACE-Brain-0.5 extends an understanding-centric model into a closed-loop foundation model. A single 8B backbone instantiates the first four functions: grounding objects and affordances, reasoning over 3D and egocentric spatial relations, decomposing instructions into subgoals, generating navigation and manipulation actions, and estimating progress for verification and recovery. To unify these capabilities without cross-task interference, we introduce SSR+, which extends Scaffold-Specialize-Reconcile with a Reactivate stage after task-vector merging. The fifth function, self-improvement, is realized by a companion framework that updates external execution state, including task schemas, spatial memory, and failure-recovery cases, from rollouts. Across fifteen benchmarks, ACE-Brain-0.5 improves over ACE-Brain-0 on 14 of 18 spatial perception and grounding benchmarks, achieves competitive navigation and manipulation performance, and provides strong progress estimation in ID and OOD settings. Together, these results mark an early step toward general Physical Agentic AI.

ACE-Brain Team Ziyang Gong, Haoming Gu, Zehang Luo et al. · 3 citations
Preprint Aug 2026

Assessing mentalization in humans and large language models

Mentalization - the ability to infer others'beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.

Aamir Sohail, Xintong Zhong, Arkady Konovalov et al. · 0 citations
Open access Aug 2026

Cognitive Entanglement: Toward a Developmental Framework of the Human-AI Coevolutionary Leap

Large language models have become routine participants in everyday cognition. Their role has widened from retrieval and text generation to helping users define problems, organize arguments, make judgments, and interpret themselves. Yet their cognitive consequences are strikingly divergent. For some users, generative AI appears to reduce critical engagement, independent judgment, and tolerance for difficulty. For others, the same class of systems becomes a medium for conceptual expansion, reflective questioning, and higher-order learning. This divergence cannot be explained by model capability alone. Mental effort is often treated as a cost to be reduced. Yet repeated delegation may also reduce opportunities to practice the processes required for independent judgment. The key issue is developmental: how sustained AI use changes users’ cognitive capacities over time. This perspective proposes cognitive entanglement as a framework for understanding the developmental consequences of sustained human-AI coupling. Cognitive entanglement refers to a relation in which human and AI activity become mutually shaping, irreducible to either party alone and organized across different developmental levels. The framework examines how repeated interaction with AI changes the ways users formulate problems, evaluate reasons, and make judgments. Unlike theories that locate the boundaries of cognition (the extended mind, enactivism) or explain the mechanisms of consciousness (global workspace, higher-order, predictive-processing, and integrated-information theories), cognitive entanglement examines whether sustained AI use preserves, weakens, or reorganizes users’ cognitive capacities. The article argues that current AI systems are often optimized for fluency, immediacy, and user satisfaction, and this may reduce the productive difficulty that supports higher-order cognitive development. If AI is to support human cognitive growth, design must move beyond answer provision and efficiency maximization toward the organization of productive human-AI relations: relations that challenge users’ initial assumptions while providing support appropriate to the task and the user’s level of expertise. The argument draws on philosophy of mind, cognitive science, and learning science, and compares divergent approaches to coupling in order to specify which forms of relation carry which developmental consequences. The concept shifts attention from AI as a tool or automation system to the developmental consequences of sustained human-AI interaction.

Xiao-kun Wu, Min Chen, Giancarlo Fortino · 0 citations