The Autonomous Agency Scale (AAS) is introduced, a behavioral framework that scores AI systems on a 0-5 lexicon across seven dimensions of agency: cognitive autonomy, temporal persistence, environmental agency, social agency, creative agency, self-awareness, and goal formation, each operationalized by falsifiable threshold tests.
Abstract
Existing AI measurement frameworks quantify cognitive capability, task automation, or catastrophic risk, but none measure autonomous agency: the extent to which a system behaves in a self-directed way. A system can saturate capability benchmarks while remaining entirely reactive, acting only when prompted and ceasing all activity when a task completes. We introduce the Autonomous Agency Scale (AAS), a behavioral framework that scores AI systems on a 0-5 lexicon across seven dimensions of agency: cognitive autonomy, temporal persistence, environmental agency, social agency, creative agency, self-awareness, and goal formation, each operationalized by falsifiable threshold tests. Every dimension is scored in two temporal bands: an Active band covering engaged, user-initiated activity, and an Ambient band covering idle periods. Ambient Level 4 is gated by the Idle-Gap Test, a counterfactual criterion (remove all triggers and observe whether internally derived activity persists) that separates self-direction from scheduled rule-following. We apply the scale to six contemporary systems spanning task agents (Claude Code, Manus, Hermes), consumer assistants (ChatGPT, Siri), and a persistent companion architecture (Airi). The two-band profile quantifies a boundary that single-score frameworks conflate: task agents reach Active composites of 2.3-2.4 while scoring 0.6-1.9 Ambient, with every idle-period behavior attributable to user-configured schedules, whereas the companion architecture, evaluated longitudinally, is the only assessed system whose idle-period behavior survives trigger removal. We discuss limitations, including single-rater provenance, developer-evaluator bias on the longitudinal assessment, and the partially operationalized self-direction boundary in the Active band.
Generative artificial intelligence (genAI) systems are increasingly integral to epistemic processes such as hypothesis generation, explanation construction, and decision-making. Although they reliably enhance performance, emerging evidence reveals a metacognitive dilemma: as external generative capacity increases, internal monitoring, calibration, and cognitive engagement may decline. This reflects a redistribution of cognitive control within distributed human-AI systems that cannot be explained by automation bias or reliance on algorithms alone. We propose the AIRIS (AI-Augmented Inquiry and Regulation in Hybrid Systems) framework to analyze this dilemma and specify where regulatory intervention can counteract it. AIRIS is a multi-level control allocation architecture specifying the conditions under which epistemic agency can be preserved in hybrid generative systems. Drawing on distributed cognition, cognitive load theory, multimedia learning, and self-regulated learning, it identifies seven interacting mechanisms through which hybrid cognition may become destabilized, from delegation and calibration drift to motivational-affective drift. Five regulatory operators (Anticipate, Interrogate, Reflect, Integrate, and Synthesize) target internal generative engagement at points of emerging instability. The architecture does not itself improve learning; it specifies what must remain in place for genAI-supported work to sustain understanding, whether through instructional design, teacher guidance, or learners'own regulation. We derive testable propositions concerning the seven mechanisms and the five operators, reframing AI augmentation as a problem of control allocation in distributed generative systems. Beyond theory, AIRIS offers a research agenda, a design framework for genAI-integrated learning environments, and a conceptual toolkit for the governance of hybrid human-AI cognition.
Jochen Kuhn, P. Gerjets, Ulrich Trautwein et al.· 0 citations
AI systems are increasingly designed in ways that lead users to perceive them as conscious. This paper provides a unified framework connecting empirical hallmarks of consciousness attribution to a structured risk taxonomy of Seemingly Conscious AI (SCAI), AI systems that exhibit hallmarks which elicit consciousness attribution from users. We survey the empirical literature to identify five such hallmarks of SCAI, spanning affective capacity, anthropomorphic features, autonomous action, self-reflective behavior, and social-interactive behavior. These provide observable, system-level proxies for this inherently subjective phenomenon, informing its design and enabling its empirical study. Drawing on this foundation, we develop a taxonomy of SCAI risks spanning risks to individuals, including emotional dependence and autonomy erosion, and societal-level harms, including human status erosion and political strife. We complement this conceptual analysis with an expert survey to assess the likelihood of each risk category. We find that risks to individuals, particularly emotional dependence and autonomy erosion, are already observable and rated as high probability, while societal risks, at a low probability, carry high potential severity and path-dependence. The single perceptual mechanism of consciousness attribution is shown to generate this heterogeneous risk surface. We then discuss the implications of these risks and map the multidisciplinary research gaps in this nascent field to inform its research agenda.
Ben Bariach, P. Schoenegger, M. Bhaskar et al.· AI and Ethics· 2 citations
Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While generative and agentic AI have demonstrated impressive capabilities across a wide range of tasks, many fundamental cognitive functions remain fragmented or weakly developed, limiting reliable operation over extended time horizons. This paper presents a taxonomy-driven survey of the major cognitive capability gaps that continue to constrain the development of Cognitive AI. The literature is organized around five dimensions: persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation. For each dimension, we review recent advances, identify recurring limitations, and discuss open research challenges. Building on these insights, we outline a conceptual Adaptive Cognitive Intelligence Architecture (ACIA) and examine emerging directions in cognition-centric evaluation. The proposed taxonomy provides a unified framework for organizing existing research, identifying unresolved challenges, and guiding the design of future cognitively capable systems. Together, the taxonomy, architectural perspective, and evaluation framework offer a roadmap for advancing AI systems that exhibit more reliable long-term reasoning, adaptive decision-making, and continual learning. The survey highlights key research opportunities toward more adaptive, reliable, and cognitively capable AI systems, providing a foundation for future progress toward Cognitive AI and, ultimately, Artificial General Intelligence (AGI).
Taye Akinrele, Sindhuja Penchala, Noorbakhsh Amiri Golilarz et al.· 0 citations
Artificial intelligence increasingly mediates work, learning, emotional support, decision making and social interaction, yet the central welfare question remains under-theorized: when does AI make human beings happier, and when does an apparent gain in the present become a loss in the future? This paper develops Dynamic Capability-Dependency Theory (DCDT), an integrative framework that treats AI-mediated happiness as a two-horizon process. The first horizon concerns acute affective relief, convenience and enjoyment. The second concerns stocks of competence, autonomy, human relatedness, meaning and dependency that accumulate through repeated use. Evidence from randomized trials, longitudinal studies and workplace deployments indicates genuine near-term benefits, including productivity gains, symptom reduction and temporary reductions in loneliness, but also shows that outcomes vary with design, usage intensity, autonomy, relational context and time horizon. DCDT formalizes these mechanisms as a dynamic state model and introduces the Temporal Well-Being Reversal condition, in which initially positive AI effects become negative after capability erosion, relational substitution or dependency accumulates. The paper derives testable propositions, specifies an empirical research program and introduces an AI-Happiness Impact Assessment that evaluates both immediate and durable effects. The central claim is not that AI is intrinsically happiness-enhancing or happiness-reducing. Rather, AI changes the production function of happiness by redistributing effort, agency, attention and relationships across time. High-value AI therefore should be judged not only by how well it satisfies a user now, but by whether repeated use leaves that user more capable, more autonomous, more connected to other humans and better able to pursue a meaningful life.
Kwan Hong Tan· Open Access Journal of Multi...· 0 citations
Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe what the ADS must demonstrably do once deployed within that domain. This gap between operating condition specification and behavioral validation represents a critical unresolved challenge in ADS safety assurance. This paper presents a structured, standards-grounded taxonomy of 21 behavioral competencies organized across three operational domains-Highway (HWY), Urban (URB), and Hub (HUB)-derived systematically from the PEGASUS six-layer model-based ODD. Each behavior is decomposed along longitudinal and lateral control axes and characterized against a four-property framework: Safety (gap maintenance, conflict avoidance, kinematic stability), Compliance (legal rules and behavioral norms), Comfort (rider dynamics and trust), and Efficiency (mission completion and product-level metrics). We further demonstrate that the crossing of ODD layer parameterizations with behavioral competency specifications yields concrete scenario families suitable for systematic behavioral testing and SOTIF coverage evidence. The taxonomy is grounded in AVSC00008202111, SAE J3237, and SAE J3016, and is validated as an operational specification layer through its deployment in a rule-enforced trajectory optimization system. The Hub domain is identified as a structurally distinct, underspecified domain warranting dedicated research attention.
Chaitanya Shinde, Hadi Hajieghrary, M. Hurtado· 0 citations
This study examines how the technological characteristics of Agentic AI relate to employees' psychological states and behavioral responses in organizations. Drawing on Job Demands-Resources (JD-R) theory and Conservation of Resources (COR) theory, we conceptualize Agentic AI technological characteristics (AITC) as a higher-order construct comprising autonomous planning, multi-tool orchestration, and proactive feedback, and propose that these characteristics function as job resources associated with greater cognitive surplus and psychological safety, which in turn relate to AI-human task reallocation. We further investigate whether these relationships differ between leaders and staff. A cross-sectional, anonymous online survey was administered via an online survey platform to participants recruited through a research panel using purposive nonprobability sampling; eligibility was restricted to adults with current or prior AI-related work experience, and screened, attention-checked responses yielded a final sample of 331 employees (171 leaders, 160 staff). All constructs were measured with validated scales adapted to the Agentic AI context on 7-point Likert scales. Data were analyzed using partial least squares structural equation modeling (PLS-SEM), applying a two-stage approach for the second-order construct, bootstrapping with 5,000 resamples for direct and indirect effects, and multi-group analysis (MGA) for leader-staff comparisons. The measurement model demonstrated acceptable reliability and validity, and the structural results were consistent with the theorized resource-gain model: AITC was positively associated with cognitive surplus, psychological safety, and task reallocation, with significant mediation pathways and stronger AITC-resource associations among leaders. By focusing on psychological mechanisms rather than technical performance, this study extends emerging Agentic AI research and offers practical implications for role-differentiated AI adoption strategies.
J. Han, Cheong Kim· Journal of Visualized Experi...· 0 citations