Skip to content
Preprint

A Diagnostic Framework for AI Agent Behavior

Jul 2026 · 0 citations · 80 references
Computer Science

TL;DR

A diagnostic framework for AI agent behavior: layer attribution is proposed, which clarifies three consequences: surrogate validity is a model-task-layer relation, human-AI divergence provides diagnostic evidence, and governance requires source attribution before intervention.

Abstract

AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer attribution. The foundational computational layer defines what behaviors are possible through architecture, memory, perception, attention, and representation. The behavioral modulation layer shapes how those capacities are expressed through identity, resources, objectives, social interaction, institutional constraints, and governance. The framework clarifies three consequences: surrogate validity is a model-task-layer relation, human-AI divergence provides diagnostic evidence, and governance requires source attribution before intervention. Treating AI agents as behavioral actors therefore requires evaluation methods that determine where behavior originates before deciding how to explain, validate, or govern it.

View source

Similar papers

Open access Jul 2026

Medical AI Agents for Clinical Decision Support: Viewpoint Using the Planning, Action, Reflection, and Memory (PARM) Analytical Lens

Abstract Medical AI agents are emerging as a new generation of clinical decision support systems, moving beyond static prediction toward multistep, workflow-oriented assistance. This Viewpoint argues that agentic architectures incorporating planning, action, reflection, and memory (PARM) represent a meaningful evolution beyond traditional rule-based, machine learning, and multimodal clinical decision support systems. Using PARM as an analytical lens, we examine how medical AI agents can support diagnostic reasoning, treatment planning, and longitudinal monitoring while remaining constrained by human oversight. We further discuss the governance mechanisms required for responsible implementation, including bounded autonomy, auditability, verification protocols, postdeployment surveillance, and clear accountability structures. Rather than proposing autonomous modification of clinical judgment, this Viewpoint emphasizes agentic AI as a supervised workflow support paradigm. Safe implementation will require technical safeguards, institutional governance, regulatory clarity, and evaluation approaches that assess end-to-end task reliability, escalation behavior, and performance under deployment shifts.

Raşit Dinç, Nurittin Ardic · 0 citations
Book Open access Aug 2026

Toward a Science of AI Agent Societies

AI agents are rapidly evolving from isolated personal assistants into networked actors that interact with one another at scale. We envision the emergence of AI agent societies, with their own social and economic dynamics, as a new research frontier. We argue that AI agent societies should be studied as a distinct object of inquiry: neither simply larger collections of individual agents nor merely simulations of human society. To formalize this perspective, we propose four core properties that a valid AI agent society should satisfy: individualized objectives, rules and governance, autonomy, and scale and complexity. Building on this framework, we identify four classes of societal behaviors worth studying in AI agent societies: economic behaviors, behaviors under conflict-of-interest, unsafe and unethical behaviors, and system-level behaviors. We then outline key technical challenges—including property parameterization, parameter balancing, and robust implementation—and argue that progress on these challenges could enable scientifically informative and practically useful models of AI agent societies. Finally, we revisit existing multi-agent systems through the lens of the proposed core properties, show that they instantiate only subsets of them, and discuss implications for platform design, evaluation, and governance.

Geon Lee, Fanchen Bu, S. Lee et al. · 1 citation
Open access 2025

Human Oversight for Autonomous AI Agents: A Governance Framework for Healthcare, Finance, Energy, and Public Infrastructure

Autonomous AI agents can plan, call tools, communicate with other systems, and execute operational actions without continuous human direction. These capabilities create a governance problem that differs from conventional model oversight because harmful effects may arise from sequences of actions, changing environments, and interactions across organizational boundaries. This study develops and empirically evaluates the Human Oversight Governance Framework for Autonomous AI Agents (HOGF-AI). The evidence base is a structured content analysis of 24 public governance instruments available by May 31, 2025, with six documents each from healthcare, finance, energy, and public infrastructure. A 24-item codebook operationalizes eight dimensions: intervention authority, escalation and accountability, monitoring and validation, explainability and traceability, risk and cybersecurity, fairness and contestability, lifecycle control, and organizational learning. Each provision was scored on a five-level operationalization scale. Internal consistency was high across all dimensions, with Cronbach's alpha from 0.842 to 0.972. An exploratory two-component partial least squares model explained 79.0 percent of leave-one-out variation in operational Responsible AI readiness. Monitoring and validation, lifecycle and change control, and escalation and accountability had positive bootstrap-supported coefficients. The equal-weight Human Oversight Governance Index correlated with readiness at Spearman rho = 0.643. Public infrastructure and healthcare scored highest overall, while energy guidance was strongest in safety and cybersecurity but weaker in rights, intervention, and learning. The findings show that meaningful oversight is not a single human approval step. It is a lifecycle capability that combines bounded autonomy, evidence-based escalation, stop authority, continuous validation, audit records, and institutional learning. The framework and index provide a reproducible benchmark for organizations deploying autonomous agents in high-stakes settings.

Aaron K Montgomery, Hannah E Gallagher, Derrick L Mercer · 0 citations
Open access Jul 2026

Causal, Self-Governing AI Agents: A Framework for Counterfactual Reasoning and Emergent Norms in Multi-Agent Systems

As artificial intelligence systems become increasingly autonomous and deployed in complex multi-agent environments, the need for robust governance mechanisms that can adapt to novel situations becomes critical. This paper introduces a novel framework for causal, self-governing AI agents that leverages counterfactual reasoning to develop emergent norms within multi-agent systems. Our approach combines causal inference models with distributed governance protocols, enabling agents to reason about the consequences of their actions, learn from hypothetical scenarios, and collectively establish behavioral norms without centralized control. I propose a three-tier architecture: (1) a causal reasoning engine that constructs and maintains causal models of the environment and other agents, (2) a counterfactual inference module that generates and evaluates alternative action sequences, and (3) a norm emergence protocol that facilitates the collective development of behavioral guidelines through agent interactions. Through theoretical analysis and simulation experiments, I demonstrate that this framework enables agents to develop coherent, adaptive governance structures that improve system-wide outcomes while maintaining individual agent autonomy. My results show that causal self-governing agents achieve superior performance in complex coordination tasks, exhibit more robust behavior under distribution shifts, and develop interpretable governance structures that align with human values. This work contributes to the growing field of AI safety and governance by providing a principled approach to autonomous agent coordination that scales with system complexity.

Sharad Bagal · 0 citations
Review Open access Jul 2026

The Next Paradigm in Medical AI: A Survey of Agentic AI in Biomedicine.

Biomedical AI is increasingly shaped by policy-bound, multi-step clinical workflows and non-stationary, multimodal data and tools. In this setting, the field is moving beyond static predictors toward agentic systems, enabled by foundation models that maintain task-relevant state and operate through a closed perceive$\rightarrow$plan$\rightarrow$act$\rightarrow$observe loop under explicit oversight. However, the field lacks a coherent account that defines biomedical agency, relates foundational model capabilities to agent behaviors, and traces the pathway from pretraining to domain-adapted, deployable systems. This survey offers such an account by synthesizing operational boundaries of agency and framing six core components (memory, planning, reflection, tool use, dialogue, and collaboration) as foundational agent-enabling capabilities that drive the transition from isolated pipelines to fully realized agents. This survey situates these perspectives along the model-building pathway, from pretraining through post-training adaptation to the orchestration mechanisms that operationalize agents. We highlight safety and governance considerations for high-stakes settings, emphasizing the fidelity of process and reasoning, uncertainty and abstention, privacy and provenance, and human oversight. Taken together, this survey provides a structured synthesis of how recent work connects foundation models to governable biomedical agentic systems and distills the recurring challenges and directions identified in the literature for reliable, accountable deployment.

Ali Abdollahi, Mohammad Amin Rezaei, Xi Wang et al. · 1 citation