Skip to content
Preprint

Automating and Scaling Behavioral Scientific Research on AI Agents

Aug 2026 · 0 citations · 72 references
Computer Science

TL;DR

AEROBAT is introduced, the first multi-agent system to automate behavioral scientific research on AI agents and demonstrates that automated behavioral scientific research on AI agents can complement and extend the reach of manual research.

Abstract

As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains manual and labor-intensive. We introduce AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents. Given an arbitrary target behavior by its user, AEROBAT automatically executes a full pipeline of behavioral scientific research---generating hypotheses about the behavior, designing and executing controlled experiments, making behavioral assessments, analyzing the results, and writing reports. For 12 target behaviors, we used AEROBAT to generate and test 73 hypotheses: designing 1,160 controlled experiments and executing 22,954 simulation rounds in total. Moderate-to-strong statistical evidence was found for 30 hypotheses, including some novel ones. In sum, our results demonstrate that automated behavioral scientific research on AI agents can complement and extend the reach of manual research.

View source

Similar papers

Preprint Aug 2026

Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems

This work argues that studying AI Scientists as human-agent systems (HAS) is both underexplored and undervalued, and calls for new research that adopts the HAS lens to develop mathematical frameworks for understanding and fostering human-AI synergy in scientific discovery.

P. Emami, Sameera Horawalavithana, T. Nguyễn et al. · 0 citations
Preprint Jul 2026

An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery

Large language model coding agents increasingly perform open-ended data modeling and analysis. These agents are stochastic and adaptive, and therefore their autonomous model discovery behavior cannot be adequately characterized by a single benchmark run. In this work, we propose an experimental design and analysis framework for systematically evaluating this discovery process, quantifying its variability, and identifying important factors. The proposed framework treats these agents as stochastic model-discovery operators, which map task-specific discovery data and an optimization target to a fitted model. Specifically, we investigate two such operators, Codex and Claude Code, under controlled experimental factors including agent's reasoning effort, task, optimization metric, and composition of training data. For each agent-task-metric combination, regression models and inference are conducted for multiple responses such as output quality, dollar cost, wall-clock time, and process complexity. Furthermore, we develop a utility-aligned canonical decomposition to characterize the dominant direction of the reasoning-effort effect and to assess whether that direction aligns with a performance-cost utility direction. The proposed framework is demonstrated on a testbed of networked word-forming games with insightful findings on reasoning effort with respect to cost and process complexity.

Hao He, Xueying Liu, C. Kuhlman et al. · 0 citations
Review Aug 2026

The Past and Future of AI Scientists

A survey of the past and future of AI Scientists: machines capable of automating science, which have the potential to transform science and create a new form of science that will create a new form of science and transform the world.

Ross D. King · 0 citations
Review Jul 2026

Beyond Component Testing: Validating Agentic AI Systems

This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems, and concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures.

Fabio Orazio Mirto, L. D’Agati, Giuseppe Tricomi et al. · 0 citations