Automating and Scaling Behavioral Scientific Research on AI Agents
Automating and Scaling Behavioral Scientific Research on AI Agents
自动化与规模化 AI 智能体行为科学研究
As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains manual and labor-intensive. 随着 AI 智能体越来越多地被部署在复杂环境中,理解其行为变得至关重要。然而,针对 AI 智能体的行为科学研究目前仍然依赖人工,且过程繁琐。
We introduce AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents. Given an arbitrary target behavior by its user, AEROBAT automatically executes a full pipeline of behavioral scientific research---generating hypotheses about the behavior, designing and executing controlled experiments, making behavioral assessments, analyzing the results, and writing reports. 我们推出了 AEROBAT,这是首个用于自动化 AI 智能体行为科学研究的多智能体系统。用户只需设定一个目标行为,AEROBAT 即可自动执行完整的行为科学研究流程——包括生成关于该行为的假设、设计并执行对照实验、进行行为评估、分析结果以及撰写报告。
For 12 target behaviors, we used AEROBAT to generate and test 79 hypotheses: designing 1,240 controlled experiments and executing 23,512 simulation rounds in total. Moderate-to-strong statistical evidence was found for 26 hypotheses, including some novel ones. 针对 12 种目标行为,我们利用 AEROBAT 生成并测试了 79 个假设:共设计了 1,240 个对照实验,并执行了 23,512 轮模拟。研究发现,其中 26 个假设具有中等到强有力的统计学证据支持,其中包括一些全新的发现。
In sum, our results demonstrate that automated behavioral scientific research on AI agents can complement and extend the reach of manual research. 总之,我们的研究结果表明,针对 AI 智能体的自动化行为科学研究可以作为人工研究的补充,并进一步扩展其研究范围。