PAWS: Policy-driven Agentic World Simulation

PAWS: Policy-driven Agentic World Simulation

Abstract: Policy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence.

摘要: 政策干预通过公共传播、机构决策和利益相关者的反应进行传播,然而,用于金融多智能体模拟的数据集很少将这些过程与时间对齐的历史证据联系起来。

We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 verified U.S. financial and economic policy episodes, 12,727 policy-linked news records, and 65,291 source-grounded stakeholder actions.

我们推出了 PAWS(政策驱动的智能体世界模拟)数据集,涵盖了 36 个经过验证的美国金融和经济政策事件、12,727 条与政策相关的新闻记录,以及 65,291 个基于来源的利益相关者行动。

Each action is linked to its supporting news and represented by a multi-layer event frame capturing its interaction mode, financial-action family and subtype, semantic attributes, and conditional mappings to external taxonomies.

每一项行动都与其支持新闻相关联,并由多层事件框架表示,该框架捕捉了其交互模式、金融行动类别与子类型、语义属性以及与外部分类法的条件映射。

Entities are resolved to normalized organizations, and actions are aligned with daily market-return context to support policy-agent simulation replay.

实体被解析为标准化的组织,行动与每日市场回报背景对齐,以支持政策智能体模拟回放。

On 2,522 stratified action samples, independent AI and human reviewers achieved 89.4% initial agreement on interaction mode, with disagreements subsequently adjudicated.

在 2,522 个分层行动样本上,独立的 AI 和人类评审员在交互模式上达成了 89.4% 的初步共识,分歧部分随后经过了裁决。

Case studies of the 2008 short-selling ban and 2001 decimalization recover documented policy timelines and associated market patterns across both dense and sparse news settings.

通过对 2008 年卖空禁令和 2001 年十进制化(decimalization)的案例研究,研究人员在新闻密集和稀疏的情况下,均恢复了记录在案的政策时间表及相关的市场模式。

A replay study further shows that high accuracy can mask failure to detect rare stakeholder actions, identifying action timing and calibration as central challenges.

回放研究进一步表明,高准确率可能会掩盖对罕见利益相关者行动检测的失败,这表明行动时机和校准是核心挑战。

PAWS provides an auditable substrate for evaluating agent influence, policy-response cascades, and action-outcome alignment in historically grounded financial simulations.

PAWS 为评估基于历史背景的金融模拟中的智能体影响力、政策响应级联以及行动与结果的一致性提供了一个可审计的基础。