FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

FinProBench:利用源自专业交付物的角色基础评估准则来评估金融 AI 智能体

Abstract: Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables. We introduce FinProBench, a benchmark for professional financial tasks, and Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives rubrics from deliverables produced by practitioners in the same role.

摘要: 评估金融 AI 智能体需要与实际专业工作相一致的评价标准。现有的评估准则方法通常从任务提示词或模型输出中提取标准,却忽略了仅在从业者交付物中可见的隐性标准。我们推出了 FinProBench,这是一个针对专业金融任务的基准测试,以及“角色基础准则构建”(Role-Grounded Rubric Construction, RGRC)——一种可复用的流水线,能够从同一角色从业者所产出的交付物中推导评估准则。

RGRC comprises four stages: Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation. Its rubrics capture tacit standards, distinguish quality levels, and transfer across tasks within a role. Before analysis, we classified 57 occupations by deliverable genre into 30 prior-rich conventional roles and 27 prior-sparse role-specialized roles.

RGRC 包含四个阶段:交付物收集、能力提取、准则合成和验证。其生成的准则能够捕捉隐性标准、区分质量等级,并可在同一角色内的不同任务间迁移。在分析之前,我们根据交付物类型将 57 种职业划分为 30 个“先验丰富型”常规角色和 27 个“先验稀疏型”专业角色。

Across all roles, Prompt-only nearly matches RGRC for conventional roles (89.2% vs. 90.7%), but RGRC substantially outperforms it for role-specialized roles (99.1% vs. 78.0%). This split indicates that prompt engineering can approximate rubrics when conventions are well represented in model priors, while professional grounding is essential for standards beyond those priors.

在所有角色中,对于常规角色,“仅提示词”(Prompt-only)方法与 RGRC 的表现几乎持平(89.2% 对 90.7%);但在专业角色中,RGRC 的表现显著优于前者(99.1% 对 78.0%)。这种差异表明,当模型先验知识中充分体现了行业惯例时,提示工程可以近似替代评估准则;而对于超出这些先验知识的标准,基于专业实践的评估则至关重要。

FinProBench is built from 1,723 curated deliverables spanning 57 occupations, 8 financial sub-industries, and 161 deliverable types, and releases an initial evaluation set of 20 complete tasks covering 20 roles in 7 sub-industries. With heterogeneous LLM judges and role-level rubrics, human deliverables rank first on average (73.7 vs. 70.3, 70.2, and 69.6 out of 100), while all four systems show overlapping 95% confidence intervals and complementary strengths. Reusing rubrics at the role level reduces estimated per-task construction effort by 6.7 times relative to authoring each rubric from scratch.

FinProBench 基于 1,723 份精选交付物构建,涵盖 57 种职业、8 个金融子行业和 161 种交付物类型。我们发布了包含 20 个完整任务的初始评估集,覆盖 7 个子行业中的 20 个角色。在使用异构大语言模型(LLM)裁判和角色级准则进行评估时,人类交付物的平均排名第一(满分 100 分,得分为 73.7,对比其他系统的 70.3、70.2 和 69.6),同时所有四个系统均显示出重叠的 95% 置信区间,且各具优势。在角色层面复用准则,相比从零开始编写每个准则,可将每个任务的预估构建工作量减少 6.7 倍。