CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
CogArena:大型语言模型认知能力结构的多方法评估
Abstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings.
摘要: 大型语言模型(LLM)的认知评分正越来越多地被总结为“按能力划分的概况”(per-ability profiles)。这些维度的评分应当在不同任务间表现出趋同性,对匹配的干预措施产生选择性响应,并能推广到用于定义这些维度的模型之外。我们引入了 CogArena,这是一个基于程序生成的 13 范式基准测试,围绕一套多方法框架构建,旨在确定认知任务评分在五个理论驱动的分组中何时值得被赋予维度标签。
Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction.
在 55 个开源权重模型中,几乎所有的范式相关性均为正,且一个共同轴解释了约一半的方差。分组内的优势较小,对评分敏感,且在不同模型家族间表现出不确定性。在一项针对来自六个家族的 12 个模型进行的独立冻结、完全交叉研究中,有针对性的支架(targeted scaffolds)显示出微小的匹配分组优势,但没有任何特定于支架的对比能在多重比较校正后保持显著,且选择性并未改善对未参与训练的模型家族的预测能力。
The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.
冻结确认标准(frozen confirmation criterion)未能通过。事后的替代措辞重复实验产生了更小的正估计值,且再次失败。总而言之,这些结果支持了一个边界结论:虽然理论对齐的提示词(Theory-aligned prompting)在测试电池中产生了一种微小的对角线趋势,但目前的证据尚无法建立稳定的五维能力概况。CogArena 提供了一种工作流程,在为模型评分贴上认知标签之前,将行为特征、协方差、匹配干预和跨家族预测结合起来进行综合评估。