HakemBench: A Turkish Benchmark of Typed Decisions
HakemBench: A Turkish Benchmark of Typed Decisions
HakemBench:土耳其语分类决策基准
Abstract: HakemBench is a Turkish benchmark of typed decisions, in which the model under test reads a text, a question and a fixed set of options and returns a probability for every option. Version 1.0 is released fully open under CC BY 4.0, with 2,346 items and 4,275 choice, yes/no and score questions in seven tracks (fact-check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support).
摘要: HakemBench 是一个土耳其语分类决策基准测试。在该基准中,受测模型需要阅读一段文本、一个问题以及一组固定的选项,并为每个选项返回一个概率值。1.0 版本已在 CC BY 4.0 协议下完全开源,包含 2,346 个条目以及 4,275 个选择题、是非题和评分题,涵盖七个领域(事实核查分类、教育、护栏机制、法律路由、内容审核、垃圾邮件与钓鱼检测,以及客户支持)。
One harness scores decision quality (macro F1), calibration (from the normalised Brier score) and selective automation (from the normalised area under the generalised risk-coverage curve), combines them by a geometric mean and reports intervals from 2,000 bootstrap draws; probes for option order, paraphrase, English translation and substituted names are reported alongside.
该基准测试框架通过决策质量(宏观 F1 分数)、校准度(基于归一化 Brier 分数)和选择性自动化(基于广义风险覆盖曲线下的归一化面积)进行评分,并通过几何平均值将这些指标结合,同时报告了 2,000 次自助抽样(bootstrap)的置信区间;此外,还报告了针对选项顺序、释义、英语翻译和名称替换的探测结果。
Most gold labels come from blind passes of one AI model family compared with the votes of a panel of large language models from other model families; they are not human-verified.
大多数黄金标签(Gold labels)来源于一个 AI 模型家族的盲测结果,并与来自其他模型家族的大型语言模型专家组的投票进行对比;这些标签未经人工验证。
On a board of 16 rows the leader scores a composite of 0.888 and the lab’s own model is 7th at 0.660. Its numbers are not blind. Earlier runs’ test results shaped its training data, so its guardrail, moderation and customer support numbers are flagged; with every model scored on the other four tracks only, its composite is 0.678, 6th of 16.
在包含 16 个条目的排行榜上,榜首模型的综合得分为 0.888,而该实验室自研的模型以 0.660 的得分位列第 7。该实验室模型的数据并非盲测结果。由于早期运行的测试结果影响了其训练数据,因此其在护栏机制、内容审核和客户支持方面的得分已被标记;若仅根据其他四个领域对所有模型进行评分,该实验室模型的综合得分为 0.678,在 16 个模型中排名第 6。