Who Wrote the Ballot: The Judgment That Happens Before the Judge Is Called
本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。
Who Wrote the Ballot: The Judgment That Happens Before the Judge Is Called. Two sites are on the shortlist for a second location. The rent is modeled, the foot traffic has been counted for a month, the fit-out has three quotes, and the spreadsheet has been open for six weeks. Both sites work. The argument in the room is genuinely about which one, and it is a real argument, decided on real numbers. What nobody says out loud is that the option to not open a second location this year, and fix the throughput problem at the first one, was never written on the list.
谁编写了选票:在法官被传唤之前发生的判断。有两个地点被列入开设第二分店的候选名单。租金已经建模,人流量已经统计了一个月,装修方案有三份报价,电子表格也已经打开了六周。两个地点都可行。会议室里的争论确实是关于选哪一个,这是一个基于真实数据的真实争论。但没人说出口的是,今年不开设第二分店、而是解决第一家店产能问题的选项,从未被写在名单上。
That is not a failure of analysis. The analysis was careful. It is a failure at a step that happens before the analysis, that is almost never recorded as a step, and that has already determined more of the outcome than every row of the spreadsheet beneath it. The step is deciding what gets to be on the ballot.
这并非分析的失败。分析过程很严谨。这是一种发生在分析之前的步骤的失败,它几乎从未被记录为一个步骤,却已经比电子表格下方的每一行数据都更能决定最终结果。这一步骤就是决定什么可以进入选票。
Two acts, and only one of them can be bought. Split any decision into two acts. The first is nomination: which options are admissible. The second is ranking: of the options admitted, which is better. They are performed by different processes, at different times, with different evidence. Only the second one is what every judgment product on the market, ours included, actually does.
两个行为,只有一个可以被购买。将任何决策拆分为两个行为。第一个是提名:哪些选项是可接受的。第二个是排序:在已接受的选项中,哪一个更好。它们由不同的流程、在不同的时间、使用不同的证据来执行。市场上所有的判断产品,包括我们自己的,实际上做的都只是第二个行为。
A judge receives the option set as an input, the way a calculator receives numbers. It cannot widen the set, because the set is not in its input space. It cannot report that the set is wrong, because nothing in the interface has a slot for that finding. What it can do is answer the question it was handed, precisely and quickly. That is not a limitation anybody is concealing. It is the shape of the category.
法官(判断工具)接收选项集作为输入,就像计算器接收数字一样。它无法扩大这个集合,因为该集合不在其输入空间内。它无法报告集合是错误的,因为界面中没有任何位置可以容纳这种发现。它能做的是精确且快速地回答被交给它的问题。这不是任何人试图隐瞒的局限性,而是这一类别的固有形态。
Earlier in this series I argued that it is a virtue of typed judgment that the set of possible answers is defined by the caller. This piece is the invoice for that virtue. Every confidence number is conditional on the list. Read a probability the way a statistician reads it and the missing part becomes obvious. A confidence of 0.9 is not a property of the world. It is a claim that among cases resembling this one, and among the options you supplied, the answer named is right about nine times in ten.
在本系列文章的前面,我曾论证过,类型化判断的一个优点是可能的答案集由调用者定义。这篇文章就是这一优点的“账单”。每一个置信度数字都以列表为前提。像统计学家那样去解读概率,缺失的部分就会变得显而易见。0.9 的置信度并非世界的某种属性。它是一种声明:在与此类似的情况下,以及在你提供的选项中,所选出的答案有十分之九的概率是正确的。
The conditioning event is doing half the work, and it is invisible in the output. The number prints to two decimals either way. So a figure can be honestly measured, published, and still describe a comparison that should never have been run. The instrument is not lying. The question was malformed before the instrument was switched on, and no amount of care inside the comparison recovers an option that was never compared.
条件事件承担了一半的工作,但在输出中却是不可见的。无论如何,数字都会打印到小数点后两位。因此,一个数据可以被诚实地测量、发布,但它所描述的比较可能根本就不该进行。仪器没有撒谎。在仪器开启之前,问题就已经被错误地构建了,而在比较过程中无论多么小心,都无法找回那个从未被纳入比较的选项。
Our published measurements are a clean illustration of how tightly the conditioning binds. They are self-run on named benchmarks, with the failures disclosed rather than quietly retried. On JudgeBench: 620 judgments, with the 6 first-verdict failures disclosed. Raw accuracy came out at 92.5% against 92.2% for a plain direct baseline. That is a tie, and we publish it as a tie, with no accuracy claim resting on it.
我们发布的测量结果清晰地展示了这种条件限制有多么紧密。它们是在指定的基准测试上自运行的,失败案例被公开而非悄悄重试。在 JudgeBench 上:进行了 620 次判断,其中 6 次首次判决失败被披露。原始准确率为 92.5%,而简单的直接基准测试为 92.2%。这是一个平局,我们将其作为平局发布,并未基于此提出任何准确性声明。
The bands are the part worth keeping: calls we reported at 90% confidence or above came back right 99.6% of the time, and the 80–90% band 94.0%. On ContextualJudgeBench, self-run under the official pairwise protocol in which a pair counts as correct only if both presentation orders are judged correctly, which is why the random floor is 25% rather than 50%, consistent accuracy is 67.1% against the benchmark’s official reference of 65.4%, with 12 orders excluded after repeated platform failures and that exclusion stated beside the result. The deliberately constructed near-tie splits sit at 46–60%.
置信区间是值得保留的部分:我们报告置信度在 90% 或以上的判断,正确率为 99.6%,80–90% 区间的正确率为 94.0%。在 ContextualJudgeBench 上,按照官方成对协议自运行(只有两种呈现顺序都被正确判断时,该对才算正确,这就是为什么随机基准是 25% 而不是 50%),一致性准确率为 67.1%,而基准测试的官方参考值为 65.4%。在多次平台故障后,有 12 个订单被排除,且该排除说明已标注在结果旁边。刻意构建的近乎平局的拆分结果处于 46–60% 之间。
Now look through that paragraph for the row that would tell you whether the right answer was among the two candidates. It is not there, and it cannot be. Every figure measures the ranking act, on questions whose admissible answers were fixed in advance by somebody who does not appear in the measurement at all. That is a genuine measurement, and it covers roughly half of what a decision is.
现在,请仔细查看那一段,寻找能告诉你正确答案是否在两个候选者之中的那一行。它不在那里,也不可能在那里。每一个数字衡量的都是排序行为,而这些问题的可接受答案是由一个根本没有出现在测量过程中的人预先设定的。这是一个真实的测量,它涵盖了一个决策大约一半的内容。
Better ranking makes a bad ballot harder to challenge. There is a comfortable story in which framing is simply the next unmeasured frontier, and better instruments will eventually reach it. The uncomfortable version is worse than that. Improving the ranking act does not merely fail to catch a bad ballot. It launders it. A pick reported at 0.91, supported by a published band table, an exclusion count and an owner’s name, is a much harder thing to challenge in a review than a shrug.
更好的排序让糟糕的选票更难被质疑。有一个令人宽慰的说法:框架设定只是下一个尚未被测量的领域,更好的工具最终会触及它。但令人不安的事实比这更糟糕。改进排序行为不仅无法发现糟糕的选票,反而是在为它“洗白”。一个置信度为 0.91 的选择,辅以已发布的置信区间表、排除计数和所有者姓名,在审查中比一个简单的耸肩要难质疑得多。
The measured apparatus confers authority on the pair it was pointed at, and the pair inherited that authority by being the only pair on the page. The better the ranking, the more legitimate the framing looks. That is a risk of our own product, not only of somebody else’s, and it is the reason this piece exists. A careful comparison is the best disguise an unexamined list has ever had. A series that has spent two weeks asking judges to publish their curves should say plainly what a curve measures: the comparison, not the choice of what to compare.
被测量的装置赋予了它所指向的选项对以权威性,而这对选项仅仅因为是页面上唯一的一对就继承了这种权威。排序越好,框架看起来就越合法。这是我们自己产品(而不仅仅是他人产品)的风险,也是这篇文章存在的原因。仔细的比较是未经审查的列表所能拥有的最好伪装。一个花了两个星期要求法官发布其曲线的系列文章,应该清楚地说明曲线衡量的是什么:是比较本身,而不是对比较对象的选择。
Where the options actually go missing. Omissions are not random, and naming the patterns is most of the defense. The status quo is not printed. Doing nothing, keeping the current system, waiting a quarter, running a two-month pilot before committing — the most common real outcomes are rarely written on a ballot, because they look like the absence of a decision rather than a decision. Once they are off the page, the deliberation proceeds as though one of the listed items must win. And it usually does.
选项实际上是在哪里丢失的?遗漏并非随机,指出这些模式是防御的大部分工作。现状通常不会被打印出来。什么都不做、维持现状、等待一个季度、在做出承诺前进行为期两个月的试点——这些最常见的真实结果很少被写在选票上,因为它们看起来像是“没有决策”而不是“一种决策”。一旦它们离开了页面,审议就会继续进行,仿佛列表中的某一项必须胜出。而且通常情况下,确实如此。
The option that costs somebody in the room is the one missing. A missing item is disproportionately the one that would require a person present to conclude that an earlier call of theirs was wrong, or that a capability they own has quietly become the problem. Omission is a political act performed in analytical vocabulary, which is exactly why it survives review: no scoring rubric can flag a row that was never created. The third way gets read.
那个会让会议室里某人付出代价的选项,往往就是缺失的那一个。缺失的项目往往不成比例地指向那些需要在场者承认其之前的判断是错误的,或者承认他们拥有的某种能力已悄然成为问题的选项。遗漏是一种用分析词汇包装的政治行为,这正是它能通过审查的原因:没有任何评分标准能标记出一行从未被创建的数据。第三种选择被解读。