Verdicts Without Annotated Evidence: Rejection Sampling or Label-Only Post-Training for Evidence Recovery?

Title: Verdicts Without Annotated Evidence: Rejection Sampling or Label-Only Post-Training for Evidence Recovery?

标题:没有标注证据的裁决:拒绝采样还是仅标签后训练用于证据恢复?

Abstract: In many review workflows the verdict is the only thing retained. The passages behind it are not marked, because that annotation costs far more than recording the decision. We measure how much of that evidence a small language model can recover when it is post-trained on the verdicts alone, with no human evidence labels at any stage.

摘要:在许多审查工作流程中,裁决是唯一被保留的内容。其背后的段落并未被标记,因为这种标注的成本远高于记录决策的成本。我们衡量了一个小型语言模型在仅针对裁决进行后训练(且在任何阶段均无人为证据标签)时,能够恢复多少此类证据。

On ContractNLI the human evidence spans are held out until evaluation. Matching the recorded verdict and agreeing with those spans are not the same thing: across six systems the two scores are only weakly related and rank the systems differently, so accuracy is a poor guide when the citations have to be reviewable.

在 ContractNLI 数据集上,人类证据跨度被保留直到评估阶段。匹配记录的裁决与认同这些跨度并非同一回事:在六个系统中,这两个分数仅有微弱的相关性,且对系统的排名也不同,因此当引用必须可审查时,准确率是一个糟糕的参考指标。

Label-only training on the bare verdict reaches accuracy 0.896 and span F1 0.564. Rejection sampling, which keeps a generated trace only when its verdict matches the record and then picks one by an automatic source-grounding score, reaches 0.797 and 0.556, against 0.747 and 0.493 before training.

仅基于裁决的标签训练达到了 0.896 的准确率和 0.564 的跨度 F1 分数。拒绝采样方法(仅在生成的追踪结果与记录的裁决匹配时保留,并通过自动来源依据分数进行选择)达到了 0.797 和 0.556,而训练前分别为 0.747 和 0.493。

Verbatim citation rises from 0.597 to 0.729 under label-only training and to 0.701 under rejection sampling. One seed on one corpus cannot say which method is better, but both improve the evidence without anyone annotating it.

逐字引用率在仅标签训练下从 0.597 上升至 0.729,在拒绝采样下上升至 0.701。单一语料库上的单一种子实验无法断定哪种方法更好,但两者都在无需人工标注的情况下改善了证据质量。