From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

从检索到类型化决策:基于生物医学句子编码器的校准系统一模型

Abstract: Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models.

摘要: 类型化决策模型(Typed decision models)通过单次前向传播回答关于文本的模式约束问题,并返回用于阈值判定的概率。我们探讨了针对检索任务训练的生物医学句子编码器是否是此类模型的良好起点。

We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head (C) and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing.

我们提出了 SBERT2S1,它将 Sentence-Transformers 编码器转换为双编码器、交叉头(C)和先验融合残差(PFR)决策模型;同时还发布了 BIODECIDE(一个生物医学类型化决策套件)以及 MEDLINE-S1(源自 NLM 索引的 243k 条训练决策数据)。

Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-set sizes, retrieval training significantly helps PFR, which keeps the retrieval prior, in 10 of 15 comparisons, but helps C in one and hurts it in five.

在六对父级检索器组合中,检索训练改善了对内容相关选项的零样本匹配效果。微调后,其效果取决于模型头:在五对组合和三种训练集规模下,检索训练在 15 次比较中有 10 次显著提升了保留检索先验的 PFR 模型性能,但对 C 模型仅有 1 次提升,且有 5 次表现下降。

A matched grid of two heads and five training objectives shows that C outperforms PFR under every objective, and that the released RLCD recipe of open System One models trails cross-entropy by 2.5-3.0 points. The deficit stems mainly from its reward normalisation, which inflates the noisy score-function term 3.6-15-fold; an unbiased leave-one-out estimator recovers most of the gap.

通过对两种模型头和五种训练目标的匹配网格分析显示,C 模型在所有目标下均优于 PFR,且已发布的开放系统一模型 RLCD 配方比交叉熵损失落后 2.5-3.0 个百分点。这一差距主要源于其奖励归一化过程,该过程将噪声评分函数项放大了 3.6-15 倍;使用无偏留一法估计器可以弥补大部分差距。

After temperature scaling, no objective is clearly better calibrated than cross-entropy. We release the code, the MEDLINE-S1 labels and a model.

在进行温度缩放(Temperature scaling)后,没有哪种目标函数在校准效果上明显优于交叉熵。我们已公开了代码、MEDLINE-S1 标签及相关模型。