Alignment Forecasting: Predicting Misalignment From Training Data
Alignment Forecasting: Predicting Misalignment From Training Data
对齐预测:从训练数据中预判模型失调
Abstract: Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training.
摘要: 在存在细微缺陷的数据上训练语言模型,有时会导致模型出现广泛的对齐失效。仅凭表面观察数据往往无法确定这种失效是否会发生,而目前我们只能在训练完成后,通过对最终模型进行审计才能发现问题。为了补充这种事后审计,我们引入了“对齐预测”(Alignment Forecasting):即在训练前预测对齐失效的任务。
Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets, and 16 failure modes.
给定一个目标模型、一个微调数据集以及诸如欺骗或谄媚等失效模式,预测器会输出微调导致该失效模式显著增加的概率。为了衡量对齐预测的进展,我们引入了 ALIGNMENTFORECASTBENCH,这是一个包含超过 5,000 个预测问题的基准测试,涵盖了 17 个目标模型、32 个数据集和 16 种失效模式。
Frontier models prompted directly perform poorly on ALIGNMENTFORECASTBENCH. We therefore propose a forecasting scaffold in which an LLM reads the dataset and rates how strongly and broadly it pushes the model toward misbehavior, and a simple learned model combines that rating with the failure mode’s base rate and the target model’s prior tendency.
直接提示前沿模型在 ALIGNMENTFORECASTBENCH 上的表现不佳。因此,我们提出了一种预测框架:由大语言模型(LLM)读取数据集,评估其在多大程度上、多大范围内推动模型产生不良行为,并由一个简单的学习模型将该评分与失效模式的基础发生率以及目标模型的先验倾向相结合。
This forecasts well above chance, and beats a model fine-tuned on the task and a simple forecaster allowed to see how weaker models behaved after fine-tuning on the same data. Its signals also flag problematic training examples that a frontier-model classifier misses. Filtering those examples out from real post-training data such as UltraChat results in more aligned models on our multiple-choice evaluation in most cases, though the benefit in open-ended conversations is unclear.
该方法的预测效果远超随机水平,且优于在任务上进行微调的模型,也优于能够观察较弱模型在相同数据上微调后表现的简单预测器。其信号还能标记出前沿模型分类器所遗漏的问题训练样本。在大多数情况下,从 UltraChat 等真实训练后数据中过滤掉这些样本,可以在我们的多项选择评估中获得更对齐的模型,尽管其在开放式对话中的益处尚不明确。
More progress is needed before forecasts can reliably guide training data curation in practice, but our results suggest that forecasting many alignment failures before training can be tractable in the SFT (Supervised Fine-Tuning) setting.
在预测能够可靠地指导实际训练数据整理之前,仍需取得更多进展,但我们的结果表明,在监督微调(SFT)场景下,在训练前预测许多对齐失效是可行的。