Register Bias in Complexity-Based Large Language Model Routing
Register Bias in Complexity-Based Large Language Model Routing
基于复杂度的语言模型路由中的语域偏见
Abstract: Large language model services increasingly route each query to one of several models of differing capability, using a cheap estimate of query complexity to send easy queries to small models and hard queries to large ones.
摘要: 大型语言模型服务正日益将每个查询路由至多个能力各异的模型之一,通过对查询复杂度的低成本估算,将简单查询发送给小型模型,将复杂查询发送给大型模型。
I show that this routing step is not register neutral: text written in a non-standard English register, African American English or the English of second-language writers, is systematically assigned a lower-capacity tier than a meaning-equivalent standard-English version of the same query.
我指出,这一路由步骤并非“语域中立”(register neutral):使用非标准英语语域(如非裔美国人英语或第二语言使用者的英语)编写的文本,在路由时会被系统性地分配到比含义相同的标准英语查询更低能力的层级。
The effect is driven by a specific, common routing signal, input length, because non-standard registers omit function words and thus look shorter and therefore simpler; other complexity signals do not carry it.
这种效应是由一个特定的、常见的路由信号——输入长度——所驱动的。因为非标准语域往往省略功能词,从而显得更短,进而被判定为更简单;而其他复杂度信号则不具备这种偏见。
I demonstrate the disparity on 37,704 authentic learner sentence pairs and on a controlled parallel corpus.
我在 37,704 对真实的学习者句子对以及一个受控平行语料库上验证了这种差异。
I then measure the quality consequence on a device, edge, and cloud model ladder and find that the harm is driven by pervasive model bias, every tier, including a frontier cloud model, answers non-standard-register queries significantly less accurately, while the marginal quality cost of the routing decision itself is not significant on this benchmark.
随后,我评估了在设备端、边缘端和云端模型梯队上的质量后果,发现这种损害是由普遍存在的模型偏见所驱动的:包括前沿云模型在内的每一个层级,在回答非标准语域查询时,准确率都显著降低;而路由决策本身带来的边际质量损失在此基准测试中并不显著。
Complexity-based routing thus compounds the exposure of the users that the models already serve worst.
因此,基于复杂度的路由加剧了那些本就处于模型服务最差端用户的劣势。