Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
语言模型认为谁有能力?职业偏见的机制分析
Abstract: Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them.
摘要: 语言模型(LMs)通常能通过行为偏见评估,但目前尚不清楚它们是不再具备产生偏见的潜在关联,还是仅仅学会了不去表达这些偏见。
In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias into two measurement points: a model’s internal representation of a user’s competence, and its observable outputs.
在这项研究中,我们证明了即使在行为偏见不可见的情况下,表征偏见往往也是可检测的。我们引入了一个因果框架,将职业偏见分解为两个测量点:模型对用户能力的内部表征,以及模型的可观察输出。
We derive steering vectors for representations of user expertise, and verify that they causally mediate model behavior in both a question-answering task and a hiring task.
我们推导出了用于用户专业知识表征的引导向量(steering vectors),并验证了它们在问答任务和招聘任务中对模型行为具有因果中介作用。
Applying this framework to several open-weight models, we find that demographic attributes, such as gender, race, and socioeconomic status, influence a model’s representation of user expertise, even in cases where behavioral metrics detect no disparity between demographics.
将该框架应用于多个开放权重模型后,我们发现性别、种族和社会经济地位等人口统计学属性会影响模型对用户专业知识的表征,即使在行为指标检测不到不同群体间存在差异的情况下也是如此。
We show that these model representations can influence downstream behavior under intervention, suggesting failure modes that behavioral metrics alone may not detect.
我们证明了这些模型表征在干预下会影响下游行为,这揭示了仅靠行为指标可能无法检测到的失效模式。