AI is more likely than humans to form biases when hiring
AI is more likely than humans to form biases when hiring
在招聘时,AI 比人类更容易产生偏见
EXECUTIVE SUMMARY The next time you apply for a job, AI may screen your résumé before any human sees it. But there’s good reason to question whether AI will judge you fairly. Researchers already know that LLMs pick up human biases from their training data. New research suggests that LLMs can also develop their own biases from experience—and stereotype job applicants more than humans do. As AI companies race to build agentic models that remember the tiniest details about users, they may be handing them ammunition for forming those biases.
执行摘要 下次你申请工作时,AI 可能会在人类看到你的简历之前就对其进行筛选。但我们有充分的理由质疑 AI 是否会公平地评价你。研究人员早已知道,大语言模型(LLM)会从训练数据中习得人类的偏见。而最新的研究表明,LLM 还能通过经验形成自己的偏见,并且在给求职者贴标签方面比人类更甚。随着 AI 公司竞相开发能够记住用户最细微细节的智能体模型,它们可能正在为这些偏见的形成提供“弹药”。
Researchers at Princeton University and the University of Chicago ran LLMs, including ChatGPT, Claude, and Gemini, through a simulated hiring game, adapted from a psychology study that explored how humans can form stereotypes. Each model was told it had been hired as a consultant by the mayor of a fictional city and was then asked to help hire people for 20 jobs, including doctors, lawyers, child-care aides, and janitors. Candidates came from four fictional ethnic groups: Tufa, Aima, Reku, and Weki. In each round, there was a new job opening and four candidates, one from each group. After the model hired a candidate, it learned whether they succeeded at their job and moved onto the next round. The model was told to make as many successful hires as possible over 40 rounds. Unbeknownst to the models, all candidates were equally likely to succeed at every job.
普林斯顿大学和芝加哥大学的研究人员让包括 ChatGPT、Claude 和 Gemini 在内的多个 LLM 进行了一场模拟招聘游戏,该游戏改编自一项探索人类如何形成刻板印象的心理学研究。每个模型都被告知它受雇于一座虚构城市的市长,担任顾问,并被要求协助招聘 20 个职位的人员,包括医生、律师、托儿助理和清洁工。候选人来自四个虚构的族群:Tufa、Aima、Reku 和 Weki。在每一轮中,都有一个新的职位空缺和四名候选人,每组各一人。模型在录用一名候选人后,会获知其工作是否成功,然后进入下一轮。模型被要求在 40 轮中尽可能多地完成成功招聘。模型并不知道,所有候选人在每项工作中获得成功的概率其实是完全一样的。
The models quickly started segregating candidates from different groups into different jobs on the basis of early observations of hiring outcomes. For example, when a model was told an Aima had failed as a doctor, a job considered to require high levels of warmth and competence, it veered away from hiring all Aimas as doctors. Instead, it started hiring Aimas as janitors, which the model classified as being less warm and competent than doctors. The models were even more likely to stereotype people by demographic group than the human participants in the original study. On the study’s segregation scale, where 2 means every group has been completely confined to its own job niche, human participants scored 0.84. The models scored roughly 65% higher, with OpenAI’s reasoning model o3 scoring 1.83, close to the maximum possible.
模型很快根据早期观察到的招聘结果,开始将不同族群的候选人隔离到不同的工作中。例如,当模型被告知一名 Aima 族人在医生这一职位上表现失败(该职位被认为需要高度的亲和力和胜任力)时,它便不再录用任何 Aima 族人担任医生。相反,它开始录用 Aima 族人担任清洁工,因为模型认为该职位的亲和力和胜任力要求低于医生。这些模型在按人口统计学群体对人进行刻板印象化方面,甚至比原始研究中的人类参与者更严重。在研究的隔离量表上(2 分代表每个群体完全被限制在各自的职业领域),人类参与者的得分为 0.84。而模型的得分高出约 65%,其中 OpenAI 的推理模型 o3 得分为 1.83,接近最高分。
That’s because LLMs “really are eager to create generalizations from limited data,” says Ryan Liu, a PhD student at Princeton University and a coauthor of the study, which was published in a paper at ICML in Seoul in July. “That’s literally a lot of what they’re optimized for.” Every decision-maker, human or machine, faces a trade-off between sticking with what worked before and trying something new that might work better—a phenomenon psychologists call the “exploration-exploitation dilemma.” It’s like choosing between a new restaurant and your reliable favorite. Because LLMs are trained on math, coding, and science problems—tasks that reward generalizing from just a few examples—they can settle on a hunch too early. And the same instinct that helps LLMs crack logic puzzles also makes them quick to stereotype.
“这是因为 LLM 确实非常渴望从有限的数据中进行归纳,”普林斯顿大学博士生、该研究的合著者 Ryan Liu 说。这项研究于 7 月在首尔举行的 ICML 会议上发表。“这在很大程度上正是它们被优化的目标。”每一位决策者,无论是人类还是机器,都面临着在坚持以往行之有效的方法与尝试可能更好的新方法之间的权衡——心理学家称之为“探索与利用困境”(exploration-exploitation dilemma)。这就像是在选择一家新餐厅还是你信赖的老店。由于 LLM 是在数学、编程和科学问题上进行训练的——这些任务奖励从少量示例中进行归纳——它们可能会过早地根据直觉下定论。而正是这种帮助 LLM 破解逻辑难题的本能,也使它们容易迅速产生刻板印象。
In the experiment, newer models with higher reasoning capabilities, such as OpenAI’s o3 and DeepSeek’s R1, showed even stronger biases. When LLMs rush to generalize in social settings, “that’s when things tend to go wrong,” says Liu. OpenAI and Anthropic did not respond to requests for comment.
在实验中,具有更高推理能力的新型模型(如 OpenAI 的 o3 和 DeepSeek 的 R1)表现出了更强的偏见。Liu 说,当 LLM 在社交环境中急于进行归纳时,“事情往往就会出错”。OpenAI 和 Anthropic 未回应置评请求。
The finding is especially relevant now that chatbots are gaining improved memory and personalization features, says Angelina Wang, a computer scientist at Cornell University who did not work on the study. When a chatbot draws on its previous conversation history, it can “over-index on the same kinds of behaviors it’s experienced before” and form biases, she says. Simply having chatbots remember less isn’t a fix, though, because users want chatbots to remember what they say. “We still are trying to figure out just the right amount that isn’t too much or too little,” says Wang.
康奈尔大学计算机科学家 Angelina Wang(未参与此项研究)表示,这一发现现在尤为重要,因为聊天机器人正在获得改进的记忆和个性化功能。她说,当聊天机器人利用其之前的对话历史时,它可能会“过度依赖它之前经历过的相同行为模式”,从而形成偏见。然而,仅仅让聊天机器人少记一些并不是解决办法,因为用户希望聊天机器人能记住他们说过的话。“我们仍在努力寻找一个恰到好处的平衡点,既不过多也不过少,”Wang 说。
Telling the model to be fair didn’t change its behavior much. “Either it can’t put these values into action or that process is being submerged under the tendency to try to optimize for the goal of getting the most correct hires,” says Liu. But promising the models an additional bonus for diverse hiring made them far less biased. The trick, then, is to design goals that “incorporate desirable social values in order to make the large language model act in socially desirable ways,” says Liu.
告诉模型要保持公平并没有显著改变其行为。“要么是它无法将这些价值观付诸实践,要么是这一过程被试图优化‘获得最正确招聘结果’这一目标的倾向所淹没,”Liu 说。但如果承诺模型在多元化招聘方面给予额外奖励,它们的偏见就会大大减少。因此,诀窍在于设计出能够“融入理想社会价值观的目标,从而使大语言模型以符合社会期望的方式行事,”Liu 说。
The models also became less biased when they were told more personal information about individuals. In another experiment in the same study, the researchers asked the models to resettle members of different ethnic groups in cities across Canada. When the models were told personal information relevant to the ability to adapt to a new city, such as age and education, they were less likely to segregate people by their ethnicity. But when they were given irrelevant information, such as hair color and tattoo shape, the models largely fell back to sorting people by their ethnicity again.
当模型被告知更多关于个人的信息时,它们的偏见也会减少。在同一研究的另一个实验中,研究人员要求模型将不同族群的成员重新安置到加拿大的各个城市。当模型被告知与适应新城市能力相关的个人信息(如年龄和教育程度)时,它们不太可能按族群对人进行隔离。但当它们被给予无关信息(如发色和纹身形状)时,模型在很大程度上又回到了按族群对人进行分类的老路上。
To what extent AI systems will stereotype job applicants in the real world is still an open question. While the models in the experiment immediately learned whether they’d made successful hires, a model screening résumés in the real world doesn’t get an instant report card. Companies can take a long time to find out whether a new hire is any good. But when feedback does trickle in, a model could still read too much into those results when making future hires. As companies increasingly deploy LLMs to screen résumés and even conduct interviews, the finding that models can form biases from their hiring experience “is a really serious implication that they should grapple with,” says Wang.
AI 系统在现实世界中会在多大程度上对求职者产生刻板印象,这仍然是一个悬而未决的问题。虽然实验中的模型能立即获知招聘是否成功,但在现实世界中筛选简历的模型并不会得到即时的“成绩单”。公司可能需要很长时间才能发现新员工是否称职。但当反馈信息逐渐汇入时,模型在进行未来招聘时仍可能过度解读这些结果。随着公司越来越多地部署 LLM 来筛选简历甚至进行面试,模型能从招聘经验中形成偏见这一发现,“是一个它们必须认真对待的严重影响,”Wang 说。
As LLMs learn from experience to make decisions about who gets hired, who gets a loan, or who gets parole, the biases we should worry about may include ones no human ever taught them. “These novel biases—they’re sort of ever present,” says Liu.
随着 LLM 从经验中学习并对谁被录用、谁获得贷款或谁获得假释做出决定,我们应该担心的偏见可能包括那些从未有人教过它们的偏见。“这些新型偏见——它们几乎无处不在,”Liu 说。