From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

从匹配模型到招聘智能体:AI 招聘系统、评估与治理的系统化叙事综述

Abstract: Artificial intelligence in recruitment has shifted the object being automated from profile pairs and ranked lists to multi-stage workflows that retrieve evidence, compare candidates, and support or execute actions. This systematized narrative review traces that development from bilateral retrieval and behavioral ranking through neural person—job matching, large language model (LLM) components, and tool-using recruiting agents.

摘要: 招聘领域的人工智能自动化对象已经发生了转变,从最初的个人资料配对和排名列表,演变为能够检索证据、比较候选人并支持或执行操作的多阶段工作流。本系统化叙事综述梳理了这一发展历程,涵盖了从双向检索和行为排序,到神经人岗匹配、大语言模型(LLM)组件,以及使用工具的招聘智能体等技术演进。

Using a purposive search and coding protocol updated through 23 July 2026, plus targeted updates through 2 September 2026, we organize 40 representative works with supporting industrial and legal sources. This synthesis is not a prevalence estimate. We analyze three coupled transitions: from similarity to reciprocal suitability, from a model to a compound workflow, and from offline prediction to evidence- and productivity-aligned evaluation.

通过截至 2026 年 7 月 23 日的定向搜索和编码协议(并辅以截至 2026 年 9 月 2 日的针对性更新),我们整理了 40 篇代表性文献,并结合了相关的行业和法律资料。本综述并非对行业普及率的评估。我们分析了三个耦合的转变:从相似性匹配到双向适配,从单一模型到复合工作流,以及从离线预测到以证据和生产力为导向的评估。

Across document understanding, retrieval, ranking, assessment, interviewing, sourcing, and human handoff, we distinguish field-, pair-, list-, case-, trajectory-, and outcome-level evidence. Persistent gaps arise because behavioral labels confound exposure, preference, and qualification; private and synthetic data limit external validity; final-output scores conceal pipeline failures; and, within the coded set, privacy is not directly evaluated and no row jointly evaluates utility, fairness, privacy, and security.

在文档理解、检索、排序、评估、面试、寻源和人工交接等环节中,我们区分了字段级、配对级、列表级、案例级、轨迹级和结果级的证据。目前仍存在持续的差距,原因在于:行为标签混淆了曝光度、偏好和资质;私有数据和合成数据限制了外部有效性;最终输出的分数掩盖了流程中的失败;此外,在所编码的研究中,隐私问题未得到直接评估,且没有任何研究同时评估了效用、公平性、隐私和安全性。

These observations describe the coded set rather than the field as a whole. We therefore introduce a staged mapping from evaluation evidence to the strongest defensible claim, together with an agenda for reciprocal, evidence-grounded, temporally controlled, selective, and auditable systems. Progress should be judged by whether workflows retrieve the right evidence, preserve uncertainty, support contestable decisions, and improve outcomes under explicit cost and risk constraints.

这些观察结果仅描述了所编码的文献集,而非整个领域。因此,我们引入了一种从评估证据到最强可辩护主张的分阶段映射方法,并提出了一项议程,旨在构建双向的、基于证据的、时间可控的、选择性的且可审计的系统。衡量进步的标准应当是:工作流是否检索到了正确的证据、是否保留了不确定性、是否支持可质疑的决策,以及在明确的成本和风险约束下是否改善了结果。