From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

📄 arXiv: 2609.04286v1 📥 PDF

作者: Ziyi Zhao, Guanzheng Wei

分类: cs.AI

发布日期: 2026-09-03

备注: 53 pages, 4 figures, 10 tables; companion literature-coding and search-log CSVs included in the source package


💡 一句话要点

系统化评估AI招聘系统以提升招聘效率与公平性

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: AI招聘 系统化评估 证据检索 候选人比较 透明性 可审计性 多阶段工作流

📋 核心要点

  1. 现有AI招聘系统在证据检索、候选人比较及决策支持方面存在不足,导致招聘效率和公平性受限。
  2. 本文提出了一种系统化的评估框架,旨在通过多阶段工作流提升招聘系统的透明度和可审计性。
  3. 通过对40项研究的分析,识别了招聘系统中的关键转变,并提出了改进建议,促进更有效的决策支持。

📝 摘要(中文)

人工智能在招聘中的应用已从简单的个人资料匹配和排名转向多阶段工作流,能够检索证据、比较候选人并支持或执行决策。本文系统化回顾了这一发展历程,涵盖了双边检索、行为排名、神经匹配、以及大型语言模型组件等技术。通过对40项代表性研究的分析,本文识别了招聘系统中的关键转变,并指出了当前方法的不足,如行为标签的混淆、数据的外部有效性限制等。最后,提出了一个基于证据的评估框架,以促进招聘系统的透明性和可审计性。

🔬 方法详解

问题定义:本文旨在解决当前AI招聘系统在证据检索和候选人比较中存在的不足,特别是行为标签的混淆和数据的外部有效性限制。

核心思路:通过系统化的评估框架,将招聘过程中的证据与决策支持相结合,提升招聘系统的透明度和可审计性。

技术框架:整体架构包括文档理解、检索、排名、评估、面试、候选人来源及人机交接等多个模块,形成一个多阶段的工作流。

关键创新:引入了从相似性到互惠适配的转变,强调了从模型到复合工作流的演变,及从离线预测到基于证据的评估的转型,显著区别于传统方法。

关键设计:在评估过程中,采用了多层次的证据分类,包括领域、配对、列表、案例、轨迹和结果级别的证据,确保评估的全面性与准确性。并且,关注隐私、效用、公平性和安全性等多维度的综合评估。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

本文通过对40项研究的系统化分析,识别出招聘系统中的多个关键转变,提出了基于证据的评估框架,旨在提升招聘决策的透明度和可审计性。具体结果显示,新的工作流能够更有效地检索证据并支持决策,提升招聘效果。

🎯 应用场景

该研究的潜在应用领域包括人力资源管理、招聘平台及相关法律合规领域。通过提升招聘系统的透明度和可审计性,能够有效改善候选人体验,促进公平招聘,进而提升企业的整体招聘效率与质量。

📄 摘要(原文)

Artificial intelligence in recruitment has shifted the object being automated from profile pairs and ranked lists to multi-stage workflows that retrieve evidence, compare candidates, and support or execute actions. This systematized narrative review traces that development from bilateral retrieval and behavioral ranking through neural person--job matching, large language model (LLM) components, and tool-using recruiting agents. Using a purposive search and coding protocol updated through 23 July 2026, plus targeted updates through 2 September 2026, we organize 40 representative works with supporting industrial and legal sources. This synthesis is not a prevalence estimate. We analyze three coupled transitions: from similarity to reciprocal suitability, from a model to a compound workflow, and from offline prediction to evidence- and productivity-aligned evaluation. Across document understanding, retrieval, ranking, assessment, interviewing, sourcing, and human handoff, we distinguish field-, pair-, list-, case-, trajectory-, and outcome-level evidence. Persistent gaps arise because behavioral labels confound exposure, preference, and qualification; private and synthetic data limit external validity; final-output scores conceal pipeline failures; and, within the coded set, privacy is not directly evaluated and no row jointly evaluates utility, fairness, privacy, and security. These observations describe the coded set rather than the field as a whole. We therefore introduce a staged mapping from evaluation evidence to the strongest defensible claim, together with an agenda for reciprocal, evidence-grounded, temporally controlled, selective, and auditable systems. Progress should be judged by whether workflows retrieve the right evidence, preserve uncertainty, support contestable decisions, and improve outcomes under explicit cost and risk constraints.