From Lexical Baselines to Agentic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework

From Lexical Baselines to Agentic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework

从词汇基准到代理式检索增强生成:基于 SFIA 框架的结构化技能与责任等级提取


Abstract: Automated skill extraction underpins workforce planning, yet most systems represent skills as flat labels with no notion of the responsibility level at which a skill is practiced. The Skills Framework for the Information Age (SFIA) captures exactly this dimension, defining 147 professional skills across seven responsibility levels, but no automated LLM-based extraction targeting SFIA has been reported.

摘要: 自动化技能提取是劳动力规划的基础,然而大多数系统仅将技能表示为扁平的标签,而缺乏对技能实践所处责任等级的概念。信息时代技能框架(SFIA)恰好捕捉到了这一维度,它定义了跨越七个责任等级的 147 项专业技能,但目前尚未有针对 SFIA 的自动化大语言模型(LLM)提取方案的相关报道。


We formalize the task as structured prediction of (skill, level) pairs from free text and ask three questions: how accurately can text be mapped onto SFIA’s closed vocabulary, which strategies reliably predict the level alongside the skill, and do agentic designs improve on simpler retrieval and prompting?

我们将该任务形式化为从自由文本中预测(技能,等级)对的结构化预测问题,并提出了三个问题:文本映射到 SFIA 封闭词汇表的准确度如何?哪些策略能够可靠地在预测技能的同时预测出等级?以及代理式设计是否比简单的检索和提示工程表现更好?


We evaluate five strategies (a lexical baseline, dense retrieval with LLM reranking, a zero-shot schema-constrained LLM, single-agent agentic RAG, and a three-agent retriever—matcher—verifier crew) against expert-mapped European ICT role profiles, all drawing on an SFIA 9 corpus built by a fully automated agentic pipeline that we release.

我们针对专家映射的欧洲 ICT 职位画像,评估了五种策略(词汇基准、带 LLM 重排序的密集检索、零样本模式约束 LLM、单代理代理式 RAG,以及由检索器-匹配器-验证器组成的三代理团队),所有策略均基于我们发布的一个由全自动代理流水线构建的 SFIA 9 语料库。


Retrieval-based matching identifies the most skills while generative strategies are markedly more precise; only strategies assigning the level as an explicit decision predict it reliably, with similarity-based selection more than twice as inaccurate; and the crew doubles latency without improving accuracy, so added agent roles do not automatically benefit closed-taxonomy matching. These results provide the first reproducible baseline for structured, level-aware skill extraction against SFIA.

基于检索的匹配识别出的技能数量最多,而生成式策略的精确度明显更高;只有将等级分配作为明确决策的策略才能可靠地预测等级,而基于相似度的选择方法其误差率高出两倍以上;此外,多代理团队在没有提高准确率的情况下使延迟翻倍,因此增加代理角色并不能自动提升封闭分类匹配的效果。这些结果为针对 SFIA 的结构化、具备等级意识的技能提取提供了首个可复现的基准。