What I learned counting 600,000 live job adverts every night

What I learned counting 600,000 live job adverts every night

我在每晚统计 60 万条实时招聘广告中学到的事

I’m building TUNAI, a job-matching assistant. Under it sits a crawler that reads live job adverts straight from employers’ own applicant tracking systems (Greenhouse, Workday, Lever, Ashby, Oracle HCM and about fifteen others) across the UK, the US, Canada and Australia. Today that is about 600,000 live adverts, recounted every night. 我正在开发一个名为 TUNAI 的职位匹配助手。其核心是一个爬虫程序,直接从英国、美国、加拿大和澳大利亚雇主自有的申请人跟踪系统(如 Greenhouse、Workday、Lever、Ashby、Oracle HCM 等约 15 个平台)抓取实时招聘广告。目前,系统每晚会重新统计约 60 万条实时广告。

Once you hold that many adverts, you start noticing things no single job seeker could. I put the odd ones on a page that updates itself: jobs.tun-ai.com/insights. Here are a few, and then the three bugs that nearly made those numbers lie. 当你掌握了如此庞大的数据量,你就会开始注意到普通求职者无法察觉的现象。我将一些有趣的发现整理在一个自动更新的页面上:jobs.tun-ai.com/insights。以下是其中一些发现,以及差点让这些数据失真的三个程序漏洞。

The numbers 34,000 live adverts say they were posted more than a year ago, 6% of the adverts that give a date. The oldest claims 2010. Another 66,000 are more than six months old. 3,100 adverts have been reposted five or more times. One has been reposted 31 times. 数据显示,有 3.4 万条实时广告发布于一年前,占所有标注日期的广告的 6%。最古老的一条甚至标注于 2010 年。另有 6.6 万条广告发布超过半年。有 3,100 条广告被重复发布了 5 次以上,其中一条甚至被重复发布了 31 次。

Friday is the busiest day: 20% of UK and US adverts go up on a Friday, and only 4% at the weekend. In the UK, 31% of timed adverts appear between 2pm and 5pm. 2,500 US job titles contain an exclamation mark. The UK has 79. Three live titles still ask for a “rockstar”, and five for a “ninja”. The hype word that won is “champion”, with 140. 周五是最繁忙的发布日:20% 的英美招聘广告在周五发布,而周末仅占 4%。在英国,31% 的定时广告出现在下午 2 点到 5 点之间。美国有 2,500 个职位名称包含感叹号,英国则有 79 个。目前仍有 3 个职位在招聘“摇滚明星”(rockstar),5 个在招聘“忍者”(ninja)。最流行的热词是“冠军”(champion),出现了 140 次。

Every figure on the page carries the number of adverts it rests on and the caveat that must travel with it. The posting dates, for example, are the employers’ own, and plenty of the old ones are evergreen adverts kept open to collect CVs rather than to fill one job. The whole table is a CSV under CC BY, also on Kaggle, Hugging Face and Zenodo. 页面上的每一个数据都标注了其统计基数以及必要的说明。例如,发布日期是雇主自己填写的,许多陈旧的广告其实是“常青”广告,其目的并非为了填补某个具体职位,而是为了持续收集简历。整个数据表以 CSV 格式提供,采用 CC BY 协议,并已发布在 Kaggle、Hugging Face 和 Zenodo 上。

Three bugs that nearly lied to me

差点让我误判数据的三个漏洞

  1. The same job, listed twice by the same employer. Oracle HCM lets one company run several career sites, and each site serves the same requisitions. Stored naively, 43,704 of 61,111 eligible Oracle adverts sat in 15,035 groups of byte-identical adverts. Merging on “similar text” had once destroyed hundreds of real vacancies, so the fix is deliberately narrow: a job points at the oldest of its group only when the source, title, place and the advert text are all identical, and both rows stay in the corpus. Only one gets a public page.

  2. 同一雇主重复发布同一职位。Oracle HCM 允许一家公司运营多个招聘网站,而每个网站都发布相同的需求。如果简单存储,在 61,111 条符合条件的 Oracle 广告中,有 43,704 条属于 15,035 组字节完全相同的重复广告。曾尝试通过“相似文本”合并,结果导致数百个真实职位空缺被误删,因此修复方案非常谨慎:只有当来源、标题、地点和广告文本完全一致时,系统才会将职位指向该组中最旧的一条,且两行数据都会保留在语料库中,但仅有一条会生成公开页面。

  3. An index that was never used. Postgres partial indexes only apply when the query’s predicate matches the index’s. Our indexes said WHERE live; our queries said live IS true. To a human those are the same; to the planner they are not, so the vector (HNSW) and GIN indexes sat unused and a match took 50 to 130 seconds. If you use partial indexes, check idx_scan on each of them. Ours read zero.

  4. 未被使用的索引。Postgres 的部分索引(partial indexes)仅在查询谓词与索引谓词匹配时生效。我们的索引定义是 WHERE live,而查询语句是 live IS true。对人类来说这两者是一样的,但对数据库查询规划器来说却不同,导致向量索引(HNSW)和 GIN 索引处于闲置状态,匹配耗时高达 50 到 130 秒。如果你使用部分索引,请检查每个索引的 idx_scan,我们的数值显示为零。

  5. A reader that stopped too early. The “how many adverts state pay” figure said 94% of US adverts give no pay. Industry research says roughly half do, so I sampled 200 of our “no pay” US adverts by hand: 43 stated pay in text we already held, and 42 of those put it after the 600th character, which is exactly where our reader stopped. US adverts tend to put the salary range near the end. The figure is withdrawn until the fix is certified.

  6. 读取器截断过早。关于“多少广告标注了薪资”的数据显示,94% 的美国广告未标注薪资。行业研究表明大约有一半标注了,于是我手动抽样了 200 条被判定为“无薪资”的美国广告:发现其中 43 条其实在已有文本中包含了薪资信息,且有 42 条将薪资放在了第 600 个字符之后,而这正是我们读取器停止读取的位置。美国广告倾向于将薪资范围放在文末。该数据已撤回,直到修复方案验证通过。

The lesson I keep relearning: when your number disagrees with everyone else’s, check your reader before you publish the surprise. 我不断重温的教训是:当你的数据与他人的结论不一致时,在发布这个“惊喜”之前,请务必先检查你的数据读取逻辑。

What it’s for

它的用途

The counting exists because the product matches people to jobs: it ranks live adverts against your CV and keeps a shortlist current. There is also a free ATS CV checker that shows how an applicant tracking system reads your CV, no sign-in needed. If a number on the insights page looks wrong to you, I would genuinely like to hear it. 进行这项统计是因为我们的产品旨在实现人岗匹配:它根据你的简历对实时广告进行排名,并保持候选名单的实时更新。此外,我们还提供一个免费的 ATS 简历检查工具,无需登录即可查看申请人跟踪系统是如何读取你的简历的。如果洞察页面上的某个数字让你觉得不对劲,我非常希望能听到你的反馈。