Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking

Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking

获 Andreessen Horowitz 支持的 Vals 旨在成为 AI 基准测试的行业标杆

Benchmarking has become the industry norm for how AI companies validate their models’ capabilities and, when the metrics swing in their favor, stand out from competitors and advertise their superiority. In other words, good benchmarks pretty much always mean good PR. 基准测试已成为 AI 公司验证模型能力的标准方式;当指标表现优异时,公司便能借此在竞争中脱颖而出,并宣传其技术优势。换句话说,优秀的基准测试几乎总是意味着良好的公关效果。

Unfortunately, companies have also figured out how to outwit legacy benchmarking systems — many of which are older, and not built to measure the capabilities of modern models. Vals, a startup formed in 2024, says that it is on a mission to fix this very imperfect system. 遗憾的是,企业也找到了绕过传统基准测试系统的方法——这些系统大多年代久远,并非为衡量现代模型的能力而设计。成立于 2024 年的初创公司 Vals 表示,其使命正是修复这一极不完善的系统。

In the span of less than two years, the company has established itself as a notable presence in the tech industry and, last year it managed to secure a seed round led by 8VC and Bloomberg Beta. Then, last month, after a period of rapid growth, it raised $40 million in a series A led by Andreessen Horowitz. 在不到两年的时间里,该公司已在科技界崭露头角。去年,它成功获得了由 8VC 和 Bloomberg Beta 领投的种子轮融资。随后,在经历了一段快速增长期后,该公司上个月又完成了由 Andreessen Horowitz 领投的 4000 万美元 A 轮融资。

Rayan Krishnan, the company’s 25-year-old co-founder, previously interned at Palantir, and, as an undergraduate at Stanford, worked for Microsoft and the school’s much lauded artificial intelligence lab. Krishnan says Vals was born from his own observations about how benchmarking was falling behind the advances of the industry it was designed to measure. 该公司 25 岁的联合创始人 Rayan Krishnan 曾在 Palantir 实习,并在斯坦福大学读本科期间为微软及该校备受赞誉的人工智能实验室工作过。Krishnan 表示,Vals 的诞生源于他个人的观察:基准测试的发展已滞后于它所要衡量的行业进步。

“We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance,” Krishnan shares. With AI being integrated into every part of society, benchmarks should really exist to verify that models can do what companies advertise they can do, Krishnan said. “我们看到大量功能强大的新模型迅速推向市场,而学术界的基准测试却跟不上这种前沿进展,”Krishnan 分享道。他认为,随着人工智能融入社会的方方面面,基准测试的真正存在意义应当是验证模型是否具备公司所宣传的能力。

Last week, the young founder showed me around his company’s two-floor office on San Francisco’s Folsom Street — an old brick building that, a century ago, served as the site of a large brewery. Instead of an industrial output of beer, the historical structure is now home to a number of different startups looking to ship the future of the tech industry. 上周,这位年轻的创始人带我参观了他公司位于旧金山 Folsom 街的两层办公室——这是一栋有着百年历史的红砖建筑,曾是一家大型啤酒厂的旧址。如今,这座历史建筑不再生产啤酒,而是成为了多家初创公司的聚集地,它们正致力于塑造科技行业的未来。

“Historically, I think evaluation has been done to evaluate intelligence in a very abstract way,” Krishnan tells me. “Like, do models know enough information to be able to take a bar exam type test?” Here, Vals seeks to differentiate itself. “从历史上看,我认为评估工作一直是以一种非常抽象的方式来衡量智能,”Krishnan 告诉我,“比如,模型掌握的信息量是否足以通过律师资格考试?”在这一点上,Vals 寻求与众不同。

While many benchmarking systems offer tests that are publicly available (this can allow a company to train its model against those tests, thus arguably cheating on their exam), Vals doesn’t publicly disclose its specific test materials. Instead of measuring an AI model’s general knowledge, Vals also evaluates models on their ability to complete complex tasks associated with specific industries like law, finance, and coding. 虽然许多基准测试系统提供公开的测试题(这可能导致公司针对这些测试训练模型,从而在考试中“作弊”),但 Vals 并不会公开其具体的测试材料。Vals 不仅衡量 AI 模型的通用知识,还评估模型在法律、金融和编程等特定行业完成复杂任务的能力。

“What we’re doing is actually looking at what are the real impacts of the models,” said Krishnan. “Can they do work that produces a product of the same quality as a human within every domain?” The idea is to check not just for positive outcomes but also for negative ones, he says. The hope is to analyze how, “if these models ran wild in the world, what the negative implications would be.” “我们实际上是在观察这些模型的真实影响,”Krishnan 说,“它们能否在每个领域产出与人类同等质量的工作成果?”他表示,其核心理念不仅是检查积极成果,还要检查消极后果。他们的目标是分析“如果这些模型在现实世界中失控,会产生什么样的负面影响”。

The capabilities that Vals is measuring are growing. In additional to more traditional industries, the startup continues to push into more unique terrain. “We have a benchmark on recursive self improvement. We’re doing some work in mental health, cybersecurity, biosecurity, and even law of armed conflict to models to understand how to apply the Geneva Convention,” Krishnan shares. Vals 正在衡量的能力范围也在不断扩大。除了更传统的行业外,这家初创公司还继续向更独特的领域进军。Krishnan 分享道:“我们有关于递归自我改进的基准测试。我们还在心理健康、网络安全、生物安全,甚至武装冲突法领域开展工作,以让模型理解如何应用《日内瓦公约》。”

Companies pay Vals to test their models, which can be an odd concept to wrap your head around. Why would a company pay to learn its model isn’t performing well? But having an effective measurement helps companies troubleshoot and improve over time. Krishnan compares their revenue model to how a student might pay the College Board to take the SAT. 企业付费让 Vals 测试其模型,这听起来可能有些难以理解。为什么公司会花钱去了解自己的模型表现不佳呢?但拥有有效的衡量标准有助于公司进行故障排查并实现持续改进。Krishnan 将他们的盈利模式比作学生付费给美国大学理事会(College Board)参加 SAT 考试。

In turn, these evaluations are becoming key decision-making factors for companies looking to acquire new AI models. The startup recently revealed that its revenue is currently eight times what it was last year. Its staff is also growing. Vals, which started the year with only eight people, has already tripled to a team of 25. 反过来,这些评估正成为企业在采购新 AI 模型时的关键决策因素。这家初创公司最近透露,其目前的收入是去年的八倍。员工人数也在增长:Vals 年初时仅有 8 名员工,如今已扩大到 25 人。

Krishnan said that as the startup grows, the plan is to relocate to a significantly bigger office, as well as to bring on an additional 10 to 15 people. The company also recently launched a program centered around providing model evaluations to federal agencies. Krishnan 表示,随着公司的发展,他们计划搬迁到更大的办公室,并增聘 10 到 15 名员工。该公司最近还启动了一项计划,专门为联邦机构提供模型评估服务。

Krishnan sees his company’s system of benchmarking as the future of how AI companies think about growing their businesses and establishing public trust. “AI companies are starting to go public. SpaceX went public. Anthropic is slated for later this year. I suspect OpenAI will be public soon. I think as AI models become a core part of the economy and are diffused more broadly, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings or talk about the prospective investments they’re going to make in AI,” he said. Krishnan 将其公司的基准测试系统视为 AI 公司思考业务增长和建立公众信任的未来。“AI 公司正开始上市。SpaceX 已经上市,Anthropic 计划在今年晚些时候上市。我怀疑 OpenAI 也很快会上市。我认为,随着 AI 模型成为经济的核心部分并得到更广泛的普及,我们所做的基准测试和评估将推动其应用,并成为这些公司提交公开文件或讨论未来 AI 投资计划的核心部分,”他说。