Popular AI leaderboard Arena nearly doubles valuation to $3.1B valuation in 10 months
Arena, which originated in 2023 as a research project at UC Berkeley that crowdsourced rankings of AI models, has raised a $200 million Series B round at a $3.1 billion valuation, it said on Thursday. This comes after the company said it reached $100 million in annualized run-rate revenue in June.
Arena 起源于 2023 年加州大学伯克利分校的一个研究项目,旨在通过众包方式对 AI 模型进行排名。该公司周四表示,已完成 2 亿美元的 B 轮融资,估值达到 31 亿美元。此前,该公司曾表示其年化收入在 6 月份已达到 1 亿美元。
The round was led by Lightspeed Venture Partners and Khosla Ventures, with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis, and others joining in. Arena previously announced a $150 million Series A in January at a $1.7 billion post-money valuation. At the time, its annualized revenue was $30 million, it said. So that means its valuation has nearly doubled in about 10 months.
本轮融资由 Lightspeed Venture Partners 和 Khosla Ventures 领投,Salesforce Ventures、01 Advisors、Dell Technologies Capital、Endeavor Catalyst、a16z、Felicis 等机构跟投。Arena 此前在 1 月份宣布了 1.5 亿美元的 A 轮融资,投后估值为 17 亿美元。当时该公司称其年化收入为 3000 万美元。这意味着其估值在约 10 个月内几乎翻了一番。
Arena provides a crowdsourced platform that is free for consumers to use. People enter prompts or request vibe-coded projects and then rate which model does it better. Arena claims it has tens of millions of monthly visitors. In September of last year, it introduced its commercial product, AI Evaluations, a service that provides model labs and enterprises with detailed performance analytics based on its community feedback.
Arena 提供了一个免费供消费者使用的众包平台。用户输入提示词或请求“氛围编码”(vibe-coded)项目,然后对哪个模型表现更好进行评分。Arena 声称其每月拥有数千万访问者。去年 9 月,该公司推出了商业产品“AI 评估”(AI Evaluations),该服务基于社区反馈为模型实验室和企业提供详细的性能分析。
The timing proved impeccable. This year, AI labs realized that their models were gaming benchmarking tests, finding ways to rack up good scores without truly earning them. At the same time, enterprises wanted help determining which model works best for their own internal needs rather than relying only on standardized benchmarks.
时机恰到好处。今年,AI 实验室意识到他们的模型正在“钻”基准测试的空子,通过各种手段获得高分,而非真正具备相应能力。与此同时,企业希望获得帮助,以确定哪种模型最适合其内部需求,而不是仅仅依赖标准化的基准测试。
“AI is advancing faster than our ability to evaluate it, and static benchmarks break down once models recognize they’re being tested,” the company said in its funding announcement. “The world needs a neutral third party to measure how safe and aligned AI actually is once it’s in the hands of real people. Arena is stepping into that role today,” it added.
“AI 的发展速度超过了我们评估它的能力,一旦模型识别出它们正在接受测试,静态基准测试就会失效,”该公司在融资公告中表示。“世界需要一个中立的第三方来衡量 AI 在真实用户手中时到底有多安全和一致。Arena 今天正在承担这一角色,”它补充道。
To that end, Arena has also added a new category to its leaderboard: alignment. This is where it ranks models based on issues like unauthorized action (taking actions it wasn’t asked to take); false attribution (wrongly crediting statements or facts to the wrong source); and what it calls “deceptive completion” (lying about completing tasks that it didn’t do).
为此,Arena 还在其排行榜中增加了一个新类别:对齐(alignment)。在该类别中,它根据诸如未经授权的操作(执行未被要求的动作)、错误归因(将陈述或事实错误地归于错误的来源)以及所谓的“欺骗性完成”(谎称完成了未实际执行的任务)等问题对模型进行排名。
Currently, a slate of OpenAI’s models are at the top of its preliminary alignment leaderboard, with Claude Opus 5.5 and Claude Fable in sixth and ninth place, respectively.
目前,OpenAI 的一系列模型位居其初步对齐排行榜榜首,Claude Opus 5.5 和 Claude Fable 分别排在第六和第九位。