Design Arena creators raise $7.9 million to bring taste to AI models
Design Arena creators raise $7.9 million to bring taste to AI models
Design Arena 创始团队融资 790 万美元,旨在为 AI 模型注入“审美”
As co-founder Grace Li tells it, her company started a few weeks before graduation in 2025, with a handful of college friends trying to make their AI game engine work. The models could make functional games, but none of the games were fun — which raised the interesting question, how can you tell if a game will be fun? 据联合创始人 Grace Li 讲述,她的公司成立于 2025 年毕业前几周,当时她和几位大学好友正试图开发一款 AI 游戏引擎。这些模型虽然能制作出功能完备的游戏,但没有一款游戏是有趣的——这引发了一个有趣的问题:你如何判断一款游戏是否好玩?
There was no substitute for human judgment, they decided, and soon they were brainstorming ways to get honest human feedback at scale. The result became Design Arena, an AI tool now used by 5.3 million people around the world. As it turned out, there were lots of AI companies looking for scalable user feedback — and many of them were willing to pay for it. 他们认为,人类的判断力是无可替代的,于是很快开始集思广益,寻找大规模获取真实人类反馈的方法。最终,Design Arena 应运而生,这款 AI 工具目前已被全球 530 万人使用。事实证明,许多 AI 公司都在寻找可扩展的用户反馈,而且其中许多公司愿意为此付费。
“It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li says. “About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history.” On Monday, the company behind Design Arena — dubbed Intelligence — announced a $7.9 million seed round led by Index Ventures with participation from Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others. “对于许多模型来说,这是在设计领域实现改进所缺失的瓶颈,”Li 说道,“大约一周后,我们与一家前沿实验室达成了第一笔重大交易,剩下的就是历史了。”周一,Design Arena 背后的公司 Intelligence 宣布完成 790 万美元种子轮融资,由 Index Ventures 领投,Conviction(Sarah Guo 和 Mike Vernal)、A*、Valkyrie 等参投。
For non-enterprise users, using Design Arena is a lot like using a sophisticated model router. There’s a ChatGPT-style window for prompts, with separate dropdowns for websites, images, and a dozen other visual formats. Once you put in the request, format, and style, you’ll be presented with a series of “A vs. B” choices until you’ve ranked the handful of outputs from best to worst. 对于非企业用户来说,使用 Design Arena 就像在使用一个复杂的模型路由工具。界面有一个类似 ChatGPT 的提示词输入窗口,并为网站、图像和其他十几种视觉格式提供了单独的下拉菜单。一旦输入请求、格式和风格,系统就会向你展示一系列“A 对 B”的选择,直到你将这几项输出结果按从好到坏进行排序。
It’s a useful service, but the real value of the platform comes from the enterprise side, where participating models can treat it as a source of endless instant feedback for their media-generating models. The users tend to be indifferent to which models they’re ranking — as Li puts it, they just want the best output they can get — so their rankings can give critical input to what users really want. 这是一项实用的服务,但该平台的真正价值在于企业端。参与的模型可以将此视为其媒体生成模型获取源源不断即时反馈的来源。用户往往不在意他们正在评估的是哪个模型——正如 Li 所言,他们只是想要得到最好的输出结果——因此,他们的排名可以为“用户真正想要什么”提供关键的参考依据。
For frontier labs, that’s a service worth paying for, Li says, adding the site is currently generating $60 million in ARR, solidifying its position as a key source of human-led evaluation data for the AI industry. Crucially, users have to log in to get their output, so Intelligence can also track how those tastes change across different continents and over time. (Li notes that web dashboards in Asia tend to have a more maximalist design style.) Li 表示,对于前沿实验室来说,这是一项值得付费的服务。她补充说,该网站目前的年度经常性收入(ARR)已达 6000 万美元,巩固了其作为 AI 行业人类主导评估数据关键来源的地位。至关重要的是,用户必须登录才能获取输出结果,因此 Intelligence 还可以追踪这些审美偏好在不同大洲和时间跨度上的变化。(Li 指出,亚洲的网页仪表盘往往具有更偏向“极繁主义”的设计风格。)
These measures are an important complement to automated benchmarks, which can operate at a greater scale but are often subject to being gamed or otherwise manipulated, as the Hugging Face breach demonstrated in dramatic fashion last week. That’s not to say that crowdsourced human feedback will be an automatic winning market. 这些措施是对自动化基准测试的重要补充。虽然自动化测试规模更大,但往往容易被“刷榜”或以其他方式操纵,正如上周 Hugging Face 遭遇的入侵事件所戏剧性地展示的那样。但这并不意味着众包的人类反馈一定会成为一个稳赢的市场。
Less than a year after launching, Yupp shuttered its doors earlier this year after raising $33 million from a16z crypto’s Chris Dixon. It too nabbed some frontier models as customers and had, it said, over 1.3 million users, but still couldn’t build a sustainable long-term business. Even so, other startups based on human evaluation seem to be thriving. LM Arena, which takes a similar approach to text-based responses, raised $150 million in a Series A in January, just four months after formally launching its paid product. 在推出不到一年后,Yupp 在从 a16z crypto 的 Chris Dixon 那里筹集了 3300 万美元后,于今年早些时候倒闭了。它也曾争取到一些前沿模型作为客户,并声称拥有超过 130 万用户,但最终仍未能建立起可持续的长期业务。即便如此,其他基于人类评估的初创公司似乎正在蓬勃发展。LM Arena 采取了类似的文本响应评估方法,在正式推出付费产品仅四个月后的 1 月份,就完成了 1.5 亿美元的 A 轮融资。