A new kind of AI model from a ChatGPT inventor is thrilling developers

A new kind of AI model from a ChatGPT inventor is thrilling developers

ChatGPT 发明者推出的一种新型 AI 模型令开发者感到振奋

ChatGPT broke Diogo Almeida’s heart. Almeida was an OpenAI researcher who helped build the chatbot and then invent reinforcement learning from human feedback (RLHF), the model-training technique perhaps most responsible for our current age of AI. But despite its capabilities, he was disappointed. “We have lightning in a bottle, and yet it is not useful,” Almeida told TechCrunch. “I’ve been battling that problem since then. It took me a while to come to the conclusion: The problem is we are optimizing for human language … We have been super good at human language for four years, but it’s not useful for automation because computers speak a different language.”

ChatGPT 让 Diogo Almeida 感到心碎。Almeida 曾是 OpenAI 的研究员,他不仅参与构建了这款聊天机器人,还共同发明了“人类反馈强化学习”(RLHF)——这项模型训练技术或许是促成当前 AI 时代的最主要功臣。然而,尽管 ChatGPT 能力出众,他却感到失望。“我们虽然掌握了核心技术,但它却并不实用,”Almeida 在接受 TechCrunch 采访时表示,“从那时起,我就一直在与这个问题作斗争。我花了一段时间才得出结论:问题在于我们一直在针对人类语言进行优化……过去四年里,我们在处理人类语言方面表现得非常出色,但这对于自动化来说并无用处,因为计算机使用的是另一种语言。”

Two years ago, Almeida left OpenAI to start TypeSafe AI, a startup trying to fix that problem. This week, the company released a new transformer-based model, Jev, that is not a large language model (LLM). It doesn’t output text, but instead produces probabilities, or what the company calls “calibrated decisions.” Eschewing language does a few things: It makes the model incredibly cheap and fast, and because users define the outputs in advance, it cannot hallucinate. Its output tokens are free, and input tokens are metered by the billion, not the million.

两年前,Almeida 离开 OpenAI 创办了 TypeSafe AI,这家初创公司旨在解决上述问题。本周,该公司发布了一款基于 Transformer 架构的新模型 Jev,它并非大语言模型(LLM)。它不输出文本,而是输出概率,即该公司所称的“校准决策”(calibrated decisions)。摒弃语言带来了几个优势:这使得模型极其廉价且快速;由于输出结果由用户预先定义,它不会产生幻觉。其输出 Token 是免费的,而输入 Token 的计费单位是十亿级,而非百万级。

Developers are taking a great interest in the product; the company briefly lost the ability to serve users from its API because demand was so high. Jev appears most useful for software automation. Thus far, software developers see it as a cheaper and more robust way to incorporate intelligence into their code. For example, Pranit Sharma, a software engineer at Vercel, a company making agentic infrastructure, said his company had used OpenAI’s ChatGPT Luna 5.6 to run a classifier to run a classifier to review commands for safety. When Vercel replaced OpenAI’s Luna with Jev, it got results five to 18 times more quickly and with greater accuracy.

开发者们对该产品表现出了浓厚兴趣;由于需求过高,该公司一度无法通过 API 为用户提供服务。Jev 在软件自动化方面似乎最为实用。到目前为止,软件开发者将其视为一种更廉价、更稳健的将智能融入代码的方式。例如,构建智能体基础设施的公司 Vercel 的软件工程师 Pranit Sharma 表示,他的公司曾使用 OpenAI 的 ChatGPT Luna 5.6 运行分类器来审查命令的安全性。当 Vercel 用 Jev 替换掉 OpenAI 的 Luna 后,其处理速度提升了 5 到 18 倍,且准确率更高。

Another developer, Bryo AI CTO Nikhil Mudholkar, tested Jev against Gemini for classifying business emails. In his test, Gemini was slightly more accurate, but 10 to 20 times more expensive. More interesting to Mudholkar were Jev’s confidence scores — “it is the only one that hands back a real probability which makes it ideal for automating workflows!!”

另一位开发者,Bryo AI 的首席技术官 Nikhil Mudholkar 在分类商务邮件的任务中对比了 Jev 和 Gemini。测试显示,Gemini 的准确率略高,但成本却是 Jev 的 10 到 20 倍。让 Mudholkar 更感兴趣的是 Jev 的置信度评分——“它是唯一能返回真实概率的模型,这使其成为自动化工作流的理想选择!!”

Besides replacing LLMs in certain use cases, the new model can also augment them, acting as a smart check on misbehavior. Using agents to monitor agents can quickly become expensive, but using Jev to do so, Almeida argues, makes sense. He sees users deploying Jev to track LLM agent traces and prevent jailbreaks. “At the end of the day, it delegates the hallucination problem a little bit to the user,” explained Armin Ronacher, the CTO of Earendil, which builds the open source model harness Pi. “The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it’s 95%, sure, then I can do something with it.”

除了在特定场景中替代 LLM 外,该新模型还可以增强 LLM,充当防止错误行为的智能检查机制。使用智能体来监控智能体很快会变得昂贵,但 Almeida 认为使用 Jev 来实现这一点非常合理。他预见用户会部署 Jev 来追踪 LLM 智能体的轨迹并防止越狱攻击。“归根结底,它将幻觉问题在一定程度上交给了用户处理,”构建开源模型工具 Pi 的 Earendil 公司首席技术官 Armin Ronacher 解释道,“用户必须判断,如果返回的概率只有 50%,那可能就像抛硬币一样,我会忽略它。但如果是 95%,那当然可以据此采取行动。”

Another potential use for Jev is model routing, Ronacher said. Predicting whether a given workload requires a specific model would be useful, but using an LLM for the job would be expensive. Jev’s low cost and speed make that kind of real-time sorting possible. And that’s Almeida’s hope. The model is named after William Stanley Jevons, the 19th-century economist whose eponymous paradox describes how the falling cost of a commodity can lead to it being used more and more. In this case, the falling cost of intelligence should lead to its widespread deployment.

Ronacher 表示,Jev 的另一个潜在用途是模型路由。预测给定的工作负载是否需要特定模型非常有用,但使用 LLM 来完成这项工作成本太高。Jev 的低成本和高速度使得这种实时分类成为可能。这正是 Almeida 的愿景。该模型以 19 世纪经济学家威廉·斯坦利·杰文斯(William Stanley Jevons)的名字命名,他提出的“杰文斯悖论”描述了某种商品的成本下降如何导致其使用量反而增加。在这种情况下,智能成本的下降应该会促使其得到广泛部署。

“We think that there’s just going to be smart software all over the place in a way that’s emergent and distributed … much more like the early internet than you know like the mega apps that people are trying to build right now,” Almeida said.

“我们认为,智能软件将以一种涌现且分布式的形式遍布各处……这更像早期的互联网,而不是人们现在试图构建的那种巨型应用,”Almeida 说道。

Almeida is tight-lipped about the model’s architecture, which outside observers suspect is built on top of an open-weight LLM. The company refers to Jev as a “System One model,” focused on intuition rather than reasoning, and specifically focused on the right task. Almeida says Jev is trained exclusively on synthetic data using a technique he calls “reinforcement learning from calibrated decisions.”

Almeida 对该模型的架构守口如瓶,外界观察者怀疑它是基于开源权重 LLM 构建的。该公司将 Jev 称为“系统一模型”(System One model),专注于直觉而非推理,并专门针对特定任务进行优化。Almeida 表示,Jev 完全使用合成数据进行训练,采用了一种他称之为“基于校准决策的强化学习”的技术。

“We made an early bet that we will be making all of our data, and that has been one of the best bets I’ve ever made in my life — better than our launch, in my opinion, better than RLHF,” he told TechCrunch. “Half of [our company] is a lab that basically owns this entire subfield of statistically well-understood synthetic data, and that is now my life joy.”

“我们很早就押注于我们将自行生成所有数据,这是我一生中做出的最好的赌注之一——在我看来,这比我们的产品发布更好,甚至比 RLHF 更好,”他告诉 TechCrunch。“我们公司有一半是实验室,基本上掌握了统计学上定义明确的合成数据这一整个细分领域,这现在是我生活的乐趣所在。”

For now, Jev stands alone as this kind of model, but Ronacher expects that competitors will spring up now that its utility is apparent. “We should have seen this earlier in many ways, but presumably because the LLMs are so cheap and subsidized, you often don’t have to be creative yet,” he said. TypeSafe itself will be building more versions of the model, in new modalities. Asked if TypeSafe is a frontier lab, Almeida said, “the main product of frontier labs is fear or hype. I would like our main product to be intelligence…[but we are] not a lab in the sense of, you know, like bet on infinite wealth, or a religion, or building God in a data center, or whatever is the thing of today.”

目前,Jev 作为此类模型独树一帜,但 Ronacher 预计,随着其效用日益显现,竞争对手将会涌现。“从很多方面来看,我们本应更早意识到这一点,但大概是因为 LLM 太便宜且有补贴,你往往还不需要发挥创造力,”他说。TypeSafe 本身将构建该模型的更多版本,并支持新的模态。当被问及 TypeSafe 是否属于前沿实验室时,Almeida 表示:“前沿实验室的主要产品是恐惧或炒作。我希望我们的主要产品是智能……但我们不是那种意义上的实验室,你知道的,比如押注于无限财富、某种宗教,或者在数据中心里构建上帝,或者诸如此类当下的流行事物。”