Clef: Open-weight decision models, and new RL fine-tuning platform

Clef: Open-weight decision models, and new RL fine-tuning platform

Over the last few weeks, there has been lots of buzz around decision models such as Typesafe AI’s Jev System One model. While classifier models have been around for some time, Jev introduces a new decision model concept into the world of AI — a model that produces bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required. 在过去的几周里,关于决策模型(如 Typesafe AI 的 Jev System One 模型)的讨论非常热烈。虽然分类模型已经存在了一段时间,但 Jev 向 AI 世界引入了一个新的决策模型概念——一种能够以低成本、快速且一致的方式产生有界结构化输出的模型,当需要决策时,它可以被添加到工作流中。

These models are capable enough to work over any set of inputs without constantly retraining the model to incorporate new classification categories. This contrasts with the world of Large Language Models (LLMs), which are largely non-deterministic, but are open-ended enough to reason and generate text and tool calls for agentic workloads. 这些模型功能强大,无需不断重新训练即可处理任何输入集,从而纳入新的分类类别。这与大型语言模型(LLM)的世界形成了鲜明对比,后者在很大程度上是不确定的,但足够开放,可以为智能体工作负载进行推理、生成文本和调用工具。

Today, we’re releasing two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI. Clef is currently the leader when evaluated against the Jev Decision Index, you can view full results on the live benchmark demo site. These models are smarter, faster, and fully Jev-API compatible, so you can experiment with these hosted models easily. We’re fully open-sourcing these models on Hugging Face under an Apache 2.0 license for you to run locally and experiment with yourselves. Lastly, we’re excited to debut our new reinforcement learning (RL) product, which allows customers to fine-tune Clef to suit their use cases as well. 今天,我们发布了两个由 Cloudflare 训练的决策模型:Clef 和 Clef-flash,它们托管在 Workers AI 上。在 Jev 决策指数(Jev Decision Index)的评估中,Clef 目前处于领先地位,您可以在实时基准测试演示网站上查看完整结果。这些模型更智能、更快速,并且完全兼容 Jev API,因此您可以轻松地对这些托管模型进行实验。我们已在 Hugging Face 上以 Apache 2.0 许可证完全开源了这些模型,供您在本地运行并自行尝试。最后,我们很高兴推出新的强化学习(RL)产品,它允许客户微调 Clef 以适应他们自己的用例。

What is a decision model?

什么是决策模型?

A decision model makes classifications to help agents decide how to act, based on certain probabilities. For example, you can pass in a customer support message (inputs) and ask if it is urgent and which team should handle it. A decision model will return typed answers with probabilities (outputs), which your code can use to route the ticket, trigger an escalation, or defer to a human. This means that a human does not necessarily need to be in the loop for agentic decisions anymore — agents can programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed. 决策模型通过一定的概率进行分类,以帮助智能体决定如何行动。例如,您可以传入一条客户支持消息(输入),并询问它是否紧急以及应该由哪个团队处理。决策模型将返回带有概率的类型化答案(输出),您的代码可以使用这些答案来路由工单、触发升级或转交给人工处理。这意味着在智能体决策过程中,人类不再必须参与其中——智能体可以以编程方式收集上下文、做出决策并执行任务,或者在需要时转交给人类。

Specifically at Cloudflare, we’ve been testing our new Clef model on our Threat Intelligence team to help us classify website domains. By giving a domain to Clef (with Browser Run) it can quickly identify categories that the domain falls under — for example, it might classify a domain with a 95% chance it is a fashion website, 85% ecommerce, <1% phishing, etc. This classification took our Clef model 2.2s to fetch, render, and classify the website. In contrast, our fastest general LLM gpt-oss-120b took 4.7s in the same workflow, and only returned two classifications. As a user, you can imagine how a 2x savings in latency and results can help us improve our threat intelligence workflows and be faster in identifying malicious or legitimate domains. Generalize this to any use case where you need to make quick programmatic decisions, and you unlock powerful agentic workflows that are able to autonomously decide, reason, and execute. 具体在 Cloudflare,我们一直在威胁情报团队测试新的 Clef 模型,以帮助我们对网站域名进行分类。通过将域名交给 Clef(配合 Browser Run),它可以快速识别该域名所属的类别——例如,它可能将某个域名分类为 95% 的概率是时尚网站,85% 是电子商务,<1% 是钓鱼网站等。我们的 Clef 模型完成该网站的获取、渲染和分类仅需 2.2 秒。相比之下,我们最快的通用 LLM gpt-oss-120b 在相同的工作流中耗时 4.7 秒,且仅返回了两个分类结果。作为用户,您可以想象延迟和结果效率提升 2 倍将如何帮助我们改进威胁情报工作流,并更快地识别恶意或合法域名。将此推广到任何需要快速进行程序化决策的用例中,您就能解锁强大的智能体工作流,使其能够自主决策、推理和执行。

In music theory, a clef is a symbol placed at the beginning of a musical staff that assigns specific pitch names to the lines and spaces. A decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it. We chose Clef as the name of our family of decision models, as it serves similar purposes, and the CF hearkens to Cloudflare. 在音乐理论中,谱号(Clef)是放置在五线谱开头的一个符号,用于为线条和间隙指定特定的音高名称。决策模型类似于音乐谱号,因为它有助于定义上下文的领域以及随之而来的后续音符(动作)。我们选择 Clef 作为我们决策模型系列的名称,因为它具有相似的用途,且“CF”也呼应了 Cloudflare。

How is Clef different from other decision models?

Clef 与其他决策模型有何不同?

Although the market is getting increasingly saturated with decision models, Clef has some unique properties that make us excited to release it to the public. First, it has a vision encoder so it’s able to take in images and classify visual content. This is different from Jev, which only does text classification today. Secondly, our model has a 64k context window (compared to Jev’s 32k), which allows users to squeeze more input state for the model to classify against. 尽管决策模型市场正变得越来越饱和,但 Clef 拥有一些独特的属性,这让我们非常兴奋能将其发布给公众。首先,它拥有视觉编码器,因此能够接收图像并对视觉内容进行分类。这与目前仅能进行文本分类的 Jev 不同。其次,我们的模型拥有 64k 的上下文窗口(相比之下 Jev 为 32k),这允许用户压缩更多的输入状态供模型进行分类。

Third, our model is accurate and powerful, scoring competitively against other decision models on the market across various quality benchmarks. We shortlisted some evaluations below that are important for decision-making as defined by the Jev Decision Index and scored some of the more popular models on the market for it. 第三,我们的模型准确且强大,在各种质量基准测试中与市场上的其他决策模型相比具有竞争力。我们在下方列出了一些对决策至关重要的评估指标(由 Jev 决策指数定义),并对市场上一些更受欢迎的模型进行了评分。