How AI decision models could change content moderation
How AI decision models could change content moderation
AI 决策模型将如何改变内容审核
As decision models spread across the industry, a company called Musubi has a new idea for how to put them to work: moderating content. On Tuesday, Musubi announced a lightweight decision model made for real-time moderation called PolicyLM-1.7B, released with open weights. 随着决策模型在整个行业中普及,一家名为 Musubi 的公司提出了一种将其投入应用的新思路:内容审核。周二,Musubi 发布了一款专为实时审核而设计的轻量级决策模型 PolicyLM-1.7B,并公开了其权重。
The idea is to take a content policy written in plain English and apply it to messages in under 50 milliseconds. Musubi’s model is designed to be similar in cost and speed to the AI classifier systems that power moderation on most social platforms — but because it has the flexibility of a modern LLM, it can apply complex policies without special training. 其核心理念是:将用通俗英语编写的内容政策,在 50 毫秒内应用到信息审核中。Musubi 的模型在成本和速度上与目前大多数社交平台所采用的 AI 分类系统相当,但由于它具备现代大语言模型(LLM)的灵活性,因此无需特殊训练即可执行复杂的政策。
Even more important, the model won’t need new training when the policy changes, allowing for human policy-setters to iterate as much as they need. As Musubi co-founder and chief AI officer Filip Jankovic sees it, it gives platform managers a way to label content proactively. “Product teams just want a better understanding of what’s happening on their platform, especially as the amount of content is exponentially increasing,” Jankovic says. “Being able to label all of that in a very scalable, customizable way is extremely useful.” 更重要的是,当政策发生变化时,该模型无需重新训练,这使得人工政策制定者可以根据需要进行多次迭代。在 Musubi 联合创始人兼首席 AI 官 Filip Jankovic 看来,这为平台管理者提供了一种主动标记内容的方法。“产品团队只是想更好地了解平台上正在发生的事情,尤其是在内容量呈指数级增长的情况下,”Jankovic 说,“能够以一种高度可扩展、可定制的方式对所有内容进行标记,是非常有用的。”
Decision models have become a hot topic in the AI world since the release of TypeSafe AI’s Jev in September, which was shortly followed by competing decision models from OpenAI and Amazon. Instead of outputting text, a decision model outputs outcome probabilities, though in this case the model outputs a binary judgement: Either the content is in the category or it isn’t. 自 9 月份 TypeSafe AI 发布 Jev 以来,决策模型已成为 AI 领域的热门话题,随后 OpenAI 和亚马逊也相继推出了竞争性的决策模型。决策模型输出的不是文本,而是结果概率;不过在本例中,模型输出的是二元判断:内容要么属于该类别,要么不属于。
By limiting the model’s output to a set of predetermined choices, decision models are able to run faster and cheaper than large language models, while still maintaining the flexibility of the transformer architecture. One early use case is reining in misbehavior by AI agents — so it’s only natural to apply the same technology to human misbehavior. 通过将模型的输出限制为一组预先确定的选项,决策模型能够比大语言模型运行得更快、成本更低,同时仍保持 Transformer 架构的灵活性。一个早期的应用案例是遏制 AI 智能体的违规行为——因此,将同样的技术应用于人类的违规行为也就顺理成章了。
Notably, Jankovic says his interest in decision models predates Jev, tracing it back to a 2024 project called GLiNER (Generalist Model for Named Entity Recognition) that deployed many of the same techniques. Still, Musubi isn’t wary of the comparison. If anything, the company is eager to use the new interest in decision models to shine a light on content moderation. “If Jev caught your eye, PolicyLM-1.7B is the same kind of model, trained specifically for content moderation, that you can run yourself,” the product announcement reads. 值得注意的是,Jankovic 表示他对决策模型的兴趣早于 Jev,可以追溯到 2024 年一个名为 GLiNER(命名实体识别通用模型)的项目,该项目部署了许多相同的技术。尽管如此,Musubi 并不排斥这种比较。相反,该公司渴望利用人们对决策模型的新兴趣,让内容审核受到更多关注。产品公告中写道:“如果 Jev 引起了你的注意,那么 PolicyLM-1.7B 就是同类型的模型,它是专门为内容审核而训练的,而且你可以自行运行。”