Felony Bench
Felony Bench
Felony Bench: A benchmark you really don’t want models to be saturated with. Felony Bench:一个你绝对不希望模型被其“饱和”的基准测试。
Learn more ↓ ModelEvaluator Score ↖ Most illegalLeast illegal ↘ 8 Anthropic 8 OpenAI 1 Meta 0 Google 0 Moonshot 了解更多 ↓ 模型评估得分 ↖ 最非法 最合法 ↘ 8 Anthropic 8 OpenAI 1 Meta 0 Google 0 Moonshot
Scores indicate count of illegal activity. Higher is… you decide. 得分代表非法活动的计数。分数越高意味着……由你来定。
| Company | Felonies | Description | Date | Source |
|---|---|---|---|---|
| 公司 | 重罪数 | 描述 | 日期 | 来源 |
| Anthropic | 1 | Exploited auth failures in an API to cancel other people’s gym classes | 8/9/2026 | ABC Australia ↗ |
| Anthropic | 1 | 利用 API 中的身份验证漏洞取消他人的健身课程 | 8/9/2026 | ABC Australia ↗ |
| Meta | 1 | Compromise of an internal account at one company | 8/5/2026 | The Information ↗ |
| Meta | 1 | 入侵某公司内部账户 | 8/5/2026 | The Information ↗ |
| Anthropic | 4 | Unauthorized use of GitHub credentials; Dependabot supply-chain attack; social engineering email campaign; public exposure of a malicious DNS server | 8/4/2026 | AISI ↗ |
| Anthropic | 4 | 未经授权使用 GitHub 凭据;Dependabot 供应链攻击;社会工程邮件活动;公开暴露恶意 DNS 服务器 | 8/4/2026 | AISI ↗ |
| OpenAI | 2 | Unauthorized use of GitHub credentials; public exposure of a malicious DNS server | 8/4/2026 | OpenAI ↗ |
| OpenAI | 2 | 未经授权使用 GitHub 凭据;公开暴露恶意 DNS 服务器 | 8/4/2026 | OpenAI ↗ |
| OpenAI | 1 | Compromise of an internal account from a misconfigured CTF evaluation | 8/4/2026 | OpenAI ↗ |
| OpenAI | 1 | 因 CTF 评估配置错误导致内部账户被入侵 | 8/4/2026 | OpenAI ↗ |
| OpenAI | 4 | Compromise of internal accounts at four companies as part of the Hugging Face incident | 7/31/2026 | OpenAI ↗ Reuters ↗ |
| OpenAI | 4 | 作为 Hugging Face 事件的一部分,入侵了四家公司的内部账户 | 7/31/2026 | OpenAI ↗ Reuters ↗ |
| Anthropic | 3 | Compromise of internal accounts at three companies | 7/30/2026 | Anthropic ↗ |
| Anthropic | 3 | 入侵了三家公司的内部账户 | 7/30/2026 | Anthropic ↗ |
| OpenAI | 1 | Compromise of Hugging Face during a model evaluation | 7/21/2026 | OpenAI ↗ |
| OpenAI | 1 | 在模型评估期间入侵 Hugging Face | 7/21/2026 | OpenAI ↗ |
Methodology 方法论
Felony Bench counts unique instances where AI agents affect third-party entities. Escaping a sandbox alone does not constitute a counted incident. It is for these reasons that Frontier Security’s Kimi K3 incident and Alibaba’s ROME incident are not counted. Felony Bench 统计的是 AI 智能体影响第三方实体的独立事件。仅逃离沙箱并不构成统计事件。正因如此,Frontier Security 的 Kimi K3 事件和阿里巴巴的 ROME 事件未被计入。