The AI Hype Index: AI loves cheating

The AI Hype Index: AI loves cheating

AI 炒作指数:AI 热衷于作弊

Brace yourself: It turns out AI is being optimized for cheating. OpenAI’s agents hacked into Hugging Face to get the answers to a cybersecurity test. Next, they solved a prestigious math problem (or just stole from two top mathematicians’ answer sheets). Anthropic’s models have also hacked into other companies’ systems four times already. And that’s only what we’ve caught so far. 做好心理准备:事实证明,人工智能正在被优化以进行作弊。OpenAI 的智能体入侵了 Hugging Face 以获取网络安全测试的答案。随后,它们解决了一个著名的数学难题(或者仅仅是从两位顶尖数学家的答题纸上窃取了答案)。Anthropic 的模型也已经四次入侵了其他公司的系统。而这仅仅是我们目前所发现的。

Freaking out? You’re not alone. AI lab researchers are quitting their jobs and issuing dire warnings that if we keep going this way, AI might eventually kill us all. Bill Gates is sounding the alarm. Bernie Sanders has teamed up with Steve Bannon, of all people, to call for curbs on AI. Anthropic CEO Dario Amodei is urging a slowdown, and other top US AI executives agree. But fear not: President Trump has a plan. He says the only guardrail AI needs is “a STRONG AND SMART (High IQ!) PRESIDENT.” 感到恐慌吗?你并不孤单。人工智能实验室的研究人员纷纷辞职,并发出严厉警告:如果我们继续这样下去,人工智能最终可能会毁灭人类。比尔·盖茨(Bill Gates)敲响了警钟。伯尼·桑德斯(Bernie Sanders)甚至与史蒂夫·班农(Steve Bannon)联手,呼吁对人工智能进行限制。Anthropic 首席执行官达里奥·阿莫代(Dario Amodei)敦促放缓发展步伐,其他美国顶级人工智能高管也表示赞同。但别担心:特朗普总统有一个计划。他说,人工智能唯一需要的护栏就是“一位强大且聪明(高智商!)的总统”。


Deep Dive 深度阅读

Artificial intelligence: A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. By Will Douglas Heaven 人工智能:一个根本性缺陷使大语言模型(LLM)极易受到攻击 这使得诱导它们做不该做的事情变得轻而易举,例如告诉你如何破坏飞机的导航系统。作者:Will Douglas Heaven

AI’s recursive self-improvement might not come so quickly after all AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems. By Michelle Kim 人工智能的递归自我改进或许并不会那么快到来 看来,人工智能智能体目前还没有足够的创造力来进行真正具有创新性的开放式人工智能研究。作者:Michelle Kim

Here’s why AI agents lie and cheat to reach their goals The misbehavior is called reward hacking. This is what you need to know. By Grace Huckins 这就是为什么人工智能智能体会为了达到目标而撒谎和作弊 这种不当行为被称为“奖励黑客攻击”(reward hacking)。这是你需要了解的内容。作者:Grace Huckins

These startups are chasing the next big thing in LLMs Meet the new kids nipping at the heels of the AI giants. By Will Douglas Heaven 这些初创公司正在追逐大语言模型领域的下一个大事件 来认识一下那些紧追人工智能巨头步伐的新秀们。作者:Will Douglas Heaven