The Download: reward hacking explained, and suspected Iranian cyberattacks
The Download: reward hacking explained, and suspected Iranian cyberattacks
《下载》:奖励黑客行为解析与疑似伊朗网络攻击
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. 这是今日份的《下载》,我们平日更新的简报,为您提供每日科技世界的动态。
Here’s why AI agents lie and cheat to reach their goals
AI 智能体为何会为了达成目标而撒谎和作弊
When two OpenAI models hacked into Hugging Face last month, they weren’t trying to make money or commit sabotage—they were just looking for answers to a test question. According to OpenAI, the models decided to solve a cybersecurity exercise by hacking out of the environment in which OpenAI had attempted to contain them and into Hugging Face’s databases, where—they reasoned—the correct answer to the problem might be stored. 上个月,当两个 OpenAI 模型入侵 Hugging Face 时,它们并非为了牟利或进行破坏,仅仅是为了寻找一道测试题的答案。据 OpenAI 称,这些模型为了解决一项网络安全练习,决定通过“越狱”——即从 OpenAI 试图限制它们的运行环境中逃逸,并入侵 Hugging Face 的数据库,因为它们推断问题的正确答案可能存储在那里。
The incident has attracted intense attention over the past couple of weeks. It’s a dramatic illustration of just how good AI models have gotten at hacking. But it’s perhaps even more striking as an example of how and why AI systems lie and cheat. Read our story explaining why AI engages in this sort of behavior—known as ”reward hacking.” —Grace Huckins 这一事件在过去几周引起了广泛关注。它生动地展示了 AI 模型在黑客攻击方面已经达到了何种水平。但更令人震惊的是,它成为了 AI 系统如何以及为何会撒谎和作弊的典型案例。阅读我们的报道,了解 AI 为何会进行这种被称为“奖励黑客”(reward hacking)的行为。——Grace Huckins
This story is from our ‘Explains’ series, where our writers untangle the complex, messy world of technology to help you understand what’s coming next. Read more from the collection. 本文来自我们的“解释”系列,我们的作者在此梳理复杂且混乱的科技世界,帮助您理解未来趋势。阅读该系列的更多内容。
The must-reads
必读精选
I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 我已为您搜罗了互联网上今天最有趣、最重要、最令人恐惧或最引人入胜的科技新闻。
-
It looks like Iran is conducting cyberattacks on US water systems That’s according to preliminary investigations on hacks in at least seven states. (NYT $) + Will this be a wake-up call? (Forbes) 伊朗似乎正在对美国供水系统进行网络攻击 根据对至少七个州发生的黑客攻击事件的初步调查显示。(《纽约时报》付费内容)+ 这会是一个警钟吗?(《福布斯》)
-
Google briefly made it easy to fake satellite images Literally the last thing the world needs right now. (NPR) + AI companies keep moving fast and breaking things. (The Atlantic $) + Apple is struggling to keep pace with incoming AI-assisted software bug reports. (FT $) 谷歌曾短暂地让伪造卫星图像变得轻而易举 这绝对是当今世界最不需要的东西。(NPR)+ AI 公司继续“快速行动,打破常规”。(《大西洋月刊》付费内容)+ 苹果公司正疲于应对源源不断的 AI 辅助软件漏洞报告。(《金融时报》付费内容)
-
Why wildfires have got so bad in Europe this summer It’s a mix of climate change, land abandonment, and outdated firefighting tactics. (New Yorker $) + How Europe can become more fire-resilient. (New Scientist $) + How much wildfire prevention is too much? (MIT Technology Review) 为何今年夏天欧洲的野火如此严重 这是气候变化、土地荒废和过时消防战术共同作用的结果。(《纽约客》付费内容)+ 欧洲如何增强防火韧性。(《新科学家》付费内容)+ 野火预防做到什么程度才算过头?(《麻省理工科技评论》)
-
Law enforcement officers are using license-plate cameras for stalking There are at least 50 examples of officers being charged with or accused of misusing them. (WP $) + Inside Chicago’s surveillance panopticon. (MIT Technology Review) 执法人员利用车牌识别摄像头进行跟踪 至少有 50 起案例显示警员因滥用这些设备而被起诉或指控。(《华盛顿邮报》付费内容)+ 走进芝加哥的监控全景监狱。(《麻省理工科技评论》)
-
China may impose more controls on its homegrown AI models They’re winning influence overseas—but create new security and political risks. (NYT $) + Silicon Valley is deeply divided over how to respond. (Rest of World) + China’s AI models have Trump’s AI world at war with itself. (MIT Technology Review) 中国可能对其本土 AI 模型实施更多管控 它们正在海外赢得影响力,但也带来了新的安全和政治风险。(《纽约时报》付费内容)+ 硅谷在如何应对这一问题上存在严重分歧。(Rest of World)+ 中国的 AI 模型让特朗普的 AI 世界陷入内斗。(《麻省理工科技评论》)
-
The vast majority of Australian teens are still on social media A lack of effective age checks means the country’s under-16s ban simply isn’t enforceable. (Reuters $) 绝大多数澳大利亚青少年仍在使用社交媒体 缺乏有效的年龄验证意味着该国针对 16 岁以下人群的禁令根本无法执行。(路透社付费内容)
-
Is it possible to make smart glasses that aren’t creepy? 👓😱 It doesn’t really look like it right now! (Wired $) 有可能制造出不让人感到毛骨悚然的智能眼镜吗?👓😱 目前看来似乎不太可能!(《连线》付费内容)
-
The US ban on robot vacuum cleaners isn’t workable It’s going to leave Americans with less choice and way higher prices. (The Verge $) 美国对扫地机器人的禁令行不通 这将导致美国消费者的选择减少,价格大幅上涨。(The Verge 付费内容)
-
YouTube just banned a bunch of ASMR artists They say they’re being unfairly caught up in rules against “sexually gratifying” content. (404 Media) YouTube 封禁了一批 ASMR 创作者 创作者们表示,他们被不公正地卷入了针对“性满足”内容的规则中。(404 Media)
-
Why Pokémon is still popular all over the world It seems to have a rare ability to both cheer us up, and bring us together. (The Guardian) 为什么宝可梦在世界各地依然受欢迎 它似乎拥有一种罕见的能力,既能让我们振作起来,又能将我们凝聚在一起。(《卫报》)
Quote of the day
今日金句
“Trump knows exactly who is responsible for this attack, and knows that other states were hit too. This is what modern warfare looks like, and it further illustrates there’s no plan to win a war with Iran.” “特朗普非常清楚谁该为这次攻击负责,也知道其他州也遭到了攻击。这就是现代战争的样子,这也进一步说明(政府)根本没有赢得对伊战争的计划。”
—Governor Tim Walz responds to Trump blaming Minnesota for cyberattacks on its own water systems, the Washington Post reports. ——据《华盛顿邮报》报道,州长蒂姆·沃尔兹(Tim Walz)回应了特朗普指责明尼苏达州应对其自身供水系统遭受的网络攻击负责的言论。
One More Thing
还有一件事
Meet the researchers testing the “Armageddon” approach to asteroid defense 认识一下那些正在测试“世界末日”式小行星防御方案的研究人员
One day a big asteroid will find itself on a collision course with Earth. If we are lucky, it’d land in the middle of the vast ocean, creating a good-size but innocuous tsunami, or in an uninhabited patch of desert. But if it has a city in its crosshairs, one of the worst natural disasters in modern times would unfold. Homes dozens of miles away would fold like cardboard. Millions of people would die. 总有一天,一颗巨大的小行星会撞向地球。如果我们幸运的话,它可能会落在广阔的海洋中央,引发一场规模尚可但无害的海啸,或者落在无人居住的沙漠地带。但如果它瞄准的是一座城市,现代史上最严重的自然灾害之一就会发生。几十英里外的房屋会像纸板一样倒塌,数百万人将失去生命。
Fortunately for all 8 billion of us, planetary defense—the science of preventing asteroid impacts—is a highly active field of research. We already know that we could ram a rock with an uncrewed spacecraft to push it away from Earth. But if that’s not enough, we could need another method, one that is notoriously difficult to test in real life: a nuclear explosion. Read our story about the scientists who, despite the odds, are trying to do exactly that. —Robin George Andrews 幸运的是,对于我们 80 亿人来说,行星防御——即预防小行星撞击的科学——是一个非常活跃的研究领域。我们已经知道,可以通过无人航天器撞击小行星来改变其轨道,使其远离地球。但如果这还不够,我们可能需要另一种方法,一种在现实中极难测试的方法:核爆炸。阅读我们的报道,了解那些尽管困难重重,却仍试图实现这一目标的科学家们。——Robin George Andrews