When can we say AI made a scientific discovery?
When can we say AI made a scientific discovery?
我们何时才能说人工智能做出了科学发现?
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Last Wednesday, Anthropic announced that earlier this year it had launched a molecular biology lab, where Claude agents read and conjecture about hard biology problems and human scientists run experiments on what they report. And this AI-powered lab, the company said, had made its first discovery. 本文最初发表于我们的 AI 每周通讯《算法》(The Algorithm)。若想第一时间在收件箱中获取此类报道,请点击此处订阅。上周三,Anthropic 宣布其在今年早些时候建立了一个分子生物学实验室。在该实验室中,Claude 智能体负责阅读并推测复杂的生物学难题,而人类科学家则针对它们报告的内容进行实验。该公司称,这个由 AI 驱动的实验室已经做出了它的第一个发现。
To understand what Anthropic says its system did, imagine you’re flipping through a library of millions of DNA sequences, amassed as scientists sequence more and more of the living world. One step toward a breakthrough might be finding a peculiar sequence that encodes an interesting enzyme, perhaps. Then you’d need to figure out what that enzyme does and, eventually, how to manipulate it to do something useful. 要理解 Anthropic 所称的系统成果,试想你正在翻阅一个包含数百万条 DNA 序列的数据库,这些数据是随着科学家对生命世界测序工作的不断深入而积累起来的。取得突破的一步可能是找到一段编码某种有趣酶的特殊序列。随后,你需要弄清楚这种酶的功能,并最终研究如何操纵它以实现某种有用的用途。
What Anthropic says its system of 950 agents found after 21 hours was not a brand-new sequence. The agents instead flagged a repeating pattern surrounding a known enzyme, a particular pattern Anthropic said hadn’t been catalogued before. But if you read through Anthropic’s announcement, which calls this pattern “reminiscent” of what led to the gene-editing technology CRISPR that “has already transformed science and medicine,” it sounds as if this army of agents really found something of note. Anthropic 表示,其由 950 个智能体组成的系统在 21 小时后发现的并非全新的序列。相反,这些智能体标记了一个围绕已知酶的重复模式,Anthropic 称该特定模式此前从未被记录过。然而,如果你仔细阅读 Anthropic 的公告,会发现它将这种模式描述为“令人联想到”促成基因编辑技术 CRISPR 的发现,并称 CRISPR “已经改变了科学和医学”。这听起来仿佛这支智能体大军真的发现了什么了不起的东西。
These claims have angered some biologists. A viral post from one, subsequently endorsed by the chair and CEO of the drugmaker Eli Lilly, said that “finding a weird cluster of genes and repeats is often the easy part. The hard part, and where the real discoveries come from, is figuring out what the system actually does.” The agents helped with some laboratory grunt work, in other words. But a discovery it is not. 这些说法激怒了一些生物学家。一位生物学家发布的一篇病毒式传播的帖子(随后得到了礼来公司董事长兼首席执行官的认可)指出:“找到一簇奇怪的基因和重复序列通常是容易的部分。真正的难点,也是真正科学发现的来源,在于弄清楚该系统到底有什么功能。”换句话说,这些智能体只是协助完成了一些实验室的苦差事,但这算不上是科学发现。
It’s a reminder that even if AI does something impressive—like finding a pattern in a mass of biological data that would be difficult to perceive with human eyes alone—the result itself may not constitute a breakthrough for science. What is novel for AI may be routine, unsurprising, or simply not that consequential to a biologist. 这提醒我们,即使 AI 做了一些令人印象深刻的事情——比如在海量生物数据中发现人类肉眼难以察觉的模式——其结果本身也未必构成科学上的突破。对 AI 而言的新颖之处,对生物学家来说可能只是常规操作、意料之中,或者根本无关紧要。
Muddying the issue further, Mario Rodríguez Mestre, a biologist at the University of Copenhagen, said over the weekend that his team had already discovered this particular pattern, the New York Times reported. Mestre, who regularly chatted with Claude in his work, wondered whether Anthropic’s team had learned from his conversations. Anthropic denies this, but Mestre says he’s stopping all use of Claude anyway. 据《纽约时报》报道,哥本哈根大学的生物学家 Mario Rodríguez Mestre 在周末表示,他的团队此前已经发现了这种特定的模式,这让问题变得更加扑朔迷离。Mestre 在工作中经常与 Claude 交流,他怀疑 Anthropic 的团队是否从他的对话中获取了信息。Anthropic 对此予以否认,但 Mestre 表示他无论如何都会停止使用 Claude。
Part of the problem here is that AI companies aren’t presenting their systems simply as tools scientists can use, like microscopes or supercomputers. They’re insisting that the AI systems are making discoveries themselves. To some, that approach is incompatible with how science actually works, with new knowledge more typically emerging from collaboration and an ever-growing arsenal of tools. It’s also making people more skeptical of genuine progress when it happens. 问题的部分原因在于,AI 公司并没有将他们的系统仅仅呈现为科学家可以使用的工具(如显微镜或超级计算机),而是坚持认为 AI 系统本身正在做出发现。在一些人看来,这种做法与科学的实际运作方式不符,因为新知识通常源于协作和不断增长的工具库。这也使得人们在真正的科学进步发生时,反而变得更加怀疑。
Whittling 200,000 candidates down to a few worth exploring is no small feat; it is legitimate scientific work. The fact that a general-purpose chatbot could do that work is notable, even if humans helped steer it and ultimately ran the experiments. But once the standard is whether Claude itself made a discovery, all that becomes evidence for one side or the other in a debate that has only two answers: breakthrough or bust. 将 20 万个候选对象筛选到几个值得探索的对象并非易事;这是合法的科学工作。一个通用聊天机器人能够完成这项工作这一事实本身就值得注意,即使有人类协助引导并最终执行了实验。但一旦标准变成了“Claude 是否亲自做出了发现”,所有这些就成了辩论双方的证据,而这场辩论只有两个答案:要么是突破,要么是失败。
Once we’re judging AI by whether it has made a discovery, it’s also tempting to shift the goalposts even after it really does seem to notch a win. Earlier this month, OpenAI said its own team agents had cracked a million-dollar problem in mathematics. But a couple of weeks later, nearly every AI skeptic in my feed was sharing an article asking whether it was the math problem that really mattered. 一旦我们开始用“是否做出发现”来评判 AI,即使它看起来确实取得了一次胜利,人们也很容易去移动球门(改变评价标准)。本月早些时候,OpenAI 表示其团队智能体攻克了一个价值百万美元的数学难题。但几周后,我信息流中几乎所有的 AI 怀疑论者都在转发一篇文章,质疑那个数学问题本身是否真的重要。
To be clear, the piece did not argue that OpenAI’s solution was wrong. Instead, it argued that the particular result may not be the one mathematicians care most about. Throw in the accusation by a mathematician that the models may have used some of his work without credit, and people are left thinking either OpenAI cheated or the solution wasn’t important anyway. Or both. 需要明确的是,那篇文章并没有争论 OpenAI 的解决方案是错误的,而是认为该特定结果可能并非数学家最关心的。再加上一位数学家指控这些模型可能在未注明出处的情况下使用了他的研究成果,人们最终认为要么是 OpenAI 作弊了,要么是这个解决方案本身就不重要。或者两者皆有。
That’s part of what concerns Lucas Harrington, the biologist who wrote the post critiquing Anthropic’s announcement. He closed with a suggestion: AI companies, he said, should “set the bar high now, so that when an AI actually discovers a fundamentally new biological mechanism, everyone appreciates how big a deal it is.” But as OpenAI’s Sam Altman and Anthropic’s Dario Amodei race to one-up each other, raising the bar for scientific breakthroughs by AI might be the last thing on their minds. 这正是撰写帖子批评 Anthropic 公告的生物学家 Lucas Harrington 所担忧的一部分。他在结尾处提出了一个建议:AI 公司应该“现在就设定高标准,这样当 AI 真正发现一种根本性的新生物机制时,每个人都能意识到这是一件多么了不起的事情。”然而,随着 OpenAI 的 Sam Altman 和 Anthropic 的 Dario Amodei 在竞争中不断试图超越对方,提高 AI 科学突破的门槛恐怕是他们最不关心的事情。