AI for science needs reasoning, not just data
AI for science needs reasoning, not just data
科学研究中的人工智能:不仅需要数据,更需要推理
Every few decades, someone announces that science has reached its end. In 1903, the revered physicist Albert Michelson wrote that the “facts of physical science have all been discovered.” In the 1980s, Stephen Hawking predicted that theoretical physics might be finished by the end of the century. 每隔几十年,总有人宣称科学已经走到了尽头。1903年,备受尊崇的物理学家阿尔伯特·迈克尔逊(Albert Michelson)写道:“物理科学的事实已经全部被发现。”20世纪80年代,斯蒂芬·霍金(Stephen Hawking)也曾预言,理论物理学可能会在本世纪末终结。
With the explosive arrival of artificial intelligence, the feeling is in the air again—this time accompanied by a Nobel Prize. In 2024, Demis Hassabis and John Jumper of Google DeepMind were awarded part of the Nobel in chemistry for their neural network AlphaFold, which predicts the three-dimensional structures of proteins by learning from thousands of experimentally measured shapes. 随着人工智能的爆发式到来,这种感觉再次弥漫开来——这一次还伴随着诺贝尔奖的加持。2024年,Google DeepMind的德米斯·哈萨比斯(Demis Hassabis)和约翰·江珀(John Jumper)因其神经网络AlphaFold获得了诺贝尔化学奖的部分奖项。AlphaFold通过学习数千种实验测得的形状,成功预测了蛋白质的三维结构。
This devilish problem had resisted systematic attacks for half a century; AlphaFold seemed to have solved it once and for all, and the world became fixated on the promise of its approach. Hassabis and his team called AlphaFold “the template for how AI can accelerate all of science to digital speed.” 这个棘手的难题在半个世纪里一直难以被系统性攻克;AlphaFold似乎一劳永逸地解决了它,全世界都沉迷于这种方法所带来的前景。哈萨比斯及其团队将AlphaFold称为“人工智能如何将所有科学研究加速至数字速度的模板”。
A wave of startups building foundation models for biology, chemistry, and materials discovery raised billions of dollars, buoyed by DeepMind’s success. AlphaFold had shown that the combination of AI and sufficient data could make groundbreaking discoveries (even if we did not understand the underlying mechanisms involved), and it seemed, once again, that a path through the rest of science was laid out before us. 受DeepMind成功的鼓舞,一波为生物学、化学和材料发现构建基础模型的初创公司筹集了数十亿美元。AlphaFold证明了人工智能与充足数据的结合可以带来突破性的发现(即使我们并不理解其中涉及的底层机制),这似乎再次表明,通往其余科学领域的道路已经铺就。
To be sure, AI will bring extraordinary changes to science, but it has become increasingly clear that AlphaFold, and things like it, may not be the best template for that metamorphosis. Though it is a profound achievement, the conditions that produced the likes of AlphaFold are rare, and the time it will take to meet those conditions in other fields will be measured in decades, not years. Instead, the acceleration of science will come about thanks to another approach: AI agents. 诚然,人工智能将为科学带来非凡的变革,但越来越明显的是,AlphaFold及其同类产品可能并不是实现这种蜕变的最佳模板。尽管这是一项深远的成就,但产生AlphaFold这类成果的条件非常罕见,而在其他领域满足这些条件所需的时间将以十年而非几年计。相反,科学的加速将归功于另一种方法:人工智能体(AI agents)。
The primary condition for AlphaFold’s success was the existence of the Protein Data Bank, a data set of roughly 170,000 experimentally validated protein structures on which DeepMind’s team could train its model. The creation of the Protein Data Bank was not simple: It took 53 years of international scientific cooperation and, by a recent estimate, roughly $21 billion worth of experimental work to assemble. AlphaFold成功的首要条件是“蛋白质数据库”(Protein Data Bank)的存在,这是一个包含约17万个经实验验证的蛋白质结构的数据集,DeepMind团队正是利用它来训练模型。蛋白质数据库的创建并不简单:它历经了53年的国际科学合作,据最近估计,其汇编工作耗资约210亿美元。
Efforts of that scale are infamously difficult to fund, next to impossible to coordinate, and hugely time-consuming to execute; they have often been unsuccessful as a result. But even in fields with the requisite cohesion and resources, and where the relevant data are not rendered inaccessible by commercial ownership, another barrier is too little discussed: the scientific impossibility of generating comparable data. 这种规模的努力在筹集资金方面极其困难,协调起来几乎不可能,执行起来也极其耗时;因此,它们往往以失败告终。但即使在具备必要凝聚力和资源、且相关数据未因商业所有权而无法获取的领域,另一个被讨论得太少的障碍是:生成可比数据的科学上的不可能性。
In the case of protein structures, the key experimental technique—protein crystallography—is an unusually replicable and dependable tool, so much so that over 25 Nobel Prizes have relied on it. But in most of experimental science, results vary more often than not. Cell lines drift. Chemicals have trace contaminants. Lab humidity changes. The creation of measured datasets that will be consistent enough, accurate enough, precise enough, and scalable enough to train a modern neural network in biology or most of chemistry would require new kinds of measurement and new standardized approaches—none of which will be ready anytime soon. 以蛋白质结构为例,其关键实验技术——蛋白质晶体学——是一种极其可复制且可靠的工具,以至于有超过25个诺贝尔奖都依赖于它。但在大多数实验科学中,结果往往各不相同。细胞系会发生漂移,化学品含有微量污染物,实验室湿度会变化。要创建足够一致、准确、精确且可扩展的测量数据集来训练生物学或大多数化学领域的现代神经网络,需要新的测量方法和新的标准化手段——而这些在短期内都无法实现。
Of course, there are a handful of fields where these requirements are met: weather forecasting, much of genomics, very limited areas of chemistry. These may see AlphaFold-style breakthroughs soon, if they haven’t already. Government support for the production and coordination of those datasets will be critical, as the US National Security Commission on Emerging Biotechnology has argued. But for most open questions in science, we will need a different plan, at least in the short term. 当然,有少数领域满足这些要求:天气预报、大部分基因组学以及非常有限的化学领域。如果尚未实现,这些领域可能很快就会看到AlphaFold式的突破。正如美国新兴生物技术国家安全委员会所主张的那样,政府对这些数据集的生产和协调的支持将至关重要。但对于科学界大多数悬而未决的问题,我们至少在短期内需要一个不同的计划。
Luckily, something quieter and more modest has begun to show promise. Scientists have always reasoned under uncertainty. Biologists working to identify new drug targets have never had perfect datasets. Instead, they combine docking calculations and known structures, factor in molecular dynamics, run a handful of binding assays, and use their judgment to weigh each method according to its particular strengths and points of failure. 幸运的是,一些更安静、更低调的方法已经开始显现出前景。科学家们总是在不确定性中进行推理。致力于识别新药物靶点的生物学家从未拥有过完美的数据集。相反,他们结合对接计算和已知结构,考虑分子动力学,运行少量结合测定,并利用自己的判断力,根据每种方法的特定优势和缺陷来权衡结果。
The skill of science is not in any single tool; it is synthesizing what many tools produce, and revising the results as the evidence comes in. This is how most working research actually proceeds. But until very recently, no software could do it. Agents now can. 科学的精髓不在于任何单一工具,而在于综合多种工具的产出,并随着证据的出现不断修正结果。这就是大多数实际研究的进行方式。但直到最近,还没有软件能做到这一点。现在,人工智能体可以了。
Simply put, an agent is an AI reasoning engine that has been given access to tools—digital or physical—and the capabilities to use them. Over the last few years, a fundamental architectural shift in AI has enabled the rapid proliferation of these programs, which are powered by large language models, dramatically reducing the need for scientifically specialized datasets. 简而言之,人工智能体是一个被赋予了数字或物理工具访问权限及其使用能力的人工智能推理引擎。在过去几年中,人工智能领域的一项根本性架构转变促成了这些程序的迅速普及。它们由大语言模型驱动,极大地降低了对科学专业数据集的需求。
For science, this technological advancement represents a foundational change: it has allowed us to create digital tools that can mimic the iterative, highly contingent process of actual research. While tools like AlphaFold apply a powerful approach to a limited question, agents are inherently generalists. They do not represent a new way to do science—instead, they digitally model the human process of discovery. 对于科学而言,这一技术进步代表着根本性的变革:它使我们能够创建数字工具,模拟实际研究中那种迭代的、高度偶然的过程。虽然像AlphaFold这样的工具是将强大的方法应用于有限的问题,但人工智能体本质上是通才。它们并不代表一种新的科学研究方式,而是以数字方式模拟了人类的发现过程。
Consider Google’s AI Co-Scientist, announced in May. Researchers gave it a one-page brief and a goal: Figure out how antibiotic resistance spreads between bacterial species, a key driver of drug-resistant infections. The system spun up sub-agents. One drafted hypotheses from the literature. Another picked them apart like a peer reviewer. A third ran tournaments to rank the strongest candidates. A fourth refined the winning hypothesis. The agent concluded that resistance genes were hitching rides on bacterial viruses, borrowing whichever virus could ferry them into a new host. The hypothesis was correct. 以5月份发布的Google AI Co-Scientist为例。研究人员给了它一份一页纸的简报和一个目标:弄清楚抗生素耐药性如何在细菌物种之间传播,这是导致耐药性感染的关键因素。该系统启动了多个子智能体:一个从文献中起草假设;另一个像同行评审员一样对其进行剖析;第三个进行竞赛以对最强的候选假设进行排名;第四个对胜出的假设进行完善。该智能体最终得出结论:耐药基因搭上了细菌病毒的“便车”,借用任何能将它们运送到新宿主的病毒进行传播。这一假设是正确的。