Don’t be fooled—LLMs don’t reason

Don’t be fooled—LLMs don’t reason

别被骗了——大语言模型并不具备推理能力

On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board in what looked like a gift to its human opponent. Move 37 in game two of the five-game match looked so absurd that some commentators thought it was a programming glitch. It wasn’t. AlphaGo won the game, ultimately triumphing 4-1 over Lee Sedol, one of the greatest professional Go players of all time. 2016年3月的一个下午,在首尔,我看着一个我参与构建的程序在围棋棋盘的第五线上落下一子,这一步看起来就像是送给人类对手的一份大礼。在五局三胜制的第二局比赛中,第37手棋看起来如此荒谬,以至于一些评论员认为这是程序故障。事实并非如此。AlphaGo赢得了那局比赛,并最终以4比1战胜了史上最伟大的职业围棋选手之一李世石。“我原以为AlphaGo是基于概率计算的,它仅仅是一台机器,”李世石事后说道,“但当我看到这一手棋时,我改变了想法。AlphaGo确实具有创造力。”

When Deep Blue defeated then reigning world chess champion Garry Kasparov in 1997, it did so by looking six to eight moves ahead per player and evaluating 200 million chess positions per second, using rules hard-coded by humans. Go is a vastly more complex game. A stone’s worth depends on how distant groups and territory unfold over dozens of moves. Computing even a fraction of the possible outcomes would take a supercomputer billions of years. To win, AlphaGo had to sense who was ahead at a glance and even invent moves no human had thought to play. That is why many accounts of AlphaGo’s match against Lee portray move 37 as a flash of pure machine intuition. But that is a misunderstanding. It was actually AlphaGo’s powers of reasoning that made this creative choice—and these are powers that today’s AI lacks. 1997年,当“深蓝”击败当时的世界象棋冠军加里·卡斯帕罗夫时,它是通过每方预判六到八步棋,并利用人类硬编码的规则每秒评估2亿个棋局位置来实现的。围棋则复杂得多。一颗棋子的价值取决于数十步棋之后棋群和地盘的演变。即使是计算出所有可能结果的一小部分,也需要超级计算机运行数十亿年。为了获胜,AlphaGo必须一眼看出谁处于领先地位,甚至要发明出人类从未想过的走法。这就是为什么许多关于AlphaGo与李世石对弈的报道将第37手棋描述为纯粹的机器直觉闪现。但这是一种误解。实际上,正是AlphaGo的推理能力做出了这一创造性的选择——而这些能力正是当今人工智能所缺乏的。

If we want future AI systems to produce trustworthy results and really novel insights in fields like science and medicine, we need to equip them with genuine reasoning capabilities of this kind. AlphaGo is made up of two systems. The first, its policy network, was trained to guess what move a strong human would play. This “intuitive” part regarded move 37 as nothing special—a play that had a roughly one in 10,000 chance of being made by an expert human player. What made AlphaGo choose it was the program’s search machinery, which looked beyond immediate plausibility and weighed the future consequences of proposed moves. It explicitly constructed and searched a game tree with thousands of branches, each representing a different possible future. 如果我们希望未来的人工智能系统能在科学和医学等领域产生值得信赖的结果和真正新颖的见解,我们就需要为它们配备这种真正的推理能力。AlphaGo由两个系统组成。第一个是它的策略网络,经过训练可以猜测人类高手会走哪一步。这个“直觉”部分认为第37手棋没什么特别的——人类专家选手走出这一步的概率大约只有万分之一。让AlphaGo选择这一步的是程序的搜索机制,它超越了眼前的合理性,权衡了所提议走法对未来的影响。它明确地构建并搜索了一棵拥有数千个分支的博弈树,每一条分支都代表着一种不同的未来可能。

A well-known theory in the behavioral sciences, popularized by Daniel Kahneman, distinguishes between two modes of human thought: System 1 is fast, gut-level, effortless; system 2, slow, step-by-step, and deliberative. AlphaGo offered a striking machine analogue of that split. Its networks supplied the hunches—this move looks promising, this position looks won—and its search supplied the deliberation, testing those hunches against the moves and countermoves that would follow. As in human cognition, neither half works alone. Intuition alone would never have opted for move 37, and brute-force search would have struggled to sieve through all the many possible moves. 行为科学中一个著名的理论(由丹尼尔·卡尼曼推广)区分了人类思维的两种模式:系统1是快速、直觉、不费力的;系统2是缓慢、循序渐进、深思熟虑的。AlphaGo提供了一个惊人的机器类比。它的网络提供了直觉——这一步看起来很有希望,这个位置看起来赢定了——而它的搜索提供了深思熟虑,通过后续的走法和应对来检验这些直觉。正如人类认知一样,这两者缺一不可。仅靠直觉永远不会选择第37手棋,而蛮力搜索也难以从所有可能的走法中筛选出最优解。

This is strikingly different from the way today’s AI models work. A large language model picks the next token, over and over. That amounts to system 1 in action—fast, associative, and surprisingly good pattern completion across almost every subject people write about. Not long after ChatGPT debuted, the field realized that language fluency alone falls short of true usefulness. The apparent solution was to make models that deliberate: Instead of answering immediately, they can now generate intermediate steps that decompose a problem, carry forward partial results, and influence subsequent reasoning—a process known as chain of thought. The gains have proved real, above all in mathematics and coding. But unlike AlphaGo’s search, this does not introduce a genuinely separate reasoning mechanism: The intermediate reasoning is still produced by the same next-token prediction process, iterated for longer before the model commits to an answer. 这与当今人工智能模型的工作方式截然不同。大语言模型只是不断地预测下一个标记(token)。这相当于系统1在运作——快速、联想式,并且在人们书写的几乎所有主题上都能出色地完成模式补全。ChatGPT问世后不久,业界就意识到,仅有语言流畅度还不足以实现真正的实用性。显而易见的解决方案是让模型进行深思熟虑:它们不再立即回答,而是可以生成中间步骤来分解问题、推进部分结果并影响后续推理——这一过程被称为“思维链”。事实证明,这种改进是有效的,尤其是在数学和编程领域。但与AlphaGo的搜索不同,这并没有引入一个真正独立的推理机制:中间的推理过程仍然是由相同的“预测下一个标记”过程产生的,只是在模型给出最终答案前进行了更长时间的迭代。

Three shortcomings prevent what chatbots do from qualifying as reasoning (in a way that a scientist might recognize). First, these models typically maintain no explicit, persistent, and inspectable epistemic state. There is no open ledger that lays out the hypotheses a model is considering, the confidence it has in various explanations, the evidence it’s weighing, and the unresolved questions it’s holding onto—all things that should be systematically revised as new information arrives. Second, they lack a clean separation between what the system knows and how it manipulates that knowledge. Knowledge and reasoning are inextricably interwoven in the weights of the neural network—there is no independent, explicitly represented set of beliefs. Third, while the chains of thought chatbots produce look like deliberation, research has demonstrated that the bots often concoct them after the fact, reaching an answer by one route but reporting another. 有三个缺陷使得聊天机器人的行为无法被归类为(科学家所认可的)推理。首先,这些模型通常不维护明确、持久且可检查的认知状态。没有一个公开的账本列出模型正在考虑的假设、它对各种解释的置信度、它正在权衡的证据以及它所持有的未决问题——所有这些都应该随着新信息的到来而系统地修订。其次,它们缺乏对“系统知道什么”与“系统如何操作这些知识”之间的清晰区分。知识和推理在神经网络的权重中交织在一起,无法分离——不存在一套独立且明确表示的信念。第三,虽然聊天机器人产生的思维链看起来像是深思熟虑,但研究表明,机器人往往是在事后编造这些过程,通过一条路径得出答案,却报告了另一条路径。

This is a problem because in the high-stakes applications we all care about, such as medicine, engineering, and scientific research, it matters not only what a system concludes but also how it arrives at its conclusion. When mistakes happen—for example, in medical diagnosis and treatment—we need to be able to pinpoint what went wrong: Was the system’s reasoning at fault, did it draw on invalid evidence, or did it make incorrect assumptions? This is why I recently left my position at Google DeepMind. I believe we need a fresh approach to machine reasoning—one that draws on AlphaGo’s architecture. AlphaGo maintains a record of what it knows about a given position: the game tree. This data structure contains all the variations, the possible futures, that AlphaGo has considered, each move and position being annotated with judgments made by its neural networks. As its reasoning progresses, AlphaGo updates the game tree and eventually synthesizes the information in it to decide which move to make. Similarly, for general reasoning a system should maintain an epistemic state that represents what the system holds as settled, what it doubts, what it has ruled out, which questions stay open. Reasoning can then be understood as a seq 这是一个问题,因为在我们都关心的关键应用领域(如医学、工程和科学研究)中,重要的不仅是系统得出了什么结论,还有它是如何得出该结论的。当错误发生时(例如在医疗诊断和治疗中),我们需要能够查明哪里出了问题:是系统的推理有误,还是它引用了无效的证据,亦或是它做出了错误的假设?这就是我最近离开Google DeepMind的原因。我相信我们需要一种新的机器推理方法——一种借鉴AlphaGo架构的方法。AlphaGo维护着它对给定棋局的认知记录:博弈树。这种数据结构包含了AlphaGo考虑过的所有变化和可能的未来,每一个走法和位置都标注了其神经网络做出的判断。随着推理的进行,AlphaGo会更新博弈树,并最终综合其中的信息来决定走哪一步。同样,对于通用推理,系统也应该维护一种认知状态,表示系统认为已确定的内容、怀疑的内容、已排除的内容以及哪些问题仍悬而未决。推理可以被理解为一种序列……