OpenAI fought dirty on career-making math problem, says NYU mathematician
OpenAI fought dirty on career-making math problem, says NYU mathematician
纽约大学数学家称 OpenAI 在争夺数学难题成果时手段卑劣
NYU mathematics professor Tristan Buckmaster announced three proofs on Tuesday with a preliminary finding on one of the major unsolved problems in theoretical mathematics. The findings, made in collaboration with Anthropic mathematician Levent Alpöge and using both Codex and Claude AI models, are significant in themselves — but they’re also accompanied by an unusual controversy surrounding OpenAI’s attempts to solve the same problem. 纽约大学数学教授 Tristan Buckmaster 周二宣布了三项证明,并对理论数学中一个主要的未解难题给出了初步研究结果。这些研究成果是与 Anthropic 公司的数学家 Levent Alpöge 合作完成的,并使用了 Codex 和 Claude AI 模型。这些发现本身意义重大,但随之而来的是一场围绕 OpenAI 试图解决同一问题的不同寻常的争议。
“There is another part of this story,” Buckmaster wrote in his statement announcing the proofs, “and one that, honestly, I very much wish I did not have to be concerned with.” According to the statement, a parallel effort by OpenAI built on their work before it became public, leading to a tangle of academic rivalries and conflicting claims. “这个故事还有另一面,”Buckmaster 在宣布这些证明的声明中写道,“老实说,我非常希望自己不必去理会这一面。”根据该声明,OpenAI 在他们的研究成果公开之前,就基于这些成果开展了平行研究,导致了一场学术竞争与相互矛盾的声明纠纷。
Shortly after the Buckmaster’s statement, OpenAI published a full proof of the Navier–Stokes existence and smoothness problem, which Buckmaster’s findings had taken steps towards. According to OpenAI, the proof was discovered by an unreleased next-generation model, which has tackled a range of different unsolved problems over the past week. All told, the week-long effort consumed 300 billion output tokens — $22.5 million worth of compute, if charged at current Astra rates. 在 Buckmaster 发表声明后不久,OpenAI 发布了关于“纳维-斯托克斯存在性与光滑性”问题的完整证明,而 Buckmaster 的研究此前已向这一目标迈出了步伐。据 OpenAI 称,该证明是由一个尚未发布的下一代模型发现的,该模型在过去一周内解决了一系列不同的未解难题。总计来看,这项为期一周的工作消耗了 3000 亿个输出 Token——如果按目前的 Astra 费率计算,相当于价值 2250 万美元的算力。
The Navier-Stokes existence and smoothness problem is one of the seven Millennium Prize problems — a set of major unsolved math problems, each carrying a $1 million bounty from Clay Mathematics Institute for the first person or group to provide a solution. The Navier-Stokes equations are widely used in fluid mechanics but poorly understood in theoretical terms. A solution would represent a significant advance in the collective understanding of mathematical physics. “纳维-斯托克斯存在性与光滑性”问题是七大“千禧年大奖难题”之一。这是一组重大的未解数学难题,克雷数学研究所为第一个提供解决方案的个人或团队设立了 100 万美元的奖金。纳维-斯托克斯方程在流体力学中被广泛应用,但在理论层面上的理解却十分有限。解决这一问题将代表人类对数学物理学的集体理解迈出重大一步。
While Buckmaster and Alpöge were finalizing their own results, they learned that “information about our progress had been passed to OpenAI.” When they contacted OpenAI, they were told that OpenAI had already achieved a full proof of the central problem. But when they asked follow-up questions about when OpenAI had begun its research into the problem and how much human input was involved, the answers became more evasive. 当 Buckmaster 和 Alpöge 即将完成他们的研究成果时,他们得知“关于我们进展的信息已被泄露给 OpenAI”。当他们联系 OpenAI 时,对方称 OpenAI 已经完成了该核心问题的完整证明。但当他们追问 OpenAI 是何时开始研究该问题以及涉及多少人工干预时,对方的回答变得闪烁其词。
“It emerged that an entire team had been working on the problem,” Buckmaster said, “and that an insane amount of compute had been used…. Eventually, it was agreed that [the first prompt] had been sent in the past few days, after information about our work had reached OpenAI.” If true, that would suggest the OpenAI team had become convinced that Buckmaster and Alpöge’s approach was the right one, and decided to use its material advantage in computing resources to reach a formal proof first. “后来发现,整个团队一直在研究这个问题,”Buckmaster 说,“而且使用了极其惊人的算力……最终,对方承认(第一个提示词)是在过去几天内发送的,也就是在我们的工作信息传到 OpenAI 之后。”如果属实,这表明 OpenAI 团队确信 Buckmaster 和 Alpöge 的研究路径是正确的,并决定利用其在计算资源上的物质优势,抢先完成正式证明。
OpenAI’s post confirms much of this timeline, specifically saying that the latest effort began on September 1, inspired by rumors that two Millenium Prize problem had been solved. Additionally, the post confirms the ongoing conversations with Buckmaster and Alpöge. OpenAI 的文章证实了这一时间线的大部分内容,明确表示最新的研究工作始于 9 月 1 日,灵感来源于关于两道“千禧年大奖难题”已被解决的传言。此外,该文章还证实了与 Buckmaster 和 Alpöge 之间正在进行的沟通。
Although the problem is widely pursued among mathematicians, the specific tactic taken by Buckmaster and his collaborator is far less common. As a result, Buckmaster found it suspicious that OpenAI ended up taking the same approach at the same time. “The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack,” Buckmaster wrote. “Almost nobody else I know of was working on it,” he continued. “It is not the direction one arrives at in a few days by giving a model the problem statement.” 尽管数学家们广泛研究这一问题,但 Buckmaster 及其合作者采取的具体策略并不常见。因此,Buckmaster 认为 OpenAI 在同一时间采取相同的方法非常可疑。“通过平滑力(smooth force)解决克雷难题的路径,即 Fefferman 问题陈述中的选项 c 和 d,是 Luis 和 Diego 开辟的路径,也是 Levent 和我悄悄选择的攻克方向,”Buckmaster 写道。“据我所知,几乎没有其他人在这方面工作,”他继续说道,“这不是给模型一个问题陈述,几天内就能得出的方向。”
While Alpöge is employed by Anthropic, he was not conducting this research on the company’s behalf. As a result, the duo used a mix of models, relying primarily on OpenAI’s Codex in their work. Even so, Alpöge’s affiliation with a rival lab seems to have been a sore point for OpenAI, and Buckmaster alleges that Bubeck asked him to remove Alpöge’s credit as part of a proposed compromise. 虽然 Alpöge 受雇于 Anthropic,但他并非代表该公司进行这项研究。因此,两人混合使用了多种模型,主要依赖 OpenAI 的 Codex 进行工作。即便如此,Alpöge 与竞争对手实验室的关系似乎成了 OpenAI 的心病,Buckmaster 声称 Bubeck 曾要求他删除 Alpöge 的署名,作为拟议妥协方案的一部分。
When Buckmaster pushed to make the dispute public, he says that Bubeck replied: “Why would you ruin your career?” Buckmaster says that when he pushed back, Bubeck followed up with: “If you don’t want me to be nice, then I don’t have to be nice.” 当 Buckmaster 坚持要公开这场纠纷时,他说 Bubeck 回复道:“你为什么要毁了自己的职业生涯?”Buckmaster 表示,当他反驳时,Bubeck 接着说:“如果你不想让我表现得友善,那我也没必要友善。”
Buckmaster also raised concerns that, because he used Codex extensively in assembling the project, information from his work could have informed OpenAI’s own efforts to solve the problem. OpenAI reserves the right to train models on Codex interactions, although users are able to opt-out. If the OpenAI team used a model trained on Buckmaster’s own Codex interactions, it’s plausible that it could have regurgitated his work when faced with a similar problem. Buckmaster 还提出担忧,由于他在项目组装过程中大量使用了 Codex,他工作中的信息可能为 OpenAI 解决该问题的努力提供了参考。OpenAI 保留使用 Codex 交互数据训练模型的权利,尽管用户可以选择退出。如果 OpenAI 团队使用了基于 Buckmaster 的 Codex 交互数据训练的模型,那么当模型面对类似问题时,很有可能会“吐出”他的研究成果。
In its own post, OpenAI downplayed the possibility that regurgitation could have been involved. “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem,” the post reads. “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).” OpenAI 在其文章中淡化了模型“吐出”数据的可能性。“我们(研究人员和智能体)在他们公开成果之前,没有通过任何方式看到他们的工作——特别是,在解决此问题时没有访问任何特定的用户数据,”文章写道。“虽然可能性不大,但我们不能排除他们使用我们产品所产生的去标识化数据有助于改进我们的模型。然而,我们的证明存在显著差异,甚至在欧拉情形下(强制与非强制),所证明的精确结果也是不同的。”
Regardless, the issue is likely to reignite the ongoing debate about AI’s role in mathematical research, and OpenAI’s specific incentives. For his part, Buckmaster seems to believe the best answer is to get as much information about the research out into the public eye. 无论如何,这一事件很可能会重新点燃关于 AI 在数学研究中作用的持续争论,以及 OpenAI 的具体动机。就 Buckmaster 而言,他似乎认为最好的回应方式是将尽可能多的研究信息公之于众。