OpenAI’s math solutions aren’t meeting the field’s standards yet

本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。

OpenAI’s math solutions aren’t meeting the field’s standards yet

When OpenAI released hundreds of claimed solutions to some of the world’s hardest math problems this week, the frontier lab said that it had consulted an advisory group of elite mathematicians to avoid the controversy that came with the last time one of its models solved a long-standing problem in the field.

本周,当 OpenAI 发布了数百个号称能解决世界上最难数学问题的方案时,这家前沿实验室表示,他们曾咨询过一个由顶尖数学家组成的顾问小组,以避免重蹈此前其模型在解决该领域长期难题时引发争议的覆辙。

But OpenAI fell short of those standards, particularly where the mathematicians emphasized the need for human understanding of a mathematical result. That’s especially concerning after a new paper highlighted gaps between the natural language and formally expressed solution to a million-dollar problem ostensibly solved by OpenAI’s models.

但 OpenAI 并未达到这些标准,特别是在数学家们强调需要人类理解数学结果的方面。在一篇新论文指出 OpenAI 模型所“解决”的一个百万美元难题在自然语言与形式化表达之间存在差距后,这一点尤为令人担忧。

The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton University’s Institute for Advanced Studies, is made up of nine prominent researchers at institutions around the world. The organization released guidelines for frontier labs solving math problems at the end of September.

由普林斯顿大学高等研究院主办的数学与人工智能顾问小组(AGMAI)由来自全球各机构的九位知名研究人员组成。该组织在 9 月底发布了针对前沿实验室解决数学问题的指导方针。

In a statement on the latest set of proofs, the AGMAI said that “it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully.” However, the organization’s first request was “to stop testing advanced mathematical problems on proprietary models.”

在关于最新一批证明的声明中,AGMAI 表示:“最终由数学界来评估我们的建议在多大程度上得到了成功遵循。”然而,该组织的首要要求是“停止在专有模型上测试高级数学问题”。

OpenAI’s release explicitly says that it is evaluating its proprietary models using open research problems in mathematics. The advisory group did not respond when asked by TechCrunch for a more thorough evaluation of OpenAI’s latest proof release.

OpenAI 的发布内容明确表示,它正在使用数学领域的开放研究问题来评估其专有模型。当 TechCrunch 要求该顾问小组对 OpenAI 最新的证明发布进行更彻底的评估时,他们没有做出回应。

The lab clearly followed some of its principles, including releasing results as soon as possible and including information about how the models reached their conclusions. But not for all of them: Just 10 of the 719 manuscripts included releases of the model’s chain of thought.

该实验室显然遵循了其中的一些原则,包括尽快发布结果,并包含有关模型如何得出结论的信息。但并非全部如此:在 719 份手稿中,只有 10 份包含了模型的思维链发布。

For papers that people don’t understand, the mathematicians suggested the proofs should be formalized — but just 42% of the proofs released by OpenAI had not undergone this process. Ultimately, it’s still not clear that OpenAI is taking “responsibility for ensuring that human understanding will follow” when releasing its proofs, in accordance to the AGMAI principles.

对于人们无法理解的论文,数学家们建议应对证明进行形式化处理——但 OpenAI 发布的证明中,只有 42% 没有经过这一过程。归根结底,目前尚不清楚 OpenAI 在发布证明时是否按照 AGMAI 的原则,承担了“确保人类能够理解后续内容”的责任。

AGMAI suggested that OpenAI should help fund the work of human mathematicians who will be required to make the lab’s solutions meaningful in any real way.

AGMAI 建议 OpenAI 应资助人类数学家的工作,因为要使该实验室的解决方案具有任何实际意义,人类数学家的参与是必不可少的。

“Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is ‘solved’, and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field,” Terence Tao, a prominent mathematician who has criticized OpenAI’s approach, wrote on social media after the release.

“问题正由人工智能提示词工程师自主解决,他们一旦达到最初的‘解决’目标,就对更广泛的领域本身毫无兴趣,而且他们对人工智能输出的理解也不足以回答关于结果的问题、进行演讲或与该领域的其他人进行互动,”曾批评 OpenAI 方法的著名数学家陶哲轩在发布后于社交媒体上写道。

That problem is exemplified by a paper released this week by mathematicians at the University of Cambridge and King’s College in London that questions the way frontier labs are approaching these challenges.

本周由剑桥大学和伦敦国王学院的数学家发表的一篇论文体现了这一问题,该论文质疑了前沿实验室处理这些挑战的方式。

When AI models solve mathematical problems, they first create a “natural language” explanation, then try to express that result in Lean, a programming language that in theory confirms the accuracy of the proof by compiling it as code.

当人工智能模型解决数学问题时,它们首先会创建一个“自然语言”解释,然后尝试用 Lean 语言表达该结果。Lean 是一种编程语言,理论上可以通过将其编译为代码来确认证明的准确性。

However, there may be problems with the way the models translate their natural language proofs into code; this paper documents at least two discrepancies between the natural language proof and the Lean code behind the solution OpenAI has offered to a problem derived from the Navier-Stokes equations that describe the complex behavior of fluids.

然而,模型将自然语言证明转换为代码的方式可能存在问题;这篇论文记录了 OpenAI 为一个源自描述流体复杂行为的纳维-斯托克斯方程的问题所提供的解决方案中,自然语言证明与背后的 Lean 代码之间至少存在两处差异。

These discrepancies don’t necessarily disprove either solution, but they do raise questions on whether we can simply rely on models to formalize their own solutions without human involvement.

这些差异并不一定能推翻任何一个解决方案,但它们确实引发了质疑:在没有人类参与的情况下,我们是否能仅仅依靠模型来形式化它们自己的解决方案。

That’s one reason that AGMAI asked OpenAI to “include machine-readable metadata correlating the natural language and formal artifacts,” something that the frontier lab did not do with these releases.

这也是 AGMAI 要求 OpenAI “包含关联自然语言和形式化制品的机器可读元数据”的原因之一,而这家前沿实验室在这些发布中并未做到这一点。

“Because of the phenomenon of mistranslations — as highlighted in this paper — the NL proof by OpenAI and other autoformalised Lean proofs should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to,” the authors of the “lost in translation” paper conclude.

“由于这篇论文中强调的误译现象,OpenAI 的自然语言证明以及其他自动形式化的 Lean 证明,在没有经过其他证明所必须经历的同行评审和审查之前,不应被直接信任,”这篇名为《迷失在翻译中》(lost in translation)的论文作者总结道。

Mathematicians stress that when new results are discovered by humans, they take responsibility for them and engage with the broader community through papers, talks, and seminars. That process increases understanding of the solutions, finds strategies that can be used to solve other problems, and allows the new knowledge to be applied in practical fields.

数学家们强调,当人类发现新成果时,他们会对其负责,并通过论文、演讲和研讨会与更广泛的社区进行交流。这一过程增加了对解决方案的理解,发现了可用于解决其他问题的策略,并使新知识能够应用于实际领域。

When a model is prompted to solve a hard problem and spits out a solution, “there is not human understanding of them at the point of release, and now the work begins,” Harvard University mathematics professor Melanie Wood told TechCrunch.

哈佛大学数学教授梅兰妮·伍德(Melanie Wood)告诉 TechCrunch:“当模型被提示解决一个难题并吐出一个解决方案时,在发布之时并没有人类对其进行理解,而现在工作才刚刚开始。”