Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs

Title: Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs

标题:纳维-斯托克斯方程在翻译中丢失:为什么对人工智能自动形式化的 Lean 验证不能保证自然语言证明的正确性

Abstract: Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI’s announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified.

摘要:自动形式化正越来越多地被用于验证数学文本,包括那些由人工智能生成的文本,例如 OpenAI 宣布的关于纳维-斯托克斯方程解的爆破证明。在此过程中,人工智能系统将文本从自然语言(NL)翻译成诸如 Lean 之类的形式语言。一旦翻译完成,用形式语言表达的论证就可以很容易地通过机器进行验证。

The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI = ∞).

本文旨在论证为什么这一过程可能无法为原始的自然语言论证提供信心,这是由于在执行语义忠实翻译时存在各种困难。特别是,我们强调了解决数学自然语言文本中歧义的问题——这对于提供语义忠实的翻译是必要的——在可解性复杂性指数(SCI)层级/算术层级中处于任意高位(SCI = ∞)。

Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI = 1). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications’.

因此,通俗地说,提供语义忠实的人工智能自动形式化比任何计算问题(包括停机问题,其 SCI = 1)都要困难。为了证明这一结果的影响,我们提供了几个在实践中人工智能将自然语言陈述和证明错误翻译成 Lean 的例子,导致自然语言证明与其 Lean “验证”之间出现不匹配。

These include OpenAI’s announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.

这些例子包括 OpenAI 宣布的纳维-斯托克斯证明。特别是,我们展示了形式化的 Lean 证明并不对应于纳维-斯托克斯方程解的爆破的自然语言证明。