Coding is not solved

Coding is not solved

编程问题尚未解决

Disclaimer: you are about to read a lot of opinions, many of them have references but some are the result of my own experience building with AI and building AI systems in the past 4 years. Regardless, beware of the straw-man fallacy: just because one argument doesn’t map to your belief system, it doesn’t mean the rest are invalid. I should also say upfront that I’m not anti-AI. If you’ve been following my work, you know that I was an early adopter of not only using LLM-powered coding tools, but building my own harness, teaching these topics and building LLM-powered products. It’s not about fear of AI but rather challenging the brain-dead narrative that asserts “coding is solved” and engineering is about “taste” now.

免责声明:你即将读到许多观点,其中许多有据可查,但也有一些是我过去 4 年在使用 AI 构建系统过程中的个人经验总结。无论如何,请警惕“稻草人谬误”:仅仅因为某个论点不符合你的信仰体系,并不意味着其余论点无效。我还要事先声明,我并不反 AI。如果你一直关注我的工作,就会知道我不仅是 LLM 编程工具的早期采用者,还构建了自己的测试框架,教授相关课题并开发过基于 LLM 的产品。我并非恐惧 AI,而是要挑战那种认为“编程已解决”、工程现在只关乎“品味”的愚蠢论调。

People who claim “LLMs can write decent code” don’t understand how code works. Sure, creation is much cheaper, but anyone who has run software in production at scale knows that maintenance, reliability, security, scalability, etc. is the majority of the cost. These are commonly known as NFR (non-functional requirements).

那些声称“大语言模型(LLM)能写出像样代码”的人,根本不理解代码是如何运作的。诚然,创建代码的成本确实降低了,但任何在大规模生产环境中运行过软件的人都知道,维护、可靠性、安全性、可扩展性等才是成本的大头。这些通常被称为非功能性需求(NFR)。

In my experience even the Functional Requirements (what the code is supposed to do) is NOT a solved problem yet. There’s a bit of Dunning-Kruger effect at place where the people who don’t read the output are more confident in it.

根据我的经验,即使是功能性需求(代码应该做什么)也远未成为一个“已解决”的问题。这里存在一种达克效应(Dunning-Kruger effect):那些不去阅读代码输出结果的人,反而对 AI 生成的代码更有信心。

As a veteran developer holding 2 engineering degrees (hardware and systems engineering), I can list 3 types of products that do not strictly require reading the code:

  1. Personal software: scratching an itch, automation, DIY patches, etc.
  2. POC (proof of concept): demonstrating technical feasibility and product viability
  3. Weaponized AI: acknowledge the risk and deliberately point it at a target to cause harm

作为一名拥有两个工程学位(硬件和系统工程)的资深开发者,我可以列举出 3 类不需要严格审查代码的产品:

  1. 个人软件:解决个人需求、自动化脚本、DIY 补丁等。
  2. POC(概念验证):用于演示技术可行性和产品生存能力。
  3. 武器化 AI:承认风险并故意将其指向目标以造成伤害。

Notice the commonality: the first 2 have high risk tolerance while the last one weaponizes the inherent risk. Most software that requires hiring and paying software engineers has low risk tolerance: healthcare, finance, automotive, defense, power plants, aviation, manufacturing… wherever a mistake can cost money, lives or legal consequences you need accountability.

注意它们的共同点:前两者具有高风险承受能力,而后者则是将固有的风险武器化。大多数需要聘请软件工程师的软件项目,其风险承受能力都很低:医疗、金融、汽车、国防、发电厂、航空、制造业……在任何一个错误都可能导致金钱损失、生命代价或法律后果的领域,你都需要问责制。

AI cannot be held accountable. It cannot suffer any consequences. The worst thing you can do to AI is to unplug it. And although it mimics human emotions (due to training data), it couldn’t care less. AI doesn’t die either. It cannot suffer a prison sentence or fines. You cannot punish AI, therefore it can never be held accountable.

AI 无法被问责。它不会承担任何后果。你能对 AI 做的最糟糕的事就是拔掉电源。尽管它(由于训练数据)模仿了人类的情感,但它根本不在乎。AI 也不会死亡。它无法被判刑或罚款。你无法惩罚 AI,因此它永远无法承担责任。

You cannot be responsible for what you can’t control either. That understanding is key to reasoning about system behavior and fixing it when the AI inevitably fails.

你也不可能为你无法控制的事物负责。这种理解对于推断系统行为以及在 AI 不可避免地发生故障时进行修复至关重要。

Why coding is NOT solved? Contrary to common narrative, coding is actually one of the last areas for the current generation of LLMs to take over!!! Allow me to elaborate: Coding is about logic. Anyone who has dealt with compiler errors knows that computers don’t give a f*** about how right you think you are. If it’s logically wrong, it doesn’t compile. Even if the syntax is fine, there are runtime errors.

为什么编程问题没有解决?与普遍的论调相反,编程实际上是当前这一代 LLM 最难接管的领域之一!!!请允许我详细说明:编程关乎逻辑。任何处理过编译器错误的人都知道,计算机根本不在乎你认为自己有多正确。如果逻辑错误,它就无法编译。即使语法正确,也可能存在运行时错误。

The reason LLMs are successful in writing code is because we’ve made a feedback loop that feeds the syntax/runtime errors back to the LLM and loops until most errors are solved or hidden.

LLM 在编写代码方面之所以成功,是因为我们建立了一个反馈循环,将语法/运行时错误反馈给 LLM,并不断循环,直到大多数错误被解决或隐藏。

LLMs are stochastic and probabilistic. The only way we could even get remotely close to making them logical is to wrap them in traditional code (known as harness), run tests, and a bunch of other techniques (e.g. CoT) but the core issue remains: LLMs struggle with logic and volume (the larger the input and the more the context window is used, the less accurate they get).

LLM 是随机且概率性的。我们唯一能让它们接近逻辑性的方法,就是用传统代码(即测试框架)包裹它们,运行测试,并结合其他技术(如思维链 CoT),但核心问题依然存在:LLM 在逻辑和处理规模上表现吃力(输入越大、上下文窗口使用越多,它们的准确性就越低)。

I’m not saying LLMs cannot generate code or maintain existing code bases. They have their utility as a tool and their capabilities are increasing in an S-curve. There is a point of diminishing return where more expensive models aren’t necessarily more productive at the rate of the price increase.

我并不是说 LLM 不能生成代码或维护现有的代码库。作为一种工具,它们有其用途,且其能力正呈 S 曲线增长。但存在一个边际收益递减点,即更昂贵的模型并不一定能带来与价格涨幅成正比的生产力提升。

Those who claim LLM-generated software is good enough:

  • Haven’t written code in ages
  • Cannot spot if their code figuratively had 6 fingers!
  • Have a low bar for what good looks like
  • Don’t care about quality or NFR
  • Have difficulty understanding an S-curve
  • Are honest: AI genuinely writes better code than them

那些声称 AI 生成的软件“足够好”的人:

  • 很久没写过代码了
  • 甚至看不出代码里是否有“六根手指”(指明显的逻辑错误)!
  • 对“好”的标准很低
  • 不关心质量或非功能性需求(NFR)
  • 难以理解 S 曲线
  • 或者很诚实:AI 确实写得比他们好

But to go ahead and extrapolate that to an entire professional industry requires a level of brain-dead thinking that’s only present in people who spend too much time with sycophantic AI.

但如果将这种现象推而广之到整个专业行业,则需要一种极其愚蠢的思维,这种思维只会出现在那些花太多时间与只会阿谀奉承的 AI 相处的人身上。