katamari architecture

katamari architecture

katamari architecture 2026-09-23 :: tags: #llm #polemic #software

or, how to build a star out of mud 或者,如何用泥巴堆出一颗恒星

User marginalia[1] on lobste.rs compared the output of poorly-directed software development LLM agents to the “katamari” from Katamari Damacy: a video game in which you roll a clump of miscellaneous objects around, sticking everything you can find to the outside. It’s an apt comparison; agentic development tends towards feature addition without attention to composition - they take the shortest path to accomplish the prompt. It has all the worst qualities of an underpaid and short-on-time human engineer. Lobste.rs 上的用户旁注[1]将缺乏有效引导的软件开发 LLM 智能体的产出,比作《块魂》(Katamari Damacy)中的“块魂”:在那款游戏中,你需要滚动一个杂物团,将所见的一切都粘在外面。这是一个贴切的类比;智能体开发往往倾向于在不考虑结构的情况下盲目添加功能——它们总是走最短路径来完成提示词的要求。这具备了一个薪水过低且时间紧迫的人类工程师所拥有的所有最糟糕的特质。

However, I think the analogy does a disservice to Katamari! The spirit clod is carefully crafted (algorithmically, sure, but the algorithm was carefully crafted) to appear thrown-together and unplanned, but the method by which they made the ball naturally round has enough complexity that they thought it warranted a patent. By contrast, LLMs don’t have the sense required to determine how to properly add whatever new feature their prompter desires, it’s just stuck on at random (modulo a probability distribution). Despite that, katamari architecture is a catchy-enough buzzword that I hope it “sticks” around! 然而,我认为这个类比对《块魂》有些不公!那个“灵魂之团”是经过精心设计的(当然是通过算法,但算法本身是精心设计的),看起来像是随意拼凑且毫无计划,但其让球体自然变圆的方法具有足够的复杂性,以至于开发者认为它值得申请专利。相比之下,LLM 并没有足够的判断力来决定如何恰当地添加提示者所期望的新功能,它们只是随机地将其粘上去(基于概率分布)。尽管如此,“块魂架构”(katamari architecture)是一个足够朗朗上口的流行词,我希望它能“粘”住并流行起来!

Katamari architecture is here, but it’s not a hopeless problem. Let’s see if we can learn from prior work. The obvious comparison is the Big Ball of Mud architecture. Foote and Yoder argue that several “forces […] conspire” to create BBoMs - time, cost, experience, skill, visibility, complexity, and scale. Considering driving forces is a good way to understand a phenomenon: steelman Chesterton before slandering his fence. “块魂架构”已经出现,但这并非无药可救。让我们看看能否从前人的工作中汲取经验。最明显的对比是“大泥球”(Big Ball of Mud)架构。Foote 和 Yoder 认为,多种“力量……共同作用”导致了“大泥球”的产生——包括时间、成本、经验、技能、可见性、复杂性和规模。考虑驱动力是理解现象的好方法:在诽谤切斯特顿的栅栏之前,先要理解其背后的逻辑(译注:指切斯特顿的栅栏原则,即在拆除栅栏前先弄清其存在的原因)。

LLMs of course run roughshod over the metaphor, since they send fences zipping across vast expanses for no intelligible reason, moving them around at random as they sycophantically make bullshit replies to your incredulous questions. But I digress; can you build a katamari from the same impulses as mudballs? 当然,LLM 粗暴地践踏了这个隐喻,因为它们会毫无理由地让栅栏在广阔的空间中飞速穿梭,在谄媚地回答你那些难以置信的问题时,随意地移动它们。但我跑题了;你能否用与“泥球”相同的驱动力来构建一个“块魂”?

Time: LLMs have plenty of time. Ostensibly they can work far faster than any human developer, 24/7 and over weekends and holidays. Time is not the issue. 时间: LLM 有充足的时间。表面上看,它们的工作速度远超任何人类开发者,可以 24/7 全天候工作,包括周末和节假日。时间不是问题。

Cost: Many companies have unlimited token budgets, whence tokenmaxing. If LLMs were any good at architecture, cost would be no object. The entire pitch of LLMs is that they’re cheaper than humans to do the same work; cost isn’t the issue. 成本: 许多公司拥有无限的 Token 预算,从而导致了“Token 最大化”。如果 LLM 在架构方面表现出色,成本根本不是问题。LLM 的核心卖点就是它们比人类完成同样的工作更便宜;成本不是问题。

Experience: Frontier LLMs are trained on (more-or-less) the sum total output of all software ever written, along with all books, blogs, and forum posts ever written. There is no architectural process that LLMs are unfamiliar with, even though the nature of the beast indicates that LLMs don’t bring up anything besides the middle-of-the-road, lowest-common-denominator ideas without being, uh, prompted. Experience isn’t the issue. 经验: 前沿 LLM 接受了(几乎)所有已编写软件的总和,以及所有书籍、博客和论坛帖子的训练。没有任何架构流程是 LLM 不熟悉的,尽管这种“野兽”的本质表明,如果不经过“提示”,它们只会提出平庸的、最低共同标准的想法。经验不是问题。

Skill: Debatable. LLMs exhibit inhumanly “spiky”[2] intelligence, similar to other automated systems. They’re impossibly good at some tasks (underspecified search through a large corpus of text) and hilariously bad at others (any number of publicised LLM epic fails, e.g. strawberry syndrome, car-wash transportation, considering a vending machine metaphysically impossible). As anyone who has to review LLM output for a living knows, the mistakes they make are not the same mistakes a human would make, and they’re much harder to spot. Lack of skill is probably contributing to katamari architecture. 技能: 有争议。LLM 表现出非人的“尖峰式”[2]智能,类似于其他自动化系统。它们在某些任务上好得不可思议(在大型文本语料库中进行模糊搜索),而在其他任务上又极其糟糕(许多公开的 LLM 史诗级失败案例,例如“草莓问题”、洗车运输问题、认为自动售货机在形而上学上是不可能的)。正如任何以审查 LLM 产出为生的人所知,它们犯的错误与人类不同,而且更难被发现。技能匮乏可能正是导致“块魂架构”的原因之一。

Visibility[3]: LLMs excel at generating vast swaths of code as far as the eye can see - which is part of the problem. No one is going to perform a close reading of a +6,000/-400 sloc PR - it’s exhausting, and it’s not going to change anything. The more LLM-generated a codebase, the worse visibility becomes. Andrej Karpathy doesn’t even read the code anymore. As Foote and Yoder put it, “[i]f the system works, and it can be shipped, who cares what it looks like on the inside?” That’s the mantra of the modern vibecoder to a T. Just surrender your cognition, embrace the clod! 可见性[3]: LLM 擅长生成一眼望不到头的海量代码——这正是问题的一部分。没有人会去仔细阅读一个增加 6000 行、删除 400 行代码的 PR——这太累人了,而且也改变不了什么。代码库中 LLM 生成的内容越多,可见性就越差。Andrej Karpathy 甚至都不再阅读代码了。正如 Foote 和 Yoder 所言:“如果系统能运行,且能发布,谁在乎它内部长什么样呢?”这正是现代“氛围编程者”(vibecoder)的座右铭。放弃你的认知吧,拥抱那个泥团!

Complexity: Most software has quite little essential complexity, and at this scale Conway’s law isn’t relevant. Complexity may explain some BBoMs, but not the LLM-generated katamaris. I imagine a plurality of truly complex domains still have mostly hand-authored software, though it’s a downward spiral at this point. 复杂性: 大多数软件几乎没有本质上的复杂性,在这种规模下,康威定律并不适用。复杂性或许能解释一些“大泥球”,但无法解释 LLM 生成的“块魂”。我想,许多真正复杂的领域仍然主要由人工编写软件,尽管目前正处于螺旋式下降中。

Change: Human effort is a natural brake on the pace of change. Automatic coding accelerates change; LLMs make implementing a new change as easy as requesting it (with some rather heavy asterisks). Unlike human change, though, LLM-effected change tends to agglomerate onto the existing architecture (such as it is) rather than cut through the heart. A heavily LLM-affected katamari codebase often has a well-designed (human-designed, typically) core, obscured by layers of stylized household objects. An agent may decide to redesign the entire system, but rarely does it have the wherewithal; a redesign or rewrite is rightly feared by software artisans as complex, painful, and interminable. LLMs going for a redesign on their own are doomed to failure. 变更: 人类的工作量是变更速度的天然刹车。自动编码加速了变更;LLM 让实现新变更变得像提出请求一样简单(尽管伴随着一些沉重的代价)。然而,与人类的变更不同,LLM 带来的变更往往倾向于堆积在现有架构(无论它是什么样)之上,而不是深入核心。一个深受 LLM 影响的“块魂”代码库通常拥有一个设计良好的(通常是人类设计的)核心,但被层层风格化的“家居杂物”所掩盖。智能体可能会决定重新设计整个系统,但它很少有能力做到;重新设计或重写被软件工匠们视为复杂、痛苦且无止境的过程,这种恐惧是合理的。LLM 自行进行的重构注定会失败。

Scale: Foote and Yoder’s point (unless I’ve misunderstood - the original is a little unclear to me) is that otherwise-skilled designers struggle to find elegance when faced with massive projects. I don’t find this applies to agents all that well, they have mediocre performance even in the small. They don’t get any better at working at a large scale, though. 规模: Foote 和 Yoder 的观点(除非我理解错了——原文对我来说有点模糊)是,即使是熟练的设计师在面对大型项目时,也难以保持优雅。我觉得这并不完全适用于智能体,它们即使在小规模任务中表现也平平。不过,它们在处理大规模任务时也不会变得更好。

Context compacted 语境压缩

The main contributors from this list are skill, visibility, and change. LLMs are simply not very good at writing code relative to a skilled human, they’re a whole new level of invisible code, and they axe the aerodynamic drag that keeps otherwise-muddy projects from collapsing into slop, ten 15,000 sloc PRs at a time. The visibility one is getting stuck in my craw. It wasn’t good enough that users couldn’t see the horrible mud inside the application, now not even the developers are reading the code? we’re so cooked chat what is to be done? LLMs are the accelerationist dream re 这份清单中的主要贡献因素是技能、可见性与变更。相对于熟练的人类,LLM 在编写代码方面确实不够出色,它们带来了全新的“不可见代码”层级,并且消除了那些原本能防止泥泞项目坍塌成一团糟的“空气阻力”,一次性抛出十个 15000 行代码的 PR。可见性问题让我如鲠在喉。用户看不见应用程序内部可怕的泥潭还不够,现在连开发者都不读代码了吗?我们彻底完了,伙计们,该怎么办?LLM 是加速主义者的梦想……