GPU Management: Why Idle GPUs Are the New Grounded Aircraft
GPU Management: Why Idle GPUs Are the New Grounded Aircraft
GPU 管理:为何闲置的 GPU 就像停飞的飞机
Utilization, not intelligence, is the next real constraint in AI. Aviation learned this the hard way. For most of the industry’s history, the number that best predicted whether an airline would survive was how much of the day each aircraft spent on the ground. The reason is structural. An aircraft’s costs accrue by the calendar hour: financing, depreciation, hull insurance, scheduled maintenance, crew contracts. Its revenue accrues only by the flight hour. Every hour spent on the ground shrinks the output side of that equation while the cost side keeps running exactly as before.
利用率,而非智能,是人工智能领域下一个真正的制约因素。航空业曾为此付出了惨痛的代价。在航空业历史的大部分时间里,预测一家航空公司能否生存的最佳指标,是每架飞机每天在地面停留的时间。其原因在于结构性因素。飞机的成本是按日历小时计算的:融资、折旧、机身保险、定期维护和机组人员合同。而其收入仅按飞行小时计算。飞机在地面停留的每一小时,都会缩减等式中的产出端,而成本端却丝毫不减。
Utilization also sits downstream of almost everything else an airline does. Turnaround discipline, network design, maintenance planning, crew rostering, and spare parts availability all eventually show up in that one number, because a broken operation underneath it keeps planes on the ground no matter what else goes right. A bigger fleet still helps. More aircraft means more available capacity, plainly and simply. But two airlines flying comparable fleets on comparable routes can end up with very different economics, and most of that gap traces back to one measurement rather than fleet size.
利用率也是航空公司几乎所有其他运营环节的下游指标。周转纪律、航线网络设计、维护计划、机组排班以及备件可用性,最终都会体现在这个数字上,因为如果底层的运营出现问题,无论其他方面做得多好,飞机都无法起飞。拥有更大的机队确实有帮助。简单来说,更多的飞机意味着更多的可用运力。然而,两家在相似航线上运营相似机队的航空公司,最终的经济效益可能大相径庭,而这种差距大部分归因于这一指标,而非机队规模。
Enterprise AI is running into the same structure, on a different piece of hardware. A GPU accrues cost by the calendar hour too, through financing, depreciation, power, and cooling, whether or not it’s doing anything useful in a given moment. Its output only accrues by the compute hour. More GPUs helps in roughly the way a bigger fleet helps an airline: real capacity, a genuine advantage, and still no guarantee of the result that actually decides who wins. Two companies with comparable GPU budgets increasingly diverge based on how much of that hardware is doing something useful at any given moment, not on how much of it either one owns. That same number, like an airline’s utilization rate, sits downstream of nearly every other infrastructure decision a company makes. Intelligence has carried the industry this far. Utilization is where the next real constraint is forming.
企业级人工智能正面临着同样的结构,只是硬件换成了 GPU。GPU 的成本同样按日历小时计算——通过融资、折旧、电力和冷却费用——无论它在某一时刻是否在执行有用的任务。而其产出仅按计算小时计算。增加 GPU 的作用大致类似于航空公司增加机队规模:它提供了实际的运力,是一种真正的优势,但并不能保证最终的胜出。两家拥有相当 GPU 预算的公司,其表现差异越来越取决于在任何给定时刻有多少硬件在执行有用的任务,而不是各自拥有多少硬件。这个数字,就像航空公司的利用率一样,处于公司几乎所有其他基础设施决策的下游。智能引领行业走到了今天,而利用率则是下一个真正的制约因素所在。
The Bottleneck Moved From Models to Compute
瓶颈从模型转移到了算力
The scarcity didn’t disappear as AI scaled. It moved up the chain, landing on a different resource entirely. The first wave of enterprise AI was won on model quality. Bigger models, trained on more compute, evaluated against tougher benchmarks: parameter count and leaderboard position dominated the conversation, and the race produced models genuinely good enough to run real enterprise workloads. That capability arrives bundled with a dependency, though. Production AI runs on specialized hardware, and today that hardware is almost entirely GPUs. GPUs are expensive, supply constrained, and in demand far beyond what’s available, and this holds even at the very top of the market.
随着人工智能规模的扩大,稀缺性并没有消失。它向上游转移,最终落在了完全不同的资源上。企业级人工智能的第一波浪潮胜在模型质量。更大的模型、基于更多算力的训练、针对更严苛基准的评估:参数规模和排行榜排名主导了当时的讨论,这场竞赛产出了足以运行真实企业工作负载的模型。然而,这种能力伴随着一种依赖性。生产级人工智能运行在专用硬件上,而今天这种硬件几乎完全是 GPU。GPU 价格昂贵、供应受限,且需求远超现有供应,即使在市场最顶端也是如此。
In 2020, Microsoft built OpenAI a dedicated supercomputer: over 10,000 GPUs and 285,000 CPU cores, reported at the time as one of the five largest systems in the world, assembled to train what became GPT-3. At the time, it looked like an almost unimaginable concentration of hardware, the kind of number that made compute look like a solved problem for whoever could get access to it. Six years later, that number reads more like a starting point than a ceiling. By 2026, even the best capitalized labs on the planet were treating compute access as a live strategic constraint rather than a settled one. Anthropic alone was running simultaneous multi-gigawatt commitments across four separate hardware platforms, Amazon, Google, Microsoft, and AMD, layered within months of one another, while Meta signed a comparable multi-gigawatt deal of its own. Spreading commitments across four vendors at once is what compute scarcity looks like when a buyer has effectively unlimited capital and still can’t get enough from any single source.
2020 年,微软为 OpenAI 构建了一台专用超级计算机:拥有超过 10,000 个 GPU 和 285,000 个 CPU 核心,当时被报道为全球五大系统之一,旨在训练后来的 GPT-3。在当时,这看起来是难以想象的硬件集中度,这种规模让算力看起来对于任何能获得访问权限的人来说都已经是一个“已解决的问题”。六年后的今天,这个数字看起来更像是一个起点而非上限。到 2026 年,即使是全球资金最雄厚的实验室,也将算力获取视为一个动态的战略制约因素,而非已解决的问题。仅 Anthropic 一家公司就同时在亚马逊、谷歌、微软和 AMD 四个独立的硬件平台上运行着数吉瓦(multi-gigawatt)的算力承诺,且这些承诺在几个月内层层叠加;与此同时,Meta 也签署了规模相当的数吉瓦协议。当买家拥有近乎无限的资金却仍无法从单一来源获得足够算力时,同时向四家供应商分散承诺,这就是算力稀缺的真实写照。
Six years apart, both events marked the frontier of what a lab needed just to stay competitive. What changed in between has less to do with AI getting more capable, and everything to do with capability no longer being the binding constraint. The same pattern shows up downstream of the labs, in a different form. Enterprises consuming these models through an API run into a pricing problem more than a hardware one. Cost scales linearly with tokens used, and that single fact separates the economics of a proof of concept from the economics of production almost completely. A PoC processing a few thousand requests a month looks affordable. The same workload at production volume can turn into a cost line that never quite clears.
相隔六年,这两个事件都标志着实验室为了保持竞争力所必须达到的前沿水平。期间发生的变化与人工智能能力的提升关系不大,而与“能力不再是核心制约因素”这一事实密切相关。同样的模式也出现在实验室的下游,只是形式不同。通过 API 使用这些模型的企业面临的更多是定价问题,而非硬件问题。成本随 Token 使用量线性增长,这一事实几乎完全将概念验证(PoC)的经济性与生产环境的经济性区分开来。每月处理几千次请求的 PoC 看起来负担得起,但同样的工作负载在生产规模下,可能会变成一笔永远无法完全结清的成本支出。
The alternative gaining ground is straightforward enough: enterprises acquiring their own GPUs and running models locally, trading a variable, linearly scaling cost for a fixed capital one. API cost rises with usage, while owned infrastructure stays close to fixed. Past the breakeven point, the trade reverses. That shift turns the GPU into infrastructure rather than a line item, sized for growth, sized for demand peaks, and therefore sized above what any given week actually needs. Which means the purchase doesn’t close the problem. It opens a new one. The day the cluster comes online, the question stops being “can we get accelerators” and becomes “can we keep them busy,” and only the first question had a procurement team assigned to it. Signing for the hardware is the part with a deadline and an owner. Keeping it off the ground is the part that quietly decides whether the deal was worth signing. These deals describe capacity commitments, not efficiency. How well that capacity gets used is a separate question, owned by different people, measured far less rigorously, and considerably further from being solved.
另一种日益流行的替代方案很简单:企业购买自己的 GPU 并在本地运行模型,用固定的资本支出取代可变的、线性增长的成本。API 成本随使用量上升,而自有基础设施的成本基本保持固定。一旦超过盈亏平衡点,这种权衡就会发生逆转。这种转变将 GPU 从一项“支出项目”变成了“基础设施”,其规模是为增长和需求高峰而设计的,因此其规模往往超过了任何特定周的实际需求。这意味着购买行为并没有解决问题,反而开启了一个新问题。集群上线的那一天,问题不再是“我们能买到加速器吗”,而是“我们能让它们保持忙碌吗”,而只有第一个问题有专门的采购团队负责。签署硬件合同是有截止日期和负责人的,而让它们“不闲置”才是决定这笔交易是否划算的关键。这些交易描述的是运力承诺,而非效率。这些运力被利用得如何是一个独立的问题,由不同的人负责,衡量标准远不严谨,且距离解决还相当遥远。