Small Models Have Arrived

Small Models Have Arrived

小型模型时代已至

AUG 26, 2026 2026年8月26日

For the past few weeks, I’ve been playing with gpt-5.6-luna. It is shockingly capable, fast, and smart. I regularly see it do ~100 tps, and rip around my codebase, email, and knowledge base. Of course, the biggest thing with luna is the cost. I’ve tried running some fairly complicated research threads, and it’s pretty tough to run up a large bill. Even having it search across thousands of emails, I end up with an API cost in the tens of cents. Courtesy of artificialanalysis.ai With GLM 5.3, we even have a new option at the Pareto frontier. 在过去的几周里,我一直在使用 gpt-5.6-luna。它的能力、速度和智能化程度令人震惊。我经常看到它的处理速度达到每秒约 100 个 token(tps),并且能高效地处理我的代码库、电子邮件和知识库。当然,luna 最显著的优势在于成本。我尝试过运行一些相当复杂的调研任务,但很难产生高额账单。即使让它搜索数千封邮件,最终的 API 成本也仅在几十分钱左右(数据来源:artificialanalysis.ai)。随着 GLM 5.3 的发布,我们在帕累托前沿(Pareto frontier)又多了一个新选择。

When doing coding work, I almost always reach for the most expensive and capable models (Fable 5, 5.6 Sol). So it’s been easy to miss the progress the small fast models have made. One thing a few investors I’ve talked with have mentioned: “It’s weird we’re not seeing more consumer AI companies. Why is that?” There’s a straightforward answer: token costs. 在进行编程工作时,我几乎总是倾向于使用最昂贵、能力最强的模型(如 Fable 5, 5.6 Sol)。因此,我很容易忽略小型快速模型所取得的进步。我与几位投资者交流时,他们都提到:“奇怪的是,我们并没有看到更多的消费级 AI 公司。这是为什么?”答案很简单:token 成本。

In the times before AI, the playbook for big consumer apps looked like this… create some sort of compelling website which is fairly cheap to run; attract a bunch of users (typically with some virality); raise money, scale to more users; create an ads marketplace. This roughly describes most of the big consumer companies (Google, Facebook, Snapchat, etc.). 在 AI 时代之前,大型消费级应用的成功路径通常是这样的:创建一个引人注目的网站,且运营成本相对低廉;吸引大量用户(通常通过病毒式传播);融资并扩展用户规模;建立广告市场。这大致描述了大多数大型消费级公司(如谷歌、Facebook、Snapchat 等)的发展历程。

But what if you want to add AI to your product? Well, now you have some real inference costs on every request! Suddenly the amount of capital required increases dramatically. A pet eval of mine is to build a daily news site, personalized to me: research @calvinfo on the internet; figure out what news they might like; build a micro-site with today’s top stories, personalized for them; search hn, reddit, twitter, etc. With the previous generation of models (Sonnet class), you’d spend ~$1 to get anywhere. Charging $30/mo is untenable for a consumer app. There’s obviously a lot we can optimize here, but if you’re charging what the WSJ or The Economist charges, you’d better be delivering similar value. But looking at luna, the results are pretty decent, and the average cost is ~$0.10. Now we’re talking! 但如果你想在产品中加入 AI 呢?现在,每一次请求都会产生实实在在的推理成本!所需的资金量会瞬间大幅增加。我曾做过一个个人评估项目:建立一个为我量身定制的每日新闻网站。它需要:在互联网上调研 @calvinfo;找出他们可能感兴趣的新闻;构建一个包含今日头条的微型网站,并为他们进行个性化推荐;搜索 HN、Reddit、Twitter 等平台。使用上一代模型(Sonnet 级别),完成这些任务大约需要 1 美元。对于消费级应用来说,每月收取 30 美元的费用是不可持续的。显然,我们可以在这里进行大量优化,但如果你收取的费用与《华尔街日报》或《经济学人》相当,你就必须提供同等的价值。然而,看看 luna,它的结果相当不错,平均成本仅为 0.10 美元左右。这才是我们想要的!

Where I think this gets even more interesting is in the world of business. My Segment co-founder Peter and I were recently comparing notes on a hike. Across his various startups, Peter has seen two kinds of work: the “IQ 180” work (some mad scientist genius type comes up with some crazy solution you’ve never thought of) and the “token spewer” work (being ultra responsive, pushing the ball forward across dozens of different fronts). 我认为这在商业领域变得更加有趣。最近在徒步时,我和 Segment 的联合创始人 Peter 交流了心得。在他的各种创业项目中,Peter 观察到两种工作类型:一种是“IQ 180”型工作(某种疯狂的科学家天才提出了你从未想过的绝妙方案);另一种是“token 喷射”型工作(保持极高的响应速度,在几十个不同的战线上推动项目进展)。

Peter runs multiple companies. Beyond Segment, he’s raised $100m+ for Charm Industrial, and just recently closed a Series A for Revoy. He’s incredibly organized and efficient with his time. And yet, Peter mentioned that ~95% of the work he does falls into bucket 2. It’s hopping on calls. Nudging people. Blocking and tackling. To be clear, Peter says his companies would be dead-in-the-water today without an IQ 180 technical mind solving the deep problems. Just that most of his work falls in bucket 2. Peter 管理着多家公司。除了 Segment,他还为 Charm Industrial 筹集了超过 1 亿美元的资金,并刚刚完成了 Revoy 的 A 轮融资。他非常善于组织和利用时间。然而,Peter 提到他所做的工作中约 95% 都属于第二类。比如参加电话会议、督促他人、解决各种琐碎的障碍。需要明确的是,Peter 表示如果没有“IQ 180”的技术大脑来解决深层次问题,他的公司今天就会陷入瘫痪。只是他大部分的工作确实属于第二类。

I think demand for “frontier-level” models is going to keep compounding. Especially for fields that require novel breakthroughs or discovery (engineering, hard science, model training). But I also think the demand for “fast/cheap/good-enough” models is just about to take off. Think of the people you interact with on a daily basis: coworkers, vendors, and customers. Nine times out of ten, you want someone who is super responsive, and just handles things for you. Most of the “human tokens” at companies today are spent this way — hiring skews heavily toward the fast/cheap/good-enough archetype. 我认为对“前沿级”模型的需求将持续增长,特别是在需要创新突破或发现的领域(如工程、硬科学、模型训练)。但我同样认为,对“快速/廉价/足够好”模型的需求即将爆发。想想你每天接触的人:同事、供应商和客户。十有八九,你希望对方反应迅速,能帮你处理好事情。如今公司里大部分的“人类 token”都花在了这方面——招聘倾向于那些快速、廉价且能胜任工作的类型。

There’s a lot of work that needs to happen to make fast/cheap/good-enough models a reality for business. New harnesses, prompt injection safety, roles, and permissions. But I’m confident we’ll figure that out. If you’re also experimenting with making small models useful, please drop me a line. 要让“快速/廉价/足够好”的模型在商业中成为现实,还有很多工作要做。例如新的框架、提示词注入安全、角色和权限管理等。但我相信我们会解决这些问题。如果你也在尝试让小型模型发挥作用,请随时联系我。