Why isn't the industry freaking out about DeepSeek 4.1 Flash?
I have been using DeepSeek 4.1 Flash for about a month, heavily, across a dozen projects. It is super capable, and orders of magnitude cheaper than the “frontier” models. When I’m mid-session, if I don’t look at the model name, I honestly could not tell you if I’m using DeepSeek or Opus. Whether it’s our conversations, the work, or the speed, I don’t notice a difference. I don’t care that there is no 4.1 “Pro”. I treat this like a frontier model because it behaves like one. I’m coming at this from my subjective usage experience, but you can see more complete benchmarks here if that floats your boat.
我使用 DeepSeek 4.1 Flash 已经大约一个月了,在十几个项目中进行了高强度使用。它的能力非常强,而且比那些“前沿”模型便宜了几个数量级。当我在使用过程中,如果不去看模型名称,我真的无法分辨自己是在用 DeepSeek 还是 Opus。无论是我们的对话、工作内容还是响应速度,我都感觉不到差异。我不在乎有没有 4.1 “Pro”版本。我把它当作前沿模型来对待,因为它表现得就像一个前沿模型。我这是基于个人的主观使用体验,如果你感兴趣,可以在这里查看更完整的基准测试。
So why aren’t the frontier labs freaking out right now? China is going to eat their lunch. They may be a month or two behind Anthropic/OpenAI, but these distilled Chinese models can handle the same workloads.
那么,为什么前沿实验室现在没有感到恐慌呢?中国正在蚕食他们的市场份额。他们可能比 Anthropic 或 OpenAI 落后一两个月,但这些经过蒸馏的中国模型完全能够处理同样的任务负载。
Sure, they stole Claude’s training, and Anthropic stole it from other people. I’m not getting into the whole who-owns-whose-data debate, because most developers aren’t thinking like that. They’re just trying to get the most bang for their buck.
当然,他们窃取了 Claude 的训练数据,而 Anthropic 也是从别人那里窃取的。我不想卷入关于谁拥有谁的数据的争论,因为大多数开发者根本不会那样想。他们只是想追求性价比最大化。
Today’s models are now good enough for high-quality unattended tasks. Chasing the latest and greatest is silly. It is fun to see the new Fable capabilities, but the tasks we throw at them are usually ridiculous (maybe even insulting) if you believe in LLM sentience. It’s like asking a math PhD to organize the files on your desktop.
如今的模型已经足以胜任高质量的无人值守任务。盲目追求最新最强是愚蠢的。看到新的 Fable 功能确实很有趣,但如果我们相信大语言模型具有感知能力,那么我们交给它们的任务通常是荒谬的(甚至可能是侮辱性的)。这就像是让一位数学博士去整理你桌面上的文件一样。
With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited. This has completely changed my way of developing. There is no shame now in spinning up mindless tasks, or exploratory UI monkey testing. And sure, go ahead and reorganize your desktop files. That will cost $0.003 instead of $1. I have rarely exceeded $1 in expected costs in a session. I try to keep my sessions tight, but sometimes they run for most of a day.
通过我每月 10 美元的 OpenCode Go 订阅,DeepSeek 基本上是无限使用的。这彻底改变了我的开发方式。现在,进行一些无脑任务或探索性的 UI 猴子测试已经没什么好羞愧的了。当然,尽管去整理你的桌面文件吧。那只会花费 0.003 美元,而不是 1 美元。我单次会话的预期成本很少超过 1 美元。我尽量保持会话简洁,但有时它们会持续运行大半天。
I even lean on 4.1 Flash for complex planning and research. For occasional critical tasks, I sometimes pull in Opus 5.5 to do a final code review, which will catch a few edge cases. Then I have DeepSeek execute the fixes. Even when I call up Opus or GLM (which seems to be drinking the same Chinese Kool-Aid as DeepSeek), it’s less about quality and capabilities and more about getting new eyes on a problem.
我甚至会依赖 4.1 Flash 进行复杂的规划和研究。对于偶尔的关键任务,我有时会引入 Opus 5.5 进行最终的代码审查,它能捕捉到一些边缘情况。然后我再让 DeepSeek 执行修复。即使当我调用 Opus 或 GLM(它似乎也喝了和 DeepSeek 一样的中国“酷爱饮料”)时,这更多是为了从不同角度审视问题,而不是为了追求质量和能力。
DeepSeek shrank the KV cache by roughly 437x compared to their V1 model. Holding that cache in GPU memory is one of the biggest costs of running long coding sessions. That’s how my all-day sessions stay under a dollar.
与 V1 模型相比,DeepSeek 将 KV 缓存缩小了大约 437 倍。将该缓存保存在 GPU 内存中是运行长时间编码会话的最大成本之一。这就是我全天候的会话成本能保持在 1 美元以下的原因。
It must be better for the environment too. Using Claude almost feels wasteful, and not just on cost: its caching means DeepSeek must be using less water and electricity. This cache magic is also how Opus 5.5 quietly got its own efficiency boost.
这对环境肯定也更好。使用 Claude 几乎让人感到浪费,不仅仅是在成本上:它的缓存机制意味着 DeepSeek 必然消耗更少的水和电。这种缓存魔法也是 Opus 5.5 能够悄然提升效率的原因。
Yes, I have frontier subscriptions. My work provides Claude, Cursor, and others. I’m not nickel-and-diming here. I’m thinking more about long-term planning, sustainability, and democratizing access to high intelligence. This is a game changer.
是的,我有前沿模型的订阅。我的工作提供了 Claude、Cursor 等工具。我不是在这里斤斤计较。我更多是在考虑长期规划、可持续性以及让高智能访问变得民主化。这是一个游戏规则的改变者。
These wins are lost on the tech industry, which thinks that if you’re not paying top dollar, it’s not worth it. FAANG wants to spend the most money for the highest intelligence. Forget it if it’s unethical, expensive, or bad for the environment or the economy. This is dog-eat-dog capitalism. This leads to people having crazy setups to load balance a dozen Claude Max subs, and complaining when they can’t get more.
科技行业忽略了这些胜利,他们认为如果不花大价钱,就不值得。FAANG(科技巨头)想要花最多的钱来获取最高的智能。如果这不道德、昂贵,或者对环境或经济有害,他们根本不在乎。这就是残酷的资本主义。这导致人们为了负载均衡十几个 Claude Max 订阅而搞出疯狂的配置,并在无法获得更多配额时抱怨。
And to the self-hosters out there, the economics of 4.1 Flash mean self-hosting is not worth it. If saving money is your goal, you will never recoup the costs. But if your concern is privacy, just wait. These cache optimizations are coming to you, and this cache magic will soon run entirely locally. Even now, 4.1 Flash is technically self-hostable, even if not practically so. Any day now.
对于那些自托管用户,4.1 Flash 的经济性意味着自托管并不划算。如果你的目标是省钱,你永远无法收回成本。但如果你关心的是隐私,那就等等吧。这些缓存优化技术即将到来,这种缓存魔法很快就能完全在本地运行。即使是现在,4.1 Flash 在技术上也是可以自托管的,尽管在实践中并不现实。指日可待。