Fable 5 – Median thinking declined in August
Fable 5 – Median thinking declined in August
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn’t explain why. Measured five different ways, August delivered dramatically fewer thinking tokens than July.
在 Anthropic 将 Fable 5 永久纳入订阅计划后,我注意到其性能出现了大幅下降。模型感觉变“笨”了,但我无法解释原因。通过五种不同的方式进行测量,八月份产生的“思考令牌”(thinking tokens)数量比七月份大幅减少。
This wasn’t a one-time drop. Reasoning fell over the entire period, and fluctuated across multi-day episodes. Some of these fluctuations aligned with specific product announcements and releases. I started to see how the model could feel great one day, and terrible the next.
这并非一次性的下降。推理能力在整个期间持续下滑,并在多日的周期中出现波动。其中一些波动与特定的产品公告和发布时间相吻合。我开始明白为什么模型有时表现极佳,而有时却表现糟糕。
I was consistently using an xhigh or max effort level, but when I looked deeper, I found that most invocations to the model were receiving little to no thinking tokens at all. And when longer thinking runs did happen, they almost never reached published benchmark levels.
我一直使用的是“超高”(xhigh)或“最大努力”(max effort)级别,但深入研究后我发现,大多数对模型的调用几乎没有获得任何思考令牌。即使出现了较长的思考过程,它们也几乎从未达到官方公布的基准水平。
This unfolded across a six-week data capture and analysis odyssey, and led to a number of surprising findings. Next time the model feels dumber, don’t ask if the model was “nerfed”. Ask about the inference regime you were served, instead.
这一发现源于长达六周的数据采集与分析过程,并得出了一些令人惊讶的结论。下次当你觉得模型变“笨”时,不要问模型是否被“削弱”(nerfed)了,而应该问问你所使用的推理机制(inference regime)究竟是什么。
Article: The Inference Gap Access to frontier models no longer guarantees access to the inference regime that can reliably reproduce frontier model capabilities. On July 1st, Anthropic restored availability to Fable 5 after…
文章:推理差距 获得前沿模型的使用权,不再等同于能够获得足以可靠地重现其前沿能力的推理机制。7 月 1 日,Anthropic 在……之后恢复了 Fable 5 的可用性。
alias@loading Genuinely think this should be illegal.
alias@loading 我真心认为这应该是非法的。