Why does Opus 5 feel worse to work with?
Why does Opus 5 feel worse to work with?
Why does Opus 5 feel worse to work with? 为什么 Opus 5 用起来感觉更差了?
In my opinion and that of the colleagues I’ve spoken with, working with Opus 5 feels like a downgrade compared to Opus 4.7, Opus 4.8, and Fable. I’m not claiming a step backwards in capabilities – it is a more capable model than Opus 4.7 and Opus 4.8 and even rivals Fable in benchmarks, yet these other models feel better to work with. 在我以及我交流过的同事看来,与 Opus 4.7、Opus 4.8 和 Fable 相比,使用 Opus 5 感觉像是一种倒退。我并不是说它的能力退步了——它确实比 Opus 4.7 和 Opus 4.8 更强大,甚至在基准测试中能与 Fable 媲美,但其他模型用起来确实感觉更好。
I believe this is because they: stop and ask questions if my intent was unclear, don’t make assumptions without checking, and don’t reinterpret or update my plans without asking. Because of this, they don’t require the careful babysitting that Opus 5 does. 我认为原因在于:当我的意图不明确时,它们会停下来提问;它们不会在未经确认的情况下做出假设;也不会在未询问的情况下擅自重新解读或更新我的计划。正因如此,它们不需要像 Opus 5 那样需要时刻小心翼翼地“照看”。
Baseless speculation 毫无根据的推测
I suspect this is the result of two compounding forces at Anthropic, and in current frontier labs in general. First, the desire to create a self-improving AI that is capable of recursively bootstrapping itself to AGI/ASI. Second, the pressure to score highly on benchmarks. 我怀疑这是 Anthropic 以及当前各大前沿实验室中两种力量共同作用的结果。首先,是创造一种能够递归式自我引导至 AGI/ASI(通用人工智能/超级人工智能)的自我改进型 AI 的愿望;其次,是在基准测试中获得高分的压力。
Although it’s an open secret that many benchmark tasks are ill-defined, unfair, hackable, or otherwise broken, a good benchmark task is self-contained. It can be solved. It doesn’t require hints, reading the task creator’s mind, or outside information to pass. That doesn’t mean a good task can only have one correct answer, just that it should score all unambiguously correct answers equally. 尽管许多基准测试任务定义不清、不公平、可被破解或存在缺陷已是公开的秘密,但一个好的基准测试任务应该是自洽的。它是可以被解决的。它不需要提示、不需要揣摩出题者的心思,也不需要外部信息就能通过。这并不意味着一个好的任务只能有一个正确答案,而是指它应该对所有明确正确的答案给予同等的评分。
Selecting for models that do well on benchmarks (and indeed training for them or on RLVR tasks in general) inherently selects for models that make bold, usually-correct assumptions in the face of ambiguity. It penalizes models with a tendency to stop and ask for clarification or direction. 筛选在基准测试中表现良好的模型(实际上是针对这些测试或 RLVR 任务进行训练),本质上是在筛选那些在面对模糊性时敢于做出大胆且通常正确假设的模型。它会惩罚那些倾向于停下来寻求澄清或指导的模型。
Unfortunately, that’s exactly what most of us want from a coding agent. Try as you might, it’s nearly impossible to get the entirety of the context, intentions, business implications, budget constraints, and what-have-you written down and accessible to a coding agent. There will invariably be ambiguity and choices to be made, and it is nice to know that an agent will stop and ask when needed. 不幸的是,这恰恰是我们大多数人对编程智能体(coding agent)的需求。无论你如何努力,几乎不可能将所有的背景信息、意图、业务影响、预算限制等方方面面都写下来并提供给智能体。模糊性和需要抉择的情况总是不可避免的,如果智能体能在需要时停下来询问,那会让人感到非常安心。
Real life just isn’t a benchmark. There isn’t a guaranteed right answer to every question, nor even a set of right answers, and with real-life consequences on the line, I do not want an agent taking its best guess! 现实生活毕竟不是基准测试。并非每个问题都有保证正确的答案,甚至连一组正确答案都没有。考虑到现实生活中的后果,我不希望智能体仅仅是凭猜测行事!