Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
为什么更强的模型反而会创造出风险更高的系统:来自金融市场中大语言模型智能体的证据
Abstract: Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes.
摘要: 大语言模型(LLMs)正被大规模部署于各类重要的现实世界系统中,涵盖从金融市场、内容审核到招聘等领域。我们研究发现,提升单个模型的能力反而可能导致系统层面的结果恶化,而非改善。
We hypothesize that shared training and architectures can lead more capable LLMs to behave more similarly, creating correlated actions that do not diversify away. We develop a general framework showing how this correlation creates a non-diversifiable risk floor and test its predictions in their predictions in financial markets using an agent-based simulation with LLM traders of varying general-purpose capability.
我们假设,共享的训练数据和架构会导致能力更强的大语言模型表现出高度相似的行为,从而产生无法通过多样化消除的关联性操作。我们开发了一个通用框架,展示了这种关联性如何形成一个不可分散的风险底线,并利用基于智能体的模拟,通过不同通用能力的大语言模型交易员在金融市场中对该预测进行了测试。
We find that: (1) frontier LLMs exhibit significantly correlated behavior that increases with capability; (2) when their shared reasoning is accurate, increasing agent participation reduces market-level risk; and (3) when agents share a common misinformation environment, the same correlated behavior becomes a liability.
研究发现:(1)前沿大语言模型表现出显著的关联行为,且这种关联性随模型能力的提升而增强;(2)当它们共享的推理逻辑准确时,增加智能体的参与度会降低市场层面的风险;(3)当智能体处于共同的错误信息环境时,同样的关联行为反而会成为一种负债。
Together, these results identify a capability paradox: improving individual models does not necessarily produce better system-level outcomes. Whether the same dynamics arise in other domains is an open empirical question.
综上所述,这些结果揭示了一个“能力悖论”:提升单个模型的能力并不一定能带来更好的系统级结果。这种动态机制是否也会出现在其他领域,仍是一个有待实证的开放性问题。