Some Large Language Models Exhibit Consistent Risk Attitudes

Some Large Language Models Exhibit Consistent Risk Attitudes

部分大型语言模型表现出一致的风险态度

Abstract: As artificial intelligence systems are deployed in open-ended, high-stakes settings, a critical dimension remains unmeasured: how perceived risk is translated into action. We test whether large language models (LLMs) exhibit systematic and consistent risk attitudes under uncertainty.

摘要: 随着人工智能系统被部署在开放式、高风险的环境中,一个关键维度尚未得到衡量:感知到的风险如何转化为行动。我们测试了大型语言模型(LLMs)在不确定性下是否表现出系统且一致的风险态度。

We introduce a cross-domain framework that decouples contextual risk belief from categorical decision, and apply it to six representative LLMs and 100 human participants across spatial navigation, clinical triage, and financial allocation tasks. Using regression models, we extract each agents belief-to-decision mapping and quantify risk sensitivity and risk attitude bias.

我们引入了一个跨领域框架,将情境风险信念与分类决策解耦,并将其应用于六个代表性 LLM 和 100 名人类参与者,涵盖空间导航、临床分诊和财务分配任务。通过回归模型,我们提取了每个智能体的“信念到决策”映射,并量化了风险敏感度和风险态度偏差。

We find that most tested LLMs exhibit (i) robust intra-task consistency, indicating stable mappings from contextual belief to risk decision within a fixed task domain; (ii) cross-domain rank-order stability, preserving relative risk posture across tasks; and (iii) a convergence toward a restricted risk-attitude distribution relative to the broader human baseline.

我们发现,大多数受测 LLM 表现出:(i) 稳健的任务内一致性,表明在固定任务领域内,从情境信念到风险决策的映射是稳定的;(ii) 跨领域排序稳定性,在不同任务中保持相对的风险姿态;以及 (iii) 相较于更广泛的人类基准,其风险态度分布趋向于收敛到一个受限的范围内。

These results reveal risk attitude as a stable and previously uncharacterized dimension of LLM behavior, establishing a foundation for evaluating and aligning AI systems in open-ended decision-making and motivating further investigation into the origins of these intrinsic behavioral dispositions.

这些结果揭示了风险态度是 LLM 行为中一个稳定且此前未被描述的维度,为评估和对齐开放式决策中的 AI 系统奠定了基础,并推动了对这些内在行为倾向起源的进一步研究。