MatrAIx: Simulating the World with 8.3 Billion Persona Agents

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

MatrAIx:用 83 亿个角色智能体模拟世界

Computer Science > Artificial Intelligence arXiv:2608.04205 (cs) [Submitted on 4 Aug 2026] 计算机科学 > 人工智能 arXiv:2608.04205 (cs) [2026 年 8 月 4 日提交]

Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. 摘要: 对人工智能系统和数字产品进行人工评估既昂贵、缓慢,又难以扩展。离线评估虽然更具可扩展性,但往往忽略了人类的多样性和交互行为。因此,我们推出了 MatrAIx,这是一个人口规模的模拟用户评估基础设施,用于测试具有异构用户的人工智能系统和数字产品。

MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. MatrAIx 包含三个核心组件:首先,Persona 8B 包含 83 亿条角色记录,由 1,290 个分类维度表示。这些记录要么是从保留相关属性的依赖图中采样而来,要么源自人工编写的档案。我们发布了一个经过质量过滤的约 100 万个角色的核心集,其中包括 599,847 条基于人类真实数据的记录和 400,000 条合成记录。

Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. 其次,MatrAIx Playground 提供了四种环境,供不同用户评估数字产品并与之交互:调查、AI 聊天机器人、网页和应用程序。第三,MatrAIx 提供了涵盖 25 个以上领域的 1,010 项应用任务,包括商业、软件、金融和医疗保健。

We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. 我们在八项代表性任务中进行了 18,189 次评估试验。角色智能体由三种大语言模型驱动:Claude Opus 4.8、GPT 5.5 和 Claude Haiku 4.5。所得反馈捕捉了决策和偏好如何随角色背景而变化,包括价格上涨后的犹豫、AI 助手失败后继续使用的意愿以及对延迟的容忍度。

We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users. 我们进行了两项主要验证研究:首先,一项 400 次试验的对照研究评估了角色在十种行为属性和所有四种环境中的依从性。在 366 次试验(91.5%)中,声明的行为得到了表达或正确的抑制。其次,人类和 LLM 评委评估了基于人类真实数据的角色的提取质量。总的来说,MatrAIx 为评估具有多样化模拟人类用户的人工智能系统和数字产品提供了一个端到端的基础设施。