CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

CUEing 用户模拟器:用于多轮基准测试的校准用户嵌入

Abstract: Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populations.

摘要: 最近的基准测试依赖于用户模拟器来评估 AI 智能体在多轮交互中的表现。虽然现有的模拟技术在人类风格和行为的表面保真度上表现良好,但要实现生态有效的交互式基准测试,还需要在模拟用户群体与真实用户群体之间,就智能体在何时以及如何失败的问题上保持一致。

We find that existing simulators lack outcome calibration: agreement with observed success rates and failure patterns when real users interact with the same agent. We introduce Calibrated User Embeddings (CUE), a framework that both encodes observed sessions and samples continuous representations, then decodes them into persona commands to steer LLMs to act as user simulators without training.

我们发现现有的模拟器缺乏结果校准:即当真实用户与同一智能体交互时,模拟器无法与观察到的成功率和失败模式保持一致。我们引入了校准用户嵌入(Calibrated User Embeddings, CUE),这是一个既能编码观察到的会话又能采样连续表示的框架,随后将其解码为角色指令,从而在无需训练的情况下引导大语言模型(LLM)充当用户模拟器。

Through this, we evaluate user-conditioned replay of past sessions and aggregate metric agreement when sampling novel personas for the same tasks. On $\tau^2$-Bench, CUEd simulators commit fewer simulator-attributed errors and more faithfully reproduce real-user agent failure modes, aggregate success rates, and outcomes for specific task-user pairs than other persona-based simulation methods.

通过这种方式,我们评估了基于用户条件的过去会话重放,以及在为相同任务采样新角色时的聚合指标一致性。在 $\tau^2$-Bench 测试中,与其他基于角色的模拟方法相比,CUEd 模拟器产生的模拟器归因错误更少,并且能更忠实地重现真实用户的智能体失败模式、聚合成功率以及特定任务-用户对的结果。

These gains coexist with competitive user fidelity as measured using metrics established in prior work. After being fit to mostly customer support interactions, the same CUEd simulators generalize to document creation, math tutoring, and casual conversation, and remain effective across different simulator LLMs without CUE retraining.

这些提升与先前研究中建立的指标所衡量的竞争性用户保真度并存。在主要针对客户支持交互进行适配后,相同的 CUEd 模拟器能够泛化到文档创建、数学辅导和日常对话等场景,并且在无需重新训练 CUE 的情况下,在不同的模拟器 LLM 上依然保持有效。