Large-Scale ChatBot Validation Through Customer Digital Twin Simulations
Large-Scale ChatBot Validation Through Customer Digital Twin Simulations
通过客户数字孪生模拟进行大规模聊天机器人验证
LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. 基于大语言模型(LLM)的聊天机器人正在改变银行等受监管领域的客户服务,但可扩展且具成本效益的验证手段仍然是实现安全部署的关键障碍。
We present a two-part contribution for large-scale chatbot validation. First, we introduce a methodology for creating high-fidelity synthetic customer agents (SCAs) as digital twins, grounded in real transactional and conversational data, that enables automatic generation and behavioral conditioning to simulate diverse customer profiles and interaction styles. 我们针对大规模聊天机器人验证提出了两部分贡献。首先,我们引入了一种创建高保真合成客户代理(SCA)作为数字孪生的方法,该方法基于真实的交易和对话数据,能够实现自动生成和行为调节,从而模拟多样化的客户画像和交互风格。
Evaluation demonstrates that SCAs achieve high semantic alignment with real customers, low hallucination rates, and successful personality trait reproduction with controllable interventions. 评估表明,SCA 与真实客户实现了高度的语义对齐,具有较低的幻觉率,并能通过可控干预成功复现个性特征。
Second, we develop an SCA-based validation framework combining automated LLM-as-a-Judge evaluation, human expert testing, and adversarial probing. Scenario-based validation across emotional states, demographic groups, and linguistic factors confirms robust performance. 其次,我们开发了一个基于 SCA 的验证框架,结合了自动化的“LLM 作为裁判”(LLM-as-a-Judge)评估、人类专家测试和对抗性探测。针对不同情绪状态、人口统计群体和语言因素的场景化验证,证实了该框架的稳健性能。
Our approach was used to validate a customer-facing chatbot at a leading UK bank, providing financial institutions with a scalable pathway toward regulatory compliance. 我们的方法已被用于验证英国一家领先银行的面向客户的聊天机器人,为金融机构提供了一条实现合规的可扩展路径。