The Model Validation Playbook for GenAI: Lessons from Banking

The Model Validation Playbook for GenAI: Lessons from Banking

生成式 AI 模型验证手册:来自银行业的经验教训

Introduction

引言

Let’s start with a recent, increasingly common scenario in the Risk Management department of large banks. Let’s say a risk model validator at a large bank opens a submission. The model is an AI assistant that reads a borrower’s financial statements, pulls relevant third-party research, and drafts the first version of a credit memo. It saves analysts several hours a week, and obviously the business wants this AI model to go live next quarter. 让我们从大型银行风险管理部门中一个近期且日益普遍的场景说起。假设某大行的风险模型验证员打开了一份提交的申请。该模型是一个人工智能助手,能够阅读借款人的财务报表,提取相关的第三方研究资料,并起草信贷备忘录的初稿。它每周能为分析师节省数小时的工作时间,显然,业务部门希望该 AI 模型能在下个季度上线。

She opens the standard validation template to start the review process. This template has been refined over a decade of regulatory examinations and worked on every scorecard, every loss forecasting model, every pricing engine she has reviewed. She reaches the first question: “Provide the development sample.” 她打开标准的验证模板开始审查流程。该模板经过了十年的监管审查打磨,在她审查过的每一个评分卡、每一个损失预测模型和每一个定价引擎中都发挥了作用。她看到了第一个问题:“提供开发样本。”

But there is no development sample. This gen AI model was trained on a corpus nobody at the bank has seen and by a vendor who won’t describe it. And that is only the first question from the remaining ninety. Model risk management was never designed for generative AI in banking. So, this is roughly where every model risk function in the banking/financial industry currently finds itself. An effective challenge on a model we cannot retrain, whose training data we cannot see, looks different from an effective challenge on a traditional scorecard. The craft shifts from replication to test design. 但根本没有开发样本。这个生成式 AI 模型是在一个银行内部无人见过的数据集上训练的,且供应商拒绝透露其细节。而这仅仅是剩余九十个问题中的第一个。银行业的模型风险管理从未为生成式 AI 设计。因此,这大致就是目前银行业/金融业中每个模型风险职能部门所处的困境。对于一个我们无法重新训练、训练数据无法查看的模型,有效的挑战方式与对传统评分卡的挑战截然不同。工作重心从“复现”转向了“测试设计”。

Why this framework matters beyond banking

为什么该框架在银行业之外同样重要

The core challenge described in this article, i.e., how to validate a system you cannot fully inspect, is now a problem for anyone deploying generative AI in a serious production context. Whether you are building a medical summarisation tool, a legal research assistant, or a customer-facing chatbot, the same questions apply: What does ‘good’ look like when there is no ground truth? How do you catch confident errors before they reach a user? 本文描述的核心挑战,即如何验证一个无法完全审查的系统,现在已成为任何在严肃生产环境中部署生成式 AI 的人所面临的问题。无论你是在构建医疗摘要工具、法律研究助手,还是面向客户的聊天机器人,同样的问题都适用:在没有“基准真值”(ground truth)的情况下,什么样的结果才算“好”?如何在错误到达用户之前捕捉到那些“自信的错误”?

The framework that follows in this article, based on risk tiering, outcome-based evaluation, robustness testing, and monitoring for silent drift, was built for banks, but it is directly transferable to any domain where the cost of being wrong matters more than the cost of being slow. 本文后续提出的框架基于风险分级、基于结果的评估、稳健性测试以及针对隐性漂移的监控。它虽是为银行构建的,但完全可以直接迁移到任何“犯错成本高于行动迟缓成本”的领域。

What model risk management in banking actually does

银行业的模型风险管理究竟在做什么

If you work in data science outside banking, this discipline may be unfamiliar. So let’s set up the context properly. Banks run on mostly traditional statistical predictive models. These models decide who gets credit and at what price. Models set how much capital the institution must hold against its loan book. Models forecast losses under hypothetical recessions, value illiquid positions, flag suspicious transactions, and determine reserves that flow directly into published financial statements. When one of these is wrong, the consequences are not an unhappy user; they are mispriced risk, understated reserves, regulatory findings/penalty, and occasionally a very large loss. 如果你在银行业之外从事数据科学工作,这个学科可能比较陌生。因此,让我们先建立正确的背景。银行主要依靠传统的统计预测模型运行。这些模型决定了谁能获得信贷以及信贷价格。模型决定了机构必须为贷款账簿持有多少资本。模型预测假设性经济衰退下的损失、评估非流动性头寸、标记可疑交易,并确定直接计入已发布财务报表的准备金。当其中一个模型出错时,后果不仅仅是用户不满意;而是风险定价错误、准备金低估、监管处罚,有时甚至是巨额亏损。

The industry learned this expensively. Credit models that assumed house prices don’t fall nationally contributed materially to the 2008 crisis. A revised risk model at one bank in 2012 understated exposure so badly that a trading loss ran into billions before anyone caught it. Regulators responded by formalising the discipline: US supervisory guidance issued in 2011 (known to everyone in the field as SR 11-7) defined model risk as the potential for adverse consequences from decisions based on incorrect or misused model output, and required banks to manage it deliberately. 银行业为此付出了昂贵的代价。那些假设全国房价不会下跌的信贷模型在 2008 年金融危机中起到了推波助澜的作用。2012 年,某银行的一个修订版风险模型严重低估了风险敞口,导致在被发现前就产生了数十亿美元的交易损失。监管机构随后将这一学科规范化:2011 年发布的美国监管指南(业内称为 SR 11-7)将模型风险定义为“基于错误或滥用模型输出的决策所带来的潜在不利后果”,并要求银行对其进行审慎管理。

The EU AI Act codifies a similar expectation for high-risk AI systems used in creditworthiness assessments, pricing, or essential banking services. Its core obligations on risk management, data governance, technical documentation, record-keeping, transparency, human oversight, and accuracy/robustness map closely onto SR 11-7’s conceptual soundness, outcomes analysis, and ongoing monitoring. For global banks, one validation framework can be structured to satisfy both regimes, but the AI Act adds explicit requirements around fundamental rights impact assessments and post-market monitoring that extend the second line’s traditional scope. 欧盟《人工智能法案》(EU AI Act)对用于信用评估、定价或基本银行服务的高风险 AI 系统提出了类似的期望。其在风险管理、数据治理、技术文档、记录保存、透明度、人工监督以及准确性/稳健性方面的核心义务,与 SR 11-7 中的“概念稳健性”、“结果分析”和“持续监控”高度吻合。对于全球性银行而言,可以构建一个统一的验证框架来同时满足这两套监管体系,但《人工智能法案》增加了关于基本权利影响评估和上市后监控的明确要求,这扩展了“第二道防线”的传统职能范围。

The model risk management structure is remarkably consistent across large institutions: 大型机构的模型风险管理结构非常一致:

Line of defenceWhoRole
FirstBusiness and model developmentBuilds the model, tests it, owns its performance and its use
SecondModel risk management/validationIndependently challenges the model before approval, and keeps challenging it
ThirdInternal auditChecks that the first two are doing their jobs
防线部门职责
第一道业务与模型开发构建模型、测试模型,对模型的性能及其使用负责
第二道模型风险管理/验证在批准前对模型进行独立挑战,并持续进行挑战
第三道内部审计检查前两道防线是否履行了职责

The second line is the part this article is about. A validator doesn’t just check arithmetic. They ask whether the modelling approach was conceptually appropriate, whether the data supported it, whether the output actually performs, whether the production implementation matches what was approved, and whether the people using the output understand its limits. Nothing goes live without their sign-off, and everything gets re-examined periodically. 本文讨论的重点是第二道防线。验证员不仅仅是检查算术。他们会询问建模方法在概念上是否合适、数据是否支持该方法、输出结果是否真正有效、生产环境的实现是否与批准的一致,以及使用输出结果的人员是否了解其局限性。没有他们的签字,任何模型都无法上线,且所有模型都会定期接受重新审查。

Three things anchor that review, and they have been stable for over a decade: conceptual soundness (is the approach defensible?), outcomes analysis (does the output hold up when tested?), and ongoing monitoring (is it still working now?). 审查工作由三点支撑,且十多年来保持稳定:概念稳健性(方法是否站得住脚?)、结果分析(输出结果在测试中是否经得起考验?)以及持续监控(它现在是否仍然有效?)。

Why Generative AI Breaks Traditional Model Validation

为什么生成式 AI 打破了传统的模型验证

Generative AI has arrived in banks faster than any modelling technology in recent memory, and not in a specific shape. It can be complaint summarisation, policy lookup, research retrieval, first drafts of credit memos, internal documentation, literally anything. 生成式 AI 进入银行的速度比近期记忆中的任何建模技术都要快,而且形式多样。它可以是投诉摘要、政策查询、研究检索、信贷备忘录初稿、内部文档,几乎可以是任何东西。

These models are attractive because they directly influence cost, but they can also be risky. Because they sit close to customers and close to credit decisions. These are exactly the places where a regulated institution has the least appetite for a wrong answer. 这些模型之所以具有吸引力,是因为它们直接影响成本,但它们也可能带来风险。因为它们处于客户接触点和信贷决策的核心位置。而这恰恰是受监管机构最无法容忍错误答案的地方。

And the validation apparatus that existed to prevent this risk no longer fits. Every question on the template assumes properties these systems don’t have. 而现有的用于防范此类风险的验证机制已不再适用。模板上的每一个问题都预设了这些系统所不具备的属性。

  1. Five Structural Breaks in Gen AI Model Validation
  2. 生成式 AI 模型验证中的五个结构性断层

Property | Why classical validation can’t absorb it 属性 | 为什么传统验证无法涵盖它 There is no m… | … (注:原文此处中断)