Your AI Adoption Lift Is a Selection Effect
Your AI Adoption Lift Is a Selection Effect
你的 AI 采用率提升其实是一种“选择效应”
Somewhere in your company there is a slide that says something like this: customers who enabled the AI assistant retain 15 points better than customers who did not. It has a bar chart. It has been in three executive reviews. It is driving next quarter’s roadmap. 在你的公司里,某处一定有一张幻灯片写着类似这样的内容:启用了 AI 助手的客户留存率比未启用的客户高出 15 个百分点。它配有一张柱状图,已经在三次高管评审中出现过,并且正在推动下一季度的产品路线图。
Nobody randomized the AI assistant. It shipped to eligible accounts, some of them turned it on, and the analytics team compared the ones that did against the ones that did not. 没有人对 AI 助手进行过随机分组测试。它被推送给符合条件的账户,其中一些账户开启了它,分析团队随后对比了开启和未开启的账户。
That comparison is not an effect. It is a description of who opts in. The feature did not make those accounts engaged. Being engaged made them adopt the feature. 这种对比并不是(功能带来的)效果,它只是对“谁选择了加入”的一种描述。并不是这个功能让这些账户变得活跃,而是因为这些账户本身活跃,才选择了这个功能。
AI features make this worse than the average opt-in feature, and it is worth being specific about why. To adopt an AI assistant, someone at the account has to notice the release, enable it, trust it enough to put it in front of their team, train people on it, fold it into a workflow, and keep using it after the novelty fades. Every one of those steps reveals something about the account: administrator engagement, executive sponsorship, technical sophistication, product maturity, organizational appetite for change. By the time an account shows up as an “adopter,” the flag is close to a proxy for organizational readiness. Readiness predicts retention on its own. The feature is riding on top of it. AI 功能让这种情况比一般的可选功能更严重,值得具体说明原因。要采用 AI 助手,账户中的某个人必须注意到发布信息、启用它、足够信任它并将其展示给团队、对员工进行培训、将其整合进工作流,并在新鲜感消退后继续使用。每一个步骤都揭示了该账户的某些特质:管理员的参与度、高管的支持、技术成熟度、产品成熟度以及组织对变革的渴望。当一个账户显示为“采用者”时,这个标签几乎成了组织准备程度的代名词。而准备程度本身就能预测留存率,该功能只是搭了顺风车。
The instinct at this point is to model the customer’s choice harder. Add covariates. Match on usage. Build a propensity score. This article argues for a different move, and it is the one sentence I would keep if everything else were cut: Don’t model the customers’ choice harder. Find variation the customers didn’t choose. 此时的直觉是更深入地建模客户的选择:增加协变量、匹配使用情况、构建倾向评分。本文主张采取不同的做法,如果只能保留一句话,我会保留这一句:不要在客户的选择上过度建模,去寻找客户无法选择的变量。
The rest of this article is about doing that with the one piece of an AI rollout that no customer picked: the eligibility rule. 本文的其余部分将探讨如何利用 AI 推送中唯一一个客户无法选择的部分来实现这一点:资格准入规则。
Define the question before the method
在方法之前先定义问题
There are three different quantities hiding inside “the effect of the AI assistant,” and the slide conflates all of them. 在“AI 助手的影响”这一概念下隐藏着三个不同的量,而那张幻灯片将它们混为一谈了。
The effect on adopters. If the accounts that turned the feature on had not adopted, how much worse would their retention have been? This is what the naive comparison is trying to estimate. It is the number the product team wants, because it describes the customers who actually experienced the feature. 对采用者的影响。 如果那些开启了该功能的账户没有采用,它们的留存率会差多少?这就是天真的对比试图估算的数值。这是产品团队想要的数字,因为它描述了实际体验过该功能的客户。
The effect on everyone. If every eligible account adopted rather than nobody adopting, how much would retention change? This is the number finance wants, because it is what a forced rollout or a default-on change would produce. It is not the same number as the effect on adopters. Whether it is larger or smaller is an empirical question about effect heterogeneity; in opt-in product settings it is often reasonable to expect adopters to benefit more, but that is an assumption, not a theorem. 对所有人的影响。 如果所有符合条件的账户都采用了该功能,而不是无人采用,留存率会发生多大变化?这是财务部门想要的数字,因为它反映了强制推送或默认开启所产生的结果。这与“对采用者的影响”不是同一个数字。它更大还是更小是一个关于效应异质性的实证问题;在可选产品设置中,通常可以合理预期采用者会受益更多,但这只是一个假设,而非定理。
The effect at the margin. For accounts right at the edge of eligibility, what does becoming eligible do to retention, and what does adopting do for the accounts that adopt because they became eligible? Those are two numbers, not one, and the distinction matters later. These are the quantities nobody asks for and the ones you can usually identify most cleanly. They are also the ones that speak to the decision most often on the table: whether to move the eligibility rule. 边际影响。 对于处于资格准入边缘的账户,获得资格对留存率有什么影响?对于因获得资格而采用该功能的账户,采用行为又有什么影响?这是两个数字,而不是一个,这种区别在后续讨论中很重要。这些是没人会问、但通常最容易识别的量。它们也最能说明决策层经常面临的问题:是否应该调整资格准入规则。
The naive comparison does not estimate any of them. It estimates the difference between ready organizations and unready ones, with a feature flag attached. 天真的对比无法估算其中任何一个。它估算的是“准备就绪的组织”与“未准备就绪的组织”之间的差异,并附带了一个功能标签。
The setup
设置
The synthetic dataset has 40,000 B2B accounts. The AI assistant is available only to accounts with 25 or more seats, a seat-count eligibility rule of the kind common in SaaS products. Among eligible accounts, adoption is voluntary. 该合成数据集包含 40,000 个 B2B 账户。AI 助手仅对拥有 25 个或更多席位的账户开放,这是一种 SaaS 产品中常见的席位数量准入规则。在符合条件的账户中,采用是自愿的。
The thing that makes this hard is a latent variable I will call engagement: how invested the account is in the product. Engaged accounts are more likely to turn on new features and more likely to renew regardless. The analyst never observes it. What the analyst observes is seats, tenure, whether the account adopted, and whether it retained six months later. 让这个问题变得困难的是一个我称之为“参与度”的潜在变量:账户对产品的投入程度。参与度高的账户更有可能开启新功能,也更有可能续订。分析师无法观察到这个变量。分析师能观察到的是席位、任期、账户是否采用以及六个月后是否留存。
The true effect baked into the simulation is +4 percentage points of 6-month retention from adopting the assistant. Retention also trends smoothly upward with account size: bigger accounts retain a little better, feature or no feature. So the observed adopter gap will mix three things: the feature effect, selection on engagement, and the fact that adopters are drawn from larger, eligible accounts. 模拟中内置的真实效果是:采用助手可带来 4 个百分点的 6 个月留存率提升。留存率也随着账户规模的增加而平稳上升:无论是否有该功能,规模较大的账户留存率都会稍好一些。因此,观察到的采用者差距将混合三件事:功能效果、参与度选择偏差,以及采用者来自规模较大且符合条件的账户这一事实。
(Python code omitted for brevity)
The full data-generating process is in the notebook. The part that matters is above: adoption and retention share a cause the analyst cannot see. 完整的数据生成过程在笔记本中。关键部分如上所述:采用和留存共享一个分析师无法看到的共同原因。
Method 1: The slide
方法 1:那张幻灯片
Business question: Do accounts that use the AI assistant retain better? 业务问题: 使用 AI 助手的账户留存率更高吗?
What it estimates: The difference in retention between adopters and non-adopters. 估算内容: 采用者与非采用者之间的留存率差异。
Identifying assumption: Adopters and non-adopters would have retained identically absent the feature. Adoption is as good as random. 识别假设: 如果没有该功能,采用者和非采用者的留存率将完全相同。即采用行为等同于随机分配。
(Python code omitted for brevity)
Fifteen points. The true effect is four. The rest is selection: mostly engagement, plus the fact that adopters come from larger, eligible accounts that already retain somewhat better. 15 个百分点。真实效果是 4 个百分点。其余部分是选择偏差:主要是参与度,加上采用者来自规模较大、符合条件且本身留存率就稍好的账户这一事实。
It does not help much to restrict the comparison to eligible accounts, which is the usual first fix. Among accounts with 25 or more seats, the adopter gap is +13.9 pp. 将对比限制在符合条件的账户内(这是通常的第一种修正方法)并没有太大帮助。在拥有 25 个或更多席位的账户中,采用者差距仍为 +13.9 个百分点。