Optimal Traffic Allocation Under Heterogeneous Variant Cost

Optimal Traffic Allocation Under Heterogeneous Variant Cost

异构变体成本下的最优流量分配

What would you do if putting a user in the treatment group cost you twice what it costs to put one in control? I can’t blame you if the first answer that comes to mind is to allocate to the treatment only half the traffic of the control arm. Twice the price per user, so buy half as many; the spend on each arm comes out even. It may feel fair, but it’s not the cheapest answer to that question. 如果将用户放入实验组的成本是放入对照组的两倍,你会怎么做?如果你的第一反应是将对照组一半的流量分配给实验组,我完全可以理解。每个用户的价格是两倍,所以购买数量减半;这样两组的支出就持平了。这听起来很公平,但它并不是解决该问题的最经济方案。

Cost-optimal allocation has to weigh two things at once: how much precision one more subject buys you, and what that subject costs; in a setup where the answer to the second depends on which arm they land in. “So what?”, you may be thinking. Done properly it can save you anywhere between 5% to 15% of the budget; in other words, from thousands to millions of savings. But at the cost of what? Grab some coffee, and let’s find out when this stops being a free lunch. 成本最优分配必须同时权衡两件事:多一个样本能为你带来多少精度提升,以及该样本的成本是多少;在实验中,后者的答案取决于样本落入哪个分组。你可能会想:“那又怎样?”如果操作得当,它可以为你节省 5% 到 15% 的预算;换句话说,节省的金额从数千到数百万不等。但代价是什么呢?喝杯咖啡,让我们来看看什么时候这就不再是“免费的午餐”了。

In this post, I walk you through: 在这篇文章中,我将带你了解:

  • The relevance, given today’s LLM-eager industry, where differing costs across arms are the norm, thinking cost-optimal splits pays off. Discounts, vouchers, cashbacks are other relevant treatments where this show good value.
  • 相关性:在当今热衷于大模型(LLM)的行业中,各分组成本差异已成常态,考虑成本最优分配是非常值得的。折扣、优惠券、返现等也是体现其价值的相关实验场景。
  • The intuition of cost of the arms plays tug-of-war with the sample variances, and how that leads to twice-the-cost-half-the-traffic NOT being a good answer to the question.
  • 直觉:各组成本与样本方差之间的博弈,以及为什么“成本翻倍、流量减半”并不是一个好的解决方案。
  • The edge cases; for when reality does not fit the math. Experiments don’t run in isolated academic minds. Incentives in business challenge this idea, rightfully.
  • 边界情况:当现实不符合数学模型时。实验并非在孤立的学术思维中运行。商业激励机制合理地挑战了这一理念。

When the treatment costs money

当实验需要成本时

Personally, I’ve rarely encountered the question of cost-optimal traffic allocation. And I kept thinking: why not? I know the math is not lying, but I also know that experimentation is rarely about the stats alone; like many things organisations, but in particular experimentation, there is a strong business and socio-technical aspect that shapes how it’s done. 就我个人而言,我很少遇到关于成本最优流量分配的问题。我一直在想:为什么呢?我知道数学不会撒谎,但我也知道实验绝不仅仅是统计学问题;就像组织中的许多事情一样,尤其是实验,其执行方式深受商业和社会技术因素的影响。

Discounting costs money; so do LLMs

折扣需要成本;大模型也是如此

In pretty much every industry, recruiting subjects for an experiment is a line in the budget. And sometimes every treatment is the line. In healthcare a new medicine that costs more than the placebo; in economics, in incentive-driven or pay-for-performance designs, every treated subject gets some form of value to nudge them toward the behaviour under study. 在几乎所有行业中,招募实验对象都是预算中的一项开支。有时,每一个实验组都是一项开支。在医疗保健领域,新药的成本高于安慰剂;在经济学中,在激励驱动或按绩效付费的设计中,每个受试者都会获得某种形式的价值,以引导他们产生研究关注的行为。

In tech, however, there’s a comfortable notion that participants are “free”. There is no one to recruit, convince, or pay; they’re already using your product, so they comply and participate whether they know it or not. 然而在科技行业,有一种普遍的舒适观念,认为参与者是“免费的”。不需要招募、说服或付费;他们已经在用你的产品,所以无论他们是否知情,他们都会顺从并参与其中。

But today half the industry is racing to duct-tape agentic features onto existing products. That’s an open invitation to expensive treatment groups: the AI experience against the old one, where the mere act of treating a user fires off a paid API call. More traditional examples are simply: vouchers, discounts, cash-backs promos, and the like. When a cost differential exists across arms, experimentation programs should start caring. 但今天,半个行业都在竞相将智能体功能“强行”拼接到现有产品上。这直接导致了昂贵的实验组出现:AI 体验对比旧体验,仅仅是让用户体验 AI 功能就会触发付费的 API 调用。更传统的例子包括:优惠券、折扣、返现促销等。当各组之间存在成本差异时,实验项目就应该开始关注这一点。

So, the notion that experimentation is (marginally) free for tech companies, is not too much off, but there is certainly a group of treatments that make cost-optimal designs an interesting path to explore. 因此,认为科技公司的实验(边际)成本为零的观点虽然不算太离谱,但确实存在一类实验,使得成本最优设计成为一个值得探索的方向。

What’s the main blocker then? Experimenting with discounts has been around forever. But so has been the need for speed. 那么主要的阻碍是什么呢?折扣实验已经存在很久了。但对速度的需求也同样存在。

Velocity, defaults, and conventions

速度、默认设置与惯例

Velocity is almost everything in digital innovation. Experiments need to come in fast to stay ahead of the competition. We can see that back in two things: the amount of research done in variance reduction techniques, and… 在数字创新中,速度几乎就是一切。实验必须快速进行才能保持竞争优势。我们可以从两件事中看出这一点:在方差缩减技术上投入的研究量,以及……

If you would be asked: why do we (ideally) split traffic 50/50? The answer would be somewhere along the lines of: that’s when power is highest; ergo, we can find the effect with the shortest runtime possible. That question gets asked far more often than what’s the cost-optimal split? 如果有人问你:为什么我们(理想情况下)要进行 50/50 的流量分配?答案通常是:因为此时统计功效最高;因此,我们能以最短的运行时间发现效应。这个问题被问到的频率远高于“什么是成本最优分配?”

The pull towards runtime/sensitivity is already baked into traditions. Without a clear protocol, or understanding, on how to govern the traffic split, it’s easy to default to standard reasoning like the 50/50. It has helped in most cases. 对运行时间/灵敏度的追求已经根植于传统之中。如果没有明确的协议或对如何管理流量分配的理解,人们很容易默认采用 50/50 这种标准逻辑。在大多数情况下,这确实有效。

It happens to be that a cost-optimal is not per se the one that leads to highest precision, and lowest runtimes. So optimising one objective, means undermining the other one. It’s tug-of-war between cost and velocity. Deciding which one to optimise requires clear understanding of the constraints and priorities; budget, experiment goal, committed timelines, etc. It’s not just a statistical problem. 恰好,成本最优并不一定意味着精度最高或运行时间最短。因此,优化一个目标就意味着削弱另一个目标。这是成本与速度之间的博弈。决定优化哪一个需要对约束条件和优先级有清晰的理解;包括预算、实验目标、承诺的时间表等。这不仅仅是一个统计学问题。

That said, we really need to understand how the stats work for this to come alive. That’s the only way that a principled conversation could be held, or perhaps started at altogether. As a data scientist, one can play a crucial role in putting the gains in the spotlights. 话虽如此,我们确实需要理解统计学原理才能让这一切发挥作用。这是进行原则性对话,或者说开启对话的唯一途径。作为一名数据科学家,可以在将这些收益置于聚光灯下方面发挥关键作用。

Back to the basics

回归基础

Say, you want to run an experiment. Every subject in the treatment arm is a few times more expensive than each one in the control arm. You have a limited budget, or simply want to pay the least for the same learning, in the same duration window, with the same total number of subjects. The question then is: which split of traffic buys you the most power per dollar? 假设你想进行一项实验。实验组中的每个受试者的成本是控制组的几倍。你的预算有限,或者只是想在相同的持续时间内、使用相同总样本量的情况下,以最低的成本获得相同的实验结论。那么问题来了:哪种流量分配方式能让你每一美元获得的统计功效最高?

What if I told you that the answer is as plain as: 如果我告诉你答案就像下面这样简单呢:

$$\frac{n_1}{n_0} = \sqrt{\frac{c_0}{c_1}}$$

The optimal sample ratio, $n_1/n_0$ (subscript 1 is treatment, 0 is control, throughout), is the square root of the inverse ratio of the arms’ marginal costs: $\sqrt{c_0/c_1}$. As pictures tell more than a thousand formulas, let’s plot that out: 最优样本比率 $n_1/n_0$(下标 1 代表实验组,0 代表对照组)是各组边际成本反比的平方根:$\sqrt{c_0/c_1}$。由于图片胜过千言万语,让我们将其绘制出来:

(Image description: The optimal sample ratio $n_1/n_0$ against the cost ratio $c_1/c_0$: the square root bends the curve, so even large cost gaps end up as modest skews. The raw ratio $c_1/c_0$ is plotted for reference.) (图片描述:最优样本比率 $n_1/n_0$ 与成本比率 $c_1/c_0$ 的关系:平方根使曲线弯曲,因此即使成本差距很大,最终也只会导致适度的偏差。图中绘制了原始比率 $c_1/c_0$ 以供参考。)

Intuition. Doubling the cost does not mean halving the allocation to the expensive arm. The twice-the-cost-half-the-subjects rule treats every subject as equally valuable, wherever they sit and however big the sample already is. 直觉:成本翻倍并不意味着将昂贵分组的分配量减半。“成本翻倍、样本减半”的规则将每个受试者视为同等价值,无论他们处于哪个分组,也无论样本量已经有多大。

Question for you: is a subject’s contribution to precision actually constant? As sample size grows… 问你一个问题:受试者对精度的贡献真的是恒定的吗?随着样本量的增加……