Who Questions What Works: When Should We Retest Our Assumptions?

Who Questions What Works: When Should We Retest Our Assumptions?

谁在质疑行之有效的方法:我们何时该重新审视自己的假设?

“What giants?” asked Sancho Panza. “Those you see over there,” replied his master, “with the long arms; sometimes they are almost two leagues long.” “Look, your grace,” Sancho responded, “those things that appear over there aren’t giants but windmills, and what looks like their arms are the sails that are turned by the wind and make the grindstone move.” “It seems clear to me,” replied Don Quixote, “that thou art not well-versed in the matter of adventures: these are giants; and if thou art afraid, move aside and start to pray whilst I enter with them in fierce and unequal combat.” — Miguel de Cervantes, Don Quixote, Part I, Chapter VIII [1]

“什么巨人?”桑丘·潘沙问道。“就是你在那边看到的那些,”他的主人回答说,“长着长胳膊的;有时它们几乎有两里格长。”“瞧,大人,”桑丘回应道,“那边出现的那些东西不是巨人,而是风车,看起来像胳膊的东西是风吹动的帆,带动磨石转动。”“在我看来很清楚,”堂吉诃德回答说,“你不精通冒险之事:这些就是巨人;如果你害怕,就退到一边去祈祷,让我与它们进行一场激烈而不平等的战斗。”——米格尔·德·塞万提斯,《堂吉诃德》,第一部,第八章 [1]

Introduction

引言

The scientific method is built around a simple idea: formulate a hypothesis, test it against reality, and decide whether the evidence supports it. In data science, we do this constantly. We run experiments, build proofs of concept, compare models, validate performance, and ask whether an idea works before investing further in it. We have become remarkably good at testing hypotheses before trusting them.

科学方法建立在一个简单的理念之上:提出假设,根据现实进行测试,并判断证据是否支持该假设。在数据科学中,我们不断地这样做。我们运行实验、构建概念验证、比较模型、验证性能,并在进一步投入之前询问一个想法是否可行。在信任假设之前,我们已经非常擅长对其进行测试。

But what happens after they work? A hypothesis that survives an experiment can become a model. A model that performs well can become a system. A successful system can become a process, and a process repeated for long enough can eventually become the way things are done. Somewhere along that path, an interesting reversal can occur: what once had to prove itself against reality can eventually become part of the lens through which we interpret it.

但当它们奏效之后会发生什么呢?一个经受住实验考验的假设可以成为模型。一个表现良好的模型可以成为系统。一个成功的系统可以成为流程,而一个流程如果重复足够长的时间,最终可能成为既定的行事方式。在这个过程中的某个节点,可能会出现一种有趣的逆转:曾经必须通过现实来证明其正确性的东西,最终可能成为我们解读现实的视角的一部分。

And reality does not stand still. Populations change, and as a consequence, data-generating processes change with them. Technologies evolve, organizations adapt, and yet past success provides a powerful reason to keep trusting the assumptions that produced it. This raises a question that goes beyond model monitoring or technical performance: How often do we retest the assumptions behind something that still seems to work?

现实并非静止不动。人口在变化,随之而来的是数据生成过程也在变化。技术在演进,组织在适应,然而过去的成功为继续信任那些产生成功的假设提供了强有力的理由。这提出了一个超越模型监控或技术性能的问题:对于那些看起来仍然有效的事物,我们多久会重新审视其背后的假设一次?

In this article, I want to explore that question at different levels, applied to data science methods, and consider a deceptively simple possibility: What if we keep adapting our models and methods without adapting the way we understand the problem?

在本文中,我想在不同层面上探讨这个问题,将其应用于数据科学方法,并考虑一种看似简单却发人深省的可能性:如果我们不断调整模型和方法,却不调整我们理解问题的方式,会怎样?

Kodak and Moody’s illustrate two different forms of the same underlying problem. Kodak represents the more familiar story: an organization struggles to move beyond the logic that made it successful. Moody’s presents a more subtle (and perhaps more dangerous) version. The organization did adapt. Models were revised, methodologies evolved, and new information was incorporated as markets changed. Yet those changes could remain bounded by the assumptions already embedded in the system [2]. Kodak asks what happens when we fail to change. Moody’s asks a harder question: what if we are changing all the time, but not changing what really matters?

柯达(Kodak)和穆迪(Moody’s)展示了同一个潜在问题的两种不同形式。柯达代表了一个更熟悉的故事:一个组织难以超越曾经使其成功的逻辑。穆迪则呈现了一个更微妙(或许更危险)的版本。该组织确实进行了调整。随着市场的变化,模型被修订,方法论在演进,新的信息也被纳入其中。然而,这些变化可能仍然局限于系统内已有的假设 [2]。柯达问的是当我们未能改变时会发生什么。穆迪则提出了一个更难的问题:如果我们一直在改变,但却没有改变真正重要的事情,那会怎样?

When Success Becomes a Constraint

当成功成为一种束缚

Everybody knows the cautionary tale of Kodak, the classic case study of a company that failed to adapt to change. But I want to use Kodak to make a point that goes beyond the usual story of technological disruption. For decades, Kodak built an extraordinarily successful business around a particular understanding of photography: cameras generated demand for film, film generated recurring revenue, and processing completed the ecosystem. Digital photography did not simply introduce a new technology. It challenged the assumptions that made that system work [5].

每个人都知道柯达的警示故事,这是一个未能适应变革的公司的经典案例研究。但我希望利用柯达来阐述一个超越技术颠覆这一常规叙事之外的观点。几十年来,柯达围绕着对摄影的一种特定理解建立了一个极其成功的业务:相机产生对胶卷的需求,胶卷产生经常性收入,而冲印服务完成了整个生态系统。数码摄影不仅仅引入了一项新技术,它挑战了使该系统得以运作的那些假设 [5]。

Experience is valuable precisely because it allows us to recognize patterns and make decisions without rediscovering everything from scratch. But there is a paradox: the more successful a particular interpretation of the world becomes, the easier it is to stop seeing it as an interpretation at all. This can be reinforced by what behavioral theory calls status quo bias: once a particular way of working is established, we become disproportionately inclined to preserve it rather than reconsider the alternatives. Past success makes that tendency even easier to justify.

经验之所以宝贵,正是因为它使我们能够识别模式并做出决策,而无需从零开始重新发现一切。但这里存在一个悖论:对世界的某种解读越成功,我们就越容易不再将其视为一种“解读”。这会被行为理论所称的“现状偏见”所强化:一旦某种特定的工作方式确立,我们就会不成比例地倾向于维持它,而不是重新考虑其他选择。过去的成功使得这种倾向更容易被合理化。

The same thing can happen in data science. Our models do not begin with algorithms. They begin with decisions about how a problem should be represented. We decide what data matters, how it should be transformed, which relationships are worth modelling, what success looks like, and which assumptions are reasonable enough to proceed. At first, these are choices. But when they work repeatedly, they become practices. Practices become processes, and processes eventually become methodology. What began as a hypothesis about how to solve a problem can quietly become the accepted way of solving it.

同样的事情也可能发生在数据科学中。我们的模型并非始于算法,而是始于关于如何表征问题的决策。我们决定哪些数据重要、如何转换数据、哪些关系值得建模、什么是成功,以及哪些假设合理到可以继续进行。起初,这些只是选择。但当它们反复奏效时,它们就变成了实践。实践变成了流程,流程最终变成了方法论。最初关于如何解决问题的假设,可能会悄然成为解决该问题的公认方式。

And this creates a more subtle kind of risk. The problem is no longer simply whether a model becomes outdated. It is whether we can keep updating the model while leaving the way we frame the problem largely untouched.

这产生了一种更微妙的风险。问题不再仅仅是模型是否过时,而是我们是否在保持问题框架基本不变的情况下,仅仅是在不断更新模型。

Models Capture Reality Through Assumptions

模型通过假设捕捉现实

A model learns from the world it has seen. The difficult question is whether it remains useful when the world changes. The question is not only how to predict Drug B, but why relationships learned from Drug A should still hold for it.

模型从它所见的世界中学习。困难的问题在于,当世界发生变化时,它是否仍然有用。问题不仅在于如何预测药物 B,还在于为什么从药物 A 中学到的关系仍然适用于药物 B。

You have probably heard this many times: a predictive model is, by definition, a simplification. It learns relationships from observations generated under particular conditions. We choose variables, define outcomes, make assumptions, and reduce a complex reality into something we can model. There is nothing inherently wrong with using historical drugs to predict the uptake of a new one [3, 4]. In fact, learning from the past is the foundation of predictive modelling.

你可能已经听过很多次了:预测模型从定义上讲就是一种简化。它从特定条件下产生的观察结果中学习关系。我们选择变量、定义结果、做出假设,并将复杂的现实简化为我们可以建模的对象。使用历史药物来预测新药的采用并没有什么本质上的错误 [3, 4]。事实上,从过去学习正是预测建模的基础。