Why You Think Like a Bayesian but Were Taught Like a Frequentist
Why You Think Like a Bayesian but Were Taught Like a Frequentist
为什么你的思维方式是贝叶斯式的,但受到的教育却是频率派的
The Schoolroom Distortion
课堂上的认知偏差
Think back to your first high school statistics class. The teacher walks to the blackboard, draws a coin, and asks a seemingly simple question: “What is the probability of flipping heads?” You and everyone else in the room answer instantly: “Fifty percent.” If pressed for an explanation, you’d probably say that if you were to flip that coin an infinite number of times, half of those flips would land on heads. 回想一下你的高中第一堂统计课。老师走到黑板前,画了一枚硬币,问了一个看似简单的问题:“抛硬币出现正面的概率是多少?”你和教室里的其他人会立刻回答:“百分之五十。”如果被追问原因,你可能会说,如果你抛无数次硬币,其中一半会是正面。
This is Frequentism, and it is the standard operating process of our educational system. It defines probability through the cold, objective lens of repeatable data. In the frequentist world, probability is understood through what would happen if we could repeat the same experiment over and over again. It is a neat, comforting framework designed for a world made of rolling dice, shuffled decks of cards, and endless time. 这就是“频率派”(Frequentism),也是我们教育体系的标准运作模式。它通过可重复数据这一冷冰冰的客观视角来定义概率。在频率派的世界里,概率是通过“如果我们能反复进行同一实验会发生什么”来理解的。这是一个整洁、令人安心的框架,专为掷骰子、洗扑克牌和拥有无限时间的理想世界而设计。
But there is a catch. The moment you step out of the classroom, the laboratory walls crumble. Real life rarely offers us the luxury of infinite trials. You cannot marry someone a thousand times to measure the probability of a happy marriage, nor can a company launch the exact same product campaign on Instagram a million times to test the engagement gain. In the messy arena of human existence, the frequentist definition of probability starts to feel less like a tool and more like a straitjacket. 但问题在于,当你走出教室的那一刻,实验室的围墙就崩塌了。现实生活很少给我们提供无限次试验的奢侈。你不可能为了衡量婚姻幸福的概率而与同一个人结婚一千次,公司也不可能在 Instagram 上发布一百万次完全相同的产品营销活动来测试参与度。在人类生存这个混乱的竞技场中,频率派对概率的定义开始显得不再像是一种工具,而更像是一件紧身衣。
The Reality: Built to be Bayesian
现实:天生的贝叶斯思维
And at this very point is where the grand illusion of our education becomes apparent: while we were trained to think like frequentists on paper, we seem to be naturally inclined to reason in Bayesian-like ways. 正是在这一点上,我们教育中的巨大幻觉显现了出来:虽然我们在纸面上被训练成像频率派那样思考,但我们似乎天生就倾向于以贝叶斯式的方式进行推理。
In the real world, probability isn’t about counting repetitions in the infinite future; it is about quantifying uncertainty in the present. This is the core of Bayesian statistics. Instead of demanding endless data before making a judgment, a Bayesian starts with a Prior, an initial belief based on past experience or intuition. Then, as new evidence arrives, they update that belief to arrive at a Posterior probability. 在现实世界中,概率不是关于计算无限未来的重复次数,而是关于量化当下的不确定性。这就是贝叶斯统计的核心。贝叶斯主义者在做出判断前,不会要求海量的数据,而是从“先验”(Prior)开始——即基于过去经验或直觉的初始信念。然后,随着新证据的出现,他们会更新这一信念,从而得出“后验”(Posterior)概率。
We don’t need a math degree to do this. The human brain often behaves like a Bayesian prediction engine. Our ancestors in the savannah didn’t have the luxury of waiting for a rustling bush to move ninety-nine times to calculate a p-value before running from a predator. They had a powerful prior: “Rustling bush equals danger.” They saw a tiny bit of new data (maybe a flicker of yellow fur) instantly updated their probability, and survived. It is not hard to see why a fast, Bayesian-like way of updating beliefs could have been useful for survival. 我们不需要数学学位就能做到这一点。人类大脑的表现往往就像一个贝叶斯预测引擎。我们在大草原上的祖先没有奢侈的时间去等待灌木丛晃动九十九次来计算一个 p 值,然后再从捕食者面前逃跑。他们拥有一个强大的先验:“灌木丛晃动等于危险。”他们看到一点点新数据(也许是黄色皮毛的一闪而过),便瞬间更新了概率并幸存下来。不难看出,这种快速的、贝叶斯式的信念更新方式为何对生存如此有用。
Of course, none of this means we are flawless statisticians. Decades of work by Kahneman and Tversky showed something paradoxical: when asked to reason about probabilities, with numbers on paper, humans are notoriously bad. We routinely ignore base rates and get textbook Bayesian problems wrong. But this is exactly the point. We run the Bayesian engine beautifully when we don’t think about it; we just can’t read its dashboard. Our intuition is the posterior; our conscious math is the bug. 当然,这并不意味着我们是完美的统计学家。卡尼曼(Kahneman)和特沃斯基(Tversky)几十年的研究揭示了一个悖论:当被要求在纸上用数字进行概率推理时,人类的表现非常糟糕。我们经常忽略基础比率,并在教科书式的贝叶斯问题上出错。但这正是重点所在:我们在不假思索时能完美地运行贝叶斯引擎,只是无法解读它的仪表盘。我们的直觉就是后验,而我们有意识的数学计算反而是那个“漏洞”。
The Supermarket and the Waiting Game
超市与等待游戏
To see this evolutionary machinery in action, we don’t need to fight off tigers. We just need to look at how we navigate ordinary, adult life. Imagine walking into a premium supermarket while traveling in a foreign country. You spot a high-end chocolate bar on a shelf, but there is no price tag. A strict frequentist approach would have a harder time here: you have zero observations for this exact item in this exact store. Theoretically the price could be two euros or two hundred, and both are equally unknown. 要观察这种进化机制的运作,我们不需要去对抗老虎,只需要看看我们如何应对普通的成年生活。想象一下,你在国外旅行时走进一家高档超市。你在货架上看到一块高端巧克力,但没有标价。严格的频率派方法在这里会遇到困难:你对这家店里的这个特定商品没有任何观察记录。理论上价格可能是两欧元,也可能是两百欧元,两者同样未知。
But you don’t panic, because your brain instantly deploys a prior. You know roughly what chocolate costs in general. You register the elegant lighting, the wooden shelving, the fact that everything else here is somehow imported. Before you have even touched the wrapper, you are carrying a probability distribution in your head, centred somewhere around four euros with a tail that would not be shocked by seven. 但你不会惊慌,因为你的大脑会瞬间部署一个先验。你大致知道巧克力的普遍价格。你注意到了优雅的灯光、木质货架,以及这里所有东西似乎都是进口的事实。在你还没触碰到包装纸之前,你的脑海中就已经形成了一个概率分布,中心大约在四欧元左右,即便价格是七欧元也不会让你感到震惊。
Notice what just happened. Nobody handed you data on this product. You built a prior out of context like the store, the country, the packaging. This is not sloppy thinking; it is the most information-efficient move available to you. Then comes the cashier and says “That will be 14 euros.” Your prediction error spikes as the number sits far out in the tail of what you were expecting. And by the time you walk out the door, you have quietly revised something bigger than the price of one chocolate bar: your whole sense of what “premium” costs in this country. The next unlabelled shelf you meet, you will guess higher or maybe you will run away. 注意刚才发生了什么。没有人给你提供关于这个产品的任何数据。你通过商店、国家、包装等背景信息构建了一个先验。这不是草率的思考,而是你所能采取的信息效率最高的行动。接着收银员说:“一共 14 欧元。”你的预测误差瞬间飙升,因为这个数字远远超出了你的预期范围。当你走出店门时,你已经悄悄修正了一个比一块巧克力价格更重要的东西:你对这个国家“高端”消费的整体认知。下次遇到没有标价的货架时,你会猜得更高,或者干脆走开。
If pricing chocolate feels too trivial, consider the higher stakes of a first date. The date goes well, the conversation flows, you laugh at the same jokes, you float home. You believe there is roughly an 80% chance they want a second date. Before falling asleep, you send a text: “I had a wonderful time tonight.” Now the evidence starts arriving. A reply in two minutes nudges your belief up towards 95% and you sleep like a baby. But two hours pass. Then five. The next morning your screen is still blank and this silence is deeply improbable under your original hypothesis. Without ever opening a statistics textbook, your brain grinds through a brutal round of updating. Eighty percent becomes fifty and fifty becomes thirty. Your prior stays exactly where it was, of course, because a prior is what you believed before. What is quietly dying overnight is your posterior. 如果给巧克力定价显得太琐碎,那就考虑一下第一次约会这种高风险的情况。约会很顺利,谈话流畅,你们一起笑话,你飘飘然地回家。你认为对方想进行第二次约会的概率大约是 80%。睡前,你发了一条短信:“今晚我很开心。”现在证据开始出现了。两分钟后的回复将你的信念推向了 95%,你睡得很香。但两小时过去了,五小时也过去了。第二天早上,屏幕依然一片空白,这种沉默在你的原始假设下是极不可能发生的。无需翻开统计学教科书,你的大脑就在进行一轮残酷的更新。80% 变成了 50%,50% 又变成了 30%。当然,你的先验保持不变,因为先验是你之前所相信的。而在一夜之间悄然消亡的,是你的后验。
A frequentist would frame the question differently. Instead of asking “what is the probability that this person likes me?”, they would ask what would happen to the response rate across repeated, comparable situations. But you don’t have a thousand attempts, of course. You have one. 频率派会以不同的方式提出问题。他们不会问“这个人喜欢我的概率是多少?”,而是会问在重复的、可比较的情况下,回复率会发生什么变化。但当然,你没有一千次尝试的机会。你只有一次。
The Mathematical Straitjacket
数学的紧身衣
If the Bayesian approach is so natural, so deeply embedded in our cognitive wiring, why does the educational system still force… 如果贝叶斯方法如此自然,如此深刻地嵌入我们的认知结构中,为什么教育体系仍然强迫……