Random rewards enrich classic game-theory contests

Random rewards enrich classic game-theory contests

随机奖励丰富了经典的博弈论竞赛

Games may be life with all the hard bits removed, but they provide a way to study why people make the choices they make. Traditional games are usually played against a static background: the rewards per outcome are constant. That limits their relevance to behavior because, in real life, the rewards and consequences of strategic choices are ever changing. Now, researchers have used a mathematical model to study a series of games that include evolving strategies and randomly varying returns.

游戏或许是去除了所有艰难困苦后的生活缩影,但它们提供了一种研究人们为何做出特定选择的方法。传统的博弈通常在静态背景下进行:即每次结果的奖励是恒定的。这限制了它们对现实行为的解释力,因为在现实生活中,战略选择的奖励和后果是不断变化的。现在,研究人员利用数学模型研究了一系列包含演化策略和随机变化收益的博弈。

A bit of history

历史背景

Perhaps the most famous game-theory contest is the prisoner’s dilemma. In the prisoner’s dilemma, a pair of thieves have been captured and are being separately interrogated by the police. If both clam up, they will be punished for a lesser crime. If one prisoner makes a deal (defects) then that prisoner gets to go free and the other gets a heavier sentence. If both make a deal, they both get an in-between punishment. The person running the game can start it with different rewards for cooperating and defecting to explore how the optimum strategy varies with reward and risk, which the players can figure out by varying the strategies across multiple rounds. Depending on the balance between the reward for staying silent (cooperating) and betrayal, the game stabilizes with everyone betraying everyone. In this simple situation, everyone loses.

博弈论中最著名的竞赛或许就是“囚徒困境”。在囚徒困境中,两名窃贼被捕并被警方分开审讯。如果两人都保持沉默,他们将因较轻的罪行受到惩罚。如果其中一名囚犯达成交易(背叛),那么该囚犯将获释,而另一名囚犯将受到更重的判决。如果两人都达成交易,他们都将受到中等程度的惩罚。博弈的发起者可以通过设置不同的合作与背叛奖励来开始游戏,从而探索最优策略如何随奖励和风险的变化而变化,玩家可以通过在多轮博弈中调整策略来摸索出这一点。根据保持沉默(合作)与背叛之间的奖励平衡,博弈最终会稳定在“所有人背叛所有人”的状态。在这种简单的情况下,每个人都是输家。

Similar dynamics can be found in games of chicken, rock-paper-scissors, and more. The evolution of strategies can lead to stable populations, bistable populations (where the population flips between two stable strategies), or limit cycles, where the population shifts continuously among multiple strategies. There is also a rich history of changing a game as it is played. Usually, these are within-game variations. For instance, you can set a limit on the amount of reward available, so strategies evolve to take into account increasingly limited resources as the number of rounds goes up. In other words, most of this older work studied situations where the player’s behavior in the current round changed the resources or rewards available for the next round.

类似的动态也存在于“胆小鬼博弈”、石头剪刀布等游戏中。策略的演化可能导致稳定的种群、双稳态种群(种群在两种稳定策略之间切换),或极限环(种群在多种策略之间持续转换)。在博弈过程中改变规则也有着丰富的历史。通常,这些属于博弈内的变化。例如,你可以限制可用奖励的总量,这样随着轮数的增加,策略会演化以考虑到日益有限的资源。换句话说,大多数早期的研究探讨的是玩家在当前轮次的行为会改变下一轮可用资源或奖励的情况。

Your influence is limited

你的影响力是有限的

In real life, though, we are driven by external factors that are not under our control. The rabbit does not control the rain that floods its burrow or the drought that kills its food. The game changes each round because the risks and rewards of each strategy change over time. The researchers recognized that incorporating these dynamics into a mathematical model might reveal new behaviors. And they were not wrong.

然而在现实生活中,我们受到无法控制的外部因素驱动。兔子无法控制淹没其洞穴的雨水,也无法控制导致其食物枯竭的干旱。博弈在每一轮都在变化,因为每种策略的风险和奖励随时间而改变。研究人员意识到,将这些动态纳入数学模型可能会揭示出新的行为模式。他们是对的。

The prisoner’s dilemma as described above has only a single stable point (everyone loses). The model quickly converges to that point, where all players adopt the same strategy within a few rounds. But the new work found that if the rewards change with time (even by a small amount), then a second stable point emerges from the model, allowing both cooperators and defectors to coexist. Under even greater variation, the defector point becomes unstable, meaning that only cooperators exist.

如上所述,囚徒困境只有一个稳定点(所有人皆输)。模型会迅速收敛到该点,即所有玩家在几轮内都采取相同的策略。但这项新研究发现,如果奖励随时间变化(即使变化很小),模型中就会出现第二个稳定点,允许合作者和背叛者共存。在更大的变化下,背叛点变得不稳定,这意味着只剩下合作者存在。

In chicken, the model’s results are a bit scarier: With no change in rewards, the stable point is that everyone swerves and we all survive. Add in just a bit of variation, and a population that does not swerve emerges. Add in more noise, and a bistable flipping between survival and crashing emerges. (Given that chicken was basically the Cold War strategy, I’m even more amazed we all survived to face our next existential challenge.)

在“胆小鬼博弈”中,模型的结果则有些令人恐惧:在奖励不变的情况下,稳定点是每个人都避让,我们都能幸存。加入一点点变化,就会出现不避让的群体。加入更多的干扰,就会出现生存与碰撞之间切换的双稳态。(考虑到“胆小鬼博弈”基本上是冷战时期的战略,我更加惊讶我们竟然都幸存了下来,去面对下一个生存挑战。)

For rock-paper-scissors, an even more complex dynamic emerges. In terms of stable and unstable points, regular rock-paper-scissors has no stable points—the strategy never settles into everyone choosing rock, for instance. Instead, everyone keeps flipping continuously among the three options. If the rewards change randomly per round, however, new stable and unstable points emerge. Depending on the case, this can cause the game to more quickly evolve to the flipping strategy or develop a limit cycle. Limit cycles emerge if the rewards are uneven (for instance, rock versus scissors gives the winner a greater payoff than paper versus rock). In this case, the population continuously cycles, with the probability of choosing rock versus paper versus scissors evolving predictably and stably over time.

对于石头剪刀布,出现了更复杂的动态。就稳定点和不稳定点而言,常规的石头剪刀布没有稳定点——例如,策略永远不会固定在每个人都出石头上。相反,每个人都在这三个选项之间持续切换。然而,如果奖励每轮随机变化,就会出现新的稳定点和不稳定点。根据具体情况,这可能导致博弈更快地演化为切换策略,或形成极限环。如果奖励不均衡(例如,石头赢剪刀的收益高于布赢石头),就会出现极限环。在这种情况下,种群会持续循环,选择石头、布和剪刀的概率随时间以可预测且稳定的方式演化。

The conclusion from all this game theorizing is that, even though the behavioral tendencies of the players may influence the game, a varying game environment can have a huge effect on the optimal strategy.

所有这些博弈论研究得出的结论是:尽管玩家的行为倾向可能会影响博弈,但多变的博弈环境对最优策略有着巨大的影响。

Life revealed through simple rules

通过简单规则揭示的生活

If we take the prisoner’s dilemma as the classic game theory tool, it yields a depressing conclusion: Cooperation is a losing strategy. Yet even under the conditions set by the prisoner’s dilemma, we observe cooperation. Why? Because there are external influences on the reward—the prisoner who defects and is released may well be retaliated against under slightly different circumstances, and not under others. Therefore, a more complex mix of strategies is likely to emerge. However, it was still surprising to see that a relatively small amount of variation in the reward structure could lead to such strong changes in behavior.

如果我们把囚徒困境作为经典的博弈论工具,它会得出一个令人沮丧的结论:合作是一种失败的策略。然而,即使在囚徒困境设定的条件下,我们也观察到了合作。为什么?因为奖励受到外部影响——背叛并获释的囚犯在稍有不同的环境下可能会遭到报复,而在其他环境下则不会。因此,更有可能出现复杂的策略组合。然而,令人惊讶的是,奖励结构中相对较小的变化竟然能导致行为如此剧烈的改变。

Game-theory models are also often used to try and understand economic behavior. I’ve always been skeptical (perhaps overly skeptical) of the insights drawn from these games, and I think this paper makes it explicit why I distrust these models. But the results here also point a way forward, where the richer dynamics that we observe in real life are replicated in games that are still quite simple.

博弈论模型也常被用来试图理解经济行为。我一直对从这些博弈中得出的见解持怀疑态度(或许过于怀疑),我认为这篇论文明确说明了我为何不信任这些模型。但这里的结果也指明了一个前进的方向,即我们在现实生活中观察到的更丰富的动态,可以在依然相当简单的博弈中得到复现。