Stable but Stuck: When Nash Equilibria Produce Bad Outcomes
Reading Eliezer Yudkowsky’s Inadequate Equilibria sent me back to a familiar problem: systems in which many participants can recognise a better outcome, yet no one can reach it safely alone. This is often where the language of Nash equilibrium enters the discussion. It is a useful concept, but only if it remains narrower than our ordinary idea of a system being “stuck”.
In John Nash’s original formulation, an equilibrium is a combination of strategies in which no player can improve their own payoff by changing strategy alone while everyone else’s strategy remains fixed. The definition says nothing about whether the outcome is fair, efficient, desirable or unique. Nor does it require the world to stop changing. It identifies a form of strategic stability: given what everyone else is doing, unilateral deviation does not pay.
That distinction explains how individually defensible choices can assemble themselves into a collectively poor result. The obstacle may not be ignorance, malice or an inability to imagine something better. It may be that improvement requires several actors to move together, while the first person to move alone accepts the cost and leaves the benefits available to everyone else.
The Equilibrium Is Not the Verdict
The standard prisoner’s dilemma gives the cleanest demonstration. In its one-shot form, two players independently choose whether to cooperate or defect. Defection gives each player a better payoff regardless of what the other chooses, so mutual defection is the unique Nash equilibrium even though mutual cooperation would leave both better off. Rationality at the level of the individual does not automatically compose into the best result for the pair.
The model is powerful because it isolates that mechanism, but its cleanliness can mislead when it is carried into the world unchanged. Governments, companies and individuals rarely meet once, anonymously, with only two available actions. They remember past conduct, make promises, build reputations and punish breaches. In iterated versions of the prisoner’s dilemma, future responses can make cooperation strategically sustainable under some conditions. Repetition does not guarantee trust, but it changes the price of betrayal.
Game theory also provides a way to describe how costly decentralised choice can become. The price of anarchy is usually expressed as a ratio comparing the social cost of a poor equilibrium with that of the system optimum. A value close to one means that self-interested choices perform nearly as well as coordination; a larger value indicates a greater collective penalty. The concept measures inefficiency, but it does not turn every equilibrium into a failure. Some games have efficient equilibria, some have several equilibria of sharply different quality, and many persistent social problems are not Nash equilibria in any useful formal sense.
That limitation matters. “Nash equilibrium” should not become an impressive synonym for any arrangement that is difficult to change. The concept identifies why unilateral movement is unattractive within a specified game. Moral and political judgement still has to come from somewhere else, as does the decision about whether the model describes the real situation well enough to help.
Nuclear Deterrence and Traffic: Analogy Versus Model
The Cold War arms race resembles a prisoner’s dilemma strongly enough to make the comparison tempting. A state that disarms while its rival retains a secure arsenal may expose itself to coercion or attack, so both sides have reasons to remain armed even while sharing an interest in avoiding unlimited expenditure and catastrophic war. Yet the analogy quickly reaches its limits. Nuclear strategy involves repeated interaction, unequal arsenals, alliances, domestic politics, imperfect information, first- and second-strike capabilities, accidental escalation and many possible degrees of armament rather than a single choice between cooperation and defection.
Thomas Schelling’s work on conflict and cooperation became influential partly because it treated such confrontations as mixed-motive situations. Opponents seek advantage, but they may also share an overriding interest in preventing the contest from destroying them both. Arms-control agreements, verification regimes, communication channels and changes to force posture do not remove self-interest. They alter which moves are credible, observable and survivable.
Traffic offers a more literal application of equilibrium reasoning. In the standard nonatomic model, each driver selects the route that appears to minimise their own travel time. At equilibrium, no single driver can improve the journey by changing routes, yet the combined flow need not minimise total travel time. Each person accounts for the delay they experience, not the additional congestion they impose on everyone behind them.
Braess’s paradox sharpens the point. Under particular network and demand conditions, adding a road can change drivers’ incentives so that the new equilibrium leaves every route slower. The extra connection is attractive to each driver when considered individually; once enough drivers make the same choice, the network performs worse. No one needs to be irrational or badly informed. The road changes the game, and the new equilibrium rewards behaviour that collectively defeats the purpose of building it.
Commons: The Rules Are Part of the Resource
Open-access fisheries create a related trap. A fisher who leaves fish in the sea bears the immediate cost of restraint but cannot be sure of receiving the future benefit, because another vessel may catch the fish first. That uncertainty encourages each participant to harvest sooner and invest in more capacity simply to protect their share. Collectively excessive effort can then reduce both the stock and the long-term return.
It is important, however, not to confuse a common-pool resource with an absence of rules. A fishery may be difficult to exclude people from and vulnerable to depletion, but “common” does not have to mean unmanaged. Elinor Ostrom’s research on governing common resources documented communities that developed durable arrangements involving clear boundaries, monitoring, graduated sanctions and participation in rule-making. Her work challenged the assumption that shared resources must either collapse, be privatised or be controlled entirely from the centre.
The deeper lesson is institutional. Participants in a social dilemma are not limited to playing the game they inherited; they may also create rules that change it. Monitoring makes restraint visible. Sanctions reduce the advantage of defecting. Defined access prevents outsiders from appropriating the gains produced by cooperation. A different equilibrium becomes possible because the resource is now embedded in a different system of expectations and consequences.
Positional Races: Work, Advertising and Credentials
Some strategic traps are driven less by the consumption of a shared resource than by relative position. Advertising can take this form when firms continue spending partly because reducing their visibility while competitors maintain theirs risks a loss of market share. Advertising also informs consumers and builds brands, so it cannot be treated as pure waste. The positional element appears when much of the expenditure is required to avoid falling behind rather than to create a comparable amount of new value.
Workplace busyness can follow similar logic. When long hours are interpreted as proof of commitment, leaving earlier may damage an employee’s prospects even when the additional time produces little. A field experiment on tournament incentives found that high-stakes competition could increase both work time and effort. That does not establish that every culture of overwork is a Nash equilibrium, or that longer hours are never productive. It does show how relative-performance rewards can make visible sacrifice individually prudent even when an organisation would prefer employees to compete on output rather than endurance.
Education and professional credentials can produce another positional race. Michael Spence’s job-market signalling model showed how education could convey information about a worker’s underlying ability even in a model where schooling did not itself raise productivity. A costly signal can be useful to the employer precisely because different types of applicant find it differently difficult to obtain.
Bryan Caplan advances a much stronger empirical claim in The Case against Education: that a large share of formal education functions as a signal of intelligence, diligence and conformity rather than as direct skill formation. I found the argument more persuasive than I expected, though its scale remains contested and its policy conclusions cannot simply be exported from the American system to countries with different tuition costs, labour markets and vocational routes.
If credentials carry a substantial signalling function, escalation becomes possible. A degree distinguishes an applicant while it is scarce; once it becomes ordinary within a candidate pool, an additional qualification may be needed to recreate the same distinction. Each applicant may be sensible to acquire it, and each employer sensible to use it as a filter, while society spends more time and money to produce roughly the same ranking. The individual cannot safely leave the race merely because the race appears wasteful.
Climate Change Is More Than One Game
Climate mitigation contains a global coordination problem on an extraordinary scale. Greenhouse-gas reductions benefit people beyond the borders of the country paying for them, so national incentives to mitigate may be weaker than the global incentive. A government can hope to enjoy reductions made elsewhere while avoiding some of the domestic political and economic cost of contributing itself.
The IPCC therefore discusses climate change as a global-commons and free-rider problem, but it also warns against treating that as the only useful frame. Climate policy is simultaneously a technological transition, a conflict over the distribution of costs and benefits, a development problem and a struggle over infrastructure that can lock societies into particular energy systems for decades.
That broader framing changes the strategic picture. Clean-energy investment may bring domestic benefits through lower air pollution, reduced exposure to imported fuels, industrial development and technological leadership. Subsidies, standards and early deployment can reduce technology costs, making later adoption easier for other countries. Conversely, poorly designed policies can concentrate costs on groups with little capacity to absorb them and provoke resistance even when the long-term national benefit is positive.
International cooperation remains necessary, but it need not take the form of every country accepting the same sacrifice at the same moment. Climate clubs, common standards, technology transfer, finance and trade arrangements can create benefits for participation or costs for free-riding. The aim is not only to persuade governments to behave more altruistically. It is to reshape the available payoffs until deeper cooperation becomes compatible with domestic interests.
Changing the Game
Equilibrium thinking can sound fatalistic because it begins from a condition in which no actor benefits from moving alone. The same definition, however, points towards the exits. Change the payoffs and previously unattractive behaviour may become rational: congestion charges can make drivers account for delay imposed on others, while fishing rights or enforceable quotas can give participants a durable interest in future stocks. Taxes, subsidies and liability rules serve the same general purpose when they bring private incentives closer to social costs.
Coordination changes what “alone” means. Treaties, unions, industry standards and collective agreements allow several actors to move together rather than demanding that one participant accept a unilateral disadvantage. Monitoring and enforcement matter because promises that cannot be verified do little to alter expectations. Repeated relationships can also support cooperation when reputation and future retaliation make a short-term gain from defection less attractive.
Sometimes the most effective intervention is to create a strategy that the original game did not contain. Better energy storage can weaken an old trade-off between variable generation and grid reliability. Route pricing can distribute traffic in ways that unaided route choice will not. A hiring process that evaluates demonstrated work rather than relying mechanically on credentials can reduce the value of educational escalation. Institutional design is often less about instructing people to choose differently than about giving them choices whose consequences fit the outcome the system is meant to produce.
This is why moral appeals and isolated acts of virtue so often disappoint. Telling drivers to avoid congestion, employees to stop performing busyness, fishers to leave more in the water or governments to prioritise global welfare does not remove the disadvantage faced by the first actor who complies alone. People may be behaving badly, but they may also be responding accurately to the incentives around them.
The more useful question is therefore not simply why participants refuse to choose the better collective outcome. It is what must change before that choice becomes individually sustainable. Stability is not proof that a system works well; it shows that the available unilateral moves do not offer a better result to the person making them. A bad equilibrium becomes escapable when the first better move is no longer a sacrifice made alone.
Comments
Post a Comment