Cheating Nash’s Game Theory, with Quantum Mechanics

The curious math of ‘quantum games’.
‘Minimal’ examples, toys, and games are my favourite way to understand a complex topic or a problem. For example, when I studied for my quantum mechanics exams, I used to boil down abstract systems to their most reduced parts. Playing with the math in these simple cases helped me understand how quantum mechanics differs from simple classical mechanics in counter-intuitive ways.
Economists do this as well. The prisoners’ dilemma is a simple thought experiment to show that acting rationally doesn’t always lead to the best outcome for society. However, if the prisoners’ are allowed to make decisions on quantum mechanical objects instead (even though they still can’t communicate with each other), we can turn the conclusion of this game on its head, putting us back in the realm of Adam Smith’s invisible hand.

Game theory is the systematic study of finding the best strategies to maximise your benefit when playing a game. In classical game theory, agents can execute the moves they wish to choose, as long as the moves are in line with the game's rules. Quantum game theory is the study of what happens when instead, players try to manipulate quantum systems. Quantum systems behave probabilistically — there is chance involved — and so agents don’t have complete control over the choices they make. This simple change of assumptions leads to amazing, counter-intuitive results which depart from traditional game theory.
Instead of assuming that the players have to decide their moves, I will assume that they can manipulate electrons with some simple magnetic mechanism to play their games. This post will explore how traditional game dynamics get shaken up and how one of the well-known results of the Prisoner’s dilemma is changed.
In the final year of my undergraduate studies, I gave this talk at Universiti Malaya, and I wanted to share a simplified version on Medium.
The Prisoners’ Dilemma
The Prisoner's dilemma is a beloved ‘minimal’ example that economists have used to study competitions between rational actors. It is a simple yet enlightening game to show that two people acting in their rational best interest do not necessarily achieve the outcome that maximises their happiness. Rather it shows that they would’ve achieved a better outcome had they worked cooperatively.
The game is simple. Suppose we have two prisoners in crime, Bonnie and Clyde. They have been caught, and the police lock them up in separate rooms for interrogation so that they cannot communicate. They are given two options: they can get either rat out their partner or stay silent.
If they both stay silent, they both get a year in prison.
If one rats the other, but that other person stays silent, the rat is completely free, and the silent person gets 10 years in prison.
If both rat out each other, they both get two years in prison.
What would you do? There are two possibilities. In the first case, suppose the other person rats you out. If this happens, ratting the other person out is a better option than staying silent. In the second case, suppose the other person stayed silent. Again, ratting the other person out is the best option. So, in this case, the best strategy is for both players to rat out the other, and hence the rational strategy gets both of them two years in prison.
Ratting each other out is a ‘bad’ equilibrium. If both actors had behaved completely rationally, in the classical economic sense of the word. However, if they cooperated, they both would have achieved a better outcome. This kind of equilibrium where both players chose the best or ‘dominant’ strategy is called the Nash equilibrium. However, it is not the best outcome they could achieve.
Economists dub the best outcome as the Pareto efficient point. The prisoner’s dilemma is such a beautiful example because it showcases a scenario where the Pareto optimal point is not the Nash equilibrium. This is a stunning and elegant example to refute the idea of Adam Smith’s invisible hand, where rationality is always in line with maximising the utility of a group.
The question is, is there a way to hack this game so that behaving rationally is maximally beneficial for both actors? It may sound convoluted, but if two criminals were to modify the game somehow, they could avoid situations where they reach sub-optimal outcomes.
Mixing in quantum mechanics
In an arXiv article posted by J. Orlin Grabbe¹, a solution is presented for the prisoners’ dilemma to make rational choices correspond to the best outcome for all. The paper models a game exactly like the prisoners’ dilemma. The difference is, instead of having each player choose whether to defect or stay silent, both actors make their decisions by manipulating electrons or ‘qubits’. The paper concludes that had the actors came up with a clever strategy of manipulating each electron, each prisoner following the rational strategy again corresponds with the socially optimal point.
To describe the first-order mechanics of an electron, physicists use a theory called quantum mechanics. Quantum mechanics is a probabilistic description of the universe that models ‘small particles’ like electrons. In line with vast amounts of experiments performed on particles, the model accurately concludes that the position and momentum of particles don’t have a single determined ‘value’ per se. It doesn't make sense to say, ‘hey, look, the particle is exactly in this spot’. Particles can only be described with random distributions that only make sense when you repeatedly repeat an experiment.
A qubit is the simplest ‘quantum’ system there is. It is an object which has a notion of being two states. In nature, a photon can be polarised up or down, so we could use this as a qubit. Alternatively, an electron can either have an up spin or a down spin. As I said above, we model these objects probabilistically regarding what state they’re in at a given time. These objects are not naturally ‘set’ in either an up or down state — they are, in fact, in a superposition of these states, only collapsing only when observed.
If I repeat an experiment to measure an electron’s spin 100 times, then even if the state of the electron is primed the same way each time, it is possible that I can observe both up and down spins. Thus, quantum mechanics as a theory is powerful because it allows us to predict the frequency of whether an electron is up or down, and we can test this experimentally by repeating experiments.
The quantum prisoners’ dilemma
In the quantum version of the prisoners' dilemma, the players each have a ‘qubit’ to manipulate. In the theory of quantum computation, qubits are manipulated through gates, which modify their states. Loosely speaking, gates allow a player to change the ‘proportion’ of up and down spin in an electron before observing it. In Grabbe’s paper, each player has access to three different gates, replacing the choices of rat or stay silent in the traditional prisoners’ dilemma.
So, each player has three choices of gates to apply to their own electron. After these gates are applied, the ‘interrogator’ then observes the state of Bonnie’s electron and Clyde’s electron. The up or down spin then decides how much jail time they are assigned.
If they both are ‘up spin’, they both get a year in prison.
If one is ‘up spin’, but the other person is ‘down’, then the ‘down spin’ prisoner is completely free, and the ‘up’ person gets 10 years in prison.
If both are ‘down’ spin out each other, they both get two years in prison.
The game is the same as the traditional prisoners’ dilemma set up, but the prisoners’ choice is different. To calculate the best strategy, the paper computes the average payout each prisoner will achieve, corresponding to all the possible choices both players combined can make (there are 8 of them).
Over multiple trials of this experiment, it turns out that the optimal strategy, in this case, is for both players to play what is called the ‘Pauli Z-gate’. Amazingly, this optimal strategy aligns with the best payout both players can receive — so it is a Nash equilibrium and Pareto efficient. Hence, rationality once again is the socially optimal thing to do.
Why do we care?
In this case, clever use of quantum mechanical networks could allow two spies caught in the prisoners’ dilemma to have a mutually beneficial set-up if they were both caught and interrogated. In any case, I think this is a really nice example as a toy of how quantum mechanics can leak into other subjects like economics and change the conclusions. John Von Neumann, the founder of modern game theory, was both a physicist and an economist — so the connectedness of these subjects is undeniable.
References
[1] An Introduction to Quantum Game Theory- Grabbe, J. arXiv:quant-ph/0506219