BACKEXPLORATIONS
AV.N.E.519.001Exploration · Interactive research note

Rock-Paper-Scissors

Can a perfectly fair game be exploited?

In a single round, there is almost nothing to exploit. Rock beats Scissors; Scissors beats Paper; Paper beats Rock. No action is stronger than the others. In a series of rounds, however, another variable appears: memory. People remember what just won and what just lost. Against a perfectly unpredictable opponent the game remains fair. Against a person, not always.

The Rule to Try

To begin, try a simple rule. Arrange the three actions in a cycle. After a win, move one step right. After a loss, move one step left.

Paper
Scissors
Rock
then Paper again
You won → rightYou lost → left
You just played Paper

What happened in that round?

Try a few rounds. Then we will examine why such a simple rule can work against a person at all.

Why It Can Work

The rule can help only when the opponent’s next choice is influenced by what just happened. Two examples make the logic concrete.

01
What happened
You played Scissors. Your opponent played Rock. You lost, so the opponent just won.
Predicted opponent move
Rock again
Your counter-move
Paper - one step left from Scissors

Winners in the experiment repeated their successful action more often. If the opponent chooses Rock again, you need Paper. In the practical cycle, Scissors → Paper is one step left. The loss-left branch therefore has the clearest connection to winner-stay evidence.

02
What happened
You played Paper. Your opponent played Rock. You won, so the opponent just lost.
Predicted opponent move
A switch from Rock to Paper
Your counter-move
Scissors - one step right from Paper

The mnemonic assumes that the losing opponent will switch from Rock to Paper. Scissors beats that predicted Paper, so Paper → Scissors is one step right. This is the less reliable part of the rule. People often change behaviour after a loss, but the direction of that change varies between people and conditions.

54,000 Games

One evening at a table is not enough to see these dependencies clearly.

From 2010 to 2014, Zhijian Wang, Bin Xu and Hai-Jun Zhou ran a Zhejiang University experiment with 360 participants. They formed 60 groups of six. Every group completed 300 rounds; in each round a computer randomly formed three pairs.

60populations
300rounds
3games per round
54,000games conducted
Construction of the full experimental total

The experiment produced 54,000 pairwise games. One six-person population with atypical behaviour was excluded from the principal analysis, leaving 354 participants, 59 populations and 53,100 games represented by the retained groups.

Balance Is Not Randomness

A player can look unpredictable in the final proportions while the order of their actions reveals a clear pattern.

A player can choose each action one-third of the time while making every next action obvious. Balance describes the totals. Independence describes whether earlier events help predict what follows.

Rock · 1/3Paper · 1/3Scissors · 1/3
The same final balance
Rock
Paper
Scissors
Rock
Paper
Scissors
Rock
Paper
Scissors
Repeat
Equal proportions; complete sequential structure

A sequence may be perfectly balanced and perfectly predictable.

What Changes When Winning Matters More?

The researchers kept the rules the same and changed how much a win was worth relative to a tie.

In some groups a win was only slightly more valuable than a tie. In others it was substantially more valuable. Behaviour changed with that incentive structure. When the difference was small, participants repeated their previous action more often. When winning became noticeably more valuable, a loss was followed more often by a move in the direction Rock → Scissors → Paper → Rock. The authors called this directed change a clockwise shift. Even in this simple game, people did not respond in exactly the same way under every condition.

These diagrams define direction only. They do not imply that both directions were observed equally in the experiment. Counter-clockwise: Rock, Paper, Scissors, Rock. Clockwise: Rock, Scissors, Paper, Rock.

In the original paper, the relative value of winning was represented by parameter a. The directional post-loss effect was especially reported for conditions where a ≥ 2. This notation is included for scientific precision; the practical rule does not require it.

The Popular Rule and the Experiment Are Not Identical

The convenient rule and the experimental result must be kept separate.

Practical opening rule

Paper → Scissors → Rock, with a win moving right and a loss moving left, is easy to remember and easy to test in an ordinary game.

What the study reported

The study found real outcome-dependent tendencies: winners often stayed, while responses after losses varied with conditions. It did not show that every person follows the practical rule exactly, and the direction after a loss is especially variable.

Use the rule as a first model of the opponent, not as a law. Keep it while the person confirms it; change the model when they do not.

From Pattern to Advantage

In an ordinary game, there is no need to calculate dozens of probabilities. Begin with Paper → Scissors → Rock → Paper: using your own result, move right after a win and left after a loss. Then watch the person.

01

Observe

Watch what the opponent does after a win and after a loss. Do they repeat a winning action or switch after losing?

02

Predict

Name the action that now seems more likely. A useful pattern should produce a specific prediction for the next round.

03

Counter

Choose the action that beats your prediction. If several predictions fail, update the model and rely less on this rule.

ObservePredictCounterUpdate

The advantage comes from a pattern the opponent repeats, and disappears when that pattern changes.

When the Rule Stops Working

The strategy does not work against the game. It works against the predictability of a particular person.

  1. One opponent may repeat a winning object, react to losses in the same way and fail to notice the pattern.
  2. Another may use a different rule, while a short run may only look like a habit by chance.
  3. A third may notice that you are reading them and deliberately change strategy.
  4. A fourth may choose independently and nearly at random, leaving no useful pattern to counter.

Reduce confidence when the evidence weakens. Adaptation is part of the strategy, not a failure of it.

The Machine Learns You

Now reverse the roles. The machine does not know your next choice. It sees only what has already happened: what you do after a win, how you react to a loss, whether you repeat an object, and whether you move around the same cycle. The first rounds tell it little. If a stable pattern appears, it gradually begins to use it. The aim is not simply to beat the machine, but to make the previous round stop telling it about the next one.

Its move is committed before your current choice is received.

This session remains in this browser tab. No move history is transmitted or stored.

So, Does It Work?

Yes - against some people.

The logic is clearest after your loss: the opponent just won, and winners often repeat a successful action. Moving left selects the counter to that possible repeat.

After your win, the rule makes a bolder assumption about how the losing opponent will switch. People often change behaviour after a loss, but the exact direction is less universal.

Use the rule as a first model of the opponent, not as a law. Keep using the pattern while their behaviour supports it. If it does not, stop before you become the predictable player.

Conclusion

Rock-Paper-Scissors contains no hidden strongest move. The advantage appears elsewhere: in a person’s response to what just happened.

The experiment produced 54,000 games. After one atypical group was excluded, the principal analysis covered 59 groups and 354 participants, representing 53,100 games. In those data, the previous result affected the probability of the next decision, while the response also varied with conditions. As an opening model, moving left after your loss counters a possible repeat by the winner; moving right after your win is a more tentative bet on how the loser will switch. Keep the model only while the opponent’s behaviour supports it.

Do not look for a sequence that beats everyone. Look for the pattern of the person playing in front of you.

Limits of the Evidence

The evidence supports a narrower claim than the popular win-right / loss-left rule.

  1. The popular win-right / loss-left rule was not directly proved by Wang et al. as a universal winning algorithm.
  2. Population-level tendencies do not determine the behaviour of an individual.
  3. Laboratory payoff conditions are not identical to a casual real-life game.
  4. The exact direction of switching after a loss varies across studies and conditions.
  5. Adaptive players can notice exploitation and change their strategy.
  6. A fifty-round browser session cannot reliably estimate every behavioural strategy.
  7. Equal action frequencies do not guarantee independence.
  8. A deterministic ‘scientific winning cycle’ is exploitable once the opponent knows it.

The evidence does not establish a universal sequence that beats all human opponents, a fixed rule followed after every result, or a stable psychological trait inferable from one short session.

Sources

Primary evidence is separated from contextual further reading. Numbered markers in the text point here.

Primary studies

  1. Wang, Z., Xu, B. & Zhou, H.-J.

    Social cycling and conditional responses in the Rock-Paper-Scissors game.

    Scientific Reports 4, 5830 · 2014

    DOI: 10.1038/srep05830
  2. Dyson, B. J., Wilbiks, J. M. P., Sandhu, R., Papanicolaou, G. & Lintag, J.

    Negative outcomes evoke cyclic irrational decisions in Rock, Paper, Scissors.

    Scientific Reports 6, 20479 · 2016

    DOI: 10.1038/srep20479
  3. Sundvall, J. & Dyson, B. J.

    Breaking the bonds of reinforcement: Effects of trial outcome, rule consistency and rule complexity against exploitable and unexploitable opponents.

    PLOS ONE 17(2), e0262249 · 2022

    DOI: 10.1371/journal.pone.0262249
  4. Arai, K., Jacob, S., Widge, A. S. & Yousefi, A.

    Deviation from Nash mixed equilibrium in repeated rock–scissors–paper reflect individual traits.

    Scientific Reports 15, 14955 · 2025

    DOI: 10.1038/s41598-025-95444-6

Further research

  1. Brockbank, E. & Vul, E.

    Repeated rock, paper, scissors play reveals limits in adaptive sequential behavior.

    Cognitive Psychology 151, 101654 · 2024

    DOI: 10.1016/j.cogpsych.2024.101654