Abstract

In society, mutual cooperation, defection, and asymmetric exploitative relationships are common. Whereas cooperation and defection are studied extensively in the literature on game theory, asymmetric exploitative relationships between players are little explored. In a recent study, Press and Dyson demonstrate that if only one player can learn about the other, asymmetric exploitation is achieved in the prisoner's dilemma game. In contrast, however, it is unknown whether such one-way exploitation is stably established when both players learn about each other symmetrically and try to optimize their payoffs. Here, we first formulate a dynamical system that describes the change in a player's probabilistic strategy with reinforcement learning to obtain greater payoffs, based on the recognition of the other player. By applying this formulation to the standard prisoner's dilemma game, we numerically and analytically demonstrate that an exploitative relationship can be achieved despite symmetric strategy dynamics and symmetric rule of games. This exploitative relationship is stable, even though the exploited player, who receives a lower payoff than the exploiting player, has optimized the own strategy. Whether the final equilibrium state is mutual cooperation, defection, or exploitation, crucially depends on the initial conditions: Punishment against a defector oscillates between the players, and thus a complicated basin structure to the final equilibrium appears. In other words, slight differences in the initial state may lead to drastic changes in the final state. Considering the generality of the result, this study provides a new perspective on the origin of exploitation in society.

Highlights

  • Equality is not achieved in society; instead, inequality among individuals is common

  • IV, we investigate the dynamics of the learning process in depth to demonstrate that a slight difference in initial strategies between the players is amplified, and this symmetry breaking results in a large payoff difference, i.e., exploitation

  • We emphasize that the equilibrium state for a repeated game is denoted by the subscript e, but it is unrelated to the equilibrium of learning dynamics discussed in the following subsection

Read more

Summary

INTRODUCTION

Equality is not achieved in society; instead, inequality among individuals is common. With the change in strategies of the players through learning, we check whether “symmetry breaking” can occur when individuals have symmetric capacities and environmental conditions For this analysis, we adopt the celebrated prisoner’s dilemma game. In the prisoner’s dilemma game, the emergence and sustainability of cooperation, even though defection is any individual player’s best choice, has been extensively investigated [1,2]. We study the well-known prisoner’s dilemma (PD) game (see Fig. 1 for the payoff matrix), in which each of two players, referred to as players 1 and 2, chooses to cooperate (C) or defect (D).

Repeated game for fixed strategies
Learning dynamics of strategies
Intuitive interpretation of the model
ANALYSIS OF LEARNING EQUILIBRIUM
Characterization of the exploitative relationship
TRANSIENT DYNAMICS TO THE LEARNING EQUILIBRIUM
Characterization of transient dynamics
Basin structure for exploitative state
SUMMARY AND DISCUSSION

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.