Abstract
Limited flexibility in behaviour gives rise to behavioural consistency, so that past behaviour is partially predictive of current behaviour. The consequences of limits to flexibility are investigated in a population in which pairs of individuals play a game of trust. The game can either be observed by others or not. Reputation is based on trustworthiness when observed and acts as a signal of behaviour in future interactions with others. Individuals use the reputation of partner in deciding whether to trust them, both when observed by others and when not observed. We explore the effects of costs of exhibiting a difference in behaviour between when observed and when not observed (i.e. a cost of flexibility). When costs are low, individuals do not attempt to signal that they will later be trustworthy: their signal should not be believed since it will always pay them to be untrustworthy if trusted. When costs are high, their local optimal behaviour automatically acts as an honest signal. At intermediate costs, individuals are very trustworthy when observed in order to convince others of their trustworthiness when unobserved. It is hypothesized that this type of strong signalling might occur in other settings.
Keywords: behavioural consistency, signalling, trust, cost of flexibility, evolutionary game theory, personality
1. Introduction
Individuals often show behavioural consistency, for example, in levels of aggression, over related but distinct circumstances [1]. There may be situations in which there is selection pressure to be consistent [2]. However, in many cases, evidence suggests that this limited plasticity is maladaptive, so that its existence requires an explanation. For example, female spiders might require a different response to a prey item they should eat than to a male with whom they should mate. They do not always seem to be able to differentiate between these two circumstances, and may be left unmated having eaten a series of possible mates [3]. This suggests their lack of flexibility is producing suboptimal behaviour, although optimal behaviour under limited information cannot be completely ruled out [4].
The number of subtly distinct circumstances encountered in the real world is vast. Evolution shapes behavioural mechanisms that must deal with this vast range. It is not clear that it is even possible to define what we mean by ‘optimal’ in the real world [5–7]. Even if we could, we would not expect evolved behaviour strategies to be exactly optimal. An optimal strategy would necessarily be enormously complex as it would have to respond flexibly and appropriately to every possible circumstance. The psychological mechanism needed to implement such a strategy probably could not evolve in the first place, and even if it did it would involve so much neuronal machinery that the maintenance costs of this machinery [8] would outweigh the benefit. For this reason, we expect the evolution of strategies implemented by psychological mechanisms of limited complexity that perform well on average but may not be exactly optimal in any circumstance [9]. It is inevitable that such mechanisms show some consistency because of their limited complexity.
It is not straightforward to predict the form of the psychological mechanisms that will evolve in the real world [9], and we make no attempt to do so here. Instead, we are concerned with the implications of lack of flexibility on the predictions of social interactions. We focus on the consequences of a lack of flexibility when there is reputation: how does the previous behaviour of a social partner in one circumstance (circumstance O) predict their behaviour in a different circumstance (circumstance U).
Suppose that individuals were completely flexible in their behaviour, responding optimally to every circumstance. Then behaviour in the two circumstance would not be correlated; individuals in circumstance U would just do the best given the circumstance, irrespective of what they did before. Thus the reputation of an individual, which acts as a summary of past behaviour, is only predictive if there is limited flexibility. When reputation is predictive, others will respond to this reputation. This then exerts a selection pressure on individuals to change their reputation by changing their behaviour, so as to affect how others respond to them in the future. Thus the degree to which reputation is predictive is crucial to the evolution and ontogeny of social behaviour. We examine the effect of the degree of flexibility on model predictions in a simple model of a game of trust. In order to capture the idea of limited flexibility, we assume that flexibility has costs [10–12]. This approach may be thought of as a simple expedient (see Discussion). In our model, reputation formation acts as a signal of future behaviour. As we will explain, costs of flexibility mean that there is a cost of deviating from the behaviour that is ‘promised’ by the signal. For instance, studies suggest that lying, a form of deception, is cognitively more demanding than telling the truth [13,14].
In the trust model, individuals pair up to play a sequence of trust games with other population members. In each game, one of the pair is randomly designated as Player 1 and the other as Player 2. Player 1 decides whether to trust their partner based on this partner’s reputation for trustworthiness. If they do trust them, Player 2 then decides whether to be trustworthy or not. Each game occurs in one of the two circumstances, O and U. Both pair members are aware of the current circumstance. In circumstance O, the interaction is observed by all other population members. In circumstance U, the interaction is only observed by the two pair members, and others do not later learn about the interaction. Thus a Player 2’s reputation for trustworthiness is based solely on their behaviour in circumstance O, but this reputation must nevertheless be used to decide whether to trust this individual in circumstance U. Player 2 pays a fitness cost for being flexible that increases with the difference in the level of trustworthiness in circumstances O and U. We vary this cost, showing that intermediate costs can produce qualitatively different behaviour to the extreme cases of zero cost (perfect flexibility) and infinite cost (no flexibility). We also investigate the cost of flexibility to Player 1.
2. The model
We consider a very large (essentially infinite) asexual population with discrete, non-overlapping generations. In each generation, there is a series of many decision epochs. At each epoch population, members are paired at random and play a trust game. This trust game has been previously analysed by a number of authors [15–18] and has the following structure (figure 1). One player is randomly assigned to the role of Player 1 and the other to Player 2. In the first stage of the game, Player 1 can either trust or reject Player 2. If Player 2 is rejected, then both individuals receive payoff s and the game ends. If Player 2 is trusted, then the game proceeds to its second stage. At the second stage, Player 2 can either cooperate with Player 1 or defect, after which the game ends. If Player 2 cooperates, both players receive payoff r. If Player 2 defects Player 1 gets nothing while Player 2 receives payoff 1. We assume that 0 < s < r < 1. For this payoff structure, it is easily seen that for a single round of this game with no reputation effects, if Player 2 is trusted they should defect. Given this, Player 1 should never trust Player 2. Thus we would predict an equilibrium at which Player 1 rejects Player 2 (and Player 2 defects if ever trusted). (This is the subgame-perfect equilibrium.) Both players then receive a payoff of s, whereas they would have both received the large payoff r had Player 1 trusted Player 2 and Player 2 cooperated. The game can thus be seen as a sequential version of the prisoner’s dilemma game.
Figure 1.
The structure and payoffs for the trust game.
In this paper, each game is played in one of two circumstance: with probability α the game is open to observation (O) by all population members, with probability 1 − α it is not observed (U) by other population members. All population members are aware of their current circumstance. The behaviour of a Player 2 is described by the two genetically determined parameters po and pu that specify the probabilities of cooperating if trusted in the two circumstances. We also refer to these quantities as the Player’s trustworthiness in the two circumstances. We assume that there have been many previous rounds of games, so that a Player 1 has been able to accurately estimate the trustworthiness of a Player 2 when observed. Thus, in a particular game Player 1 knows the Player 2’s po value, so that this value acts as the signal of Player 2’s behaviour if trusted. Since behaviour in circumstance U is not observed, Player 1 does not know pu, and can only infer its value from the value of po. The behaviour of Player 1 is described by the two genetically determined parameters zo and zu. When the game is observed Player 2 is only trusted if po ≥ zo. When the game is unobserved Player 2 is trusted only if po ≥ zu, so that whether Player 2 is trusted when unobserved is based on the previous behaviour of Player 2 when in role 2 and observed.
An individual has four genes specifying zo, zu, po and pu. The fitness of an individual is the long-term mean reward per game multiplied by the factor
| 2.1 |
Here, the constants kz and kp determine the cost of being flexible when in role 1 and 2, respectively. To consider evolutionary dynamics, we take the values of each of the parameters zo, zu, po, pu to lie on a discrete grid. We then compute how the proportion of population members at each grid point combination changes over generations. The population is followed over many generations until proportions settle down to equilibrium values. The results presented are for this equilibrium. A detailed specification of the dynamics and computational procedure is given in the electronic supplementary material.
3. Results
When an interaction is observed, a Player 1 should trust a Player 2 if and only if their probability of cooperating satisfies po ≥ s/r [17]. Thus in order to be trusted in future when observed, the Player 2 should have a probability of cooperating when observed such that po ≥ s/r. We refer to s/r as the trustworthiness threshold. Since the payoff to Player 2 if trusted decreases with their trustworthiness (because r < 1), we would not expect Player 2’s probability of cooperating to exceed this threshold unless there are reputational reasons to do so. Thus we expect po = s/r unless Player 2s are behaving to impress future partners that they are trustworthy when not observed.
In our evolutionary simulations, there is a discrete grid of values. The probability of acceptance is also smoothed close to acceptance thresholds (electronic supplementary material). Thus instead of the precise values predicted, evolved mean values satisfy zo < s/r < po (figure 2a,c). There is variation about these means maintained by mutation (electronic supplementary material): mean values reflect a mutation–selection balance. The selection pressure is, however, asymmetric. For example, if zo is low, then the probability of wrongly trusting a Player 2 is small since few Player 2s have low po values. By contrast, if zo > s/r then there is a high probability of wrongly rejecting a Player 2. As a consequence of these asymmetric selection pressures, the mean of zo may be significantly below s/r and the mean of po may be significantly above this threshold. This results in almost all Player 2s being trusted when observed (figure 2e). These results hold for all values of the cost of flexibility.
Figure 2.
Evolved mean population characteristics as a function of the cost of flexibility for Player 2, kp (shown on a logarithmic scale). There is no cost of flexibility for Player 1 (kz = 0). (a) Acceptance threshold for Player 1 when observed. (b) Acceptance threshold for Player 1 when unobserved. (c) Probability Player 2 cooperates when observed. (d) Probability Player 2 cooperates when unobserved. (e) Probability Player 2 is trusted when observed. (f) Probability Player 2 is trusted when not observed. (g) Probability Player 2 cooperates when trusted (observed). (h) Probability Player 2 cooperates when trusted (unobserved). Solid line: α = 0.3. Dashed line: α = 0.7. Dotted line: critical acceptance threshold . s = 0.24, r = 0.60.
Figure 2 considers the case in which flexibility by Player 2 is costly but flexibility by Player 1 is not. When the cost is small (kp < 0.1), Player 2s are only sufficiently trustworthy to be trusted when observed (po just greater than the trustworthiness threshold , figure 2c). When not observed they are rather untrustworthy (pu less than , figure 2d), so that Player 1s set such a high acceptance threshold (figure 2b) that Player 2s are rarely trusted (figure 2f).
When the cost to Player 2 is intermediate (0.1 < kp < 1), a Player 2 with po just greater than s/r is liable to have pu < s/r and should not be trusted when unobserved. By contrast, if po is much greater than s/r, then we might expect pu > s/r since otherwise the Player 2 would incur high costs. Thus only a Player 2 with high po should be trusted when unobserved. This means that there is selection on Player 2s to be very trustworthy when observed in order to convince Player 1s that they are also trustworthy when unobserved. Our computations bear out this intuitive argument (figure 2c).
When the flexibility cost is large (kp > 1) Player 2s are again only sufficiently trustworthy to be trusted when observed (po just greater than s/r, figure 2c). Since Player 2s pay a high cost from flexibility most have approximately the same probability of being trustworthy when not observed as when observed (po ≈ pu, figure 2c,d). Thus Player 1s apply the same acceptance threshold in the two circumstances (zu ≈ zo, figure 2a,b), so that Player 2s tend to be trusted when not observed (figure 2f ), and are indeed reasonably trustworthy; i.e. those that are trusted unobserved have a probability of being trustworthy that exceeds s/r (figure 2h)
Our central result—that trustworthiness when observed should be highest at intermediate costs—appears to be robust under changes in parameter values. In particular, it still occurs when s and r are changed, holding the ratio s/r at 0.4 (electronic supplementary material, figure S1), and it occurs when s/r is changes (electronic supplementary material, figure S2). The result is also predicted by a simple analytic model (electronic supplementary material, figures S4, S5, S6).
Lowering α increases the proportion of interactions that are unobserved and hence the importance to Player 2s of being trusted when unobserved. Thus, as figure 2c and electronic supplementary material, figures S1(c) and S2(c) illustrate effects become stronger. The analytic model of electronic supplementary material also illustrates that effects are weaker for large α (electronic supplementary material, figures S4, S5, S6).
In order to highlight the effects of costs of flexibility for Player 2, in figure 2, we assumed no cost of flexibility for Player 1. It seems reasonable, however, that there will be similar flexibility costs in both roles. Figure 3 illustrates the case where the two costs are equal. When costs are small they have little effect, so that results are similar in figures 2 and 3. At intermediate cost, there is selection to make zo and zu similar. Since Player 1s need to be very choosy when unobserved, with flexibility costs they are more choosy when observed (compare figures 2a and 3a), so that Player 2s are even more trustworthy when observed (compare figures 2c and 3c). This also means that more of those that are trusted when observed are actually trustworthy (compare figures 2g and 3g), with this effect being particularly strong when most interactions are observed (α = 0.7). Finally, at high costs, since zo and zu were already similar in the kz = 0 case, results are little changed when a cost to flexibility of Player 1 is introduced.
Figure 3.
Results corresponding to figure 2 when there are also costs of flexibility to Player 2, with kz = kp.
4. Discussion
In social interactions, the best action of an individual typically depends on that of other individuals. Consequently, when an individual must choose its action before the behaviour of others is known, any information that can help predict their behaviour will be advantageous. Here, we have investigated how past behaviour provides information. We have shown that limited flexibility in behaviour results in individuals having to exhibit high levels of cooperation (trustworthiness) when other can observe them so that they will be trusted in future interactions that do not have an audience, and so cannot affect their reputation. This conclusion might also apply to other phenomena. For example, the propensity to be angry or retaliate as a result of an infringement when observed might be correlated with the level of this behaviour when not observed. For intermediate levels of flexibility, one might then expect a very strong reaction when observed in order to credibly signal that the individual will retaliate if there is an infringement when not observed.
When the past behaviour of an individual partially predicts current behaviour a feedback loop is set up: population members evolve to pay attention to reputations and respond to them; which then selects for individuals to change their reputation through changing their behaviour in order to change how others respond to them [19,20]. When reputation is important, the action of an individual has two effects: (i) it gives an immediate payoff, and (ii) it affects future payoffs through its effect on the response of others to the individual in the future. There is often a trade-off in the choice of action, and it may be advantageous to reduce the current payoff if this is more than made up for in terms of increased future payoffs. In our example, the payoff to a Player 2 in a single round is maximized by the player being as untrustworthy as possible (always defecting). Behaving in this way when observed does not evolve because it results in a poor reputation, leading to rejection by future partners and hence a poor future payoff.
The previous behaviour of an individual will often have been observed in circumstances that are distinct, but related to, the current circumstances. The value of reputation then depends crucially on the consistency of behaviour across contexts. In our example, the behaviour of a Player 2 when observed acts as signal about their behaviour when not observed. When an individual is not observed their behaviour cannot affect their reputation and hence affect future payoffs. Thus when there are no costs of or constraints on flexibility, any signal is worthless: if trusted when not observed it is always best defect, maximizing the current payoff.
When there are sufficient costs to being flexible, the situation is different. The trustworthy behaviour of a Player 2 when observed can then signal that the trustworthiness of this player when unobserved exceeds the trustworthiness threshold, and should be trusted. It is costly to produce this signal, since the payoff to a Player 2 when trusted decreases with their trustworthiness, but this initial cost is not relevant to the honesty of the signal. Rather it is the cost of a future deviation from what is promised that is critical: if a Player 2 signals high trustworthiness then it would not be in the interests of the Player 2 to be too untrustworthy when unobserved. The situation is therefore what Kirmani & Rao [21] refer to as default-contingent signalling; that is signalling in which a cost is paid if the signaller deviates from the behaviour that is indicated by the signal. This signalling paradigm has been highlighted in a biological context by Enquist [22] and Hurd [23].
When the cost of flexibility is high in our trust model, a Player 2 that is only slightly more trustworthy than the trustworthiness threshold when observed is giving a reliable signal they will exceed this threshold when not observed. When costs are intermediate it requires a stronger signal when unobserved to convey the same message. Finally, for low costs, it would still be worth a Player 2 that was completely trustworthy when observed to be below the trustworthiness threshold when unobserved, so that Player 2s no longer attempt to signal their behaviour when unobserved.
The existence and strength of costs of flexibility in behaviour has been debated (e.g. [10–12]). Although costs plays a central role in the analysis of the specific model we have presented, we see lack of flexibility rather than the effect of costs as the key to the real biology of social reputations. In our model, costs are just a way of constraining flexibility in a model that assumes that evolution produces optimal behaviour. We have argued, however, that the real world is too complex to produce exactly optimal behaviour, and that a lack of flexibility is an inevitable consequence of the evolution of mechanisms. Thus a more realistic model would evolve some class of behavioural mechanism in a suitably complex environment [9]. A lack of flexibility would then automatically arise, even without costs. For example, it might arise through the internalization of social norms [24–26]. Such models should be investigated in the future. Detailed results might then depend on the class of mechanisms and environments considered, although we would expect our main conclusion that signals should be amplified to still hold.
Supplementary Material
Acknowledgements
We thank Pat Barclay, Andy Higginson, Alasdair Houston, Olof Leimar, Sasha Dall and an anonymous reviewer for their comments on a previous version of this manuscript.
Data accessibility
This article has no additional data.
Competing interests
We declare we have no competing interests.
Funding
Z.B. was financed by the Higher Education Institutional Excellence Programme (NKFIH-1150-6/2019) of the Ministry of Innovation and Technology in Hungary, within the framework of the DE-FIKP Behavioural Ecology Research Group thematic programme of the University of Debrecen.
References
- 1.Sih A, Bell A, Johnson JC. 2004. Behavioral syndromes: an ecological and evolutionary overview. Trends Ecol. Evol. 19, 372–378. ( 10.1016/j.tree.2004.04.009) [DOI] [PubMed] [Google Scholar]
- 2.Wolf M, Van Doorn G, Weissing FJ. 2011. On the coevolution of social responsiveness and behavioural consistency. Proc. R. Soc. B 278, 440–448. ( 10.1098/rspb.2010.1051) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Arnqvist G, Henriksson S. 1997. Sexual cannibalism in the fishing spider and a model for the evolution of sexual cannibalism based on genetic constraints. Evol. Ecol. 11, 255–273. ( 10.1023/A:1018412302621) [DOI] [Google Scholar]
- 4.Fawcett TW, Fallenstein B, Higginson AD, Houston AI, Mallpress DEW, Trimmer PC, McNamara JM. 2014. The evolution of decision rules in complex environments. Trends Cogn. Sci. 18, 153–164. ( 10.1016/j.tics.2013.12.012) [DOI] [PubMed] [Google Scholar]
- 5.Savage LJ. 1972. The foundation of statistics. New York, NY: Dover Publications. [Google Scholar]
- 6.Todd PM, Gigerenzer G. 2000. Précis of Simple heuristics that make us smart. Behav. Brain Sci. 23, 727–780. ( 10.1017/S0140525X00003447) [DOI] [PubMed] [Google Scholar]
- 7.Binmore K. 2009. Rational decisions. Princeton, NJ: Princeton University Press. [Google Scholar]
- 8.Bullmore E, Sporns O. 2012. The economy of brain network organization. Nat. Rev. Neurosci. 13, 336–349. ( 10.1038/nrn3214) [DOI] [PubMed] [Google Scholar]
- 9.McNamara JM, Houston AI. 2009. Integrating function and mechanism. Trends Ecol. Evol. 24, 670–675. ( 10.1016/j.tree.2009.05.011) [DOI] [PubMed] [Google Scholar]
- 10.Auld JR, Agrawal AA, Relyea RA. 2010. Re-evaluating the costs and limits of adaptive phenotypic plasticity. Proc. R. Soc. B 277, 503–511. ( 10.1098/rspb.2009.1355) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.DeWitt TJ, Sih A, Wilson DL. 1998. Costs and limits of phenotypic plasticity. Trends Ecol. & Evol. 13, 77–81. ( 10.1016/S0169-5347(97)01274-3) [DOI] [PubMed] [Google Scholar]
- 12.Relyea RA. 2002. Costs of phenotypic plasticity. Am. Nat. 159, 272–282. ( 10.1086/338540) [DOI] [PubMed] [Google Scholar]
- 13.Suchotzki K, Verschuere B, Van Bockstaele B, Ben-Shakhar G, Crombez G. 2017. Lying takes time: a meta-analysis on reaction time measures of deception. Psychol. Bull. 143, 428–453. ( 10.1037/bul0000087) [DOI] [PubMed] [Google Scholar]
- 14.Ofen N, Whitfield-Gabrieli S, Chai XJ, Schwarzlose RF, Gabrieli JDE. 2017. Neural correlates of deception: lying about past events and personal beliefs. Soc. Cogn. Affect. Neurosci. 12, 116–127. ( 10.1093/scan/nsw151) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Frank RH. 1988. Passions within reason. New York, NY: Norton. [Google Scholar]
- 16.Guth W, Kliemt H. 2000. Evolutionarily stable cooperative commitments. Theory Decis. 49, 197–221. ( 10.1023/A:1026570914311) [DOI] [Google Scholar]
- 17.McNamara JM, Houston AI. 2002. Credible threats and promises. Phil. Trans. R. Soc. Lond. B 357, 1607–1616. ( 10.1098/rstb.2002.1069) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.McNamara JM, Stephens PA, Dall SRX, Houston AI. 2009. Evolution of trust and trustworthiness: social awareness favours personality differences. Proc. R. Soc. B 276, 605–613. ( 10.1098/rspb.2008.1182) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.McNamara JM, Doodson P. 2015. Reputation can enhance or suppress cooperation through positive feedback. Nat. Commun. 6, 6134 ( 10.1038/ncomms7134) [DOI] [PubMed] [Google Scholar]
- 20.McNamara JM, Leimar O. 2020. Game theory in biology. Oxford, UK: Oxford University Press. [Google Scholar]
- 21.Kirmani A, Rao AR. 2000. No pain, no gain: a critical review of the literature on signaling unobservable product quality. J. Market. 64, 66–79. ( 10.1509/jmkg.64.2.66.18000) [DOI] [Google Scholar]
- 22.Enquist M. 1985. Communication during aggressive interactions with particular reference to variation in choice of behaviour. Anim. Behav. 33, 1152–1161. ( 10.1016/S0003-3472(85)80175-5) [DOI] [Google Scholar]
- 23.Hurd PL. 1995. Communication in discrete action-response games. J. Theor. Biol. 174, 217–222. ( 10.1006/jtbi.1995.0093) [DOI] [Google Scholar]
- 24.Akçay E, Van Cleve J. In press. Internalizing cooperative norms in group-structured populations. In Social cooperationand conflict: biological mechanisms at the interface (eds W Wilczynski, S Brosnan). Cambridge, UK: Cambridge University Press.
- 25.Gavrilets S, Richerson PJ. 2017. Collective action and the evolution of social norm internalization. Proc. Natl Acad. Sci. USA 114, 6068–6073. ( 10.1073/pnas.1703857114) [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Gintis H. 2003. The Hitchhiker’s guide to altruism: gene-culture coevolution, and the internalization of norms. J. Theor. Biol. 220, 407–418. ( 10.1006/jtbi.2003.3104) [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
This article has no additional data.



