Skip to main content
eLife logoLink to eLife
. 2026 Mar 31;14:RP106840. doi: 10.7554/eLife.106840

The self-interest of adolescents overrules cooperation in social dilemmas

Xiaoyan Wu 1,2,3,†, Hongyu Fu 1,2,†, Gökhan Aydogan 4, Chunliang Feng 5, Shaozheng Qin 1,2, Yi Zeng 2,6,7,8,9, Chao Liu 1,2,✉
Editors: Xiaosi Gu10, Andre F Marquand11
PMCID: PMC13038262  PMID: 41913590

Abstract

Cooperation is essential for success in society. Research consistently showed that adolescents are less cooperative than adults, which is often attributed to underdeveloped mentalizing that limits their expectations of others. However, the internal computations underlying this reduced cooperation remain largely unexplored. This study compared cooperation between adolescents and adults using a repeated Prisoner’s Dilemma Game. Adolescents cooperated less than adults, particularly after their partner’s cooperation. Computational modeling revealed that adults increased their intrinsic reward for reciprocating when their partner continued cooperating, a pattern absent in adolescents. Both computational modeling and self-reported ratings showed that adolescents did not differ from adults in building expectations of their partner’s cooperation. Therefore, the reduced cooperation appears driven by a lower intrinsic reward for reciprocity, reflecting a stronger motive to prioritize self-interest, rather than a deficiency in predicting others’ cooperation in social learning. These findings provide insights into the developmental trajectory of cooperation from adolescence to adulthood.

Research organism: Human

eLife digest

In everyday life, people often face choices between pursuing their own interests and cooperating with others. Cooperation helps individuals achieve shared goals and build positive relationships, but it often requires sacrificing some immediate personal benefit. Adolescence is a critical period during which young people learn how to manage friendships and collaborate in groups. Many studies have shown that teenagers tend to cooperate less than adults, but the reasons for this remain unclear. Do adolescents struggle to recognise when others are willing to cooperate, or do they recognise these intentions but choose to prioritise their own benefit?

Wu et al. aimed to understand why adolescents cooperate less than adults, both in their behaviour and in the decision-making processes that underlie it. Specifically, they examined whether adolescents fail to recognise cooperative behaviour from others, or whether they do recognise it but are more tempted than adults to take advantage of the situation for personal gain.

To investigate this, the researchers compared the behaviour of teenagers and adults in a repeated cooperation game. In this game, two players could either cooperate for mutual benefit or attempt to gain more for themselves at their partner’s expense. The results showed that teenagers cooperated less than adults, particularly after their partner had just cooperated. Importantly, teenagers and adults were equally accurate in estimating how cooperative their partner was. This suggests that adolescents recognise when others are willing to cooperate but feel less motivated to reciprocate.

These findings may help teachers, parents, and those designing school programmes better support teenagers’ social development. The study of Wu et al. suggests that it may be useful not only to help adolescents understand others’ intentions, but also to strengthen the value they place on fairness and on reciprocating cooperation when others behave kindly. Future research should explore whether similar patterns occur in real-life interactions and across more diverse groups of young people.

Introduction

Cooperation among individuals facilitates the achievement of shared goals and enhances overall group efficiency (Fehr and Fischbacher, 2003; Nowak, 2006). For individuals, cooperation skills are key to success in society; this ability is not innate but gradually acquired through socialization (Warneken, 2018). Successful cooperation requires individuals to prioritize the common purpose over their personal interests, focusing on collective goals (Sachs et al., 2004). Experimental psychology has often used the Prisoner’s Dilemma Game (PDG; Axelrod and Hamilton, 1981) to study human cooperative behaviors. Extensive research has adapted the PDG into a repeated version to explore how people respond to interactive cooperation (Andreoni and Miller, 1993; Embrey et al., 2018), requiring individuals to adjust their responses dynamically to others and simulating real-life cooperation more closely (Axelrod and Hamilton, 1981). In such social dilemmas, individuals face a trade-off between immediate rewards from defection and long-term benefits from cooperation (Rilling et al., 2002). Decision-making in these situations is thought to engage mentalizing abilities, which are functions related to theory of mind that enable individuals to form expectations about others’ cooperative intentions (Rilling et al., 2004).

Cooperation is not an innate skill but is gradually cultivated and refined through socialization (House et al., 2020). Adolescence, in particular, marks a critical developmental phase in the transition to independent social roles (Steinberg, 2005). Studies using the PDG consistently show that adolescents cooperate less than adults (Belli et al., 2012; Nava et al., 2023; Taheri et al., 2018). This reduced cooperation is often attributed to an underdeveloped theory of mind, which may lead adolescents to underestimate others’ trustworthiness and willingness to cooperate in social learning (Gutiérrez-Roig et al., 2014; Fett et al., 2014).

However, there are findings that may not support this hypothesis. For example, a previous study found that adolescents’ lower cooperation, compared to adults, emerges only when following a partner’s cooperation. Conversely, when the partner defected, adolescents’ cooperative behaviors resembled those of adults (Gutiérrez-Roig et al., 2014). Similarly, a Trust Game study (Fett et al., 2014) reported a comparable pattern: adolescents invested less (measured as trust behavior) than adults only when the partner was cooperative. When faced with a non-cooperative partner, both adolescents and adults consistently reduced their trust behaviors. These findings suggest that adolescents, like adults, are able to adjust their behavior in response to others’ actions. This selective reduction in adolescent cooperation implies that factors beyond deficits in mentalizing may be at play. Adolescents may prioritize maximizing immediate rewards over long-term reciprocity (Rilling et al., 2002). When confident that their partner will cooperate, defection may become the optimal strategy for maximizing self-interest. This hypothesis remains untested, but computational modeling could provide a valuable approach for examining the underlying mental processes behind these behavioral variations (Farrell and Lewandowsky, 2010; Wu et al., 2024a; Wu et al., 2024b).

This study aimed to investigate variations in cooperative behavior between adolescents and adults and to explore the mental processes underlying these differences using computational modeling. Based on legal criteria for majority and prior empirical work, we adopt 18 years as the boundary between adolescence and adulthood (Icenogle et al., 2019; Tervo-Clemmens et al., 2023). A total of 127 adolescents and 134 adults participated in the study, playing a repeated Prisoner’s Dilemma Game (rPDG) with a presumed human partner, whose behavior was predetermined by a computer program (see Figure 1a). The program ensured consistent conditions across age groups. To enhance realism, variability was introduced into the computer-simulated partner’s behavior. The rPDG provides a symmetric and simultaneous framework that isolates the motivational conflict between self-interest and joint welfare, avoiding the sequential trust and reputation dynamics characteristic of asymmetric tasks such as the Trust Game (Rilling et al., 2002; King-Casas et al., 2005). Based on the standard payoff matrix of the rPDG (Figure 1b), mutual cooperation maximizes collective interests, while defection maximizes self-interest from an individual perspective. Our focus was on how adolescents respond to their partner’s consistent cooperation and defection, aiming to identify potential mental variables contributing to adolescents’ lower cooperation.

Figure 1. Experiment setup and behavioral results.

Figure 1.

(a) Partner’s cooperation probability: in the first half of the 120 trials, the partner cooperated 78% of the time; in the second half, cooperation alternated between 20% and 80%. (b) Payoff matrix: payoffs are 4 for mutual cooperation, 2 for mutual defection, 0 for cooperation when the other defects, and 6 for defecting when the other cooperates. (c) Trial illustration: after a 0.5 s fixation, participants choose a shape (triangle for cooperation, square for defection) within 4 s and see both players’ choices for 1.5 s. (d, e) Post hoc comparisons: d and e show the participants’ cooperation probability on the y-axis. The x-axis represents the consistency of the partner’s actions in previous trials (t-1: last trial, t-1,2: last two trials, t-1,2,3: last three trials). Large red (adolescents) and blue (adults) dots indicate mean probabilities, with black error bars for standard error (SE). Gray dots represent mean probabilities across trials, and green error bars show predicted cooperation rates with SE. Notes: n.s.p>0.05; *p<0.05; **p<0.01; ***p<0.001.

We developed computational models to investigate the dynamic variables guiding cooperative decisions in the rPDG. The model explicitly incorporates both expectations of the partner’s cooperation and the intrinsic reward of reciprocity. A basic reinforcement learning (RL) algorithm was used to model participants’ dynamic expectations regarding the partner’s cooperation. Drawing on research on asymmetric reward learning in adolescents (Palminteri et al., 2016; Rosenbaum et al., 2022), we included asymmetric updating for positive (better-than-expected) and negative (worse-than-expected) outcomes. We identified the asymmetric RL learning model as the winning model that best explained the cooperative decisions of both adolescents and adults. Participants’ expectations were modeled as a trial-by-trial dynamic variable, represented by parameter p. Following previous studies (Fareri et al., 2012; Fareri et al., 2015), a non-monetary reward for cooperation, represented by parameter ω, reflects individual preferences for mutual cooperation. The term p×ω quantifies the intrinsic reward for reciprocity.

We hypothesize that adolescents will exhibit lower overall cooperation compared to adults, consistent with previous studies (Belli et al., 2012; Nava et al., 2023; Taheri et al., 2018). Specifically, we expect adolescents to demonstrate reduced cooperation after their partner’s cooperation but not following defection (Gutiérrez-Roig et al., 2014; Fett et al., 2014). Furthermore, we aim to explore whether this lower conditional cooperation is driven by inappropriate expectations of their partners (represented by p), a reduced intrinsic reward for reciprocity (represented by p×ω), or a combination of both.

Results

Adolescents exhibit lower cooperation than adults following partner cooperation, but not defection

In each trial of the rPDG, as shown in Figure 1c, participants were presented with two choices: a triangle representing cooperation and a square representing defection. The choice associated with each symbol was randomly balanced across participants. They were informed that they were playing the game simultaneously with another partner. After making their decision, participants were shown both their own choice and that of their partner. We performed a generalized linear mixed model (GLMM1) analysis (see Appendix 1—table 1) to examine the effects of each independent variable and their interactions on the decision to cooperate or defect.

Consistent with most previous studies (Belli et al., 2012; Nava et al., 2023; Taheri et al., 2018), adolescents cooperated less than adults (b of group = 0.79, 95% CI = [0.311, 1.270], p = 0.001; Figure 1d). Following the interaction of group × previous trial × partner’s choice (b of interaction = 0.24, 95% CI = [0.126, 0.361], p < 0.001), we found that adolescents showed significantly less cooperation compared to adults only after the partner’s cooperation (t(259)group=−2.84, p = 0.005, BF10=6.01).

However, such a difference was not significant after the partner’s defection (t(259)group=−1.86, p=0.064, BF10=0.69; Figure 1d). We also found that adults increased cooperation in response to their partners’ consistent cooperation (the partner cooperated once vs. the partner cooperated thrice: t(266)adults=−2.50, p = 0.013, BF10=2.56), but this pattern was not observed in adolescents (t(252)adolescents=−1.18, p = 0.239, BF10=0.27, see Figure 1d).

Nevertheless, both groups significantly decreased cooperation in response to the partner’s continual defection (the partner defected once vs. the partner defected twice: t(266)adults=4.46, p < 0.001, BF10 >103, t(252)adolescents=2.78, p = 0.006, BF10=5.21; the partner defected once vs. the partner defected thrice: t(266)adults=5.56, p < 0.001, BF10 >103 for adults, t(252)adolescents=4.32, p < 0.001, BF10=761.12 for adolescents, Figure 1e).

Asymmetric RL learning in the social reward model best explains cooperative decisions of adolescents and adults

Computational modeling was used to simulate participants’ mental processes during the rPDG. Starting with a baseline model that assumed decisions were made through random selection (Model 1), we compared several alternatives: a win-stay and loss-shift model (Model 2), a reward learning model (Model 3), an inequality aversion model (Model 4), and a social reward model (Model 5). Among these, the social reward model outperformed the others. We then compared a basic RL algorithm (Model 6), an influence learning rule (Model 7), and an asymmetric RL learning rule (Model 8) within the social reward framework. The asymmetric RL learning model best explained the cooperative decisions of both adolescents and adults (see Figure 2a for adolescents and Figure 2b for adults; methods for details). Model recovery analysis indicated that the asymmetric RL learning within the social reward model was distinguishable from the other models (Figure 2c) and accurately captured the behaviors of both adolescents (Figure 2d) and adults (Figure 2e). The overlap between Models 4 and 5 likely arises because neither model incorporates a learning mechanism, making them less able to account for trial-by-trial adjustments in this dynamic task. For further validation of the best-fitting model, see Appendix 1—figure 1 for model predictions, Appendix 1—figure 2 for the distributions of free parameters, Appendix 1—figure 3 for parameter recovery, Appendix 1—figure 4 for partial correlation matrices among parameters, and Appendix 1—figure 5 for group-level posterior distributions from the hierarchical Bayesian estimation for the best-fitting model.

Figure 2. Computational modeling.

Figure 2.

(a, b) Model comparisons for adolescents and adults, respectively. The y-axis represents model fitness based on the Akaike Information Criterion with a correction for sample size (AICc; Hurvich and Tsai, 1989). For each participant, the model with the lowest AICc served as a reference to compute ΔAICc by subtracting it from the AICc of other models (ΔAICc=AICcx−AICclowest). A lower ΔAICc indicates a better model fit. Protected exceedance probability (PEP) is a group-level measure that assesses the likelihood of each model’s superiority over the others (Rigoux et al., 2014). (c) Model recovery analysis. Each model was used to generate 100 synthetic datasets, and for each dataset, model fitting and comparison were performed. Each column corresponds to one generative model, and each row corresponds to one fitting model. The color in each cell indicates the probability that the synthetic datasets generated by the model in the column were best fit by the model in the row, with a darker color denoting a higher probability. (d, e) Model prediction. Sample illustration of the best-fitting model prediction versus data for adolescents and adults, respectively.

Distinct learning rates and social preferences between adolescents and adults in repeated cooperation

Although the asymmetric RL learning in the social reward model best explained the behaviors of both adolescents and adults, the two groups exhibited distinct learning dynamics and social preferences for cooperation. Specifically, adolescents applied a higher positive learning rate (α+, t(259)=2.95, p = 0.003, BF10=8.02, Figure 3a) to update better-than-expected prediction errors, and a lower negative learning rate (α−, t(259)=−2.62, p = 0.009, BF10=3.46, Figure 3b) for worse-than-expected prediction errors.

Figure 3. Learning rates and social preferences.

Figure 3.

(a–d) Comparison between adolescents and adults for positive learning rate (α+), negative learning rate (α−), social preference (ω), and inverse temperature (β), respectively. (e–h) Correlation between age and positive learning rate, negative learning rate, social preference, and inverse temperature, respectively. Notes: *p<0.05; **p<0.01.

Additionally, a positive correlation was found between participants’ age and the negative learning rate (α−, r=0.21, p < 0.001, Figure 3f), while no significant correlation was observed with the positive learning rate (α+, r=−0.09, p = 0.16, Figure 3e).

Furthermore, adolescents displayed a weaker preference for cooperation compared to adults (ω, t(259)=−3.03, p = 0.003, BF10=9.92, Figure 3c), and their social preferences for cooperation increased with age (r=0.20, p < 0.001, Figure 3g). Additionally, adolescents exhibited a higher inverse temperature parameter compared to adults, indicating they were more sensitive to utility differences between cooperation and defection (β, t(259)=2.14, p = 0.034, BF10=1.17, Figure 3d). This sensitivity decreased with age, as shown by a negative correlation with age (r=−0.17, p = 0.007, Figure 3h).

Adolescents compared to adults show no inappropriate expectations but less intrinsic reward for reciprocity

To further explore what underlies the observed decrease in cooperation among adolescents, we focused on two hidden trial-by-trial updating variables: the partner cooperation expectation (p) and the intrinsic reward for reciprocity (p×ω). Additionally, participants’ self-reported cooperativeness scores, assessed every 15 trials, provided further insight into their subjective estimation of the partner’s willingness to cooperate.

Partner cooperation expectation

We performed a linear mixed model (LMM1, Appendix 1—table 2) on partner cooperation expectation to assess the effects of each independent variable and their interactions. Following the interaction of group × previous trial × partner’s choice (b of interaction = 0.03, 95% CI = [0.022, 0.038], p < 0.001), we found that the partner cooperation expectation for both adolescents and adults increased with the partner’s consistent cooperation (the partner cooperated once vs. the partner cooperated twice: t(252)adolescents=−2.81, p = 0.005, BF10=5.75, t(266)adults=−4.45, p < 0.001, BF10 >103; the partner cooperated once vs. the partner cooperated thrice: t(252)adolescents=−3.69, p < 0.001, BF10=78.00, t(266)adults=−6.23, p < 0.001, BF10 >103; Figure 4a). Additionally, expectations decreased with the partner’s consistent defection (the partner cooperated once vs. the partner cooperated twice: t(252)adolescents=4.44, p < 0.001, BF10 >103, t(266)adults=7.02, p < 0.001, BF10 >103; the partner cooperated once vs. the partner cooperated thrice: t(252)adolescents=5.60, p < 0.001, BF10 >103, t(266)adults=8.40, p < 0.001, BF10 >103; Figure 4b). These results showed that both adolescents and adults held very similar expectations toward their partner’s cooperation and did not have significant differences between the groups (b of group = –0.04, 95% CI = [–0.102, 0.021], p = 0.198).

Figure 4. Analysis of hidden variables from the best-fitting model.

Figure 4.

(a, b) Post-hoc comparison of LMM1: interaction of group × previous trial × partner’s choice. The y-axis shows participants’ expectations of partner cooperation probability (p) from the best-fitting model. (c, d) Self-reported cooperativeness: normalized scores on partner cooperativeness for two orders of partner cooperation probability, with adolescents (orange-red line) and adults (blue line). Scores were assessed on a 0–9 scale and normalized to 0–1. The dotted line indicates the presumed partner’s cooperation probability, with mean values and standard errors shown. (e, f) Post-hoc comparison of LMM3: interaction of group × previous trial × partner’s choice. The y-axis shows participants’ intrinsic reward for reciprocity (p×ω) from the best-fitting model. The x-axis represents the consistency of the partner’s actions in previous trials (t-1: last trial, t-1,2: last two trials, t-1,2,3: last three trials). Colored dots with error bars indicate mean values with standard errors for adolescents (orange-red) and adults (blue), while small gray dots represent individual participants. Notes: n.s.p>0.05; *p<0.05; **p<0.01; ***p<0.001.

Moreover, we performed an LMM2 (Appendix 1—table 3) analysis on participants’ self-reported scores regarding the cooperativeness of their partners to examine the effects of each independent variable and their interactions. In line with the expectation of partner cooperation, we observed minimal discrepancy in the self-reported scores on partner cooperativeness between adolescents and adults. Neither the main effect of group nor the interaction achieved statistical significance (b of group = 0.17, 95% CI = [–0.51, 0.85], p = 0.616; b of interaction = 0.38, 95% CI = [–0.052, 0.812], p = 0.085; Figure 4c–d). These results provide evidence that adolescents did not differ from adults in assessing their partner’s cooperation.

Intrinsic reward for reciprocity

We performed an LMM3 (Appendix 1—table 4) on the intrinsic reward for reciprocity to assess the effects of each independent variable and their interactions. We found that adolescents appreciated reciprocity less than adults did (b of group = 0.52, 95% CI = [0.224, 0.816], p < 0.001).

Following the interaction of group × previous trial × partner’s choice (b of interaction = 0.37, 95% CI = [0.318, 0.424], p < 0.001), unlike adults, adolescents did not increase their intrinsic reward for reciprocity in response to the partner’s consistent cooperation (the partner cooperated once vs. the partner cooperated twice: t(252)adolescents=−0.96, p = 0.336, BF10=0.21, t(266)adults=−2.13, p = 0.034, BF10=1.15; the partner cooperated once vs. the partner cooperated thrice: t(252)adolescents=−1.38, p = 0.170, BF10=0.34, t(266)adults=−3.08, p = 0.002, BF10=11.63; Figure 4e).

Moreover, under consistent defection by the partner, evidence for the one-versus-two last trials comparison was inconclusive in adolescents but supported a decrease in adults (t(252)adolescents=1.99, p = 0.047, BF10=0.90, t(266)adults=−2.71, p = 0.007, BF10=4.27). Importantly, in the one-versus-three last trials comparison, both adolescents and adults consistently showed a decrease in intrinsic reward for reciprocity (t(252)adolescents=2.64, p = 0.009, BF10=3.66, t(266)adults=3.37, p < 0.001, BF10=27.47; Figure 4f).

In brief, adolescents did not deviate in forming expectations about their partner’s willingness to cooperate, but they showed lower social preferences for cooperation and a reduced intrinsic reward for reciprocity. Specifically, compared to adults, adolescents displayed less intrinsic reward for reciprocity and did not increase it in response to consistent cooperation, although their reactions to consistent defection tended to be similar to those of adults.

Discussion

Cooperation lies at the heart of societal functioning, facilitating the achievement of shared goals and fostering social harmony. In this study, we sought to deepen our understanding of the developmental aspects of cooperation by examining differences in cooperative behavior between adolescents and adults in the context of the rPDG. Our findings shed light on the cognitive and affective processes underlying these behaviors, offering insights into the mechanisms driving cooperative decision-making across different developmental stages.

Consistent with many previous studies (Fett et al., 2014; Gutiérrez-Roig et al., 2014; Westhoff et al., 2020), our results showed that adolescents exhibited lower levels of cooperation compared to adults. However, such lower cooperation was not generally observed during the task, but selectively occurred after their partner cooperated in the previous rounds. Moreover, our results showed that adults increased cooperation in response to their partner’s consistent cooperation; such a pattern was not observed in adolescents. However, both age groups decreased cooperation in response to consistent partner defection, indicating shared responses to non-cooperative behavior.

Our results suggest that the lower levels of cooperation observed in adolescents stem from a stronger motive to prioritize self-interest rather than a deficiency in predicting others’ cooperation in social learning. In both, the expectation of partner’s cooperation estimated from computational modeling and the self-reported measurements, adolescents did not exhibit significant differences from adults. However, adolescents exhibited a weaker preference for (conditional) cooperation compared to adults, resulting in a reduced intrinsic reward for reciprocity. The results are consistent with prior research (Crone and Dahl, 2012; Do et al., 2017; Pfeifer and Berkman, 2018; van den Bos et al., 2010, van den Bos et al., 2011), suggesting that adolescents prioritize immediate gains over long-term benefits, potentially undermining the benefits of cooperation. This tendency aligns with earlier findings that adolescents exhibit heightened sensitivity to reward feedback (Blakemore and Mills, 2014; Crone and Dahl, 2012; Davis et al., 2023; Do et al., 2017; van den Bos et al., 2011; van Duijvenvoorde et al., 2015), which may influence their decision-making in cooperative interactions. Overall, these findings indicate that adolescents’ lower cooperation is unlikely to be driven solely by strategic considerations, but may instead reflect differences in the valuation of others’ cooperation or reduced motivation to reciprocate. Although defection is the payoff-dominant strategy in the PDG, the selective pattern of adolescents’ cooperation and the model comparison results indicate that their reduced cooperation cannot be fully explained by strategic incentives, but rather reflects weaker valuation of social reciprocity.

It has been acknowledged that individuals update positive and negative outcomes by different weights in social cooperation, and such asymmetric learning process can be modeled by a basic RL algorithm with both positive and negative learning rates (Garrett and Daw, 2020; Rosenbaum et al., 2022). In this study, we find that an asymmetrical RL algorithm in a social reward model provided best model fits of the behaviors of both adolescents and adults. Adolescents demonstrated a larger positive learning rate, but a smaller negative learning rate compared to adults, suggesting heightened sensitivity to positive feedback from cooperative behavior and reduced sensitivity to negative feedback from defection. This asymmetrical learning pattern may drive adolescents to focus more on self-beneficial social signals, maximizing immediate gains in response to cooperative behavior. These findings align with (van den Bos et al., 2011), which highlight adolescents’ heightened sensitivity to immediate rewards and less stable trusting behavior compared to adults. Adolescents also showed higher inverse temperature values (β), indicating greater sensitivity to expected value and more value-based choice behavior. Together, these findings suggest that the differentiation between positive and negative learning rates changes with age, reflecting more selective feedback sensitivity in development, while higher (β) values in adolescents indicate greater value sensitivity. This interpretation remains tentative and requires further validation in future research.

Adolescence is characterized by increased self-discovery and egocentrism (Pfeifer and Berkman, 2018; Ting et al., 2019), leading individuals to prioritize immediate gains over long-term benefits. Consistent with this, the higher value sensitivity (β) observed in adolescents suggests a stronger focus on immediate utility during cooperative exchanges. Consequently, adolescents may be more inclined toward self-serving motives in sustained social interactions (Pfeifer and Berkman, 2018). However, these tendencies are not static; as individuals mature into adulthood, their socio-emotional capacities continue to develop (Worthman and Trang, 2018), enabling a more balanced integration of short-term rewards and long-term social outcomes (Crone and Dahl, 2012; Wu et al., 2023).

It is important to note some limitations of this study. First, we used artificial opponents with pre-determined cooperation patterns to better control the stimuli. While this approach allowed us to isolate specific motivations for cooperation (financial vs. social rewards), it is possible that participants might behave differently in more natural settings. Our study serves as an initial step in understanding cooperation motivations in adolescents and adults, and future research could explore these behaviors in more real-world contexts. Second, our study employed the rPDG as the primary task to directly capture cooperation in symmetric multi-round interactions. However, because it is a zero-sum framework that structurally incentivizes defection as the dominant strategy, the rPDG may influence choices beyond participants’ intrinsic preferences. For example, one potential interpretation of adolescents’ lower cooperation is that they adopt a strategic response to the payoff structure, through leveraging defection as the more rewarding strategy within the game. If this account holds, adolescents should exhibit lower cooperation across all rounds. However, we find that adolescents and adults exhibit similar behavioral patterns when partners defect. By contrast, adolescents cooperate less than adults when partners cooperate, and their cooperation does not increase significantly even when partners cooperate consecutively. Although this pattern is consistent with the interpretation that adolescents’ lower cooperation reflects a relatively more self-interested motivation, stronger conclusions about age differences in cooperative preferences require further examination in tasks with varied structures. Third, although both age groups were recruited from Beijing and nearby regions, minimizing major regional and cultural variation, adolescents and adults may still differ in socioeconomic status, financial independence, and social experience. Such contextual differences could interact with developmental processes in shaping cooperative behavior and reward valuation. Future research with demographically matched samples or explicit measures of socioeconomic background will help disentangle biological from sociocultural influences.

In conclusion, our study contributes to an understanding of the developmental aspects of cooperation and the cognitive-affective processes underlying cooperative decision-making. By examining differences in cooperative behavior between adolescents and adults in the rPDG and integrating computational modeling, we offer valuable insights into the mechanisms driving cooperative behavior across different developmental stages. These findings have implications for promoting prosocial behaviors and designing effective socialization interventions during adolescence. By highlighting the importance of reciprocity, our findings offer insights into the developmental trajectory of cooperation from adolescence to adulthood and provide practical implications for enhancing cooperative interactions in real-world contexts.

Materials and methods

Participants

A total of 261 participants took part in the current study, consisting of 127 adolescents (n=127, aged 14–17 years, mean ± SD: 16.13±0.63, 44 females) and 134 adults (n=134, aged 18–30 years, mean ± SD: 21.63±2.88, 79 females). No a priori power analysis was conducted. The sample size was determined based on previous studies investigating cooperation behaviors in adolescents and adults. Adolescents were recruited from a local high school, and adults were recruited through advertisements on a university campus forum. Written informed consent was obtained from all adult participants. For adolescents, written informed consent was obtained from their legal guardians, and assent was obtained from the adolescents themselves prior to participation. Participants were included if they had normal or corrected-to-normal vision and no history of psychiatric or neurological illness. Exclusion criteria included any self-reported diagnosed psychiatric or neurological disorder. No participants dropped out of the experiment, and all collected data were included in the statistical analyses. This study was approved by the Ethics Committee of Beijing Normal University (Approval Nos. CNL_A_0001_009 and RB_A_0003_202001). All procedures were conducted in accordance with the Declaration of Helsinki. Participants received monetary compensation based on their task performance (see rPDG for details).

Experimental procedure

All participants completed the experiments in a laboratory setting with multiple participants present. They were informed that they were participating in a multiple-round interaction game with an anonymous partner. In the instructions section, we referred to the interaction game as the rPDG and refrained from using the terms ‘cooperate’ and ‘defect’ to minimize the influence of social expectations, biases, and promote comparability across studies. Participants were instructed to believe that their partner was also playing the game at the same time. Compensation for their participation was based on the tokens earned during the game, with 10% of the rounds randomly selected for payment calculation at an exchange rate of 1 token to 1 yuan. Participants were explicitly informed in advance about this incentive mechanism. Prior to the formal experiment, participants underwent a quiz and several practice rounds to ensure a full understanding of the task. Following the experiment, participants completed a Social Value Orientation (SVO) task to assess their prosocial personality traits. The entire procedure lasted approximately 60 min. Blinding was not applicable in this study, as all participants interacted with a computer-controlled partner. To minimize potential bias, the partner’s behavior patterns and stimulus meanings were randomized across participants. A detailed protocol is available upon request.

The repeated prisoner’s dilemma game

Similarly to the classic version of PDG, rPDG involves two players. Consistent with the standard payoff matrix of the PDG (Figure 1b), when both players cooperated (defected), they each received four tokens (two tokens). If the players made different decisions, the one who cooperated received 0 tokens, while the one who defected received 6 tokens. Participants were told that their partner was another human participant in the laboratory and that they would interact with the same partner across all rounds. However, in reality, the actions of the partner were predetermined by a computer program. This setup allowed for a clear comparison of the behavioral responses between adolescents and adults. Participants were not informed of the total number of rounds in the rPDG. In order to enhance the realism of the partner’s response, we manipulated the variability in the partner’s decision making. The partner’s cooperation probability remained stable at 78% for half of the trials. In the other half of the trials, the partner’s cooperation probability varied, switching between 20%, 80%, and 20% for each set of 20 trials. The order of these two sessions was counterbalanced between participants. During the rPDG, participants were asked every 15 rounds to evaluate their partner’s cooperativeness using a 10-point scale, where 0 represents ‘no cooperation’ and 9 represents ‘very high cooperation’. The question posed to the participants was ‘How cooperative do you think your partner is at the moment?’.

Behavioral data analysis

All statistical analyses were conducted in MATLAB R2023a (RRID:SCR_001622). GLMM was implemented using the ‘fitglme’ function in MATLAB. Interaction contrasts were performed for significant interactions and, when higher-order interactions were not significant, pairwise or sequential contrasts were performed for significant main effects. Post hoc comparisons were conducted using Bayes factor analyses with MATLAB’s bayesFactor Toolbox version v3.0, with a Cauchy prior scale σ=0.707 (Krekelberg, 2024).

GLMM1: Participant’s choices (cooperate or defect) of all trials are the dependent variable; fixed effects include an intercept, the main effects of group (adolescents or adults), previous trial (last one trial, last two trials, and last three trials), partner’s choice (cooperation or defection), and all possible interaction effects of the independent variables. Gender (male and female) and timing (trial number from 1 to 120) were also included as the control variables. Random effects include correlated random slopes of group, previous trial, partner’s choice, gender, trial number, and random intercept for participants. The group, previous trial, partner’s choice, and gender are the category variables. The trial number is a continuous variable. See Appendix 1—table 1 for the statistical results of GLMM1.

LMM2: Participants’ self-reported score on partner’s cooperativeness is the dependent variable; fixed effects include an intercept, the main effects of group (adolescents or adults), the order of the sessions (regarding the partner’s cooperation involved fixed 78% cooperation probability, followed by shifting into 20%, 80%, and 20% for each 20, or vice versa), the interaction of group ×order. Gender (male and female) and timing (trial number from 1 to 120) were also included as the control variables. Random effects include correlated random slopes of group, gender, timing, and random intercept for participants. The group, previous trial, partner’s choice, and gender are the category variables. The trial number is a continuous variable. See Appendix 1—table 3 for the statistical results of LMM2.

Behavioral modeling

We systematically developed models based on various assumptions regarding participants’ decision-making processes in the rPDG.

Model 1: the baseline model

We modeled each participant’s choices in each trial (i.e. whether to cooperate) as outcomes from a Bernoulli distribution, where the cooperation probability is controlled by a parameter, b∈[0,1]. For each participant, the probabilities of cooperation (q(cooperation)) and defection (q(defection)) are denoted as follows:

q(cooperation)=b (1)
q(defection)=1−b (2)

Model 2: win-stay and loss-shift model

The model assumes that individuals adopt a tit-for-tat strategy in decision-making. Participants are likely to repeat their previous choice with a probability of 1−ε2 if they won, and ε2 if they lost in the last trial, where ε represents the choice variability. Winning and losing are defined based on the payoff outcomes of 4 or 6 (win) and 0 or 2 (loss), respectively.

qt+1=qt(1−ε2)δ+qt(ε2)(1−δ) (3)

where qt denotes the probability of repeating the previous choice and 1−qt denotes the probability of shifting to another option at trial t.

Model 3: reward learning model

This model assumes that participants make decisions by comparing the values of choosing cooperation and defection. The values of the two options are updated using an RL algorithm:

Vct+1=Vct+α(Rt−Vct) (4)
Vdt+1=Vdt+α(Rt−Vdt) (5)

where Vc (Vd) denotes the value of cooperation (defection) option. R represents the reward feedback, which can be 0, 2, 4, or 6, depending on the payoff matrix. Rt−Vct (Rt−Vdt) represents the reward prediction error for the cooperation (defection) option, and α is the learning rate. Participants’ choices are modeled by a softmax function:

q(cooperate)t=11+eβ(Vdt−Vct) (6)

where qt denotes the participants’ probability of cooperation and β denotes the inverse temperature. The lower the value of the inverse temperature, the greater the sensitivity to the different values between options.

Model 4: inequality aversion model

The model assumes that participants’ decisions aim to reduce both disadvantageous and advantageous inequality between themselves and their partners:

Uct=cself−φmax(cother−cself,0)−νmax(cself−cother,0) (7)
Udt=dself−φmax(dother−dself,0)−νmax(dself−dother,0) (8)

where Uc (Ud) denotes the utility of cooperation (defection). φ represents aversion to disadvantageous inequality and ν represents aversion to advantageous inequality. cself and cother denote the expected payoffs for cooperation to oneself and the partner, respectively, while dself and dother denote the expected payoffs for defection to oneself and the partner. p denotes participants’ partner cooperation expectation. The model assumes that participants did not update the inferred cooperation probability based on feedback; p is fixed at 0.5.

Based on the payoff matrix, the payoffs for participants and their partners are calculated using the following functions:

cself=4p (9)
cother=6−2p (10)
dself=4p+2 (11)
dother=2−2p (12)

Participants’ choices are modeled by a softmax function:

q(cooperate)t=11+eβ(Udt−Uct) (13)

where qt denotes the participants’ probability of cooperation and β denotes the inverse temperature. The lower the value of the inverse temperature, the greater the randomness in decisions.

Model 5: social reward model

The model assumes that participants make decisions by comparing the expected payoff of cooperation and defection based on the payoff matrix and an additional subjective bonus from cooperation:

Uct=p(4+ω) (14)
Uct=4p+2 (15)

where ω represents an additional social reward associated with cooperation.

Model 6: social reward model with RL algorithm

The model, building on Model 5, assumed that participants update their expectations of partner cooperation trial-by-trial, based on the partner’s previous decisions, using a basic RL algorithm:

Uct=pt(4+ω) (16)
Udt=4pt+2 (17)

where pt denotes participants’ expectation of partner cooperation probability at trial t and is updated by the following function:

pt+1=pt+α(Pt−pt) (18)

where α is the learning rate applied to the prediction error, (Pt−pt) represents the partner’s decision at trial t, equating to 1 if the partner cooperates and 0 if the partner defects.

Model 7: social reward model with influence model

The model is based on Model 6 and includes an additional assumption that participants update their expectation of the partner’s cooperation by considering not only the partner’s previous decisions but also the influence of their own previous decisions on the partner’s subsequent decisions. This aspect is referred to as second-order belief and is updated by the following function:

pt+1=pt+α(Pt−pt)+κ(Qt−qt′) (19)
qt′=2ωln⁡(1pt−1)1βω (20)

where Qt represents the participants’ decision at trial t, equating to 1 if the participants cooperate and 0 if participants defect. qt′ represents the participants’ inferred cooperation probability of themselves from the partner’s perspective in trial t, which was inferred from function 13. Therefore, (Qt−qt′) denotes the second-order prediction error, and κ is the second-order learning rate that governs the updating of second-order belief.

Model 8: social reward model with asymmetric RL rule

The model, based on Model 6, assumes that participants asymmetrically update positive expectation errors (better than expected) and negative prediction errors (worse than expected) using two distinct learning rates:

pt+1=pt+α+δ(PE)+α−(1−δ)PE (21)
PE=Pt−pt (22)
δ={1,if PE>00,if PE<0 (23)

Model fitting and model comparison

We used maximum likelihood estimation to fit models to each participant’s choices across all trials. The likelihood function, based on the binomial distribution, captured the association between each participant’s choices and each model’s predictions. To minimize the negative log-likelihood, we employed MATLAB’s (MathWorks) fmincon function. To enhance the likelihood of finding the global minimum, we repeated the parameter search process 500 times, using different starting points. In addition, we tested Model 9 (social reward model with Pearce–Hall learning, i.e., dynamic learning rate; see Appendix Analysis for details; and also see Appendix 1—figure 6).

For model evaluation, we first used the AICc, which accounts for the model’s complexity and the number of observed data points (Hurvich and Tsai, 1989). The second metric was the protected exceedance probability from group-level Bayesian model selection (Rigoux et al., 2014), providing a measure of the likelihood that a specific model is superior to other models under consideration. We chose to use the AIC as the metric of goodness-of-fit for model comparison for the following statistical reasons. First, BIC is derived based on the assumption that the ‘true model’ must be one of the models in the limited model set one compares (Burnham and Anderson, 2002; Gelman and Shalizi, 2013), which is unrealistic in our case. In contrast, AIC does not rely on this unrealistic ‘true model’ assumption and instead selects out the model that has the highest predictive power in the model set (Gelman et al., 2014). Second, AIC is also more robust than BIC for finite sample size (Vrieze, 2012).

The log-likelihood is calculated as the following function:

L=∑t=1Tlog⁡(qt|α+,α−,ω,β) (24)

where qt represents the probability of participants’ decision at trial t, equating to q(cooperation) if participants cooperate and q(defection) if participants defect.

Model identifiability and parameter recovery analyses

We performed a model identifiability analysis to ensure that model comparisons were not compromised by model misidentification. For each model, we generated synthetic datasets using parameters estimated from the data of all participants. We then fitted each alternative model to its corresponding synthetic dataset and identified the best-fitting model through model comparison. To test robustness, we repeated this procedure 100 times, calculating the percentage of instances where each model was recognized as the best model across all synthetic datasets generated by that specific model. Consistently high percentages indicated model identifiability. Additionally, we assessed parameter recovery for the best-fitting model (model 8: social reward model with an asymmetric RL rule). This assessment involved calculating the Pearson correlation between the parameters estimated from the 100 synthetic datasets (recovered parameters) and the parameters used to generate these datasets. A higher correlation coefficient between the recovered and the estimated parameters suggested non-redundancy in the parameter space (Appendix 1—figure 3).

Hidden mental variables analysis

LMM1: Participants’ expectation of partner’s cooperation probability that estimated from the winning model, the variable p, is the dependent variable. The fixed and random effects remain the same as GLMM1. See Appendix 1—table 2 for the statistical results of LMM1.

LMM3: Participants’ intrinsic reward for reciprocity that is estimated from the winning model, p×ω, are the dependent variable. The fixed and random effects remain the same as GLMM1. See Appendix 1—table 4 for the statistical results of LMM3.

Acknowledgements

This project has received funding from the Brain Science and Brain-like Intelligence Technology - National Science and Technology Major Project (2021ZD0200500), the National Natural Science Foundation of China (32441109, 32271092, 32130045), the National Social Science Foundation (25VRC015), the Open Research Fund of the State Key Laboratory of Cognitive Neuroscience and Learning (CNLYB2404), the Beijing Major Science and Technology Project under Contract (Z241100001324005), and the Opening Project of the State Key Laboratory of General Artificial Intelligence (SKLAGI20240P06). We thank Christian C Ruff and Xiangjuan Ren for insightful discussions.

Appendix 1

Analysis

Model 9: Social reward model with Pearce–Hall (PH) learning algorithm

To examine whether trial-by-trial learning rate adaptation improves model performance, we extended the social reward model by incorporating a Pearce–Hall (PH) dynamic learning rule (Pearce and Hall, 1980; Li et al., 2011). In this model, participants update their belief about their partner’s cooperation probability (r^t) using a learning rate αt that changes adaptively as a function of recent prediction errors:

r^t+1=r^t+αt(rt−r^t),αt+1=αt+λ(|PEt|−αt),

where rt denotes the observed partner cooperation (or reward), PEt=rt−r^t, and λ controls the rate at which αt adapts to the magnitude of recent prediction errors. This formulation allows the learning rate to increase following surprising outcomes and decrease when outcomes are predictable, independent of the valence of feedback. As in previous models, decision values were computed based on the expected utilities of cooperation (Ucoop=r^t(4+ω)) and defection (Udef=r^t×4+2), and choice probabilities were obtained via a softmax function governed by the inverse temperature parameter β. Model comparison indicated that the PH dynamic learning model did not outperform the best-fitting asymmetric RL model (Model 8) in either age group (see Appendix 1—figure 6).

Hierarchical Bayesian estimation

To complement the maximum likelihood estimation analyses, we additionally implemented a hierarchical Bayesian estimation for the best-fitting model. The hierarchical model incorporated both group-level (adolescent and adult) and individual-level structures to improve the stability and identifiability of parameter estimation.

At the individual level, each participant s was characterized by four parameters: the positive learning rate (α+s), negative learning rate (α−s), social reward weight (ωs), and inverse temperature (βs). Each parameter was assumed to follow a normal distribution centered on the corresponding group-level mean with a group-specific standard deviation:

αs+∼N(μαg[s]+,σαg[s]+),αs−∼N(μαg[s]−,σαg[s]−),ωs∼N(μωg[s],σωg[s]),βs∼N(μβg[s],σβg[s]),

where g[s]∈{adolescent,adult} denotes the group of participant s.

At the group level, the hyperparameters μ and σ were assigned weakly informative priors to regularize estimation while allowing flexibility across age groups. The likelihood for each trial t followed a Bernoulli distribution based on the model-predicted cooperation probability ps,t, computed via the softmax (logistic) transformation of the utility difference:

ys,t∼Bernoulli(ps,t),ps,t=11+exp⁡[−βs(Us,tcoop−Us,tdef)],

where Us,tcoop and Us,tdef denote the expected utilities of cooperation and defection, respectively, updated trial-by-trial using asymmetric learning rates:

PEs,t=rs,t−r^s,t−1,r^s,t={r^s,t−1+αs+⋅PEs,t,if PEs,t>0,r^s,t−1+αs−⋅PEs,t,if PEs,t<0.

The hierarchical Bayesian model was implemented in Stan (Stan Development Team, 2023) and fitted using the sampling() function in RStan. Four independent Markov chains were run with 4000 iterations each, including 1000 warm-up samples for adaptation. Convergence was evaluated using the potential scale reduction statistic R^, with all parameters showing R^<1.01, indicating good convergence and stable sampling (see Appendix 1—figure 7a). Trace plots for the group-level parameters (α+, α−, ω, and β) further confirmed well-mixed chains (see Appendix 1—figure 7b).

Posterior group-level parameter estimates (μ parameters) were then directly compared between adolescents and adults. The hierarchical Bayesian results closely replicated those obtained from the MLE approach, demonstrating consistent age-group differences in parameters (see Appendix 1—figure 5).

Robustness analyses and extensions of main findings

To further examine the robustness of the findings, we first reconduct GLMM1 and LMM1–3 by replacing the categorical Group factor (adolescents vs. adults) with a continuous age predictor, yielding GLMMsup1 and LMMsup1-3. Overall, the patterns replicated those observed for the Group effect: the Previous trial ×Partner’s choice ×Age interaction was significant in GLMMsup1 (b=0.02, 95% CI = [0.001, 0.036], p=0.038, see Appendix 1—table 5) and LMMsup3 (b=0.04, 95% CI = [0.031, 0.046], p<0.001, see Appendix 1—table 8), and marginally significant in LMMsup1 (b=0.0012, 95% CI = [−1.63×10−7, 0.002], p=0.050, see Appendix 1—table 6). Consistent with the nonsignificant effect of the group in LMM2 (see Appendix 1—table 3), the corresponding age effect in LMMsup2 (see Appendix 1—table 7) was also not significant (b=0.03, 95% CI = [–0.168, 0.222], p=0.784). Furthermore, based on GLMM1, LMM1, and LMM3, we further included the phase factor (stable versus changing phase) as a fixed effect with random slopes, yielding GLMMsup2 and LMMsup4–5. This specification accounts for potential effects of the experimental manipulation of the partner’s cooperation, which differed between the first and second halves of the trials. We also replicated the main results: the Previous trial ×Partner’s choice ×Age interaction was significant in GLMMsup2 (b=0.25, 95% CI = [0.128, 0.366], p<0.001, see Appendix 1—table 9), LMMsup4 (b=0.03, 95% CI = [0.022, 0.037], p<0.001, see Appendix 1—table 10), and LMMsup5 (b=0.37, 95% CI = [0.314, 0.417], p<0.001, see Appendix 1—table 11). Besides, to account for the potential influence of individual differences in prosociality on participants’ cooperative behavior and the reward for reciprocity, we also extended GLMM1 and LMM3 by adding the measured SVO as a fixed effect with random slopes, yielding GLMMsup3 and LMMsup6. The results showed that higher SVO positively predicted greater cooperation (b=0.02, 95% CI = [0.007, 0.027], p<0.001, see Appendix 1—table 12), whereas its effect on the reward for reciprocity was not significant (b=0.006, 95% CI = [–0.002, 0.014], p=0.137, see Appendix 1—table 13). Importantly, the primary findings remained unchanged after controlling for SVO.

Finally, we conducted an exploratory analysis to examine the relationship between participants’ expectations of partner’s cooperation probability and their own cooperative behavior. Specifically, we estimated GLMMsup4, in which participants’ choices served as the dependent variable. The fixed effects model included trial number, gender, group, cooperation expectation, and the interaction between group and cooperation expectation. Random effects included correlated random slopes of trial number, gender, group, cooperation expectation, and a random intercept for participants. We showed that participants’ cooperation expectations positively predicted cooperative behavior (b=7.90, 95% CI = [7.258, 8.459], p<0.001, see Appendix 1—table 14), indicating that their cooperation depends on their estimates of partner’ belief. In addition, we found the interaction between group and cooperation expectation was not significant (b=0.015, 95% CI = [–0.358, 0.387], p=0.938), suggesting that this social learning process likely operates similarly in adolescents and adults and is not driven by age-group differences.

Appendix 1—table 1. Statistical results for cooperation decision (GLMM1).

Fixed effects Estimated beta value SE t value p value
(Intercept) 0.47 0.22 2.10 p=0.036
Timing –0.001 0.001 –1.01 p=0.311
Gender 0.14 0.15 0.93 p=0.355
Group 0.79 0.24 3.23 p=0.001
Previous trial –1.81 0.08 –23.43 p<0.001
Partner’s choice –0.73 0.10 –7.21 p<0.001
Group × Previous trial –0.41 0.10 –3.89 p<0.001
Group × Partner’s choice –0.19 0.11 –1.77 p=0.076
Previous trial × Partner’s choice 1.12 0.04 25.43 p<0.001
Group × Previous trial × Partner’s choice 0.24 0.06 4.05 p<0.001

Coding of variables. Trial number: integer sequence from 2 to 120; gender: female = 0, male = 1; group: adolescents = 0, adults = 1; previous trials: last one trial = 1, last two trials = 2, last three trials = 3; partner’s choice: cooperation = 1, defection = 0.

Appendix 1—table 2. Statistical results for partner cooperation expectation (LMM1).

Fixed effects Estimated beta value SE t value p value
(Intercept) 0.59 0.03 17.52 p<0.001
Trial number <0.001 <0.001 0.02 p=0.982
Gender –0.01 0.03 –0.36 p=0.718
Group –0.04 0.03 –1.29 p=0.198
Previous trial –0.25 0.01 –48.97 p<0.001
Partner’s choice –0.06 0.01 –4.90 p<0.001
Group × Previous trial –0.04 0.01 –6.10 p<0.001
Group × Partner’s choice –0.02 0.01 –2.73 p=0.006
Previous trial × Partner’s choice 0.16 0.002 53.89 p<0.001
Group × Previous trial × Partner’s choice 0.03 0.004 7.31 p<0.001

Coding of variables is consistent with Appendix 1—table 1.

Appendix 1—table 3. Statistical results for self-reported perceived partner cooperativeness (LMM2).

Fixed effects Estimated beta value SE t value p value
(Intercept) 5.82 0.27 21.93 p<0.001
Group 0.17 0.35 0.50 p=0.616
Gender –0.44 0.11 –3.84 p<0.001
Order of sessions 0.09 0.15 0.60 p=0.549
Rating number –0.13 0.02 –5.55 p<0.001
Group × Order 0.38 0.22 1.72 p=0.085

Coding of variables. Group: adolescents = 0, adults = 1; gender: female = 0, male = 1; order of sessions: stable to volatile = 0, volatile to stable = 1; rating number: integer sequence from 1 to 8.

Appendix 1—table 4. Statistical results for intrinsic reward for reciprocity (LMM3).

Fixed effects Estimated beta value SE t value p value
(Intercept) 2.43 0.14 17.04 p<0.001
Trial number –0.001 0.001 –0.60 p=0.551
Gender 0.08 0.13 0.59 p=0.554
Group 0.52 0.15 3.44 p<0.001
Previous trial –1.23 0.03 –36.18 p<0.001
Partner’s choice –0.27 0.08 –3.24 p=0.001
Group × Previous trial –0.57 0.05 –11.95 p<0.001
Group × Partner’s choice –0.24 0.05 –4.43 p=0.006
Previous trial × Partner’s choice 0.79 0.02 40.64 p<0.001
Group × Previous trial × Partner’s choice 0.37 0.02 13.70 p<0.001

Coding of variables is consistent with Appendix 1—table 1.

Appendix 1—table 5. Statistical results for cooperation decision with age (GLMMsup1).

Fixed effects Estimated beta value SE t value p value
(Intercept) –1.55 0.73 –2.12 p=0.034
Trial number –0.001 0.001 –0.94 p=0.346
Gender –0.06 0.18 –0.34 p=0.732
Previous trial –0.49 0.15 –3.30 p=0.001
Partner’s choice –0.60 0.33 –1.83 p=0.068
Age 0.10 0.04 2.49 p=0.013
Previous trial × Partner’s choice 0.89 0.17 5.10 p=<0.001
Previous trial × Age –0.02 0.01 –2.00 p=0.045
Partner’s choice × Age –0.01 0.02 –0.68 p=0.499
Previous trial × Partner’s choice × Age 0.02 0.01 2.07 p=0.038

Age was treated as a continuous variable, and all other variables were coded as in Appendix 1—table 1.

Appendix 1—table 6. Statistical results for partner cooperation expectation with age (LMMsup1).

Fixed effects Estimated beta value SE t value p value
(Intercept) 0.76 0.12 6.52 p<0.001
Trial number −2.53×10−7 0.0002 –0.002 p=0.999
Gender –0.01 0.03 –0.32 p=0.746
Previous trial –0.10 0.01 –10.51 p<0.001
Partner’s choice –0.03 0.03 –1.05 p=0.294
Age –0.014 0.006 –2.27 p=0.023
Previous trial × Partner’s choice 0.15 0.01 13.24 p<0.001
Previous trial × Age −4.76×10−7 0.0004 0.001 p=0.999
Partner’s choice × Age –0.002 0.001 –1.74 p=0.082
Previous trial × Partner’s choice × Age 0.0012 0.0006 1.96 p=0.050

Coding of variables is consistent with Appendix 1—table 5.

Appendix 1—table 7. Statistical results for self-reported perceived partner cooperativeness with age (LMMsup2).

Fixed effects Estimated beta value SE t value p value
(Intercept) 5.18 1.89 2.74 p=0.006
Gender –0.37 0.19 –1.96 p=0.050
Order of sessions 2.31 1.21 1.90 p=0.057
Age 0.03 0.10 0.27 p=0.784
Rating number –0.13 0.02 –6.65 p<0.001
Order × Age –0.10 0.07 –1.57 p=0.116

Age was treated as a continuous variable, and all other variables were coded as in Appendix 1—table 3.

Appendix 1—table 8. Statistical results for partner cooperation expectation with age (LMMsup3).

Fixed effects Estimated beta value SE t value p value
(Intercept) 2.02 0.58 3.49 p<0.001
Trial number –0.001 0.001 –0.60 p=0.546
Gender –0.07 0.13 –0.50 p=0.619
Previous trial –0.22 0.06 –3.59 p<0.001
Partner’s choice –0.11 0.18 –0.62 p=0.533
Age 0.02 0.03 0.61 p=0.540
Previous trial × Partner’s choice 0.25 0.08 3.28 p=0.001
Previous trial × Age –0.02 0.003 –5.24 p<0.001
Partner’s choice × Age –0.02 0.01 –1.85 p=0.065
Previous trial × Partner’s choice × Age 0.04 0.004 9.82 p<0.001

Coding of variables is consistent with Appendix 1—table 5.

Appendix 1—table 9. Statistical results for cooperation decision with phase (GLMMsup2).

Fixed effects Estimated beta value SE t value p value
(Intercept) 1.06 0.24 4.34 p<0.001
Trial number –0.01 0.001 –4.46 p<0.001
Gender 0.17 0.15 1.10 p=0.271
Group 0.60 0.19 3.15 p=0.002
Previous trial –0.64 0.04 –16.49 p<0.001
Partner’s choice –0.75 0.10 –7.73 p<0.001
Phase –0.66 0.10 –6.45 p<0.001
Group × Previous trial –0.16 0.05 –3.15 p=0.002
Group × Partner’s choice –0.19 0.11 –1.77 p=0.077
Previous trial × Partner’s choice 1.04 0.04 23.44 p<0.001
Group × Previous trial × Partner’s choice 0.25 0.06 4.07 p<0.001

Phase was dummy-coded (0=stable, 1=changing), and all other variables were coded as in Appendix 1—table 1.

Appendix 1—table 10. Statistical results for partner cooperation expectation with phase (LMMsup4).

Fixed effects Estimated beta value SE t value p value
(Intercept) 0.77 0.03 28.09 p<0.001
Trial number −5.86×10−6 9.25×10−5 –0.06 p=0.949
Gender –0.02 0.03 –0.96 p=0.337
Group –0.06 0.03 –2.36 p=0.018
Previous trial –0.08 0.002 –33.62 p<0.001
Partner’s choice –0.07 0.01 –6.65 p<0.001
Phase –0.14 0.01 –19.16 p<0.001
Group × Previous trial –0.013 0.003 –4.02 p<0.001
Group × Partner’s choice –0.025 0.007 –3.11 p=0.002
Previous trial × Partner’s choice 0.15 0.003 50.67 p<0.001
Group × Previous trial × Partner’s choice 0.03 0.004 7.40 p<0.001

Coding of variables is consistent with Appendix 1—table 9.

Appendix 1—table 11. Statistical results for intrinsic reward for reciprocity with phase (LMMsup5).

Fixed effects Estimated beta value SE t value p value
(Intercept) 3.57 0.14 25.57 p<0.001
Trial number –0.0011 0.0006 –2.01 p=0.045
Gender 0.02 0.13 0.18 p=0.860
Group 0.25 0.13 1.96 p=0.050
Previous trial –0.38 0.02 –22.73 p<0.001
Partner’s choice –0.33 0.07 –4.59 p<0.001
Phase –0.83 0.06 –14.76 p<0.001
Group × Previous trial –0.19 0.02 –8.79 p<0.001
Group × Partner’s choice –0.27 0.05 –5.18 p<0.001
Previous trial × Partner’s choice 0.70 0.02 37.41 p<0.001
Group × Previous trial × Partner’s choice 0.37 0.03 13.99 p<0.001

Coding of variables is consistent with Appendix 1—table 9.

Appendix 1—table 12. Statistical results for cooperation decision with SVO (GLMMsup3).

Fixed effects Estimated beta value SE t value p value
(Intercept) –0.63 0.22 –2.88 p=0.004
Trial number –0.002 0.001 –1.16 p=0.247
Gender 0.18 0.14 1.27 p=0.202
Group 0.56 0.18 3.11 p=0.002
Previous trial –0.71 0.04 –18.17 p<0.001
Partner’s choice –0.77 0.10 –7.41 p<0.001
SVO 0.02 0.005 3.35 p<0.001
Group ×Previous trial –0.14 0.05 –2.84 p=0.005
Group ×Partner’s choice ×Age –0.16 0.11 –1.44 p=0.150
Previous trial ×Partner’s choice 1.13 0.04 25.18 p<0.001
Group ×Previous trial ×Partner’s choice 0.25 0.06 4.06 p<0.001

SVO was treated as a continuous variable, and all other variables were coded as in Appendix 1—table 1.

Appendix 1—table 13. Statistical results for intrinsic reward for reciprocity with SVO (LMMsup6).

Fixed effects Estimated beta value SE t value p value
(Intercept) 2.03 0.15 13.73 p<0.001
Trial number –0.001 0.00 1 –0.71 p=0.479
Gender 0.02 0.12 0.18 p=0.855
Group 0.26 0.13 2.07 p=0.039
Previous trial –0.45 0.02 –25.75 p<0.001
Partner’s choice –0.31 0.08 –3.65 p<0.001
Phase 0.006 0.004 1.49 p=0.137
Group × Previous trial –0.18 0.02 –8.02 p<0.001
Group × Partner’s choice –0.17 0.06 –3.03 p=0.002
Previous trial × Partner’s choice 0.80 0.02 40.26 p<0.001
Group × Previous trial × Partner’s choice 0.37 0.03 13.38 p<0.001

Coding of variables is consistent with Appendix 1—table 12.

Appendix 1—table 14. Statistical results for cooperation decision predicted by cooperation expectation (GLMMsup4).

Fixed effects Estimated beta value SE t value p value
(Intercept) –4.57 0.35 –12.92 p<0.001
Trial number –0.002 0.001 –2.25 p=0.024
Gender 0.20 0.28 0.71 p=0.475
Group 1.37 0.34 4.06 p<0.001
Cooperation expectation 7.90 0.33 24.05 p<0.001
Group × Cooperation expectation 0.01 0.19 0.08 p=0.938

Coding of variables. Trial number: integer sequence from 2 to 120; gender: female = 0, male = 1; group: adolescents = 0, adults = 1. Cooperation expectation was treated as a continuous variable.

Appendix 1—figure 1. Model prediction.

Appendix 1—figure 1.

This figure compares the empirical cooperation probabilities and the model-predicted values for adults (a) and adolescents (b). The x-axis represents the trial number, and the y-axis represents the mean cooperation probability across participants. The shaded areas indicate the 95% confidence intervals.

Appendix 1—figure 2. Distributions of estimated parameters from the best-fitting model.

Appendix 1—figure 2.

Each panel displays one parameter. The histograms and their kernel fits are represented by color bars and curves, respectively. Red indicates participants in the adolescent sample, and blue denotes those in the adult sample. Parameters have been transformed into a log scale for enhanced visualization.

Appendix 1—figure 3. Parameter recovery for the best-fitting model.

Appendix 1—figure 3.

Each panel represents one parameter. Each dot corresponds to one virtual participant. The value of r indicates Pearson’s correlation coefficient between the true values (estimated from the participants) and the recovered parameters.

Appendix 1—figure 4. Partial correlation matrices among parameters for the best-fitting model.

Appendix 1—figure 4.

The upper-triangular cells show partial correlations for adults, and the lower-triangular cells show partial correlations for adolescents. Each cell shows the partial Pearson correlation coefficient (controlling for the other parameters). Colors range from green (negative) to violet (positive), with the color bar spanning [–1,1]. Notes: n.s. p > 0.05; *p < 0.05; **p < 0.01; ***p < 0.001.

Appendix 1—figure 5. Group-level posterior distributions from the hierarchical Bayesian estimation for adolescents and adults.

Appendix 1—figure 5.

Posterior densities are shown separately for adolescents (red) and adults (blue). Δ values indicate the posterior mean difference (Adult – Adolescent) with 95% credible intervals (CrI) and Bayesian p values. Compared with adolescents, adults exhibited higher positive learning rates (α+) and lower negative learning rates (α−), suggesting greater differentiation between learning from positive and negative feedback. Adults also showed lower inverse temperature (β), indicating more exploratory decision behavior, and higher social reward weight (ω), reflecting greater valuation of reciprocity or social outcomes. Notes: n.s. p > 0.05; *p < 0.05; **p < 0.01; ***p < 0.001.

Appendix 1—figure 6. Model comparison results for (a) adults and (b) adolescents, including the newly added M9 (Social Reward and Pearce–Hall learning).

Appendix 1—figure 6.

Lower ΔAICc values indicate better model fits. The dynamic learning rate model (Model 9: Social Reward model with dynamic RL algorithm) did not outperform the best-fitting model (Model 8) in either group.

Appendix 1—figure 7. Convergence diagnostics for the hierarchical Bayesian model.

Appendix 1—figure 7.

(a) Distribution of R^ (Rhat) values across all model parameters. The majority of R^ values are below the conservative convergence threshold of 1.01 (red dashed line), indicating stable and well-mixed MCMC chains. The gray shaded area highlights the region where R^≤1.01. (b) Trace plots for the group-level parameters (four chains) in adolescents (left, red box) and adults (right, blue box). Each line represents the sampled posterior values of one chain across iterations, with overlapping traces and stable fluctuations confirming adequate convergence and mixing for all key parameters (α+, α−, ω, β).

Funding Statement

The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.

Contributor Information

Chao Liu, Email: liuchao@bnu.edu.cn.

Xiaosi Gu, Yale University, United States.

Andre F Marquand, Radboud University Nijmegen, Netherlands.

Funding Information

This paper was supported by the following grants:

  • Brain Science and Brain-like Intelligence Technology - National Science and Technology Major Project 2021ZD0200500 to Chao Liu.

  • National Natural Science Foundation of China 32441109 to Chao Liu.

  • National Social Science Foundation 25VRC015 to Chao Liu.

  • Open Research Fund of the State Key Laborary of Cognitive Neuroscience and Learning CNLYB2404 to Chao Liu.

  • Beijing Major Science and Technology Project Z241100001324005 to Chao Liu.

  • Open Research Fund of the State Key Laborary of General Artificial Intelligence SKLAGI20240P06 to Chao Liu.

  • National Natural Science Foundation of China 32130045 to Chao Liu.

  • National Natural Science Foundation of China 32271092 to Chao Liu.

Additional information

Competing interests

No competing interests declared.

Author contributions

Conceptualization, Resources, Data curation, Software, Formal analysis, Validation, Investigation, Visualization, Methodology, Writing – original draft, Project administration, Writing – review and editing.

Methodology, Writing – original draft, Writing – review and editing.

Writing – review and editing.

Validation, Visualization, Methodology, Writing – original draft, Writing – review and editing.

Conceptualization, Resources, Writing – review and editing.

Resources, Validation.

Resources, Supervision, Funding acquisition, Writing – review and editing.

Ethics

This study was approved by the Ethics Committee of Beijing Normal University (Approval Nos. CNL_A_0001_009 and RB_A_0003_202001). Written informed consent was obtained from all adult participants and from both adolescent participants and their legal guardians prior to participation. The study was conducted in accordance with the Declaration of Helsinki. No identifiable personal information is included in this manuscript.

Additional files

MDAR checklist

Data availability

All data and analysis code required to reproduce the main results are publicly available at Zenodo (https://doi.org/10.5281/zenodo.15046430; Wu, 2026). The source code is maintained at GitHub (https://github.com/xiaoyanwu2024/Adolescents_SelfInterest_Cooperation).

The following dataset was generated:

Wu X. 2025. The Self-Interest of Adolescents Overrules Cooperation in Social Dilemmas. Zenodo.

References

  1. Andreoni J, Miller JH. Rational cooperation in the finitely repeated Prisoner’s dilemma: experimental evidence. The Economic Journal. 1993;103:570. doi: 10.2307/2234532. [DOI] [Google Scholar]
  2. Axelrod R, Hamilton WD. The evolution of cooperation. Science. 1981;211:1390–1396. doi: 10.1126/science.7466396. [DOI] [PubMed] [Google Scholar]
  3. Belli SR, Rogers RD, Lau JYF. Adult and adolescent social reciprocity: experimental data from the Trust Game. Journal of Adolescence. 2012;35:1341–1349. doi: 10.1016/j.adolescence.2012.05.004. [DOI] [PubMed] [Google Scholar]
  4. Blakemore SJ, Mills KL. Is adolescence a sensitive period for sociocultural processing? Annual Review of Psychology. 2014;65:187–207. doi: 10.1146/annurev-psych-010213-115202. [DOI] [PubMed] [Google Scholar]
  5. Burnham KP, Anderson DR. Model Selection and Multimodel Inference: A Practical Information-Theoretic Approach. Springer; 2002. [DOI] [Google Scholar]
  6. Crone EA, Dahl RE. Understanding adolescence as a period of social-affective engagement and goal flexibility. Nature Reviews. Neuroscience. 2012;13:636–650. doi: 10.1038/nrn3313. [DOI] [PubMed] [Google Scholar]
  7. Davis MM, Modi HH, Skymba HV, Finnegan MK, Haigler K, Telzer EH, Rudolph KD. Thumbs up or thumbs down: neural processing of social feedback and links to social motivation in adolescent girls. Social Cognitive and Affective Neuroscience. 2023;18:nsac055. doi: 10.1093/scan/nsac055. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Do KT, Guassi Moreira JF, Telzer EH. But is helping you worth the risk? defining prosocial risk taking in adolescence. Developmental Cognitive Neuroscience. 2017;25:260–271. doi: 10.1016/j.dcn.2016.11.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Embrey M, Fréchette GR, Yuksel S. Cooperation in the finitely repeated Prisoner’s Dilemma*. The Quarterly Journal of Economics. 2018;133:509–551. doi: 10.1093/qje/qjx033. [DOI] [Google Scholar]
  10. Fareri DS, Chang LJ, Delgado MR. Effects of direct social experience on trust decisions and neural reward circuitry. Frontiers in Neuroscience. 2012;6:148. doi: 10.3389/fnins.2012.00148. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Fareri DS, Chang LJ, Delgado MR. Computational substrates of social value in interpersonal collaboration. The Journal of Neuroscience. 2015;35:8170–8180. doi: 10.1523/JNEUROSCI.4775-14.2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Farrell S, Lewandowsky S. Computational models as aids to better reasoning in psychology. Current Directions in Psychological Science. 2010;19:329–335. doi: 10.1177/0963721410386677. [DOI] [Google Scholar]
  13. Fehr E, Fischbacher U. The nature of human altruism. Nature. 2003;425:785–791. doi: 10.1038/nature02043. [DOI] [PubMed] [Google Scholar]
  14. Fett AKJ, Gromann PM, Giampietro V, Shergill SS, Krabbendam L. Default distrust? An fMRI investigation of the neural development of trust and cooperation. Social Cognitive and Affective Neuroscience. 2014;9:395–402. doi: 10.1093/scan/nss144. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Garrett N, Daw ND. Biased belief updating and suboptimal choice in foraging decisions. Nature Communications. 2020;11:3417. doi: 10.1038/s41467-020-16964-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Gelman A, Shalizi CR. Philosophy and the practice of Bayesian statistics. The British Journal of Mathematical and Statistical Psychology. 2013;66:8–38. doi: 10.1111/j.2044-8317.2011.02037.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis. Chapman and Hall/CRC; 2014. [DOI] [Google Scholar]
  18. Gutiérrez-Roig M, Gracia-Lázaro C, Perelló J, Moreno Y, Sánchez A. Transition from reciprocal cooperation to persistent behaviour in social dilemmas at the end of adolescence. Nature Communications. 2014;5:4362. doi: 10.1038/ncomms5362. [DOI] [PubMed] [Google Scholar]
  19. House BR, Kanngiesser P, Barrett HC, Broesch T, Cebioglu S, Crittenden AN, Erut A, Lew-Levy S, Sebastian-Enesco C, Smith AM, Yilmaz S, Silk JB. Universal norm psychology leads to societal diversity in prosocial behaviour and development. Nature Human Behaviour. 2020;4:36–44. doi: 10.1038/s41562-019-0734-z. [DOI] [PubMed] [Google Scholar]
  20. Hurvich CM, Tsai CL. Regression and time series model selection in small samples. Biometrika. 1989;76:297–307. doi: 10.1093/biomet/76.2.297. [DOI] [Google Scholar]
  21. Icenogle G, Steinberg L, Duell N, Chein J, Chang L, Chaudhary N, Di Giunta L, Dodge KA, Fanti KA, Lansford JE, Oburu P, Pastorelli C, Skinner AT, Sorbring E, Tapanya S, Uribe Tirado LM, Alampay LP, Al-Hassan SM, Takash HMS, Bacchini D. Adolescents’ cognitive capacity reaches adult levels prior to their psychosocial maturity: Evidence for a “maturity gap” in a multinational, cross-sectional sample. Law and Human Behavior. 2019;43:69–85. doi: 10.1037/lhb0000315. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. King-Casas B, Tomlin D, Anen C, Camerer CF, Quartz SR, Montague PR. Getting to know you: reputation and trust in a two-person economic exchange. Science. 2005;308:78–83. doi: 10.1126/science.1108062. [DOI] [PubMed] [Google Scholar]
  23. Krekelberg B. Matlab toolbox for bayes factor analysis. v3.0Zenodo. 2024 doi: 10.5281/zenodo.13744717. [DOI]
  24. Li J, Schiller D, Schoenbaum G, Phelps EA, Daw ND. Differential roles of human striatum and amygdala in associative learning. Nature Neuroscience. 2011;14:1250–1252. doi: 10.1038/nn.2904. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Nava F, Margoni F, Herath N, Nava E. Age-dependent changes in intuitive and deliberative cooperation. Scientific Reports. 2023;13:4457. doi: 10.1038/s41598-023-31691-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Nowak MA. Five rules for the evolution of cooperation. Science. 2006;314:1560–1563. doi: 10.1126/science.1133755. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Palminteri S, Kilford EJ, Coricelli G, Blakemore SJ. The computational development of reinforcement learning during adolescence. PLOS Computational Biology. 2016;12:e1004953. doi: 10.1371/journal.pcbi.1004953. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Pearce JM, Hall G. A model for Pavlovian learning: variations in the effectiveness of conditioned but not of unconditioned stimuli. Psychological Review. 1980;87:532–552. [PubMed] [Google Scholar]
  29. Pfeifer JH, Berkman ET. The development of self and identity in adolescence: Neural evidence and implications for a value-based choice perspective on motivated behavior. Child Development Perspectives. 2018;12:158–164. doi: 10.1111/cdep.12279. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Rigoux L, Stephan KE, Friston KJ, Daunizeau J. Bayesian model selection for group studies - revisited. NeuroImage. 2014;84:971–985. doi: 10.1016/j.neuroimage.2013.08.065. [DOI] [PubMed] [Google Scholar]
  31. Rilling JK, Gutman DA, Zeh TR, Pagnoni G, Berns GS, Kilts CD. A neural basis for social cooperation. Neuron. 2002;35:395–405. doi: 10.1016/S0896-6273(02)00755-9. [DOI] [PubMed] [Google Scholar]
  32. Rilling JK, Sanfey AG, Aronson JA, Nystrom LE, Cohen JD. The neural correlates of theory of mind within interpersonal interactions. NeuroImage. 2004;22:1694–1703. doi: 10.1016/j.neuroimage.2004.04.015. [DOI] [PubMed] [Google Scholar]
  33. Rosenbaum GM, Grassie HL, Hartley CA. Valence biases in reinforcement learning shift across adolescence and modulate subsequent memory. eLife. 2022;11:e64620. doi: 10.7554/eLife.64620. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Sachs JL, Mueller UG, Wilcox TP, Bull JJ. The evolution of cooperation. The Quarterly Review of Biology. 2004;79:135–160. doi: 10.1086/383541. [DOI] [PubMed] [Google Scholar]
  35. Stan Development Team Stan reference manual. 2.33StanCon. 2023 https://mc-stan.org
  36. Steinberg L. Cognitive and affective development in adolescence. Trends in Cognitive Sciences. 2005;9:69–74. doi: 10.1016/j.tics.2004.12.005. [DOI] [PubMed] [Google Scholar]
  37. Taheri M, Rotshtein P, Beierholm U. The effect of attachment and environmental manipulations on cooperative behavior in the prisoner’s dilemma game. PLOS ONE. 2018;13:e0205730. doi: 10.1371/journal.pone.0205730. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Tervo-Clemmens B, Calabro FJ, Parr AC, Fedor J, Foran W, Luna B. A canonical trajectory of executive function maturation from adolescence to adulthood. Nature Communications. 2023;14:6922. doi: 10.1038/s41467-023-42540-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Ting F, He Z, Baillargeon R. Toddlers and infants expect individuals to refrain from helping an ingroup victim’s aggressor. PNAS. 2019;116:6025–6034. doi: 10.1073/pnas.1817849116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. van den Bos W, Westenberg M, van Dijk E, Crone EA. Development of trust and reciprocity in adolescence. Cognitive Development. 2010;25:90–102. doi: 10.1016/j.cogdev.2009.07.004. [DOI] [Google Scholar]
  41. van den Bos W, van Dijk E, Westenberg M, Rombouts SARB, Crone EA. Changing brains, changing perspectives: the neurocognitive development of reciprocity. Psychological Science. 2011;22:60–70. doi: 10.1177/0956797610391102. [DOI] [PubMed] [Google Scholar]
  42. van Duijvenvoorde ACK, Huizenga HM, Somerville LH, Delgado MR, Powers A, Weeda WD, Casey BJ, Weber EU, Figner B. Neural correlates of expected risks and returns in risky choice across development. The Journal of Neuroscience. 2015;35:1549–1560. doi: 10.1523/JNEUROSCI.1924-14.2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Vrieze SI. Model selection and psychological theory: a discussion of the differences between the Akaike information criterion (AIC) and the Bayesian information criterion (BIC) Psychological Methods. 2012;17:228–243. doi: 10.1037/a0027127. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Warneken F. How children solve the two challenges of cooperation. Annual Review of Psychology. 2018;69:205–229. doi: 10.1146/annurev-psych-122216-011813. [DOI] [PubMed] [Google Scholar]
  45. Westhoff B, Molleman L, Viding E, van den Bos W, van Duijvenvoorde ACK. Developmental asymmetries in learning to adjust to cooperative and uncooperative environments. Scientific Reports. 2020;10:21761. doi: 10.1038/s41598-020-78546-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Worthman CM, Trang K. Dynamics of body time, social time and life history at adolescence. Nature. 2018;554:451–457. doi: 10.1038/nature25750. [DOI] [PubMed] [Google Scholar]
  47. Wu X, Zhu R, Gong X, Luo Y, Liu C. Social incentives foster cooperation through guilt aversion: An effect that diminishes with primary psychopathic traits. PsyCh Journal. 2023;12:389–398. doi: 10.1002/pchj.641. [DOI] [PubMed] [Google Scholar]
  48. Wu X, Fu H, Zhang T, Bao D, Hu J, Zhu R, Feng C, Gu R, Liu C. A cognitive computational mechanism for mutual cooperation: The roles of positive expectation and social reward. Acta Psychologica Sinica. 2024a;56:1299. doi: 10.3724/SP.J.1041.2024.01299. [DOI] [Google Scholar]
  49. Wu X, Ren X, Liu C, Zhang H. The motive cocktail in altruistic behaviors. Nature Computational Science. 2024b;4:659–676. doi: 10.1038/s43588-024-00685-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Wu X. Xiaoyanwu2024/adolescents-selfinterest-cooperation. v2.1.0Zenodo. 2026 doi: 10.5281/zenodo.18551834. [DOI]

eLife Assessment

Xiaosi Gu 1

This important work investigates cooperative behaviors in adolescents using a repeated Prisoner's Dilemma game. The approach used in the study is solid. The impact of this work could be further enhanced with more rigorous modelling procedures and more modeling selection/comparison details, as well as by framing the findings in terms of the specific game-theoretic context, rather than general cooperation. Findings from this study will be of interest to developmental psychologists, economists, and social psychologists.

Reviewer #1 (Public review):

Anonymous

Summary:

Wu and colleagues aimed to explain previous findings that adolescents, compared to adults, show reduced cooperation following cooperative behaviour from a partner in several social scenarios. The authors analysed behavioural data from adolescents and adults performing a zero-sum Prisoner's Dilemma task and compared a range of social and non-social reinforcement learning models to identify potential algorithmic differences. Their findings suggest that adolescents' lower cooperation is best explained by a reduced learning rate for cooperative outcomes, rather than differences in prior expectations about the cooperativeness of a partner. The authors situate their results within the broader literature, proposing that adolescents' behaviour reflects a stronger preference for self-interest rather than a deficit in mentalising.

Strengths:

The work as a whole suggests that, in line with past work, adolescents prioritise value accumulation, and this can be, in part, explained by algorithmic differences in weighted value learning. The authors situate their work very clearly in past literature, and make it obvious the gap they are testing and trying to explain. The work also includes social contexts which move the field beyond non-social value accumulation in adolescents. The authors compare a series of formal approaches that might explain the results and establish generative and model-comparison procedures to demonstrate the validity of their winning model and individual parameters. The writing was clear, and the presentation of the results was logical and well-structured.

Weaknesses:

I had some concerns about the methods used to fit and approximate parameters of interest. Namely, the use of maximum likelihood versus hierarchical methods to fit models on an individual level, which may reduce some of the outliers noted in the supplement, and also may improve model identifiability.

There was also little discussion given the structure of the Prisoner's Dilemma, and the strategy of the game (that defection is always dominant), meaning that the preferences of the adolescents cannot necessarily be distinguished from the incentives of the game, i.e. they may seem less cooperative simply because they want to play the dominant strategy, rather than a lower preferences for cooperation if all else was the same.

The authors have now addressed my comments and concerns in their revised version.

Appraisal & Discussion:

Overall, I believe this work has the potential to make a meaningful contribution to the field. Its impact would be strengthened by more rigorous modelling checks and fitting procedures, as well as by framing the findings in terms of the specific game-theoretic context, rather than general cooperation.

Comments on revisions:

Thank you to the authors for addressing my comments and concerns.

Reviewer #2 (Public review):

Anonymous

Summary:

This manuscript investigates age-related differences in cooperative behavior by comparing adolescents and adults in a repeated Prisoner's Dilemma Game (rPDG). The authors find that adolescents exhibit lower levels of cooperation than adults. Specifically, adolescents reciprocate partners' cooperation to a lesser degree than adults do. Through computational modeling, they show that this relatively low cooperation rate is not due to impaired expectations or mentalizing deficits, but rather a diminished intrinsic reward for reciprocity. A social reinforcement learning model with asymmetric learning rate best captured these dynamics, revealing age-related differences in how positive and negative outcomes drive behavioral updates. These findings contribute to understanding the developmental trajectory of cooperation and highlight adolescence as a period marked by heightened sensitivity to immediate rewards at the expense of long-term prosocial gains.

Strengths:

Rigid model comparison and parameter recovery procedure. Conceptually comprehensive model space. Well-powered samples.

Weaknesses:

A key conceptual distinction between learning from non-human agents (e.g., bandit machines) and human partners is that the latter are typically assumed to possess stable behavioral dispositions or moral traits. When a non-human source abruptly shifts behavior (e.g., from 80% to 20% reward), learners may simply update their expectations. In contrast, a sudden behavioral shift by a previously cooperative human partner can prompt higher-order inferences about the partner's trustworthiness or the integrity of the experimental setup (e.g., whether the partner is truly interactive or human). The authors may consider whether their modeling framework captures such higher-order social inferences. Specifically, trait-based models-such as those explored in Hackel et al. (2015, Nature Neuroscience)-suggest that learners form enduring beliefs about others' moral dispositions, which then modulate trial-by-trial learning. A learner who believes their partner is inherently cooperative may update less in response to a surprising defection, effectively showing a trait-based dampening of learning rate.

This asymmetry in belief updating has been observed in prior work (e.g., Siegel et al., 2018, Nature Human Behaviour) and could be captured using a dynamic or belief-weighted learning rate. Models incorporating such mechanisms (e.g., dynamic learning rate models as in Jian Li et al., 2011, Nature Neuroscience) could better account for flexible adjustments in response to surprising behavior, particularly in the social domain.

Second, the developmental interpretation of the observed effects would be strengthened by considering possible non-linear relationships between age and model parameters. For instance, certain cognitive or affective traits relevant to social learning-such as sensitivity to reciprocity or reward updating-may follow non-monotonic trajectories, peaking in late adolescence or early adulthood. Fitting age as a continuous variable, possibly with quadratic or spline terms, may yield more nuanced developmental insights.

Finally, the two age groups compared-adolescents (high school students) and adults (university students)-differ not only in age but also in sociocultural and economic backgrounds. High school students are likely more homogenous in regional background (e.g., Beijing locals), while university students may be drawn from a broader geographic and socioeconomic pool. Additionally, differences in financial independence, family structure (e.g., single-child status), and social network complexity may systematically affect cooperative behavior and valuation of rewards. Although these factors are difficult to control fully, the authors should more explicitly address the extent to which their findings reflect biological development versus social and contextual influences.

Comments on revisions:

The authors have addressed most of my previous comments adequately. I only have a minor question: The models with some variations of RL seem to have very similar AIC. What were the authors' criteria in deciding which model is the "winning" model when several models have similar AIC? Are there ways of integrating models with similar structures into a "model family"? Alternatively, is it possible that different models fit better for different subgroups of participants (e.g., high schoolers vs. college students)?

eLife. 2026 Mar 31;14:RP106840. doi: 10.7554/eLife.106840.4.sa3

Author response

Xiaoyan Wu 1, Hongyu Fu 2, Gökhan Aydogan 3, Chunliang Feng 4, Shaozheng Qin 5, Yi Zeng 6, Chao Liu 7

The following is the authors’ response to the previous reviews

Public Reviews:

Reviewer #1 (Public review):

Summary:

Wu and colleagues aimed to explain previous findings that adolescents, compared to adults, show reduced cooperation following cooperative behaviour from a partner in several social scenarios. The authors analysed behavioural data from adolescents and adults performing a zero-sum Prisoner's Dilemma task and compared a range of social and non-social reinforcement learning models to identify potential algorithmic differences. Their findings suggest that adolescents' lower cooperation is best explained by a reduced learning rate for cooperative outcomes, rather than differences in prior expectations about the cooperativeness of a partner. The authors situate their results within the broader literature, proposing that adolescents' behaviour reflects a stronger preference for self-interest rather than a deficit in mentalising.

Strengths:

The work as a whole suggests that, in line with past work, adolescents prioritise value accumulation, and this can be, in part, explained by algorithmic differences in weighted value learning. The authors situate their work very clearly in past literature, and make it obvious the gap they are testing and trying to explain. The work also includes social contexts that move the field beyond non-social value accumulation in adolescents. The authors compare a series of formal approaches that might explain the results and establish generative and modelcomparison procedures to demonstrate the validity of their winning model and individual parameters. The writing was clear, and the presentation of the results was logical and wellstructured.

We thank the reviewer for recognizing the strengths of our work.

Weaknesses:

(Q1) I also have some concerns about the methods used to fit and approximate parameters of interest. Namely, the use of maximum likelihood versus hierarchical methods to fit models on an individual level, which may reduce some of the outliers noted in the supplement, and also may improve model identifiability.

We thank the reviewer for this suggestion. Following the comment, we added a hierarchical Bayesian estimation. We built a hierarchical model with both group-level (adolescent group and adult group) and individual-level structures for the best-fitting model. Four Markov chains with 4,000 samples each were run, and the model converged well (see Figure supplement 7)

We then analyzed the posterior parameters for adolescents and adults separately. The results were consistent with those from the MLE analysis (see Figure 2—figure supplement 5). These additional results have been included in the Appendix Analysis section (also see Figure supplement 5 and 7). In addition, we have updated the code and provided the link for reference. We appreciate the reviewer’s suggestion, which improved our analysis.

(Q2) There was also little discussion given the structure of the Prisoner's Dilemma, and the strategy of the game (that defection is always dominant), meaning that the preferences of the adolescents cannot necessarily be distinguished from the incentives of the game, i.e. they may seem less cooperative simply because they want to play the dominant strategy, rather than a lower preferences for cooperation if all else was the same.

We thank the reviewer for this comment and agree that adolescents’ lower cooperation may partly reflect a rational response to the incentive structure of the Prisoner’s Dilemma.

However, our computational modeling explicitly addressed this possibility. Model 4 (inequality aversion) captures decisions that are driven purely by self-interest or aversion to unequal outcomes, including a parameter reflecting disutility from advantageous inequality, which represents self-oriented motives. If participants’ behavior were solely guided by the payoff-dominant strategy, this model should have provided the best fit. However, our model comparison showed that Model 5 (social reward) performed better in both adolescents and adults, suggesting that cooperative behavior is better explained by valuing social outcomes beyond payoff structures.

Besides, if adolescents’ lower cooperation is that they strategically respond to the payoff structure by adopting defection as the more rewarding option. Then, adolescents should show reduced cooperation across all rounds. Instead, adolescents and adults behaved similarly when partners defected, but adolescents cooperated less when partners cooperated and showed little increase in cooperation even after consecutive cooperative responses. This pattern suggests that adolescents’ lower cooperation cannot be explained solely by strategic responses to payoff structures but rather reflects a reduced sensitivity to others’ cooperative behavior or weaker social reciprocity motives. We have expanded our Discussion to acknowledge this important point and to clarify how the behavioral and modeling results address the reviewer’s concern.

“Overall, these findings indicate that adolescents’ lower cooperation is unlikely to be driven solely by strategic considerations, but may instead reflect differences in the valuation of others’ cooperation or reduced motivation to reciprocate. Although defection is the payoffdominant strategy in the Prisoner’s Dilemma, the selective pattern of adolescents’ cooperation and the model comparison results indicate that their reduced cooperation cannot be fully explained by strategic incentives, but rather reflects weaker valuation of social reciprocity.”

Appraisal & Discussion:

(Q3) The authors have partially achieved their aims, but I believe the manuscript would benefit from additional methodological clarification, specifically regarding the use of hierarchical model fitting and the inclusion of Bayes Factors, to more robustly support their conclusions. It would also be important to investigate the source of the model confusion observed in two of their models.

We thank the reviewer for this comment. In the revised manuscript, we have clarified the hierarchical Bayesian modeling procedure for the best-fitting model, including the group- and individual-level structure and convergence diagnostics. The hierarchical approach produced results that fully replicated those obtained from the original maximumlikelihood estimation, confirming the robustness of our findings. Please also see the response to Q1.

Regarding the model confusion between the inequality aversion (Model 4) and social reward (Model 5) models in the model recovery analysis, both models’ simulated behaviors were best captured by the baseline model. This pattern arises because neither model includes learning or updating processes. Given that our task involves dynamic, multi-round interactions, models lacking a learning mechanism cannot adequately capture participants’ trial-by-trial adjustments, resulting in similar behavioral patterns that are better explained by the baseline model during model recovery. We have added a clarification of this point to the Results:

“The overlap between Models 4 and 5 likely arises because neither model incorporates a learning mechanism, making them less able to account for trial-by-trial adjustments in this dynamic task.”

(Q4) I am unconvinced by the claim that failures in mentalising have been empirically ruled out, even though I am theoretically inclined to believe that adolescents can mentalise using the same procedures as adults. While reinforcement learning models are useful for identifying biases in learning weights, they do not directly capture formal representations of others' mental states. Greater clarity on this point is needed in the discussion, or a toning down of this language.

We sincerely thank the reviewer for this professional comment. We agree that our prior wording regarding adolescents’ capacity to mentalise was somewhat overgeneralized. Accordingly, we have toned down the language in both the Abstract and the Discussion to better align our statements with what the present study directly tests. Specifically, our revisions focus on adolescents’ and adults’ ability to predict others’ cooperation in social learning. This is consistent with the evidence from our analyses examining adolescents’ and adults’ model-based expectations and self-reported scores on partner cooperativeness (see Figure 4). In the revised Discussion, we state:

“Our results suggest that the lower levels of cooperation observed in adolescents stem from a stronger motive to prioritize self-interest rather than a deficiency in predicting others’ cooperation in social learning”.

(Q5) Additionally, a more detailed discussion of the incentives embedded in the Prisoner's Dilemma task would be valuable. In particular, the authors' interpretation of reduced adolescent cooperativeness might be reconsidered in light of the zero-sum nature of the game, which differs from broader conceptualisations of cooperation in contexts where defection is not structurally incentivised.

We thank the reviewer for this comment and agree that adolescents’ lower cooperation may partly reflect a rational response to the incentive structure of the Prisoner’s Dilemma. However, our behavioral and computational evidence suggests that this pattern cannot be explained solely by strategic responses to payoff structures, but rather reflects a reduced sensitivity to others’ cooperative behavior or weaker social reciprocity motives. We have expanded the Discussion to acknowledge this point and to clarify how both behavioral and modeling results address the reviewer’s concern (see also our response to Q2).

(Q6) Overall, I believe this work has the potential to make a meaningful contribution to the field. Its impact would be strengthened by more rigorous modelling checks and fitting procedures, as well as by framing the findings in terms of the specific game-theoretic context, rather than general cooperation.

We thank the reviewer for the professional comments, which have helped us improve our work.

Reviewer #2 (Public review):

Summary:

This manuscript investigates age-related differences in cooperative behavior by comparing adolescents and adults in a repeated Prisoner's Dilemma Game (rPDG). The authors find that adolescents exhibit lower levels of cooperation than adults. Specifically, adolescents reciprocate partners' cooperation to a lesser degree than adults do. Through computational modeling, they show that this relatively low cooperation rate is not due to impaired expectations or mentalizing deficits, but rather a diminished intrinsic reward for reciprocity. A social reinforcement learning model with asymmetric learning rate best captured these dynamics, revealing age-related differences in how positive and negative outcomes drive behavioral updates. These findings contribute to understanding the developmental trajectory of cooperation and highlight adolescence as a period marked by heightened sensitivity to immediate rewards at the expense of long-term prosocial gains.

Strengths:

(1) Rigid model comparison and parameter recovery procedure.

(2) Conceptually comprehensive model space.

(3) Well-powered samples.

We thank the reviewer for highlighting the strengths of our work.

Weaknesses:

(Q1) A key conceptual distinction between learning from non-human agents (e.g., bandit machines) and human partners is that the latter are typically assumed to possess stable behavioral dispositions or moral traits. When a non-human source abruptly shifts behavior (e.g., from 80% to 20% reward), learners may simply update their expectations. In contrast, a sudden behavioral shift by a previously cooperative human partner can prompt higher-order inferences about the partner's trustworthiness or the integrity of the experimental setup (e.g., whether the partner is truly interactive or human). The authors may consider whether their modeling framework captures such higher-order social inferences. Specifically, trait-based models-such as those explored in Hackel et al. (2015, Nature Neuroscience)-suggest that learners form enduring beliefs about others' moral dispositions, which then modulate trial-bytrial learning. A learner who believes their partner is inherently cooperative may update less in response to a surprising defection, effectively showing a trait-based dampening of learning rate.

We thank the reviewer for this thoughtful comment. We agree that social learning from human partners may involve higher-order inferences beyond simple reinforcement learning from non-human sources. To address this, we had previously included such mechanisms in our behavioral modeling. In Model 7 (Social Reward Model with Influence), we tested a higher-order belief-updating process in which participants’ expectations about their partner’s cooperation were shaped not only by the partner’s previous choices but also by the inferred influence of their own past actions on the partner’s subsequent behavior. In other words, participants could adjust their belief about the partner’s cooperation by considering how their partner’s belief about them might change. Model comparison showed that Model 7 did not outperform the best-fitting model, suggesting that incorporating higher-order influence updates added limited explanatory value in this context. As suggested by the reviewer, we have further clarified this point in the revised manuscript.

Regarding trait-based frameworks, we appreciate the reviewer’s reference to Hackel et al. (2015). That study elegantly demonstrated that learners form relatively stable beliefs about others’ social dispositions, such as generosity, especially when the task structure provides explicit cues for trait inference (e.g., resource allocations and giving proportions). By contrast, our study was not designed to isolate trait learning, but rather to capture how participants update their expectations about a partner’s cooperation over repeated interactions. In this sense, cooperativeness in our framework can be viewed as a trait-like latent belief that evolves as evidence accumulates. Thus, while our model does not include a dedicated trait module that directly modulates learning rates, the belief-updating component of our best-fitting model effectively tracks a dynamic, partner-specific cooperativeness, potentially reflecting a prosocial tendency.

(Q2) This asymmetry in belief updating has been observed in prior work (e.g., Siegel et al., 2018, Nature Human Behaviour) and could be captured using a dynamic or belief-weighted learning rate. Models incorporating such mechanisms (e.g., dynamic learning rate models as in Jian Li et al., 2011, Nature Neuroscience) could better account for flexible adjustments in response to surprising behavior, particularly in the social domain.

We thank the reviewer for the suggestion. Following the comment, we implemented an additional model incorporating a dynamic learning rate based on the magnitude of prediction errors. Specifically, we developed Model 9: Social reward model with Pearce–Hall learning algorithm (dynamic learning rate), in which participants’ beliefs about their partner’s cooperation probability are updated using a Rescorla–Wagner rule with a learning rate dynamically modulated by the Pearce–Hall (PH) Error Learning mechanism. In this framework, the learning rate increases following surprising outcomes (larger prediction errors) and decreases as expectations become more stable (see Appendix Analysis section for details).

The results showed that this dynamic learning rate model did not outperform our bestfitting model in either adolescents or adults (see Figure supplement 6). We greatly appreciate the reviewer’s suggestion, which has strengthened the scope of our analysis. We now have added these analyses to the Appendix Analysis section (also Figure Supplement 6) and expanded the Discussion to acknowledge this modeling extension and further discuss its implications.

(Q3) Second, the developmental interpretation of the observed effects would be strengthened by considering possible non-linear relationships between age and model parameters. For instance, certain cognitive or affective traits relevant to social learning-such as sensitivity to reciprocity or reward updating-may follow non-monotonic trajectories, peaking in late adolescence or early adulthood. Fitting age as a continuous variable, possibly with quadratic or spline terms, may yield more nuanced developmental insights.

We thank the reviewer for this professional comment. In addition to the linear analyses, we further conducted exploratory analyses to examine potential non-linear relationships between age and the model parameters. Specifically, we fit LMMs for each of the four parameters as outcomes (α+, α-, β, and ω). The fixed effects included age, a quadratic age term, and gender, and the random effects included subject-specific random intercepts and random slopes for age and gender. Model comparison using BIC did not indicate improvement for the quadratic models over the linear models for α+ (ΔBICquadratic-linear = 5.09), α-(ΔBICquadratic-linear = 3.04), β (ΔBICquadratic-linear = 3.9), or ω (ΔBICquadratic-linear = 0). Moreover, the quadratic age term was not significant for α+, α−, or β (all ps > 0.10). For ω, we observed a significant linear age effect (b = 1.41, t = 2.65, p = 0.009) and a significant quadratic age effect (b = −0.03, t = −2.39, p = 0.018; see Author response image 1). This pattern is broadly consistent with the group effect reported in the main text. The shaded area in the figure represents the 95% confidence interval. As shown, the interval widens at older ages (≥ 26 years) due to fewer participants in that range, which limits the robustness of the inferred quadratic effect. In consideration of the limited precision at older ages and the lack of BIC improvement, we did not emphasize the quadratic effect in the revised manuscript and present these results here as exploratory.

Author response image 1. Linear and quadratic model fits showing the relationship between age and the ω parameter, with 95% confidence intervals.

Author response image 1.

(Q4) Finally, the two age groups compared - adolescents (high school students) and adults (university students) - differ not only in age but also in sociocultural and economic backgrounds. High school students are likely more homogenous in regional background (e.g., Beijing locals), while university students may be drawn from a broader geographic and socioeconomic pool. Additionally, differences in financial independence, family structure (e.g., single-child status), and social network complexity may systematically affect cooperative behavior and valuation of rewards. Although these factors are difficult to control fully, the authors should more explicitly address the extent to which their findings reflect biological development versus social and contextual influences.

We appreciate this comment. Indeed, adolescents (high school students) and adults (university students) differ not only in age but also in sociocultural and socioeconomic backgrounds. In our study, all participants were recruited from Beijing and surrounding regions, which helps minimize large regional and cultural variability. Moreover, we accounted for individual-level random effects and included participants’ social value orientation (SVO) as an individual difference measure.

Nonetheless, we acknowledge that other contextual factors, such as differences in financial independence, socioeconomic status, and social experience—may also contribute to group differences in cooperative behavior and reward valuation. Although our results are broadly consistent with developmental theories of reward sensitivity and social decisionmaking, sociocultural influences cannot be entirely ruled out. Future work with more demographically matched samples or with socioeconomic and regional variables explicitly controlled will help clarify the relative contributions of biological and contextual factors. Accordingly, we have revised the Discussion to include the following statement:

“Third, although both age groups were recruited from Beijing and nearby regions, minimizing major regional and cultural variation, adolescents and adults may still differ in socioeconomic status, financial independence, and social experience. Such contextual differences could interact with developmental processes in shaping cooperative behavior and reward valuation. Future research with demographically matched samples or explicit measures of socioeconomic background will help disentangle biological from sociocultural influences.”

Reviewer #3 (Public review):

Summary:

Wu and colleagues find that in a repeated Prisoner's Dilemma, adolescents, compared to adults, are less likely to increase their cooperation behavior in response to repeated cooperation from a simulated partner. In contrast, after repeated defection by the partner, both age groups show comparable behavior.

To uncover the mechanisms underlying these patterns, the authors compare eight different models. They report that a social reward learning model, which includes separate learning rates for positive and negative prediction errors, best fits the behavior of both groups. Key parameters in this winning model vary with age: notably, the intrinsic value of cooperating is lower in adolescents. Adults and adolescents also differ in learning rates for positive and negative prediction errors, as well as in the inverse temperature parameter.

Strengths:

The modeling results are compelling in their ability to distinguish between learned expectations and the intrinsic value of cooperation. The authors skillfully compare relevant models to demonstrate which mechanisms drive cooperation behavior in the two age groups.

We thank the reviewer’s recognition of our work’s strengths.

Weaknesses:

(Q1) Some of the claims made are not fully supported by the data:

The central parameter reflecting preference for cooperation is positive in both groups. Thus, framing the results as self-interest versus other-interest may be misleading.

We thank the reviewer for this insightful comment. In the social reward model, the cooperation preference parameter is positive by definition, as defection in the repeated rPDG always yields a +2 monetary advantage regardless of the partner’s action. This positive value represents the additional subjective reward assigned to mutual cooperation (e.g., reciprocity value) that counterbalances the monetary gain from defection. Although the estimated social reward parameter ω was positive, the effective advantage of cooperation is Δ=p×ω−2. Given participants’ inferred beliefs p, Δ was negative for most trials (p×ω<2), indicating that the social reward was insufficient to offset the +2 advantage of defection. Thus, both adolescents and adults valued cooperation positively, but adolescents’ smaller ω and weaker responsiveness to sustained partner cooperation suggest a stronger weighting on immediate monetary payoffs.

In this light, our framing of adolescents as more self-interested derives from their behavioral pattern: even when they recognized sustained partner cooperation and held high expectations of partner cooperation, adolescents showed lower cooperative behavior and reciprocity rewards compared with adults. Whereas adults increased cooperation after two or three consecutive partner cooperations, this pattern was absent among adolescents. We therefore interpret their behavior as relatively more self-interested, reflecting reduced sensitivity to the social reward from mutual cooperation rather than a categorical shift from self-interest to other-interest, as elaborated in the Discussion.

(Q2) It is unclear why the authors assume adolescents and adults have the same expectations about the partner's cooperation, yet simultaneously demonstrate age-related differences in learning about the partner. To support their claim mechanistically, simulations showing that differences in cooperation preference (i.e., the w parameter), rather than differences in learning, drive behavioral differences would be helpful.

We thank the reviewer for raising this important point. In our model, both adolescents and adults updated their beliefs about partner cooperation using an asymmetric reinforcement learning (RL) rule. Although adolescents exhibited a higher positive and a lower negative learning rate than adults, the two groups did not differ significantly in their overall updating of partner cooperation probability (Fig. 4a-b). We then examined the social reward parameter ω, which was significantly smaller in adolescents and determined the intrinsic value of mutual cooperation (i.e., p×ω). This variable differed significantly between groups and closely matched the behavioral pattern.

Following the reviewer’s suggestion, we conducted additional simulations varying one model parameter at a time while holding the others constant. The difference in mean cooperation probability between adults and adolescents served as the index (positive = higher cooperation in adults). As shown in the Author response image 2, decreases in ω most effectively reproduced the observed group difference (shaded area), indicating that age-related differences in cooperation are primarily driven by variation in the social reward parameter ω rather than by others.

Author response image 2. Simulation results showing how variations in each model parameter affect the group difference in mean cooperation probability (Adults – Adolescents).

Author response image 2.

Based on the bestfitting Model 8 and parameters estimated from all participants, each line represents one parameter (i.e., α+, α-, ω, β) systematically varied within the tested range (α±:0.1–0.9; ω, β:1–9) while other parameters were held constant. Positive values indicate higher cooperation in adults. Smaller ω values most strongly reproduced the observed group difference, suggesting that reduced social reward weighting primarily drives adolescents’ lower cooperation.

(Q3) Two different schedules of 120 trials were used: one with stable partner behavior and one with behavior changing after 20 trials. While results for order effects are reported, the results for the stable vs. changing phases within each schedule are not. Since learning is influenced by reward structure, it is important to test whether key findings hold across both phases.

We thank the reviewer for this thoughtful and professional comment. In our GLMM and LMM analyses, we focused on trial order rather than explicitly including the stable vs. changing phase factor, due to concerns about multicollinearity. In our design, phases occur in specific temporal segments, which introduces strong collinearity with trial order. In multi-round interactions, order effects also capture variance related to phase transitions.

Nonetheless, to directly address this concern, we conducted additional robustness analyses by adding a phase variable (stable vs. changing) to GLMM1, LMM1, and LMM3 alongside the original covariates. Across these specifications, the key findings were replicated (see GLMMsup2 and LMMsup4–5; Tables 9-11), and the direction and significance of main effects remained unchanged, indicating that our conclusions are robust to phase differences.

(Q4) The division of participants at the legal threshold of 18 years should be more explicitly justified. The age distribution appears continuous rather than clearly split. Providing rationale and including continuous analyses would clarify how groupings were determined.

We thank the reviewer for this thoughtful comment. We divided participants at the legal threshold of 18 years for both conceptual and practical reasons grounded in prior literature and policy. In many countries and regions, 18 marks the age of legal majority and is widely used as the boundary between adolescence and adulthood in behavioral and clinical research. Empirically, prior studies indicate that psychosocial maturity and executive functions approach adult levels around this age, with key cognitive capacities stabilizing in late adolescence (Icenogle et al., 2019; Tervo-Clemmens et al., 2023). We have clarified this rationale in the Introduction section of the revised manuscript.

“Based on legal criteria for majority and prior empirical work, we adopt 18 years as the boundary between adolescence and adulthood (Icenogle et al., 2019; Tervo-Clemmens et al., 2023).”

We fully agree that the underlying age distribution is continuous rather than sharply divided. To address this, we conducted additional analyses treating age as a continuous predictor (see GLMMsup1 and LMMsup1–3; Tables S1-S4), which generally replicated the patterns observed with the categorical grouping. Nevertheless, given the limited age range of our sample, the generalizability of these findings to fine-grained developmental differences remains constrained. Therefore, our primary analyses continue to focus on the contrast between adolescents and adults, rather than attempting to model a full developmental trajectory.

(Q5) Claims of null effects (e.g., in the abstract: "adults increased their intrinsic reward for reciprocating... a pattern absent in adolescents") should be supported with appropriate statistics, such as Bayesian regression.

We thank the reviewer for highlighting the importance of rigor when interpreting potential null effects. To address this concern, we conducted Bayes factor analyses of the intrinsic reward for reciprocity and reported the corresponding BF10 for all relevant post hoc comparisons. This approach quantifies the relative evidence for the alternative versus the null hypothesis, thereby providing a more direct assessment of null effects. The analysis procedure is now described in the Methods and Materials section:

“Post hoc comparisons were conducted using Bayes factor analyses with MATLAB’s bayesFactor Toolbox (version v3.0, Krekelberg, 2024), with a Cauchy prior scale σ = 0.707.”

(Q6) Once claims are more closely aligned with the data, the study will offer a valuable contribution to the field, given its use of relevant models and a well-established paradigm.

We are grateful for the reviewer’s generous appraisal and insightful comments.

Recommendations for the authors:

Reviewer #1 (Recommendations for the authors):

(1) I commend the authors on a well-structured, clear, and interesting piece of work. I have several questions and recommendations that, if addressed, I believe will strengthen the manuscript.

We thank the reviewer for commending the organization of our paper.

(2) Introduction: - Why use a zero-sum (Prisoner's Dilemma; PD) versus a mixed-motive game (e.g. Trust Task) to study cooperation? In a finite set of rounds, the dominant strategy can be to defect in a PD.

We thank the reviewer for this helpful comment. We agree that both the rationale for using the repeated Prisoner’s Dilemma (rPDG) and the limitations of this framework should be clarified. We chose the rPDG to isolate the core motivational conflict between selfinterest and joint welfare, as its symmetric and simultaneous structure avoids the sequential trust and reputation dependencies/accumulation inherent to asymmetric tasks such as the Trust Game (King-Casas et al., 2005; Rilling et al., 2002).

Although a finitely repeated rPDG theoretically favors defection, extensive prior research shows that cooperation can still emerge in long repeated interactions when players rely on learning and reciprocity rather than backward induction (Rilling et al., 2002; Fareri et al., 2015). Our design employed 120 consecutive rounds, allowing participants to update expectations about partner behavior and to establish stable reciprocity patterns over time. We have added the following clarification to the Introduction:

“The rPDG provides a symmetric and simultaneous framework that isolates the motivational conflict between self-interest and joint welfare, avoiding the sequential trust and reputation dynamics characteristic of asymmetric tasks such as the Trust Game (Rilling et al., 2002; King-Casas et al., 2005)”

(3) Methods:

Did the participants know how long the PD would go on for?

Were the participants informed that the partner was real/simulated?

Were the participants informed that the partner was going to be the same for all rounds?

We thank the reviewer for the meticulous review work, which helped us present the experimental design and reporting details more clearly. the following clarifications: I. Participants were not informed of the total number of rounds in the rPDG. This prevented endgame expectations and avoided distraction from counting rounds, which could introduce additional effects. II. Participants were told that their partner was another human participant in the laboratory. However, the partner’s behavior was predetermined by a computer program. This design enabled tighter experimental control and ensured consistent conditions across age groups, supporting valid comparisons. III. Participants were informed that they would interact with the same partner across all rounds, aligning with the essence of a multiround interaction paradigm and stabilizing partner-related expectations. For transparency, we have clarified these points in the Methods and Materials section:

“Participants were told that their partner was another human participant in the laboratory and that they would interact with the same partner across all rounds. However, in reality, the actions of the partner were predetermined by a computer program. This setup allowed for a clear comparison of the behavioral responses between adolescents and adults. Participants were not informed of the total number of rounds in the rPDG.”

(4) The authors mention that an SVO was also recorded to indicate participant prosociality. Where are the results of this? Did this track game play at all? Could cooperativeness be explained broadly as an SVO preference that penetrated into game-play behaviour?

We thank the reviewer for pointing this out. We agree that individual differences in prosociality may shape cooperative behavior, so we conducted additional analyses incorporating SVO. Specifically, we extended GLMM1 and LMM3 by adding the measured SVO as a fixed effect with random slopes, yielding GLMMsup3 and LMMsup6 (Tables 12–13). The results showed that higher SVO was associated with greater cooperation, whereas its effect on the reward for reciprocity was not significant. Importantly, the primary findings remained unchanged after controlling for SVO. These results indicate that cooperativeness in our task cannot be explained solely by a broad SVO preference, although a more prosocial orientation was associated with greater cooperation. We have reported these analyses and results in the Appendix Analysis section.

(5) Why was AIC chosen rather an BIC to compare model dominance?

Sorry for the lack of clarification. Both the Akaike Information Criterion (AIC, Akaike, 1974) and Bayesian Information Criterion (BIC, Schwarz, 1978) are informationtheoretic criterions for model comparison, neither of which depends on whether the models to be compared are nested to each other or not (Burnham et al., 2002). We have added the following clarification into the Methods.

“We chose to use the AICc as the metric of goodness-of-fit for model comparison for the following statistical reasons. First, BIC is derived based on the assumption that the “true model” must be one of the models in the limited model set one compares (Burnham et al., 2002; Gelman & Shalizi, 2013), which is unrealistic in our case. In contrast, AIC does not rely on this unrealistic “true model” assumption and instead selects out the model that has the highest predictive power in the model set (Gelman et al., 2014). Second, AIC is also more robust than BIC for finite sample size (Vrieze, 2012).”

(6) I believe the model fitting procedure might benefit from hierarchical estimation, rather than maximum likelihood methods. Adolescents in particular seem to show multiple outliers in a^+ and w^+ at the lower end of the distributions in Figure S2. There are several packages to allow hierarchical estimation and model comparison in MATLAB (which I believe is the language used for this analysis; see https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1007043).

We thank the reviewer for this helpful comment and for referring us to relevant methodological work (Piray et al., 2019). We have addressed this point by incorporating hierarchical Bayesian estimation, which effectively mitigates outlier effects and improves model identifiability. The results replicated those obtained with MLE fitting and further revealed group-level differences in key parameters. Please see our detailed response to Reviewer#1 Q1 for the full description of this analysis and results.

(7) Results: Model confusion seems to show that the inequality aversion and social reward models were consistently confused with the baseline model. Is this explained or investigated? I could not find an explanation for this.

The apparent overlap between the inequality aversion (Model 4) and social reward (Model 5) models in the recovery analysis likely arises because neither model includes a learning mechanism, making them unable to capture trial-by-trial adjustments in this dynamic task. Consequently, both were best fit by the baseline model. Please see Response to Reviewer #1 Q3 for related discussion.

(8) Figures 3e and 3f show the correlation between asymmetric learning rates and age. It seems that both a^+ and a^- are around 0.35-0.40 for young adolescents, and this becomes more polarised with age. Could it be that with age comes an increasing discernment of positive and negative outcomes on beliefs, and younger ages compress both positive and negative values together? Given the higher stochasticity in younger ages (\beta), it may also be that these values simply represent higher uncertainty over how to act in any given situation within a social context (assuming the differences in groups are true).

We appreciate this insightful interpretation. Indeed, both α+ and α- cluster around 0.35–0.40 in younger adolescents and become increasingly polarized with age, suggesting that sensitivity to positive versus negative feedback is less differentiated early in development and becomes more distinct over time. This interpretation remains tentative and warrants further validation. Based on this comment, we have revised the Discussion to include this developmental interpretation.

We also clarify that in our model β denotes the inverse temperature parameter; higher β reflects greater choice precision and value sensitivity, not higher stochasticity. Accordingly, adolescents showed higher β values, indicating more value-based and less exploratory choices, whereas adults displayed relatively greater exploratory cooperation. These group differences were also replicated using hierarchical Bayesian estimation (see Response to Reviewer #1 Q1). In response to this comment, we have added a statement in the Discussion highlighting this developmental interpretation.

“Together, these findings suggest that the differentiation between positive and negative learning rates changes with age, reflecting more selective feedback sensitivity in development, while higher β values in adolescents indicate greater value sensitivity. This interpretation remains tentative and requires further validation in future research.”

(9) A parameter partial correlation matrix (off-diagonal) would be helpful to understand the relationship between parameters in both adolescents and adults separately. This may provide a good overview of how the model properties may change with age (e.g. a^+'s relation to \beta).

We thank the reviewer for this helpful comment. We fully agree that a parameter partial correlation matrix can further elucidate the relationships among parameters. Accordingly, we conducted a partial correlation analysis and added the visually presented results to the revised manuscript as Figure 2-figure supplement 4.

(10) It would be helpful to have Bayes Factors reported with each statistical tests given that several p-values fall within the 0.01 and 0.10.

We thank the reviewer for this important recommendation. We have conducted Bayes factor analyses and reported BF10 for all relevant post hoc comparisons. We also clarified our analysis in the Methods and Materials section:

“Post hoc comparisons were conducted using Bayes factor analyses with MATLAB’s bayesFactor Toolbox (version v3.0, Krekelberg, 2024), with a Cauchy prior scale σ = 0.707.”

(11) Discussion: I believe the language around ruling out failures in mentalising needs to be toned down. RL models do not enable formal representational differences required to assess mentalising, but they can distinguish biases in value learning, which in itself is interesting. If the authors were to show that more complex 'ToM-like' Bayesian models were beaten by RL models across the board, and this did not differ across adults and adolescents, there would be a stronger case to make this claim. I think the authors either need to include Bayesian models in their comparison, or tone down their language on this point, and/or suggest ways in which this point might be more thoroughly investigated (e.g., using structured models on the same task and running comparisons: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0087619).

We thank the reviewer for the comments. Please see our response to Reviewer 1 (Appraisal & Discussion section) for details.

Reviewer #2 (Recommendations for the authors):

(1) The authors may want to show the winning model earlier (perhaps near the beginning of the Results section, when model parameters are first mentioned).

We thank the reviewer for this suggestion. We agree that highlighting the winning model early improves clarity. Currently, we have mentioned the winning model before the beginning of the Results section. Specifically, in the penultimate paragraph of the Introduction we state:

“We identified the asymmetric RL learning model as the winning model that best explained the cooperative decisions of both adolescents and adults.”

Reviewer #3 (Recommendations for the authors):

(1) In addition to the points mentioned above, I suggest the following:

Clarify plots by clearly explaining each variable. In particular, the indices 1 vs. 1,2 vs 1,2,3 were not immediately understandable.

We thank the reviewer for this suggestion. We agree that the indices were not immediately clear. We have revised the figure captions (Figure 1 and 4) to explicitly define these terms more clearly:

“The x-axis represents the consistency of the partner’s actions in previous trials (t−1: last trial; t−1,2: last two trials;t−1,2,3: last three trials).”

(2) It's unclear why the index stops at 3. If this isn't the maximum possible number of consecutive cooperation trials, please consider including all relevant data, as adolescents might show a trend similar to adults over more trials.

We thank the reviewer for raising this point. In our exploratory analyses, we also examined longer streaks of consecutive partner cooperation or defection (up to four or five trials). Two empirical considerations led us to set the cutoff at three in the final analyses. First, the influence of partner behavior diminished sharply with temporal distance. In both GLMMs and LMMs, coefficients for earlier partner choices were small and unstable, and their inclusion substantially increased model complexity and multicollinearity. This recency pattern is consistent with learning and decision models emphasizing stronger weighting of recent evidence (Fudenberg & Levine, 2014; Fudenberg & Peysakhovich, 2016). Second, streaks longer than three were rare, especially among some participants, leading to data sparsity and inflated uncertainty. Including these sparse conditions risked biasing group estimates rather than clarifying them. Balancing informativeness and stability, we therefore restricted the index to three consecutive partner choices in the main analyses, which we believe sufficiently capture individuals’ general tendencies in reciprocal cooperation.

(3) The term "reciprocity" may not be necessary. Since it appears to reflect a general preference for cooperation, it may be clearer to refer to the specific behavior or parameter being measured. This would also avoid confusion, especially since adolescents do show negative reciprocity in response to repeated defection.

We thank you for this comment. In our work, we compute the intrinsic reward for reciprocity as p × ω, where p is the partner cooperation expectation and ω is the cooperation preference. In the rPDG, this value framework manifests as a reciprocity-derived reward: sustained mutual cooperation maximizes joint benefits, and the resulting choice pattern reflects a value for reciprocity, contingent on the expected cooperation of the partner. This quantity enters the trade-off between Ucooperation and Udefection and captures the participant’s intrinsic reward for reciprocity versus the additional monetary reward payoff of defection. Therefore, we consider the term “reciprocity” an acceptable statement for this construct.

(4) Interpretation of parameters should closely reflect what they specifically measure.

We thank the reviewer for pointing this out. We have refined the relevant interpretations of parameters in the current Results and Discussion sections.

(5) Prior research has shown links between Theory of Mind (ToM) and cooperation (e.g., Martínez-Velázquez et al., 2024). It would be valuable to test whether this also holds in your dataset.

We thank the reviewer for this thoughtful comment. Although we did not directly measure participants’ ToM, our design allowed us to estimate participants’ trial-by-trial inferences (i.e., expectations) about their partner’s cooperation probability. We therefore treat these cooperation expectations as an indirect representation for belief inference, which is related to ToM processes. To test whether this belief-inference component relates to cooperation in our dataset, we further conducted an exploratory analysis (GLMMsup4) in which participants’ choices were regressed on their cooperation expectations, group, and the group × cooperation-expectation interaction, controlling for trial number and gender, with random effects. Consistent with the ToM–cooperation link in prior research (MartínezVelázquez et al., 2024), participants’ expectations about their partner’s cooperation significantly predicted their cooperative behavior (Table 14), suggesting that decisions were shaped by social learning about others’ inferred actions. Moreover, the interaction between group and cooperation expectation was not significant, indicating that this inference-driven social learning process likely operates similarly in adolescents and adults. This aligns with our primary modeling results showing that both age groups update beliefs via an asymmetric learning process. We have reported these analyses in the Appendix Analysis section.

(6) More informative table captions would help the reader. Please clarify how variables are coded (e.g., is female = 0 or 1? Is adolescent = 0 or 1?), to avoid the need to search across the manuscript for this information.

We thank the reviewer for raising this point. We have added clear and standardized variable coding in the table notes of all tables to make them more informative and avoid the need to search the paper. We have ensured consistent wording and formatting across all tables.

(7) I hope these comments are helpful and support the authors in further strengthening their manuscript.

We thank the three reviewers for their comments, which have been helpful in strengthening this work.

References

(1) Fudenberg, D., & Levine, D. K. (2014). Recency, consistent learning, and Nash equilibrium. Proceedings of the National Academy of Sciences of the United States of America, 111(Suppl. 3), 10826–10829. https://doi.org/10.1073/pnas.1400987111.

(2) Fudenberg, D., & Peysakhovich, A. (2016). Recency, records, and recaps: Learning and nonequilibrium behavior in a simple decision problem. ACM Transactions on Economics and Computation, 4(4), Article 23, 1–18. https://doi.org/10.1145/2956581

(3) Hackel, L., Doll, B., & Amodio, D. (2015). Instrumental learning of traits versus rewards: Dissociable neural correlates and effects on choice. Nature Neuroscience, 18, 1233– 1235. https://doi.org/10.1038/nn.4080

(4) Icenogle, G., Steinberg, L., Duell, N., Chein, J., Chang, L., Chaudhary, N., Di Giunta, L., Dodge, K. A., Fanti, K. A., Lansford, J. E., Oburu, P., Pastorelli, C., Skinner, A. T.Sorbring, E., Tapanya, S., Uribe Tirado, L. M., Alampay, L. P., Al-Hassan, S. M.,Takash, H. M. S., & Bacchini, D. (2019). Adolescents’ cognitive capacity reaches adult levels prior to their psychosocial maturity: Evidence for a “maturity gap” in a multinational, cross-sectional sample. Law and Human Behavior, 43(1), 69–85. https://doi.org/10.1037/lhb0000315

(5) Krekelberg, B. (2024). Matlab Toolbox for Bayes Factor Analysis (v3.0) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.13744717

(6) Martínez-Velázquez, E. S., Ponce-Juárez, S. P., Díaz Furlong, A., & Sequeira, H. (2024). Cooperative behavior in adolescents: A contribution of empathy and emotional regulation? Frontiers in Psychology, 15,1342458. https://doi.org/10.3389/fpsyg.2024.1342458

(7) Tervo-Clemmens, B., Calabro, F. J., Parr, A. C., et al. (2023). A canonical trajectory of executive function maturation from adolescence to adulthood. Nature Communications, 14, 6922. https://doi.org/10.1038/s41467-023-42540-8

(8) King-Casas, B., Tomlin, D., Anen, C., Camerer, C. F., Quartz, S. R., & Montague, P. R. (2005). Getting to know you: reputation and trust in a two-person economic exchange. Science, 308(5718), 78-83. https://doi.org/10.1126/science.1108062

(9) Rilling, J. K., Gutman, D. A., Zeh, T. R., Pagnoni, G., Berns, G. S., & Kilts, C. D. (2002).A neural basis for social cooperation. Neuron, 35(2), 395-405. https://doi.org/10.1016/s0896-6273(02)00755-9

(10) Fareri, D. S., Chang, L. J., & Delgado, M. R. (2015). Computational substrates of social value in interpersonal collaboration. Journal of Neuroscience, 35(21), 8170-8180. https://doi.org/10.1523/JNEUROSCI.4775-14.2015

(11) Akaike, H. (2003). A new look at the statistical model identification. IEEE transactions on automatic control, 19(6), 716-723. https://doi.org/10.1109/TAC.1974.1100705

(12) Schwarz, G. (1978). Estimating the dimension of a model. The annals of statistics, 461464. https://doi.org/10.1214/aos/1176344136

(13) Burnham, K. P., & Anderson, D. R. (2002). Model selection and multimodel inference: A practical information-theoretic approach (2nd ed.). Springer.https://doi.org/10.1007/b97636

(14) Gelman, A., & Shalizi, C. R. (2013). Philosophy and the practice of Bayesian statistics. British Journal of Mathematical and Statistical Psychology, 66(1), 8–38. https://doi.org/10.1111/j.2044-8317.2011.02037.x

(15) Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., & Rubin, D. B. (2014). Bayesian data analysis (3rd ed.). Chapman and Hall/CRC. https://doi.org/10.1201/b16018

(16) Vrieze, S. I. (2012). Model selection and psychological theory: A discussion of the differences between the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC). Psychological Methods, 17(2), 228–243. https://doi.org/10.1037/a0027127

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Data Citations

    1. Wu X. 2025. The Self-Interest of Adolescents Overrules Cooperation in Social Dilemmas. Zenodo. [DOI] [PMC free article] [PubMed]

    Supplementary Materials

    MDAR checklist

    Data Availability Statement

    All data and analysis code required to reproduce the main results are publicly available at Zenodo (https://doi.org/10.5281/zenodo.15046430; Wu, 2026). The source code is maintained at GitHub (https://github.com/xiaoyanwu2024/Adolescents_SelfInterest_Cooperation).

    The following dataset was generated:

    Wu X. 2025. The Self-Interest of Adolescents Overrules Cooperation in Social Dilemmas. Zenodo.


    Articles from eLife are provided here courtesy of eLife Sciences Publications, Ltd

    RESOURCES