Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2015 Oct 12;112(43):13407–13410. doi: 10.1073/pnas.1507527112

Receipt of reward leads to altered estimation of effort

Arezoo Pooresmaeili a,1, Aurel Wannig b, Raymond J Dolan a,c,d
PMCID: PMC4629341  PMID: 26460026

Significance

Retrospective reevaluation of effort is a pervasive aspect of everyday life, such as when we assess our professional satisfaction after knowing the ensuing outcomes. Previous studies have focused on the interaction of effort and reward when a choice is to be made, whereas retrospective interactions have been largely ignored. Here we show that humans revise their estimation of effort after receiving a reward. When rewarded more than average, subjects tend to overestimate their effort, with a converse effect observed for low rewards. The size of this bias depends strongly on the contingency between reward magnitude and task difficulty and is dynamically adjusted when changes occur in these contingencies. These results reveal a sophisticated mechanism to cope with reward–effort inconsistencies.

Keywords: effort, reward, retrospective, Bayesian, cue integration

Abstract

Effort and reward jointly shape many human decisions. Errors in predicting the required effort needed for a task can lead to suboptimal behavior. Here, we show that effort estimations can be biased when retrospectively reestimated following receipt of a rewarding outcome. These biases depend on the contingency between reward and task difficulty and are stronger for highly contingent rewards. Strikingly, the observed pattern accords with predictions from Bayesian cue integration, indicating humans deploy an adaptive and rational strategy to deal with inconsistencies between the efforts they expend and the ensuing rewards.


Aye, and I saw Sisyphus in violent torment, seeking to raise a monstrous stone with both his hands.

Homer, Book XI of The Odyssey

The adage “it was well worth the effort” highlights an assumed interdependency between attainment of reward and retrospective effort assignment. Despite its ubiquity we know little about the nature of these retrospective effort estimations. Previous studies have focused on the interaction of effort and reward as costs and benefits when a choice is to be made (1–5). Whether receipt of reward influences a retrospective estimation of effort is not known. Intuitively, we assume to have immediate and unbiased access to internal representations of how much “effort” we expended in an endeavor. Here, we demonstrate that retrospective estimation of effort is strongly affected by the amount of monetary reward attained and, as such, is profoundly biased. This bias adheres to established principles of Bayesian cue integration (6–8) and, on this basis, is not irrational.

In a behavioral experiment, participants pressed two buttons on a keyboard to push a ball up a virtual ramp and rated their experienced physical effort in each trial (Fig. 1; see also SI Materials and Methods for additional information regarding the task). The ball rolled back by a constant amount on each frame of the display, hence simulating a gravity force that varied so as to manipulate task difficulty (n = 6 difficulty levels, adjusted individually for each participant). Successful trials where participants managed to push the ball all of the way up the ramp were rewarded. Reward was contingent upon task difficulty (with values drawn from six Gaussian distributions with means from 1.5 to 6.5 cents) and the strength of this contingency varied across different blocks of the experiment (SDs of 1.2 or 2.5 cents). Additionally, we included a control experiment in which reward receipt was unrelated to task difficulty (SD = ∞).

Fig. 1.

Fig. 1.

Participants were asked to move a ball up a ramp by engaging in fast, alternating key presses. A gravity force was simulated, displacing the ball backward by a constant amount on each display frame. We used six levels of task difficulty, corresponding to the amount of ball displacement per time frame. After the ball was successfully pushed all of the way to the top of the ramp, participants received a monetary reward, where reward amount was contingent upon task difficulty. The strength of this contingency was varied in two separate blocks. Reward receipt information was either displayed before (90%) or after the rating of effort (10% of trials, not shown here). Subjects rated their effort by shifting the position of a sliding bar. At the end of each trial, they were asked to indicate whether they had seen a color change in the ball. All intervals in a trial were self-paced except for outcome reward display, which was in view for 2–3 s.

The reward information was presented either before or after the rating of effort (in 90% and 10% of trials, respectively). Trials in which reward was shown after the estimation of effort served as a reference, because here subjects are not influenced by preceding reward information. Participants were instructed to pay attention to all information presented in a trial, including a brief color change of the ball (50% of the trials), which they needed to detect on each trial. This manipulation was implemented to distract subjects from the true purpose of experiment, discouraging ad hoc strategies that might link effort and reward (also see SI Materials and Methods). Twenty-six participants (15 females, age 20–39 y, mean: 27.07 ± 5.1 y) took part in our main experiment. Two participants were excluded from the analysis because in our debriefing they mentioned they had not paid attention to the reward magnitudes during the experiment. Fourteen participants (nine females, age 21–35 y, mean: 27.5 ± 4.1 y) participated in the control experiment. Participants gave oral and written consent for their attendance. The study was approved by the local ethics committee of Berlin Charité University Hospital.

In trials where reward information was presented before effort rating, we examined whether reward magnitude influenced the effort estimation (Fig. 2). We measured the regression slope between trial-by-trial variations in reward (Ri − µr, with Ri being the reward on each trial in cents and µr being the mean reward of each difficulty level) and estimated effort (Ei − µe, with Ei being the estimated effort on each trial and µe being the mean estimated effort of each difficulty level; see also Fig. S1). In both blocks with different reward contingencies (red vs. blue bars in Fig. 2), there was a significant relationship between reward variation and estimated effort (Wilcoxon sign rank test, P = 0.00094 for SD = 1.2 cents and P = 0.002 for SD = 2.5 cents). This effect of reward on estimated effort was stronger when reward variance was smaller, that is, when reward was highly contingent upon the task difficulty (mean slopes of 0.012 and 0.005 for SD = 1.2 and SD = 2.5 cents, respectively; Wilcoxon sign rank test for the difference of both slopes, P = 0.004). This result was highly consistent across subjects (Fig. S2). We observed the same pattern of results when reward magnitude was balanced across blocks using a stratification method (Fig. S3). In a control experiment, where reward was randomly varied and unrelated to the task difficulty (SD = ∞), regression slopes did not differ from zero (Wilcoxon sign rank test, P = 0.15; for individual data see Fig. S4). These results indicate that reward influences effort estimation only when there is a reliable relationship with task difficulty, and hence a reliable relationship with the true exertion level subjects expend while pushing the ball.

Fig. 2.

Fig. 2.

(A) Relationship between trial-by-trial fluctuations in reward and retrospective estimates of expended effort in a typical subject tested with high (SD = 1.2 cents, data shown in red) and low (SD = 2.5 cents, data shown in blue) reward contingencies. The estimated effort Ei is normalized to each subject’s maximum estimated effort in the whole experiment, and the mean effort µi of each difficulty level is subtracted. The same procedure is implemented in relation to rewards Ri and their means µr. The slope β of the linear regression is larger for the higher reward contingency. In a control experiment, we tested whether random variations of reward (SD = ∞, data shown in black) affect estimates of effort. Here, slopes did not significantly differ from zero (P = 0.15; Wilcoxon sign rank test.) (B) Regression slopes of all individual subjects in both experiments (colors as in A). Average slopes differ from zero only when reward is contingent on task difficulty (P levels from Wilcoxon sign rank test). (C) Average regression slopes across subjects are higher when reward is more contingent on task difficulty than when less contingent (P = 0.004, Wilcoxon sign rank test). Error bars denote SEM. Regression results are based on a robust regression analysis (“robustfit” in MATLAB with default settings) that minimizes the effect of potential outliers.

Fig. S1.

Fig. S1.

Mean effort (µi) for each of the six difficulty levels i. µi increases monotonically with task difficulty. In our analysis, to assess the effect of reward independently of task difficulty, µi is subtracted from participants’ rated efforts (Fig. 2A). The resulting effort fluctuations (Ei − µi) are therefore only influenced by variations of reward (also see Fig. 2 and Fig. S3).

Fig. S2.

Fig. S2.

Correlation between variations in rewards and estimated effort in individual subjects. Data in red correspond to the block with stronger contingency between reward and task difficulty (SD = 1.2), and data in blue correspond to the block with a weaker contingency (SD = 2.5).

Fig. S3.

Fig. S3.

Regression slopes of individual subjects (A) and average regression slopes (B) for stratified data (colors as in Fig. 2). Difference in contingency between reward and task difficulty entailed that reward spread could differ across blocks (compare reward spreads in Fig. 2A), which might in turn affect regression slopes. We therefore used a stratification method to balance reward magnitudes across blocks and checked whether the observed slope differences do still hold. Stratification was done by defining six equally spaced bins between minimum and maximum reward (Ri − µr). We randomly removed surplus trials until trial number was the same for the two blocks in every bin. We recomputed the regression slopes for the stratified data. As shown in this figure, all our results did also hold for the stratified data (see also Fig. 2 B and C). In all our regression analyses (in the main text as well as in SI Materials and Methods), we used a robust regression analysis (‟robustfit” in MATLAB with the default bisquare weighing function). This was done to minimize the contribution of potential outliers on the regression slopes.

Fig. S4.

Fig. S4.

Rewards that are not contingent on task difficulty (SD = ∞) have no effect on estimated effort.

We next compared our behavioral data to predictions arising out of five separate computational models (Fig. 3). At the time when participants estimate their effort, the requirement was to retrospectively recall their true effort level. This recalled effort, Em (gray distribution in Fig. 3A), is predictive of the true effort but is corrupted by noise, as reflected in the variance σm2. When reward is correlated with task difficulty, reward magnitude yields by itself an independent effort estimation Er with uncertainty σr2 (red distribution in Fig. 3A). In each trial, Er and Em can differ by a certain amount (ΔE ≠ 0). The models we consider are based on distinct ways in which a conflict between these two informational sources (Er and Em) might be resolved.

Fig. 3.

Fig. 3.

(A) Two types of information can be combined to derive the estimated effort (Ê): Em is the recalled effort that is distorted by memory noise (σm); Er is the most likely effort level given the reward magnitude in each trial, its variability σr being dependent on the contingency of reward and task difficulty. According to Bayes optimal models, the influence of each signal on Ê depends on its reliability. In blocks with low contingency, σr is large, and Er, has a weaker influence on the estimate Ê, whereas for high reward contingency Ê is closer to Er (Upper and Lower, respectively). In each trial, Em and Er differ by a certain amount ΔE (ΔE ≠ 0). (B) Model comparison: BIC weights indicate the weight of evidence in favor of each model. (C) Using the weight ratio ωm/ωr derived from the data in one block (squares) and the known reward variability σr of both blocks, the weight ratio ωm/ωr in the other block (circles) can be accurately predicted (see also Supporting Information). Error bars denote SEM.

In all models, Em is computed based on the trials where reward was presented after the estimation of effort (i.e., where the estimation of effort is unaffected by reward information). Er is computed based on the posterior probability distribution of task difficulty given an obtained reward (for details see SI Materials and Methods). Model 1 (memory only) assumes that participants solely rely on the recalled effort (Em) and completely ignore reward information. Hence, the effort estimate E^ is equal to Em multiplied by a scaling factor km:

E^=EmKm. [1]

By contrast, model 2 (reward only) relies completely on reward information, whereas the recalled effort (Em) is disregarded:

 E^=ErKr. [2]

Model 3 assumes that both Em and Er contribute to effort estimation with E^ being a weighted average (WA) of Em and Er:

E^=ωEm+(1−ω)Er. [3]

Models 1–3 all assume that information regarding each signal’s variance is not explicitly exploited by the participants. However, model 4 (Bayes optimal, BO) assumes that a Bayesian optimal strategy is used by the participants where signals are weighted based on their respective reliability or inverse variance:

E^=ωmEm+ωrEr, [4]

where ωm=1−ωr and ωr= 1/σr2/(1/σm2+1/σr2).

Similar reliability-based Bayesian models have been used previously to explain integration of sensory cues during perceptual decision making (6–8). The variances σm2 and σr2 are derived from the data and reward probability distributions. It is, however, debatable whether σm2 is indeed inferred correctly using the trials where reward was presented after the rating of effort. Instead, the true variance σm2 might be a multiple of the variance in these trials. Therefore, model 5 (adapted Bayes optimal, aBO) is a modified version of model 4, assuming that the variance of Em (σm2) is scaled by a free parameter k (σm^2=kσm2):

E^=1σr21kσm2+1σr2Em+(1−1σr21kσm2+1σr2)Er. [5]

We evaluated these models by computing their maximum likelihood fits to the trial-by-trial data of individual subjects, measuring the quality of fits by Bayesian information criterion (BIC, SI Materials and Methods). BIC weights shown in Fig. 3B are the weight of evidence in favor of each model (9, 10). The average BIC weights are highest for the Bayesian weighted averaging model (see also Table S1). Importantly, in most subjects the model that merely relied on the memorized effort alone (without assuming any reward influence) performed considerably worse in terms of explaining the data (Table S2).

Table S1.

Modeling results averaged across subjects

Model No. of parameters BIC ΔBIC BIC weight
1 (MO) 1 −113.53 46.52 0.0689
2 (RO) 1 −69.07 90.97 0.0001
3 (WA) 1 −155.37 4.67 0.1277
4 (BO) 0 −158.7085 1.34 0.6177
5 (aBO) 1 −156.4995 3.55 0.1855

Table S2.

Modeling results for individual subjects (n = 24)

BIC BIC weights
Subjects MO RO WA BO aBO MO RO WA BO aBO
1 −175.03 −167.45 −197.27 −202.25 −197.05 0.0000 0.0000 0.0715 0.8644 0.0641
2 −205.87 −182.26 −261.87 −268.74 −263.78 0.0000 0.0000 0.0289 0.8962 0.0750
3 −33.622 7.0498 −30.45 −35.568 −32.806 0.2215 0.0000 0.0453 0.5859 0.1473
4 −81.158 −111.84 −156.97 −162.61 −158.21 0.0000 0.0000 0.0509 0.8544 0.0947
5 −71.123 −69.995 −137.42 −139.4 −137.58 0.0000 0.0000 0.2094 0.5641 0.2264
6 −142.58 −138.78 −225.23 −229.3 −224.42 0.0000 0.0000 0.1075 0.8210 0.0715
7 10.735 72.853 −29.768 −35.432 −30.125 0.0000 0.0000 0.0522 0.8855 0.0624
8 9.7384 39.692 −18.855 −24.8 −19.406 0.0000 0.0000 0.0458 0.8940 0.0603
9 −104.23 −63.194 −162.66 −168.19 −166.11 0.0000 0.0000 0.0445 0.7054 0.2501
10 −606.66 −521.47 −673.45 −668.16 −670.93 0.0000 0.0000 0.7382 0.0523 0.2094
11 −95.44 −17.079 −128.83 −128.54 −128.71 0.0000 0.0000 0.3566 0.3075 0.3359
12 −123.41 −59.395 −115.25 −108.39 −118.97 0.8881 0.0000 0.0150 0.0005 0.0965
13 −191.23 −34.226 −185.17 −194.97 −190.03 0.1236 0.0000 0.0060 0.8026 0.0679
14 −106.22 −162.15 −190.13 −194.66 −191.64 0.0000 0.0000 0.0781 0.7555 0.1663
15 −184.42 −26.206 −181.8 −184.45 −179.51 0.4212 0.0000 0.1137 0.4289 0.0362
16 −120.16 −53.503 −180.91 −186.64 −183.57 0.0000 0.0000 0.0448 0.7863 0.1689
17 8.6609 27.09 −43.955 −50.88 −47.31 0.0000 0.0000 0.0261 0.8339 0.1399
18 −127.8 −179.65 −190.3 −189.39 −188.4 0.0000 0.0024 0.4935 0.3131 0.1910
19 −79.793 −14.468 −101.56 −104.51 −107.21 0.0000 0.0000 0.0450 0.1967 0.7583
20 −68.548 55.834 −105.27 −109.71 −104.93 0.0000 0.0000 0.0905 0.8330 0.0764
21 −109 −27.06 −122.41 −125.6 −121.03 0.0002 0.0000 0.1555 0.7665 0.0778
22 −15.643 69.586 −42.175 −46.313 −41.178 0.0000 0.0000 0.1050 0.8313 0.0638
23 188.52 199.45 111.75 107.82 113.27 0.0000 0.0000 0.1158 0.8298 0.0544
24 −300.66 −300.62 −359.09 −358.31 −366.35 0.0000 0.0000 0.0254 0.0172 0.9575

Bold numbers demarcate the best model for each subject.

Model 4 (BO) also holds that the ratio of weights ωm/ωr derived from the data in one block should be predictive of the ratio of weights in the other block, respectively, given the known variances σr2 of the reward signal in each block, and assuming the uncertainty of memory σm2 to be constant (see also Supporting Information). Fig. 3C shows that this prediction provides a good match to the data (P > 0.5, Wilcoxon sign rank, for comparison of observed and predicted ωm/ωr). The reliance of participants on the reward signal can thus be accurately predicted by its variance. Finally, the variance of joint estimates was on average smaller than the variance of each signal alone, supporting that the Bayes optimal strategy participants use for signal integration improves their general precision in effort estimation (Fig. S5).

Fig. S5.

Fig. S5.

(A) Estimated effort (E^), shown in magenta, is computed based on Em (shown in gray) and Er (shown in red). Each of these estimates has its corresponding uncertainty reflected in the SD of the distribution (σe, σm, and σr, respectively). Bayesian optimal integration predicts σe to be smaller than both σm and σr. (B) SDs σe, σm, and σr, averaged across all subjects and both reward contingency blocks. σe is smaller than σm (P = 0.003) and σr (P = 0.01, Wilcoxon sign rank test).

In this study, we show that after receiving reward information, humans revise their estimation of effort required for its attainment. The strength of correlation between task difficulty and obtained reward had a profound influence on retrospective effort estimates. Importantly, participants were adept at adjusting their estimation of effort when reward contingencies changed within the same experimental session.

We note that in our experiments more difficult trials required larger number of key presses and therefore took more time to complete than was the case for easy trials, entailing a decreased reward density per unit time. Therefore, participants’ estimation of their effort might also be influenced by their estimation of trial time or perceived reward density. Although we cannot rule out a contribution from this factor, we would suggest that a dependency on time might be an inherent feature of effort estimations, because highly demanding tasks usually entail longer realization times.

How do these findings extend to real-life situations? In many instances, the effort we expend is closely tethered to consequential outcomes. For example, in the context of performance-based pay employees are remunerated in proportion to the degree to which a task is accomplished (11–13). Accordingly, there is a strong prior in many societies that rewards (wages) are contingent upon effort (labor invested). Equally, in fields such as sports (14, 15) and education (16–18), there is a common belief in a contingency between effort and success. In our study we show that such a prior is deployed by humans when they retrospectively evaluate effort. However, a direct comparison of the impact of real-life contingencies and the contingencies used in our paradigm is still missing. Moreover, in our task design and description participants may have paid more attention to reward, and the information conveyed by it, than is the case in real life. Future studies are needed to reveal whether, and to what extent, our findings generalize to other situations including real-life scenarios.

A question arises as to whether a flexible estimation of past effort serves a functional role. We suggest that rethinking one’s efforts after receiving rewards constitutes a metacognitive ability (19) to negotiate uncertainties in effort–reward relationships. In everyday life, rewards are usually correlated with effort, but the strength of this correlation may change. The here reported acute adjustment of effort estimation can greatly aid goal-directed behavior when decisions are based on online monitoring of environmental factors. Indeed, failure of such mechanisms might lead to occupational disorders such as burn-out syndrome that are thought to be related to perception of an effort–reward imbalance (20, 21).

The distorted effort rating observed in the current study also bears resemblance to hindsight bias (22), reflecting a tendency to change recalled probability estimations once outcomes are known. As in a hindsight bias, retrospective change in effort estimation might reflect a general tendency to reshape memory contents to make them fit with an updated knowledge base (23). A large number of other cognitive biases have been described in the past that also reflect humans’ deviation from rationality (24). Similarly, perceptual decisions are prone to deviate from the veridical as seen in phenomena such as sensory illusions (7, 25). Recent theoretical work has suggested that these “biases” reflect humans’ ability to deal with the uncertainties in the world using the probabilistic structure of the environment (26, 27). Therefore, seemingly erroneous judgments are not only very common in the course of evolution (28) but are also optimal and rational (29–31). The finding that the impact of rewards on retrospectively evaluated effort conforms to a Bayesian rule of cue integration is in line with these previous studies, extending them to a domain with relevance to most aspects of our daily life.

We have shown that human subjects either under- or overestimate their past effort when rewards are smaller or larger than average, respectively. Whether a similar tendency influences normative beliefs, for example regarding the distribution of wealth in society, is a potentially important avenue of further investigation. For example, individuals with higher incomes who are exposed to greater-than-average rewards might have an inflated perception of the effort they expended to acquire their wealth. Conversely, those with low income might have the opposite perception. One might also speculate that such a biased perception of effort could contribute to stabilizing inequality. Indeed, the increasing equality gaps in many societies reinforce the importance of gaining insight into the complex interplay between retrospective assignments and the emergence of socioeconomic norms.

SI Materials and Methods

Stimuli and Task.

Stimuli were produced with MATLAB and the Psychophysics Toolbox (32, 33). Each trial consisted of a stimulus (ball and ramp), a reward, and an effort rating display all shown on a black background (Fig. 1). The stimulus display contained the ball (radius: three visual degrees), initially at the starting part of the ramp (ramp length: 19 visual degrees; both ball and ramp had a light gray color). The ball was displaced up the ramp with consecutive alternate key presses (left and right arrow keys) until it reached the upper plateau. Each key press resulted in a constant amount of displacement (0.87 visual degrees per key press) and was counteracted by a gravity force of variable strength that displaced the ball backward. To determine the levels of gravity force, at the beginning of the experiment we asked each subject to push the ball up the ramp by pressing both keys alternately and consecutively as fast as they possibly could. Ninety percent of the gravity force necessary to counteract the maximum number of key presses in a limited time (10 s) determined the maximum gravity force used in the experiment. Based on this individualized estimate, six equally spaced gravity levels were defined and used in the experiment. A trial was aborted if key presses did not occur fast enough (maximum pause allowed was 2 s). If participants were able to successfully push the ball all of the way up, they received a monetary reward, with the amount contingent on task difficulty (gravity force level). Reward magnitude was defined based on six Gaussians with means of 1.5, 2.5, 3.5, 4.5, 5.5, and 6.5 cents, using two different SDs (SD of 1.2 or 2.5 cents) which determined the contingency between reward magnitude and task difficulty. Reward display consisted of a pie chart that depicted subjects’ reward as a proportion of maximum reward possible and a number that showed the reward in Arabic numerals. Effort rating display consisted of a slider, and participants were instructed to set the slider at a position that represented their experienced effort during a trial proportionate to the maximum effort they had ever experienced during the experiment. In 90% of trials, the reward display was shown immediately after the stimulus display, whereas in 10% of trials it occurred after the rating of effort.

The main experiment was done in two separate blocks in which different reward contingencies (SD of 1.2 or 2.5 cents) were used. Participants did not receive any instruction regarding the strength of contingency between task difficulty and reward and were only told that a change in the strength of this relationship would occur across blocks. The order of these blocks was counterbalanced across subjects. Owing to the random assignment of this order (i.e., whether the contingency of the first block was high or low), the direction of change across blocks (from weak to strong or vice versa) was also unknown to the experimenters. Therefore, participants had to experience the relationship between task difficulties and reward by themselves, in the absence of any prior knowledge regarding the strength of contingencies. Subjects first performed 12 training trials before the start of each block so that they were acquainted with the task and the reward–effort relationship. These trials were not included in our analyses. Each block consisted of six smaller miniblocks, each consisting of 36 trials. Participants could take a pause and rest between the miniblocks. In the control experiment, rewards were randomly chosen and varied between 0.5 and 5 cents, without any relationship to the difficulty levels.

In pilot experiments, some participants reported that to work out a relationship between task difficulty and reward they had adopted different ad-hoc strategies. Because task difficulty and reward were the only parameters that varied across trials, participants had presumably focused exclusively on relating the two. In the main experiment, to distract subjects from its true purpose we introduced a second task into the main paradigm. Thus, subjects were also asked to report whether they had seen a brief (21 ms or three frames) color change on the ball (to green, red, or blue) at the end of each trial, which occurred in 50% of the trials. We debriefed the participants after the experiments and note that none reported that the main question of the study involved effort and reward relationships, nor did they now report using ad-hoc strategies.

Computational Modeling.

Our models are based on two independent effort estimates: an estimate purely based on the subjective memory of the effort spent in a trial (Em), which is unaffected by reward information, and an estimate purely based on the information conveyed by reward magnitude (Er).

Em was derived from a Gaussian distribution with mean μm and variance σm2, which were computed based on the trials in which estimation of effort was unaffected by reward (i.e., where effort was rated after receiving reward). Effort estimates of these unbiased trials will be referred to as Eu. The average μm at a difficulty level di was calculated as the median of a participant’s rated efforts Eu in that difficulty level. The variance σm2 across trials t was calculated as

σm2=1n∑t=1n(Eut−μm)2. [S1]

To estimate Em in a given trial, a random value was then drawn from a Gaussian distribution with mean μm and variance σm2.

Er is the most likely effort level given a certain reward. It was derived from a Gaussian distribution with mean μr and variance σr2. To infer μr, we computed the most probable difficulty level d^ from all levels di, given a reward r, which can be obtained maximizing the posterior probability function

P(di|r)=P(r|di)P(di)/P(r), [S2]

where P(di) and P(r) are the probability of each difficulty level and each reward magnitude, respectively, and P(r|di) is the probability of a given reward at each difficulty level. The average estimate μr of all trials with a given d^ was computed as the median of estimated efforts Eu at that difficulty level. The variance σr2 across all trials t was approximated as

σr2=1n∑t=1n(μr−μm)2. [S3]

To estimate Er in a given trial, a random value was then drawn from a Gaussian distribution with mean μr and variance σr2.

In model 4 (BO), for each block, the corresponding σr2 was used; σm2 was assumed to remain unchanged across blocks. Model 5 (aBO) challenges the assumption that the variance σm2 is correctly reflected in the trials where participants rate their effort before receiving reward: The variance in these trials is influenced by motor noise, which might not be incorporated in σm. Moreover, at the time when subjects retrospectively evaluate their effort, the variance of their estimates might be affected by unknown factors (such as increased memory noise). Therefore, in model aBO, a free parameter (k) is used to scale σm2. All other aspects of model 5 remained the same as in model 4. In all models, we only included trials with a conflict between Em and Er, that is, where the estimate of task difficulty d^ provided by reward differed from the actual difficulty level. All trials of the two blocks (reward SD of 1.2 and 2.5) were modeled at once.

Model Comparison.

Models one, two, three, and five each had one free parameter: km, kr, ω, and k, respectively. Model 4 (BO) had no free parameter. The fit of each model (n = 5) to the data was evaluated by computing log-likelihood and BIC:

BIC=−2LL+ln(n)m, [S4]

where LL is the model log-likelihood and m is the number of free parameters. The individual BIC values contain arbitrary constants and are very much affected by sample size. We therefore rescaled BIC to obtain ∆BICs (10, 34), where

 ΔBIC=BICi−BICmin [S5]

and BICmin is the minimum of the N different BICi values. This transformation forces the best model to have ∆BIC = 0, whereas the rest of the models have positive values. The Bayesian weights are then derived from ∆BIC:

ωi=exp(−0.5ΔBICi)∑n−1N(−0.5ΔBICn). [S6]

The Bayesian weights (ωi) of all of the models in a set sum up to 1 and indicate the probability for each model to be the best model for the data.

Prediction of Bayesian Weights Using Reward Contingencies.

When the memory signal Em and the reward signal Er yield different effort estimates (ΔE ≠ 0), participants integrate both signals into an estimate Ê, using the weights ωm and ωr, respectively. The ratio of these weights can be directly derived from the data (see also Fig. 3A):

ωrωm=|Ê−Em|ΔE|Ê−Er|ΔE. [S7]

Model 4 also holds that the weights of the reward and memory signal are inversely proportional to their relative variance:

ωrωm=σm2σr2. [S8]

Assuming that the memory variance σm2 remains the same in the two blocks with different reward contingencies, we can compute the variances σr12 and σr22 as

σr12ωr1ωm1=σr22ωr2ωm2
ωr1ωm1=σr12σr22ωr2ωm2. [S9]

The variances σr12 and σr22 can be inferred from the known contingencies between reward and task difficulty (Eq. S2). Thus, Eq. S9 involves a direct prediction of the ratio of weights in one block from the ratio of weights in the other block (as shown in Fig. 3C).

Acknowledgments

We thank Alexandra Klein, Rafaela Wahl, and Valerie Keller for their help with the data collection, and Ulf Toelch for his advice on the modeling analysis. This work was performed while Ray Dolan was a Visiting Einstein Fellow at the Humboldt-Universität, Berlin School of Mind and Brain and was supported by Wellcome Trust Senior Investigator Award 098362/Z/12/Z (to R.J.D.).

Footnotes

The authors declare no conflict of interest.

This article is a PNAS Direct Submission. P.W.G. is a guest editor invited by the Editorial Board.

This article contains supporting information online at www.pnas.org/lookup/suppl/doi:10.1073/pnas.1507527112/-/DCSupplemental.

References

  • 1.Pessiglione M, et al. How the brain translates money into force: A neuroimaging study of subliminal motivation. Science. 2007;316(5826):904–906. doi: 10.1126/science.1140459. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Meyniel F, Sergent C, Rigoux L, Daunizeau J, Pessiglione M. Neurocomputational account of how the human brain decides when to have a break. Proc Natl Acad Sci USA. 2013;110(7):2641–2646. doi: 10.1073/pnas.1211925110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Skvortsova V, Palminteri S, Pessiglione M. Learning to minimize efforts versus maximizing rewards: Computational principles and neural correlates. J Neurosci. 2014;34(47):15621–15630. doi: 10.1523/JNEUROSCI.1350-14.2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Kurniawan IT, Guitart-Masip M, Dolan RJ. Dopamine and effort-based decision making. Front Neurosci. 2011;5:81. doi: 10.3389/fnins.2011.00081. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Wardle MC, Treadway MT, Mayo LM, Zald DH, de Wit H. Amping up effort: Effects of d-amphetamine on human effort-based decision-making. J Neurosci. 2011;31(46):16597–16602. doi: 10.1523/JNEUROSCI.4387-11.2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Ernst MO, Banks MS. Humans integrate visual and haptic information in a statistically optimal fashion. Nature. 2002;415(6870):429–433. doi: 10.1038/415429a. [DOI] [PubMed] [Google Scholar]
  • 7.Alais D, Burr D. The ventriloquist effect results from near-optimal bimodal integration. Curr Biol. 2004;14(3):257–262. doi: 10.1016/j.cub.2004.01.029. [DOI] [PubMed] [Google Scholar]
  • 8.Körding KP, Wolpert DM. Bayesian integration in sensorimotor learning. Nature. 2004;427(6971):244–247. doi: 10.1038/nature02169. [DOI] [PubMed] [Google Scholar]
  • 9.Burnham KP, Anderson DR. Model Selection and Multimodel Inference: A Practical Information-Theoretic Approach. Springer; New York: 2015. [Google Scholar]
  • 10.Wagenmakers EJ, Farrell S. AIC model selection using Akaike weights. Psychon Bull Rev. 2004;11(1):192–196. doi: 10.3758/bf03206482. [DOI] [PubMed] [Google Scholar]
  • 11.Lazear EP. Salaries and piece rates. J Bus. 1986;59(3):405–431. [Google Scholar]
  • 12.Stiglitz JE. Alternative theories of wage determination and unemployment in LDC’s: The labor turnover model. Q J Econ. 1974;88(2):194–227. [Google Scholar]
  • 13.Nalebuff BJ, Stiglitz JE. Prizes and incentives: Towards a general theory of compensation and competition. Bell J Econ. 1983;14(1):21–43. [Google Scholar]
  • 14.Van-Yperen NW, Duda JL. Goal orientations, beliefs about success, and performance improvement among young elite Dutch soccer players. Scand J Med Sci Sports. 1999;9(6):358–364. doi: 10.1111/j.1600-0838.1999.tb00257.x. [DOI] [PubMed] [Google Scholar]
  • 15.Si G, Rethorst S, Willimczik K. Causal attribution perception in sports achievement: A cross-cultural study on attributional concepts in Germany and China. J Cross Cult Psychol. 1995;26(5):537–553. [Google Scholar]
  • 16.Frieze IH, Snyder HN. Children’s beliefs about the causes of success and failure in school settings. J Educ Psychol. 1980;72(2):186–196. [PubMed] [Google Scholar]
  • 17.Tollefson N. Classroom applications of cognitive theories of motivation. Educ Psychol Rev. 2000;12(1):63–83. [Google Scholar]
  • 18.Livengood J. Students’ motivational goals and beliefs about effort and ability as they relate to college academic success. Res Higher Educ. 1992;33(2):247–261. [Google Scholar]
  • 19.Fleming SM, Dolan RJ, Frith CD. Metacognition: Computation, biology and function. Philos Trans R Soc Lond B Biol Sci. 2012;367(1594):1280–1286. doi: 10.1098/rstb.2012.0021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Siegrist J. Adverse health effects of high-effort/low-reward conditions. J Occup Health Psychol. 1996;1(1):27–41. doi: 10.1037//1076-8998.1.1.27. [DOI] [PubMed] [Google Scholar]
  • 21.de Jonge J, Bosma H, Peter R, Siegrist J. Job strain, effort-reward imbalance and employee well-being: A large-scale cross-sectional study. Soc Sci Med. 2000;50(9):1317–1327. doi: 10.1016/s0277-9536(99)00388-3. [DOI] [PubMed] [Google Scholar]
  • 22.Roese NJ, Vohs KD. Hindsight bias. Perspect Psychol Sci. 2012;7(5):411–426. doi: 10.1177/1745691612454303. [DOI] [PubMed] [Google Scholar]
  • 23.Hoffrage U, Hertwig R, Gigerenzer G. Hindsight bias: A by-product of knowledge updating? J Exp Psychol Learn Mem Cogn. 2000;26(3):566–581. doi: 10.1037//0278-7393.26.3.566. [DOI] [PubMed] [Google Scholar]
  • 24.Kahneman D, Slovic P, Tversky A. Judgment Under Uncertainty: Heuristics and Biases. Cambridge Univ Press; New York: 1982. [Google Scholar]
  • 25.Geisler WS, Kersten D. Illusions, perception and Bayes. Nat Neurosci. 2002;5(6):508–510. doi: 10.1038/nn0602-508. [DOI] [PubMed] [Google Scholar]
  • 26.Trommershäuser J. Biases and optimality of sensory-motor and cognitive decisions. Prog Brain Res. 2009;174:267–278. doi: 10.1016/S0079-6123(09)01321-1. [DOI] [PubMed] [Google Scholar]
  • 27.Summerfield C, Tsetsos K. Do humans make good decisions? Trends Cogn Sci. 2015;19(1):27–34. doi: 10.1016/j.tics.2014.11.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Marshall JA, Trimmer PC, Houston AI, McNamara JM. On evolutionary explanations of cognitive biases. Trends Ecol Evol. 2013;28(8):469–473. doi: 10.1016/j.tree.2013.05.013. [DOI] [PubMed] [Google Scholar]
  • 29.Fennell J, Baddeley R. Uncertainty plus prior equals rational bias: An intuitive Bayesian probability weighting function. Psychol Rev. 2012;119(4):878–887. doi: 10.1037/a0029346. [DOI] [PubMed] [Google Scholar]
  • 30.Trimmer PC, et al. Decision-making under uncertainty: Biases and Bayesians. Anim Cogn. 2011;14(4):465–476. doi: 10.1007/s10071-011-0387-4. [DOI] [PubMed] [Google Scholar]
  • 31.Bach DR, Dolan RJ. Knowing how much you don’t know: A neural organization of uncertainty estimates. Nat Rev Neurosci. 2012;13(8):572–586. doi: 10.1038/nrn3289. [DOI] [PubMed] [Google Scholar]
  • 32.Brainard DH. The Psychophysics Toolbox. Spat Vis. 1997;10(4):433–436. [PubMed] [Google Scholar]
  • 33.Pelli DG. The VideoToolbox software for visual psychophysics: Transforming numbers into movies. Spat Vis. 1997;10(4):437–442. [PubMed] [Google Scholar]
  • 34.Burnham KP, Anderson DR. Multimodel inference: Understanding AIC and BIC in model selection. Sociol Methods Res. 2004;33(2):261–304. [Google Scholar]

Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES