Abstract
The rat is a common animal model used to uncover the neural underpinnings of decision-making and their disruption in psychiatric illness. Here, we ask if rats can perform a decision-making task that assesses self-control by delayed gratification in the context of diminishing returns. In this task, rats could choose to press one of two levers. One lever was associated with a fixed delay (FD) schedule that delivered reward after a fixed time delay (10s). The other lever was associated with a progressive delay (PD) schedule; the delay increased by a fixed amount of time (1 s) after each PD lever press. Rats were tested under two conditions: 1) a reset condition where rats could reset the PD schedule back to its initial 0s delay by pressing the FD lever and 2) a no-reset condition where resetting the PD schedule was unavailable. We found that rats adapted behavior within reset sessions by delaying gratification to obtain more reward in the long run. That is, they selected the FD lever with a longer delay to reset the PD delay back to zero prior to the equality point, thus achieving more reward over the course of the session. These results are consistent with other species, demonstrating that rats can also maximize the net rate of reward by selecting an option that is not immediately beneficial. Moreover, use of this task in rodents might provide insights into how the brain governs normal and abnormal behavior, as well as treatments that can improve self-control.
Keywords: decision-making, self-control, diminishing returns, delay, reward, rat
One endeavor of behavioral neuroscientists is to understand the neural mechanisms of self-control and their disruption in psychiatric disease. To accomplish this, researchers have taken advantage of cross-species tasks (e.g., reversal learning, go/no-go tasks, stop-signal tasks and delay-discounting tasks) that allow for investigation of such mental processes at multiple levels of investigation (Bari, & Robbins, 2013; Brockett et al., 2018). These tasks have led to great insights into how the brain governs normal and abnormal behavior, as well as treatments that can improve self-control. A good example of such progress comes from delay-discounting (aka intertemporal choice), a task that pits the choice of a small immediate reward against the choice of a larger delayed reward in order to quantify an individual’s level of impulsivity (Bickel, et al. 2007).
Although delay-discounting has provided invaluable information to the field, other tasks might prove to be more beneficial in assessing self-control (Hayden, et al. 2016). One potentially interesting candidate that also examines sensitivity to delayed reward is the Diminishing Returns task, first developed in chimpanzees (Pan troglodytes) by Hodos and Trumbule (1967). Much like delay-discounting, the value of future reward is decreased by the expected delay that precedes it. However, during Diminishing Returns tasks the delay preceding reward is constantly changing based on the animal’s selection. Specifically, during performance of Diminishing Returns tasks the persistent selection of one of two options results in a declining rate of gain either in the number of responses necessary to obtain reward (i.e., progressive ratio; PR) or how long the animal has to wait for reward (i.e., progressive delay; PD). The other option is ‘fixed’ in that the same amount of work or time is required to obtain rewards across the entire experiment. Importantly, the fixed option is always initially set higher than the progressive option, thus the optimal strategy is to start by selecting the progressive option and then switch to the fixed option when the cost (amount of work or time) is equal between the two (i.e., equality point).
To provide an example, let us consider a situation where an experimental animal has two routes to reward, one that is fixed, leading to reward after 10s (Fixed Delay; FD-10) and another that is progressive, starting with a 0s delay that increases progressively by 1s every time it is chosen (Progressive Delay; PD-1). The optimal strategy under these conditions is to select the PD option 11 times and then switch to the FD option for the remainder of the experiment. This is similar to choosing between two equally preferred restaurants based on how quick the service is. If service became progressively slower at one, then you would eventually switch to the other restaurant. In the wild, this concept relates to animals foraging water holes or food supplies; the more often one site is visited, the longer it would take to obtain nourishment. Not surprisingly, animals perform well at such tasks.
What is surprising and, in our opinion, more interesting, is that animals also perform well in a version of the task where selection of the fixed option ‘resets’ the cost associated with the PD option. In the example above, this would be a situation where selection of the FD option resets the delay associated with the PD option back to zero. This manipulation sets up a scenario where the short term consequences favor switching at the equality point (i.e., after 11 selections of the PD option), but long-term consequences favor switching much earlier (i.e., switching is optimal after 4 selections of the PD option). This would be like preemptively switching restaurants long before service became too slow with the notion that it might cause the owner to make improvements. Performing optimally on this task requires self-control and consideration of the outcome of future choices. It also demonstrates how selection of an immediately unfavorable option can have long-term benefits (delayed gratification), similar to many real-life scenarios of complex decision-making that currently are not properly addressed by intertemporal choice tasks. Remarkably, despite the immediate cost of selecting the fixed option (e.g., wait 10s instead of 4s because it resets the PD back to 0 s), chimpanzees, humans, and pigeons have been shown to switch prior to the equality point, at near optimal levels (Hackenberg & Axtell, 1993; Hackenberg & Hineline, 1992; Hodos & Trumbule, 1967; Jacobs & Hackenberg, 1996).
Given this literature it is clear that Diminishing Returns tasks can be an important tool for understanding the neural mechanisms of complex decision-making. In order to work towards that goal, however, this task must first be established in an animal model conducive to studies with modern techniques from the field of neuroscience. To address this issue, we show that rats also adjust their behavior under ‘reset’ conditions during performance of three different time-based Diminishing Returns tasks.
Materials and Methods
Subjects
Six female and six male Sprague-Dawley rats (236–267g) were purchased from Charles River, and individually housed in a temperature and humidity-controlled vivarium on a diurnal light cycle (12 h light/dark). Rats were food restricted and maintained at 85% of their initial weight. Following experiment completion, rats were euthanized via isoflurane overdose. Rats were tested at the University of Maryland in accordance with National Institute of Health (NIH) guidelines and the Institutional Animal Care and Use Committee (IACUC) at the University of Maryland, College Park.
Behavioral apparatus
All behavioral procedures were conducted in two modular shuttle boxes (ENV-010MD, Med Associates, St. Albans, VT). Each shuttle box was contained within a sound and light-attenuating cabinet. The end wall of each box was equipped with a stimulus light, retractable lever, and pellet trough connected to a pellet dispenser (Med Associates ENV-221M, ENV-112CM, ENV-200R2M-6, and ENV-203M-45, respectively). The food troughs were positioned on the left half of the end-walls (when facing the wall from inside the chamber), the retractable levers were positioned in the lower right so that when the levers were extended, they were ≈ 2cm above the grid floor (Med Associates ENV-010MB-GF), and the stimulus lights were positioned immediately above the levers.
Diminishing Returns tasks
After shaping (see supplemental), rats performed three different permutations of the Diminishing Returns task. These different variations consisted of: a PD-1/FD-10 schedule pair where the progressive lever delay incremented 1s with each press, while the fixed lever maintained a delay at 10 s; a PD-2/FD-20 schedule pair where the progressive delay increment was changed to 2s and the fixed delay was increased to 20 s; and a PD-1/FD-20 schedule pair where the PD increment was 1s and the FD delay remained at 20 s. Across all Diminishing Returns variations, each 1 h session was initiated with insertion of both levers into the box. When one lever was pressed, both levers were retracted, and the cue light adjacent to the pressed lever was illuminated until two seconds after the reward was delivered to the pellet trough on the same side. If a rat did not press a lever within 30s of their extension into the box the trial was recorded as an omission, and a 30s inter-trial interval was initiated.
The PD-1/FD-10 schedule lasted for 20 days (Figure 1A). Half of the rats performed 10 days of reset and then 10 days of no-reset, while these reset/no-reset conditions were reversed for the other half of subjects. Lever assignments were additionally switched after 10 days in order to prevent side biases in lever-responding. As illustrated in Figure 1B, the PD lever delivered rewards with a progressively increasing reward that maximized to 50 s, while the FD lever delivered reward after a constant 10s delay. During reset sessions, pressing the FD lever reset the delay on the PD lever to 0s (Fig 1B; ‘Reset’). During no-reset sessions (Fig 1B; ‘No-reset’), pressing the FD-10 lever had no additional consequence beyond pellet delivery. That is, the delay on the PD lever continued to increase even after the FD lever was pressed. Due to a pellet dispenser malfunction data from sessions 5–7 were not included in any analyses.
Figure 1. Task schematic and training timeline.

Note. a Summary of training and testing schedule. For the FD-10/PD-1 task, rats were counterbalanced within subjects among reset and no-reset conditions. In the other two iterations of the Diminishing Returns task, each condition had 6 rats (3F, 3M). b In the Diminishing Returns task, rats choose between two levers. One lever delivers reward after a fixed delay (FD) schedule, and the other delivers reward on a progressive delay (PD) schedule where the delay increases by a fixed amount of time with each lever press. In the reset condition (right), the rat can reset the PD delay back to zero by selecting the FD lever. In the no-reset condition (left), the PD delay continues to increase regardless of FD lever selection. Optimal performance entails switching from the PD to the FD after only 4 PD presses – when the delay on the current trial is 3s. This sequence of choices yields the highest arithmetic rate of reinforcement, computed by dividing the number of reinforcers in the sequence by the cumulated time. The mathematical demonstration of the rationale of why 4 lever presses followed by reset is the optimal strategy is further illustrated in Supplemental Figure 1.
For the next 5 days, rats worked for rewards on a PD-2/FD-20 schedule so that reward delivery on the PD lever was initially 0s but was incremented by 2 s, and the FD lever consistently delivered reward after a 20s delay. The reset conditions and PD/FD lever assignments for subjects were maintained from the last 10 days of the PD-1/FD-10 Diminishing Returns task so that there were 6 rats (3M, 3F) in each reset/no-reset condition. On the last 5 days, rewards were delivered on a PD-1/FD-20 schedule, which returned the PD increment to 1s while the FD delay remained at 20 s. This schedule also had 6 rats (3M, 3F) in each reset/no-reset condition.
Data Analyses
Optimal responding on each Diminishing Returns schedule under both reset and no-reset conditions is defined as the behavior that maximizes the reward gained in a single session (Hackenberg & Axtell 1993). For the first pair of schedules, PD-1 and FD-10, under reset conditions optimal responding entails pressing the PD lever four times consecutively and then pressing the FD lever once, to reset the PD schedule (Figure 1B; ‘Reset’). Thus, under these conditions an optimally-behaving rat will respond on the PD lever with an 80% probability and the FD lever with a 20% probability. Under no-reset conditions (Figure 1B; ‘No-reset’), optimal behavior entails eleven presses on the PD lever (i.e., delay of 10s following the last press), followed by all other presses on the FD lever. Thus, the optimal probabilities for responding on each lever under no-reset conditions depends on the number of trials completed within that session (the optimal PD probability equals eleven divided by the total number of completed trials). For the second pair of schedules PD-2 and FD-20, optimal behavior is the same as PD-1 and FD-10. For PD-1 and FD-20, optimal behavior under reset conditions entailed pressing the FD lever once after six consecutive presses on the PD lever, thus optimal probabilities of responding on the PD and FD levers were 85.7% and 14.3%, respectively. Under no-reset conditions for this pair of schedules optimal behavior involves initially responding on the PD lever 21 times followed by all other responses on the FD lever, and again the optimal probabilities depend on the total number of completed trials within the session. These predictions for optimal Diminishing Returns performance come from Charnov’s Marginal Value Theorem (MVT), which describes optimal behavior while foraging in a habitat that contains patches that can be depleted of their resources (Charnov, 1976;Addicott et al., 2015; Pleasants, 1989; McGinley, 1984).
Optimal behavior is illustrated in Figure 1. Rats should begin each session by responding on the PD lever because the initial delay on this lever is 0 s, which is less than the FD (e.g., 10s). Under no-reset conditions the optimal number of responses to make on the PD lever before switching to the FD lever is the number of responses that results in a PD that equals the FD. Afterwards, rats should continue to respond on the FD lever and never again respond on the PD lever, because the delay for doing so will exceed the FD. To determine the optimal number of responses to make on the PD lever before switching to the FD lever under reset conditions we first consider that the FD lever should only ever be pressed once in a row, because it will reset the PD schedule, making the delay for selecting the PD lever 0 s. Then we determine the average reward rate achieved by pressing the PD lever X number of times before pressing the FD lever. These determinations were made starting with X=0 and incrementing X until the maximum average reward rate was identified.
To compare a subject’s median number of consecutive presses on each lever to a) population values and b) optimal behavior predicted by Charnov’s MVT, we used lognormal generalized linear mixed models (GLMMs) to analyze the data both across and within sessions. We chose to analyze medians as opposed to means because they better represent non-symmetric distributions like the lognormal distribution (Bijma, Jonker, & Van der Vaart, 2017; Forbes, et al. 2011). For analyses of the dependent variables by session, we considered for inclusion in the final model a fixed effect of the reset-condition, a discrete fixed effect of the session, and an interaction between these effects. We additionally considered a random intercept and random effect of the session for each subject, or for each subject within each reset condition when applicable. For analyses of the dependent variables by trial, we considered for inclusion in the final model a fixed effect of the reset-condition, a continuous covariate of the trial number, and an interaction between these effects, as well as a random intercept and random effect of the trial number for each subject, or for each subject within each reset condition when applicable.
The statistical models were analyzed for main effects and interactions using F-tests (Stroup 2012). Significant main effects of reset condition, in the absence of an interaction, were further explored with a t-test. Interactions between reset condition and session or trial were further explored using t-tests to conduct a simple effects comparison of the reset conditions within each session or trial. Wald goodness-of-fit tests were used to compare statistical models to the optimal values, and to make comparisons between models. All statistical analyses were conducted using SAS 9.4 (SAS Institute, Cary, NC). Only data from the Diminishing Returns tasks were analyzed. The distribution for each model was determined by comparing Aikaike’s Information Criterion scores, a measure of model parsimony, for the available distributions in the GLIMMIX procedure using the default link functions.
Results
PD-1 and FD-10 schedule pair
As a first pass to determine if rats performed appropriately during reset and no-reset sessions, we first asked if rats responded more or less on the PD lever than the FD over the entire session. Generally, the PD lever should be selected more often because delays to reward remain short due to the resetting of the delay to zero upon selection of the FD lever. Indeed, we found that under reset conditions a rat was significantly more likely to respond on the PD lever (86.4% ± 1.3%) than the FD lever for data pooled over all sessions (Wald test; p < 0.0001; Figure 2A dark gray thick solid vs. light gray thin solid), and for each individual session (80.4% ± 3% to 91% ± 1.8%, all ps < 0.0001; Supplemental Figure 2A). Remarkably, this included session 1 (85.2% ± 3.5%) and 11 (81.3% ± 3%), when rats first learned or switched task contingencies, respectively. Thus, rats quickly adapted to the reset contingencies within single sessions.
Figure 2. Rats adjusted behavior under reset conditions of the three Diminishing Returns tasks.

Note. Error bars and solid lines display means ± standard errors. Data was not recorded on days 5–7 of the FD-10/PD-1 Diminishing Returns task due to a pellet dispenser malfunction. a Overall probability of responding on PD-1 or FD-10 levers. b Probability of responding by trial.
Transparent sections show means ± standard errors. c The overall number of consecutive presses that a typical rat will emit on the PD-1 and FD-10 levers. d-f Same as A-C but for the FD-20/PD-2 task. g-I Same as A-C but for the FD-20/PD-1 task.
The opposite results were obtained during no-reset sessions; rats were significantly more likely to respond on the FD lever (67.3% ± 2.6%) than the PD lever when pooling data over all sessions (Wald test; p < 0.0001; Figure 2A dark gray thick dashed vs. light gray thin dashed), and when examining individual sessions 3–20 (64.3% ± 5% to 80.1% ± 4.9%; sessions 3–9, ps < 0.0001; session 10, p < 0.001; sessions 11–20, ps < 0.0001; Supplemental Figure 2A). Only during sessions 1 (42.8% ± 6.2%) and 2 (54% ± 7.6%) did rats not significantly respond more on the FD lever during no-reset sessions. Thus, after 2 sessions, rats responded appropriately during no-reset sessions.
Next, we examined behavior across trials within a session. Here we focused on sessions 10 and 20, after rats had several days’ experience with both reset and no-reset contingencies. As expected from above, rats’ frequency of responding on the two levers differed during reset and no-reset sessions (reset condition (F(1, 22) = 19.53, p = 0.0002; trial (F(1, 22) = 22.32, p = 0.0001); and reset condition by trial interaction (F(1, 22) = 41.31, p < 0.0001)). Under reset conditions a rat was significantly more likely to respond on the PD lever than the FD lever over all trials (85.4% ± 1.5%; Wald test: p < 0.0001) and on each individual trial (83.2% ± 3.4% to 89.3% ± 4.3%; Wald tests: all ps < 0.0001; Figure 2B dark gray thick solid vs. light gray thin solid), further suggesting that rats appropriately tracked delay contingencies throughout the entire session
Under no-reset conditions rats should begin each session by selecting the PD lever more often and then switch to the FD lever once delays on the two levers are equal (equality point). Indeed, we found that on trials 1–6 (42.3% ± 5% to 45% ± 4.8%; Wald tests: trials 1–3 (p < 0.01); 4–6 (p < 0.05)) rats were more likely to choose the PD lever over the FD lever. However, it was not until trial 24 that rats showed significantly higher probabilities of selecting the FD lever (54.9% ± 4.5% to 96.6% ± 2%; trials 24–25 (ps < 0.05), 26–28 (ps < 0.01), 29–30 (ps < 0.001) through the remainder of trials (all ps < 0.0001); Figure 2B dark gray thick dashed vs. light gray thin dashed). This result suggests that rats responded appropriately early in sessions but took longer to commit to the FD after delays became equal.
The analyses above demonstrate that rats generally adapt behavior appropriately during both reset and no-reset sessions. We next asked how well behavior matched optimal decision-making. Charnov’s MVT states that an animal forages optimally if it changes patches when the marginal rate of reward in the patch equals the average rate of reward in the habitat (Charnov, 1976). To test if a typical rat under each reset condition had learned the average rate of reward in the conditioning chamber, we first determined the optimal probability for selecting the PD lever for reset conditions, as calculated based on the methods of Hackenberg & Axtell (1993). Under the PD-1/FD-10 schedule, optimal responding during reset sessions was to press the PD lever four times consecutively and then press the FD lever once to reset the PD schedule. Thus, under these conditions an optimally behaving rat should respond on the PD lever with an 80% probability and the FD lever with a 20% probability. Under no-reset conditions, the optimal probabilities for responding on each lever depend on the number of trials completed within that session (the optimal PD probability equals eleven divided by the total number of completed trials).
Across sessions, the probability for selecting the PD lever under reset conditions (86.4% ± 1.5%) was significantly greater than optimal when analyzed over all sessions (Wald test; p < 0.0001; Figure 2A dark gray thick solid bar vs. dark gray thick solid dashes). During sessions 1, 4, 11–14, and 20, however, responding (80.4% ± 3% to 86.4 ± 3%) was not significantly different from optimal (Supplemental Figure 2A). During the final reset session, the probabilities for selecting the PD lever were not significantly different from optimal for the first 66 trials (83.2% ± 3.1% to 85.4% ± 2.4%) and were higher than optimal for trials 67–160 (85.4% ± 2.4% to 89.3% ± 3.7%; ps < 0.05; Figure 2B dark gray thick solid line vs. dark gray thick solid dashes). Thus, during many sessions and trials, rats’ performance was not significantly different from optimal, however performance was suboptimal late in sessions due to over selection of the PD lever.
During no-reset sessions, performance was suboptimal over all sessions (67.3% ± 2.8% FD probability; Wald test; p < 0.0001; Figure 2A dark gray thick dashed bar vs. dark gray thick dashed dashes), and during each individual session (42.8% ± 6.2% to 80.1% ± 4.9%; all ps < 0.0001, Supplemental Figure 2A). During the final sessions (i.e., sessions 10 and 20), the observed probabilities were significantly different from optimal for the first 108 trials (Figure 2B dark gray thick dashed line vs. dark gray thick dashed dashes), 42.3% ± 4.9% to 88.6% ± 3.2%; trials 89–93 (ps < 0.001), 94–100 (ps < 0.01), and 101–107 (ps < 0.05), ps > 0.05 for the last 53 trials). Thus, overall rats performed suboptimal during no-reset sessions, and only achieved optimal performance (88.8% ± 3.2% to 96.6% ± 1.7%) late in the session.
The above analyses all suggest that during reset sessions rats adjusted behavior compared to no-reset sessions by selecting PD and FD levers close to 80% and 20%, respectively. However, these probabilities of responding might not explain the sequence of choices a rat made during a session. That is, 80/20 probabilities could result from 8 PD presses followed by 2 FD presses, not 4 PD presses followed by 1 FD press. To address this concern, we used a lognormal GLMM to statistically model the number of consecutive responses that a typical rat emitted on the PD lever before switching to the FD lever under both reset conditions across sessions (Figure 2C). Wald tests comparing the typical values to the optimal values reveal that under reset conditions the median number of responses (4.6 ± 0.4) on the PD lever before switching was not significantly different from optimal (p = 0.097). While most individual sessions were not significantly different from optimal (3.5 ± 0.5 to 5.9 ± 1.1; Supplemental Fig 2B; Supplemental Fig 3A), sessions 3 (6 ± 1.2, p < 0.05) and 15 (6 ± 1, p < 0.05) were. Across individual bouts of pressing on the PD lever within the final session, we observed no significant differences from optimal (4.6 ± 1 to 5.9 ± 5.3; Supplemental Figure 2C).
PD-2 and FD-20 schedule pair
Charnov’s MVT states that an optimally foraging animal will seek a new patch when the marginal rate of reward in the current patch equals the average rate of reward for that habitat (Charnov, 1976). These comparisons between the marginal rate of reward and average rate of reward reflect the ratio of PD to FD delays, not the actual values of the delays themselves, thus changing the reinforcement schedules associated with each lever, while leaving the ratio the same, should result in similar patterns of behavior. To test this hypothesis, during days 11–15 of testing we switched the concurrent schedule from PD-1/FD-10 to PD-2/FD-20.
As above, the optimal probability for selecting the PD lever under reset conditions is 80%, and the optimal probability for the FD lever is 20%. During reset sessions, the probability for selecting the PD lever was greater than the probability for selecting the FD lever analyzed over all sessions, and it was significantly greater than optimal values (85.8% ± 1.3%; Wald test; p < 0.001; Figure 2D; solid). Analyzing individual PD-2/FD-20 sessions during reset conditions revealed significant differences only in sessions 2 (86.4% ± 2.2%; p < 0.05), and 3–4 (88.7% ± 2.3% and 88% ± 2.1%; p < 0.01; Supplemental Figure 2D). However, analyzing the final PD-2/FD-20 session by trial revealed that during reset sessions the probabilities for selecting the PD and FD levers were not significantly different from optimal for any trial (78.6% ± 2.6% to 86.1% ± 7.1%; all ps > 0.05; Figure 2E; solid).
Under no-reset conditions, the optimal probabilities for selecting the PD or FD lever was significantly different from the observed probabilities over all sessions (70% ± 3.1% FD probability; Wald test; p < 0.0001; Figure 2D; dashed), and each individual session (67% ± 6.9% to 74.3% ± 3.3%; sessions 1 and 5, p < 0.0001; sessions 2–4, ps < 0.001; Supplemental Figure 2D). Analysis of individual trials found behavior was significantly different from optimal up to trial 80 (51.6% ± 4.5% to 85.9% ± 2.9%; trials 1–66 (ps < 0.0001), 67–70 (ps < 0.001), 71–74 (ps < 0.01), and 75–79 (ps < 0.05); Figure 2E; dashed).
Similar to the PD-1 and FD-10 schedule pair, optimal responding under reset conditions for the PD-2 and FD-20 schedule pair involves repeatedly emitting four consecutive responses on the PD lever followed by a single response on the FD lever. We used Wald tests to determine if the pattern of responding produced by a typical rat differs from the optimal pattern. Under reset conditions the median number of responses on the PD lever before switching was significantly different from optimal when analyzed across sessions (4.6 ± 0.3; p < .05; Figure 2F). On a session by session basis no significant differences occurred during sessions 1, 2, and 5 (3.7 ± 0.4 to 4.8 ± 0.8) but did occur during sessions 3 and 4 (5.7 ± 1 and 5.5 ± 0.7; ps < 0.05; Supplemental Figure 2E; Supplemental Figure 3B). During the final session, however, we observed no bouts of pressing on the PD lever that differed significantly from optimal (3.3 ± .4 to 5.2 ± 4.3; Supplemental Figure 2F). Thus, as above, during many sessions and trials rats adjusted behavior appropriately during reset conditions, but tended to overly choose the PD lever late in sessions performing in a suboptimal manner.
PD-1 and FD-20 schedule pair
Lastly, we had rats perform a third concurrent PD-1/FD-20 schedule to determine if rats could adapt response patterns when the ratio of delays unexpectedly changed. Optimal behavior under reset conditions entailed pressing the FD lever once after six consecutive presses on the PD lever, thus optimal probabilities for responding on the PD and FD levers were 85.7% and 14.3%, respectively. Over all reset sessions the probability for selecting the PD lever was significantly greater than the probability for selecting the FD lever and higher than optimal values (89% ± 0.9%; Wald test; p < 0.001; Figure 2G; solid; Supplemental Figure 2G). Analyzing the final PD-1/FD-20 session by trial revealed that during reset sessions the probabilities for selecting the PD and FD levers were not significantly different from optimal for any trial (87.5% ± 1.2% to 83.5% ± 4.7%; all ps > 0.05; Figure 2H; solid).
Under no-reset conditions, the optimal probability for selecting the FD lever was significantly different from optimal over all sessions (64% ± 1.6%; Wald test; p < 0.0001; Figure 2G; dashed), in each individual session (69.3% ± 4.2% to 56.3% ± 3%; session 3, p < 0.01; all other sessions ps < 0.0001; Supplemental Figure 2G), and was significantly different from optimal during the first 98 trials (30.4% ± 3.7% to 83.6% ± 2.8%; trials 1–87 (ps < 0.0001), 88–89 (ps < 0.001), 90–93 (ps < 0.01), 94–97 (ps < 0.05), Figure 1H; dashed).
Examining response patterns, under reset conditions the median number of responses on the PD lever before switching was not significantly different from optimal when analyzed over all sessions (6.4 ± 0.4; p = 0.333; Figure 2I). Most individual sessions did not result in a significant difference from optimal performance (5.4 ± 0.5 to 7.2 ± 1), except for session 3 (7.2 ± 0.6; p < 0.05; Supplemental Figure 2H and Supplemental Figure 3C). During the final PD-1/FD-20 session, bouts of pressing on the PD lever did not differ significantly from optimal (6.3 ± 0.8 to 4.5 ± 0.9; Supplemental Figure 2I). Thus, as under the two other reinforcement schedules, during the large majority of sessions and trials, rats adjusted behavior appropriately during reset conditions, switching before the equality point. However, suboptimal performance was observed in the form of selecting the PD left too often.
Discussion
Previously, researchers have observed near-optimal responding in pigeons, monkeys, and humans during various diminishing returns tasks (Hackenberg & Axtell, 1993; Hackenberg & Hineline, 1992; Hodos & Trumbule, 1967; Jacobs & Hackenberg, 1996). Yet, the neural basis of this behavior is still unknown. The goal of this study was to determine if rats could also delay-gratification to perform optimally in this task. By establishing this behavior in the rat model future work can examine neural mechanisms underlying the ability of rats to exhibit self-control by switching earlier in the PD sequence during reset compared to no-reset sessions despite the short-term costs, and investigate how this behavior and associated neural mechanisms are disrupted in models of psychiatric illness and aging. While delay-discounting and gambling tasks effectively demonstrate how changing value can shift an animal’s pursuit of reward, they do not address the computations behind selecting an immediately unfavorable option in order to benefit an animal in the long-term over many trials (Bickel et al., 2007; Ferland et al., 2018; Hayden, 2016). Thus, by adapting the time-based version of the Diminishing Returns task to rats, we have now established the rat as an animal model that can be utilized to study decision-making in the context of time-based diminishing returns.
Performance during diminishing returns tasks are thought to reflect innate, evolutionarily conserved decision-making processes that maximizes energy intake, as hypothesized by optimal foraging theory (Pyke, Pulliam, & Charnov, 1977), however it might be argued that rats simply followed a particular sequence of responses that produced the highest rate of reward. Even if this was true, we would argue that rats must have been keeping track of the relative rate of reward, at least initially, to learn the most beneficial response sequence. Further, given that rats performed so well immediately after changes in lever-reward contingencies, our results suggest that they were not simply repeating a learned sequence of habitual actions but were adjusting to the context of their current environment.
Interestingly, although behavior did not significantly differ from optimal across many analyses, sessions, schedules and trials, suboptimal behavior was often observed later within sessions, at least during FD-10/ PD-1 and FD-20/ PD-2 schedules. It is unclear why this might be; however, we speculate that this reflects satiation or perhaps lost interest, self-control, arousal or attention as trial blocks progress. Regardless of what the explanation might be, we anticipate that neural correlates related to reward value and response selection will mirror within-session changes in choice selection.
Unlike reset sessions, behavior was suboptimal during no-reset sessions early within sessions (i.e., it took them more trials to commit to the FD lever). It is unclear as to why a typical rat under no-reset conditions did not behave more optimally, especially because it seems like the more simplistic task. One explanation might be that rats were treating the blocks of trials as some sort of mid-session reversal during which they used time as a discriminative cue to alter behavior. Use of time would not be as precise as monitoring reward delays, which may lead to poorer performance, however it would also result in a steady increase in FD choices over trials within a session as we observed. Another possible explanation is that choosing between two, and only two, foraging patches with relatively frequent and quickly obtained rewards was not a common problem that rats faced in evolutionary history. Speculation aside, it is important to point out that rats did perform fairly well during no-reset sessions, choosing the FD lever far more often than the PD lever, especially later in each session during which they approached near 100% responding on the FD lever. Not fully committing to one FD lever early in each session might reflect some sort of exploratory behavior that is innate or that is potentially fostered by numerous contingency changes across training and testing.
For the analysis and interpretations that we describe in this article we have assumed that there is no noise in the decision-process. However, there must be some error in the judgements made by the rats, which could explain suboptimal behavior. One approach to estimate error argues that value estimates in the brain are expected to follow Weber’s law (Gallistel, 2017). If one assumes a 10–20% Weber fraction, then the “optimal decision” will be expected to deviate from the proposed solution by up to the number of presses in which the above long-term reward rate deviates by up to 10–20%. Another approach argues that there will be some error in a rate calculation due to Weber’s law in time perception, and that this error will also be approximately scalar (Namboodiri, Mihalas, & Shuler, 2014). In this framework, the Weber fraction might depend on the intervals the rats generally experience in the task. Notably, in either formalism, the behavior of rats is well within the error rate in the reset condition. However, in the non-reset condition, it seems that 20% of error in the reward rate calculation occurs when PD is pressed up to 22 or so times, which is close to the observed number in the non-reset condition. Further, it may also well be that in the non-reset conditions, which result in much longer intervals experienced by the animals, the error in time perception might be supra-scalar, potentially due to the fact that the time resolution needs to be optimized for the shorter durations. Jazayeri and Shadlen (2010), also argue for more biases in the decision in the non-reset condition. Thus, considering variability in performance across the conditions and noise in the decision process, rats perform the task extremely well.
For us, as behavioral neuroscientists interested in animal models of psychiatric disorders, the most interesting aspect of this task is that the reset manipulation sets up a scenario where the short-term consequences favor switching at the equality point (e.g., after 11 selections of the PD option), but long-term consequences favor switching much earlier (e.g., 4 selections of the PD option). Remarkably, despite the immediate cost of selecting the fixed option (e.g., wait 10s instead of 4s because it resets the PD back to 0), multiple species have been shown to switch prior to the equality point, at near optimal levels (Hackenberg & Axtell, 1993; Hackenberg & Hineline, 1992; Hodos & Trumbule, 1967; Jacobs & Hackenberg, 1996). MVT suggests that animals accomplish successful foraging during this task by tracking the average rate of reward, a function that is common across many tasks that probe behavioral vigor, risk sensitivity, labor–leisure trade-offs, self-control, and delay discounting (Cools et al., 2011; Constantino & Daw, 2015; Daw & Touretzky, 2002; Gallistel & Gibbon, 2000; Guitart-Masip et al., 2011; Kacelnik, 1997; Keramati, Dezfouli, & Piray, 2011; Kurzban, Duckworth, Kable, & Myers, 2012; Niv et al., 2006; Niv et al., 2007; Niyogi et al., 2014). Since decision-making of this nature shares the same dependence on tracking of average reward rate it is possible that they engage common neural circuits as the time-based diminishing returns tasks described here.
Conclusion
Here, we successfully establish a time-based Diminishing Returns task in the rat model. Establishing this task in rats opens the door to study the underlying neural processes of self-control, delayed-gratification and decision-making in an experimentally tractable animal. Further, testing rat models of psychiatric disorders, aging, and addiction in this task may yield valuable insights into their etiologies, symptomatologies, and potential treatment strategies (Lamichhane et al., 2020).
Supplementary Material
Acknowledgments
This work was supported by the National Institute on Drug Abuse [DA031695, 2018-2022]. The funding source had no involvement with the study design, research, or preparation of the manuscript.
We would like to thank Dr. Todd Braver and Dr. William Hodos for their thoughtful input on experimental design and preparation of the manuscript.
Footnotes
The authors declare no biomedical financial interests or potential conflicts of interest.
References
- Addicott MA, Pearson JM, Kaiser N, Platt ML, & McClernon FJ (2015). Suboptimal foraging behavior: A new perspective on gambling. Behavioral Neuroscience, 129, 656–665. doi: 10.1037/bne0000082 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bari A, & Robbins TW (2013). Inhibition and impulsivity: Behavioral and neural basis of response control. Progress in Neurobiology, 108, 44–79. doi: 10.1016/j.pneurobio.2013.06.005 [DOI] [PubMed] [Google Scholar]
- Bickel WK, Miller ML, Yi R, Kowal BP, Lindquist DM, & Pitcock JA (2007). Behavioral and neuroeconomics of drug addiction: Competing neural systems and temporal discounting processes. Drug and Alcohol Dependence, 90, S85–S91. doi: 10.1016/j.drugalcdep.2006.09.016 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bijma F, Jonker M, & Van der Vaart A (2017). An introduction to mathematical statistics. Amsterdam University Press. [Google Scholar]
- Brockett AT, Pribut HJ, Vázquez D, & Roesch MR (2018). The impact of drugs of abuse on executive function: Characterizing long-term changes in neural correlates following chronic drug exposure and withdrawal in rats. Learning and Memory, 25, 461–473. doi: 10.1101/lm.047001.117 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Charnov EL (1976). Optimal foraging: The marginal value theorem. Theoretical Population Biology, 9, 129–136. 10.1016/0040-5809(76)90040-X [DOI] [PubMed] [Google Scholar]
- Constantino SM, & Daw ND (2015). Learning the opportunity cost of time in a patch-foraging task. Cognitive, Affective, & Behavioral Neuroscience, 15, 837–853. doi: 10.3758/s13415-015-0350-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cools R, Nakamura K, & Daw ND (2011). Serotonin and dopamine: Unifying affective, activational, and decision functions. Neuropsychopharmacology, 36, 98–113. doi: 10.1038/npp.2010.121 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Daw ND, & Touretzky DS (2002). Long-term reward prediction in TD models of the dopamine system. Neural Computation, 14, 2567–2583. https://psycnet.apa.org/doi/10.1162/089976602760407973 [DOI] [PubMed] [Google Scholar]
- Ferland JN, Adams WK, Murch S, Wei L, Clark L, & Winstanley CA (2018). Investigating the influence of ‘losses disguised as wins’ on decision making and motivation in rats. Behavioral Pharmacology, 29, 732–744. doi: 10.1097/FBP.0000000000000455 [DOI] [PubMed] [Google Scholar]
- Forbes C, Evans M, Hastings N, & Peacock B (2011). Lognormal distribution. In Statistical distributions (pp. 131–134). Hoboken, NJ: John Wiley & Sons. [Google Scholar]
- Gallistel CR (2018). Finding numbers in the brain. Philosophical Transactions of the Royal Society B. 373: 20170119. 10.1098/rstb.2017.0119 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gallistel CR, & Gibbon J (2000). Time, rate, and conditioning. Psychological Review, 107, 289–344. doi: 10.1037/0033-295x.107.2.289 [DOI] [PubMed] [Google Scholar]
- Guitart-Masip M, Beierholm UR, Dolan R, Duzel E, & Dayan P (2011). Vigor in the face of fluctuating rates of reward: An experimental examination. Journal of Cognitive Neuroscience, 23, 3933–3938. https://psycnet.apa.org/doi/10.1162/jocn_a_00090 [DOI] [PubMed] [Google Scholar]
- Hackenberg TD, & Axtell SA (1993). Human’s choices in situations of time-based diminishing returns. Journal of the Experimental Analysis of Behavior, 59, 445–470. 10.1901/jeab.1993.59-445 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hackenberg TD, & Hineline PN (1992). Choice in situations of time-based diminishing returns: Immediate versus delayed consequences of action. Journal of the Experimental Analysis of Behavior, 57, 67–80. doi: 10.1901/jeab.1992.57-67 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hayden BY (2016). Time discounting and time preference in animals: A critical review. Psychonomic Bulletin and Review, 23, 39–53. 10.3758/s13423-015-0879-3 [DOI] [PubMed] [Google Scholar]
- High R (2019) Zero-inflated and zero-truncated count data models with the NLMIXED procedure. Paper presented at the SAS Global Forum, Washington D.C. Paper 3768–2019. [Google Scholar]
- Hodos W, & Trumbule GH (1967). Strategies of schedule preference in chimpanzees. Journal of the Experimental Analysis of Behavior, 10, 503–514. https://psycnet.apa.org/doi/10.1901/jeab.1967.10-503 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jacobs EA, & Hackenberg TD (1996). Humans’ choices in situations of time-based diminishing returns: Effects of fixed-interval duration and progressive-interval step size. Journal of the Experimental Analysis of Behavior, 65, 5–19. doi: https://dx.doi.org/10.1901%2Fjeab.1996.65-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jazayeri M & Shadlen MN (2010). Temporal context calibrates interval timing. Nature Neuroscience, 13, 1020–1026. 10.1038/nn.2590 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kacelnik A (1997). Normative and descriptive models of decision-making: Time discounting and risk sensitivity. In Bock GR & Cardew G (), Characterizing human psychological adaptations (pp. 51–70). West Sussex, England: John Wiley & Sons Ltd. [DOI] [PubMed] [Google Scholar]
- Keramati M, Dezfouli A, & Piray P (2011). Speed/accuracy trade-off between the habitual and the goal-directed processes. PLoS Computational Biology, 7, e1002055. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kurzban R, Duckworth A, Kable JW, & Myers J (2012). An opportunity cost model of subjective effort and task performance. Behavioral and Brain Sciences, 36, 697–698. doi: 10.1017/S0140525X12003196 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lamichhane B, Di Rosa E, Green L, Myerson J, & Braver TS (2020). Examining delay gratification in healthy aging. Behavioural Processes, 176, 104125. 10.1016/j.beproc.2020.104125 [DOI] [PMC free article] [PubMed] [Google Scholar]
- McGinley MA (1984). Central place foraging for nonfood items: Determination of the stick size-value relationship of house building materials collected by eastern woodrats. The American Naturalist, 123, 841–853. doi: 10.1086/284243 [DOI] [Google Scholar]
- Niv Y, Daw N, & Dayan P (2006). How fast to work: Response vigor, motivation and tonic dopamine. In Weiss Y, Scholkopf B, & Platt J (), Advances in neural information processing systems (Vol. 18, pp. 1019–1026). Cambridge, MA: MIT Press. [Google Scholar]
- Niv Y, Daw ND, Joel D, & Dayan P (2007). Tonic dopamine: Opportunity costs and the control of response vigor. Psychopharmacology, 191, 507–520. doi: 10.1007/s00213-006-0502-4 [DOI] [PubMed] [Google Scholar]
- Niyogi RK, Breton YA, Solomon RB, Conover K, Shizgal P, & Dayan P (2014). Optimal indolence: A normative microscopic approach to work and leisure. Interface, 11, 20130969. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pleasants J (1989). Optimal Foraging by Nectarivores: A Test of the Marginal-Value Theorem. The American Naturalist, 134, 51–71. https://www-jstor-org.proxy-um.researchport.umd.edu/stable/2462275 [Google Scholar]
- Pyke G, Pulliam H, & Charnov E (1977). Optimal foraging: A selective review of theory and tests. The Quarterly Review of Biology, 52, 137–154. https://www-jstor-org.proxy-um.researchport.umd.edu/stable/2824020 [Google Scholar]
- Stroup WW (2012). Generalized linear mixed models: Modern concepts, methods and applications. Boca Raton, FL: CRC Press. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
