Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 May 15.
Published in final edited form as: Neurobiol Learn Mem. 2025 May 15;220:108063. doi: 10.1016/j.nlm.2025.108063

Female rats retain goal-directed planning of action sequences after acute stress despite changes in planning structure and action sequence execution

Russell Dougherty a,1, Eric A Thrailkill a,b,c, Sarah Van Horn a, Auny Kussad a, Donna J Toufexis a
PMCID: PMC12302003  NIHMSID: NIHMS2084270  PMID: 40381721

Abstract

When making decisions under stress, organisms tend to deliberate less and rely on automatic habits. Prior investigation into the influence of stress on decision-making has primarily viewed goal-direction and habit as independent and competitive sources of control in static environments. The effects of acute stress on the integration of goal-direction and habit in hierarchical planning to solve dynamic tasks remain unclear. Here, our aim was to assess whether stress prompted the usage of habitual action sequences over the selection of discrete goal-directed actions in a serial decision task. We trained 16 female Long Evans rats in a two-stage binary choice task and performed two probe tests, one following acute restraint stress and one under control conditions, to identify how stress affected higher-level planning of behavior and intermediate action structures. We found that under both stressed and control conditions, rats exhibited goal-directed planning of habitual action sequences. However, following stress, rats showed a greater tendency to reiterate action sequences independent of reinforcement, indicating that stress may induce an aversion to exploration in action planning. Stress also increased the latency between responses – degrading action sequence integrity despite conserving their overall structure and performance. Taken together, these findings suggest that although acute stress does not disrupt the overall macrostructure of behavior in two-stage decision-making, it does alter the microstructure of goal-directed and habitual control individually. Further, these results imply that the extent to which stress impairs goal-direction in female rats may depend on the incentive structure and attentional demands of the decision environment.

Keywords: action sequences, decision-making, habit, goal-directed action, planning, reinforcement learning, stress

1. Introduction

In predictable environments organisms often use habitual forms of behavior as a means to conserve cognitive resources (Bouton, 2021; Linnebank et al., 2018; Wood et al., 2002). When stressed, organisms also sacrifice costly forms of decision-making for habits to liberate cognitive resources for escaping perceived threats (Hermans et al., 2014). The dichotomy between habitual and goal-directed modes of decision-making is formalized in dual process theories of instrumental behavior, which impose a tension between a habitual system driven by stimulus-response associations, and a goal-directed system governed by action-outcome associations (Adams & Dickinson, 1981; Dickinson & Balleine, 1994). Stress induces a bias towards habitual control in instrumental conditioning experiments (Dias-Ferreira et al., 2009; Dougherty et al., 2024; Schwabe & Wolf, 2009, 2010). However, emerging research has increasingly advanced the idea that goal-directed and habitual systems operate hierarchically and may both be involved in goal selection, planning, and response performance (Ballard et al., 2024; Balleine & Dezfouli, 2019; Du et al., 2022; Favila et al., 2024; Ferguson et al., 2024; Frölich et al., 2023; Morris & Cushman, 2019). These theories suggest that rather than directly competing for control over behavior, responses may be organized through collaborative integration of goal-directed and habitual processes.

Hierarchical theories suggest that habits are the result of sequentially performed actions becoming “chunked” into a functional unit with repeated performance (Balleine & Dezfouli, 2019; Dezfouli et al., 2014; Smith & Graybiel, 2013, 2016). Action chunking has been identified in various settings including motor skill learning, free-operant procedures, and spatial navigation (Dezfouli & Balleine, 2013, 2019; Halbout et al., 2019; Keele, 1968; Nissen & Bullemer, 1987; Pew, 1966; Smith & Graybiel, 2013; van Elzelingen et al., 2022). By unitizing actions that are executed together into a sequence, a course of action can be evaluated based on the expected value of the sequence as a whole (e.g., ∑R1 + R2 + R3) (Dezfouli & Balleine, 2012). Once initiated, the action sequence is performed to completion and hence becomes insensitive to changes within the sequence. Functionally, this reduces time spent in deliberation over the outcome of each action within the sequence. Such a process captures characteristics often attributed to habits, such as reduced latency and more efficient performance (Garr & Delamater, 2019). A hierarchical controller selects between performing an action sequence versus using goal-directed evaluation of individual actions. Once selected, each action in a chunked action sequence serves as the eliciting stimulus for its successor. Thus, whether actions occur as discrete (goal-directed) responses or as a (habitual) chunked sequence can be assessed by introducing test trials in which accurate performance requires re-evaluation of the response after initiating the sequence (Daw et al., 2011; Dezfouli & Balleine, 2013, 2019).

Hierarchical theories of habitual behavior suggest that chunked action sequences may underly habitual behaviors and provide a parsimonious account of how habits can be integrated within overarching goal-directed control (Balleine & Dezfouli, 2019; Dezfouli & Balleine, 2012, 2013). However, little is known about how stress may affect hierarchical control over goal-direction and habit in instrumental behavior. Here, we extend this theoretical approach to investigate how acute stress influences the tendency toward goal-directed and habitual action sequences.

Stress can influence several types of computations involved in serial decision-making (Cremer et al., 2021; Otto, Gershman, et al., 2013; Otto, Raio, et al., 2013; Radenbach et al., 2015; Raio et al., 2020), but no studies have tested whether stress promotes action chunking. The goal of this experiment was to test the influence of stress on whether rats perform habitual action sequences during serial decision-making. To accomplish this, rats received a within-subject evaluation of acute (60 min) restraint stress on performance in a two-stage decision-making task (Dezfouli & Balleine, 2019). The procedure is diagrammed in Figure 1. Stage 1 consists of a choice between two response alternatives (R1 and R2) that initiate a sequence (R1 ➜ R2 or R2 ➜ R1). Rats learn to choose in stage 1 based on the outcome of recent trials (rewarded or nonrewarded) and select the stage 1 action likely to earn reward. In stage 2, rats must select the appropriate subsequent action that earns the reinforcer (i.e., R1 ➜ R2 or R2 ➜ R1). Rats can learn to make a correct stage 2 response according to either its regularity of reinforcement following the stage 1 choice, or according to its association with reward based on the stage 2 discriminative stimulus (Figure 1A). This task can include test trials (Probe trials) to evaluate whether action sequences are being performed using an action chunking strategy or by evaluating each individual action. On these trials, the discriminative stimulus for the stage 2 action is switched in order to assess whether the rat completes the action sequence initiated in stage 1 or uses the discriminative stimulus to select the appropriate stage 2 action to earn a reward. Thus, the probe trials can discern between habitual action sequences where stage 1 and 2 actions are executed as a chunked unit without separate evaluation of consequences, or as discrete successive choices where outcomes or discriminative stimuli are used to separately guide action selection in each stage. We hypothesized that stress would increase the hierarchical goal-directed selection of habitual action sequences, and that stress would interfere with rats’ ability to flexibly select discrete actions after initiating a sequence.

Figure 1.

Figure 1.

Two-stage task. Note. A) Rats received intermixed discrimination training where they were expected to learn the association between the light stimulus (S1 or S2), its associated lever press (R1 or R2), and reward (+). B and C) Rats were then trained in a two-stage decision task with two trial types. In each trial type only one stage 2 S:R pair was “reward-possible”. The “reward-possible” S:R pair changed probabilistically during the task – requiring the rat to adjust its stage 1 choice through experience with recent trials. D) Example showing adjustment of response trajectory to reach the new “reward-possible” S:R pair following a reversal in trial type. RST = Stay, RSW = Switch.

2. Materials and Methods

2.1. Subjects

16 Female Long Evans rats (Charles River, Quebec; 75–90 days old at the time of arrival) were used in this experiment. Sample size was based on other rodent studies with similar tasks (Dezfouli & Balleine, 2019; Miller et al., 2017, 2022). Rats were pair-housed with unlimited access to water in a climate-controlled colony room maintained at 23°C with a 12-hour light/dark cycle (7:00 – 19:00 light-on period) and were handled each day. Experimental sessions were conducted daily at approximately 08:30. Rats were maintained at 85% of their free-feeding weights and received supplemental feeding approximately 4 hr after each session. All procedures were approved by the Institutional Animal Care and Use Committee at the University of Vermont.

2.2. Apparatus

Training procedures took place in 8 standard operant chambers (model ENV-008-VP; Med Associates, St. Albans, VT) housed within individual sound-attenuating cabinets. In the center of the right-facing chamber wall was a food cup (5 cm × 5 cm) into which a pellet dispenser (ENV-203–45) delivered 45-mg sucrose pellets (Bio-Serv). Entries to the food cup were recorded by an infrared beam. Two retractable levers (ENV-112CM) were positioned to the left and right of the food cup. One panel light was installed on the left-facing chamber wall, centered, and directly below the chamber ceiling. A sonalert module was mounted in the center of the right-facing chamber wall directly below the ceiling and delivered an auditory tone when activated. Each chamber was illuminated by a single 150 lumen red LED bulb mounted to the roof of the sound-attenuating cabinet. Ventilation fans provided background noise. Experimental events were controlled, and data was recorded by a computer located in an adjacent room.

The apparatus for acute restraint stress was a clear plexiglass cylindrical restraining device 9 cm × 15 cm (D × H; Braintree Scientific Inc., Braintree, MA) placed on the center of a brightly lit table located in a different room of the laboratory.

2.3. Procedure

The experiment timeline is diagrammed in Figure 2. Prior to the experiment, rats were assigned to an operant chamber in which they received all training and testing procedures. All sessions began 2 min after rats were loaded into their chambers. For all phases other than magazine training, sessions began with the insertion of the lever(s) into the chamber, and sessions ended with lever retraction.

Figure 2.

Figure 2.

Experimental timeline. Note. Sequence of events during the experiment.

2.3.1. Magazine training

On the first two days, all rats received sessions in which one sucrose pellet was delivered into the food cup every 60 s, on average (variable-time 60 s schedule; VT 60 s). Sessions terminated after the delivery of 30 pellets. Pellet delivery was signaled by the sound of the pellet striking the food cup; no other stimuli were presented. Lever manipulanda were not extended during these sessions.

2.3.2. Discrimination training

On the next day, all rats received two single lever training sessions in which the left lever or the right lever was available and each lever press delivered the reinforcer (fixed ratio 1; FR 1). Each session transitioned from free-operant to discrete-trial operant training following 20 reinforced lever presses. Discrete trial training consisted of presentations of either a steady (S1) or blinking (0.2 s-on, 0.2-s off; 5 Hz; S2) panel light stimulus separated by a variable intertrial interval (ITI) that averaged 35-s. Trials ended either with a pellet delivered following a lever press (FR 1) or without pellet delivery after 60 s. One session was administered for each stimulus-lever pair (e.g., steady light and right lever; S1:R2+) and sessions were separated by 60 min. Session order and stimulus-lever combinations were counterbalanced across rats. Sessions ended after 30 reinforcers were earned in the discrete-trial phase (50 total) or 60 min elapsed, whichever occurred first.

The next day, all rats received two more sessions with each lever. Sessions were identical to Day 1 except without the free operant phase; instead, sessions began with discrete-trial operant training. These sessions ended after 30 reinforcers or after 60 min elapsed.

On the following three days, rats received one session per day consisting of intermixed S1 and S2 trials with both levers extended into the chamber (S1:[R2+, R1-] and S2:[R1+, R2-]) in which they were expected to learn the association between the light stimulus, its associated lever press, and reward (see Figure 1A). Trials were separated by a variable 35-s ITI and could end with either a correct lever press (R2 in S1 or R1 in S2) or after 10 s without a correct lever press. Incorrect lever presses did not terminate the trial, meaning that rats had the opportunity to attempt the opposite lever in the event of an initial incorrect press. Sessions ended after 30 reinforcers were earned (15 in each S) or 60 min elapsed.

2.3.3. Two-stage task training

Next, rats received training on the two-stage task diagrammed in Figure 1B-D. Stage 1 was signaled by the tone (S0). In S0, rats could respond on one of two available levers, R1 (e.g., left lever press) or R2 (e.g., right lever press) to turn off S0 (FR 1) and present one of two stage 2 stimuli, S1 or S2 (steady or blinking panel lights). In stage 2 pressing the lever opposite to that pressed in S0 (R1➜ S1:R2+; R2➜ S2:R1+) would earn a reinforcer (FR 1). Prior to each S0 onset, the program selected only one of the sequences to be reinforced. During the session, there was a probability that the computer would reverse the reinforced sequence from the current sequence (e.g., R1➜R2+, R2➜R1-) to the opposite sequence (R2➜R1+, R1➜R2-). Thus, reinforcement in stage 2 required the ability to switch stage 1 actions based on the outcome of the preceding trial. A press to either lever ended stage 2 (turned S1 or S2 off) and initiated the next trial by turning on S0. Rats received one session per day, terminating after 60 reinforced trials or 60 min, whichever occurred first. Mapping of stage 1 action transitions (R1➜R2; R2➜R1) to stage 2 discriminative stimuli (S1, S2) was deterministic and counterbalanced across rats.

As illustrated in Figure 1D, when the probability of reversal is low, the reward-possible state (S1 or S2) is likely to remain the same from trial to trial. Optimal behavior entails repeating the same stage 1 action on the next trial to transition to the previously rewarded state (e.g., S0:R1➜S1:R2+). However, if no reward was earned in the previous trial (e.g., S0:R1➜S1:R2-), the optimal choice is to switch stage 1 actions and transition to the other stage 2 state (e.g., S2), as it was now likely that this state is reward-possible (e.g., S0:R2➜S2:R1+). Therefore, to maximize reinforcement in this task, responses must remain sensitive to these “reversals” and adjust by either “staying” on the same stage 1 action or “switching” to the other based on the most recent trial outcome (reinforced/nonreinforced) in a goal-directed manner.

Rats received three blocks of daily training sessions across which the probability of reversal was successively adjusted. Acquisition was measured in terms of the log odds of repeating the S0 action following reward (stay; log odds > 0 indicates greater likelihood of repeating the S0 action). Additional measures of acquisition included reaction time (RT) to assess automaticity in sequential actions within the trial, discrimination accuracy, trial attempt rates, and percent of rewarded responses to assess overall task performance across training. In Block 1 (sessions 1–7), the ITI was 1 s and the probability of a reversal was 0.50. In Block 2 (sessions 8–11), the ITI was increased to 3 s and the probability of reversal was 0.50. Finally, in Block 3 (sessions 12–29), the ITI was 0 s and the probability of reversal was 0.14 following four reinforced trials. This final stipulation meant that rats necessarily experienced each lever press sequence because a reversal could not occur until they had successfully completed the reinforced sequence four times. Block 3 training continued until the mean log odds of repeating the S0 action after earning reward was greater than zero for at least three sessions.

2.3.4. Probe tests and acute stress

Probe tests were conducted to assess whether rats were using chunked action sequences to approximate optimal stay/switch behavior. As in the final block of training, the reinforced sequence reversed probabilistically across the test session. However, to discern chunked action sequence versus discrete choice strategies, probe trials could occur with a probability of .20. On a probe trial, diagrammed in Figure 3A-B, the response in S0 (e.g., R1) would lead to the other stage 2 state (e.g., S2 instead of S1). Thus, the usual or “common” transitions were the same as during two-stage training (S0:R1➜S1:R2 and S0:R2➜S2:R1) and occurred with a probability of 0.80. Probe trials were contrastingly “rare” transitions and occurred with a probability of 0.20. Since reinforcement could occur only if the stage 2 response corresponded to the discriminative stimulus, the action sequence (e.g., R1➜R2, or R2➜R1) would be nonreinforced on rare transition trials (diagrammed in Figure 3C). To further illustrate, common transition trials (e.g., S0:R1➜S1:R2+) allowed actions to be concatenated into “S0:R1➜R2+”. This action sequence could therefore be selected based on the expected value of the sequence – bypassing the evaluation of constituent within-sequence actions signaled by S1 or S2. In contrast, on rare transition trials (e.g., S0:R1➜S2:R1+), reinforcement is earned by taking the out-of-sequence action in response to the unexpected S2 discriminative stimulus. Probe tests ended after 60 reinforced trials or 60 min, whichever occurred first.

Figure 3.

Figure 3.

Probe test. Note. A and B) Task structure during the probe tests. Transitions from stage 1 to stage 2 states were shifted from deterministic to probabilistic, where “Common” transitions remained the same as experienced during two-stage task training, and reinforcement on “Rare” transitions required rats to adjust their stage 2 response according to the S (S1 or S2) of that state. This changeallowed discernment of discrete choice and habitual sequence strategies shown in C). Habitual sequences (R1→R2, or R2→R1) would be nonreinforced on “Rare” transition trials.

All rats received two probe tests, one following 60 min of restraint stress (Stress) and the other under non-stressed (Control) conditions. Acute restraint stress induces robust physiological and behavioral stress responses in both male and female rats (Aloisi et al., 1998; Reis et al., 2011). The restraint stress took place immediately before the test session in an adjacent brightly lit laboratory room to which the rats had no prior exposure. For the control treatment, rats were handled for a similar duration to that of the stress group during transportation to the restraint apparatus (3–5 min) but left in their home cages in the colony room. A recent study from our laboratory found no disruption of estrous cycling when using the same restraint method (Dougherty et al., 2024); therefore, we omitted estrous cycle tracking here. All rats received stress and control treatments prior to a probe test session in counterbalanced order, with cage mates always receiving the same treatment. Stress and control probe tests were separated by three sessions of two-stage retraining (same methods as prior two-stage training), one session per day, to wash out effects of exposure to the first test session.

2.4. Experimental design and statistical analyses

2.4.1. Discrimination training

Discrimination accuracies (measured as reinforcers earned divided by total trials) during the intermixed discrimination training sessions were analyzed using a repeated measures analysis of variance (ANOVA) with session entered as the within-subjects factor.

2.4.2. Two-stage task training

To assess change in trial-by-trial behavior across training sessions, S0 and stage 2 responses were analyzed using a generalized linear mixed model (GLMM) analysis with logistic regression for each training block. GLMM models account for hierarchically structured data (i.e., correlated trial-by-trial responses of each rat) by treating trial responses and outcomes as observations nested within animals. In this framework, the relationship between factors such as reinforcement on the preceding trial and current trial behavior can be estimated at the population level while accounting for random variability between rats within the sample. This approach reduces the risk of a Type I error due to interdependence in within-subject data, increases statistical power, and allows for a more nuanced assessment of behavior than possible using other methods such as repeated measures ANOVA for session-level summary statistics (Sommet & Morselli, 2017; Yu et al., 2022). In all GLMM analyses described, intercept was included as a random effect varying across rats to allow the probability of the outcome variable (e.g., log odds of staying on same stage 1 response) to randomly vary between animals. Thus, the model accounts for individual variability when estimating the relationship between the predictors and the outcome variable – providing more robust inferences of their overall effects. Using this framework, responses in S0 and stage 2 (stay [1] or switch [0]) from each training block were separately regressed on reward earned on the previous trial (reward [1] or no reward [0]), session, and a reward by session interaction term. Task acquisition was assessed using the reward by session interaction term, with a significant positive interaction indicating increasing tendency to repeat reinforced actions from preceding trials and switch actions following omission of reward as sessions progressed. Stage 1 and 2 RTs were regressed on session to assess changes across training. Two-stage task retraining data was analyzed in the same manner except “session” was replaced with a binary variable “retraining” (training [0] or retraining [1]) to assess changes in responses and RT compared to the final three sessions of training.

Changes in session-level performance statistics including discrimination accuracy, trials attempted per minute, and percent of rewarded trials were also assessed as additional indicators of task acquisition and were each analyzed using a repeated measures ANOVA in the same manner as for discrimination training (described above). Sphericity of the data was assessed using Mauchly’s test (p < 0.05), and Greenhouse-Geisser (ε < 0.75) or Huynh-Feldt (ε > 0.75) and corrections to degrees of freedom were applied when sphericity was violated. Session-level performance statistics for two-stage task retraining were compared to the final three training sessions with paired samples t-tests.

2.4.3. Probe tests

Trial reward rates (percent of rewarded trials) and trial attempt rates were compared across stress conditions in paired samples t-tests. Responses in S0 were regressed on reward earned on the previous trial, stress (stressed [1] or non-stressed [0]), and a reward by stress interaction term. GLMM with logistic regression was also used for the analysis of stage 2 responses during the probe tests. Predictors included S0 response, reward earned on the previous trial, stress, and interaction terms for S0 response by reward and three-way S0 response by reward by stress, and the outcome was whether the stage 2 response matched the previous trial (stay or switch). Only those trials in which the stage 2 state differed from that of the previous trial were entered into this analysis. This allowed us to evaluate the likelihood of repeating stage 2 responses after earning reward on the preceding trial and repeating the same S0 action; that is, repeating the previously rewarded action sequence despite entering a different stage 2 state. Thus, a significant S0 response × reward interaction was used to assess the usage of habitual action sequences, and the S0 response × reward × stress three-way interaction was used to determine whether the use of action sequences differed between stressed and non-stressed tests. By including only trials in which the stage 2 state differed from the preceding trial, we assessed the extent to which the stage 2 action was controlled by the stage 1 response (i.e., in an chunked action sequence) and whether rats could decouple the sequence to earn reinforcement on trials in which a rare transition led to a different discriminative stimulus in the stage 2 state. For both GLMM analyses, intercept and slopes for each predictor were included as random effects varying across rats.

Discrimination accuracy during stage 2 of the task was analyzed using a repeated measures ANOVA with stress (stressed vs. non-stressed) and trial type (common vs. rare transition) entered as within-subjects factors. Median S0 and stage 2 RTs per rat were also analyzed in a repeated measures ANOVA with stress and stage (S0 vs. stage 2) as within-subjects factors. Planned contrasts were computed between stressed and non-stressed tests.

Statistical analyses were conducted in SPSS (IBM, version 29) and Excel (Microsoft Corporation, 2024). For all GLMMs, tests of fixed effects and coefficients were performed with robust estimation of the parameter estimates covariance matrices. RTs were log (base 10) transformed prior to ANOVA analyses to correct for violations of normality. 95% confidence intervals are reported. Alpha was set at p < 0.05 for all tests.

3. Results

3.1. Intermixed discrimination training

Results of discrimination training are shown in Figure 4A. Rats increased discrimination accuracy across sessions. This observation was supported by a repeated-measures ANOVA with a significant main effect of session, F(2, 30) = 7.98, p = .002, ηp2 = .35. Mean accuracy exceeded 90% on the final day of discrimination training (Day 3: M = 93.27%, SD = 7.44).

Figure 4.

Figure 4.

Training results. Note. A) Mean discrimination accuracy (percent correct) from intermixed discrimination training. B) Mean logarithm of the odds ratio (log odds) of staying on the same stage 1 and stage 2 action after being rewarded on the preceding trial. Log odds > 0 indicates greater likelihood of staying than switching after earning reward. Log odds = 0 indicates equal preference for both actions. Statistical reporting (NS = nonsignificant, * = p < .05, ** = p < .01, *** = p < .001) is for reward by session interaction for each block. Statistical reporting for retraining data shows difference between training sessions 27–29 and retraining sessions. C) Reaction time in stage 1 and stage 2 of the task. D) Mean stage 2 discrimination accuracy during the task. E) Trials per minute and percent of rewarded trials during the task. F) Log odds of staying on the same stage 1 action following reward averaged across the final three sessions (27–29) of training. Dots indicate individual data points. Block 1: ITI = 1s, P(reversal) = 0.50; Block 2: ITI = 3s, P(reversal) = 0.50; Block 3: ITI = 0s, P(reversal) = 0.14 following 4 reinforced trials. Error bars indicate standard error of the mean (SEM).

3.2. Two-stage task

3.2.1. Block 1

The results of two-stage task training are shown in Figure 4B-F. In Block 1, reward earned on the preceding trial was a significant negative predictor of the S0 action, β = −0.85, 95% CI [−1.19, −0.50], SE = 0.17, p < .001, indicating a greater likelihood of switching S0 actions following reward. Controlling for the effect of reward on the preceding trial, session was a significant positive predictor of staying on the same S0 action, β = 0.07, 95% CI [0.05, 0.10], SE = 0.01, p < .001, meaning that the tendency of rats to repeat the same S0 action trial-by-trial increased across sessions. The reward by session interaction, however, was not significant (β = 0.01, p = .64), suggesting that rats were not becoming increasingly likely to repeat S0 actions following reward across sessions (Figure 4B, Block 1).

Rats showed a similar pattern of results for stage 2 behavior in Block 1. Earning reward on the previous trial was significantly associated with switching stage 2 actions from the preceding trial, β = −0.30, 95% CI [−0.59, −0.02], SE = 0.15, p = .038. Effects of session and the reward by session interaction were not significant (respectively: β = 0.03, p = .07; β = −0.02, p = .47), indicating that rats did not increase their trial-by-trial tendency to repeat stage 2 actions across sessions or increase their likelihood of repeating stage 2 actions following reward.

RTs in each stage of the task are shown in Figure 4C. In Block 1, RTs for Stage 1 (β = −15.56, 95% CI [−28.21, −2.91], SE = 6.45, p = .016) and stage 2 (β = −16.17, 95% CI [−23.91, 8.43], SE = 3.95, p < .001) actions both significantly decreased across sessions, likely reflecting increased automaticity in action selection. Likewise, discrimination accuracies significantly increased across Block 1 sessions (Figure 4D) as indicated by a repeated measures ANOVA with a significant main effect of session, F(3.45, 51.76) = 27.41, p < .001, ηp2 = .65. Main effects of session were also significant in ANOVAs for trials attempted per minute (Figure 4E), F(3.00, 44.98) = 16.21, p < .001, ηp2 = .52, and percent of rewarded trials, F(6, 90) = 16.53, p < .001, ηp2 = .52, showing that rats increased their trial attempt and success rates during Block 1.

3.2.2. Block 2

In Block 2, the effect of reward on staying on the same S0 action increased (Figure 4B, Block 2), however it was not significant (β = 0.43, p = .33). Effects of session and reward by session interaction were also nonsignificant, smallest p = .13. Similarly, effects of reward, session, and reward by session interaction for stage 2 actions were also not significant (smallest p = .87), and RTs (Figure 4C, Block 2) in both S0 and stage 2 did not significantly change across sessions (respectively: β = 39.08, p = .24; β = 11.98, p = .34). Repeated measures ANOVAs evaluating discrimination accuracies (Figure 4D, Block 2), trial attempt rates (Figure 4E, Block 2), and percent of rewarded trials found no main effects of session for any of these measures, F’s < 1. Collectively, these results suggested that the changes in task structure imposed during Block 2 (ITI increased from 1s to 3s) did not encourage the acquisition of optimal stay/switch behavior or greater reward rates.

3.2.3. Block 3

In Block 3, the ITI was reduced to 0 s and the probability of reversal was set to 0.14. Here, the effect of previous reward on repeating S0 actions was significant, β = −3.21, 95% CI [−3.78, −2.64], SE = 0.29, p < .001, but negative, reflecting the fact that for the majority of training the mean log odds of staying on the same stage 1 action remained below zero (Figure 4B, Block 3). The effect of session was also significant and negative, β = −.04, 95% CI [−0.06, −0.02], SE = 0.01, p < .001, indicating decreasing perseveration across session and increasing response variability in S0. Importantly, the reward by session interaction was significant and positive, β = 0.13, 95% CI [0.09, 0.16], SE = 0.02, p < .001, suggesting that the probability of repeating the same S0 action following reward increased across sessions (Figure 4B, Block 3). Ultimately, 12 of 16 (75%) rats finished Block 3 with log odds ratios greater than zero when averaged across the final three training sessions (Mdn = 0.40, IQR = −0.04, 0.95; Figure 4F).

A similar pattern of behavior was found for stage 2 actions during Block 3 (Figure 4B, Block 3). The effect of previous reward on repeating stage 2 actions was significant but negative, β = −2.26, 95% CI [−2.85, −1.67], SE = 0.30, p < .001, as was the effect of session, β = −.04, 95% CI [−0.05, −0.02], SE = 0.01, p < .001. Again, the model indicated a significant positive reward by session interaction, β = 0.11, 95% CI [0.08, 0.15], SE = 0.02, p < .001, suggesting that as sessions progressed, rats were increasingly likely to repeat stage 2 actions following reward on the preceding trial and switch stage 2 actions when reward was omitted.

Models for both S0 and stage 2 RT indicated that RTs did not significantly change across sessions during Block 3 (S0: β = 0.78, p = .48; Stage 2: β = −2.10, p = .07; Figure 4C, Block 3). Random intercepts for RT varied significantly between rats in both S0 (Z = 2.68, p = .007) and stage 2 (Z = 2.61, p = .009) suggesting that rats showed substantial RT variability in both stages. Repeated measures ANOVAs for both discrimination accuracies and trial attempt rates found no significant main effect of session, largest F = 1.76 (Figure 4D-E, Block 3). However, the percentage of rewarded trials per session increased across Block 3, suggesting that as rats became increasingly likely to repeat actions that were rewarded on the preceding trial, they likewise earned reward on a greater proportion of trials overall (Figure 4E, Block 3). This result was supported by a repeated measures ANOVA with a main effect of session, F(4.95, 74.22) = 12.66, p < .001, ηp2 = .46.

3.3. Two-stage task retraining

Two-stage task retraining data is shown in Figures 4B-E. The GLMM analysis suggested that rats did not significantly differ in their likelihood of repeating S0 responses after earning reward between training end and retraining sessions (β = 0.11, p = .45; Figure 4B), nor did they differ in their S0 RTs (β = −7.64, p = .16; Figure 4C). However, rats were significantly less likely to repeat stage 2 responses after earning reward during retraining compared to training end, β = −0.50, 95% CI [−0.83, −0.18], SE = 0.17, p = .003; Figure 4B), possibly due to experience with the new stage 2 transition probabilities introduced during the probe test. Stage 2 RTs also significantly decreased during retraining, β = −16.32, 95% CI [−23.89, −8.75], SE = 3.86, p < .001; Figure 4C.

Discrimination accuracy did not significantly differ between training end (M = 77.88%, SD = 7.68) and retraining (M = 79.54%, SD = 5.90; t(15) = 1.06, p =.31; Figure 4D). Trial attempt rates were significantly greater during retraining (M = 19.70, SD = 3.64) compared to training end (M = 18.52, SD = 3.60; t(15) = 3.47, p =.003; Figure 4E). Reward rates also significantly increased during retraining (M = 49.29%, SD = 8.08) compared to training end (M = 44.77, SD = 10.16; t(15) = 2.95, p =.01; Figure 4E). Overall, these results suggested that rats modestly improved task speed and reward rates during the retraining period and showed increased response diversity in stage 2 of the task.

3.4. Probe tests

Results of the probe tests are shown in Figure 5. When stressed, rats attempted significantly fewer trials per min (M = 19.16, SD = 4.57) compared to non-stressed (M = 22.80, SD = 5.08; t(15) = −5.63, p < .001; Figure 5A). Obtained reward rates were slightly lower when stressed (M = 36.60%, SD = 5.74) compared to non-stressed (M = 38.81%, SD = 7.76); although, this difference was not statistically significant, t(15) = −1.91, p = .08 (Figure 5B).

Figure 5.

Figure 5.

Probe test performance. Note. A) Trials attempted per min. B) Percent of trials rewarded. C) Percentage of trials where the stage 1 action from the preceding trial was repeated as a function of whether the preceding trial was rewarded (reward/no reward), and whether the test occurred under stressed or non- stressed conditions. Str. = stressed, N-Str. = non-stressed. D) Regression coefficients from the GLMM model predicting stage 1 stay/switch behavior based on reward earned on the preceding trial (Reward), stress condition (Stress), and a reward by stress interaction assessing whether the effect of reward on stage 1 stay/switch behavior differed between stress conditions. E) Percentage of trials where the stage 2 action from the preceding trial was repeated as function of whether the preceding trial was rewarded (reward/no reward), whether the stage 1 action from the preceding trial was repeated (S0 stay/ S0 switch), and stress condition. F) Regression coefficients from the GLMM model predicting stage 2 stay/switch behavior based on reward earned on the preceding trial (Reward), whether the stage 1 action from the preceding trial was repeated (S0 Stay), stress condition (Stress), a reward by S0 stay interaction term assessing whether the effect of reward on stage 2 stay/switch behavior depended on stage 1 choice, and an interaction term assessing whether this effect depended on stress condition. G) Stage 2 discrimination accuracy (percent correct) during stressed and non-stressed probe tests as a function of whether the trial was common or rare. H) RTs (s) during stressed and non-stressed probe tests for stage 1 and stage 2 of the task. Dots represent individual data points. ns = nonsignificant, * = p < .05, ** = p < .01, *** = p < .001. Error bars indicate 95% confidence intervals.

Figures 5C and 5D show responses in S0 in the probe tests. Across stressed and non-stressed test conditions, reward earned on the preceding trial was a significant positive predictor of repeating the same S0 action, β = 0.42, 95% CI [0.10, 0.73], SE = 0.16, p = .01. This indicates that S0 actions during the probe tests were likely guided by goal-directed tracking of trial-to-trial reinforcement. Stress significantly increased the likelihood of repeating the S0 action, β = 0.19, 95% CI [0.06, 0.32], SE = 0.07, p = .005, indicating that rats became more perseverative when stressed. The reward by stress interaction was not significant (β = −0.07, p = .62), suggesting that reward earned on preceding trials guided behavior to a similar extent under stress and in its absence. The variance in the random effect of reward was significant (Z = 3.00, p = .003), consistent with the heterogeneity observed at the end of Block 3.

Stage 2 action data from probe trials is shown in Figures 5E and 5F. Effects of staying on the same S0 action, reward earned on the preceding trial, and stress were not significant (smallest p = .45). However, the S0 action by reward interaction was significant, β = 0.60, 95% CI [0.06, 1.13], SE = 0.27, p = .03, suggesting that across stress conditions, rats tended to repeat stage 2 actions after earning reward on the preceding trial and repeating the same S0 action– thereby repeating the previously rewarded action sequence despite entering a different stage 2 state. This suggests that performance of the action sequence was not sensitive to the changed stage 2 state and was consistent with a chunked action sequence. The S0 by reward interaction did not significantly differ between stressed and non-stressed conditions (β = −0.18, p = .29), indicating that rats used action sequences to a similar extent in the stressed and non-stressed probe tests.

Stage 2 accuracy data is shown in Figure 5G. Mean accuracies for common transition trials were similar under stressed (M = 76.44%, SD = 8.37) and non-stressed conditions (M = 78.04%, SD = 7.02). For rare trials, discrimination accuracies were higher when stressed (M = 31.41%, SD = 10.79) compared to non-stressed (M = 24.66%, SD = 8.21). Confidence interval upper bounds for accuracies in both stress conditions fell below 50%, suggesting that rats consistently failed to adjust their action sequence once initiated on rare trials. The ANOVA indicated significant main effects of stress, F(1, 15) = 5.12, p = .039, ηp2 = .25, and transition type, F(1, 15) = 264.87, p < .001, ηp2 = .95, suggesting that stage 2 accuracy was increased under stress, and higher on common trials. The stress by trial type interaction was not significant, F(1, 15) = 3.40, p = .09, ηp2 = .18, however planned contrasts suggested that common-transition accuracy was similar in stressed and non-stressed tests, t(15) = 0.85, p = .41, but rare-transition accuracy was significantly higher in the stressed test, t(15) = −2.21, p = .043. This suggests that stress primarily affected performance by increasing discrimination accuracy on rare transition trials.

Figure 5H shows RT data from S0 and stage 2. Mean S0 RTs were longer under stressed conditions (M = 1.64s, SD = 0.71) than non-stressed (M = 1.40s, SD = 0.56), as were stage 2 RTs (Stressed: M = 0.85s, SD = 0.20; Non-stressed: M = 0.76s, SD = 0.16). The analysis found significant main effects of stress, F(1, 15) = 9.47, p = .008, ηp2 = .39, and stage, F(1, 15) = 85.70, p < .001, ηp2 = .85, suggesting that RTs were longer overall when stressed compared to non-stressed, and were longer during stage 1 of the task compared to stage 2. The stress by stage interaction was not significant, F < 1.

4. Discussion

In female rats, two-stage response patterns suggested that action planning was guided by online tracking of trial-to-trial reinforcement, consistent with goal-directed planning of action sequences. This was observed in within-subject performance under stressed and non-stressed conditions. Once selected, action sequences were performed without evaluation at intermediate steps, as a chunked or habitual sequence. Evidence of habitual action sequences was observed on probe trials during the stressed and non-stressed tests. Therefore, the results suggest that acute stress did not interfere with female rats’ hierarchical goal-directed selection of habitual action sequences or prompt the deconstruction of the chunked action into a goal-directed sequence.

In the probe tests, rats tended to repeat the stage 1 action if the preceding trial ended in reward. Goal-directed planning of action sequences was evidenced by rats switching stage 1 actions in response to the omission of reward. Such behavior requires monitoring of trial-to-trial reinforcement to guide choice to repeat the same stage 1 action following rewarded but not nonrewarded trials.

Probe tests further revealed habitual execution of action sequences. First, repeating the stage 1 action after earning reward significantly predicted repeating the same stage 2 action, even when a different stage 2 state had been entered as signaled by a distinct discriminative stimulus (rare transitions). This suggests that rats completed the initiated sequence and did not withhold the second action in the sequence or adjust their stage 2 choice to correspond to the contingency occasioned by the discriminative stimulus.

The analysis revealed several effects of stress on the goal-directed planning and execution of action sequences. First, in models that controlled for the influence of reward on the preceding trial, rats tended to repeat the stage 1 action. This could suggest that following stress, rats exhibited a decreased tendency to explore action sequences and preferentially selected the same stage 1 action as on the preceding trial. Stress-induced changes in action selection have been observed in explore-exploit behavior with rats and humans (Harms, 2017; Lenow et al., 2017; Matisz et al., 2021). One hypothesis suggests that stressful events promote negative appraisals of the environment that generalize to decision-making. In this framework, positive environmental appraisals strongly predict exploration of new actions, and negative appraisals can facilitate a decision bias towards exploiting the resources at hand (Constantino & Daw, 2015).

Stress increased the latency between responses within action sequences and allowed rats to make more correct discriminations on rare trials. However, the prevailing tendency to make the incorrect discrimination suggests that stress likely did not induce a complete shift to deliberate choice on rare trials. Nonetheless, stress may have decreased the number of errors in stage 2 by decoupling the timing of stage 1 and stage 2 actions and thus allowing the discriminative stimuli to influence response selection at stage 2. Taken together, the latency results suggest that acute stress influenced the microstructure of goal-directed planning and habitual action sequence execution rather than altering the general behavioral macrostructure.

Although our task structure and training methodology was based on Dezfouli and Balleine (2019), who found evidence of step-by-step deliberation in male rats, the present experiment did not find evidence of deliberation at stage 2 under non-stressed or stressed conditions. Several differences in the initial training method may have influenced the results. First, intermixed discrimination training may have encouraged response switching as rats had the opportunity to attempt multiple responses within a 10s window if their first response was incorrect. This may have facilitated action sequences composed of lever switching. Other studies suggest that action sequences learned in one stage of a task can transfer to subsequent stages and serve as the building blocks for future behaviors (Ferguson et al., 2024). Second, we explored different reversal probabilities during initial training phases. Experience with a .50 reversal probability in Blocks 1 and 2 could have further facilitated response switching. Perhaps more importantly, unpredictability may have interfered with active monitoring of trial-to-trial reinforcement because either response sequence was reinforced with equal probability. Nonetheless, rats acquired accurate performance in Block 3 and maintained it during and between tests. Since the stimulus-response mappings remained the same in Block 3, response sequences could be selected as habitual action sequences by a goal-directed planner. Thus, it remains possible that stress could have a different effect when the character of the sequence reflects goal direction. Future studies may also address how reversal probability influences the tendency for rats to perform habitual versus goal-directed action sequences as revealed by probe tests (cf. Dezfouli & Balleine, 2019). Whether stress alters how animals switch from deliberation over single instrumental actions to performing them in a chunked habitual sequence likely depends on a multiple training variables.

Differences in action sequence behavior by male rats in Dezfouli & Balleine (2019) and the female rats of the present study could also be due to sex. Prior work by our laboratory has shown that female rats shift from goal-directed to habitual control over instrumental responding with less training than males (Schoenberg et al., 2019); although, the extent to which this sex difference extends to instrumental sequences of actions has yet to be tested. One possibility is that female rats may show a greater tendency to chunk sequential actions into habitual sequences than males, who may instead use discrete successive choices under similar conditions. This too represents an important direction for future study.

The present results add to recent studies that challenge the notion that acute stress causes a complete switch from goal-directed to habitual modes of behavior (Buabang et al., 2022; Otto, Raio, et al., 2013; Radenbach et al., 2015; Smeets et al., 2023). Female rats can retain goal-directed planning following stress when task incentive structure encourages active monitoring of feedback. Intact goal-direction following stress, as found here, diverges from previous findings of enhanced habit following stress in free operant tasks (Braun & Hauber, 2013; Dias-Ferreira et al., 2009; Dougherty et al., 2024). Differences in task-related attentional demands may influence performance when stressed, as the heightened state of vigilance produced by acute stress could interfere with responding as rats divert attention toward outward monitoring of the environment for threat (Blanchard et al., 1990; van Marle et al., 2009).

A role for attention to the response in habit learning has been proposed by Bouton and colleagues (Bouton, 2021; Thrailkill et al., 2018, 2021). Simple free operant tasks may require little monitoring of the response because the conditioning context reliably predicts one response will earn reward (Bouton, 2021; see also Kosaki & Dickinson, 2010). As a result, defensive monitoring can occur with little cost to obtained reinforcement. Such conditions may facilitate insensitivity to outcome devaluation following acute stress even after minimal free-operant training (e.g. Dougherty et al., 2024). Such a perspective suggests that stress enhances habitual control over the response by increasing defensive monitoring which may in turn reduce resources available for processing of the behavior, and potentially the devalued outcome. In contrast, concurrent choice tasks, which may be thought to include two-stage tasks, require active monitoring of the response-outcome contingencies. In a two-stage task, the reward associated with selecting an entire course of actions remains unpredictable and requires constant attention to action selection in stage 1 to maintain reinforcement. Interestingly, fewer resources are needed in stage 2, which is highly predictable and was well-learned prior to the probe tests. Under stress this may allow rats to allocate attentional resources to defensive monitoring without decreasing reinforcement rate. Consistent with the present results, defensive monitoring induced by stress may reduce trial attempt rates and alter the timing of habitual action sequences at the sequence level while leaving goal-directed planning intact.

Several limitations of this study should be made clear. Training included two early exploratory phases (Blocks 1 and 2) to establish performance on each lever and explore sequence learning under conditions when reversals were unpredictable to the animal. As noted above, this preliminary experience could have resulted in latent inhibition, where strategies learned during Blocks 1 and 2 (reversal p = .50) could have interfered with performance in Block 3 (reversal p = .14). These initial phases were experienced by all animals and the effect of stress was tested within subject. Nonetheless, future studies could leave out these exploratory phases to avoid possible history effects and further clarify how stress influences performance. An additional limitation relates to the lack of a direct measurement of the physiological stress response. The inclusion of physiological measures of acute stress is an important area for future work.

Collectively, the present findings suggest that female rats can retain goal-directed planning of sequential actions following stress despite exhibiting less diversity in action selection and weaker coupling of actions within a sequence under stressed conditions. The present findings highlight the importance of testing stress effects in diverse contexts and the important functional role of hierarchical selection of action sequences (Dezfouli & Balleine, 2012; Smith & Graybiel, 2016). These results also extend previous work on the effects of stress on goal-directed and habitual behaviors in male rats by characterizing the influence of acute stress in females – a previously unstudied area (Braun & Hauber, 2013; Dias-Ferreira et al., 2009; Taylor et al., 2014). Next steps in this line of work should seek to identify potential sex differences in the influence of stress on hierarchical organization of goal-directed and habitual responding. As stress-induced pathologies of decision-making play a critical role in public health, further study is needed to understand the environmental, associative, neural, and molecular mechanisms for how and when stress influences goal-directed selection of action sequences.

Highlights.

  • Female rats maintained hierarchical planning structure following acute stress

  • Acute stress induced systematic differences in two-stage task performance

  • Stress increased reaction time and response variability within action sequences

  • Attentional demands of the decision task may modulate goal-direction under stress

Funding

This research was supported by funding from the University of Vermont Department of Psychological Science. EAT was supported by K01-DA044456 and P20-GM103644 from the National Institutes of Health.

Footnotes

CRediT authorship contribution statement

Russell Dougherty: Conceptualization, Methodology, Formal Analysis, Investigation, Data Curation, Writing – Original Draft, Writing – Review & Editing, Visualization, Project Administration. Eric A. Thrailkill: Conceptualization, Methodology, Software, Writing – Review & Editing, Supervision. Sarah Van Horn: Investigation. Auny Kussad: Investigation. Donna J. Toufexis: Conceptualization, Resources, Writing – Review & Editing, Supervision, Funding Acquisition.

Disclosures

The authors report there are no competing interests to declare.

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Data availability statement

Data is available in the following repository: https://doi.org/10.17632/v3y4bnmds3.1

References

  1. Adams CD, & Dickinson A. (1981). Instrumental Responding following Reinforcer Devaluation. The Quarterly Journal of Experimental Psychology Section B, 33(2b), 109–121. 10.1080/14640748108400816 [DOI] [Google Scholar]
  2. Aloisi AM, Ceccarelli I, & Lupo C. (1998). Behavioural and hormonal effects of restraint stress and formalin test in male and female rats. Brain Research Bulletin, 47(1), 57–62. 10.1016/S0361-9230(98)00063-X [DOI] [PubMed] [Google Scholar]
  3. Ballard IC, Waskom M, Nix KC, & D’Esposito M. (2024). Reward Reinforcement Creates Enduring Facilitation of Goal-directed Behavior. Journal of Cognitive Neuroscience, 1–16. 10.1162/jocn_a_02150 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Balleine BW, & Dezfouli A. (2019). Hierarchical Action Control: Adaptive Collaboration Between Actions and Habits. Frontiers in Psychology, 10, 2735. 10.3389/fpsyg.2019.02735 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Blanchard RJ, Blanchard DC, Rodgers J, & Weiss SM (1990). The characterization and modelling of antipredator defensive behavior. Neuroscience & Biobehavioral Reviews, 14(4), 463–472. 10.1016/S0149-7634(05)80069-7 [DOI] [PubMed] [Google Scholar]
  6. Bouton ME (2021). Context, attention, and the switch between habit and goal-direction in behavior. Learning & Behavior, 49(4), 349–362. 10.3758/s13420-021-00488-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Braun S, & Hauber W. (2013). Acute stressor effects on goal-directed action in rats. Learning & Memory, 20(12), 700–709. 10.1101/lm.032987.113 [DOI] [PubMed] [Google Scholar]
  8. Buabang EK, Boddez Y, Wolf OT, & Moors A. (2022). The role of goal-directed and habitual processes in food consumption under stress after outcome devaluation with taste aversion. Behavioral Neuroscience. 10.1037/bne0000439 [DOI] [PubMed] [Google Scholar]
  9. Constantino S, & Daw ND (2015). Learning the opportunity cost of time in a patch-foraging task. Cognitive, Affective & Behavioral Neuroscience, 15(4), 837–853. 10.3758/s13415-015-0350-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Cremer A, Kalbe F, Gläscher J, & Schwabe L. (2021). Stress reduces both model-based and model-free neural computations during flexible learning. NeuroImage, 229, 117747. 10.1016/j.neuroimage.2021.117747 [DOI] [PubMed] [Google Scholar]
  11. Daw ND, Gershman SJ, Seymour B, Dayan P, & Dolan RJ (2011). Model-based influences on humans’ choices and striatal prediction errors. Neuron, 69(6), 1204–1215. 10.1016/j.neuron.2011.02.027 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Dezfouli A, & Balleine BW (2012). Habits, action sequences and reinforcement learning. The European Journal of Neuroscience, 35(7), 1036–1051. 10.1111/j.1460-9568.2012.08050.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Dezfouli A, & Balleine BW (2013). Actions, Action Sequences and Habits: Evidence That Goal-Directed and Habitual Action Control Are Hierarchically Organized. PLOS Computational Biology, 9(12), e1003364. 10.1371/journal.pcbi.1003364 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Dezfouli A, & Balleine BW (2019). Learning the structure of the world: The adaptive nature of state-space and action representations in multi-stage decision-making. PLOS Computational Biology, 15(9), e1007334. 10.1371/journal.pcbi.1007334 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Dezfouli A, Lingawi NW, & Balleine BW (2014). Habits as action sequences: Hierarchical action control and changes in outcome value. Philosophical Transactions of the Royal Society B: Biological Sciences, 369(1655), 20130482. 10.1098/rstb.2013.0482 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Dias-Ferreira E, Sousa JC, Melo I, Morgado P, Mesquita AR, Cerqueira JJ, Costa RM, & Sousa N. (2009). Chronic stress causes frontostriatal reorganization and affects decision-making. Science (New York, N.Y.), 325(5940), 621–625. 10.1126/science.1171203 [DOI] [PubMed] [Google Scholar]
  17. Dickinson A, & Balleine B. (1994). Motivational control of goal-directed action. Animal Learning & Behavior, 22(1), 1–18. 10.3758/BF03199951 [DOI] [Google Scholar]
  18. Dougherty R, Thrailkill EA, Mohammed Z, VonDoepp S, Hilton-Vanosdall E, Charette S, Van Horn S, Quirk A, Kraus A, & Toufexis DJ (2024). Acute stress facilitates habitual behavior in female rats. Physiology & Behavior, 275, 114456. 10.1016/j.physbeh.2024.114456 [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Du Y, Krakauer JW, & Haith AM (2022). The relationship between habits and motor skills in humans. Trends in Cognitive Sciences, 26(5), 371–387. 10.1016/j.tics.2022.02.002 [DOI] [PubMed] [Google Scholar]
  20. Favila N, Gurney K, & Overton PG (2024). Role of the basal ganglia in innate and learned behavioural sequences. Reviews in the Neurosciences, 35(1), 35–55. 10.1515/revneuro-2023-0038 [DOI] [PubMed] [Google Scholar]
  21. Ferguson LA, Matamales M, Nolan C, Balleine BW, & Bertran-Gonzalez J. (2024). Adaptation of sequential action benefits from timing variability related to lateral basal ganglia circuitry. iScience, 27(3), 109274. 10.1016/j.isci.2024.109274 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Frölich S, Esmeyer M, Endrass T, Smolka MN, & Kiebel SJ (2023). Interaction between habits as action sequences and goal-directed behavior under time pressure. Frontiers in Neuroscience, 16. https://www.frontiersin.org/articles/10.3389/fnins.2022.996957 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Garr E, & Delamater AR (2019). Exploring the relationship between actions, habits, and automaticity in an action sequence task. Learning & Memory, 26(4), 128–132. 10.1101/lm.048645.118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Halbout B, Marshall AT, Azimi A, Liljeholm M, Mahler SV, Wassum KM, & Ostlund SB (2019). Mesolimbic dopamine projections mediate cue-motivated reward seeking but not reward retrieval in rats. eLife, 8, e43551. 10.7554/eLife.43551 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Harms MB (2017). Stress and Exploitative Decision-Making. Journal of Neuroscience, 37(42), 10035–10037. 10.1523/JNEUROSCI.2169-17.2017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Hermans EJ, Henckens MJAG, Joëls M, & Fernández G. (2014). Dynamic adaptation of large-scale brain networks in response to acute stressors. Trends in Neurosciences, 37(6), 304–314. 10.1016/j.tins.2014.03.006 [DOI] [PubMed] [Google Scholar]
  27. Keele SW (1968). Movement control in skilled motor performance. Psychological Bulletin, 70(6, Pt.1), 387–403. 10.1037/h0026739 [DOI] [Google Scholar]
  28. Kosaki Y, & Dickinson A. (2010). Choice and contingency in the development of behavioral autonomy during instrumental conditioning. Journal of Experimental Psychology: Animal Behavior Processes, 36(3), 334–342. 10.1037/a0016887 [DOI] [PubMed] [Google Scholar]
  29. Lenow JK, Constantino SM, Daw ND, & Phelps EA (2017). Chronic and Acute Stress Promote Overexploitation in Serial Decision Making. The Journal of Neuroscience, 37(23), 5681–5689. 10.1523/JNEUROSCI.3618-16.2017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Linnebank FE, Kindt M, & de Wit S. (2018). Investigating the balance between goal-directed and habitual control in experimental and real-life settings. Learning & Behavior, 46(3), 306–319. 10.3758/s13420-018-0313-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Matisz CE, Badenhorst CA, & Gruber AJ (2021). Chronic unpredictable stress shifts rat behavior from exploration to exploitation. Stress, 24(5), 635–644. 10.1080/10253890.2021.1947235 [DOI] [PubMed] [Google Scholar]
  32. Miller KJ, Botvinick MM, & Brody CD (2017). Dorsal hippocampus contributes to model-based planning. Nature Neuroscience, 20(9), Article 9. 10.1038/nn.4613 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Miller KJ, Botvinick MM, & Brody CD (2022). Value representations in the rodent orbitofrontal cortex drive learning, not choice. eLife, 11, e64575. 10.7554/eLife.64575 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Morris A, & Cushman F. (2019). Model-Free RL or Action Sequences? Frontiers in Psychology, 10, 2892. 10.3389/fpsyg.2019.02892 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Nissen MJ, & Bullemer P. (1987). Attentional requirements of learning: Evidence from performance measures. Cognitive Psychology, 19(1), 1–32. 10.1016/0010-0285(87)90002-8 [DOI] [Google Scholar]
  36. Otto AR, Gershman SJ, Markman AB, & Daw ND (2013). The Curse of Planning: Dissecting multiple reinforcement learning systems by taxing the central executive. Psychological Science, 24(5), 10.1177/0956797612463080. https://doi.org/10.1177/0956797612463080 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Otto AR, Raio CM, Chiang A, Phelps EA, & Daw ND (2013). Working-memory capacity protects model-based learning from stress. Proceedings of the National Academy of Sciences, 110(52), 20941–20946. 10.1073/pnas.1312011110 [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Pew RW (1966). Acquisition of hierarchical control over the temporal organization of a skill. Journal of Experimental Psychology, 71(5), 764–771. 10.1037/h0023100 [DOI] [PubMed] [Google Scholar]
  39. Radenbach C, Reiter AMF, Engert V, Sjoerds Z, Villringer A, Heinze H-J, Deserno L, & Schlagenhauf F. (2015). The interaction of acute and chronic stress impairs model-based behavioral control. Psychoneuroendocrinology, 53, 268–280. 10.1016/j.psyneuen.2014.12.017 [DOI] [PubMed] [Google Scholar]
  40. Raio CM, Konova AB, & Otto AR (2020). Trait impulsivity and acute stress interact to influence choice and decision speed during multi-stage decision-making. Scientific Reports, 10(1), 7754. 10.1038/s41598-020-64540-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Reis DG, Scopinho AA, Guimarães FS, Corrêa FMA, & Resstel LBM (2011). Behavioral and Autonomic Responses to Acute Restraint Stress Are Segregated within the Lateral Septal Area of Rats. PLoS ONE, 6(8), e23171. 10.1371/journal.pone.0023171 [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Schoenberg HL, Sola EX, Seyller E, Kelberman M, & Toufexis DJ (2019). Female rats express habitual behavior earlier in operant training than males. Behavioral Neuroscience, 133(1), 110–120. 10.1037/bne0000282 [DOI] [PubMed] [Google Scholar]
  43. Schwabe L, & Wolf OT (2009). Stress Prompts Habit Behavior in Humans. The Journal of Neuroscience, 29(22), 7191–7198. 10.1523/JNEUROSCI.0979-09.2009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Schwabe L, & Wolf OT (2010). Socially evaluated cold pressor stress after instrumental learning favors habits over goal-directed action. Psychoneuroendocrinology, 35(7), 977–986. 10.1016/j.psyneuen.2009.12.010 [DOI] [PubMed] [Google Scholar]
  45. Smeets T, Ashton SM, Roelands SJAA, & Quaedflieg CWEM (2023). Does stress consistently favor habits over goal-directed behaviors? Data from two preregistered exact replication studies. Neurobiology of Stress, 23, 100528. 10.1016/j.ynstr.2023.100528 [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Smith KS, & Graybiel AM (2013). A Dual Operator View of Habitual Behavior Reflecting Cortical and Striatal Dynamics. Neuron, 79(2), 361–374. 10.1016/j.neuron.2013.05.038 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Smith KS, & Graybiel AM (2016). Habit formation. Dialogues in Clinical Neuroscience, 18(1), 33–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Sommet N, & Morselli D. (2017). Keep Calm and Learn Multilevel Logistic Modeling: A Simplified Three-Step Procedure Using Stata, R, Mplus, and SPSS. International Review of Social Psychology, 30(1), 203–218. 10.5334/irsp.90 [DOI] [Google Scholar]
  49. Taylor SB, Anglin JM, Paode PR, Riggert AG, Olive MF, & Conrad CD (2014). Chronic stress may facilitate the recruitment of habit- and addiction-related neurocircuitries through neuronal restructuring of the striatum. Neuroscience, 280, 231–242. 10.1016/j.neuroscience.2014.09.029 [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Thrailkill EA, Michaud NL, & Bouton ME (2021). Reinforcer Predictability and Stimulus Salience Promote Discriminated Habit Learning. Journal of Experimental Psychology. Animal Learning and Cognition, 47(2), 183. 10.1037/xan0000285 [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Thrailkill EA, Trask S, Vidal P, Alcalá JA, & Bouton ME (2018). Stimulus Control of Actions and Habits: A Role for Reinforcer Predictability and Attention in the Development of Habitual Behavior. Journal of Experimental Psychology. Animal Learning and Cognition, 44(4), 370–384. 10.1037/xan0000188 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. van Elzelingen W, Warnaar P, Matos J, Bastet W, Jonkman R, Smulders D, Goedhoop J, Denys D, Arbab T, & Willuhn I. (2022). Striatal dopamine signals are region specific and temporally stable across action-sequence habit formation. Current Biology, 32(5), 1163–1174.e6. 10.1016/j.cub.2021.12.027 [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. van Marle HJF, Hermans EJ, Qin S, & Fernández G. (2009). From Specificity to Sensitivity: How Acute Stress Affects Amygdala Processing of Biologically Salient Stimuli. Biological Psychiatry, 66(7), 649–655. 10.1016/j.biopsych.2009.05.014 [DOI] [PubMed] [Google Scholar]
  54. Wood W, Quinn JM, & Kashy DA (2002). Habits in everyday life: Thought, emotion, and action. Journal of Personality and Social Psychology, 83(6), 1281–1297. 10.1037/0022-3514.83.6.1281 [DOI] [PubMed] [Google Scholar]
  55. Yu Z, Guindani M, Grieco SF, Chen L, Holmes TC, & Xu X. (2022). Beyond t test and ANOVA: Applications of mixed-effects models for more rigorous statistical analysis in neuroscience research. Neuron, 110(1), 21–35. 10.1016/j.neuron.2021.10.030 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data is available in the following repository: https://doi.org/10.17632/v3y4bnmds3.1

RESOURCES