Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Jan 1.
Published in final edited form as: J Exp Anal Behav. 2023 Dec 27;121(1):88–96. doi: 10.1002/jeab.894

Habit and persistence

Mark Bouton 1
PMCID: PMC10842266  NIHMSID: NIHMS1948558  PMID: 38149526

Abstract

Voluntary behaviors (operants) can come in two varieties: Goal-directed actions, which are emitted based on the remembered value of the reinforcer, and habits, which are evoked by antecedent cues and performed without the reinforcer’s value in active memory. The two are perhaps most clearly distinguished with the reinforcer-devaluation test: Goal-directed actions are suppressed when the reinforcer is separately devalued and responding is tested in extinction, but habitual behaviors are not. But what is the function of habit learning? Habits are often thought to be strong and unusually persistent. The present selective review examines this idea by asking whether habits identified by the reinforcer-devaluation test are more resistant to extinction, resistant to the effects of other contingency change, vulnerable to relapse, resistant to the weakening effects of context change, or permanently in place once they are learned. Surprisingly little evidence supports the idea that habits are permanent or more persistent. Habits are more context-specific than goal-directed actions are. Methods that make behavior persistent do not necessarily work by encouraging habit. The function of habit learning may not be to make a behavior strong or more persistent but to make it automatic and efficient in a particular context.

Keywords: goal direction, habit, operant behavior, persistence, reinforcer devaluation


Many disciplines in behavioral science (e.g., associative learning, behavioral neuroscience, social psychology, economics, and computer science) now suppose that voluntary (operant) behavior can take separable goal-directed and habitual forms. Casually speaking, a goal-directed action (hereafter, “action”) appears deliberate and goal oriented (thus the name), whereas a habit is more reflexive, automatic, and emitted as a response to trigger stimuli without a goal “in mind.” (See below for a more formal behavioral definition.) Habits are adaptive at least in part because they allow performance to be automatic and efficient; they require few cognitive resources and make it easier to navigate and multitask in a complex world. William James, an early student of habit, recognized this function in his 1890 chapter discussing it. He said, using language of the day, “the more of the details of our daily life we can hand over to the effortless custody of automatism, the more our higher powers of mind will be set free for their own proper work” (James, 1890, p. 122). However, habits are also often assumed to be strong and unusually persistent. For example, a prominent feature of a smoker’s or drug abuser’s bad habits is that they are supposed to be hard to change. And when we encourage development of healthy behaviors, training them as habits might help make them “stick” (cf. Poldrack, 2021; Wood, 2019). In this short and selective review, I will look at laboratory evidence, often collected in my own laboratory, to ask whether habitual behaviors are necessarily stronger and stickier than goal-directed ones. To anticipate the argument, the evidence suggests they are not; contrary to the sticky-habit notion, a behavior’s persistence is hard to predict from its status as habit or goal-directed action. The function of habit learning, I conclude, is not to make a behavior strong or more persistent but to make it automatic and efficient in the right context.

Distinguishing between habit and goal direction

The distinction between habitual and goal-directed behavior in animal learning labs is supported most clearly by the reinforcer-devaluation test. In that test, an animal first learns an operant behavior, and then in a separate step the experimenter devalues the reinforcer by conditioning a taste aversion to it or prefeeding and satiating the animal on it (or on a control reinforcer) just before the test. After either type of devaluation, the rate of operant responding is tested in extinction, so no reinforcers are delivered during the test. If the operant is goal-directed, response rate adjusts downward, as if the animal knows the reinforcer that the action produces is no longer any good. In contrast, if the behavior is a habit, even complete devaluation of the reinforcer (so that the animal now rejects it) has no effect on the response; the animal continues to perform the response as if nothing has changed. To repeat, the final test is conducted in extinction, so the animal never has a chance to associate the response with the devalued reinforcer in experience. And the reinforcer that is devalued must be the one that was specifically associated with the operant response (e.g., Bouton, Allan, et al., 2021; Colwill & Rescorla, 1985b). Thus, when reinforcer devaluation suppresses the response in the extinction test, the animal is thought to remember the reinforcer that was associated with the response and then respond according to its current value. This adjustment of behavior according to the changed value of the goal goes to the heart of what we mean by goal direction. And the fact that a habitual response continues regardless of reinforcer devaluation makes it seem blind and relatively cognitionless.

The reinforcer-devaluation effect is consistent with a long history of research in learning theory supporting the idea that reinforcers (or the expectation of them) motivate behavior as well as reinforce it (e.g., Bouton & Balleine, 2019). This view has traditionally been supported by latent learning effects and by the rapid contrast effects that are seen when the size or value of the reinforcer shifts (Crespi, 1942; Flaherty, 1996; Tinklepaugh, 1928; Tolman & Honzik, 1930). Besides its theoretical implications, the reinforcer-devaluation effect and the action–habit distinction have practical implications for behavioral practitioners. If a goal-directed action is affected by separate manipulation of the reinforcer’s value, then reinforcer revaluation treatments can expand the range of things that can be done to help correct behavioral excesses or deficits. That is, we might be able to decrease the strength of an unwanted behavior by separately reducing the reinforcer’s value or increase the strength of a wanted one by separately increasing it. And importantly, because habits are not so affected by reinforcer revaluation, understanding the conditions that make operants become goal directed or habitual is critical for understanding and predicting behavioral control.

Operants appear to be goal directed early in their training but become habitual after extended practice (e.g., Adams, 1982; Dickinson et al., 1995; Thrailkill & Bouton, 2015). This conversion of action into habit can be hastened by exposure to drugs of abuse (e.g., Corbit et al., 2012, 2014; Furlong et al., 2014; Nelson & Killcross, 2006). Although the latter effects may be mediated by the direct effects of these substances on the dorsomedial striatum (e.g., Becchi et al., 2022; Shipman & Corbit, 2022), a brain substrate of goal direction, there are currently three behavioral accounts of how habit is learned with overtraining. One is Thorndike’s (1911) law of effect: He proposed that the effect of a reinforcer is to strengthen a connection between the stimulus (S) and the response (R), with the more the better. The S-R connection is the habit. Because he never considered what we now call goal direction, Thorndike’s purely S-R view of instrumental behavior does not explain the transition from action to habit or the behavioral suppression observed in the reinforcer-devaluation test—which theorists assume depends on learning the relation between R and its outcome (O; e.g., Balleine & O’Doherty, 2010; Colwill & Rescorla, 1985b; de Wit & Dickinson, 2009). A second view of how an operant becomes a habit is the rate-correlation view (Dickinson, 1989; Perez & Dickinson, 2020). In this account, which is designed to explain habit acquisition in free-operant experiments in which organisms respond freely and repeatedly for intermittently scheduled rewards, an operant becomes habitual as it becomes regular and the experienced correlation between behavior rate and reinforcer rate becomes weak—a feature of responding on interval schedules of reinforcement (Baum, 1973). A third approach claims that habit develops as the reinforcer becomes well-predicted by discriminative stimuli in the environment (Bouton, 2021; Thrailkill et al. 2018, 2021). When the reinforcer becomes unsurprising, the animal pays less attention to its behavior, just as organisms pay less attention to cues associated with well-predicted reinforcers (e.g., Kaye & Pearce, 1984; Hogarth et al., 2008; Pearce & Hall, 1980; Pearce & Mackintosh, 2010; see Bouton, 2021, for more extended discussion). Consistent with this third account, discriminated operants become habitual when the discriminative stimulus occasions a reinforcer every time it is presented as opposed to an unpredictable half the time (Thrailkill et al., 2018, 2021)—the number of response–reinforcer pairings being equated. And a habit can return to goal-directed status if the outcome is made surprising, for example, by introducing a new reinforcer (Bouton et al., 2020). Neither type of result is predicted by the law of effect or the rate-correlation view.

Are habits unusually sticky and persistent?

Habits may seem persistent in the reinforcer-devaluation test in the sense that they fail to adjust when the reinforcer is separately devalued and the habit is then tested in extinction. But notice that the reinforcer-devaluation test is an analytical test designed to distinguish between habit and goal-direction in the laboratory; its crucial conditions (“offline” or response-independent devaluation of the reinforcer and then a test of the response without the reinforcer in extinction) are probably rare in the natural world. And as Hogarth (2018, 2020) has noted, when a habit is actually paired with the reinforcer after the reinforcer has been devalued (rather than tested in extinction), the animal readily stops performing the habitual response (e.g., Dickinson, 1985; Dickinson et al., 1983; Thrailkill & Bouton, 2015). Habits are not immune to modification when they are directly consequated this way. And it is difficult to know whether habits are slower than actions to adjust because a goal-directed action is already suppressed by devaluation before the action can be paired with the devalued reinforcer. However, a behavior may also be considered persistent if it is resistant to extinction, resistant to the effects of other contingency change, vulnerable to relapse after it has been eliminated, resistant to the weakening effects of context change, or enduringly expressed in behavior once it is learned. In the present article, I thus ask whether goal-directed actions and habits actually differ on these different dimensions.

Resistance to extinction

A first point to make is that two goal-directed actions can have different strengths and potential persistence. Conceptually, one can think of the hypothetical association between R and O as having strong or weak strength, as it might after a lot of exposure to the R-O contingency or just a little bit, respectively. If we make conventional assumptions about how R-O strength might change when the contingency between R and O changes, say, as it does in extinction, performance controlled by a stronger association will reach zero later than a performance controlled by a weaker association (e.g., Rescorla & Wagner, 1972). And if the reinforcer (O) is itself very highly valued, the action that yields it might be more persistent than one that yields a reinforcer that is less valued. Hogarth (e.g., 2018, 2020) has made such an argument about addictive behavior—it can be persistent because the goal’s value is very high and goal direction strong. A similar idea seems embedded in the concept of “reinforcer pathology” (e.g., Bickel et al., 2014); a reinforcer that has high demand and excessively high valuation may also create behavior that is resistant to change. It seems clear that a strong goal-directed action may be tougher to change than a weaker one.

Of course, learning theorists have long known that a behavior’s reinforcement history also massively influences behavioral persistence. In the partial-reinforcement extinction effect, a behavior that has been reinforced only some of the time turns out to be more resistant to extinction (and thus more persistent) than one that has been reinforced every time. The most successful explanations of the partial-reinforcement extinction effect (e.g., see Mackintosh, 1974, for review) have explained it by emphasizing that partial reinforcement reinforces the response under conditions that better resemble those of extinction. According to Capaldi’s sequential theory (1967, 1994), partial reinforcement reinforces responding while the organism remembers preceding nonreinforced trials—making responding persist when the animal encounters an extended run of nonreinforced trials. Similarly, according to Amsel’s frustration theory (Amsel, 1958, 1962), partial reinforcement reinforces responding when the organism is frustrated because of preceding nonreinforced trials—making responding persist when the organism encounters frustration in extinction. In either of these views, behavior will persist when the conditions of training approximate—and generalize well to—the conditions of testing. This powerful rule is presumably orthogonal to the response’s status as an action or habit. To illustrate, while we were discovering that habits develop primarily when discriminative cues make the reinforcer predictable, Thrailkill et al. (2018) obtained a paradoxical but relevant result: A partially reinforced action (that was sensitive to reinforcer devaluation) was slower to extinguish than a continuously reinforced habit (that was insensitive to reinforcer devaluation). Slower extinction of an action than a habit contradicts the notion that habit is more persistent. The partial-reinforcement extinction effect, which illustrates the principle that persistence is enhanced when you train in conditions that generalize to the test, was more relevant in predicting persistence than the habit–action distinction.

Was habit ever supposed to be more persistent in extinction? In none of our own experiments have habits been particularly slow to extinguish, and this may be a general rule (e.g., see Dickinson et al., 1995, 1998). The idea also seems a little paradoxical, because if habits are relatively context specific (e.g., Thrailkill & Bouton, 2015) and if a series of extinction trials generates context change (as the just-mentioned theories of Capaldi and Amsel suggest), then we might have grounds for thinking that they might be less persistent in extinction than goal-directed actions are. Perhaps such contextual change created by extinction is a cause of the overtraining extinction effect, in which overtrained responses can actually extinguish faster than responses that are more moderately trained (e.g., Ison, 1962; Siegel & Wagner, 1963; Wagner, 1963). On the other hand, evidence to be discussed below actually suggests that overtraining does not destroy a behavior’s initial status as action (e.g., Bouton, 2021) and that goal direction can be revealed in performance as the expression of the habit component is lost with context change (Steinfeld & Bouton, 2020, 2021). Assuming that an observed response can be controlled by both habit and goal direction, goal direction emerging in extinction as habit declines more quickly might allow responding in its general sense to seem persistent. Interestingly, there is at least one report of a habit showing emerging action properties (increased sensitivity to reinforcer devaluation) as extinction proceeds (Dezfouli et al., 2014).

Resistance to the effects of contingency degradation

Based on the results of Uhl (1973), Dickinson et al. (1998) noted that a behavior’s sensitivity to the suppressing effects of an omission schedule (in which responding now prevents or delays a scheduled noncontingent reward) might be a more sensitive measure of a behavior’s persistence than resistance to extinction. This led to early experiments suggesting contingency degradation as another measure of habit: habits are supposed to be more resistant than actions to the suppressive effects of adding an omission contingency. In two influential experiments (Dickinson et al., 1998), different groups of rats were overtrained or only moderately trained on two lever-press responses. Then a new reinforcer was added and presented freely every 30 s on average. However, making one of the two responses postponed the free reinforcer. This omission contingency caused more suppression of the omission response than the control response, but the suppression was only evident in the moderately trained group and not the overtrained group. Overtraining thus caused more persistence—that is, resistance to change in response to the new omission contingency. However, an important issue here is whether the overtrained response was actually a habit as defined here—there was no devaluation test to confirm that overtraining had created a habit. Recall that stronger goal-directed actions can in principle be more resistant to change than weaker ones. And the issue is important here because the method of concurrently training two responses in each group is one that actually appears to prevent the development of habit (e.g., Colwill & Rescorla, 1985a, 1988; Kosaki & Dickinson, 2010). Thus, the results could have been due to stronger goal direction, not habit, in the overtrained group.

The resistance-to-omission test is often considered an alternative index of habit (e.g., Dezfouli & Balleine, 2012; Schreiner et al., 2020). But the possibility that a strong action can also be slower to adjust to contingency change than a weaker one suggests that care must be taken in using it. I am aware of only one experiment showing that habits and actions explicitly confirmed as such by the results of a reinforcer-devaluation test differ in their sensitivity to an omission contingency. Specifically, in a theoretical paper, Dezfouli and Balleine (2012) reported an unpublished study in which different groups were overtrained or moderately trained on a single lever-press response and given a devaluation test that confirmed that overtraining had created habit. Then, after some retraining, the groups received the reinforcer noncontingently, but in one overtrained and one moderately trained group, the noncontingent reinforcers were delayed by the lever press. The overtrained response that was empirically established as a habit was indeed less affected by omission than the moderately trained one. But this sort of within-experiment habit confirmation is rare. We need more of it because, when considering the effects of contingency degradation on a response, it is best to remember that a strong goal-directed action could be more resistant to contingency change than a weaker one and thus be mistaken for a habit. Similarly, “compulsive” insensitivity to a positive punishment contingency in animals self-administering drugs can likewise have causes other than habit (e.g., Lüscher et al., 2020; McNally et al., 2023).

Vulnerability to “relapse”

Another sense of the persistence of a learned behavior is that it can return or relapse after the behavior has been eliminated by extinction, punishment, or omission training (e.g., Bouton, 2019). Phenomena like renewal, reinstatement, resurgence, and spontaneous recovery give another sense in which a learned behavior can seem sticky. Are habits more vulnerable than goal-directed behavior to these relapse effects? The question has not been addressed extensively. But Steinfeld and Bouton (2020) did test renewal after extinction of free-operant lever pressing when rats were first given a moderate amount of operant training, which created a goal-directed action as confirmed by reinforcer-devaluation tests, or more extensive training, which created a habit as confirmed by the same kind of test. Both actions and habits showed strong ABA renewal: When the responses were reinforced in Context A and then extinguished in Context B, they returned when the response was tested again in Context A. Moreover, the action renewed as an action (it was sensitive to reinforcer devaluation) and the habit renewed as habit (it was insensitive to reinforcer devaluation) at the time of the test. Actions and habits were also vulnerable to ABC renewal: When lever pressing was reinforced in Context A and extinguished in Context B, responding was renewed in a third (neutral, but equally familiar) context, Context C. But here both action-trained and habit-trained lever pressing renewed as actions; renewal of both was sensitive to reinforcer devaluation during the renewal test in Context C. The results are relevant to the present discussion for at least two reasons. First, goal-directed responding seemed no less vulnerable to renewal than habitual responding was. It was thus no less persistent. And second, if anything, goal direction was arguably more persistent than habit in the sense that goal direction returned and replaced habit in the renewal tests in Context C.1 Habit is generally attenuated when the context is changed, whereas goal-direction does not seem to be (see also Steinfeld & Bouton, 2021; Thrailkill & Bouton, 2015). Which is more persistent?

Resistance to the effects of context change

The fact that context change reduces habit more than goal direction does ironically suggests a sense in which a behavior’s goal-directed status takes priority: If a behavior is habitual in one context but generally goal directed in others, then goal direction, not habit, would seem to be the behavior’s default mode. And interestingly, if a learned behavior’s resistance to disruption by context change were taken as an indication of its persistence (cf. Schmidt & Bjork, 1992), then habit might even be considered less persistent than goal direction.

It is worth further considering how one can defeat the detrimental effect that context change has on operant responding generally (e.g., Bouton et al., 2011, 2014; Trask & Bouton, 2018; Trask et al., 2017). Theoretically, effects like the partial-reinforcement extinction effect work because the partial-reinforcement procedure enhances generalization to extinction, the conditions of testing in that paradigm, and thus make responding more persistent. In an analogous way, operant responding generalizes better to a new context if the response is trained in multiple contexts (e.g., Todd et al., 2012; Trask & Bouton, 2018). The treatment thus makes operant behavior more persistent in the sense that it is more resistant to the effects of context change. Several explanations of this effect are possible. In one, learning in multiple contexts increases the strength of the response (e.g., as explained for goal direction below; see also “retrieval strength,” e.g., Bjork & Bjork, 1992). In a second account, if we assume that every context is made up of a large set of stimuli or contextual elements, then training in multiple contexts would increase the number of contextual elements that gain control of the response and thus increase the odds that a new context might contain shared elements that have already gained control over the response (e.g., Estes, 1955; Trask & Bouton, 2018).

It is interesting to note that, at a theoretical level, training an operant response in multiple contexts could influence goal-direction and habit processes quite differently. First, consider the effects of multiple-context training on habit. If we accept the evidence that habit occurs when contemporary cues predict the reinforcer well (Thrailkill et al., 2018, 2020), multiple-context training will impede conversion to habit because each transfer to a new context will reduce the extent to which the reinforcer is predicted by contextual cues. Indeed, if one wants habit to develop rapidly, it would be best to reinforce the response consistently in the same context. Now consider the effects of multiple-context training on the acquisition of goal direction. Let us assume again that the strength of goal direction is represented by the strength of the R-O association, which theories of associative learning expect will grow as long as there is prediction error—that is, as long as the reinforcer is surprising (e.g., Rescorla & Wagner, 1972; Pearce, 1994; Wagner, 2008). Training methods that enhance prediction error will thus enhance goal-direction learning. As just implied, the context is a concurrent stimulus that also predicts the outcome (i.e., there is background S-O learning). There is evidence that R-O and S-O learning both occur together and compete with one another, as two simultaneous Pavlovian signals are expected to compete for S-O (e.g., Mackintosh & Dickinson, 1979; see Dickinson, 1980). Given this, if an operant behavior is trained in a series of new contexts, there will be less S-O present after each context switch and thus less competition with (and “blocking” of) R-O learning. Thus, training in multiple contexts will theoretically enhance goal-direction learning. The general point is that, compared with training in a single context, training in multiple contexts might facilitate goal-direction learning (and enhance generalization to new contexts, as above) at the same time it decreases habit learning. Again, a method that can increase persistence does not have much to do with habit.2

The theoretical effects of training in multiple contexts are probably more nuanced than this. If operant training extensively occurs in a repeating set of multiple contexts (e.g., ABCABCABC …), prediction error will eventually approach zero in each context. The method will thus allow habit to emerge eventually and perhaps become connected to more contextual elements, as above. Perhaps habit can be encouraged to generalize this way. The picture is incomplete and speculative. But the point, again, is that although training an operant behavior in multiple contexts can cause it to generalize better to new contexts, the principle does not necessarily have much to do with the behavior’s habitual or goal-directed status.

Habit conversion as a behavioral end point

In at least one recent model of goal direction and habit, habit acquisition replaces goal direction only when goal direction weakens as the experienced correlation between behavior rate and reinforcer rate declines (Perez & Dickinson, 2020). No memory of goal direction is assumed, and the habit that develops is thus expected to permanently replace it. Although it is rarely explicitly said this way, an emphasis on habits in addiction might also be taken to imply that the conversion of an action into a habit is sticky and difficult to change. However, Bouton (2021, in press) has summarized evidence that this is not true of habits that develop with food reinforcers. A behavior that has been a habit can revert to action status with several environmental manipulations, suggesting a more dynamic and flexible view.

As a further test of the findings suggesting that habit develops when environmental cues make the reinforcer predictable (Thrailkill et al., 2018, 2021), we asked whether making the reinforcer surprising again might return a habit back to action (Bouton et al., 2020). The prediction was based on earlier work in Pavlovian learning suggesting that reduced attention to a CS that has perfectly predicted its outcome can return if the outcome is made surprising (e.g., Kaye & Pearce, 1984; Hall & Pearce, 1982). Consistent with the earlier research, when an extensively trained lever-pressing response that was demonstrably a habit unexpectedly produced a new outcome in a final training session (we switched grain pellet to sucrose pellet or vice versa), the response became sensitive to the reinforcer-devaluation effect and thus made the habit goal directed again (Bouton et al., 2020). Relatedly, when the test was preceded by unexpected prefeeding with an irrelevant food pellet, a habit also became goal directed again. These findings converged on other work in which we found that, after extensive training with one response, either intermixed training with a new response in a different context or merely presenting the reinforcer freely there could also turn the habit back to action (again as judged by reinforcer-devaluation tests; Trask et al., 2020; see also Shipman et al., 2018). As a group, the results suggested that, consistent with the attentional view, habits can return to action status with a number of manipulations that entailed exposure to unexpected reinforcers. The original action knowledge had not been erased when the response became a habit.

Other work suggests that the context can control a behavior’s status as goal directed or habitual. I already mentioned that when Steinfeld and Bouton (2020) trained either goal-directed or habitual lever pressing (with moderate or extended training on an RI 30-s schedule) in Context A and then extinguished the response in Context B, responding renewed as goal directed (sensitive to reinforcer devaluation) when responding was tested in Context C. This was additional evidence that habits do not lose their potential action status, which might be renewed with context change. To put the idea to a more direct test, Steinfeld and Bouton (2021) gave rats a modest amount of lever-press training in Context A (to create goal direction there) before giving the response additional and more extended training in Context B (to turn it into a habit there). When the response was then tested in A and B (in a counterbalanced order), the behavior was habitual in Context B but goal directed in Context A. And when, after the same moderate and extensive training in A then B, we tested the response Context B and Context C (an equally familiar context, but one where the lever press had never been trained), the response was again habitual in the habit context (B) but was goal directed in the new context. The most straightforward explanation is that after a goal-directed behavior becomes a habit, goal direction remains available but is expressed according to the context. Goal direction transfers across contexts (e.g., from Context A to C), whereas habit does not transfer out of B. Goal direction can return to performance with a context change; with our methods, the conversion to habit is clearly not permanent.

The fact that a behavior can seem to switch between habit and action so readily is not anticipated by the idea that habit is a behavioral end point (e.g., Perez & Dickinson, 2020). But our behavioral findings are consistent with those of other research. Most relevant is the underlying behavioral neuroscience research. Goal direction is thought to be controlled by the dorsomedial striatum and the prelimbic cortex, whereas habit is controlled by the dorsolateral striatum and infralimbic cortex (e.g., Balleine, 2019; Balleine & O’Doherty, 2010). And crucially, suppressing habit by inactivating either infralimbic cortex or the dorsolateral striatum can return a habit to goal-directed status (e.g., Coutureau & Killcross, 2003; Yin et al., 2006)—just as our behavioral manipulations did. It should be acknowledged that behavioral theorists have correspondingly posited separate goal-direction and habit memory systems (e.g., de Wit & Dickinson, 2009), with an unexplored potential to activate them selectively. And reinforcement learning theorists have generated models that imagine activation of goal-directed (model based) and habitual (model free) systems based on different “arbitration” criteria (e.g., Daw et al., 2005; Lee et al., 2014; Dezfouli et al., 2014).

Bouton (2021) emphasized the parallel between the conversion of a response from goal directed to habitual and other retroactive interference paradigms, such as extinction, where a new performance replaces an older one. Decades of research on extinction (e.g., Bouton, Maren, et al., 2021) suggest that extinction does not erase the original learning but instead suppresses it in a context-specific way. The same appears true of habit learning. Bouton (2021) noted that a simple contextual “switch” analogous to the one thought to control the expression of extinction may also control the expression of habit. As just noted, habit acquisition does not destroy the original action learning. In the context of habit, habit inhibits goal direction. But in other contexts (which may be defined quite broadly, e.g., see Bouton, 2019), habit is lost and goal-direction—freed from inhibition by habit—is expressed again. This is a simple but sensible way for the habit/goal-direction system to be organized: in conditions where habit has paid off, behavior is habitual and exploits what works in that environment, but in other contexts, the relative flexibility of goal direction is intact.

CONCLUSION

This selective review has uncovered little evidence to support the idea that habits are unusually strong or difficult to change. As I noted earlier, it is true that habits are more persistent in the sense that they do not adjust when the reinforcer’s value separately changes—as we know from the reinforcer-devaluation test—but they adjust when the averted outcome is made contingent on the behavior. And there is little reason to think that habits are necessarily more resistant than goal-directed behaviors to extinction or other types of contingency change. Although the evidence is sparse, both habits and actions are vulnerable to renewal effects. And habit is more affected, not less affected, by changing the context. The expression of habit is not necessarily permanent; habits can switch back into action status with a number of manipulations. Finally, even classic methods that appear to make operant behaviors more persistent (partially reinforcing them as a hedge against extinction or omission, Thrailkill, 2023, or training them in multiple contexts as a hedge against context change) do not seem to operate by encouraging habit. It would be a mistake to think that habits as defined by insensitivity to reinforcer devaluation are inherently stronger, more durable, or more persistent. Goal-directed actions can be persistent, too.

This conclusion reopens the question of what makes habit adaptive and functional. Based on the above, habit learning may not be a mechanism that functions to make behavior strong or durable. Instead, it is a process that makes a behavior that is effective in a specific setting efficient and automatic. These attributes of habitual performance are just as important as “persistence.” Thus, the practitioner should encourage the development of habit and automaticity in an effort to maintain healthy behavior, such as regular exercise or diabetes self-management (e.g., Cummings et al., 2021). And the efficiency of habits—for example, by virtue of their gluing or chunking different behaviors together (e.g., Dezfouli & Balleine, 2012; Graybiel, 2008; Smith et al., 2012; Turner et al., 2022)—makes them easier to deploy rapidly. The overall picture is that habits are efficient and automatic behaviors that exploit or capitalize on reinforcer regularities in certain contexts. They may need no more functional value than that.

ACKNOWLEDGMENTS

I thank John Green, Mike Steinfeld, Eric Thrailkill, and Travis Todd for comments.

FUNDING INFORMATION

Much of the research discussed here was supported by Grant DA RO1 033123 from the National Institutes of Health.

Footnotes

CONFLICT OF INTEREST STATEMENT

The author declares no conflicts of interest.

ETHICS APPROVAL

No human or animal subjects were used to produce this article.

1

In this issue, Fujimaki et al. (2023) report parallel results in the resurgence paradigm: When a goal-directed action or a habit was extinguished and a new response was reinforced to replace it, they both “resurged” when the replacement response was extinguished. But just as in ABC renewal (Steinfeld & Bouton, 2020), action and habit both resurged in the form of a goal-directed action—during resurgence testing, both recovered responses were sensitive to reinforcer devaluation.

2

Interestingly, a model-based reinforcement-learning algorithm (e.g., see Daw et al., 2005) might make a different prediction here. Essentially, every context switch might create a new “state” the organism needs to learn about, slowing down the acquisition of a model-based representation of goal-directed action. However, the finding that goal direction is context independent (e.g., Steinfeld & Bouton, 2020, 2021; Thrailkill & Bouton, 2015) suggests that a new physical context is not necessarily a new state in this sense.

REFERENCES

  1. Adams CD (1982). Variations in the sensitivity of instrumental responding to reinforcer devaluation. The Quarterly Journal of Experimental Psychology B: Comparative and Physiological Psychology, 34B(2), 77–98. 10.1080/14640748208400878 [DOI] [Google Scholar]
  2. Amsel A (1958). The role of frustrative nonreward in noncontinuous reward situations. Psychological Bulletin, 55(2), 102–119. 10.1037/h0043125 [DOI] [PubMed] [Google Scholar]
  3. Amsel A (1962). Frustrative nonreward in partial reinforcement and discrimination learning: Some recent history and a theoretical extension. Psychological Review, 69(4), 306–328. 10.1037/h0046200 [DOI] [PubMed] [Google Scholar]
  4. Balleine BW (2019). The meaning of behavior: Discriminating reflex and volition in the brain. Neuron, 104(1), 47–62. 10.1016/j.neuron.2019.09.024 [DOI] [PubMed] [Google Scholar]
  5. Balleine BW, & O’Doherty JP (2010). Human and rodent homologies in action control: corticostriatal determinants of goal-directed and habitual action. Neuropsychopharmacology, 35(1), 48–69. 10.1038/npp.2009.131 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Baum WM (1973). The correlation-based law of effect. Journal of the Experimental Analysis of Behavior, 20(1), 137–153. 10.1901/jeab.1973.20-137 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Becchi S, Hood J, Kendig MD, Mohammadkhani A, Shipman ML, Balleine BW, Borgland SL, & Corbit LH (2022). Food for thought: diet-induced impairments to decision-making and amelioration by N-acetylcysteine in male rats. Psychopharmacology, 239(11), 3495–3506. 10.1007/s00213-022-06223-4 [DOI] [PubMed] [Google Scholar]
  8. Bickel WK, Johnson MW, Koffarnus MN, MacKillop J, & Murphy JG (2014). The behavioral economics of substance use disorders: reinforcement pathologies and their repair. Annual Review of Clinical Psychology, 10, 641–677. 10.1146/annurev-clinpsy-032813-153724 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Bjork RA, & Bjork EL (1992). A new theory of disuse and old theory of stimulus fluctuation. In Healy A, Kosslyn S, & Shiffrin R (Eds.), From learning processes to cognitive processes: Essays in honor of William K. Estes. Erlbaum. [Google Scholar]
  10. Bouton ME (2019). Extinction of instrumental (operant) learning: interference, varieties of context, and mechanisms of contextual control. Psychopharmacology, 236(1), 7–19. 10.1007/s00213-018-5076-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Bouton ME (2021). Context, attention, and the switch between habit and goal-direction in behavior. Learning & Behavior, 49(4), 349–362. 10.3758/s13420-021-00488-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Bouton ME (in press). Situating habit and goal direction in a general view of instrumental behavior. In Vandaele Y (Ed.), Habits: Their definition, neurobiology, and role in addiction. Springer Nature; Switzerland. [Google Scholar]
  13. Bouton ME, Allan SM, Tavakkoli A, Steinfeld MR, & Thrailkill EA (2021). Effect of context on the instrumental reinforcer devaluation effect produced by taste-aversion learning. Journal of Experimental Psychology: Animal Learning and Cognition, 47(4), 476–489. 10.1037/xan0000295 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Bouton ME, & Balleine BW (2019). Prediction and control of operant behavior: What you see is not all there is. Behavior Analysis: Research and Practice, 19(2), 202–212. 10.1037/bar0000108 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Bouton ME, Maren S, & McNally GP (2021). Behavioral and neurobiological mechanisms of Pavlovian and instrumental extinction learning. Physiological Reviews, 101(2), 611–681. 10.1152/physrev.00016.2020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Bouton ME, Broomer MC, Rey CN, & Thrailkill EA (2020). Unexpected food outcomes can return a habit to goal-directed action. Neurobiology of Learning and Memory, 169, Article 107163. 10.1016/j.nlm.2020.107163 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Bouton ME, Todd TP, & León SP (2014). Contextual control of discriminated operant behavior. Journal of Experimental Psychology: Animal Learning and Cognition, 40(1), 92–105. 10.1037/xan0000002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Bouton ME, Todd TP, Vurbic D, & Winterbauer NE (2011). Renewal after the extinction of free operant behavior. Learning & Behavior, 39(1), 57–67. 10.3758/s13420-011-0018-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Capaldi EJ (1967). A sequential hypothesis of instrumental learning. In Spence KW & Spence JT (Eds.), The psychology of learning and motivation: 1 (pp. 1–65). Academic Press. [Google Scholar]
  20. Capaldi EJ (1994). The sequential view: From rapidly fading stimulus traces to the organization of memory and the abstract concept of number. Psychonomic Bulletin & Review, 1(2), 156–181. 10.3758/BF03200771 [DOI] [PubMed] [Google Scholar]
  21. Colwill RM, & Rescorla RA (1985a). Postconditioning devaluation of a reinforcer affects instrumental responding. Journal of Experimental Psychology: Animal Behavior Processes, 11(1), 120–132. 10.1037/0097-7403.11.1.120 [DOI] [PubMed] [Google Scholar]
  22. Colwill RM, & Rescorla RA (1985b). Instrumental responding remains sensitive to reinforcer devaluation after extensive training. Journal of Experimental Psychology: Animal Behavior Processes, 11(4), 520–536. 10.1037/0097-7403.11.4.520 [DOI] [Google Scholar]
  23. Colwill RM, & Rescorla RA (1988). The role of response-reinforcer associations increases throughout extended instrumental training. Animal Learning & Behavior, 16(1), 105–111. 10.3758/BF03209051 [DOI] [Google Scholar]
  24. Corbit LH, Chieng BC, & Balleine BW (2014). Effects of repeated cocaine exposure on habit learning and reversal by N-acetylcysteine. Neuropsychopharmacology, 39(8), 1893–1901. 10.1038/npp.2014.37 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Corbit LH, Nie H, & Janak PH (2012). Habitual alcohol seeking: Time course and the contribution of subregions of the dorsal striatum. Biological Psychiatry, 72(5), 389–395. 10.1016/j.biopsych.2012.02.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Coutureau E, & Killcross S (2003). Inactivation of the infralimbic prefrontal cortex reinstates goal-directed responding in overtrained rats. Behavioural Brain Research, 146(1–2), 167–174. 10.1016/j.bbr.2003.09.025 [DOI] [PubMed] [Google Scholar]
  27. Crespi LP (1942). Quantitative variation of incentive and performance in the white rat. The American Journal of Psychology, 55, 467–517. 10.2307/1417120 [DOI] [Google Scholar]
  28. Cummings C, Benjamin NE, Prabhu HY, Cohen LB, Goddard BJ, Kaugars AS, Humiston T, & Lansing AH (2022). Habit and diabetes self-management in adolescents with Type 1 diabetes. Health Psychology, 41(1), 13–22. 10.1037/hea0001097 [DOI] [PubMed] [Google Scholar]
  29. Daw N, Niv Y & Dayan P (20050. Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control. Nature Neuroscience, 8, 1704–1711, 10.1038/nn1560 [DOI] [PubMed] [Google Scholar]
  30. de Wit S, & Dickinson A (2009). Associative theories of goal-directed behaviour: A case for animal-human translational models. Psychological Research, 73(4), 463–476. 10.1007/s00426-009-0230-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Dezfouli A, & Balleine BW (2012). Habits, action sequences and reinforcement learning. The European Journal of Neuroscience, 35(7), 1036–1051. 10.1111/j.1460-9568.2012.08050.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Dezfouli A, Lingawi NW, & Balleine BW (2014). Habits as action sequences: Hierarchical action control and changes in outcome value. Philosophical Transactions of the Royal Society B: Biological Sciences, 369(1655), Article 20130482. 10.1098/rstb.2013.0482 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Dickinson A (1980). Contemporary animal learning theory. Cambridge University Press. [Google Scholar]
  34. Dickinson A (1985). Actions and habits: The development of behavioural autonomy. Philosophical Transactions of the Royal Society B: Biological Sciences, 308, 67–78. 10.1098/rstb.1985.0010 [DOI] [Google Scholar]
  35. Dickinson A (1989). Expectancy theory in animal conditioning. In Klein SB & Mowrer RR (Eds.), Contemporary learning theories: Pavlovian conditioning and the status of traditional learning theory (pp. 279–308). Lawrence Erlbaum Associates. [Google Scholar]
  36. Dickinson A, Balleine B, Watt A, Gonzalez F, & Boakes RA (1995). Motivational control after extended instrumental training. Animal Learning & Behavior, 23(2), 197–206. https://psycnet.apa.org/doi/10.3758/BF03199935 [Google Scholar]
  37. Dickinson A, Nicholas DJ, & Adams CD (1983). The effect of the instrumental training contingency on susceptibility to reinforcer devaluation. The Quarterly Journal of Experimental Psychology B: Comparative and Physiological Psychology, 35(B–1), 35–51. 10.1080/14640748308400912 [DOI] [Google Scholar]
  38. Dickinson A, Squire S, Varga Z, & Smith JW (1998). Omission learning after instrumental pretraining. The Quarterly Journal of Experimental Psychology B: Comparative and Physiological Psychology, 51B(3), 271–286. 10.1080/713932679 [DOI] [Google Scholar]
  39. Estes WK (1955). Statistical theory of spontaneous recovery and regression. Psychological Review, 62(3), 145–154. 10.1037/h0048509 [DOI] [PubMed] [Google Scholar]
  40. Flaherty CF (1996). Incentive relativity. Cambridge University Press. [Google Scholar]
  41. Fujimaki S, Hu T, & Kosaki Y (2023). Resurgence of goal-directed actions and habits. Journal of the Experimental Analysis of Behavior. Advance online publication. 10.1002/jeab.884 [DOI] [PubMed] [Google Scholar]
  42. Furlong TM, Jayaweera HK, Balleine BW, & Corbit LH (2014). Binge-like consumption of a palatable food accelerates habitual control of behavior and is dependent on activation of the dorsolateral striatum. The Journal of Neuroscience, 34(14), 5012–5022. 10.1523/JNEUROSCI.3707-13.2014 [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Graybiel AM (2008). Habits, rituals, and the evaluative brain. Annual Review of Neuroscience, 31, 359–387. 10.1146/annurev.neuro.29.051605.112851 [DOI] [PubMed] [Google Scholar]
  44. Hall G, & Pearce JM (1982). Restoring the associability of a pre-exposed CS by a surprising event. The Quarterly Journal of Experimental Psychology B: Comparative and Physiological Psychology, 34(3b), 127–140. 10.1080/14640748208400881 [DOI] [Google Scholar]
  45. Hogarth L (2018). A critical review of habit theory of drug dependence. In Verplanken B (Ed.), The psychology of habit: Theory, mechanisms, change, and contexts (pp. 325–342). Springer. [Google Scholar]
  46. Hogarth L (2020). Addiction is driven by excessive goal-directed drug choice under negative affect: Translational critique of habit and compulsion theory. Neuropsychopharmacology, 45(5), 720–735. 10.1038/s41386-020-0600-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Hogarth L, Dickinson A, Austin A, Brown C, & Duka T (2008). Attention and expectation in human predictive learning: The role of uncertainty. The Quarterly Journal of Experimental Psychology, 61(11), 1658–1668. 10.1080/17470210701643439 [DOI] [PubMed] [Google Scholar]
  48. Ison JR (1962). Experimental extinction as a function of number of reinforcements. Journal of Experimental Psychology, 64(3), 314–317. 10.1037/h0042858 [DOI] [Google Scholar]
  49. James W (1890). The principles of psychology. Henry Holt & Co. [Google Scholar]
  50. Kaye H, & Pearce JM (1984). The strength of the orienting response during Pavlovian conditioning. Journal of Experimental Psychology: Animal Behavior Processes, 10(1), 90–109. 10.1037/0097-7403.10.1.90 [DOI] [PubMed] [Google Scholar]
  51. Kosaki Y, & Dickinson A (2010). Choice and contingency in the development of behavioral autonomy during instrumental conditioning. Journal of Experimental Psychology: Animal Behavior Processes, 36(3), 334–342. 10.1037/a0016887 [DOI] [PubMed] [Google Scholar]
  52. Lee SW, Shimojo S, & O’Doherty JP (2014). Neural computations underlying arbitration between model-based and model-free learning. Neuron, 81(3), 687–699. 10.1016/j.neuron.2013.11.028 [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Lüscher C, Robbins TW, & Everitt BJ (2020). The transition to compulsion in addiction. Nature Reviews Neuroscience, 21(5), 247–263. 10.1038/s41583-020-0289-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Mackintosh NJ (1974). The psychology of animal learning. Academic Press. [Google Scholar]
  55. Mackintosh NJ, & Dickinson A (1979). Instrumental (Type II) conditioning. In Dickinson A & Boakes RA (Eds.), Mechanisms of learning and motivation: A memorial volume to Jerzy Konorski (pp. 143–169). Erlbaum. [Google Scholar]
  56. McNally GP, Jean-Richard-Dit-Bressel P, Millan EZ, & Lawrence AJ (2023). Pathways to the persistence of drug use despite its adverse consequences. Molecular Psychiatry, 28(6), 2228–2237. 10.1038/s41380-023-02040-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Nelson A, & Killcross S (2006). Amphetamine exposure enhances habit formation. The Journal of Neuroscience, 26(14), 3805–3812. 10.1523/JNEUROSCI.4305-05.2006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Pearce JM (1994). Similarity and discrimination: A selective review and a connectionist model. Psychological Review, 101(4), 587–607. 10.1037/0033-295X.101.4.587 [DOI] [PubMed] [Google Scholar]
  59. Pearce JM, & Hall G (1980). A model for Pavlovian learning: Variations in the effectiveness of conditioned but not of unconditioned stimuli. Psychological Review, 87(6), 532–552. 10.1037/0033-295X.87.6.532 [DOI] [PubMed] [Google Scholar]
  60. Pearce JM, & Mackintosh NJ (2010). Two theories of attention: A review and a possible integration. In Mitchell CJ & Le Pelley ME (Eds.), Attention and associative learning (pp. 11–39). Oxford University Press. [Google Scholar]
  61. Perez OD, & Dickinson A (2020). A theory of actions and habits: The interaction of rate correlation and contiguity systems in free-operant behavior. Psychological Review, 127(6), 945–971. 10.1037/rev0000201 [DOI] [PubMed] [Google Scholar]
  62. Poldrack R (2021). Hard to break: Why our brains make habits stick. Princeton University Press. [Google Scholar]
  63. Rescorla RA, & Wagner AR (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In Black AH & Prokasy WF (Eds.), Classical conditioning II: Current research and theory (pp 64–99). Appleton-Century-Crofts. [Google Scholar]
  64. Schmidt RA, & Bjork RA (1992). New conceptualizations of practice: Common principles in three paradigms suggest new concepts for training. Psychological Science, 3, 207–217. 10.1111/j.1467-9280.1992.tb00029.x [DOI] [Google Scholar]
  65. Schreiner DC, Renteria R, & Gremel CM (2020). Fractionating the all-or-nothing definition of goal-directed and habitual decision-making. Journal of Neuroscience Research, 98(6), 998–1006. 10.1002/jnr.24545 [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Shipman ML, Corbit LH (2022). Diet-induced deficits in goal-directed control are rescued by agonism of group II metabotropic glutamate receptors in the dorsomedial striatum. Translational Psychiatry, 12, Article 42. 10.1038/s41398-022-01807-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Shipman ML, Trask S, Bouton ME, & Green JT (2018). Inactivation of prelimbic and infralimbic cortex respectively affects minimally-trained and extensively-trained goal-directed actions. Neurobiology of Learning and Memory, 155, 164–172. 10.1016/j.nlm.2018.07.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  68. Siegel S, & Wagner AR (1963). Extended acquisition training and resistance to extinction. Journal of Experimental Psychology, 66(3), 308–310. 10.1037/h0041325 [DOI] [PubMed] [Google Scholar]
  69. Smith KS, Virkud A, Deisseroth K, & Graybiel AM (2012). Reversible online control of habitual behavior by optogenetic perturbation of medial prefrontal cortex. Proceedings of the National Academy of Sciences, 109(46), 18932–18937. 10.1073/pnas.1216264109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Steinfeld MR, & Bouton ME (2020). Context and renewal of habits and goal-directed actions after extinction. Journal of Experimental Psychology: Animal Learning and Cognition, 46(4), 408–421. 10.1037/xan0000247 [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Steinfeld MR, & Bouton ME (2021). Renewal of goal direction with a context change after habit learning. Behavioral Neuroscience, 135(1), 79–87. 10.1037/bne0000422 [DOI] [PMC free article] [PubMed] [Google Scholar]
  72. Thorndike EL (1911). Animal Intelligence. Macmillan. [Google Scholar]
  73. Thrailkill EA (2023). Partial reinforcement extinction and omission effects in the elimination and recovery of discriminated operant behavior. Journal of Experimental Psychology: Animal Learning and Cognition, 49(3), 194–207. 10.1037/xan0000354 [DOI] [PMC free article] [PubMed] [Google Scholar]
  74. Thrailkill EA, & Bouton ME (2015). Contextual control of instrumental actions and habits. Journal of Experimental Psychology: Animal Learning and Cognition, 41(1), 69–80. 10.1037/xan0000045 [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Thrailkill EA, Michaud NL, & Bouton ME (2021). Reinforcer predictability and stimulus salience promote discriminated habit learning. Journal of Experimental Psychology: Animal Learning and Cognition, 47(2), 183–199. 10.1037/xan0000285 [DOI] [PMC free article] [PubMed] [Google Scholar]
  76. Thrailkill EA, Trask S, Vidal P, Alcalá JA, & Bouton ME (2018). Stimulus control of actions and habits: A role for reinforcer predictability and attention in the development of habitual behavior. Journal of Experimental Psychology: Animal Learning and Cognition, 44(4), 370–384. 10.1037/xan0000188 [DOI] [PMC free article] [PubMed] [Google Scholar]
  77. Tinklepaugh OL (1928). An experimental study of representative factors in monkeys. Journal of Comparative Psychology, 8(3), 197–236. 10.1037/h0075798 [DOI] [Google Scholar]
  78. Todd TP, Winterbauer NE, & Bouton ME (2012). Effects of the amount of acquisition and contextual generalization on the renewal of instrumental behavior after extinction. Learning & Behavior, 40(2), 145–157. 10.3758/s13420-011-0051-5 [DOI] [PubMed] [Google Scholar]
  79. Tolman EC, & Honzik CH (1930). Introduction and removal of reward, and maze performance in rats. University of California Publications in Psychology, 4, 257–275. [Google Scholar]
  80. Trask S, & Bouton ME (2018). Retrieval practice after multiple context changes, but not long retention intervals, reduces the impact of a final context change on instrumental behavior. Learning & Behavior, 46(2), 213–221. 10.3758/s13420-017-0304-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  81. Trask S, Shipman ML, Green JT, & Bouton ME (2017). Inactivation of the prelimbic cortex attenuates context-dependent operant responding. The Journal of Neuroscience, 37(9), 2317–2324. 10.1523/JNEUROSCI.3361-16.2017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  82. Trask S, Shipman ML, Green JT, & Bouton ME (2020). Some factors that restore goal-direction to a habitual behavior. Neurobiology of Learning and Memory, 169, Article 107161. 10.1016/j.nlm.2020.107161 [DOI] [PMC free article] [PubMed] [Google Scholar]
  83. Turner KM, Svegborn A, Langguth M, McKenzie C, & Robbins TW (2022). Opposing Roles of the dorsolateral and dorsomedial striatum in the acquisition of skilled action sequencing in rats. The Journal of Neuroscience, 42(10), 2039–2051. 10.1523/JNEUROSCI.1907-21.2022 [DOI] [PMC free article] [PubMed] [Google Scholar]
  84. Uhl CN (1973). Eliminating behavior with omission and extinction after varying amounts of training. Animal Learning & Behavior, 1(3), 237–240. 10.3758/BF03199082 [DOI] [Google Scholar]
  85. Wagner AR (1963). Overtraining and frustration. Psychological Reports, 13(3), 717–718. 10.2466/pr0.1963.13.3.717 [DOI] [Google Scholar]
  86. Wagner AR (2008). Evolution of an elemental theory of Pavlovian conditioning. Learning & Behavior, 36(3), 253–265. 10.3758/LB.36.3.253 [DOI] [PubMed] [Google Scholar]
  87. Wood W (2019). Good habits, bad habits: The science of making positive changes that stick. Farrar, Straus, and Giroux. [Google Scholar]
  88. Yin HH, Knowlton BJ, & Balleine BW (2006). Inactivation of dorsolateral striatum enhances sensitivity to changes in the action-outcome contingency in instrumental conditioning. Behavioural Brain Research, 166(2), 189–196. 10.1016/j.bbr.2005.07.012 [DOI] [PubMed] [Google Scholar]

RESOURCES