Abstract
Three experiments explored how training reinforcement schedule and context influence the elimination and recovery of human operant behavior. In Experiment 1, participants learned a discriminated operant response in Context A before the response was eliminated with extinction in Context B. They then received a final test in each context. Groups were trained with a discriminative stimulus that predicted a reinforced response on either every trial (Continuous Reinforcement; CRF) or some of the trials (Partial Reinforcement; PRF). Extinction was slower following PRF training (a partial reinforcement extinction effect; PREE) and extinguished responding increased when tested in Context A (ABA renewal). Experiment 2 further confirmed the PREE was obtained equally whether extinction occurred in the training context (Context A) or a new context (Context B) which is consistent with trial-based accounts of the PREE. Experiment 3 used the same design as Experiment 1 to evaluate the influence of training reinforcement on response elimination with an omission contingency. Across the omission training phase in Context B, the decrease in responding occurred more slowly in the PRF-trained group in comparison to the CRF-trained group, perhaps the first demonstration of what might be termed a partial reinforcement omission effect. Again, ABA renewal was observed in Context A. Training reinforcement schedule therefore had a similar influence on response elimination with extinction and omission. Elimination and recovery of human instrumental behavior, with extinction or omission, are influenced by training reinforcement schedule and context.
Keywords: Discriminated operant, extinction, human, omission, online, partial reinforcement
Few human laboratory studies connect to factors known to influence the elimination, persistence, and recovery of instrumental behavior in nonhuman animals. Instrumental behavior can be reduced or eliminated with at least three methods. First, extinction refers to the removal of the reinforcing outcome that followed the response during training and the corresponding reduction in response frequency over time (Bouton, 2019). Second, omission training, sometimes referred to as negative punishment or differential reinforcement of other behavior, arranges presentation of the reinforcing outcome contingent upon a period of nonoccurrence, or withholding, of the instrumental response (e.g., Uhl & Garcia, 1969). The present study is focused on extinction and omission, but positive punishment also deserves mention as a third response elimination method and refers to the added presentation of an aversive outcome following the response (Bouton & Schepers, 2015; Church, 1963, 1969; Jean-Richard-dit-Bressel et al., 2021). Each response elimination method is effective for reducing the instrumental behavior of nonhuman and human animals alike and may involve common learning mechanisms (Bouton, 2019; Bouton & Broomer, 2023).
The response elimination methods just described do not destroy the original learning (Bouton et al., 2021). Instead, the procedures result in new learning that interferes with the original learning and depends on the response elimination context for expression. For example, Nakajima et al. (2002) trained rats to press a lever to earn a food-pellet reinforcer in a distinct context (Context A) before eliminating the response by introducing an omission contingency in a second context (Context B). Contexts were distinguished by unique olfactory, visual, and tactile cues. The eliminated response increased (renewed) when rats were tested in the training context, Context A. Importantly, renewal was observed even though the omission contingency continued to operate during the test in Context A.
Renewal of free-operant responding has been found after elimination with extinction and punishment in nonhumans and humans (Bouton et al., 2011; Bouton & Schepers, 2015; Ritchey et al., 2021a). Recent human laboratory studies have addressed renewal and other response recovery phenomena after omission training (Vila et al., 2020, 2022). Of note is a recent report by Finch and colleagues (Finch et al., 2022), who examined renewal of free-operant responding after omission training in the training context (ABA renewal) as well as in a new context (Context C). Each form of renewal (ABA and ABC) was observed even though the omission contingency continued to operate during the test. This finding is consistent with a theoretical perspective that emphasizes the role of interference; that is, instrumental response elimination involves learning to inhibit the response which interferes with, but does not necessarily weaken, the original excitatory learning supporting the response (Bouton, 2019; Bouton et al., 2016). More specifically, ABC renewal suggests that changing the context removes a source of response inhibition that is specific to the response elimination context.
Although evidence connects response elimination and recovery in humans and nonhumans, less is known about the generality of factors that influence response elimination. Nonhuman animal studies provide particularly strong support for the role of the pre-elimination schedule of reinforcement in the persistence of instrumental behavior during response elimination (Capaldi, 1967, 1994; Mackintosh, 1974; Thrailkill et al., 2016, 2018). The partial-reinforcement extinction effect (PREE) refers to the finding that response elimination during extinction occurs more slowly if the response was trained with a partial reinforcement (PRF) schedule (i.e., the response or trial is inconsistently followed by the reinforcing outcome) than if it were trained under a continuous reinforcement (CRF) schedule (i.e., each response or trial is consistently followed by the reinforcer). Several theories suggest that the PREE is partly the result of the greater generalization decrement experienced when extinction is introduced after CRF compared to PRF training (Amsel, 1962, 1992; Capaldi, 1967, 1994; Harris, 2019; Mackintosh, 1974). That is, experience with nonreinforced (N) responses (or trials) during PRF training makes extinction more like the training conditions. However, theoretical perspectives differ on why PRF training increases generalization. Amsel (1962, 1992) suggested that an emotional response evoked by N trials, frustration, is present when the response is subsequently reinforced (R trials) during PRF training. Frustration present with reinforced responding on R trials allows N trials to energize responding during extinction. Capaldi (1967, 1994) alternatively suggested that recent N trials are present in memory on the subsequent R trials during PRF training. Here, N trials acquire an excitatory association with the reinforced response. Amsel and Capaldi’s accounts of PREE were developed with instrumental responses and tend to emphasize excitatory discriminative stimulus control over the response. Such accounts can be contrasted with an account that suggest the PREE results from a comparison of rates of reinforcement in the trial and in the background (Gallistel & Gibbon, 2000). Extinction requires the organism to detect a change in reinforcement rate in the trial. After PRF training, the organism must experience more trials to detect the change because a lower per-trial reinforcement rate requires more nonreinforced trials to detect.
Based on more recent studies of Pavlovian extinction, a third account of the PREE has been proposed that differs from the accounts just described. Like Capaldi, Harris (2019) suggested that organisms learn about R and N trials, and that PRF establishes a discriminative relationship between nonreinforcement and subsequent reinforcement. In contrast to earlier accounts that suggested N trials have nonspecific effects on arousal and memory which may generalize to more than one stimulus (Amsel, 1962; Capaldi, 1967), Harris (2019) suggested that PRF training results in organisms learning that the specific stimulus predicts reinforcement and nonreinforcement. Evidence for this view comes from studies that arranged within-subject comparisons of appetitive conditioning to multiple PRF stimuli (Chan & Harris, 2019). One important piece of evidence supporting stimulus-specific encoding of N trials is that a greater number of N trials does not result in a greater post-trial expectation of the reinforcer (Harris & Kwok, 2018). Encoding of N trials as specific to the PRF-trained stimulus and separate from the CRF-trained stimulus can also account for instances of within-subject PREE (Harris, 2019; see also Bouton, Woods, & Todd, 2014).
While the PREE is well-established in nonhuman Pavlovian and instrumental learning, few studies have examined the PREE in human participants. One recent study reported an instrumental PREE with children after free-operant training (ages 8–12; De Meyer et al., 2019; see also Witte, 1977). The present experiments were conducted to strengthen connections between processes known to influence instrumental response elimination, persistence, and recovery in nonhumans and humans. A discriminated operant method was developed to allow 1) additional measures of behavioral control by the experimental contingencies, and 2) connect to everyday behaviors which tend to be under some form of antecedent stimulus control (Skinner, 1953). The designs are sketched in Table 1. Experiment 1 addressed the influence of PRF training on the extinction of a discriminated operant response with human participants in in-person and online settings. Experiment 2 directly examined the effect of context on the instrumental PREE. Experiment 3 further extended the method to assess the effect of PRF training on response elimination with omission training. Experiments 1 and 3 assessed recovery of the eliminated response in an ABA renewal design in which response training and elimination took place in distinct contexts (A and B) and the response was then tested in each context (Bouton et al., 2011). Of added interest for theories of the PREE is that no prior investigations have addressed the generality of learning in PRF to response elimination with omission training. The results provide insight to the generality and potential mechanisms of reinforcement schedule and context effects on behavioral persistence and recovery.
Table 1.
Experimental Designs
| Experiment | Group | Phase 1 (Training) | Phase 2 (Response limination) | Phase 3 (Test) |
|---|---|---|---|---|
| 1a & 1b | CRF | A: SR+ | B: SR− | A: SR− vs. B: SR− |
| PRF | A: SR+ | B: SR− | A: SR− vs. B: SR− | |
| 2 | CRF-A CRF-B | A: SR+ | A: SR− B: SR- | --- |
| PRF-A PRF-B | A: SR+ | A: SR− B: SR- | --- | |
| 3 | CRF | A: SR+ | B: SR−, O if no R | A: SR−, O if no R vs. B: SR−, O if no R |
| PRF | A: SR+ | B: SR−, O if no R | A: SR−, O if no R vs. B: SR−, O if no R |
Note: CRF is continuous reinforcement, PRF is partial reinforcement, S is discriminative stimulus, R is response, O is reinforcing outcome, A is Context A, B is Context B, + is reinforced, - is nonreinforced, --- Experiment 2 did not include Phase 3. Test order was counterbalanced.
Experiment 1
The first experiment had three goals. First, it established a method for studying discriminated operant behavior in human participants. Second, it examined the effect of PRF training on instrumental extinction. And third, it evaluated response recovery after extinction. The experimental design is shown in Table 1.
Participants completed a computer task adapted from one previously reported by this laboratory (Thrailkill et al., 2019; Thrailkill & Alcalá, 2022) and by others (Morris et al., 2015; 2022; Quail et al., 2017). The task consisted of opportunities to earn points associated with a preferred snack food by pressing a keyboard button. There were three phases. Phase 1 was discriminated operant training in which a target button press was reinforced in the presence of a discriminative stimulus (S) but not its absence. Presentations of S, or trials, were separated by intertrial intervals during which all button presses were nonreinforced. In Phase 2, S presentations continued but target presses in them were no longer reinforced (extinction). Phases 1 and 2 occurred in distinct contextual stimuli (background color and tone combinations). Extinction remained in effect in Phase 3, when responding in S was tested in each context in a counterbalanced order (a within-subject test; cf. Bouton et al., 2011).
Participants were randomly assigned to one of two Phase-1 training conditions defined by two schedules of reinforcement. For one group, the schedule guaranteed that a single reinforcer could be earned contingent on a target-button press in every S (Group Continuous Reinforcement; CRF). For a second group, the schedule arranged the possibility of one or more reinforcers but on only half the S presentations (Group Partial Reinforcement; PRF). The method was modeled on a procedure studied with rats (Thrailkill et al., 2018, Experiment 4). Target responding was expected to extinguish more rapidly and to a lower level in Group CRF than in Group PRF. In the test, responding was also expected to be greater in the training context (Context A) than in the extinction context (Context B) based on previous observations in nonhuman (e.g., Bouton et al., 2011) and human free-operant renewal (e.g., Ritchey et al., 2021a). The experiment was conducted twice: Once in-person with undergraduates in a computer laboratory, and second with an online sample of individuals recruited on a crowdsourcing platform. Thus, in addition to extending the methods often used to study relapse and behavioral persistence in nonhumans to humans, the experiment extends in-person laboratory methods and materials to an online research setting.
Method
Transparency and Openness
Effort has been made to comply with the eight Transparency and Openness Promotion research planning and reporting guidelines. All materials used have been cited. Data and study materials including code are available upon request. The experimental design and analyses were not preregistered. All sample size determinations, criteria for data exclusions, manipulations and behavioral measures are described.
Experiment 1a
Participants.
19 undergraduate students (14 female) enrolled in introductory psychology courses at the University of Vermont (UVM) participated for course credit. Students signed up anonymously to participate via a recruitment website maintained by the UVM Department of Psychological Science. Students did not have prior experience with research or greater than introductory knowledge of psychology. A target sample size of 20 was based on a previous study of response recovery in this laboratory (Thrailkill et al., 2019). To reduce variability in level of hunger, participants were instructed not to eat for 3 hr prior to their appointment. Participants were not asked to verify whether they had eaten. Screening excluded students who reported food allergies. Participants ranged from 18 to 26 years in age. Each participant provided informed consent and the UVM institutional review board approved all procedures and materials.
Materials and apparatus.
All procedures took place in a room that contained a table with a computer (Dell Optiplex 755), 43-cm (diagonal) monitor, keyboard, and mouse. A QWERTY keyboard with extended number pad was positioned 19 cm from the edge of the table, and the monitor was positioned 47 cm from the edge, with the bottom 15 cm from the table’s surface. Green stickers were affixed to keys “1”, “2”, “3”, “4”, “5”, “6”, “7”, “8”, and “9” to the number keys on the number pad located on the right side of the keyboard. This created a 3 by 3 square of green buttons. A blue sticker covered the “>” button that served to advance instruction screens (see below). Button presses were defined as the press and release of a key. Experimental events and data were controlled and collected by programs written with Microsoft Visual Studio 2019 (Richmond, WA). In addition to instructions and counts of outcomes earned, the program could display two cartoon images of a vending machine (white on black background). One measured 5.5 cm by 10.0 cm (w by h) on the screen and the other measured 2.75 cm by 5.0 cm (w by h). The vending machine images were always centered on the screen. Snack food reinforcers (M&M’s, Wavy Lay’s potato chips, or Bare Fuji red apple chips) were present in small bowls arranged in front of their identifying containers on a second table located on the wall behind the computer table. Bowls were not visible while seated at the computer. Images of the snack foods in bowls could be presented on the screen as reinforcers. When displayed, the snack images measured 5 cm by 4 cm (w by h) and were centered on the screen. Contextual stimuli were created with combinations of background screen colors and low-volume continuous tones presented through headphones (blue and 500 Hz, red and 200 Hz, and yellow and 350 Hz). Participants were randomly assigned to one of three pairs of context stimuli prior to the task (red and yellow, blue and red, yellow and blue).
Procedure.
Demographic information including age and gender was collected prior to introduction to the experimental room. Participants were directed to the content of the bowls on the second table and encouraged to sample the snacks before taking a seat in front of the computer. Bowls were not visible to participants when seated at the computer station. Once seated at the computer, participants rated their current level of hunger and the pleasantness of the three snack foods on a 7-point Likert scale. The ensuing experimental task consisted of three sequential phases that were completed in a single laboratory visit that lasted approximately 20 min.
Phase 1 (Discriminated Operant Training).
In the initial phase, participants could press a button to earn their highest-rated snack. The task consisted of the presentation of the vending machine image on the computer monitor. Participants were instructed to use only one finger from their dominant hand to press buttons. Instructions based on Uengoer et al. (2013) and Thrailkill et al. (2019) were read aloud by an experimenter and presented as a series of instruction screens (the full instructions are provided in the online Supplemental Material that accompanies this article). The instructions emphasized the goal of learning what button to press and when to press it to get as many snack points as possible.
Once the participant initiated the experiment, the experimenter left the room. No other instructions were given. For all participants, the task began with the presentation of the small vending machine and the randomly assigned contextual stimuli. Presentations of S were defined by presenting the larger vending machine image which made the vending machine appear to move closer on the screen. During this S, pressing the target button could earn snack points, indicated by a 1-s presentation of an image of the preferred (i.e., highest rated) snack item. The vending machine returned to its initial size (2.75 cm × 5.0 cm) to appear further away on the screen after 6 s. A variable 10-s intertrial interval (ITI) (range: 3 s to 18 s) separated each S presentation. The target response was defined as pressing the center button on the number pad (“5”). The surrounding buttons were defined as control or “adjacent” responses.
Participants were randomly assigned to training conditions distinguished by two reinforcement schedules. For Group Continuous Reinforcement (CRF), the onset of S initiated a variable interval (VI) randomly selected from a list (mean 3 s; range: 2 to 4 s). After the interval elapsed, the first response on the target button (“5”) resulted in a snack point. Upon offset of the snack point, the trial continued, but no further reinforcers were possible. For Group Partial Reinforcement (PRF), S onset initiated a 1-s timer that queried a uniform 1-in-6 probability to randomly determine reinforcer availability (a random-interval 6-s schedule). Once the program indicated a reinforcer was available, the next target button press would present the reinforcer and restart the schedule timer. More than one reinforcer could occur on a trial, but a reinforcer could not be earned if an interval longer than 6 s was the first interval selected. This schedule arranged approximately half the S presentations to be nonreinforced (cf. Thrailkill et al., 2018). The vending machine image disappeared during the reinforcer presentation. The training phase began with an ITI, ended following 24 S presentations, and lasted approximately 6 min.
Phase 2 (Extinction).
After Phase 1, the vending machine was replaced with a message “Please wait, the experiment will continue shortly”. After 4 s, the screen background color and audio tone frequency changed from red/200 Hz to yellow/350 Hz, blue/500 Hz to red/200 Hz, or yellow/350 Hz to blue/500 Hz, and the vending machine appeared in the far position. The color/tone combinations were selected to be highly contrasting. Phase 2 then consisted of 32 presentations of S during which all button presses were recorded but had no consequence (extinction). No snack images were presented. The S and ITI durations were identical to the training phase. Phase 2 lasted approximately 9 min.
Phase 3 (Test).
After Phase 2, the vending machine was replaced with “Please wait, the experiment will continue shortly” for 4 s. Next, participants received 4 additional extinction presentations of S in either the extinction context (Context B) or the training context (Context A). Following 4 trials, the vending machine was again replaced with the 4-s “Please wait, the experiment will continue shortly” message. Participants then received a final block of 4 S presentations in either Context A or Context B depending on whether the first test was in Context B or Context A, respectively. Participants were randomly assigned to receive the test in Context A-Context B or Context B-Context A order. The test trials and ITI durations were determined in the same manner as the extinction and training phases. Responses were recorded but had no consequences. The test phase lasted approximately 2 min.
Experiment 1b
Participants.
Participants were recruited from the crowdsourcing service Prolific. Inclusion criteria were current location in United States, Canada, United Kingdom, or Australia, above 18 years old, and a work-approval rate above 50 percent. Participants were ineligible if they reported significant problems reading text (literacy difficulties), current dieting, chronic disease, diagnosed heart condition, mental health, or illness condition. These criteria were applied by the Prolific platform and were selected prior to advertising the study opportunity. A total of 56,107 participants matched criteria and had been active in the past 90 days. The training group effect in Experiment 1a was entered into a power analysis. The analysis indicated that 25 participants would be needed for power to detect a significant during Phase 2. Based on this result and concern for variability among crowdsourced participants (e.g., Ritchey et al., 2021b), a recruitment target was set at 30 participants per group. Participants were paid individually based on their time to complete the study at a rate of USD $9.00 per hour.
Materials and apparatus.
Participants completed all study materials on their own internet-connected devices. A desktop or laptop computer was required. Mobile phones and tablets were not allowed as they do not generally have an external keyboard. Instructions directed participants to accept the study only if they had approximately 30 min to complete the task without interruption and to calibrate the sound on their computer speakers or headphones to a low setting between 20–30%. Once the participant clicked the link to begin the task a 60-min timer started. Participants that exceeded the time limit were paid but not included. Programs were written in PsychoPy and hosted by Pavlovia.org.
Procedure.
After accepting the task, participants were directed to the study on their browser. Participants first viewed an instruction screen, then rated three snack foods (M&M’s, potato chips, and popcorn). After rating the snack foods, participants received instruction screens describing the task. These showed images of the vending machine and a keyboard with the “Q”, “W”, “E”, “A”, “S”, “D”, “Z”, “X”, and “C” highlighted yellow and the other buttons shaded grey. Note this change was intended to maintain a 3-by-3 block of buttons and accommodate keyboards without an extended number pad. The instructions were the same as those in Experiment 1a and participants progressed through instruction sections, presented one-paragraph-per-screen, at their own pace.
The training and extinction phases were identical to that described in Experiment 1a. The extinction phase was extended to 48 S presentations with the intention to reduce target responding to as low a level as possible. The test proceeded in the same manner as Experiment 1a. The experiment required approximately 21 min to complete.
Data analysis
Button presses were recorded during the ITI and during S. Of interest was the number of presses in S that occurred prior to the delivery of the reinforcer and the number of presses in the 3 s period prior to S onset. A 3-s pre-S period was chosen to provide a baseline of similar duration to the average period preceding a reinforced response in S. Button presses were categorized a “Target” and “Adjacent”, with the reinforced button, “number pad 5” in Experiment 1a and “S” in Experiments 1b, 2, and 3, defined as the target. The buttons immediately surrounding the target, number pad buttons “1”, “2”, “3”, “4”, “6”, “7”, “8”, and “9” or “Q”, “W”, “E”, “A”, “S”, “D”, “Z”, “X”, and “C”, were categorized as adjacent.
Two comparisons provided evidence of behavioral control by the experimental contingencies. First, the number of responses prior to the first reinforcer in S was compared to the number of responses in the 3-s pre-S period is a measure of discriminative control of the response by the S. Second, the mean number of responses on the adjacent buttons in the S and pre-S periods was analyzed. Whether participants increased responses on the target button relative to the adjacent buttons indicates discrimination of the target from the adjacent buttons. Adjacent button presses remained low in each experiment and are presented as mean and standard error (SEM). The two comparisons provide evidence of learning which button to press and when to press it.
Analysis of variance with training group (CRF, PRF) as between-subject factor and 4-trial block as a within-subject factor was used to compare target responding in each experimental phase. For all statistical tests, the alpha level was set at .05. Effect sizes and their confidence intervals are reported for significant tests when relevant to the hypotheses.
Results
Experiment 1a
One participant did not press any buttons during Phase 1 and was excluded. Participants were between 18 and 26 years old. The final sample included 10 (6 female) participants in Group CRF and 8 (7 female) in Group PRF.
Discriminated operant training.
The left panel of Figure 1 shows the results of the discriminated operant training phase. Participants in each group learned to press the target button during the S presentations. Target responding increased across blocks and pre-S target responses remained low in each group. Adjacent button and pre-S responses remained low and were similar in each group. These observations were supported by Group (CRF, PRF) by 4-trial Block (8) ANOVAs that compared target responding in the S and pre-S periods. During S, there was a significant effect of Block, F(7, 112) = 7.47, MSE = 39.65, p < .001, and no effect of group or interaction, largest F(7, 112) = 1.03. The same analysis applied to pre-S target responding found no significant effects, largest F(1, 16) = 1.79, MSE = 1.19. In the final Phase 1 block, mean adjacent responses were 1.1 (SEM = 0.4) and 1.4 (SEM = 0.4), and 0.0 (SEM = 0.0) and 0.1 (SEM = 0.1) in Groups CRF and PRF in the trial and pre-trial periods, respectively. Groups CRF and PRF earned a mean of 25.0 (SEM = 2.1) and 16.4 (SEM = 2.7) reinforcers during the training phase, respectively. This difference was significant, F(1, 16) = 6.65, MSE = 49.74, p = .020.
Figure 1.

Results Experiment 1a.
Note. Target button presses during the training (Phase 1; Left), response elimination (Phase 2; Center), and Test (Phase 3; Right) in Experiment 1a. “CRF” is continuous reinforcement, “PRF” is partial reinforcement, “Target” is target button presses during the discriminative stimulus (S). “Pre-target” is target button presses during the 3-s period prior to S onset, “Pre O” is responses in S before the first reinforcer delivery. Error bars are the standard error of the mean and only appropriate for between group comparisons.
Extinction.
The center panels of Figure 2 show the results of the extinction phase. While target responding remained elevated during S, each group reduced the number of target button presses across blocks of S presentations. Group PRF maintained a higher rate of target pressing than Group CRF. A Group by Block ANOVA supported these observations. During S, Group PRF made more target responses than Group CRF across extinction blocks, F(1, 16) = 5.80, MSE = 159.22, p = .028. A significant effect of block supported the decrease in target responding across blocks, F(7, 112) = 3.89, MSE = 15.74, p < .001. The group by block interaction was not significant, F < 1. The same analysis applied to pre-trial periods, found greater pre-S target responding in Group PRF, F(1, 16) = 4.52, MSE = 0.72, p = .049, and no effect of block or interaction, largest F(7, 112) = 1.99, MSE = 0.08. In the final Phase 2 block, adjacent responses were 1.9 (SEM = 0.4) and 1.6 (SEM = 0.3), and 0.0 (SEM = 0.0) and 0.1 (SEM = 0.0) in Groups CRF and PRF in the S and pre-S periods.
Figure 2.

Results of Experiment 1b
Note. Target button presses during the training (Phase 1; Left), response elimination (Phase 2; Center), and Test (Phase 3; Right) in Experiment 1b. “CRF” is continuous reinforcement. “PRF” is partial reinforcement. Error bars are the standard error of the mean.
Test.
The results of the test phase are shown in the right panel of Figure 1. During S, target responding was greater in the acquisition context, Context A, than in the extinction context, Context B, F(1, 16) = 18.75, MSE = 18.71, p < .001, η = .54, 95% C.I. [.16, .72]. Although target responding was numerically greater in Group PRF, the group effect and group by context interaction were not statistically significant, largest F(1, 16) = 3.68, MSE = 18.71, p = .073. During the pre-S periods, responding was similar in each context. The same analysis found greater responding in Group PRF, F(1, 16) = 9.24, MSE = 0.08, p = .008, and no significant effects involving context, Fs < 1. Adjacent responses during test trials were 2.0 (SEM = 0.4) and 2.4 (SEM = 0.5) in Context B, and 2.2 (SEM = 0.3) and 1.5 (SEM = 0.4) in Context A in Groups CRF and PRF. For pre-trial periods, adjacent responses were 0.1 (SEM = 0.1) and 0.6 (SEM = 0.3) in Context B, and 0.3 (SEM = 0.3) and 0.1 (SEM = 0.1) in Context A in Groups CRF and PRF, respectively.
Experiment 1b
Five participants with blocked I.P. addresses were excluded as location could not be identified (1 in Group CRF and 4 in Group PRF). Prior to analysis, participants that did not show evidence of learning the discriminated operant response were excluded. The learning criterion was a difference of greater than or equal to 1 response in favor of the target button in the final block of Phase 1. Participants that made an average of less than 1 more response on the target button than on the adjacent buttons were excluded. This resulted in 2 and 13 exclusions in Groups CRF and PRF, respectively. One participant in Group CRF was removed for target responding greater than 3 standard deviations (SD) above the group mean in the final extinction block. In the final sample, there were 19 individuals in Group CRF (11 female, age = 31.1 years, SEM = 0.5), and 21 in Group PRF (8 female, age = 28.6 years, SEM = 0.4).
Discriminated operant training.
The left panel of Figure 2 shows the responses across blocks of discriminated operant training. Target responding in S increased while pre-S responses remained low. For target responding during S, a Training by 4-trial Block ANOVA found a significant effect of block, F(7, 266) = 10.05, MSE = 32.85, p < .001, and no significant group effect or interaction, largest F(1, 38) = 2.58, MSE = 152.60. During pre-S periods, the same analysis found an increase in target presses across blocks, F(7, 266) = 4.63, MSE = 6.51, p < .001, and no significant effects involving group, largest F(7, 266) = 1.34. In the final Phase 1 block, adjacent responses were 0.5 (SEM = 0.2) and 1.1 (SEM = 0.5), and 0.4 (SEM = 0.2) and 0.5 (SEM = 0.2) in Groups CRF and PRF in the S and pre-S periods, respectively. Groups CRF and PRF earned a mean of 27.3 (SEM = 0.7) and 16.8 (SEM = 1.4) reinforcers during the training phase. This difference was statistically significant, F(1, 38) = 42.14, MSE = 26.37, p < .001.
Extinction.
The results of the extinction phase are shown in the center panel of Figure 2. Across blocks, target responding in S decreased more slowly in Group PRF. For target responding in S, a Group by 4-trial Block ANOVA found a significant group by block interaction, F(11, 418) = 1.95, MSE = 30.70, p = .032, η = .05, 95% C.I. [.00, .07], and no effects of group or block, largest F(1, 38) = 1.51, MSE = 753.60. The interaction was consistent with the fact that the block effect was present in Group CRF, F(11, 198) = 2.23, MSE = 19.91, p = .014, and not Group PRF, F(11, 220) = 1.49, MSE = 40.42, p = .138. The same analysis applied to target responding in the pre-S periods found no significant effects, largest F(1, 38) = 2.62, MSE = 8.12. In the final Phase 2 block, adjacent responses were 1.2 (SEM = 0.2) and 2.0 (SEM = 0.3), and 0.3 (SEM = 0.1) and 0.5 (SEM = 0.2) in Groups CRF and PRF in the S and pre-S periods.
Test.
Data from the test phase are shown in the right panel of Figure 2. Each group made more target button presses in S when in the acquisition context (Context A) than the extinction context (Context B). The effect of context was significant, F(1, 38) = 8.84, MSE = 18.69, p = .005, η = .19, 95% C.I. [.02, .39], and the group effect and interaction did not reach significance, largest F(1, 38) = 3.92, MSE = 64.54. In pre-S periods, neither group nor context effects on target responding were significant, largest F(1, 38) = 1.52, MSE = 3.60. There was a group by context interaction suggesting pre-S target responding was different between groups across tests, F(1, 38) = 4.67, MSE = 0.75, p = .037. Adjacent responses during S were 1.6 (SEM = 0.3) and 2.1 (SEM = 0.2) in Context B, and 2.1 (SEM = 0.4) and 2.1 (SEM = 0.4) in Context A in Groups CRF and PRF. For pre-S, adjacent responses in Groups CRF and PRF were 0.2 (SEM = 0.1) and 0.4 (SEM = 0.1) in Context B, and 0.2 (SEM = 0.1) and 0.5 (SEM = 0.2) in Context A.
Discussion
In both experiments, participants reliably acquired the discriminated operant response. Clear discriminative control over the target response was observed as a reliable increase above pre-S levels during S prior to the presentation of the reinforcer. Participants also learned to discriminate the target button from the adjacent keyboard buttons, a second index of discriminative control. Although responding was similar at the end of Phase 1, schedule of reinforcement influenced acquisition. More participants in Group PRF failed to meet the acquisition criterion than Group CRF. Group PRF also earned fewer reinforcers than Group CRF during training. Consistent with previous nonhuman and human studies, however, schedule of reinforcement during training influenced extinction of the instrumental response. Despite not being reinforced, target responding remained elevated in the PRF groups compared to CRF groups. There was also evidence of ABA renewal of target responding in the test phase. Renewal was specific to the target response during S and the test context did not influence responding on adjacent buttons or the target during the pre-S periods. Finally, findings from the in-person undergraduate student sample were replicated with an independent remote crowdsourced sample.
It is notable that the results demonstrate an instrumental PREE outside of the acquisition context. As explained below, recent evidence suggests that the PREE might be at least partly context specific. However, Experiment 1 did not compare extinction in Context A to extinction in Context B. Experiment 2 was conducted to address this possibility.
Experiment 2
Studies with rats suggest that discriminated operant responding is partly context specific (Bouton et al., 2014b). Some support for the response, be it discriminative control by the S or the context itself, is lost with a context shift (see also Bouton et al., 2011; Thrailkill & Bouton. 2015). Experiment 2 examined whether changing the context influences the discriminated operant PREE.
Theoretical accounts of PREE might predict a context shift to have a proportional influence on CRF and PRF trained responses since the strength of discriminative control is assumed to be similar after similar amounts of CRF and PRF training (e.g., Capaldi, 1994). Recent evidence from discriminated operant experiments with rats suggests that CRF training can lead to the response becoming insensitive to outcome devaluation which is a hallmark of stimulus-response control (Adams & Dickinson, 1981). In contrast, after the same number of training trials and reinforced responses, PRF training leaves responding sensitive to outcome devaluation and therefore categorized as goal-directed and under response-outcome control (Thrailkill et al., 2018; 2021). Other evidence suggests that, whereas responding under response-outcome control transfers well across contexts, responding under stimulus-response control is context-specific and weakened by a context switch (Thrailkill & Bouton, 2015). Taken together, these findings suggest that PRF trained responses could be more context-general than CRF trained responses. A context switch could weaken a CRF-trained response more than a PRF-trained response and enhance the PREE in Context B. Given the evidence from the rat studies that CRF and PRF training result in different contents of learning, the potential influence of training on the degree to which the discriminated operant PREE transfers across contexts is unclear.
Experiment 2 (outlined in Table 1) therefore extended the method from Experiment 1b to compare extinction after CRF versus PRF training in either the training context (Context A) or a different context (Context B). If the PREE depends on the expectation of reinforcement and nonreinforcement encoded during training as suggested by theoretical accounts, then a similar PREE would be observed in Contexts A and B. Alternatively, if differences in stimulus-response versus response-outcome learning influence the PREE, then the PREE may interact with the extinction context.
Method
Participants
Participants were recruited from Prolific using the same criteria as Experiment 1b. The target sample size of 120 was based on Experiment 1b and determined to provide adequate power to detect statistically significant differences between the groups.
Materials
The same method was used as described in Experiment 1b.
Procedure
As in Experiment 1b, participants that did not complete the study materials within 60 min of accepting the task were paid but not included.
Phase 1 (Discriminated operant training).
Participants completed Phase 1 with continuous (Groups CRF-A and CRF-B) or partial reinforcement (Groups PRF-A and PRF-B) in the same manner as Experiment 1b. All aspects of the procedure were identical.
Phase 2 (Response elimination).
Phase 2 consisted of 48 presentations of the 6-s S. Groups CRF-B and PRF B received the same treatment as Groups CRF and PRF in Experiment 1b consisting of extinction in a context that differed from the training context in terms of the background screen color and tone frequency. Groups CRF-A and PRF-A received the identical treatment but with the same context stimuli as those present during Phase 1. Target and adjacent button responding was recorded but had no other scheduled consequences.
Results
A total of 128 participants consented and were randomly assigned to the conditions. Out of these, 104 completed the task (27 in Group CRF-A, 23 in Group CRF-B, 25 in Group PRF-A, and 29 in Group PRF-B). Twenty-five participants (14 from CRF and 11 from PRF groups) were excluded for failing to meet the performance criteria applied in Experiment 1. Other reasons for failing to complete the materials included exceeding the time limit to complete the materials, internet connection problems, and browser/computer crashes. The final sample included 24 (9 Female, 1 not reported, mean age 31.1 years [SEM = 1.9]) in Group CRF-A, 24 (11 Female, 2 not reported, mean age 36.0 years [SEM = 3.4]) in Group CRF-B, 26 (18 Female; mean age 30.4 years [SEM = 1.6]) in Group PRF-A, and 25 (18 Female, 2 not reported; mean age 32.8 years [SEM = 2.2]) in Group PRF-B.
Discriminated operant training
The left panels of Figure 3 shows target responses across blocks of discriminated operant training. Target responding in S increased while pre-S and adjacent button responses remained low. For target responding during S, a Training by 4-trial Block ANOVA found a significant effect of block, F(7, 665) = 38.89, MSE = 14.12, p < .001, and no significant group effect or interaction, largest F(7, 665) = 1.78. During pre-S periods, the same analysis found an increase in target presses across blocks, F(7, 665) = 5.39, MSE = 4.07, p < .001, and no significant effects involving group, largest F(1, 95) = 1.92, MSE = 25.33. In the final Phase 1 block, adjacent responses in the S were 0.4 (SEM = 0.2), 0.6 (SEM = 0.1), 0.8 (SEM = 0.1), and 0.8 (SEM = 0.1) in Groups CRF-A, CRF-B, PRF-A, and PRF-B. Adjacent responses in pre-S periods were 0.3 (SEM = 0.1), 0.2 (SEM = 0.1), 0.3 (SEM = 0.1), and 0.4 (SEM = 0.1) in Groups CRF-A, CRF-B, PRF-A, and PRF-B, respectively. Groups CRF-A, CRF-B, PRF-A, and PRF-B earned a mean of 26.9 (SEM = 1.3), 25.4 (SEM = 1.2), 19.4 (SEM = 0.8), and 19.1 (SEM = 1.5) reinforcers during the training phase. The difference in number of reinforcers earned was significant based on CRF and PRF training, F(1, 95) = 33.19, MSE = 35.39, p < .001, but not extinction context, Fs < 1.
Figure 3.

Results of Experiment 2.
Note. Target button presses during the training (Left) and response elimination (Right) in Experiment 2. “CRF” is continuous reinforcement, “PRF” is partial reinforcement, “A” is the training context, and “B” is a different context. See text for more details. Error bars are the standard error of the mean.
Extinction
The results of the extinction phase are shown in the right panels of Figure 3. Across blocks, target responding in S decreased more slowly in Groups PRF-A and PRF-B in comparison to the CRF groups. For target responding in S, a Training (CRF, PRF) by Extinction Context (A, B) by 4-trial Block ANOVA found significant effects of block, F(11, 1,045) = 9.30, MSE = 18.98, training by block interaction, F(11, 1,045) = 2.45, p = .005, η = .03, 95% C.I. [.00, .04], and extinction context by block interaction, F(11, 1,045) = 4.07, p < .001, η = .04, 95% C.I. [.01, .06]. That Context by Block interaction took the form of initially higher responding in Context A than Context B. Importantly, there was no evidence of an interaction between Training and Context, Fs < 1. The same analysis applied to target responding in the pre-S periods found a significant Extinction Context by Block interaction, F(11, 1,045) = 1.85, MSE = 5.67, p = .043, and no other significant effects, largest F(1, 95) = 1.63, MSE = 21.20.
Extinction in Context B reduced responding at the beginning of the extinction phase. This was examined in each training group with Extinction Context by Block ANOVAs. For CRF groups, there was a significant effect of block, F(11, 506) = 9.13, p < .001, and significant extinction context by block interaction, F(11, 506) = 1.97, MSE = 18.72, p = .035, η = .04, 95% C.I. [.00, .06]. The same analysis on PRF groups found a significant block effect, F(11, 539) = 2.52, p = .004, and extinction context by block interaction, F(11, 539) = 3.07, MSE = 19.22, p < .001, η = .06, 95% C.I. [.01, .08].
Discussion
Participants acquired the discriminated operant response under CRF and PRF schedules without incident. Consistent with the loss of some discriminative control over the response observed with rats (Bouton et al., 2014b), the results show clear evidence of the context specificity of discriminated operant responding. There was an immediate reduction in the response in groups that received extinction in Context B compared to groups extinguished in Context A. At the same time, the PREE was observed in each context and did not appear different in Context B in comparison to Context A. This finding is consistent with theories of the PREE under the assumption that CRF and PRF training results in excitatory stimulus control of the response that transfers equally well across contexts. Greater stimulus-response learning may have been present after CRF, but this did not appear to influence the PREE after the context switch. A recent study with rats found that a response under clear stimulus-response control after CRF training extinguished more rapidly than a response under clear response-outcome control trained with PRF (Thrailkill et al., 2018, Experiment 4). That is, the content of learning differed after discriminated operant training with CRF versus PRF but did not interact with the PREE when tested in the training context. Future research is needed to assess the motivational control of instrumental behavior after CRF and PRF training in humans.
So far, the results document the influence of partial reinforcement on discriminated operant extinction and renewal. One feature is that target responding was never completely abolished at the end of Phase 2. This is consistent with observations of incomplete extinction in human free-operant procedures (e.g., Weiner, 1970). Factors such as demand characteristics, task duration, or other task features may contribute to continued responding during extinction (e.g., Rosenfarb et al., 1992). Experiment 3 therefore developed omission training as a method to eliminate responding to a greater degree than extinction and examine the influence of PRF on response elimination with an omission schedule. In addition to a greater reduction in the response, it would be worth knowing whether responding eliminated with omission training was still sensitive to the PREE and renewal (Nakajima et al., 2002; Rey et al., 2020). A greater reduction of responding could also imply stronger response inhibition learning during omission training (Rey et al., 2020) which in turn could more effectively reduce persistence after PRF training.
Experiment 3
The goal of Experiment 3 was to extend the method developed in Experiment 1 to study how Phase 1 training conditions influence response elimination with omission training. The design is shown in Table 1. It followed Experiment 1 except for the response elimination method. In Phase 1, groups received either CRF or PRF training in Context A. In Phase 2, each group received an omission contingency in Context B. Finally, all participants received a test in each context in counterbalanced order with the omission contingency still in effect (cf. Nakajima et al., 2002). As before, renewal of the target response was expected when each group was tested in Context A. Importantly, the influence of PRF on omission training has not been examined previously in humans or nonhumans.
Method
Participants
Participants were recruited from Prolific using the same criteria as Experiments 1b and 2.
Materials
Participants completed demographic items presented on Gorilla Experiment Builder (Anwyl-Irwine et al., 2020). Randomization and quota tools were used to randomize assignment of participants to experimental conditions and to control accrual based on demographic characteristics. The design assigned a similar number of male- and female-identifying participants to each group (other-identifying participants were randomly assigned to a group). After completing the demographic items on Gorilla, participants were directed to the experimental materials hosted on Pavlovia.org.
Procedure
As in Experiments 1b and 2, participants that did not complete the study materials within 60 min of accepting the task were paid but not included.
Phase 1 (Discriminated operant training).
Participants completed Phase 1 with continuous (Group CRF) or partial reinforcement (Group PRF) in the same manner as Experiments 1b and 2. All other aspects of the procedure were identical except that the ITI was shortened to from 12 s to 6 s, on average. This change was made to reduce the time required to complete the study. Although a shortened ITI could result in greater context conditioning and slow acquisition of the operant discrimination, there was no basis to expect the change would have a differential influence on sensitivity to the omission contingency for Groups CRF and PRF.
Phase 2 (Response elimination).
Phase 2 consisted of 24 presentations of the 6-s S in Context B. A reinforcer was scheduled to be presented in the final second of the 6-s S presentation. A single response on the target button during S canceled the reinforcer presentation. Target responding had no other scheduled consequences. Phase 2 was shortened to 24 trials to reduce the amount of time required to complete the study.
Phase 3 (Test).
Tests were conducted the same manner as in Experiment 1b with the exception that the reinforcers were scheduled as they were for during Phase 2; that is, the omission contingency continued to operate during the test. The entire experiment required approximately 11 min to complete.
Results
A total of 93 participants consented and were randomly assigned to the conditions. Out of these, 63 completed the task (32 in Group CRF and 31 in Group PRF). Reasons for failing to complete the materials included exceeding a quota for group assignment, exceeding the time limit to complete the materials, internet connection problems, and browser/computer crashes. The same performance criteria were applied prior to analysis as in Experiments 1 and 2. A total of 14 participants did not meet the acquisition criterion (7 in Group CRF and 7 in Group PRF). The final sample included 25 (14 Female, mean age 28.5 years [SEM = 1.6]) in Group CRF and 24 (13 Female; mean age 33.8 years [SEM = 1.9]) in Group PRF.
Discriminated operant training
Acquisition took place without incident and is shown in the left panel of Figure 4. Target responding increased across blocks and was influenced by the training condition. That is, Group CRF made more responses than Group PRF. A Group by Block ANOVA supported these observations. There were significant effects of group, F(1, 47) = 5.24, MSE = 196.05, p = .027, η = .10, 95% C.I. [.00, .27], and block, F(7, 329) = 21.03, MSE = 24.26, p < .001, and no interaction, F < 1. The same analysis applied to pre-S target responding found a significant effect of block, F(7, 329) = 3.56, MSE = 1.31, p = .001, and no significant effects involving group, largest F(1, 47) = 3.44, MSE = 4.73. Adjacent responses were 0.5 (SEM = 0.2) and 0.3 (SEM = 0.2), and 0.7 (SEM = 0.2) and 0.1 (SEM = 0.0) in S and pre-S periods in Groups CRF and PRF, respectively. Groups CRF and PRF earned a mean of 27.2 (SEM = 0.9) and 18.6 (SEM = 1.1) reinforcers during the training phase. This difference was statistically significant, F(1, 47) = 36.16, MSE = 25.15, p < .001.
Figure 4.

Results of Experiment 3.
Note. Target button presses during the training (Phase 1; Left), response elimination (Phase 2; Center), and Test (Phase 3; Right) in Experiment 3. “CRF” is continuous reinforcement. “PRF” is partial reinforcement. Error bars are the standard error of the mean.
Response elimination
The center panel of Figure 4 shows the results of the response elimination phase. For Group CRF, target responding was reduced to a low level by the end of the phase. In contrast, although target responding occurred at a lower rate during training, target responding remained elevated across blocks in Group PRF. These observations were supported by a Group by Block ANOVA. There was a significant effect of block, F(5, 235) =12.88, MSE = 21.31, p < .001, and a significant group by block interaction, F(5, 235) = 4.19, p = .001, η= .08, 95% C.I. [.01, .14]. The same analysis of pre-trial target responses found a significant effect of block, F(5, 235) = 2.58, MSE = 0.35, p = .027, and no other significant effects, largest F(5, 235) = 2.21. Adjacent responses were 2.9 (SEM = 0.3) and 0.2 (SEM = 0.1), and 1.8 (SEM = 0.3) and 0.1 (SEM = 0.0) in S and pre-S periods in Groups CRF and PRF, respectively. With the omission contingency in effect, Groups CRF and PRF earned a mean of 16.8 (SEM = 1.5) and 11.2 (SEM = 1.8) reinforcers. This difference was statistically significant, F(1, 47) = 5.83, MSE = 66.67, p = .020.
Test
The results of the tests are shown on the right in Figure 4. Target responding was greater in Context A than Context B, suggesting renewal, and this difference appeared stronger in Group CRF, although that comparison was not reliable. A Group by Context ANOVA found a significant effect of context, F(1, 47) = 13.87, MSE = 12.55, p < .001, η = .23, 95% C.I. [.05, .41]; the other effects were not significant, largest F(1, 47) = 3.58, MSE = 12.55. The same analysis applied to the pre-trial target responding found a significant interaction, F(1, 47) = 5.09, MSE = 0.25, p = .029, and no other significant effects, Fs < 1. Group pre-S responding did not differ statistically in either context, largest F(1, 47) = 2.44, MSE = 0.37, p = .125. Adjacent responses during S were 3.1 (SEM = 0.3) and 2.2 (SEM = 0.3) in Context B, and 2.8 (SEM = 0.3) and 2.1 (SEM = 0.3) in Context A in Groups CRF and PRF, respectively. For pre-S periods, adjacent responses in Groups CRF and PRF were 0.2 (SEM =0.1) and 0.1 (SEM = 0.0) in Context B, and 0.4 (SEM = 0.1) and 0.1 (SEM = 0.0) in Context A, respectively. Groups CRF and PRF earned 3.2 (SEM = 0.3) and 2.5 (SEM = 0.4) reinforcers in Context B, and 2.5 (SEM = 0.3) and 2.2 (SEM = 0.3) reinforcers in Context A. The difference in reinforcers across test blocks was statistically significant, F(1, 47) = 6.04, MSE = 0.88, p = .018. The number of reinforcers did not differ by group and there was no interaction, largest F(1, 47) = 1.75, MSE = 3.91.
An additional analysis examined target responding in the test expressed as the difference between each test block and responding in the final block of Phase 2. Differences in responding were −0.14 (SEM = 0.50) and −0.01 (SEM = 0.52), and 4.24 (SEM = 1.02) and 1.30 (SEM = 0.75) in Groups CRF and PRF in Context B and A, respectively. A Group by Context ANOVA found a significant context effect, F(1, 45) = 15.01, MSE = 12.68, p < .001, and group by context interaction, F(1, 45) = 4.36, p = .042. Follow-up comparisons found a significant context effect in Group CRF, F(1, 22) = 17.81, MSE = 12.39, p < .001, but not in Group PRF, F(1, 22) = 1.59, MSE = 12.95, p = .219.
Discussion
Participants again learned the discriminated operant task. In contrast to Experiment 1, responding was lower in Group PRF than Group CRF during Phase 1. Differences in training response rate have been observed between PRF and CRF in prior studies (e.g., Pittenger et al., 1988), and may reflect stronger conditioning in Group CRF. However, this difference reversed in Phase 2, suggesting that PRF training resulted in more resistance to response elimination with the omission contingency. Response rates in Group PRF remained elevated across blocks and crossed over Group CRF which decreased to low levels by the end of Phase 2. There was also evidence that PRF training resulted in fewer reinforcers earned under the omission contingency. In the test, responding in each group renewed in Context A and remained suppressed in Context B. Relative to responding on the final Phase 2 block, the test results suggested stronger renewal in Group CRF. However, this should be taken as preliminary since responding differed between groups just prior to the test. The results clearly suggest that, like extinction, response elimination with an omission contingency was influenced by PRF training. To my knowledge, this is the first observation of a Partial Reinforcement Omission Effect (PROE). The result is consistent with theory suggesting that nonreinforced trials in PRF training increase response generalization when reinforcers are omitted (Capaldi, 1967; 1994; Harris, 2019).
General Discussion
The current experiments examined the influence of the reinforcement schedule used during training on elimination and renewal of human discriminated operant responding. Experiment 1 developed and validated a method for studying the acquisition and extinction of discriminated operant responding with humans participating in in-person and online research settings. In each setting, participants reliably learned to increase the target button press during the S to earn reinforcers and to discriminate the target button from the adjacent buttons. Acquisition was robust with consistent (i.e., CRF) and inconsistent (i.e., PRF) reinforcers arranged for target responding in S. Extinction of target responding in S proceeded more slowly across blocks of trials following PRF compared to CRF training. However, testing S in Context A caused target responding to renew in comparison to Context B. This result was observed with in-person laboratory and remote crowdsourced samples and is consistent with other demonstrations of ABA renewal of extinguished operant responding (Bouton, 2019; Finch et al., 2022; Ritchey et al., 2021a, 2021b; Robinson & Kelley, 2020; Vila et al., 2020, 2022).
Experiment 2 then assessed the influence of the extinction context on the instrumental PREE. There was clear evidence of a reduction in discriminated operant responding corresponding to the change in context, a context-switch effect (Bouton et al., 2014b; Rosas et al., 2013). There was also clear evidence of the PREE in both contexts. Extinction was always slower following PRF compared to CRF training.
Experiments 1 and 2 suggest that the instrumental PREE transfers remarkably well across contexts. Instrumental responses are partly context specific (Abiero & Bradfield, 2021; Bouton et al., 2014; Thrailkill & Bouton, 2015; Trask & Bouton, 2018), and the loss of operant responding across contexts can be, at least partly, attributed to loss of control by the context or other discriminative stimuli (Thrailkill & Bouton, 2015). This effect was demonstrated in Experiment 2. Accounts of the PREE might predict a context switch to affect CRF- and PRF trained responses similarly (Capaldi, 1967, 1994). However, this also assumes the CRF and PRF training result in a similar form of discriminative stimulus control which may not be true given recent evidence of stronger stimulus-response learning following CRF training (Thrailkill et al., 2018, 2021). This raised the possibility of the context switch producing a similar loss of stimulus control for CRF and PRF but having a differential effect on the CRF-trained response due to stronger stimulus-response learning. The present data did not support this prediction: A PREE was observed in Context B in Experiments 1 and 2, and the effect was similar when compared in trained (Context A) and untrained (Context B) contexts in Experiment 2. While that result should be taken as preliminary given the null finding, there was also clear evidence of a context change weakening the discriminated operant response in CRF- and PRF-trained groups. This is positive evidence of context specificity of discriminative control of the response. There are reasons to predict a context switch to not weaken the PREE as Pavlovian conditioned responses tend to transfer very well, and often perfectly, across contexts (Bouton & King, 1983; Bouton & Peck, 1990; Rosas et al., 2013). Consistent with this analysis, a context switch did not weaken the PREE in a study of rats autoshaped lever pressing (Boughner & Papini, 2006). Information about reinforcement encoded during discriminated operant training seems to transfer across contexts to influence instrumental response elimination (Capaldi, 1967, 1994; Harris, 2019). Nonetheless, additional experiments are needed to further dissociate control by context-specific stimulus-response and context-general response-outcome processes in human discriminated operants and understand the implications for behavioral persistence outside the laboratory.
Experiment 3 introduced omission training in Phase 2 as an alternative method of response elimination. Although omission training has been studied with free-operant (Finch et al., 2022; Nakajima et al, 2002; Rey et al., 2020; Vila et al, 2020, 2022) and Pavlovian sign-tracking procedures (e.g., Killeen, 2003; Sanabria et al., 2006; Williams & Williams, 1969), the present study is the first to show that an omission contingency can reduce a discriminated operant response to low levels. Training reinforcement schedule influenced response elimination with omission training, the PROE. There are several explanations for why PRF training resulted in less sensitivity to the omission contingency relative to CRF training. One possibility is that, after PRF training, reinforcer signaling by N trials interfered with detection of the changed response-outcome contingency (Capaldi, 1967, 1994). A second possibility is that CRF training allowed faster detection of the changed contingency. If supported, this would suggest CRF training facilitates detection of positive, extinction, and omission contingencies. A third possibility is that PRF training weakens or interferes with the ability to detect changes in causal relationships between actions and outcomes. A full understanding of the relative importance of the stimulus-outcome and response-outcome contingencies in response elimination requires further study.
Another important finding is that the omission contingency did not prevent the ABA renewal effect (see also Finch et al., 2022; Nakajima et al., 2002). The findings adds to results with rats showing that responses eliminated by omission training are as vulnerable to renewal as responses eliminated by extinction (Rey et al., 2020). Renewal was observed in Experiments 1 and 3 and suggests that extinction and omission training may involve a common form of context-specific learning about the response (Bouton, 2019; Bouton et al., 2021). Practically, the implication may be that relapse is just as likely following interventions that include omission contingencies to eliminate target behavior as those that use extinction (e.g., Higgins et al., 1991, 2004). Several lines of evidence suggest this implication extends to other response suppression methods including positive punishment and reinforcement of alternative behaviors (e.g., Bouton & Schepers, 2015; Broomer & Bouton, 2022; Thrailkill et al., 2019).
Experiment 3 extended the analysis of omission to include the effect of training reinforcement schedule. The experiment is the first to document the influence of PRF on the effectiveness of omission training, the PROE. The influence of training reinforcement schedule on recovery after omission was also examined. There was some evidence to suggest PRF training influenced renewal. Here again, further research and replication is needed to determine whether this result was due to a response-scale difference at the end of training or due to the combination of PRF and omission training.
Remarkable consistency was observed between in-person and crowdsourced versions of the discriminated operant task. This is a growing area of methodological development with potential to facilitate data collection on topics of practical and theoretical significance (see e.g., Lee & Lovibond, 2021; Ritchey et al., 2021a; b; Robinson & Kelley, 2020). While the area is promising, best practices are developing and methods for addressing known and unknown sources of bias and variability have not been consistent across studies (e.g., Agley et al., 2021; Mellis & Bickel, 2019). One possible weakness of the present experiments was that, although datasets with blocked I.P. addresses were removed, other forms of activity such as inattentive responding could impact data quality. The present experiments are unique in that performance criteria were developed from in-person observations and replicated in the online samples with little modification to the method. Experiments 2 and 3 then manipulated variables with different participants from the online population while replicating parts of Experiment 1. Internal replication is a feature of the experimental psychology approach that may apply particularly well to the development of online data collection methods.
In summary, the present experiments studied the effects of training schedule of reinforcement and context on the elimination and renewal of discriminated operant responding in human participants. Discriminated operants are particularly relevant for understanding factors that influence human behavior because most of our everyday behavior is under antecedent stimulus control (Skinner, 1953). The history of reinforcement associated with the discriminative stimulus influenced persistence during extinction and omission training but did little to influence on the renewal effect. These findings connect the understanding of influences on behavioral persistence and relapse from nonhumans to humans and outline several areas to pursue in future studies. Reinforcer predictability and renewal are likely to have a pervasive influence on the effectiveness of behavior reduction interventions.
Supplementary Material
Acknowledgments
This project was supported by the National Institutes of Health awards K01-DA044456 from the National Institute on Drug Abuse and P20-GM103644 from the National Institute on General Medical Sciences. Data and materials are available upon request. Portions of the results were presented at the 2021 meetings of the Eastern Psychological Association, Association for Behavior Analysis International, and American Psychological Association. I thank Mark Bouton, Julian Kafka, and Margaret Palmiero for assistance and helpful discussions during the preparation of this manuscript.
References
- Abiero AR, & Bradfield LA (2021). The contextual regulation of goal-directed actions. Current Opinion in Behavioral Sciences, 41, 57–62. 10.1016/j.cobeha.2021.03.022 [DOI] [Google Scholar]
- Adams CD, & Dickinson A (1981). Actions and habits: Variations in associative representations during instrumental learning. In Spear NE & Miller RR (Eds.), Information processing in animals: Memory mechanisms (pp. 143–165). Erlbaum. [Google Scholar]
- Agley J, Xiao Y, Nolan R, & Golzarri-Arroyo L (2021). Quality control questions on Amazon’s Mechanical Turk (MTurk): A randomized trial of impact on the USAUDIT, PHQ-9, and GAD-7. Behavior Research Methods, 1–13. 10.3758/s13428-021-01665-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Amsel A (1962). Frustrative nonreward in partial reinforcement and discrimination learning: Some recent history and a theoretical extension. Psychological Review, 69(4), 306–328. 10.1037/h0046200 [DOI] [PubMed] [Google Scholar]
- Amsel A (1992). Frustration Theory. Cambridge, England: Cambridge University Press [Google Scholar]
- Anwyl-Irvine AL, Massonnié J, Flitton A, Kirkham N, & Evershed JK (2020). Gorilla in our midst: An online behavioral experiment builder. Behavior Research Methods, 52(1), 388–407. 10.3758/s13428-019-01237-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boughner RL, & Papini MR (2006). Survival of the partial reinforcement extinction effect after contextual shifts. Learning and Motivation, 37(4), 304–323. 10.1016/j.lmot.2005.09.001 [DOI] [Google Scholar]
- Bouton ME (2019). Extinction of instrumental (operant) learning: interference, varieties of context, and mechanisms of contextual control. Psychopharmacology 236, 7–19. 10.1007/s00213-018-5076-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bouton ME, & Broomer MC (2023). Learning to stop responding. Behavioural Processes, 104830. 10.1016/j.beproc.2023.104830 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bouton ME, & King DA (1983). Contextual control of the extinction of conditioned fear: Tests for the associative value of the context. Journal of Experimental Psychology: Animal Behavior Processes, 9(3), 248–265. 10.1037/0097-7403.9.3.248 [DOI] [PubMed] [Google Scholar]
- Bouton ME, Maren S, & McNally GP (2021). Behavioral and neurobiological mechanisms of Pavlovian and instrumental extinction learning. Physiological Reviews. 101, 611–681. 10.1152/physrev.00016.2020 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bouton ME, & Peck CA (1989). Context effects on conditioning, extinction, and reinstatement in an appetitive conditioning preparation. Animal Learning & Behavior, 17(2), 188–198. 10.3758/BF03207634 [DOI] [Google Scholar]
- Bouton ME, & Schepers ST (2015). Renewal after the punishment of free operant behavior. Journal of Experimental Psychology: Animal Learning and Cognition, 41(1), 81–90. 10.1037/xan0000051 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bouton ME, Todd TP, & León SP (2014). Contextual control of discriminated operant behavior. Journal of Experimental Psychology: Animal Learning and Cognition, 40(1), 92–105. 10.1037/xan0000002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bouton ME, Todd TP, Vurbic D, & Winterbauer NE (2011). Renewal after the extinction of free operant behavior. Learning & Behavior, 39, 57–67. 10.3758/s13420-011-0018-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bouton ME, Trask S, & Carranza-Jasso R (2016). Learning to inhibit the response during instrumental (operant) extinction. Journal of Experimental Psychology: Animal Learning and Cognition, 42(3), 246–258. 10.1037/xan0000102 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Broomer MC, & Bouton ME (2022). A comparison of renewal, spontaneous recovery, and reacquisition after punishment and extinction. Learning & Behavior, 1–12. 10.3758/s13420-022-00552-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Capaldi EJ (1967). A sequential hypothesis of instrumental learning. In Spence KW & Spence JT (Eds.), Psychology of Learning and Motivation (Vol. 1, pp. 67–156). New York: Academic Press. [Google Scholar]
- Capaldi EJ (1994). The sequential view: From rapidly fading stimulus traces to the organization of memory and the abstract concept of number. Psychonomic Bulletin & Review, 1(2), 156–181. 10.3758/BF03200771 [DOI] [PubMed] [Google Scholar]
- Chan CKJ, & Harris JA (2019). The partial reinforcement extinction effect: The proportion of trials reinforced during conditioning predicts the number of trials to extinction. Journal of Experimental Psychology: Animal Learning and Cognition, 45(1), 43–58. 10.1037/xan0000190 [DOI] [PubMed] [Google Scholar]
- Church RM (1963). The varied effects of punishment on behavior. Psychological Review, 70(5), 369–402. 10.1037/h0046499 [DOI] [PubMed] [Google Scholar]
- Church RM (1969). Response suppression. In Campbell BA & Church RM (Eds.), Punishment and Aversive Behavior (pp. 111–156). New York, NY: Appleton-Century-Crofts. [Google Scholar]
- De Meyer H, Beckers T, Tripp G, & Van der Oord S (2019). Reinforcement contingency learning in children with ADHD: Back to the basics of behavior therapy. Journal of Abnormal Child Psychology, 47(12), 1889–1902. 10.1007/s10802-019-00572-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- Finch KR, Williams CL, & Kestner KM (2022). ABA and ABC Renewal during Ongoing Omission Training. The Psychological Record, 1–21. 10.1007/s40732-022-00524-y [DOI] [Google Scholar]
- Gallistel CR, & Gibbon J (2000). Time, rate, and conditioning. Psychological Review, 107(2), 289–344. 10.1037/0033-295X.107.2.289 [DOI] [PubMed] [Google Scholar]
- Harris JA (2019). The importance of trials. Journal of Experimental Psychology: Animal Learning and Cognition, 45(4), 390–404. 10.1037/xan0000223 [DOI] [PubMed] [Google Scholar]
- Harris JA, & Kwok DWS (2018). The probability of reinforcement per trial affects posttrial responding and subsequent extinction but not within-trial responding. Journal of Experimental Psychology: Animal Learning and Cognition, 44(1), 23–35. 10.1037/xan0000158 [DOI] [PubMed] [Google Scholar]
- Higgins ST, Delaney DD, Budney AJ, Bickel WK, Hughes JR, Foerg F, & Fenwick JW (1991). A behavioral approach to achieving initial cocaine abstinence. American Journal of Psychiatry, 148, 1218–1224. 10.1176/ajp.148.9.1218 [DOI] [PubMed] [Google Scholar]
- Higgins ST, & Silverman KE (1999). Motivating Behavior Change Among Illicit-Drug Abusers: Research on Contingency Management Interventions (pp. xv–399). American Psychological Association. [Google Scholar]
- Jean-Richard-dit-Bressel P, Lee JC, Liew SX, Weidemann G, Lovibond PF, & McNally GP (2021). Punishment insensitivity in humans is due to failures in instrumental contingency learning. Elife, 10, e69594. 10.7554/eLife.69594 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Killeen PR (2003). Complex dynamic processes in sign tracking with an omission contingency (negative automaintenance). Journal of Experimental Psychology: Animal Behavior Processes, 29(1), 49–61. 10.1037/0097-7403.29.1.49 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lee JC, & Lovibond PF (2021). Individual differences in causal structures inferred during feature negative learning. Quarterly Journal of Experimental Psychology, 74(1), 150–165. 10.1177/1747021820959286 [DOI] [PubMed] [Google Scholar]
- Mackintosh NJ (1974). The Psychology of Animal Learning. Academic Press. [Google Scholar]
- Morris RW, Quail S, Griffiths KR, Green MJ, & Balleine BW (2015). Corticostriatal control of goal-directed action is impaired in schizophrenia. Biological Psychiatry, 77(2), 187–195. 10.1016/j.biopsych.2014.06.005 [DOI] [PubMed] [Google Scholar]
- Mellis AM, & Bickel WK (2020). Mechanical Turk data collection in addiction research: Utility, concerns and best practices. Addiction, 115(10), 1960–1968. 10.1111/add.15032 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nakajima S, Urushihara K, & Masaki T (2002). Renewal of operant performance formerly eliminated by omission or noncontingency training upon return to the acquisition context. Learning and Motivation, 33(4), 510–525. 10.1016/S0023-9690(02)00009-7 [DOI] [Google Scholar]
- Pittenger DJ, Pavlik WB, Flora SR, & Kontos J (1988). Analysis of the partial reinforcement extinction effect in humans as a function of sequence of reinforcement schedules. The American Journal of Psychology, 371–382. 10.2307/1423085 [DOI] [Google Scholar]
- Quail SL, Laurent V, & Balleine BW (2017). Inhibitory Pavlovian–instrumental transfer in humans. Journal of Experimental Psychology: Animal Learning and Cognition, 43(4), 315–324. 10.1037/xan0000148 [DOI] [PubMed] [Google Scholar]
- Rey CN, Thrailkill EA, Goldberg KL, & Bouton ME (2020). Relapse of operant behavior after response elimination with an extinction or an omission contingency. Journal of the Experimental Analysis of Behavior, 113(1), 124–140. 10.1002/jeab.568 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ritchey CM, Kuroda T, & Podlesnik CA (2021a). Evaluating effects of context changes on resurgence in humans. Behavioural Processes, 104563. 10.1016/j.beproc.2021.104563 [DOI] [PubMed] [Google Scholar]
- Ritchey CM, Kuroda T, Rung JM, & Podlesnik CA (2021b). Evaluating extinction, renewal, and resurgence of operant behavior in humans with Amazon Mechanical Turk. Learning and Motivation, 74, 101728. 10.1016/j.lmot.2021.101728 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Robinson TP, & Kelley ME (2020). Renewal and resurgence phenomena generalize to Amazon’s Mechanical Turk. Journal of the Experimental Analysis of Behavior, 113(1), 206–213. 10.1002/jeab.576 [DOI] [PubMed] [Google Scholar]
- Rosas JM, Todd TP, & Bouton ME (2013). Context change and associative learning. Wiley Interdisciplinary Reviews: Cognitive Science, 4(3), 237–244. 10.1002/wcs.1225 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rosenfarb IS, Newland MC, Brannon SE, & Howey DS (1992). Effects of self‐generated rules on the development of schedule‐controlled behavior. Journal of the Experimental Analysis of Behavior, 58(1), 107–121. 10.1901/jeab.1992.58-107 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sanabria F, Sitomer MT, & Killeen PR (2006). Negative automaintenance omission training is effective. Journal of the Experimental Analysis of Behavior, 86(1), 1–10. 10.1901/jeab.2006.36-05 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Skinner BF (1953). Science and Human Behavior. New York, NY: Macmillan. [Google Scholar]
- Thrailkill EA, & Alcalá JA (2022). Relapse after incentivized choice treatment in humans: A laboratory model for studying behavior change. Experimental and Clinical Psychopharmacology, 30(2), 220–234. 10.1037/pha0000443 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thrailkill EA, Ameden WC, & Bouton ME (2019). Resurgence in humans: Reducing relapse by increasing generalization between treatment and testing. Journal of Experimental Psychology: Animal Learning and Cognition, 45(3), 338–349. 10.1037/xan0000209 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thrailkill EA, & Bouton ME (2015). Contextual control of instrumental actions and habits. Journal of Experimental Psychology: Animal Learning and Cognition, 41(1), 69–80. 10.1037/xan0000045 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thrailkill EA, Kacelnik A, Porritt F, & Bouton ME (2016). Increasing the persistence of a heterogeneous behavior chain: Studies of extinction in a rat model of search behavior of working dogs. Behavioural Processes, 129, 44–53. 10.1016/j.beproc.2016.05.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thrailkill EA, Michaud NL, & Bouton ME (2021). Reinforcer predictability and stimulus salience promote discriminated habit learning. Journal of Experimental Psychology: Animal Learning and Cognition, 47(2), 183–199. 10.1037/xan0000285 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Thrailkill EA, Trask S, Vidal P, Alcalá JA, & Bouton ME (2018). Stimulus control of actions and habits: A role for reinforcer predictability and attention in the development of habitual behavior. Journal of Experimental Psychology: Animal Learning and Cognition, 44(4), 370–384. 10.1037/xan0000188 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Trask S, & Bouton ME (2018). Retrieval practice after multiple context changes, but not long retention intervals, reduces the impact of a final context change on instrumental behavior. Learning & Behavior, 46(2), 213–221. 10.3758/s13420-017-0304-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- Uengoer M, Lotz A, & Pearce JM (2013). The fate of redundant cues in human predictive learning. Journal of Experimental Psychology: Animal Behavior Processes, 39(4), 323–333. 10.1037/a0034073 [DOI] [PubMed] [Google Scholar]
- Uhl CN, & Garcia EE (1969). Comparison of omission with extinction in response elimination in rats. Journal of Comparative and Physiological Psychology, 69(3), 554–562. 10.1037/h0028243 [DOI] [Google Scholar]
- Vila J, Rojas-Iturria F, & Bernal-Gamboa R (2020). ABA renewal and spontaneous recovery of operant performance formerly eliminated by omission training. Learning and Motivation, 70, 101631. 10.1016/j.lmot.2020.101631 [DOI] [Google Scholar]
- Vila J, Rojas-Iturria F, & Bernal-Gamboa R (2022). Comparison of operant behavior renewal after elimination by extinction or differential reinforcement of other behavior. Behavior Analysis: Research and Practice, 22(3), 238–245. 10.1037/bar0000192 [DOI] [Google Scholar]
- Weiner H (1970). Instructional control of human operant responding during extinction following fixed‐ratio conditioning. Journal of the Experimental Analysis of Behavior, 13(3), 391–394. 10.1901/jeab.1970.13-391 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Williams DR, & Williams H (1969). Auto‐maintenance in the pigeon: sustained pecking despite contingent non‐reinforcement. Journal of the Experimental Analysis of Behavior, 12(4), 511–520. 10.1901/jeab.1969.12-511 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Witte KL (1977). Children’s resistance to extinction: Two tests of the discrimination hypothesis. Bulletin of the Psychonomic Society, 9(4), 262–264. 10.3758/BF03336994 [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
