Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Aug 31;63(9):e70387. doi: 10.1111/psyp.70387

Distinct Neural Dynamics Underlying Reward Expectation Formation and Outcome Processing: Insights From ERP and Time‐Frequency Analyses

Matthew D Bachman 1,, Kaya Scheman 2, René San Martin 3, Scott A Huettel 4, Marty G Woldorff 4,5
PMCID: PMC13530306  PMID: 42675605

ABSTRACT

Reward expectations are fundamental to theories of decision making and reinforcement learning. While prior research has focused on how expectations can influence outcome processing, far fewer studies have investigated how these expectations are actually formed. To address this gap, we measured EEG activity from participants as they completed a two‐stage binary‐choice task designed to separate the formation of expectations about outcome probabilities from the processing of actual reward outcomes. To more fully examine the neural mechanisms underlying each stage, we measured the Reward Positivity (RewP) and P3b time‐domain event‐related potential (ERP) components, as well as the delta‐ and theta‐band activity underlying each ERP. During the Outcome Probability stage, participants learned the likelihood of their choice winning on that trial. Each measure of RewP‐latency activity (ERP, delta, theta) was larger for outcomes that were certain to occur, but each measure diverged in its relationship to outcome valence. Conversely, all P3b‐latency measures were increased for losses that were certain to occur. Notably, changes in RewP‐Theta, not in ERP components, provided the earliest marker of sensitivity to certain losses. At the Actual Outcome stage, the RewP and P3b‐ERPs were larger for unexpected wins, consistent with theories of reward prediction errors and context updating. Delta activity generally followed the patterns observed in its temporally matched ERP but displayed an inconsistent relationship with outcome valence, suggesting that it may reflect contextualized feedback processing rather than a specific win‐related signal. Theta was insensitive to outcome valence and only sensitive to expectations at longer latencies, indicating a shift from valence‐sensitive processing during expectation formation to a broader role in monitoring expectancy violations. Together, the results underscore the importance of temporally and functionally distinguishing between the expectation formation and outcome phases, while demonstrating the value of multimethodological analytical approaches to fully capture the dynamic nature of reward‐based decision‐making.

Keywords: EEG, reward expectations, reward positivity, reward prediction error, time‐frequency analyses

Impact

Understanding reward processing requires examining not just how outcomes are evaluated, but how expectations are formed beforehand. Using a combination of traditional ERP and time‐frequency analyses, we show that the neural mechanisms underlying these two stages are oftentimes distinct from one another, underscoring the importance of treating them as distinguishable events rather than a single unified process. Critically, time‐frequency measures captured nuances that time‐domain ERPs alone would have missed, underscoring the utility of multimethod approaches in reward research.

1. Introduction

Expectations about future outcomes influence our interactions with the world around us. Accurate expectations can maximize the likelihood of future positive outcomes (i.e., rewards) and minimize the potential for negative outcomes (i.e., losses); conversely, inaccurate expectations may result in missed opportunities that signal the need for further learning. Consequently, cognitive scientists and neuroscientists have spent considerable effort studying the ways in which expectations shape how humans and other animals process reward outcomes (for reviews, see: Glazer et al. 2018; San Martín 2012). Far less is known, however, about the neural mechanisms underlying the formation of reward expectations. Here, we separate the formation of reward expectations from the processing of subsequent outcomes, while assessing the underlying neural dynamics using electroencephalography (EEG).

Many studies have leveraged the high temporal resolution of EEG recordings of brain activity to investigate the cognitive events evoked in rapid succession during reward outcome processing. Most such work has focused on the event‐related potential (ERP) response known as the Reward Positivity (RewP), which has been thought to reflect a reward prediction error that contributes to reinforcement learning (Krigolson 2018; San Martín 2012; Walsh and Anderson 2012). The RewP is a positive‐polarity deflection in response to feedback indicating a win (vs. a loss) that typically peaks around 200 to 350 ms over frontocentral electrodes (Proudfit 2015). Many studies have shown that the RewP is typically larger for unexpected outcomes and larger still for unexpected wins (Holroyd et al. 2008, 2011), thus aligning with its theoretical interpretation as a reward‐prediction‐error signal. The RewP is thought to be primarily generated from the anterior cingulate cortex (ACC). This association was initially based on EEG‐based source localization analyses (Gehring and Willoughby 2002) and is now further supported from studies using fMRI (Amiez et al. 2013), simultaneous EEG/fMRI (Becker et al. 2014; Hauser, Iannaccone, Ball, et al. 2014; Hauser, Iannaccone, Stämpfli, et al. 2014) and intracranial EEG in humans (Oerlemans et al. 2025; Smith et al. 2015). As such, the RewP has become a useful index of ACC activity associated with reward outcomes, and studies of this component have played a substantial role in the refinement of theories concerning the role of the ACC in learning (Alexander and Brown 2011; Holroyd and Coles 2002; Shenhav et al. 2013).

Another key ERP component related to outcome processing is the longer‐latency P3b (see Polich 2007, for a general review of the P3b). The reward P3b is thought to reflect the attention‐driven evaluative processing and updating of outcome‐related information (Glazer et al. 2018; San Martín 2012; San Martín et al. 2013), with a greater updating process generating a larger P3b. The P3b follows the RewP in time (latency ~350–600 ms) and has a maximal distribution at electrode locations over the parietal scalp. While prior studies report inconsistent effects of outcome valence (i.e., wins vs. losses) on P3b amplitude, other key factors such as greater magnitude outcomes and higher risk consistently elicit larger P3b responses (Chandrakumar et al. 2018; San Martín 2012; San Martín et al. 2013). The P3b is also sensitive to expectations such that it is typically larger for unexpected wins (Hajcak et al. 2005; Wu and Zhou 2009), consistent with its postulated role in an information‐updating process following feedback and in predicting future behavioral adjustments (Chandrakumar et al. 2018; San Martín 2012; San Martín et al. 2013).

Most EEG studies of reward processing have focused exclusively on the reward outcome processing stage, rather than other important stages of reward processing; this imbalance has been criticized in several recent review papers (Glazer et al. 2018; Meyer et al. 2021). For example, many studies have detailed how reward expectations can influence reward outcome processing, but very little work has investigated the neural processing of cues that form expectations. Moreover, the formation of expectations will necessarily influence the subsequent evaluation of rewards (e.g., as better or worse than expected), yet there is little reason to assume that a single neural mechanism supports both processes; for example, reward‐related cues evoking the RewP and P3b are not significantly correlated with their outcome‐evoked counterparts (Novak et al. 2016; Novak and Foti 2015; Pornpattananangkul and Nusslock 2015). Thus, there is a pressing need to understand the mechanisms by which expectations are formed, including the underlying neural processes.

To our knowledge, only two studies have specifically investigated the processing of cues that create reward expectations (Yu et al. 2011; Zheng et al. 2020). These studies primarily focused on the degree of information about the likely outcome (i.e., certainty), and the likely valence of the outcome (i.e., win or loss). However, these studies evaluated only a small subset of neural processes, subsequently limiting our broader understanding of these mechanisms. For example, Yu et al. (2011) compared the expectation generation RewP and outcome‐evoked RewP but did not report on other ERPs such as the longer‐latency P3b, leaving unanswered questions about which computations are uniquely supported by the RewP. Zheng et al. (2020) studied both the RewP and P3b during the expectation cue, but did not investigate any components from the outcome feedback stage, thereby preventing comparisons between different stages. In the present study, we conducted a more comprehensive examination of both expectation formation processing and final outcome processing. This goal was pursued by combining ERP analysis approaches with time‐frequency analyses, a relatively underutilized approach that can yield additional insights into associated cognition beyond that which can be gained from traditional ERPs alone.

Measuring and disentangling the RewP and P3b components is challenging due to their temporal overlap, as the latter part of the earlier‐latency RewP component can be difficult to distinguish from the rise of the later P3b. However, time‐frequency analyses have demonstrated that the RewP and the P3b‐ERPs reflect distinct combinations of frontocentral theta (3–7 Hz) and parietal delta (0.5–3 Hz) activity (Bernat et al. 2011, 2015; Cavanagh 2015; Cavanagh et al. 2012; Foti et al. 2011; Watts et al. 2017). Moreover, the correlation between the RewP and P3b‐ERP is much higher than the correlation between frontocentral theta and parietal delta (Bernat et al. 2011, 2015), suggesting that these time‐frequency measures may better separate the constituent underlying processes compared to the time‐domain ERPs.

Several time‐frequency studies have shown that delta and theta are sensitive to different aspects of reward outcome processing. Theta activity is generally larger for losses and sensitive to the primary aspects of outcomes (i.e., valence), while delta is typically larger for wins and can be sensitive to multiple secondary aspects of outcome, such as magnitude, reward context, and outcome history (Bachman et al. 2021; Bernat et al. 2008, 2011, 2015; Nelson et al. 2011; Watts and Bernat 2018). Two papers have also demonstrated that reward expectations influence outcome‐evoked delta activity more than theta (Watts et al. 2017; Zheng et al. 2020). However, the composition of these components during the formation of the reward expectations remains unknown, presenting a notable gap in our understanding. This gap is particularly evident given the limited ERP research on this topic of reward expectation formation, along with the conceptual and methodological limitations noted above.

The present study provides two key advances towards the understanding of reward expectations. First, we isolated the neural processing underlying the generation of reward expectations by separating that process from reward outcome evaluation in a two‐stage reward‐prediction task. In the initial Outcome Probability stage, a cue provided information about the likelihood of winning on that trial, thereby forming or shaping participants' expectations about winning or losing on that trial. Thereafter, in the Actual Outcome stage, participants learned whether they had won or lost on that trial. This design allows us to separate assessment of the formation of reward expectations from their subsequent influence on evaluation of outcomes. Second, we investigated the dynamic formation of expectations by examining both traditional time‐domain ERPs as well as more novel time‐frequency measures at each task stage.

To help clarify the goals of the analyses for this two‐stage reward‐processing paradigm, we have listed all of our main hypotheses in Table 1 below. The rationale and supporting papers for each hypothesis can be seen in Appendix S1. Some of our main hypotheses were as follows: during the Outcome Probability screen, we expected to see that the RewP‐ERP elicited by the cue would be larger for when that cue indicated that a win was certain (i.e., 100% confident) relative to all other predicted outcomes, and that the P3b‐ERP would be larger for certain outcomes (i.e., 100% or 0% confident). We additionally hypothesized that theta activity would be highest for certain losses (i.e., 0% confident), while delta activity would be largest for certain wins. During the Actual Outcome stage, we hypothesized that the RewP‐ERP and P3b‐ERP would be largest for unexpected wins. For our time‐frequency measures at this stage, we hypothesized that delta activity would be larger for unexpected wins and that theta activity would be larger for losses.

TABLE 1.

Hypothesized results.

Neural component Stage
Stage 1: Outcome Probability Stage 2: Actual Outcome
RewP
ERP Largest for certain wins Largest for unexpected wins
Theta Largest for certain losses Largest for losses; insensitive to expectations
Delta Largest for certain wins Largest for unexpected wins
P3b
ERP Largest for certain outcomes Largest for unexpected wins
Theta Largest for certain losses Largest for losses; insensitive to expectations
Delta Largest for certain outcomes Largest for unexpected wins

Note: The hypothesized effects for each analysis.

2. Methods

2.1. Participants

Thirty‐six healthy volunteers participated in this study. Data from four of these participants were rejected from further analysis due to the presence of excessive EEG artifacts (described below in EEG acquisition and processing) that led to the rejection of > 50% of the epochs. Thus, our final sample consisted of 32 participants (Age: 21.1 ± 3.6 years; 20 females). Participants gave their written consent before the study began. At the end of the experiment, they received both a base compensation, which consisted of either 1 course credit/h or a monetary payment of $15/h, along with an additional monetary bonus that was related to their performance on the task (average bonus: $12 ± $4.8, ranging from $2 to $23). All participants had normal color vision (as tested by Ishihara 1925) and normal or corrected‐to‐normal visual acuity. The study was approved by the Duke Campus Institutional Review Board.

2.2. Design of Two‐Stage Reward Prediction Task

Each trial began with a choice screen, where a green circle and a purple circle were presented, stacked one above the other on the vertical meridian of the screen, above and below a central fixation cross (Figure 1). Participants had up to 1.5 s to select which colored circle they thought would win on that trial by pressing one of the two right‐sided shoulder buttons on a game controller with their right hand. The computer acknowledged their decision by highlighting their choice on the screen for a duration varying randomly between 800 ms and 1000 ms. The respective position of the colored circles relative to the fixation cross (i.e., above vs. below) was randomized for each trial.

FIGURE 1.

FIGURE 1

Experiment design. On each trial, participants selected which of two colored circles (i.e., green or purple) they thought would win on that trial. After making their selection, they received information about the outcome in two stages. First, they were presented with the Outcome Probability screen, which displayed the probability that each color would win on that trial, as depicted by the ratio of green to purple colored circles. In the example in Figure 1, the ratio of colored circles indicates that there was a 75% chance that green would win and a 25% chance that purple would win on this trial. In the second stage of outcome‐related information (i.e., the Actual Outcome screen), participants learned the actual outcome of the trial when the color was removed from seven of the eight circles on the Actual Outcome screen, leaving the winning color circle (purple here).

The participants then received information about the outcome of the trial in two temporally separate stages. First, they saw an Outcome Probability screen, which informed participants how probable it was that their chosen color would win on that trial. This screen consisted of eight circles that were each colored either green or purple and were arranged in a circle around the fixation cross. The ratio of the number of green circles and purple circles indicated the likelihood as to which color would win when the actual outcome of that trial would be revealed (e.g., six purple circles and two green circles indicated a 75% likelihood of winning if purple had been chosen and a 25% likelihood of winning if green had been chosen, etc.). The possible set of ratios was set in increments of two circles and included the possibility that all of the circles could be the same color (e.g., 8 green circles and 0 purple circles, indicating a 100% and 0% likelihood of winning for each respective color choice). The Outcome Probability screen displayed for a random duration of 1000 to 1200 ms before transitioning to the Actual Outcome stage, in which participants learned the winning color for that trial. We purposefully designed the transition between these two screens to ensure that participants could find the actual outcome at a fast and consistent rate, without unnecessary changes in sensory information. The actual outcome information was indicated by removing the color from seven of the circles (i.e., seven changing to white from purple and green), such that the remaining colored circle of the eight indicated the winning color for that trial. The Actual Outcome screen lasted for 800 ms before moving on to an intertrial interval blank screen, which was a fixation cross only and remained on the screen for 800 to 1000 ms until the next trial.

Participants gained one point for selecting the winning color on a trial and lost one point for selecting the losing color. Pressing a button on any screen besides the choice screen would also lose one point and display a warning. Participants were told that these points would accumulate throughout the experiment and be converted into a monetary bonus at the end of the study, but were not told the conversion factor. The bonus money was determined by dividing the total number of points by ten and rounding up to the next dollar.

Before beginning the main task, participants received thorough instructions about the experimental procedures; in addition, most participants chose to complete an optional 10‐trial practice run. The main task was conducted in 28 blocks of 20 trials each and took a total of ~45 min runtime to complete, excluding optional breaks. Participants were told that, in each block, one of the colors would be set to win somewhat more frequently than the other, and that identifying that color could lead to more points and a larger bonus payment. The winning ratios were always set to 70%/30%, with the more frequently winning color in each block determined randomly at the start of the block. While each block was designed to favor one color over the other, it also included trials where the information provided during the Outcome Probability stage indicated that the unfavored color was more likely to win on that trial. The experiment was programmed and presented using the Presentation software suite (Neurobehavioral Systems Inc., Berkeley, CA, www.neurobs.com).

2.3. EEG Acquisition and Processing

Each participant sat 60 cm away from a 61‐cm monitor in a dimly lit, sound‐attenuated, electrically shielded room. Participants were fit with an elastic EEG cap equipped with 64 Ag/AgCL active electrodes (actiCAP; Brain Vision LLC, Morrisville, NC). The caps were custom designed for extended coverage of the full head, from slightly above the eyebrows to below the inion posteriorly (Woldorff et al. 2002). Electrode impedances were kept below 15 kΩ and referenced to the right mastoid during recording. The EEG signal was recorded using a three‐staged cascaded integrator‐comb filter with a corner lo‐pass value of 130 Hz and sampled at a rate of 500 Hz per channel (actiCHamp; Brain Vision LLC Morrisville, NC). Participants were instructed to maintain their fixation on a central cross that was present throughout the entire experiment. Each participant's eye movements were monitored using vertical and horizontal EOG channels for later artifact correction and rejection, along with a closed‐circuit zoom‐lens camera.

Offline processing of the EEG data was conducted using the EEGLAB Toolbox (Delorme and Makeig 2004) in MATLAB (MATLAB and Statistics Toolbox Release 2016a, The MathWorks Inc., Natick, Massachusetts 2016). First, a 40‐Hz non‐causal low‐pass filter was applied to the data. The data were then down‐sampled to 250 Hz per channel and filtered with a 0.01 Hz non‐causal high‐pass filter. Data were then re‐referenced to the algebraic average of the two mastoid electrodes. Independent component analysis (ICA) was applied to correct for artifacts from eye movements and eye blinks. Electrodes that had lost good connectivity during the recording or were excessively noisy, as determined by visual inspection of the data, were excluded from that individual participant's dataset before the ICA correction. After correcting for such eye artifacts, any electrode channels that had been previously removed for being excessively noisy were reintroduced back into the dataset using a spherical‐spline interpolation solution (Perrin et al. 1989). The artifact‐corrected data were then segmented from −1000 ms before to +2000 ms after the onset of both the Outcome Probability screen and the Actual Outcome screen, and these sets of epochs were each baseline corrected from −200 to 0 ms. Epochs that still contained residual eye blink or horizontal eye movement artifacts were then excluded, using two separate step functions that identified rapid changes in EEG amplitude associated with such eye activity. Lastly, any epochs in which activity in any channel exceeded ±100 μV were removed from analysis. Further information about the retention rate of epochs across each condition can be found in Appendix S2.

2.4. Behavioral Measures

Our behavioral measures consisted of choice time and choice accuracy as a function of the trial's position within each block. Choice accuracy was calculated as whether the participant chose the optimal color (i.e., the one most likely to win in that block), regardless of whether that individual trial resulted in a win or loss. Choice time was calculated as the time between the onset of the choice screen and the participant's selection. These measures were averaged over the trial's position within each block (e.g., the average choice time and choice accuracy for every 1st trial of each block of 20, every 2nd trial of each block of 20, and so on).

2.5. ERP Measures

For each ERP component we took the mean activity across electrode clusters that have been commonly associated with that ERP. For the RewP we analyzed data in a cluster of 7 electrodes centered around the frontocentral site FCz, where this component tends to be largest (Krigolson 2018); the P3b was extracted from a cluster of 7 electrodes centered around site Pz, where it tends to be largest (Polich 2007; Chandrakumar et al. 2018).

The activity for each ERP component was calculated as the mean amplitude across a time‐window. The start and end points of this time‐window were set by first collapsing the data across certain conditions and inspecting the resulting averages. We took this approach because we expected that some of the different features of the Outcome Probability and Actual Outcome screens might lead to inherent differences in the timing of various ERP components. For example, the structure of the Actual Outcome screen is likely to trigger a shift in spatial attention that is not required for the Outcome Probability screen, thereby potentially delaying some of the ERP components.

We calculated the mean activity for the RewP for each valence by first separately collapsing the data across probabilities. As such, we derived the mean response amplitude for a win and then subtracted the corresponding mean response amplitude for a loss (again, for each of the two informational stages). This approach is commonly used for the RewP to reduce the influence of other temporally nearby ERPs that are commonly evoked in reward processing paradigms (e.g., the P200 and P3b; Krigolson 2018; Luck 2014). Furthermore, we were not interested in studying the outcome valence of the RewP per se, which has been robustly shown to be larger for wins than losses (Proudfit 2015; San Martín 2012) but were instead interested in whether certainty/expectations could modulate the RewP; similarly, we were interested in whether certainty/expectations also interacted with outcome valence. Using this approach, we determined that the RewP occurred on average between 232 and 332 ms following the presentation of the Outcome Probability screen (derived, as noted above, after collapsing across the various outcome possibilities), and between 252 and 452 ms following the presentation of the Actual Outcome screen (similarly collapsed). As noted above, the later time window for the RewP evoked by the Actual Outcome screen likely resulted from the need to shift one's attention to the winning color before being able to fully process it.

For the P3b component, we collapsed the data across all conditions, given that this component has not generally shown a consistent relationship with outcome valence (Chandrakumar et al. 2018). Based on this collapsed data, the time window for the P3b was set between 400 to 800 ms in response to the Outcome Probability screen and 500 to 800 ms in response to the Actual Outcome screen.

2.6. Time‐Frequency Activity Measures

Our time‐frequency analyses employed a similar approach to that of prior time‐frequency studies of reward processing (e.g., Bachman et al. 2021; Watts et al. 2017). For each participant we time‐lock‐averaged the EEG data separately for each condition and electrode. We then applied time‐frequency decompositions to these averaged signals, extracting evoked (i.e., phase‐locked) time‐frequency activity using a binomial reduced interface distribution variant of Cohen's class of time‐frequency transforms, which provided a uniform resolution in both time and frequency (Bernat et al. 2005). The time‐frequency decompositions were created with a resolution of 64 time bins per second and 2 frequency bins per Hz. From these time‐frequency representations we extracted averaged power regions of interest (ROI) based on a priori time, frequency, and spatial location characteristics. For the time dimension we selected windows that most closely approximated the start and end time points used for the RewP and P3b in each phase. Delta and theta measures were distinguished based on their relative frequency range. The activity occurring between 0.5 to 3.5 Hz was defined as delta, while the activity occurring between 4.5 to 7.5 Hz was defined as theta. A 1‐Hz gap was purposefully left between these two frequency bands to optimize our ability to distinguish differences in their relationships with experimental conditions. For delta we extracted activity from the parietal cluster of electrodes used to measure the P3b, while for theta we extracted activity from the frontocentral cluster of electrodes used to measure the RewP.

2.7. Behavioral Analyses

We conducted several behavioral analyses to assess how quickly participants learned which color was most likely to win in that block. To do this, we used two generalized linear mixed‐effects models using two dependent variables: proportion of optimal choices (i.e., selecting the color that was most likely to win in that block) and choice time. For predictors, we input which color they chose on each trial (green or purple) as well as the trial number within each block (1 through 20). Similar studies of probabilistic learning have shown that learning across trials is not necessarily linear (van den Berg et al. 2019). To account for the possibility of non‐linear learning across the block, we fit both linear and quadratic polynomial terms to the trial number within each block. The mixed‐effects models were implemented using the glme function in MATLAB with random slopes and intercepts. Statistical significance was defined as p < 0.05, and any statistical tests that resulted in a p‐value < 1 × 10−10 are reported at that level, given the limitations on the precision of our statistical analyses.

2.8. Neural Analyses

We assessed the sensitivities of the various neural measures to outcome certainty and expectations by submitting them to repeated‐measures analyses of variance (rmANOVAs) using JASP 0.96 (JASP Team 2026). More specifically, we analyzed these components' sensitivities to Outcome Certainty using a 2 × 2 rmANOVA ([(outcome Valence: win vs. loss) × (Certainty: Certain and Uncertain)]). The Certain condition at the Actual Outcome event consisted of trials where the outcome had already been determined on the prior Outcome Probability screen (i.e., 100% win or 0% loss). The Uncertain conditions were collapsed together across the remaining probabilities that still pointed towards a likely valenced outcome (25% or 75% chance of winning). The 50% condition was excluded from this analysis as it did not give participants an expectation about whether they would likely win or lose.

Our analyses of outcome Expectancy utilized 2 × 3 rmANOVAs [(outcome Valence: win vs. loss) × (Expectancy: expected outcome, 50/50 outcome, or unexpected outcome)]. Significance for each rmANOVA was defined as p < 0.05. Significant main effect terms were followed with pairwise comparisons while significant interaction terms were followed with paired sample t‐tests that were corrected for multiple comparisons using the Holm‐Bonferroni method (Aickin and Gensler 1996; Holm 1979). All p‐values that were corrected in this manner are indicated as p HB .

3. Results

3.1. Behavioral Results

We first fit a generalized linear mixed‐effects model to understand how the selected color (green or purple) and trial number within a block (1 to 20) predicted the participant's likelihood of choosing the color that was more likely to result in a win (i.e., the optimal choice). These results can be seen in Figure 2A. Results showed a significant effect of trial number for both the linear term (slope = 0.22, t(17,773) =12.03, p < 1 × 10−10) and the quadratic term (slope = −0.01, t(17,773) = −8.15, p < 1 × 10−10). This indicated that participants were able to learn to identify and choose the optimal color over the course of that block. As expected, we found no significant effect of whether the green or purple circle was the chosen color (slope = 0.03, t(17,773) = 0.25, p = 0.804), and we found no significant interactions between the chosen color and the linear term (slope = −0.02, t(17,773) = −0.79, p = 0.429) or quadratic term (slope = 0.00, t(17,773) = 0.72, p = 0.470), confirming that there were no effects of color upon the participants' ability to learn the optimal option.

FIGURE 2.

FIGURE 2

Behavioral Results. (A) Proportion of optimal choices as a function of when the trial occurred within a block. Participants tended to learn which color was most likely to result in a win over the first half of each block and then plateau at 80% optimal choices for the remaining trials. (B) Participants took longer to make a choice on the first trial before making choices at a consistent speed across the remainder of the block.

We then set up an analogous generalized linear mixed‐effects model for choice time (Figure 2B). There was a significant effect of trial number for both the linear (slope = −26.06, t(17,706) = −18.55, p < 1 × 10−10) and quadratic term (slope = 1.05, t(17,706) = 16.13, p < 1 × 10−10). Visual inspection of the data led us to conduct an exploratory analysis using a model that excluded the first trial. The resulting linear and quadratic terms were no longer significant (p's ≥ 0.600), indicating that this effect was indeed being primarily driven by the first trial. There was no significant effect of the chosen color (slope = −1.85, t(17,706) = −0.20, p = 0.843), and no significant interaction between chosen color and trial number (color choice × trial number linear term slope = 1.43, t(17,706) = 0.72, p = 0.483; color choice × trial number quadratic term slope = −0.10, t(17,706) = −1.10, p = 0.273). In summary, participants were slow to make their first choice in a block before continuing at a faster but consistent speed for the remainder of the block.

3.2. Outcome Probability: RewP‐ERP and Time‐Frequency Components

During the Outcome Probability screen, we extracted the RewP‐ERP (Figure 3A) as well as the corresponding RewP‐Delta and RewP‐Theta time‐frequency regions of interest (Figure 3B) for each condition. We then submitted each of these measures to 2 × 2 rmANOVAs of the predicted Valence of the final outcome (Win vs. Loss) by the degree of Certainty about that outcome (Certain outcome: 100% and 0%; Uncertain outcome: 75% and 25%). The full results of these rmANOVAs are summarized in Table 2 and described in greater detail below.

FIGURE 3.

FIGURE 3

RewP activity during the Outcome Probability screen. (A) Raw ERPs extracted during this stage, where the time range extracted is indicated using a shaded vertical yellow bar. The cluster of electrodes used for the plots and statistical analyses are indicated with a purple circle. (B) Averaged time‐frequency (TF) activity during the same period, plotted at electrode FCz. Dashed boxes indicate the time and frequency ranges for the Delta and Theta ROIs. The plots below indicate the spatial distribution of these components. Delta power exhibited a more centroparietal distribution, while Theta was largest over frontocentral sites. (C) Results of the 2 × 2 rmANOVA, where the RewP was separately larger for predicted wins and Certain outcomes. (D) rmANOVA for Theta, which demonstrated a selective increase in activity for Certain losses. (E) Results for the rmANOVA for Delta, which was larger for Certain outcomes over Uncertain ones.

TABLE 2.

Statistical summary of RewP activity during the Outcome Probability screen.

Model term description rmANOVAs F‐values
RewP‐ERP RewP‐Theta RewP‐Delta
Valence 5.04* 7.59* 1.49
Win > Loss Loss > Win
Certainty 6.13* 7.58* 29.35***
Certain > Uncertain Certain > Uncertain Certain > Uncertain
Interaction 1.03 10.00* 0.68

Certain: Loss > Win

Uncertain: No difference.

Note: The F‐values for statistically significant terms are bolded. *p < 0.05; **p < 0.01; ***p < 0.001. Degrees of freedom for each term were (1, 31).

For the RewP‐ERP (Figure 3C; Table 2, left) we found a significant effect of outcome Valence (p = 0.032), indicating that, as expected, the RewP‐ERP was larger for a screen predicting an eventual win relative to one predicting an eventual loss. There was also a significant effect of Certainty (p = 0.019), indicating that the RewP was larger for Certain outcomes (i.e., 0% or 100%) vs. Uncertain ones (75% or 25%). Notably, there was no significant interaction between outcome Valence and Certainty (p = 0.317; i.e., the lines in Figure 3C are parallel). Thus, the RewP was separately sensitive to both outcome Valence (higher for predicted wins) and Certainty (higher for Certain outcomes), with the largest RewPs thus being generated by guaranteed wins.

We next turned to the rmANOVA for the average Theta activity matched to the time range and spatial region used for the RewP‐ERP extracted during the Outcome Probability screen (Figure 3D; Table 2, center). We found a main effect of Valence (p = 0.01), where RewP‐Theta activity was larger for predicted losses. A significant main effect of Certainty (p = 0.01) also suggested greater activity for Certain outcomes. There was also a significant interaction term (p = 0.003; non‐parallel lines in Figure 3D). Post hoc tests revealed there were only significant differences between wins and losses for the Certain condition (t(31) = −3.27, p HB  = 0.003) but not the Uncertain condition (t(31) = −0.45, p HB  = 0.659). This was driven by an increase in Theta activity from Uncertain losses (i.e., when there was only a 25% chance of winning) to Certain losses (t(31) = 4.30, p HB < 0.001). In conclusion, changes in RewP‐Theta activity were driven by predicted losses but only when the predicted loss was Certain. Notably, this interaction effect for win/loss vs. Certainty/Uncertainty can be seen for the Theta activity during the RewP time range, but not for the RewP‐ERP during the same time period, suggesting that the time‐frequency analysis revealed additional information not provided by the traditional ERP analyses (Figure 3C vs. D).

For RewP‐Delta in the Outcome Probability stage (Figure 3E; Table 2, right), we found no main effect of Valence (p = 0.23), but there was a significant main effect of Certainty (p < 0.001), where activity was larger for Certain outcomes. There was no significant interaction term (p = 0.416; non‐parallel lines in Figure 3E). Thus, Delta activity in the RewP time‐range was only driven by Certainty and not by predicted‐outcome Valence, or by interactions with predicted‐outcome Valence.

3.3. Outcome Probability: P3b‐ERP and Time‐Frequency Components

We next conducted analogous 2 × 2 rmANOVAs on neural activity in the P3b time range during the Outcome Probability screen. This included the P3b‐ERP (Figure 4A) as well as its corresponding P3b‐Theta and Delta time‐frequency ROIs (Figure 4B). Table 3 presents a summary of the rmANOVAs, which will be elaborated upon in the paragraphs below.

FIGURE 4.

FIGURE 4

P3b activity during the Outcome Probability screen. (A) Raw ERPs extracted during this stage, where the purple circle indicates the electrodes used for plotting and statistical analyses. The time range extracted is indicated using a shaded vertical yellow bar. (B) Averaged time‐frequency (TF) activity during the same period, plotted at electrode Pz. Dashed boxes indicate the time and frequency ranges for the Delta and Theta ROIs. Delta power exhibited a more centroparietal distribution, while Theta was largest over frontocentral sites. (C) Results of the rmANOVA, where the P3b‐ERP was larger for Certain losses. (D) The rmANOVA for Theta, which revealed a significant increase in Theta activity for Certain losses. (E) Results for the rmANOVA for Delta, which was particularly larger for Certain losses.

TABLE 3.

Statistical summary of P3b activity during the Outcome Probability screen.

Model term description rmANOVAs F‐values
P3b‐ERP P3b‐Theta P3b‐Delta
Valence 18.79*** 8.56** 21.61***
Loss > Win Loss > Win Loss > Win
Certainty 137.90*** 0.84 19.80***
Certain > Uncertain Certain > Uncertain
Interaction 11.54** 10.97** 24.94***

Certain: Loss > Win

Uncertain: No difference

Certain: Loss > Win

Uncertain: No difference

Certain: Loss > Win

Uncertain: No difference

Note: The F‐values for statistically significant terms are bolded. *p < 0.05; **p < 0.01; ***p < 0.001. Degrees of freedom for each term were (1, 31).

The 2 × 2 rmANOVA of Valence and Certainty for the Outcome Probability P3b (Figure 4C; Table 3, left) showed significant main effects for predicted‐outcome Valence (p < 0.001), where the P3b was larger for predicted losses. There was also a significant main effect of outcome Certainty (p < 0.001), indicating that the P3b was larger for predicted outcomes that were Certain. Critically, there was also a significant interaction effect of these two factors on the P3b (p = 0.002; non‐parallel lines in Figure 4C). Post hoc tests found that there was a significant difference between win and loss activity in the Certain condition (t(31) = −4.85, p HB < 0.001) but not between the Uncertain win and Uncertain loss (t(31) = −2.04, p HB  = 0.050). In sum, the P3b was larger for losses, especially Certain ones.

The rmANOVA for P3b‐Theta activity during the Outcome Probability screen (Figure 4D; Table 3, center) did not reveal a main effect of Certainty (p = 0.367), but there was a main effect of Valence that was driven by increased Theta activity for predicted losses (p = 0.006). There was also a significant interaction term (p = 0.002; non‐parallel lines in Figure 4D). Post hoc tests revealed a similar pattern as seen in the earlier‐latency RewP‐Theta activation: there was a highly significant difference between Certain loss and Certain win activity (t(31) = −3.71, p HB < 0.001) but no significant difference between Uncertain outcomes (i.e., likely win vs. likely loss; t(31) = −0.82, p HB = 0.417). This effect was driven by an increase in theta loss activity from Uncertain to Certain predicted outcomes (t(31) = −2.31, p HB  = 0.028). Thus, P3b‐Theta activity was driven by changes in Certain predicted wins vs. Certain predicted losses.

The rmANOVA for P3b‐Delta activity (Figure 4E; Table 3, right) revealed main effects of Valence (p < 0.001) that indicated increased activity for predicted losses, as well as a main effect of Certainty (p < 0.001) where Certain outcomes generated more P3b‐Delta activity. There was also a significant interaction term (p < 0.001; non‐parallel lines in Figure 4E). Post hoc tests found that there was a significant difference in P3b‐Delta activity between predicted Certain wins vs. predicted Certain losses (t(31) = −5.17, p HB < 0.001) but not for Uncertain ones (i.e., likely win vs. likely loss; t(31) = −2.03, p HB  = 0.051). Thus, P3b‐Delta activity increased for both predicted likely losses and Certain outcomes, with the strongest effects for Certain losses.

3.4. Actual Outcome: RewP‐ERP and Time‐Frequency Components

Our next set of analyses examined neural activity occurring in response to the presentation of the Actual Outcome screen, when participants learned the winning color for that trial. We utilized 2 × 3 rmANOVAs consisting of the outcome's Valence (win vs. loss) and the participant's expectations of the indicated outcome based on the preceding outcome probability screen (Expectancy: expected, 50/50%, and unexpected). We first applied these rmANOVAs to the RewP‐ERP (Figure 5A) as well as to the RewP‐Delta and Theta‐ROIs (Figure 5B). Table 4 summarizes the statistical results from each of the rmANOVAs, and the paragraphs below describe each analysis in greater detail.

FIGURE 5.

FIGURE 5

RewP activity in response to presentation of the Actual Outcome for the trial. (A) Raw ERPs extracted during this stage, where the purple circle indicates the electrodes used for plotting and statistical analyses. The time range extracted is indicated using a shaded vertical yellow bar. (B) Averaged time‐frequency activity during the same period, plotted at electrode FCz. Dashed boxes indicate the time and frequency ranges for the Delta and Theta ROIs. Delta power exhibited a more centroparietal distribution, while Theta was largest over frontocentral sites. (C) Results of the 2 × 3 rmANOVA, where the RewP was significantly smaller for more expected wins. (D) The Theta rmANOVA found no significant variation in this period. (E) Results for the rmANOVA for Delta, for which win activity was larger for more unexpected outcomes while loss activity was larger when there were equal expectations about winning or losing.

TABLE 4.

Statistical summary of RewP activity during the Actual Outcome screen.

Model term description rmANOVAs F‐values
RewP‐ERP RewP‐Theta RewP‐Delta
Valence 36.81*** 1.97 9.57**
Win >Loss Win > Loss
Expectancy 9.41*** 2.19 15.95***
Unexpected > Expected Unexpected > Expected
Interaction 15.97*** 1.33 12.46***

Win: Smallest for Expected

Loss: No difference

9

Win: Smallest for Expected

Loss: Largest for 50/50

Note: The F‐values for statistically significant terms are bolded. *p < 0.05; **p < 0.01; ***p < 0.001. Degrees of freedom for the Valence term were (1, 31), and (2, 62) for the Expectancy and interaction terms.

The rmANOVA for the RewP‐ERP during the Actual Outcome screen (Figure 5C; Table 4, left) revealed a significant effect of both Valence (p < 0.001) and Expectancy (p < 0.001), as well as a significant interaction between these two factors (p < 0.001; non‐parallel lines in Figure 5C). Post hoc tests found that the RewP for wins was sensitive to expectations, such that win activity was smallest for expected wins (Expected Win vs. 50/50 Win: t(31) = −4.60, p HB < 0.001; 50/50 Win vs. Unexpected Win: t(31) = −2.66, p HB  = 0.012). Conversely, RewP activity for losses did not differ across expectation levels (p HB s > 0.332). These results indicate that RewP activity was sensitive to expectations set up by the preceding Outcome Probability stage, but only when they involved win outcomes.

The rmANOVA for RewP‐Theta activity (Figure 5D; Table 4, center) during this Actual Outcome phase revealed no main effects of either Valence (p = 0.170) or Expectancy (p = 0.120), nor was there a significant interaction term (p = 0.271). This indicates that the neurocognitive processes reflected by the RewP‐Theta during this stage were unaffected by these factors of the final outcome.

For RewP‐Delta during the Actual Outcome stage (Figure 5E; Table 4, right) we first found significant main effects of Valence (p = 0.004) that indicated greater activity for wins. There was also a significant effect of Expectancy (p < 0.001), as well as a significant interaction term (p < 0.001; non‐parallel lines in Figure 5E). Post hoc tests revealed an interesting pattern of activity. Win activity was smallest when this outcome was expected but was not significantly different between 50/50 and unexpected wins (Expected Win vs. 50/50 Win: t(31) = −6.20, p HB < 0.001; 50/50 Win vs. Unexpected Win: t(31) = −0.07, p HB  = 0.943). Conversely, loss activity was largest for the 50/50 condition (Expected Loss vs. 50/50 Loss: t(31) = −2.51, p HB  = 0.035; 50/50 Loss vs. Unexpected Loss: t(31) = 3.22, p HB  = 0.009). Collectively, the effects on the RewP‐Delta activity for wins were driven by decreases in activity for expected outcomes, while loss activity was highest during the most ambiguous outcomes.

3.5. Actual Outcome: P3b‐ERP and Time‐Frequency Components

Our final set of analyses tested for how outcome Valence and outcome expectations modulated the ERP (Figure 6A) and time‐frequency components during the P3b latency (Figure 6B) in response to the presentation of the Actual Outcome screen. More specifically, we applied 2 × 3 rmANOVAs of the outcome's Valence (win vs. loss) and the participant's expectations of the indicated outcome based on the preceding outcome probability screen (Expectancy: expected, 50/50%, and unexpected) to each of these three neural measures. Table 5 presents a summary of the statistical output followed by a detailing of the precise pattern of results.

FIGURE 6.

FIGURE 6

P3b‐latency activity in response to presentation of the Actual Outcome screen. (A) Raw ERPs extracted during this stage, plotted from the cluster of electrodes encapsulated by the purple circle. The shaded vertical yellow bar indicates the time range extracted for the P3b. (B) Averaged time‐frequency activity during the same latency period, plotted at electrode Pz. Dashed boxes indicate the time and frequency ranges for the Delta and Theta ROIs. The spatial plots indicated a more parietal distribution for Delta and a more frontocentral distribution for Theta. (C) Results of the 2 × 2 rmANOVA on the P3b‐ERP. There was more P3b activity for unexpected compared to expected wins. (D) Corresponding rmANOVA for Theta, which was smaller for expected outcomes. (E) Results for the rmANOVA for Delta, where activity was smaller for wins and for expected outcomes.

TABLE 5.

Statistical summary of P3b activity during the Actual Outcome screen.

Model term description rmANOVAs F‐values
P3b‐ERP P3b‐Theta P3b‐Delta
Valence 18.17*** 0.39 20.38***
Loss > Win Loss > Win
Expectancy 15.03*** 3.33* 9.43***
Smallest for Expected Smallest for Expected Smallest for Expected
Interaction 13.01*** 0.92 3.08

Win: Smallest for Expected

50/50: Loss > Win

Note: The F‐values for statistically significant terms are bolded. *p < 0.05; **p < 0.01; ***p < 0.001. Degrees of freedom for the Valence term were (1, 31), and (2, 62) for the Expectancy and interaction terms.

The 2 × 3 rmANOVA of the P3b during the Actual Outcome screen (Figure 6C; Table 5, left) revealed a significant main effect of Valence (p < 0.001), which indicated that loss activity was generally larger (i.e., more ERP positivity) than that for wins. A significant main effect of Expectancy (p < 0.001) revealed that expected outcomes elicited less activity compared to the 50/50 and unexpected conditions. There was also a significant interaction term (p < 0.001; non‐parallel lines in Figure 6C). Post hoc comparisons revealed that P3b activity for wins grew larger as they became more unexpected (Expected vs. 50/50: t(31) = −4.66, p HB < 0.001; 50/50 vs. Unexpected: t(31) = −2.13, p HB  = 0.041). Loss activity was also larger in the 50/50 condition than for expected (t(31) = 2.54, p HB  = 0.049) and unexpected losses (t(31) = 2.43, p HB  = 0.049). Altogether, win activity scaled as a function of unexpectedness, while loss activity increased due to ambiguity.

We next conducted parallel 2 × 3 rmANOVAs for the P3b‐Theta component (Figure 6D; Table 5, center). Theta activity in this stage revealed no main effects of Valence (p = 0.538), but there was a main effect of Expectancy (p = 0.042). Post hoc comparisons revealed that Theta activity generally grew as the outcome was progressively more unexpected, but these comparisons did not survive correction for multiple comparisons (Expected vs. 50/50: t(31) = −2.11, p HB  = 0.086; 50/50 vs. Unexpected: t(31) = −0.92, p HB  = 0.365). There was no significant interaction between Valence and Expectancy (p = 0.406). In sum, the neurocognitive processes reflected by P3b‐Theta activity were sensitive to expectations but not outcome valence.

For P3b‐Delta during the Actual Outcome stage (Figure 6E; Table 5, right) we found significant main effects of Valence (p < 0.001) that indicated greater Delta activity for losses vs. wins. There was also a significant effect of Expectancy (p < 0.001), where activity was smaller for expected outcomes. However, there was no significant interaction between Valence and Expectancy (p = 0.053; mostly parallel lines in Figure 6E). Thus, Delta activity was independently smaller for wins and for expected outcomes.

3.6. Summary of the Predicted and Observed Results

Given the numerous effects examined, the tables below provide a summary of the hypothesized effects and corresponding observed results for the Outcome Probability (Table 6) and Actual Outcome stages (Table 7).

TABLE 6.

Comparison of hypothesized and observed results: Outcome Probability stage.

Neural component Hypothesis Observed results
RewP
ERP Largest for Certain wins As hypothesized
Theta Largest for Certain losses As hypothesized
Delta Largest for Certain wins Largest for Certain outcomes
P3b
ERP Largest for Certain outcomes Largest for Certain losses
Theta Largest for Certain losses As hypothesized
Delta Largest for Certain outcomes Largest for Certain losses

Note: Nuanced differences between the hypothesized and observed results are bolded.

TABLE 7.

Comparison of hypothesized and observed results: Actual Outcome stage.

Neural component Hypotheses Observed results
RewP
ERP Largest for unexpected wins As hypothesized
Theta Largest for losses; insensitive to expectations Insensitive to all factors
Delta Largest for unexpected wins As hypothesized
P3b
ERP Largest for unexpected wins As hypothesized
Theta Largest for losses; insensitive to expectations Sensitive to expectations and not outcome valence
Delta Largest for unexpected wins Largest for unexpected outcomes

Note: Nuanced differences between the hypothesized and observed results are bolded.

4. Discussion

In this study, we examined the temporal dynamics of how reward expectations are formed and subsequently shape the processing of the reward outcomes. While the paragraphs below describe each of our findings in greater detail, here we highlight two key advances. First, we show that these two stages of information processing are distinct and modulated by different forms of information: expectation formation is most strongly shaped by more certain information, especially for losses, whereas outcome processing is most responsive to unexpected outcomes, especially for wins. Second, we demonstrate that traditional time‐domain ERPs and evoked time‐frequency measures provide distinct insights into the neural mechanisms underlying these processes. While parietal delta activity often mirrored the patterns observed in its temporally matched ERP, frontocentral theta frequently captured unique sensitivities to reward information. Together, these findings provide a comprehensive overview of the neural processes through which reward expectations are formed and how they influence subsequent outcome processing. These findings collectively demonstrate the value of broadening one's approach to investigating reward processing, both in terms of which stages are examined and the methodological tools used to uncover their respective neural mechanisms.

We first investigated neural activity underlying the RewP in response to the Outcome Probability screen, where participants were given information that formed their expectations about their likelihood of winning or losing on that trial. All our measures (the RewP‐ERP, RewP‐Delta, and RewP‐Theta) were larger for predicted outcomes that were certain vs. uncertain. These results suggest that participants treated these certain outcomes as if they had already won or lost, since this information was fully predictive of what they would experience during the subsequent Actual Outcome stage. Our ERP findings here align with that of Yu et al. (2011), who found that the cue‐RewP was larger for more certain outcomes and in particular, certain wins. While all of our measures in response to the Outcome Probability screen displayed consistent relationships with certainty, they diverged in their specific relationship to predicted outcome valence. More specifically, there was more RewP‐ERP activity when a win was likely, greater RewP‐Theta activity when a loss was likely, and no change in RewP‐Delta to either likely wins or losses. Collectively, these results indicate that while certainty engages a common evaluative process, that process diverges based on whether the anticipated outcome seems promising or bleak.

Frontocentral theta was also the only component at the Outcome Probability stage to be sensitive to the interaction of these factors, displaying more activity for Certain losses but not significantly differentiating between Uncertain losses and Uncertain wins. This finding aligns with a large body of work reporting theta's particular sensitivity to losses but less so to other secondary factors such as the magnitude of the reward or its relative context (Bernat et al. 2015; Watts and Bernat 2018). However, finding that only Certain losses elicited an increased theta response suggests that this component does not treat all loss information in the same manner. Instead, our findings suggest that likely losses are treated as if they can still generate a win, even when the odds are stacked against them. Notably, we only identified an interaction between outcome valence and certainty for the frontocentral RewP‐Theta and not for the RewP‐ERP. This may reflect one of the purported advantages of these oscillatory analyses, in that they isolate targeted signals from a mixture of other signals that normally get amalgamated together in an averaged ERP. This is particularly relevant when discussing the RewP‐ERP and frontocentral theta, both of which are thought to be primarily generated from the ACC (Oerlemans et al. 2025; Smith et al. 2015). However, a key methodological weakness of the RewP‐ERP is that it is a combination of both frontocentral theta activity and parietal delta activity (Bernat et al. 2011, 2015; Watts et al. 2017), with the source of latter having been modeled as arising from the basal ganglia (Foti et al. 2015). This may explain why our frontocentral theta measures here were able to pick up on an interaction between outcome valence and certainty that was otherwise lost by the RewP‐ERP. This methodological advantage may also help to adjudicate the various competing theories concerning ACC function, which have frequently relied on RewP‐ERP research (Alexander and Brown 2011; Holroyd and Coles 2002; Shenhav et al. 2013).

For the longer‐latency P3b evoked in response to the Outcome Probability screen, we found that all of the neural measures were larger for certain losses over certain wins, with no significant differences between uncertain loss and uncertain win outcomes. Both of the certain outcomes effectively marked an early end to that trial, but Certain losses may have also signaled that participants needed to adjust their future choices. Notably, our P3b‐latency delta results had a different relationship with predicted outcome valence than what we had hypothesized, showing more activity for a prediction of a likely loss instead of a likely win. One possible reason for this difference may lie within structural differences between our task and those reported in previous studies. Many past reports of delta activity used tasks where each trial was independent from the rest, while our experimental design included a learning component over a block. Parietal delta activity has been hypothesized to reflect, at least in part, the overall value of the decision‐making environment (Bachman et al. 2021; Watts and Bernat 2018). Our behavioral results suggest that participants could learn the more rewarding color rather quickly; consequently, any form of loss feedback may have been surprising to participants and suggested the need to further refine their understanding of the task. This greater allocation of attention towards evaluating their circumstances may have been reflected in an elevated amount of parietal delta activity; however, this finding will need to be confirmed by future work.

Our next set of analyses targeted neural activity in response to the Actual Outcome screen, starting with activity occurring in the RewP latency range. Our ERP measure in this latency range was found to be largest for unexpected wins, aligning with a large body of previous work suggesting that this component reflects a quantitative level of reward prediction error (Holroyd et al. 2008, 2011; Walsh and Anderson 2012). Notably, however, we found several surprising results within our time‐frequency measures. While RewP‐latency theta activity was insensitive to expectations (as hypothesized), we were surprised to find that it was also insensitive to the valence of the outcome. Similarly, we predicted that delta would be largest for unexpected wins, replicating past work (Watts et al. 2017), but we did not expect to find that delta loss activity would also be sensitive to expectations. Instead, we found that delta activity for losses was largest when expectations were the most ambiguous (i.e., when the Outcome Probability screen had been the 50/50 condition). Zheng et al. (2020) also found that delta activity may be sensitive to ambiguity but still reported higher activity for wins over losses. The difference seen here may once again be due to the learning component of our task, where participants could use their past choices to learn which colored circle would win more frequently in the block. As a result, any sort of feedback information—even negative outcomes—could be useful in informing their future choices. Within this view, the worst outcome would be a loss outcome with a 50/50 expectation, as it provides no reward while also providing the least amount of useful information that could help optimize future choices. While parietal delta has traditionally been associated with learning from positive outcomes rather than negative ones (Glazer et al. 2018), our findings suggest that its sensitivity to valence may vary depending on the decision‐making environment, and more specifically, which circumstances may motivate the greatest amount of motivated learning.

Our final set of analyses centered on the P3b latency range in response to the Actual Outcome screen. The P3b‐ERP was larger for unexpected wins, again replicating past work suggesting that these outcomes generated a larger updating of outcome‐related information, a process that has been linked to the P3b (Cohen et al. 2007; San Martín et al. 2013; Watts et al. 2017). However, the P3b‐ERP activity was also larger for predicted ambiguous outcomes, in a similar manner as the earlier‐latency RewP‐Delta activity. For the P3b‐latency Delta we had hypothesized that it would be larger for unexpected wins. While our results indicated that it was indeed larger for unexpected outcomes, it was also larger for losses rather than wins, and these two factors did not interact with one another. Our parietal delta results broadly demonstrated an inconsistent relationship with outcome valence, suggesting that this measure may reflect a more complex evaluative process that accounts for environmental circumstances, rather than a narrow, win‐related signal. Longer‐latency frontocentral theta was found to be larger for unexpected outcomes but was insensitive to outcome valence. Theta's lack of sensitivity to outcome valence at both the early and late time ranges of this final outcome stage may speak to a larger debate concerning the brain region from which it has been thought to originate, namely the anterior cingulate cortex (Foti et al. 2015; Tsujimoto et al. 2006). A major point of disagreement between several prominent theories of ACC function is whether it is sensitive only to violations of predictions regardless of outcome valence (e.g., Alexander and Brown 2011), or whether it is also sensitive to outcome valence (e.g., Holroyd and Coles 2002; Shenhav et al. 2013). Here we found that frontocentral theta was sensitive to predicted outcome valence only during the expectation generation stage, which may reflect a large, initial adjustment/formation of expectations for that trial. After this expectation is set, it may only be additionally sensitive to unexpected feedback, regardless of whether it is positive or negative. In other words, our findings suggest that this component is only sensitive to outcome valence when expectations are being set.

One limitation of our research is that we only focused on evoked (i.e., phase‐locked) time frequency activity and did not study any non‐phase‐locked activity. This is consistent with the same methodological choices made in most prior investigations of delta and theta activity during reward outcome processing (Bachman et al. 2021; Bernat et al. 2008, 2011, 2015; Ellis et al. 2018; Foti et al. 2015; Nelson et al. 2011; Watts et al. 2017, 2018; Watts and Bernat 2018). A major benefit of this approach is that evoked delta and theta can account for the majority of the variance in the time‐domain ERP RewP and P3b (Bernat et al. 2011, 2015; Watts et al. 2017), thus permitting a closer comparison of these two types of measures. However, this does not provide any insight into the relative contributions of non‐phase locked oscillations, and moreover, whether these contributions mirror or diverge from those seen in their evoked counterparts. This general limitation extends to most time‐frequency studies conducted in this field; consequently, future research should aim to report a more comprehensive comparison of these different time‐frequency measures.

In summary, our study offers a broader view of the neural mechanisms underlying both the formation of reward expectations and the processing of the eventual outcome in the context of those expectations. In many cases there was a correspondence between the results of one or both time‐frequency components to the temporally overlapping ERP, particularly between the long‐latency P3b‐ERP and the concurrent delta activity. However, time‐frequency measures also provided distinct insights that were not reflected in the temporally corresponding ERP. Finally, our findings suggest that the processes underlying the formation of expectations are distinguishable from those underlying the evaluation of outcomes, with distinct neural dynamics and different sensitivity to information.

Author Contributions

Matthew D. Bachman: methodology, software, data curation, formal analysis, investigation, visualization, writing – original draft. René San Martin: conceptualization, methodology, software, writing – review and editing. Kaya Scheman: data curation, investigation, writing – review and editing. Scott A. Huettel: conceptualization, supervision, resources, writing – review and editing. Marty G. Woldorff: conceptualization, supervision, resources, project administration, writing – review and editing.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Appendix S1: Supporting papers for each hypothesis.

Appendix S2: Proportion of retained epochs by block, condition, and stage.

PSYP-63-e70387-s001.docx (44.9KB, docx)

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.

References

  1. Aickin, M. , and Gensler H.. 1996. “Adjusting for Multiple Testing When Reporting Research Results: The Bonferroni vs Holm Methods.” American Journal of Public Health 86, no. 5: 726–728. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Alexander, W. H. , and Brown J. W.. 2011. “Medial Prefrontal Cortex as an Action‐Outcome Predictor.” Nature Neuroscience 14, no. 10: 1338–1344. 10.1038/nn.2921. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Amiez, C. , Neveu R., Warrot D., Petrides M., Knoblauch K., and Procyk E.. 2013. “The Location of Feedback‐Related Activity in the Midcingulate Cortex Is Predicted by Local Morphology.” Journal of Neuroscience 33, no. 5: 2217–2228. 10.1523/JNEUROSCI.2779-12.2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Bachman, M. D. , Watts A. T. M., Collins P., and Bernat E. M.. 2021. “Sequential Gains and Losses During Gambling Feedback: Differential Effects in Time‐Frequency Delta and Theta Measures.” Psychophysiology 20, no. 3: e13907. 10.1111/psyp.13907. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Becker, M. P. I. , Nitsch A. M., Miltner W. H. R., and Straube T.. 2014. “A Single‐Trial Estimation of the Feedback‐Related Negativity and Its Relation to BOLD Responses in a Time‐Estimation Task.” Journal of Neuroscience 34, no. 8: 3005–3012. 10.1523/JNEUROSCI.3684-13.2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Bernat, E. M. , Nelson L. D., and Baskin‐Sommers A. R.. 2015. “Time‐Frequency Theta and Delta Measures Index Separable Components of Feedback Processing in a Gambling Task.” Psychophysiology 52, no. 5: 626–637. 10.1111/psyp.12390. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Bernat, E. M. , Nelson L. D., Holroyd C. B., Gehring W. J., and Patrick C. J.. 2008. “Separating Cognitive Processes With Principal Components Analysis of EEG Time‐Frequency Distributions.” Advanced Signal Processing Algorithms, Architectures, and Implementations XVIII 70: 70740S. 10.1117/12.801362. [DOI] [Google Scholar]
  8. Bernat, E. M. , Nelson L. D., Steele V. R., Gehring W. J., and Patrick C. J.. 2011. “Externalizing Psychopathology and Gain–Loss Feedback in a Simulated Gambling Task: Dissociable Components of Brain Response Revealed by Time‐Frequency Analysis.” Journal of Abnormal Psychology 120, no. 2: 352–364. 10.1037/a0022124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Bernat, E. M. , Williams W. J., and Gehring W. J.. 2005. “Decomposing ERP Time–Frequency Energy Using PCA.” Clinical Neurophysiology 116, no. 6: 1314–1334. 10.1016/j.clinph.2005.01.019. [DOI] [PubMed] [Google Scholar]
  10. Cavanagh, J. F. 2015. “Cortical Delta Activity Reflects Reward Prediction Error and Related Behavioral Adjustments, but at Different Times.” NeuroImage 110: 205–216. 10.1016/j.neuroimage.2015.02.007. [DOI] [PubMed] [Google Scholar]
  11. Cavanagh, J. F. , Zambrano‐Vazquez L., and Allen J. J. B.. 2012. “Theta Lingua Franca: A Common Mid‐Frontal Substrate for Action Monitoring Processes.” Psychophysiology 49, no. 2: 220–238. 10.1111/j.1469-8986.2011.01293.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Chandrakumar, D. , Feuerriegel D., Bode S., Grech M., and Keage H. A. D.. 2018. “Event‐Related Potentials in Relation to Risk‐Taking: A Systematic Review.” Frontiers in Behavioral Neuroscience 12: 111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Cohen, M. X. , Elger C. E., and Ranganath C.. 2007. “Reward Expectation Modulates Feedback‐Related Negativity and EEG Spectra.” NeuroImage 35, no. 2: 968–978. 10.1016/j.neuroimage.2006.11.056. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Delorme, A. , and Makeig S.. 2004. “EEGLAB: An Open Source Toolbox for Analysis of Single‐Trial EEG Dynamics Including Independent Component Analysis.” Journal of Neuroscience Methods 134, no. 1: 9–21. 10.1016/j.jneumeth.2003.10.009. [DOI] [PubMed] [Google Scholar]
  15. Ellis, J. S. , Watts A. T. M., Schmidt N., and Bernat E. M.. 2018. “Anxiety and Feedback Processing in a Gambling Task: Contributions of Time‐Frequency Theta and Delta.” Biological Psychology 136: 1–12. 10.1016/j.biopsycho.2018.05.001. [DOI] [PubMed] [Google Scholar]
  16. Foti, D. , Weinberg A., Bernat E. M., and Proudfit G. H.. 2015. “Anterior Cingulate Activity to Monetary Loss and Basal Ganglia Activity to Monetary Gain Uniquely Contribute to the Feedback Negativity.” Clinical Neurophysiology 126, no. 7: 1338–1347. 10.1016/j.clinph.2014.08.025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Foti, D. , Weinberg A., Dien J., and Hajcak G.. 2011. “Event‐Related Potential Activity in the Basal Ganglia Differentiates Rewards From Nonrewards: Temporospatial Principal Components Analysis and Source Localization of the Feedback Negativity.” Human Brain Mapping 32, no. 12: 2207–2216. 10.1002/hbm.21182. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Gehring, W. J. , and Willoughby A. R.. 2002. “The Medial Frontal Cortex and the Rapid Processing of Monetary Gains and Losses.” Science 295, no. 5563: 2279–2282. 10.1126/science.1066893. [DOI] [PubMed] [Google Scholar]
  19. Glazer, J. E. , Kelley N. J., Pornpattananangkul N., Mittal V. A., and Nusslock R.. 2018. “Beyond the FRN: Broadening the Time‐Course of EEG and ERP Components Implicated in Reward Processing.” International Journal of Psychophysiology, Reward and Feedback Processing: State of the Field, Best Practices and Future Directions 132: 184–202. 10.1016/j.ijpsycho.2018.02.002. [DOI] [PubMed] [Google Scholar]
  20. Hajcak, G. , Holroyd C. B., Moser J. S., and Simons R. F.. 2005. “Brain Potentials Associated With Expected and Unexpected Good and Bad Outcomes.” Psychophysiology 42, no. 2: 161–170. 10.1111/j.1469-8986.2005.00278.x. [DOI] [PubMed] [Google Scholar]
  21. Hauser, T. U. , Iannaccone R., Ball J., et al. 2014. “Role of the Medial Prefrontal Cortex in Impaired Decision Making in Juvenile Attention‐Deficit/Hyperactivity Disorder.” JAMA Psychiatry 71, no. 10: 1165–1173. [DOI] [PubMed] [Google Scholar]
  22. Hauser, T. U. , Iannaccone R., Stämpfli P., et al. 2014. “The Feedback‐Related Negativity (FRN) Revisited: New Insights Into the Localization, Meaning and Network Organization.” NeuroImage 84: 159–168. [DOI] [PubMed] [Google Scholar]
  23. Holm, S. 1979. “A Simple Sequentially Rejective Multiple Test Procedure.” Scandinavian Journal of Statistics 47: 65–70. [Google Scholar]
  24. Holroyd, C. , and Coles M.. 2002. “The Neural Basis of Human Error Processing: Reinforcement Learning, Dopamine, and the Error‐Related Negativity.” Psychological Review 109, no. 4: 679–709. [DOI] [PubMed] [Google Scholar]
  25. Holroyd, C. B. , Krigolson O. E., and Lee S.. 2011. “Reward Positivity Elicited by Predictive Cues.” Neuroreport 22, no. 5: 249–252. 10.1097/WNR.0b013e328345441d. [DOI] [PubMed] [Google Scholar]
  26. Holroyd, C. B. , Pakzad‐Vaezi K. L., and Krigolson O. E.. 2008. “The Feedback Correct‐Related Positivity: Sensitivity of the Event‐Related Brain Potential to Unexpected Positive Feedback.” Psychophysiology 45, no. 5: 688–697. 10.1111/j.1469-8986.2008.00668.x. [DOI] [PubMed] [Google Scholar]
  27. Ishihara, S. 1925. “Tests for Colour‐Blindness, Handaya, Tokyo: Hongo Harukicho, 1917.” Google Scholar.
  28. JASP Team . 2026. JASP (Version 0.96) [Computer Software]. JASP Team. [Google Scholar]
  29. Krigolson, O. E. 2018. “Event‐Related Brain Potentials and the Study of Reward Processing: Methodological Considerations.” International Journal of Psychophysiology, Reward and Feedback Processing: State of the Field, Best Practices and Future Directions 132: 175–183. 10.1016/j.ijpsycho.2017.11.007. [DOI] [PubMed] [Google Scholar]
  30. Luck, S. J. 2014. An Introduction to the Event‐Related Potential Technique. MIT press. [Google Scholar]
  31. Meyer, G. M. , Marco‐Pallarés J., Boulinguez P., and Sescousse G.. 2021. “Electrophysiological Underpinnings of Reward Processing: Are We Exploiting the Full Potential of EEG?” NeuroImage 242: 118478. 10.1016/j.neuroimage.2021.118478. [DOI] [PubMed] [Google Scholar]
  32. Nelson, L. D. , Patrick C. J., Collins P., Lang A. R., and Bernat E. M.. 2011. “Alcohol Impairs Brain Reactivity to Explicit Loss Feedback.” Psychopharmacology 218, no. 2: 419–428. 10.1007/s00213-011-2323-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Novak, B. K. , Novak K. D., Lynam D. R., and Foti D.. 2016. “Individual Differences in the Time Course of Reward Processing: Stage‐Specific Links With Depression and Impulsivity.” Biological Psychology 119: 79–90. 10.1016/j.biopsycho.2016.07.008. [DOI] [PubMed] [Google Scholar]
  34. Novak, K. D. , and Foti D.. 2015. “Teasing Apart the Anticipatory and Consummatory Processing of Monetary Incentives: An Event‐Related Potential Study of Reward Dynamics.” Psychophysiology 52, no. 11: 1470–1482. 10.1111/psyp.12504. [DOI] [PubMed] [Google Scholar]
  35. Oerlemans, J. , Alejandro R. J., Van Roost D., et al. 2025. “Unravelling the Origin of Reward Positivity: A Human Intracranial Event‐Related Brain Potential Study.” Brain: A Journal of Neurology 148, no. 1: 199–211. 10.1093/brain/awae259. [DOI] [PubMed] [Google Scholar]
  36. Perrin, F. , Pernier J., Bertrand O., and Echallier J. F.. 1989. “Spherical Splines for Scalp Potential and Current Density Mapping.” Electroencephalography and Clinical Neurophysiology 72, no. 2: 184–187. [DOI] [PubMed] [Google Scholar]
  37. Polich, J. 2007. “Updating P300: An Integrative Theory of P3a and P3b.” Clinical Neurophysiology 118, no. 10: 2128–2148. 10.1016/j.clinph.2007.04.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Pornpattananangkul, N. , and Nusslock R.. 2015. “Motivated to Win: Relationship Between Anticipatory and Outcome Reward‐Related Neural Activity.” Brain and Cognition 100: 21–40. 10.1016/j.bandc.2015.09.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Proudfit, G. H. 2015. “The Reward Positivity: From Basic Research on Reward to a Biomarker for Depression.” Psychophysiology 52, no. 4: 449–459. 10.1111/psyp.12370. [DOI] [PubMed] [Google Scholar]
  40. San Martín, R. 2012. “Event‐Related Potential Studies of Outcome Processing and Feedback‐Guided Learning.” Frontiers in Human Neuroscience 6: 304. 10.3389/fnhum.2012.00304. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. San Martín, R. , Appelbaum L. G., Pearson J. M., Huettel S. A., and Woldorff M. G.. 2013. “Rapid Brain Responses Independently Predict Gain Maximization and Loss Minimization During Economic Decision Making.” Journal of Neuroscience 33, no. 16: 7011–7019. 10.1523/JNEUROSCI.4242-12.2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Shenhav, A. , Botvinick M. M., and Cohen J. D.. 2013. “The Expected Value of Control: An Integrative Theory of Anterior Cingulate Cortex Function.” Neuron 79, no. 2: 217–240. 10.1016/j.neuron.2013.07.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Smith, E. H. , Banks G. P., Mikell C. B., et al. 2015. “Frequency‐Dependent Representation of Reinforcement‐Related Information in the Human Medial and Lateral Prefrontal Cortex.” Journal of Neuroscience 35, no. 48: 15827–15836. 10.1523/JNEUROSCI.1864-15.2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. The MathWorks Inc . 2016. MATLAB Version: 9.0.0 (R2016a). MathWorks Inc. https://www.mathworks.com/. [Google Scholar]
  45. Tsujimoto, T. , Shimazu H., and Isomura Y.. 2006. “Direct Recording of Theta Oscillations in Primate Prefrontal and Anterior Cingulate Cortices.” Journal of Neurophysiology 95, no. 5: 2987–3000. [DOI] [PubMed] [Google Scholar]
  46. van den Berg, B. , Geib B. R., San Martín R., and Woldorff M. G.. 2019. “A Key Role for Stimulus‐Specific Updating of the Sensory Cortices in the Learning of Stimulus–Reward Associations.” Social Cognitive and Affective Neuroscience 14, no. 2: 173–187. 10.1093/scan/nsy116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Walsh, M. M. , and Anderson J. R.. 2012. “Learning From Experience: Event‐Related Potential Correlates of Reward Processing, Neural Adaptation, and Behavioral Choice.” Neuroscience & Biobehavioral Reviews 36, no. 8: 1870–1884. 10.1016/j.neubiorev.2012.05.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Watts, A. T. M. , Bachman M. D., and Bernat E. M.. 2017. “Expectancy Effects in Feedback Processing Are Explained Primarily by Time‐Frequency Delta Not Theta.” Biological Psychology 129: 242–252. 10.1016/j.biopsycho.2017.08.054. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Watts, A. T. M. , and Bernat E. M.. 2018. “Effects of Reward Context on Feedback Processing as Indexed by Time‐Frequency Analysis.” Psychophysiology 55, no. 9: e13195. 10.1111/psyp.13195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Watts, A. T. M. , Tootell A. V., Fix S. T., Aviyente S., and Bernat E. M.. 2018. “Utilizing Time‐Frequency Amplitude and Phase Synchrony Measure to Assess Feedback Processing in a Gambling Task.” International Journal of Psychophysiology, Reward and Feedback Processing: State of the Field, Best Practices and Future Directions 132: 203–212. 10.1016/j.ijpsycho.2018.04.013. [DOI] [PubMed] [Google Scholar]
  51. Woldorff, M. , Liotti M., Seabolt M., Busse L., Lancaster J., and Fox P.. 2002. “The tem‐Poral Dynamics of the Effects in Occipital Cortex of Visual‐Spatial Selective Attention.” Brain Research Cognitive Brain Research 15: 1–15. [DOI] [PubMed] [Google Scholar]
  52. Wu, Y. , and Zhou X.. 2009. “The P300 and Reward Valence, Magnitude, and Expectancy in Outcome Evaluation.” Brain Research 1286: 114–122. 10.1016/j.brainres.2009.06.032. [DOI] [PubMed] [Google Scholar]
  53. Yu, R. , Zhou W., and Zhou X.. 2011. “Rapid Processing of Both Reward Probability and Reward Uncertainty in the Human Anterior Cingulate Cortex.” PLoS One 6, no. 12: e29633. 10.1371/journal.pone.0029633. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Zheng, Y. , An T., Li Q., and Xu J.. 2020. “Distinct Electrophysiological Correlates Between Expected Reward and Risk Processing.” Psychophysiology 57, no. 10: e13638. 10.1111/psyp.13638. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix S1: Supporting papers for each hypothesis.

Appendix S2: Proportion of retained epochs by block, condition, and stage.

PSYP-63-e70387-s001.docx (44.9KB, docx)

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.


Articles from Psychophysiology are provided here courtesy of Wiley

RESOURCES