Abstract
To build an understanding of our world, we make inferences about the connections between our actions, experiences, and the environment. This process, state inference, requires an agent to guess the current state of the world given a set of observations. During value-based decision-making, a growing body of evidence implicates the orbitofrontal cortex (OFC) and the hippocampus (HPC) in the process of contextualizing information and identifying links between stimuli, actions, and outcomes. However, the neural mechanisms driving these processes in primates remain unknown. To investigate how OFC and HPC contribute to state inference, we recorded simultaneously from both regions while two male monkeys (Macaca mulatta) performed a probabilistic reversal learning task, where reward contingencies could be captured by two task states. Using population-level decoding, we found neural representations of task state in both OFC and HPC that remained stable within each trial but strengthened with learning as monkeys adapted to reversals. Subjects also appeared to use their understanding of task structure to anticipate reversals, evidenced by anticipatory neural representations of the upcoming task state.
Keywords: hippocampus, neurophysiology, orbitofrontal cortex, reversal learning
Significance Statement
Orbitofrontal cortex (OFC) and the hippocampus (HPC) are implicated in the process of contextualizing information and identifying links between stimuli, actions, and outcomes. However, limited work has been done in nonhuman primates to bridge the gap between rodent and human models. Here, we show that task state is represented in both OFC and HPC in nonhuman primates. These representations remain stable within trials, but evolve with learning and anticipate upcoming task changes, equipping subjects to adapt choices to reversals in reward contingencies. Our results support the theory that OFC–HPC interactions are important for flexible, goal-directed decision-making, and provide insight into how OFC and HPC participate in decision making when information is not explicitly provided, but must instead be inferred.
Introduction
In a reinforcement learning framework, task state is defined as the set of landmarks and rules that characterize an agent’s current environment (Sutton and Barto, 1998; Niv, 2019). To make optimal decisions in a dynamic world, animals must make inferences about their environment and identify task states that can explain recent observations and inform future decisions (Gershman et al., 2015). These states can then be linked together to form an abstract cognitive map that specifies the likelihood of one state leading to another state (Behrens et al., 2018; Whittington et al., 2020; Knudsen and Wallis, 2022). By constructing distinct state representations that capture features of the environment that dictate whether actions will be rewarding, we can learn to effectively behave and adapt to change (Langdon et al., 2019).
The hippocampus (HPC) has long been theorized to store a cognitive map that flexibly encodes spatiotemporal and causal relationships (Tolman, 1948; O'Keefe and Nadel, 1978; McKenzie et al., 2014; Eichenbaum, 2017; Sanders et al., 2020; Whittington et al., 2020, 2022). More recently, this process has also been linked to the orbitofrontal cortex (OFC; Wikenheiser and Schoenbaum, 2016; Niv, 2019; Elston and Wallis, 2025). For example, neuroimaging studies have shown that the ventromedial prefrontal cortex, a brain area adjacent to OFC, is the only brain area from which task states can be decoded (Schuck et al., 2016). Rodent work has shown that both OFC and HPC neurons encode task states (Zhou et al., 2019a,b, 2021; Basu et al., 2021). However, it is not clear how readily these findings translate to primates. There have been dramatic changes in the size and complexity of OFC across mammalian evolution (Wise, 2008; Preuss and Wise, 2022). Primates possess more frontal areas than rodents, and only primates possess a granular and dysgranular frontal cortex, which includes areas 11 and 13 in OFC. Concomitant with these changes in the frontal cortex, there have also been changes in the structure of HPC across the course of mammalian evolution, particularly in the size and complexity of anterior HPC in primates (Insausti, 1993 ), which is the part of HPC that connects with frontal cortex (Barbas and Blatt, 1995; Aggleton et al., 2015). Furthermore, there is limited evidence that OFC neurons encode task states in primates (Kennerley et al., 2009; Wallis and Rich, 2011; Cai and Padoa-Schioppa, 2014; Balewski et al., 2023). The primary driver of OFC neuronal activity concerns value and decision-making (Padoa-Schioppa and Assad, 2006; Rich and Wallis, 2016), not spatiotemporal contingencies that would be more relevant to encoding task states and their relationships.
Despite striking similarities between the proposed functions of the OFC and the HPC, few studies have tested their contributions to state inference during value-based learning and decision-making. One task that can be used for this is reversal learning. Recent theoretical explanations of reversal learning have argued that it is not simply the unlearning of the original association and the learning of the new contingency, but rather that subjects are learning about states of the world and their relationships, akin to the concept of a cognitive map (Wilson et al., 2014; Jang et al., 2015). Specifically, the subject learns that there are two “states” of the world, one in which stimulus A is rewarded and B is not, and one in which stimulus B is rewarded and A is not (Wilson et al., 2014; Jang et al., 2015). Previous studies have shown that monkeys performing reversal learning tasks do have representations of distinct task states in neuronal activity in both OFC (Saez et al., 2015) and HPC (Bernardi et al., 2020). While this supports the model of reversal learning as consisting of two distinct task states, whether there are distinct contributions of OFC and HPC remains unclear.
Here, we trained two monkeys (Macaca mulatta) to perform a probabilistic reversal learning task where they were challenged to adapt their decision-making behavior to uncued changes in reward contingencies. We used behavioral modeling to show that monkeys were utilizing state-based learning strategies rather than simple reinforcement learning. We recorded simultaneously from OFC and HPC and found encoding of task state in both areas. We then used population-level decoding to further explore the dynamics of these task state representations and examine how they interact with neural representations of reward.
Materials and Methods
Experimental model and subject details
All procedures were carried out as specified in the National Research Council guidelines and approved by the Animal Care and Use Committee and the University of California, Berkeley. Two male rhesus macaques (Subjects D and V, respectively) aged 8 and 11 years old, weighing 9.8 and 10 kg at the time of recording, were used in the current study. Subjects sat head-fixed in a primate chair (Crist Instrument) while eye movements were tracked with an infrared eye-tracking system (SR Research). Stimulus presentation and behavioral conditions were controlled using the MonkeyLogic toolbox (Hwang et al., 2019). Subjects had unilateral recording chambers implanted, centered over the posterior frontal lobe and parietal lobe.
Task design
Subjects performed a probabilistic reversal learning task where they were required to choose between two differently valued pictures (free choice, 85% of trials) or a single picture (forced choice, 15% of trials). Forced choice trials were included to prevent side bias and ensure that subjects experienced outcomes associated with all picture options. All pictures were isoluminant, presented against a dark gray background. To initiate the trial subjects had to maintain fixation on a central cue for 500 ms. After a 400 ms delay, pictures (2° × 2°) appeared 8° on either side of the fixation cue. Subjects used saccades to choose the left or right picture, indicating their choice by maintaining fixation for 500 ms. After making a choice, the unchosen picture disappeared, and an apple juice reward was delivered probabilistically after a 400 ms delay. Trials were separated by a 1,500 ms intertrial interval.
Both animals received two sets of two unique images each session, with novel pictures used each day (two unique pairs and four unique images per session). They were trained on these novel pictures the day prior to recording. There were two sets of two pictures each, where one set was associated with a low probability of reward (Subject D: 40%, Subject V: 30 or 40%) and the other was associated with a high probability of reward (90% for both subjects). The magnitude and type of reward (diluted apple juice) remained consistent across all trials. Free choice trials always consisted of one high-value option and one low value option, and all combinations of high and low value options were used. The images were balanced for whether they were presented on the left or right side of the screen.
Reversals were triggered once the animal reached a learning criterion where they chose the more valuable option on 70% or more over the last 30 trials and they had to maintain this performance level for 10 trials. Once this criterion was met, subjects entered a phase where an uncued pseudo-probabilistic reversal could occur within 10 trials so long as accuracy remained above 60% throughout this phase. These postcriterion trials enabled the subject to exploit what they had learned, ensuring they did not become frustrated with constant reversals.
Neurophysiological recording
Subjects were fitted with head positioners and imaged in a 3 T MRI scanner. From the MR images, we constructed 3D models of each subjects’ skull and target brain areas (Paxinos et al., 2000). Subjects were implanted with custom radiolucent recording chambers fabricated from polyether ether ketone (PEEK). During each recording session, up to three multisite linear probes (32-channel V probes with 100 μm contact spacing, Plexon) were lowered into OFC (areas 11 and 13) and HPC (CA1 and CA2/3) simultaneously (up to six electrodes total). Unique electrode trajectories were defined for each session in custom software, and the appropriate microdrives were 3D printed (Form 2 and 3, Formlabs; Knudsen et al., 2019). Lowering depths were derived from the MR images and verified from neurophysiological signals via gray/white matter transitions. Neural signals were digitized using a Plexon OmniPlex system, with continuous spike-filtered signals (200 Hz–6 kHz) acquired at 40 kHz and local field-filtered signals acquired at 1 kHz.
We recorded neuronal activity over the course of nine sessions for Subject D and 12 sessions for Subject V. We restricted our analysis to neurons with a mean firing rate across the session >1 Hz. Spike sorting was performed manually (Offline Sorter, Plexon). To ensure adequate isolation of individual units, we excluded neurons where >0.2% of spikes were separated by <1,100 ms. We transformed single neuron activity into a binary time series at 1 ms resolution, where 1 indicated the presence of a spike and 0 the absence.
Behavioral modeling
To determine how our subjects were learning the task, we fit learning models to the subjects’ choice behavior. We tested three different models: (1) a simple reinforcement learning model (Sutton and Barto, 1998), (2) a reinforcement learning model that incorporated knowledge that the four images were organized into two pairs, and (3) a reinforcement learning model that incorporated a state belief (Rodriguez et al., 1999).
For the standard reinforcement learning model (RLM-4), we estimated the subject’s value, Q1…4, of each individual picture, n, on each trial, t. Values were updated based on the outcome, R, of previously chosen picture n and the learning rate, α, which dictates how much weight prediction errors have on value updates:
Learning rate ranges were given boundaries of 0.01–0.95. The model’s choices were made by using the standard softmax-activation function:
where Pi,t is the probability of choosing the picture i over picture j on trial t and β is the inverse temperature parameter that determines the slope of the choice function and hence the noisiness of choices. The inverse temperature parameter was given boundaries of 0.01–10. Throughout the paper, we use the term V to indicate the objective value of the picture (the actual probability of reward) and Q to indicate the subjective value of the pictures (the subject’s estimate of the value of the picture) derived from the RL model.
The second reinforcement learning model (RLM-2) was identical to the first, except there were only two Q values, one for set A and one for set B.
The third model (RLM-State) updated state beliefs, S1 and S2, instead of Q values. These states corresponded to the two states shown in Figure 1b. For example, if the subject chose picture A and received a reward, then S1 would be increased, since it is more likely that this occurrence indicates State 1 is in effect rather than State 2, while S2 would be decreased. In all other respects, the model works identically to the prior models, except S replaces Q.
Figure 1.
Task design. a, Subjects fixated on a central cue to initiate each trial. After acquiring fixation, the central cue disappeared. After a 400 ms delay, they were presented with one or two pictures (forced and free choice trials, respectively). To select a picture and receive a probabilistic juice reward, monkeys made a saccade and fixated the picture for 500 ms. The outcome of the choice (juice or no juice) was revealed after 400 ms. b, Subjects learned the values of two sets of pictures, where pictures within each set had the same value. After each reversal, the values corresponding to each set switched. Value was defined here as reward probability.
We compared the fit of the models using Akaike’s Information Criterion (AIC; Akaike, 1974), and to compare these values across sessions with differing numbers of trials, we converted the AIC values into AIC weights, which measures the relative likelihood of a given model relative to the other models. We first subtract the lowest AIC value from each model’s AIC value, to give Δ1…K, where K is the total number of models. We then calculate the AIC weight, w1…K, using the following equation:
Single neuron regression analysis
For each neuron, in overlapping 100 ms windows shifted by 25 ms, we performed the following linear regression on firing rates (FR) aligned to picture onset:
We added a parameter for reward when performing linear regression on firing rates aligned to outcome:
with binary variables for reward (+1 reward, −1 no reward), state (+1 state 1, −1 state 2), chosen side (+1 left, −1 right), chosen value (+1 high, −1 low), and trial type (+1 free, −1 forced). Chosen picture was dummy coded (+1 for chosen picture, −1 for all three unchosen options). We included trial number as a noise parameter to absorb potential variance due to neuronal drift over the recording session. Significance was defined as maintaining p < 0.001 for 100 ms (four consecutive time bins). We validated this threshold by finding the proportion of units labeled as significantly predicted by chosen side during the 750 ms before and 200 ms after fixation, before pictures appeared on the screen. Only 20 of 1,492 units, or 1%, met these criteria during this window, so we proceeded to use this cutoff to detect units significantly predicted by all factors in our regression model.
Task state decoding with single trial resolution
For each session, we trained a decoding algorithm using linear discriminant analysis (LDA) to derive trial-by-trial estimates of task state from neuronal activity. We trained the decoder on trials where behavioral performance was high since these would be more likely to correspond with neural representations of the current task state. Thus, we restricted training trials to high-performance trials, defined as trials where running behavioral accuracy exceeded 60% across 30 trials. To account for nonstationarity in firing rates caused by neuronal drift, we split trials into overlapping bins for LDA training, where each bin contained ∼25% of all available high-performance trials. Training bins were then stepped by 50 trials and the LDA decoder training repeated.
Within each bin, training trials were balanced to include an equal number of trials corresponding to each state. This involved randomly downsampling trials to match the state with the fewest number of trials. For each bin, decoders were trained on 200 ms sliding windows of neural activity, stepped by 50 ms. In summary, training data for each bin consisted of an [n training trials × n timestamps × n units] matrix of firing rates. To reduce the dimensionality of the input features in each window, we then performed principal components analysis across trials and restricted LDA inputs to the top principal components that explained 95% of variance.
To assess the performance of our decoder, we used a leave-one-out (LOO) cross-validation procedure for each bin. Iterating over trials within each time bin, we held out one trial, trained the decoder on the remaining trials, and then tested on the held-out trial along with a 175-trial window buffer around the current bin. This allowed for every trial in the session to be tested. Because training trials were randomly removed in the downsampling procedure, to ensure that as many trials from each session contributed to the analysis as possible, we repeated the entire procedure 25 times and averaged posteriors across all instances where a trial was tested. We defined decoder confidence, C, as the difference between the posterior probabilities, P, of decoding the state of the current block relative to the state of the previous block, averaged across decoder iterations, N:
Very occasionally (D: 1 of 133 windows, V: 6 of 196 windows), there was an extended sequence of trials where only one state label was present—since these periods tended to indicate stretches of low task motivation and poor performance (thereby preventing a reversal from occurring), we excluded these trials from our training sets.
Statistics
All statistical tests are described in the main text or the corresponding figure legends. Error bars and shading indicate standard error of the mean (SEM) unless otherwise specified. All terms in regression models were normalized and had maximum variance inflation factors of 1.4. The coefficient of partial determination (CPD) was used to quantify the percentage of overall variance uniquely explained by each term. All comparisons were two tailed.
Results
Behavioral analysis
We trained two monkeys (Subjects D and V) on a probabilistic reversal task (Fig. 1). Subject D completed 13,889 trials over nine sessions (mean, 1,488 ± 105), while Subject V completed 20,543 trials over 12 sessions (mean, 1,712 ± 71). This equated to a mean of 16 ± 2 reversals per session for D and 13 ± 2 reversals per session for V (Fig. 2a–c). To calculate the number of trials for subjects to reach our learning criterion, we excluded forced choice trials, exploitation trials, and the last block of the session where subjects’ motivation was reduced and they did not reach the learning criterion. The mean number of trials to reach criterion were 53 ± 4 trials for D and 88 ± 6 trials for V (Fig. 2b).
Figure 2.

Behavior. a, Probability of choosing the high value option across an example session (Subject V). Vertical gray lines mark reversal points. The dashed black line is chance (50%) and the dashed green line is the learning criterion. b, Number of trials to reach a state reversal learning criterion for each reversal block. Data points are individual reversal blocks and red lines indicate the mean. c, Number of reversals per session. Data points are individual sessions and red lines indicate the mean. d, Mean (±SEM) normalized AIC values of three different RLM models tested, weighted, and averaged across sessions. The model incorporating a state belief was strongly favored for both subjects.
Our behavioral modeling showed that in both subjects our model that incorporated a state belief outperformed the other models. Comparing the AIC values for each subject over each session, we found that the model that incorporated state belief was the best fitting model in 8/9 or 89% of the sessions in D and 12/12 or 100% of the sessions in V (Fig. 2d; Tables 1, 2).
Table 1.
Model AIC values
| Subject | Session | RLM 4 | RLM 2 | RLM state |
|---|---|---|---|---|
| D | 1 | 1,885 | 1,909 | 1,511 |
| D | 2 | 821 | 826 | 744 |
| D | 3 | 1,768 | 1,782 | 1,439 |
| D | 4 | 1,394 | 1,412 | 1,116 |
| D | 5 | 1,811 | 1,855 | 1,552 |
| D | 6 | 1,199 | 1,244 | 1,241 |
| D | 7 | 1,741 | 1,808 | 1,435 |
| D | 8 | 1,126 | 1,153 | 895 |
| D | 9 | 2,129 | 2,216 | 1,880 |
| V | 1 | 2,333 | 2,348 | 1,662 |
| V | 2 | 1,865 | 1,878 | 1,284 |
| V | 3 | 2,381 | 2,399 | 1,706 |
| V | 4 | 2,257 | 2,258 | 1,740 |
| V | 5 | 1,910 | 1,903 | 1,401 |
| V | 6 | 1,729 | 1,762 | 1,355 |
| V | 7 | 2,280 | 2,304 | 1,703 |
| V | 8 | 2,038 | 2,049 | 1,631 |
| V | 9 | 1,424 | 1,440 | 1,059 |
| V | 10 | 2,124 | 2,137 | 1,791 |
| V | 11 | 1,675 | 1,712 | 1,482 |
| V | 12 | 1,382 | 1,418 | 1,023 |
Bold indicates best model.
Table 2.
Estimated model parameters
| Subject | Session | RLM 4 alpha | RLM 2 alpha | RLM state alpha | RLM 4 beta | RLM 2 beta | RLM state beta |
|---|---|---|---|---|---|---|---|
| D | 1 | 0.21 | 0.14 | 0.44 | 1.4 | 1.6 | 3.4 |
| D | 2 | 0.36 | 0.27 | 0.30 | 3.0 | 3.3 | 1.2 |
| D | 3 | 0.29 | 0.20 | 0.45 | 1.7 | 2.0 | 1.8 |
| D | 4 | 0.40 | 0.38 | 0.30 | 1.9 | 2.1 | 1.2 |
| D | 5 | 0.41 | 0.29 | 0.49 | 2.6 | 2.8 | 8.2 |
| D | 6 | 0.58 | 0.40 | 0.30 | 3.1 | 3.3 | 1.2 |
| D | 7 | 0.41 | 0.32 | 0.30 | 2.4 | 2.2 | 1.2 |
| D | 8 | 0.39 | 0.29 | 0.30 | 2.6 | 2.7 | 1.2 |
| D | 9 | 0.42 | 0.29 | 0.30 | 2.2 | 2.1 | 1.2 |
| Mean | 0.39 | 0.29 | 0.35 | 2.3 | 2.5 | 2.3 | |
| S.E. | 0.034 | 0.027 | 0.027 | 0.19 | 0.20 | 0.78 | |
| V | 1 | 0.09 | 0.07 | 0.27 | 1.5 | 1.3 | 0.0 |
| V | 2 | 0.05 | 0.24 | 0.56 | 1.3 | 0.8 | 6.3 |
| V | 3 | 0.10 | 0.04 | 0.49 | 1.2 | 1.2 | 5.5 |
| V | 4 | 0.02 | 0.03 | 0.45 | 1.4 | 1.3 | 0.3 |
| V | 5 | 0.07 | 0.05 | 0.30 | 1.6 | 2.0 | 1.2 |
| V | 6 | 0.08 | 0.07 | 0.60 | 2.7 | 2.4 | 10.0 |
| V | 7 | 0.08 | 0.07 | 0.30 | 1.9 | 1.8 | 1.2 |
| V | 8 | 0.11 | 0.07 | 0.35 | 1.3 | 1.2 | 2.0 |
| V | 9 | 0.11 | 0.08 | 0.52 | 1.7 | 1.8 | 6.2 |
| V | 10 | 0.09 | 0.07 | 0.25 | 2.3 | 2.4 | 1.3 |
| V | 11 | 0.13 | 0.11 | 0.33 | 3.1 | 3.0 | 7.6 |
| V | 12 | 0.10 | 0.11 | 0.49 | 2.3 | 1.8 | 6.1 |
| Mean | 1.4 | 1.6 | 2.0 | 2.5 | 2.9 | 4.0 | |
| S.E. | 0.008 | 0.016 | 0.035 | 0.18 | 0.18 | 0.96 |
Neural representations of task state strengthen with learning and in anticipation of change
We recorded single neurons from OFC (D: N = 224, V: 346) and HPC (D: 342, V: 580), using up to three acute multisite probes per region per session (Fig. 3a). For D, we recorded a mean of 63 ± 8 neurons per session (OFC: 25 ± 3 neurons, HPC: 38 ± 5 neurons) and a mean of 77 ± 7 neurons from V (OFC: 29 ± 3 neurons, HPC: 48 ± 4 neurons).
Figure 3.
Task parameters encoded in OFC and HPC. a, Reconstruction of OFC and HPC recording sites on coronal slices. The size of each circle represents the approximate number of neurons recorded at that location. AP (anterior-posterior) locations are in millimeters, relative to the interaural line. b, Percentage of neurons encoding different task parameters during different task epochs. Selectivity was defined as a significant beta weight (p < 0.01) for at least four consecutive 100 ms time bins. Lighter bars represent OFC neurons and darker bars represent HPC neurons. Asterisks indicate that there was a significant difference in the prevalence of selective neurons between the areas (chi-squared test, *p < 0.05, **p < 0.01, ***p < 0.001).
To examine the relationship between learning, choice behavior, and neuronal firing rates across and within trials, we performed linear regressions on firing rates aligned to the time of picture onset and reward outcome. We fit firing rates in sliding time windows with task state, choice direction, chosen value, chosen picture, trial type (forced or free), and reward as predictors (Fig. 3b; see Materials and Methods). The firing rates of most neurons were significantly predicted by at least one model term in one task epoch in both OFC (D: 177/224 or 79%, V: 193/346 or 56%) and HPC (D: 300/342 or 88%, V: 395/580 or 68%).
Figure 4 illustrates examples of single neuron selectivity in OFC and HPC. The firing rate of some neurons was driven by a single factor, such as the reward-selective neuron in OFC (Fig. 4a). However, many neurons in OFC (D: 92/224 or 41%, V: 76/346 or 22%) and HPC (D: 166/342 or 49%, V: 146/580 or 25%) encoded more than one factor. For example, the OFC neuron in Figure 4b encoded both the state and the choice response, while the HPC neuron in Figure 4d encoded the choice response and the reward outcome. Many neurons also encoded the current state and maintained this information across all epochs of the task (Fig. 4c,e). In Subject D, these neurons tended to be more prevalent in HPC than OFC (Fig. 3b), but they were equally prevalent in both areas in Subject V.
Figure 4.
Single neuron examples. Spike density histograms recorded from single neurons in OFC (a–c) and HPC (d–f), including significant encoding of (from top to bottom) reward outcome, task state, chosen value, and chosen side, across trial epochs (fixation, picture onset, and reward outcome). Lines depict firing rates (mean ± SEM) averaged over conditions within each factor. Red bars at the top of each panel indicate periods of significance determined by the regression.
To examine how state information responded to reversals, we used LDA to decode trial-by-trial estimates of state from population-level neuronal activity (see Materials and Methods). Mean decoder accuracy was above chance (50%) for all sessions and remained relatively constant across trial events. In general, decoders trained on HPC activity were more accurate (V: 60.5% ± 2.7%, D: 61% ± 1.6%) than those trained on OFC (V: 56.5% ± 2%, D: 55.5% ± 1.1%). However, we also typically had larger neuronal sample sizes in HPC. When we downsampled HPC ensembles to match those in OFC, there was no consistent difference in decoder performance between the two areas across sessions (p > 0.1, permutation test for all epochs in both subjects).
To determine how state representations in OFC and HPC relate to behavior, we examined how state decoding evolved over the course of learning. Blocks of trials (defined as all the trials for a given reversal) varied in decoder accuracy and confidence, but generally, decoder accuracy evolved in parallel with behavior (Fig. 5a). Immediately following a state reversal, behavioral performance usually dipped below chance as subjects continued erroneously choosing the best option for the prereversal state. Once subjects detected the state change, their performance improved, eventually meeting the criterion required to trigger the next reversal. Decoder confidence (defined as the normalized mean posterior probability of decoding the current state over the previous state, averaged across postevent timestamps) showed similar dynamics, gradually increasing over the first half of the block as the subject learned that the state had changed (Fig. 5b). More surprisingly, decoder confidence decreased during the second half of the block despite behavioral performance being near or above criterion by that point in learning. This suggests that subjects are beginning to represent the next task state in anticipation of the upcoming reversal. In other words, they recognize that a period of consistent reward (from making optimal choices) will always soon be followed by a state reversal, and this anticipation is reflected in the state represented.
Figure 5.
Decoding of state across learning. a, Top, Normalized (z-scored) mean posterior probability of decoding the current state over the previous state, averaged across timestamps. Bottom, Posterior probability of decoding state 1 over state 2, where 100% state 1 is gold and 100% state 2 is blue. Solid bars above the heatmap show the true state label for each block. Here, we show all trials from one example session (Subject V) where decoders were trained on HPC population activity synced to picture onset. b, To equate learning across blocks with different numbers of trials, we split each block into bins containing 10% of the trials (trial decile). We then plotted state decoding as a function of trial decile and fit generalized linear regression models to the first five deciles and second five deciles. Data points show the mean (±SEM) confidence for each decile, averaged across blocks (pooled across sessions). Black lines illustrate best-fit lines. Solid lines and filled points represent the picture onset epoch, and dashed lines and open circles represent the outcome epoch. Fits were significant in both subjects for both brain areas and task epochs (linear regression, p < 0.01).
Combined, these findings suggest that the population activity in OFC and HPC encodes information about task state and that these representations strengthen as the animal becomes more certain that its observations align with that state. Once performance rises to the point where the animal begins to suspect that a reversal may happen soon, OFC and HPC begin to represent the upcoming task state in anticipation of the next block.
Neural state representations drive choice behavior
We next asked whether the speed with which the monkeys adapted to the new task state after a reversal correlated with neural state decoder dynamics. Because the number of trials subjects took to trigger reversals varied appreciably, we were able to sort blocks into one of two groups. “Fast” blocks were defined as the 25% of blocks where subjects reached the learning criterion quickest, whereas “slow” blocks were the 25% of blocks where subjects took longest to reach the criterion (Fig. 6a). The mean number of trials to reach learning criterion for fast blocks was 32 ± 2 trials for D and 42 ± 1 trials for V. For slow blocks, it was 102 ± 13 trials for D and 207 ± 16 trials for V.
Figure 6.
State decoding correlates with learning dynamics. a, Mean probability of choosing the high-value option following reversal for fast blocks (blue: shortest 25% of blocks) and slow blocks (gray: longest 25%). For each group, we only plotted the data to the point where less than half the blocks had reached the learning criterion. The horizontal dashed black line shows chance performance (50%) and the horizontal dashed red line shows the learning criterion. b, Example session showing mean decoder confidence for fast and slow blocks across trials at the time of the reward outcome. Dark blue indicates greater confidence in the old, now-incorrect state than the current state, while red indicates confidence in the new, correct state. For visualization purposes, we smoothed over trials using a two-trial window. Data shown is from HPC in subject D. c, d, Correlation between learning and (c) OFC and (d) HPC decoder performance. Behavioral performance was measured as the number of trials to reach the learning criterion (maintaining 70% accuracy over 30 trials for 10 consecutive trials). Decoder performance was measured as the number of trials required to exceed 70% of the peak posterior probability of decoding for the correct task state in the current block. Datapoints show individual blocks combined across sessions for each subject.
We visualized the dynamic adaptation of population activity to the new task state by making heatmaps of decoder confidence across time and trials for fast and slow groups (Fig. 6b). In fast blocks, the decoder rapidly adjusted to report the relevant state and its confidence in this report progressively increased. In slow blocks, the decoder took longer to report the relevant state, and its confidence also took longer to increase.
To quantify this effect, we correlated behavioral performance with changes in the state reported by the decoder. Specifically, we examined whether the posterior probability of decoding the current state over the new state evolved in pace with behavioral performance, such that the decoder gained confidence in the current task state more slowly during blocks where subjects were slow to adapt their choice behavior (and vice versa). We defined performance thresholds for behavior and the decoder and identified the number of trials it took to hit that threshold within each block. For behavior, we set the threshold to the learning criterion used to trigger probabilistic reversals during recording sessions: choosing the more valuable option on 70% or more of 30 trials and maintaining this performance level for 10 trials. For the decoder, we set the threshold to 70% of the peak decoder confidence (defined as the absolute value of the difference between the probability of decoding the current state relative to the previous state).
The number of trials it took an OFC decoder to reach threshold was significantly correlated with the number of trials it took the subject to reach learning criteria for a given block (Fig. 6c). This suggests that the speed with which OFC updates the representation of task state to match the current block informs how quickly behavior adapts after reversals. The results from HPC were inconsistent across subjects, with a significant correlation between behavioral performance and the HPC decoder in Subject V, but not in Subject D.
Discussion
We observed encoding of the current task state in both OFC and HPC that was evident throughout the trial and that was independent of stimulus presentation and reward outcome. The strength with which task state was represented in both regions predicted choice behavior, such that neuronal decoding of state and behavioral performance increased together in the trials following a reversal in reward contingencies. Subjects also appeared to use their understanding of task structure to anticipate reversals, evidenced by anticipatory neural representations of the upcoming task state.
Many neurons in OFC and HPC were also sensitive to the interaction of reward outcome and expected value, and the encoding of these neurons strengthened with learning of the reward contingencies. However, unlike the representations of task state, there was no evidence that value expectations were anticipating future reversals. Thus, although it appeared that neurons preemptively encoded the upcoming task state, they continued to represent the current expected values until evidence (i.e., failure to reliably receive reward for previously rewarding choices) pushed them to update those.
Representing rules, states, and beliefs in the frontal cortex
Our findings are consistent with recent modeling results investigating how monkeys solve reversal learning tasks. When monkeys are exposed to repeated reversals, they learn to reverse more and more quickly, and this can be modeled as the development of a Bayesian prior that reversals can occur (Jang et al., 2015). Our results suggest that this can also lead to neural activity that attempts to predict the likelihood of a reversal as performance approaches a point that typically triggers a reversal. Our results are also consistent with recent ideas arguing that OFC is important for representing task states (Stalnaker et al., 2015; Wikenheiser and Schoenbaum, 2016; Niv, 2019). For example, many neuropsychological results from OFC lesions can be modeled by assuming disruption of the representation of task state (Wilson et al., 2014).
One challenge to these ideas is that there are many semantically related ideas to states, including rules (Wallis et al., 2001), categories (Freedman et al., 2001; Miller et al., 2003), contexts (Asaad et al., 2000), and beliefs (Zhu et al., 2012). Implementation of these functions has typically not been ascribed to OFC, but rather lateral prefrontal cortex (Miller et al., 2003). One distinction that has been proposed argues that states reflect abstract representations of the observable and unobservable properties of the current environment, while rules specify mappings from conditions to actions (Wilson et al., 2014). However, it can be difficult to determine how to apply such distinctions to behavioral tasks to make predictions. For example, in visual discrimination learning lesions of the lateral prefrontal cortex impair switching attentional set, while lesions of OFC impair switching of reward contingencies (Collins et al., 1998). It is unclear why an attentional set would be more clearly related to actions than reward.
More recently, we have proposed that OFC may have a more restricted role than representing task states (Knudsen and Wallis, 2022). We argue that behavioral tasks can be represented as state-transition graphs, which define the spatiotemporal and causal contingencies of how one state relates to another. HPC may play an important role in constructing and representing such graphs, although other prefrontal areas, such as lateral and medial prefrontal cortex, may also play an important role in the process. OFC is responsible for representing the current goal and using the state-transition graph to calculate the value of states in relation to the goal. This allows the subject to choose optimally in any specific state and to flexibly respond when goals change. We recently showed evidence for this interpretation. We trained animals on a task where a cue indicated whether stimulus–reward contingencies were in one mapping or a reversed mapping. HPC (but not OFC) encoded the state cue and communicated this information to OFC via theta synchronization so that OFC could calculate state-dependent values (Elston and Wallis, 2025).
Hippocampal and orbitofrontal contributions and potential interactions
Previous studies investigating the roles of OFC and HPC in representing task states during reversal learning have largely focused on each region individually. When OFC and HPC are compared, it was often across animal cohorts and limited to rodent models (Wikenheiser and Schoenbaum, 2016; Zhou et al., 2019a,b). Here, we recorded from OFC and HPC simultaneously during reversal learning in primates, enabling us to study how both regions encoded task states in parallel, within the same subject and recording session. We found neural representations of task state in both regions, which remained stable within trials while strengthening with learning. The trial-by-trial dynamics of state encoding appeared similar across OFC and HPC, and we did not see consistent differences between regions in terms of representation strength or timing. Precisely how and when OFC and HPC interact during learning to use models of task state to guide behavior remains an open question.
Studies in humans (Backus et al., 2016) and nonhuman primates (Brincat and Miller, 2015) have shown that synchrony between prefrontal cortex and HPC, quantified by the coupling strength of oscillations in the theta frequency band, predict learning in value-based and memory-based decision-making tasks. Previous work in our lab causally demonstrated that HPC-driven theta oscillations in OFC are crucial for updating choice preferences in the face of changing reward contingencies (Knudsen and Wallis, 2020), raising the possibility that OFC–HPC interactions are important for creating, updating, or employing representations of task state to guide action selection. One intriguing direction for future research could be examining interactions between OFC and HPC during the learning of state-transition graphs or their modification to better understand their individual contributions to constructing and using task state representations.
Interpretational issues
Prominent theories of HPC organization have argued for a dorsal-ventral (in rodents) or posterior-anterior (in primates) gradient of organization, with the dorsal/posterior HPC involved in cognitive processes, such as representing spatial information for navigation and memory, whereas the ventral/anterior HPC is more important for processing affective information (Moser and Moser, 1998; Fanselow and Dong, 2010; Strange et al., 2014). For example, there is a gradient within hippocampus with respect to the resolution of place cells, with the dorsal/posterior HPC encoding finer-grained spatial representations than the ventral/anterior HPC (Kjelstrup et al., 2008). In contrast, neurons in ventral/anterior HPC are more likely to differentiate locations with respect to reward contingencies (Royer et al., 2010). There is also a gradient in HPC with respect to its connections to the frontolimbic cortex, including OFC, with an increased density of connections toward the ventral/anterior HPC (Barbas and Blatt, 1995). However, more recent studies have shown that reward encoding is evident throughout HPC (Yun et al., 2023). Future studies specifically comparing subregions of OFC and HPC during state inference could more precisely locate task state representations in the brain and guide further research on how state representations are created, communicated, and updated.
An additional challenge is determining the behavioral strategy that the animals use to solve the task. Although reversal learning can be modeled as a two-state inference task (Wilson et al., 2014; Bartolo and Averbeck, 2020), there is considerable debate regarding how subjects actually learn the task (Izquierdo et al., 2017). Furthermore, even when a fully model-based behavioral strategy is theoretically available, animals (including humans) tend to employ a blend of model-based, model-free, and other processes (Glascher et al., 2010; Daw et al., 2011; Collins and Cockburn, 2020). In our study, we found evidence that subjects used a model of the task to generalize across paired stimuli, where the outcome of choosing one picture informed the likelihood of choosing the paired picture. However, expected value, or the probability that selecting a given picture will lead to reward, is difficult to pull apart from task state in a reversal learning task where “state” is defined by the current reward contingencies. Future studies in rodents and nonhuman primates could pilot more complex, but still consistently achievable, tasks that more clearly separate task state from the factors that define it.
Conclusion
Modeling tasks as a series of transitions through specific states is a computationally tractable and concise way to describe goal-directed behavior, but its precise implementation in the brain remains unclear. We showed that neural activity in OFC and HPC may be important, not only for representing current states, but also potentially anticipating future states. Understanding these processes could shed light on OFC damage and dysfunction which typically results in deficits in anticipating the consequences of one’s actions.
References
- Aggleton JP, Wright NF, Rosene DL, Saunders RC (2015) Complementary patterns of direct amygdala and hippocampal projections to the macaque prefrontal cortex. Cereb Cortex 25:4351–4373. 10.1093/cercor/bhv019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Akaike H (1974) A new look at the statistical model identification. IEEE Trans Autom Control 19:716–723. 10.1109/TAC.1974.1100705 [DOI] [Google Scholar]
- Asaad WF, Rainer G, Miller EK (2000) Task-specific neural activity in the primate prefrontal cortex. J Neurophysiol 84:451–459. 10.1152/jn.2000.84.1.451 [DOI] [PubMed] [Google Scholar]
- Backus AR, Schoffelen JM, Szebenyi S, Hanslmayr S, Doeller CF (2016) Hippocampal-prefrontal theta oscillations support memory integration. Curr Biol 26:450–457. 10.1016/j.cub.2015.12.048 [DOI] [PubMed] [Google Scholar]
- Balewski ZZ, Elston TW, Knudsen EB, Wallis JD (2023) Value dynamics affect choice preparation during decision-making. Nat Neurosci 26:1575–1583. 10.1038/s41593-023-01407-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Barbas H, Blatt GJ (1995) Topographically specific hippocampal projections target functionally distinct prefrontal areas in the rhesus monkey. Hippocampus 5:511–533. 10.1002/hipo.450050604 [DOI] [PubMed] [Google Scholar]
- Bartolo R, Averbeck BB (2020) Prefrontal cortex predicts state switches during reversal learning. Neuron 106:1044–1054.e4. 10.1016/j.neuron.2020.03.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Basu R, Gebauer R, Herfurth T, Kolb S, Golipour Z, Tchumatchenko T, Ito HT (2021) The orbitofrontal cortex maps future navigational goals. Nature 599:449–452. 10.1038/s41586-021-04042-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Behrens TEJ, Muller TH, Whittington JCR, Mark S, Baram AB, Stachenfeld KL, Kurth-Nelson Z (2018) What is a cognitive map? Organizing knowledge for flexible behavior. Neuron 100:490–509. 10.1016/j.neuron.2018.10.002 [DOI] [PubMed] [Google Scholar]
- Bernardi S, Benna MK, Rigotti M, Munuera J, Fusi S, Salzman CD (2020) The geometry of abstraction in the hippocampus and prefrontal cortex. Cell 183:954–967.e21. 10.1016/j.cell.2020.09.031 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Brincat SL, Miller EK (2015) Frequency-specific hippocampal-prefrontal interactions during associative learning. Nat Neurosci 18:576–581. 10.1038/nn.3954 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cai X, Padoa-Schioppa C (2014) Contributions of orbitofrontal and lateral prefrontal cortices to economic choice and the good-to-action transformation. Neuron 81:1140–1151. 10.1016/j.neuron.2014.01.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Collins AGE, Cockburn J (2020) Beyond dichotomies in reinforcement learning. Nat Rev Neurosci 21:576–586. 10.1038/s41583-020-0355-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Collins P, Roberts AC, Dias R, Everitt BJ, Robbins TW (1998) Perseveration and strategy in a novel spatial self-ordered sequencing task for nonhuman primates: effects of excitotoxic lesions and dopamine depletions of the prefrontal cortex. J Cogn Neurosci 10:332–354. 10.1162/089892998562771 [DOI] [PubMed] [Google Scholar]
- Daw ND, Gershman SJ, Seymour B, Dayan P, Dolan RJ (2011) Model-based influences on humans’ choices and striatal prediction errors. Neuron 69:1204–1215. 10.1016/j.neuron.2011.02.027 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eichenbaum H (2017) On the integration of space, time, and memory. Neuron 95:1007–1018. 10.1016/j.neuron.2017.06.036 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Elston TW, Wallis JD (2025) Context-dependent decision-making in the primate hippocampal-prefrontal circuit. Nat Neurosci 28:374–382. 10.1038/s41593-024-01839-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fanselow MS, Dong HW (2010) Are the dorsal and ventral hippocampus functionally distinct structures? Neuron 65:7–19. 10.1016/j.neuron.2009.11.031 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Freedman DJ, Riesenhuber M, Poggio T, Miller EK (2001) Categorical representation of visual stimuli in the primate prefrontal cortex. Science 291:312–316. 10.1126/science.291.5502.312 [DOI] [PubMed] [Google Scholar]
- Gershman SJ, Norman KA, Niv Y (2015) Discovering latent causes in reinforcement learning. Curr Opin Behav Sci 5:43–50. 10.1016/j.cobeha.2015.07.007 [DOI] [Google Scholar]
- Glascher J, Daw N, Dayan P, O'Doherty JP (2010) States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning. Neuron 66:585–595. 10.1016/j.neuron.2010.04.016 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hwang J, Mitz AR, Murray EA (2019) NIMH MonkeyLogic: behavioral control and data acquisition in MATLAB. J Neurosci Methods 323:13–21. 10.1016/j.jneumeth.2019.05.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Insausti R (1993) Comparative anatomy of the entorhinal cortex and hippocampus in mammals. Hippocampus 3:19–26. 10.1002/hipo.1993.4500030705 [DOI] [PubMed] [Google Scholar]
- Izquierdo A, Brigman JL, Radke AK, Rudebeck PH, Holmes A (2017) The neural basis of reversal learning: an updated perspective. Neuroscience 345:12–26. 10.1016/j.neuroscience.2016.03.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jang AI, Costa VD, Rudebeck PH, Chudasama Y, Murray EA, Averbeck BB (2015) The role of frontal cortical and medial-temporal lobe brain areas in learning a Bayesian prior belief on reversals. J Neurosci 35:11751–11760. 10.1523/JNEUROSCI.1594-15.2015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kennerley SW, Dahmubed AF, Lara AH, Wallis JD (2009) Neurons in the frontal lobe encode the value of multiple decision variables. J Cogn Neurosci 21:1162–1178. 10.1162/jocn.2009.21100 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kjelstrup KB, Solstad T, Brun VH, Hafting T, Leutgeb S, Witter MP, Moser EI, Moser MB (2008) Finite scale of spatial representation in the hippocampus. Science 321:140–143. 10.1126/science.1157086 [DOI] [PubMed] [Google Scholar]
- Knudsen EB, Balewski ZZ, Wallis JD (2019) A model-based approach for targeted neurophysiology in the behaving non-human primate. Int IEEE EMBS Conf Neural Eng 2019:195–198. 10.1109/NER.2019.8716968 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Knudsen EB, Wallis JD (2020) Closed-loop theta stimulation in the orbitofrontal cortex prevents reward-based learning. Neuron 106:537–547.e4. 10.1016/j.neuron.2020.02.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Knudsen EB, Wallis JD (2022) Taking stock of value in the orbitofrontal cortex. Nat Rev Neurosci 23:428–438. 10.1038/s41583-022-00589-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Langdon AJ, Song M, Niv Y (2019) Uncovering the ‘state': tracing the hidden state representations that structure learning and decision-making. Behav Processes 167:103891. 10.1016/j.beproc.2019.103891 [DOI] [PMC free article] [PubMed] [Google Scholar]
- McKenzie S, Frank AJ, Kinsky NR, Porter B, Riviere PD, Eichenbaum H (2014) Hippocampal representation of related and opposing memories develop within distinct, hierarchically organized neural schemas. Neuron 83:202–215. 10.1016/j.neuron.2014.05.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Miller EK, Nieder A, Freedman DJ, Wallis JD (2003) Neural correlates of categories and concepts. Curr Opin Neurobiol 13:198–203. 10.1016/S0959-4388(03)00037-0 [DOI] [PubMed] [Google Scholar]
- Moser MB, Moser EI (1998) Functional differentiation in the hippocampus. Hippocampus 8:608–619. 10.1002/(SICI)1098-1063(1998)8:6<608::AID-HIPO3>3.0.CO;2-7 [DOI] [PubMed] [Google Scholar]
- Niv Y (2019) Learning task-state representations. Nat Neurosci 22:1544–1553. 10.1038/s41593-019-0470-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- O'Keefe J, Nadel L (1978) The hippocampus as a cognitive map. Oxford: Oxford University Press. [Google Scholar]
- Padoa-Schioppa C, Assad JA (2006) Neurons in the orbitofrontal cortex encode economic value. Nature 441:223–226. 10.1038/nature04676 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Paxinos G, Huang X, Toga AW (2000) The rhesus monkey brain in stereotaxic coordinates. San Diego, CA: Academic press. [Google Scholar]
- Preuss TM, Wise SP (2022) Evolution of prefrontal cortex. Neuropsychopharmacology 47:3–19. 10.1038/s41386-021-01076-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rich EL, Wallis JD (2016) Decoding subjective decisions from orbitofrontal cortex. Nat Neurosci 19:973–980. 10.1038/nn.4320 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rodriguez AC, Parr R, Koller D (1999) Reinforcement learning using approximate belief states. Adv Neural Inf Process Syst 12:1036–1042. 10.5555/3009657.3009803 [DOI] [Google Scholar]
- Royer S, Sirota A, Patel J, Buzsaki G (2010) Distinct representations and theta dynamics in dorsal and ventral hippocampus. J Neurosci 30:1777–1787. 10.1523/JNEUROSCI.4681-09.2010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Saez A, Rigotti M, Ostojic S, Fusi S, Salzman CD (2015) Abstract context representations in primate amygdala and prefrontal cortex. Neuron 87:869–881. 10.1016/j.neuron.2015.07.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sanders H, Wilson MA, Gershman SJ (2020) Hippocampal remapping as hidden state inference. Elife 9:e51140. 10.7554/eLife.51140 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schuck NW, Cai MB, Wilson RC, Niv Y (2016) Human orbitofrontal cortex represents a cognitive map of state space. Neuron 91:1402–1412. 10.1016/j.neuron.2016.08.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stalnaker TA, Cooch NK, Schoenbaum G (2015) What the orbitofrontal cortex does not do. Nat Neurosci 18:620–627. 10.1038/nn.3982 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Strange BA, Witter MP, Lein ES, Moser EI (2014) Functional organization of the hippocampal longitudinal axis. Nat Rev Neurosci 15:655–669. 10.1038/nrn3785 [DOI] [PubMed] [Google Scholar]
- Sutton RS, Barto AG (1998) Reinforcement learning: an introduction (adaptive computation and machine learning). Cambridge: MIT Press. [Google Scholar]
- Tolman EC (1948) Cognitive maps in rats and men. Psychol Rev 55:189–208. 10.1037/h0061626 [DOI] [PubMed] [Google Scholar]
- Wallis JD, Anderson KC, Miller EK (2001) Single neurons in prefrontal cortex encode abstract rules. Nature 411:953–956. 10.1038/35082081 [DOI] [PubMed] [Google Scholar]
- Wallis JD, Rich EL (2011) Challenges of interpreting frontal neurons during value-based decision-making. Front Neurosci 5:124. 10.3389/fnins.2011.00124 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Whittington JCR, Muller TH, Mark S, Chen G, Barry C, Burgess N, Behrens TEJ (2020) The Tolman-Eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation. Cell 183:1249–1263.e23. 10.1016/j.cell.2020.10.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Whittington JCR, McCaffary D, Bakermans JJW, Behrens TEJ (2022) How to build a cognitive map. Nat Neurosci 25:1257–1272. 10.1038/s41593-022-01153-y [DOI] [PubMed] [Google Scholar]
- Wikenheiser AM, Schoenbaum G (2016) Over the river, through the woods: cognitive maps in the hippocampus and orbitofrontal cortex. Nat Rev Neurosci 17:513–523. 10.1038/nrn.2016.56 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wilson RC, Takahashi YK, Schoenbaum G, Niv Y (2014) Orbitofrontal cortex as a cognitive map of task space. Neuron 81:267–279. 10.1016/j.neuron.2013.11.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wise SP (2008) Forward frontal fields: phylogeny and fundamental function. Trends Neurosci 31:599–608. 10.1016/j.tins.2008.08.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yun M, Hwang JY, Jung MW (2023) Septotemporal variations in hippocampal value and outcome processing. Cell Rep 42:112094. 10.1016/j.celrep.2023.112094 [DOI] [PubMed] [Google Scholar]
- Zhou J, Montesinos-Cartagena M, Wikenheiser AM, Gardner MPH, Niv Y, Schoenbaum G (2019a) Complementary task structure representations in hippocampus and orbitofrontal cortex during an odor sequence task. Curr Biol 29:3402–3409.e3. 10.1016/j.cub.2019.08.040 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhou J, Gardner MPH, Stalnaker TA, Ramus SJ, Wikenheiser AM, Niv Y, Schoenbaum G (2019b) Rat orbitofrontal ensemble activity contains multiplexed but dissociable representations of value and task structure in an odor sequence task. Curr Biol 29:897–907.e3. 10.1016/j.cub.2019.01.048 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhou J, Jia C, Montesinos-Cartagena M, Gardner MPH, Zong W, Schoenbaum G (2021) Evolving schema representations in orbitofrontal ensembles during learning. Nature 590:606–611. 10.1038/s41586-020-03061-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhu L, Mathewson KE, Hsu M (2012) Dissociable neural representations of reinforcement and belief prediction errors underlie strategic learning. Proc Natl Acad Sci U S A 109:1419–1424. 10.1073/pnas.1116783109 [DOI] [PMC free article] [PubMed] [Google Scholar]





