SUMMARY
We often remember the consequences of past choices to adapt to changing circumstances. Recalling past events requires the hippocampus (HPC), and using stimuli to anticipate outcome values requires the orbitofrontal cortex (OFC).1–3 Spatial reversal tasks require both structures to navigate newly rewarded paths.4,5 Both HPC place6 and OFC value cells7,8 fire in phase with theta (4–12 Hz) oscillations. Both structures are described as cognitive maps: HPC maps space9 and OFC maps task states.10 These similarities imply that OFC-HPC interactions are crucial for using memory to predict outcomes when circumstances change, but the mechanisms remain largely unknown. To investigate possible interactions, we simultaneously recorded ensembles in OFC and CA1 as rats learned spatial reversals in a plus maze. Striking interactions occurred only while rats learned their first reversal: CA1 population vectors predicted changes in OFC activity but not vice versa, OFC spikes phase locked to hippocampal theta oscillations, mixed pairs of CA1 and OFC neurons fired together within single theta cycles, and CA1 led OFC spikes by ~30 ms. After the new contingency became familiar, CA1 ensembles stably represented distinct spatial paths, whereas OFC ensembles developed more generalized goal arm representations in different paths to identical rewards. These frontotemporal interactions, engaged selectively when new task features inform decision-making, suggest a mechanism for linking novel episodes with expected outcomes, when HPC signals trigger “cognitive remapping” by OFC.11
Graphical Abstract

In brief
Riceberg et al. show that CA1 ensembles strongly interact with and entrain OFC ensembles only when learning a spatial reversal for the first time. Interactions vanish once associations become familiar. The results suggest a mechanism for linking novel episodes with expected reward outcomes.
RESULTS
CA1 and orbitofrontal cortex (OFC) single units and local field potentials (LFPs) were recorded simultaneously as rats (n = 4) performed a familiar spatial discrimination and then learned a “first” and “unfamiliar” reversal in the same testing session. Subsequent sessions recorded activity in both structures as rats performed familiar discriminations and reversals (Figure 1A). Learning curves defined by Smith’s method6,12,13 quantified the probability of a correct choice for each trial, and the first trial to exceed chance defined the “learning trial” for each contingency block. Rats required more trials to learn the first reversal than subsequent reversals (Figures 1B and S1A–S1C; trials to criterion [TTC], reversal session 3 block interaction, F = 39.89, p < 0.01; Tukey post hoc Rev, first versus familiar, p < 0.01; Smith learning trial, reversal session 3 block interaction, F = 37.12, p < 0.01; Tukey post hoc Rev, first versus familiar, p < 0.01). Performance did not vary by start arm (Figure S1D) but did correlate strongly with both trial number and elapsed time (Figures S1E–S1G).
Figure 1. Experimental design and behavioral performance.

(A) Rats were trained to find food in one goal arm in a plus maze (e.g., “West”) and then implanted with tetrodes targeting ventrolateral OFC and dorsal CA1 of the hippocampus. Schematics depict sagittal and coronal brain sections 2.4 mm lateral, 3.7 mm anterior, and 3.8 mm posterior from bregma, respectively. After recovery from surgery, each rat was retrained to find reward in the pre-surgical location. Neural activity was recorded as rats learned to find reward in a new location and subsequently familiar locations, i.e., reversal learning. Rewarded locations switched when rats reached 17/20 correct.
(B) Rats learn to find reward in a newly rewarded location (first reversal) more slowly than a previously rewarded location (familiar reversal), requiring more trials to reach both the criterion trial (trials to criterion [TTC] ~54 first reversal versus ~24 familiar reversal) and the learning trial (~38 first reversal versus ~9 familiar reversal) determined by the Smith algorithm (STAR Methods). ID, initial discrimination; Rev, reversal. Error bars show SEM. ***p < 0.01.
See also Figure S1.
Neural signals that inform memory discriminations should predict choices. We therefore quantified how well unit ensembles predicted choices using support vector machines (SVMs) that categorized population vectors (PVs) recorded in the start arms before rats chose east or west goal arms for each correct trial. CA1 PVs predicted 85% and OFC PVs predicted ~80% of correct choices (Figure 2A); error trials were categorized poorly (Figure S2A). By predicting the goal of each trial before the discriminative response, OFC and CA1 signals could inform downstream brain regions, including one another, of pending choices.
Figure 2. Goal decoding and Granger prediction in CA1 and OFC ensembles.

(A) Support vector machines (SVMs) categorized goal choices in single trials from OFC (left) and CA1 (right) PVs recorded in the start arms. Each dot represents a single ensemble; filled dots show ensembles that decoded single trial choices better than chance cf. goal-shuffled permutation tests.
(B) The history of CA1 activity helped predict changes in OFC activity as rats learned the first reversal, but not after stable performance was attained. CA1 activity had a small but significant influence on OFC states during stable performance as rats performed a familiar reversal. Asterisks (***p < 0.01, *p < 0.05) above columns indicate significantly different from chance defined by permutation tests (dashed horizontal lines).
See also Figure S2.
CA1 and OFC contribute to learning unfamiliar spatial reversals4 and could do so either in parallel or interactively. Bidirectional interactions could either integrate CA1 journey codes5 with OFC processes that compute spatial expected outcomes14 or integrate OFC outcome expectancies with CA1 computations to differentiate journeys. To assess potential interactions between structures during learning, we analyzed sequential changes in PVs recorded in the start arm using Granger prediction.15 We divided neural activity by session (first versus familiar reversals) and learning stage, categorizing trials before and after the learning trial as “early learning” (EL) or “stable performance” (SP). We calculated Granger values in triplets of sequential trials, using PVs recorded in the first two trials to predict PVs in the third, e.g., trials 1 and 2 predict trial 3, trials 2 and 3 predict trial 4, etc. If during learning CA1 modulates OFC, then sequences of CA1 PVs should predict changes in OFC PVs; conversely, if OFC modulates CA1, then sequences of OFC PVs should predict changes in CA1. We found that CA1 PVs on the previous two trials increased the explained variance in OFC PVs on the current trial by ~40% only as rats learned the first reversal (actual versus shuffled Granger values, p < 0.0001; Figure 2B). In contrast, OFC PVs did not predict CA1 activity during the same trials (3% increase, p > 0.05). During SP of the first reversal neither structure’s activity predicted changes in the other (all p > 0.05). During SP of familiar reversals CA1 PVs predicted 5% of the variance in OFC activity (p < 0.05; Figures 2B and S2B); OFC never predicted CA1 activity. Before rats chose a goal, CA1 predicted OFC activity and not vice versa, and the strongest prediction occurred when learning requires the OFC.
Granger predictions analyzed sequences of trials each separated by many seconds, a timescale that could reflect the outcome of CA1-OFC interactions, but not their neuronal mechanisms. Communication between brain circuits requires precise spike timing so that pre- and post-synaptic activity is synchronized within milliseconds, timing that may be coordinated by LFPs. OFC and CA1 neurons phase-lock to local theta (4–12 Hz) LFPs, so we investigated how OFC phase locking to CA1 LFPs varied with learning. We first quantified the theta phase of all spikes recorded in the start arms and computed a mean phase for each unit and session (Figures 3A and S3F). CA1 units phase locked to CA1 theta throughout all sessions. In contrast, the population of OFC units phase locked to CA1 theta only as rats learned the first reversal (Figure 3A, blue lines; population Rayleigh16 tests, R1, learning R = 4.6, p < 0.01, n = 57; other conditions, R ≤ 0.92, p > 0.05). Pairwise-phase consistency17 showed similar results (Figure S3E, top). The population of OFC spikes was not significantly phase locking to CA1 theta during stable performance or subsequent reversals (all p > 0.4; Figure S3F). We next quantified phase locking of single OFC units. Unlike the population as a whole, only 1 of 57 single OFC units phase locked significantly to CA1 theta as rats learned the first reversal (p < 0.05; Figure 3A, left, dark black dots). During SP, however, 8 of 55 single OFC units phase locked to CA1 theta (Figure 3B, left; χ2 test, p < 0.05, X = 4.59; Figures S3G and S3E, bottom). Neither CA1 theta power (Figure S3A), OFC unit firing rates (Figure S3B), nor OFC predictive coding (Figure S3H) covaried with OFC unit phase locking. OFC phase locking to CA1 theta changed in opposite directions at the population and the single-unit levels of analysis during learning and memory performance. The population of OFC spikes aligned with hippocampal theta only as rats learned an unfamiliar path to reward, when both structures are required for the reversal; single OFC units phase locked to theta only afterward.
Figure 3. Theta dynamics across learning.

(A) OFC spikes were most likely to occur around the trough of CA1 theta when rats learned the first reversal. CA1 theta phase is double plotted along the horizontal axis. The bright blue lines show the phase distribution of the mean phase of all OFC units (proportion of OFC units’ mean phase, right vertical axis). Overall OFC spiking was significantly phase locked to CA1 theta as rats learned their first reversal (bold blue distribution) but not during other conditions. Each dot shows the magnitude of each OFC unit’s mean theta phase locking (MRL, mean resultant length, left vertical axis), with black dots indicating significant theta-locking for that unit by Rayleigh tests.
(B) Significantly fewer OFC units significantly phase locked to CA1 theta during early learning than during stable performance of the first reversal (1/57 versus 8/55, *χ2 = 4.59, p < 0.05).
(C) Coactive pairs of neurons in two example sessions (first and familiar reversals). Each node in the perimeter of the connectome represents a single unit from CA1 (circles) or OFC (diamonds). The probability of coactivity within theta cycles is shown by the color of the line connecting nodes (warmer = more likely paired activity). Coactivity was most likely when rats learned the first reversal (left connectome).
(D) Coactive pairs across all sessions. Within first (left) and familiar (right) reversal sessions, trials were subdivided into learning and stable performance. For all unit pairs, coactivity (vertical axis) was most likely while rats learned the first reversal. Error bars show SEM. *p < 0.001.
See also Figure S3.
Learning entails sub-second modulation of spike timing, e.g., the cross-correlations of CA1 spike pairs within theta cycles18 and place cell replay during sharp wave ripples.19 If learning depends on coordinated spiking by CA1 and OFC neurons, then their single units should be coactive, especially as rats learn the first reversal. We therefore estimated synchronized activity among all pairs of simultaneously recorded units in each theta cycle as rats approached the choice point. Ensemble spike trains were subdivided by single theta cycles into bins, and the coactivity of each unique pair of units was computed for each bin with a nonlinear recursive Bayesian method.20 To measure the probability that pairs of units fired within the same theta cycle, the coactivity values were averaged across theta cycles for each trial. Coactive pairs of CA1, OFC, and mixed pairs of CA1 and OFC units fired together most often as rats learned the first reversal, and less so as they adapted to familiar contingencies (Figures 3C and 3D). Figure 3C shows the coactivity of all unit pairs in two typical ensembles, one recorded during the first and the other during a familiar reversal. Pairwise coactivity was highest as rats learned the first reversal, specifically before the learning trial for all ensembles (Figures 3D, S3C, and S3D; ANOVA coactivity, reversal familiarity * learning stage, F4,1299 = 3.02, p < 0.01). Coactivity differed by reversal familiarity (F1,1299 = 8.05, p < 0.005), task phase (F1,1299 = 6.21, p < 0.05), and learning stage (F4,1299 = 5.6, p < 0.001), but not by pair region (F2,1299 = 1.76, p = 0.17). No other interactions were significant (F < 0.13, p > 0.95). Coactivity values were unrelated to firing rates (Figure S3B). Coactivity in ~33-ms gamma cycles (but not in ~250-ms delta cycles; data not shown) displayed similar dynamics (Figure S3C), Pairs of CA1 and OFC units that fire in gamma frequency nested within theta cycles could provide a neuronal mechanism for each structure to modulate the other through synaptic plasticity.21 Inter-spike interval (ISI) histograms show that CA1 spikes were more likely to precede OFC spikes by <30 ms than vice versa during first but not familiar reversals (Figure S2C), spike timing that could let CA1 modulate OFC.
If CA1 indeed modulates OFC when rats learn a new contingency, what is the effect? We conjectured that OFC integrated CA1 signals to represent a new path to a familiar reward—an expected outcome associated with location. We assessed changing neuronal representations by comparing spatial decoding in ensembles recorded during the first and subsequent reversals. Bayesian decoding quantified the probability that a rat occupied each maze position for every 125-ms PV sampled throughout each trial. CA1 PVs22 predicted occupied locations accurately and consistently across sessions (CA1 decoding index [Σdiagonal] ANOVAs: F1,192 < 3.1, p > 0.05; Figures 4A and 4B, right). OFC predictions were less accurate and less stable than CA1 (~2/5 the resolution: Σdiagonal, CA1 > OFC, t = 21.98, p < 0.001; Figure 4B, red right axis). OFC decoding accuracy was reduced in part because OFC PVs represented both goal arms, i.e., decoding the East while the rat was in the West and vice versa (Figure 4A, dashed boxes illustrated only for Familiar CA1). OFC decoding of the unoccupied goal arm increased after the first reversal session (ANOVA decoding index: reversal, F1,195 = 5.21, p < 0.05; arm, F1,195 = 9.58, p < 0.005; reversal * arm, F1,195 = 8.0, p < 0.01, first versus familiar reversal in goal arm, Bonferroni p < 0.01; Figure 4B), and did so gradually (Figures S4A and S4B) as both goal arms became familiar. Though neither local nor non-local CA1 goal decoding reliably predicted OFC goal decoding, CA1 prospective decoding predicted OFC goal decoding most strongly during the first reversal (delocalized CA1 goal arm decoding and localized OFC goal arm decoding, r = 0.16, p = 0.15).
Figure 4. Bayesian decoding of position in CA1 and OFC ensembles.

(A) Bayesian decoding quantified spatial probability distributions (y axis) based on ensemble activity in OFC (left) and CA1 (right) for each centimeter pixel along the maze (x axis) across learning (top to bottom). Dotted boxes, illustrated only for familiar CA1 decoding, enclose corresponding spatial locations in the other start or goal arm (“alternate-arm” locations—e.g., 5 cm mark along the North arm when at 5 cm mark along the South arm).
(B) CA1 representations stabilized but OFC spatial representations generalized, particularly in the goal arm during familiar reversals. Alternate-arm modulation measured the degree to which activity generalized across start or goal arm locations. As expected, CA1 activity faithfully and stably reflected current location (right y axis, red dots). To a lesser extent, OFC activity also signaled current location (red dots). Dotted bars show that goal generalization increased in OFC familiar cf. first reversals. *p < 0.01. Error bars show SEM.
See also Figure S4.
DISCUSSION
OFC and CA1 activity coordinated dynamically as rats learned their first spatial reversal, a task impaired by OFC dysfunction:4 CA1 ensembles predicted OFC dynamics across trials, OFC spikes phase locked to CA1 theta oscillations, pairs of CA1 and OFC neurons fired within the same theta and gamma cycles, and CA1 spikes preceded OFC spikes by ~30 ms. After learning, OFC ensemble dynamics, spike phase locking, coactivity, and ISI lags relative to CA1 decreased, and OFC spatial representations of opposite goal arms became more similar. The results suggest that OFC circuits may improve reversal learning by integrating the history of rewarded spatial episodes signaled by CA1 into paths toward expected outcomes. These frontotemporal interactions could link episodic memories and expected outcomes in distributed cognitive maps of task2 and memory23 spaces.
CA1 and OFC ensembles each predicted correct choices in single trials but represented locations differently. CA1 ensembles distinguish places,9 temporal intervals,22 motives,24 and other salient task features, including overlapping journeys to different spatial25,26 and auditory goals.27 Here, CA1 ensembles maintained distinct representations of spatial journeys while OFC representations generalized paths in the two goal arms that led to the same reward, external features that shared functional significance. Elsewhere, OFC neurons generalized and CA1 neurons discriminated identical odors presented in different sequences to the same reward.5 Generalizing different paths to the same outcome may provide a mechanism for learning sets or schemas28 that support cognitive flexibility. OFC activity is crucial for identifying rewards,29 and has been linked to hypothetical,30 imagined,31 forgone,32 and regretted33 outcomes. By integrating multiple predictors of common outcomes, the OFC could categorize environmental features that guide choices. If CA1-OFC interactions are needed to associate unfamiliar paths with reward, then similar dynamics should re-emerge when rats learn spatial discriminations and reversals on a new maze. If the OFC computes associations between spatial paths and specific outcomes, then OFC representations should become more distinct as rats learn that each path leads to a different reward. Future experiments will test if such contingencies let OFC predict CA1 dynamics.
Phase locking to and paired spiking within theta and gamma cycles may support CA1-OFC communication. As rats learned the new reversal contingency OFC spikes were phase locked to CA1 theta. After rats learned the contingency overall phase locking ceased even as the proportion of phase-locked single units increased. The transition from inconsistent phase locking by many neurons to reliable locking by a few may reflect competitive learning, as widely distributed weak synaptic weights transition via activity-dependent mechanisms to a small subset of strong weights. OFC spiking coordinated by CA1 theta and gamma may associate active CA1 and OFC cells during learning. Subsequently, the small set of theta locked OFC units could maintain an established link between the two circuits. Similar dynamics in CA1-OFC synchrony should occur whenever new paths to reward are learned,34 whenever OFC signals facilitate learning.35
Beyond phase locking of OFC spikes to CA1 LFPs, pairs of OFC and CA1 neurons fired in the same theta and gamma cycles, and this coactivity changed with learning. Both homogeneous and mixed pairs of CA1 and OFC neurons fired more often within LFP cycles when rats learned a new contingency than when rats adopted or switched to a familiar contingency. Different mechanisms may synchronize unit coactivity and modulate unit phase locking. CA1 phase locking did not covary with pairwise coactivity, as CA1 units were more often coactive during early reversal learning then subsequently, when phase locking was consistent. The dissociation of phase locking and coactivity could indicate that different mechanisms coordinate spike timing in CA1 and OFC, particularly when CA1 “remaps”36 and “theta flickers”37 as animals learn. Spike timing coordinated by hippocampal theta could provide a mechanism for OFC-CA1 communication that links behavioral episodes to their outcomes.
We hypothesized that OFC-HPC interactions would reflect “guided activation,”38 that PFC modulation of HPC activity would select prospective codes that inform memory retrieval. Instead, CA1 modulated OFC activity as rats approached the choice point, not vice versa, and did so only as rats learned an unfamiliar contingency—precisely when both violations of reward expectations and distinctness of spatial episodes peak in salience. Spatial journeys coded by CA1 could become integrated with reward signals by OFC circuits that predict outcomes associated with locations. Spatial locations, analogous to other (e.g., olfactory) discriminative stimuli, would be associated with reward history and integrated by OFC into paths to expected outcomes.2 More generally, HPC signals could let OFC link any episodic features with outcomes.
OFC-CA1 interactions are gated, occurring when animals adapt to new salient task features. Here, coordinated activity peaked when rats learned their first reversed reward contingency. In monkeys, OFC theta power increased as subjects fixated on discriminative stimuli, theta phase alignment between OFC and HPC LFPs increased as new stimulus-outcome associations were learned, and disrupting phase alignment impaired learning.8 In rats performing a two-odor go/no-go discrimination reversal task, phase locking of OFC spikes to local theta LFPs anticipated sucrose delivery, strengthened with learning, and predicted correct outcomes.7 Phase locking of OFC units to CA1 LFPs could reflect coherence of CA1 and OFC theta LFPs,7,14 the effects of theta-synchronized inputs from CA1 to OFC, or input synchronized by basal forebrain structures.39
Tolman proposed that people and rats learn cognitive maps, “indicating routes and paths and environmental relationships” that guide behavior.40 O’Keefe and Dostrovsky proposed that the HPC implements cognitive maps.9 Wilson et al. described the OFC as a cognitive map of task space10 that interacts with the HPC2 to support cognitive flexibility. The current results jibe with proposals that the OFC is needed for “cognitive remapping,” establishing or updating models of task space, rather than for tracking familiar contingencies.11 Here, CA1 activity guided OFC remapping as rats learned a new path to reward, and OFC decoded separate spatial paths with partially overlapping representations, consistent with a map of task space informed by maze locations. If other prefrontal circuits map other task dimensions, then, e.g., maps of abstract rules computed by the PFC could inform task and memory spaces, with active maps coordinated by cognitive demand.
STAR✩METHODS
RESOURCE AVAILABILITY
Lead contact
Further information and requests should be directed to the lead contact, Justin Riceberg (jriceberg@gmail.com).
Materials availability
This study did not generate new unique reagents.
Data and code availability
Data acquired and code developed in this study are made publicly available at https://github.com/learningandmemorylab/Curr_Bio_2022, with more detail available upon reasonable request.
EXPERIMENTAL MODEL AND SUBJECT DETAILS
Adult male Long-Evans rats (N = 4) weighed 275–300 grams at the start of experiments and were housed individually in a colony room with a 12h light/dark cycle. The rats were acclimated to the colony for a week, food-restricted to no less than 80% of their ad libitum body weight, and maintained on a restricted diet for the duration of the experiment. After 5 days of food restriction, animals were handled by the experimenter for 20 minutes per day for 5 days to acclimate them to human contact. Subjects were implanted with tetrode bundles targeting neurons in the ventrolateral orbitofrontal cortex (OFC) and CA1 sub-region of the dorsal hippocampus (Figure 1A). All procedures were performed in accordance with Institutional Animal Care and Use Committee guidelines and those established by the National Institutes of Health.
METHOD DETAILS
Behavioral testing
Apparatus
A plus-shaped maze made of wood was painted gray with four arms (59 cm long, 6.5 cm wide, with edges 2 cm high) that met in the center at 90° angles. Food cups at the end of each arm were recessed 0.5 cm, and an inaccessible piece of food was kept below a mesh screen. Each of two opposing arms was designated as North and South start arms; the two orthogonal arms were designated East and West goal arms. A rectangular waiting platform made of white-painted wood (30×35 cm) stood beside the maze in a room with several visual cues on the walls. The platform and the maze were elevated 81 cm above the floor.
Behavioral training
The rats were handled for 5 days and acclimated to the testing room and the maze where they found scattered chocolate sprinkles. Inaccessible food rewards were distributed throughout the maze to minimize the use of odor cues to guide foraging. The rats were then trained in a spatial win-stay task. Each trial began when the experimenter placed food in the designated goal arm, picked up a rat from the waiting platform, and placed it at the end of a pseudorandomly selected start arm facing away from the maze center. The experimenter used similar movements while placing food in the food cups to avoid cueing its location to the rat. The rats were trained initially to find chocolate sprinkles at the end of the West arm. Self-correction was permitted during this stage. After the rat consumed the reward, the experimenter picked up the rat and placed it on the waiting platform for 5–10 seconds between trials. After a rat made 8 consecutive correct “Go West” responses it was assigned to surgery (Figure 1).
Spatial reversal tasks
As in pre-training, the experimenter put food in the west goal arm. Each trial began when the rat was placed in a start arm and ended after the rat either consumed the food or reached the end of the unrewarded arm, when the rat was returned to the waiting platform. A choice was counted when the rat put all four paws into a goal arm; self-correction was not permitted. Each rat was given 5–8 days of discrimination retraining, and tetrodes were lowered concurrently toward CA1 and the OFC. Unit recording began after a rat met the 17/20 correct trial criterion in the “go West” discrimination and stable units were detected in both brain regions. The same contingency held during the first recording sessions and initial block of reversal learning, which was divided into blocks of trials. In each block, the same goal arm was rewarded until the rat met criterion performance (trials to criterion, TTC), 17/20 correct trials (40–97 total trials). The opposite goal arm was rewarded in the next block of trials, so that each daily testing session included 2–4 blocks and 1–3 reversals. The dataset included 971 trials.
Learning speed was measured by TTC. Learning curves were quantified using a Bayesian expectation-maximization algorithm that uses all trials in a contingency block to calculate the probability, with confidence intervals (CI), that a rat will select the correct goal on each trial.17 The Smith algorithm was used to assign trials into either early learning or stable performance.6,12,13 Early learning included trials before the animal performed better than chance (95% CI ≤ 0.5). Stable performance included trials after this point (lower CI > 0.5). These operational definitions categorized trials for physiological analyses to compare interactions between the OFC and the hippocampus during standard levels of learning and stable performance.
Electrophysiology
Electrode drives and placement
Hyperdrives with 24 independently movable tetrodes were built in-house. Tetrodes were spun from 12.5 mm nichrome wire (Kanthal Precision Technologies), loaded into the hyperdrive, cut, and gold plated until the impedance on each wire was approximately 200 kΩ measured at 1000 Hz. During implantation, the electrode interface board (EIB 36 24TT; Neuralynx) was connected via 0.003” stainless steel wire to four ground screws distributed across the skull as well as two reference screws implanted above the cerebellum. The implant coordinates in mm from Bregma were CA1: AP 3.6, ML 2.0; OFC: AP +3.8, ML 2.2. See the section on general surgical procedures for more detail. The tetrodes were lowered 1.4 mm into the cortex after surgery and were not moved again for at least one week. The tetrodes were advanced slowly (mm/day) toward the recording target, reaching either CA1 or OFC after 2.5–3 weeks. The proximity of CA1 tetrodes to the pyramidal cell layer was estimated by sharp wave / ripple profiles.41 Tetrodes were lowered to the OFC (~4 mm) by turning the microdrive screws ~14 times. To ensure recording stability, tetrodes were not adjusted for > 6 hours before behavior testing and recording.
Recording apparatus
Multiunit spike and LFPs were acquired using a Digital Lynx SX (Neuralynx). An electrode interface board connected tetrodes to a headstage containing unity gain amplifiers to minimize cable motion artifact; the headstage was connected to the system amplifiers with thin-wire tethers. Unit activity was filtered between 600 and 9000 Hz and digitized at 32,000 Hz prior to online spike detection. For each tetrode, amplitude thresholds were manually set on each wire to maximize the signal-to-noise ratio for spike detection. When the amplitude on a single wire rose above the threshold, the signal waveform around the threshold crossing was saved (along with a timestamp) for all four wires. The waveforms were sorted offline into single units. LFPs signals were sampled continuously at 2000 Hz with a band-pass filter (1–512 Hz) and recorded with the active electrode referenced directly to a skull screw implanted above the cerebellum or to another tetrode implanted in the brain. In the latter case, the reference tetrode was also recorded with respect to a skull screw implanted above the cerebellum so that the signal from the active electrode with respect to the skull could be recovered offline by subtraction. The position of LEDs mounted on the headstage was recorded by an overhead video camera, digitized (30 Hz, 640×480 pixels), converted to time-stamped XY coordinates, and stored for offline analysis by the Cheetah recording system.
Surgical procedures
Each animal was given its daily allotment of rat chow ~2 hr before surgery, and was then anesthetized in a Plexiglas chamber with 5% isoflurane delivered at a rate of 1 L/min. Once deeply anesthetized, the rat was given ketoprofen subcutaneously (3 mg/kg; Sigma) to minimize postoperative pain, and its head was shaved and positioned in a Kopf stereotax. Isoflurane (1%–3%) delivered via a nose cone maintained anesthesia during surgery. The scalp was cleaned with Povidone-Iodine (Dynarex) and anesthetized with 0.7 cc of lidocaine / epinephrine (0.5% / 1:200,000; Hospira). Core body temperature was monitored and maintained using a rectal probe and heating pad (part number: ATC 1000; World Precision Instruments). Sterile normal saline (1 cc) was delivered subcutaneously every hour to maintain hydration, and an ophthalmic ointment (Puralube vet ointment; Dechra) was applied to the animal’s eyes and reapplied hourly; additionally, the eyes were covered for the duration of surgery. Twenty min after ketoprofen administration, the scalp was resected and the skull was cleaned using distilled water and a dilute solution of hydrogen peroxide. Seven holes were drilled into the thicker parts of the skull for stainless steel bone screws (part number: 40-77-8; FHC) to stabilize the implants and serve as ground or reference. The skull was cleaned again and lambda and bregma were set to the same DV level. Target sites were measured and marked; thin layers of Metabond (Parkell) and Panavia (Kuraray) were applied to the skull to increase implant stability, excluding the areas around the implant site. Burr holes were drilled through the skull above the implant target sites (in mm from Bregma, CA1: −3.6AP, 2.0ML; OFC: 3.8AP, 2.2ML) to expose the dura, and a stereo microscope was used to ensure that it was fully exposed and clean. The dura was incised with micro-scissors and retracted, and the exposed cortex was kept clean and hydrated with normal saline until the implant was lowered into place. The craniotomy around the implant was sealed with Kwik-Cast (WPI), and the entire implant was then secured to the skull with dental acrylic (Coltene/Whaledent).
After surgery, animals were returned to a clean home cage containing a wet mash of rat chow. For the three days following the procedure the mash was infused with a Meloxicam suspension (1 mg/kg; Boehringer Ingelheim) to minimize pain due to swelling, and the rats were monitored to ensure they consumed the food. Animals generally responded well to this treatment and returned to preoperative levels of activity after four days. Animals were not handled by the experimenter for 7 days post-surgery.
QUANTIFICATION AND STATISTICAL ANALYSIS
Behavioral analyses
Spatial task behavioral performance is illustrated in Figure 1. Maze behavior was categorized in four possible journey types: northeast (NE), north-west (NW), south-east (SE), and south-west (SW). To analyze spatial firing correlates, the X,Y video coordinates stored for each journey was converted to a linear sequence of distances from the starting point. The video coordinates of reliable trajectories for all trials of a given journey type were fit with a polynomial ranging between order 5 and 7 and used to derive a canonical trajectory onto which data from individual trials were projected. Goal choices in single trials were predicted by unit activity recorded on the start arm from 22 to 54 cm marks, after the rat turned toward the maze center and before it entered the choice point. To determine the extent to which firing correlates were influenced by different behaviors on the start arm, we analyzed differences in heading angles and running speed. Heading angles were assessed in 1 cm increments along the segment of the start arm (25 to 55 cm) used for the decoding analyses described below. The mean heading angle was calculated at each position for each journey type, and the differences between East and West journeys were calculated for the North and South start arms separately. A null difference distribution was generated by shuffling trial labels for each start arm and recalculating the heading angle differences 1000 times. Visual inspection of the data revealed that the differences were von Mises distributed, and the parameters of the distribution were calculated accordingly.16,42 The probability that the actual heading angle difference was obtained by chance alone was evaluated according to the null difference PDF, and the threshold for a significant difference was set at alpha = 0.05, false-discovery-rate corrected at each position. Running speed was assessed on a trial-by-trial basis to generate a distribution of velocities at each point on the maze. Positions on the maze within a trial in which animal’s speed fell outside of two standard deviations of the mean for that position were removed from subsequent analyses.
Spike sorting
Spikes recorded on individual tetrodes were clustered into functional units, with single neurons as putative sources. Spike waveform parameters describing the shapes and relative amplitudes of the waveforms across all four tetrode wires were selected to define the basis for a space in which clustering was performed. Semi-automatic clustering was performed in several steps starting with KlustaKwik (http://klustakwik.sourceforge.net/). After noise clusters (e.g., from chewing artifact) were manually removed, the remaining data were automatically clustered again and then edited to identify well-segregated clusters of spikes that were assigned to functional units. The stability of individual units was assessed using a combination of waveform features and firing rates. The mean spike amplitude for a given unit was calculated in 20 equal temporal bins spanning the entire recording session for each tetrode wire. Any unit showing a significant Spearman rank correlation between time and amplitude on any wire (alpha = 0.05, uncorrected) and an accompanying drift in mean firing rate (alpha = 0.05, uncorrected) was rejected from further analysis. Visual inspection revealed that the statistical approach was conservative and units that seemed stable by eye were rejected; the included units had markedly stable waveforms.
Units with stable waveforms were clustered into putative pyramidal cell and interneuron groups separately for each region analyzed based on waveform features and firing rate.41,43 Spike asymmetry and firing rate were the best discriminators, with interneurons having higher firing rates and more asymmetric waveforms and making up approximately 5% of the total number of identified units. The total dataset included 14 OFC ensembles totaling 279 units (202 stable putative pyramidal cells), and 14 CA1 ensembles totaling 222 units (150 stable putative pyramidal cells). All analyses of neural representations included putative pyramidal units and excluded units with firing rates < 2% of the unit with the highest firing rate. The number of units included in each analysis is described in the relevant sections of the main text.
Support vector machines
Support vector machines44 (SVMs) quantified the extent to which activity in OFC and CA1 ensembles recorded in the start arm distinguished between animals’ pending goal choices. SVMs were trained to distinguish two classes of data (‘go east’ vs. ‘go west’; MATLAB statistics toolbox R2018b, Mathworks). Ensembles had to include >2 units that fired on the start arm for SVM analysis. For each such ensemble (14 OFC, 7 CA1), single-unit firing rates were z-score normalized across trials, and each trial’s vector was normalized to unit length. To account for doublet and triplet interactions among units, the SVMs were fit using an inhomogeneous polynomial kernel, the effect of which is equivalent to replacing each trial vector with the most general second or third degree polynomial in the individual unit firing rates, and each SVM was then optimized to maximize the leave-one-out cross validation accuracy (see below). The mean expansion order required for goal prediction was 2.24 (s.d. = 1.15) for CA1 activity, and 2.44 (s.d. = 1.3) for OFC activity. A given SVM was fit to a dataset with the exception of a single trial, which was subsequently input to the SVM to classify. The cross-validation accuracy was the proportion of trials correctly classified. To determine if the SVMs were classifying trial goals by discovering task structure or unrelated noise, we repeated the above procedure using the same data and parameter sets after shuffling the trial labels (e.g., the current goal or current start arm). The leave-one-out cross-validation decoding accuracy following 1000 separate shuffles of the data determined if the observed decoding accuracy was better than chance, defined here as > 95% of the shuffles.
Spike-phase locking to CA1 theta rhythm
The preferred firing phase of units and populations were measured with respect to the hippocampal theta oscillation by calculating the phase angle coincident with each spike. Units that fired ≥21 spikes on the start arm were included in the phase locking analysis. The distribution of phase angles for each unit or population of spikes was quantified by the Rayleigh test for non-uniformity. Filtered CA1 LFPs (4–12 Hz) were Hilbert transformed to extract the instantaneous phase of the theta oscillation, and the endpoints of each theta cycle were detected as phase ‘wraparounds’ from pi to –pi. The phase corresponding to each spike time was interpolated linearly between these endpoints to avoid the non-uniform distribution of Hilbert transform values caused by the sawtooth shape of theta oscillations. For the population locking, each eligible unit’s mean phase was considered, to avoid biasing the population statistics from differential firing rates and hence spike-count of different units. The pairwise-phase consistency (PPC) method17 verified Rayleigh tests for phase uniformity. Briefly, each spike is assigned a theta phase (as above) and the consistency of phase-offsets between all spike pairs is calculated (see equations 8–10 in Vinck et al.17). To assess statistical significance, the resulting PPC values were compared to PPCs generated from random spike phases assigned to identical spike counts.
Coordinated spike timing: state space analysis
Learning depends upon neural plasticity that alters both the rate and timing of action potentials. To assess the extent to which learning modified precise spike timing, we used a nonlinear recursive Bayesian method20 to quantify synchronized activity in pairs of neurons recorded in the start arm and measured how coactivity changed as rats learned and performed spatial discriminations and reversals. Ensemble spike trains of N units were subdivided into bins defined by single theta cycles, and the spike synchrony rate of each unique pair of units yI,j was computed for each bin time bin t as
where i; j are single units ∈ N, fi;j is the product xi * xj, where xn are binary elements = 1 if unit n spiked within bin t, 0 if not, and Xt;l is the number of N-tuple binary variables in bin t of trial l (from Equation 6 in Shimazaki et al.20). Each coactivity bin was used to update the parameters of the log-linear model (Equation 1 in Shimazaki et al.20).
Because the model analyzed spike pairs, logp(x) describes the probability that a given pair of units fired synchronously, the θa parameters correspond to the activation state of the spiking pattern for the single unit xa. To distinguish between terms, we will use q to indicate coactivity and theta to indicate the 4–12 Hz LFP. The θab parameters correspond to pairwise coactivity between unit a and unit b, and j(q) is a normalization function. The pairwise coactivity in each theta cycle was represented by a matrix of activation states, with each matrix element describing the activation state of a given unit pair. E.g., q1 Is the activation state of unit 1, qN is the activation state of unit N, and q1;N is the pairwise coactivity of units 1 and N.
The mean and standard deviation of the unique set of pairwise activity states (θij) in the lower triangular matrix in each theta cycle was used to exclude values less than 2 standard errors of the mean. For each trial, the included coactivity values were averaged across theta cycles to compute the probability that pairs of units fired within the same theta cycle. The mean and standard deviation of θij was calculated separately for pairs of CA1 units, OFC units, and mixed pairs of CA1 and OFC units. Learning-related changes were assessed by comparing coactivity in trials assigned to one of five learning stages defined by the Smith algorithm with a 4-way ANOVA (task phase: initial discrimination/reversal; reversal type: first/familiar; learning stages: 1–5; and pair type: OFC-OFC, CA1-CA1, or OFC-CA1). To test if firing rates within structure accounted for the coactivity, firing rates in each structure were analyzed similarly (S3B). In contrast to theta (Figures 3C and 3D) and gamma (S3C), time bins based on delta (~250ms bins) rhythms failed to converge.
Granger prediction
Granger prediction15 assesses the temporally directed statistical relationship between two time series by testing if the recent history of one time series predicts changes in a target time series beyond that predicted by the history of the target series itself. The Granger value is the log ratio of the residual variances for the model incorporating only one time series (the target) and the model incorporating both. The higher the Granger value, the greater the prediction gained by including the second time series. Prediction does not entail causation, however, and Granger values do not identify the neural mechanisms that drive observed predictions. Firing rates of individual units were z-score normalized across trials and smoothed (1/2 trial standard deviation) to minimize variance due to place coding. The smoothing procedure linearly interpolated missing trials (e.g., due to poor video tracking), padded individual units’ trialwise rate vectors with a time-reversed copy of itself, and convolved each vector with a Gaussian via multiplication in the frequency domain. The padding procedure prevented mixing of information from the first and last trials of the recording session. The data were transformed back into the time domain, the padded portion of the vector was discarded, and interpolated data points removed. From the smoothed data, an activity state was defined for each trial as the dot product of its PV with the mean vector of the last two trials of the previous contingency using the kernel trick.12 We assessed the three trials that immediately followed a rule change, and used the activity in the first two trials to predict activity in the third, then the next two to predict the fourth, and so on, concatenating the dependent and independent variables for the model.
Granger predictions were compared to chance using permutation tests that shuffled the trial order of the non-target series and calculating the Granger value 1000 times to generate a null distribution. By removing the temporal correspondence between the two series, shuffling tested the extent to which any increase in explained variance of the target was due to the additional independent variables provided by the non-target series. Actual Granger values were compared to the shuffled values to calculate the corresponding probability. A similar procedure assessed differences between Granger values: the within-trial correspondence between sequences was preserved (e.g., CA1t-1, CA1t-2, OFCt-1, OFCt-2), and observations were shuffled between models prior to calculating Granger values (e.g., from the model testing OFC’s influence on CA1 to the model testing CA1’s influence on OFC). The difference in Granger values calculated after each shuffle generated a null difference distribution against which observed differences were compared. Floor or ceiling effects were not detected since the null distribution values revealed a broad range of granger values (both higher and lower than the actual).
Granger prediction: Bayes information criterion
The dependent variable (DV) for each observation was the similarity of a given trial’s ensemble activity to that observed at the end of the previous learning epoch (i.e., ID, R1, or R2). The independent variables (iVs) were the similarity measures for the two previous trials for either the target time series alone or both time series. The number of iVs (2) was determined by minimizing the Bayes’ information criterion (BIC) with respect to trial number. The BIC aids in model selection by balancing model fit with parsimony (i.e., the number of parameters). The BIC was calculated using the following equation:
where n is the number of observations in the model, RSS is the residual sum of squares, and k is the number of parameters in the model that includes both time series.
BAYESIAN POSITION DECODING
To estimate how well OFC and CA1 ensembles represented locations on a behavioral timescale, we used a Bayesian classifier19,45 topredict the animal’s position from the firing rates of each cell measured every 125 milliseconds. The posterior probability (Prob) of the animal’s position (pos) across M total positions given a time window t containing spikes (spikes) from N units is
where j is position ∈ M,
fi(pos) is the expected firing rate of unit i in position pos, and ni is the number of spikes from unit i in the time window t, here set to 33 milliseconds. The rats’ horizontal and vertical coordinates from the video tracker were expressed in terms of distance from the start of each journey. The linear paths in each arm were then divided into 2cm bins and concatenated, and the Bayesian classifier assigned the probability of the animal’s position in each bin based on the spikes occurring during each time window. Position reconstruction error was defined as the sum of the diagonal in the actual vs. decoded probability matrix divided by the sum of all probabilities. Alternate-arm-location decoding (3) was defined as the difference between decoding in current location and the corresponding location in the other start arm or goal arm. For example, we compared the decoded probability that the rat was in the 5th bin of the South start arm when it occupied that location to the decoded probability that the rat was in the 5th bin of the North start arm (Figure 4). A decoding index quantified the difference between the actual and corresponding-arm-location representations for each bin: 1-[(actual-other)/(actual+other)], and ANOVAs compared decoding indices for trials in different learning phases.
Spatial information
The spatial information (bits/spike) was calculated according to Skaggs and Mcnaughton:46
where the linearized maze was divided into nonoverlapping 1-cm spatial bins i=1:N, pi is the occupancy probability of bin i, li is the mean firing rate for bin i, and l is the overall mean firing rate of the unit.
Inter-spike intervals
Inter-spike intervals (ISI) were calculated between co-active CA1-OFC unit pairs for each learning phase (early vs stable) and session (first vs familiar). Qualified units fired at least 20 start-arm spikes in a given condition. Only ISIs <400ms were considered for analysis.
Supplementary Material
KEY RESOURCES TABLE
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Chemicals, peptides, and recombinant proteins | ||
| Forane (isoflurane) | Baxter | NDC: 10019-360-60 |
| Ketoprofen | Sigma | K1751 |
| Lidocaine/Epinephrine | Hospira | NDC: 0409-0996-01 |
| Puralube Vet Ointment | Dechra | NDC: 17033-211-38 |
| Metabond | Parkell | S380 |
| Panavia | Kuraray | 488KA |
| Experimental models: Organisms/strains | ||
| Long Evans Rats | Charles River | 006 |
| Software and algorithms | ||
| Matlab 2018b | Mathworks | RRID: SCR_001622 |
| Python version 3.8 | Python Software Foundation | RRID: SCR_008394 |
| Other | ||
| 96 Channel DigitalLynx 16SX | Neuralynx | https://neuralynx.com/products/digital_data_acquisition_systems/digital_lynx_16sx |
| 12.5 um nichrome wire | Kanthal Precision Technology | RO-800 |
| Stainless steel bone screws | FHC | 40-77-8 |
Highlights.
CA1 ensembles guide evolving OFC representations when learning a new path
OFC activity synchronizes to CA1 theta rhythms when learning a new path
CA1-OFC unit coactivity predominates early learning and fades thereafter
OFC goal representations generalize after learning
ACKNOWLEDGMENTS
We thank Maojuan Zhuang, Linda Barenboim, Pablo Martin, and Pierre Enel for technical support and Michael Goodman, Erin Rich, and Ioana Carcea for valuable comments and suggestions. This work was supported by NIMH grants MH073689 and MH065658, the Icahn School of Medicine at Mount Sinai, and Albany Medical College.
Footnotes
SUPPLEMENTAL INFORMATION
Supplemental information can be found online at https://doi.org/10.1016/j. cub.2022.06.010.
DECLARATION OF INTERESTS
The authors declare no competing interests.
REFERENCES
- 1.Schoenbaum G, Roesch MR, Stalnaker TA, and Takahashi YK (2009). A new perspective on the role of the orbitofrontal cortex in adaptive behave-iour. Nat. Rev. Neurosci 10, 885–892. 10.1038/nrn2753. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Wikenheiser AM, and Schoenbaum G (2016). Over the river, through the woods: cognitive maps in the hippocampus and orbitofrontal cortex. Nat. Rev. Neurosci 17, 513–523. 10.1038/nrn.2016.56. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Zhou J,Montesinos-Cartagena M,Wikenheiser AM,Gardner MPH,Niv Y, and Schoenbaum G (2019). Complementary task structure representations in hippocampus and orbitofrontal cortex during an odor sequence task. Curr. Biol 29, 3402–3409.e3. 10.1016/j.cub.2019.08.040. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Riceberg JS, and Shapiro ML (2012). Reward stability determines the contribution of orbitofrontal cortex to adaptive behavior. J. Neurosci 32, 16402–16409. 10.1523/jneurosci.0776-12.2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Ferbinteanu J, and Shapiro ML (2003). Prospective and retrospective memory coding in the hippocampus. Neuron 40, 1227–1239. 10.1016/s0896-6273(03)00752-9. [DOI] [PubMed] [Google Scholar]
- 6.Guise KG, and Shapiro ML (2017). Medial prefrontal cortex reduces memory interference by modifying hippocampal encoding. Neuron 94, 183–192.e8. 10.1016/j.neuron.2017.03.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.van Wingerden M, Vinck M, Lankelma J, and Pennartz CMA (2010). Theta-band phase locking of orbitofrontal neurons during reward expectancy. J. Neurosci 30, 7078–7087. 10.1523/jneurosci.3860-09.2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Knudsen EB, and Wallis JD (2020). Closed-loop theta stimulation in the orbitofrontal cortex prevents reward-based learning. Neuron 106, 537–547.e4. 10.1016/j.neuron.2020.02.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.O’Keefe J, and Dostrovsky J (1971). The hippocampus as a spatial map. Preliminary evidence from unit activity in the freely-moving rat. Brain Res 34, 171–175. 10.1016/0006-8993(71)90358-1. [DOI] [PubMed] [Google Scholar]
- 10.Wilson RC, Takahashi YK, Schoenbaum G, and Niv Y (2014). Orbitofrontal cortex as a cognitive map of task space. Neuron 81, 267–279. 10.1016/j.neuron.2013.11.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Gardner MPH, and Schoenbaum G (2021). The orbitofrontal cartographer. Behav. Neurosci 135, 267–276. 10.1037/bne0000463. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Rich EL, and Shapiro M (2009). Rat prefrontal cortical neurons selectively code strategy switches. J. Neurosci 29, 7208–7219. 10.1523/jneurosci.6068-08.2009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Smith AC, Frank LM, Wirth S, Yanike M, Hu D, Kubota Y, Graybiel AM, Suzuki WA, and Brown EN (2004). Dynamic analysis of learning in behavioral experiments. J. Neurosci 24, 447–461. 10.1523/jneurosci.2908-03.2004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Young JJ, and Shapiro ML (2011). Dynamic coding of goal-directed paths by orbital prefrontal cortex. J. Neurosci 31, 5989–6000. 10.1523/jneurosci.5436-10.2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Cohen MX (2014). Analyzing Neural Time Series (MIT Press; ). [Google Scholar]
- 16.Fisher NI (1993). Statistical Analysis of Circular Data (Cambridge University Press; ). [Google Scholar]
- 17.Vinck M, van Wingerden M, Womelsdorf T, Fries P, and Pennartz CMA (2010). The pairwise phase consistency: a bias-free measure of rhythmic neuronal synchronization. Neuroimage 51, 112–122. 10.1016/j.neuroimage.2010.01.073. [DOI] [PubMed] [Google Scholar]
- 18.Shapiro ML, and Ferbinteanu J (2006). Relative spike timing in pairs of hippocampal neurons distinguishes the beginning and end of journeys. Proc. Natl. Acad. Sci. USA 103, 4287–4292. 10.1073/pnas.0508688103. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Pfeiffer BE, and Foster DJ (2013). Hippocampal place-cell sequences depict future paths to remembered goals. Nature 497, 74–79. 10.1038/nature12112. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Shimazaki H, Amari S.i., Brown EN, and Grün S (2012). State-space analysis of time-varying higher-order spike correlation for multiple neural spike train data. PLoS Comp. Biol 8. e1002385. 10.1371/journal.pcbi.1002385. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Larson J, Wong D, and Lynch G (1986). Patterned stimulation at the theta frequency is optimal for the induction of hippocampal long-term potentiation. Brain Res 368, 347–350. 10.1016/0006-8993(86)90579-2. [DOI] [PubMed] [Google Scholar]
- 22.MacDonald CJ, Lepage KQ, Eden UT, and Eichenbaum H (2011). Hippocampal “time cells” bridge the gap in memory for discontiguous events. Neuron 71, 737–749. 10.1016/j.neuron.2011.07.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Eichenbaum H (2018). What versus where: non-spatial aspects of memory representation by the hippocampus. Curr. Top. Behav. Neurosci 37, 101–117. 10.1007/7854_2016_450. [DOI] [PubMed] [Google Scholar]
- 24.Kennedy PJ, and Shapiro ML (2009). Motivational states activate distinct hippocampal representations to guide goal-directed behaviors. Proc. Natl. Acad. Sci. USA 106, 10805–10810. 10.1073/pnas.0903259106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Wood ER, Dudchenko PA, Robitsek RJ, and Eichenbaum H (2000). Hippocampal neurons encode information about different types of memory episodes occurring in the same location. Neuron 27, 623–633. 10.1016/s0896-6273(00)00071-4. [DOI] [PubMed] [Google Scholar]
- 26.Frank LM, Brown EN, and Wilson M (2000). Trajectory encoding in the hippocampus and entorhinal cortex. Neuron 27, 169–178. 10.1016/s0896-6273(00)00018-0. [DOI] [PubMed] [Google Scholar]
- 27.Aronov D, Nevers R, and Tank DW (2017). Mapping of a non-spatial dimension by the hippocampal-entorhinal circuit. Nature 543, 719–722. 10.1038/nature21692. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Zhou J, Jia C, Montesinos-Cartagena M, Gardner MPH, Zong W, and Schoenbaum G (2021). Evolving schema representations in orbitofrontal ensembles during learning. Nature 590, 606–611. 10.1038/s41586-020-03061-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.McDannald MA, Lucantonio F, Burke KA, Niv Y, and Schoenbaum G (2011). Ventral striatum and orbitofrontal cortex are both required for model-based, but not model-free, reinforcement learning. J. Neurosci 31, 2700–2705. 10.1523/jneurosci.5499-10.2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Abe H, and Lee D (2011). Distributed coding of actual and hypothetical outcomes in the orbital and dorsolateral prefrontal cortex. Neuron 70, 731–741. 10.1016/j.neuron.2011.03.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Takahashi YK, Chang CY, Lucantonio F, Haney RZ, Berg BA, Yau HJ, Bonci A, and Schoenbaum G (2013). Neural estimates of imagined outcomes in the orbitofrontal cortex drive behavior and learning. Neuron 80, 507–518. 10.1016/j.neuron.2013.08.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Steiner AP, and Redish AD (2014). Behavioral and neurophysiological correlates of regret in rat decision-making on a neuroeconomic task. Nat. Neurosci 17, 995–1002. 10.1038/nn.3740. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Sommer T, Peters J, Gläscher J, and Büchel C (2009). Structure-function relationships in the processing of regret in the orbitofrontal cortex. Brain Struct. Funct 213, 535–551. 10.1007/s00429-009-0222-8. [DOI] [PubMed] [Google Scholar]
- 34.Schoenbaum G, Nugent SL, Saddoris MP, and Setlow B (2002). Orbitofrontal lesions in rats impair reversal but not acquisition of go, no-go odor discriminations. NeuroReport 13, 885–890. 10.1097/00001756-200205070-00030. [DOI] [PubMed] [Google Scholar]
- 35.Zhou J, Gardner MPH, Stalnaker TA, Ramus SJ, Wikenheiser AM, Niv Y, and Schoenbaum G (2019). Rat orbitofrontal ensemble activity contains multiplexed but dissociable representations of value and task structure in an odor sequence task. Curr. Biol 29, 897–907.e3. 10.1016/j.cub.2019.01.048. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Bahar AS, Shirvalkar PR, and Shapiro ML (2011). Memory-guided learning: CA1 and CA3 neuronal ensembles differentially encode the commonalities and differences between situations. J. Neurosci 31, 12270–12281. 10.1523/jneurosci.1671-11.2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Jezek K, Henriksen EJ, Treves A, Moser EI, and Moser MB (2011). Theta-paced flickering between place-cell maps in the hippocampus. Nature 478, 246–249. 10.1038/nature10439. [DOI] [PubMed] [Google Scholar]
- 38.Miller EK, and Cohen JD (2001). An integrative theory of prefrontal cortex function. Annu. Rev. Neurosci 24, 167–202. 10.1146/annurev.neuro.24.1.167. [DOI] [PubMed] [Google Scholar]
- 39.Cape EG, and Jones BE (2000). Effects of glutamate agonist versus procaine microinjections into the basal forebrain cholinergic cell area upon gamma and theta EEG activity and sleep-wake state. Eur. J. Neurosci 12, 2166–2184. 10.1046/j.1460-9568.2000.00099.x. [DOI] [PubMed] [Google Scholar]
- 40.Tolman EC (1948). Cognitive maps in rats and men. Psychol. Rev 55, 189–208. 10.1037/h0061626. [DOI] [PubMed] [Google Scholar]
- 41.Csicsvari t., Hirase H, Czurkó A, Mamiya A, and Buzsáki G (1999). tFast Network Oscillations in the Hippocampal CA1 Region of the Behaving Rat. J. Neurosci 19, RC20. 10.1523/jneurosci.19-16-j0001.1999. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Riceberg JS, and Shapiro ML (2017). Orbitofrontal cortex signals expected outcomes with predictive codes when stable contingencies promote the integration of reward history. J. Neurosci 37, 2010–2021. 10.1523/jneurosci.2951-16.2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Barthó P, Hirase H, Monconduit L, Zugaro M, Harris KD, and Buzsáki G (2004). Characterization of neocortical principal cells and interneurons by network interactions and extracellular features. J. Neurophysiol 92, 600–608. 10.1152/jn.01170.2003. [DOI] [PubMed] [Google Scholar]
- 44.Smola AJ, and Schölkopf B (2004). A tutorial on support vector regression. Stat. Comput 14, 199–222. 10.1023/b:stco.0000035301.49549.88. [DOI] [Google Scholar]
- 45.Davidson TJ, Kloosterman F, and Wilson MA (2009). Hippocampal replay of extended experience. Neuron 63, 497–507. 10.1016/j.neuron.2009.07.027. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Skaggs WE, and Mcnaughton BL (1996). Replay of neuronal firing sequences in rat hippocampus during sleep following spatial experience. Science 271, 1870–1873. 10.1126/science.271.5257.1870. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data acquired and code developed in this study are made publicly available at https://github.com/learningandmemorylab/Curr_Bio_2022, with more detail available upon reasonable request.
