Summary
Endogenous neuropeptides are uniquely poised to regulate neuronal activity and behavior across multiple timescales. Traditional studies ascribing neuropeptide contributions to behavior lack spatiotemporal precision. The endogenous opioid dynorphin is a neuropeptide highly enriched in the dorsal striatum, a region critical for goal-directed behavior. However, the functional role of endogenous dynorphin-KOR signaling on goal-directed behavior is unknown. Here, we report that local, time-locked dynorphin release from dorsomedial striatum medium spiny neurons (MSNs) is necessary and sufficient for goal-directed behavior using a suite of modern approaches including conditional deletions, neuropeptide biosensor detection, two-photon imaging and time-locked optogenetic manipulations of neuropeptide release. We discovered that glutamatergic axon terminals from the basolateral amygdala evoke striatal dynorphin release, resulting in feed-forward retrograde presynaptic GPCR inhibition to promote behavior. Collectively, our findings isolate a causal role for endogenous neuropeptide release at rapid timescales, and subsequent GPCR activity for promoting goal-directed behavior.
Keywords: goal-directed behavior, dorsal striatum, neuropeptides, opioids, dynorphin, kappa opioid receptor, in vivo two-photon calcium imaging, in vivo fiber photometry
Graphical Abstract

In Brief:
Cell type–specific, time-locked release of the opioid neuropeptide dynorphin in the dorsomedial striatum is both necessary and sufficient for goal-directed behavior. Inputs from the basolateral amygdala trigger dynorphin release that engages retrograde Gi-coupled presynaptic inhibition via the kappa opioid receptor, revealing a rapid neuropeptide mechanism for shaping behavior.
Introduction
Most behaviors animals perform are “goal-directed”. At its simplest, goal-directed behavior is an animal’s ability to associate and sustain the predictive relationship between an action and a subsequent outcome. This behavior is fundamental to survival and relies on learning across multiple timescales – of seconds (fast) in which animals look for causal relationships between actions and outcomes, and days (slow) where animals learn the associated probabilities between these causal actions and outcomes. Dysfunctions in goal-directed behavior are implicated in neuropsychiatric disorders including obsessive compulsive disorder1, depression2, chronic stress3 and substance use disorders (SUDs)4.
The dorsal striatum plays an integral role in enabling goal-directed behavior in rodents5, non-human primates6,7 and humans8 with the medial portion of dorsomedial striatum (DMS) implicated most specifically9,10. Canonically, functional roles for goal-directed behavior are ascribed to largely two cell-types – the direct pathway medium spiny projection neurons (MSNs) expressing the dopamine D1 receptor/dynorphin (D1 MSNs) and the indirect pathway neurons expressing the dopamine D2 receptor/Enkephalin (D2 MSNs)11,12, implicated in goal-directed action-outcome behaviors13–16. However, either the timescale of manipulation studies do not match the behavior or lack the capabilities to determine activity across large populations of DMS neurons.
Neuropeptides are expressed widely in the brain and canonically thought to be released in a diffuse manner17, exerting their effects via their cognate G protein-coupled receptors (GPCRs), engaging ion channels (fast) and cAMP production (slow) to impact neuronal plasticity, and gene transcription18. The neuropeptide dynorphin is expressed broadly across the brain19 and exclusively in D1 MSNs in the striatum20, and has been implicated in stress regulation, dysphoria, anxiety-like behavior and drug-seeking behavior21–24. Dynorphin causes long-lasting changes in brain activity via the inhibitory Gαi-coupled kappa opioid receptor (KOR)25 and dampens neuronal activity over slower timescales (minutes to hours)22. However, Gαi-coupled G protein-coupled receptor signaling, can also work at faster timescales via coupling to calcium channels and potassium channels26,27. This inhibitory function could be highly relevant for stabilizing animal behavioral sequences, as well as promoting learning through retrograde feedback28,29.
Gαi -coupled GPCRs are expressed widely across different neuronal compartments and influence neuronal activity via both pre and post-synaptic mechanisms30. KOR is expressed on dendrites, cell bodies, and presynaptic axon terminals, suggesting dyn-KOR signaling could modulate postsynaptic neuron activity and/or presynaptic neurotransmitter release31,32. In vitro electrophysiology studies have shown that KOR activation in the striatum inhibits the presynaptic release from a variety of inputs33,34, including D1 MSNs via dyn-KOR control of the basolateral amygdala (BLA) inputs35. KOR mRNA is enriched in >50% of BLA neurons36. While decades of work has established the BLA as an important hub for goal-directed learning37–41, and recent studies suggest a role for DMS-projecting BLA neurons in behavior41,42, whether dynorphin-KOR signaling influences BLA-DMS projections is unclear.
Here, we sought to decode how neuropeptides dynamically control essential inputs to the basal ganglia to shape goal-directed behavior. We report that dynorphin release in the DMS at strikingly fast timescales promotes goal-directed behavior. We observe that rapid fluctuations in DMSpdyn neuron activity and dynorphin release evolve as animals learn goal-directed behavior and encode subsequent goal-directed action in a trial-to-trial manner. Finally, we isolate a foundational mechanism by which inhibitory neuropeptide-GPCR signaling influences incoming excitatory neuronal activity from the BLA, providing a homeostatic mechanism for the DMS to stabilize goal-directed behaviors.
Results
DMS dynorphin is necessary for acquiring goal-directed behavior.
We first determined the necessity of dynorphin for acquiring goal-directed behavior (Fig. 1A). Trained in a self-paced operant conditioning paradigm, mice progressed toward making significantly higher active and lower inactive nosepokes on day 5 (trained), compared to day 1 (early), while consistently consuming sucrose on all trials. Mice then underwent extinction training during which sucrose delivery was omitted on all trials; mice rapidly decreased their active responses and approaches to the sucrose receptacle across 5 days of extinction. Collectively, these features were represented as an increase in a summary metric, the “operant index”, accounting for action vigor, action discrimination and outcome consumption (See Methods for detail).
Figure 1. DMS dynorphin is necessary for acquiring goal-directed behavior.

(A) Schematic of viral injection in the DMS and operant behavior.
(B) Confocal images of Ctrl (top) and DMSpdyn-cKO (bottom) with ISH for DAPI (blue) and DRD1 (green), and DAPI (blue) and Pdyn (majenta) for sections of Striatum (left, 10X), DMS (middle, 40X) and Nac (right), and Quantification of ISH.
(C) Operant learning (Simple Linear Regression: n=9 Ctrl mice, R2=0.5782, p<0.0001****. n=12 DMSpdyn-cKO mice, R2=0.2808, p=0.0009***. Difference in Slopes – p=0.0229*).
(D) Operant Reversal (Simple Linear Regression: n=5 Ctrl mice, R2=0.7663, p<0.0001****. n=5 DMSpdyn-cKO mice, R2=0.7333, p<0.0001****. Difference in Slopes – p=0.0133*).
(E) Operant FR-3 (Simple Linear Regression: n=5 Ctrl mice, R2=0.2814, p=0.0419*. n=5 DMSpdyn-cKO mice, R2=0.04635, p>0.05. Difference in Intercepts – p=0.0029##).
(F) Operant Extinction (Simple Linear Regression: n=5 Ctrl mice, R2=0.7928, p<0.0001****. n=5 DMSpdyn-cKO mice, R2=0.1324, p>0.05. Difference in Slopes – p=0.0086**).
(G) Operant U50,488 injection (n=5 Ctrl, DMSpdyn-cKO mice; Two Way ANOVA, genotype x treatment p<0.0037**. Multiple comparisons – Ctrl veh vs. U50, p>0.05; DMSpdyn-cKO veh vs. U50, p=0.0003***).
(H) Schematic of viral injection and fiber implant in the DMS.
(I) Left – Learning (Simple Linear Regression: n=6 Ctrl mice, R2=0.8036, p<0.0001****. n=5 DMSpdyn-cKO mice, R2=0.3523, p=0.0197*. Difference in Slopes – p=0.0003***). Right - Extinction (Simple Linear Regression: n=6 Ctrl mice, R2=0.7947, p<0.0001****. n=5 DMSpdyn-cKO mice, R2=0.2682, p=0.0480*. Difference in Slopes – p=0.0011**)
(J-L) Mean fluorescence and heatmap raster plots during early, trained and extinction operant behavior (representative animals) for Ctrl (dark) and DMSpdyn-cKO (light).
(M) Normalized Peak z-score values during cue+outcome (2–10s) of Ctrl (Left: n=6 mice; One Way ANOVA, p=0.0004***. Multiple comparisons – early vs. trained, p<0.0001****, trained vs. extinction, p=0.0029**) and DMSpdyn-cKO (n=5 mice; One Way ANOVA, p>0.05).
(N) Peak z-score values during cue+outcome (2–10s) subtracted from a baseline (−5 to −1s) (n=6 Ctrl, 5 DMSpdyn-cKO mice; Two Way ANOVA, genotype x day p<0.0037**. Multiple comparisons – Early, p>0.05; Trained, p<0.0001****; Extinction, p=0.0359*).
Preprodynorphin, a precursor for dynorphin25 was conditionally removed from the DMS by bilateral injections of AAV5-CMV-Cre recombinase in Pdynlox/lox (DMSpdyn-cKO) or WT mice (Ctrl) (Fig. 1A, S1A). In situ hybridization of striatal slices showed a significant reduction of Pdyn mRNA selectively in the DMS from DMSpdyn-cKO mice, but not in the NAc, compared to Ctrl mice (Fig. 1B). DMSpdyn-cKO mice consumed the same amount of sucrose under food-restriction in their homecage (Fig. S1B). In contrast, they consumed fewer pellets during Pavlovian conditioning (Fig. S1C). Further, during operant conditioning, DMSpdyn-cKO mice were slower to learn goal-directed behavior (Fig. 1C), performing fewer operant actions and consuming less rewards compared to the controls (Fig. S1D). This deficit translated to other operant contingencies where DMSpdyn-cKO displayed deficits in both operant reversal learning (Fig. 1D, S1E) and fixed ratio-3 responding (Fig. 1E, S1F). Additionally, DMSpdyn-cKO mice extinguished their operant responding slower than Ctrl animals (Fig. 1F, S1G). Strikingly, a single i.p injection of a KOR agonist, U50,488 (5 mg/kg; 30 minutes prior to behavior) rescued these deficits in DMSpdyn-cKO mice while behavior in Ctrl mice remained unaffected (Fig. 1G, S1H).
Next, we determined when dynorphin is released in the DMS during goal-directed behavior, using a genetically-encoded fluorescent sensor (κLight1.3a) that we recently characterized to be selective for dynorphin and other KOR agonists43. We injected Pdynlox/lox (DMSpdyn-cKO) or WT mice (Ctrl) mice with AAV5-CAG-DIO-κLight1.3a and AAV5-CMV-Cre recombinase and implanted an optic fiber into the DMS (Fig. 1H, S1I). DMSpdyn-cKO mice showed significant reductions in operant learning and extinction (Fig. 1I, S1J–K). Before behavior, we characterized κlight1.3a sensor performance. Following an injection with U50,488 (i.p., 10 mg/kg), we observed significant agonist-induced increases in κLight-mediated fluorescence in both DMSpdyn-cKO and Ctrl mice, relative to their baseline period (0–5 minutes) (Fig. S1L–M). We determined κLight-dependent fluorescence as a measure of dynorphin release in Ctrl and DMSpdyn-cKO mice during goal-directed behavior. We observed time-locked elevations in κLight-dependent fluorescence after mice nosepoked and approached rewards in Ctrl mice (Fig. 1J–K, S1N). This increase was significantly higher following training and significantly diminished after extinction (Fig. 1M-left). Importantly, DMSpdyn-cKO mice showed no significant changes in κLight-dependent fluorescence across learning and extinction (Fig. 1J–N, 1M-right, S1N), demonstrating κLight selectivity to dynorphin. To determine if dynorphin release is not due to the cue, the same Ctrl mice underwent operant conditioning without the cue period, where rewards were delivered at a 2s delay following the nosepoke. Mice performed equivalent number of reinforced nosepokes and rewards without the cue (Fig. S1O–P). We found time-locked elevations in κLight-dependent fluorescence following a nosepoke in both conditions (Fig. S1Q). Although peak magnitudes were unchanged (Fig. S1R), the area under the curve was significantly reduced in mice that underwent operant conditioning without the cue (Fig. S1S). Altogether, our results indicate that increases in κLight-dependent fluorescence in the DMS in vivo are selective for dynorphin, evolve across operant learning and increase rapidly following action during the anticipation for reward.
DMS dynorphin release promotes goal-directed action.
Our results indicated that DMS dynorphin release is enhanced immediately following a goal-directed action. To further illuminate its contributions across trials, we injected KOR-cre mice with AAV5-CAG-DIO-κLight1.3a and implanted an optic fiber (Fig. 2A, S2A) into the DMS. Mice were screened with an injection of KOR agonist U50,488 (i.p., 10 mg/kg) alone, or a pre-treatment with a short-acting, reversible selective KOR antagonist aticaprant44 (i.p., 5 mg/kg) 30 minutes prior to U50,488. We observed significant, agonist-induced increases in κLight-mediated fluorescence, blocked by pre-treatment with aticaprant (Fig. S2B–C). Mice that showed significant elevations to U50 (6/8 mice) were food-restricted and underwent behavior. Mice showed no significant elevations in κLight-mediated fluorescence to the cue, reward delivery, or consumption during Pavlovian conditioning (Fig. S2D–F). Upon goal-directed learning and extinction (Fig. 2B, S2G), mice showed a significant increase in DMS dynorphin release in anticipation of the outcome that was rapidly extinguished following extinction (Fig. 2C–F). To uncover putative patterns in dynorphin release in a trial-to-trial manner within and across sessions (Fig. 2C–E), we trained a generalized linear regression model (GLM) to determine whether dynorphin release encoded specific experimental variables – trial identity (trial, green), preceding inter-trial interval (ITI, light purple), subsequent inter-trial interval (ITI+1, dark purple) and time to consume reward following a nosepoke (eat, pink) (Fig. 2G). Following model training, we dropped individual variables to ascertain the impact on model variance, with a significant reduction implying selective contributions of that variable. Full model variance on early sessions of all mice was poor, with no appreciable changes after dropping out individual variables (Fig. S2H). Remarkably, full model variance was significantly negatively impacted by only the subsequent ITI (ITI+1, dark purple) on trained sessions (Fig. 2H), further apparent as a significant negative correlation between the peak magnitude of dynorphin release and the subsequent ITI (ITI+1) (Fig. 2I–K). There was no impact on model variance during extinction (Fig. S2I).
Figure 2. DMS dynorphin release promotes goal-directed action.

(A) Top - Schematic of viral injection and optic fiber implantation in the DMS. Bottom - 20X Confocal image of optic fiber implant in the DMS with DAPI, expressing κLight 1.3a.
(B) Top - Operant behavior schedule during photometry. Bottom: Operant behavior (n=6 mice, Left, Learning Simple Linear Regression: R2=0.8382, p<0.0001. Right, Extinction Simple Linear Regression: R2=0.5800, p=0.0002).
(C-E) Mean fluorescence and heatmap raster plots during early, trained and extinction operant behavior (representative animal).
(F) Normalized Peak z-score values during cue+outcome (2–10s) (n=6mice; One Way ANOVA, p<0.0001****. Multiple comparisons – early vs. trained, p=0.0307*, trained vs. extinction, p=0.0006***, early vs. extinction, p=0.0094**).
(G) Schematic of Generalized Linear Model (GLM) analyses for trial-trial peak z-score values during cue+outcome and behavioral variables – trial number (trial, green), preceding inter-trial interval (ITI, light purple), subsequent inter-trial interval (ITI+1, dark purple) and latency to consume the reward pellet following a reinforced active nosepoke (eat, pink).
(H) Proportion of variance following behavioral variable drop-out in the trained sessions (n=6 mice, 199 trials; Full Model R2=0.4414, Trial R2=0.4755, p>0.05, ITI R2=0.4724, p>0.05, ITI+1 R2=0.0751, p=0.047*, Eat R2=0.4257, p>0.05).
(I) Mean Fluorescence for three consecutive trials with short ITIs (top) and two consecutive trials with long ITIs (bottom) time-locked to active nosepokes from a representative animal.
(J) Heat maps for trial-trial normalized ITIs (left) and trial-trial peak z-score values during cue+outcome (right) from a representative animal.
(K) Correlation between normalized ITI+1 and peak z-score during cue+outcome (Second Order Polynomial Quadratic Fit, R2=0.5408, p<0.0001****).
(L) Top - Schematic of viral injection and optic fiber implantation in the DMS. Bottom - 20X Confocal image of optic fiber implant in the DMS with DAPI, expressing ChR2.
(M) Operant behavior during photoactivation (n=8 Ctrl,8 dyn-cre mice; Two Way ANOVA, stim x genotype p<0.0009***. Multiple comparisons – ctrl off vs. on, p>0.05; dyn-cre off vs. on, p<0.0001****).
(N) Inter-trial Intervals for operant behavior during photoactivation (n=8 Ctrl,8 dyn-cre mice; Two Way ANOVA, stim x genotype p=0.0005***. Multiple comparisons – ctrl off vs. on, p>0.05; dyn-cre off vs. on, p=0.0005***).
(O) Extinction during days 1–5 (Left, Simple Linear Regression: n=4mice, Off - R2=0.6401, p=0.0018**, On – R2=0.001827, p>0.05. Difference in intercepts – p=0.0035##), and days 6–10 (Left, Simple Linear Regression: n=4mice, Off - R2=0.6401, p<0.0001****, On – Simple Linear Regression, R2=0.008760, p>0.05. Difference in Slopes – p=0.0356*).
(P) Operant behavior during photoactivation with aticaprant injection (n=8 dyn-cre mice; Two Way ANOVA, stim x treatment p<0.0059**. Multiple comparisons – veh off vs. on, p=0.0007***; aticaprant off vs. on, p>0.05).
Next, we determined the contribution of dynorphin release to sustain goal-directed behavior. We injected Pdyn-cre (DMSpdyn) or WT (Ctrl) mice with AAV5-EF1α-DIO-ChR2-YFP and implanted them with an optic fiber (Fig. 2L, S2J). Following training, mice received 20 Hz, 5 ms pulsewidth, 465 nm light stimulation during the anticipation window (7 seconds after an active nosepoke), to mimic the period of dynorphin release we previously observed. DMSpdyn mice showed a significant increase in their operant index compared to Ctrl upon photo-activation (Fig. 2M, S2K) and a significant reduction in their ITI compared to Ctrl (Fig. 2N). Importantly, sucrose-naïve DMSpdyn or WT mice did not display reinforcement behavior when nosepoking just for stimulation for 7 seconds (Fig. S2L-left), but did for 1 second stimulation (Fig. S2L-right). Next, we split the DMSpdyn mice into two groups for extinction – one that received photo-stimulation and the other that did not. We found that the group that received photo-stimulation maintained their operant responding (Fig. 2O-left, S2M). Upon switching the two groups, we observed that the group that previously received stimulation rapidly extinguished operant behavior, while the group that displayed prior extinction now regained their operant responding (Fig. 2O-right, S2M). Finally, we i.p injected DMSpdyn mice with aticaprant (5 mg/kg, 30 minutes prior to behavioral session) and found that this eliminated the increase in operant index upon stimulation (Fig. 2P). Interestingly, this effect was largely mediated by a reduction in reward consumption as animals still increased their active nosepokes upon stimulation (Fig. S2N). Collectively, these results suggest that DMSpdyn activity during the anticipation window promotes subsequent goal-directed behavior.
DMSpdyn neuron activity promotes goal-directed action.
Our results indicate that DMS dynorphin release influences goal-directed action in a trial-to-trial manner. To measure the activity of DMS neurons during goal-directed action-outcome behavior and their impact on trial engagement, we imaged the activity of over 10,000 neurons using 2-photon calcium imaging through implanted microprisms during head-fixed operant behavior. We injected D1R-CreXAi14 mice (D1R-TdTomato mice) with a virus packaging a genetically-encoded calcium sensor, AAVDJ-hsyn-GCaMP6s and implanted a microprism in the DMS (Fig. 3A, S3A), to co-register cells that were positive for both GCaMP and TdTomato. We determined that ~25% of the individually-tracked neurons were D1/pdyn-positive (DMSpdyn) (Fig. 3B). Following lick training for 10% sucrose and Pavlovian conditioning to associate a tone to sucrose delivery, mice were trained on a previously characterized45 self-paced operant task to rotate a wheel in an “active” direction to obtain the cue and the reward, with rotations in the “inactive” direction yielding nothing. Mice progressed to biasing their rotations towards the active direction across 8 days of learning (Fig. S3B) and increase the frequency of their overall operant responses (Fig. S3C,D). Further, when sucrose delivery was omitted, mice rapidly reduced their behavior (Fig. S3B). We adapted our existing operant index summary metric used during freely-moving behavior to capture the progression of action-outcome sequences across individual animal variance (Fig. 3C, See Methods).
Figure 3. DMSpdyn neuron activity promotes goal-directed action.

(A) Top - Schematic of viral injection and prism implantation in the DMS. Bottom – 20X and 40X confocal images of prism implant in the DMS with DAPI, expressing GCaMP6s and Td-Tomato.
(B) Top – Representative standard deviation images of GCaMP6s (left) and Td-tomato (right) fluorescence. Bottom – tracked and co-registered GCaMP6s+Td-tomato neurons and proportion of tracked GCaMP6s+ Td-tomato neurons across 3 mice.
(C) Top - Schematic of experiment and operant behavior schedule using a head-fixed wheel-based task. Bottom - Operant Behavior (n=3mice, Simple Linear Regression: Learning - R2=0.7388, p=0.003**, Extinction - R2=0.9837, p=0.0001****).
(D-F) Spectral clustering classification of early, trained and extinction operant data across 3 mice showing fluorescence activity traces (top) and heat maps (bottom) of cells.
(G) Peak z-score quantification of clusters - Action (549 cells; One Way ANOVA, p<0.0041**. Multiple comparisons – trained vs. early, p<0.0122*, trained vs. extinction, p<0.0030**), Action+Cue (487 cells; One Way ANOVA, p<0.0196*. Multiple comparisons – trained vs. early, p=0.0344*, trained vs. extinction, p=0.0169*), Cue+Reward clusters (332 cells; One Way ANOVA, p<0.0485*. Multiple comparisons – trained vs. early, p>0.05, trained vs. extinction, p=0.0344*), Reward (345 cells; One Way ANOVA, p<0.0302*. Multiple comparisons – trained vs. early, p>0.05, trained vs. extinction, p=0.0239*), and Inactive (967 cells; One Way ANOVA, p=0.0153*. Multiple comparisons – trained vs. early, p=0.01*, trained vs. extinction, p>0.05).
(H) Schematic of Artificial Recurrent Neural Network (RNN) modelling for decoding trial-trial DMSpdyn neuron activity to predict ITI length split into quintiles based on length using the activity of the preceding trial (trial N; ITI+1), subsequent trial (trial N+1; ITI), or both.
(I) Accuracy of RNN model for early (n=3 mice, 38 trials, 2680 cells; shuffled vs. real neural data, unpaired t test, p>0.05), trained (n=3 mice, 182 trials, 2680 cells; shuffled vs. real neural data, unpaired t test, p<0.0001****) and extinction (n=3 mice, 70 trials, 2680 cells; shuffled vs. real neural data, unpaired t test, p>0.05).
(J) Accuracy of RNN model for shortest 1/5th ITIs (Left, n=3 mice, 37 trials, 2680 cells; One Way ANOVA, p<0.0001****. Multiple comparisons – trained vs. early, p=0.0103*, trained vs. extinction, p=0.0002***) and longest 1/5th ITIs (Right, n=3 mice, 37 trials, 2680 cells; One Way ANOVA, p>0.05).
(K) Schematic of identifying neurons belonging to specific clusters based on the accuracy of decoding ITI length using the RNN model.
(L) Accuracy of RNN model after dropping out specific clusters for shortest one-fifth of ITIs length in trained operant (n=3 mice, 37 trials; Dashed Line – Full Model Accuracy, One Way ANOVA, p=0.0030**. Multiple comparisons – full vs. action, p>0.05, full vs. action+cue, p=0.0268*, full vs. cue+reward, p>0.05, full vs. reward, p>0.05, full vs. inactive, p>0.05).
(M) Schematic describing hypothesis: DMSpdyn neuron activity and dynorphin release evolves as animals learn goal-directed behavior, segregating their activity in response to distinct behavioral variables (left). Upon learning, DMSpdyn neuron activity and dynorphin release influences engagement in goal-directed action-outcome behavior (right).
Simultaneously, we imaged the activity of 3552 individually-tracked DMSpdyn neurons across multiple weeks (Fig. 3B). From trial-averaged data on day 8, we observed that DMSpdyn activity correlated to each of the three behavioral variables across a trial window (0–11s) – action (wheel rotation, 0–5s), anticipation (cue delivery, 5–8s) or reward (delivery and consumption, 8–11s). We found that a significantly higher proportion of DMSpdyn neurons were inhibited during the trial window encompassing action, cued anticipation and reward (0–11s) during the early (day 1) and extinction (day 9) sessions, compared to the trained (day 8) session (Fig. S3E). Among these neurons, a greater proportion were inactive/quiescent during the early and extinction sessions (Fig. S3F). Furthermore, these active neurons had a reduced trial-averaged amplitude compared to the trained session (Fig. S3G). Therefore, we used spectral clustering46,47 to group individual neurons based on their activity on the trained operant day (day 8) with the active cells (2680/3552 cells). Principle component analysis using each frame in the trial window as a principal component revealed that 4 principal components explained 90% of the variance (Fig. S3H); unbiased spectral clustering revealed five distinct clusters with the highest sillhouette score (Fig. S3I). The five clusters possessed activity during one or more defined behavioral variables – 20.5% during action (Action), 18.2% during the transition between action and cue (Action+Cue), 12.4% during cue and reward delivery (Cue+Reward), 12.9% during reward consumption (Reward), and 36% relatively inactive during these periods (Inactive) (Fig. 3E). Peak activity of neurons in each of these clusters was significantly higher in the window where activity was observed, compared to others (Fig. S3J–M). We then applied the same spectral clustering classification to the early and extinction days. We found that the same DMSpdyn neurons displayed significantly different and/or reduced patterns of activity early in operant learning (Fig. 3D,G). Furthermore, this activity diminished as animals decreased their behavior during extinction (Fig. 3F,G). Next, we tested whether DMSpdyn neuron activity in each trial influenced the inter-trial interval across learning like we observed with DMS dynorphin release. We trained an artificial recurrent neural network (RNN)48 to ascertain whether DMSpdyn activity could decode the length of the preceding (ITI) or subsequent (ITI+1) inter-trial interval (Fig. 3H). Training the RNN model on trial activity preceding (Fig. S3N) or succeeding (Fig. S3O) an ITI separately resulted in poor decoding accuracy on all sessions, near indistinguishable from shuffled neural data. However, training the model on activity data from both resulted in a significantly higher decoding accuracy relative to shuffled neural data on the trained session alone, but not in the early and extinction session (Fig. 3I). Notably, decoding accuracy improved significantly for the shortest ITI bin on the trained session but not for the longest (Fig. 3J), and not the early and extinction sessions (Fig. S3P), when binning it is in fifths based on length. Next, we asked whether specific clusters showed better decoding depending ITI length. Neurons belonging to all five clusters showed increased accuracy depending on the length of the inter-trial interval (Fig. S3Q); we observed a significant change in decoding accuracy for the shortest ITI bin using a cluster drop-out approach only for the action+cue cluster, compared to all others (Fig. 3L). We trained the same RNN model on trial activity during the last day of Pavlovian conditioning, where animals received the same tone followed by sucrose delivery across 60 trials at a variable, experimenter-imposed ITI. Here, we found poor decoding accuracies when training the model on ITI or ITI+1 alone, or both (Fig. S3R). These data suggest that upon operant learning, individual DMSpdyn neuron activity can influence the length of the interval and that DMSpdyn neuron activity and dynorphin release during anticipation of reward promote future re-engagement in action-outcome behavior (Fig. 3M).
BLAKOR neurons project to DMSpdyn neurons and are necessary for goal-directed behavior.
To determine the locus of action of KOR signaling in the DMS, we injected KORlox/lox mice with AAV2retro-Cre recombinase in the DMS (DMSretroKOR-cKO) (Fig. S4A–B). DMSretroKOR-cKO mice showed deficits in goal-directed behavior, despite heightened action vigor early on, compared to WT (Ctrl) mice (Fig. S4A). DMSretroKOR-cKO mice consumed the same number of rewards as Ctrls in their homecage and during Pavlovian conditioning (Fig. S4C,D). Since AAV2retro-Cre recombinase also removes KOR from neurons in the DMS, we injected KORlox/lox mice with AAV5-CMV-Cre recombinase in the DMS (DMSKOR-cKO) (Fig. S4E–F). DMSKOR-cKO mice did not show any significant changes compared to Ctrls during goal-directed behavior (Fig. S4E), in the number of rewards consumed in their homecage or Pavlovian conditioning (Fig. S4G,H).
To determine circuit-specific contributions of KOR signaling, we injected WT mice with AAV2retro-Cre recombinase and used fluorescent in situ hybridization in regions that project to the DMS (Fig. 4A). We found that ~20% of BLA neurons expressed Cre mRNA, overlapped entirely with CamKII mRNA, and over 60% of Cre+ neurons also expressed KOR mRNA (Fig. 4A-top). Additionally, we found that over 80% of Cre+ neurons also expressed Vglut1 mRNA (Fig. 4A-bottom), which is the predominant vesicular glutamate transporter found in the BLA49. Anterograde viral tracing in either KOR-Cre or Vglut1-Cre mice injected with AAV5-EF1α-DIO-ChR2-YFP showed YFP+ terminals in the DMS from either BLA Vglut1- (Fig. 4B-left) or KOR-expressing neurons (Fig. 4B-right).
Figure 4. BLAKOR neurons preferentially project to DMSpdyn neurons and are necessary for goal-directed behavior.

(A) Top - Schematic of viral injection in the DMS. Middle - 20X (left) and 40X (middle) Confocal image of BLA section with ISH for DAPI, CamKII, KOR and Cre, and (right) Quantification of ISH. Bottom - 20X (left) and 40X (middle) Confocal image of BLA section with ISH for DAPI, CamKII, Vglut1 and Cre, and (right) Quantification of ISH.
(B) Top - Schematic of viral injection in the BLA. 20X (middle) and 40X (bottom) confocal image of DMS sections stained for DAPI, with Vglut1+ (left) or KOR+ (right) BLA fibers expressing ChR2.
(C) Schematic of viral injection in the BLA.
(D) Input-output curve of optically-evoked EPSCs from DMSpdyn and D1(−) DMS neurons (n=5 mice, 16 cells; Two Way ANOVA, Cell-type p<0.0014**. Multiple comparisons – DMSpdyn vs. D1(−) at 2.7 mW, p=0.0240, at 5.7 mW p=0.0282*).
(E) Normalized optically-evoked EPSCs from DMSpdyn and D1(−) DMS neurons following U69 (n=4 mice, 8 cells; DMSpdyn - paired t test, p<0.0313*; D1(−) - paired t test, p<0.0002***).
(F) Normalized optically-evoked EPSCs from DMS neurons following U69 (n=5 mice, 17 cells, D1-TdTom; n=4 mice, 6 cells, KOR-cKO; D1-TdTom vs. KOR-cKO - unpaired t test, p<0.0235*).
(G) Representative traces for Veh and U69 in DMS (red) and KOR-cKO (brown).
(H) Left - Schematic of viral injection in the BLA and schematic of operant behavior. Right - 20X (left) and 40X (middle) Confocal image of Control (top) and BLAKOR-cKO (bottom) BLA section with ISH for DAPI, CamKII, KOR and overlap, and (right) Quantification of ISH.
(I) Left – Learning (Simple Linear Regression: n=6 Ctrl mice, R2=0.6587, p<0.0001****. n=8 BLAKOR-cKO mice, R2=0.2152, p=0.0224*. Difference in Slopes – p=0.0043**). Right – Extinction (Simple Linear Regression: n=6 Ctrl mice, R2=0.5057, p=0.0009***. n=8 BLAKOR-cKO mice, R2=0.0440, p>0.05. Difference in Slopes – p=0.001**).
To isolate if BLA terminals are functionally connected to the DMS and the impact of KOR signaling, we used ex vivo whole-cell patch clamp electrophysiology. We injected D1R-TdTomato or KORlox/lox mice with AAV5-CamKIIα-ChR2-YFP and AAV5-CMV-Cre recombinase in the BLA to record optically evoked EPSCs from either D1/dyn (DMSpdyn) or D1(−) neurons (Fig. 4C). We found that BLA terminal stimulation evoked oEPSCs in both populations, albeit with a significantly higher amplitude in DMSpdyn neurons (Fig. 4D). Furthermore, we bath applied the KOR agonist U69 to slices from D1R-TdTomato or KORlox/lox mice. We found that U69 inhibited oEPSCs equivalently from both neuronal populations (Fig. 4E). Importantly, U69 did not significantly dampen oEPSCs from neurons in KORlox/lox compared to D1R-TdTomato mice (Fig. 4F–G).
Next, we determined the impact of KOR deletion from the BLA on goal-directed behavior. We injected KORlox/lox mice with AAV5-CMV-Cre recombinase in the BLA (BLAKOR-cKO) (Fig. 4H) and observed that these mice consumed the same amount of sucrose in the homecage (Fig. S4I), but fewer rewards during Pavlovian conditioning (Fig. S4J). Moreover, BLAKOR-cKO mice were slower to learn, sustain and extinguished their operant responding compared to controls (Fig. 4I, S4K–L). These results indicate that glutamatergic BLA terminals in the DMS are negatively regulated by presynaptic KORs, and BLA KOR is necessary for goal-directed behavior.
BLAVglut1-DMS terminals are necessary and sufficient for promoting goal-directed behavior.
Our results suggest a homeostatic mechanism of regulation of glutamatergic BLA-DMS terminals during goal-directed behavior by dynorphin-KOR signaling. To test this hypothesis in vivo, we first measured the activity of Vglut1-expressing BLA-DMS terminals (BLAvglut1-DMS) using fiber photometry, by injecting Vglut1-cre mice with AAVDJ-EF1α-DIO-GCaMP6s and implanting an optic fiber in the DMS (Fig. 5A, S5A). As mice progressed through Pavlovian conditioning (Fig. S5B), we observed a significant inhibition of BLAvglut1-DMS fluorescence following cue+reward delivery (Fig. S5C–E). As animals learned operant conditioning (Fig. 5B), we observed a similar reduction in fluorescence during cue+reward (Fig. 5C,D,F), but strikingly, a significant ramping of GCaMP activity as the animals performed nosepokes (Fig. 5C,D,F). This activation during action and inhibition during cue+reward were eliminated during extinction (Fig. 5D,E,F). Next, we ascertained whether trial-to-trial changes in BLAvglut1-DMS activity influenced re-engagement in behavior, similar to DMS dynorphin release (Fig. 2). We performed simple linear regressions between the peak magnitude during action, or the peak minimum during cue+reward of BLAvglut1-DMS activity with the preceding (ITI) or subsequent ITI (ITI+1) (Fig. 5G–J). We observed a significant negative correlation between BLAvglut1-DMS activity during action with ITI (Fig. 5G), and between the reduction in activity during cue+reward with ITI+1 (Fig. 5J). We found no correlations of activity to ITIs in the early or extinction sessions (Fig. 5G–J).
Figure 5. BLA-DMS terminal dynamics are sufficient and necessary for goal-directed action.

(A) Top - Schematic of viral injection in the BLA and optic fiber implantation in the DMS. Bottom - 20X Confocal image of optic fiber implant in the DMS and BLA with DAPI, expressing GCaMP6s.
(B) Top - Operant behavior schedule during photometry. Bottom – Operant (n=6 mice; Simple Linear Regression: Learning, R2=0.7781, p<0.0001****. Extinction, R2=0.4251, p=0.0033**).
(C-E) Mean fluorescence and heatmap raster plots during early, trained and extinction operant behavior from a representative animal.
(F) Top – Normalized Peak z-score values during action (−20–0s, top) (n=6 mice; One Way ANOVA, p=0.0026**. Multiple comparisons – early vs. trained, p=0.0005***, trained vs. extinction, p=0.0092**, early vs. extinction, p>0.05), and cue+outcome (2–10s, bottom) (n=6 mice; p=0.0002*. Multiple comparisons – early vs. trained, p=0.0002***, trained vs. extinction, p=0.0027*, early vs. extinction, p>0.05).
(G) Simple Linear Regression between normalized ITI vs. peak z-score during action across early (n=6 mice, 78 trials, R2=0.002, p>0.05), trained (n=6 mice, 290 trials, R2=0.4733, p<0.0001****) and extinction (n=6mice, 526 trials, R2=0.0055, p>0.05).
(H) Simple Linear Regression between normalized ITI vs. peak z-score during cue+outcome across early (n=6 mice, 78 trials, R2=0.00525, p>0.05), trained (n=6mice, 290 trials, R2=0.0024, p>0.05) and extinction sessions (n=6mice, 526 trials, R2=0.0014, p>0.05).
(I) Simple Linear Regression between normalized ITI+1 vs. peak z-score during action across early (n=6 mice, 72 trials, R2=0.0209, p>0.05), trained (n=6 mice, 284 trials, R2=0.0204, p>0.05) and extinction (n=6mice, 520 trials, R2=0.0016, p>0.05).
(J) Simple Linear Regression between normalized ITI+1 vs. peak z-score during action across early (n=6mice, 72 trials, R2=0.007, p>0.05), trained (n=6mice, 284 trials, R2=0.3744, p<0.0001****) and extinction sessions (n=6mice, 520 trials, R2=0.0003, p>0.05).
(K) Left - Schematic of viral injection in the BLA and optic fiber implantation in the DMS. Right - 20X Confocal image of optic fiber implant in the DMS (top) and BLA (bottom) with DAPI, expressing ChR2.
(L) Active nosepokes for 1 second self-stimulation (n=8 vglut1-cre mice; paired t test, p<0.0001****).
(M) Operant behavior during photoactivation at action (n=4 Ctrl,8 vglut1-cre mice; Two Way ANOVA, stim x genotype p>0.05).
(N) Operant behavior during photoactivation at cue+outcome (n=4 Ctrl,8 Vglut1-cre mice; Two Way ANOVA, stim x genotype p=0.005**. Multiple comparisons – ctrl off vs. on, p>0.05; vglut1-cre off vs. on, p=0.0008***).
(O) Inter-trial Intervals for operant behavior during photoactivation at cue+outcome (n=4 Ctrl,8 Vglut1-cre mice; Two Way ANOVA, stim x genotype p=0.0041**. Multiple comparisons – ctrl off vs. on, p>0.05; vglut1-cre off vs. on, p<0.0001****).
(P) Left - Schematic of viral injection in the BLA and optic fiber implantation in the DMS. Right - 20X Confocal image of optic fiber implant in the DMS (top) and BLA (bottom) with DAPI, expressing PPO.
(Q) Operant behavior during photoinhibition at whole session (n=4 Ctrl,9 vglut1-cre mice; Two Way ANOVA, stim x genotype p=0.0079**. Multiple comparisons – ctrl off vs. on, p>0.05; vglut1-cre off vs. on, p<0.0001****).
(R) Operant behavior during photoinhibition at cue+outcome (n=4 Ctrl,9 vglut1-cre mice; Two Way ANOVA, stim p=0.0005***, stim x genotype p>0.05. Multiple comparisons – ctrl off vs. on, p>0.05; vglut1-cre off vs. on, p=0.0002***).
(S) Inter-trial Intervals for operant behavior during photoinhibition at cue+outcome (n=4 Ctrl,9 Vglut1-cre mice; Two Way ANOVA, stim x genotype p=0.005***. Multiple comparisons – ctrl off vs. on, p>0.05; vglut1-cre off vs. on, p<0.0001****).
(T) Extinction behavior during photoinhibition at whole session (Simple Linear Regression: n=5 vglut1-cre mice; Off - R2=0.6151, p=0.0005***. On – R2=0.2822, p=0.0416*. Difference in Slopes – p=0.0261*).
Next, we dissected differing contributions of BLAvglut1-DMS terminals to behavior via optogenetic manipulation. We injected Vglut1-Cre mice with AAV5-EF1α-DIO ChR2-YFP in the BLA and implanted optical fibers in the DMS to photo-activate BLAvglut1-DMS terminals (Fig. 5K, S5F). Sucrose-naïve mice nosepoked repeatedly just for photo-activation (20 Hz, 5ms pulse-width, 1mW laser power), suggesting that their activity is reinforcing (Fig. 5L). After operant training, mice either received light delivery triggered by the nosepoke, or during cue+reward delivery. Although animals receiving nosepoke-triggered stimulation made more active nosepokes, they consumed the same number of rewards (Fig. S5G), resulting in no difference in their operant index (Fig. 5M). Conversely, photo-activation during cue+reward resulted in a reduction in their operant index and an increase in their ITI (Fig. 5N–O, S5H). In parallel, we injected Vglut1-Cre mice with Cre-dependent parapinopsin (PPO)50 in the BLA and implanted optical fibers in the DMS to spatiotemporally mimic Gi-coupled GPCR inhibition of BLAvglut1-DMS terminals (Fig. 5P, S5L). Optical inhibition during the entire session resulted in a significant reduction in the operant index (Fig. 5Q, S5I). Furthermore, photo-inhibition during cue+reward delivery to mimic the KOR-mediated reduction in their activity (Fig 4) caused a significant increase in operant behavior, and a concomitant reduction in their ITIs (Fig. 5R–S, S5J). Next, mimicking terminal Gi/o-mediated inhibition significantly decreased extinction behavior in PPO animals compared to controls (Fig. 5T, S5K). Altogether, our results show that BLAvglut1-DMS terminal activity is engaged upon learning, necessary and sufficient for goal-directed behavior, and is under the control of Gi/o GPCR-mediated inhibition during outcome, thereby invigorating action.
BLA terminals DMSpdyn neuron activity, retrograde dynorphin release and dynorphin-KOR signaling to promote goal-directed action
Our results thus far suggest that DMS dynorphin release during reward anticipation leads to KOR signaling at BLA-DMS terminals to shape and sustain goal-directed behaviors. To determine if DMSpdyn activity is due to BLA terminal activation, we injected D1R-TdTomato mice we used in Fig.3 (expressing AAVDJ-hsyn-GCaMP6s and containing a microprism implant), with AAV9-CamKIIα-rsChRmine (rs-ChRmine) or AAV5-hsyn-mCherry (as a control) in the BLA to stimulate BLA terminals while imaging activity from the DMS (Fig. 6A). We performed sequential, spiral stimulation at discrete points across the field of view (spirals of 5–10 μm in diameter at 20Hz, 5ms pulse-widths for 500 ms, every 10 seconds at 10 mW of laser power, sequentially tiling the entire FOV), and in counterbalanced sessions, applied the same stimulation protocol with the laser switched off, resulting only in the shutter opening (Baseline sessions). We grouped neurons in the stim session that displayed activity higher than that during the baseline session as “responsive”, and the rest as “unresponsive” to BLA terminal stimulation. Using this protocol, we identified ~52.8% of DMSpdyn neurons that were responsive to BLA terminal stimulation compared to “unresponsive”, “baseline” and control (Fig. 6C,D, S6B). We then determined the relative proportions of these BLA-stimulated DMSpdyn neurons based on their cluster classification during operant behavior. While “responsive” DMSpdyn neurons were represented in each behavioral cluster, they were significantly enriched in the action+cue cluster (Fig. 6E, S6C–D), suggesting that DMSpdyn neurons active to action+cue are preferentially responsive to BLA input.
Figure 6. BLA terminals DMSpdyn neuron activity, retrograde dynorphin release and dynorphin-KOR signaling to promote goal-directed action.

(A) Top - Schematic of viral injection in the BLA and DMS, and microprism implantation in the DMS. Bottom - Schematic of experimental design to test whether BLA terminals are sufficient to activate DMSpdyn neurons (left) and traces from individual DMSpdyn neurons in the “Stim” and “Baseline” conditions (right).
(B) Mean fluorescence and heatmap raster plots of DMSpdyn neurons classified as “Responsive”, “Unresponsive”, or “Baseline” (representative animal, 1172 neurons).
(C) Mean fluorescence and heatmap raster plots of neurons from Control (representative animal, 440 neurons).
(D) Peak z-score quantification (Left: n=2 mice, 2 sessions each; One Way ANOVA, p<0.0001****. Multiple comparisons – responsive vs. unresponsive, p<0.0001****, responsive vs. baseline, p<0.0001****. Right: (n=2 mice, 2 sessions each; ChR vs. Ctrl, unpaired t test, p<0.0001****).
(E) Left – Schematic of identities of responsive DMSpdyn neurons in clusters. Right - Proportion of responsive DMSpdyn neurons (n=2 mice; One Way ANOVA, p<0.0044**. Multiple comparisons – Action+Cue vs. Action, p=0.0158*, Action+Cue vs. Cue, p=0.0368*, Action+Cue vs. Reward, p=0.0294*, Action+Cue vs. Inactive, p=0.0009***).
(F) Top - Schematic of viral injection in the BLA and DMS, and optic fiber implantation in the DMS.
(G) Peak z-score values following stim (2–10s), after baseline subtraction (−5 to −1s) (n=7 mice each; Ctrl vs. DMSpdyn-cKO, unpaired t test, p=0.0019**).
(H) Mean fluorescence and heatmap raster plots following BLA terminal stimulation in Ctrl (dark) and DMSpdyn-cKO (light; n=7 mice each).
(I) Top - Schematic of viral injection in the BLA and optic fiber implantation in the DMS. Bottom - Schematic of experimental design.
(J) Operant behavior (n=4 mice; vehicle vs. aticaprant, paired t test, p>0.05).
(K) Peak z-score values following stim period (2–7s), after baseline subtraction (−5 to −1s) (n=4 mice; vehicle vs. aticaprant, paired t test, p=0.0186*).
(L) Mean fluorescence and heatmap raster plots following injection during operant behavior of Vehicle (dark) and aticaprant (light; n=4mice).
(M) Top - Schematic of viral injection in the BLA and DMS, and optic fiber implantation in the DMS. Bottom - Schematic of experimental design.
(N) Operant behavior during photoactivation at cue+outcome (n=3 Ctrl, 4 Pdyn-cre mice; Two Way ANOVA, stim x genotype p=0.0068**. Multiple comparisons – ctrl off vs. on, p>0.05; Pdyn-cre off vs. on, p=0.0026**).
(O) Inter-trial Intervals for operant behavior during photoactivation at cue+outcome (n=3 Ctrl, 4 Pdyn-cre mice; Two Way ANOVA, stim x genotype p=0.0005***. Multiple comparisons – ctrl off vs. on, p>0.05; Pdyn-cre off vs. on, p=0.0002***).
(P) Mean fluorescence and heatmap raster plots during operant behavior with no stimulation (dark) and stimulation in Ctrl (light; n=3mice).
(Q) Normalized Peak z-score quantification of no stim vs. stim during action period (−5–0s) (n=3 mice; paired t test, p>0.05).
(R) Normalized Minimum z-score quantification of no stim vs. stim during cue+outcome period (2–10s) (n=3 mice; paired t test, p>0.05).
(S) Mean fluorescence and heatmap raster plots during operant behavior with no stimulation (dark) and stimulation in Pdyn-Cre (light; n=3 mice).
(T) Normalized Peak z-score quantification of no stim vs. stim during action period (−5–0s) (n=4 mice; paired t test, p=0.0354*).
(U) Normalized Minimum z-score quantification of no stim vs. stim during cue+outcome period (2–10s) (n=4 mice; paired t test, p=0.0059*).
We then determined if BLA terminal activity was sufficient for DMS dynorphin release by multiplexing BLA terminal stimulation with fiber photometry to measure κLight1.3a fluorescence in vivo (Fig. 6F). WT (Ctrl) or Pdynlox/lox (DMSpdyn-cKO) from Fig. 1 were also injected with AAV5-DIO-ChRimson in the BLA (Fig. 6F). Mice received 1s of 635 nm light stimulation (20 Hz, 5 ms pulse-width, 2mW laserpower) at a variable interval to stimulate BLA-DMS terminals. Ctrl mice showed a significant increase in κLight fluorescence, while significantly blunted in DMSpdyn-cKO mice (Fig. 6G,H). To determine if this is specific to BLAKOR terminals, we injected KOR-cre mice with AAV5-EF1α-DIO-ChRimson-TdTomato in the BLA and AAV5-EF1α-DIO-κLight1.3a in the DMS, and implanted optic fibers in the DMS (Fig. 6I, S6B). Since BLAvglut1-DMS terminal stimulation is reinforcing, mice nosepoked for 635 nm light stimulation (20Hz, 5ms pulse-width, 2mW laser power) in counterbalanced sessions, treated with vehicle or aticaprant (5 mg/kg, i.p) (Fig. 6I). Mice displayed robust nosepoking for BLAKOR-DMS terminal stimulation in both conditions (Fig. 6J). We observed a significant reduction in baseline subtracted (−5 to −1s of the window) κLight fluorescence following nosepokes during KOR antagonism (Fig. 6K, L).
Next, we ascertained whether retrograde dynorphin-KOR signaling at BLA-DMS terminals can potentiate KOR-mediated BLA terminal inhibition to shape goal-directed behavior. We multiplexed DMSpdyn stimulation with fiber photometry to measure BLA-DMS terminal fluorescence in vivo. We injected WT (Ctrl) or pdyn-cre mice with AAVDJ-CamKIIα-GCaMP6s in the BLA and AAV5-EF1α-DIO-ChRimson-TdTomato in the DMS, and implanted optic fibers in the DMS (Fig. 6M, S6B). We stimulated DMSpdyn neurons for 60s with 635 nm wavelength light (20Hz, 5ms pulse-width, 2mW laser power) and found a significant inhibition of BLACamKII-DMS terminal activity upon DMSpdyn stimulation, blunted by KOR antagonism via aticaprant i.p injection (Fig. S6F–H). Following conditioning, we triggered photo-activation during cue+reward to mimic dynorphin release and the reduction in BLA-DMS terminal activity (Fig. 1,2 and 5) in counter-balanced sessions (Fig. 6O). Pdyn-cre mice enhanced their operant behavior during stimulation resulting in a shorter ITI between action-outcome trials, while Ctrl mice did not (Fig. 6N,O, S6I). Whereas we saw no changes in BLACamKII-DMS upon stimulation in Ctrls (Fig. 6P–R), DMSpdyn stimulation enhanced the magnitude of BLACamKII-DMS activation during action (Fig. 6S,T) and the magnitude of inhibition during cue+reward (Fig. 6S,U). Altogether, our results suggest that BLA terminal activity stimulates DMSpdyn activity, dynorphin release and retrograde KOR signaling at BLA-DMS terminals to promote goal-directed behavior.
Dynorphin-KOR signaling at BLA-DMS terminals is necessary for acquiring goal-directed behavior
To determine whether dyn-KOR signaling impacts BLA-DMS terminal activity and goal-directed behavior, we injected Vglut1-Cre mice with either AAV1-CMV-DIO-SaCas9-U6-sgROSA (Vglut1-CresgROSA) or a previously validated51,52 AAV1-CMV-DIO-SaCas9-U6-sgOPRK1 (Vglut1-Cresgoprk1) and AAVDJ-EF1α-DIO-GCaMP6s and implanted an optic fiber in the DMS (Fig. 7A, S7A). We confirmed efficient knock-down of KOR mRNA in the BLA in Vglut1-Cresgoprk1 compared to Vglut1-CresgROSA (Fig. 7A), similar to previously published work51. Vglut1-Cresgoprk1 mice consumed the same amount of sucrose in the homecage (Fig. S7B), yet fewer rewards during Pavlovian conditioning compared in Vglut1-CresgROSA mice (Fig. S7C). Vglut1-Cresgoprk1 mice displayed significantly blunted acquisition of goal-directed behavior (Fig. 7B, S7D). and reduced activity during action (Fig. 7C,D), and cue+outcome (Fig. 7C,E), relative to a baseline period (−10 to −5s of the timewindow) compared to Vglut1-CresgROSA mice.
Figure 7. Dynorphin-KOR signaling at BLA-DMS terminals is necessary for acquiring goal-directed behavior.

(A) Top - Schematic of viral injection in the BLA and optic fiber implantation in the DMS. Bottom: 10X (left) and 40X (right) confocal images of Vglut1-CresgROSA (top) and Vglut1-Creoprk1 (bottom) BLA section with ISH for DAPI (blue) and Vglut1 (green), and DAPI (blue) and KOR (majenta), and quantification of ISH.
(B) Operant learning (Simple Linear Regression: n=4 Vglut1-CresgROSA mice, R2=0.9173, p<0.0001****. n=6 Vglut1-Creoprk1 mice, R2=0.6745, p<0.0001****. Difference in Slopes – p<0.0001****)
(C) Mean fluorescence and heatmap raster plots during operant behavior (n=4 Vglut1-CresgROSA, 6 Vglut1-Creoprk1 mice).
(D) Peak z-score values during action (−5–0s), subtracted from baseline (−10 to −5s) (n=4 Vglut1-CresgROSA, 6 Vglut1-Creoprk1 mice; unpaired t test, p=0.0042**).
(E) Minimum z-score values during cue+outcome (2–10s), subtracted from baseline (−10 to −5s) (n=4 Vglut1-CresgROSA, 6 Vglut1-Creoprk1 mice; unpaired t test, p=0.0483*).
(F) Left - Schematic of viral injection in the BLA and DMS, and optic fiber implantation in the DMS. Right - 20X Confocal image of optic fiber implant in the DMS (top) and BLA (bottom) with DAPI, expressing GCaMP6s.
(G) Operant learning (Simple Linear Regression: n=4 Ctrl mice, R2=0.8551, p<0.0001****. n=7 DMSpdyn-cKO mice, R2=0.3656, p=0.0037**. Difference in Slopes – p<0.0001****).
(H) Mean fluorescence and heatmap raster plots during operant behavior (n=4 Ctrl mice, n=7 DMSpdyn-cKO mice).
(I) Peak z-score values during action (−5–0s), subtracted from baseline (−10 to −5s) (n=4 Ctrl, 7 DMSpdyn-cKO mice; unpaired t test, p=0.0001***).
(J) Minimum z-score values during cue+outcome (2–10s), subtracted from baseline (−10 to −5s) (n=4 Ctrl, 7 DMSpdyn-cKO mice; unpaired t test, p=0.0497*).
(K) Schematic of viral injection in the BLA and DMS, and optic fiber implantation in the DMS.
(L) 10X (left) and 40X (right) confocal images of DMSTScKO Striatum section with ISH for DAPI (blue) and DRD1 (green), DAPI (blue) and Pdyn (majenta), and DAPI (blue) and Cre (gold). Bottom - Quantification of ISH.
(M) Operant learning (Simple Linear Regression: n=4 Ctrl mice, R2=0.4692, p=0.014*. n=6 DMSTScKO mice, R2=0.6239, p<0.0001****. Difference in Slopes – p=0.0204*).
Next, we injected WT (Ctrl) or pdynlox/lox mice with AAV5-CMV-Cre recombinase in the DMS bilaterally to delete DMSpdyn and AAVDJ-CaMKIIα-GCaMP6s in the BLA, and implanted an optic fiber in the DMS unilaterally (BLACaMKII-DMSpdyn-cKO) (Fig. 7F, S7E). Mice lacking DMSpdyn consumed the same amount of sucrose under food-restriction in their homecage but fewer pellets during Pavlovian conditioning (Fig. S7F, S7G). We observed similar prior deficits (Fig. 1) in goal-directed behavior in BLACaMKII-DMSpdyn-cKO mice relative to controls (Fig. 7G, S7H). Strikingly, we found a concomitant drastic reduction in the magnitude of BLACaMKII-DMSpdyn-cKO fluorescence to action and cue+reward, relative to a baseline period (−10 to −5s of the timewindow) (Fig. 7H–J).
Finally, to determine whether this impact is BLA-DMS specific, we injected WT (Ctrl) or pdynlox/lox mice with a transynaptic AAV1-FDIO-Cre recombinase in the BLA bilaterally, followed by AAV1-FLP in the DMS bilaterally to delete dynorphin selectively from DMSpdyn neurons receiving BLA input (BLA-DMSTScKO) (Fig. 7K, S7I). Our ISH results confirmed significant DMSpdyn knockdown, including efficient transynaptic labeling of Cre in the DMS (Fig. 7L). BLA-DMSTscKO mice consumed the same amount of sucrose in the homecage, yet significantly lesser during Pavlovian conditioning (Fig. S7J, S7K). BLA-DMSTscKO mice showed significantly attenuated goal-directed behavior during operant conditioning compared to Ctrl mice (Fig. 7M, S7L). Collectively, these results lead us to conclude that modulation of BLA-DMS activity by dynorphin-KOR signaling consequently shapes goal-directed behavior.
Discussion
It is well-established that neuropeptide-GPCR signaling can shape neural activity to have a sustained impact on behavior. However, the timescales of their action, a mechanistic understanding of how they control neural activity, and the behaviors they subsequently modulate is unclear. In this manuscript, we uncover an endogenous neuropeptidergic GPCR-mediated homeostatic mechanism for the control of neural activity in the DMS to sculpt goal-directed behavior across multiple timescales.
The DMS is essential for goal-directed behavior in rodents, non-human primates and humans53,54. Despite knowing that a large percentage of D1 MSNs express dynorphin in the DMS20, dynorphin has largely been used as a marker for these cells. Only recently have we been able to detect dynorphin levels55, allowing for the possibility of definitive attributions to the role of dynorphin-KOR signaling to behavior. Here, along with conditional deletions of dynorphin production and a novel dynorphin biosensor κLight1.3a56–58, our results suggest that an ongoing recruitment of dynorphin shapes goal-directed learning (Fig. 1). Our studies draw similar parallels to that observed with drug self-administration59, where prior work showed slow and sustained elevations in pdyn mRNA levels following drug self-administration behavior60–62, and consequentially, activating63,64 or antagonizing65,66 KOR impacts drug-seeking and reinstatement. Our studies also revealed a contribution of dynorphin release at strikingly fast timescales, promoting subsequent engagement in action-outcome sequences (Fig. 2). Such roles for a signaling molecule to predict future engagement in a task have been attributed to dopamine67, but as yet unreported for a neuropeptide. Furthermore, using implantable microprisms in the DMS68, allowing us to image and track thousands of neurons per animal over weeks, we demonstrated that DMSpdyn neuron activity is recruited for segregation into clusters defined by the distinct behavioral variables involved, corroborating prior reports for action and outcome encoding in striatal MSNs in non-human primates7,69,70 and rodents14–16,71–73. Importantly, we also showed that DMSpdyn neuron activity can predict future engagement in action-outcome sequences, underscoring concordance to our observations with DMS dynorphin release (Fig. 3). These experiments also corroborate prior studies on roles for striatal neurons in decoding an animal’s task engagement15, or the initiation of action-outcome sequences74.
It is important to note that recent studies have provided results that present alternative findings to those in this manuscript. One study showed that exposure to norBNI systemically, a long-lasting KOR antagonist, results in enhanced operant learning75. Another study reported that animals lacking pdyn in all D1R-expressing neurons from birth resulted in increased flexibility of operant behavior76. In our work, dynorphin expression was manipulated only in the DMS during adulthood. Several brain regions express pdyn and KOR, and multiple other brain regions in addition to the dorsal striatum show dyn/D1R overlap77–80. Furthermore, whereas infusion of KOR agonists into multiple brain regions produced aversion, infusion into the dorsal striatum did not81, suggesting a putative different role for dyn-KOR control in the dorsal striatum. Ultimately, these reported findings do not conflict with the results in this study and serve to further implicate endogenous opioid signaling in the modulation of adaptive behavior, which only a few studies have done80,82.
Prior evidence in ex-vivo preparations suggests that opioids may function to dampen neuronal circuit activity in a retrograde manner35,64,83. Among the regions that project to the DMS, the BLA is enriched in KOR expression36,84, projects extensively across the striatum85 including the dorsal striatum41,42,86,87 and has been implicated as a hub for motivated behaviors40,88. We found a significant population of KOR+ DMS-projecting BLA neurons that preferentially project to D1/dyn neurons in the DMS and are under dyn-KOR control. Furthermore, KOR deletion from the BLA resulted in a profound reduction in goal-directed behavior (Fig. 4). Remarkably, we found a bimodal pattern of activity at BLA terminals and once again observed specific modulation for engagement in behavior, wherein increased activity during action was influenced by engagement in the current trial, while reductions during outcome promoted future action-outcome sequences (Fig. 5). Other studies observing activity in this projection have only been conducted at the level of the cell bodies in the BLA41,42. While both studies report an engagement in activity, they did not observe reductions, suggesting that BLA terminals may be under differential regulation, independent of BLA soma activity.
Finally, we observed that BLA terminal regulation via retrograde dynorphin-KOR signaling is sufficient (Fig. 6) and necessary (Fig. 7) for goal-directed behavior. Whereas retrograde release of dynorphin has been posited as a mechanism of action to modulate dopamine release in the striatum23,52,64, this study explores its impact on circuits important for reward processing. Our study demonstrates the synergy between activity, neuropeptide release and retrograde signaling resulting in the refinement of behavior. Herein, (i) BLA terminal excitation during action promotes the activity of cue-selective DMSpdyn neurons and subsequent release of DMSdyn during cued anticipation, (ii) thereby inhibiting BLA terminals via retrograde dynorphin-KOR inhibition, and (iii) affording the stability of DMSpdyn neuron selectivity to encode distinct variables to sustain engagement in goal-directed behavior. More specifically, our study raises key questions for additional follow up experiments - What is the functional role for dynorphin-KOR signaling in goal-directed behavior? Why is dynorphin released during cued anticipation and why does it promote future engagement in goal-directed action? As we observed an increase in dynorphin across learning (Fig. 1 and 2), our results suggest that dynorphin could be signaling the salience of the context animals find themselves in to obtain rewards. Alternatively, dynorphin may contribute to determining outcome value in the context of action-outcome learning. While our studies explore local DMS dynorphin release, this does not preclude the contributions of extra-striatal dynorphin, which warrants further study. Finally, the role of local neuropeptide-GPCR regulation of excitatory terminals as a generalizable mechanism to promote neuronal activity and behavior warrants further study.
In summary, we ascribe a dynamic role for the opioid neuropeptide dynorphin via a retrograde Gi-GPCR inhibition mechanism whereby dynorphin-KOR controls the activity of BLA projections to the DMS, ultimately promoting action-outcome behavior. These efforts illuminate a window for dynorphin-KOR action during these behaviors and provide necessary insight to guide the much-anticipated therapeutic development for the treatment of neuropsychiatric disorders which impact related motivated behaviors.
Limitations of the study
A notable limitation of our study is the use of fluorescence intensity-dependent measurements via genetically-encoded sensors for dynorphin (κLight1.3a) and calcium (GCaMP6s). As these measurements are relative, they do not provide direct measurements for dynorphin or changes in extracellular calcium unequivocally across slow timescales. Further, these limitations make comparisons across animals more challenging. Hence, we draw conclusions from our results regarding the timescale of effects on molecule, circuit and behavior based on the nature of the manipulation and the inclusion of appropriate controls, while recognizing the need for better tools in this domain. Indeed, the advent of fluorescence lifetime-compatible genetically-encoded biosensors89,90 and better biosensor-independent tools55,91 would help us resolve these changes more accurately. Additionally, studies observing reductions in in GCaMP6s from axon terminals have been less prominent, but have been reported at excitatory terminals into the nucleus accumbens92,93. While these reductions capture possible dampening of neuronal activity by Gαi GPCR-coupled mechanisms, via the inhibition of calcium channels, they limit our understanding of other Gαi GPCR-coupled mechanisms of inhibition, namely the reduction of vesicular glutamate release (fast), and the reduction in cAMP production via kinase inhibition and subsequent decreases in intracellular calcium (slow). Indeed, studies are already underway with glutamate sensors94, kinase sensors95,96 and cAMP sensors97 to ascertain the impact of GPCR signaling on neuronal activity and behavior at both slow and fast timescales. Furthermore, our approach to co-register td-Tomato positive cells in the field of view during two-photon imaging. yielded a percentage of DMSpdyn neurons lower than that reported previously98, and us (Fig. 1B), likely attributable to a combination of the efficiency of genetically-encoded fluorophore expression and limited detection using two-photon imaging with td-Tomato’s efficiency. Finally, even though our results support a role for dynorphin-KOR signaling at fast timescales to promote goal-directed action-outcome sequences, it does not preclude the possibility of these effects arising only across slow timescales or in other circuits relevant to the striatum. Therefore, future studies warrant isolating the role of dynorphin-KOR signaling, specifically at rapid timescales via time-locked, signaling-specific manipulations such as opto-opioidGPCRs99 or time-locked excision of opioid precursors100. Furthermore, future studies also warrant the impact of long-term plasticity mechanisms that goal-directed behavioral learning may induce at BLA-DMS synapses, and other opioid-sensitive striatal synapses76,101.
Resource Availability
Lead Contact
Further information and requests should be directed to the lead contact, Dr. Michael R. Bruchas (mbruchas@uw.edu)
Materials Availability
No new materials were generated in this paper.
Data and Code Availability
All data reported in this paper will be shared by the lead contact upon request.
All original code has been deposited at Zenodo at https://zenodo.org/records/21286480
and is publicly available as of the date of publication.
Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.
STAR Methods Text
Experimental Models and Study Participant Details
Animals
Adult (18–35 g) male and female wildtype (WT), Pdynlox/lox (Pdyn cKO), Oprk1-Cre (KOR-cre), Pdyn-IRES-Cre (Pdyn-Cre), Ai14 x DrD1-Cre, and Oprk1lox/lox (KOR cKO) mice were group housed, given access to food pellets and water ad libitum, and maintained on a 12 hr:12 hr light:dark cycle (lights off at 9:00 AM, lights on at 9:00 PM). All mice were kept in a sound-attenuated, isolated holding facility one week prior to surgery, post-surgery, and throughout the duration of the behavioral assays to minimize stress. For cell-type conditional deletion and optogenetic experiments we used age-matched Cre- cage, littermate and WT controls. Unless otherwise noted, animals had ad libitum access to food and water. Any variation from these approaches was due to behavioral attrition from off-target injections/implants or headcap failures. All animals were drug and test naive, individually assigned to specific experiments as described, and not involved with other experimental procedures. Statistical comparisons did not detect any significant differences between male and female mice and were therefore combined to complete final group sizes. Statistical and behavioral comparisons did not detect any changes across strains in operant behavior, and all animals involving comparisons began food/water-restriction at the same time and experimented on in parallel at the same time. All animals were monitored for health status daily and before experimentation for the entirety of the study. All procedures were approved by the Animal Care and Use Committee of the University of Washington and conformed to US National Institutes of Health guidelines.
Method Details
Stereotaxic Surgery
All coordinates, viruses and implant type for experiments are listed in Table S1. After mice were acclimated to the holding facility for at least seven days, the mice were anaesthetized in an induction chamber (1%−4% isoflurane) and placed into a stereotaxic frame (Kopf Instruments, model 1900) where they were mainlined at 1%−2% isoflurane. For mice receiving viral injections followed by microprism implants, we used a Nanoject II (Drummond Scientific) to inject 4 × 300 nL of virus at a rate of 100 nL/min. For all other viral injections, a blunt needle (86200, Hamilton Company) syringe was used to deliver 400 nL of virus at a rate of 100 nL/min either in the DMS or BLA. For mice receiving intracranial implants (i.e., microprism implants, fiber photometry or optogenetic optic fibers), a hole was drilled above the site of interest, and the implant was slowly lowered to the coordinates. Cannulas were secured to the skull using one bone screw and super glue (Lang Dental). All other implants were secured using MetaBond (C & B Metabond). For mice undergoing head-fixed experiments, a head-ring was placed over the implant before securing them and the implants with MetaBond. For further detail on 1.5 × 1.5 × 8mm microprism (OptoSigma) implantation, please refer to47.
Freely moving operant behavior
For freely moving experiments, animals were food-restricted 6–8 weeks following surgery to maintain their body weight at ~85%. If undergoing self-stimulation experiments, animals were tethered and placed in an operant chamber (MedPC). Animals nosepoked into illuminated nosepoke ports, randomly designated “active” or “inactive”, with the active poke resulting in a 1s or 7s delivery of 465 nm (1–5mW laser power) or 635 nm (1–2mW laser power) laser light at 20Hz, 5 ms pulse-width. For experiments involving sucrose, animals were tethered (if tethering was required), placed in an operant chamber (MedPC) and underwent magazine training for random delivery of sucrose pellets (BioServ) in a receptacle for 3 days. Following this, animals underwent Pavlovian conditioning to associate a 5s houselight succeeded by sucrose pellet delivery for 5 days. Mice then underwent operant conditioning, where they nosepoked into illuminated nosepoke ports, randomly designated “active” or “inactive”, with the active poke resulting in a 2s timeout, then 5s houselight, followed by sucrose delivery for 5 days. When specified, animals then underwent operant reversal, where the nosepoke ports were reversed, fixed ratio-3, where they performed 3 nosepokes instead of 1 for cue and reward delivery of the same time, and progressive ratio testing where the nosepokes required escalated exponentially as we previously described107. If undergoing PR, animals were returned to fixed ratio-1 for 3 days until stable responding was achieved and then underwent operant extinction for 5 days. Here, an active nosepoke yielded cue delivery, but no sucrose delivery. Intraperitoneal injections during behavior, when indicated, were performed 30 minutes prior to the start of the session.
Operant Index for Freely-moving Behavior
To accurately capture all components of action-outcome behavior, we devised an operant index considering action discrimination (Active – Inactive nosepokes), action vigor (total nosepokes) and sucrose pellets consumed. This was normalized to a hypothetical index under ideal conditions involving the active nosepokes required to achieve the maximum outcomes in Pavlovian conditioning (40 trials), half the number of inactive nosepokes to achieve discrimination and the maximum outcomes consumed in Pavlovian conditioning (40 trials). For extinction, we used the number of approaches the animal made to the sucrose hopper instead of outcomes consumed. For transparency, all active and inactive nosepokes, rewards consumed and approaches to hopper are quantified in the supplemental figures. Unless mentioned otherwise, “early” is day 1, “trained” is day 5, and “extinction” is day 1 of extinction.
In Vivo Fiber Photometry
Fiber photometry recordings were made throughout the entirety of 60-minute Operant testing. Prior to recording, an optic fiber was attached to the implanted fiber using a ferrule sleeve (Doric, ZR_2.5). For κLight1.3a biosensor recordings, two LEDs were used to excite κLight1.3a. A 531-Hz sinusoidal LED light (Thorlabs, LED light: M490F3; LED driver: DC4104) was bandpass filtered (490 ± 20 nm, Doric, FMC6) to excite κLight1.3a and evoke κLight1.3a-dependent emission. A 211-Hz sinusoidal LED light (Thorlabs, C_LED:435; LED driver: DC4104) was bandpass filtered (435 ± 10 nm, Doric, FMC6) to excite κLight1.3a and evoke κLight1.3a-independent isosbestic control emission. For GcaMP6s recordings, two LEDs were used to excite GCaMP6s. A 531-Hz sinusoidal LED light (Thorlabs, LED light: M470F3; LED driver: DC4104) was bandpass filtered (470 ± 20 nm, Doric, FMC4) to excite GCaMP6s and evoke Ca2+-dependent emission. A 211-Hz sinusoidal LED light (Thorlabs, LED light: M405FP1; LED driver: DC4104) was bandpass filtered (405 ± 10 nm, Doric, FMC4) to excite GCaMP6s and evoke Ca2+-independent isosbestic control emission. Prior to recording, a 120 s period of excitation, where the majority of baseline drift is observed with 435/405 nm and 490/470 nm light, was used to allow the signal to normalize. Laser intensity for the 490/470 nm and 435/405 nm wavelength bands were measured at the tip of the optic fiber and adjusted to ~30 μW before each day of recording. κLight1.3a or GCaMP6s fluorescence traveled through the same optic fiber before being bandpass filtered (525 ± 25 nm, Doric, FMC6 or FMC4), transduced by a femtowatt silicon photoreceiver (Newport, 2151) and recorded by a real-time processor (TDT, RZ10). The envelopes of the 531-Hz and 211-Hz signals were extracted in real-time by the TDT program Synapse at a sampling rate of 1017.25 Hz. For the ChrimsonR stimulation experiments, a 635 nm laser was used with a custom filter cube (Doric) at 1–2 mW intensity to deliver red light through the tip of the same optic fiber used to excite GCaMP6s similar to108.
Head-fixed Operant Behavior
Experiments were performed as previously described48. In brief, animals were signal checked for GcaMP6s dynamics 4–6 weeks after surgery by securing them in the OHRBETS platform. Animals were water-restricted for 1 week and maintained at ~85% body weight. Animals then underwent sipper training for 10% sucrose for 3 days, followed by Pavlovian conditioning for 3 days to associate a 3s tone followed by 3s access to sucrose via sipper extension, then operant conditioning to rotate a wheel for 8 days. Rotating the wheel a half-turn in the “active” contingency yielded the tone and sucrose availability, the “inactive” contingency yielded nothing. Tone delivery also resulted in a brake applied to the wheel. This was followed by operant extinction where active rotations resulted in the tone and sipper extension, but no sucrose for 2 days. Direction of wheel rotation was counterbalanced to implant hemisphere.
Operant Index for Head-Fixed Behavior
As above with freely moving operant behavior, to accurately capture all components of an action-outcome behavior, we devised a summary metric called the operant index taking into account action discrimination (Active – Inactive wheel rotations), action vigor (total wheel rotations) and outcomes consumed (every trial the animal consumed sucrose). We then normalized this value to a hypothetical index under ideal conditions involving the active rotation contingencies required to achieve the maximum outcomes in Pavlovian conditioning (40 trials), half the number of inactive rotation contingencies to achieve discrimination and the maximum outcomes consumed in Pavlovian conditioning (40 trials). For extinction, we used the number of trials the animals performed a lick at the sipper instead of outcomes consumed. For transparency, active-inactive wheel rotation data is reported in the supplemental figures. Unless mentioned otherwise, “early” is day 1, “trained” is day 8, and “extinction” is day 1 of extinction.
Two-Photon Calcium Imaging
For details on how imaging and longitudinal tracking was performed, please refer to Hjort & Gowrishankar et al47. Animals used in Hjort & Gowrishankar et al.,47 for sucrose consumption were used for Pavlovian and operant conditioning here. Prior to each behavior session, animals were placed in the OHRBETS platform outfitted with a Thorlabs goniometer (TTR001/M, Thorlabs) to facilitate levelling of the imaging plane. Imaging was conducted on a Bruker 2p+ (Bruker) at 920 nm using the Cousa objective (20mm working distance; Pacifica Optics). Following identification of the same field of view (FOV) to enable longitudinal tracking, animals underwent 15 minute sessions of lick training, Pavlovian conditioning, Operant conditioning or Operant extinction where imaging was acquired at 7.5 Hz using resonant galvos (4 frame averaging). At the end of each session, a 5 minute static recording was acquired at at 7.5 Hz using resonant galvos (4 frame averaging) at 1080 nm wavelength in the same FOV to register td-Tomato postitive cells. Imaged cells were tracked longitudinally by concatenating recording sessions and running them through Suite2p (HHMI, Janelia Research Campus). Sorted cells were manually verified as well. GcaMP6s and Td-Tomato cells were overlaid in ImageJ/Fiji and quantified for overlap. Spectral clustering, peak analysis, RNN and GLM analyses were performed using custom Python code based on49. All data was was z-scored using the mean fluorescence and standard deviation of a predefined time window (1s prior active wheel rotation). Data are presented as z-score across −1s to 11s, with 0s signifying the start of wheel rotation in the trial.
For targeted sequential photostimulation experiments of BLA axons in conjunction with DMS two-photon imaging, a second laser path with a 1040nm high powered femtosecond laser (Spirit One, SpectraPhysics) was used with a pair of galvanometers to generate spiral montages across the FOV, similar to109,110. Spirals of 5–10 μm in diameter were generated at 20Hz, 5ms pulse-widths for 500 ms, every 10 seconds at 10 mW of laser power, sequentially tiling the entire FOV. In parallel, animals also underwent “baseline” sessions resulting in shutter opening for the same frequency, but with the laser off. Data was normalized to a ~2 minute baseline prior to stimulation. All the cells in the entire FOV were analyzed following stimulation (or shutter opening, for “baseline” sessions. No distance threshold was applied for the analysis as we did not previously determine the extent of arborization of BLA terminals in the DMS. “Responsive” cells were determined based on whether their activity was higher than the peak activity in the “baseline” sessions; cells that had activity equal to or below the “baseline” session were termed “unresponsive”. These sessions were then concatenated with the recordings from the trained operant session and analyzed in Suite2p for tracking to determine the identities of the neurons enriched in behavioral clusters. Point spread functions to assess the utility for microprims for spatial stimulation were assessed in47. All data was was z-scored using the mean fluorescence and standard deviation of a predefined time window (1s prior to laser stimulation). Data are presented as z-score across −5s to 5s, with 0s signifying the start of laser stimulation.
Patch-Clamp Electrophysiology
Coronal brain slices were prepared at 250 μM on a vibrating Leica VT1000S microtome using standard procedures. Mice were anesthetized with Isoflurane, and transcardially perfused with ice-cold and oxygenated cutting solution consisting of (in mM): 93 N-Methyl-D-glucamine (NMDG), 2.5 KCL, 20 HEPES, 30 NaHCO3, NaH2PO4, 10 MgSO4⋅7H20, 0.5 CaCl2⋅2H20, 25 glucose, 3 Na+-pyruvate, 5 Na+-ascorbate, and 5 N-acetylcysteine. Following collection of coronal sections, the brain slices were transferred to a 34°C chamber containing oxygenated cutting solution for a 10-minute recovery period. Slices were then transferred to a holding chamber consisting of (in mM) 92 NaCl, 2.5 KCl, 20 HEPES, 2 MgSO4⋅7H20, 1.2 NaH2PO4, 30NaHCO3, 2 CaCl2⋅2H20, 25 glucose, 3 Na-pyruvate, 5 Na-ascorbate, 5 N-acetylcysteine and were allowed to recover for ≥ 30 min. For recording, slices were perfused with oxygenated artificial cerebrospinal fluid (ACSF; 31–33°C; 300–303 milliosmols) consisting of (in mM): 113 NaCl, 2.5 KCl, 1.2 MgSO4⋅7H20, 2.5 CaCl2⋅6H20, 1 NaH2PO4, 26 NaHCO3, 20 glucose, 3 Na+-pyruvate, 1 Na+-ascorbate, at a flow rate of 2–3ml/min and visualized using differential interference contrast through a 40x water-immersion objective mounted on an upright microscope (Olympus BX51WI). DMS neurons were initially voltage clamped in whole-cell configuration using borosilicate glass pipettes (2–4MΩ). For recordings of excitatory currents, pipettes were filled with internal solution containing (in mM): 125 K+-gluconate, 4 NaCl, 10 HEPES, 4 MgATP, 0.3 Na-GTP, and 10 Na-phosphocreatine (pH 7.30–7.35). The patch pipette also included 50 μM picrotoxin to block GABAA currents. Following break-in to the cell, we waited ≥ 3 minutes to allow for exchange of internal solution and stabilization of membrane properties. Neurons with an access resistance of > 30MΩ or that exhibited greater than a 20% change in access resistance during the recording were not included in our datasets. For all voltage clamp experiments, neurons were held at −70mV. To assess connectivity between the BLA and the DMS, voltage clamp recordings were performed from cells located near eYFP-expressing axons withing the DMS. For optogenetic recordings of input/output curves, we used a Thorlabs LEDD1B T-Cube driver and obtained separate recordings of 470nm wavelength oEPSCs / oIPSCs at 7 output levels corresponding to 5.7, 2.7, 1.6, 1, 0.5, 0.2, and 0.1 mW of LED intensity with a constant pulse width of 5ms. U69 (1uM) washes were conducted after collecting a stable evoked oEPSC baseline for 7 minutes. Data acquisition occurred at 10 kHz sampling rate through a MultiClamp 700B amplifier connected to a Digidata 1440A digitizer (Molecular Devices). Data were processed using Clampfit v11.0.3.03 (Molecular Devices) and analyzed using GraphPad Prism v8.3.0. All tests were two-sided and corrected for multiple comparisons or unequal variance where appropriate.
To verify that we were not recording from fast spiking parvalbumin interneurons, we obtained action potential recordings to segregate MSNs from fast spiking interneurons based on AP frequency and AP width111,112,113. To ensure we were not recording from ChAT interneurons, we recorded Sag currents, which are mediated through HCN channels which are minimally expressed on MSNs114,115. Furthermore, we obtained recordings of afterhyperpolarization currents, as ChAT neurons exhibited a characteristic pattern of prolonged decay kinetics compared to MSNs116. These metrics gave us reasonable confidence that the D1- neurons that we recorded from were indeed MSNs, and ostensibly D2+. Data acquisition occurred at 10 kHz sampling rate through a MultiClamp 700B amplifier connected to a Digidata 1440A digitizer (Molecular Devices). Data were processed using Clampfit v11.0.3.03 (Molecular Devices) and analyzed using GraphPad Prism v8.3.0. All tests were two-sided and corrected for multiple comparisons or unequal variance where appropriate.
Tissue processing
Unless otherwise stated, animals were transcardially perfused with 0.1 M phosphate-buffered saline (PBS) and then 40 mL 4% paraformaldehyde (PFA). Brains were dissected and post-fixed in 4% PFA overnight and then transferred to 30% sucrose solution for cryoprotection. Brains were sectioned at 40 μM on a microtome and stored in a 0.01M phosphate buffer at 4°C prior to immunohistochemistry and tracing experiments. For behavioral cohorts, viral expression and optical fiber placements were confirmed before inclusion in the presented datasets.
RNAscope Fluorescent In Situ Hybridization
Following rapid decapitation of WT, Pdynfl/fl or Oprk1fl/fl mice brains were rapidly frozen in 100mL −50°C isopentane and stored at −80°C. Coronal sections corresponding to the site of interest or injection plane used in the behavioral experiments were cut at 20uM at −20°C and thaw-mounted onto SuperFrost Plus slides (Fisher). Slides were stored at −80°C until further processing. Fluorescent in situ hybridization was performed according to the RNAscope 2.0 Fluorescent Multiple Kit User Manual for Fresh Frozen Tissue (Advanced Cell Diagnostics, Inc.). Briefly, sections were fixed in 4% PFA, dehydrated, and treated with pretreatment 4 protease solution. Sections were then incubated for target probes for mouse calcium calmodulin kinase II (CamKIIα, accession number NM_009792.3), vesicular glutamate transporter 1 (slc17a7, accession number NM_182993.2), dopamine D1 receptor (DrD1, accession number NM_010076.3) kappa opioid receptor (Oprk1, accession number NM_001204371.1) and prodynorphin (Pdyn, accession number NM_018863.3). All target probes were obtained from Advanced Cell Diagnostics. Following probe hybridization, sections underwent a series of probe signal amplification steps followed by incubation of fluorescently labeled robes designed to target the specific channel associated with the probes. Slides were counterstained with DAPI, and coverslips were mounted with Vectashield Hard Set mounting medium (Vector Laboratories. Images were obtained on an Olympus Fluoview 3000 confocal microscope and analyzed with HALO software. To analyze the images, each image was opened in the HALO software. DAPI positive cells were then registered and used as markers for individual cells. A positive cell consisted of an area within the radius of a DAPI nuclear staining that measured at least 3 positive pixels for receptor probes, or 10 total positive pixels for neurotransmitter probes. Two - three separate slices from the DMS or BLA were used for each animal and that total is presented in the data.
Quantification and Statistical analyses
Behavioral Analysis
All behavior data was averaged and plotted in GraphPad Prism 8.0 (Graphpad, La Jolla, CA). Data are expressed as mean ± SEM and analyzed using Linear Regression, Student’s t test, one-way ANOVA or a two-way repeated-measures ANOVA followed by post hoc tests as appropriate. Unless stated otherwise, N represents number of animals and are included in the figure legends. All statistical analyses are included in the figure legends. Statistical significance was taken as *p < 0.05, **p < 0.01, and ***p < 0.001 and ****p < 0.0005. D’Agostino-Pearson’s tests were performed when necessary to determine normal distribution of data, and F statistics were used to determine equal variance. For across animal comparisons using t tests, Welch’s t tests were performed assuming the standard deviation was not equal.
Photometry Analysis
Custom MATLAB scripts were developed for analyzing fiber photometry data in context of mouse behavior and can be accessed via GitHub. Analyses for κLight1.3a and GCaMP6s fluorescence were identical. To account for slow photobleaching artifacts, a double exponential curve (dec) was fit to the raw isosbestic 405/435 and 470/490 excitation signal, and this fit was subtracted from the raw isosbestic and excitation to obtain dec_fits for both. Then, to account for motion correction and differences in light intensity, the dec_isosbestic was refit to the dec_excitation to obtain refit_isosbestic. This refit_isosbestic was then subtracted from the dec_excitation to obtain a motion corrected signal. This is what is typically done for cpGFP-based sensors such as dLight117. Finally, this time series for the entire session was converted to dF/F. This dF/F was then used to align to specific behavioral events, over a window of 30s before and 30s after the behavioral event of interest (reinforced nosepoke for operant conditioning, cue onset for Pavlovian conditioning) and z-scored to the mean and standard deviation of this total 60s time window. For operant behavior, data are presented as z-score across −5s to 10s (κLight1.3a) or −10s to 20s (GcaMP6s), with 0s signifying a reinforced nosepoke. For Pavlovian, data are presented as z-score across −30s to 30s, with 0s signifying the start of cue delivery. Normalized or baseline-subtracted peak z-score values were obtained using MATLAB 9.6 (The MathWorks, Natick, MA) across the time windows indicated above and plotted in GraphPad Prism 8.0 (Graphpad, La Jolla, CA). Data are expressed as mean ± SEM and analyzed using Student’s t test, one-way ANOVA or a two-way repeated-measures ANOVA followed by post hoc tests as appropriate. Unless stated otherwise, N represents number of animals and are included in the figure legends. All statistical analyses are included in the figure legends. Statistical significance was taken as *p < 0.05, **p < 0.01, and ***p < 0.001 and ****p < 0.0005. D’Agostino-Pearson’s tests were performed when necessary to determine normal distribution of data, and F statistics were used to determine equal variance. For across animal comparisons using t tests, Welch’s t tests were performed assuming the standard deviation was not equal.
GLM Analysis
To determine changes in peak photometry signals caused by behavioral variables, we used a Gaussian generalized linear model (GLM) and can be accessed via GitHub. In the GLM, each available dependent (trial number: I1, preceding inter-trial interval: I2, subsequent inter-trial interval: I3, and latency to consume pellet after a nosepoke: I4) was taken as an explanatory variable to predict the fluorescence detected. Therefore, the final GLM equation was:
Where F(t) represents the peak photometry signal detected, β represents the model coefficients and I represents each explanatory (dependent) variable, defined above. B0 is the intercept term and ϵ is the error term. To determine the specific impact of each explanatory variable on the model’s ability to predict fluorescence, we used a drop-out approach, whereby we removed each predictor from the GLM equation one at a time. For every model, we performed 5-fold cross validation, thus utilizing 80% of the data for training, and 20% of the data for testing on each run. We then first calculated the model’s coefficient of determination (R2) when all predictors were used. Next, we removed predictors individually, one at a time, and re-calculated the model’s coefficient of determination. This approach allowed us to isolate the role of each explanatory variable in terms of the full model’s predictive power, as we obtained a distribution of R2 values for each model across each of the 5 folds. Folds were split identically for the full model and each subsequent model with one predictor variable removed, and this approach ensured that the entirety of the dataset would be tested on at some point. We additionally applied this approach to data from the early, trained, and extinction sessions to assess the impact of each explanatory variable during those sessions. Proportion of variance data collected for the full model and ther drop-out of each variable were plotted in GraphPad Prism 8.0 (Graphpad, La Jolla, CA). Data are expressed as mean ± SEM and analyzed using one-way ANOVA followed by post hoc tests as appropriate. N represents number of trials and are included in the figure legends. All statistical analyses are included in the figure legends. Statistical significance was taken as *p < 0.05, **p < 0.01, and ***p < 0.001 and ****p < 0.0005. D’Agostino-Pearson’s tests were performed when necessary to determine normal distribution of data, and F statistics were used to determine equal variance.
Two-Photon Calcium Imaging Analyses
Imaged cells were tracked longitudinally by concatenating all the recording sessions per animal (Pavlovian and operant conditioning, and operant extinction) and analyzing them through Suite2p118,119 (HHMI, Janelia Research Campus). Sorted cells were manually verified as well and cells displaying no activity, or no appreciable fluctuations in activity above neuropil-extracted activity were removed. Following tracking through Suite2p, GcaMP6s and Td-Tomato cells were overlaid in ImageJ/Fiji and quantified for overlap. This resulted in fluorescence over time series files that were used to generate peri-event histograms for each cell, across every trial. This was done individually for each animal, each session. Next, peri-event histograms were generated across 97 frames (12s) by averaging across all the trials, z-scored using the mean fluorescence and standard deviation of a predefined time window (1s prior to the initiation of active wheel rotation). This trial-averaged peri-event histogram was used for our spectral clustering analyses below. Spectral clustering was performed using Python code based on49 and can be accessed via GitHub. Trial-averaged peri-event histograms first underwent dimensionality reduction using Principal Component Analyses (PCA) to determine the number PCs that explain 90% of the variance (4 in the case of this dataset (Fig. S3H). We used spectral clustering as it has been showed to produce stable results for high dimensional datasets49,50. Spectral clustering was performed using the Scikit-learn function sklearn. cluster. spectralclustering by setting the maximum number of clusters as 20, and a nearest neighbor parameter sweep at 1-(total number of neurons), varied by a factor of 50. This yielded a silhouette score reflecting the accuracy of clustering across all possible cluster and neighbor combinations (Fig. S3I). Spectral clustering was carried out on individual animals first to ensure that the classification and silhouette scores were similar before performing it on the entire dataset. Once clustering was performed, neurons were assigned a unique label for the cluster they belonged to. This classification from the trained day was then applied to the early and extinction days to generate fluorescence traces and heat map raster plots. Data are presented as z-score across −1s to 11s, with 0s signifying the start of wheel rotation in the trial. For spiral stimulation experiments, the sessions were concatenated along with the recordings from the trained operant session and analyzed in Suite2p for tracking to determine the identities of the neurons enriched in behavioral clusters. Point spread functions to assess the utility for microprims for spatial stimulation were assessed in47. Per-event histograms were obtained from the fluorescence time-series as trial-averaged data across a 10s window, aligned to each stimulation timepoint. All data was was z-scored using the mean fluorescence and standard deviation of a predefined time window (1s prior to laser stimulation). Data are presented as z-score across −5s to 5s, with 0s signifying the start of laser stimulation. Peak z-score values were obtained across the time windows indicated above and plotted in GraphPad Prism 8.0 (Graphpad, La Jolla, CA). Data are expressed as mean ± SEM and analyzed using one-way ANOVA followed by post hoc tests as appropriate. N represents number of animals and are included in the figure legends. All statistical analyses are included in the figure legends. Statistical significance was taken as *p < 0.05, **p < 0.01, and ***p < 0.001 and ****p < 0.0005. D’Agostino-Pearson’s tests were performed when necessary to determine normal distribution of data, and F statistics were used to determine equal variance.
RNN Decoding
To determine if inter-trial interval (ITI) could be decoded from neurological activity of the preceding and/or subsequent trials, we trained an artificial Recurrent Neural Network (RNN), specifically a Gated Recurrent Unit (GRU) model and can be accessed via GitHub. To achieve this, we first binned ITI values into quintiles (1/5th) separately for each animal. We additionally standardized all training data as implemented in scikit-learn’s StandardScaler to ensure stable training of the neural network. To prevent data leakage, in all cases test data were standardized using the mean and standard deviation derived from the training set. We then trained the model to use the activity of the preceding trial, the subsequent trial, or both, to predict the bin of the corresponding ITI. Unless otherwise stated, all models were trained and tested on trials from retained cells on the trained day. For all further steps, the model was trained with the activity of both the preceding and subsequent trials, as this resulted in the highest accuracy. To utilize the RNN’s ability to consider temporal dependencies in the data, we input features as part of 2 distinct time steps, relative to the ITI: the activity of the trial preceding the ITI, and the activity of the trial succeeding the ITI. The network was constructed with an initial GRU layer with 256 units, followed by 2 hidden GRU layers with 128, and 64 units respectively. We then placed a hidden layer with 32 units and Rectified Linear Unit (ReLU) activation function. Finally, the output layer was constructed as a Dense layer with 5 units and a softmax activation function. AdamW120 was used as the optimizer. We used categorical cross entropy as the loss function, which aims to maximize the log likelihood of the correct class being predicted121. During the training process, model parameters were continually refined to minimize the loss function. The entirety of the model was implemented using tensorflow and keras. As a method of regularization, we used dropout122 and placed these layers after the input and all hidden layers. To determine whether dropout improved model performance, and what the ideal dropout rate was, we used a self-implemented grid search where we tested values in the range (0, 0.1, 0.3, 0.5, 0.7). We additionally tested learning rates in the range (10–1…10–5) to determine the optimal set of parameters for model performance. We trained the model with 80% of the data, using 10% of the training dataset for validation during training, and held out the final 20% as test data. The split between training and test data was consistent across each set of parameters tested. The combination of values which provided the best overall accuracy during final testing was used for all further models. To ensure that we trained each model for the ideal number of epochs, we monitored validation loss, and used EarlyStopping as implemented in the keras callback with a patience of 10. Thus, when validation loss does not improve for 10 consecutive epochs, the model stops training. The test dataset and all future data evaluated by the model utilizes the weights from the model with the lowest validation loss. To ensure model generalizability, we performed 5-fold cross validation on the reported RNN model. 10% of the training data was utilized as a validation dataset for each fold, facilitating the previously discussed EarlyStopping mechanism. This scheme ensures that each unique combination of input and output will be in the test dataset exactly once. From the cross validated results, we computed the proportion of correct predictions overall, as well as the proportion of correct predictions for data belonging to each class (ITI bin), relative to all predictions for that class. All reported metrics are computed from model predictions on all test data across all folds, which encompasses the entirety of the dataset used in the model. We then mapped the cluster identity of each cell from the spectral clustering to each of our cross-validated model’s predictions. We performed 1000 random shuffles for each reported model in total and utilized this as the null distribution for statistical testing. Finally, to further assess the impact of each cluster on decoding ability, we evaluated model performance after removing each cluster one at a time. Before performing these analyses, we randomly subsampled a number of cells from each cluster equal to the number of cells in the smallest cluster. This was to ensure that the decoding ability of each cluster was assessed independently of cluster size. We then trained the model with all the cells to establish baseline model metrics. To determine the importance of each cluster in the decoding, we trained the model after individually removing each cluster and assessing the impact of each drop-out on reported model metrics. Model architecture, training, testing, and cross-validation was performed in the same manner as for the full model. For all scripts, random seeds were used to ensure reproducibility. RNN accuracy values obtained were plotted in GraphPad Prism 8.0 (Graphpad, La Jolla, CA). Data are expressed as mean ± SEM and analyzed using Student’s t test and one-way ANOVA followed by post hoc tests as appropriate. N represents number of trials and are included in the figure legends. All statistical analyses are included in the figure legends. Statistical significance was taken as *p < 0.05, **p < 0.01, and ***p < 0.001 and ****p < 0.0005. D’Agostino-Pearson’s tests were performed when necessary to determine normal distribution of data, and F statistics were used to determine equal variance.
Supplementary Material
Key Resources Table
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Antibodies | ||
| Chicken-anti-GFP | Abcam | Ab13970 |
| Bacterial and Virus Strains | ||
| AAV-DJ-hsyn-GCaMP6s | Stanford University Gene Vector and Viral Core | N/A |
| AAV5-CAG-DIO-κLight1.3a | Tian Lab, Bruchas Lab | N/A |
| AAV5-CMV-myc-NLS-Cre | The Hope Center Viral Core – Washington University at St. Louis | N/A |
| AAV5-EF1α-ChR2-EYFP | Bruchas Lab | N/A |
| AAV2retro-CMV-myc-NLS-Cre | The Hope Center Viral Core – Washington University at St. Louis | N/A |
| AAV-DJ-EF1α-DIO-GCAMP6s | Stanford University Gene Vector and Viral Core | N/A |
| AAV5-EF1α-DIO-PPO-Venus | Bruchas Lab | N/A |
| AAV-DJ-CamKIIα-GCaMP6s | Stanford University Gene Vector and Viral Core | N/A |
| AAV1-CMV-DIO-SaCas9-U6-sgROSA | Zweifel Lab | N/A |
| AAV1-CMV-DIO-SaCas9-U6-sgOPRK1 | Zweifel Lab | N/A |
| AAV1- EF1α-DIO-Cre | Addgene | CAT#121675-AAV1 |
| AAV1- EF1α-FLPO | Addgene | CAT#55637-AAV5 |
| AAV5-EF1α-DIO-ChRimsonR-tdTomato | Bruchas Lab | N/A |
| AAV8-CamKII-rsChRmine-oScarlet | Stanford University Gene Vector and Viral Core | N/A |
| Biological Samples | N/A | |
| Chemicals, Peptides, and Recombinant Proteins | ||
| VECTASHIELD Hardset Antifade Mounting Medium | Vector Laboratories | CAT#H-1400 |
| VECTASHIELD Hardset Antifade Mounting Medium with DAPI | Vector Laboratories | CAT#H-1800 |
| LY 2456302, Aticaprant hydrochloride | NIDA Drug Supply Program | N/A |
| Naloxone hydrocholoride | Tocris | CAT#0599; CAS: 357-08-4 |
| Critical Commercial Assays | ||
| RNAscope Fluorescent Multiplex Kit 2.0 | Advanced Cell Diagnostics | CAT#320850 |
| Mm-OPRK1 | Advanced Cell Diagnostics | CAT#316111 |
| Mm-Pdyn | Advanced Cell Diagnostics | CAT#318771 |
| Mm-Slc17a7 | Advanced Cell Diagnostics | CAT#503511 |
| Mm-CamKIIα | Advanced Cell Diagnostics | CAT#445231 |
| Mm-DrD1 | Advanced Cell Diagnostics | CAT#461901 |
| Deposited Data | N/A | |
| Experimental Models: Cell Lines | N/A | |
| Experimental Models: Organisms/Strains | ||
| Ai14 x DrD1-cre | This paper, Bred in house | N/A |
| KOR-cre | (Cai et al., 2016)102 | Gift from Charles Chavkin, Strain No: 035045 (Jackson Laboratories) |
| Pdyn-IRES-Cre | (Krashes et al., 2014)103 | Gift from Dr. Richard Palmiter, Strain No: 027958 (Jackson Laboratories) |
| Pdyn fl/fl | (Yang et al., 2023)104 | Gift from Charley Chavkin, Strain No: 037400 (Jackson Laboratories) |
| KORfl/fl | (Ehrich et al., 2015)105 | Gift from Charley Chavkin, Strain No: 030076 (Jackson Laboratories) |
| Vglut1-Cre | (Harris et al., 2014)106 | Gift from Larry Zweifel, Strain No: 037512 (Jackson Laboratories) |
| Oligonucleotides | N/A | |
| Recombinant DNA | N/A | |
| Software and Algorithms | ||
| Suite2p | HHMI, Janelia Research Campus | https://github.com/MouseLand/suite2p |
| FIJI/ImageJ | NIH | https://imagej.net/software/fiji/ |
| MATLAB | Mathworks | https://www.mathworks.com/products/matlab.html |
| Med-PC V Software Suite | Med-Associates Inc. | https://www.med-associates.com/med-pc-v/ |
| Ethovision 10 | Noldus | https://www.noldus.com/ethovision-xt |
| Synapse | Tucker-Davis Technologies | https://www.tdt.com/docs/hardware/rz10x-lux-integrated-processor/ |
| PRISM 8 | Graphpad | https://www.graphpad.com/ |
| HALO Image Analysis | Indica Labs | https://indicalab.com/halo/ |
| Clampfit v11.0.3.03 | Molecular Devices | https://www.moleculardevices.com/products/axon-patch-clamp-system/acquisition-and-analysis-software/pclamp-software-suite |
| Olympus Fluoview 3000 | Olympus Life Science | https://www.olympus-lifescience.com/en/laser-scanning/fv4000/ |
| Illustrator CS6 | Adobe | https://www.adobe.com/products/illustrator.html |
| Analysis pipelines for photometry and 2p data | Github | https://zenodo.org/records/21286480 |
| Other | ||
| Entended Right Angle Microprisms | OptoSigma | https://www.optosigma.com/eu_en/optics/prisms/right-angle-prism/extended-right-angle-microprisms-two-photon-endoscopic-imaging-protected-aluminum-coated-hypotenuse.html |
| Ultima 2P Plus w/ NeuraLight 3D Spatial Light Modulator | Bruker | https://www.bruker.com/en/products-and-solutions/fluorescence-microscopy/multiphoton-microscopes/ultima-2pplus.html |
| Insight X3 Tunable Ultrafast Laser | Bruker | https://www.spectra-physics.com/en/f/insight-x3-tunable-laser |
| Spirit High Power Femtosecond Lasers | Bruker | https://www.spectra-physics.com/en/f/spirit-femtosecond-laser |
| RZ-10X Lux Integrated Processor | Tucker Davis Technologies | https://www.tdt.com/docs/hardware/rz10x-lux-integrated-processor/ |
Highlights:
DMS dynorphin production is necessary for goal-directed behavior
Rapid DMS dynorphin release promotes engagement in goal-directed behavior
DMSpdyn neuron activity influences engagement in goal-directed behavior
Dynorphin-KOR modulation at BLA-DMS terminals promotes goal-directed behavior
Acknowledgements
We thank Taylor Hobbs, Carina Pizzano and Bailey Wells for animal colony maintenance. We thank Azra Suko for lab management and virus preparation. We thank Scott Ng-Evans for operations support and Jazmyne Fosha, and Lusine Eyde for programmatic support. We thank Vijay Namboodiri for original Python code and feedback, and Christian Pederson for original MATLAB code for analyses. We thank Howard Fields, Larry Zweifel, Daniel Castro, Sean Piantadosi and Marta Trzeciak for valuable discussions regarding and comments on the manuscript. R.G is funded by NIH grant K99DA058709. M.H is funded by NIH grant F31DA053706. G.D.S is funded by NSF grant 1934288 and NIH R37DA032750. M.R.B. and R.G. were partially supported by NIH grants R37DA033396, DA048736, 5P50MH119467 (5821) and the Weill Neurohub. M.R.B and G.D.S and microscopy support were also supported by NIH grant P30DA048736.
Footnotes
Declaration of Interests
The authors declare no competing interests.
References:
- 1.Gillan CM, and Robbins TW. (2014). Goal-directed learning and obsessive–compulsive disorder. Philos. Trans. R. Soc. B Biol. Sci. 369, 20130475. 10.1098/rstb.2013.0475. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Ironside M, Amemori K-I, McGrath CL, Pedersen ML, Kang MS, Amemori S, Frank MJ, Graybiel AM, and Pizzagalli DA. (2020). Approach-Avoidance Conflict in Major Depressive Disorder: Congruent Neural Findings in Humans and Nonhuman Primates. Biol. Psychiatry 87, 399–408. 10.1016/j.biopsych.2019.08.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Yoshida K, Drew MR, Kono A, Mimura M, Takata N, and Tanaka KF. (2021). Chronic social defeat stress impairs goal-directed behavior through dysregulation of ventral hippocampal activity in male mice. Neuropsychopharmacology 46, 1606–1616. 10.1038/s41386-021-00990-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Hogarth L. (2020). Addiction is driven by excessive goal-directed drug choice under negative affect: translational critique of habit and compulsion theory. Neuropsychopharmacology 45, 720–735. 10.1038/s41386-020-0600-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Balleine BW, Delgado MR, and Hikosaka O. (2007). The Role of the Dorsal Striatum in Reward and Decision-Making. J. Neurosci. 27, 8161–8165. 10.1523/JNEUROSCI.1554-07.2007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Hikosaka O, Sakamoto M, and Usui S. (1989). Functional properties of monkey caudate neurons. III. Activities related to expectation of target and reward. J. Neurophysiol. 61, 814–832. 10.1152/jn.1989.61.4.814. [DOI] [PubMed] [Google Scholar]
- 7.Hollerman JR, Tremblay L, and Schultz W. (1998). Influence of reward expectation on behavior-related neuronal activity in primate striatum. J. Neurophysiol. 80, 947–963. 10.1152/jn.1998.80.2.947. [DOI] [PubMed] [Google Scholar]
- 8.O’Doherty JP, Deichmann R, Critchley HD, and Dolan RJ. (2002). Neural responses during anticipation of a primary taste reward. Neuron 33, 815–826. 10.1016/s0896-6273(02)00603-7. [DOI] [PubMed] [Google Scholar]
- 9.Yin HH, Ostlund SB, Knowlton BJ, and Balleine BW. (2005). The role of the dorsomedial striatum in instrumental conditioning. Eur. J. Neurosci. 22, 513–523. 10.1111/j.1460-9568.2005.04218.x. [DOI] [PubMed] [Google Scholar]
- 10.Corbit LH, and Janak PH. (2010). Posterior dorsomedial striatum is critical for both selective instrumental and Pavlovian reward learning. Eur. J. Neurosci. 31, 1312–1321. 10.1111/j.1460-9568.2010.07153.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Gerfen CR, and Surmeier DJ. (2011). Modulation of striatal projection systems by dopamine. Annu. Rev. Neurosci. 34, 441–466. 10.1146/annurev-neuro-061010-113641. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Shan Q, Ge M, Christie MJ, and Balleine BW. (2014). The acquisition of goal-directed actions generates opposing plasticity in direct and indirect pathways in dorsomedial striatum. J. Neurosci. Off. J. Soc. Neurosci. 34, 9196–9201. 10.1523/JNEUROSCI.0313-14.2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Gremel CM, and Costa RM. (2013). Orbitofrontal and striatal circuits dynamically encode the shift between goal-directed and habitual actions. Nat. Commun. 4, 2264. 10.1038/ncomms3264. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Bloem B, Huda R, Sur M, and Graybiel AM. (2017). Two-photon imaging in mice shows striosomes and matrix have overlapping but differential reinforcement-related responses. eLife 6, e32353. 10.7554/eLife.32353. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Bloem B, Huda R, Amemori K, Abate AS, Krishna G, Wilson AL, Carter CW, Sur M, and Graybiel AM. (2022). Multiplexed action-outcome representation by striatal striosome-matrix compartments detected with a mouse cost-benefit foraging task. Nat. Commun. 13, 1541. 10.1038/s41467-022-28983-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Peak J, Chieng B, Hart G, and Balleine BW. (2020). Striatal direct and indirect pathway neurons differentially control the encoding and updating of goal-directed learning. eLife 9, e58544. 10.7554/eLife.58544. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Smith SJ, Hawrylycz M, Rossier J, and Sümbül U. (2020). New light on cortical neuropeptides and synaptic network plasticity. Curr. Opin. Neurobiol. 63, 176–188. 10.1016/j.conb.2020.04.002. [DOI] [PubMed] [Google Scholar]
- 18.van den Pol AN. (2012). Neuropeptide transmission in brain circuits. Neuron 76, 98–115. 10.1016/j.neuron.2012.09.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Cahill C, Tejeda HA, Spetea M, Chen C, and Liu-Chen L-Y. (2022). Fundamentals of the Dynorphins / Kappa Opioid Receptor System: From Distribution to Signaling and Function. Handb. Exp. Pharmacol. 271, 3–21. 10.1007/164_2021_433. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Reiner A, and Anderson KD. (1990). The patterns of neurotransmitter and neuropeptide co-occurrence among striatal projection neurons: conclusions based on recent findings. Brain Res. Brain Res. Rev. 15, 251–265. 10.1016/0165-0173(90)90003-7. [DOI] [PubMed] [Google Scholar]
- 21.Shippenberg TS, Zapata A, and Chefer VI. (2007). Dynorphin and the pathophysiology of drug addiction. Pharmacol. Ther. 116, 306–321. 10.1016/j.pharmthera.2007.06.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Bruchas MR, Land BB, and Chavkin C. (2010). The dynorphin/kappa opioid system as a modulator of stress-induced and pro-addictive behaviors. Brain Res. 1314, 44–55. 10.1016/j.brainres.2009.08.062. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Tejeda HA, and Bonci A. (2019). Dynorphin/kappa-opioid receptor control of dopamine dynamics: Implications for negative affective states and psychiatric disorders. Brain Res. 1713, 91–101. 10.1016/j.brainres.2018.09.023. [DOI] [PubMed] [Google Scholar]
- 24.Limoges A, Yarur HE, and Tejeda HA. (2022). Dynorphin/kappa opioid receptor system regulation on amygdaloid circuitry: Implications for neuropsychiatric disorders. Front. Syst. Neurosci. 16. 10.3389/fnsys.2022.963691. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Chavkin C, James IF, and Goldstein A. (1982). Dynorphin is a specific endogenous ligand of the kappa opioid receptor. Science 215, 413–415. 10.1126/science.6120570. [DOI] [PubMed] [Google Scholar]
- 26.Al-Hasani R, and Bruchas MR. (2011). Molecular mechanisms of opioid receptor-dependent signaling and behavior. Anesthesiology 115, 1363–1381. 10.1097/ALN.0b013e318238bba6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Corder G, Castro DC, Bruchas MR, and Scherrer G. (2018). Endogenous and Exogenous Opioids in Pain. Annu. Rev. Neurosci. 41, 453–473. 10.1146/annurev-neuro-080317-061522. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.MARDER E. (2012). NEUROMODULATION OF NEURONAL CIRCUITS: BACK TO THE FUTURE. Neuron 76, 1–11. 10.1016/j.neuron.2012.09.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Nair A, Karigo T, Yang B, Ganguli S, Schnitzer MJ, Linderman SW, Anderson DJ, and Kennedy A. (2023). An approximate line attractor in the hypothalamus encodes an aggressive state. Cell 186, 178–193.e15. 10.1016/j.cell.2022.11.027. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Atwood BK, Lovinger DM, and Mathur BN. (2014). Presynaptic long-term depression mediated by Gi/o-coupled receptors. Trends Neurosci. 37, 663–673. 10.1016/j.tins.2014.07.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Meshul CK, and McGinty JF. (2000). Kappa opioid receptor immunoreactivity in the nucleus accumbens and caudate-putamen is primarily associated with synaptic vesicles in axons. Neuroscience 96, 91–99. 10.1016/s0306-4522(99)90481-5. [DOI] [PubMed] [Google Scholar]
- 32.Svingos AL, Colago EE, and Pickel VM. (1999). Cellular sites for dynorphin activation of kappa-opioid receptors in the rat nucleus accumbens shell. J. Neurosci. Off. J. Soc. Neurosci. 19, 1804–1813. 10.1523/JNEUROSCI.19-05-01804.1999. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Hjelmstad GO, and Fields HL. (2003). Kappa Opioid Receptor Activation in the Nucleus Accumbens Inhibits Glutamate and GABA Release Through Different Mechanisms. J. Neurophysiol. 89, 2389–2395. 10.1152/jn.01115.2002. [DOI] [PubMed] [Google Scholar]
- 34.Mu P, Neumann PA, Panksepp J, Schlüter OM, and Dong Y. (2011). Exposure to Cocaine Alters Dynorphin-Mediated Regulation of Excitatory Synaptic Transmission in Nucleus Accumbens Neurons. Biol. Psychiatry 69, 228–235. 10.1016/j.biopsych.2010.09.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Tejeda HA, Wu J, Kornspun AR, Pignatelli M, Kashtelyan V, Krashes MJ, Lowell BB, Carlezon WA, and Bonci A. (2017). Pathway- and Cell-Specific Kappa-Opioid Receptor Modulation of Excitation-Inhibition Balance Differentially Gates D1 and D2 Accumbens Neuron Activity. Neuron 93, 147–163. 10.1016/j.neuron.2016.12.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Nygard SK, Hourguettes NJ, Sobczak GG, Carlezon WA, and Bruchas MR. (2016). Stress-Induced Reinstatement of Nicotine Preference Requires Dynorphin/Kappa Opioid Activity in the Basolateral Amygdala. J. Neurosci. 36, 9937–9948. 10.1523/JNEUROSCI.0953-16.2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Baxter MG, and Murray EA. (2002). The amygdala and reward. Nat. Rev. Neurosci. 3, 563–573. 10.1038/nrn875. [DOI] [PubMed] [Google Scholar]
- 38.Ostlund SB, and Balleine BW. (2008). Differential Involvement of the Basolateral Amygdala and Mediodorsal Thalamus in Instrumental Action Selection. J. Neurosci. 28, 4398–4405. 10.1523/JNEUROSCI.5472-07.2008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Parkes SL, and Balleine BW. (2013). Incentive Memory: Evidence the Basolateral Amygdala Encodes and the Insular Cortex Retrieves Outcome Values to Guide Choice between Goal-Directed Actions. J. Neurosci. 33, 8753–8763. 10.1523/JNEUROSCI.5071-12.2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Wassum KM, and Izquierdo A. (2015). The basolateral amygdala in reward learning and addiction. Neurosci. Biobehav. Rev. 57, 271–283. 10.1016/j.neubiorev.2015.08.017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Courtin J, Bitterman Y, Müller S, Hinz J, Hagihara KM, Müller C, and Lüthi A. (2022). A neuronal mechanism for motivational control of behavior. Science 375, eabg7277. 10.1126/science.abg7277. [DOI] [PubMed] [Google Scholar]
- 42.Giovanniello JR, Paredes N, Wiener A, Ramírez-Armenta K, Oragwam C, Uwadia HO, Yu AL, Lim K, Pimenta JS, Vilchez GE, et al. (2025). A dual-pathway architecture for stress to disrupt agency and promote habit. Nature 640, 722–731. 10.1038/s41586-024-08580-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Dong C, Gowrishankar R, Jin Y, He XJ, Gupta A, Wang H, Sayar-Atasoy N, Flores RJ, Mahe K, Tjahjono N, et al. (2024). Unlocking opioid neuropeptide dynamics with genetically encoded biosensors. Nat. Neurosci. 27, 1844–1857. 10.1038/s41593-024-01697-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Lowe SL, Wong CJ, Witcher J, Gonzales CR, Dickinson GL, Bell RL, Rorick-Kehn L, Weller M, Stoltz RR, Royalty J, et al. (2014). Safety, tolerability, and pharmacokinetic evaluation of single- and multiple-ascending doses of a novel kappa opioid receptor antagonist LY2456302 and drug interaction with ethanol in healthy subjects. J. Clin. Pharmacol. 54, 968–978. 10.1002/jcph.286. [DOI] [PubMed] [Google Scholar]
- 45.Gordon-Fennell A, Barbakh JM, Utley MT, Singh S, Bazzino P, Gowrishankar R, Bruchas MR, Roitman MF, and Stuber GD. (2023). An open-source platform for head-fixed operant and consummatory behavior. eLife 12, e86183. 10.7554/eLife.86183. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Namboodiri VMK, Otis JM, van Heeswijk K, Voets ES, Alghorazi RA, Rodriguez-Romaguera J, Mihalas S, and Stuber GD. (2019). Single-cell activity tracking reveals that orbitofrontal neurons acquire and maintain a long-term memory to guide behavioral adaptation. Nat. Neurosci. 22, 1110–1121. 10.1038/s41593-019-0408-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Hirokawa J, Vaughan A, Masset P, Ott T, and Kepecs A. (2019). Frontal cortex neuron types categorically encode single decision variables. Nature 576, 446–451. 10.1038/s41586-019-1816-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Glaser JI, Benjamin AS, Chowdhury RH, Perich MG, Miller LE, and Kording KP. (2020). Machine Learning for Neural Decoding. eNeuro 7. 10.1523/ENEURO.0506-19.2020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Fremeau RT, Troyer MD, Pahner I, Nygaard GO, Tran CH, Reimer RJ, Bellocchio EE, Fortin D, Storm-Mathisen J, and Edwards RH. (2001). The Expression of Vesicular Glutamate Transporters Defines Two Classes of Excitatory Synapse. Neuron 31, 247–260. 10.1016/S0896-6273(01)00344-0. [DOI] [PubMed] [Google Scholar]
- 50.Copits BA, Gowrishankar R, O’Neill PR, Li J-N, Girven KS, Yoo JJ, Meshik X, Parker KE, Spangler SM, Elerding AJ, et al. (2021). A photoswitchable GPCR-based opsin for presynaptic inhibition. Neuron 109, 1791–1809.e11. 10.1016/j.neuron.2021.04.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Fellinger L, Jo YS, Hunker AC, Soden ME, Elum J, Juarez B, and Zweifel LS. (2021). A midbrain dynorphin circuit promotes threat generalization. Curr. Biol. CB 31, 4388–4396.e5. 10.1016/j.cub.2021.07.047. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Gordon-Fennell L, Farero RD, Burgeno LM, Murray NL, Abraham AD, Soden ME, Stuber GD, Chavkin C, Zweifel LS, and Phillips PEM. (2023). Kappa Opioid Receptors in Mesolimbic Terminals Mediate Escalation of Cocaine Consumption. bioRxiv, 2023.12.21.572842. 10.1101/2023.12.21.572842. [DOI] [Google Scholar]
- 53.O’Doherty J, Dayan P, Schultz J, Deichmann R, Friston K, and Dolan RJ. (2004). Dissociable Roles of Ventral and Dorsal Striatum in Instrumental Conditioning. Science 304, 452–454. 10.1126/science.1094285. [DOI] [PubMed] [Google Scholar]
- 54.Balleine BW. (2019). The Meaning of Behavior: Discriminating Reflex and Volition in the Brain. Neuron 104, 47–62. 10.1016/j.neuron.2019.09.024. [DOI] [PubMed] [Google Scholar]
- 55.Al-Hasani R, Wong J-MT, Mabrouk OS, McCall JG, Schmitz GP, Porter-Stransky KA, Aragona BJ, Kennedy RT, and Bruchas MR. (2018). In vivo detection of optically-evoked opioid peptide release. eLife 7, e36520. 10.7554/eLife.36520. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Rappleye M, Gordon-Fennel A, Castro DC, Matarasso AK, Zamorano CA, Stine C, Wait SJ, Lee JD, Siebart JC, Suko A, et al. (2022). Opto-MASS: a high-throughput engineering platform for genetically encoded fluorescent sensors enabling all-optical in vivo detection of monoamines and opioids. Preprint at bioRxiv, 10.1101/2022.06.01.494241 https://doi.org/10.1101/2022.06.01.494241. [DOI] [Google Scholar]
- 57.Tian L, Dong C, Gowrishankar R, Jin Y, He X, Gupta A, Wang H, Atasoy N, Flores-Garcia R, Mahe K, et al. (2023). Unlocking opioid neuropeptide dynamics with genetically-encoded biosensors. Preprint, 10.21203/rs.3.rs-2871083/v1 https://doi.org/10.21203/rs.3.rs-2871083/v1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Zhou X, Stine C, Prada PO, Fusca D, Assoumou K, Dernic J, Bhat MA, Achanta AS, Johnson JC, Jadhav S, et al. (2023). Development of a genetically-encoded sensor for probing endogenous nociceptin opioid peptide release. Preprint at bioRxiv, 10.1101/2023.05.26.542102 https://doi.org/10.1101/2023.05.26.542102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Wee S, and Koob GF. (2010). The role of the dynorphin-kappa opioid system in the reinforcing effects of drugs of abuse. Psychopharmacology (Berl.) 210, 121–135. 10.1007/s00213-010-1825-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Hurd YL, Brown EE, Finlay JM, Fibiger HC, and Gerfen CR. (1992). Cocaine self-administration differentially alters mRNA expression of striatal peptides. Brain Res. Mol. Brain Res. 13, 165–170. 10.1016/0169-328x(92)90058-j. [DOI] [PubMed] [Google Scholar]
- 61.Hurd YL, and Herkenham M. (1993). Molecular alterations in the neostriatum of human cocaine addicts. Synap. N. Y. N 13, 357–369. 10.1002/syn.890130408. [DOI] [PubMed] [Google Scholar]
- 62.Fagergren P, Smith HR, Daunais JB, Nader MA, Porrino LJ, and Hurd YL. (2003). Temporal upregulation of prodynorphin mRNA in the primate striatum after cocaine self-administration. Eur. J. Neurosci. 17, 2212–2218. 10.1046/j.1460-9568.2003.02636.x. [DOI] [PubMed] [Google Scholar]
- 63.Redila VA, and Chavkin C. (2008). Stress-induced reinstatement of cocaine seeking is mediated by the kappa opioid system. Psychopharmacology (Berl.) 200, 59–70. 10.1007/s00213-008-1122-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Ehrich JM, Phillips PEM, and Chavkin C. (2014). Kappa Opioid Receptor Activation Potentiates the Cocaine-Induced Increase in Evoked Dopamine Release Recorded In Vivo in the Mouse Nucleus Accumbens. Neuropsychopharmacology 39, 3036–3048. 10.1038/npp.2014.157. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Wee S, Orio L, Ghirmai S, Cashman JR, and Koob GF. (2009). Inhibition of kappa opioid receptors attenuated increased cocaine intake in rats with extended access to cocaine. Psychopharmacology (Berl.) 205, 565–575. 10.1007/s00213-009-1563-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Walker BM, Zorrilla EP, and Koob GF. (2011). Systemic κ-opioid receptor antagonism by nor-binaltorphimine reduces dependence-induced excessive alcohol self-administration in rats. Addict. Biol. 16, 116–119. 10.1111/j.1369-1600.2010.00226.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Mohebi A, Wei W, Pelattini L, Kim K, and Berke JD. (2024). Dopamine transients follow a striatal gradient of reward time horizons. Nat. Neurosci. 27, 737–746. 10.1038/s41593-023-01566-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Hjort MM, Gowrishankar R, Tian L, Gordon-Fennell A, Namboodiri VMK, Bruchas MR, and Stuber GD. (2024). Microprisms enable enhanced throughput and resolution for longitudinal tracking of neuronal ensembles in deep brain structures. Neurophotonics 11, 033407. 10.1117/1.NPh.11.3.033407. [DOI] [Google Scholar]
- 69.Lau B, and Glimcher PW. (2007). Action and Outcome Encoding in the Primate Caudate Nucleus. J. Neurosci. 27, 14502–14514. 10.1523/JNEUROSCI.3060-07.2007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Histed MH, Pasupathy A, and Miller EK. (2009). Learning Substrates in the Primate Prefrontal Cortex and Striatum: Sustained Activity Related to Successful Actions. Neuron 63, 244–253. 10.1016/j.neuron.2009.06.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Kimchi EY, and Laubach M. (2009). Dynamic Encoding of Action Selection by the Medial Striatum. J. Neurosci. 29, 3148–3159. 10.1523/JNEUROSCI.5206-08.2009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Tai L-H, Lee AM, Benavidez N, Bonci A, and Wilbrecht L. (2012). Transient stimulation of distinct subpopulations of striatal neurons mimics changes in action value. Nat. Neurosci. 15, 1281–1289. 10.1038/nn.3188. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Matamales M, McGovern AE, Mi JD, Mazzone SB, Balleine BW, and Bertran-Gonzalez J. (2020). Local D2- to D1-neuron transmodulation updates goal-directed learning in the striatum. Science 367, 549–555. 10.1126/science.aaz5751. [DOI] [PubMed] [Google Scholar]
- 74.Martinez MC, Zold CL, Coletti MA, Murer MG, and Belluscio MA. Dorsal striatum coding for the timely execution of action sequences. eLife 11, e74929. 10.7554/eLife.74929. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Farahbakhsh ZZ, Song K, Branthwaite HE, Erickson KR, Mukerjee S, Nolan SO, and Siciliano CA. (2023). Systemic kappa opioid receptor antagonism accelerates reinforcement learning via augmentation of novelty processing in male mice. Neuropsychopharmacology 48, 857–868. 10.1038/s41386-023-01547-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Yang R, Tuan RRL, Hwang F-J, Bloodgood DW, Kong D, and Ding JB. (2023). Dichotomous regulation of striatal plasticity by dynorphin. Mol. Psychiatry 28, 434–447. 10.1038/s41380-022-01885-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Drago J, Gerfen CR, Lachowicz JE, Steiner H, Hollon TR, Love PE, Ooi GT, Grinberg A, Lee EJ, and Huang SP. (1994). Altered striatal function in a mutant mouse lacking D1A dopamine receptors. Proc. Natl. Acad. Sci. 91, 12564–12568. 10.1073/pnas.91.26.12564. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Perreault ML, Hasbi A, Alijaniaram M, Fan T, Varghese G, Fletcher PJ, Seeman P, O’Dowd BF, and George SR. (2010). The Dopamine D1-D2 Receptor Heteromer Localizes in Dynorphin/Enkephalin Neurons. J. Biol. Chem. 285, 36625–36634. 10.1074/jbc.M110.159954. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Kim J, Zhang X, Muralidhar S, LeBlanc SA, and Tonegawa S. (2017). Basolateral to Central Amygdala Neural Circuits for Appetitive Behaviors. Neuron 93, 1464–1479.e5. 10.1016/j.neuron.2017.02.034. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Wang H, Flores RJ, Yarur HE, Limoges A, Bravo-Rivera H, Casello SM, Loomba N, Enriquez-Traba J, Arenivar M, Wang Q, et al. (2024). Prefrontal cortical dynorphin peptidergic transmission constrains threat-driven behavioral and network states. Preprint at bioRxiv, 10.1101/2024.01.08.574700 https://doi.org/10.1101/2024.01.08.574700. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Bals-Kubik R, Ableitner A, Herz A, and Shippenberg TS. (1993). Neuroanatomical sites mediating the motivational effects of opioids as mapped by the conditioned place preference paradigm in rats. J. Pharmacol. Exp. Ther. 264, 489–495. [PubMed] [Google Scholar]
- 82.Abraham AD, Casello SM, Schattauer SS, Wong BA, Mizuno GO, Mahe K, Tian L, Land BB, and Chavkin C. (2021). Release of endogenous dynorphin opioids in the prefrontal cortex disrupts cognition. Neuropsychopharmacology 46, 2330–2339. 10.1038/s41386-021-01168-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Castro DC, Oswell CS, Zhang ET, Pedersen CE, Piantadosi SC, Rossi MA, Hunker AC, Guglin A, Morón JA, Zweifel LS, et al. (2021). An endogenous opioid circuit determines state-dependent reward consumption. Nature 598, 646–651. 10.1038/s41586-021-04013-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Crowley NA, Bloodgood DW, Hardaway JA, Kendra AM, McCall JG, Al-Hasani R, McCall NM, Yu W, Schools ZL, Krashes MJ, et al. (2016). Dynorphin Controls the Gain of an Amygdalar Anxiety Circuit. Cell Rep. 14, 2774–2783. 10.1016/j.celrep.2016.02.069. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Kelley AE, Domesick VB, and Nauta WJ. (1982). The amygdalostriatal projection in the rat--an anatomical study by anterograde and retrograde tracing methods. Neuroscience 7, 615–630. 10.1016/0306-4522(82)90067-7. [DOI] [PubMed] [Google Scholar]
- 86.Lee IB, Lee E, Han N-E, Slavuj M, Hwang JW, Lee A, Sun T, Jeong Y, Baik J-H, Park J-Y, et al. (2024). Persistent enhancement of basolateral amygdala-dorsomedial striatum synapses causes compulsive-like behaviors in mice. Nat. Commun. 15, 219. 10.1038/s41467-023-44322-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Corbit LH, Leung BK, and Balleine BW. (2013). The Role of the Amygdala-Striatal Pathway in the Acquisition and Performance of Goal-Directed Instrumental Actions. J. Neurosci. 33, 17682–17690. 10.1523/JNEUROSCI.3271-13.2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Namburi P, Al-Hasani R, Calhoon GG, Bruchas MR, and Tye KM. (2016). Architectural Representation of Valence in the Limbic System. Neuropsychopharmacol. Off. Publ. Am. Coll. Neuropsychopharmacol. 41, 1697–1715. 10.1038/npp.2015.358. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Ma P, Chen P, Tilden EI, Aggarwal S, Oldenborg A, and Chen Y. (2024). Fast and slow: Recording neuromodulator dynamics across both transient and chronic time scales. Sci. Adv. 10, eadi0643. 10.1126/sciadv.adi0643. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Lodder B, Kamath T, Savenco E, Röring B, Siegel M, Chouinard JA, Lee SJ, Zagoren C, Rosen P, Hartman I, et al. (2025). Absolute measurement of fast and slow neuronal signals with fluorescence lifetime photometry at high temporal resolution. Neuron 113, 3554–3566.e7. 10.1016/j.neuron.2025.08.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Kennedy RT. (2013). Emerging trends in in vivo neurochemical monitoring by microdialysis. Curr. Opin. Chem. Biol. 17, 860–867. 10.1016/j.cbpa.2013.06.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Reed SJ, Lafferty CK, Mendoza JA, Yang AK, Davidson TJ, Grosenick L, Deisseroth K, and Britt JP. (2018). Coordinated Reductions in Excitatory Input to the Nucleus Accumbens Underlie Food Consumption. Neuron 99, 1260–1273.e4. 10.1016/j.neuron.2018.07.051. [DOI] [PubMed] [Google Scholar]
- 93.Iyer ES, Vitaro P, Wu S, Muir J, Tse YC, Cvetkovska V, and Bagot RC. (2025). Reward integration in prefrontal-cortical and ventral-hippocampal nucleus accumbens inputs cooperatively modulates engagement. Nat. Commun. 16, 3573. 10.1038/s41467-025-58858-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.Aggarwal A, Negrean A, Chen Y, Iyer R, Reep D, Liu A, Palutla A, Xie ME, MacLennan BJ, Hagihara KM, et al. (2025). Glutamate indicators with increased sensitivity and tailored deactivation rates. Preprint at bioRxiv, 10.1101/2025.03.20.643984 https://doi.org/10.1101/2025.03.20.643984. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Chen Y, Saulnier JL, Yellen G, and Sabatini B. (2014). A PKA activity sensor for quantitative analysis of endogenous GPCR signaling via 2-photon FRET-FLIM imaging. Front. Pharmacol. 5. 10.3389/fphar.2014.00056. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Ma L, Jongbloets BC, Xiong W-H, Melander JB, Qin M, Lameyer TJ, Harrison MF, Zemelman BV, Mao T, and Zhong H. (2018). A Highly Sensitive A-Kinase Activity Reporter for Imaging Neuromodulatory Events in Awake Mice. Neuron 99, 665–679.e5. 10.1016/j.neuron.2018.07.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Zhang SX, Kim A, Madara JC, Zhu PK, Christenson LF, Lutas A, Kalugin PN, Jin Y, Pal A, Tian L, et al. (2023). Competition between stochastic neuropeptide signals calibrates the rate of satiation. Preprint at bioRxiv, 10.1101/2023.07.11.548551 https://doi.org/10.1101/2023.07.11.548551. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Gerfen CR, Engber TM, Mahan LC, Susel Z, Chase TN, Monsma FJ, and Sibley DR. (1990). D1 and D2 dopamine receptor-regulated gene expression of striatonigral and striatopallidal neurons. Science 250, 1429–1432. 10.1126/science.2147780. [DOI] [PubMed] [Google Scholar]
- 99.Siuda ER, Copits BA, Schmidt MJ, Baird MA, Al-Hasani R, Planer WJ, Funderburk SC, McCall JG, Gereau RW, and Bruchas MR. (2015). Spatiotemporal control of opioid signaling and behavior. Neuron 86, 923–935. 10.1016/j.neuron.2015.03.066. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Morikawa K, Furuhashi K, de Sena-Tomas C, Garcia-Garcia AL, Bekdash R, Klein AD, Gallerani N, Yamamoto HE, Park S-HE, Collins GS, et al. (2020). Photoactivatable Cre recombinase 3.0 for in vivo mouse applications. Nat. Commun. 11, 2141. 10.1038/s41467-020-16030-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101.Atwood BK, Kupferschmidt DA, and Lovinger DM. (2014). Opioids induce dissociable forms of long-term depression of excitatory inputs to the dorsal striatum. Nat. Neurosci. 17, 540–548. 10.1038/nn.3652. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Cai X, Huang H, Kuzirian MS, Snyder LM, Matsushita M, Lee MC, Ferguson C, Homanics GE, Barth AL, and Ross SE. (2016). Generation of a KOR-Cre Knockin Mouse Strain to Study Cells Involved in Kappa Opioid Signaling. Genes. N. Y. N 2000 54, 29–37. 10.1002/dvg.22910. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Krashes MJ, Shah BP, Madara JC, Olson DP, Strochlic DE, Garfield AS, Vong L, Pei H, Watabe-Uchida M, Uchida N, et al. (2014). A Novel Excitatory Paraventricular Nucleus to AgRP Neuron Circuit that Drives Hunger. Nature 507, 238–242. 10.1038/nature12956. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Yang R, Tuan RRL, Hwang F-J, Bloodgood DW, Kong D, and Ding JB. (2023). Dichotomous Regulation of Striatal Plasticity by Dynorphin. Mol. Psychiatry 28, 434–447. 10.1038/s41380-022-01885-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 105.Ehrich JM, Messinger DI, Knakal CR, Kuhar JR, Schattauer SS, Bruchas MR, Zweifel LS, Kieffer BL, Phillips PEM, and Chavkin C. (2015). Kappa Opioid Receptor-Induced Aversion Requires p38 MAPK Activation in VTA Dopamine Neurons. J. Neurosci. 35, 12917–12931. 10.1523/JNEUROSCI.2444-15.2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Harris JA, Hirokawa KE, Sorensen SA, Gu H, Mills M, Ng LL, Bohn P, Mortrud M, Ouellette B, Kidney J, et al. (2014). Anatomical characterization of Cre driver mice for neural circuit mapping and manipulation. Front. Neural Circuits 8. 10.3389/fncir.2014.00076. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Parker KE, Pedersen CE, Gomez AM, Spangler SM, Walicki MC, Feng SY, Stewart SL, Otis JM, Al-Hasani R, McCall JG, et al. (2019). A Paranigral VTA Nociceptin Circuit that Constrains Motivation for Reward. Cell 178, 653–671.e19. 10.1016/j.cell.2019.06.034. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Al-Hasani R, Gowrishankar R, Schmitz GP, Pedersen CE, Marcus DJ, Shirley SE, Hobbs TE, Elerding AJ, Renaud SJ, Jing M, et al. (2021). Ventral tegmental area GABAergic inhibition of cholinergic interneurons in the ventral nucleus accumbens shell promotes reward reinforcement. Nat. Neurosci. 24, 1414–1428. 10.1038/s41593-021-00898-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109.Yang W, Carrillo-Reid L, Bando Y, Peterka DS, and Yuste R. (2018). Simultaneous two-photon imaging and two-photon optogenetics of cortical circuits in three dimensions. eLife 7, e32671. 10.7554/eLife.32671. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Piantadosi SC, Zhou ZC, Pizzano C, Pedersen CE, Nguyen TK, Thai S, Stuber GD, and Bruchas MR. (2024). Holographic stimulation of opposing amygdala ensembles bidirectionally modulates valence-specific behavior via mutual inhibition. Neuron 112, 593–610.e5. 10.1016/j.neuron.2023.11.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 111.Koos T, Tepper JM, and Wilson CJ. (2004). Comparison of IPSCs evoked by spiny and fast-spiking neurons in the neostriatum. J. Neurosci. Off. J. Soc. Neurosci. 24, 7916–7922. 10.1523/JNEUROSCI.2163-04.2004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.Planert H, Szydlowski SN, Hjorth JJJ, Grillner S, and Silberberg G. (2010). Dynamics of synaptic transmission between fast-spiking interneurons and striatal projection neurons of the direct and indirect pathways. J. Neurosci. Off. J. Soc. Neurosci. 30, 3499–3507. 10.1523/JNEUROSCI.5139-09.2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113.Kondabolu K, Doig NM, Ayeko O, Khan B, Torres A, Calvigioni D, Meletis K, Koós T, and Magill PJ. (2023). A Selective Projection from the Subthalamic Nucleus to Parvalbumin-Expressing Interneurons of the Striatum. eNeuro 10. 10.1523/ENEURO.0417-21.2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Kawaguchi Y. (1993). Physiological, morphological, and histochemical characterization of three classes of interneurons in rat neostriatum. J. Neurosci. Off. J. Soc. Neurosci. 13, 4908–4923. 10.1523/JNEUROSCI.13-11-04908.1993. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Bennett BD, and Wilson CJ. (1998). Synaptic regulation of action potential timing in neostriatal cholinergic interneurons. J. Neurosci. Off. J. Soc. Neurosci. 18, 8539–8549. 10.1523/JNEUROSCI.18-20-08539.1998. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 116.Oswald MJ, Oorschot DE, Schulz JM, Lipski J, and Reynolds JNJ. (2009). IH current generates the afterhyperpolarisation following activation of subthreshold cortical synaptic inputs to striatal cholinergic interneurons. J. Physiol. 587, 5879–5897. 10.1113/jphysiol.2009.177600. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 117.Simpson EH, Akam T, Patriarchi T, Blanco-Pozo M, Burgeno LM, Mohebi A, Cragg SJ, and Walton ME. (2024). Lights, fiber, action! A primer on in vivo fiber photometry. Neuron 112, 718–739. 10.1016/j.neuron.2023.11.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 118.Pachitariu M, Stringer C, Dipoppa M, Schröder S, Rossi LF, Dalgleish H, Carandini M, and Harris KD. (2017). Suite2p: beyond 10,000 neurons with standard two-photon microscopy. Preprint at bioRxiv, 10.1101/061507 https://doi.org/10.1101/061507. [DOI] [Google Scholar]
- 119.Stringer C, Ki C, DelGrosso N, LaFosse P, Zhang Q, and Pachitariu M. (2026). Extracting large-scale neural activity with Suite2p. Preprint at bioRxiv, 10.64898/2026.02.04.703741 https://doi.org/10.64898/2026.02.04.703741. [DOI] [Google Scholar]
- 120.Loshchilov I, and Hutter F. (2019). Decoupled Weight Decay Regularization. Preprint at arXiv, 10.48550/arXiv.1711.05101 https://doi.org/10.48550/arXiv.1711.05101. [DOI] [Google Scholar]
- 121.Terven J, Cordova-Esparza D-M, Romero-González J-A, Ramírez-Pedraza A, and Chávez-Urbiola EA. (2025). A comprehensive survey of loss functions and metrics in deep learning. Artif. Intell. Rev. 58, 195. 10.1007/s10462-025-11198-7. [DOI] [Google Scholar]
- 122.Srivastava N, Hinton G, Krizhevsky A, Sutskever I, and Salakhutdinov R. (2014). Dropout: A Simple Way to Prevent Neural Networks from Overfitting. J. Mach. Learn. Res. 15, 1929–1958. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data reported in this paper will be shared by the lead contact upon request.
All original code has been deposited at Zenodo at https://zenodo.org/records/21286480
and is publicly available as of the date of publication.
Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.
