Abstract
Successful cooperation requires dynamic integration of social cues. However, the neural mechanisms supporting this complex process remain unknown. Here, we reveal that the primate dorsomedial prefrontal cortex (dmPFC) implements a gaze-dependent social evidence accumulation process to guide cooperative decisions in freely moving marmoset dyads. A drift-diffusion process in which the partner’s action variability is accumulated by social gaze best explains the cooperative actions of the actor. Single-neuron recordings in dmPFC revealed a direct neural correlate: the slope of predictive ramping activity mapped directly onto the rate of evidence accumulation, while baseline firing, modulated by prior outcomes, mapped onto the initial bias. At the population level, the geometry of dmPFC neural trajectories reflected the strength of social evidence and was linked to cooperative success. Together, these findings establish a multi-level neural mechanism for transforming active sensing into a decision variable, linking a canonical computation to cooperative behavior in a naturalistic setting.
Introduction
Cooperation is a defining feature of advanced social behavior. To cooperate successfully, individuals must monitor the actions of their partners, infer intentions, and decide whether to act in a coordinated manner (1). A central challenge lies in transforming social cues—particularly through gaze, which reveals attention and intention (2, 3)—into decision variables that shape action selection.
Evidence accumulation frameworks provide a powerful explanation for how the brain integrates information to form a decision (4, 5). Such mechanisms are well-established in domains such as value-based choice (6–8) and, most notably, perceptual decision-making (5, 9). In perceptual tasks, neural correlates of evidence accumulation have been identified in regions like the dorsolateral prefrontal cortex (dlPFC) (10) and the lateral intraparietal area (LIP) (11), where the neural activity encodes the strength of incoming sensory evidence. However, two fundamental challenges remain. First, it is unclear whether evidence accumulation is a core, conserved computational mechanism that extends to the domain of complex social cognition. Second, because prior studies have relied on highly constrained, head-restrained task structures, it is unknown if these accumulation processes operate during naturalistic, unconstrained social interactions.
To address these challenges, we designed a cooperation task where freely moving pairs of marmosets coordinate their actions for mutual reward (12), allowing us to investigate the process of social decision-making in a naturalistic setting. While recent work has identified the use of gaze-dependent strategies during these cooperative interactions (13), the precise manner in which gaze is used to integrate available social information and the neural mechanisms underlying this integration process remain unknown. We hypothesized that during these naturalistic social interactions, the animals self-organize their behavior into discrete, self-paced episodes (henceforth “trials”), each consisting of a period of gaze-dependent social evidence accumulation culminating in a cooperative decision. Conceptually, we defined each trial as a complete behavioral epoch, beginning with the initiation of movement leading to a pull action and ending after the outcome of that action was delivered and evaluated. To investigate the neural underpinnings of this process, we recorded from the dorsomedial prefrontal cortex (dmPFC), a key node in the social brain network (14–16). Given the dmPFC’s involvement in encoding others’ actions and outcomes (17), we hypothesized that it plays a key role in the accumulation of social evidence. By combining behavioral modeling with single-neuron and population-level analyses, we provide the first evidence demonstrating that the dmPFC implements a social gaze-dependent evidence accumulation process to transform social information into decision variables guiding cooperation (Fig. 1A).
Figure 1. Conceptual hypothesis, behavioral metrics, and drift diffusion model.
(A) Conceptual hypothesis. In the mutual cooperation (MC) task, marmoset dyads are required to pull their levers within a 1-second time window to receive a juice reward for both. We hypothesize that the social gaze of the actor animal collects social evidence about a partner’s cooperative intention, which is accumulated by the dorsomedial prefrontal cortex to guide action decisions. (B) Top: Definition of eight continuous behavioral variables. Bottom: Example traces for one animal in one session. (C) Distribution of inter-pull intervals. (D) Definition of the social gaze accumulation as the area under curve (AUC) of the social gaze probability. Dashed rectangle shows 4 s functional window of analysis. Pink shaded areas highlight social gaze duration (social gaze probability> 0). (E) Definition of the trial start and response time (RT). (F) Definition of the holistic representation of the partner’s actions via principal component analysis (PCA). (G) Structure of the drift diffusion model. (H) Coefficient estimates from the main model (M1) showing that partner action (PC1) contributes only when social gaze is present. (I) Definition of the social evidence. Wilcoxon signed-rank test for panel H (* p < 0.05).
Results
A quantitative framework for studying naturalistic cooperation
To quantify the animals’ behavior in the mutual cooperation task, we first used markerless facial tracking to extract eight continuous behavior variables for each marmoset (18, 19), capturing gaze direction, face position, and movement dynamics (Fig. 1B; Supp. Mov. S1). Analysis of the intervals between consecutive pulls by each animal revealed a rhythmic pattern, supporting the characterization of the interactions as a series of self-paced trials (Fig. 1C). For both animals from which we recorded the neural activity (hereafter, the ‘actors’), the majority of these inter-pull intervals were longer than 4 seconds, a timescale that informed the functional window for our subsequent analyses. Within each trial, we defined social gaze as the actor looking at the partner’s face, and social gaze accumulation as the area under the curve of the social gaze probability in this functional window preceding a pull action (Fig. 1D).
To objectively define the start of each trial, we sought the single most powerful and unique behavioral predictor of a future pull. Using a Generalized Linear Model (GLM), we found that the linear speed (LS) of the actor’s face was a consistently high-importance predictor in both animals (Supp. Fig. S1A). Furthermore, a Variance Inflation Factor (VIF) analysis confirmed that LS had low collinearity with the other variables, indicating it provided unique predictive information (Supp. Fig. S1B). Based on this evidence, we used the ramping onset of face linear speed to define the start of each trial, and operationalized the duration from this starting point to the pull action as the response time (RT) for that trial (Fig. 1E).
With this quantitative framework in place, we hypothesized that the actor animal utilizes a holistic representation of the partner’s actions that would collectively allow it to infer the partner’s cooperative intention. To define such a representation, we performed Principal Component Analysis (PCA) on the partner’s eight behavioral variables. The first principal component (PC1) captured a large fraction of the behavioral variance and thus served as a summary of the partner’s actions (Fig. 1F). This pipeline thus allowed us to deconstruct continuous behavior into a structured, trial-based format suitable for studying the social evidence accumulation process.
Gaze-dependent drift-diffusion model explains cooperative decisions
To formalize the decision-making process underlying cooperation, we fit the animals’ response times and choices using a drift-diffusion model (DDM) (20), which decomposes behavior into several key parameters: drift rate (, the speed of accumulation), the starting point (, the initial bias), and non-decision time (Fig. 1G; see Methods). Critically, we formulated four competing models to explicitly test how social gaze mediates the integration of partner information (Supp. Fig. S2A). Model comparison using the Deviance Information Criterion (DIC) revealed that a model which allows partner evidence to be integrated only during period of social gaze (M1; Fig. 1H) provided the best fit compared to alternate models assuming either no gaze modulation (M2), or an additive (M3), or multiplicative (M4) effect of gaze accumulation (Supp. Fig. S2B). This result supports a gaze-dependent integration mechanism, where information about the partner’s actions is specifically sampled during periods of social attention.
Analysis of the coefficients from the best fitting model (M1) provided crucial insight into the nature of the social evidence being accumulated. We found that the coefficient for the variability of the partner’s actions was consistently and significantly negative in both animals (Fig. 1H). This indicates that the variability of the partner’s action (PC1) was the major social information being accumulated, where a noisier or more ambiguous action (higher standard deviation) leads to a smaller drift rate, slowing the decision process. A control model using a single partner variable (partner’s linear speed, LS), instead of the holistic PC1, did not perform better and its key coefficients were not significant, demonstrating that the decision process relies on a rich, holistic representation of the partner’s actions (Supp. Fig. S2C–D). We therefore defined social evidence as a measure of variability in partner’s action (PC1), with lower variability (less noise) corresponding to stronger evidence (Fig. 1I; see Methods).
dmPFC neurons show predictive, mixed selectivity for social variables
To investigate the neural basis of this social gaze-dependent process, we recorded single-unit activity from the dmPFC of the actor marmosets using a custom-built, wireless recording system (Fig. 2A; Supp. Mov. S1). We implemented a GLM (Fig. 2B; see Methods) to formally characterize the diversity of observed neural responses (Supp. Fig. S3A). Many dmPFC neurons exhibited mixed selectivity, including a core subpopulation of “triple-encoding” neurons simultaneously represented all three key task variables: the actor’s pull, the actor’s gaze, and the partner’s actions (Fig. 2C).
Figure 2. dmPFC neurons show mixed selectivity.
dmPFC neurons integrate self, partner, pull and gaze variables, and encode both predictive and reactive components of the actions. (A) Wireless recording system design: (A1) illustration of the headcap design; (A2) example video frame highlighting the wireless recording apparatus on one of the animals during task performance; (A3) estimation of the dmPFC recording sites; (A4) example of action potentials (black ticks) from the 64-channel probe aligned at a pull action. (B) GLM pipeline for defining tuning properties and encoding types. (C) Distribution of selectivity across neurons reveals mixed coding for pull, gaze, and partner action variables, supporting flexible social computations. (D) Distribution of encoding categories for pull and gaze neurons. R, reactive; P, predictive. (E) Subset of “triple-encoding” neurons’ coding categories for pull and gaze actions. (F) Predictive encoding is enriched among triple-encoding neurons. Chi-squared test for panel C and Fisher’s exact test for panel F (* p < 0.05, ** p <0.01, *** p< 0.001).
We next investigated the temporal dynamics of this neural encoding, refining our GLM to classify tuning as either “predictive” (neural spikes preceding behaviors) or “reactive” (neural spikes reacting to behaviors) (Fig. 2B; see Methods). We found distinct subpopulations with purely predictive or reactive profiles, as well as many that exhibited both (Fig. 2D–E; Supp. Fig. S3B). Critically, we reasoned that the triple-encoding neurons, encoding the most comprehensive information, would be primarily involved in integrating social information, thereby leading to the actor’s own actions. A statistical comparison confirmed this intuition: the triple-encoding subpopulation was significantly enriched for predictive encoding of pull and gaze actions relative to the broader population of task-responsive neurons (Fig. 2F). Together, these findings show that dmPFC neurons integrate social information and self-actions, with the most informative signals emerging before cooperative actions rather than in response to them.
Neural ramping activity reflects evidence accumulation
Having established that dmPFC neurons show predictive, mixed selectivity, we next asked whether their activity exhibits the dynamic signatures of the evidence accumulation process itself. A key feature of a neural accumulator is ramping activity, wherein the slope is inversely related to RT. We found precisely this relationship in the dmPFC population: the average slope of the pre-pull firing rate was steeper for faster responses and shallower for slower ones (Fig. 3A). Critically, these firing rate slopes scaled with the strength of social evidence (Fig. 3B). A significant subpopulation of neurons conjunctively encoded both RT and social evidence in their firing rate slopes (Fig. 3C). Crucially, this encoding of social evidence by differential ramping was strongly linked to behavioral outcome: this relationship was robust on successful trials but significantly weaker on failed trials (Fig. 3D). We further found a strong correspondence between this dynamic slope-based encoding and the static tuning properties of the neurons (Supp. Fig. S4A). Moreover, this slope-encoding population showed a trend of being enriched for predictive tuning, consistent with an anticipatory, rather than reactive, process (Supp. Fig. S4B).
Figure 3. Neural firing rate slopes encode both response time and social evidence.
(A) Mean normalized firing rate of all dmPFC neurons aligned at the pull action and separating into three quantiles of the RTs (left), and correlations between dmPFC firing rate slopes and RT (right). (B) Mean normalized firing rate of dmPFC neurons aligned at the trial start and separating into three quantiles of the social evidence (left), and correlations between dmPFC firing rate slopes and RT social evidence (right). (C) Significant proportions of neurons encode both factors simultaneously. (D) Neuron numbers with and without significant correlations between firing rate slopes and social evidence in successful vs failed pulls. (E) Neural DDM linking firing rate slope to drift rate and baseline firing to starting point . (F) The difference in the mean baseline firing rates following successful pulls vs. failed pulls. (G) Behavioral DDM model incorporating outcome-dependent starting points. Wilcoxon signed-rank test for panels A, B, E, F and G, Chi-squared test for panels C and E, and McNemar test for panel D (* p < 0.05, ** p <0.01, *** p< 0.001). For the red datapoints, Pearson’s correlation test for panels A and B, 95% Credible Interval (CI) of the posterior distributions for panel E, and Mann-Whitney U Test for panel F.
Neural activity maps directly onto computational model parameters
To formally connect the observed neural dynamics to the computational model, we constructed a neural DDM that incorporated both the firing rate slopes and baseline firing rates as predictors. The results revealed a direct mapping: trial-by-trial firing rate slopes were a significant positive predictor of the DDM’s drift rate, while baseline firing rates quantified during the non-decision time predicted the starting point (Fig. 3E). This starting point bias was itself influenced by recent history, in which a failed prior trial was associated with higher baseline firing (Fig. 3F). To test the hypothesis suggested by this neural result—that post-error changes in baseline activity have a direct behavioral consequence—we returned to our best behavioral model (M1) and added the previous trial’s outcome (1 for success pulls and 0 for failed pull) as a predictor of the starting point. This updated model (M1+) confirmed the hypothesis, revealing a significant negative coefficient for this relationship (Fig. 3G). This finding provides a direct behavioral correlate for the neural activity, showing that the outcome of the prior trial systematically biased the starting point of the current decision process, potentially reflecting a post-error increase in motivation. These analyses provide a direct bridge between the firing rate and computational levels, demonstrating that the dynamic and static components of dmPFC activity map onto the core parameters of the evidence accumulation process.
Population dynamics link neural processing to cooperative success
While single-neuron activity correlates with key decision variables, cognitive functions are ultimately implemented by the coordinated, dynamic activity of large neural populations. Therefore, to understand how the dmPFC as a whole processes social evidence, we examined the collective neural activity by analyzing population trajectories at the level of single trials (Fig. 4A). We found that the geometry of these neural population trajectories, specifically their path variability, defined as the local tortuosity of the trajectory (21), was significantly negatively correlated with the level of social evidence (Fig. 4B,C). This effect was most prominent early in the trial during the initial phase of the decision process (Fig. 4B,E). Crucially, this relationship between the geometry of the neural population and the strength of social evidence was also tightly linked to behavioral outcome. At the population level, the strong relationship between path variability and social evidence was present only on successful trials and was absent on failed trials (Fig. 4D). This finding at the population level mirrored our observations with single neurons, where the encoding of social evidence by firing rate slopes was robust on successful trials and significantly weaker on failed trials (Fig. 3D). Together, these results demonstrate that the effective neural processing of social cues in dmPFC is critical for guiding successful cooperation.
Figure 4. Population dynamics reflect social gaze accumulation.
(A) Definition of path variability (tortuosity) of the single-trial neural population trajectory, example session. (B) Heatmaps of the temporal structure of the path variability showing that social evidence shapes dmPFC population dynamics during social decisions, averaged across sessions. (C) Correlation between the strength of social evidence and path variability of the neural trajectory across sessions. (D) Correlation between social evidence and path variability of the neural trajectory in successful and failed pulls. (E) Average change of path variability over time in individual trials. Wilcoxon signed-rank test for panels C and D (* p < 0.05, ** p <0.01), and paired t-test with cluster-based correction for panel E (red segments p < 0.05).
Discussion
Our study provides a multi-level account of how the primate dmPFC orchestrates cooperation in a naturalistic behavioral setting. We show that the dmPFC implements a social gaze-dependent evidence accumulation process, transforming actively sampled social cues into a decision variable. By identifying this canonical computation in freely moving animals without any imposed task structure, our work bridges the gap between classic decision neuroscience and the complexity of real-world social interactions.
Our findings connect behavior, computation, and neural activity across multiple scales. A gaze-dependent DDM best explained cooperative decisions, establishing that evidence about a partner’s actions is integrated only when actively sampled. We then identified a direct neural correlate for this process. At the single-neuron level, the predictive ramping of dmPFC neurons encoded both the response time and the strength of social evidence. These neural dynamics mapped directly onto the DDM’s core parameters, with firing rate slopes predicting the drift rate and baseline activity—updated by recent trial outcomes—predicting the starting point bias. Finally, this mechanism was reflected at the population level, where the geometry of neural trajectories tracked the level of social evidence. Crucially, this signature was present mostly on successful trials, cementing the link between the neural mechanism and the behavioral outcome.
Evidence accumulation has long been established as a canonical computation for perceptual decisions (5), with clear neural signatures in sensorimotor areas such as LIP (11) and dlPFC (10). Our findings extend this framework into the social domain, identifying the dmPFC as a key locus for this more complex form of accumulation. Unlike perceptual tasks, where evidence is often externally driven, social evidence about another’s intentions is uncertain and must be actively sampled via gaze (2). The dmPFC is uniquely suited for this computation (22), integrating information about the partner’s actions with the actor’s own social attention.
A significant advance of our study is the application of a detailed, mechanistic approach to neural function in freely moving, socially interacting primates. While the field has begun to embrace the study of more naturalistic behaviors (23), as exemplified by recent work in rodents (24), tree shrews (25) and macaques (26), research on the detailed computational mechanism has been lacking, facing both technical hurdles in recording from unconstrained animals and analytical challenges in interpreting complex, continuous behavior. We addressed the technical challenges by developing a novel modular head chamber (27), which improved surgical success rates. We also used movable probes, in contrast to fixed chronic implants (28), to sample diverse neural populations across sessions. To overcome the analytical challenges, we employed a sophisticated computational framework, including the GLM pipelines (29–32), to deconstruct the continuous behavioral stream into meaningful, quantifiable variables. By identifying computational signatures of decision-making in this unconstrained paradigm, we show that even without the rigid structure of a classic task, marmosets spontaneously organize their behavior in a way that is amenable to an evidence accumulation process. This demonstrates the robustness of this neural algorithm, suggesting it is a fundamental process that is flexibly deployed to make decisions in dynamic environments.
Our results also highlight several avenues for future research. The influence of the previous trial’s outcome on the DDM starting point, mediated by baseline dmPFC firing, suggests a rapid, experience-dependent plasticity that could be interpreted as an alerting signal or a shift in motivation following failure (33–35). Future studies employing neural manipulation could also directly test the causal role of this mechanism; for instance, microstimulation delivered during the gaze accumulation period could potentially disrupt the evidence integration process and alter the cooperative decision (36).
In conclusion, by linking behavior, modeling, and neural dynamics across multiple scales, we established a comprehensive framework for how the brain transforms active sensing into a decision variable to guide cooperation. This neurocomputational framework provides a mechanistic grounding for economic models of strategic interaction (37), a potential biomarker for social deficits in psychiatry (38), and offers new biological insights for designing the next generation of cooperative multi-agent artificial systems (39).
Materials and Methods
Animal subjects
Two common marmosets (Callithrix jacchus) participated in the study: a 7-year-old female (Animal K) and a 6.5-year-old male (Animal D). Animal K participated in 21 sessions with four different partners (2 females and 2 males), and Animal D participated in 29 sessions with three different partners (3 females). A detailed comparison of behavior and neural activity across different partners is beyond the scope of this study. However, to confirm that our findings were not confounded by partner identity, we performed alternative analyses using linear mixed-effects models, which included partner identity as a random effect. These confirmatory models did not change the statistical conclusions of the primary analyses reported throughout the manuscript. The animals were pair-housed, maintained on a 12-hour light-dark cycle, and food and water were removed from their home cages for 1–3 hours prior to testing. All experimental procedures were approved by the Yale Institutional Animal Care and Use Committee and complied with the National Institutes of Health Guide for the Care and Use of Laboratory Animals.
Experimental task
The experimental apparatus and training procedures were identical to those described in previous studies (12, 13). The animals had been previously trained to perform a mutual cooperation (MC) task. Both marmosets in the dyads were required to pull their respective levers within a 1-second time window to successfully receive a reward. A pull was registered when the lever passed a predefined position threshold, and its timestamp was recorded by the task-control computer at a 10 kHz sampling rate. A successful, coordinated trial resulted in the delivery of a 0.2 mL liquid reward (marshmallow fluff solution) to both animals, which was immediately preceded by an auditory tone (2 KHz) to signal success. A failed trial occurred if the pulls were not synchronized within the 1-second window. This task was designed to be performed in a naturalistic setting, with no explicitly defined or implemented temporal structure; instead, the animals engaged in a continuous series of self-paced cooperative decisions. The data analyzed in this study were collected after the animals had been fully trained on this task, representing established cooperative performance.
Behavioral Data Acquisition and Processing
Continuous behavior was recorded at 30 frames/sec from three synchronized cameras (GoPro Hero10) positioned at the front of the apparatus. The cameras were controlled via a shared remote control. To align the video data with the event timestamps from the task-control computer, we used two redundant signals, both controlled by the task computer: an audible tone at the start of the session for coarse alignment, and an LED that flashed every 20 seconds for fine alignment, providing a reliable and recurring synchronization signal captured by the cameras. We used the multi-animal version of DeepLabCut 2.0 to perform markerless pose estimation (19, 40). Six facial key points were labeled on 500 frames: two ears, two eyes, the central blaze, and the mouth. The DeepLabCut neural network was trained on these labels for 150,000 iterations until the tracking error reached a stable minimum. The 2D coordinates from the three cameras were then used to triangulate the 3D position of each key point using Anipose (18). This 3D reconstruction allowed for the precise calculation of each animal’s head position and orientation in space throughout the session. From the 3D facial coordinates, we derived the eight continuous behavioral variables used for all subsequent analyses (Fig. 1B):
: Three angular variables measuring the direction of gaze relative to the other animal, the reward tube, and the lever.
, , : Three distance variables measuring the distance to the other animal, the reward tube, and the lever.
: The face linear speed, calculated from the movement of the head’s 3D position.
: The gaze vector angular speed, capturing the rate of change in head orientation.
Similar to earlier work (13), head gaze was defined as a 15-degree cone extending perpendicularly from the facial plane formed by the tracked key points. Social gaze was determined to have occurred whenever one animal’s gaze cone intersected with the other animal’s face.
Implant Design and Surgical Processing
We designed a novel, multi-component head implant to facilitate wireless neural recordings in freely moving marmosets (Fig. 2A). The system was custom-designed in CAD software (Shapr3D) to fit each animal’s cranium (obtained from pre-operative CT scan) and consists of two main parts. The first is a 3D-printed medical-grade titanium base (Materialise NV) that protects the surgical site where the skull is exposed. An integrated headpost on the base allows for brief periods of stable head-fixation. The second part is a 3D-printed recording chamber (nylon; Protolabs), which consists of two sealed and separated chambers. The probe chamber houses and protects a miniature screw microdrive (nanodrive; Cambridge Neurotechnologies) and the probe, while the electronics chamber protects the headstage chips needed for digitization and signal amplification. A silicone-based sealant (Kwik-Sil; World Precision Instruments, LLC) separates the chambers, allowing the probe to remain sterile while the data logger is connected to the headstage electronics. This modular design is advantageous for animal welfare, as it permits a multi-step surgical approach with shorter individual procedures and simplifies implant cleaning and maintenance.
Enabled by this modular implant design, the surgical implantation followed a multi-step process. First, the titanium base was implanted. In a subsequent procedure, a small opening was made in the skull within the protected base, and a nanodrive holding a 64-channel linear array probe (NeuroNexus) was mounted inside the probe chamber. The nanodrive allowed movement of the probe along the superior-inferior axis, while the mounting apparatus holding the nanodrive allowed repositioning along the anterior-posterior and mediolateral axes. The target cortical area, the dorsomedial prefrontal cortex (dmPFC), was identified using pre-operative CT scans and stereotaxic coordinates to guide the initial placement. The probe was lowered to a site within the dmPFC during the surgery, and the probe track location was confirmed post-operatively by co-registering a CT scan with the probe implanted against a digital marmoset brain atlas (41).
Neural Recording and Signal Processing
During each experimental session, which lasted approximately 15 minutes, we used a wireless eCube headstage system (White Matter LLC) to record single-unit activity while animals were freely moving and performing the Mutual Cooperation (MC) task (Animal K: 490 neurons, Animal D: 464 neurons). The data logger was attached at the start and removed at the end of each session while the animals were briefly head-fixed, allowing for untethered social interaction during the task. Synchronization between the wireless recordings and the task-control computer was managed by a DAQ board (National Instruments Corp.) connected to the logger’s transceiver. At the start of each recording, the logger sent a pulse that was captured by the NI board and timestamped by the task computer for initial alignment. To correct for any temporal drift, the task computer also sent a continuous 10 kHz signal through the NI board to the logger for auto-correction. Electrophysiological signals were acquired at a 20 kHz sampling rate. The raw data was processed offline: action potentials were first automatically detected and clustered using Kilosort4 (42), and the resulting clusters were then manually curated using phy, an open-source Python library, to isolate well-defined single units for all subsequent analyses.
Following spike-sorting, all subsequent analyses were performed on neural data that was first downsampled from 20 kHz to 30 Hz to match the camera’s frame rate. Continuous firing rates were estimated by counting the number of spikes in 33.3 ms (1/30 s) bins and then smoothing the resulting histogram with a Gaussian kernel. This continuous firing rate trace was then z-scored using the mean and standard deviation of that neuron’s activity across the entire session to obtain a normalized firing rate. For the neural DDM, the trial-by-trial slope of firing rate was calculated by fitting a linear regression model to the normalized (z-scored) firing rate trace over the trial’s duration. For the neural GLM (see below), the model was fit directly to the 30 Hz spike count train for each neuron.
Computational Modeling of Decision-Making
To prepare the behavioral data for modeling, we first defined trial epochs and key predictor variables. Continuous behavioral data were segmented into discrete trials centered on pull actions. The start of each trial was marked by the onset of the actor’s face linear speed (LS), with response time (RT) defined as the duration until the subsequent pull. The justification for this marker is detailed in the supplementary materials (Supp. Fig S1). To create a holistic measure of the partner’s actions, we performed Principal Component Analysis (PCA) on their eight behavioral variables. These variables were then used to fit a hierarchical Bayesian drift-diffusion model (DDM) to formalize the evidence accumulation process. The DDM decomposes behavior into distinct cognitive parameters: the drift rate , representing the speed of evidence accumulation; the starting point , representing an initial bias; and the non-decision time . To explicitly test our hypothesis about how social gaze mediates the integration of partner information, we formulated and compared four competing models (Supp. Fig S2A):
- M1 (Gaze-gated model): A model in which partner information influences the drift rate only during periods of social gaze. Periods of social gaze were defined as the time points in the functional window when social gaze probability was greater than zero.
- M2 (No gaze modulation): A model where partner information is integrated regardless of gaze.
- M3 (Additive gaze effect): A model where social gaze (measured as social gaze accumulation, ) has a separate, additive effect on the drift rate.
- M4 (Multiplicative gaze effect): A model where social gaze multiplicatively scales the influence of partner information.
- M1+ (Gaze-gated model with history effect): An extension of M1 where the starting point is modulated by the outcome of the previous trial, which is 1 if the previous trial is successful, and 0 if it failed. The drift rate is formulated identically to M1.
- M5 (Gaze-gated model with single behavioral variable): A control model with the same gaze-gated structure as M1, but using the partner’s linear speed instead of the holistic action variable (PC1, ).
To link the cognitive model to neural activity, we constructed a neural DDM. This model used trial-by-trial neural features to predict the DDM parameters. The drift rate was predicted by the slope of a neuron’s firing rate, while the starting point was predicted by the neuron’s mean baseline firing rate. This baseline activity was calculated during the non-decision time period , using the duration estimated for each trial from the fitted parameters of model M1. The specific relationships were:
All DDM models were fit using the HDDM (Hierarchical Drift-Diffusion Model) Python package (20), which uses Markov Chain Monte Carlo (MCMC) sampling to estimate posterior distributions of the model parameters. For each model, we generated 2000 samples from the posterior distribution following a burn-in period of 1000 iterations, which were discarded. We confirmed model convergence by visually inspecting the MCMC chains. The final coefficient values reported are the mean of the posterior distributions from the collected samples.
Social Evidence Estimation
We defined the social evidence for a given trial by subtracting the standard deviation of the partner’s action (PC1) during periods of social gaze from the maximum value observed across all trials, a linear transformation where lower variability (less noise) corresponds to stronger evidence:
Statistical Modeling of Behavior and Neural Tuning
Behavioral GLM for Trial Start Identification
To objectively identify the behavioral variable that most reliably predicted the initiation of a pull action, we used a Generalized Linear Model (GLM) with a logistic link function, fit to the eight continuous behavioral variables prepared at a 30 Hz sampling rate. The model was fit separately for each behavioral session. Within each session, the data was split into an 80% training set and a 20% testing set. The model was trained to predict the probability of a pull event from the eight continuous behavioral variables. To capture temporal dynamics, each variable trace was convolved with a basis set of 10 temporal kernels spanning the 4-second window preceding each pull action (-4s to 0s). To determine the importance of each variable, we used a leave-one-out (LOO) procedure. The variable importance was quantified as the drop in the Area Under the Receiver Operating Characteristic curve (AUROC), calculated on the 20% testing set, when that variable was removed from the model. This entire procedure was repeated for 1000 bootstrap iterations to assess the statistical significance of the drop in AUROC for each variable.
Neural GLM for Single-Neuron Tuning Properties
We used a Poisson GLM to characterize the tuning of individual dmPFC neurons by fitting the model to the 30 Hz downsampled spike count train for each neuron. The model was fit separately for each neuron using an 80/20 train/test split. The model predicted the number of spikes in each 30 Hz time bin based on four main predictors: the actor’s pull action, social gaze, the partner’s action (PC1), and reward delivery. The reward delivery variable was included as a control to ensure that observed neural activity was not confounded by the motor action of drinking. Each of the four predictors was convolved with a basis set of 10 temporal kernels spanning an 8-second window around each event (-4s to +4s), creating a total of 40 weighted features for the full model.
To determine a neuron’s overall selectivity, we performed an LOO comparison where all 10 kernels for a given variable were excluded. A neuron was considered significantly tuned if the AUROC of the full model on the testing set was significantly better than that of the reduced model. Furthermore, to distinguish between predictive and reactive activity, we assessed them separately based on their temporal relationship to the event under consideration. This was achieved by splitting the 10 temporal kernels into a predictive set (spanning −4s to 0s pre-event) and a reactive set (spanning 0s to +4s post-event). For instance, to test for predictive tuning, we compared the full model to a reduced model where only the 5 predictive kernels for a given variable were excluded. A neuron was classified as having predictive tuning if the full model’s AUROC on the testing set was significantly better than that of this specific reduced model.
Population Dynamics
To examine the collective dynamics of the dmPFC population, we carried out single-trial population analyses based on existing methods (21). This analysis was restricted to sessions with more than 5 simultaneously recorded neurons (up to 40 neurons per session; 19 sessions in Animal K, 26 sessions in Animal D). For each of these sessions, the normalized (z-scored) firing rates of the population were arranged into an activity matrix. We then performed PCA on each session’s data separately and focused on the neural state space defined by the first three principal components (PCs).
Before calculating geometric measures, each 3D trajectory was preprocessed. First, because trials had different durations, each trajectory was normalized by being reparameterized by its arc length. This process resamples the trajectory so that each path consists of a uniform 100 time points, making them comparable. Second, these reparameterized trajectories were smoothed using a rolling Gaussian window to isolate the large-scale geometry of the path from small-scale neural jitter.
On these preprocessed trajectories, we calculated the following two local geometric measures (Fig. 4A):
- Path Direction (Curvature, ): Curvature measures how quickly the trajectory deviates from a straight line. It is formally defined as the magnitude of the rate of change of the curve’s unit tangent vector with respect to arc length .
- Path Variability (Tortuosity, ): Tortuosity captures the “wiggleness” or complexity of the trajectory. We defined it as the squared ratio of the rate of change of curvature with respect to arc length to the curvature itself.
These continuous measures were estimated numerically from the discrete, preprocessed trajectories. Derivatives were approximated using a rolling window of 5 time points, and the term was calculated using the gradient of the curvature trace. Both the final curvature and tortuosity traces were themselves smoothed with a Gaussian kernel for final analysis.
Statistical Analysis
All statistical analyses were performed using custom scripts written in Python with the SciPy and statsmodels libraries. All tests were two-sided, and results were considered significant at p < 0.05. Shaded regions in plots represent the standard error of the mean (SEM).
Comparison of Proportions: A Fisher’s exact test was used to assess the enrichment of predictive tuning in triple-encoding neurons (Fig. 2F). A Chi-squared test was used for other comparisons of categorical data (Fig. 2C, 3C, 3E).
Paired Comparisons: To compare the presence of significant slope encoding on successful versus failed trials within the same neuronal population (Fig. 3D), we used the McNemar test for paired nominal data.
One-Sample Tests: Distributions of coefficients and correlation values were tested for significance against a median of zero using the non-parametric Wilcoxon signed-rank test (Fig. 1H, 3A–B, 3F–G, 4C–D).
Time-Series Analysis: To assess the significance of the group-average time course of path variability (Fig. 4E), we used a permutation test with cluster-based correction. The null hypothesis data was generated by creating one “shuffled” trace for each session. To do this, we first took the real firing rate trace for each neuron in each trial and performed a circular shuffle. The PCA and path variability calculations were then re-run on this shuffled single-trial data to generate the final shuffled trace for each session. The group of real session traces was then compared to the group of shuffled traces using a paired t-test at each time point, with cluster-based correction to control for multiple comparisons.
Confirmatory Analyses: To ensure that results were not confounded by the use of different partners across sessions, all primary statistical analyses were confirmed using linear mixed-effects models with partner identity included as a random effect. These confirmatory models yielded the same statistical conclusions, and therefore, we kept the original results.
Successful and failed pull comparison: Since the successful and failed pulls had different numbers of occurrences, to match the statistical power, we used all failed pull trials and randomly picked the same number of successful pull trials for each session, and performed pair-wise comparisons.
Supplementary Material
Acknowledgments:
This work was supported by the ‘Biological Sciences Training Program (BSTP)’ training grant (T32 MH014276 to W.S.), the National Institute of Mental Health (R21 MH126072, RF1 MH138396 to S.W.C.C., A.S.N., M.P.J.), the Simons Foundation Autism Research Initiative (SFARI 875855 to S.W.C.C., A.S.N., M.P.J.), Wu Tsai Institute at Yale University (to S.W.C.C., A.S.N., M.P.J., W.S.), a National Eye Institute core grant for vision research (P30 EY026878 to Yale University) and the National Science Foundation Graduate Research Fellowship (DGE2139841 to O.C.M.). We thank the veterinary and husbandry staff at Yale for their excellent animal care, and Anthony DeSimone from the Yale Machine Shop for advice on the neural recording system design.
Footnotes
Conflict of Interest: None
References
- 1.Burkart J. M., van Schaik C. P., Cognitive consequences of cooperative breeding in primates? Anim Cogn 13, 1–19 (2010). [DOI] [PubMed] [Google Scholar]
- 2.Emery N. J., The eyes have it: the neuroethology, function and evolution of social gaze. Neurosci Biobehav R 24, 581–604 (2000). [Google Scholar]
- 3.Frischen A., Bayliss A. P., Tipper S. P., Gaze cueing of attention: Visual attention, social cognition, and individual differences. Psychol Bull 133, 694–724 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ratcliff R., Smith P. L., Brown S. D., McKoon G., Diffusion Decision Model: Current Issues anc History. Trends in Cognitive Sciences 20, 260–281 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Gold J. I., Shadlen M. N., The neural basis of decision making. Annual Review of Neuroscience 30, 535–574 (2007). [Google Scholar]
- 6.Busemeyer J. R., Gluth S., Rieskamp J., Turner B., Cognitive and Neural Bases of Multi-Attribute, Multi-Alternative, Value-based Decisions. Trends in Cognitive Sciences 23, 251–263 (2019). [DOI] [PubMed] [Google Scholar]
- 7.Krajbich I., Armel C., Rangel A., Visual fixations and the computation and comparison of value in simple choice. Nature Neuroscience 13, 1292–1298 (2010). [DOI] [PubMed] [Google Scholar]
- 8.Krajbich I., Rangel A., Multialternative drift-diffusion model predicts the relationship between visual fixations and choice in value-based decisions. P Natl Acad Sci USA 108, 13852–13857 (2011). [Google Scholar]
- 9.Hanks T. D., Summerfield C., Perceptual Decision Making in Rodents, Monkeys, and Humans. Neuron 93, 15–31 (2017). [DOI] [PubMed] [Google Scholar]
- 10.Kim J. N., Shadlen M. N., Neural correlates of a decision in the dorsolateral prefrontal cortex of the macaque. Nat Neurosci 2, 176–185 (1999). [DOI] [PubMed] [Google Scholar]
- 11.Shadlen M. N., Newsome W. T., Neural basis of a perceptual decision in the parietal cortex (area LIP) of the rhesus monkey. Journal of Neurophysiology 86, 1916–1936 (2001). [DOI] [PubMed] [Google Scholar]
- 12.Meisner O. C. et al. , Development of a Marmoset Apparatus for Automated Pulling to study cooperative behaviors. Elife 13, (2024). [Google Scholar]
- 13.Meisner O. C. et al. , Diverse and flexible strategies enable successful cooperation in marmoset dyads. Curr Biol, (2025). [Google Scholar]
- 14.Ruff C. C., Fehr E., The neurobiology of rewards and values in social decision making. Nature Reviews Neuroscience 15, 549–562 (2014). [DOI] [PubMed] [Google Scholar]
- 15.Gangopadhyay P., Chawla M., Dal Monte O., Chang S. W. C., Prefrontal-amygdala circuits in social decision-making. Nat Neurosci 24, 5–18 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Isoda M., The Role of the Medial Prefrontal Cortex in Moderating Neural Representations of Self and Other in Primates. Annual Review of Neuroscience, Vol 44, 2021 44, 295–313 (2021). [Google Scholar]
- 17.Noritake A., Ninomiya T., Isoda M., Social reward monitoring and valuation in the macaque brain. Nature Neuroscience 21, 1452-+ (2018). [DOI] [PubMed] [Google Scholar]
- 18.Karashchuk P. et al. , Anipose: A toolkit for robust markerless 3D pose estimation. Cell Reports 36, (2021). [Google Scholar]
- 19.Lauer J. et al. , Multi-animal pose estimation, identification and tracking with DeepLabCut. Nat Methods 19, 496–504 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wiecki T. V., Sofer I., Frank M. J., HDDM: Hierarchical Bayesian estimation of the Drift-Diffusion Model in Python. Front Neuroinform 7, (2013). [Google Scholar]
- 21.Monsalve-Mercado M. M., Stine G. M., Shadlen M. N., Miller K. D., The geometry of the neural state space of decisions. bioRxiv, (2025). [Google Scholar]
- 22.Dal Monte O. et al. , Widespread implementations of interactive social gaze neurons in the primate prefrontal-amygdala networks. Neuron 110, 2183-+ (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Fan S. Q., Dal Monte O., Chang S. W. C., Levels of naturalism in social neuroscience research. Iscience 24, (2021). [Google Scholar]
- 24.Jiang M. et al. , Neural basis of cooperative behavior in biological and artificial intelligence systems. Science, eadw8151 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Jiang M. et al. , Evolution and neural representation of mammalian cooperative behavior. Cell Rep 37, 110029 (2021). [DOI] [PubMed] [Google Scholar]
- 26.Franch M. et al. , Visuo-frontal interactions during social learning in freely moving macaques. Nature, (2024). [Google Scholar]
- 27.Jendritza P., Klein F. J., Fries P., Multi-area recordings and optogenetics in the awake, behaving marmoset. Nat Commun 14, 577 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Walker J. D. et al. , Chronic wireless neural population recordings with common marmosets. Cell Rep 36, 109379 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Calhoun A. J., Pillow J. W., Murthy M., Unsupervised identification of the internal states that shape natural behavior. Nature Neuroscience 22, 2040-+ (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Coen P., Xie M., Clemens J., Murthy M., Sensorimotor Transformations Underlying Variability in Song Intensity during Courtship. Neuron 89, 629–644 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Coen P. et al. , Dynamic sensory cues shape song structure in. Nature 507, 233-+ (2014). [DOI] [PubMed] [Google Scholar]
- 32.Li J. W., Aoi M. C., Miller C. T., Representing the dynamics of natural marmoset vocal behaviors in frontal cortex. Neuron 112, (2024). [Google Scholar]
- 33.Purcell B. A., Kiani R., Neural Mechanisms of Post-error Adjustments of Decision Policy in Parietal Cortex. Neuron 89, 658–671 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Gold J. I., Law C. T., Connolly P., Bennur S., The Relative Influences of Priors and Sensory Evidence on an Oculomotor Decision Variable During Perceptual Learning. Journal of Neurophysiology 100, 2653–2668 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Odoemene O., Pisupati S., Nguyen H., Churchland A. K., Visual Evidence Accumulation Guides Decision-Making in Unrestrained Mice. J Neurosci 38, 10143–10155 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Fan S. Q., Dal Monte O., Nair A. R., Fagan N. A., Chang S. W. C., Closed-loop microstimulations of the orbitofrontal cortex during real-life gaze interaction enhance dynamic social attention. Neuron 112, (2024). [Google Scholar]
- 37.Camerer C., Behavioral game theory : experiments in strategic interaction. The Roundtable series in behavioral economics (Russell Sage Foundation; Princeton University Press, New York; Princeton, N.J., 2003), pp. xv, 550 p. [Google Scholar]
- 38.Kennedy D. P., Adolphs R., The social brain in psychiatric and neurological disorders. Trends in Cognitive Sciences 16, 559–572 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Zhu C. X., Dastani M., Wang S. H., A survey of multi-agent deep reinforcement learning with communication. Auton Agent Multi-Ag 38, (2024). [Google Scholar]
- 40.Mathis A. et al. , DeepLabCut: markerless pose estimation of user-defined body parts with deep learning. Nature Neuroscience 21, 1281-+ (2018). [DOI] [PubMed] [Google Scholar]
- 41.Woodward A. et al. , The Brain/MINDS 3D digital marmoset brain atlas (vol 5, 180009, 2018). Sci Data 9, (2022). [Google Scholar]
- 42.Pachitariu M., Sridhar S., Pennington J., Stringer C., Spike sorting with Kilosort4. Nat Methods 21, (2024). [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.




