Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Mar 6.
Published in final edited form as: Appl Anim Behav Sci. 2025 Mar 1;285:106578. doi: 10.1016/j.applanim.2025.106578

Establishing a Predictable Cue for Catches to Reduce Reactivity to Management Events for Captive Rhesus Macaques (Macaca mulatta)

Alexander J Pritchard 1,2,*, Rose A Blersch 1,2, Amy C Nathman 1, Eli R DeBruyn 1, Julia A Salamango 1, Emily M Dura 1, Brianne A Beisner 1, Jessica J Vandeleest 1,2, Brenda McCowan 1,2
PMCID: PMC12952906  NIHMSID: NIHMS2149014  PMID: 41778142

Abstract

Psychological duress can emerge from the perceived lack of predictability such that, in captive circumstances, reliable signals for aversive events can afford animals with the opportunity to behaviorally and physiologically prepare. Does a reliable and unique signal cue for an aversive management event reduce reactivity to management events that share unreliable cues? We recorded animal responses to management events near, or involving, outdoor-housed rhesus macaques (Macaca mulatta) in two large mixed-sex groups, with experimental periods that introduced a signal coupled to catch events. Management events varied in the severity and magnitude of animal responses. Our results validated that catches were more disruptive than management events that indirectly involved animal subjects, yet were comparable to management events involving direct interactions. Signal use reduced aversive responses to more routine management events that shared unreliable cues with catches. Due to the abundance of these routine events, we assert that the value of change with the implementation of the signal provided a detectable improvement across multiple measures of disruption.

Introduction

Perceived control and predictability of an event are critical qualities determining activation of a stress response (Romero & Wingfield, 2015). In captivity, the number of environmental variables under a subject’s control are considerably limited (Morgan & Tromborg, 2007). Individuals cannot escape a stressor beyond the enclosure (Morgan & Tromborg, 2007) or cannot disperse from disruptive social conditions (Pritchard et al., 2024). Signaling upcoming aversive events establishes predictability and facilitates control by providing an opportunity for behavioral and physiological preparation (Badia et al., 1979; Bassett & Buchanan-Smith, 2007). Subjects have been shown to choose signaled aversive events over unpredictable events (Bassett & Buchanan-Smith, 2007), and prefer signaled temporally-variable shocks over unsignaled temporally-fixed shocks (Badia et al., 1975). Although signaling is presented as a viable method of behavioral management to alleviate psychological distress in captivity, implementing such an approach in an applied setting with socially housed animals has not been well-demonstrated. This is an important caveat given this is a preferred mode of housing for gregarious animals, such as primates.

With regards to animal management, “negative (aversive) events should be made predictable” (p.240, Basset & Buchanan-Smith, 2017) yet few nonhuman primate studies have implemented and examine the effect of predictability for aversive management events, or for stimuli that share cues with aversive events. This latter stipulation is important, as aversive events (e.g., enclosure entry, animal capture) can share cues with otherwise neutral or positive events (e.g., feeding, providing enrichment) – which events are aversive likely depends on the routine of management. In zoo-housed brown capuchins (Sapajus apella), ‘knock’ signals have been used to discriminate door usage for enclosure entry from door usage for other management activities (Rimpley & Buchanan-Smith, 2013). This signal reduced anxiety-related behaviors relative to a baseline period sans signal. Among single-housed rhesus macaques (Macaca mulatta), the use of unique ‘doorbell’ signals for three types of management events resulted in lower rates of vocalizations (Gottlieb et al., 2013). These indoor single-housed subjects, however, were argued to exhibit reduced control or stimulation relative to group- and outdoor-housed animals – with uncertainty in how these outcomes would scale to large groups (Gottlieb et al., 2013).

In group settings, establishing predictability is important as multiple stimuli might co-occur across events or contexts, such that animals might associated unreliable cues with aversive events. For instance, catches may share unreliable cues (e.g., donning protective DuPont Tyvek® gowns [gowns, hereafter], vehicle use) with management events that occur more regularly (e.g., feeding, health checks, cleaning). Unreliable cues facilitate fear generalization (Dunsmoor & Paz, 2015) whereby cues linked to aversive events generalize fear responses to less invasive events. A predictable signal would be anticipated to disambiguate these unreliable cues from catching. Thus, establishing a unique predictable cue for aversive events is expected to reduce reactivity to unreliable cues. This approach likely requires fear inhibition on behalf of the subjects, which can manifest via: extinction of responses to unreliable cues, promoting perceptual threat discrimination, and/or learning when the environment is safe (Bassett & Buchanan-Smith, 2007; Clark et al., 2019; Milad & Quirk, 2012; Struyf et al., 2015). Signals of averse stressors can facilitate understanding of when the environment is safe (i.e., signal absence) (Bassett & Buchanan-Smith, 2007; Greiveldinger et al., 2007). As the signal is coupled to the averse stimuli, simple perceptual discrimination is expected to be promoted for unreliable cues (Jovanovic & Norrholm, 2011). Group settings, however, complicate the practical implementation of signaling procedures because signals are not individual specific and social cues could play an additional role in learning – implementation of such a signal would rely on both associative (Heyes, 2012) and social learning (Whiten, 2000) processes.

We implemented a signal in two large outdoor-housed groups of rhesus macaques to establish predictability for a notable aversive event: catches. Catches were selected because they share unreliable cues with other management events. Additionally, catches are characterized as distressing and potentially dangerous high energy events with group clustering, running, frequent vocalizations, and excreting of wastes (Personal Observations; Luttrell et al., 1994). They are routine management events in which staff enter the enclosure, and isolate a single individual for capture. This procedure is commonly used in routine management for social, veterinary, and research purposes which provides sufficient opportunities to learn the signal.

We had two interrelated research questions: how do management events vary in the severity of disruption for a group? And: how does the experimental introduction of a signal, predictably coupled to an aversive management event, alter the disruption severity across events? Informed by cumulative staff experiences and discussion, we expected events to be ordered from high-to-low disruption: catches, animal releases, health checks, technicians in gowns, technician vehicles, technicians in uniforms, and Primate Behavior & Health Service (PBHS) vehicles (see S1 for management event descriptions) (Figure 1). While the signal was used, we expected management events that involved a direct interaction with the subject group (i.e., releases, health checks) would exhibit reduced animal aversion (Figure 1). We had similar expectations for technician vehicle and gown events, with higher uncertainty due to the diverse husbandry of surrounding grounds and enclosures. Finally, uniformed technician, PBHS vehicle, and other management events did not typically have unreliable cues shared with catches; thus, we did not expect a strong signal effect. Signaled catches would be, either, less or more disruptive. Animal subjects could have the opportunity to behaviorally and physiologically prepare for catches (Badia et al., 1979; Bassett & Buchanan-Smith, 2007). The intervening time between signal and catches, however, could increase anticipatory behaviors (Bassett & Buchanan-Smith, 2007; Howell et al., 2023; Mistlberger, 1994; Polizzi di Sorrentino et al., 2010; Waitt & Buchanan-Smith, 2001).

Figure 1.

Figure 1.

Conceptual representation of project expectations. A) A bidirectional arrow represents a hypothetical continuum of disruptiveness, provoking variation in aversive behaviors. B) The different management events are expected to fall on different locations on this continuum, with some events being more disruptive than others. For simplicity, we have visualized this as rank ordered with no variation. C) When the signal is implemented, we expect that different management events will have distinct outcomes for the direction of the observable effect.

Methods

Research was completed at the California National Primate Research Center (CNPRC) at the University of California, Davis. This work was approved by the University of California, Davis’ Institutional Animal Care and Use Committee. Two studies were performed, each on a unique animal group. Within each study, subjects were housed in half-acre outdoor enclosures in a large mixed-sex group. Animal groups have access to food and water ad libitum, the former of which is distributed twice a day. Enclosures include play structures, mirrors, A-frames, benches, and sprinklers as well as heaters. Doors for the enclosures were located at the front of the enclosure, nearest the road. The frequency of necessary, routine, and irregular management events varied by research project needs, season, staff schedules, event type, and animal behavior. Technician vehicle events occur very frequently, but follow a temporal rhythm throughout the day that tracks with nearly all routine and ad libitum management events such as feeding, health checks, project schedules, etc. The donning of gowns can occur with the entry of any enclosure for any reason – including cleaning, moving equipment, catches, releases, morning health checks, etc. If staff are not in gowns, then they are always in uniform – thus, this event is nearly a catch-all for technicians present near the enclosure. Health check events occur twice a day accompanied with the provisioning of seed forage: once in the early morning and once in the afternoon, with morning checks also including enclosure entry. Catches and releases vary in their frequency depending on the health status of the animals, the number of animals in an enclosure, the season, and research project requirements. ‘Multiple’ events were scored for instances with multiple concurrent cues from among the following: technician vehicle, uniform, PBHS vehicle, other events, and adjacent enclosure. ‘Other’ types of disruptive events were assigned when unconventional stimuli were present that observers anticipated might have a disruptive potential for the animals. A priori expectations for multiple events and other types of disruptive events were not set.

This research was originally intended to function as two identical replicates. In the intervening time between the first replicate and the second, however, we were able to accommodate changes to the protocol that we estimated would aid implementation of signal learning. We term these replicates Study I and Study II, but emphasize that they were both intended a priori to address the original hypotheses. Formal analyses did not occur between the studies, and changes were made based on observer observations; not using the results reported here. Study I occurred from September 7 to December 15, 2023; Study II occurred March 5 to June 4, 2024. For both studies we collected data for six hours a day (9:00–12:00, 13:00–16:00), four days a week.

Study I Design

Study I’s group ‘A’ consisted of 149 study subjects (96 females; 53 males). Data collection included group-level observations of all animals in the enclosure. In Study I, we had a Baseline period, a Training period where we introduced the signal paired with catches but did not resume formal data collection, an Experimental period where we resumed data collection with signaled catches, and a Follow-up period where we discontinued the signal (Figure 2). Our three periods of data collection each lasted four weeks. Training was initially set to last two weeks, as used in similar work (Rimpley & Buchanan-Smith, 2013). During the second week, we made the decision to extend training to three weeks due to the group size and high number of catches at the start of training.

Figure 2.

Figure 2.

Timelines for the studies’ design, chronologically proceeding from left-to-right. Periods are represented with the colored blocks and are sized according to their duration of data collection. Days are annotated below blocks and show the number of data days in each period (not calendar duration). Block placement along the horizontal dashed line indicates whether that period had signaled catches (bottom blocks) or not (top blocks). Block fill is used to emphasize different periods. Hue is used in Study II’s experimental period to emphasize the two phases.

The signal was a remotely triggered wireless notification system with an extended range that produced a chime sound and a flashing red light (model: ERA-UTDCR; SKU: ERA-UT-KITS; Safeguard Supply, Kennesaw, GA). The receiver was housed in a weatherproof box – open on one side – positioned at the front of the enclosure nearest the main staff road. Technicians would trigger the signal when they left the office for a catch, which was nearly ~400m by road, yet ~100m directly. During Baseline, the technicians would still trigger the signal, observers could see the signal light though it was not visible or audible to the subjects during this time. Observers would initiate recording of the catch event when the signal was triggered.

As catch events could be relatively infrequent, we coordinated to have mock catches in the event that necessitated catches did not occur. Necessitated catches included traditional management events, such as: animal medical care, inspection to assess need for care, or colony movement of animals for other research projects. During mock catches, an individual would be selected in consultation with animal care technicians. They would coordinate to catch the subject, bring the animal to the front of the enclosure, and then release the subject instead of placing them in a box for transport. We did not have mock catches if sufficient necessitated catches occurred. For the first week of Training, mock catches were scheduled to occur to obtain at least two signaled catches each study day (morning and afternoon). This frequency was tapered to one a day for the rest of Training, but were cancelled if we had a necessitated catch that day. During Experimental and Follow-up we implemented sufficient mock catches to achieve a minimum of two events a week.

Study II Design

Study II’s group ‘B’ consisted of 150 study subjects (85 females; 65 males). Study II was largely similar to the protocol outlined with Study I with a few changes. The biggest change was extending the Experimental period in lieu of the Follow-up and interim Training periods (Figure 2). We divided the Experimental period into two analytical phases, with the first subsuming training effects. Thus, Baseline lasted four weeks, while each Experimental phase lasted five weeks. This change was largely due to observations that led us to believe that the higher frequency of catch events in Study I’s Training may have interfered with effective attentiveness to the signal. Thus, we reduced the frequency of mock catches to occur only if we needed to obtain a minimum of two catches in a week throughout the whole of Study II. Necessitated catches occurred as needed and always precluded the need for mock catches.

In addition, Group ‘B’ was housed ~100m by road to the office where catches would be initiated. Therefore, the time technicians spent in transit before a catch-up was short, relative to Study I. Thus, we altered the time of signaling to occur when staff were out of sight near the office, but before they donned protective gowns. Staff would then prepare to arrive within two minutes of the signal being pressed (the duration of the signal). If staff were unable to arrive within two minutes, then they would leave the area of the enclosure and signal, before returning to the enclosure. We streamlined data collection by omitting recording events involving only PBHS vehicles, as observers reported that animals did not react aversively to these events. This was logical as PBHS staff wore different uniforms than technicians and duties included provision of treats or enrichment.

Management Events and Measurements

Two trained observers (reliable at a threshold of Krippendorf’s α ≥ 0.85) alternated, by week, collecting data. Data collection consisted of whole-group animal observations, whereby observers would begin recording data upon the initiation of a management event. Events would be initiated by observer identification of a management event, but animal subjects often behaviorally responded to events before they were visible or audible – these events would be coded as well. We recorded whether animal subjects observed the event first (i.e., behaviorally responded), or the observer did; an inevitable design given that the subjects were very attentive and could react before observers could see or hear event cues. Management events could be labelled based on several identifiers (see S1 for descriptions) that were not mutually exclusive. For example, uniformed and gowned staff could both arrive in a vehicle and initiate a catch. During data collection, observers would select all relevant criteria, but, to simplify analysis, we assigned a single criterion to events that were scored with more than one criterion using the following priority order: catch, release, health, gowns, multiple events, technician vehicle, uniform, PBHS vehicle, other events, and adjacent enclosure (Supplementary Text 1). We made this decision prior to data exploration based on expectations guided by strong pooled and cumulative staff knowledge of the system. Assigning a priority for effects was the easiest way to avoid dichotomous fixed effects with multi-way interactions across events. After completing this process, we removed two unlabeled events in Study I and adjacent cage events had been subsumed by higher priority events. For Study II, we removed two unlabeled events and one adjacent enclosure event; the majority of these events were represented by other criteria higher in the priority order. Also for Study II, we omitted the uniform and other events as these types of events did not occur in Experimental Phase 1.

Scoring management events

During management events, we recorded several group-level measures on 30-second intervals, after the onset of an event. On interval marks, observers scored: response ratings, several dichotomous behaviors that were used to create a composite reactivity score, movement scores, and density scores. Intervals were continually added to a single event until the first of two termination criteria were met, either: 10 intervals had been collected, or at least 50% of the animals returned to their typical behavior patterns (a ‘relaxed’ state). During catch or release events, data collection would be paused once any technician entered the enclosure; data collection would resume upon the technicians’ exit at the relevant interval. This was because animal behavior with technician(s) in the enclosure was not comparable to behavior prior to or after entry due to animal reactions with the presence of staff and due to staff actively engaging the animals during these events.

Response ratings were a 1–5 rating of the severity of the animals’ response to an event throughout each interval. A rating of ‘1’ was indicative of the least severe reaction of ‘no response’, while ‘5’ was indicative of the most severe reaction typified by ‘a rapid response with fast locomotion vertically or laterally away from the disruption by a high majority of animals (>75%), sustained across the disruption sampling interval. The animals condense into one area, typically farthest from the source of the disruption. Inclusive of vocalizations and general arousal.’ The mid-point of ‘3’ was defined as a ‘moderate response typified by numerous animals (~50%) moving at a moderate speed with a few animals (<10%) responding with fast locomotion (jumping/fast move). Responses are occasionally accompanied by vocalizations. Individuals usually condense to nearby elevated areas, and a small majority (close to 50%) watch the event.’ We defined ‘2’ and ‘4’ as intermediate relative to the end- and mid-points set from the other ratings (‘1’, ‘3’, and ‘5’). This rating was a summative measure inclusive of the observers’ enclosure-side assessment of the prolongment and speed of the animals’ response, the number of animals that responded, distance the animals moved, and emotionality of the responding animals. Thus, this measure subsumed aspects of emotionality, avoidance, and magnitude of the response.

Reactivity scores were an additive score summarizing dichotomously coded behavioral responses from each interval, that were all expected a priori to reflect higher reactive responsiveness to disruptions. During each interval, observers dichotomously coded four behaviors: ‘cooing’ vocalizations, ‘alarm bark’ vocalizations, ‘monkeys watching’, and ‘vertical space use’. ‘Cooing’ and ‘alarm bark’ were each separately scored as a ‘1’ if any animals emitted the respective vocalization at any point during the relevant interval. ‘Monkeys watching’ were scored as a ‘1’ if at least 5 subjects watched the event for 3 or more seconds at any point during the relevant interval. Finally, we scored ‘vertical space use’ according to whether the majority of the subjects were on elevated spaces (i.e., higher than the lowest benches, which were approximately the height of a female monkey). After data collection, we summed each of these dichotomous scores within each interval, such that every interval had a reactivity score ranging from 0–4. These scores likely reflect emotionality, as these behaviors are generally elicited by and directed towards the disturbance, but are not capable of physically altering the animals’ circumstances in response to the stressor (e.g., escape).

Movement scores were the maximum percentage of animals that were observed moving during each interval. Observers scored each interval using one of four quartiles: <25%, 25–50%, 50–75%, and >75% of animal subjects moving. These scores likely measure active avoidance, as animals typically locomote away from the direction of the disruption.

Density scores were tallied at the end of each interval by ranking the density of the monkeys in each third of the enclosure. That is, at the end of an interval, observers would rank the section with the highest density of monkeys as a ‘1’, the second densest section as a ‘2’, and the section with the fewest monkeys as a ‘3’. As we were principally interested in ‘retreating’ behavior from the entry point at front of the enclosure, we limited our subsequent analyses to the ranking of the front third of the enclosure as a having a relative density of ‘1’, ‘2’, or ‘3’. These scores measure the relative abundance of animals. The front of the enclosure is generally preferred, when undisturbed, due to food hopper, shade, and heater availability, with high visibility of most neighboring enclosures and roadside activity. Scores of ‘3’ are indicative of avoidance of the front of the enclosure, relative to scores of ‘1’, with ‘2’s as intermediate.

Describing Responses During Catches

During Study I, we did not record group level responses during catch events themselves – indeed, we paused data logging until after technicians exited the enclosure. Thus, prior to beginning Study II, we implemented a recording protocol for a second observer to begin when technicians entered the enclosure for a catch. The aim was to document fine-grained qualitative changes in animal behavior that might be difficult to capture with formalized data collection techniques. We implemented a semi-structured open answer survey, which was linked to major events that occurred during catch events. The observer logged the timing of nine occurrences (e.g., enclosure entry, boxing of animal for transport) during each catch event while also responding to 14 questions with short answers. The subset of analyzed prompts is included in the Supplementary Materials (Supplementary Text 2).

Model Construction

We ran analyses with Bayesian regression models using Stan (Bürkner, 2017, 2018, 2021) in the R programming language (R Core Team, 2022). Outcomes were response ratings, reactivity scores, movement scores, or density scores, for which we compiled cumulative family models with flexible thresholds – logit-linked μ and log-linked discrimination. All disruption response variables were treated as ordered factors. For response rating, reactivity score, and movement score models: high estimates indicate higher disruption or aversion. For density score models: high estimates indicate lower disruption or aversion. For all models we included event type as an interaction with study periods or Experimental phases. We also included four dichotomous fixed effects, which scored a one if: the animals detected the event before the observers, the signal was activated, a technician approached the enclosure, or a technician entered the enclosure. We included a smoothed spline for event duration (interval counts) to accommodate the likelihood that events with many intervals might see a progressive change in the response variables. We included a random effect for hour and for date. This approach places the emphasis on differences at the level of the study period or phase because we anticipated variation across events and in daily patterns. We did not nest these random effects as we anticipated consistent hourly management schedules across dates, but also anticipated that particular dates might be more disruptive than others. The majority of events lasted one interval, so we omitted event ID as a random effect. We included discrimination terms for event type or period if they improved model fit. Discrimination terms allow variance across factor groupings, such as event types. This is logical given our expectation that some event types might simply be more disruptive – and perhaps include few or no low scores. The inclusion of discrimination terms facilitates greater model flexibility of the distribution for included fixed effects (Bürkner & Vuorre, 2019). We examined variation explained by our models using Bayesian pseudo-R2 estimates, suitable for ordinal models through the use of a continuous latent variable, which can be interpreted as analogous to Bayesian R2 estimates (Gelman et al., 2019; McKelvey & Zavoina, 1975).

Models were run with weakly informative priors, with 5,000 iterations, 2,000 warmup draws, and a thin of two across four chains, resulting in 6,000 post-warmup draws. As model estimates are relevant to the reference group, we conducted post hoc comparisons to examine our predictions. Posterior comparisons make within-model contrasts tractable. Fits were good according to R^ values, effective sample sizes, and posterior predictive checks (Supplementary Tables 18, Supplementary Figures 2a9b).

Results

Interpreting Response Ratings

Prior to proceeding with formal analysis, we examined the response ratings versus variation in reactivity and movements scores, to aid in interpreting response rating model outcomes. The response ratings showed nuance in their comparability with our reactivity and movement scores. Reactivity scores showed a clear shift in their proportions from a response rating of ‘1’ to ‘2’ (Figure 3). The comparable changes in reactivity scores, however, plateaued such that higher scores did not covary with further increases in response ratings. Movement scores, however, showed little to no change between response ratings of ‘1’ and ‘2’ (Figure 3). Though higher movement scores exhibited greater concordance with increases in response ratings for both studies. Given reactivity scores covaried with lower response scores, and movement scores covaried with higher response scores; we recoded the reactivity and movement scores into numeric values, and then combined them into a single metric (Figure 3). This combined metric exhibited similar variance to the response ratings, indicating that response ratings likely shared partial variance with both movement and reactivity scores.

Figure 3.

Figure 3.

Plots comparing variables of interest, with rectangles sized according to the proportion of the data represented in each bin. Response ratings (y-axis) relative to the x-axis of reactivity scores (top row), movement scores (middle row), and combined scores (bottom row) – a summed aggregate of reactivity and movement scores. The two studies are shown in the left (Study I) and right columns (Study II). See Supplementary Figure 1 for a colored version of the combined score figure.

Descriptive Information of Catches

Both studies had catches distributed throughout the periods that were predominantly related to general management processes – i.e., not mock catches. During Study I, a total of 74 catches were performed. The majority of these catches were observed (89%) and were necessitated by general management processes (84%). The animals experienced 20 catches during Baseline, 25 during Training, 15 during Experimental, and 14 during Follow-up. During Study II, a total of 65 catches were performed. The majority of these catches were observed (65%) and were necessitated by general management processes (91%). The animals experienced 24 catches during Baseline, 21 during Experimental Phase 1, and 20 during Phase 2. Thus, the two studies were similar in the number of catches and their distribution across study periods or phases.

Describing Responses During Catches

Here we report qualitative descriptions of animal responses to the signal during catches, which are important to understand learning processes that were not possible to capture with our disruption data protocol. During Study II, we collected qualitative data during 36 catch events: 10 during Baseline, 15 during Experimental Phase 1, and 11 during Phase 2 – seven events had two or more animals caught. These data provide context on complex processes that are difficult to measure in a large social group with ongoing social associative learning. During catches, technicians spent an average duration in the enclosure of 7m:12s, though this average was lower when catching a single animal (6m:21s). Among single animal catches, the duration in the enclosure was similar across study phases (≤ 1s difference between mean durations).

During Baseline, none of the animals altered their behavior when the signal was activated – which is expected as the signal was not visible or audible from inside the enclosure during Baseline, but the period when it would be activated was known to observers prior to technician arrival. In the first week of Experimental Phase 1 we had one signal event and ‘few’ animals interrupted their present behavior when the signal was activated. As the study advanced, observers scored an increase in the number of animals interrupting their behavior, from ‘few’ to ‘half’, before becoming ‘many’ through Experimental Phase 2. Through Phase 1, we saw a decrease in the resumption of behavior after signaling such that, by Phase 2, none of the animals resumed their behavior and many (exceeding 75% of animals in 45% of events) watched the technician office during Experimental Phase 2. Even so, we never observed more than 25% of animals watching the signal itself. After the technicians exited the enclosure, animals did not immediately resume their behavior from before the event. Thus, the management routine of catches did not markedly change, as indicated by the similar count and duration of catches, yet the animals qualitatively altered their behavior in response to the signal.

Model Summaries

Our fixed effects, without considering the interaction between event type and study period, provide insight into direct actions applied to the enclosure. That is, we generally observed more aversive responses in intervals when the animals detected the event, the signal was activated, or when technicians approached the enclosure. Though these dynamics exhibited some variation depending on the measure and study. In Study I and II, response ratings were higher with technician approach, signal activation, and animal detection of the event; technician enclosure entry did not alter ratings (Supplementary Table 1 & 2). Reactivity scores (i.e., composite scores of vocalizations, vertical space use, and watching) were higher with technician approach and animal detection of the event across both studies. In Study II, signal activation resulted in higher reactivity scores, but signal activation did not alter reactivity in Study I; enclosure entry did not change scores for either study (Supplementary Table 3 & 4). Movement scores were higher with signal activation and animal detection of the event, while enclosure entry decreased movement scores across both studies (Supplementary Table 5 & 6). Technician approach increased movement scores in Study I, but had no effect in Study II. Density scores were lower with signal activation while enclosure approach did not change scores across both studies (Supplementary Table 7 & 8). In Study I, density scores were lower with animal detection, but there was no effect in Study II. In Study II, density scores were lower with enclosure entry, but there was no effect in Study I. In summary, animal detection of events were consistent predictors of more aversive responses across studies and measures (except density scores in Study II), signal activation and technician approach were nearly consistent predictors of more aversive responses, while enclosure entries were less consistent in their predictive contribution – though did predict lower movement scores.

For Study I, our models’ conditional pseudo-R2 estimates ranged from 0.15 to 0.43 (mean = 0.35) and the marginal pseudo-R2 estimates ranged from 0.11 to 0.40 (mean = 0.31) (Supplementary Table 9). For Study II, our models’ conditional pseudo-R2 estimates ranged from 0.19 to 0.70 (mean = 0.43) and the marginal pseudo-R2 estimates ranged from 0.13 to 0.65 (mean = 0.38) (Supplementary Table 9). Thus, the majority of variation explained by our models was attributable to our fixed effects. Note that additional descriptive information regarding interval counts across the types of management events is provided in the supplementary (Supplementary Text 3).

Management Event Differences

Here we present tests of whether management events exhibit differences in response severity within each of our four measures. To do so, we performed post hoc pairwise comparisons using the emmeans package (Lenth, 2024) across our four measures of disruption for each of our studies (Supplementary Tables 10 & 11). We extracted the HPD intervals from the model posteriors and plotted the events represented in both studies to interrogate whether event types showed stable patterns across studies and measures (Figure 4). We present whole model outcomes averaged across the study periods and phases, but include the same comparisons between studies for Baselines and the Experimental period with Experimental Phase 1 in the Supplementary (Supplementary Figures 10 & 11).

Figure 4.

Figure 4.

Plots comparing model estimates (points) across management event types (color) with HPD intervals (error bars) across each of the studies (Study I on x-axis, Study II on y-axis), for each of the four measures (clockwise from upper left: response ratings, reactivity scores, density scores, and movement scores).

Across Study I and II, response ratings and movement scores exhibited similar relative placement in the ordering of events (Figure 4). Catch, release, and health clustered; gown was positioned intermediately, with multiple and technician vehicle events having the lowest estimates in response ratings and reactivity scores. Reactivity scores exhibited a pronounced separation of release events from the remaining events. Catch, gown, health, and multiple event types clustered, with technician vehicle types falling slightly below these estimates – albeit still overlapping with many of the event types. Density score estimates exhibited the highest uncertainty with wide credible intervals for intercepts, potentially attributable to difficulty in distinguishing thresholds (Supplementary Figures 1215). Even so, density score estimates for catch and gown events were similar, and release was intermediate to technician vehicle and multiple events. Health events were a special case in the density scores, as this event had the lowest estimates in Study I, but the highest in Study II. Post hoc interrogation of the health events between the studies showed that Study II was almost exclusively represented by afternoon health checks, while Study I had a higher incidence of AM health checks (Supplementary Figure 16). This is relevant as morning checks were typified by enclosure entry, with seed forage distributed inside after; afternoon checks were typified by seed forage distributed at the front of the enclosure with no enclosure entry.

Limiting our focus to Study I’s pairwise comparisons for response ratings, movement scores, and reactivity scores (Supplementary Table 10): catch, release, and health had high scores or ratings (i.e., were more disruptive), relative to the other event types. Health was lower than catch (for movement) or release (for reactivity). Gowns had higher scores than the majority of the remaining event types, though multiple and other events were similar depending on the measure. For response ratings and movement scores, the technician vehicle was similar to the multiple events, and multiple was similar to other events, but technician vehicle had higher reactivity and movement measurements relative to other events. Gown, technician vehicle, and multiple events were higher than uniform events. PBHS vehicle events were inconsistent across models but were often situated between uniform and technician vehicle events. In summary, the management events, rank ordered from most to least disruptive, were: catch and release – with the former higher in movement, but lower in reactivity; health; gowns, technician vehicle, multiple, other; PBHS vehicle; and uniformed.

In Study I, the ordering of management events differed based on density scores (Supplementary Table 10): uniform events had the highest rank-ordered densities of individuals in the front (i.e., least disruptive), equivalent to the PBHS vehicle type. Technician vehicle events were similar to multiple, other, and PBHS vehicle events, but had higher density scores relative to the catch, release, health, and gown events. Gown, multiple, other, release, and PBHS vehicle events were indistinguishable from each other. Health had similar densities to catch. For Study I density scores, the management events were rank ordered from most to least disruptive as: health, catch, gown, release, other, multiple, technician vehicle, PBHS vehicle, and uniformed.

Limiting our focus to Study II’s pairwise comparisons for response ratings, movement scores, and reactivity scores (Supplementary Table 11): catch, release, and health had high reactivity scores or response ratings, relative to the other event types. Depending on the measure, health was lower than catch or release. Gowns had higher scores than the remaining two event types. For response ratings and reactivity scores, the multiple event type was higher than the technician vehicle event; though they were equivalent for the movement scores. In summary, the management events, rank ordered from most to least disruptive, were: release, catch, health, gowns, multiple, and technician vehicle. Results were more consistent in rank order between reactivity scores and response ratings, relative to movement scores, despite the greater overlap in variance for reactivity scores.

As with Study I, the ordering of Study II’s management events also differed based on density scores (Supplementary Table 11): gown events had the lowest rank-ordered densities of individuals in the front, equivalent to catch events. Catch and technician vehicle events had lower density scores relative to health and multiple types. Health and multiple types were similar to each other, as well as release types. For Study II density scores, the management events were rank ordered from most to least disruptive as: gowns, catches, technician vehicle, release, multiple, and health.

Overall, patterns across our response ratings and movement scores aligned with our expectations, reactivity scores violated our expectations with an emphasis on discriminating release events from the remaining events, while density scores had high uncertainty for separating out the management events.

Signal’s Effect on Responses to Events

Finally, we present tests of whether the experimental introduction of a signal, coupled to an aversive management event, altered the disruption severity across events. We performed post hoc posterior comparisons between each of the periods, within each management event, using the hypothesis function in the brms package (Bürkner, 2017, 2018, 2021) across our four measures of disruption for each of our studies. Posterior comparisons (Figure 5 & 6) can be interpreted as meaningful if they markedly diverge from zero – i.e., no difference between the periods. Marked divergence was determined with an alpha of 0.025 or 0.050 for directional and bidirectional predictions, respectively.

Figure 5.

Figure 5.

Posterior draw density plots for Study I comparing the Experimental periods to Baseline or Follow-up (color hue). Densities reflect the difference between Experimental and the other period (x-axis) such that lower values indicate that Experimental estimates were lower. Bolded outlines highlight credible non-zero differences between periods in the expected direction. Plots show posterior draws for the four measures (clockwise from upper left: response ratings, reactivity scores, density scores, and movement scores). Y-axes are the management event types.

Figure 6.

Figure 6.

Posterior draw density plots for Study II comparing the two Experimental phases (1 & 2; color hue) to Baseline. Densities reflect the difference between each Experimental Phase and Baseline (x-axis) such that lower values indicate that Experimental estimates were lower. Bolded outlines highlight credible non-zero differences between periods in the expected direction. Plots show posterior draws for the four measures (clockwise from upper left: response ratings, reactivity scores, density scores, and movement scores). Y-axes are the management event types.

In Study I, based on response ratings and reactivity scores, technician vehicle events precipitated less aversive responses during the Experimental period (Figure 5; Supplementary Table 12). Based on density scores and relative to the Experimental period, catch events were more aversive during the Baseline period and release events were more aversive during Follow-up. In Study II, numerous event types exhibited a reduction in aversive behaviors across the Experimental Phases, with a more pronounced effect in Phase 2 (Figure 6; Supplementary Table 13). Relative to Baseline, technician vehicle events were less aversive based on reactivity scores in Experimental Phase 2 and density scores across both phases. Relative to Baseline, gown events were less aversive in Experimental Phase 1 based on response ratings and in Phase 2 based on response ratings, reactivity scores, as well as density scores. Relative to Baseline, health events were less aversive in Experimental Phase 2 based on response ratings and density scores. Catch event density scores in Experimental Phase 1 were credibly less averse than in Baseline. The remaining events did not exhibit differences in predicted directions, though multiple events in Experimental Phase 2 were lower in aversion based on density scores, relative to Baseline. Thus, across both studies, technician vehicle events were most consistent in exhibiting a reduced response with the signal. Gown events also exhibited a strong effect of the signal across multiple measures in Study II.

We did not formally test whether there were associations in the opposite direction of our predictions – that is, whether animals exhibited more aversive responses. Even so, numerous event types in Study I (health, release, and uniform) exhibited evidence of behavioral changes in the opposite direction of our predictions (i.e., Baseline or Follow-up were less aversive than the Experimental period) (Figure 5; Supplementary Table 12). In Study II, reactivity scores for the release events were likely to be higher in the Experimental Phase 2, relative to Baseline (Figure 6; Supplementary Table 13). For Study II, this was the only evidence that was directionally contradictory to our expectations.

To determine whether our Experimental period had latent effects that persisted after ceasing the signal, we tested whether the Study I Follow-up period estimates differed from Baseline (Supplementary Table 14). As this post hoc analysis was exploratory, we report directional outcomes. For response ratings, Baseline had lower estimates for health, release, and gown events. For movement scores, Baseline had lower estimates for gown and both types of vehicle events. Finally, for density scores, Baseline had higher estimates for release, gown, and technician vehicle events. Thus, gown events were consistently more aversive during Follow-up relative to Baseline. These estimate differences suggest that latent effects did exist and were indicative of greater disruption and aversion after ceasing the signal. In summary, the signal reduced aversive responses for some events in both studies, these effects did not persist after the cessation of the signal in Study I, but showed evidence of strengthening as signal use continued in Study II.

Discussion

The disruptive ordering of our management events was broadly similar to the expected ordering based on the authors’ cumulative experience, for three of our four measures, supporting a priori expectations of disruptive severity. The introduction of a signal linked to an aversive management event (i.e., catches) reduced responsiveness and reactivity to other management events, especially those sharing unreliable cues with catches such as technician vehicle or DuPont Tyvek® gown events. This effect was stronger in Study II and strengthened as the study progressed. Relative to Study I, Study II had a longer Experimental period and fewer catches early in signal implementation. Results from Study I, however, importantly contribute evidence that the reduced response to adverse events ceased after signals were no longer provided.

Management Event Differences

For response rating and movement scores, the disruptive ordering of events broadly aligned with our expectations: catch, health, and release events were all disruptive, multiple and technician vehicle events were not overly disruptive, and gown events were situated intermediately between these two groupings. This relative placement is logical for the highest scoring events as they shared a high likelihood of staff engaging the enclosure. Indeed, enclosure approaches also predicted increased response ratings and, for Study I, movement scores. Enclosure approach or entry is often identified as an aversive event (Gottlieb et al., 2013; Rimpley & Buchanan-Smith, 2013; Theil et al., 2017), yet the relative severity of response is not always reported. Indeed, in our work, enclosure entry did not show increased scores and, instead, predicted a reduction in movement across both studies. Importantly, variance was high between many of the events – unsurprising given the diverse nature of management activities and group responses, which likely include intergroup responsivity.

Reactivity scores violated our expectations with animal releases clearly separating from catches and health checks. Post hoc investigations of the raw data suggest that this separation was likely attributable to the proportion of alarm barks, which were markedly higher in the animal release events. In Study I and II, the proportions of release events with alarm barks exceeded six times the mean of the remaining events (Study I: 0.37 versus 0.06M±0.04sd; Study II: 0.61 versus 0.10M±0.07sd). Alarm barks, as warning calls (Lindburg, 1971), can parsimoniously be expected to be goal-directed towards evoking conspecific responses (Fischer & Price, 2017). Physiological arousal of an animal has been shown to alter production of such calls (Bercovitch et al., 1995). Drawing upon this evidence, we posit that animals could be more willing to vocalize during releases due to a lower perception of personal threat, relative to catches. Furthermore, animals could have higher perceived danger for released conspecifics, and engage in goal-directed calling for conspecifics.

Though considering the relative proportion of animals in the front of the enclosure was logical from an applied perspective, our density scores were difficult to interpret. This difficulty was likely due to poor discrimination between event types: density scores were primarily represented by scores of ‘1’, signifying that the front of the enclosure was densely occupied – though catch events had many more ‘3’s than the remaining types of events (Supplementary Figures 5a & 9a). This distinction is logical as an avoidance response for catches, because staff entry points are at the front of enclosure. Similarly, we observed high front densities scores for health events during Study II, but not Study I. Post hoc explorations suggest that this distinction was attributable to more morning health events in Study I (Supplementary Figure 15). Morning and afternoon health events differ in their nature with enclosure entry typical in the former, but not the latter (Supplementary Text 1).

Signal’s Effect on Responses to Events

Our work established predictability through the introduction of a signal for a highly aversive management event (catches). Signal implementation reduced aversive responses to management events that shared cues (vehicles and gowns) with catches. Study I showed that this benefit ceased after the signal was removed. Study II showed that the benefit strengthened, and became more generalized to other events, with continued use of the signal. From a management perspective, this evidence is indicative of a positive change towards reducing aversive behaviors for management events. This action is anticipated to occur through the establishment of predictability for an aversive event (Bassett & Buchanan-Smith, 2007) through the disambiguation of unreliable cues that occur across other management events. Importantly, we were able to implement this change using a strongly aversive event in a group setting, with simple modifications to animal management protocols: pressing a button before a catch.

Applied Recommendations

We coupled our signal to a management event that we expected would be the most aversive, following recommendations (Basset & Buchanan-Smith, 2017). We rationalized that the moderate or high stress of the catches themselves would increase receptivity to conditioned learning (Sandi & Pinelo-Nava, 2007). We note, however, that if clear and unique signals already exist for an aversive event, it is unlikely that adding a new signal would be of utility. Catch events were also selected because they shared unreliable cues with other routine events that subsequently exhibited the greatest change with the signal. The success of our approach was most strongly visible in less invasive management events that shared unreliable cues with catches. Even so, more invasive events (i.e., events involving engaging or entering the enclosure) also exhibited lower scores on our measures towards the end of the Study II. This is logical because these events share a greater number of cues with catches and were relatively disruptive to begin with.

If a predictable signal cannot be established frequently or clearly, then its utility would be questionable. That is, sufficient catches need to occur in association with the signal to facilitate learning. The frequency of events must be informed by the management system, species, and events. Indeed, we refined the frequency of events based on the system. We had implemented mock catches to maximize associative (Heyes, 2012) and social (Whiten, 2000) learning opportunities. We became uncertain, during Study I’s data collection, whether the number of catches was increasing duress. Analyses later revealed that Study I catch and release events had higher densities of individuals in the front of the enclosure during Experimental, relative to Baseline. This finding could be interpreted as a habituation effect resulting from catches in Training or Experimental. Alternatively, the signal could have provided sufficient warning for individuals to prepare for the impending catch (Badia et al., 1979; Bassett & Buchanan-Smith, 2007). Nevertheless, we made the decision to extend Study II’s Experimental period, to give more time for learning and increase the interval between catches. Although subjects experienced a similar number of signaled catches between the studies (40 in Study I, 41 in Study II), this change seemed to be an improvement given most of our effects were stronger in Study II’s Experimental Phase 2. It is unclear whether this effect is attributable to the greater duration, lower rate of catches, or other effects.

Enclosure conditions and species typical behaviors are important to consider prior to signal implementation. The two studies had different animal groups housed in distinct locations. Group ‘B’ might have had higher certainty of other events or event cues due to greater visibility of the management office. Seasonal variation might have also altered groups’ capacity to learn or respond to stressors. For example, the fall breeding season is a time of heightened sociosexual intensity with concurrent increases in aggression (Theil et al., 2017; Wilson & Boelkins, 1970).

Finally, we utilized several measures of disruption that exhibited some covariance and some divergence. For instance, in Study I, movement scores did not exhibit similar patterns to reactivity and response ratings for technician vehicles events. This lack of concordance is reconcilable given reactivity and response scores covaried more at lower values, while movement scores covaried with higher response scores (Figure 1). Thus, reactivity measures may provide a more nuanced scale for events with lower levels of disruption. Alternatively, reactivity and response ratings may provide more insight into emotionality rather than gross avoidance behavior. We do not advocate, however, for a single measure as better than any other measure. Rather, studies should incorporate multiple measures.

Conclusion

We implemented a reliable signal for aversive management events (Bassett & Buchanan-Smith, 2007) to disambiguate unreliable cues from routine enclosure management events (Gottlieb et al., 2013; Rimpley & Buchanan-Smith, 2013). Signal implementation reduced aversive responses to several less invasive management events, including technician vehicle events. The frequency of these routine events indicates that predictable signals have a high capacity to provide a strong quality of life improvement in captive animal welfare. For instance, technician vehicle events accounted for 27.68% and 26.17% of all intervals in Study I and II, respectively, with a frequency over twelve times that for catch and release events, our most disruptive events. Thus, our work shows a group-level improvement in reducing

Supplementary Material

Supplementary

Acknowledgements

We greatly appreciate the flexibility, attentiveness, and eagerness shown by CNPRC’s animal care and management staff during the implementation of this project. This work was funded under the National Institutes of Health’s #5R24-OD030036 (BM) and #P51-OD011107 (CNPRC) grants.

Data Availability

Data and analytical files needed to replicate manuscript results are provided via the Dryad data repository: https://doi.org/10.5061/dryad.h18931zw8

References

  1. Badia P, Harsh J, & Abbott B (1979). Choosing between predictable and unpredictable shock conditions: Data and theory. Psychological Bulletin, 86(5), 1107–1131. 10.1037/0033-2909.86.5.1107 [DOI] [Google Scholar]
  2. Badia P, Harsh J, & Coker CC (1975). Choosing between fixed time and variable time shock. Learning and Motivation, 6(2), 264–278. 10.1016/0023-9690(75)90027-2 [DOI] [Google Scholar]
  3. Bassett L, & Buchanan-Smith HM (2007). Effects of predictability on the welfare of captive animals. Applied Animal Behaviour Science, 102(3), 223–245. 10.1016/j.applanim.2006.05.029 [DOI] [Google Scholar]
  4. Bercovitch FB, Hauser MD, & Jones JH (1995). The endocrine stress response and alarm vocalizations in rhesus macaques. Animal Behaviour, 49(6), 1703–1706. 10.1016/0003-3472(95)90093-4 [DOI] [Google Scholar]
  5. Bürkner P-C (2017). brms: An R Package for Bayesian Multilevel Models Using Stan. Journal of Statistical Software, 80(1), 1–28. 10.18637/jss.v080.i01 [DOI] [Google Scholar]
  6. Bürkner P-C (2018). Advanced Bayesian Multilevel Modeling with the R Package brms. The R Journal, 10(1), 395–411. 10.32614/RJ-2018-017 [DOI] [Google Scholar]
  7. Bürkner P-C (2021). Bayesian Item Response Modeling in R with brms and Stan. Journal of Statistical Software, 100(5), 1–54. 10.18637/jss.v100.i05 [DOI] [Google Scholar]
  8. Bürkner P-C, & Vuorre M (2019). Ordinal Regression Models in Psychology: A Tutorial. Advances in Methods and Practices in Psychological Science, 2(1), 77–101. 10.1177/2515245918823199 [DOI] [Google Scholar]
  9. Clark JW, Drummond SPA, Hoyer D, & Jacobson LH (2019). Sex differences in mouse models of fear inhibition: Fear extinction, safety learning, and fear–safety discrimination. British Journal of Pharmacology, 176(21), 4149–4158. 10.1111/bph.14600 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Dunsmoor JE, & Paz R (2015). Fear Generalization and Anxiety: Behavioral and Neural Mechanisms. Biological Psychiatry, 78(5), 336–343. 10.1016/j.biopsych.2015.04.010 [DOI] [PubMed] [Google Scholar]
  11. Fischer J, & Price T (2017). Meaning, intention, and inference in primate vocal communication. Neuroscience & Biobehavioral Reviews, 82, 22–31. 10.1016/j.neubiorev.2016.10.014 [DOI] [PubMed] [Google Scholar]
  12. Gelman A, Goodrich B, Gabry J, & Vehtari A (2019). R-squared for Bayesian Regression Models. The American Statistician, 73(3), 307–309. 10.1080/00031305.2018.1549100 [DOI] [Google Scholar]
  13. Gottlieb DH, Coleman K, & McCowan B (2013). The effects of predictability in daily husbandry routines on captive rhesus macaques (Macaca mulatta). Applied Animal Behaviour Science, 143(2), 117–127. 10.1016/j.applanim.2012.10.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Greiveldinger L, Veissier I, & Boissy A (2007). Emotional experience in sheep: Predictability of a sudden event lowers subsequent emotional responses. Physiology & Behavior, 92(4), 675–683. 10.1016/j.physbeh.2007.05.012 [DOI] [PubMed] [Google Scholar]
  15. Heyes C (2012). Simple minds: A qualified defence of associative learning. Philosophical Transactions of the Royal Society B: Biological Sciences, 367(1603), 2695–2703. 10.1098/rstb.2012.0217 [DOI] [Google Scholar]
  16. Howell CP, Warnes C, & Bayliss PA (2023). The influence of feeding routines on the behavior of zoo-housed Sulawesi crested black macaques (Macaca nigra). Zoo Biology, 42(6), 757–765. 10.1002/zoo.21790 [DOI] [PubMed] [Google Scholar]
  17. Jovanovic T, & Norrholm SD (2011). Neural Mechanisms of Impaired Fear Inhibition in Posttraumatic Stress Disorder. Frontiers in Behavioral Neuroscience, 5. 10.3389/fnbeh.2011.00044 [DOI] [Google Scholar]
  18. Lenth R (2024). Emmeans: Estimated Marginal Means, aka Least-Squares Means. R Package v 1.10.2 [Computer software]. https://CRAN.R-project.org/package=emmeans
  19. Lindburg DG (1971). The rhesus monkey in North India: An ecological and behavioral study. Primate Behavior: Developments in Field and Laboratory Research, 2, 1–106. [Google Scholar]
  20. McKelvey RD, & Zavoina W (1975). A statistical model for the analysis of ordinal level dependent variables. The Journal of Mathematical Sociology, 4(1), 103–120. 10.1080/0022250X.1975.9989847 [DOI] [Google Scholar]
  21. Milad MR, & Quirk GJ (2012). Fear Extinction as a Model for Translational Neuroscience: Ten Years of Progress. Annual Review of Psychology, 63(1), 129–151. 10.1146/annurev.psych.121208.131631 [DOI] [Google Scholar]
  22. Mistlberger RE (1994). Circadian food-anticipatory activity: Formal models and physiological mechanisms. Neuroscience & Biobehavioral Reviews, 18(2), 171–195. 10.1016/0149-7634(94)90023-X [DOI] [PubMed] [Google Scholar]
  23. Morgan KN, & Tromborg CT (2007). Sources of stress in captivity. Applied Animal Behaviour Science, 102(3–4), 262–302. 10.1016/j.applanim.2006.05.032 [DOI] [Google Scholar]
  24. Polizzi di Sorrentino E, Schino G, Visalberghi E, & Aureli F (2010). What time is it? Coping with expected feeding time in capuchin monkeys. Animal Behaviour, 80(1), 117–123. 10.1016/j.anbehav.2010.04.008 [DOI] [Google Scholar]
  25. Pritchard AJ, Beisner BA, Nathman A, & McCowan B (2024). Social stability via management of natal males in captive rhesus macaques (Macaca mulatta). Journal of Applied Animal Welfare Science, 1–18. 10.1080/10888705.2024.2303679 [DOI] [Google Scholar]
  26. R Core Team. (2022). R: A language and environment for statistical computing. (Version 4.2.2) [Computer software] R Foundation for Statistical Computing. https://www.R-project.org/ [Google Scholar]
  27. Rimpley K, & Buchanan-Smith HM (2013). Reliably signalling a startling husbandry event improves welfare of zoo-housed capuchins (Sapajus apella). Applied Animal Behaviour Science, 147(1), 205–213. 10.1016/j.applanim.2013.04.017 [DOI] [Google Scholar]
  28. Romero LM, & Wingfield JC (2015). Models of Stress. In Tempests, Poxes, Predators, and People (pp. 69–114). Oxford University Press. https://global.oup.com/academic/product/tempests-poxes-predators-and-people-9780195366693?cc=us&lang=en& [Google Scholar]
  29. Struyf D, Zaman J, Vervliet B, & Van Diest I (2015). Perceptual discrimination in fear generalization: Mechanistic and clinical implications. Neuroscience & Biobehavioral Reviews, 59, 201–207. 10.1016/j.neubiorev.2015.11.004 [DOI] [PubMed] [Google Scholar]
  30. Theil JH, Beisner BA, Hill AE, & McCowan B (2017). Effects of Human Management Events on Conspecific Aggression in Captive Rhesus Macaques (Macaca mulatta). Journal of the American Association for Laboratory Animal Science, 56(2), 122–130. [PMC free article] [PubMed] [Google Scholar]
  31. Waitt C, & Buchanan-Smith HM (2001). What time is feeding?: How delays and anticipation of feeding schedules affect stump-tailed macaque behavior. Applied Animal Behaviour Science, 75(1), 75–85. 10.1016/S0168-1591(01)00174-5 [DOI] [Google Scholar]
  32. Whiten A (2000). Primate Culture and Social Learning. Cognitive Science, 24(3), 477–508. 10.1207/s15516709cog2403_6 [DOI] [Google Scholar]
  33. Wilson AP, & Boelkins RC (1970). Evidence for seasonal variation in aggressive behaviour by Macaca mulatta. Animal Behaviour, 18, 719–724. 10.1016/0003-3472(70)90017-5 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary

Data Availability Statement

Data and analytical files needed to replicate manuscript results are provided via the Dryad data repository: https://doi.org/10.5061/dryad.h18931zw8

RESOURCES