Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2017 Oct 1.
Published in final edited form as: Law Hum Behav. 2016 May 30;40(5):524–535. doi: 10.1037/lhb0000198

Adversarial Allegiance: The Devil is in the Evidence Details, Not Just on the Witness Stand

Bradley D McAuliff 1, Jeana L Arter 2
PMCID: PMC5036989  NIHMSID: NIHMS786278  PMID: 27243362

Abstract

This study examined the potential influence of adversarial allegiance on expert testimony in a simulated child sexual abuse case. A national sample of 100 witness suggestibility experts reviewed a police interview of an alleged 5 year-old female victim. Retaining party (prosecution, defense) and interview suggestibility (low, high) varied across experts. Experts were very willing to testify, but more so for the prosecution than the defense when interview suggestibility was low and vice versa when interview suggestibility was high. Experts' anticipated testimony focused more on pro-defense aspects of the police interview and child's memory overall (negativity bias), but favored retaining party only when interview suggestibility was low. Unlike prosecution-retained experts who shifted their focus from pro-defense aspects of the case in the high suggestibility interview to pro-prosecution aspects in the low suggestibility interview, defense experts did not. Blind raters' perceptions of expert focus mirrored those findings. Despite an initial bias toward retaining party, experts' evaluations of child victim accuracy and interview quality were lower in the high versus low interview suggestibility condition only. Our data suggest that adversarial allegiance exists, that it can (but not always) influence how experts process evidence, and that it may be more likely in cases involving evidence that is not blatantly flawed. Defense experts may evaluate this type of evidence more negatively than prosecution experts due to negativity bias and positive testing strategies associated with confirmation bias.

Keywords: expert testimony, adversarial allegiance, negativity bias, confirmation bias, child sexual abuse, child victim testimony, suggestibility


The use of expert witnesses in jury trials is common, especially when evidence is technical or difficult to understand (Groscup, Penrod, Studebaker, Huss, & O'Neil, 2002). Experts are assumed to be neutral, objective parties who disseminate information to jurors; however, the adversarial system may bias experts in favor of the side that retained them (prosecution or defense). This tendency is known as adversarial allegiance (Murrie et al., 2008).

Basic information about expert testimony, such as frequency and type, has been difficult to obtain (Krafka, Dunn, Johnson, Cecil, & Miletich, 2002). The absence of public information about expert testimony and cases in which experts have intentionally provided incorrect information has intensified lay beliefs that experts at best are susceptible to bias and at worst are “hired guns” or “whores of the court” (Cooper & Neuhaus, 2000; Hagen, 1997). Lay people are not alone in their cynicism regarding expert testimony. Judges have lamented the use of expert testimony for more than a century. According to Mnookin (2008), the underlying message from the bench is resoundingly clear: “Expert witnesses in court are often not deserving of our confidence. Their conclusions cannot be relied upon, and their words cannot be trusted” (p. 1010). A primary source of judicial concern regarding party-retained expert testimony is partisanship or adversarial allegiance.

Adversarial Allegiance

The construct Murrie and his colleagues (2008) coined as “adversarial allegiance” has been conceptualized differently over the years. Zusman and Simon (1983) used the term “forensic identification” to describe the process by which neutral experts' involvement with litigants or attorneys fosters a viewpoint that subtly influences their subsequent examination and evaluation in the case. “Without intending to or even being aware of it, they [experts] tend to emphasize findings and patterns that support “their side””(p. 1304). Rogers (1987) contemplated how perceived agency might contaminate and skew an expert's conclusions in favor of a perceived client, and Brodsky (2013) noted a similar “pull to affiliate” that can affect experts' opinions and commitment to legal outcomes that favor the retaining party.

What can be distilled from these overlapping conceptualizations is that retention by or affiliation with a party in a legal proceeding may create bias that influences the expert's thoughts, feelings, and behavior in favor of the retaining or affiliated party. Of course experts, like any other witness, can intentionally and consciously bias their testimony in favor of one party. However, such behavior would be unethical and violate established professional practice guidelines (e.g., Ethical Principles of Psychologists and Code of Conduct, 2010; Specialty Guidelines for Forensic Psychology, 2013). Also procedural safeguards such as cross-examination, opposing expert testimony, and even the threat of prosecution in extreme cases hopefully should minimize the presence of deliberate bias in the courtroom. As a result, researchers have concentrated primarily on the unintentional and unconscious bias stemming from adversarial allegiance.

Systematic research on adversarial allegiance is surprisingly sparse (see Murrie & Boccaccini, 2015, for a review). On the whole, there is more evidence supporting a retaining-party bias (Murrie et al., 2008; Murrie et al., 2009; Otto, 1989; Zusman & Simon, 1983) than not (Beckham, Annis, & Gusafson, 1989). Yet we must keep in mind that this handful of studies has relied quite heavily on expert testimony by forensic evaluators in actual civil cases. This methodology, although irrefutably high in ecological validity, limits researchers' ability to randomly assign participants to conditions in which the retaining party and other critical variables of interest are manipulated systematically. Retaining parties, experts, and their conclusions must be taken “as is” and doing so may limit the internal validity of the research.

One remarkable study recently overcame the limitations inherent in field research (Murrie, Boccaccini, Guarnera & Rufino, 2013). Those researchers paid actual forensic psychologists and psychiatrists to review the same authentic offender case files, but deceived some participants to believe they were consulting for the prosecution and others to believe they were consulting for the defense. Participants met in person with a confederate attorney and thought they were taking part in a large-scale forensic consultation funded by a special defense unit or a public-defender service. Strong evidence of adversarial allegiance among the experts emerged (effect sizes up to d = 0.85). Those who believed they were prosecution consultants assigned higher risk scores to offenders whereas those who believed they were defense consultants assigned lower risk scores.

In the present study, we sought to advance our scientific understanding of adversarial allegiance. We used an experimental paradigm to randomly assign a previously unstudied population (witness suggestibility experts) from across the United States to conditions that systematically varied the retaining party in a simulated criminal child sexual abuse case. We also manipulated certain features of the evidence—specifically, whether the police interview of the alleged child victim was low or high in suggestibility—to determine whether this variable moderated adversarial bias. No other published research has examined how evidence features might interact with retaining party to influence adversarial allegiance. Identifying the mechanisms underlying adversarial allegiance “is a crucial next step in understanding allegiance and, in turn, intervening to reduce allegiance” (Murrie & Boccaccini, 2015, p. 51).

Expert Testimony on Suggestibility

Suggestibility is the extent to which certain cognitive, social, and developmental factors influence an individual's ability to encode, store, retrieve, and report an event (Ceci & Bruck, 1993). Scholars have studied suggestibility since the early 1900s; however, an unprecedented surge in research occurred after extreme allegations of child sexual abuse in preschool daycare settings surfaced in the early 1990s (e.g., McMartin Preschool case in California, Kelly Michaels case in New Jersey). Since that time, our scientific understanding of how the accuracy of memory can be influenced by suggestive questions has increased dramatically. For example, scholars in the field know that younger (versus older) children answering leading (versus open-ended questions) from a high (versus low) authority interviewer are less accurate (McAuliff & Kovera, 2007).

Social scientists have developed evidence-based strategies and protocols to overcome the challenges inherent in interviewing children. Most guidelines recommend that children be interviewed privately in an age-appropriate, child-friendly setting to minimize stress and avoid potential contamination from other adults (Saywitz, Lyon, & Goodman, 2011). Building rapport with children, giving them permission to say “I don't know” and “I don't understand,” and asking them to correct the interviewer when appropriate can increase the amount and accuracy of information provided by children (Lamb, Orbach, Hershkowitz, Esplin, & Horowitz, 2007). Supportive interviewer behaviors such smiling and relaxed body posture help children feel more comfortable and better resist misleading questions (Bottoms, Quas, & Davis, 2007). Perhaps most importantly, interviewers should ask more open-ended questions and fewer closed-ended questions (Lyon, 2010). This includes using non-suggestive invitations (“Tell me why you're here today”) and open-ended prompts (“Tell me more about that” and “What happened next?”) while avoiding yes/no and forced-choice questions.

Experts may impart their knowledge of witness suggestibility and interview protocols to jurors in court. Although jurors understand age-related trends in suggestibility, they lack crucial knowledge about how factors such as leading questions and interviewer authority can increase suggestibility and reduce accuracy (Buck, Warren, Bruck, & Kuehnle, 2014; McAuliff & Kovera, 2007; Quas, Thompson, & Clarke-Stewart, 2005). Expert testimony on these issues has been shown to improve jurors' understanding and decision-making. One study examined the effects of interview quality and expert testimony on mock jurors' evaluations of a child victim's forensic interview (Buck, London, & Wright, 2011). Mock jurors who received expert testimony were sensitive to poor interview quality and acquitted the defendant more often than mock jurors who received no expert testimony. Additional research has shown that expert testimony can improve jurors' understanding of children's testimonial demeanor (Kovera, Gresham, Borgida, Gray, & Regan, 1997), prompted reports (Laimon & Poole, 2008), hearsay testimony (Nunez, Gray, & Buck, 2012), and recovered memories (Buck & Warren, 2009).

The need for and helpfulness of expert testimony on suggestibility make it a fitting backdrop to examine the potential effects of adversarial allegiance. Do experts in a child sexual abuse case selectively focus their testimony on aspects of the police interview that favor the retaining party? Do prosecution-retained experts evaluate child accuracy and interview quality more favorably than experts retained by the defense or vice versa? Our study is the first to provide answers to these important questions.

Overview and Hypotheses

We examined the potential influence of adversarial allegiance on different aspects of expert testimony in a simulated child sexual abuse case. Experts were asked by the prosecution or defense to read a description of a police officer's low or high suggestibility interview of a 5 year-old girl who alleged inappropriate sexual touching by her stepfather. Experts then indicated their willingness to testify, described the issues they would address in their testimony, and rated the child victim's accuracy, trustworthiness, and police interview quality.

We developed four hypotheses. First, with respect to experts' willingness to testify, we did not anticipate any differences solely as a function of retaining party. Adversarial allegiance is a bias that stems from one's retention by or affiliation with a party in a legal proceeding and therefore should emerge after experts agree to testify for the prosecution or defense, but not before. Nevertheless, it is possible and perhaps even likely that experts take into account both retaining party and case evidence (e.g., police interview suggestibility) when deciding whether they are willing to testify. If true, then experts asked by the prosecution should be more willing to testify when interview suggestibility is low (strong evidence for the prosecution) than experts asked by the defense (weak evidence for the defense) and vice versa when interview suggestibility is high (Hypothesis 1).

Once experts agree to testify, adversarial allegiance should affect how they review case evidence in preparation for trial testimony. We hypothesized that prosecution-retained experts would focus on a larger number and higher proportion of pro-prosecution aspects of the police interview and child's memory (Hypothesis 2) whereas defense-retained experts would focus on a larger number and higher proportion of pro-defense aspects (Hypothesis 3). Our fourth hypothesis addressed the conclusions experts draw after reviewing the evidence. We hypothesized that experts' judgments of child victim accuracy, trustworthiness, and interview quality would vary as a function of retaining party. Prosecution-retained experts should rate the child and interview more favorably than defense-retained experts (Hypothesis 4).

In addition to these main effects, we predicted that Hypotheses 2, 3, and 4 would be qualified by a Retaining Party × Interview Suggestibility interaction such that when interview suggestibility was low, prosecution-retained experts would be more likely to provide testimony that favored the prosecution (both in terms of actual and perceived focus) and to rate the child victim and interview quality more positively than defense-retained experts. Conversely, when interview suggestibility was high, we predicted defense-retained experts would be more likely to provide testimony that favored the defense and to rate the child victim and interview quality more negatively than prosecution-retained experts.

Method

Expert Sample

We randomly selected five states from three geographic regions (15 states total) in the western, central, and eastern United States to identify potential expert witnesses to participate in the study. For each state, we used the terms “child sexual abuse,” “witness suggestibility,” and “forensic interviews,” to search online expert databases maintained by courts (e.g., http://www.lacourt.org/division/criminal/pdf/witnesses.pdf) or by contacting local court clerks directly. In addition, we reached out to attendees of the 26th Annual San Diego Conference on Child and Family Maltreatment with expertise in witness suggestibility and forensic interviews. The final list of potential respondents included 330 experts from 32 different states. Experts were contacted via electronic and regular mail (Dillman, 2000) and were offered $25 in exchange for their participation.

Forty-eight percent (n = 158) of experts responded in some way. One hundred provided complete responses, and 58 reported they were not qualified to serve as an expert on the topic of witness suggestibility. In part, the large number of “not qualified” respondents was because the court databases we consulted to identify experts were organized using very broad categories (e.g., “child sexual abuse”) and not by specialty (e.g., medical doctor, sexual abuse nurse examiner). As a result, our initial list of potential respondents included experts who were qualified to testify in child sexual abuse cases, but who were not experts on witness suggestibility or forensic interviewing per se. Thirteen invitation emails or letters were returned as undeliverable (4%) and 159 experts did not respond in any way (48%).

Thirty-seven percent of the final sample of experts who provided complete responses resided in California, 7% resided in Texas, and the remaining 56% were from 30 other states. Most respondents were female (63%) and they had earned doctorate degrees (46%), specialized in clinical psychology or child development (38%), and averaged 55 years in age. Sixty-seven percent of experts had been asked to testify in a child sexual abuse case, 61% had agreed to testify, and 57% had actually testified. All of the 33 experts who had never been asked to testify indicated they were qualified to do so. Experts testified for the prosecution an average of 28 times (Median = 5) and for the defense an average of 20 times (Median = 3). Twenty-two percent of experts reported having testified for both the prosecution and defense during their careers.

Procedure

We adapted several procedures from the Total Design Method of survey research (Dillman, 2000) to increase the response rate including multiple mailings and reminder notices at predetermined time intervals. Previous studies (Kovera & McAuliff, 2000; McAuliff & Kovera, 2007) using similar methods with experts and legal professionals have achieved response rates ranging from 40% to 60%.

We began by sending a contact postcard that described the study and outlined future correspondence. One week later, we mailed and emailed the invitation letter to each expert. If they chose to participate, they were asked to click on a web link provided in the email/letter that took them to the survey posted on www.psychsurveys.org. We provided a separate link for those who did not wish to participate in the survey. One week after the invitation letter, we mailed/emailed a reminder postcard to experts who had not completed the survey. One month after the initial contact postcard was sent, we mailed/emailed a final letter that emphasized the importance of receiving completed surveys from as many experts as possible in order for the sample results to generalize to the population.

We randomly assigned the final expert sample to one of four experimental conditions (n = 25) in which retaining party (prosecution, defense) and interview suggestibility (low, high) systematically varied. We ran a series of chi-square and univariate analyses of variance on several key variables to ensure that random assignment was successful. No statistically significant differences emerged for participant age, gender, education, or level of interaction with children across the four experimental conditions. All p values ranged from .06 to .66.

On average, experts spent 18.57 minutes (Median = 15 minutes) completing the study that appeared on eight different webpages: informed consent, police interview description, willingness to serve as an expert, anticipated testimony focus, child and interview ratings, manipulation check, demographic items, and payment information.

Stimulus Materials

Respondents read that they were being sought for potential expert testimony on behalf of the prosecution or defense in a child sexual abuse case. They read one of two descriptions of a 30-minute police interview of an alleged child victim. The high suggestibility interview condition described an interview conducted by a police officer in her full uniform using a standard interrogation room with nothing but a table, two chairs facing each other, and a third chair where the child's mother sat during the interview. The police officer told the child it was important to answer all of the questions and the interview consisted mainly of direct questions about information such as what the perpetrator did and said during the incident. In the low suggestibility interview condition, respondents were informed the police officer changed out of her formal uniform and conducted the interview in a “child friendly” room with comfortable furniture and games. The child's mom observed the interview through a one-way mirror in an adjacent room. The officer began by explaining the interview process and stated that it was important to say “I don't know” or “I don't understand” if the child did not know the answer or understand the question. The interview consisted mainly of open-ended questions followed by direct questions when necessary to clarify the child's responses.

In both conditions, the child disclosed that her stepfather made her touch his penis and “move it around” while she was watching television with him in the living room. She said he had never done anything like that to her before.

Dependent Measures

After reading the police interview description, respondents indicated their willingness to serve as an expert on witness suggestibility in the case by answering a 7-point Likert-type scale (1 = Certainly No, 4 = Neutral, 7 = Certainly Yes). An open-ended question followed, asking experts to describe what aspects of the police officer's interview and the child's memory they would choose to focus on if called to testify in the case. Experts then rated the child and interview by answering a series of 7-point Likert-type items (1 = Not at All, 4 = Neutral, 7 = Extremely). Three items measured experts' evaluations of child victim accuracy (“In your opinion, how accurate is Anna's memory for the events that she described?”, “How reliable was Anna's testimony?”, “How easily influenced is Anna by suggestive or misleading questions from the interviewer?”*), two items measured child victim trustworthiness (“How likely is it that Anna was telling the truth?”, “How believable was Anna's testimony”), and three items measured police interview quality (“Was Officer Olsen's interview of Anna biased?”*, “Was Officer Olsen's interview of Anna successful at obtaining the truth?”, “Was Officer Olsen's interview of Anna suggestive?”*). We averaged experts' ratings of the child and interview to create three composite measures: Child Victim Accuracy (Cronbach's alpha = .74), Child Victim Trustworthiness (Cronbach's alpha = .87), and Interview Quality (Cronbach's alpha = .81). We reverse scored three items (indicated by “*”) for inclusion in the composite variables. Higher values represented more positive evaluations of the child and interview.

Experts concluded their participation by answering a manipulation check question asking which attorney requested their expert testimony in the case (District Attorney/Prosecution or Defense Attorney) and then by completing a series of demographic items about their training, prior experiences as expert witnesses, and level of interaction with children.

We developed a coding scheme to identify specific topics in experts' responses to the open-ended question about what aspects of the police interview and child's memory they would focus on during their testimony. The final coding scheme consisted of 36 topics, one-third of which were features we deliberately varied between the low and high suggestibility levels of the police interview manipulation that experts read. These 12 topics focused on the police officer (experience, dress, introduction), the interview room (atmosphere, mom's presence), and interview (“don't know” “don't understand” and “answer all questions” ground rules, child's initial reluctance, encouragement, use of protocol, question type). We also performed an independent reading of the open-ended question responses to identify topics beyond our interview suggestibility manipulation that experts commonly reported they would focus on when testifying. This process yielded an additional 24 topics across both readers that we included in the final coding scheme.

Two graduate research assistants who were blind to the study's hypotheses and experimental conditions coded experts' responses to the open-ended question. Coding was done in two phases. In Phase 1, the research assistants read each expert's entire open-ended response to make an overall gut-level determination whether it favored the defense or prosecution (-1 = Pro-defense, 0 = Neutral, 1 = Pro-prosecution). Cohen's kappa was .83 across ratings for all 100 experts on this Perceived Expert Focus variable.

In Phase 2, we provided each research assistant with the pre-determined final coding scheme of 36 topics. The research assistants independently evaluated whether each topic was present (0 = Absent, 1 = Present) and if so, whether the expert's statement about the topic favored the defense or prosecution (-1 = Pro-defense, 0 = Neutral, 1 = Pro-prosecution). Examples of pro-prosecution statements included: “Child-friendly room,” “Introduction of Susan instead of Officer Olsen,” “Mother viewed interview from a different room,” and “The officer took time to build rapport with the child.” Examples of pro-defense statements included: “Interview took place in an interrogation room,” “Officer appeared in full uniform,” “Doubtful that a 5 year-old would use the word penis,” “The child was not instructed that it was okay to not answer all the questions.” Examples of neutral statements included: “How questions were asked,” “Was the child interviewed by a forensic interviewer” and “Essential to have audio or video or transcript of interview.” Cohen's kappa was .76 (range = .72 to .82 for each topic category) across ratings for all 100 experts on the 36 topics coded as present or absent and either pro-defense, neutral, or pro-prosecution. All coding discrepancies between raters were resolved through discussion prior to final data analysis.

We calculated four dependent measures of actual expert focus based on experts' responses regarding their anticipated testimony: (1) Prosecution Focus Count (total number of pro-prosecution statements); (2) Prosecution Focus Proportion (total number of pro-prosecution statements divided by the total sum of pro-prosecution and pro-defense statements); (3) Defense Focus Count (total number of pro-defense statements); and (4) Defense Focus Proportion (total number of pro-defense statements divided by the total sum of pro-prosecution and pro-defense statements).

Results

Manipulation Checks

Experts were sensitive to the Retaining Party and Interview Suggestibility manipulations. Within the prosecution-retained condition, experts were more likely to correctly identify the prosecution as the retaining party (88%) compared with the defense (12%), X2 (1, N = 50) = 28.88, p = .001, θ = .76. Within the defense-retained condition, experts were more likely to correctly identify the defense as the retaining party (84%) than the prosecution (16%), X2 (1, N = 50) = 23.12, p = .001, θ = .68. We included all experts in the final data analyses to ensure adequate statistical power and random assignment to experimental condition. We also viewed this as a more conservative test of adversarial allegiance than excluding potentially neutral experts who were not attentive to and/or concerned about the retaining party. In the end, including versus excluding experts based on this manipulation check did not affect the statistical significance of the main effects and interactions we observed—only the estimated size of the effects. Experts also were sensitive to the interview suggestibility manipulation as seen in the child victim accuracy and interview quality results we report shortly.

To ensure there were no systematic differences in how often experts testified for the prosecution or defense across all experimental conditions, we conducted a 2 Retaining Party (prosecution, defense) × 2 Interview Suggestibility (low, high) multivariate analysis of variance (MANOVA) on the self-report data from experts regarding how often they had previously testified for each side. Using the Pillai's Trace criterion multivariate statistic, there were no statistically significant differences for retaining party, Mult. F(2, 23) = .18, p = .84, partial η2 = .02; interview suggestibility, Mult. F (2, 23) = .05, p = .95, partial η2 = .01; or the interaction, Mult. F (2, 23) = 1.61, p = .22, partial η2 = .12, indicating random assignment was successful. Because the data were positively skewed, we performed a logarithmic transformation to satisfy the normality assumption and then re-ran the MANOVA. Once again, there were no statistically significant differences for retaining party, Mult. F(2, 23) = .04, p = .96, partial η2 = .001; interview suggestibility, Mult. F(2, 23) = .12, p = .89, partial η2 = .01; or the interaction, Mult. F (2, 23) = 3.11, p = .06, partial η2 = .21.

Data Analytic Approach

We conducted a series of two-way analyses of variance (ANOVAs) to examine the effects of Retaining Party and Interview Suggestibility on the willingness to testify, prosecution focus, defense focus, and perceived expert focus dependent measures. We conducted a two-way MANOVA on the child victim accuracy, child victim trustworthiness, and interview quality dependent measures and used the Pillai's Trace criterion multivariate statistic to test the significance of all main effects and interactions. We followed-up any significant multivariate effects with the appropriate univariate F-tests. See Tables 1 and 2 for all means and univariate effects.

Table 1. Means and Main Effects of Retaining Party and Interview Suggestibility on Dependent Measures.

Means (SD) Univariate Effects
Dependent Measures Main Effect Prosecution Defense Low Suggestibility High Suggestibility F df p d 95% CI
Willingness to Testify Retaining Party 5.50 (2.03) 5.68 (2.11) 0.21 1, 96 0.65 0.09
Interview Suggestibility 5.62 (2.15) 5.56 (2.00) 0.02 1, 96 0.88 0.03
Prosecution Focus Count Retaining Party 1.29 (1.79) 0.36 (0.79) 10.67 1, 83 < 0.001* 0.68 0.40, 0.96
Interview Suggestibility 1.44 (1.79) 0.25 (0.69) 17.36 1, 83 < 0.001* 0.89 0.61, 1.17
Prosecution Focus Proportion Retaining Party 0.47 (0.48) 0.14 (0.30) 15.05 1, 75 < 0.001* 0.84 0.75, 0.92
Interview Suggestibility 0.52 (0.47) 0.11 (0.28) 21.75 1, 75 < 0.001* 1.07 0.99, 1.16
Defense Focus Count Retaining Party 2.02 (2.58) 2.74 (1.95) 1.60 1, 83 0.21 0.32
Interview Suggestibility 1.21 (1.67) 3.50 (2.31) 26.69 1, 83 < 0.001* 1.15 0.73, 1.57
Defense Focus Proportion Retaining Party 0.53 (0.48) 0.86 (0.30) 15.05 1, 75 < 0.001* 0.84 0.75, 0.92
Interview Suggestibility 0.48 (0.47) 0.89 (0.28) 21.75 1, 75 < 0.001* 1.07 0.99, 1.16
Perceived Expert Focus Retaining Party -0.09 (.92) -0.62 (0.61) 11.20 1, 87 < 0.001* 0.69 0.53, 0.84
Interview Suggestibility -0.02 (.85) -0.70 (0.63) 20.01 1, 87 < 0.001* 0.92 0.76, 1.07
Child Victim Accuracy Retaining Party 4.03 (1.09) 3.80 (0.82) 1.54 1, 94 0.22 0.24
Interview Suggestibility 4.25 (0.87) 3.60 (0.97) 12.06 1, 94 < 0.001* 0.71 0.53, 0.89
Child Victim Trustworthiness Retaining Party 4.37 (1.12) 4.23 (0.99) 0.46 1, 94 0.50 0.13
Interview Suggestibility 4.44 (1.08) 4.17 (1.03) 1.47 1, 94 0.23 0.26
Police Interview Quality Retaining Party 3.88 (1.33) 3.56 (1.12) 2.05 1, 94 0.16 0.27
Interview Suggestibility 4.35 (1.06) 3.11 (1.08) 32.96 1, 94 < 0.001* 1.17 0.96, 1.38

Notes:

*

Difference between means statistically significant at p < .05. Willingness to Testify: 1 = Certainly No, 4 = Neutral, 7 = Certainly Yes. Prosecution Focus Count = Total number of pro-prosecution statements regarding anticipated testimony. Prosecution Focus Proportion = Total number of pro-prosecution statements divided by the total sum of pro-prosecution and pro-defense statements. Defense Focus Count = Total number of pro-defense statements regarding anticipated testimony. Defense Focus Proportion = Total number of pro-defense statements divided by the total sum of pro-prosecution and pro-defense statements. Perceived Expert Focus: -1 = Pro-defense, 0 = Neutral, 1 = Pro-prosecution. Child Victim Accuracy, Child Victim Trustworthiness, and Police Interview Quality: 1 = Not at All, 4 = Neutral, 7 = Extremely.

Table 2. Means, Retaining Party × Interview Suggestibility Interaction Effects, and Follow-ups on Dependent Measures.

Dependent Measure Prosecution Mean (SD) Defense Mean (SD) F df p d 95% CI
Willingness to Testify 11.45 1, 96 < 0.001*
Low Suggestibility 6.20 (1.41)a 5.04 (2.59)b 4.29 1, 96 0.04* 0.74 0.22, 1.27
High Suggestibility 4.80 (2.33)a 6.32 (1.25)b 7.37 1, 96 0.01* 0.83 0.32, 1.34
Prosecution Focus Count 11.25 1, 83 < 0.001*
Low Suggestibility 2.21 (1.96)a 0.47 (0.91)b 21.54 1, 83 < 0.001* 1.12 0.61, 1.17
High Suggestibility 0.24 (0.70) 0.26 (0.69) 0.004 1, 83 0.95 0.03
Prosecution Focus Proportion 11.27 1, 75 < 0.001*
Low Suggestibility 0.75 (0.41)a 0.19 (0.35)b 25.47 1, 75 < 0.001* 1.48 1.36, 1.60
High Suggestibility 0.13 (0.33) 0.09 (0.24) 0.14 1, 75 0.71 0.15
Defense Focus Count 1.87 1, 83 0.18
Low Suggestibility 0.71 (1.37) 1.84 (1.84) 3.43 1, 83 0.07 0.73
High Suggestibility 3.52 (2.84) 3.48 (1.76) 0.006 1, 83 0.94 0.03
Defense Focus Proportion 11.27 1, 75 < 0.001*
Low Suggestibility 0.25 (0.41)a 0.81 (0.35)b 25.47 1, 75 < 0.001* 1.48 1.36, 1.60
High Suggestibility 0.87 (0.33) 0.91 (0.35) 0.14 1, 75 0.71 0.15
Perceived Expert Focus 4.96 1, 87 0.03
Low Suggestibility 0.36 (0.81)a -0.45 (0.67)b 16.05 1, 87 < 0.001* 1.11 0.90, 1.32
High Suggestibility -0.62 (0.61) -0.78 (0.52) 0.59 1, 87 0.45 0.26

Notes: Means sharing different superscripts within each row were statistically significant at p < .05. Willingness to Testify: 1 = Certainly No, 4 = Neutral, 7 = Certainly Yes. Prosecution Focus Count = Total number of pro-prosecution statements regarding anticipated testimony. Prosecution Focus Proportion = Total number of pro-prosecution statements divided by the total sum of pro-prosecution and pro-defense statements. Defense Focus Count = Total number of pro-defense statements regarding anticipated testimony. Defense Focus Proportion = Total number of pro-defense statements divided by the total sum of pro-prosecution and pro-defense statements. Perceived Expert Focus: -1 = Pro-defense, 0 = Neutral, 1 = Pro-prosecution.

Hypothesis 1: Willingness to Testify

Neither the main effect of Retaining Party nor Interview Suggestibility was statistically significant. Experts were very willing to testify in the case whether asked by the prosecution or defense or whether interview suggestibility was low or high (Table 1). The interaction was statistically significant. Experts were more willing to testify for the prosecution than the defense in the low suggestibility interview condition; however, the opposite was true for the high suggestibility interview condition—experts were more willing to testify for the defense than the prosecution (Table 2).

Of note, there were nine experts who indicated they were “certainly” unwilling to testify in the case but completed the remaining dependent measures we report next. When we included those data in our analyses, the effects of Retaining Party and Interview Suggestibility were identical to when we excluded the experts who were unwilling to testify. Only the actual values for the test statistics, means, and effect size estimates varied slightly; therefore, we kept all experts in the final data analyses.

Hypotheses 2 & 3: Expert Testimony Focus

Actual Prosecution Focus Count and Proportion

The main effects of Retaining Party and Interview Suggestibility were statistically significant. Prosecution-retained experts provided a larger number and higher proportion of pro-prosecution statements about the police interview and child's memory than did defense-retained experts (Table 1). The same effect emerged for the low versus high suggestibility interview. These main effects were qualified by a statistically significant interaction. Prosecution-retained experts provided a larger number and higher proportion of pro-prosecution statements about the case than did defense-retained experts, but only when interview suggestibility was low (Table 2).

Actual Defense Focus Count

Only the main effect of Interview Suggestibility was statistically significant. Experts provided a larger number of pro-defense statements about the case when interview suggestibility was high versus low (Table 1).

Actual Defense Focus Proportion

The main effects of Retaining Party and Interview Suggestibility were statistically significant. Defense-retained experts provided a higher proportion of pro-defense statements about the police interview and child's memory than did prosecution-retained experts (Table 1). The same effect emerged for the high versus low suggestibility interview. These main effects were qualified by a statistically significant interaction. Defense-retained experts provided a higher proportion of pro-defense statements about the case than did prosecution-retained experts, but only when interview suggestibility was low (Table 2).

Perceived Expert Focus

The main effects of Retaining Party and Interviewer Suggestibility were statistically significant. Blind raters perceived more pro-defense focus in experts' anticipated testimony when experts were retained by the defense versus prosecution and in the high versus low interview suggestibility conditions (Table 1). These main effects were qualified by a statistically significant interaction. When interview suggestibility was low, raters perceived more pro-prosecution focus for prosecution-retained experts and more pro-defense focus for defense-retained experts. No differences emerged when interview suggestibility was high—raters perceived both experts were more pro-defense focused rather than pro-prosecution (Table 2).

Table 3 describes the topics experts most commonly reported when asked what aspects of the police interview and child's memory they would focus on if called to testify. Overall, experts focused on pro-defense aspects of the case three-times more often than pro-prosecution aspects. Experts most frequently cited the police officer, interview room, and interview itself (68%) whereas fewer discussed issues pertaining to the abuse act (13%), disclosure (11%), or the child victim (8%).

Table 3. Frequencies and Percentages for Expert Testimony Topics Reported.
Topic Category Topic Pro-Prosecution n (%) Pro-Defense n (%) Neutral n (%) Topic Category n Topic n
Officer & Interview Room 16 (17.0%) 74 (78.7%) 4 (4.3%) 94
Training 2 (15.4%) 10 (76.9%) 1 (7.7%) 13
Dress* 6 (18.8%) 25 (78.1%) 1 (3.1%) 32
Atmosphere* 7 (23.3%) 22 (73.3%) 1 (3.3%) 30
Mom's presence* 1 (5%) 17 (90%) 1 (5%) 19
Interview 33 (28.2%) 75 (64.1%) 9 (7.7%) 117
“Don't Know”* 4 (66.7%) 2 (33.3%) 0 (0%) 6
Answer all questions* 1 (16.7%) 5 (83.3%) 0 (0%) 6
Children's initial reluctance* 2 (28.6%) 5 (71.4%) 0 (0%) 7
Encouragement* 1 (16.7%) 5 (83.3%) 0 (0%) 6
Use of protocol* 4 (50%) 4 (50%) 0 (0%) 8
Question type* 20 (40.8%) 26 (53.1%) 3 (6.1%) 49
Age-appropriate questions 0 (0%) 7 (87.5%) 1 (12.5%) 8
Use of props 1 (20%) 3 (60%) 1 (20%) 5
Video recording 0 (0%) 12 (80%) 3 (20%) 15
Use of the word “penis” 0 (0%) 6 (85.7%) 1 (14.3%) 7
Disclosure 2 (8.0%) 16 (64.0%) 7 (28.0%) 25
Delay 0 (0%) 3 (42.9%) 4 (57.1%) 7
Prior disclosure 2 (28.6%) 2 (28.6%) 3 (42.9%) 7
Spontaneous disclosure 0 (0%) 11 (100%) 0 (0%) 11
Abuse 7 (20.6%) 12 (35.3%) 15 (44.1%) 34
History 2 (18.2%) 4 (36.4%) 5 (45.5%) 11
Act itself 4 (25%) 6 (37.5%) 6 (37.5%) 16
Witnesses 1 (14.3%) 2 (28.6%) 4 (57.1%) 7
Child 5 (22.7%) 14 (63.6%) 3 (13.6%) 22
Age 2 (15.4%) 9 (69.2%) 2 (15.4%) 13
Memory 3 (33.3%) 5 (55.6%) 1 (11.1%) 9
Total 63 (22%) 191 (63%) 38 (15%) 292 292

Note:

*

denotes topic manipulated in the low versus high suggestibility police interview.

Topics coded with total frequencies < 5 do not appear, but included: Officer experience* (n = 1), Officer introduction* (n = 3), “Don't understand” interview instruction* (n = 3), Interview content (n = 2), Interview length (n = 4), Prior conversations about abuse (n = 3), Events precipitating disclosure involving family (n = 4), Events precipitating disclosure involving friends (n = 1), Child's demeanor (n = 4), Defendant's prior record (n = 3), Physical evidence (n = 4), Medical exam (n = 2), Forensic interview (n = 1), and Child's ability to distinguish truth/lie (n = 3).

Recall twelve topics from our coding scheme directly touched on aspects of the interview that we varied between the low and high suggestibility interview conditions. Experts mentioned interview question type the most often (15%) with slightly more of experts' responses focusing on the police officer's use of direct (pro-defense, 53%) versus open-ended (pro-prosecution, 41%) questions during the interview. Experts also commented on the police officer's attire (10%), interview atmosphere (9%), mom's presence during the interview (6%), and video recording (5%). Unlike interview question type, experts' focus on these topics was disproportionately pro-defense. They took note of when the police officer was in uniform, interviewed the child victim in an interrogation room with her mother present, and did not video record the interview. In contrast, experts focused least frequently on the police officer's prior experience (0.3%), her introduction to the child (1%), giving the child permission to say “I don't understand” (1%), and the content of the child's allegation (0.6%).

We also broke down the descriptive data by experimental condition to determine whether experts focused on different topics as a function of retaining party and interview suggestibility. See the Online Supplementary Material for a table that presents the 11 most frequently reported topics (60% of total) in descending order. We calculated the difference between the number of pro-prosecution statements minus the number of pro-defense statements for each topic in all four experimental conditions. Irrespective of retaining party, difference scores should be positive for the low suggestibility interview conditions (indicating more pro-prosecution topics reported) and negative for the high suggestibility conditions (indicating more pro-defense topics reported). Indeed, this was the case with one notable exception: defense-retained experts in the low interview suggestibility condition consistently made more pro-defense statements for the question type, mom's presence, act itself, video recording, training, age, spontaneous disclosure, memory, and use of protocol topics (9 of the 11 most frequently reported topics).

Hypothesis 4: Child Victim Accuracy, Trustworthiness, and Police Interview Quality

A two-way MANOVA on the Child Victim Accuracy, Trustworthiness, and Police Interview quality dependent measures revealed a statistically significant main effect for Interview Suggestibility only, Mult. F(3, 92) = 13.93, p = .001, partial η2 = .31. Neither the main effect of Retaining Party nor Retaining Party × Interview Suggestibility interaction was statistically significant [Mult. F(3, 92) = .81, p = .49, partial η2 = .03. and Mult. F(3, 92) = 1.35, p = .26, partial η2 = .04, respectively].

Child Victim Accuracy

Follow-up univariate tests revealed a statistically significant effect for Interview Suggestibility. Experts found the child victim to be less accurate when interview suggestibility was high versus low (Table 1).

Child Victim Trustworthiness

Experts' evaluations of the child victim's trustworthiness were neutral in both interview conditions and did not differ at a statistically significant level (Table 1).

Police Interview Quality

The main effect of Interview Suggestibility was statistically significant. Experts rated the quality of the police interview more negatively when suggestibility was high rather than low (Table 1).

Discussion

The purpose of this study was to determine whether retaining party and certain evidence features (namely, police interview suggestibility) influenced different aspects of expert testimony in a simulated child sexual abuse case. The data revealed that the answer to this question depends in part on the exact measures examined.

Hypothesis 1: Willingness to Testify

As predicted, experts asked by the prosecution were more willing to testify when interview suggestibility was low than experts asked by the defense and vice versa when interview suggestibility was high. This finding makes sense: Experts should be more willing to testify when they have something to say and believe their testimony will help jurors understand the evidence. Presumably this would be the case for prosecution experts reviewing a low suggestibility interview (“The interview was sound and does not raise concerns about the child's accuracy”) and for defense experts reviewing a high suggestibility interview (“The interview was unsound and raises concerns about the child's accuracy”). This notion dovetails with the legal concepts of “relevance” and “helpfulness” set forth in the Federal Rules of Evidence. Simply put, if the expert evidence in question could make a difference in the case (Rule 401) and aids jurors' fact-finding mission (Rule 702) then it should be admitted (Rule 402). Experts in our study may have perceived their testimony as being more relevant to the case and more helpful to jurors when the evidence favored the party soliciting their testimony than when it did not.

Hypotheses 2 & 3: Expert Testimony Focus

Experts focused more on pro-defense aspects of the case overall, but favored the retaining party only when interview suggestibility was low. Blind raters detected this bias. These results did not support our crossover interaction hypothesis. Prosecution-retained experts were more pro-prosecution (both in terms of the number and proportion of statements reported) and defense-retained experts were more pro-defense (proportion of statements only) when interview suggestibility was low but not high. This evidence of adversarial allegiance is consistent with previous research (Murrie et al., 2008, 2009, 2013; Otto, 1989; Zusman & Simon, 1983).

One potential theoretical explanation for the adversarial allegiance in our study is confirmation bias, which refers to “an inclination to retain, or a disinclination to abandon, a currently favored hypothesis” (Klayman, 1995, p. 386). A key determinant in this process is information gathering and assimilation. Wason's (1960) “rule discovery” paradigm demonstrated that people engage in positive testing strategies in which they search for evidence that confirms, rather than disconfirms, a current belief. Clinical students testing the diagnosis that a confederate “patient” lacked self-control recalled more behaviors that were consistent with low self-control, even though the patient exhibited an equal number of inconsistent and consistent behaviors (Strohmer & Shivy, 1994). Medical students and residents who reviewed patient case histories favored a suggested diagnosis more often and identified more physical features consistent with that diagnosis compared with a nonsuggested, but equally plausible, diagnosis (Leblanc, Brooks, & Norman, 2002).

Much like the clinicians and medical students in previous research, experts in our study engaged in a positive testing strategy consistent with confirmation bias. Yet clinical and medical students were provided a specific hypothesis to test (i.e., whether a patient had a particular medical or psychological condition), but experts in our study were not—they simply were asked by the prosecution or defense to review case materials and testify. This highlights a vexing aspect of adversarial allegiance that we touched on in the Introduction. Experts appear to have developed their own hypotheses and implemented a positive testing strategy that was influenced by retaining party even though they were not explicitly instructed to do so. This effect cannot be attributed to pre-existing individual differences (recall experts were randomly assigned to condition and there was no systematic difference in how often they had testified for the prosecution or defense in the past).

A more plausible explanation involves the nature and transparency of the adversarial legal system itself. Attorneys do not routinely seek testimony from experts that will undermine their cases. It seems reasonable to expect (and our data suggest) that an expert asked to testify for the prosecution infers that the task at hand is to find strengths in the police interview that support the child's accuracy whereas a defense expert infers just the opposite. If true, the adversarial allegiance effects observed here should be even larger in cases when attorneys make their preferred findings and conclusions explicitly known to experts.

Yet confirmation bias alone cannot entirely explain the adversarial allegiance that emerged in our study. Recall interview suggestibility moderated the effects of retaining party on what aspects of the police officer's interview and the child's memory that experts reported they would focus on if called to testify in the case. To better understand this novel finding, we must first take a step back and look at the expert testimony focus data more broadly. Experts, irrespective of retaining party and interview suggestibility, focused on pro-defense aspects of the case three-times more often than pro-prosecution aspects. Even though the proportion of pro-prosecution statements was larger for experts asked to testify by the prosecution (M = 0.47) than by the defense (M = 0.14), overall experts still focused on a higher proportion of pro-defense aspects of the case (M = 0.53 prosecution experts, M = 0.86 defense experts). Also the number of pro-defense statements prosecution- and defense-retained experts made did not differ at a statistically significant level.

These findings are consistent with a negativity bias (Rozin & Royzman, 2001) or a “bad is stronger than good” effect (Baumeister, Bratslavsky, Finkenauer, & Vohs, 2001) that researchers have observed in a variety of judgment and information processing tasks. In essence, negative stimuli attract more attention, receive greater weight in evaluations, and are recalled more frequently than positive stimuli (Meffert, Chung, Joiner, Waks, & Garst, 2006). Classic work by Fiske (1980) revealed that participants who viewed photographs of people paid more attention to negative than positive behaviors when forming impressions. A related study used a modified Stroop paradigm and presented participants with personality trait adjectives (Pratto & John, 1991). Bad traits attracted more attention than good traits as evidenced by slower reaction times on the Stroop task (naming the ink color of the trait displayed) and participants were twice as likely to remember bad traits rather than good. Different theoretical explanations for negativity bias exist, with those most relevant to the present study emphasizing the perceptual saliency, informativeness, and diagnosticity of negative over positive information (Skowronski & Carlson, 1989).

People's penchant for negative information helps explain why experts disproportionately focused on pro-defense aspects of the case. Pro-defense is synonymous with unreliable evidence, and in our simulated child sexual abuse case, the key evidence against the defendant was the police officer's interview and the child's memory. Experts retained by both sides appear to have been naturally inclined to focus on weaknesses rather than strengths in how the police officer interviewed the child victim and what she said in response. This negativity bias may have been enhanced by the fact that it is easier to define, and therefore pinpoint, examples of what constitutes a bad versus good interview. A single flaw (inadequate ground rules, no narrative practice, excessive direct questions) can compromise the quality of an entire interview, but an entire interview must be practically flawless to be considered good by some experts. In this sense, the deck is somewhat stacked against professionals who interview children and may be subject to cross-examination and opposing expert testimony at trial.

Experts' negativity bias, in conjunction with confirmation bias, provides one explanation for why adversarial allegiance only emerged when interview suggestibility was low. The high interview suggestibility condition contained (by design) blatant, fundamental violations of sound forensic interviewing practice. Here a positive testing strategy, whether fueled by pro-prosecution bias (“The interview is sound”) or pro-defense bias (“The interview is unsound,”) is relatively easy to reject or accept because there are so many flaws in the interview. Moreover, experts' negativity bias (“What's wrong with this interview?”) complemented the process resulting in similarly high proportions of pro-defense statements irrespective of retaining party. In other words, the flawed interview and experts' negativity bias was enough to outweigh any pro-prosecution bias: Bad Interview + Negativity Bias > Pro-Prosecution bias.

In contrast, the high quality of the low suggestibility interview actually complemented experts' pro-prosecution bias and was enough to overcome the negativity bias: Good Interview + Pro-Prosecution Bias > Negativity Bias. Prosecution-retained experts shifted their focus from pro-defense to pro-prosecution aspects of the case, but defense experts did not. They continued to provide a higher proportion of pro-defense statements even though the low suggestibility interview was better than the high suggestibility interview. Here Pro-Defense Bias + Negativity Bias > Good Interview. Why the different results for prosecution experts who read a bad interview and defense experts who read a good interview? The answer may lie in our earlier argument that judging an interview to be bad is inherently an easier, more objective task than judging an interview to be good. The latter almost always leaves room for improvement and hence criticism.

Hypothesis 4: Child Victim Accuracy, Trustworthiness, and Police Interview Quality

The data did not support our fourth hypothesis. Experts' evaluations of child victim accuracy and interview quality were more negative in the high versus low interview suggestibility condition only. None of the differences we observed for experts' evaluations of child victim trustworthiness were statistically significant.

These findings are more encouraging than we have discussed thus far. Even though adversarial allegiance emerged on actual and perceived expert testimony focus (process-oriented variables), experts did not evaluate the child's accuracy, trustworthiness, or police interview quality (outcome-oriented variables) in a manner that favored the retaining party. The lack of adversarial allegiance on these measures is consistent with previous research involving forensic evaluators' verdicts in criminal (Beckham et al., 1989) and civil (Otto, 1989) cases, but is at odds with work by Zusman and Simon (1983) and Murrie et al. (2008, 2009, 2013). These mixed results raise the question of why experts in our study, despite their initial retaining party bias, were able to render objective evaluations of the child and police interview whereas participants in other studies did not. We believe Wilson and Brekke's (1994) discussion of mental contamination and correction may provide some preliminary answers.

Mental contamination is defined as “the process whereby a person has an unwanted judgment, emotion, or behavior because of mental processing that is unconscious or uncontrollable” (Wilson & Brekke, 1994, p. 117). To avoid contaminated judgments, a person must be aware of the unwanted mental processing, be motivated to correct it, know the direction and extent of bias, and have sufficient control over the processing to overcome the bias. Differences on any one or more of these dimensions between our experts and participants in earlier research could explain the inconsistent results. For example, several of the studies examined experts' psychiatric evaluations (Zusman & Simon, 1983) and actuarial risk assessment measure scores (Murrie et al. 2008, 2009) in actual civil cases. Undoubtedly, the information processing demands for these experts were much higher compared with the experts in our study who read a brief stimulus. The increased processing demands may have limited experts' awareness or ability to control the bias in the real-world settings.

These and other differences aside, we recognize our conclusion that experts in our study were aware, motivated, knowledgeable, and able to control bias is largely, if not completely, inductive. It makes sense then to consider whether our study's methodological features may have influenced the lack of adversarial allegiance on the child accuracy and police interview quality measures. None of the usual suspects that call into question the validity of null effects (unsuccessful manipulation checks, low statistical power, or restricted response range) were present in our study, so we are confident our data reflect a true lack of differences and not statistical artifacts. The brevity of the stimulus may raise concerns about experimental demand characteristics or a socially desirable response bias. If experts realized the purpose of the study and sought to provide data that confirmed the existence of adversarial allegiance, it seems reasonable to expect that they would have done so across all measures, and they did not. Similarly, if experts modified their responses to appear socially desirable, it seems they would have masked their adversarial allegiance on all items, not just their evaluations of the child and police interview. We think a more plausible explanation is that even though adversarial allegiance affected experts' anticipated testimony focus, they were mindful of this bias and evaluated the child in a non-partisan manner by focusing exclusively on interview suggestibility.

Limitations

We must address certain limitations before considering the implications of our findings and directions for future research. First, despite our attempt to maximize the response rate by following Dillman's (2000) Total Design Method, more experts (48%) failed to respond in any way than those who provided complete responses (30%). Even though our final sample included responses from experts in 32 different states, to characterize these findings as representative of all or even most experts would be misleading. We do not know whether experts who responded differ in meaningful ways from those who did not respond.

Second, the ecological validity of our study was lower than past adversarial allegiance research using forensic evaluators in actual cases or simulations (Zusman & Simon, 1983; Murrie et al. 2008, 2009, 2013). Experts in our study read they were being sought for potential expert testimony by the prosecution or defense—they had no contact with actual attorneys, only read a brief summary of a police interview rather than viewing a complete video or transcript, and received a nominal payment. Moreover, experts in our study did not actually testify and were not subjected to cross-examination, which might provide an even stronger test of adversarial allegiance. Whether the effects we observed replicate in more ecologically valid research or generalize to actual cases remains to be seen.

Lastly, we realize experts in witness suggestibility cases are not allowed to testify about ultimate opinion issues such as whether a specific child is accurate or whether a particular police interview is sound. Nevertheless experts still form personal opinions about these issues and often they are asked to express their opinions to jurors in court through hypotheticals posed by the attorneys (Kovera et al., 1997). Hypotheticals, by design, mirror the case facts and therefore experts' opinions about actual versus hypothetical child victims and police interviews are virtually one in the same.

Implications and Future Research

Our adversarial allegiance results have implications for the social scientific and legal communities. No other published research has examined how evidence features might interact with retaining party to influence adversarial allegiance (Murrie & Boccaccini, 2015). We did and observed that adversarial allegiance influenced experts only when the evidence did not contain egregious errors. In this sense, the devil is in the evidence details, not just on the witness stand.

Future research examining different features of other types of evidence and experts with varying degrees of prior testimony for the prosecution, defense, or both will help advance the state of social science on adversarial allegiance. Our study also suggests that the distinction between process- and outcome-oriented variables is important. Focusing exclusively on one or the other in our study would have dramatically changed the conclusions we drew. Researchers should strive to include both types of variables in future work so that we can make more sophisticated conclusions about the effects of adversarial allegiance on how experts examine evidence, as well as what they conclude.

With respect to the legal community, our results suggest that overly simplistic conclusions about whether adversarial allegiance exists and why are dangerous. In reality, this phenomenon is quite complex and depends on myriad factors. Based on our data, we know that the evidence features and the type of measures matter; however, more work is needed before drawing definitive conclusions for legal professionals. Tentatively we can suggest to judges and attorneys that adversarial allegiance exists, that it can (but not always) influence how experts process evidence, and that it may be more likely in cases involving evidence that is not blatantly flawed. What is striking about this conclusion is that from a statistical standpoint, experts are more likely to encounter evidence that rests at the middle of the quality distribution than either extreme end. Completely good or completely bad police interviews of children are much less common than interviews that are “sort of” good or bad. From this perspective, adversarial allegiance is probably more common than previously thought.

The likely prevalence of adversarial allegiance highlights the potential role of legal safeguards in minimizing its undesirable effects. Procedural safeguards such as cross-examination, opposing expert testimony, and even the threat of prosecution in extreme cases should significantly reduce the presence of purposeful bias in the courtroom. Court-appointed experts, who testify on behalf of the court instead of one side or the other, may be one way to combat adversarial allegiance, yet research indicates they are used infrequently (Krafka et al., 2002) and that they may not be the panacea once thought (Mnookin, 2008).

Lastly, experts called to testify in child sexual abuse cases should be mindful of the potential for negativity bias and confirmation bias. These biases were present in our study and differentially influenced how experts retained by opposing parties—particularly the defense—examined the case evidence. Unlike prosecution-retained experts who shifted their focus from pro-defense aspects of the case in the high suggestibility interview to pro-prosecution aspects in the low suggestibility interview, defense experts did not. They continued to provide a higher proportion of pro-defense statements even though the low suggestibility interview was better than the high suggestibility interview. These data and our argument that negativity bias and confirmation bias lead to different outcomes based on evidence quality highlight the need for defense experts to be especially mindful of and vigilant against the tendency to evaluate evidence that is not blatantly flawed more negatively. Countering the negativity bias with a “positivity” bias—that is making a conscious effort to search for strengths, as well as weaknesses, in evidence—may be one way to achieve this goal.

That said, we must not overlook that experts in our study were able to correctly distinguish between a low versus high suggestibility police interview of a child and that adversarial allegiance did not significantly influence their evaluations of the child's accuracy and quality of the police interview. Experts' understanding of witness suggestibility and jurors' lack thereof (Buck et al., 2014; McAuliff & Kovera, 2007; Quas et al., 2005) demonstrates that expert testimony on these issues should satisfy the helpfulness requirement of FRE Rule 702 and therefore be admissible in court. Courts that routinely disallow witness suggestibility expert testimony on the grounds that it is not helpful to jurors would be wise to reconsider their reasoning accordingly.

Supplementary Material

Online Supplementary Table: Frequencies and Differences for Most Commonly Reported Expert Testimony Topics by Experimental Condition

Acknowledgments

The first author was supported by Award Number R15HD065651 from the Eunice Kennedy Shriver Institute of Child Health and Human Development during the writing of this manuscript. The content is solely the responsibility of the authors and does not necessarily represent the official views of the Eunice Kennedy Shriver Institute of Child Health and Human Development or the National Institutes of Health. The second author was supported by a Grant in-Aid from the American Psychology-Law Society and a grant from the Office of Graduate Studies at California State University, Northridge. Portions of this research were presented at the 2013 meeting of the American Psychology-Law Society in Portland, OR. The authors would like to thank Andrew Ainsworth, Joshua Lapin, and Scott Plunkett for their time and invaluable contributions to this research. We also appreciate Nick Scurich's insightful comments on an earlier draft. Any errors that remain are the fault of the authors.

Contributor Information

Bradley D. McAuliff, Department of Psychology, California State University, Northridge

Jeana L. Arter, Department of Psychology, California State University, Northridge

References

  1. American Psychological Association. Amendments to the 2002 ‘Ethical principles of psychologists and code of conduct’. American Psychologist. 2010;65:493. doi: 10.1037/a0020168. [DOI] [PubMed] [Google Scholar]
  2. American Psychological Association. Specialty guidelines for forensic psychology. American Psychologist. 2013;68:7–19. doi: 10.1037/a0029889. [DOI] [PubMed] [Google Scholar]
  3. Baumeister RF, Bratslavsky E, Finkenauer C, Vohs KD. Bad is stronger than good. Review of General Psychology. 2001;5:323–370. doi: 10.1037/1089-2680.5.4.323. [DOI] [Google Scholar]
  4. Beckham JC, Annis LV, Gustafson DJ. Decision making and examiner bias in forensic expert recommendations for not guilt by reason of insanity. Law & Human Behavior. 1989;13:79–87. [Google Scholar]
  5. Bottoms BL, Quas JA, Davis SL. The influence of interviewer-provided social support on children's suggestibility, memory, and disclosures. In: Pipe ME, Lamb ME, Orbach Y, Cederborg AC, editors. Child sexual abuse: Disclosure, delay, and denial. Mahwah, NJ: Lawrence Erlbaum; 2007. pp. 135–158. [Google Scholar]
  6. Brodsky SL. Testifying in court: Guidelines and maxims for the expert witness. 2nd. Washington DC: American Psychological Association; 2013. [Google Scholar]
  7. Buck JA, London K, Wright DB. Expert testimony regarding child witnesses: Does it sensitize jurors to forensic interview quality? Law & Human Behavior. 2011;35:152–164. doi: 10.1007/s10979-010-9228-2. [DOI] [PubMed] [Google Scholar]
  8. Buck JA, Warren AR. Expert testimony in recovered memory trials: Effects on mock jurors' opinions, deliberations and verdicts. Applied Cognitive Psychology. 2009;24:495–512. doi: 10.1002/acp.1569. [DOI] [Google Scholar]
  9. Buck JA, Warren AR, Bruck M, Kuehnle K. How common is “common knowledge” about child witnesses among legal professionals? Comparing interviewers, public defenders, and forensic psychologists with laypeople. Behavioral Sciences & the Law. 2014;32:867–883. doi: 10.1002/bsl.2150. [DOI] [PubMed] [Google Scholar]
  10. Ceci SJ, Bruck MB. Suggestibility of the child witness: A historical review and synthesis. Psychological Bulletin. 1993;113:403–439. doi: 10.1037/0033-2909.113.3.403. [DOI] [PubMed] [Google Scholar]
  11. Cooper J, Neuhaus IM. The “hired gun” effect: Assessing the effect of pay, frequency of testifying, and credentials on the perception of expert testimony. Law & Human Behavior. 2000;24:149–171. doi: 10.1023/A:1005476618435. [DOI] [PubMed] [Google Scholar]
  12. Dillman DA. Mail and internet surveys: The tailored design Method. New York: Wiley; 2000. [Google Scholar]
  13. Federal Rules of Evidence. St. Paul Minnesota: West: 2001. [Google Scholar]
  14. Fiske ST. Attention and weight in person perception: The impact of negative and extreme behavior. Journal of Personality & Social Psychology. 1980;38:889–906. doi: 10.1037/0022-3514.38.6.889. [DOI] [Google Scholar]
  15. Groscup JL, Penrod SD, Studebaker CA, Huss MT, O'Neil KM. The effects of Daubert on the admissibility of expert testimony in state and federal criminal cases. Psychology, Public Policy, & Law. 2002;8:339–372. doi: 10.1037/1076-8971.8.4.339. [DOI] [Google Scholar]
  16. Hagen MA. Whores of the court: The fraud of psychiatric testimony and the rape of American justice. New York, NY: Harper Collins; 1997. [Google Scholar]
  17. Klayman J. Varieties of confirmation bias. Psychology of Learning and Motivation. 1995;32:385–418. doi: 10.1016/S0079-7421(08)60315-1. [DOI] [Google Scholar]
  18. Kovera MB, Gresham AW, Borgida E, Gray E, Regan PC. Does expert psychological testimony inform or influence juror decision making? A social cognitive analysis. Journal of Applied Psychology. 1997;82:178–191. doi: 10.1037/0021-9010.82.1.178. [DOI] [PubMed] [Google Scholar]
  19. Kovera MB, McAuliff BD. The effects of peer review and evidence quality on judge evaluations of psychological science: Are judges effective gatekeepers? Journal of Applied Psychology. 2000;85:574–586. doi: 10.1037/0021-9010.85.4.574. [DOI] [PubMed] [Google Scholar]
  20. Krafka C, Dunn MA, Johnson MT, Cecil JS, Miletich D. Judge and attorney experiences, practices, and concerns regarding expert testimony in federal civil trials. Psychology, Public Policy, & Law. 2002;8:309–332. doi: 10.1037//1076-8971.8.3.309. [DOI] [Google Scholar]
  21. Laimon RL, Poole DA. Adults usually believe young children: The influence of eliciting questions and suggestibility presentations on perceptions of children's disclosures. Law & Human Behavior. 2008;32:489–501. doi: 10.1007/s10979-008-9127-y. [DOI] [PubMed] [Google Scholar]
  22. Lamb ME, Orbach Y, Hershkowitz I, Esplin P, Horowitz D. A structured interview protocol improves the quality and informativeness of investigative interviews with children: A review of research using the NICHD Investigative Interview Protocol. Child Abuse & Neglect. 2007;21:1201–1231. doi: 10.1016/j.chiabu.2007.03.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Leblanc VR, Brooks LR, Norman GR. Believing is seeing: The influence of a diagnostic hypothesis on the interpretation of clinical features. Academic Medicine. 2002;77:S67–S69. doi: 10.1097/00001888-200210001-00022. [DOI] [PubMed] [Google Scholar]
  24. Lyon TD. Investigative interviewing of the child. In: Duquette DN, Haralambie AM, editors. Child welfare law and practice. 2nd. Denver, CO: Bradford; 2010. pp. 87–109. [Google Scholar]
  25. McAuliff BD, Kovera MB. Estimating the effects of misleading information on witness accuracy: Can experts tell jurors something they don't already know? Applied Cognitive Psychology. 2007;21:849–870. doi: 10.1002/acp.1301. [DOI] [Google Scholar]
  26. Meffert MF, Chung S, Joiner AJ, Waks L, Garst J. The effects of negativity and motivated information processing during a political campaign. Journal of Communication. 2006;56:27–51. doi: 10.1111/j.1460-2466.2006.00003.x. [DOI] [Google Scholar]
  27. Mnookin J. A cross-disciplinary look at scientific truth: What's the law to do? Expert evidence, partisanship, and epistemic competence. Brooklyn Law Review. 2008;73:1009–1033. [Google Scholar]
  28. Murrie DC, Boccaccini MT. Adversarial allegiance among expert witnesses. Annual Reviews of Law and Social Science. 2015;11:37–55. doi: 10.1146/annurev-lawsocsci-120814-121714. [DOI] [Google Scholar]
  29. Murrie DC, Boccaccini MT, Johnson JT, Janke C. Does interrater (dis)agreement on psychopathy checklist scores in sexually violent predator trials suggest partisan allegiance in forensic evaluation? Law & Human Behavior. 2008;32:352–362. doi: 10.1007/s10979-007-9097-5. [DOI] [PubMed] [Google Scholar]
  30. Murrie DC, Boccaccini MT, Turner DB, Meeks M, Woods C, Tussey C. Rater (dis)agreement on risk assessment measure in sexually violent predator proceedings. Psychology, Public Policy & Law. 2009;15:19–53. doi: 10.1037/a0014897. [DOI] [Google Scholar]
  31. Murrie DC, Boccaccini MT, Guarnera LA, Rufino KA. Are forensic experts biased by the side that retained them? Psychological Science. 2013;24:1889–1897. doi: 10.1177/0956797613481812. [DOI] [PubMed] [Google Scholar]
  32. Nunez N, Gray J, Buck JA. Educative expert testimony: A one-two punch can affect jurors' decisions. Journal of Applied Social Psychology. 2012;42:535–559. doi: 10.1111/j.1559-1816.2011.00782.x. [DOI] [Google Scholar]
  33. Otto RK. Bias and expert testimony of mental health professionals in adversarial proceedings: A preliminary Investigation. Behavioral Sciences & the Law. 1989;7:267–273. doi: 10.1002/bsl.2370070210. [DOI] [Google Scholar]
  34. Pratto F, John OP. Automatic vigilance: The attention-grabbing power of negative social information. Journal of Personality & Social Psychology. 1991;61:380–291. doi: 10.1037/0022-3514.61.3.380. [DOI] [PubMed] [Google Scholar]
  35. Quas JA, Thompson WC, Clarke-Stewart KA. Do jurors ‘know’ what isn't so about child witnesses? Law & Human Behavior. 2005;29:425–503. doi: 10.1007/s10979-005-5523-8. [DOI] [PubMed] [Google Scholar]
  36. Rogers R. Ethical dilemmas in forensic evaluations. Behavioral Sciences & the Law. 1987;5:149–160. doi: 10.1002/bsl.2370050207. [DOI] [PubMed] [Google Scholar]
  37. Rozin P, Royzman EB. Negativity bias, negativity dominance, and contagion. Personality & Social Psychology Review. 2001;5:296–320. doi: 10.1207/S15327957PSPR0504_2. [DOI] [Google Scholar]
  38. Saywitz KJ, Lyon TD, Goodman GS. Interviewing children. In: Myers JEB, editor. The APSAC handbook on child maltreatment. 3rd. Newbury Park, CA: Sage; 2011. pp. 337–360. [Google Scholar]
  39. Skowronski JJ, Carlston DE. Negativity and extremity biases in impression formation: A review of explanations. Psychological Bulletin. 1989;105:131–142. doi: 10.1037/0033-2909.105.1.131. [DOI] [Google Scholar]
  40. Strohmer DC, Shivy VA. Bias in counselor hypothesis testing: Testing the robustness of counselor confirmatory bias. Journal of Counseling & Development. 1994;73:191–197. doi: 10.1002/j.1556-6676.1994.tb01735.x. [DOI] [Google Scholar]
  41. Wason PC. On the failure to eliminate hypotheses in a conceptual task. Quarterly Journal of Experimental Psychology. 1960;12:129–140. doi: 10.1080/17470216008416717. [DOI] [Google Scholar]
  42. Wilson TD, Brekke N. Mental contamination and mental correction: Unwanted influences on judgments and evaluations. Psychological Bulletin. 1994;116:117–142. doi: 10.1037/0033-2909.116.1.117. [DOI] [PubMed] [Google Scholar]
  43. Zusman J, Simon J. Difference in repeated psychiatric examinations of litigants to a lawsuit. American Journal of Psychiatry. 1983;140:1300–1304. doi: 10.1176/ajp.140.10.1300. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Online Supplementary Table: Frequencies and Differences for Most Commonly Reported Expert Testimony Topics by Experimental Condition

RESOURCES