Abstract
Our voice provides salient cues about how confident we sound, which promotes inferences about how believable we are. However, the neural mechanisms involved in these social inferences are largely unknown. Employing functional magnetic resonance imaging, we examined the brain networks and individual differences underlying the evaluation of speaker believability from vocal expressions. Participants (n = 26) listened to statements produced in a confident, unconfident, or “prosodically unmarked” (neutral) voice, and judged how believable the speaker was on a 4‐point scale. We found frontal–temporal networks were activated for different levels of confidence, with the left superior and inferior frontal gyrus more activated for confident statements, the right superior temporal gyrus for unconfident expressions, and bilateral cerebellum for statements in a neutral voice. Based on listener's believability judgment, we observed increased activation in the right superior parietal lobule (SPL) associated with higher believability, while increased left posterior central gyrus (PoCG) was associated with less believability. A psychophysiological interaction analysis found that the anterior cingulate cortex and bilateral caudate were connected to the right SPL when higher believability judgments were made, while supplementary motor area was connected with the left PoCG when lower believability judgments were made. Personal characteristics, such as interpersonal reactivity and the individual tendency to trust others, modulated the brain activations and the functional connectivity when making believability judgments. In sum, our data pinpoint neural mechanisms that are involved when inferring one's believability from a speaker's voice and establish ways that these mechanisms are modulated by individual characteristics of a listener. Hum Brain Mapp 38:3732–3749, 2017. © 2017 Wiley Periodicals, Inc.
Keywords: cerebellum, feeling of (un)knowing, functional connectivity, social inference, trustworthiness
INTRODUCTION
The human brain has evolved to recognize how the speaker communicates (e.g., with a lower pitch, louder, faster), what they communicate (e.g., a feeling of knowing or unknowing), and why (e.g., to establish trust, demonstrate knowledge, mitigate social commitment). Deciphering and inferring meaning from vocal and other nonverbal cues is of great importance to adaptive social functioning [Marsh and Blair, 2008; Hooker and Park, 2002; McCann and Peppé, 2003]. For example, judging how (un)believable another person is drives generosity and trust when making economic decisions [e.g., van't Wout and Sanfey, 2008] and the selection of cooperation partners in daily life [e.g., Cosmides and Tooby, 1992]. However, the neural mechanisms responsible for evaluating the interpersonal stance, attitudes, or mental state of a speaker from their voice have received little attention [Hensel et al., 2013; Mitchell and Ross, 2013; Pell, 2007]. This study aims to foster understanding of how social inferences are made from vocal cues by using functional magnetic resonance imaging (fMRI). Specifically, our goal was to illuminate the neural systems involved in evaluating speaker believability and trust, in contexts where the speaker provides explicit vocal cues signaling their level of confidence (confident or unconfident expressions) and in other ubiquitous social situations in which believability must be inferred from a “neutral” (prosodically unmarked) speaking voice [Hensel et al., 2013].
Neurocognitive Systems Supporting Social Inference From Communicative Signals
The neural mechanisms underlying social inference have been mainly investigated in the context of interpreting visual action representations (e.g., grasp a cup) or verbally described behaviors referring to task goals and long‐term intentions [e.g., prepare a wedding, see Van Overwalle, 2009; Van Overwalle and Baetens, 2009 for reviews; Van Overwalle and Vandekerckhove, 2013]. Current evidence suggests that two neurocognitive systems support social inference: the mentalizing network and the mirroring network [Van Overwalle, 2009; Van Overwalle and Baetens, 2009]. The mirroring network involves the reading of other person's nonverbal behaviors and movements, and recognizing the goal of the perceived action by matching it to a representation in memory of our own action [e.g., Keysers and Gazzola, 2007]. It allows one to rapidly and intuitively sense the other person's goals on the basis of low‐level behavioral input [e.g., Calvo‐Merino et al., 2005] and includes the anterior intraparietal sulcus (aIPS) and the premotor cortex (PMC) [Gallese et al., 2004; Keysers and Gazzola, 2007; Tunik et al., 2007]. The aIPS and PMC receive input from the superior temporal sulcus (STS), which serves to parse the motion or auditory input (speech) into a sequence of coherent and meaningful (temporal) units [Redcay, 2008].
The mentalizing (or “theory of mind”) network enables humans to understand mental states, in oneself or others, which underlie overt behavior [e.g., Amodio and Frith, 2006; Gallagher and Frith, 2003]. It requires the separation of one's own mental perspective from that of others. This system includes the posterior medial structure precuneus (PC) and the medial prefrontal area of mPFC, and the temporo‐parietal junction (TPJ). The TPJ seems crucial for the representation of temporary goals and intentions [Saxe and Powell, 2006; Van Overwalle and Baeten, 2009]. The mPFC is responsible for reasoning about actions and judgments, including goals and intentions [Keysers and Gazzola, 2007; Van der Cruyssen et al., 2009], and is more related to the attribution of enduring traits and qualities about the self and other people [Van Overwalle, 2009].
When believability inferences are made from nonverbal stimuli such as faces, Winston et al. [2002] discovered that the right STS is responsible for coding unbelievable faces when the believability level was judged explicitly, while the right amygdala and right insula were involved regardless of whether the participants' attention was guided to these features or not. Differential activations observed when attention was guided towards or away from the stimuli suggest that believability judgments engender an automatic, emotional process and intentional social inferences [Adolphs, 2002; see Bzdok et al., 2011a, 2011b for a review]. Recent work further demonstrated that when one evaluated a character's believability based on their dynamic facial/bodily expressions during the answer of a trivia question, the activity in the mentalizing network (TPJ; mPFC) was correlated [Kuhlen et al., 2015]. These findings point to neural circuitry that is likely involved in believability judgments, although few studies have used evidence from vocal behavior to understand how the mirroring and mentalizing networks are engaged when accessing a speaker's mental state and evaluating their believability.
Neural Circuits of Processing Sociocommunicative Meaning From Vocal Signals
Recent advancements in fMRI have revealed the neural pathways underlying the processing of vocal expressions of emotional signals [Frühholz et al., 2016; Schirmer and Kotz, 2006; Wildgruber et al., 2009]. The superior temporal gyrus (STG) registers the unfolding of relevant acoustic affective cues over time, and serves to integrate this information in the form of an auditory percept; the inferior frontal gyrus (IFG) then functions to integrate emotionally relevant sound features provided by the STG via dorsal and ventral connections, allowing sounds to be categorized according to their social meaning and affective weights [Frühholz and Grandjean, 2013b; Frühholz et al., 2015; Sammler et al., 2015; Schirmer and Kotz, 2006]. In addition, the mPFC is involved in vocal emotion processing; it supports emotional and social functions related to interpersonal communication and understanding, permitting increased emotional appraisal and evaluation processes (especially the dorsal‐caudal part) [Ethofer et al., 2009; Wildgruber et al., 2009]. Subcortical regions such as the basal ganglia (BG), especially the dorsal BG (caudate), are involved in temporal prediction and decoding of acoustic variation in vocal expressions to facilitate the decoding of emotional meaning [Kotz and Schwartze, 2010; Kotz et al., 2009; Pell and Leonard, 2003]. Recent studies have also suggested a link between cerebellar function [Van Overwalle et al., 2014; Schwartze and Kotz, 2016] and social perception of speech signals; for example, the cerebellum may temporally code the auditory input and prepare segmented information in speech for further integration [Schwartze and Kotz, 2016]. The cerebellum seems to be particularly activated in tasks that involve abstracting and mentalizing other's behaviors, in particular those described in terms of traits or permanent characteristics that are not concrete [Baetens et al., 2013; Van Overwalle et al., 2014].
However, as underscored earlier, vocal signals convey more than emotions. Studies on processing social meaning in the tone of voice have focused on vocal attractiveness [Bestelmeyer et al., 2012], warnings and hesitation [Hellbernd and Sammler, 2016; Sammler et al., 2015], politeness [Pell, 2007], confidence [Jiang and Pell, 2015; Pell, 2007], and sincerity [Rigoulot et al., 2014]. Clinical studies suggest a general role for the right hemisphere and BG in prosodic change that induces mental state processing [Monetta et al., 2008; Pell, 2007]. For example, Pell [2007] provided the first evidence that focal right hemisphere damage impairs the ability to infer graded meanings from vocally expressed confidence or politeness when evaluating recorded statements (see Monetta et al. [2008] for related insights from Parkinson's disease). Using a passive listening task, an fMRI study revealed that increased levels of vocal attractiveness (independently judged) were associated with increased activations in the left middle occipital cortex and bilateral fusiform; in contrast, decreased levels of vocal attractiveness were associated with BOLD signal increases in bilateral superior temporal gyrus (STG) and right inferior frontal gyrus (IFG) [Bestelmeyer et al., 2012]. In another study comparing responses to vocal expressions of basic emotion (e.g., happiness, anger) with more complex, mixed emotions (pride, guilt, boredom; Alba‐Ferrara et al. [2011]) from utterances that took the form of numbered lists, the mPFC, frontal operculum, and left insula were activated to a greater extent for the complex emotions.
One of the critical functions of vocal signals is to imply speaker trustworthiness and believability [Belin et al., 2004; Schweinberger et al., 2014], cues that can shape certain aspects of subsequent discourse [such as persuasion, Chebat and Hedhi, 2007; Gelinas‐Chebat et al., 1996]. This raises the question of what neural systems are recruited by confidence‐related vocal cues to generate believability inferences in the absence of visual information. Vocal cues that mark differences in a speaker's confidence level, or “feeling of knowing” as they speak, are known to have distinct acoustic configurations; high speaker confidence is associated with short and infrequent pauses, increased loudness and speaker rate, and typically a fall in intonation contour at the end of the utterance. In contrast, low speaker confidence (doubt) is associated with rising intonation and elevated pitch, changes in vocal quality, and more frequent pause [Jiang and Pell, 2017, 2015]. Listeners also routinely evaluate the believability of those who make statements in a relatively neutral tone, that is, from utterances which do not contain salient vocal cues that betray the speaker's feeling of knowing as they speak (prosodically unmarked utterances are comparatively faster and exhibit less pitch and loudness modulation than confidence expressions, Jiang and Pell, 2017]. Recent EEG studies show that the brain rapidly attunes to the value of explicit vocal cues in speech that refer to a speaker's confidence level, differentiating highly confident and unconfident voices in the first 200 ms of speech processing, emphasizing that vocal confidence expressions are highly salient and quickly incorporated into a speaker representation [Jiang and Pell, 2015, 2016a, 2016b]. Moreover, neural responses are qualitatively different and occur much later when listeners evaluate speaker confidence from the same statements produced in a prosodically unmarked tone, implying that this context for inferring speaker believability relies on distinct neurocognitive mechanisms [Jiang and Pell, 2015], although additional work using complementary neuroinvestigative methods is needed.
This Study
This study is the first to document the neural underpinnings that allow listeners to infer believability from vocal expressions, building on previous work that described vocal confidence using perceptual acoustic [Jiang and Pell, 2017], EEGs [Jiang and Pell, 2015, 2016a, 2016b], and neuropsychological approaches [Monetta et al., 2008; Pell, 2007]. Three principal questions were explored.
First, we sought new insights into the behavioral and neural substrates which differentiate vocal expressions of different levels of confidence in a believability inference task, given the intimate link between vocal expressed confidence and speaker believability [Buller & Buergoon, 1986; Demeure et al., 2011]. Here, listeners were asked to rate the speaker's believability from “prosodically marked” statements containing overt vocal expressions of confidence (confident, unconfident) and identical “prosodically unmarked” statements produced in a neutral manner that did not attempt to vocally validate or disconfirm the content of the utterance. Behaviorally, we predicted that both confident and neutral expressions would be rated as more believable than unconfident expressions [Jiang and Pell, 2017, 2015], and that the unmarked neutral statements might sound similar or equal in believability to confident statements, given that they were judged to be “close‐to‐confident” in a previous study [Jiang and Pell, 2015]. Evaluating prosodically marked (confident vs unconfident) expressions should lead to differential activations in the IFG and right STG [Frühholz and Grandjean, 2013b]. The right IFG participates in prosodic processing [Kotz and Paulmann, 2011; Sammler et al., 2015; Schirmer and Kotz, 2006], especially in tasks emphasizing explicit evaluation of vocal information [Wildgruber et al., 2009; Ethofer et al., 2009], whereas the right STG is known to detect acoustic changes underlying vocal expressions [Frühholz et al., 2016]. Although confident and prosodically unmarked neutral expressions should both promote impressions of high confidence and believability [Jiang and Pell, 2017, 2015], it is predicted that social impressions formed in each voice condition will implicate distinct neural circuitry. Unlike a confident voice, which provides salient cues that facilitate social believability inferences, neutral voices can be interpreted in various ways in different social contexts and should impose additional inferential demands [Hensel et al., 2013]. We therefore hypothesized that person mentalizing and mirroring networks [Van Overwalle, 2009] would be activated to a greater extent for prosodically unmarked versus vocal confidence expressions in our task.
The second objective of our study was to establish the neural networks that lead to (un)believability inferences about a speaker from vocal cues. To this aim, we undertook analyses that would identify regions parametrically associated with the participants' responses, either positively or negatively [Kuhlen et al., 2015]. We predicted that believability judgments would be strongly associated with the parietal network. As part of the dorsal attention network [Liao et al., 2010], the parietal regions are involved in action observation and imitation [Molenberghs et al., 2012], showing larger activation for statements judged to be more believable. For example, the left PoCG has been linked to lie perception, showing larger activation for statements judged to be less believable [Wu et al., 2011]. The BG, especially the bilateral caudate, may also play an important role in believability inferences; BG activity is known to increase according to one's trust bias toward a communicative partner (Black vs White individual) [Stanley et al., 2012] and functional impairment of BG negatively impacts on the derivation of speaker meaning from vocal cues [Monetta et al., 2011].
Third, we sought to characterize the role of mentalizing and mirroring networks in speech‐related believability judgments, as has been investigated in reference to facial expressions and bodily postures [Kuhlen et al., 2015]. In conditions that are inference‐demanding, we predicted that the circuitry activated by believability judgments would overlap with key brain regions engaged by these social inference networks. Moreover, individual differences that affect social interactions, such as one's interpersonal sensitivity to others [Davis, 1983] and tendency to trust [Rotter, 1967], are likely to predict networks for decoding vocal expressions that promote believability inferences [Frank et al., 2015; Jiang and Pell, 2016a, 2016b; Li et al., 2014]. Including analyses of individual differences will provide new insight regarding how personal characteristics relevant to inferring believability affect the neural systems for rendering these judgments based on vocal expressions.
METHODS
Participants
Twenty‐six right‐handed native English speakers participated in the fMRI study (See Table 1 for demographic profiles). No participants reported a history of neurological or psychiatric disorder, nor any serious medical condition. None reported a history of hearing or speech‐language disorder. Widely used personality inventories including an Interpersonal Reactivity Index Scale (IRI) [Davis, 1983] and Interpersonal Trust Scale (ITS) [Rotter, 1967] were administered after fMRI scanning. The IRI scale includes 28 items of 5‐point Likert Scale from “does not describe me well” to “describes me very well.” The four subscales measure cognitive (perspective‐taking, fantasy) and affective (empathic concern, personal distress) aspects of interpersonal sensitivity [Davis, 1983]. The ITS scale includes 25 items of 5‐pt Likert Scale from “strongly agree” to “strongly disagree.” The ITS measures one's expectation that the behavior, promises, or statements of other individuals can be relied upon. Items cover a range of social interactions with different individuals, including parents, sales people, the judiciary, people in general, political figures, and media. Most items deal with the credibility of social agents, but some cover general optimism about the future of society. The internal consistency (Cronbach's alpha) was 0.82 for IRI total score (M = 93.92 ± 12.25) and 0.81 for ITS (M = 79.56 ± 11.27). The study was performed according to the ethical procedure approved by the Montreal Neurological Institute and Hospital and the McGill University Faculty of Medicine. Written informed consent was obtained from all participants at study onset.
Table 1.
Demographic profiles of the participants in the fMRI study
| Listener Sex | Number of participants | Age (years) | Education Experience (years) | ||||
|---|---|---|---|---|---|---|---|
| Mean | Max | Min | Mean | Max | Min | ||
| Female | 14 | 22.00 | 30 | 18 | 15.29 | 18 | 13 |
| Male | 12 | 23.42 | 29 | 20 | 16.67 | 19 | 14 |
Materials and Designs
Ninety statements produced in a confident, unconfident, or prosodically unmarked voice by two Canadian‐English speakers (1 female and 1 male) were selected from a perceptually validated vocal database of confidence expressions [Jiang and Pell, 2014; 2017]. In previous studies, these recordings showed robust differences in both their acoustic profiles and neurophysiological responses [Jiang and Pell, 2015, 2016a, 2016b]. All stimuli were short personal‐knowledge statements (They are renting a cottage) or personal opinions (She'll do a good job). The intended confidence was validated by an independent group of listeners who did not participate in the fMRI study (Table 2). Another 180 stimuli produced in other English accents (Canadian‐French and Australian) were part of the design. All 270 stimuli were fully mixed and presented with the critical vocal expressions in each run. The results for these accented vocal expressions were not the focus of this study and are reported in a companion study (Jiang, Sanford, Pell, in preparation). Each participant received a different sequence of stimuli. The sequences were pseudorandomized such that no more than three consecutive trials were spoken in the same level of confidence. Following each vocal stimulus, the participants were asked to judge how believable the speaker was on a 4‐point scale from 1 “not at all believable” to 4 “very much believable.” The meaning of the scale was reversed in half of the participants and balanced by listener sex.
Table 2.
Validation speaker confidence rating and online speaker believability rating in each type of vocal expression
| Confidence Rating | Believability Rating | |||
|---|---|---|---|---|
| Mean | SD | Mean | SD | |
| Confident | 4.42 | 0.38 | 2.88 | 1.05 |
| Unconfident | 1.78 | 0.52 | 1.83 | 1.04 |
| Prosodically unmarked | 3.84 | 0.29 | 3.10 | 0.98 |
Scanning Parameters
Scanning was performed on the same 3‐T Siemens Imager with a 32‐channel head coil. Structural T1‐weighted anatomical images were first acquired for anatomical reference for each participant (repetition time (TR)/echo time (TE)/inversion time (TI) = 2300/2.98/900 ms, flip angle = 9°, and voxel size = 1 × 1 × 1 mm3). Six functional echo‐planar runs were then acquired for each participant. Each functional run contained 41 slices with 53 volumes with whole‐head interleaved acquisition (TR/TE = 8000/30 ms, flip angle = 90°, and voxel size: 3.5 × 3.5 × 3.5 mm3). Each functional scan used a sparse‐sampling paradigm, which minimizes the influence of the BOLD response due to scanner noise [Belin et al., [Link], 1999]. This paradigm takes advantage of the 4–6 s delay in the hemodynamic response peak following the stimulus [Gaab et al., 2007].
Task Procedure
Volumes were acquired every 8 s and lasted for 2.5 s after the presentation of the vocal expressions. Each vocal stimulus was presented during the silent periods between acquisitions. The vocal stimuli were preceded by a fixation of 0.5 s and followed by a 1 s rating scale, signaling the participant to rate how believable the speaker sounded for that trial. The participants were instructed to respond between the signaling screen and the fixation of the next trial. The vocal stimuli varied in length (M = 1.67 s, from 1.27 to 3.09 s). The onset of vocal stimuli was jittered such that the center point of each vocal expression was 4.25 s from the middle of the subsequent scanning period [Rodd et al., 2005]. Thirty null events were randomly mixed with the vocal events, including three null events at the start of each run [Bach et al., 2008; Fecteau et al., 2004]. Null events had the same trial length as the vocal stimuli, except that no sound was played. The whole scanning session consisted of six runs with the same stimuli composition, with each lasting about 7 min in its entirety.
Data Analysis
To analyze the fMRI data, tools from the Functional Magnetic Resonance Imaging of the Brain software library (FSL) were utilized [Smith et al., 2004]. The following preprocessing steps were applied to all scans for each participant across the runs: brain extraction [Smith, 2002]; motion correction using MCFLIRT [Jenkinson and Smith, 2001]; spatial smoothing with a 5 mm FWHM isotropic Gaussian kernel; and high‐pass temporal filtering at 1/100 Hz. For each run, the first three volumes were removed to allow for stabilization of magnetization. Each participant's fMRI scan was then linearly registered to their corresponding T1‐weighted structural MRI using a six‐parameter rigid‐body transformation. This was followed by nonlinear registration to the Montreal Neurological Institute (MNI)‐152 standard space [Mazziotta et al., 2001]. The resulting transformations were combined to transform all functional images into the MNI152 standard space. Owing to unacceptable registration quality to the standard space, one participant was removed from further analysis.
A whole‐brain general linear model (GLM)‐based statistical analysis of the blood‐oxygen‐level‐dependent signal (BOLD) was performed on a voxel‐wise basis [Worsley and Friston, 1995]. The first‐level analysis used FSL FILM with local autocorrelation correction for GLM time series analysis [Woolrich et al., 2001]. Two separate models were built to assess BOLD signal as a function of the confidence levels and participant's believability responses. The first model focused on the correlation between the BOLD signal and different levels of confidence. Each vocal stimulus was labeled as confident, unconfident or prosodically unmarked (i.e., neutral expression). These stimuli were modeled as durational events. Here, four contrasts were assessed in both directions (“Confident vs Unconfident,” “Confident vs prosodically unmarked,” “Unconfident vs prosodically unmarked,” and “Confident + Unconfident vs prosodically unmarked”). The second model regressed the BOLD signal against the participant's believability response. Activations in both positive and negative correlations were tested. The participant's responses were modeled as an impulse response at the onset of the vocal stimuli [Jiang and Pell, 2015]. For all models, a binary variable indicating a null event was included as a covariate. Each variable of interest was convolved with the double‐gamma hemodynamic response function (HRF).
As each participant had multiple sessions with similar stimuli composition, a second‐level analysis was performed to combine whole‐brain statistical maps from the first‐level GLM time series analysis for each participant. This utilized FSL FMRIB's Local Analysis of Mixed Effects (FLAME) to perform fixed effects analysis modeling [Beckmann et al., 2003; Woolrich et al., 2004]. Final group level analysis used FLAME to perform mixed‐effects analysis with automatic outlier deweighting to capture the mean group effect in both positive and negative directions, and the effect of personality, indexed by the ITS and IRI scores, on the underlying BOLD signals.
Psychophysiological Interaction (PPI) Analysis
The whole‐brain GLM analysis assessing the participant's believability response identified activity in right superior parietal lobule (SPL) and left posterior central gyrus (PoCG) as regions that were significantly associated with more and less believable judgments, respectively. Here, a psychophysiological interaction (PPI) analysis was conducted to assess the functional connectivity with the participant's believability response [O'Reilly et al., 2012]. We were interested in the neural networks that could be functionally connected to these regions and facilitate certain believability outcomes, for example, those involved in derivation of sociomeanings from vocal expressions (e.g., BG) [Monetta et al., 2011; Stanley et al., 2012].
To perform the PPI analysis, a 9‐mm‐diameter sphere was defined centered at the peak activations in the right SPL and left PoCG. The physiological activity (i.e., time series BOLD signal) from both regions was extracted, and the BOLD signal correlation with the participant's believability response was considered psychological regressor. The first‐level analysis found the interaction between the underlying physiological activity and psychological regressor. Similar to the whole‐brain GLM analysis, the second‐level analysis combined the statistical maps across runs for each participant, while the third‐level analysis captured the mean group effect and its correlations with the ITS and IRI scores.
For all whole‐brain GLM and PPI analysis, areas of significant activation were identified using cluster thresholding with a Z cutoff of 1.96. All models were corrected for multiple comparisons at P < 0.05 using Gaussian random field theory [Worsley et al., 2002]. Percentage signal changes for each condition in models which displayed significant effects are shown in the Supplementary Material (Supporting Information, Table S1).
Assessing Hypothesized Network and Region Involvement
To determine if the significant activations from the whole‐brain GLM and PPI analysis involved hypothesized mirroring and mentalizing networks, dorsal and ventral pathways in prosody perception, and anterior and posterior superior, medial and inferior temporal lobes, we defined masks based on coordinates from prior reports and meta‐analysis [Van Overwalle and Baetens, 2009; Van Overwalle et al., 2014; Sammler et al., 2015; Desikan et al., 2006]. Masks were defined as a sphere with a diameter of 9 mm centered at the reported coordinates for the networks and pathways, while temporal lobe regions were segmented based on the Harvard–Oxford Cortical Structural Atlas [Desikan et al., 2006]. All masks were applied to the corrected whole‐brain results, simply to see if hypothesized networks, pathways, and regions were activated under particular effects [Kuhlen et al., 2015; Poldrack, 2007]. Tables III‐V show the regions, along with their MNI coordinate, used in this study.
RESULTS
Behavioral Ratings of Speaker Believability
Linear mixed effects modeling revealed that unconfident statements had significantly lower believability ratings to confident (b = −1.00, t = −4.57, P < 0.0001, R 2 = 0.22) and prosodically unmarked (b = −1.23, t = −5.66, P < 0.0001, R 2 = 0.22, Table 2) expressions. Participant responses for confident and prosodically unmarked expressions did not significantly differ. Ratings of speaker confidence interacted with the interpersonal trust score (b = 0.04, t = 1.96, P = 0.05, R 2 = 0.18); ITS negatively predicted believability ratings for prosodically unmarked expressions (b = −0.03, t = −2.65, P = 0.02, R 2 = 0.21) and positively predicted the ratings for unconfident expressions (b = 0.03, t = 2.55, P = 0.02, R 2 = 0.23). These findings suggest that participants with less commitment of trusting others rated prosodically unmarked expressions as less believable, and unconfident expressions as more believable. No interaction was revealed between speaker confidence and Interpersonal Reactivity Index score.
fMRI Results
BOLD response to confident, unconfident, and prosodically unmarked expressions
The direct comparison between confident and unconfident expressions revealed significant activations in the left superior frontal gyrus (SFG), left inferior frontal gyrus (IFG) (orbitalis), and right supplementary motor area (SMA) (Fig. 1A). Significant activations in the right STG and right Heschl's Gyrus (HG) were revealed when unconfident was compared with confident expressions (Fig. 1B). Table 3 showed the activation details.
Figure 1.

Activation maps of the contrasts: (A) confident vs unconfident expressions; (B) unconfident vs confident expressions; (C) prosodically marked vs prosodically unmarked expressions; (D) unconfident vs prosodically unmarked expressions. All activations survived the threshold at cluster level z > 1.96, P < 0.05 (GRF corrected). [Color figure can be viewed at http://wileyonlinelibrary.com]
Table 3.
Activations of peaks for functional contrasts of different confident expressions and list of coordinates of regions of mask from prior reports and meta‐analysis and their involvement in the contrast
| Effects of expressed confidence | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Brain region | Number of voxels | z score | x | y | z | Region of mask | x | y | z |
| Confident vs unconfident expression | |||||||||
| L Frontal superior G. | 362 | 3.29 | −14 | 50 | 34 | ‐ | |||
| L Inferior orbital frontal G. | 168 | 2.98 | −46 | 32 | −2 | ‐ | |||
| R Supplementary motor area | 151 | 3.09 | 8 | −22 | 54 | ‐ | |||
| Unconfident vs confident expression | |||||||||
| R Heschl G. | 113 | 2.98 | 50 | −8 | 8 | ‐ | |||
| Unconfident vs prosodically unmarked expression | |||||||||
| R Superior temporal G. | 175 | 3.27 | 60 | −14 | −4 | ‐ | |||
| Prosodically marked vs unmarked expression | |||||||||
| L Medial orbital frontal G. | 260 | 3.3 | −16 | 36 | −6 | ‐ | |||
| R Inferior orbital frontal G. | 139 | 3.26 | 22 | 24 | −12 | ‐ | |||
| Prosodically unmarked vs marked expression | |||||||||
| R Cuneus | 1842 | 4.04 | 10 | −70 | 32 | L Cerebellum VIa | −11 | −70 | −36 |
| R Cerebellum | 566 | 3.51 | 26 | −38 | −24 | R Cerebellum VIa | 9 | −78 | −43 |
| L Intraparietal sulcusb | −33 | −46 | 34 | ||||||
| Prosodically unmarked vs confident expression | |||||||||
| R Cuneus | 1600 | 3.86 | 12 | −70 | 32 | R Anterior inferior parietal sulcusc | 39 | −44 | 49 |
| R Fusiform G./R Cerebellum | 446 | 3.08 | 30 | −46 | −14 | L Cerebellum VIa | −11 | −70 | −36 |
| Prosodically unmarked vs unconfident expression | |||||||||
| L Cuneus | 435 | 3.68 | −4 | −68 | 26 | ‐ | |||
| R Cerebellum | 331 | 3.26 | 24 | −36 | −24 | ‐ | |||
| L Cerebellum | 230 | 3.19 | −14 | −36 | −20 | Cerebellum VIa | −11 | −70 | −36 |
| L Paracentral L. | 150 | 3.25 | −18 | −20 | 64 | ‐ | |||
| Interaction effects with IRI total score | |||||||||
| Prosocially marked vs unmarked expression × increasing IRI total score | |||||||||
| R Caudate | 277 | 3.2 | 16 | 0 | 22 | ||||
| R Frontal Superior G. | 157 | 3.16 | 12 | 18 | 48 | ‐ | |||
When BOLD responses were compared between statements with overt, prosodically marked vocal cues (i.e., confident and unconfident) versus prosodically unmarked expressions, selective activations to vocally expressed confidence were observed in the left orbital medial prefrontal cortex (omPFC) and right IFG (orbitalis, Fig. 1C). When confident and unconfident expressions were separately compared with unmarked expressions, significant activations in the right superior temporal gyrus (STG) were found for unconfident versus prosodically unmarked expressions (Fig. 1D). No differences were revealed when confident and prosodically unmarked expressions were compared.
When prosodically unmarked statements were compared with marked expressions, significant activations were found in the medial temporo‐occipital regions, including bilateral calcarine and cuneus, and right cerebellum (Fig. 2A). These activations were part of the person mentalizing and mirroring networks (cerebellum VI) and prosody pathway (left intraparietal sulcus (IPS)). Comparing prosodically unmarked versus confident expressions revealed significant activations in the bilateral cuneus, right fusiform, and right cerebellum (Fig. 2B), whereas comparing prosodically unmarked versus unconfident expressions increased responses in the left cuneus, bilateral cerebellum (4/5), and left paracentral lobule (Fig. 2C). Significant effects were found in the person mentalizing network (left cerebellum VI) for both contrasts, mirroring network (right aIPS) when prosodically unmarked was compared to confident, and mirroring network (right cerebellum VI) when prosodically unmarked was compared to unconfident.
Figure 2.

Activation maps of the contrasts: (A) Prosodically unmarked vs. prosodically marked expressions; (B) prosodically unmarked vs confident expressions; (C) prosodically unmarked vs unconfident expressions. All activations survived the threshold at cluster‐level z > 1.96, P < 0.05 (GRF corrected). [Color figure can be viewed at http://wileyonlinelibrary.com]
BOLD Signal Correlation With Participant's Believability Response
When increased believability ratings were modeled, activations in the right SPL, extending to the right PoCG, were associated with increased speaker believability (Fig. 3A and Table 4). These activations overlapped with the aIPS, a region commonly associated with the mirroring network. PPI analysis assessed the regions functionally connected to the right SPL during believability judgment. Here, bilateral caudate, pallidum, ACC, and the right medial superior frontal gyrus (mSFG) were revealed to be correlated with the activity in the right SPL (Fig. 3B).
Figure 3.

Activation maps showing (A) regions surviving parametric analysis of increased (red) and decreased (blue) believability rating; (B) the functional connectivity surviving parametric analysis of decreased believability rating with left posterior central gyrus (PoCG) as the seed region; and (C) the functional connectivity surviving parametric analysis of increased believability rating using right superior parietal lobule (SPL) as seed region. All activations survived the threshold at cluster‐level z > 1.96, P < 0.05 (GRF corrected). [Color figure can be viewed at http://wileyonlinelibrary.com]
Table 4.
Activations of peaks for parametric effects of believability rating
| Parametric effects of believability rating | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Brain region | Number of voxels | Z score | x | y | z | Region of mask | x | y | z |
| Increasing believability rating | |||||||||
| R Superior parietal G. | 169 | 2.66 | 40 | −46 | 58 | R Anterior inferior parietal S.a | 39 | −40 | 49 |
| Decreasing Believability Rating | |||||||||
| L Post central G. | 181 | 2.69 | −40 | −22 | 58 | ‐ | |||
With less believable judgments, increased activations were observed in the left PoCG extending to the left SPL (Fig. 3a). Connectivity between left PoCG and bilateral SMA (extending into paracentral lobule) was found as believability decreased (Fig. 3C).
Individual Difference in Interpersonal Sensitivity Score (IRI) and Interpersonal Trust Score (ITS)
The contrast between prosodically marked and unmarked expressions revealed a positive interaction with IRI score in right caudate and SMA extending into right SFG, suggesting that the frontostriatal circuit is involved in evaluating overt confidence‐related vocal cues in those participants who display higher interpersonal sensitivity (Fig. 4A) (Table 5). Individual differences were shown for the functional connectivity that was positively and negatively correlated with the believability rating separately. The connectivity between right SPL and right caudate (extending into right IFG) was positively modulated by IRI score; participants who were more sensitive to interpersonal relationships showed enhanced connectivity to make more believable judgments (Fig. 4B). Applying masks onto this result revealed that these activations were part of the dorsal prosody pathway (right IFG). The connectivity between left PoCG and bilateral caudate was also positively modulated by IRI score, with those who displayed higher interpersonal sensitivity exhibiting enhanced connectivity to render judgments of reduced believability (Fig. 4C). The connectivity between left PoCG and bilateral ACC (which extended to left vmPFC) was positively modulated by ITS score; those who showed lower interpersonal trust displayed enhanced connectivity to render judgments of unbelievability (Fig. 4D). Applying masks onto this result revealed an activation in the mentalizing network (mPFC).
Figure 4.

Activation maps of (A) the regions surviving the contrast between prosodically marked and unmarked expression which were positively modulated by IRI total score; (B) the functional connectivity surviving the parametric analysis of increased believability with right SPL as seed region, which was modulated by IRI total score; (C,D) the functional connectivity surviving the parametric analysis of decreased believability with the left PoCG as the seed region, which was modulated by IRI total score (C) and ITS total score (D). All activations survived the threshold at cluster‐level z > 1.96, P < 0.05 (GRF corrected). [Color figure can be viewed at http://wileyonlinelibrary.com]
Table 5.
Activations of peaks for the PPI analysis of the parametric effects of believability rating, and the interaction between these effects and IRI and ITS scores
| Parametric effects of believability rating | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Brain region | Number of voxels | Z score | x | y | z | Region of mask | x | y | z |
| Increasing believability rating: R superior parietal G. as seed region | |||||||||
| L Pallidum/L Amygdala | 1088 | 3.21 | −22 | −2 | −2 | ‐ | |||
| R Anterior cingulate C./R medial superior frontal G. | 1046 | 3.03 | 4 | 52 | 24 | Medial Prefrontal cortexb | 0 | 53 | 22 |
| R Caudate | 717 | 3 | 3 | 16 | 6 | ‐ | |||
| Decreasing believability rating: L Post central G. as seed region | |||||||||
| Supplementary motor area | 1312 | 3.99 | 0 | −24 | 60 | ‐ | |||
| Interaction effects with IRI total score | |||||||||
| Increasing believability rating × increasing IRI total score: R superior parietal G. as seed region | |||||||||
| R Caudate | 845 | 3.25 | 24 | 22 | 18 | Right inferior frontal G.a | 42 | 8 | 31 |
| Decreasing believability rating × increasing IRI total score: L post central G. as seed region | |||||||||
| L Caudate | 2015 | 3.51 | −26 | 16 | 14 | ‐ | |||
| R Caudate | 1351 | 3.27 | 18 | −4 | 26 | ‐ | |||
| Interaction effects with ITS total score | |||||||||
| Decreasing believability rating × increasing ITS total score: L post central G. as seed region | |||||||||
| L Anterior cingulate C./L medial orbital frontal G. | 1734 | 3.63 | −2 | 42 | 2 | Medial prefrontal cortexb | 0 | 53 | 22 |
DISCUSSION
Using fMRI, we examined the neural correlates underlying speaker believability evaluations from vocal confidence expressions produced by Canadian speakers of English. Our statements expressing speaker confidence varied only in the tone of voice and all referred to personal (rather than shared) knowledge, which means that no semantic (factual) knowledge could be used to infer believability. Overall, statements produced in a confident voice were judged to be more believable than in an unconfident voice, whereas prosodically unmarked expressions—perceived as “close‐to‐confident” in a previous study [Jiang & Pell, 2015]—were rated to be just as believable as confident expressions. The dissociation of confident and unconfident expressions when evaluating believability confirms that vocal confidence cues provide evidence of the correctness or truth value of a speaker's statement and the reliability of a person [Scherer et al., 1973], whereas doubt (lack of confidence) is marked by cues that supply signs of untrustworthiness or lack of credibility [Kuhlen et al., 2015]. These findings emphasize that vocally expressed confidence is a salient source of information about speaker believability [Buller & Buergoon, 1986; Demeure et al., 2011].
The GLM analysis based on neural data revealed dissociated brain activations underlying different levels of vocally expressed confidence, with increased frontal activity associated with confident expressions, and increased temporal activities with unconfident expressions. In contrast, medial temporo‐occipital regions and cerebellum were engaged by prosodically unmarked statements versus those with overt vocal cues. The parametric analysis revealed the involvement of right SPL/PoCG in the increased believability inference, and that of left PoCG in the decreased believability inference, regardless of the form of the vocal stimulus. The PPI analysis revealed that interpersonal sensitivity modulated the connectivity between the bilateral PoCG and the BG (caudate). Below, we focus on four issues related to the evaluation of speaker believability: (1) frontal‐temporal networks for judging confident/unconfident voices; (2) role of cerebellum for judging prosodically unmarked voices for believability; (3) neural networks associated with believability inference; and (4) role of the BG and individual differences.
Inferring Believability From Confident/Unconfident Voices—Frontal‐Temporal Activities
Our first question is what neural networks function to decode vocal expressions of confidence. Increased activation in the left SFG and IFG (orbitalis) for confident expressions suggests the involvement of processes for salience and relevance detection and cognitive control mechanisms in the integration of these cues [e.g., Jiang and Pell, 2016b; Kouneiher et al., 2009]. A confident voice is particularly relevant to decoding speaker believability given that our task encouraged explicit evaluation of emotive information in the voice [Wildgruber et al., 2005; Ethofer et al., 2009], thus activating orbital lateral prefrontal lobe. Looking at scalp‐recorded EEGs, confident voices elicited an enhanced positive response at 200 ms [Jiang and Pell, 2015], suggesting early prioritization of attention to explicit vocal cues that signal a speaker's feeling of knowing vs. unknowing. Moreover, when listeners had to resolve incongruent messages that began with a confident vocal expression (I'm sure) followed by an unconfident statement (he has access to the building), a source‐localization analysis on ERP responses identified the left SFG and left pre‐SMA; this finding suggests that violating expectancies created by a confident voice requires a mental‐representation update and enhanced attentional control to allow perceptual adjustment [Frühholz and Grandjean, 2013a, 2013b; Jiang and Pell, 2016b]. Here, expressions with overt vocal cues of confidence were mixed with prosodically unmarked expressions, and as the latter were perceived as close to confident, this may well have increased the complexity of a perceptual decision [Jiang and Pell, 2015]. These findings argue for enhanced attentional effort in controlling the believability response and integrating vocal expressions with lexical context or task demand in the case of confident speech [Kotz et al., 2013; Schirmer and Kotz, 2006; Wildgruber et al., 2009]. Observed increases in right SMA activity for confident expressions could be linked to operations for speech act detection in the dorsal pathway of prosody perception [Sammler et al., 2015]; for example, overt vocal confidence cues may be interpreted as an act of verbal persuasion [Jiang and Pell, 2015]. Given that right SMA was engaged more by participants with higher Interpersonal Reactivity (IRI) scores, our data show that a person's interpersonal awareness predicts the extent to which they detect underlying speech acts, and possibly, other interpersonal meanings intended by vocal expressions [Ethofer et al., 2006; Grandjean et al., 2005; Kotz et al., 2013; Wildgruber et al., 2005].
The right STG and HG were more activated by unconfident expressions compared to confident and prosodically unmarked expressions, although no difference was found between the latter two types. The lack of difference between confident and prosodically unmarked expressions is consistent with current literature suggesting that, in the context of representing a speaker's feeling of knowing, prosodically unmarked voices often lead to similar conclusions about the perceived level of confidence and believability of speakers as explicit confident expressions [Jiang and Pell, 2015; Table 3). The right STG (and its adjacent Heschl's gyrus) encodes suprasegmental information (e.g., the intonation contour of vocal expressions) as salient and acoustically complex socioemotional events [Grandjean et al., 2005; Kotz et al., 2006; Schirmer and Kotz, 2006; Wildgruber et al., 2006]. Here, the right STG activity for unconfident expressions overlapped the anterior portion of the ventral pathway in prosody perception [Sammler et al., 2015], which analyses slow spectral changes such as melody or pitch [Schönwiesner et al., 2005; Zatorre and Belin, 2001]. Acoustically, the speaker's lack of knowledge evolves over a larger time‐scale and exhibits marked changes in acoustic features along this scale (e.g., deviations in speech rate, hesitations, and increased pitch height), thus primarily activating the STG during the integration of emotive vocal information.
Inferring Believability From Prosodically Unmarked Voices: Cerebellar Activity
Prosodically unmarked expressions, despite lacking explicit vocal cues of confidence [Jiang and Pell, 2014], were rated as “close‐to‐confident” in one study [Jiang and Pell, 2015] and just as believable as confident expressions in this experiment overall. However, data argue that these judgments arise from distinct neural processes when statements are produced in a prosodically unmarked tone. Prosodically unmarked expressions elicited an increased right‐lateralized, delayed positive deflection when compared to confident, unconfident, and close‐to‐confident voices [Jiang and Pell, 2015]; this suggests that when listeners evaluate confidence or believability, the speaker meaning is derived much later from prosodically unmarked statements in light of the mismatch between overt vocal features of confidence and task requirements. Here, prosodically unmarked expressions elicited greater activation in medial temporo‐occipital regions, including right fusiform and left middle occipital gyrus (MOG), and cerebellum, when compared to confident or unconfident expressions [Baetens et al., 2013; Bestelmeyer et al., 2012; Ma et al., 2013].
There could be two possible accounts for the bilateral cerebellar activity observed when prosodically unmarked and prosodically marked expressions were compared. First, the cerebellum could be engaged in behavioral/embodied responses to affective sounds, especially those that lead to involuntary, immediate motor responses [Frühholz et al., 2016; Laricchiuta et al., 2015; Zald and Pardo, 2002]. However, this idea runs counter to our data, which suggest that the prosodically unmarked voice reduces involuntary responses and places top–down demands on the listener to uncover the implication of a nonexpressive voice and its relevance to speaker believability. Jiang and Pell [2015] reported that prosodically unmarked utterances uniquely elicited a late positivity (beginning 900 ms after the onset of the vocal expression) and this neural response can be moderated by a listener's sex and their level of interpersonal sensitivity [Jiang and Pell, 2016a, 2016b]. These data underscore that a time‐consuming inference about the underlying sociopragmatic function of prosodically unmarked statements is required. An alternative, more compatible view ascribes a role for the cerebellum in high‐order social inference‐making [Van Overwalle et al., 2014]. The (posterior) cerebellar lobules are often recruited bilaterally or unilaterally in concert with cortical regions in studies utilizing salient socioemotional stimuli (such as emotional nonsense syllable pairs, Kotz et al. [2013]; see Stoodley and Schmahmann [2010] for a review). Van Overwalle has proposed that the cerebellum (especially the posterior portion) reflects the demand of construal level, with intensified cerebellar activity when construing persons or events with larger social distance (e.g., contrasting impersonal vs personal/self‐action) or temporal distance (e.g., contrasting distant vs close past events) [Salmi et al., 2010; Van Overwalle et al., 2014]. This activity also increases in situations which raise questions about an agent's intentions (e.g., when a communicative partner uses unfamiliar movements to reflect deception, Van Overwalle et al. [2014]).
We have predicted that the absence of prosodic marking may encourage mentalizing and mirroring networks to participate in the inferential process. In line with the second account, applying masks from previous studies corroborated that cerebellar activation in prosodically unmarked versus marked conditions overlaps with regions used in typical person mentalizing tasks [Van Overwalle et al., 2014]. Compared with confidence and doubt, prosodically unmarked voices do not contain marked acoustic cues that bias a feeling of another's knowing, which could impose greater demands on the cerebellum to abstract the speaker (un) believability. The observation of aIPS suggests further involvement of the mirror system to infer speaker believability from prosodically unmarked statements; this system provides rapid input about the goals and intentions of observed actions to the mentalizing system for conscious reflection [Van Overwalle and Baetens, 2009]. Many previous studies looking at how actions lead to inferences about high‐level goals have used verbal descriptions, which could limit the extent to which the mirror system has been implicated in such tasks [Yao et al., 2011; cf. Brück et al., 2014]. The mirroring network could be more involved in processing intentions and dispositions for which the meaning is derived against the context, such as sincerity and sarcasm [Cheang and Pell, 2008; Jiang and Pell, 2015, 2016b; Rigoulot et al., 2014], and “neutrality” in a believability judgment task here.
Neural Networks Associated With Increasing/Decreasing Judgments of Believability
Our parametric analysis modeling the believability responses tested whether parietal activations are modulated as listeners inferred increasing versus decreasing believability from the voice. Consistent with our hypothesis and preceding works, statements judged to be more believable led to increases in the right PoCG extending into the right SPL, whereas statements judged to be less believable increased activation in the left PoCG/left SPL. The right SPL has been reported in other tasks requiring social understanding, such as evaluating people [Hensel et al., 2013; Zysset et al., 2002], uncovering lies [Harada et al., 2009], detecting embarrassment and guilt [Takahashi et al., 2004], reasoning about other's minds [Lissek et al., 2008], and moral dilemmas [Bzdok, Schilbach et al., 2012]. As bilateral PoCG seems to be sensitive to voices with varying texture regularity or acoustic properties [Bestelmeyer et al., 2012], it is likely that right PoCG mediates the inference that speakers are more believable from vocal cues such as pitch height and spectrotemporal regularity [Bestelmeyer et al., 2012]. Interestingly, the left PoCG is often activated during tasks for recognizing the mismatch between what is said and what is fact [Jiang and Pell, 2016b] and during lie perception [Wu et al., 2011]. Given our finding that left PoCG, and its connectivity with bilateral SMA, increased as a function of decreasing interpersonal believability—i.e., when a speaker's voice casts doubt on the validity of a statement—it can be said that this region is crucial for recognizing speech acts with a disposition of lacking honesty or trustworthiness on the part of the speaker.
It might be questioned whether speaker‐related acoustic factors other than confidence expressions influenced speaker believability. Overall, male speakers were judged to be slightly more believable than female speakers, although only for confident expressions (b = 0.26, t = 3.07, P = 0.005, Supporting Information, Table S2). However, supplementary analyses show that this speaker bias could not explain the patterns associated with speaker believability when activation differences due to the sex of the voice were modeled in the bilateral temporal gyrus and in right superior frontal gyrus (Supporting Information, Fig. S1). Largely different neural networks were activated by pitch, intensity, and durational measures and by speaker believability (Supporting Information, Table S3 and Fig. S2). Although we found that activity in the right PoCG was negatively associated with certain acoustic features of the stimuli (pitch and voice quality), no significant mediation of the PoCG response was shown by these acoustic measures in relation to the believability judgment. The role of the bilateral parietal inferential network in building the representation of speaker characteristics based on acoustic cues awaits future examination.
Functional connectivity with these parietal activations provides additional information about the neurocognitive processes that influence the believability inference. As speakers are judged to be more believable, connectivity increased between the right SPL and the ACC and right mSFG. The existence of a linear relationship between right SPL and executive control regions, vital for initiating top–down attentional control and adjusting performance based on contextual demands [Botnivick et al., 2004], points to the gradual need for cognitive control in mediating circumstances when speakers are judged to be highly believable [Mitchell, 2013]. This may be especially true when social evaluations are made in a continuous (rather than binary) manner, as in our task. When a speaker was judged to be less believable, networks that mediate attentional resources during social reasoning (dorsal mPFC and ACC) were also more synchronized with the left PoCG, although this synchronization was dependent on an individual's tendency to trust others (e.g., trustworthiness and attractiveness) [Bzdok, Langner et al., 2012; Hensel et al., 2013]. Our data supply new evidence that these frontal mechanisms guide graded decisions about the meaning of vocal cues, while showing that the ACC synchronizes with the right SPL when a speaker sounds more believable, and with the left PoCG when a speaker sounds less believable. The observation that parietal regions responding to increased vs. decreased impressions of interpersonal believability are differentially lateralized awaits further elaboration and replication.
The application of masks from the meta‐analysis [Van Overwalle & Baetens, 2009] also confirmed that interactions between the mirroring and mentalizing networks vary according to whether speakers are judged to be more versus less believable. The more believable a speaker sounds, the more functionally synchronized were signals in the two networks (right aIPS and mPFC, respectively); however, only mentalizing regions (mPFC) were implicated in conditions when the speaker was judged to be less believable. Given that the mirroring and mentalizing systems arrive at social judgments in different ways (e.g., simulation of actions vs abstraction of durable traits), it is likely that forming a strong impression of interpersonal believability places greater demands on regions that support high‐level construal processes to infer durable personal characteristics of a speaker, using simulated knowledge of actual vocal behavior as input [Baetens et al., 2013; Van Overwalle and Baetens, 2009]. In contrast, demands on the mentalizing network when judging that a speaker is not believable only increased for participants who do not tend to trust others (low ITS score). It has been noted that self‐other discrepancies are generally larger for those who do not tend to believe others [Hoffman et al., 1996]. Pending additional work, our findings suggest that for individuals who display low interpersonal trust in their daily lives, judging that a speaker lacks credibility not only depends on accurate decoding of salient vocal cues that mark the speaker's lack of confidence, but also high‐level inferences about the underlying intentions of the speaker.
Individual Differences and the Role of Basal Ganglia
Investigating individual differences in our study sheds new light on how contextual details, such as personality characteristics, affect social inferences based on vocal speech cues. Our discussion underscores that individual factors, such as differences in interpersonal awareness and the propensity to trust others, critically modulate neural networks underlying believability judgments and social reasoning. Of note, we found that listeners who display higher interpersonal reactivity (IRI score) selectively engage connections between the right caudate/right IFG with the right SPL when making believability judgments, suggesting that access to communicative meanings from the voice depends on a listener's interpersonal awareness and general sensitivity to social cues and relations [Jiang and Pell, 2016a, 2016b; Sammler et al., 2015]. In vocal communication, the caudate is sensitive to temporal patterns in sound and provides temporal predictions that reinforce meaning [Frühholz et al., 2016; Pell and Leonard, 2003; Weninger et al., 2013], while the right IFG helps to form an integrated interpretation of prosodic patterns and their meaning as part of the dorsal pathway of prosody perception [Frühholz and Grandjean, 2013b; Sammler et al., 2015]. The right IFG in linguistic or nonlinguistic affective speech also connects temporal regions through ventral pathways [Friederici, 2011; Frühholz et al., 2015]. Our data imply that subcortical mechanisms for making predictions about the significance of temporal speech cues, and for weighing and integrating sound information in the IFC to make social judgments, are harnessed to a greater extent by listeners who report higher interpersonal sensitivity in their daily lives.
Recent EEG studies demonstrate that individuals who display higher interpersonal reactivity (IRI score) also exhibit a stronger delayed positivity (around 900–1600 ms postonset of speech) when processing utterances that encode a speaker's feeling of unknowing [Jiang and Pell, 2016a], conflicting messages in expressed certainty [Jiang and Pell, 2016b], or which lack specificity about who an uttered person refers to [Jiang and Zhou, 2015]. One common feature of these studies, which all report an increased response in listeners who are arguably more “socially aware,” is that contextual cues are available for the listeners to reconcile speaker meaning in face of processing difficulties. If one combines these observations with our current finding that listeners with higher interpersonal sensitivity recruit the right caudate/IFG to a greater extent, we can further speculate that the BG/frontal–striatal–dorsal system plays a central role in facilitating comprehension of socioemotional meanings in the voice that contribute to different facets of social perception and theory‐of‐mind inferences [Pell et al., 2014]. While functional damage to this system is known to impair social evaluation and perception of vocal cues [Dara et al., 2008; Monetta et al., 2008; Paulmann et al., 2011; Van Lancker and Sidtis, 1992], it appears that individuals who are highly attuned to social relations can actively recruit these mechanisms in service of complex social judgments such as interpersonal believability and trust [Stanley et al., 2012].
CONCLUSION AND FUTURE DIRECTION
In this study, we examined how a speaker's commitment toward what is being said (e.g., confidence and doubt) modulates the brain response when inferring speaker believability. We found that dissociated frontotemporal networks are engaged while processing confident and unconfident vocal expressions, while additional cerebellar regions were recruited when processing “prosodically unmarked” vocal expressions. When higher and lower believability judgments were made, the right SPL and left PoCG were more activated and collaborated more intimately with the regions potentially responsible for recognizing speech acts, deployment of attentional resources, and derivation of social meanings from vocal expressions. Our findings highlight the significance of the neural systems underlying social inference from the human voice. It should be noted that impressions of vocal confidence (based on perceptual data from our validation study) [Jiang and Pell, 2014] do not always map onto corresponding impressions of speaker believability; approximately 68% of confident expressions in our experiment were scored as highly believable, whereas 72% of unconfident expressions were scored as unbelievable. This emphasizes that inferring believability is at least partially distinct from the intended confidence level communicated by the speaker, as supported by arguments raised by observations in the imaging data.
Two modes of processing have been proposed to support social inference: the implicit process, which is inaccessible to consciousness and control; and the explicit, which is accessible to awareness, introspection, and control [Lieberman, 2007; Van Overwalle and Van Dekerckhove, 2013]. Neurocognitive evidence shows that the explicit and implicit evaluation of descriptions of a person's traits or personal disposition involve different temporal dynamics [Van der Cruiyssen et al., 2009; Van Duynslaeger et al., 2007] and engage similar neural regions [Hensel et al., 2013; Van Overwalle and Van Dekerckhove, 2013]. We employed an explicit task in which the personality traits of the speaker (believable or unbelievable) are judged based on implicit processing of vocal cues that signal a speaker's commitment towards a described fact or opinion, and possibly other speaker‐related features [Van Overwalle and Van Dekerckhove, 2013]. Arguably, listeners conduct an evaluation of speaker confidence as an initial strategy using the frontal–temporal network in prosody perception, which is followed by further evaluation of believability in the functional network involving connections between the medial prefrontal and parietal regions. This process is dependent on a listener's interpersonal sensitivity, tendency to trust, and general attitudes toward people around them. This hypothesis awaits verification by future research that combines EEG/fMRI, that uses effective connectivity approaches such as dynamic causal modeling [Van Ackeren et al., 2016], and in tasks involving different levels of construal as listeners infer a speaker's believability. Examining how networks for evaluating speaker believability are engaged when listening to individuals with atypical or pathological speaking behaviors (e.g., foreign‐accented speech or motor speech disorders) [Wilson, 2015] also merits future attention.
Supporting information
Supporting Information
ACKNOWLEDGMENTS
This study was funded by a James McGill Professor award to M. D. Pell, and the McGill McLaughlin Scholarship to X. Jiang. We thank Kelly Hennegan and Dr Pan Liu for their precious help in testing participants, and Michael Ferreira for his technical consultation. The authors declare no conflict of interest regarding the publication of this article.
Contributor Information
Xiaoming Jiang, Email: xiaoming.jiang@mail.mcgill.ca.
Marc D. Pell, Email: marc.pell@mcgill.ca.
REFERENCES
- Adolphs R (2002): Trust in the brain. Nature 5:192–193. [DOI] [PubMed] [Google Scholar]
- Alba‐Ferrara L, Hausmann M, Mitchell R, Weis S (2011): The neural correlates of emotional prosody comprehension: Disentangling simple from complex emotion. PLoS One 6:e28701. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Amodio D, Frith C (2006): Meeting of minds: The medial frontal cortex and social cognition. Nat Rev Neurosci 7:268–277. [DOI] [PubMed] [Google Scholar]
- Bach D, Grandjean D, Sander D, Herdener M, Strik W, Seifritz E (2008): The effect of appraisal level of processing of emotional prosody in meaningless speech. NeuroImage 42:919–927. [DOI] [PubMed] [Google Scholar]
- Baetens K, Ma N, Steen J, Van Overwalle F (2013): Involvement of the mentalizing network in social and non‐social high construal. Soc Cogn Affect Neurosci 9:817–824. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Beckmann CF, Jenkinson M, Smith SM (2003): General multilevel linear modeling for group analysis in fMRI. Neuroimage 20:1052–1063. [DOI] [PubMed] [Google Scholar]
- P Belin, R Zatorre, R Hoge, A Evans, B Pike (1999): Event‐related fMRI of the auditory cortex. NeuroImage 10:417–429. [DOI] [PubMed] [Google Scholar]
- Belin P, Fecteau S, Bédard C (2004): Thinking the voice: Neural correlates of voice perception. Trends Cogn Sci 8:129–135. [DOI] [PubMed] [Google Scholar]
- Bzdok D, Langner R, Hoffstaedter F, Turetsky B, Zilles K, Eickhoff S (2012): The modular neuroarchitecture of social judgments on faces. Cereb Cortex. 22:945–961. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bestelmeyer P, Latinus M, Bruckert L, Crabbe F, Belin P (2012): Implicitly perceived vocal attractiveness modulates prefrontal cortex activity. Cereb Cortex 22:1263–1270. [DOI] [PubMed] [Google Scholar]
- Buller D, Buergoon J (1986): The effects of vocalics and nonverbal sensitivity on compliance: A replication and extension. Hum Commun Res 13:126–144. [Google Scholar]
- Brück C, Kreifelts B, Göβling‐Arnold C, Wertheimer J, Wildgruber D (2014): ‘Inner voices’: The cerebral representation of emotional voice cues described in literary texts. Soc Cogn Affect Neurosci 9:1819–1827. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bzdok D, Langner R, Hoffstaedter F, Turetsky B, Zilles K, Eickhoff S (2012): The modular neuroarchitecture of social judgments on faces. Cereb Cortex. 22:945–961. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bzdok D, Langner R, Caspers S, Kurth F, Habel U, Zilles K, Laird A, Eickhoff S (2011a): ALE meta‐analysis on facial judgments of trustworthiness and attractiveness. Brain Struct Funct 215:209–223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bzdok D, Langner R, Hoffstaedter F, Turetsky B, Zilles K, Eickhoff S (2011b): The modular neuroarchitecture of social judgments on faces. Cereb Cortex 22:951–961. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bzdok D, Schilbach L, Vogeley K, Schneider K, Laird A, Langner R, Eickhoff S (2012): Parsing the neural correlates of moral cognition: ALE meta‐analysis on morality, theory of mind, and empathy. Brain Struct Funct 217:783–796. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Calvo‐Merino B, Glaser D, Grèzes J, Passingham R, Gaggard P (2005): Action observation and acquired motor skills: An fMRI study with expert dancers. Cereb Cortex 15:1243–1249. [DOI] [PubMed] [Google Scholar]
- Cheang H, Pell M (2008): The sound of sarcasm. Speech Commun 50:366–381. [Google Scholar]
- Chebat J, Hedhi K (2007): Voice and persuasion in a banking telemarketing context. Percept Mot Skills 104:419–437. [DOI] [PubMed] [Google Scholar]
- Chen JL, Rae C, Watkins K (2012): Learning to play a melody: An fMRI study examining the formation of auditory‐motor associations. Neuroimage 59:1200–1208. [DOI] [PubMed] [Google Scholar]
- Cosmides L, Tooby J (1992): Cognitive adaptations for social exchange In: Barkow J, Cosmides L, Tooby J, editors. The Adapted Mind: Evolutionary Psychology and the Generation of Culture. New York: Oxford University Press. [Google Scholar]
- Dara C, Monetta L, Pell MD (2008): Vocal emotion processing in Parkinson's disease: Reduced sensitivity to negative emotions. Brain Res 1188:100–111. [DOI] [PubMed] [Google Scholar]
- Davis M (1983): Measuring individual differences in empathy: Evidence for a multidimensional approach. J Pers Soc Psychol 44:113–126. [Google Scholar]
- Demeure V, Niewiadomski R, Pelachaud C (2011): How is believability of a virtual agent related to warmth, competence, personification, and embodiment. Presence 20:431–448. [Google Scholar]
- Desikan RS, Ségonne F, Fischl B, Quinn BT, Dickerson BC, Blackler D, Buckner RL, Dale AM, Maguire RP, Hyman BT, Albert MS, Killiany RJ (2006): An automated labeling system for subdividing the human cerebral cortex on MRI scans into gyral based regions of interest. Neuroimage 31:968–980. [DOI] [PubMed] [Google Scholar]
- Ethofer T, Anders S, Erb M, Herbert C, Wiethoff S, Kissler J, Grodd W, Wildgruber D (2006): Cerebral pathways in processing of affective prosody: A dynamic causal modeling study. NeuroImage 30:580–587. [DOI] [PubMed] [Google Scholar]
- Ethofer T, Van De Ville D, Scherer K, Vuilleumier P (2009): Decoding of emotional information in voice‐sensitive cortices. Curr Biol 19:1028–1033. [DOI] [PubMed] [Google Scholar]
- Fecteau S, Armony J, Joanette Y, Belin P (2004): Is voice processing species‐specific in human auditory cortex? An fMRI study. NeuroImage 23:840–848. [DOI] [PubMed] [Google Scholar]
- Frank C, Baron‐Cohen S, Ganzel B (2015): Sex differences in the neural basis of false‐belief and pragmatic language comprehension. NeuroImage 105:300– 311. [DOI] [PubMed] [Google Scholar]
- Friederici A (2011): The brain basis of language processing: From structure to function. Physiol Rev 91:1357–1392. [DOI] [PubMed] [Google Scholar]
- Frühholz S, Gschwind M, Grandjean D (2015): Bilateral dorsal and ventral fiber pathways for the processing of affective prosody identified by probabilistic fiber tracking. NeuroImage 109:27–34. [DOI] [PubMed] [Google Scholar]
- Frühholz S, Grandjean D (2013a): Amygdala subregions differentially respond and rapidly adapt to threatening voices. Cortex 49:1394–1403. [DOI] [PubMed] [Google Scholar]
- Frühholz S, Grandjean D (2013b): Multiple subregions in superior temporal cortex are differentially sensitive to vocal expressions: A quantitative meta‐analysis. Neurosci Biobehav Rev 37:24–35. [DOI] [PubMed] [Google Scholar]
- Frühholz S, Trost W, Kotz S (2016): The sound of emotions – Towards a unifying neural network perspective of affective sound processing. Neurosci Biobehav Rev 68:96–110. [DOI] [PubMed] [Google Scholar]
- Gaab N, Gabrieli JDE, Glover GH (2007): Assessing the influence of scanner background noise on auditory processing. I. An fMRI study comparing three experimental designs with varying degrees of scanner noise. Hum Brain Mapp 28:703–720. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gallagher H, Frith C (2003): Functional imaging of ‘theory of mind’. Trends Cogn Sci 7:77–83. [DOI] [PubMed] [Google Scholar]
- Gallese V, Keysers C, Rizzolatti G (2004): A unifying view of the basis of social cognition. Trends Cogn Sci 8:396–403. [DOI] [PubMed] [Google Scholar]
- Grandjean D, Sander D, Pourtois G, Schwartz S, Seghier M, Scherer K, Vuilleumier P (2005): The voices of wrath: Brain responses to angry prosody in meaningless speech. Nat Neurosci 8:145–146. [DOI] [PubMed] [Google Scholar]
- Harada T, Itakura S, Xu F, Lee K, Nakashita S, Saito D, Sadato N (2009): Neural correlates of the judgment of lying: A functional magnetic resonance imaging study. Neurosci Res 63:24–34. [DOI] [PubMed] [Google Scholar]
- Hellbernd N, Sammler D (2016): Prosody conveys speaker's intentions: Acoustic cues for speech act perception. J Mem Lang 88:70–86. [Google Scholar]
- Hensel L, Bzdok D, Müller V, Zilles K, Eickhoff S (2013): Neural correlates of explicit social judgments on vocal stimuli. Cereb Cortex 25:1152–1162. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hoffman E, McCabe K, Smith V (1996): Social distance and other‐regarding behavior in dictator games. Am Econ Rev 86:653–660. [Google Scholar]
- Hooker C, Park S (2002): Emotion processing and its relationship to social functioning in schizophrenia patients. Psychiatry Res 112:41–50. [DOI] [PubMed] [Google Scholar]
- Jenkinson M, Smith SM (2001): A global optimization method for robust affine registration of brain images. Med Image Anal 5:143–156. [DOI] [PubMed] [Google Scholar]
- Jiang X, Pell MD (2014): Encoding and decoding confidence information in speech Proceedings of the 7th International Conference in Speech Prosody (Social and Linguistic Speech Prosody) 576−579. [Google Scholar]
- Jiang X, Pell MD (2017): The sound of confidence and doubt. Speech Commun 88:106–126. [Google Scholar]
- Jiang X, Pell MD (2015): On how the brain decodes vocal cues about speaker confidence. Cortex 66:9–34. [DOI] [PubMed] [Google Scholar]
- Jiang X, Pell MD (2016a): Neural responses towards a speaker's feeling of (un)knowing. Neuropsychologia 81:79–93. [DOI] [PubMed] [Google Scholar]
- Jiang X, Pell MD (2016b): The feeling of another's knowing: How “mixed messages” in speech are reconciled. J Exp Psychol Hum Percept Perform 42:1412–1428. [DOI] [PubMed] [Google Scholar]
- Jiang X, Zhou X (2015): Who is respectful? Effects of social context and individual empathic ability on ambiguity resolution during utterance comprehension. Front Psychol 6:1588. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Keysers C, Gazzola V (2007): Integrating simulation and theory of mind: From self to social cognition. Trends Cogn Sci 11:194–196. [DOI] [PubMed] [Google Scholar]
- Kotz S, Kalberlah C, Bahlmann J, Friederici A, Haynes J (2013): Predicting vocal emotion expressions from the human brain. Hum Brain Mapp 34:1971–1981. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kotz S, Meyer M, Paulmann S (2006): Lateralization of emotional prosody in the brain: An overview and synopsis on the impact of study design. Prog Brain Res 156:286–294. [DOI] [PubMed] [Google Scholar]
- Kotz S, Paulmann S (2011): Emotion, language and the brain. Lang Linguist Compass 5:108–125. [Google Scholar]
- Kotz S, Schwartze M (2010): Cortical speech processing unplugged: A timely subcortico‐cortical framework. Trends Cogn Sci 14:392–399. [DOI] [PubMed] [Google Scholar]
- Kotz S, Schwartze M, Schmidt‐Kassow M (2009): Non‐motor basal ganglia functions: A review and proposal for a model of sensory predictability in auditory language perception. Cortex 45:982–990. [DOI] [PubMed] [Google Scholar]
- Kouneiher F, Charron S, Koechlin E (2009): Motivation and cognitive control in the human prefrontal cortex. Nat Neurosci 12:939–945. [DOI] [PubMed] [Google Scholar]
- Kuhlen AK, Bogler C, Swerts M, Haynes J‐D (2015): Neural coding of assessing another person's knowledge based on nonverbal cues. Soc Cogn Affect Neurosci 10:729–734. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Laricchiuta D, Petrosini L, Picerni E, Cutuli D, Iorio M, Chiapponi C, Caltagirone C, Piras F, Spalletta G (2015): The embodied emotion in cerebellum: A neuroimaging study of alexithemia. Brain Struct Funct 220:2275–2287. [DOI] [PubMed] [Google Scholar]
- Liao W, Chen H, Feng Y, Mantini D, Gentili C, Pan Z, Ding J, Duan X, Qiu C, Lui S, Gong Q, Zhang W (2010): Selective aberrant functional connectivity of resting state networks in social anxiety disorder. NeuroImage 52:1549–1558. [DOI] [PubMed] [Google Scholar]
- Li S, Jiang X, Yu H, Zhou X (2014): Cognitive empathy modulates the processing of pragmatic constraints during sentence comprehension. Soc Cogn Affect Neurosci 9:1166– 1174. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lieberman M (2007): Social cognitive neuroscience: A review of core processes. Annu Rev Psychol 58:259–289. [DOI] [PubMed] [Google Scholar]
- Lissek S, Peters S, Fuchs N, Witthaus H, Nicolas V, Tegenthoff M, Juckel G, Brüne M (2008): Cooperation and deception recruit different subsets of the theory‐of‐mind network. PLoS One 3:e2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ma N, Baetens K, Vandekerckhove M, Kestemont J, Fias W, Van Overwalle F (2013): Traits are represented in the medial prefrontal cortex: An fMRI adaptation study. Soc Cogn Affect Neurosci 9:1185–1192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Marsh A, Blair R (2008): Deficits in facial affect recognition among antisocial populations: A meta‐analysis. Neurosci Biobehav Rev 32:454–465. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mazziotta J, Toga A, Evans AC, Fox P, Lancaster J, Zilles K, Woods R, Paus T, Simpson G, Pile B, Holmes C, Collins DL, Thompson PM, MacDonald D, Iacoboni M, Schormann T, Amunts K, Palomero‐Gallagher N, Geyer S, Parsons L, Narr K, Kabani N, Goualher GL, Boomsma D, Cannon T, Kawashima R, Mazoyer B (2001): A probabilistic atlas and reference system for the human brain: International Consortium for Brain Mapping (ICBM). Philos Trans R Soc Lond B Biol Sci 356:1293–1322. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McCann J, Peppé S (2003): Prosody in autism spectrum disorders: A critical review. Int J Lang Commun Disord 38:325–350. [DOI] [PubMed] [Google Scholar]
- Mitchell R (2013): Further characterisation of the functional neuroanatomy associated with prosodic emotion decoding. Cortex 49:1722–1732. [DOI] [PubMed] [Google Scholar]
- Mitchell R, Ross E (2013): Attitudinal prosody: What we know and directions for future study. Neurosci Biobehav Rev 37:471–479. [DOI] [PubMed] [Google Scholar]
- Molenberghs P, Cunnington R, Mattingley J (2012): Brain regions with mirror properties: A meta‐analysis of 125 human fMRI studies. Neurosci Biobehav Rev 36:341–349. [DOI] [PubMed] [Google Scholar]
- Monetta L, Cheang HS, Pell MD (2008): Understanding speaker attitudes from prosody by adults with Parkinson's disease. J Neuropsychol 2:415–430. [DOI] [PubMed] [Google Scholar]
- O'Reilly JX, Woolrich MW, Behrens TEJ, Smith SM, Johansen‐Berg H (2012): Tools of the trade: Psychophysiological interactions and functional connectivity. Soc Cogn Affect Neurosci 7:604–609. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Paulmann S, Ott D, Kotz S (2011): Emotional speech perception unfolding in time: The role of the basal ganglia. PLoS One 6: e17694. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pell MD (2007): Reduced sensitivity to prosodic attitudes in adults with focal right hemisphere brain damage. Brain Lang 101:64–79. [DOI] [PubMed] [Google Scholar]
- Pell MD, Leonard C (2003): Processing emotional tone from speech in Parkinson's disease: A role for the basal ganglia. Cogn Affect Behav Neurosci 3:275–288. [DOI] [PubMed] [Google Scholar]
- Pell MD, Monetta L, Rothermich K, Kotz S, Cheang H, McDonald S (2014): Social perception in adults With Parkinson's disease. Neuropsychology 28:905–916. [DOI] [PubMed] [Google Scholar]
- Poldrack R (2007): Region of interest analysis for fMRI. Soc Cogn Affect Neurosci 2:67–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Redcay E (2008): The superior temporal sulcus performs a common function for social and speech perception: Implications for the emergence of autism. Neurosci Biobehav Rev 32:123–142. [DOI] [PubMed] [Google Scholar]
- Rigoulot S, Fish K, Pell MD (2014): Neural correlates of inferring speaker sincerity from white lies: An event‐related potential source localization study. Brain Res 1565:48– 62. [DOI] [PubMed] [Google Scholar]
- Rodd J, Davis M, Johnsrude I (2005): The neural mechanisms of speech comprehension: fMRI studies of semantic ambiguity. Cereb Cortex 15:1261–1269. [DOI] [PubMed] [Google Scholar]
- Rotter J (1967): A new scale for the measurement of interpersonal trust. J Personal 35:651–665. [DOI] [PubMed] [Google Scholar]
- Sammler D, Grosbras M, Anwander A, Bestelmeyer PEG, Belin P (2015): Dorsal and ventral pathways for prosody. Curr Biol 25:3079–3085. [DOI] [PubMed] [Google Scholar]
- Saxe R, Powell L (2006): It's the thought that counts specific brain regions for one component of theory of mind. Psychol Sci 17:692–699. [DOI] [PubMed] [Google Scholar]
- Scherer K, London H, Wolf J (1973): The voice of confidence: Paralinguistic cues and audience evaluation. J Res Pers 7:31–44. [Google Scholar]
- Schirmer A, Kotz S (2006): Beyond the right hemisphere: Brain mechanisms mediating vocal emotional processing. Trends Cogn Sci 10:24–30. [DOI] [PubMed] [Google Scholar]
- Smith S (2002): Fast robust automated brain extraction. Hum Brain Mapp 17:143–155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smith SM, Jenkinson M, Woolrich MW, Beckmann CF, Behrens TE, Johansen‐Berg H, Bannister PR, Luca MD, Drobnjak I, Flitney DE, Niazy RL, Saunders J, Vickers J, Zhang Y, Stefano ND, Brady JM, Matthews PM (2004): Advances in functional and structural MR image analysis and implementation as FSL. NeuroImage 23:S208–S219. [DOI] [PubMed] [Google Scholar]
- Schönwiesner M, Rübsamen R, von Cramon D (2005): Spectral and temporal processing in the human auditory cortex—Revisited. Ann N Y Acad Sci 1060:89–92. [DOI] [PubMed] [Google Scholar]
- Schwartze M, Kotz S (2016): Contributions of cerebellar event‐based temporal processing and preparatory function to speech perception. Brain Lang 161:28–32. [DOI] [PubMed] [Google Scholar]
- Schweinberger S, Kawahara H, Simpson A, Skuk V, Zäske R (2014): Speaker perception. Cogn Sci 5:15–25. [DOI] [PubMed] [Google Scholar]
- Stanley D, Soko‐Hessner P, Fareri D, Perino M, Delgado M, Banaji M, Phelps E (2012): Race and reputation: Perceived racial group trustworthiness influences the neural correlates of trust decisions. Philos Trans R Soc Lond B Biol Sci 367:744–753. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stoodley C, Schmahman J (2010): Evidence for topographic organization in the cerebellum of motor control versus cognitive and affective processing. Cortex 46:831–844. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Takahashi H, Yahata N, Koeda M, Matsuda T, Asai K, Okubo Y (2004): Brain activation associated with evaluative processes of guilt and embarrassment: an fMRI study. NeuroImage 23:967–974. [DOI] [PubMed] [Google Scholar]
- Van Ackeren M, Smaragdi A, Rueschemeyer S (2016): Neuronal interactions between mentalizing and action systems during indirect request processing. Soc Cogn Affec Neurosci 11:1402–1410. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Van der Cruyssen L, Van Duynslaeger M, Cortoos A, Van Overwalle F (2009): ERP time course and brain areas of spontaneous and intentional goal inferences. Soc Neurosci 4:165–184. [DOI] [PubMed] [Google Scholar]
- Van Duynslaeger M, Van Overwalle F, Verstraeten E (2007): Electrophysiological time course and brain areas of spontaneous and intentional trait inferences. Soc Cogn Affect Neurosci 2:174– 188. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Van Lancker D, Sidtis J (1992): The identification of affective prosodic stimuli by left and right hemisphere‐damaged subjects. J Speech Hear Res 35:963–970. [DOI] [PubMed] [Google Scholar]
- Van Overwalle F (2009): Social cognition and the brain: A meta‐analysis. Hum Brain Mapp 30:829–858. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Van Overwalle F, Baetens K (2009): Understanding others' actions and goals by mirror and mentalizing systems: A meta‐analysis. Neuroimage 48:564–584. [DOI] [PubMed] [Google Scholar]
- Van Overwalle F, Baetens K, Mariën P, Vandekerchkhove M (2014): Social cognitive and the cerebellum: A meta‐analysis of over 350 fMRI studies. Neuroimage 86:554–572. [DOI] [PubMed] [Google Scholar]
- Van Overwalle F, Vandekerckhove M (2013): Implicit and explicit social mentalizing: Dual processes driven by a shared neural network. Front Hum Neurosci 7:560. Article [DOI] [PMC free article] [PubMed] [Google Scholar]
- Van 't Wout M, Sanfey A (2008): Friend or foe: The effect of implicit trustworthiness judgments in social decision making. Cognition 108:796–803. [DOI] [PubMed] [Google Scholar]
- Weninger F, Eyben F, Schuller B, Mortillaro M, Scherer K (2013): On the acoustics of emotion in audio: What speech, music and sound have in common. Front Psychol 4:292. Article [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wildgruber D, Ackermann H, Kreifelts B, Ethofer T (2006): Cerebral processing of linguistic and emotional prosody: fMRI studies. Prog Brain Res 156:249–268. [DOI] [PubMed] [Google Scholar]
- Wildgruber D, Ethofer T, Grandjean D, Kreifelts B (2009): A cerebral network model of speech prosody comprehension. Int J Speech Lang Pathol 11:277–281. [Google Scholar]
- Wildgruber D, Riecker A, Hertrich I, Erb M, Grodd W, Ethofer T, Ackermann H (2005): Identification of emotional intonation evaluated by fMRI. NeuroImage 24:1233–1241. [DOI] [PubMed] [Google Scholar]
- Wilson L (2015): Evaluating the believability of standardized patients portraying aphasia Master Thesis. University of Washington. [Google Scholar]
- Winston J, Strange B, O'Doherty J, Dolan R (2002): Automatic and intentional brain responses during evaluation of trustworthiness of faces. Nat Neurosci 5:277–283. [DOI] [PubMed] [Google Scholar]
- Woolrich MW, Behrens TEJ, Beckmann CF, Jenkinson M, Smith SM (2004): Multilevel linear modelling for fMRI group analysis using Bayesian inference. Neuroimage 21:1732–1747. [DOI] [PubMed] [Google Scholar]
- Woolrich MW, Ripley BD, Brady M, Smith SM (2001): Temporal autocorrelation in univariate linear modeling of fMRI data. Neuroimage 14:1370–1386. [DOI] [PubMed] [Google Scholar]
- Worsley K, Liao C, Aston J, Petre V, Duncan G, Morales F, Evans A (2002): A general statistical analysis for fMRI data. NeuroImage 15:1–15. [DOI] [PubMed] [Google Scholar]
- Worsley KJ, Friston KJ (1995): Analysis of fMRI time‐series revisted‐Again. Neuroimage 2:173–181. [DOI] [PubMed] [Google Scholar]
- Wu D, Loke IC, Xu F, Lee K (2011): Neural correlates of evaluations of lying and truth‐telling in different social contexts. Brain Res 1389:115–124. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yao B, Belin P, Scheepers C (2011): Silent reading of direct versus indirect speech activates voice‐selective areas in the auditory cortex. J Cogn Neurosci 23:3146–3152. [DOI] [PubMed] [Google Scholar]
- Zald D, Pardo J (2002): The neural correlates of aversive auditory stimulation. NeuroImage 16:746–753. [DOI] [PubMed] [Google Scholar]
- Zatorre R, Belin P (2001): Spectral and temporal processing in human auditory cortex. Cereb Cortex 11:946–953. [DOI] [PubMed] [Google Scholar]
- Zysset S, Huber O, Ferstl E, von Cramon D (2002): The anterior frontomedian cortex and evaluative judgment: An fMRI study. Neuroimage 15:983–991. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supporting Information
