Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Feb 17;16:9388. doi: 10.1038/s41598-026-39668-0

Behavioral and neural aftermath of right inferior frontal cortex disruption on ambiguous vocal emotion decisional processes

Leonardo Ceravolo 1,2,✉, Marius Moisa 3,4, Didier Grandjean 1,2, Christian Ruff 2,4, Sascha Frühholz 4,5,6
PMCID: PMC13003003  PMID: 41702981

Abstract

The evaluation of socio-affective auditory information is accomplished by the primary auditory cortex in collaboration with limbic regions and the inferior frontal cortex (IFC)—the latter often recruited during affective voice classification. IFC activity was observed either for coding sensory or perceptual ambiguity or representing more exclusively categorical processes such as object and affect information. Here, we presented ‘clear’ (or ‘pure’) and emotionally ‘ambiguous’ affective speech to two groups of human participants and asked them to explicitly categorize these voices’ emotion. During the task, wholebrain functional magnetic resonance imaging was acquired and combined with transcranial magnetic stimulation, specifically continuous theta-burst stimulation (cTBS; target site: right IFC pars triangularis) between the first and subsequent task runs. Contrary to our hypotheses, offline right IFC disruption through cTBS did not reveal improved accuracy to classify ‘ambiguous’ voices, and instead led to reduced auditory cortical activity and increased fronto-limbic connectivity for ‘clear’ affective speech. Although the expected impact of right IFC cTBS on behavioral performance did not occur, our imaging data still indicate activity patterns modified by the cTBS procedure. This was especially true in the right anterior superior temporal lobe, parietal cortex and superior frontal cortex. Due to the absence of conclusive behavioral results, we discuss our data and study limitations to hopefully allow for clearer design and methods for future work on the important topic of psychological and neural processes underlying vocal emotion decisional processes.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-39668-0.

Keywords: Affective neuroscience, Speech, Voice, Inferior frontal cortex, Decision making

Subject terms: Neuroscience, Psychology, Psychology

Introduction

The neural processing and classification of acoustic information involves a cortical neural network beyond the mere auditory decoding in the auditory cortex (AC)1. The core of this neural network especially for decoding acoustic and social information conveyed by acoustic voice signals involves an integrated functioning of the AC together with various regions in the inferior frontal cortex (IFC)2,3. These regions are especially relevant for extracting socio-affective information from voice signals4–9. In the latter case, this auditory-frontal network is often accompanied by neural activity in the limbic system, primarily—but not exclusively—located in the amygdala10,11.

Within this broad auditory-frontal-limbic network for voice signals processing and voice information decoding12–14, the involvement of the IFC in decoding emotions from affective speech has been consistently reported3,9,14, but its functional role in the neural network and its functional contribution to voice signal processing and especially in vocal affect decoding remains debated. Recent studies highlighted the functional relevance of the left IFC to elaborate evaluations and classification of socio-affective information in voices15,16. These processes are supposed to happen down-stream to the acoustical analysis in the AC2,17,18 with different bilateral IFC subregions coding for the complexity of the evaluation and classification task12,13,19. According to these studies, the bilateral IFC would represent rather domain-general task, evaluation, and classification difficulty, and would show increased activity for any challenging (affective) evaluation and classification tasks. This notion would support the related view on the IFC as a region that exerts some top-down control on the AC depending on certain task requirements3,12 and explicit sound classifications demands3. The latter seems especially relevant when classifying ambiguous voices and affect information20, which requires that acoustical analyses in AC are supplemented by functional support by neural processing resources of the IFC9, bilaterally. In fact, emotionally ambiguous voices roughly contain the same acoustic information as non-ambiguous voices yet they are perceived differently21,22. This view is also coherent with recent work on decisional processes taking place in the bilateral IFC pars triangularis when human participants had to categorize non-human primate vocalizations23 that can be considered ‘difficult’ to classify24,25—hence potentially perceptually ambiguous. According to this literature, the bilateral IFC therefore mostly fulfills the higher-level function of classification and categorization of vocal signals.

Unlike the aforementioned evidence, some other studies following from animal research have shown a more concrete ‘psychoacoustic’ role of the IFC instead of representing abstract decisional task demands, as some bilateral IFC/ventral prefrontal cortex subregions seem to be directly involved in the perceptual discrimination of different types of conspecific vocalizations19,26,27. According to these studies, the bilateral IFC would function pre-dominantly as a domain-specific higher-order auditory processing node that represents sound information in relation to categorical certainty or discriminability14, especially during explicit voice signal classification tasks3. Especially, this line of research would predict higher bilateral IFC activity for clear as opposed to ambiguous affective voices, as clear voices would allow a more direct categorical representation. Given this literature about the role of the bilateral IFC in processing affective voices—and the previous tendency to focus on the right hemisphere for emotion28, more knowledge is critical to clarify the functional role(s) of the IFC in general and of the right IFC in particular. We therefore aimed at clarifying the involvement of the right IFC in situations of classification and decisions, when participants were exposed to emotionally ambiguous vocal signals—a proxy for ambiguous social interactions in Humans.

To mechanistically investigate the functional role of the IFC in the classification of socio-affective voice information, an experimentally induced alteration of right IFC (pars triangularis) activity provides the required causality, albeit not exhaustively testing lateralization due to the unilateral procedure. Emotionally ambiguous vocal stimuli might trigger computations in the bilateral IFC given their challenging perceptual and cognitive processing, while this may not be the case for less challenging vocal material as in our ‘clear’ or ‘emotionally pure’ voice stimuli. Both cases might additionally require integrated IFC-STC (STC: superior temporal cortex) functioning in terms of neural connectivity3, either for top-down IFC-to-STC facilitations or as co-representations of categorical affective information29. The use of transcranial magnetic stimulation (TMS) appears as a reliable and non-invasive option to inhibit a proper functioning of computations performed in the IFC, especially in this context. Continuous theta burst stimulation (cTBS), a patterned pulse stimulation, as well as repetitive TMS (rTMS) are known to create reversible neural activity alteration30. Although the neural effects induced by the cTBS procedure are much faster and potentially longer-lasting than those of rTMS, the cTBS procedure was rather scarcely used for social and especially voice processing studies, even less in combination with functional magnetic resonance imaging (fMRI). Previous studies using cTBS procedure revealed an impairment of voice identity recognition but not in the case of affect discrimination tasks when cTBS was applied over the premotor cortex31. Similar results were observed in a combined cTBS-fMRI study targeting the premotor cortex in vocal affect processing, showing that cTBS triggered differential brain patterns in fronto-parietal, parahippocampal, and right IFC brain regions32.

Compared to cTBS, the use of rTMS revealed more mixed results regarding socio-affective processing from voices, yielding to unsuccessful outcomes in right STC33 and bilateral IFC22. Some methodological variations may explain these inconsistencies, most importantly regarding the type of TMS procedure (rTMS vs. cTBS) and especially large variations in IFC target regions. Interindividual differences may also account for these heterogenous data.

Given this variability concerning socio-affective classification from affective speech in the existing literature, we designed a three-alternative classification task on clear (pure emotions) and emotionally ambiguous voices while brain activity was influenced—and recorded—in a combined cTBS-fMRI setup. The right IFC target region for cTBS was chosen a priori based on previous—and above-mentioned—research by us and others, since it was an almost perfect match to right pars triangularis IFC observed in voice prosody categorization (see previous work by us17 and a thorough review19). Additionally, this specific right IFC region fitted well with the suprasegmental specialization attributed to this lateralized region19–we were also financially limited to stimulate only one region, i.e., one experimental group (more details are reported in the Methods). We then refined the target region based on the highest peak of ambiguous voice processing in the previously acquired control group for the best task-related precision. In order to address more precisely the cognitive and evaluative role of the right IFC in such contexts28, we created stimuli containing a gradual blending of angry and fearful emotions to manipulate affective ambiguity expressed in these voices (see the Methods). Anger and fear are very distinctive and well-recognized emotions expressed in vocal affect, and a blending of anger and fear leads to considerable affective ambiguity given their opposite nature for behavioral adaptations in listeners34. In an ideal world, we would have also included other emotions but since positive and negative emotions do not morph well and since one full session for one participant lasted already about three hours, we had to make a choice regarding the emotions to be included. We also did not want to create ambiguous voices referenced to neutral expressions35 in order to have a continuum of emotional ambiguity. In one group, cTBS was applied to the right IFC (experimental group, ‘EG’), while in another group cTBS was applied to the vertex as a control brain site (control group, ‘CG’).

Taking into consideration the literature introduced above, we hypothesized that: (a) cTBS over the right IFC pars triangularis would alter speech affect classifications and increase response speed especially for ambiguous affective voices. This first hypothesis relies on previous literature on the role of the right IFC in challenging perceptual and cognitive processing of voice signals18,19,22,28; (b) cTBS over the right IFC should also lead to a distinct and short-lived brain activity ‘reorganization’—illustrated by a decrease—of brain activity while classifying the most ambiguous affective voices, especially within the right AC/STC14,17,18 and in the limbic system (i.e., amygdala) for affect processing10; (c) finally, we expected altered functional connectivity between IFC and the limbic and auditory cortical system during cTBS to the right IFC, given the important status of the IFC in the neural network of socio-affective voice processing17,18,36. Because this analysis was exploratory, we did not direct our hypothesis and will report such data if any results survive a statistical threshold corrected for multiple comparisons.

Results

Preamble

As mentioned above, the present study used a combined—but not concurrent—cTBS-fMRI setup to non-invasively and temporarily alter brain functioning in the right IFC pars triangularis while human participants of the EG (Figure 2) and the CG (cTBS over the vertex; Figure 2) performed a three-alternative forced choice task on emotionally clear and emotionally ambiguous affective voices. These conditions were created by blending angry and fearful affective bursts—matching for actor identity, including: 90% fear and 10% anger (F90), 70% fear and 30% anger (F70), 50% fear and 50% anger (AF50), 30% fear and 70% anger (A70), 10% fear and 90% anger (A90). See Figure 1C for the evaluation of the initial stimuli by an independent sample, among which we selected the five mentioned blends. F90, F70, A90 and A70 were considered as ‘clear’ vocal emotion while AF50 voices illustrated the most ‘ambiguous’ category in our analyses and in the discussion. Neutral, unmorphed affective bursts were also included as a control. For all participants of both groups, the task was split into four runs, with the first run being performed before any cTBS procedure (‘base’ run or ‘cTBSpre’, as opposed to Run 1-3 or ‘cTBSpost’). See Methods for a detailed description of all technical aspects.

Fig. 2.

Fig. 2

Neuroimaging results of the affect classification task for the CG with a focus on the IFC post- versus. pre- cTBS. (A) Whole-brain data showing control group (CG) activations (black circle: left IFC maxima) for the [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast with continuous theta burst stimulation (cTBS) over the vertex, left hemisphere. (B) Activations for the same contrast in the right hemisphere, showing global maxima activity in the right inferior frontal gyrus pars triangularis (red circle, dashed black outline), used as the target region for the cTBS procedure of the EG. (C) Percentage of signal change using a cube of 27 voxels adjacent to the peak voxel—extracted according to activity maxima of the [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast for the CG—for both groups in the left (MNI xyz [-58, 20, 8]; black circle) and right (MNI xyz [10,34,54]; red circle, dashed black outline) IFGtri for each morphing condition. (D) TMS coil positioning for the cTBS procedure in both groups (CG: vertex, red circle with continuous black outline, MNI xyz [0, -20, 80]; EG: right IFC within the IFGtri, red circle with dashed black outline, MNI xyz [10,34,54]), decided after the CG was scanned in whole. Wholebrain activations are reported at p<.005 uncorrected with k>59 voxels, equivalent to FWE cluster correction for multiple comparisons of p<.05. The colorbar represents t-statistics. Inferior frontal cortex delineated in white with subregions delineated in black using the automated anatomical labelling (‘aal’) atlas. CG: control group; EG: experimental group; TMS: transcranial magnetic stimulation; IFC: inferior frontal cortex; IFGop: inferior frontal gyrus pars opercularis; IFGtri: inferior frontal gyrus pars triangularis; IFGorb: inferior frontal gyrus pars orbitalis; cTBS: continuous theta burst stimulation. Brain activations were displayed on brain renders of the CONN toolbox (v.22)37 and the figure was designed and created by the first author in Inkscape software (v.1.4.2; https://inkscape.org/).

Fig. 1.

Fig. 1

Experimental timeline, example affective voices, and behavioral results for initial stimulus selection and the vocal affect categorization task. (A) Experimental timeline for one participant, showing the different procedures for each group and detailing the three-alternative forced choice task (3-AFC) with examples of stimuli including spectrograms of neutral and morphed voices with percentage of each morphed emotion (anger, fear). (B) Probability of an ‘anger’ response as a function of the morphing procedure for each run, per group for the vocal affect categorization task. (C) Following the initial morphing procedure, we had an independent sample of nineteen right-handed participants (10 female, 9 male, mean age 32.2, SD 5.2) evaluate the emotionally blended stimuli. Participants were asked to categorize each stimulus (n=110; 10 trials for each of the 11 morphing levels) by a keypress (1: the voice expresses fear; 2: the voice epxresses anger; the keys were counterbalanced across participants). The line plots illustrate on the y axis the probability of an ‘anger’ response (emotion categorization response; black circles) and reaction times (black squares) according to each morphing level (x axis) with errorbars representing the standard error of the mean (SEM). (D) Averaged reaction times results for the vocal affect categorization task, for each morphing level, run and group. Error bars represent the standard error of the mean (SEM). aMT active motor threshold; cTBS continuous theta-burst stimulation (transcranial magnetic stimulation procedure); rIFCtri right inferior frontal cortex (pars triangularis) target region (MNI xyz 54 34 10) based a priori and on CG data (see TMS site localization section); CG control group; EG experimental group; ITI inter-trial interval; cTBS continuous theta burst stimulation; Probab. Probability; resp response. ***p<.001.

Right IFC as target region for cTBS procedure

As mentioned in the introduction, decisions on vocal affect are typically associated with activity in bilateral IFC. In the current study, we acquired the data of the CG first and we chose the right IFC, namely the right inferior frontal gyrus pars triangularis [IFGtri, MNI xyz 54,34,10; Figure 2AB], as target region for cTBS. This more specific localization of the best task-based target region was based a priori on previous literature17,19 and therefore on our own empirical data, namely the observation that the right IFGtri showed the highest activity for clear as opposed to ambiguous voice processing in the CG especially when contrasting cTBSpost compared to cTBSpre runs. Noteworthy here is the fact that our hypotheses were formulated on the basis of the inverse contrast—namely, ambiguous voices—and we therefore focused on the region rather than on the expected contrast itself, as reported in existing literature17,19. We however discuss this aspect in the limitations of the study, since this aspect alone hinders our interpretations (see Discussion and Limitations). The homologous IFC region in the left hemisphere (Figure 2AC) responded to a lesser extent to our contrast of interest and was therefore not targeted by the cTBS procedure—note also that including the left IFC as target would have required another group of participants and limited funding excluded that option, unfortunately. For the CG, the cTBS procedure targeted the vertex [MNI xyz 0 -20 80]. Coil positioning is reported in Figure 2D for each group and was meticulously placed for each participant with the help of a neuronavigation apparatus.

Behavioral results

We first assessed the effects of cTBS applied to the right IFC pars triangularis—3-3.5min before the beginning of the second run of the task, see Methods—on the decisional patterns during the classification of emotionally clear and emotionally ambiguous voice stimuli—as well as neutral voices (‘no-interest’ control condition). The behavioral classification data were parametrized along response speeds and decision categories in the three-alternative task (anger, fear, neutral). We wanted to test our first hypothesis according to which participants of the CG would perform better and potentially faster at categorizing blended voices as opposed to those of the EG who received cTBS over the right IFC—but we could not confirm this hypothesis since no such effect was observed, see below. As mentioned above, neutral voices were used as baseline control trials and we did not have any hypothesis concerning these, so we modelled these trials in all our analyses but discarded them from behavioral and neuroimaging planned contrasts. Performance was high for both groups and no significant difference was observed (χ2(1)=0.08, p=1; see Table 1). In the first statistical model with the response for an anger choice on affective trials as the dependent variable (Figure 1B), we observed a significant Group × Morph interaction (F(4, 13815)=20.14, p<.0001, η2p=.006), indicating group differences in emotion recognition across morphing levels. Additionally, a Group × Run interaction was observed (F(3, 13815)=3.25, p=.021, η2p=.0007), suggesting temporal dynamics in cTBS effects. Post-hoc contrasts revealed no simple effects for Group X Run (all p>.10) while a crossover interaction pattern for Group X Morphing was observed: at 10% morphing (fear-dominant), the EG showed impaired emotion recognition relative to the CG (d=0.25, 95% CI [0.05, 0.44], p=.013). Conversely, at 90% morphing (anger-dominant), the EG showed a trend toward enhanced recognition (d=-0.19, 95% CI [-0.38, 0.00], p=.055). No group differences emerged at intermediate morphing levels, namely in the most emotionally ambiguous voices (30-70%; all p>.24) and contrary to our hypothesis. This suggests emotion-specific rather than more general cTBS effects. The three-way interaction was not significant (F(12, 13815)=1.01, p>.10). For this reason, the two-way interactions mentioned above cannot help to distinguish from effects emerging before or after cTBS, and illustrate more general variability between groups. This model explained 24.63 % of the variance including both fixed and random effects (R2c= 0.2463) and 9.95 % of the variance solely for fixed effects (R2m= 0.9950).

Table 1.

Accuracy and reaction times data for neutral trials, for each group and run.

cTBSpre run cTBSpost Run 1 cTBSpost Run 2 cTBSpost Run 3
CG acc 87.50% (6.55) 87.78% (4.96) 88.05% (5.55) 88.05% (5.55)
CG rt 1032ms (188) 928ms (201) 918ms (213) 944ms (208)
EG acc 93.23% (6.74) 97.94% (5.10) 97.06% (5.71) 97.94% (5.10)
EG rt 1112ms (192) 1043ms (204) 1061ms (186) 1137ms (206)

Value: mean (SD). acc: accuracy; rt: reaction times; ms: milliseconds.

In the second statistical model with the reaction times as the dependent variable (Figure 1D), we observed a significant main effect of Run (F(3, 13815)=17.97, p<.0001, η2p=.004), indicating that response speed changed across sessions, likely reflecting learning effects for both groups. A Group × Run interaction was observed (F(3, 13815)=3.33, p=.019, η2p=.0007) suggesting temporal dynamics in cTBS effects on response speed. Post-hoc contrasts revealed that the EG showed a trend towards faster responses immediately after stimulation (Run1: CG-EG=152ms, d=0.32, p=.074), with this speed ‘advantage’ disappearing in the subsequent runs (Run2: 113ms, d=0.24, p=.183; Run3: 99ms, d=0.21, p=.244). However, unlike the accuracy data, no Group × Morph interaction emerged for reaction times (F(4, 13815)=0.54, p=.704, η2p<.001), indicating that the emotion-specific cTBS effects observed in accuracy were not accompanied by corresponding changes in response latency, and that again this result reflects a difference between groups mostly independent from the cTBS procedure. Again, the three-way interaction was not significant (F(12, 13815)=0.66, p>.10). This model explained 31.12 % of the variance including both fixed and random effects (R2c= 0.3112) and 1.84 % of the variance solely for fixed effects (R2m= 0.1835). For both models, the remaining effects are reported in Table 2.

Table 2.

Response and traction times statistics.

Response
Effects F df1 df2 p eta_p cohens_f
1 Group 0.004 1 33.0 0.9529 0.0001 0.010
2*** Morph 28.457 4 36.0 0.0000 0.7597 1.778
3 Run 1.383 3 13815.1 0.2459 0.0003 0.017
4*** Group:Morph 20.136 4 13815.0 0.0000 0.0058 0.076
5* Group:Run 3.253 3 13815.1 0.0208 0.0007 0.027
6 Morph:Run 0.965 12 13815.0 0.4804 0.0008 0.029
7 Group:Morph:Run 1.006 12 13815.0 0.4401 0.0009 0.030
RT
Effects F df1 df2 p eta_p cohens_f
1 Group 1.784 1 33 0.1908 0.0513 0.233
2 Morph 1.764 4 36 0.1576 0.1638 0.443
3*** Run 17.969 3 13815 0.0000 0.0039 0.062
4 Group:Morph 0.544 4 13815 0.7036 0.0002 0.013
5* Group:Run 3.329 3 13815 0.0187 0.0007 0.027
6 Morph:Run 0.674 12 13815 0.7783 0.0006 0.024
7 Group:Morph:Run 0.662 12 13815 0.7899 0.0006 0.024

RT: reaction times; Group: group factor (CG;EG); Morph: Morphing factor (percentage of anger/fear: 10/90,30/70, 50/50, 70/30, 90/10); Run: Run factor (base run; Run1, Run2, Run3); F: Fisher F statistics; df: degrees of freedom; eta_p: partial eta-squared; cohens_f: Cohen’s f. *p<.05, ***p<.001.

Voice-sensitive activations in bilateral auditory cortex

To determine the sample-specific regions in bilateral AC that are generally sensitive to voice compared to other types of auditory signals, we analyzed the data of a functional voice localizer scan specific to our sample. During this scan, participants listened to vocal and non-vocal sounds, and we found higher activity in bilateral STC and IFC (Figure 3AB; Table S4) when contrasting vocal against non-vocal sounds. The observed pattern of activations are similar to previous reports (“voice areas”38 or VA). IFC subregions are also delineated in the maps of Figure 3.

Fig. 3.

Fig. 3

Sample-specific (N=35) activations for the VA localizer task. (A) Voice areas and IFC regions (black outline) in the left hemisphere for the [Vocal > Non-vocal] contrast. (B) Voice areas and IFC regions (black outline) in the right hemisphere for the [Vocal > Non-vocal] contrast. Wholebrain activations are reported at p<.005 uncorrected with k>59 voxels, equivalent to a FWE cluster correction for multiple comparisons of p<.05. Colorbars represent t-statistics. IFGop inferior frontal gyrus pars opercularis; IFGtri inferior frontal gyrus pars triangularis; IFGorb inferior frontal gyrus pars orbitalis; INS insula; STC superior temporal cortex; STS superior temporal sulcus; MTC middle temporal cortex; DLPFC dorsolateral prefrontal cortex; MFG middle frontal gyrus. Brain activations were displayed on brain renders of the CONN toolbox (v.22)37 and the figure was designed and created by the first author in Inkscape software (v.1.4.2; https://inkscape.org/).

Functional brain activations for the classification of morphed voices as a function of cTBS

We used the sample-specific cortical definition of the VA to determine functional activations in the main experiment that were located inside as well as outside of this general voice processing network.

Therefore, our contrast of interest was computed to uncover brain activity relating to [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] and its opposite [AF50 > A90, A70, F90, F70] x [cTBSpost > cTBSpre], per group first and then between groups. We found enhanced brain activity in the former but not in the latter contrast. In fact, the [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast yielded enhanced brain activity in the STC (within the temporal voice areas), bilateral IFC especially in the pars triangularis for the CG (Figure 4BE) and EG (Figure 4CF), separately.

Fig. 4.

Fig. 4

Neuroimaging results of the vocal affect classification task with outlines of sample-specific voice areas. (A) Whole-brain data showing between-group activations for the [CG > EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast with continuous theta burst stimulation (cTBS) over the vertex (red circle, black outline) for the control group (CG) as opposed to over the right inferior frontal gyrus (IFG; red circle, dotted outline) for the experimental group (EG). (B) Left hemisphere activations for [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast for the CG. (C) Left hemisphere activations for [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast for the EG. (D) Percentage of signal change (±SEM) extracted in the right mSTG* [MNI xyz 58, 0, -12] for the main contrast ([CG > EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre]). (E) Right hemisphere activations for the [CG > EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast for the control group. (F) Right hemisphere activations for the [CG > EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast for the experimental group. Sample-specific voice areas (VA) are outlined in each panel in black. a/m/p anterior/mid/posterior; op/tri pars opercularis/triangularis; MFG middle frontal gyrus; STC superior temporal cortex; MTC middle temporal cortex; SPL superior parietal lobule; INS insula; PTe planum temporale; SMG supramarginal gyrus; STS superior temporal sulcus. Wholebrain activations are reported at p<.005 uncorrected with k>59 voxels, equivalent to FWE cluster correction for multiple comparisons of p<.05. Colorbars represent t-statistics. Brain activations were displayed on brain renders of the CONN toolbox (v.22)37 and the figure was designed and created by the first author in Inkscape software (v.1.4.2; https://inkscape.org/).

Computing the three-way interaction between the group, morphing, and run factors revealed activity within the VA (Figure 4A), located in anterior and mid STC as well as middle temporal cortex (MTC), and in superior parietal lobule and superior frontal gyrus only in the right hemisphere ([CG > EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre]; Figure 4A). The [EG > CG] x [A90, A70, F90, F70 > AF50], [EG > CG] x [AF50 > A90, A70, F90, F70] x [cTBSpost > cTBSpre] and the [CG > EG] x [AF50] x [cTBSpost > cTBSpre] contrasts did not yield any above-threshold brain activity. Peak coordinates and statistical information are reported in Table 3. These results are the exact opposite to what we hypothesized—according to which cTBS over the right IFC should lead to a modified brain activity patterns while classifying the most emotionally ambiguous voices in the STC (see Figure 4D), and we did not observe any above-threshold activity in the amygdala or limbic cortex.

Table 3.

MRI coordinates for contrasts of interest.

[CG > EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre]
Region label Hemisphere MNI X MNI Y MNI Z T-value Cluster size (voxel count)
SPL R 28 -42 48 4.63 186
STG R 58 0 -12 3.74 84
MFG R 20 20 56 3.73 73
Cuneus R 18 -80 32 3.53 120
[CG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre]
Region label Hemisphere MNI X MNI Y MNI Z T-value Cluster size (voxel count)
STG R 58 -26 2 5.70 549
Cerebellum L -20 -74 -36 4.75 318
Insula R 32 20 -14 4.69 218
STG L -50 -16 -6 4.49 686
STS R 54 -4 -14 4.08 252
AMY R 20 -2 -12 3.84 107
IFGtri R 50 46 -4 3.78 105
[EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre]
Region label Hemisphere MNI X MNI Y MNI Z T-value Cluster size (voxel count)
IFGtri R 54 34 10 5.21 324
Cerebellum L -14 84 -14 4.66 4273
STS R 58 -32 2 4.60 822
STS R 54 -4 -12 4.38 372
STG L -54 -42 8 4.25 626
IFGtri L -58 20 8 4.19 713
STG L -50 -20 -6 3.82 66
MFG L -4 64 28 3.47 98

Statistical threshold: p<.05 FWE cluster correction (voxelwise p<.005 uncorrected, k>59).

SPL superior parietal lobule; STG superior temporal gyrus; MFG middle frontal gyrus; STS superior temporal sulcus; AMY amygdala; IFGtri inferior frontal gyrus pars triangularis; L: left hemisphere; R: right hemisphere.

Functional connectivity data for morphed voice classification as a function of cTBS

ROI-specific functional connectivity analyses were computed in order to assess the Group X Morphing X Run interaction on the organization of functional networks of emotionally ambiguous voice classification. These analyses revealed anti-coupling in the right amygdala and left mid STC (Figure 5AB) as well as coupling in the left amygdala and right IFC (Figure 5BC) triggered by cTBS on the right IFC (contrast: [EG > CG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre]). This result indicated that when the EG as compared to the CG classified emotionally clear as opposed to emotionally ambiguous voices as a function of the cTBS procedure, linear negative correlation between left mid STC and right amygdala increased. On the other hand, linear association between the left amygdala and right IFC (the cTBS target region for the EG) increased. The cTBS procedure on the right IFC therefore seems to enhance functional connectivity between the bilateral amygdala and subparts of the VA and IFC. Such data were used to explore hypothesis (c) but reveal no causal relation whatsoever between any of the correlating regions, nor a causal or directional link between conditions/emotions and this connected network.

Fig. 5.

Fig. 5

Seed-to-seed functional connectivity results for the vocal affect categorization task. (A) Anti-coupled functional connectivity was found between the left mid superior temporal cortex (mSTC) and the right amygdala (AMY) for the [EG > CG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast. (B) Summary of coupled and anti-coupled functional connectivity data and statistical values. (C) Coupled functional connectivity between the left amygdala (AMY) and the right inferior frontal cortex (IFC) for the [EG > CG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] contrast. Colorbar shows two-tailed t-statistics. Statistical threshold: p<.05 FDR corrected for multiple comparisons at the seed level. Brain activations were displayed on brain renders of the CONN toolbox (v.22)37 and the figure was designed and created by the first author in Inkscape software (v.1.4.2; https://inkscape.org/).

Since functional connectivity data were computed using bivariate correlations between neural nodes, inversing the contrast to [CG > EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] yields to a sign inversion (coupling in the right amygdala and left mid STC; anti-coupling in the left amygdala and right IFC) and is therefore not illustrated.

Discussion

The present study had the general aim of understanding the causal role of the right inferior frontal cortex in classifying and representing emotionally clear and emotionally ambiguous affective bursts, voices. A combined cTBS-fMRI procedure was therefore used to non-invasively and temporarily alter activity in the right IFG pars triangularis for half of our participants (EG), while the other participants (CG) received cTBS over a control region—namely the vertex, selected because it did not alter any speech- or attention-related neural activity relevant for affective sound processing and classification39. Even though behavioral differences were observed between groups, the absence of a significant behavioral triple interaction goes contrary to our first hypothesis and thus prevents us from drawing any causal conclusions on decisional processes during the affect classification task. This target region was chosen a priori because it is a commonly observed region across numerous studies on voice processing, categorization, and classification18,19,28,36,40 and then refined empirically because it was predominantly recruited during the processing of affective voices in the CG. In our study, the cTBS procedure therefore had an impact on brain activations but not on behavioral data, therefore not allowing us to validate our primary behavioral hypothesis, unfortunately. Right IFC cTBS highlighted a reduced recruitment of STC regions for emotionally clear as opposed to ambiguous voices. Functional connectivity analyses for emotionally clear versus emotionally ambiguous voices revealed between-group differences, such that right IFC cTBS led to anti-coupling between the mid STC and right amygdala and enhanced coupling between left amygdala and right IFC, the latter region being the target of the cTBS procedure in the EG. As mentioned in the Results section above, our three hypotheses could not be confirmed, especially the first hypothesis pertaining to decisional differences between groups as a function of the cTBS procedure. We discuss these aspects below in the hope of highlighting inconsistencies in our design and data that will serve future work on this topic.

Absence of behavioral effects as a function of right IFC cTBS

The morphing procedure used to create affective ambiguity allowed us to test the role of the right IFC for affective voice classification challenges (as induced by emotionally ambiguous voices). Accordingly, we hypothesized that right IFC alteration by the cTBS procedure would affect the classification of emotionally ambiguous affective voices, namely those with a blend of 50% anger and fear. This hypothesis was formulated because this voice category would represent peak decisional difficulty as compared to emotionally clear voices. As mentioned above, these data were inconclusive and the required triple interaction was not significant. In a recent study on the categorization of affective human and nonhuman primate vocalizations, we observed an implication of the IFC—especially the pars triangularis—for both the probability of a correct and an incorrect classification, linked to categorization difficulty25. OFC activity was specifically correlated with the probability of a correct categorization. Such result in the IFC favors the viewpoint of an integration of both decisional difficulty and accurate categorization within the IFC, and since our data cannot be interpreted due to the absence of significant cTBS-related behavioral results, future work should test more specifically subregions of the IFC in addition to alteration of OFC activity—probably using another stimulation technique due to potentially painful muscle activation—in the context of emotionally ambiguous voice categorization.

This study cannot therefore bring new light to the causal role of the right IFC in such context, although previous work highlighted its impact in the cognitive evaluation and judgement of affective voices9,18,19. The IFC is also known to be sensitive to nonverbal vocalizations—as we used in our task—in both children41 and nonhuman primates42, supporting sound classification43 and higher-order auditory representations27. The neural effects of cTBS are usually the largest immediately after the application, and the effects are known to decay over time—they can last up to 60 minutes after this specific type of patterned stimulation44. If the right IFC is assumed to be a neural node for lifting affective voice processing on a cognitive-focused level19,28,45,46 that demands processing efforts47—as observed in spoken word recognition and phonetic competition48–50, inhibition of the right IFC with cTBS might loosen this cognitive ‘focus’ to facilitate processing51 with presumably better processing efficiency52. This is pure speculation here, and we look forward to future work testing such possibility, potentially using a more computational modeling perspective and using bilateral or lateralized stimulation procedures.

The relevance of this right IFC subregion as targeted by our cTBS procedure for complex affective evaluation and classification tasks is highlighted by previous studies. The IFC is a rather large region19 and it includes several subregions or subparts such as its more superior part, pars opercularis, the more inferior and anterior part, pars triangularis and the most ventral part, pars orbitalis that is located next to the mid orbitofrontal cortex. In our study, the location of the right IFC was exclusively within the pars triangularis in the right hemisphere. In another study, this specific subregion of the right IFC was recruited when more complex affective voice categorizations as opposed to simpler discrimination was performed, also sometimes labelled ‘unbiased’ versus ‘biased’ perceptual decision making, respectively12. Such results are therefore in line with our neuroimaging data, especially when considering the fact that an affect classification task was employed in our procedure. Affective voice discrimination as opposed to categorization would on the other hand depend on a different IFG subregion, namely the bilateral IFG pars opercularis6,23, to which only residual current could be transmitted due to the cTBS procedure and coil location in our study. According to our results and to the literature12,14,19,53, the right and potentially the left IFG pars triangularis54 could therefore be highly selective to categorization, especially for decisional uncertainty55,56. This is backed by two studies that emphasized the importance of the target location for a TMS procedure by pointing toward slower or faster responses according precise IFC disruption in the pars orbitalis57 or opercularis40, respectively.

Impact of right IFC cTBS on brain activity for vocal affect decision-making

Considering our neuroimaging data and when looking at groups separately (CG, EG), bilateral IFG pars triangularis activity was observed specifically for emotionally clear as compared to emotionally ambiguous voices in both groups separately, but to our surprise not in the inverse contrast. This observation is contrary to the hypothesis of the IFC being pre-dominantly a brain node coding and regulating decisional challenges especially during sensory ambiguity, at least in the present study. The IFC is active when making socio-affective decisions on voice signals, but not simple acoustic decisions3. It is therefore primarily recruited when decisions may be flagged with categorical uncertainty and/or choice precaution25,55,56. The bilateral functional significance of the IFC could be influenced solely by right hemisphere cortical interference through cTBS, as already shown using TMS on the left prefrontal cortex58. This observation is in line with the patterns of post-cTBS activations found in the IFC in our study, namely large bilateral IFC activity in the CG vs. focused right-lateralized activity in the EG. Such assumptions could also be tested in the future by using intermittent—as opposed to continuous—TBS over the right and/or left IFG pars triangularis, known to enhance brain activity and not alter it30,59. Combining intermittent and continuous TBS on the bilateral IFC during emotionally clear and emotionally ambiguous voices in a recognition or categorization task would allow to specifically test the role of this region more exhaustively in such context. Such assumption would help clarify the presumed inter-hemispheric plasticity observed in the IFC60, with stimulation intensity having a potential crucial impact on current transmission in the brain tissues58.

Numerous other functions have been proposed for the IFC, and some of them could also apply to our data61. For instance, behavioral classifications imply not only categorical associations but also action and motor functioning as well as executive functions, most notably inhibition. Inhibition was shown to causally rely on the right IFC, with direct current stimulation over the IFC causing an ‘activation of inhibition’62. Such topic was also previously reviewed in detail and while the distinct role of prefrontal cortex subregions was initially questioned63, the specific role of the IFC would be to implement some cautionary mechanisms over immediate response tendencies, especially the right IFC pars opercularis and triangularis64. According to this literature, it is therefore possible that cTBS over the right IFG pars triangularis enabled the release of inhibition in the EG. Again, this possible effect should be investigated in the future and is only highly speculative here. Indeed, response inhibition and inhibition of immediate categorical decisions should be considered in future work on voice recognition and categorization studies. In the context of social interactions through both emotionally ambiguous and clear or ‘pure’ voices, processing modes can be more cognitive- or intuition-driven and this distinction has been documented in the literature65,66. These processes could potentially be on a temporal continuum67, with intuition-driven processing supposed to be more efficient68, especially when relying on prior expertise69. The role(s) of the bilateral and especially right IFC in this context should be investigated, especially through the lens of inhibition-related mechanisms and their impact on voice decisional processes.

The anterior STC is a major brain node in the analysis of affective speech11,70 and language71, and shows functional70 and structural72,73 connections to the right IFC. Especially, the between-group contrasts highlighted above-threshold voxels in the right mid-to-anterior STC for CG compared to EG, especially when comparing the cTBS-dependent classification of emotionally clear as opposed to emotionally ambiguous voices—but again, surprisingly not for the inverse contrast. The role of the STC in affective speech decoding is well documented by us and others4,14,17,28,36,70,74,75 and among this vocal emotion literature—including speech, pseudowords and affective bursts, some studies isolated activity enhancement in the right anterior STC for the categorization of female-ambiguous voices by male participants76. Right anterior STC activity was also enhanced as a function of trial-level acoustic distance between voices when categorizing morphed male and female voices to create gender-ambiguous stimuli77. Perceived gender ambiguity did however recruit the bilateral IFC and the posterior and anterior cingulate cortex. These results were interpreted by the authors as a two-stage process for gender voice classification, with auditory feature extraction in the anterior STC and voice categorization in the bilateral IFC, in a similar fashion described by Schirmer and Kotz for processing and making decisions on affective speech28. This interpretation is an interesting viewpoint that could have contributed to the understanding of our results, had they followed our hypotheses, although our morphing procedure targets voice emotion, not gender. Importantly, it also means that mid and anterior STC neuron populations have the ability to influence classification and decision processes in the IFC and that at some point altered functioning in the IFC could be overcome or at least reduced by right anterior STC activity and its feature-selectivity function.

A last point of discussion concerns the large activation of the bilateral anterior insula in the EG. While the insula was repeatedly reported in social cognition and more specifically in the embodiment of emotion by us78 and others79,80, it is also a core region for pain81. Stimulating the anterior insula was indeed shown to increase pain thresholds, with participants subsequently feeling less pain82. Since IFC pars triangularis stimulation could have been painful for some participants due to involuntary facial muscular stimulation during cTBS, this enhanced activity only in the EG could reflect some mechanism that was used by the brain to keep the pain at large. Alternatively, activity in the anterior insula in the EG could also have been enhanced by cTBS to the right IFC. This is again speculation but it should definitely be investigated in the future, especially since this region is highly involved in bodily self-consciousness and interoception, concepts that pertain particularly well to the embodiment of emotion in voice signals78 and therefore to vocal emotion perception and to a lesser extent, categorization.

Functional connectivity as a function of IFC cTBS, between-groups

Our functional connectivity results also highlight the importance of the limbic system for processing affective speech. The right amygdala showed decreased functional connectivity with the left STC in the EG vs. CG for the classification of emotionally clear and emotionally ambiguous voices, while the left amygdala was significantly anti-coupled with the right IFC. These results were not expected as is—and therefore partly contradict our last hypothesis, since we expected an effect of cTBS on brain networks of ambiguous rather than non-ambiguous voice categorization. However, they are coherent with existing literature on vocal affect processing10,17 and argue for a wider role of the limbic system in affective voice classification10 and feature extraction83 depending on the emotional ambiguity of the voice signal. In fact, intuitive processing of vocal anger74,84 would take place in the amygdala following which feature extraction would be completed in the STC28. According to Schirmer and Kotz28, the last of the three-stage process of vocal affect processing would consist of an evaluative process in the IFC. Our functional connectivity results, especially the coupling between the left amygdala and the right IFG pars triangularis as a function of cTBS and morphing, suggests a potential enhancement in feedback generation from automatic processing—in the amygdala—and final decision leading to categorization, taking place in the right IFC. However, this interpretation lacks the backing of behavioral effects related to cTBS, and should therefore be clarified in future work.

Material and methods

Participants

Forty healthy volunteers took part in the fMRI study. Sample size was therefore calculated using ‘G*Power 3’ software85 with standard values for the architecture of our study (effect size of 0.5, power of 0.95 and alpha of 0.05) and 80% of explained variance was reached with a sample of 20 participants per group. We hence recruited 40 participants matching our inclusion criteria. Half of them were randomly assigned to the control group (CG) while the remaining participants were assigned to the experimental group (EG). For each group, 10 female and 10 male participants were included. Five participants were excluded from the final sample due to corrupted log files (N=2) and the impossibility to determine a reliable motor threshold (N=3). Therefore, the final groups included 18 participants for the CG (mean age=22.35, SD=1.66, 8 female) and 17 participants for the EG (mean age=25.05, SD=5.18, 7 female) with no significant between-group age difference (F(1,34)=3.66, p=.065). Given the dropout we encountered and the subsequent reduction of the sample size to N=35 (NEG=17, NCG=18), explained variance was reduced to 68.5%, as recalculated using ‘G*Power 3’ by keeping the other parameters identical (effect size of 0.5, power of 0.95 and alpha=0.05). To confirm the a priori right IFC pars triangularis target region for the cTBS procedure, the CG was acquired first and results confirmed previous work on the matter (see Figure 4E). Participants were informed of all aspects of the experiment before giving their informed written consent to take part in the study. They were also informed that they could abort the experiment and quit the study at any time without justification. The study was conducted according to the Declaration of Helsinki and approved by the Ethics Committee of the University of Zürich, Switzerland.

Stimuli: initial evaluation and selection

To select stimuli for the affective forced choice task, we created blends of vocally expressed anger and fear stimuli using a voice morphing procedure1,2 on a set of commonly available and validated vocal affective bursts, namely the Montreal affective voice database or ‘MAV’86. From this database we selected five male and five female voices. We created blends of affective voices ranging from 100% of one emotion to 100% of the other emotions in steps of 10% morphing, resulting in 11 different stimulus categories (Figure 1C) and 110 stimuli in total: 60-100% fear proportions (F60, F70, F80, F90, F100), 60-100% anger proportions (A60, A70, A80, A90, A100), and the 50/50 mix of anger and fear in the condition AF50. This morphing procedure was implemented using the STRAIGHT toolbox scripts (https://github.com/meccaLeccaHi/voice_morphing) running in Matlab 2018a (The Mathworks Inc., Natick, MA, USA).

For selecting the appropriate stimuli and morphing rates for the main experiment, we asked an independent sample of participants (N=19) to classify these stimuli in a 2AFC task as portraying either ‘anger’ or ‘fear’. Our aim was to select one emotionally ambiguous morphing level as well as four other gradually less ambiguous emotional voices for a total of five morphed voices. Using generalized linear mixed-effects (‘glmer’) modelling with the participants’ choices as dependent variable, illustrating the probability of an ‘anger’ response. Fixed effects included the Morphing factor with eleven morphing levels of the voice stimuli (anger/fear percentage: 100/0, 90/10, 80/20, 70/30, 60/40, 50/50, 40/60, 30/70, 20/80, 10/90, 0/100) with N=10 trials per morphing level (N=110 trials per participant; each stimulus is repeated only once for each participant). Random effects included in this order: Participant, Age, Gender, Stimulus identity, Speaker identity. Each trial was presented in a fully randomized order and not in a ‘type1-index1’ manner, although we realized a posteriori this may have impacted a) evaluation since trial order effects in such paradigm were observed in the past20 and b) our fMRI results due to probable carry-over effects between trials sequence87. The same model was used to analyze reaction time data—these were of no interest for stimulus selection—with reaction times as dependent variable. All reported p-values of these analyses use a Bonferroni correction for multiple comparisons implemented in package ‘lmerTest’. We observed a significant main effect of morphing on the categorization (χ2(10)=1094.3, p<.0001); reaction time data were of no major interest here but are reported in Figure 1C and Table S1. We used planned comparisons to explore which morphing levels would be more suitable—i.e., accurately evaluated—and contrasts showed significant differences between any pair of morphing levels (all p<.0001, see Figure 1, Table S2). Since the AF50 condition embodied the highest affective ambiguity, we selected it as well as two less ambiguous morphing levels per emotion in equal steps of 20% morphing, that is A90, A70, F90 and F70. According to this selection criteria, the stimuli for the main experiment consisted of expressions of vocal affect containing 10%, 30%, 50%, 70%, or 90% of fear and therefore 90%, 70%, 50%, 30%, and 10% of anger at the same time, respectively (labels: A90, A70, AF50, F70, F90; Figure 1A). The 10% conditions referred to rather clearly expressed vocal affect, the 50% condition was of highest affective ambiguity, and the 30% morphing level referred to an intermediate level of ambiguity. This model explained 40.25% of the variance including both fixed and random effects (R2c=0. 4025) and 30.46% of the variance solely for fixed effects (R2m=0.3046). In the main fMRI task, participants therefore listened to these five expressions of morphed vocal affect as well as to neutral vocalizations, and they were asked to classify them as ‘anger’, ‘fear’, or ‘neutral’, see below.

Vocal affect categorization task

Our stimuli consisted of affective bursts (“Aah”) from the Montreal Affective Voices database86, selected from the initial evaluation of a larger sample of stimuli. See previous section. Voices were pronounced by five female and five male actors. They expressed either a neutral emotional tone or a mix of anger and fear with a varying percentage of each emotion (percentage of anger/fear): 90/10, 70/30, 50/50, 30/70, 10/90 (Figure 1), respectively labelled A90, A70, AF50, F70, F90 in the text. Each voice was morphed using the same actor to avoid creating a strange identity that would add a critical confound to the emotional morphing procedure. As mentioned above, this morphing procedure was implemented using the STRAIGHT toolbox scripts (https://github.com/meccaLeccaHi/voice_morphing) running in Matlab 2018a (The Mathworks Inc., Natick, MA, USA).

For each run of the vocal affect categorization task, we therefore had a total of five conditions of interest involving a morphing of anger and fear emotions (24 trials each) and one control condition involving neutral voices only (10 trials). Each voice had a duration of 500 ms to 1200 ms, the order of which being pseudo-randomized for each run and participant (Figure 1A). Each run included 24 trials of each conditions of interest for a grand total of 72 trials and 30 neutral voice trials across the three experimental runs and the control run.

The experiment took place at the University Hospital of Zürich, Switzerland, using the research-dedicated whole-body magnetic resonance imaging (MRI) scanner (Philips Achieva 3T) of the Laboratory for Social and Neural Systems Research (Department of Economics, University of Zürich, Zürich, Switzerland). Participants were first taken to a transcranial magnetic stimulation (TMS) room in order to determine their individual motor threshold through a standard procedure (see TMS procedure below). They were then taken to the MRI room and comfortably installed in the scanner. The MRI session started with one base run, followed by the TMS procedure and the three experimental runs (Figure 1A). At the end of the control run, the participants stayed on the scanner table. The table was taken halfway out and the TMS apparatus was brought to the participant. For the CG, continuous theta burst stimulation (cTBS) was administered for 40 seconds over the vertex (MNI xyz [0, -20, 80]) because it was previously successfully used as a good control site for TMS protocols39 and used successfully in recent studies88,89. For the EG, cTBS was administered to the right inferior frontal cortex (rIFC; MNI xyz [10,34,54]), for 40 seconds as well (see detailed TMS procedure below). As soon as the TMS procedure ended (for both groups), necessary care was taken to relocate the participant as fast and as smoothly as possible inside the scanner. This procedure was used to minimize the time between stimulation and the start of the first experimental run, in order to have the most reliable TMS effect over the rIFC for the following runs across participants. The average duration between the end of the TMS procedure and the start of the first experimental run was 3.33 minutes (SD=0.47) for the CG and 3.00 minutes (SD=0.52) for the EG. No statistical difference was observed regarding this duration between groups (F(1,34)=0.33, p=.57).

For all runs, the task of the participants was to explicitly categorize the emotional tone of the voice that was presented in each trial by a key press using three buttons on an MRI-compatible response box (Current Designs Inc., PA, USA). Button mapping was fully randomized between participants. Buttons represented the following options: Anger, Fear, Neutral. Participants were instructed to respond as fast and as accurately as possible during the 2 sec blank screen that appeared after each trial (see Figure 1A).

Behavioral data analysis

To take into account within- and between-subject variance for the initial stimulus evaluation and for the main cTBS-fMRI task—for the dependent variable of each model—we used lmerTest90 and lme4 packages91 in R Studio92 to perform mixed-effects modelling of the data. For all analyses, models were tested using type II Wald Chi-square and type III Fisher tests using the ‘Anova’ function of the ‘car’ package93. We report effect sizes using the ‘MuMIn’94 package based on two indicators, a marginal and a conditional R2 (R2m and R2c, respectively). R2m reflects the variance explained by the fixed factors, R2c the variance explained by the entire model (both fixed and random effects). Partial eta-squared and Cohen’s f effect sizes are also reported for each model effect and each computed post-hoc tests.

Vocal affect categorization task

For the main cTBS-fMRI task, mixed-effects modelling of both the responses (model 1, linear modelling, linear mixed-effects model ‘lmer’) and reaction times (model 2, linear modelling, linear mixed-effects model ‘lmer’) of our participants were computed, for each trial. Data included our conditions of interest, namely the 90/10, 70/30, 50/50, 30/70, 10/90 anger/fear morphed voices—labelled A90, A70, AF50, F70, F90, respectively, in the manuscript—and the control voices expressing neutral content (no morphing, filler condition of no-interest). For reaction times data and since raw values were not normally distributed, the log of the reaction times was used. For model 1, raw responses were used to determine the choice or category for an anger response for each trial. This variable was the dependent variable (RESP) while the log of the reaction times RTs_log) was the dependent variable for model 2. For both models, fixed effects included, in this order, the interaction between Group, Morphing and Run while Participant, Age and Gender were introduced as random effects in addition to Stimulus and Speaker identity (StimulusID and SpeakerID, respectively). The formulae are the following:

graphic file with name d33e2790.gif

and

graphic file with name d33e2794.gif

These models allowed us to test our hypothesis according to which participants of the CG would perform significantly better at categorizing morphed voices, considering the TMS procedure: [CG > EG] x [AF50 > A90, A70, F90, F70] x [cTBSpost > cTBSpre] and its inverse regarding morphing [CG > EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre]. For these two contrasts, the weighting of the vector for the Morphing factor was similar to the one used for fMRI data, namely [6*(AF50) > -2*(A90), -1*(A70), -2*(F90), -1*(F70)] and [-6*(AF50) > 2*(A90), 1*(A70), 2*(F90), 1*(F70)]. A contrast targeting specifically the most ambiguous voices: [CG > EG] x [AF50] x [cTBSpost > cTBSpre].

All reported p-values of these behavioral analyses use a Bonferroni correction for multiple comparisons implemented in package ‘lmerTest’.

Voice-sensitive areas localizer task

Stimuli

Auditory stimuli consisted of sounds from a variety of sources. Vocal stimuli were obtained from 47 speakers: 7 babies, 12 adults, 23 children and 5 older adults. Stimuli included 20 runs of vocal sounds and 20 runs of non-vocal sounds. Vocal stimuli within a run could be either speech 33%: words, non-words, foreign language or non-speech 67%: laughs, sighs, various onomatopoeia. Non-vocal stimuli consisted of natural sounds 14%: wind, streams, animals 29%: cries, gallops, the human environment 37%: cars, telephones, airplanes or musical instruments 20%: bells, harp, instrumental orchestra. The paradigm, design and stimuli were obtained through the Voice Neurocognition Laboratory website (http://vnl.psy.gla.ac.uk/resources.php). Stimuli were presented at an intensity that was kept constant throughout the experiment 70 dB sound-pressure level.

Experimental procedure, paradigm

Participants were instructed to actively listen to the sounds, both vocal and non-vocal. The distinction between vocal and non-vocal runs was not revealed to the participants who were therefore naïve to run organization and timing. The silent inter-run interval between each run of either vocal or non-vocal was 8s long. Task duration was about 10 minutes in total. Wholebrain, sample-specific result outline of this task for the vocal > non-vocal contrast of interest is reported in Figure 3 and voice areas are outlined in black in Figure 4.

TMS procedure

As aforementioned, participants were stimulated either over the rIFC (EG) or over the vertex (CG) by means of standard cTBS44. First, we determined the active motor threshold (aMT) individually for each participant by stimulating the primary motor cortex in the left hemisphere. The aMT was defined as the percent of maximum stimulator output (mean intensity was 49.3%, SD 6.52%) required to elicit a motor-evoked potential larger than 200μV from the contralateral first dorsal interosseous muscle in five out of ten TMS pulses. During the determination of the aMT the participants exerted a constant pressure between the index finger and the thumb of about 20% of the maximum force44. For the cTBS protocol in the MRI room, the stimulation intensity was set to 80% of the aMT (mean intensity was 39.94%, SD 5.21%). Due to unwanted muscle contraction and discomfort at the right IFC location—more specifically on the inferior frontal gyrus pars triangularis (IFGtri), percentage of aMT for the EG was further reduced by 10%, leading to lower amplitude stimulation for the EG compared to the CG in the MRI scanner (CG: mean intensity=42.84, SD=4.78; EG: mean intensity=36.87, SD=3.75; F(1,34)=16.80, p<.001).

On the bed of the MRI scanner, the participants received cTBS with an MR-compatible coil (MRi-B91 coil, MagVenture A/S, Farum, Denmark). Stimulation site and coil orientation (see TMS site localization; Figure 2) were marked on a fixed cap by means of a TMS Neuronavigation system (BrainSight 2, Rogue Research Inc., Canada). The cTBS stimulation protocol comprised bursts of 3 stimuli at 50Hz that were repeated with a frequency of 5Hz for 40s, resulting in a total of 600 pulses. The implemented cTBS protocol is thought to reduce the excitability of the stimulated brain region for about 60min44.

TMS site localization

We determined the stimulation sites using individual T1-weighted structural scans and TMS Neuronavigation system (BrainSight 2, Rogue Research Inc., Canada). We used data from the CG—acquired before those of the EG—to define the coordinates of the rIFC (MNI xyz 54 34 10) even though they originated from the opposite to the hypothesized contrast driven by ambiguous voices, coordinates that also corresponded to those observed in a previous study on the role of the rIFC for emotional judgments22 and this location also overlapped with vocal judgments in general28 and vocal emotion processing in the inferior frontal gyrus19. This target region was therefore favored for these reasons. For each participant, we transformed the rIFC peak coordinates into the native space of the individual structural scan using the parameter estimates for spatial normalization of the anatomical scan performed in SPM12. The TMS coil was positioned tangentially to the cortical surface over the rIFG, with the handle perpendicular to the rIFC. As a control site we used the vertex, which was defined as the meeting point of the pre- and post-central sulcus in the interhemispheric fissure. For the control group, the TMS coil was positioned tangentially to the cortical surface over vertex (MNI xyz [0, -20, 80]), with the handle pointing in a posterior direction (see Figure 2D).

MRI data acquisition

Imaging data acquisition was performed at the Laboratory for Social and Neural Systems research of the University of Zürich, on a Philips Achieva 3T whole-body scanner equipped with an eight channel MR head coil. Four runs of 10 min each were collected for each participant (1 base run, 3 experimental runs). Each run contained 200 volumes (voxel size = 3 x 3 x 3mm3, 0.5mm gap, matrix size = 80 x 80, TR/TE = 2100/30ms, flip angle = 79, parallel imaging factor = 1.5, 35 slices acquired in ascending order for full coverage of the brain). High-resolution T1-weighted 3D turbo field echo structural scans were acquired and used for image registration and normalization (181 sagittal slices, matrix size = 256 x 256, voxel size = 1mm3, TR/TE/TI = 8.3/2.26/181ms).

MRI data analysis

Functional images were analyzed with Statistical Parametric Mapping software (SPM12, Wellcome Trust Centre for Neuroimaging, London, UK, http://www.fil.ion.ucl.ac.uk/spm). Preprocessing steps included realignment to the first volume of the time series, slice timing, normalization to the Montreal Neurological Institute (MNI)95 space using the DARTEL toolbox96 and spatial smoothing with an isotropic Gaussian filter of 8mm full width at half maximum. To remove low frequency components, we used a high-pass filter with a cutoff frequency of 1/128s. Anatomical locations were defined with a standardized MNI coordinate database using xjView toolbox (https://www.nitrc.org/projects/xjview).

Voice-sensitive areas localizer task

For the voice-sensitive areas localizer task, a general linear model was used to compute first-level statistics, in which each run was modelled by using a run function and was convolved with the hemodynamic response function, time-locked to the onset of each run. Separate regressors were created for each condition (vocal and non-vocal; Condition factor). Finally, six motion parameters were included as regressors of no interest to account for movement in the data. The Condition regressors were used to compute simple contrasts for each participant, leading to a main effect of vocal and non-vocal material at the first-level of analysis [1 0] for vocal, [0 1] for non-vocal. These simple contrasts were then taken to a flexible factorial second-level analysis in which there were two factors: the Participant factor with independence set to yes, variance set to unequal and the Condition factor with independence set to no, variance set to unequal.

Vocal affect categorization task

For the experimental runs (cTBSpost) and the base run (cTBSpre) of the main task, we used a first-level general linear model, in which each stimulus display was modelled by using a stick function and was convolved with the hemodynamic response function. Events were time-locked to the onset of the voice stimuli. Separate regressors were created for each condition of interest (five conditions with percentage of anger/percentage of fear: 90/10, 70/30, 50/50, 30/70, 10/90 or respectively A90, A70, AF50, F70, F90) and for neutral voices (regressor of no-interest). We thus had five regressors of interest including 24 trials each per run (base and experimental runs, respectively cTBSpre and cTBSpost runs) and 72 trials each in total for the 3 experimental runs, in addition to the neutral voice condition as non-interest regressors (10 trials per run, 40 trials in total). Moreover, six motion parameters were included as regressors of no interest to account for movement in the data. Our design matrix was therefore as follows: Anger/Fear 90/10 (A90), Anger/Fear 70/30 (A70), Anger/Fear 50/50 (AF50), Anger/Fear 30/70 (F70), Anger/Fear 10/90 (F90), Neutral, Movement parameters, Constant term; 11 columns in total per run. We therefore had four sessions per model (base run, experimental run 1, 2, 3), leading to 44 columns in the design matrix. Each regressor of interest (Anger/Fear conditions) was used to compute contrasts for each participant (first-level statistics). First-level contrasts were computed to highlight the difference between difficult (more ambiguous) and easier (less ambiguous) trials for the emotional categorization of voices in all runs, especially the experimental runs. This procedure was decided based on the fact that our TMS inhibitory effect of the rIFC would last approximately 60 min44 and our runs were 10 min each for a total of 30 min, thus clearly within the bounds of the TMS effect. The contrast of interest was therefore the following: [A50 > A90, A70, F90, F70] x [cTBSpost > cTBSpre] (contrast vector: -2 -1 6 -1 -2 for each experimental run) and we computed its inverse regarding morphing, [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] (contrast vector: 2 1 -6 1 2 for each experimental run) for each group separately as well as a vector specifically for the most ambiguous voices ([AF50] x [cTBSpost > cTBSpre], contrast vector: 0 0 1 0 0 for each experimental run). The weighted contrasts were used to better characterize the level of morphing of the voice stimuli as a function of the BOLD signal.

First-level contrast results of each participant were then averaged by group at the second-level using a two-sample t-test analysis. Using this procedure allowed us to test for an effect of cTBS stimulation over the vertex (CG) as opposed to cTBS over the rIFC (EG), our region of interest thought to be responsible for sensitive, accurate vocal emotional judgments. We therefore looked at the interaction between our contrasts of interest between groups: [CG > EG] x [AF50 > A90, A70, F90, F70] x [cTBSpost > cTBSpre] (contrast vector: 1 -1), [CG > EG] x [A90, A70, F90, F70 > AF50] x [cTBSpost > cTBSpre] (contrast vector: -1 1), [CG > EG] x [AF50] x [cTBSpost > cTBSpre] (contrast vector: 1 -1). Second-level statistical analyses of the main task assumed that Participants (Factor 1) were independent whereas Conditions (Factor 2) were not. Variance estimation was set to unequal for all factors in order to consider the inhomogeneous variance of the data.

For the voice-sensitive areas localizer task, simple contrasts were then taken to a flexible factorial second-level analysis in which there were two factors: the Participant factor with independence set to yes, variance set to unequal and the Voice factor with independence set to no, variance set to unequal.

All wholebrain activations are reported at a threshold of p<.005 (uncorrected) and a cluster extent threshold of k > 59 voxels, equivalent to a Family-Wise Error correction for multiple comparison of p<.05 at the cluster level. This threshold was based on the final FWHM of the data (11.9, 11.9, 11.1 mm), using the ‘3dClustSim’ function in AFNI (http://afni.nimh.nih.gov/afni) software97, using a non-parametric method with 10’000 iterations to estimate the necessary cluster extent thresholding for side-to-side voxels (NN-2 option). ‘3dClustSim’ reports a cluster extent threshold for each specified statistical p-value and follows the assumption that neighboring voxels are part of a similar functional response pattern, rather than a completely different and independent measure as implied by the family-wise error correction at the voxel level. Inferior frontal cortex was delineated (Figure 2) using the automated anatomical labelling (‘aal’) atlas98.

Functional connectivity analysis

Seed-to-seed functional analysis was performed for all runs (base and experimental) using the CONN toolbox37 version 19.b implemented in Matlab 9.0 (The MathWorks, Inc., Natick, MA, USA). Functional connectivity analyses were computed using as seeds each region of interest (ROI) overlapping with results from studies targeting the decoding of emotional prosody in the lateral, superior temporal cortex28, medial temporal lobe10 and more specifically the role of the IFC in vocal emotion processing19,22. We therefore ended up including eight ROI (bilateral IFC, bilateral anterior, mid and posterior STC, bilateral amygdala) in our seed-to-seed, generalized psychophysiological interaction analysis. Spurious sources of noise were estimated and removed using the automated toolbox preprocessing algorithm, and the residual BOLD time-series was band-pass filtered using a low frequency window (0.008 < f < 0.09 Hz). Correlation maps were then created for each condition of interest by taking the residual BOLD time-course for each condition from atlas regions of interest and computing bivariate Pearson’s correlation coefficients between the time courses of each voxel of each ROI of the atlas, averaged by ROI. We used generalized psychophysiological interaction (gPPI) measures, representing the level of task-modulated connectivity between ROI or between ROI and voxels. gPPI is computed using a separate multiple regression model for each target (ROI). Each model includes three predictors: 1) task effects convolved with a canonical hemodynamic response function (psychological factor); 2) each seed ROI BOLD time series (physiological factor) and 3) the interaction term between the psychological and the physiological factors, the output of which is regression coefficients associated with this interaction term. Finally, group-level analyses were performed on these regression coefficients to assess for main effects within-group for contrasts of interest in seed-to-seed and seed-to-voxel analyses. Results of these analyses reveal only undirected and functional, but not causal and/or directed connectivity. Type I error was controlled by the use of seed-level FDR correction with p<.05 two tailed to correct for multiple comparison.

Limitations

While we did our best to design a study and methodology specifically tailored to answer our research question, we should mention several crucial limitations that could definitely have impacted our results and/or absence thereof. As a first limitation, sample size should be discussed: in fact, even though we calculated the sample size as a function of 80% power with a conservative alpha value of 5%, we had some dropout and power was reduced to about 68.5%. Therefore, we could have added more participants to each group if the study was not limited in time and especially funding. With this dropout in mind, the minimum detectable effect size is ~73.4, meaning that only strong effects could reliably be detected in our study with the final sample size. For this reason, the null findings should be treated very carefully and not interpreted as an absence of result, but rather as non-conclusive data due to low power. Second, we cannot exclude that other regions within or outside of the inferior frontal cortex may play a causal role in making decisions on emotionally ambiguous voices—nor can we exclude a significant degree of IFC heterogeneity and anatomical variations between participants. Again, this would have implied other groups of participants—one group per region to be stimulated by cTBS—and it was impossible budget-wise but such procedure should definitely be addressed by future work. Also, the fact that the clearest and most emotionally ambiguous voices were impacted in the mSTC by cTBS but not more intermediate levels should be investigated in detail. The fact that we altered right IFC, but not left of simultaneously both left and right IFC in this study also raises some questions as to plasticity and especially laterality effects linked to the cTBS procedure. This aspect could also explain why wholebrain differences between groups were only significant in the right but not the left hemisphere. Linked to that aspect, we should also consider that cTBS over the vertex for the CG could have altered the default mode network, which could in turn modify the assessment of affect-related content. The cTBS effect could also have decayed faster than expected, leading to a more modest altering of right IFC functioning throughout each session. The existence of carry-over effects could also have clouded the results, with trial sequences and potentially inappropriate randomization being of particular relevance here. Third, one can wonder whether other emotions would have led to similar observations, namely positive or more complex emotions. Our choice of morphing angry and fearful voices was however made for a smooth morphing process. Therefore, we cannot be sure that the observed neural effects of cTBS would generalize to other emotions or other emotion categories. Lastly, we cannot exclude fundamental differences between our groups and between participants, independently of the TMS procedure, related to decision-making and vocal affect processing. This aspect seems to emerge from the behavioral response data, especially. Different speed-accuracy tradeoff strategies may also have been used by the two groups. Such interindividual differences could in fact be observed and should obviously be the main topic of future studies on the matter.

Conclusion

Our data may unfortunately not provide a more detailed and mechanistic picture of the functional role of the right IFC pars triangularis in social sound cognition, especially in voice signal classification along socio-affective dimensions. We could not draw data-backed conclusions between our groups of participants by applying cTBS to the right IFC in our human sample. Decisional responses on emotionally clear and emotionally ambiguous voices become faster immediately after cTBS, but this result may only highlight learning. More importantly, no significant triple interaction between groups, runs and morphing could be observed behaviorally. Neuroimaging data allow for a more functional reorganization of brain activity following right IFC cTBS but for emotionally clear rather than for emotionally ambiguous voices, showing significantly lower activity in the auditory cortex, and higher fronto-limbic connectivity—these results are again opposite to our hypotheses. Future work should clarify the causal role(s) of the bilateral IFC, as well as implement continuous and intermittent TBS procedures combined with fMRI to disentangle the behavioral and neural effects of alteration vs. enhancement of IFC activity in the context of voice-related decision making, especially in emotionally ambiguous situations.

Supplementary Information

Author contributions

LC helped program the tasks, collected the behavioral and neuroimaging data, analyzed the data, created and edited the figures and wrote the manuscript. MM collected TMS data and neuroimaging data and helped write the methods of the manuscript. DG helped design the study. CR helped design the study, especially the neuroimaging part to make it compatible with the TMS procedure. SF designed the study, programmed the tasks, collected part of the data and helped design the figures. All authors reviewed and edited the manuscript.

Funding

The present study was supported by the Swiss National Science Foundation (SNSF 105314_146559/1) and the National Center for Affective Sciences (51NF40-104897). SF receives additional support from the SNSF (PP00P1_157409/1 and PP00P1_183711/1).

Data availability

The data and codes can be found in the open, FAIR-compliant ‘Yareta’ repository. The specific, permanent link for raw behavioral and MRI data/code is the following: 10/nxf4, while the link for the concatenated behavioral data and R analysis code is: 10.26037/yareta:pxsatt3fcjbd7m5m4542kmk4mq.

Declarations

Comepting interest

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Staib, M. & Frühholz, S. Cortical voice processing is grounded in elementary sound analyses for vocalization relevant sound patterns. Progress Neurobiol.200, 101982 (2020). [DOI] [PubMed] [Google Scholar]
  • 2.Rauschecker, J. P. & Scott, S. K. Maps and streams in the auditory cortex: nonhuman primates illuminate human speech processing. Nat. Neurosci.12, 718–724 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Roswandowitz, C., Swanborough, H. & Frühholz, S. Categorizing human vocal signals depends on an integrated auditory-frontal cortical network. Hum. Brain Mapp.42, 1503 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Ceravolo, L., Frühholz, S. & Grandjean, D. Proximal vocal threat recruits the right voice-sensitive auditory cortex. Soc. Cognit. Affect. Neurosci.11, 793–802 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Ceravolo, L., Frühholz, S., Pierce, J., Grandjean, D. & Péron, J. Basal ganglia and cerebellum contributions to vocal emotion processing as revealed by high-resolution fMRI. Sci. Rep.11, 10645 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Dricu, M., Ceravolo, L., Grandjean, D. & Frühholz, S. Biased and unbiased perceptual decision-making on vocal emotions. Sci. Rep.7, 16274. 10.1038/s41598-017-16594-w (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Frühholz, S., Ceravolo, L. & Grandjean, D. Specific brain networks during explicit and implicit decoding of emotional prosody. Cereb. Cortex22, 1107–1117. 10.1093/cercor/bhr184 (2012). [DOI] [PubMed] [Google Scholar]
  • 8.Frühholz, S. & Grandjean, D. Multiple subregions in superior temporal cortex are differentially sensitive to vocal expressions: a quantitative meta-analysis. Neurosci. Biobehav. Rev.37, 24–35. 10.1016/j.neubiorev.2012.11.002 (2013). [DOI] [PubMed] [Google Scholar]
  • 9.Swanborough, H., Staib, M. & Frühholz, S. Neurocognitive dynamics of near-threshold voice signal detection and affective voice evaluation. Sci. Adv.6, eabb3884 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Frühholz, S., Trost, W. & Grandjean, D. The role of the medial temporal limbic system in processing emotions in voice and music. Progress Neurobiol.123, 1–17 (2014). [DOI] [PubMed] [Google Scholar]
  • 11.Frühholz, S., Trost, W. & Kotz, S. A. The sound of emotions—towards a unifying neural network perspective of affective sound processing. Neurosci. Biobehav. Rev.68, 96–110 (2016). [DOI] [PubMed] [Google Scholar]
  • 12.Dricu, M., Ceravolo, L., Grandjean, D. & Frühholz, S. Biased and unbiased perceptual decision-making on vocal emotions. Sci. Rep.7, 1–16 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Dricu, M. & Frühholz, S. A neurocognitive model of perceptual decision-making on emotional signals. Hum. Brain Mapp.41, 1532–1556 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Frühholz, S. & Schweinberger, S. R. Nonverbal auditory communication-evidence for integrated neural systems for voice signal production and perception. Progress Neurobiol.199, 101948 (2020). [DOI] [PubMed] [Google Scholar]
  • 15.Fecteau, S., Armony, J. L., Joanette, Y. & Belin, P. Sensitivity to voice in human prefrontal cortex. J. Neurophysiol.94, 2251–2254 (2005). [DOI] [PubMed] [Google Scholar]
  • 16.Schirmer, A., Zysset, S., Kotz, S. A. & von Cramon, D. Y. Gender differences in the activation of inferior frontal cortex during emotional speech perception. NeuroImage21, 1114–1123 (2004). [DOI] [PubMed] [Google Scholar]
  • 17.Frühholz, S., Ceravolo, L. & Grandjean, D. Specific brain networks during explicit and implicit decoding of emotional prosody. Cerebral cortex22, 1107–1117 (2012). [DOI] [PubMed] [Google Scholar]
  • 18.Frühholz, S. & Grandjean, D. Towards a fronto-temporal neural network for the decoding of angry vocal expressions. Neuroimage62, 1658–1666 (2012). [DOI] [PubMed] [Google Scholar]
  • 19.Frühholz, S. & Grandjean, D. Processing of emotional vocalizations in bilateral inferior frontal cortex. Neurosci. Biobehav. Rev.37, 2847–2855 (2013). [DOI] [PubMed] [Google Scholar]
  • 20.Bestelmeyer, P. E., Maurage, P., Rouger, J., Latinus, M. & Belin, P. Adaptation to vocal expressions reveals multistep perception of auditory emotion. J. Neurosci.34, 8098–8105 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Hoekert, M., Bais, L., Kahn, R. S. & Aleman, A. Time course of the involvement of the right anterior superior temporal gyrus and the right fronto-parietal operculum in emotional prosody perception. PLoS One3, e2244 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Hoekert, M., Vingerhoets, G. & Aleman, A. Results of a pilot study on the involvement of bilateral inferior frontal gyri in emotional prosody perception: an rTMS study. BMC Neurosci.11, 93 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Debracque, C., Ceravolo, L., Clay, Z., Grandjean, D. & Gruber, T. Categorization and discrimination of human and non-human primate affective vocalizations: Investigation of frontal cortex activity through fNIRS. Imaging Neurosci.3, imag_a_00480 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Ceravolo, L., Debracque, C., Gruber, T. & Grandjean, D. Sensitivity of the human temporal voice areas to nonhuman primate vocalizations. Elife10.7554/eLife.108795.1 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Ceravolo, L., Debracque, C., Pool, E., Gruber, T. & Grandjean, D. Frontal mechanisms underlying primate calls recognition by humans. Cereb. Cortex Commun.4, tgad019 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Averbeck, B. B. & Romanski, L. M. Probabilistic encoding of vocalizations in macaque ventral lateral prefrontal cortex. J. Neurosci.26, 11023–11033 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Cohen, Y. et al. A functional role for the ventrolateral prefrontal cortex in non-spatial auditory cognition. Proc. Natl. Acad. Sci.106, 20045–20050 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Schirmer, A. & Kotz, S. A. Beyond the right hemisphere: brain mechanisms mediating vocal emotional processing. Trends Cognit. Sci.10, 24–30 (2006). [DOI] [PubMed] [Google Scholar]
  • 29.Steiner, F., Bobin, M. & Frühholz, S. Auditory cortical micro-networks show differential connectivity during voice and speech processing in humans. Commun. Biol.4, 1–10 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Oberman, L., Edwards, D., Eldaief, M. & Pascual-Leone, A. Safety of theta burst transcranial magnetic stimulation: a systematic review of the literature. J. Clin. Neurophysiol.28, 67 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Banissy, M. J. et al. Suppressing sensorimotor activity modulates the discrimination of auditory emotions but not speaker identity. J. Neurosci.30, 13552–13557 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Agnew, Z. K., Banissy, M. J., McGettigan, C., Walsh, V. & Scott, S. K. Investigating the neural basis of theta burst stimulation to premotor cortex on emotional vocalization perception: a combined TMS-fMRI study. Front. Hum. Neurosci.12, 150 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Jiahui, G. et al. Normal voice processing after posterior superior temporal sulcus lesion. Neuropsychologia105, 215–222 (2017). [DOI] [PubMed] [Google Scholar]
  • 34.Whiting, C. M., Kotz, S. A., Gross, J., Giordano, B. L. & Belin, P. The perception of caricatured emotion in voice. Cognition200, 104249 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Giordano, B. L. et al. The representational dynamics of perceived voice emotions evolve from categories to dimensions. Nat. Hum. Behav.5, 1203–1213 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Grandjean, D. Brain networks of emotional prosody processing. Emotion Rev.13, 34 (2020). [Google Scholar]
  • 37.Whitfield-Gabrieli, S. & Nieto-Castanon, A. Conn: a functional connectivity toolbox for correlated and anticorrelated brain networks. Brain Connect2, 125–141. 10.1089/brain.2012.0073 (2012). [DOI] [PubMed] [Google Scholar]
  • 38.Belin, P., Zatorre, R. J., Lafaille, P., Ahad, P. & Pike, B. Voice-selective areas in human auditory cortex. Nature403, 309–312 (2000). [DOI] [PubMed] [Google Scholar]
  • 39.Jung, J., Bungert, A., Bowtell, R. & Jackson, S. R. Vertex stimulation as a control site for transcranial magnetic stimulation: a concurrent TMS/fMRI study. Brain Stimul.9, 58–64 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Suran, T., Rumiati, R. I. & Piretti, L. The contribution of the left inferior frontal gyrus in affective processing of social groups. Cognit. Neurosci.10, 186–195 (2019). [DOI] [PubMed] [Google Scholar]
  • 41.Grossmann, T., Oberecker, R., Koch, S. P. & Friederici, A. D. The developmental origins of voice processing in the human brain. Neuron65, 852–858 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Romanski, L. M. & Averbeck, B. B. The primate cortical auditory system and neural representation of conspecific vocalizations. Ann. Rev. Neurosci.32, 315–346 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Cohen, Y. E., Hauser, M. D. & Russ, B. E. Spontaneous processing of abstract categorical information in the ventrolateral prefrontal cortex. Biol. Lett.2, 261–265 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Huang, Y.-Z., Edwards, M. J., Rounis, E., Bhatia, K. P. & Rothwell, J. C. Theta burst stimulation of the human motor cortex. Neuron45, 201–206 (2005). [DOI] [PubMed] [Google Scholar]
  • 45.Brück, C., Kreifelts, B. & Wildgruber, D. Emotional voices in context: a neurobiological model of multimodal affective information processing. Phys. Life Rev.8, 383–403 (2011). [DOI] [PubMed] [Google Scholar]
  • 46.Ethofer, T. et al. Cerebral pathways in processing of affective prosody: a dynamic causal modeling study. Neuroimage30, 580–587 (2006). [DOI] [PubMed] [Google Scholar]
  • 47.Verbruggen, F., Aron, A. R., Stevens, M. A. & Chambers, C. D. Theta burst stimulation dissociates attention and action updating in human inferior frontal cortex. Proc. Natl. Acad. Sci.107, 13966–13971 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Rogers, J. C. & Davis, M. H. Inferior frontal cortex contributions to the recognition of spoken words and their constituent speech sounds. J. Cognit. Neurosci.29, 919–936 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Xie, X. & Myers, E. Left inferior frontal gyrus sensitivity to phonetic competition in receptive language processing: a comparison of clear and conversational speech. J. Cognit. Neurosci.30, 267–280 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Zhuang, J., Tyler, L. K., Randall, B., Stamatakis, E. A. & Marslen-Wilson, W. D. Optimally efficient neural systems for processing spoken language. Cereb. Cortex24, 908–918. 10.1093/cercor/bhs366 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Zander, T., Horr, N. K., Bolte, A. & Volz, K. G. Intuitive decision making as a gradual process: investigating semantic intuition-based and priming-based decisions with fMRI. Brain Behav.6, e00420. 10.1002/brb3.420 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Dippel, G. & Beste, C. A causal role of the right inferior frontal cortex in implementing strategies for multi-component behaviour. Nat. Commun.6, 6587 (2015). [DOI] [PubMed] [Google Scholar]
  • 53.Gruber, T. et al. Human discrimination and categorization of emotions in voices: a functional near-infrared spectroscopy (fNIRS) study. Front. Neurosci.14, 570 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Lupyan, G., Mirman, D., Hamilton, R. & Thompson-Schill, S. L. Categorization is modulated by transcranial direct current stimulation over left prefrontal cortex. Cognition124, 36–49 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Nastase, S., Iacovella, V. & Hasson, U. Uncertainty in visual and auditory series is coded by modality-general and modality-specific neural systems. Hum. Brain Mapp.35, 1111–1128. 10.1002/hbm.22238 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Toelch, U., Bach, D. R. & Dolan, R. J. The neural underpinnings of an optimal exploitation of social information under uncertainty. Soc. Cognit. Affect. Neurosci.9, 1746–1753. 10.1093/scan/nst173 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Keuken, M. et al. The role of the left inferior frontal gyrus in social perception: an rTMS study. Brain Res.1383, 196–205 (2011). [DOI] [PubMed] [Google Scholar]
  • 58.Nahas, Z. et al. Unilateral left prefrontal transcranial magnetic stimulation (TMS) produces intensity-dependent bilateral effects as measured by interleaved BOLD fMRI. Biol. Psychiatr.50, 712–720 (2001). [DOI] [PubMed] [Google Scholar]
  • 59.Mix, A., Benali, A., Eysel, U. T. & Funke, K. Continuous and intermittent transcranial magnetic theta burst stimulation modify tactile learning performance and cortical protein expression in the rat differently. Euro. J. Neurosci.32, 1575–1586. 10.1111/j.1460-9568.2010.07425.x (2010). [DOI] [PubMed] [Google Scholar]
  • 60.Hartwigsen, G. et al. Perturbation of the left inferior frontal gyrus triggers adaptive plasticity in the right homologous area during speech production. Proc. Natl. Acad. Sci.110, 16402–16407 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Liakakis, G., Nickel, J. & Seitz, R. Diversity of the inferior frontal gyrus—a meta-analysis of neuroimaging studies. Behav. Brain Res.225, 341–347 (2011). [DOI] [PubMed] [Google Scholar]
  • 62.Jacobson, L., Javitt, D. C. & Lavidor, M. Activation of inhibition: diminishing impulsive behavior by direct current stimulation over the inferior frontal gyrus. J. Cognit. Neurosci.23, 3380–3387 (2011). [DOI] [PubMed] [Google Scholar]
  • 63.Aron, A. R., Robbins, T. W. & Poldrack, R. A. Inhibition and the right inferior frontal cortex. Trends Cognit. Sci.8, 170–177 (2004). [DOI] [PubMed] [Google Scholar]
  • 64.Aron, A. R., Robbins, T. W. & Poldrack, R. A. Inhibition and the right inferior frontal cortex: one decade on. Trends Cognit. Sci.18, 177–185 (2014). [DOI] [PubMed] [Google Scholar]
  • 65.Gilovich, T. & Griffin, D. Heuristics and biases: The psychology of intuitive judgment (Cambridge university press, 2002). [Google Scholar]
  • 66.Kahneman, D. Maps of bounded rationality: a perspective on intuitive judgment and choice. Nobel Prize Lect.8, 351–401 (2002). [Google Scholar]
  • 67.Hogarth, R. M. Intuition: a challenge for psychological research on decision making. Psychol. Inq.21, 338–353. 10.1080/1047840X.2010.520260 (2010). [Google Scholar]
  • 68.Hammond, K. R., Hamm, R. M., Grassia, J. & Pearson, T. Direct comparison of the efficacy of intuitive and analytical cognition in expert judgment. IEEE Trans. Syst. Man Cybern.17, 753–770. 10.1109/TSMC.1987.6499282 (1987). [Google Scholar]
  • 69.Dane, E., Rockmann, K. W. & Pratt, M. G. When should I trust my gut? linking domain expertise to intuitive decision-making effectiveness. Organ. Behav. Hum. Decis. Process.119, 187–194 (2012). [Google Scholar]
  • 70.Ethofer, T. et al. Emotional voice areas: anatomic location, functional properties, and structural connections revealed by combined fMRI/DTI. Cereb. Cortex22, 191–200 (2012). [DOI] [PubMed] [Google Scholar]
  • 71.Fedorenko, E., Ivanova, A. A. & Regev, T. I. The language network as a natural kind within the broader landscape of the human brain. Nat. Rev. Neurosci.25, 1–24 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Ethofer, T. et al. Functional responses and structural connections of cortical areas for processing faces and voices in the superior temporal sulcus. Neuroimage76, 45–56 (2013). [DOI] [PubMed] [Google Scholar]
  • 73.Frühholz, S., Gschwind, M. & Grandjean, D. Bilateral dorsal and ventral fiber pathways for the processing of affective prosody identified by probabilistic fiber tracking. Neuroimage109, 27–34 (2015). [DOI] [PubMed] [Google Scholar]
  • 74.Grandjean, D. et al. The voices of wrath: brain responses to angry prosody in meaningless speech. Nat. Neurosci.8, 145–146 (2005). [DOI] [PubMed] [Google Scholar]
  • 75.Witteman, J., Van Heuven, V. J. & Schiller, N. O. Hearing feelings: a quantitative meta-analysis on the neuroimaging literature of emotional prosody perception. Neuropsychologia50, 2752–2763 (2012). [DOI] [PubMed] [Google Scholar]
  • 76.Sokhi, D. S., Hunter, M. D., Wilkinson, I. D. & Woodruff, P. W. Male and female voices activate distinct regions in the male brain. Neuroimage27, 572–578 (2005). [DOI] [PubMed] [Google Scholar]
  • 77.Charest, I., Pernet, C., Latinus, M., Crabbe, F. & Belin, P. Cerebral processing of voice gender studied using a continuous carryover fMRI design. Cereb. Cortex23, 958–966 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Selosse, G., Grandjean, D. & Ceravolo, L. Neural correlates of embodied and vibratory mechanisms associated with emotional prosody production. Soc. Cognit. Affect. Neurosci.20, 084 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Pollatos, O., Gramann, K. & Schandry, R. Neural systems connecting interoceptive awareness and feelings. Hum. Brain Mapp.28, 9–18 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Zhang, Y. et al. The roles of subdivisions of human insula in emotion perception and auditory processing. Cereb. Cortex29, 517–528 (2019). [DOI] [PubMed] [Google Scholar]
  • 81.Wiech, K. et al. Anterior insula integrates information about salience into perceptual decisions about pain. J. Neurosci.30, 16324–16331 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Lutz, A., McFarlin, D. R., Perlman, D. M., Salomons, T. V. & Davidson, R. J. Altered anterior insula activation during anticipation and experience of painful stimuli in expert meditators. Neuroimage64, 538–546 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Pannese, A., Grandjean, D. & Frühholz, S. Amygdala and auditory cortex exhibit distinct sensitivity to relevant acoustic features of auditory emotions. Cortex85, 116–125 (2016). [DOI] [PubMed] [Google Scholar]
  • 84.Sander, D. et al. Emotion and attention interactions in social cognition: brain regions involved in processing anger prosody. Neuroimage28, 848–858 (2005). [DOI] [PubMed] [Google Scholar]
  • 85.Faul, F., Erdfelder, E., Lang, A.-G. & Buchner, A. G* Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behav. Res. Methods39, 175–191 (2007). [DOI] [PubMed] [Google Scholar]
  • 86.Belin, P., Fillion-Bilodeau, S. & Gosselin, F. The Montreal Affective Voices: A validated set of nonverbal affect bursts for research on auditory affective processing. Behav. Res. Methods40, 531 (2008). [DOI] [PubMed] [Google Scholar]
  • 87.Aguirre, G. K. Continuous carry-over designs for fMRI. Neuroimage35, 1480–1494 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Hill, C. A. et al. A causal account of the brain network computations underlying strategic social behavior. Nat. Neurosci.20, 1142–1149 (2017). [DOI] [PubMed] [Google Scholar]
  • 89.Obeso, I., Moisa, M., Ruff, C. C. & Dreher, J.-C. A causal role for right temporo-parietal junction in signaling moral conflict. Elife7, e40671 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Kuznetsova, A., Brockhoff, P. B., & Christensen, R. H. lmerTest package: tests in linear mixed effects models. Journal of statistical software, 82, 1–26 (2017).
  • 91.Bates, D., Mächler, M., Bolker, B., & Walker, S. Fitting linear mixed-effects models using lme4. Journal of statistical software, 67, 1–48 (2015).
  • 92.Team, R. C. R: A Language and Environment for Statistical Computing https://www. R-project. org. (2014).
  • 93.Fox, J. et al. The car package. R foundation for statistical computing (2007).
  • 94.Barton, K. & Barton, M. K. Package ‘mumin’. Version1, 18 (2015). [Google Scholar]
  • 95.Collins, D., Neelin, P., Peters, T. & Evans, A. Automatic 3D intersubject registration of MR volumetric data in standardized talairach space. J. Comput. Assist Tomogr.18, 192–205 (1994). [PubMed] [Google Scholar]
  • 96.Ashburner, J. A fast diffeomorphic image registration algorithm. NeuroImage38, 95–113. 10.1016/j.neuroimage.2007.07.007 (2007). [DOI] [PubMed] [Google Scholar]
  • 97.Cox, R. W. AFNI: software for analysis and visualization of functional magnetic resonance neuroimages. Comput. Biomed. Res.29, 162–173 (1996). [DOI] [PubMed] [Google Scholar]
  • 98.Tzourio-Mazoyer, N. et al. Automated anatomical labeling of activations in SPM using a macroscopic anatomical parcellation of the MNI MRI single-subject brain. Neuroimage15, 273–289 (2002). [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The data and codes can be found in the open, FAIR-compliant ‘Yareta’ repository. The specific, permanent link for raw behavioral and MRI data/code is the following: 10/nxf4, while the link for the concatenated behavioral data and R analysis code is: 10.26037/yareta:pxsatt3fcjbd7m5m4542kmk4mq.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES