Abstract
Recent research has extensively reported the phenomenon of inter-brain neural coupling between speakers and listeners during speech communication. Yet, the specific speech processes underlying this neural coupling remain elusive. To bridge this gap, this study estimated the correlation between the temporal dynamics of speaker–listener neural coupling with speech features, utilizing two inter-brain datasets accounting for different noise levels and listener’s language experiences (native vs. non-native). We first derived time-varying speaker–listener neural coupling, extracted acoustic feature (envelope) and semantic features (entropy and surprisal) from speech, and then explored their correlational relationship. Our findings reveal that in clear conditions, speaker–listener neural coupling correlates with semantic features. However, as noise increases, this correlation is only significant for native listeners. For non-native listeners, neural coupling correlates predominantly with acoustic feature rather than semantic features. These results revealed how speaker–listener neural coupling is associated with the acoustic and semantic features under various scenarios, enriching our understanding of the inter-brain neural mechanisms during natural speech communication. We therefore advocate for more attention on the dynamic nature of speaker–listener neural coupling and its modeling with multilevel speech features.
Keywords: speaker–listener neural coupling, acoustic feature, semantic feature, functional near-infrared spectroscopy (fNIRS), naturalistic speech
Introduction
Interpersonal neural coupling, representing the coupling or synchrony in neural activities between individuals, emerges prominently during verbal communication between speakers and listeners (Stephens et al. 2010, Jiang et al. 2012, Schoot et al. 2016). It is observed across various demographics, from infants to the elderly (Wass et al. 2020, Li et al. 2021, Liu et al. 2021), and among individuals with different linguistic experiences, including non-native speakers (Li et al. 2023). Such speaker–listener neural coupling emerges from extensive brain regions, including the perisylvian language network, as well as areas in the prefrontal cortex and subcortical structures (Yeshurun et al. 2021, Kelsen et al. 2022). Moreover, the strength of this coupling could predict the quality or outcome of speech communication (Stephens et al. 2010, Li et al. 2021). Accordingly, speaker–listener neural coupling is proposed as a fundamental mechanism underlying interpersonal transmission of speech information between speakers and listeners (Stephens et al. 2010, Jiang et al. 2021) and plays a mechanistic role in interpersonal communication (Hasson and Frith 2016). However, considering the rich and multiple information, e.g. acoustics and semantics, contained in speech, the intricate speech processes underlying speaker–listener neural coupling remain largely elusive.
Several studies have attempted to address this issue by manipulating particular speech features to discern their influence on neural coupling (Stephens et al. 2010, Dikker et al. 2014, Liu et al. 2019). For instance, Dikker et al. (2014) manipulated the semantic predictability of the speech materials and observed increased speaker–listener neural coupling in the posterior superior temporal gyrus when the speech was more semantically predictable. However, the relationship between speaker–listener neural coupling and the processing of speech features is still not well understood. This is primarily because existing studies have focused on the coupling strength averaged throughout speech, neglecting the temporal dynamics of speaker–listener neural coupling. Notably, contemporary findings literally suggest that this neural coupling is not constant during natural speech communication (Mayo and Gordon 2020). Some researchers further proposed a hypothesis that the strength of speaker–listener neural coupling should vary with fluctuations of speech features during communication (Perez and Davis 2023). However, no empirical evidence has validated this hypothesis.
To capture this dynamic interplay, we propose a new approach that seeks to correlate the temporal fluctuations of speaker–listener neural coupling with the dynamics of speech features. The time-varying nature of inter-brain neural coupling and its relationship with external stimuli or experimental tasks have been long noted by researchers. Hasson et al. (2004) have revealed that the activation of the fusiform gyrus, where significant inter-brain neural coupling emerged among participants watching the same movie, peaked at face-related images. It indicated that this neural coupling originated from the dynamic processing of facial stimuli. King-Casas et al. (2005) have discovered a more synchronized inter-brain relationship between communicators over the ongoing trials, which was associated with changes in interpersonal trust during social interaction. More recently, some research has attempted to segment the communication or interaction durations into different events and then compare the strength of neural coupling during corresponding periods (Jiang et al. 2015, Pan et al. 2018). For example, Jiang et al. (2015) classified interaction processing into two types of behaviors, i.e. verbal communication and nonverbal communication, which rapidly alternated in seconds. This study revealed that speaker–listener neural coupling from the frontal lobe was significantly higher during verbal communication as compared to nonverbal communication, suggesting its specificity to verbal content. Following the same logic, given the great fluctuation of multiple speech features throughout natural speech, the modeling between the temporal dynamics of speaker–listener neural coupling and multilevel speech features would deepen the knowledge about the speech processing behind this inter-brain neural mechanism. However, past research has not explored their relationship, possibly due to the great challenge of encoding high-level speech features, including semantics. The recent development of the natural language processing (NLP) algorithm has demonstrated the potential to provide a quantitative description of semantics in a manner similar to human understanding and computation (Goldstein et al. 2022; Grand et al. 2022), thus paving the way for computing its correlational relationship with neural coupling.
The aim of the present study is to investigate the intricate speech processing involved in the inter-brain neural mechanisms during interpersonal communication. To achieve this, we computed the correlation between the temporal fluctuations of speaker–listener neural coupling with speech features throughout interpersonal communication, utilizing two inter-brain datasets collected from speakers and listeners (Li et al. 2021, 2023). These datasets were based on functional near-infrared spectroscopy (fNIRS) measurement and accounted for different noise levels and listener’s language experiences (native vs. non-native), which enabled an in-depth examination of the time-varying speaker–listener neural coupling under varying conditions. In this study, both acoustic feature (envelope) and word-level semantic features (surprisal and entropy) were extracted from the speaker’s speech. The semantic features were automatically calculated by NLP algorithms. Then, the correlation coefficients between temporal-varying speaker–listener neural coupling and speech features were examined. It is hypothesized that the time-varying speaker–listener neural coupling significantly correlates with speech features. Moreover, their correlational relationship is modulated by both the noise level and the listener’s language experience. Specifically, referring to previous behavior patterns of people’s robust comprehension of native speech in noisy conditions (Golestani et al. 2013), native speaker–listener neural coupling is hypothesized to correlate with high-level semantic features regardless of the noise levels. In contrast, considering the non-native disadvantage in semantic processing in noisy environments (Scharenborg and van Os 2019), speaker–listener neural coupling for non-native listeners could correlate to semantic features in situations with no or mild noise. When the noise level increases, this coupling may correlate with low-level acoustic features, indicating that non-native listeners could only track the acoustics, instead of the content of the speaker. These explorations are expected to uncover the acoustic and semantic processing underlying speaker–listener neural coupling across various scenarios.
Materials and methods
fNIRS data, stimuli, and experimental design
Data were obtained from two previous inter-brain datasets (Li et al. 2021, 2023). Both datasets adopted the same sequential fNIRS-based inter-brain approach (Redcay and Schilbach 2019), in which both speaker and listener participants were recruited, and the speaker’s experiment was completed prior to the listener’s experiment. The speakers were recruited to give Chinese narratives based on given topics. The listeners were then instructed to listen to the speaker’s narratives under different noise levels. The two datasets shared the same speaker’s experiment and contained two groups of listeners with different language experiences (native vs. non-native Chinese listeners) for the listener’s experiment.
Participants
Speaker participants consisted of six native Chinese speakers (three females and three males; age range 21–24 years old). They were instructed to deliver 90-s Chinese narratives based on given topics. Listener participants consisted of two groups of listeners. The first group of listeners was from the first dataset (Li et al. 2021), which included 15 native Chinese listeners (eight females and seven males; age range 18–24 years old). The second group of listeners was from the second dataset (Li et al. 2023), which included 15 Korean listeners (nine females and six males; age range 18–24 years old). They have learned Chinese as a non-native language for 4–9 years after the age of 12 years, and all passed the Chinese Proficiency Test Level VI (the official Chinese language test for non-native speakers), indicating a high capability of fluent Chinese communication. All participants self-reported being right-handed, with normal hearing and normal or corrected-to-normal vision.
Speaker’s experiment
The speaker’s task was to tell a narrative story in standard Mandarin Chinese based on the daily topics adapted from the National Mandarin Proficiency Test. Examples of the topics were “What is your favorite trip?” and “What is your favorite hobby?”. They were required to speak for ∼90 s. The content of the narratives was based on the speaker’s personal experience, which was thus unfamiliar to the listeners. All audio was recorded with a regular microphone at a sampling rate of 44 100 Hz in a sound-attenuated room. All speakers had received years of professional training in broadcasting or hosting, which guaranteed the quality of the spoken narratives.
A total of 32 Chinese narratives spoken by speakers were selected for the listener’s experiment. Two speakers (one male and one female) contributed eight stories each, and the other four speakers (two males and two females) contributed four stories each. During the listener’s experiments, these speakers’ audios were modified by adding white noise and processed into four different noise levels: no noise (NN) condition and three noisy conditions with signal-to-noise ratio (SNR) equaling 2, −6, and −9 dB. All processed audio was then matched in terms of the average root-mean-square sound pressure level. Besides, choice questions about the content of the narrative were prepared for listeners’ experiments to estimate their comprehension performance in different noise levels.
Listener’s experiment
The experiment design was consistent between the two groups of listeners. Each listener participant was instructed to listen to all 32 Chinese narratives. These narratives were presented in 32 separate trials, with each trial consisting of one narrative played at one of the four noise levels (NN, SNRs of 2, −6, and −9 dB). After listening to the narrative, they had to rate the clarity, intelligibility of the audio on a 7-point Likert scale, and complete four four-choice questions. The order in which the narratives were presented and the assigned noise levels were randomized or counterbalanced across listeners. Notably, both groups of listeners were able to understand the narratives at the highest noise level (−9 dB): the native listener participants had an average accuracy of 71.3% ± 11.2% (mean ± SE) in answering the choice questions, while the non-native listener participants had an average accuracy of 51.0% ± 4.2% (mean ± SE). Both were higher than the chance level of 25%, indicating that the participants had a moderate level of comprehension of the narratives.
fNIRS data
All participants’ fNIRS signals were recorded from the same 36 channels, covering the typical speech-related brain regions located in the bilateral fronto-parieto-temporal cortex and prefrontal cortex. The raw fNIRS data underwent a two-step preprocessing to remove possible motion artifacts. First, the artifact-related principal components were removed, and the remaining principal components were back-projected to reconstruct the cleaned fNIRS signals (Yücel et al. 2014). Secondly, the numerical outliers within the cleaned data were further identified and corrected by a cubic spline interpolation method (Scholkmann et al. 2010). The preprocessed data during the speech task session (speaking or listening) and the resting-state session were then segmented for subsequent speaker–listener neural coupling analysis.
Significant speaker–listener neural coupling
The significant speaker–listener neural coupling was identified based on the following steps (Li et al. 2021, 2023). First, wavelet transform coherence (WTC) was used to estimate the correlation between the listener’s fNIRS data and the corresponding speaker’s fNIRS data. It could return the coherence values as a function of frequency and time (Grinsted et al. 2004) and was done for each trial and for each channel combination. The non-analytic Morlet wavelet (ω0 = 6, smallest scale s0 = 1/6, spacing between scales ds = 0.4875) was used for the WTC calculation. Next, the coherence values were temporally averaged over the entire 90-s narrative and Fisher-z transformed. Then, for each listener, the coherence values of eight trials for each noise level were averaged. The coherence values for the resting-state condition were calculated using a similar approach. The above computation yielded the speaker–listener neural coupling in 1296 channel combinations (36 × 36 channels) × 75 frequency bins (0.01–0.7 Hz) and five conditions (the four noise levels and resting-state condition) for each listener.
At the group level, a repeated-measures ANOVA (rmANOVA, within-participant factor: task condition) was employed to identify whether the speaker–listener neural coupling was modulated by the task condition for each channel combination and each frequency bin. A nonparametric cluster-based permutation method was used to account for the multiple comparisons. It was suggested to be advantageous over the classical correction methods, e.g. false discovery rate (FDR) correction method, for its suitability for neural data by considering the continuity of possible statistical effects over the spectral or spatial dimension (Maris and Oostenveld 2007, Sassenhagen and Draschkow 2019). In the present study, neighboring frequency bins with an uncorrected P-value of <.05 were combined into clusters, with the sum of their F-value as the cluster-level statistics. The cluster was only generated in the spectral dimension and not over the spatial dimension, i.e. channel or channel combinations, to avoid potential confounding of global physiological noise. A null distribution was further obtained by shuffling the labels of the five conditions (n = 1000 permutations). It was formed by the maximum cluster-level statistics from each shuffling and then used for obtaining the P-values for each real cluster (Li et al. 2021). Those clusters whose P-value was <.05 were identified as significant. These clusters were further devoted to post hoc analysis to examine whether their neural coupling was significantly higher at four noise levels than at baseline. The above calculation was carried out separately for both native and non-native listener groups.
Significant speaker–listener neural coupling was found at all four noise levels compared to the resting-state baseline (Li et al. 2021, 2023). Notably, while all speaker–listener neural coupling related to the same frequency band at 0.01–0.032 Hz, distinct inter-brain patterns were identified with significant effects during native and non-native speech processing. For the native context, a total of 24 channel combinations were identified, where the listener’s neural activities from the left inferior frontal gyrus (IFG), right middle temporal gyrus (MTG), and angular gyrus (AG) were significantly coupled to the speaker’s neural activities from a broad range of brain regions, including left superior frontal gyrus, bilateral supramarginal gyrus, bilateral middle frontal gyrus (MFG), bilateral AG, and right postcentral gyrus (postCG). For the non-native context, a total of 10 channel combinations were identified, where the listener’s neural activities from the right MFG, precentral gyrus, postCG, superior temporal gyrus (STG), MTG, and left IFG were coupled to the speaker’s right STG or postCG, showing right lateralization in both the speaker’s and the listener’s neural activities. The brain regions and locations of all channel combinations are shown in Fig. 1d. The speaker–listener neural couplings from these clusters were selected for this study. The detailed information on these clusters, including channel combinations and frequency bands, can be found in the Supplementary material. Notably, the temporal dynamics of speaker–listener neural coupling were extracted over the entire 90-s trial duration without averaging across the time dimension.
Figure 1.

(a) Analysis framework. (b) Speech features extracted from the speaker’s speech. (c) Correlation between speech features. (d) Channel combinations showing significant speaker–listener neural coupling.
Extraction of speech features
Both acoustic and semantic features were extracted from the speaker’s narratives. The audio narratives were first automatically converted into text using the iFlyrec software (iFlytek Co., Ltd, Hefei, Anhui), which was then manually checked for accuracy. The text was then manually segmented into words according to “The Segmentation Guidelines for the Penn Chinese Treebank” (https://repository.upenn.edu/cgi/viewcontent.cgi?article=1038&context=ircs_reports), and the onset time and offset times of each word were also manually labeled in the iFlyrec software, with a minimum labeling timescale of 20 ms.
Acoustic feature
The amplitude envelope was one of the most commonly used acoustic features in recent literature (Brodbeck et al. 2018, Armeni et al. 2019, Broderick et al. 2019, Etard and Reichenbach 2019, Russo et al. 2022). It was obtained by a Hilbert transform. To ensure consistency with the calculation of word-level semantic features and to better match the temporal resolution of the hemodynamic signals at the second level, as measured by fNIRS, the envelope values of all time points within each word were averaged. This averaged envelope was considered to reflect the overall loudness or intensity of the speech at the word level.
Semantic features
Surprisal and entropy values were calculated as semantic features of narratives. In particular, surprisal describes how unexpected the current word is, given the previous words and context (Willems et al. 2016, Armeni et al. 2017). By formalizing the sentence as a sequence of words: w1, w2, …, wt, the surprisal for the current word wt is defined by the formula below, where P(wt| w1, …, wt−1) represents the conditional probability to the upcoming word wt with the processing of the previous t − 1 word (w1, …, wt−1).
Entropy measures the degree of uncertainty or unpredictability of the next word, based on the probability distribution of all possible upcoming words (Willems et al. 2016, Armeni et al. 2017). Entropy is defined as a function of the probability distribution of all possible words, denoted by W. The formula for calculating entropy is as follows:
![]() |
Surprisal and entropy are two metrics commonly used in recent literature to describe the semantic features of narratives (Willems et al. 2016, Armeni et al. 2017, Brodbeck et al. 2018, Aurnhammer and Frank 2019). Both metrics serve as measures of word prediction. Next, word prediction is not only an important mechanism for human language processing (Pickering and Garrod 2007, 2013, Hickok et al. 2011, Kutas and Federmeier 2011) but also one of the fundamental NLP tasks for processing semantic information in the texts (Armeni et al. 2017, Goldstein et al. 2022, Heilbron et al. 2022). The probability of P(wt| w1, …, wt−1) in the above formulas can be automatically estimated by NLP models, which are typically trained on large corpora to estimate the co-occurrence of arbitrary word sequences (Armeni et al. 2017). In this study, we used a classical Chinese-based natural language model (Bengio et al. 2003, Kingma and Ba 2014) to compute the surprisal and entropy values for each word in the speaker’s narratives. The model was trained on the People’s Daily corpus, which consisted of 534 246 words for model training, 66 781 words for cross-validation, and 66 781 words for testing.
Correlations between acoustic feature and semantic features of all speakers’ narratives were calculated, as shown in Fig. 1c. We found weak but statistically significant correlations between the envelope and entropy (Pearson’s r = 0.044, P = .0004, r2 = 0.002) or surprisal (Pearson’s r = −0.051, P < .0001, r2 = 0.003). Besides, a negative correlation was observed between entropy and surprisal (Pearson’s r = −0.273, P < .0001, r2 = 0.075). The relationship between these speech features will be further discussed in the Discussion section.
Data analysis
The correlations between the significant speaker–listener neural coupling and the speech features were analyzed. First, for each trial and for each listener participant, the Pearson correlation coefficient r between the temporal dynamics of speaker–listener neural coupling from a single channel combination and a speech feature was calculated, as shown in Fig. 1a. Subsequently, the r-values were Fisher-z transformed. They were then averaged over eight trials for each noise level and further averaged over all listener participants. Next, for each noise level, a one-sample t-test was used to examine the significance of the r-values from all channel combinations. A rmANOVA was also applied to analyze the r-values in four noise levels. A post hoc analysis was further performed to determine which comparison was significant if the result of the rmANOVA was significant. Considering the marked difference in the inter-brain patterns of speaker–listener neural coupling between native and non-native listeners, their analysis was conducted separately and without comparing their r-values.
A permutation test was further used to validate the significance of the r-values for each noise level in native and non-native conditions. As the r-values from all channel combinations showed similar trends in the four noise levels, we first averaged the speaker–listener neural coupling from all the channel combinations and then calculated its correlation with the speech features. The true r-values were obtained using the same procedures as mentioned earlier. We then shuffled the speech features of each trial and recomputed the r-values. A null distribution of r-values was obtained by shuffling 5000 times, which was used to compute the statistical significance (P-value) of the true r-values. Those P-values from multiple computations (two listener groups × three speech features × four noise levels) were further corrected by the FDR method (Benjamini and Hochberg 1995).
Results
Speaker–listener neural coupling during native speech communication correlates with semantic features
The t-test revealed significant correlations between native speaker–listener neural coupling and semantic features, i.e. entropy and surprisal. For entropy, the correlation was significant only in NN [t(23) = −5.35; P < .05, FDR corrected; Cohen’s d = 1.09], but not in the three noisy levels (P > .05, FDR corrected). For surprisal, the correlations were significant in all four noise levels [t(23) = −4.10, −11.63, −9.38, −9.09; P < .05, FDR corrected; Cohen’s d = 0.84, 2.37, 1.91, 1.86]. These significant correlations were all negative, implying higher speaker–listener neural coupling when semantics were more predictable.
The rmANOVA showed that noise level affected the correlation between the speaker–listener neural coupling and semantic features [entropy: rmANOVA F(3, 69) = 5.89, P = .001; surprisal: rmANOVA F(3, 69) = 3.57, P = .018]. For entropy, the correlation was significantly more negative in NN than in the other three noisy levels (post hoc t-test P < .05, FDR corrected). Meanwhile, for surprisal, the correlation was significantly more negative in −9 dB than in NN (post hoc t-test, P < .05, FDR corrected). The above results are shown in Fig. 2a.
Figure 2.

Correlation results between native speaker–listener neural coupling and speech features. (a) Each light line represents the results of one channel combination. (b) Results of the permutation test.
Note: In (a), the bold line represents the averaged correlation of all channel combinations. The asterisks indicate significant correlations (P < .05, FDR corrected). The horizontal lines indicate significant pairwise differences (post hoc t-test, P < .05, FDR corrected). In (b), the asterisks indicate significant correlations based on permutation test (P < .05, FDR corrected).
Table 1 displays the correlation coefficients of the native speaker–listener neural coupling averaged across all 24 channel combinations and speech features. The permutation test confirmed that the correlation for entropy in NN and the correlations for the surprisal in all four noise levels were significant (P < .05, FDR corrected). The remaining results were not significant (P > .05, FDR corrected). The results are shown in Fig. 2b.
Table 1.
The correlations between speaker–listener neural coupling and acoustic feature (envelope) and semantic features (entropy, surprisal) in various conditions.
| Envelope | Entropy | Surprisal | ||
|---|---|---|---|---|
| Native speech | NN | 0.004 (0.006) | −0.018 (0.008)* | −0.019 (0.006)* |
| 2 dB | −0.004 (0.008) | <0.001 (0.005) | −0.020 (0.007)* | |
| −6 dB | −0.003 (0.007) | <0.001 (0.006) | −0.026 (0.008)* | |
| −9 dB | −0.004 (0.006) | −0.002 (0.008) | −0.024 (0.008)* | |
| Non-native speech | NN | 0.004 (0.008) | −0.023 (0.008)* | −0.018 (0.012)* |
| 2 dB | 0.003 (0.011) | −0.014 (0.009)* | −0.022 (0.007)* | |
| −6 dB | −0.002 (0.008) | −0.004 (0.008) | −0.010 (0.008) | |
| −9 dB | 0.014 (0.008)* | −0.010 (0.008) | −0.008 (0.008) |
The standardized deviation of the correlation coefficients was indicated in the brackets. The bold values and asterisks (*) indicate significant correlations that passed the permutation test (FDR corrected).
Speaker–listener neural coupling during native speech communication does not correlate with envelope
As shown in Fig. 2, the correlation between the native speaker–listener neural coupling and envelope was not significant in all four noise levels (P > .05, FDR corrected). The result of the rmANOVA was also not significant [rmANOVA F(3, 69) = 0.29, P = .82].
Non-native speaker–listener neural coupling correlates with semantic features
The t-test revealed significant negative correlations between non-native speaker–listener neural coupling and semantic features, as shown in Fig. 3a. For entropy, the correlation was significant in NN and −2 dB [t(9) = −5.87, −7.38; P < .05, FDR corrected; Cohen’s d = 1.86, 2.34], but not in −6 and −9 dB (P > .05, FDR corrected). For surprisal, the correlations were significant in NN, 2 dB, and −6 dB [t(9) = −5.12, −8.05, −3.22, P < .05, FDR corrected; Cohen’s d = 1.61, 2.55, 1.02], but not in −9 dB (P > .05, FDR corrected).
Figure 3.

Correlation results between non-native speaker–listener neural coupling and speech features. (a) Each light line represents the results of one channel combination. The bold line represents the averaged correlation of all channel combinations. The asterisks indicate significant correlations (P < .05, FDR corrected). The horizontal lines indicate significant pairwise differences (post hoc t-test, P < .05, FDR corrected). For surprisal, the correlations in NN, 2 dB, and −6 dB were significant. For entropy, the correlations in NN and 2 dB were significant. In −9 dB, the correlation between non-native speaker–listener neural coupling and the acoustic envelope was significant, suggesting a higher neural coupling with increasing acoustic amplitude. (b) Results of the permutation test. The neural coupling from all channel combinations was averaged. Significant results are marked by asterisks (P < .05, FDR corrected). Correlations for surprisal and entropy in NN and 2 dB and for envelope in −9 dB were significant.
The rmANOVA showed that the noise level affected the correlation between non-native speaker–listener neural coupling and semantic features [entropy: rmANOVA F(3, 27) = 9.35, P < .001; surprisal: rmANOVA F(3, 27) = 3.71, P = 0.023]. For entropy, the correlations were more negative in NN and 2 dB than in −6 and −9 dB (post hoc t-test, P < .05, FDR corrected). For surprisal, the correlations were more negative in 2 and −6 dB than in −9 dB (post hoc t-test, P < .05, FDR corrected).
The correlations between the non-native speaker–listener neural coupling averaged across all 10 channel combinations and the speech features are shown in Table 1. The results of the permutation test demonstrated that the correlations with entropy and surprisal in NN and 2 dB were significant (P < .05, FDR corrected), while the remaining results were not significant (P > .05, FDR corrected), as shown in Fig. 3b.
Non-native speaker–listener neural coupling correlates with the acoustic envelope in the high noise level
The t-test also revealed a significant correlation between non-native speaker–listener neural coupling and envelope in −9 dB [t(9) = 3.19; P < .05, FDR corrected; Cohen’s d = 1.01]. This correlation was positive, indicating a higher speaker–listener neural coupling when the acoustic amplitude was larger. However, this correlation was not significant in the other three levels (P > .05, FDR corrected). The results of the permutation test confirmed that the neural coupling averaged across the 10 channel combinations was positively correlated with envelope only in −9 dB, as shown in Fig. 3b and Table 1. However, there was no significant difference between the correlations in four noise levels [rmANOVA F(3, 27) = 1.27, P = .30].
Discussion
The present study aims to identify the speech processing underlying the speaker–listener neural coupling during speech communication. We computed the correlation between the temporal dynamics of speaker–listener neural coupling with speech features under various conditions, utilizing two inter-brain datasets accounting for different noise levels and listener’s language experiences (native vs. non-native). Results showed that speaker–listener neural coupling was associated with acoustic (envelope) and semantic (entropy and surprisal) features. Moreover, this relationship was modulated by both noise level and listener’s language experience. Specifically, native speaker–listener neural coupling correlated with surprisal regardless of noise level. Meanwhile, non-native speaker–listener neural coupling correlated with semantic features (both entropy and surprisal) in the absence or presence of low background noise, but only correlated with the acoustic envelope when the noise became strong. Overall, our study revealed that speaker–listener neural coupling is associated with the dynamics of acoustic and semantic features of speech, providing comprehensive insights into the inter-brain neural mechanisms during natural speech communication. We thus advocate for more attention to the time-varying nature of speaker–listener neural coupling and further exploration of the relationship between speaker–listener neural coupling and multilevel speech features.
The present study found that native speaker–listener neural coupling was correlated with semantic features rather than acoustic feature, suggesting that this speaker–listener neural coupling emerged from the semantic processing shared between speakers and listeners. These findings were consistent with our hypotheses and with previous literature. On the one hand, the native speaker–listener neural coupling was from the listener’s left IFG, MTG, and AG. These regions have been widely reported to be involved in semantic processing of speech (Lau et al. 2008, Friederici 2012, Golestani et al. 2013). Following this line, the speaker–listener neural coupling from these regions would naturally reflect the semantic processing, instead of acoustic processing, shared between speakers and listeners. On the other hand, the speaker–listener neural coupling was negatively correlated with semantic features, i.e. surprisal and entropy. A high surprisal value meant that the current word did not match the expectation or prediction; a high entropy value meant a high degree of uncertainty about what the next word would be. The negative correlation thus indicated that speaker–listener neural coupling was higher when the semantics were more predictable or more congruent with the prediction, which was consistent with the findings of two existing inter-brain studies (Dikker et al. 2014, Russo et al. 2022). Furthermore, as both surprisal and entropy estimated the word-level semantics from a predictive coding framework (Willems et al. 2016, Armeni et al. 2017), these findings provide additional support for theories proposing that the listener’s neural coupling to the speaker arises from the inference or prediction to the speaker’s words (Hasson and Frith 2016, Perez and Davis 2023).
More interestingly, the correlation between native speaker–listener neural coupling and surprisal still remained significant and even became stronger in noisy environments, implying a robust semantic processing and its greater contribution to the neural coupling in noisy conditions. These results were also in good agreement with previous evidence that semantic cues or contextual information could improve people’s speech comprehension in noisy environments (Golestani et al. 2013, Bhandari et al. 2021, Rysop et al. 2021). However, while both entropy and surprisal estimated the semantic features of natural speech, this neural coupling was not associated with entropy in noisy conditions. One possible explanation for this discrepancy was that surprisal and entropy focused on different stages of semantic processing. Entropy described the possibilities of upcoming words from a prospective standpoint, whereas surprisal retrospectively integrated the current word into a context-based prediction to estimate the actual outcome of the semantic prediction (Willems et al. 2016, Armeni et al. 2017). Thus, the null results of the correlation between entropy and neural coupling suggest that speech-in-noise comprehension is not dependent on the forward prediction alone, but on integrating both internal prediction and the perception of external stimuli.
Moreover, the present study found that language experience modulated the correlation between speaker–listener neural coupling and speech features in noisy conditions. Specifically, the neural coupling of non-native speaker–listeners exhibited correlations with both surprisal and entropy levels under mild background noise conditions. These results diverged from native listeners, whose neural coupling solely correlated with surprisal. This difference could contribute to the inadequate language experience of non-native listeners. In specific, non-native listeners showed a relative disadvantage in precisely perceiving the actual semantic input as compared to native listeners (Scharenborg and van Os 2019). They might rely more on the prospective prediction to maintain their speech-in-noise comprehension, leading to significant correlation results for entropy. Moreover, these correlations were not significant when the noise became strong, indicating that the relationship between non-native speaker–listener neural coupling and semantic features was more vulnerable to background noise compared to native speaker–listener neural coupling. These results were consistent with previous behavioral observations that non-native speakers benefited little from semantic or contextual information to improve their speech-in-noise performance (Scharenborg and van Os 2019, Borghini and Hazan 2020), suggesting a non-native disadvantage of semantic processing in noisy environments. Meanwhile, non-native speaker–listener neural coupling was found to correlate with acoustic feature at high noise levels, which demonstrated a great reliance on acoustic processing for non-native speech comprehension in noisy environments. It was in accordance with previous neuroimaging evidence showing that non-native comprehension was closely related to neural responses from low-level auditory brain regions (Bidelman and Dexter 2015, Zinszer et al. 2022), as well as behavioral observations that acoustic cues enhanced the non-native speech-in-noise comprehension (Borghini and Hazan 2020). Due to limited language experience, the non-native listeners showed disadvantages in semantic processing (Hervais-Adelman et al. 2014) and might adaptively allocate more cognitive resources to acoustic processing in order to improve their speech-in-noise performance (Song and Iverson 2018, Reetzke et al. 2021). Under this circumstance, the non-native speaker–listener neural coupling would be more closely associated with the low-level acoustic envelope rather than the high-level semantic features. Additionally, we observed a significant yet modest correlation between acoustic and semantic features within the context of the present Chinese speech materials. It aligns with previous studies conducted on different languages, such as English (Brodbeck et al. 2018, Broderick et al. 2019). This correlation may indicate an interplay among various speech features inherent to natural speech (Hamilton and Huth 2020). Nevertheless, given the observed small effect size (r2 < 0.005), it appears unlikely that this acoustic–semantic correlation would significantly influence our modeling between neural coupling and specific speech features. This is particularly the case when the distinct modeling outcomes for acoustic and semantic features were observed across varying noise levels and among diverse populations. Further research is warranted to fully elucidate the intricate relationship between acoustics and semantics.
Above all, the present study shed light on the speech processes underlying speaker–listener neural coupling by exploring the correlation between the neural coupling and speech features. It deepened our knowledge of the inter-brain neural mechanisms involved in natural speech communication. Specifically, while prior studies have emphasized inter-brain neural coupling as a neural foundation for effective speech communication (Stephens et al. 2010, Jiang et al. 2012, Li et al. 2021, 2023), they have not elucidated the speech processing underlying this inter-brain neural coupling. Our study addresses this gap by modeling the temporal dynamics of speaker–listener neural coupling with multilevel speech features. The results provide empirical support for the hypothesis that inter-brain neural coupling is intricately associated with the processing of both acoustic and semantic information (Hasson and Frith 2016). Besides, by comparing the results under various listening conditions and from two groups of listeners with different language experiences, we further highlighted how this association was modulated by the depth of speech processing. It enriches our understanding of how the brain adapts to various adverse environments to maintain speech comprehension. Furthermore, our correlational modeling not only offers a feasible approach for unraveling the computational processing behind the inter-brain neural coupling but also underscores the necessity for more research attention on the time-varying nature and temporal dynamics of the inter-brain neural coupling.
It is important to point out that the speaker–listener neural coupling in this study occurred in a very-low frequency band of 0.01–0.032 Hz. Meanwhile, the dominant rhythm of speech, which is particularly at the syllable level, typically operated ∼4 Hz (Poeppel and Assaneo 2020). The fluctuations of acoustic-phonemic features could be even faster, exceeding 10 Hz. It thus becomes evident that this slow neural coupling might not directly represent the shared encoding of word-level acoustics or semantics. Rather, this very-low frequency band aligned most closely with the temporal duration of high-level speech structures, such as sentences or discourses, which could last for tens of seconds (Baldassano et al. 2017). Therefore, the neural coupling observed in this study might reflect the speech processing at these higher levels, extending beyond the time scale of a single word. An alternative explanation for this very-low frequency band of speaker–listener neural coupling is rooted in the hypothesis that low-frequency oscillations emerge from switching of basic metastable states to maximize the computation and communication of the whole neural system (Palva and Palva 2018). In this line, the speaker–listener neural coupling might indicate a shared computational state for particular speech processing. However, the physiological basis and functional role of very low-frequency neural rhythms remained elusive. Future studies are needed to investigate the relationship between speaker–listener neural coupling from multiple frequency bands and speech features at different temporal scales, which would help to further elucidate the principles, foundations, and functions of interpersonal synchrony (Mayo and Gordon 2020).
There are several ways to further expand upon our findings. First, our study did not distinguish the speaker–listener neural coupling from different brain regions. We found that speaker–listener neural coupling from various regions displayed a similar correlation to speech features, although some functional magnetic resonance imaging (fMRI) studies have discovered that speaker–listener neural coupling from different brain regions might correspond to distinct speech processes (Stephens et al. 2010, Liu et al. 2020). The lack of heterogeneity across different brain regions could be attributed to several factors, including the very low-frequency band of this neural coupling and the limited coverage and spatial resolutions of fNIRS. To overcome these limitations, future studies could employ neuroimaging technologies with a broader coverage or higher spatial resolution, such as magnetoencephalography and fMRI. Secondly, the correlation between the speaker–listener neural coupling and speech features was not particularly strong, suggesting that the speech features explained a relatively low variance of this neural coupling. Howbeit, these correlation values were comparable to the results reported in existing literature that modeled the relationship between speech features and single-brain neural activities measured by scalp electroencephalogram (EEG) (Brodbeck et al. 2018, Broderick et al. 2018) or fMRI (Heilbron et al. 2022). To improve the explanatory power of the correlation relationship, more advanced neural technologies with higher SNRs, such as intracranial EEG (Goldstein et al. 2022) and single neural unit technology, could be implemented. Thirdly, to investigate the specific frequency band demonstrating significant speaker–listener neural coupling, the wavelet parameters used for WTC computation were chosen to provide a relatively good spectral resolution. However, this necessitated involving multiple time points both preceding and following the target words, resulting in a less time-localized estimation. Given the accumulating knowledge about the effective frequency band for neural coupling, future studies could refine parameters or employ alternative computational methods to achieve a more precise estimation of the temporal dynamics of neural coupling. Lastly, as a pioneering exploration, this study simply calculated the correlation coefficient between speaker–listener neural coupling and speech features. More advanced mathematical methods and algorithms, such as encoding and decoding models (Hamilton and Huth 2020), could be introduced to better capture the dynamics, directions, and causality of this relationship.
Supplementary Material
Contributor Information
Zhuoran Li, Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing 100084, China; Tsinghua Laboratory of Brain and Intelligence, Tsinghua University, Beijing 100084, China; Department of Psychiatry, University of Iowa Carver College of Medicine, Iowa City, IA 52242, United States; Stead Family Department of Pediatrics, University of Iowa Carver College of Medicine, Iowa City, IA 52242, United States.
Bo Hong, Tsinghua Laboratory of Brain and Intelligence, Tsinghua University, Beijing 100084, China; Department of Biomedical Engineering, School of Medicine, Tsinghua University, Beijing 100084, China.
Guido Nolte, Department of Neurophysiology and Pathophysiology, University Medical Center Hamburg Eppendorf, Hamburg 20246, Germany.
Andreas K Engel, Department of Neurophysiology and Pathophysiology, University Medical Center Hamburg Eppendorf, Hamburg 20246, Germany.
Dan Zhang, Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing 100084, China; Tsinghua Laboratory of Brain and Intelligence, Tsinghua University, Beijing 100084, China.
Supplementary data
Supplementary data is available at SCAN online.
Conflict of interest
None declared.
Funding
This work was supported by the National Natural Science Foundation of China (NSFC) (T2341003 and 61977041), and the NSFC and the German Research Foundation (DFG) in project Crossmodal Learning (NSFC 61621136008/DFG TRR-169/C1, B1).
Data availability
The data underlying this article will be shared on reasonable request to the corresponding author.
References
- Armeni K, Willems RM, Frank SL. Probabilistic language models in cognitive neuroscience: promises and pitfalls. Neurosci Biobehav Rev 2017;83:579–88. 10.1016/j.neubiorev.2017.09.001 [DOI] [PubMed] [Google Scholar]
- Armeni K, Willems RM, Van den Bosch A et al. Frequency-specific brain dynamics related to prediction during language comprehension. NeuroImage 2019;198:283–95. [DOI] [PubMed] [Google Scholar]
- Aurnhammer C, Frank SL. Evaluating information-theoretic measures of word prediction in naturalistic sentence reading. Neuropsychologia 2019;134:107198. 10.1016/j.neuropsychologia.2019.107198 [DOI] [PubMed] [Google Scholar]
- Baldassano C, Chen J, Zadbood A et al. Discovering event structure in continuous narrative perception and memory. Neuron 2017;95:709–721e705. 10.1016/j.neuron.2017.06.041 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bengio Y, Ducharme R, Vincent P et al. A neural probabilistic language model. J Mach Learn Res 2003;3:1137–55. [Google Scholar]
- Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Series B (Methodological) 1995;57:289–300. [Google Scholar]
- Bhandari P, Demberg V, Kray J. Semantic predictability facilitates comprehension of degraded speech in a graded manner. Front Psychol 2021;12:714485. 10.3389/fpsyg.2021.714485 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bidelman GM, Dexter L. Bilinguals at the “cocktail party”: dissociable neural activity in auditory-linguistic brain regions reveals neurobiological basis for nonnative listeners’ speech-in-noise recognition deficits. Brain Lang 2015;143:32–41. 10.1016/j.bandl.2015.02.002 [DOI] [PubMed] [Google Scholar]
- Borghini G, Hazan V. Effects of acoustic and semantic cues on listening effort during native and non-native speech perception. J Acoust Soc Am 2020;147:3783. 10.1121/10.0001126 [DOI] [PubMed] [Google Scholar]
- Brodbeck C, Hong LE, Simon JZ. Rapid transformation from auditory to linguistic representations of continuous speech. Curr Biol 2018;28:3976–3983e3975. 10.1016/j.cub.2018.10.042 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Broderick MP, Anderson AJ, Di Liberto GM et al. Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narrative speech. Curr Biol 2018;28:803–809e803. 10.1016/j.cub.2018.01.080 [DOI] [PubMed] [Google Scholar]
- Broderick MP, Anderson AJ, Lalor EC. Semantic context enhances the early auditory encoding of natural speech. J Neurosci 2019;39:7564–75. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dikker S, Silbert LJ, Hasson U et al. On the same wavelength: predictable language enhances speaker-listener brain-to-brain synchrony in posterior superior temporal gyrus. J Neurosci 2014;34:6267–72. 10.1523/JNEUROSCI.3796-13.2014 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Etard O, Reichenbach T. Neural speech tracking in the theta and in the delta frequency band differentially encode clarity and comprehension of speech in noise. J Neurosci 2019;39:5750–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Friederici AD. The cortical language circuit: from auditory perception to sentence comprehension. Trends Cognit Sci 2012;16:262–68. 10.1016/j.tics.2012.04.001 [DOI] [PubMed] [Google Scholar]
- Goldstein A, Zada Z, Buchnik E et al. Shared computational principles for language processing in humans and deep language models. Nat Neurosci 2022;25:369–80. 10.1038/s41593-022-01026-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Golestani N, Hervais-Adelman A, Obleser J et al. Semantic versus perceptual interactions in neural processing of speech-in-noise. NeuroImage 2013;79:52–61. 10.1016/j.neuroimage.2013.04.049 [DOI] [PubMed] [Google Scholar]
- Grand G, Blank IA, Pereira F et al. Semantic projection recovers rich human knowledge of multiple object features from word embeddings. Nat Hum Behav 2022;6:975–987. 10.1038/s41562-022-01316-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Grinsted A, Moore JC, Jevrejeva S. Application of the cross wavelet transform and wavelet coherence to geophysical time series. Nonlinear Proc Geoph 2004;11:561–66. 10.5194/npg-11-561-2004 [DOI] [Google Scholar]
- Hamilton LS, Huth AG. The revolution will not be controlled: natural stimuli in speech neuroscience. Lang Cogn Neurosci 2020;35:573–82. 10.1080/23273798.2018.1499946 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hasson U, Frith CD. Mirroring and beyond: coupled dynamics as a generalized framework for modelling social interactions. Philos Trans R Soc Lond B Biol Sci 2016;371:20150366. 10.1098/rstb.2015.0366 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hasson U, Nir Y, Levy I et al. Intersubject synchronization of cortical activity during natural vision. Science 2004;303:1634–40. 10.1126/science.1089506 [DOI] [PubMed] [Google Scholar]
- Heilbron M, Armeni K, Schoffelen JM et al. A hierarchy of linguistic predictions during natural language comprehension. Proc Natl Acad Sci USA 2022;119:e2201968119. 10.1073/pnas.2201968119 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hervais-Adelman A, Pefkou M, Golestani N. Bilingual speech-in-noise: neural bases of semantic context use in the native language. Brain Lang 2014;132:1–6. 10.1016/j.bandl.2014.01.009 [DOI] [PubMed] [Google Scholar]
- Hickok G, Houde J, Rong F. Sensorimotor integration in speech processing: computational basis and neural organization. Neuron 2011;69:407–22. 10.1016/j.neuron.2011.01.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jiang J, Chen C, Dai B et al. Leader emergence through interpersonal neural synchronization. Proc Natl Acad Sci USA 2015;112:4274–79. 10.1073/pnas.1422930112 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jiang J, Dai B, Peng D et al. Neural synchronization during face-to-face communication. J Neurosci 2012;32:16064–69. 10.1523/JNEUROSCI.2926-12.2012 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jiang J, Zheng L, Lu C. A hierarchical model for interpersonal verbal communication. Soc Cogn Affect Neurosci 2021;16:246–55. 10.1093/scan/nsaa151 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kelsen BA, Sumich A, Kasabov N et al. What has social neuroscience learned from hyperscanning studies of spoken communication? A systematic review. Neurosci Biobehav Rev 2022;132:1249–62. 10.1016/j.neubiorev.2020.09.008 [DOI] [PubMed] [Google Scholar]
- King-Casas B, Tomlin D, Anen C et al. Getting to know you: reputation and trust in a two-person economic exchange. Science 2005;308:78–83. 10.1126/science.1108062 [DOI] [PubMed] [Google Scholar]
- Kingma D, Ba J. Adam: a method for stochastic optimization. arXiv:1412.6980. 2014. https://arxiv.org/abs/1412.6980
- Kutas M, Federmeier KD. Thirty years and counting: finding meaning in the N400 component of the event-related brain potential (ERP). Annu Rev Psychol 2011;62:621–47. 10.1146/annurev.psych.093008.131123 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lau EF, Phillips C, Poeppel D. A cortical network for semantics: (de)constructing the N400. Nat Rev Neurosci 2008;9:920–33. 10.1038/nrn2532 [DOI] [PubMed] [Google Scholar]
- Li Z, Hong B, Wang D et al. Speaker-listener neural coupling reveals a right-lateralized mechanism for non-native speech-in-noise comprehension. Cereb Cortex 2023;33:3701–14. 10.1093/cercor/bhac302 [DOI] [PubMed] [Google Scholar]
- Li Z, Li J, Hong B et al. Speaker-listener neural coupling reveals an adaptive mechanism for speech comprehension in a noisy environment. Cereb Cortex 2021;31:4719–29. 10.1093/cercor/bhab118 [DOI] [PubMed] [Google Scholar]
- Liu L, Ding X, Li H et al. Reduced listener–speaker neural coupling underlies speech understanding difficulty in older adults. Brain Struct Funct 2021;226:1571–84. 10.1007/s00429-021-02271-2 [DOI] [PubMed] [Google Scholar]
- Liu L, Zhang Y, Zhou Q et al. Auditory-articulatory neural alignment between listener and speaker during verbal communication. Cereb Cortex 2020;30:942–51. 10.1093/cercor/bhz138 [DOI] [PubMed] [Google Scholar]
- Liu W, Branigan HP, Zheng L et al. Shared neural representations of syntax during online dyadic communication. NeuroImage 2019;198:63–72. 10.1016/j.neuroimage.2019.05.035 [DOI] [PubMed] [Google Scholar]
- Maris E, Oostenveld R. Nonparametric statistical testing of EEG- and MEG-data. J Neurosci Methods 2007;164:177–90. 10.1016/j.jneumeth.2007.03.024 [DOI] [PubMed] [Google Scholar]
- Mayo O, Gordon I. In and out of synchrony-behavioral and physiological dynamics of dyadic interpersonal coordination. Psychophysiology 2020;57:e13574. 10.1111/psyp.13574 [DOI] [PubMed] [Google Scholar]
- Palva S, Palva JM. Roles of brain criticality and multiscale oscillations in temporal predictions for sensorimotor processing. Trends Neurosci 2018;41:729–43. 10.1016/j.tins.2018.08.008 [DOI] [PubMed] [Google Scholar]
- Pan Y, Novembre G, Song B et al. Interpersonal synchronization of inferior frontal cortices tracks social interactive learning of a song. NeuroImage 2018;183:280–90. 10.1016/j.neuroimage.2018.08.005 [DOI] [PubMed] [Google Scholar]
- Perez A, Davis MH. Speaking and listening to inter-brain relationships. Cortex 2023;159:54–63. 10.1016/j.cortex.2022.12.002 [DOI] [PubMed] [Google Scholar]
- Pickering MJ, Garrod S. Do people use language production to make predictions during comprehension? Trends Cognit Sci 2007;11:105–10. 10.1016/j.tics.2006.12.002 [DOI] [PubMed] [Google Scholar]
- Pickering MJ, Garrod S. An integrated theory of language production and comprehension. Behav Brain Sci 2013;36:329–47. 10.1017/S0140525X12001495 [DOI] [PubMed] [Google Scholar]
- Poeppel D, Assaneo MF. Speech rhythms and their neural foundations. Nat Rev Neurosci 2020;21:322–34. 10.1038/s41583-020-0304-4 [DOI] [PubMed] [Google Scholar]
- Redcay E, Schilbach L. Using second-person neuroscience to elucidate the mechanisms of social interaction. Nat Rev Neurosci 2019;20:495–505. 10.1038/s41583-019-0179-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reetzke R, Gnanateja GN, Chandrasekaran B. Neural tracking of the speech envelope is differentially modulated by attention and language experience. Brain Lang 2021;213:104891. 10.1016/j.bandl.2020.104891 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Russo AG, De Martino M, Elia A et al. Negative correlation between word-level surprisal and intersubject neural synchronization during narrative listening. Cortex 2022;155:132–49. 10.1016/j.cortex.2022.07.005 [DOI] [PubMed] [Google Scholar]
- Rysop AU, Schmitt LM, Obleser J et al. Neural modelling of the semantic predictability gain under challenging listening conditions. Hum Brain Mapp 2021;42:110–27. 10.1002/hbm.25208 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sassenhagen J, Draschkow D. Cluster‐based permutation tests of MEG/EEG data do not establish significance of effect latency or location. Psychophysiology 2019;56:e13335. 10.1111/psyp.13335 [DOI] [PubMed] [Google Scholar]
- Scharenborg O, van Os M. Why listening in background noise is harder in a non-native language than in a native language: a review. Speech Commun 2019;108:53–64. 10.1016/j.specom.2019.03.001 [DOI] [Google Scholar]
- Scholkmann F, Spichtig S, Muehlemann T et al. How to detect and reduce movement artifacts in near-infrared imaging using moving standard deviation and spline interpolation. Physiol Meas 2010;31:649. 10.1088/0967-3334/31/5/004 [DOI] [PubMed] [Google Scholar]
- Schoot L, Hagoort P, Segaert K. What can we learn from a two-brain approach to verbal interaction? Neurosci Biobehav Rev 2016;68:454–59. 10.1016/j.neubiorev.2016.06.009 [DOI] [PubMed] [Google Scholar]
- Song J, Iverson P. Listening effort during speech perception enhances auditory and lexical processing for non-native listeners and accents. Cognition 2018;179:163–70. 10.1016/j.cognition.2018.06.001 [DOI] [PubMed] [Google Scholar]
- Stephens GJ, Silbert LJ, Hasson U. Speaker-listener neural coupling underlies successful communication. Proc Natl Acad Sci USA 2010;107:14425–30. 10.1073/pnas.1008662107 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wass SV, Whitehorn M, Haresign IM et al. Interpersonal neural entrainment during early social interaction. Trends Cognit Sci 2020;24:329–42. 10.1016/j.tics.2020.01.006 [DOI] [PubMed] [Google Scholar]
- Willems RM, Frank SL, Nijhof AD et al. Prediction during natural language comprehension. Cereb Cortex 2016;26:2506–16. 10.1093/cercor/bhv075 [DOI] [PubMed] [Google Scholar]
- Yeshurun Y, Nguyen M, Hasson U. The default mode network: where the idiosyncratic self meets the shared social world. Nat Rev Neurosci 2021;22:181–92. 10.1038/s41583-020-00420-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yücel MA, Selb J, Cooper RJ et al. Targeted principle component analysis: a new motion artifact correction approach for near-infrared spectroscopy. J Innov Opt Health Sci 2014;7:1350066. 10.1142/S1793545813500661 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zinszer BD, Yuan QM, Zhang ZQ et al. Continuous speech tracking in bilinguals reflects adaptation to both language and noise. Brain Lang 2022;230:105128. 10.1016/j.bl.2022.105128 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data underlying this article will be shared on reasonable request to the corresponding author.

