Abstract
Visual speech, such as lipreading, facilitates spoken word recognition, but the neural mechanisms underlying audiovisual speech perception remain poorly understood. Visual cues may disambiguate fine-grained articulatory features during early perceptual stages or instead integrate with speech at more categorical, phoneme-level stages. To test how speech representations are modulated by visual input, we analyzed intracranial electroencephalography (iEEG) signals recorded from 12 epilepsy patients performing an audiovisual speech perception task. Participants perceived 16 monosyllabic words presented in auditory-only, visual-only, or congruent audiovisual formats. Words were constructed from four onset consonants (/b/, /g/, /m/, /n/) and four rimes (vowel nucleus and any coda consonants). We examined event-related potentials (ERP) in superior temporal gyrus (STG) and trained support vector machine (SVM) classifiers to decode word identity from neural activity at individual electrodes. Discrete and continuous confusion matrices captured complementary changes in classification accuracy and normalized inverse classification loss, a continuous proxy for classifier confidence. Decoding performance was hierarchically evaluated at the word, phoneme, and phonetic feature levels to determine the representations affected by visual speech. Congruent audiovisual speech increased classifier confidence for phoneme-level representations and improved decoding accuracy at both the word and phoneme level, without corresponding effects on phonetic features. Time-resolved analyses further revealed earlier successful decoding for audiovisual than auditory-only speech, with audiovisual enhancement primarily observed for onset consonants rather than rimes. Together, these findings suggest that visual speech sharpens primarily categorical phoneme representations in STG, with accelerated speech processing and improved word recognition emerging as downstream consequences of phoneme-level enhancement.
Keywords: Multisensory integration, speech perception, iEEG, neural decoding
Introduction
Speech perception is a special case of multisensory processing (Sumby & Pollack, 1954). Although the normal-hearing population relies predominantly on audition, incongruent visual information reshapes the auditory phoneme percept (Alsius et al., 2018; McGurk & MacDonald, 1976), and congruent cues enhance speech comprehension in optimal and challenging environments (Ross et al., 2007; Sumby & Pollack, 1954). Despite extensive behavioral evidence, how visual speech alters auditory cortical processing remains unclear.
Visual speech modulates auditory processing across temporal and spectral scales (Cao et al., 2024; Karthik et al., 2021; Mégevand et al., 2020), and silent lipreading alone activates auditory cortex (Pekkola et al., 2005). Recent intracranial electroencephalography (iEEG) and fMRI results suggest that beyond the modulatory effects on auditory processing, visemes—the categorical units of speech-related lip movements (Fisher, 1968)—may be encoded in auditory cortex (Audenhaege et al., 2025; Karthik et al., 2024). Visual information is thus relayed into auditory cortex, yet its interaction with linguistic representations is underexplored.
Linguistic representations (DeWitt & Rauschecker, 2013; Gwilliams, Bhaya-Grossman, et al., 2025; Hickok & Poeppel, 2004, 2007; Holt & Lotto, 2010) and their organization in superior temporal gyrus (STG) (Bhaya-Grossman et al., 2025; de Heer et al., 2017) recapitulate the hierarchical structure of speech. Speech sounds are described along composable articulatory dimensions, which define phonetic features such as place of articulation (POA, where the vocal tract constricts airflow) and manner of articulation (MOA, how airflow is constricted). Phonemes are language-specific minimal units distinguishing word meanings. In STG, phonetic features are represented by local spectrotemporal tunings (Leonard et al., 2024; Mesgarani et al., 2014), whereas phonemes are categorically organized at the population level (Chang et al., 2010). To achieve stable yet flexible speech perception despite variability (Kleinschmidt & Jaeger, 2015), listeners integrate available cues to resolve ambiguous phonemes or restore phonetic features (Cole, 1973; Gwilliams et al., 2018; Leonard et al., 2016). Computational models further support efficient multisensory cue combination (Chandrasekaran, 2017; Körding et al., 2007; Magnotti & Beauchamp, 2017), raising the question of where in the hierarchy congruent visual cues are integrated.
While all levels of the psycholinguistic hierarchy are susceptible to visual influence (Bernstein & Liebenthal, 2014; Campbell, 2008), visual speech is often thought to enrich phonetic features, the acoustic consequences of articulator states (Grant & Walden, 1996; Summerfield, 1979). Accordingly, facial cues bias perceived POA even when acoustics are identical (Green & Kuhl, 1989), and McGurk-style paradigms show systematic shifts in phonetic encoding (Brancazio et al., 2003; Shahin et al., 2018). Consistent with this feature-based account, neurophysiologically, visual speech produces articulator-specific facilitation of early auditory processing proportional to its phonetic predictiveness (van Wassenhove et al., 2005).
Alternatively, with meaning-rich stimuli, the goal of audiovisual speech is intelligibility (Sumby & Pollack, 1954), i.e., improving phoneme and word perception rather than refining features. Visual speech alone supports word identification (Bernstein et al., 2000), and intracranial evidence shows stronger audiovisual benefits for mouth-leading than voice-leading words (Karas et al., 2019). Visual speech may therefore directly modulate phoneme-level processing by suppressing incompatible alternatives to restrict lexical candidates (Luce, 1986; Tye-Murray et al., 2007). However, as phonetic features underlie phoneme category membership, the two levels are intertwined, and phoneme judgments remain sensitive to acoustic similarity (Samuel, 1981), making it challenging to isolate the representational level at which visual input shapes word perception.
Here we test how visual speech interacts with auditory representations at phonetic-feature, phoneme, and word levels during word identification. We recorded iEEG from STG in epilepsy patients hearing and viewing monosyllabic words and decoded word identity using support vector machines. Decoding was evaluated hierarchically with confusion and confidence matrices, capturing discrete changes in classifier decisions and continuous representational gains, respectively. Congruent visual speech increased phoneme- and word-level decoding accuracy and phoneme-level classifier confidence, without comparable benefits for isolated phonetic features. Visual speech thus appears to enhance the inter-category separability of phoneme representations in human STG, with word-level gains emerging downstream from enhanced phoneme separability.
Methods
Experimental model and participant details
Participants
All experimental procedures were approved by the Institutional Review Boards (IRBs) at the University of Michigan (UM) and Henry Ford (HF) hospitals. IEEG data were recorded from 13 patients undergoing clinical evaluations for intractable epilepsy. One patient was excluded from the analysis due to a lack of STG coverage, resulting in a final sample of n = 12 patients (5 female; mean age = 30.0, std = 9.36 years; 8 right-handed, 2 left-handed, and 2 with no clear hand preference). All patients were native speakers of American English and provided written informed consent to participate in the research. Electrode type and implant location were decided solely based on clinical needs. All 12 analyzed patients had sEEG depth electrodes. 11 patients were implanted only with depth electrodes, with 5-mm center-to-center spacing for UM patients and 3.5-mm spacing for HF patients. One patient additionally had a subdural ECoG grid with 10 mm center-to-center spacing implanted above the left STG. Two patients had bilateral STG coverage, whereas the remaining 10 had left STG coverage only. For visualization purposes, all electrodes were projected onto the left hemisphere in MNI space (Fig. 1C). No anatomical abnormalities affecting STG regions were identified.
Figure 1: Task Schematic and iEEG Data Overview.

(A) Trial structure. Participants heard auditory-only (A) or audiovisual (AV) word lists and identified the initial consonant of the final word. In AV trials, visual speech began 250 ms before auditory onset. (B) Hierarchical decoding model. The X-axis shows labels predicted by the SVM classifiers, and the Y-axis shows the true labels. Diagonal cells indicate correct word classification; off-diagonal shading denotes errors that nevertheless preserved phoneme- or phoneticfeature-level information or were incorrect at all levels. (C) Electrode coverage projected onto the left hemisphere in MNI space. Blue electrodes were included in the analysis, gray electrodes were excluded, and red shading highlights superior temporal gyrus (STG) grey and white matter. Electrodes from patients with bilateral coverage were projected onto the left hemisphere for visualization. (D) Event-related potentials (ERPs) from two representative left STG electrodes for words beginning with /b/, /g/, /m/, or /n/. Responses show phonetic-feature organization, with similar responses for nasal (/m/, /n/) and plosive (/b/, /g/) consonants. Blue boxes indicate the 0–250 ms window used for SVM classification; anatomical insets show electrode locations.
Method details
Experimental design
Participants were tested at their bedside in the Epilepsy Monitoring Units using a laptop running Psychtoolbox (Brainard, 1997; Kleiner et al., 2007). The speech stimuli consisted of word lists with varying lengths, ranging from 4 to 16 words to prevent participants from attending to only the end of the list. Each list was generated by randomly selecting words from a pool of 16 monosyllabic words (for experimental setup, see Fig. 1A; for the word set, see Table S1). These words were constructed from combinations of four onset consonants (/b/, /g/, /m/, /n/) and four rimes (vowel nucleus and coda consonants; /ey-t/, /uw-n/, /ae-sh/, /ih-l/). The four onset consonants were evenly divided between plosive (/b/, /g/), and nasal (/m/, /n/) manners of articulation.
Stimuli were video recorded from a female speaker (frame rate: 59.94 fps). For each word, a single exemplar utterance was selected for presentation in three conditions: unimodal auditory (A; speech linearly mixed with pink noise, with speech and noise waveform amplitudes scaled at a 70:30 ratio, presented with a static grey screen), unimodal visual (V; video of the speaker’s face accompanied by auditory pink noise), and multimodal audiovisual (AV; synchronized auditory and visual speech, with visual onset preceding auditory onset by 250 ms). In the A and AV conditions, time 0 was defined as auditory onset, ensuring a consistent temporal reference for auditory processing across conditions. In the V condition, time 0 was instead defined as visual onset to better align neural activity associated with visual speech processing. Word durations ranged 576–1054 ms, with onset consonants ranging 70–196 ms and rimes ranging 434–937 ms. Audio and video files were trimmed or padded to be 750 ms. Within each word list, successive words were separated by a 300-ms inter-word interval (ISI) measured from the auditory offset of one word to the auditory onset of the next. During the ISI, the final frame of the visual stimulus remained on screen until the visual onset of the subsequent word.
Participants were instructed to identify the initial consonant of the final word in each word list by button press with their dominant hand (four-alternative forced choice). Response time limits were adaptive, initially set to 2.5 s and adjusted in 250 ms steps based on performance, with bounds of 1.5 and 4 s. For each patient, trials from all three conditions were fully randomized. As testing duration varied across patients due to clinical and fatigue-related constraints, the number of repetitions differed across participants and conditions. Each word was presented at least eight times per condition (8–35 repetitions in A, and 8–32 repetitions in AV). 11 of the 12 patients completed all three conditions. One patient completed an earlier version of the task that only contained A and AV conditions.
Data acquisition and preprocessing
IEEG data were acquired at 4096 Hz with the Natus Quantum Amplifier (7 patients; University of Michigan Hospital, resampled to 1024 Hz) or 1000 Hz with the Nihon Kohden Amplifier (5 patients; Henry Ford Hospital). Surface reconstruction, electrode registration, and projection to Montreal Neurological Institute (MNI) coordinates were performed by aligning each patient’s preoperative T1 MRI to their postoperative CT using FreeSurfer and custom MATLAB scripts (Brang et al., 2016).
Drift was removed from each electrode by fitting and subtracting a third-order polynomial, followed by high-pass filtering at 0.1 Hz. Notch filtering at 60 Hz was applied to remove power line noise. No additional low-pass or band-pass filtering was applied, so the full broadband signal from 0.1 Hz to the Nyquist frequency (UM: 512 Hz; HF: 500 Hz) was retained for ERP computation. High-gamma power (HGp; 70–150 Hz) was extracted using wavelet decomposition for complementary analyses reported in the Supplementary Materials. Excessively noisy electrodes were identified and excluded if their signal variance deviated by more than 5 standard deviations in either direction from the mean across all electrodes. The remaining channels were manually inspected and channels with residual signal artifacts were removed. To minimize the influence of volume conduction on localized neural responses, bipolar re-referencing was applied. IEEG signals from adjacent electrodes were subtracted and assigned to a linearly interpolated spatial coordinate. Bipolar electrodes were included in analyses only if their nearest FreeSurfer volumetric anatomical label (projected onto the individual patient’s MRI) corresponded to STG (Desikan et al., 2006) and the signal from the electrode exhibited speech responsiveness. Speech responsiveness was defined as both a significant ERP response within 250 ms of sound onset (i.e., at least one non-zero ERP time point; p_adj < 0.05, FDR-corrected) and significant decoding performance when unimodal auditory and multimodal audiovisual trials were pooled (p < 0.001; for details, see Support Vector Machine Decoding). All electrode screening steps were completed in patients’ individual anatomical space, and electrode coordinates were projected into MNI space for group-level visualization only. For the A and AV analyses, 158 electrodes met the anatomical criterion, of which 87 also passed the functional selection and were retained for subsequent analyses. For the V analyses, 142 electrodes met the anatomical criterion, with 81 retained after functional selection. The number of analyzed bipolar electrodes per patient ranged from 1 to 15 contacts. For A and AV analyses, patients had a mean of 7.25 analyzed contacts (std = 4.63; n = 12); for V analyses, patients had a mean of 6.75 analyzed contacts (std = 5.08; n = 11). Across patients, the final A and AV electrode set comprised 82 depth and 5 subdural electrodes, whereas the V electrode set comprised 76 depth and 5 subdural electrodes (see Fig. 1C).
Quantification and statistical analysis
ERP pattern inspection
Although high-gamma power (HGp) is a commonly used index of local cortical processing in iEEG research, audiovisual speech effects in STG are often expressed in lower-frequency neural activity. In particular, visual speech has been shown to provide lateral or feedback inputs into auditory cortex through phase-resetting and cross-modal temporal alignment of low frequency oscillations (Bastos et al., 2015; Luo et al., 2010; Mégevand et al., 2020; Schroeder et al., 2008), and recent iEEG studies have shown that pre-articulatory visual speech cues modulate low-frequency activity in STG (Karthik et al., 2021, 2024). ERPs were therefore selected as the primary measure because they are particularly sensitive to these low-frequency audiovisual processes. Analyses using HGp were performed for complementary inquiry (see Supplementary Materials).
The bipolar-referenced time series were segmented into 1.5-second epochs around the onset of the word ([−500, 1000 ms], with 0 marking auditory onset of the word). The epoch window was set to capture both preparatory speech-related movements preceding auditory onset and sustained post-onset neural responses during word processing. Epochs were averaged to compute four ERPs within each condition, corresponding to the four initial consonants of the word, and tested for significance against the pre-stimulus baseline period ([−300, −250 ms]). The baseline window was selected to occur after the auditory offset of the previous word and before the visual onset of articulatory movements in the AV condition, minimizing contamination from anticipatory speech-related activity. Fig. 1D shows ERP groupings in the unimodal auditory condition from 2 sample electrodes. All four ERP waveforms peaked at a similar latency, approximately 200 ms after word onset. Notably, the ERPs elicited by the two nasals (/m/ and /n/) and the two plosives (/b/ and /g/) closely tracked each other within their respective groups in both timing and waveform morphology.
Support Vector Machine decoding
Electrode-level word decoding
Support Vector Machine (SVM) classifiers were trained to decode the 16 words from single-trial ERP patterns for each patient. For each trial, the feature vector consisted of ERP amplitudes sampled at each time point between 0 and 250 ms after word onset from a single electrode. The decoding window was selected based on a previous implementation of the same task paradigm using a different set of monosyllabic words, which showed robust speech-related ERP responses within this time range (Karthik et al., 2024). This resulted in N-dimensional feature vectors per trial, where N corresponded to the number of sampled time points within the decoding window (N = 256 for UM patients and N = 250 for HF patients). Classification was performed at the electrode level using an 8-fold multi-class linear SVM classifier (MATLAB function ‘fitcecoc’). Multiclass classification was implemented using the one-vs-one approach, where a separate binary SVM classifier was trained to distinguish each pair of classes. With 16 classes, this approach yields a total of 120 binary classifiers. Cross-validation was stratified so that each fold contained a balanced proportion of each word label.
Classification accuracy was computed as the k-fold mean proportion of correctly predicted labels across all held-out trials. Accuracy was first calculated for each electrode and then averaged across electrodes within each patient for group-level analysis. With 16 classes per condition, the theoretical chance level was 6.25%. For each electrode, statistical significance of classification accuracy was assessed using a one-tailed binomial test (MATLAB function ‘binocdf’), comparing observed accuracy to chance. The minimum accuracy required for statistical significance depended on the number of trials and was determined individually for each electrode.
In addition to overall classification accuracy, we further evaluated decoding performance with confusion matrices and classifier confidence. Classifier confidence was quantified using the negated average binary classification loss (‘NegLoss’) returned by the MATLAB function ‘kfoldPredict’, where higher values indicated smaller error magnitude and therefore higher classifier confidence. Trial-wise NegLoss values were converted by row-wise softmax into probability-like distributions across the 16 classes, removing any additive offset shared across classes. For each electrode, these per-trial distributions were then averaged across all trials of each word to form a 16-by-16 continuous error-magnitude matrix, paralleling the structure of the discrete confusion matrix. Both confidence matrices and confusion matrices were generated for each electrode and then averaged within patients for group-level analysis.
Electrode-level rime decoding
To test whether audiovisual enhancement extended to rime information beyond onset consonants, additional SVM classifiers were trained to decode rime labels from ERP patterns. The decoding procedure was identical to that used for word-label decoding, except that the analysis window was time-locked to the offset of the onset consonants, with window length defined as the median rime length across all 16 words (702 ms). Phoneme timing was estimated using the Montreal Forced Aligner (McAuliffe et al., 2017), then manually verified and corrected. We first performed a 16-class rime decoding analysis in which classifiers were trained to predict the rime label associated with each of the 16 words. To further dissociate rime representations from whole-word identity, an additional 4-class SVM analysis was conducted in which separate classifiers were trained to discriminate among the four rimes regardless of the preceding consonants. Statistical significance for rime decoding performance was assessed using one tailed binomial tests following the same procedure used for word-label decoding, except that theoretical chance performance was 6.25% for the 16-class classifiers and 25% for the 4-class classifiers.
Patient-level time-resolved decoding
To examine the temporal dynamics of decoding performance, time-resolved SVM analyses were additionally performed from −500 to 1000 ms relative to auditory onset. For these analyses, iEEG data were resampled to 500 Hz, yielding decoding estimates every 2 ms. For each patient and time point, the feature vector consisted of ERP amplitudes across all analyzed electrodes. Classifiers were trained and tested independently at each time point, producing patient-level time-resolved decoding accuracy curves for each condition. Decoding accuracies were then smoothed within each patient using a 100 ms (i.e., 50 time points) moving average implemented with the MATLAB function ‘movmean’. Statistical significance against chance performance at each time point was assessed using one-sample t-tests across patients and corrected for multiple comparisons using false discovery rate (FDR) correction. Successful decoding onset was defined as the first run of at least 20 ms in which group accuracy exceeded empirical chance (one-sided t-test, FDR-corrected across conditions and time points, p < 0.05). To compare successful decoding latencies between AV and A conditions, we used a jackknifebased procedure (Kiesel et al., 2008; Miller et al., 1998), estimating onsets from leave-one-out averages across n patients and evaluating the AV–A latency difference with a jackknife-corrected paired t-test against n − 1 degrees of freedom.
Hierarchical model for information processing
To answer the question of which levels of linguistic representation are susceptible to visual speech, we designed a hierarchical model to interpret patterns in the group-level confusion matrix and confidence matrix (see Fig. 1B for model structure). Rather than training separate classifiers for word-, phoneme-, and feature-level representations, we used a single word-identification model and examined the structure of its classification errors. This approach mirrors the behavioral task completed by participants and enables characterization of how different linguistic representations contribute to word-level decisions, thus providing a closer analogue to human speech perception. The diagonal elements of each matrix represent trials in which the word label was decoded correctly (word-level accuracy), while off-diagonal cells were categorized into three types to capture different levels of decoding accuracy: (1) phoneme-level accuracy, where the initial consonant was correctly decoded but the remainder of the word label was incorrect; (2) feature-level accuracy, where the initial consonant was misclassified as another consonant sharing the same manner of articulation (e.g., a word starting with ‘m’ classified as one starting with ‘n’); and (3) other errors, where neither the initial consonant nor its manner of articulation matched the true label.
This exclusive hierarchical structure enabled dissociation of decoding performance at the word, phoneme, and feature levels. Because the word set crossed four onset consonants with four rimes, each representational level was estimated from a different number of confusion-matrix cells: 16 cells contributed to the word level (the matrix diagonal), 48 to the phoneme level (correct onset, incorrect rime), and 64 to the feature level (correct shared manner of articulation, incorrect onset and rime). To determine whether audiovisual enhancement differed statistically across representational levels, audiovisual accuracy gains (AV − A) were first analyzed using a linear mixed-effects model with representational level as a fixed effect and subject as a random effect. Follow-up Wilcoxon signed-rank tests were then conducted separately at each of the four representational levels for both the confusion and confidence matrices to characterize the effect of visual speech on decoding accuracy and classifier error-magnitude within each level.
Results
Since only the final word of each word list required a behavioral response, behavioral data were available for 1,022 of 10,285 iEEG epochs (9.9%). Participants remained engaged throughout the task, responding on the majority of trials across conditions (response rates: A: M = 93.9%, std = 6.0%; AV: M = 94.9%, std = 4.8%; V: M = 82.7%, std = 23.7%). Identification accuracy was high when auditory speech was present (A: M = 87.6%, std = 13.4%; AV: M = 90.5%, std = 15.2%) but substantially lower in the visual-only condition (M = 42.0%, std = 15.4%). Performance in the unimodal visual condition remained significantly above chance (0.25) (Wilcoxon signed-rank, W(12) = 75, p = 0.001, r = 0.923), indicating that participants were able to extract some speech information from visual input alone. Although AV accuracy was numerically higher than A accuracy, this difference did not reach statistical significance (W(13) = 53, p = 0.151, r = 0.359).
Speech-responsive electrodes in STG were selected for group-level analysis using an orthogonal channel selection procedure. Meaningful channels were identified as those with overall decoding accuracy significantly above chance, calculated across both unimodal auditory (A) and audiovisual (AV) trials combined (p < 0.001). Subsequent statistical comparisons of decoding accuracy between AV and A conditions were performed exclusively on this pre-selected set of channels (See Fig. 1C for the distribution of all selected channels).
We first tested whether decoding accuracy for word identity was higher in the AV condition compared to the A condition (Fig. 2A). Overall, electrode-level SVM decoding accuracy was 0.169 (std = 0.069) for AV, 0.158 (std = 0.068) for A, and 0.068 (std = 0.009) for V. Wilcoxon signed rank test between AV and A revealed that AV accuracy was significantly higher than A (W(12) = 67, p = 0.013, r = 0.634).
Figure 2: Visual speech enhances word- and phoneme-level decoding accuracy.

(A) SVM decoding accuracy for 16-word classification across 12 patients in the auditory-only (A), audiovisual (AV), and visual-only (V) conditions. Left, mean accuracy in each condition; each colored dot represents one patient. Right, AV – A accuracy difference. Dashed lines indicate chance (0.0625) and zero difference, respectively. AV showed significantly higher decoding accuracy than A. (B) Group-averaged confusion matrices for A, AV, and V. The V condition is shown both rescaled to match the A/AV range and on its original scale. Axes show true versus predicted labels. Black 4×4 boxes mark correct initial-consonant (phoneme-level) decoding; colored boxes mark correct phonetic-feature decoding (purple, nasal; green, plosive). (C) AV – A accuracy differences at each hierarchical level (see Fig. 1B). Statistically significant differences between AV and A were observed for both word- and phoneme-level decoding.
To further determine which level of information was affected by congruent visual speech input, we used confusion matrices to decompose classifier performance in addition to examining correct responses. Even when word identity was misclassified, the predicted word could still share the target’s phoneme or phonetic features, allowing us to quantify information retained at each level (Fig. 2B–C). Audiovisual gain differed significantly across representational levels, according to the linear mixed-effects model with representational level as a fixed effect and subject as a random effect, F(3, 44) = 4.420, p = 0.008, R2marginal = 0.216, R2conditional = 0.216. Wilcoxon signedrank tests were then conducted at each of the four levels. A significant enhancement in the AV condition was observed at the word level (A: 0.158, AV: 0.169; W(12) = 67, p = 0.013, r = 0.634) and the phoneme level (A: 0.075, AV: 0.080; W(12) = 64, p = 0.026, r = 0.566). No significant differences were found at the feature level (A: 0.068, AV: 0.067; W(12) = 35, p = 0.633, r = −0.091) or for other types of error (A: 0.043, AV: 0.040; W(12) = 12, p = 0.987, r = −0.611). Notably, the absence of an audiovisual effect at the feature level cannot be attributed to limited measurement precision at that level. The phonetic category was estimated from more confusion-matrix cells than either the phoneme or word level (64 versus 48 and 16 cells, respectively), so any genuine audiovisual modulation of phonetic-feature representations would have been estimated with comparable or greater statistical precision than the effects detected at the word and phoneme levels.
Classification accuracy provides a thresholded readout of representational changes, registering an effect only when sharpening is sufficient to convert a misclassification into a correct one. More gradual improvements in representational separability may therefore remain undetected if neural representations move toward the correct category but fail to cross a decision boundary and change the final classification outcome. To test for such sub-threshold effects, we next examined classifier confidence, quantified by the distance from the decision boundaries (Fig. 3). As with the confusion matrix analysis, the linear mixed-effect model revealed a significant effect of representational level on audiovisual gain, F(3, 44) = 3.286, p = 0.029, R2marginal = 0.170, R2conditional = 0.170, and follow-up Wilcoxon signed-rank tests were conducted for each of the four levels. In this case, a significant enhancement in the AV condition was observed only at the phoneme level (A: 0.066, AV: 0.067; W(12) = 71, p = 0.005, r = 0.725). No significant differences were found at the word level (A: 0.074, AV: 0.075; W(12) = 60, p = 0.055, r = 0.476), feature level (A: 0.066, AV: 0.065; W(12) = 38, p = 0.545, r = −0.023), or for other error types (A: 0.058, AV: 0.043; W(12) = 22, p = 0.912, r = −0.385). For completeness and comparison to auditory iEEG work, decoding and hierarchical analyses were also performed using HGp (70–150 Hz) with the same pipeline. HGp-based analyses showed a qualitatively similar but weaker overall pattern compared with ERP-based analyses (Fig. S1).
Figure 3: Visual speech enhances phoneme-level confidence.

(A) Group-averaged confidence matrices for unimodal auditory (A), audiovisual (AV), and unimodal visual (V) conditions. Classifier confidence was derived from normalized inverse classification loss, with higher values indicating lower classification error and a better match between the neural activity pattern and a given class. (B) Statistical comparison of hierarchical model structure. The figure setup is identical to that in Fig. 2B–C, except that the values plotted here reflect classifier confidence rather than classification accuracy. Classifier confidence was quantified as being inversely related to the distance from the decision boundaries (See Methods). Statistically significant differences between AV and A were only observed for phoneme-level decoding.
Given that audiovisual enhancement was observed for word- and phoneme-level decoding accuracy, as well as phoneme-level confidence, but not for phonetic-feature representations, we hypothesized that strengthened categorical phoneme representations may contribute to the observed improvements in word-level classification. To test this hypothesis, we fit a linear mixed-effects model predicting word decoding gain from increased phoneme confidence while accounting for variability across patients. There was a non-significant positive relationship between increased phoneme confidence and word decoding gain (b = 2.579, SE = 2.14, t(85) = 1.205, p = 0.232, 95% CI [−1.677, 6.834]). Although this analysis did not provide direct statistical support, the positive coefficient was consistent with the proposed link between clarified phoneme-level representations and improved word-level decoding.
To further characterize the temporal dynamics of audiovisual enhancement, and to test whether AV benefits extended beyond onset consonants to facilitate decoding of subsequent rime information, patient-level time-resolved decoding analyses were performed at 2 ms intervals (Fig. 4A). Above-chance word-level decoding emerged earlier in the AV condition than in the A condition, starting at −2 ms and 32 ms relative to auditory onset, respectively. A jackknife-based latency comparison confirmed the earlier onset for AV than A decoding, mean AV–A difference = −35 ms, t(11) = −2.124, p = 0.029. Decoding remained elevated throughout much of the postauditory-onset interval. Peak decoding accuracy was also higher in the AV (0.168) condition relative to the A (0.164), W(12) = 67, p = 0.013, r = 0.634. In contrast, decoding accuracy in V emerged later (118 ms) and remained comparatively modest throughout the trial (peak accuracy = 0.076).
Figure 4: Visual speech accelerates word decoding with limited rime enhancement.

(A) Group-averaged time-resolved SVM decoding accuracy for word labels in the auditory-only (A), audiovisual (AV), and visual-only (V) conditions across 12 patients, plotted relative to auditory onset (0 ms). Additional dashed lines indicate visual onset of the current word (−250 ms) and auditory offset of the preceding word (−300ms, grey shaded region). Colored shaded regions represent ±1 SEM across participants; colored horizontal bars indicate time points significantly above chance, p < 0.05, FDR-corrected). AV showed earlier decoding onset and higher peak accuracy than A, whereas V decoding emerged later and remained weaker. (B) Electrode-level SVM decoding accuracy for rime labels in the A, AV, and V conditions across 11 patients. One patient was excluded because no bipolar contacts met the speech-responsive criterion (significant decoding for A and AV combined, p < .001). The figure format is the same as in Fig. 2A. AV rime decoding was numerically higher than A but did not differ significantly.
Given that time-resolved decoding accuracy remained significantly above chance throughout the trial in both the A and AV conditions, we next examined whether audiovisual enhancement was also present for rime-level information (Fig. 4B). Electrode-level 16-class SVM decoding accuracy for rimes was numerically higher in AV (0.266) than that for A (0.263) at the group level, although this effect did not reach statistical significance (W(11) = 51, p = 0.062, r = 0.483). To determine whether this trend toward an AV advantage reflected facilitated rime processing or instead was driven by improved whole-word decoding, we conducted an additional analysis in which classifiers discriminated among the 4 rimes irrespective of the preceding consonants. Averaged patient-level time-resolved decoding performance was significantly above chance in both the A and AV conditions (Fig. 5A; peak decoding accuracy in A: 0.407, AV: 0.410, V: 0.283; significance latency relative to auditory onset: A = 76 ms, AV = 78 ms, V = 472 ms). However, no evidence for an audiovisual advantage was observed in decoding onset latency (jackknife-based comparison: mean AV-A difference = 44 ms, t(11) = 0.134, p = 0.552), peak decoding accuracy (W(12) = 44, p = 0.367, r = 0.113), or in the electrode-level decoding accuracy (Fig. 5B; W(11) = 33, p = 0.517, r = 0.0). Together, these findings suggest that the AV-related improvement observed in the 16-class rime analysis was more likely attributable to enhanced whole-word representations than to rime information itself.
Figure 5: Rime enhancement reflects word-level audiovisual advantage.

(A) Group-averaged time-resolved 4-class SVM decoding accuracy for rime identity in the auditory-only (A), audiovisual (AV), and visual-only (V) conditions. The figure format is the same as in Fig. 4A. (B) Electrode-level 4-class SVM decoding accuracy for rime identity across the 11 included patients in the A, AV, and V conditions. The figure format is the same as in Fig. 4B, except that decoding was performed across the four rimes irrespective of onset consonant. Although decoding in A and AV remained above chance, no significant audiovisual advantage was observed in either the time-resolved or electrode-level analysis.
Discussion
Congruent visual input facilitates speech perception, but this behavioral enhancement can arise through multiple mechanisms, including changes in temporal processing (Cao et al., 2024; McGrath & Summerfield, 1985; Mégevand et al., 2020), attention allocation (ten Oever et al., 2014; Zion Golumbic et al., 2012; Zion Golumbic et al., 2013), cue weighting (Shahin et al., 2018), speech parsing (Chandrasekaran et al., 2009), and potential integration of spectral (Bröhl et al., 2022; Plass et al., 2020) and visemic information (Karthik et al., 2021) derived from visual cues. In addition to vision-driven modulation of auditory responses during speech perception (Okada et al., 2013), recent evidence demonstrates that visemic identities are directly represented in the auditory cortex in a manner analogous to phonemes (Audenhaege et al., 2025; Karthik et al., 2024), suggesting a direct interaction with linguistic representations in the auditory cortex. However, it remains unclear whether such audiovisual integration alters speech representations encoded by neuronal populations in STG, and if so, at which level of linguistic representation these effects emerge. In the present study, we showed enhanced neural decoding accuracy for word identity in audiovisual relative to unimodal auditory conditions, which is consistent with classic findings from the speech perception literature and extend prior findings to neural representations of audiovisual speech. Importantly, our results suggest that the improvement in word-level decoding accuracy may arise from enhanced confidence in phoneme-level representations.
Although the increase in phoneme-level classifier confidence and the improvement in word-level decoding accuracy were observed in different metrics and at different levels of the hierarchical model, they likely reflect a common underlying change in representational structure. Classification accuracy is a thresholded measure that captures discrete representational changes only when they alter classification outcomes, whereas classifier confidence is sensitive to more subtle continuous changes in representational separability. Visual speech may therefore first augment phoneme representations, increasing confidence before affecting decoding accuracy. As phoneme representations become increasingly distinct, these changes eventually improve both phoneme- and word-level classification. Since the words in our word set were primarily distinguished by onset phoneme identity, increased phoneme separability would be expected to propagate to word-level decoding. Consistent with this interpretation, the positive coefficient in the LME analysis suggests that the phoneme-confidence and word-decoding effects may reflect successive stages of the same representational process.
The absence of any audiovisual effect at the phonetic-feature level further constrains the nature of this representational change. Importantly, this finding does not suggest that phonetic features are weakly represented in STG. In fact, the confusion and confidence matrices showed clear phonetic-feature grouping in both unimodal auditory and audiovisual conditions, consistent with prior work showing that phonetic features are among the most salient dimensions encoded in human STG (Leonard et al., 2024; Mesgarani et al., 2014). Thus, the dominant representational structure in the present task remained organized by phonetic features, as expected for an auditory-dominant speech task.
However, our central question was which level of representation is modified by congruent visual speech. From this perspective, the lack of such a feature-level effect supports a genuine representational dissociation: congruent visual speech selectively increased the discriminability of phoneme representations without corresponding enhancement of the already robust phonetic feature organization. This supports the view that audiovisual facilitation operates on categorical, lexically relevant units rather than on the acoustic-articulatory features from which those categories are formed (Shahin et al., 2018). Intracranial recordings from pSTG directly support this notion, showing that visual speech suppresses auditory responses most for mouth-leading words, an effect attributed to inhibition of incompatible phoneme representations (Karas et al., 2019). Reduced overall responses and sharper category separation are compatible outcomes of the same competitive process. This interpretation aligns well with prior research using masking (Shahin et al., 2012) and speech degradation (Rennig & Beauchamp, 2022), suggesting that visual speech cues facilitate the recovery of higher-level linguistic information and improve intelligibility of speech in noise. Likewise, silent lip-reading studies show that visual speech processing extracts visemes, the visual analogues of phonemes, dissociably from lip movements themselves (Nidiffer et al., 2023). Taken together, these findings support the view that visual information contributes to the identification of abstract, meaning-related units or provides complementary cues to linguistic content during word processing, rather than simply recovering missing acoustic or articulatory details.
The temporal dynamics of whole-word decoding reported here showed earlier successful decoding in the audiovisual condition relative to the unimodal auditory condition. This finding supports a predictive role for visual speech. As articulatory movements typically precede the voice, visual cues constrain predictions for upcoming phonetic content, shortening auditory response latencies in proportion to visual predictiveness (van Wassenhove et al., 2005) via a fast route conveying visual predictions directly to auditory cortex (Arnal et al., 2009). The temporal window for audiovisual integration is wide and asymmetric (van Wassenhove et al., 2007), aligning with the multi-scale integration windows proposed for speech processing (Poeppel et al., 2008). Collectively, these mechanisms allow visual input to sharpen auditory representations before the acoustic signal is fully available, consistent with the earlier word-level decoding observed here. Notably, patient-level time-resolved decoding analyses showed robust rime decoding in both audiovisual and unimodal auditory conditions, indicating that sustained speech information remained decodable throughout the trial. However, significant group-level audiovisual enhancement in the electrode-level decoding analysis was observed for onset consonants, but not for rimes. These findings are consistent with previous work showing that visual speech primarily benefits the initial stages of auditory speech processing, whereas sustained speech processing relies more heavily on tracking the acoustic envelope (Cao et al., 2024). This distinction is further supported by converging intracranial evidence. Words with a visual head start show the largest visual benefit (Karas et al., 2019), and human STG distinguishes onset from sustained responses (Hamilton et al., 2018).
Additionally, we found that 16-class rime decoding, which preserved coarticulatory information between onset consonants and rimes, showed a greater numerical audiovisual advantage than the 4-class rime decoding analysis. This pattern suggests that visual speech may be particularly informative for distinguishing transitional acoustic-phonetic patterns rather than isolated phoneme categories, consistent with the inherently coarticulated nature of continuous speech (Schwartz & Savariaux, 2014; Venezia et al., 2016). Within the current experiment, the stronger audiovisual effect observed in the 16-class analysis further suggests that the apparent rime advantage was largely driven by enhanced whole-word decoding rather than improved decoding of isolated rime segments. More broadly, these findings are consistent with renewed interest in the role of coarticulation in natural speech processing (Kent & Minifie, 1977; Moreira et al., 2025).
While our decoding results suggest that congruent visual information leads to better word identity decoding through more confident phoneme categorization, the underlying neural computations responsible for this enhancement remain unclear. Given that visual information is already present in the auditory cortex before sound onset (Karthik et al., 2024), future studies could adopt an encoding perspective to directly investigate the content and timing of congruent visual speech representation. Additionally, examining the spectral profile of these effects across frequency bands may help clarify the neural mechanism by which visemic information is transformed into phoneme representations (Karthik et al., 2021; Schroeder et al., 2008).
The present study focused exclusively on STG. Future work incorporating regions implicated in audiovisual integration (e.g., pSTS: Audenhaege et al., 2025; Erickson et al., 2014; Karthik et al., 2024; Rennig & Beauchamp, 2022; Zhu & Beauchamp, 2017) and word retrieval (e.g., IFG: Yu et al., 2025) would provide a more comprehensive picture of information flow during multisensory speech perception. Specifically, the absence of an audiovisual effect at the phoneticfeature level in STG does not preclude such effects elsewhere in the speech network. We would predict stronger audiovisual enhancement of phonetic-feature representations in the pSTS, where visual articulatory information may be integrated before being converted into phoneme representations in STG (Zhang et al., 2025). Within STG, a sharp functional boundary separates anterior and posterior subregions in multisensory speech processing (Ozker et al., 2017, 2018), so audiovisual effects may also vary along this axis. Additional limitations include restricted phoneme coverage and imbalanced word frequencies, both necessary for the controlled word set. These constraints, however, do not compromise our central inquiry into how visual information influences speech perception, as identical stimuli were presented across conditions. A broader characterization of phoneme and viseme confusion patterns in English (Cutler et al., 2004; Fisher, 1968) will be necessary to draw more generalizable conclusions about how visemes and phonemes interact to support successful speech perception. Furthermore, extending the current analysis from word identification to longer temporal scales could yield a deeper understanding of the multisensory hierarchy underlying natural speech perception (Gwilliams, Marantz, et al., 2025; Heilbron et al., 2022; O’Sullivan et al., 2021).
In sum, the present study provides new evidence that congruent visual information enhances phoneme processing in STG, with increased confidence in phoneme identity leading to more accurate word perception. These findings are consistent with the hypothesis that visual information primarily supports the extraction of higher-level linguistic representations or provides supplementary cues for speech comprehension, rather than simply enhancing acoustic or articulatory information, encouraging future work toward a multisensory framework for speech perception.
Supplementary Material
Significance Statement.
Seeing a speaker’s face improves speech perception, especially in noise, but the neural representations altered by visual speech remain unclear. Using intracranial recordings from human superior temporal gyrus, we show that congruent visual speech does not simply amplify all speech-related information. Instead, visual input selectively enhances the separability of confusable phoneme-level representations and improves word identity decoding, while leaving the robust phonetic-feature organization largely intact. These findings clarify the distinction between what visual speech modulates and the representational structure of speech in auditory cortex, suggesting that audiovisual facilitation primarily acts on categorical speech representations that more directly support word recognition.
Acknowledgments:
This study was supported by NIH grants R01DC020717, and R01NS094399. We sincerely thank the patients for generously contributing their time and effort to this research.
Footnotes
Conflict of interest statement: The authors declare no competing financial interests.
Data/code availability:
Code will be publicly available upon publication of this article. Data are available upon reasonable request.
References
- Alsius A., Paré M., & Munhall K. G. (2018). Forty Years After Hearing Lips and Seeing Voices: The McGurk Effect Revisited. Multisensory Research, 31(1–2), 111–144. 10.1163/22134808-00002565 [DOI] [PubMed] [Google Scholar]
- Arnal L. H., Morillon B., Kell C. A., & Giraud A.-L. (2009). Dual Neural Routing of Visual Facilitation in Speech Processing. Journal of Neuroscience, 29(43), 13445–13453. 10.1523/JNEUROSCI.3194-09.2009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Audenhaege A. V., Mattioni S., Cerpelloni F., Gau R., Szmalec A., & Collignon O. (2025). Phonological Representations of Auditory and Visual Speech in the Occipito-temporal Cortex and Beyond. Journal of Neuroscience, 45(26). 10.1523/JNEUROSCI.1415-24.2025 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bastos A. M., Vezoli J., Bosman C. A., Schoffelen J.-M., Oostenveld R., Dowdall J. R., De Weerd P., Kennedy H., & Fries P. (2015). Visual areas exert feedforward and feedback influences through distinct frequency channels. Neuron, 85(2), 390–401. 10.1016/j.neuron.2014.12.018 [DOI] [PubMed] [Google Scholar]
- Bernstein L. E., & Liebenthal E. (2014). Neural pathways for visual speech perception. Frontiers in Neuroscience, 8. 10.3389/fnins.2014.00386 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bernstein L. E., Tucker P. E., & Demorest M. E. (2000). Speech perception without hearing. Perception & Psychophysics, 62(2), 233–252. 10.3758/BF03205546 [DOI] [PubMed] [Google Scholar]
- Bhaya-Grossman I., Leonard M. K., Zhang Y., Gwilliams L., Johnson K., Lu J., & Chang E. F. (2025). Shared and language-specific phonological processing in the human temporal lobe. Nature, 1–12. 10.1038/s41586-025-09748-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Brainard D. H. (1997). The Psychophysics Toolbox. Spatial Vision, 10(4), 433–436. [PubMed] [Google Scholar]
- Brancazio L., Miller J. L., & Paré M. A. (2003). Visual influences on the internal structure of phonetic categories. Perception & Psychophysics, 65(4), 591–601. 10.3758/BF03194585 [DOI] [PubMed] [Google Scholar]
- Brang D., Dai Z., Zheng W., & Towle V. L. (2016). Registering imaged ECoG electrodes to human cortex: A geometry-based technique. Journal of Neuroscience Methods, 273, 64–73. 10.1016/j.jneumeth.2016.08.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bröhl F., Keitel A., & Kayser C. (2022). MEG Activity in Visual and Auditory Cortices Represents Acoustic Speech-Related Information during Silent Lip Reading. eNeuro, 9(3). 10.1523/ENEURO.0209-22.2022 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Campbell R. (2008). The processing of audio-visual speech: Empirical and neural bases. Philosophical Transactions of the Royal Society B: Biological Sciences, 363(1493), 1001–1010. 10.1098/rstb.2007.2155 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cao C. Z., Stacey W. C., Wasade V. S., Towle V. L., Tao J. X., Wu S., Issa N. P., & Brang D. (2024). Visual speech enhances auditory onset timing and envelope tracking through distinct mechanisms (p. 2024.11.23.624953). bioRxiv. 10.1101/2024.11.23.624953 [DOI] [Google Scholar]
- Chandrasekaran C. (2017). Computational principles and models of multisensory integration. Current Opinion in Neurobiology, Neurobiology of Learning and Plasticity, 43, 25–34. 10.1016/j.conb.2016.11.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chandrasekaran C., Trubanova A., Stillittano S., Caplier A., & Ghazanfar A. A. (2009). The Natural Statistics of Audiovisual Speech. PLOS Computational Biology, 5(7), e1000436. 10.1371/journal.pcbi.1000436 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chang E. F., Rieger J. W., Johnson K., Berger M. S., Barbaro N. M., & Knight R. T. (2010). Categorical speech representation in human superior temporal gyrus. Nature Neuroscience, 13(11), 1428–1432. 10.1038/nn.2641 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cole R. A. (1973). Listening for mispronunciations: A measure of what we hear during speech. Perception & Psychophysics, 13(1), 153–156. 10.3758/BF03207252 [DOI] [Google Scholar]
- Cutler A., Weber A., Smits R., & Cooper N. (2004). Patterns of English phoneme confusions by native and non-native listeners. The Journal of the Acoustical Society of America, 116(6), 3668–3678. 10.1121/1.1810292 [DOI] [PubMed] [Google Scholar]
- de Heer W. A., Huth A. G., Griffiths T. L., Gallant J. L., & Theunissen F. E. (2017). The Hierarchical Cortical Organization of Human Speech Processing. The Journal of Neuroscience, 37(27), 6539–6557. 10.1523/JNEUROSCI.3267-16.2017 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Desikan R. S., Ségonne F., Fischl B., Quinn B. T., Dickerson B. C., Blacker D., Buckner R. L., Dale A. M., Maguire R. P., Hyman B. T., Albert M. S., & Killiany R. J. (2006). An automated labeling system for subdividing the human cerebral cortex on MRI scans into gyral based regions of interest. NeuroImage, 31(3), 968–980. 10.1016/j.neuroimage.2006.01.021 [DOI] [PubMed] [Google Scholar]
- DeWitt I., & Rauschecker J. P. (2013). Wernicke’s area revisited: Parallel streams and word processing. Brain and Language, 127(2), 181–191. 10.1016/j.bandl.2013.09.014 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Erickson L. C., Zielinski B. A., Zielinski J. E. V., Liu G., Turkeltaub P. E., Leaver A. M., & Rauschecker J. P. (2014). Distinct cortical locations for integration of audiovisual speech and the McGurk effect. Frontiers in Psychology, 5. 10.3389/fpsyg.2014.00534 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fisher C. G. (1968). Confusions Among Visually Perceived Consonants. Journal of Speech and Hearing Research, 11(4), 796–804. 10.1044/jshr.1104.796 [DOI] [PubMed] [Google Scholar]
- Grant K. W., & Walden B. E. (1996). Evaluating the articulation index for auditory–visual consonant recognition. The Journal of the Acoustical Society of America, 100(4), 2415–2424. 10.1121/1.417950 [DOI] [PubMed] [Google Scholar]
- Green K. P., & Kuhl P. K. (1989). The role of visual information in the processing of place and manner features in speech perception. Perception & Psychophysics, 45(1), 34–42. 10.3758/bf03208030 [DOI] [PubMed] [Google Scholar]
- Gwilliams L., Bhaya-Grossman I., Zhang Y., Scott T., Harper S., & Levy D. (2025). Computational Architecture of Speech Comprehension in the Human Brain. Annual Review of Linguistics, 11(1), 209–226. 10.1146/annurev-linguistics-031120-111245 [DOI] [Google Scholar]
- Gwilliams L., Linzen T., Poeppel D., & Marantz A. (2018). In Spoken Word Recognition, the Future Predicts the Past. The Journal of Neuroscience, 38(35), 7585–7599. 10.1523/JNEUROSCI.0065-18.2018 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gwilliams L., Marantz A., Poeppel D., & King J.-R. (2025). Hierarchical dynamic coding coordinates speech comprehension in the human brain. bioRxiv: The Preprint Server for Biology, 2024.04.19.590280. 10.1101/2024.04.19.590280 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hamilton L. S., Edwards E., & Chang E. F. (2018). A Spatial Map of Onset and Sustained Responses to Speech in the Human Superior Temporal Gyrus. Current Biology, 28(12), 1860–1871.e4. 10.1016/j.cub.2018.04.033 [DOI] [PubMed] [Google Scholar]
- Heilbron M., Armeni K., Schoffelen J.-M., Hagoort P., & de Lange F. P. (2022). A hierarchy of linguistic predictions during natural language comprehension. Proceedings of the National Academy of Sciences, 119(32), e2201968119. 10.1073/pnas.2201968119 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hickok G., & Poeppel D. (2004). Dorsal and ventral streams: A framework for understanding aspects of the functional anatomy of language. Cognition, Towards a New Functional Anatomy of Language, 92(1), 67–99. 10.1016/j.cognition.2003.10.011 [DOI] [PubMed] [Google Scholar]
- Hickok G., & Poeppel D. (2007). The cortical organization of speech processing. Nature Reviews Neuroscience, 8(5), 393–402. 10.1038/nrn2113 [DOI] [PubMed] [Google Scholar]
- Holt L. L., & Lotto A. J. (2010). Speech perception as categorization. Attention, Perception & Psychophysics, 72(5), 1218–1227. 10.3758/APP.72.5.1218 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karas P. J., Magnotti J. F., Metzger B. A., Zhu L. L., Smith K. B., Yoshor D., & Beauchamp M. S. (2019). The visual speech head start improves perception and reduces superior temporal cortex responses to auditory speech. eLife, 8, e48116. 10.7554/eLife.48116 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karthik G., Cao C. Z., Demidenko M. I., Jahn A., Stacey W. C., Wasade V. S., & Brang D. (2024). Auditory cortex encodes lipreading information through spatially distributed activity. Current Biology, 34(17), 4021–4032.e5. 10.1016/j.cub.2024.07.073 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karthik G., Plass J., Beltz A. M., Liu Z., Grabowecky M., Suzuki S., Stacey W. C., Wasade V. S., Towle V. L., Tao J. X., Wu S., Issa N. P., & Brang D. (2021). Visual speech differentially modulates beta, theta, and high gamma bands in auditory cortex. European Journal of Neuroscience, 54(9), 7301–7317. 10.1111/ejn.15482 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kent R. D., & Minifie F. D. (1977). Coarticulation in recent speech production models. Journal of Phonetics, 5(2), 115–133. 10.1016/S0095-4470(19)31123-4 [DOI] [Google Scholar]
- Kiesel A., Miller J., Jolicœur P., & Brisson B. (2008). Measurement of ERP latency differences: A comparison of single-participant and jackknife-based scoring methods. Psychophysiology, 45(2), 250–274. 10.1111/j.1469-8986.2007.00618.x [DOI] [PubMed] [Google Scholar]
- Kleiner M., Brainard D. H., Pelli D., Ingling A., Murray R., & Broussard C. (2007). What’s new in Psychtoolbox-3. Perception, 36, 1–16. 10.1068/v070821 [DOI] [Google Scholar]
- Kleinschmidt D. F., & Jaeger T. F. (2015). Robust speech perception: Recognize the familiar, generalize to the similar, and adapt to the novel. Psychological Review, 122(2), 148–203. 10.1037/a0038695 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Körding K. P., Beierholm U., Ma W. J., Quartz S., Tenenbaum J. B., & Shams L. (2007). Causal Inference in Multisensory Perception. PLOS ONE, 2(9), e943. 10.1371/journal.pone.0000943 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leonard M. K., Baud M. O., Sjerps M. J., & Chang E. F. (2016). Perceptual restoration of masked speech in human cortex. Nature Communications, 7(1), 13619. 10.1038/ncomms13619 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leonard M. K., Gwilliams L., Sellers K. K., Chung J. E., Xu D., Mischler G., Mesgarani N., Welkenhuysen M., Dutta B., & Chang E. F. (2024). Large-scale single-neuron speech sound encoding across the depth of human cortex. Nature, 626(7999), 593–602. 10.1038/s41586-023-06839-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Luce P. A. (1986). Neighborhoods of Words in the Mental Lexicon. Research on Speech Perception. Technical Report No. 6. https://eric.ed.gov/?id=ED353610 [Google Scholar]
- Luo H., Liu Z., & Poeppel D. (2010). Auditory Cortex Tracks Both Auditory and Visual Stimulus Dynamics Using Low-Frequency Neuronal Phase Modulation. PLOS Biology, 8(8), e1000445. 10.1371/journal.pbio.1000445 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Magnotti J. F., & Beauchamp M. S. (2017). A Causal Inference Model Explains Perception of the McGurk Effect and Other Incongruent Audiovisual Speech. PLOS Computational Biology, 13(2), e1005229. 10.1371/journal.pcbi.1005229 [DOI] [PMC free article] [PubMed] [Google Scholar]
- McAuliffe M., Socolof M., Mihuc S., Wagner M., & Sonderegger M. (2017). Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi. Proc. Interspeech 2017, 498–502. 10.21437/Interspeech.2017-1386 [DOI] [Google Scholar]
- McGrath M., & Summerfield Q. (1985). Intermodal timing relations and audio-visual speech recognition by normal-hearing adults. The Journal of the Acoustical Society of America, 77(2), 678–685. 10.1121/1.392336 [DOI] [PubMed] [Google Scholar]
- McGurk H., & MacDonald J. (1976). Hearing lips and seeing voices. Nature, 264(5588), 746–748. 10.1038/264746a0 [DOI] [PubMed] [Google Scholar]
- Mégevand P., Mercier M. R., Groppe D. M., Golumbic E. Z., Mesgarani N., Beauchamp M. S., Schroeder C. E., & Mehta A. D. (2020). Crossmodal Phase Reset and Evoked Responses Provide Complementary Mechanisms for the Influence of Visual Speech in Auditory Cortex. Journal of Neuroscience, 40(44), 8530–8542. 10.1523/JNEUROSCI.0555-20.2020 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mesgarani N., Cheung C., Johnson K., & Chang E. F. (2014). Phonetic Feature Encoding in Human Superior Temporal Gyrus. Science, 343(6174), 1006–1010. 10.1126/science.1245994 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Miller J., Patterson T., & Ulrich R. (1998). Jackknife-based method for measuring LRP onset latency differences. Psychophysiology, 35(1), 99–115. 10.1111/1469-8986.3510099 [DOI] [PubMed] [Google Scholar]
- Moreira J. P. C., Carvalho V. R., Mendes E. M. A. M., Fallah A., Sejnowski T. J., Lainscsek C., & Comstock L. (2025). An open-access EEG dataset for speech decoding: Exploring the role of articulation and coarticulation. Scientific Data, 12(1), 1017. 10.1038/s41597-025-05187-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nidiffer A. R., Cao C. Z., O’Sullivan A., & Lalor E. C. (2023). A representation of abstract linguistic categories in the visual system underlies successful lipreading. NeuroImage, 282, 120391. 10.1016/j.neuroimage.2023.120391 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Okada K., Venezia J. H., Matchin W., Saberi K., & Hickok G. (2013). An fMRI Study of Audiovisual Speech Perception Reveals Multisensory Interactions in Auditory Cortex. PLOS ONE, 8(6), e68959. 10.1371/journal.pone.0068959 [DOI] [PMC free article] [PubMed] [Google Scholar]
- O’Sullivan A. E., Crosse M. J., Liberto G. M. D., Cheveigné A. de, & Lalor E. C. (2021). Neurophysiological Indices of Audiovisual Speech Processing Reveal a Hierarchy of Multisensory Integration Effects. Journal of Neuroscience, 41(23), 4991–5003. 10.1523/JNEUROSCI.0906-20.2021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ozker M., Schepers I. M., Magnotti J. F., Yoshor D., & Beauchamp M. S. (2017). A Double Dissociation between Anterior and Posterior Superior Temporal Gyrus for Processing Audiovisual Speech Demonstrated by Electrocorticography. Journal of Cognitive Neuroscience, 29(6), 1044–1060. 10.1162/jocn_a_01110 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ozker M., Yoshor D., & Beauchamp M. S. (2018). Converging Evidence From Electrocorticography and BOLD fMRI for a Sharp Functional Boundary in Superior Temporal Gyrus Related to Multisensory Speech Processing. Frontiers in Human Neuroscience, 12. 10.3389/fnhum.2018.00141 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pekkola J., Ojanen V., Autti T., Jääskeläinen I. P., Möttönen R., Tarkiainen A., & Sams M. (2005). Primary auditory cortex activation by visual speech: An fMRI study at 3 T. NeuroReport, 16(2), 125. 10.1097/00001756-200502080-00010 [DOI] [PubMed] [Google Scholar]
- Plass J., Brang D., Suzuki S., & Grabowecky M. (2020). Vision perceptually restores auditory spectral dynamics in speech. Proceedings of the National Academy of Sciences, 117(29), 16920–16927. 10.1073/pnas.2002887117 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Poeppel D., Idsardi W. J., & van Wassenhove V. (2008). Speech perception at the interface of neurobiology and linguistics. Philosophical Transactions of the Royal Society B: Biological Sciences, 363(1493), 1071–1086. 10.1098/rstb.2007.2160 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rennig J., & Beauchamp M. S. (2022). Intelligibility of audiovisual sentences drives multivoxel response patterns in human superior temporal cortex. NeuroImage, 247, 118796. 10.1016/j.neuroimage.2021.118796 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ross L. A., Saint-Amour D., Leavitt V. M., Javitt D. C., & Foxe J. J. (2007). Do You See What I Am Saying? Exploring Visual Enhancement of Speech Comprehension in Noisy Environments. Cerebral Cortex, 17(5), 1147–1153. 10.1093/cercor/bhl024 [DOI] [PubMed] [Google Scholar]
- Samuel A. G. (1981). The role of bottom-up confirmation in the phonemic restoration illusion. Journal of Experimental Psychology: Human Perception and Performance, 7(5), 1124–1131. 10.1037/0096-1523.7.5.1124 [DOI] [PubMed] [Google Scholar]
- Schroeder C. E., Lakatos P., Kajikawa Y., Partan S., & Puce A. (2008). Neuronal oscillations and visual amplification of speech. Trends in Cognitive Sciences, 12(3), 106–113. 10.1016/j.tics.2008.01.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schwartz J.-L., & Savariaux C. (2014). No, There Is No 150 ms Lead of Visual Speech on Auditory Speech, but a Range of Audiovisual Asynchronies Varying from Small Audio Lead to Large Audio Lag. PLOS Computational Biology, 10(7), e1003743. 10.1371/journal.pcbi.1003743 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shahin A. J., Backer K. C., Rosenblum L. D., & Kerlin J. R. (2018). Neural Mechanisms Underlying Cross-Modal Phonetic Encoding. Journal of Neuroscience, 38(7), 1835–1849. 10.1523/JNEUROSCI.1566-17.2017 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shahin A. J., Kerlin J. R., Bhat J., & Miller L. M. (2012). Neural restoration of degraded audiovisual speech. NeuroImage, 60(1), 530–538. 10.1016/j.neuroimage.2011.11.097 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sumby W. H., & Pollack I. (1954). Visual Contribution to Speech Intelligibility in Noise. The Journal of the Acoustical Society of America, 26(2), 212–215. 10.1121/1.1907309 [DOI] [Google Scholar]
- Summerfield Q. (1979). Use of Visual Information for Phonetic Perception. Phonetica. 10.1159/000259969 [DOI] [PubMed] [Google Scholar]
- ten Oever S., Schroeder C. E., Poeppel D., van Atteveldt N., & Zion-Golumbic E. (2014). Rhythmicity and cross-modal temporal cues facilitate detection. Neuropsychologia, 63, 43–50. 10.1016/j.neuropsychologia.2014.08.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tye-Murray N., Sommers M., & Spehar B. (2007). Auditory and Visual Lexical Neighborhoods in Audiovisual Speech Perception. Trends in Amplification, 11(4), 233–241. 10.1177/1084713807307409 [DOI] [PMC free article] [PubMed] [Google Scholar]
- van Wassenhove V., Grant K. W., & Poeppel D. (2005). Visual speech speeds up the neural processing of auditory speech. Proceedings of the National Academy of Sciences, 102(4), 1181–1186. 10.1073/pnas.0408949102 [DOI] [PMC free article] [PubMed] [Google Scholar]
- van Wassenhove V., Grant K. W., & Poeppel D. (2007). Temporal window of integration in auditory-visual speech perception. Neuropsychologia, Advances in Multisensory Processes, 45(3), 598–607. 10.1016/j.neuropsychologia.2006.01.001 [DOI] [PubMed] [Google Scholar]
- Venezia J. H., Thurman S. M., Matchin W., George S. E., & Hickok G. (2016). Timing in audiovisual speech perception: A mini review and new psychophysical data. Attention, Perception, & Psychophysics, 78(2), 583–601. 10.3758/s13414-015-1026y [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yu L., Dugan P., Doyle W., Devinsky O., Friedman D., & Flinker A. (2025). A leftlateralized dorsolateral prefrontal network for naming. Cell Reports, 44(5). 10.1016/j.celrep.2025.115677 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Y., Magnotti J. F., Zhang X., Wang Z., Yu Y., Davis K. A., Sheth S. A., Chen H. I., Yoshor D., & Beauchamp M. S. (2025). Stereoelectroencephalography Reveals Neural Signatures of Multisensory Integration in the Human Superior Temporal Sulcus during Audiovisual Speech Perception. Journal of Neuroscience, 45(42). 10.1523/JNEUROSCI.1037-25.2025 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhu L. L., & Beauchamp M. S. (2017). Mouth and Voice: A Relationship between Visual and Auditory Preference in the Human Superior Temporal Sulcus. Journal of Neuroscience, 37(10), 2697–2708. 10.1523/JNEUROSCI.2914-16.2017 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zion Golumbic E. M., Poeppel D., & Schroeder C. E. (2012). Temporal context in speech processing and attentional stream selection: A behavioral and neural perspective. Brain and Language, 122(3), 151–161. 10.1016/j.bandl.2011.12.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zion Golumbic E. M., Ding N., Bickel S., Lakatos P., Schevon C. A., McKhann G. M., Goodman R. R., Emerson R., Mehta A. D., Simon J. Z., Poeppel D., & Schroeder C. E. (2013). Mechanisms Underlying Selective Neuronal Tracking of Attended Speech at a “Cocktail Party.” Neuron, 77(5), 980–991. 10.1016/j.neuron.2012.12.037 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Code will be publicly available upon publication of this article. Data are available upon reasonable request.
