Abstract
Successful speech perception requires listeners to bin continuous acoustic information into discrete phonetic categories. However, some people maintain within-category acoustic information (gradient) while others discard category-irrelevant information (discrete) during perception. Listeners also vary in how consistently they label speech sounds and more gradient/consistent labeling has been linked with better speech-in-noise (SIN) perception. Here, we test how neuroanatomical properties of the brain’s major speech-language and auditory pathways relate to individual differences in speech categorization and SIN processing. We measured phonetic categorization and SIN comprehension via phoneme labeling and QuickSIN tasks. Diffusion-weighted imaging (DWI) with probabilistic tractography estimated axonal density within the bilateral arcuate fasciculi and brainstem-cortical auditory projections. Anatomical morphology (surface area, gray matter volume, thickness) was also quantified in the adjacent frontotemporal cortical areas and midbrain. Behaviorally, we found more consistent categorizers had better performance on the QuickSIN. DWI showed that more gradient listeners had greater white matter density in the left arcuate fasciculus and brainstem-cortical auditory pathways, while better SIN performance was predicted by denser white matter in the brainstem-cortical auditory pathways. Morphometric results revealed more consistent listening was associated with greater cortical thickness in right superior temporal gyrus and more gradient listening was associated with greater surface area in right pars opercularis. We infer that individual differences in phonetic categorization relate to SIN comprehension and are at least partially explained by neuroanatomical properties of the auditory-linguistic brain.
Keywords: Categorization, consistency, gradience, speech-in-noise, diffusion-weighted imaging (DWI), magnetic resonance imaging (MRI)
1. Introduction
Listeners differ in how they categorize speech sounds. They vary in how gradiently vs. discretely they label speech sounds (Kapnoula et al., 2021; Kapnoula et al., 2017; Kong & Edwards, 2016; Myers et al., 2024; Rizzi & Bidelman, 2024) and in how they weight within- vs. between-category information to inform their perceptual identification (Kapnoula et al., 2017; McMurray, 2022; Rizzi & Bidelman, 2024). Though all listeners have simultaneous access to both between- and within-category information (Beach et al., 2021; McMurray, 2022; McMurray et al., 2008; Pisoni & Tash, 1974; Toscano et al., 2018), a more gradient listener uses fine-grained acoustic information, while a more discrete listener more heavily weights abstract categorical representations. This weighting results in more gradient listeners having a more linear mapping of acoustics to perception and more discrete listeners having a perceptual space that is strongly warped onto category labels. The degree of perceptual gradience in a listener’s responses can be quantified as the slope of their identification curves derived from labeling speech sounds along an acoustic-phonetic continuum. As in most categorical perception experiments, shallower slopes represent more gradient listening and steeper slopes represent more discrete listening (e.g., Bidelman et al., 2026a; Kapnoula et al., 2021; Kapnoula et al., 2017; Rizzi & Bidelman, 2024).
Similarly, listeners can vary in how consistently they categorize speech sounds (Kim et al., 2025a; Kim et al., 2025b; Kim et al., 2024; Rizzi & Bidelman, 2025). For example, some listeners choose the same label for a speech sound across trials, while others might use different labels across repeated presentations of the same token. Typical categorical perception task paradigms using a two-alternative forced choice (2AFC) task confound the perceptual properties of gradience and consistency: a shallow identification slope could derive from a consistent gradient listener or an inconsistent discrete listener. Moreover, consistency cannot be directly measured with a 2AFC because the nature of the binary response imposed by a 2AFC conflates gradience with consistency (Apfelbaum et al., 2022; Kapnoula et al., 2017; McMurray, 2022). A listener with a shallower identification curve slope from a 2AFC could either be a more gradient or a noisier (less consistent) responder, resulting in averaged responses that are closer to chance. However, consistency and gradience can both be assessed during phoneme labeling with a visual analog scale (VAS) paradigm which avoid binary labels and allows more graded responding to an acoustic-phonetic continuum (Apfelbaum et al., 2022). Thus, individual differences in speech categorization can be more comprehensively assessed when listeners label speech tokens from an acoustic-phonetic continuum using a VAS.
It has been hypothesized that gradience and consistency in auditory categorization might also relate to other important listening skills, including speech-in-noise (SIN) processing. Gradience might reflect the ability to maintain and use probabilistic information during speech perception (Clayards et al., 2008) and thus may reflect higher-level representational strategy. Preserving subphonemic detail with increased gradience could be useful to recover from misinterpretations of ambiguous input (Kapnoula et al., 2021; 2017; Kutlu et al., 2024; Wong et al., 2026), or cope with talker variability. Flexible speech perception allows listeners to adapt their acoustic-phonetic mapping when faced with additional listening challenges such as talker variability (e.g., non-native accented speech), coarticulation, and noise (Assmann & Summerfield, 2004; Bent, 2015; Viswanathan & Kelty-Stephen, 2018). Prior work has shown more gradient listeners have more flexible perception (Kapnoula et al., 2021; Kapnoula et al., 2017; Kutlu et al., 2024; Wong et al., 2026), suggesting they may be better able to adapt to challenges posed by noise-degraded listening. However, findings investigating this hypothesis have been somewhat mixed, with some studies finding gradient listening benefits SIN perception (Myers et al., 2024; Rizzi & Bidelman, 2024) while others fail to observe a relationship between gradience and SIN skills (Kapnoula et al., 2021; Kapnoula et al., 2017; Rizzi & Bidelman, 2025). These mixed findings may be due to high inter-subject variability in gradience especially when assessed across different phonetic contrasts and at the word and sentence level (Kim et al., 2025b; Myers et al., 2024). In contrast, consistency seems to be a more trait-like index of perceptual categorization that remains stable within listener even when assessed across varied phonetic contrasts (Kim et al., 2025b). Consistency might reflect the precision or stability of the speech categorization system and internal noise of the decision process and index fidelity (lower-level precision) of auditory-phonetic encoding rather than the representational format itself (McMurray, 2022). More consistent categorizers have better language and reading skills (Kim et al., 2025a; 2026a), suggesting consistency is necessary for higher-level language comprehension beyond phonetic categorization. Consistency may also be a more sensitive measure of between-subject differences in categorization as it can measure subtle trial-by-trial variance in a listeners’ responses that are lost when averaging data across many trials (Kim et al., 2025b). Thus, if a relationship does exist between categorization skills and SIN perception, consistency may be a stronger, more stable predictor of SIN understanding than gradience (Myers et al., 2024; Rizzi & Bidelman, 2025). More consistent listeners might have a more robust perceptual readout that is less degraded by noise than in an inconsistent listener. Using EEG, we have recently shown that perceptual consistency is related to increased neural consistency in how the brain encodes speech sounds on a trial-by-trial basis (Rizzi et al., 2026). This provides a functional mechanism to account for differences in behavioral categorization. More consistent categorization has also been linked to improved SIN performance (Myers et al., 2024; Rizzi & Bidelman, 2025; Rizzi et al., 2026). While our other work has examined functional/electrophysiological correlates of perceptual gradience and consistency (Rizzi & Bidelman, 2024; Rizzi et al., 2026), it is unclear how underlying anatomical brain structures or pathways might also relate to these perceptual properties of speech categorization.
In this vein, neuroimaging studies have revealed that phonetic categories are clearly represented in inferior frontal gyrus (IFG) (Binder et al., 2004; Blumstein et al., 2005; Lee et al., 2012; Luthra et al., 2019; Myers, 2007; Myers et al., 2009; Rogers & Davis, 2017; Toscano et al., 2018). However, some studies have observed categorical organization in neural responses even earlier in the neural hierarchy in superior temporal gyrus (STG) (Bidelman & Walker, 2019; Chang et al., 2010; Desai et al., 2008; Liebenthal et al., 2005; Mesgarani et al., 2014) and auditory brainstem (Carter & Bidelman, 2023; Rizzi & Bidelman, 2023). Neural responses to speech become more abstract along the auditory pathway, with some neurons selectively responding to speech over nonspeech in STG (Benson et al., 2001; Binder et al., 2000; Humphries et al., 2014; Zatorre et al., 1992) but not primary auditory cortex (Hamilton et al., 2021). And within IFG, the pars opercularis and triangularis are particularly involved in categorical computations (Lee et al., 2012; Mahmud et al., 2021; Papoutsi et al., 2009). Such findings are consistent with the idea that phonetic categories are likely coded at higher levels of the auditory system and downstream from initial sound arrival in A1. While categorical representations are observable in human brainstem FFRs, they have only been observed under attentive states, suggesting category coding is not intrinsic to brainstem, per se, but rather inherited from top-down corticofugal tuning of subcortical activity from higher cortical structures (Bidelman et al., 2013; Carter & Bidelman, 2023; Rizzi & Bidelman, 2023). Together, it is clear that a network of auditory and linguistic brain regions including brainstem, STG, and IFG are relevant to acoustic-phonetic mapping in speech perception.
To date, few studies have investigated relationships between the morphology of auditory-linguistic brain structures, the white matter tracts connecting them, and variance in categorization gradience and consistency. Notably, Fuhrmeister and Myers (2021) found reduced surface area in right middle frontal gyrus predicted more gradient fricative labeling and increased gyrification of bilateral transverse temporal gyri (Heschl’s gyri) predicted less consistent speech categorization, supporting the notion that frontal regions are sensitive to categorical information. While Fuhrmeister and Myers (2021) used a VAS to assess phoneme labeling, the task used discretized steps which provided 7 anchor points (with a 7-token continuum) for listeners to place their response. This design makes it difficult to differentiate between response profiles of listeners attempting to bin tokens into smaller, yet discrete categories, and those utilizing a more continuous listening strategy. Likewise, a listener may appear more consistent under this task if they have anchors to guide their responses. Thus, the structure-behavior relations observed in Fuhrmeister and Myers (2021) could be confounded. In the present study, we use a fully continuous response paradigm to better elucidate the relationships between structural properties of the brain’s auditory-linguistic pathways and perception gradience and consistency.
While volumetric findings in Fuhrmeister and Myers (2021) were restricted to surface area and local gyrification, other metrics of anatomical morphology can also be extracted from MRI brain images including cortical thickness and gray matter volume. There are subtle differences in these measures, with surface area reflecting the number of cell columns in the cortical region and thickness reflecting the number and density of cells within those columns (Rakic, 1988). Gray matter volume is the combination of thickness and surface area, representing both the number of cellular columns and the number of cells within those columns. There are distinct genetic bases (Panizzon et al., 2009; Winkler et al., 2010) and developmental trajectories (Wierenga et al., 2014) for cortical thickness and surface area, suggesting they could differentially relate to behavior. Furthermore, cortical thickness and surface area are phenotypically independent both globally and regionally (Panizzon et al., 2009; Winkler et al., 2010), emphasizing they may have distinct contributions to perception. Thus, measuring all three metrics can aid in a deeper understanding of structure-behavior relationships.
While studies show that fronto-temporal brain regions are critically involved in categorization (e.g., Alho et al., 2016; Bidelman & Walker, 2019; Blumstein et al., 2005; Husain et al., 2006b; Lee et al., 2012; Myers et al., 2009), none to our knowledge have explicitly investigated how connectivity between these structures relate to phonetic categorization skills. White matter (WM) tract density can be readily measured using diffusion weighted imaging (DWI), a form of magnetic resonance imaging (MRI) that measures the amount of water diffusion in the brain in each cardinal direction. In general, water molecules move more easily within and along than through axons. DWI scans can then be used with probabilistic tractography algorithms to construct a map of the major fiber tracks in the brain. From fiber tracking, different metrics of WM density and integrity, such as quantitative anisotropy (QA), can be estimated to understand the microstructural properties of the neuroanatomy. QA provides an estimate of WM density that is more robust to crossing fiber tracks than some other diffusion metrics (Shen et al., 2015; Yeh et al., 2013).
Given the role of STG and IFG in categorization, it is plausible that WM density of arcuate fasciculus (AF) could relate to phonetic categorization. The classical (direct) AF pathway connects Wernicke’s (posterior STG) and Broca’s (IFG) areas with an additional indirect pathway running through the parietal lobe (Catani et al., 2005). Specifically, the indirect pathway can be divided into an anterior segment connecting Broca’s area to the inferior parietal lobe and a posterior segment connecting the inferior parietal lobe to Wernicke’s area (Catani et al., 2005). The first reported function of the AF was for repeating verbal information, an ability that is impaired in conduction aphasia (Wernicke, 1874). However, more modern views acknowledge broader functions of the AF. In the dual stream model of speech, the AF provides direct anatomical connection for the dorsal stream which supports the mapping of acoustic signals to frontal articulatory networks (Hickok & Poeppel, 2007). This dorsal stream is largely left-lateralized and underlies sub-lexical speech sound segmentation and short-term phonological memory and guides articulation, while the ventral stream plays a larger role in comprehension (Hickok & Poeppel, 2004, 2007). Though not originally proposed as a direct function of the dorsal stream, functional neuroimaging studies have revealed patterns of connectivity suggesting phonetic categorization relies on the dorsal stream (Chevillet et al., 2013; Rauschecker, 2012; Zaehle et al., 2008). Thus, commonly observed functional activations of IFG and STG observed in categorization studies likely involve the AF. Supporting this notion, studies have demonstrated microstructure of the AF is related to sensitivity (Perron et al., 2021; Tremblay et al., 2019) and accuracy (Li et al., 2021) of CV discrimination in noise. Microstructure of the AF also relates to phonological processing and dyslexia in left hemisphere (Aeby et al., 2013; Langer et al., 2015; Lebel & Beaulieu, 2009; Perdue et al., 2025; Vandermosten et al., 2012; Yeatman et al., 2011) and to pitch perception in right hemisphere (Chen et al., 2018; Loui et al., 2009; Loui et al., 2011). Thus, it is plausible the AF similarly supports the conversion of acoustic to phonetic information during speech categorization.
WM microstructure along the canonical auditory pathway connecting lower processing centers (brainstem) to primary auditory cortex could also relate to speech processing and categorization. However, few studies have examined white matter tractography of the lemniscal auditory pathways. Among the handful of reports, declines in WM of the brainstem-cortical projections have been linked to tinnitus and hearing loss (Koops et al., 2021; Svobodová et al., 2024). Though it is plausible that denser WM bundles in the central auditory pathways might provide a higher fidelity readout to support speech perception in noise or retain subphonemic acoustic detail, no studies have assessed whether brainstem-cortical WM microstructure relates to auditory perception. Moreover, how categorization gradience/consistency and SIN perception relate to language (AF) pathways also remains unknown. If categorization relates more strongly to auditory system anatomy, this would suggest that acoustic-to-phonetic mapping is a somewhat automatic process and depends on integrity of low-level hearing pathways. On the contrary, if categorization more strongly relates to language pathway anatomy (AF), this would argue that acoustic-to-phonetic processing requires later control mechanisms downstream from the canonical auditory system. If categorization relates to microstructure of both auditory and language WM tracts, it may rely on an interplay between the strength of early auditory encoding and later involvement of auditory-motor integration as implied by the dual stream theory of speech-language processing (Hickok & Poeppel, 2004). It is also likely that gradience and consistency arise from distinct underlying neural mechanisms. Though there have been somewhat mixed findings across studies, it seems these two measures are somewhat separable perceptually (Honda et al., 2024; Kapnoula et al., 2017; Kim et al., 2025; Myers et al., 2024; Rizzi & Bidelman, 2024, 2025), indicating they may relate to distinct neural structures or mechanisms. If gradience and consistency have different relationships with the structure of white matter tracts, this would be further evidence to support that gradience and consistency arise from distinct neural origins.
Here, we extended the functional results of our prior work (Rizzi & Bidelman, 2024; Rizzi et al., 2026) by investigating whether there is also a structural basis to categorization skills (gradience and consistency) and SIN processing. We used MRI volumetrics and DWI neuroimaging to quantify the anatomical morphology and microstructure connectivity within and between major auditory and language regions in cortex (Heschl’s gyrus, STG, and IFG subdivisions) and the subcortical auditory pathways. Based on findings from Fuhrmeister and Myers (2021) that more discrete listeners had greater volume in middle frontal gyrus, we hypothesized that more discrete listeners would also have larger regional volumetrics in IFG due to its crucial role in phonetic categorization. Fuhrmeister and Myers (2021) also demonstrated effects of consistency on the local gyrification of Heschl’s gyrus (though not on other volumetric measures) and functional neuroimaging has suggested auditory cortical regions may be involved in categorization consistency (Rizzi et al., 2026). We therefore expected more consistent listeners would have larger volumetrics in STG but not Heschl’s gyrus, due to its more prominent involvement in categorization. We further hypothesized that density of both AF and auditory brainstem-cortical fiber tracts would relate to variance in categorization skills, emphasizing the role of both the auditory system and later integration via the dorsal stream for phonetic categorization. Specifically, we expected that more gradient listeners who have greater sensitivity to fine acoustic-phonetic details would have denser brainstem-cortical projections and AF, as implied by previous studies (Perron et al., 2021; Tremblay et al., 2019). Relationships between consistency and white matter density were more exploratory in nature, though we hypothesized that consistency and gradience would relate to structure of different tracts, suggesting distinct neural mechanisms. Finally, we hypothesized that denser WM in both auditory and language pathways would also relate to better SIN processing, suggesting an anatomical mechanism for the relationship between categorization and SIN ability.
2. Materials and Methods
The sample included a subset of listeners from our companion EEG experiments who had no contraindications for MRI scanning (EEG data are reported in Rizzi et al., 2026). This included N = 31 young adults (22.65 ± 4.75 years; 9 male, 22 female) with 16.32 ± 2.77 years of education and 8.10 ± 5.96 years of formal music training. Participants were mostly right-handed (75% ± 41% Edinburgh Handedness Inventory; Oldfield, 1971). All participants had normal hearing and had English as their first language. Participants provided written informed consent in accordance with a protocol approved by the Institutional Review Board at Indiana University and were paid $15 an hour for their time.
2.1. Behavioral measures
Behavioral data were collected as described in Rizzi et al. (2026). Listeners labeled 100 ms vowels along an acoustic-phonetic continuum from /u/-/a/ using a visual analog scale (VAS). We chose to use a vowel continuum to promote more gradient responding and thus better reflect individual differences in gradience since vowels tend to be perceived less discretely than consonant-vowel contrasts (Bidelman et al., 2025; Eimas, 1963; Fry et al., 1962; Pisoni, 1973). Stimuli were generated via cascade formant filter synthesis in MATLAB and varied along a first formant frequency (F1) continuum changing from 430 Hz (/u/) to 730 Hz (/a/) and were otherwise acoustically identical (F0: 150 Hz, F2: 1090 Hz, F3: 2350 Hz) (see Bidelman et al., 2013; Carter & Bidelman, 2023; Rizzi et al., 2026). Stimuli were 100 ms in duration, were gated with 10 ms ramps, and were presented at 80 dB SPL over shielded insert headphones (ER-2; Etymotic Research). Phoneme labeling occurred during EEG recording in the first lab visit (further details of EEG recording can be found in Rizzi et al., 2026). We fit individual subject identification curves with a sigmoid P = 1/[1 + e−β1(x−β0)] and estimated parameters using the psignifit function (Schütt et al., 2016) in MATLAB (v2024a). To quantify gradience, we estimated the psychometric slope (β1) from VAS labeling during EEG recording. Consistency was quantified as 1 – σ of VAS responses (distance clicked along the scale). We assessed SIN ability for each listener using the average SNR Loss score from 2 lists of the QuickSIN (Killion et al., 2004) administered binaurally via MATLAB at 70 dB HL over Sennheiser HD 280 circumaural headphones. Listeners’ QuickSIN scores are equal to the number of correctly repeated keywords (five in each sentence and six sentences in each list) subtracted from 25.5, representing their SNR loss, how much higher the SNR must be relative to clinical norms to achieve 50% keyword recall.
2.2. MRI
We collected 3D whole-brain T1-weighted anatomical volumes from each participant (MPRAGE; TE/TR=2.7 ms/2400 ms, 160 axial slices, voxel size=1×1×1 mm3, slice thickness=0.8 mm; FOV=256mm, FA=80) using the 3T Siemens Magnetom Prisma scanner housed in the IU Imaging Research Facility. MRI data were converted to ezBIDS format (https://brainlife.io/ezbids/; Levitas et al., 2024).
2.3. Segmentation
We used the FreeSurfer (v7.3.1; http://surfer.nmr.mgh.harvard.edu/) comprehensive recon-all pipeline (Fischl, 2012) to perform cortical reconstruction and volumetric segmentation on the T1-weighted structural MRI scans using. This automated pipeline includes motion correction (Reuter et al., 2010), removal of non-brain tissue and skull stripping via hybrid watershed/surface deformation (Fischl et al., 2004), automated Talairach transformation, segmentation of white and gray matter volumetric structures (Desikan et al., 2006; Fischl et al., 2004), and cortical surface reconstruction (Dale et al., 1999). MRI processing was conducted on the Indiana University high-throughput Quartz supercomputing cluster (92 compute nodes, each equipped with two 64-core AMD EPYC 7742 2.25 GHz CPUs and 512 GB of RAM).
From the full-brain FreeSurfer output (i.e., aparc + aseg stats table from recon-all), we measured cortical thickness (mm), surface area (mm2), and gray matter volume (mm3) (Dale et al., 1999; Fischl & Dale, 2000; Fischl et al., 1999; Fischl et al., 2004) from each region defined in the Desikan-Killiany atlas parcellation (Desikan et al., 2006). FreeSurfer morphometrics have good test-retest reliability across scanners and various field strengths (Han et al., 2006; Reuter et al., 2012). ROIs for analysis included the eight regions needed to test our hypothesis: bilateral superior temporal gyrus (STG), bilateral transverse temporal gyrus (TTG/Heschl’s gyrus), and bilateral pars opercularis (PO) and pars triangularis (PT) of the IFG. We did not analyze brainstem volume since other metrics (thickness, surface area) cannot be computed from the subcortical atlas.
2.4. Diffusion weighted imaging (DWI)
We used a multishell DWI scheme with two runs of DWI acquisitions with opposite phaseencoding directions (TR/TE=3516/88 ms, SMS acceleration factor=4, b-values=1000 and 2500 s/mm2, sampling directions=76 and 74, plus 5 volumes with b=0, isotropic in plane resolution=1.5 × 1.5 × 1.5 mm3 with whole-brain coverage, total scan time=10 mins). The two DWI scans were combined and corrected for susceptibility and Eddy current distortions using the FSL software suite (Jenkinson et al., 2012). Deterministic fiber tracking was then performed using DSI Studio software (version “Chen”) (https://dsi-studio.labsolver.org/) (Yeh et al., 2016; Yeh et al., 2010). The diffusion data were reconstructed in the MNI space using q-space diffeomorphic reconstruction (Yeh & Tseng, 2011) to obtain the spin distribution function (Yeh et al., 2010). A diffusion sampling length ratio of 1.25 was used. The output resolution in diffeomorphic reconstruction was 1.5 mm isotropic. The restricted diffusion was quantified using restricted diffusion imaging (Yeh et al., 2017). The tensor metrics were calculated using DWI with b-value lower than 1750 s/mm2. There was one scan with poor quality that was excluded from further analysis, resulting in 30 total scans.
2.5. Fiber tracking
We computed fiber tracks restricted to the bilateral auditory brainstem-cortical pathways based on the HCP tractography atlas as defined in DSI studio (Yeh et al., 2018) and ICBM 152 adult template brain landmarks (Mazziotta et al., 1995). Auditory tracks were generated by seeding ROIs in the “brainstem” and “auditory sensory” regions of the CerebrA and Campbell atlases, respectively. This constraint produced similar streamlines to the in vivo human auditory brainstem white matter atlas described by Sitek et al. (2019). Autotrack was used to automatically identify auditory fibers with a distance tolerance of 24 mm in the ICBM152 space. The anisotropy threshold was randomly selected between 0.5 and 0.7 Otsu threshold. The angular threshold was randomly selected from 45 to 90°. The step size was set to voxel spacing. A total of up to 5000 streamlines were generated for each pathway. Auditory tracks with length shorter than 5 or longer than 100 mm were discarded. We also computed tracks restricted to bilateral arcuate fasciculus. AF tracks with lengths shorter than 30 mm or longer than 300 mm were discarded. Autotolerance for tracking was set at 24. We then extracted the mean quantitative anisotropy (QA) of each tract bundle per participant, reflecting the overall density of each anatomical pathway. Left and right hemispheres were analyzed separately.
2.6. Statistical Analysis
Unless otherwise specified, we used linear mixed effects models (lme4 package version 1.1–32 in R version 4.2.1; Bates et al., 2015) to assess differences in neural (MRI volumetrics and WM tract QA) and behavioral (categorization consistency, gradience, and QuickSIN score) dependent variables. Full model details are reported below. Pairwise contrasts were Tukey-adjusted to account for multiple comparisons and family-wise error corrections were employed where appropriate. Degrees of freedom for mixed models were estimated using Satterthwaite’s method. Handedness (p > 0.85) and years of music training (p > 0.05) did not relate to behavioral measures in our sample and were therefore not included in our statistical models.
3. Results
3.1. Behavioral data
Categorization consistency and SIN perception were correlated [r(28) = −0.39, p = 0.038]. More consistent categorizers had better SIN scores relative to inconsistent categorizers (Fig. 1).
Figure 1.

More consistent categorizers in vowel labeling perform better on the QuickSIN. Shading = 95% CI.
3.2. Structural data
Descriptive statistics of surface area, gray matter volume, and cortical thickness of auditory-linguistic ROIs can be found in supplementary material (Supplementary Fig. 1 and Table 1). We assessed whether cortical morphology was related to our behavioral outcomes (QuickSIN score, categorization gradience). Since structural MRI measures of area, volume, and thickness reflect different anatomical properties but are inherently colinear, we built separate linear regressions for each MRI measure (z-scored within-measure) per ROI. This approach resulted in 72 models (3 behavioral outcomes × 3 MRI measures × 8 ROIs). For each ROI, we applied a Holm family-wise correction for the family of regressions predicting each behavioral measure with the different anatomical metrics (e.g., QuickSIN ~ STGSA; QuickSIN ~ STGGMV; QuickSIN ~ STGCT). This approach was taken because different behavioral measures and different ROIs addressed different questions and thus was intended to balance Type I and Type II errors.
No left hemisphere structure predicted behavioral outcomes including categorization (all pHolm > 0.248) and QuickSIN scores (all pHolm > 0.332). Surface area in RH PO (IFG) predicted categorization gradience [F(1, 27) = 7.27, p = 0.012, pHolm = 0.036, ] (Fig. 2A) with more gradient listeners having greater right PO surface area. Right hemisphere STG thickness also marginally predicted categorization consistency [F(1, 27) = 5.43, p = 0.0275, pHolm = 0.082, ] (Fig. 2B) with more consistent listeners having thicker right STG. Categorization consistency was not predicted by MRI measures in any other right hemisphere ROI (all other pHolm > 0.368).
Figure 2. Structure-behavior relationships.

(A-B) Estimated effects from models described in text with corrected and uncorrected p values shown. (A) More gradient listeners have larger surface area in RH IFG (pars opercularis). (B) More consistent listeners have greater cortical thickness in RH STG (auditory cortical region), though this effect does not remain significant when family-wise error correction is applied. Shading = 95% CI.
3.3. Tractography data
Grand average tractography of the arcuate fasciculus and brainstem-cortical projections are shown in Fig. 3 and mean QA across tracts is shown in Supplementary Fig. 2. We examined whether QA across tracts related to SIN performance. We used a linear fixed effects model predicting QuickSIN with fixed effects of QA from each tract across brainstem and cortical levels [QuickSIN ~ QAAF_Left + QAAF_Right + QABS_Left + QABS_Right]. However, variance inflation factors (VIFs: 8.0–13.9) indicated high multicollinearity among the predictors. To address this collinearity, we collapsed across hemispheres to investigate whether QuickSIN was predicted by QA at the brainstem and/or cortical level [QuickSIN ~ QABS_mean + QAAF_mean]. This model reduced VIFs to >5.7, indicating more acceptable levels of multicollinearity (Kim, 2019). The model revealed that QA of the brainstem-cortical projections (but not AF) predicted QuickSIN [F(1, 27) = 4.45, p = 0.044, ], indicating greater density in the brainstem fiber tracts predicted better (i.e., lower) QuickSIN scores (Fig. 4B). To further investigate hemispheric differences, we also analyzed left and right pathways separately [i.e., QuickSIN ~ QABS_Left + QABS_Right; QuickSIN ~ QAAF_Left + QAAF_Right]. Though our first model indicated QA of brainstem-cortical projections predicted QuickSIN scores, these models revealed QA did not predict QuickSIN scores in any tract individually (all p > 0.08).
Figure 3.

Grand average tractography of bilateral arcuate fasciculus (A-B) and brainstem-cortical projections (C).
Figure 4. Structure-behavior relations between DWI tractography and categorization/SIN perception.

(A) Categorization slope effects. White matter density (QA) shows opposite relationships to perceptual gradience across hemispheres. More gradient listeners have denser white matter in LH brainstem-cortical projections and AF, but less dense RH brainstem-cortical projections. (B) Listeners with better SIN scores have denser white matter in brainstem-cortical projections averaged across hemispheres. Shading = 95% CI.
We next investigated whether QA predicted categorization gradience (identification slopes). We again started with a full model predicting behavioral slope from QA at cortical and subcortical levels, as above. VIFs in this model also indicated high multicollinearity (VIF=8 to 14). To improve multicollinearity issues as well as assess possible hemispheric asymmetries, we analyzed left and right pathways separately [i.e., slope ~ QABS_Left + QABS_Right; slope ~ QAAF_Left + QAAF_Right]. For each of these models, VIFs remained under 2.65, indicating acceptable multicollinearity. For the brainstem model, we found significant main effects of both RH [F(1, 27) = 5.18, p = 0.031, ] and LH [F(1, 27) = 9.38, p = 0.005, ] QA on slope. The direction of these effects was inverse across hemispheres with a positive estimate for RH and a negative estimate for LH. That is, more gradient listeners had denser brainstem-cortical projections in the LH but sparser projections in RH (Fig. 4A). For the AF model, there was a significant main effect of LH QA [F(1, 27) = 6.93, p = 0.014, ], whereby greater QA in the left arcuate predicted more gradient categorization (Fig. 4A).
We similarly assessed a linear regression model with QA from each tract predicting categorization consistency. Diagnostic VIFs ranged from 8.1 to 14.0. As with slopes, we built separate models for each anatomical level (i.e., consistency ~ QABS_Left + QABS_Right; consistency ~ QAAF_Left + QAAF_Right). Multicollinearity was acceptable from these models (VIF < 2.65). However, QA did not predict categorization consistency in either model (all p > 0.62).
4. Discussion
We assessed whether individual differences in speech categorization and SIN perception related to structural morphology and anatomical connectivity between major nodes of the brain’s auditory-linguistic network including the arcuate fasciculus and central auditory pathways. We found that more gradient listeners had greater surface area in right IFG, while more consistent listeners had thicker right STG, suggesting perceptual gradience and consistency may be supported by linguistic- and auditory-centric brain regions, respectively. Gradience also correlated with white matter tract density differentially across hemispheres. At a cortical level, more gradient listeners had denser white matter tracts in LH arcuate fasciculus. At a subcortical level, gradient listeners also showed denser projections of the auditory brainstem pathways in LH and sparser projection in RH compared to more discrete listeners. Denser auditory projections were further associated with better QuickSIN scores.
4.1. Structural Findings
4.1.1. Categorization gradience and consistency relate to volumetrics of auditory-linguistic brain regions.
By examining volumetric properties of listeners’ MRIs, we found that distinct auditory and linguistic brain regions related to different perceptual constructs of phonetic speech categorization. The relationship between brain structure and gradience was constrained to frontal ROIs (right PO), while that with consistency was constrained to temporal ROIs (right STG). Although entirely structural in nature, these findings are consistent with prior functional neuroimaging studies demonstrating frontal language regions, including PO, respond categorically (i.e., non-gradiently) to speech sound continua (Alho et al., 2016; Bidelman & Walker, 2019; Husain et al., 2006a; Lee et al., 2012; Luthra et al., 2019; Myers, 2007; Myers et al., 2009), as well as our recent EEG data showing perceptual consistency is functionally localized to activity from auditory cortex (Rizzi et al., 2026). That our volumetric findings were restricted to the right hemisphere is likely due to the use of vowels. Vowels are distinguished by steady-state spectral information and rely on longer integration time windows that better engage right hemispheric auditory processing (Boemio et al., 2005; Poeppel, 2003; Sininger & Bhatara, 2012; Sininger & Cone-Wesson, 2004; Zatorre & Belin, 2001). A similar rightward lateralization was observed relating fricative categorization to structural brain morphology (Fuhrmeister & Myers, 2021).
Our findings support prior work demonstrating individual differences in perceptual gradience vs. consistency might be segregated at the neuroanatomical level–supported by structural properties of frontal vs. auditory regions, respectively (Fuhrmeister & Myers, 2021). However, we hypothesized that more discrete listeners would have greater surface area in frontal regions based on findings from Fuhrmeister and Myers (2021) who demonstrated a relationship in right middle frontal gyrus (MFG). Husain et al. (2006a) also showed greater activation of MFG when categorizing nonspeech relative to speech sounds. While activation in MFG are common during tasks with categorical decisions (Mahmud et al., 2021), many studies, including our data here, emphasize the importance of adjacent IFG in speech categorization (e.g., Alho et al., 2016; Blumstein et al., 2005; Lee et al., 2012; Meyers et al., 2008; Myers et al., 2009). In right IFG (pars opercularis), we found an opposite relationship where increased surface area predicted more gradient responding. Similarly, Golestani et al. (2011) demonstrated larger left pars opercularis volume correlated with the number of years an individual had with phonetic transcription experience. Though they did not assess auditory categorization explicitly, expert phoneticians presumably have greater sensitivity to fine-grained acoustic detail (i.e., more gradience) which may allow more flexible categorization when transcribing acoustic input into an orthographic form.
We also found categorization consistency assessed under a VAS task was positively related to cortical thickness of right STG supporting our initial hypothesis that more consistent listeners have larger volumetrics in auditory cortex. Though, we note this relation did not survive family-wise error correction so the effect should be interpreted cautiously. Similarly, Fuhrmeister and Myers (2021) found less consistent listeners had increased gyrification of bilateral transverse temporal gyri (i.e., HG). Though we did not find a relationship between consistency and morphology of HG in our analysis, we only investigated relationships with surface area, cortical thickness, and gray matter volume, which were similarly not related to behavior in Fuhrmeister and Myers (2021). Though gyrification and cortical thickness are distinct measures, they are often negatively related in MRI morphometry studies in healthy middle aged adults (Gautam et al., 2015). Similarly, children with dyslexia have reduced cortical thickness and increased local gyrification in occipitotemporal regions, suggesting an inverse relationship between these metrics related to reading ability (Williams et al., 2017). Thus, the seemingly opposite relationships with consistency observed here with cortical thickness and in Fuhrmeister and Myers (2021) with local gyrification are easy to reconcile. The overall structure of auditory cortical regions, both decreased gyrification and increased cortical thickness, predict a more consistent categorizer. Gyrification of transverse HG is developed in utero and changes little with environmental factors (Golestani et al., 2011). However, cortical thickness can change with social-environmental factors such as socioeconomic status (Piccolo et al., 2016) and perceptual experience. For instance, STG thickness increases after foreign language learning (Mårtensson et al., 2012), simultaneous language interpretation training in multilinguals (Hervais-Adelman et al., 2017), and balance training (Rogge et al., 2018). Similarly, children who participated in musical training had slowed rates of cortical thinning in STG relative to controls (Habibi et al., 2020). Thus, it is possible that some listeners are more predisposed to be perceptually consistent which may relate to an anatomical morphotype with less local gyrification of transverse HG. In contrast, those who become more perceptually consistent through auditory experience have increased cortical thickness in STG. In other words, specific structural differences observed between these studies could arise from distinct origins (i.e., structural predispositions vs. experience-dependent changes in auditory perceptual skills). Our sample did not have a spread in language experience. While there was a spread in musical experience, music training was not a significant factor in any of our models when added as a covariate. When interpreted collectively with our findings, it appears consistency in categorizing speech sounds relates to structural properties of auditory cortical regions, though to different structural properties across primary and secondary regions.
4.1.2. Categorization gradience and consistency may be independent processes.
The finding that gradience relates to structure of IFG while consistency relates to structure of STG suggests that these two perceptual processes of speech categorization might be subserved by distinct neural regions. Whether categorization gradience and consistency are independent attributes of behavior has been equivocal across studies (Honda et al., 2024; Kapnoula et al., 2017; Kim et al., 2025a; Myers et al., 2024; Rizzi & Bidelman, 2024, 2025); some reports show gradience is correlated with consistency in phoneme labeling while others do not. Though structural neuroimaging alone cannot discern whether perceptual processes rely on functionally distinct neural mechanisms, the anatomy of auditory cortical and brainstem regions seems to at least partially explain functional elements of auditory encoding (Bidelman et al., 2026b) and their correspondence with different anatomical regions supports the notion that they reflect independent constructs in categorization. Functionally, consistency may be more localized to auditory cortical regions (or even subcortex as suggested by Rizzi et al., 2026), whereas gradience may arise from more frontal brain areas. Prior work has shown activity in auditory cortical regions relates to the accuracy of sound identification (i.e., the quality of speech representation), while inferior frontal activity relates to the speed of identification, suggesting a functional dissociation between sensory and decision-based processes in the brain’s fronto-temporal pathways (Binder et al., 2004). If consistency arises earlier in the auditory system than gradience, gradience may be a higher-level process related to stimulus decision. That is, consistency might arise from more automatic, sensory processing. This could explain why perceptual consistency seems to be more trait-like (Kim et al., 2025b), while perceptual gradience is less stable across different phonetic contrasts (Kapnoula et al., 2021; Kim et al., 2025b; Myers et al., 2024). Though, Kapnoula and McMurray (2021) found gradient listeners had more gradient representations of voice onset time in N1 ERP amplitudes whereas discrete listeners had more categorical representations, suggesting early auditory encoding could be modulated by gradience. Future studies using both structural and functional neuroimaging could further elucidate whether gradience and consistency are independent.
4.2. Structural connectivity findings
4.2.1. Denser WM in auditory pathways predicts better SIN understanding.
DWI tractography showed that listeners with better SIN understanding had denser white matter in bilateral auditory brainstem-cortical projections that comprise the canonical central auditory system pathways. Our hemispheric analysis revealed that QA in brainstem-cortical projections from either hemisphere alone did not predict QuickSIN scores, suggesting this relationship is not strongly lateralized, and denser white matter bilaterally is more important for understanding SIN than density within a single hemisphere. Fiber tracking does not assess functional connectivity between brain regions nor the directionality of electrophysiological signaling (only form, size, and shape of the anatomy). Thus, whether the relation between the auditory brainstem projections and SIN processing observed here is due to afferent or efferent neurophysiological function cannot be determined from our purely structural DWI data. Denser auditory pathways could support stronger and more efficient “bottom-up” encoding of speech and/or stronger efferent feedback from the descending corticofugal pathways that provide greater “top-down” control over auditory processing at lower levels. Theoretically, both the ascending and descending systems are critical for SIN processing. Robust SIN understanding requires strong bottom-up encoding to maximize the quality of sensory representation (e.g., Anderson et al., 2012; Anderson et al., 2011; Coffey et al., 2017; Parbery-Clark et al., 2009a). Indeed, reduced afferent functional connectivity between auditory midbrain and cortex correlates with poorer SIN comprehension as measured by the QuickSIN (Bidelman et al., 2019). Several electrophysiological studies in animals and humans also suggest that the corticofugal efferent system, including the cortical-brainstem connections measured here, assist in noise-degraded speech listening (Asilador & Llano, 2021; de Boer & Thornton, 2008; Price & Bidelman, 2021). Thus, the structure-behavior relationship we observe here could support either enhanced afferent or efferent functional connectivity, both of which might aid SIN perception. Further studies should systematically investigate relationships between directional brainstem-auditory tractography, functional connectivity, and SIN processing.
Contrary to our hypothesis, we did not observe a relationship between SIN performance and AF density. Although limited, prior work has suggested denser or higher integrity AF relates to better SIN perception, motivating our analysis of this pathway with DWI tractography (Li et al., 2021; Perron et al., 2021; Tremblay et al., 2019). However, these previous anatomical studies only assessed CV discrimination in noise (Li et al., 2021; Perron et al., 2021; Tremblay et al., 2019) rather than sentence-level SIN processing as used here. Given the documented role of the AF in phonological processing (Lebel & Beaulieu, 2009; Perdue et al., 2025; Tremblay et al., 2019; Vandermosten et al., 2012; Yeatman et al., 2011), it is possible sentence-level SIN perception, assessed by the QuickSIN, recruits different neural regions than simpler phonetic discrimination in noise tasks used in prior work (Li et al., 2021; Perron et al., 2021; Tremblay et al., 2019). Most studies also examined participants with extensive musical training, and musicians are known to have enhanced SIN perception (for review, see Hennessy et al., 2022; Maillard et al., 2023) and stronger AF (Halwani et al., 2011; Li et al., 2021; Moore et al., 2017; Oechslin et al., 2009; Perron et al., 2021). For example, Li et al. (2021) observed a relationship between SIN and AF microstructure but only in highly trained musicians while Perron et al. (2021) did not observe this effect in amateur singers (vs. non-singers). Similarly, while SIN advantages have been reported for musicians (e.g., Bidelman & Krishnan, 2010; Bidelman & Yoo, 2020; Hennessy et al., 2022; Maillard et al., 2023; Parbery-Clark et al., 2009b; Parbery-Clark et al., 2011; Slater et al., 2015; Yoo & Bidelman, 2019), this effect is not always observed (e.g., Boebinger et al., 2015; Madsen et al., 2019; Ruggles et al., 2014; Yeend et al., 2017). We did not explicitly recruit musicians in our study. Thus, it is possible we were unable to observe a relationship between SIN and AF given the relative homogeneity in these measures among out musically naïve listeners.
4.2.2. Denser WM in LH auditory and language pathways predicts more gradient listening.
With regard to anatomical correlates of categorization, we found significant asymmetries in the cortical auditory-linguistic pathways. More gradient listeners had denser AF in left hemisphere. This finding largely agrees with our hypotheses. We predicted denser WM in left hemisphere auditory and language tracts would correspond to more gradient listening. Phonetic categorization is largely a LH process with more leftward lateralized auditory and language regions specialized for phonetic category processing (Blumstein et al., 2005; Desai et al., 2008; Husain et al., 2006a; Joanisse et al., 2007; Liebenthal et al., 2005; Wolmetz et al., 2011). One study found that sensitivity on CV discrimination was related to diffusivity in bilateral AF, suggesting more developed AF in both hemispheres predicts higher sensitivity to phonetic details (Tremblay et al., 2019). Similar findings were reported by Perron et al. (2021) for left AF. Though their tasks were not canonical categorization tasks, it is possible more gradient listeners in these studies were more sensitive to the acoustic details differentiating CV contrasts (e.g., Perron et al., 2021; Tremblay et al., 2019). This could explain why gradient listening was associated with denser WM in LH auditory-language pathways. This finding also mirrors results from our recent functional (EEG) neuroimaging studies which showed that more gradient listeners had stronger left-lateralized speech activations in the temporal lobe (Rizzi & Bidelman, 2024).
The leftward asymmetry of this effect suggests that auditory perceptual gradience might depend on the dorsal speech processing stream (Hickok & Poeppel, 2004). The dorsal stream is a functional pathway that runs from left STG through the temporoparietal junction with termination in IFG. A major anatomical component of the dorsal network is the AF which is traditionally involved in sensorimotor integration and mapping sound to articulatory representations (Hickok & Poeppel, 2004, 2007). Functional neuroimaging work has suggested dorsal stream involvement in speech categorization (Alho et al., 2016; Chevillet et al., 2013; Lee et al., 2012). We extend these prior results by establishing a structural basis for such effects. Moreover, our findings lead us to infer that the dorsal stream not only contributes to phonetic categorization very broadly, but more importantly, to how nuanced (gradient) a listener maps speech sounds to their identity.
Notably, we also found structural asymmetries within the auditory brainstem-cortical pathways. Though some studies have tracked changes in the central auditory brainstem pathways in relation to hearing acuity and tinnitus (Koops et al., 2021; Svobodová et al., 2024), none to our knowledge have explored whether their density maps to perceptual processing. Here, we show that more gradient listeners have denser left but sparser right WM in their auditory tracts, supporting our hypothesis that more continuous modes of perception would relate to auditory system neuroanatomy. Presumably, denser WM in auditory system may allow for more robust auditory encoding and precise capture of fine-grained acoustic details for later perceptual judgements. Interestingly, the lateralization of the brainstem data is also internally consistent with the cortical results: more gradient listeners have denser WM along the left auditory brainstem pathways that persists into left lateralized language circuitry at the cortical level.
4.3. Limitations
It is possible that other brain regions which we did not analyze relate to categorization behavior. For instance, the planum temporale is another auditory cortical region which has been involved in phonetic learning and has been related to phonetic categorization performance (Callan et al., 2003; Elmer et al., 2013; Griffiths & Warren, 2002). The planum temporale is not included in the Desikan-Killiany parcellation atlas using Freesurfer, informing our choice to not include it as an ROI in our investigation. Future studies could use more fine-grained atlases to measure gross anatomy of this and other smaller auditory cortical regions to investigate the specificity of the relationship observed here between auditory cortical morphology and phonetic categorization.
Further, we analyzed gradience and consistency independently, though it is not entirely clear whether these are fully independent constructs. Prior work has suggested that highly gradient but inconsistent listeners may have noisier processing rather than truly highly gradient processing (Apfelbaum et al., 2022). An advantage of using the VAS to measure categorization behavior is that both consistency and gradience can be measured to generate a fuller profile of a listener’s categorization behavior (Apfelbaum et al., 2022; Kutlu et al., 2024; Sorensen et al., 2024). Prior empirical studies have had mixed findings regarding whether gradience and consistency are related within listeners (Kapnoula et al., 2017; Kim et al., 2025; Myers et al., 2024; Rizzi & Bidelman, 2024, 2025). We chose to analyze these measures independently due to the smaller sample size of our sample and we found gradience and consistency were not correlated (r(28) = 0.19, p = 0.32). However, interpretations of gradience without accounting for consistency should be interpreted more cautiously, as it is possible some gradient listeners here have noisier processing or responses.
Though a sample size of 31 is comparable to or larger than many other MRI/fMRI studies of phonetic categorization (Blumstein et al., 2005; Elmer et al., 2013; Liu & Wang, 2025; Myers & Swan, 2012), it is possible that our sample was not large enough to detect effects of individual differences across all of the variables we measured. Our sample size was based on a power analysis to detect a d = 1.25 effect size (80% power, paired t-test, α=0.05) based on our prior effect sizes from behavioral differences in slope performed for our EEG experiment (see Rizzi et al., 2026). However, it is possible smaller effects were not adequately powered. Similarly, it is possible that there are inflated false-positive findings due to the numerous behavioral and neural measures used in our hypothesis testing. While we controlled for family-wise error using Holm-corrected p-values for our morphological statistics and Tukey-adjusted p-values for multiple comparisons in linear mixed models, it is possible that there were other false-positive findings in the results. It is also possible that our findings are not generalizable across different speech continua. While it seems that consistency is more stable across different phonetic contrasts, gradience can vary within subject across phonetic contrasts (Bidelman et al., 2025; Kim et al., 2025; Kim et al., 2026b; Myers et al., 2024). Thus, whether the neural bases of vowel categorization are stable across phonetic contrasts is not known. Future studies could measure categorization across multiple phonetic contrasts and with a larger sample size to determine how generalizable these findings are.
5. Conclusions
We measured volumetrics (MRI) and structural connectivity (DWI) within and between major auditory and language brain and evaluated their relationship to auditory-perceptual categorization and SIN abilities. Structurally, we found gradience and consistency in speech-sound labeling related to volumetric measures in distinct brain regions; increased gradience predicted by greater surface area in right IFG and increased consistency predicted by thicker right STG. These anatomical findings imply that a consistent perceptual readout of the speech signal may be more automatic and supported by auditory cortical regions, while gradience may arise from higher-order frontal regions, which could be more influenced by attentional processing.
Structural tractography demonstrated denser WM in the auditory system was related to both improved SIN perception and categorization gradience, emphasizing the role of central auditory pathway integrity for multiple auditory perceptual skills. Specifically, increased gradience was predicted by denser left-lateralized WM in both the auditory and language tracts, suggesting these LH tracts may play a role in retaining subphonemic detail up to the level of IFG for phonetic categorization.
This work expands on our prior functional EEG experiments (Rizzi & Bidelman, 2024; Rizzi et al., 2026) by revealing structural-behavioral relationships that also underly individual differences in listening strategies as they relate to speech categorization. The MRI/DWI results herein highlight the importance of the central auditory system pathways and dorsal speech processing stream in accounting for gradience and consistency in auditory category mapping. That brain-behavior relationships were largely left lateralized across subcortical and cortical fiber tracts implies an integrated network of auditory and language pathways that underly phonetic categorization gradience.
Supplementary Material
Highlights.
More consistent categorization predicted thicker right superior temporal gyrus.
More gradient categorization predicted greater right pars opercularis surface area.
Denser white matter in auditory-linguistic pathways related to more gradience and better speech-in-noise perception.
Acknowledgements:
The authors thank Tessa Bent, Jennifer Lentz, and Samantha Gustafson for their comments on earlier versions of this manuscript.
Funding:
This work was supported by the National Institutes of Health/National Institute on Deafness and Other Communication Disorders R01DC016267 (awarded to G. M. B.) and F31DC023124 (awarded to R. R.). This project was also supported by the Indiana Clinical and Translational Sciences Institute (CTSI), funded in part by grant #UL1TR002529 from the National Institutes of Health, National Center for Advancing Translational Sciences, Clinical and Translational Sciences Award.
Footnotes
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
Conflict of interest: The authors have no relevant financial or non-financial interests to disclose.
Ethical approval: This study was performed in line with the principles of the Declaration of Helsinki. Approval was granted by the Institutional Review Board of Indiana University (No. 15650, approved 12/9/25).
Data Availability:
The datasets generated during the current study are available from the corresponding author on reasonable request.
References
- Aeby A, De Tiège X, Creuzil M, David P, Balériaux D, Van Overmeire B, Metens T, & Van Bogaert P. (2013). Language development at 2 years is correlated to brain microstructure in the left superior temporal gyrus at term equivalent age: A diffusion tensor imaging study. Neuroimage, 78, 145–151. 10.1016/j.neuroimage.2013.03.076 [DOI] [PubMed] [Google Scholar]
- Alho J, Green BM, May PJC, Sams M, Tiitinen H, Rauschecker JP, & Jääskeläinen IP. (2016). Early-latency categorical speech sound representations in the left inferior frontal gyrus. Neuroimage, 129, 214–223. 10.1016/j.neuroimage.2016.01.016 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Anderson S, Parbery-Clark A, White-Schwoch T, & Kraus N. (2012). Aging affects neural precision of speech encoding. The Journal of Neuroscience, 32(41), 14156–14164. 10.1523/jneurosci.2176-12.2012 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Anderson S, Parbery-Clark A, Yi H-G, & Kraus N. (2011). A Neural Basis of Speech-in-Noise Perception in Older Adults. Ear and Hearing, 32(6), 750–757. 10.1097/AUD.0b013e31822229d3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Apfelbaum KS, Kutlu E, McMurray B, & Kapnoula EC. (2022). Don’t force it! Gradient speech categorization calls for continuous categorization tasks. The Journal of the Acoustical Society of America, 152(6), 3728–3745. 10.1121/10.0015201 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Asilador A, & Llano DA. (2021). Top-Down Inference in the Auditory System: Potential Roles for Corticofugal Projections [Review]. Frontiers in Neural Circuits, Volume 14 - 2020. 10.3389/fncir.2020.615259 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Assmann P, & Summerfield Q. (2004). The Perception of Speech Under Adverse Conditions. In Greenberg S, Ainsworth WA, Popper AN, & Fay RR. (Eds.), Speech Processing in the Auditory System (pp. 231–308). Springer; New York. 10.1007/0-387-21575-1_5 [DOI] [Google Scholar]
- Bates D, Mächler M, Bolker B, & Walker S. (2015). Fitting Linear Mixed-Effects Models Using lme4. Journal of Statistical Software, 67(1), 1–48. 10.18637/jss.v067.i01 [DOI] [Google Scholar]
- Benson RR, Whalen DH, Richardson M, Swainson B, Clark VP, Lai S, & Liberman AM. (2001). Parametrically dissociating speech and nonspeech perception in the brain using fMRI. Brain and Language, 78(3), 364–396. [DOI] [PubMed] [Google Scholar]
- Beach SD, Ozernov-Palchik O, May SC, Centanni TM, Gabrieli JDE, & Pantazis D. (2021). Neural Decoding Reveals Concurrent Phonemic and Subphonemic Representations of Speech Across Tasks. Neurobiology of Language, 2(2), 254–279. 10.1162/nol_a_00034 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bent T. (2015). Development of perceptual flexibility International Congress of Phonetic Sciences. [Google Scholar]
- Bidelman GM, Bernard F, & Skubic K. (2025). Hearing in categories and speech perception at the “cocktail party”. PLoS One, 20(1), e0318600. 10.1371/journal.pone.0318600 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bidelman GM, Eisenhut Z, Borowski L, Rizzi R, & Pisoni DB. (2026a). Elliptical speech reveals the use of broad phonetic categories aids noise-degraded speech perception. bioRxiv. 10.64898/2026.01.02.695202 [DOI] [Google Scholar]
- Bidelman GM, & Krishnan A. (2010). Effects of reverberation on brainstem representation of speech in musicians and non-musicians. Brain Research, 1355, 112–125. 10.1016/j.brainres.2010.07.100 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bidelman GM, Moreno S, & Alain C. (2013). Tracing the emergence of categorical speech perception in the human auditory system. Neuroimage, 79, 201–212. 10.1016/j.neuroimage.2013.04.093 [DOI] [PubMed] [Google Scholar]
- Bidelman GM, Price CN, Shen D, Arnott SR, & Alain C. (2019). Afferent-efferent connectivity between auditory brainstem and cortex accounts for poorer speech-in-noise comprehension in older adults. Hearing Research, 382, 107795. 10.1016/j.heares.2019.107795 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bidelman GM, Stirn JR, Rizzi R, MacLean JA, & Cheng H. (2026b). Auditory Brainstem–Cortical Anatomy Relates to the Magnitude of Frequency-Following Responses (FFRs) and Event-Related Potentials (ERPs) Coding Speech-in-Noise. Neuroimaging, 1(1), 6. https://www.mdpi.com/3042-8807/1/1/6 [Google Scholar]
- Bidelman GM, & Walker B. (2019). Plasticity in auditory categorization is supported by differential engagement of the auditory-linguistic network. Neuroimage, 201, 116022. 10.1016/j.neuroimage.2019.116022 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bidelman GM, & Yoo J. (2020). Musicians Show Improved Speech Segregation in Competitive, Multi-Talker Cocktail Party Scenarios [Brief Research Report]. Frontiers in Psychology, Volume 11 - 2020. 10.3389/fpsyg.2020.01927 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Binder JR, Frost JA, Hammeke TA, Bellgowan PS, Springer JA, Kaufman JN, & Possing ET. (2000). Human temporal lobe activation by speech and nonspeech sounds. Cerebral Cortex, 10(5), 512–528. [DOI] [PubMed] [Google Scholar]
- Binder JR, Liebenthal E, Possing ET, Medler DA, & Ward BD. (2004). Neural correlates of sensory and decision processes in auditory object identification. Nature Neuroscience, 7(3), 295–301. 10.1038/nn1198 [DOI] [PubMed] [Google Scholar]
- Blumstein SE, Myers EB, & Rissman J. (2005). The Perception of Voice Onset Time: An fMRI Investigation of Phonetic Category Structure. Journal of Cognitive Neuroscience, 17(9), 1353–1366. 10.1162/0898929054985473 [DOI] [PubMed] [Google Scholar]
- Boebinger D, Evans S, Rosen S, Lima CF, Manly T, & Scott SK. (2015). Musicians and non-musicians are equally adept at perceiving masked speech. The Journal of the Acoustical Society of America, 137(1), 378–387. 10.1121/1.4904537 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boemio A, Fromm S, Braun A, & Poeppel D. (2005). Hierarchical and asymmetric temporal sensitivity in human auditory cortices. Nature Neuroscience, 8(3), 389–395. 10.1038/nn1409 [DOI] [PubMed] [Google Scholar]
- Callan DE, Tajima K, Callan AM, Kubo R, Masaki S, & Akahane-Yamada R. (2003). Learning-induced neural plasticity associated with improved identification performance after training of a difficult second-language phonetic contrast. Neuroimage, 19(1), 113–124. 10.1016/S1053-8119(03)00020-X [DOI] [PubMed] [Google Scholar]
- Carter JA, & Bidelman GM. (2023). Perceptual warping exposes categorical representations for speech in human brainstem responses. Neuroimage, 269, 119899. 10.1016/j.neuroimage.2023.119899 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Catani M, Jones DK, & ffytche DH. (2005). Perisylvian language networks of the human brain. Annals of Neurology, 57(1), 8–16. 10.1002/ana.20319 [DOI] [PubMed] [Google Scholar]
- Chang EF, Rieger JW, Johnson K, Berger MS, Barbaro NM, & Knight RT. (2010). Categorical speech representation in human superior temporal gyrus. Nat Neurosci, 13(11), 1428–1432. 10.1038/nn.2641 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen X, Zhao Y, Zhong S, Cui Z, Li J, Gong G, Dong Q, & Nan Y. (2018). The lateralized arcuate fasciculus in developmental pitch disorders among mandarin amusics: left for speech and right for music. Brain Structure and Function, 223(4), 2013–2024. 10.1007/s00429-018-1608-2 [DOI] [PubMed] [Google Scholar]
- Chevillet MA, Jiang X, Rauschecker JP, & Riesenhuber M. (2013). Automatic phoneme category selectivity in the dorsal auditory stream. Journal of Neuroscience, 33(12), 5208–5215. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Clayards M, Tanenhaus MK, Aslin RN, & Jacobs RA. (2008). Perception of speech reflects optimal use of probabilistic speech cues. Cognition, 108(3), 804–809. 10.1016/j.cognition.2008.04.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Coffey EBJ, Chepesiuk AMP, Herholz SC, Baillet S, & Zatorre RJ. (2017). Neural Correlates of Early Sound Encoding and their Relationship to Speech-in-Noise Perception [Original Research]. Frontiers in Neuroscience, Volume 11 - 2017. 10.3389/fnins.2017.00479 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dale AM, Fischl B, & Sereno MI. (1999). Cortical surface-based analysis. I. Segmentation and surface reconstruction. Neuroimage, 9(2), 179–194. 10.1006/nimg.1998.0395 [DOI] [PubMed] [Google Scholar]
- de Boer J, & Thornton ARD. (2008). Neural Correlates of Perceptual Learning in the Auditory Brainstem: Efferent Activity Predicts and Reflects Improvement at a Speech-in-Noise Discrimination Task. The Journal of Neuroscience, 28(19), 4929–4937. 10.1523/jneurosci.0902-08.2008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Desai R, Liebenthal E, Waldron E, & Binder JR. (2008). Left Posterior Temporal Regions are Sensitive to Auditory Categorization. Journal of Cognitive Neuroscience, 20(7), 1174–1188. 10.1162/jocn.2008.20081 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Desikan RS, Segonne F, Fischl B, Quinn BT, Dickerson BC, Blacker D, Buckner RL, Dale AM, Maguire RP, Hyman BT, Albert MS, & Killiany RJ. (2006). An automated labeling system for subdividing the human cerebral cortex on MRI scans into gyral based regions of interest. Neuroimage, 31(3), 968–980. 10.1016/j.neuroimage.2006.01.021 [DOI] [PubMed] [Google Scholar]
- Eimas PD. (1963). The Relation between Identification and Discrimination along Speech and Non-Speech Continua. Language and Speech, 6(4), 206–217. 10.1177/002383096300600403 [DOI] [Google Scholar]
- Elmer S, Hänggi J, Meyer M, & Jäncke L. (2013). Increased cortical surface area of the left planum temporale in musicians facilitates the categorization of phonetic and temporal speech sounds. Cortex, 49(10), 2812–2821. 10.1016/j.cortex.2013.03.007 [DOI] [PubMed] [Google Scholar]
- Fischl B. (2012). FreeSurfer. Neuroimage, 62(2), 774–781. 10.1016/j.neuroimage.2012.01.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fischl B, & Dale AM. (2000). Measuring the thickness of the human cerebral cortex from magnetic resonance images. Proceedings of the National Academy of Sciences of the United States of America, 97(20), 11050–11055. 10.1073/pnas.200033797 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fischl B, Sereno MI, & Dale AM. (1999). Cortical surface-based analysis. II: Inflation, flattening, and a surface-based coordinate system. Neuroimage, 9(2), 195–207. 10.1006/nimg.1998.0396 [DOI] [PubMed] [Google Scholar]
- Fischl B, van der Kouwe A, Destrieux C, Halgren E, Segonne F, Salat DH, Busa E, Seidman LJ, Goldstein J, Kennedy D, Caviness V, Makris N, Rosen B, & Dale AM. (2004). Automatically parcellating the human cerebral cortex. Cereberal Cortex, 14(1), 11–22. [DOI] [PubMed] [Google Scholar]
- Fuhrmeister P, & Myers EB. (2021). Structural neural correlates of individual differences in categorical perception. Brain and Language, 215, 104919. 10.1016/j.bandl.2021.104919 [DOI] [PubMed] [Google Scholar]
- Fry DB, Abramson AS, Eimas PD, & Liberman AM. (1962). The Identification and Discrimination of Synthetic Vowels. Language and Speech, 5(4), 171–189. 10.1177/002383096200500401 [DOI] [Google Scholar]
- Gautam P, Anstey KJ, Wen W, Sachdev PS, & Cherbuin N. (2015). Cortical gyrification and its relationships with cortical volume, cortical thickness, and cognitive performance in healthy mid-life adults. Behavioural Brain Research, 287, 331–339. 10.1016/j.bbr.2015.03.018 [DOI] [PubMed] [Google Scholar]
- Golestani N, Price CJ, & Scott SK. (2011). Born with an Ear for Dialects? Structural Plasticity in the Expert Phonetician Brain. The Journal of Neuroscience, 31(11), 4213–4220. 10.1523/jneurosci.3891-10.2011 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Griffiths TD, & Warren JD. (2002). The planum temporale as a computational hub. Trends in Neurosciences, 25(7), 348–353. 10.1016/S0166-2236(02)02191-4 [DOI] [PubMed] [Google Scholar]
- Habibi A, Ilari B, Heine K, & Damasio H. (2020). Changes in auditory cortical thickness following music training in children: converging longitudinal and cross-sectional results. Brain Structure and Function, 225(8), 2463–2474. 10.1007/s00429-020-02135-1 [DOI] [PubMed] [Google Scholar]
- Halwani GF, Loui P, Rueber T, & Schlaug G. (2011). Effects of Practice and Experience on the Arcuate Fasciculus: Comparing Singers, Instrumentalists, and Non-Musicians [Original Research]. Frontiers in Psychology, volume 2 - 2011. 10.3389/fpsyg.2011.00156 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hamilton LS, Oganian Y, Hall J, & Chang EF. (2021). Parallel and distributed encoding of speech across human auditory cortex. Cell, 184(18), 4626–4639.e4613. 10.1016/j.cell.2021.07.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Han X, Jovicich J, Salat D, van der Kouwe A, Quinn B, Czanner S, Busa E, Pacheco J, Albert M, Killiany R, Maguire P, Rosas D, Makris N, Dale A, Dickerson B, & Fischl B. (2006). Reliability of MRI-derived measurements of human cerebral cortical thickness: the effects of field strength, scanner upgrade and manufacturer. Neuroimage, 32(1), 180–194. 10.1016/j.neuroimage.2006.02.051 [DOI] [PubMed] [Google Scholar]
- Hennessy S, Mack WJ, & Habibi A. (2022). Speech-in-noise perception in musicians and non-musicians: A multi-level meta-analysis. Hear Res, 416, 108442. 10.1016/j.heares.2022.108442 [DOI] [PubMed] [Google Scholar]
- Hervais-Adelman A, Moser-Mercer B, Murray MM, & Golestani N. (2017). Cortical thickness increases after simultaneous interpretation training. Neuropsychologia, 98, 212–219. 10.1016/j.neuropsychologia.2017.01.008 [DOI] [PubMed] [Google Scholar]
- Hickok G, & Poeppel D. (2004). Dorsal and ventral streams: a framework for understanding aspects of the functional anatomy of language. Cognition, 92(1), 67–99. 10.1016/j.cognition.2003.10.011 [DOI] [PubMed] [Google Scholar]
- Hickok G, & Poeppel D. (2007). The cortical organization of speech processing. Nature Reviews Neuroscience, 8(5), 393–402. 10.1038/nrn2113 [DOI] [PubMed] [Google Scholar]
- Honda CT, Clayards M, & Baum SR. (2024). Individual differences in the consistency of neural and behavioural responses to speech sounds. Brain Research, 1845, 149208. 10.1016/j.brainres.2024.149208 [DOI] [PubMed] [Google Scholar]
- Humphries C, Sabri M, Lewis K, & Liebenthal E. (2014). Hierarchical organization of speech perception in human auditory cortex. Frontiers in Neuroscience, 8, 406. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Husain FT, Fromm SJ, Pursley RH, Hosey LA, Braun AR, & Horwitz B. (2006a). Neural bases of categorization of simple speech and nonspeech sounds. Human Brain Mapping, 27(8), 636–651. 10.1002/hbm.20207 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Husain FT, McKinney CM, & Horwitz B. (2006b). Frontal cortex functional connectivity changes during sound categorization. NeuroReport, 17(6), 617–621. https://journals.lww.com/neuroreport/fulltext/2006/04240/frontal_cortex_functional_connectivity_changes.12.aspx [DOI] [PubMed] [Google Scholar]
- Jenkinson M, Beckmann CF, Behrens TE, Woolrich MW, & Smith SM. (2012). FSL. Neuroimage, 62(2), 782–790. 10.1016/j.neuroimage.2011.09.015 [DOI] [PubMed] [Google Scholar]
- Joanisse MF, Zevin JD, & McCandliss BD. (2007). Brain Mechanisms Implicated in the Preattentive Categorization of Speech Sounds Revealed Using fMRI and a Short-Interval Habituation Trial Paradigm. Cerebral Cortex, 17(9), 2084–2093. 10.1093/cercor/bhl124 [DOI] [PubMed] [Google Scholar]
- Kapnoula EC, Edwards J, & McMurray B. (2021). Gradient Activation of Speech Categories Facilitiates Listeners’ Recovery From Lexical Graden Paths, But Not Perception of Speech-in-Noise. Human Perception and Performance, 47(4), 578–595. 10.1037/xhp0000900 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kapnoula EC, & McMurray B. (2021). Idiosyncratic use of bottom-up and top-down information leads to differences in speech perception flexibility: Converging evidence from ERPs and eye-tracking. Brain and Language, 223, 105031. 10.1016/j.bandl.2021.105031 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kapnoula EC, Winn MB, Kong EJ, Edwards J, & McMurray B. (2017). Evaluating the Sources and Functions of Gradiency in Phoneme Categorization: An Individual Differences Approach. Human Perception and Performance, 43(9), 1594–1611. 10.1037/xhp0000410 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Killion MC, Niquette PA, Gudmundsen GI, Revit LJ, & Banerjee S. (2004). Development of a quick speech-in-noise test for measuring signal-to-noise ratio loss in normal-hearing and hearing-impaired listeners. J Acoust Soc Am, 116(4 Pt 1), 2395–2405. 10.1121/1.1784440 [DOI] [PubMed] [Google Scholar]
- Kim H, Klein-Packard J, Sorensen E, Oleson J, Tomblin B, & McMurray B. (2025a). Speech categorization consistency is associated with language and reading abilities in school-age children: Implications for language and reading disorders. Cognition, 263, 106194. 10.1016/j.cognition.2025.106194 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kim H, McMurray B, Sorensen E, & Oleson J. (2025b). The consistency of categorization-consistency in speech perception. Psychonomic Bulletin & Review. 10.3758/s13423-025-02700-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kim H, Tomblin JB, & McMurray B. (2024). Speech Categorization Consistency Predicts Overall Language Abilities. PsyArXiv. 10.31234/osf.io/u46pj [DOI] [Google Scholar]
- Kim H, Kim W-J, McMurray B, & Yim D. (2026a). Speech Categorization Consistency Predicts Language and Reading Abilities in Korean School-Age Children. Journal of Speech, Language, and Hearing Research, 69(5), 2046–2066. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kim H, Yang T-H, & Lee J. (2026b). Individual differences in speech categorization and perceptual cue reliance across phonological contrasts. Scientific Reports. 10.1038/s41598-026-58920-1 [DOI] [Google Scholar]
- Kim JH. (2019). Multicollinearity and misleading statistical results. Korean J Anesthesiol, 72(6), 558–569. 10.4097/kja.19087 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kong EJ, & Edwards J. (2016). Individual differences in categorical perception of speech: Cue weighting and executive function. Journal of Phonetics, 59, 40–57. 10.1016/j.wocn.2016.08.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Koops EA, Haykal S, & van Dijk P. (2021). Macrostructural Changes of the Acoustic Radiation in Humans with Hearing Loss and Tinnitus Revealed with Fixel-Based Analysis. The Journal of Neuroscience, 41(18), 3958–3965. 10.1523/jneurosci.2996-20.2021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kutlu E, Baxelbaum K, Sorensen E, Oleson J, & McMurray B. (2024). Linguistic diversity shapes flexible speech perception in school age children. Scientific Reports, 14(1), 28825. 10.1038/s41598-024-80430-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Langer N, Peysakhovich B, Zuk J, Drottar M, Sliva DD, Smith S, Becker BLC, Grant PE, & Gaab N. (2015). White Matter Alterations in Infants at Risk for Developmental Dyslexia. Cerebral Cortex, 27(2), 1027–1036. 10.1093/cercor/bhv281 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lebel C, & Beaulieu C. (2009). Lateralization of the arcuate fasciculus from childhood to adulthood and its relation to cognitive abilities in children. Human Brain Mapping, 30(11), 3563–3573. 10.1002/hbm.20779 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lee Y-S, Turkeltaub P, Granger R, & Raizada RDS. (2012). Categorical Speech Processing in Broca’s Area: An fMRI Study Using Multivariate Pattern-Based Analysis. The Journal of Neuroscience, 32(11), 3942–3948. 10.1523/jneurosci.3814-11.2012 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Levitas D, Hayashi S, Vinci-Booher S, Heinsfeld A, Bhatia D, Lee N, Galassi A, Niso G, & Pestilli F. (2024). ezBIDS: Guided standardization of neuroimaging data interoperable with major data archives and platforms. Scientific Data, 11(1), 179. 10.1038/s41597-024-02959-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li X, Zatorre RJ, & Du Y. (2021). The Microstructural Plasticity of the Arcuate Fasciculus Undergirds Improved Speech in Noise Perception in Musicians. Cerebral Cortex, 31(9), 3975–3985. 10.1093/cercor/bhab063 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu S, & Wang S. (2025). A multimodal neuroimaging dataset for investigating speech perceptual normalization. Scientific Data, 12(1), 1893. 10.1038/s41597-025-06183-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liebenthal E, Binder JR, Spitzer SM, Possing ET, & Medler DA. (2005). Neural Substrates of Phonemic Perception. Cerebral Cortex, 15(10), 1621–1631. 10.1093/cercor/bhi040 [DOI] [PubMed] [Google Scholar]
- Loui P, Alsop D, & Schlaug G. (2009). Tone Deafness: A New Disconnection Syndrome? The Journal of Neuroscience, 29(33), 10215–10220. 10.1523/jneurosci.1701-09.2009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Loui P, Li HC, & Schlaug G. (2011). White matter integrity in right hemisphere predicts pitch-related grammar learning. Neuroimage, 55(2), 500–507. 10.1016/j.neuroimage.2010.12.022 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Luthra S, Guediche S, Blumstein SE, & Myers EB. (2019). Neural substrates of subphonemic variation and lexical competition in spoken word recognition. Language, Cognition and Neuroscience, 34(2), 151–169. 10.1080/23273798.2018.1531140 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Madsen SMK, Marschall M, Dau T, & Oxenham AJ. (2019). Speech perception is similar for musicians and non-musicians across a wide range of conditions. Scientific Reports, 9(1), 10404. 10.1038/s41598-019-46728-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mahmud MS, Yeasin M, & Bidelman GM. (2021). Data-driven machine learning models for decoding speech categorization from evoked brain responses. Journal of Neural Engineering, 18(4), 046012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maillard E, Joyal M, Murray MM, & Tremblay P. (2023). Are musical activities associated with enhanced speech perception in noise in adults? A systematic review and meta-analysis. Curr Res Neurobiol, 4, 100083. 10.1016/j.crneur.2023.100083 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mårtensson J, Eriksson J, Bodammer NC, Lindgren M, Johansson M, Nyberg L, & Lövdén M. (2012). Growth of language-related brain areas after foreign language learning. Neuroimage, 63(1), 240–244. 10.1016/j.neuroimage.2012.06.043 [DOI] [PubMed] [Google Scholar]
- Mazziotta JC, Toga AW, Evans A, Lancaster JL, & Fox PT. (1995). A probabilistic atlas of the human brain: Theory and rationale for its development. Neuroimage, 2, 89–101. [DOI] [PubMed] [Google Scholar]
- McMurray B. (2022). The myth of categorical perception. The Journal of the Acoustical Society of America, 152(6), 3819–3842. 10.1121/10.0016614 [DOI] [PMC free article] [PubMed] [Google Scholar]
- McMurray B, Aslin RN, Tanenhaus MK, Spivey MJ, & Subik D. (2008). Gradient sensitivity to within-category variation in words and syllables. Journal of Experimental Psychology: Human Perception and Performance, 34(6), 1609–1631. 10.1037/a0011747 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mesgarani N, Cheung C, Johnson K, & Chang EF. (2014). Phonetic Feature Encoding in Human Superior Temporal Gyrus. Science, 343(6174), 1006–1010. doi: 10.1126/science.1245994 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Meyers EM, Freedman DJ, Kreiman G, Miller EK, & Poggio T. (2008). Dynamic Population Coding of Category Information in Inferior Temporal and Prefrontal Cortex. Journal of Neurophysiology, 100(3), 1407–1419. 10.1152/jn.90248.2008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Moore E, Schaefer RS, Bastin ME, Roberts N, & Overy K. (2017). Diffusion tensor MRI tractography reveals increased fractional anisotropy (FA) in arcuate fasciculus following music-cued motor training. Brain and Cognition, 116, 40–46. 10.1016/j.bandc.2017.05.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Myers E, Phillips M, & Skoe E. (2024). Individual differences in the perception of phonetic category structure predict speech-in-noise performance. The Journal of the Acoustical Society of America, 156(3), 1707–1719. 10.1121/10.0028583 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Myers EB. (2007). Dissociable effects of phonetic competition and category typicality in a phonetic categorization task: An fMRI investigation. Neuropsychologia, 45(7), 1463–1473. 10.1016/j.neuropsychologia.2006.11.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Myers EB, Blumstein SE, Walsh E, & Eliassen J. (2009). Inferior Frontal Regions Underlie the Perception of Phonetic Category Invariance. Psychological Science, 20(7), 895–903. 10.1111/j.1467-9280.2009.02380.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Myers EB, & Swan K. (2012). Effects of Category Learning on Neural Sensitivity to Non-native Phonetic Categories. Journal of Cognitive Neuroscience, 24(8), 1695–1708. 10.1162/jocn_a_00243 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Oechslin MS, Imfeld A, Loenneker T, Meyer M, & Jäncke L. (2009). The plasticity of the superior longitudinal fasciculus as a function of musical expertise: a diffusion tensor imaging study. Front Hum Neurosci, 3, 76. 10.3389/neuro.09.076.2009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Oldfield RC. (1971). The assessment and analysis of handedness: the Edinburgh inventory. Neuropsychologia, 9(1), 97–113. 10.1016/0028-3932(71)90067-4 [DOI] [PubMed] [Google Scholar]
- Panizzon MS, Fennema-Notestine C, Eyler LT, Jernigan TL, Prom-Wormley E, Neale M, Jacobson K, Lyons MJ, Grant MD, Franz CE, Xian H, Tsuang M, Fischl B, Seidman L, Dale A, & Kremen WS. (2009). Distinct Genetic Influences on Cortical Surface Area and Cortical Thickness. Cerebral Cortex, 19(11), 2728–2735. 10.1093/cercor/bhp026 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Papoutsi M, de Zwart JA, Jansma JM, Pickering MJ, Bednar JA, & Horwitz B. (2009). From phonemes to articulatory codes: an fMRI study of the role of Broca’s area in speech production. Cereb Cortex, 19(9), 2156–2165. 10.1093/cercor/bhn239 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Parbery-Clark A, Skoe E, & Kraus N. (2009a). Musical Experience Limits the Degradative Effects of Background Noise on the Neural Processing of Sound. The Journal of Neuroscience, 29(45), 14100–14107. 10.1523/jneurosci.3256-09.2009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Parbery-Clark A, Skoe E, Lam C, & Kraus N. (2009b). Musician Enhancement for Speech-In-Noise. Ear and Hearing, 30(6), 653–661. 10.1097/AUD.0b013e3181b412e9 [DOI] [PubMed] [Google Scholar]
- Parbery-Clark A, Strait DL, Anderson S, Hittner E, & Kraus N. (2011). Musical Experience and the Aging Auditory System: Implications for Cognitive Abilities and Hearing Speech in Noise. PLoS One, 6(5), e18082. 10.1371/journal.pone.0018082 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Perdue MV, Geeraert BL, Manning KY, Dewey D, & Lebel C. (2025). Phonological decoding ability is associated with fiber density of the left arcuate fasciculus longitudinally across reading development. Developmental Cognitive Neuroscience, 72, 101537. 10.1016/j.dcn.2025.101537 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Perron M, Theaud G, Descoteaux M, & Tremblay P. (2021). The frontotemporal organization of the arcuate fasciculus and its relationship with speech perception in young and older amateur singers and non-singers. Human Brain Mapping, 42(10), 3058–3076. 10.1002/hbm.25416 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Piccolo LR, Merz EC, He X, Sowell ER, Noble KG, & Pediatric Imaging NGS. (2016). Age-Related Differences in Cortical Thickness Vary by Socioeconomic Status. PLoS One, 11(9), e0162511. 10.1371/journal.pone.0162511 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pisoni DB. (1973). Auditory and phonetic memory codes in the discrimination of consonants and vowels. Perception & Psychophysics, 13(2), 253–260. 10.3758/BF03214136 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pisoni DB, & Tash J. (1974). Reaction times to comparisons within and across phonetic categories. Perception & Psychophysics, 15(2), 285–290. 10.3758/BF03213946 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Poeppel D. (2003). The analysis of speech in different temporal integration windows: cerebral lateralization as ‘asymmetric sampling in time’. Speech Communication, 41(1), 245–255. 10.1016/S0167-6393(02)00107-3 [DOI] [Google Scholar]
- Price CN, & Bidelman GM. (2021). Attention reinforces human corticofugal system to aid speech perception in noise. Neuroimage, 235, 118014. 10.1016/j.neuroimage.2021.118014 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rakic P. (1988). Specification of cerebral cortical areas. Science, 241(4862), 170–176. 10.1126/science.3291116 [DOI] [PubMed] [Google Scholar]
- Rauschecker JP. (2012). Ventral and dorsal streams in the evolution of speech and language [Perspective]. Frontiers in Evolutionary Neuroscience, Volume 4 - 2012. 10.3389/fnevo.2012.00007 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reuter M, Rosas HD, & Fischl B. (2010). Highly accurate inverse consistent registration: a robust approach. Neuroimage, 53(4), 1181–1196. 10.1016/j.neuroimage.2010.07.020 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reuter M, Schmansky NJ, Rosas HD, & Fischl B. (2012). Within-subject template estimation for unbiased longitudinal image analysis. Neuroimage, 61(4), 1402–1418. 10.1016/j.neuroimage.2012.02.084 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rizzi R, & Bidelman GM. (2023). Duplex perception reveals brainstem auditory representations are modulated by listeners’ ongoing percept for speech. Cerebral Cortex, 33(18), 10076–10086. 10.1093/cercor/bhad266 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rizzi R, & Bidelman GM. (2024). Functional benefits of continuous vs. categorical listening strategies on the neural encoding and perception of noise-degraded speech. Brain Research, 1844, 149166. 10.1016/j.brainres.2024.149166 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rizzi R, & Bidelman GM. (2025). Consistency in phonetic categorization predicts successful speech-in-noise perception. JASA Express Letters, 5(12). 10.1121/10.0041846 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rizzi R, Stirn JR, Eisenhut Z, & Bidelman GM. (2026). Perceptual consistency in phoneme categorization is driven by neural consistency and predicts improved speech-in-noise performance [preprint]. bioRxiv. 10.64898/2026.07.02.736174 [DOI] [Google Scholar]
- Rogers JC, & Davis MH. (2017). Inferior Frontal Cortex Contributions to the Recognition of Spoken Words and Their Constituent Speech Sounds. Journal of Cognitive Neuroscience, 29(5), 919–936. 10.1162/jocn_a_01096 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rogge A-K, Röder B, Zech A, & Hötting K. (2018). Exercise-induced neuroplasticity: Balance training increases cortical thickness in visual and vestibular cortical regions. Neuroimage, 179, 471–479. 10.1016/j.neuroimage.2018.06.065 [DOI] [PubMed] [Google Scholar]
- Ruggles DR, Freyman RL, & Oxenham AJ. (2014). Influence of Musical Training on Understanding Voiced and Whispered Speech in Noise. PLoS One, 9(1), e86980. 10.1371/journal.pone.0086980 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schütt HH, Harmeling S, Macke JH, & Wichmann FA. (2016). Painfree and accurate Bayesian estimation of psychometric functions for (potentially) overdispersed data. Vision Research, 122, 105–123. 10.1016/j.visres.2016.02.002 [DOI] [PubMed] [Google Scholar]
- Shen C-Y, Tyan Y-S, Kuo L-W, Wu CW, & Weng J-C. (2015). Quantitative Evaluation of Rabbit Brain Injury after Cerebral Hemisphere Radiation Exposure Using Generalized q-Sampling Imaging. PLoS One, 10(7), e0133001. 10.1371/journal.pone.0133001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sininger YS, & Bhatara A. (2012). Laterality of basic auditory perception. Laterality, 17(2), 129–149. 10.1080/1357650x.2010.541464 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sininger YS, & Cone-Wesson B. (2004). Asymmetric cochlear processing mimics hemispheric specialization. Science, 305(5690), 1581. 10.1126/science.1100646 [DOI] [PubMed] [Google Scholar]
- Sitek KR, Gulban OF, Calabrese E, Johnson GA, Lage-Castellanos A, Moerel M, Ghosh SS, & De Martino F. (2019). Mapping the human subcortical auditory system using histology, postmortem MRI and in vivo MRI at 7T. Elife, 8. 10.7554/eLife.48932 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Slater J, Skoe E, Strait DL, O’Connell S, Thompson E, & Kraus N. (2015). Music training improves speech-in-noise perception: Longitudinal evidence from a community-based music program. Behavioural Brain Research, 291, 244–252. 10.1016/j.bbr.2015.05.026 [DOI] [PubMed] [Google Scholar]
- Sorensen E, Oleson J, Kutlu E, & McMurray B. (2024). A Bayesian hierarchical model for the analysis of visual analogue scaling tasks. Statistical Methods in Medical Research, 33(6), 953–965. 10.1177/09622802241242319 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Svobodová V, Profant O, Škoch A, Tintěra J, Tóthová D, Chovanec M, Čapková D, & Syka J. (2024). The effect of aging, hearing loss, and tinnitus on white matter in the human auditory system revealed with fixel-based analysis [Original Research]. Frontiers in Aging Neuroscience, Volume 15 - 2023. 10.3389/fnagi.2023.1283660 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Toscano JC, Anderson ND, Fabiani M, Gratton G, & Garnsey SM. (2018). The time-course of cortical responses to speech revealed by fast optical imaging. Brain and Language, 184, 32–42. 10.1016/j.bandl.2018.06.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tremblay P, Perron M, Deschamps I, Kennedy-Higgins D, Houde J-C, Dick AS, & Descoteaux M. (2019). The role of the arcuate and middle longitudinal fasciculi in speech perception in noise in adulthood. Human Brain Mapping, 40(1), 226–241. 10.1002/hbm.24367 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vandermosten M, Boets B, Poelmans H, Sunaert S, Wouters J, & Ghesquière P. (2012). A tractography study in dyslexia: neuroanatomic correlates of orthographic, phonological and speech processing. Brain, 135(3), 935–948. 10.1093/brain/awr363 [DOI] [PubMed] [Google Scholar]
- Viswanathan N, & Kelty-Stephen DG. (2018). Comparing speech and nonspeech context effects across timescales in coarticulatory contexts. Attention, Perception, & Psychophysics, 80(2), 316–324. 10.3758/s13414-017-1449-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wernicke C. (1874). Der aphasische Symptomenkomplex. In Wernicke C. (Ed.), Der aphasische Symptomencomplex: Eine psychologische Studie auf anatomischer Basis (pp. 1–70). Springer; Berlin Heidelberg. 10.1007/978-3-642-65950-8_1 [DOI] [Google Scholar]
- Wierenga LM, Langen M, Oranje B, & Durston S. (2014). Unique developmental trajectories of cortical thickness and surface area. Neuroimage, 87, 120–126. 10.1016/j.neuroimage.2013.11.010 [DOI] [PubMed] [Google Scholar]
- Williams VJ, Juranek J, Cirino P, & Fletcher JM. (2017). Cortical Thickness and Local Gyrification in Children with Developmental Dyslexia. Cerebral Cortex, 28(3), 963–973. 10.1093/cercor/bhx001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Winkler AM, Kochunov P, Blangero J, Almasy L, Zilles K, Fox PT, Duggirala R, & Glahn DC. (2010). Cortical thickness or grey matter volume? The importance of selecting the phenotype for imaging genetics studies. Neuroimage, 53(3), 1135–1146. 10.1016/j.neuroimage.2009.12.028 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wolmetz M, Poeppel D, & Rapp B. (2011). What does the right hemisphere know about phoneme categories? J Cogn Neurosci, 23(3), 552–569. 10.1162/jocn.2010.21495 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wong BWL, Samuel AG, & Kapnoula EC. (2026). The role of speech perception gradiency in L1 versus L2 spoken-word recognition. Journal of Experimental Psychology: Human Perception and Performance, 52(5), 553–572. 10.1037/xhp0001399 [DOI] [PubMed] [Google Scholar]
- Yeatman JD, Dougherty RF, Rykhlevskaia E, Sherbondy AJ, Deutsch GK, Wandell BA, & Ben-Shachar M. (2011). Anatomical Properties of the Arcuate Fasciculus Predict Phonological and Reading Skills in Children. Journal of Cognitive Neuroscience, 23(11), 3304–3317. 10.1162/jocn_a_00061 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yeend I, Beach EF, Sharma M, & Dillon H. (2017). The effects of noise exposure and musical training on suprathreshold auditory processing and speech perception in noise. Hearing Research, 353, 224–236. 10.1016/j.heares.2017.07.006 [DOI] [PubMed] [Google Scholar]
- Yeh FC, Badre D, & Verstynen T. (2016). Connectometry: A statistical approach harnessing the analytical potential of the local connectome. Neuroimage, 125, 162–171. 10.1016/j.neuroimage.2015.10.053 [DOI] [PubMed] [Google Scholar]
- Yeh FC, Liu L, Hitchens TK, & Wu YL. (2017). Mapping immune cell infiltration using restricted diffusion MRI. Magnetic Resonance in Medicine, 77(2), 603–612. 10.1002/mrm.26143 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yeh FC, Panesar S, Fernandes D, Meola A, Yoshino M, Fernandez-Miranda JC, Vettel JM, & Verstynen T. (2018). Population-averaged atlas of the macroscale human structural connectome and its network topology. Neuroimage, 178, 57–68. 10.1016/j.neuroimage.2018.05.027 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yeh FC, & Tseng WY. (2011). NTU-90: a high angular resolution brain atlas constructed by q-space diffeomorphic reconstruction. Neuroimage, 58(1), 91–99. 10.1016/j.neuroimage.2011.06.021 [DOI] [PubMed] [Google Scholar]
- Yeh FC, Verstynen TD, Wang Y, Fernández-Miranda JC, & Tseng WY. (2013). Deterministic diffusion fiber tracking improved by quantitative anisotropy. PLoS One, 8(11), e80713. 10.1371/journal.pone.0080713 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yeh FC, Wedeen VJ, & Tseng WY. (2010). Generalized q-sampling imaging. IEEE Transactions on Medical Imaging, 29(9), 1626–1635. 10.1109/tmi.2010.2045126 [DOI] [PubMed] [Google Scholar]
- Yoo J, & Bidelman GM. (2019). Linguistic, perceptual, and cognitive factors underlying musicians’ benefits in noise-degraded speech perception. Hearing Research, 377, 189–195. 10.1016/j.heares.2019.03.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zaehle T, Geiser E, Alter K, Jancke L, & Meyer M. (2008). Segmental processing in the human auditory dorsal stream. Brain Research, 1220, 179–190. 10.1016/j.brainres.2007.11.013 [DOI] [PubMed] [Google Scholar]
- Zatorre RJ, & Belin P. (2001). Spectral and temporal processing in human auditory cortex. Cereb Cortex, 11(10), 946–953. 10.1093/cercor/11.10.946 [DOI] [PubMed] [Google Scholar]
- Zatorre RJ, Evans AC, Meyer E, & Gjedde A. (1992). Lateralization of phonetic and pitch discrimination in speech processing. Science, 256(5058), 846–849. 10.1126/science.1589767 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets generated during the current study are available from the corresponding author on reasonable request.
