Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2012 Aug 13;109(35):14265–14270. doi: 10.1073/pnas.1200632109

Sound-sized segments are significant for Mandarin speakers

Qingqing Qu 1,1, Markus F Damian 1, Nina Kazanina 1
PMCID: PMC3435182  PMID: 22891321

Abstract

Do speakers of all languages use segmental speech sounds when they produce words? Existing models of language production generally assume a mental representation of individual segmental units, or phonemes, but the bulk of evidence comes from speakers of European languages in which the orthographic system codes explicitly for speech sounds. By contrast, in languages with nonalphabetical scripts, such as Mandarin Chinese, individual speech sounds are not orthographically represented, raising the possibility that speakers of these languages do not use phonemes as fundamental processing units. We used event-related potentials (ERPs) combined with behavioral measurement to investigate the role of phonemes in Mandarin production. Mandarin native speakers named colored line drawings of objects using color adjective-noun phrases; color and object name either shared the initial phoneme or were phonologically unrelated. Whereas naming latencies were unaffected by phoneme repetition, ERP responses were modulated from 200 ms after picture onset. Our ERP findings thus provide strong support for the claim that phonemic segments constitute fundamental units of phonological encoding even for speakers of languages that do not encode such units orthographically.

Keywords: speech production, phonological planning, phoneme priming


Language production involves successive planning stages. A communicative message is formed at the initial stage of conceptual preparation, followed by a stage of grammatical and lexical selection during which lexical items are retrieved from the mental lexicon and slotted into a grammatical structure. The focus of the current study is a subsequent stage in the speech planning process, the phonological encoding stage, during which a selected lexical item is given its sound form, which then enables its articulation. Over the past decades, many studies have investigated this stage with the aim to identify functional units that underlie word-form encoding, (i.e., units that play a fundamental and necessary role in phonological encoding of lexical items). Existing theories (15) differ considerably on details but converge on the assumption that phonological encoding involves selection and sequencing of abstract sound-sized segmental units or phonemes, which have been postulated as units for encoding words in long-term memory (6). Furthermore, most frameworks assume syllable-sized production units, either as abstract structural frames that are filled with segmental content (7, 8) or as a mental syllabary that prespecifies articulatory “gestural scores” (9, 10).

Existing frameworks are largely based on Indo-European languages, and despite some universalities, it cannot be assumed that all aspects generalize across all world languages. Indeed, in languages with alphabetical scripts, such as English, the phoneme is a prominent (perhaps the most prominent) processing unit. By contrast, recent studies that have investigated phonological encoding in Mandarin Chinese, a language with a nonalphabetical script, claim that the most prominent unit of spoken word production is a syllable, whereas phonological segments are less relevant or perhaps not relevant at all (11, 12). These studies leave open the question whether an independent role of the phoneme can be demonstrated for Chinese speakers. In the present study, we investigated this issue with electrophysiological measures.

Evidence for a critical role of phonemes in spoken production in European languages comes from speech error analyses, which indicate that the majority of phonological errors involve the insertion, deletion, substitution, or exchange of segments (2, 1315) (e.g., York library→lork yibrary, reading list→leading list). By contrast, errors involving whole syllables (e.g., napkin→kinnap) or single phonological features (e.g., blue→plue) are rather rare (1517). The fact that “holistic” phonemes are critically involved in speech errors highlights the psychological reality of segment-sized mental representations in language production. Support for the phoneme as a production unit also comes from a number of experimental studies, most prominently studies that used the “implicit priming” technique. Speakers are asked to produce a small number of target words repeatedly based on visually presented associated cues. Within an experimental block, all target responses either share part of their phonological form (e.g., cat, cup, corn; “homogeneous” condition) or they do not (e.g., cat, door, pipe; “heterogeneous” condition). A basic finding is that responses are faster in homogeneous blocks than in heterogeneous blocks when there is segmental form overlap in word-initial position (1820 for Dutch, French, and English, respectively). Importantly, overlap that involves very similar featural composition but a different phonological segment (e.g., /b/ and /p/, identical except for the voicing feature) generated no priming (21), which underscores the importance of segments as central planning units of spoken word production. More recent evidence comes from a picture-naming task in which English speakers named colored line drawings of simple objects with adjective-noun phrases (22, 23). When colors and objects were chosen such that the initial phoneme of both constituents matched (e.g., green goat, red rug), naming latencies were faster than when the same colors and objects were combined without overlap (e.g., red goat, green rug). The phoneme repetition effect was found even when the overlapping phoneme did not occupy the same position within the word (e.g., green flag) and has been taken to reflect the coactivation of phoneme-sized segments. Because facilitation is found despite considerable acoustic variations (i.e., the phoneme /g/ in green and flag has a very distinct acoustic realization), the finding supports involvement of abstract segmental representations in speech production.

The studies discussed above thus provide strong evidence for segmental representations as central units of spoken production in Western European languages that use alphabetical scripts. It is less clear whether sound-sized segments play a role in production for nonalphabetical languages, such as Mandarin Chinese. Chinese has a logographic writing script in which each orthographically distinct unit (character) maps onto a syllable; importantly, there are no orthographic representations that specifically correspond to individual phonemes. The prominence of syllables is also apparent in spoken language. Contrary to most alphabetical languages, Chinese has relatively few syllable types (∼400 not counting tone and 1,200 counting tone), extremely clear syllable boundaries, and no resyllabification (11). This cross-linguistic variation in orthography and phonology might entail fundamental differences in which phonological units are realized in speech production across languages.

The idea that syllables rather than phonemes are the primary unit of phonological encoding in Chinese has been supported by several studies. For example, in contrast to what has been found in alphabetical languages, a significant proportion of phonological speech errors in Chinese are syllable-sized, whereas segmental errors are quite rare (24, 25). Furthermore, benefits of word-initial phoneme overlap that have been reliably documented in European languages using implicit priming (see above) have not been found in similar tasks with Chinese speakers, even though such benefits did emerge with a word-initial syllable overlap (11, 12). To account for these cross-linguistic differences in priming, O’Seaghdha et al. (12) proposed a proximate unit principle according to which the first selectable phonological units below the level of the word or morpheme (so-called “proximate units”) vary across languages, and planning and execution of phonological encoding are crucially dependent on them. Proximate units in European languages are phonemes; hence, segment-based priming is found in English and Dutch, for example. Proximate units in Mandarin Chinese are syllables; hence, priming is restricted to syllabic overlap.

An arising issue is whether Chinese speakers mentally represent subsyllabic units, specifically whether a role of segment-sized representations can be identified. O’Seaghdha et al. tend toward a limited role for phonemes: “Mandarin speakers naturally intend to produce syllables, perhaps to the exclusion of subsyllabic ingredients” (ref. 12, p. 285). Indeed, evidence for psychological validity of subsyllabic units, particularly phonemes, in nonalphabetical languages is scant and inconclusive. Chinese speakers who lack experience with an alphabetical script have difficulty in manipulating speech sounds in phonological awareness tasks (e.g., repeat a spoken word with the initial phoneme deleted), whereas Chinese speakers who have been exposed to an alphabetical system can perform such tasks more readily (26). This finding may be taken to conclude that phonemes are artifacts resulting solely from experience with an alphabetically organized orthographic system. However, such a conclusion needs to be taken with caution, because metalinguistic phonological awareness tasks readily might be solved using orthography-based strategies; hence, they may not index mental representations that are relevant in naturalistic speech processing (27). In a picture-word interference task with Cantonese speakers, Wong and Chen (28, 29) manipulated form overlap between picture and word (e.g., onset, rhyme, tone). Critically, and in contrast to what is typically found with speakers of European languages, they found no facilitation from segmental overlap for Chinese speakers and concluded that individual segments are unimportant in spoken word preparation. However, it is not clear whether single-segment overlap results in significant facilitation even in languages with alphabetical script, because in parallel studies in European languages, form-related picture-word pairs typically overlap in more than a single segment (30, 31).

In the current study, we used event-related potentials (ERPs) to examine whether phonemes are a representational unit in Chinese spoken production. ERPs reflect brain activity that is time-locked to an external stimulus and, because of their millisecond resolution, make it possible to explore the detailed course of a cognitive process as it unfolds in time. To date, few studies have utilized ERPs to explore spoken production using overt naming tasks (e.g., picture naming), largely attributable to the fact that ERPs are susceptible to articulation-related artifacts. However, several recent studies have shown that brain activity can be reliably measured at least until 400 ms after picture onset (3237, reviewed in 38). The present study targeted the phonological encoding stage, which is estimated to take place 275–400 ms after picture onset (39). For such relatively early stages of speech preparation, ERPs are relatively artifact-free and provide a window into stages of processing that are not easily detectable with behavioral measures. Our study used a picture-naming task in which participants name a colored object with an adjective-noun phrase (e.g., red train). With English speakers, a repeated initial phoneme between the color and object names (e.g., green goat) accelerates naming onset latencies, and this effect is attributed to an abstract segmental processing level (22, 23). Chinese speakers named colored objects, and word-initial segmental overlap between the color and object name was manipulated (Fig. 1A). If segmental overlap affects ERPs or naming latencies, it would suggest that phonemes constitute an important processing unit in Chinese spoken production.

Fig. 1.

Fig. 1.

(A) Colored object displays in the two experimental conditions were named by 27 Chinese speakers using a color adjective-noun phrase. In the phonologically related condition, color and object names had an identical initial phoneme in Chinese. In the phonologically unrelated condition, initial phonemes were different. Numbers indicate tone for each; a neutral tone is not indicated. (B) Behavioral data show no effect of phonological relatedness on naming latencies (gray bars, left axis) or error rates (Inline graphic, right axis). Error bars represent 95% confidence intervals. (C) Grand average ERPs from the same 27 native Chinese speakers for the phonologically related (black line) and phonologically unrelated (gray line) conditions at six ROIs (shown as red filled circles on the electrode layout on the left): left-anterior (electrodes F7, F5, and FC5), midanterior (electrodes Fz, FCz, and Cz), right-anterior (electrodes F6, F8, and FC6), left-posterior (electrodes P7, P5, and CP5), midposterior (electrodes CPz, Pz, and POz), and right-posterior (electrodes CP6, P8, and P6). The onset of a picture is represented by 0 ms. In the posterior regions, the phonologically unrelated condition was significantly more negative in the 200- to 300-ms time interval after picture onset (blue shading), indicating that an initial phoneme overlap affects early speech preparation in Chinese speakers. An opposite effect of relatedness was found in the 300- to 400-ms window in the anterior regions (purple shading).

Results

Naming Latencies and Error Rates.

Behavioral data are summarized in Fig. 1B. Mean latencies were virtually identical in the two conditions (phonologically related: 977 ms, phonologically unrelated: 976 ms; F1 and F2 < 1), and so were error rates (phonologically related: 7.0%, phonologically unrelated: 7.0%; F1 and F2 < 1).

ERPs.

Grand average ERP waveforms are displayed in Fig. 1C for six regions of interest (ROIs) chosen for the analysis (Methods). The statistical results for eight 100-ms time intervals are summarized in Table S1. No main effects or interactions involving relatedness were found in the baseline intervals (−200 to −100 ms and −100 to 0 ms; F < 1). In the 0- to 100-ms interval or in the 100-to 200-ms interval, the only (marginally) significant effect involving the factor relatedness was the relatedness × anteriority interaction [0–100 ms: F(1,26) = 2.94, P < 0.098; 100–200 ms: F(1,26) = 9.36, P = 0.005]. Pairwise comparisons revealed no significant effect of relatedness, however, when anterior and posterior ROIs were considered separately (P > 0.153) or at any ROI (P > 0.110).

In the 200- to 300-ms time window, the relatedness × anteriority interaction was significant [F(1,26) = 21.34, P < 0.001] and the main effects of relatedness and other interactions involving relatedness were not significant (F < 1.33). Follow-up pairwise comparisons demonstrated that the effect of relatedness was significant in the posterior regions (P = 0.034), as a result of more positive ERPs for the phonologically related condition relative to the phonologically unrelated condition, but not significant in the anterior regions (P > 0.525). Planned pairwise comparisons at each ROI revealed a significant effect of relatedness in the midposterior region (P = 0.018) and a marginally significant effect in the right posterior region (P = 0.078). To explore this effect in more detail, a larger set of 7 middle posterior electrodes (CPz, Pz, POz, Oz, P1, P2, and PO3) was analyzed. Relatedness had a significant effect [F(1,26) = 7.15, P = 0.013] and did not interact with electrodes (F < 1). Relatedness had a significant effect when an even larger set of 26 posterior electrodes (TP7, CP5, CP3, CP1, CPz, CP2, CP4, CP6, TP8, P7, P5, P3, P1, Pz, P2, P4, P6, P8, PO7, PO3, POz, PO4, PO8, O1, Oz, and O2) was selected [relatedness: F(1,26) = 5.29, P = 0.030; relatedness × electrode; F(25,650) = 2.04, P = 0.085].

In the 300- to 400-ms time window, the relatedness × anteriority interaction was significant [F(1,26) = 10.48, P = 0.003] and the relatedness × anteriority × laterality interaction was marginally significant [F(2,52) = 3.02, P = 0.064], but the main effects of relatedness and the relatedness × laterality interaction were not (F < 1). Follow-up pairwise comparisons demonstrated that the effect of relatedness was significant in the anterior regions (P = 0.047), as a result of more negative ERPs for the phonologically related condition relative to the phonologically unrelated condition, but not significant in the posterior regions (P > 0.678). Analyses per ROI revealed a significant effect of relatedness at the right anterior sites (P = 0.037) and a marginally significant effect of relatedness in the midanterior region (P = 0.065). This pattern was confirmed via additional analyses on a larger set of 9 anterior electrodes (FT7, FC5, FC3, FC1, FCz, FC2, FC4, FC6, and FT8): Relatedness had a significant effect [(F(1,26) = 4.91, P = 0.036] and did not interact with electrodes (F < 1). A similar pattern was obtained when an even larger set of 27 anterior electrodes (F7, F5, F3, F1, Fz, F2, F4, F6, F8, FT7, FC5, FC3, FC1, FCz, FC2, FC4, FC6, FT8, T7, C5, C3, C1, Cz, C2, C4, C6, and T8) was analyzed [relatedness: F (1,26) = 3.53, P = 0.071; relatedness × electrode: F < 1]. Hence, in the 300- to 400-ms interval, phonological overlap in a single initial phoneme yielded significantly more negative ERPs across anterior electrodes.

In the 400- to 500-ms and 500- to 600-ms intervals, the only significant effect involving the factor relatedness was the relatedness × anteriority interaction [400–500 ms: F(1,26) = 4.98, P = 0.035; 500–600 ms: F(1,26) = 8.60, P = 0.007]. However, no significant effect of relatedness was found when anterior and posterior channels were analyzed separately (P > 0.219) or at any ROI (P > 0.380).

Discussion

The present study measured behavioral responses (naming onset latencies and error rates) and ERPs from Chinese speakers to elucidate whether the phoneme is a functional unit in Chinese spoken production, by manipulating word-initial phonemic overlap in color adjective-noun phrases. Behavioral responses showed no effect of phoneme overlap and are consistent with previous findings from different tasks in which Chinese speakers showed no phoneme-based priming, in contrast to speakers of European languages, for whom robust segmental priming was observed. By contrast, our ERP data provide clear evidence that phonemes are an important processing unit for Chinese speakers: Phoneme repetition modulated ERPs from ∼200 ms after picture onset. More specifically, phoneme repetition elicited significantly more positive ERPs in the posterior regions 200–300 ms after picture onset and more negative ERPs in the anterior regions 300–400 ms after picture onset, relative to no-repetition trials. Thus, at the broadest level, although syllables may be proximate units of spoken production for Chinese speakers (12), phonological segments nevertheless have a robust impact on speech preparation.

The interval of 200–400 ms that hosts ERP effects in our study is in agreement with an estimated time course of phonological encoding and internal monitoring by Indefrey and Levelt (39) and with previous ERP studies on overt production. For instance, Strijkers et al. (37) observed that manipulation of lexical frequency of objects in a picture-naming task, commonly assumed to reside at the phonological form level, yielded ERP differences ∼180 ms after picture presentation. Eulitz et al. (33) explored the time course of phonological encoding by comparing overt picture naming with passive viewing of the same pictures and found differential ERPs in the 275- to 400-ms interval. In a picture-word interference task, facilitation effect from phonological relatedness occurred in a similar time frame, between 250 and 400 ms (32).

In the present study, in the 200- to 300-ms time window, we found more positive ERPs in the phonologically related condition over posterior electrode sites. Because of the absence of previous ERP studies on language production that used a colored picture-naming task, we are unable to assess this finding against a closely matched counterpart. However, it is fully in line with the results from the delayed picture-naming task of Jescheniak et al. (34) in which German participants heard an auditory word while preparing a picture name. The word was either phonologically related to the picture name (i.e., shared the initial consonant and vowel, as in Schema-Schere) or phonologically unrelated; the ERPs were more positive in the phonologically related condition broadly across posterior regions. In the present study, we interpret the posterior ERP effect in the 200- to 300-ms window as reflecting facilitation from phoneme repetition arising during phonological encoding proper. Segments that are repeatedly retrieved in the planning of a phrase may be easier to retrieve; alternatively, two word forms may prime each other via shared phonemes at the segmental level (23). Although our findings cannot distinguish between these processing mechanisms, it is important to note that both mechanisms capitalize on involvement from segmental representations during the phonological encoding stage.

We propose that more negative ERP amplitudes in the phonologically related condition in the 300- to 400-ms interval reflect internal speech monitoring for phonological appropriateness/correctness. Self-monitoring is an essential property of speaking and involves an internal loop that monitors abstract phonological codes and an external loop that monitors self-generated speech output (39). Internal self-monitoring operates on a syllabified phonological representation (40), which is estimated to be available around 355 ms poststimulus for a single word naming (39); this time estimate is therefore consistent with the time frame of 300–400 ms observed in our study. A proposed underlying mechanism for self-monitoring is that during planning an utterance production, speakers verify whether the appropriate phonemes have been retrieved and sequenced correctly. Repeated phonemes create a tendency for (near-)adjacent speech sounds to be misordered (e.g., left hemisphere is prone to a slip, such as heft lemisphere, because “left” and “hemisphere” share the vowel “e”) (41). Hence, when phonemes are repeated within a portion of an utterance, as in the adjective-noun phrase in the phonologically related condition in this study, the monitoring system is arguably under higher load to prevent speech errors compared with when there is no phoneme repetition, leading to more negative ERPs. Hence, a facilitatory effect of phoneme repetition during phonological encoding in the 200- to 300-ms interval is reversed during subsequent self-monitoring in the later 300- to 400-ms interval. A similar dissociation between early and later time windows has been reported in previous studies on overt language production (42, 43), and in line with the argument presented here, later effects are typically attributed to self-monitoring.

The present findings provide important insights into the phonological representations of speakers of nonalphabetical languages, such as Mandarin Chinese. The key finding that phoneme repetition affected Chinese speakers’ ERPs in the 200- to 400-ms time interval constitutes clear evidence for the role of phonemes as processing units in Chinese spoken production. As outlined previously, Chinese characters map onto syllables, whereas there are no graphical representations that map specifically onto individual phonemes. This has led to recent claims by O’Seaghdha et al. (12) that syllabic representations are of central importance at the phonological encoding stage in Chinese, whereas segmental representations might be less or not at all relevant. Similarly, in theoretical phonology, unambiguous evidence in support of phonemic units in Mandarin Chinese is surprisingly difficult to find, which led to claims that such units are superfluous (44). A rare piece of phonological evidence comes from [k] ∼ [tɕ] alternation in Mandarin, with [tɕ] appearing only before [i] and [j] and [k] never appears before [i]/[j]. (45). Our experimental results highlight the importance that phonological segments have even for Chinese speakers’ phonological encoding. An apparent benefit of using phonemic units that seems to hold universally across languages is that they enable an efficient representation of thousands of lexical forms in the mental lexicon via no more than several dozens of phonemes (whereas at least several hundred syllables would be needed). Nevertheless, the conclusion that phonemes are nonsuperfluous in Mandarin Chinese does not contradict the assumption that syllables may be particularly important planning units in Chinese (more so than for speakers of European languages). In that respect, we broadly concur with the view of O’Seaghdha et al. (12) that in Chinese, syllables are primary selectable phonological units below the level of word and phonemic segments are retrieved via mediation of syllable. Once a syllable frame is accessed, the segments are retrieved and linked to relevant positions inside the syllable frame. In other words, Mandarin and English speakers are likely to have different proximate units for speech production (syllable vs. phoneme). Furthermore, our statement that phonemes play a functional role for phonological encoding in Chinese should not be overinterpreted to suggest that phonemes have an identical effect at each stage of speech production in Chinese vs. English. The scope of our findings is that they reveal phonemes to be operational at the phonological encoding stage in Chinese (in contrast to theoretical and experimental claims suggesting otherwise).

It is also essential to highlight that our results argue specifically for a functional role of phonemes at the stage of phonological encoding in Mandarin Chinese and should not be attributed to a later motor/articulatory stage of production. Phoneme repetition influenced neural responses of Mandarin speakers as early as 200 ms after picture onset, and the early latency of the effect makes it highly unlikely that it emerged as a result of repetition of the same articulatory gestures in the phonologically related condition. There are also independent reasons to believe that the effects of phoneme repetition in implicit priming or in word-naming tasks must originate at the phonological encoding stage. O’Seaghdha et al. (12) point out that effects resulting from repetition of motor/articulatory gestures should hold universally across languages. However, phoneme repetition results in decreased naming latencies in Western languages but not in Mandarin Chinese, thus undermining an explanation that is based on motor considerations. Further support comes from Damian and Dumay (23), who reported a delayed colored picture-naming task in which English participants were instructed to prepare their response fully but to begin articulation only after a cue appeared. This procedure is typically taken to allow speakers to complete all cognitive processes involved in identification and phonological encoding of the target before the onset of motor execution (46, 47). In this study, the typical facilitation effect attributable to phoneme repetition between color and object name (as in “green goat”) was eliminated, which confirms that the effect visible in nondelayed picture naming is not grounded in articulation and must be attributed to earlier stages of speech production, such as phonological encoding.

As far as methodology is concerned, in the present study, we found that phoneme repetition did not affect participants’ behavioral responses yet modulated their ERPs. Consideration of the behavioral findings only would have led to an erroneous conclusion that phonemes are not relevant functional units in Chinese spoken production. This highlights the advantage of combining behavioral measurement with temporally fine-grained neuropsychological measurement. Our study thus adds to a growing body of previous research on different aspects of language processing showing that because of its fine-grained temporal resolution, electrophysiology may provide insight into distinct (usually earlier) stages of cognitive processing than behavioral measurement (48, 49).

What could be the reason for the absence of a behavioral phoneme repetition effect in Chinese speakers, in contrast to a robust effect previously found in English speakers (22, 23)? Differences in experimental design could be one potential reason. Although the current study was designed to be as similar as possible to the English version, some procedural differences remained [e.g., 8 colors were used in this study, but only 3 or 4 were used by Damian and Dumay (22, 23)]. To explore whether such procedural differences are relevant, we conducted an additional behavioral experiment with English materials and participants that was as similar as possible to the Chinese version and obtained a significant 45-ms priming effect of phoneme repetition (SI Text). Hence, contrasting behavioral findings in English vs. Chinese are not a mere artifact of procedural differences.

Instead, whether or not phoneme overlap generates behaviorally measurable priming is a function of language and its properties. Earlier, while interpreting our ERP findings, we discussed two stages of speech planning: phonological encoding and self-monitoring. Whereas we have no prior reasons to believe that the latter stage varies cross-linguistically, our opinion is that the former stage is not uniform across languages as a result of differences in the hierarchy of proximate units among languages. With phonemes being the most prominent processing units in Western languages, it is plausible that phoneme-based facilitation is so strong and pervasive that it outweighs inhibitory effects from phoneme repetition during self-monitoring, resulting in a net facilitative effect in response latencies. In Chinese, on the other hand, phoneme-based facilitation during phonological encoding is comparatively weak and balanced out by self-monitoring, resulting in a null behavioral effect. This account is rather speculative but generates clear predictions concerning a parallel version of our study conducted in English speakers: The early, positive component of the relatedness effect should be substantially more pronounced than in the current Chinese version, but the late negative component should be comparable. [Note that our central claim that phonological segments are functional units during Chinese phonological encoding remains valid whether or not this prediction is borne out.]

Finally, it is important to note that participants in our study had acquired Chinese orthography via use of pinyin, a Romanized transcription of the Chinese character system widely adopted in Chinese education in recent decades. In addition, they had experience of an alphabetical script via their second language, English. Hence, it may be argued that exposure to these script systems entailed sensitivity to segmental representations that may not be representative of Chinese speakers who were never exposed to alphabetical scripts (or perhaps of illiterate individuals). It is difficult to test this possibility because of the widespread educational use of pinyin in mainland China [we note that the Taiwanese participants in the studies by Chen et al. (11) and O’Seaghdha et al. (12) also had most likely acquired Chinese characters via a phonetic transcription system]. Research on Cantonese speakers from Hong Kong may provide some insight into the issue because this population learns characters without the aid of a segmental coding system. However, such individuals are likely to have exposure to English as a second language from relatively early on, which, again, may artificially induce sensitivity to segmental representations.

To conclude, our ERP results provide evidence for the claim that phonemes constitute fundamental functional units of speech production even for languages that do not use an alphabetical script and are unlikely to be artifacts arising solely via exposure with orthography. Most broadly, this finding highlights the universality of sound-sized segments for linguistic systems across the world’s languages.

Methods

Participants.

Participants were 27 native speakers of Mandarin Chinese (13 females; age range: 22–34 y, mean = 24.7 y) from the University of Bristol. All were right-handed, had normal or corrected-to-normal vision, and were not color-blind. All participants had lived in the United Kingdom, on average, for 1.8 y (range: 0.5–4 y).

Materials and Design.

Eight colors (red, yellow, black, blue, brown, green, orange, and pink) were used, and 24 line drawings of objects with no canonical color were chosen from the Snodgrass and Vanderwart (50) picture set. All color names in Chinese were monosyllabic, and all picture names were disyllabic; there was no orthographic overlap between any of the object and color names (SI Appendix). As in English, adjectives precede nouns in Chinese. In the phonologically related condition, each object was presented in a color, such that the object and color name shared a single initial consonantal phoneme (i.e., there was no overlap in the following vowel). In the phonologically unrelated condition, colors and objects were recombined to avoid any phonemic overlap between the object name and color name. Hence, each color was combined with three objects to form 24 phonologically related color–object pairings and with three other objects to form 24 unrelated pairings. Each participant was presented with four blocks of 48 trials with each of the 24 related and 24 unrelated color–object combinations appearing exactly once in each block in a randomized order (192 trials in total). The order of trials within each block was randomized.

Procedure.

The experimental session was administered in Mandarin Chinese by a native Chinese speaker (Q.Q.). Participants were tested individually in an electrically shielded booth. Participants were first asked to familiarize themselves with the experimental stimuli by viewing them in a booklet, with the expected name printed underneath each object. Subsequently, participants were told that they would see the objects in different colors presented on a computer screen and their task was to name them with an adjective-noun combination (e.g., 绿盒子, /lü4he2zi/, “green box”). Participants received a practice trial comprising 16 objects (8 colors repeated twice). Subsequently, four experimental blocks of 48 trials were presented, separated by a short break. The experiment lasted approximately 25 min.

Each trial started with the prompt “Ready?” (“准备好了”?). On a keypress, the prompt was replaced with a blank screen (1,000 ms), followed by a color drawing of an object in the center of the screen against a white background. The drawing disappeared once the participant initiated a verbal response or after a time-out of 5,000 ms. The intertrial interval was 2,000 ms. Participants were instructed to produce a verbal description of the colored object as quickly and accurately as possible. No feedback was provided.

The experiment was delivered using Presentation software (Neurobehavioral Systems). Participants were asked to refrain from moving or blinking during picture presentation to minimize EEG artifacts.

EEG Recordings.

EEG signals were recorded from 64 Ag/AgCl electrodes fitted on an elasticated cap according to the extended International 10–20 system relative to a common FCz reference. Impedances were kept at <5 kΩ. EEG recordings were sampled at 1,000 Hz and filtered off-line using a 40-Hz low-pass (zero-phase) filter. Recordings were analyzed using BESA software (BESA 5.3, BESA GmbH). Blinks were corrected using BESA automatic artifact correction (51). The EEG recordings were segmented into 800-ms epochs relative to picture onset, which included a 200-ms prestimulus interval. Epochs containing amplitudes exceeding ±120 μV were rejected (approximately 4% of all epochs). The 600-ms poststimulus interval was chosen to minimize contamination of EEG signals from speech articulation and was calculated by subtracting a typical response execution time of 300 ms (39) from the average response of around 900 ms (range: 706–1,312 ms) in the present study.

Data Analysis.

Trials in which speakers produced an incorrect color or object name, or a dysfluency (e.g., stuttering, utterance repairs) were excluded from the behavioral and ERP analyses (7.0% of all trials). For both analyses, we additionally rejected trials with naming onset latencies faster than 600 ms (4.6%) and exceeding 3 SDs from the participant’s mean (1.7%). Naming onset latencies were analyzed using a repeated measures ANOVA that included relatedness as a fixed variable and participants (F1) and items (F2) as random variables.

Continuous EEG signals were analyzed in 800-ms-long epochs. Artifact-contaminated epochs were excluded before averaging (3.8% of all trials). The remaining epochs (82.9% of all trials) were baselined using a prestimulus baseline and averaged by condition. Mean amplitudes were calculated separately for each participant and each condition in six time windows (0–100, 100–200, 200–300, 300–400, 400–500, and 500–600 ms). To provide a comprehensive picture of ERP effects, we conducted statistical analyses using six ROIs, with each representing an average of three electrodes: left-anterior (electrodes: F5, F7, and FC5), midanterior (Fz, FCz, and Cz), right-anterior (F6, F8, and FC6), left-posterior (P5, P7, and CP5), midposterior (CPz, Pz, and POz), and right-posterior (P6, P8, and CP6). In this ROI analysis that enabled us to probe the scalp distribution of ERP differences, mean amplitudes from each time window were entered into a 2 × 2 × 3 repeated measures ANOVA with the factors relatedness (phonologically related/phonologically unrelated), anteriority (anterior/posterior), and laterality (left/middle/right). Midposterior electrodes (CPz, Pz, and POz) were identified on the basis of previous research on phonological relatedness in speech production (33, 34) as a likely locus of the effect. In the ERP analyses, Greenhouse–Geisser correction was applied where appropriate.

Supplementary Material

Supporting Information

Acknowledgments

We thank Bill Idsardi for helpful feedback and TingTing Xu for help in data collection. This study was supported by the Faculty of Science, University of Bristol (N.K.) and by Grant SG102008 from the British Academy (to M.F.D.).

Footnotes

The authors declare no conflict of interest.

This article is a PNAS Direct Submission.

This article contains supporting information online at www.pnas.org/lookup/suppl/doi:10.1073/pnas.1200632109/-/DCSupplemental.

References

  • 1.Levelt WJM, Roelofs A, Meyer AS. A theory of lexical access in speech production. Behav Brain Sci. 1999;22:1–38, discussion 38–75. doi: 10.1017/s0140525x99001776. [DOI] [PubMed] [Google Scholar]
  • 2.Dell GS. A spreading-activation theory of retrieval in sentence production. Psychol Rev. 1986;93:283–321. [PubMed] [Google Scholar]
  • 3.Rapp B, Goldrick M. Discreteness and interactivity in spoken word production. Psychol Rev. 2000;107:460–499. doi: 10.1037/0033-295x.107.3.460. [DOI] [PubMed] [Google Scholar]
  • 4.Caramazza A. How many levels of processing are there in lexical access? Cogn Neuropsychol. 1997;14:177–208. [Google Scholar]
  • 5.Dell GS. The retrieval of phonological forms in production: Tests of predictions from a connectionist model. J Mem Lang. 1988;27:124–142. [Google Scholar]
  • 6.Baudouin de Courtenay J. 1972. A Baudouin de Courtenay Anthology; The Beginnings of Structural Linguistics, ed and trans Stankiewicz E. (Indiana Univ Press, Bloomington, IN)
  • 7.Costa A, Sebastian-Galles N. Abstract phonological structure in language production: Evidence from Spanish. J Exp Psychol Learn Mem Cogn. 1998;24:886–903. [Google Scholar]
  • 8.Sevald CA, Dell GS, Cole J. Syllable structure in speech production: Are syllables chunks or schemas? J Mem Lang. 1995;34:807–820. [Google Scholar]
  • 9.Cholin J, Levelt WJM, Schiller NO. Effects of syllable frequency in speech production. Cognition. 2006;99:205–235. doi: 10.1016/j.cognition.2005.01.009. [DOI] [PubMed] [Google Scholar]
  • 10.Levelt WJM, Wheeldon L. Do speakers have access to a mental syllabary? Cognition. 1994;50:239–269. doi: 10.1016/0010-0277(94)90030-2. [DOI] [PubMed] [Google Scholar]
  • 11.Chen J-Y, Chen T-M, Dell GS. Word-form encoding in Mandarin Chinese as assessed by the implicit priming task. J Mem Lang. 2002;46:751–781. [Google Scholar]
  • 12.O’Seaghdha PG, Chen J-Y, Chen T-M. Proximate units in word production: Phonological encoding begins with syllables in Mandarin Chinese but with segments in English. Cognition. 2010;115:282–302. doi: 10.1016/j.cognition.2010.01.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Garrett MF. In: Errors in Linguistic Performance: Slips of the Tongue, ear, Pen, and Hand. Fromkin VA, editor. New York: Academic; 1980. pp. 263–271. [Google Scholar]
  • 14.Meringer R, Mayer K. 1895-1978. Versprechen und Verlesen: Eine psychologisch-linguistische Studie (Göschen, Stuttgart and John Benjamins, Amsterdam)
  • 15.Shattuck-Hufnagel S. In: Sentence Processing: Psycholinguistic Studies Presented to Merrill Garrett. Cooper WE, Walker ECT, editors. Hillsdale, NJ: Erlbaum; 1979. pp. 295–342. [Google Scholar]
  • 16.Fromkin VA. The nonanomalous nature of anomalous utterances. Language. 1971;47:27–52. [Google Scholar]
  • 17.Shattuck-Hufnagel S. In: The Production of Speech. MacNeilage PF, editor. Springer, New York: 1983. pp. 109–136. [Google Scholar]
  • 18.Meyer AS. The time course of phonological encoding in language production: Phonological encoding inside a syllable. J Mem Lang. 1991;30:69–89. [Google Scholar]
  • 19.Alario FX, Perre L, Castel C, Ziegler JC. The role of orthography in speech production revisited. Cognition. 2007;102:464–475. doi: 10.1016/j.cognition.2006.02.002. [DOI] [PubMed] [Google Scholar]
  • 20.Damian MF, Bowers JS. Effects of orthography on speech production in a form-preparation paradigm. J Mem Lang. 2003;49:119–132. [Google Scholar]
  • 21.Roelofs A. Phonological segments and features as planning units in speech production. Lang Cogn Process. 1999;14:173–200. [Google Scholar]
  • 22.Damian MF, Dumay N. Time pressure and phonological advance planning in spoken production. J Mem Lang. 2007;57:195–209. [Google Scholar]
  • 23.Damian MF, Dumay N. Exploring phonological encoding through repeated segments. Lang Cogn Process. 2009;24:685–712. [Google Scholar]
  • 24.Chen J-Y. Inline graphic些国语的自然语误及其分类.华文世界 [A small corpus of speech errors in Mandarin Chinese and their classification] The World of Chinese Language. 1993;69:26–41. [Google Scholar]
  • 25.Chen J-Y. Syllable errors from naturalistic slips of the tongue in Mandarin Chinese. Psychologia. 2000;43:15–26. [Google Scholar]
  • 26.Read C, Zhang YF, Nie HY, Ding BQ. The ability to manipulate speech sounds depends on knowing alphabetic writing. Cognition. 1986;24:31–44. doi: 10.1016/0010-0277(86)90003-x. [DOI] [PubMed] [Google Scholar]
  • 27.Goswami U. In the beginning was the rhyme? A reflection on Hulme, Hatcher, Nation, Brown, Adams, and Stuart. J Exp Child Psychol. 2002;82:47–57, discussion 58–64. doi: 10.1006/jecp.2002.2673. [DOI] [PubMed] [Google Scholar]
  • 28.Wong AW-K, Chen H-C. Processing segmental and prosodic information in Cantonese word production. J Exp Psychol Learn Mem Cogn. 2008;34:1172–1190. doi: 10.1037/a0013000. [DOI] [PubMed] [Google Scholar]
  • 29.Wong AW-K, Chen H-C. What are effective phonological units in Cantonese spoken word planning? Psychon Bull Rev. 2009;16:888–892. doi: 10.3758/PBR.16.5.888. [DOI] [PubMed] [Google Scholar]
  • 30.Damian MF, Martin RC. Semantic and phonological codes interact in single word production. J Exp Psychol Learn Mem Cogn. 1999;25:345–361. doi: 10.1037//0278-7393.25.2.345. [DOI] [PubMed] [Google Scholar]
  • 31.Starreveld PA. On the interpretation of onsets of auditory context effects in word production. J Mem Lang. 2000;42:497–525. [Google Scholar]
  • 32.Dell’acqua R, et al. ERP evidence for ultra-fast semantic processing in the picture-word interference paradigm. Front Psychol. 2010;1:177. doi: 10.3389/fpsyg.2010.00177. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Eulitz C, Hauk O, Cohen R. Electroencephalographic activity over temporal brain areas during phonological encoding in picture naming. Clin Neurophysiol. 2000;111:2088–2097. doi: 10.1016/s1388-2457(00)00441-7. [DOI] [PubMed] [Google Scholar]
  • 34.Jescheniak JD, Schriefers H, Garrett MF, Friederici AD. Exploring the activation of semantic and phonological codes during speech planning with event-related brain potentials. J Cogn Neurosci. 2002;14:951–964. doi: 10.1162/089892902760191162. [DOI] [PubMed] [Google Scholar]
  • 35.Koester D, Schiller NO. Morphological priming in overt language production: Electrophysiological evidence from Dutch. Neuroimage. 2008;42:1622–1630. doi: 10.1016/j.neuroimage.2008.06.043. [DOI] [PubMed] [Google Scholar]
  • 36.Laganaro M, et al. Electrophysiological correlates of different anomic patterns in comparison with normal word production. Cortex. 2009;45:697–707. doi: 10.1016/j.cortex.2008.09.007. [DOI] [PubMed] [Google Scholar]
  • 37.Strijkers K, Costa A, Thierry G. Tracking lexical access in speech production: Electrophysiological correlates of word frequency and cognate effects. Cereb Cortex. 2010;20:912–928. doi: 10.1093/cercor/bhp153. [DOI] [PubMed] [Google Scholar]
  • 38.Ganushchak LY, Christoffels IK, Schiller NO. The use of electroencephalography in language production research: A review. Front Psychol. 2011;2:208. doi: 10.3389/fpsyg.2011.00208. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Indefrey P, Levelt WJM. The spatial and temporal signatures of word production components. Cognition. 2004;92:101–144. doi: 10.1016/j.cognition.2002.06.001. [DOI] [PubMed] [Google Scholar]
  • 40.Wheeldon LR, Levelt WJM. Monitoring the time course of phonological encoding. J Mem Lang. 1995;34:311–334. [Google Scholar]
  • 41.Dell GS. Representation of serial order in speech: Evidence from the repeated phoneme effect in speech errors. J Exp Psychol Learn Mem Cogn. 1984;10:222–233. doi: 10.1037//0278-7393.10.2.222. [DOI] [PubMed] [Google Scholar]
  • 42.Maess B, Friederici AD, Damian MF, Meyer AS, Levelt WJM. Semantic category interference in overt picture naming: An MEG study. J Cogn Neurosci. 2002;14:455–462. doi: 10.1162/089892902317361967. [DOI] [PubMed] [Google Scholar]
  • 43.Schiller NO, Bles M, Jansma BM. Tracking the time course of phonological encoding in speech production: An event-related brain potential study. Brain Res Cogn Brain Res. 2003;17:819–831. doi: 10.1016/s0926-6410(03)00204-0. [DOI] [PubMed] [Google Scholar]
  • 44.Silverman D. A Critical Introduction to Phonology: Of Sound, Mind, and Body. London and New York: Continuum; 2006. [Google Scholar]
  • 45.Dunbar E, Idsardi WJ. Review of “A critical introduction to phonology: of sound, mind, and body.”. Phonology. 2010;27:325–331. [Google Scholar]
  • 46.Kemeny S, et al. Temporal dissociation of early lexical access and articulation using a delayed naming task—An FMRI study. Cereb Cortex. 2006;16:587–595. doi: 10.1093/cercor/bhj006. [DOI] [PubMed] [Google Scholar]
  • 47.Rastle K, Croot KP, Harrington JM, Coltheart M. Characterizing the motor execution stage of speech production: Consonantal effects on delayed naming latency and onset duration. J Exp Psychol Hum Percept Perform. 2005;31:1083–1095. doi: 10.1037/0096-1523.31.5.1083. [DOI] [PubMed] [Google Scholar]
  • 48.Pylkkänen L, Stringfellow A, Marantz A. Neuromagnetic evidence for the timing of lexical activation: An MEG component sensitive to phonotactic probability but not to neighborhood density. Brain Lang. 2002;81:666–678. doi: 10.1006/brln.2001.2555. [DOI] [PubMed] [Google Scholar]
  • 49.Thierry G, Wu YJ. Brain potentials reveal unconscious translation during foreign-language comprehension. Proc Natl Acad Sci USA. 2007;104:12530–12535. doi: 10.1073/pnas.0609927104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Snodgrass JG, Vanderwart M. A standardized set of 260 pictures: Norms for name agreement, image agreement, familiarity, and visual complexity. J Exp Psychol Hum Learn. 1980;6:174–215. doi: 10.1037//0278-7393.6.2.174. [DOI] [PubMed] [Google Scholar]
  • 51.Berg P, Scherg M. A multiple source approach to the correction of eye artifacts. Electroencephalogr Clin Neurophysiol. 1994;90:229–241. doi: 10.1016/0013-4694(94)90094-9. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supporting Information
1200632109_sapp.pdf (76KB, pdf)

Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES