Abstract
A “multidimensional phoneme identification” (MPI) model is proposed to account for vowel perception by cochlear implant users. A multidimensional extension of the Durlach-Braida model of intensity perception, this model incorporates an internal noise model and a decision model to account separately for errors due to poor sensitivity and response bias. The MPI model provides a complete quantitative description of how listeners encode and combine acoustic cues, and how they use this information to determine which sound they heard. Thus, it allows for testing specific hypotheses about phoneme identification in a very stringent fashion. As an example of the model’s application, vowel identification matrices obtained with synthetic speech stimuli (including “conflicting cue” conditions [Dorman et al., J. Acoust. Soc. Am. 92, 3428–3432 (1992)]) were examined. The listeners were users of the “compressed-analog” stimulation strategy, which filters the speech spectrum into four partly overlapping frequency bands and delivers each signal to one of four electrodes in the cochlea. It was found that a simple model incorporating one temporal cue (i.e., an acoustic cue based only on the time waveforms delivered to the most basal channel) and spectral cues (based on the distribution of amplitudes among channels) can be quite successful in explaining listener responses. The new approach represented by the MPI model may be used to obtain useful insights about speech perception by cochlear implant users in particular, and by all kinds of listeners in general.
PACS numbers: 43.64.Me, 43.66.Ba, 43.71.Es, 43.71.Cq [RVS]
INTRODUCTION
Cochlear implants (CIs) are extremely helpful to postlingually deafened listeners. These devices are excellent sensory aids to lipreading, and many patients obtain enough phonetic information from their devices to achieve significant levels of auditory-only, open-set word recognition.
However, the range of speech perception performance shown by cochlear implant users remains wide, and even users who perform relatively well may benefit from alternative speech processing strategies (Wilson et al., 1991). Unfortunately, our basic understanding of the sensory, perceptual, and cognitive mechanisms CI users employ to perceive speech still lags behind the substantial clinical benefit that postlingually deafened adults with CIs obtain with their devices. In particular, we do not know the exact combination of acoustic dimensions that are employed by CI users to understand speech, how sensory information about the different acoustic dimensions is represented and combined, and how that information is used to perform speech identification. These questions are important for two reasons. From a clinical standpoint, increased basic knowledge of the perceptual mechanisms employed by CI users will likely result in better ways to help them understand speech, and to guide the search for better speech processing strategies. In addition, knowledge about how listeners with an impoverished input signal (such as that provided by a CI) understand speech may shed light on the perceptual mechanisms used by listeners with normal hearing.
Stimulation strategies for cochlear implants are frequently classified according to the type of waveform delivered by the device. Although most CI users today use pulsatile stimulation strategies, the use of analog strategies (in particular the simultaneous analog stimulation or SAS strategy, as implemented in the Clarion device, manufactured by Advanced Bionics Corp. in Sylmar, CA) is on the rise. Subjects in this study were users of the compressed analog (CA) stimulation strategy, which is very similar to SAS and has been used successfully to provide substantial speech perception to postlingually deaf subjects with multichannel cochlear implants. Figure 1 illustrates how the CA stimulation strategy is implemented in the Ineraid multichannel CI. The incoming speech signal is filtered into four overlapping frequency bands, with crossover frequencies at approximately 700 Hz, 1.4 kHz, and 2.3 kHz. The filters are broad, with slopes of 12 dB per octave. An analog representation of each filter output is delivered to an intracochlear electrode. Stimulation amplitudes for each electrode are referenced to the individual listener’s thresholds in each electrode. The listener can perform minor volume and sensitivity adjustments using the speech processor’s control knobs, but there are no frequency-dependent adjustments available (i.e., no “tone” control). In the Ineraid CI used by subjects in this study, intracochlear electrodes are placed at 4-mm intervals, with the most basal electrode placed about 8–10 mm from the base of the cochlea and the most apical 20–22 mm from the base. Electrodes closer to the base of the cochlea are associated with filters that have higher cutoff frequencies. Using this strategy, a typical patient can have a fluent one-on-one conversation (with the help of lipreading) and can also achieve significant speech reception scores without lipreading (cf. Tye-Murray and Tyler, 1989; Dorman et al., 1989; Rabinowitz et al., 1992).
FIG. 1.

Simplified block diagram of the “compressed analog” stimulation strategy for cochlear implants.
There are different kinds of acoustic cues that may be used by listeners with the CA strategy. One possibility is that at least the first formant (F1) is encoded by a temporal code, that is, by neurons discharging synchronously with the period of F1. This is certainly plausible: examination of the waveforms output by the lowest frequency channel of the Ineraid device in response to different vowels (see Fig. 2) reveals that the F1 period is well represented and could conceivably be employed by Ineraid cochlear implant users to discriminate some vowels. For example, the waveforms delivered to channel 1 in response to /u/, /ɛ/, and /æ/ are quite distinct, reflecting the periodicity of different F1 values: 350 Hz for /u/, 500 Hz for /ɛ/, and 700 Hz for /æ/ (see Fig. 2).
FIG. 2.

Channel 1 output waveform in response to vowels /u/, /ɛ/, and /æ/.
Psychophysical studies of the rate pitch (i.e., the pitch associated with the repetition rate of a periodic stimulation waveform) of electrical stimuli lend further support to the idea that CA users may obtain frequency information from the waveform sent to channel 1: CI users detect increases in pitch when the frequency of a pulse train or a sinusoid is increased up to 300–500 Hz, and some subjects still detect pitch differences when rate is increased up to 1000 Hz (Shannon, 1993). Therefore, the studies of sinusoid waveform discrimination cited above suggest that cochlear implant users may have the temporal processing capabilities to extract some F1 information from a low-pass-filtered version of the speech signal, delivered to a single electrode. Further support for temporal encoding of F1 frequency is found in experiments where single channel cochlear implant users identify vowels at above-chance levels (Rosen and Ball, 1986; Tyler et al., 1989; Rabinowitz and Eddington, 1995) because, in these experiments, F1 information obtained from the stimulation waveform is essentially the only reliable cue for vowel identification.
Another mechanism that users of the CA strategy may employ to identify vowels is by comparing the amplitude of stimulation delivered to different electrodes. Since the spectra of different vowels have peaks in unique locations of the F1-F2 plane (Peterson and Barney, 1952; Hillenbrand et al., 1995), the stimulation amplitudes delivered to each electrode are vowel dependent. For example, the top panels of Fig. 3 show stimulation amplitudes in each channel of the Ineraid device, in response to different vowels. The /u/ sound (right panel) has very low F1 and F2 frequencies, with practically all the F1 energy and a substantial amount of the F2 energy falling in the low-frequency filter associated with channel 1. Consequently, stimulation amplitude for channel 1 in response to /u/ is much higher than other channel amplitudes. The vowel /æ/ has higher formants than /u/ and F1, for example, is 700 Hz. This means that F1 energy for /æ/ goes partly to channel 1 and partly to channel 2. Consequently, the amplitudes in channels 1 and 2 are quite similar for /æ/, and channels 3 and 4 receive more energy than with /u/ as the input. In summary, stimulation amplitudes for each channel depend on vowel identity and also on the loudness of the input. A loudness change has the same effect on all channel amplitudes when they are measured in dB. Therefore, ratios of channel amplitude remain largely unaffected by variations in input loudness, and may be better cues for vowel identity than loudness levels in individual channels.
FIG. 3.

The top panel shows stimulation waveforms delivered by the Ineraid implant to each channel in response to the vowels /æ/ and /u/. The bar charts indicate the rms amplitude in each channel. The bottom panel shows a “conflicting-cue” vowel, where the waveforms corresponding to /æ/ are amplified or attenuated so that the rms amplitude in each channel is the same as for /u/.
While both types of cues (temporal and channelamplitude) are potentially useful, the specific cue or combination of cues that are actually employed by users of the CA strategy remains to be determined. In an interesting study designed to address this question, Dorman et al. (1992) presented synthetic “conflicting-cue” vowels to a group of subjects. The synthesized waveforms presented to each channel specified one vowel, but the amplitudes were manipulated to specify a different vowel (see Fig. 3). In general, subjects perceived the vowel that was specified by channel amplitudes rather than the one specified by the electrical stimulation waveforms (see Table I), leading Dorman et al. to conclude that users of the CA strategy mostly rely on information contained in channel amplitudes to identify vowels. However, responses to two of the six vowels presented in that study suggest that cues other than channel amplitude were also employed by these subjects. The vowel with /æ/ waveforms and /ɛ/ amplitudes elicited a roughly equal number of /æ/ and /ɛ/ responses and, surprisingly, the vowel with /u/ waveforms and /æ/ amplitudes elicited a majority of /ɛ/ responses.
Table I.
Vowel identification by six Ineraid users (from Dorman et al., 1992). “Conflicting-cue” vowels with stimulation waveforms corresponding to vowel x and channel amplitudes corresponding to vowel y are denoted “temp x, amp y.” The last column lists the predominant response given by subjects in response to a conflicting-cue vowel. Amp indicates that subjects heard the vowel specified by channel amplitudes at least 66% of the time. Amp-temp indicates that subjects heard the vowel specified by channel amplitudes about half of the time, and they heard the vowel specified by the waveforms delivered to each channel the other half of the time. Neither indicates that, most of the time, subjects heard a vowel different from thatspecified by channel amplitudes and from that specified by channel waveforms.
| Stimulus | Response /u/ | Response /ɛ/ | Response /æ/ | Type of response |
|---|---|---|---|---|
| /u/ | 99 | 0 | 1 | |
| /ɛ/ | 0 | 96 | 4 | |
| /æ/ | 0 | 8 | 92 | |
| temp /u/, amp /ɛ/ | 0 | 84 | 16 | Amp |
| temp /u/, amp /æ/ | 0 | 63 | 37 | Neither |
| temp /ɛ/, amp /u/ | 93 | 3 | 3 | Amp |
| temp /ɛ/, amp /æ/ | 10 | 23 | 67 | Amp |
| temp /æ/, amp /u/ | 90 | 7 | 3 | Amp |
| temp /æ/, amp /ɛ/ | 0 | 47 | 53 | Amp-temp |
The aim of the current study was to determine the cues employed by users of the Ineraid cochlear implant to identify vowels, as well as the specific way in which these cues are combined. These questions were addressed with a novel modeling approach. A mathematical model is proposed that explains vowel identification based on estimates of psychophysical performance along the acoustic dimensions hypothesized to be relevant. The mathematical framework of the proposed model is a multidimensional extension of Durlach and Braida’s single-dimensional model of loudness perception (Durlach and Braida, 1969; Braida and Durlach, 1972), which is in turn based on signal detection theory and on earlier work by Thurstone (1927a,b), among others. The multidimensional model is conceptually similar to that proposed by Braida (1991) to explain integration of visual and auditory information in consonant identification. The confusion matrices generated by the model were compared to the Dorman et al. data.
I. METHODS
A. The data
The data to be fit by the model came from the identification experiment conducted by Dorman et al. (1992). The stimuli were /u/, /æ/, /ɛ/, and six “conflicting-cue” vowels where the temporal waveforms delivered to each channel specified one vowel but channel gains were manipulated so that the rms amplitude of each channel specified another vowel. The vowels were steady state, 200 ms long, and they were generated with the KLATT synthesizer (Klatt and Klatt, 1990). Six subjects took part in an identification test where the possible responses were /u/, /æ/, and /ɛ/. These subjects were selected for their superior vowel identification performance. Figure 3 illustrates an example of a conflictingcue vowel created with channel waveforms corresponding to /æ/ and rms amplitudes corresponding to /u/ (denoted as [temp /æ/, amp /u/]). Table II lists the principal acoustic characteristics of all the vowels in this study.
Table II.
Physical characteristics of the stimuli used by Dorman et al. (1992).
| Stimulus | Ch 1 rms (dB) | Ch 2 rms (dB) | Ch 3 rms (dB) | Ch 4 rms (dB) | F1 (Hz) | F2 (Hz) | F3 (Hz) |
|---|---|---|---|---|---|---|---|
| /u/ | 15 | 7 | 6 | 1 | 350 | 1250 | 2200 |
| /ɛ/ | 10 | 6 | 12 | 9 | 500 | 1700 | 2500 |
| /æ/ | 11 | 10 | 9 | 6 | 700 | 1500 | 2400 |
| temp /u/, amp /ɛ/ | 10 | 6 | 12 | 9 | 350 | 1250 | 2200 |
| temp /u/, amp /æ/ | 11 | 10 | 9 | 6 | 350 | 1250 | 2200 |
| temp /ɛ/, amp /u/ | 15 | 7 | 6 | 1 | 500 | 1700 | 2500 |
| temp /ɛ/, amp /æ/ | 11 | 10 | 9 | 6 | 500 | 1700 | 2500 |
| temp /æ/, amp /u/ | 15 | 7 | 6 | 1 | 700 | 1500 | 2400 |
| temp /æ/, amp /ɛ/ | 10 | 6 | 12 | 9 | 700 | 1500 | 2400 |
Table I shows Dorman et al.’s data, averaged across the six subjects. The top three lines simply show that subjects were able to identify the normal, nonconflicting-cue vowels quite successfully. The bottom six lines show subject responses to the conflicting-cue vowels. Four of these lines show that the preponderant response was based on the channel-amplitude cues. For example, when subjects were presented with vowel [temp /ɛ/, amp /u/], they responded /u/ 93% of the time. Similar responses (based on channelamplitude cues) were obtained for vowels [temp /æ/, amp /u/] and [temp /u/, amp /ɛ/]. The three other conflicting-cue vowels present a more mixed picture; [temp /ɛ/, amp /æ/] received 67% /æ/ responses (corresponding to its channel amplitudes) but also a substantial number of other responses; [temp /æ/, amp /ɛ/] received roughly equal numbers of responses according to temporal and channel-amplitude cues, and, finally (and most interestingly), when subjects were presented with a vowel that combined the temporal waveforms of /u/ and the channel amplitudes of /æ/ the preponderant response was /ɛ/, a vowel that had neither the temporal characteristics nor the channel amplitudes of the stimulus.
B. The multidimensional phoneme identification (MPI) model
model Svirsky (1991; Svirsky and Svirsky, 1992) has proposed a mathematical model that predicts phoneme identification based on a listener’s discrimination along specified perceptual dimensions. The model incorporates an internal noise model to account for basic sensitivity, a decision model that allows for response bias, and a multidimensional perceptual space. The MPI model is a multidimensional extension of the Braida and Durlach (1972) model of intensity resolution in one-interval paradigms, which is in turn based on earlier work by Thurstone (1927a,b), among others. In the MPI model, the percept associated with each stimulus presentation is represented as a point in a multidimensional space where all dimensions are statistically independent. Due to perceptual noise, the same stimulus elicits somewhat different percepts every time it is presented to a subject. The model postulates that all percepts elicited by different presentations of the same token are normally distributed along each perceptual dimension. The mean of each multidimensional Gaussian distribution is determined by the physical characteristics of the stimulus, and the standard deviations along each dimension are equal to the listener’s justnoticeable difference (jnd) along the relevant perceptual dimension. In other words, large jnd’s along the relevant perceptual dimensions result in greater uncertainty as to the location of a given stimulus in perceptual space. This is modeled with a broader Gaussian distribution for this stimulus, as exemplified in Fig. 4 for a specific two-dimensional perceptual space. The two perceptual dimensions in Fig. 4 are first formant frequency (which the subject estimates based on the waveform presented to channel 1) and the ratio of amplitudes presented to channels 4 and 1. Three Gaussian distributions are depicted in each panel, corresponding to the vowels /u/, /æ/ and /ɛ/. The top panel shows distributions that may be found with an average subject whose jnd’s along each dimension are large enough that some overlap exists. The overlap between different distributions indicates that this subject will make some identification errors due to his insufficient sensitivity along the relevant perceptual dimensions.
FIG. 4.

Example of a specific two-dimensional perceptual space. The two perceptual dimensions are first formant frequency and the ratio of amplitudes presented to channels 4 and 1. Three Gaussian distributions are depicted in each panel, corresponding to the vowels /u/, /æ/, and / ɛ /. The vertical dimension represents the probability that a given vowel will cause a percept in that location of the perceptual space. The bottom panel shows distributions with smaller standard deviations, as would be found in a “star” subject with relatively small jnd’s along both perceptual dimensions. The top panel shows distributions that may be found with an average subject whose jnd’s along each dimension are double those of the star subject. Overlap between these distribution means that the “average” subject would make some vowel identification errors.
The bottom panel shows distributions with smaller standard deviations, as would be found in a “star” subject with excellent discrimination capabilities (i.e., jnd’s that are half the size of the normal subject’s) along both perceptual dimensions. Given the smaller jnd’s, there is no overlap between the three distributions in this example.
The model postulates that subjects establish their responses by partitioning the perceptual space into nonoverlapping response regions, one region for each possible response. Consequently, each point in the perceptual space belongs to one, and only one, of the response regions. Subject responses are determined by the response region where a given percept falls. In principle, the perceptual space could be partitioned in an infinite number of ways. However, in this study it was assumed that subjects partition the perceptual space in an optimal way (i.e., it was hypothesized that subjects show no response bias). Consequently, perceptual performance is limited only by the overlap between distributions corresponding to different vowels (as illustrated in the bottom panel of Fig. 4). To summarize, the model can be described in a mathematically precise fashion as follows. According to the internal noise model, in a perceptual space with m dimensions, the Gaussian probability function S associated with stimulus Ei is
| Formula 1. |
where xj is the value of stimulus Ei along dimension j, Tij is the average value of stimulus i over dimension j, and jnd is the subject’s just-noticeable difference (jnd) along dimension j. Each nth presentation of stimulus Ei results in a sensation which is modeled as a point which varies stochastically in a multidimensional space, following the Gaussian distribution S(Ei). We shall denote this point as S(Ei, n). The stochastic variation of S(Ei, n) arises from a combination of “sensation noise,” which is a measure of the observer’s sensitivity to stimulus differences along the relevant dimension, and “memory noise,” which is related to uncertainty in the observer’s internal representation of the experimental context. Note that these two kinds of noise (or uncertainty) are different from the variability resulting from separate utterances of the same speech sound, which can be more conveniently modeled by employing one Gaussian distribution for each utterance. Once the value of S(Ei, n) has been determined, a decision model is applied to determine the subject’s response to the stimulus. The decision model associates a “response center” Rk with each possible response, thus creating a partition of the multidimensional space into response regions (one region for each response center). Response region rk consists of the points that are closer to Rk than to any other response center. When calculating distances we assume that the multidimensional space is Euclidean and that the dimensions are orthogonal. Response regions determine a subject’s responses: when a stimulus falls in response region rk (or, equivalently, when a stimulus is closer to response center Rk than to any other response center), the subject’s response is k. One interpretation of the response center concept is that it reflects a subject’s expected sensation in response to a stimulus (e.g., a prototype of the subject’s phoneme category). If the percept associated with a stimulus is close to the expected sensation, the stimulus is identified correctly. If the percept corresponding to stimulus “1” is closer to the expected sensation for stimulus “2” than to the expected sensations for any other stimulus, the subject responds “2” instead of “1,” thus making an identification error. To generalize this concept and make it precise, cell ik in the response matrix (i.e., the percentage of “k” responses to the “i” stimulus) is determined by the multiple integral of distribution Si over the region rk. In other words, the predicted response matrix for a set of stimuli Ei is obtained by integrating all the S(Ei) distributions over each multidimensional response region.
Formula 2.
Predicted response-matrix cell as the integral of distribution Si over response region rk.
To use the MPI model the following steps must be taken. First, hypothesize what are the relevant perceptual dimensions. Second, measure the mean location of each phoneme along each postulated perceptual dimension (this is uniquely determined by the physical characteristics of the stimuli and the selected perceptual dimensions). Third, estimate (or better yet, measure) the subjects’ just-noticeable difference (jnd) along each perceptual dimension, using appropriate psychophysical tests. The fourth step is to estimate (or measure) the response bias along each postulated perceptual dimension. This step may be considered optional if we assume (as in the present study) that the subject is an ideal observer. Note that these steps differ in nature. Specifying the perceptual dimensions is a basically arbitrary (although, one would hope, well informed) decision that may be validated or contradicted by the fit between model output (i.e., a predicted confusion matrix) and the real data. The jnd’s and response bias, on the other hand, are measurable psychophysical parameters. Finally, stimulus values along each dimension are measurable physical quantities that are uniquely determined once the dimensions are chosen. In this study, the four steps described above were handled as follows. Two kinds of perceptual dimensions were considered: a “temporal” dimension that is related to information presented to an individual channel (specifically, this dimension is F1, the first formant frequency, evaluated from the waveform presented to channel 1) and a number of “channel-amplitude” dimensions that are encoded by information presented to different channels (i.e., the ratios of amplitudes presented to channels 2, 3, and 4 with respect to channel 1 amplitude). The temporal dimension is named F1 and the channel-amplitude dimensions are named A2/A1, A3/A1, and A4/A1. These kinds of dimensions (or others conceptually very similar) had already been proposed in the Dorman et al. (1992) study. The second step to implement the MPI model is to determine the location of each stimulus along each dimension (i.e., the value of F1, A2/A1, and the other dimensions for each one of the vowels in this study). All this information can be determined from the data in Table II, which describes the physical characteristics of the stimuli. The third step involves determining or estimating the listener’s jnd’s along each specified acoustic dimension. In the present study, jnd’s were free parameters used to optimize the model’s fit to the data. It was assumed that all channel-amplitude jnd’s were the same, so the number of free parameters was two: a jnd for F1 and a channelamplitude jnd. The F1 jnd was allowed to vary between 50 and 300 Hz, in 10-Hz steps, and the channel-amplitude jnd was allowed to vary between 0.1 and 5 dB, in 0.1 dB steps. Finally, it was assumed that listeners were ideal observers, and thus response bias was zero. The model was implemented by numerically integrating the Gaussian distributions corresponding to a given stimulus and jnd [see Eq. (2)] over the three response regions corresponding to the three possible responses in the experiment: /u/, /æ/, and /ɛ/. In total, nine Gaussian distributions (corresponding to the nine stimuli used in the experiment) were integrated over three regions each (corresponding to the three possible responses), to obtain a 9×3 response matrix that attempted to fit the data shown in Table I. Because each row in the matrix adds up to 100%, the data to be fit by the model has 18 degrees of freedom and the number of free parameters (as discussed above) is 2. Finally, the whole process was repeated for several other choices of dimensions. The first two choices were models with only one type of dimension: temporal-only (i.e., F1) and channel-amplitude-only (i.e., A2/A1, A3/A1, A4/A1).
Although simple examination of the original data set reveals that these dimension choices are insufficient to explain all the experimental data, it is interesting to show how the MPI model fails when an inappropriate set of dimensions is chosen. Then, the MPI model was run using all possible subsets of the channel-amplitude dimensions combined with the temporal dimension, to determine whether a subset of the fourdimensional model may be just as effective in fitting the data.
II. RESULTS
Table III shows the best fitting matrix generated by the four-dimensional model, i.e., the matrix with minimum rms difference with respect to the actual data (shown in Table I). This best fit was obtained for jnd values of 120 Hz for F1 and 2.6 dB for channel amplitude ratios. The jnd for F1 is broadly consistent with the literature on rate pitch perception by cochlear implant users (Townshend et al., 1987) and the jnd for channel amplitude ratios is comparable to (but higher than) data on jnd’s for intensity discrimination, which range from a fraction of a dB to 1–2 dB, depending on presentation level, task performed by the subject, etc. (HochmairDesoyer, 1981; Douek et al., 1977; Fourcin et al., 1979).
Table III.
Best fit obtained with a four-dimensional model that includes one temporal dimension (F1) and three channel-amplitude dimensions (A2/A1, A3/A1, and A4/A1). This fit was obtained for jnd values of 120 Hz for F1 and 2.6 dB for the channel-amplitude ratios. The fitted matrix is quite close to the data: it does not have any errors greater than 20 percentage points, and it has only four cells with errors between 10 and 20 percentage points ~indicated in underlined numbers!. The predominant type of response predicted by the model is indicated in the right column. The second most frequent type of response is also indicated between parentheses, when it exceeds 30% of the total responses. The type of response for the predicted data is quite close to that observed in the listeners (see Table I).
| Stimulus | Response /u/ | Response /ɛ/ | Response /æ/ | Type of response |
|---|---|---|---|---|
| /u/ | 99 | 0 | 1 | |
| /ɛ/ | 0 | 95 | 5 | |
| /æ/ | 0 | 5 | 95 | |
| temp /u/, amp /ɛ/ | 0 | 98 | 2 | Amp |
| temp /u/, amp /æ/ | 9 | 58 | 33 | Neither |
| temp /ɛ/, amp /u/ | 98 | 0 | 2 | Amp |
| temp /ɛ/, amp /æ/ | 2 | 31 | 67 | Amp (temp) |
| temp /æ/, amp /u/ | 88 | 0 | 12 | Amp |
| temp /æ/, amp /ɛ/ | 0 | 64 | 36 | Amp (temp) |
This is reasonable because even though discrimination of amplitude ratios in different electrodes may be mediated by estimating the separate loudness of each electrode, it is still a more complex task than intensity discrimination in a single electrode.
The predicted matrix is quite close to the observed data: it does not have any errors greater than 20 percentage points, and it has very few cells with errors between 10 and 20 percentage points. The mean square error is less than 8%. Furthermore, the model explains the subjects’ excellent performance with the natural vowels; it explains the three conflicting-cue vowels where subjects give a preponderance of responses (80% or more) according to channel amplitudes; it also explains why [temp /ɛ/, amp /æ/] should receive about two thirds of /æ/ responses (based on its channel amplitudes), and some /ɛ/ responses (based on the temporal waveform of channel 1); it explains why [temp /æ/, amp /ɛ/] should receive a substantial number of /ɛ/ responses based on channel amplitudes, and many /æ/ responses based on the temporal waveforms; and, most importantly, the model explains why the vowel with the temporal waveforms of /u/ and the channel amplitudes of /æ/ ([temp /u/, amp /æ/]) should receive a majority of /ɛ/ responses and some /æ/ responses. Table IV shows the results obtained with different dimension choices. As expected, the temporal-only and the channel-amplitude-only models resulted in a much poorer fit than the four-dimensional model. However, two other choices of dimensions resulted in fits that were just as good as that obtained with the four-dimensional model. One of these choices included the F1, A2/A1 and A4/A1 dimensions and the other one included only F1 and A4/A1. The differences between predictions obtained with the fourdimensional model and these two alternative choices of dimensions were quite minor. On one hand, the rms differences between predicted and observed matrices for the alternative choices of dimensions were as small as or even slightly smaller than those for the four-dimensional model, but on the other hand, the four-dimensional model resulted in fewer cells with errors greater than 10%.
Table IV.
Characteristics of the best-fit matrices obtained with different choices of perceptual dimensions. The first four columns indicate which dimensions were used; the next two columns indicate the jnd values that yielded the best fit; and the last three columns list different measures of fit between predicted matrices and observed data: rms difference, number of cells with errors between 10 and 20 percentage points, and number of cells with errors greater than 20 percentage points.
| A2/A1 | A3/A1 | A4/A1 | F1 | Amplitude jnd | Frequency jnd | rms | Cells 10–20% | Cells >20% |
|---|---|---|---|---|---|---|---|---|
| X | X | X | X | 2.6 | 120 | 7.3 | 4 | 0 |
| X | N/A | 260 | 38.8 | 6 | 19 | |||
| X | X | X | 4.9 | N/A | 14.7 | 3 | 6 | |
| X | X | 1.2 | 110 | 9.0 | 4 | 2 | ||
| X | X | 1.7 | 130 | 8.9 | 5 | 2 | ||
| X | X | 1.9 | 130 | 7.3 | 5 | 0 | ||
| X | X | X | 2.3 | 130 | 7.6 | 5 | 0 | |
| X | X | X | 2.7 | 140 | 7.1 | 5 | 0 | |
| X | X | X | 2.5 | 120 | 7.7 | 5 | 0 |
III. DISCUSSION
It was already clear from the original study by Dorman et al. that listeners must have used some combination of temporal and channel-amplitude cues to identify vowels, but many questions remained unanswered. Why were there so many /ɛ/ responses to the stimulus [temp /u/, amp /æ/], given that /ɛ/ had neither the temporal nor the channel-amplitude characteristics of the stimulus? Why didn’t this happen with any of the other five conflicting-cue vowels? Why did the stimulus [temp /æ/, amp /ɛ/] receive a substantial number of responses following the temporal characteristics, as well as many responses following the channel-amplitude characteristics of the stimulus, and why didn’t this happen with the other conflicting-cue vowels? Why did three of the conflicting-cue vowels receive 84% or more responses according to channel-amplitude cues, and why didn’t this happen with the other conflicting-cue vowels? Why were listeners able to identify the nonmodified vowels so well? The main finding of this study is that an extremely simple model incorporating one temporal cue and three channel-amplitude cues was successful in explaining all these questions in a strictly quantitative fashion. Moreover, the best fit was achieved with jnd parameter values that seem quite plausible.
It should be noted that the data set under analysis is consistent with several different choices of dimensions, as indicated in Table IV. All these choices include the F1 dimension and different subsets of channel-amplitude dimensions. However, not just any subset of channel-amplitude dimensions combined with the F1 dimension results in equally good fits: some subsets result in greater rms difference between observed and predicted data, and in some cells having errors greater than 20 percentage points. Future studies of individual listeners should explore whether the dimensions employed in a listening task may be subject dependent or even task dependent.
It may be particularly interesting to explore why the model correctly fit the seemingly strange responses to [temp /u/, amp /æ/]. It is rather difficult to visualize this in a space with four dimensions, but it can be illustrated with a twodimensional space (see Fig. 5 where the dimensions are F1 and A4/A1, the ratio of amplitudes in channels 4 and 1).
FIG. 5.

Filled circles show the location of the three normal vowels used in the Dorman et al. experiment, in a two-dimensional perceptual space that incorporates the temporal dimension and one of the channel-amplitude dimensions used in this study. The heavy black lines indicate the boundaries between the three response regions, which correspond to the three possible responses: /u/, /ɛ/, and /æ/. The empty circle shows the location of the conflicting-cue vowel [temp /u/, amp /æ/]. As the figure shows, this conflicting-cue vowel is closer to /ɛ/ than to any of the other two vowels. Assuming jnd’s of 120 dB for F1 and 2.6 dB for A4/A1, equal geometric distances in this figure represent equal perceptual distances.
Filled circles show the location of the three normal vowels used in the Dorman et al. experiment. The empty circle shows the location of the conflicting-cue vowel [temp /u/, amp /æ/]. As the figure shows, this conflicting-cue vowel is closer to /ɛ/ than to either of the other two vowels. This result is also true in the tetradimensional perceptual space used by the model, explaining why the observed responses are a necessary consequence of the postulated perceptual space.
If the specific dimensions proposed in this study were indeed the ones employed by listeners who use the CA stimulation strategy, these listeners would have difficulty identifying vowels uttered with extreme values of spectral tilt, due to its different effect on the amplitude of the different stimulation channels. This unfortunate inability to perform talker normalization in the face of spectral tilt differences would be a direct consequence of the CI listeners’ limited ability to identify the frequency of a spectral peak, which forces them to rely on more indirect ways of determining vowel identity. Normal listeners, using their ability to discriminate fine differences in formant frequency, are known to identify vowels accurately even in the face of variability in spectral tilt, breathiness, loudness, and many other acoustic parameters. Conceivably, CI users could adapt to the spectral tilt of a single talker by determining the talker’s long-term spectral tilt and adjusting the response centers of different vowels accordingly. However, this strategy would not be as useful when trying to identify vowels in a multitalker test.
More refined versions of mathematical models such as these may be useful both to answer basic research questions and for clinical use. One example of possible clinical use may be the fitting of newer speech processors like that of the Clarion device, which can implement the state-of-the-art continuous Interleaved sampling strategy as well as a version of the compressed-analog strategy (named SAS) in a variety of electrode configurations. The best method that clinicians have at their disposal among the different strategies and electrode configurations that may be available with a specific implant is to let the patient use each strategy or electrode configuration for at least a few weeks (for training purposes) and measure perceptual performance at the end of the training period. In order to rule out possible learning effects, a reversal design (A-B-A) is necessary, further complicating the testing. Models like the ones proposed here may be used to predict a subject’s maximum achievable perceptual performance with a given strategy or electrode configuration by measuring the subject’s psychophysical performance along the relevant perceptual dimensions and using it as input to the model. There is some evidence that psychophysical performance does not change dramatically after cochlear implantation (Brown et al., 1995; Svirsky et al., 1999). This is in contrast to speech perception, which requires weeks or even months to reach asymptote after implantation. The psychophysical tests would not require weeks or months of training, so the whole process of determining which strategy to use would be considerably shortened. The MPI model has already been used in several preliminary studies to explain vowel perception by users of formant-extracting pulsatile cochlear implants (Svirsky, 1991; Svirsky and Svirsky, 1992; Blamey and Svirsky, 1993) and to explain vowel and consonant perception by users of the SPEAK strategy, a vocoderlike scheme used with cochlear implants (Svirsky and Meyer, 1997, 1998). In the case of vowel perception, the model was able to predict the large majority of vowel pairs that were or were not confused by subjects. This was achieved both for individual and for group data, and it was done using one free parameter when jnd data were not available or no free parameters when it was possible to obtain jnd data from the subjects. Group consonant data were also fit by the model quite well: most of the consonant pairs that were or were not confused in an 18-consonant confusion matrix were predicted by the model, using three free parameters.
The traditional approach to investigate the relation between psychophysical variables and speech perception by CI users has been to perform correlational analyses between those variables and speech perception scores, usually with modest results (Tyler et al., 1982; Hochmair-Desoyer et al., 1985; Shannon, 1989). More complex psychophysical parameters have resulted in higher correlations with speech perception scores from CI users, and may hold more promise (Collins et al., 1994; Nelson et al., 1995; Dorman et al., 1996). Although these studies have provided important information, any correlations between psychophysics and speech perception, by their very nature, cannot explain the mechanisms CI users employ to identify speech sounds. A correlation does not imply causality, it may be a byproduct of both the psychophysical parameter and the speech perception scores being correlated with a third variable. Simply stating that speech perception is related to one specific psychophysical variable does not explain how listeners may actually use acoustic information to arrive at a higher-level decision involving categorization and labeling of speech sounds. In addition, finding correlations between speech perception and psychophysics does not explain how listeners may combine various acoustic cues to label the speech sounds they heard.
Finally, performing a linear correlation involves the underlying hypothesis of a linear relation between the psychophysical parameter and the speech perception score. Consequently, it may not be reasonable to hypothesize or expect a high correlation between psychophysical and speech perception measures, even when there is an underlying relation between the two variables, because it is unlikely that this relation will be truly linear.
In contrast to the simple calculation of linear correlations, the MPI model provides new ways to test competing hypotheses about the psychophysical dimensions that underlie speech perception by CI users. The MPI mathematical description states in precise terms what the relevant information is that CI users extract from the acoustic signal, how they combine information from different perceptual dimensions, and how they use this information to identify speech sounds. It is important to point out that the prediction provided by the MPI model is not simply a speech perception score on a given test, but an entire confusion matrix. In other words, the MPI model makes specific predictions as to which pairs of vowels or consonants should be more easily confused by a given CI user. Because the MPI model makes such specific predictions about patterns of perceptual behavior, its hypotheses are more easily falsifiable than those used in the simple correlational approach. Consequently, the approach represented by the MPI model may be helpful in advancing our understanding of the role of sensory discrimination abilities and their relation to speech perception by CI users.
Although the present study has shown the feasibility of using the MPI model to fit group data, one of the most interesting potential uses of the model is the examination of individual data. This can be done either using free parameters (as was done in this study) or by measuring the listeners’ jnd’s along each dimension and then generating predictions without free parameters. A preliminary study (Meyer et al., 1999) has shown that the model’s predictions based on the listeners’ own jnd’s may represent an overestimate of actual performance. This suggests that a listener’s ability to extract and integrate acoustic information from different sources in real time may be limited not only by psychophysical factors, but also by more central, cognitive factors. Studies are underway that will employ the MPI model to assess its validity in fitting individual data and the role of higher order cognitive factors in phoneme identification.
ACKNOWLEDGMENTS
This research was funded by NIDCD Grant No. R01- DC03937, Contract No. N01-DC-2-2402 and a grant from CONICYT (Uruguay). Ashesh Shah provided valuable help modifying figures and running various versions of the model. Important suggestions were made by many colleagues, in particular Don Eddington, Bill Rabinowitz, Mike Dorman, Joe Tierney, Melanie Matthies, Ted Meyer, Karen I. Kirk, David Pisoni, Peter Blamey, an anonymous reviewer, Rosalie Uchanski and (last but not least) Lou Braida.
References
- Blamey PJ, and Svirsky MA (1993). “Identification of vowels and stimulation channels by cochlear implant users,” Speech Communication Group Working Papers, Vol. IX, pp. 35–62. [Google Scholar]
- Braida LD (1991). “Crossmodal Integration in the Identification of Consonant Segments,” Q. J. Exp. Physiol. (1981) 43A(3), 647–677. [DOI] [PubMed] [Google Scholar]
- Braida LD, and Durlach NI (1972). “Intensity Perception II. Resolution in one interval paradigms,” J. Acoust. Soc. Am 51, 483–502. [Google Scholar]
- Brown CJ, Abbas PJ, Bertschy M, Tyler RS, Lowder M, Takahashi G, Purdy S, and Gantz BJ (1995). “Longitudinal assessment of physiological and psychophysical measures in cochlear implant users,” Ear Hear. 16(5), 439–449. [DOI] [PubMed] [Google Scholar]
- Collins LM, Wakefield GH, and Feinman GR (1994). “Temporal pattern discrimination and speech recognition under electrical stimulation,” J. Acoust. Soc. Am 96, 2731–2737. [DOI] [PubMed] [Google Scholar]
- Dorman MF, Smith L, Smith M, and Parkin J (1992). “The coding of vowel identity by patients who use the Ineraid cochlear implant,” J. Acoust. Soc. Am 92, 3428–3432. [DOI] [PubMed] [Google Scholar]
- Dorman MF, Smith LM, Smith M, and Parkin JL (1996). “Frequency discrimination and speech recognition by patients who use the Ineraid and continuous interleaved sampling cochlear-implant signal processors,” J. Acoust. Soc. Am 99, 1174–1184. [DOI] [PubMed] [Google Scholar]
- Dorman MF, Hannley MT, Dankowski K, Smith L, and McCandless G (1989). “Word recognition by 50 patients fitted with the Symbion multichannel cochlear implant,” Ear Hear. 10(1), 44–49. [DOI] [PubMed] [Google Scholar]
- Douek E, Fourcin AJ, Moore BCJ, and Clarke GP (1977). “A new approach to the cochlear implant,” Proc. R. Soc. Med 70, 379–383. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Durlach NI, and Braida LD (1969). “Intensity perception. I. Preliminary theory of intensity resolution,” J. Acoust. Soc. Am 46, 372–383. [DOI] [PubMed] [Google Scholar]
- Fourcin AJ, Rosen SM, Moore BC, Douek EE, Clarke GP, Dodson H, and Bannister LH (1979). “External electrical stimulation of the cochlea: clinical, psychophysical, speech-perceptual and histological findings,” Br. J. Audiol 13, 85–107. [DOI] [PubMed] [Google Scholar]
- Hillenbrand J, Getty LA, Clark MJ, and Wheeler K (1995). “Acoustic characteristics of American English vowels,” J. Acoust. Soc. Am 97, 3099–3111. [DOI] [PubMed] [Google Scholar]
- Hochmair-Desoyer IJ, Hochmair ES, and Stiglbrunner HK (1985). “Psychoacoustic Temporal Processing and Speech Understanding in Cochlear Implant Patients,” in Cochlear Implants, edited by Schindler RA and Merzenich MM, pp. 291–303. [Google Scholar]
- Hochmair-Desoyer IJ, Hochmair ES, Burian K, and Fischer RE (1981). “Four years of experience with cochlear prostheses,” Med. Prog. Technol 8, 107–119. [PubMed] [Google Scholar]
- Klatt DH, and Klatt LC (1990). “Analysis, synthesis, and perception of voice quality variations among female and male talkers,” J. Acoust. Soc. Am 87, 820–857. [DOI] [PubMed] [Google Scholar]
- Meyer TA, Svirsky MA, Kaiser AR, Simmons PM, and Lai TT (1999). “Predicting consonant perception with the Multidimensional Phoneme Identification (MPI) model for individual cochlear implant users,” ARO 22, 719. [Google Scholar]
- Nelson DA, Van Tassell DJ, Schroder AC, Soli S, and Levine S (1995). “Electrode ranking of ‘place pitch’ and speech recognition in electrical hearing,” J. Acoust. Soc. Am 98, 1987–1999. [DOI] [PubMed] [Google Scholar]
- Peterson GE, and Barney HL (1952). “Control methods used in a study of the vowels,” J. Acoust. Soc. Am 24, 175–184. [Google Scholar]
- Rabinowitz WM, Eddington DK, Delhorne LA, and Cuneo PA (1992). “Relations among different measures of speech reception in subjects using a cochlear implant,” J. Acoust. Soc. Am 92, 1869–1881. [DOI] [PubMed] [Google Scholar]
- Rabinowitz WM, and Eddington DK (1995). “Effects of channel-to-electrode mappings on speech reception with the Ineraid cochlear implant,” Ear Hear. 16(5), 450–458. [DOI] [PubMed] [Google Scholar]
- Rosen S, and Ball V (1986). “Speech perception with the Vienna extracochlear single-channel implant: a comparison of two approaches to speech coding,” Br. J. Audiol 20(1), 61–83. [DOI] [PubMed] [Google Scholar]
- Shannon RV (1989). “Detection of gaps in sinusoids and pulse trains by patients with cochlear implants,” J. Acoust. Soc. Am 85, 2587–2592. [DOI] [PubMed] [Google Scholar]
- Shannon RV (1993). “Psychophysics,” in Cochlear implants: Audiological Foundations, edited by Tyler RS (Singular, San Diego: ). [Google Scholar]
- Svirsky MA (1991). “A mathematical model of vowel perception by users of pulsatile cochlear implants,” Presented at the 22nd Annual Neural Prosthesis Workshop, Bethesda, MD, 22–24 October. [Google Scholar]
- Svirsky MA, and Meyer TA (1997). “A mathematical model of vowel perception by cochlear implantees who use the SPEAK stimulation strategy,” presented at the Twentieth ARO Midwinter Meeting, St. Petersburg, FL. [Google Scholar]
- Svirsky MA, and Meyer TA (1998). “A mathematical model of consonant perception by cochlear implant users with the SPEAK strategy,” J. Acoust Soc. Am 103, 2977(A). [DOI] [PMC free article] [PubMed] [Google Scholar]
- Svirsky MA, Meyer TA, Kaiser AR, Basalo S, Silveira A, Suarez H, Lai TT, and Simmons PM (1999). “Learning how to perceive vowels with a cochlear implant: The role of discrimination and labeling,” Assoc. Res. Otolaryngol. Abs, p. 720. [Google Scholar]
- Svirsky MA, and Svirsky SH (1992). “A multidimensional mathematical model of vowel perception by users of pulsatile cochlear implants,” J. Acoust. Soc. Am 92, 2416A–2417A. [Google Scholar]
- Thurstone LL (1927a). “A law of comparative judgement,” Psychol. Rev 34, 273–286. [Google Scholar]
- Thurstone LL (1927b). “Psychophysical analysis,” Am. J. Psychol 38, 368–389. [PubMed] [Google Scholar]
- Townshend B, Cotter N, Van Compernolle D, and White RL (1987). “Pitch perception by cochlear implant subjects,” J. Acoust. Soc. Am 82, 106–115. [DOI] [PubMed] [Google Scholar]
- Tye-Murray N, and Tyler RS (1989). “Auditory consonant and word recognition skills of cochlear implant users,” Ear Hear. 10, 292–298. [DOI] [PubMed] [Google Scholar]
- Tyler RS, Tye-Murray N, Moore BCJ, and McCabe B (1989). “Synthetic two-formant vowel perception by some of the better cochlear implant patients,” Audiology 28, 301–315. [DOI] [PubMed] [Google Scholar]
- Tyler RS, Summerfield Q, Wood EJ, and Fernandes MA (1982). “Psychoacoustic and phonetic temporal processing in normal and hearing-impaired listeners,” J. Acoust. Soc. Am 72, 740–752. [DOI] [PubMed] [Google Scholar]
- Wilson BS, Finley CC, Lawson DT, Wolford RD, Eddington DK, and Rabinowitz WM (1991). “Better speech recognition with cochlear implants,” Nature (London) 352, 236–238. [DOI] [PubMed] [Google Scholar]
