Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2020 Sep 1.
Published in final edited form as: Ear Hear. 2019 Sep-Oct;40(5):1098–1105. doi: 10.1097/AUD.0000000000000696

Effects of Reverberation on the Relationship between Compression Speed and Working Memory for Speech-in-Noise Perception

Paul Reinhart 1, Pavel Zahorik 2, Pamela Souza 1,3
PMCID: PMC6688967  NIHMSID: NIHMS1035074  PMID: 31025984

Abstract

Objectives:

Previous work has suggested that when listening in modulated noise, individuals benefit from different wide dynamic range compression (WDRC) speeds depending on their working memory ability. Reverberation reduces the modulation depth of signals and may impact the relationship between WDRC speed and working memory. The purpose of this study was to examine this relationship across a range of reverberant conditions.

Design:

Twenty-eight older listeners with mild-to-moderate sensorineural hearing impairment were recruited in the current study. Individual working memory was measured using a reading span test. Sentences were combined with noise at two signal-to-noise ratios (2 and 5 dB SNR), and reverberation was simulated at a range of reverberation times (0.00, 0.75, 1.50, and 3.00 seconds). Speech intelligibility was measured in listeners when listening to the sentences processed with simulated fast-acting and slow-acting WDRC conditions.

Results:

There was a significant relationship between WDRC speed and working memory with minimal or no reverberation. Consistent with previous research, this relationship was such that individuals with high working memory had higher speech intelligibility with fast-acting WDRC, and individuals with low working memory performed better with slow-acting WDRC. However, at longer reverberation times there was no relationship between WDRC speed and working memory.

Conclusions:

Consistent with previous studies, results suggest that there is an advantage of tailoring WDRC speed based on an individual’s working memory under anechoic conditions. However, the present results further suggest that there may not be such a benefit in reverberant listening environments due to reduction in signal modulation.

INTRODUCTION

Speech perception in noise remains one of the most challenging tasks for hearing aid users. As such, a great deal of research has gone into investigating optimal hearing aid processing strategies specifically intended for improving speech perception in noise. One strategy for improved speech perception in noise is to tailor hearing aid signal processing based on an individual’s working memory.

Working memory is a short-term processing mechanism used in the simultaneous processing and storage of incoming information during complex cognitive tasks, such as speech perception (Baddeley, 2003). It has been argued that working memory is engaged during speech perception, particularly in situations in which the auditory signal is deficient due to any combination of internal (e.g., hearing loss) or external (e.g., background noise) sources of distortion (Rönnberg et al. 2008, 2013). Because working memory is a limited-capacity system, listeners with lower working memory may be at a disadvantage compared to those with higher working memory in certain challenging listening situations. Consistent with the modifications resulting from hearing aid signal processing, previous work has demonstrated a link between working memory and amplified speech perception (see Souza et al. 2015 for a recent review).

One instance where working memory has been found to affect amplified speech perception is for speech that has been modified using wide dynamic range compression (WDRC) amplification. WDRC amplification provides time-varying gain based on the momentary intensity level of the incoming signals within any number of channels. To prevent rapid gain fluctuations, the speed of the time-varying gain function is partially controlled by a release time parameter. The release time controls the rate of increase in gain as a result of momentary decreases in the input intensity level. Fast-acting WDRC (i.e., processed with short release times) provides rapid gain adjustments which amplify the low-intensity portions of the signal to an audible level. However, in doing so, fast-acting WDRC distorts the modulation characteristics of the signal (Jenstad and Souza, 2005; Davies-Venn et al. 2009; Reinhart et al. 2016). Conversely, slow-acting WDRC (i.e., processed with long release times) adjusts the gain more slowly. The expected result is lower overall audibility of the low-intensity speech segments but less modulation distortion. Previous studies have demonstrated that individuals with higher working memory perform better in background noise with fast-acting WDRC, whereas individuals with poorer working memory perform better in background noise with slow-acting WDRC (Gatehouse et al. 2003; Lunner and Sundewall-Thorén, 2007; Lunner et al. 2009; Souza and Sirow, 2014; Ohlenforst et al. 2016). It is hypothesized that the reason for this is because individuals with greater working memory are able to benefit from the improved audibility of speech processed with fast-acting WDRC, while being able to compensate for the modulation distortion by using top-down processing (Rönnberg et al. 2008). In contrast, individuals with poorer working memory do not have sufficient cognitive resources to compensate for the modulation distortion caused by fast-acting WDRC and subsequently perform better with slow-acting WDRC.

However, the relationship between working memory and WDRC speed for speech perception in noise appears to be limited to cases in which the background noise is modulated. Lunner and Sundewall-Thorén (2007) examined the relationship between working memory and speech intelligibility with fast-acting vs. slow-acting WDRC in both modulated and unmodulated noise conditions. They reported a significant relationship between working memory and WDRC speed for speech intelligibility only in the modulated noise condition. While the exact reason for this is not fully known, it is likely related to how modulated noises provide glimpses of speech information during low amplitude portions of the noise. During these glimpses, listeners receive a segment of the speech signal. Fast-acting WDRC is expected to provide greater audibility of this underlying speech segment than would be provided by slow-acting WDRC. A listener with greater working memory resources may be better able to reconstruct the speech signal from these glimpsed, disjointed components of the signal than a listener with fewer working memory resources. While the results of Lunner and Sundewall-Thoren (2007) suggest that signal modulation characteristics impact the relationship between working memory and WDRC speed, the extent to which variation in modulation characteristics in the real world will impact this relationship is not known.

In partial examination of this issue, Ohlenforst et al. (2016) investigated the role of modulation variation on the relationship between working memory and WDRC speed by varying the number of background noise talkers. Using 1-talker, 2-talker, and 6-talker modulated noises, Ohlenforst and colleagues found that the relationship between working memory and speech intelligibility was present even in the least-modulated noise condition (i.e., 6-talker modulated noise). However, they also found that while the effect was present, it was significantly smaller with the 6-talker noise (less modulated) condition than the 1-talker noise (more modulated) condition. This suggests that the degree of signal modulation may impact the benefit of fitting WDRC speed based on individual working memory. Nevertheless, the authors concluded that truly steady-state, unmodulated maskers are rarely present in the real world and thus this effect is still likely generalizable to the real world.

One previously unaccounted-for factor is that most everyday listening situations contain reverberation. Reverberation occurs when acoustic energy reflections off of features in an environment persist even after the original sound source has ceased. As a consequence of this persistence of acoustic energy, reverberation distorts the transmission of acoustic information, including reducing the modulation depth of signals (Houtgast and Steeneken, 1985; Reinhart et al. 2016). In other words, reverberation causes modulated signals (including background noise) to become less modulated. This will likely decrease the glimpsing opportunities necessary for the relationship between working memory and speech intelligibility with fast-acting vs. slow-acting WDRC. For this reason we hypothesize that the presence of reverberation modifies the benefit of the cognition-based WDRC recommendation for speech intelligibility in noise.

In support of this hypothesis, Reinhart et al. (2017) performed acoustic analyses examining the relationship between reverberation and WDRC speed on speech-in-noise signals. The authors quantified the change in signal-to-noise ratio (SNR) as a result of fast-acting WDRC vs. slow-acting WDRC across a range of reverberant conditions. Consistent with previous work, there was a greater change in SNR with fast-acting WDRC than slow-acting WDRC in anechoic conditions (Naylor and Johannesson, 2009; Alexander and Masterson, 2015). However, as the amount of reverberation in the speech-in-noise increased, the acoustic difference between fast-acting WDRC and slow-acting WDRC decreased. This suggests that reverberation reduces signal modulation, causing signals processed by varying WDRC speeds to become more acoustically similar than they would be in anechoic conditions. Thus, differences in performance across WDRC speeds may be reduced or even eliminated in more reverberant conditions. However, change in SNR is only one acoustic effect of WDRC processing of speech-in-noise signals. It is not known from pure acoustic analyses what the net effect of reverberation will be on listener perception.

The purpose of the present experiment was to examine the relationship between working memory and WDRC compression speed on speech across a range of reverberant conditions. It was predicted that under anechoic conditions (i.e., without reverberation) individuals with higher working memory will benefit more from fast-acting WDRC (compared to slow-acting WDRC), whereas individuals with lower working memory will benefit more from slow-acting WDRC (compared to fast-acting WDRC). However, this relationship will potentially decrease or even be eliminated with increasing reverberation due to the signals becoming less modulated.

MATERIALS AND METHODS

Participants

Twenty-eight older adults with sensorineural hearing impairment who were not current hearing aid users participated in the study (mean age = 73.3 years, range 60 to 85; 17 males, 11 females). Air conduction thresholds were measured at 250–8000 Hz octave frequencies and inter-octaves at 3000 and 6000 Hz. Bone conduction testing was performed at octave frequencies 250–4000 Hz. Participants presented with no more than a single air-bone gap ≥15 dB. Participants had symmetrical hearing loss defined as no more than a 10 dB difference in pure-tone average (thresholds 500, 1000, 2000 Hz) between ears. The Northwestern University Institutional Review Board approved all study procedures. Participants completed an informed consent process prior to participation, and they were compensated for their time.

Reading Span Test

Working memory was assessed using an English-language version of the Reading Span Test (RST) originally developed by Rönnberg et al. (1989). The RST taxed working memory resources by requiring individuals to simultaneously process and store sequential, incoming information. The test materials consisted of 54 sentences. Half of the sentences made semantic sense (e.g., “the captain saw his boat”) and half of the sentences did not make semantic sense (e.g., “the spider biked home”). Sentences were displayed on a 26-inch computer monitor in three clusters (e.g., “the spider – biked – home”) with intervals of 0.8 seconds between each cluster. Sentences were presented in sets ranging from 3 to 6 sentences per set. Participants were required to complete three tasks during the RST: (1) to read the words aloud as they flashed across the screen, (2) at the end of each sentence to make a semantic judgment of whether the particular sentence made semantic sense, (3) at the end of each sentence set, to recall the first or last word of each sentence within that set of sentences. The participant was blinded prior to seeing the set of sentences to whether the first or last word would be prompted. Whether the experimenter asked for the first word or last word of each sentence within a set was pseudo-randomized for each participant, such that first-word and last-word recall conditions occurred an equal number of times over the course of the test. The final RST score was the percentage of first or last words correctly recalled by the test participants out of the 54 sentences.

The distribution of results on the RST are depicted in Figure 1. For a portion of the analyses listeners were split into high and low working memory groups on the basis of their performance on the RST. Based on previous studies in a similar sample population, a cut-off criterion of 41% was used in the present study (Arehart et al. 2013; Ohlenforst et al. 2016). Thirteen individuals were classified as having high working memory (mean RST = 48.2%, SD = 4.9), and fifteen individuals were classified as having low working memory (mean RST = 30.4%, SD = 3.8). Mean participant audiograms for both the left and right ears for both groups can be seen in Figure 2. There were no significant differences observed between high and low working memory groups in age [t(26) = .93, p =.36] or degree of hearing loss, as quantified by 4-frequency pure-tone average (mean of thresholds for both ears at 500–4000 Hz octaves) [t(26) = .04, p = .97].

Figure 1:

Figure 1:

Distribution of scores on Reading Span Test. Each circle represents an individual’s performance, and boxplot shows overall group distribution. The whiskers extend to the most extreme data points not defined as outliers (+/− 2.7 standard deviations from the mean). Individuals were split into High/Low working memory groups based on 41% cutoff score.

Figure 2:

Figure 2:

Mean air-conduction thresholds for each working memory group for the left and right ears. Error bars represent +/− 1 standard deviation.

Stimuli and Processing

Sentence stimuli were taken from the IEEE corpus (IEEE, 1969). Low-context IEEE sentences (e.g., “The ripe taste of cheese improves with age”) were used because the lengths of the sentence stimuli are likely to tax working memory, as well as to ensure that amplitude modulation in the signals will capture the fluctuations of the WDRC processor simulation. Sentences were locally-recorded from a single male talker from the Greater Chicago area in order to control for differences in regional dialect (McCloy et al. 2015). Stimuli were processed in three stages to yield the final set: (1) combined with noise, (2) reverberation simulation, (3) WDRC simulation.

Noise:

Sentence stimuli were first combined with 1-talker modulated noise from the International Collegium for Rehabilitative Audiology (ICRA; Dreschler et al. 2001). ICRA noise was used due to its previous use in the literature exploring the relationship between working memory and WDRC speed (Gatehouse et al. 2003; Lunner and Sundewall-Thorén, 2007; Ohlenforst et al. 2016). The ICRA noise is an artificial signal with speech-like spectral and temporal properties but lacks informational content. The noise preceded the onset of the speech stimuli by 2.0 seconds in order to engage the WDRC processor prior to sentence presentation. On the basis of pilot testing to avoid floor and ceiling effects when combined with reverberation, the background noise was added at 2 and 5 dB SNR. Additionally, these SNRs represent the range of typical SNRs in everyday listening environments (Hodgson et al. 2007). The speech level was fixed at 65 dB SPL and the noise level varied to yield the final SNR levels. Multiple SNRs were used to increase generalizability of results.

Reverberation Simulation:

The second stage of signal processing was to process the speech-in-noise signals using a reverberation simulation. Reverberation was simulated using a MATLAB (Natick, MA) simulation developed and validated by Zahorik (2009) which produced binaural room response simulations. Briefly, the simulation used an image-source model (Allen and Berkley, 1979) to simulate the direct sound and first 500 reflections within a hypothetical room. The direction and delay of late, reverberant responses was estimated based on the source-listener location and dimensions of the modeled room. The attenuation of these late reverberant components were modeled using independent Gaussian noise samples with separate decay functions based on octave band (125–4000 Hz) absorption coefficients of the hypothetical room. Lastly, direct, early, and late reverberant components were spatially-rendered using a non-individualized head-related transfer function.

In the present simulation, source-listener distance was fixed, while the room size and absorptive properties of the reflective surfaces were varied to produce a range of reverberation conditions. Source-listener distance was fixed at 1.4 m which is a typical conversational distance. In the real world, larger rooms tend to produce longer reverberation times. To reflect this, room size was increased incrementally with increasing reverberation. In total, four reverberation conditions were simulated. See Table 1 for a summary of the simulated room sizes for each reverberation time and other details of the simulation. Clarity index (C50) was calculated as the logarithmic ratio of direct sound and early reflections arriving within the first 50 ms, to the late reflections arriving after 50 ms. A higher clarity index is expected to yield higher perceived clarity because there is less reflected energy to mask out the direct energy from subsequent speech portions. Additionally the Speech Transmission Index (STI) was calculated to examine the degree to which the transmission of speech modulations were degraded as a result of reverberation convolution. STI was computed based on standard techniques summarized in Houtgast and Steeneken (1985).

Table 1:

Room size details for each of the reverberation time conditions. Broadband reverberation time was quantified as the mean reverberation time (T60) of octave bands 125–4000 Hz. Similarly, clarity index was calculated as the mean clarity index (C50) of octave bands 125–4000 Hz. Speech Transmission Index (STI) was computed based on standard techniques summarized in Houtgast and Steeneken (1985).

Simulated Room Size (length×width×height) Broadband Reverberation Time Clarity Index (C50) Speech Transmission Index
Free field 0.00 s 1.00
5.7m×4.3m×2.6m 0.75 s 7.79 .77
8.6m×6.5m×3.9m 1.50 s 4.82 .71
12.9m×9.8m×5.9m 3.00 s 6.21 .75

WDRC Simulation:

The third and final stage of stimuli processing was to simulate the effects of varying WDRC speed processing on the reverberant speech-in-noise signals. This was achieved using a modified MATLAB-based simulation developed by Kates (2008). Briefly, the program implemented a six-channel filter bank with center frequencies of 250, 500, 1000, 2000, 4000, 6000 Hz and bandwidths 0–375 Hz, 375–750 Hz, 750–1500 Hz, 1500–3000 Hz, 3000–5000 Hz, 5000-Nyquist frequency, respectively. The filter bank was followed by a peak detector that reacted to within-band fluctuations in signal level. Increases in signal level were followed using an attack time which was set to 10 ms. Decreases in signal level were followed using a release time parameter. Two WDRC speed conditions were simulated using a release time of either 12 ms (fast-acting WDRC) or 1500 ms (slow-acting WDRC). In both conditions, the compression threshold was set at 45 dB SPL, and a compression ratio of 2:1 was used similar to previous WDRC simulation research (e.g., Ohlenforst et al., 2016). Stimuli levels into the compressor were set such that the speech level was fixed at 65 dB SPL.

Speech Task and Procedure

Speech intelligibility was the outcome variable used in the study. Speech intelligibility testing was conducted with the participant seated in a double-walled sound booth. Signals were presented binaurally via Etymotic ER-2 insert phones (Elk Grove Village, IL). Signal playout level was calibrated such that the speech level was presented at 65 dB SPL, and NAL shaping (Byrne and Dillon, 1986) was subsequently applied across the six filter bank channels according to the participant’s hearing thresholds. Presentation blocks were organized by WDRC speed, and block order was randomized for each participant. One second prior to each sentence presentation there was a 1000 Hz pure tone of 250 ms duration in order to alert participants of the imminent sentence presentation. Following sentence presentation participants were asked to repeat the sentence back as best as they could. They were encouraged to guess even if they were not certain of what they heard or only got a part of the sentence. Participant response time was unlimited. Responses were recorded on the basis of 5 keywords per sentence (e.g., “The ripe taste of cheese improves with age”) by a single experimenter to ensure consistent scoring. Listeners did not receive feedback during testing. Stimuli presentation and response recording were controlled by a custom-made MATLAB program. There were 10 sentences per condition for a total of 160 sentences (2 WDRC speeds * 4 reverberation times * 2 SNRs * 10 repetitions). Breaks were given in 20-minute intervals to minimize participant fatigue.

RESULTS

Mean intelligibility scores (in percent correct) for each WDRC speed, SNR, and reverberation time are displayed in Figure 3. Mean intelligibility scores for each condition ranged from 37% to 81% which suggests no substantial floor or ceiling effects in the intelligibility data. Intelligibility was generally poorer in the 2 dB SNR condition than the 5 dB SNR condition. Intelligibility was also generally poorer in the reverberant conditions than the anechoic (0.00 s reverberation time) condition, consistent with the STI values for each reverberation time condition. For statistical analyses, several transformations were performed on the data. First, intelligibility data were first transformed to rationalized arcsine units (RAU) in order to stabilize variance across the performance scale (Studebaker, 1985). Next, data were averaged between the two SNR conditions because the role of SNR was not a central research question and preliminary analyses suggested no interaction with SNR. Lastly, similar to previous studies (Lunner and Sundewall-Thorén, 2007; Ohlenforst et al. 2016), data were reduced by subtracting the scores for the slow-acting WDRC speed from the scores for the fast-acting WDRC speed within-subjects.

Figure 3:

Figure 3:

Mean intelligibility scores (in percent correct) as a function of reverberation time at each SNR. The error bars represent +/− 1 standard error of the mean.

Group Analyses

Data were initially analyzed using a group-split approach similar to those previously used to examine the relationship between working memory and performance with varying WDRC speeds (Lunner and Sundewall-Thorén, 2007; Souza and Sirow, 2014; Reinhart and Souza, 2016). The resulting RAU difference scores are shown in Figure 4. Positive values indicated that listener performance was better with fast-acting WDRC. Conversely, negative values indicated that listener performance was better with slow-acting WDRC. In all the conditions except for the 3.00 second reverberation time, the high working memory group performed better with the fast-acting WDRC, while the low working memory group performed better with the slow-acting WDRC. With a 3.00 second reverberation time, the high working memory group had similar performance with the two WDRC schemes, while the low working memory group performed slightly better with the fast-acting WDRC scheme. For both groups, the mean difference scores were greater in the environments with 0.00 and 0.75 second reverberation times compared to those with longer reverberation times.

Figure 4:

Figure 4:

Mean difference in speech intelligibility scores (RAU-transformed fast-acting WDRC speech intelligibility minus RAU-transformed slow-acting WDRC speech intelligibility) for high and low working memory groups as a function of reverberation time. Positive scores reflect better performance with fast-acting WDRC, and negative scores reflect better performance with slow-acting WDRC. The error bars represent +/− 1 standard error of the mean.

The transformed speech intelligibility scores were analyzed using a 2-way repeated measures analysis of variance (RM-ANOVA) with one within-subjects variable (reverberation time) and one between-subjects variable (working memory group). All assumptions of the model were met. The main effect of reverberation time was not significant [F(3,78)=.396, p=.756, partial η2 =.015]. There was a significant main effect of working memory group [F(1,26)=6.202, p=.019 partial η2 =.193] which suggests that individuals with different working memory performed significantly different with either fast-acting or slow-acting WDRC. However, there was also a significant working memory group×reverberation time interaction [F(3,78)=3.809, p=.013, partial η2 =.128] so these main effects should be interpreted with qualification.

To further explore the working memory group×reverberation time interaction, post hoc independent-samples t-tests were conducted between the difference RAU scores for the High and Low working memory groups at each reverberation time condition with a Bonferroni correction. These results indicated that there was a significant difference in performance with either fast-acting or slow-acting WDRC in the 0.00 second [t(26)=3.532, p=.002] and 0.75 second [t(26)=2.524, p=.018] reverberation time conditions. There was no difference between the groups in either the 1.50 second or 3.00 second reverberation time conditions (both p>.050).

Individual Analyses

Regression analyses were performed to quantify the amount of variance explained by working memory, as variance explained may provide further insight into the potential utility of measuring working memory in a clinical setting. That is, even though group effects may be significant it may be more clinically significant to consider listener working memory for conditions in which working memory accounts for a higher proportion of the variance compared to conditions in which working memory does not account for as much variance in performance.

To explore these relationships, multiple hierarchical regression models were calculated. The dependent variable in each model was the transformed speech intelligibility data. A separate analysis was conducted at each of the reverberation time conditions. The primary predictor of interest was working memory, while also considering the effects of hearing loss (quantified as average of thresholds at 500, 1000, 2000, and 4000 Hz across both ears) and age. The predictors were entered into the models sequentially, in the order of PTA, age, and then working memory. See Table 2 for the results of these analyses.

Table 2:

Results of regression analyses examining listener factors (pure-tone average, age, and working memory) associated with transformed speech intelligibility difference score.

Reverberation Time (s) Variable ΔR2 F p
0.00 PTA .005 .130 .721
Age .010 .188 .830
Working memory .350 4.601 .011
0.75 PTA .079 2.219 .148
Age .022 1.397 .266
Working memory .315 5.696 .004
1.50 PTA .012 .316 .579
Age .058 .935 .406
Working memory .041 .992 .413
3.00 PTA .064 1.772 .195
Age .063 1.818 .183
Working memory .021 1.388 .271

For the 0.00 second reverberation time condition, neither participant age nor PTA were significant predictors. In that condition, working memory was a significant factor (p=.011) and explained 35% of the variance in performance. Similarly, in the 0.75 second reverberation time condition working memory was the only significant factor (p=.004). However, in this condition with slightly less modulation the amount of variance explained by working memory decreased to 31.5%. In both the 1.50 and 3.00 second reverberation time conditions, none of the listener factors were significant predictors (all p>.050).

DISCUSSION

The purpose of the present experiment was to examine the relationship between working memory and WDRC compression speed on speech-in-noise perception across a range of reverberant conditions. Speech intelligibility was measured in older listeners with hearing impairment for sentences in noise with varying amounts of reverberation processed by fast-acting and slow-acting WDRC. Consistent with previous research in anechoic conditions, high working memory was associated with improved speech intelligibility using fast-acting WDRC, whereas low working memory was associated with improved speech intelligibility using slow-acting WDRC. This is consistent with the hypothesis that fast-acting WDRC amplifies brief speech segments during modulation of the noise signal. While this distorts the speech signal, individuals with a higher working memory are able to allocate cognitive resources to construct the speech message from these disjointed speech segments using lexical and contextual information. In contrast, individuals with less working memory resources do not have sufficient cognitive resources to perform this compensatory decoding and instead perform better with slow-acting WDRC which causes less distortion.

The magnitude of this effect without any reverberation was slightly larger in the current study compared to previous studies that used a similar design with low context sentence stimuli (Gatehouse et al. 2003; Lunner et al. 2007; Ohlenforst et al. 2016). This is potentially due to differences in the release time values used between the fast-acting and slow-acting WDRC conditions. Previous studies have used release times of 40 and 640 ms for their fast-acting and slow-acting WDRC, respectively. In contrast, the current study used more extreme release time values of 12 and 1500 ms. It is likely that using more disparate release times would exacerbate the acoustic differences between fast-acting and slow-acting conditions. Thus, the perceptual implications for individuals with varying working memory resources would reflect this when comparing between more disparate release times.

Effects of Reverberation

The presence of reverberation significantly affected the relationship between working memory and performance with fast-acting vs. slow-acting WDRC. In reverberant conditions, the benefit of tailoring WDRC speed based on working memory was reduced (in the 0.75 second reverberation time condition) or eliminated (in the 1.50 and 3.00 second reverberation time conditions). This effect of reverberation on the relationship between working memory and WDRC speed was consistent with our hypothesis, because reverberation reduces the modulation of the noise signal (as evidenced by the STI values) which prevents the WDRC gain function from amplifying the underlying speech components through the noise.

The degree to which reverberation disrupts the relationship between working memory and WDRC speed is dependent on the amount of reverberation and subsequent modulation reduction. At the mild reverberation time (0.75 seconds) there may have still been some modulation of the noise remaining which would provide those glimpsing opportunities for the WDRC gain function to amplify speech segments. Thus, there was still some benefit of tailoring WDRC speed based on working memory (Figure 4). However, due to the overall reduction of signal modulation the benefit and amount of variance in performance explained by examining working memory was reduced compared to the anechoic condition which maintained perfect signal modulation (Table 2). That is, greater amounts of reverberation were likely to completely fill in the noise modulation which prevented the WDRC gain function from increasing during the noise modulation to amplify segments of the speech. This meant that in the higher reverberation time conditions, there was no significant benefit of tailoring WDRC speed based on working memory.

Overall, these effects of reverberation on the relationship between working memory and WDRC speed are consistent with previous acoustic analyses. Reinhart et al. (2017) concluded that fast-acting and slow-acting WDRC had similar effects on the SNR of a speech-in-noise signal as a function of increasing reverberation. It stands to reason that if slow-acting and fast-acting WDRC are acoustically similar in reverberant conditions, then listener perception between those processing conditions should also be similar, as seen in the present study.

While the current study suggests there may be a benefit of tailoring WDRC speed based on working memory at mild reverberation times, there may also be an interaction with noise characteristics. The current study used a 1-talker modulated noise which was highly modulated and provided maximal glimpsing opportunities. It is possible that with a less modulated noise source (e.g., 6-talker modulated noise) in which the relationship between working memory and WDRC speed is already reduced (Ohlenforst, 2016), a mild amount of reverberation would be sufficient to further reduce the signal modulations such that there would be no benefit of tailoring WDRC speed based on working memory. That is, if the present experiment was conducted with 2- or 6-talker modulated noise, then the benefit may have been nonsignificant at the 0.75 second reverberation time.

It is also of interest how different speech and noise materials may affect the relationship between working memory and WDRC speed in reverberant environments. The present study used low context sentences; however, much of the communication that occurs in everyday listening situations is rich in context. Varying linguistic context has been seen to interact with working memory (e.g., Cox & Xu, 2010). Similarly, varying informational masking in the background noise is also likely tax working memory to varying extents (Sörqvist & Rönnberg, 2012). Future work is required to more fully explore the relationships among linguistic properties of speech and noise, WDRC speed, working memory, and reverberation.

Clinical Implications and Future Directions

Overall these results suggest that when considering how to improve listener communication in noise in specific situations it may be important to consider the room acoustics of the listener’s environment. Tailoring the WDRC speed based on individual working memory may provide real-world benefit in environments without much reverberation (i.e., certain restaurants). However, this strategy may not be as beneficial for listeners in more reverberant environments (i.e., theaters or places of worship). This may be especially true since listeners are often seated well outside of the critical distance from the speech source (picture an audience member in the back row of an auditorium) where the reverberation will be more distortive (i.e., yielding lower STI). In these moderately reverberant environments where modulation transmission is more greatly reduced, a listener may experience better speech-in-noise benefit using a different rehabilitation strategy. For example, remote microphones would not only preserve the signal modulation by reducing the amplification of reverberant energy but would more importantly improve the SNR of the signal (Boothroyd, 2004).

There are several commercially available dereverberation algorithms designed to improve listener communication in reverberation (e.g., Fabry and Tehorz, 2005), but the acoustic and behavioral benefits of those algorithms have not been explored in the scientific literature to our knowledge. It is not known whether the parallel use of a dereverberation algorithm would restore the modulation of the speech and noise signals in such a way that the relationship between working memory and WDRC speed would reemerge. The potential benefits of dereverberation algorithms and how they might interact with other hearing aid signal processing algorithms (e.g., WDRC, digital noise reduction, directional microphones) requires further research.

Conclusions

Consistent with previous studies, these results suggest the potential for benefit of tailoring WDRC speed based on individual cognitive ability but only for certain environments. In anechoic conditions, individuals with high working memory performed better with fast-acting WDRC, whereas individuals with low working memory performed better with slow-acting WDRC. However, this effect was diminished in mildly reverberant conditions and eliminated at higher reverberation times. Overall, there should be caution when attempting to generalize expectation of the benefits of tailoring WDRC speed based on individual cognition to the real world where reverberation varies.

Acknowledgements

The authors would like to thank Tim Schoof for help with calibration, Laura Mathews and Melissa Sherman for assistance with participant recruitment, and Thomas Lunner for providing the Reading Span Test. This research was funded by the National Institutes of Health Grants F31 DC015373 to Paul Reinhart, R01 DC008168 to Pavel Zahorik, and R01 DC006014 to Pamela Souza.

Financial Disclosures/Conflicts of Interest:

This research was funded by the National Institutes of Health Grants F31 DC015373 to Paul Reinhart, R01 DC008168 to Pavel Zahorik and R01 DC006014 to Pamela Souza.

REFERENCES

  1. Alexander JM and Masterson K, 2015. Effects of WDRC release time and number of channels on output SNR and speech recognition. Ear Hear, 36(2), e35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Arehart KH, Souza P, Baca R et al. 2013. Working memory, age and hearing loss: Susceptibility to hearing aid distortion. Ear Hear, 34(3), 251. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Baddeley A, 2003. Working memory and language: An overview. J Commun Disord, 36(3), 189–208. [DOI] [PubMed] [Google Scholar]
  4. Boothroyd A, 2004. Hearing aid accessories for adults: The remote FM microphone. Ear Hear, 25(1), 22–33. [DOI] [PubMed] [Google Scholar]
  5. Byrne D and Dillon H, 1986. The National Acoustic Laboratories’(NAL) new procedure for selecting the gain and frequency response of a hearing aid. Ear Hear, 7(4), 257–265. [DOI] [PubMed] [Google Scholar]
  6. Cox RM, & Xu J (2010). Short and long compression release times: speech understanding, real-world preferences, and association with cognitive ability. J Am Acad Audiol, 21(2), 121–138. [DOI] [PubMed] [Google Scholar]
  7. Davies-Venn E, Souza P, Brennan M et al. 2009. Effects of audibility and multichannel wide dynamic range compression on consonant recognition for listeners with severe hearing loss. Ear Hear, 30(5), 494. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Dreschler WA, Verschuure H, Ludvigsen C et al. 2001. ICRA Noises: Artificial Noise Signals with Speech-like Spectral and Temporal Properties for Hearing Instrument Assessment: Ruidos ICRA: Señates de ruido artificial con espectro similar al habla y propiedades temporales para pruebas de instrumentos auditivos. Audiol, 40(3), 148–157. [PubMed] [Google Scholar]
  9. Fabry D and Tehorz J, 2005. A Hearing System that Can Bounce Back from Reverberation: A new hearing aid function has been created that reduces the influence of reverberation in both directional and omnidirectional microphone modes. Hear Rev, 12(10), 48. [Google Scholar]
  10. Gatehouse S, Naylor G and Elberling C, 2003. Benefits from hearing aids in relation to the interaction between the user and the environment. Int J Audiol, 42(sup1), 77–85. [DOI] [PubMed] [Google Scholar]
  11. Houtgast T, and Steeneken HJ (1985). A review of the MTF concept in room acoustics and its use for estimating speech intelligibility in auditoria. J Acoust Soc Am, 77(3), 1069–1077. [Google Scholar]
  12. IEEE. IEEE recommended practice for speech quality measurements; IEEE Report No. 297; 1969. [Google Scholar]
  13. Jenstad LM and Souza PE, (2005). Quantifying the effect of compression hearing aid release time on speech acoustics and intelligibility. J Speech Lang Hear Res, 48(3), 651–667. [DOI] [PubMed] [Google Scholar]
  14. Kates JM, (2008). Dynamic-Range Compression (eds.). Digital hearing aids. Plural publishing, 221–262 [Google Scholar]
  15. Lunner T and Sundewall-Thorén E, 2007. Interactions between cognition, compression, and listening conditions: Effects on speech-in-noise performance in a two-channel hearing aid. J Am Acad Audiol, 18(7), 604–617. [DOI] [PubMed] [Google Scholar]
  16. McCloy DR, Wright RA, and Souza PE (2015). Talker versus dialect effects on speech intelligibility: A symmetrical study. Lang and Speech, 58(3), 371–386. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Naylor G and Johannesson RB, (2009). Long-term signal-to-noise ratio at the input and output of amplitude-compression systems. J Am Acad Audiol, 20(3), 161–171. [DOI] [PubMed] [Google Scholar]
  18. Ohlenforst B, Souza PE and MacDonald EN, (2016). Exploring the relationship between working memory, compressor speed, and background noise characteristics. Ear Hear, 37(2), 137–143. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Reinhart PN and Souza PE, (2016). Intelligibility and Clarity of Reverberant Speech: Effects of Wide Dynamic Range Compression Release Time and Working Memory. J Speech Lang Hear Res, 59(6), 1543–1554. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Reinhart PN, Souza PE, Srinivasan NK et al. (2016). Effects of reverberation and compression on consonant identification in individuals with hearing impairment. Ear Hear, 37(2), 144–152. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Reinhart P, Zahorik P, & Souza PE (2017). Effects of reverberation, background talker number, and compression release time on signal-to-noise ratio. J Acoust Soc Am, 142(1), EL130–EL135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Rönnberg J, Arlinger S, Lyxell B et al. 1989. Visual Evoked Potentials Relation to Adult Speechreading and Cognitive Function. J Speech Lang Hear Res, 32(4), 725–735. [PubMed] [Google Scholar]
  23. Rönnberg J, Rudner M, Foo C et al. 2008. Cognition counts: A working memory system for ease of language understanding (ELU). Int J Audiol, 47(sup2), S99–S105. [DOI] [PubMed] [Google Scholar]
  24. Rönnberg J, Lunner T, Zekveld A, et al. (2013). The Ease of Language Understanding (ELU) model: theoretical, empirical, and clinical advances. Front Syst Neurosci, 7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Sörqvist P, & Rönnberg J, (2012). Episodic long-term memory of spoken discourse masked by speech: what is the role for working memory capacity?. J Speech Lang Hear Res, 55(1), 210–218. [DOI] [PubMed] [Google Scholar]
  26. Souza PE and Sirow L, 2014. Relating working memory to compression parameters in clinically fit hearing aids. Am J Audiol, 23(4), 394–401. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Souza P, Arehart K and Neher T, 2015. Working memory and hearing aid processing: Literature findings, future directions, and clinical applications. Front Psychol, 6, 1894. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Studebaker GA, 1985. A rationalized arcsine transform. J Speech Lang Hear Res, 28(3), 455–462. [DOI] [PubMed] [Google Scholar]
  29. Zahorik P, 2009. Perceptually relevant parameters for virtual listening simulation of small room acoustics. J Acoust Soc Am, 126(2), 776–791. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES