Skip to main content
PLOS Digital Health logoLink to PLOS Digital Health
. 2026 Jul 13;5(7):e0000975. doi: 10.1371/journal.pdig.0000975

Multi-dimensional evaluation of user-operated audible contrast sensitivity tests towards efficient hearing healthcare services

Raul Sanchez-Lopez 1,2,3,*, Johannes Zaar 2,4, Søren Laugesen 1
Editor: Michael Winter5
PMCID: PMC13362146  PMID: 42441667

Abstract

The user-operated audiometry (UAud) project aims at introducing an automated system for user-operated audiometric testing into everyday clinical practice. Here, we focus on the Audible Contrast Threshold (ACT) test, which has been proposed as a language-independent alternative to aided speech-in-noise tests. Five distinct user-operated ACT (U-ACT) test candidates were evaluated in terms of performance, reliability, and usability against two established test benchmarks. The study involved 28 participants with diverse hearing and cognitive abilities. The results demonstrated that the primary factor influencing the reliability of test results was the type of task. In terms of usability, the participants reported positive experience with all test candidates, with the Yes/No task resulting in the lowest perceived difficulty and the best understanding. Overall, a test candidate with a custom adaptive procedure inspired by audiologists’ experiences with ACT was chosen for the UAud protocol as a user-operated test of hearing-in-noise perception.

Author summary

Evaluating how well a person can detect subtle differences in time and frequency, known as contrast sensitivity, is crucial for diagnosing hearing-in-noise difficulties. Traditional clinical assessments often require direct clinical supervision, and intensive patient training, which can limit their accessibility. To address this, we developed and evaluated a series of user-operated, automated testing methods that allow individuals to measure their own contrast sensitivity independently.

We compared five different self-test configurations, varying by the type of interface and response buttons used, against established, researcher-led and audiologist-driven clinical benchmarks. Through a comprehensive, multi-dimensional evaluation spanning psychoacoustic accuracy, test-retest reliability, instruction efficacy, and user experience research, our findings demonstrate that a simplified, automated approach allows participants to complete the test quickly and accurately without clinical oversight. Crucially, this user-operated paradigm proved universally accessible; successful task comprehension, robust interaction quality, and highly reproducible thresholds were achieved entirely independent of a participant’s age or cognitive background. The chosen user-operated audible contrast threshold test offers a way to complement automatic standard pure-tone audiometry and make audiology services more efficient. We anticipate this approach will significantly lower the barriers to comprehensive hearing health assessments and the early detection of hearing deficits.

Introduction

The total number of people living with some degree of hearing loss worldwide is projected to rise from approximately 1.57 billion in 2019 to 2.45 billion by 2050. This represents an increase of 56.1% more individuals who will need hearing healthcare services [1]. Furthermore, it is expected that more than 44% of individuals aged 70 years and older will experience hearing deficits. With this perspective, and an increasing workforce shortage in hearing healthcare [2], there exists an urgent need to improve efficiency and reduce costs in hearing healthcare to ensure access to quality services [1,3,4]. The currently used hearing tests require active participation of a hearing care professional, and represent a bottleneck that effectively limits access to hearing healthcare [5,6]. To meet the demand for hearing assessments, the availability of well-trained audiologists and clinical facilities is crucial [7]. However, audiologists are responsible for various important tasks beyond hearing diagnostic tests such as hearing-aid fitting and verification as well as counselling [8]. The User-operated Audiometry (UAud) project seeks to introduce an automated system that enables user-operated audiometric testing within everyday clinical practice, aiming to support audiologists in prioritizing hearing healthcare services and thereby allowing them to treat more patients and provide better hearing rehabilitation [9].

The most prevalent form of auditory impairment is presbycusis, or age-related hearing loss, which typically affects individuals over the age of 55 [10]. This progressive condition significantly impacts communication and overall quality of life. Notably, speech perception abilities are closely linked to cognitive factors such as working memory [1113]. Furthermore, cognitive decline is strongly associated with auditory deprivation; unaddressed hearing loss has been identified as the leading modifiable risk factor for dementia [14,15]. Within the scope of the UAud project, individuals with comorbid hearing loss and cognitive decline may encounter difficulties with user-operated testing. Elderly people with cognitive impairment generally require more time and more assistance to complete tasks compared to healthy patients when performing behavioral tests or using technology in general [16,17]. It is, therefore, essential to design automated protocols that are accessible to a broad population, regardless of their cognitive profile [18,19].

Currently, audiological assessments are mainly based on pure-tone audiometry, which consists of the detection of pure tones played at different frequencies and sound pressure levels to obtain the patient’s hearing thresholds [8]. However, there has been a growing focus on additionally evaluating supra-threshold abilities [11,20], such as the patient’s ability to discriminate speech in the presence of noise when the speech is fully audible [21]. The clinical evaluation of realistic, or ecologically valid, aided speech-in-noise perception presents several challenges, such as the need for multi-loudspeaker setups and speech material in multiple languages [22,23]. This has motivated the search for alternative tests capable of capturing the fundamental abilities employed in challenging ‘cocktail-party’ scenarios, which are characterized by multiple interfering sound sources including speech [2426]. However, despite decades of research, the clinical characterization of hearing deficits beyond elevated hearing thresholds remains a challenge.

One promising approach involves the detection of spectro-temporal modulations (STM) imposed on broadband sounds [2628]. These modulations create spectral and temporal ripples that mimic the rapid fluctuations of natural speech. A key advantage of STM stimuli is the precise control of modulation levels within classical psychoacoustic adaptive procedures [28,29]. This paradigm is analogous to contrast sensitivity testing in the visual domain, which utilizes spatial and temporal patterns [30,31]. Because the ability to discriminate subtle cues in noisy speech is strongly associated with STM sensitivity [32,33], these tests have been proposed both as a proxy for speech-in-noise perception in ecologically valid environments and as a predictor of hearing-aid benefit [34].

Recently, the Audible Contrast Threshold (ACT) test [35] was introduced as a quick, simple, and clinically viable manual version of an STM sensitivity test [28,33,34]. The results of ACT and its precursor, the STM test [34], showed high correlations with speech reception thresholds measured in hearing-in-noise tests. The speech reception threshold is defined as the speech sound level, relative to the interfering noise, at which an individual can repeat 50% of the speech material (sentences in this case) in an adaptive test. Notably, this correlation was stronger when using realistic speech-in-noise conditions with spatially distributed speech interferers [28]. Although the addition of ACT in current clinical practice is less resource-demanding than other time-consuming audiometric tests, an automated implementation of ACT offers significant benefits for busy clinical services, making it a practical choice alongside the automatic method for testing auditory sensitivity (AMTAS) [6,36].

In the present study, we therefore evaluated distinct candidate user-operated tests of audible contrast sensitivity (U-ACT), which assess the supra-threshold hearing-in-noise abilities of the patient with language-independent stimuli and can thus complement the pure-tone audiometric testing. Both the ACT and the STM tests were considered here as test benchmarks, to which the different U-ACT test candidates were compared. While ACT is a test administered by an audiologist, STM is a user-operated test conducted using a research platform for psychophysical experiments. For the U-ACT tests we considered different tasks and procedures. In this context, a task is the type of patient action that produces a response in the test (e.g., pressing a button upon detecting an event), while a procedure is the method used to seek the threshold, for example, decreasing the contrast level (the strength of the modulation) by a certain value every time there is a response until no further responses are obtained.

The main goal of the present study was to identify an optimal candidate for a user-operated ACT test that is clinically feasible and adheres to the stated objectives. The objectives for a successful user-operated test were:

  • 1. The test should be easy to perform such that the patient immediately understands what needs to be done.

  • 2. The test should not be unnecessarily long; it must allow patients to maintain attention and performance effectively.

  • 3. The test must be reliable and repeatable, demonstrating good agreement with standard or well-accepted alternatives (benchmarks).

  • 4. Overall, the test should be accessible to a broad group of patients without requiring any special modifications for subpopulations, including those with low cognitive abilities.

  • 5. The test should have an optimal quality of interaction so the patients can perform the test seamlessly.

In the present study, we proposed a user-centered multi-dimensional evaluation which aims at providing evidence of the feasibility and adequacy of the new user-operated test in a realistic case scenario and with a diverse group of participants.

Experimental design and study framework

Multidimensional user-centered evaluation

The study design was inspired by the Roles Activities Materials Environment System (RAMES) framework which supports user-centered evaluations in research studies [37]. The objective of the framework is to assess all the factors affecting the outcomes of the evaluation, knowing the interaction among the different elements. Additionally, the design was guided by the ISO/IEC 25022:2016 [38] standard considering the quality in use with focus on the context coverage. The context coverage is related to efficacy, efficiency, and satisfaction and it is defined as the extent to which the system can be used by people without specific knowledge, skills, or experience. The definition of the best candidate was based on the criteria summarized in Table 1.

Table 1. Criteria for the definition of best U-ACT test candidate.

Psychoacoustics User interaction Audiometric groups Cognition groups
High Reliability

Low training effect
High success rate

Low time on task

Low self-reported Difficulty
All groups can perform the test All groups can perform the test

Study participants

The study sample comprised 24 hearing-impaired (HI) participants with diverse hearing and cognitive abilities, all of them over 50 years old, along with 6 young normal-hearing (NH) listeners aged below 30 years. The HI participants were divided into two groups according to their cognitive abilities. To assess the audiological characteristics, the participants underwent standard pure-tone audiometry as part of the protocol. Furthermore, their cognitive abilities were assessed using the reversed digit span [39,40] (RDS) test, which specifically evaluates the working memory performance of the participant.

Study protocol

The study protocol involved two visits, each comprising two blocks. Within each block, participants completed one run of each of five user-operated ACT test candidates, presented in random order. Data from all four blocks were utilized for various investigations into the feasibility of the test candidates. Block 1 evaluated instruction efficacy, blocks 2 & 4 assessed measurement reliability and quality, while block 3 provided a fresh run after one week. This fresh run was intended to simulate a follow-up visit where the test is repeated, but it is not the first time that the participant experiences the test.

In the present study, we specifically focused on exploring the utility of language-independent instructions in user-operated tests. To achieve this, we implemented a set of non-verbal instructions consisting of four main steps: 1) putting the headphones on, 2) familiarization, 3) training, and 4) demonstration. These instructions were refined through multiple iterations with naïve users interacting with the system [41] who were not part of the final study sample. During this pre-investigation, we identified key graphical signifiers (a set of icons) that facilitated correct instruction usage while allowing for user mistakes. After the pre-investigations, we defined efficacy based on a primary success rate reflecting whether the threshold obtained in block 1 was comparable to the one obtained in the remaining blocks. Additionally, we monitored the frequency of support interventions by tracking participant requests for written instructions, supplementary visual cues, or examiner assistance.

Within the two visits, two benchmark tests of the participants’ ability to detect spectro-temporal modulations were used: (i) a research version (STM) [34] and (ii) the clinically viable ACT test [35]. STM is user-operated and uses a 3-alternative forced choice task and a transformed up-down adaptive procedure [29], while ACT is administered by an audiologist. The ACT test consists of a running stimulus of consecutive 1-second sound segments. The audiologist activates the target stimulus which adds a spectro-temporal modulation with a specific contrast level to the sound, and follows a Hughson-Westlake procedure (similar to the pure-tone audiometry) [42] to seek the threshold of the patient based on their responses. The test benchmarks were tested between blocks 1 and 2 in visit 1, and between blocks 3 and 4 in visit 2.

Participants completed a subjective evaluation immediately after each test (except for the initial block). The evaluation consisted of a series of statements rated using a continuous visual analog scale (VAS) [43] from “completely disagree” to “completely agree”. The statements were divided into two categories: test attributes and self-confidence with the test. Test attributes included difficulty, perceived length, demand, and boredom. Self-confidence questions aimed at determining if participants understood the instructions, found the test useful, thought that the results were good, and what their overall opinion of the test was. To ensure the reliability of the subjective usability metrics, the experimental design utilized both positive and negative linguistic framing for the attributes (wording). This approach was implemented to control and assess the stability of participants’ ratings.

User-operated Audible Contrast Threshold test candidates

We designed and implemented a total of five U-ACT test candidates which differed in terms of task and procedure. In a non-automatic test, the task is related to the instructions given to the participant while the procedure consists of step-by-step actions followed by the examiner. In user-operated tests, the procedure is automated. Each U-ACT candidate was numbered from 0 (U-ACT-0) through 4 (U-ACT-4). U-ACT-0 mirrors the standard ACT test by using the same task and procedure (but in an automatic manner), while U-ACT-4 aligns closely with the STM test’s task and procedure. The present multidimensional evaluation aims to compare three tests termed U-ACT-1, U-ACT-2, and U-ACT-3, sharing a custom psychoacoustic procedure but differing in task configuration. These tests are evaluated alongside U-ACT-0 and U-ACT-4, which closely resemble the established benchmarks.

Fig 1 depicts the five U-ACT candidates showing their progression from quick-and-simple ACT-like (U-ACT-0, blue background), to the research-oriented STM-like (U-ACT-4, orange background) designs. The three intermediate candidates (green background) differ only in their task configuration. Note that the 1-button task is only compatible with a continuous stimulus. In terms of the stimulus presentation, U-ACT-1 and U-ACT-0 used a “running stimulus” paradigm requiring responses ‘when’ the target was presented. As depicted in Fig 1, the user interface consisted of a single response button. The test candidates U-ACT-2, U-ACT-3, and U-ACT-4 employed a sequential presentation of three stimuli with brief pauses between intervals. The sequential presentation tests differed in terms of task: in U-ACT-2 the first and third intervals were always references, while the second interval contained either a target or a (catch-trial) reference (a ‘which’ task) and the user interface accordingly consisted of two buttons representing a target (left button) and a reference (right button). U-ACT-3 and U-ACT-4 applied a 3-Alternative Forced Choice (3-AFC) task (a ‘where’ task) where the target is in one of the three intervals and the user interface consists of three buttons corresponding to the three intervals. In terms of procedure, U-ACT-1, U-ACT-2 and U-ACT-3 share a Bayesian-based custom tracking procedure (see Methods section) instead of typical up-down adaptive procedures. The procedure was inspired by the experience of audiologists with the manual ACT, and it utilizes a Bayesian approach to decrease the uncertainty of the measure (the participant’s threshold) with each new stimulus presentation.

Fig 1. Illustration of the essential aspects of the user-operated Audible Contrast Threshold (U-ACT) test candidates, in terms of stimulus, task, user interface, procedure, and similarity with the test benchmarks.

Fig 1

In the user-operated domain, we introduced a visual communication system in which the target stimulus is referred to as an ‘ambulance’, and the reference sound is described as a ‘wave’. From left to right, the test candidates are designated with an increasing number related to their resemblance with the clinical Audible Contrast Threshold (ACT) test and spectro-temporal modulation detection (STM) test.

Results

Participants’ hearing and cognitive abilities

Fig 2 shows the results from pure-tone audiometric assessments, cognitive assessments based on RDS scores, and participants’ ages. The HI participants were divided into two groups based on their RDS scores: a group of 11 people with nominally higher cognitive abilities (HIhi) and a group of 11 with lower cognitive abilities (HIlo). This categorization is visually represented by a dashed line in the corresponding panel (third panel from the left) of Fig 2.

Fig 2. Summary of the participant characteristics.

Fig 2

First and second panels from the left represent the average pure-tone thresholds (bold lines) with maximum and minimums (whiskers) and raw data (thin lines) corresponding to each of the three participant groups (NH: Normal hearing in green, HIhi: Hearing-impaired with higher cognitive abilities in orange, HIlo: Hearing-impaired with lower cognitive abilities in red). Third panel from the left: Reverse Digit Span (RDS) scores, which were the basis for splitting the hearing-impaired group into HIhi and HIlo (separation indicated by black dashed line). Rightmost panel: Age, for the three participant groups. Boxes indicate the 25th, 50th (median), and 75th percentiles; whiskers represent the range of non-outlier data, and symbols (+) denote values exceeding 1.5 times the interquartile range.

The average hearing thresholds among participants with hearing loss showed no noteworthy differences between the two HI groups, with average thresholds below 1 kHz of about 20 dB hearing level (HL) and a gradual increase in thresholds at higher frequencies (between 40 and 60 dB HL). On average, the hearing thresholds were slightly lower in the left ear than in the right ear in HI listeners. This was an incidental finding, as there were no specific symmetry requirements for audiograms during recruitment. The NH group scored between 14 and 22 on the RDS (maximum score of 28), averaging approximately 17.8 points. The participants with hearing loss and higher cognitive abilities scored between 14 and 21 whereas the ones with lower cognitive abilities scored between 6 and 12. In terms of age, the HIlo group was slightly older (median = 72 years) than the HIhi group (median = 66 years).

Test benchmarks are comparable

The results from the two benchmarks across all participants are displayed and compared in Fig 3.

Fig 3. Comparison between the spectro-temporal modulation detection (STM) test and the Audible Contrast Threshold (ACT) test.

Fig 3

Left panel: Scatter plot of agreement between tests. Middle panel: Bland-Altman plot of differences between tests. STM* denotes that the thresholds have been transformed in dB nCL for better comparison with ACT. Right panel: Differences in duration. Boxes represent the interquartile difference and whiskers the maximum and minimum values not considering outliers. Red crosses indicate outliers.

The left panel shows the STM results as a function of the ACT results. ACT (and the result of the U-ACT variants) is measured in normalized contrast level (nCL), a quantity that maps the physical magnitude (modulation level) to perceived contrast. Thus, an ACT score of 0 dB nCL represents the median performance in listeners with normal hearing [35] and an elevated threshold indicates a contrast loss. In contrast, the STM test result is expressed in modulation depth in dB full scale (FS), where 0 dB FS corresponds to fully modulated and it is equivalent to 16 dB nCL [35]. For the Bland-Altman analysis, the dB-FS results from STM were transformed to dB nCL by adding 16 dB. The intraclass correlation coefficient between the STM and ACT test results was very high (ICC(2,k) = 0.87) indicating excellent agreement. The participants’ thresholds spanned a range of -2.95 to 12.9 dB nCL, with a mean value of 1.8 dB nCL. The differences, depicted in the middle panel, were smaller than ±2 dB with a bias of 1.73 dB showing that ACT values were slightly higher than STM results, which is consistent with Zaar/Simonsen et al. [35]. This result demonstrates the equivalence and agreement of both benchmark tests despite the substantial difference in testing time (Fig 3, right panel), with ACT requiring much shorter testing time in line with Zaar/Simonsen et al. [35].

Task comprehension and success across diverse populations

To evaluate the efficacy of non-verbal instructions, a test run in block 1 was considered successful if the ACT value obtained was within 4 dB of the average of results obtained in blocks 2–4. The decision to use 4 dB as a criterion was taken because it is two steps apart in the clinical procedure (employing a stepsize of 2 dB) and is considered within the limits of agreement of the test [35]. Block 1 served as an experimental block where some participants were expected to perform less well due to the absence of training or previous explanations and lack of familiarity with the manual ACT test at this point in time.

Table 2 shows the success rate of each U-ACT test candidate across the three participant groups. The NH group demonstrated high success rates for most U-ACT candidates, while the HI sub-groups showed lower performance, particularly those with lower cognitive abilities HIlo (45–63%). Notably, the 3-AFC task (U-ACT-3 and -4) achieved the highest success rates across both HI groups. For those paradigms, 7/11 participants with lower cognitive abilities and 9/11 with higher cognitive abilities performed reliably without prior knowledge and only with the help of the non-verbal instructions. The test with the lowest success rate in the two HI sub-groups was U-ACT-2 while the U-ACT-1, with running stimulus, showed low success rate only in the group with lower cognitive abilities. Note that U-ACT-2 is the only candidate that has catch trials.

Table 2. Success rates obtained with the different candidate paradigms in the three participant groups. Definition of success: U-ACT-x in block 1 was less than 4 dB apart from the average U-ACT-x result obtained in the three subsequent blocks.

U-ACT-0 U-ACT-1 U-ACT-2 U-ACT-3 U-ACT-4
NH 100% 100% 100% 83% 100%
HI hi 72% 63% 54% 81% 100%
HI lo 54% 45% 54% 63% 63%

Participants frequently utilized support features during the first block. In 31% of the test runs (44/140), participants accessed written instructions; 18% (25/140) required additional visual cues, and 11% (15/140) requested assistance from the examiner. The majority of ‘call the examiner’ requests were attributed to the user interface (UI) design at the end of the demonstration phase, where some participants repeated the demo multiple times after failing to locate the ‘start test’ button. By the second visit (Block 3), support requirements decreased: written instructions were accessed in only 13% of runs (18/140) and visual cues in 5% (7/140), with zero instances of examiner intervention.

Effect of cognitive abilities, task, and procedure on the measured thresholds

To explore the effect of task, procedure, and the likely influence of cognitive abilities, we analyzed the differences between the thresholds obtained from all five U-ACT test candidates and those from the manual ACT, with which the participants were also tested. Fig 4 shows the threshold differences for each of the three groups of participants. The 1-button tasks (U-ACT-0 and -1) with running stimuli showed minimal differences compared to the manual ACT values. In contrast, tasks involving a 3-interval sequence (U-ACT-2-4) yielded significantly lower thresholds than those from the manual test, in agreement with Zaar/Simonsen et al. [35]. In the analysis of variance (ANOVA) with a mixed effects model, no significant differences were observed across the participant groups suggesting that any differences in mean performance among ACT and the various U-ACT tests were largely independent of the participants’ cognitive abilities.

Fig 4. Differences between user-operated Audible Contrast Thresholds (U-ACT) and manually obtained Audible Contrast Thresholds (ACT).

Fig 4

In the boxplots, the horizontal lines indicate the median, boxes indicate interquartile difference, and vertical lines indicate the range between the max and minimum values within ±1.5 times the interquartile range (whiskers). Individual data points outside of the whiskers are also indicated.

Examining the overall data, task-specific differences were evident: there were significant mean threshold differences between the 1-button tasks (U-ACT-0 and U-ACT-1) and the 3-interval yes/no choice of U-ACT-2 (t = 8.96, p < 0.001) and between the 1-button tasks and the 3-AFC tasks, i.e., U-ACT-3 and -4, (t = 15.17, p < 0.001). No significant differences were observed among the tests with sequential presentation (U-ACT-2, U-ACT-3 and U-ACT-4). Notably, the self-paced sequential presentation allows the participant to reflect on the response after the presentation of the three intervals, while the 1-button task is done while there is a running stimulus and it has a fixed response window (0.2-1.6 seconds after the onset of the target stimulus).

Quality of measurements: Reliability, repeatability, and agreement with benchmarks

The test-retest reliability of the U-ACT test candidates was evaluated using the thresholds obtained in the second block of each visit (that is block 2 in visit 1 and block 4 in visit 2). For the analysis of the reliability and repeatability, three different metrics were employed: the inter-class correlation (ICC) [44], the coefficient of repeatability (CoR) [45] from the Bland-Altman plots [46,47] and a metric based on information retrieval named Fractional Rank Precision (FRP) [48]. While ICC and FRP are quantified on a scale from 0 to 1, where 1 indicates ideal reliability, the CoR represents the range within which repeated measurements are likely to fall. The CoR is thus expressed in dB, and a smaller range indicates a more reproducible measure. These metrics provided a comprehensive assessment of the tests’ consistency.

Table 3 shows the reliability results of the U-ACT test candidates and the two benchmarks. The STM test demonstrated the highest reliability. U-ACT-2 emerged as the second-best test, and thus the best U-ACT candidate in terms of reliability, across all considered metrics. The manual 1-button ACT was substantially more reliable than its user-operated counterparts (U-ACT-0 and U-ACT-1).

Table 3. Test-retest reliability of the user-operated Audible Contrast Threshold (U-ACT) test candidates and test benchmarks analyzed using data from blocks 2 and 4. The three metrics are the inter-class correlation (ICC), coefficient of repeatability (CoR), and fractional rank precision (FRP). The rightmost column summarizes the average test duration. Bold letters indicate the best values among the U-ACT test candidates.

Test paradigm ICC (2,1) CoR (dB nCL) FRP (ratio 0–1) Average Duration (s)
U-ACT-0 0.71 3.72 0.75 90
U-ACT-1 0.69 4.11 0.70 97
U-ACT-2 0.84 3.27 0.78 *202
U-ACT-3 0.80 3.59 0.73 *145
U-ACT-4 0.78 3.62 0.74 *292
ACT 0.83 3.51 0.75 82
STM 0.87 3.04 0.75 324

*In hindsight, U-ACT-2,3 and 4 included an unnecessary 1-second pause, which added about 25 seconds to the reported test duration.

To assess the agreement between U-ACT candidates and benchmarks, the same metrics were applied to the average thresholds obtained across blocks 2, 3, and 4. Table 4 highlights that the U-ACT-2 test candidate showed remarkable agreement with both benchmarks ACT and STM achieving an ICC > 0.9, a low coefficient of repeatability (lower than the one corresponding to ACT, i.e., CoR < 3.5) and the best fractional rank precision (FRP = 0.8) compared to the other candidates. Although the other U-ACT test candidates showed an acceptable agreement with both benchmarks, it is remarkable that U-ACT-4 exhibited an excellent agreement with the STM test (ICC(2,k) = 0.97) despite being shorter and using larger step sizes (see Methods).

Table 4. Agreement between the U-ACT test candidates and the test benchmarks analyzed using the mean across blocks 2, 3, and 4. The three metrics were the inter-class correlation (ICC), coefficient of repeatability (CoR) and fractional rank precision (FRP).The best value for each metric is highlighted in bold font.

ICC (2,k) CoR (dB nCL) FRP (ratio 0–1)
Agreement ACT STM ACT STM ACT STM
U-ACT-0 0.86 0.69 3.46 4.18 0.77 0.69
U-ACT-1 0.88 0.76 3.70 4.37 0.78 0.66
U-ACT-2 0.93 0.94 3.16 2.90 0.77 0.82
U-ACT-3 0.80 0.95 4.15 2.81 0.66 0.81
U-ACT-4 0.86 0.97 3.77 2.43 0.69 0.83

Quality of interaction and seamless user experience

Figs 5 and 6 display the standardized data (see Methods) of the usability evaluation for the five U-ACT candidate tests.

Fig 5. Subjective evaluation of the usability of the user-operated Audible Contrast Threshold (U-ACT) candidate tests (U-ACT-0 to U-ACT-4).

Fig 5

Standardized subjective ratings of the attributes related to the test: Difficulty, Length, Demand, and Boredom. Boxplots depict the median, interquartile ranges, and whiskers as in Fig 4. The colors represent Positive (yellow) or Negative (blue) wording. Asterisks indicate statistically significant differences between groups (* p < 0.05, ** p < 0.01, *** p < 0.001). The significance of the attribute wording is shown at the bottom left corners.

Fig 6. Evaluation of the usability of the user-operated Audible Contrast Threshold (U-ACT) candidate tests.

Fig 6

Average subjective ratings of the attributes related to self-confidence with the tests: Understanding of Task, Usefulness, Goodness of Results, and Overall opinion. Boxplots depict the median, interquartile ranges, and whiskers as in Fig 4. The colors represent Positive (yellow) or Negative (blue) wording. Asterisks indicate statistically significant differences between groups (n.s. p > 0.05, * p < 0.05, ** p < 0.01, *** p < 0.001).

Fig 5 compares subjective ratings of attributes related to the test experience. Lower ratings indicate better results, meaning less difficult, shorter, less demanding, or less boring. The wording, which reflects a positive or negative formulation of the attribute, had a significant impact on all four attributes. This suggests that the phrasing of statements can affect ratings. For example, participants found U-ACT-2 to be less “difficult” than U-ACT-4 but still equally “easy”. The most significant results showed that U-ACT-4 was perceived as significantly longer than the other tests and U-ACT-2 was less demanding than U-ACT-0 and U-ACT-4.

Fig 6 shows subjective ratings reflecting attributes related to participants’ confidence with the tests. Higher ratings indicate better results, meaning that the participant understood the task, found the test useful, believed they did well, or had a positive overall opinion of the test. From these attributes, only the overall opinion did not differ significantly in terms of wording. U-ACT-3 was perceived as significantly less useful than U-ACT-1 and U-ACT-2. Additionally, participants reported that the results obtained using U-ACT-2 were significantly better than the ones from U-ACT-0, U-ACT-3 or U-ACT-4.

Discussion

The present study consisted of a multi-dimensional evaluation of various test candidates for a user-operated auditory test. The study aimed to generate quantitative data on test reliability, observations from tracking user interactions with the system, and users’ self-reported opinions about usability. Instructions are critical for obtaining meaningful results in user-operated tests. Initial results from block 1, conducted without additional instructions, indicated that a substantial proportion of hearing-impaired listeners may struggle with understanding their task, particularly those with lower cognitive abilities. All test candidates showed good agreement with benchmarks and good test-retest reliability. While the procedures and tasks differed from classical psychoacoustic tests [28,29], results were comparable to benchmarks in terms of perceptual evaluation results. Some insights emerged, such as perceived differences in test difficulty and length among candidates, and a difference in results based on attribute wording, which showed a strong and clear significance in user evaluations.

The chosen test candidate

The study evaluated user interaction in a realistic scenario and aimed to identify the “best candidate” for user-operated ACT. The selected candidate was U-ACT-2, which exhibited high reliability and repeatability, with excellent agreement compared to ACT and STM benchmarks. Overall, the test was appropriate in terms of length and acceptable, in the sense that the users’ evaluation of U-ACT-2 was average or better than average in all dimensions considered. While generally easy to perform due to its simplified target/reference discrimination task, U-ACT-2 initially presented comprehension challenges for some participants during the first block. However, addressing this through improved instructions and combined verbal/non-verbal cues might be sufficient [49].

Relation to other studies

Previous studies have proposed quick and reliable automated tests of spectro-temporal modulation perception [5052]. In most cases, the process of developing and validating the test has been meticulously done in various steps, adding complexity in each new study towards a realistic application. For example, a spectral ripple discrimination test was developed and further optimized for clinical applications [50,51]. However, there is no clinical evidence for the use of the test for aural rehabilitation. Evaluations of the STM task in the Portable Automated Rapid Testing (P.A.R.T) [53] platform by Gallun, Lelo de Larrea-Mancera and colleagues initially focused on stimulus generation [54] and population-specific performance [52,55]. Subsequent research has significantly expanded this scope to include formal psychoacoustic modeling via the Adaptive Scan algorithm to improve time efficiency and rigorous usability testing to ensure the platform’s viability for at-home auditory assessment [55,56]. In contrast, we have directly tested U-ACT in a realistic environment, with a diverse sample in terms of auditory and cognitive abilities.

The evaluation methods of user-operated audiological tests are mainly focused on accuracy and reliability [57]. The automated pure-tone audiometry protocol utilized in the UAud project (AMTAS) is supported by a substantial body of evidence validating its reliability and clinical effectiveness [36]. Previous multi-center studies have demonstrated high correlation between AMTAS-derived thresholds and traditional manual audiometry [58], including validations of at-home settings [59]. Furthermore, the UAud project has examined the impact of AMTAS results on hearing-aid fitting outcomes [60]. However, comprehensive usability data remain scarce, particularly regarding user experience of older adults. This gap is critical, as cognitive decline may significantly and negatively impact the successful completion of user-operated diagnostics in the overwhelmingly elderly target population. The present study addresses this by including older adults with lower cognitive abilities to determine if U-ACT is accessible to a broad population regardless of cognitive profile, evaluating also the use of instructions and the quality of interaction for five alternative test variants. To our knowledge, this is the first study in audiology to select and evaluate an auditory test using a multimodal approach centered on the user experience and guided by the RAMES framework.

Subjective evaluation

Subjective evaluations are known to have response biases and systematic errors that can distort survey outcomes. In our study, we chose a visual analogue scale because it offers higher sensitivity to small changes than discrete scales and provides continuous data suitable for more powerful statistical analysis [61]. However, it is known to have leniency bias, end-aversion bias, and directional bias (such as pseudo-neglect) [62]. The challenge is to control how the participants make use of the scale to ensure they accurately translate their sensory experiences into the linear format without clustering responses or reinterpreting anchors.

In our study, the inclusion of both positive and negative item wordings was a deliberate methodological choice designed to mitigate acquiescence bias (the tendency of respondents to agree with statements regardless of their content) and directional response bias. By employing dual formulations for the same attribute, we sought to enhance the internal validity of the subjective ratings. A significant difference between U-ACT candidates across both wordings thus indicates a robust perceived difference in usability. In contrast, when the wording factor itself is significant, it highlights a framing effect, suggesting that participant judgments are sensitive to how the question is posed. Furthermore, by doubling the observations per attribute through these dual formulations, we increased the statistical power of the usability evaluation.

The response bias has been investigated in several studies and different mitigation approaches have been tried, such as using intermediate ticks in the scale [63], combination with other techniques [62], or complex statistical modelling [64]. In our study, the combination of standardization of the data and the redundancy obtained through use of different wording allowed us to identify subtle differences between the test candidates.

Strengths and limitations of the study

The UAud approach aims to implement user-operated hearing assessments in healthcare services. We proposed a protocol for multi-dimensional evaluation involving reliability, efficiency, efficacy, and subjective evaluations. Key strengths of the study include a diverse participant selection regarding hearing losses and cognitive abilities, and a comprehensive evaluation using a simple two-visit protocol.

While the multi-dimensional evaluation was valuable, there are limitations to consider. Firstly, participants were sourced from the university’s database, potentially biasing results due to prior hearing research experience affecting task interactions. Besides, the cognitive screening was only based on one aspect of a specific ability (i.e., working memory). Although other types of assessments were considered [65,66], we decided to use a robust test that has been previously used at the Technical University of Denmark (DTU) in various studies [40], rather than a test including aspects like attention or executive function that could be time consuming or less reliable. Furthermore, young normal-hearing listeners with low cognitive abilities might be an interesting population to evaluate the test in terms of reliability and usability. Secondly, repeated assessments of the same listening ability within each block may have positively influenced results through training effects. This was a known limitation that we could not disentangle in the analysis. As an alternative, we could have used a different experimental design such as a randomized trial with one group per test candidate, however, this approach would have resulted in other limitations since each individual would only have used one of the candidates. Additionally, the protocol lacked speech intelligibility in noise tests as a reference measure, as included in Zaar et al. [34]. The U-ACT test is expected to be equivalent to the clinical ACT and have a similar predictive power so it can be considered a proxy for speech-in-noise tests. However, this should be explored in a separate study comparing the two contrast tests with the hearing-in-noise test with a clinically relevant population.

Outlook

The auditory tests used for audiological diagnosis should provide a reliable estimate of a patient’s hearing status; however, manual and user-operated tests need not be based on identical tasks or procedures [6]. While the reliability of test results is influenced by user behavior, it is the responsibility of both the audiologist (human) and the system (represented by the system designer) to present the test in a manner that minimizes influencing factors unrelated to hearing status, such as attention, understanding of the task, and fatigue [67,68]. In user-operated tests, it is reasonable to think that this difference in the allowed response time is the most influential factor. The results obtained here support the use of sequential presentation in user-operated tests (as also demonstrated in AMTAS studies [36]), while emphasizing the importance of the examiner’s role in manually operated ACT. In research, the current version of U-ACT opens the possibility to test other populations (pediatric, cognitively impaired) using a single and similar framework for audibility and audible contrast tests, and to explore the use of U-ACT as a remote test (as in the case of the audiometry [69]). Overall, U-ACT can help to further investigate the role of contrast loss in auditory perception and communication handicap.

Conclusion

The present study utilized a multi-dimensional evaluation framework to compare user-operated audible contrast sensitivity tests in terms of feasibility, reliability, and accessibility towards the implementation of the U-ACT in hearing healthcare services.

  • The test proved easy to perform, patients immediately understood the instructions; this is supported by the fact that examiner interventions dropped to zero instances by the second visit and participants reported the highest “understanding” ratings for the preferred test paradigm, U-ACT-2.

  • The test duration was not unnecessarily long; U-ACT-0 and U-ACT-2 were significantly more efficient than traditional methods, completing in an average of 90 and 202 seconds respectively, compared to over 324 seconds for the STM benchmark.

  • The test results were reliable and repeatable, demonstrating excellent agreement with clinical benchmarks. This was specifically true for the U-ACT-2 Yes/No candidate which achieved an ICC of 0.84 (test-retest) and a high correlation with the manual ACT benchmark (ICC > 0.90).

  • The test was accessible to a broad group of patients, with successful completion independent of participant subpopulation or cognitive profile. Statistical analysis verified that the cognition grouping did not significantly influence measured thresholds or the reliability of the results.

  • The quality of interaction was optimal, enabling patients to perform the test seamlessly; subjective data confirmed the U-ACT-2 Yes/No candidate was perceived as the least difficult and most usable interface among all tested variants.

Methods

Study design and participants

Participant recruitment.

The potential participants of the study were recruited via purposive convenience sampling from the database of the Hearing Systems Section at DTU. In the recruitment process, care was taken to ensure a balanced number of male and female participants, and a reasonable variety of hearing losses and cognitive abilities as reflected in their audiometry and RDS scores. The potential participants were divided into three groups already during the recruitment based on the information contained in the database. The study aimed to include a total of 30 participants, stratified into three groups: six young NH controls, twelve HI listeners with high cognitive abilities (HIhi), and twelve HI listeners with low cognitive abilities (HIlo). We intentionally recruited 6 NH listeners below 30 years old which is sufficient to establish a stable baseline for interaction quality [70]. Cognitive classification was based on the RDS scores; the high-ability group included scores from 14 to 21, while the low-ability group included scores from 6 to 12. The only exclusion criteria were a hearing loss above the range where the ACT can be performed (i.e., exceeding 95 dB HL) or an inability to operate a tablet due to dexterity or vision problems. Thus, a total of 30 participants with diverse hearing and cognitive abilities were recruited for the study. Two participants abandoned the study after the first visit and 28 participants completed the protocol. This study was carried out in accordance with the ethical approval of the Danish Science-Ethics Committee granted by the Capital Region Committee (H-16036391). All participants gave written informed consent and were offered an economic compensation for their participation.

Equipment and testing environment.

The study was conducted in the laboratory of the Hearing Systems Section at the DTU, specifically in the audiology clinic (which resembles a typical audiology clinic). This facility includes a sound-isolated booth equipped with a clinical audiometer, videoscope, middle-ear analyzer, and computer for conducting hearing-in-noise tests and cognitive assessments. For the present study, we employed a clinical research prototype of Audible Contrast Threshold testing implemented in MATLAB (Version 9.8, The MathWorks Inc.). The prototype consists of an RME UC soundcard (RME Audio, Haimhausen, Germany) connected to a Lake People G103-P MKII amplifier (Lake People, Überlingen, Germany). The prototype included a pair of headphones Radioear DD450 (Middelfart, Denmark), and a push button. For the user-operated tests, we used an iPad 8th generation MYL92KN/A (Apple Inc., Cupertino, CA, USA) connected to the computer functioning as an external screen using the software Splashtop Wired Xdisplay (Splashtop Inc., San Jose, CA, USA).

Protocol.

The participants were invited for two visits that occurred at least one week and no more than four weeks apart. In the first visit, they were introduced to the study and were informed that the test involved a usability test and that we wanted the participants to interact with the system. After this briefing, participants underwent a standard pure-tone audiometry examination administered by a trained audiologist before they conducted the main experiment. The second visit began with the administration of the RDS test followed by the main experiment. The main experiments consisted of a total of four testing blocks (two blocks in each visit), where participants evaluated five alternative U-ACT candidates, with candidates presented in balanced order across groups and blocks. In each visit, both test benchmarks (that is, ACT and STM) were administered in between the two testing blocks once the examiner had explained the test to the participant. After block 4, semi-structured interviews were conducted with each participant. The interview included questions about the opinions of having user-operated tests in the clinic and the utility of the U-ACT to complement the audiometry. The qualitative results of these interviews are not within the scope of the present investigation.

Auditory tests of audible contrast sensitivity

Stimuli.

The stimuli used in all the tests performed in the study were identical to those used in the ACT as described by Zaar/Simonsen et al. [35], which has been thoroughly explained in previous studies using STM detection [34]. In general terms, the stimuli consist of 1-second-long broadband noise carriers (354–2000 Hz) – serving as reference stimuli –, onto which spectral modulations of 2 cycles per octave and temporal modulations of 4 Hz are imposed to create the target stimuli. The experimental variable was the modulation depth (m), which was transformed into dB FS by 20 log(m) for the STM test, and into dB nCL for all the ACT tests (both manual and user operated).

Test benchmarks.

Clinical assessment of contrast sensitivity has emerged as an important diagnostic tool for evaluating a patient’s ability to extract subtle acoustic information from complex sounds, and its use in hearing rehabilitation is supported by evidence [34,35]. In this study, we compare the performance of U-ACT candidates with two benchmarks that have been previously used in relevant studies [34,35]. The first benchmark, referred to as the STM detection test [28,34], is a psychoacoustic measure that utilizes a three-alternative forced choice (3-AFC) task and a 2-down 1-up procedure that approximates the 70.7%-point on the psychometric function [29]. The psychometric function represents the probability of detecting the target sound and it is often modelled as a logistic regression function with a midpoint (threshold) and a slope that define the performance for any level of the tracking variable. The second benchmark is the clinical version of the Audible Contrast Threshold (ACT) [35] test, which has been widely used in previous studies.

Test tasks.

Three tasks were considered for the U-ACT test candidates. These were:

  1. 1-button task: The stimuli are presented in 1-second-long successive waves of noise, and occasionally the wave contains the target stimulus (i.e., it is modulated), thus differing from the other (unmodulated) reference waves. We use the term ‘wave’ because the 1-second-long noise has a fade-in and fade-out and it is followed immediately by another 1-second-long noise so it sounds like waves in the ocean. A large virtual button, referred to as the ‘1-button’, was placed on the touch screen occupying 40% of the screen’s area at the bottom. We implemented a fixed response window (starting 200 ms after stimulus onset and ending 600 ms after offset, for a total of 1.6 s) to categorize trials as HIT or MISS. This window mirrors the parameters used in the manual ACT [35], assuming comparable response times between physical buttons and touch screens. If the participant did not touch the screen, the trial was marked as MISS, whereas the trial was marked as a HIT if the participant touched the screen during the response window. The screen touches outside of the response window were registered as false positives. This task was used in U-ACT-0 and U-ACT-1.

  2. Yes/No task: The stimuli were presented in triplets, the first and last intervals were always references, while the second interval could be either a reference or a target stimulus. The user’s task was to press either the button depicting an ambulance when detecting a target, or the button depicting a wave indicating the detection of a reference, i.e., no detection of a target. This task resembles the one used in AMTAS [36] and includes catch trials with feedback when the participant gets caught pressing the button when there was no target stimulus presented. This was done to ensure that the performance and attention is maintained during the test. The catch trials consisted of trials containing a reference stimulus in the second interval. There were 5 catch trials which were assigned at the beginning of the test to be presented before specific trials. The first catch trial is always between the first and second trials, the remaining were assigned randomly. The feedback was presented only when the participant was caught in a catch trial which is expected to minimize the bias and have less influence on the results of the actual test [71,72]. There were no restrictions on response time since the stimuli were presented sequentially. This task was used in U-ACT-2.

  3. 3-AFC task: The stimuli were presented in triplets. The screen showed three virtual buttons with numbers 1–3, which were highlighted sequentially in synchrony with the playback of the three intervals. The target stimulus was randomly placed in one of the three intervals. The task of the participant was to indicate which of the three intervals was different from the other two by pressing the button numbered with that interval. If no differences were heard, guessing was compulsory to advance the test. There were no restrictions on response time since the stimuli were presented sequentially. This task was used in U-ACT-3 and U-ACT-4.

Test procedures.

Three procedures were considered for the U-ACT test candidates:

  • Modified Hughson-Westlake procedure [42]: This procedure was identical to that used in the ACT test, but automatic. After each response (either HIT or MISS), the contrast level of the next target stimulus was adjusted by, respectively, decreasing or increasing the contrast level. The step size was 2 dB nCL. After a HIT, the contrast level was decreased by 2 steps (i.e., 4 dB nCL). After a MISS, the contrast level was increased by 1 step (i.e., 2 dB nCL). Following the ‘ascending method’ [73] used in pure-tone audiometry, testing ceased when three out of five positive responses occurred at a single contrast level, provided each followed a non-response. This procedure was used in U-ACT-0.

  • Custom Bayesian procedure: The presented procedure is an adaptation of the BAYES-FIG [74]. In their procedure, Remus and Collins combined Bayesian inference with the Fisher information matrix and Theory of Optimal Experiments. Overall, the Bayesian inference aims to reduce the uncertainty by obtaining new evidence (a response at a certain level) and updating the probability of a given combination of midpoint and slope of the psychometric function to be the true result. The Fisher information matrix is obtained using the derivative of the psychometric function with respect to the midpoint and the slope and estimating which value of the tracking variable is the one that should be tested next to maximize the information provided by the new evidence. However, our method incorporates distinct variations:

    • Bayesian updates: Unlike BAYES-FIG, our method utilizes the last two trials (n-1 and n) to update probabilities, rather than relying solely on the last trial as in other psychoacoustic procedures based on Bayes theorem. This was done under the assumption that using two trials makes the estimation more robust to momentary lapses in participant attention. This assumption was tested and confirmed using Monte Carlo simulations.

    • Informational gain function: Instead of using Fisher’s informational gain, our method employs a different information gain function that is derived solely from the slope’s derivative.

    • Optimal next CL algorithm: In the BAYES-FIG the selection of the level for the next trial is based on the maximum of the information gain function. In contrast, our method alternates between the maximum and minimum values corresponding to the 82% and 18% points of the psychometric function, respectively, on even and odd trials. The exact points of the psychometric function are defined by the partial derivative of the psychometric function with respect to the slope.

    • Exception inclusion: To mirror the manual test experiences reported by audiologists, our procedure incorporates additional catch trials (used in the Yes/No task) or presentations above threshold to maintain patient attention and ensure consistent performance. These presentations occurred based on erratic patient responses that might suggest that the patient has lost the internal reference of the sound that they have to detect.

    • Catch Trials: In the Yes/No task, catch trials were included in the procedure. By default, the system considers 4–5 catch trials assigned randomly to specific trial numbers with 4–8 trials in between.

Overall, the Bayesian procedure consisted of 25 trials that started with 8 pre-test trials for a first approximation of the threshold using the modified Hughson-Westlake method and then continued alternating easier and more difficult presentations using the Optimal next CL algorithm. This procedure is employed in U-ACT-1,2 and 3.

  • Modified 2-down 1-up procedure: This procedure was based on the 2-down 1-up described by Levitt [29], with a slight modification to make it shorter. Thus, the contrast level of the stimulus was decreased only if there were two consecutive correct responses and increased otherwise. The entire procedure consisted of two iterations (sub-runs), each of them consisting of 4 upper reversals with decreasing step size. An upper reversal is a change of direction that takes place in ascending direction after two consecutive positive responses. At the start, the step size was 8 dB after the first upper reversal it decreased to 4 dB, and after the second to 2 dB. The procedure then continued until there was a correct response after a negative response. The second iteration started 4 dB nCL above the last presentation level and the procedure was repeated identically. This procedure was employed in U-ACT-4

Post-processing and threshold estimates.

For each test run, the levels of the tracking variable and the participant responses constituted a trace. In the context of adaptive testing procedures, a trace is a step-by-step history of the tracking variable from the start and until it converges to the threshold. The trace was fitted to a psychometric function [75]:

ψ=1001+e4s50·(x50x),

where the slope (s50) and midpoint (x50) are the parameters, x is the contrast level and ψ is the % correct. In the fitting process, the slope was kept fixed (s50= 0.2518  dB1) and x was defined in the closed interval [-8, 16] with a step size of 0.1. The variable x was thus allowed to go below the minimum value used with ACT (-4 dB nCL) to ensure accurate representation of the entire psychometric function, particularly for cases where threshold levels were very low. For all the ACT tests (manual and user-operated) the ACT value was taken as the ψ=70.216 %-point on the psychometric function, following the recommendations from Zaar/Simonsen et al. [35]. The STM benchmark followed the same procedure as in the Zaar et al. [28,34] studies (i.e., a transformed 2-down 1-up procedure [29] with a final step size of 1 dB). However, the test procedure of the STM was extended by 8 more reversals compared to the studies of Zaar et al. [28,34] which allows an investigation of the likely effects of fatigue (not considered here). The threshold estimate (STMth) was the resulting mean across the values of the first 8 reversals (4 upper and 4 lower) at a step size of 1 dB In this case, the threshold corresponds to the ψ=70.7%-point [29].

Non-verbal instructions

The instructions consisted of a succession of assignments and were designed so they only contained non-verbal audiovisual information [76]. The participants had to go through four instruction phases, each of which comprised a specific assignment and a goal that they had to achieve by themselves. However, the participants could make use of different types of support during the instruction phases. 1) Help: If the user pressed a question mark in the bottom left corner of the touch screen, minimalistic written information was added to the screen. 2) Written instructions: If the user pressed the blue “information” button, thorough written instruction appeared on the screen. 3) Call the examiner: The user always had the option to call the examiner for task explanation. The participants could go backwards and forwards and were not obliged to complete the assignment that was presented in each step of the instruction phase. A cartoon of the instructions and main visual elements is depicted in Fig 7.

Fig 7. User workflow for the U-ACT test.

Fig 7

The participant progresses from a welcome screen through headphone setup, sound familiarization, and training, before starting the test demonstration. The interface uses an “Ambulance” as the target-sound icon and a “Wave” as the reference icon, with on-screen help always available. The participant can make use of 3 types of help: extra information, written instructions, or call the examiner to obtain verbal instructions.

The instructions aimed at achieving four goals for the participant:

  1. Correct headphone placement. The screen displayed an image of the headphones indicating which side corresponds to which ear.

  2. Become familiar with the stimuli: The screen showed an ambulance and a wave, each of them with a button with an icon representing a loudspeaker.

  3. Try out a detection test. The screen showed three buttons. The button with a speaker was enabled and played a sound when pressed. The participant was then required to press one of the other two buttons (ambulance or wave). After the button was pressed by the participant, a sign appeared on the screen indicating whether their choice was correct or incorrect.

  4. Watch a demonstration. The participants should watch a short video showing how to perform the task.

The use of the instructions was investigated in block 1. All the events (i.e., next phase starts, go forward, go back, etc.), button presses, and response times were captured by the system.

Subjective evaluation

The evaluation was carried out using the touch screen. The data collection process involved the use of a visual analogue continuous scale (VAS), following each run in blocks 2, 3, and 4. On the screen, a statement was written starting always with “The test was …” and below it, a continuous slider served to collect the judgments. To mimic a Likert scale, the slide had the three labels: “strongly disagree”, “neutral” and “strongly agree” and the result was a number between -50 and 50. The collected attributes included difficulty, self-perceived duration, demanding, boredom, as well as evaluations on the test’s outcome such as understanding the instructions, usefulness, goodness of results, and overall impression. These attributes were assessed twice: once using positive statements like “The test was easy”, and another time employing negative statements such as “The test was difficult”.

Analyses

Statistical analysis of threshold differences.

All statistical analyses were carried out using linear mixed-effects models and based on an analysis of variance (ANOVA) with subsequent post-hoc analyses of pairwise differences. The analyses were run in RStudio (Version 12; Posit team, 2024) using the packages lmerTest and emmeans. More details about the investigations can be found in the supplementary material.

For the investigation of the training effects and the effects of task and procedure on the thresholds, we made use of the difference between the U-ACT test candidate results and the benchmark ACT value of each individual participant (denoted ∆ACT). This allowed a fairer comparison of the training effects. For this purpose, block 1 was disregarded since the purpose of that block was to evaluate the instructions and the participants did not have any previous experience with the test candidates at this point. The effect of the remaining repetitions was included in the factor block.

Statistical models included all the possible fixed effects and their interaction with the participant group while participant ID was included as a random effect. The models had the following form in R notation:

∆ACT ~ group*(task + procedure + block) + (1 | ID).

Estimated marginal means (EMMs) for the fixed factors were calculated with Satterthwaite’s degrees-of-freedom approximation. Pairwise contrasts between task conditions were evaluated using Bonferroni-corrected t-tests to account for multiple comparisons. In case of significant interactions, an analysis of the subset of the data related to specific conditions or groups was performed.

Test-retest reliability metrics.

To evaluate the test-retest performance of the test candidates and the agreement with the test benchmarks, the following metrics were considered:

  • Intra-Class Correlation (ICC) [44,48]: This metric is commonly used in test-retest reliability studies. It measures consistency and agreement across different occasions, conditions or procedures, on the same subjects and it makes use of mixed models. For the measure of absolute test-retest reliability, we employed a 2-way random effect model (2 different measurements) with a single rater measurement (the measurement is not the average of various repetitions). In the manuscript we refer to this as ICC(2,1). For the comparisons between tests, we employed a 2-way random effects model, to quantify the absolute agreement, among multiple repetitions. In the manuscript we refer to this as ICC(2,k) [44,48].

  • Mean difference and Coefficient of Repeatability (CoR) [45,47]: These are obtained from Bland & Altman plots used, for example, in the middle panel in Fig 3, where the mean difference represents the bias, and the CoR defines the limits within which 95% of the differences between paired measurements lie. Mathematically, CoR=1.962 *  within subject standard deviation.

  • Fractional Rank Precision (FRP) [77]: This is a metric used to assess the precision of a method in ranking objects or samples and thus related to the mean average precision, which is a standard tool for information retrieval. The higher the FRP value, the better the precision in rankings.

Postprocessing and data modelling of subjective evaluation data.

Prior to the analysis, data from the subjective evaluations underwent preprocessing to account for variations in the use of the VAS scale among participants and to focus on the differences between U-ACT test candidates. To standardize the ratings, each raw judgement (ratingi,b) was transformed using a participant-block-specific Z-score:

Zratingi,b= ratingi,b μi,b σi,b,

where μi,b  and σi,b represent the mean and standard deviation, respectively, calculated across all rated attributes but separately for each participant i within each experimental block b.

Additionally, both types of wording for each attribute (positive and negative wording) were transformed so that they shared the same sign and represented by the factor wording. For example, difficulty remained unchanged while easiness was multiplied by -1. Subsequently, linear models were created for each attribute in the following format:

Attribute = group + candidate + wording

However, the group factor was not significant in any of the attribute models and was thus removed. The analysis was followed by pairwise contrasts between UACT test candidates using Bonferroni-corrected t-tests for comparison. Note that we use a linear model here with no random effects because of the standardization.

Acknowledgments

This work was carried out at the Technical University of Denmark. We thank T. Dau, C. van Oosterhout and R.S. Sørensen for their support. The participants for their insights, time and effort, and L.B. Simonsen for her contributions to the custom procedure and J. Undurraga for his help with model diagnostics. The reverse digit span was originally implemented by J. Hjortkjaer and J. Märcher-Rørsted. We want to acknowledge the discussion among the UAud project partners at SDU and OUH, and several colleagues from Demant A/S, and DTU who tried out earlier versions of the tests, especially S. Santurette, D. Mistry.

Artificial Intelligence statement: During the preparation of this work, the authors used Gemini (version 2.5 Pro) and Mistral (version Nemo Instruct) to enhance readability and correct grammatical and typographical errors. The authors reviewed and edited all AI-generated suggestions and take full responsibility for the final content of the manuscript.

Data Availability

The data and code used in the analyses of this manuscript is available in the Zenodo repository https://doi.org/10.5281/zenodo.15730189. Most additional data can be made available on request.

Funding Statement

This work was supported by Innovation Fund Denmark (https://innovationsfonden.dk/) Grand Solutions 9090-00089B (UAud project) and the William Demant Foundation case number 19-1251 (https://www.williamdemantfonden.dk/). SL and JZ are permanent employees of Interacoustics A/S and Demant A/S respectively. The employment of RS-L and consultancy services, dissemination, APCs, travel expenses and materials were funded by the mentioned sources. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.GBD 2019 Hearing Loss Collaborators. Hearing loss prevalence and years lived with disability, 1990-2019: findings from the Global Burden of Disease Study 2019. Lancet. 2021;397(10278):996–1009. doi: 10.1016/S0140-6736(21)00516-X [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Shen Y, Zhou T, Zou W, Zhang J, Yan S, Ye H, et al. Global, regional, and national burden of hearing loss from 1990 to 2021: findings from the 2021 global burden of disease study. Ann Med. 2025;57(1):2527367. doi: 10.1080/07853890.2025.2527367 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Clark JL, Swanepoel DW. The World Report on Hearing - a new era for global hearing care. Int J Audiol. 2021;60(3):161. doi: 10.1080/14992027.2021.1881318 [DOI] [PubMed] [Google Scholar]
  • 4.Garuccio J, Ukert B, Arnold M, Phillips S, Pesko MF. Using supply and demand to identify shortages in the hearing health care professional workforce. JAMA Otolaryngol Head Neck Surg. 2025;151(9):868–73. doi: 10.1001/jamaoto.2025.2112 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Dillard LK, Der CM, Laplante-Lévesque A, Swanepoel DW, Thorne PR, McPherson B, et al. Service delivery approaches related to hearing aids in low- and middle-income countries or resource-limited settings: a systematic scoping review. PLOS Glob Public Health. 2024;4(1):e0002823. doi: 10.1371/journal.pgph.0002823 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Margolis RH, Morgan DE. Automated pure-tone audiometry: an analysis of capacity, need, and benefit. Am J Audiol. 2008;17:109–13. doi: 10.1044/1059-0889(2008/07-004 [DOI] [PubMed] [Google Scholar]
  • 7.Windmill IM, Freeman BA. Demand for audiology services: 30-yr projections and impact on academic programs. J Am Acad Audiol. 2013;24:407–16. doi: 10.3766/jaaa.24.5.7 [DOI] [PubMed] [Google Scholar]
  • 8.Parmar BJ, Rajasingam SL, Bizley JK, Vickers DA. Factors affecting the use of speech testing in adult audiology. Am J Audiol. 2022;31(3):528–40. doi: 10.1044/2022_AJA-21-00233 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Sidiras C, Sanchez-Lopez R, Pedersen ER, Sørensen CB, Nielsen J, Schmidt JH. User-Operated Audiometry Project (UAud) - introducing an automated user-operated system for audiometric testing into everyday clinic practice. Front Digit Health. 2021;3:724748. doi: 10.3389/fdgth.2021.724748 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Gates GA, Mills JH. Presbycusis. The Lancet. 2005;366:1111–20. [DOI] [PubMed] [Google Scholar]
  • 11.Houtgast T, Festen JM. On the auditory and cognitive functions that may explain an individual’s elevation of the speech reception threshold in noise. Int J Audiol. 2008;47(6):287–95. doi: 10.1080/14992020802127109 [DOI] [PubMed] [Google Scholar]
  • 12.Rönnberg J, Lunner T, Ng EHN, Lidestam B, Zekveld AA, Sörqvist P, et al. Hearing impairment, cognition and speech understanding: exploratory factor analyses of a comprehensive test battery for a group of hearing aid users, the n200 study. Int J Audiol. 2016;55(11):623–42. doi: 10.1080/14992027.2016.1219775 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Schoof T, Rosen S. The role of auditory and cognitive factors in understanding speech in noise by normal-hearing older listeners. Front Aging Neurosci. 2014;6:307. doi: 10.3389/fnagi.2014.00307 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Livingston G, Huntley J, Sommerlad A, Ames D, Ballard C, Banerjee S, et al. Dementia prevention, intervention, and care: 2020 report of the Lancet Commission. Lancet. 2020;396(10248):413–46. doi: 10.1016/S0140-6736(20)30367-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Cantuaria ML, Pedersen ER, Waldorff FB, Wermuth L, Pedersen KM, Poulsen AH, et al. Hearing loss, hearing aid use, and risk of dementia in older adults. JAMA Otolaryngol Head Neck Surg. 2024;150(2):157–64. doi: 10.1001/jamaoto.2023.3509 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Mata R, Schooler LJ, Rieskamp J. The aging decision maker: cognitive aging and the adaptive selection of decision strategies. Psychol Aging. 2007;22(4):796–810. doi: 10.1037/0882-7974.22.4.796 [DOI] [PubMed] [Google Scholar]
  • 17.Contreras-Somoza LM, Irazoki E, Toribio-Guzmán JM, de la Torre-Díez I, Diaz-Baquero AA, Parra-Vidales E, et al. Usability and user experience of cognitive intervention technologies for elderly people with MCI or dementia: a systematic review. Front Psychol. 2021;12:636116. doi: 10.3389/fpsyg.2021.636116 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Zhang Y, Zhu X, Yang S, Wong AKC, Chen X. Digital health technologies applied in patients with early cognitive change: scoping review. J Med Internet Res. 2025;27:e82881. doi: 10.2196/82881 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Pastucha M, Gos E, Kochanek K, Skarżyński H, Jedrzejczak WW. Usability of a hearing test mobile app across generations. PLoS One. 2025;20(7):e0327726. doi: 10.1371/journal.pone.0327726 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Grant KW, Walden BE, Summers V, Leek MR. Introduction: auditory models of suprathreshold distortion in persons with impaired hearing. J Am Acad Audiol. 2013;24(4):254–7. doi: 10.3766/jaaa.24.4.2 [DOI] [PubMed] [Google Scholar]
  • 21.Plomp R. Auditory handicap of hearing impairment and the limited benefit of hearing aids. J Acoust Soc Am. 1978;63(2):533–49. doi: 10.1121/1.381753 [DOI] [PubMed] [Google Scholar]
  • 22.Keidser G, et al. The quest for ecological validity in hearing science: what it is, why it matters, and how to advance it. Ear Hear. 2020;41:5S. doi: 10.1097/AUD.0000000000000944 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Akeroyd MA, et al. International Collegium of Rehabilitative Audiology (ICRA) recommendations for the construction of multilingual speech tests: ICRA Working Group on Multilingual Speech Tests. Int J Audiol. 2015;54:17–22. doi: 10.3109/14992027.2015.1030513 [DOI] [PubMed] [Google Scholar]
  • 24.Johannesen PT, Pérez-González P, Kalluri S, Blanco JL, Lopez-Poveda EA. The influence of cochlear mechanical dysfunction, temporal processing deficits, and age on the intelligibility of audible speech in noise for hearing-impaired listeners. 2016. doi: 10.1177/2331216516641055 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Strelcyk O, Dau T. Relations between frequency selectivity, temporal fine-structure processing, and speech reception in impaired hearing. J Acoust Soc Am. 2009;125(5):3328–45. doi: 10.1121/1.3097469 [DOI] [PubMed] [Google Scholar]
  • 26.Bernstein JGW, Mehraei G, Shamma S, Gallun FJ, Theodoroff SM, Leek MR. Spectrotemporal modulation sensitivity as a predictor of speech intelligibility for hearing-impaired listeners. J Am Acad Audiol. 2013;24(4):293–306. doi: 10.3766/jaaa.24.4.5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Chi T, Gao Y, Guyton MC, Ru P, Shamma S. Spectro-temporal modulation transfer functions and speech intelligibility. J Acoust Soc Am. 1999;106(5):2719–32. doi: 10.1121/1.428100 [DOI] [PubMed] [Google Scholar]
  • 28.Zaar J, Simonsen LB, Dau T, Laugesen S. Toward a clinically viable spectro-temporal modulation test for predicting supra-threshold speech reception in hearing-impaired listeners. Hear Res. 2023;427:108650. doi: 10.1016/j.heares.2022.108650 [DOI] [PubMed] [Google Scholar]
  • 29.Levitt H. Transformed up-down methods in psychoacoustics. J Acoust Soc Am. 1971;49(2):Suppl 2:467+. [PubMed] [Google Scholar]
  • 30.Pelli DG, Bex P. Measuring contrast sensitivity. Vision Res. 2013;90:10–4. doi: 10.1016/j.visres.2013.04.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Dorr M, et al. Next-generation vision testing: the quick CSF. Curr Dir Biomed Eng. 2015;1:131–4. doi: 10.1515/CDBME-2015-0034 [DOI] [Google Scholar]
  • 32.Mehraei G, Gallun FJ, Leek MR, Bernstein JGW. Spectrotemporal modulation sensitivity for hearing-impaired listeners: dependence on carrier center frequency and the relationship to speech intelligibility. J Acoust Soc Am. 2014;136(1):301–16. doi: 10.1121/1.4881918 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Bernstein JGW, Danielsson H, Hällgren M, Stenfelt S, Rönnberg J, Lunner T. Spectrotemporal modulation sensitivity as a predictor of speech-reception performance in noise with hearing aids. Trends Hear. 2016;20:2331216516670387. doi: 10.1177/2331216516670387 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Zaar J, Simonsen LB, Laugesen S. A spectro-temporal modulation test for predicting speech reception in hearing-impaired listeners with hearing aids. Hear Res. 2024;443:108949. doi: 10.1016/j.heares.2024.108949 [DOI] [PubMed] [Google Scholar]
  • 35.Zaar J, Simonsen LB, Sanchez-Lopez R, Laugesen S. The Audible Contrast Threshold (ACT) test: a clinical spectro-temporal modulation detection test. Hear Res. 2024;453:109103. doi: 10.1016/j.heares.2024.109103 [DOI] [PubMed] [Google Scholar]
  • 36.Margolis RH, Glasberg BR, Creeke S, Moore BCJ. AMTAS: automated method for testing auditory sensitivity: validation studies. Int J Audiol. 2010;49(3):185–94. doi: 10.3109/14992020903092608 [DOI] [PubMed] [Google Scholar]
  • 37.Larusdottir MK, Gulliksen J, Hallberg N. RAMES – Framework supporting user centred evaluation in research and practice. Behav Inf Technol. 2019;38:132–49. doi: 10.1080/0144929X.2018.1519034 [DOI] [Google Scholar]
  • 38.ISO/IEC 25022: Systems and software engineering – Systems and software Quality Requirements and Evaluation (SQuaRE) - Measurement of quality in use. ISO/IEC. 2016. [Google Scholar]
  • 39.blackburn HL, benton AL. Revised administration and scoring of the digit span test. J Consult Psychol. 1957;21(2):139–43. doi: 10.1037/h0047235 [DOI] [PubMed] [Google Scholar]
  • 40.Fuglsang SA, Märcher-Rørsted J, Dau T, Hjortkjær J. Effects of sensorineural hearing loss on cortical synchronization to competing speech during selective attention. J Neurosci. 2020;40(12):2562–72. doi: 10.1523/JNEUROSCI.1936-19.2020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Ericsson KA, Simon HA. How to study thinking in everyday life: contrasting think-aloud protocols with descriptions and explanations of thinking. Mind Cult Act. 1998;5:178–86. doi: 10.1207/s15327884mca0503_3 [DOI] [Google Scholar]
  • 42.Carhart R, Jerger JF. Preferred method for clinical determination of pure-tone thresholds. J Speech Hear Disord. 1959;24:330–45. doi: 10.1044/jshd.2404.330 [DOI] [Google Scholar]
  • 43.Jebb AT, Ng V, Tay L. A review of key Likert scale development advances: 1995–2019. Front Psychol. 2021;12. doi: 10.3389/fpsyg.2021.637547 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Koo TK, Li MY. A Guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016;15(2):155–63. doi: 10.1016/j.jcm.2016.02.012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Bartlett JW, Frost C. Reliability, repeatability and reproducibility: analysis of measurement errors in continuous variables. Ultrasound Obstet Gynecol. 2008;31(4):466–75. doi: 10.1002/uog.5256 [DOI] [PubMed] [Google Scholar]
  • 46.Bland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. Int J Nurs Stud. 2010;6. doi: 10.1016/j.ijnurstu.2009.10.001 [DOI] [PubMed] [Google Scholar]
  • 47.Giavarina D. Understanding Bland Altman analysis. Biochem Med (Zagreb). 2015;25(2):141–51. doi: 10.11613/BM.2015.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Dorr M, Elze T, Wang H, Lu Z-L, Bex PJ, Lesmes LA. New precision metrics for contrast sensitivity testing. IEEE J Biomed Health Inform. 2018;22(3):919–25. doi: 10.1109/JBHI.2017.2708745 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Sørensen CB, Pedersen C, Pedersen ER, Schmidt JH, Nielsen J. Observed issues and their prevalence when patients conduct user-operated audiometry: International Symposium on Auditory and Audiological Research. The Conference was ISAAR2023, Nyborg, Denmark, 2023.
  • 50.Aronoff JM, Landsberger DM. The development of a modified spectral ripple test. J Acoust Soc Am. 2013;134(2):EL217-22. doi: 10.1121/1.4813802 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Landsberger DM, Dwyer RT, Stupak N, Gifford RH. Validating a quick spectral modulation detection task. Ear Hear. 2019;40(6):1478–80. doi: 10.1097/AUD.0000000000000713 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Lelo de Larrea-Mancera ES, Stavropoulos T, Hoover EC, Eddins DA, Gallun FJ, Seitz AR. Portable Automated Rapid Testing (PART) for auditory assessment: validation in a young adult normal-hearing population. J Acoust Soc Am. 2020;148(4):1831. doi: 10.1121/10.0002108 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Lelo de Larrea-Mancera ES, Koerner TK, Bologna WJ, Momtaz S, Menon KN, Carrillo A, et al. At-home auditory assessment using Portable Automated Rapid Testing (PART) to understand self-reported hearing difficulties. Trends Hear. 2025;29:23312165251397373. doi: 10.1177/23312165251397373 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Stavropoulos TA, Isarangura S, Hoover EC, Eddins DA, Seitz AR, Gallun FJ. Exponential spectro-temporal modulation generation. J Acoust Soc Am. 2021;149(3):1434. doi: 10.1121/10.0003604 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Lelo de Larrea-Mancera ES, Stavropoulos T, Carrillo AA, Cheung S, He YJ, Eddins DA, et al. Remote auditory assessment using Portable Automated Rapid Testing (PART) and participant-owned devices. J Acoust Soc Am. 2022;152(2):807. doi: 10.1121/10.0013221 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Lelo de Larrea-Mancera ES, Stavropoulos T, Carrillo AA, Menon KN, Hoover EC, Eddins DA, et al. Validation of the adaptive scan method in the quest for time-efficient methods of testing auditory processes. Atten Percept Psychophys. 2023;85(8):2797–810. doi: 10.3758/s13414-023-02743-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Shojaeemend H, Ayatollahi H. Automated audiometry: a review of the implementation and evaluation methods. Healthc Inform Res. 2018;24:263. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Eikelboom RH, Swanepoel DW, Motakef S, Upson GS. Clinical validation of the AMTAS automated audiometer. Int J Audiol. 2013;52(5):342–9. doi: 10.3109/14992027.2013.769065 [DOI] [PubMed] [Google Scholar]
  • 59.Margolis RH, Killion MC, Bratt GW, Saly GL. Validation of the Home Hearing Test™. J Am Acad Audiol. 2016;27(5):416–20. doi: 10.3766/jaaa.15102 [DOI] [PubMed] [Google Scholar]
  • 60.Pedersen C, et al. Comparison of hearing-aid effectiveness based on user-operated versus traditional audiometry: a randomised clinical trial. Int J Audiol. 2024;:1–8. doi: 10.1080/14992027.2024.2434897 [DOI] [PubMed] [Google Scholar]
  • 61.Voutilainen A, Pitkäaho T, Kvist T, Vehviläinen-Julkunen K. How to ask about patient satisfaction? The visual analogue scale is less vulnerable to confounding factors and ceiling effect than a symmetric Likert scale. J Adv Nurs. 2016;72(4):946–57. doi: 10.1111/jan.12875 [DOI] [PubMed] [Google Scholar]
  • 62.Sung Y-T, Wu J-S. The Visual Analogue Scale for Rating, Ranking and Paired-Comparison (VAS-RRP): a new technique for psychological measurement. Behav Res Methods. 2018;50(4):1694–715. doi: 10.3758/s13428-018-1041-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.García-Pérez MA, Alcalá-Quintana R. Accuracy and precision of responses to visual analog scales: inter- and intra-individual variability. Behav Res Methods. 2023;55(8):4369–81. doi: 10.3758/s13428-022-02021-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Kersten P, White PJ, Tennant A. Is the pain visual analogue scale linear and responsive to change? An exploration using Rasch analysis. PLoS One. 2014;9(6):e99485. doi: 10.1371/journal.pone.0099485 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Nasreddine ZS, Phillips NA, Bédirian V, Charbonneau S, Whitehead V, Collin I, et al. The Montreal Cognitive Assessment, MoCA: a brief screening tool for mild cognitive impairment. J Am Geriatr Soc. 2005;53(4):695–9. doi: 10.1111/j.1532-5415.2005.53221.x [DOI] [PubMed] [Google Scholar]
  • 66.Diniz BSO, Yassuda MS, Nunes PV, Radanovic M, Forlenza OV. Mini-mental State Examination performance in mild cognitive impairment subtypes. Int Psychogeriatr. 2007;19(4):647–56. doi: 10.1017/S104161020700542X [DOI] [PubMed] [Google Scholar]
  • 67.Poling GL, Kunnel TJ, Dhar S. Comparing the accuracy and speed of manual and tracking methods of measuring hearing thresholds. Ear Hear. 2016;37(5):e336-40. doi: 10.1097/AUD.0000000000000317 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Mahomed F, Swanepoel DW, Eikelboom RH, Soer M. Validity of automated threshold audiometry: a systematic review and meta-analysis. Ear Hear. 2013;34(6):745–52. doi: 10.1097/01.aud.0000436255.53747.a4 [DOI] [PubMed] [Google Scholar]
  • 69.Nielsen J, Landauer TK. A mathematical model of the finding of usability problems. In: Proceedings of the SIGCHI conference on Human factors in computing systems - CHI ’93, 1993. p. 206–13. 10.1145/169059.169166 [DOI]
  • 70.MIkhalevskaya MB, FInkel’ NV. The influence of feedback on the detectability of a signal. Sov Psychol. 1986;25:50–60. doi: 10.2753/RPO1061-0405250150 [DOI] [Google Scholar]
  • 71.Hoglund EM, Feth LL. Bayesian analysis of the impact of incomplete feedback on response bias. In: Sydney, Australia, 2023. 10.1121/2.0001851 [DOI]
  • 72.Suh MJ, et al. Improving accuracy and reliability of hearing tests: an exploration of international standards. J Audiol Otol. 2023;27:169–80. PMC10603284 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Remus JJ, Collins LM. Comparison of adaptive psychometric procedures motivated by the theory of optimal experiments: simulated and experimental results. J Acoust Soc Am. 2008;123(1):315–26. doi: 10.1121/1.2816567 [DOI] [PubMed] [Google Scholar]
  • 74.Brand T, Kollmeier B. Efficient adaptive procedures for threshold and concurrent slope estimates for psychophysics and speech intelligibility tests. J Acoust Soc Am. 2002;111(6):2801–10. doi: 10.1121/1.1479152 [DOI] [PubMed] [Google Scholar]
  • 75.Zabalbeascoa P. The nature of the audiovisual text and its parameters. Benjamins Translation Library. John Benjamins Publishing Company; 2008. p. 21–37. doi: 10.1075/btl.77.05zab [DOI] [Google Scholar]
  • 76.McGraw KO, Wong SP. Forming inferences about some intraclass correlation coefficients. Psychol Methods. 1996;1:30–46. doi: 10.1037/1082-989X.1.1.30 [DOI] [Google Scholar]
  • 77.Margolis RH, Saly GL, Le C, Laurence J. Qualind: a method for assessing the accuracy of automated tests. J Am Acad Audiol. 2007;18(1):78–89. doi: 10.3766/jaaa.18.1.7 [DOI] [PubMed] [Google Scholar]
PLOS Digit Health. doi: 10.1371/journal.pdig.0000975.r001

Decision Letter 0

Zhao Ni, Michael Winter

1 Dec 2025

-->PDIG-D-25-00510-->-->Evaluating user-operated tests of audible contrast sensitivity towards efficient hearing healthcare services-->-->PLOS Digital Health-->--> -->-->Dear Dr. Sanchez-Lopez,-->--> -->-->Thank you for submitting your manuscript to PLOS Digital Health. After careful consideration, we feel that it has merit but does not fully meet PLOS Digital Health's publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.-->--> -->-->Please submit your revised manuscript by Jan 30 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at digitalhealth@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pdig/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.-->--> -->-->Please include the following items when submitting your revised manuscript:-->-->* A rebuttal letter that responds to each point raised by the editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. This file does not need to include responses to any formatting updates and technical items listed in the 'Journal Requirements' section below.-->-->* A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.-->-->* An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.-->--> -->-->If you would like to make changes to your financial disclosure, competing interests statement, or data availability statement, please make these updates within the submission form at the time of resubmission. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.-->--> -->-->We look forward to receiving your revised manuscript.-->--> -->-->Kind regards,-->--> -->-->Michael Winter-->-->Academic Editor-->-->PLOS Digital Health-->--> -->-->Leo Anthony Celi-->-->Editor-in-Chief-->-->PLOS Digital Health-->-->orcid.org/0000-0001-6712-6626-->--> -->-->Journal Requirements: -->--> -->-->If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. -->--> -->-->Additional Editor Comments (if provided): -->--> -->--> -->--> -->-->[Note: HTML markup is below. Please do not edit.]-->--> -->-->Reviewers' Comments: -->--> -->-->Reviewer's Responses to Questions

-->Comments to the Author

1. Does this manuscript meet PLOS Digital Health’s publication criteria? Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe methodologically and ethically rigorous research with conclusions that are appropriately drawn based on the data presented.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->2. Has the statistical analysis been performed appropriately and rigorously?-->

Reviewer #1: I don't know

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->3. Have the authors made all data underlying the findings in their manuscript fully available (please refer to the Data Availability Statement at the start of the manuscript PDF file)?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception. The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS Digital Health does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: No

**********

-->5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #1: This research and invention focused on overlooked area of care, and it has the potential to significantly enhance the accessibility of hearing loss diagnosis and treatment. I, therefore, understand the significance of this paper. The paper is crafted carefully and the experiments were conducted systematically and thoroughly. However, there are critical issues that are missed in this paper. First, order or presentation make the paper difficult to understand and fully evaluate. the methodology should come before the results, discussion and conclusion, which makes reading and comprehension of the paper easier. However, on this paper the methodology was the last part of the paper. Second, the sample technique wasn't explicitly mentioned, and it is not clear what measures were taken to ensure the randomness of the sample. Third, there were arbitrary thresholds without sufficient justifications. For example, in 1-button task, 1.6 second was used to categorize miss and hit, however, there is no sufficient explanation why this number was selected and what was the investigator try to achieve by timing the tasks. In addition, subsequent tasks, such as yes/no task and 3-AFC tasks, weren't timed, but adequate justifications weren't provided why these tasks weren't timed. Fourth, the discussion section of the paper was too sparse and the findings of the experiments weren't deeply explored or explained further in the context of existing literature. I also think the diversity of the participants, real clinical setting testing and the use of multi-dimensional protocols was the strength of the paper.

Reviewer #2: Thank you for giving me the opportunity to review your article.

1. The authors mentioned that there were "Some insights emerged, such as

perceived differences in test difficulty and length among candidates, and a difference in

results based on attribute wording which had a strong and clear significance in user

evaluations." (line 331). It might be worth elaborating if these differences were the result of candidates having different hearing/cognitive abilities, or if the distribution of candidates who have different experiences in test difficulty and length is similar across the different hearing/cognitive abilities groups.

2. The authors mentioned that the participants had an option to "Call the examiner: The user

always had the option to call the examiner for task explanation". It might be worth stating whether any participants called the examiner for help, and if it the calls for help affected the rating of the test experience.

Reviewer #3: This paper develops and evaluates user-operated alternatives to the ACT test. While it is relevant to PLOS Digital Health and the research overall does seem sound, I have several significant issues with the work that I don’t think can be addressed with a revision. Specifically, I think the paper needs significant rewrites, and additional experiments need to be run before I can recommend the paper for publication. Please note that the examples I use to illustrate my points are not exhaustive; there are usually more cases of the issues I point out.

Poor writing and organization:

First, this paper is rather difficult to follow at both a structural and a sentence level. Regarding the latter, while sentences are not grammatically incorrect, there is a lot of awkward phrasing that makes for a choppy and, at times, verbose read. Just taking the sentence beginning at the end of line 409 as an example, “We aimed at a total of 30 participants where there would be a control group formed of 6 young NH listeners (<30 years old), a group of people with hearing loss and with high cognitive abilities (i.e., higher scores in the RDS test, HIhi.), and another group of participants with hearing loss with lower cognitive abilities (i.e., lower scores in the RDS test HIlo)” could much more cleanly rephrased as “We aimed to recruit 30 participants for this study. This comprised 6 NH listeners, 12 listeners with hearing loss and high cognitive abilities (Hihi), and 12 listeners with hearing loss and low cognitive abilities (Hilo). Cognitive ability was determined based on RDS scores, with the high group scoring between 14 and 21 and the low group scoring between 6 and 12.”

Secondly, definitions for concepts that need to be explained frequently come after their use, so I had to read the sentence, find the definition, then go back and read the sentence again to understand what the authors were getting at. One example is lines 46-52, where the authors introduce the ACT test as a version of the STM test, neither of which is really introduced or defined (more on this later). They both involve speech reception thresholds, which are defined in the next sentence, meaning that a reader has to start this whole section over again to understand what they are reading, with the context from the later sentences. As another example, on line 91, RDS scores are introduced, but it's only in the Methods that they are noted as a test for cognitive ability (lines 407-408). As a final example here, the RAMES framework is mentioned for the first time on line 356, and even the acronym is not spelled out until line 391, where it’s just cited and not defined. I bring up these sentence-level issues to highlight the structural issues with the paper, namely that it’s not particularly accessible to readers who aren’t intimately familiar with this specific area of research. In general, it seems as though the authors struggled to write a paper where the results preceded the methods, as I frequently had to go to the Methods section to understand what was being presented in the results (this is highlighted by the fact that several times the authors stated “(see Methods section)” after portions of the Results section) and the Results section is frequently bloated by half descriptions of the Methods employed to achieve the Results. As examples, lines 87-90, 91-93, 116-122, and the entire section starting from 146 are Methods, not results. The authors need to go back and firmly delineate what goes into Methods and what goes into Results.

In general, I feel like the paper needs a significant rewrite for clarity. In particular, I would recommend that the authors read other papers that place the Results section before the Methods section and pay particular attention to how information is introduced and which sections contain which information. One particular point of emphasis: the authors really should clearly introduce the ACT and STM tests in the introduction, as these are foundational to their work, and without a clear understanding of those, readers who are not familiar with them are going to have a frustrating experience.

Study population issues:

I also have issues with the study population. I can understand why the population of hearing-impaired participants is low, as I can see how that would be a more difficult population to recruit. However, there are only 6 individuals who comprise the control group, and the only requirement for being a part of the control group is that they are below 30 years old and do not have hearing loss. There’s no reason for this group of participants to be that low, especially given that this study was conducted in a university setting. The authors should conduct the study with additional control group participants and re-run their analysis. Furthermore, is age just a proxy for hearing ability? Was one of the requirements for being in the hearing impaired group being over 50, and one of the requirements for being in the normal hearing group being under 30? Or was this just how recruitment panned out? The authors should clarify whether they recruited specifically for age or not, and if they did, they should justify why they did so.

Also, I have several issues with the inclusion of cognitive abilities. Why was this a point of emphasis? The authors should justify this with citations in the introduction. Also, why are RDS scores a fair evaluation of cognitive abilities? To me, this is a rather narrow view of cognitive ability, and there are many other tests; the authors need to justify the use of this.

A note on the language here. The statement “Since one of the study's goals was to obtain an evaluation from a diverse group of participants with varying hearing and cognitive abilities...” is confusing to me. Which of the five goals the authors laid out does this refer to? My best guess is Goal 4, but if that’s the case, evaluation on a diverse group of participants is a means of ensuring Goal 4; the goal itself is not to evaluate on a diverse group of participants. Really, I think that whole sentence can be removed, and the section can simply begin with something like “Figure 1 shows the results from…” which removes some of the Methods language from the Results section as I noted earlier. Also, the statement “We aimed at a total of 30 participants where there would be a control group formed by 6 young NH listeners” (lines 409-411) is confusing. Why did the authors aim to recruit only 30 participants? Why specifically 6 participants? Is this just a language issue, where a statement such as “We recruited a total of 30 participants that comprised 6 young NH listeners…” is what the authors meant?

Introduction issues:

The introduction generally lacks citations. There are only 10 citations throughout, and several claims really should be backed up with citations. For instance, this sentence: “The currently used audiometric tests, which are time-consuming and require active participation of a hearing care professional, represent a bottleneck that effectively limits access to hearing healthcare.” (lines 26-28), have health care professionals or patients noted that they are time consuming and represent a bottleneck? Also, this sentence: “However, there has been a growing focus on additionally evaluating supra-threshold abilities, such as the patient’s ability to discriminate speech in the presence of noise when the speech is fully audible.” (lines 38-40), are there any statistics/studies to show this shift in focus? Also, this sentence: “This has motivated the search for alternative tests capable of capturing the fundamental abilities employed in challenging 'cocktail-party’ scenarios, which are characterized by multiple interfering sound sources including speech.” (lines 43-45), are there other studies that have developed and evaluated alternative tests? Also, I think it might make the introduction stronger if the first sentence was separated into multiple sentences; the authors could include actual statistics on the scope and impact of hearing loss to really emphasize the need for improving accessibility and reducing cost.

Figure issues:

I don’t think the combination of the text, Figure 3, and Figure 3’s caption does a good job of explaining the differences between the test candidates. The authors should rework these to make it clearer what is changing between each candidate. In particular, I was looking to see if only one thing had changed between candidates, to understand what variable was being tested with each candidate, and it seems like that’s the case? I found this hard to follow.

For Figures 5-6, there’s quite a bit going on, and a lot of significance between various groups. Fundamentally, I don’t understand why positive and negative wordings were used; I appreciate that phrasing of statements can affect ratings, but there’s no meaningful follow-up on how this should influence how I interpret participants’ views on the candidates. For example, if I’m understanding correctly, in that first panel in Figure 5, there is a significant difference in the difficulty of candidates 2 and 3 when negative wording is used, but the difference is not significant when positive wording is used. So, is candidate 2 better in terms of difficulty? It just seems like these figures are introducing a lot of complexity, and all that the authors are getting out of it is a muddier picture. I will note that the point of a Discussion section is to contextualize results, and I was looking for some clarity there, but I did not find it. Also, I don’t understand what the word “wording” in the bottom left of each panel means. It also seems to cover up some data points in some panels.

Other issues:

Line 84: “A diverse group of participants” is a weird title for a subsection. I would recommend rephrasing to something like “Study participants”

Lines 87-93, the language is a bit confusing here. Initially, I read it as 28 HI participants were recruited, and an additional 6 NH participants were recruited, which made the division into 11 and 11 confusing. Please clean up the language here.

Figure 1 would be cleared if Right and Left were Right Ear and Left Ear.

Line 105, “two groups” should be “two HI groups”

Line 193, how many users were used for this refinement? These users were separate from the 28 participants that were used to evaluate the system, right?

Line 435, what iPad model was used?

**********

-->6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

Do you want your identity to be public for this peer review?  If you choose “no”, your identity will remain anonymous but your review may still be made public.

For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #1: No

Reviewer #2: No

Reviewer #3: No

**********

-->--> -->-->[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]-->--> -->-->Figure resubmission: -->--> -->--> -->While revising your submission, we strongly recommend that you use PLOS’s NAAS tool (https://ngplosjournals.pagemajik.ai/artanalysis) to test your figure files. NAAS can convert your figure files to the TIFF file type and meet basic requirements (such as print size, resolution), or provide you with a report on issues that do not meet our requirements and that NAAS cannot fix.--> -->

After uploading your figures to PLOS’s NAAS tool - https://ngplosjournals.pagemajik.ai/artanalysis, NAAS will process the files provided and display the results in the "Uploaded Files" section of the page as the processing is complete. If the uploaded figures meet our requirements (or NAAS is able to fix the files to meet our requirements), the figure will be marked as "fixed" above. If NAAS is unable to fix the files, a red "failed" label will appear above. When NAAS has confirmed that the figure files meet our requirements, please download the file via the download option, and include these NAAS processed figure files when submitting your revised manuscript.-->-->--> -->-->Reproducibility: -->--> -->-->To enhance the reproducibility of your results, we recommend that authors of applicable studies deposit laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols-->

PLOS Digit Health. doi: 10.1371/journal.pdig.0000975.r003

Decision Letter 1

Zhao Ni, Michael Winter, Zhao Ni, Michael Winter

30 Mar 2026

-->PDIG-D-25-00510R1-->

-->Evaluating user-operated tests of audible contrast sensitivity towards efficient hearing healthcare services-->

-->PLOS Digital Health-->

--> -->

-->Dear Dr. Sanchez-Lopez,-->

--> -->

-->Thank you for submitting your manuscript to PLOS Digital Health. After careful consideration, we feel that it has merit but does not fully meet PLOS Digital Health's publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.-->

--> -->

-->Please submit your revised manuscript by Apr 29 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at digitalhealth@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pdig/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.-->

--> -->

-->Please include the following items when submitting your revised manuscript:-->

-->* A letter that responds to each point raised by the editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. This file does not need to include responses to any formatting updates and technical items listed in the 'Journal Requirements' section below.-->

-->* A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.-->

-->* An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.-->

--> -->

-->If you would like to make changes to your financial disclosure, competing interests statement, or data availability statement, please make these updates within the submission form at the time of resubmission. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.-->

--> -->

-->We look forward to receiving your revised manuscript.-->

--> -->

-->Kind regards,-->

--> -->

-->Michael Winter-->

-->Academic Editor-->

-->PLOS Digital Health-->

--> -->

-->Leo Anthony Celi-->

-->Editor-in-Chief-->

-->PLOS Digital Health-->

-->orcid.org/0000-0001-6712-6626-->

--> -->

--> -->

-->Journal Requirements: -->

--> -->

-->If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. -->

--> -->

-->Additional Editor Comments (if provided): -->

--> -->

--> -->

--> -->

-->[Note: HTML markup is below. Please do not edit.]-->

--> -->

-->Reviewers' Comments: -->

--> -->

-->Reviewer's Responses to Questions

-->Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.-->

Reviewer #1: All comments have been addressed

Reviewer #2: (No Response)

Reviewer #3: (No Response)

**********

-->2. Does this manuscript meet PLOS Digital Health’s publication criteria? Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe methodologically and ethically rigorous research with conclusions that are appropriately drawn based on the data presented.-->

Reviewer #1: Yes

Reviewer #2: (No Response)

Reviewer #3: Yes

**********

-->3. Has the statistical analysis been performed appropriately and rigorously?-->

Reviewer #1: I don't know

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->4. Have the authors made all data underlying the findings in their manuscript fully available (please refer to the Data Availability Statement at the start of the manuscript PDF file)?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception. The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

-->5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS Digital Health does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: (No Response)

**********

-->6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #1: The structure of the paper is improved significantly, enhancing readability of the study. Sentence level issues are also resolved, and enabled easy understanding of the paper. Compelling clarifications were provided with ample citations which strengthened the discussions on the paper. The technical parts are also adequately explained.

Reviewer #2: I would like to thank the authors for their thoughtful and constructive response to the previous round of peer review.

An aim of the study is to ensure that "the test should be accessible to a broad group of patients without requiring any special modifications for subpopulations, including those with low cognitive abilities.", Upon reviewing the section on participants recruited, there are participants who are hearing-impaired with both higher and lower cognitive abilities. However there seems to be only participants with normal hearing and higher cognitive abilities, and there seems to be a lack of participants with normal hearing abilities and lower cognitive abilities. It would be good if the authors could elaborate on the choice of participants, especially as a key strength of the study mentioned was inclusion of "a diverse participant selection regarding hearing losses and cognitive abilities".

Thank you!

Reviewer #3: The authors have done an excellent job of responding to reviewer feedback and making appropriate revisions. I'm generally satisfied with the updated manuscript, but I do have a couple of comments.

1. While I appreciate the thought behind adding a new section after the introduction and before the results, it is effectively a methods section, and it is a bit strange to have two methods sections. I will note that while PLOS's submission guidelines do state that the preference is for the Methods section to come after the Discussion, it does state "To provide flexibility, however, authors are also able to include the Materials and Methods section before the Results section or before the Discussion section." Given there's quite a bit of context required to follow the Results section, I might recommend just combining the "Experimental design and study framework" section and "Methods" section into one section that follows the Introduction and precedes the Results. The Methods section is also quite long, and I would recommend considering moving non-essential definitions to supplemental materials. If the authors really would like to keep the methods after the discussion, I don't think adding a pseudo-methods section is the way to go; rather, the authors should work in the needed context into the introduction/results section.

2. A suggestion to improve readability: At the end of the introduction, the authors lay out five objectives, but these aren't referred to again after that. Given the number of results, it may help to refer back to these objectives in the results (and maybe also the discussion) so readers can quickly understand what objectives are being addressed by each finding. With this approach, I'd also recommend making the list of objectives in the introduction into an actual numbered list to give it the visual standout so that readers know to pay attention to it.

**********

-->7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

Do you want your identity to be public for this peer review?  If you choose “no”, your identity will remain anonymous but your review may still be made public.

For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #1: No

Reviewer #2: No

Reviewer #3: No

**********

-->

--> -->

-->[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]-->

--> -->

-->Figure resubmission: -->

--> -->

-->

-->While revising your submission, we strongly recommend that you use PLOS’s NAAS tool (https://ngplosjournals.pagemajik.ai/artanalysis) to test your figure files. NAAS can convert your figure files to the TIFF file type and meet basic requirements (such as print size, resolution), or provide you with a report on issues that do not meet our requirements and that NAAS cannot fix.-->

-->

After uploading your figures to PLOS’s NAAS tool - https://ngplosjournals.pagemajik.ai/artanalysis, NAAS will process the files provided and display the results in the "Uploaded Files" section of the page as the processing is complete. If the uploaded figures meet our requirements (or NAAS is able to fix the files to meet our requirements), the figure will be marked as "fixed" above. If NAAS is unable to fix the files, a red "failed" label will appear above. When NAAS has confirmed that the figure files meet our requirements, please download the file via the download option, and include these NAAS processed figure files when submitting your revised manuscript.-->

-->

--> -->

-->Reproducibility: -->

--> -->

-->To enhance the reproducibility of your results, we recommend that authors of applicable studies deposit laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols-->

PLOS Digit Health. doi: 10.1371/journal.pdig.0000975.r005

Decision Letter 2

Zhao Ni, Michael Winter, Zhao Ni, Michael Winter, Zhao Ni, Michael Winter

16 Jun 2026

Multi-dimensional evaluation of user-operated audible contrast sensitivity tests towards efficient hearing healthcare services

PDIG-D-25-00510R2

Dear Sanchez-Lopez,

We are pleased to inform you that your manuscript 'Multi-dimensional evaluation of user-operated audible contrast sensitivity tests towards efficient hearing healthcare services ' has been provisionally accepted for publication in PLOS Digital Health.

Before your manuscript can be formally accepted you will need to complete some formatting changes, which you will receive in a follow-up email from a member of our team.

Please note that your manuscript will not be scheduled for publication until you have made the required changes, so a swift response is appreciated.

IMPORTANT: The editorial review process is now complete. PLOS will only permit corrections to spelling, formatting or significant scientific errors from this point onwards. Requests for major changes, or any which affect the scientific understanding of your work, will cause delays to the publication date of your manuscript.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they'll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact digitalhealth@plos.org.

Thank you again for supporting Open Access publishing; we are looking forward to publishing your work in PLOS Digital Health.

Best regards,

Michael Winter

Academic Editor

PLOS Digital Health

***********************************************************

Additional Editor Comments (if provided):

Reviewer Comments (if any, and for reference):

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    Attachment

    Submitted filename: Response Letter_R1.pdf

    pdig.0000975.s002.pdf (220.3KB, pdf)
    Attachment

    Submitted filename: Response Letter_R2.pdf

    pdig.0000975.s003.pdf (108.2KB, pdf)

    Data Availability Statement

    The data and code used in the analyses of this manuscript is available in the Zenodo repository https://doi.org/10.5281/zenodo.15730189. Most additional data can be made available on request.


    Articles from PLOS Digital Health are provided here courtesy of PLOS

    RESOURCES