Abstract
No LLMs (Large Language Models) have yet been evaluated for understanding picture reports. Pure-tone audiograms, the gold standard for hearing loss assessment, are technical and often incomprehensible to patients without specialist interpretation. We conducted a blinded, multicenter evaluation of eight LLMs across diagnostic, interpretive, and recommendation tasks using 140 audiogram reports, assessed by clinicians and lay reviewers. The study revealed that DeepSeek-V3 achieved the highest diagnostic accuracy (severity: 67.00% ; type: 54.00%), R1 proved most suitable for general readership (FKGL: 6.41). The general public perceived significant benefits from all models in comprehension and emotional support, with Gemini 2.0 Flash/Thinking scoring higher. Challenges remain in understanding pathological mechanisms and controlling hallucinations. While current general-purpose LLMs cannot replace the diagnostic capabilities of physicians, they may serve as effective auxiliary tools for translating specialized audiogram data into structured, patient-accessible interpretations, with particular relevance for populations facing limited access to hearing-care services.
Subject terms: Health care, Medical research, Signs and symptoms
Introduction
As a globally prevalent health issue, hearing loss has been clearly defined as a significant public health concern1–3. The World Health Organization (WHO) estimates in its latest World Report on Hearing that nearly 2.5 billion people worldwide will experience some degree of hearing loss by 20504. Hearing loss is now recognized as the most prevalent sensory organ disability worldwide5, ranking third in the Global Burden of Disease (GBD) Years Lived with Disability (YLD) list, surpassed only by low back pain and migraine, and first among sensory impairments. Pure Tone Audiometry (PTA) is the gold standard for assessing hearing loss and functional testing6. Pure tone audiogram reports, generated from PTA results, are hearing health care graphical reports documenting a subject’s auditory responses to pure tone signals of varying frequencies and intensities. They serve as critical evidence for assessing auditory function, determining the degree, type, and location of hearing loss, and are widely used in otologic disease diagnosis, occupational health monitoring, and other fields7,8.
Currently, hearing health services face a dual challenge: on one hand, specialized hearing resources remain relatively scarce globally; on the other hand, traditional pure-tone audiogram reports, with their highly technical charts and concise diagnostic conclusions, often lack specific guidance and explanations tailored to individual patients. This leaves most individuals struggling to accurately understand their own hearing status and unable to receive clear recommendations on improving communication abilities9, such as through hearing aid use. The emergence of LLMs presents a potential solution. LLMs possess the capability to process and understand human language, as well as generate coherent question-answering responses10, and have played a positive role in enhancing patient-centered healthcare services. LLMs play an increasingly critical role in enhancing patient care quality (e.g., clinical decision support11, patient communication, and health education12) and optimizing healthcare processes (e.g., intelligent triage and guidance13,14, medical record, and coding automation15,16, appointment scheduling and resource allocation17). Their effective evaluation is therefore paramount.
Existing studies have explored LLM applications in interpreting reports across multiple medical specialties, such as radiology18, ophthalmology19,20, laboratory medicine21, and pathology22. However, according to benchmarks and tasks compiled by Li et al.23, prior research has predominantly focused on text summarization (TS) and dialog generation (DG) tasks, with minimal attention paid to evaluating general-purpose large models on medical image captioning (IC) tasks. This study will conduct the first investigation of general LLMs in otolaryngology-head and neck surgery (OHS) for IC tasks, focusing on whether LLMs can comprehend pure-tone audiogram reports and generate accurate, logical, and linguistically sound hearing interpretations and personalized recommendations that provide practical assistance to patients. This will systematically evaluate their potential value in aiding patients’ understanding of their hearing status and supporting hearing health decisions.
Results
Accuracy of LLMs in hearing diagnosis and hallucinations
The accuracy rates of the 8 LLMs models varied in hearing loss diagnosis. When diagnosing the degree of hearing loss based on provided audiograms (see Fig. 1, Supplementary Table 1), DeepSeek-V3 achieved the highest accuracy at 67.00%, followed by DeepSeek-R1 at 64.50%, while the remaining four models fell below 50%. Regarding the diagnosis of the type of hearing loss (e.g., conductive, sensorineural, or mixed), only DeepSeek-V3 achieved an accuracy rate above 50% (54.00%), while Gemini 2.0 Flash recorded the lowest rate (32.00%). Given the suboptimal diagnostic accuracy in the main experiments, this study conducted two supplementary experiments (see Supplementary Table 1). Results showed that diagnostic accuracy significantly improved (p < 0.01) for all six models except the two DeepSeek variants. Furthermore, the improvement in accuracy for diagnosing hearing severity was generally superior to that for diagnosing hearing loss type (p < 0.01).
Fig. 1. Diagnostic accuracy of LLMs in hearing test reports.
This bar chart presents the accuracy rates of various models in classifying two aspects of hearing loss: degree of hearing loss (light blue) and nature of hearing loss (dark blue). Each line or point style corresponds to a distinct model, including open-source LLMs Kimi-V1, Kimi-K1.5, Deepseek-V3, Deepseek-R1, and closed-source LLMs ChatGPT-4o, ChatGPT-o3, Gemini 2.0 Flash, and Gemini 2.0 Flash Thinking.
Additionally, we quantified hallucinations in each LLMs’ output. Among 100 responses, both Kimi-V1 and ChatGPT-4o generated the highest number of hallucinatory reports (n = 24, 24%), while Gemini 2.0 Flash Thinking produced the fewest (n = 4, 4%). Detailed results are presented in Supplementary Fig. 1.
Readability of LLM-generated content
Supplementary Table 2 presents the results of readability analysis for eight LLMs. From the FRE perspective, content generated by Kimi-V1 scored highest (66.05, 95% CI 63.08–68.90), while text produced by Kimi-K1.5 had the lowest FRE score (44.71, 95% CI 38.90–50.50). Regarding FKGL, DeepSeek generated the lowest average FKGL scores (V3: 6.83, 95% CI 6.40–7.30, R1: 6.41, 95% CI 5.93–6.98), while Kimi-K1.5 produced the highest FKGL scores (10.17, 95% CI 8.90–11.40). Regarding response length, Gemini 2.0 Flash produced the longest texts (21,885.54 words, 95% CI 20,864.00–22,899.25) among the eight LLMs, while DeepSeek-V3 generated the shortest texts (8,131.60 words, 95% CI 7675.50–8618.00) (Supplementary Fig. 2), with statistically significant differences (p < 0.01).
Experts’ comprehensive evaluation of LLMs
Using human expert consensus assessment as the gold standard, we evaluated the performance of these eight models in diagnostic processes, interpretation, and recommendations. Additionally, we performed a Cohen’s kappa analysis on the consistency of expert assessments. The specific Kappa values are presented in Supplementary Table 3. Based on the evaluation thresholds by JR Landis and GG Koch24, the Cohen’s kappa for expert assessments falls within the range of strong agreement, indicating minimal variation in expert consensus and consistent reliability of the expert assessment metrics. Figures 2 and 4, along with Supplementary Table 4, illustrate the response characteristics of the eight models as assessed by experts. In overall accuracy evaluation, DeepSeek-V3 and R1 achieved higher average scores of 3.35 (95% CI 3.14–3.56, p < 0.01) and 3.22 (95% CI 3.00–3.43, p < 0.01), while Gemini2.0 Flash and 2.0Flash Thinking scored relatively lower at 2.38 (95% CI 2.21–2.55, p < 0.01) and 2.45 (95% CI 2.26–2.64, p < 0.01). Scores for the diagnostic reasoning section indicated that DeepSeek-V3 and R1 remained competitive, with average scores of 3.35 (95% CI 3.14–3.56, p < 0.01) and 3.24 (95% CI 3.00–3.43, p < 0.01), respectively, while Gemini 2.0 Flash scored the lowest (2.38, 95% CI 2.21–2.55, p < 0.01). Furthermore, the results of the comprehensive assessment of report interpretation indicated that DeepSeek-V3 and Gemini 2.0 Flash performed best, with average scores of 3.19 (95% CI 2.99–3.39, p < 0.01) and 3.14 (95% CI 2.96-3.32, p < 0.01), respectively. DeepSeek-R1 followed closely (3.08, 95% CI 2.91–3.25, p < 0.01), while Kimi-V1 received the lowest performance score (2.45, 95% CI 2.29–2.61, p < 0.01). Finally, the helpfulness assessment in the recommendation section showed DeepSeek-V3 scored highest (3.30, 95% CI 3.08–3.52, p < 0.01), while Kimi-V1 scored lowest (2.69, 95% CI 2.51–2.87, p < 0.01).
Fig. 2. Results of expert evaluations of LLMs output text.
The four panels (a–d) display violin plots comparing the following indicators across models: a The Accuracy Indicator, b The Comprehension and Reasoning Ability Indicator, c The Comprehensiveness Indicator, d The Helpfulness Indicator. Each plot shows the distribution of scores (with the median and interquartile range) for models including Kimi-V1, Kimi-K1.5, DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, ChatGPT-01, Gemini2.0 Flash, and Gemini2.0 Flash Thinking. This violin plot displays the distribution of expert evaluation scores across different large language models. Each violin shape reflects the density distribution of the data.
Fig. 4. Radar chart of cross-model evaluation dimensions from expert and public perspectives.
a the comprehensive performance evaluation results of China LLMs; b the comprehensive performance evaluation results of U.S. LLMs. The pink box indicates dimensions evaluated from the expert perspective, while the blue box represents dimensions evaluated from the public perspective. The plotted indicators include accuracy, understanding and reasoning, comprehensiveness, helpfulness, comprehensibility, empathy, perceived value, and satisfaction, with models such as Kimi-V1, Kimi-K1.5, DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, ChatGPT-o1, Gemini2.0 Flash, and Gemini2.0 Flash Thinking.
Public evaluation
Figures 3 and 4, along with Supplementary Table 5, present the overall user experience of eight LLMs from a public perspective. Results indicate that Gemini 2.0 Flash Thinking achieved the highest comprehensibility score (4.25, 95% CI 4.14–4.35, p < 0.01), while Kimi-K1.5 scored the lowest (3.71, 95% CI 3.59–3.82, p < 0.01). Regarding empathy, Gemini models demonstrated superior performance with scores of 3.77 (95% CI 3.67–3.87, p < 0.01) and 3.75 (95% CI 3.65–3.85, p < 0.01), respectively. DeepSeek-R1 scored the lowest at 3.33 (95% CI 3.19–3.46, p < 0.01). In the assessment of perceived value, ChatGPT-4o scored highest (3.87, 95% CI 3.76–3.97, p < 0.01), while DeepSeek-R1 scored lowest (3.61, 95% CI 3.51–3.71, p < 0.01). Regarding satisfaction performance, Gemini 2.0 Flash achieved the highest average score (3.83, 95% CI 3.76–3.91, p < 0.01), while DeepSeek-R1 scored the lowest (3.44, 95% CI 3.34–3.54, p < 0.01).
Fig. 3. Public evaluation results of LLMs output text.
The four panels (a-d) display violin plots comparing the following indicators across models: a The Comprehensibility Indicator, b The Empathy Indicator, c The Perceived value Indicator, d The Satisfaction Indicator. Each plot shows the distribution of scores (with the median and interquartile range) for models including Kimi-V1, Kimi-K1.5, DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, ChatGPT-01, Gemini2.0 Flash, and Gemini2.0 Flash Thinking. This violin plot illustrates the distribution of public evaluation scores across various LLMs. Each violin shape reflects the density distribution of the data.
Discussion
To our knowledge, this study represents the first systematic evaluation of large language models’ ability to comprehend pure-tone audiogram reports. Using a blinded cross-sectional design, we assessed patient-facing hearing diagnosis, interpretation, and personalized recommendation generation capabilities of eight models, which span China and the US, encompass both open-source and closed-source systems, and cover general-purpose and reasoning-oriented large models, across 140 audiogram reports collected from multiple centers. Results indicate that the DeepSeek series models achieved the highest diagnostic accuracy. Expert evaluations demonstrated DeepSeek’s superior professional performance. In multidimensional assessments of patient assistance, multiple models showed potential auxiliary value, with Gemini scoring higher overall than comparable mainstream models like ChatGPT. This study provides a critical baseline for future clinical applications in hearing health.
LLMs still exhibit significant limitations in pure-tone audiometry diagnosis and are highly dependent on the format of input information. Although some models (such as DeepSeek-V3) achieve an accuracy rate of 67% in diagnosing the degree of hearing loss, their performance in diagnosing the nature of hearing loss remains suboptimal (with a maximum accuracy of only 54%). This outcome falls far below LLMs performance in text-based tasks like neurodegenerative diseases25, ophthalmology26, and radiological lung cancer staging27, indicating their limited capability in processing structured chart information and suggesting that relying solely on LLMs for hearing diagnosis lacks clinical reliability. Notably, after introducing supplementary experiments (e.g., providing information combining structured numerical data), the diagnostic accuracy of all six models except DeepSeek significantly improved (p < 0.01). We attribute this to potential reasoning errors caused by information overload. This aligns with prior research highlighting current LLMs’ challenges in integrating image and text information, with generally limited accuracy in medical image localization and anomaly detection tasks28,29. This demonstrates that LLMs' performance in medical diagnosis depends not only on model capabilities but also closely on task characteristics and input data types.
This study demonstrates that current large language models exhibit a high risk of hallucinations when interpreting audiology reports, potentially misleading diagnostic and intervention recommendations. Among the 100 reports evaluated, Kimi-V1 and ChatGPT-4o showed a hallucination rate of 24%, while Gemini 2.0 Flash Thinking had the lowest rate (4%). This stands in stark contrast to the 0.26% hallucination rate observed in Eric Steimetz et al.‘s study of GPT-4 interpreting pathology reports30. These errors often occur systematically in critical areas such as diagnostic inferences and intervention recommendations. For example, diagnostic errors include fabricating hearing threshold values (e.g., incorrectly reporting a pure-tone average) or inventing non-standard hearing classification criteria. Recommendation errors are more common and concerning, such as suggesting non-existent hearing aid models, recommending devices that are technically incompatible with the patient’s audiometric profile (e.g., recommending a device with insufficient gain for a severe-to-profound loss), or proposing inappropriate management pathways. These errors collectively provide users with substantively erroneous and potentially harmful information. Multiple specialty-specific studies (e.g., nursing31 and oncology32,33) also indicate that while LLMs perform well in primary care settings, they may generate inaccurate content when handling complex specialty issues, underscoring the need for careful evaluation in medical applications.
The readability of LLM outputs is critical for patient comprehension of reports. This study further quantitatively assessed the readability of text outputs from various models, revealing differences in readability when LLMs generate interpretations of hearing reports. Results indicate that only two metrics from the DeepSeek series (FKGL: 6.41, 95%CI 5.93–6.98; FRE: 63.70, 95%CI 59.50–67.30) simultaneously met the National Institutes of Health (NIH) and the American Medical Association (AMA) recommended reading levels for patient materials (FKGL 6-7, FRE 60-70)34,35, indicating concise and understandable responses more suitable for non-specialist users. In contrast, Kimi-K1.5’s FKGL corresponds to a high school level (10.17, 95% CI 8.90–11.40), potentially hindering comprehension among users with low health literacy36. Regarding response length, Gemini 2.0 Flash produced the longest responses (21,885.54 words, 95% CI 20,864.00–22,899.25), potentially affecting reader patience; DeepSeek-V3 generated the most concise responses (8,131.60 words, 95% CI 7,675.50–8,618.00) without sacrificing key information. Future LLMs must balance information density with readability to avoid compromising professional value through excessive simplification or verbosity.
The results of this study indicate that while LLMs demonstrate potential in interpreting pure-tone audiogram reports, their accuracy remains far below that of clinical practitioners. Expert evaluations reveal that the DeepSeek series models exhibit significant professional advantages. DeepSeek-V3 outperformed R1 in accuracy (3.35, 95% CI 3.14–3.56 and 3.22, 95% CI 3.00–3.43), comprehension and reasoning ability (3.35, 95% CI 3.09-3.61 and 3.24, 95% CI 2.99-3.49), and helpfulness (3.30, 3.08–3.52 and 3.25, 95% CI 3.03–3.47) than other models (p < 0.01). Its reports systematically cover key clinical elements such as audiogram interpretation, curve classification, and urgency of intervention, while providing actionable recommendations. In contrast, Gemini 2.0 Flash demonstrated acceptable comprehensiveness (3.14, 95% CI 2.96–3.32) but the lowest accuracy (2.38, 95% CI 2.21–2.55) and reasoning ability (2.38, 95% CI 2.17–2.59), indicating that content richness does not equate to reliable medical judgment. Key issues included misjudging hearing thresholds, failing to correctly map WHO classification criteria, and misinterpreting air-bone gap values, leading to erroneous assessments of the type of hearing loss (e.g., conductive, sensorineural, or mixed). Kimi-V1 and ChatGPT-4o exhibited limitations in comprehensiveness and explanatory depth, consistent with prior research37, indicating that current LLMs still face challenges in providing comprehensive and reliable medical explanations.
From the perspective of patient experience, certain models demonstrate unique value in enhancing the patient experience. Notably, unlike clinical experts who rigorously assess LLM performance against professional standards, general users prioritize comprehensibility, linguistic friendliness, psychological comfort in communication, and ease of information access. Research indicates that Gemini 2.0 Flash Thinking achieved the highest score for comprehensibility (4.25, 95% CI 4.14–4.35), while the Gemini series also demonstrated strong empathy scores (3.75, 95% CI 3.65–3.85 and 3.77, 95% CI 3.67–3.87), indicating that its generated content aligns more closely with patients’ communication needs and emotional expectations rather than solely emphasizing medical expertise. While DeepSeek-R1 leads in professional evaluation, it scored lowest in user satisfaction (3.44, 95% CI 3.34–3.54) and empathy (3.33, 95% CI 3.19–3.46), reflecting its high terminology density and inadequate emotional expression, which hinder public comprehension. This highlights patients’ heightened sensitivity to experiential, patient-centered features of AI outputs (e.g., clarity, comprehensibility, and interactivity), which traditional medical explanations often fail to address. However, it should be noted that high comprehensibility does not necessarily equate to reliable medical judgments, suggesting that optimizing a single dimension may lead to imbalances in clinical value. Therefore, we recommend that LLMs applied in hearing health education and communication assistance should balance both the accuracy and reliability of information with the emotional resonance of communication.
This study responds to the World Health Organization’s mission of “hearing health for all8.” However, in clinical practice, due to the scarcity of specialized audiology resources and the high workload of physicians in outpatient settings, it is challenging to provide patients with thorough interpretation of examination reports and detailed health recommendations. Compared to traditional hearing reports that only include diagrams and diagnostic conclusions, LLMs can help patients quickly develop a scientific understanding of their own auditory function by converting specialized data into accessible language as a potential auxiliary tool. This approach prevents anxiety caused by delayed interpretation or misunderstanding, while also reducing the risk of “neglecting early intervention due to insufficient awareness.” Especially in regions lacking audiological diagnostic resources, information provided by LLMs is highly likely to become patients’ primary source of medical explanations, partially compensating for insufficient professional communication. However, it must be made clear that patients’ high satisfaction with model outputs should not lead them to equate the results with clinical diagnoses38,39. Therefore, it is essential to incorporate prominent risk warnings in system design, emphasizing that the model serves only as an auxiliary interpretation rather than a medical conclusion. Thus, in clinical practice, LLMs may serve more practically as supplementary tools rather than independent decision-makers. When receiving LLM interpretations, patients should clearly understand that these represent only auxiliary opinions, with final diagnoses and treatment decisions relying on professional medical judgment.
There are also some shortcomings in this study. Firstly, data representativeness. This study only included pure-tone audiometry reports from Chinese populations and did not encompass medical histories, speech audiometry, or multimodal data (e.g., ABR and OAE). Additionally, the dataset demonstrated an uneven distribution of hearing loss severities, with mild-to-moderate cases predominating, aligning with population trends where mild and moderate hearing loss are more common than severe or profound cases4,5,40. Moreover, this pattern aligns with epidemiological trends observed in the general Chinese population. A meta-analysis review on the prevalence of hearing loss among Chinese seniors indicates that moderate hearing loss (41%) is more prevalent than mild (27%) or severe (10.1%) hearing loss, suggesting that early-stage hearing loss is more common41. This imbalance may limit the generalizability of our findings to populations with more severe hearing loss, which are less frequently observed in clinical practice. While this is unlikely to impact the primary comparative analysis, it underscores the need for validation in more representative cohorts.
Second, the manual evaluation may be influenced by human factors. Although all tasks were independently completed by blinded assessors who were unaware of which output corresponded to which LLM, subjective biases from the assessors could still affect the objectivity and consistency of the evaluation results.
Third, the evaluation dimensions were incomplete and did not quantify the model’s actual impact on clinical decision-making (e.g., intervention delays caused by misdiagnosis rates).
Fourth, our analysis was limited to web-based models for consistency; thus, performance may differ from the same LLMs accessed through their native application or other deployment environments. Furthermore, these platforms did not provide precise parameters used, and online applications may utilize shared data for future training, raising questions about their suitability for research settings.
Fifth, the ongoing rapid evolution of LLMs may limit the validity of current findings when newer versions are released.
In summary, this study represents the first evaluation of LLMs’ ability to comprehend specialized medical chart reports, focusing on their potential application in interpreting pure-tone audiograms for patients. It covers model-generated patient-facing hearing diagnoses, explanatory interpretations, and personalized health recommendations. The findings indicate that while existing models, including DeepSeek-V3, still exhibit significantly lower diagnostic accuracy than professional audiologists and produce varying degrees of “hallucination” outputs, rendering them unsuitable for independent clinical diagnosis, they demonstrate clear auxiliary potential and patient value in translating specialized audiograms into patient-understandable report interpretations and health recommendations. This study not only addresses the gap in evaluating LLM applications within audiology but also reveals the critical impact of task characteristics and input formats on LLM medical performance. It highlights the current limitations of LLMs in achieving professional accuracy in chart interpretation while underscoring their potential for generating patient-specific explanations and recommendations. Through optimized prompt engineering, enhanced domain knowledge integration, and the development of human-AI collaborative systems, LLMs hold promise as vital auxiliary tools to bridge global hearing healthcare gaps and mitigate inequitable distribution of medical resources. Their core value lies in improving patient comprehension and health information accessibility, thereby providing essential technological support toward achieving “accessible, inclusive, and precise” universal hearing healthcare8.
Methods
The workflow of this study is shown in Fig. 5. For the detailed design of the overall framework, please refer to Supplementary Fig. 3. As the evaluation targets a chatbot, the report follows the Chatbot Evaluation Report Tool (CHART) guidelines (BMJ, August 2025)42, with the full checklist in Supplementary Table 6. The study protocol was prospectively registered on the Open Science Framework43.
Fig. 5. Overall workflow for evaluating pure-tone audiograms based on large language models.
(1) Data Resources and Processing: Collected 140 pure-tone audiogram reports from two institutions (Second Affiliated Hospital of Zhejiang University SAHZU and Hangzhou Hearing Center), each containing personal information, bilateral audiograms, and clinical diagnoses. For privacy and blind evaluation purposes, de-identified audiograms (300dpi) were regenerated from the original data. (2) LLMs and Preparation: Select 8 LLMs from OpenAI, Google, DeepSeek, and Moonshot (4 open-source + 4 closed-source); Forty reports were randomly selected for expert prompt design and gold standard establishment (not included in the formal experiment). (3) LLMs Input and Output: Researchers posed three consecutive questions to each model using standardized prompts and sample reports, collecting: a Diagnosis (left/right ear “degree + nature,” requiring reasoning process); b Interpretation (patient-friendly simplified report); c Recommendation (professional opinion). (4) Evaluation Methods: Based on the QUEST framework, each model’s responses are evaluated through both objective and subjective assessments.
Data resources and processing
The pure-tone audiogram reports for this study were obtained from: the Affiliated Hospital of Zhejiang University School of Medicine (SAHZU) and Hangzhou Huitin (International) Hearing and Balance Center. A total of 140 Chinese pure-tone audiogram reports (including bilateral audiograms and corresponding clinical diagnoses) were randomly collected between October 2024 and December 2024. The specific data collection process and inclusion/exclusion criteria are detailed in Supplementary Note 1. Seventy reports originated from SAHZU, and 70 from Huitin (International) Audiology and Balance Center. Clinical diagnoses in the pure-tone audiogram reports included the degree and nature of bilateral hearing loss. We categorized each ear based on hearing loss severity; the specific sample size distribution is shown in Supplementary Table 7. Based on the original hearing data, we regenerated standardized, anonymized pure-tone audiogram reports without diagnostic outcomes. Specific examples are shown in Supplementary Fig. 4. This study received ethical approval from the SAHZU Faculty of Medicine Ethics Committee(2025–0027). According to the SAHZU Faculty of Medicine Ethics Committee regulations, the use of these anonymized data does not require approval or informed consent. The Institutional Review Board has conditionally approved sharing data access with third parties.
LLMs selection and preparation
This study selected four latest open-source and closed-source LLMs from four leading companies in China(DeepSeek, Moonshotand) the US (OpenAI, Google). Specific model version dates are listed in Supplementary Table 8.
We randomly sampled 40 reports from the experimental dataset for prompt design and gold standard evaluation. To prevent data leakage, these reports were not reused in subsequent experiments. On one hand, one large model expert (GXW) and two experts with relevant certifications and extensive clinical experience (CHT, MC) conducted multi-round pre-experiments using the 40 cases in Hangzhou, China (January 3–15, 2025). On the other hand, Experts provided prompts and contextual information to the large language model based on existing literature and clinical experience37,44–49. After each round, expert feedback was collected, and prompts and experimental design were adjusted accordingly. Specific prompts are detailed in Supplementary Table 9. Separately, two experts (CHT and MC) independently interpreted the 40 reports and provided recommendations, independent of the LLMs outputs, establishing the gold standard for the interpretation and recommendation sections. Additionally, the gold standard for the diagnosis section was the original diagnostic findings.
LLMs input and output
The collection of input and output results for LLMs was conducted by two trained graduate students (QCF and TTZ) from January 22, 2025, to March 20, 2025. The specific operational procedure was as follows: First, the LLMs generated diagnostic results based on prompts, including assessments of hearing loss severity in both ears and the nature of the hearing loss (e.g., left ear: moderate hearing loss; left ear: sensorineural hearing loss; and the right ear had severe hearing loss with a mixed hearing loss nature). The LLMs were also required to provide the diagnostic reasoning process associated with these diagnoses. Second, the LLMs were tasked with interpreting and providing recommendations based on the pure-tone audiogram report, delivering open-ended explanations and professional advice in a manner accessible to the general public. Additionally, to prevent random responses from the LLMs, each model’s temperature parameter was set to 0, and offline mode was used. This process employed a continuous question-and-answer format, posing three diagnostic, interpretive, and recommendation queries to the LLMs via an online platform. All generated content was collected into structured response reports (100 in total), with specific examples available in Supplementary Table 10. All collected data were ultimately reviewed and verified by the third graduate student (MYX).
Evaluation methods
We extracted certain subjective metrics from the QUEST framework proposed by Tam et al.50 and combined them with the most commonly used objective metrics in LLMs evaluation, including accuracy, readability, and hallucination-related indicators28,51, to form an evaluation metric system for diagnosing, interpreting, and recommending LLMs-generated content, as shown in Table 1. To ensure the objectivity of the evaluation, this study employed a blind evaluation design, where the information about the models to be evaluated was concealed from the evaluators to control potential evaluation biases.
Table 1.
Evaluation Indicators and Standards
| Evaluation indicators | Evaluator | Evaluation criteria | Evaluation results | |
|---|---|---|---|---|
| Objective evaluation | Diagnostic accuracy | 1 graduate student in audiology (MYX) | Evaluate only the diagnostic findings section. Compare with the original report’s diagnostic results and record the percentage of correctly diagnosed cases among 200 ears (including accuracy rates for diagnosing hearing loss severity and nature). Additionally, we have incorporated two supplementary tests: a pure-tone audiogram based on raw data (hearing thresholds for each frequency band in both ears) and a combined report featuring both audiograms and tables. (see Supplementary Figs. 5 and 6, 300 dpi format). | Percentage (0–100%) |
| Readability-FRE53 | Web-based tools (readable.com) | Only evaluate the interpretation section, calculated based on sentence length and lexical difficulty. A higher score indicates a shorter text with a simpler structure. | value (1–100) | |
| Readability-FKGL54,55 | Only the interpretation section will be evaluated. This scoring assesses the text’s “grade level” based on sentence length and lexical complexity. | Value (0–18) | ||
| Readability-text length | Response length for statistical diagnosis, interpretation, and recommendations | Value (word count) | ||
| Hallucination56 | 2 experts (CHT, MC) | Evaluate the content of the diagnostic process, interpretation, and recommendations. Did the LLMs generate medical findings not present in the original report, or fabricate nonexistent medical data? | Categories (presence or absence) | |
| Subjective assessment | Accuracy57 | Expert evaluation (2 experts: CHT, MC) | Evaluate the content of the diagnostic process, interpretation, and recommendations. The primary focus is on assessing whether responses are factually correct, precise, and error-free. | 5-point Likert scale |
| Understanding and reasoning58 | The evaluation focuses solely on the diagnostic component, primarily assessing whether the large model correctly understands the problem and infers the correct outcome. This includes fundamental comprehension abilities, logical reasoning capabilities, and clinical diagnostic support. | 5-point Likert scale | ||
| Comprehensive59 | Evaluate only the interpretation and recommendation sections. The primary focus is on assessing the completeness of the responses provided by the LLMs. Does the response cover all key aspects, offering a comprehensive overview or detailed insights? | 5-point Likert scale | ||
| Helpfulness60 | Evaluate only the recommendation section. This refers to the applicability and practicality of the response. Determine whether the answer holds real-world value and whether it is actionable and relevant to the user’s question. | 5-point Likert scale | ||
| Comprehensibility/understandability61 | Public evaluation (3 volunteers without medical backgrounds: MJL, YF, and YCQ) | The ease with which users can comprehend the information and content provided by the LLMs. | 5-point Likert scale | |
| Empathy62 | The ability of an LLMs to empathize with others’ emotions, thoughts, and circumstances, and to respond appropriately with emotional sensitivity. | 5-point Likert scale | ||
| Perceived value/Usefulness63 | Does the information or content provided by the LLMs genuinely prioritize the user’s perspective, address their core needs, and deliver tangible benefits? | 5-point Likert scale | ||
| Satisfaction64 | Users’ subjective perceptions and level of recognition regarding the overall output quality and practicality of LLMs. | 5-point Likert scale |
Statistical analysis
Statistical analyses were performed using IBM SPSS Statistics version 29.0.1.052. Descriptive statistics are reported as mean ± standard deviation for continuous variables, with confidence intervals calculated using Wilson’s scoring method. For the continuous readability metrics, the Friedman test was employed as a robust alternative after confirming that the normality assumption for parametric repeated-measures ANOVA was violated. Since Likert scale scores are ordinal and do not meet the normality assumption required for parametric tests, the Friedman test was used to compare the overall performance of different large language models, supplemented by Nemenyi post-hoc tests. Mann-Whitney U tests were applied for independent sample comparisons. Expert assessment consistency was analyzed using Cohen’s Kappa. P < 0.05 were considered statistically significant.
Supplementary information
Acknowledgements
This study extends sincere gratitude to three public volunteers without medical backgrounds (Mengjie Liu, Yue Fan, and Yachun Qi) for their invaluable contributions. They systematically evaluated content generated by large language models from a public user perspective, providing indispensable patient-centric data for this research. Furthermore, their review and feedback on the draft manuscript ensured the research findings are clearly communicated to non-specialist readers. Their insightful participation significantly enhanced the rigor and practical relevance of this research. This work was supported by the National Natural Science Foundation of China (grant numbers 81871455 and 72574006), Zhejiang Provincial Natural Science Foundation of China (grant number LY22H180001), Municipal Natural Science Foundation of Beijing of China (grant number 7222306).
Author contributions
J.B.L. and J.L. conceived the study. Q.C.F. and T.T.Z. participated in inputting LLMs and collecting output results. M.Y.X. and X.M.L. were responsible for reviewing all data. M.C. and C.H.T. conducted professional evaluations of LLM outputs. P.X. and G.X.W. performed data analysis. Q.C.F. and X.Y.J. completed manuscript translation, while Z.Y.L. and J.K.H. were responsible for creating and designing figures and tables. J.K.H. and J.B.L. provided criticism, suggestions, and revisions for the article. M.Y.X. drafted the initial manuscript, and J.L. performed critical revisions.
Data availability
Due to the sensitive nature of the data (e.g., patient information), access is restricted. Data are available from the corresponding author upon reasonable request and with permission from the Human Research Ethics Committee, Second Affiliated Hospital, School of Medicine, Zhejiang University.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Jun Liang, Mengyao Xing, Peng Xiang, Guixuan Wang.
These authors jointly supervised this work: Chenghua Tian, Jianbo Lei.
Contributor Information
Chenghua Tian, Email: 20071044@zcmu.edu.cn.
Jianbo Lei, Email: jblei@hsc.pku.edu.cn.
Supplementary information
The online version contains supplementary material available at 10.1038/s41746-026-02537-1.
References
- 1.Alshehri, K. A., Alqulayti, W. M., Yaghmoor, B. E. & Alem, H. Public awareness of ear health and hearing loss in Jeddah, Saudi Arabia. South Afr. J. Commun. Disord.66, e1–e6 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Głowacka, M. D. et al. Knowledge of parents from urban and rural areas vs. prevention methods of hearing loss threats seen as challenges for public health. Ann. Agric. Environ. Med.24, 157–161 (2017). [DOI] [PubMed] [Google Scholar]
- 3.Ferndale, D., Watson, B. & Munro, L. Hearing loss as a public health matter-why not everyone wants their deafness or hearing loss cured. Aust. N. Z. J. Public Health37, 594–595 (2013). [DOI] [PubMed] [Google Scholar]
- 4.Sensory Functions, Disability and Rehabilitation (SDR). World Report on Hearing. https://www.who.int/publications/i/item/9789240020481 (2021).
- 5.GBD 2019 Hearing Loss Collaborators Hearing loss prevalence and years lived with disability, 1990-2019: findings from the Global Burden of Disease Study 2019. Lancet397, 996–1009 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Lee, S. Y., Seo, H. W., Jung, S. M., Lee, S. H. & Chung, J. H. Assessing the accuracy and reliability of application-based audiometry for hearing evaluation. Sci. Rep.14, 7359 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Greenberg, J. et al. Cultivating resiliency in patients with neurofibromatosis 2 who are deafened or have severe hearing loss: a live‑video randomized control trial. J. Neurooncol.145, 561–569 (2019). [DOI] [PubMed] [Google Scholar]
- 8.Curhan, S. G. WHO world hearing forum: guest editorial. Ear and hearing care: a global public health priority. Ear Hear.40, 1–2 (2019). [DOI] [PubMed] [Google Scholar]
- 9.Gotlieb, R. et al. Accuracy in patient understanding of common medical phrases. JAMA Netw. Open5, e2242972 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.The Lancet Digital Health, null Large language models: a new chapter in digital health. Lancet Digit. Health6, e1 (2024). [DOI] [PubMed] [Google Scholar]
- 11.Goh, E. et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat. Med.31, 1233–1238 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Aydin, S., Karabacak, M., Vlachos, V. & Margetis, K. Large language models in patient education: a scoping review of applications in medicine. Front. Med.11, 1477898 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Masanneck, L. et al. Triage performance across large language models, ChatGPT, and untrained doctors in emergency medicine: comparative study. J. Med. Internet Res.26, e53297 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Williams, C. Y. K. et al. Use of a large language model to assess clinical acuity of adults in the emergency department. JAMA Netw. Open7, e248895 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Duggan, M. J. et al. Clinician experiences with ambient scribe technology to assist with documentation burden and efficiency. JAMA Netw. Open8, e2460637 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Roberts, K. Large language models for reducing clinicians’ documentation burden. Nat. Med.30, 942–943 (2024). [DOI] [PubMed] [Google Scholar]
- 17.Friedman, A. B., Delgado, M. K. & Weissman, G. E. Artificial intelligence for emergency care triage-much promise, but still much to learn. JAMA Netw. Open7, e248857 (2024). [DOI] [PubMed] [Google Scholar]
- 18.Mallio, C. A., Sertorio, A. C., Bernetti, C. & Beomonte Zobel, B. Large language models for structured reporting in radiology: performance of GPT-4, ChatGPT-3.5, Perplexity and Bing. Radiol. Med.128, 808–812 (2023). [DOI] [PubMed] [Google Scholar]
- 19.Liu, X. et al. Uncovering language disparity of ChatGPT on retinal vascular disease classification: cross-sectional study. J. Med. Internet Res.26, e51926 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Longwell, J. B. et al. Performance of large language models on medical oncology examination questions. JAMA Netw. Open7, e2417641 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Ullah, E., Parwani, A., Baig, M. M. & Singh, R. Challenges and barriers of using large language models (LLM) such as ChatGPT for diagnostic medicine with a focus on digital pathology - a recent scoping review. Diagn. Pathol. 19, 43 (2024). [DOI] [PMC free article] [PubMed]
- 22.Lee, D. et al. Using large language models to automate data extraction from surgical pathology reports: retrospective cohort study. JMIR Form. Res.9, e64544 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Li, Z. et al. Leveraging large language models for NLG evaluation: advances and challenges. in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 16028–16045 (Association for Computational Linguistics, Miami, Florida, USA, 2024). 10.18653/v1/2024.emnlp-main.896.
- 24.Landis, J. R. & Koch, G. G. The measurement of observer agreement for categorical data. Biometrics33, 159–174 (1977). [PubMed] [Google Scholar]
- 25.Koga, S., Martin, N. B. & Dickson, D. W. Evaluating the performance of large language models: ChatGPT and Google Bard in generating differential diagnoses in clinicopathological conferences of neurodegenerative disorders. Brain Pathol. Zur. Switz.34, e13207 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Mikhail, D. et al. DeepSeek-R1 vs OpenAI o1 for ophthalmic diagnoses and management plans. JAMA Ophthalmol. e252918, 10.1001/jamaophthalmol.2025.2918 (2025). [DOI] [PMC free article] [PubMed]
- 27.Lee, J. E. et al. Lung cancer staging using chest CT and FDG PET/CT free-text reports: comparison among three ChatGPT large language models and six human readers of varying experience. Am. J. Roentgenol.223, e2431696 (2024). [DOI] [PubMed] [Google Scholar]
- 28.Suh, P. S. et al. Comparing diagnostic accuracy of radiologists versus GPT-4V and Gemini Pro Vision using image inputs from diagnosis please cases. Radiology312, e240273 (2024). [DOI] [PubMed] [Google Scholar]
- 29.Zhu, Q. et al. How well do multi-modal LLMs interpret CT scans? An auto-evaluation framework for analyses. J. Biomed. Inform. 168, 104864 (2025). [DOI] [PubMed]
- 30.Steimetz, E. et al. Use of artificial intelligence chatbots in interpretation of pathology reports. JAMA Netw. Open7, e2412767 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Li, J., Dada, A., Puladi, B., Kleesiek, J. & Egger, J. ChatGPT in healthcare: a taxonomy and systematic review. Comput. Methods Prog. Biomed.245, 108013 (2024). [DOI] [PubMed] [Google Scholar]
- 32.Zhou, Z. et al. ChatGPT in oncology diagnosis and treatment: applications, legal and ethical challenges. Curr. Oncol. Rep.27, 336–354 (2025). [DOI] [PubMed] [Google Scholar]
- 33.Mundinger, A. Artificial intelligence in breast oncology. Holist. Integr. Oncol.4, 33 (2025). [Google Scholar]
- 34.Clear & Simple | National Institutes of Health (NIH). https://www.nih.gov/institutes-nih/nih-office-director/office-communications-public-liaison/clear-communication/clear-simple.
- 35.Rooney, M. K. et al. Readability of patient education materials from high-impact medical journals: a 20-year analysis. J. Patient Exp.8, 2374373521998847 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Ho, B., Hong, E. M. & Benson, B. E. Assessing and improving the effectiveness of online patient education materials on essential vocal tremor: a comprehensive evaluation. J. Voice10.1016/j.jvoice.2024.02.021 (2024). [DOI] [PubMed]
- 37.Suárez, A. et al. Beyond the Scalpel: assessing ChatGPT’s potential as an auxiliary intelligent virtual assistant in oral surgery. Comput. Struct. Biotechnol. J.24, 46–52 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.He, Z. et al. Quality of answers of generative large language models versus peer users for interpreting laboratory test results for lay patients: evaluation study. J. Med. Internet Res.26, e56655 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Sorin, V. et al. Large language models and empathy: systematic review. J. Med. Internet Res.26, e52597 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Cantuaria, M. L. et al. Hearing loss, hearing aid use, and risk of dementia in older adults. JAMA Otolaryngol. Head. Neck Surg.150, 157–164 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Xian, Y., Gao, J., Chen, H. et al. Meta-analysis of the prevalence of hearing loss among the elderly in China. Mod. Prev. Med.49, 2451–2458 (2022). [Google Scholar]
- 42.CHART Collaborative Reporting guidelines for chatbot health advice studies: explanation and elaboration for the Chatbot Assessment Reporting Tool (CHART). BMJ390, e083305 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.M. Xing, T. Zhou & Q. Fang OSF | Evaluation of the diagnostic, interpretive and advisory capabilities of LLM in pure-tone hearing reports: a multicenter cross-sectional study. https://osf.io/bvjgt/.
- 44.Williams, C. Y. K., Miao, B. Y., Kornblith, A. E. & Butte, A. J. Evaluating the use of large language models to provide clinical recommendations in the emergency department. Nat. Commun.15, 8236 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Bedi, S. et al. Testing and evaluation of health care applications of large language models: a systematic review. JAMA333, 319–328 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Huo, B. et al. Large language models for chatbot health advice studies: a systematic review. JAMA Netw. Open8, e2457879 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Johri, S. et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nat. Med.31, 77–86 (2025). [DOI] [PubMed] [Google Scholar]
- 48.Sandmann, S. et al. Benchmark evaluation of DeepSeek large language models in clinical decision-making. Nat. Med.31, 2546–2549 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Shool, S. et al. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Med. Inform. Decis. Mak.25, 117 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Tam, T. Y. C. et al. A framework for human evaluation of large language models in healthcare derived from literature review. NPJ Digit. Med.7, 258 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Chang, Y. et al. A survey on evaluation of large language models. ACM Trans. Intell. Syst. Technol.15, 1–45 (2024). [Google Scholar]
- 52.IBM Corp. Downloading IBM SPSS Statistics 29.0.2.0. https://www.ibm.com/support/pages/downloading-ibm-spss-statistics-29020.
- 53.Flesch R. A new readability yardstick. https://psycnet.apa.org/record/1949-01274-001 (1948). [DOI] [PubMed]
- 54.Coleman, M. & Liau, T. L. A computer readability formula designed for machine scoring. J. Appl. Psychol.60, 283–284 (1975). [Google Scholar]
- 55.Fry E. Fry’s readability graph: clarifications, validity, and extension to Level 17 on JSTOR. https://www.jstor.org/stable/40018802 (1977).
- 56.Farquhar, S., Kossen, J., Kuhn, L. & Gal, Y. Detecting hallucinations in large language models using semantic entropy. Nature630, 625–630 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Bazzari, F. H. & Bazzari, A. H. Utilizing ChatGPT in telepharmacy. Cureus16, e52365 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Singhal, K. et al. Large language models encode clinical knowledge. Nature620, 172–180 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Lechien, J. R., Georgescu, B. M., Hans, S. & Chiesa-Estomba, C. M. ChatGPT performance in laryngology and head and neck surgery: a clinical case-series. Eur. Arch. Otorhinolaryngol.281, 319–333 (2024). [DOI] [PubMed] [Google Scholar]
- 60.Peng, W. et al. Evaluating AI in medicine: a comparative analysis of expert and ChatGPT responses to colorectal cancer questions. Sci. Rep.14, 2840 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Chen, J. et al. A comparative analysis of large language models on clinical questions for autoimmune diseases. Front. Digit. Health7, 1530442 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Lee, J., Park, S., Shin, J. & Cho, B. Analyzing evaluation methods for large language models in the medical field: a scoping review. BMC Med. Inform. Decis. Mak.24, 366 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Yun, J. Y., Kim, D. J., Lee, N. & Kim, E. K. A comprehensive evaluation of ChatGPT consultation quality for augmentation mammoplasty: a comparative analysis between plastic surgeons and laypersons. Int. J. Med. Inf.179, 105219 (2023). [DOI] [PubMed] [Google Scholar]
- 64.Choi, J. et al. Availability of ChatGPT to provide medical information for patients with kidney cancer. Sci. Rep.14, 1542 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Due to the sensitive nature of the data (e.g., patient information), access is restricted. Data are available from the corresponding author upon reasonable request and with permission from the Human Research Ethics Committee, Second Affiliated Hospital, School of Medicine, Zhejiang University.





