ABSTRACT
Background
Artificial intelligence (AI) technologies, particularly large language models (LLMs) such as ChatGPT, are increasingly utilised in medical education and clinical information retrieval. Nevertheless, their capacity to accurately reproduce recommendations from established clinical practice guidelines (CPGs) has not been thoroughly examined. The present study evaluated the concordance between responses generated by GPT‐5 and recommendations contained in German CPGs addressing oral potentially malignant disorders (OPMDs) and oral carcinomas (OCs).
Methods
A cross‐sectional analytical comparison was performed between GPT‐5 outputs and German CPG recommendations available as of October 2025. Individual guideline statements were entered verbatim into GPT‐5, which was asked to confirm or reject the statements. To assess methodological robustness, inverted versions of the same statements were additionally tested. GPT‐5 was accessed through the free version without internet connectivity to ensure that responses originated solely from the model's internal training data. Accuracy was defined as the proportion of correctly classified statements. Concordance between guideline content and model responses was quantified using Cohen's 𝝹.
Results
Two German CPGs comprising 111 recommendations were included: the S2k guideline for OPMDs (15 recommendations) and the S3 guideline for OCs (96 recommendations). GPT‐5 correctly affirmed all authentic recommendations and rejected all inverted statements. Agreement between guideline statements and GPT‐5 responses was perfect when the original recommendations were analysed (𝝹 = 1.0) and remained very high when both original and inverted statements were evaluated jointly (𝝹 = 0.96). The majority of references cited within the guidelines were published in English (> 93%) and originated from outside Germany (> 77%).
Conclusion
When guideline recommendations were presented verbatim, GPT‐5 demonstrated complete concordance with German oral oncology CPGs. These findings indicate that the model is capable of recognising and retrieving established guideline information. However, this experimental design evaluates recognition of existing statements rather than autonomous clinical reasoning. At present, LLMs should therefore be regarded primarily as educational and informational tools rather than a replacement for expert clinical judgement in oral oncology.
Keywords: artificial intelligence, clinical practice guideline, GPT‐5, oral oncology
1. Introduction
Artificial intelligence (AI) has rapidly become integrated into numerous aspects of modern healthcare, including medical education, research and clinical information management. Among recent developments, large language models (LLMs) such as ChatGPT, developed by OpenAI, have attracted considerable attention. These systems are increasingly used by clinicians and trainees to summarise scientific literature, clarify medical concepts and access guideline‐based information [1]. Despite their growing popularity, questions remain regarding the reliability, transparency and clinical validity of their responses.
As LLMs are progressively incorporated into clinical information workflows, systematic evaluation of their performance has become essential. In particular, assessing whether such systems can accurately reproduce established guideline recommendations represents an important step in determining their potential utility within evidence‐based medicine.
Earlier investigations examining previous ChatGPT versions have reported heterogeneous results, with reported accuracy ranging from approximately 36%–90%, depending on the complexity of the task, methodological design, languages and medical speciality examined [2]. Notably, many of these studies focused on case‐based reasoning or clinical decision‐making scenarios, whereas others examined simpler tasks involving retrieval of factual knowledge or reproduction of guideline contents.
Clinical practice guidelines (CPGs) constitute an important cornerstone of evidence‐based clinical care. They synthesise available scientific evidence and expert consensus into structured recommendations designed to guide diagnostic and therapeutic decisions. In Germany, such guidelines are developed within the framework of the Association of the Scientific Medical Societies in Germany (AWMF), which differentiates among several methodological levels, including consensus‐based S2k guidelines and fully evidence‐ and consensus‐based S3 guidelines.
Because guideline recommendations are typically formulated as explicit, structured statements, they provide a convenient reference standard for assessing whether AI systems can recognise and reproduce established medical knowledge. However, verifying the correctness of existing guideline statements differs fundamentally from generating independent clinical recommendations in real‐world situations.
To date, the performance of GPT‐5—released in August 2025—has not been examined with respect to German oral oncology CPGs. The present study therefore investigated the level of concordance between GPT‐5 responses and recommendations contained in German guidelines addressing oral potentially malignant disorders (OPMDs) and oral carcinomas (OCs). Rather than evaluating complex clinical reasoning, the primary objective was to determine whether GPT‐5 can accurately recognise and verify structured guideline statements. Additionally, this study demonstrates a straightforward methodological framework that may be used to assess LLM concordance with guideline recommendations.
2. Materials and Methods
A cross‐sectional comparative study design was employed to evaluate the concordance between GPT‐5 responses and recommendations contained in German CPGs addressing OPMDs and OCs that were available as of October 2025. Ethical approval was not required because the investigation relied exclusively on publicly available guideline documents and did not involve patient data or human participants.
Each recommendation contained in the selected guidelines was entered verbatim into GPT‐5. The model was then asked to determine whether the statement represented correct medical information according to current knowledge reflected in German CPGs. To minimise variability arising from external information retrieval, the free version of GPT‐5 without internet connectivity was used so that responses were generated exclusively from the model's internal training corpus. This version was intentionally selected because it reflects the level of access available to the general public.
To strengthen methodological reliability, two investigators (J.H. and P.P.) independently conducted the evaluation. In addition to testing the original guideline statements, inverted (‘false’) versions of the same statement were generated. These inverted statements were created by reversing the principal recommendation while maintaining grammatical coherence and clinical plausibility—for example, by substituting recommended interventions with non‐recommended alternatives or by reversing indications and contraindications. Any discrepancies between the two evaluators were resolved through consensus.
It should be acknowledged that entering guideline statements verbatim may predispose the model to recognise previously encountered textual content. Consequently, the present methodological approach primarily evaluates recognition of structured statements rather than independent clinical reasoning or decision‐support capabilities.
Accuracy was defined as the proportion of statements correctly classified by GPT‐5, including both confirmation of true recommendations and rejection of inverted statements.
The principal predictor variable was the source of information (CPG vs. GPT‐5), whereas the primary outcome variable was response accuracy. Additional study variables were categorised as follows:
-
–
CPG‐related factors: year of publication, guideline validity period, methodological level (German S‐classification), type of recommendation, number of English‐language references and number of references originating outside Germany.
-
–
GPT‐5‐related factors: explanatory output provided by the model when confirming or rejecting statements.
Agreement between GPT‐5 responses and guideline recommendations was quantified using Cohen's 𝝹 statistic. Although responses were deterministic in the present design, 𝝹 was calculated to provide a quantitative measure of concordance between the two information sources. A significance threshold of p ≤ 0.05 was applied for selected descriptive comparisons.
3. Results
Two German CPGs met the inclusion criteria: the 2019 S2k guideline addressing OPMDs (valid until August 2024) [3] and the 2021 S3 guideline addressing OCs (valid until February 2026) [4]. Together, these documents contained 111 recommendations, including 15 (13.5%) related to OPMDs and 96 (86.5%) related to OCs.
The S2k guideline focused predominantly on diagnostic recommendations, whereas the S3 guideline contained a larger proportion of therapeutic recommendations.
When the original guideline statements were entered into GPT‐5, the model correctly confirmed all recommendations. Consequently, agreement between GPT‐5 responses and guideline content was perfect (𝝹 = 1.0; standard error [SE]: 0.0; 95% confidence interval [CI]: 1.0–1.0).
When the analysis included both original and inverted statements, GPT‐5 also correctly rejected all inverted versions. The resulting overall concordance remained extremely high (𝝹 = 0.96; SE: 0.02; 95% CI: 0.93–1.0).
Because no discrepancies were observed in the primary outcome variable, additional multivariable statistical modelling was not considered informative.
Analysis of the bibliographic sources cited within the guidelines revealed that the vast majority were published in English (93.5%–98.8%). Furthermore, between 77% and 91% of the cited studies originated from outside Germany. The proportion of German references was lower in the OC guideline compared with the OPMD guideline. Table 1 summarises the main characteristics of the included guidelines.
TABLE 1.
Characteristics of the included German clinical practice guidelines (CPGs) for oral potentially malignant disorders (OPMDs) and oral carcinomas (OCs).
| Characteristic | OPMD guideline (S2k) | OC guideline (S3) |
|---|---|---|
| Year of publication | 2019 | 2021 |
| Validity period | Until August 2024 | Until February 2026 |
| Level of evidence (German S‐grading) | S2k | S3 |
| Total recommendations | 15 (13.5%) | 96 (86.5%) |
| Nature of recommendations | ||
| Diagnostic | 14 (93.3%) | 23 (24.0%) |
| Therapeutic | 1 (6.7%) | 58 (60.4%) |
| Prevention/risk factors | 0 | 6 (6.3%) |
| Surveillance/follow‐up | 0 | 9 (9.4%) |
| References cited in CPGs | ||
| Total references | 62 | 593 |
| References in English | 58 (93.5%) | 586 (98.8%) |
| References originating outside Germany | 48 (77.4%) | 539 (90.9%) |
| GPT‐5 responses | ||
| Correct confirmation of true recommendations | 15/15 (100%) | 96/96 (100%) |
| Correct rejection of inverted statements | 15/15 (100%) | 96/96 (100%) |
Note: Categorical variables are presented as number (percentage). The table summarises descriptive characteristics of the included CPGs and the corresponding GPT‐5 responses. Inferential statistics were omitted because the study primarily evaluated concordance rather than comparative hypothesis testing.
4. Discussion
The present study examined whether GPT‐5 is capable of recognising and verifying structured recommendations contained in German CPGs related to oral oncology. When guideline statements were entered verbatim, the model demonstrated complete concordance with the original recommendations and accurately rejected inverted statements.
These findings suggest that GPT‐5 can reliably recognise structured guideline statements and retrieve previously established medical knowledge. Nevertheless, it is important to emphasise that the present experimental design primarily evaluates recognition and information retrieval rather than independent clinical reasoning or decision‐making capabilities.
The methodological design used in this study inherently favours high levels of agreement. Presenting recommendations verbatim increases the probability that similar textual content may have been encountered during model training or within related biomedical literature. Consequently, the finding should not be interpreted as evidence that GPT‐5 is capable of independently generating guideline‐compliant recommendations in complex clinical contexts.
Previous investigations examining LLM performance in medical applications have reported highly variable results depending on the type of task, evaluation methods, languages and medical disciplines considered [2]. Studies involving case‐based reasoning, diagnostic interpretation or therapeutic planning typically demonstrate lower accuracy than tasks that primarily require retrieval of factual information or reproduction of guideline content [5, 6, 7, 8].
Comparable patterns have also been described in studies comparing LLM outputs with CPGs across different medical fields. In these investigations, models often reproduced structured recommendations with reasonable accuracy but showed less consistent performance when confronted with complex clinical scenarios requiring interpretation or judgement. For example, research in musculoskeletal rehabilitation, pain management and digital health has demonstrated that LLM responses may align with guideline statements while still necessitating careful clinical oversight when applied to decision‐support contexts [9, 10, 11].
The near‐perfect concordance observed in the present analysis therefore most likely reflects the structured nature of the input rather than genuine reasoning ability. Head and neck malignancies represent a major global health concern, and their management has been extensively documented in the biomedical literature. Previous studies comparing ChatGPT outputs with recommendations from the National Comprehensive Cancer Network (NCCN) guidelines for head and neck cancer have reported concordance rates of approximately 92%–95% [12, 13]. The complete concordance observed here may partially reflect the direct presentation of guideline statements rather than responses to open‐ended clinical questions.
Another noteworthy observation relates to the predominance of English‐language literature within the analysed German CPGs. More than 90% of the cited references were published in English, and most originated from outside Germany. Given that LLM training datasets consist largely of English‐language text, this linguistic distribution may facilitate recognition of guideline information derived from internationally published research and guidelines. The findings therefore highlight the substantial influence of global biomedical research on national guideline development.
From an educational perspective, LLMs may offer potential value as supplementary tools for rapid information retrieval and guideline familiarisation. However, their outputs should always be interpreted critically and verified by qualified clinicians before being applied in clinical practice.
Several limitations should be acknowledged. First, the study examined only two German oral oncology guidelines, which limits the generalisability of the findings to other medical disciplines. Second, the methodological design involved verbatim input of guideline statements and therefore assessed recognition rather than responses to realistic clinical queries. Third, the investigation did not evaluate case‐based reasoning, diagnostic interpretation or therapeutic decision‐making.
Future studies should therefore examine multiple LLMs across a broader range of clinical scenarios, including patient‐specific questions and complex guideline interpretation tasks. Such research will be essential to determine the appropriate role of AI systems within future evidence‐based clinical decision‐support environments.
5. Conclusion
When guideline recommendations were presented verbatim, GPT‐5 demonstrated complete concordance with German CPGs addressing oral oncology. These findings indicate that the model possesses a strong capacity to recognise and retrieve established guideline information from its training corpus. However, the present study primarily evaluated recognition of predefined statements rather than independent clinical reasoning or guideline interpretation within real clinical scenarios.
Accordingly, current LLMs should be regarded primarily as educational and informational support tools that may facilitate guideline familiarisation. They should not be considered as substitutes for expert clinical judgement. Further investigations incorporating case‐based clinical scenarios and comparisons across multiple AI systems will be necessary to clarify the potential role of LLMs in future evidence‐based clinical decision support.
Author Contributions
Julius Hirsch: methodology, software, validation, resources, investigation, writing – original draft, writing – review and editing and visualisation. Keskanya Subbalekha: conceptualisation, software, validation, formal analysis, data curation, writing – original draft, writing – review and editing and visualisation. Chatpong Tangmanee: conceptualisation, methodology, software, validation, resources, formal analysis, data curation, writing – original draft, writing – review and editing and visualisation. Christian Stoll: conceptualisation, methodology, validation, formal analysis, data curation, writing – review and editing, visualisation, supervision and project administration. Poramate Pitak‐Arnnop: conceptualisation, methodology, software, validation, resources, investigation, formal analysis, data curation, writing – original draft, writing – review and editing, visualisation, supervision and project administration.
Funding
The authors have nothing to report.
Ethics Statement
The authors have nothing to report.
Consent
The authors have nothing to report.
Conflicts of Interest
The authors declare no conflicts of interest.
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
References
- 1. Bagde H., Dhopte A., Alam M. K., and Basri R., “A Systematic Review and Meta‐Analysis on ChatGPT and Its Utilization in Medical and Dental Research,” Heliyon 9 (2023): e23050, 10.1016/j.heliyon.2023.e23050. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Wei Q., Yao Z., Cui Y., Wei B., Jin Z., and Xu X., “Evaluation of ChatGPT‐Generated Medical Responses: A Systematic Review and Meta‐Analysis,” Journal of Biomedical Informatics 151 (2024): 104620, 10.1016/j.jbi.2024.104620. [DOI] [PubMed] [Google Scholar]
- 3. AWMF , “S2k‐Leitlinie Diagnostik und Management von Vorläuferläsionen des oralen Plattenepithelkarzinoms in der Zahn‐, Mund‐ und Kieferheilkunde. Version 2.1,” accessed October 14, 2025, https://register.awmf.org/de/leitlinien/detail/007‐092.
- 4. AWMF , “S3‐Leitlinie Diagnostik und Therapie des Mundhöhlenkarzinoms. Version 3.0,” accessed October 14, 2025, https://register.awmf.org/de/leitlinien/detail/007‐100OL.
- 5. Sah M. and Gogebakan K., “Comparing ChatGPT Responses With Clinical Practice Guidelines for Diagnosis, Prevention, and Treatment of Diabetes,” in 2nd International Congress of Electrical and Computer Engineering, ed. Seyman M. N. (Springer, 2024), 10.1007/978-3-031-52760-9_5. [DOI] [Google Scholar]
- 6. Dagher T., Dwyer E. P., Baker H. P., Kalidoss S., and Strelzow J. A., ““Dr. AI Will See You Now”: How Do ChatGPT‐4 Treatment Recommendations Align With Orthopaedic Clinical Practice Guidelines?,” Clinical Orthopaedics and Related Research 482 (2024): 2098–2106, 10.1097/CORR.0000000000003234. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Li C. P., Jakob J., Menge F., Reißfelder C., Hohenberger P., and Yang C., “Comparing ChatGPT‐3.5 and ChatGPT‐4's Alignments With the German Evidence‐Based S3 Guideline for Adult Soft Tissue Sarcoma,” iScience 27 (2024): 111493, 10.1016/j.isci.2024.111493. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Carl N., Schramm F., Haggenmüller S., et al., “Large Language Model Use in Clinical Oncology,” npj Precision Oncology 8 (2024): 240, 10.1038/s41698-024-00733-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Rossettini G., Cook C., Palese A., Pillastrini P., and Turolla A., “Pros and Cons of Using Artificial Intelligence Chatbots for Musculoskeletal Rehabilitation Management,” Journal of Orthopaedic and Sports Physical Therapy 53 (2023): 728–734, 10.2519/jospt.2023.12000. [DOI] [PubMed] [Google Scholar]
- 10. Gianola S., Bargeri S., Castellini G., et al., “Performance of ChatGPT Compared to Clinical Practice Guidelines in Making Informed Decisions for Lumbosacral Radicular Pain: A Cross‐Sectional Study,” Journal of Orthopaedic and Sports Physical Therapy 54 (2024): 222–228, 10.2519/jospt.2024.12151. [DOI] [PubMed] [Google Scholar]
- 11. Rossettini G., Bargeri S., Cook C., et al., “Accuracy of ChatGPT‐3.5, ChatGPT‐4o, Copilot, Gemini, Claude, and Perplexity in Advising on Lumbosacral Radicular Pain Against Clinical Practice Guidelines: Cross‐Sectional Study,” Frontiers in Digital Health 7 (2025): 1574287, 10.3389/fdgth.2025.1574287. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Marchi F., Bellini E., Iandelli A., Sampieri C., and Peretti G., “Exploring the Landscape of AI‐Assisted Decision‐Making in Head and Neck Cancer Treatment: A Comparative Analysis of NCCN Guidelines and ChatGPT Responses,” European Archives of Oto‐Rhino‐Laryngology 281 (2024): 2123–2136, 10.1007/s00405-024-08525-z. [DOI] [PubMed] [Google Scholar]
- 13. Washington C. J., Abouyared M., Karanth S., et al., “The Use of Chatbots in Head and Neck Mucosal Malignancy Treatment Recommendations,” Otolaryngology and Head and Neck Surgery 171 (2024): 1062–1068, 10.1002/ohn.818. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
