Abstract
Introduction and aims
Large language models (LLMs), such as ChatGPT and Gemini, are increasingly being used in medical domains, including dental diagnostics. Despite advancements in image-based deep learning systems, LLM diagnostic capabilities in oral and maxillofacial surgery (OMFS) for processing multi-modal imaging inputs remain underexplored. Radiolucent jaw lesions represent a particularly challenging diagnostic category due to their varied presentations and overlapping radiographic features. This study evaluated diagnostic performance of ChatGPT 4o and Gemini 2.5 Pro using real-world OMFS radiolucent jaw lesion cases, presented in multiple-choice (MCQ) and short-answer (SAQ) formats across 3 imaging conditions: panoramic radiography only, panoramic + CT, and panoramic + CT + pathology.
Methods
Data from 100 anonymized patients at Wonkwang University Daejeon Dental Hospital were analyzed, including demographics, panoramic radiographs, CBCT images, histopathology slides, and confirmed diagnoses. Sample size was determined based on institutional case availability and statistical power requirements for comparative analysis. ChatGPT and Gemini diagnosed each case under 6 conditions using 3 imaging modalities (P, P+C, P+C+B) in MCQ and SAQ formats. Model accuracy was scored against expert-confirmed diagnoses by 2 independent evaluators. McNemar's and Cochran's Q tests evaluated statistical differences across models and imaging modalities.
Results
For MCQ tasks, ChatGPT achieved 66%, 73%, and 82% accuracies across the P, P+C, and P+C+B conditions, respectively, while Gemini achieved 57%, 62%, and 63%, respectively. In SAQ tasks, ChatGPT achieved 34%, 45%, and 48%; Gemini achieved 15%, 24%, and 28%, respectively. Accuracy improved significantly with additional imaging data for ChatGPT; ChatGPT consistently outperformed Gemini across all conditions (P < .001 for MCQ; P = .008 to < .001 for SAQ). MCQ format, which incorporates a human-in-the-loop (HITL) structure, showed higher overall performance than SAQ.
Conclusion
ChatGPT demonstrated superior diagnostic performance compared to Gemini in OMFS diagnostic tasks when provided with richer multimodal inputs. Diagnostic accuracy increased with additional imaging data, especially in MCQ formats, suggesting LLMs can effectively synthesize radiographic and pathological data.
Clinical Relevance
LLMs have potential as diagnostic support tools for OMFS, especially in settings with limited specialist access. Presenting clinical cases in structured formats using curated imaging data enhances LLM accuracy and underscores HITL integration. Although current LLMs show promising results, further validation using larger datasets and hybrid AI systems are necessary for broader contextualised, clinical adoption.
Keywords: Large language models, Artificial intelligence, Odontogenic cysts, Odontogenic tumors, Decision making
Introduction
Large-language models (LLMs), such as GPT-based architectures, have rapidly evolved in their ability to comprehend and generate domain-specific knowledge. These models have demonstrated notable performance in complex professional tasks, including clinical diagnostics, board examinations, and even question generation for educational use.1, 2, 3, 4 In particular, their applications in dental education and assessment have been investigated, such as solving questions on national licensing exams2,3 and generating relevant clinical scenarios.1
Artificial intelligence has emerged as a transformative force in dentistry, offering unprecedented opportunities to enhance diagnostic accuracy, treatment planning, and patient care across multiple dental specialties.5,6 The integration of AI technologies, particularly in oral and maxillofacial surgery (OMFS), represents a paradigm shift toward more precise and efficient clinical decision-making processes.
Deep learning has substantially contributed to the development of diagnostic support systems for OMFS. Previous studies have used convolutional neural networks (CNNs) and other deep learning architectures to assess osseointegration,7 predict surgical difficulty,8,9 determine indications for orthognathic surgery,10 identify cracked teeth,11 and detect cranio-spinal discrepancies.12 These tools typically utilize panoramic, cephalometric, and CT images to perform predictive analyses.
Radiolucent jaw lesions present particular diagnostic challenges in OMFS due to their diverse etiology, overlapping radiographic presentations, and the complexity of differential diagnosis. These lesions encompass a broad spectrum of pathologies, including odontogenic cysts, odontogenic tumors, and various inflammatory conditions, each requiring distinct treatment approaches. The selection of radiolucent lesions for this study was based on their clinical frequency, diagnostic complexity, and the need for comprehensive multimodal imaging evaluation in clinical practice.
Despite these developments, direct comparisons among LLMs in OMFS diagnosis, particularly in the context of multimodal image interpretation, remain underexplored. No previous studies have examined whether providing LLMs with richer imaging inputs (e.g., panoramic + CT + pathology) yields substantially different diagnostic outcomes across structured and open-ended tasks.
This study aimed to evaluate the diagnostic performance of 2 state-of-the-art LLMs, ChatGPT 4o and Gemini 2.5 Pro, using a structured question set based on real-world OMFS cases, particularly those with frequent radiolucent lesions of the jaw. Each model answered questions under 3 imaging conditions—panoramic only (P), panoramic + CT (P+C), and panoramic + CT + pathology (P+C+B)—across both the multiple-choice question (MCQ) and short-answer question (SAQ) formats. By integrating insights from recent LLM and deep learning studies, this study seeks to clarify the role of generative AI in the OMFS diagnostic workflow.
Methods
Ethical approval
This study was reviewed by the Public Institutional Review Board of the Ministry of Health and Welfare, Republic of Korea, and was exempt from ethical review (IRB No. P01-202506-01-017, approved on June 11, 2025). This study was a retrospective analysis of anonymized patient data. No personally identifiable information was used, and there was no direct interaction with human participants.
Study design
This study used data from 100 patients who underwent biopsy at Wonkwang University Daejeon Dental Hospital, Republic of Korea (Figure 1). The sample size of 100 cases (Table 1) was determined based on institutional case availability over the study period and statistical power requirements for comparative analysis between 2 independent groups, providing adequate power (>80%) to detect clinically meaningful differences in diagnostic accuracy (effect size ≥0.3) with α=0.05.
Fig. 1.
Flow diagram showing the process by which large language models diagnose radiolucent lesions of the jaw in this study. MCQ, Multiple-choice question; SAQ, Short-answer question; P, panoramic radiograph only; P+C, panoramic radiograph + CT image; P+C+B, panoramic radiograph + CT image + histopathologic slide.
Table 1.
Distribution of radiolucent lesions used in this study.
| Diagnosis | count |
|---|---|
| Dentigerous cyst | 24 |
| Radicular cyst | 21 |
| Simple bone cyst | 11 |
| Odontogenic keratocyst | 15 |
| Ameloblastoma | 9 |
| Glandular odontogenic cyst | 6 |
| Nasopalatine duct cyst | 5 |
| Postoperative maxillary cyst* | 5 |
| Buccal bifurcation cyst | 1 |
| Lateral periodontal cyst | 1 |
| Mucous retention cyst⁎⁎ | 1 |
| Paradental cyst | 1 |
| Total | 100 |
Postoperative maxillary cyst refers to maxillary cysts developing after surgical procedures, though this term requires standardization with current WHO classification.
The mucous retention cyst case presented with significant radiolucent characteristics on imaging, justifying its inclusion in this study despite typical classification differences.
The following information was collected from each patient: sex, age, panoramic radiograph, cone-beam CT (CBCT) image, histopathology slide image, and final pathological diagnosis. Each case was diagnosed and named based on the definitive diagnosis made by an oral pathology specialist following lesion excision from patients who visited Wonkwang University Daejeon Dental Hospital. All diagnoses were subsequently verified by a second independent oral pathology specialist to ensure diagnostic accuracy and minimize subjectivity bias.
Two LLMs, ChatGPT 4o and Gemini 2.5 Pro were evaluated and tested using their latest available versions, as of July 13, 2025. This study will be conducted between July 1 and 31, 2025.
Image preprocessing and technical specifications
All images underwent standardized preprocessing to ensure consistency. Panoramic radiographs were normalized to 1024 × 512 pixel resolution, and CBCT images were presented as standardized axial, coronal, and sagittal views. Histopathology slides were captured at 40 × magnification and cropped to regions of interest (ROI) measuring 512 × 512 pixels. Color normalization was applied to histopathology images using standard H&E staining parameters. All lesion sites on radiographic images were visually emphasized using red circles for clarity.
Token limits were set at 4,096 tokens for input prompts and 1,024 tokens for model responses to ensure consistent processing across both platforms.
Question generation
Data from 100 patients were processed into 2 formats: MCQ and SAQ (Figure 2). Each format was prepared under 3 different imaging conditions: panoramic radiograph only (P); panoramic radiograph + CT image (P+C); and panoramic radiograph + CT image + histopathological slide (P+C+B). The MCQ version contained 7 answer choices (simple bone cyst, dentigerous cyst, radicular cyst, nasopalatine duct cyst, odontogenic keratocyst, ameloblastoma, and none of these). The order of image presentation was randomized for each case to minimize sequence bias.
Fig. 2.
Example of diagnosing radiolucent lesions of the jaw by large language models based on age, sex, panoramic radiography, computed tomography, and histopathologic findings.
Information presentation and model prompting
Each LLM was provided with the full set of questions for each version, with 3 independent runs performed for each condition to account for inherent variability in generative output. The final accuracy scores represent the majority consensus across the 3 runs. The following prompts were used:
For MCQ format: "From the following 100 patient cases, select the appropriate diagnosis. Each question includes 7 choices. The provided images are (panoramic radiograph / CT image / histopathologic slide)."
For SAQ format: "Based on the given information and images, infer the most likely diagnosis. The provided images are (panoramic radiograph / CT image / histopathologic slide)."
Performance evaluation and statistical analysis
LLM performance was evaluated using 2 primary criteria. First, for quantitative assessment, each response was classified as either correct or incorrect based on a predefined answer key. For SAQ responses, a standardized grading rubric was developed that awarded full credit for exact diagnostic matches, partial credit for diagnostically relevant broader categories (e.g., "odontogenic cyst" for specific cyst types), and no credit for incorrect diagnoses. Abbreviations and synonymous terms were accepted as correct. Two independent evaluators graded all SAQ responses, with inter-rater agreement assessed using Cohen's kappa coefficient (κ = 0.89, indicating excellent agreement).
Statistical analysis was conducted to compare the performance of the 2 models and to evaluate within-model variation across imaging modalities. McNemar's test was used to assess paired categorical differences between the models, while Cochran's Q test was used to evaluate differences among the 3 imaging conditions (P, P+C, and P+C+B) within each model. These non-parametric tests were chosen as they are appropriate for paired categorical data and do not require assumptions of normality or independence across conditions.
Results
Under the MCQ conditions, ChatGPT achieved 66% accuracy in the panoramic-only (P) condition, 73% in the panoramic + CT (P+C) condition, and 82% in the panoramic + CT + pathology (P+C+B) condition. Gemini showed accuracies of 57%, 62%, and 63%, respectively, under the same conditions. The accuracy of ChatGPT increased by 7 percentage points when transitioning from P to P+C, and by 9 percentage points from P+C to P+C+B. Gemini showed increases of 5 and 1 percentage points, respectively. Differences between stages were statistically significant for ChatGPT according to McNemar's test (P→P+C, P < .001; P+C→P+C+B, P = .004), while for Gemini, only the first transition showed significance (P→P+C, P = .004; P+C→P+C+B, P = .265).
In SAQ conditions, ChatGPT achieved 34%, 45%, and 48% accuracy in the P, P+C, and P+C+B conditions, respectively. Gemini recorded accuracies of 15%, 24%, and 28%, respectively. ChatGPT's accuracy increased by 11 percentage points from P to P+C and by 3 percentage points from P+C to P+C+B. Gemini showed an increase of 9 percentage points from P to P+C and 4 percentage points from P+C to P+C+B. According to McNemar's test, the differences were statistically significant for ChatGPT in both transitions (P→P+C, P < .001; P+C→P+C+B, P = .009). For Gemini, both transitions showed significant differences (P→P+C, P = .002; P+C→P+C+B, P = .045).
In all conditions, ChatGPT showed higher accuracy than Gemini, and this difference was statistically significant across all MCQ (P, P+C, P+C+B; P < .001) and SAQ conditions (P, P+C, P+C+B; P = .008, P < .001, P < .001, respectively). Both models showed a trend of increased accuracy as additional imaging data were provided, although the magnitude of the increase varied by model and condition (Figure 3 and Table 2).
Fig. 3.
Performance comparison of large language models for radiolucent jaw lesions: (A) comparison of answer accuracy between ChatGPT 4o and Gemini 2.5 Pro in MCQ format, (B) comparison of answer accuracy between ChatGPT 4o and Gemini 2.5 Pro in SAQ format, (C) comparison of answer accuracy of ChatGPT 4o between MCQ and SAQ formats; (D) comparison of answer accuracy of Gemini 2.5 Pro between MCQ and SAQ formats, (E) condition-wise comparison of ChatGPT 4o's answer accuracy in MCQ format, (F) condition-wise comparison of Gemini 2.5 Pro's answer accuracy in MCQ format, (G) condition-wise comparison of ChatGPT 4o's answer accuracy in SAQ format, and (H) condition-wise comparison of Gemini 2.5 Pro's answer accuracy in SAQ format.
Statistical significance was indicated using the following notation: 'n.s.' for not significant (P ≥ .05), '' for P < .05, '' for P < .01, and '' for P < .001. Higher numbers of asterisks indicate a higher level of statistical significance.
Table 2.
Performance of large language models for frequent radiolucent lesion of jaw.
| ChatGPT 4o | Gemini 2.5 pro | |||
|---|---|---|---|---|
| MCQ | SAQ | MCQ | SAQ | |
| P | 66% | 34% | 57% | 15% |
| PC | 73% | 45% | 62% | 24% |
| PCB | 82% | 48% | 63% | 28% |
MCQ, multiple-choice question; SAQ, short-answer question; P, panoramic radiograph only; P+C, panoramic radiograph + CT image; P+C+B, panoramic radiograph + CT image + histopathological slide.
Discussion
LLMs are increasingly being explored for diagnostic applications in dentistry by expert users and general clinicians, owing to their growing accessibility. In oral pathology, Tassoker investigated the diagnostic ability of ChatGPT in 123 challenging oral and maxillofacial cases, highlighting that while LLMs show considerable promise in recognizing pathology patterns, their consistency varies across lesion types and presentation formats.¹³ Our study builds upon this by systematically evaluating LLMs under controlled multi-modal image inputs and structured/unstructured formats, thereby adding granularity to the evidence on when and how LLMs perform best.
The superior performance of ChatGPT compared to Gemini observed in our study may be attributed to several factors. ChatGPT's training architecture and multimodal integration capabilities appear more optimized for medical image analysis tasks. Additionally, differences in training datasets, with ChatGPT potentially having greater exposure to medical and pathological content, may contribute to its enhanced diagnostic accuracy. The consistently lower performance of Gemini across all conditions suggests fundamental differences in how these models process and integrate visual and contextual information in medical contexts.
However, a core challenge with LLMs is the phenomenon of hallucinations, where models produce plausible but incorrect content.13,14 This underscores the importance of human-in-the-loop (HITL) systems, which integrate expert oversight in AI-assisted workflows.15,16 In this study, MCQs—structured formats that inherently provide contextual constraints—acted as a form of HITL and resulted in higher diagnostic accuracy compared to open-ended SAQs. This aligns with earlier findings that structured prompts improve reliability.1,2
Our findings support the broader applications of AI in oral pathology. Rewthamrongsris et al. compared LLMs and CNNs in diagnosing oral lichen planus and demonstrated that LLMs benefited significantly from example-guided prompts and structured differentials, whereas CNNs were more consistent in pixel-based classification tasks.17 This highlights a complementary dynamic: CNNs excel in raw image analysis, whereas LLMs offer interpretive flexibility, especially when enhanced with guided clinical reasoning.
Schmidl et al. applied deep learning models for the classification of oral cancer and leukoplakia and found that image-based AI can achieve high diagnostic accuracy when trained on well-labeled datasets.18 Unlike these CNN-based models, which require extensive labeled training data, our LLM approach operates under a zero-shot inference paradigm, requiring no task-specific model training. This advantage makes LLMs adaptable to clinics with a limited AI infrastructure or domain-specific datasets.
Although LLMs showed improvement with the addition of CT and histopathological images, the rationale behind their decisions remains opaque. Furthermore, reliance on static training cutoffs can lead to outdated outputs unless continuously updated.19 The performance disparity between ChatGPT and Gemini also reveals that LLMs do not generalize equally, even under identical conditions. This discrepancy was also reported by Tassoker, who noted variations between GPT and other models in interpreting nuanced maxillofacial lesions.20
Privacy and data governance represent critical considerations for clinical implementation. The use of diagnostic AI in clinical environments raises ethical issues, particularly when real patient imaging data are uploaded to external platforms.21 Proper anonymization and institutional safeguards are essential before the integration of LLMs into clinical practice. Future implementations should consider local deployment options or secure API configurations to maintain patient confidentiality while leveraging AI capabilities.
From a clinical perspective, our results suggest that LLMs, particularly ChatGPT 4o, can serve as effective adjunctive tools for OMFS diagnosis, especially when image data are rich and structured guidance is available. Structured LLM-based tools can enhance the diagnostic capacity of general practitioners in regions lacking access to oral and maxillofacial radiologists. Importantly, these systems should be deployed with an awareness of their limitations and in concert with human oversight.
This study has several important limitations that affect the generalizability of findings. First, the dataset was derived from a single institution, which may limit the external validity of results across different populations and imaging protocols. Second, the distribution of lesion types showed considerable imbalance, with common lesions like dentigerous cysts and radicular cysts representing nearly half of all cases, while rare entities had minimal representation. This imbalance may have influenced model performance and limits the applicability of findings to less common pathologies.
Third, the MCQ format included a "none of these" option that encompassed several ground-truth diagnoses not explicitly listed among the choices. This design limitation may have artificially inflated error rates for certain rare lesions and complicated error interpretation. Fourth, the use of single histopathology images per case may not fully represent the histological complexity of some lesions, potentially limiting diagnostic accuracy.
Conclusions
ChatGPT 4o demonstrated superior diagnostic capabilities compared to Gemini 2.5 Pro in OMFS diagnostic tasks across all formats and imaging conditions. Incorporating multimodal image inputs notably enhanced performance, particularly for ChatGPT. The structured MCQ format consistently outperformed open-ended SAQ format, highlighting the importance of human-in-the-loop integration in AI-assisted diagnosis.
Although promising, LLMs still have limitations in terms of their interpretability and reasoning transparency. Future research directions should focus on developing hybrid systems that integrate deep-learning image analysis with natural language understanding for comprehensive diagnostic support. Multi-institutional validation studies with larger, more diverse datasets are needed to establish the clinical utility and safety of LLMs in OMFS practice.
Specific next steps include: (1) development of explainable AI frameworks that provide transparent diagnostic reasoning, (2) creation of hybrid CNN-LLM systems that combine spatial accuracy with contextual interpretation, (3) establishment of standardized evaluation protocols for AI diagnostic tools in oral pathology, and (4) investigation of LLM performance across broader categories of maxillofacial pathology including soft tissue lesions and mixed radiopaque-radiolucent entities.
Ethics approval and consent to participate
This study was approved by the public Institutional Review Board (IRB) of South Korea (P01-202506-01-017). Due to the absence of personal identification data, the IRB waived the need for written or verbal informed consent from the subjects.
Availability of data and materials
The data used in this study can be made available, if required, within the regulation boundaries for data protection. They are available from the corresponding author (B.C.K.) on reasonable request.
Author contributions
The study was conceived by B.C.K., who conducted the experiments. K.G.K. conducted the experiments. K.G.K. and B.C.K. generated the data. K.G.K. and B.C.K. analyzed and interpreted the data. K.G.K. and B.C.K. prepared the manuscript. K.G.K. and B.C.K. have read and approved the final version of the manuscript.
Funding
This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korean Government (MSIT) (RS-2024-00451221). The funding bodies played no role in the study design, data collection, analysis, interpretation, or the writing of the manuscript.
Conflict of interest
None disclosed.
References
- 1.Kim K., Mun S.B., Kim Y.J., Kim B.C., Kim KG. How valuable are the questions and answers generated by large language models in oral and maxillofacial surgery? PLOS One. 2025;20(5) doi: 10.1371/journal.pone.0322529. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Kim W., Kim B.C., Yeom HG. Performance of large language models on the Korean dental licensing examination: a comparative study. Int Dent J. 2025;75(1):176–184. doi: 10.1016/j.identj.2024.09.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Zong H., Wu R., Cha J., Wang J., Wu E., Li J., et al. Large language models in worldwide medical exams: platform development and comprehensive analysis. J Med Internet Res. 2024;26 doi: 10.2196/66114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Chen Y., Huang X., Yang F., Lin H., Lin H., Zheng Z., et al. Performance of ChatGPT and Bard on the medical licensing examinations varies across different cultures: a comparison study. BMC Med Educ. 2024;24(1):1372. doi: 10.1186/s12909-024-06309-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Samaranayake L., Tuygunov N., Schwendicke F., Osathanon T., Khurshid Z., Boymuradov S.A., Cahyanto A. The transformative role of artificial intelligence in dentistry: a comprehensive overview. part 1: fundamentals of AI, and its contemporary applications in dentistry. Int Dent J. 2025;75(2):383–396. doi: 10.1016/j.identj.2025.02.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Tuygunov N., Samaranayake L., Khurshid Z., Rewthamrongsris P., Schwendicke F., Osathanon T., Yahya NA. The transformative role of artificial intelligence in dentistry: a comprehensive overview part 2: the promise and perils, and the international dental federation communique. Int Dent J. 2025;75(2):397–404. doi: 10.1016/j.identj.2025.02.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Oh S., Kim Y.J., Kim J., Jung J.H., Lim H.J., Kim B.C., et al. Deep learning-based prediction of osseointegration for dental implant using plain radiography. BMC Oral Health. 2023;23(1):208. doi: 10.1186/s12903-023-02921-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Yoo J.H., Yeom H.G., Shin W., Yun J.P., Lee J.H., Jeong S.H., et al. Deep learning based prediction of extraction difficulty for mandibular third molars. Sci Rep. 2021;11(1):1954. doi: 10.1038/s41598-021-81449-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Shin W., Yeom H.G., Lee G.H., Yun J.P., Jeong S.H., Lee J.H., et al. Deep learning based prediction of necessity for orthognathic surgery of skeletal malocclusion using cephalogram in Korean individuals. BMC Oral Health. 2021;21(1):130. doi: 10.1186/s12903-021-01513-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Jeong S.H., Woo M.W., Shin D.S., Yeom H.G., Lim H.J., Kim B.C., et al. Three-dimensional postoperative results prediction for orthognathic surgery through deep learning-based alignment network. J Pers Med. 2022;12(6):998. doi: 10.3390/jpm12060998. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Mun S.B., Kim J., Kim Y.J., Seo M.S., Kim B.C., Kim KG. Deep learning-based prediction of indication for cracked tooth extraction using panoramic radiography. BMC Oral Health. 2024;24(1):952. doi: 10.1186/s12903-024-04721-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Jeong S.H., Yun J.P., Yeom H.G., Kim H.K., Kim BC. Deep-Learning-Based detection of cranio-spinal differences between skeletal classification using cephalometric radiography. Diagnostics. 2021;11(4):591. doi: 10.3390/diagnostics11040591. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Ji Z., Lee N., Frieske R., Yu T., Su D., Xu Y., et al. Survey of hallucination in natural language generation. ACM Comput Surv. 2023;55(12):1–38. [Google Scholar]
- 14.Xu Z., Jain S., Kankanhalli M. Hallucination is inevitable: an innate limitation of large language models. arXiv. 2025 [Google Scholar]
- 15.Amirizaniani M., Yao J., Lavergne A., Okada E.S., Chadha A., Roosta T., et al. LLMAuditor: a framework for auditing large language models using human-in-the-loop. arXiv. 2024 [Google Scholar]
- 16.Wang J, Guo B, Chen L. Human-in-the-loop machine learning: a macro-micro perspective.
- 17.Rewthamrongsris P., Burapacheep J., Phattarataratip E., Kulthanaamondhita P., Tichy A., Schwendicke F., et al. Image-based diagnostic performance of LLMs vs CNNs for oral lichen planus: example-guided and differential diagnosis. Int Dent J. 2025;75(4) doi: 10.1016/j.identj.2025.100848. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Schmidl B., Hütten T., Pigorsch S., et al. Artificial intelligence for image recognition in diagnosing oral and oropharyngeal cancer and leukoplakia. Sci Rep. 2025;15(1):3625. doi: 10.1038/s41598-025-85920-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Cheng J., Marone M., Weller O., Lawrie D., Khashabi D., Durme BV. Dated data: tracing knowledge cutoffs in large language models. arXiv. 2024 [Google Scholar]
- 20.Tassoker M. Exploring ChatGPT's potential in diagnosing oral and maxillofacial pathologies: a study of 123 challenging cases. BMC Oral Health. 2025;25(1):1187. doi: 10.1186/s12903-025-06444-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Yao Y., Duan J., Xu K., Cai Y., Sun Z., Zhang Y. A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly. High-Confid Comput. 2024;4(2) [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data used in this study can be made available, if required, within the regulation boundaries for data protection. They are available from the corresponding author (B.C.K.) on reasonable request.



