Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Apr 2;24(8):e1087–e1095. doi: 10.1111/ddg.70039

Beyond black‐box AI: Comparing ChatGPT‐4 interpretability and accuracy to CNNs in melanocytic lesions diagnosis

Ofer Reiter 1,2, Cristian Navarrete‐Dechent 3, Mor Atlas 4, Nir Nathansohn 5, Yaron Ben Mordehai 6, Tomer Mimouni 2, Romi Gleicher 7, Mahdi Awwad 8, Itay Cohen 9, Ziad Khamaysi 7,10, Jonathan Shapiro 6,
PMCID: PMC13435119  PMID: 41925624

Summary

Background

Artificial intelligence (AI) algorithms have advanced and recently shown high accuracy in diagnosing skin cancer from dermoscopic images. This study compared the diagnostic performance of the large language model ChatGPT‐4 with that of specialized convolutional neural network (CNN)‐based models in analyzing melanocytic lesions.

Patients and Methods

A cross‐sectional comparative study was conducted using 117 dermoscopic images. The performance of ChatGPT‐4 was assessed under two conditions: diagnosing lesions directly without annotations and diagnosing after annotating dermoscopic features. Results were compared with CNN‐based models (YPSONO and ResNet) and human expert evaluations. The confusion matrices of all the models were calculated in addition to the diagnostic accuracy, sensitivity, specificity, and interobserver agreement (Cohen's Kappa).

Results

ChatGPT‐4 achieved 92 % sensitivity, 89 % specificity, and an accuracy of 89.7 % in direct diagnosis. When annotations were required, sensitivity and specificity dropped to 68 % and 64 %, respectively. Agreement with experts on dermoscopic patterns was minimal (Cohen's Kappa = 0.13). ChatGPT‐4 outperformed CNN models in direct diagnosis but exhibited notable limitations in describing dermoscopic features.

Conclusions

ChatGPT‐4 demonstrated promising potential for accurate melanoma versus nevus classification without annotations, surpassing CNN‐based models. However, its limited ability to describe dermoscopic features accurately highlights the need for further research and training.

Keywords: Annotations, artificial intelligence, dermoscopy, diagnostic accuracy, GPT, interobserver agreement, melanoma, nevus

INTRODUCTION

In 2017, Esteva et al. published a landmark research letter in Nature, demonstrating that Artificial Intelligence (AI) algorithms can diagnose skin cancer from dermoscopic images. 1 Since then, numerous studies have investigated computer vision convolutional neural network (CNN) algorithms, demonstrating that these can achieve diagnostic accuracy with dermoscopic images that is non‐inferior and sometimes superior to that of human experts. 2 , 3 , 4 Unlike previous methodologies, these algorithms do not rely on predefined diagnostic criteria. Instead, they are trained on thousands of labeled images, enabling them to develop their own diagnostic criteria within the hidden layers of the networks, which remain opaque and incomprehensible to humans. 5 Additionally, these algorithms are not generative, as they do not produce clinical descriptions of dermoscopic images.

Over the past two years, generative AI algorithms, specifically generative pre‐trained transformers (GPTs), have gained popularity. These algorithms are distinguished by their ability to generate new content and provide insights that extend beyond predefined datasets. This developmental leap deserves individual scrutiny, as the assumptions, methods, and implications differ substantially between generations of AI. One of the most renowned algorithms in this category is ChatGPT (OpenAI), which has already demonstrated impressive capabilities in solving medical challenge problems. 6 While ChatGPT itself is primarily a text‐based model, OpenAI has introduced multimodal models, including ChatGPT‐4, which can process both text and images. 7 However, OpenAI does not disclose the specifics of how these models are trained, including which images or databases were used in the training process.

In recent studies 8 , 9 the diagnostic accuracy of ChatGPT Vision was evaluated using dermoscopic images retrieved from publicly available image databases. Their findings indicated that ChatGPT's diagnostic accuracy was inferior to that of other AI algorithms. However, in one of these studies, 9 the authors suggested that ChatGPT may still hold potential in supporting clinicians by providing detailed dermoscopy descriptions, though such applications must be approached with caution, as current models raise important concerns regarding data protection, patient consent, and transparency in the storage and use of input data.

Moreover, studies have shown that there is an overreliance on generative AI in many fields 10 , 11 including medicine. Notably, physicians have been shown to align their diagnosis with AI‐generated suggestions, 12 which may lead to overreliance on incorrect AI advice. 13 Therefore, there is a need to calibrate the trust of physicians in these systems 14 and understand the quality of the generative AI performance in the field of dermatology. This study aims to evaluate ChatGPT4's performance (GPT‐4‐turbo version available at the time of analysis) in analyzing dermoscopic images of melanocytic lesions, providing a dermoscopic description and diagnosis, and to compare its performance with that of available online CNN models (Ypsono and ResNet).

Given the rapid pace of advancement in large language models, newer versions beyond ChatGPT‐4 may offer improved performance. At the time this study was conducted, ChatGPT‐4 was the most widely used and accessible multimodal generative model of ChatGPT. As with any AI research, the results should be interpreted as a snapshot of the technology at that specific point in time. Despite inevitable future improvements, such analyses remain valuable for understanding the capabilities and limitations of generative AI in dermatology, providing a reference point against which subsequent models can be compared.

PATIENTS AND METHODS

This cross‐sectional comparative study was approved by the institutional review board at Maccabi Health Services under protocol 0012‐24‐MHS.

Data

Given that the databases used to train ChatGPT‐4 have not been made publicly available, the dermoscopic images analyzed in this study were randomly selected from the private database of a dermatology practice. All images were obtained using the D200evo digital dermoscopy camera (Canfield, USA) as part of total‐body photography, with a resolution of 3096 × 2080 pixels (approximately 6.4 megapixels). The dpi value shown in the image metadata refers only to display settings and does not indicate the level of detail in the digital image. All dermoscopic images used in this study were fully anonymized prior to analysis. They consisted exclusively of dermoscopic photographs without any associated metadata, patient identifiers, or total‐body context. Because these images do not include anatomical landmarks, facial features, tattoos, or other identifiable characteristics, they meet the criteria for actual anonymization, preventing patient re‐identification. This ensures compliance with relevant data privacy and regulatory standards, supporting the secure and ethical use of image data for artificial intelligence analysis.

Inclusion criteria were: (1) availability of a dermoscopic image; (2) diagnosis of a melanocytic lesion, either benign or malignant. Diagnoses were based on histopathologic examination (melanoma vs nevi and dysplastic nevi) or on at least 12 months of dermoscopic follow‐up without any observed changes (nevi and dysplastic nevi). Exclusion criteria were: (1) unclear dermoscopic image; (2) duplicates (different dermoscopic images of the same lesion); (3) images of lesions located on the face, palms, soles, and mucosal surfaces.

All images were annotated in a consensus review by two expert dermatologists (C.N.D. and O.R.), who were blinded to the final diagnosis. The following parameters were evaluated: (1) Overall pattern: a categorical variable with the following options: reticular; globular; reticular‐globular; central homogeneous with peripheral network; central homogeneous with peripheral globules; central network with peripheral globules; structureless homogeneous; starburst; two‐component; multicomponent; diffuse negative network; and other. (2) Global organization: classified as organized or disorganized. (3) Presence of dermoscopic features according to Liopyris et al.: 15 network (typical or atypical); dots (regular or irregular); globules (regular, irregular, or rim); blotch (regular or irregular); branched streaks; peppering; scar‐like depigmentation; shiny white structures; angulated lines; negative pigment network; blue‐white veil; milky red areas; tan structureless areas; and vessels (comma, corkscrew, dotted, linear irregular, polymorphous, or milky red globules).

Analysis

For analysis by ChatGPT‐4, images were uploaded in their original resolution and evaluated between April 30, 2024, and May 10, 2024, using a chatbot specifically designed for this purpose. The configuration of the GPT chatbot is described in the online supplementary material (Supplement S1). The chatbot was asked to evaluate the lesion images for their overall pattern (organized or disorganized) and dermoscopic features, as were the dermatologists.

The performance of ChatGPT‐4 was evaluated using two distinct approaches: First, its accuracy in annotations and diagnoses was assessed. For each dermoscopic image, all annotations made by ChatGPT‐4 were compared to those provided by human readers. In addition, ChatGPT‐4 was asked to provide a diagnosis for each lesion, which was then compared to the reference standard diagnosis established by histopathology or clinical follow‐up. The second approach evaluated diagnostic performance without annotations using the prompt: “Please describe the dermoscopic lesion in a professional way and provide a differential diagnosis sorted by most probable to least probable.” To ensure reproducibility and avoid the effects of few‐shot learning, each dermoscopic image was analyzed in a separate ChatGPT session, using the same standardized prompt text for all cases. The two prompting approaches (with annotations and direct diagnosis) were conducted in independent sessions for each image. The “chat history & training” option in ChatGPT settings was disabled before analysis to prevent any storage or reuse of the input data for model training. This approach ensured that each analysis was independent and that outputs for one image could not influence responses for subsequent images.

ChatGPT‐4 generated final diagnoses directly from dermoscopic images and was compared with (1) the reference standard; (2) the publicly available YPSONO CNN model (https://dermonaut.meduniwien.ac.at/ypsono), selected as an external benchmark due to its established performance and availability as a pre‐trained, ready‐to‐use system requiring no additional training; and (3) a ResNet‐based CNN trained by us on the International Skin Imaging Collaboration dataset (ISIC; https://www.isic‐archive.com/), chosen because it represents a widely used architecture in medical image analysis and allows performance evaluation using a model optimized for our cohort. Images were uploaded individually to YPSONO.

For the ResNet model, 6,582 melanoma and 20,142 benign nevus images were resized to 224 × 224 px, normalized, and augmented (rotations, flips, crops). The dataset was partitioned into training (80 %) and test (20 %) sets, with 10 % of the training data for validation. A pre‐trained ResNet was fine‐tuned using cross‐entropy loss, the Adam optimizer, and a learning‐rate scheduler, then evaluated on the same images assessed by YPSONO and ChatGPT‐4.

Performance of the four models (ResNet, YPSONO, ChatGPT‐4 with and without annotations) on 117 images was summarized by sensitivity, specificity, positive predictive value (PPV), overall accuracy, and Cohen's κ. Sensitivity was the proportion of correctly identified melanomas; specificity, correctly identified benign nevi; PPV, true melanomas among cases flagged as melanoma; and accuracy, the proportion of all correct classifications.

Statistical analysis

Cohen's Kappa coefficient was used to measure interobserver agreement between each model's diagnoses and the ground truth diagnoses. Kappa values range from –1 to 1, where 1 indicates perfect agreement, 0 indicates agreement no better than chance, and negative values indicate disagreement. The models’ performances were compared based on their sensitivity, specificity, PPV, diagnostic accuracy, and Kappa scores to evaluate their relative strengths and weaknesses in analyzing dermoscopic images.

Moreover, we calculated the confidence interval (CI) as part of the results analysis to assess the statistical reliability of GPT's model's accuracy estimate. We applied the standard formula for the confidence interval of a proportion: p±Z1a·p^(1p^)n, where p is the observed accuracy and n is the sample size. The data were analyzed using SPSS, version 26.0 for Windows (SPSS, Inc.).

RESULTS

One hundred and fifty images were randomly selected from the database, including 42 histologically diagnosed melanomas (39 superficial spreading, three lentigo maligna), 46 histologically diagnosed nevi, and an additional 62 nevi that had follow‐up. Thirty‐three images were excluded due to either low quality (n = 7), lesions located on the face, palms, soles, or mucosa (n = 9), or duplicates (n = 17). A total of 117 images were included in the final analysis, comprising 92 nevi (59 followed for 12 months and 33 biopsied – 17 dysplastic, nine junctional, and seven compound) and 25 melanomas (22 superficial spreading and 3 lentigo maligna). All patients were Israeli Jewish individuals with Fitzpatrick skin phototypes I–III.

ChatGPT, with annotations, correctly diagnosed 17 of the 25 melanomas (sensitivity = 68 %, specificity = 64 %, precision = 34 %, and F1 = 0.45). The overall diagnostic accuracy was 64.96 %. ChatGPT demonstrated higher specificity for nevi diagnosed by the absence of changes during follow‐up compared to nevi diagnosed histopathologically (71.2 % versus 51.5 %; p = 0.03).

Concordance rate between ChatGPT‐4 and the human dermoscopy experts was 24 % in the overall dermoscopic patterns, with a slight interobserver agreement (Cohen's Kappa = 0.13; 95 % CI 0.04–0.21). Human dermoscopic experts marked most lesions as reticular, while ChatGPT marked most lesions as globular. The globular pattern showed the highest concordance rate (61.5 %) between experts and ChatGPT. None of the lesions identified as reticular‐globular, two‐component, or central network with peripheral globules by the experts were marked as such by ChatGPT‐4. Conversely, ChatGPT marked five lesions as “other,” whereas the human readers did not mark any lesions in this category (Table 1). Twenty‐eight percent of the nevi that were marked as structureless by the human readers were marked as globular by ChatGPT‐4, 27 % of central homogeneous with peripheral network nevi were marked as reticular nevi, and 19 % of reticular nevi were marked as multicomponent (Table 2).

TABLE 1.

Comparison of human and ChatGPT performance in identifying dermoscopic patterns.

N = 117 Human, n (%) ChatGPT, n (%) Interobserver agreement (95% CI) Sensitivity* (%) Specificity* (%) PPV* (%) Diagnostic accuracy* (%)
OVERALL PATTERNS
Reticular 43 (36.8) 18 (15.4) 0.13 (0.04–0.21) 18.6 66.7 44.4 23.9
Globular 13 (11.1) 24 (20.5) 61.5 55.6 33.3
Other 0 (0) 5 (4.3) N/A 84.8 0
Structureless‐homogeneous 18 (15.4) 12 (10.3) 27.8 76.7 41.7
Central Homogeneous with peripheral network 15 (12.8) 10 (8.5) 6.7 75 10
Reticular‐globular 2 (1.7) 4 (3.4) 0 87.5 0
Two‐component pattern 3 (2.6) 5 (4.3) 0 84.8 0
Multicomponent 17 (14.5) 16 (13.7) 23.5 66.7 25
Starburst 2 (1.7) 4 (3.4) 50 90 25
Central homogeneous with peripheral globules 0 (0) 9 (7.7) N/A 75.7 0
Central network with peripheral globules 2 (1.7) 8 (6.8) 0 77.8 0
Diffuse negative network 2 (1.7) 2 (1.7) 50 96.4 50
ORGANIZED yes/no 34 (29.1) 56 (47.9) 0.13 (–0.04–0.3) 58.8 56.6 35.7 57.3
DERMOSCOPIC STRUCTURES
Network 89 (76.1) 63 (53.8) 0.32 (0.17–0.48) 64 78.6 90.5 67.5
Network: typical 30 (25.6) 23 (19.7) 0.01 (–0.17–0.18) 20 80.4 26.1 65
Network: atypical 59 (50.4) 40 (34.2) 0.2 (0.03–0.37) 44.1 75.9 65 59.8
Globules 51 (43.6) 61 (52.1) 0.12 (–0.06–0.29) 58.8 53 49.2 55.6
Globules: regular 11 (9.4) 24 (20.5) 0.11 (–0.08–0.31) 36.4 81.1 16.7 76.9
Globules: irregular 40 (34.2) 37 (31.6) 0.17 (–0.02–0.35) 42.5 74 45.9 63.2
Globules: rim 11 (9.4) 0 (0) N/A 0 100 N/A 90.6
Dots 62 (53) 48 (41) 0.05 (–0.12–0.23) 43.5 61.8 56.3 52.1
Dots: regular 16 (13.7) 6 (5.1) 0.02 (–0.15–0.19) 6.3 95 16.7 82.9
Dots: irregular 46 (39.3) 42 (35.9) 0.2 (0.02–0.38) 47.8 71.8 52.4 62.4
Blotch 15 (12.8) 43 (36.8) 0.06 (–0.09–0.21) 46.7 64.7 16.3 67.5
Blotch: regular 8 (6.8) 7 (6) −0.07 (–0.1 to –0.03) 0 93.6 0 87.2
Blotch: irregular 7 (6) 36 (30.8) −0.06 (–0.16–0.04) 14.3 68.2 2.8 65
Branched streaks 9 (7.7) 4 (3.4) 0.11 (–0.16–0.38) 11.1 97.2 25 90.6
Peppering 29 (24.8) 16 (13.7) 0.16 (–0.03–0.36) 24.1 89.8 43.8 73.5
Scar‐like depigmentation 1 (0.9) 1 (0.9) −0.01 (–0.02–0) 0 99.1 0 98.3
Shiny‐white structures 3 (2.6) 7 (6) −0.04 (–0.07 to –0.01) 0 93.9 0 91.5
Angulated lines 1 (0.9) 2 (1.7) −0.01 (–0.03–0) 0 98.3 0 97.4
Negative network 7 (6) 5 (4.3) 0.12 (–0.17–0.42) 14.3 96.4 20 91.5
Blue‐white veil 0 (0) 3 (2.6) N/A N/A 97.4 0 97.4
Milky‐red areas 0 (0) 0 (0) N/A N/A 100 N/A 100
Tan structureless areas 3 (2.6) 2 (1.7) −0.02 (–0.04–0) 0 98.2 0 95.7
Vessels 28 (23.9) 2 (1.7) 0.04 (–0.07–0.14) 3.6 98.9 50 76.1
Vessels: comma 4 (3.4) 0 (0) N/A 0 100 N/A 96.6
Vessels: corkscrew 0 (0) 0 (0) N/A N/A 100 N/A 100
Vessels: dotted 20 (17.1) 2 (1.7) −0.03 (–0.07–0.1) 0 97.9 0 81.2
Vessels: linear irregular 6 (5.1) 0 (0) N/A 0 100 N/A 94.9
Vessels: polymorphous 3 (2.6) 0 (0) N/A 0 100 N/A 97.4
Vessels: milky red 1 (0.9) 0 (0) N/A 0 100 N/A 99.1

Abbr.: CI, confidence interval; N, number; N/A, not available; PPV, positive predictive value

Sensitivity, specificity, positive predictive value, and diagnostic accuracy were calculated for ChatGPT's annotations compared with human annotations considered the ground truth.

TABLE 2.

Concordance between human and ChatGPT‐4 annotations of dermoscopic patterns.

graphic file with name DDG-24-e1087-e001.jpg Reticular Globular Other Structureless CHPN Reticulo‐globular Two components Multi component Starburst CHPG CNPG Negative network Total
Reticular, n (%) 8 (18.6) 7 (16.3) 1 (2.3) 3 (7) 4 (9.3) 0 (0) 2 (4.7) 8 (18.6) 2 (4.7) 3 (7) 5 (11.6) 0 (0) 43
Globular, n (%) 1 (7.7) 8 (61.5) 1 (7.7) 1 (7.7) 0 (0) 1 (7.7) 0 (0) 1 (7.7) 0 (0) 0 (0) 0 (0) 0 (0) 13
Other, n (%) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0
Structureless, n (%) 2 (11.1) 5 (27.8) 1 (5.6) 5 (27.8) 1 (5.6) 0 (0) 0 (0) 1 (5.6) 1 (5.6) 2 (11.1) 0 (0) 0 (0) 18
CHPN, n (%) 4 (26.7) 1 (6.7) 2 (13.3) 1 (6.7) 1 (6.7) 2 (13.3) 0 (0) 2 (13.3) 0 (0) 1 (6.7) 1 (6.7) 0 (0) 15
Reticular‐globular, n (%) 0 (0) 1 (50) 0 (0) 1 (50) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 2)
Two components, n (%) 0 (0) 0 (0) 0 (0) 0 (0) 1 (33.3) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 1 (33.3) 1 (33.3) 3
Multi‐component, n (%) 3 (17.6) 1 (5.9) 0 (0) 0 (0) 2 (11.8) 1 (5.9) 3 (17.6) 4 (23.5) 0 (0) 2 (11.8) 1 (5.9) 0 (0) 17
Starburst, n (%) 0 (0) 0 (0) 0 (0) 1 (50) 0 (0) 0 (0) 0 (0) 0 (0) 1 (50) 0 (0) 0 (0) 0 (0) 2
CHPG, n (%) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0
CNPG, n (%) 0 (0) 0 (0) 0 (0) 0 (0) 1 (50) 0 (0) 0 (0) 0 (0) 0 (0) 1 (50) 0 (0) 0 (0) 2
Negative network, n (%) 0 (0) 1 (50) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 1 (50) 2
Total, n (%) 18 (15.4) 24 (20.5) 5 (4.3) 12 (10.3) 10 (8.5) 4 (3.4) 5 (4.3) 16 (13.7) 4 (3.4) 9 (7.7) 8 (6.8) 2 (1.7) 117

Abbr.: CHPG, central homogeneous with peripheral globules; CHPN, central homogeneous with peripheral network; CNPG, central network with peripheral globules; GPT, generative pre‐trained transformer; n, number

All percentages were calculated based on the row totals shown in the last column.

There was a 57 % concordance in the annotation of lesions as organized versus disorganized, with slight interobserver agreement (Cohen's Kappa = 0.13; 95 % CI –0.3 to 0.3) between the human readers and ChatGPT. The interobserver agreement for annotating specific dermoscopic structures varied, ranging from no agreement to fair agreement (Cohen's Kappa ranging from ‐0.04 to 0.32). The presence of a pigment network had the highest interobserver agreement. However, ChatGPT‐4 showed little agreement with human dermoscopic experts on distinguishing between typical and atypical pigment networks. Features with the lowest interobserver agreement included shiny‐white structures, tan structureless areas, and angulated lines (Table 1).

The performances of the various models, including ChatGPT‐4 (with and without annotations), the YPSONO model, and the ResNet‐based convolutional neural network (CNN), were compared using confusion matrices generated for each approach and presented in Table 3. A total of 117 dermoscopic images were analyzed by all models. The confusion matrix with a sum of 112 indicates that five images encountered errors during processing and were excluded from the analysis.

TABLE 3.

Comparison of confusion matrices of different models for melanoma identification on the study test set.

GPT4 + annotations ChatGPT‐4 without annotations YPSONO ResNet Model
All Bx Follow
TP 17 23 23 23 18 19
FP 33 10 9 2 17 17
FN 8 2 2 2 7 6
TN 59 82 24 57 70 70
Total 117 117 58 84 112 112
Accuracy 64.9% 89.7% 81.0% 95.0% 78.6% 78%
Sensitivity 68% 92% 92% 92% 72% 76%
Specificity 64% 89% 73% 97% 80% 80%
PPV 34% 70% 72% 92% 51% 53%
Kappa 0.24 0.73 0.63 0.89 0.46 0.49
F1 0.45 0.79 0.81 0.92 0.6 0.62

Abbr.: Bx, histopathologic examination (biopsy); FN, false negative; FP, false positive; F1, F1 score; GPT, generative pre‐trained transformer; PPV, positive predictive value; TN, true negative; TP, true positive

The confidence interval of the accuracy of the model was calculated: p±Z1a·p^(1p^)n=0.897±1.96·0.897(10.897)11789.7%±5.5%. Thus, the accuracy of the model is within the range [84.2 %, 95.2 %] with 95 % confidence.

The ResNet‐based model, trained on the ISIC image database, achieved a sensitivity of 76 % with a Kappa score of 0.49. The YPSONO model achieved a sensitivity of 72.0 % with a Kappa score of 0.4592. These models processed 112 images, with five images being excluded due to errors.

ChatGPT‐4's performance varied significantly depending on whether it was tasked with generating annotations. When generating descriptions and annotations prior to diagnosis, the model achieved a sensitivity of 68.0 % with a Kappa score of 0.24, successfully analyzing 117 images. In contrast, when ChatGPT‐4 was used to directly provide a diagnosis without generating annotations, it achieved the highest sensitivity of 92.0 % with a Kappa score of 0.73, successfully analyzing all 117 images without any errors. Two sub‐analyses were performed to further evaluate ChatGPT‐4's performance. In the first, we included only melanomas and benign nevi that were biopsied (excluding lesions that were followed for 12 months). In the second, we included only melanomas and benign nevi that were followed for 12 months (excluding those that were biopsied). A complete table of the sub‐analysis and the performances of the four models is presented in Table 3.

DISCUSSION

The results of this study provide insights into the capabilities and limitations of ChatGPT‐4 in analyzing dermoscopic images of melanocytic lesions. ChatGPT‐4 performs best in diagnosing melanoma when directly providing a diagnosis without generating annotations, with significantly higher sensitivity and Kappa scores compared to all other models. As noted in the introduction, LLMs are not specifically optimized for dermoscopic annotations; introducing an intermediate textual step may therefore add noise and bias, ultimately reducing diagnostic accuracy. Moreover, the clinician‐defined dermoscopic categories commonly used in practice may not align with the way LLMs process information and could, in fact, impair their diagnostic performance when users are prompted to rely on them. Recently, domain‐specific medical large language models (e.g., Med‐PaLM, MedLM) have emerged, underscoring the need to evaluate whether specialized training improves performance on dermatology tasks compared to general‐purpose models.

ChatGPT‐4 demonstrated an overall diagnostic accuracy of 65–89.7 % which aligns with findings from previous studies on ChatGPT's diagnostic capabilities. 8 The ResNet‐based model developed for this study and the YPSONO model showed similar performance. However, the YPSONO model demonstrated a slightly higher sensitivity but a lower Kappa score than the ResNet model. These findings suggest that while ChatGPT‐4 achieves high diagnostic accuracy under certain conditions, its performance varies depending on the method of interaction with the images. In addition, several total‐body photography platforms (e.g., FotoFinder ATBM, Canfield VECTRA WB360) now incorporate deep‐learning algorithms capable of analyzing dermoscopic images within the mole‐mapping workflow. The reported diagnostic performance of these systems in melanoma detection is similar to the results obtained by ChatGPT in this study, with sensitivities ranging from 90 % to 95 % and specificities between 64 % and 77 %. 16 , 17

However, the accuracy of ChatGPT‐4 when diagnosis is based on annotations is significantly lower than that of other CNN AI algorithms, which have shown much higher diagnostic accuracy for melanoma across a broader range of lesion types, often including eight or more different categories. 2 , 3 , 4 It is important to note that the images used in this study were of atypical and concerning lesions selected for digital dermoscopic monitoring in a skin cancer screening center. This selection may have affected the results, and ChatGPT‐4's performance might have been better if the images included banal benign melanocytic nevi rather than suspicious and dysplastic nevi. Supporting this interpretation, ChatGPT‐4 achieved higher performance diagnosing lesions that were only monitored (i.e., not biopsied), with specificity increasing from 89 % to 97 %, Cohen's κ rising to 0.89, and the F1 score improving to 0.92. Notably, there were no melanomas among the followed lesions or among the nevi that were biopsied, and those included in the biopsy group were inherently more suspicious. These lesions likely exhibited more concerning dermoscopic features, which contributed to a higher rate of false‐positive melanoma diagnoses. This pattern is also reflected in the precision metrics, with a positive predictive value of 92 % in the follow‐up group compared to 72 % in the biopsy group. Together, these findings suggest that ChatGPT‐4 performs particularly well in confirming benignity in lower‐suspicion contexts, whereas in higher‐suspicion scenarios it may tend to overcall melanoma, prioritizing sensitivity at the expense of specificity.

The study revealed only slight interobserver agreement between ChatGPT‐4 and human dermatologists regarding overall dermoscopic patterns. For specific dermoscopic features, ChatGPT‐4's performance varied, with interobserver agreement ranging from none to fair. It has been previously shown that while expert dermoscopists may commonly agree on the most likely diagnosis of a lesion depicted in a dermoscopic image, their agreement rate on the presence and localization of specific dermoscopic features is much lower. 15 Interestingly, similar to the findings of agreement between human expert dermoscopists, the agreement rate in the current study between human experts and ChatGPT was higher for the presence of a pigment network and globules and lower for tan structureless areas and angulated lines. Additionally, there was little agreement between ChatGPT and the experts on whether a pigment network is typical or atypical, mirroring the disagreement often seen among human experts. 15 , 18 , 19

There are two main differences between the CNNs used in previous studies and ChatGPT‐4. First, CNNs are trained explicitly for dermoscopic image recognition with the sole purpose of providing diagnosis or risk assessment. In contrast, it is unclear whether ChatGPT‐4 was trained on a dermoscopic image database labeled with diagnoses, as it serves multiple purposes across various fields. Second, unlike CNNs, which typically do not provide explanations for their predictions and do not rely on human‐defined criteria, ChatGPT generates textual descriptions of recognized dermoscopic features and explains its diagnostic reasoning. This may enhance physicians’ confidence in the model's judgment. 13 Given its accessibility and ability to produce detailed dermoscopic descriptions, ChatGPT may therefore have particular potential for clinical use – especially as AI‐generated explanations have been shown to increase trust in algorithms for melanoma diagnosis. 20

When initiating this study, we hypothesized that GPT, as a generative AI model, would have the advantage of providing accurate dermoscopic descriptions of melanocytic lesions, thereby supporting teaching and training. The results of the current study indicate that while the GPT algorithm has a limited ability to describe dermoscopic images at the level of expert dermoscopists, it also shows great potential. Training a GPT on exemplary dermoscopic images may improve its accuracy in recognizing dermoscopic structures and patterns. Such a trained GPT could serve as a valuable tool for teaching students and residents and for automating dermoscopic analyses, thereby reducing the need for multiple readers to annotate large numbers of images. However, caution is warranted when using this general tool, especially due to its accessibility and comfort of use. AI‐generated judgments have been shown to influence physicians’ clinical decision‐making, potentially leading over time to overreliance, 13 which may raise concerns within the framework of medical practice standards.

Limitations

All images were sourced from a single practice and represented equivocal lesions selected for biopsy or dermoscopic digital monitoring by the physician, making them outlier lesions rather than representative benign nevi commonly seen in daily practice. Only images and diagnoses were available for analysis, with no demographic data about participants.

We lack detailed knowledge about the training set used for ChatGPT, preventing us from drawing precise conclusions about its capabilities and limitations. Moreover, the relatively small sample size constrains generalizability. Future research with larger datasets could reduce the accuracy range and provide a more balanced representation of dermoscopic patterns. Finally, our images did not include lesions located on the face, palms, soles, or mucosal surfaces. Thus, the exclusions may have induced improved results. Another consideration is the emergence of domain‐specialized medical LLMs, which may address some of the limitations inherent in using general‐purpose models such as ChatGPT‐4. While our study, conducted between April and May 2024, evaluated a general‐purpose multimodal version of ChatGPT‐4, more recent developments have introduced health‐focused large language models and dedicated benchmarking frameworks, including Google's Med‐PaLM and MedLM family, recent MedGemma releases, and OpenAI's HealthBench for clinically aligned evaluation. These specialized models, often fine‐tuned on curated medical datasets, may demonstrate improved domain‐specific reasoning and diagnostic performance. Future research should therefore include direct comparisons between general‐purpose and domain‐specialized models, ideally with multimodal vision‐language capabilities, using dermoscopy‐specific datasets to better assess their potential clinical utility.

CONCLUSIONS

ChatGPT‐4 demonstrated strong diagnostic accuracy in distinguishing melanomas from nevi when directly analyzing dermoscopic images without prior annotations, surpassing conventional CNN‐based models. However, its ability to accurately describe dermoscopic features and agree with expert dermatologists remains limited. Future research and targeted training using detailed, high‐quality dermoscopic datasets may improve its descriptive accuracy, expanding its utility in clinical decision support, dermatological education, and research. These findings highlight both the promise and the caution required in utilizing ChatGPT‐4 as an adjunct diagnostic tool, emphasizing the necessity of critical clinical judgment.

CONFLICT OF INTEREST STATEMENT

None.

Supporting information

Supplementary information

DDG-24-e1087-s001.docx (13.8KB, docx)

REFERENCES

  • 1. Esteva A, Kuprel B, Novoa RA, et al. Dermatologist‐level classification of skin cancer with deep neural networks. Nature. 2017;542(7639):115‐118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Tschandl P, Codella N, Akay BN, et al. Comparison of the accuracy of human readers versus machine‐learning algorithms for pigmented skin lesion classification: an open, web‐based, international, diagnostic study. Lancet Oncol. 2019;20(7):938‐947. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Dick V, Sinz C, Mittlbock M, et al. Accuracy of Computer‐Aided Diagnosis of Melanoma: A Meta‐analysis. JAMA Dermatol. 2019;155(11):1291‐1299. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Salinas MP, Sepulveda J, Hidalgo L, et al. A systematic review and meta‐analysis of artificial intelligence versus clinicians for skin cancer diagnosis. NPJ Digit Med. 2024;7(1):125. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Murdoch WJ, Singh C, Kumbier K, et al. Definitions, methods, and applications in interpretable machine learning. Proc Natl Acad Sci U S A. 2019;116(44):22071‐22080. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Nori H KN, McKinney SM, Carignan D, Horvitz E. Capabilities of GPT‐4 on Medical Challenge Problems 2023: Available from: http://arxiv.org/abs/2303.13375. March 2023. [Last accessed October 9, 2025].
  • 7. Openai . ChatGPT Can Now See, Hear, and Speak 2023. 25‐09‐2023. Available from: https://openai.com/index/chatgpt‐can‐now‐see‐hear‐and‐speak. [Last accessed October 9, 2025].
  • 8. Perlmutter JW, Milkovich J, Fremont S, et al. Beyond the Surface: Assessing GPT‐4's Accuracy in Detecting Melanoma and Suspicious Skin Lesions From Dermoscopic Images. Plast Surg (Oakv) Published online February 18, 2025, 10.1177/22925503251315489. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Shifai N, van Doorn R, Malvehy J, et al. Can ChatGPT vision diagnose melanoma? An exploratory diagnostic accuracy study. J Am Acad Dermatol. 2024;90(5):1057‐1059. [DOI] [PubMed] [Google Scholar]
  • 10. Samuel Dahan RB, Liang D, Zhu X. A call for an Open‐source Legal Language Model 2023: Available from: https://ssrn.com/abstract=4587092 or 10.2139/ssrn.4587092. 28/06/2024. [Last accessed October 9, 2025]. [DOI]
  • 11. Chunpeng Zhai SWLDL. The effects of over‐reliance on AI dialogue systems on students' cognitive abilities: a systematic review. Smart Learning Environments. 2024;11(1):28. [Google Scholar]
  • 12. Rezazade Mehrizi MH, Mol F, Peter M, et al. The impact of AI suggestions on radiologists' decisions: a pilot study of explainability and attitudinal priming interventions in mammography examination. Sci Rep. 2023;13(1):9230. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Prinster D, Mahmood A, Saria S, et al. Care to Explain? AI Explanation Types Differentially Impact Chest Radiograph Diagnostic Performance and Physician Trust in AI. Radiology. 2024;313(2):e233261. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Gu H, Yang C, Magaki S, et al. Majority voting of doctors improves appropriateness of AI reliance in pathology. Int J Hum‐Comput Stud 2024,190:103315. [Google Scholar]
  • 15. Liopyris K, Navarrete‐Dechent C, Marchetti MA, et al. Expert Agreement on the Presence and Spatial Localization of Melanocytic Features in Dermoscopy. J Invest Dermatol. 2024;144(3):531‐539. e513. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Haenssle HA, Fink C, Toberer F, et al. Man against machine reloaded: performance of a market‐approved convolutional neural network in classifying a broad spectrum of skin lesions in comparison with 96 dermatologists working under less artificial conditions. Ann Oncol. 2020;31(1):137‐143. [DOI] [PubMed] [Google Scholar]
  • 17. Cerminara SE, Cheng P, Kostner L, et al. Diagnostic performance of augmented intelligence with 2D and 3D total body photography and convolutional neural networks in a high‐risk population for melanoma under real‐world conditions: A new era of skin cancer screening? Eur J Cancer. 2023;190:112954. [DOI] [PubMed] [Google Scholar]
  • 18. Posch C. Melanocytic lesions: How to navigate variations in human and artificial intelligence. J Eur Acad Dermatol Venereol. 2024;38(5):792‐793. [DOI] [PubMed] [Google Scholar]
  • 19. Goessinger EV, Cerminara SE, Mueller AM, et al. Consistency of convolutional neural networks in dermoscopic melanoma recognition: A prospective real‐world study about the pitfalls of augmented intelligence. J Eur Acad Dermatol Venereol. 2024;38(5):945‐953. [DOI] [PubMed] [Google Scholar]
  • 20. Chanda T, Hauser K, Hobelsberger S, et al. Dermatologist‐like explainable AI enhances trust and confidence in diagnosing melanoma. Nat Commun. 2024;15(1):524. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary information

DDG-24-e1087-s001.docx (13.8KB, docx)

Articles from Journal Der Deutschen Dermatologischen Gesellschaft are provided here courtesy of Wiley

RESOURCES