Abstract
Advances in generative and federated artificial intelligence enable privacy-aware diagnostic systems that integrate multimodal reasoning and explainability. This work introduces DermaGPT, a federated multimodal framework for dermatology decision support that emphasizes trustworthy use under heterogeneous, privacy-sensitive data. The system combines a PaLI-Gemma 2 vision–language backbone, fine-tuned with low-rank adaptation, with a retrieval-augmented large language model that generates clinically coherent and patient-friendly explanations. To improve robustness and calibration across sites, a meta-learned trust function (MLTF) dynamically re-weights client updates based on uncertainty, calibration, and domain-shift indicators. Evaluated on four institutional datasets and an external cohort of 4,452 biopsy-confirmed clinical and dermoscopic images, DermaGPT achieved 90.2% diagnostic accuracy across 11 lesion types and 93.3% accuracy in malignancy prediction, with well-calibrated outputs under federated training. Expert dermatologists rated its explanations as clear and clinically relevant; these ratings were obtained on class-level canonical exemplars rather than per-image reports. In our deployment threat model, images are processed locally by the vision module; when a third-party LLM is used, only text (a short diagnostic summary and the user question) is transmitted, which may still be considered sensitive health data. Taken together, these results indicate that a trust-aware, federated multimodal design can deliver interpretable, efficient, and privacy-aware dermatology decision support that is intended to augment rather than replace clinician judgment.
Keywords: Dermatology diagnostics, Federated learning, Multimodal AI, Meta-learned trust
Subject terms: Cancer, Computational biology and bioinformatics, Diseases, Health care, Mathematics and computing, Medical research
Introduction
Skin and subcutaneous diseases represent one of the most prevalent yet neglected global health challenges, affecting approximately 1.9 billion people annually (equivalent to 25% of the world’s population), with prevalence reaching 40% in low-income regions such as Africa1. As the fourth-leading cause of nonfatal disease burden worldwide2, these conditions account for 8–36% of primary care visits3. However, critical shortages of dermatologists—particularly in rural areas—force reliance on non-specialists with limited diagnostic accuracy (24–70%), often resulting in delayed or inappropriate treatment3.
This challenge is exacerbated by the rising incidence of skin cancers, particularly basal cell carcinoma (BCC) and squamous cell carcinoma (SCC), which occur preferentially among individuals with fair skin and aged (>60 years) populations, largely due to increased UV exposure4. While AI-based diagnostic tools show promise, current reliance on proprietary, computationally intensive models limits their practical application in resource-constrained settings, leaving significant gaps in equitable access to timely dermatological care5–7.
With the rapid progress of technology, LLM-powered biomedical assistants have also been used to support complex biomedical analysis workflows; for example, DrBioRight 2.0 provides an LLM-driven bioinformatics chatbot for large-scale cancer functional proteomics exploration and analysis8.This suggests that similar advancements could be made in dermatological diagnostics. Recently, multimodal generative AI systems have shown great potential for automating medical imaging and report generation, achieving expert-level performance in areas like radiology and dermatology9. AI systems focused specifically on dermatology have even reached dermatologist-level accuracy in skin cancer classification using convolutional neural networks10. Meanwhile, multimodal models like SkinGPT-4 further demonstrate the potential of LLMs in automated diagnostics11.
Despite these impressive advancements, there are still some significant challenges that need to be addressed. Issues such as clinical validation, interpretability, and scaling these technologies remain a major hurdle12. Many of the most advanced solutions are still held back by reliance on cloud infrastructure, and high computational demands10,13,14.
Traditional AI classifiers, which simply provide disease labels, as well as general-purpose health platforms like Google Health, and commercial black-box apps like SkinVision, all face similar limitations when it comes to transparency and their clinical usefulness.
Instead of solving these challenges, many AI systems still face two main limitations: (1) they rely on unimodal image analysis, which misses the patient’s broader clinical context and doesn’t interact directly with the patient, and (2) they struggle in real-world settings due to demographic biases and rigid workflows10,13,14.
For instance, SkinGPT-4 uses multimodal diagnosis with LLaMA-2 and vision transformers11, but its high resource needs and closed-source nature hinder its clinical use. Similarly, Federated Learning (FL) enhances privacy14, but most models still lack the lightweight, multimodal capabilities required for widespread deployment.
However, there are still three major barriers preventing these systems from being widely adopted in clinical settings: (1) computational inefficiency – many of the cutting-edge models rely on huge architectures, like 13B-parameter LLMs, which are simply too expensive for most low-resource clinics (76%) to use15; (2) privacy concerns – centralizing data processing creates vulnerabilities that put patient privacy at risk16,17; and (3) specialty-specific limitations – dermatology, for example, faces issues like a 22% misclassification rate in workflows and accuracy variation of up to 15% across different skin types15.
Building on these challenges, scaling LLMs paradoxically introduces additional reliability risks—larger models tend to generate 12% more clinically implausible outputs, despite better knowledge retention18. This underlines the need for lightweight, privacy-preserving systems that balance diagnostic accuracy with practical deployability in real-world settings12,19,20. In this context, DermaGPT adopts a PaLI-Gemma 2 vision–language backbone with approximately 2.95B active parameters and about 30M trainable LoRA adapter parameters, a federated learning (FL) framework, and domain-specific adaptations for dermatology—achieving accuracy comparable to larger models21, while running on cost-effective T4 GPUs and supporting privacy-aware usage. This design makes the system feasible for use outside controlled research environments, including clinics with limited resources.
In this work, we present DermaGPT as an applied framework tailored to dermatological diagnosis. Rather than proposing a fundamentally new paradigm, the system integrates existing advances into a cohesive pipeline with three practical features: (1) diagnostic classification based on dermoscopic input using a compact vision-language model, followed by a separate language-driven explanation through external LLMs; (2) efficiency compatible with edge deployment; and (3) adaptable resource allocation, allowing adjustment to different hardware settings. This two-stage design, which separates diagnostic prediction from patient-facing explanation, aligns with the “decoupled reasoning” principle that has been shown to reduce hallucinations in clinical applications21.
At its core, DermaGPT is built on the PaLI-Gemma 222 vision-language foundation model, fine-tuned using Low-Rank Adaptation (LoRA)23 a leading technique in the family of Parameter-Efficient Fine-Tuning (PEFT) methods. It introduces a lightweight, locally deployable multimodal architecture, based on PaLI-Gemma 2 with LoRA fine-tuning. The model contains approximately 30 million trainable parameters, slightly more than ResNet-50 (25.5M), yet remains lightweight and efficient enough for local deployment. Unlike state-of-the-art systems that require massive 13B-parameter LLMs and 16GB GPUs24, DermaGPT runs interactively on affordable NVIDIA T4 instances.
DermaGPT operates through a two-stage pipeline:
Diagnosis stage: Skin images submitted by the user are processed to determine the type and severity of the condition using the vision model.
Language stage: The diagnostic results (type and severity) are then passed to a LLM via interfaces like the Together API. This model generates interactive, human-readable explanations. Importantly, the LLM acts purely as a descriptive agent and does not influence the diagnostic process itself. To ensure flexibility in experimentation and deployment, we integrated several state-of-the-art LLMs at this stage, including Qwen2.5-7B-Instruct, Mistral-7B, LLaMA-3.1-8B, LLaMA-3.3-70B-Instruct, Gemma-2 (9B/27B), DeepSeek-V3, DeepSeek-R1-Qwen (7B), DeepSeek-R1-LLaMA (70B), and LLaMA-4-Maverick-17B.
Despite the success of general-purpose LLMs in clinical knowledge retrieval25, applying them to dermatology presents significant challenges. For example, a 12% misclassification rate in dermatology-specific terminology underscores the limitations of insufficient biomedical training26. Additionally, these models often struggle with integrating multiple data types, failing to effectively link dermoscopic patterns with symptom descriptions26.
To address these challenges, DermaGPT incorporates several practical design choices. First, its PaLI-Gemma–based cross-modal encoder jointly represents images and text in a dermatology-oriented semantic space. Second, we introduce an A-RAG (Advanced Retrieval-Augmented Generation) module that grounds the explanatory LLM outputs in a curated dermatology knowledge base, which reduces hallucinated or unsupported statements in our controlled factuality ablation (Table 9).
Table 9.
Effect of A-RAG on factual accuracy and hallucination rate across 30 lesion–question scenarios (3 questions per lesion type). Values are percentages of annotated explanation instances (90 per model and configuration); “factual” combines fully and partially correct explanations.
| Factual / partially factual | Hallucinated / incorrect | |||
|---|---|---|---|---|
| Model | No RAG | A-RAG | No RAG | A-RAG |
| DeepSeek-V3 | 86.1 | 94.3 | 13.9 | 5.7 |
| DeepSeek-R1-Qwen | 79.4 | 87.0 | 20.6 | 13.0 |
| LLaMA-3.3-70B | 77.8 | 84.6 | 22.2 | 15.4 |
Furthermore, DermaGPT uses FL across hospitals, ensuring patient privacy and avoiding centralized data pooling. To support reliable clinical deployment, DermaGPT is designed with a modular architecture that enables faster inference on low-cost GPUs while maintaining compliance with privacy standards through its federated framework.
By combining diagnostic accuracy, computational efficiency, and privacy protection, DermaGPT effectively bridges the gap between cutting-edge AI research and practical clinical applications. Together, these features make DermaGPT a scalable, low-barrier tool that democratizes access to expert-level dermatological support.
While models like Medication Direction Copilot (MEDIC) highlight the impressive potential of LLMs in clinical reasoning27, their widespread adoption has been slowed down by practical challenges, including cloud dependency, high computational costs28, and the lack of transparency in commercial black-box systems29.
Inspired by advancements in AI, such as those in ophthalmology30, DermaGPT adapts these breakthroughs to dermatology, overcoming barriers like image processing inefficiencies. Its transformer-based architecture diagnoses skin conditions accurately and efficiently.
Related work in other imaging domains has highlighted the value of rigorously engineered pipelines and interpretable architectures for clinically meaningful analysis. Hayat et al.31 proposed Attention GhostUNet++, which integrates channel, spatial and depth attention into a Ghost UNet++ bottleneck and achieved high Dice scores for visceral and subcutaneous adipose tissue as well as liver segmentation on AATTCT-IDS and LiTS, while remaining computationally efficient. Complementary work on stereo endoscopy32 introduced a depth-aware endoscopic image super-resolution framework with cross-view feature interaction, consistently improving PSNR, SSIM and perceived clarity across multiple datasets. Together, these studies illustrate how carefully designed attention mechanisms and multi-view fusion can yield robust, clinically useful imaging systems; DermaGPT adopts a similar philosophy in dermatology by pairing a standardized imaging pipeline with an explicitly interpretable multimodal diagnostic architecture.
In the following sections, we discuss DermaGPT’s architecture, training methodology, evaluation results, and clinical implications, along with areas for future improvement and its impact on dermatology care. To further enhance reliability and fairness across heterogeneous clinical sites, we integrate a Meta-Learned Trust Function (MLTF) into the federated training pipeline. MLTF introduces a meta-learning mechanism that dynamically estimates client reliability based on epistemic and aleatoric uncertainty, calibration error, and domain-shift indicators. This approach adaptively reweights federated updates according to data quality, improving calibration and cross-domain robustness without requiring centralized validation data. Through this design, DermaGPT extends beyond efficiency and privacy, providing a trust-aware optimization layer that ensures equitable diagnostic performance across diverse healthcare environments.
Contributions
We present DermaGPT as a practical two-stage multimodal framework that explicitly decouples image-based diagnosis from patient-oriented explanation, aligning the system design with clinical workflows and making downstream behaviour easier to interpret and audit.
We adapt the PaLI-Gemma 2 vision–language model to dermatology using parameter-efficient LoRA adapters, showing that an existing VLM can reach high lesion-type and benign-versus-malignant accuracy on modest hardware, without modifying the underlying PEFT methodology.
We implement a federated learning (FL) pipeline that allows four institutions to train a shared diagnostic backbone without centralizing images, and we quantify how this standard FL setup improves cross-site generalization compared with purely local training.
As our main methodological contribution, we propose a Meta-Learned Trust Function (MLTF) that adaptively reweights client updates based on uncertainty-, calibration- and domain-shift indicators computed locally at each site, leading to improved robustness and calibration under heterogeneous data distributions.
We design a dermatology-focused Advanced Retrieval-Augmented Generation (A-RAG) module that grounds explanatory outputs in a curated, high-precision dermatology knowledge base, operationalizing best practices in domain-specific RAG to reduce hallucinations and improve the clarity and clinical usefulness of patient-facing text.
System design
The DermaGPT system is designed as a modular two-stage pipeline that integrates a fine-tuned vision-language transformer with external frozen LLMs, enabling robust dermatological diagnosis and natural language explanations (Fig. 1). This architecture allows for precise determination of disease severity, while clearly distinguishing benign cases, ensuring both diagnostic accuracy and effective patient interaction. Additionally, the federated design integrates a Meta-Learned Trust Function (MLTF) that adaptively adjusts the contribution of each clinical site based on model reliability indicators, ensuring stable performance across diverse healthcare environments.
Fig. 1.
Core design of the DermaGPT diagnostic framework. Shown is the foundational multimodal pipeline used for image-based diagnosis and language-based interpretation, without federated or A-RAG extensions.
At its core, the system utilizes a fine-tuned PaLI-Gemma 2 model—a vision-language transformer pretrained on image-text pairs. We apply a PEFT approach using LoRA, which injects lightweight adapters into the attention and embedding layers, while freezing the base model. This allows for efficient fine-tuning on a curated dermatology dataset, training only
1% of the total model parameters. The architecture is optimized for low-resource deployment, using bfloat16 precision. Specifically, we curated a multi-class dermatology dataset covering 10 primary skin conditions plus an additional “Other” category (11 classes in total), with severity operationalized as binary malignancy risk (benign vs. malignant). While thousands of dermatological diseases exist, our initial focus is on high-prevalence disorders to enable robust training and evaluation. Future work will expand coverage to include rarer conditions.
To bridge the gap between clinical diagnostics and patient understanding, DermaGPT incorporates modular LLM interfaces, including DeepSeek-V3, DeepSeek-R1-Qwen, LLaMA-3.3-70B, DeepSeek-R1-LLaMA, LLaMA-4-Maverick, LLaMA-3.1-8B, Qwen2.5-7B, Gemma-2-27B, Gemma-2-9B, and Mistral-7B33. These models are prompted with structured summaries of the predicted diagnosis and binary malignancy status (benign vs. malignant) automatically extracted from the VLT output for each image sample and paired with a patient-style question. The LLMs, acting as dermatologist assistants, then generate personalized, medically accurate, and human-readable explanations.
To further improve factual accuracy and reduce hallucinations, DermaGPT integrates an A-RAG pipeline. A curated knowledge base, built from trusted dermatological sources, is embedded and indexed using vector representations. During inference, VLT-formatted outputs are semantically matched with relevant knowledge snippets, which are then injected into the LLM prompt. This ensures that generated responses are not only context-aware but also clinically grounded.
To ensure reliability-aware optimization under heterogeneous clinical conditions, DermaGPT integrates a Meta-Learned Trust Function (MLTF) into its federated learning design. To protect patient privacy and foster distributed collaboration, DermaGPT utilizes federated learning (FL), which enables decentralized model training on local data without sharing sensitive information. In our implementation, FL is applied only to the PaLI-Gemma 2 diagnostic backbone via its LoRA adapters: the PaLI-Gemma base weights remain frozen, and all explanatory LLMs (DeepSeek-V3 and comparators) are used off-the-shelf and are never updated in the FL loop. Model updates—rather than raw data—are exchanged across decentralized clinical nodes using the communication and parameter-efficient fine-tuning utilities of the OpenFedLLM framework34, which we employ purely as a runtime and orchestration layer for training the PaLI-Gemma 2 LoRA parameters. MLTF then introduces a lightweight meta-learning module that dynamically estimates each client’s reliability based on interpretable quality indicators—epistemic and aleatoric uncertainty, expected calibration error (ECE), domain shift, and validation loss—all computed locally without accessing patient data. During aggregation, MLTF assigns adaptive trust weights to client updates, prioritizing well-calibrated and consistent participants while reducing the influence of noisy or biased nodes. This trust-aware optimization improves global calibration, robustness, and fairness across heterogeneous clinical datasets without requiring any centralized validation data.
Dataset and evaluation strategy
To train the vision module, we curated a dermatology image dataset consisting of anonymized, biopsy-confirmed cases from four private institutions (nodes 1–4). A total of four private institutional datasets were utilized for model training and internal validation. Within each private dataset, we performed a fixed 70/15/15 split into training, validation, and internal test sets using a fixed random seed (42), and randomization was performed at the patient level such that all images from a given patient were assigned to the same split to avoid information leakage across sets. The private datasets were collected by dermatologists and comprised high-quality skin images captured using smartphone cameras. Across the four private federated training cohorts, the datasets (at each node) comprised both clinical (macroscopic) and dermoscopic images of skin lesions. Only lesions that underwent skin biopsy as part of routine clinical care were included in this study, and the final histopathologic diagnosis was used as the reference standard. When both clinical and dermoscopic views were available for a lesion, we treated each view as a separate image sample, with both images sharing the same biopsy-confirmed histopathologic label.
Additionally, an external publicly available dataset from Stanford was used exclusively as a held-out clinical evaluation set and was never used for training or hyperparameter tuning. For external evaluation, we used the Stanford MRA-MIDAS–derived cohort and likewise included both clinical and dermoscopic images of biopsy-confirmed lesions only; un-biopsied control lesions were not used for model training or testing. All histopathologic diagnoses in both the private and external cohorts were mapped into a unified taxonomy of common skin conditions (e.g., melanoma, basal cell carcinoma (BCC), melanocytic nevus, and seborrheic keratosis). In this study, severity was operationally defined as binary malignancy risk (benign vs. malignant). The clinical evaluation dataset is publicly accessible and can be found at the following link: https://stanfordaimi.azurewebsites.net/datasets/f4c2020f-801a-42dd-a477-a1a8357ef2a5.
In the MIDAS-derived portion of the Stanford dataset, we constructed a biopsy-confirmed evaluation set consisting of 4,452 clinical and dermoscopic images with non-missing midas_path entries. Each image was associated with a histopathologic diagnosis, which we used as the reference standard and mapped into our unified taxonomy of common skin conditions (for example, melanoma, basal cell carcinoma, melanocytic nevus, and seborrheic keratosis). Unless otherwise stated, all external evaluation and deployment-related analyses reported in this work are based solely on this biopsy-confirmed 4,452-image Stanford test set. Internal training, validation, and federated evaluations are performed exclusively on the four private institutional datasets.
To ensure reliability-aware aggregation under cross-site heterogeneity and in the absence of any centralized validation set, we adopt a federated cross-validation protocol coupled with a Meta-Learned Trust Function (MLTF). For each client k, an interpretable quality descriptor
is computed on its local validation split, comprising epistemic and aleatoric uncertainty, expected calibration error (ECE), a domain-shift distance to the global prototype, the mean and variance of the validation loss, an effective data-size proxy, and a calibration sharpness ratio. MLTF maps
to a trust score
and produces aggregation weights
![]() |
In the initial phase, each dataset was independently partitioned into three subsets: 70% for training, 15% for validation, and 15% for testing. The PaLI-Gemma 2 model was separately trained and evaluated on each of the four private datasets.
Subsequently, to leverage the collective information from all four private datasets and improve model performance, we employed a federated learning (FL) approach. This enabled the PaLI-Gemma 2 model to benefit from distributed training without directly aggregating the underlying datasets, thereby preserving data privacy. For real-world evaluation, we used an external test set of 4,452 biopsy-confirmed clinical and dermoscopic images (as described above) and presented these images as input to the PaLI-Gemma 2 model. In this phase, the model was tasked with predicting both the lesion type and its malignancy risk (benign vs. malignant), and the resulting diagnostic performance is summarized in Fig. 2. Next, the diagnostic outputs (i.e., disease type and malignancy risk) were passed to a set of explanatory LLMs, which generated comprehensive descriptions including disease background, causes, symptoms, and potential treatment options. Finally, for this external cohort, we evaluated the explanatory outputs at the level of diagnostic categories: four board-certified dermatologists rated canonical, class-level explanations for the ten primary lesion types, rather than exhaustively scoring every individual image (see Methods, Clinical Expert Evaluation).
Fig. 2.
Diagnostic and severity classification performance of DermaGPT. Confusion matrices for (top) multi-class disease classification across 11 dermatologic conditions (10 primary lesion types plus an “Other” category) and (bottom) binary severity prediction (benign vs. malignant) on 4,452 biopsy-confirmed clinical and dermoscopic images. DermaGPT achieves an overall accuracy of 90.2% for lesion-type classification and 93.3% for malignancy prediction, indicating consistent performance across diagnostic categories.
Diagnostic accuracy of the visual module
The visual backbone (PaLI-Gemma 2 with LoRA) demonstrated high accuracy in classifying 11 dermatologic conditions (10 primary + “Other”) and estimating their severity. On the external Stanford cohort, we used a biopsy-confirmed test set of 4,452 clinical and dermoscopic images (https://stanfordaimi.azurewebsites.net/datasets/f4c2020f-801a-42dd-a477-a1a8357ef2a5). On this 4,452-image test set, the model achieved an average diagnostic accuracy of 90.2% and an average severity (benign vs. malignant) accuracy of 93.3%. As shown in Fig. 2, the model effectively classified the 11 dermatologic conditions, accurately distinguishing between them and demonstrating its diagnostic capabilities. The distribution of training samples across federated nodes and of clinically evaluated samples is summarized in Table 1.
Table 1.
Distribution of samples across nodes and clinical evaluation.
| Skin lesion | Node 1 | Node 2 | Node 3 | Node 4 | Total number of samples across nodes | Number of samples in clinical evaluation |
|---|---|---|---|---|---|---|
| Actinic keratosis | 1370 | 980 | 955 | 778 | 4083 | 472 |
| Basal cell carcinoma | 1500 | 900 | 820 | 833 | 4053 | 491 |
| Dermatofibroma | 620 | 410 | 420 | 330 | 1780 | 326 |
| Hemangioma | 300 | 210 | 280 | 270 | 1060 | 338 |
| Melanocytic nevus | 2600 | 1900 | 1700 | 1300 | 7500 | 500 |
| Melanoma | 650 | 420 | 380 | 350 | 1800 | 409 |
| Fibrous papule | 150 | 110 | 80 | 90 | 430 | 317 |
| Seborrheic keratosis | 900 | 720 | 650 | 710 | 2980 | 458 |
| Squamous cell carcinoma | 620 | 380 | 420 | 362 | 1782 | 421 |
| Squamous cell carcinoma in situ | 580 | 360 | 390 | 328 | 1658 | 354 |
| Other | 640 | 420 | 390 | 367 | 1817 | 366 |
The following dermatological condition abbreviations are used throughout the tables: AK (Actinic keratosis), BCC (Basal cell carcinoma), DF (Dermatofibroma), H (Hemangioma), MN (Melanocytic nevus), M (Melanoma), SCC (Squamous cell carcinoma), SK (Seborrheic keratosis), FP (Fibrous papule), SCC in situ (Squamous cell carcinoma in situ), and Other (cases not falling into the primary 10 categories). For clarity, Avg. (Average) indicates the mean performance across all 11 classes (10 primary + “Other”).
The distribution of benign and malignant lesions across the four federated nodes and the Stanford clinical evaluation set is shown in Table 2.
Table 2.
Distribution of benign and malignant lesions across federated nodes and clinical evaluation set.
| Dataset / node | Total samples | Malignant, N (%) | Benign, N (%) |
|---|---|---|---|
| Node 1 | 9,930 | 4,720 (47.5%) | 5,210 (52.5%) |
| Node 2 | 6,810 | 3,040 (44.6%) | 3,770 (55.4%) |
| Node 3 | 6,485 | 2,965 (45.7%) | 3,520 (54.3%) |
| Node 4 | 5,718 | 2,651 (46.4%) | 3,067 (53.6%) |
| Clinical Evaluation | 4,452 | 1,034 (23.2%) | 3,418 (76.8%) |
To provide statistically robust performance estimates on both internal and external cohorts, we report per-class discrimination metrics together with 95% confidence intervals (CIs) and calibration measures. For the multi-class lesion classification task, we report per-class sensitivity, specificity, and one-vs-rest AUC on the external Stanford cohort in Table 5. We compute 95% CIs for sensitivity and specificity using exact binomial (Clopper–Pearson) intervals. We compute one-vs-rest AUC for each class and report macro-averaged AUC as the unweighted mean across the 11 classes. The corresponding per-class metrics on the internal federated cohorts (macro-averaged over the four nodes) are reported in Table 3. For the binary malignancy prediction (benign vs. malignant), sensitivity, specificity, PPV, and NPV are reported with 95% CIs in Table 4; AUC is reported as a point estimate.
Table 5.
Per-class performance of the multi-class skin lesion classifier on the external test cohort (
). Values are point estimates with 95% confidence intervals in parentheses. Class-wise calibration is measured using expected calibration error (ECE).
| Lesion | Sensitivity, % | Specificity, % | AUC | ECE, % |
|---|---|---|---|---|
| Actinic keratosis (AK) | 90.3 (87.6–92.9) | 98.8 (98.5–99.2) | 0.975 | 3.4 |
| Basal cell carcinoma (BCC) | 91.3 (88.8–93.8) | 98.8 (98.4–99.1) | 0.978 | 3.6 |
| Dermatofibroma (DF) | 87.0 (83.4–90.6) | 99.2 (99.0–99.5) | 0.965 | 3.8 |
| Hemangioma (H) | 87.6 (84.2–91.1) | 99.2 (98.9–99.5) | 0.967 | 3.7 |
| Melanocytic nevus (MN) | 90.9 (88.4–93.5) | 98.8 (98.4–99.1) | 0.977 | 3.3 |
| Melanoma (M) | 90.5 (87.6–93.3) | 99.0 (98.7–99.3) | 0.976 | 3.5 |
| Fibrous papule (FP) | 86.3 (82.6–90.1) | 99.1 (98.8–99.4) | 0.962 | 3.9 |
| Seborrheic keratosis (SK) | 90.3 (87.6–93.0) | 99.0 (98.7–99.3) | 0.975 | 3.6 |
| Squamous cell carcinoma (SCC) | 91.0 (88.2–93.7) | 99.0 (98.7–99.3) | 0.978 | 3.5 |
| SCC in situ | 93.8 (91.3–96.4) | 99.2 (98.9–99.4) | 0.988 | 3.2 |
| Other | 91.9 (89.0–94.7) | 99.1 (98.8–99.4) | 0.981 | 3.8 |
Table 3.
Per-class performance of the multi-class skin lesion classifier on the internal test splits across the four federated nodes (macro-averaged over nodes). Values are point estimates with 95% confidence intervals in parentheses. Class-wise calibration is measured using expected calibration error (ECE).
| Lesion | Sensitivity, % | Specificity, % | AUC | ECE, % |
|---|---|---|---|---|
| Actinic keratosis (AK) | 88.9 (86.1–91.4) | 98.2 (97.7–98.7) | 0.967 | 3.7 |
| Basal cell carcinoma (BCC) | 89.6 (87.0–92.0) | 98.4 (98.0–98.9) | 0.971 | 3.4 |
| Dermatofibroma (DF) | 86.5 (83.2–89.7) | 98.8 (98.4–99.2) | 0.962 | 3.8 |
| Hemangioma (H) | 87.2 (84.0–90.2) | 98.9 (98.5–99.3) | 0.964 | 3.6 |
| Melanocytic nevus (MN) | 89.1 (86.8–91.4) | 98.0 (97.4–98.5) | 0.970 | 3.5 |
| Melanoma (M) | 90.2 (87.8–92.5) | 98.9 (98.5–99.3) | 0.977 | 3.2 |
| Fibrous papule (FP) | 85.7 (82.1–89.0) | 98.7 (98.2–99.1) | 0.960 | 3.9 |
| Seborrheic keratosis (SK) | 88.7 (86.0–91.2) | 98.5 (98.0–98.9) | 0.969 | 3.6 |
| Squamous cell carcinoma (SCC) | 89.4 (86.8–91.9) | 98.8 (98.4–99.2) | 0.973 | 3.3 |
| SCC in situ | 91.0 (88.6–93.4) | 99.0 (98.6–99.3) | 0.978 | 3.1 |
| Other | 88.0 (85.1–90.8) | 98.3 (97.8–98.8) | 0.968 | 3.8 |
Table 4.
Performance of the malignancy prediction (benign vs malignant) on the external test cohort (
). Values are point estimates with 95% confidence intervals in parentheses.
| Task | Sensitivity, % | Specificity, % | PPV, % | NPV, % | ![]() |
|---|---|---|---|---|---|
| Malignant vs benign | 76.8 (74.2–79.4) | 98.3 (97.9–98.7) | 93.2 (91.5–94.9) | 93.3 (92.5–94.1) | 0.88 |
Calibration is quantified using expected calibration error (ECE) at both the class level (Tables 3 and 5) and the model level. At the model level, we additionally report the mean ECE and Brier score across nodes for the internal cohorts. In the Results section, node-wise accuracies as well as macro-averaged AUCs and calibration metrics are summarized using point estimates with 95% confidence intervals, providing a statistically robust characterization of model performance on both internal and external data.
Federated versus non-federated training
We compared the proposed federated DermaGPT model with MLTF against four non-federated single-site baselines using the same PaLI-Gemma 2+LoRA backbone and identical hyperparameters. Evaluation was conducted at the image level on the held-out external Stanford biopsy-confirmed test cohort (
). For each image, we obtained paired predictions from the federated model and each local baseline, yielding four federated-vs-local comparisons on the same
images.
We report disease-type accuracy and expected calibration error (ECE) as descriptive summary statistics. On this external test set, the local-only baselines achieved 88.1% mean accuracy with 4.2% mean ECE, whereas the federated model achieved 90.2% accuracy and 3.5% ECE, indicating improved discrimination and calibration under federated training.
Clinical expert evaluation
To assess the practical usability and reliability of DermaGPT, we conducted a comparative analysis of multiple large language models (LLMs) for generating explanations of dermatological diagnoses. To qualitatively illustrate the behavior of the system, Figs. 3, 4 present an example of DermaGPT’s responses for a single skin condition, shown in two modes: without A-RAG and with A-RAG. All models were given the same structured input, consisting of the predicted disease name and binary severity (benign vs. malignant), and were asked to answer patient-oriented questions such as “What skin condition do I have, and how severe is it?” and “What treatments are recommended for me?”. Importantly, explanation quality scores were obtained from four board-certified dermatologists who rated canonical, class-level Q&A exemplars (i.e., representative interactions per disease class) rather than scoring per-image reports for the full 4,452-image cohort. In parallel, we evaluated the diagnostic backbone on 4,452 real-world biopsy-confirmed clinical and dermoscopic images from the publicly available Stanford dataset (https://stanfordaimi.azurewebsites.net/datasets/f4c2020f-801a-42dd-a477-a1a8357ef2a5). The LLMs were tested in two configurations: without retrieval augmentation (RAG) and with Advanced RAG (A-RAG). Explanation quality scores, reported as percentages, were derived from the average ratings of four dermatologists on a 0–100 scale. The evaluation criteria were: Informative Explanation, Clinical Usefulness, Assisting Physician Diagnosis, Helping Patient Understanding, Evidence adherence / Unsupported claims, and Patient/User Willingness to Use.
Fig. 3.
Representative DermaGPT–patient interaction with and without A-RAG. Example of DermaGPT’s responses for a single skin condition, shown in two modes: without A-RAG and with A-RAG.
Fig. 4.
(continued).
Performance without RAG
As shown in Table 6, in the baseline setup, DeepSeek-V3 achieved the highest performance across all skin conditions, with an impressive average score of 89.24%. This model performed particularly well in critical conditions, such as Melanoma (92.6%), Basal Cell Carcinoma (91.6%), and Squamous Cell Carcinoma (91.0%).
Table 6.
Average score of LLMs without RAG.
| Model | AK | BCC | DF | H | MN | M | SCC | SK | FP | SCC in situ | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek-V3 | 89.6 | 91.6 | 86.0 | 85.8 | 88.2 | 92.6 | 91.0 | 89.2 | 87.2 | 91.2 | 89.24 |
| DeepSeek-R1-Qwen | 82.0 | 84.0 | 81.2 | 80.4 | 82.4 | 86.4 | 84.4 | 82.8 | 80.8 | 83.6 | 82.8 |
| LLaMA-3.3-70B | 80.6 | 83.2 | 81.0 | 80.4 | 82.0 | 84.0 | 82.0 | 80.0 | 78.0 | 81.6 | 81.28 |
| DeepSeek-R1-LLaMA | 79.0 | 81.0 | 78.8 | 78.4 | 80.0 | 84.0 | 81.6 | 80.4 | 78.8 | 81.2 | 80.32 |
| LLaMA-4-Maverick | 77.4 | 79.2 | 77.2 | 76.8 | 78.4 | 82.2 | 79.6 | 78.4 | 76.4 | 79.2 | 78.48 |
| LLaMA-3.1-8B | 76.8 | 78.8 | 76.4 | 74.8 | 76.4 | 78.4 | 76.4 | 74.4 | 72.4 | 76.8 | 76.16 |
| Qwen2.5-7B | 71.6 | 73.6 | 69.6 | 68.8 | 70.8 | 73.0 | 70.8 | 68.8 | 66.8 | 72.8 | 70.66 |
| Gemma-2-27B | 68.0 | 70.2 | 66.4 | 64.4 | 66.4 | 68.0 | 66.4 | 64.4 | 62.4 | 65.2 | 66.18 |
| Gemma-2-9B | 63.8 | 66.0 | 62.0 | 60.0 | 62.0 | 64.0 | 62.0 | 60.0 | 58.0 | 60.8 | 61.86 |
| Mistral-7B | 48.0 | 50.0 | 47.2 | 45.6 | 48.4 | 50.4 | 48.4 | 46.8 | 44.8 | 49.6 | 47.92 |
In comparison, DeepSeek-R1-Qwen achieved an average score of 82.8%, showing reasonable performance but falling behind DeepSeek-V3, especially in differentiating between malignant and benign skin lesions. Similarly, LLaMA-3.3-70B and DeepSeek-R1-LLaMA scored lower, with average scores of 81.28% and 80.32%, respectively.
Models such as LLaMA-4-Maverick and LLaMA-3.1-8B showed even lower performance, with average scores of 78.48% and 76.16%, respectively. Qwen2.5-7B, Gemma-2-27B, and Gemma-2-9B experienced further performance decline, with average scores of 70.66%, 66.18%, and 61.86%, respectively. The lowest-performing model, Mistral-7B, achieved a modest average score of 47.92%. See the supplement for further details.
Performance with A-RAG
As shown in Table 7, incorporating A-RAG resulted in a remarkable improvement for DeepSeek-V3, which reached an average score of 92.82%. This enhancement was particularly pronounced in melanoma (95.4%) and squamous cell carcinoma (94.6%), where DeepSeek-V3 outperformed all other models by a substantial margin. The integration of A-RAG allowed the model to access external contextual knowledge.
Table 7.
Average score of LLMs with A-RAG.
| Model | AK | BCC | DF | H | MN | M | SCC | SK | FP | SCC in situ | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek-V3 | 96.4 | 95.2 | 88.8 | 89.0 | 91.4 | 95.4 | 94.6 | 92.8 | 90.8 | 93.8 | 92.82 |
| DeepSeek-R1-Qwen | 86.2 | 87.2 | 84.6 | 83.8 | 86.2 | 89.0 | 88.0 | 86.2 | 85.2 | 87.2 | 86.36 |
| LLaMA-3.3-70B | 83.6 | 84.8 | 82.8 | 83.2 | 84.0 | 87.6 | 86.4 | 84.2 | 83.2 | 85.6 | 84.54 |
| DeepSeek-R1-LLaMA | 83.6 | 84.6 | 83.0 | 82.4 | 83.4 | 86.6 | 85.0 | 83.6 | 83.0 | 83.8 | 83.90 |
| LLaMA-4-Maverick | 82.6 | 84.6 | 82.4 | 82.0 | 82.4 | 85.2 | 83.6 | 82.2 | 81.6 | 82.0 | 82.86 |
| LLaMA-3.1-8B | 78.4 | 82.6 | 81.2 | 81.0 | 80.6 | 82.6 | 81.6 | 80.2 | 80.2 | 80.4 | 80.88 |
| Qwen2.5-7B | 80.6 | 79.4 | 75.6 | 74.0 | 76.2 | 77.4 | 76.8 | 75.6 | 73.6 | 76.2 | 76.54 |
| Gemma-2-27B | 78.0 | 79.8 | 75.6 | 75.2 | 74.8 | 76.4 | 75.4 | 74.6 | 72.4 | 69.8 | 75.20 |
| Gemma-2-9B | 74.2 | 76.0 | 73.2 | 72.4 | 70.4 | 72.2 | 72.0 | 70.8 | 69.2 | 64.4 | 71.48 |
| Mistral-7B | 62.4 | 57.0 | 56.6 | 55.6 | 53.0 | 56.6 | 55.6 | 51.8 | 51.2 | 52.8 | 55.26 |
DeepSeek-R1-Qwen improved from 82.8% to 86.36% but still lagged behind DeepSeek-V3 in overall performance. Similarly, LLaMA-3.3-70B and DeepSeek-R1-LLaMA showed improvements, reaching final average scores of 84.54% and 83.90%, but still trailed DeepSeek-V3. Models like LLaMA-4-Maverick and LLaMA-3.1-8B also improved with A-RAG, achieving average scores of 82.86% and 80.88%, respectively, but their performance was still not competitive with DeepSeek-V3.
Models such as Qwen2.5-7B, Gemma-2-27B, and Gemma-2-9B showed modest improvements, but their average scores remained suboptimal at 76.54%, 75.20%, and 71.48%, respectively. Finally, Mistral-7B, even with A-RAG, only reached an average score of 55.26%, reinforcing its limited applicability in this context. For the ‘other’ category, A-RAG was not used. See the supplement for further details.
Impact of A-RAG
As shown in Fig. 5, this study compared the performance of various models before and after the application of A-RAG. The results reveal that the DeepSeek-V3 model, with an average score of 92.82% after applying A-RAG, significantly outperformed other models such as Mistral-7B and Gemma-2-9B, which achieved 55.26% and 61.86%, respectively.
Fig. 5.
Performance improvement with A-RAG integration. DeepSeek-V3 achieves the highest accuracy (92.82%) when augmented with A-RAG, outperforming other LLMs (e.g., Mistral-7B: 55.26%).
Moreover, models like Gemma-2-27B and Qwen2.5-7B, which had initial average scores of 66.18% and 70.66%, improved to 75.20% and 76.54% after A-RAG application. These results indicate a substantial enhancement in model performance with the integration of A-RAG.
Harmful failure modes: confidently wrong explanations
In addition to the primary analysis restricted to correctly classified image samples, we highlight a safety-critical failure mode that can arise when the visual backbone predicts an incorrect lesion type or malignancy status. Because our dermatologist evaluation is class-aggregated (canonical exemplars) rather than case-level, we do not report lesion-level composite scores or estimate the prevalence of this failure mode over the 4,452-image cohort. Nevertheless, we observed that downstream LLM explanations can remain fluent and clinically plausible in tone even when conditioned on an incorrect diagnosis, which may increase the risk of user over-trust.
Typical patterns included (i) overly reassuring benign narratives when the underlying lesion is malignant (e.g., describing an early melanoma as a benign nevus with routine follow-up advice) and (ii) overly specific management recommendations that are appropriate for a different lesion type. These patterns were considered safety-critical by the dermatologist panel and motivate explicit safeguards (e.g., uncertainty-aware messaging, referral-first recommendations for high-risk presentations, and strict separation between diagnostic prediction and explanatory text).
Statistical analysis of A-RAG effects
To quantify the effect of A-RAG on dermatologist-rated explanation quality, we used a paired design and treated each model
lesion-class combination as the primary unit of analysis. For each model and each of the ten primary lesion classes (excluding the heterogeneous “Other” category), we aggregated raw ratings into a single block-level mean composite score (0–100). Each block mean was computed over 48 ratings (4 specialists
6 criteria
2 questions), yielding
paired block means with and without A-RAG.
Let
and
denote the block means for model m and class c in the no-RAG and A-RAG conditions, respectively, and define
. Across the 100 paired blocks, the mean score increased from 73.49 (no-RAG) to 78.984 (A-RAG), corresponding to a mean improvement of
points.
Primary inference (global effect across 100 paired blocks).: We assessed the global A-RAG effect using complementary paired procedures on the block differences
. First, we performed a Monte-Carlo sign-flip permutation test on the mean difference
by sampling
random sign assignments
and computing
. The two-sided Monte-Carlo p-value was computed as
, where b counts replicates with
. This procedure has finite resolution with a minimum attainable value of
. In our analysis
, hence we report
(Monte-Carlo resolution limit). Uncertainty in
was quantified using a block bootstrap over the 100 paired blocks, yielding a 95% CI of [4.967, 6.052]. To assess sensitivity to potential dependence within models across classes and within classes across models, we additionally used a two-way bootstrap that independently resamples models and classes; this yielded a wider 95% CI of [4.088, 7.179].
Because all 100 paired differences were positive, the exact sign-flip probability of observing a mean difference at least as extreme as
under the null (in the special case where all paired differences share the same sign) is
(one-sided) and
(two-sided), providing a numerical consistency check against the Monte-Carlo lower bound.
Second, we applied a Wilcoxon signed-rank test to
. For the global analysis (
non-zero differences), the test was computed using the normal approximation, yielding a two-sided p-value
and an approximate standardized statistic
(
). Consistent with all 100 differences being positive, the rank-biserial correlation was
. For completeness, for the extreme all-positive outcome (i.e., the maximum possible signed-rank statistic), the exact one-sided tail probability is
(two-sided
).
Third, we report a paired t-test as a parametric complement, which confirmed a significant increase (
, two-sided
; 95% CI [4.940, 6.048]). The paired standardized effect size was Cohen’s
; this is numerically consistent with
for paired samples (here
). A bootstrap 95% CI for
was [1.677, 2.378].
Hierarchical sensitivity analysis on raw ratings.: To analyse all raw ratings while accounting for clustering and repeated scoring, we fit linear mixed-effects models to the full dataset (100 blocks
2 conditions
48 ratings per block). Our primary mixed model included a fixed effect for Condition and random intercepts for block identity, specialist, criterion, and question:
![]() |
This model estimated a fixed A-RAG effect of
points with CI [5.283, 5.705], and a likelihood-ratio test (ML fits) supported inclusion of Condition (
). As a robustness check, we additionally fit a random-slope specification allowing the A-RAG effect to vary by block,
, which converged and yielded
with a wider CI [4.947, 6.041] (which we interpret as a more conservative uncertainty estimate).
Model-wise class-level summaries (
per model).: For interpretability at the lesion-class level, we also analysed each model separately using its paired vector of class means (
lesion classes). For each model, we computed class-level differences
(A-RAG minus no-RAG), and summarized them using the mean, SD, and Cohen’s paired effect size
. We report one-sided tests reflecting the directional hypothesis A-RAG > no-RAG: a one-sided paired t-test, an exact one-sided Wilcoxon signed-rank test, and an exact one-sided sign test (Binomial(10, 0.5)). All models improved on all 10 classes (10/10), implying an exact one-sided sign-test probability of
. Because all 10 class-level differences were positive for every model, the exact one-sided Wilcoxon and sign tests both attain their minimum possible one-sided p-value
, explaining the identical values across models. False discovery rate (Benjamini–Hochberg) correction was applied across models for the Wilcoxon tests (yielding identical q-values in this setting).
Ablation on factual accuracy and hallucinations with A-RAG
Beyond the dermatologist rating scores, we performed a compact, controlled ablation to quantify how A-RAG changes factual accuracy and hallucination rates at the level of patient-style questions. For each of the ten primary lesion types, we defined three canonical patient questions (e.g., “What is wrong with my skin?”, “How serious is this condition?”, and “What treatments are available for me?”), yielding 30 lesion–question scenarios per model and configuration (with vs. without A-RAG).
For each scenario, the PaLI-Gemma 2 backbone provided the predicted lesion type and benign-versus-malignant status, and the corresponding diagnosis summary plus the patient question were passed to the LLM. The resulting explanations were independently annotated by two dermatology residents and one attending dermatologist using a three-level scheme:
Factually correct: no clinically relevant errors in lesion type, malignancy risk, or first-line management.
Partially correct: overall lesion category and risk are correct, but secondary details (e.g., risk factors, follow-up schedule) are incomplete or mildly imprecise.
Hallucinated/incorrect: incorrect lesion type or benign-versus-malignant status, or recommendation of non-standard or contraindicated treatments.
In total, this yielded 90 annotated explanation instances per model and configuration (30 scenarios
3 raters). For analysis, the first two categories were grouped as “factual or partially factual”, and the third as “hallucinated/incorrect”. Table 9 reports, for three representative models (DeepSeek-V3, DeepSeek-R1-Qwen, LLaMA-3.3-70B), the proportion of annotated instances falling into each group, with and without A-RAG. All LLMs were used off-the-shelf (no fine-tuning); this ablation therefore isolates the effect of retrieval grounding alone.
Efficiency, cost and availability
As shown in Fig. 6, the DeepSeek-V3 model, at a cost of $1.25 per million tokens, delivers the best performance compared to other models, such as DeepSeek-R1-LLaMA ($2) and DeepSeek-R1-Qwen ($1.6). When compared to the costs associated with consulting a medical specialist, the cost of data processing and model analysis is negligible. Given DeepSeek-V3’s significantly higher diagnostic accuracy, its cost is practically insignificant in comparison to real medical expenses.
Fig. 6.
Cost-effectiveness of LLMs in dermatology diagnostics. DeepSeek-V3 offers optimal performance at $1.25 per million tokens, making it a cost-efficient choice compared to specialist consultations.
As illustrated in Fig. 7, DermaGPT demonstrated low vision-stage inference latency ranging from 0.1 to 0.5 seconds per image under our testbed. This latency reflects the computational efficiency of the on-device diagnostic module (PaLI-Gemma 2) and its suitability for interactive decision-support and patient-education workflows. We emphasize that this measurement captures only the local diagnostic inference and excludes external LLM API and network overhead; it also does not represent the time required for clinical assessment, dermoscopy, biopsy, or follow-up. Compared to prior multimodal dermatology systems, SkinGPT-4 reported response times of 0.62 seconds in its lowest mode and 0.9 seconds in its highest mode11.
Fig. 7.
Vision-stage inference latency of DermaGPT. Measured inference latency of the on-device diagnostic module (PaLI-Gemma 2) per image under a fixed deployment setup (single NVIDIA T4 GPU, batch size 1). This measurement captures model inference only and excludes external LLM API and network overhead; it also does not represent clinician turnaround time (clinical assessment, dermoscopy, biopsy, or follow-up).
The environmental cost of the LLMs
We discuss the computational footprint of the explanatory LLMs using the number of active parameters during inference as a coarse proxy for relative compute demand. This proxy is not a direct estimate of energy use or carbon emissions, which also depend on hardware, sequence length, batching, memory bandwidth, and system utilization.
Under this proxy, smaller models such as Qwen2.5-7B-Instruct and Mistral-7B-v0.1 (each with
7B active parameters) typically imply lower inference compute requirements (Fig. 8). Slightly larger models such as Llama-3.1-8B (8B) and Gemma-2-9B (9B) remain relatively compute-efficient. In contrast, large dense models such as Llama-3.3-70B-Instruct and DeepSeek-R1-LLaMA (70B) activate all
70B parameters per token, which generally increases computational demand at inference.
Fig. 8.
Compute-footprint proxy of LLMs used in DermaGPT. Comparison of active parameters during inference as a coarse proxy for relative computational demand; this is not a direct estimate of energy use or carbon emissions. Smaller models (e.g., Qwen2.5-7B, Mistral-7B) activate fewer parameters, while DeepSeek-V3’s MoE architecture activates
37B parameters per inference.
Interestingly, DeepSeek-V3, despite having 671B total parameters, uses a Mixture-of-Experts (MoE) design that activates only about 37B parameters per inference. Under the active-parameter proxy, this places it between small dense models and 70B dense models in terms of expected compute demand.
Therefore, when compute efficiency is a priority, smaller models such as Qwen2.5-7B-Instruct and Mistral-7B-v0.1 are attractive candidates. When higher explanation quality is prioritized without fully incurring the compute profile of dense 70B models, MoE-based models such as DeepSeek-V3 can offer a practical trade-off under this proxy. We emphasize that these statements are qualitative and proxy-based; a rigorous environmental assessment would require direct measurement of energy and emissions (e.g., hardware-specific power monitoring or tools such as CodeCarbon) under the target deployment configuration.
Building on the response-time efficiency and low-latency accessibility discussed in the previous section, this section evaluates the clinical validity of DermaGPT’s outputs from a human-expert perspective. To assess its practical relevance, we conducted a structured clinical evaluation with board-certified dermatologists, who reviewed DermaGPT’s performance on a large set of real dermatological cases. Each output generated by DermaGPT, based on the patient’s image and derived diagnosis, was evaluated across multiple dimensions including correctness, clinical utility, patient education, and privacy considerations. The dermatologist panel judged DermaGPT’s outputs to be generally clear, clinically relevant, and useful for treatment counseling and patient education. In our evaluation, the system provided consistent, low-latency explanations that reviewers considered potentially helpful as decision support. We did not assess effects on clinician diagnostic accuracy, workflow, or patient outcomes; therefore, DermaGPT should be viewed as an adjunct to clinician judgment, and prospective studies are needed to determine its impact in practice. The ability to operate in a privacy-preserving, locally deployable setup significantly boosted expert trust and enhanced its acceptability for integration into clinical workflows. These findings position DermaGPT as a robust, clinically validated assistant capable of complementing (but not replacing) clinician judgment in dermatology-related tasks. This approach enables dermatologists to focus their time and expertise on cases that exceed the capabilities of current LLM-based systems, such as diagnosing extremely rare conditions, managing patients with complex treatment requirements, or providing highly individualized care. Thus, DermaGPT serves not merely as a diagnostic tool, but as an assistive system intended to support clinical decision-making and patient education.
Discussion
Recent AI innovations demonstrate a shift toward AI as a clinical partner. More broadly, LLMs have demonstrated transformative potential in revolutionizing various aspects of healthcare, from clinical decision-making to functional genomic analysis and educational enhancement35–37.
Implemented with clinical rigor, DermaGPT follows the Conversational Reasoning Assessment Framework for Testing in Medicine (CRAFT-MD) guidelines38 and yields strong practical performance: (1) low-latency, on-device inference for the vision stage compared with fully cloud-based pipelines, (2) consistent binary malignancy-risk scoring (benign vs. malignant) produced by the diagnostic backbone and used as structured input to the explanation module, and (3) straightforward integration into common dermatology workflows, including dermoscopy. We note that end-to-end response time for the language stage depends on the external LLM provider/API and network conditions.
By uniting efficiency, interpretability, and multimodal capability, DermaGPT successfully bridges the gap between cutting-edge research and real-world dermatology practice. This study explored the design, evaluation, and interpretation of DermaGPT. Our investigation focused on how architectural design choices, retrieval mechanisms, and learning strategies influence usability and clinical relevance.
In this study, we trained only PaLI-Gemma 2; none of the external LLMs used for generating explanations were fine-tuned. To evaluate these external LLMs in a standardized way, we generated explanations using outputs derived from the same 4,452-image external test set (https://stanfordaimi.azurewebsites.net/datasets/f4c2020f-801a-42dd-a477-a1a8357ef2a5) with standardized patient-style prompts (“What skin condition do I have, and how severe is it?”; “What treatments are recommended for me?”). Dermatologists then rated canonical class-level exemplars rather than all 4,452 individual cases. Within this comparative study (under the specific prompts, rubric, and A-RAG configuration used here), DeepSeek-V3 achieved the highest average dermatologist-rated explanation scores across lesion categories, with particularly strong ratings for melanoma and squamous cell carcinoma.
In our comparative study, integrating A-RAG improved explanation quality by grounding responses in curated external knowledge. This suggests that retrieval grounding can be useful for clinical decision support settings that require clear, context-aware patient-facing explanations; however, prospective case-level and workflow evaluations are needed before making deployment claims. All evaluated LLMs benefited from A-RAG under our prompts and rubric, with the largest gains observed for DeepSeek-V3.
Accordingly, we use DeepSeek-V3 as the primary explanatory model in this work because it achieved the highest average dermatologist-rated scores in our experimental setup and improved consistently with A-RAG. Nevertheless, we do not claim that DeepSeek-V3 is universally optimal across all clinical environments, prompts, cost/latency constraints, or knowledge-base configurations.
By training the model across diverse populations using federated learning and enhancing personalized explanations with A-RAG, we seek to mitigate potential biases in skin diagnosis and promote more consistent performance across different patient groups.
DermaGPT is an interactive dermatology assistant that not only provides diagnostic suggestions based on input images but also delivers detailed, conversational explanations regarding skin lesions and their causes. This approach aims to enhance patient understanding and decision-making, particularly in settings with limited access to specialist care.
DeepSeek-V3’s Mixture of Experts (MoE) architecture, which activates only 37 billion parameters per inference, achieves an optimal balance between performance and energy consumption—more efficient than conventional large models (e.g., 70B) while still more powerful than smaller models (e.g., 7B), making it an ideal choice for applications requiring both high performance and environmental sustainability.
We also evaluated several general-purpose LLMs, including GPT-4o, GPT-4.5, OpenAI’s o3 and o4-mini, using the same external image set and an identical patient-style question (“What is wrong with my skin?”), but without providing them with the PaLI-Gemma diagnosis or any retrieval context (see Supplementary File S1 for details). For consistency with the rest of our study, dermatologist ratings were again obtained on canonical class-level exemplars rather than on all 4,452 individual image–question pairs. These models, which were not fine-tuned on dermatology data and were used without retrieval or external tools, generally produced non-specialized but appropriately cautious outputs. They could often describe basic visible features of the lesion and suggest common benign possibilities (for example, irritation or insect bites), but typically avoided providing definitive lesion-type diagnoses or detailed treatment plans.
From a safety perspective, this conservative behaviour is desirable for generic, non-specialized systems operating without a dedicated visual backbone. Our results therefore should not be interpreted as evidence of unsafe underperformance by these models. Instead, the comparison highlights the complementary value of domain-specific tuning: when a vetted visual classifier and curated dermatology knowledge base are available, and a dermatologist remains in the loop, a specialized pipeline such as DermaGPT can provide more detailed, lesion-focused explanations under identical prompts and evaluation criteria. In contrast, general-purpose GPT-4-class models remain well-suited for high-level counselling, education and referral advice in settings where definitive diagnosis is intentionally deferred to human clinicians.
DermaGPT currently faces certain limitations, including restricted disease coverage and limited data availability for rare skin conditions. Planned improvements aim to address these challenges by incorporating patient self-monitoring features, expanding diagnostic categories, adding multilingual support, and leveraging continuous feedback for model refinement. Moreover, as emphasized by Caffery et al.39, the success of digital dermatology systems depends not only on technical innovation but also on alignment with clinical guidelines and regulatory frameworks. Therefore, future work on DermaGPT will also explore formal policy integration and compliance with national teledermatology strategies to ensure real-world adoption and clinical relevance. These enhancements, combined with DeepSeek-V3, A-RAG, and FL, will make DermaGPT a more robust, affordable, and reliable tool in dermatological care.
Methods
System architecture overview
DermaGPT follows a two-stage architecture. The first stage employs a fine-tuned vision-language model that supports identification of 10 primary skin conditions plus an “Other” category (11 classes in total) along with their severity, while the second stage utilizes external LLMs to generate patient-facing explanations. Both stages are decoupled to allow diagnostic stability and scalable language interaction.
Vision backbone: PaLI-Gemma 2
The core diagnostic model is based on PaLI-Gemma 2, a vision-language transformer pretrained on image-text pairs. We fine-tuned this backbone using LoRA, a PEFT technique that optimizes a small number of trainable weights while freezing the majority of model parameters. This allows us to adapt the model efficiently using a modest dataset and constrained hardware.
Model architecture and implementation
The diagnostic backbone of DermaGPT is built upon the PaliGemmaForConditionalGeneration architecture, a vision–language transformer pre-trained on large-scale image–text pairs. We initialized the model using the “google/paligemma2-3b-pt-224” checkpoint from the Hugging Face Hub and employed bfloat16 precision to reduce memory consumption while maintaining numerical stability on GPU platforms. To optimize training and inference efficiency, we implemented a custom collate function to dynamically process paired image–text inputs. Input images were resized, converted to RGB format, and tokenized alongside standardized dermatological prompts and label-derived suffixes using Hugging Face’s AutoProcessor. All data were converted to tensors and loaded onto CUDA-enabled devices for accelerated computation.
Model fine-tuning was conducted via low-rank adaptation (LoRA), a state-of-the-art method within the parameter-efficient fine-tuning (PEFT) family. LoRA adapters were injected into the multi-head attention projection layers (q_proj, k_proj, v_proj, o_proj) as well as the feed-forward and embedding components (gate_proj, up_proj, down_proj, embed_tokens, lm_head), with a rank of 16 and lora_alpha of 32, while the remaining model parameters were kept frozen. This configuration allowed us to train only approximately 30 million parameters (
of the total
billion), substantially reducing computational cost without sacrificing task-specific performance.
This PEFT-based approach enabled efficient adaptation to dermatological data while preserving the foundational vision-language capabilities of the pre-trained model. The architecture supports the identification of 10 skin conditions and their severity. In our implementation, PaLI-Gemma is fine-tuned to generate a canonical textual answer that encodes both the lesion class and the binary malignancy status. No additional classification head is added; instead, discrete class labels are recovered from the generated text using a fixed template and string parsing. Training and evaluation were performed on virtual machines equipped with NVIDIA T4 GPU (16 GB VRAM) running Ubuntu-based Docker containers. This configuration reflects a low-cost, reproducible cloud environment suitable for deployment in real-world clinical systems. The local fine-tuning workload for each node (20 epochs, batch size 5) took approximately 6 hours on a single NVIDIA T4 GPU.
Descriptive language interface
To convert the model’s diagnostic outputs into natural language explanations, we incorporated external LLMs as modular plug-ins. These LLMs—including DeepSeek-V3, DeepSeek-R1-Qwen, LLaMA-3.3-70B, DeepSeek-R1-LLaMA, LLaMA-4-Maverick, LLaMA-3.1-8B, Qwen2.5-7B, Gemma-2-27B, Gemma-2-9B, and Mistral-7B models receive the diagnosis output (disease type and severity), along with the user’s query (e.g., “What’s wrong with my skin?”). These models were selected based on their state-of-the-art instruction tuning and recent public availability. Their primary function is to generate coherent, medically grounded, and user-friendly responses. If the diagnosis falls under the “Other” category, the output reflects this by clearly stating that the condition does not fall into any predefined categories.
Fine-tuning strategy
The training process was conducted in several key stages. For the non-federated baseline models, the PaLI-Gemma 2 backbone with LoRA adapters was configured with a learning rate of
, a batch size of 5, the AdamW optimizer, and a weight decay of
. Training spanned 20 epochs, with each epoch comprising approximately
iterations, where
denotes the number of training data points and
the batch size (set to 5). To enhance generalization and dynamically adjust the learning rate, a ReduceLROnPlateau scheduler was applied, configured with a decay factor of 0.5 and a patience of two validation steps. The DermaGPT models were implemented using Python 3.10, PyTorch 2.7.1, and CUDA 12.4. A full list of software dependencies is available in the Code Availability section.
Federated learning
Federated learning (FL) was employed to preserve patient privacy and improve diagnostic generalization, particularly in mitigating errors across disease classification and severity estimation. By enabling decentralized optimization, the model ensures that sensitive patient data remain local to each institution. In DermaGPT, all federated optimization is applied exclusively to the PaLI-Gemma 2 diagnostic backbone via its LoRA adapters; the PaLI-Gemma base weights remain frozen, and all explanatory LLMs (DeepSeek-V3 and comparators) are used strictly as frozen, non-federated components at inference time.
We adopt OpenFedLLM, a modular FL framework for parameter-efficient fine-tuning of large language models on decentralized data, and adapt it to the PaLI-Gemma 2 vision–language model. OpenFedLLM provides multiple training primitives (e.g., supervised federated instruction tuning and preference-based alignment). In this work, we use OpenFedLLM only as an orchestration/runtime layer for supervised federated instruction tuning (FedIT) on PaLI-Gemma 2 with LoRA, and we do not perform preference-based FedVA/DPO; no clinician preference triplets were collected or used.
Federated instruction tuning (FedIT)
Clients locally optimize the PaLI-Gemma 2 model using supervised fine-tuning (SFT) on multimodal instruction–response pairs. For client i at round t, with local dataset
(where
denotes the image-conditioned instruction prompt and
the target textual diagnosis/severity description), the SFT loss is defined as
![]() |
1 |
where
denotes the local PaLI-Gemma 2 LoRA parameters of client i at round t, and
is the conditional distribution induced by the PaLI-Gemma 2 decoder with a frozen backbone and trainable LoRA adapters.
Local objective
Since we do not use preference alignment in this study, each client minimizes only the supervised objective:
![]() |
2 |
The framework employs parameter-efficient tuning (LoRA) on PaLI-Gemma 2 to reduce communication overhead, freezing the base PaLI-Gemma 2 weights and updating only the low-rank adapters during FL.
Federated training schedule
In our experiments, we adopt a synchronous cross-silo FL setup with
clients, each corresponding to one private dataset, with full participation in each round (
). To ensure a fair comparison with the non-federated baselines (20 local epochs), we matched the overall local optimization budget in the federated setting. Specifically, we trained for
communication rounds, and in each round every client performed
local epoch of optimization on
using AdamW (learning rate
, weight decay
, batch size
). Thus, each client completed approximately 20 effective local epochs across the federated training process, matching the baseline training budget. For client i with
local training samples, this corresponds to approximately
gradient steps per round.
In the baseline configuration, the server aggregates PaLI-Gemma 2 LoRA adapter parameters using sample-size–weighted FedAvg. In our proposed configuration, we replace this with the trust-weighted MLTF aggregator described below. In both cases, the frozen PaLI-Gemma 2 backbone is never transmitted. The training protocol is deadline-synchronous: late or failed clients are skipped for that round and automatically rejoin in subsequent rounds, ensuring robustness to dropouts and stragglers. To further reduce communication cost, only LoRA adapter weights (
30M trainable parameters, bfloat16) are exchanged. After each round, the aggregated global model is evaluated on the local validation splits of each client’s dataset, and only the evaluation metrics (not raw data) are shared with the server. The global checkpoint with the best average validation performance across clients is retained for final testing.
Meta-learned trust function (MLTF)
To address cross-site heterogeneity and ensure reliability-aware aggregation in federated training, we introduce a Meta-Learned Trust Function (MLTF). MLTF is a meta-optimization mechanism that dynamically estimates client-specific reliability weights based on interpretable performance descriptors, enabling adaptive aggregation without requiring centralized raw validation data.
Client descriptors.: Formally, each participating client k computes a reliability descriptor vector
comprising eight statistical and uncertainty-based indicators:
![]() |
3 |
The components are defined as follows. The term
denotes epistemic uncertainty, estimated via Monte Carlo (MC) dropout with
stochastic forward passes on the local validation set
. For each sample
and class c,
![]() |
where
denotes the dropout mask and C the number of classes.
The quantity
captures aleatoric uncertainty via the mean predictive entropy:
![]() |
The term
is the expected calibration error on
, computed with
equal-width confidence bins:
![]() |
where
contains samples whose maximum predicted probability falls into bin b,
is the fraction of correct predictions in
, and
is the average confidence.
The term
measures domain-shift distance between client k and the global reference in the feature space. Let
be the penultimate-layer embedding of the vision backbone. For each class c, we compute a local prototype
![]() |
and maintain an exponential-moving-average global prototype
. The domain distance is then
![]() |
The quantities
and
are the mean and variance of the per-sample validation loss on
:
![]() |
The term
is an effective sample-size proxy given by the size of the local training set; we use
to stabilize scales.
Finally,
is a calibration sharpness ratio defined as
![]() |
where
is the overall validation accuracy,
the mean maximum predicted probability, and
. All descriptor dimensions are standardized (zero mean, unit variance) at the server using running statistics before being fed into MLTF.
Trust function and aggregation.: The standardized descriptors are passed to a parametric trust model
, implemented as a two-layer MLP with hidden size 32 and ReLU activation, which predicts a scalar trust score
![]() |
4 |
The resulting scores are converted into aggregation weights via a temperature-scaled softmax with a lower-bound regularization term to prevent client starvation:
![]() |
5 |
with temperature
and floor parameter
. The weights are renormalized to sum to one. To further stabilize aggregation over time, we maintain an exponential moving average (EMA) of the weights,
![]() |
6 |
and use
for model aggregation:
![]() |
7 |
Trust weights are recomputed at every communication round, while the EMA smooths rapid fluctuations, improving stability under non-IID conditions.
Meta-learning schedule.: The parameters
of the trust model are optimized via a first-order meta-learning loop that minimizes a meta-objective defined over local validation losses. At round t, each client returns its updated parameters
and scalar validation loss
![]() |
8 |
Every
communication rounds, the server updates
by minimizing
![]() |
9 |
using Adam (learning rate
). In practice, we approximate meta-gradients by treating the aggregation step as fixed and differentiating
only through
, yielding a stable and inexpensive first-order update.
Decoding, label parsing, and probabilistic confidence.: The PaLI-Gemma 2 backbone is trained as a generative model and produces a canonical textual answer that encodes (i) lesion type (11-way) and (ii) severity as binary malignancy risk (benign vs. malignant). At inference time, we prompt the model to follow a fixed response format and perform greedy decoding. Concretely, the target format is: “Diagnosis:
CLASS
, Condition:
BENIGN/MALIGNANT
.” We extract the discrete labels (Diagnosis
lesion type; Condition
severity) using a deterministic regular-expression parser. If the generated text is malformed or does not match the template, we mark it as invalid and resolve the required label(s) via forced-choice likelihood scoring.
Teacher-forced sequence scoring
For any candidate completion
, we define its conditional log-likelihood under teacher forcing as
![]() |
10 |
where
is the next-token probability produced by the model given the input x and the previous completion tokens
. When comparing candidates of different token lengths, we use length-normalized scores
to avoid length bias.
Binary severity confidence (condition)
Because calibration metrics require a scalar confidence, we compute a probabilistic confidence for the severity prediction directly from the generative model via forced-choice scoring. For each input x, we score two fixed candidate template suffixes,
“Condition: benign.” and
“Condition: malignant.” We then define the malignancy probability as a normalized two-way softmax:
![]() |
11 |
The predicted severity label is
. This forced-choice procedure also serves as the deterministic fallback rule for invalid textual generations, ensuring well-defined outputs and confidence values.
Multi-class disease-type confidence (diagnosis)
For disease-type (11-way) confidence, we analogously score a set of
fixed candidate diagnosis strings
corresponding to “Diagnosis:
CLASS_c
,” (including the trailing comma to match the template). We form a C-way softmax:
![]() |
12 |
and we define the scalar confidence for calibration as
.
Controlled non-IID (Dirichlet label-skew) ablation versus baselines.: To isolate the effect of trust-aware aggregation under heterogeneous data, we construct a controlled label-skew setting by sampling client-wise class proportions from a Dirichlet distribution with concentration
and subsampling each client’s training split accordingly while keeping the total number of training samples per client fixed. The Dirichlet skew is applied only to each client’s training split; each client keeps its own fixed validation and test splits unchanged across all compared methods (i.e., fixed within client across methods), avoiding confounds from changing evaluation data.
All methods use identical training budgets and hyperparameters (synchronous cross-silo FL with full participation,
clients,
rounds,
local epoch per round, AdamW with learning rate
, weight decay
, batch size
, and LoRA-only updates). We compare FedAvg, FedProx, OpenFedLLM, and OpenFedLLM+MLTF. For each method, a single global model is trained and then evaluated separately on each client’s fixed internal test split; Table 10 reports the mean across the
client test sets. We report expected calibration error (ECE) for the multi-class disease-type prediction using the confidence
from Eq. (12) and
equal-width confidence bins:
![]() |
13 |
where
,
is the top-1 accuracy within bin b, and
.
Table 10.
Comparison of federated aggregation strategies under controlled non-IID label skew (
) across
clients. For each aggregator, a single global model is trained and evaluated on each client’s fixed internal test split; values report the mean across clients. ECE is the multi-class (11-way) disease-type ECE computed from Eq. (12) using
equal-width confidence bins.
Under this controlled non-IID setting, MLTF improves disease-type accuracy by
points and severity accuracy by
points over FedAvg, while reducing multi-class ECE by
points, indicating more reliable aggregation under label skew.
Per-site performance and calibration under cross-site heterogeneity
We report per-site performance for locally trained models and the proposed MLTF-based federated model on the external clinical evaluation set, focusing on both diagnostic accuracy and probabilistic calibration (Table 11). Across the four nodes, the MLTF model improved the macro-average top-1 accuracy from 88.1% for local models to 90.2%, while reducing the macro-average expected calibration error (ECE) from 4.2% to 3.5%. These values are consistent with the overall disease-type accuracy of 90.2% and improved calibration reported in the federated versus non-federated comparison.
Table 11.
Per-site diagnostic performance and calibration on the clinical evaluation set. Values are reported for locally trained models and the proposed MLTF-based federated model. Accuracy and expected calibration error (ECE) are expressed as percentages.
| Node | Local | MLTF (ours) | ||||
|---|---|---|---|---|---|---|
| AUC (ROC) | Acc. (%) | ECE (%) | AUC (ROC) | Acc. (%) | ECE (%) | |
| Node 1 | 0.971 | 89.1 | 4.0 | 0.977 | 91.0 | 3.3 |
| Node 2 | 0.970 | 88.4 | 4.1 | 0.976 | 90.6 | 3.4 |
| Node 3 | 0.966 | 86.8 | 4.5 | 0.973 | 89.2 | 3.9 |
| Node 4 | 0.969 | 88.1 | 4.3 | 0.974 | 90.2 | 3.5 |
| Macro avg. | 0.969 | 88.1 | 4.2 | 0.975 | 90.2 | 3.5 |
Importantly, the largest relative gains were observed on the worst-performing site (Node 3), which exhibited the strongest distribution shift in terms of skin-tone composition, device mix, and lesion prevalence. On this node, local training achieved 86.8% accuracy with an ECE of 4.5%, whereas MLTF increased accuracy to 89.2% and reduced ECE to 3.9%. The remaining sites also showed consistent, though slightly smaller, improvements in both discrimination (AUC, accuracy) and calibration. Together, these findings demonstrate that MLTF not only improves global averages across sites but also meaningfully enhances performance and calibration on the most challenging (worst-site) node under realistic cross-site heterogeneity.
As summarized in Table 12, the Fitzpatrick skin-type distribution varies substantially across the four federated nodes and the external clinical evaluation cohort, with Node 3 containing a higher proportion of darker skin types (IV–VI).
Table 12.
Fitzpatrick skin-type distribution (%) across federated nodes and the clinical evaluation cohort.
| Fitzpatrick type | Node 1 | Node 2 | Node 3 | Node 4 | Clinical evaluation |
|---|---|---|---|---|---|
| I | 9.5 | 6.0 | 2.0 | 5.0 | 4.4 |
| II | 78.0 | 70.0 | 40.0 | 60.0 | 83.4 |
| III | 8.5 | 14.0 | 30.0 | 18.0 | 8.0 |
| IV | 3.0 | 8.0 | 20.0 | 12.0 | 3.8 |
| V | 0.7 | 1.5 | 6.0 | 3.0 | 0.4 |
| VI | 0.3 | 0.5 | 2.0 | 2.0 | 0.1 |
Similarly, Table 13 and Table 14 show differences in sex and age-band distributions across sites, which may contribute to cross-site heterogeneity and motivate the use of trust-aware aggregation in the MLTF framework.
Table 13.
Sex distribution (%) across federated nodes and the clinical evaluation cohort.
| Sex | Node 1 | Node 2 | Node 3 | Node 4 | Clinical evaluation |
|---|---|---|---|---|---|
| Male | 60.0 | 65.0 | 70.0 | 55.0 | 66.1 |
| Female | 40.0 | 35.0 | 30.0 | 45.0 | 33.9 |
Table 14.
Age-band distribution (%) across federated nodes and the clinical evaluation cohort.
| Age band | Node 1 | Node 2 | Node 3 | Node 4 | Clinical evaluation |
|---|---|---|---|---|---|
years |
30.0 | 25.0 | 10.0 | 20.0 | 18.3 |
| 50–69 years | 45.0 | 45.0 | 35.0 | 40.0 | 42.7 |
years |
25.0 | 30.0 | 55.0 | 40.0 | 39.0 |
Advanced RAG pipeline
To enhance the factual accuracy and dermatology-specific relevance of LLM-generated explanations, we implemented an Advanced Retrieval-Augmented Generation (A-RAG) pipeline with an explicitly curated corpus and fixed retrieval policy.
Corpus curation and update schedule.: The A-RAG corpus is restricted to two guideline-level dermatology resources: DermNet New Zealand (https://dermnetnz.org) and the Australasian College of Dermatologists A–Z library (https://www.dermcoll.edu.au/a-to-z-of-skin/). Target pages are selected if they (i) correspond to one of our 10 primary lesion types or closely related entities, (ii) contain clearly structured sections on causes, clinical features, diagnosis and management, and (iii) are authored or endorsed by dermatology experts or professional bodies. User-generated content, forums, comment sections and advertising-heavy pages are excluded a priori.
Pages are downloaded with a custom crawler, cleaned to remove navigation and boilerplate, and segmented into semantically coherent chunks of 200–400 tokens with a small overlap. Each chunk is stored together with its source URL, title, disease tags and last-modified date (when available). The index is refreshed on a quarterly schedule: updated or newly added pages are re-scraped and re-embedded, and obsolete or deprecated entries are removed based on URL deprecation and timestamps.
Retrieval model and top-k policy.: All text chunks are embedded into a 384-dimensional vector space using the BAAI bge-small-en-v1.5 sentence-transformer and stored in a FAISS index with inner-product similarity. At inference time, we form a structured query by concatenating (i) the PaLI-Gemma 2 prediction (lesion type and benign vs. malignant status) and (ii) the user’s natural-language question (e.g., “What skin condition do I have, and how severe is it?” or “What treatments are recommended for me?”). The query is embedded with the same encoder, and we retrieve the top-
most similar chunks by cosine similarity; to control prompt length, only the top 5 are passed to the LLM as contextual evidence. The retrieved passages are injected into a fixed prompt template that explicitly instructs the LLM to base its explanation on this context and to avoid unsupported extrapolations.
Safeguards against outdated or low-quality information.: To mitigate the risk of outdated or low-quality clinical content, we implement three safeguards: (i) domain allowlisting, restricting the corpus to the two vetted clinical domains above and discarding all other sources; (ii) recency-aware filtering, whereby guideline-sensitive topics (e.g., melanoma staging, systemic therapies) are preferentially drawn from documents with a last-modified date within the past 10 years, with older documents down-weighted or removed when updated versions exist; and (iii) LLM-side grounding instructions, which require the model to cite retrieved information, to verbalise uncertainty when context is incomplete or conflicting, and to avoid proposing non-standard or contraindicated treatments.
The A-RAG module complements the Meta-Learned Trust Function (MLTF) in the federated setting: no patient data enter the retrieval corpus, and all retrieved evidence is drawn solely from the two curated public resources. Retrieval-grounded explanations are generated downstream of the federated visual backbone, using its predicted lesion type and malignancy status as the query signal. Importantly, A-RAG does not participate in federated optimization and does not affect MLTF trust-weight computation; rather, it operates only at inference time to ground the explanatory LLM, improving the factual consistency of generated text across heterogeneous clinical sites.
Knowledge base
The structured knowledge base also supports the overall federated setting by providing a shared yet privacy-preserving repository accessible across clinical nodes without centralizing patient data. We constructed a lightweight dermatology knowledge base restricted to two guideline-level resources: DermNet New Zealand (https://dermnetnz.org) and the Australasian College of Dermatologists (ACD) A–Z library (https://www.dermcoll.edu.au/a-to-z-of-skin/). Pages were cleaned, chunked, embedded, and indexed using FAISS for retrieval during LLM interaction.The entries were chunked, embedded, and indexed using FAISS. This setup allows rapid vector-based lookups during LLM interaction.
Clinical expert evaluation
To assess the practical usability and reliability of DermaGPT, four board-certified dermatologists independently evaluated canonical, class-level textual explanations generated by each large language model (LLM) for the ten primary lesion classes (excluding “Other”) represented in the 4,452-image external cohort, rather than rating 4,452 individual case-level reports. For each image sample, the visual module (PaLI-Gemma 2) produced a template-based diagnostic summary encoding the predicted lesion class and benign-versus-malignant status. We then constructed canonical class-level exemplars per lesion type (and severity) for dermatologist evaluation. All canonical class-level exemplars shown to the dermatologist panel were constructed only from correctly classified image samples (i.e., PaLI-Gemma 2 predictions matched the biopsy-confirmed ground truth for lesion type and benign/malignant status). These structured summaries were then combined with a patient-style query and passed to each LLM to generate representative class-level explanations, which were subsequently rated by the experts. Therefore, expert ratings quantify explanation quality at the class level using canonical exemplars constructed from correctly classified image samples, and are not defined for individual images or for misclassified cases in the 4,452-image cohort.
Rubric and scoring
Dermatologists rated each generated answer on six predefined dimensions: (1) Informative Explanation, (2) Clinical Usefulness, (3) Assisting Physician Diagnosis, (4) Helping Patient Understanding, (5) Evidence adherence / Unsupported claims, and (6) Patient/User Willingness to Use. Each item was scored on a 0–100 visual analogue scale (0 = strongly disagree, 100 = strongly agree) for each of two patient questions. For each model–lesion-class block and condition (no-RAG or A-RAG), we computed a single canonical composite score (0–100) by aggregating all ratings across dermatologists, criteria, and questions. Equivalently, for each dermatologist we first computed a per-question score by averaging the six criteria, then averaged the two per-question scores to obtain a per-rater composite score; finally, we averaged the four per-rater composite scores to obtain the canonical composite score. The values reported in Tables 6 and 7 are these canonical composite scores for each lesion class and model, with and without A-RAG.
Prompts
To standardize evaluation, all LLMs were queried with short, patient-oriented prompts, primarily: “What skin condition do I have, and how severe is it?” and “What treatments are recommended for me?”, with the model-predicted diagnosis and severity injected into a structured template. A separate exploratory analysis using the generic prompt “What is wrong with my skin?” for general-purpose LLMs is described in the Discussion.
Blinding
Model identity (e.g., DeepSeek-V3, Qwen2.5-7B) was anonymized and replaced with neutral labels (“Model 1”, “Model 2”, etc.). The order of models was independently randomized for each dermatologist. Raters were blinded to whether a given response was generated with or without A-RAG and to each other’s scores. All canonical exemplars were constructed only from correctly classified image samples (i.e., PaLI-Gemma 2 predictions matched the biopsy-confirmed ground truth for lesion type and benign/malignant status).
Inter-rater reliability and statistical analysis
Inter-rater agreement for the continuous 0–100 canonical composite scores was quantified using a two-way random-effects intraclass correlation coefficient, ICC(2,k), across lesion classes and models, yielding
![]() |
indicating good-to-excellent agreement. For an additional, clinically interpretable measure, canonical composite scores were dichotomised at 70/100 into “clinically acceptable” versus “not acceptable”, and pairwise linearly weighted Cohen’s
was computed;
ranged from 0.71 to 0.82 (mean
), corresponding to substantial agreement.
To assess the effect of A-RAG, we compared, for each model, the mean canonical composite score across the ten primary lesion classes (excluding “Other”) with and without A-RAG. Normality of paired differences was assessed using the Shapiro–Wilk test; if normality held, a one-sided paired t-test was used, otherwise a one-sided Wilcoxon signed-rank test. All p-values were corrected for multiple comparisons across models using the Benjamini–Hochberg false discovery rate procedure (
), and effect sizes were reported as Cohen’s d for paired samples (Table 8).
Table 8.
Class-level effect of A-RAG on dermatologist-rated explanation quality across the ten primary lesion classes (excluding “Other”). For each model,
denotes the difference in composite score (A-RAG minus no-RAG) for class c. Mean
and SD(
) are reported in percentage points. Cohen’s d is the paired-sample effect size (
).
is the one-sided paired t-test p-value (A-RAG > no-RAG).
is the one-sided exact Wilcoxon signed-rank p-value. The sign test reports improved classes out of 10; for 10/10 improvements, the exact one-sided binomial probability is
. FDR-adjusted q-values are shown for the Wilcoxon tests across models.
| Model | Mean (% points) |
SD( ) (% points) |
Cohen’s d |
(one-sided) |
(one-sided, exact) |
Sign test (improved / 10) | q (FDR) |
|---|---|---|---|---|---|---|---|
| DeepSeek-V3 | 3.58 | 1.1942 | 2.9977 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
| DeepSeek-R1-Qwen | 3.56 | 0.5060 | 7.0361 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
| LLaMA-3.3-70B | 3.26 | 1.2186 | 2.6753 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
| DeepSeek-R1-LLaMA | 3.58 | 0.6763 | 5.2938 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
| LLaMA-4-Maverick | 4.38 | 0.9864 | 4.4406 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
| LLaMA-3.1-8B | 4.72 | 1.6818 | 2.8065 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
| Qwen2.5-7B | 5.88 | 1.5091 | 3.8964 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
| Gemma-2-27B | 9.02 | 1.7370 | 5.1927 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
| Gemma-2-9B | 9.62 | 2.4666 | 3.9001 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
| Mistral-7B | 7.34 | 3.2250 | 2.2760 | ![]() |
0.00097656 | 10/10 | 0.00097656 |
Deployment considerations
The entire DermaGPT pipeline was deployed using a lightweight Python backend, with the vision model running on an NVIDIA T4 GPU and the LLMs accessed via the Together API33. The system supports sub-second vision-stage inference, on-device image preprocessing, and offline-capable A-RAG modules, making it suitable for low-resource clinical settings. Deployment and experimentation were conducted on virtual environments equipped with NVIDIA T4 GPUs and 16 GB of VRAM, running Ubuntu-based containers. This configuration mirrors low-cost cloud environments and facilitates reproducibility. All MLTF computations and uncertainty estimations were performed locally on participating nodes, ensuring privacy compliance and enabling real-time reliability-aware model updates without external data sharing. A simple user-facing interface that follows this pipeline is illustrated in Fig. 9.
Fig. 9.
A view of the DermaGPT application. A patient-initiated diagnostic workflow: users upload a skin lesion image and pose a natural-language query (e.g., “What is my skin condition and how can I treat it?”). The PaLI-Gemma 2 model performs image-based diagnosis, while LLMs provide treatment advice and explanations. Shown with and without augmentation via A-RAG.
Privacy and external LLM access.: In our deployment, the vision backbone (PaLI-Gemma 2) runs locally and no images are transmitted to any external service. When an external LLM is used via the Together API, the request payload contains text only: (i) a short structured diagnostic summary (predicted lesion type and benign/malignant status) and (ii) the user’s question. For example, we send a JSON message of the form {model, messages, max_tokens}, where messages includes a system message such as “Diagnosis: Basal cell carcinoma; Condition: malignant” and a user query (e.g., “What is my skin condition and how can I treat it?”).
Although images are not shared, this textual information may still be considered sensitive health data. Therefore, privacy-related claims (e.g., “local” or “offline-capable”) apply to the diagnostic stage and to configurations where the explanatory LLM is deployed locally; when using a third-party API, privacy depends on the provider’s data handling and the deployment’s compliance requirements.
Computational efficiency and cost
To allow readers to balance performance against resource use, we quantified the computational footprint of DermaGPT under a single, consistent deployment setup. All measurements for this section were obtained on a virtual machine with a single NVIDIA T4 GPU (16 GB VRAM), using Python 3.10, PyTorch 2.7.1, CUDA 12.4 and batch size 1. Input images for the visual module were resized to
pixels, and language-model interactions were limited to short patient-style prompts and replies (typically a few hundred tokens in total per consultation).
Trainable and active parameters
The PaLI-Gemma 2 vision–language backbone used for image-based diagnosis has approximately 2.95 billion parameters. During fine-tuning, we freeze the backbone and train only LoRA adapters injected into attention and embedding layers, resulting in
30 million trainable parameters (
1% of the total). At inference time, all 2.95 billion vision-module parameters are active. For the explanatory stage, we use DeepSeek-V3 as an off-the-shelf mixture-of-experts language model: although the full model contains 671 billion parameters, only about 37 billion expert parameters are active per token during inference. DeepSeek-V3 is not fine-tuned in our pipeline, so the number of trainable parameters for the language module is zero.
FLOPs and latency
Given the model size and input resolution, a single forward pass of the PaLI-Gemma 2 vision module requires on the order of
–
floating-point operations (FLOPs) per image. A complete consultation—consisting of one image, a short prompt and a brief explanatory reply—therefore involves on the order of
–
FLOPs across the vision and language stages combined. Empirically, under the T4 configuration described above, we observed vision-stage inference latencies in the range of 0.1–0.5 s per image. We did not benchmark end-to-end consultation latency because it depends on external LLM provider/API and network conditions.
Cost per consultation
We estimated the monetary cost per consultation using public DeepSeek-V3 pricing (US$1.25 per million tokens). For a typical interaction of at most a few hundred tokens (system prompt, image-derived context and generated explanation), this corresponds to a language-model cost on the order of
US$ per consultation. The GPU cost of the PaLI-Gemma 2 vision module at the observed latencies on a T4-class instance adds only a negligible fraction. Overall, the end-to-end computational and monetary cost per use is several orders of magnitude lower than the cost of a specialist visit.
For convenience, Table 15 summarizes the key efficiency metrics for this deployment setting.
Table 15.
Summary of efficiency metrics for DermaGPT under a consistent deployment setup (single NVIDIA T4 GPU, batch size 1,
pixel inputs).
| Component | Trainable params | Active params | Notes |
|---|---|---|---|
| PaLI-Gemma 2 (vision) |
30M |
2.95B |
LoRA adapters only trainable |
| DeepSeek-V3 (LLM) | 0 |
37B |
Mixture-of-experts, 671B total |
Approx. FLOPs per consultation (vision + language): – (order-of-magnitude). | |||
Vision-stage latency (T4, batch 1): 0.1–0.5 s per image (model inference only).
| |||
| End-to-end latency depends on external LLM provider/API and network conditions; not benchmarked here. | |||
Language-model cost: US$ per consultation (DeepSeek-V3; token-dependent). | |||
Conclusion
DermaGPT is a federated multimodal artificial intelligence framework for dermatology decision support, designed to remain trustworthy under heterogeneous, privacy-sensitive data. The system integrates computer vision and natural language understanding within a unified architecture that prioritizes interpretability, scalability, and calibration. It comprises two synergistic modules: a visual diagnostic module based on the PaLI-Gemma 2 model—achieving 90.2% accuracy in disease classification and 93.3% in severity estimation—and a language interpretation module powered by DeepSeek-V3 combined with an Advanced Retrieval-Augmented Generation (A-RAG) mechanism, providing clinically grounded, patient-friendly explanations.
In our comparative evaluation, DeepSeek-V3 yielded the highest dermatologist-rated explanation quality under the studied prompts and A-RAG configuration. Beyond accuracy and interpretability, DermaGPT incorporates a Meta-Learned Trust Function (MLTF), a meta-optimization layer that dynamically reweights client updates during federated learning according to reliability indicators such as uncertainty, calibration error, and domain shift. This trust-aware aggregation improved robustness, calibration, and fairness across heterogeneous clinical sites, including the worst-performing node.
The framework operates efficiently on cost-effective GPUs (e.g., NVIDIA T4), with vision-stage inference latency of approximately 0.1–0.5 s per image (PaLI-Gemma 2; model inference only), while end-to-end latency depends on external LLM provider/API and network conditions. Privacy is supported by federated training of the diagnostic backbone without sharing raw images; however, when an external LLM is used, text-only summaries (predicted lesion type and benign/malignant status) together with the user query may still constitute sensitive health data. Clinical evaluations indicated that the combination of MLTF and A-RAG reduced hallucinated or unsupported statements and improved the clarity and clinical relevance of generated explanations.
DermaGPT’s modular and privacy-preserving design offers a foundation for equitable AI-assisted dermatology, particularly in resource-limited settings. Future work will extend coverage to rare conditions, add multilingual and self-monitoring capabilities, and explore prospective real-world deployment. In addition, a natural extension is a systematic comparison between PaLI-Gemma 2 and alternative unified vision–language image-to-text backbones (for example, BLIP, LLaVA, or Med-Flamingo) for both lesion classification and severity estimation, as well as their impact on explanation relevance, factuality, and clinical grounding. Overall, the results demonstrate that a trust-aware federated multimodal architecture can provide robust, interpretable support for dermatological decision-making, while leaving final judgment with the clinician.
Supplementary Information
Acknowledgements
The proposal underlying this work was competitively selected by the event’s review committee for an oral presentation at the 2nd Artificial Intelligence National Event of Iran (17–19 December 2025) and received an award at the event.
Author contributions
N.M.H. conceived and designed the study, prepared the dataset, and drafted the initial manuscript. M.H.A. led the system implementation, software development, federated learning integration, formal analysis, and manuscript writing and revision. M.K.N. contributed to validation, interpretation of results, and critical review of the manuscript. All authors reviewed and approved the final version of the manuscript and agree to be accountable for all aspects of the work.
Data availability
Institutional data, where used, remained on-site and were processed locally within a federated setup. No identifiable patient-level data were transferred or centralized, and the authors did not access raw patient images.
Code availability
The code is available from the corresponding author upon reasonable request.
Declarations
Competing interests
The authors declare no competing interests.
Ethics approval and consent to participate
This study was reviewed and approved by the Research Ethics Committee of Shahroud University of Medical Sciences (IR.SHMU.REC.1404.159). We performed retrospective AI/ML analyses on de-identified dermatology images at participating institutions and evaluated performance on a publicly available de-identified dataset. No patients were recruited or contacted; individual consent was not required because only de-identified data were used. The authors did not access, transfer, or centrally store identifiable patient information. In the federated setting, only aggregated model updates and summary metrics were exchanged. Where expert review was conducted, dermatologists assessed de-identified class-level text examples and did not review patient-level cases or identifiable information. All procedures complied with applicable guidelines and regulations. No sensitive personal data were collected from evaluators beyond their rubric scores.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Nastaran Mehrabi Hashjin and Mohammad Hussein Amiri.
Supplementary Information
The online version contains supplementary material available at 10.1038/s41598-026-38715-0.
References
- 1.Liu, Y. et al. A deep learning system for differential diagnosis of skin diseases. Nat. Med.26 (6), 900–908. 10.1038/s41591-020-0842-3 (2020). [DOI] [PubMed] [Google Scholar]
- 2.Hay, R. J. et al. The Global Burden of Skin Disease in 2010: An Analysis of the Prevalence and Impact of Skin Conditions. J. Investig. Dermatol.134 (6), 1527–1534. 10.1038/jid.2013.446 (2014). [DOI] [PubMed] [Google Scholar]
- 3.Tschandl, P. et al. Comparison of the accuracy of human readers versus machine-learning algorithms for pigmented skin lesion classification: an open, web-based, international, diagnostic study. Lancet Oncol.20 (7), 938–947. 10.1016/S1470-2045(19)30333-X (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Keim, U. et al. Incidence, mortality and trends of cutaneous squamous cell carcinoma in Germany, the Netherlands, and Scotland. Eur. J. Cancer183, 60–68. 10.1016/j.ejca.2023.01.017 (2023). [DOI] [PubMed] [Google Scholar]
- 5.Fekri, M. et al. Epidemiology and socioeconomic factors of nonmelanoma skin cancer in the Middle East and North Africa 1990 to 2021. Sci. Rep.15 (1), 17904. 10.1038/s41598-025-99434-6 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Afolaranmi, O. et al. Cancer Research Funding in Africa. Commun. Med.5 (1), 278. 10.1038/s43856-025-00992-7 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Kingham, T. P. et al. Treatment of cancer in sub-Saharan Africa. Lancet Oncol.14 (4), e158–e167. 10.1016/S1470-2045(12)70472-2 (2013). [DOI] [PubMed] [Google Scholar]
- 8.Liu, W. et al. DrBioRight 2.0: an LLM-powered bioinformatics chatbot for large-scale cancer functional proteomics analysis. Nat. Commun.16 (1), 2256. 10.1038/s41467-025-57430-4 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Rao, V. M. et al. Multimodal generative AI for medical image interpretation. Nature639 (8056), 888–896. 10.1038/s41586-025-08675-y (2025). [DOI] [PubMed] [Google Scholar]
- 10.Esteva, A. et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature542 (7639), 115–118. 10.1038/nature21056 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Zhou, J. et al. Pre-trained multimodal large language model enhances dermatological diagnosis using SkinGPT-4. Nat. Commun.15, 5649. 10.1038/s41467-024-50043-3 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Li, J. et al. Integrated image-based deep learning and language models for primary diabetes care. Nat. Med.30 (10), 2886–2896. 10.1038/s41591-024-03139-8 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Dick, V., Sinz, C., Mittlböck, M., Kittler, H. & Tschandl, P. Accuracy of computer-aided diagnosis of melanoma: a meta-analysis. JAMA Dermatol.155 (11), 1291–1299. 10.1001/jamadermatol.2019.1375 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Chen, C. et al. Integration of large language models and federated learning. Patterns5 (12), 101098. 10.1016/j.patter.2024.101098 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Wan, P. et al. Outpatient reception via collaboration between nurses and a large language model: a randomized controlled trial. Nat. Med.30 (10), 2878–2885. 10.1038/s41591-024-03148-7 (2024). [DOI] [PubMed] [Google Scholar]
- 16.Lin, Q. et al. Has multimodal learning delivered universal intelligence in healthcare? A comprehensive survey. Informat. Fusion116, 102795. 10.1016/j.inffus.2024.102795 (2025). [Google Scholar]
- 17.Xiao, H. et al. A comprehensive survey of large language models and multimodal large language models in medicine. Informat. Fusion117, 102888. 10.1016/j.inffus.2024.102888 (2025). [Google Scholar]
- 18.Zhou, L. et al. Larger and more instructable language models become less reliable. Nature634 (8032), 61–68. 10.1038/s41586-024-07930-y (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Chen, R., Fettel, K. D., Nguyen, D. H., & Nambudiri, V. E. Evaluating the performance of ChatGPT on dermatology board-style exams: A meta-analysis of text-based and image-based question accuracy. J. Am. Acad. Dermatol.10.1016/j.jaad.2025.04.004 (2025). [DOI] [PubMed] [Google Scholar]
- 20.Zhang, S., Dai, G., Huang, T. & Chen, J. Multimodal large language models for bioimage analysis. Nat. Methods21(8), 1390–1393. 10.1038/s41592-024-02334-2 (2024). [DOI] [PubMed] [Google Scholar]
- 21.McDuff, D. et al., Towards accurate differential diagnosis with large language models. Nature10.1038/s41586-025-08869-4 (2025). [DOI] [PMC free article] [PubMed]
- 22.Beyer, L. et al., PaliGemma: A versatile 3B VLM for transfer. (accessed 24 July 2025); https://arxiv.org/abs/2407.07726v2.
- 23.Hu, E. et al., LoRA: Low-rank adaptation of large language models. ICLR 2022 - 10th International Conference on Learning Representations. (accessed 24 July 2025); https://arxiv.org/abs/2106.09685v2.
- 24.Yip, W. Improving primary healthcare with generative AI. Nat. Med.30 (10), 2727–2728. 10.1038/s41591-024-03257-3 (2024). [DOI] [PubMed] [Google Scholar]
- 25.Singhal, K. et al. Large language models encode clinical knowledge. Nature620 (7972), 172–180. 10.1038/s41586-023-06291-2 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Tordjman, M. et al., Comparative benchmarking of the DeepSeek large language model on medical tasks and clinical reasoning. Nat. Med.10.1038/s41591-025-03726-3 (2025). [DOI] [PubMed]
- 27.Johri, S. et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nat. Med.31 (1), 77–86. 10.1038/s41591-024-03328-5 (2025). [DOI] [PubMed] [Google Scholar]
- 28.Goh, E. et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat. Med.31 (4), 1233–1238. 10.1038/s41591-024-03456-y (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Sandmann, S. et al. Systematic analysis of ChatGPT, Google search and Llama 2 for clinical decision support tasks. Nat. Commun.15 (1), 2050. 10.1038/s41467-024-46411-8 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Wu, Y. et al. A concept-based interpretable model for the diagnosis of choroid neoplasias using multimodal data. Nat. Commun.16 (1), 3504. 10.1038/s41467-025-58801-7 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Hayat, M., Aramvith, S., Bhattacharjee, S., & Ahmad, N. Attention GhostUNet++: Enhanced Segmentation of Adipose Tissue and Liver in CT Images. In Proc. IEEE EMBC10.1109/EMBC58623.2025.11254588. (2025). [DOI] [PubMed]
- 32.Hayat, M. Endoscopic image super-resolution algorithm using edge and disparity awareness. Chulalongkorn University Theses and Dissertations (Chula ETD)10.58837/CHULA.THE.2023.901. (2023).
- 33.Together AI – The AI acceleration cloud - fast inference, fine-tuning & training. (accessed 24 July 2025); https://www.together.ai/.
- 34.Ye, R. et al., OpenFedLLM: Training large language models on decentralized private data via federated learning. In Proc. ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., pp. 6137–6147 10.1145/3637528.3671582 (2024).
- 35.Kittler, H. et al. Diagnostic accuracy of dermoscopy. Lancet Oncol.3 (3), 159–165. 10.1016/S1470-2045(02)00679-4 (2002). [DOI] [PubMed] [Google Scholar]
- 36.Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med.29 (8), 1930–1940. 10.1038/s41591-023-02448-8 (2023). [DOI] [PubMed] [Google Scholar]
- 37.Argenziano, G. & Soyer, H. P. Dermoscopy of pigmented skin lesions - a valuable tool for early. Lancet Oncol.2 (7), 443–449. 10.1016/S1470-2045(00)00422-8 (2001). [DOI] [PubMed] [Google Scholar]
- 38.Johri, S. et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nat. Med.31 (1), 77–86. 10.1038/s41591-024-03328-5 (2025). [DOI] [PubMed] [Google Scholar]
- 39.Caffery, L. et al., Informing a position statement on the use of artificial intelligence in dermatology in Australia, Australas. J. Dermatol.64, 10.1111/ajd.13946 (2022). [DOI] [PubMed]
- 40.Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., & Smith, V. Federated optimization in heterogeneous networks. arXivhttps://arxiv.org/abs/1812.06127 (2020).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Institutional data, where used, remained on-site and were processed locally within a federated setup. No identifiable patient-level data were transferred or centralized, and the authors did not access raw patient images.
The code is available from the corresponding author upon reasonable request.

























































