Skip to main content
Proceedings (Baylor University. Medical Center) logoLink to Proceedings (Baylor University. Medical Center)
. 2026 Jun 10;39(5):837–844. doi: 10.1080/08998280.2026.2683945

Evaluating the quality of artificial intelligence responses to psoriasis-related clinical and patient questions: a comparative study of ChatGPT, Gemini, and Microsoft Copilot

Gozde Ulutaş Demirbas a, Esin Diremsizoglu b,✉, Abdullah Demirbas b
PMCID: PMC13523945  PMID: 42267786

Abstract

Background

Patient use of artificial intelligence (AI) chatbots for dermatologic information is increasing, but their performance on psoriasis-related questions across clinically distinct domains remains unclear. We compared ChatGPT (GPT-5.3 Instant), Gemini (Gemini 3 Flash), and Microsoft Copilot using a multidimensional scoring framework.

Methods

Fifty-four psoriasis-related questions were submitted to each model across diagnostic (n = 12), treatment (n = 12), and patient-question (n = 30) categories. Three board-certified dermatologists independently scored responses for accuracy, evidence consistency, completeness, and clinical safety (maximum score, 8).

Results

Interrater agreement was substantial to almost perfect (κ = 0.743 for ChatGPT, 0.830 for Gemini, and 0.844 for Copilot). Overall mean scores differed significantly: 7.36 ± 0.97, 7.77 ± 0.77, and 7.05 ± 1.09, respectively (Friedman P < 0.001). No difference was observed for diagnostic questions (P = 0.37). Gemini outperformed both models in treatment (8.00 ± 0.00) and patient questions (P < 0.001), while ChatGPT and Copilot did not differ in treatment. Differences were driven by completeness and accuracy, not clinical safety or evidence consistency. Gemini also had the highest rate of high-reliability responses (92.6%).

Conclusions

All models showed high clinical safety, but Gemini provided the most complete and highest-quality responses. The observation that models may provide accurate yet clinically incomplete responses, particularly for treatment content, emphasizes the need for physician oversight when AI-generated information is used in dermatological practice.

Keywords: Artificial intelligence, ChatGPT, clinical decision support, dermatology, Gemini, large language models, Microsoft Copilot, patient education, psoriasis


Psoriasis is a chronic, immune-mediated inflammatory skin disease affecting approximately 2% to 3% of the global population, with a disproportionate impact on quality of life, psychological well-being, and cardiometabolic comorbidities.1,2 Despite the availability of highly effective therapies, real-world disease management remains suboptimal: patients frequently experience diagnostic delays, limited access to dermatological expertise, and insufficient health literacy regarding their disease and treatment options.3,4 These challenges have positioned artificial intelligence (AI)-based tools as increasingly consulted sources of medical information for both patients and clinicians.

ChatGPT, Gemini, and Microsoft Copilot have emerged as the most widely used large language model (LLM)-based AI platforms through which patients and healthcare professionals seek medical information.5,6 Contemporary LLMs have demonstrated the ability to pass standardized medical licensing examinations7 and specialty-level dermatology assessments,8 reflecting a capacity for clinical knowledge synthesis that extends beyond simple information retrieval. Comparative evaluations across multiple medical domains have consistently shown intermodel performance differences that vary by question type and clinical context,9,10 underscoring the importance of specialty-specific, multiplatform benchmarking. In dermatology, ChatGPT responses to common patient queries have shown notable inconsistencies for complex or guideline-dependent topics including psoriasis management,11 and a recent rapid review highlighted persistent concerns about accuracy and completeness across dermatological applications.12

Psoriasis-specific evaluation of LLMs remains limited. Available studies have assessed single platforms in isolation, evaluating ChatGPT responses to patient concerns13,14 or comparing its outputs against published meta-analytic estimates for systemic therapies,15 and none has incorporated clinical safety, evidence consistency, accuracy, and completeness as independently scored domains. No study has simultaneously compared multiple leading platforms across both clinician-oriented and patient-oriented psoriasis question types.

To address this gap, we compared the response quality of ChatGPT, Gemini, and Microsoft Copilot across 54 psoriasis-related questions spanning diagnostic, treatment, and patient education domains, independently rated by three board-certified dermatologists. Our aim was to provide criterion-level evidence on the relative strengths and limitations of these platforms, informing their responsible use as adjuncts to psoriasis information delivery in clinical and patient-facing settings.

METHODS

This was a cross-sectional comparative study evaluating the quality of AI-generated responses to psoriasis-related questions, following published recommendations for evaluating LLM performance in clinical settings.16 As the study did not involve human participants or patient data, formal ethics review was not required.

AI models

All models were accessed via their standard free-tier web interfaces in March 2026. The versions available at that time were ChatGPT (GPT-5.3 Instant, OpenAI), Google Gemini (Gemini 3 Flash, Google DeepMind), and Microsoft Copilot (Microsoft Corporation; accessed via the free-tier consumer web interface; the precise underlying model version was not independently verifiable at the time of access). Identical prompts were submitted to each model without prior context loading or iterative refinement, reflecting typical end-user behavior.

Question set development

Fifty-four questions were developed through a structured consensus process by a panel of board-certified dermatologists, informed by clinical experience and a review of common psoriasis-related inquiries encountered in practice. Questions were classified into three domains: (1) diagnostic questions (n = 12), covering clinical features, differential diagnosis, biopsy indications, comorbidity screening, and severity scoring; (2) treatment questions (n = 12), addressing topical therapies, phototherapy, conventional systemic agents, biologics, and pregnancy considerations; and (3) patient questions (n = 30), reflecting common real-world inquiries. Patient questions were further divided into four subcategories: disease knowledge (n = 8), triggering factors (n = 7), daily life (n = 7), and treatment/prognosis (n = 8) (Supplementary Table 1).

Evaluation

Each AI response was independently evaluated by three dermatologists, who were blinded to each other’s assessments, using a four-criterion scoring instrument adapted from previously validated frameworks16,17: (1) Accuracy: scientific correctness (0 = incorrect; 1 = partially correct; 2 = fully correct and current); (2) Evidence Consistency: alignment with American Academy of Dermatology (AAD) and European Academy of Dermatology and Venereology (EADV) guidelines (0 = inconsistent; 1 = partially consistent; 2 = fully consistent); (3) Completeness: coverage of clinically important aspects (0 = grossly incomplete; 1 = partial; 2 = comprehensive); (4) Clinical Safety: absence of potentially harmful recommendations (0 = potentially harmful; 1 = minor risk; 2 = safe). The maximum total score per response was 8 points, calculated as the mean of the three raters’ independent scores. The evaluation dimensions were drawn from the reporting framework of Huo et al16; the scoring structure, reliability classification thresholds, and precedent for guideline alignment as a discrete criterion were adopted from Walker et al.17 Evidence Consistency was operationalized against the AAD-NPF guidelines for biologic therapy18 and the EuroGuiDerm/EADV guidelines for systemic psoriasis treatment,19 which served as the primary reference standards. Scoring anchors were constructed by the study panel. Scores were classified as high reliability (7–8), moderate (5–6), low (3–4), or unreliable (0–2).17

Statistical analysis

Interrater reliability was assessed using Fleiss’ kappa (κ) for each evaluation criterion independently, interpreted as follows: <0.21 slight; 0.21–0.40 fair; 0.41–0.60 moderate; 0.61–0.80 substantial; >0.80 almost perfect.20 The Shapiro-Wilk test was performed to assess score distributions. The Friedman test was used for three-way comparisons within each question category; significant results were followed by Wilcoxon signed-rank post hoc tests with Bonferroni correction (adjusted α = 0.02). The same approach was applied at the criterion level. All analyses were performed using Python 3.12 (SciPy 1.13). Statistical significance was set at P < 0.05.

RESULTS

Interrater reliability

Interrater agreement was substantial to almost perfect across all model–criterion combinations (Supplementary Table 2). Mean κ values were 0.743 for ChatGPT, 0.830 for Gemini, and 0.844 for Copilot. The lowest agreement was observed for ChatGPT’s Evidence Consistency rating (κ = 0.474, moderate). Almost perfect agreement (κ = 1.000) was achieved for the Safety criterion across all three models.

Overall performance and category-level differences

Mean total quality scores were 7.36 ± 0.97 (median 8.00) for ChatGPT, 7.77 ± 0.77 (median 8.00) for Gemini, and 7.05 ± 1.09 (median 7.33) for Copilot, with a statistically significant difference across models (Friedman P < 0.001) (Figure 1, Table 1, Supplementary Table 3). All three pairwise comparisons reached significance after Bonferroni correction: ChatGPT vs. Gemini (P = 0.002), ChatGPT vs. Copilot (P = 0.005), and Gemini vs. Copilot (P < 0.001).

Figure 1.

Bar chart comparing mean total quality scores for ChatGPT, Gemini, and MS Copilot across categories: Diagnostic, Treatment, Patient questions, Overall. Grouped bar chart showing mean total quality scores (max. 8) for ChatGPT, Gemini, and MS Copilot across Diagnostic (ChatGPT 7.50, Gemini 7.44, MS Copilot 7.25), Treatment (ChatGPT 6.92, Gemini 8.00, MS Copilot 6.03), Patient questions (ChatGPT 7.48, Gemini 7.81, MS Copilot 7.38), and Overall (ChatGPT 7.36, Gemini 7.77, MS Copilot 7.05). Gemini shows consistently higher scores in Treatment, Patient questions, and Overall, while MS Copilot has the lowest scores.

Mean total quality scores by question category. Bars represent mean total scores (maximum = 8) averaged across three raters.

Table 1.

Mean total quality scores, Friedman test results, and Bonferroni-corrected post hoc pairwise comparisons by question categorya

Category n Mean ± SD (median)
Friedman P ChatGPT vs Gemini ChatGPT vs Copilot Gemini vs Copilot
ChatGPT Gemini MS Copilot
Diagnostic questions 12 7.50 ± 1.02 (8.00) 7.44 ± 1.02 (8.00) 7.25 ± 1.07 (8.00) 0.368 — — —
Treatment questions 12 6.92 ± 0.96 (6.67) 8.00 ± 0.00 (8.00)d 6.03 ± 0.90 (6.00) <0.001c 0.016 b 0.078 ns 0.001b
Patient questions (total) 30 7.48 ± 0.89 (8.00) 7.81 ± 0.77 (8.00) 7.38 ± 0.91 (8.00) <0.001c 0.010b 0.102 ns 0.006b
Disease knowledge 8 7.92 ± 0.22 (8.00) 8.00 ± 0.00 (8.00) 7.92 ± 0.22 (8.00) 0.368 — — —
Triggering factors 7 7.57 ± 0.43 (7.67) 8.00 ± 0.00 (8.00) 7.57 ± 0.43 (7.67) 0.018* 0.125 nse 1.000 nse 0.125 nse
Daily life 7 7.52 ± 0.71 (8.00) 8.00 ± 0.00 (8.00) 7.29 ± 0.72 (7.33) 0.037* 0.250 nse 0.500 nse 0.125 nse
Treatment and prognosis 8 6.92 ± 1.34 (7.50) 7.29 ± 1.36 (8.00) 6.75 ± 1.31 (7.00) 0.526 — — —
Overall 54 7.36 ± 0.97 (8.00) 7.77 ± 0.77 (8.00) 7.05 ± 1.09 (7.33) <0.001c 0.002b 0.005b <0.001b

aMaximum score = 8; n = 54 questions. Values are mean ± standard deviation (SD) (median) based on the average of three independent rater scores. Indented rows represent patient question subcategories. Post hoc comparisons (Wilcoxon signed-rank, Bonferroni-corrected α = 0.02) were performed only when the Friedman test was significant. —indicates post hoc not applicable; ns, not significant after Bonferroni correction.

bP < 0.02 (Bonferroni-corrected).

cP < 0.001.

dGemini achieved a uniform score of 8.00 across all 12 treatment questions (zero variance).

eFriedman test significant but no pairwise comparison reached significance after Bonferroni correction, likely due to limited statistical power (n = 7).

No significant difference was found for diagnostic questions (P = 0.37). For treatment questions (P < 0.001), Gemini achieved a perfect mean of 8.00 ± 0.00 (median 8.00) across all 12 questions, significantly outperforming ChatGPT (6.92 ± 0.96, median 6.67; post hoc P = 0.02) and Copilot (6.03 ± 0.90, median 6.00; post hoc P = 0.001); the difference between ChatGPT and Copilot did not reach significance (P = 0.08). For patient questions (P < 0.001), Gemini achieved the highest scores (7.81 ± 0.77, median 8.00) compared to ChatGPT (7.48 ± 0.89, median 8.00; post hoc P = 0.01) and Copilot (7.38 ± 0.91, median 8.00; post hoc P = 0.006); the ChatGPT vs. Copilot comparison was not significant (P = 0.10).

Patient question subcategory analysis

No significant intermodel differences were found for disease knowledge (P = 0.37) or treatment/prognosis (P = 0.53) (Table 1). Friedman tests were significant for triggering factors (P = 0.02) and daily life (P = 0.04); however, no individual pairwise comparison reached significance after Bonferroni correction in either subcategory (all post hoc P ≥ 0.13). The treatment/prognosis subcategory exhibited the highest score variability across all models (SD: 1.31–1.36).

Criterion-level analysis

Significant intermodel differences were found for Completeness (P < 0.001) and Accuracy (P < 0.001) across all categories except diagnostics (Table 2, Figure 2). The largest completeness gap was observed in treatment questions, where Gemini achieved 2.00/2 versus Copilot’s 1.11/2 (P < 0.001). No significant differences were found for Evidence Consistency (P = 0.17) or Clinical Safety (P = 0.61) in any category or subcategory, with all three models averaging ≥1.91/2 for both criteria.

Table 2.

Criterion-level Friedman test results comparing AI models on each evaluation dimension (0–2 points per criterion)

Criterion Category P ChatGPT mean Gemini mean Copilot mean
Accuracy Overall (n = 54) <0.001 1.772 1.932 1.654
Diagnostic (n = 12) 0.37 1.806 1.778 1.667
Treatment (n = 12) <0.001 1.528 2.000 1.194
Patient (n = 30) 0.01 1.856 1.967 1.833
Evidence Consistency Overall (n = 54) 0.17 1.951 1.969 1.926
Diagnostic (n = 12) — 1.944 1.944 1.944
Treatment (n = 12) 0.17 1.917 2.000 1.806
Patient (n = 30) — 1.967 1.967 1.967
Completeness Overall (n = 54) <0.001 1.667 1.926 1.525
Diagnostic (n = 12) 0.37 1.833 1.806 1.722
Treatment (n = 12) <0.001 1.472 2.000 1.111
Patient (n = 30) <0.001 1.678 1.944 1.611
Safety Overall (n = 54) 0.61 1.963 1.944 1.944
Diagnostic (n = 12) — 1.917 1.917 1.917
Treatment (n = 12) 0.37 2.000 2.000 1.917
Patient (n = 30) 0.37 1.967 1.933 1.967

Scores represent the mean of three independent raters. Bold P values denote statistical significance at P < 0.05 (Friedman test). Post hoc pairwise comparisons (Bonferroni-corrected α = 0.0167) were performed when Friedman test was significant. —indicates not applicable, Friedman χ2 could not be computed because all three AI models received identical scores (zero variance).

Figure 2.

Four radar charts compare ChatGPT, Gemini, and MS Copilot on Accuracy, Evidence consistency, Completeness, and Safety, with values ranging from 0 to 2. The figure features four radial charts arranged in a 2x2 grid, comparing performance metrics across questions (Q1 to Q49) for Accuracy, Evidence consistency, Completeness, and Safety. Each chart has a scale from 0 (center) to 2 (outer rim) and displays dark grey lines for ChatGPT, medium grey for Gemini, and light grey for MS Copilot. Key highlights include high values for ChatGPT at Q1 across all categories, with lower values seen at Q31, Q37, and Q43. Safety metrics show higher performance nearing the outer rim compared to other metrics.

Criterion-level performance profiles of three AI models across all 54 questions. Radar plots display mean scores (0–2) for (a) Accuracy, (b) Evidence Consistency, (c) Completeness, and (d) Safety per question. Q numbers correspond to question indices (D1–D12, T1–T12, P1–P30).

Reliability classification

Gemini achieved the highest proportion of high-reliability responses (92.6%), followed by ChatGPT (75.9%) and Copilot (63.0%) (Figure 3). In the treatment category, all 12 Gemini responses (100%) were classified as high reliability, versus 41.7% for ChatGPT and 8.3% for Copilot. No response was classified as unreliable across any model, and all three models achieved 100% high-reliability ratings for the disease knowledge and triggering factors subcategories.

Figure 3.

Bar chart comparing reliability percentages for ChatGPT, Gemini, and MS Copilot. A horizontal stacked bar chart shows reliability percentages for three tools: ChatGPT (75.9% high, 20.4% moderate), Gemini (92.6% high), and MS Copilot (63.0% high, 31.5% moderate). Gemini has the highest high reliability, while MS Copilot has the lowest high reliability and highest moderate reliability.

Reliability classification of AI responses by model. Stacked bars show the proportion of high (7–8), moderate (5–6), and low (3–4) reliability responses (n = 54 per model). No response was classified as unreliable.

DISCUSSION

This study systematically compared ChatGPT, Gemini, and Copilot across 54 psoriasis-related queries spanning diagnostic, treatment, and patient education domains. Gemini demonstrated superior overall performance, achieving the highest mean quality score and the largest proportion of high-reliability responses (92.6%), followed by ChatGPT (75.9%) and Copilot (63.0%). Intermodel differences were confined to completeness and accuracy; no significant variation was observed for clinical safety or evidence consistency across any question category. These findings suggest that while contemporary LLMs have converged on safety-compliant psoriasis responses, platform choice remains clinically meaningful, particularly in the treatment domain.

Performance did not, however, diverge uniformly across question categories. Diagnostic questions showed no intermodel variation, consistent with the literature. Established diagnostic criteria such as the Auspitz sign, Koebner phenomenon, and PASI scoring are widely covered in the medical literature and well represented in LLM training datasets, and comparable ceiling-level performance has been reported for diagnostic queries in other dermatological conditions.21,22 The same pattern extended to the disease knowledge subcategory of patient questions, where all three models achieved perfect high-reliability ratings. Across these question types, platform choice appears to have little bearing on response quality. It should be noted that all subcategory-level comparisons involve small sample sizes (n = 7–8), and significant Friedman test results for triggering factors and daily life should be interpreted with caution given the elevated risk of type I error at these group sizes; no pairwise comparison reached significance after Bonferroni correction in either subcategory.

Treatment questions produced the most pronounced intermodel divergence. Gemini achieved a mean score of 8.00/8 across all 12 treatment queries, with every response classified as high reliability. In contrast, only 41.7% of ChatGPT and 8.3% of Copilot responses met this threshold, with both platforms frequently omitting biologic treatment sequencing, contraindication hierarchies, and monitoring requirements despite maintaining high accuracy and safety ratings. This pattern is consistent with dermatology-specific evaluations: accuracy fell to 56% for acne treatment queries23; treatment reliability was the weakest category in a rosacea-focused assessment21; and a vitiligo-specific evaluation found perfect accuracy on descriptive questions but notable inconsistency on treatment recommendations for special populations.22 These findings suggest a structural limitation of current LLMs in treatment-domain content, likely arising from two converging factors: biologic guidelines undergo frequent revision and specificity-requiring content is substantially underrepresented in general training corpora7,24; and models optimized for harm avoidance tend to generate conservative, partial responses when confronted with guideline-dependent queries, a tradeoff documented in medical LLM evaluations where nearly half of responses omitted clinically important content without domain-specific alignment.25 Gemini’s superiority across all treatment items may reflect its web-integrated retrieval architecture, though direct architectural comparisons remain outside the scope of this study. The absence of score variance across 12 treatment questions warrants acknowledgment: while this may reflect ceiling performance, rater consensus bias driven by Gemini’s consistently structured output style cannot be fully excluded. Gemini employs real-time web retrieval integrated into its response generation; Google’s own evaluation of Gemini’s medical capabilities demonstrated that grounding responses in dynamically retrieved web sources reduces uncertainty and improves coverage of guideline-dependent content compared to static training data alone.26 This architectural feature may be particularly consequential in the treatment domain, where biologic sequencing, contraindication hierarchies, and monitoring protocols undergo frequent revision and are therefore susceptible to staleness in models relying exclusively on fixed training corpora.7,24 Whether this represents a durable structural advantage or a version-specific finding cannot be determined from a cross-sectional evaluation; longitudinal benchmarking across model updates will be necessary to establish whether Gemini’s treatment-domain superiority is reproducible and architecturally attributable.

In contrast, all three models performed equivalently on clinical safety and evidence consistency, averaging ≥1.91/2 on both criteria with perfect interrater agreement for safety, consistent with prior dermatological multiplatform evaluations.21,25 This convergence likely reflects the harm-avoidance weighting inherent in contemporary LLM training.7 Critically, the same optimization that suppresses harmful outputs may structurally constrain completeness: in a medical LLM evaluation, content omission rates fell substantially only after domain-specific alignment, suggesting that safety and completeness are jointly shaped by the same training process.24 A patient consulting an AI chatbot for psoriasis treatment may therefore receive content that avoids harm while omitting relevant biologic options or contraindications, with no signal that anything is missing, a distinction that standard safety scoring instruments, including ours, are not designed to detect. The documented decline in medical disclaimer rates across successive model generations further narrows the gap between apparent safety and genuine clinical accountability.27

Interrater agreement in this study was substantially higher than in comparable dermatological chatbot evaluations, where weighted Cohen’s kappa values of 0.44 and 0.40–0.57 have been reported.23,25 This likely reflects the structured, criterion-based scoring instrument used here, which decomposed response quality into four independently rated dimensions rather than relying on holistic judgment. The exception was ChatGPT’s Evidence Consistency criterion (κ = 0.474), which remained at moderate agreement despite the structured approach. Assessing guideline adherence without explicit citations requires interpretive judgment that varies among raters regardless of the instrument used,20 and this finding highlights a methodological challenge that is unlikely to be resolved without standardized citation requirements for LLM-generated medical content.

These findings carry practical implications for dermatologists whose patients increasingly consult AI chatbots between clinical encounters.3,4 The observation that all three models achieved high clinical safety and guideline consistency suggests that AI-generated psoriasis information is unlikely to cause direct harm; however, the substantial completeness gap, particularly in the treatment domain, means that patients relying solely on AI responses may receive accurate but incomplete information, potentially underestimating available therapeutic options or monitoring requirements. Dermatologists should be aware that patients consulting Gemini for treatment-related questions may arrive better informed than those using ChatGPT or Copilot, and that platform choice may therefore influence the quality of shared decision-making conversations. Proactively directing patients toward higher-performing platforms, while emphasizing that AI-generated content requires clinical contextualization, may represent a pragmatic interim strategy until standardized quality assurance frameworks for medical AI are established.16

LIMITATIONS

Several limitations warrant acknowledgment. First, AI models are continuously updated; findings reflect a specific version and timepoint and may not generalize to current or future iterations. Second, the question set was developed by a single center and may not represent the full breadth of real-world clinical queries. Third, our evaluation did not assess readability, response length, empathy, or comprehensibility for patients. Fourth, patient question subcategory sample sizes (n = 7–8) limit statistical power, and results for triggering factors (P = 0.02) and daily life (P = 0.04) should be interpreted with caution given the risk of type I error at these group sizes. Fifth, our scoring instrument did not capture patient-perceived usefulness or health literacy alignment. Sixth, raters were not blinded to model identity, which represents a potential source of performance-expectation bias; however, the high interrater reliability across all models (κ = 0.743–0.844) suggests scoring was predominantly driven by response content rather than model recognition. Finally, the reliability thresholds applied (high: 7–8; moderate: 5–6; low: 3–4; unreliable: 0–2) were carried over from Walker et al17 to facilitate cross-study comparability and have not been independently validated for psoriasis-specific content.

CONCLUSIONS

All three models demonstrated high clinical safety and guideline consistency across psoriasis-related queries, yet significant intermodel differences in completeness, particularly for treatment content, indicate that Gemini’s performance advantage carries clinical relevance. These findings suggest that AI chatbots may serve as useful adjuncts for psoriasis patient education, but response completeness remains a critical limitation that warrants careful attention before routine deployment in clinical settings. Future research should prioritize real-world patient evaluations, longitudinal model tracking, and readability assessment to comprehensively characterize the clinical utility of AI chatbots in dermatology.

Supplementary Material

Supplemental Material

Disclosure statement/Funding

The authors report no funding or conflict of interest. Data are available from the corresponding author upon reasonable request.

References

  • 1.Rendon A, Schäkel K.. Psoriasis pathogenesis and treatment. Int J Mol Sci. 2019;20(6):1475. doi: 10.3390/ijms20061475. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.World Health Organization . Global Report on Psoriasis. Geneva: WHO; 2016. [Google Scholar]
  • 3.Barlow R, Bewley A, Gkini MA.. AI in psoriatic disease: scoping review. JMIR Dermatol. 2024;7:e50451. doi: 10.2196/50451. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Smith P, Johnson CE, Haran K, et al. Advancing psoriasis care through artificial ıntelligence: a comprehensive review. Curr Dermatol Rep. 2024;13(3):141–147. doi: 10.1007/s13671-024-00434-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Thirunavukarasu AJ, Ting DSJ, Elangovan K, et al. Large language models in medicine. Nat Med. 2023;29(8):1930–1940. doi: 10.1038/s41591-023-02448-8. [DOI] [PubMed] [Google Scholar]
  • 6.Huo B, Boyle L, Marfo N, et al. Large language models for chatbot health advice studies: a systematic review. JAMA Netw Open. 2025;8(2):e2457879. doi: 10.1001/jamanetworkopen.2024.57879. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Gilson A, Safranek CW, Huang T, et al. How does ChatGPT perform on the United States Medical Licensing Examination? JMIR Med Educ. 2023;9:e45312. doi: 10.2196/45312. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Lewandowski M, Łukowicz P, Świetlik D, Barańska-Rybak W.. ChatGPT-3.5 and ChatGPT-4 dermatological knowledge level based on the Specialty Certificate Examination. Clin Exp Dermatol. 2024;49(7):686–691. doi: 10.1093/ced/llad255. [DOI] [PubMed] [Google Scholar]
  • 9.Tepe M, Emekli E.. Assessing the responses of large language models (ChatGPT-4, Gemini, and Microsoft Copilot) to frequently asked questions in breast imaging. Cureus. 2024;16(5):e59960. doi: 10.7759/cureus.59960. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Rossettini G, Bargeri S, Cook C, et al. Accuracy of ChatGPT-3.5, ChatGPT-4o, Copilot, Gemini, Claude, and Perplexity in advising on lumbosacral radicular pain against clinical practice guidelines. Front Digit Health. 2025;7:1574287. doi: 10.3389/fdgth.2025.1574287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Ferreira AL, Chu B, Grant-Kels JM, Ogunleye T, Lipoff JB.. Evaluation of ChatGPT dermatology responses to common patient queries. JMIR Dermatol. 2023;6:e49280. doi: 10.2196/49280. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Khamaysi Z, Awwad M, Jiryis B, Bathish N, Shapiro J.. The role of ChatGPT in dermatology diagnostics. Diagnostics (Basel). 2025;15(12):1529. doi: 10.3390/diagnostics15121529. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Huang C, Hong D, Chen X.. ChatGPT in medicine: evaluating psoriasis patient concerns. Skin Res Technol. 2024;30(4):e13680. doi: 10.1111/srt.13680. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Tamer F, Polat M.. Does ChatGPT help patients access reliable and comprehensive information about psoriasis? Proc (Bayl Univ Med Cent). 2025;38(5):658–661. doi: 10.1080/08998280.2025.2518854. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Lam Hoai XL, Simonart T.. Comparing meta-analyses with ChatGPT in the evaluation of the effectiveness and tolerance of systemic therapies in moderate-to-severe plaque psoriasis. J Clin Med. 2023;12(16):5410. doi: 10.3390/jcm12165410. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Huo B, Collins G, Chartash D, et al. Reporting guideline for Chatbot Health Advice studies: the CHART statement. BMC Med. 2025;23(1):447. doi: 10.1186/s12916-025-04274-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Walker HL, Ghani S, Kuemmerli C, et al. Reliability of medical information provided by ChatGPT: assessment against clinical guidelines and patient information quality instrument. J Med Internet Res. 2023;25:e47479. doi: 10.2196/47479. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Menter A, Strober BE, Kaplan DH, et al. Joint AAD-NPF guidelines of care for the management and treatment of psoriasis with biologics. J Am Acad Dermatol. 2019;80(4):1029–1072. doi: 10.1016/j.jaad.2018.11.057. [DOI] [PubMed] [Google Scholar]
  • 19.Nast A, Smith C, Spuls PI, et al. EuroGuiDerm guideline on the systemic treatment of psoriasis vulgaris – Part 1: treatment and monitoring recommendations. J Eur Acad Dermatol Venereol. 2020;34(11):2461–2498. doi: 10.1111/jdv.16915. [DOI] [PubMed] [Google Scholar]
  • 20.Landis JR, Koch GG.. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159–174. doi: 10.2307/2529310. [DOI] [PubMed] [Google Scholar]
  • 21.Yan S, Du D, Liu X, et al. Assessment of the reliability and clinical applicability of ChatGPT’s responses to patients’ common queries about rosacea. Patient Prefer Adherence. 2024;18:249–253. doi: 10.2147/PPA.S444928. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Su J, Yang X, Li X, et al. Evaluating large language models for accuracy and completeness of vitiligo patient education: a comparative analysis. Clin Cosmet Investig Dermatol. 2025;18:2757–2767. doi: 10.2147/CCID.S552979. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Lakdawala N, Channa L, Gronbeck C, et al. Assessing the accuracy and comprehensiveness of ChatGPT in offering clinical guidance for atopic dermatitis and acne vulgaris. JMIR Dermatol. 2023;6:e50409. doi: 10.2196/50409. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172–180. doi: 10.1038/s41586-023-06291-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Olszewski R, Watros K, Mańczak M, Owoc J, Jeziorski K, Brzeziński J.. Assessing the response quality and readability of chatbots in cardiovascular health, oncology, and psoriasis: a comparative study. Int J Med Inform. 2024;190:105562. doi: 10.1016/j.ijmedinf.2024.105562. [DOI] [PubMed] [Google Scholar]
  • 26.Saab K, Tu T, Weng WH, et al. Capabilities of Gemini models in medicine. ArXiv. 2024; 10.48550/arXiv.2404.18416. [DOI] [Google Scholar]
  • 27.Sharma S, Alaa AM, Daneshjou R.. A longitudinal analysis of declining medical safety messaging in generative AI models. NPJ Digit Med. 2025;8(1):592. doi: 10.1038/s41746-025-01943-1. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Material

Articles from Proceedings (Baylor University. Medical Center) are provided here courtesy of Baylor University Medical Center

RESOURCES