Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2025 Aug 27;93(5):370–378. doi: 10.1111/cod.70011

Evaluation of Artificial Intelligence‐Assisted Diagnosis of Skin Erythema in a Patch Test

Seoyoung Kim 1,2, Hyunsik Hwang 3, Mihyun Oh 1, Jieun Han 1, Sodam Park 1, Soyoung Lee 1, Goun Kim 1, Sungwon Cho 3, Dong Hun Lee 4, Jae Youl Cho 2,
PMCID: PMC12503976  PMID: 40859876

ABSTRACT

Background

The patch test evaluates skin erythema, infiltration, papules and vesicles following exposure to various substances, including metals, cosmetics and medicines. Accurate evaluation of these conditions requires consistent skin score assessments, precise visual grading and minimal inter‐expert variability.

Objectives

This study aimed to develop a skin irritation artificial intelligence model based on the YOLOv5x object detection framework to automatically detect skin irritation from the patch test images for multiple test substances.

Methods

Patch test images were collected with test sites marked to enable the YOLOv5x algorithm to locate the samples. An expert assigned a score to each sample (0–4) for training and validation. The model was trained using 83 629 data points. Evaluation and validation were performed with 1312 and 1536 data points, respectively.

Results

The model achieved an overall accuracy of 0.983 at both 24 and 48 h, with an F1 score (harmonic mean of recall and precision) of 0.982. The areas under the curve (AUCs) for scores 0, 1 and 2 were 0.914, 0.838 and 0.865, respectively. The sensitivity for a score of 0 was 0.997.

Conclusion

These findings suggest that this AI model effectively supports and classifies skin irritation, thereby facilitating faster and more accurate dermatological evaluations.

Keywords: AI, human patch test, skin irritation, YOLOv5x


An AI model based on YOLOv5x accurately detected and graded skin reactions from patch test images. It achieved 98.3% accuracy and high sensitivity across scores. This approach enables faster, consistent and automated dermatological assessments.

graphic file with name COD-93-370-g001.jpg

1. Introduction

Patch testing is among the most reliable and widely accepted methods for evaluating the potential of cosmetic products to induce skin irritation and allergic reactions [1, 2, 3]. This approach is essential for assessing product safety prior to market release and plays a critical role in safeguarding consumers from adverse dermatological effects.

Skin reactions are graded based on standardised scales, such as the ICDRG and ESCD [4]. Although visual assessment methods are widely used, they have some inherent limitations. One primary limitation is the subjectivity of visual evaluations, which can lead to inconsistent results due to variability in interpretation and assessment by different observers [5]. Moreover, mild reactions may be overlooked, resulting in potential false‐negative outcomes [6].

Despite its effectiveness as a diagnostic tool, patch testing is subject to observer bias because interpretation of results by observers can vary. For example, inter‐rater variability has been identified in the discrimination between uncertain and irritant reactions, as well as in the distinction between uncertain and weak reactions. To mitigate this, continuous standardisation and training are recommended [4].

Another significant limitation is the influence of skin colour and pigmentation on the visibility of reactions. It is challenging to detect subtle erythema in individuals with darker skin, which could compromise the assessment accuracy. Consequently, it is imperative that observers be adequately trained to recognise and account for these pigmentation differences [7].

To address these limitations, numerous studies have actively explored the use of AI as an assistive diagnostic tool in the field of dermatology, with growing interest in its potential for practical clinical application. This shift has attracted significant scientific interest owing to advances in machine learning (ML) and deep learning (DL). Chan et al. demonstrated the feasibility of using ML to automate the detection of skin reactions and improve diagnostic workflows [8]. Hall et al. further highlighted the application of DL to democratise patch testing, an approach that could enhance accessibility and consistency in clinical dermatology [9]. Moreover, Vezakis et al. evaluated various pre‐processing techniques and modalities for detecting skin reactions via DL, emphasising the importance of advanced image pre‐processing and algorithm refinement in optimising the accuracy of automatic detection systems [10].

These investigations highlight the revolutionary potential of AI‐powered diagnostics in dermatology, particularly for improving the accuracy, consistency and accessibility of epicutaneous patch testing.

In this study, we developed a model that overcomes the limitations of the convolutional neural network (CNN) to classify skin irritation from images using the object detection algorithm YOLO v5 (You Only Looking Once version 5).

2. Methods

2.1. Participants

We analysed patch test data collected at our institution between 2020 and 2023. Participants were recruited from Seoul and Gyeonggi Province based on voluntary interest in the study. The main inclusion criteria were: adults aged 18 years or older who had healthy skin; were able to provide written informed consent; and were able to attend follow‐up visits while complying with all study procedures. Exclusion criteria included: severe atopic dermatitis or other chronic skin conditions affecting test interpretation; recent use of systemic corticosteroids or immunosuppressants within 3 months, or topical corticosteroids or immunomodulators within 1 month; pregnancy or breastfeeding; serious systemic illnesses such as autoimmune diseases, malignancies, or significant kidney or liver dysfunction; uncontrolled chronic conditions like asthma, diabetes, or hypertension; history of severe allergic reactions to cosmetic products; anticipated excessive UV exposure during the test period; and known hypersensitivity or irritation to patch test adhesives or occlusive materials. This study adhered to the Declaration of Helsinki and was approved by the Amorepacific Institutional Review Board (IRB approval no. 2019‐1SR‐N085R, 2020‐1SR‐N20R, 2022‐1SR‐N55R, 2023‐1SR‐C92). Written informed consent was obtained from all the participants.

2.2. Patch Test Procedure

Van der Bend chambers (Brielle, Netherlands) were used for patch testing. The chambers were made of hypoallergenic, medical, non‐woven polyester and each contained chromatographic filter paper. Each sample (0.020 mL) was placed in a Van der Bend chamber.

Cosmetics or its ingredients were placed on the upper back of the participants and closed for 24 h. The evaluation was performed 1 and 24 h after patch removal. The skin response was graded using a 5‐point scale according to a modified version of the Frosch and Kligman and CTFA safety testing guidelines: 0, No visible reaction; 1, Slight erythema, either spotty or diffuse; 2, Moderate uniform erythema; 3, Intense erythema with edema; and 4, Intense erythema with edema and vesicle. Clinical photographs were captured using a CANON EOS 6D MARK II camera (Canon, Tokyo, Japan) in a controlled environment to document the examination and reaction sites, regardless of the presence or absence of a reaction. A white diffuser was used to soften the flash, and two flash units were fixed at an approximately 45° angle above the camera to evenly illuminate the back area of the panellists. All images were taken under consistent conditions to minimise variation.

2.3. Data Acquisition and Preprocessing

All images were captured at a consistent resolution of 4160 × 2768 pixels to accurately detect subtle erythematous changes at patch test sites. Image preprocessing was conducted to prepare data for the YOLOv5x deep learning model. Instead of manually cropping images by reaction, the LabelImg annotation tool was used to draw bounding boxes around each patch test site, precisely localising the corresponding skin reaction regions. Trained clinicians assessed the degree of erythema on a standardised scale from 0 to 4 and, to reduce inter‐rater variability and ensure diagnostic accuracy, all evaluators possessed 8 to 15 years of clinical training in patch testing, during which they consistently applied standardised evaluation criteria and engaged in cross‐validation. For the initial dataset, four independent reviewers evaluated each image and in cases of disagreement, a senior evaluator made the final decision. These scores were linked to the respective bounding boxes within the annotation tool. To ensure annotation quality, images with unclear, ambiguous or otherwise unsuitable regions for training were excluded from the dataset.

Ultimately, a total of 83 629 images were collected, with 1312 and 1536 images allocated to the evaluation and validation sets, respectively. For model training and inference, all images were resized to 1280 × 1280 pixels to balance computational efficiency and preservation of key skin features.

2.4. Model Development

As an object detection model for reading skin irritation, the YOLOv5x architecture was adopted and transfer learning was conducted. YOLOv5x, developed by Ultralytics (Maryland, USA), is an object detection model that employs the CSPNet (Cross Stage Partial Network) structure for fast and efficient detection.

The backbone network of YOLOv5x used in this study consisted of focus, CSP1_X and spatial pyramid pooling (SPP) layers designed to effectively extract features at various scales. The focus layer enhances the computational speed by splitting the input image into patches for processing, whereas the CSP1_X structure increases the computational efficiency and prevents overfitting by dividing and merging the feature maps. The SPP layer further enhances network performance by integrating spatial information through various pooling sizes.

In the network neck, a Path Aggregation Network structure is applied to effectively fuse features extracted at multiple scales, enhancing the detection performance for objects of varying sizes.

The model training was conducted from scratch without using pretrained weights. This approach allowed for better extraction of features suitable for the specific domain of skin irritation detection. Pretrained models often have biases toward general datasets; hence, this study excluded such biases to build a model optimised for the domain.

Hyperparameters were adjusted during the model optimisation process (Appendix A).

2.5. Performance Metrics

The principal metrics used to assess the model performance include true positives (TPs), false negatives (FNs) and false positives (FPs). Precision, calculated as TP/(TP + FP), indicates the proportion of predicted positives that are actually positive, with high precision implying few false‐positive predictions. Recall (sensitivity), expressed as TP/(TP + FN), measures the proportion of true‐positive predictions among all actual positives in the dataset. The F1‐score is the harmonic mean of precision and recall, and it provides a single metric that balances the trade‐off between these two measures. This was calculated using the formula 2TP/(2TP + FP + FN). The F1‐score considers both FPs and FNs, making it useful for imbalanced datasets in which precision and recall are skewed.

2.6. Inter‐Rater Reliability

To evaluate the inter‐rater reliability, readings from five expert evaluators and an AI model were compared and analysed for 32 participants and 3072 samples. The consistency of these readings was assessed using Kappa statistics and Kendall's coefficient of concordance. These analyses were conducted using the Minitab 18.

3. Results

3.1. Participant Information

A total of 3203 individuals participated in the tests, including repeat panellists. The cohort had an average age of 42.5 years, comprising 2815 females (mean age: 43.3 years) and 388 males (mean age: 37.2 years).

3.2. Model Development

Model development was performed according to the data preprocessing, training and optimisation workflow (Figure 1). Upon completion of the data preprocessing, we obtained 83 629 expert‐graded images. Of the 83 629 images, 80 327 (96.1%) were graded as 0, 3020 (3.6%) as 1 and 258 (0.3%) as 2. These data were used for model training (Table 1).

FIGURE 1.

FIGURE 1

Flow chart of the study design for model development. Data preprocessing, training and optimisation workflow.

TABLE 1.

Datasets for individual test sites used in training, evaluation and validation.

Reaction Training (n = 83 629) Evaluation Validation
24 h (n = 1312) 48 h (n = 1312) 24 h (n = 1536) 48 h (n = 1536)
0 80 327 (96.1%) 1256 (95.7%) 1255 (95.7%) 1480 (96.4) 1484 (96.6)
1 3020 (3.6%) 26 (2%) 31 (2.4%) 53 (3.5) 47 (3.1)
2 258 (0.3%) 27 (2.1%) 25 (1.9%) 3 (0.2) 5 (0.3)
3 15 (0.002%) 1 (0.1%) 0 (0%) 0 0
4 9 (0.001%) 2 (0.2%) 1 (0.1%) 0 0

Note: Twenty four‐hour data were obtained 1 h after patch removal, while 48 h data were obtained 24 h later.

3.3. Model Performance

To validate the performance of the trained model, we acquired an evaluation dataset consisting of 2624 samples. The efficacy of the model was assessed using receiver operating characteristic (ROC) curve and area under the curve (AUC) metrics (Figure 2). The model demonstrated an accuracy of 0.983 for the entire dataset. The precision scores were 0.992, 0.650 and 0.905 for grades 0, 1 and 2, respectively. The recall scores were 0.997 for grade 0, 0.684 for grade 1 and 0.731 for grade 2. The F1 scores, which provide a balanced measure of the model performance, exceeded 0.98 across all datasets (Table 2).

FIGURE 2.

FIGURE 2

Receiver operating characteristic (ROC) curve and area under the curve (AUC) for evaluating model performance (Blue line: score 0, Yellow line: score 1, Green line: score 2). (a) Dataset for 24 h. (b) Dataset for 48 h. (c) Dataset for 24 h and 48 h.

TABLE 2.

Confusion matrix and performance metrics for each reaction score.

Score Predicted classification
0 1 2 3 4
Actual classification 0 Accuracy False positive
1 False negative Accuracy Predicted as positive with intensity classification discrepancy
2 Accuracy
3 Predicted as positive with intensity classification discrepancy Accuracy
4 Accuracy
24 h Predicted classification
0 1 2 3 4
Actual classification 0 1252 4 0 0 0
1 11 15 0 0 0
2 0 5 22 0 0
3 0 0 1 0 0
4 0 1 1 0 0
48 h Predicted classification
0 1 2 3 4
Actual classification 0 1251 3 1 0 0
1 7 24 0 0 0
2 1 8 16 0 0
3 0 0 0 0 0
4 0 0 1 0 0
Dataset Accuracy Precision_by scores Recall_by scores F1_score
0 1 2 0 1 2
24 h 0.982 0.991 0.600 0.917 0.997 0.577 0.815 0.981
48 h 0.984 0.994 0.686 0.889 0.997 0.774 0.640 0.983
Overall 0.983 0.992 0.650 0.905 0.997 0.684 0.731 0.982

Note: When analyzing the model results, it is important to distinguish between cases where negative results are misclassified as positive (or vice versa), and cases where the result is correctly classified as positive but the classification intensity are incorrect (in color shades).

3.4. Inter‐Rater Reliability of the Model Validation Study

We examined the concordance between the model and four experts by assessing the inter‐rater agreement using 3072 images. The accuracy rates for the four experts were 98.01% (97.46, 98.48), 97.01% (96.34, 97.58), 97.04% (96.38, 97.61) and 97.46% (96.84, 97.99). Of 3072 cases assessed for concordance between researchers and AI, the scores for 2906 (94.6%) cases were in agreement. Both the Kappa statistic and Kendall's coefficient of concordance indicated p values of 0.0000 (Table 3; Figure 3).

TABLE 3.

Inter‐rater reliability of experts and model prediction.

Expert Accuracy (95% CI) Inter‐rater reliability (Fleiss' k)
Model_Predicted 97.36 (96.73, 97.90) 0.63
1 98.01 (97.46, 98.48) 0.69
2 97.01 (96.34, 97.58) 0.62
3 97.04 (96.38, 97.61) 0.52
4 97.46 (96.84, 97.99) 0.55

FIGURE 3.

FIGURE 3

Model‐to‐expert visual scoring agreement.

3.5. Deep Learning‐Based Automated Detection of Skin Irritation in Patch Testing

We developed an integrated system that includes the following three key features to efficiently evaluate and manage skin irritation tests. We developed a history management tool that systematically manages and operates a database of irritation test histories to enable effective test system management. Furthermore, the system automatically detected sample attachment locations from the uploaded image files and evaluated the erythema reaction intensity for each area to provide real‐time analysis results. By adding a score correction feature to detect erroneous readings, we automated the skin irritation assessment process while allowing experts to review and make judgements (Figure 4).

FIGURE 4.

FIGURE 4

Deep learning‐based automated detection of skin irritation in patch testing.

4. Discussion

Advancements in AI technology are being leveraged to diagnose skin conditions such as acne and eczema. Training on a large dataset of skin images demonstrated an accuracy comparable to that of dermatologists. In a recent study, an AI system called AcneDet was developed to automatically detect acne and classify its severity using facial images captured by smartphones. The system achieved an average accuracy of approximately 0.85 [11]. Another study developed an automatic version of the SCORAD using a state‐of‐the‐art CNN to analyse skin lesion images and assess the severity of atopic dermatitis. This system demonstrated evaluation results similar to those of human experts while reducing inter‐rater variability [12].

In addition to these areas of research, AI has been used to predict contact dermatitis, with recent studies highlighting the potential of ML techniques to enhance diagnosis and accuracy in this field. A study conducted by Chan [8] developed a CNN model aimed at predicting allergic contact dermatitis from patch test images, achieving a remarkable accuracy of 99.5%. Another study by Ravishankar [13] utilised a CNN model to differentiate between reactions and non‐reactions; the latter category encompassed those showing no response in the patch test and suspected irritant reactions, achieving an accuracy of 90.1%. Both studies implemented a method in which images obtained from participants were used for reaction classification. These promising results suggest that CNN models can significantly enhance the diagnostic process for allergic or contact dermatitis, offering high accuracy in distinguishing different types of skin reactions [14].

In contrast to these studies, we developed a method for training without cropping images. We also trained the model on a large number of images using the object detection algorithm YOLO v5 to improve the accuracy.

Confusion matrix analysis revealed that the AI model demonstrated outstanding precision and recall in identifying grade 0 reactions, achieving above 99%, yet its ability to distinguish between adjacent erythema grades—particularly slight (grade 1) and moderate (grade 2)—remained limited. At the 24‐h evaluation, a substantial portion (42.3%, 11/26) of true grade 1 reactions was underestimated as grade 0; this underestimation decreased to 22.6% (7/31) by 48 h. Conversely, grade 2 reactions were increasingly misclassified as grade 1 over time, rising from 22.7% (5/22) at 24 h to 36.0% (9/25) at 48 h, indicating a growing tendency to underestimate moderate erythema as slight erythema as time progresses.

Consequently, our model achieved an overall accuracy of 0.983 with an F1 score of 0.982. Among the training data, 96.1% were graded as 0, with a classification accuracy of 0.992 for this grade. The accuracy for grade 2 was 0.905. However, for grade 1 images, relatively lower accuracy (0.650) and recall (0.684) were observed, reflecting inconsistency in the training response values. These misclassifications and inconsistencies likely stem from challenges in differentiating slight erythema from transient irritant responses immediately after patch removal, the subjective nature of interpreting borderline cases between grades 1 and 2, differences among experts in erythema assessment, irritation biases in surrounding skin areas and false‐positive skin reactions induced by coloured products such as lipstick.

Notably, the accuracy for grade 1 classification improved markedly from 57.7% at 24 h to 77.4% at 48 h, suggesting that transient irritant effects resolved over time, allowing clearer allergic reaction presentations. In contrast, the accuracy for grade 2 decreased from 81.5% to 64.0%, implying that moderate erythema may fade or become less distinct as time passes, complicating accurate interpretation.

Despite these limitations, the model's high performance in accurately detecting grade 0 reactions is clinically significant, as it helps reduce false negatives and minimises the risk of misjudgement.

In the inter‐rater reliability of the model validation study, our AI model demonstrated the same level of consistency as trained professionals, with 94.6% agreement (2906 of three cases) between the trained experts and AI. This was supported by both the Kappa statistic and Kendall's coefficient of concordance, with p values of 0.0000 confirming the utility of the model.

Although the model demonstrated high accuracy, several limitations remain that should be acknowledged and addressed in future studies. As the training dataset consisted of images taken at 1 and 24 h after patch removal, the model has limited predictive capability for identifying delayed skin reactions that may develop beyond this timeframe. The incorporation of a second or delayed reading, typically conducted at 72‐ or 96‐h post‐application, may enhance the identification of transient irritant reactions and enable clearer differentiation from true allergic responses by capturing delayed hypersensitivity reactions. This approach has the potential to reduce the incidence of false‐positive results frequently observed with the conventional 48‐h reading. In addition, by collecting a sufficient number of cases exhibiting mechanical irritation, borderline reactions, pigmentary changes and the angry‐back phenomenon, and accurately labelling them to expand the dataset, the model can be trained to exclude these confounding factors from the scoring process, thereby effectively reducing false‐positive reactions.

Moreover, this model was specifically developed for the evaluation of skin irritation induced by cosmetic ingredients and products. However, due to the infrequent occurrence of grade 3 or 4 reactions in cosmetic patch testing, it was not possible to derive meaningful performance metrics for these higher severity grades, representing a limitation in the current study. Additionally, assessment indicators for patch‐test grades 3 or 4, such as edema and vesicles, are difficult to verify using images.

Furthermore, the trained images consisted only of Fitzpatrick skin types of Koreans among whom, according to the study by Youn et al., type III is the most common, with types III, IV and V together accounting for 88.8% of the total population. This poses limitations in interpreting a wide range of skin types [15].

To address these limitations and further improve model performance, future research should emphasise the refinement of reaction grade classification and the extension of applicability across varied clinical environments. To improve the accuracy and precision of grade 1 reaction detection, we propose several approaches.

First, addressing class imbalance through targeted data augmentation may enhance model performance. Specifically, acquiring additional data for underrepresented classes by collecting more clinical images or re‐evaluating existing subclinical cases can improve dataset balance and quality. Furthermore, object‐level selective augmentation where only lesion areas labelled with specific grades are extracted and augmented using techniques such as contrast enhancement, HSV (hue, saturation and value) colour space adjustments, minor rotations, or noise injection can diversify the training data. These augmented lesion patches can then be reintegrated into original or synthetic backgrounds to further enrich the dataset. Furthermore, to improve detection accuracy on darker skin tones, targeted data augmentation that includes images from a range of skin tones and colour space adjustments accounting for pigmentation differences should be employed. This approach enhances the model's ability to detect subtle erythema and other dermatological features that are less visually distinct on darker skin.

Second, advanced image processing and feature extraction methods may increase detection sensitivity. Incorporating multi‐channel image analysis with texture‐based features, such as Local Binary Patterns (LBP), can help emphasise subtle distinctions between erythematous and normal skin regions, potentially improving classification accuracy.

Third, improving inter‐rater reliability is crucial for consistent grade labelling. This can be achieved through additional training and standardisation protocols for expert evaluators, clearer documentation of grading criteria and the development of a comprehensive reference image library to enhance reading consistency. The validation strategy will be expanded to assess inter‐rater variability among dermatologists across multiple clinical sites, thereby enhancing the robustness and clinical relevance of the model. Additionally, to enhance interpretability beyond providing prediction results, visualisation techniques such as Grad CAM and Saliency Maps will be implemented to highlight key reactive areas and clearly present the AI's decision making process. Incorporating a system that collects clinical background information, including patient history, allergy records and product usage will enable the AI to more accurately assess the clinical relevance of the reactions.

Finally, since the model was developed using data captured at a consistent resolution (4160 × 2768) from a single institution, it does not fully reflect the variability of real‐world clinical environments. To address this limitation, future studies will incorporate image datasets collected from multiple institutions. Additionally, performance evaluations using images captured by more accessible devices, such as smartphone cameras, will be conducted to expand the model's applicability in telemedicine.

The AI‐based erythema reading model developed in this study demonstrates significant potential to enhance the efficiency of evaluations while minimising inter‐rater variability, thereby enabling more objective and consistent assessments. Moreover, the integration of the proposed future improvements is expected to further increase the accuracy and reliability of patch test reaction grading, ultimately broadening the model's applicability across a variety of clinical environments.

Author Contributions

Seoyoung Kim: conceptualization, investigation, validation, data curation, writing – original draft, writing – review and editing, methodology, project administration, funding acquisition. Hyunsik Hwang: conceptualization, methodology, data curation, validation, investigation, writing – original draft. Mihyun Oh: data curation, validation. Jieun Han: data curation, validation. Sodam Park: data curation, validation. Soyoung Lee: data curation, validation. Goun Kim: data curation, validation. Sungwon Cho: supervision. Dong Hun Lee: supervision. Jae Youl Cho: supervision.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A.

YOLOv5x Hyperparameters Used in This Study

Hyperparameters Values
lr0 (Initial learning rate) 0.01
lrf (Final OneCycleLR learning rate) 0.01
momentum (SGD momentum/Adam beta1) 0.937
weight_decay (Optimizer weight decay) 0.0005
warmup_epochs (Warmup epochs) 3
warmup_momentum (Warmup initial momentum) 0.8
warmup_bias_lr (Warmup initial bias learning rate) 0.1
box (Box loss gain) 0.05
cls (Classification loss gain) 0.5
cls_pw (Classification BCELoss positive weight) 1
obj (Object loss gain) 1.0
obj_pw (Object BCELoss positive weight) 1
iou_t (IoU training threshold) 0.2
anchor_t (Anchor‐multiple threshold) 4
fl_gamma (Focal loss gamma) 0
hsv_h (Image HSV‐Hue augmentation fraction) 0.015
hsv_s (Image HSV‐Saturation augmentation fraction) 0.7
hsv_v (Image HSV‐Value augmentation fraction) 0.4
degrees (Image rotation) 0
translate (Image translation fraction) 0.1
scale (Image scale gain) 0.5
shear (Image shear degrees) 0
perspective (Image perspective fraction) 0.0001
flipud (Image flip up‐down probability) 0
fliplr (Image flip left–right probability) 0
mosaic (Image mosaic probability) 1
mixup (Image mixup probability) 0.1
copy_paste (Segment copy‐paste probability) 0.1
mosaic (Image mosaic probability) 1
mixup (Image mixup probability) 0.1
copy_paste (Segment copy‐paste probability) 0.1

Kim S., Hwang H., Oh M., et al., “Evaluation of Artificial Intelligence‐Assisted Diagnosis of Skin Erythema in a Patch Test,” Contact Dermatitis 93, no. 5 (2025): 370–378, 10.1111/cod.70011.

Data Availability Statement

Research data are not shared.

References

  • 1. Garg T., Agarwal S., Chander R., Singh A., and Yadav P., “Patch Testing in Patients With Suspected Cosmetic Dermatitis: A Retrospective Study,” Journal of Cosmetic Dermatology 17, no. 1 (2018): 95–100. [DOI] [PubMed] [Google Scholar]
  • 2. Cosmetics Europe , “Product Test Guidelines for the Assessment of Human Skin Compatibility,” (1997).
  • 3. Johansen J. D., Mahler V., Lepoittevin J. P., and Frosch P. J., Contact Dermatitis, 6th ed. (Springer, 2021). [Google Scholar]
  • 4. Johansen J. D., Aalto‐Korte K., Agner T., et al., “European Society of Contact Dermatitis Guideline for Diagnostic Patch Testing ‐ Recommendations on Best Practice,” Contact Dermatitis 73, no. 4 (2015): 195–221. [DOI] [PubMed] [Google Scholar]
  • 5. Basketter D., Reynolds F., Rowson M., Talbot C., and Whittle E., “Visual Assessment of Human Skin Irritation: A Sensitive and Reproducible Tool,” Contact Dermatitis 37, no. 5 (1997): 218–220. [DOI] [PubMed] [Google Scholar]
  • 6. Arora P., Brumley C., and Hylwa S., “Clinical Relevance of Doubtful Reactions in Patch Testing: A Single‐Centre Retrospective Study,” Contact Dermatitis 90, no. 6 (2024): 607–612. [DOI] [PubMed] [Google Scholar]
  • 7. Nguyen L., Parker L., Hennessy K., Shah N., and Cohen G., “Comparison of Patch Testing Results of White and Black Patients,” Journal of Clinical and Aesthetic Dermatology 17, no. 6 (2024): 55–57. [PMC free article] [PubMed] [Google Scholar]
  • 8. Chan W. H., Srivastava R., Damaraju N., et al., “Automated Detection of Skin Reactions in Epicutaneous Patch Testing Using Machine Learning,” British Journal of Dermatology 185, no. 2 (2021): 456–458. [DOI] [PubMed] [Google Scholar]
  • 9. Hall M. R., Weston A. D., Wieczorek M. A., et al., “An Automated Approach for Diagnosing Allergic Contact Dermatitis Using Deep Learning to Support Democratization of Patch Testing,” Mayo Clinic Proceedings: Digital Health 2, no. 1 (2024): 131–138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Vezakis I. A., Lambrou G. I., Kyritsi A., Tagka A., Chatziioannou A., and Matsopoulos G. K., “Detecting Skin Reactions in Epicutaneous Patch Testing With Deep Learning: An Evaluation of Pre‐Processing and Modality Performance,” Bioengineering 10, no. 8 (2023): 924. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Huynh Q. T., Nguyen P. H., Le H. X., et al., “Automatic Acne Object Detection and Acne Severity Grading Using Smartphone Images and Artificial Intelligence,” Diagnostics 12, no. 8 (2022): 1879. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Medela A., Mac Carthy T., Aguilar Robles S. A., Chiesa‐Estomba C. M., and Grimalt R., “Automatic SCOring of Atopic Dermatitis Using Deep Learning: A Pilot Study,” JID Innovations 2, no. 3 (2022): 100–107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Ravishankar A., Heller N., and Bigliardi P. L., “Demonstration of Convolutional Neural Networks to Determine Patch Test Reactivity,” Dermatitis 35, no. 2 (2024): 144–148. [DOI] [PubMed] [Google Scholar]
  • 14. McMullen E., Grewal R., Storm K., et al., “Diagnosing Contact Dermatitis Using Machine Learning: A Review,” Contact Dermatitis 91, no. 3 (2024): 186–189. [DOI] [PubMed] [Google Scholar]
  • 15. Youn J. I., Choe Y. B., Park S. B., et al., “The Fitzpatrick Skin Type in Korean People,” Korean Journal of Dermatology 38, no. 7 (2000): 920–927. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Research data are not shared.


Articles from Contact Dermatitis are provided here courtesy of Wiley

RESOURCES