To the Editor,
Allergy diagnostics has taken an innovative turn with the introduction of the skin prick automated test (SPAT) [1, 2]. For respiratory allergy, skin prick testing (SPT) and serum specific IgE (sIgE) measurement are the gold standard methods to detect IgE‐mediated sensitisation [3]. Despite its high sensitivity, manual SPT is subject to operator‐dependent variability and prone to human errors. SPAT addresses these issues by standardizing the entire procedure, applying a fixed amount of allergen and a controlled prick force, in combination with AI‐assisted readout to improve efficiency [1].
Previous research has demonstrated that SPAT, compared to manual SPT, shows lower intra‐subject variability [1], higher consistency, and reduced patient discomfort [4]. When applying the validated 4.5‐mm cut‐off, SPAT shows equivalent diagnostic accuracy in detecting birch pollen and house dust mite allergies [2].
More recently, AI has been integrated into SPAT to provide wheal measurement suggestions, further increasing standardization and efficiency. Seys et al. described how the AI algorithm was trained (n = 651), validated (n = 217), and independently tested (n = 95) [5]. A strong correlation was observed between the physician‐measured and the AI‐measured longest wheal diameter (Pearson r = 0.83; p < 0.0001), with 5.8% of AI measurements adjusted by physicians, resulting in a change in test interpretation in 0.5% of cases.
To evaluate external validity in clinical practice, real‐world SPAT data from 37 hospitals across 5 European countries were assessed. This cohort comprises test results from 10,756 patients (126,526 pricks) with respiratory (96.7%), food (3.3%), insect venom (0.06%) and drug (0.01%) allergens. AI‐measured longest wheal diameters were compared with physician‐verified AI measurements, confirming a strong correlation (Pearson r = 0.96, p < 0.0001; Figure 1). No wheal was detected by AI in 7.9% (10,031) of pricks, with only 0.3% (380) of pricks corresponding to a physician measurement of ≥ 4.5 mm. In total, 6.1% of the wheals were enlarged by the physician (median (IQR): +1.0 mm (+0.5; +1.6)), and 2.9% reduced (median (IQR): −0.8 mm (−1.8; −0.4)), resulting in a change in test interpretation—from negative to positive or vice versa—in 1.7% and 0.4% of the cases, respectively. An analysis of potential confounders is provided in the Supporting Information (Table S3–S5 and Figure S3).
FIGURE 1.

Scatterplot of physician's and AI measurement of longest wheal diameter (116,495 pricks). Pearson r = 0.96 (p < 0.0001).
This large real‐world analysis demonstrates that the AI model is broadly generalizable and shows acceptable clinical performance. The low rate of interpretation‐affecting adjustments is reassuring. The observed high correlation between AI and physician‐verified AI measurements suggests that the AI model performs consistently across different hospitals. This is particularly relevant in comparison to manual SPT, where inter‐ and intra‐observer variability has historically been a major limitation [6]. We observed that the total number of physician adjusted wheals (9.0%) in the current study was lower than the inter‐observer readout variability (median: 19.8%) reported previously [5]. Hence, the physician and AI agree more often than different physicians do. While blinding physicians to AI output would ideally minimize anchoring bias, the real‐world design of this study reflects routine clinical practice, where AI measurements are intentionally used to inform physician review rather than replace independent assessment.
These results support safe integration of the AI‐assisted readout model into daily practice as a decision‐support tool. Rather than replacing clinical judgment, the AI model functions as a reliable first reader, allowing physicians to focus on cases that require expert review [5]. By continuing use across an increasing number of patient populations, we aim to further improve performance, particularly regarding skin type variations where slightly lower performance was observed (Table S4 and Roux et al. [7]).
In conclusion, SPAT device and its AI‐assisted readout method reduce variability, improve consistency, and increase time efficiency compared with manual SPT. This real‐world evidence supports implementation in routine clinical practice, with clear benefits for physicians, healthcare systems, and patient care.
Author Contributions
S.F.S., S.G., and L.V.G. conceived the study, designed the data collection protocol, and performed the primary data analysis. J.W. and S.F.S. drafted the manuscript. S.F.S., D.L., S.G., J.W., and L.V.G. contributed to data acquisition and interpretation of results. D.L. and S.F.S. provided statistical support and reviewed the analytical approach. I.B., G.C., J.C., T.D., G.D.G., J.D.M., N.F., J.H., F.H., S.H., E.H., V.H., B.L., W.L., P.L., C.L., Y.M., T.S.N., H.N., R.O., B.C., M.P., M.S., C.S., S.T., O.V., K.V.G., V.V., A.‐S.V., A.W., R.D., S.G., A.M.C., and L.V.G. critically revised the manuscript for important intellectual content. All authors contributed to and approved the final version of the manuscript.
Funding
This work was supported by Hippo Dx.
Conflicts of Interest
S.F.S., R.D., D.L. and S.G. are employees of Hippocreates BV. S.F.S., R.D., D.L., S.G., L.V.G., and I.B. hold shares of Hippocreates BV. AMC reports grants, speaker honoraria, consultancy or advisory fees and/or research support and other, all via Technical University of Munich from Allergopharma, ALK Abello, Astra Zeneca, Bencard/Allergen Therapeutics, GSK, Novartis, Hippo Dx, LETI, Roche, Zeller, Sanofi, Regeneron, Thermo Fisher, European Institute of Technology (EIT Health) and Federal Ministry of Research and Education Germany. Other authors declare no conflicts of interest.
Supporting information
Table S1: Overview of tested allergens.
Table S2: Overview of hospitals per country.
Table S3: Impact of sex.
Table S4: Impact of skin type.
Table S5: Impact of allergen type.
Table S6: Patient characteristics.
Figure S1: Representative composite image with measurement of the longest wheal diameter.
Figure S2: Bland–Altman plot of AI versus physician‐verified AI measurement.
Figure S3: Accuracy versus number of tests stratified per hospital.
Acknowledgements
Open Access funding provided by Medizinische Universitat Wien.
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
References
- 1. Gorris S., Uyttebroek S., Backaert W., et al., “Reduced Intra‐Subject Variability of an Automated Skin Prick Test Device Compared to a Manual Test,” Allergy 78, no. 5 (2023): 1366–1368. [DOI] [PubMed] [Google Scholar]
- 2. Seys S. F., Gherasim A., Odul F., et al., “Validation of the Skin Prick Automated Test (SPAT) Cut‐Off Value in Birch Pollen and House Dust Mite Allergic Rhinitis Patients,” Allergy 80, no. 12 (2025): 3302–3309. [DOI] [PubMed] [Google Scholar]
- 3. Gureczny T., Heindl B., Klug L., Wantke F., Hemmer W., and Wöhrl S., “Allergy Screening With Extract‐Based Skin Prick Tests Demonstrates Higher Sensitivity Over In Vitro Molecular Allergy Testing,” Clinical and Translational Allergy 13, no. 2 (2023): e12220. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Seys S. F., Roux K., Claes C., et al., “Skin Prick Automated Test Device Offers More Reliable Allergy Test Results Compared to a Manual Skin Prick Test,” Rhinology Journal 62, no. 2 (2024): 216–222. [DOI] [PubMed] [Google Scholar]
- 5. Seys S. F., Hox V., Chaker A. M., et al., “Artificial Intelligence (AI)‐Assisted Readout Method for the Evaluation of Skin Prick Automated Test Results,” Nature Communications 16, no. 1 (2025): 8637. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. McCann W. A. and Ownby D. R., “The Reproducibility of the Allergy Skin Test Scoring and Interpretation by Board‐Certified/Board‐Eligible Allergists,” Annals of Allergy, Asthma & Immunology 89, no. 4 (2002): 368–371. [DOI] [PubMed] [Google Scholar]
- 7. Roux K., Seys S. F., Hox V., et al., “Impact of Real‐World Confounders on the Accuracy of an AI Model to Support Read Out of Skin Prick Automated Test Results,” Rhinology, ahead of print, June 3, 2026, 10.4193/Rhin25.634. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Table S1: Overview of tested allergens.
Table S2: Overview of hospitals per country.
Table S3: Impact of sex.
Table S4: Impact of skin type.
Table S5: Impact of allergen type.
Table S6: Patient characteristics.
Figure S1: Representative composite image with measurement of the longest wheal diameter.
Figure S2: Bland–Altman plot of AI versus physician‐verified AI measurement.
Figure S3: Accuracy versus number of tests stratified per hospital.
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
