ABSTRACT
Skin tears (ST) are common traumatic wounds, particularly among older adults, that can lead to complications if not accurately assessed and classified. The International Skin Tear Advisory Panel (ISTAP) classification system is widely used internationally; however, no validated Persian version currently exists. To culturally adapt, and evaluate the clinimetric properties of the Persian version of the ISTAP Classification System. This methodological study was conducted from February to May 2025 in multiple phases. After forward–backward translation and expert review, face and content validity were assessed. Criterion validity was assessed by comparing nurses' classifications with expert consensus using weighted Cohen's kappa coefficient. Construct validity was examined using the known‐groups method, comparing skin tear frequency and severity between 30 elderly patients with impaired mobility and 30 younger adults without impaired mobility. Reliability was evaluated using Fleiss' kappa coefficient for multiple raters, and weighted Cohen's kappa coefficient for inter‐rater and intra‐rater agreement. Diagnostic accuracy indices, including sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), positive likelihood ratio (LR+), negative likelihood ratio (LR–), odds ratio (OR) and the area under the receiver operating characteristic curve (AUC), were calculated for each skin tear type. Content validity was excellent (content validity ratio (CVR): 0.82–1.00; item‐level content validity index (I‐CVI): 0.91–1.00; scale‐level content validity index (S‐CVI/Ave): 0.94). Criterion validity showed almost perfect agreement with experts (weighted κ = 0.902, p < 0.001). Construct validity was supported by significant group differences in skin tear frequency (Fisher's exact p = 0.001) and severity (t(58) = 2.12, p = 0.039). Reliability was substantial to almost perfect across analyses (Fleiss' κ = 0.8447; inter‐rater weighted κ = 0.66; intra‐rater weighted κ = 0.86). Diagnostic accuracy was excellent for all types (AUC = 0.99), with sensitivity 97.5%–99.2%, specificity 98.4%–99.6%, PPV 97.5%–99.3%, NPV 98.1%–99.6% and very high OR and LR values. The Persian version of the ISTAP Classification System demonstrated excellent validity, reliability and diagnostic accuracy, supporting its use as a standardised tool for assessing ST in Persian‐speaking healthcare settings.
Keywords: cross‐cultural adaptation, diagnostic accuracy, ISTAP classification, psychometric validation, skin tears
Summary
The Persian ISTAP Classification System was successfully translated and culturally adapted.
Face, content, criterion and construct validity confirmed the clarity and clinical relevance of the Persian version.
The tool showed excellent inter‐rater and intra‐rater reliability across all categories.
Strong agreement with expert assessments supports its diagnostic validity.
The Persian ISTAP is quick, practical, and suitable for routine clinical use and research in Iran.
1. Introduction
Skin tears (ST) represent a frequently encountered type of acute wound that, if not appropriately treated, can develop into more serious chronic wounds [1]. According to the International Skin Tear Advisory Panel (ISTAP), these injuries are defined as ‘traumatic wounds resulting from mechanical forces, including adhesive removal’, and they vary in severity but do not penetrate beyond the subcutaneous layer [2]. ST is linked to various risk factors, including extremes of age (neonatal and elderly) [3], dehydration, malnutrition characterised by reduced serum albumin levels, pharmacological treatments such as corticosteroids and anticoagulants, poor activities of daily living or impaired mobility, critical or chronic disease and mechanical forces associated with routine care [4]. In many instances, these injuries remain under‐reported. They are observed in all types of healthcare environments and are especially prevalent among older adults, infants and individuals with critical or long‐term health conditions [5, 6]. While they can affect any area of the body, they are most commonly seen on the arms and legs [3, 7, 8].
Incidence rates of ST range from 2.2% to 92%, with the highest rates found in long‐term care facilities [1, 9, 10]. According to a meta‐analysis of 23 studies involving a total of 59 489 older adults, the prevalence of ST was 6.5% [6]. These variations are linked to differences in patient profiles, prevention and care practices, staff training and the lack of standardised assessment methods [1]. No previous studies in Iran have specifically assessed ST. The only related research, conducted in Iranian neonatal intensive care units, reported that 72% of neonates experienced skin injuries, with bruises being the most common type [11].
One of the major challenges in managing ST has been the lack of a global tool for their identification and assessment [12]. ST is frequently misidentified or under‐reported due to their not being recognised as a distinct wound category [13]. Poor diagnostic accuracy of ST can lead to delayed healing, complications, longer hospital stays and higher costs, reducing care quality [14]. Early use of a valid international classification tool is essential for effective management [2].
So far, three different tools have been created to classify ST [15, 16, 17]. The Payne‐Martin classification system, introduced in 1990, was the first to categorise ST into three categories and five types based on wound morphology [15]. The authors later revised some definitions and assessed the system's internal validity, external validity and utility—expressing concern over limited practical use [15]. White suggested that the limited use of the classification system was mainly due to insufficient awareness and familiarity among clinicians, particularly in Australian healthcare settings [18]. To improve applicability, Carville and colleagues developed the STAR classification system, modifying the original model and validating it through consensus [16]. However, its use in practice remains limited, with few supporting studies. Both the Payne‐Martin and STAR classification systems were considered too complex for routine clinical use and failed to achieve wide adoption [19, 20]. Additionally, the Payne‐Martin system was never psychometrically evaluated [15]. To address the need for a simpler, more practical tool, the ISTAP consensus panel developed and validated a user‐friendly system that classifies ST into three types: Type 1 (no skin/flap loss), Type 2 (partial skin loss) and Type 3 (total skin loss) [19]. Since its inception, the ISTAP classification system has been translated into 15 languages and its clinimetric properties have been evaluated in multiple regions, including Spanish [21], Denmark [22], Sweden [23], French Canada [24], Brazil [4], Chile [4] and Australia [15]. A global validation study involving 1601 healthcare professionals from 44 countries further confirmed its reliability and diagnostic accuracy [1].
Currently, there is a notable lack of research on ST within Iranian healthcare settings. No studies have been identified that focus on the prevalence, assessment or management of ST in Iran. Moreover, there is no evidence of the ISTAP Classification System being adapted or validated in the Persian language. This gap underscores the need for localised research and the development of culturally and linguistically appropriate tools to improve the diagnosis and management of ST in Iran.
Given the growing awareness of wound care needs in Iran and the importance of using standardised, evidence‐based tools, developing a culturally and linguistically adapted Persian version is essential to ensure accurate assessment, documentation and management of ST in Persian‐speaking healthcare settings.
2. Methods
2.1. Study Designs
This methodological study aimed to culturally adapt the ISTAP Classification System into Persian and assess its clinimetric properties—specifically its validity and reliability—between February and May 2025. The study was conducted in multiple phases involving diverse participant groups (Figure 1).
FIGURE 1.

Flow chart of the study.
2.2. Study Population/Sampling/Settings
In total, 147 healthcare providers—including nurses and specialists—as well as 60 patients were recruited through convenience sampling. Participants were voluntarily recruited from critical care units, intensive care units, surgical, internal medicine and psychiatric wards of hospitals in Iran between April and July 2025.
2.3. Adaptation the Persian Version of ISTAP Classification System and Qualitative Content Validity
The Persian adaptation of the ISTAP Classification System was conducted following the standardised methodology outlined by Beaton et al. [25] and Ramada‐Rodilla et al. [26], encompassing five key steps:
Forward translation: two independent bilingual translators, both native Persian speakers fluent in English separately translated the original ISTAP classification tool into Persian.
Synthesis of translations: the research team reviewed and compared both forward translations, resolving any discrepancies to generate a single, synthesised Persian version.
Back translation: two different translators, blinded to the original instrument, independently back‐translated the synthesised Persian version into English to verify conceptual equivalence.
Expert review: a panel of five experts—including two wound care specialists, one nursing methodologist, one clinical linguist and one professional translator—evaluated all translation versions for semantic, idiomatic, experiential and conceptual equivalence. Experts were purposefully selected based on their experience in wound care, instrument translation and cross‐cultural adaptation. The number of five experts was determined according to Beaton et al. [25] and Ramada‐Rodilla et al. [26], who recommend a panel of four to six experts for expert committee review in the translation phase to ensure diverse perspectives while maintaining manageability of consensus. This step corresponded to the qualitative phase of content validity, focusing on item clarity, linguistic equivalence and cultural appropriateness prior to quantitative validation. Qualitative content validity refers to the process of evaluating an instrument's items for clarity, relevance, comprehensibility and cultural appropriateness through expert judgement and feedback [27].
Pre‐testing (pilot testing): the pre‐final Persian version was pilot‐tested with 10 wound care nurses. Participants first received a 30‐min training on skin tear risk factors, prevention, management and the adapted classification system. Then, they independently classified 30 standardised photographs (from the original ISTAP validation set [19]) equally representing Types 1, 2 and 3. These images was also used for the Chilean Spanish [21], Danish [22] and Swedish [23] translation and validation of the classification system. Images were displayed randomly via PowerPoint, and participants had 30 s to classify each photo using a structured response form. The session lasted approximately 1 h.
3. Clinimetric Properties
3.1. Face Validity
The qualitative face validity of the Persian version of the ISTAP Classification System was evaluated through cognitive interviews with a panel of 10 experts, including wound care nurses and clinical specialists. The experts were purposefully selected based on their professional experience in wound management, clinical education and familiarity with skin tear assessment. This purposeful selection ensured that participants had sufficient clinical and theoretical expertise to evaluate the clarity, comprehensibility and cultural relevance of each item within the context of Iranian healthcare practice.
3.2. Quantitative Content Validity
To quantitatively assess the content validity of the Persian version of the ISTAP Classification System, we employed three indices: content validity ratio (CVR), item‐level content validity index (I‐CVI) and scale‐level content validity index using the average method (S‐CVI/Ave). Eleven experts specialising in wound care management participated in this evaluation. They were purposefully selected from academic and clinical settings based on their professional experience in wound assessment, education and instrument validation. For CVR, each expert rated the necessity of each item as ‘essential’, ‘useful but not essential’, or ‘not necessary’. CVR was calculated using Lawshe's formula: CVR = (n e − N/2)/(N/2). According to Lawshe's table, with 11 experts, a CVR value of 0.59 or higher is considered acceptable [28]. To estimate I‐CVI, the experts assessed each item's relevance on a four‐point Likert scale: 1 (not relevant) to 4 (highly relevant). I‐CVI was computed as the proportion of experts rating the item as either 3 or 4. Based on the recommendation of Lynn, for a panel of six or more experts, an I‐CVI of 0.78 or higher is acceptable [29]. Further, S‐CVI/Ave was calculated across all items. An S‐CVI/Ave of 0.90 or higher indicates excellent content validity for the overall scale. According to Polit and Beck, an S‐CVI/Ave of 0.90 or higher indicates excellent content validity for the overall scale [30].
3.3. Criterion Validity
Concurrent validity was assessed by comparing the classifications made using the Persian version of the ISTAP Classification System to a reference standard established by expert consensus, as recommended in validation methodology when no absolute gold standard is available [25, 31, 32].
A panel of 10 wound care experts—including nurses, dermatologists and academic clinicians—received structured training on the original English version of the ISTAP classification system using official guidelines and sample cases. After training, the experts independently classified 45 standardised skin tear cases (clinical photos + brief vignettes). The consensus classification for each case was determined using the majority rule and served as the reference standard.
Separately, 10 nurses trained in the Persian ISTAP version classified the same cases. Agreement between their classifications and the expert consensus was evaluated using Weighted kappa, with κ ≥ 0.70 considered acceptable [33].
The use of standardised photographs ensured consistent visual conditions, equal exposure of all raters to identical wound characteristics and facilitated controlled comparison between raters and the expert reference standard. Although reliability in live clinical assessments may sometimes be higher than in photographic evaluations, the photograph‐based approach is widely accepted in early validation studies because it minimises patient‐related variability and ethical challenges associated with repeated wound assessments in clinical settings [1, 4, 21, 22].
3.4. Construct Validity
Construct validity was evaluated using the known‐groups method. Based on Cohen's guidelines (1988), a minimum sample size of 32 participants (16 per group) is recommended to detect a large effect size (w = 0.50) with 80% statistical power at a 5% significance level [34]. Accordingly, 60 hospitalised patients were recruited and divided into two groups: elderly patients with impaired mobility or dependence (n = 30) and younger adults without impaired mobility or dependence (n = 30). All participants or their proxies provided informed consent before participating in the study.
The frequency and severity of ST were measured and compared between these groups. Severity of ST was classified into three types (Type 1 to Type 3), coded numerically from 1 to 3. Statistical analyses, including chi‐square tests and independent samples t‐tests, were used to assess whether the tool could discriminate between groups expected to differ in skin tear occurrence and severity, thus supporting its construct validity.
The Shapiro–Wilk test was conducted to assess the normality of the severity scores. The results indicated no significant deviation from normality (W = 0.981, p = 0.485), suggesting that the severity variable is approximately normally distributed.
3.5. Reliability Analysis
Inter‐rater reliability was assessed to evaluate the reliability of the Persian version of the ISTAP Classification System. The sample size for the inter‐rater reliability was determined based on Nunnally's recommendation [35], which suggests recruiting a minimum of five participants per item or instrument to ensure adequate statistical power. Given that the ISTAP classification system has three items, we recruited 30 other nurses. Each participant was asked to classify the standardised ISTAP photograph set according to the type of injury, following the procedures established in the pilot study. Informed consent was obtained from all participants prior to their involvement in the study. The classification tasks were conducted at the participants' respective workplaces to facilitate convenience and ensure a comfortable environment.
To evaluate inter‐rater reliability in a real clinical context, two trained nurses independently assessed 20 cases of ST in hospitalised patients using the Persian version of the ISTAP Classification System. The number of raters and cases was limited due to logistical and ethical considerations, including the need to minimise patient manipulation and ensure uniform wound exposure conditions during evaluation. Employing two experienced raters enabled a controlled and standardised comparison, minimising variability related to differing levels of clinical expertise or interpretation. Although including a larger number of nurses and observations would enhance generalizability, similar small‐scale inter‐rater assessments have been reported in previous ISTAP validation studies [1, 4, 21, 22].
To evaluate the consistency of raters' assessments over time, intra‐rater reliability was measured by another 30 nurses using weighted Cohen's kappa coefficient. Each nurse independently rated a set of three standardised photographs (one from each category: Type 1, Type 2 and Type 3 ST) on two separate occasions spaced 2 weeks apart. The ratings from the first and second assessments were compared for each rater individually. Both unweighted and weighted Cohen's kappa statistics were calculated to assess the degree of agreement beyond chance, with weighted kappa accounting for the severity of disagreement between categories.
3.6. Diagnostic Accuracy and Agreement Assessment
To evaluate the diagnostic accuracy of nurses' classification of ST, the expert consensus served as the gold standard. Each of the 15 standardised skin tear photographs was independently classified by 20 nurses and five experts using the ISTAP Classification System, which categorises tears into three types.
For each skin tear type (Type 1, Type 2 and Type 3), we converted the classifications into binary outcomes (presence vs. absence of that type) to calculate diagnostic accuracy measures. Sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), likelihood ratios (LR+ and LR–), odds ratios (OR) and the area under the receiver operating characteristic curve (AUC) were computed to assess nurses' ability to correctly identify each type compared to expert ratings.
The sensitivity quantified the proportion of expert‐confirmed cases correctly identified by nurses, while specificity measured the proportion of non‐cases correctly excluded. PPV and NPV reflected the probability that nurses' positive and negative classifications were accurate, respectively. Likelihood ratios and odds ratios provided additional measures of diagnostic performance strength.
3.7. Statistical Analysis
All statistical analyses were performed using Stata/MP version 17. Descriptive statistics were used to summarise participant characteristics, expressed as means and standard deviations (SD) for continuous variables and frequencies with percentages for categorical variables. Content validity was evaluated using the CVR, I‐CVI and S‐CVI/Ave, with acceptable thresholds set at CVR ≥ 0.59 [28], I‐CVI ≥ 0.78 [29] and S‐CVI/Ave ≥ 0.90 [30]. Criterion validity was assessed by comparing nurses' classifications using the Persian ISTAP Classification System with expert consensus, with agreement quantified using weighted Cohen's kappa. Construct validity was examined using the known‐groups approach, comparing skin tear frequency between elderly patients with impaired mobility and younger adults without impaired mobility using Pearson's chi‐square or Fisher's exact test as appropriate, and comparing mean severity scores with independent samples t‐tests after confirming normality using the Shapiro–Wilk test. Reliability was assessed through Fleiss' kappa for multiple raters, weighted Cohen's kappa for two‐rater inter‐rater agreement and weighted Cohen's kappa for intra‐rater agreement across two assessments. The kappa statistic quantifies agreement on a scale ranging from 0 to 1, where 0 signifies agreement expected by chance and 1 indicates perfect agreement [33]. Landis and Koch (1977a, 165) provide the following interpretation for intermediate values: less than 0.00 reflects poor agreement; 0.00–0.20, slight; 0.21–0.40, fair; 0.41–0.60, moderate; 0.61–0.80, substantial; and 0.81–1.00, almost perfect agreement. Diagnostic accuracy was calculated for each skin tear type (Types 1–3) against the expert consensus gold standard, including sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), likelihood ratios (LR+ and LR–), odds ratios (OR) and the area under the receiver operating characteristic (ROC) curve (AUC), all reported with 95% confidence intervals (CIs). A two‐tailed p < 0.05 was considered statistically significant.
4. Result
4.1. Cultural Adaptation
The Persian translation of the ISTAP Classification System was completed successfully following standardised cross‐cultural adaptation guidelines. Minor linguistic adjustments were made during expert panel review to ensure conceptual equivalence and cultural appropriateness. The final Persian version of the ISTAP Classification System (Figure 2) was reviewed by the ISTAP team for conceptual and linguistic equivalence and subsequently approved via email correspondence.
FIGURE 2.

Persian version of ISTAP classification system with both Persian and English translated versions. * A skin tear is an injury resulting from mechanical forces, such as friction, shear or the removal of adhesive materials. The severity of the tear can vary depending on its depth but does not penetrate beyond the subcutaneous tissue layer. ** In the context of skin tears, a flap refers to a segment of skin—comprising the epidermis and/or dermis—that becomes partially or completely detached from its original position due to shear, friction or blunt trauma. This should not be mistaken for skin tissue that is deliberately separated for medical procedures, such as surgical skin grafts.
4.2. Face and Content Validity
Face validity was confirmed. For qualitative content validity, the items consistently exhibited simplicity and clarity during both face and content validity assessments, and no revisions were deemed necessary. As Table 1 shows, all CVR scores ranged from 0.82 to 1.00, indicating substantial agreement between experts on the necessity of the selected items. The I‐CVI values ranged from 0.91 to 1.00 and the S‐CVI/Ave value was calculated as 0.94. The content validity and expert agreement on the relevance and representativeness of the translation tool were excellent.
TABLE 1.
Content validity results for the Persian version of the ISTAP classification system (N = 11 experts).
| Item no. | Item description | CVR | I‐CVI | S‐CVI/Ave |
|---|---|---|---|---|
| 1 | Type 1: no skin loss | 1.00 | 1.00 | 0.94 |
| 2 | Type 2: partial flap loss | 0.91 | 0.91 | |
| 3 | Type 3: total flap loss | 0.82 | 0.91 |
Abbreviations: CVR, content validity ratio; I‐CVI, item‐level content validity index; S‐CVI/Ave, scale‐level content validity index.
4.3. Concurrent Validity
The agreement between nurse classifications using the Persian ISTAP system and the expert consensus reference standard was excellent. As shown in Table 2, nurses correctly classified the majority of skin tear cases in accordance with expert consensus: For Type 1 skin tears, four out of five cases were correctly classified. For Type 2 skin tears, 17 out of 18 cases were correctly classified. For Type 3 skin tears, 21 out of 22 cases were correctly classified. The overall observed weighted agreement was 96.67%, while the expected agreement by chance was 66.00%. The weighted kappa coefficient was 0.902 (SE = 0.120), indicating almost perfect agreement between nurses and experts. This agreement was statistically significant (Z = 7.50, p < 0.0001), confirming the high concurrent validity of the Persian ISTAP Classification System. These results demonstrate that the Persian version reliably reproduces expert classifications in a controlled photographic setting, supporting its use in both research and clinical practice.
TABLE 2.
Classification agreement between nurses and experts.
| Nurse classification | Expert classification | Total (nurse) | ||
|---|---|---|---|---|
| 1 | 2 | 3 | ||
| 1 | 4 | 0 | 0 | 4 |
| 2 | 1 | 17 | 1 | 19 |
| 3 | 0 | 1 | 21 | 22 |
| Total (expert) | 5 | 18 | 22 | 45 |
| Observed weighted agreement | 96.67% | |||
| Expected agreement | 66.00% | |||
| Weighted kappa | 0.9020 | |||
| Standard error | 0.1202 | |||
| Z‐value | 7.50 | |||
| p value | < 0.0001 | |||
4.4. Construct Validity
The distribution of skin tear frequency differed significantly between hospitalised elderly patients and younger adults (Pearson chi‐square (4) = 15.60, p = 0.004; Fisher's exact test, p = 0.001). Among the elderly group (n = 30), the most common frequencies were 0 (33.3%), 1 (23.3%) and 2 (30.0%) ST. In contrast, the younger group (n = 30) predominantly had zero ST (76.7%), with very few patients having more than one skin tear (Table 2). This confirms that elderly patients experience significantly more frequent ST compared with younger adults.
Pearson's chi‐squared test on the classification distribution found a statistically significant difference between the two categories (p > 0.051) and this finding supports the construct validity of the Persian ISTAP (Table 3). The mean severity score for the elderly group (n = 30) was 1.97 (SD = 0.81), while the younger group (n = 30) had a mean severity of 1.53 (SD = 0.78). The difference in mean severity scores between the groups was 0.43 (95% CI: 0.02–0.84). This difference was statistically significant, t(58) = 2.12, p = 0.039 (two‐tailed), indicating that elderly patients had significantly higher severity scores compared to younger adults(Table 4).
TABLE 3.
The difference between skin tears' frequency by patient type.
| Group | Number of ST | Total | ||||
|---|---|---|---|---|---|---|
| 0 | 1 | 2 | 3 | 4 | ||
| Elderly | 10 (33.3%) | 7 (23.3%) | 9 (30.0%) | 3 (10.0%) | 1 (3.3%) | 30 |
| Younger | 23 (76.7%) | 6 (20.0%) | 1 (3.3%) | 0 (0.0%) | 0 (0.0%) | 30 |
| Total | 33 (55.0%) | 13 (21.7%) | 10 (16.7%) | 3 (5.0%) | 1 (1.7%) | 60 |
Note: Pearson chi‐square (4) = 15.60, p = 0.004; Fisher's exact test, p = 0.001.
TABLE 4.
Comparison of mean skin tear severity scores between elderly and younger patients.
| Group | N | Mean | Std. dev. | Std. error | 95% confidence interval for mean |
|---|---|---|---|---|---|
| Elderly | 30 | 1.97 | 0.81 | 0.15 | 1.66–2.27 |
| Younger | 30 | 1.53 | 0.78 | 0.14 | 1.24–1.82 |
| Difference (elderly − younger) | 0.43 | 0.20 | 0.02–0.84 |
Note: t = 2.1175, degrees of freedom = 58, p = 0.039.
4.5. Reliability
The agreement was excellent across all categories, with kappa values of 0.8531 (Z = 97.46, p < 0.001) for Category 1, 0.8170 (Z = 93.33, p < 0.001) for Category 2 and 0.8637 (Z = 98.67, p < 0.001) for Category 3. The combined overall Fleiss' kappa was 0.8447 (Z = 136.46, p < 0.001), indicating a high level of consistency among raters in classifying the photos (Table 5). These results demonstrate excellent inter‐rater reliability for the classification system used.
TABLE 5.
Inter‐rater reliability across 30 raters according to the type of skin tears.
| Outcome | Kappa | Z‐value | p |
|---|---|---|---|
| 1 | 0.8531 | 97.46 | < 0.001 |
| 2 | 0.8170 | 93.33 | < 0.001 |
| 3 | 0.8637 | 98.67 | < 0.001 |
| Combined | 0.8447 | 136.46 | < 0.001 |
According to Table 6, the inter‐rater reliability of the Persian version of the ISTAP Classification System was assessed using weighted Cohen's kappa to account for partial agreement between categories. The observed agreement between the two nurses was 87.5%, while the expected agreement by chance was 63.75%. The weighted kappa coefficient was 0.66 (SE = 0.16), indicating substantial agreement beyond chance. This result was statistically significant (Z = 3.97, p < 0.001), demonstrating good consistency between raters in classifying ST using the Persian ISTAP Classification System.
TABLE 6.
Inter‐rater reliability between two raters according to the type of skin tears of hospitalised patients.
| Nurse 2 | Total | ||||
|---|---|---|---|---|---|
| Type 1 ISTAP | Type 2 ISTAP | Type 3 ISTAP | |||
| Nurse 1 | Type 1 ISTAP | 3 | 1 | 1 | 5 |
| Type 2 ISTAP | 1 | 9 | 0 | 10 | |
| Type 3 ISTAP | 0 | 1 | 4 | 5 | |
| Total | 4 | 11 | 5 | 20 | |
The intra‐rater reliability between the two assessment times was also evaluated using weighted Cohen's kappa to account for the degree of disagreement between categories. The weighted observed agreement was 93.89%, with an expected agreement by chance of 55.68%. The weighted kappa coefficient was 0.86 (SE = 0.08), indicating excellent agreement beyond chance when considering the severity of disagreement. This result was statistically significant (Z = 10.41, p < 0.001), demonstrating strong consistency of ratings by the same raters across the two time points with consideration for partial agreement (Table 7).
TABLE 7.
Intra‐rater reliability across 30 raters according to the type of skin tears.
| Rating_Time1 | Type 1 | Type 2 | Type 3 | Total |
|---|---|---|---|---|
| Type 1 | 29 | 2 | 1 | 32 |
| Type 2 | 0 | 26 | 5 | 31 |
| Type 3 | 0 | 2 | 25 | 27 |
| Total | 29 | 30 | 31 | 90 |
4.6. Diagnostic Accuracy and Agreement Assessment
Diagnostic performance was assessed separately for Types 1, 2 and 3 ST, with expert consensus serving as the gold standard. As Table 8 indicates, for Type 1, nurses demonstrated excellent sensitivity (97.5%; 95% CI: 86.8–99.9) and specificity (99.6%; 95% CI: 97.9–100.0), with an area under the ROC curve (AUC) of 0.99 (Figure 3). The positive predictive value (PPV) of 97.5% indicates that when nurses classified a tear as Type 1, they were correct nearly all the time. The negative predictive value (NPV) of 99.6% shows strong confidence that non‐Type 1 cases were accurately excluded. The likelihood ratio for a positive test (LR+) was very high at 253.50, meaning a positive nurse classification was over 250 times more likely in true Type 1 cases than others. The low likelihood ratio for a negative test (LR– = 0.03) shows that a negative nurse classification reliably rules out Type 1 tears. The odds ratio (OR) of 10 101 further confirms strong diagnostic discrimination.
TABLE 8.
Diagnostic accuracy measures for Types 1,2 and 3 of the skin tears.
| Measure | Type 1 | Type 2 | Type 3 |
|---|---|---|---|
| Prevalence Pr(A) | 13.3% (9.7–17.7) | 47.3% (41.6–53.2) | 39.3% (33.8–45.1) |
| Sensitivity Pr(+|A) | 97.5% (86.8–99.9) | 97.9% (94.0–99.6) | 99.2% (95.4–100.0) |
| Specificity Pr(−|N) | 99.6% (97.9–100.0) | 99.4% (96.5–100.0) | 98.4% (95.3–99.7) |
| ROC area | 0.99 (0.96–1.00) | 0.99 (0.97–1.00) | 0.99 (0.98–1.00) |
| Likelihood ratio (+) | 253.50 (35.82–1793.97) | 154.66 (21.92–1091.32) | 60.15 (19.58–184.79) |
| Likelihood ratio (−) | 0.03 (0.00–0.17) | 0.02 (0.01–0.07) | 0.01 (0.00–0.06) |
| Odds ratio | 10101.00 (747.51–—) | 7274.33 (879.92–—) | 6981.00 (844.07–—) |
| PPV Pr(A|+) | 97.5% (86.8–99.9) | 99.3% (96.1–100.0) | 97.5% (92.9–99.5) |
| NPV Pr(N|−) | 99.6% (97.9–100.0) | 98.1% (94.6–99.6) | 99.4% (96.9–100.0) |
FIGURE 3.

ROC curve showing diagnostic accuracy of nurses' classification for Types 1,2 and 3 of the skin tears.
For Type 2, similarly high accuracy was found: sensitivity 97.9% (95% CI: 94.0–99.6), specificity 99.4% (95% CI: 96.5–100.0) and AUC 0.99. PPV was 99.3% and NPV 98.1%, showing high correctness in positive and negative classifications, respectively. LR+ of 154.66 indicates a very strong increase in the likelihood of true Type 2 classification when the nurse diagnosis is positive, while LR– of 0.02 suggests a negative result almost rules out Type 2. The OR of 7274 confirms excellent diagnostic strength.
For Type 3, sensitivity reached 99.2% (95% CI: 95.4–100.0) and specificity 98.4% (95% CI: 95.3–99.7), with an AUC of 0.99. PPV of 97.5% and NPV of 99.4% indicate accurate positive and negative classifications. The LR+ was 60.15, and LR– was 0.01, demonstrating very strong diagnostic value of nurse classification. An OR of 6981 further supports excellent discrimination.
Overall, the very high PPV and NPV values indicate that the ISTAP Classification System is both accurate and reliable for ruling in and ruling out each skin tear type. The large odds ratios and likelihood ratios reinforce the strong diagnostic power of the ISTAP Classification System.
5. Discussion
This study aimed to culturally adapt and psychometrically validate the Persian version of the ISTAP Classification System. The findings demonstrated excellent content validity, strong criterion and construct validity, substantial to almost perfect reliability and outstanding diagnostic accuracy. These results collectively confirm that the Persian version maintains the conceptual clarity and measurement integrity of the original ISTAP classification system.
The content validity indices indicate strong expert agreement on item clarity, necessity and relevance of the Persian version of the ISTAP classification system. Importantly, our study evaluated both qualitative and quantitative aspects of content validity, whereas some previous validations assessed only qualitative content validity [1, 4, 21, 24] or did not report a content validity evaluation step at all [22]. All studies that assessed the content validity, confirmed its validity, qualitatively. Among studies that examined content validity, all confirmed the tool's appropriateness qualitatively, supporting its conceptual clarity across different cultural contexts. In addition to the ISTAP classification system, our content validity assessment also included the definitions of ‘ST’ and ‘skin flap’ as updated by LeBlanc et al. [36] and Van Tiggelen et al. [1]. Establishing a clear definition for ‘ST’ and ‘skin flap’ is crucial, as interpretations can vary by educational background [1]. For instance, in reconstructive surgery, a ‘skin flap’ refers to tissue deliberately detached for grafting and repair [37]. An internationally accepted definition of ‘skin flap’ in the context of ‘ST’ will reduce confusion and promote best practices [38].
Criterion validity assessment in our study demonstrated an almost perfect agreement between nurses and the expert consensus. Similarly, Van Tiggelen et al. [1], in their multinational study, reported comparable levels of agreement during the diagnostic accuracy and agreement phases. Additionally, da Silva et al. confirmed concurrent criterion validity by showing a perfect total correlation (r = 1) between the Brazilian Portuguese ISTAP classification system, its adapted version, and the Skin Tear Audit Research (STAR) tool [4].
Construct validity was supported by the known‐groups comparison, which demonstrated significantly higher frequency and severity of ST among elderly patients with impaired mobility compared to younger adults without mobility limitations. To date, no prior studies on the adaptation of the ISTAP classification system have reported construct validity analyses. Our study thus provides novel evidence in this regard. This finding aligns with previous research showing that frail elderly and dependent patients experience higher rates of ST [3, 8, 39, 40, 41, 42], with advanced age, frailty and immobility, undernutrition, some chronic diseases consistently identified as key predictors of increased skin tear severity.
Reliability testing in our study showed substantial to almost perfect agreement across different raters and time points. These values align with those reported in the Van Tiggelen et al. global study with moderate to substantial agreement [1] and are similar to the Chilean Spanish with moderate to substantial agreement [21], and French Canadian validations with substantial agreement [24].
In comparison with the previous studies, our estimated reliability testing was higher; the reason may be attributable to variation in case complexity or to differences in rater clinical experience, where novice raters occasionally showed lower consistency.
The diagnostic accuracy of the Persian ISTAP was exceptionally high across all types, with sensitivity, specificity, predictive values and AUC all comparable to or exceeding those in prior validations. In the Van Tiggelen et al. study, the mean sensitivity for identifying Type 1 ST versus Types 2 and 3 was 88% (95% CI: 0.87–0.88), with a mean specificity of 92% (95% CI: 0.92–0.93). In comparison, sensitivity and specificity were somewhat reduced when differentiating Type 2 from Types 1 and 3, as well as when distinguishing Type 3 from Types 1 and 2. The Brazilian Portuguese validation reported similarly strong accuracy, though with slightly lower positive likelihood ratios, possibly reflecting sample heterogeneity [1]. The high level of agreement likely reflects the tool's user‐friendliness [19], particularly aided by clear definitions of ST and skin flaps [1].
One reason our indices exceeded those of previous studies may be that we used a standardised set of clear skin tear photos provided by the ISTAP team, selecting wounds without necrotic tissue and with well‐defined skin flaps for easier classification. Van Tiggelen et al. [1] noted that diagnostic accuracy and agreement might be higher in live assessments rather than photographic evaluations. Accurate classification requires wound cleansing, debridement of necrotic tissue and flap reapproximation—steps that are challenging to assess from photographs alone [2, 17]. Additionally, we evaluated inter‐rater reliability with hospitalised patients, recognising that clinical skin assessments may provide more accurate alternatives. Given the complex aetiology of ST [2], accurate identification and classification demand substantial knowledge and experience. Our study setting included a wound care clinic, and staff had undergone multiple wound care training sessions, providing them with foundational knowledge that likely improved classification accuracy. Additional studies are necessary to assess the impact of education and training on improving skin tear assessment and classification abilities in healthcare professionals.
5.1. Strengths and Limitations
This study has several strengths. It is the first to translate, culturally adapt and psychometrically validate the ISTAP Classification System for Persian‐speaking healthcare settings, using a comprehensive methodological approach that included both qualitative and quantitative content validity assessment. Multiple aspects of validity (content, construct and criterion) and reliability (inter‐rater and intra‐rater) were rigorously evaluated, and diagnostic accuracy indices were calculated for each classification type, providing a thorough evaluation rarely reported in previous validation studies. The inclusion of diverse participant groups, such as wound care experts and clinical nurses, enhances the generalizability of findings to real‐world practice.
However, some limitations should be considered. The study sample for construct validity was limited to hospitalised patients from selected clinical settings, which may not fully represent community or long‐term care populations. Additionally, while standardised photographs were used for part of the validation process, the classification of live patients in uncontrolled environments may present additional challenges not fully captured here. Furthermore, we used two nurses in inter‐rater reliability; however, including a larger number of nurses and observations would enhance generalizability, Finally, although the study demonstrated strong clinimetric properties, future research should evaluate the tool's impact on patient outcomes and its usability in various healthcare settings, including rural and resource‐limited environments.
6. Conclusion
The Persian version of the ISTAP Classification System demonstrated excellent content validity, strong criterion and construct validity, substantial to almost perfect reliability and outstanding diagnostic accuracy. These findings are consistent with international validation studies and confirm that the Persian adaptation retains the conceptual clarity, usability and diagnostic precision of the original instrument. By providing a standardised and culturally adapted tool, this study addresses a critical gap in wound assessment for Persian‐speaking healthcare settings. Implementation of the Persian ISTAP classification system in clinical practice can improve the accuracy of skin tear classification, enhance documentation quality and facilitate more effective prevention and management strategies. Future research should assess its impact on clinical outcomes and explore its applicability in different care environments, including community and long‐term care settings.
Funding
The authors have nothing to report.
Ethics Statement
This study was approved by the Ethics Committee of Bam University of Medical Sciences (Ethics Code: IR.MUBAM.REC.1403.098) and conducted in accordance with the ethical principles outlined in the Declaration of Helsinki. Permission to translate and culturally adapt the instrument into Persian was obtained from the ISTAP team. We appreciate their contribution in sharing standardised skin tear photographs for our study.
Consent
Before participating, all subjects were fully informed about the study's objectives, assured of data confidentiality and made aware of their right to withdraw at any time. Written informed consent was obtained from all participants prior to enrolment.
Conflicts of Interest
The authors declare no conflicts of interest.
Acknowledgements
The authors sincerely thank all participants—both patients and healthcare professionals—who generously shared their time and insights for this study. We also appreciate the expert panel members for their crucial feedback during the translation and validation stages. Special gratitude is extended to the Clinical Research Development Unit at Baqiyatallah Hospital, Baqiyatallah University of Medical Sciences, Tehran, Iran, for their ongoing support and collaboration throughout the study. Additionally, we acknowledge the assistance of ChatGPT AI (OpenAI) in enhancing the clarity and structure of the manuscript's English language.
Jafari M., Nassehi A., Dehi M., Jamshidi Z., and Jafari‐Oori M., “Validation and Clinimetric Properties of Persian Version of the ISTAP Classification System,” International Wound Journal 23, no. 2 (2026): e70800, 10.1111/iwj.70800.
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
References
- 1. Van Tiggelen H., LeBlanc K., Campbell K., et al., “Standardizing the Classification of Skin Tears: Validity and Reliability Testing of the International Skin Tear Advisory Panel Classification System in 44 Countries,” British Journal of Dermatology 183, no. 1 (2020): 146–154. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. LeBlanc K., Campbell K. E., Wood E., and Beeckman D., “Best Practice Recommendations for Prevention and Management of Skin Tears in Aged Skin: An Overview,” Journal of Wound Ostomy & Continence Nursing 45, no. 6 (2018): 540–542. [DOI] [PubMed] [Google Scholar]
- 3. Najafi‐Ghezeljeh T., Ghasemifard F., and Jafari‐Oori M., “The Effects of a Multicomponent Fall Prevention Intervention on Fall Prevalence, Depression, and Balance Among Nursing Home Residents,” Nursing and Midwifery Studies 8, no. 2 (2019): 78–84. [Google Scholar]
- 4. da Silva C. V., Campanili T. C., Freitas N. O., LeBlanc K., Baranoski S., and Santos V. L. G., “ISTAP Classification for Skin Tears: Validation for Brazilian Portuguese,” International Wound Journal 17, no. 2 (2020): 310–316. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. LeBlanc K. and Baranoski S., “Skin Tears: Finally Recognized,” Advances in Skin & Wound Care 30, no. 2 (2017): 62–63. [DOI] [PubMed] [Google Scholar]
- 6. Xu J., Xiong Y., Yan H., Zhou Z., Wen J., and Wang S., “Prevalence and Influencing Factors of Skin Tears in Older Adults: A Systematic Review and Meta‐Analysis,” Geriatric Nursing 61 (2025): 491–498. [DOI] [PubMed] [Google Scholar]
- 7. Fan S., Jiang H., Shen J., et al., “Prediction Models for Skin Tears in the Elderly: A Systematic Review and Meta‐Analysis,” Geriatric Nursing 59 (2024): 103–112. [DOI] [PubMed] [Google Scholar]
- 8. Jafari Oori M., Najafi Ghezlzah T., Mehrtak M., Nasiri K., and Aryapoor S., “The Effect of a Multidimensional Fall Prevention Program on Static and Dynamic Balance in Nursing Homes in Tehran,” Nursing and Midwifery Journal 13, no. 5 (2015): 367–376. [Google Scholar]
- 9. Strazzieri‐Pulido K. C., Peres G. R. P., Campanili T. C. G. F., and de Gouveia Santos V. L. C., “Incidence of Skin Tears and Risk Factors: A Systematic Literature Review,” Journal of Wound, Ostomy, and Continence Nursing 44, no. 1 (2017): 29–33. [DOI] [PubMed] [Google Scholar]
- 10. LeBlanc K. A., “Skin Tear Prevalence, Incidence and Associated Risk Factors in the Long‐Term Care Population” (PhD diss., Queen's University, 2017).
- 11. Ahmadizadeh L., Valizadeh L., Farshi M. R., et al., “Skin Injuries in Neonates Admitted to Three Iranian Neonatal Intensive Care Units,” Journal of Neonatal Nursing 28, no. 3 (2022): 159–163. [Google Scholar]
- 12. LeBlanc K., Baranoski S., Holloway S., Langemo D., and Regan M., “A Descriptive Cross‐Sectional International Study to Explore Current Practices in the Assessment, Prevention and Treatment of Skin Tears,” International Wound Journal 11, no. 4 (2014): 424–430. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Bernatchez S. F. and Bichel J., “The Science of Skin: Measuring Damage and Assessing Risk,” Advances in Wound Care 12, no. 4 (2023): 187–204. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Van Tiggelen H. and Beeckman D., “Skin Tears Anno 2022: An Update on Definition, Epidemiology, Classification, Aetiology, Prevention and Treatment,” Journal of Wound Management 23, no. 2 (2022): 38–51. [Google Scholar]
- 15. Payne R. and Martin M., “Defining and Classifying Skin Tears: Need for a Common Language,” Ostomy/Wound Management 39, no. 5 (1993): 16–20. [PubMed] [Google Scholar]
- 16. Carville K., Lewin G., Newall N., et al., “STAR: A Consensus for Skin Tear Classification,” Primary Intention: The Australian Journal of Wound Management 15, no. 1 (2007): 18–21, 24–28. [Google Scholar]
- 17. LeBlanc K., Baranoski S., Christensen D., et al., “International Skin Tear Advisory Panel: A Tool Kit to Aid in the Prevention, Assessment, and Treatment of Skin Tears Using a Simplified Classification System,” Advances in Skin & Wound Care 26, no. 10 (2013): 459–476. [DOI] [PubMed] [Google Scholar]
- 18. White W., “Skin Tears: A Descriptive Study of the Opinions, Clinical Practice and Knowledge Base of RNs Caring for the Aged in High Care Residential Facilities,” Primary Intention: The Australian Journal of Wound Management 9, no. 4 (2001): 138–149. [Google Scholar]
- 19. LeBlanc K., Baranoski S., Holloway S., and Langemo D., “Validation of a New Classification System for Skin Tears,” Advances in Skin & Wound Care 26, no. 6 (2013): 263–265. [DOI] [PubMed] [Google Scholar]
- 20. LeBlanc K. and Baranoski S., “Skin Tears: State of the Science: Consensus Statements for the Prevention, Prediction, Assessment, and Treatment of Skin Tears,” Advances in Skin & Wound Care 24, no. 9 (2011): 2–15. [DOI] [PubMed] [Google Scholar]
- 21. Hevia H., Ríos L., Bailey C., LeBlanc K., and Santos V. L. C. G., “Cultural Adaptation and Reliability of the ISTAP Skin Tear Classification System to Chilean Spanish,” Journal of Wound Care 30, no. 5 (2021): S16–S22. [DOI] [PubMed] [Google Scholar]
- 22. Skiveren J., Bermark S., LeBlanc K., and Baranoski S., “Danish Translation and Validation of the International Skin Tear Advisory Panel Skin Tear Classification System,” Journal of Wound Care 24, no. 8 (2015): 388–392. [DOI] [PubMed] [Google Scholar]
- 23. Källman U., Kimberly L. B., and Bååth C., “Swedish Translation and Validation of the International Skin Tear Advisory Panel Skin Tear Classification System,” International Wound Journal 16, no. 1 (2019): 13–18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Chaplain V., Labrecque C., Woo K. Y., and Le Blanc K., “French Canadian Translation and the Validity and Inter‐Rater Reliability of the ISTAP Skin Tear Classification System,” Journal of Wound Care 27, no. 9 (2018): S15–S20. [DOI] [PubMed] [Google Scholar]
- 25. Beaton D. E., Bombardier C., Guillemin F., and Ferraz M. B., “Guidelines for the Process of Cross‐Cultural Adaptation of Self‐Report Measures,” Spine 25, no. 24 (2000): 3186–3191. [DOI] [PubMed] [Google Scholar]
- 26. Ramada‐Rodilla J. M., Serra‐Pujadas C., and Delclós‐Clanchet G. L., “Cross‐Cultural Adaptation and Health Questionnaires Validation: Revision and Methodological Recommendations,” Salud Pública de México 55, no. 1 (2013): 57–66. [DOI] [PubMed] [Google Scholar]
- 27. Roebianto A., Savitri S. I., Aulia I., Suciyana A., and Mubarokah L., “Content Validity: Definition and Procedure of Content Validation in Psychological Research,” Testing, Psychometrics, Methodology in Applied Psychology 30, no. 1 (2023): 5–18. [Google Scholar]
- 28. Lawshe C. H., “A Quantitative Approach to Content Validity,” Personnel Psychology 28, no. 4 (1975): 563–575. [Google Scholar]
- 29. Lynn M. R., “Determination and Quantification of Content Validity,” Nursing Research 35, no. 6 (1986): 382–386. [PubMed] [Google Scholar]
- 30. Polit D. F. and Beck C. T., “The Content Validity Index: Are You Sure You Know What's Being Reported? Critique and Recommendations,” Research in Nursing & Health 29, no. 5 (2006): 489–497. [DOI] [PubMed] [Google Scholar]
- 31. Mokkink L. B., Terwee C. B., Patrick D. L., et al., “The COSMIN Checklist for Assessing the Methodological Quality of Studies on Measurement Properties of Health Status Measurement Instruments: An International Delphi Study,” Quality of Life Research 19, no. 4 (2010): 539–549. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Terwee C. B., Prinsen C. A., Chiarotto A., et al., “COSMIN Methodology for Evaluating the Content Validity of Patient‐Reported Outcome Measures: A Delphi Study,” Quality of Life Research 27, no. 5 (2018): 1159–1170. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Landis J. R. and Koch G. G., “The Measurement of Observer Agreement for Categorical Data,” Biometrics 33 (1977): 159–174. [PubMed] [Google Scholar]
- 34. Cohen J., Statistical Power Analysis for the Behavioral Sciences (Routledge, 2013). [Google Scholar]
- 35. Nunnally J. C., “Psychometric Theory—25 Years Ago and Now,” Educational Researcher 4, no. 10 (1975): 7–21. [Google Scholar]
- 36. LeBlanc K. and Baranoski S., “Skin Tears: The Underappreciated Enemy of Aging Skin,” Wounds International 9 (2018): 6–10. [Google Scholar]
- 37. Sun Q., He Y., Liu K., Fan S., Parrott E. P., and Pickwell‐MacPherson E., “Recent Advances in Terahertz Technology for Biomedical Applications,” Quantitative Imaging in Medicine and Surgery 7, no. 3 (2017): 345–355. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Nokaneng E., Heerschap C., and Thayer D., “Best Practice Recommendations for the Prevention and Management of Skin Tears in Aged Skin,” 2025.
- 39. Rayner R., Carville K., Leslie G., and Roberts P., “A Review of Patient and Skin Characteristics Associated With Skin Tears,” Journal of Wound Care 24, no. 9 (2015): 406–414. [DOI] [PubMed] [Google Scholar]
- 40. Koyano Y., Nakagami G., Iizaka S., et al., “Exploring the Prevalence of Skin Tears and Skin Properties Related to Skin Tears in Elderly Patients at a Long‐Term Medical Facility in Japan,” International Wound Journal 13, no. 2 (2016): 189–197. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Peres G. R. P., Bandeira da Silva C. V., Strazzieri‐Pulido K. C., and de Gouveia Santos V. L. C., “Skin Tears in Older Adult Residents of Long‐Term Care Facilities: Prevalence and Associated Factors,” Journal of Wound Care 31, no. 6 (2022): 468–478. [DOI] [PubMed] [Google Scholar]
- 42. Kaçmaz H. Y., Karadağ A., Kahraman H., Döner A., Ödek Ö., and Akın S., “The Prevalence and Factors Associated With Skin Tears in Hospitalized Older Adults: A Point Prevalence Study,” Journal of Tissue Viability 31, no. 3 (2022): 387–394. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
