Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2022 Jul 18;79(1):28–42. doi: 10.1002/jclp.23416

Florida Obsessive‐Compulsive Inventory and Children's Florida Obsessive Compulsive Inventory: A reliability generalization meta‐analysis

Alejandro Sandoval‐Lentisco 1, Rubén López‐Nicolás 1, José Antonio López‐López 1, Julio Sánchez‐Meca 1,
PMCID: PMC10084361  PMID: 35849418

Abstract

Background

The Florida Obsessive Compulsive Inventory (FOCI) and its pediatric version, the Children's Florida Obsessive Compulsive Inventory (C‐FOCI), are instruments for evaluating obsessive‐compulsive symptomatology.

Method

A reliability generalization meta‐analysis was conducted to estimate an average reliability of the scores and to identify study characteristics that explained the heterogeneity among scores. Using Kuder–Richardson 20 (KR‐20) and Cronbach's α, a total of 23 and 20 independent samples were included in the meta‐analysis for the FOCI and C‐FOCI.

Results

We found an average KR‐20 of 0.826 for the FOCI's Symptom Checklist and an α of 0.882 FOCI's Symptom Severity. An average KR‐20 of 0.740 was found for the C‐FOCI's Symptom Checklist, while an average α of 0.794 was found for the C‐FOCI's Symptom Severity. Moderator analyses showed that the source of the coefficients (i.e., whether they were reported by the authors of the primary study or estimated by the meta‐analysts) was an important variable for the FOCI Symptom Severity, and that the focus of the study (i.e., whether it was psychometric or applied) and the sample size were relevant for the C‐FOCI Symptom Checklist.

Conclusions

Considering that the FOCI and C‐FOCI are scales characterized by their brevity and ease of use, and the reliabilities obtained here, both scales are well suited for screening purposes.

Keywords: Cronbach's α coefficient, Florida Obsessive Compulsive Inventory, meta‐analysis, obsessive‐compulsive disorder, reliability generalization

1. INTRODUCTION

Obsessive‐compulsive disorder (OCD) is characterized by the presence of recurrent irrepressible thoughts and/or repeated behaviors. It is among the most frequent psychological disorders (Rosa‐Alcázar et al., 2008) and it is associated with functional impairment (Markarian et al., 2010) and a negative impact on quality of life (Subramaniam et al.,, 2013). The diagnostic and statistical manual of mental disorders (DSM‐5; American Psychiatric Association, 2013) included OCD under a new diagnostic category, “Obsessive‐compulsive and related disorders,” along with other related disorders such as body dysmorphic disorder, hoarding disorder, or trichotillomania.

There are more than 10 instruments available to measure OCD symptomatology in adults and children (McGuire et al., 2017). Among those, the Florida Obsessive Compulsive Inventory (FOCI; Storch et al., 2007) stands out from the rest of instruments because it is the only scale that briefly assesses both OCD symptoms and their severity in a simple, self‐report format, whereas other self‐report tests take longer to complete and/or only assess the presence of symptoms and not their severity. The FOCI comprises two scales: Symptom Checklist and Symptom Severity. Symptom Checklist is a dichotomous scale that evaluates the presence/absence of 10 common obsessions and 10 common compulsions, whereas Symptom Severity consists of 5 items that use a 5 ‐point scoring system ranging from 0 to 4, with higher scores reflecting greater severity. Furthermore, the Symptom Severity scale is not content‐dependent, which makes FOCI a more sensitive instrument to measure less common OCD symptoms and, recently, it was proven to be sensitive to clinical change as well (Veale et al., 2021). More recently, Storch et al. (2009) developed a pediatric version of the scale, the Children's Florida Obsessive Compulsive Inventory (C‐FOCI), which also includes the Symptom Checklist and Symptom Severity subscales. The first of these C‐FOCI subscales measures the presence/absence of 17 obsessions and compulsions that are frequent among young people with OCD, whereas the latter uses the same five items as the original FOCI, with small vocabulary changes. The FOCI and the C‐FOCI have been adapted and validated to other languages such as Thai (Saipanish et al., 2015), Chinese (Cao et al., 2021), and Spanish (Piqueras et al., 2016).

In its original validation study, the FOCI showed good internal consistency. Measured with the Kuder–Richardson 20 (KR‐20) formula, the reliability of the Symptom Checklist scores was 0.83 and using Cronbach's α for the Symptom Severity, score reliability was 0.89. Furthermore, scores showed good concurrent validity as they correlated, respectively, moderately and strongly (Symptom Checklist: r = 0.47, Symptom Severity: r = 0.89) with another OCD scale, the Yale‐Brown Obsessive‐Compulsive Scale (Y‐BOCS; Goodman et al., 1989), and also moderately correlated with the Beck Depression Inventory (Beck et al., 1961; with Symptom Checklist: r = 0.40, with Symptom Severity: r = 0.70), which can be expected to be often present in OCD (cf. Storch et al., 2007). Similarly, in the original validation study of the C‐FOCI, scores showed an acceptable internal consistency, with KR‐20 = 0.76 for Symptom Checklist and α = 0.79 for Severity Scale. Likewise, C‐FOCI moderately correlated with other OCD scales (Children's Yale‐Brown Obsessive Compulsive Scale [CY‐BOCS]; Scahill et al., 1997; with Symptom Checklist: r = 0.318, with Symptom Severity: r = 0.498) and scales of depression (Children's Depression Inventory Short Form; Kovacs, 1992; with Symptom Checklist: r = 0.351, with Symptom Severity: r = 0.405) and anxiety (Multidimensional Anxiety Scale for Children; March et al., 1997; with Symptom Checklist r = 0.607, with Symptom Severity r = 0.396; Storch et al., 2009).

Although in the literature we often find misleading statements such as “this test is reliable,” reliability, however, is not an inherent property of a test and it can vary each time the test is applied to a different sample of participants (Crocker & Algina, 1986; Streiner & Norman, 2008). Thus, anytime a study makes use of a scale, authors should report a reliability estimate with the data at hand (Appelbaum et al., 2018). Unfortunately, researchers often fail to report reliability estimates based on their own participants' scores, and instead it is common to find references to the reliability obtained in the original test validation study. With this practice, which is known as reliability induction, researchers erroneously assume that their data will offer the same reliability. Previous studies have reported reliability induction rates for psychological tests to be between 54% and 78.6% (Sánchez‐Meca et al., 2015; Vacha‐Haase et al., 1999; Whittington, 1998). Checking the reliability of the scores of newer test applications is important not only to ensure that the measure itself is reliable but also because reliability affects the effect sizes. If the test scores obtain a lower reliability, it may attenuate effect sizes that are calculated with that instrument (Vacha‐Haase et al., 2000).

Nevertheless, we can use meta‐analytic methods to obtain a representative reliability value for an instrument by integrating different reliability estimates obtained across studies. This is often referred to as reliability generalization (RG; Vacha‐Haase, 1998). Furthermore, if there is heterogeneity across reliability estimates based on the same test, an RG meta‐analysis allows us to examine whether some characteristics of the studies (i.e., moderators) might account for the variability of the reliability coefficients (Henson & Thompson, 2002; Rodriguez & Maeda, 2006; Sánchez‐Meca et al., 2013). Examples of study characteristics that can affect the reliability are the mean and variability of test scores, the target population (community, undergraduate, clinical), or whether the original version or an adaptation (to different idioms, cultures, or countries) of the test was applied.

1.1. Objectives

The aim of this study was to conduct an RG meta‐analysis to (a) estimate an average reliability for the FOCI and C‐FOCI by integrating all the studies that have applied the scales and reported a reliability estimate with the data at hand; (b) assess whether there is a large heterogeneity between reliability estimates for the same instrument and, if so, perform moderator analyses to identify study characteristics that account for such variability, and (c) estimate the reliability induction rate for the FOCI and C‐FOCI.

2. METHOD

Review methods and reporting were performed according to the Reliability Generalization Meta‐Analysis (REGEMA) guidelines (Sánchez‐Meca et al., 2021). The REGEMA checklist for the present meta‐analysis is available at https://osf.io/zdhvk/. We did not preregister a study protocol.

2.1. Participants

2.1.1. Literature search

To identify studies that could be relevant for the RG meta‐analysis, we systematically searched Google Scholar, ProQuest, PsycInfo, and PubMed from 2007, the year the FOCI was published, to November 2020. We used the term “Florida Obsessive Compulsive Inventory” as the keyword to be found anywhere in the article. With this keyword we were able to identify studies that applied both the FOCI and the C‐FOCI. In addition, references cited in studies that had applied any of these scales were also consulted to identify additional studies that applied the FOCI or the C‐FOCI.

2.1.2. Eligibility criteria

To be included in this RG meta‐analysis, studies had to meet three criteria: (a) to be an empirical study that applied the FOCI or the C‐FOCI to one or more samples of participants; (b) to report any reliability estimate (internal consistency, temporal stability, interrater agreement) with the data at hand or to include enough information to allow us calculate them manually; and (c) to be written in either English or Spanish language. In addition, for studies that applied the FOCI or the C‐FOCI and did not report any reliability estimate with the data at hand, the corresponding author was contacted and, if they sent any reliability estimate, the study was also included in the meta‐analysis. N = 1 or case series studies were excluded. Participant samples could belong to any kind of target population (community, clinical, or subclinical) and papers could be either published or unpublished (e.g., a PhD thesis dissertation).

2.2. Instruments

A coding form was developed to extract relevant study characteristics and reliability estimates. The coding form is available at https://osf.io/yqvhs/ and contains information about all coded variables.

2.3. Procedure

For the reliability estimates, we extracted the KR‐20 for the Symptom Checklist and Cronbach's α for the Symptom Severity subscales. Besides, for those that had a reliability estimate available, we extracted data on the following potential moderators: (a) mean and standard deviation of the scores in each scale; (b) mean and standard deviation of participants' age; (c) study sample size; (d) gender distribution (% male); (e) publication year. (f) sample ethnicity (% Caucasian); (g) target population (community, undergraduate students, clinical, or mixed); (h) location of the study (country, continent); (i) test version (original or other); (j) administration format (paper‐and‐pencil or online); (k) study aim (psychometric or applied); (l) if psychometric, the focus of the psychometric study (FOCI or other). The coding form was applied to the studies that reported any reliability estimate of the FOCI or the C‐FOCI with the data at hand. In addition, from studies that did not report reliability estimates several characteristics of the samples were also extracted: target population (clinical vs. nonclinical), mean and SD of total test scores, mean age, gender distribution (% male), and ethnicity (% of Caucasians). These variables were extracted to conduct a sensitivity analysis to explore whether the composition and variability of the inducing and reporting studies were similar. The information extraction was carried out by one author (A. S. L.).

2.4. Data analysis

Separate meta‐analyses were conducted for each subscale of the FOCI and C‐FOCI tests. Random‐effects models were assumed as heterogeneity among the reliability estimates was expected. The reliability coefficients were weighted by the inverse variance, using the restricted maximum likelihood method to estimate the between‐studies variance. The 95% confidence bounds were computed according to the improved method proposed by Hartung and Knapp (2001). To normalize the sampling distribution and to stabilize the sampling variances of the reliability coefficients, we applied Bonett's transformation (Bonett, 2002) to the internal consistency reliability coefficients: L i  = ln(1 − |α i |), where αi denotes the reliability coefficient of each study, ln is the natural logarithm, and L i is the transformed coefficient. After the statistical integration, the L i coefficients were back‐transformed to the reliability coefficient metric (αi  = 1 − eLi) to facilitate interpretation and, as a sensitivity analysis, we also conducted the meta‐analyses with the untransformed coefficients (Sánchez‐Meca et al., 2013).

Heterogeneity among reliability estimates was assessed in each meta‐analysis using the Q statistic, the I 2 index and the calculation of 95% credibility intervals to estimate a plausible range of reliability estimates for each of the (or any future) studies (Riley et al., 2011). I 2 values of approximately 25%, 50%, and 75% have been suggested before as reflecting low, moderate, and large heterogeneity, respectively (Higgins et al., 2003), although such benchmarks were proposed in the context of clinical trials and hence they might not be applicable to the context of RG studies. If there was evidence of heterogeneity, we performed moderator analyses to identify which of the study characteristics accounted for this heterogeneity. We applied mixed‐effects meta‐regression models for continuous variables and subgroup analyses for the categorical variables with the improved method proposed by Knapp and Hartung (2003). R 2 indices were calculated to estimate the amount of variance accounted for by each moderator (López‐López et al., 2014). Furthermore, we selected the moderators that were either statistically significant or explained at least 10% of the variance (i.e., R 2 > 0.10) and made moderator models with all of their possible combinations for every scale that presented considerable heterogeneity. To compare which of the models fits best the data, we made use of an Information Theory approach based on the Akaike Information Criterion (AIC). For each model, we calculated the corrected version of the AIC (AICC), recommended for small sample sizes (Symonds & Moussalli, 2011), using the maximum likelihood estimate. All statistical analyses were conducted using the metafor package in R (Viechtbauer, 2010) and the script is available at https://osf.io/65zpj/.

3. RESULTS

3.1. Study selection

Figure 1 summarizes the entire selection process through a flowchart. A total of 440 references were initially identified. We then removed duplicated results and, based on abstract and title, excluded references that were not empirical studies. Next, we full‐text reviewed 189 references, from which 133 did not apply the FOCI or C‐FOCI. Forty‐two articles applied the FOCI, from which 13 studies (30.9%) reported some reliability estimate with the data at hand and were included in the meta‐analysis. The remaining 29 studies (69.1%) induced the reliability from previous applications of the test (by omission or by report). However, four studies reported the mean and standard deviation of the scores, enabling us to use the KR‐20 formula for the Symptom Checklist subscale, and we also obtained reliability estimates from further three studies after contacting authors via email, including those in the meta‐analysis as well and adding up to a total of 20 studies and 23 independent samples for the FOCI. Furthermore, 14 studies applied the C‐FOCI, from which 11 studies (78.6%) reported some reliability estimate with the data at hand and were included in the meta‐analysis. The remaining three studies (21.4%) induced reliability from previous applications of the test (by omission or by report). However, one study reported the mean and standard deviation of the scores, allowing us to calculate KR‐20 for the Symptom Checklist subscale, and one author sent us a reliability estimate, so these studies were also included in the meta‐analysis, adding up to a total of 13 studies with 21 independent samples. References from the studies that were included in the analyses are available at https://osf.io/6tvna/. The total sample size for the RG study of the FOCI was N = 4076 for the Symptom Checklist and N = 5898 for the Symptom Severity subscales, with range 18–986 (mean = 226; SD = 260) and 47–1224 (mean = 281; SD = 305), respectively. Most studies were carried out in North America (52.2%), followed by Asia (26.1%), Oceania (13.0%), and Europe (8.7%). The majority of studies included only people with a clinical condition (56.52%), followed by undergraduate/graduate nonclinical students (34.78%), nonundergraduate/graduate people without any known clinical condition (4.35%) and a mix of them (4.35%).

Figure 1.

Figure 1

Flowchart of the study selection process

With regard to the C‐FOCI, the total sample size was N = 11,327 for the Symptom Checklist and N = 11,276 for the Symptom Severity subscales, with range 51–4293 (mean = 566; SD = 937) and 82–4293 (mean = 593; SD = 955), respectively. Most studies were carried out in Europe (65%), followed by North America (20%), South America (5%), Asia (5%), and Oceania (5%). The majority of studies included only people with a clinical condition (65%), followed by nonundergraduate/graduate people without any known clinical condition (25%) and a mix of them (10%).

Finally, concerning the reliability induction rates, 69.1% of the studies that applied the FOCI induced the reliability by omission (not making any reference to the reliability of the scale) or by report (explicitly making reference to the reliability shown in previous applications). Regarding the C‐FOCI, 21.4% of the studies that applied it induced the reliability by omission or report. The database can be found at https://osf.io/pnfwq/.

3.2. Mean reliability and heterogeneity

Although our purpose was to synthesize all types of reliability coefficients, studies that applied the FOCI or the C‐FOCI only reported α or KR‐20 coefficients. Thus, our results were based on these types of internal consistency coefficients. Table 1 presents the main results for each of the FOCI and C‐FOCI subscales. The number of reliability coefficients extracted from the studies ranged from 15 coefficients for FOCI's Symptom Severity subscale to 20 coefficients for C‐FOCI's Symptom Checklist subscale. Because similar results were obtained using transformed or untransformed coefficients, only the results with the untransformed coefficients are presented here. Figures 2, 3, 4, 5 display forest plots of the reliability coefficients of each meta‐analyzed subscale.

Table 1.

Mean reliability coefficients, 95% confidence and credibility/prediction intervals, and heterogeneity statistics for each subscale

Total scale k α + 95% CI 95% Cr.I Q I 2
LL UL LL UL
FOCI Checklist 17 0.826 0.815 0.838 0.794 0.859 26.243 40.57
FOCI Severity 15 0.882 0.861 0.903 0.816 0.948 118.633* 90.04
C‐FOCI Checklist 20 0.740 0.714 0.765 0.640 0.840 201.804* 89.40
C‐FOCI Severity 19 0.794 0.744 0.844 0.582 1a 319.149* 98.50

Abbreviations: α +, mean coefficient α; I 2, heterogeneity index; k, number of studies; LL and UL, lower and upper limits of the 95% confidence and credibility/prediction interval for α +; Q, Cochran's heterogeneity Q statistic.

a

The limit estimated (1.007) exceeded the range of an α coefficient, therefore the value was truncated at 1.

*

p < 0.0001.

Figure 2.

Figure 2

Forest plot of the KR‐20 coefficients reported for FOCI Symptom Checklist. The lines at both sides of the diamond show the limits of the 95% prediction interval. FOCI, Florida Obsessive Compulsive Inventory; KR‐20, Kuder–Richardson 20.

Figure 3.

Figure 3

Forest plot of the α coefficients reported for FOCI Symptom Severity. The lines at both sides of the diamond show the limits of the 95% prediction interval. FOCI, Florida Obsessive Compulsive Inventory.

Figure 4.

Figure 4

Forest plot of the KR‐20 coefficients reported for C‐FOCI Symptom Checklist. The lines at both sides of the diamond show the limits of the 95% prediction interval. C‐FOCI, Children's Florida Obsessive Compulsive Inventory; KR‐20, Kuder–Richardson 20.

Figure 5.

Figure 5

Forest plot of the α coefficients reported for C‐FOCI Symptom Severity. The lines at both sides of the diamond show the limits of the 95% prediction interval. C‐FOCI, Children's Florida Obsessive Compulsive Inventory.

Regarding FOCI's Symptom Checklist, a mean KR‐20 coefficient of 0.826 (95% CI [0.815, 0.838]) was found, ranging from 0.78 to 0.86. Although the Q statistic was not significant, there was some evidence of heterogeneity among the true reliability coefficients (Q(16) = 26.24, p = 0.064, I 2  = 40.57, τ 2  = 0.0002, 95% CrI [0.794, 0.859]). For FOCI's Symptom Severity, a mean α coefficient of 0.882 (95% CI [0.861, 0.903]) was found, ranging from 0.83 to 0.92; the Q statistic indicated heterogeneity among the true reliability coefficients (Q(14) = 118.63, p < 0.0001, I 2  = 90.04, τ 2  = 0.0008, 95% Cr.I [0.816, 0.948]). For C‐FOCI's Symptom Checklist, a mean KR‐20 coefficient of 0.74 (95% CI [0.714, 0.765]) was found, ranging from 0.61 to 0.82; the Q statistic reflected heterogeneity among the true reliability coefficients (Q(19) = 201.80, p < 0.0001, I 2  = 89.4, τ 2  = 0.0021, 95% Cr.I [0.640, −840]). For C‐FOCI's Symptom Severity, a mean α coefficient of 0.794 (95% CI [0.744, 0.844]) was found, ranging from 0.49 to 0.89; the Q statistic indicated heterogeneity among the true reliability coefficients (Q(18) = 319.15, p < 0.0001, I 2  = 98.50, τ 2  = 0.0097, 95% Cr.I [0.582, 1]).

3.3. Analysis of moderator variables

Supporting Information: Tables 1–4 available at https://osf.io/hgzmk  present the results of the meta‐regression models applied for the continuous moderators for the FOCI and the C‐FOCI. None of them exhibited a statistically significant relationship with the reliability coefficients. Nonetheless, for the C‐FOCI Symptom Checklist, there was weak evidence of sample size and standard deviation of the scores being important moderators (see Supporting Information: Table S3). In particular, the larger the sample size and the SD of test scores, the larger the reliability.

Weighted analysis of variances on the transformed reliability coefficients for each categorical moderator variable were also carried out. Full moderator analysis results are available at https://osf.io/hgzmk/. There was strong evidence of a difference when comparing the mean reliability coefficients grouped by the source from where the coefficient was obtained (reported in the paper vs. obtained from the authors upon request) for FOCI Symptom Severity (F (1, 13) = 9.492; p = 0.002; R 2  = 0.41; Supporting Information: Table S6). Studies that reported the coefficients exhibited higher reliability coefficients (α + = 0.89; k = 13) than studies that did not report it and had to be obtained upon request (α + = 0.80; k = 2). In addition, we found weak evidence of differences when comparing the mean reliability coefficients grouped by the main scale analyzed in the studies with a psychometric focus (C‐FOCI vs. other) for C‐FOCI Symptom Checklist (F (1, 10) = 4.681; p = 0.056; R 2  = 0.26; Supporting Information: Table S7). Studies whose psychometric focus was the C‐FOCI had higher reliability coefficients (α+ = 0.78; k = 5) than studies whose psychometric focus was another scale (α+ = 0.72; k = 7).

3.4. Model comparison

From all the possible moderator combinations, only the five best models (according to the AICC) are presented (see Table 2). Because no moderator was statistically significant nor explained more than 10% of the variance for the FOCI Symptom Checklist, we did not perform any model comparison for this scale. For the FOCI Symptom Severity, the model that best fits our data was the one including the test adaptation and the source of the estimates as moderators, with an Akaike weight of 0.419. Models containing the source only and test adaptation and source as moderators achieved similar but slightly worse fit, all of which contained the source as a moderator. For C‐FOCI Symptom Checklist, the model with the lowest AICC was the one containing the sample size only as moderator, with an Akaike weight of 0.55, achieving slightly better fit than the model with the intercept only. Last, for C‐FOCI Symptom Severity, the best model was the one containing gender distribution (% male) as moderator, also achieving slightly better fit than the model with the intercept only.

Table 2.

Five best‐ranked models for each of the scales according to AICC

Candidate models df logLik AICc
i
wi
FOCI Symptom Severity
Ver + Src 4 2.242 7.5 0 0.419
Src 3 0.172 7.8 0.32 0.356
Cnt + Src 5 3.169 10.3 2.81 0.103
Ppl + Src 5 3.015 10.6 3.12 0.088
Ver 3 −2.161 12.5 4.99 0.035
C‐FOCI Symptom Checklist
Siz 3 6.422 −5.3 0 0.550
Int 2 4.441 −4.2 1.17 0.307
Frm 4 6.449 −2.2 3.11 0.116
Frm + Siz 5 6.582 1.1 6.47 0.022
Cnt a 6 7.327 3.8 9.15 0.006
C‐FOCI Symptom Severity
Gen 3 −7.895 23.8 0
Int 2 −10.72 26.2 2.4

Abbreviations: Cnt, continent; Frm, administration format; Gen, gender distribution (% males); Int, intercept only; Ppl, target population; Siz, sample size; Src, estimates source; Ver, test version.

a

Although Continent exhibited an R 2 = 0.09, it was also included.

4. DISCUSSION

We performed an RG meta‐analysis to characterize how reliability of the FOCI and the C‐FOCI varies across studies. The FOCI showed an average reliability of 0.826 for the Symptom Checklist subscale and 0.882 for the Symptom Severity subscale, whereas the C‐FOCI showed an average reliability of 0.740 for the Symptom Checklist subscale and 0.794 for the Symptom Severity subscale. Acording to Cicchetti (1994), these values can be considered as good (FOCI) and fair (C‐FOCI) reliability coefficients, respectively. If we compare these values to those obtained in other RG meta‐analyses of other OCD scales such as the Y‐BOCS, the original Padua Inventory (PI) and their two revisions (the Padua Inventory Revision [PI–R], and the Padua Inventory–Washington State University Revision [PI–WSUR]), the Maudsley Obsessional Compulsive Inventory and the Dimensional Obsessive‐Compulsive Scale (López Nicolás et al., 2021; López‐Pina et al., 2015; Núñez‐Núñez et al., 2022; Rubio‐Aparicio et al., 2020; Sánchez‐Meca et al., 20112017), we see that they have similar or slightly better psychometric properties. Nonetheless, we want to again emphasize that the FOCI and C‐FOCI are the only OCD scales that briefly examine both OCD symptoms and their severity, while also being brief, simple, and self‐reported. The average reliabilities obtained here are above 0.80 for the FOCI and 0.70 for the C‐FOCI. Considering the guidelines given by Nunnally and Bernstein (1994), where 0.70 is considered an acceptable standard for instruments used in basic research and quick assessments, 0.80 is suitable for research comparing the scores of different groups, and 0.90 is needed for clinical settings where diagnostics or other important decisions are made, the average reliabilities obtained in the present study make both scales seem especially relevant to be applied on community population for screening purposes of OCD symptoms and, additionally for the FOCI, also to be applied when conducting a research involving group comparisons.

A large heterogeneity among coefficients was found for both FOCI subscales and in the C‐FOCI Symptom Checklist subscale so we performed moderator analyses to identify which study characteristics could be explaining this variability. For continuous moderators, we found that none of them were statistically associated with the reliability coefficients, although there was some evidence that sample size and standard deviation of the scores accounted for a substantial part of the variability among the reliability estimates. The standard deviation of the scores has been previously found to be a source of systematic variation of the reliability coefficients (Botella et al., 2010). Thus, what might be surprising here is that, even though in many cases the standard deviation explained a considerable portion of the variance, it did not reach statistical significance. Considering the scarce number of studies that have reported the standard deviation of the test scores, this lack of statistical significance might be due to a low statistical power. Another aspect to take into consideration is the fact that most studies only applied the FOCI or C‐FOCI to a particular sample (either people with a clinical condition or without one, but not both). This might be artificially restricting the standard deviations, making them very similar across studies and, therefore, preventing their potential role as a moderator variable. For the categorical moderators, we found that, in the FOCI's Symptom Severity subscale, the source from where the coefficient was obtained (either reported in the paper or requested to the authors) was a significant moderator. Notably, studies that reported the coefficients showed higher reliability compared to studies where our team requested them from the authors. This finding can be interpreted as indicating the existence of reporting bias of the reliability coefficients, such that studies are less likely to report the reliability coefficient obtained with the data at hand when the observed value is small. We also found that, in C‐FOCI's Symptom Checklist, studies whose psychometric focus was the C‐FOCI showed higher reliability coefficients than studies whose psychometric focus was another scale. Lastly, another remarkable result was that the administration format was not a significant moderator in any of the scales, that is, there was no difference between the reliability estimates when they were obtained in a pencil and paper format compared to when they were obtained online. Thus, both the FOCI and the C‐FOCI might be used as online screening tools to assess OCD symptoms easily and quickly.

Model comparison also showed that, for the C‐FOCI Symptom Checklist, sample size might be a relevant moderator of the reliability coefficients, with a positive association between sample size and the magnitude of the reliability coefficients. Last, for the C‐FOCI Symptom Severity, gender might also be important as a moderator, since the model containing gender achieved the best fit. Specifically, there appears to be a negative trend toward the percentage of males in the sample and the reliability scores.

Finally, we also calculated the reliability induction rate for both scales. For the FOCI, we found a rather high induction rate (69.05% of the studies induced the reliability). This finding is in line with previous appraisals of this practice (Sánchez‐Meca et al., 2015; Vacha‐Haase et al., 1999; Whittington, 1998). Considering that this scale is rather recent, this result suggests that the practice of inducing reliability of test scores from previous studies is still extended. On the other hand, only 21.43% of the studies that applied the C‐FOCI induced the reliability by omission or report. This considerably lower rate of induction (compared to the FOCI and other scales) might be due to the fact that half of the studies that applied the C‐FOCI had a psychometric focus (by contrast, only 5.25% of the studies that applied the FOCI had a psychometric focus), and it is reasonable to expect that psychometric studies of a test will calculate and report reliability coefficients.

4.1. Limitations

Although many studies have applied the FOCI or C‐FOCI in their research, the number of studies that have reported a reliability estimate with the data at hand is considerably smaller. This, together with the fact that missing data on potential moderator variables was common, may have limited the results of the moderator analyses. Another limitation was the failure of authors to report important study characteristics whose potential influence on the reliability estimates can be investigated in an RG meta‐analysis. In particular, many studies did not report the mean and standard deviation of test scores, two very important moderators in the context of RG studies.

On another note, this RG meta‐analysis was based on Cronbach's α (and KR‐20, a special case of α). Because Cronbach's α does not require multiple test administrations to estimate the reliability of a scale, it is the most commonly reported reliability coefficient (Hogan et al., 2000; Scherer & Teo, 2020). However, using Cronbach's α as a measure of reliability presents some drawbacks. First, this coefficient is not invariant to the numbers of items in the scale, larger number of items will inflate Cronbach's α (Tavakol & Dennick, 2011). Second, it relies on some assumptions that are usually not met in practice (τ‐equivalence, uncorrelated item errors, unidimensionality, etc.), resulting in an over or underestimation of the reliability (McNeish, 2018). For this reason, Cronbach's α is being criticized and it is advised that future primary studies should, instead, make use of composite reliability measures such as the omega coefficients or measures of maximal reliability such as the coefficient H, allowing future RG meta‐analyses to be based on those coefficients (Scherer & Teo, 2020). Nonetheless, studies published to date have seldom used these alternative reliability coefficients, so an RG meta‐analysis based on them is not yet feasible for most measurement instruments.

5. CONCLUSIONS

Our findings show that the FOCI and C‐FOCI instruments present, on average, good and acceptable reliability values, respectively. Given the brevity and ease of use of these two scales and that they are able to simultaneously assess OCD symptoms and their severity, both tests can be recommended for the screening of OCD symptoms. For the FOCI, we found a large rate of reliability induction, showing how this practice is still extended even in recent empirical studies using psychometric tests. For the C‐FOCI, we found a much lower rate of reliability induction, although this result is probably accounted by the fact that half of the studies that applied the C‐FOCI had a psychometric focus. Thus, it is still necessary to emphasize that studies that make use of a psychometric scale should estimate reliability with the data at hand, rather than insisting on the erroneous practice of reporting estimates obtained in previous applications of the instrument.

CONFLICT OF INTEREST

The authors declare no conflict of interest.

PEER REVIEW

The peer review history for this article is available at https://publons.com/publon/10.1002/jclp.23416

Supporting information

Supplementary Information

ACKNOWLEDGMENTS

This study was funded by Agencia Estatal de Investigación (Government of Spain) and by FEDER Funds, AEI/10.13039/501100011033, grants no. PID2019‐104033GA‐I00 and PID2019‐104080GB‐I00.

Sandoval‐Lentisco, A. , López‐Nicolás, R. , López‐López, J. A. , & Sánchez‐Meca, J. (2023). Florida Obsessive‐Compulsive Inventory and Children's Florida Obsessive Compulsive Inventory: A reliability generalization meta‐analysis. Journal of Clinical Psychology, 79, 28–42. 10.1002/jclp.23416

DATA AVAILABILITY STATEMENT

The data that support the findings of this study are openly available at https://osf.io/zxn2k/.

REFERENCES

  1. American Psychiatric Association . (2013). Diagnostic and statistical manual of mental disorders (5th ed.). 10.1176/appi.books.9780890425596 [DOI]
  2. Appelbaum, M. , Cooper, H. , Kline, R. B. , Mayo‐Wilson, E. , Nezu, A. M. , & Rao, S. M. (2018). Journal article reporting standards for quantitative research in psychology: The APA Publications and Communications Board task force report. American Psychologist, 73(1), 3. 10.1037/amp0000191 [DOI] [PubMed] [Google Scholar]
  3. Beck, A. T. , Ward, C. H. , Mendelson, M. , Mock, J. , & Erbaugh, J. (1961). An inventory for measuring depression. Archives of General Psychiatry, 4(6), 561–571. 10.1001/archpsyc.1961.01710120031004 [DOI] [PubMed] [Google Scholar]
  4. Bonett, D. G. (2002). Sample size requirements for testing and estimating coefficient alpha. Journal of Educational and Behavioral Statistics, 27(4), 335–340. 10.3102/10769986027004335 [DOI] [Google Scholar]
  5. Botella, J. , Suero, M. , & Gambara, H. (2010). Psychometric inferences from a meta‐analysis of reliability and internal consistency coefficients. Psychological Methods, 15(4), 386. 10.1037/a0019626 [DOI] [PubMed] [Google Scholar]
  6. Cao, X. , Gao, R. , Liu, Y. , Zhou, Y. , Wang, J. , Chen, Y. , Wang, Z. , Guzick, A. G. , Goodman, W. K. , Storch, E. A. , & Fan, Q. (2021). The reliability and validity of the florida obsessive‐compulsive Inventory in a chinese clinical sample. Journal of Obsessive‐Compulsive and Related Disorders, 28, 100623. 10.1016/j.jocrd.2021.100623 [DOI] [Google Scholar]
  7. Cicchetti, D. V. (1994). Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessments instruments in psychology. Psychological Assessment, 6, 284–290. 10.1037/1040-3590.6.4.284 [DOI] [Google Scholar]
  8. Crocker, L. M. , & Algina, J. (1986). Introduction to classical and modern test theory. Holt, Rinehart & Winston. [Google Scholar]
  9. Goodman, W. K. , Price, L. H. , Rasmussen, S. A. , Mazure, C. , Fleischmann, R. L. , Hill, C. L. , & Charney, D. S. (1989). The Yale‐Brown Obsessive Compulsive Scale: I. Development, use, and reliability. Archives of General Psychiatry, 46(11), 1006–1011. 10.1001/archpsyc.1989.01810110048007 [DOI] [PubMed] [Google Scholar]
  10. Hartung, J. , & Knapp, G. (2001). On tests of the overall treatment effect in the meta‐analysis with normally distributed responses. Statistics in Medicine, 20, 1771–1782. 10.1002/sim.791 [DOI] [PubMed] [Google Scholar]
  11. Henson, R. K. , & Thompson, B. (2002). Characterizing measurement error in scores across studies: Some recommendations for conducting “reliability generalization” studies. Measurement and Evaluation in Counseling and Development, 35(2), 113–127. 10.1080/07481756.2002.12069054 [DOI] [Google Scholar]
  12. Higgins, J. P. T. (2003). Measuring inconsistency in meta‐analyses. BMJ, 327(7414), 557– 560. 10.1136/bmj.327.7414.557 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Hogan, T. P. , Benjamin, A. , & Brezinski, K. L. (2000). Reliability methods: A note on the frequency of use of various types. Educational and Psychological Measurement, 60(4), 523–531. 10.1177/00131640021970691 [DOI] [Google Scholar]
  14. Knapp, G. , & Hartung, J. (2003). Improved tests for a random effects meta‐regression with a single covariate. Statistics in Medicine, 22(17), 2693–2710. 10.1002/sim.1482 [DOI] [PubMed] [Google Scholar]
  15. Kovacs, M. (1992). Children's depression inventory manual. Multi‐Health Systems. [Google Scholar]
  16. López‐López, J. A. , Marín‐Martínez, F. , Sánchez‐Meca, J. , Van den Noortgate, W. , & Viechtbauer, W. (2014). Estimation of the predictive power of the model in mixed‐effects meta‐regression: A simulation study. British Journal of Mathematical and Statistical Psychology, 67(1), 30–48. 10.1111/bmsp.12002 [DOI] [PubMed] [Google Scholar]
  17. López‐ Nicolás, R. , Rubio‐ Aparicio, M. , López‐ Ibáñez, C. , & Sánchez‐ Meca, J. (2021). A reliability generalization meta‐analysis of the Dimensional Obsessive‐Compulsive Scale. Psicothema, 33, 481–489. 10.7334/psicothema2020.455 [DOI] [PubMed] [Google Scholar]
  18. López‐Pina, J. A. , Sánchez‐Meca, J. , López‐López, J. A. , Marín‐Martínez, F. , Núñez‐Núñez, R. M. , Rosa‐Alcázar, A. I. , Gómez‐Conesa, A. , & Ferrer‐Requena, J. (2015). The yale–brown obsessive compulsive scale: A reliability generalization meta‐analysis. Assessment, 22(5), 619–628. [DOI] [PubMed] [Google Scholar]
  19. March, J. S. , Parker, J. D. , Sullivan, K. , Stallings, P. , & Conners, C. K. (1997). The Multidimensional Anxiety Scale for Children (MASC): Factor structure, reliability, and validity. Journal of the American Academy of Child & Adolescent Psychiatry, 36(4), 554–565. 10.1097/00004583-199704000-00019 [DOI] [PubMed] [Google Scholar]
  20. Markarian, Y. , Larson, M. J. , Aldea, M. A. , Baldwin, S. A. , Good, D. , Berkeljon, A. , & McKay, D. (2010). Multiple pathways to functional impairment in obsessive–compulsive disorder. Clinical Psychology Review, 30(1), 78–88. 10.1016/j.cpr.2009.09.005 [DOI] [PubMed] [Google Scholar]
  21. McGuire, J. F. , Storch, E. A. , & Goodman, W. (2017). Clinical rating scales for OCD. In Pittenger C. (Ed.), Obsessive‐compulsive disorder: Phenomenology, pathophysiology, and treatment (pp. 137–146). Oxford University Press. [Google Scholar]
  22. McNeish, D. (2018). Thanks coefficient alpha, we'll take it from here. Psychological Methods, 23(3), 412. 10.1037/met0000144 [DOI] [PubMed] [Google Scholar]
  23. Núñez‐Núñez, R. M. , Rubio‐Aparicio, M. , Marín‐Martínez, F. , Sánchez‐Meca, J. , López‐Pina, J. A. , & López‐López, J. A. (2022). A reliability generalization meta‐analysis of the padua inventory‐revised (PI‐R). International Journal of Clinical and Health Psychology, 22(1), 100277. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Nunnally, J. C. , & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw‐Hill. [Google Scholar]
  25. Piqueras, J. A. , Rodríguez‐Jiménez, T. , Ortiz, A. G. , Moreno, E. , Lázaro, L. , & Storch, E. A. (2016). Factor structure, reliability, and validity of the spanish version of the children's florida obsessive compulsive Inventory (C‐FOCI). Child Psychiatry & Human Development, 48(1), 166–179. 10.1007/s10578-016-0661-4 [DOI] [PubMed] [Google Scholar]
  26. Riley, R. D. , Higgins, J. P. T. , & Deeks, J. J. (2011). Interpretation of random effects meta‐analyses. British Medical Journal, 342, d549. 10.1136/bmj.d549 [DOI] [PubMed] [Google Scholar]
  27. Rodriguez, M. C. , & Maeda, Y. (2006). Meta‐analysis of coefficient alpha. Psychological Methods, 11(3), 306. 10.1037/1082-989X.11.3.306 [DOI] [PubMed] [Google Scholar]
  28. Rosa‐Alcázar, A. I. , Sánchez‐Meca, J. , Gómez‐Conesa, A. , & Marín‐Martínez, F. (2008). Psychological treatment of obsessive–compulsive disorder: A meta‐analysis. Clinical Psychology Review, 28(8), 1310–1325. 10.1016/j.cpr.2008.07.001 [DOI] [PubMed] [Google Scholar]
  29. Rubio‐Aparicio, M. , Núñez‐Núñez, R. M. , Sánchez‐Meca, J. , López‐Pina, J. A. , Marín‐Martínez, F. , & López‐López, J. A. (2020). The Padua Inventory–Washington State University Revision of obsessions and compulsions: A reliability generalization meta‐analysis. Journal of Personality Assessment, 102(1), 113–123. 10.1080/00223891.2018.1483378 [DOI] [PubMed] [Google Scholar]
  30. Saipanish, R. , Hiranyatheb, T. , Jullagate, S. , & Lotrakul, M. (2015). A study of diagnostic accuracy of the florida obsessive‐compulsive Inventory – Thai version (FOCI‐T). BMC Psychiatry , 15(1). 10.1186/s12888-015-0643-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Sánchez‐Meca, J. , López‐López, J. A. , & López‐Pina, J. A. (2013). Some recommended statistical analytic practices when reliability generalization studies are conducted. British Journal of Mathematical and Statistical Psychology, 66(3), 402–425. 10.1111/j.2044-8317.2012.02057.x [DOI] [PubMed] [Google Scholar]
  32. Sánchez‐Meca, J. , Marín‐Martínez, F. , López‐López, J. A. , Núñez‐Núñez, R. M. , Rubio‐Aparicio, M. , López‐García, J. J. , López‐Pina, J. A. , Blázquez‐Rincón, D. M. , López‐Ibáñez, C. , & López‐Nicolás, R. (2021). Improving the reporting quality of reliability generalization meta‐analyses: The REGEMA checklist.  Research Synthesis Methods, 12(4), 516–536. 10.1002/jrsm.1487 [DOI] [PubMed] [Google Scholar]
  33. Sánchez‐Meca, J. , López‐Pina, J. A. , López‐López, J. A. , Marín‐Martínez, F. , Rosa‐Alcázar, A. I. , & Gómez‐Conesa, A. (2011). The Maudsley Obsessive‐Compulsive Inventory: A reliability generalization meta‐analysis. International Journal of Clinical and Health Psychology, 11(3), 473–493. https://www.redalyc.org/articulo.oa?id=56019881004 [Google Scholar]
  34. Sánchez‐Meca, J. , Rubio‐Aparicio, M. , López‐Pina, J. A. , Núñez‐Núñez, R. M. , & Marín‐Martínez, F. (2015, July 22–24). El fenómeno de la inducción de la fiabilidad en Ciencias Sociales y de la Salud (The reliability induction phenomenon in the Social and Health Sciences) [Conference presentation). XIV Congress of Methodology of the Social and Health Sciences, Palma de Mallorca, Spain . [Google Scholar]
  35. Sánchez‐Meca, J. , Rubio‐Aparicio, M. , Núñez‐Núñez, R. M. , López‐Pina, J. , Marín‐Martínez, F. , & López‐López, J. A. (2017). A reliability generalization Meta‐Analysis of the Padua inventory of obsessions and compulsions. The Spanish Journal of Psychology, 20, 1–15. 10.1017/sjp.2017.65 [DOI] [PubMed] [Google Scholar]
  36. Scahill, L. , Riddle, M. A. , McSwiggin‐Hardin, M. , Ort, S. I. , King, R. A. , Goodman, W. K. , & Leckman, J. F. (1997). Children's Yale‐Brown obsessive compulsive scale: Reliability and validity. Journal of the American Academy of Child & Adolescent Psychiatry, 36(6), 844–852. 10.1097/00004583-199706000-00023 [DOI] [PubMed] [Google Scholar]
  37. Scherer, R. , & Teo, T. (2020). A tutorial on the meta‐analytic structural equation modeling of reliability coefficients. Psychological Methods, 25(6), 747. 10.1037/met0000261 [DOI] [PubMed] [Google Scholar]
  38. Storch, E. A. , Bagner, D. , Merlo, L. J. , Shapira, N. A. , Geffken, G. R. , Murphy, T. K. , & Goodman, W. K. (2007). Florida obsessive‐compulsive inventory: Development, reliability, and validity. Journal of Clinical Psychology, 63(9), 851–859. 10.1002/jclp.20382 [DOI] [PubMed] [Google Scholar]
  39. Storch, E. A. , Khanna, M. , Merlo, L. J. , Loew, B. A. , Franklin, M. , Reid, J. M. , & Murphy, T. K. (2009). Children's Florida obsessive compulsive inventory: Psychometric properties and feasibility of a self‐report measure of obsessive–compulsive symptoms in youth. Child Psychiatry and Human Development, 40(3), 467–483. 10.1007/s10578-009-0138-9 [DOI] [PubMed] [Google Scholar]
  40. Streiner, D. L. , & Norman, G. R. (2008). Health measurement scales: A practical guide to their development and use (4th ed.). Oxford University Press. 10.1093/acprof:oSo/9780199231881.001.0001 [DOI] [Google Scholar]
  41. Subramaniam, M. , Soh, P. , Vaingankar, J. A. , Picco, L. , & Chong, S. A. (2013). Quality of life in obsessive‐compulsive disorder: Impact of the disorder and of treatment. CNS Drugs, 27(5), 367–383. 10.1007/s40263-013-0056-z [DOI] [PubMed] [Google Scholar]
  42. Symonds, M. R. , & Moussalli, A. (2011). A brief guide to model selection, multimodel inference and model averaging in behavioural ecology using Akaike's information criterion. Behavioral Ecology and Sociobiology, 65(1), 13–21. 10.1007/s00265-010-1037-6 [DOI] [Google Scholar]
  43. Tavakol, M. , & Dennick, R. (2011). Making sense of Cronbach's alpha. International Journal of Medical Education, 2, 53–55. 10.5116/ijme.4dfb.8dfd [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Vacha‐Haase, T. (1998). Reliability generalization: Exploring variance in measurement error affecting score reliability across studies. Educational and Psychological Measurement, 58, 6–20. 10.1177/0013164498058001002 [DOI] [Google Scholar]
  45. Vacha‐Haase, T. , Kogan, L. R. , & Thompson, B. (2000). Sample compositions and variabilities in published studies versus those in test manuals: Validity of score reliability inductions. Educational and Psychological Measurement, 60(4), 509–522. 10.1177/00131640021970682 [DOI] [Google Scholar]
  46. Vacha‐Haase, T. , Ness, C. , Nilsson, J. , & Reetz, D. (1999). Practices regarding reporting of reliability coefficients: A review of three journals. The Journal of Experimental Education, 67(4), 335–341. 10.1080/00220979909598487 [DOI] [Google Scholar]
  47. Veale, D. , Simkin, V. , Orme, K. , & Grant, N. (2021). Defining reliable change, treatment response and remission on the Florida Obsessive Compulsive Inventory. Journal of Obsessive‐Compulsive and Related Disorders, 29, 100635. 10.1016/j.jocrd.2021.100635 [DOI] [Google Scholar]
  48. Viechtbauer, W. (2010). Conducting meta‐analyses in R with the metafor Package. Journal of Statistical Software , 36(3). 10.18637/jss.v036.i03 [DOI] [Google Scholar]
  49. Whittington, D. (1998). How well do researchers report their measures? An evaluation of measurement in published educational research. Educational and Psychological Measurement, 58(1), 21–37. 10.1177/0013164498058001003 [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Information

Data Availability Statement

The data that support the findings of this study are openly available at https://osf.io/zxn2k/.


Articles from Journal of Clinical Psychology are provided here courtesy of Wiley

RESOURCES