Abstract
Background
Standard setting is a critical component of assessment in health professions education, ensuring fairness, defensibility, and validity in determining minimum competence. Among criterion-referenced approaches, the Angoff method and its variants are widely applied. However, the evidence base on their comparative performance remains fragmented, with no recent quantitative synthesis available. This systematic review and meta-analysis aimed to evaluate the outcomes of Angoff methods across different educational and assessment contexts.
Methods
A systematic electronic search was conducted for studies that applied Angoff or its variants [Modified, Yes/No, Three-level, Group (Yes/No), or Mastery] in health professions assessments and reported at least one of the following: cut score, pass rate, or inter-rater reliability. Pooled estimates were calculated using generalized linear mixed models and random-effects meta-analysis. Odds ratio (OR) was used as the effect estimate for pass rate in comparative studies. Meta-regression analyses explored heterogeneity by number of judges, and test length.
Results
A total of 91 studies were included in the systematic review. Mixed treatment comparison pooled estimates revealed significantly higher pass rates with Angoff (OR: 7.48) compared to the conventional fixed method. Mastery Angoff (88.2%) and Mastery Angoff with reality check (86.9%) were associated with higher cut scores, resulting in substantially lower pass rates (61.44% for Mastery Angoff). The Modified Angoff with reality check achieved excellent reliability (r = 0.917) while the Angoff (Yes/No) yielded the weakest and most variable results (r = 0.536), a finding confirmed by bootstrap analyses. Meta-regression indicated that each additional judge was associated with a 0.19-percentage point increase in cut scores (p = 0.003) while each additional test item was associated with a 0.05-percentage point increase in the pass rate (p = 0.001).
Conclusion
For educators and policymakers, these findings underscore that the Angoff method is not a single technique but a flexible family of approaches with distinct outcomes. The choice of variants allows for the deliberate setting of more lenient or stringent standards, with Modified Angoff incorporating a reality check offering the most reliable and defensible results. This flexibility is a key strength but also a critical limitation, as an uninformed choice can lead to unintended pass rates. Therefore, selecting a specific Angoff variant must be a conscious decision aligned with the assessment’s purpose and context.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12909-025-08300-6.
Keywords: Standard setting method, Angoff, Modified angoff, Mastery angoff
Introduction
The process of determining minimum competence in health professional education examinations is a fundamental component of ensuring fairness, validity, and defensibility. Standard setting methods (SSMs) provide the foundation for establishing cut scores that separate competent candidates from those who have not yet achieved the required level of proficiency [1]. Unlike arbitrary fixed pass marks, SSMs are grounded in systematic procedures aligned with curriculum goals and professional context, thereby safeguarding both learners and society [2].
SSMs are broadly categorized as criterion-referenced or norm-referenced. Criterion-referenced approaches assess whether an examinee has achieved an absolute level of competence, which aligns with the ethos of professional education [3]. In contrast, norm-referenced methods define competence relative to a cohort, which risks passing underqualified candidates in a strong group or failing competent ones in a weak group [3]. Consequently, the educational community emphasizes criterion-referenced approaches to minimize false positives and false negatives [4]. Among criterion-referenced methods, the Angoff method is the most widely implemented and researched in health professions education [5]. It involves a panel of experts estimating the probability that a minimally competent candidate would answer each test item correctly [6]. Over time, several variants have been introduced to improve their practicality and reliability. These include the Modified Angoff (independent probability estimates) [7], the Yes/No method (a dichotomous choice) [8], the Three-level method (“Yes,” “No,” or “Maybe”) [9], and approaches incorporating a “reality check” with candidate performance data.
Despite its widespread use, systematic evaluations comparing the effectiveness and reliability of different Angoff variants remain limited. The only existing systematic review is from 2003 and was constrained by the limited evidence available at that time [10]. Given the evolving nature of assessments and the emphasis on evidence-based standard setting, a contemporary and comprehensive synthesis is urgently needed. Therefore, this systematic review and meta-analysis was undertaken to compare Angoff methods in health professional education. The objectives of this review were to: Compare the cut scores derived from different Angoff variants; examine the pass rates resulting from these methods; evaluate the inter-rater reliability (IRR) of different Angoff approaches; and identify key factors (number of judges and test length) affecting these outcomes.
Methods
Search strategy and study design
This nested work was conducted as part of a systematic review and meta-analysis on standard-setting approaches, registered with the Open Science Framework [11]. We systematically searched the following electronic databases for eligible studies: Medline (via PubMed), Cochrane CENTRAL, Google Scholar, and ERIC, with the last search completed on June 30, 2025. For Medline, the search string was Angoff [tiab], while in ERIC, the same strategy was applied with the additional filter “higher education.” No restrictions regarding publication year or language were imposed. Reference lists of relevant studies and previously published meta-analyses were also screened for additional eligible articles. Studies were included if they involved any form of the Angoff method, either as single-arm applications or as comparator studies contrasting Angoff approaches with fixed cut-off methods. The review was conducted in line with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) framework [12].
Eligible studies comprised any design that evaluated Angoff-based approaches. Both single-arm studies implementing Angoff methods and comparative studies examining Angoff against other standard-setting methods were considered.
Population: Undergraduate or postgraduate learners in medicine or other health professions, where Angoff-based methods were used to set performance standards.
Intervention: Various Angoff methods such as traditional Angoff, Modified Angoff, Angoff (Yes/No), Angoff Group (Yes/No), Three-level Angoff and Mastery Angoff methods (Table 1) were included in this review. For all methods, “reality checks” were considered if judges were provided with performance data from current or previous cohorts before finalizing cut scores.
Comparator: Conventional fixed pass-mark methods or any of the norm-reference methods (such as mean-1 SD, mean-1.5 SD or mean-2SD) in comparator studies. Studies that have compared Angoff methods with other SSM were included and only the data for Angoff methods were considered in this meta-analysis and were pooled separately.
Outcomes: Pass rates, cut scores, and IRR.
Table 1.
Description of Angoff methods used in this review
| Angoff methods | Description |
|---|---|
| Angoff (Traditional Angoff) | Judges collectively agree on the likelihood that a minimally competent (borderline) candidate would answer each item correctly. This standard defines the threshold for basic proficiency. |
| Modified Angoff | Judges independently provide their estimates of the probability that borderline candidates would answer correctly. |
| Angoff (Yes/No) | Judges independently state whether a borderline candidate would likely answer each item correctly (“yes”) or not. |
| Angoff Group (Yes/No) | Judges reach a group consensus on whether a borderline candidate would correctly answer (“yes”) or not. |
| Three-level Angoff | Judges independently classify each item as “yes,” “maybe,” or “no” with respect to whether a borderline candidate would answer correctly. |
| Mastery Angoff | This variant sets a higher standard. Instead of a borderline candidate, judges estimate the probability that a candidate demonstrating advanced proficiency or mastery (one who has fully achieved the learning outcomes) would answer each item correctly. This establishes a cut score for excellence rather than minimal competence. |
Study procedure
Two reviewers independently performed the literature search and data extraction. Extracted variables included: study identifier, design, learner level (undergraduate/postgraduate), number of examinees, item count, assessment format (MCQs, OSCEs/OSPEs, SAQs), number and expertise of judges, and reported outcomes (pass rates, cut scores, inter-rater reliability). Discrepancies were resolved through discussion. Publication bias could not be assessed as none of the comparisons had at least ten studies [13]. A random-effects model was employed to generate pooled estimates for direct, indirect, and mixed treatment comparisons of pass rates. Direct pooled estimates came from head-to-head studies, indirect estimates were based on common comparators, and mixed estimates combined both sources. Odds ratios (ORs) with 95% confidence intervals (CIs) were used as the effect estimates for pass rates from comparative studies, while pooled mean with 95% CI for other pooled estimates. Inconsistency between direct and indirect estimates was assessed using the H statistic, classified as low (< 3), moderate (3–6), or high (>6) for the mixed treatment comparison estimates [14]. Subgroup analyses were performed according to discipline (medicine, nursing, dentistry), item type (MCQ vs. OSCE), and learner level (undergraduate vs. postgraduate). Leave-one-out sensitivity analyses were performed by excluding one study at a time to examine robustness of findings. For outcomes with significant overall pooled effects, bootstrap meta-analysis was conducted using the DerSimonian–Laird random-effects model with 1,000 replicates. Cumulative meta-analyses were performed by chronologically adding studies. Pooled estimates from comparative studies generated using mixed treatment comparisons were generated using MetaXL© [15]. Study quality was assessed with a modified Medical Education Research Study Quality Instrument (MMERSQI) [16], omitting the “validity” and “outcomes” items as they were not applicable with the maximum score of 65. The certainty of evidence was graded using a modified GRADE approach [17]. Sensitivity analyses and bootstrap methods (for 1000 iterations) were performed for the outcomes and variability in the pooled estimates with removal of studies for each SSM (for sensitivity analysis) and resampling pooled estimates for Angoff variants (for bootstrap analyses) were assessed.
Three separate meta-analyses were conducted using R’s metafor package to synthesize data on pass rates, cut scores, and inter-rater reliability estimates obtained from all studies including those that did not compare the Angoff methods with conventional fixed or norm-reference methods.
Pass Rates: Proportions were pooled using a generalized linear mixed model with a logit link and a random study effect. Method type was a fixed effect for comparisons, and meta-regression assessed the influence of the number of judges and items.
Cut Scores: The cut scores were expressed as percentages in individual studies. We used inverse-variance weights for approximating the standard error (SE) for each study’s cut score. Based on the few studies that mentioned standard deviation of item-level judgments, a median estimate of 4.31% was used as a representative conservative and commonly observed estimate of judge variability. The standard error for each study was subsequently calculated using the formula SE = 4.31/√J, where J is the number of judges participating in the standard-setting exercise for that study. The variance for each study was then derived as V = SE². A random-effects meta-analysis was conducted using the generic inverse-variance method to pool the cut scores, accounting for anticipated heterogeneity beyond sampling error arising from differences in tests, panels, and contexts.
Inter-Rater Reliability: Reliability coefficients (r) were Fisher’s z-transformed for pooling in a random-effects model, then back-transformed for interpretation. The variance was calculated as 1/(I-3), with I being the number of items. Subgroup analysis and meta-regression were used to compare methods and assess the influence of judges (J) and items (I).
Results
Search results
The search strategy identified 232 studies, of which 91 [18–108] were included in the systematic review and 88 in the meta-analysis (Fig. 1). Key characteristics of the included studies are presented in the Electronic Supplementary File 1. Across learner populations, most studies assessed undergraduate students (n = 55, 60.4%), followed by postgraduates, residents, and fellows (n = 36, 39.6%). The majority were conducted in medical specialties (n = 67, 73.6%), followed by dentistry (n = 8, 8.8%). In terms of assessment format, MCQs were the most frequently used (n = 42, 46.2%), followed by OSCE/OSPE formats (n = 24, 26.4%). The most applied methods were the Modified Angoff (n = 34, 37.4%), followed by Angoff (n = 22, 24.2%) and Angoff (Yes/No) (n = 15, 16.5%). The number of judges varied widely with a median of 11, typically consisting of clinical faculty, professors, or trained examiners with substantial teaching or assessment experience. Reported outcomes included cut scores (n = 80, 87.9%) and pass rates (n = 57, 62.6%), while IRR was assessed in 37 studies (40.7%). The median (range) of modified MMERSQI scores for included studies was 43 (35–47).
Fig. 1.
PRISMA flow diagram. A total of 91 studies were included in this systematic review
Pooled estimates for pass rates from comparative studies
A key aspect of the comparative analysis is that it isolates the effect of the SSM. In the studies included, different methods (such as Angoff variants, fixed standards) were applied to the same examination and the same cohort of students. This approach inherently controls for the actual performance level of the cohort, as any difference in the resulting pass rate can be more confidently attributed to the method used to set the cut score, rather than to differences in student ability. Therefore, the pooled OR presented below reflect the relative leniency or stringency of the standard-setting methods themselves. Twenty-one studies (22,622 examinees) were pooled for analysis of overall pass rates, with most comparing Angoff against conventional fixed methods (Fig. 2A). Mixed treatment comparison pooled estimates revealed significantly higher pass rates with Angoff (OR: 7.48; 95% CI: 3.24–17.28) and norm-referenced mean–2SD methods (OR: 1.79; 95% CI: 1.10–2.91) (Fig. 2B). Heterogeneity was mild (H = 1).
Fig. 2.
Network and forest plots for pass rate from comparative studies. A: Network plot of Angoff methods included for pass rates in the comparative studies; and B: Forest plot of mixed treatment comparison pooled estimates of pass rates of Angoff and norm-reference methods with conventional fixed method. The network plot depicts the relationship between SSMs that are assessed in the included studies. The sizes of green circles represent the relative number of studies assessing each SSM. The forest plot represents the mixed treatment comparison pooled estimates that were generated both by direct and indirect comparison methods. The blue circles represent the point estimates, and the blue lines represent the 95% CI for these pooled estimates. Vertical black line represents the line of no difference in the risk of pass rate. Angoff method was observed with significantly higher pass rate compared to conventional fixed method of standard setting
Subgroup analyses for pass rates
Details of the pooled estimates in various subgroups are outlined in the Electronic Supplementary File 2. The overarching pattern indicated a tendency towards more lenient standards (higher pass rates) with Angoff methods for undergraduate learners and in OSCE/OSPE formats. In contrast, a trend towards stricter standards (lower pass rates) was observed at the postgraduate level and within the dentistry discipline (Fig. 3). The specific Angoff variant used was a critical determinant of stringency. The Modified Angoff method, particularly when used with a reality check, was consistently associated with more lenient standards across several subgroups, including in medical disciplines and for assessments with 11–25 items. Smaller panels (fewer than 5 judges) and shorter tests (fewer than 10 items) were consistently associated with more lenient standards.
Fig. 3.
Forest plots of pass rate in comparative studies in various sub-groups. Forest plots of mixed treatment comparison pooled estimates for pass rates amongst undergraduate students (A); postgraduate students (B); MCQ (C); OSPE (D); Medicine (E); Dentistry (F); Dentistry (G); Items ≤ 10 (H); Items 11 to 25 (I); Items > 25 (J); Judges ≤ 5 (K); Judges 6 to 10 (L); and Judges ≥ 11 (M). The forest plot represents the mixed treatment comparison pooled estimates that were generated both by direct and indirect comparison methods. The blue circles represent the point estimates, and the blue lines represent the 95% CI for these pooled estimates. Vertical black line represents the line of no difference in the risk of pass rate
Pooled estimates of cut scores
The forest plot of cut scores across methods is shown in Fig. 4A. Mastery Angoff and Mastery Angoff with reality check were associated with significantly higher cut scores compared to other Angoff methods. Moderator analysis using bubble plots (Fig. 4B) demonstrated a significant association between the number of judges and cut scores (regression coefficient = 0.19, p = 0.003), but no significant association with the number of test items (regression coefficient = − 0.01, p = 0.2). Each additional judge was associated with a 0.19-percentage point increase in cut score.
Fig. 4.
Forest and Bubble plots for cut scores. A: Forest plot and B: Bubble plots for moderator analysis. Square boxes represent point pooled estimates and horizontal lines indicate 95% CI for each method in the forest plot. Each bubble represents an individual study or data point, with its size proportional to study variance. The analysis explores how the scale of the standard-setting process (number of judges) depicted in blue circles and the length of the examination (number of test items) indicated in green circles influence the resulting passing standard. The trend line (red color with number of judges and blue color with the number of test items) illustrates the overall correlation between the variables
Pooled estimates for pass rates
The forest plot for pass rates with various SSMs assessed in this study is depicted in Fig. 5A. Overall, these SSMs yield an average pass rate of 80.85% (95% CI: 75.86, 84.99). Among the various methods, Angoff with reality check, norm-referenced methods and the Angoff Group (Yes/No) variant tend to produce higher pass rates (95.8%, 95.43% and 93.36%, respectively). In contrast, stricter standards are set by the Mastery Angoff method and a conventional fixed 70% cutoff, resulted in substantially lower pass rates of 61.44% and 33.40%, respectively. The moderator analysis (Fig. 5B) revealed a significant association with the number of test items (regression coefficient: 0.05; and p-value: 0.001) but not with the number of judges (regression coefficient: −0.01; and p-value: 0.8). This indicates that each additional test item is associated with 0.05%-point increase in the pass rate.
Fig. 5.
Forest and Bubble plots for pass rates. A: Forest plot and B: Bubble plots for moderator analysis. Blue circles represent point pooled estimates and horizontal lines indicate 95% CI for each method in the forest plot. Each bubble represents an individual study or data point, with its size proportional to study variance. The analysis explores how the scale of the standard-setting process (number of judges) depicted in blue circles and the length of the examination (number of test items) indicated in green circles influence the resulting passing standard. The trend line (red color bold line) illustrates the overall correlation between the variables
Pooled estimates for IRR coefficients
The forest plot for IRR (Fig. 6A) demonstrated that incorporating a reality check substantially improved reliability. The Modified Angoff with reality check achieved excellent reliability (r = 0.917). In contrast, the Angoff (Yes/No) yielded the weakest and most variable results (r = 0.536), suggesting that oversimplification compromises accuracy. These findings support the robustness of Angoff and Modified Angoff methods for defensible standard setting, whereas variants with fewer supporting studies warrant further validation. Moderator analysis (Fig. 6B) showed no significant association between IRR and either the number of judges (regression coefficient = 0.02, p = 0.26) or the number of test items (regression coefficient = − 0.01, p = 0.2).
Fig. 6.
Forest and Bubble plots for inter-rater reliability coefficients. A: Forest plot and B: Bubble plots for moderator analysis. Square boxes represent point pooled estimates and horizontal lines indicate 95% CI for each method in the forest plot. Each bubble represents an individual study or data point, with its size proportional to study variance. The analysis explores how the scale of the standard-setting process (number of judges) depicted in blue circles and the length of the examination (number of test items) indicated in green circles influence the resulting passing standard. The trend line (black color bold line) illustrates the overall correlation between the variables
Sensitivity analyses
Statistical sensitivity analyses for outcomes are depicted in Fig. 7. Regarding the cut scores (Fig. 7A), the Angoff (Yes/No) variant demonstrated the highest variability (mean = 69.6%, SD = 18.8), indicating its cut scores are highly inconsistent and context dependent. Other methods with considerable variability included the Angoff with reality check (SD = 17.5) and the Three-level Angoff (SD = 13.2). The traditional Angoff method also showed high variability (SD = 15.0). More stable methods included the Modified Angoff (mean = 61.0%, SD = 9.1) and the Mastery Angoff (mean = 88.2%, SD = 7.2), which consistently set a high, stringent standard.
Fig. 7.
Sensitivity analyses for outcomes. Box plots depicting the statistical variability of sensitivity analysis for cut scores (A), pass rates (B) and IRR coefficient (C). The end-lines in the boxes indicate first and third quartiles, and the middle line in the box indicates median values for each SSM
Regarding the sensitivity analysis of pass rates, a clear stratification between lenient and stringent standard-setting methods was observed. The most lenient methods, resulting in the highest average pass rates, were Angoff with reality check and the norm-referenced (mean-2SD) approach (both 95.8%), followed by a conventional fixed 60% cutoff (90.8%). In direct contrast, the most stringent methods, which produced the lowest pass rates, were the conventional fixed 70% standard (34.7%), the averaged norm-referenced (mean-1SD) method (45.7%), and Three-level Angoff (52.4%). Furthermore, the consistency of these outcomes varied substantially across studies. While some methods like the Three-level Angoff were predictably stringent on average, they exhibited high variability, indicating their outcomes are highly context-dependent.
Bootstrap analyses
The variability of the pooled estimates following bootstrapping for Angoff methods for all outcomes are depicted in Fig. 8. Mastery Angoff followed by Modified Angoff with reality check were observed with higher cut scores and reduced variability compared to other Angoff variants (Fig. 8A). Bootstrapping of the IRRs revealed that Modified Angoff with reality check and Mastery Angoff were associated with high coefficients (highly reliable) while Angoff (Yes/No) was the least reliable with huge variability (Fig. 8C).
Fig. 8.
Bootstrap analyses for Angoff methods for all the outcomes. Bootstrap reliability plots depicting the variability of pooled estimates following resampling for cut scores (A), pass rates (B) and IRR coefficient (C). Blue circles represent the mean estimate, and the horizontal red lines represent 95% CI following resampling bootstrap methods
Grading the strength of evidence
A summary of key estimates for cut scores, pass rates and IRR for various Angoff variants can be found in Table 2. This meta-analysis reveals a clear trade-off between the stringency and reliability of different Angoff methods. The Mastery Angoff method is the most stringent, producing the highest cut score and the lowest pass rate, while methods like the Angoff Group (Yes/No) and Angoff with reality check are the most lenient, resulting in the highest pass rates. Crucially, the choice of variant profoundly impacts the defensibility of the standard, as shown by the IRR. The Modified Angoff with a reality check achieved excellent reliability, making it a robust choice for high-stakes assessment, whereas the Angoff (Yes/No) method demonstrated the poorest reliability, underscoring the risk of its practicality. However, the strength of evidence for key comparisons for pass rates was observed to be very low (Table 3) mainly constrained by the imprecise estimates, inability to rule out the publication bias and heterogeneity amongst the included studies.
Table 2.
Summary of pooled estimates for pass rates, cut scores and IRR for various Angoff methods
| Method | Pass rate (%) | Cut scores | IRR |
|---|---|---|---|
| Modified Angoff | 71.28 [58.37–81.46] | 61.22 [58.81, 63.63] | 0.78 [0.71, 0.83] |
| Angoff (Yes/No) | 85.31 [76.71–91.11] | 69.61 [62.38, 76.84] | 0.54 [0.26, 0.79] |
| Angoff | 80.14 [69.06–87.93] | 64.47 [60.07, 68.87] | 0.71 [0.58, 0.81] |
| Mastery Angoff | 61.44 [39.48–79.54] | 88.22 [84.12, 92.32] | 0.85 [0.53, 0.97] |
| Modified Angoff with reality check | 69.78 [48.03–85.22] | 66.7 [63.38, 70.02] | 0.92 [0.88, 0.94] |
| Angoff (Yes/No) with reality check | 83.83 [78.8–87.87.8.87] | 60.15 [54.09, 66.21] | 0.76 [0.76, 0.76] |
| Three-level Angoff | 73.54 [3.25–99.57] | 58.54 [45.66, 71.42] | 0.7 [0.65, 0.75] |
| Angoff Group (Yes/No) | 93.36 [77.71–98.27] | 58.12 [52.7, 63.54] | Not available |
| Angoff with reality check | 95.80 [95.06–96.44] | 78.65 [54.45, 102.9] |
Values are represented in pooled mean [95% CI]
Table 3.
Strength of evidence for pass rate for key Angoff methods from comparative studies
| Pass rate compared to conventional fixed standard setting method | Comparative risks of pass rate per 1000 examinees (95% confidence intervals) | Effect estimates (OR) and the quality of evidence for mixed treatment comparisons | |
|---|---|---|---|
| Assumed risk1 | Corresponding risk | ||
| Overall pass rate with Angoff | 838 |
975 (944 to 989) |
7.48 [3.24, 17.28]; ⊕⊝⊝⊝; Very low2, 3, 4 |
| Pass rate with Angoff amongst undergraduate students | 714 |
969 (924 to 988) |
12.59 [4.86, 32.59]; ⊕⊝⊝⊝; Very low2, 3, 4 |
| Pass rate with Modified Angoff with MCQ test items | 967 |
1000 (998 to 1000) |
286.75 [14.34, 5733.13]; ⊕⊝⊝⊝; Very low2, 3, 4 |
| Pass rate with Angoff with OSPE/OSCE test items | 364 |
878 (736 to 949) |
12.59 [4.86, 32.59]; ⊕⊝⊝⊝; Very low2, 3, 4 |
| Pass rate with Modified Angoff with reality check in medicine discipline | 821 |
1000 (979 to 1000) |
672.59 [10.35, 43707.5]; ⊕⊝⊝⊝; Very low2, 3, 4 |
| Pass rate with Modified Angoff in medicine discipline | 821 |
997 (854 to 1000) |
78.22 [1.28, 4788.22]; ⊕⊝⊝⊝; Very low2, 3, 4 |
1 Median risk in the conventional fixed method
1 Downgraded two levels as publication bias could not be assessed/ruled out
2 Downgraded one levels for serious limitations in the precision of the estimates
3 Downgraded one level for serious limitations in not assessing the publication bias
4 Downgraded one level for serious limitations due to heterogeneity of the included studies
Very low: We have very little confidence in the effect estimate
Discussion
Key findings
This systematic review and meta-analysis of 91 studies demonstrate the Angoff method is best understood not as a single technique but as a flexible family of methods, whose validity and impact are directly shaped by its implementation. The main conclusions are:
Angoff methods generally produce significantly higher pass rates compared to conventional fixed standards (e.g., a fixed 70% cutoff).
The specific variant used is a primary determinant of stringency and reliability, with the Modified Angoff often setting more lenient standards and the Yes/No variant resulting in stricter outcomes, especially for postgraduates.
The Modified Angoff with a reality check and the Mastery Angoff were identified as the most reliable variants, whereas the Angoff (Yes/No) was the least reliable.
Context matters profoundly, with outcomes significantly moderated by the learner’s level (undergraduate vs. postgraduate), academic discipline (e.g., medicine vs. dentistry), and assessment format.
Methodological parameters influence the standard, as the number of judges positively influenced cut scores, and test length was associated with pass rates.
The certainty of this evidence is graded as very low, limited mainly by the imprecision of estimates and the inability to rule out publication bias.
Comparison with existing literature
The selection of a SSM is a high-stakes decision with profound implications for educational fairness, program accreditation, and public policy. A defensible cut score ensures that competent candidates are not erroneously failed (minimizing false negatives) and that those who are not yet competent do not progress (minimizing false positives), thereby upholding the fundamental social contract of the health professions. Arbitrary or poorly justified standards can undermine the validity of an entire assessment, lead to challenges in accreditation reviews, and result in inequitable policy decisions affecting the careers of countless learners and the safety of the public they will serve. Therefore, the choice between Angoff variants, as detailed in this review, is not merely a methodological preference but a core component of responsible educational governance. Our findings must be interpreted within the context of the inherent pedagogical and conceptual challenges of judgment-based standard setting. The Angoff method rests on several complex assumptions: that a panel of experts can consistently conceptualize a hypothetical “minimally competent” or “borderline” candidate; that their individual judgments of item-level performance for this candidate are accurate; and that the process avoids circularity, where judges’ perceptions are unconsciously anchored by their knowledge of student performance rather than an independent ideal. Furthermore, the final standard is inevitably influenced by a host of external factors, including the quality of teaching, student support systems, and the design of the assessment items themselves. Our quantitative synthesis does not resolve these foundational issues, but it operates upon this landscape, demonstrating that even within these shared constraints, specific procedural choices lead to significantly different and more or less reliable outcomes.
The finding that Angoff methods generally produce higher pass rates than conventional fixed standards align with and extends the existing critique of arbitrary cut-scores. Fixed pass marks are inherently arbitrary, failing to account for variable exam difficulty and the specific competence level of a cohort [109]. The Angoff method, by its judgmental and item-centric nature, introduces necessary flexibility, allowing the standard to be tailored to the assessment itself. This likely explains its widespread adoption, as judges may inherently account for subtle item flaws or realistic expectations of a minimally competent candidate, which a fixed percentage cannot [110]. However, our meta-analysis reveals a critical nuance that challenges the view of Angoff as a monolithic method. The most critical contribution of the present study is the stark contrast between Angoff variants. While the simplified Angoff (Yes/No) method is often promoted for its practicality, our findings suggest this may come at a significant cost to reliability, a concern that has been raised anecdotally but is now quantitatively substantiated [8]. By reducing judgment to a dichotomous choice, it may strip away the nuanced deliberation central to other variants, potentially introducing error and arbitrariness.
Conversely, our results validate the use of the Modified Angoff method with a reality check, confirming and quantifying its superiority in achieving excellent reliability. This finding provides robust empirical support for best-practice recommendations that have been based on smaller-scale studies [7]. The “reality check,” where judges review actual candidate performance and psychometric data, anchors subjective judgments in empirical evidence. From a validity perspective, this process directly strengthens the argument based on Kane’s framework, supporting the extrapolation inference by linking judgments to performance data and the decision inference by providing a transparent, evidence-based rationale for the cut score [111]. Our observation that the number of judges positively predicts cut scores also aligns with a recent large-scale simulation, which found that panel size is a key factor for precision, with diminishing returns after 15–30 judges depending on test length [112].
These findings yield direct and actionable implications for educational practice and standard-setting procedures. To ensure a psychometrically defensible and reliable process, the Modified Angoff method incorporating a reality check is recommended as the preferred variant. Conversely, the Angoff (Yes/No) method should be employed with caution and only alongside extensive judge training, given its propensity for lower reliability. Regarding panel composition, while a larger number of judges may marginally increase the cut score, a well-trained panel of 8 to 12 members is likely sufficient, as there is no strong evidence that larger panels further enhance reliability. This underscores that investment in comprehensive judge training is more critical than merely expanding the panel’s size. Ultimately, these decisions must be guided by contextual awareness; the choice of standard-setting method must be intentional and account for the learner level (e.g., anticipating stricter standards for postgraduates), the academic discipline, and the assessment format, affirming that a single, universal approach is not appropriate.
Strengths, limitations and way forward
This study is the first systematic review and meta-analysis to comprehensively evaluate the various Angoff standard setting methods from 91 studies published cross-disciplinary inclusion (medicine, dentistry, nursing). A rigorous methodology was adopted, including a pre-registered protocol, comprehensive database searches across multiple sources, and adherence to PRISMA guidelines. Robust statistical approaches such as generalized linear mixed models, random-effects meta-analysis, subgroup analyses, and meta-regression were employed to address heterogeneity and identify contextual moderators. The inclusion of multiple Angoff variants (traditional, modified, Yes/No, three-level, mastery) and the application of validated quality assessment tools (MMERSQI, modified GRADE) further enhance the credibility and reproducibility of findings.
Several limitations must be acknowledged. First, the inherent heterogeneity in how primary studies implement and report their methods (such as differing judge training protocols, variations in executing a “reality check”) introduces noise into the meta-analysis. Second, the field is dominated by studies in medical education, limiting the generalizability of findings to other health professions, though we have identified dentistry as a distinct outlier. Finally, as with all meta-analyses, our conclusions are based on aggregated study-level data rather than individual participant data, which constrains deeper causal inference.
Future research should prioritize well-designed, multicenter, and adequately powered trials that directly compare different Angoff variants in diverse health professional education settings. Standardized reporting of key outcomes such as cut scores, pass rates, and IRR would improve comparability and facilitate meta-analytic synthesis, particularly when pooling from single group studies. Incorporating psychometric modeling and simulation studies may further clarify the impact of judge number, item pool size, and examinee population on the robustness of Angoff-derived standards. Additionally, evaluating Angoff methods in emerging assessment formats such as virtual simulations, and particularly in the complex context of workplace-based assessments (WBAs) and programmatic assessment frameworks, will be a crucial and necessary challenge for the field. Establishing international consensus on best practices and minimum methodological standards for applying and reporting Angoff procedures would strengthen the evidence base and guide educators in selecting defensible and context-appropriate standard setting approaches.
Conclusion
In conclusion, this review affirms the value of the Angoff method as a robust, flexible, and widely applicable approach to standard setting in health professions education. However, it definitively shows that the Angoff method is not a single, uniform technique. Its outcomes, the pass rates, cut scores, and reliability, are profoundly shaped by procedural choices. The choice of variants allows for the deliberate setting of more lenient or stringent standards, with Modified method incorporating a reality check offering the most reliable and defensible results. The variant selected, the composition and training of the panel, the integration of empirical feedback, and the context of the assessment are not mere details but are active determinants of the standard itself. Therefore, the practice of setting standards must evolve from simply “using the Angoff” to making a series of evidence-based decisions about how to use it most appropriately for a given context. Future research should focus on validating these standards against external outcome measures and further exploring the sociological and cognitive processes that underline expert judgment in these critical panels.
Supplementary Information
Acknowledgements
We wish to acknowledge ChatGPT for improving the language clarity and grammar in this manuscript.
Authors’ contributions
KS: Conceived the idea; KS and GS: Data curation, analysis and interpretation; KS: Wrote the first draft of the manuscript; and KS and GS: Involved in critical revisions and final acceptance of the manuscript. The authors confirm that we have substantially contributed to the conception and design of the review article and interpreting the relevant literature and have been involved in writing the review article and revising it for intellectual content.
Funding
This paper was not funded.
Data availability
The data is available with the corresponding author and shall be shared upon request.
Declarations
Competing interests
The authors declare no competing interests.
ORCID ID
0000-0003-3811-6503.
Conflict of interest
The authors have no relevant affiliations or financial involvement with any organization or entity with a financial interest in or financial conflict with the subject matter or materials discussed in the manuscript. This includes employment, consultancies, honoraria, stock ownership or options, expert testimony, grants or patents received or pending, or royalties.
Ethics approval
This study was carried out on a publicly available database due to which Ethics approval was not required.
Consent to publish
declaration: not applicable.
Consent to participate
declaration: not applicable.
Clinical trial number
not applicable.
Authors’ contributions statement
KS: Conceived the idea; KS and GS: Data curation, analysis and interpretation; KS: Wrote the first draft of the manuscript; and KS and GS: Involved in critical revisions and final acceptance of the manuscript. The authors confirm that we have substantially contributed to the conception and design of the review article and interpreting the relevant literature and have been involved in writing the review article and revising it for intellectual content.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.De Champlain AF. Standard setting methods in medical education: high-stakes assessment. Understanding medical education: Evidence, theory, and practice. 2018;3:347–59. [Google Scholar]
- 2.Mortaz Hejri S, Jalili M. Standard setting in medical education: fundamental concepts and emerging challenges. Med J Islam Repub Iran. 2014;28:34. [PMC free article] [PubMed] [Google Scholar]
- 3.Ricker KL. Setting cut-scores: a critical review of the Angoff and modified Angoff methods. Alberta J Educ Res. 2006;52(1):53–64. [Google Scholar]
- 4.Hassan S. Standard setting in medical education: Standards, Methods, and Psychometrics. In: Global medical education in normal and challenging times. Cham: Springer Nature Switzerland. 2024:137–50.
- 5.Saaiq M. Standard setting methods for the assessment of knowledge and skills in medical education. J Health Prof Edu Innov. 2024;1:14–20. [Google Scholar]
- 6.Mubuuke AG, Mwesigwa C, Kiguli S. Implementing the Angoff method of standard setting using postgraduate students: Practical and affordable in resource-limited settings. Afr J Health Prof Educ. 2017;9(4):171–175. [DOI] [PMC free article] [PubMed]
- 7.Modified-Angoff Method; Set a Defensible Cutscore. Available at: https://assess.com/modified-angoffmethod/#:~:text=The%20modified%2DAngoff%20method%20is,difficulty%20and%20he%20intended%20population. Accessed on 27 Aug 2025.
- 8.Park J. Possibility of using the yes/no Angoff method as a substitute for the percent Angoff method for estimating the cutoff score of the Korean medical licensing examination: a simulation study. J Educ Eval Health Prof. 2022;19:23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Liu M, Liu KM. Setting pass scores for clinical skills assessment. Kaohsiung J Med Sci. 2008;24(12):656–63. [DOI] [PMC free article] [PubMed]
- 10.Hurtz GM, Auerbach MA. A meta-analysis of the effects of modifications to the Angoff method on cutoff scores and judgment consensus. Educ Psychol Meas. 2003;63(4):584–601.
- 11.Standard setting methods in health professional education: Systematic review and meta-analysis. Available at: https://osf.io/udpja. Aaccess date: 27 Sept 2025. [DOI] [PMC free article] [PubMed]
- 12.Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, Shamseer L, Tetzlaff JM, Akl EA, Brennan SE, Chou R. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021;372. [DOI] [PMC free article] [PubMed]
- 13.Sterne JA, Harbord RM. Funnel plots in meta-analysis. Stata J. 2004;4(2):127–41.
- 14.Dias S, Caldwell DM. Network meta-analysis explained. Arch Dis Child Fetal Neonatal Ed. 2019;104(1):F8–F12. [DOI] [PMC free article] [PubMed]
- 15.Xia M, Zheng G, Mukherjee S, Shokouhi M, Neubig G, Awadallah AH. MetaXL: Meta representation transformation for low-resource cross-lingual learning. arXiv preprint arXiv:2104.07908. 2021.
- 16.Al Asmri M, Haque MS, Parle J. A modified medical education research study quality instrument (MMERSQI) developed by Delphi consensus. BMC Med Educ. 2023;23(1):63. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Brignardello-Petersen R, Mustafa RA, Siemieniuk RA, Murad MH, Agoritsas T, Izcovich A, et al. GRADE approach to rate the certainty from a network meta-analysis: addressing incoherence. J Clin Epidemiol. 2019;108:77–85. [DOI] [PubMed] [Google Scholar]
- 18.Abd-Rahman AN, Baharuddin IH, Abu‐Hassan MI, Davies SJ. A comparison of different standard‐setting methods for professional qualifying dental examination. J Dent Educ. 2021;85(7):1210–6. [DOI] [PubMed] [Google Scholar]
- 19.Yousefi Afrashteh M. Comparison of the validity of bookmark and Angoff standard setting methods in medical performance tests. BMC Med Educ. 2021;21(1):1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Albin CS, Pergakis MB, Sigman EJ, Bhatt NR, Hutto SK, Koneru S, Osehobo EM, Vizcarra JA, Morris NA. Education research: junior neurology residents achieve competency but not mastery after a brief acute ischemic stroke simulation course. Neurol® Educ. 2023;2(2):e200071. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Alston GL, Haltom WR. Reliability of a minimal competency score for an annual skills mastery assessment. Am J Pharm Educ. 2013;77(10):211. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Anderson HG Jr, Nelson AA. Reliability and credibility of progress test criteria developed by alumni, faculty, and mixed alumni-faculty judge panels. Am J Pharm Educ. 2011;75(10):200. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Ansari R, Ab Manan N, Mahat NA, Omar NS, Abdul Latiff A, Idris S, et al. Standard setting methods in objective structured clinical examination (OSCE): a comparative study of five methods. Journal of Medical Education Development. 2024;17(56):87–96. [Google Scholar]
- 24.Bahammam LA. Cutoff score evaluation of undergraduate dental curriculum. Med Teach. 2017;39(sup1):S33-6. [DOI] [PubMed] [Google Scholar]
- 25.Barsuk JH, Cohen ER, Wayne DB, McGaghie WC, Yudkowsky R. A comparison of approaches for mastery learning standard setting. Acad Med. 2018;93(7):1079–84. [DOI] [PubMed] [Google Scholar]
- 26.Basu S, Roberts C, Newble DI, Snaith M. Competence in the musculoskeletal system: assessing the progression of knowledge through an undergraduate medical course. Med Educ. 2004;38(12):1253–60. [DOI] [PubMed] [Google Scholar]
- 27.Bick JS, Wanderer JP, Myler CS, Shaw AD, McEvoy MD. Standard Setting for Clinical Performance of Basic Perioperative Transesophageal Echocardiography. Anesthesiology. 2017;126(4):718–28. [DOI] [PubMed] [Google Scholar]
- 28.Boursicot KA, Roberts TE, Pell G. Standard setting for clinical competence at graduation from medical school: a comparison of passing scores across five medical schools. Adv Health Sci Educ Theory Pract. 2006;11(2):173–83. [DOI] [PubMed] [Google Scholar]
- 29.Brunk I, Schauber S, Georg W. Do they know too little? An inter-institutional study on the anatomical knowledge of upper-year medical students based on multiple choice questions of a progress test. Annals of Anatomy-Anatomischer Anzeiger. 2017;209:93–100. [DOI] [PubMed]
- 30.Burr S, Martin T, Edwards J, Ferguson C, Gilbert K, Gray C, Hill A, Hosking J, Johnstone K, Kisielewska J, Milsom C. Standard setting anchor statements: a double cross-over trial of two different methods. MedEdPublish. 2021;10:32. [DOI] [PMC free article] [PubMed]
- 31.Carlson J, Tomkowiak J, Knott P. Simulation-based examinations in physician assistant education: a comparison of two standard-setting methods. J Physician Assist Educ. 2010;21(2):7–14. [DOI] [PubMed] [Google Scholar]
- 32.Chan SC, Amin SM, Lee TW. Implementing standard setting into the conjoint MAFP/FRACGP part 1 examination–process and issues. Malaysian Family Physician: Official J Acad Family Physicians Malaysia. 2016;11(2–3):2. [PMC free article] [PubMed] [Google Scholar]
- 33.Clauser JC, Hambleton RK, Baldwin P. The effect of rating unfamiliar items on Angoff passing scores. Educ Psychol Meas. 2017;77(6):901–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Cusimano MD, Rothman AI. The effect of incorporating normative data into a criterion-referenced standard setting in medical education. Acad Med. 2003;78(10):S88-90. [DOI] [PubMed] [Google Scholar]
- 35.Dalum J, Christidis N, Myrberg IH, Karlgren K, Leanderson C, Englund GS. Are we passing the acceptable? Standard setting of theoretical proficiency tests for foreign-trained dentists. Eur J Dent Educ. 2023;27(3):640–9. [DOI] [PubMed] [Google Scholar]
- 36.Dimanche K, Klatt EC, Angle SM Jr. Predictive validity evidence of Yes-No Angoff standard setting in a Pre-Clinical medical school curriculum. BMC Med Educ. 2025;25(1):384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Downing SM, Lieska NG, Raible MD. Establishing passing standards for classroom achievement tests in medical education: a comparative study of four methods. Acad Med. 2003;78(10):S85-7. [DOI] [PubMed] [Google Scholar]
- 38.Dwivedi NR, Vijayashankar NP, Hansda M, Dubey AK, Nwachukwu F, Curran V, et al. Comparing standard setting methods for objective structured clinical examinations in a Caribbean medical school. J Med Educ Curric Dev. 2020;7:2382120520981992. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Dwyer T, Wright S, Kulasegaram KM, Theodoropoulos J, Chahal J, Wasserstein D, et al. How to set the bar in competency-based medical education: standard setting after an objective structured clinical examination (OSCE). BMC Med Educ. 2016;16(1):1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Fu CP, Chi HY, Li MW, Liu CH, Yeh JH, Wang CC, Su CT. Development of an objective structured clinical examination station for pediatric occupational therapy and an evaluation of its quality. Am J Occup Therapy. 2022;76(2):7602205010. [DOI] [PubMed] [Google Scholar]
- 41.George S, Haque MS, Oyebode F. Standard setting: comparison of two methods. BMC medical education. 2006;6(1):46. [DOI] [PMC free article] [PubMed]
- 42.Goss BD, Ryan AT, Waring J, Judd T, Chiavaroli NG, O’Brien RC, Trumble SC, McColl GJ. Beyond selection: the use of situational judgement tests in the teaching and assessment of professionalism. Acad Med. 2017;92(6):780–4. [DOI] [PubMed] [Google Scholar]
- 43.Hasty BN, Lau JN, Tekian A, Miller SE, Shipper ES, Merrell SB, Lee EW, Park YS. Validity evidence for a knowledge assessment tool for a mastery learning scrub training curriculum. Acad Med. 2020;95(1):129–35. [DOI] [PubMed]
- 44.Hess B, Subhiyah RG, Giordano C. Convergence between cluster analysis and the Angoff method for setting minimum passing scores on credentialing examinations. Eval Health Prof. 2007;30(4):362–75. [DOI] [PubMed] [Google Scholar]
- 45.Huang GC, Newman LR, Schwartzstein RM, Clardy PF, Feller-Kopman D, Irish JT, et al. Procedural competence in internal medicine residents: validity of a central venous catheter insertion assessment instrument. Acad Med. 2009;84(8):1127–34. [DOI] [PubMed] [Google Scholar]
- 46.Isezuo S, Kadiri S, Arogundade F, Ogunbiyi A, Bello B, Kolawole W, Ojo O, Ohwovoriole A, Obasohan A, Ogunniyi A, Gwaram B. Assessment reform: moving from fixed passing scores to standard setting based passing scores. Med Teacher. 2025:1–0. 10.1080/0142159X.2025.2515982. [DOI] [PubMed]
- 47.Joseph MN, Chang J, Buck SG, Auerbach MA, Wong AH, Beardsley TD, Reeves PM, Ray JM, Evans LV. A novel application of the modified Angoff method to rate case difficulty in simulation-based research. Simul Healthc. 2021;16(6):e142–50. [DOI] [PubMed] [Google Scholar]
- 48.Jalili M, Hejri SM, Norcini JJ. Comparison of two methods of standard setting: the performance of the three‐level Angoff method. Med Educ. 2011;45(12):1199-208. [DOI] [PubMed]
- 49.Kamath MG, Pallath V, Ramnarayan K, Kamath A, Torke S, Gonsalves J. Standard setting of objective structured practical examination by modified Angoff method: A pilot study. Natl Med J India. 2016;29(3):160. [PubMed] [Google Scholar]
- 50.Kang Y. Evaluating the cutoff score of the advanced practice nurse certification examination in Korea. Nurse Educ Pract. 2022;63:103407. [DOI] [PubMed] [Google Scholar]
- 51.Kardong-Edgren S, Mulcock PM. Angoff method of setting cut scores for high-stakes testing: foley catheter checkoff as an exemplar. Nurse Educ. 2016;41(2):80–2. [DOI] [PubMed] [Google Scholar]
- 52.Kaufman DM, Mann KV, Muijtjens AM, van der Vleuten CP. A comparison of standard-setting procedures for an OSCE in undergraduate medical education. Acad Med. 2000;75(3):267–71. [DOI] [PubMed] [Google Scholar]
- 53.Kim J, Yang JS. How to improve reliability of cut-off scores in dental competency exam: a comparison of rating methods in standard setting. Eur J Dent Educ. 2020;24(4):734–40. [DOI] [PubMed] [Google Scholar]
- 54.Kim DH, Kang YJ, Park HK. Possibility of independent use of the yes/no Angoff and Hofstee methods for the standard setting of the Korean medical licensing examination written test: a descriptive study. J Educ Eval Health Prof. 2022. 10.3352/jeehp.2022.19.33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Klein MR, Schmitz ZP, Adler MD, Salzman DH. Simulation-based mastery learning improves emergency medicine residents’ ability to perform temporary transvenous cardiac pacing. West J Emerg Med. 2022;24(1):43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Konge L, Clementsen P, Larsen KR, Arendrup H, Buchwald C, Ringsted C. Establishing pass/fail criteria for bronchoscopy performance. Respiration. 2012;83(2):140–6. [DOI] [PubMed] [Google Scholar]
- 57.Kramer A, Muijtjens A, Jansen K, Düsman H, Tan L, Van Der Vleuten C. Comparison of a rational and an empirical standard setting procedure for an OSCE. Med Educ. 2003;37(2):132–9. [DOI] [PubMed] [Google Scholar]
- 58.Leask R, Cronje T, Holm DE, van Ryneveld L. Comparing veterinary students’ performance with cut-scores determined using a modified individual Angoff method featuring bloom’s taxonomy. Vet Rec. 2020;187(12):e121. [DOI] [PubMed] [Google Scholar]
- 59.Lee M, Hernandez E, Brook R, Ha E, Harris C, Plesa M, et al. Competency-based standard setting for a high-stakes objective structured clinical examination (OSCE): validity evidence. MedEdPublish. 2018;7:200. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Levett-Jones T, Andersen P, Bogossian F, Cooper S, Guinea S, Hopmans R, et al. A cross-sectional survey of nursing students’ patient safety knowledge. Nurse Educ Today. 2020;88:104372. [DOI] [PubMed] [Google Scholar]
- 61.Lypson ML, Downing SM, Gruppen LD, Yudkowsky R. Applying the bookmark method to medical education: standard setting for an aseptic technique station. Med Teach. 2013;35(7):581–5. [DOI] [PubMed] [Google Scholar]
- 62.Mäkinen M, Axelsson Å, Castrén M, Nurmi J, Lankinen I, Niemi-Murola L. Assessment of CPR-D skills of nursing students in two institutions: reality versus recommendations in the guidelines. Eur J Emerg Med. 2010;17(4):237–9. [DOI] [PubMed] [Google Scholar]
- 63.Malakooti MR, McBride ME, Mobley B, Goldstein JL, Adler MD, McGaghie WC. Mastery of status epilepticus management via simulation-based learning for pediatrics residents. J Grad Med Educ. 2015;7(2):181–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Mikhaeil-Demo Y, Barsuk JH, Culler GW, Bega D, Salzman DH, Cohen ER, et al. Use of a simulation-based mastery learning curriculum for neurology residents to improve the identification and management of status epilepticus. Epilepsy Behav. 2020;111:107247. [DOI] [PubMed] [Google Scholar]
- 65.Miller DT, Zaidi HQ, Sista P, Dhake SS, Pirotte MJ, Fant AL, Salzman DH. Creation and implementation of a mastery learning curriculum for emergency department thoracotomy. West J Emerg Med. 2020;21(5):1258. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Mitzman J, Reynolds M, Panchal A, Yee J. A Pilot Study of a Simulation-Based Mastery Learning Procedural Curriculum for Pediatric Emergency Medicine Fellows. Pediatric Emergency Care. 2024;40:10–97. [DOI] [PubMed] [Google Scholar]
- 67.Montgomery GP, Crockford DN, Hecker K. The coordinators of psychiatric education (COPE) residency in-training exam: a preliminary psychometric assessment. Acad Psychiatry. 2010;34(3):203–7. [DOI] [PubMed] [Google Scholar]
- 68.Moreno-López R, Hope D. Can borderline regression method be used to standard set osces in small cohorts? Eur J Dent Educ. 2022;26(4):686–91. [DOI] [PubMed] [Google Scholar]
- 69.Morrison H, McNally H, Wylie C, McFaul P, Thompson W. The passing score in the objective structured clinical examination. Med Educ. 1996;30(5):345–8. [DOI] [PubMed] [Google Scholar]
- 70.Nitsche JF, Butler TR, Shew AW, Jin S, Brost BC. Optimizing the amount of simulation training used to teach vaginal delivery skills to medical students. Int J Gynaecol Obstet. 2018;140(1):123–7. [DOI] [PubMed] [Google Scholar]
- 71.O’Neill TR, Marks CM, Reynolds M. Re-evaluating the NCLEX-RN® passing standard. J Nurs Meas. 2005;13(2):147–67. [DOI] [PubMed] [Google Scholar]
- 72.Oyeronke EO, Oyinlola CE, Adenike OF, Nwabueze AE. Standard-setting methods for assessment in a post-graduate medical college. Niger Postgrad Med J. 2024;31(3):263–8. [DOI] [PubMed] [Google Scholar]
- 73.Park YS, Kamin C, Son D, Kim G, Yudkowsky R. Differences in expectations of passing standards in communication skills for pre-clinical and clinical medical students. Patient Educ Couns. 2019;102(2):301–8. [DOI] [PubMed] [Google Scholar]
- 74.Park J, Ahn DS, Yim MK, Lee J. Comparison of standard-setting methods for the Korean radiological technologist licensing examination: Angoff, Ebel, bookmark, and hofstee. J Educ Eval Health Prof. 2018. 10.3352/jeehp.2018.15.32. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Park J, Yim MK, Kim NJ, Ahn DS, Kim YM. Similarity of the cut score in test sets with different item amounts using the modified Angoff, modified Ebel, and Hofstee standard-setting methods for the Korean medical licensing examination. J Educ Eval Health Prof. 2020;17:28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Park J. Possibility of using the yes/no Angoff method as a substitute for the percent Angoff method for estimating the cutoff score of the Korean medical licensing examination: a simulation study. J Educ Eval Health Prof. 2022;19:23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Prenner SB, McGaghie WC, Chuzi S, Cantey E, Didwania A, Barsuk JH. Effect of trainee performance data on standard-setting judgments using the mastery Angoff method. J Grad Med Educ. 2018;10(3):301–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Prince KJ, Scherpbier AJ, Van Mameren H, Drukker J, Van Der Vleuten CP. Do students have sufficient knowledge of clinical anatomy? Med Educ. 2005;39(3):326–32. [DOI] [PubMed] [Google Scholar]
- 79.Reid F, Power A, Stewart D, Watson A, Zlotos L, Campbell D, et al. Piloting the united Kingdom ‘prescribing safety assessment’with pharmacist prescribers in Scotland. Res Social Adm Pharm. 2018;14(1):62–8. [DOI] [PubMed] [Google Scholar]
- 80.Rider AC, Miller DT, Ashenburg N, Duanmu Y, Lobo V, Schertzer K, Sebok‐Syer SS. Using a simulated model and mastery learning approach to teach the ultrasound‐guided serratus anterior plane block to emergency medicine residents: a pilot study. AEM Educ Train. 2021;5(3):e10525. [DOI] [PMC free article] [PubMed]
- 81.Rodriguez O, Sánchez-Ismayel A. Development and implementation of an objective structured clinical examination (OSCE) of the subject of surgery for undergraduate students in an institution with limited resources. MedEdPublish. 2021;10:97. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Sam AH, Millar KR, Westacott R, Melville CR, Brown CA. Standard setting very short answer questions (VSAQs) relative to single best answer questions (SBAQs): does having access to the answers make a difference? BMC Med Educ. 2022;22(1):640. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Schneid SD, Armour C, Park YS, Yudkowsky R, Bordage G. Reducing the number of options on multiple-choice questions: response time, psychometrics and standard setting. Med Educ. 2014;48(10):1020–7. [DOI] [PubMed] [Google Scholar]
- 84.Schoonheim-Klein M, Muijtjens A, Habets L, Manogue M, Van Der Vleuten C, Van der Velden U. Who will pass the dental OSCE? Comparison of the Angoff and the borderline regression standard setting methods. Eur J Dent Educ. 2009;13(3):162–71. [DOI] [PubMed] [Google Scholar]
- 85.Schroedl CJ, Frogameni A, Barsuk JH, Cohen ER, Sivarajan L, Wayne DB. Impact of simulation-based mastery learning on resident skill managing mechanical ventilators. ATS Scholar. 2021;2(1):34–48. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Senthong V, Chindaprasirt J, Sawanyawisuth K, Aekphachaisawat N, Chaowattanapanit S, Limpawattana P, Choonhakarn C, Sookprasert A. Group versus modified individual standard-setting on multiple-choice questions with the Angoff method for fourth-year medical students in the internal medicine clerkship. Adv Med Educ Pract. 2013;27:195–200. [DOI] [PMC free article] [PubMed]
- 87.Siddiqui NY, Tarr ME, Geller EJ, Advincula AP, Galloway ML, Green IC, et al. Establishing benchmarks for minimum competence with dry lab robotic surgery drills. J Minim Invasive Gynecol. 2016;23(4):633–8. [DOI] [PubMed] [Google Scholar]
- 88.Smith S, Lobo V, Anderson KL, Gisondi MA, Sebok-Syer SS, Duanmu Y. A randomized controlled trial of simulation‐based mastery learning to teach the extended focused assessment with sonography in trauma. AEM Educ Train. 2021;5(3):e10606. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Stern DT, Ben-David MF, Champlain AD, Hodges B, Wojtczak A, Roy Schwarz M. Ensuring global standards for medical graduates: a pilot study of international standard-setting. Med Teach. 2005;27(3):207–13. [DOI] [PubMed] [Google Scholar]
- 90.Talente G, Haist SA, Wilson JF. A model for setting performance standards for standardized patient examinations. Eval Health Prof. 2003;26(4):427–46. [DOI] [PubMed] [Google Scholar]
- 91.Tappan RS, Roth HR, McGaghie WC. Using simulation-based mastery learning to achieve excellent learning outcomes in physical therapist education. J Phys Therapy Educ. 2023;39:10–97. [DOI] [PubMed]
- 92.Tavakol M, O’Brien D, Stewart C. Determining intra-standard-setter inconsistency in the Angoff method using the three-parameter item response theory. Int J Med Educ. 2023;14:123. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Taylor J, Curtis SD, Onge ES, Egelund EF, Venugopalan V, Whalen K. Implementation of standard setting for high-stakes objective structured clinical examinations. Curr Pharm Teach Learn. 2024;16(6):465–8. [DOI] [PubMed] [Google Scholar]
- 94.Teitelbaum EN, Soper NJ, Santos BF, Rooney DM, Patel P, Nagle AP, Hungness ES. A simulator-based resident curriculum for laparoscopic common bile duct exploration. Surgery. 2014;156(4):880–93. [DOI] [PubMed]
- 95.Toal GG, Gisondi MA, Miller NM, Sebok-Syer SS, Avedian RS, Dixon WW. Simulation-based mastery learning to teach distal radius fracture reduction. Simul Healthc. 2021;16(6):e176-80. [DOI] [PubMed] [Google Scholar]
- 96.Varkey P, Natt N, Lesnick T, Downing S, Yudkowsky R. Validity evidence for an OSCE to assess competency in systems-based practice and practice-based learning and improvement: a preliminary investigation. Acad Med. 2008;83(8):775–80. [DOI] [PubMed] [Google Scholar]
- 97.Verhoeven, der Steeg V, Scherpbier, Muijtjens, Verwijnen, Der Vleuten V. Reliability and credibility of an Angoff standard setting procedure in progress testing using recent graduates as judges. Med Educ. 1999;33(11):832–7. [DOI] [PubMed]
- 98.Verhoeven BH, Verwijnen GM, Muijtjens AM, Scherpbier AJ, Van der Vleuten CP. Panel expertise for an Angoff standard setting procedure in progress testing: item writers compared to recently graduated students. Med Educ. 2002;36(9):860–7. [DOI] [PubMed] [Google Scholar]
- 99.Wayne DB, Fudala MJ, Butter J, Siddall VJ, Feinglass J, Wade LD, McGaghie WC. Comparison of two standard-setting methods for advanced cardiac life support training. Acad Med. 2005;80(10):S63–6. [DOI] [PubMed]
- 100.Wayne DB, Barsuk JH, Cohen E, McGaghie WC. Do baseline data influence standard setting for a clinical skills examination? Acad Med. 2007;82(10):S105–8. [DOI] [PubMed] [Google Scholar]
- 101.Wayne DB, Butter J, Cohen ER, McGaghie WC. Setting defensible standards for cardiac auscultation skills in medical students. Acad Med. 2009;84(10):S94–6. [DOI] [PubMed] [Google Scholar]
- 102.Yim M. Comparison of results between modified-Angoff and bookmark methods for estimating cut score of the Korean medical licensing examination. Korean J Med Educ. 2018;30(4):347. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Yim MK, Shin S. Using the Angoff method to set a standard on mock exams for the Korean nursing licensing examination. J Educ Eval Health Prof. 2020;17:14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Yousef MK, Alshawwa L, Tekian A, Park YS. Challenging the arbitrary cutoff score of 60%: standard setting evidence from preclinical operative dentistry course. Med Teach. 2017;39(sup1):S75–9. [DOI] [PubMed] [Google Scholar]
- 105.Yousef MK, Alshawwa LA, Farsi JM, Tekian A, Yoon Soo P. Determining defensible cut-off scores for dental courses. Eur J Dent Educ. 2020;24(2):186–92. [DOI] [PubMed] [Google Scholar]
- 106.Yousuf N, Violato C, Zuberi RW. Standard setting methods for pass/fail decisions on high-stakes objective structured clinical examinations: a validity study. Teach Learn Med. 2015;27(3):280–91. [DOI] [PubMed] [Google Scholar]
- 107.Yudkowsky R, Downing SM, Wirth S. Simpler standards for local performance examinations: the yes/no Angoff and whole-test Ebel. Teach Learn Med. 2008;20(3):212–7. [DOI] [PubMed] [Google Scholar]
- 108.Yudkowsky R, Tumuluru S, Casey P, Herlich N, Ledonne C. A patient safety approach to setting pass/fail standards for basic procedural skills checklists. Simul Healthc. 2014;9(5):277–82. [DOI] [PubMed] [Google Scholar]
- 109.Meskauskas JA, Norcini JJ. Standard-setting in written and interactive (oral) specialty certification examinations: issues, models, methods, challenges. Eval Health Prof. 1980;3(3):321–60. [DOI] [PubMed] [Google Scholar]
- 110.Cizek G, Bunch M. The Angoff method and Angoff variations. In: Cizek G, Bunch M, editors. Standard setting. Thousand Oaks, California: SAGE Publications, Inc.; 2007a. pp. 81–96. [Google Scholar]
- 111.Kane MT. Validating the interpretations and uses of test scores. J Educ Meas. 2013;50(1):1–73. [Google Scholar]
- 112.Shulruf B, Wilkinson T, Weller J, Jones P, Poole P. Insights into the Angoff method: results from a simulation study. BMC Med Educ. 2016;16(1):134. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data is available with the corresponding author and shall be shared upon request.








