Skip to main content
PLOS One logoLink to PLOS One
. 2023 Oct 5;18(10):e0292306. doi: 10.1371/journal.pone.0292306

Ranking versus rating in peer review of research grant applications

Robyn Tamblyn 1,2,3,*, Nadyne Girard 1, James Hanley 2, Bettina Habib 1, Adrian Mota 4, Karim M Khan 4,5,6, Clare L Ardern 7,8
Editor: Julian D Cortes9
PMCID: PMC10553257  PMID: 37796852

Abstract

The allocation of public funds for research has been predominantly based on peer review where reviewers are asked to rate an application on some form of ordinal scale from poor to excellent. Poor reliability and bias of peer review rating has led funding agencies to experiment with different approaches to assess applications. In this study, we compared the reliability and potential sources of bias associated with application rating with those of application ranking in 3,156 applications to the Canadian Institutes of Health Research. Ranking was more reliable than rating and less susceptible to the characteristics of the review panel, such as level of expertise and experience, for both reliability and potential sources of bias. However, both rating and ranking penalized early career investigators and favoured older applicants. Sex bias was only evident for rating and only when the applicant’s H-index was at the lower end of the H-index distribution. We conclude that when compared to rating, ranking provides a more reliable assessment of the quality of research applications, is not as influenced by reviewer expertise or experience, and is associated with fewer sources of bias. Research funding agencies should consider adopting ranking methods to improve the quality of funding decisions in health research.

Introduction

For decades, the allocation of public funds for research has been predominantly based on peer review. In complex areas of endeavor, such as the medical sciences, it is assumed that the assessment of quality and potential impact is best done by peers who have the in-depth knowledge needed to evaluate the quality of both the scientific team and the proposed research. Although peer review is supported by the academic community [1], it has many vocal critics [2, 3]. There is evidence that peer review is conservative and less likely to fund higher risk, innovative projects, or those that involve a multidisciplinary team of investigators [48]. Perhaps the most distressing aspect of peer review for applicants is the low level of agreement among experts about the quality of the application. The reliability of peer review rating varies from 0.20 to 0.61 (intra-class correlation coefficient) in different studies [914], being somewhat higher for the basic medical sciences (0.41) than the applied sciences (0.32) [11]. The variability in rating translates into a lack of consistency in decision-making about which applications should be funded. Depending on the subset of reviewers selected to review an application, variability in scoring alters the funding decision in approximately 17% to 35% of proposals [10, 1519].

Most funding agencies ask reviewers to rate an application on some form of ordinal scale from poor to excellent, for individual criteria, such as research approach, as well as an overall application score. This form of rating, referred to as “absolute judgment”, requires a reviewer to judge an application against an internalized concept of an “ideal” application. Experienced reviewers who have reviewed hundreds of applications would be expected to have a more robust and stable concept of the “ideal” application than less experienced reviewers [20]. This hypothesis was indirectly supported by a small study of public health reviewers where training improved reliability from an intra-class correlation of 0.61 to 0.89, with the effect being greatest for inexperienced reviewers [12]. In contrast to “absolute” judgment, “relative” judgements do not rely on an internal concept of the ideal application [20]. Instead, reviewers are asked to rank applications from the highest quality to the lowest. Relative judgments do not rely on a stable internalized concept of the ideal application. They are expected to be more reliable, as they are less sensitive to differences in the experience of the reviewers, as well as systematic differences in the harshness or leniency of the reviewer [9, 18, 21] and thus should reduce random variation in judgements. Even within-person variation in judgement appears to be more stable with ranking. For example, comparative ranking of preferences is 20% more stable over time than rating of preferences [22].

Ranking compared to rating also appears to reduce bias in judgements in cross-national studies [23]. This finding is of particular interest as there is evidence of systematic bias in peer review of grant applications. Female applicants who are as qualified and productive as male applicants, are scored systematically lower in both postdoctoral and grant applications [11, 2426], and receive less NIH funding [27]. Similar systematic biases in judgement are observed for black compared to white applicants in NIH competitions, biases that appear to be related to differences in how individual criteria are rated [28, 29].

An opportunity to compare the reliability and potential for bias of rating and ranking occurred when the Canadian Institutes of Health Research (CIHR) introduced a new funding program and scoring approach to support the research of leading scientists and rising stars: the foundation funding program [30]. In this program, reviewers scored and then ranked the applicants in two phases: first the quality of the applicant, and then for the highest-ranking applicants, the quality of the research approach. We evaluated whether ranking as compared to rating improved the reliability of peer review and reduced potential sources of bias for two distinct aspects of the quality of the application.

Methods

Design, population and data sources

A historical cohort study was conducted that included all applications submitted to the CIHR foundation funding program between 2014 when it was first launched to 2017. For each application, the principal investigator’s CIHR number, sex, age, institution, co-investigators, application title, scientific domain (biomedical, clinical, health services, population health), requested budget and duration of funding were retrieved. For each application, we also identified the CIHR number of the review committee chair, all reviewers that were approached to review, and the reviewers’ self-rated conflict of interest, and expertise to review. The study was approved by CIHR senior executive management and CIHR legal counsel. All data were stored on secured servers at McGill University and after linkages, all nominal data were removed.

The foundation funding program

The objective of the foundation funding program was to fund the research programs of leading scientists, both at the early and later career stages, for a period of 5–7 years to provide sustained support and flexibility for innovative, high impact research. In the three-phase review process, top ranked applications in phase 1 were invited to apply to phase 2. The first two phases were conducted remotely through a web platform. In the third phase, applications and reviews were discussed at a multidisciplinary in-person meeting where a decision was made about funding. There were no fixed panels of reviewers. Instead, at each phase of this application-centric review, the application was assigned a set of expert reviewers from the CIHR College of Reviewers, a pool of over 4,000 Canadian and International Scientists who had previously reviewed CIHR funding applications [31], by scientific chairs. A different set of chairs and reviewers was selected for each phase.

Review process and assessment of quality

In the first phase, “virtual” chairs were assigned a set of 1 to 50 applications. The chairs selected 4–5 expert reviewers per application from the College of Reviewers. Reviewers were assigned between 5 to 20 applications to review, and reviewers received their assignments from 1 to 11 chairs. Reviewers assessed the caliber of the applicant using four criteria: vision, impact, productivity, and leadership. Each criterion was rated based on the applicant’s curriculum vitae and a two-page summary of their contributions in each of these domains. Ratings were done using letter grades (poor, fair, good, excellent, excellent+, excellent++, outstanding, outstanding+, outstanding++). Letter grades were converted to a score from 0 (poor) to 28 (outstanding++) per criterion with a maximum score of 112. All the applications rated by a given reviewer were then ranked by score. The reviewer then adjusted the application rank to reflect their judgement of the best to the worst applications and to eliminate ties when applicable. As reviewers rated a different number of applications, ranks were standardized by converting them to percentiles. The mean of all reviewers’ percentile ranks was the final score for an application.

In the second phase, a second set of “virtual” chairs were assigned a set of 1 to 45 phase 2 applications, and they selected 4–5 expert reviewers from the College of Reviewers. Reviewers assessed the quality of the program of research according to the following criteria: research concept, research approach, expertise, mentorship, and environmental support based on a 10-page application. Each criterion was graded using the same letter grades as in phase 1, with a maximum score of 140. However, the criteria were weighted differently, with 50% of the weight assigned to the research concept and approach, 40% to expertise and mentoring, and 10% to environmental support. Similar to phase 1, letter grades were converted to a score that was used to establish a preliminary rank order. The reviewer then modified the ranked applications from best to worst. The reviewer-adjusted ranks were converted to percentiles, and the mean of all reviewer ranks was the final score for the application. The 5 reviewers for an application did not necessarily have the same applications to review so it was assumed that they would, on average, have an equivalent distribution of poor, good and excellent applications.

In both phase 1 and 2, applications were flagged for asynchronous discussion if there was differences in scores and ranks of greater than one standard deviation, or if the virtual chair identified comments/ questions raised by one or more reviewers that warranted discussion.

Outcomes

To assess the reliability and potential bias in rating and ranking methods of assessment, we used the overall rating and percentile rank scores assigned to an application by the different sets of reviewers and review processes in phase 1 and 2. In addition, for the rating method, we assessed reliability using reviewers’ assessment of the individual criteria to determine if there was greater agreement among reviewers for some criteria than others. To assess potential predictors of reliability, we calculated the between-rater variance for an application and used the log of this variance as the outcome variable. As we had no “gold standard” assessment of the quality of the application, we used measures of the applicant’s scientific productivity to assess potential predictors of bias among applicants with equivalent productivity.

Potential predictors of reliability and bias

Applicant characteristics

Applicant age, sex, and institution were retrieved from the demographic information provided in the application. Applicant scientific productivity was measured using the H-index. The H-Index measures the impact of the applicant’s cumulative research contributions based on citations, allowing unbiased comparison among applicants competing for the same resources [32]. To calculate the H-index, we used the principal applicant’s first, middle, and last name and institution to retrieve all publications from the Web of Science database where the applicant was listed as an author up to and including the year in which the application was submitted. For each publication, we retrieved the citation reports by linking the ISSN of the journal to the Journal Citation Record file. When there was no recorded ISSN, we used the full and abbreviated journal name to make the link. The applicant was classified as an early career investigator if the application was submitted within 5 years of their first university/institute research appointment, as a dedicated stream of funding was available for this group. As the H-index can be positively biased by the number of collaborators on an applicant’s publications [33], we also counted the number of unique collaborators associated with the applicants publications.

Application characteristics

The year in which the application was submitted was included as the foundation scheme was a new program and the review process was expected to improve with applicant and reviewer experience. As there was a change in policy after the first competition in 2014 to provide equivalent opportunity for male and female applicants to progress to phase 2, we classified the application submission year as 2014 or after 2014 [25]. The content domain of the application was included, as reliability tends to be better for biomedical applications compared to clinical, health services and population health applications [11, 13].

Review characteristics

As the role of the chair is to facilitate unbiased high quality reviews, discussion and consensus on the quality of an application, where possible, we measured, for each application, three characteristics of the chair that may influence their performance. These included: 1) the number of reviewers the chair was responsible for, 2) the number applications the chair was assigned, and 3) the number of years of experience the chair had since 2000 as a reviewer. With respect to reviewers, for each application, we measured: 1) the number of reviewers assigned to an application, 2) the mean number of years of review experience since 2000 of all reviewers assigned to an application, as greater experience should be associated with better reliability for rating [20], 3) the sex mix of reviewers, as female reviewers score systematically lower than male reviewers so a mix of male and female reviewers may decrease reliability [11], 4) the proportion of reviewers with high self-assessed expertise to review an application, 4) the proportion of reviewers whose prior applications were in the same scientific domain as the applicant, and 5) the workload of the reviewers of an application, measured as the mean of the number of applications each reviewer was assigned to review. With respect to the review process, we measured whether there was on-line discussion (yes/no) using data retrieved from the on-line review system. We also measured the proportion of reviewers who declined to review an application because of conflicts of interest, as conflicts in the review panel have been associated with higher scores [5, 11, 24, 34, 35], a phenomenon that may not apply when reviews are done virtually rather than in-person.

Data analysis

To assess reliability of rating and ranking, we estimated the intra-class correlation coefficient (ICC) for the overall rating and percentile ranking scores separately for applications in the phase 1 and phase 2 review stages. The ICC for ratings of individual criteria were also estimated. Bootstrapping was used to estimate 95% confidence intervals. To estimate the association between applicant, application and review characteristics and within-rater variance of application scores, we used generalized estimating equation multiple linear regression to account for clustering of applications within applicants (i.e. repeated submissions). Application was the unit of analysis. To assess potential bias in review by rating compared to ranking, we created two multivariate models, one using overall mean rating of the application as the outcome and the second using overall mean percentile ranking. The H index of the principal applicant and the log of the number of collaboraters were included in the model to assess the impact of other attributes of the applicant, application and review process that may have influenced scores among applicants with equivalent scientific productivity. As the H-index in prior research has been shown to have a non-linear relationship to the application score [11], we tested whether including the quadratic term for the H-index improved model fit. All measured characteristics of the applicant, application and review process were included in the model. Among scientists with equivalent scientific productivity, we tested the hypothesis that the overall rating and ranking score may be modified by applicant sex or age by including the two-way interaction terms (H-index* applicant sex, H-index*applicant age), in separate multivariate models. All analyses were conducted using SAS, version 9.41M5.

Results

Study population characteristics

In the 4 years the foundation program was offered, 3,156 applications were submitted by 2,249 investigators. Female investigators accounted for 32.6% of applications and 38.3% of applicants were aged 41 to 50 years (Table 1). Early career investigators submitted 29.8% of applications. The mean H-index of applicants was 13.5 (SD 9.9), and 24.3% had an H-index of greater than 19 in the year they applied. The mean number of collaboraters in phase 1 was 507.5 (SD 663.7) and in phase 2 it was 763.1 (SD 860.4). Most applications were submitted in 2014 (42.6%) or 2015 (28.8%) and were in the basic science domain (58.5%).

Table 1. Applicant, application, and review characteristics.

Phase 1 Leadership & Vision Phase 2 Research Methods Phase 1 Leadership & Vision Phase 2 Research Methods
N (%) N (%) N (%) N (%)
Applicant Characteristics Reviewer Workload
Sex Chair–# of reviewers, mean (SD) 31.1 (13.8) 28.5 (17.1)
    Male 2126 (67.4%) 794 (72.4%) Chair–# of applications, mean (SD) 23.3 (10.9) 13.6 (10.2)
    Female 1030 (32.6%) 302 (27.6%) Reviewer–# of applications, mean (SD) 12.6 (3.2) 9.5 (2.1)
Age Reviewer characteristics
    <40 years old 807 (25.6%) 201 (18.3%) CIHR Experience
    41–50 years old 1208 (38.3%) 380 (34.7%) < = 2 years 442 (14.0%) 315 (28.7%)
    51–60 years old 830 (26.3%) 347 (31.7%) >2–3 1724 (54.6%) 485 (44.3%)
    >60 years 311 (9.9%) 168 (15.3%) >3–4 810 (25.7%) 214 (19.5%)
Early Career Investigator >4 180 (5.7%) 82 (7.5%)
    Yes 939 (29.8%) 219 (20.0%) % of Female Reviewers
    No 2217 (70.2%) 877 (80.0%) <50% 2010 (63.7%) 814 (74.3%)
H-Index 50%-79% 875 (27.7%) 225 (20.5%)
    < = 6 792 (25.1%) 137 (12.5%) 80% more 271 (8.6%) 57 (5.2%)
    >6-< = 12 854 (27.1%) 213 (19.4%) Mean age of the reviewer (SD) 48 (9.5) 50 (9.7)
    >12-< = 19 744 (23.6%) 280 (25.5%) Review Process
    >19 766 (24.3%) 466 (42.5%) % Reviewers with High Expertise
N of Collaborators >80% 317 (10.0%) 281 (25.6%)
    < = 122 794 (25.2%) 137 (12.5%) 60%-<80% 536 (17.0%) 208 (19.0%)
    >122-< = 302 786 (24.9%) 195 (17.8%) <60% 2303 (73.0%) 607 (55.4%)
    >302-< = 643 789 (25.0%) 317 (28.9%) % Reviewers with Applications in Same Domain
    >643 787 (24.9%) 447 (40.8%) >80% 504 (15.9%) 310 (28.3%)
Application Characteristics 60%-<80% 589 (18.7%) 145 (13.2%)
Year Submitted <60% 2063 (65.4%) 641(58.5%)
    2014 1343 (42.6%) 445 (40.6%) On-Line Discussion
    2015 910 (28.8%) 260 (23.7%) Yes 607 (19.2%) 396 (36.1%)
    2016 600 (19.0%) 228 (20.8%) No 2549 (80.8%) 700 (63.9%)
    2017 303 (9.6%) 163 (14.9%) Conflicts of Reviewers Approached
Scientific Domain None 720 (22.8%) 35 (3.2%)
    Biomedical 1847 (58.5%) 639 (58.3%) At least 1 2436 (77.2%) 1061 (96.8%)
    Clinical 590 (18.7%) 223 (20.3%)
    HSR 315 (10.0%) 101 (9.2%)
    PPH 395 (12.5%) 132 (12.0%)

CIHR = Canadian Institutes of Health Research, SD = standard deviation, HSR = Health services research, PPH = Population & public health

Phase 1, Leadership & Vision: Number of principal applicants = 2,249, Number of applications = 3,156

Phase 2, Research Methods & Environment: Number of principal applicants = 863, Number of applications = 1,096

Virtual Chairs were responsible for a mean of 23.3 applications and 31.1 reviewers in phase 1, and 13.6 applications and 28.5 reviewers in phase 2 (Table 1). The majority of applications in phase 1 (66.9%) and phase 2 (75.0%) were reviewed by 5 reviewers. Each reviewer had a mean of 12.6 (range: 5–20) applications to review in phase 1, and 9.5 (range: 4–14) in phase 2. The majority of reviewers had 2 to 3 years of CIHR review experience. In phase 1, 73.0% of applications had reviewers with less than 60% higher expertise to review; 55.4% in phase 2. There was on-line discussion for 19.2% of applications in phase 1 and 36.1% of applications in phase 2.

Application rating and ranking characteristics

The mean overall score for phase 1 was 85.1 out of 128 and percentile rank was 49.4%, and for phase 2 it was 115.3 and 50.0%, respectively (Table 2). The correlation between the overall rating and rank in phase 1 was 0.83, and in phase 2 it was 0.78. The highest scoring criterion in phase 1 was for leadership (mean 21.9) and in phase 2 it was for the environment (mean 24.7). The item to total correlations for criterion rating in phase 1 was (Chronbach alpha) 0.91, and in phase 2 it was 0.86.

Table 2. Application scores and percentile ranks.

Application Score Mean (SD)
Phase 1 Leadership & Vision
Initial Score 85.1 (10.5)
 Leadership 21.9 (3.0)
 Vision 20.2 (3.8)
 Productivity 21.6 (3.1)
 Impact 21.4 (3.1)
Percentile Rank 49.4 (16.6)
Phase 2 Research Methods
Initial Score 115.3 (10.8)
 Research Concept 22.2 (3.1)
 Research Approach 21.0 (3.4)
 Expertise 24.2 (2.5)
 Mentorship 23.2 (2.9)
 Environment 24.7 (2.3)
Percentile Rank 50.0 (21.7)

Reliability of rating and ranking

The overall reliability of assessment (ICC) for phase 1 was 0.54 for rating and 0.59 for ranking (Table 3). For individual criteria rated in phase 1, the highest ICC was for productivity (0.50) and the lowest was for vision (0.35). In phase 1, on average, reviewers had to break ties in overall score for 31.7% (SD 20.0%) of their applications, and 84% of reviewers had to break at least one tie. For Phase 2, the overall reliability was 0.25 for rating and 0.38 for ranking. For the 1096 (34.7%) of applicants invited to apply to phase 2, reliability was highest for research approach (0.23), and lowest for support (0.12). In phase 2 reviewers had to break ties in 28.3% (SD 24.1%) of their applications, and 75.4%of reviewers had to break at least one tie.

Table 3. Reliability (ICC) of rating and ranking.

Characteristics ICC (95% CI) Variance Between Applications Rater Variance Within Applications
Phase 1 Leadership & Vision
Percentile Rank 0.59 (0.58; 0.61) 3181.10 (3076.72; 3289.27) 406.49 (394.38; 418.69)
Initial Rating 0.54 (0.52; 0.55) 1010.88 (956.37; 1065.68) 156.22 (150.61; 162.10)
Rating Criteria
 Impact 0.46 (0.44; 0.48) 69.16 (65.26; 73.01) 13.63 (13.10; 14.20)
 Leadership 0.48 (0.47; 0.50) 74.66 (70.41; 78.74) 13.74 (13.15; 14.34)
 Productivity 0.50 (0.48; 0.51) 78.00 (73.50; 82.62) 13.87 (13.32; 14.45)
 Vision 0.35 (0.33; 0.37) 70.83 (67.07; 74.26) 19.89 (19.23; 20.54)
Phase 2 Research Methods
Percentile Rank 0.38 (0.35; 0.40) 2614.48 (2445.57; 2773.62) 670.83 (641.82; 700.09)
Initial Rating 0.25 (0.22; 0.28) 501.15 (458.34; 544.38) 191.28 (177.63; 204.82)
Rating Criteria
 Expertise 0.19 (0.16; 0.22) 22.65 (20.71; 24.67) 10.62 (9.77; 11.51)
 Mentorship 0.17 (0.14; 0.21) 26.55 (23.65; 29.69) 13.22 (12.36; 14.17)
 Research Approach 0.23 (0.20; 0.26) 43.78 (40.42; 47.46) 17.98 (16.86; 19.17)
 Research Concept 0.21 (0.18; 0.24) 37.07 (33.83; 40.22) 16.09 (15.02; 17.20)
 Support 0.12 (0.09; 0.15) 15.44 (13.73; 17.25) 9.37 (8.62; 10.17)

ICC = Intra-class correlation coefficient

Applicant, application, and review characteristics associated with within rater variance in rating and ranking

In phase 1, the applicant’s H-index was significantly associated with rater variance—for each one point increase in the H-index, the rater variance was reduced by -4,19% (p<0.01) in rating and by -2.25% (p<0.01) in ranking (Table 4). Compared to applications in biomedical science, there was significantly greater rater variance of 25.8% (p = 0.01) for ranking but not for rating (8.63%, p = 0.13) for those in applied science. Review characteristics influenced variance of rating but not ranking. Rater variance decreased when a greater proportion of the reviewers had more experience (-9.62%, p<0.01), were in the same scientific domain (-30.05%, p<0.01), and were more likely to have conflicts (-20.51% p<0.01). A greater number of applications assigned to a reviewer and increasing reviewer age were associated with a significant increase in reviewer variance. In contrast, on-line discussion was the only review characteristic associated with increased rater variance in ranking.

Table 4. Association between applicant, application and review characteristics and rater variance for the two methods: Rating and ranking.

Phase 1 Leadership & Vision Phase 2 Research Methods
Variance Rating Variance Percentile Rank Variance Rating Variance Percentile Rank
Percent change (95% CI) P-Value Percent change (95% CI) P-Value Percent change (95% CI) P-Value Percent change (95% CI) P-Value
Applicant Characteristics
    H index -4.19 (-4.79; -3.58) <0.01 -2.45 (-3.95; -0.93) <0.01 -1.40 (-2.1; -0.7) <0.01 -0.65 (-2.25; 0.98) 0.43
    Total N collaborators (log) 2.43 (1.67; 3.19) <0.01 1.41 (0.11; 2.74) 0.03 1.13 (-0.1; 2.3) 0.06 0.41 (-1.09; 1.93) 0.59
    PI age -1.15 (-1.76; -0.54) <0.01 -0.65 (-1.69; 0.40) 0.22 -0.75 (-1.6; 0.1) 0.09 -1.44 (-2.85; -0.00) 0.05
    Sex
        Male Ref Ref
        Female 4.51 (-4.84; 14.79) 0.36 -13.17 (-26.63; 2.77) 0.10 15.95 (0.2; 34.2) 0.05 8.59 (-5.99; 25.43) 0.26
    Early Career
        No Ref Ref
        Yes 9.30 (-2.19; 22.13) 0.12 0.63 (-17.34; 22.51) 0.95 -8.23 (-24.97; 12.25) 0.4 -21.28 (-41.0; 4.9) 0.1
Application Characteristics
    Scientific Domain
        Basic Science Ref Ref
        Applied Science 8.60 (-2.19; 20.59) 0.12 25.83 (5.59; 49.95) 0.01 36.93 (14.9; 63.2) <0.01 10.92 (-5.26; 29.85) 0.20
    Competition Year
        2014 Ref Ref
        2015 + 7.17 (-12.63; 31.47) 0.51 40.55 (-4.25; 106.31) 0.08 -41.04 (-54.1; -24.3) <0.01 -13.21 (-36.5; 18.6) 0.37
Reviewer Workload
# of reviewers per chair 0.04 (-0.32; 0.41) 0.81 0.09 (-0.48; 0.66) 0.76 0.17 (-0.2; 0.5) 0.38 0.09 (-0.35; 0.53) 0.69
Mean # applications per reviewer 4.93 (1.49; 8.48) <0.01 5.05 (-0.68; 11.11) 0.08 3.00 (-2.5; 8.8) 0.29 -1.02 (-8.43; 6.98) 0.80
Reviewer Characteristics
    Mean age reviewers 1.17 (0.41; 1.94) <0.01 0.01 (-1.25; 1.28) 0.99 -1.74 (-2.9; -0.6) <0.01 -0.01 (-1.18; 1.17) 0.98
    % female reviewers 0.26 (-15.23; 18.57) 0.98 -2.27 (-29.54; 35.57) 0.89 21.63 (-9.6; 63.7) 0.2 30.71 (-8.19; 86.08) 0.14
    Mean years of experience for reviewer -9.62 (-15.09; -3.80) <0.01 0.62 (-10.56; 13.20) 0.92 -4.75 (-13.6; 5.0) 0.33 6.65 (-7.01; 22.32) 0.36
    Pct high expertise reviewers -14.23 (-27.35; 1.26) 0.07 -15.58 (-34.96; 9.59) 0.20 -29.14 (-45.3; -8.2) 0.01 -35.35 (-53.90; -9.32) 0.01
    % reviewers with applications in same domain -30.05 (-41.36; -16.55) <0.01 -10.69 (-34.38; 21.54) 0.47 -15.26 (-36.4; 12.9) 0.26 -33.75 (-51.65; -9.22) 0.01
Review Process
    Conflicts on review panel
        None Ref Ref Ref Ref
        At least one -20.51 (-27.49; -12.86) <0.01 12.54 (-3.47; 31.21) 0.13 -26.99 (-48.8; 4.1) 0.08 -16.39 (-38.50; 13.68) 0.25
    On-Line Discussion
        No Ref Ref Ref Ref
        Yes 0.07 (-11.28; 12.88) 0.99 164.15 (124.1; 211.2) <0.01 24.92 (4.3; 49.6) 0.02 96.58 (48.4; 160.4) <0.01

The multiple regression models were specified as follows

Phase 1 & 2: Y(variancerating)=x0(intercept)+x1(Hindex)+x2(PIage)+x3(femalePI)+x4(earlycareer)+x5(appliedscience)+x6(2015competition)+x7(#reviewers/chair)+x8(mean#applicationsperreviewer)+x9(meanagereviewers)+x10(%femalereviewers)+x11(meanreviewerexperience)+x12(%highexpertisereviewers)+x13(%reviewersinsamescientificdomain)+x14(conflictonthepanel)+x15(onlinediscussion)+x16(logncollaborators)+x(residualvariance)

Phase 1 & 2: Y(varianceranking)=x0(intercept)+x1(Hindex)+x2(PIage)+x3(femalePI)+x4(earlycareer)+x5(appliedscience)+x6(2015competition)+x7(#reviewers/chair)+x8(mean#applicationsperreviewer)+x9(meanagereviewers)+x10(%femalereviewers)+x11(meanreviewerexperience)+x12(%highexpertisereviewers)+x13(%reviewersinsamescientificdomain)+x14(conflictonthepanel)+x15(onlinediscussion)+x16(logncollaborators)+x(residualvariance)

Similar to phase 1, in phase 2, rating compared to ranking was more likely to be influenced by review, applicant, and application characteristics (Table 4). Variance in rating was greater for female applicants and those whose applications were in the applied compared to the biomedical science domain. A higher H-index was associated with a reduction in variance, but only significantly so for rating compared to ranking. Reductions in the variance for both rating and ranking were seen in the subsequent 2015–2017 competitions compared to 2014, although only significant for rating. Variance in rating and ranking was reduced when a greater proportion of reviewers had high expertise. On-line discussion was a marker of disagreement, being associated with increased variance for both rating and ranking.

Applicant, application and review characteristics associated with systematic bias in rating and ranking

In the assessment of possible bias, we found the expected positive association between the applicant’s H-index and overall application rank and score for both phase 1 and 2 (Table 5). The quadratic H-index term was significant, indicating that the H-index was associated with higher application scores, the slope of the increase decreasing at the upper end of the distribution. After adjusting for H-index, number of collaboraters and other application and review characteristics, older applicants received higher scores and ranks in phase 1 but applicant age had no impact in phase 2. Early career investigators received lower scores and ranks in both phase 1 and 2, although not significantly so for ranking in phase 1. Female applicants received significantly lower ratings in phase 1, but there was no impact of applicant sex in phase 2 for either rating or ranking. Of interest, there was a significant interaction between H-index and applicant sex in phase 1 rating (Fig 1). Female applicants were rated lower than male applicants with lower values of the H-index, with rating of female applicants being equivalent or higher than male applicants at the upper end of the H index distribution. None of the other hypothesized interactions between applicant and reviewer sex were significant.

Table 5. Association between applicant, application and review characteristics and final application score for the two methods: Rating and ranking.

Phase 1 Leadership & Vision Phase 2 Research Methods
Final Rating Final Percentile Rank Final Rating Final Percentile Rank
Estimate (95% CI) P-Value Estimate (95% CI) P-Value Estimate (95% CI) P-Value Estimate (95% CI) P-Value
Applicant Characteristics  
    H index 1.31 (1.12; 1.50) <0.01 2.17 (1.86; 2.48) <0.01 0.20 (0.00; 0.40) 0.05 0.61 (0.10; 1.12) 0.02
    H index * H index -0.01 (-0.02; -0.01) <0.01 -0.02 (-0.03; -0.01) <0.01 -0.14 (-0.28; 0.01) 0.07 -0.47 (-0.85; -0.08) 0.02
    Total N collaborators (log) -0.46 (-0.57; 0.34) <0.011 -0.87 (-1.07; -0.67) <0.01 0.00 (-0.00; 0.00) 0.92 0.00 (-0.01; 0.01) 0.88
    PI age 0.14 (0.07; 0.21) <0.01 0.32 (0.19; 0.46) <0.01 0.02 (-0.08; 0.11) 0.74 0.07 (-0.17; 0.31) 0.58
    Sex
        Male Ref Ref Ref Ref
        Female -1.62 (-2.77; -0.46) <0.01 -1.44 (-3.59; 0.72) 0.199 -0.17 (-1.74; 1.40) 0.83 1.67 (-2.21; 5.54) 0.4
    Early Career
        No Ref Ref Ref Ref
        Yes -2.66 (-4.21; -1.11) <0.01 -0.04 (-2.82; 2.73) 0.98 -3.23 (-5.46; -1.01) < 0.01 -7.21 (-13.01; -1.41) 0.01
Application Characteristics
    Scientific Domain
        Basic Science Ref Ref Ref Ref
        Applied Science -0.36 (-1.57; 0.84) 0.55 0.38 (-2.02; 2.78) 0.76 -4.65 (-6.49; -2.81) < 0.01 -4.44 (-9.05; 0.17) 0.06
    Competition Year
        2014 Ref Ref Ref Ref
        2015 + -0.11 (-2.19; 1.96) 0.92 -11.42 (-15.55; -7.29) <0.01 2.93 (0.49; 5.38) 0.02 -7.97 (-14.39; -1.55) 0.02
Reviewer Workload
    # of reviewers per chair 0.11 (0.07; 0.14) <0.01 0.33 (0.25; 0.40) <0.01 0.03 (0.00; 0.07) 0.04 0.04 (-0.05; 0.13) 0.34
    Mean # applications per reviewer -0.42 (-0.76; -0.08) 0.02 0.53 (-0.14; 1.19) 0.12 -0.24 (-0.82; 0.34) 0.42 0.51 (-0.92; 1.94) 0.48
Reviewer Characteristics
    Mean age reviewers -0.13 (-0.21; -0.05) <0.01 -0.26 (-0.42; -0.09) <0.01 0.16 (0.03; 0.28) 0.01 0.17 (-0.15; 0.50) 0.29
    % female reviewers 0.59 (-1.27; 2.45) 0.53 1.48 (-2.19; 5.16) 0.43 -1.57 (-4.42; 1.29) 0.28 -0.54 (-8.02; 6.94) 0.89
    Mean years of experience for reviewer -0.45 (-1.00; 0.09) 0.10 -2.12 (-3.34; -0.89) <0.01 0.34 (-0.50; 1.19) 0.43 -1.22 (-3.55; 1.10) 0.3
    Pct high expertise reviewers 1.45 (-0.11; 3.01) 0.077 -0.15 (-3.52; 3.23) 0.93 3.29 (0.94; 5.65) < 0.01 6.17 (-0.09; 12.44) 0.05
    % reviewers with applications in same domain 0.92 (-0.75; 2.59) 0.28 -1.10 (-4.76; 2.56) 0.55 1.97 (-0.74; 4.68) 0.15 2.38 (-4.83; 9.60) 0.52
Review Process
    Conflicts on the review panel
        None Ref Ref Ref Ref
        At least 1 6.28 (5.14; 7.41) <0.01 10.69 (8.62; 12.75) <0.01 4.37 (-0.09; 8.84) 0.06 9.19 (-0.96; 19.34) 0.08
    On-Line Discussion
        No Ref Ref Ref Ref
        Yes 1.74 (0.82; 2.65) <0.01<0.012 2.39 (-0.02; 4.80) 0.052 0.29 (-1.22; 1.79) 0.71 5.76 (1.15; 10.37) 0.01

The multiple regression models were specified as follows

Phase 1 & 2: Y(finalrating)=x0(intercept)+x1(Hindex)+x2(Heindex*Hindex)+x3(PIage)+x4(femalePI)+x5(earlycareer)+x6(appliedscience)+x7(2015competition)+x8(#reviewers/chair)+x9(mean#applicationsperreviewer)+x10(meanagereviewers)+x11(%femalereviewers)+x12(meanreviewerexperience)+x13(%highexpertisereviewers)+x14(%reviewersinsamescientificdomain)+x15(conflictonthepanel)+x16(onlinediscussion)+x17(logncollaborators)+x(residualvariance)

Phase 1 &2: Y(finalranking)=x0(intercept)+x1(Hindex)+x2(Heindex*Hindex)+x3(PIage)+x4(femalePI)+x5(earlycareer)+x6(appliedscience)+x7(2015competition)+x8(#reviewers/chair)+x9(mean#applicationsperreviewer)+x10(meanagereviewers)+x11(%femalereviewers)+x12(meanreviewerexperience)+x13(%highexpertisereviewers)+x14(%reviewersinsamescientificdomain)+x15(conflictonthepanel)+x16(onlinediscussion)+x17(logncollaborators)+x(residualvariance)

Fig 1. Association between H-index and final application rating, by sex.

Fig 1

Review characteristics influenced both the overall rating of an application and percentile rank in phase 1 and 2 (Table 5). With respect to workload, the number of reviewers the chair was responsible for was positively associated with higher ratings and ranking of an application in phase 1, and marginally higher ratings but not ranking in phase 2. The number of applications the reviewer was assigned was associated with significantly lower ratings but had no impact on ranking. This association was only evident in phase 1, where the mean number of applications per reviewer was 12.6 compared to 9.5 for phase 2. With respect to reviewer characteristics, having a higher percentage of reviewers with high expertise was associated with higher scores and ranks in phase 2. In contrast, reviewers whose applications were in the same domain as the applicant did not influence application score or rank. Although the sex mix of the reviewers did not significantly influence application score or rank, older reviewers were more likely to score and rank an application lower than younger reviewers in phase 1, as were reviewers with more experience. Finally, on-line discussion of an application was associated with higher scores and ranks in phase 1 and 2. Of note, application scores and ranks were higher in both phase 1 and 2 when one or more invited reviewers declared a conflict, although only significant for phase 1.

Discussion

This study provided one of the first opportunities to compare the reliability and potential sources of bias of two methods of peer review scoring—rating and ranking—in a national funding competition at the Canadian Institutes of Health Research. We found that ranking was more reliable than rating, and less susceptible to characteristics of the review panel such as level of expertise and experience for both reliability and potential sources of bias. However, both rating and ranking penalized early career investigators and favoured older applicants. Sex bias was evident only for rating and when the applicant’s H-index was at the lower end of the H-index distribution.

Theoretically, the reliability of rating would be more likely influenced by reviewer experience than that of ranking, as more experienced reviewers would have a more stable conceptual model of the “ideal application” against which an application could be judged [20]. Our findings are consistent with this hypothesis. Greater reviewer experience was positively associated with smaller rater variance for rating, but had no influence on ranking. Of note, rater variance was also smaller when there was less uncertainty of an applicant’s track record, with lower levels of rater variance for more senior investigators and those with a higher H-index. Rater variance was significantly higher for early career investigators, although these effects were less pronounced for ranking compared to rating. Consistent with prior research [9, 18, 21, 36], reliability was higher for assessment of the researcher (phase 1), than for the research project (phase 2) for both rating and ranking. In this study, lower reliabilities for the research project may also be related to the selection of only the best applications to proceed to phase 2. Even though applications were rated by an independent set of reviewers in phase 2, using a different set of criteria, the quality of the candidates and applications selected to phase 2 may be more homogenious than in phase 1 making it more difficult for reviewers to distinguish amongst higher quality proposals. In this context, ranking outperformed rating, likely because it forced reviewers to distinguish among applications given the same score by breaking ties among applications with the same absolute score. This would effectively increase the between application variance relative to the within rater variance thereby increasing reliability. The mean number of applications with ties was substantial in both phase 1 and 2 and most reviewers had at least one tie that needed to be broken in both phases. The exercise of tie breaking likely allowed reviewers to reflect on more nuanced differences between applications, and possibly to correct for more extreme high or low ratings of some applicants.

As the majority of funding agencies use peer review rating to assess the quality of an application and determine funding, the reliability of peer review should be improved by using a ranking rather than a rating system. Ranking would also provide greater flexibility in recruiting reviewers from various disciplines with different levels of experience without compromising reliability. Ranking may also confer reductions in potential sources of bias in application assessment. In our study ranking compared to rating significantly reduced the negative bias in scoring for female and early career applicants.

In this study, the reliability of ranking could have been improved if reviewers had been given an equivalent number of applications to review. To adjust for inequities in the number of applications reviewed, ranks were converted to percentile ranks, which added some “noise” to the application score, reducing the ICC based on phase 1 ranking from 0.6 to 0.59. For example, a ranking of second best for a reviewer who had 5 applications to review and rank would result in a percentile rank of 80%, whereas a rank of second for a reviewer who had 12 applications would result in a percentile rank of 91.7%. Another factor that may have contributed to random measurement error in ranking is an unequal distribution of the quality of applications assigned to each reviewer, a challenge that should be addressed in future research of peer-review ranking methods.

Similar to other studies [5], we found that older applicants received higher scores and ranks compared to younger applicants with the same H-index. Also, extending to others’ findings [37], we found female applicants were more likely to receive lower scores and ranks on their applications. Our analysis provides some insight into the reason for these differences, as they pertain only to male and female applicants with the same H-index but at the lower end of the H-index distribution. At higher values of the H-index, female applicants perform the same or better than males in terms of final application ranks and scores. The policy implications of these findings are important, as these unconscious biases in rating will influence those at early career stages, potentially discouraging female applicants from continuing to pursue a scientific career. The Matthew effect, where early failure in application for research funding among equivalent candidates discourages subsequent efforts for funding, may be one of the reasons why female scientists are less likely to progress to more senior positions in their research career [24, 38, 39]. Indeed, many countries have noted an over-representation of males in senior academic and scientific positions and awards, leading to proactive policies to support equity, diversity and inclusion in all areas of science [4042].

With respect to the review process, we found that on-line discussion of an application was associated with higher rater variance, which may be expected, as efforts would be selectively made to resolve differences in review opinion among reviewers with very divergent scores and ranks. Of interest, on-line discussion had a positive effect, significantly increasing the scores and ranks for applications that were discussed on-line compared to those that were not. The other interesting finding was the effect of conflicts. In prior research, we and others have shown that conflicts on a panel with an application are associated with higher scores [5, 11, 24, 34, 35], even when those in conflict do not participate in scoring or the discussion. It has been assumed that this phenomenon is related to cronyism, the influence of professional networks or group dynamics within the peer review meetings [5, 24, 34, 35]. We also note the same positive effects of conflicts on scores, but as reviews were done asynchronously and virtually, and no reviewer in conflict was involved in scoring or discussion, these prior theories for the positive effects of conflicts are not supported. We speculate that applications where there are reviewer conflicts are being submitted by scientists in highly productive groups or networks, and if this is the case, conflicts may be a marker of the quality of an application. This hypothesis needs to be evaluated in future research.

There are important limitations to consider in interpreting our findings. First, we had no direct measurement of the quality of an application. We used the principal applicant’s H-index as a proxy for application quality, and while it was strongly associated with the final application score and rank, we have no direct measure if its validity. Second, we do not know what the impact is of rating an application first and then ranking, as opposed to ranking and then rating. We suspect that we may have underestimated the benefits of ranking, and future research should address this potential bias. Finally, we had no information about other applicant characteristics, such as race or disability, which have been associated with bias in assessment [29, 43].

In conclusion, ranking compared to rating provides a more reliable assessment of the quality of research applications, was not as influenced by reviewer expertise or experience, and was associated with fewer potential sources of bias.

Data Availability

The datasets analyzed in this study are held by the Canadian Institutes for Health Research (CIHR) and are not publicly available due to privacy and legal restrictions. Researchers wishing to obtain access to these data need to contact the Vice-President of Research Programs-Operations at CIHR (christian.baron@cihr-irsc.gc.ca) to obtain approval to access de-identified data on foundation funding program applications submitted between 2014 and 2017.

Funding Statement

Funding was provided by the Canadian Institutes of Health Research (CIHR). The study sponsor approved the use of the data, and the original manuscript as per CIHR policy. CIHR had no role in the design of the study, the analysis of interpretation of the data, or the the review and approval of the final published manuscript.

References

  • 1.Bornmann L. Scientific peer review. Annual Review of Information Science and Technology. 2011;45(1):197–245. [Google Scholar]
  • 2.Roy R. Funding Science: The Real Defects of Peer Review and An Alternative To It. Science, Technology, & Human Values. 1985;10(3):73–81. [Google Scholar]
  • 3.Bendiscioli S. The troubles with peer review for allocating research funding. EMBO reports. 2019;20(12):e49472. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Luukkonen T. Conservatism and risk-taking in peer review: Emerging ERC practices. Research Evaluation. 2012;21(1):48–60. [Google Scholar]
  • 5.Guthrie S, Ghiga I, Wooding S. What do we know about grant peer review in the health sciences? F1000Res. 2017;6:1335. doi: 10.12688/f1000research.11917.2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Azoulay P, Graff Zivin JS, Manso G. National Institutes of Health Peer Review: Challenges and Avenues for Reform. Innovation Policy and the Economy. 2013;13:1–22. [Google Scholar]
  • 7.Langfeldt L, Kyvik S. Researchers as evaluators: tasks, tensions and politics. Higher Education. 2011;62(2):199–212. [Google Scholar]
  • 8.Fang FC, Casadevall A. Research Funding: the Case for a Modified Lottery. mBio. 2016;7(2):e00422. doi: 10.1128/mBio.00422-16 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Jayasinghe UW, Marsh HW, Bond N. A multilevel cross-classified modelling approach to peer review of grant proposals: the effects of assessor and researcher attributes on assessor ratings. Journal of the Royal Statistical Society: Series A (Statistics in Society). 2003;166(3):279–300. [Google Scholar]
  • 10.Fogelholm M, Leppinen S, Auvinen A, Raitanen J, Nuutinen A, Väänänen K. Panel discussion does not improve reliability of peer review for medical research grant proposals. Journal of Clinical Epidemiology. 2012;65(1):47–52. doi: 10.1016/j.jclinepi.2011.05.001 [DOI] [PubMed] [Google Scholar]
  • 11.Tamblyn R, Girard N, Qian CJ, Hanley J. Assessment of potential bias in research grant peer review in Canada. Canadian Medical Association Journal. 2018;190(16):E489–E99. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Sattler DN, McKnight PE, Naney L, Mathis R. Grant Peer Review: Improving Inter-Rater Reliability with Training. PLOS ONE. 2015;10(6):e0130450. doi: 10.1371/journal.pone.0130450 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Reinhart M. Peer review of grant applications in biology and medicine. Reliability, fairness, and validity. Scientometrics. 2009;81(3):789–809. [Google Scholar]
  • 14.Pier EL, Brauer M, Filut A, Kaatz A, Raclaw J, Nathan MJ, et al. Low agreement among reviewers evaluating the same NIH grant applications. Proceedings of the National Academy of Sciences. 2018;115(12):2952–7. doi: 10.1073/pnas.1714379115 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Graves N, Barnett AG, Clarke P. Funding grant proposals for scientific research: retrospective analysis of scores by members of grant review panel. BMJ. 2011;343:d4797. doi: 10.1136/bmj.d4797 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Mayo NE, Brophy J, Goldberg MS, Klein MB, Miller S, Platt RW, et al. Peering at peer review revealed high degree of chance associated with funding of grant applications. Journal of Clinical Epidemiology. 2006;59(8):842–8. doi: 10.1016/j.jclinepi.2005.12.007 [DOI] [PubMed] [Google Scholar]
  • 17.Clarke P, Herbert D, Graves N, Barnett AG. A randomized trial of fellowships for early career researchers finds a high reliability in funding decisions. Journal of Clinical Epidemiology. 2016;69:147–51. doi: 10.1016/j.jclinepi.2015.04.010 [DOI] [PubMed] [Google Scholar]
  • 18.Marsh HW, Jayasinghe UW, Bond NW. Improving the peer-review process for grant applications: reliability, validity, bias, and generalizability. Am Psychol. 2008;63(3):160–8. doi: 10.1037/0003-066X.63.3.160 [DOI] [PubMed] [Google Scholar]
  • 19.Jerrim J, Vries R. Are peer reviews of grant proposals reliable? An analysis of Economic and Social Research Council (ESRC) funding applications. The Social Science Journal. 2023;60(1):91–109. [Google Scholar]
  • 20.Saaty TL. Rank from comparisons and from ratings in the analytic hierarchy/network processes. European Journal of Operational Research. 2006;168(2):557–70. [Google Scholar]
  • 21.Jayasinghe UW, Marsh HW, Bond N. A new reader trialapproach to peer review in funding research grants: An Australian experiment. Scientometrics. 2006;69(3):591–606. [Google Scholar]
  • 22.Jones N, Brun A, Boyer A, editors. Improving reliability of user preferences: Comparing instead of rating. 2011 Sixth International Conference on Digital Information Management; 2011. 26–28 Sept. 2011. [Google Scholar]
  • 23.Harzing A-W, Baldueza J, Barner-Rasmussen W, Barzantny C, Canabal A, Davila A, et al. Rating versus ranking: What is the best way to reduce response and language bias in cross-national research? International Business Review. 2009;18(4):417–32. [Google Scholar]
  • 24.Wennerås C, Wold A. Nepotism and sexism in peer-review. Nature. 1997;387(6631):341–3. doi: 10.1038/387341a0 [DOI] [PubMed] [Google Scholar]
  • 25.Witteman HO, Hendricks M, Straus S, Tannenbaum C. Are gender gaps due to evaluations of the applicant or the science? A natural experiment at a national funding agency. The Lancet. 2019;393(10171):531–40. doi: 10.1016/S0140-6736(18)32611-4 [DOI] [PubMed] [Google Scholar]
  • 26.Kaatz A, Lee YG, Potvien A, Magua W, Filut A, Bhattacharya A, et al. Analysis of National Institutes of Health R01 Application Critiques, Impact, and Criteria Scores: Does the Sex of the Principal Investigator Make a Difference? Acad Med. 2016;91(8):1080–8. doi: 10.1097/ACM.0000000000001272 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Eloy JA, Svider PF, Kovalerchik O, Baredes S, Kalyoussef E, Chandrasekhar SS. Gender differences in successful NIH grant funding in otolaryngology. Otolaryngol Head Neck Surg. 2013;149(1):77–83. doi: 10.1177/0194599813486083 [DOI] [PubMed] [Google Scholar]
  • 28.Erosheva Elena A, Grant S, Chen M-C, Lindner Mark D, Nakamura Richard K, Lee Carole J. NIH peer review: Criterion scores completely account for racial disparities in overall impact scores. Science Advances.6(23):eaaz4868. doi: 10.1126/sciadv.aaz4868 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Ginther DK, Schaffer WT, Schnell J, Masimore B, Liu F, Haak LL, et al. Race, ethnicity, and NIH research awards. Science. 2011;333(6045):1015–9. doi: 10.1126/science.1196783 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Foundation Grant: Overview: Canadian Institutes of Health Research; [Available from: https://cihr-irsc.gc.ca/e/49798.html.
  • 31.College of Reviewers—Membership List: Canadian Institutes of Health Research; [Available from: https://cihr-irsc.gc.ca/e/51148.html.
  • 32.Hirsch JE. An index to quantify an individual’s scientific research output. Proceedings of the National Academy of Sciences. 2005;102(46):16569–72. doi: 10.1073/pnas.0507655102 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Vinkler P. Impact of the number and rank of coauthors on h-index and π-index. The part-impact method. Scientometrics. 2023;128(4):2349–69. [Google Scholar]
  • 34.Sandström U, Hällsten M. Persistent nepotism in peer-review. Scientometrics. 2008;74(2):175–89. [Google Scholar]
  • 35.Abdoul H, Perrey C, Tubach F, Amiel P, Durand-Zaleski I, Alberti C. Non-Financial Conflicts of Interest in Academic Grant Evaluation: A Qualitative Study of Multiple Stakeholders in France. PLOS ONE. 2012;7(4):e35247. doi: 10.1371/journal.pone.0035247 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Jayasinghe UW, Marsh HW, Bond N. Peer Review in the Funding of Research in Higher Education: The Australian Experience. Educational Evaluation and Policy Analysis. 2001;23(4):343–64. [Google Scholar]
  • 37.Bornmann L, Mutz R, Daniel H-D. Gender differences in grant peer review: A meta-analysis. Journal of Informetrics. 2007;1(3):226–38. [Google Scholar]
  • 38.Bol T, de Vaan M, van de Rijt A. The Matthew effect in science funding. Proceedings of the National Academy of Sciences. 2018;115(19):4887–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Strengthening Canda’s Research Capacity: The Gender Dimension Ottawa, Canada: Council of Canadian Academies (CCA); 2012. [Google Scholar]
  • 40.Bilimoria D, Liang X. Gender Equity in Science and Engineering: Advancing Change in Higher Education. 1 ed: Routledge; 2011. [Google Scholar]
  • 41.Shannon G, Jansen M, Williams K, Cáceres C, Motta A, Odhiambo A, et al. Gender equality in science, medicine, and global health: where are we at and why does it matter? The Lancet. 2019;393(10171):560–9. [DOI] [PubMed] [Google Scholar]
  • 42.Coe IR, Wiley R, Bekker L-G. Organisational best practices towards gender equality in science and medicine. The Lancet. 2019;393(10171):587–93. doi: 10.1016/S0140-6736(18)33188-X [DOI] [PubMed] [Google Scholar]
  • 43.Swenor BK, Munoz B, Meeks LM. A decade of decline: Grant funding for researchers with disabilities 2008 to 2018. PLOS ONE. 2020;15(3):e0228686. doi: 10.1371/journal.pone.0228686 [DOI] [PMC free article] [PubMed] [Google Scholar]

Decision Letter 0

Luís A Nunes Amaral

1 Nov 2022

PONE-D-22-19788Ranking versus Rating in Peer Review of Research Grant ApplicationsPLOS ONE

Dear Dr. Tamblyn,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. The reviewer raises concerns about the soundness of the statistical analysis that need to be addressed.

Please submit your revised manuscript by Dec 16 2022 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Luís A. Nunes Amaral, Ph.D.

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at 

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and 

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please change "female” or "male" to "woman” or "man" as appropriate, when used as a noun (see for instance https://apastyle.apa.org/style-grammar-guidelines/bias-free-language/gender).

3. In your Data Availability statement, you have not specified where the minimal data set underlying the results described in your manuscript can be found. PLOS defines a study's minimal data set as the underlying data used to reach the conclusions drawn in the manuscript and any additional data required to replicate the reported study findings in their entirety. All PLOS journals require that the minimal data set be made fully available. For more information about our data policy, please see http://journals.plos.org/plosone/s/data-availability.

"Upon re-submitting your revised manuscript, please upload your study’s minimal underlying data set as either Supporting Information files or to a stable, public repository and include the relevant URLs, DOIs, or accession numbers within your revised cover letter. For a list of acceptable repositories, please see http://journals.plos.org/plosone/s/data-availability#loc-recommended-repositories. Any potentially identifying patient information must be fully anonymized.

Important: If there are ethical or legal restrictions to sharing your data publicly, please explain these restrictions in detail. Please see our guidelines for more information on what we consider unacceptable restrictions to publicly sharing data: http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions. Note that it is not acceptable for the authors to be the sole named individuals responsible for ensuring data access.

We will update your Data Availability statement to reflect the information you provide in your cover letter.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Partly

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: No

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: No

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: No

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: In this manuscript, Tamblyn et al. analyzed an interesting dataset about grant peer-review in Canada. The main purpose of the analysis is to compare the reliability and bias of peer-review under two different scenarios: ranking versus rating. The main conclusion is that ranking seems to increase reliability while reducing bias. Furthermore, the authors also found some factors correlated with variance and outcome of the review.

My comments are listed below:

Comments about conceptual issues:

1. Perhaps the most convincing result in the paper is that ranking increases the reliability of the review compared to raw scores (Table 3). However, the mechanism articulated by the authors is a bit vague to me. The authors claimed that the benefit of ranking stems from removing the requirement for a stable internalized standard, which is difficult for inexperienced reviewers compared to experienced reviewers. This seems an overly complicated explanation. A more straightforward explanation would be that ranking controls for leniency (1) variation among evaluators, so the evaluation is more robust. Either way, I recommend a more thorough discussion of potential mechanisms and perhaps lay out future tests that can distinguish between these hypotheses.

2. Compared to the results about reliability, the results about reducing bias (table 5) are much less convincing. The main reason is that the quality proxy for proposals is H-index. While this might be partially justified in the first stage evaluation (since the first stage appears to be evaluating people, not the grant proposal), it is important to acknowledge that H-index can be a signal of prestige instead of researcher quality, so the results in table 5 might as well reflect that ranking and rating weighted different kinds of bias differently;

3. Given that the authors found relatively convincing evidence of ranking reducing reliability versus reducing bias, it might be beneficial to clearly contrast these two concepts in the introduction, as many readers can lump the two concepts together and believe they are concordant measures of peer-review quality. Sometimes, larger variance in evaluation can be good for better evaluation outcomes since pooling different opinions can result in the ‘wisdom of the crowd (2,3)’;

Comments about statistics:

1. I have a minor concern about ICC measures for reliability (I might be wrong since I am not an expert on this measure). If I remember correctly, the ICC requires some normality assumption (4). This is likely true for raw rating score but might not be valid for percentile ranking, as one would expect the percentile ranking would follow a uniform distribution between 0-1. Or maybe it will be approximately normal after taking the mean?

2. If I understand correctly, the regressions in Tables 4 and 5 are also used to search for factors that are significantly associated with variance (another way of measuring reliability) and outcome in different stages and schemes. If this is the case, the authors are effectively conducting multiple testing. In this case, a good practice is to perform multiple testing corrections, even for multiple regression (5). At the very least, authors can display the p-values with higher precision so that interested readers can gauge the results when needed. Along the same lines, if figure 1 shows the only significant interaction terms, then the (adjusted) p-value of the interaction term should be demonstrated, as figure 1 is what remains among the several interaction terms tested as described in the main text. Besides, to show the non-linear effect in figure 1, confidence intervals should be plotted;

Comment about presentation:

1. Finally, the organization of the paper should be substantially improved for readability. For example, a schematic figure showing the reviewing process of the CIHR program would be helpful. Furthermore, the formula for multiple regressions conducted should be explicitly shown instead of relying on a verbal description. In addition, the results would be much easier to read if divided into different sections (with titles) according to the conclusions.

Overall, I recommend minor revisions before further consideration.

References:

(1) Sampat, B., & Williams, H. L. (2019). How do patents affect follow-on innovation? Evidence from the human genome. American Economic Review, 109(1), 203-36.

(2) Shi, F., Teplitskiy, M., Duede, E., & Evans, J. A. (2019). The wisdom of polarized crowds. Nature human behaviour, 3(4), 329-336.

(3) Sun, M., Barry Danfa, J., & Teplitskiy, M. (2022). Does double-blind peer review reduce bias? Evidence from a top computer science conference. Journal of the Association for Information Science and Technology, 73(6), 811-819.

(4) Nakagawa, S., & Schielzeth, H. (2010). Repeatability for Gaussian and non-Gaussian data: a practical guide for biologists. Biological Reviews, 85(4), 935-956.

(5) Perrett, D. J. M. J. J., Schaffer, J., Piccone, A., & Roozeboom, M. (2006). Bonferroni adjustments in tests for regression coefficients. Multiple Linear Regression Viewpoints, 32(1), 1-6.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2023 Oct 5;18(10):e0292306. doi: 10.1371/journal.pone.0292306.r002

Author response to Decision Letter 0


5 Dec 2022

We have included a detailed response to the reviewer's comments in the attached files.

We have also amended the data availability statement and have modified the formatting of the manuscript (final version only, not tracked version) to fit the style requirements of PLOS ONE.

Attachment

Submitted filename: Response to reviews 2022-11-09.docx

Decision Letter 1

Julian D Cortes

10 Apr 2023

PONE-D-22-19788R1Ranking versus rating in peer review of research grant applicationsPLOS ONE

Dear Dr. Tamblyn,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

See comments below. 

Please submit your revised manuscript by May 25 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Julian D. Cortes

Academic Editor

PLOS ONE

Additional Editor Comments:

Dear author/s, thanks for submitting your work to PLoS ONE,

I contrasted the assessment of two reviewers of your work. Considering that the overall assessment resulted in a major and a minor revision, the article still needs further adjustments, particularly in the literature review which will enrich the discussion of the findings, and the methodology/statistical method applied.

Contrasting reviewers’ assessment with PLoS ONE’s requisites for publication, the article should be strengthened in the following terms:

• Experiments, statistics, and other analyzes are performed to a high technical standard and are described in sufficient detail.

• Conclusions are presented appropriately and are supported by the data.

I hope you can incorporate the above suggestions to improve your already valuable work.

Sincerely,

Julián D. Cortés

Associate Editor

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

Reviewer #2: (No Response)

Reviewer #3: (No Response)

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Partly

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: No

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: No

Reviewer #2: No

Reviewer #3: No

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The authors have addressed all my comments. I have no further questions. The only suggestion is that the formula presentation could possibly be improved via LaTex. But I assume this can be done in the final proofreading process.

Reviewer #2: The authors suggest that the peer review process used in determining the funding priority of submitted proposals has been plagued by reports of poor reliability and that one potential solution may lie in examining other schemas of evaluation, including a ranking process, which may be more reliable and may be less prone to bias. To examine this, the authors compare the reliability associated with application rating with those of application ranking in 3,156 applications that have been submitted to the Canadian Institutes of Health Research funding agency. In general, the authors found that ranking was more reliable than

rating and “less susceptible to the characteristics of the review panel, such as level of

expertise and experience, for both reliability and potential sources of bias.” This natural experiment is important and the authors should be lauded for taking this analysis on, as access to these types of data are limited. In general, this is an important area, and improving upon current review processes is important not only to provide more consistent output but for the credibility of the scientific process itself.

First, the differences observed in the reliability measures for rating and ranking are an important finding. While it seems the rating and ranking are independent processes, they are also highly correlated. Is there an assumption here that reviewers use the same criteria for ranking and rating? If they are different, do the authors have any insight into the ways they diverge? It may be helpful to add a little language in the discussion addressing what the authors think might be the differences in the decision-making processes peer reviewers use for rating vs ranking?

Also, the reliabilities for rating and ranking are somewhat similar for stage 1 and better than that of stage 2 (possibly because reviewers in stage 1 are mostly discriminating between good versus bad proposals?). For stage 2, reliability for both rating and ranking is worse, but the difference between them has grown, suggesting ranking is more clearly advantageous over rating for stage 2 (perhaps because judgements at this stage are more difficult; i.e. it is more difficult to discriminate between good and great applications via a rating mechanism?). Do the authors feel the utility of ranking over rating improves for proposals that are in the good/great zone? This is important as this may be a range where peer review is the least reliable or at least has the most difficulty in discriminating between proposals and may be most prone to subjectivity and bias.

Another area that might be undersold here a bit is the inherent power of ranking for tie-breaking. Could the authors elaborate on this in the analysis (e.g. how many rating ties are broken by the ranking process?).

While a good portion of the analysis is spent on bias, it is not entirely clear that ranking is devoid of bias or if it is less susceptible to bias, why that is in the context of peer review? Maybe the authors might add a little language in the discussion to address this.

A minor concern, the authors state that “Ranking would also provide greater flexibility in recruiting reviewers with different levels of experience and expertise without compromising reliability.” But is the goal to reduce expertise levels of reviewers?

Also, the authors mention the online discussion increased ranking variance “With respect to the review process, we found that on-line discussion of an application was associated with higher rater variance, which may be expected, as efforts would be selectively made to resolve differences in review opinion among reviewers with very divergent scores” Maybe the authors could provide more description about how proposals were selected for online discussion (I’m presuming there was a cut-off for discrepant scores?). If one looked at scores before and after discussion, would you would see a resolving of scores, i.e. scores coming together (reduced variance)?

Reviewer #3: The authors present the work titled: "Ranking versus rating in peer review of research grant applications." Though the topic is of great relevance, it has several flaws that authors should address to enhance the quality of the manuscript.

First, the literature review omits several relevant aspects that the reader should know to understand the full scope and implications of the research. Note, however, that the literature on grant funding evaluations is enormous. A potential contribution in this regard relates to a comprehensive literature review. Unfortunately, the literature review provided by the authors is weak. The authors could have included several other factors to illustrate or explain their empirical strategy. For example, the Matthew effect (Bol et al, 2018), gender effects (Eloy et al, 2013), the use of multi-criteria approaches (Oztaysi et al, 2017), the effects per discipline (Jerrim & Vries, 2020); implementation considerations (Neta et al 2015), or previous recommendations for reliability and validity (Marsh et al 2008) are just a few points of an extensive list of missed points.

Another concern is the statistical treatment. In particular, the authors have missed potential problems associated with endogeneity in the data. Endogeneity broadly refers to situations in which an explanatory variable (e.g., the H-Index of the applicant) correlates with the error term of the regression equation (tables 4 and 5). In this case, if authors do not control for how many co-authors an applicant has worked with in previous works, this might increase the error term of the regression equation because H-index correlates with the number of co-authors (Vinkler, 2023).

Due to the issues described above, the manuscript can not be accepted in its current form and the authors should address them before an additional consideration for the study.

SUGGESTED REFERENCES

Bol, T., de Vaan, M., & van de Rijt, A. (2018). The Matthew effect in science funding. Proceedings of the National Academy of Sciences, 115(19), 4887-4890.

Eloy, J. A., Svider, P. F., Kovalerchik, O., Baredes, S., Kalyoussef, E., & Chandrasekhar, S. S. (2013). Gender differences in successful NIH grant funding in otolaryngology. Otolaryngology--Head and Neck Surgery, 149(1), 77-83.

Jerrim, J., & Vries, R. D. (2020). Are peer-reviews of grant proposals reliable? An analysis of Economic and Social Research Council (ESRC) funding applications. The Social Science Journal, 1-19.

Marsh, H. W., Jayasinghe, U. W., & Bond, N. W. (2008). Improving the peer-review process for grant applications: reliability, validity, bias, and generalizability. American psychologist, 63(3), 160.

Neta, G., Sanchez, M. A., Chambers, D. A., Phillips, S. M., Leyva, B., Cynkin, L., ... & Vinson, C. (2015). Implementation science in cancer prevention and control: a decade of grant funding by the National Cancer Institute and future directions. Implementation Science, 10, 1-10.

Oztaysi, B., Onar, S. C., Goztepe, K., & Kahraman, C. (2017). Evaluation of research proposals for grant funding using interval-valued intuitionistic fuzzy sets. Soft Computing, 21, 1203-1218.

Vinkler, P. (2023). Impact of the number and rank of coauthors on h-index and π-index. The part-impact method. Scientometrics, 1-21.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

Reviewer #3: Yes: Juan Carlos Correa Nuñez

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2023 Oct 5;18(10):e0292306. doi: 10.1371/journal.pone.0292306.r004

Author response to Decision Letter 1


8 May 2023

Responses below have also been appended as an attached Word document.

Reviewer #1: The authors have addressed all my comments. I have no further questions. The only suggestion is that the formula presentation could possibly be improved via LaTex. But I assume this can be done in the final proofreading process.

--> The equations in the table footnotes have been re-formatted in LaTex

Reviewer #2: The authors suggest that the peer review process used in determining the funding priority of submitted proposals has been plagued by reports of poor reliability and that one potential solution may lie in examining other schemas of evaluation, including a ranking process, which may be more reliable and may be less prone to bias. To examine this, the authors compare the reliability associated with application rating with those of application ranking in 3,156 applications that have been submitted to the Canadian Institutes of Health Research funding agency. In general, the authors found that ranking was more reliable than rating and “less susceptible to the characteristics of the review panel, such as level of expertise and experience, for both reliability and potential sources of bias.” This natural experiment is important and the authors should be lauded for taking this analysis on, as access to these types of data are limited. In general, this is an important area, and improving upon current review processes is important not only to provide more consistent output but for the credibility of the scientific process itself.

--> Thank-you for your comments and support for this work.

First, the differences observed in the reliability measures for rating and ranking are an important finding. While it seems the rating and ranking are independent processes, they are also highly correlated. Is there an assumption here that reviewers use the same criteria for ranking and rating? If they are different, do the authors have any insight into the ways they diverge? It may be helpful to add a little language in the discussion addressing what the authors think might be the differences in the decision-making processes peer reviewers use for rating vs ranking?

--> The advantage of this natural experiment was that the criteria used to rate and rank applications were the same. While we have no information on the decision-making process of reviewers, we have hypothesized possible mechanisms that could be operating in the ranking process that would lead to improved reliability and reduction in bias to the discussion.

Also, the reliabilities for rating and ranking are somewhat similar for stage 1 and better than that of stage 2 (possibly because reviewers in stage 1 are mostly discriminating between good versus bad proposals?). For stage 2, reliability for both rating and ranking is worse, but the difference between them has grown, suggesting ranking is more clearly advantageous over rating for stage 2 (perhaps because judgements at this stage are more difficult; i.e. it is more difficult to discriminate between good and great applications via a rating mechanism?). Do the authors feel the utility of ranking over rating improves for proposals that are in the good/great zone? This is important as this may be a range where peer review is the least reliable or at least has the most difficulty in discriminating between proposals and may be most prone to subjectivity and bias.

Another area that might be undersold here a bit is the inherent power of ranking for tie-breaking. Could the authors elaborate on this in the analysis (e.g. how many rating ties are broken by the ranking process?).

--> This is an excellent point and we have added this information on tie breaking to the results reported relative to Table 3. As the reviewer suspects, tie breaking was a frequent occurrence. In phase 1, the average number of ties that needed to be broken by a reviewer were for 31.7% of their application reviews and only 16% of reviewers had no ties that needed to be broken. In phase 2,ties needed to eb broken in 28.3% of applications and by 75.4% of reviewers.

While a good portion of the analysis is spent on bias, it is not entirely clear that ranking is devoid of bias or if it is less susceptible to bias, why that is in the context of peer review? Maybe the authors might add a little language in the discussion to address this.

--> We agree. In our discussion we indicated that ranking may confer reductions in potential sources of bias, we clarified that this comment referred to our findings related to applicant gender and early career status where there was an appreciable difference in the performance of rating and ranking.

A minor concern, the authors state that “Ranking would also provide greater flexibility in recruiting reviewers with different levels of experience and expertise without compromising reliability.” But is the goal to reduce expertise levels of reviewers?

--> We realize this sentence is misleading and have revised it as follows..

“Ranking would also provide greater flexibility in recruiting reviewers from various disciplines with different levels of experience without compromising reliability”

Also, the authors mention the online discussion increased ranking variance “With respect to the review process, we found that on-line discussion of an application was associated with higher rater variance, which may be expected, as efforts would be selectively made to resolve differences in review opinion among reviewers with very divergent scores” Maybe the authors could provide more description about how proposals were selected for online discussion (I’m presuming there was a cut-off for discrepant scores?). If one looked at scores before and after discussion, would you would see a resolving of scores, i.e. scores coming together (reduced variance)?

--> In the Methods section, we have augmented the description of the peer review process to indicate how applicants were selected for on-line discussion. The reviewer has raised an important point about the potential consequences of discussion on rater variance. Unfortunately only the final rating and ranking were retained by the peer review system so this question cannot be addressed.

Reviewer #3: The authors present the work titled: "Ranking versus rating in peer review of research grant applications." Though the topic is of great relevance, it has several flaws that authors should address to enhance the quality of the manuscript.

First, the literature review omits several relevant aspects that the reader should know to understand the full scope and implications of the research. Note, however, that the literature on grant funding evaluations is enormous. A potential contribution in this regard relates to a comprehensive literature review. Unfortunately, the literature review provided by the authors is weak. The authors could have included several other factors to illustrate or explain their empirical strategy. For example, the Matthew effect (Bol et al, 2018), gender effects (Eloy et al, 2013), the use of multi-criteria approaches (Oztaysi et al, 2017), the effects per discipline (Jerrim & Vries, 2020); implementation considerations (Neta et al 2015), or previous recommendations for reliability and validity (Marsh et al 2008) are just a few points of an extensive list of missed points.

--> We thank the reviewer for identifying additional references that were not identified in our literature search, many of which were relevant to the research question addressed in this manuscript. The literature review outlined in the introduction was related to the specific question of reliability and bias in grant applications using two methods, rating and ranking. We cover both the theoretical literature on these forms of assessment, and the results of applications in other contexts as well as in grant review. A comprehensive literature review on all that is known about grant peer review was beyond the scope or intent of this manuscript. Two comprehensive reviews have already been published on this topic, both of which are cited in this manuscript (Guthrie, 2018; Shepherd, 2018).

Another concern is the statistical treatment. In particular, the authors have missed potential problems associated with endogeneity in the data. Endogeneity broadly refers to situations in which an explanatory variable (e.g., the H-Index of the applicant) correlates with the error term of the regression equation (tables 4 and 5). In this case, if authors do not control for how many co-authors an applicant has worked with in previous works, this might increase the error term of the regression equation because H-index correlates with the number of co-authors (Vinkler, 2023).

--> This is a very good point. We have redone the analysis to include the number of unique collaborators listed in the applicant’s publications to address this problem of unmeasured covariates in the error term. The correlation between the H index and number of collaborators is reasonably high (r=0.5). Both the H index and number of collaborators are significantly associated with rater variance and the score in the same direction. However, there is no impact of including number of collaborators on estimates of applicant, reviewer and review characteristics on variance and bias. The new results are presented in Table 4 and 5.

Attachment

Submitted filename: Response to Reviewers-PLOS-2023-05-04.docx

Decision Letter 2

Julian D Cortes

26 Jul 2023

PONE-D-22-19788R2Ranking versus rating in peer review of research grant applicationsPLOS ONE

Dear Dr. Tamblyn,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

==============================

Dear authors,

I’ve been handling your manuscript over the last four months after receiving an editor transfer request. I’m aware of the more than a year reviewing process and, therefore, admire your commitment during this year of revisions and adjustment to your article, following reviewers’ minor and major suggestions.

The first round of revisions received minor and overall favorable concepts. However, in the process of seeking new reviewers to check upon those revisions, before the last round the article received a “major revision,” that you already completed, and in the last round, unfortunately, the article received a “reject” concept.

To assume a nuanced position not only based on the most recent reviews but also those of the first versions of the manuscript managed by the previous editor, I will consider that you could address the concerns of the reviewer below. Mind that there are multiple methodological considerations that might require running analyses and interpret and discuss results with the literature suggested by the reviewer.

I look forward to your revised manuscript. 

==============================

Please submit your revised manuscript by Sep 09 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Julian D. Cortes

Academic Editor

PLOS ONE

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #4: (No Response)

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #4: No

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #4: No

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #4: No

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #4: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #4: Based on the title and abstract, this paper appears to be aimed at studying ranking and rating in grant peer review. The authors use a dataset from the Canadian Institutes of Health Research to compare their ranking and rating procedures for grant proposal peer review quality assessement. The main conclusion of the paper is that funding agencies should consider adopting ranking methods to improve the quality of funding decisions in health research.

The authors consider an important problem and attempt to study it using a large grant peer review data set from a reputable funding agency. The question of comparative reliability in ranking and rating is interesting. However, the actual quality assessment procedure that is being studied relative to standard rating is not ranking; it could be more accurately described as “tie-breaking score-based rank percentile” procedure. Therefore, the manuscript’s title and abstract claims are misleading.

Is the manuscript technically sound, and do the data support the conclusions?

There are substantial technical concerns with the manuscript that do not allow one to determine if the data support the conclusions about the use of ranking in grant peer review. The most important concerns include:

1. Although the title and abstract state that reliability of ratings and rankings are compared in the manuscript, that is not the case. Instead, the authors compare reliability of ratings and rank percentiles, obtained after reviewers were asked to break ties in their ratings. Rank percentiles are substantially different from rankings and have distinct properties. Many desirable properties of rankings, such as the direct comparison of proposals without reference to a scale, are lost when converting ranks to rank percentiles.

2. Relatedly, some conclusions noted by the authors are the direct mathematical result of the construction of their “ranking” data. Such conclusions include: (1) the authors note on line 240 that the mean rank percentile is 50%, which is a consequence of their data conversion that would hold for any data source. (2) Similarly, the lack of statistical significance for coefficients in the rank percentile model is likely the result of the “zero-sum” construction of rank percentiles, and not necessarily a true lack of statistical significance.

3. Although authors note that “inclusion of only top scoring applications in phase 2 reduced the reliability of assessment”, they still make the conclusion on lines 356-360 that IRR was found to be smaller for the assessment of research project quality (phase 2) than that of researchers (phase 1). This comparison is flawed. That IRR decreased in phase 2 is primarily a mathematical result of removing the estimated “worst” proposals (see Erosheva, Martinkova, and Lee, 2021).

4. On lines 138-139, the authors note “it was assumed that the distribution of the quality of applications assigned to each reviewer would be equivalent.” What is precisely meant by the phrase “the distribution of the quality of applications…would be equivalent”? Whatever the formal description of this assumption might be, it appears to be quite stringent and unlikely to hold, especially if there was substantial variability in the number of proposals assessed by each reviewer. The authors report the mean numbers of 12.6 and 9.5 of proposals assessed by each reviewer in phases 1 and 2, respectively, but do not comment on the range. Given this implausible assumption, the authors should explore its influence on the results.

5. Comparing the reliability (variance) of ratings and rankings is ultimately challenging given the ordinal properties of rankings. The authors’ attempts to make this comparison, while commendable in goal, are not statistically appropriate.

6. On numerous instances, we believe the authors use the phrase “percent rank” when referring to “rank percentile”. These are different concepts.

7. Interpretation of regression coefficients should be done conditionally on all other variables being held constant.

References:

Erosheva, E. A., Martinková, P., & Lee, C. J. (2021). When zero may not be zero: A cautionary note on the use of inter‐rater reliability in evaluating grant peer review. Journal of the Royal Statistical Society: Series A (Statistics in Society)

Has the statistical analysis been performed appropriately and rigorously?

Please refer to the concerns raised in answering the previous question.

Have the authors made all data underlying the findings in their manuscript fully available?

The data will not be publicly available citing privacy concerns. We suggest the authors attempt to release a de-identified subset of the data publicly, as has been done with similar data from the Swiss NSF in Heyard et al. (2021), the National Institutes of Health in Erosheva et al. (2020), and the American Institute of Biological Sciences in Pearce and Erosheva (2022).

References:

Heyard, R., Ott, M., Salanti, G., & Egger, M. (2022). Rethinking the funding line at the Swiss national science foundation: Bayesian ranking and lottery. Statistics and Public Policy, 9(1), 110-121.

Erosheva, E. A., Grant, S., Chen, M. C., Lindner, M. D., Nakamura, R. K., & Lee, C. J. (2020). NIH peer review: Criterion scores completely account for racial disparities in overall impact scores. Science Advances, 6(23), eaaz4868.

Pearce, M., & Erosheva, E. A. (2022). A unified statistical learning model for rankings and scores with application to grant panel review. Journal of Machine Learning Research, 23(210), 1-33.

Is the manuscript presented in an intelligible fashion and written in standard English?

Yes.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #4: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2023 Oct 5;18(10):e0292306. doi: 10.1371/journal.pone.0292306.r006

Author response to Decision Letter 2


29 Aug 2023

Dear Editor,

Thank you for providing us with the opportunity to respond to comments from a new reviewer of our manuscript. If I am to understand correctly, our responses to the previous sets of reviewers were acceptable but the most recent reviewer has raised new issues not previously identified by the prior sets of reviewers.

Unfortunately, the critique provided by this reviewer was based on a series of assumptions that are incorrect. They include the following:

1. The reviewer has assumed that the unit of analysis is the reviewer (comment 1,2). This is incorrect. The unit of analysis is the application. We measured the characteristics of the application and its score and percentile rank based on the mean of ratings and percentile ranks provided by all reviewers of the application.

2. The reviewer has assumed that there is a range restriction in scores in phase 2, as was the case in the Erosheva paper that was cited by the reviewer (comment 3). This is incorrect. The applications in phase 2 were scored again, using a different set of reviewers and criteria to assess the quality of the research program using the the full score range for their assessment. Our analysis is based on ALL applications, both those that were funded and those that were unsuccessful in both Phase 1 and Phase 2.

3. The reviewer assumed that the data were ordinal, likely because of confusion about the unit of analysis (comment 5). This is incorrect. The data are continuous as evidenced by the histogram and rug plot shown in response to this comment. Each application received a rating allowing one decimal place of a scale of 0-112 for phase 1 and 0-140 for phase 2 and a ranking from each of the 5 reviewers which was converted by CIHR to a percentile ranking to adjust for differences in the number of applications rated by each reviewer. The arithmetic mean of these values became the application score and percentile rank.

4. The reviewer assumed that only bivariate analysis was conducted (comment 7). This is incorrect. In the analysis section we indicated that multivariate regression was used to estimate the associations between application, applicant and reviewer characteristics and each outcome, and provided the regression models in the footnotes of each table. By definition, multivariate models provide estimates of the independent association of a given variable, controlling for all other variables in the model.

Our possibly unclear description of the methodology may have led, in part, to this series of incorrect assumptions. Therefore, we have revised the methods section of the paper to improve clarity and avoid these misunderstandings by future readers.

Our detailed responses to each comment follow.

The authors consider an important problem and attempt to study it using a large grant peer review data set from a reputable funding agency. The question of comparative reliability in ranking and rating is interesting. However, the actual quality assessment procedure that is being studied relative to standard rating is not ranking; it could be more accurately described as “tie-breaking score-based rank percentile” procedure. Therefore, the manuscript’s title and abstract claims are misleading.

Is the manuscript technically sound, and do the data support the conclusions?

There are substantial technical concerns with the manuscript that do not allow one to determine if the data support the conclusions about the use of ranking in grant peer review. The most important concerns include:

1. Although the title and abstract state that reliability of ratings and rankings are compared in the manuscript, that is not the case. Instead, the authors compare reliability of ratings and rank percentiles, obtained after reviewers were asked to break ties in their ratings. Rank percentiles are substantially different from rankings and have distinct properties. Many desirable properties of rankings, such as the direct comparison of proposals without reference to a scale, are lost when converting ranks to rank percentiles.

We have clarified our description of the process of ranking in the methodology and provided the following example in an on-line appendix. We are unsure what the reviewer means by the comment “Many desirable properties of rankings, such as the direct comparison of proposals without reference to a scale, are lost when converting ranks to rank percentiles”. The grant peer reviewers never worked with percentile ranks, only with the rank order of applications. They were asked to order all applications from best to worst, breaking ties as appropriate. The ranking process had no scale. We provide examples to illustrate how the grant peer reviewers changed the rank order of the applications and how they broke ties.

It was only after the grant peer reviewer completed the ranking exercise that CIHR converted the ranks to percentile ranks. This conversion was necessary because reviewers of the same application may have reviewed a different number of applications. The mean of the percentile ranks provided by each reviewer of the application was the final application score.

Appendix A provides three de-identified examples from the data set used in the analysis (please see attached reviewer response for clearer formatting).

In Example 1, Reviewer A reviewed 14 applications and scored them from 18 to 106 for the highest score application. The reviewer was presented with the initial ranking based on score. Two applications were tied with a score of 70. The reviewer broke this tie providing application #6 with an adjusted rank of 8 and application #7 with a rank of 9. Reviewer A also changed the rank order of application #13 and #14, from a rank of first to application #14 to second and application #13 was ranked first. CIHR converted the adjusted ranks to rank percentiles. If reviewer was the unit of analysis the mean rank percentile, by definition, would be 50% (as illustrated). However, in our study application was the unit of analysis. To obtain an overall score the application the mean of the percentile ranks of each reviewer was calculated. In example 1, for application #6, there were 5 reviewers with initial scores ranging from 70 to 96, and percentile ranks from 41.177 to 75.00 providing an overall mean percentile rank score of 54.549.

UNIT OF ANALYSIS IS THE REVIEWER

Example Scores and percentile ranks for REVIEWER A and the 14 applications he was assigned to:

Application number Reviewer pin Reviewers initial score Reviewer INITIAL Rank Reviewer ADJUSTED rank Reviewers percentile rank

1 A 18 13 14 0.000

2 A 48 12 13 7.692

3 A 54 11 12 15.385

4 A 60 10 11 23.077

5 A 62 9 10 30.769

6 A 70 8 8 46.154

7 A 70 8 9 38.462

8 A 74 7 7 53.846

9 A 80 6 6 61.539

10 A 84 5 5 69.231

11 A 88 4 4 76.923

12 A 92 3 3 84.615

13 A 98 2 1 100.000

14 A 106 1 2 92.308

Mean percentile rank 50.000

UNIT OF ANALYSIS IS THE APPLICATION

Scores and percentile ranks for APPLICATION no 6 from the 5 REVIEWERS assigned to review it & Mean percentile rank:

Application number Reviewer pin Reviewers initial score Reviewers percentile rank Mean percentile rank

6 a 70 46.154 54.549

6 b 80 41.177 54.549

6 c 84 66.667 54.549

6 d 84 43.750 54.549

6 e 96 75.000 54.549

Mean percentile rank 54.549

In Example #2, Reviewer Z reviewed 16 applications, scoring them from 64 to 106 with one tied score for application #15 and #16 of 76. The reviewer broke this tie ordering application #15 with a rank of 11 and #16 with a rank of 12. The lowest ranked application was #11 with an adjusted rank of 16 and a percentile rank of 0.00. Application #11 had 5 reviewers who scored the application between 48 and 84 with percentile ranks ranging from 0.00 to 62.5. The mean percentile rank score for this application was 28.56.

Example Scores and percentile ranks for REVIEWER Z and the 16 applications he was assigned to:

Application number Reviewer pin Reviewers initial score Reviewer INITIAL Rank Reviewer ADJUSTED rank Reviewers percentile rank

11 Z 64 15 16 0.000

12 Z 68 14 15 6.667

13 Z 70 13 14 13.333

14 Z 72 12 13 20.000

15 Z 76 11 11 33.333

16 Z 76 11 12 26.667

17 Z 86 10 10 40.000

18 Z 88 9 9 46.667

19 Z 90 8 8 53.333

20 Z 92 7 7 60.000

21 Z 96 6 6 66.667

22 Z 98 5 5 73.333

23 Z 100 4 4 80.000

24 Z 102 3 3 86.667

25 Z 104 2 2 93.333

26 Z 106 1 1 100.000

Mean percentile rank 50.000

Scores and percentile ranks for APPLICATION no 11 from the 5 REVIEWERS assigned to review it & Mean percentile rank:

Application number Reviewer pin Reviewers initial score Reviewers percentile rank Mean percentile rank

11 Z 64 0.000 28.560

11 P 62 13.333 28.560

11 Q 84 62.500 28.560

11 T 70 31.250 28.560

11 Y 48 35.714 28.560

Mean percentile rank 28.559

In Example #3, Reviewer T reviewed 13 applications, scoring them from 78 to 112. The reviewer broke two ties involving application #36 and #37 as well as #38 and #39. In addition, the reviewer changed the rank order of the top three applications, moving the third ranked to the second. The top ranked application for Reviewer T (application #43) had 5 reviewers, four of whom ranked the application as first (rank percentile=100.00) in the pool of applications they reviewed, providing an overall mean percentile rank from the 5 reviewers of 87.50 for application #43.

Example Scores and percentile ranks for REVIEWER T and the 13 applications he was assigned to:

Application number Reviewer pin Reviewers initial score Reviewer INITIAL Rank Reviewer ADJUSTED rank Reviewers percentile rank

31 T 78 11 13 0.000

32 T 84 10 12 8.333

33 T 88 9 11 16.667

34 T 96 8 10 25.000

35 T 98 7 9 33.333

36 T 100 6 7 66.667

37 T 100 6 8 58.333

38 T 102 5 5 50.000

39 T 102 5 6 41.667

40 T 104 4 4 75.000

41 T 106 3 2 91.667

42 T 108 2 3 83.333

43 T 112 1 1 100.000

Mean percentile rank 50.000

Scores and percentile ranks for APPLICATION no 43 from the 5 REVIEWERS assigned to review it & Mean percentile rank:

Application number Reviewer pin Reviewers initial score Reviewers percentile rank Mean percentile rank

43 T 112 100.000 97.500

43 EE 106 87.500 97.500

43 Q 100 100.000 97.500

43 A 112 100.000 97.500

43 V 104 100.000 97.500

Mean percentile rank 97.500

2. Relatedly, some conclusions noted by the authors are the direct mathematical result of the construction of their “ranking” data. Such conclusions include: (1) the authors note on line 240 that the mean rank percentile is 50%, which is a consequence of their data conversion that would hold for any data source.

The mean rank percentile is, by construction, is 50% when the unit of analysis is the reviewer as shown in our three examples Our study used the application as the unit of analysis. The final score for the application was the mean of the percentile ranks provided by all reviewers of the application. As illustrated in the three examples, the mean percentile rank of the application is, of course, not 50%.

We did not conduct any conversion of the data. The rating, ranking, adjusted ranking, and percentile ranking of each reviewer for an application, and the final mean ranking for each application is provided by CIHR. We have all ranks and percentile ranks for all reviewers for all applications for each stage of the competition. No grants were excluded from the analysis.

(2) Similarly, the lack of statistical significance for coefficients in the rank percentile model is likely the result of the “zero-sum” construction of rank percentiles, and not necessarily a true lack of statistical significance.

The zero-sum construction of rank percentiles would only apply if we were using reviewer as the unit of analysis (see above examples). We wish to point out that the zero-sum construction does not apply when using application as the unit of analysis where the outcome is the mean of all reviewers’ percentile ranks.

3. Although authors note that “inclusion of only top scoring applications in phase 2 reduced the reliability of assessment”, they still make the conclusion on lines 356-360 that IRR was found to be smaller for the assessment of research project quality (phase 2) than that of researchers (phase 1). This comparison is flawed. That IRR decreased in phase 2 is primarily a mathematical result of removing the estimated “worst” proposals (see Erosheva, Martinkova, and Lee, 2021).

We apologize for the lack of clarity in the presentation of results. We think the reviewer has assumed that the same phase 1 score was used to calculate the IRR for applications that were successful in reaching phase 2. This is not correct. In phase 2, a different set of reviewers was selected and they used a different set of criteria and scoring system to rate the quality of the research program, using the full range of possible scores. We have revised this paragraph describing the results to simply report that the ICC was lower for phase 2. In the discussion we have outlined likely reasons; 1) that the quality of candidates selected to submit their research program to phase 2 was more homogeneous, making it more difficult for reviewers to distinguish amongst higher quality proposals, and/or 2) that there is less agreement among reviewers about what constitutes a high quality research program than there is for evaluating the quality of the candidate (that was evaluated in phase 1).

4. On lines 138-139, the authors note “it was assumed that the distribution of the quality of applications assigned to each reviewer would be equivalent.” What is precisely meant by the phrase “the distribution of the quality of applications…would be equivalent”? Whatever the formal description of this assumption might be, it appears to be quite stringent and unlikely to hold, especially if there was substantial variability in the number of proposals assessed by each reviewer. The authors report the mean numbers of 12.6 and 9.5 of proposals assessed by each reviewer in phases 1 and 2, respectively, but do not comment on the range. Given this implausible assumption, the authors should explore its influence on the results.

In the “ideal” world, all reviewers would review each application submitted to a given program/committee, and the average rank for these applications would be assigned to the application. However this is a completely impractical. Nor is it possible to randomly allocate applications to reviewers whereby the random assignment would ensure that an equivalent number of poor, good and excellent applications would be reviewed by each reviewer. The priority in CIHR’s assignment of applications to reviewers was to optimize the match between reviewer expertise and the content of the application. A priori, neither CIHR or the chair of a committee would have knowledge of the quality of the application, therefore there was no possibility of systematic bias in assignment (for example where only the poor applications would go to one reviewer and the excellent applications to another). Therefore, it is assumed that, on balance, reviewers would obtain a mix of applications of varying quality. Any violation of this assumption would contribute to random errors in measurement and lead to an under-estimate of the reliability of ranking. We have noted this limitation in the discussion. The range in the number of applications assigned to a reviewer has been added to the description of Table 1 results.

5. Comparing the reliability (variance) of ratings and rankings is ultimately challenging given the ordinal properties of rankings. The authors’ attempts to make this comparison, while commendable in goal, are not statistically appropriate.

We include the histogram and rug plots of the log of the inter-rater variance for individual application scores and ranks for analyses reported in Table 3, to clarify that the data are normally distributed and continuous. Confusion about the nature of the data may have occurred if the reviewer assumed that the reviewer was the unit of analyses in the study. In our study, application is the unit of analysis. Each application has four to five reviewers, each of whom scored the application.

Distribution of the Variance of Ranks between Raters of an Application (see attached reviewer response for graph)

Distribution of the Variance of Scores between Raters of an Application (see attached reviewer response for graph)

6. On numerous instances, we believe the authors use the phrase “percent rank” when referring to “rank percentile”. These are different concepts.

This was our error, and we have corrected our terminology throughout the manuscript.

7. Interpretation of regression coefficients should be done conditionally on all other variables being held constant.

In the analysis section we indicated that multivariate regression was used to estimate the associations between application, applicant and reviewer characteristics and each outcome, and provided the regression models in the footnotes of each table. By definition, multivariate models provide estimates of the independent association of a given variable, controlling for all other variables in the model.

Is the manuscript presented in an intelligible fashion and written in standard English?

Yes.

References:

Erosheva, E. A., Martinková, P., & Lee, C. J. (2021). When zero may not be zero: A cautionary note on the use of inter‐rater reliability in evaluating grant peer review. Journal of the Royal Statistical Society: Series A (Statistics in Society)

Has the statistical analysis been performed appropriately and rigorously?

Please refer to the concerns raised in answering the previous question.

Have the authors made all data underlying the findings in their manuscript fully available?

The data will not be publicly available citing privacy concerns. We suggest the authors attempt to release a de-identified subset of the data publicly, as has been done with similar data from the Swiss NSF in Heyard et al. (2021), the National Institutes of Health in Erosheva et al. (2020), and the American Institute of Biological Sciences in Pearce and Erosheva (2022).

References:

Heyard, R., Ott, M., Salanti, G., & Egger, M. (2022). Rethinking the funding line at the Swiss national science foundation: Bayesian ranking and lottery. Statistics and Public Policy, 9(1), 110-121.

Erosheva, E. A., Grant, S., Chen, M. C., Lindner, M. D., Nakamura, R. K., & Lee, C. J. (2020). NIH peer review: Criterion scores completely account for racial disparities in overall impact scores. Science Advances, 6(23), eaaz4868.

Pearce, M., & Erosheva, E. A. (2022). A unified statistical learning model for rankings and scores with application to grant panel review. Journal of Machine Learning Research, 23(210), 1-33.

Attachment

Submitted filename: Editor response + review response.docx

Decision Letter 3

Julian D Cortes

6 Sep 2023

PONE-D-22-19788R3Ranking versus rating in peer review of research grant applicationsPLOS ONE

Dear Dr. Tamblyn,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

==============================

Dear author/s, thanks for submitting your work to PLoS ONE,

All major conceptual and methodological reviews of the article have been addressed. Still, please consider the following minor revisions:

- Add a space (•) before cited references. This applies to all the manuscript (e.g., from “investigators(4-8)” to “investigators•(4-8)”

- Authors mention in section “Applicant characteristics” that “For each publication, we retrieved the citation reports and assigned the impact factor for each journal by linking the ISSN of the journal to the Journal Citation Record file. When there was no recorded ISSN, we used the full and abbreviated journal name to make the link.” It is not clear why the authors sourced the impact factor of journals if the computation of the H-Index do not require such a value (H-Index=for a set of articles N of an author and defining ci as the number of citations corresponding to an article i then ordering the set of articles in decreasing order according to the number of citations).

In addition, although at this point it might be out of the article scope as it is, would be to include an H-Index based indicator that controls for an authors number of collaborators which is the hm index (For a set of articles N with ci the number of citations for the article i and ai the number of corresponding authors, the cumulative sum of the inverse of the number of authors is proposed as the effective rank . Then, sorting the set of articles in decreasing order according to the number of citations: https://doi.org/10.1016/j.joi.2008.05.001). The advantage is to compute an H index based indicator that controls for the number of collaborators throughout a researcher's career.

- PLoS ONE strongly advocates for a data availability and reproducible results as much as possible. Authors stated in “Data availability” item: “No - some restrictions will apply” considering the Canadian Institutes for Health Research (CIHR) policy. However, it might be available by request to the Vice-President of Research Programs-Operations at CIHR. Ultimately, I strongly suggest, if possible, to the authors to share in a public/institutional repository the code scripts produced for the analysis so other authors with the approval from the Canadian Institutes for Health Research (CIHR) could be able to reproduce/replicate/crowdsource the results of the study. Sincerely, Julian D. Cortés

Associated editor

==============================

Please submit your revised manuscript by Oct 21 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Julian D. Cortes

Academic Editor

PLOS ONE

Journal Requirements:

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

[Note: HTML markup is below. Please do not edit.]

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2023 Oct 5;18(10):e0292306. doi: 10.1371/journal.pone.0292306.r008

Author response to Decision Letter 3


14 Sep 2023

Dear author/s, thanks for submitting your work to PLoS ONE,

All major conceptual and methodological reviews of the article have been addressed. Still, please consider the following minor revisions:

- Add a space (•) before cited references. This applies to all the manuscript (e.g., from “investigators(4-8)” to “investigators•(4-8)”

We have corrected the spacing issue throughout the manuscript.

- Authors mention in section “Applicant characteristics” that “For each publication, we retrieved the citation reports and assigned the impact factor for each journal by linking the ISSN of the journal to the Journal Citation Record file. When there was no recorded ISSN, we used the full and abbreviated journal name to make the link.” It is not clear why the authors sourced the impact factor of journals if the computation of the H-Index do not require such a value (H-Index=for a set of articles N of an author and defining ci as the number of citations corresponding to an article i then ordering the set of articles in decreasing order according to the number of citations).

This was our oversight as we had another measure we were working on for other projects that required the journal impact factor. Thank-you so much for picking up this error. We have corrected it in our revised manuscript.

In addition, although at this point it might be out of the article scope as it is, would be to include an H-Index based indicator that controls for an authors number of collaborators which is the hm index (For a set of articles N with ci the number of citations for the article i and ai the number of corresponding authors, the cumulative sum of the inverse of the number of authors is proposed as the effective rank . Then, sorting the set of articles in decreasing order according to the number of citations: https://doi.org/10.1016/j.joi.2008.05.001). The advantage is to compute an H index based indicator that controls for the number of collaborators throughout a researcher's career.

Thank you for this suggestion for a minor revision. However, based on the comment from a previous peer review, we have decided to keep the number of collaborators as a separate variable that we adjust for in the models to transparently address the reviewer’s concerns about endogeneity.

Prior Reviewer: Another concern is the statistical treatment. In particular, the authors have missed potential problems associated with endogeneity in the data. Endogeneity broadly refers to situations in which an explanatory variable (e.g., the H-Index of the applicant) correlates with the error term of the regression equation (tables 4 and 5). In this case, if authors do not control for how many co-authors an applicant has worked with in previous works, this might increase the error term of the regression equation because H-index correlates with the number of co-authors (Vinkler, 2023).

Our response to the prior review: This is a very good point. We have redone the analysis to include the number of unique collaborators listed in the applicant’s publications to address this problem of unmeasured covariates in the error term. The correlation between the H index and number of collaborators is reasonably high (r=0.5). Both the H index and number of collaborators are significantly associated with rater variance and the score in the same direction. However, there is no impact of including number of collaborators on estimates of applicant, reviewer and review characteristics on variance and bias. The new results are presented in Table 4 and 5.

- PLoS ONE strongly advocates for a data availability and reproducible results as much as possible. Authors stated in “Data availability” item: “No - some restrictions will apply” considering the Canadian Institutes for Health Research (CIHR) policy. However, it might be available by request to the Vice-President of Research Programs-Operations at CIHR. Ultimately, I strongly suggest, if possible, to the authors to share in a public/institutional repository the code scripts produced for the analysis so other authors with the approval from the Canadian Institutes for Health Research (CIHR) could be able to reproduce/replicate/crowdsource the results of the study.

We have noted that the SAS code used to conduct the analyses will be available at McGill -Dataverse curated by the McGill University Library (McGill University Dataverse (borealisdata.ca) within the next 2 weeks

Sincerely,

Julian D. Cortés

Associated editor

Attachment

Submitted filename: Response to Reviews-september-2023.docx

Decision Letter 4

Julian D Cortes

18 Sep 2023

Ranking versus rating in peer review of research grant applications

PONE-D-22-19788R4

Dear Dr. Tamblyn,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Julian D. Cortes

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

I am satisfied with the revisions carried out based on earlier feedback. Therefore, I recommend your paper for acceptance, subject to the usual final formatting checks required by the editorial office.

Reviewers' comments:

Acceptance letter

Julian D Cortes

25 Sep 2023

PONE-D-22-19788R4

Ranking versus rating in peer review of research grant applications

Dear Dr. Tamblyn:

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department.

If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org.

If we can help with anything else, please email us at plosone@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Professor Julian D. Cortes

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    Attachment

    Submitted filename: Response to reviews 2022-11-09.docx

    Attachment

    Submitted filename: Response to Reviewers-PLOS-2023-05-04.docx

    Attachment

    Submitted filename: Editor response + review response.docx

    Attachment

    Submitted filename: Response to Reviews-september-2023.docx

    Data Availability Statement

    The datasets analyzed in this study are held by the Canadian Institutes for Health Research (CIHR) and are not publicly available due to privacy and legal restrictions. Researchers wishing to obtain access to these data need to contact the Vice-President of Research Programs-Operations at CIHR (christian.baron@cihr-irsc.gc.ca) to obtain approval to access de-identified data on foundation funding program applications submitted between 2014 and 2017.


    Articles from PLOS ONE are provided here courtesy of PLOS

    RESOURCES