Skip to main content
PLOS Biology logoLink to PLOS Biology
. 2026 Sep 8;24(9):e3003986. doi: 10.1371/journal.pbio.3003986

Discovery Stack Pilot demonstrates the feasibility and outcomes of a scientist-designed peer-review model that separates quality and impact

Maureen A McGargill 1, Beiyun C Liu 2, Michael S Kuhns 3, Daniel Mucida 4,5, Isabella Rauch 6, Lauren B Rodda 6, Meghan A Koch 7,8, Hugo Gonzalez Velozo 9,10,11, Ken Cadwell 12,13, Tanya S Freedman 14,15,16, Tiffany C Scharschmidt 17, Richard Sever 18,19, Jose Ordovas-Montanes 20,21,22,23,24, Sara Suliman 25, Andrew Oberst 8, Brooke Runnette 1, Matthew F Krummel 25,*
Editor: Stephen Curry26
PMCID: PMC13614658  PMID: 42709878

Abstract

Peer review serves as the cornerstone of scientific quality control. Yet, the current journal-centric system is hindered by long timelines, high publication costs, inconsistent review quality, systemic biases, and editorial gatekeeping. Notably, the system relies on misaligned measures of impact that are tethered to journal branding and conflate scientific rigor (Quality) with perceived significance (Impact). Here, we report findings from the Discovery Stack Pilot Study, which tested a scientist-designed, journal-independent peer review model. The Discovery Stack model integrates in-line reviewer comments to promote constructive feedback and separately evaluates scientific Quality and Impact using defined criteria. To examine feasibility and effectiveness, manuscripts were reviewed in parallel with traditional journal review. A total of 162 reviews were completed, and survey data from 86 participants were analyzed. The results showed that reviewers effectively evaluated Quality and Impact as separate dimensions, with Quality scores being more consistent across reviewers than Impact scores. Participants strongly supported the core elements of the Discovery Stack model and expressed enthusiasm for its broader adoption to enhance transparency, efficiency, and value in peer review. Future studies will explore integrating this model into a digital platform for reviewing and curating scientific discoveries to improve the production and dissemination of high-quality research.


The modern form of scientific publishing often incentivizes overstating results to maximize perceived impact at the expense of ensuring rigor. This study tests the Discovery Stack model, which seeks to realign peer review with its core purpose of upholding scientific rigor by separating the assessment of quality and impact.

Introduction

Manuscript publication is the primary way researchers share new discoveries. Peer review prior to publication remains the central mechanism for evaluating the scientific rigor and validity of these findings [1]. Researchers depend on this process to guide future studies, validate results, and uphold their professional reputation. Likewise, funding agencies, regulatory bodies, academic institutions, and industry stakeholders rely on peer-reviewed research, and the prestige of the journals in which it is published, to inform critical decisions on funding, policy, faculty promotion, media communication, product development, and patient care. In short, peer review underpins nearly every aspect of how scientific knowledge is generated, communicated, applied, and valued.

Despite its central role, the current journal-centric peer review and scientific publishing system often fails to meet the needs of scientists and society [2–6]. A key limitation is the lack of a clear distinction between scientific Quality and Impact. Quality refers to the rigor and reproducibility of the data supporting a study’s conclusions, whereas Impact reflects the extent to which the findings advance understanding. When these dimensions are blurred, high-impact or hyped findings can overshadow weak evidence, while rigorous but incremental work is undervalued. Moreover, perceived Impact is often inferred from journal prestige rather than the intrinsic merit of the research itself [7,8]. Despite longstanding concerns regarding its validity, journal prestige is typically approximated using the Journal Impact Factor (JIF), a proprietary metric calculated by Clarivate [9–12]. The JIF represents the mean number of citations received by articles published in a journal over a two-year period. However, citation distributions within journals are highly skewed, with a relatively small proportion of papers accounting for a large fraction of total citations. One analysis estimated that ~15% of articles generate 50% of citations, meaning that the journal average poorly reflects the influence of most individual publications [13]. Furthermore, review articles or a single highly cited paper can substantially skew the JIF [12]. Importantly, JIF may not be a reliable proxy for the methodological quality or reliability of individual publications. For example, higher JIFs are associated with higher retraction rates [11,14]. However, this pattern may partly reflect greater scrutiny and readership of articles published in highly visible journals, which increases the likelihood that errors are detected, rather than a difference in underlying research quality. Collectively, these dynamics distort how scientific contributions are valued, reinforcing journal reputation over scientific merit and undervaluing confirmatory studies that are essential for establishing confidence in foundational discoveries.

The traditional publishing process is also slow and inefficient. Manuscripts are considered by only one journal or journal family at a time, and each review cycle may take months. Across disciplines, the average time from submission to acceptance is ~6 months, but frequently extends well beyond a year [15,16].

The quality, bias, and transparency of peer review are additional concerns. Reviews vary widely in depth and rigor, and critical flaws are sometimes missed [2–5,16,17]. The lack of formal training and standardized guidelines contributes to this inconsistency [18]. Reviewer anonymity, while intended to promote objectivity, can also shield bias and hostility from accountability. Together, these challenges undermine the effectiveness of peer review as a mechanism for quality control and diminish its value to authors.

Compounding these challenges, researchers perform peer review labor without compensation, while also paying to publish and access scientific literature. This model is inequitable, frustrating, and increasingly unsustainable, as publication numbers continue to rise without a proportional increase in the number of available reviewers [15,19].

Although the growing prevalence of open-access preprint servers has improved accessibility to new findings [8,20], these platforms often lack meaningful peer review or metrics of rigor. Consequently, readers face a new challenge, information overload, without effective mechanisms to evaluate, search, or filter studies based on Quality or Impact, making it difficult to determine which studies to trust and prioritize.

To address these challenges and accelerate scientific progress, the peer review process must be strengthened and modernized. We posit that an improved system should emphasize a “peer-improvement” mindset, positioning reviewers as collaborators focused on strengthening scientific rigor and benefiting the scientific community, rather than functioning as journal consultants determining binary publication eligibility. Such a system should apply standardized metrics to assess a study’s Quality and Impact as separate dimensions [21]. Moreover, because Impact evolves over time through ongoing evaluation and influence on subsequent studies, this metric should remain dynamic and independent of journal branding.

The Discovery Stack Pilot Study tested the feasibility and effectiveness of a new peer review model built on these principles. The pilot evaluated the outcomes of separately assessing manuscript Quality and Impact, applying standardized metrics to independently measure and report both dimensions, and using an in-line commenting tool to facilitate constructive contextual feedback directly to authors and readers. The pilot also explored mechanisms for improving transparency, accountability, and timeliness, aiming to shorten the time from submission to dissemination. Findings from the Discovery Stack Pilot provide a foundation for further refinement and optimization of approaches to improve scientific publishing and peer review.

Results

Discovery Stack model

The Discovery Stack Pilot was designed to test a structured, multi-phase peer review process that independently evaluated scientific Quality from potential Impact using standardized assessment criteria. The process included three sequential phases: 1) Quality Review, 2) Author Response, and 3) Impact Review (Fig 1A). Each phase was built on the previous one, with Impact reviewers able to view compiled Quality review feedback and author responses. To ensure sufficient evaluations for each manuscript, reviewers who completed the Quality review phase also performed an independent assessment of the manuscript’s Impact using a separate evaluation form. Additional Impact-only reviewers were then recruited to increase the number and diversity of Impact assessments. Detailed procedures for each phase are provided in Materials and methods.

Fig 1. The Discovery Stack model and manuscript enrollment.

Fig 1

(A) The Discovery Stack model includes three sequential phases: 1) Quality Review, 2) Author Response, and 3) Impact Review. Image created in Canva. (B) Eighteen manuscripts were reviewed in the pilot. Manuscripts are grouped according to recruitment source and author participation, with the number of completed Quality and Impact reviews for each manuscript depicted. **Manuscript D118 was posted on SSRN rather than bioRxiv and was excluded from some analyses due to compatibility issues between SSRN and Hypothes.is. *D165 and D173 were included in the pilot without author participation. (C) Number of Quality and Impact reviews completed per manuscript. The data underlying these figures are available in S1 Table and https://doi.org/10.5281/zenodo.2004122.

Manuscripts were eligible for enrollment if they: (1) were posted to a preprint server such as bioRxiv, which enabled testing of in-line comments through the Hypothes.is platform; (2) had been submitted to a traditional journal, allowing for comparison with conventional peer review; and (3) were in the fields of immunology or cancer biology, enabling us to leverage the subject expertise of the scientific advisory board (SAB) and participating reviewers.

To identify eligible manuscripts, initial invitations were sent to 116 individuals who had previously registered to participate as authors or reviewers in the Discovery Stack Pilot, which resulted in five enrolled manuscripts. Outreach was then expanded to corresponding authors of bioRxiv preprints that met the above criteria. The most recent preprints were prioritized to maximize the likelihood that manuscripts had not yet been accepted for journal publication. Using these criteria, 118 bioRxiv authors were contacted, resulting in 11 more enrolled manuscripts. Additionally, two bioRxiv preprints for which no author response was received were included in the pilot, as reviewers with appropriately matched expertise had already been identified and agreed to participate.

In total, 18 manuscripts were enrolled (Fig 1B), generating 164 reviews: 50 Quality and 114 Impact reviews (Fig 1B, 1C). Each manuscript received at least two Quality and five Impact reviews, with an average of 2.8 Quality and 6.5 Impact reviews per manuscript. Because we hypothesized that Impact assessments would be inherently more subjective and variable, we aimed to recruit three Quality reviewers and six Impact reviewers per manuscript. Most manuscripts met or exceeded this goal, demonstrating the feasibility of enrolling manuscripts, recruiting reviewers, and completing both Quality and Impact assessments.

The Discovery Stack Pilot was designed as a feasibility study to evaluate the implementation of this model under real-world conditions. The primary objectives were to assess recruitment, workflow execution, and whether the model could generate actionable insights into key aspects of peer review, including the separation of Quality and Impact and participant perceptions of standardized metrics. As such, results should be interpreted in the context of a pilot study. In addition, because enrollment relied on voluntary participation from both previously enrolled Discovery Stack participants and authors of recent bioRxiv preprints, the study population may have been enriched for individuals more receptive to alternative peer review approaches. This, together with the relatively small cohort size, should be considered when interpreting perception-based outcomes. Larger, more systematically recruited cohorts will be required to confirm and extend these findings.

Reviewers successfully distinguished Quality from Impact

A central feature of the Discovery Stack model is the independent evaluation of scientific Quality and Impact using structured, standardized assessment criteria. To implement this framework, reviewers completed assessment forms evaluating defined attributes of scientific Quality and Impact. The Quality Assessment Form consisted of six short-answer questions and 13 Likert-scale items addressing experimental design, controls, statistical analysis, reproducibility, and whether the data supported the stated conclusions (Fig 2). The Impact Assessment Form included a short-answer question, a multiple-choice question identifying features contributing to a manuscript’s Impact, and five Likert-scale items evaluating transformative potential, generalizability, technological advancement, and mechanistic insight (Fig 3). To generate quantitative metrics for filtering and prioritizing manuscripts by Quality and Impact, composite scores were generated from reviewer responses. Each question was weighted according to its relative importance to either the Quality or Impact of a manuscript (S1A, S1B Fig). Weighted responses from each reviewer were averaged to yield a single composite Quality and a single composite Impact score per reviewer per manuscript. Composite scores across reviewers were then averaged to produce a final Quality and Impact score for each manuscript (S1 Table; see Materials and methods for full scoring details).

Fig 2. Quality Review Assessment Form.

Fig 2

The Quality Review Assessment Form included (A) six short-answer questions, followed by (B) a set of Likert-scale items designed to evaluate specific attributes regarding the rigor and reproducibility of the data. (C) Scores from all reviewers were compiled, graphed, and shared with the authors.

Fig 3. Impact review assessment form.

Fig 3

(A) The Impact Review Assessment Form included: a short-answer question asking whether specific strengths or weaknesses identified in the Quality Review influenced their perception of the manuscript’s Impact, a multiple-choice question allowing reviewers to select features they believed contributed to the manuscript’s Impact, and (B) five Likert-scale questions assessing key dimensions of the manuscript’s Impact. (C) Multiple choice and (D) Likert-scale responses were compiled, graphed, and shared with the authors.

To determine whether reviewers evaluated Quality and Impact as distinct dimensions, we examined the relationship between composite Quality and Impact across manuscripts. First, composite Quality and Impact scores for each manuscript were plotted according to the JIF of the journal to which each manuscript was published or submitted, which served as a proxy for the author’s perception of Impact (Fig 4A). This visualization showed several manuscripts rated high in Quality but low in Impact, suggesting that reviewers were able to uncouple their assessment of Impact from Quality. To formally test this, we quantified the relationship between these two dimensions. Because each manuscript received more Impact than Quality reviews, the primary analysis was restricted to reviewers who provided both Quality and Impact scores for the same manuscript to minimize differences driven by unequal sample sizes and reviewer variability. Pearson’s correlation showed a significant positive association between Quality and Impact (r = 0.73, p = 0.0006), and linear regression confirmed that Quality significantly predicted Impact (Fig 4B; β = 0.70, R2 = 0.53, p = 0.0006). As expected, poor-Quality manuscripts are unlikely to be considered impactful. However, residual analysis indicated substantial divergence. Ten of 18 manuscripts (56%) had Impact scores outside the 95% confidence interval (CI) (Fig 4B), with residuals ranging from −0.99 to +0.53 Impact points (Fig 4C). The standard deviation (SD) of the residuals (SD = 0.41) further demonstrated variation around the regression line, indicating that reviewers’ Impact evaluations frequently diverged from predictions based solely on Quality. Analyses including all reviewer scores yielded comparable results (r = 0.78, p = 0.0001; β = 0.68, R2 = 0.60, p = 0.0001).

Fig 4. Effective separation of Quality and Impact with strong support for standardized metrics.

Fig 4

(A) Composite Quality and Impact scores arranged by JIF of the journal the manuscript was published in. *Manuscripts that were not published in a journal at the time of the final analysis are denoted by an asterisk. (B) The relationship between composite Quality and Impact scores for each manuscript was tested using Pearson’s correlation and linear regression, restricting analysis to reviewers who completed both assessments. The shaded region shows the 95% confidence interval (CI) of the regression line. (C) Residuals from the linear regression model to assess variation in Impact not explained by Quality. Shaded region represents ±1 standard deviation (SD = 0.41). (D) Manuscripts were grouped into tertiles based on Quality scores and Impact scores were compared across tertiles using Kruskal–Wallis tests (H = 10.4, p = 0.0043), followed by Dunn’s multiple comparisons. (E, F) Pearson’s correlation was used to evaluate the relationship between Quality (E) or Impact (F) scores and the JIF of the journal to which each manuscript was submitted. The submission journal of D165 and D173 was unknown, excluding them from the analysis. (G–J) Survey data from Quality/Impact reviewers (n = 42), Impact-only reviewers (n = 33), and authors (n = 11) were tallied, and the percentage of each group selecting each response is shown. The data underlying these figures are available in S1 Table and https://doi.org/10.5281/zenodo.2004122.

To further assess the independence between Impact and Quality, manuscripts were grouped into tertiles by ranked composite Quality scores, and mean Impact scores were compared across groups. Mean Impact scores increased with higher Quality, but significant differences were observed only between the lowest and highest tertiles, with substantial overlap between adjacent groups (Fig 4D). The residual variation (SD = 0.41) was nearly as large as the average difference in Impact between tertiles (0.51–0.53), indicating that manuscripts with comparable Quality scores frequently differed in their Impact ratings.

Next, we examined whether reviewer scores aligned with the JIFs to which the manuscripts were submitted. It is important to note that Discovery Stack reviewers were unaware of which journals the authors selected. Quality scores did not significantly correlate with JIFs (Fig 4E; r = −0.42, p = 0.11), whereas Impact scores were significantly correlated (Fig 4F; r = −0.51, p = 0.045). Again, analyses including all reviewer scores yielded comparable results (Quality r = −0.44, p = 0.092; Impact r = −0.57, p = 0.021). These findings suggest that reviewers’ perceptions of Impact aligned more closely with the authors’ expectations of significance than their Quality assessments. Although the modest sample size may have limited our ability to detect a statistically significant association between Quality scores and JIFs, these findings are consistent with evidence that JIF is a poor surrogate for scientific reliability [11]. Similarly, when manuscripts were grouped into JIF tiers based on the submitted or published journal, manuscripts receiving better Impact scores generally aligned with higher JIF tiers, whereas Quality scores showed weaker correspondence (S2 Fig). Together, these results demonstrate that reviewers distinguished between Quality and Impact as separate but complementary dimensions of manuscript evaluation.

A limitation of this analysis is that comparisons between Discovery Stack scores of revised manuscripts to JIFs in which manuscripts were ultimately published were not possible, as reviewers only assessed initial submissions. In addition, at the time of the final analysis, only 14 of 18 manuscripts were accepted for journal publication, which further limited comparisons between Discovery Stack scores and the JIF of the journals that ultimately accepted the manuscripts. Additionally, it is possible that reviewers inferred the tier of journal selected by authors based on formatting of the preprint.

Widespread endorsement for standardized metrics and separate evaluation

To evaluate participant perceptions following completion of the pilot study, reviewers and authors completed surveys assessing key features of the Discovery Stack model. Since separating Quality from Impact was a central component of the framework, we examined whether participants believed the Discovery Stack model effectively supported this distinction and whether doing so enhanced the review process. The vast majority of reviewers (93%) agreed or strongly agreed that the model effectively supported separate assessments of Quality and Impact (Fig 4G). Moreover, 85% of all participants agreed that separating these dimensions led to more constructive and insightful evaluations than traditional reviews (Fig 4H).

We also examined perceptions of rating standardized attributes as a means to establish quantitative metrics for scientific rigor and Impact. Support for standardized metrics was compelling, with 90% of participants agreeing that standardized Quality ratings could generate meaningful metrics of rigor (Fig 4I), and 82% agreeing that standardized Impact ratings could capture perceived significance (Fig 4J). Reviewer support exceeded author support for both metrics (92% versus 72% for Quality; 86% versus 55% for Impact), though the author sample was smaller than the reviewer sample (n = 11 versus n = 75). These findings demonstrate broad endorsement of two core elements of the Discovery Stack model: (1) the separation of Quality and Impact, and (2) the use of standardized metrics to increase transparency, improve review Quality, and reduce reliance on journal branding as a measure of scientific value. An important consideration is that participant support for standardized metrics may also reflect endorsement of the structured assessment framework itself, including the use of clearly defined evaluation criteria to guide reviewer assessments of Quality and Impact. Such guidance is uncommon in traditional peer review and may have contributed to perceptions of improved clarity, consistency, and transparency.

Because participation in this pilot was voluntary, and recruitment relied in part on prior engagement and professional networks, we considered the possibility that responses may have been influenced by prior familiarity with the Discovery Stack model. To assess this, survey responses from participants new to the study (NEW) were compared to those from participants previously enrolled in the pilot (DSP-enrolled). Across survey questions, responses were highly similar between new and previously enrolled participants, with comparable median scores and uniformly small effect sizes (S2 Table). After adjustment for multiple comparisons, only one question showed a statistically significant difference in responses, while all others were not significant. Visualization of mean differences with bootstrap CIs further demonstrated that most comparisons were centered near zero with overlapping CIs, indicating minimal differences between groups (S3A, S3B Fig). Together, these results suggest that prior familiarity with the Discovery Stack model did not meaningfully influence participant responses. Nevertheless, because participation in the study was voluntary, participants may be more receptive to alternative models of peer review than the broader scientific community, which should be considered when interpreting these results.

Impact reviews exhibit greater variability than Quality reviews

Visual inspection of side-by-side Quality and Impact scores suggested that, within individual manuscripts, Impact scores varied more than Quality scores (Fig 4A). Therefore, we compared reviewer score dispersion using three complementary metrics: SD, range (maximum − minimum score), and interquartile range (IQR; difference between the 75th and 25th percentile). Across all three metrics, Impact scores consistently exhibited greater dispersion than Quality scores (Fig 5A–5C). To account for unequal numbers of Impact and Quality reviews per manuscript, we randomly subsampled the Impact scores to match the number of Quality scores for each manuscript and recalculated the variability metrics across 5,000 iterations. Bootstrap distributions of the differences in variability between Impact and Quality scores were shifted above zero across all three metrics, consistent with greater variability in Impact assessments even after subsampling Impact reviews to match the number of Quality reviews per manuscript (Fig 5D; S3 Table). These findings underscore the value of evaluating Quality and Impact as distinct dimensions and support the need for a greater number of Impact reviewers to capture the broader range of perspectives on scientific significance.

Fig 5. Impact reviews exhibit greater variability than Quality reviews.

Fig 5

Variability between composite Quality and Impact scores, within each manuscript, was quantified using (A) standard deviation (SD), (B) range, and (C) interquartile range (IQR), and compared using Wilcoxon matched-pairs signed-rank test. p < 0.05*; p < 0.001***. (D) A subsampling approach was performed to compare variability between Impact and Quality scores. For each manuscript, Impact scores were randomly subsampled to match the number of Quality scores, and variability metrics were recalculated for each subsampled dataset. The subsampling procedure was repeated for 5,000 iterations to generate empirical distributions of the mean differences (Impact - Quality) for each variability metric. Confidence intervals and one-sided p-values were derived from these distributions, with p-values defined as the proportion of iterations in which the mean difference (Impact - Quality) was less than or equal to zero (S3 Table). Manuscript D136 was excluded as an outlier (z-score = −3.14, > 3 SD from mean difference) and D118 was excluded as additional Impact reviewers were not recruited due to Hypothes.is compatibility issues, resulting in a final sample size of n = 16 manuscripts for panels A–D. The data underlying these figures are available in S3 Table and https://doi.org/10.5281/zenodo.2004122.

Identity disclosure may be associated with greater perceived transparency and higher Impact scores

To promote transparency, reviewers were encouraged to disclose their identity to authors and co-reviewers, although anonymity remained an option to preserve the integrity of the review process. Approximately half of reviewers disclosed their identity: 50% of Quality reviewers and 54% of Impact reviewers (S4A Fig). Interestingly, a greater proportion of trainees (73%) than principal investigators (PIs) (47%) identified themselves (S4B Fig). The proportion of identified reviewers varied substantially across manuscripts (S4C Fig; range: 0%–100%), suggesting that factors such as authorship or perceived study Quality may have influenced identity disclosure decisions. The most common reasons for remaining anonymous were familiarity with the authors and concern about professional repercussions (S4D Fig).

Among the 11 authors who responded, eight (73%) agreed that the Discovery Stack review was more transparent than traditional review (S4E Fig). Authors whose manuscripts had a higher proportion of identified reviewers tended to perceive greater transparency (S4F Fig). Given the limited number of author responses, these findings should be considered preliminary observations rather than definitive evidence.

Quality reviewers’ feedback was shared with authors and Impact-only reviewers, allowing both groups to evaluate whether reviewer identity influenced scoring. Most Impact-only reviewers (89%) reported observing no clear differences in scores between identified and anonymous reviewers (S4G Fig). The small number of author responses were mixed, with six of 10 reporting no clear difference between anonymous and identified reviewers, while the remaining four stated that identified reviews were more constructive than anonymous reviews. These mixed responses warrant further study in larger cohorts.

To directly evaluate whether reviewer identity influenced scores, we compared individual scores from anonymous and identified reviewers across all manuscripts (unpaired) and within manuscripts (paired). The paired analysis assessed whether, for a given manuscript, scores differed between anonymous and identified reviewers. There were no significant differences in Quality scores between the two groups (S4I, S4J Fig), but Impact scores were significantly higher among identified reviewers (S4K, S4L Fig). These findings were consistent in the unpaired analysis across all manuscripts and within manuscripts in the paired analysis. Additionally, the higher Impact scores from identified reviewers were not explained by reviewer career stage (S4M Fig).

In summary, identity disclosure was associated with higher Impact scores and may influence perceptions of transparency. Given the small number of author responses, these observations should be considered preliminary. Nevertheless, they raise the possibility that reviewers may be more likely to identify themselves when giving favorable evaluations, or alternatively, that agreeing to disclose identity may incline reviewers toward softer Impact assessments. These observations warrant further investigation in larger studies of how identity disclosure influences both the review process and its perception.

Discovery Stack model delivers a better experience than traditional review

A central goal of the Discovery Stack Pilot was to assess whether participants believed that the model offered a better experience than traditional peer review. Overall, a strong majority (83%) rated their experience as “much better” or “slightly better”, while only 3.5% rated it as worse (Fig 6A). Positive ratings were highest among Impact-only reviewers (94%), followed by authors (82%; 9 of 11), and Quality/Impact reviewers (74%). It is important to consider that unlike traditional peer review, the Discovery Stack Pilot did not involve editorial accept or reject decisions, which may contribute to the more favorable author perceptions of the review experience. The lower satisfaction among Quality/Impact reviewers may reflect the additional effort required to complete both review phases and learn the Hypothes.is platform.

Fig 6. High satisfaction with the Discovery Stack model.

Fig 6

(A) Participants rated their overall experience with the Discovery Stack model compared to traditional review. The percent of authors (n = 11), Quality/Impact reviewers (n = 42), and Impact-only reviewers (n = 33) selecting each rating is shown. (B) Authors (n = 11) and Quality/Impact reviewers (n = 42) were asked whether the Discovery Stack model fostered a greater “Peer-Improvement” mindset than traditional peer review. (C) Participants were asked in an open-ended question what aspects of the Discovery Stack model they found most beneficial compared to traditional review. A total of 74 responses were collected, categorized into key thematic areas, and the percent of participants (n = 86) who mentioned each theme is shown. All responses are shown. (D) Participants identified aspects of the model they found most challenging. A total of 56 open-ended responses were collected, categorized into thematic categories, and the percentage of participants (n = 86) mentioning each challenge is shown. All responses are shown. The data underlying these figures are available at https://doi.org/10.5281/zenodo.2004122.

Only three participants reported a worse experience, citing the need to consult setup instructions, perceived platform complexity, or uncertainty about how their reviews affected manuscript outcomes. These challenges are typical of new systems and are expected to diminish with familiarity. Additionally, because the pilot ran in parallel to traditional review, authors were not required to respond to Discovery Stack feedback, which limited reviewers’ insight into the impact of their efforts.

A core tenet of the Discovery Stack model is to promote a “peer-improvement” mindset, encouraging reviewers to provide constructive, actionable feedback that enhances scientific rigor. This principle was operationalized through multiple components of the model, including reviewer onboarding that explicitly emphasized the value of constructive, peer-improvement-oriented feedback, as well as the structured Quality and Impact assessment forms, in-line annotation, and the separation of Quality and Impact review phases. Most participants (75%) agreed that the Discovery Stack model fostered this mindset more effectively than traditional reviews, with stronger agreement among reviewers (81%) than authors (55%; 6 of 11) (Fig 6B). Because the survey assessed perceptions of the model as a whole, this result likely reflects the combined contributions of these components, including the reviewer briefing.

To gain qualitative insight, participants were asked open-ended questions about the most beneficial and most challenging aspects of the Discovery Stack model compared to traditional review. Consistent with the positive ratings, more participants cited benefits (n = 74) than challenges (n = 56). The most common benefits were in-line commenting (35%), separation of Quality and Impact reviews (20%), and standardized questions and scoring (18.6%) (Fig 6C). The most frequent challenges were setting up and learning Hypothes.is (20%), Hypothes.is limitations (9%), and time demands (11%) (Fig 6D).

Although challenges with the Hypothes.is tool were most frequently cited, in-line commenting was also the most reported benefit, highlighting both its value as a core feature and the need for technical enhancements. With 82% of participants reporting a better experience than traditional review, these findings provide strong support for Discovery Stack and its potential for broader adoption with continued optimization.

In-line commenting improves review clarity, collegiality, and efficiency

Given that in-line commenting was both the most cited benefit and key area for improvement, we evaluated its effectiveness in improving review clarity, collegiality, and efficiency. The Hypothes.is tool was selected because it enabled contextual annotation of bioRxiv-hosted preprints within private groups, allowing reviewers to leave feedback directly on the manuscript text. Responses from both reviewers and authors strongly supported this feature. Most authors and Impact-only reviewers (71%) agreed that in-line comments were more collegial and constructive than traditional reviews (Fig 7A). Notably, all authors that responded (10 of 10) found in-line comments easier to respond to and more helpful for identifying needed revisions than traditional review summaries (Fig 7B).

Fig 7. In-line commenting improves review clarity, collegiality, and efficiency.

Fig 7

(A) Authors (n = 8-9) and Impact-only reviewers (n = 32) were asked whether they were able to read the Quality reviewers’ in-line Hypothes.is comments without difficulty, and whether they found the comments more constructive and helpful than those typically received through traditional review. (B) Authors (n = 8-9) were asked whether responding to reviewers’ in-line comments would be easier than addressing traditional review comments, and whether the in-line format helped clarify which aspects of the paper needed improvement. (C) Quality/Impact reviewers (n = 40-42) were asked a series of questions focused on providing in-line comments using the Hypothes.is platform. The data underlying these figures are available at https://doi.org/10.5281/zenodo.2004122.

Among Quality/Impact reviewers, 93% agreed that in-line comments made it easier to highlight errors in logic or clarity, 86% felt in-line commenting improved the constructiveness of feedback, and 71% preferred addressing comments directly to the authors, believing that their in-line comments would be more helpful than a traditional review summary (Fig 7C). Although the Hypothes.is tool was new to most reviewers, 73% reported that it was easy to use, and 60% agreed that in-line commenting improved review efficiency. Additionally, 67% agreed that in-line commenting helped them adopt a more collegial tone (Fig 7C). Together, these findings indicate that in-line commenting is both feasible and effective, providing clearer, more constructive, and more collegial feedback than traditional review.

Authors find Discovery Stack Quality reviews rigorous and helpful

We next evaluated whether the Discovery Stack model effectively directed reviewers to focus on scientific rigor by asking authors about their perception of the Quality reviewer feedback. Most authors (82%; 9 of 11) agreed that Quality reviews focused on scientific rigor, and 73% (8 of 11) found the standardized Quality and Impact scores helpful for understanding how their manuscript was evaluated (Fig 8A). Likewise, 82% (9 of 11) agreed the Quality Assessment Form effectively summarized the key strengths and weaknesses, and more than 60% (5 of 8) reported revising their manuscript based on reviewers’ feedback. In-line commenting was also well received, with seven of eight (88%) authors agreeing that the opportunity to interact with reviewers through Hypothes is provided a good method to clarify feedback and expedite the review process (Fig 8A).

Fig 8. Discovery Stack enables rigorous Quality review and supports meaningful Impact assessment.

Fig 8

(A) Authors (n = 8–11) were asked to rate their level of agreement with statements related to the feedback they received during the Quality review. (B) Authors (n = 11) were asked to rate the overall usefulness of feedback they received from the Discovery Stack review compared to traditional peer review. (C) Quality/Impact (n = 42) and Impact-only (n = 32) reviewers responded to questions focused on their experience with the Impact assessment. (D) Impact-only reviewers (n = 32) were asked whether the Quality reviewer's feedback from the assessment form and in-line comments aided their assessment of manuscript Impact. The data underlying these figures are available at https://doi.org/10.5281/zenodo.2004122.

When asked to compare the overall Discovery Stack feedback to traditional review, of 11 authors, five rated it better, five rated it similar, and only one rated it worse. (Fig 8B). These findings indicate that Discovery Stack delivered rigorous, constructive, and actionable Quality reviews. Although limited by a small author sample (n = 11), the results support the model's potential to improve consistency and utility of peer review.

Impact assessment viewed as useful and strengthened by prior Quality review

Next, we examined participants’ impressions of the Impact review and its value in assessing the broader significance and potential influence of a manuscript’s findings. Most reviewers (77%) reported that evaluating Impact independently made it easier to assess significance (Fig 8C), consistent with the model’s premise that separating Quality and Impact enables a more nuanced evaluation.

Reviewers also found the Impact review to be well-designed and intuitive; 93% agreed the steps were clear and manageable, 88% found the Impact Assessment Form helpful for estimating a manuscript’s “must-read” value, and 89% agreed the form included essential criteria for evaluating significance (Fig 8C). Nearly all respondents (92%) agreed that identifying the criteria applicable to each manuscript was essential for an accurate evaluation (Fig 8C). Simultaneously, reviewers acknowledged challenges of reliably assessing Impact, with 58% agreeing that assessing Impact requires more than three reviewers, and 89% agreeing that a study’s true Impact often becomes clear only over time, highlighting the importance of dynamic, evolving Impact metrics.

To facilitate the separate evaluation, the pilot was designed so that Quality reviews were completed before the Impact assessment. This allowed Impact reviewers to consider Quality review feedback while evaluating significance. More than 90% of Impact-only reviewers agreed that access to the Quality reviews improved their ability to evaluate Impact (Fig 8D). These findings indicate that reviewers viewed the Impact assessment as well-structured, valuable, and strengthened by prior Quality review.

Reviewer engagement relies on outreach, familiarity, and trainee involvement

Recruiting reviewers is a persistent challenge in peer review and a critical barrier to timely evaluations [15,22]. This challenge was amplified for the Discovery Stack Pilot, as most researchers were unfamiliar with the review model and each manuscript required six reviewers. Acceptance rates varied dramatically depending on reviewers’ familiarity with the study and whether they received personal outreach. Among individuals with no prior connection to the pilot, the acceptance rate was only 8.6% (S5A Fig). In contrast, reviewers already enrolled in the pilot accepted invitations at a much higher rate of 59% (p < 0.0001, Fisher’s Exact test).

To improve reviewer participation among new reviewers, members of the SAB sent follow-up emails to individuals in their respective networks, which substantially raised the acceptance rate among new reviewers to 53% (S5A Fig). Personal outreach increased acceptance odds by more than 12-fold (Odds Ratio (OR) = 12.2, p < 0.0001), while prior enrollment increased odds by 15-fold (OR = 15.2, p < 0.0001). With these strategies in place, for the 17 bioRxiv manuscripts reviewed in the pilot, 283 review invitations were sent (160 for Quality/Impact and 123 for Impact-only reviews), yielding an overall acceptance rate of 33% (S5B, S5C Fig), which is similar to the 32% acceptance rate reported in 2022 by Clarivate’s ScholarOne platform, which supports over 8,000 journals [22]. Overall, 90% of accepting reviewers were either familiar with the study or personally contacted by someone in their network. Survey data reinforced this trend, with 70% of reviewers and 64% (7 of 11) of authors reporting they heard about the pilot study prior to participating or alternatively received outreach from a colleague after the initial reviewer request (S5D Fig). These findings underscore that outreach and professional networks substantially improve reviewer engagement.

Although trainees who reviewed independently of their advisors represented a small proportion of reviewers (28%; S5E Fig), their acceptance rate was markedly higher than that of PIs. After adjusting for familiarity and outreach, trainees remained 15 times more likely to accept review invitations (OR = 15.3, p < 0.0001; S5F Fig). These findings highlight that actively training and recruiting trainees may be an effective strategy for increasing reviewer engagement.

Workflow analyses identify operational barriers to expedited peer review

Another persistent challenge in peer review is the prolonged duration between manuscript submission and publication, which slows the dissemination of scientific findings. Accordingly, an important objective of the Discovery Stack Pilot was to evaluate whether the review workflow could accelerate peer review while maintaining rigorous evaluation. Because the Discovery Stack model included separate Quality and Impact review phases, total review duration was not directly comparable to traditional journal review. However, the Quality review phase was designed to approximate the initial stage of journal review and therefore provided a basis for operational comparison.

The importance of developing more efficient review workflows was underscored by the publication timelines of manuscripts included in the pilot. At the time of the final analysis, 14 of 18 manuscripts had been published in journals. The mean time from submission to publication for manuscripts published in journals was 13 months. The remaining manuscripts were still under review or undergoing revisions, averaging 22 months since submission (Fig 9A).

Fig 9. Workflow Analyses identify operational barriers to expedited peer review.

Fig 9

(A) For manuscripts still under review at the time of the final analysis (08/12/26), the elapsed time from author-reported submission date to analysis date is plotted. For manuscripts published in journals, time from submission to publication reported by the journal is shown. The submission date for manuscript D165 is unknown, so it is excluded from this analysis. (B) Days required to recruit three Quality reviewers, complete Quality reviews, recruit three additional Impact-only reviewers, and complete Impact reviews is shown for each manuscript. D118 was excluded from analyses in B-D as additional Impact reviewers were not recruited due to Hypothes.is compatibility issues. (C) Total duration of the Quality phase (including recruitment and review completion), the Impact phase, and combined review time for each manuscript. (D) Average recruitment and review times per manuscript were analyzed with Kruskal–Wallis and Dunn’s multiple comparison post-test. p < 0.05*; p < 0.01**. (E) Authors (n = 11) reported the duration of initial journal review, which was compared to the Quality review duration using Wilcoxon matched-pairs signed-rank test. (F) Authors (n = 11) reported whether they received Discovery Stack Quality reviews faster than the initial traditional review. (G) Individual Quality or Impact review times were compared using the Mann-Whitney test (p = 0.0002***). (H) Individual Quality Review times were plotted by manuscript. D118 was excluded. (I) Quality/Impact reviewers (n = 40) estimated the time to complete the review, including reading the pre-print, commenting in Hypothes.is, completing both the Quality and Impact Assessment Forms. Impact-only reviewers (n = 33) estimated the time to read the pre-print along with Quality reviewers’ feedback and complete the Impact Assessment Form. (J) Quality/Impact reviewers (n = 42) rated the efficiency of Discovery Stack review, considering both the time spent and value it provided to authors, compared to traditional peer review. The data underlying these figures are available in S1 Table and https://doi.org/10.5281/zenodo.2004122.

Across the pilot, the average time from reviewer recruitment to completion of both the Quality and Impact reviews was 86 days (range: 42–125) (Fig 9B, 9C). The Quality phase averaged 50 days and the Impact phase 33 days, with reviewer recruitment averaging 15–16 days per phase. The most time-consuming component was completion of Quality reviews, which averaged 35 days (Fig 9D). Among the 11 manuscripts for which authors reported initial review times, Discovery Stack Quality review duration was comparable to the author-reported duration of initial journal review (52 versus 63 days) (Fig 9E). Although the mean Discovery Stack Quality review time was modestly shorter than the author-reported duration of initial journal review, variability between manuscripts was substantial, with several manuscripts receiving Discovery Stack reviews faster than journal review, whereas others were slower. Consistent with this variability, 55% (6 of 11) of authors who completed the surveys reported receiving Discovery Stack Quality reviews faster than journal reviews (Fig 9F).

To understand why the overall review process was not faster, given that reviewers agreed to complete reviews within 14 days, we examined factors contributing to delays in review completion. Individual turnaround times were generally close to the expected timelines, with Quality reviews averaging 18 days and Impact reviews 14 days (Fig 9G). Moreover, 61% (roughly two out of three) of Quality reviewers and 78% of Impact reviewers submitted their reviews within 3 days of the deadline. Visualization of individual reviewer times per manuscript demonstrated that delays were typically due to a single late reviewer rather than widespread delays across all reviewers (Fig 9H).

To investigate whether the time required to complete each review contributed to delays in review completion, reviewers were asked to estimate their time investment. Quality/Impact reviewers spent an average of 3.5 ± 0.97 hours per review, and Impact-only reviewers spent 1.7 ± 0.84 hours (Fig 9I). When asked to evaluate the efficiency of the review process, considering both the time invested and the value provided to authors, 91% of Quality/Impact reviewers rated it as equal to or more efficient than traditional review (Fig 9J).

Together, these findings indicate that streamlining the reviewer workflow alone is insufficient to substantially shorten review duration. Most reviewers adhered to the expected timeline, but isolated late reviews drove the overall delays observed at the manuscript level. Meaningful reductions in review time will therefore require operational interventions specifically targeting the late-reviewer problem, including pre-recruitment of additional reviewers, modest compensation tied to timely submission, and structured contingency plans for replacement when delays exceed defined thresholds.

Participant perspectives on future publishing models

To gauge how participants viewed alternative publishing approaches, we surveyed attitudes about needed reforms and desirable platform features. Open-ended responses most frequently identified inefficiencies in the peer review process, unconstructive or biased reviews, and excessive reviewer demands (S6A–S6C Fig). These concerns aligned closely with the goals of the Discovery Stack model, particularly the use of structured evaluation criteria, separation of Quality from Impact, in-line commenting, and a peer-improvement mindset.

Participants expressed strong interest in using a new platform that applies scientist-developed metrics to evaluate and curate peer-reviewed research (Fig 10A). Respondents also indicated willingness to submit an original research paper to such a platform, but this willingness was lower than using a platform to curate and read research, which likely reflects the continued influence of traditional journals as the key metric of researcher productivity.

Fig 10. Strong interest in a new scientist-driven publishing platform.

Fig 10

(A) Participants (n = 85) were asked whether they would be interested in using a new platform that utilizes scientist-developed metrics to evaluate and curate peer-reviewed research, and whether they would submit an original research paper to a new platform that utilizes scientist-developed metrics to assess Quality and Impact. (B) Participants (n = 85) were asked to rank the importance of features that would influence their likelihood of using a novel platform for evaluating and curating peer-reviewed and peer-approved research. Features were rated from “absolutely essential” to “not important”. The percent of respondents selecting each rank is shown. The data underlying these figures are available at https://doi.org/10.5281/zenodo.2004122.

When asked which features would most influence adoption of a new platform, participants prioritized a “peer-improvement” mindset (88%), a “peer-approved” status for validated manuscripts (87%), faster publication timelines (84%), a scientist-owned platform free of journal revenue and branding (81%), reviewer feedback metrics (77%), scientist-developed Quality metrics (77%), and evolving Impact metrics based on ongoing community input (72%) (Fig 10B). Additional requested features included institutional recognition by funding and promotion committees, improved access to underlying data and code, and moderation of comments to maintain constructive review Quality.

Given the modest cohort size and potential for selection bias towards participants receptive to alternative peer review models, these findings should be interpreted cautiously. Nevertheless, they suggest that several of the core features tested in the Discovery Stack Pilot align with researchers’ priorities and highlight the need for institutional recognition of the platform to overcome hesitancy to depart from traditional journals.

Discussion

The Discovery Stack Pilot evaluated the feasibility and value of a scientist-designed peer review model that separates scientific rigor (Quality) from perceived significance (Impact), uses standardized metrics, incorporates in-line commenting, and promotes a peer-improvement mindset. Conducting reviews in parallel with traditional journal reviews enabled a direct comparison of feasibility and effectiveness. Several important findings emerged. Notably, reviewers successfully evaluated Quality distinctly from Impact, supporting the conceptual separation of these dimensions. Impact assessments showed greater variability than Quality assessments, highlighting the importance of broader reviewer input when evaluating scientific significance and reinforcing the rationale for independently evaluating these dimensions. Participants also strongly endorsed core elements of the model, including structured evaluation criteria, standardized metrics, and the separation of Quality from Impact. Together, these findings support the feasibility of the Discovery Stack framework and provide a foundation for larger-scale evaluation.

A major outcome of this pilot was empirical support for separating Quality from Impact, as we had previously proposed separating these two dimensions to address the long-standing problem that perceptions of significance can mask concerns about rigor, and vice versa [21]. Although Quality and Impact were positively associated, as expected, when poor rigor limited perceived importance, residual and tertile analyses revealed frequent divergence. Moreover, Impact scores, but not Quality scores, significantly correlated with JIFs. Although this observation should be interpreted cautiously given the pilot study's modest sample size, it is consistent with growing evidence that JIF is a poor proxy for the methodological rigor and reliability of individual studies [11]. JIFs also have important conceptual and practical limitations as measures of the Quality or Impact of individual publications [9–14]. Because citation distributions are highly skewed, the arithmetic mean used to calculate JIFs may not accurately represent the citation performance of a typical article within a journal [12–14]. Additionally, publishers can negotiate which article types are included in the denominator, with independently verified discrepancies of up to 19% compared to Clarivate's published figures [10]. Importantly, studies suggest that journals with higher JIFs do not consistently publish more reliable research [10,14]. For example, higher JIFs are associated with higher retraction rates, although this relationship may partly reflect the greater visibility and scrutiny of articles published in highly cited journals [10,14]. Data from the Discovery Stack Pilot demonstrate that reviewer-assessed Quality captured information distinct from conventional journal-based JIFs, which supports the rationale for evaluating scientific rigor independently of perceived significance. Importantly, most reviewers (93%) agreed that the Discovery Stack model effectively supported separate assessments, and 85% found that separating these dimensions led to more constructive and insightful reviews. Separating Quality and Impact would substantially benefit science by assigning value to high-Quality work, independent of novelty, creating space for careful replication studies, incremental clarifications, and contrarian findings that the current novelty-oriented model often overlooks.

Data from the pilot also revealed a greater dispersion in Impact scores relative to Quality, reflecting inherent subjectivity in evaluating significance, which is sensitive to field coverage, methodological preferences, translational focus, and individual research agendas. Operationally, this finding supports the use of a larger number of Impact reviewers per manuscript to capture a representative distribution of viewpoints.

The pilot also evaluated the feasibility and utility of standardized metrics. Participants strongly endorsed metrics for both Quality (90%) and Impact (82%). This support likely reflects not only the resulting quantitative scores but also endorsement of the structured framework itself, including the use of clearly defined evaluation criteria, which is uncommon in traditional peer review. In addition to providing data to generate metrics, the Likert-scale questions aligned reviewer attention to the core elements of rigor (methodology, controls, statistics, and concordance between data and conclusions) or significance (extent of advance in scientific understanding, filling of a knowledge gap, “must-read” value), supporting more consistent evaluation across manuscripts.

One potential long-term application of standardized Quality metrics would be the development of empirically derived thresholds for scientific rigor. With substantially larger datasets and appropriate validation of score calibration, reproducibility, and predictive performance, it may become possible to identify score ranges that reliably distinguish manuscripts meeting predefined standards of scientific rigor. Although the present pilot was not designed to establish such thresholds, it demonstrates the feasibility of generating standardized Quality scores that could serve as the foundation for future investigation. Like traditional peer review, revised manuscripts would undergo additional rounds of evaluation, with Quality and Impact scores updated following each revision. The final Quality score could serve as an indicator of the rigor and reliability of the evidence supporting a manuscript’s conclusions.

In contrast, because scientific significance evolves as discoveries are replicated, extended, and applied, an important long-term goal of the Discovery Stack framework would be the development of dynamic Impact metrics rather than fixed assessments at the time of publication. Such metrics could integrate three complementary sources of information: 1) expert assessments of significance generated at the time of review, 2) citation data from bibliometric databases, capturing formal scientific uptake, and 3) broader indicators of influence, including readership, downloads, reference manager saves, press coverage, policy citations, and social media engagement. This framework would require rigorous empirical evaluation of its strengths and limitations as each of these components has limitations. While overreliance on metrics alone risks oversimplification [23], these hazards can be mitigated when metrics are based on clearly defined attributes, coupled with narrative feedback and in-line annotations, and routinely evaluated.

The Discovery Stack model builds upon several recent innovations in peer review that seek to improve the evaluation and communication of scientific research [24]. Soundness-only peer review, pioneered by journals such as PLOS ONE and later adopted by other journals, aims to evaluate methodological rigor while leaving judgments of novelty and significance to the broader scientific community after publication. However, studies suggest that reviewers and editors often continue to incorporate traditional assessments of importance and novelty into their recommendations, highlighting the practical difficulty of separating these dimensions in a conventional editorial framework [24]. The Discovery Stack model addresses this challenge by explicitly evaluating and scoring both dimensions using defined criteria.

The publish-review-curate model, implemented by eLife and now used on other platforms [25], represents another important innovation by distinguishing the strength of evidence from the significance of findings, while emphasizing transparent evaluation after publication. The Discovery Stack model shares this objective but extends the concept by generating standardized Quality and Impact metrics that allow readers to filter and prioritize work by rigor and significance at scale, rather than evaluating manuscripts on a journal-by-journal basis.

Open review models represent another innovation now used by many journals to increase transparency by publishing reviewer comments alongside manuscripts. While this approach provides valuable context and insight into the review process, readers must still evaluate individual reviews to identify rigorous and influential studies. As the volume of published research continues to grow, this becomes increasingly challenging. In addition, some journals have incorporated in-line annotation tools during peer review of manuscripts, demonstrating the feasibility of integrating contextual comments directly within manuscripts [26]. The Discovery Stack model builds upon existing approaches by combining transparent review, independent Quality and Impact assessments, in-line annotation, metrics based on standardized criteria, and a framework for dynamic post-publication evaluation of scientific influence. Together, these features enable readers to filter, search, and prioritize manuscripts by rigor and significance at scale, rather than relying on journal prestige or reading the full text of every review.

Registered Reports are another innovation that seeks to improve research rigor by evaluating study design before results are known, reducing publication and reporting biases [27]. This approach is particularly well suited to hypothesis-driven studies with predefined experimental protocols, whereas the Discovery Stack model provides a complementary framework for evaluating the Quality and Impact of completed research across a broad range of study types.

Timeliness remains one of the most significant challenges in peer review. Despite emphasizing deadlines and streamlining reviews by using in-line comments instead of lengthy narratives, the Discovery Stack Quality review phase averaged about 8 weeks, matching the initial journal review rather than shortening it as intended. Individual reviewers generally adhered to the expected timelines, and most delays were caused by a single reviewer. Replacing a late reviewer rarely shortens the overall review time because confirming a replacement takes time, and the new reviewer requires time to complete the review. While many journals address this by relying on two reviewers instead of three, this approach risks weakening the overall Quality of the evaluation. The Discovery Stack Pilot demonstrated that reducing the review duration requires more than workflow simplification. Rather, clear accountability and pre-planned backup coverage are required. Practical steps include confirming four reviewers at the outset, recruiting “alternate” reviewers, and rewarding timely submissions with monetary compensation and recognition. Notably, a recent study combining pre-recruitment of expert reviewers, compensation, and strict deadlines achieved review turnaround times of less than 1 week, illustrating how innovative recruitment strategies and incentives can expedite the review process [28].

Reviewer recruitment and reviewer burden are practical considerations for adoption of the Discovery Stack model, as reviewer recruitment remains a major bottleneck in peer review and adds substantially to total review time [15,22]. In our pilot, confirming three Quality reviewers averaged 2 weeks per manuscript, a nontrivial delay that accounted for one quarter of the 8-week Quality review phase. Rather than recruiting nine fully independent reviewers per manuscript (three Quality and six Impact), Quality reviewers also performed a separate Impact assessment of the same manuscript. Subsequently, three additional Impact-only reviewers were recruited specifically to increase the number and diversity of Impact evaluations. Importantly, the Impact-only review phase was designed to reduce reviewer burden by focusing exclusively on scientific significance and allowing reviewers to reference prior Quality reviewer comments addressing rigor and methodology. Consistent with this streamlined role, Impact-only reviewers reported spending an average of 1.7 hours per review, substantially less than the 3.5 hours reported by reviewers to complete the Quality assessment. Thus, additional Impact reviewers add less to the total reviewer burden than their number alone would suggest. Nevertheless, reviewer recruitment remains an important consideration for future implementations of the model. Reducing the reviewer bottleneck will ultimately depend on increasing the reviewer pool. Our recruitment analyses provide important guidance for future deployment of the Discovery Stack model. Familiarity with the project and direct outreach markedly improved reviewer acceptance rates, suggesting that current enthusiasm for the model may not yet extend broadly across the wider research community and that broader implementation will require strategies to engage scientists beyond early adopters and existing professional networks. Additionally, trainees were especially responsive to accepting review assignments. These insights suggest that visible recognition, modest compensation, and formal training for early-career scientists may help expand the reviewer pool, while also providing opportunities to develop constructive peer review skills and establish scientific reputations. Modest reviewer compensation was intentionally incorporated into the pilot because equitable compensation for peer-review labor is a central principle of the Discovery Stack model. However, the present study was not designed to determine the relative contributions of compensation, professional outreach, prior familiarity, and scientific interest to reviewer participation. Notably, 73% of compensated reviewers (55 of 75) elected to donate their honoraria rather than accept it, suggesting that financial compensation was not the sole motivation for many participants. Nonetheless, identifying sustainable funding sources to support reviewer compensation will be an important consideration for implementation of alternative peer-review models.

Despite recruitment challenges, survey responses from newly recruited participants were broadly similar to those from previously engaged participants, suggesting that support for the Discovery Stack model was not solely driven by prior familiarity with the project. Nevertheless, because participation was voluntary and recruitment relied on existing networks, these findings should be interpreted cautiously. Within this cohort, survey responses revealed broad dissatisfaction with current publishing norms and strong enthusiasm for reform. Accordingly, 92% of participants indicated a willingness to use a scientist-designed platform to evaluate and curate peer-reviewed research. Features most likely to drive adoption align with the Discovery Stack model, including a peer-improvement mindset, a “peer-approved” status to indicate scientific validation, faster timelines to publication, scientist ownership, and reviewer feedback metrics. These results suggest the Discovery Stack model provides a promising foundation for further refinement and evaluation.

Although willingness to submit manuscripts was strong (71%), it lagged behind readiness to use a platform for evaluation and curation, reflecting continued dependence on journal prestige and JIFs for career advancement and funding decisions. Despite the limitations of JIFs, these metrics remain the primary mechanism used by funding and academic institutions to assess the value and impact of scientific research. The reliance on journal-based JIFs presents a major barrier to the adoption of new publishing models. Accordingly, successful implementation of scientist-designed peer review platforms will require not only technical innovation but also broader acceptance by funding agencies, academic institutions, and other stakeholders. Encouragingly, movements promoting improved research assessment are emerging [7,23,29], and the Discovery Stack model, with its independent evaluation of scientific rigor and Impact, may provide a practical complement to these initiatives.

This study has limitations. The modest sample size limits statistical power and the strength of survey-based conclusions. The focus on immunology and cancer biology preprints, selected from a predefined eligibility pool rather than randomly sampled across bioRxiv, limits generalizability of the findings to other scientific disciplines. In addition, the pilot was not designed to incorporate explicit stratification by manuscript or author characteristics, including geographic distribution, institutional diversity, or other demographic features. Because participation was voluntary and leveraged prior engagement and professional networks, the study population may be enriched for participants more receptive to alternative peer review models, which further constrains generalizability of participant perceptions. In addition, surveys were completed by a single representative author from each participating manuscript, typically the senior author or first author responsible for the review process. Consequently, this pilot was not designed to evaluate whether perceptions differed according to authorship role or career stage. Future studies should include larger numbers of authors from individual research teams to determine whether perspectives vary among senior authors, first authors, and other contributors. Other limitations include incomplete publication outcomes for some manuscripts, reliance on an external annotation tool, the evaluation of mostly initial submissions, which limited both the analysis of score dynamics across revisions and direct comparison of Discovery Stack scores with the JIFs of journals in which manuscripts were ultimately published. These constraints are typical of feasibility studies and highlight the need to address these limitations in future iterations through larger and more diverse cohorts.

In summary, the Discovery Stack Pilot demonstrated that reviewers reliably evaluated Quality and Impact as separate dimensions, and that Impact, by its nature, varies more than Quality. The model’s core elements of standardized metrics, in-line annotation, peer-improvement mindset, and distinct Quality and Impact evaluations enhanced clarity and constructiveness, earning considerable support. Overall, these results provide a practical foundation for continued refinement and future studies to assess the performance, scalability, and broader applicability of the Discovery Stack model across more diverse scientific communities.

Materials and methods

Study design

The Discovery Stack Pilot consisted of three sequential phases designed to test a structured, peer review process that separately evaluated Quality and Impact. These phases were: 1) Quality Review, 2) Author Response, and 3) Impact Review.

Quality review phase

The Quality review phase provided a structured, standardized evaluation of each manuscript’s scientific rigor. Reviewers were instructed that Quality referred to the appropriateness of experimental design, methodology, controls, statistical analyses, and sample sizes, as well as whether the data supported the stated conclusions. To support consistency and emphasize that improving scientific Quality was the primary objective of this phase, reviewers received a Quality Assessment Checklist (S7 Fig) directing them to evaluate four key areas: 1) whether conclusions were adequately supported by the data and free of contradictions, 2) soundness of the experimental design, controls, methodology, and statistical analyses, 3) data integrity and reproducibility, including sample sizes and number of replicates, and 4) clarity of presentation.

Quality reviewers provided feedback through two complementary mechanisms. First, in-line comments were added directly to the manuscript preprint using the Hypothes.is tool, mirroring how scientists typically provide feedback to colleagues during manuscript preparation. Comments were initially posted in private groups visible only to the editor, then compiled and shared in groups visible to all reviewers and the authors. Comments from reviewers who wished to remain anonymous were de-identified before sharing.

Second, reviewers completed a standardized Quality Assessment Form consisting of six short-answer questions and multiple Likert-scale questions (Fig 2A, 2B). This form was designed to capture summary-level assessments and test the feasibility and utility of using structured, quantitative metrics to evaluate scientific rigor. All assessment forms were designed, distributed, and collected using the Formaloo platform.

Author response phase

Following submission of Quality reviews, feedback was compiled and shared with authors and other reviewers. In-line comments were shared by inviting authors and reviewers to a shared Hypothes.is group containing all reviewer annotations. Quality Assessment Form feedback was compiled in a summary PDF that included short-answer responses and graphical representations of the Likert-style questions (Fig 2C). Both in-line comments and assessment form feedback were labeled by reviewer. Reviewers who disclosed their identity were named, while anonymous reviewers were designated “Reviewer #”. Authors were encouraged, though not required, to reply directly to reviewers in Hypothes.is. Authors were not required to conduct additional experiments for the pilot.

Impact review phase

Impact was defined as the extent to which a study advances scientific understanding, addresses critical knowledge gaps, influences multiple fields, or has therapeutic relevance. Reviewers were informed that a manuscript could be of high Quality yet have limited Impact if its contribution was incremental or relevant to only a small audience.

To ensure sufficient evaluations for each manuscript, Quality reviewers were also asked to provide a separate assessment of the manuscript’s potential Impact. After completing the Quality Assessment Form, reviewers were directed to a separate Impact Assessment Form including questions evaluating transformative potential, generalizability, technological advancement, and mechanistic insight (Fig 3A–3D). This structured framework enabled Impact to be evaluated independently of Quality while capturing the diversity of perspectives on scientific significance.

Impact-only review phase

Since assessment of Impact is inherently subjective, additional reviewers were recruited to focus solely on the Impact evaluation. The goal was to obtain a minimum of three Quality and six Impact reviews per manuscript. Impact-only reviewers were recruited after the Quality phase was completed and they received the manuscript, Quality reviewers' in-line comments, and assessment form feedback. Impact-only reviewers completed the same Impact Assessment Form described above (Fig 3A–3D).

Reviewer experience, training, expectations, and compensation

Reviewers were selected based on their subject matter expertise, identified through author recommendations, nominations by members of the SAB, previously enrolled participants, or PubMed searches by the editor. While most reviewers were PIs, trainees also completed reviews with prior endorsement from their PI or as part of collaborative reviews. Most reviewers brought substantial experience to the process: 83% previously reviewed 10 or more manuscripts, 49% reported over a decade of experience submitting and reviewing papers, and 80% had been engaged in peer review for at least 5 years (S5G, S5H Fig).

Before receiving review materials, reviewers were invited to attend a brief Zoom onboarding session or receive detailed instructions via email. Most Quality reviewers (88%) attended a Zoom onboarding meeting, whereas most Impact-only reviewers (90%) received detailed instructions via email. Onboarding covered Hypothes.is setup, pilot goals, and emphasized the value of a peer-improvement mindset, defined as offering clear, constructive, and actionable feedback to improve scientific rigor and suggest additional experiments only when essential to support the conclusions, feasible for the research group, and within the scope of the study. Quality reviewers agreed to complete reviews within 14 days and Impact-only reviewers within 10 days. Timeliness was communicated during reviewer recruitment, reinforced during onboarding sessions, emphasized in the review instructions, and reiterated in follow-up reminder emails. To acknowledge their contributions and reflect our commitment to a platform that compensates reviewers, Quality reviewers were offered $30, and Impact-only reviewers were offered $20 per review. Reviewers had the option to donate their compensation to Solving For Science rather than accept payment. Of the 75 reviewers who completed the surveys, 55 (73%) elected to donate their compensation.

Manuscript enrollment

Five manuscripts were initially enrolled following email invitations to previously enrolled participants. Subsequently, manuscripts were recruited from recent bioRxiv postings based on predefined criteria, including 1) relevance to immunology or cancer biology and 2) recent posting date. Using these criteria, 118 authors were contacted, resulting in the enrollment of 11 more manuscripts. Two additional bioRxiv preprints were included in the pilot despite lack of author response, as reviewers with appropriately matched expertise had already been identified and agreed to participate. In total, 18 manuscripts were reviewed in the Discovery Stack Pilot. One manuscript (D118) was hosted on SSRN rather than bioRxiv and was excluded from some analyses due to incompatibility of Hypothes.is and SSRN.

At the conclusion of the pilot, 11 of 18 authors (61%) completed the author feedback survey, which captured information about the submission status of their manuscripts. Of these 11 manuscripts, eight were first-time submissions to a traditional journal, while three were revised drafts following one round of journal revision. As of August 12, 2026, 14/18 manuscripts reviewed in the pilot had been accepted for publication in a traditional journal. Upon completion of the pilot, surveys were sent to a single author from each manuscript, either the senior author or the first author as middle authors are typically less involved in the review process. Nine of the 11 responding authors were PIs and senior authors, while the remaining two author respondents were first authors (a postdoc and graduate student).

Development and validation of composite Quality and Impact metrics

Quality and Impact metrics were developed using standardized assessment forms containing Likert-scale questions rating defined attributes of each dimension on a 1–5 scale (1 = strongly agree, best; and 5 = strongly disagree, worst (Figs 2, 3; S1A–S1D).

Quality scores were derived from 13 Likert-scale questions addressing experimental design, controls, statistical analysis, reproducibility, and data support for conclusions. Each question was weighted according to its relative importance, and weighted responses were averaged to yield a single composite Quality score per review (S1A Fig). All manuscripts received at least two composite Quality scores; most received three. To evaluate the effect of weighting, weighted and unweighted averages were compared to the average score of a key item (“Overall, the data support the conclusions presented in the paper”), which served as a benchmark of overall rigor. Weighted scores aligned with unweighted averages but trended toward the benchmark (S1C Fig).

Impact scores were generated from one multiple-choice question and five Likert-scale items. In the multiple-choice question, reviewers selected features contributing to a manuscript’s Impact (Fig 3). Each feature was weighted by importance, and the weighted sums were normalized to a 1–5 scale and combined with the weighted average of the five Likert-scale responses to generate a single composite Impact Score for each review. All manuscripts received at least five composite Impact scores, and most (15 of 17) received six or more. Unweighted and weighted averages, with and without the multiple-choice question, were compared and showed similar results (S1D Fig).

Most manuscripts reviewed in the pilot were initial submissions. However, three manuscripts were revised versions after one round of review. Composite Quality and Impact scores of initial and revised submissions were compared to determine if revised manuscripts should be analyzed separately. No significant differences, or even trends toward higher scores were observed in the revised manuscripts, so all manuscripts were grouped together for further analyses (S1E, S1F Fig).

Composite Quality and Impact score analysis

Associations between Quality and Impact scores were tested using Pearson’s correlation and linear regression models in GraphPad Prism. Analyses were performed using all available Quality and Impact scores, and separately using only scores from reviewers who evaluated both Quality and Impact.

Variability between Quality and Impact scores was assessed using SD, range, and IQR. Quartiles were calculated using the inclusive quantile definition (QUARTILE.INC in Excel) and IQR was defined as the difference between the 75th and 25th percentiles (Q3-Q1). Differences were compared using the Wilcoxon matched-pairs signed-rank test. Manuscript D136 was excluded as an outlier (z-score = −3.14, > 3 SD from mean difference). To account for unequal reviewer numbers, a subsampling approach was implemented using R. For each manuscript, Impact scores were randomly subsampled to match the number of Quality scores, and variability metrics were recalculated for each subsampled dataset. The subsampling procedure was repeated for 5,000 iterations to generate empirical distributions of the mean differences (Impact - Quality) for each variability metric. CIs and one-sided p-values were derived from these distributions, with p-values defined as the proportion of iterations in which the mean difference (Impact - Quality) was less than or equal to zero.

Survey design and analysis

At the completion of the pilot study, surveys were distributed to 93 individuals who participated as authors, Quality/Impact reviewers, or Impact-only reviewers. Author surveys were distributed only to the author who enrolled the manuscript in the Discovery Stack Pilot, as these individuals were directly engaged with the Discovery Stack review process and this approach avoided disproportionate weighting of manuscripts with large author lists. Surveys were generated and distributed using the Formaloo platform. Three distinct surveys were developed, each tailored to the participant’s role. Questions on each survey are available at https://doi.org/10.5281/zenodo.20041225. Individuals who served in multiple roles received a separate survey for each role. In total, 101 survey invitations were sent, and 86 completed responses were received, yielding an 85% completion rate.

Responses included a mix of Likert-scale, multiple-choice, and open-ended formats. Quantitative responses were summarized as percentages of total respondents per question. Open-ended responses were coded thematically, and frequencies were tallied to identify common themes.

To assess potential bias associated with prior familiarity with the study, responses from participants NEW were compared to those from previously enrolled participants (DSP-enrolled) using Mann-Whitney U tests. Analyses were restricted to survey items with sufficient responses across groups. Questions that were only administered to authors were excluded due to the limited number of author respondents (NEW, n = 7; DSP-enrolled, n = 4). For all other questions, responses from all participant roles (Quality/Impact reviewers, Impact-only reviewers, authors) were pooled within the NEW and DSP-enrolled groups and compared. Lower scores corresponded to more favorable responses. P-values were adjusted using the Benjamini–Hochberg procedure. Effect sizes (r) were calculated from the standardized Mann-Whitney statistic as r = Z/√N, where Z was derived using the normal approximation with continuity and tie corrections and N is the total number of observations. These analyses were performed in R.

Supporting information

S1 Fig. Calculation and comparison of Quality and Impact composite scores.

Composite Quality and Impact review scores were calculated to condense the full set of ratings from each review into a single metric. (A) Composite Quality scores were derived from 13 Likert-scale questions that were weighted based on relative importance. Composite Quality score = ∑ (Question Score x Importance weight)/ ∑ (Importance weight). (B) Composite Impact scores were based on responses to one multiple-choice question and five Likert-scale items. To integrate responses across two question types, each feature in the multiple-choice question was assigned a weight reflecting its relative importance to Impact, the weighted sum of selected features was calculated and then normalized to a 1−5 scale using a linear transformation. Transformed Q2 = 6 - (1 + ((∑Q2–1)*4/(36−1))). In this scale, selecting no features corresponds to a score of 5 (lowest Impact) and selecting features totaling the maximum score achieved (36) corresponds to a score of 1 (highest Impact). This transformed Q2 value was then combined with the weighted average of the five Likert-scale responses to generate a single composite Impact score for each review. (C) The weighted Quality composite score, unweighted average of Likert-scale responses, and the average score for the first statement (Overall, the data support the conclusions presented in the paper) were compared across manuscripts. (D) The weighted composite Impact score (including Q2), the unweighted average of Q2–Q7, and the unweighted average of Q3-Q7 were compared across manuscripts. (E, F) Weighted composite scores for (E) Quality or (F) Impact were compared between manuscripts enrolled as initial submissions versus revised versions. The version of three manuscripts (D165, D173, and D178) was unknown, so are excluded. No significant differences were observed by Mann–Whitney test. The data underlying these figures are available at https://doi.org/10.5281/zenodo.2004122.

(TIFF)

pbio.3003986.s001.tiff (1.3MB, tiff)
S2 Fig. Alluvial plot of reviewer Quality and Impact scores by submitted JIF tier.

Manuscripts are grouped on the left by the JIF of the journal to which they were (A) submitted or (B) published, binned into three tiers: High (JIF ≥ 42.5), Mid (JIF 15.7–27.6), and Low (JIF ≤ 9.1). Tier boundaries were defined by the actual JIFs. Within each tier, manuscripts were ordered from top to bottom first by JIF then combined Quality + Impact score. Each curve connects a manuscript's tier position on the left to its average reviewer score on the right. Reviewer scores range from 1 (best) to 5 (worst), with lower scores indicating stronger reviews. Because Quality and Impact scores span different ranges across the dataset, each metric is displayed on its own independently scaled right-hand axis, each beginning 0.1 units below the lowest recorded score for that metric. Flatter curves indicate greater agreement between JIF tier and reviewer score. The data underlying these figures are S1 Table and https://doi.org/10.5281/zenodo.20041225.

(TIFF)

pbio.3003986.s002.tiff (1.8MB, tiff)
S3 Fig. Comparison of survey responses between newly recruited (NEW) and previously enrolled (DSP-enrolled) participants.

(A) Points represent the difference in mean Likert scores between participants new to the study (NEW) and those previously enrolled (DSP-enrolled) for each survey question. Horizontal lines indicate 95% bootstrap confidence intervals. Lower scores correspond to more favorable responses. The vertical line at zero indicates no difference between groups. (B) Distribution of responses for each survey question among NEW and DSP- enrolled participants. For each question, responses from NEW participants are shown in the upper bar and DSP-enrolled responses are in the lower bar. Bars are centered on the neutral response category, with more favorable responses extending to the left and less favorable responses to the right. The data underlying these figures are available in S2 Table and https://doi.org/10.5281/zenodo.20041225.

(TIFF)

pbio.3003986.s003.tiff (455.7KB, tiff)
S4 Fig. Identity disclosure may be associated with greater perceived transparency and higher Impact scores.

(A) The percent of Quality (n = 50) or Impact (n = 114) reviews for which the reviewer's  identity was disclosed or remained anonymous. (B) The percent of PIs (n = 129) and trainees (n = 33) who disclosed their identity or remained anonymous. (C) The percent of identified reviewers per manuscript. (D) Reviewers who remained anonymous (n = 35) were asked a follow-up multiple-choice question regarding the reasons that influenced their choice. (E) Authors (n = 11) were asked if they perceived the Discovery Stack review process to be more transparent than traditional review. (F) Transparency ratings for authors that completed surveys (n = 11) were compared to percent of identified reviewers for each manuscript using Pearson’s correlation. (G) Authors (n = 10) and Impact-only reviewers (n = 18) were asked whether they noticed a difference in scoring between identified and anonymous reviewers (H) Authors (n = 10) were asked follow-up questions regarding their perception of feedback from identified and anonymous reviewers. (I, K) Composite (I) Quality and (K) Impact scores were compared between identified and anonymous reviewers across all manuscripts using the Mann-Whitney test (p = 0.0082** for Impact). (J, L) Within each manuscript, the average (J) Quality and (L) Impact scores from identified and anonymous reviewers were compared using the Wilcoxon matched-pairs signed rank test (p = 0.026* for Impact). Manuscripts with no identified reviewers (Quality n = 2, Impact n = 1) or no anonymous reviewers (Quality n = 3; Impact n = 3) were excluded. (M) Composite Quality and Impact scores were compared between PIs (Quality n = 37; Impact n = 92) and trainees (Quality n = 12; Impact n = 21). Differences were tested using the Mann–Whitney test. The data underlying these figures are available at https://doi.org/10.5281/zenodo.20041225.

(TIFF)

pbio.3003986.s004.tiff (1.1MB, tiff)
S5 Fig. Reviewer engagement depends on outreach, networks, and trainee involvement.

(A) The positive response rate was compared between individuals who were already familiar with the study (DSP-enrolled), new to the study (New to DSP), or new but contacted by a known colleague on the Scientific Advisory Board (New to DSP + email) using two-sided Fisher’s exact tests (p < 0.0001). Odds ratios (ORs) and 95% CIs were calculated from the contingency table. P-values were adjusted for multiple pairwise comparisons using the Bonferroni correction. Error bars represent 95% CIs. DSP-enrolled versus New to DSP (OR=15.2, 95% CI: 7.1–32.7, p < 0.0001, adjusted), New to DSP + email versus New to DSP (OR=12.2, 95% CI: 5.8–25.7, p < 0.0001, adjusted), DSP-enrolled versus New to DSP + email significant (OR=1.25, 95% CI: 0.64–2.4, p = 0.61, adjusted). (B) Number of reviewer invitations per manuscript to recruit Quality/Impact and Impact-only reviewers. (C) The percent of individuals that responded “yes”, “no”, or did not respond to reviewer invitations per manuscript. The overall acceptance was 33%. (D) Survey respondents indicated whether they were unfamiliar with the study, previously enrolled, or recruited by a colleague. (E) Percent of Quality (n = 48) or Impact (n = 110) reviewers at each career stage. Among Quality reviewers, 73% were Principal Investigators (PIs) while 80% of Impact reviewers were PIs. Only independent trainee reviewers (i.e., not co-reviewing with a PI) were included in these percentages. Other includes industry positions. (F) A multivariable logistic regression model estimated the odds of reviewer acceptance as a function of position (PI versus trainee) and participant type (DSP-enrolled, New to DSP, New to DSP + email) as predictors. Only independent trainee reviewers, not co-reviewers, were included in the analysis. OR and 95% CI were obtained by exponentiating the logistic regression coefficients. Statistical significance was assessed using Wald tests (OR = 15.3, 95% CI: 4.2–56.2, p < 0.0001). (G) Survey participants reported their review experience. (H) The number of years of peer review experience reported by participants. The data underlying these figures are available at https://doi.org/10.5281/zenodo.20041225.

(TIFF)

pbio.3003986.s005.tiff (804.8KB, tiff)
S6 Fig. Top priorities for improved scientific publishing.

(A) Participants (n = 86) were asked what they believe needs to change about the current scientific publishing and peer review systems. A total of 65 open-ended responses were collected, categorized into key thematic areas, and tallied. The number of respondents mentioning each concern is shown. (B) Participants were also asked to imagine that scientific publishing didn’t exist and to identify three core principles they would prioritize if building a system from scratch. A total of 53 open-ended responses were collected, categorized by theme, and tallied. (C) Responses across both questions were averaged and ranked to highlight the most broadly recognized priorities for reform. The data underlying these figures are available at https://doi.org/10.5281/zenodo.20041225.

(TIFF)

pbio.3003986.s006.tiff (1.3MB, tiff)
S7 Fig. Quality assessment checklist.

The Quality Assessment checklist was given to Quality/Impact reviewers to highlight key elements to consider when evaluating the Quality of a manuscript. It offers a comprehensive framework for reviewing conclusions, assessing experimental design, ensuring data integrity, and evaluating clarity and scholarly analysis. This guide contains resources adapted from [30] https://doi.org/10.5281/zenodo.5484087.

(TIFF)

pbio.3003986.s007.tiff (1.3MB, tiff)
S1 Table. Composite reviewer scores, JIFs, and journal review timelines per manuscript.

Individual reviewer composite scores and average composite scores are shown for both Quality and Impact evaluations for each manuscript. Lower scores are more favorable evaluations. The table also shows the JIF of the journal to which each manuscript was initially submitted and ultimately published (when applicable), as well as the duration of journal review for manuscripts published in journals.

(XLSX)

pbio.3003986.s008.xlsx (12KB, xlsx)
S2 Table. Comparison of survey responses between newly recruited (NEW) and previously enrolled (DSP-enrolled) participants.

Survey responses from (NEW) and those previously enrolled (DSP-enrolled) were compared using Mann-Whitney U tests. Analyses were restricted to survey items with sufficient responses across groups. Questions that were only administered to authors were excluded due to the limited number of author respondents (NEW, n = 7; DSP-enrolled, n = 4). For all other questions, responses from all participant roles (Quality/Impact reviewers, Impact-only reviewers, authors) were pooled within the NEW and DSP-enrolled groups and compared. Lower scores corresponded to more favorable responses. P-values were adjusted using the Benjamini–Hochberg procedure. Effect sizes (r) were calculated from the standardized Mann–Whitney statistic as r = Z/√N, where Z was derived using the normal approximation with continuity and tie corrections and N is the total number of observations.

(DOCX)

pbio.3003986.s009.docx (22KB, docx)
S3 Table. Bootstrap summary comparing variability of Quality and Impact reviewer scores.

Mean differences (Impact - Quality), confidence intervals, and one-sided p-values derived from 5,000 bootstrap iterations in which Impact scores were randomly subsampled to match the number of Quality scores for each manuscript. One-sided p-values were defined as the proportion of iterations in which the mean difference (Impact - Quality) was less than or equal to zero.

(XLSX)

pbio.3003986.s010.xlsx (9.3KB, xlsx)

Acknowledgments

We thank Solving For Science for support associated with undertaking this study. We thank Igor Brodsky, Nicole Scharping, Julien Gaillard, Ananda Goldrath, Nikhil Joshi, Savan Ram, Carla Rothlin, Sunny Shin, and Joe Sun for helpful advice during the planning, implementation, and analysis of the pilot.

Abbreviations

CI

confidence interval

IQR

interquartile range

JIF

Journal Impact Factor

NEW

new to the study

ORs

Odds ratios

PIs

Principal Investigators

SAB

Scientific Advisory Board

SD

standard deviation

Data Availability

Source data are available in Supplementary Tables or at https://doi.org/10.5281/zenodo.20041225, as indicated in the Figure Legends. The code for analyses is available at https://github.com/mmcgargi/Discovery-Stack-Analysis and archived at https://doi.org/10.5281/zenodo.22087956.

Funding Statement

MM received partial support from Solving For Science (solvingfor.org). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. No additional specific support was received for this work.

References

  • 1.Ghasemi A, Mirmiran P, Kashfi K, Bahadoran Z. Scientific publishing in biomedicine: a brief history of scientific journals. Int J Endocrinol Metab. 2022;21(1):e131812. doi: 10.5812/ijem-131812 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Jefferson T, Rudin M, Brodney Folse S, Davidoff F. Editorial peer review for improving the quality of reports of biomedical studies. Cochrane Database Syst Rev. 2007;2007(2):MR000016. doi: 10.1002/14651858.MR000016.pub3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Smith R. Peer review: a flawed process at the heart of science and journals. J R Soc Med. 2006;99(4):178–82. doi: 10.1177/014107680609900414 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Superchi C, Hren D, Blanco D, Rius R, Recchioni A, Boutron I, et al. Development of ARCADIA: a tool for assessing the quality of peer-review reports in biomedical research. BMJ Open. 2020;10(6):e035604. doi: 10.1136/bmjopen-2019-035604 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Superchi C, González JA, Solà I, Cobo E, Hren D, Boutron I. Tools used to assess the quality of peer review reports: a methodological systematic review. BMC Med Res Methodol. 2019;19(1):48. doi: 10.1186/s12874-019-0688-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Tennant JP, Ross-Hellauer T. The limitations to our understanding of peer review. Res Integr Peer Rev. 2020;5:6. doi: 10.1186/s41073-020-00092-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Jancovich L, Pitches C, Stevenson D. Failures in impact evaluation. Res Eval. 2025;34:rvaf033. [Google Scholar]
  • 8.Pagliaro M. Publishing scientific articles in the digital era. OSJ. 2020;5(3). doi: 10.23954/osj.v5i3.2617 [DOI] [Google Scholar]
  • 9.Sick of impact factors. Reciprocal Space. Available from: https://occamstypewriter.org/scurry/2012/08/13/sick-of-impact-factors/. Accessed 2026 June 17.
  • 10.Brembs B, Button K, Munafò M. Deep impact: unintended consequences of journal rank. Front Hum Neurosci. 2013;7:291. doi: 10.3389/fnhum.2013.00291 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Brembs B. Prestigious science journals struggle to reach even average reliability. Front Hum Neurosci. 2018;12:37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Dimitrov JD, Kaveri SV, Bayry J. Metrics: journal’s impact factor skewed by a single paper. Nature. 2010;466(7303):179. doi: 10.1038/466179b [DOI] [PubMed] [Google Scholar]
  • 13.Seglen P. The skewness of science. J Am Soc Inf Sci. 1992;43(9):628. doi: 10.1002/(SICI)1097-4571(199210)43:9<628::AID-ASI5>3.0.CO;2-0 [DOI] [Google Scholar]
  • 14.Fang FC, Casadevall A. Retracted science and the retraction index. Infect Immun. 2011;79(10):3855–9. doi: 10.1128/IAI.05661-11 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Hanson MA, Barreiro PG, Crosetto P, Brockington D. The strain on scientific publishing. Quant Sci Stud. 2024;5:823–43. [Google Scholar]
  • 16.Huisman J, Smits J. Duration and quality of the peer review process: the author’s perspective. Scientometrics. 2017;113(1):633–50. doi: 10.1007/s11192-017-2310-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Nelson L, Ye H, Schwenn A, Lee S, Arabi S, Hutchins BI. Robustness of evidence reported in preprints during peer review. Lancet Glob Health. 2022;10(11):e1684–7. doi: 10.1016/S2214-109X(22)00368-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Lyons-Warren AM, Aamodt WW, Pieper KM, Strowd RE. A structured, journal-led peer-review mentoring program enhances peer review training. Res Integr Peer Rev. 2024;9(1):3. doi: 10.1186/s41073-024-00143-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Aczel B, Szaszi B, Holcombe AO. A billion-dollar donation: estimating the cost of researchers’ time spent on peer review. Res Integr Peer Rev. 2021;6(1):14. doi: 10.1186/s41073-021-00118-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Sever R, Hindle S, Roeder T, Fereres S, Fernández Gayol O, Ghosh S, et al. BioRxiv: the preprint server for biology. bioRxiv. Available from: https://www.biorxiv.org/content/10.1101/833400v1. 2019. Accessed 2025 September 11. [Google Scholar]
  • 21.Krummel M, Blish C, Kuhns M, Cadwell K, Oberst A, Goldrath A, et al. Universal principled review: a community-driven method to improve peer review. Cell. 2019;179(7):1441–5. doi: 10.1016/j.cell.2019.11.029 [DOI] [PubMed] [Google Scholar]
  • 22.Dance A. Stop the peer-review treadmill. I want to get off. Nature. 2023;614(7948):581–3. doi: 10.1038/d41586-023-00403-8 [DOI] [PubMed] [Google Scholar]
  • 23.Helmer S, Blumenthal DB, Paschen K. What is meaningful research and how should we measure it? Scientometrics. 2020;125(1):153–69. doi: 10.1007/s11192-020-03649-5 [DOI] [Google Scholar]
  • 24.Spezi V, Wakeling S, Pinfield S, Fry J, Creaser C, Willett P. Let the community decide? The vision and reality of soundness-only peer review in open-access mega-journals. J Doc. 2018;74:137–61. [Google Scholar]
  • 25.Richter FC, Gea‐Mallorquí E, Mortha A, Ruffin N, Vabret N. The preprint club. EMBO Rep. 2023;24(6). doi: 10.15252/embr.202357258 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Nicholson DN, Rubinetti V, Hu D, Thielk M, Hunter LE, Greene CS. Examining linguistic shifts between preprints and publications. PLoS Biol. 2022;20(2):e3001470. doi: 10.1371/journal.pbio.3001470 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Chambers CD, Tzavella L. The past, present and future of registered reports. Nat Hum Behav. 2022;6(1):29–42. doi: 10.1038/s41562-021-01193-7 [DOI] [PubMed] [Google Scholar]
  • 28.Gorelick DA, Clark A. Fast & fair peer review: a pilot study demonstrating feasibility of rapid, high-quality peer review in a biology journal. bioRxiv. 2025. Available from: https://www.biorxiv.org/content/10.1101/2025.03.18.644032v1 [Google Scholar]
  • 29.Rushforth A. Beyond impact factors? Lessons from the Dutch attempt to transform academic research assessment. Res Eval. 2025;34:rvaf035. [Google Scholar]
  • 30.Foster A, Hindle S, Murphy KM, Saderi D. Open reviewers reviewer guide; 2021.

Decision Letter 0

Roland Roberts

26 Nov 2025

Dear Dr Krummel,

Thank you for submitting your manuscript entitled "Discovery Stack Pilot: Feasibility and Outcomes of a Scientist-Designed Peer Review Model Separating Quality and Impact" for consideration as a Meta-Research Article by PLOS Biology. Please accept my apologies for the delayed incurred while I was travelling.

Your manuscript has now been evaluated by the PLOS Biology editorial staff, and I'm writing to let you know that we would like to send your submission out for external peer review.

However, before we can send your manuscript to reviewers, we need you to complete your submission by providing the metadata that is required for full assessment. To this end, please login to Editorial Manager where you will find the paper in the 'Submissions Needing Revisions' folder on your homepage. Please click 'Revise Submission' from the Action Links and complete all additional questions in the submission questionnaire.

Once your full submission is complete, your paper will undergo a series of checks in preparation for peer review. After your manuscript has passed the checks it will be sent out for review. To provide the metadata for your submission, please Login to Editorial Manager (https://www.editorialmanager.com/pbiology) within two working days, i.e. by Nov 28 2025 11:59PM.

If your manuscript has been previously peer-reviewed at another journal, PLOS Biology is willing to work with those reviews in order to avoid re-starting the process. Submission of the previous reviews is entirely optional and our ability to use them effectively will depend on the willingness of the previous journal to confirm the content of the reports and share the reviewer identities. Please note that we reserve the right to invite additional reviewers if we consider that additional/independent reviewers are needed, although we aim to avoid this as far as possible. In our experience, working with previous reviews does save time.

If you would like us to consider previous reviewer reports, please edit your cover letter to let us know and include the name of the journal where the work was previously considered and the manuscript ID it was given. In addition, please upload a response to the reviews as a 'Prior Peer Review' file type, which should include the reports in full and a point-by-point reply detailing how you have or plan to address the reviewers' concerns.

During the process of completing your manuscript submission, you will be invited to opt-in to posting your pre-review manuscript as a bioRxiv preprint. Visit http://journals.plos.org/plosbiology/s/preprints for full details. If you consent to posting your current manuscript as a preprint, please upload a single Preprint PDF.

Feel free to email us at plosbiology@plos.org if you have any queries relating to your submission.

Kind regards,

Roli Roberts

Roland Roberts, PhD

Senior Editor

PLOS Biology

rroberts@plos.org

Decision Letter 1

Roland Roberts

19 Mar 2026

Dear Dr Krummel,

Thank you for your patience while your manuscript "Discovery Stack Pilot: Feasibility and Outcomes of a Scientist-Designed Peer Review Model Separating Quality and Impact" was peer-reviewed at PLOS Biology. It has now been evaluated by the PLOS Biology editors and by three independent reviewers.

You'll see that reviewer #1 raises significant concerns about the robustness of your study, while reviewers #2 and #3 are more favourably disposed, seeing a valuable contribution even in this relatively preliminary state. During post-review cross-commenting, the latter reviewers maintained this position (while recognising rev #1's points), and in balance we share their opinion that there is value in considering your paper further. Naturally you should attempt to address reviewer #1's most pressing concerns (e.g. re weak statistical power and selection bias), whether by new analyses or flagging limitations, and in particular to acknowledge/emphasise the very preliminary nature of your findings (and it may well be that during the past 2-3 months you have accrued more data which may be able to strengthen your conclusions...).

In light of the reviews, which you will find at the end of this email, we would like to invite you to revise the work to thoroughly address the reviewers' reports.

Given the extent of revision needed, we cannot make a decision about publication until we have seen the revised manuscript and your response to the reviewers' comments. Your revised manuscript is likely to be sent for further evaluation by all or a subset of the reviewers.

In addition to these revisions, you will need to complete some formatting changes, which you will receive in a follow up email. A member of our team will be in touch with a set of requests shortly.

We expect to receive your revised manuscript within 3 months. Please email us (plosbiology@plos.org) if you have any questions or concerns, or would like to request an extension.

At this stage, your manuscript remains formally under active consideration at our journal; please notify us by email if you do not intend to submit a revision so that we may withdraw it.

**IMPORTANT - SUBMITTING YOUR REVISION**

Your revisions should address the specific points made by each reviewer. Please submit the following files along with your revised manuscript:

1. A 'Response to Reviewers' file - this should detail your responses to the editorial requests, present a point-by-point response to all of the reviewers' comments, and indicate the changes made to the manuscript.

*NOTE: In your point-by-point response to the reviewers, please provide the full context of each review. Do not selectively quote paragraphs or sentences to reply to. The entire set of reviewer comments should be present in full and each specific point should be responded to individually, point by point.

You should also cite any additional relevant literature that has been published since the original submission and mention any additional citations in your response.

2. In addition to a clean copy of the manuscript, please also upload a 'track-changes' version of your manuscript that specifies the edits made. This should be uploaded as a "Revised Article with Changes Highlighted" file type.

*Re-submission Checklist*

When you are ready to resubmit your revised manuscript, please refer to this re-submission checklist: https://plos.io/Biology_Checklist

To submit a revised version of your manuscript, please go to https://www.editorialmanager.com/pbiology/ and log in as an Author. Click the link labelled 'Submissions Needing Revision' where you will find your submission record.

Please make sure to read the following important policies and guidelines while preparing your revision:

*Published Peer Review*

Please note while forming your response, if your article is accepted, you may have the opportunity to make the peer review history publicly available. The record will include editor decision letters (with reviews) and your responses to reviewer comments. If eligible, we will contact you to opt in or out. Please see here for more details:

https://blogs.plos.org/plos/2019/05/plos-journals-now-open-for-published-peer-review/

*PLOS Data Policy*

Please note that as a condition of publication PLOS' data policy (http://journals.plos.org/plosbiology/s/data-availability) requires that you make available all data used to draw the conclusions arrived at in your manuscript. If you have not already done so, you must include any data used in your manuscript either in appropriate repositories, within the body of the manuscript, or as supporting information (N.B. this includes any numerical values that were used to generate graphs, histograms etc.). For an example see here: http://www.plosbiology.org/article/info%3Adoi%2F10.1371%2Fjournal.pbio.1001908#s5

*Blot and Gel Data Policy*

We require the original, uncropped and minimally adjusted images supporting all blot and gel results reported in an article's figures or Supporting Information files. We will require these files before a manuscript can be accepted so please prepare them now, if you have not already uploaded them. Please carefully read our guidelines for how to prepare and upload this data: https://journals.plos.org/plosbiology/s/figures#loc-blot-and-gel-reporting-requirements

*Protocols deposition*

To enhance the reproducibility of your results, we recommend that if applicable you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Thank you again for your submission to our journal. We hope that our editorial process has been constructive thus far, and we welcome your feedback at any time. Please don't hesitate to contact us if you have any questions or comments.

Sincerely,

Roli Roberts

Roland Roberts, PhD

Senior Editor

PLOS Biology

rroberts@plos.org

------------------------------------

REVIEWERS' COMMENTS:

Reviewer #1:

[identifies himself as Ross Mounce]

Those who seek to critique the current academic publication system really ought to make sure that their work is robust, well designed, and well evidenced. In my opinion this manuscript falls very short in terms of scientific quality.

The experimental design of this manuscript is obviously flawed for many different reasons including:

A) insufficient sample size

B) underexplained highly subjective recruitment, making it prone to possible bias

C) duration: trying to analyse the experiment too early before all the results are known

D) lack of demographic transparency or stratification in the sample recruitment

Many of these experimental design issues could have been pointed out and corrected before any data was collected (and effort wasted?) if the researchers behind this study had submitted their experimental design as a registered report for peer review. I am thus also unsurprised to note that the authors appear to make no mention to or reference of registered reports in their manuscript. The authors might want to read a review article on registered reports: Chambers, C.D., Tzavella, L. The past, present and future of Registered Reports. Nat Hum Behav 6, 29-42 (2022). https://doi.org/10.1038/s41562-021-01193-7

I shall try to explain the perceived experimental design flaws in more detail below:

A) **Sample size**

The recruitment of merely 18 preprints, only 17 of which could be included in all analyses is bafflingly small and makes this research statistically underpowered assuming the effect sizes are small (e.g. if Cohen's d = 0.2). I can see that this is labelled as a "pilot" and thus can't be expected to be an industry-wide experiment but equally I'd ask what good is such a tiny and insufficiently described experiment? Did the authors do any kind of prospective power analysis before collecting data to see how many preprints they ought to recruit to get meaningful results?

The small overall sample size is further undermined by the fact that the sample was recruited in three different ways. On of the recruitment methods recruited just three preprints, so a sample size of just three for that subset. Does the recruitment method not have an impact? I think it does. What sample size for each of the recruitment methods would you need to prove me wrong?

What if I were to assert that two of the recruited preprints were highly atypical papers? This is why you need a larger N for the number of recruited preprints, to accommodate for any outliers that might be in the sample.

B) **Underexplained highly subjective recruitment**

Why were three different recruitment methods employed?

1)Editor emailed DSP-enrolled authors

2)Editor emailed BioRxiv authors

3)Selected from BioRxiv - authors did not enroll

In category 3 can we please have more detail on how and why three were "selected by the pilot's scientific advisory board and reviewed without author input". There were over 43,000 preprints posted to BioRxiv in 2024, why were the three selected arbitrarily chosen out of thousands of other possible choices?

I suggest a more rigorous experimental design would have been to use a random number generator to randomly select bioRxiv preprints for recruitment into this experiment. "Three additional bioRxiv manuscripts were selected by the scientific advisory board…" leaves this study open to accusations of cherry-picking and bias in how these three extra preprints were selected.

Another bias is in category 1 ("Editor emailed DSP-enrolled authors") this is textbook self-selection bias. These DSP-enrolled authors have already shown interest and engagement in the Discovery Stack Project and so these are included? Would it not be a stronger test of the discovery stack model to recruit authors/preprints who have not heard of or previously engaged with the discovery stack model? The inclusion of DSP-friendly or DSP-aware authors in the experiment invalidates some of the results such as on page 10:

"Moreover, 84.9% of authors and reviewers agreed that separating these dimensions led to more constructive and insightful evaluations than traditional reviews"

because of self-selection bias. You're sampling from people who like the idea and so it's unsurprising that these authors might agree with reviewers who have also agreed to test the discovery stack model. Homophily.

C) **Duration**

There's a remarkable degree of impatience displayed in this manuscript: "only six of 18 manuscripts had been accepted at the time of analysis". Why not simply wait a few more months before publishing an analysis so that the ultimate journal venue of publication can be analysed for >90% of the recruited preprints? Why rush to publish so prematurely? I think knowing the final publication venue for >90% of the recruited preprints would be valuable data to include in any serious analysis of this model.

D) ** lack of demographic transparency or stratification in the sample recruitment **

The 18 preprints that form the core of this work appear to have been anonymised with labels such as DSPM118 and DSPL178. Why the anonymisation? If these are preprints published online at BioRxiv then they are published and known. I don't see why they need to be anonymised. It decreases my trust in this research that I don't know the exact identity of the 18 preprints that were recruited.

As I think we all recognise, no two scientific preprints are the same - a preprint isn't a simple substitutable unit.

What are the demographic variations of each of these preprints?

i.e. How many authors are named on each preprint?

How many pages are each preprint?

How many different institutions are affiliated with each paper? E.g. (1-100+)

Which countries are the listed author institutions on each preprint?

You can't call the discovery stack model a success if it's only been piloted with US & UK based academics… There needs to be explicit sociopolitical stratification built into the experimental design of the testing.

Aside from the issues with the experiment design, I also take issue with the near total absence of the importance of software in the discussion of Quality.

The authors state: "Quality refers to the rigor and reproducibility of the data supporting a study's conclusions"

Is that it? Really? Data only?

What about the critical role of software? Software should really get a mention as being part and parcel of research quality. As has been said elsewhere "The future of research software is the future of research" [1] 7 out of 10 UK-based researchers agree with the statement that it's impossible to conduct research without software. [2]

[1] Chue Hong NP, Aragon S, Hettrick S, Jay C. The future of research software is the future of research. Patterns (N Y). 2025 Jul 11;6(7):101322. doi: 10.1016/j.patter.2025.101322. PMID: 40926964; PMCID: PMC12416084.

[2] Simon Hettrick (2014) https://www.software.ac.uk/blog/its-impossible-conduct-research-without-software-say-7-out-10-uk-researchers

I specifically also examined the Quality Assessment Checklist given in supplementary figure 4 - there was no mention of software or code here either. Reading the Quality Assessment Checklist I inferred that it was devised and/or written by wet lab biologists: "identify experiments not repeated a standard of three times" being quite an obvious tell in this regard. If there is a desire to expand the discovery stack model to further disciplines then specific items like this will need to be revised...

I note that the word "code" features in Supplementary Figure 3 panel B and panel C and a throwaway mention on page 21 of the manuscript as something that was actually suggested by survey participants (well done to those participants that mentioned it!): "Other suggestions included easier access to raw data, code, and protocols; moderation of comments and vetting of users and reviewers; and an initial quality screen before review."

The discovery stack model absolutely must be amended to include data and software/code (with parity) within the consideration of research Quality. Examining the quality of research with no regard to its underlying software/code is madness.

Poor code can completely invalidate research and thus code must be inspected for quality. See e.g. the famous Reinhart and Rogoff debacle in economics

Herndon, Thomas & Ash, Michael & Pollin, Robert. (2013). Does High Public Debt Consistently Stifle Economic Growth? A Critique of Reinhart and Rogoff. Cambridge Journal of Economics. 38. 257-279. 10.1093/cje/bet075

The unwise disregard of code is evident within these authors own manuscript. I note with displeasure that on page 33 of the manuscript being reviewed the authors write:

"Bootstrap resampling was performed using custom Python scripts to assess the SD, range, and IQR of Quality and Impact scores"

Where are these Python scripts? This reviewer would very much like to see these Python scripts before being able to decide on the quality of this manuscript. It would be good if the authors could share the code behind their analyses at a trusted data repository such as Zenodo or Dryad.

*For the sake of reproducibility, what version of the GraphPad Prism software did the authors use and what version of NumPy?

*Where is there a copy of the exact survey instrument deposited, detailing all the questions and possible answers that were distributed to the 93 individuals (page 33). This does not seem to be contained within any of the supplementary figures.

*Curiously the author feedback survey is said to have been completed by 11 of 18 authors (61%). Preprints in the fields of Immunology and Cancer Biology do not tend to be authored by just a single person. Surely the feedback survey should poll ALL the authors of each preprint in order to get the views of all the authors? I would thus expect the total number of authors to be something like more than 100 not just 18.

*I note the manuscript only cites 19 other sources. Considering the topic I find it a little bit insufficient. More comparison and acknowledgement needs to be made to both registered reports and the publish, review, curate (PRC) model as relevant alternatives (that share important similarities) to the discovery stack model. Aczel et al (2021) should be cited to highlight the financial and time cost of current peer review systems.

Aczel, B., Szaszi, B. & Holcombe, A.O. A billion-dollar donation: estimating the cost of researchers' time spent on peer review. Res Integr Peer Rev 6, 14 (2021). https://doi.org/10.1186/s41073-021-00118-2

*Lack of discussion of the relevance or not of Journal Impact Factor ™ numbers created and sold by Clarivate. The authors describe Journal Impact Factor numbers as merely "flawed". I think this is somewhat of an understatement.

See e.g.

Curry, S. (2012, August 13). Sick of Impact Factors. Reciprocal Space. https://doi.org/10.59350/xqrv5-7bv94

Brembs B, Button K, Munafò M. Deep impact: unintended consequences of journal rank. Front Hum Neurosci. 2013 Jun 24;7:291. doi: 10.3389/fnhum.2013.00291. PMID: 23805088; PMCID: PMC3690355.

Dimitrov JD, Kaveri SV, Bayry J. Metrics: journal's impact factor skewed by a single paper. Nature. 2010 Jul 8;466(7303):179. doi: 10.1038/466179b. PMID: 20613817.

*The lack of significant correlation of (submitted) Journal Impact Factor ™ numbers with the Quality assessment is a fascinating result and should be discussed with reference to Brembs work:

Brembs B. Prestigious Science Journals Struggle to Reach Even Average Reliability. Front Hum Neurosci. 2018 Feb 20;12:37. doi: 10.3389/fnhum.2018.00037. Erratum in: Front Hum Neurosci. 2018 Oct 09;12:376. doi: 10.3389/fnhum.2018.00376. PMID: 29515380; PMCID: PMC5826185.

I am delighted that the authors get a result yet again confirming that the research quality of individual papers has no significant correlation with (submitted) Journal Impact Factor (Figure 2E). That is a key finding that more scientists and policymakers should take on board. Although the low sample size again is a bit unfortunate… Will this hold up with a larger sample size?

Reviewer #2:

[identifies himself as Stephen Curry]

This is a very interesting manuscript reporting the results of a pilot study of a peer review model that separates the assessment of quality and impact and enables the used of in-line reviewer comments. Although the scale of the pilot is relatively small (18 manuscripts assessed; 162 quality and/or impact reviews obtained) and focused on just two biomedical subfields (cancer biology and immunology), it has been performed and reported very thoroughly - albeit a little too thoroughly in some repects, given the small numbers in the surveys that make up a large part of the analysis. That said, overall, it makes for a valuable contribution to the literature on peer review which, event at this late stage, remains something of a black box.

There are a number of weaknesses in the study that I think should be addressed before publication.

The authors do a good job of showing that reviewers are able to meaningfully distinguish quality from impact (p8-9), but I wanted to see a clearer analysis of the claimed support for standardized metrics (p10). Although the methods for arriving at quantitative scores for quality and impact are explained clearly in the Methods section (i.e. using separate questionnaires with Likert-Scale questions), I think a brief outline of the assessment process should be presented alongside the results. It strikes me that the reported support for standardized metrics could well also arise from the fact that the questionnaires provide clearly defined criteria for the quality and impact assessments and that it is this clarity of criteria (still very rare in normal practice) that appealed to authors, alongside the scoring mechanism for arriving at the metrics. In fact, I think the use of defined criteria is such an important feature of the review protocol at the heart of the study that Supplementary Figures 6 and 7 should be promoted to the main body of the manuscript, rather than being relegated to the supplementary information, where they will more easily be missed.

Given the very small number of authors surveyed in the study, I think the manuscript authors need to be much more circumspect about concluding that "Identify disclosure enhances perceived transparency" (p11-12). They also place emphasis on the observation that 4/10 authors "stated that identified reviews were more constructive than anonymous reviews", seemingly overlooking their finding that 6/10 authors found there was no difference or that reviewers were less constructive. I think the finding here is that the results are inconclusive! In any case it should be emphasised that findings based on such small numbers must be regarded as extremely preliminary.

In many places the survey results are quoted to 3 significant figures. Given the small number of participants involved (<100), I think that a maximum of 2 significant figures would be more appropriate. This would also have the advantage of making the manuscript more readable.

On p13-14 the section titled "Discovery Stack Platform Delivers a Better Experience than Traditional Review" discusses authors' and reviewers' perceptions of their new platform. When discussing the authors' responses, I think it would be worth considering the fact that the review process did not produce an accept/reject decision and discussing whether that might colour author perceptions.

At the top of p14 it is stated that "Most participants (75%) agreed that the platform fostered this mindset more effectively than traditional reviews". However, I would question whether the platform is solely responsible for this perception. In the Methods it is revealed that the briefing given to reviewers "emphasized the value of a peer-improvement mindset". At the very least, this instruction would appear to be a confounder of the analysis. The authors should remind the reader in presenting the results that this instruction was given and comment on its likely impact.

On p18-19 there is a section titled "Expedited Review Is Achievable with Reviewer Accountability and Backup Plans". The purpose of this section is not clear to me. The authors report review times for quality and impact assessments, but then claim that their pilot cannot be compared directly to traditional review. They then proceed to discuss the impact of single late reviews and the need for contingency plans. In the absence of an experiment designed to directly compare traditional review with the Discovery Stack model, I don't see that this section adds any great value to the manuscript.

The final two sections of the results (p20-21) probe the surveys for their interest in new publishing models and their views on desirable features. Again, given the small sample size, I don't think too much can be made of these results. Moreover, the answers to questions about using 'scientist-developed metrics to assess quality and impact' must be difficult to interpret if the form of the metrics is not discussed.

We are told also in the last section (p21) that 71.8% of respondents would prioritise "evolving impact metrics based on ongoing community input". Maybe so - and I agree that impact is very difficult to assess a priori - but how would such "community input" be gathered? In the absence of any realistic prospect for a mechanism, such aspirations come across as somewhat naïve. I think here and in the discussion the authors should discuss the practical obstacles to "dynamic impact metrics".

The Discussion section also raises the possibility that in time "a threshold for high-Quality papers may emerge" from the use of standardized metrics. Although the authors are careful to acknowledge the risks associated with metrics, they don't outline a process for re-scoring quality on their platform. As presented, it produces a quality score for the submitted, not the revised manuscript. Re-scoring would add a further burden on reviewers, as would any attempt to keep track of impact months or years after publication. In any case, does this add substantial value above and beyond the practice of open review (in which reviews are published)? This should be discussed.

The discussion of the "broad dissatisfaction with current publishing norms" seems to me to add nothing new to our understanding of present woes and could be truncated. Happy to be corrected if I have missed something.

Overall, I have the impression that the numerous and detailed analyses presented in this paper end up muddying or masking what I think are its most important messages. These are the demonstrable value of a review protocol that separates quality from impact, AND assesses these using multiple, clearly defined criteria scored in a standardised way. The second important innovation is the utility of introducing in-line reviewer comments. I think some of the weaker or less important results sections could be truncated (or relegated to the supplementary information) and this would allow the central findings to come through more strongly.

It would also help to have a stronger focus on the practicalities, given the well-known pressures on the peer review system. A protocol that needs 2-3 quality reviewers and 5-6 impact reviewers doesn't seem likely to take off. Would it be more realistic to have fewer reviewers, who both score quality and impact (using the different criteria) and down-weight the impact score given the greater uncertainty associated with that assessment? A sharper consideration of the practicalities would in my view enhance the quality of the present manuscript and augment the probability that it will have a significant impact!

Reviewer #3:

[identifies himself as Ludo Waltman]

My review report is available online: https://prereview.org/reviews/18915960.

Ludo Waltman

March 8, 2026

FLAT TEXT VERSION OF THAT REVIEW:

This paper presents the outcomes of a pilot study of a new approach to peer review of research articles. In this new approach, the quality and impact of an article are evaluated separately. I found the paper highly interesting to read. As pointed out by the authors, established approaches to peer review are facing major challenges. There is a strong need for studying alternative approaches to peer review.

My comments on the paper are fairly minor.

My most important comment is that the idea of separating the evaluation of quality and impact is not new. The authors should acknowledge existing approaches based on this idea. Let me highlight two approaches.

First, an increasing number of journals are using a so-called soundness-only approach to peer review. These journals ask reviewers to evaluate the soundness (i.e., quality) of an article and to refrain from evaluating the impact of an article. Soundness-only peer review was pioneered by PLOS One. Nowadays it is also used by various other journals, such as Scientific Reports. There is also literature studying soundness only peer review. See https://doi.org/10.1108/JD-06-2017-0092.

Second, the journal eLife uses a publish-review-curate model in which an explicit distinction is made between the strength of evidence (i.e., quality) and the significance of the findings (i.e., impact). Strength of evidence and significance of the findings are each assessed on a five- or six-point scale. There seem to be close connections between the eLife model and the Discovery Stack model proposed by the authors. The authors should discuss these connections.

A minor comment relates to the way in which the authors use the word ‘published’. The authors for instance state: “At the time of analysis, only six of 18 manuscripts had been published. The mean from submission to publication for the published manuscripts was 231 days.” This is confusing. Articles can be published (i.e., be made publicly available) not only in journals but also on other platforms, such as preprint servers. In this case, all 18 manuscripts have already been published on a preprint server. Therefore, instead of “only six of 18 manuscripts had been published”, the authors should say “only six of 18 manuscripts had been published in a journal”. Likewise, “from submission to publication” should be “from submission to publication in a journal”.

Finally, I was unable to find the data underlying the analyses presented in the paper. In the interest of transparency and reproducibility, publishing the data is crucial.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they did not use generative AI to come up with new ideas for their review.

Decision Letter 2

Roland Roberts

3 Aug 2026

Dear Max,

Thank you for your patience while we considered your revised manuscript "Discovery Stack Pilot: Feasibility and Outcomes of a Scientist-Designed Peer Review Model Separating Quality and Impact" for publication as a Meta-Research Article at PLOS Biology. This revised version of your manuscript has been evaluated by the PLOS Biology editors, the Academic Editor and the original reviewers.

Based on the reviews, we are likely to accept this manuscript for publication, provided you satisfactorily address the remaining points raised by the reviewers, and the following data and other policy-related requests.

IMPORTANT - please attend to the following:

a) We need your Title to contain an active verb and to avoid a colon (these tend to result in title truncation by some aggregators). We suggest changing it to: "A pilot study reveals the feasibility and outcomes of Discovery Stack, a scientist-designed peer peview model that separates quality and impact"

b) Please attend to the remaining comments from the reviewers.

c) Please confirm that this process, which involved human subjects in a trial peer review process, did not require ethical approval.

d) We note that you specify one author as receiving "specific" support for this project. Please also declare any non-specific support which was used for this project.

e) Please address my Data Policy requests below; specifically, we need you to supply the numerical values underlying Figs 1BC, 4ABCDEFGHIJ, 5ABCD, 6ABCD, 7ABC, 8ABCD, 9ABCDEFGHIJ, 10AB, S1CDEF, S2AB, S3ABCDEFGHIJKLM, S4ABCDEFGH, S5ABC, S7AB, either as a supplementary data file or as a permanent DOI’d deposition.

f) Please cite the location of the data clearly in all relevant main and supplementary Figure legends, e.g. “The data underlying this Figure can be found in S1 Data” or “The data underlying this Figure can be found in https://zenodo.org/records/XXXXXXXX; ; also update your Data Availability Statement accordingly.

g) Please make any custom code available, either as a supplementary file or as part of your data deposition.

h) Please include the URLs of your funders in the Financial Disclosure statement.

As you address these items, please take this last chance to review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the cover letter that accompanies your revised manuscript.

In addition to these revisions, you may need to complete some formatting changes, which you will receive in a follow up email. A member of our team will be in touch with a set of requests shortly. If you do not receive a separate email within a few days, please assume that checks have been completed, and no additional changes are required.

We expect to receive your revised manuscript within two weeks.

To submit your revision, please go to https://www.editorialmanager.com/pbiology/ and log in as an Author. Click the link labelled 'Submissions Needing Revision' to find your submission record. Your revised submission must include the following:

- a cover letter that should detail your responses to any editorial requests, if applicable, and whether changes have been made to the reference list

- a Response to Reviewers file that provides a detailed response to the reviewers' comments (if applicable, if not applicable please do not delete your existing 'Response to Reviewers' file.)

- a track-changes file indicating any changes that you have made to the manuscript.

NOTE: If Supporting Information files are included with your article, note that these are not copyedited and will be published as they are submitted. Please ensure that these files are legible and of high quality (at least 300 dpi) in an easily accessible file format. For this reason, please be aware that any references listed in an SI file will not be indexed. For more information, see our Supporting Information guidelines:

https://journals.plos.org/plosbiology/s/supporting-information

*Published Peer Review History*

Please note that you may have the opportunity to make the peer review history publicly available. The record will include editor decision letters (with reviews) and your responses to reviewer comments. If eligible, we will contact you to opt in or out. Please see here for more details:

https://plos.org/published-peer-review-history/

*Press*

Should you, your institution's press office or the journal office choose to press release your paper, please ensure you have opted out of Early Article Posting on the submission form. We ask that you notify us as soon as possible if you or your institution is planning to press release the article.

*Protocols deposition*

To enhance the reproducibility of your results, we recommend that if applicable you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Please do not hesitate to contact me should you have any questions.

Sincerely,

Roli

Roland Roberts, PhD

Senior Editor

rroberts@plos.org

PLOS Biology

------------------------------------------------------------------------

ETHICS STATEMENT:

-- Please include information about the form of consent (written/oral) given for research involving human participants. All research involving human participants must have been approved by the authors' Institutional Review Board (IRB) or an equivalent committee, and must have been conducted according to the principles expressed in the Declaration of Helsinki.

------------------------------------------------------------------------

DATA POLICY:

You may be aware of the PLOS Data Policy, which requires that all data be made available without restriction: http://journals.plos.org/plosbiology/s/data-availability. For more information, please also see this editorial: http://dx.doi.org/10.1371/journal.pbio.1001797

Note that we do not require all raw data. Rather, we ask that all individual quantitative observations that underlie the data summarized in the figures and results of your paper be made available in one of the following forms:

1) Supplementary files (e.g., excel). Please ensure that all data files are uploaded as 'Supporting Information' and are invariably referred to (in the manuscript, figure legends, and the Description field when uploading your files) using the following format verbatim: S1 Data, S2 Data, etc. Multiple panels of a single or even several figures can be included as multiple sheets in one excel file that is saved using exactly the following convention: S1_Data.xlsx (using an underscore).

2) Deposition in a publicly available repository. Please also provide the accession code or a reviewer link so that we may view your data before publication.

Regardless of the method selected, please ensure that you provide the individual numerical values that underlie the summary data displayed in the following figure panels as they are essential for readers to assess your analysis and to reproduce it: Figs 1BC, 4ABCDEFGHIJ, 5ABCD, 6ABCD, 7ABC, 8ABCD, 9ABCDEFGHIJ, 10AB, S1CDEF, S2AB, S3ABCDEFGHIJKLM, S4ABCDEFGH, S5ABC, S7AB. NOTE: the numerical data provided should include all replicates AND the way in which the plotted mean and errors were derived (it should not present only the mean/average values).

IMPORTANT: Please also ensure that figure legends in your manuscript include information on where the underlying data can be found, and ensure your supplemental data file/s has a legend.

Please ensure that your Data Statement in the submission system accurately describes where your data can be found.

------------------------------------------------------------------------

CODE POLICY

Per journal policy, if you have generated any custom code during the course of this investigation, please make it available without restrictions. Please ensure that the code is sufficiently well documented and reusable, and that your Data Statement in the Editorial Manager submission system accurately describes where your code can be found. More information on our Code Policy, what and how to share can be found here: https://journals.plos.org/plosbiology/s/code-availability

Please note that we cannot accept sole deposition of code in GitHub, as this could be changed after publication. However, you can archive this version of your publicly available GitHub code to Zenodo. Once you do this, it will generate a DOI number, which you will need to provide in the Data Accessibility Statement (you are welcome to also provide the GitHub access information). See the process for doing this here: https://docs.github.com/en/repositories/archiving-a-github-repository/referencing-and-citing-content

------------------------------------------------------------------------

DATA NOT SHOWN?

- Please note that per journal policy, we do not allow the mention of "data not shown", "personal communication", "manuscript in preparation" or other references to data that is not publicly available or contained within this manuscript. Please either remove mention of these data or provide figures presenting the results and the data underlying the figure(s).

------------------------------------------------------------------------

REVIEWERS' COMMENTS:

Reviewer #1:

[identifies himself as Ross Mounce]

Thank you for the extensive and thoughtful revision. I appreciate the care with which you have engaged with the reviewers' comments and the substantial effort invested in clarifying the study design, expanding the discussion, improving transparency, and making the underlying data and code available.

Having read both the original review reports and your detailed responses, I believe the manuscript has improved considerably. In particular, the additions regarding recruitment, the explicit framing of the work as a feasibility pilot, the new analyses addressing familiarity bias, the provision of code and survey instruments, and the expanded discussion of related initiatives (e.g., soundness-only review, publish-review-curate approaches, and Registered Reports) all strengthen the manuscript and address many of the concerns raised previously.

That said, I remain somewhat cautious about the strength of several conclusions and would encourage further moderation of the manuscript's framing in a few areas.

**1. Pilot framing and generalizability | Overclaiming**

The revised manuscript does a much better job of describing itself as a feasibility study, and I commend the authors for explicitly acknowledging recruitment challenges, selection effects, and the limited size of the cohort. However, I still feel that some passages occasionally drift from "evidence of feasibility" toward broader claims regarding the supposed superiority or future success of the model.

The central contribution of this paper is not that the Discovery Stack model has been proven superior to existing peer review systems. Rather, it is that the authors have demonstrated the feasibility of recruiting 18 independent co-authors of preprints into implementing a new review process that separates assessments of Quality and Impact, incorporates structured criteria, and uses in-line commenting. The manuscript is strongest when it stays focused on these achievements.

For instance:

*the final sentence of the Discussion, the claim that the system "is ready for expansion and adoption" goes beyond what a pilot of 18 manuscripts can demonstrate. The study shows feasibility, not readiness for broad deployment.

*Quote from the revised manuscript: "In future implementations of the platform, a threshold for high-Quality papers may emerge, above which manuscripts would be designated as rigorous and trustworthy."

oThis is an interesting hypothesis, but it implies a level of validation that has not yet occurred. The pilot demonstrates that reviewers can generate Quality scores; it does not yet establish that those scores are sufficiently calibrated, reproducible, or predictive to support a threshold denoting trustworthy science.

*Quote from new text in the revised manuscript: "More examples would need to be produced but we consider it likely that a system like the one we tested would in the long run be at least as predictive of the long-term citation/value of a given piece of science as the journal impact factor."

oThis is tremendously speculative about the Discovery Stack model. I think it also manages to misrepresent the science of studies about Journal Impact Factor and its relation to the citedness of individual papers. Journal Impact Factor is a poor predictor of the citedness of individual papers within a journal. As shown by: Seglen, P.O. (1997). "Why the impact factor of journals should not be used for evaluating research." BMJ 314: 498-502. Zhang, L., Rousseau, R., & Sivertsen, G. (2017). "Science deserves to be judged by its contents, not by its wrapping." PLOS ONE 12:e0174205. Larivière, V. et al. (2016). "A simple proposal for the publication of journal citation distributions." and others...

**2. Recruitment bias remains an important limitation**

I appreciate the additional analyses comparing previously engaged participants and newly recruited participants. The finding that responses were broadly similar is useful and reassuring. Nevertheless, I do not think these analyses entirely eliminate concerns about participation bias.

The manuscript now documents that reviewer acceptance rates were dramatically higher among individuals already familiar with the project or contacted through existing networks. Indeed, the study reveals that personal outreach and prior engagement were the dominant determinants of participation.

To me, this is an interesting finding. It suggests that enthusiasm for alternative peer-review systems may not yet extend broadly across the wider research community and that adoption barriers remain substantial.

In the response to the reviewer comments, the authors disclose that only a single-author of each multi-author enrolled manuscript responded to the surveys. This masks variation in the role of each author in multi-author teams and how that maps to responses each author might give to the survey questions. What if the survey-responding author was a first-author, what if the survey-responding author was the PI on the grant and/or the last-author, what if the survey-responding author was a middle author amongst twenty other middle authors, who did the some of the pipetting work (minor involvement). It seems odd to treat all authors on a preprint as "the same", when particularly in terms of power, influence, and positionality towards change in scholarly communication, there are very important differences between researchers in different stages of their careers. This was not controlled for in the experimental design. Unfortunate.

**3. dynamic Impact metrics**

The revised discussion provides a much clearer explanation of how dynamic Impact metrics might work in practice. I found this addition helpful.

That said, implementation remains speculative. Citation counts, altmetrics, expert evaluations, downloads, policy citations, and social engagement all have known limitations and biases. I would encourage the authors to make clear that the proposed framework remains aspirational and untested rather than implying that an accepted methodology already exists.

### Overall assessment

My overall view is considerably more positive than in the previous round.

While I remain unconvinced that this study provides strong evidence for the effectiveness of the Discovery Stack model at scale, I do believe it now provides useful evidence that 18 anonymised preprints have undergone a Discovery Stack paid-review process. The manuscript is more transparent-ish, more balanced, better contextualized within the broader literature, and substantially more reproducible than the version originally reviewed.

However, one should not overlook the impact of paying peer-reviewers on the assessment of the model feasibility ($30 for each Quality review, $20 for each Impact review). Was it only feasible because you were offering direct financial payment? Would this level of payment be feasible if scaled up? Who exactly would provide the funding to pay for the payments to peer-reviewers and would they be happy to pay these extra fees, at a time when the United States Office of Management and Budget is proposing to stop paying from federal funds for any publication-related costs?[1] I'm not at all surprised that many happy volunteers were found to pilot this model given that payment was offered. Ker-ching! People like to earn more money. The understandable desire from an individual to earn more money may perhaps mask actual enthusiasm for the scholarly benefit(s) or disbenefit(s) of the model?

[1] https://sparcopen.org/our-work/2026-proposed-2cfr200-updates-faqs/ & https://www.federalregister.gov/documents/2026/05/29/2026-10817/regulation-for-federal-financial-assistance

Sidenotes:

A stylistic point. I would advise against lower-casing mentions of Clarivate's proprietary Journal Impact Factor ™ (JIF) numbers. Lower-casing gives the false impression that JIF numbers are somehow transparent, reproducible, or sensibly calculated. They are not. They are proprietary, negotiable, statistically-illiterate numbers which only Clarivate has the right to assign to journals as their official (Clarivate-awarded) Journal Impact Factor™ . References throughout to 'journal impact factor' or, even worse 'impact factor' [failing to mention the word journal, which neglects to explicitly inform that it is a journal-level number] may mislead readers into thinking that these are generic numbers created following some kind of rational and statistically-sound method - they are NOT!

Prior examples of using in-line commenting for peer-review of preprints include one by myself in 2021 https://greenelab.github.io/annorxiver_manuscript/ (see the Hypothesis logo in the top-right corner). This was journal-organized peer-review. The journal which organized it was PLOS Biology (!). The journal eventually published it here: https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3001470

These in-line Hypothes.is comments were incorporated into and extensively referenced in the submitted first-round peer review report https://journals.plos.org/plosbiology/article/peerReview?id=10.1371/journal.pbio.3001470

I'm sure there are many other examples, either more systematically implemented at a journal, or at an earlier date than 2021, such as at the American Geophysical Union (Hypothes.is and the Geophysical Electronic Manuscript Submission [GEMS] in editorial review) and JMIR Publications (Hypothes.is and preprints). The point is that in-line commenting for peer-review of preprints during journal-organized peer review is not completely new to the Discovery Stack model. I just wanted to make that clear.

I re-ran the code supplied at github to check the reproducibility of the work. Yes, every single line. It is well-documented and the results generated match what is in the manuscript. It was interesting to see from this code-level review that the sample size in the Quality vs Impact Dispersion analyses (figures 5A to 5D) is on just 16 manuscripts, not 18. This is somewhat disclosed in the figure legend text for figure 5, although it would be nice to make it more consistent with the other figure legends to spell it out with "n=16".

Reviewer #2:

[identifies himself as Stephen Curry]

I am satisfied by the way that the authors have addressed the comments I raised in relation to their original submission.

One change to the revised manuscript does not pass muster. The claim on page 21 that 'a system like the one we tested would in the long run be at least as predictive of the long-term citation/value of a given piece of science as the journal impact factor' is not reasonable because the JIF - an average for the journal - is not a good predictor of scientific value of a the individual articles therein. This claim should be deleted since it adds nothing of merit.

Reviewer #3:

[identifies himself as Ludo Waltman]

I consider the revised paper to be suitable for publication in the journal.

I still have a few minor comments:

"In fact, growing evidence suggests that the quality of published research inversely correlates with journal impact factor, which is evidenced by the higher retraction rates in journals with high impact factors (11, 14).": I wonder how strong this evidence is. Couldn't the inverse correlation be the result of articles in high impact factor journals receiving more attention than articles in low impact factor journals? This then increases the probability of the detection of errors in articles in high impact factor journals and, consequently, the probability of retraction of these articles.

"The lack of significant correlation between Quality scores and journal impact factor is consistent with evidence that journal rank is a poor surrogate for scientific reliability (11).": I see the lack of a statistically significant correlation as a reflection of the small sample size. For a large enough sample size, it seems likely that the correlation will be statistically significant (but not necessarily substantively significant).

"Journal impact factors are conceptually and statistically flawed (9-14).": Impact factors are causing lots of problems, but I don't agree they are statistically flawed. In my view, commonly used arguments about impact factors being statistically flawed are themselves statistically flawed. See https://doi.org/10.12688/f1000research.23418.2.

Finally, I would like to note that I was slightly disappointed that the authors have not uploaded their revised article to bioRxiv. The version of the article on bioRxiv is outdated. As a result, I am unable to publish this review report online. I hope the authors will still update their article on bioRxiv.

Ludo Waltman

July 30, 2026

Decision Letter 3

Roland Roberts

20 Aug 2026

Dear Max,

Thank you for the submission of your revised Meta-Research Article "Discovery Stack Pilot demonstrates the feasibility and outcomes of a scientist-designed peer-review model that separates quality and impact" for publication in PLOS Biology. On behalf of my colleagues and the Academic Editor, Stephen Curry, I'm pleased to say that we can in principle accept your manuscript for publication, provided you address any remaining formatting and reporting issues. These will be detailed in an email you should receive within 2-3 business days from our colleagues in the journal operations team; no action is required from you until then. Please note that we will not be able to formally accept your manuscript and schedule it for publication until you have completed any requested changes.

IMPORTANT:

a) You'll see that I've taken the liberty of changing your Title into sentence case, which is journal style; I did this to ensure that "Discovery Stack Pilot" will be capitalised correctly by our Production team.

b) I also took the liberty of including in your "competing interests statement" the fact that your co-author Ken Cadwell is a member of our editorial board. Your statement now reads, “I have read the journal’s policy and the authors of this manuscript have the following competing interests: KC is a member of PLOS Biology’s Editorial Board. The other authors declare that no competing interests exist."

c) I've asked my colleagues to include the following further request alongside their own: "Many thanks for depositing your code in Github. However, because Github depositions can be readily changed or deleted, please make a permanent DOI’d copy (e.g. in Zenodo) and provide this URL."

Please take a minute to log into Editorial Manager at http://www.editorialmanager.com/pbiology/, click the "Update My Information" link at the top of the page, and update your user information to ensure an efficient production process.

PRESS: We frequently collaborate with press offices. If your institution or institutions have a press office, please notify them about your upcoming paper at this point, to enable them to help maximise its impact. If the press office is planning to promote your findings, we would be grateful if they could coordinate with biologypress@plos.org. If you have previously opted in to the early version process, we ask that you notify us immediately of any press plans so that we may opt out on your behalf.

We also ask that you take this opportunity to read our Embargo Policy regarding the discussion, promotion and media coverage of work that is yet to be published by PLOS. As your manuscript is not yet published, it is bound by the conditions of our Embargo Policy. Please be aware that this policy is in place both to ensure that any press coverage of your article is fully substantiated and to provide a direct link between such coverage and the published work. For full details of our Embargo Policy, please visit http://www.plos.org/about/media-inquiries/embargo-policy/.

Thank you again for choosing PLOS Biology for publication and supporting Open Access publishing. We look forward to publishing your study.

Sincerely,

Roli

Roland G Roberts, PhD, PhD

Senior Editor

PLOS Biology

rroberts@plos.org

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Fig. Calculation and comparison of Quality and Impact composite scores.

    Composite Quality and Impact review scores were calculated to condense the full set of ratings from each review into a single metric. (A) Composite Quality scores were derived from 13 Likert-scale questions that were weighted based on relative importance. Composite Quality score = ∑ (Question Score x Importance weight)/ ∑ (Importance weight). (B) Composite Impact scores were based on responses to one multiple-choice question and five Likert-scale items. To integrate responses across two question types, each feature in the multiple-choice question was assigned a weight reflecting its relative importance to Impact, the weighted sum of selected features was calculated and then normalized to a 1−5 scale using a linear transformation. Transformed Q2 = 6 - (1 + ((∑Q2–1)*4/(36−1))). In this scale, selecting no features corresponds to a score of 5 (lowest Impact) and selecting features totaling the maximum score achieved (36) corresponds to a score of 1 (highest Impact). This transformed Q2 value was then combined with the weighted average of the five Likert-scale responses to generate a single composite Impact score for each review. (C) The weighted Quality composite score, unweighted average of Likert-scale responses, and the average score for the first statement (Overall, the data support the conclusions presented in the paper) were compared across manuscripts. (D) The weighted composite Impact score (including Q2), the unweighted average of Q2–Q7, and the unweighted average of Q3-Q7 were compared across manuscripts. (E, F) Weighted composite scores for (E) Quality or (F) Impact were compared between manuscripts enrolled as initial submissions versus revised versions. The version of three manuscripts (D165, D173, and D178) was unknown, so are excluded. No significant differences were observed by Mann–Whitney test. The data underlying these figures are available at https://doi.org/10.5281/zenodo.2004122.

    (TIFF)

    pbio.3003986.s001.tiff (1.3MB, tiff)
    S2 Fig. Alluvial plot of reviewer Quality and Impact scores by submitted JIF tier.

    Manuscripts are grouped on the left by the JIF of the journal to which they were (A) submitted or (B) published, binned into three tiers: High (JIF ≥ 42.5), Mid (JIF 15.7–27.6), and Low (JIF ≤ 9.1). Tier boundaries were defined by the actual JIFs. Within each tier, manuscripts were ordered from top to bottom first by JIF then combined Quality + Impact score. Each curve connects a manuscript's tier position on the left to its average reviewer score on the right. Reviewer scores range from 1 (best) to 5 (worst), with lower scores indicating stronger reviews. Because Quality and Impact scores span different ranges across the dataset, each metric is displayed on its own independently scaled right-hand axis, each beginning 0.1 units below the lowest recorded score for that metric. Flatter curves indicate greater agreement between JIF tier and reviewer score. The data underlying these figures are S1 Table and https://doi.org/10.5281/zenodo.20041225.

    (TIFF)

    pbio.3003986.s002.tiff (1.8MB, tiff)
    S3 Fig. Comparison of survey responses between newly recruited (NEW) and previously enrolled (DSP-enrolled) participants.

    (A) Points represent the difference in mean Likert scores between participants new to the study (NEW) and those previously enrolled (DSP-enrolled) for each survey question. Horizontal lines indicate 95% bootstrap confidence intervals. Lower scores correspond to more favorable responses. The vertical line at zero indicates no difference between groups. (B) Distribution of responses for each survey question among NEW and DSP- enrolled participants. For each question, responses from NEW participants are shown in the upper bar and DSP-enrolled responses are in the lower bar. Bars are centered on the neutral response category, with more favorable responses extending to the left and less favorable responses to the right. The data underlying these figures are available in S2 Table and https://doi.org/10.5281/zenodo.20041225.

    (TIFF)

    pbio.3003986.s003.tiff (455.7KB, tiff)
    S4 Fig. Identity disclosure may be associated with greater perceived transparency and higher Impact scores.

    (A) The percent of Quality (n = 50) or Impact (n = 114) reviews for which the reviewer's  identity was disclosed or remained anonymous. (B) The percent of PIs (n = 129) and trainees (n = 33) who disclosed their identity or remained anonymous. (C) The percent of identified reviewers per manuscript. (D) Reviewers who remained anonymous (n = 35) were asked a follow-up multiple-choice question regarding the reasons that influenced their choice. (E) Authors (n = 11) were asked if they perceived the Discovery Stack review process to be more transparent than traditional review. (F) Transparency ratings for authors that completed surveys (n = 11) were compared to percent of identified reviewers for each manuscript using Pearson’s correlation. (G) Authors (n = 10) and Impact-only reviewers (n = 18) were asked whether they noticed a difference in scoring between identified and anonymous reviewers (H) Authors (n = 10) were asked follow-up questions regarding their perception of feedback from identified and anonymous reviewers. (I, K) Composite (I) Quality and (K) Impact scores were compared between identified and anonymous reviewers across all manuscripts using the Mann-Whitney test (p = 0.0082** for Impact). (J, L) Within each manuscript, the average (J) Quality and (L) Impact scores from identified and anonymous reviewers were compared using the Wilcoxon matched-pairs signed rank test (p = 0.026* for Impact). Manuscripts with no identified reviewers (Quality n = 2, Impact n = 1) or no anonymous reviewers (Quality n = 3; Impact n = 3) were excluded. (M) Composite Quality and Impact scores were compared between PIs (Quality n = 37; Impact n = 92) and trainees (Quality n = 12; Impact n = 21). Differences were tested using the Mann–Whitney test. The data underlying these figures are available at https://doi.org/10.5281/zenodo.20041225.

    (TIFF)

    pbio.3003986.s004.tiff (1.1MB, tiff)
    S5 Fig. Reviewer engagement depends on outreach, networks, and trainee involvement.

    (A) The positive response rate was compared between individuals who were already familiar with the study (DSP-enrolled), new to the study (New to DSP), or new but contacted by a known colleague on the Scientific Advisory Board (New to DSP + email) using two-sided Fisher’s exact tests (p < 0.0001). Odds ratios (ORs) and 95% CIs were calculated from the contingency table. P-values were adjusted for multiple pairwise comparisons using the Bonferroni correction. Error bars represent 95% CIs. DSP-enrolled versus New to DSP (OR=15.2, 95% CI: 7.1–32.7, p < 0.0001, adjusted), New to DSP + email versus New to DSP (OR=12.2, 95% CI: 5.8–25.7, p < 0.0001, adjusted), DSP-enrolled versus New to DSP + email significant (OR=1.25, 95% CI: 0.64–2.4, p = 0.61, adjusted). (B) Number of reviewer invitations per manuscript to recruit Quality/Impact and Impact-only reviewers. (C) The percent of individuals that responded “yes”, “no”, or did not respond to reviewer invitations per manuscript. The overall acceptance was 33%. (D) Survey respondents indicated whether they were unfamiliar with the study, previously enrolled, or recruited by a colleague. (E) Percent of Quality (n = 48) or Impact (n = 110) reviewers at each career stage. Among Quality reviewers, 73% were Principal Investigators (PIs) while 80% of Impact reviewers were PIs. Only independent trainee reviewers (i.e., not co-reviewing with a PI) were included in these percentages. Other includes industry positions. (F) A multivariable logistic regression model estimated the odds of reviewer acceptance as a function of position (PI versus trainee) and participant type (DSP-enrolled, New to DSP, New to DSP + email) as predictors. Only independent trainee reviewers, not co-reviewers, were included in the analysis. OR and 95% CI were obtained by exponentiating the logistic regression coefficients. Statistical significance was assessed using Wald tests (OR = 15.3, 95% CI: 4.2–56.2, p < 0.0001). (G) Survey participants reported their review experience. (H) The number of years of peer review experience reported by participants. The data underlying these figures are available at https://doi.org/10.5281/zenodo.20041225.

    (TIFF)

    pbio.3003986.s005.tiff (804.8KB, tiff)
    S6 Fig. Top priorities for improved scientific publishing.

    (A) Participants (n = 86) were asked what they believe needs to change about the current scientific publishing and peer review systems. A total of 65 open-ended responses were collected, categorized into key thematic areas, and tallied. The number of respondents mentioning each concern is shown. (B) Participants were also asked to imagine that scientific publishing didn’t exist and to identify three core principles they would prioritize if building a system from scratch. A total of 53 open-ended responses were collected, categorized by theme, and tallied. (C) Responses across both questions were averaged and ranked to highlight the most broadly recognized priorities for reform. The data underlying these figures are available at https://doi.org/10.5281/zenodo.20041225.

    (TIFF)

    pbio.3003986.s006.tiff (1.3MB, tiff)
    S7 Fig. Quality assessment checklist.

    The Quality Assessment checklist was given to Quality/Impact reviewers to highlight key elements to consider when evaluating the Quality of a manuscript. It offers a comprehensive framework for reviewing conclusions, assessing experimental design, ensuring data integrity, and evaluating clarity and scholarly analysis. This guide contains resources adapted from [30] https://doi.org/10.5281/zenodo.5484087.

    (TIFF)

    pbio.3003986.s007.tiff (1.3MB, tiff)
    S1 Table. Composite reviewer scores, JIFs, and journal review timelines per manuscript.

    Individual reviewer composite scores and average composite scores are shown for both Quality and Impact evaluations for each manuscript. Lower scores are more favorable evaluations. The table also shows the JIF of the journal to which each manuscript was initially submitted and ultimately published (when applicable), as well as the duration of journal review for manuscripts published in journals.

    (XLSX)

    pbio.3003986.s008.xlsx (12KB, xlsx)
    S2 Table. Comparison of survey responses between newly recruited (NEW) and previously enrolled (DSP-enrolled) participants.

    Survey responses from (NEW) and those previously enrolled (DSP-enrolled) were compared using Mann-Whitney U tests. Analyses were restricted to survey items with sufficient responses across groups. Questions that were only administered to authors were excluded due to the limited number of author respondents (NEW, n = 7; DSP-enrolled, n = 4). For all other questions, responses from all participant roles (Quality/Impact reviewers, Impact-only reviewers, authors) were pooled within the NEW and DSP-enrolled groups and compared. Lower scores corresponded to more favorable responses. P-values were adjusted using the Benjamini–Hochberg procedure. Effect sizes (r) were calculated from the standardized Mann–Whitney statistic as r = Z/√N, where Z was derived using the normal approximation with continuity and tie corrections and N is the total number of observations.

    (DOCX)

    pbio.3003986.s009.docx (22KB, docx)
    S3 Table. Bootstrap summary comparing variability of Quality and Impact reviewer scores.

    Mean differences (Impact - Quality), confidence intervals, and one-sided p-values derived from 5,000 bootstrap iterations in which Impact scores were randomly subsampled to match the number of Quality scores for each manuscript. One-sided p-values were defined as the proportion of iterations in which the mean difference (Impact - Quality) was less than or equal to zero.

    (XLSX)

    pbio.3003986.s010.xlsx (9.3KB, xlsx)
    Attachment

    Submitted filename: Response_to_Reviewers_MFK.docx

    pbio.3003986.s013.docx (62.3KB, docx)
    Attachment

    Submitted filename: Response_to_Reviewers2nd.docx

    pbio.3003986.s014.docx (36.5KB, docx)

    Data Availability Statement

    Source data are available in Supplementary Tables or at https://doi.org/10.5281/zenodo.20041225, as indicated in the Figure Legends. The code for analyses is available at https://github.com/mmcgargi/Discovery-Stack-Analysis and archived at https://doi.org/10.5281/zenodo.22087956.


    Articles from PLOS Biology are provided here courtesy of PLOS

    RESOURCES