Abstract
Objectives
Our main objective is to assess the inter-reviewer reliability (IRR) reported in published systematic literature reviews (SLRs). Our secondary objective is to determine the expected IRR by authors of SLRs for both human and machine-assisted reviews.
Methods
We performed a review of SLRs of randomised controlled trials using the PubMed and Embase databases. Data were extracted on IRR by means of Cohen’s kappa score of abstract/title screening, full-text screening and data extraction in combination with review team size, items screened and the quality of the review was assessed with the A MeaSurement Tool to Assess systematic Reviews 2. In addition, we performed a survey of authors of SLRs on their expectations of machine learning automation and human performed IRR in SLRs.
Results
After removal of duplicates, 836 articles were screened for abstract, and 413 were screened full text. In total, 45 eligible articles were included. The average Cohen’s kappa score reported was 0.82 (SD=0.11, n=12) for abstract screening, 0.77 (SD=0.18, n=14) for full-text screening, 0.86 (SD=0.07, n=15) for the whole screening process and 0.88 (SD=0.08, n=16) for data extraction. No association was observed between the IRR reported and review team size, items screened and quality of the SLR. The survey (n=37) showed overlapping expected Cohen’s kappa values ranging between approximately 0.6–0.9 for either human or machine learning-assisted SLRs. No trend was observed between reviewer experience and expected IRR. Authors expect a higher-than-average IRR for machine learning-assisted SLR compared with human based SLR in both screening and data extraction.
Conclusion
Currently, it is not common to report on IRR in the scientific literature for either human and machine learning-assisted SLRs. This mixed-methods review gives first guidance on the human IRR benchmark, which could be used as a minimal threshold for IRR in machine learning-assisted SLRs.
PROSPERO registration number
CRD42023386706.
Keywords: Systematic Review, Randomized Controlled Trial, Surveys and Questionnaires
STRENGTHS AND LIMITATIONS OF THIS STUDY.
First assessment of threshold of agreement between human reviewers of systematic literature reviews.
First reference for a threshold of agreement for machine learning-assisted systematic literature reviews.
Under-reporting of inter-reviewer reliability metrics may not accurately reflect the true agreement.
Sample size of the survey is small, undermining both the internal and external validity.
Introduction
Evidence-based medicine (EBM) is an integration of clinical expertise combined with the best available evidence from systematic research. The aim of EBM is to combine comprehensive evidence with patient’s values to inform decision-making for the individual care pathway and the development of clinical (treatment) guidelines.1 Synthesis of available evidence by means of a systematic literature review (SLR) forms the foundation to this type of informed medical decision-making, making it one of the most important sources of evidence.2 The rigorous character of SLRs combined with the increasing volume of evidence and the need for systematic updates to prevent evidence to become outdated3 4 has put excessive pressure on researchers involved in evidence generation and assessment. The potential impact of outdated and incomplete health information goes beyond the field of evidence generation and is likely to result in suboptimal treatment of patients. As a result, automating aspects of the SLR process could lead to better and up-to-date informed medical decision-making, and thus indirectly improve the health of individual patients and entire populations.
A combination of human and machine learning automation efforts has been proposed to reduce the workload of conducting SLRs, and potentially enhance screening and data extraction quality. Machine learning can be used for both fully automated or assisted screening and eligibility assessment, as well as to support data extraction efforts, and has shown promising potential over the recent years.5 A recent survey found a 32% uptake of automation tools among systematic review practitioners, but the survey identified a lack of published evidence on the tool’s benefit as one possible cause of the relatively low uptake.6 The absence of transparency was also mentioned as a barrier to use automation tools. Another concern is machine learning’s compatibility with established methodology in evidence synthesis. The need for rigorously produced, disseminated and easily accessed evidence of machine learning validation is one of the key aspects to advance the field of machine learning in evidence synthesis. The demonstration of accuracy versus human classification is one of the first steps in the significant introduction of machine learning into evidence synthesis.7
In terms of accuracy of agreement between reviewers of SLRs, Cohen’s kappa is the primary inter-reviewer reliability (IRR) score often used to assess the agreement between two reviewers.8 Cohen’s kappa is defined as the relative observed agreement among reviewers, corrected for the probability of chance of agreement. There is also a level of agreement scale attached to the Cohen’s kappa score ranging from everything under 0.20 as none to slight agreement, 0.21–0.39 as minimal agreement, 0.40–0.59 as weak agreement, 0.60–0.79 as moderate agreement, 0.80–0.90 as strong agreement and everything above 0.90 as almost perfect agreement.9 Notably, the Cohen’s kappa score can also be used to analyse machine learning approaches’ performance in comparison to human classification in automation tasks. For example, the Cohen’s kappa score was used to evaluate the performance of a machine learning model on automated classification of patient-based age-related macular degeneration severity using colour fundus photographs.10 According to our knowledge, both an assessment and a reference standard of the inter-reviewer reliability between human literature reviewers or between human and machine learning automated IRR is yet to be determined.
One way to potentially determine how well the researchers involved in a SLR understood the topic being investigated is the reporting and level of the disagreements between the researchers.11 This can be used as a proxy for the quality of the evidence summarised in such SLRs. As a first step to understanding the potential improvements machine approaches can bring, it is important to consider the standard of human executed SLRs. To define the accuracy of machine learning approaches in SLRs, it is necessary to compare machine learning automation approaches for systematic review screening with the human performed SLRs using the human–human agreement as a reference. In this way, the human inter-reviewer reliability can be used as a reference to validate the introduction of machine learning automation in SLRs.
When looking at a similar problem in setting the standards for the introduction of machine learning automation, autonomous self-driving cars show similar high stakes in terms of human lives and the automation of jobs.12 One of the problems encountered with setting safety standards for self-driving cars is the better-than-average effect. At the level of individual risk assessment, most drivers perceive themselves to be safer than the average driver.12 Most drivers want self-driving cars that are safer than their own perceived ability to drive safely before they would feel reasonably safe riding in a self-driving vehicle or buying a self-driving vehicle, all other things being equal.12 To our knowledge, the better-than-average effect is unknown in the introduction of machine learning automation in evidence synthesis.
The aim of this work is to assess the level of agreement of human-executed SLRs and to assist in setting objectives for machine learning algorithms and creating a benchmark for determining the level of conflicts between machine learning algorithm classification and human classification. Therefore, our main objective is to assess the IRR reported in systematic reviews and analyse how it relates with the overall systematic review quality, reviewer experience and size of the screening task. Our secondary objective is to determine the expected IRR by authors of SLRs and to determine the potential existence of the better-than-average effect between human and automated SLRs.
Methods
This mixed-methods review consists of two parts to create a comprehensive synthesis of quantitative and qualitative data of IRR for screening and data extraction in SLRs. In the first part of this study, we performed a review of SLRs of randomised controlled trials (RCTs). The Pitts web application (www.pitts.ai) was used for both the literature screening as well as the data extraction.13 The machine learning components of the Pitts web application were not employed for any steps of this systematic review. To limit the scope of the review and to increase the comparability of the included studies, we only included systematic reviews of RCTs. We followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guideline14 for reporting for this part of the study (online supplemental appendix A). In the second part of this research, we surveyed authors of SLRs on their expectations of machine learning automation and human performed inter-reviewer reliability in SLRs. For the survey, we used the Qualtrics platform, a user-friendly, feature-rich, web-based survey tool which allows users to build, distribute and analyse online surveys. We followed the Strengthening the Reporting of Observational Studies in Epidemiology guideline15 for reporting of observational studies for this part of the study (online supplemental appendix B).
bmjopen-2023-076912supp001.pdf (58.2KB, pdf)
bmjopen-2023-076912supp002.pdf (112.1KB, pdf)
Systematic literature review
Selection criteria for eligible studies
For title and abstract, and full-text screening, the Sample, Phenomenon of Interest, Design, Evaluation, Research type framework16 (online supplemental appendix C) was used to guide the selection of the keywords and organise the inclusion and exclusion criteria. We included SLRs of RCTs with double-blind screening that reported on inter-reviewer reliability of screening or data extraction decisions. We restricted the search to publications in the past 5 years.
bmjopen-2023-076912supp003.pdf (64.1KB, pdf)
Studies were included if they met the following inclusion criteria
-
Systematic review articles of RCTs in which:
Two or more reviewers were involved during literature screening or data extraction.
Level of agreement between reviewers was reported.
The studies screening process was double-blinded.
The number of inclusion and excluded studies was reported.
Kappa score or other inter-reviewer agreement metrics was reported for abstract screening or full-text screening, or overall literature screening, or double data extraction was reported. Kappa score or other inter-reviewer agreement metrics can be calculated for any of the above items.
Studies were excluded if they met the following exclusion criteria
-
Systematic review.
Where not two or more reviewers were involved in literature screening or data extraction.
Which does not report level of agreement between reviewers.
In which the study screening was not double-blinded or not clearly reported.
In which the data extraction was not double-blinded or not clearly reported.
Where the number of included and excluded studies was not reported.
The following types of studies were excluded: diagnostic test reviews, individual patient data meta-analysis, scoping review, realist reviews, systematic review of health economic evaluation, empirical studies, case reports, narratives, letters to editors, genome-wide association meta-analysis, umbrella review and mixed-methods review.
Search strategy and study selection
The databases PubMed and Embase were searched using the search terms presented below. The search was limited in terms of language (Dutch and English) and time (publication between January 2017 and December 2022). The search was performed on 7 December 2022. Two reviewers performed the title and abstract literature screening independently. The articles that fulfilled the eligibility criteria were retrieved as full text. The full-text screening was performed again by two reviewers independently. A third reviewer was asked to resolve the disagreement between the two authors during the title and abstract screening as well as the full-text screening. Agreement between reviewers will be presented via the Cohen’s kappa score, for abstract screening, full-text screening and data extraction separately.
Search query PubMed via https://pubmed.ncbi.nlm.nih.gov/ (416 hits):
(("Systematic Review" [Publication Type] OR "Meta-Analysis" [Publication Type]) AND ("inter rater agreement" OR "inter-rater reliability" OR "IRR" OR "percent agreement" OR "percentage agreement" OR "reviewers agreed" OR no disagreement OR "Cohen's kappa" OR "Cohen’s Kappa statistic" OR "Cohen’s kappa coefficient" OR "Cohen’s Κ" OR "kappa test" OR "kappa*")) AND (randomised controlled trial OR randomized controlled trial OR rct OR randomized control trial OR randomised control trial) AND (y_5[Filter])
Search query Embase via https://www.embase.com/%23advancedSearch/default (454 hits):
('meta analysis'/exp OR ’systematic review'/exp OR 'meta-analysis' OR 'metaanalysis') AND ('inter rater agreement' OR 'inter-rater reliability' OR 'irr' OR 'percent agreement' OR 'percentage agreement' OR 'reviewers agreed' OR 'no disagreement' OR 'cohen/s kappa' OR 'cohen/s kappa statistic' OR 'cohen/s kappa coefficient' OR 'cohen/s κ' OR 'kappa test' OR kappa*) AND ('randomised controlled trial' OR 'rct' OR 'randomized control trial' OR 'randomised control trial' OR 'randomized controlled trial') AND [2017–2022]/py.
Data extraction and data synthesis
Two reviewers independently performed the data extraction in a custom-made data form. If necessary, data were calculated from the available information in the article. If there were any inconsistencies during data extraction, a third author was consulted. If necessary, data were calculated from the available information in the article. If data weres missing, we tried to contact the author to retrieve the missing information. The following data were extracted: study characteristics, review team size, screening data (ie, number of papers screened for title/abstract, number of papers retained after title/abstract screening, number of studies retained following full-text screening, and reporting of inclusion and exclusion criteria), study quality (A MeaSurement Tool to Assess systematic Reviews (AMSTAR) 217) and Cohen’s kappa score or other inter-reviewer reliability metrics of title/abstract screening, full-text screening and/or data extraction. In addition, one of the reviewers assessed if the protocol of the included SLR was registered in a publicly available database such as PROSPERO.
For the quality assessment of the included SLRs, we performed the AMSTAR 217 checklist on all the included SLRs. AMSTAR 2 is a critical appraisal tool for SLRs, which consists of 16 questions. The assessment of quality results was done by following the approach recommended by the authors. The quality of each SLR was deemed:
High; no or one non-critical weakness: The systematic review provides an accurate and comprehensive summary of the results of the available studies that address the question of interest.
Moderate; more than one non-critical weakness: The systematic review has more than one weakness but no critical flaws. It may provide an accurate summary of the results of the available studies that were included in the review.
Low; one critical flaw with or without non-critical weaknesses: The review has a critical flaw and may not provide an accurate and comprehensive summary of the available studies that address the question of interest.
Critically low; more than one critical flaw with or without non-critical weaknesses: The review has more than one critical flaw and should not be relied on to provide an accurate and comprehensive summary of the available studies.
Survey
The list of questions included in the survey is presented in online supplemental appendix D. The survey questions are intended to collect data on the participants’ expectations on the Cohen’s kappa score of SLRs between two human reviewers or a human reviewer and a machine learning agent, the better-than-average expectations of the participant, and the participants’ own experience on presenting the Cohen’s kappa score of their SLRs as primary outcome. In addition, we collected the participants’ scientific experience, experience with SLRs and experience with automation in SLRs as potential effect modifiers. Participants in this survey were also asked to reflect on the study objectives. We contacted authors of SLRs included in our review via email to complete our survey. In addition, we used a snow-balling approach to identify authors of SLRs from our own network. We followed up with the authors once after the initial contact. Informed consent was given by all authors who participated in the survey. The survey was open between 31 January 2023 and 17 February 2023, in which we aimed for a response rate of 10% of the targeted SLR authors (in the past 5 years).
bmjopen-2023-076912supp004.pdf (61.7KB, pdf)
Patient and public involvement
No patient involved.
Data analysis
We calculated the mean, variance and SD statistics over the extracted IRR metrics. The analysis of variance test was used to determine if a significant difference in the Cohen’s kappa exists between the different levels of review team size, screening data size and AMSTAR 2 ratings. An alpha value of <0.001 was used to determine statistical significance. All survey metrics were presented against the different aspects of the reviewer (ie, reviewer experience, reviewer ability, reviewer background, publication experience and academic qualification of the reviewer) in a colour-coded heat grid to observe trends in the association with human and machine learning-assisted reviewing expected IRR. All questions of the survey needed to be answered to avoid missing data. Only surveys that were answered fully were analysed.
Results
Systematic literature review
A total of 836 records were identified based on applying the search query after removal of duplicates. As part of the study selection procedure, we excluded 423 records during abstract screening and another 363 studies were excluded during full-text screening. In total, 45 articles met the eligibility criteria and were included in this review. Figure 1 gives an overview of the flow diagram of the systematic review. The primary reason for exclusion at full-text level was that the full-text study did not report on the level of agreement between reviewers (n=307). An overview of the included studies and the extracted data are given in online supplemental appendix E. The Cohen’s kappa scores between reviewers for the abstract screening and full-text screening were both 0.72.
Figure 1.
Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) flow diagram 2020 for systematic reviews.14 The number represents the total respondents for the combination of answer categories, a darker colour represents a higher number of respondents.
bmjopen-2023-076912supp005.pdf (71.1KB, pdf)
Among the included SLRs, on average a team of 2.45 (SD=1.17) reviewers screened 3386 abstracts (SD=8880), 129 full texts (SD=170) and included 41 articles (SD=71). The average Cohen’s kappa score was 0.82 (SD=0.11, n=12) for abstract screening, 0.77 (SD=0.18, n=14) for full-text screening, 0.86 (SD=0.07, n=15) for the whole screening process and 0.88 (SD=0.08, n=16) for data extraction. Most of the studies (84.4%, n=38/45) reported on the inclusion and exclusion criteria. However, the AMSTAR 2 rating was critically low in almost all the studies included (91.1%, n=41/45). No association was observed between the IRR reported and review team size, items screened and quality of the SLR as reported by the AMSTAR 2 rating as well as the number of ‘no’s’ reported in the rating.
Survey
In total, 37 respondents completed the survey (table 1). The full survey results can be found in online supplemental appendix F. Academic qualifications were primarily Ph.D. or doctorate degree (59.5%, n=22/37), but with an almost evenly spread of H-indexes. Most of the respondents have their professional activity in academia (83.8%, n=31/37), followed by industry (10.8%, n=4/37), government (5.4%, n=2/37) and healthcare organisations (5.4%, n=2/37).
Table 1.
Characteristics of survey respondents
| Characteristics | No of respondents (%), n=37 |
| Academic qualification | |
| Associate degree | 1 (2.7) |
| Bachelor’s degree | 3 (8.1) |
| Master’s degree | 11 (29.7) |
| Ph.D. or doctorate degree | 22 (59.5) |
| Number of publications | Median=32 (25th–75th percentile= 8–100) |
| H-Index | |
| Between 0 and 5 | 10 (27.0) |
| Between 6 and 10 | 7 (18.9) |
| Between 11 and 15 | 5 (13.5) |
| Between 16 and 20 | 6 (16.2) |
| Higher than 20 | 9 (24.3) |
| Sector | |
| Academia | 31 (83.8) |
| Industry | 4 (10.8) |
| Government | 2 (5.4) |
| Healthcare | 2 (5.4) |
| Scientific experience | |
| Between 0 and 5 years | 8 (21.6) |
| Between 6 and 10 years | 11 (29.7) |
| Between 11 and 15 years | 5 (13.5) |
| Between 16 and 20 years | 5 (13.5) |
| Higher than 20 years | 8 (21.6) |
| Publication experience SLRs | |
| Between 0 and 1 SLRs | 10 (27.0) |
| Between 2 and 4 SLRs | 9 (24.3) |
| Between 5 and 7 SLRs | 7 (18.9) |
| Between 8 and 10 SLRs | 3 (8.1) |
| More than 10 SLRs | 8 (21.6) |
| Machine learning experience SLRs | |
| I am far below average | 14 (37.8) |
| I am below average | 9 (24.3) |
| I am average | 11 (29.7) |
| I am above average | 2 (5.4) |
| I am far above average | 1 (2.7) |
| Screening ability (ie, time spend per article and overall precision) | |
| I am far below average | 0 (0.0) |
| I am below average | 1 (2.7) |
| I am average | 19 (51.4) |
| I am above average | 13 (35.1) |
| I am far above average | 4 (10.8) |
| Data extract ability (ie, time spend per article and overall precision) | |
| I am far below average | 0 (0.0) |
| I am below average | 2 (5.4) |
| I am average | 16 (43.2) |
| I am above average | 16 (43.2) |
| I am far above average | 3 (8.1) |
SLRs, systematic literature reviews.
bmjopen-2023-076912supp006.pdf (724.5KB, pdf)
Both scientific experience and experience with SLR publications of respondents were evenly distributed among the different categories (table 1). However, experience with machine learning within SLRs was reported to be below average (self-assessed) for most of the respondents (62.2%, n=23/37). Both the screening ability and the data extraction ability (ie, time spent per article and overall precision) were estimated by themselves to be above average for the larger part of the respondents (97.3%, n=36/37 and 94.6%, n=35/37). The survey results for the expected IRR showed overlapping expected Cohen’s kappa values compared with the SLR performed ranging between approximately 0.6–0.9, indicating a moderate to strong agreement between reviewers (figures 2 and 3). In general, respondents expect a higher-than-average IRR for machine learning-assisted SLRs compared with human based SLRs in both the screening and the data extraction. No trend was observed between reviewer experience (ie, scientific experience, publication experience SLRs, machine learning experience SLRs, and screening and data extraction ability) and expected IRR in both human based SLR and machine learning-assisted SLR.
Figure 2.
Lowest acceptable agreement expressed in Cohen’s kappa score between two reviewers (human–human or human–machine learning agent) for double-blinded literature abstract screening, full-text screening and data extraction decisions. The number represents the total respondents for the combination of answer categories, a darker colour represents a higher number of respondents.
Figure 3.
Lowest acceptable agreement expressed in Cohen’s kappa score between two reviewers (human–human or human–machine learning agent) for double-blinded literature abstract screening, full-text screening and data extraction decisions to be published in a scientific journal.
When comparing figures 2 and 3, the IRR (human–human or human–machine learning agent) acceptability in general and the acceptability to be published in a scientific journal was addressed. Similar trends were observed in both human–human and human–machine learning agent expected IRR, in which more respondents prefer not to have a minimal Cohen’s kappa score when addressing acceptability to be published in scientific journals. The absolute Cohen’s kappa score between general acceptability and acceptability to be published in a scientific journal remains similar for both human and machine learning-assisted acceptability for the respondents.
The majority of the respondents indicated that automated literature screening systems should be above average with respect to their ability to screen (72.97%) or extract (70.27%) accurately before they should be used for peer-reviewed and journal-published SLRs. This is in line with the number of respondents who indicated that automated literature screening systems should perform above average before the respondent would feel confident using the system (75.67%), compared with performing average (18.92%) or respondents that would never use an automated literature screening system (5.41%). Among the respondents, 45.95% indicated that they reported before on IRR for either screening or data extraction in their SLRs. Among reasons to not report on IRR, respondents indicated that they were not aware of the need of reporting such metrics or the software used was not capable of retrieving such information. An important common response on not reporting IRR was the fact that the reviewers included a third assessor to solve disagreement and the IRR only reflects part of the consensus-building process. Misinterpretation of the IRR as a quality measure of the reviewing process was mentioned as another underlying reason to not record or report IRR in SLRs.
Discussion
The aim of this study was to assess the IRR reported in human performed SLRs and expected IRR of SLRs authors of both human and machine learning-assisted SLRs. The findings from our review on SLRs of RCTs show that there is moderate to strong agreement between reviewers in published reviews. The data from our survey show expected moderate to strong agreement between reviewers, and a trend of higher agreement throughout the screening process. On average, respondents of our survey expect a higher-than-average IRR agreement for machine learning-assisted SLR compared with human based SLR. The association between reported IRR agreement between reviewers and review team size, items screened and quality of the SLR, and reviewer experience and expected IRR was absent, likely because of the low sample size.
The findings of our review highlight that human reviewers are able to reach moderate to strong agreement, but not perfect agreement, despite the use of certain mechanisms, such as clearly defined study selection criteria to improve agreement outcomes. The reviewer’s decision is likely dependent on individual characteristics and interpretation of the reviewing process. Some may strictly follow the inclusion and exclusion criteria, while others may not be as strict and be more inclusive, this leads to disagreement and a resultant suboptimal inter-reviewer agreement score.18 The survey showed an increase in expected agreement as the screening process progresses, with the expected agreement of abstract screening being lower relative to full-text screening and data extraction. This might be attributed to the fact that authors expect that the understanding of the inclusion and exclusion criteria increases as the work progresses thereby minimising the disagreement, and that full-text papers contain much more information than abstracts. Additionally, the discussions held among the reviewers to settle the disagreement through consensus or a third party in the early stages of the review might help one reviewer understand the decision-making behaviour of the others and foster harmony in the later stages. However, such a trend was not observed in the SLRs of RCTs, where IRR was comparable among the different review steps. An explanation could be that the learning effect of the reviewers as described above is cancelled out by the higher quantity of data you can have disagreement on.
SLRs are trusted to inform clinical practices and public health policies and have been used for this purpose for a very long time. In case of deviation from the conventional practice, it is very likely that stakeholders may have concerns with a new method that might be considered not as effective. Lacking trust from users is a major barrier to the adoption of machine learning in the SLRs process.19 Therefore, in order to forward the adoption of machine learning in SLRs, researchers, funding agencies and policy-makers must be convinced that the new norm is either equal to or better than current practice.19 The early adoption of machine learning for SLRs is already well documented in literature20–22 and, therefore, it is crucial to investigate reviewers’ expectations for the use of machine learning. This can impact how much reviewers can rely on machine learning in comparison to their own and other human reviewers' performance. In that regard, our study can be used to set a perceived performance benchmark for machine learning vis-à-vis the human reviewers’ performance. A finding of our research is that in order for machine learning to be widely integrated in the systematic review process, it should demonstrate above-average IRR agreement and above-average ability to carefully screen literature and extract data. Considering the IRR found in our review (moderate agreement between human reviewers), this would mean that machine learning reviews should have a minimal strong agreement.
To our knowledge, this is the first attempt to assess the threshold of agreement between human-executed reviewers to guide the process of setting threshold for machine learning-assisted SLRs. However, the results should be seen in light of its limitations. It is to assume that there is publication bias present in the reporting of IRR in SLRs. If reported, the IRR level in SLRs would likely be higher compared with non-reported IRR. We found scarce reporting on Cohen’s kappa scores for IRR in SLRs of RCTs. About 83% of the articles in our SLRs were excluded during the full-text review because they did not contain any IRR data. However, it is unclear whether the authors did not conduct an IRR at all or why they did not report it in their publication. This may be partially because the Cochrane Review Guide and other major systematic review guidelines do not require conducting an IRR or reporting data on IRR for literature screening.23 The publication bias in this particular instance would be more pronounced as the IRR measures the agreement level of authors’ who themselves are directly involved in reporting the finding. Therefore, it is important to note that the IRR reported in SLRs may be inflated and so, not accurately reflect the true IRR. In those studies that reported Cohen’s kappa statistics, variability in the kappa scores was found. Differences in researcher experience in combination with the difficulty of the screening task attributed to the fineness of the inclusion and exclusion criteria, time constraints on reviewing, and level of preparation including pretest and training for screening and data extraction, likely drove variations in inter-reviewer kappa estimates, but it was not possible to assess this in our review. As reported in our survey an important common response on not to report was the fact that the reviewers included a third assessor to solve disagreement and the IRR only reflects part of the consensus-building process. Misinterpretation of the IRR as a quality measure of the reviewing process was mentioned as another underlying reason to not record or report IRR in SLRs. The process of agreement between humans is likely to be more iterative than the final IRR measure presented in the final publication. However, the use of AI in machine-assisted reviews, where the AI model introduction is often trained or retrained, could be seen as an iterative process by itself as well. In that sense, both human–human IRR reported in SLRs, and machine-assisted IRR in SLRs could be facing the same interpretation issues.
A sizeable portion of respondents to the survey expressed the opinion that there should not be a minimum acceptable agreement level (expressed in kappa score) for screening and data extraction decisions or for publishing a systematic review in a scientific journal, indicating that reviewers will continue to question the appropriateness of IRRs for the benchmarking. To increase confidence in the accuracy of the overall review, it is crucial to ensure reliability during screening and data extraction.18 SLRs involve a number of subjective decisions that need to be recorded along the way from screening to data extraction in order to ensure transparency and replicability.18 24 It is only in this sense that the significance of IRR can be fully appreciated. Due to the lack of prior research on the topic, it is impossible to compare the findings of the current study with those of other studies, which is another limitation of our study. Finally, the sample size both for the SLRs and survey is small, undermining both the internal and external validity of the study. We recommend the future study with a larger sample size to generate more accurate results. Future SLRs should report IRR kappa scores as a best scientific practice to showcase how well the researchers involved in SLRs understood the work they did. These anticipated future results could build further on the quality of human-executed SLRs and the value of machine learning-assisted methods for conducting SLRs. In addition, future validation and direct comparison of machine-assisted reviews versus the human reviewers’ performance should place the IRR threshold in the context of the real-world performance as opposed to the indirect comparison which is made in this study. Accuracy measures of the reviewing process should be further explored to guide the process and evaluation of introduction of machine learning in evidence generation and evidence synthesis.
Currently, it is not common to report on IRR in the scientific literature for either human or machine learning-assisted SLRs. Human performed SLRs likely show a moderate agreement between reviewers, while authors expect machine learning-assisted SLRs to perform better. This mixed-methods review gives first guidance on the human IRR benchmark, which could be used as a minimal threshold for IRR in machine learning-assisted SLRs. A minimal strong agreement between reviewers of machine learning-assisted SLRs is recommended to ensure overall acceptance of machine learning in SLRs.
Supplementary Material
Footnotes
Contributors: Conceptualisation of this study was done by PH, AW, SA, LQ, ML, MP, CB and JvdS. The design of the methodology was done by PH, AW, SA, LQ, ML, MP, CB and JvdS. The abstract and full-text screening was done by AW and JvdS, with CB for disagreement resolution. The data extraction was performed by AW, PH, with CB for disagreement resolution. The statistical analysis was performed by RdJ and JJM. Writing, reviewing and editing was done by all authors. All authors have read and agreed to the published version of the manuscript. JvdS acts as the guarantor for this publication.
Funding: This work was funded by F. Hoffmann-La Roche, Basel, Switzerland (grant number: N/A).
Competing interests: MP and CB reported stock ownership in Health-Ecore B.V. LQ, ML and SA are employed by Roche. SA and ML reported stock ownership in Roche. Roche provided funding for this research to Health-Ecore and Pitts. PH and JJM reported ownership in Pitts. The other authors declare that they have no further competing interests related to this specific study and topic.
Patient and public involvement: Patients and/or the public were not involved in the design, or conduct, or reporting, or dissemination plans of this research.
Provenance and peer review: Not commissioned; externally peer reviewed.
Supplemental material: This content has been supplied by the author(s). It has not been vetted by BMJ Publishing Group Limited (BMJ) and may not have been peer-reviewed. Any opinions or recommendations discussed are solely those of the author(s) and are not endorsed by BMJ. BMJ disclaims all liability and responsibility arising from any reliance placed on the content. Where the content includes any translated material, BMJ does not warrant the accuracy and reliability of the translations (including but not limited to local regulations, clinical guidelines, terminology, drug names and drug dosages), and is not responsible for any error and/or omissions arising from translation and adaptation or otherwise.
Data availability statement
No data are available.
Ethics statements
Patient consent for publication
Consent obtained directly from patient(s).
Ethics approval
This study involves human participants but patient anonymity is guaranteed by the use of a unique anonymous identifier, hence the ethical approval for observational studies has been waived. Participants gave informed consent to participate in the study before taking part.
References
- 1.Sackett DL, Rosenberg WM, Gray JA, et al. Evidence based medicine: what it is and what it isn't. BMJ 1996;312:71–2. 10.1136/bmj.312.7023.71 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Gough D, Elbourne D. Systematic research synthesis to inform policy, practice and democratic debate. Soc Policy Soc 2002;1:225–36. 10.1017/S147474640200307X [DOI] [Google Scholar]
- 3.Shojania KG, Sampson M, Ansari MT, et al. How quickly do systematic reviews go out of date? A survival analysis. Ann Intern Med 2007;147:224–33. 10.7326/0003-4819-147-4-200708210-00179 [DOI] [PubMed] [Google Scholar]
- 4.Elliott JH, Synnot A, Turner T, et al. Living systematic review: 1. introduction-the why, what, when, and how. J Clin Epidemiol 2017;91:23–30. 10.1016/j.jclinepi.2017.08.010 [DOI] [PubMed] [Google Scholar]
- 5.Cierco Jimenez R, Lee T, Rosillo N, et al. Machine learning computational tools to assist the performance of systematic reviews: a mapping review. BMC Med Res Methodol 2022;22:322. 10.1186/s12874-022-01805-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.van Altena AJ, Spijker R, Olabarriaga SD. Usage of automation tools in systematic reviews. Res Synth Methods 2019;10:72–82. 10.1002/jrsm.1335 [DOI] [PubMed] [Google Scholar]
- 7.Arno A, Elliott J, Wallace B, et al. The views of health guideline developers on the use of automation in health evidence synthesis. Syst Rev 2021;10:16. 10.1186/s13643-020-01569-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Park CU, Kim HJ. Measurement of inter-rater reliability in systematic review. Hanyang Med Rev 2015;35:44. 10.7599/hmr.2015.35.1.44 [DOI] [Google Scholar]
- 9.McHugh ML. Interrater reliability: the Kappa Statistic. Biochem Med (Zagreb) 2012;22:276–82. [PMC free article] [PubMed] [Google Scholar]
- 10.Peng Y, Dharssi S, Chen Q, et al. Deepseenet: a deep learning model for automated classification of patient-based age-related macular degeneration severity from color fundus photographs. Ophthalmology 2019;126:565–75. 10.1016/j.ophtha.2018.11.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Belur J, Tompson L, Thornton A, et al. Interrater reliability in systematic review methodology. Sociol Methods Res 2018:004912411879937. 10.1177/0049124118799372 [DOI] [Google Scholar]
- 12.Nees MA. Safer than the average human driver (who is less safe than me)? Examining a popular safety benchmark for self-driving cars. J Safety Res 2019;69:61–8. 10.1016/j.jsr.2019.02.002 [DOI] [PubMed] [Google Scholar]
- 13.Pitts . Living systematic review software. Available: https://pitts.ai/ [Accessed 24 Nov 2022].
- 14.Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. Syst Rev 2021;10:89. 10.1186/s13643-021-01626-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Elm E von, Altman DG, Egger M, et al. Strengthening the reporting of observational studies in epidemiology (STROBE) statement: guidelines for reporting observational studies. BMJ 2007;335:806–8. 10.1136/bmj.39335.541782.AD [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Cooke A, Smith D, Booth A. Beyond PICO: the SPIDER tool for qualitative evidence synthesis. Qual Health Res 2012;22:1435–43. 10.1177/1049732312452938 [DOI] [PubMed] [Google Scholar]
- 17.Shea BJ, Reeves BC, Wells G, et al. AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of Healthcare interventions, or both. BMJ 2017;358:j4008. 10.1136/bmj.j4008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Belur J, Tompson L, Thornton A, et al. Interrater reliability in systematic review methodology: exploring variation in coder decision-making. Sociol Methods Res 2021;50:837–65. 10.1177/0049124118799372 [DOI] [Google Scholar]
- 19.O’Connor AM, Tsafnat G, Thomas J, et al. A question of trust: can we build an evidence base to gain trust in systematic review automation technologies. Syst Rev 2019;8:143. 10.1186/s13643-019-1062-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Cohen AM, Hersh WR, Peterson K, et al. Reducing workload in systematic review preparation using automated citation classification. J Am Med Inform Assoc 2006;13:206–19. 10.1197/jamia.M1929 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Howard BE, Phillips J, Miller K, et al. SWIFT-review: a text-mining workbench for systematic review. Syst Rev 2016;5:87. 10.1186/s13643-016-0263-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Liao J, Ananiadou S, Currie LG, et al. Automation of citation screening in pre-clinical systematic reviews. Neuroscience [Preprint] 2018. 10.1101/280131 [DOI]
- 23.Higgins JPT, Thomas J, Chandler J, et al., eds. Cochrane Handbook for Systematic Reviews of Interventions version 6.3 (updated February 2022). Cochrane, 2022. Available: www.training.cochrane.org/handbook [Google Scholar]
- 24.McHugh ML. Interrater reliability: the Kappa statistic. Biochem Med (Zagreb) 2012;22:276–82. [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
bmjopen-2023-076912supp001.pdf (58.2KB, pdf)
bmjopen-2023-076912supp002.pdf (112.1KB, pdf)
bmjopen-2023-076912supp003.pdf (64.1KB, pdf)
bmjopen-2023-076912supp004.pdf (61.7KB, pdf)
bmjopen-2023-076912supp005.pdf (71.1KB, pdf)
bmjopen-2023-076912supp006.pdf (724.5KB, pdf)
Data Availability Statement
No data are available.



