Skip to main content
Systematic Reviews logoLink to Systematic Reviews
letter
. 2024 Oct 25;13:269. doi: 10.1186/s13643-024-02682-2

Leveraging artificial intelligence to enhance systematic reviews in health research: advanced tools and challenges

Lixia Ge 1,, Rupesh Agrawal 2, Maxwell Singer 3, Palvannan Kannapiran 1, Joseph Antonio De Castro Molina 1, Kiok Liang Teow 1, Chun Wei Yap 1, John Arputhan Abisheganaden 1,4
PMCID: PMC11504244  PMID: 39456077

Abstract

Artificial Intelligence (AI) is transforming systematic reviews (SRs) in health research by automating processes such as study screening, data extraction, and quality assessment. This perspective highlights recent advancements in AI tools that enhance efficiency and accuracy in SRs. It discusses the benefits, challenges, and future directions of AI integration, emphasising the need for human oversight to ensure the reliability of AI outputs in evidence synthesis and decision-making in healthcare.

Keywords: Artificial intelligence, Systematic review, Health research

Introduction

Systematic reviews (SRs) are crucial for synthesising evidence from multiple studies to inform clinical practice and future research. However, they are labour-intensive, particularly in the phases of study screening and data. Since 2005, artificial intelligence (AI) tools have gained increasing attention [1] for their ability to automate these processes, offering increased efficiency and accuracy [2]. This letter outlines recent developments in AI for different SR processes, highlights commonly used tools, and discusses the challenges and future directions.

AI tools for research question development and search strategy

Developing a clear research question and a comprehensive search strategy to identify relevant studies is a crucial yet time-consuming step in conducting SRs. AI tools like OpenAI’s ChatGPT can assist by generating PICO-based research questions and tailored search strings for databases like PubMed and Embase [2]. Additionally, ChatGPT can create custom code to automate the search and retrieval process using the National Center for Biotechnology Information E-utilities application programming interface (API).

Platforms like searchrefiner employ automation tools to identify frequent MeSH terms from selected references and categorise them into health condition, treatment, and study design, which are subsequently used to develop Boolean queries [3, 4]. Such AI tools save time and enhance the thoroughness of search strategies, helping ensure relevant studies are not overlooked. However, careful human oversight remains essential to ensure alignment with research goals.

AI in study searching and screening

Study searching and screening are laborious but essential processes of SRs as a significant portion of retrieved records often turn out to be irrelevant. AI-powered tools can greatly reduce this workload by automating or semi-automating relevance assessment using machine learning (ML) techniques.

SR software platforms like Nested Knowledge and DistillerSR have integrated search engine functionalities, allowing the automatic import of records from databases via respective APIs and efficiently filtering duplicates. However, not all the database APIs are embedded in these platforms, and delays in API updates can affect accuracy.

Several software tools support titles and abstract screening, with Abstrackr and Rayyan being the most validated [5, 6]. Abstrackr helps save time with minimal risk of missing relevant records by assisting one or two reviewers [7]. Rayyan, a widely used web-based software platform, integrates AI tools to facilitate the screening process. Table 1 lists commonly used AI-enabled software and platforms developed or enhanced in the past 5 years.

Table 1.

The commonly used AI tools in different systematic review processes

Tools (version/latest update year) AI-assisted SR process Description Machine learning paradigm and algorithms Availability and cost
searchrefiner (2019) Developing search strategies An open-source, end-to-end web-based search strategy refining package to objectively identify terms and develop search strategies.

Learning paradigm: ensemble learning

Classifier: SVM, LDA

Model inputs: references / identified studies (title, abstract)

Feature extraction: MeSH terms

Free, under further development; https://ielab.io/searchrefiner
Scispace (2024) Literature search; Data extraction A platform designed to facilitate the process of academic research, management, and organisation of research papers.

Learning paradigm: supervised learning

Feature extraction: NER, part-of-speech tagging, and sentiment analysis

Deep learning using transformers (e.g., BERT, GPT)

Computer Vision technique: image recognition, Optical Character Recognition

Free for Basic with limited features; USD$144/year for Premium; USD$60–96/year per user under Labs&Universities, depending on number of users;

https://typeset.io/library

Elicit.org (v1.9/2024) Literature search; Data extraction; Synthesis A comprehensive web-based software application designed to support systematic reviews and other types of evidence synthesis. Large language models, advanced NLP techniques, process-based ML approach, chain-of-thought prompting, and fine-tuned models

Free for Basic plan, USD$120/year for Plus, USD$499/year for Pro;

Elicit.com

SWIFT-Review (v1.43/2019) Identifying research question; Screening; Data extraction; Synthesis An interactive desktop application providing numerous tools to assist with problem formulation and discovering important terms and phrases, as well as literature prioritisation.

Learning paradigm: unsupervised and supervised learning

Classifier: LDA

Feature extraction: TF–IDF

Free; https://www.sciome.com/swift-review/
DistillerSR (v2.31.0/2024) Searching; Screening; Data extraction A web-based systematic review management software designed to streamline and automate various aspects of the systematic review process.

Learning paradigm: active learning

Classifier: not publicly specified, allows user to create custom classifiers

Model inputs: title, abstract

Feature extraction: Word2Vec and other NLP techniques

Label options: include, exclude

USD$3120 for DistillerSR Professional (need to write into sales team)
Nested Knowledge (2024) Searching; Screening; Synthesis Offers a comprehensive software platform for systematic literature review and meta-analysis. It comprises two parts which work in tandem. Search, screen, extract data, and complete critical appraisal with AutoLit®. Visualise, analyse, publish, and share insights with Synthesis.

Learning paradigm: supervised and unsupervised learning

Classifier: not publicly specified, allows user to create custom classifiers

Model inputs: n-grams, keywords, and phrases from abstracts and full texts

Feature extraction: embeddings, NER

Label options: include, exclude

Free for first review; https://nested-knowledge.com/
Rayyan (modified everyday) Screening A web-based semi-automated tool for systematic review screening, supporting collaborative work and machine learning prioritisation, highlighting PICO.

Learning paradigm: active learning

Classifier: not publicly specified

Model inputs: user-provided key terms and references (title and abstract)

Feature extraction: unigrams, bigrams, keywords, and other advanced NLP techniques

Label options: include, exclude, maybe

Free with limited features for the Student plan,

USD$99.96/year for the Professional plan; https://new.rayyan.ai/

ASReview LAB (v1.6.2/2024) Screening An open-source machine learning-aided pipeline (installation of Python is required).

Learning paradigm: active learning

Classifier: NB; SVM; DNN; LR; LSTM-base; LSTM-pool; RF

Model inputs: piece of text (for example, title and abstract)

Feature extraction: Doc2Vec; embedding IDF, TF–IDF, sBERT

Label options: relevant; irrelevant

Free
Covidence Screening A web-based software platform designed to streamline and manage the systematic review process. It facilitates collaboration and simplifies the entire workflow, from screening and data extraction to quality assessment and synthesis.

Learning paradigm: active learning

Classifier: not publicly specified

Model inputs: not publicly specified

Feature extraction: not publicly specified

Label options: Yes, Maybe, No

Free trial

Single review: USD$289/year

Package (up to 3 reviews): USD$867/year

SWIFT-Active Screener (v1.21/2023) Screening A web-based application leveraging machine learning algorithms to prioritise and streamline the process of screening large volumes of references.

Learning paradigm: active learning

Classifier: not publicly specified

Model inputs: user-provided key terms and citation (abstract, title, keywords)

Feature extraction: not publicly specified

Label options: include, exclude

Commercial https://www.sciome.com/swift-activescreener/
Colandr (v2.0/2024) Screening; Data extraction Open-source machine-learning assisted online platform for conducting reviews and syntheses of text-based evidence.

Learning paradigm: active learning

Classifier: SVM with SGD learning

Model inputs: user-provided key terms and citation (abstract, title, keywords)

Feature extraction: Word2Vec

Label options: include, exclude

Free; https://www.colandrapp.com/
EPPI reviewer 6 (v6.15.3.0/2024) Screening; Synthesis An application designed to support all types of literature reviews.

Learning paradigm: active learning

Classifier: SVM

Model inputs: user-provided key terms and citation (abstract, title, keywords)

Feature extraction: Word2Vec

Label options: relevant, irrelevant

“Not for profit” service, monthly subscription-based with one-month free trial; https://eppi.ioe.ac.uk/cms/Default.aspx?tabid=2914
Abstrackr (beta/2021) Screening; Data extraction A semi-automated, open-source, web-based interface for screening and data extraction.

Learning paradigm: active learning

Classifier: SVM

Model inputs: user-provided keywords (relevant/irrelevant with degree of confidence); citations

Feature extraction: unigrams, bigrams, TF–IDF

Label options: relevant; borderline; irrelevant

Free; http://abstrackr.cebm.brown.edu/
RobotReviewer (2018) Data extraction; Quality assessment A machine learning system which takes RCT reports to automatically retrieve information describing study design, population, intervention, comparator, and outcomes, and automates quality assessment of RCTs using the Cochrane Risk of Bias tool.

Learning paradigm: supervised learning

Classifier: not publicly specified

Feature extraction: various NLP techniques

Free; https://www.robotreviewer.net/
Scispace’s GPT Data extraction; Quality assessment; Report writing A web-based tool leveraging advanced AI to support researchers throughout academic research and academic writing process. Transformer-based deep learning Free version with basic features, USD$240/year for subscription plan; https://chatgpt.com/g/g-NgAcklHd8-scispace

As a starting point we used the Table 1 in Rens van et al.’s paper [8] that describes existing tools implementing active learning for systematic reviewing

Abbreviations: BOW bag of words, DNN dense neural network, Doc2Vec document to vector, IDF inverse document frequency, GDPR general data protection regulation, GPT generative pre-trained transformers, LDA latent dirichlet allocation, NB naive Bayes, NER named entity recognition, NLP natural language processing, LR logistic regression, LSTM long short-term memory, MeSH medical subject headings, ML machine learning, RCT randomised control trial, RF random forests, sBERT sentence bidirectional encoder representations from transformers, SGD stochastic gradient descent, SVM support vector machine, TF–IDF term frequency–inverse document frequency, Word2Vec words to vector

The performance of these tools depends on the effectiveness of the underlying ML paradigms and algorithms, the quality and quantity of data used for training, and the degree of human involvement [9]. Studies have demonstrated that well-developed AI tools, though not yet fully automated, can match or even surpass human reviewers in screening efficiency, accuracy, and thus accelerating the SR process. Hence, AI application in this process is among the most developed and explored. However, the effectiveness of AI in screening largely depends on factors such as the quality of training data and the extent of human oversight or verification.

AI for data extraction

Data extraction is a manual, error-prone task that AI tools are increasingly automating using natural language processing (NLP) and ML algorithms. Tools like RobotReviewer have high accuracy in extracting relevant details from randomised control trials (RCTs) [10]. Nested Knowledge and Abstrackr (Table 1) highlight relevant sections for data extraction, while SciSpace and Elicit.org offer intuitive interfaces to automatically extract data from multiple papers and present in tables. SciSpace also provides citations for extracted data and enables PDF interactions via Copilot for validation. SciSpace’s GPT can extract data by interacting with individual uploaded PDFs through specific prompts. Although not formally validated, these tools show promise in improving efficiency and saving time, though variability in content identification may affect reproducibility and accuracy.

Despite their advancements, AI tools continue to face challenges with nuanced interpretations and poorly reported data. Their accuracy and reliability depend on factors such as algorithms, the quality and diversity of training data, the complexity of source documents, and the precision of prompts. While AI reduces data extraction time, it still requires cautious use and human oversight to ensure accuracy.

AI for quality assessment

AI tools can streamline quality assessment by applying specific criteria to studies. For instance, RobotReviewer uses NLP to identify relevant text and assess the risk of bias in RCTs, achieving an accuracy rate of 71–78.3% [11], though it is less precise than Cochrane reviews. An AI extension for the Prediction model Risk Of Bias Assessment Tool (PROBAST) is under development to evaluate bias in clinical prediction model studies [12]. Additionally, SciSpace’s GPT can provide quality assessment for uploaded papers with justifications when properly prompted. While promising, these tools still lack the precision of manual assessments by experienced reviewers, highlighting the need for further development and validation.

AI for synthesis and summarisation

Traditional ML and NLP tools are not capable of fully automating data synthesis [10]. However, advanced AI tools offer valuable assistance in aggregating information from large volumes of research evidence. Platforms like Nested Knowledge and DistillerSR (Table 1) automatically generate and update PRISMA diagrams and synthesis tables, while ChatGPT and SciSpace’s GPT can assist with academic writing. While AI tools cannot fully automate this process and the generated summaries must be cross-checked by humans, it can significantly aid in organising and summarising research findings, accelerating the preparation of SR reports.

Challenges and future directions

AI tools have shown great potential in streamlining various SR processes, yet challenges persist. The risk of AI-induced biases, particularly in sensitive fields like healthcare, raises concerns. To address this, robust evaluation frameworks and standards are needed to ensure these tools produce reliable and unbiased results. Transparency in AI algorithms is also crucial for fostering trust and reproducibility.

The advancement of large language model (LLM) technology, exemplified by models like ChatGPT, offers new avenues for automating and enhancing SRs by improving the understanding and processing of complex language data. APIs enable LLMs to be integrated into code for efficient, large-scale data processing. A ChatGPT v4.0 API script achieved 96% screening specificity and 93% sensitivity when prompted with inclusion and exclusion criteria [13], showcasing the capability of base models without fine-tuning. Consecutive scripts using LLM APIs could automate SRs from screening to report writing. However, errors compound across steps, with a 5% error per step resulting in an overall accuracy of about 81.5%. This highlights the need to minimise errors at each stage and consider human-in-the-loop mechanisms for complex tasks like SRs.

While AI tools, particularly LLMs, offer immense potential to streamline various SR processes, the role of AI must be carefully defined to avoid misuse. Clear guidelines for AI integration are essential to prevent AI-generated content from slipping through peer-review processes. AI tools should augment, not replace, the meticulous work of researchers. For example, LLMs can assist in generating initial drafts, but human experts must review, refine, and validate these outputs to ensure scientific rigour and accuracy. Additionally, integrating checks such as provenance tracking and flagging AI-generated sections can enhance accountability. Furthermore, AI-generated summaries or syntheses should always be cross-checked against source data by experienced reviewers to prevent the propagation of errors. Embedding human oversight at each stage ensures that the final SR remains credible and unbiased.

Integrating critically appraised LLM-based AI tools with existing SR software could further optimise the process, although creating effective and efficient prompts remains a challenge. Collaboration between AI developers, SR methodologists, and healthcare experts is essential to refine these tools to maximise their potential.

Conclusion

AI holds transformative potential to enhance SRs by increasing efficiency and quality, but human judgment remains essential to ensure reliability. Defining the complementary roles of AI and human reviewers will help maintain SR integrity while leveraging AI’s efficiencies. Continued development and validation of AI tools are vital to fully realise their benefits. By integrating AI judiciously with ongoing human oversight, researchers can conduct more efficient, accurate, and comprehensive SRs, advancing evidence-based healthcare.

Acknowledgements

Not applicable.

Abbreviations

AI

Artificial intelligence

API

Application programming interface

GPT

Generative pre-trained transformer

LLM

Large language model

ML

Machine learning

NLP

Natural language processing

RCT

Randomised control trials

SR

Systematic review

Authors’ contributions

LG and JADCM conceptualised the study. LG prepared Table 1, tested out the tools, and drafted and revised the manuscript. RA and MS substantially revised the manuscript with valuable inputs. PK, KLT, and CWY provided comments on AI and ML techniques. JADCM and JAA supervised the work. All authors reviewed and approved the final manuscript.

Funding

No funding was received for the study.

Data availability

Not applicable.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare that they have no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Cohen AM, Hersh WR, Peterson K, Yen P-Y. Reducing workload in systematic review preparation using automated citation classification. J Am Med Inform Assoc. 2006;13:206–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Fabiano N, et al. How to optimize the systematic review process using AI tools. JCPP Advances. 2024;4: e12234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Scells H, Zuccon G, Koopman B, Clark J, et al. A computational approach for objectively derived systematic review search strategies. In: Jose JM, et al., editors. Advances in Information Retrieval. Cham: Springer International Publishing; 2020. 10.1007/978-3-030-45439-5_26. [Google Scholar]
  • 4.Scells H, Zuccon G. searchrefiner: a query visualisation and understanding tool for systematic reviews. In: Proceedings of the 27th ACM International Conference on Information and Knowledge Management. New York: Association for Computing Machinery; 2018. p. 1939–42. 10.1145/3269206.3269215. [Google Scholar]
  • 5.Harrison H, Griffin SJ, Kuhn I, Usher-Smith JA. Software tools to support title and abstract screening for systematic reviews in healthcare: an evaluation. BMC Med Res Methodol. 2020;20:7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Hamel C, et al. An evaluation of DistillerSR’s machine learning-based prioritization tool for title/abstract screening–impact on reviewer-relevant outcomes. BMC Med Res Methodol. 2020;20:256. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Gates A, et al. The semi-automation of title and abstract screening: a retrospective exploration of ways to leverage Abstrackr’s relevance predictions in systematic and rapid reviews. BMC Med Res Methodol. 2020;20:139. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.van de Schoot R, et al. An open source machine learning framework for efficient and transparent systematic reviews. Nat Mach Intell. 2021;3:125–33. [Google Scholar]
  • 9.O’Connor AM, et al. A question of trust: can we build an evidence base to gain trust in systematic review automation technologies? Syst Rev. 2019;8:143. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Marshall IJ, Wallace BC. Toward systematic review automation: a practical guide to using machine learning tools in research synthesis. Syst Rev. 2019;8:163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Marshall IJ, Kuiper J, Wallace BC. RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials. J Am Med Inform Assoc. 2016;23:193–201. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Collins GS, et al. Protocol for development of a reporting guideline (TRIPOD-AI) and risk of bias tool (PROBAST-AI) for diagnostic and prognostic prediction model studies based on artificial intelligence. BMJ Open. 2021;11: e048008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Li M, Sun J, Tan X. Evaluating the effectiveness of large language models in abstract screening: a comparative analysis. Syst Rev. 2024;13:219. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Not applicable.


Articles from Systematic Reviews are provided here courtesy of BMC

RESOURCES