Abstract
Purpose of Review
To analyze the efficacy and efficiency of current large language models (LLMs), specifically GPT-5 in screening titles and abstracts for three review topics within different subspecialties in orthopedics.
Recent Findings
Python scripts were developed to call on the GPT-5 model via OpenAIs application programming interface (API). Two human reviewers simultaneously performed screening based on the same inclusion and exclusion criteria. Performance metrics such as specificity, sensitivity, accuracy, positive predictive value (PPV), negative predictive values (NPV), and F1 scores for GPT-5 were calculated based on a gold-standard inclusion and exclusion list developed by a third human adjudicator. Efficiency metrics included total cost and time to completion for each task. The number of titles and abstracts to screen ranged between 668 and 1,131 amongst the three review topics. All performance metrics were above 92.3% amongst all three topics, with sensitivities ranging from 94.1%-100%. Time to completion ranged between 38.5-174.3 minutes. Cost ranged from $1.32-$3.73USD
Summary
GPT-5 demonstrated exceptional accuracy, sensitivity, specificity, PPV, NPV, and F1 scores in automating title and abstract screening for three orthopedic systematic review topics in three different subspecialties. Results are similar to previous studies investigating the role of AI for screening, specifically increased accuracy and time-to-completion relative to humans. The average rate of screening ranged from 6.5-17.4 abstracts per minute and the average price ranged from $0.002-$0.0036USD per abstract, suggesting a high degree of efficiency compared to current standards.
Keywords: GPT-5, Artificial intelligence, Systematic review, Automation, Large language model, LLM
Introduction
Systematic reviews are widely considered to be an integral part of evidence-based practice due to robust methodological standards and guidelines [1]. However, these standards lead to time and resource intensive practices, namely in screening and extraction [1]. In orthopaedics, the screening process typically assesses hundreds, if not thousands of articles, while authors are required to maintain high levels of sensitivity [2]. This necessary precision can lead to a typical time from registration to completion ranging from 6 to 24 months [3, 4]. While these robust guidelines have been accepted globally, many of the organizations who identified such standards have recognized the potential impact which machine learning can have on accelerating these processes if the same level of robustness can be maintained [5].
The exploration of systematic review screening automation through machine learning initially yielded mixed results with a potential reduction of recall of up to 5% [4]. High recall is a mandatory standard enforced within systematic reviews and is thus a requisite for automated screening [cite]. However, the introduction of large language models (LLMs) has expanded the potential for higher degrees of automated recall. Studies have demonstrated that LLMs can screen studies with accuracies ranging from 82% to over 90% [6, 7]. When compared against human screeners, LLMs have demonstrated similar decision making and accuracy when screening studies [8]. Incorporating LLMs have not only performed with strong accuracy, but significantly reduced the workload and time to complete screening [7]. GPT-4 was a common high performer amongst LLMs within these studies, yielding the highest levels of accuracy and sensitivity [7]. However, most studies have investigated LLMs for systematic review screening within computer science journals [7, 8], with no study to date investigating this within the field of orthopedic surgery.
With the release of GPT-5, there is massive potential for heightened performance when utilizing this technology for title and abstract screening within the field of orthopedic surgery. As outlined by the International Collaboration for the Automation of Systematic Reviews (ICASR), it is integral that automation does not come at the expense of high accuracy and sensitivity. Therefore, validation of LLMs for this task is essential to transition to a future where this technology is widely adopted. Therefore, this study aims to assess the efficacy and efficiency of GPT-5 relative to humans in screening titles and abstracts for three different review topics within three different orthopedic subspecialties, each focusing on a different screening objective. It is hypothesized that GPT-5 will have strong performance metrics and will be able to complete tasks with a high-degree of efficiency.
Methods
Study Design and Overview
Three independent evaluations of GPT-5 and human reviewers for title and abstract screening in systematic reviews were conducted. Each evaluation targeted a distinct research objective: (1) identifying a niche patient or surgical population, (2) identifying specific study designs, and (3) identifying abstracts reporting on a specific outcome. To increase generalizability, each objective and evaluation was within a distinct orthopedic subspecialty (sports medicine, trauma, and arthroplasty).
Data Sources and Abstract Sets
Abstracts for each task were retrieved from comprehensive literature searches using PubMed. Alternative databases weren’t used as the focus of the study was to evaluate the ability of LLMs to correctly include and exclude titles/abstracts in comparison with humans as opposed to truly simulating the systematic review process. A full description of the search criteria is described in Table 1.
Table 1.
PubMed search strategies
| Date Searched | Search Criteria | |
|---|---|---|
| Topic 1: Sports Medicine - Clinical outcomes after revision ACLR with quadriceps tendon autograft | July 1, 2025 | (ACL OR anterior cruciate ligament) AND (quad*) AND (revised OR re-operation OR revision OR reoperation OR repeat OR second*) |
| Topic 2: Trauma - RCTs comparing clinical outcomes after operative vs. nonoperative management of midshaft clavicle fractures | July 12, 2025 | (midshaft OR middle third or mid-shaft) AND (clavicle OR clavicles OR clavicular) AND (fracture OR fractures) AND (surgical OR nonsurgical OR operative OR nonoperative OR open reduction internal fixation OR intramedullary nailing OR intramedullary nail OR plate OR plating) |
| Topic 3 Arthroplasty - Dislocation after posterior-approach primary THA | July 19, 2025 | (total hip arthroplasty OR total hip replacement OR THA OR THR) AND (posterior OR posterior approach) AND (dislocation OR instability OR prosthetic dislocation OR hip dislocation) |
Objective 1: Identifying a Niche Patient Population (Sports Medicine)
The first objective was to have GPT-5 and human reviewers identify abstracts that report on clinical outcomes after revision anterior cruciate ligament reconstruction (ACLR) using quadriceps tendon autografts. PubMed was queried on July 1, 2025, yielding 668 total results. Inclusion criteria were:
Patients undergoing revision ACLR.
Patients receiving a quadriceps tendon autograft.
Primary studies involving human patients, and reporting on any clinical outcomes from treatment.
> 5 human patients.
Exclusion criteria were:
Studies not reporting on revision ACLR with a quadriceps tendon autografts.
Studies not specifying clinical outcomes in patients undergoing revision ACLR with quadriceps tendon autografts.
Systematic reviews, meta-analyses, book chapters, editorials, commentaries, surgical technique papers without patient data, and biomechanical, animal, or cadaveric studies.
Objective 2: Identifying Specific Study Designs (Trauma)
The second objective was to have GPT-5 and human reviewers identify randomized controlled trials (RCTs) comparing operative and nonoperative management of mid-shaft clavicle fractures. PubMed was queried on July 12, 2025, yielding 1,041 total results. Inclusion criteria were:
Human patients with clavicle fractures specifically located in the middle third or mid-shaft of the bone.
RCTs comparing surgical fixation of the fracture with non-surgical or conservative management.
Reporting on clinical outcomes.
Exclusion criteria were:
Unclear location of the clavicle fracture or focus on distal-third or proximal-third fractures.
Studies only comparing two surgical techniques or two non-operative techniques.
Studies that are secondary analyses of previous RCTs or protocols of RCTs.
Any non-randomized primary design.
Systematic reviews, meta-analyses, book chapters, editorials, commentaries, surgical technique papers without patient data, and biomechanical, animal, or cadaveric studies.
Objective 3: Identifying Specific Outcomes (Arthroplasty)
The third objective was to have GPT-5 and human reviewers identify abstracts that report on postoperative dislocation after posterior-based approaches to primary total hip arthroplasty (THA). PubMed was queried on July 19, 2025, yielding 1,131 total results. Inclusion criteria were:
Patients undergoing primary THA.
A posterior-based approach was used either as the sole approach or as one of multiple approaches with posterior results reported separately.
Dislocation reported as one outcome for the posterior approach cohort or posterior subgroup within a mixed-group study.
Study must be a primary clinical research article involving actual human patients reporting on clinical outcomes.
> 10 patients undergoing posterior-based THA, either if stated in abstract or if it can be reasonably be inferred (e.g. X patients were randomized into two groups, with one group being the posterior group).
Exclusion criteria were:
Patients undergoing hemiarthroplasty, revision THA, or conversion THA.
Studies not reporting on outcomes for posterior-based THA.
Dislocation not reported as an outcome after posterior-based THA.
Systematic reviews, meta-analyses, book chapters, editorials, commentaries, surgical technique papers without patient data, and biomechanical, animal, or cadaveric studies.
Human Review Process
Two human reviewers independently screened each title and abstract for eligibility based on the previously outlined inclusion and exclusion criteria on Covidence online software (Veritas Health Innovation, Melbourne, Australia). Discrepancies between the two reviewers were flagged as conflicts. Pre-conflict resolution inter-rater agreement was calculated using Cohen’s kappa with 95% confidence intervals (CI).
Pre-Processing of Abstracts
For each topic, database exports from PubMed searches described above were downloaded and parsed using a custom Python script. The script simplified and organized the export data by PMID, DOI, title, abstract, publication date, first author, and journal into a plain-text structure, cleaning up formatting artifacts such as line breaks or continuation spaces. This step was critical in ensuring that the LLM received a clean, uniform text without extra meta-data.
GPT-5 Screening
A Python-based NLP script was developed by the primary author to screen titles and abstracts using GPT-5 (OpenAI, San Francisco, California) via the OpenAI Application Programming Interface (API).
The above text file generated from the pre-processing phase was parsed by the NLP script. The NLP script read a “prompt” that consisted of the inclusion and exclusion criteria to create decisions on each abstract within the pre-processed text file. System prompts can be found in Table 2.
Table 2.
System prompts provided to GPT-5
| Prompt | |
|---|---|
| Topic 1: Sports Medicine - Clinical outcomes after revision ACLR with quadriceps tendon autograft |
““"You are helping to screen studies for a systematic review. Here are the inclusion criteria: 1. The study must involve **patients who underwent revision anterior cruciate ligament (ACL) reconstruction**, defined as a **second or subsequent ACL surgery performed after failure of a previous ACL reconstruction**. It is **not sufficient** for the abstract to merely mention revision ACL as a possible indication or future application — the study must include actual patients who underwent revision ACL reconstruction. 2. The study must involve **patients who received ACL reconstruction using a quadriceps tendon graft**, described as “quadriceps tendon”, “quad tendon”, or “quadriceps autograft”. It is **not sufficient** for the abstract to simply mention the technique or its suitability — it must report actual use of this graft in patients. 3. The study must be a **primary clinical research article involving actual human patients**. It must report clinical outcomes or complications from treatment. Exclude reviews, systematic reviews, meta-analyses, book chapters, editorials, commentaries, surgical technique papers without patient data, and biomechanical or cadaveric studies. 4. The study must report on **more than 5 human patients** (i.e., exclude case reports or small case series with 5 or fewer patients). 5. The study must involve **human subjects only**. Based on the abstract below, respond with a one-word decision: **YES** if all inclusion criteria are met, **NO** if not. ““” |
| Topic 2: Trauma - RCTs comparing clinical outcomes after operative vs. nonoperative management of midshaft clavicle fractures |
““You are helping to screen studies for a systematic review. INCLUDE (YES) if ALL of the following are CLEARLY TRUE based on the abstract: 1) POPULATION: - Human patients with clavicle fractures specifically located in the **middle third (midshaft)** of the bone. - The abstract must **explicitly** state this location using any reasonable wording (e.g., “midshaft”, “mid-shaft”, “middle third”, “middle-third”). - EXCLUDE if location is unclear, unspecified, or focuses only on distal-third or proximal-third fractures. 2) INTERVENTION & COMPARATOR: - Surgical fracture fixation (e.g., plating, open reduction internal fixation, intramedullary nailing) **versus** non-surgical or conservative management (e.g., sling, bracing, immobilization, non-operative care). - EXCLUDE if it compares two surgical techniques (e.g., plate vs. nail, locking vs. non-locking plate) or two non-operative approaches. 3) STUDY DESIGN: - Randomized Controlled Trial (RCT) with patients randomized **within the study described in the abstract**. - Accept wording such as “patients were randomized” or “prospective randomized study” only if it is clear that randomization occurred as part of the present study. - EXCLUDE if: • It uses or reanalyzes data from a previous RCT without new patient randomization. • It is a secondary analysis, subgroup study, post hoc analysis, follow-up of a prior RCT, registry, or database study. • It is an RCT protocol describing design only (no outcomes). • It is any non-randomized study. 4) OUTCOMES: - Reports **clinical outcomes** (e.g., function, pain, healing, union rate, complications, patient-reported outcome measures). - EXCLUDE if it reports **only** radiographic findings or surgical techniques without clinical outcome data. --- ALWAYS EXCLUDE: - Systematic reviews, meta-analyses, editorials, commentaries, surgical technique papers without patient data, biomechanical studies, cadaveric studies, or animal studies. TIE-BREAK RULE FOR RANDOMIZATION: If the abstract mentions randomization **and** also indicates it is a secondary/post hoc/subgroup/follow-up/re-analysis of a prior trial (or uses a registry/database from a prior trial), answer NO **unless** it clearly states that **new** patients were randomized in this study and clinical outcomes from that new randomization are reported. --- OUTPUT INSTRUCTIONS: Respond with EXACTLY one word: YES (if included) or NO (excluded) |
| Topic 3 Arthroplasty - Dislocation after posterior-approach primary THA |
““"You are helping to screen studies for a systematic review. INCLUDE (YES) if ALL of the following are TRUE based on the abstract: 1) POPULATION SURGICAL PROCEDURE: - Patients underwent PRIMARY total hip arthroplasty (THA) or PRIMARY total hip replacement (THR) - EXCLUDE if patients underwent hemiarthroplasty or revision total hip arthroplasty 2) POPULATION SURGICAL APPROACH: - The posterior approach was used, either as the sole approach OR - as one of multiple approaches with posterior results reported separately OR - where it is reasonably inferrable from the abstract that a posterior-based approach was used (e.g. known synonyms such as “mini-posterior”, “posterolateral”, “Southern”, “Moore”, or description of approach consistent with posterior) 3) OUTCOME: - Dislocation is reported as an outcome for the posterior approach cohort OR for the posterior subgroup within a mixed-approach study - This includes numerical data (rates, number of events) OR qualitative statements indicating that dislocation rates/events were higher or lower in the posterior group - Synonyms you may see include “instability”, “hip dislocation”, “prosthetic dislocation” - EXCLUDE studies that only talk about dislocation but do not specify the rate or events or qualitative description of events or rates 4) STUDY TYPE: - The study must be a **primary clinical research article involving actual human patients, reporting on clinical outcomes - EXCLUDE reviews, systematic reviews, meta-analyses, book chapters, editorials, animal studies, commentaries, surgical technique papers without patient data, and biomechanical or cadaveric studies. 5) PATIENT NUMBER: - The study must have GREATER than 10 patients with patients undergoing posterior-approach total hip arthroplasty - The number of patients in the posterior group may be stated or may be inferred in studies where allocation suggests > 10 posterior patients (e.g. X patients were randomized into two groups, with one group being the posterior group) - EXCLUDE if there is 10 or less patients undergoing posterior-approach total hip arthroplasty or if number of patients in the posterior group can’t be confirmed or reasonably inferred --- --- OUTPUT INSTRUCTIONS: Respond with EXACTLY one word: YES (if included) or NO (excluded) ““” |
ACLR = anterior cruciate ligament reconstruction, RCT = randomized controlled trials, THA = total hip arthroplasty
Gold-Standard Inclusion and Exclusion Decisions
Abstracts included by both humans and by GPT-5 were considered true positives while those excluded were considered true negatives. Both did not undergo adjudication by the third senior human reviewer. All conflicts between human reviewers and all discrepancies between the LLM and humans were adjudicated by the third reviewer. This was used to generate a list of gold standard inclusion and exclusion decisions to test GPT-5 performance.
Model Performance
Each inclusion or exclusion decision made by GPT-5 was treated as a binary classification task: whether the decision was correctly or incorrectly included or excluded relative to the gold-standard. Performance was evaluated with negative predictive value (NPV), positive predictive value (PPV), sensitivity, specificity, accuracy, and F1 score based on true positives (TP), false positives (FP), true negatives (TN), and false negative (FN). A decision was reported as a TP if GPT-5 correctly included the study, while a decision was reported as TN if GPT-5 correctly excluded the study. A FP was considered a decision where GPT-5 incorrectly included a study, while a FN was when GPT-5 incorrectly excluded a study. Performance metrics were reported as percentages and were calculated as following for each review topic:
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
Efficiency Outcomes
To assess efficiency and feasibility, runtime and API usage cost were tracked. The total time elapsed for parsing the abstracts within each review was recorded using timestamps from the beginning to the completion of the batch process. API usage costs were calculated using OpenAI’s available per-token pricing for GPT-5 model as of August 2025. The equation was as follows: cost = (input tokens x 1.25/1,000,000) + (output tokens x 10/1,000,000). A token is the smallest unit of text that the model is able to process or generate.
Results
Objective 1: Identifying a Niche Patient Population (Sports Medicine)
Of the 668 abstracts screened, GPT-5 classified 26 as eligible for inclusion, while two human reviewers classified 19 as eligible before conflict resolution. There were six abstracts with conflicting decisions between the human reviewers, thus the pre-resolution kappa was 0.859 (95%CI, 0.748–0.970).
Among 19 abstracts marked eligible by both human reviewers, GPT-5 included 18. The single abstract excluded by GPT-5 was incorrectly included by reviewers as it described revision ACLR with quadriceps tendon autografts, but only reported on practice trends without clinical outcomes. Of the six conflicts between the human reviewers, GPT-5 included five and excluded one. Upon adjudication by the hird reviewer, the excluded abstract was correctly excluded for focusing on rectus tendon autograft use. Of the five conflicts that were included by GPT-5, four were correct, and one was incorrectly included because the study involved only five patients.
GPT-5 also included three abstracts that both human reviewers had excluded. Adjudication by the third reviewer determined that two of these were correctly included while one was incorrectly included, as there were only two patients who underwent quadriceps tendon revision ACLR.
After final adjudication by the third human reviewer, which served as the gold standard for inclusion/exclusion decisions, 26 abstracts met the inclusion criteria and 642 were excluded. Based on this reference, GPT-5 correctly identified 24 TPs, 642 TNs, with two FPs and no FNs. This corresponded to a sensitivity of 100%, specificity of 99.7%, PPV of 92.3%, NPV of 100%, accuracy of 99.7%, and F1 score of 95.9%. The total cost to screen all 668 abstracts was $1.32USD ($0.0020USD/abstract), and the task was completed in 38.5 min (17.4 abstracts/minute) (Fig. 1; Table 3).
Fig. 1.
Efficacy metrics for GPT-5 in screening titles and abstracts for three different orthopedic systematic review topics in the fields of sports, trauma, and arthroplasty
Table 3.
GPT-5 performance metrics
| Total Number of Abstracts | Sensitivity | Specificity | PPV | NPV | Accuracy | F1 Score | Time to Completion | Cost | |
|---|---|---|---|---|---|---|---|---|---|
| Topic 1: Sports Medicine - Clinical outcomes after revision ACLR with quadriceps tendon autograft | 668 | 100% | 99.7% | 92.3% | 100% | 99.7% | 95.9% | 38.5 min | $1.32USD |
| Topic 2: Trauma - RCTs comparing clinical outcomes after operative vs. nonoperative management of midshaft clavicle fractures | 1,041 | 94.1% | 100% | 100% | 99.9% | 99.9% | 97.0% | 74.8 min | $2.20USD |
| Topic 3 Arthroplasty - Dislocation after posterior-approach primary THA | 1,131 | 99.2% | 98.2% | 93.8% | 99.8% | 98.4% | 96.4% | 174.3 min | $3.73USD |
PPV = positive predictive value, NPV = negative predictive value, USD = United States Dollars, ACLR = anterior cruciate ligament reconstruction, RCT = randomized controlled trials, THA = total hip arthroplasty
Objective 2: Identifying Specific Study Designs (Trauma)
Of the 1,041 abstracts screened, GPT-5 classified 16 as eligible for inclusion, while two human reviewers also classified 16 as eligible before conflict resolution. There were four abstracts with conflicting decisions between the human reviewers, thus the pre-resolution kappa was 0.887 (95%CI 0.777–0.997).
Among 16 abstracts marked eligible by both human reviewers, GPT-5 included 14. Of the two of 16 abstracts that GPT-5 excluded, one was incorrectly included by reviewers as it was described as a surgical technique study as opposed to a true RCT. However, one study was incorrectly excluded by GPT-5, despite the abstract clearly stating that it was an RCT comparing plate osteosynthesis with figure-of-eight bracing. Of the four conflicts between the human reviewers, GPT-5 correctly included two and excluded two upon adjudication by the third reviewer.
After final adjudication by the third human reviewer, which served as the gold standard for inclusion and exclusion decisions, 17 abstracts met the inclusion criteria and 1,024 were excluded. Based on this reference, GPT-5 correctly identified 16 TPs, 1,024 TNs, with no FPs and one FN. This corresponded to a sensitivity of 94.1%, specificity of 100%, PPV of 100%, NPV of 99.9%, accuracy of 99.9%, and F1 score of 97.0%. The total cost to screen all 1,041 abstracts was $2.20USD ($0.0021USD/abstract), and the task was completed in 74.8 min (13.9 abstracts/minute) (Fig. 1; Table 3).
Objective 3: Identifying Specific Outcomes (Arthroplasty)
Of the 1,131 abstracts screened, GPT-5 classified 256 as eligible for inclusion, while two human reviewers also classified 180 as eligible before conflict resolution. There were 64 abstracts with conflicting decisions between the humans, thus the pre-resolution kappa was 0.814 (95%CI 0.770–0.858).
Among 180 abstracts marked eligible by both human reviewers, GPT-5 included 175. Of the five abstracts that GPT-5 excluded, four were incorrectly included by the reviewers, with one being incorrectly excluded by GPT-5. Of the 64 conflicts between the human reviewers, 43 and 15 were correctly included and excluded by the AI, respectively. Of the six remaining conflicts, five were incorrectly included by GPT-5 and one was incorrectly excluded.
GPT-5 also included 33 abstracts that both human reviewers excluded. After adjudication by the third reviewer, twenty-two of these were correctly included by GPT-5 and incorrectly excluded by the human reviewers, while 11 were incorrectly included by GPT-5 and correctly excluded by humans.
Therefore after final adjudication to create the gold standard for inclusion and exclusion decisions, 242 abstracts met the inclusion criteria and 889 were excluded. Based on this reference, GPT-5 correctly identified 240 TPs, 873 TNs, with 16 FPs and two FN. Comparatively, pre-conflict resolution, human reviewers had 22 FNs. The GPT-5 performance metrics included a sensitivity of 99.2%, specificity of 98.2%, PPV of 93.8%, NPV of 99.8%, accuracy of 98.4%, and F1 score of 96.4%. The total cost to screen all 1,041 abstracts was $3.73USD ($0.0036USD/abstract), and the task was completed in 174.28 min (6.5 abstracts/minute) (Fig. 1; Table 3).
Discussion
The primary finding of this analysis was that GPT-5 can achieve excellent performance in automated title and abstract screening for orthopaedic systematic reviews across three distinct objectives, in three different areas of orthopedics. Across all tasks, the sensitivity of GPT-5 ranged from 94.1% to 100.0%, specificity from 98.2% to 100.0%, PPV from 92.3% to 100.0%, NPV 99.8%−100%, and F1 scores from 95.9% to 97.0%. In all objectives, GPT-5 matched or outperformed human reviewers in sensitivity, notably reducing false negatives compared with human reviewers in a broad screening task within the field of arthroplasty focusing on isolating abstracts reporting on a specific outcome. Screening costs were minimal ($1.32-$3.73 USD per review of 668-1,131 abstracts) and processing times were markedly shorter than typical human workflows (38.5–174.28 min), highlighting both performance robustness and operational efficiency.
The findings from this analysis align with emerging literature on the use of LLMs for the purpose of systematic review screening. One previous analysis utilized a custom LLM to perform title and abstract and full-text screening for a systematic review on vitamin D, finding a specificity of 99.6%, a NPV of 100%, and a PPV of 25.6% [9]. The ability to minimize false negatives were grossly similar to the current findings, however, their false positive rate as suggested by the PPV, was much lower [9]. The comparison arm for this analysis was human reviewers using Rayyan AI, which utilizes the current gold-standard for AI in systematic review tasks, where manual review is used to train a machine learning model that labels future studies as “most likely to exclude”, “likely to exclude”, “undecided”, “likely to include”, or “most likely to include” based on what it learned [9]. Another study utilized DistillerAI, another tool that utilizes machine learning to train a model based on human reviewers manually screening a few studies [8]. While this has shown to have high sensitivity and specificity, it still requires users to have to manually screen studies [8], compared to the current described strategy that just requires users to input their inclusion and exclusion criteria. Prior work using a “zero-shot” strategy used GPT-3.5 to do a similar task, but while gaining in efficiency, there were much more frequent false positives and negatives, especially when inclusion criteria was complex [10]. Another review of available tools used to improve systematic review efficiency consisting of 103 studies only reported on one analysis that used an LLM [11]. GPT-5 demonstrated consistent performance, minimizing false negatives and positives across a variety of complex, nuanced screening challenges, highlighting improvements in the model’s reasoning and contextual comprehension. This performance suggests that in the field of orthopedics, LLMs have become advanced enough to potentially become the mainstay of how titles and abstracts are screened for systematic reviews.
Screening is a repetitive, rules-driven, and text-dependent process, making it the ideal candidate for automation with LLMs. Humans often may screen 1000 s of abstracts, introducing the potential for cognitive fatigue or variability in decision making between reviewers, which is eliminated with an LLM. Arguably the biggest benefit of using AI to screen titles and abstracts is its ability to dramatically shorten the process, accelerating timelines for completion of project ideas. This scalability is especially valuable in fields such as orthopedics where there is a rapid increase in publication volumes within the literature [12], and reviews may run the risk of becoming “outdated” before publication if there is a large lapse between date of search and publication. With screening rates as high as 17.4 abstracts per minute, the efficiency of AI far surpasses that of human reviewers. From a practical standpoint, the GPT-5 pipeline is also low-cost, with the most expensive run being $3.73 dollars for 1,131 abstracts. The system requires no specialized machine learning infrastructure beyond access to the OpenAI API and Python, making it accessible to most research groups and even low-resource settings. Moreover, the structured, reproducible decision-making process enhances transparency, as GPT-5’s outputs can be archived and audited in a way that enhances methodological rigour and reproducibility in systematic reviews.
The predominant strength of this study includes the use of gold-standard adjudication to benchmark AI and human performance, the evaluation of three distinct and clinically relevant screening tasks amongst three different orthopedic subspecialties, and the measurement of both accuracy and efficiency outcomes. By incorporating varying levels of inclusion complexity, from population to outcome-specific, the study tested the adaptability of GPT-5 across multiple screening demands and review types. It is expected that these results would be replicated for reviews in a variety of different medical and surgical specialties, however future work is encouraged to confirm this generalizability. Future work is also encouraged to integrate full-text screening capabilities and evaluate the impact of prompt engineering strategies on model performance. Unfortunately, given the novelty of LLMs, there is no gold-standard way to design a prompt, and much of the task is dependent on trial and error. However, this study has disclosed the full prompts for each screening task that resulted in strong performance, therefore, can be reliably used to model future prompts for related similar topics. Given these results, it is possible that systematic review title and abstract screening can be done with LLMs, when used via APIs in a controlled setting, with the aim of minimizing hallucinations. While there aren’t established benchmarks for understanding what entails strong AI performance, performance metrics that are consistently above 90%, specifically for sensitivity and specificity suggest readiness for transition to AI-based screening without human validation. Author recommendations for using AI for title and abstract screening are listed in Appendix Table 4.
Table 4.
Author's recommendations for using artificial intelligence and large language models for automated systematic review screening
| Domain | Recommendation and Implementation Notes |
|---|---|
| Goals | Optimize for sensitivity, would prefer to minimize false negatives (missing studies) as opposed to minimizing false positives |
| Model Optimization | Choosing the correct model is important. Models focusing on thinking and reasoning are preferred over models focused on speed |
| Prompt Optimization | Detailed prompting is essential. Detailed explanations of what should be included and excluded for each stage of the PICOS (population, intervention, comparison outcome, study design) framework helps maximize performance. |
| API Usage | To create a controlled environment, usage of an API is important as opposed to inputting searches into a large language model web-platform |
There are few limitations from this analysis. First, the study focused exclusively on orthopedic topics and PubMed-derived abstracts. Performance is expected to be similar using abstracts from alternative databases, however, a system that integrates abstracts from multiple databases and removes duplicates prior to passing them through the AI is necessary to mirror existing platforms such as Covidence and Rayyan. Second, an analysis on how different prompts affect model performance was not performed as it was not the focus of the paper. However, in order to standardize using AI for systematic review screening, understanding how to correctly prompt engineer to optimize performance is essential. Third, only three orthopedic subspecialties were chosen for this analysis, despite multiple others within the field (e.g. spine, pediatrics, hand, etc.). However, it is highly unlikely that there would be differences in model performance based on subspecialty given the equivalent performance between the three chosen ones. Finally, while the number of FNs were low, it is unclear what led to GPT-5 wrongly excluding certains ones.
Conclusion
GPT-5 demonstrated exceptional accuracy, sensitivity, specificity, PPV, NPV, and F1 scores in automating title and abstract screening for three orthopedic systematic review topics in three different subspecialties, focusing on identifying a specific patient population, study designs, and outcomes. The average rate of screening ranged from 6.5 to 17.4 abstracts per minute and the average price ranged from $0.002-$0.0036USD per abstract, suggesting a high degree of efficiency compared to current standards. Current LLMs offer the ability to transform the way systematic reviews are performed.
Key References
- Affengruber L, van der Maten MM, Spiero I, Nussbaumer-Streit B, Mahmić-Kaknjo M, Ellen ME, et al. An exploration of available methods and tools to improve the efficiency of systematic review production: a scoping review. BMC Med Res Methodol. 2024:24:210.https://doi.org/10.1186/s12874-024-02320-4.
- ○ A comprehensive systematic review outlining modern tools used to automate systematic reviews, highlighting how few studies have utilized chatGPT for the purpose of title and abstract screening.
- Zsidai B, Kaarre J, Hilkert A-S, Narup E, Senorski EH, Grassi A, et al. Accelerated evidence synthesis in orthopaedics—the roles of natural language processing, expert annotation and large language models. J Exp Orthop. 2023:10:99. https://doi.org/10.1186/s40634-023-00662-4.
- ○ Narrative review article highlighting the potential of large language models as a natural language processor in orthopedic research
Appendix
Author Contributions
PV : idea conception, code generation, data analysis, writing, reviewing, editing HS and LB: human screening, writing, reviewing, editing MDB: writing, reviewing, editing ORA and JK: writing, reviewing, editing, supervision.
Funding
No funding was received.
Data Availability
Data can be made upon reasonable request to prushoth.vivekanantha@medportal.ca.
Declarations
Compliance with Ethical Standards
This research did not require ethical approval given the lack of human or animal subjects included.
Competing interests
The authors declare no competing interests.
Human and Animal Rights and Informed Consents
This article does not contain any studies with human or animal subjects performed by any of the authors.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Nussbaumer-Streit B, Ellen M, Klerings I, Sfetcu R, Riva N, Mahmić-Kaknjo M, et al. Resource use during systematic review production varies widely: a scoping review. J Clin Epidemiol. 2021;139:287–96. 10.1016/j.jclinepi.2021.05.019. [DOI] [PubMed] [Google Scholar]
- 2.Sambunjak D, Franić M. Steps in the undertaking of a systematic review in orthopaedic surgery. Int Orthop. 2012;36:477–84. 10.1007/s00264-011-1460-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Beller E, Clark J, Tsafnat G, Adams C, Diehl H, Lund H, et al. Making progress with the automation of systematic reviews: principles of the international collaboration for the automation of systematic reviews (ICASR). Syst Rev. 2018;7:77. 10.1186/s13643-018-0740-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Tsafnat G, Glasziou P, Choong MK, Dunn A, Galgani F, Coiera E. Systematic review automation technologies. Syst Rev. 2014;3:74. 10.1186/2046-4053-3-74. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Kolaski K, Logan LR, Ioannidis JPA. Guidance to best tools and practices for systematic reviews. JBJS Rev. 2023;11:e23.00077. 10.2106/JBJS.RVW.23.00077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Syriani E, David I, Kumar G. Screening articles for systematic reviews with ChatGPT. J Comput Lang. 2024;80:101287. 10.1016/j.cola.2024.101287. [Google Scholar]
- 7.Li M, Sun J, Tan X. Evaluating the effectiveness of large language models in abstract screening: a comparative analysis. Syst Rev. 2024;13:219. 10.1186/s13643-024-02609-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Burns JK, Etherington C, Cheng-Boivin O, Boet S. Using an artificial intelligence tool can be as accurate as human assessors in level one screening for a systematic review. Health Inf Libr J. 2024;41:136–48. 10.1111/hir.12413. [DOI] [PubMed] [Google Scholar]
- 9.Trad F, Yammine R, Charafeddine J, Chakhtoura M, Rahme M, El-Hajj Fuleihan G, et al. Streamlining systematic reviews with large language models using prompt engineering and retrieval augmented generation. BMC Med Res Methodol. 2025;25:130. 10.1186/s12874-025-02583-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Issaiy M, Ghanaati H, Kolahi S, Shakiba M, Jalali AH, Zarei D, et al. Methodological insights into ChatGPT’s screening performance in systematic reviews. BMC Med Res Methodol. 2024;24:78. 10.1186/s12874-024-02203-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Affengruber L, van der Maten MM, Spiero I, Nussbaumer-Streit B, Mahmić-Kaknjo M, Ellen ME, et al. An exploration of available methods and tools to improve the efficiency of systematic review production: a scoping review. BMC Med Res Methodol. 2024;24:210. 10.1186/s12874-024-02320-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Zsidai B, Kaarre J, Hilkert A-S, Narup E, Senorski EH, Grassi A, et al. Accelerated evidence synthesis in orthopaedics—the roles of natural language processing, expert annotation and large language models. J Exp Orthop. 2023;10:99. 10.1186/s40634-023-00662-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Data can be made upon reasonable request to prushoth.vivekanantha@medportal.ca.







