Skip to main content
Journal of Managed Care & Specialty Pharmacy logoLink to Journal of Managed Care & Specialty Pharmacy
. 2026 Sep;32(9):1062–1075. doi: 10.18553/jmcp.2026.32.9.1062

Comparative performance of large language models for appraising bias in real-world evidence studies

Chijioke M Okeke 1,2, Regina Nechi 3, Javeria Khalid 2, Ibraheem M Karaye 4, J Douglas Thornton 1,2, Ismaeel Yunusa 5,6,
PMCID: PMC13505367  PMID: 42640121

Abstract

BACKGROUND:

Real-world evidence (RWE) is increasingly used to inform regulatory and payer policy decisions and health technology assessment, yet appraising the methodological credibility of RWE studies remains time-intensive and requires specialized expertise. The appraisal task could involve using an appraisal tool that provides a structured approach for evaluating bias in observational studies of comparative effectiveness and safety. Large language models (LLMs) may offer a scalable means to support this appraisal process, but their performance on structured bias assessment tasks has not been fully characterized.

OBJECTIVE:

To compare the performance of LLMs from 6 major artificial intelligence (AI) technology providers against human expert assessments in appraising bias in published RWE studies using the Appraisal of Potential Bias in Real-World Evidence Studies framework.

METHODS:

We conducted a comparative diagnostic accuracy study evaluating 40 LLMs from OpenAI, Anthropic, Google, xAI, Meta, and DeepSeek. Ten published RWE studies representing diverse pharmacoepidemiological designs and data sources were appraised by each LLM using a structured chain-of-thought prompt with conditional rubric injection based on the Appraisal of Potential Bias in Real-World Evidence Studies framework. Two independent human reviewers with pharmacoepidemiology training evaluated each study, with a third adjudicator resolving disagreements to establish the reference standard. LLM performance was assessed using overall accuracy and macro-averaged precision, recall, and F1 scores. Assessment time was compared between models and benchmarked against human reviewers. Bootstrap method was used to construct 95% CI for performance measures.

RESULTS:

Across 280 item-level assessments per model (10 studies × 28 items), overall accuracy ranged from 12.9% to 66.1%. The highest-performing model was Claude-Sonnet-4.6 (66.1%), followed by o3 (65.4%) and Gemini-3.1-pro-preview (65.0%). Macro-averaged F1 scores ranged from 30.9% to 66.9%; o3 achieved the highest F1 score (66.9%), followed by GROK-4 (65.5%) and Gemini-3.1-pro-preview (65.4%). Human reviewers required an average of 61.05 minutes per study; all LLMs completed assessments substantially faster, with average time per study ranging from 0.80 to 17.22 minutes relative to humans.

CONCLUSIONS:

LLMs hold considerable promise for automating methodological appraisal of RWE studies; however, their performance is variable and model dependent. Their greatest value may lie in enhancing efficiency and supporting human-led appraisal as decision-support tools rather than replacing expert review. Future research should assess performance across larger, more diverse RWE study collections and evaluate output reproducibility across repeated runs.

Plain language summary

People who decide how medical treatments are used and paid for depend on real-world evidence (RWE) studies, which show how well treatments work. If studies were not done right, people making these decisions might base decisions on incorrect information. Making sure studies were done well is important, but that takes a lot of time and special skills. Artificial intelligence tools we tested may help experts check studies faster but cannot do it without experts.

Implications for managed care pharmacy

Payers increasingly rely on RWE in formulary decision-making, coverage determinations, and health technology assessments, yet examining the methodological credibility of these studies is time-consuming. Our study tested 40 artificial intelligence tools that may help with RWE quality evaluation. Although the tools performed faster than human reviewers, human oversight remains essential because the tools’ accuracy is inconsistent. Pharmacy and outcomes research teams may consider them as decision-support tools for preliminary triage of RWE submissions.


Real-world evidence (RWE) is assuming a central role in regulatory and health technology assessment (HTA) decision-making. In December 2025, the US Food and Drug Administration removed a longstanding requirement that sponsors submit identifiable individual patient data when using real-world data sources, clearing the path for deidentified databases, including insurance claims, electronic health records, and national registries, to support drug and device applications.1 In parallel, the Academy of Managed Care Pharmacy updated its Format for Formulary Submissions in 2024 to require real-world evidence in product dossiers and, in 2025, published payer-focused standards for evaluating RWE in coverage, formulary, and reimbursement decisions.24 These datasets can encompass millions of patients and make it feasible to study rare exposures, uncommon outcomes, and long-term safety signals that traditional randomized trials are ill-suited to capture.57 At the same time, the absence of randomization means that selection bias, confounding, misclassification, and incomplete documentation remain persistent threats to validity.810

As the volume and complexity of RWE submissions grow, regulators, HTA bodies, payers, and journal reviewers face a practical bottleneck: appraising the methodological credibility of these studies is time-intensive and demands specialized expertise.11,12 The Appraisal of Potential Bias in Real-World Evidence Studies (APPRAISE) framework, developed by a working group of the International Society for Pharmacoepidemiology in collaboration with HTA experts, addresses the standardization side of this problem.13 APPRAISE comprises 28 structured items across 9 bias domains and yields transparent, reproducible assessments of study validity. However, applying it still requires careful reading, domain-by-domain judgment, and reconciliation among reviewers, a workflow that does not scale easily when large numbers of RWE submissions require appraisal.13

Large language models (LLMs), deep learning systems trained on vast datasets, have demonstrated expanding capabilities in health care–related tasks, from clinical decision support to pharmacovigilance signal detection and peer-review assistance.1418 Their ability to process lengthy documents, extract methodological details, and reason about study design makes them plausible candidates for structured bias appraisal. However, LLMs differ considerably in architecture, parameter count, training data, and reasoning strategy, and whether these differences translate into meaningful variation in the quality of bias appraisal has not been systematically examined. To date, most investigations of LLMs in evidence evaluation have focused on automating citation screening in systematic reviews, with comparatively little attention to the more cognitively demanding task of structured methodological assessment.1924 To address this gap, we evaluated LLMs from 6 major artificial intelligence (AI) technology providers (OpenAI, Anthropic, Google, xAI, Meta, and DeepSeek) in their capacity to conduct APPRAISE-based bias assessments of published RWE studies, benchmarked against a human-adjudicated reference standard.

Methods

STUDY DESIGN

We conducted a comparative diagnostic accuracy study comparing LLM-generated bias assessments of published RWE studies against a human-adjudicated reference standard. The study did not involve human subjects; the University of South Carolina Institutional Review Board deemed it exempt from review.

APPRAISAL FRAMEWORK

All assessments were conducted using the APPRAISE framework, which comprises 28 structured items organized into 3 primary bias domains (design-related, misclassification, and confounding), further categorized into 9 subdomains: (1) time-related bias, (2) selection bias, (3) exposure misclassification, (4) outcome misclassification, (5) reverse causation, (6) detection bias, (7) informative censoring, (8) unmeasured and residual confounding, and (9) other sources of bias.13 Each item requires the assessor to evaluate whether a specific source of bias is present, absent, or uncertain, based on information reported in the study. Items 1 and 2 capture descriptive study characteristics; items 3 through 28 address specific bias domains and were used as the primary basis for classification performance metrics.

SELECTION OF REAL-WORLD EVIDENCE STUDIES

Ten published RWE studies were selected as test cases for bias appraisal.2534 We selected studies using a purposive sampling approach designed to maximize methodological heterogeneity relevant to bias appraisal. Specifically, studies were chosen to vary across 4 key dimensions: (1) observational design type (eg, retrospective cohort and,medical record or chart review); (2) data source category (eg, administrative claims, electronic health records, national disease registries); (3) sample size range (spanning smaller to large populations); and (4) analytic approaches to confounding (eg, propensity score matching, inverse probability weighting, and multivariable regression). This diversity was intended to expose models to varied sources of potential bias across different study contexts (Table 1).

TABLE 1.

Characteristics of Assessed Real-World Evidence Studies

Article number Study title Author, year Study design Dataset used Outcome assessed (safety/effectiveness) Sample size
1 Comparative effectiveness and safety of single inhaler triple therapies for chronic obstructive pulmonary disease: new user cohort study Feldman et al, 2024 Retrospective cohort Optum Clinformatics Data Mart Safety and effectiveness 87,751
2 Glucagon-like peptide-1 receptor agonists and incidence of dementia among older adults with type 2 diabetes Inoue et al, 2025 Retrospective cohort (target trial emulation) Medicare fee-for-service Safety 14,295
3 Stroke and myocardial infarction with contemporary hormonal contraception: real-world, nationwide, prospective cohort study Yonis et al, 2024 Retrospective cohort Danish national registry Safety 2,025,691
4 Risk of major congenital malformations associated with first- trimester exposure to topical antifungal medications: a large claims database study Kunitoki et al, 2025 Retrospective cohort JMDC Claims Database Safety 12,472
5 Clinical and safety outcomes with GLP-1 receptor agonists and SGLT2 inhibitors in type 1 diabetes: a real-world study Edwards et al, 2022 Retrospective medical record review UT Electronic Medical Records Safety and efficacy 104
6 Association of proton pump inhibitors with risk of dementia- a pharmacoepidemiological claims data analysis Gomm et al, 2016 Retrospective cohort Claims Database Safety 218,493
7 Association between use of non–vitamin K oral anticoagulants, with and without concurrent medications, and risk of major bleeding in nonvalvular atrial fibrillation Chang et al, 2017 Retrospective cohort Taiwan National Health Insurance Database Safety 91,330
8 Outcomes of anastrozole, letrozole, and exemestane in patients with postmenopausal breast cancer Dumas et al, 2025 Retrospective cohort (target trial emulation) French National Medicoadministrative Data Efficacy 148,436
9 Alzheimer disease treatment with acetylcholinesterase inhibitors and incident age-related macular degeneration Sutton et al, 2024 Retrospective cohort Veteran Affairs Database Safety 21,823
10 ADHD drug treatment and risk of suicidal behaviours, substance misuse, accidental injuries, transport accidents, and criminality: emulation of target trials Zhang et al, 2025 Retrospective cohort (target trial emulation) Sweden National Registries Safety 148,581

HUMAN REFERENCE STANDARD

Two independent reviewers with training in pharmacoepidemiology appraised each study at the item level using the framework. In cases of disagreement, a third reviewer served as adjudicator, reviewing the study and both initial assessments before rendering a final determination. The adjudicated responses constituted the reference standard against which all LLM outputs were compared.

LLM SELECTION AND CONFIGURATION

We evaluated 40 LLMs from 6 AI technology providers, selected to represent a broad spectrum of architectures, parameter scales, and capability tiers, ranging from flagship reasoning models to lightweight, efficiency-oriented variants.35 The OpenAI models spanned the GPT-5, GPT-4.1, GPT-4o, GPT-4 Turbo, and o-series families, including flagship models (GPT-5.4, GPT-5.2, GPT-5.1), compact variants (GPT-5.4-mini, GPT-5.4-nano, GPT-4.1 mini, GPT-4.1 nano, GPT-4o mini), general-purpose multimodal models (GPT-4o, GPT-4 Turbo), and reasoning-focused models (o3, o3-mini, o4-mini).36,37 Anthropic models covered the Claude 4.6, Claude 4.5, Claude 4.1, and Claude 4 families, including high-capacity (Opus), balanced (Sonnet), and efficiency-oriented (Haiku) variants (eg, Claude-Opus-4.6, Claude-Sonnet-4.6, Claude-Opus-4.5, Claude-Sonnet-4.5, Claude-Haiku-4.5, Claude-Opus-4.1, Claude-Sonnet-4).38 For xAI, we included Grok models across the 3 and 4 series, including beta and newer configurations (eg, Grok-3-beta, Grok-3 mini-beta, Grok-4, Grok-4.1 FR/FNR, and Grok-4.2 FR/FNR, as well as Grok-code-fast 1).39 From Google, we evaluated Gemini models spanning the 2.5 and preview 3.x families, including both high-capability and lightweight variants (eg, Gemini-2.5-pro, Gemini-2.5-flash, Gemini-2.5-flash-lite, Gemini-3.1-pro-preview, Gemini-3.1-flash-lite-preview, Gemini-3-flash-preview).40 For Meta, we included Llama-series models with varying parameter scales and instruction-tuned configurations (eg, Llama-3.3-70B-versatile, Llama-3.1-8B-instant, Llama-4-scout-17B-16E-instruct).41 Lastly, from DeepSeek, we included the chat and reasoner variants, representing conversational and reasoning-optimized architectures, respectively.42

ASSESSMENT PROCEDURE

Each of the 10 studies was submitted to every LLM for appraisal using a 3-layer prompting architecture designed to elicit structured, evidence-based reasoning. First, a system prompt established the model’s role as an expert research methodologist specializing in bias assessment using the APPRAISE framework, with instructions to provide evidence-based assessments grounded in the study text rather than in general knowledge or the study’s own conclusions. Second, for each of the 28 items, the model received a per-question evaluation prompt containing the question stem, allowable response options, and the study text (truncated to 50,000 characters). The prompt instructed the model to select the most appropriate response option, provide detailed reasoning citing specific methodological features of the study, and flag insufficient information by selecting “Unclear” or “Unsure” where applicable. This structured answer-then-reasoning format constituted the chain-of-thought component of the strategy.

The third layer of prompting pipeline involved conditional injection of the domain-specific interpretive rubric corresponding to the model’s selected response. For example, if the model selected “present” for a time-related bias item, the corresponding domain rubric was immediately introduced. This required the model to evaluate its reasoning against the framework’s considerations before the assessment was finalized. This conditional rubric injection ensured that interpretive guidance varied with each response rather than being presented as static text. We also varied the evaluation prompt by asking the models to interpret the RWE studies using only relevant passages retrieved from a vector database (a system that stores text as searchable numerical representations) rather than the full text of the studies. In this configuration, models were explicitly instructed to base their assessment solely on the retrieved content. The complete prompt templates and all domain-specific rubrics are provided in Supplementary Appendix A (1.9MB, pdf) (available in online article).

Model responses were captured as structured outputs at the item level to facilitate direct comparison with the human reference standard. To operationalize this workflow, we developed an R application, using the Shiny web framework, that automated the pipeline from document upload through text extraction, prompt construction, application programming interfaces (API) submission, response parsing, and structured output generation.43 The application supported models from all 6 AI technology providers through their respective APIs, with available models detected at runtime based on API key availability. The workflow automation application and its outputs are described in further detail in Supplementary Appendix B (1.9MB, pdf) .

OUTCOMES AND STATISTICAL ANALYSIS

The primary outcome was the degree to which each LLM’s item-level responses agreed with the human-adjudicated reference standard. Overall accuracy, defined as the proportion of items for which the LLM and adjudicated human responses matched, provided a global measure of agreement pooled across all 10 studies (with a total of 280 assessments per model). Because global accuracy can obscure how well a model performs across individual response categories, particularly when some categories are more common than others, we also computed precision, recall, and F1 scores averaged equally across all response categories (macro-averaging) for items 3 through 28 (260 assessments per model). These metrics were computed at the item level across items 3 to 28 to account for differences in question structure or applicability. The F1 score, which balances a model’s ability to correctly assign a category with its ability to capture all items belonging to that category, served as the primary measure of classification performance.44,45 Lastly, we compared assessment efficiency by recording the time each model required to complete the full evaluation relative to the average time recorded for human reviewers. We applied bootstrapping approach to estimate 95% CIs for accuracy, precision, recall and F1 measures.46,47 All analyses were conducted using SAS version 9.4.

ADDITIONAL ANALYSES

In addition to pooled performance metrics, we conducted stratified analyses by main bias domain (design-related, misclassification, and confounding) to evaluate domain-specific variation in model performance. We also assessed model performance at the individual article level to examine consistency across studies. Furthermore, we evaluated interrater reliability between the 2 primary human reviewers prior to adjudication using preadjudication reviewer-level data. Agreement was quantified using Cohen’s kappa statistic.48

Results

Across the 10 included studies, 280 APPRAISE item-level assessments were evaluated per model.

ACCURACY

Model performance varied considerably across the evaluated LLMs. Overall accuracy ranged from 12.9% to 66.1%. The highest-performing model was Claude-Sonnet-4.6, which achieved an accuracy of 66.1%, followed by o3 (65.4%) and Gemini-3.1-pro-preview (65.0%). The models demonstrating the lowest accuracy were GPT-4o-mini (12.9%) and GROK-4.2 FNR (13.6%) (Figure 1). Figure 2 shows model accuracy plotted against model release date.

FIGURE 1.

Heatmap of LLM Performance Across Accuracy, Precision, Recall, and F1 Score, Grouped by AI Provider

FIGURE 1

All scores are percentages. Color intensity reflects performance level, with green indicating higher scores and red indicating lower scores. Human assessors serve as reference (average 61.05 minutes/assessment; no numeric performance metrics reported).

AI = artificial intelligence; LLM = large language model.

FIGURE 2.

Accuracy of LLMs Over Time, Stratified by Frontier Type

FIGURE 2

Each point represents a model plotted at its release date and accuracy score. Step-function lines trace the performance frontier for closed-source (yellow) and open-source (blue) models over time. Accuracy is expressed as the percentage of 280 Appraisal of Potential Bias in Real-World Evidence Studies items correctly classified per model.

CLASSIFICATION PERFORMANCE

F1 scores ranged from 30.9% to 66.9%, reflecting considerable variability in balanced classification performance across models. The highest F1 score was achieved by o3 (66.9%), followed by GROK-4 (65.5%), Gemini-3.1-pro-preview (65.4%), and Claude-Sonnet-4.6 (63.3%). Moderate-performing models, including GPT-5.1, o4-mini Gemini-2.5-pro, and GROK-code-fast 1, had F1 scores generally ranging between 58.0% and 59.0%. Lower-performing models, including GROK-3-beta, GROK-4.2 FNR, and GROK-4.2 FR, had F1 scores below 35.0%; the lowest was recorded for GPT-4o mini (30.9%) (Figure 2). Precision ranged from 25.0% to 68.1% and recall from 22.9% to 61.4% across all models. Among the top performers, o3, GROK-4, and Gemini-3.1-pro-preview demonstrated relatively balanced precision and recall, which contributed to their higher F1 scores (Figures 1 and 3).

FIGURE 3.

Precision-Recall Profiles of LLMs, Grouped by AI Provider

FIGURE 3

Each point represents one model. Curved lines denote F1-score iso-contours; dashed lines mark the overall median recall and precision across all 40 models. Quadrant labels indicate the directional interpretation of each region.

AI = artificial intelligence; LLM = large language model.

ASSESSMENT TIMES

Human reviewers required an average of 61.05 minutes to assess each study. All evaluated LLMs were significantly faster, with average time ranging from 0.80 to 17.22 minutes relative to the human benchmark. Most models from OpenAI, Anthropic, Meta, and DeepSeek had very fast assessment times per study. For example, GPT-5.4-nano, GPT-5.4-mini, GPT-4.1, and GPT-4o each had assessment time averaging between 1.57 and 2.44 minutes, and Claude-Opus-4.5 and Claude-Haiku-4.5 averaged between 5.32 and 5.58 minutes. GROK-4 had the slowest time of any evaluated model at an average of 17.22 minutes, while Meta and Gemini models, including Llama-4 Scout, Llama-3.1-8B, Gemini-2.5-flash-lite, and GPT-4.1 nano, achieved the fastest assessment time averaging between 0.80 and 1.57 minutes per study (Figure 4).

FIGURE 4.

Accuracy vs Average Assessment Time, Grouped by AI Provider

FIGURE 4

Each point represents one model. Dashed lines mark the overall median accuracy and median assessment time. Quadrant labels indicate the practical interpretation of each region. Human reviewers averaged 61.05 minutes per assessment.

AI = artificial intelligence.

ADDITIONAL ANALYSIS

Bias domain-stratified analyses demonstrated that model performance was largely consistent across bias domains and aligned with the primary results, with no meaningful deviations observed (Supplementary Appendix C, Table C1 (1.9MB, pdf) ). Similarly, article-level analyses showed consistent model performance across the included studies (Supplementary Appendix C, Tables C2-C5 (1.9MB, pdf) ). Interrater agreement between the 2 human reviewers prior to adjudication was kappa = 0.67 (95% CI = 0.60-0.74), indicating substantial agreement (Supplementary Appendix C, Figure C1 (1.9MB, pdf) ).

Discussion

The substantial variability in performance observed across the 40 evaluated LLMs was not random. It reflects meaningful and largely predictable differences in how these models were designed and trained, and in the volume of text they can process in a single pass. Taken together, the findings suggest that although LLMs offer considerable efficiency advantages over manual appraisal, their accuracy remains mixed and model dependent, with direct implications for how they might be integrated into RWE evaluation workflows.

No evaluated model exceeded 68% overall accuracy, a ceiling that reflects the inherent cognitive demands of structured bias appraisal. Many items require the assessor to infer whether a bias is likely based on what the study authors did not report, a form of reasoning that depends as much on epidemiological expertise as on the study text itself. This challenge is not unique to LLMs. In this study, human reviewers required adjudication by a third reviewer to resolve a few disagreements in a subset of items, which indicates that the task carries a degree of subjectivity even for trained experts. Applying general methodological knowledge to item-level judgments about specific, incompletely reported studies remains an imperfect process for current-generation models.4952 The capacity of a model to process a full study text in a single pass, quantified as its context window size, also constrained performance in ways that varied across the models evaluated.53 Models with larger context windows can simultaneously access methodological details distributed across an entire study, from the design section through the statistical analysis plan and supplementary materials. Those with smaller or less efficiently utilized context windows may lose access to earlier portions of the text when processing later APPRAISE items, leading to inconsistent judgments within the same assessment.54,55

The highest overall accuracy was achieved by Claude-Sonnet-4.6, with a decent performance also observed with Claude-Opus-4.6, both developed by Anthropic, as well as the o3 model from OpenAI. These Anthropic models are trained using a process called Constitutional AI, in which the model is first taught to evaluate and critique its own outputs against an explicit set of principled criteria, and then refined through reinforcement learning from human feedback, a training technique in which human raters score model responses to guide further improvement.56,57 The self-evaluation component of this training process shares structural features with the appraisal task itself. Both require the model to assess information against defined criteria and justify a judgment. This alignment between training procedure and task demands may partly explain why Claude Opus and Sonnet variants performed well. These models also represent Anthropic’s highest-capacity tier in terms of parameter count, computational resources used during training (training compute), and context window size, which collectively support the sustained, multistep reasoning required across the 28 sequential APPRAISE items.

Despite having a high overall accuracy, Claude-Sonnet-4.6 recorded a comparatively lower F1 score than some models with lower overall accuracy. This divergence reflects a phenomenon in classification tasks where some response categories occur more often than others. A model that has learned to favor the most commonly correct response category will accumulate high accuracy on frequent items while systematically underperforming on less common categories. Because the distribution of correct responses across the items is not uniform, models calibrated toward dominant categories can achieve high global accuracy without performing consistently across all bias domains. The strong performance of the o3 model may be explained by their design as reasoning models, which OpenAI describes as being trained with large-scale reinforcement learning on chain-of-thought and intended for multistep analysis across text, code, and images. They are characterized as models built for complex reasoning, technical writing, instruction-following, and problems that require sustained analysis rather than rapid surface-level generation.58

GROK-4 and GROK-code-fast 1 achieved both high overall accuracy and high F1 scores, placing them among the most consistently performing models evaluated. The developer of the Grok model family, xAI, has emphasized truth-directed reasoning as a training objective, with explicit efforts to discourage the model from defaulting to responses that are merely expected or socially preferred.59 In a structured bias appraisal context, this training disposition may reduce the tendency toward response category anchoring, contributing to the more balanced performance observed across APPRAISE domains. Supporting this interpretation, GROK-4 exhibited the smallest time reduction of any evaluated model relative to human reviewers. Rather than indicating inefficiency, this longer per-item processing time is consistent with more deliberate, token-intensive reasoning for each appraisal question.60 GROK-4 also supports a very large context window, enabling it to retain access to the full study text throughout the 28-item assessment without losing earlier methodological details.

The poor-performing models (such as Claude-Haiku-4.5 and GROK-4.2 FR) share a common developmental characteristic. They are efficiency-optimized variants produced through knowledge distillation, a process in which a compact model is trained to replicate the output patterns of a much larger model, rather than learning independently from raw training data.61 Correctly classifying 28 sequentially presented items, each requiring conceptual discrimination between distinct bias types within the same study, demands reasoning depth that distilled models structurally lack. These models also typically operate with smaller context windows than their full-capacity counterparts, meaning that in the context of a 50,000-character study submission, they may have processed only a portion of the relevant text effectively when assessing later items. The performance gradient within Google’s Gemini family, from Gemini-2.5-flash-lite and Gemini-3.1-flash-lite-preview to Gemini-2.5-pro and Gemini-3.1-pro-preview, provides a within-developer illustration of the cost of sacrificing reasoning depth (the number of sequential processing steps a model performs before generating a response) for inference speed. Flash and lite variants achieve faster response times through architectural modifications including reduced attention layers, lower numerical precision, and speculative decoding techniques. These modifications curtail the model’s capacity to integrate information distributed across a long document while simultaneously applying different evaluative rubrics to successive items. Context window differences within the Gemini family have a similar consequence. Full-capacity pro and preview models support much larger context windows than their flash-lite counterparts, meaning they can more reliably retain access to methodological details from early study sections when evaluating bias domains later in the appraisal sequence.

The domain-stratified analyses demonstrated that model performance was stable across key bias domains, including design-related bias, misclassification, and confounding, with results mirroring the primary analyses. This consistency suggests that the models were not significantly disproportionately driven by performance in any single domain, addressing a common concern in aggregate evaluations where strong performance in one category may mask deficiencies in others. Instead, the findings support the robustness of model performance across heterogeneous methodological constructs that require distinct forms of reasoning and evidence appraisal. Similarly, the article-level analyses indicated minimal variability in performance across individual studies, despite differences in study design, data sources, and reporting quality. This consistency reinforces the generalizability of the models’ performance and suggests that their ability to identify and classify bias is not sensitive to study-specific characteristics. These sensitivity analyses findings strengthen confidence in the reliability of the models when applied across diverse pharmacoepidemiologic contexts. Interrater agreement between the 2 human reviewers prior to adjudication was substantial. This level of agreement shows the interpretive complexity of bias evaluation, even among trained reviewers. Importantly, the presence of nonperfect agreement also further demonstrates the importance of adjudication and expert oversight, particularly for nuanced or ambiguous bias domains.

These findings provide some implications. The divergence between overall accuracy and F1 scores across models has direct implications for model selection in applied RWE appraisal, because the 2 metrics capture meaningfully different aspects of classification performance. Where the primary objective is rapid screening of large volumes of RWE submissions to identify the most prevalent methodological concerns, overall accuracy in a high-performing model is the more relevant criterion. Conversely, where the objective is comprehensive bias appraisal in which accurate identification of rare or complex bias domains carries equal weight, as in detailed HTA or regulatory submissions, models demonstrating higher F1 scores, may offer superior epistemic reliability despite lower global accuracy. Selecting models on the basis of the metric most aligned with the intended application is therefore not merely a statistical consideration but a matter of consequential methodological choice.

The efficiency advantage demonstrated across all evaluated LLMs is substantial and consistent with a growing body of evidence on the capacity of LLMs to accelerate structured evidence synthesis tasks.6264 All LLMs completed assessments substantially faster than human reviewers. Notably, the model with the least average time, GROK-4, also achieved among the highest accuracy and F1 scores, suggesting that its longer per-item processing reflects deeper reasoning rather than computational inefficiency. This inverse relationship between speed and classification balance is an important consideration for workflow design. In time-sensitive or high-volume contexts, such as rapid systematic reviews or preliminary regulatory triage, faster models offering moderate accuracy may represent an acceptable operational tradeoff. In high-stakes applications where classification reliability across all bias domains is essential, accepting longer processing times in exchange for more balanced and trustworthy outputs is the more defensible choice. These considerations collectively support a tiered approach to LLM integration in RWE appraisal with high-capacity reasoning models reserved for comprehensive assessments that require sustained reasoning, and smaller, faster models for initial screening where the primary objective is volume reduction. Regardless of the model selected, LLM-generated assessments should be interpreted with human oversight, particularly for items requiring inferential judgment about unreported study features or those for which the APPRAISE rubric accommodates meaningful evaluative discretion.51

STRENGTHS

This study has notable strengths. First, it represents one of the most comprehensive comparative evaluations of LLMs for structured bias appraisal to date, encompassing 40 models from 6 providers and spanning the full spectrum of available capability tiers, from flagship reasoning models to lightweight, efficiency-optimized variants. This breadth enables mechanistic comparisons across architectures and training approaches that narrower evaluations cannot support. Second, the use of APPRAISE as the evaluative framework confers real-world credibility, as it is a peer-reviewed, internationally developed tool actively used in HTA and payer decision-making contexts. Third, the human reference standard was established through a rigorous dual-reviewer process with formal third-party adjudication, which is more methodologically stringent than the single-rater or consensus-only designs used in many prior LLM evaluation studies.

Fourth, performance was assessed across multiple complementary metrics, including overall accuracy, precision, recall, and F1 scores, which together revealed important distinctions in model behavior that would have been obscured by accuracy alone. Fifth, the 10 test studies represented diverse observational designs and data sources, exposing models to varied bias profiles rather than a homogeneous set of methodological challenges. Sixth, the structured 3-layer prompting architecture, including the conditional rubric injection strategy, is fully documented in Supplementary Appendix A (1.9MB, pdf) , supporting direct replication and iterative refinement in future work.

LIMITATIONS

These findings should be interpreted in light of the following limitations. First, although the included studies were selected using a purposive sampling approach to maximize heterogeneity across study designs, analytic methods, data sources, and other key methodological dimensions, the sample may not fully capture the breadth of variability observed across all pharmacoepidemiology studies. As such, the findings should be interpreted with caution, particularly with respect to generalizability beyond the included study contexts. Second, although human adjudication served as the reference standard, human assessments are inherently subject to variability. This was mitigated through a dual-reviewer process with third-party adjudication of disagreements, although some degree of evaluative subjectivity cannot be ruled out. Third, study texts were truncated at 50,000 characters for input to each model. For longer RWE publications, relevant methodological details appearing beyond this threshold were inaccessible to the model during assessment, which may have affected performance on items dependent on information reported later in the study. Fourth, the LLMs evaluated represent specific model versions available at the time of the study. Developers update models continuously, and performance characteristics may differ in subsequent versions. Findings should therefore be interpreted as a point-in-time benchmark rather than a durable characterization of any given provider.

Fifth, the study evaluated final item-level classification agreement but did not assess the validity or coherence of the model’s stated reasoning. Agreement with the reference standard does not confirm that the underlying reasoning pathway was methodologically sound, a distinction that is important for the auditability and trustworthiness of LLM outputs in applied appraisal workflows. Sixth, given the stochastic (ie, probabilistic, nondeterministic) nature of LLM outputs, classification responses may vary across repeated runs using identical prompts, and the reproducibility of assessments was not evaluated. Assessing response stability across iterations remains an important direction for future research. All misclassifications were treated equally regardless of degree of discordance. For example, a model selecting “Unclear” when the reference was “Yes” represents a lesser error than selecting “No,” yet both were penalized identically. Lastly, because the study’s primary aim was to compare LLM performance against a human reference standard using a standardized prompting approach, prompt ablation or sensitivity analyses varying the prompting architecture were not conducted. Evaluating how alternative prompting strategies affect model output is an important question that would benefit from further dedicated investigation. Despite these limitations, our study provides timely evidence on the value of LLMs in supporting structured bias assessment in pharmacoepidemiologic research.

Conclusions

LLMs demonstrate considerable promise for supporting the methodological appraisal of RWE studies, yet their performance remains variable and model dependent. Our findings indicate that model architecture, training approach, and context window capacity jointly determine the suitability of any particular model for structured bias assessment. Although no single model excelled across all performance dimensions, all evaluated models completed assessments substantially faster than human reviewers. The greatest practical value of LLM integration into RWE appraisal therefore appears to lie in efficiency gains. In this role, LLMs are best positioned as decision-support tools that complement expert judgment rather than supplant it, particularly given that no model approached the accuracy levels expected of trained pharmacoepidemiologists. Future research should evaluate performance across larger and more diverse RWE study collections, examine reproducibility across repeated model runs, and explore whether hybrid human-LLM workflows can reliably improve appraisal throughput without compromising methodological rigor.

Disclosures

The authors have no disclosures to report.

References

  • 1.U.S. Food and Drug Administration . FDA Eliminates Major Barrier to Using Real-World Evidence in Drug and Device Application Reviews. FDA News Release. Accessed February 17, 2026. https://www.fda.gov/news-events/press-announcements/fda-eliminates-major-barrier-using-real-world-evidence-drug-and-device-application-reviews
  • 2.AMCP Format for Formulary Submissions 5.0 . J Manag Care Spec Pharm. 2024;30(4-b Suppl):1-64. doi: 10.18553/jmcp.2024.30.4-b.s1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Lockhart CM, Powers E, Sweet B, Gleason PP, Brixner D. AMCP real-world evidence standards: overcoming barriers to using real-world evidence in US payer decision-making. J Manag Care Spec Pharm. 2025;31(12):1230-6. doi: 10.18553/jmcp.2025.25108 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Brixner D, Biskupiak J, Oderda G, et al. Payer perceptions of the use of real-world evidence in oncology-based decision making. J Manag Care Spec Pharm. 2021;27(8):1096-105. doi: 10.18553/jmcp.2021.27.8.1096 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Bartlett VL, Dhruva SS, Shah ND, Ryan P, Ross JS. Feasibility of using real-world data to replicate clinical trial evidence. JAMA Netw Open. 2019;2(10):e1912869. doi: 10.1001/jamanetworkopen.2019.12869 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Liu F, Panagiotakos D. Real-world data: a brief review of the methods, applications, challenges and opportunities. BMC Med Res Methodol. 2022;22(1):287. doi: 10.1186/s12874-022-01768-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Liu J, Barrett JS, Leonardi ET, et al. Natural history and real-world data in rare diseases: applications, limitations, and future perspectives. J Clin Pharmacol. 2022;62(Suppl 2):S38-S55. doi: 10.1002/jcph.2134 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Mc Cord KA, Al-Shahi Salman R, Treweek S, et al. Routinely collected data for randomized trials: promises, barriers, and implications. Trials. 2018;19(1):29. doi: 10.1186/s13063-017-2394-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Lu CY. Observational studies: a review of study designs, challenges and strategies to reduce confounding. Int J Clin Pract. 2009;63(5):691-7. doi: 10.1111/j.1742-1241.2009.02056.x [DOI] [PubMed] [Google Scholar]
  • 10.Morton SC, Costlow MR, Graff JS, Dubois RW. Standards and guidelines for observational studies: quality is in the eye of the beholder. J Clin Epidemiol. 2016;71:3-10. doi: 10.1016/j.jclinepi.2015.10.014 [DOI] [PubMed] [Google Scholar]
  • 11.Makady A, van Veelen A, Jonsson P, et al. Using real-world data in health technology assessment (HTA) practice: a comparative study of five HTA agencies. PharmacoEconomics. 2018;36(3):359-68. doi: 10.1007/s40273-017-0596-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Murphy LA, Akehurst R, Cunningham D, de Pouvourville G, Solà-Morales O. Real-world evidence to support health technology assessment and payer decision making: is it now or never? Int J Technol Assess Health Care. 2025;41(1):e20. doi: 10.1017/S0266462325000145 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Bykov K, Jaksa A, Lund JL, et al. APPRAISE: a tool for appraising potential for bias in real-world evidence studies on medication effectiveness or safety. Value Health. 2025;28(12):1849-56. doi: 10.1016/j.jval.2025.07.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Raiaan MAK, Mukta MSH, Fatema K, et al. A review on large language models: architectures, applications, taxonomies, open issues and challenges. IEEE Access. 2024;12:26839-74. doi: 10.1109/ACCESS.2024.3365742 [DOI] [Google Scholar]
  • 15.Kumar P. Large language models (LLMs): survey, technical frameworks, and future challenges. Artif Intell Rev. 2024;57(10):260. doi: 10.1007/s10462-024-10888-y [DOI] [Google Scholar]
  • 16.Lin C, Kuo CF. Roles and potential of large language models in healthcare: a comprehensive review. Biomed J. 2025;48(5):100868. doi: 10.1016/j.bj.2025.100868 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Yang R, Tan TF, Lu W, Thirunavukarasu AJ, Ting DSW, Liu N. Large language models in health care: development, applications, and challenges. Health Care Sci. 2023;2(4):255-63. doi: 10.1002/hcs2.61 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Sun Z. Large language models in peer review: challenges and opportunities. Scientometrics. 2025;130(10):5503-46. doi: 10.1007/s11192-025-05440-w [DOI] [Google Scholar]
  • 19.Fabiano N, Gupta A, Bhambra N, et al. How to optimize the systematic review process using AI tools. JCPP Adv. 2024;4(2):e12234. doi: 10.1002/jcv2.12234 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Blaizot A, Veettil SK, Saidoung P, et al. Using artificial intelligence methods for systematic review in health sciences: a systematic review. Res Synth Methods. 2022;13(3):353-62. doi: 10.1002/jrsm.1553 [DOI] [PubMed] [Google Scholar]
  • 21.van Dijk SHB, Brusse-Keizer MGJ, Bucsán CC, van der Palen J, Doggen CJM, Lenferink A. Artificial intelligence in systematic reviews: promising when appropriately used. BMJ Open. 2023;13(7):e072254. doi: 10.1136/bmjopen-2023-072254 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Tsou AY, Treadwell JR, Erinoff E, Schoelles K. Machine learning for screening prioritization in systematic reviews: comparative performance of Abstrackr and EPPI-Reviewer. Syst Rev. 2020;9(1):73. doi: 10.1186/s13643-020-01324-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Van der Mierden S, Tsaioun K, Bleich A, Leenaars CHC. Software tools for literature screening in systematic reviews in biomedical research. ALTEX. 2019;36(3):508-17. doi: 10.14573/altex.1902131 [DOI] [PubMed] [Google Scholar]
  • 24.Johnson N, Phillips M. Rayyan for systematic reviews. Journal of Electronic Resources Librarianship. 2018;30(1):46-8. doi: 10.1080/1941126X.2018.1444339 [DOI] [Google Scholar]
  • 25.Edwards K, Li X, Lingvay I. Clinical and safety outcomes with GLP-1 receptor agonists and SGLT2 inhibitors in type 1 diabetes: a real-world study. J Clin Endocrinol Metab. 2023;108(4):920-30. doi: 10.1210/clinem/dgac618 [DOI] [PubMed] [Google Scholar]
  • 26.Wu CY, Alkabbani W, Shah BR, et al. Comparative dementia risk with GLP1 receptor agonists, SGLT2 inhibitors, or DPP4 inhibitors: a population-based cohort study. Alzheimers Res Ther. 2025;17(1):269. doi: 10.1186/s13195-025-01929-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Feldman WB, Suissa S, Kesselheim AS, et al. Comparative effectiveness and safety of single inhaler triple therapies for chronic obstructive pulmonary disease: new user cohort study. BMJ. 2024;387:e080409. doi: 10.1136/bmj-2024-080409 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Inoue K, Saliba D, Gotanda H, et al. Glucagon-like peptide-1 receptor agonists and incidence of dementia among older adults with type 2 diabetes: a target trial emulation. Ann Intern Med. 2025;178(9):1258-67. doi: 10.7326/ANNALS-24-02648 [DOI] [PubMed] [Google Scholar]
  • 29.Shapiro SB, Yin H, Yu OHY, Rej S, Suissa S, Azoulay L. Glucagon-like peptide-1 receptor agonists and risk of suicidality among patients with type 2 diabetes: active comparator, new user cohort study. BMJ. 2025;388:e080679. doi: 10.1136/bmj-2024-080679 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Zhang L, Zhu N, Sjölander A, et al. ADHD drug treatment and risk of suicidal behaviours, substance misuse, accidental injuries, transport accidents, and criminality: emulation of target trials. BMJ. 2025;390:e083658. doi: 10.1136/bmj-2024-083658 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Sutton SS, Magagnoli J, Cummings TH, Hardin JW, Ambati J. Alzheimer disease treatment with acetylcholinesterase inhibitors and incident age-related macular degeneration. JAMA Ophthalmol. 2024;142(2):108-14. doi: 10.1001/jamaophthalmol.2023.6014 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Dumas E, Hamy AS, Wanis KN, et al. Outcomes of anastrozole, letrozole, and exemestane in patients with postmenopausal breast cancer. JAMA Netw Open. 2025;8(12):e2550842. doi: 10.1001/jamanetworkopen.2025.50842 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Chang SH, Chou IJ, Yeh YH, et al. Association between use of non-vitamin K oral anticoagulants with and without concurrent medications and risk of major bleeding in nonvalvular atrial fibrillation. JAMA. 2017;318(13):1250-9. doi: 10.1001/jama.2017.13883 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Gomm W, von Holt K, Thomé F, et al. Association of proton pump inhibitors with risk of dementia: a pharmacoepidemiological claims data analysis. JAMA Neurol. 2016;73(4):410-6. doi: 10.1001/jamaneurol.2015.4791 [DOI] [PubMed] [Google Scholar]
  • 35.Ferreira L, Silva M, Costa T, et al. Bringing foundation models to the edge with efficient deployment strategies. Preprints. Preprint posted online July 26, 2025. doi: 10.36227/techrxiv.175356750.08746502/v1 [DOI]
  • 36.Open AI . Developers. Models. Accessed February 17, 2026. https://developers.openai.com/api/docs/models
  • 37.Roumeliotis KI, Tselikas ND. ChatGPT and Open-AI models: a preliminary review. Future Internet. 2023;15(6):192. doi: 10.3390/fi15060192 [DOI] [Google Scholar]
  • 38.Claude API. Docs. Models overview: Choosing a model. Accessed February 17, 2026. https://platform.claude.com/docs/en/about-claude/models/overview
  • 39.xAI . Models and Pricing. Accessed February 17, 2026. https://docs.x.ai/developers/models
  • 40.Gemini API . Gemini models. Accessed February 17, 2026. https://ai.google.dev/gemini-api/docs/models
  • 41.Meta AI . Models and libraries. Accessed February 17, 2026. https://ai.meta.com/resources/models-and-libraries/
  • 42.DeepSeek API Docs . Lists Models. Accessed February 17, 2026. https://api-docs.deepseek.com/api/list-models
  • 43.Doi J, Potter G, Wong J, Alcaraz I, Chi P. Web application teaching tools for statistics using R and Shiny. Technology Innovations in Statistics Education. 2016;9(1). doi: 10.5070/T591027492 [DOI] [Google Scholar]
  • 44.Yacouby R, Axman D.Probabilistic extension of precision, recall, and F1 score for more thorough evaluation of classification models. In: Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems. Association for Computational Linguistics; 2020:79-91. doi: 10.18653/v1/2020.eval4nlp-1.9 [DOI] [Google Scholar]
  • 45.Hinojosa Lee MC, Braet J, Springael J. Performance metrics for multilabel emotion classification: comparing micro, macro, and weighted F1-scores. Appl Sci (Basel). 2024;14(21):9863. doi: 10.3390/app14219863 [DOI] [Google Scholar]
  • 46.DiCiccio TJ, Efron B. Bootstrap confidence intervals. Stat Sci. 1996;11(3). doi: 10.1214/ss/1032280214 [DOI] [Google Scholar]
  • 47.Efron B. Better bootstrap confidence intervals. J Am Stat Assoc. 1987;82(397):171-85. doi: 10.1080/01621459.1987.10478410 [DOI] [Google Scholar]
  • 48.Warrens MJ. New interpretations of Cohen’s kappa. J Math. 2014;2014:1-9. doi: 10.1155/2014/203907 [DOI] [Google Scholar]
  • 49.Bommasani R, Hudson DA, Adeli E, et al. On the opportunities and risks of foundation models. arXiv. Preprint posted online 2021. doi: 10.48550/ARXIV.2108.07258 [DOI]
  • 50.Bechny M, Fiorillo L, van der Meer J, et al. Beyond accuracy: a framework for evaluating algorithmic bias and performance, applied to automated sleep scoring. Sci Rep. 2025;15(1):21421. doi: 10.1038/s41598-025-06019-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Sokolova M, Lapalme G. A systematic analysis of performance measures for classification tasks. Inf Process Manage. 2009;45(4):427-437. doi: 10.1016/j.ipm.2009.03.002 [DOI] [Google Scholar]
  • 52.Liang P, Bommasani R, Lee T, et al. Holistic evaluation of language models. arXiv. Published online 2022. doi: 10.48550/ARXIV.2211.09110 [DOI] [PubMed]
  • 53.Iqbal A, Shahid A, Roman M, Afzal MT, Hassan UU. Optimising window size of semantic of classification model for identification of in-text citations based on context and intent. PLoS One. 2025;20(3):e0309862. doi: 10.1371/journal.pone.0309862 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Ratner N, Levine Y, Belinkov Y, et al. Parallel context windows for large language models. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics; 2023:6383-402. doi: 10.18653/v1/2023.acl-long.352 [DOI] [Google Scholar]
  • 55.Peng B, Quesnelle J, Fan H, Shippole E. YaRN: efficient context window extension of large language models. arXiv. Preprint posted online 2023. doi: 10.48550/ARXIV.2309.00071 [DOI]
  • 56.Torgbi Agbemabiese W. Toward constitutional autonomy in AI systems: a theoretical framework for aligned agentic intelligence. IEEE Access. 2026;14:11385-402. doi: 10.1109/ACCESS.2026.3654907 [DOI] [Google Scholar]
  • 57.Huang S, Siddarth D, Lovitt L, et al. Collective constitutional AI: aligning a language model with public input. In: The 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM; 2024:1395-417. doi: 10.1145/3630106.3658979 [DOI] [Google Scholar]
  • 58.Open AI . Developers: Models. Accessed April 14, 2026. https://developers.openai.com/api/docs/models/o3
  • 59.Rayner M, Chatgpt C-Lara-Instance . How woke is Grok? Empirical evidence that xAI’s Grok aligns closely with other frontier models. Published online 2025. doi:10.13140/RG.2.2.35804.45449
  • 60.Wilhelm P, Wittkopp T, Kao O.Beyond test-time compute strategies: advocating energy-per-token in LLM inference. In: Proceedings of the 5th Workshop on Machine Learning and Systems. ACM; 2025:208-15. doi: 10.1145/3721146.3721953 [DOI] [Google Scholar]
  • 61.Alkhulaifi A, Alsahli F, Ahmad I. Knowledge distillation in deep learning and its applications. PeerJ Comput Sci. 2021;7:e474. doi: 10.7717/peerj-cs.474 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Marshall IJ, Kuiper J, Wallace BC. RobotReviewer: evaluation of a system for automatically assessing bias in clinical trials. J Am Med Inform Assoc. 2016;23(1):193-201. doi: 10.1093/jamia/ocv044 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Wang Z, Cao L, Danek B, Jin Q, Lu Z, Sun J. Accelerating clinical evidence synthesis with large language models. NPJ Digit Med. 2025;8(1):509. doi: 10.1038/s41746-025-01840-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Gartlehner G, Kahwati L, Nussbaumer-Streit B, et al. From promise to practice: challenges and pitfalls in the evaluation of large language models for data extraction in evidence synthesis. BMJ Evid Based Med. 2025;30(6):385-9. doi: 10.1136/bmjebm-2024-113199 [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Journal of Managed Care & Specialty Pharmacy are provided here courtesy of Academy of Managed Care Pharmacy

RESOURCES