ABSTRACT
Objective
To explain discrepant findings for fluoxetine's efficacy in three influential network meta‐analyses (NMAs) of treatments for pediatric depression, which led to conflicting clinical recommendations.
Design
Critical appraisal and re‐analysis of three published NMAs.
Data Sources
NMAs published in two Lancet journals and by Cochrane, together with the trial datasets reported therein.
Data Synthesis
We compared efficacy estimates for fluoxetine versus placebo across NMAs. We identified and assessed an outlying trial included only in the Lancet NMAs using the INSPECT‐SR instrument, and re‐analyzed the NMAs with and without this trial.
Results
The larger effects reported in the Lancet NMAs (SMD −0.51, 95% CrI −0.99 to −0.03; and −0.51, −0.84 to −0.18) were driven by an inconsistent fluoxetine–placebo–nortriptyline loop, which the original NMA authors could not explain. We identified the cause of the inconsistency as a small outlier trial of fluoxetine versus nortriptyline that reported an implausibly large effect size (SMD > 4) favoring fluoxetine, and which was not included in the Cochrane NMA. Excluding this trial from the Lancet NMA datasets resolved the inconsistency and yielded efficacy estimates for fluoxetine that closely matched the Cochrane NMA (SMD −0.20, 95%CI −0.28 to −0.11). The outlier trial also showed multiple methodological concerns suggesting low trustworthiness.
Conclusion
Discrepancies between the three NMAs were explained by the indirect influence of a single small trial with extreme and unreliable results. Removing this trial reconciled the Lancet NMAs with the Cochrane NMA, yielding a more reliable estimate of fluoxetine efficacy versus placebo. It also resolved the inconsistency. This case illustrates how inclusion of a single small problematic trial can substantially distort the clinically important results of NMAs. Our findings may alter the clinical risk/benefit assessment of fluoxetine for this indication.
Other
No specific funding was involved in the study.
Keywords: adolescents, antidepressants, biases, children, fluoxetine, network meta analysis
Summary
We found that one small clinical trial with untrustworthy results led to an overestimation of how well fluoxetine works in treating depression in children and adolescents.
Removing the untrustworthy trial resolved discrepant findings for how well fluoxetine works in three major reviews that have shaped treatment guidelines.
Without the untrustworthy trial, all three reviews indicate that the difference between fluoxetine and placebo is small and unimportant in reducing depression symptoms in children and adolescents.
Plain Language Summary
Why three major reviews disagreed about fluoxetine for depression in children and why its overall effectiveness was over‐estimated by including an untrustworthy trial.
What Is Known About the Efficacy of Fluoxetine?
Fluoxetine is recommended in guidelines for the treatment of depression in children and adolescents. Three large analyses (called network meta‐analyses) have pooled all the clinical trials of fluoxetine for depression in this population to estimate how well it works. However, these analyses came to different conclusions about how well fluoxetine works. Two of the analyses, both published in Lancet journals, reported that fluoxetine may reduce depression symptoms better than placebo (sugar pills). In contrast, the third analysis, published by Cochrane, indicated that the difference between fluoxetine and placebo was small and unimportant.
What Did We Want to Find Out?
We aimed to explain why the analyses yielded different results for fluoxetine. This is important, because such analyses should report similar results, and doctors and young people might make different decisions about using fluoxetine depending on which result they were guided by.
What Did We Do?
We examined the results of the three analyses and also the results of the individual trials that were included in each of them, to see how each trial informed the overall result.
What Did We Find?
We found one small trial that was included in the two Lancet analyses, but was not included in the Cochrane analysis, which reported an implausibly large effect for fluoxetine. We examined this trial and found other features that indicated that its data were untrustworthy, such as other highly improbable results and numerical inconsistencies.
We repeated all three analyses to reproduce the original results and then repeated the Lancet analyses without the untrustworthy trial. We found that the two Lancet analyses had overestimated how well fluoxetine works because of the inclusion of the suspicious trial with a large effect. Without this trial, the results of the Lancet analyses became similar to those of the Cochrane analysis, with all three analyses indicating that fluoxetine probably makes little to no difference in reducing depressive symptoms. Removing the untrustworthy trial also made the results of the Lancet analyses more robust according to technical criteria for assessing the reliability of findings in research of this kind.
What Were the Reactions to Our Findings?
We sent our concerns to the journal that published the results of the untrustworthy trial, and to the journals that published the analyses that used it, but no action has been taken to correct the original study or the analyses.
What Does This Mean for Research and Clinical Practice?
Earlier reviews have overestimated the efficacy of fluoxetine. In fact, its efficacy in treating depression in young people is small and comparable to that of a placebo. To enable shared decision‐making, clinicians and patients should be informed of fluoxetine's actual efficacy, and treatment guidelines should be updated accordingly.
Our findings also show that in summarizing results of many trials a single small trial can have a major impact on the most clinically important results, and that current practice for doing such analyses does not reliably identify trials with untrustworthy results.
1. Introduction
Treatment recommendations rely on high‐quality and trustworthy evidence and its synthesis in systematic reviews and meta‐analyses. Discrepant findings in evidence syntheses require explanation, particularly if the findings have differing clinical implications. It seems that this problem applies to fluoxetine for the treatment of pediatric depression.
Depression affects a significant proportion of children and adolescents and a recent review estimated the global point prevalence at 8% [1]. Worldwide, pharmacotherapy for this indication is initiated by a range of clinicians, including Family Physicians, Pediatricians, Psychiatrists, and nonmedical prescribers [2]. Antidepressant prescription rates in children and adolescents are ever increasing and fluoxetine is among the most frequently prescribed antidepressants [3, 4, 5, 6], likely because fluoxetine is recommended as the preferred medication in major guidelines [7].
Major guidelines and similar publications [7, 8, 9, 10] base their recommendations for fluoxetine as first‐line medication on two network meta‐analyses (NMAs) published in Lancet group journals by Cipriani et al. [11] and Zhou et al. [12] (Henceforth simply referred to as Cipriani and Zhou). Cipriani reported a standardized mean difference (SMD) for depression symptom reduction for fluoxetine versus placebo of −0.51 (95% Credible Interval [CrI] −0.99 to −0.03). Zhou reported a nearly identical SMD −0.51 (95% CrI −0.84 to −0.18). The confidence in the efficacy findings for fluoxetine was rated as “very low” in both NMAs.
In contrast, the more recent NMA by Hetrick et al. (henceforth simply Hetrick) and published by Cochrane, produced a different result [13]. Efficacy results for depression symptom reduction were restricted to trials using the most common symptom rating scale, the Children's Depression Rating Scale‐Revised (CDRS‐R, range 17 to 113 points), and Hetrick did not report an SMD. Instead they reported a mean difference (MD) of −2.8 points on the CDRS‐R (95% Confidence Interval [CI] −4.12 to −1.56); but using a standard deviation (SD) of 14.471 as suggested by Hetrick, the SMD can be calculated as −0.20 (−0.28 to −0.11). The confidence in the finding was rated as “moderate.”
While the confidence intervals between the Lancet and Cochrane NMAs overlap and statistical significance was found in all three NMAs, the point‐estimate and the lower end of the confidence interval were reduced by approximately half in the Cochrane NMA. Moreover, fluoxetine's efficacy was now clinically equivalent with placebo (even when considering the whole range of the confidence interval) according to the predetermined criteria used by Hetrick, i.e., ±5 CDRS‐R points (approximately equivalent to SMD 0.35). The appropriate range for equivalence can be debated, but Hetrick's threshold is at the smaller end of any such range, both by Hetrick's own criteria (they suggested a range at least 50% of 10–15 points as a minimum) and by common estimates of the minimal important difference in antidepressant research [14]. Consequently, recommendations for fluoxetine as a first line treatment are at odds with the findings of the Cochrane NMA. In fact, the findings arguably support the view that fluoxetine should not be used at all, an argument that the authors acknowledged in their conclusion but refrained from endorsing themselves.
To our best knowledge, these discrepancies have not been acknowledged or explained in the scientific literature, clinical guidelines, or the Cochrane NMA itself, despite some overlap of authorship in the three NMAs.
Furthermore, from a methodological point of view, discrepant findings between NMAs assessing the same treatments require explanation, because NMAs are supposed to be reproducible evidence syntheses. Potential reasons for differences between NMAs include varying inclusion criteria or statistical methods, and errors. For example we have frequently identified what we call the “standard error,” that is, when the standard error is conflated with the standard deviation, leading to erroneously large effect sizes. Another explanation of the discrepancy would be the inclusion in the more recent NMA of more recent studies with smaller effect estimates. Furthermore, more recent trials often have multiple arms, with a novel drug, placebo, and fluoxetine as an established comparator drug. This reduces the probability of receiving placebo and may change expectancy effects and unblinding bias [7]. However, when we investigated this in detail, the effect of more recent trials that were included in the more recent Cochrane NMA was not sufficient to explain the discrepant findings in the three NMAs [7]. Therefore, we aimed to further explore why the fluoxetine‐placebo effect sizes differed in the Lancet NMAs and Hetrick by critically examining and re‐analyzing the NMAs.
2. Methods
2.1. Data Collection
RL, FN, and MP examined the NMAs and sought associated data and code. The complete dataset for Cipriani is freely available through the Oxford University Research Archive. The dataset for Zhou is freely available through Mendeley. For Hetrick, the full dataset was not publicly available. We requested this from the lead author and from Cochrane, but the full dataset had not been retained and only a partial dataset was provided (Sarah Hetrick and Nick Meader, personal communication, 27 May 2025). MP and RL compiled a dataset of the studies that were included in Hetrick, combining the partial data provided by the authors with additional data that we extracted from the Zhou dataset (provided online for details), and following the preferences for data sources and types that were stated by Hetrick.
2.2. Critical Analysis of an Outlier Trial
When a trial with an extreme effect size and various other features that undermined its credibility was identified among the datasets of the NMAs, we formally assessed it using outlier detection tests and the INSPECT‐SR (INveStigating ProblEmatic Clinical Trials in Systematic Reviews) framework [15]. We contacted the authors of the study, the journal editor, and the publisher, to raise our concerns.
2.3. Critical Analysis of the Three NMAs
We adopted a constructivist approach, engaging in a critical analysis of the above materials, followed by a collaborative discussion among three of the authors (RL, MP, FN). We structured our analysis around NMA methodology, network geometry, and inconsistency [16].
2.4. Re‐Analysis of the Three NMAs
2.4.1. Effect Measures
We re‐analyzed efficacy estimates for fluoxetine as assessed by change in depression symptoms with placebo as reference group, because it was for this outcome that we observed discrepancies across the three NMAs. To facilitate comparison, we used SMDs in the synthesis and presentation of results.
2.4.2. Synthesis Methods
One of the authors (MP) undertook a full reanalysis of the original Cipriani dataset, including sensitivity analysis to control for the effect of a single trial with outlying results—see below. When discrepancies persisted between the original NMA results and our reproduced results using the full dataset, MP involved an expert in Bayesian statistics NMA (GvV). We then conducted a re‐analysis of all three NMAs.
The dataset from Cipriani provided the relevant information (mean change from baseline, SD, sample size) for all arms and trials and was thus ready to use for the re‐analysis. The dataset from Zhou included similar information; however, there was variation in the assignment of treatment arms to results without supporting information to explain the assignments. Therefore, we inspected each trial in the dataset to ensure matching of treatment arms to results. For re‐analysis of Hetrick we created a dataset as described above. In trials with multiple doses we combined arms following the formula provided in the Cochrane Handbook [17]. There was no missing data. Estimates of direct and indirect summary measures were compiled and presented through tabulation.
2.4.3. Analyses
For all re‐analyses, we initially aimed to achieve computational reproducibility by using the same data and code as in the published articles. We twice contacted the corresponding authors of all three studies to request the full code used in their analyses (Cipriani and Zhou published only part of their code; Hetrick published none). We did not receive any code. Therefore, we wrote an analytical code, which is available via the Open Science Framework (OSF).
We conducted Bayesian NMAs using a random effects model for the Lancet studies’ datasets, using all available information including details of priors to emulate their analyses (vague priors with a uniform distribution between 0 and 1 for the heterogeneity SD, and a normal with mean 0 and SD = 100 for all other prior distributions). We used R version 4.3.0 [18] with R's “gemtc” package version 1.1 [19]. Model convergence was assessed with visual inspection of trace‐plots and Gelman plots. We could not achieve perfect computational reproducibility so we conducted a sensitivity analysis using a frequentist approach to assess the robustness of the Bayesian results for Cipriani and Zhou.
We conducted a frequentist random effects NMA for the Hetrick dataset, as in the original publication, using R's “netmeta” package version 3.2‐0 [20].
To assess the impact of a single study with outlying results on efficacy estimates, heterogeneity, and consistency, we re‐evaluated the two Bayesian NMAs with and without this study. We also ran tests of outliers by investigating standardized residuals and Mahalanobis distance with R's NMAoutlier package [21] for the Cipriani and Zhou data.
2.4.4. Open Science and Reporting
This exploratory study did not have a pre‐registered protocol. All data and code are either publicly available (links in text) or on the OSF. While there is no specific reporting guideline for this type of analysis, we adapted the PRISMA‐NMA framework for reporting findings [16].
2.5. Patient Involvement
No patients were involved in the planning of our re‐analysis. The results of our study were discussed with the external committee of the RestorRes project (Research Integrity in Biomedical Research), which brings together a diverse range of stakeholders, including patient representatives, civil society representatives, legal experts, journalists, research integrity officers, and researchers. The committee discussed the reasons provided by The Lancet for rejecting the manuscript, particularly the argument that the findings would have a limited impact on the overall conclusions, and that responsibility lay with the original journal rather than with The Lancet. The importance of informing the authors of the affected meta‐analyses was emphasized, in order to give them the opportunity to make appropriate corrections. The committee also suggested contacting the authors, editors, and publisher of the trial in question. Finally, the committee highlighted “the need to await validation through peer review before any public communication of the results.”
3. Results
3.1. Critical Analysis of the Three NMAs
3.1.1. NMA Methodology
Table 1 shows the key characteristics of the three NMAs regarding their fluoxetine‐placebo efficacy findings. Cipriani and Zhou used a Bayesian approach while Hetrick used a frequentist one. This would not be expected to account for substantial differences in a single finding, particularly when the Bayesian analyses used vague priors.
Table 1.
Characteristics of three network meta‐analyses of medication for treatment of depression in children and adolescents focusing on fluoxetine versus placebo efficacy findings for symptom reduction.
| Cipriani et al. (2016) | Zhou et al. (2020) | Hetrick et al. (2021) | |
|---|---|---|---|
| Interventions considered | Antidepressants | Antidepressants and psychotherapy | New generation antidepressants |
| Efficacy Outcomes Included for Depression Symptoms | CDRS‐S, CDI, BDI | “Standardized depressive symptom scales” | CDRS‐R only |
| Total # Studies | 34 | 71 | 26 |
| # Studies including a fluoxetine arm | 10 | 14 | 9 (8 used the CDRS) |
| # Studies including fluoxetine and placebo arms | 8 | 9 | Not stated; we calculated 9 |
| Fluoxetine vs. placebo network SMD | −0.51 (95% CrI −0.99 to −0.03) | −0.51 (95% CrI −0.84 to −0.18) | −0.20* (95% CI −0.28 to −0.11) using −2.84 points on CDRS‐R and SD = 14.471 |
| Pre‐defined equivalence range | None | None | ±5 points on CDRS‐R |
| Provided direct and indirect estimates | In appendix | In appendix | No |
| Confidence rating for fluoxetine vs. placebo SMD | Very low | Very low | Moderate |
SMD is for change in depression rating scales.
Hetrick et al. reported only mean difference (MD) and did not report an SMD; the SMD here is calculated using a standard deviation (SD) 14.471, as the authors suggest based on their previous version of the NMA.
We considered whether differences in study populations would explain the inconsistent findings, particularly the possibility that Hetrick included more recent studies with direct comparisons of fluoxetine versus placebo that were not available for the earlier NMAs. Two such studies (Findling et al. [22, 23]) were included only in Hetrick, both of which used fluoxetine as a control arm alongside placebo and an investigational drug (vilazodone and vortioxetine respectively). Findling (2020) reported an MD for fluoxetine versus placebo of −2.3 points (p = .14) on the CDRS‐R, while Findling (2022) reported an MD of −3.73 points (p = .015) favoring fluoxetine [22, 23]. However, in our pairwise meta‐analytic aggregation of fluoxetine trials over time, published elsewhere [7], these two studies had no notable effect on the point estimate.
3.1.2. Network Geometry
We did not identify any major differences in the network geometry for medication. Zhou included studies of psychotherapy as well as medication. In all three NMAs, placebo was the most common comparator for drug trials and fluoxetine the most studied drug.
3.1.3. Inconsistency
Both Lancet NMAs reported inconsistency between the direct and indirect results for fluoxetine versus placebo, identified using the node‐splitting method (separating direct and indirect estimates for each comparison), i.e., the direct and indirect estimates were statistically significantly different. The supporting material for both NMAs showed that there were very large differences between the direct and indirect results for the fluoxetine‐notriptyline‐placebo loop (summarized in Table 2). The effect sizes for the direct comparison between nortriptyline versus fluoxetine, and for the indirect comparison between nortriptyline versus placebo, were very large (SMD ≈ 4).
Table 2.
Evaluation of the inconsistency by node‐splitting model in Cipriani et al. and Zhou et al. for the comparisons comprising the fluoxetine‐nortriptyline‐placebo loop for efficacy on depression symptoms.
| Direct | Indirect | Difference | ||||||
|---|---|---|---|---|---|---|---|---|
| Study | Comparison | SMD | SE | SMD | SE | SMD | SE | p |
| Cipriani 2016 | PBO‐FLX | 0.26 | 0.14 | 1.40 | 0.48 | −1.14 | 0.50 | 0.0209 |
| PBO‐NOR | 0.11 | 0.25 | −3.96 | 0.62 | 4.07 | 0.67 | <0.001 | |
| NOR‐FLX | 4.22 | 0.62 | 0.15 | 0.26 | 4.07 | 0.67 | <0.0001 | |
| Zhou 2020 | PBO‐FLX | 0.24 | 0.12 | 1.26 | 0.30 | −1.02 | 0.33 | 0.0019 |
| PBO‐NOR | 0.12 | 0.29 | −3.93 | 0.65 | 4.06 | 0.71 | 0.0000 | |
| NOR‐FLX | 4.22 | 0.64 | 0.17 | 0.30 | 4.06 | 0.71 | 0.0000 | |
The inconsistent results included a very large direct result for the comparison of fluoxetine versus nortriptyline, SMD 4.22 in both Lancet NMAs, and we noted that effects of this size are unheard of in studies of antidepressants. The largest effect size reported in a 2018 NMA of medication for treatment of depression in adults, including drug‐drug comparisons, was amitriptyline versus placebo, with SMD (pooled) −0.48 [24], and in a review of various psychiatric drug treatments the maximum reported effect size was SMD 1.12 [25]. The pattern of inconsistency in the fluoxetine‐nortriptyline‐placebo loop indicated that the efficacy result for fluoxetine versus placebo had been indirectly informed by this very large SMD for fluoxetine versus nortriptyline.
3.1.4. Identifying the Source of the Large Effect Size in the Lancet NMAs
To identify the source of the large SMD for the direct fluoxetine versus nortriptyline comparison in the Lancet NMAs we reviewed the included studies to find head‐to‐head trials of these drugs. In both NMAs there was only one such study, Attari et al. [26]. This was a small trial of 40 patients randomized to fluoxetine (n = 20) or nortriptyline (n = 20) without a placebo arm, undertaken at a single site in Iran. The study used the self‐rated Children's Depression Inventory (CDI; range 0–54) to measure change in depression symptoms. It reported a mean change score of −10.9 (SD 2.6) for fluoxetine and −2.6 (SD 0.8) for nortriptyline, giving an MD on the CDI of 8.35 points favoring fluoxetine, corresponding to SMD −4.24 (95% CI −5.37 to −3.11, see OSF).
To further scrutinize this large effect size we considered Hetrick's subgroup analysis of three studies that used the CDI. The raw mean difference for drug versus placebo was −1.30 points (95% CI −5.87 to 3.27) for fluoxetine, −0.43 (−2.91 to 2.05) for paroxetine, and −0.28 (−3.72 to 3.16) for citalopram. The difference of 8.35 points reported by Attari et al. for fluoxetine versus nortriptyline is obviously much greater. We then considered the Attari et al. effect size in the context of the other effect sizes in the Lancet NMA datasets (for all rating scales). For the Cipriani dataset, Attari's effect size of −4.24 is more than 4 standard deviations from the pooled mean effects and more than four times the next largest effect size of 0.9 (Figure 1). Attari's outlier status was further confirmed by Mahalanobis plot and standardized residuals plot (Figures S2 and S3).
Figure 1.

Effect sizes in Cipriani et al. (2016).
We examined the supporting material of the Lancet NMAs to see if attempts were made to identify outlying studies. The only assessment we could identify were funnel plots, which are primarily used to identify publication bias but may secondarily identify outliers. However, in both Lancet NMAs the effect sizes had been centered on the pooled comparison‐specific effects, not on the pooled effect for the full dataset, so the comparison for nortriptyline versus fluoxetine was plotted at zero because Attari et al. was the only study for this comparison (i.e., it was compared to itself).
Next, the possibility was considered that the findings of Attari et al. could be explained by nortriptyline being genuinely (and greatly) inferior to placebo, but this was not supported by pairwise comparisons of nortriptyline versus placebo in the Lancet NMAs, both of which reported SMD −0.11 (95%CrI −0.55 to 0.34). This result was informed by two trials of nortriptyline versus placebo (Geller et al. [27], n = 31; Geller et al. [28], n = 60).
3.1.5. Critically Assessing the Outlier Trial
Considering the extremity of the effect size reported by Attari et al., we proceeded to analyze this study for other concerning features that might bring into question the credibility of its findings. To guide this analysis we used the preliminary version of INSPECT‐SR. We identified concerning features in two of four domains, particularly regarding contradictions in the data (see Table 3). The statistical results reported in the study were incompatible with each other, i.e., impossible. The extreme effect size based on change scores was more than four times larger than the effect size based on endpoint scores (SMD = −0.93 [95% CI −1.59 to −0.28], see OSF), and also differed from effect‐sizes calculated from baseline and endpoint values and assuming a range of correlations. To produce the reported effect size the correlation between pre‐post scores would have to be almost perfect (r > 0.97; see Figure S1 in Appendix and OSF) and this is not plausible. Furthermore, we noted unusually small SDs for the change scores and larger SDs for the groups based on pre‐ and post‐scores. These features call into question the trustworthiness of the data in Attari et al.
Table 3.
INSPECT‐SR Checklist for the trial by Attari et al. [26].
| Check | Response | Comments |
|---|---|---|
|
Domain: Inspecting post‐publication notices Overall domain judgment: No concerns | ||
| Does the study have an associated retraction? | No | |
| Does the study have an associated expression of concern or other relevant post publication notice? | No | The study is not listed in PubMed and has no Digital Object Identifier, thus it is not possible to raise concern via platforms such as PubPeer |
| Do other studies by the research team highlight causes for concern (associated retractions, expressions of concern, relevant post‐publication notices?) | No | |
|
Domain: Inspecting conduct, governance and transparency Overall domain judgment: Some concerns | ||
| Are there concerns relating to ethical approval? | Yes | Ethical approval not mentioned |
| Are there concerns relating to the timing or absence of study registration? | Unclear | No registration but study was published 2006 |
| Are there important inconsistencies between the publication and the registration documents? | N/A | No protocol is available |
| Is the recruitment of participants implausible? | No | 40 patients in 2 years at a single center |
| Are the reported methods implausible considering the reported resources? | Unclear | No information on the resources |
|
Inspecting text and figures Overall domain judgment: No concerns | ||
| Is there any plagiarized text, or text that is incompatible with the study? | No | |
| Is there evidence of manipulation or duplication of figures? | No | |
|
Inspecting results in the study Overall domain judgment: Serious concerns | ||
| Are there any unexplained discrepancies between reported data and participant inclusion criteria? | No | |
| Are numbers of participants allocated to each group implausible given the allocation method? | No | |
| Are any baseline data implausible? | No | |
| Are there any discrepancies between results reported in figures, tables, and text? | Yes | Minimal differences between text and data in Table 1: Text states no difference between groups for ethnicity; ethnicity is not reported in the table |
| Are the numbers of participants lost to follow‐up implausible? | No | |
| Are there any unexplained inconsistencies in the numbers of participants? | Yes | “The dropout in nortriptyline group and fluoxetine group were 4 and 2 respectively, during the study. So, they were replaced by new cases.” Replacement of dropouts is rarely done in RCTs, and no information is given if this was pre‐specified or done post‐hoc. Moreover, it is not clear if replacements were randomized |
| Are any outcome data, including estimated treatment effects, implausible? | Yes |
SMD 4.3 based on change scores is implausible and substantially deviates from the SMD 0.93 based on endpoint scores To produce the reported effect size the correlation between pre‐post scores would have to be almost perfect (r > 0.97). SD 0.8 for change score in fluoxetine group is implausible The small SD 0.8 for the fluoxetine group suggests that the response was very similar for all the patients in the group; this is unlikely to be compatible with the following: “At the endpoint (8th week), 10 cases (50%) didn't meet the criteria of Major Depression based on DSM‐IV in fluoxetine group” SDs for pre‐post groups and implausibly different from SDs for change scores |
| Are the means and variances of integer data impossible? | No | All reported mean values were checked with the GRIMMER test |
| Are there errors in statistical results? | Yes |
The p‐value of the t‐test for the main result is given as p = .004, but it should be 0.0052, and this cannot be explained with rounding issues. SDs for pre‐post scores are unlikely to be compatible with SD scores for change scores |
| Are any other contradictions implied by the data? | No | No |
| Are there inconsistencies in descriptions of methods and results across publications describing the study? | Yes | Unexplained replacement of dropouts |
| Overall study judgment: Serious concerns | ||
However, even if concerns about trustworthiness are set aside, Attari et al. is obviously an outlier. The Cochrane Handbook recommends that, “It is advisable to perform analyses both with and without outlying studies as part of a sensitivity analysis” [17]. No such sensitivity analysis was undertaken in either of the Lancet NMAs. We noted that because Hetrick excluded studies of tricyclics from their analysis, Attari et al. would not have been included. We therefore hypothesized that this study might explain the discrepant findings for the efficacy of fluoxetine between the Lancet NMAs and Hetrick, and proceeded to undertake a sensitivity analysis of the Lancet NMA datasets with and without this study.
3.1.6. Re‐Analysis of the Three NMAs With and Without the Outlier Trial
We re‐analyzed the Cipriani and Zhou datasets with and without Attari et al. using a Bayesian approach, to see how this study affected the fluoxetine efficacy estimates. Table 4 shows the results of the original analyses and all re‐analyses with and without Attari et al. We were able to closely reproduce the original findings of Cipriani and Zhou, except for some indirect estimates that were larger in our analyses. We reproduced the statistically significant inconsistency between direct and indirect comparisons for fluoxetine for both NMAs.
Table 4.
Results of the re‐analysis of Cipriani et al. for fluoxetine efficacy versus placebo with and without Attari et al.
| Study | Analysis | Direct efficacy estimate SMD (95% CrI/CI) | Indirect efficacy estimate SMD (95% CrI/CI) | p‐value for the test of inconsistency | Combined efficacy estimate SMD (95% CrI/CI) |
|---|---|---|---|---|---|
| Cipriani 2016 | Original NMA | −0.26 (−0.54, 0.02) | −1.40 (−2.34 to −0.46) | 0.021 | −0.51 (−0.99 to −0.03) |
| Re‐analysis with Attari | −0.24 (−0.68 to 0.20) | −2.1 (−3.3 to −1.1) | 0.00278 | −0.51 (−0.99 to −0.03) | |
| Re‐analysis without Attari | −0.27 (−0.46 to −0.07) | −0.14 (−0.58 to 0.86) | 0.72269 | −0.26 (−0.43 to −0.09) | |
| Zhou 2020 | Original NMA | −0.24 (−0.48 to 0.00) | −1.26 (−1.85 to −0.67) | 0.002 | −0.51 (−0.84 to −0.18) |
| Re‐analysis with Attari | −0.23 (−0.56 to 0.09) | −1.9 (−2.70 to −1.10) | 0.00014 | −0.49 (−0.82 to −0.16) | |
| Re‐analysis without Attari | −0.24 (−0.47 to −0.01) | −0.68 (−1.3 to 0.02) | 0.23341 | −0.29 (−0.50 to −0.07) | |
| Hetrick 2021 | Original NMA (raw CDRS score) | Not reported | Not reported | Not reported | −2.80 (−4.12 to −1.56) |
| Re‐analysis raw CDRS score | −3.04 (−4.47 to −1.61) | −4.68 (−12.17 to 2.82) | 0.6739 | −3.09 (−4.50 to −1.69) | |
| Re‐analysis SMD | −0.26 (−0.38 to −0.14) | −0.39 (−1.01 to 0.24) | 0.7009 | −0.27 (−0.38 to −0.15) |
Note: Efficacy estimates are given as SMDs for Cipriani and Zhou and CDRS‐points for Hetrick. Intervals in brackets are 95% credible intervals (CrI) in Bayesian analysis or 95% confidence intervals (CI) in frequentist analysis. CrIs from the “Original NMA” were calculated using ±1.96 multiplied with the reported SE; the SMD of the original Hetrick et al. finding was calculated using the suggested SD by Hetrick (see text for details).
Notably, exclusion of the trial by Attari et al. resolved the inconsistency for all three comparisons comprising the fluoxetine‐nortriptyline‐placebo loop. The network efficacy estimates for fluoxetine versus placebo were substantially smaller without Attari et al., with SMD −0.26 (−0.43 to −0.09) for Cipriani, and −0.29 (−0.50 to −0.07) for Zhou, i.e., much closer to the result of Hetrick, which we had calculated as −0.2 using the MD and suggested SD from the original NMA, and which was −0.27 (−0.38 to −0.15) in our re‐analysis using our reconstructed Hetrick dataset. For the frequentist approach we also found comparable results in the sensitivity analysis for Cipriani and Zhou (Table S1).
In our re‐analysis of Hetrick there was no inconsistency between direct and indirect comparisons for fluoxetine versus placebo (SMD −0.26 and −0.39, p = 0.70), in‐keeping with the original NMA.
3.1.7. Raising Concerns
We twice contacted the authors of Attari et al. to raise our concerns about the validity of their data but received no response. We then twice contacted the editor of the journal in which it was published, and received a response which said the editor had forwarded our concerns to the authors. We have since followed this up with three further emails to journal editors and received no response. We then twice emailed the publisher, Wolters Kluwer, and received no response. We then contacted Wolters Kluwer through a professional social networking platform and received a response that a subsidiary organization of Wolters Kluwer, MedKnow, held publishing responsibility for the journal. An email address was provided for MedKnow. We contacted MedKnow three times by email and at the time of writing have received no response.
The journal in which Attari et al. was published has no Digital Object Identifier and is not listed in PubMed so we are unable to post a commentary on, for example, PubPeer. We submitted our findings (as an earlier version of this article) to Lancet Psychiatry, who sent our manuscript for peer review and shared our findings with The Lancet and with The Lancet Group Research Integrity Team. Based on the reports of two reviewers, The Lancet Group journals rejected our manuscript and advised us that they felt no further action was required on their part and they considered the matter closed. They suggested that if we had concerns about Attari et al. we should raise these with the relevant authors or editors.
4. Discussion
4.1. Principal Findings
We have explained the discrepant findings between the Lancet and Cochrane NMAs: the inclusion in the former of Attari et al., which increased the overall estimate for fluoxetine versus placebo, through indirect comparisons. In our re‐analysis we found that removing this trial resolved the discrepant findings for fluoxetine's efficacy, and it resolved the inconsistency between direct and indirect estimates in the Lancet NMAs. An important finding from our analysis is that one small outlying trial can substantially bias the overall results in an NMA, including the most clinically impactful results.
The question arises of why Attari was not identified as a problematic trial in the original Lancet publications. While Attari et al. was rated as having a moderate risk of bias by Cipriani and an unclear risk of bias by Zhou, two important features of this trial remained undetected. First, the extreme effect size in Attari is clearly a statistical outlier. Second, it also has multiple concerning features which undermine its reliability so significantly that it may need to be corrected or retracted. The inconsistency it created was acknowledged as problematic in both NMAs, and they offered various hypotheses as explanation, but the cause was not identified and consequently sensitivity analysis was not done.
Our re‐analysis supports Cochrane's recommendation to run sensitivity analyses when there are statistical outliers, but shows that obvious outliers can remain unacknowledged even in meta‐analyses by large teams of experienced researchers. Furthermore, our findings suggest that existing risk‐of‐bias analyses are not sufficient to detect problematic trials. Systematic reviews typically assess the methodological quality of the included trials, but do not consider whether the data themselves can be trusted. The importance of assessing the trustworthiness of findings, using a tool such as INSPECT‐SR [15], is increasingly recognized as an essential part of ensuring the integrity of research syntheses [29].
Our re‐analysis also has implications regarding the evaluation of fluoxetine's efficacy. While the confidence or credible intervals in all three NMAs excluded the null‐effect, the point‐estimates and lower margin of the credible interval in the Lancet NMAs were reduced by about 50% from the original results, in line with the findings by Hetrick. This efficacy estimate puts fluoxetine much more clearly in the range of clinical equivalence with placebo in all three NMAs, using both the Hetrick definition and using established estimations in antidepressant research [14].
Our experience of sharing our findings with authors and editors indicates that significant barriers remain to facilitating critical analysis and debate about evidence syntheses and problematic trials. Our experience of communicating our concerns to the editors of the journals that published the NMAs that included Attari et al. raises important questions about the responsibility of authors and editors of evidence syntheses. If editors have no responsibility to alert readers to such concerns, but the authors of the original trials ignore correspondence, it is not clear what recourse there is (or should be) for researchers when they encounter such findings.
4.2. Strengths and Weaknesses
We were unable to replicate all numerical estimates exactly, and we did not have access to all original code and data used in the three NMAs. Discrepancies in our results are minor, however, and are likely to be caused by small variations in statistical modeling (e.g., different methods for node‐splitting [30]). For Hetrick, there may have been minor variations in our data extraction compared to the original NMA. However, we were able to reproduce the principal conclusions and inferences, i.e., the different magnitude of the effect estimates across the NMAs, and the existence of inconsistencies between direct and indirect estimates in the Lancet NMAs. These issues were resolved by the exclusion of Attari et al. and this was confirmed in sensitivity analysis using both Bayesian and frequentist approaches.
Several other studies have reported on the effects of problematic trials in evidence syntheses, with sometimes substantial changes in effect sizes or even a change of direction of the effect [31, 32]. We are not aware of a systematic review of how the inclusion/exclusion of problematic trials in meta‐analyses can alter treatment recommendations, and of how the publishing system corrects (or does not correct) systematic reviews and meta‐analyses when problematic trials are detected after publication. However, the negative effects of problematic trials can only be mitigated effectively if authors and editors involved in the affected evidence syntheses take responsibility for making corrections [33].
4.3. Implications for Practice and Research and Unanswered Questions
Our results may help clinicians and guideline developers who use evidence syntheses to estimate the efficacy of fluoxetine and weigh up the risks and benefits of its use in this population. Our re‐analyses give a lower point estimate with narrower certainty intervals that exclude much of the lower portion of the original Lancet NMA credible intervals, and more closely accord with the results of Hetrick. Furthermore, our results bring the estimates for the Lancet NMA datasets (excluding Attari et al.) largely within the range of clinical equivalence as defined by Hetrick.
Our findings emphasize that confidence ratings in risk of bias analyses may not communicate effectively the degree to which the data underlying NMA results may be unreliable. Furthermore, confidence ratings are often only a detail in the results section or supporting information. It remains an open question how researchers can avoid such important contextual information being lost in evidence summaries and guidelines. Sometimes “low confidence” should be replaced with “no confidence.”
Our re‐analysis has implications for the broader field of evidence synthesis and medical publishing. Our findings illustrate how the inclusion of a single small problematic trial can substantially distort the clinically important results of NMAs.
Our analysis also highlights the importance of post‐publication peer review and how transparency in the process of analysis, including the sharing of datasets and code, is important to facilitate this process. Our experience of raising concerns with authors and editors highlights the difficulty of ensuring transparency and critical appraisal of the published evidence when concerns about trials and their impact on evidence syntheses arise. While the need for self‐correction in the publishing system is acknowledged, open questions remain as to how this process can be enforced and expedited, especially in cases where treatment recommendations—and patient care—are directly affected.
Author Contributions
Richard Lyus: conceptualization, writing – original draft, writing – review and editing, data curation, validation. Florian Naudet: writing – review and editing, supervision. Gert van Valkenhoef: formal analysis, writing – review and editing. Martin Plöderl: formal analysis, conceptualization, writing – original draft, data curation, validation, writing – review and editing.
Funding
The authors have nothing to report.
Conflicts of Interest
All authors have completed the Unified Competing Interest form and declare: RL and MP have no conflicts to declare. FN has no relation with any pharmaceutical company but received funding from the French National Research Agency (ANR‐23‐CE36‐0006, for his work on research integrity), the French ministry of health and the French ministry of research. He is a work package leader in the OSIRIS project (Open Science to Increase Reproducibility in Science). The OSIRIS project has received funding from the European Union's Horizon Europe research and innovation program under the grant agreement No. 101094725. He is a work package leader for the doctoral network MSCA‐DN SHARE‐CTD (HORIZON‐MSCA‐2022‐DN‐01 101,120,360), funded by the EU. GvV is an employee of the Cochrane Collaboration. However, Cochrane was not involved with, and did not support, this work in any way.
Copyright/Licence for Publication
The Corresponding Author has the right to grant on behalf of all authors and does grant on behalf of all authors, a worldwide licence to the Publishers and its licensees in perpetuity, in all forms, formats and media (whether known now or created in the future), to (i) publish, reproduce, distribute, display and store the Contribution, (ii) translate the Contribution into other languages, create adaptations, reprints, include within collections and create summaries, extracts and/or, abstracts of the Contribution, (iii) create any other derivative work(s) based on the Contribution, (iv) to exploit all subsidiary rights in the Contribution, (v) the inclusion of electronic links from the Contribution to third party material where‐ever it may be located; and (vi) licence any third party to do any or all of the above.
Data Sharing
All statistical code is available via the OSF at https://osf.io/jyq5w/. The data for Cipriani is available through the Oxford University Research Archive at https://ora.ox.ac.uk/objects/uuid:a4d1186b-6be4-486a-b9ce-ce431ccd545b. The dataset for Zhou et al. is freely available through Mendeley at https://data.mendeley.com/datasets/kw6nmfn2tb/1, the version we used for our analysis is publicly available here https://docs.google.com/spreadsheets/d/12R6YFEXsoIewi3OXMFx5D4StHDpiO-i6iGtuefB9Na8/edit?gid=1032532817#gid=1032532817. The data we used for the re‐analysis of Hetrick et al. is publicly available here https://docs.google.com/spreadsheets/d/1AyUTNtlQa5KKGu-0zJrvKHv47Bzw4V7R5Ib3VXXlGgA/edit?usp=sharing.
Transparency Statement
The lead author (the manuscript's guarantor) MP affirms that the manuscript is an honest, accurate, and transparent account of the study being reported; no important aspects of the study have been omitted, and discrepancies from the study as originally planned/registered have been explained.
Supporting information
Supporting File 1
Supporting File 2
Acknowledgments
We thank Cipriani et al. and Zhou et al. for publicly sharing their data. We thank Jack Wilkinson, lead author of the INSPECT‐SR template, for sharing the latest version of the template and providing feedback on our assessment of Attari et al. Open Access funding provided by Paracelsus Medizinische Privatuniversitat.
Data Availability Statement
The data that support the findings of this study are openly available in OSF at https://osf.io/jyq5w/.
References
- 1. Shorey S., Ng E. D., and Wong C. H. J., “Global Prevalence of Depression and Elevated Depressive Symptoms Among Adolescents: A Systematic Review and Meta‐Analysis,” British Journal of Clinical Psychology 61 (2022): 287–305, 10.1111/bjc.12333. [DOI] [PubMed] [Google Scholar]
- 2. Lester T. R., Herrmann J. E., Bannett Y., Gardner R. M., Feldman H. M., and Huffman L. C., “Anxiety and Depression Treatment in Primary Care Pediatrics,” Pediatrics 151 (2023): e2022058846, 10.1542/peds.2022-058846. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Chua K.‐P., Volerman A., Zhang J., Hua J., and Conti R. M., “Antidepressant Dispensing to US Adolescents and Young Adults: 2016–2022,” Pediatrics 153 (2024): e2023064245, 10.1542/peds.2023-064245. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Wesselhoeft R., Jensen P. B., Talati A., et al., “Trends in Antidepressant Use Among Children and Adolescents: A Scandinavian Drug Utilization Study,” Acta Psychiatrica Scandinavica 141 (2020): 34–42, 10.1111/acps.13116. [DOI] [PubMed] [Google Scholar]
- 5. Cao T. X. D., Fraga L. F. C., Fergusson E., et al., “Prescribing Trends of Antidepressants and Psychotropic Coprescription for Youths in UK Primary Care, 2000‐2018,” Journal of Affective Disorders 287 (2021): 19–25, 10.1016/j.jad.2021.03.022. [DOI] [PubMed] [Google Scholar]
- 6. Klau J., Bernardo C. D. O., Gonzalez‐Chica D. A., Raven M., and Jureidini J., “Trends in Prescription of Psychotropic Medications to Children and Adolescents in Australian Primary Care from 2011 to 2018,” Australian & New Zealand Journal of Psychiatry 56 (2022): 1477–1490, 10.1177/00048674211067720. [DOI] [PubMed] [Google Scholar]
- 7. Plöderl M., Lyus R., Horowitz M. A., and Moncrieff J., “The Loss of Efficacy of Fluoxetine in Pediatric Depression: Explanations, Lack of Acknowledgment, and Implications for Other Treatments,” Journal of Clinical Epidemiology 189 (2026): 112016, 10.1016/j.jclinepi.2025.112016. [DOI] [PubMed] [Google Scholar]
- 8. Taylor D. M., Barnes T. R., and Young A. H., The Maudsley Prescribing Guidelines in Psychiatry. John Wiley & Sons, 2025. 15th ed. [Google Scholar]
- 9. Walter H. J., Abright A. R., Bukstein O. G., et al., “Clinical Practice Guideline for the Assessment and Treatment of Children and Adolescents with Major and Persistent Depressive Disorders,” Journal of the American Academy of Child and Adolescent Psychiatry 62 (2023): 479–502, 10.1016/j.jaac.2022.10.001. [DOI] [PubMed] [Google Scholar]
- 10. Korczak D. J., Westwell‐Roper C., and Sassi R., “Diagnosis and Management of Depression in Adolescents,” Canadian Medical Association Journal 195 (2023): E739–E746, 10.1503/cmaj.220966. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Cipriani A., Zhou X., Del Giovane C., et al., “Comparative Efficacy and Tolerability of Antidepressants for Major Depressive Disorder in Children and Adolescents: A Network Meta‐Analysis,” Lancet 388 (2016): 881–890, 10.1016/S0140-6736(16)30385-3. [DOI] [PubMed] [Google Scholar]
- 12. Zhou X., Teng T., Zhang Y., et al., “Comparative Efficacy and Acceptability of Antidepressants, Psychotherapies, and Their Combination for Acute Treatment of Children and Adolescents With Depressive Disorder: A Systematic Review and Network Meta‐Analysis,” Lancet Psychiatry 7 (2020): 581–601, 10.1016/S2215-0366(20)30137-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Hetrick S. E., McKenzie J. E., Bailey A. P., et al., “New Generation Antidepressants for Depression in Children and Adolescents: A Network Meta‐Analysis,” Cochrane Database of Systematic Reviews 5, no. 5 (2021): CD013674, 10.1002/14651858.CD013674.pub2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Hengartner M. P. and Plöderl M., “Estimates of the Minimal Important Difference to Evaluate the Clinical Significance of Antidepressants in the Acute Treatment of Moderate‐to‐Severe Depression,” BMJ Evidence‐based Medicine 27 (2021): 60–73, 10.1136/bmjebm-2020-111600. [DOI] [PubMed] [Google Scholar]
- 15. Wilkinson J. D., Heal C., Flemyng E., et al., “INSPECT‐SR: A Tool for Assessing Trustworthiness of Randomised Controlled Trials,” BMJ 394 (2026): e00611, 10.1136/bmj-2026-100611. [DOI] [PubMed] [Google Scholar]
- 16. Hutton B., Salanti G., Caldwell D. M., et al., “The PRISMA Extension Statement for Reporting of Systematic Reviews Incorporating Network Meta‐Analyses of Health Care Interventions: Checklist and Explanations,” Annals of Internal Medicine 162 (2015): 777–784, 10.7326/M14-2385. [DOI] [PubMed] [Google Scholar]
- 17. Higgins J. P. T., Thomas C. J., Chandler J., et al., Cochrane Handbook for Systematic Reviews of Interventions Version 6.5. Cochrane, 2024. [Google Scholar]
- 18. R Core Team R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, 2025. [Google Scholar]
- 19. Van Valkenhoef G. and Kuiper J., gemtc: Network Meta‐Analysis Using Bayesian Methods. CRAN, 2025. [Google Scholar]
- 20. Balduzzi S., Rücker G., Nikolakopoulou A., et al., “netmeta. An R Package for Network Meta‐Analysis Using Frequentist Methods,” Journal of Statistical Software 106 (2023): 1–40, 10.18637/jss.v106.i02. [DOI] [Google Scholar]
- 21. Petropoulou M., Schwarzer G., Panos A., et al. NMAoutlier: Detecting Outliers in Network Meta‐Analysis. CRAN, 2019. [Google Scholar]
- 22. Findling R. L., McCusker E., and Strawn J. R., “A Randomized, Double‐Blind, Placebo‐Controlled Trial of Vilazodone in Children and Adolescents With Major Depressive Disorder With Twenty‐Six‐Week Open‐Label Follow‐Up,” Journal of Child and Adolescent Psychopharmacology 30 (2020): 355–365, 10.1089/cap.2019.0176. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Findling R. L., DelBello M. P., Zuddas A., et al., “Vortioxetine for Major Depressive Disorder in Adolescents: 12‐Week Randomized, Placebo‐Controlled, Fluoxetine‐Referenced, Fixed‐Dose Study,” Journal of the American Academy of Child and Adolescent Psychiatry 61 (2022): 1106–1118, 10.1016/j.jaac.2022.01.004. [DOI] [PubMed] [Google Scholar]
- 24. Cipriani A., Furukawa T. A., Salanti G., et al., “Comparative Efficacy and Acceptability of 21 Antidepressant Drugs for the Acute Treatment of Adults With Major Depressive Disorder: A Systematic Review and Network Meta‐Analysis,” Lancet 391 (2018): 1357–1366, 10.1016/S0140-6736(17)32802-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Leucht S., Hierl S., Kissling W., Dold M., and Davis J. M., “Putting the Efficacy of Psychiatric and General Medicine Medication Into Perspective: Review of Meta‐Analyses,” British Journal of Psychiatry 200 (2012): 97–106, 10.1192/bjp.bp.111.096594. [DOI] [PubMed] [Google Scholar]
- 26. Attari A., Moghaddam F. Y., Hasanzadeh A., et al., “Comparison of Efficacy of Fluoxetine With Nortriptyline in Treatment of Major Depression in Children and Adolescents: A Double‐Blind Study,” Journal of Research in Medical Sciences 11 (2006): 24–30. [Google Scholar]
- 27. Geller B., Cooper T. B., Graham D. L., Marsteller F. A., and Bryant D. M.. “Double‐Blind Placebo‐Controlled Study of Nortriptyline in Depressed Adolescents Using a ‘Fixed Plasma Level’ Design,” Psychopharmacology Bulletin 26 (1990): 85–90. [PubMed] [Google Scholar]
- 28. Geller B., Cooper T. B., Graham D. L., Fetner H. H., Marsteller F. A., and Wells J. M.. “Pharmacokinetically Designed Double‐Blind Placebo‐Controlled Study of Nortriptyline in 6‐ to 12‐Year‐Olds With Major Depressive Disorder,” Journal of the American Academy of Child & Adolescent Psychiatry 31 (1992): 34–44. [DOI] [PubMed] [Google Scholar]
- 29. Weibel S., Popp M., Reis S., Skoetz N., Garner P., and Sydenham E., “Identifying and Managing Problematic Trials: A Research Integrity Assessment Tool for Randomized Controlled Trials in Evidence Synthesis,” Research Synthesis Methods 14 (2023): 357–369, 10.1002/jrsm.1599. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Van Valkenhoef G., Dias S., Ades A. E., and Welton N. J., “Automated Generation of Node‐Splitting Models for Assessment of Inconsistency in Network Meta‐Analysis,” Research Synthesis Methods 7 (2016): 80–93, 10.1002/jrsm.1167. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Graña Possamai C., Cabanac G., Perrodeau E., Ghosn L., Ravaud P., and Boutron I., “Inclusion of Retracted Studies in Systematic Reviews and Meta‐Analyses of Interventions: A Systematic Review and Meta‐Analysis,” JAMA Internal Medicine 185 (2025): 702, 10.1001/jamainternmed.2025.0256. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Xu C., Fan S., Tian Y., et al., “Investigating the Impact of Trial Retractions on the Healthcare Evidence Ecosystem (VITALITY Study I): Retrospective Cohort Study,” BMJ 389 (2025): e082068, 10.1136/bmj-2024-082068. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Gross C. P., Flanagin A., Perencevich E. N., and Inouye S. K., “Mitigating the Impact of Retracted Studies in the Medical Literature,” JAMA Internal Medicine 185 (2025): 621, 10.1001/jamainternmed.2025.0251. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supporting File 1
Supporting File 2
Data Availability Statement
The data that support the findings of this study are openly available in OSF at https://osf.io/jyq5w/.
