Meta-analyses are widely regarded as the highest level of evidence because they combine multiple studies to provide more precise estimates of treatment effects. This expectation is justified when the studies address essentially the same clinical question with comparable methodological rigor. In practice, however, investigators must make clinical and methodological assumptions regarding which studies should be combined and how differences in study credibility should be handled. Consequently, pooled estimates are not simply extracted from the literature but are constructed through these assumptions. Meta-analysis is therefore a powerful analytical tool rather than an automatic source of truth. Under some circumstances, quantitative synthesis may obscure rather than clarify the contribution of the most credible individual studies.
Recent systematic reviews evaluating melatoninergic agents for delirium prevention in critically ill patients illustrate this challenge. Chaves and colleagues performed a comprehensive systematic review including both melatonin and ramelteon and reported findings supportive of a preventive effect of melatoninergic therapy [1]. In contrast, our group restricted the analysis to melatonin and concluded that current evidence was insufficient to recommend routine prophylactic administration [2]. Rather than asking which meta-analysis is “correct”, we believe the more informative question is why two rigorous reviews addressing a similar clinical question reached different conclusions.
A first source of divergence concerns clinical heterogeneity, which determines study eligibility. Chaves et al. combined trials evaluating melatonin and ramelteon, whereas our review focused exclusively on melatonin [1, 2]. Although these agents belong to the same pharmacological family, differences in receptor affinity, pharmacokinetics and potentially mechanisms relevant to delirium pathophysiology make their quantitative combination a reasonable, but not self-evident, methodological choice [3]. Nevertheless, this difference appears insufficient to explain the apparent discrepancy. Chaves et al. also reported a subgroup analysis restricted to melatonin trials, allowing direct comparison with our review. The resulting effect estimates were remarkably similar (RR 0.77 versus 0.86), with largely overlapping confidence intervals (0.62–0.97 vs. 0.71–1.04). Thus, differences in clinical eligibility criteria explain only part of the observed discrepancy.
The remaining differences appear to arise primarily from methodological heterogeneity. Once clinically eligible studies have been identified, investigators must decide how differences in study credibility should be incorporated into the analysis. Two equally defensible strategies are commonly adopted. One includes all clinically eligible studies, maximizing precision while acknowledging differences in study quality during interpretation. The other prioritizes studies at low risk of bias, accepting a potential loss of precision in exchange for greater protection against systematic error. Neither strategy is inherently correct, simply reflecting different balances between protection against false-positive and false-negative conclusions. Consequently, analyses stratified according to risk of bias should not be viewed merely as sensitivity analyses performed to confirm a preferred primary result. Rather, they provide an alternative, methodologically justified perspective on the same body of evidence, and disagreement between them should be regarded as an important finding requiring explanation.
Our review adopted the second strategy, considering the analysis restricted to low-risk-of-bias randomized trials as primary while presenting the broader analysis as secondary [2]. Chaves et al. pooled all eligible trials irrespective of risk of bias while assessing study quality using ROB-2 [1]. Both approaches are methodologically defensible. However, because nearly 40% of the melatonin trials included in the primary analysis were judged at high risk of bias, stratification according to the authors’ own ROB-2 assessment would have provided important additional information.
To explore this issue, we performed a post hoc stratified analysis of the melatonin trials included by Chaves et al. using their published ROB-2 assessments. The apparent treatment benefit was largely confined to studies at high risk of bias (RR 0.46, 95% CI 0.29–0.73), whereas studies at low risk of bias yielded a more conservative estimate (RR 0.86, 95% CI 0.70–1.05), remarkably similar to that observed in our review (RR 0.89, 95% CI 0.73–1.09). The interaction between risk-of-bias strata was statistically significant (P for interaction = 0.01). Thus, much of the apparent disagreement between the two reviews disappeared once methodological heterogeneity was explicitly explored.
The comparison between these reviews illustrates a broader principle. When clinically coherent analyses based on different methodological assumptions produce similar results, confidence in the evidence increases. Conversely, when conclusions depend on whether studies at high risk of bias are included, the discrepancy itself becomes an important finding. Rather than selecting the preferred pooled estimate, investigators should first seek to understand the origin of the disagreement.
One practical approach is to identify studies exerting the greatest influence on the pooled estimate. Leave-one-out analyses performed on both datasets produced remarkably consistent findings. Sequential omission of the trials by Yin and Shi substantially attenuated the apparent treatment benefit, whereas omission of the PRO-MEDIC trial by Wibrow et al. shifted the pooled estimate in the opposite direction [4]. These observations do not invalidate any individual study; instead, they identify the trials that deserve the closest clinical and methodological scrutiny. Careful comparison may reveal differences in patient selection, intervention protocols, outcome assessment or methodological conduct that are not immediately apparent from the pooled estimate alone.
The principal lesson extends beyond melatonin and delirium prevention. Meta-analysis should not be regarded as the endpoint of critical appraisal but as the beginning of a deeper investigation whenever different, methodologically defensible analyses produce conflicting conclusions. Exploration of methodological heterogeneity, particularly through analyses stratified according to risk of bias, should become an integral part of evidence interpretation rather than a sensitivity analysis consulted only after the primary result has been accepted. When these analyses lead to different conclusions, the discrepancy itself should increase uncertainty rather than be dismissed because one analysis was designated as secondary. We suggest that the robustness of conclusions across clinically and methodologically defensible analyses deserves greater consideration when judging confidence in pooled estimates. From a guideline perspective, such uncertainty would rarely justify a strong recommendation based on efficacy alone. The final recommendation should instead reflect the balance of potential benefits, harms, resource use, feasibility, patient values and preferences, to determine whether a conditional recommendation is appropriate or whether the evidence should simply be considered insufficient.
Acknowledgements
Not applicable.
Author contributions
IL and DP wrote the manuscript together and contributed equally.
Funding
None.
Data availability
No datasets were generated or analysed during the current study.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
IL received fees from presentations from Viatris and AOP Health. No other potential conflict of interest relevant to this article was reported.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Chaves-Filho A, Maria F, Medradoa T, Zambonia J, Quinne. Viviane Cordeiro Veigaf, Theodoros Mavridisg. Effectiveness and Safety of Melatoninergic Agonists in Preventing Delirium in the ICU: An Updated Dose–Response Meta-Analysis of Randomized Controlled Trials. Critical Care. 2026. [DOI] [PMC free article] [PubMed]
- 2.Lakbar I, Poole D, Delamarre L, Chanques G, Pensier J, Monet C, Belafia F, Capdevila M, De Jong A, Jaber S. Melatonin and delirium in the intensive care units: a systematic review and meta-analysis of randomized controlled trials. Intensive Care Med. 2025;6. 10.1007/s00134-025-08143-1. [DOI] [PubMed]
- 3.Spadoni G, Bedini A, Lucarini S, Mor M, Rivara S. Pharmacokinetic and pharmacodynamic evaluation of ramelteon: an insomnia therapy. Expert Opin Drug Metab Toxicol. 2015;11(7):1145–56. 10.1517/17425255.2015.1045487. [DOI] [PubMed] [Google Scholar]
- 4.Wibrow B, Martinez FE, Myers E, Chapman A, Litton E, Ho KM, Regli A, Hawkins D, Ford A, Van Haren FMP, Wyer S, McCaffrey J, Rashid A, Kelty E, Murray K, Anstey M. Prophylactic melatonin for delirium in intensive care (Pro-MEDIC): a randomized controlled trial. Intensive Care Med. 2022;48(4):414–25. 10.1007/s00134-022-06638-9. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No datasets were generated or analysed during the current study.
