Skip to main content
NPJ Digital Medicine logoLink to NPJ Digital Medicine
letter
. 2026 Jun 13;9:452. doi: 10.1038/s41746-026-02911-z

Reassessing the evidence linking clinical leadership to AI deployment outcomes

Henry Bair 1,✉
PMCID: PMC13264604  PMID: 42288669

Abstract

Li et al. report that clinician last authorship is associated with greater “impact” in AI deployment trials. This commentary argues that current evidence is insufficient to attribute this to clinical leadership per se. Concerns include outcome definition in a literature where 88% of trials are positive, statistical instability from only 13 non-events, authorship as a noisy proxy for leadership, and structural confounding by trial design and geography.

Subject terms: Health care, Medical research, Scientific community


As AI deployment studies in healthcare accumulate, the field is beginning to draw organizational lessons from a literature in which positive findings are common and study designs are heterogeneous. In an integrative analysis of 105 studies, Li et al. report that clinician last authorship is associated with statistically significant impact in published AI deployment trials1. The current evidence, however, may not yet be sufficient to attribute that association to clinical leadership itself. Using the summary data provided by Li et al., this commentary argues that the reported pattern is also consistent with publication bias, instability of sparse-data modeling, differences in trial design, and variation in cross-disciplinary expertise. It additionally proposes alternative interpretations and clearer reporting standards for future AI deployment studies.

Reinterpreting the reported association between leadership and “impact”

Li et al. synthesized 105 clinical studies of AI deployment, largely randomized clinical trials, and classified last authors as either clinicians or technologists. They found that 94% (80/85) of clinician-led studies versus 60% (12/20) of technologist-led studies reported a statistically significant effect of AI. Logistic regression adjusting for AI- and workflow-related variables yielded an odds ratio (OR) of 7.79 (P = 0.039) for clinician versus technologist last authors; a robustness check restricted to randomized trials reported an even larger OR of 19.90 (P = 0.047). On this basis, the authors suggest that clinical leadership is associated with more impactful AI deployment in healthcare.

Leadership and team structure are likely to be crucial in realizing the promise of AI in healthcare. However, the data and methods currently available support a more cautious interpretation. In my view, the observed association is also consistent with four distinct issues: (i) the construct used to define “impact”, (ii) publication and selective-reporting processes that shape which studies enter the literature, (iii) instability of the statistical model under severe data sparsity, and (iv) systematic differences in trial design and team expertise that are only partially captured by the clinician–technologist dichotomy.

Outcome definition and publication bias constrain inference

First, the outcome analyzed by Li et al. is itself a constructed binary variable: “impact” was defined as whether a study reported at least one statistically significant effect (P ≤ 0.05) or a favorable conclusion. Across the 105 studies, 92 (88%) were classified as having significant impact, with only 13 deemed non-impactful. This distribution is far more lopsided than is typical for randomized trials in other clinical domains and suggests substantial publication and selective-reporting bias, as the authors themselves note is a concern in this literature. In such a setting, the outcome being modeled may therefore approximate the “probability that a published paper reports a positive result” more closely than the underlying clinical impact of AI.

Dichotomizing highly heterogeneous trials by a single significance threshold also ignores effect sizes, precision, multiplicity, and the clinical importance of the endpoints being measured, all of which influence whether a study is labeled impactful independently of any substantive benefit. The limitations of treating P < 0.05 as a binary indicator of success have been extensively discussed in the methodological literature2. Even before any regression model is fit, this makes the outcome difficult to interpret as a measure of true clinical benefit.

A separate and independent concern is the statistical model used to estimate the association. When the event of interest is rare, standard logistic regression with many covariates can produce unstable estimates and exaggerated odds ratios3. In Li et al.’s study, only 13 studies contribute information about the “negative” outcome, yet the multivariable model includes numerous predictors (leadership, centers, AI model type, AI origin, clinical setting, region, comparison type, and leadership roles). This yields fewer than two non-events per parameter, falling below many commonly recommended thresholds for reliable logistic regression, and is consistent with the extremely wide confidence intervals reported (e.g., 1.62–910.19 for the RCT-only analysis). Under these conditions, modest misclassification of outcomes (for instance, borderline P values or differences in how authors phrase conclusions) or small changes in the set of included studies could substantially alter the estimated effect of leadership background. Penalized methods such as Firth-corrected or rare-events logistic regression are often recommended in such sparse settings and could be considered in future analyses3,4. The reported odds ratios are therefore more safely interpreted as suggestive rather than definitive, and the primary contribution of the study is hypothesis-generating.

Authorship order is a noisy proxy for leadership and expertise

The central exposure in Li et al.’s study is last authorship, assumed to indicate leadership in AI deployment within each study team. While this assumption aligns with many clinical publishing conventions, it is less reliable in multidisciplinary AI work, where patterns of credit differ across fields and regions. Computer science and engineering traditions often emphasize first authorship; some journals list authors alphabetically; and “co-senior” or “co-corresponding” authorship can diffuse leadership across multiple individuals.

Moreover, the classification of authors into “clinicians” and “technologists” masks important hybrid roles. Li et al. treat authors with both MD and PhD degrees, or with PhDs in clinical disciplines such as radiology, as clinicians. In contemporary AI deployment projects, such cross-trained physician–informaticians frequently serve as boundary spanners who understand both clinical workflow and algorithmic nuance, and may embody precisely the kind of diversified experience that supports effective strategic decision-making5,6. If such hybrid leaders are systematically more successful and also more likely to be labeled as “clinician last authors,” then the analysis conflates professional identity with cross-disciplinary expertise.

Li et al. also introduce “organizational leadership roles” (e.g., director, chief) based on current titles retrieved from public profiles as of 2025, restricted to first and last authors. These roles did not show a significant association with impact in their model. However, the timing and scope of such positions may not align with the period of AI deployment, and operational leaders (e.g., nursing or administrative champions) may not appear as first or last authors. Collectively, these factors suggest that last-author discipline is an imperfect and potentially noisy proxy for the leadership and expertise attributes believed to influence AI implementation.

Trial design and geography are likely structural confounders

Li et al. observe that technologist-led studies disproportionately used an “AI versus routine care” comparison, whereas clinician-led studies more often evaluated “AI-assisted versus unassisted” clinicians (75% vs 35% of trials in the former category). This design choice is not merely a covariate; it fundamentally shapes the clinical question being asked. Trials of AI-assisted clinicians test augmentation (does decision support improve clinician performance?), whereas AI-versus-routine-care trials often test substitution or triage (can an AI replace or pre-screen clinicians?). There are strong a priori reasons to expect higher success rates in augmentation trials than in substitution trials, regardless of leadership background.

At the same time, Li et al. show that impact rates differ substantially by region, with studies conducted in Asia reporting significant effects more frequently than those from other regions, and that clinician last authorship is more prevalent in Asia and North America than elsewhere. Regional variation in regulatory regimes, health-system readiness, reimbursement, and journal audience likely affects both the feasibility of certain comparison types and the tendency to publish equivocal results.

Taken together, these patterns suggest a causal structure in which leader background influences trial design and setting, and where region and specialty shape both leadership opportunities and the probability of publishing positive findings. In such a directed acyclic graph, trial design and region lie on pathways between leadership and observed “impact” and may also act as confounders. Simply adjusting for them in a sparse logistic regression may not fully account for this structure and could either under- or over-estimate the role of leadership background.

From “who leads” to “how teams lead”

Despite these limitations, Li et al.’s work is a valuable prompt to think more rigorously about team structure in AI deployment. The question their data can more safely address is not whether clinical leaders are inherently more effective than technologist leaders, but rather how current norms around authorship, trial design, and publication in AI-in-healthcare research shape the appearance of success. A stronger causal interpretation would require reanalysis that separates augmentation from substitution trials, uses modeling approaches appropriate for sparse outcomes, and complements bibliometric authorship classification with more granular assessment of who actually led workflow redesign, implementation, and evaluation.

I propose three practical directions for future work:

  1. Deployment-specific leadership descriptors. Most academic journals now require author contribution statements, often using the CRediT taxonomy. The more specific recommendation here is that AI deployment studies report, in addition, who led clinical workflow design, local model adaptation or validation, implementation, frontline training, evaluation, and organizational change. Capturing cross-trained “hybrid” expertise, rather than only clinical versus technical degrees, could better align operational reality with analytical categories7.

  2. Pre-specified endpoints, effect sizes, and preregistration. Building on existing scoping reviews of AI trials in clinical practice, which have highlighted publication bias, endpoint heterogeneity, and limited patient-important outcomes8,9, future meta-research should classify outcomes (structural, process, patient-important) and analyse effect magnitudes, not only statistical significance. But effect-size reporting alone will not resolve selective publication bias. Beyond the established requirement to register randomized trials, preregistration of AI deployment studies more broadly—including prospective cohort, stepped-wedge, interrupted time-series, and before-and-after designs—would help distinguish prespecified primary outcomes from post hoc favorable reporting.

  3. Prospective organizational evaluation using common protocols. Rather than relying solely on cross-sectional analyses of published papers, health systems or multicenter consortia could prospectively follow AI deployment efforts using a common protocol that records team composition, whether leadership is clinical, technical, or paired, comparator choice, end-user involvement, governance structure, adoption, fidelity, workflow workarounds and sustainability over time. Embedding qualitative interviews or process evaluations alongside quantitative outcomes would make it possible to study how teams lead in practice, not only who appears in the last-author position.

In conclusion, Li et al. provide an important first attempt to link leadership background with the outcomes of AI deployment studies in healthcare. Building on their own acknowledgement of publication bias and authorship limitations, the current evidence is best interpreted as raising hypotheses about how teams are configured and how trials are designed, rather than as definitive evidence that clinical leadership per se increases the likelihood of AI impact. Until analyses are available that more carefully define “impact,” address sparse-data modeling, and characterize leadership structure more directly, the findings are better viewed as descriptive associations within the published literature than as causal evidence. Clarifying this distinction will be essential for informing workforce development and organizational strategies as health systems scale up AI-enabled care.

Acknowledgements

No financial support was received for this work.

Author contributions

H.B. contributed to the conceptualization and writing of this manuscript.

Data availability

No datasets were generated or analysed during the current study.

Competing interests

The author declares no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Li, Q. et al. The impact of leadership on AI deployment study outcomes in healthcare: an integrative analysis. npj Digit. Med.8, 799 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Greenland, S. et al. Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. Eur. J. Epidemiol.31, 337–350 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.King, G. & Zeng, L. Logistic regression in rare events data. Polit. Anal.9, 137–163 (2001). [Google Scholar]
  • 4.Heinze, G. & Schemper, M. A solution to the problem of separation in logistic regression. Stat. Med.21, 2409–2419 (2002). [DOI] [PubMed] [Google Scholar]
  • 5.Hambrick, D. C. Upper echelons theory: An update. Acad. Manag. Rev.32, 334–343 (2007). [Google Scholar]
  • 6.Hambrick, D. C. & Mason, P. A. Upper echelons: The organization as a reflection of its top managers. Acad. Manag. Rev.9, 193–206 (1984). [Google Scholar]
  • 7.Hall, K. L. et al. The science of team science: A review of the empirical evidence and research gaps on collaboration in science. Am. Psychol.73, 532–548 (2018). [DOI] [PubMed] [Google Scholar]
  • 8.Han, R. et al. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review. Lancet Digit. Health6, e367–e373 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Yang, R. et al. Disparities in clinical studies of AI-enabled applications from a global perspective. npj Digit. Med.7, 209 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No datasets were generated or analysed during the current study.


Articles from NPJ Digital Medicine are provided here courtesy of Nature Publishing Group

RESOURCES