Stein et al 1 provide an integrative overview of the expanding landscape of biomarker research in psychiatry. They emphasize the importance of a nuanced perspective on the strengths and limitations of large‐scale analytics, and the need for advancements in methodology, study design, and the conceptualization of mental illness.
With a still limited understanding of core pathological processes, psychiatry actually faces a particularly pronounced challenge among medical fields in identifying biomarkers with robust and clinically meaningful performance – a difficulty underscored by recent large‐scale evaluations showing that even advanced machine‐learning approaches fail to yield generalizable neuroimaging‐based diagnostic biomarkers for major depressive disorder 2 .
Influenced by yet unresolved conceptual questions and the availability of large retrospective datasets, psychiatric biomarker research has often utilized case‐control designs, which, despite their methodological validity, may be less informative regarding clinically actionable tasks such as treatment selection or relapse prediction.
In other medical fields, including oncology and cardiology, biomarker and predictive research is embedded in more clearly defined biological substrates and, importantly, in established intervention frameworks – factors that, as noted by Stein et al, critically influence the clinical utility of large‐scale data analytics. This raises the question of whether psychiatry faces a uniquely intractable complexity in biomarker discovery and clinical translation, or whether valuable insights can be drawn from the translational trajectories of these other disciplines.
In oncology, biomarker research is anchored in an arguably more obvious cellular pathology compared to psychiatry, yet faces similar challenges due to pronounced biological heterogeneity 3 . For example, individual cancer types may harbor multiple histological and molecular subtypes, limiting prediction from available samples. Despite this, numerous highly effective oncological biomarkers, approved by the US Food and Drug Administration (FDA) – predominantly genomic and molecular companion diagnostics – are being used in clinical practice 4 .
Their clinical translation illustrates a conceptual shift from a “mechanism‐complete” perspective towards clinically actionable endpoints and an iterative refinement of etiological understanding. Newer markers with direct patient benefit may inform therapeutic decisions – for example, mutations conferring sensitivity or resistance to targeted therapies. These advancements, rooted in the acceptance of biological heterogeneity as an inherent limitation, may provide psychiatry with a blueprint where translational strategy is directed towards clinical applicability and an iterative refinement of etiological understanding.
A further example for a “non‐linear” translational pathway is offered by developments in cardiology, where the utility of multimodal artificial intelligence (AI) algorithms for risk of heart failure or arrhythmia was driven by the a priori definition of useful clinical decision points. Focusing on who would benefit from intensified monitoring or anticoagulation evaluation, rather than diagnosing the underlying pathophysiology, facilitated the development of FDA‐cleared, AI‐assisted electrocardiographic models for prediction of atrial fibrillation risk that are available for use in clinical practice 5 . Similarly, emerging AI applications for heart‐failure risk stratification target clinically actionable endpoints, such as the need to initiate or escalate guideline‐directed therapy. The clinical translation of these applications was supported by intervention pathways that were already established, even if the underlying pathophysiology was only partially understood.
A third example comes from pediatrics, where growth charts have long had high clinical utility despite not being tied to disease‐specific mechanistic understanding. These are used to compare measures such as body mass index, height, or head circumference against age‐ and sex‐specific reference distributions. Clinically meaningful deviations from typical developmental trajectories then guide decisions about further diagnostic evaluation or intervention. While the analogous “normative modeling” approaches are widely used in neuroscience, linking them to concrete clinical decision rules could be an important step towards their future clinical application.
Unlike conventional, supervised machine learning, normative models capture reference distributions and usually do not contain information on clinical endpoints directly. In psychiatry, where different deviations may be associated with the same clinical phenomenon, normative models may capture clinically relevant change better than direct prediction, allowing them, as suggested by Stein et al, to generate real clinical value.
Although personalized medicine frameworks from other medical disciplines may not be directly applicable in psychiatry, directing biomarker discovery more towards clinical action may be an important step to facilitate translation. As emphasized by Stein et al, advances in reproducible and standardized data acquisition are needed to support multi‐site analytics and longitudinal robustness. This should go hand‐in‐hand with the definition of concrete, clinically actionable endpoints 6 , such as early treatment switching, or intensified monitoring to prevent relapse. These endpoints need to be operationalized regarding their application contexts, intervention pathways, and expected benefits and risks.
Baseline risks under usual care can inform the required effect of a predictive model to have a positive impact on clinical practice. This can guide the definition of actionable decision thresholds – e.g., based on the predicted non‐response probability. AI model performance metrics can be supplemented with measures more closely tied to the clinical application contexts.
For example, the frequently used area under the curve (AUC) measures the ability of a model to differentiate between individuals in different outcome groups at all possible model thresholds. However, in practice, decisions are made at specific thresholds, and the metric does not indicate whether such decision would meaningfully change the outcome probability for a given patient. The evaluation of clinical relevance requires additional absolute risk‐based measures at specific decision thresholds, complemented by decision‐analytic concepts, such as the net benefit, that summarize whether model‐guided decisions are likely to improve patient outcomes 7 . Similarly, positive and negative predictive values indicate how reliably high‐ or low‐risk classifications translate into true outcomes in the population to which the model is applied.
These examples illustrate that the clinical utility of AI models depends on the context of their application and the underlying event rate in the population where they find application. Considering this already during study design could allow aligning model development more closely to clinical decision making. Linking predictions to actionable clinical endpoints also supports the interpretability of AI model outputs at the clinical level, which is an important requirement for successful clinical translation 8 .
The methodological standardization alluded to by Stein et al will then form a critical element for staged validation schemes where AI models need to be tested in comparable external cohorts, followed by prospective clinical trials. These validation strategies will profit from early engagement with regulatory bodies, to ensure alignment with emerging qualification pathways and conformity requirements for AI‐based clinical tools, laying the groundwork for integration into electronic health records, clinical guidelines, and reimbursement structures 9 .
Assessments of patient benefit should actively incorporate patient perspectives, for example through co‐development with patient representative groups, to ensure that clinical decisions supported by AI tools align with patient values and priorities. They should also be accompanied by a deep understanding of the consequences of false positive and false negative predictions, that can only be established in representative and sufficiently heterogeneous target populations.
Advances made through a structured translational strategy focused around clinically actionable endpoints may lead to more rapid implementation of tools with real benefit for affected individuals and, in parallel, to an iterative conceptual refinement of mechanistic theories of mental disorders. As Stein et al argue, AI and related big‐data methodologies will function as powerful instruments only to the extent that they are coupled with a clearly articulated and effective translational strategy.
This work was supported by the Hector Foundation II.
REFERENCES
- 1. Stein DJ, Kessler RC, Torous J et al. World Psychiatry 2026;25:388‐413. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Winter NR, Blanke J, Leenings R et al. JAMA Psychiatry 2024;81:386‐95. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Passaro A, Al Bakir M, Hamilton EG et al. Cell 2024;187:1617‐35. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Kulasingam V, Prassas I, Diamandis EP. NPJ Precis Oncol 2017;1:17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Niazai A, Jamil H, Hameed M et al. BMC Cardiovasc Disord 2025;25:849. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Meehan AJ, Lewis SJ, Fazel S et al. Mol Psychiatry 2022;27:2700‐8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Van Calster B, Collins GS, Vickers AJ et al. Lancet Digit Health 2025;7:100916. [DOI] [PubMed] [Google Scholar]
- 8. Rajpurkar P, Chen E, Banerjee O et al. Nat Med 2022;28:31‐8. [DOI] [PubMed] [Google Scholar]
- 9. Busch F, Kather JN, Johner C et al. NPJ Digit Med 2024;7:210. [DOI] [PMC free article] [PubMed] [Google Scholar]
