Skip to main content
Annals of General Psychiatry logoLink to Annals of General Psychiatry
. 2026 May 9;25:52. doi: 10.1186/s12991-026-00671-4

Diagnostic accuracy of machine learning approaches for suicide‑related outcomes: a meta‑analysis

Parisa Kohnepoushi 1, Maryam Afraie 2, Hamza Rahmani 1, Seyedeh Asrin Seyedoshohadaei 3, Babak Ghadirzadeh 1, Yousef Moradi 4,✉
PMCID: PMC13326310  PMID: 42106827

Abstract

Objective

This diagnostic test accuracy meta-analysis aimed to provide clinically interpretable estimates (sensitivity, specificity, likelihood ratios (LR), predictive value (PV) and post‑test probabilities) for machine‑learning (ML) models predicting suicide‑related outcomes.

Methods

A systematic search of PubMed, Embase, PsycINFO, and Web of Science identified studies published between January 2010 and December 2024. Eligible studies used a single-gate design (cross-sectional or longitudinal), included at least 100 participants, and reported diagnostic performance metrics (e.g., sensitivity, specificity, area under the receiver operating curve (AUC)) for ML algorithms. Models examined included Support Vector Machines (SVM), Logistic Regression, Random Forest, XGBoost, Artificial Neural Networks (ANN), Gradient Boosting, and Ensemble approaches. Two reviewers independently extracted data. Pooled estimates of sensitivity, specificity, PPV, NPV, and AUC were calculated using a bivariate random-effects model. Risk of bias was assessed using QUADAS-2.

Results

Of 500 screened records, 22 studies met inclusion criteria. Ensemble models demonstrated the highest pooled AUC (0.95; 95% CI, 0.92–0.96), with specificity of 0.97 (95% CI, 0.95–0.98) and sensitivity of 0.50 (95% CI, 0.29–0.71). Gradient-boosting and ensemble approaches showed strong discriminative performance overall; model-specific estimates are reported in the Results. At a pre-test probability of 25%, post-test probabilities for a positive result ranged from 64% (Logistic Regression) to 88% (Ensemble Models).

Conclusion

Machine-learning approaches demonstrated promising diagnostic accuracy for suicide-related outcomes in heterogeneous clinical populations. However, because primary studies rarely reported diagnosis- or outcome-specific performance, these findings should not be assumed to generalize across specific disorders or to suicide mortality. Future research should incorporate diagnosis-stratified and outcome-resolved validation to clarify clinical applicability.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12991-026-00671-4.

Keywords: Machine learning, Suicidal ideation, Suicide attempts, Suicide mortality, Diagnostic accuracy, Meta-analysis

Introduction

Suicide is a critical public health issue, contributing significantly to premature mortality globally. In order to increase the effectiveness of clinical intervention and prevention, accurate identification of individuals at risk for suicidal ideation, attempts, or death is of great importance [1]. Traditional suicide risk assessment tools, relying on clinical judgment and risk factor checklists, often demonstrate limited sensitivity and specificity; with predictive accuracy only marginally better than chance [2]. For instance, current literature shows that up to 95% of patients classified as high-risk do not die by suicide, whilst many who do were previously deemed low-risk; highlighting the clinical inadequacy of these tools [3, 4].

Various studies have evaluated the effectiveness of traditional suicide risk prediction models [3–5]. These models utilize clinical interviews, psychiatric evaluations, plus demographic factors to predict the likelihood of future suicide attempts alongside providing valuable insights into suicide risk factors [5]. A 2017 meta-analysis by Franklin et al. evaluated quantitative models for suicide risk prediction from the past fifty years. The analysis determined that the models’ performance is only slightly better than random chance [6, 7]. This poor accuracy may compromise patient safety by misclassifying risk levels. Large et al. (2016) discovered that 95% of patients deemed high risk did not die by suicide, while many of those who did were initially classified as low risk [8]. These findings suggest that traditional tools for suicide risk assessment lack adequacy in clinical settings. Furthermore, the low incidence rate of suicide complicates accurate predictions [9, 10]. Machine learning (ML), on the other hand, offers a promising alternative by leveraging complex, non-linear patterns in high-dimensional data, such as electronic health records and psychological assessments, in order to enhance the accuracy level of prediction [11–13]. Recent studies suggest ML models outperform traditional methods in identifying suicide-related outcomes, as they provide a scalable and adaptable approach for clinical decision-making [7, 11]. However, prior reviews have been limited by reliance on suboptimal metrics (e.g., odds ratios (ORs)) and study designs, which fail to assess longitudinal predictive validity [14]. Additionally, risks of bias, such as data leakage and overfitting, remain underexplored [14]. Contrary to traditional methods, ML models are more flexible and scalable since they are not only capable of processing large amounts of data but also adaptable to different data sources [12]. These models demonstrate potential for improving suicide risk detection by effectively representing complex relationships between risk factors and outcomes, thereby providing a more accurate and reliable tool for clinical decision-making [12].Existing syntheses have been predominantly narrative and have often pooled discrimination metrics across heterogeneous study designs, limiting the ability to draw conclusions about predictive validity and clinical applicability; the present approach directly addresses these methodological gaps.

According to current literature, more robust metrics, such as sensitivity, specificity, and area under the receiver operating curve (AUC), better evaluate model accuracy for clinical practice [15–17]. However, many reviews relied on cross-sectional studies, limiting their ability to assess true predictive validity [18, 19]. Longitudinal studies are therefore essential for evaluating future suicide-related outcomes [18, 19]. Additionally, prior meta-analyses often overlooked risks of bias, such as data leakage, improper feature selection, and overfitting, which might have inflated the performance estimates reported by them [20, 21].

This systematic review and meta-analysis evaluates the performance of ML models in predicting suicide-related outcomes, focusing on robust metrics (e.g., sensitivity, specificity, the receiver operating characteristic curve [22]). By addressing gaps in prior studies and incorporating recent evidence, the present study aims to guide the integration of ML into clinical practice, attempting to improve suicide risk assessment.

This review advances the literature by restricting quantitative synthesis to longitudinal studies in which predictors precede outcomes, thereby aligning estimates with predictive validity; by employing a bivariate random-effects diagnostic test accuracy model to synthesize sensitivity, specificity, and likelihood ratios rather than pooling AUCs alone; and by presenting post-test probabilities to facilitate clinical interpretation.

Methods

This systematic review and meta-analysis was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy (PRISMA-DTA) guidelines in order to evaluate the diagnostic accuracy of ML models in predicting suicide-related outcomes (suicidal ideation, attempts, and mortality) [23].

Protocol registration

The present review protocol was registered prospectively in the International Prospective Register of Systematic Reviews (PROSPERO) under the registration number CRD420251159042.

Information sources and search strategy

PubMed, Embase, PsycINFO, and Web of Science was searched thoroughly for studies published between January 2010 and December 2024. The start year (2010) was selected a priori to coincide with the emergence and uptake of modern machine-learning methods in psychiatric research and the broader availability of digitized electronic health records, enabling model development and validation at scale. The search strategy utilized in the present study combined terms related to machine learning (e.g., “machine learning”, “artificial intelligence”, “predictive modeling”), suicide-related outcomes (e.g., “suicide”, “suicidal ideation”, “suicide attempt”), and diagnostic accuracy (e.g., “sensitivity”, “specificity”, “AUC”) using Boolean connectors. Moreover, reference lists of included studies and relevant reviews were manually screened for additional eligible studies.

Eligibility criteria and study selection

Studies were included based on the PIRT framework (Participants, Index Test, Reference Standard, Target Condition) in order to evaluate the diagnostic accuracy of ML models regarding the prediction of suicide-related outcomes. Screening and eligibility decisions were undertaken independently by two reviewers, with conflicts resolved by consensus and, upon requirement, third-reviewer adjudication. The inclusion and exclusion criteria are as follows:

Inclusion criteria

Participants

Studies involving human participants from clinical settings (e.g., psychiatric inpatients, outpatients) or community populations (e.g., general population, veterans, adolescents), with a minimum sample size of 100 to ensure robust diagnostic accuracy estimates. No restrictions were applied to age, gender, or geographic location. Where available, it was planned to stratify diagnostic performance by primary psychiatric diagnosis; however, separable diagnostic subgroup metrics were infrequently reported across primary studies.

Index test

ML models (specifically Support Vector Machine (SVM), Logistic Regression, Random Forest, XGBoost, Artificial Neural Networks (ANN), Gradient Boosting, and Ensemble Models) used to predict suicide-related outcomes, reporting at least one diagnostic accuracy metric (sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), positive likelihood ratio (LR+), negative likelihood ratio (LR-), or area under the receiver operating characteristic curve [22]). Furthermore, models must have employed internal or external validation (e.g., k-fold cross-validation, hold-out testing).

Reference standard

Validated measures of suicide-related outcomes, such as suicidal ideation (e.g., Columbia-Suicide Severity Rating Scale, Beck Scale for Suicide Ideation), suicide attempts (e.g., documented in electronic health records or clinical registries), or suicide mortality (e.g., confirmed via national death registries or coroner reports).

Target condition

Suicide-related outcomes (suicidal ideation, suicide attempts, or suicide mortality) assessed amongst studies with a minimum follow-up of 6 months in order to evaluate predictive validity. Studies must be peer-reviewed, English-language publications from January 2010 to December 2024. Since outcome-specific diagnostic metrics (for ideation, attempts, and mortality) were infrequently and inconsistently reported across primary studies, pre-specified pooling across suicide-related outcomes and planned outcome-specific subgroup meta-analyses was pre-specified only if ≥ 3 studies per outcome reported separable sensitivity/specificity; this threshold was not met.

In practice, analyzable diagnostic indices (paired sensitivity/specificity with AUC) were predominantly available in studies published from 2022 onward, reflecting improvements in reporting standards [12, 15].

Data extraction

Two reviewers independently extracted data using a standardized form, including study characteristics (author, year, country, sample size, population), ML model details (algorithm type, input features, validation method), diagnostic accuracy metrics (sensitivity, specificity, PPV, NPV, LR+, LR-, AUC), study design (e.g., longitudinal follow-up duration) and risk of bias indicators. Missing data were requested from study authors where feasible. Two independent reviewers extracted all data in duplicate using a piloted form. The reviewers were distinct individuals and worked independently in order to minimize abstraction bias and transcription error. Discrepancies were resolved by consensus; a third senior reviewer adjudicated when consensus was not reached. For each study, the named algorithm was mapped to the pre-specified taxonomy (e.g., GBM/GBT to Gradient Boosting; XGBoost to XGBoost; RF to Random Forest; LR to Logistic Regression; SVM to SVM; ANN to ANN; stacking/voting to Ensemble models).

Risk of bias and quality assessment

The Quality Assessment of Diagnostic Accuracy Studies (QUADAS-2) tool was used to assess risk of bias and applicability concerns across four domains: patient selection, index test, reference standard, and flow and timing; resulting in each domain being rated as low, high, or unclear risk [24]. Particular attention was given to patient selection and flow/timing to avoid inflated accuracy from suboptimal two-gate case–control designs, consistent with DTA guidance. Risk-of-bias judgments were performed in duplicate by two independent reviewers. The reviewers were different individuals from the data extraction pairings, and each completed assessments independently to reduce bias and enhance reproducibility. Any disagreements were reconciled by consensus, with escalation to a third reviewer for arbitration when necessary.

Model taxonomy and selection

Algorithm families were defined a priori in order to reflect commonly applied paradigms in the included literature and to enable clinically interpretable contrasts. Logistic regression represented the linear/probabilistic baseline; support vector machines represented margin-based kernel methods; random forests represented tree-based bagging; gradient boosting represented tree-based boosting; artificial neural networks represented deep/connectionist models and study-level ensembles captured stacking or voting approaches that combined multiple base learners. Classes were retained for quantitative pooling only when ≥ 3 longitudinal studies reported compatible diagnostic metrics.

Rationale for distinguishing XGBoost

XGBoost was analyzed separately from ‘generic’ gradient boosting because many primary reports implemented XGBoost specifically and labeled it as such. And also because of the fact that its regularized objective, second-order (Newton) tree boosting, shrinkage, column subsampling, and sparsity-aware split finding can yield performance and calibration profiles that differ from other gradient-boosting implementations [25, 26]. Where study counts were insufficient for separate pooling, XGBoost results were described narratively.

Definition of ensemble models

In the present review, ‘ensemble models’ refers solely to study-level stacking/blending or voting meta-models that combine predictions from heterogeneous base learners (e.g., LR, SVM, RF, GBM). Algorithms whose learning strategy is intrinsically an ensemble, Random Forest (bagging) and Gradient Boosting/XGBoost (boosting), were analyzed within their respective families and not classified as ‘ensemble models’. This avoids double counting and preserves interpretable contrasts across algorithm families.

Data synthesis and analysis

A bivariate random-effects model was employed to estimate pooled sensitivity, specificity, PPV, NPV, LR+, and LR– for each ML model (SVM, Logistic Regression, Random Forest, XGBoost, ANN, Gradient Boosting, Ensemble Models), accounting for the correlation between sensitivity and specificity. Summary receiver operating characteristic (SROC) curves were constructed to calculate AUC and assess discriminative ability, with prediction and confidence contours to evaluate consistency. Bivariate boxplots were generated to examine heterogeneity in sensitivity and specificity estimates across studies. Fagan’s nomograms were used to calculate post-test probabilities based on a 25% pre-test probability, reflecting clinical scenarios with moderate suicide risk. Likelihood matrices were derived to visualize the separation of positive and negative test outcomes. Heterogeneity was assessed using the I² statistic and Cochran’s Q test. Data synthesis was performed using the Midas and Metadta package in STATA (version 18). Forest plots and HSROC curves were generated to visualize diagnostic accuracy. A p-value < 0.05 was considered statistically significant for meta-regression. Moreover, Diagnostic subgroup meta-analyses were pre-specified depending upon ≥ 3 studies per diagnostic category reporting separable indices; this criterion was not met, precluding formal diagnosis-stratified pooling. On the other hand, between-study heterogeneity attributable to different decision thresholds was accommodated using hierarchical models (bivariate/HSROC), which capture the sensitivity–specificity trade-off across studies. Accordingly, cut-off variability, rather than design label alone, was anticipated to be the principal source of heterogeneity.

Results

A total of 500 records were identified through international database searches, with no additional records from other registers. After removing 160 duplicate entries, 340 records were screened by title, resulting in the exclusion of 228. The remaining 112 records were screened by abstract, and 36 were excluded at this stage. 76 full-text articles were assessed for eligibility, of which 54 were excluded; 29 due to reporting various outcomes or effect sizes not aligned with the review objective, and 25 due to being inappropriate study types such as cross-sectional studies, letters, or case reports. Ultimately, 22 studies [27–47] met the inclusion criteria and were therefore included in the final review (Fig. 1 and Supplementary File, Table S1). Most of included cohorts were diagnostically heterogeneous, and primary reports seldom provided outcome-specific diagnostic indices stratified by psychiatric diagnosis alongside several studies reporting composite ‘suicide-related outcomes’ without separable metrics; thus, averting the diagnosis-resolved subgroup or sensitivity meta-analyses.

Fig. 1.

Fig. 1

PRISMA 2020 flow diagram for a systematic review, which included searches of databases and registers only

Quality assessment of included studies

The studies included in this meta-analysis were evaluated for quality through the QUADAS-2 tool. The evaluation was conducted by two independent reviewers, and any disagreements were addressed and resolved through a consensus approach. The majority of studies exhibited a low risk concerning the index test, which applies to the validation of machine learning models; as well as the reference standard, which involves the application of established measures for assessing suicide risk. Nonetheless, the variability observed in patient selection, as well as in flow and timing, was more pronounced. Several studies have raised concerns regarding the representativeness of the samples and the consistency of follow-up procedures. These factors contributed to moderate or high risk of bias in some studies (Supplementary File, Figure S1). Observed heterogeneity was consistent with threshold variation across studies, a recognized driver in DTA meta-analyses [48].

Support vector machine (SVM)

The diagnostic outputs for the SVM model (Fig. 2) collectively demonstrate strong classification performance. The SROC curve shows an area under the curve (AUC) of 0.89(CI 95%: 0.79–0.85), indicating high discriminative ability. The prediction and confidence contours are narrow, suggesting consistent performance across studies. Furthermore, the bivariate boxplot supports this by showing firmly clustered sensitivity and specificity estimates which have no significant outliers, thereby indicating low heterogeneity and reliable generalization. The forest plot on the other hand demonstrates pooled sensitivity of 0.77 (95% CI: 0.68–0.84) and specificity of 0.77 (95% CI: 0.53–0.91) and the Fagan’s nomogram reflects a meaningful clinical impact; given a pre-test probability of 25%, a positive SVM test result increases the post-test probability to 72%, while a negative result reduces it to 9%. This model reaches a positive likelihood ratio (LR+) of 6.4 and a negative likelihood ratio (LR–) of 0.31 while the complementary likelihood matrix confirms these values; emphasizing on the strength of SVM in converting test outcomes into useful clinical information (Fig. 2).

Fig. 2.

Fig. 2

Logistic Regression performance across included longitudinal studies. (A) SROC curve with 95% confidence and prediction regions. (B) Forest plot of study-level sensitivity and specificity with 95% CIs. (C) Bivariate boxplot showing between-study dispersion in logit sensitivity/specificity. (D) Fagan’s nomogram translating pooled likelihood ratios to post-test probabilities at prespecified pretest risks. (E) Likelihood-ratio matrix (LR + and LR−) across plausible thresholds

Logistic regression

The logistic regression model (Fig. 3) shows considerably inferior performance across all five diagnostic outputs. The SROC curve, reflecting greater variability among studies, shows an AUC of 0.86 (95% CI:0.69–0.77). This is also seen in the bivariate boxplot; while having more distributed points, it shows a heterogeneous threshold selection among studies. The forest plot, on the other hand, presents a pooled sensitivity of 0.54 (95% CI:0.46–0.62) and specificity of 0.87 (95% CI:0.78–0.92); which suggests more variable performance. Moreover, Fagan’s nomogram shows a post-test probability increase to 64% after a positive test with an LR + of 4.2 and LR– of 0.41; which indicates clinical benefit. Furthermore, the likelihood matrix complements this finding with modest separation between positive and negative post-test probabilities, demonstrating the limited extent of diagnostic shift (Fig. 3).

Fig. 3.

Fig. 3

SVM performance across included longitudinal studies.Panel mapping: (A) SROC; (B) Forest plot; (C) Bivariate boxplot; (D) Fagan’s nomogram; (E) Likelihood-ratio matrix

Random forest

The random forest model’s overall diagnostic profile is shown in Fig. 4. The SROC curve here shows an AUC of 0.92 (95% CI: 0.78–0.84), indicating strong accuracy and repeatability. Furthermore, the bivariate boxplot shows data points as centered and symmetrically distributed, indicating minimal variability across studies. With a pooled sensitivity of 0.59 (95% CI: 0.48–0.69) and pooled specificity of 0.91 (95% CI: 0.83–0.95), the forest plot confirms previous findings. These results point to a strong equilibrium between actual positive and actual negative rates. Conversely, the Fagan’s nomogram indicates that with an LR + of 7.1 and an LR- of 0.27, post-test probability increases to 81% after a positive result (from a 25% pre-test). Furthermore, the likelihood matrix supports clinical integration by showing significant variation between positive and negative results and by reflecting these ratios (Fig. 4).

Fig. 4.

Fig. 4

Random Forest performance Panel mapping: (A) SROC; (B) Forest plot; (C) Bivariate boxplot; (D) Fagan’s nomogram; (E) Likelihood-ratio matrix

XGBoost

XGBoost shows the highest overall diagnostic performance among individual models (Fig. 5). The SROC curve yields an AUC of 0.94 (95% CI: 0.83–0.89) highlighting excellent generalizability. Moreover, the bivariate boxplot reveals compact, centralized distribution with minimal variance, indicating even diagnostic thresholds. In the forest plot, on the other hand, pooled sensitivity is 0.71 (95% CI: 0.60–0.80) and pooled specificity is 0.85 (95% CI: 0.78–0.90); indicating minimal heterogeneity. The Fagan’s nomogram shows that a positive test raises post-test probability to 86%, supported by an LR + of 8.9 and LR– of 0.24. The likelihood matrix, however, illustrates this separation visually, confirming this model’s strong diagnostic capacity (Fig. 5).

Fig. 5.

Fig. 5

Gradient Boosting performance Panel mapping: (A) SROC; (B) Forest plot; (C) Bivariate boxplot; (D) Fagan’s nomogram; (E) Likelihood-ratio matrix

Artificial neural networks (ANN)

Artificial Neural Network models demonstrate strong performance, as illustrated in Fig. 6, however, with a somewhat higher degree of variability compared to tree-based methods. The SROC curve reports an AUC of 0.95 (95% CI: 0.92–0.96), with broader prediction bands than XGBoost. The bivariate boxplot shows a wider spread, especially in specificity, implying moderate heterogeneity across studies. The forest plot reveals a pooled sensitivity of 0.50 (95% CI: 0.29–0.71) and a pooled specificity of 0.97 (95% CI: 0.95–0.98), showing broader confidence intervals compared to the Random Forest model. Fagan’s nomogram shows a post-test probability of 69%; followed by an LR + of 5.3 and LR– of 0.35, which indicates useful yet not optimal clinical interpretation. Similarly, the likelihood matrix illustrates these moderate values; thereby affirming the significance of the ANN model, specifically in contexts where non-linear relationships play a crucial role (Fig. 6).

Fig. 6.

Fig. 6

XGBoost performance Panel mapping: (A) SROC; (B) Forest plot; (C) Bivariate boxplot; (D) Fagan’s nomogram; (E) Likelihood-ratio matrix

Gradient boosting

Gradient Boosting, shown in Fig. 7, offers consistent yet slightly lower diagnostic power. The SROC curve yields an AUC of 0.89 (95% CI: 0.78–0.85); revealing slightly wider prediction contours than the XGBoost model. Furthermore, the bivariate boxplot shows mild asymmetry, suggesting modest heterogeneity in threshold effects. Regarding the forest plot, pooled sensitivity was reported as 0.79 (95% CI: 0.68–0.86) and pooled specificity as 0.71 (95% CI: 0.53–0.84). Fagan’s nomogram demonstrates a post-test probability increase to 71% which was supported by an LR + of 5.7 and LR– of 0.33. Furthermore, the likelihood matrix demonstrates a significant, yet mild, distinction between outcomes, which highlights the need for cautious clinical application (Fig. 7).

Fig. 7.

Fig. 7

ANN performance Panel mapping: (A) SROC; (B) Forest plot; (C) Bivariate boxplot; (D) Fagan’s nomogram; (E) Likelihood-ratio matrix

Ensemble models

Ensemble models defined as stacking/blending or voting meta-models combining heterogeneous learners (excluding Random Forest and Gradient Boosting/XGBoost, which are reported separately), shown within Fig. 8, provide the strongest diagnostic profile overall. The SROC curve presents the highest AUC at 0.95 (95% CI: 0.92–0.96) with extremely narrow prediction and confidence contours indicating exceptional reproducibility. The bivariate boxplot reveals a dense and symmetrical cluster of estimates confirming minimal heterogeneity. The forest plot demonstrates a moderate pooled sensitivity of 0.50 (95% CI: 0.29–0.71) and a high pooled specificity of 0.97 (95% CI: 0.95–0.98). The Fagan’s nomogram shows a striking rise in post-test probability to 88%, from a 25% baseline, accompanied by an LR + of 9.2 and LR– of 0.19, suggesting outstanding diagnostic utility. The likelihood matrix demonstrates a significant distinction between positive and negative test results, thereby reinforcing the model’s status as the most clinically relevant among those evaluated (Fig. 8).

Fig. 8.

Fig. 8

Ensemble model performance (stacking/blending/voting) Panel mapping: (A) SROC; (B) Forest plot; (C) Bivariate boxplot; (D) Fagan’s nomogram; (E) Likelihood-ratio matrix. Random Forest and Gradient Boosting/XGBoost are reported separately and not included in this ensemble category

Discussion

With growing interest in applying ML models to psychiatric risk prediction, particularly in studies published after 2021, there is a need to evaluate how these models perform across different algorithmic families and to assess their potential suitability for clinical integration [12, 49–53]. This systematic review and meta-analysis aimed to provide an updated and comprehensive assessment of ML models used to predict suicide-related outcomes, including suicidal ideation, attempts, and deaths. Overall, ML models demonstrated strong discriminative performance in longitudinal settings, with several algorithmic classes showing promising diagnostic characteristics [52, 53]. However, interpretation of comparative performance across algorithms should remain cautious, as direct within-study comparisons were limited and reporting practices varied considerably.

According to current theoretical frameworks, suicide is understood as a context-dependent and multifactorial phenomenon shaped by dynamic interactions among biological, psychological, and sociocultural factors [54]. ML algorithms are capable of modeling nonlinear and high-dimensional relationships without assuming linearity or independence among predictors [55]. This structural flexibility may help explain the diagnostic performance observed in several ML models; however, improved accuracy in research settings does not imply clinical readiness. Appropriate external validation, calibration, and evaluation of clinical utility remain essential before any model can be considered for implementation. The theoretical alignment between ML architectures and the complex nature of suicide risk highlights potential value, but this potential should be interpreted cautiously and within the limits of available evidence [6].

Ensemble models demonstrated high AUC values and relatively consistent specificity across included studies. These characteristics may reflect the ability of ensemble strategies to combine complementary strengths of individual algorithms [12, 50, 56, 57]. At the same time, lower pooled sensitivity indicates a trade-off that may limit applicability in settings where minimizing false negatives is critical. These observations should be interpreted as descriptive findings rather than evidence of superiority [12, 50, 56–58]. XGBoost also showed strong performance metrics in several studies, with relatively consistent threshold behavior. Nonetheless, variations in outcome definitions, predictor sets, and validation strategies across primary studies limit the extent to which general conclusions can be drawn. Because direct head-to-head comparisons were uncommon, these results should not be interpreted as evidence of universal advantage over other algorithms [59, 60]. Logistic regression demonstrated lower sensitivity in pooled estimates, a pattern that aligns with literature noting challenges faced by linear models in capturing complex interactions. However, differences in study design, feature selection, and outcome definitions may also contribute to these findings. Although logistic regression showed high specificity, the reduced sensitivity may limit its utility in contexts where missed cases carry substantial clinical consequences [61, 62]. SVM demonstrated balanced sensitivity and specificity, with an AUC of 0.89. SVM performance is influenced by its ability to handle high‑dimensional data, although kernel selection and feature tuning may be required to achieve consistent results across datasets [63]. As with other models, the evidence base remains heterogeneous, and conclusions should be drawn cautiously. Random Forest models showed a balance between specificity and moderate sensitivity, with a pooled AUC of 0.92. Lower sensitivity may be related to overfitting in high‑dimensional or limited datasets or to calibration issues. While the ensemble nature of Random Forests provides robustness to noise, their applicability depends on context, data quality, and validation procedures rather than inherent model characteristics [64].

Public health implications

ML models may support suicide-prevention efforts by providing risk estimates that can assist in prioritizing individuals for further assessment or intervention [65, 66]. When appropriately validated, likelihood ratios and post-test probabilities generated by these models can contribute to risk-stratified decision-making within existing public health frameworks [67]. In settings with limited mental health resources, such tools may help guide more focused screening or follow-up activities [68].

Despite these potential benefits, the low base rate of suicide substantially limits the positive predictive value of any model, making population-level screening challenging [69, 70]. Models differ in their balance of sensitivity and specificity, and these trade-offs influence their suitability for different public health purposes [71]. For example, models with higher specificity may reduce false positives but risk missing individuals who require support, whereas models with more balanced performance may be more appropriate for broader surveillance. Integrating ML tools into public health systems requires careful consideration of clinical workflows, equity in access to care, and the potential consequences of false negatives or false positives [72, 73]. ML models should complement, not replace, clinical judgment and established prevention strategies. With cautious implementation and adequate validation, ML-based approaches may offer incremental support to suicide-prevention initiatives, particularly in resource-constrained environments [74]. Although overall findings were encouraging, the review also showed that model performance varied substantially across studies. This variability is likely influenced by differences in study design, sample characteristics, outcome definitions, and data sources [62]. Studies using electronic health records often differ meaningfully from those relying on social media data or self-reported surveys, particularly in timing, measurement detail, and variable availability [67]. Threshold selection and classification cut-offs represent an additional source of heterogeneity, as inconsistent thresholds can inflate or reduce sensitivity and specificity estimates [75].

Concerns regarding risk of bias in primary studies also remain. Only a minority of studies applied formal bias-assessment tools such as PROBAST or QUADAS-2, and issues such as data leakage, suboptimal feature selection, and lack of external validation were frequently unaddressed [76, 77]. These limitations may lead to overly optimistic performance estimates and restrict generalizability. The findings highlight the need for rigorous internal and external validation and greater transparency in feature engineering.

While many ML models, particularly complex ensemble and neural network approaches, show promising predictive capabilities, their internal decision-making processes are often difficult to interpret [78]. In clinical psychiatry, this “black-box” nature raises concerns about trust, accountability, and ethical deployment [79]. Interpretability methods such as SHAP or LIME may help clarify feature contributions and support clinician understanding [80]. Incorporating such tools is important not only for model validation but also for ensuring ethically informed use that respects patient autonomy.

Limitations of the review

While this meta-analysis provides important insights, several limitations should be acknowledged. Considerable heterogeneity across study populations, data sources (e.g., electronic health records, surveys), and machine learning model specifications (e.g., feature selection, validation methods) may have influenced pooled estimates of sensitivity, specificity, and AUC. Subgroup analyses and meta-regression were initially planned to explore potential sources of heterogeneity; however, these analyses were not feasible due to insufficient reporting and lack of stratified data across primary studies. Unmeasured factors, such as differences in clinical settings, outcome definitions, or model training procedures, may therefore have contributed to the observed variability, particularly for models such as Artificial Neural Networks (ANN), which demonstrated wider confidence intervals. Another important limitation concerns outcome granularity. Most primary studies did not report separate sensitivity, specificity, or AUC estimates for suicidal ideation, suicide attempts, and suicide mortality. As a result, suicide-related outcomes were pooled, and outcome-specific subgroup meta-analyses could not be conducted. Additionally, applicability to specific diagnostic groups was restricted by limited reporting. Because most studies did not disaggregate diagnostic performance by conditions such as depression, bipolar disorder, schizophrenia/psychosis, or personality disorders, diagnosis-stratified analyses were not possible.

Third, the low base rate of suicide-related outcomes, even in high-risk populations, poses a challenge for diagnostic accuracy, particularly for PPV. This may have impacted the clinical utility of models with high specificity but lower sensitivity, such as ensemble models and Random Forest, potentially underestimating their effectiveness in low-prevalence settings.

Finally, restricting the search to 2010–2024 aligns with the modern ML and EHR era but might introduce time-lag bias; earlier studies seldom applied contemporary ML methods or reported paired diagnostic metrics suitable for pooling.

Conclusion

In summary, this review indicates that machine‑learning models—including ensemble and gradient‑boosting approaches—show generally encouraging diagnostic performance for suicide‑related outcomes in longitudinal and diagnostically diverse settings. However, because most primary studies did not report outcome‑ or diagnosis‑specific performance, these findings should be interpreted cautiously, and future research should provide stratified validation across diagnostic groups and outcome types. Although several models demonstrated promising accuracy, their clinical utility depends on more than statistical performance. Variability in study quality, threshold definitions, and risk of bias underscores the need for rigorous internal and external validation before these tools can be reliably applied in practice. The effective use of ML in suicide prevention will require careful integration into existing clinical and public health systems, along with attention to transparency, interpretability, and ethical considerations. Finally, due to limited reporting of outcome‑specific diagnostic indices, the present results should not be extrapolated across different suicide‑related outcomes. In particular, findings based on ideation or attempts cannot be assumed to reflect performance for suicide mortality. Outcome‑resolved evaluations are needed to determine whether ML models can accurately and ethically support mortality‑specific risk prediction.

Supplementary material

Supplementary material 1. (124.7KB, docx)

Acknowledgements

We would like to thank all the authors whose articles have been used in this meta-analysis.

Abbreviations

OR

Odds ratio

CI

Confidence interval

AUC

Area under the receiver operating curve

ML

Machine learning

PRISMA-DTA

The preferred reporting items for systematic reviews and meta-analyses of diagnostic test accuracy

PIRT framework

Participants, index test, reference standard, target condition

SVM

Support vector machine

ANN

Artificial neural networks

PPV

Positive predictive value

NPV

Negative predictive value

QUADAS-2

The quality assessment of diagnostic accuracy studies

SROC

Summary receiver operating characteristic

LR+

Positive likelihood ratio

LR–

Negative likelihood ratio

ANNs

Artificial neural networks

Author contributions

YM and SAS: concept development (provided idea for the research). BGH, HR, and PK: search strategy. PK, BGH, MA, and PK: data extraction. YM and SAS: supervision. BGH, MA, PK, and YM: analysis/interpretation. All authors: writing (each contributed substantially to drafting and revising the manuscript).

Funding

None.

Data availability

Data are available from the corresponding author upon reasonable request.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

All authors agree to publish.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Naghavi M. Global, regional, and national burden of suicide mortality 1990 to 2016: systematic analysis for the Global Burden of Disease Study 2016. BMJ. 2019;364:l94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Saab MM, et al. Suicide and Self-Harm Risk Assessment: A Systematic Review of Prospective Research. Archives Suicide Res. 2022;26(4):1645–65. [DOI] [PubMed] [Google Scholar]
  • 3.Ryan EP, Oquendo MA. Suicide Risk Assessment and Prevention: Challenges and Opportunities. Focus (Am Psychiatr Publ). 2020;18(2):88–99. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Bentley KH, et al. Implementing Machine Learning Models for Suicide Risk Prediction in Clinical Practice: Focus Group Study With Hospital Providers. JMIR Form Res. 2022;6(3):e30946. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Walsh CG, et al. Risk Model-Guided Clinical Decision Support for Suicide Screening: A Randomized Clinical Trial. JAMA Netw Open. 2025;8(1):e2452371. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Franklin JC, et al. Risk factors for suicidal thoughts and behaviors: A meta-analysis of 50 years of research. Psychol Bull. 2017;143(2):187–232. [DOI] [PubMed] [Google Scholar]
  • 7.Whiting D, Fazel S. How accurate are suicide risk prediction models? Asking the right questions for clinical practice. Evid Based Ment Health. 2019;22(3):125–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Large M, et al. Meta-Analysis of Longitudinal Cohort Studies of Suicide Risk Assessment among Psychiatric Patients: Heterogeneity in Results and Lack of Improvement over Time. PLoS ONE. 2016;11(6):e0156322. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Seyedsalehi A, Fazel S. Suicide risk assessment tools and prediction models: new evidence, methodological innovations, outdated criticisms. BMJ Ment Health, 2024. 27(1). [DOI] [PMC free article] [PubMed]
  • 10.Grendas LN, et al. Incidence rate of suicidal behavior stratified by diagnosis among high-risk patients. Psychiatry Res. 2025;343:116310. [DOI] [PubMed] [Google Scholar]
  • 11.Kaminsky Z, et al. Machine Learning-Based Suicide Risk Prediction Model for Suicidal Trajectory on Social Media Following Suicidal Mentions: Independent Algorithm Validation. J Med Internet Res. 2024;26:e49927. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Ehtemam H, et al. Role of machine learning algorithms in suicide risk prediction: a systematic review-meta analysis of clinical studies. BMC Med Inf Decis Mak. 2024;24(1):138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Linthicum KP, Schafer KM, Ribeiro JD. Machine learning in suicide science: Applications and ethics. Behav Sci Law. 2019;37(3):214–22. [DOI] [PubMed] [Google Scholar]
  • 14.Zang C, et al. Accuracy and transportability of machine learning models for adolescent suicide prediction with longitudinal clinical records. Transl Psychiatry. 2024;14(1):316. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Kusuma K, et al. The performance of machine learning models in predicting suicidal ideation, attempts, and deaths: A meta-analysis and systematic review. J Psychiatr Res. 2022;155:579–88. [DOI] [PubMed] [Google Scholar]
  • 16.Corke M, et al. Meta-analysis of the strength of exploratory suicide prediction models; from clinicians to computers. BJPsych Open. 2021;7(1):e26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Schafer KM, et al. A direct comparison of theory-driven and machine learning prediction of suicide: A meta-analysis. PLoS ONE. 2021;16(4):e0249833. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Pepe MS, et al. Limitations of the odds ratio in gauging the performance of a diagnostic, prognostic, or screening marker. Am J Epidemiol. 2004;159(9):882–90. [DOI] [PubMed] [Google Scholar]
  • 19.Steyerberg EW, et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21(1):128–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Horwitz AG, et al. Using machine learning with intensive longitudinal data to predict depression and suicidal ideation among medical interns over time. Psychol Med. 2023;53(12):5778–85. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Bracher-Smith M, Crawford K, Escott-Price V. Machine learning for genetic prediction of psychiatric disorders: a systematic review. Mol Psychiatry. 2021;26(1):70–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Torbahn G, et al. Surgery for the treatment of obesity in children and adolescents. Cochrane Database Syst Rev. 2022;9(9):pCd011740. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Frank RA, Bossuyt PM, McInnes MD. Systematic reviews and meta-analyses of diagnostic test accuracy: the PRISMA-DTA statement. 2018, Radiological Society of North America. pp. 313–314. [DOI] [PubMed]
  • 24.Whiting P, et al. The development of QUADAS: a tool for the quality assessment of studies of diagnostic accuracy included in systematic reviews. BMC Med Res Methodol. 2003;3:1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Wiens M, et al. A Tutorial and Use Case Example of the eXtreme Gradient Boosting (XGBoost) Artificial Intelligence Algorithm for Drug Development Applications. Clin Transl Sci. 2025;18(3):e70172. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. 2016. 785–794.
  • 27.Su R, John JR, Lin P-I. Machine learning-based prediction for self-harm and suicide attempts in adolescents. Psychiatry Res. 2023;328:115446. [DOI] [PubMed] [Google Scholar]
  • 28.Wright-Berryman J, et al. Virtually screening adults for depression, anxiety, and suicide risk using machine learning and language from an open-ended interview. Front Psychiatry. 2023;14:1143175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Min S, et al. Acoustic analysis of speech for screening for suicide risk: machine learning classifiers for between-and within-person evaluation of suicidality. J Med Internet Res. 2023;25:e45456. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Nielsen SD, et al. Prediction models of suicide and non-fatal suicide attempt after discharge from a psychiatric inpatient stay: A machine learning approach on nationwide Danish registers. Acta psychiatrica Scandinavica. 2023;148(6):525–37. [DOI] [PubMed] [Google Scholar]
  • 31.Broadbent M, et al. A machine learning approach to identifying suicide risk among text-based crisis counseling encounters. Front Psychiatry. 2023;14:1110527. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Huang Y, et al. Comparison of three machine learning models to predict suicidal ideation and depression among Chinese adolescents: a cross-sectional study. J Affect Disord. 2022;319:221–8. [DOI] [PubMed] [Google Scholar]
  • 33.Haghish E, Czajkowski NO, von Soest T. Predicting suicide attempts among Norwegian adolescents without using suicide-related items: a machine learning approach. Front Psychiatry. 2023;14:1216791. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Lee H, et al. Machine learning–based prediction of suicidality in adolescents with allergic rhinitis: derivation and validation in 2 independent Nationwide cohorts. J Med Internet Res. 2024;26:e51473. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Aldhyani TH, et al. Detecting and analyzing suicidal ideation on social media using deep learning and machine learning models. Int J Environ Res Public Health. 2022;19(19):12635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Agne NA, et al. Predictors of suicide attempt in patients with obsessive-compulsive disorder: an exploratory study with machine learning analysis. Psychol Med. 2022;52(4):715–25. [DOI] [PubMed] [Google Scholar]
  • 37.Park H, Lee K. Using boosted machine learning to predict suicidal ideation by socioeconomic status among adolescents. J personalized Med. 2022;12(9):1357. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Haghish E, et al. Suicide attempt risk predicts inconsistent self-reported suicide attempts: a machine learning approach using longitudinal data. J Affect Disord. 2024;355:495–504. [DOI] [PubMed] [Google Scholar]
  • 39.Nordin N, et al. An explainable predictive model for suicide attempt risk using an ensemble learning and Shapley Additive Explanations (SHAP) approach. Asian J psychiatry. 2023;79:103316. [DOI] [PubMed] [Google Scholar]
  • 40.Kim S, Lee K. The effectiveness of predicting suicidal ideation through depressive symptoms and social isolation using machine learning techniques. J Personalized Med. 2022;12(4):516. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Cansel N, et al. Interpretable estimation of suicide risk and severity from complete blood count parameters with explainable artificial intelligence methods. Psychiatria Danubina. 2023;35(1):62–72. [DOI] [PubMed] [Google Scholar]
  • 42.Reeves M, Bhat HS, Goldman-Mellor S. Resampling to address inequities in predictive modeling of suicide deaths. BMJ health care Inf. 2022;29(1):e100456. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Na K-S, Geem ZW, Cho S-E. The development of a suicidal ideation predictive model for community-dwelling elderly aged > 55 years. Neuropsychiatr Dis Treat. 2022;18:163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Li Q, Liao K. A multimodal prediction model for suicidal attempter in major depressive disorder. PeerJ. 2023;11:e16362. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Huang R, et al. Exploring the role of first-person singular pronouns in detecting suicidal ideation: A machine learning analysis of clinical transcripts. Behav Sci. 2024;14(3):225. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Shortreed SM, et al. Complex modeling with detailed temporal predictors does not improve health records-based suicide risk prediction. NPJ Digit Med. 2023;6(1):47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Kharrat GZ. Explainable artificial intelligence models for predicting risk of suicide using health administrative data in Quebec. PLoS ONE. 2024;19(4):e0301117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Arredondo Montero J. Diagnostic Test Accuracy Meta-Analysis: A Practical Guide to Hierarchical Models. J Surg Res. 2025;315:768–81. [DOI] [PubMed] [Google Scholar]
  • 49.Le Glaz A, et al. Machine Learning and Natural Language Processing in Mental Health: Systematic Review. J Med Internet Res. 2021;23(5):e15708. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Amanollahi M, et al. Machine learning applied to the prediction of relapse, hospitalization, and suicide in bipolar disorder using neuroimaging and clinical data: A systematic review. J Affect Disord. 2024;361:778–97. [DOI] [PubMed] [Google Scholar]
  • 51.Liu L, et al. Predictive Performance of Machine Learning for Suicide in Adolescents: Systematic Review and Meta-Analysis. J Med Internet Res. 2025;27:e73052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Spittal MJ, et al. Machine learning algorithms and their predictive accuracy for suicide and self-harm: Systematic review and meta-analysis. PLoS Med. 2025;22(9):e1004581. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Yıldız E. Machine Learning and Artificial Intelligence in Suicide Prevention: A Bibliometric Analysis of Emerging Trends and Implications for Nursing. Issues Ment Health Nurs. 2025;46(7):672–84. [DOI] [PubMed] [Google Scholar]
  • 54.Tio ES, Misztal MC, Felsky D. Evidence for the biopsychosocial model of suicide: a review of whole person modeling studies using machine learning. Front Psychiatry. 2023;14:1294666. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Oh M-Y, et al. Machine Learning–Based Explainable Automated Nonlinear Computation Scoring System for Health Score and an Application for Prediction of Perioperative Stroke: Retrospective Study. J Med Internet Res. 2025;27:e58021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Hunik L, et al. The role and utility of artificial intelligence and machine learning for diagnostic prediction in general practice. Eur J Gen Pract. 2026;32(1):2620908. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Kazienko P, Lughofer E, Trawiński B. Hybrid and ensemble methods in machine learning J. UCS special issue. J Univers Comput Sci. 2013;19(4):457–61. [Google Scholar]
  • 58.Dietterich TG. Ensemble methods in machine learning. in International workshop on multiple classifier systems. 2000. Springer.
  • 59.Mitchell R et al. Xgboost: Scalable gpu accelerated learning. arXiv preprint arXiv:1806.11248, 2018.
  • 60.Chen T. XGBoost: A Scalable Tree Boosting System. Cornell University; 2016.
  • 61.Evangelia C et al. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol, 2019. 110. [DOI] [PubMed]
  • 62.Christodoulou E, et al. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol. 2019;110:12–22. [DOI] [PubMed] [Google Scholar]
  • 63.Battineni G, Chintalapudi N, Amenta F. Machine learning in medicine: Performance calculation of dementia prediction by support vector machines (SVM). Inf Med Unlocked. 2019;16:100200. [Google Scholar]
  • 64.Maindola M et al. Utilizing random forests for high-accuracy classification in medical diagnostics. in. 2024 7th International Conference on Contemporary Computing and Informatics (IC3I). 2024. IEEE.
  • 65.Pigoni A, et al. Machine learning and the prediction of suicide in psychiatric populations: a systematic review. Translational psychiatry. 2024;14(1):140. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Belsher BE, et al. Prediction models for suicide attempts and deaths: a systematic review and simulation. JAMA psychiatry. 2019;76(6):642–51. [DOI] [PubMed] [Google Scholar]
  • 67.Su C, et al. Machine learning for suicide risk prediction in children and adolescents with electronic health records. Translational psychiatry. 2020;10(1):413. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Pauker SG, Kassirer JP. The threshold approach to clinical decision making. N Engl J Med. 1980;302(20):1109–17. [DOI] [PubMed] [Google Scholar]
  • 69.Kuhn E, et al. Interdisciplinary perspectives on digital technologies for global mental health. PLOS Global Public Health. 2024;4(2):e0002867. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Aguilera A. Digital technology and mental health interventions: Opportunities and challenges. Arbor. 2015;191(771):a210. [Google Scholar]
  • 71.Franklin JC, et al. Risk factors for suicidal thoughts and behaviors: A meta-analysis of 50 years of research. Psychol Bull. 2017;143(2):187. [DOI] [PubMed] [Google Scholar]
  • 72.Van Smeden M, et al. Clinical prediction models: diagnosis versus prognosis. J Clin Epidemiol. 2021;132:142–5. [DOI] [PubMed] [Google Scholar]
  • 73.Chen L. Overview of clinical prediction models. Annals translational Med. 2020;8(4):71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Obermeyer Z, et al. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447–53. [DOI] [PubMed] [Google Scholar]
  • 75.Reitsma JB, et al. Bivariate analysis of sensitivity and specificity produces informative summary measures in diagnostic reviews. J Clin Epidemiol. 2005;58(10):982–90. [DOI] [PubMed] [Google Scholar]
  • 76.Wolff RF, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51–8. [DOI] [PubMed] [Google Scholar]
  • 77.Moons KG, et al. PROBAST: a tool to assess risk of bias and applicability of prediction model studies: explanation and elaboration. Ann Intern Med. 2019;170(1):W1–33. [DOI] [PubMed] [Google Scholar]
  • 78.Steyerberg EW, Harrell FE Jr. Prediction models need appropriate internal, internal-external, and external validation. J Clin Epidemiol. 2015;69:245. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1(5):206–15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Lundberg SM, Lee S-I. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst, 2017. 30.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary material 1. (124.7KB, docx)

Data Availability Statement

Data are available from the corresponding author upon reasonable request.


Articles from Annals of General Psychiatry are provided here courtesy of BMC

RESOURCES