ABSTRACT
Background: Timely identification of individuals at risk for developing PTSD following trauma is crucial for providing targeted preventive interventions. Machine learning techniques show promise for deriving accurate prognostic screening instruments. However, accurate externally validated prognostic screening instruments for broad application in trauma-exposed civilians are not yet available. Moreover, it remains unknown whether prognostic screening instrument accuracy may be improved if developed in a sex-stratified manner.
Objective: We aimed to develop an externally validated prognostic PTSD screening instrument based on self-report information obtained within 2 months post-trauma in two independent cohorts of recently trauma-exposed civilians, using machine learning techniques allowing for extraction of a short screener. We examined whether separate models for males and females improved prognostic accuracy compared to sex-combined models.
Methods: Prognostic machine learning models (CART and XGBoost) were developed in a longitudinal cohort of N = 327 adults (38% females) requiring evaluation of (suspected) serious injury by an emergency department. External validation was performed in another longitudinal cohort of N = 466 adults (57% females) referred for emotional, practical or legal victim support following crime or traffic accidents. PTSD status at 1 year post-trauma was based on CAPS-IV for internal and PCL-5 for external validation.
Results: During internal validation, all models achieved excellent accuracy (AUC/sensitivity/specificity > 0.90). During external validation, sufficient accuracy was only achieved for the sex-combined XGBoost model (AUC = 0.73, sensitivity = 0.69, specificity = 0.68), including 22 items of demographic and health characteristics, trauma characteristics, peri-traumatic distress or dissociation, post-traumatic cognitions, PTSD symptoms and social support.
Conclusion: We developed an accurate externally validated short prognostic screening instrument for PTSD based on self-report questions that is applicable to a broad population of recently trauma-exposed civilians. This novel instrument enables timely identification of individuals at risk for PTSD following trauma, and research into early targeted interventions to prevent long-term PTSD for civilians following trauma.
KEYWORDS: Posttraumatic stress disorder (PTSD), trauma, longitudinal, machine learning, sex, prediction, social support, posttraumatic cognitions, peritraumatic dissociation, peritraumatic distress
HIGHLIGHTS
We developed an accurate externally validated short screening instrument for PTSD risk 1 year post-trauma.
It includes 22 self-report questions obtained within 2 months post-trauma and is applicable to a broad population of recently trauma-exposed civilians.
This novel instrument enables timely identification of individuals at risk for PTSD following trauma.
Abstract
Antecedentes: La identificación oportuna de las personas con riesgo de desarrollar TEPT tras un trauma es crucial para brindar intervenciones preventivas específicas. Las técnicas de aprendizaje automático son prometedoras para la obtención de instrumentos de cribado pronóstico precisos. Sin embargo, aún no se dispone de instrumentos de cribado de pronóstico precisos y validados externamente para su aplicación generalizada en civiles expuestos a traumas. Además, se desconoce si la precisión de los instrumentos de cribado pronóstico podría mejorarse si se desarrollaran estratificando por sexo.
Objetivo: Nuestro objetivo fue desarrollar un instrumento de cribado pronóstico para el TEPT, validado externamente, basado en información auto-reportada obtenida en los dos meses posteriores al trauma en dos cohortes independientes de civiles recientemente expuestos a traumas, utilizando técnicas de aprendizaje automático que permitieran la extracción de un cuestionario breve. Examinamos si los modelos separados para hombres y mujeres mejoraban la precisión pronóstica en comparación con los modelos combinados por sexo.
Métodos: Se desarrollaron modelos de aprendizaje automático pronóstico (CART y XGBoost) en una cohorte longitudinal de N = 327 adultos (38 % mujeres) que requerían evaluación por lesiones graves (sospechadas) en un servicio de urgencias. La validación externa se realizó en otra cohorte longitudinal de N = 466 adultos (57 % mujeres) derivados para recibir apoyo emocional, práctico o legal a víctimas de delitos o accidentes de tráfico. El diagnóstico de TEPT al año del trauma se basó en la CAPS-IV para la validación interna y en la PCL-5 para la validación externa.
Resultados: Durante la validación interna, todos los modelos alcanzaron una precisión excelente (AUC/sensibilidad/especificidad > 0.90). Durante la validación externa, solo el modelo XGBoost combinado por sexo alcanzó una precisión suficiente (AUC = 0.73, sensibilidad = 0.69, especificidad = 0.68), incluyendo 22 ítems de características demográficas y de salud, características del trauma, malestar o disociación peritraumática, cogniciones postraumáticas, síntomas de TEPT y apoyo social.
Conclusión: Desarrollamos un instrumento de cribado pronóstico breve, preciso y validado externamente para el TEPT, basado en preguntas de autoinforme y aplicable a una amplia población de civiles expuestos recientemente a un trauma. Este novedoso instrumento permite la identificación oportuna de personas en riesgo de desarrollar TEPT tras un trauma, así como la investigación de intervenciones tempranas y específicas para prevenir el TEPT a largo plazo en civiles que han sufrido un trauma.
PALABRAS CLAVE: Trastorno de estrés postraumático (TEPT), trauma, longitudinal, aprendizaje automático, sexo, predicción, apoyo social, cogniciones postraumáticas, disociación peritraumática, malestar peritraumático
Most individuals experience at least one potentially traumatic event (PTE) in their lifetime (e.g. 81.5% lifetime PTE prevalence in the Netherlands, Hoeboer et al., 2025). For a considerable proportion of exposed individuals, a traumatic event results in developing posttraumatic stress disorder (PTSD; e.g. 11.1% lifetime PTSD prevalence in the Netherlands, Hoeboer et al., 2025). PTSD is characterized by intrusions, avoidance, negative mood and cognitions, and alterations in arousal and reactivity (APA, 2013). Moreover, PTSD has been associated with a high risk of co-morbid anxiety and depression symptoms, reduced well-being and quality of life, and increased health-care and work-related costs (e.g. Geraerds et al., 2019; Karchoud et al., 2024a, 2024b; Kessler, 2000). We also observed these adverse psychological, functional and economic outcomes in our recent long-term data, together with a 4.8% PTSD prevalence 12–15 following trauma exposure, further underscoring the considerable risk of long-term PTSD (Karchoud et al., 2024a, 2024b). As PTSD by definition can only develop after traumatic events, the initial period post-trauma presents an opportunity for preventive interventions to reduce (long-term) PTSD and its associated adverse outcomes (e.g. Karchoud et al., 2024a, 2024b; Kessler, 2000). There are different approaches to determine which individuals should receive preventive interventions (U.S. Institute of Medicine Committee on Prevention of Mental Disorders; see Mrazek & Haggerty, 1994). Increasing evidence suggests that preventive interventions for PTSD should not be provided to all trauma-exposed individuals (i.e. universal interventions), as it is more promising to target preventive interventions to those at high risk for developing PTSD (i.e. selective interventions) or those with substantial early PTSD symptoms (i.e. indicated interventions; e.g. Ennis et al., 2018; Garcia & Delahanty, 2017). For example, systematic reviews of preventive interventions for individuals exposed to traumatic events showed that targeted preventive interventions resulted in better improvements in mental health outcomes compared to preventive interventions delivered to all trauma-exposed individuals (Bisson et al., 2021; Ennis et al., 2018). In order to target preventive interventions towards individuals who need them most and are most likely to benefit, it is crucial to timely and accurately identify individuals at risk for developing PTSD following trauma.
Previous studies identified early post-trauma risk and protective factors for later development of PTSD across various domains, such as demographic, socio-economic, psychiatric, psychosocial, biological, trauma history and environmental factors (see e.g. Tortella-Feliu et al., 2019 for umbrella review of systematic reviews and meta-analyses on risk factors for PTSD). A large body of studies has examined such risk and protective factors using traditional statistical approaches investigating linear associations (Tortella-Feliu et al., 2019, e.g. van der Mei et al., 2020). Currently, machine learning methods are increasingly applied to predict PTSD based on these previously identified risk and protective factors (see systematic reviews and meta-analysis Blekic et al., 2025 and Vali et al., 2025). These machine learning studies demonstrate that these previously observed risk and protective factors are indeed relevant for making individual predictions (Blekic et al., 2025). Multiple studies across various civilian populations have shown good classification accuracy of internally validated prognostic models, meaning models were derived in a subset of their sample (i.e. training set) and accuracy was subsequently tested in a separate set of their sample (i.e. test set; Schultebraucks et al., 2021; Hinrichs et al., 2019; Galatzer-Levy et al., 2017; Papini et al., 2018). A recent meta-analytic study showed a pooled AUC (Area Under the Receiver Operator Curve) of 0.81 for these internal validation-based prognostic models (Vali et al., 2025), which supports machine learning as a promising computational method to derive prognostic screening instruments for PTSD. However, such internally validated models have a considerable risk of overfitting, meaning that the model may have learned patterns specific to the training data which do not generalize well to new data. External validation (i.e. testing in independent sample) is necessary to ensure that derived screening models are generalizable beyond their original samples and the resulting screening instrument will perform well upon implementation in practice (Altman et al., 2009; Vali et al., 2025). There are currently only a few studies that performed external validation of prognostic PTSD models, generally showing poor accuracy (pooled AUC = 0.59; Vali et al., 2025). This discrepancy in accuracy between internal and external validation illustrates that prognostic models can appear accurate during internal validation yet fail to perform adequate in new samples during external validation, thereby restricting their applicability in practice. This emphasizes the need for accurate externally validated models. A potential strategy to prevent overfitting to the training data during internal validation and achieve higher accuracy during external validation is to use less complex machine learning models (Zhang et al., 2022).
Successful implementation of a screening instrument in practice depends not only on its accuracy but also on its feasibility. The previously mentioned externally validated classification models mainly rely on acute biomedical assessments and information retrieved from hospital patients’ records, limiting their applicability beyond acute medical care settings (Schultebraucks et al., 2021; Hinrichs et al., 2019; Galatzer-Levy et al., 2017; Papini et al., 2018). Including only self-report items would likely promote large-scale applicability within a broader population, as this would allow recently trauma-exposed individuals to complete the assessment independently without the need for a healthcare professional. This approach could also increase uptake and acceptability of a screening instrument by promoting user empowerment, engagement, and a sense of self-control (Huygens et al., 2016). Previous machine learning studies based on self-reported data early post-trauma were able to predict subsequent PTSD with fair accuracy in internal validation sets (AUC > 0.70; Kim et al., 2023; Papini et al., 2023). However, data within these models were collected either already at the emergency department (ED) during the first 24 h post-trauma limiting the applicability beyond acute medical care settings (AUC = 0.79 Kim et al., 2023), or in highly specific populations such as military personnel limiting the generalizability to broader trauma-exposed populations (AUC = 0.74; Papini et al., 2023). None of the models based solely on self-reported data have been externally validated in recently trauma-exposed civilians (Vali et al., 2025). Thus, there is still a need for an accurate externally validated risk screening instrument for end-point PTSD status based on self-reported data that is both feasible to implement in practice and is widely applicable to a broad range of trauma-exposed civilians. This may include purposefully identifying a minimal number of risk and protective factors into prognostic screening models to ensure they remain practical and efficient for real-world use (Vali et al., 2025). Moreover, no study has yet examined whether stratifying prognostic models by sex and/or gender improves accuracy. Given the previously observed difference in PTSD prevalence and differential prognostic value of various risk and protective factors for PTSD between men and women (Haering et al., 2024a; 2024b; 2024c; Tolin & Foa, 2008), applying sex-stratified models may offer a more precise approach to PTSD risk assessment.
The aim of the current study was to develop a short, accurate externally validated prognostic PTSD risk screening instrument based on self-report information in recently trauma-exposed civilians. We included 180 prognostic variables based on previously found risk and protective factors for PTSD (see e.g. Tortella-Feliu et al., 2019 for umbrella review). Within this study we first derived prognostic models for end-point PTSD status at 1 year post-trauma in ED patients after (suspected) serious injury based on self-report information collected within the first 2 months post-trauma. Our goal was to find a balance between the complexity of machine learning models and the usability of a future screening instrument. Subsequently, we externally validated the derived prognostic models for end-point PTSD status in a sample of recent victims of traffic accidents and crime, using the same self-report questions, time periods since trauma for the prognostic information and the outcome, and algorithms. We developed prognostic models in females and males separately as well as a sex-combined model to examine whether sex-differential screening instruments may be relevant for improving early PTSD risk detection.
1. Methods
1.1. Participants and study design
1.1.1. Model development sample
The model development sample was derived from the TraumaTIPS cohort (‘The Incidence, Prediction and Prevention of Post-trauma Psychopathology Study’; see Mouthaan et al., 2014 for more details). This sample consisted of N = 327 adults (37.6% female; mean age = 44.14 years, SD = 15.51) transported for medical evaluation of (suspected) serious injury by ambulance or helicopter to a level-1 emergency department (ED) of former hospitals in Amsterdam, the Netherlands (Academic Medical Center and VU University Medical Center, currently merged into Amsterdam University Medical Center) between 2005 and 2008. Participants were followed up to 1 year post-trauma. Inclusion criteria were: age 18 years or older; proficiency in Dutch; exposure to traumatic event according to Diagnostic and Statistical Manual of Mental Disorders 4th edition (DSM-IV) PTSD A1 criterion. Exclusion criteria were: current severe psychiatric symptoms (psychosis or schizophrenia; severe personality disorders; injuries resulting from deliberate self-harm); moderate-severe traumatic brain injury; permanent residency outside the Netherlands. The TraumaTIPS cohort study was approved by the Medical Ethics Review Committee of both hospitals (registration numbers 05-054# 05.17.0504; 06/039).
1.1.2. External validation sample
The external validation sample was derived from the 2-ASAP cohort (‘Towards Accurate Screening and Prevention for PTSD; see Karchoud et al., 2024a, 2024b for more details). This sample consisted of N = 466 adults (57.4% female; mean age = 46.69 years, SD = 17.97) referred (mostly by the police) for emotional, practical or legal support following crime or traffic accidents to Victim Support Netherlands (Slachtofferhulp Nederland) between 2022 and 2023. Victim Support Netherlands is the largest non-profit organization in the Netherlands to provide nationwide emotional support as well as practical and legal support related to the criminal justice and damage compensation process after accidents, crimes or calamities. Victim Support Netherlands provides support in person, by telephone or online. While they help individuals recognize and manage stress-related complaints, they do not provide professional psychological treatment and instead refer victims to their general practitioner if indicated. Participants were followed up to 1 year post-trauma. Inclusion criteria were: age 18 years or older; experience of a traumatic event according to Diagnostic and Statistical Manual of Mental Disorders 5th edition (DSM-5) PTSD A criterion maximally 2 months post-trauma at baseline; in the form of direct exposure to events with an acute onset; external cause of a civilian nature; and the potential to lead to serious physical injury. Exclusion criteria were: evidence of homocidality or suicidality (i.e. having attempted to kill oneself or another person); injuries due to intentional self-inflicted injury; evidence of ongoing or repeated trauma exposure, such as ongoing domestic violence; evidence of an inability to understand study procedures, risks or being otherwise unable to give informed consent; evidence of being unable to reliably follow protocol (including visual or cognitive or physical impairment precluding completion of protocol); impairment in ability to use or no regular access to e-mail and internet-connected smartphone, tablet or computer; insufficient understanding of Dutch language to follow protocol. The 2-ASAP cohort study was approved by the Medical Ethics Review Committee of Amsterdam UMC (registration number 2022.0030).
The 2-ASAP cohort included more participants who were exposed to traumatic events involving physical assault, while the TraumaTIPS cohort included more occupational, domestic or recreational accidents (Pearson Chi-square = 160.05, p < .001). Participants of the 2-ASAP cohort were also less often injured as a result of their traumatic event compared to those in the TraumaTIPS cohort (Pearson Chi-square = 22.37, p < .001). Furthermore, the 2-ASAP cohort included more females (Pearson Chi-Square = 29.55, p < .001), more participants who were currently employed (Pearson Chi-Square = 134.95, p < .001), in a relationship without living together (Pearson Chi-Square = 10.45, p = .034), and with higher education levels (Pearson Chi-Square = 61.82, p < .001) compared to the TraumaTIPS cohort. There were no significant cohort differences in parental status (p = .606) and age (p = .50). See Table 1 for sample characteristics at baseline and end-point PTSD measurements of both the model development sample (TraumaTIPS Cohort: ED Amsterdam UMC) and external validation sample (2-ASAP Cohort: Victim Support Netherlands).
Table 1.
Sample characteristics of model development sample (TraumaTIPS cohort: ED Amsterdam UMC) and external validation sample (2-ASAP cohort: Victim Support Netherlands) at baseline assessment < 2 months post-trauma.
| Model Development sample (TraumaTIPS cohort, N = 327) n |
External Validation sample (2-ASAP cohort, N = 446) n |
|
|---|---|---|
| Sex (females) | 123 (37.6%) | 256 (57.4%) |
| Age in years at baseline, M (SD) | 44.14 (15.51) | 46.69 (17.97) |
| Relationship status at baseline | ||
| Married/cohabitating/committed relationship | 213 (65.1%) | 324 (72.6%) |
| Divorced/widowed | 27 (8.3%) | 26 (5.8%) |
| No committed relationship | 87 (26.6%) | 96 (21.5%) |
| Children (yes) at baseline | 184 (56.3%) | 260 (58.3%) |
| Currently employed at baseline | 73 (22.3%) | 307 (68.8%) |
| Education, highest completed | ||
| Primary education/high school/secondary education | 143 (43.7%) | 128 (28.8%) |
| Secondary vocational education | 90 (27.5%) | 106 (23.8%) |
| Higher vocational education or University | 77 (23.6%) | 205 (46%) |
| Trauma type | ||
| Traffic accident | 221 (67.6%) | 298 (66.8%) |
| Physical assault | 9 (2.7%) | 131 (30.4%) |
| Occupational/domestic/recreational accident | 93 (28.4%) | 14 (3.8%) |
| Other (including e.g. natural disaster, fire or explosion) | 4 (1.2%) | 0 |
| Injuries due to index trauma (yes) | 312 (95.4%) | 378 (84.8%) |
| PTSD symptom severity at 1 year post-trauma, M (SD) | 16.01 (17.56) | 12.48 (13.03) |
| End-point PTSD at 1 post-trauma (yes) | 34 (8.3%) | 48 (10.8%) |
1.2. Procedures
1.2.1. Model development sample
After medical evaluation in the ED, potential participants for the TraumaTIPS cohort were identified by screening hospital patient records regarding the inclusion and exclusion criteria. Further eligibility screening was performed in the hospital or via telephone within 72 h post-trauma. At the baseline assessment (T0), participants were screened for the exclusion criteria of current severe psychiatric symptoms using the Mini International Neuropsychiatric Interview (MINI; Plus version 5.0; Sheehan et al., 1998; Van Vliet & De Beurs, 2007), and provided written and oral informed consent. Participants completed self-report questionnaires on potential risk and protective factors for PTSD at baseline and the first follow-up within 2 months post-trauma (range = 1–60 days, M = 23.57, SD = 13.44). They received the questionnaire on paper. Participants were followed up to 1 year post-trauma with PTSD symptoms assessed via diagnostic interview using the Clinician-Administered PTSD scale for DSM-IV (CAPS-IV; Hovens et al., 1994; Weathers et al., 2004) at 3, 9 and 12 months. We used the CAPS-IV at 12 months post-trauma (range = 316–744 days, M = 427.33, SD = 69.32) as outcome variable for end-point PTSD status.
1.2.2. External validation sample
Victim Support Netherlands identified potential participants for the 2-ASAP cohort via client records and sent them a letter via postal services, inviting them to contact the research team at Amsterdam UMC if they were interested in study participation. Potential participants who expressed interest in study participation, received additional study information and were called for eligibility screening (T0). Upon meeting all inclusion criteria and none of the exclusion criteria, informed consent was obtained through postal services. After inclusion, participants completed a baseline (T1) assessment within 2 months post-trauma (range = 16–60 days, M = 46.61, SD = 8.15), during which they completed the same self-report questionnaires on potential risk and protective factors for PTSD as in the model development sample. They received a personal link to the online questionnaire in Castor Electronic Data Capture (EDC). Participants were followed up to 1 year post-trauma with PTSD symptoms assessed using the PTSD checklist for DSM-5 (PCL-5; Boeschoten et al., 2014; Blevins et al., 2015; Hoeboer et al., 2024) at 3, 6, 9 and 12 months post-trauma. We used the PCL-5 at 12 months post-trauma (range = 360–434 days, M = 367.07, SD = 11.56) as outcome variable for end-point PTSD status.
1.3. Measures
1.3.1. Outcome variable: end-point PTSD status
CAPS-IV. The Dutch version of the CAPS-IV was used to assess end-point PTSD status in the model development sample (Hovens et al., 1994; Weathers et al., 2004). The CAPS-IV consists of 17 items that correspond to DSM-IV PTSD symptom criteria (5 items for re-experiencing; 7 items for avoidance; 5 items for hyperarousal), assessing both frequency and intensity of each symptom in the past month on a 4-point Likert scale, ranging from 0 ‘absent’ to 4 ‘extremely’. PTSD symptom severity total scores were calculated by summing frequency and intensity scores for all 17 items (range 0-136, with higher scores reflecting higher symptom severity). The CAPS-IV has excellent internal consistency (Hovens et al., 1994; Weathers et al., 2004). We used a cut-off total score of 45 as indicative of a probable PTSD diagnosis (Weathers et al., 1999).
PCL-5. The Dutch version of the PCL-5 was used to assess end-point PTSD status in the external validation sample (Blevins et al., 2015; Boeschoten et al., 2014; Boeschoten et al., 2018). The PCL-5 consists of 20 items that corresponds to DSM-5 PTSD symptom criteria (5 items for intrusions; 2 items for avoidance; 7 items for negative alterations in cognitions and mood; 6 items for hyperarousal), assessing how much participants have been bothered by each symptom on a 5-point Likert scale ranging from 0 ‘not at all’ to 5 ‘extremely’ in the past month. PTSD symptom severity total scores were calculated by summing all item scores (range 0-80, with higher scores reflecting higher symptom severity; Chronbach’s α current sample = 0.94). We used a previously established cut-off score of 29, which was identified as most accurate threshold for estimating probable PTSD prevalence based on the CAPS-5 (Hoeboer et al., 2024). This validation study was conducted in a follow-up study of the current TraumaTIPS cohort and showed high convergence between self-reported (i.e. PCL-5) and clinical assessment of PTSD symptom severity total scores (i.e. CAPS-5; Hoeboer et al., 2024; see also Hoeboer et al., 2025).
1.3.2. Prognostic variables: early-post trauma predictors
We evaluated 180 prognostic variables that were assessed early post-trauma (<2 months) from the following domains: demographic and health characteristics; medical and psychiatric history; psychological and physical health symptoms; current trauma characteristics and related acute emotions, posttraumatic cognitions and symptoms; perceived social support; prior trauma exposure. See Supplementary file A for an overview of all included prognostic variables.
1.3.3. Sex
Within the TraumaTIPS cohort sex assigned at birth (i.e. 62.4% male, 37.6% female) was retrieved from hospital records. Within the 2-ASAP cohort we assessed both sex assigned at birth and gender identity. We asked participants their sex assigned at birth (i.e. 42.6% male, 57.4% female), and whether they self-identified as men (43%), women (56.7%) or otherwise (0.2%). These percentages indicate substantial overlap between sex assigned at birth and gender identity within this sample. We chose to focus on sex assigned at birth for consistency with the TraumaTIPS cohort.
1.4. Statistical analyses
The statistical analyses plan was pre-registered on OSF (Karchoud et al., 2024a, 2024b). Statistical analyses were performed using R Version 3.6.1 and IBM SPSS Statistics Version 28.0. We reported all information necessary for quality assessment (i.e. risk of bias and applicability) of our prediction models using the ‘Prediction model Risk of Bias ASsement Tool’ (PROBAST) guidelines (Moons et al., 2019).
1.4.1. Missing data imputation
See Supplementary file A for an overview of missing data in the model development and external validations samples.
Model Development Sample. The percentage of missing data for prognostic variables in the model development sample ranged from 0% to 29.7%. Missing data was imputed prior to splitting the data into test/training sets (Tang & Ishwaran, 2017), using the non-parametric Random Forest algorithm to impute both categorical and continuous predictor data (Stekhoven & Bühlmann, 2012). We used the R package ‘missForest’ with 5 iterations (Stekhoven & Bühlmann, 2012). While Stekhoven and Bühlmann (2012) calculated Normalized Root-Mean-Squared Error (NRMSE) to evaluate performance of imputation, we chose not to standardize our variables given that variable distributions may differ between the model development and external validation samples. Performance of imputation was good based on the proportion of Falsely Classified (PFC = 0.12) and Root-Mean-Squared Error (RMSE = 0.63, higher scores due to non-standardized data; Stekhoven & Bühlmann, 2012).
External Validation Sample. The percentage of missing data of predictors variables in the external validation sample ranged from 0% to 5.8%, except for one variable measuring working hours per week with a missingness of 51.8%, because many participants who worked did not specify the follow-up question of amount of working hours per week. Missing data was imputed the same way as in the model development sample, using the R package ‘missForest’ with 4 iterations (Stekhoven & Bühlmann, 2012). Given that the external validation sample contained limited missing data (see Supplementary File A for an overview), the imputation procedure is unlikely to have influenced the external validation results. Performance of imputation was good based on the proportion of Falsely Classified (PFC = 0.01) and Root-Mean-Squared Error (RMSE = 0.92, higher scores due to non-standardized data; Stekhoven & Bühlmann, 2012).
1.4.2. Development of prognostic machine learning models
We developed prognostic machine learning models to predict end-point PTSD status, separately for males and females, and for both sexes combined. We conducted random synthetic oversampling of the minority class (i.e. those with PTSD) in males and females separately to ensure balanced classes of the outcome variables, as recommended to deal with imbalanced classes (López et al., 2013). For the sex combined sample, we applied oversampling within each sex separately to ensure an equal proportion of males and females (50%), in order to avoid biased model estimations for the underrepresented sex (Langeland & Olff, 2024). We used the ‘imbalance’ package in R for Majority Weighted Minority Oversampling Technique (MWMOTE; extension of synthetic minority over-sampling technique (SMOTE) algorithm; Cordón et al., 2018). The samples were randomly split into a 65% training set for model building and 35% test set for internal validation of the derived model.
We followed a stepwise approach of conducting different machine learning models to achieve good classification accuracy for internal and external validation. For internal validation, we aimed to achieve good accuracy (AUC = 0.80, sensitivity = 0.75, specificity = 0.70) as the minimally acceptable accuracy for our prognostic models. For external validation, we aimed to achieve fair accuracy (AUC = 0.70) and a sensitivity of 0.65 and specificity of 0.60, considering that external validation generally results in lower accuracy (Siontis et al., 2015; Vali et al., 2025). We choose to perform tree-based models, given results of a meta-analytic study showing that tree-based models generally performed better in predicting PTSD than other models (Vali et al., 2025). We evaluated 180 potential risk and protective factors as input variables in the algorithms, from which the algorithm derived a minimal feature set consisting of the variables included in the final models. First, we performed the Classification And Regression Tree (CART; Breiman et al., 2017), a pragmatic tree-based algorithm suited for our goal to identify a minimal set of variables in the prognostic models. We performed CART using the ‘rpart’ package in R (Therneau et al., 2015). We build an initial tree using 5-fold cross-validation on the training set, while avoiding excessive complexity of the model to avoid the risk of overfitting (i.e. maximum depth of tree = 10; complexity parameter = 0.005; minimum number of observations required to split a node = 5). Subsequently, we used a pruning strategy to optimize the amount of features within the models by determining the complexity parameter based on the ‘1-SE’ rule (Breiman et al., 2017). Second, we proceeded with a more advanced tree-based algorithm, Extreme Gradient Boosting (XGBoost; Chen et al., 2015), an ensemble tree-based method that combines multiple decision trees to improve classification accuracy through iterative boosting. We performed XGBoost using the ‘caret’ and ‘xgboost’ packages in R (Chen et al., 2019; Kuhn, 2008). We build a tree using 5-fold cross-validation on the training set, while avoiding excessive complexity of the model to avoid the risk of overfitting and to identify a minimal feature set (i.e. number of trees = 50; maximum depth of tree = 1; eta learning rate = 0.2; gamma complexity regularization = 3; subsample size per iteration = 0.7; percentages features per tree = 0.5; minimal child weight for tree node = 2). The final models generated from the training sets were applied to the test sets to calculate the accuracy parameters. We calculated the overall accuracy (i.e. ratio correct predictions and all predictions); AUC (Area Under the Receiver Operator Curve; i.e. capability to distinguish different classes); sensitivity (i.e. proportion true positives correctly classified); specificity (i.e. proportion true negatives correctly classified); precision (i.e. ratio true positives and all predictions); recall (i.e. ratio true positives and actual positive predictions).
1.4.3. External validation of derived models
External validation was performed on the derived models of end-point PTSD status for females and males separately and in a sex-combined sample. The algorithms generated from the model development sample were applied to the external validation test set to calculate the same accuracy parameters as for internal validation (i.e. overall accuracy; AUC; sensitivity; specificity; precision; recall).
1.4.4. SHapley Additive exPlanations (SHAP)
To provide interpretations of our derived machine learning model, we applied SHapley Additive exPlanations (SHAP), using the ‘shapviz‘ and ‘kernelshap’ packages in R (Lundberg & Lee, 2017). SHAP is used for decision tree-based non-linear models, such as XGBoost, to allow for interpretable insights into how individual variables contribute to the model’s prediction (Lundberg & Lee, 2017). We assessed the relative importance of each predictor in the final model (i.e. SHAP values), and the direction of its influence on the predicted outcome (i.e. feature values).
2. Results
2.1. Model development and validation
See Table 2 for an overview of the accuracy parameters for internal and external validation of CART and XGBoost prognostic models for end-point PTSD status 1 year post-trauma in males, females, and sex combined.
Table 2.
Accuracy of CART and XGBoost prognostic models for end-point PTSD status in males, females, and sex combined.
| n variables in model | Internal validation Accuracy (95% CI) |
Internal validation AUC; Sensitivity; Specificity |
Internal validation Precision; Recall |
External validation Accuracy (95% CI) |
External validation AUC; Sensitivity; Specificity |
External validation Precision; Recall |
|
|---|---|---|---|---|---|---|---|
| CART model | |||||||
| Females | 4 | 95% (0.88–0.99) | 0.95; 0.98; 0.93 | 0.93; 0.98 | 80% (0.75–0.85) | 0.45; 0.00; 0.90 | 0.00; 0.00 |
| Males | 4 | 97% (0.93–0.99) | 0.97; 1.00; 0.94 | 0.94; 1.00 | 86% (0.81–0.91) | 0.55; 0.16; 0.94 | 0.23; 0.16 |
| Sex combined | 6 | 94% (0.89–0.97) | 0.96; 0.96; 0.93 | 0.85; 0.96 | 81% (0.77–0.84) | 0.68; 0.31; 0.87 | 0.22; 0.31 |
| XGBoost model | |||||||
| Females | 19 | 95% (0.88–0.99) | 0.98; 0.98 0.93 | 0.93; 0.98 | 82% (0.76–0.86) | 0.57; 0.03; 0.92 | 0.05; 0.03 |
| Males | 19 | 96% (0.92–0.99) | 1.00; 1.00; 0.93 | 0.93; 1.00 | 80% (0.74–0.85) | 0.74; 0.53; 0.83 | 0.26; 0.53 |
| Sex combined | 22 | 93% (0.89–0.94) | 0.98; 0.92; 0.93 | 0.86; 0.89 | 68% (0.63–0.72) | 0.73; 0.69; 0.68 | 0.20; 0.69 |
2.1.1. Internal validation
We achieved excellent accuracy for all sex-specific and sex-combined models in the internal validation test sets, using CART (overall accuracy = 94-97%, AUC = 0.95–0.97, sensitivity 0.96–1.00, specificity 0.93–0.94) and XGBoost (overall accuracy = 93-96%, AUC = 0.98–1.00, sensitivity 0.92–1.00, specificity 0.93).
2.1.2. External validation
External validation of the CART sex-combined model resulted in poor accuracy (AUC = 0.68; specificity = 0.87; sensitivity = 0.16). This model performed better than the separate models for males (AUC = 0.55, sensitivity = 0.16, specificity = 0.94) and females (AUC = 0.45, sensitivity = 0.00, specificity = 0.90).
External validation of the XGBoost sex-combined model resulted in fair accuracy (AUC = 0.73) with a specificity of 0.69 and sensitivity of 0.68. This model performed better than the separate models for males with fair accuracy (AUC = 0.74) but a poor sensitivity of 0.53 and a specificity of 0.83; and females with poor accuracy (AUC = 0.57, sensitivity = 0.03, specificity = 0.92).
2.2. Prognostic screener for end-point PTSD status
The sex-combined XGBoost model was the only prognostic model that achieved sufficient external validation accuracy and was thus used to obtain the prognostic screening instrument for end-point PTSD status. The model included 22 questions from the following domains: demographic and health characteristics; current trauma characteristics; peri-traumatic distress or dissociation; post-traumatic cognitions; PTSD symptoms; social support. See Table 3 for an overview of each predictor per domain included in the final derived prognostic screener for end-point PTSD status. See Figure 1 for SHAP values indicating the relative importance of each predictor in the final derived model and the direction of its influence on the chance of developing end-point PTSD and of developing no end-point PTSD. See supplementary file B for instructions to use the algorithm that was used to derive the model.
Table 3.
Overview of variables per domain included in the final prognostic screener for end-point PTSD status (sex-combined XGBoost model).
| Included variables in prognostic screener for end-point PTSD status per domain |
|---|
|
Demographics and health characteristics Sex; Age at trauma; Hours of weekly exercise before trauma; Use of psychopharmacological medication before trauma; The number of glasses of alcohol consumed on a typical drinking day in the past year (AUDIT Q2); |
|
Current trauma characteristics Duration of hospital stay in days; Duration of trauma in minutes; Sustained head injury during trauma; |
|
Peri-traumatic distress or dissociation Lost track of what was going on by blanking out or spacing out, or not feeling part of what was going on during or immediately after trauma (PDEQ Q1); Felt ashamed of emotional reactions during or immediately after trauma (PDI Q6); Felt helpless during trauma; |
|
Post-traumatic cognitions Thoughts since trauma that the event happened because of the way I acted (PTCI Q1); Thoughts since trauma that people cannot be trusted (PTCI Q7); Thoughts since trauma that I felt isolated and set apart from others (PTCI Q23); |
|
PTSD symptoms Felt irritable and angry in the past week (IES-R Q4) |
|
Social support Received support for emotional problems (SSLD Q2); Received informative support (SSLD Q6); Received support for social companionship (SSLD Q5); Received too little support for social companionship (SSLD Q5 discrepancy between desired and received); Received too little for esteem support (SSLD Q3 discrepancy between desired and received); Received too much support for social companionship (SSLD Q5 discrepancy between desired and received); Received too much daily emotional support (SSLD Q1 discrepancy between desired and received). |
Note: AUDIT: Alcohol Use Disorders Identification Test (Bush et al., 1998); PDEQ: Peritraumatic Dissociative Experiences Questionnaire (Marmar et al., 2004); PDI: Peritraumatic Distress Inventory (Brunet et al., 2001); PTCI: Posttramatic Cognitions Inventory (Foa et al., 1999); IES-R: Impact of Event Scale Revised (Weiss, 2007); VVV: Verkorte Vermoeidsheidsvragenlijst (Alberts et al., 1997); HADS: Hospital Anxiety and Depression Scale (Spinhoven et al., 1997); SSLD: Sociale Steun Lijst Discrepanties (van Sonderen, 1997).
Figure 1.
SHAP values indicating the relative importance of each predictor in the final derived sex-combined XGBoost model and the direction of its influence on end-point PTSD risk.
Note. SHAP values indicating the relative importance of each predictor in the final derived model (i.e. higher absolute values on x-axis indicating greater importance; predictors on y-axis ranked from most to least important), and the direction of its influence (i.e. low to high feature values indicated by colour per participant) on the chance of developing end-point PTSD (i.e. positive SHAP values on x-axis) and of developing no end-point PTSD (i.e. negative SHAP values on x-axis).
3. Discussion
The aim of the study was to develop an externally validated prognostic screening instrument for PTSD that is applicable to recently trauma-exposed civilians. To achieve this, we used machine learning techniques to extract a short prognostic screening instrument for end-point PTSD status 1 year post-trauma based on self-report information obtained within 2 months post-trauma. The prognostic models were first developed in an emergency department cohort primarily exposed to traffic accidents and occupational, domestic or recreational accidents involving (suspected) serious injury. Subsequently, we performed external validation on the derived models to test the generalizability in a cohort from a national victim support organization with more heterogeneous types of trauma (i.e. more physical assault and also traffic accidents), and who were less likely to be injured, and less severely injured. In addition to recruitment centre and trauma type, the cohorts also differed in socio-demographic characteristics, including sex, education level, relationship status, and current employment status. We also compared the accuracy of separate models for males and females and sex-combined models to examine whether sex-differential screening instruments may be relevant for improving prognostic screening performance. The only model with fair accuracy for external validation was the sex-combined XGBoost model (AUC = 0.73; specificity = 0.69; sensitivity = 0.68). The differences between cohorts further underscores our model’s likely generalizability to different populations of recently trauma-exposed civilians, thereby highlighting its potential for real-world implementation.
Within our study we included 180 predictors (i.e. self-report items) in the algorithm based on previously established risk and protective factors for PTSD (see systematic review Tortella-Feliu et al., 2019). Our final derived screening instrument included 22 self-report items. The most important predictor for PTSD risk was feeling helpless during trauma. We found multiple predictors related to demographic, health and trauma characteristics, which were all already well-documented risk factors for PTSD, particularly females and younger individuals are often considered at higher risk for PTSD, as well as trauma severity indicators such as duration of hospital stay (Tolin & Foa, 2008; Tortella-Feliu et al., 2019). Additionally, the same or similar predictors as in our model have also been reported in a recent systematic review of longitudinal machine learning studies on predictors of PTSD (i.e. sex; younger age; previous psychological treatment; peri-traumatic dissociation and distress; post-traumatic cognitions; PTSD symptoms; alcohol use and perceived social support; Blekic et al., 2025). Notably, some of these predictors were also supported in external validation studies (i.e. younger age; psychiatric history; alcohol use; peri-traumatic distress and dissociation; and social support; Blekic et al., 2025). Within these studies external validation was performed in a similar sample in which the model was built, both ED cohorts (Schultebraucks et al., 2020) or comparable military deployment cohorts (Karstoft et al., 2020; Papini et al., 2023). Our study, including external validation in a different cohort type, extends these findings by demonstrating that similar predictors hold in different populations of recently trauma-exposed civilians.
A recent systematic review of machine learning studies examined whether predictors identified using these data-driven methods align with leading theoretical models of PTSD (Blekic et al., 2025). They reported multiple predictors that we also found within this study regarding psychiatric history; peri-traumatic distress and dissociation; as well as social support; that all align with the cognitive model of PTSD (Ehlers & Clark, 2000). Our model additionally identified predictors of post-traumatic cognitions (i.e. mistrust in others, self-blame and feelings of isolation) that align with both the cognitive model (Ehlers & Clark, 2000) as well as the social cognitive model (Sharp et al., 2012), emphasizing maladaptive understandings of the self- and other that hinder social support and increase the vulnerability of developing PTSD. This is also consistent with our derived predictors related to discrepancies in received and desired social support, with both too little and too much perceived social support as risk factor for PTSD. This social support domain was most represented in the final model. On the one hand, this is not surprising as there is robust evidence for the importance of social support as protective factor for trauma-related disorders (Brewin et al., 2000; Maercker, 2025; Santos et al., 2025; Wang et al., 2021). On the other hand, it may be considered surprising that the final model only included one PTSD symptom as relevant (i.e. felt irritable and angry). Our results suggest that early risk screening for PTSD based on risk and protective factors across multiple domains may be more accurate, rather than relying solely on acute PTSD symptoms alone as captured by validated PTSD symptom screeners such as the Impact Event Scale Revised (IES-R; Christianson & Marren, 2012). Future research should directly compare the accuracy of our derived multi-domain data-driven prognostic screening instrument based on machine learning algorithms to the accuracy of such validated PTSD symptom screeners that assess current PTSD, screening instrument based on theoretical constructs, and other best practice screening instruments that are currently available for PTSD, such as the primary care PTSD screen for DSM-5 (PC-PTSD-5; Prins et al., 2016); Global Psychotrauma Screen (GPS; Frewen et al., 2021; Olff et al., 2020). Whilst our application of SHAP provided a first step into understanding the relative contributions of the 22 items within the derived algorithm, future research could also apply network analyses to investigate whether identified risk and protective factors are similary associated amongst individuals with and without high risk for PTSD. Taken together, our findings demonstrate considerable convergence in predictors with previous machine learning studies and theoretical models for PTSD, while also extending the literature by highlighting the utility of self-report questions in a broad population of recently trauma-exposed civilians.
Within this study we tried to find a balance between model complexity and practicality for use of a screening instrument with fewer items. We started with a relatively simple tree-based algorithm (i.e. CART), however this did not generalize well as we found poor accuracy during external validation. Subsequently we performed an ensemble method that combines multiple decision trees which improves generalization to unseen cases (i.e. XGBoost algorithm; Chen et al., 2015). The sex-combined XGBoost model achieved higher accuracy for external validation (AUC = 0.73) than the pooled prevalence of the few studies that performed external validation of prognostic PTSD models thus far (pooled AUCs = 0.59; Vali et al., 2025). The poorer accuracy of external validation observed in previous studies has been proposed to result from overfitting during internal validation due to the use of complex machine learning models that are overfitted to the training set (Zhang et al., 2022). Although other studies also performed XGBoost, we simplified the model by adjusting complexity parameters during model development to limit the number of items selected by the machine learning model (e.g. maximum depth of a tree). We did not build more complex models (e.g. allowing deeper trees), as this would have conflicted with the study aim to develop a concise screening instrument. Moreover, this also prevented the model from overfitting to the training set. Although our derived model achieved fair prognostic accuracy for external validation (AUC = 0.73), the model achieved modest sensitivity (0.69), specificity (0.68), recall (0.69) and precision (0.20). This means that while the model correctly identified the majority of individuals who develop PTSD (i.e. true positives), it also incorrectly classified a substantial number of individuals at risk who ultimately do not develop PTSD (i.e. false positives). While this low precision is not ideal, it is not necessarily problematic in a screening context, as these individuals could still experience subclinical PTSD symptoms or other adverse psychological outcomes and as such may still benefit from preventive interventions. Additionally, the clinical relevance of subclinical PTSD and its potential to cause distress and impairment is now acknowledged by the DSM-5 (i.e. under ‘Other Specified Trauma- and Stressor-Related Disorder’). Thus, even with limited precision, the model may serve a valuable role in identifying those who are most likely to benefit from preventive intervention.
Future research should keep improving classification accuracy by further updating these prognostic machine learning models. We built our prognostic models based on previously established risk and protective factors that were established at the time of data collection in the model development sample (TraumaTPS cohort; Mouthaan et al., 2014), but it may also be valuable to explore other potential risk and protective factors for PTSD, for example those identified in the systematic review of machine learning studies with external validation sets (Blekic et al., 2025). Within this study we used self-report questions to promote large-scale applicability of the derived screening instrument for PTSD risk. In particular, it would be interesting to examine other predictors assessed via self-report questions, such as those related to sleep problems or coping strategies (Blekic et al., 2025).
Given that sex assigned at birth was found as predictor in our final model, future research could benefit from including sex-specific and gender-specific risk factors in stratified models, for example by examining factors related to ovarian steroid hormonal variation across the female lifespan (Wiseman et al., 2023) or related to types of interpersonal trauma more often experienced by women such as sexual assault (Hoeboer et al., 2025). Sex-specific models performed well for internal validation, but poor in the external validation phase, particularly for females compared to males. This could be explained overfitting due to the reduced sample size following stratification by sex (e.g. Larracy et al., 2021). We encourage other researchers to examine adequately powered sex-stratified machine learning models including sex- and gender-specific risk factors alongside sex-combined models. Besides improving prognostic accuracy, accurate sex-stratified models may also offer valuable insights into potential sex-specific risk and protective factors (see systematic review and meta-analysis on sex/gender differences in PTSD risk factors; Haering et al., 2024a, 2024b, 2024c).
The current study has several limitations. Our initial aim was to predict distinct courses of PTSD symptoms over time using latent trajectory analyses rather than a single end-point status of PTSD (see protocol study Karchoud et al., 2024a, 2024b). However, this turned out not to be feasible due to substantial differences in emerging latent PTSD symptom trajectories and model fit between the internal and external validation samples, which would have resulted in predicting fundamentally different outcomes across cohorts. In order to be able to perform external validation, we therefore focused on predicting end-point PTSD status instead, which still served the aim of the study to identify individuals at risk for developing PTSD over time. However, the potential impact of the use of different instruments to assess end-point PTSD status in the internal and external validation cohorts must be noted. Although previous research demonstrated high convergence between self-reported (i.e. PCL-5) and clinical assessment (i.e. CAPS-5) of PTSD symptom severity total scores (Hoeboer et al., 2024), the use of different outcome measures may have introduced variability in the classification of PTSD cases. Given that the model has been trained to capture PTSD based on the CAPS-IV assessment, the model may not fully capture PTSD based on the PCL-5 assessment. This may have affected the model’s accuracy during external validation, potentially missing individuals with PTSD based on the PCL-5 or falsely classifying individuals with PTSD based on the CAPS-IV. However, even if both cohorts had used clinician-rated assessments (i.e. CAPS), this would still have resulted in a discrepancy, as the internal cohort was assessed using the CAPS-IV based on DSM-IV criteria, while the external cohort would require the CAPS-5 based on DSM-5. Moreover, differences in the timing of data collection of the risk and protective factors for PTSD may have influenced model performance during external validation. Although both assessed within 2 months post-trauma, the predictors in the model development sample were assessed immediately post-trauma while this was a few weeks later in the external validation cohort. Potentially, the differences in the timing of the data collection of the included risk and protective factors in the model development and external validation samples may have influenced the PTSD symptom presentation, as the external validation sample was assessed at a later time point post-trauma and participants may therefore have been in a different phase of symptom progression or recovery. While these discrepancies between the model development and external validation cohorts may have introduced lower prediction accuracy, they also underscore the robustness of the derived model across diverse assessment contexts. Last, while the hospital records used for our model development and internal validation cohort most likely contained information on biological sex, it is not entirely certain whether sex or gender was assessed.
For future implementation purposes, we tested our derived prognostic models in another independent cohort of recently trauma-exposed civilians. We also focused on deriving short and interpretable models from large-scale self-report data. This minimizes overwhelming users with excessive item burden when filling in the screening instrument, supporting its potential use as a general prognostic screener for widespread implementation in real-world settings. This facilitates identification of individuals at risk for developing end-point PTSD 1 year post-trauma, enabling targeted preventive interventions for PTSD. The screener including the algorithm developed in this study will be freely available. We will pilot the clinical utility in a forthcoming randomized controlled trial (RCT) indicated early intervention for PTSD towards those recognized to be at high PTSD risk using the screener.
4. Conclusion
We developed a short screening instrument for PTSD risk in recently trauma-exposed civilians based on a machine learning algorithm using information derived from self-report questions. By externally validating the derived prognostic machine learning model, we demonstrated its generalizability across different populations of recently trauma-exposed civilians. This strengthens the potential for real-world implementation of our screening instrument for early PTSD risk screening. With this study, we take a step towards targeted early interventions for civilians in the aftermath of a traumatic event.
Supplementary Material
Funding Statement
This study was funded by the Netherlands Organization for Health Research and Development (ZonMw #62300038; #636340004) and Achmea Stichting Slachtoffer en Samenleving (SASS).
Disclosure statement
No potential conflict of interest was reported by the author(s).
Data and code availability
The data of the study and the code to produce the results described in this paper are available at Open Science Framework (OSF; https://osf.io/59eax/). The 2-ASAP and TraumaTIPS cohorts are registered in the FAIR Traumatic Stress Data Sets library of the Global Collaboration on Traumatic Stress (GCTS).
Supplemental Material
Supplemental data for this article can be accessed online at https://doi.org/10.1080/20008066.2025.2594266.
References
- Alberts, M., Smets, E. M. A., Vercoulen, J. H. M. M., Garssen, B., & Bleijenberg, G. (1997). ‘Verkorte vermoeidheidsvragenlijst': een praktisch hulpmiddel bij het scoren van vermoeidheid. [PubMed]
- Altman, M. B., Stinauer, M. A., Javier, D., Smith, B. D., Herman, L. C., Pytynia, M. L., … Roeske, J. C. (2009). Validation of temporal optimization effects for a single fraction of radiation in vitro. International Journal of Radiation Oncology*Biology*Physics, 75(4), 1240–1246. 10.1016/j.ijrobp.2009.06.076 [DOI] [PubMed] [Google Scholar]
- APA . (2013). Diagnostic and statistical manual of mental disorders (5th ed.). American Psychiatric Association. [Google Scholar]
- Bisson, J. I., Wright, L. A., Jones, K. A., Lewis, C., Phelps, A. J., Sijbrandij, M., … Roberts, N. P. (2021). Preventing the onset of post traumatic stress disorder. Clinical Psychology Review, 86, 102004. 10.1016/j.cpr.2021.102004 [DOI] [PubMed] [Google Scholar]
- Blekic, W., D’Hondt, F., Shalev, A. Y., & Schultebraucks, K. (2025). A systematic review of machine learning findings in PTSD and their relationships with theoretical models. Nature Mental Health, 139–158. 10.1038/s44220-024-00365-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Blevins, C. A., Weathers, F. W., Davis, M. T., Witte, T. K., & Domino, J. L. (2015). The posttraumatic stress disorder checklist for DSM-5 (PCL-5): Development and initial psychometric evaluation. Journal of Traumatic Stress, 28(6), 489–498. 10.1002/jts.22059 [DOI] [PubMed] [Google Scholar]
- Boeschoten, M. A., Bakker, A., Jongedijk, R., & Olff, M. (2014). PTSS checklist voor de DSM-5 [PTSD checklist for DSM-5]. Arq Nationaal Psychotrauma Centrum. 30. Blevins CA. [Google Scholar]
- Boeschoten, M. A., Van der Aa, N., Bakker, A., Ter Heide, F. J. J., Hoofwijk, M. C., Jongedijk, R. A., … Olff, M. (2018). Development and evaluation of the Dutch clinician-administered PTSD scale for DSM-5 (CAPS-5). European Journal of Psychotraumatology, 9(1). 10.1080/20008198.2018.1546085 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Breiman, L., Friedman, J., Olshen, R. A., & Stone, C. J. (2017). Classification and regression trees. Routledge. [Google Scholar]
- Brewin, C. R., Andrews, B., & Valentine, J. D. (2000). Meta-analysis of risk factors for posttraumatic stress disorder in trauma-exposed adults. Journal of Consulting and Clinical Psychology, 68(5), 748. 10.1037/0022-006X.68.5.748 [DOI] [PubMed] [Google Scholar]
- Brunet, A., Weiss, D. S., Metzler, T. J., Best, S. R., Neylan, T. C., Rogers, C., … Marmar, C. R. (2001). The Peritraumatic Distress Inventory: A proposed measure of PTSD criterion A2. American Journal of Psychiatry, 158(9), 1480–1485. 10.1176/appi.ajp.158.9.1480 [DOI] [PubMed] [Google Scholar]
- Bush, K., Kivlahan, D. R., McDonell, M. B., Fihn, S. D., Bradley, K. A., & Ambulatory Care Quality Improvement Project (ACQUIP). (1998). The AUDIT alcohol consumption questions (AUDIT-C): An effective brief screening test for problem drinking. Archives of Internal Medicine, 158(16), 1789–1795. 10.1001/archinte.158.16.1789 [DOI] [PubMed] [Google Scholar]
- Chen, T., He, T., Benesty, M., Khotilovich, V., Tang, Y., Cho, H., … & Zhou, T. (2015). Xgboost: extreme gradient boosting. R Package Version 0.4-2, 1(4), 1–4. [Google Scholar]
- Chen, T., He, T., Benesty, M., & Khotilovich, V. (2019). Package ‘xgboost’. R version, 90(166), 40. [Google Scholar]
- Christianson, S., & Marren, J. (2012). The impact of event scale-revised (IES-R). Medsurg Nursing, 21(5), 321–322. [PubMed] [Google Scholar]
- Cordón, I., García, S., Fernández, A., & Herrera, F. (2018). Imbalance: Oversampling algorithms for imbalanced classification in R. Knowledge-Based Systems, 161, 329–341. 10.1016/j.knosys.2018.07.035 [DOI] [Google Scholar]
- Ehlers, A., & Clark, D. M. (2000). A cognitive model of posttraumatic stress disorder. Behaviour Research and Therapy, 38(4), 319–345. 10.1016/S0005-7967(99)00123-0 [DOI] [PubMed] [Google Scholar]
- Ennis, N., Sijercic, I., & Monson, C. M. (2018). Internet-delivered early interventions for individuals exposed to traumatic events: Systematic review. Journal of Medical Internet Research, 20(11), e280. 10.2196/jmir.9795 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Foa, E. B., Ehlers, A., Clark, D. M., Tolin, D. F., & Orsillo, S. M. (1999). The posttraumatic cognitions inventory (PTCI): Development and validation. Psychological Assessment, 11(3), 303. 10.1037/1040-3590.11.3.303 [DOI] [Google Scholar]
- Frewen, P., McPhail, I., Schnyder, U., Oe, M., & Olff, M. (2021). Global Psychotrauma Screen (GPS): Psychometric properties in two internet-based studies. European Journal of Psychotraumatology, 12(1), 1881725. 10.1080/20008198.2021.1881725 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Galatzer-Levy, I. R., Ma, S., Statnikov, A., Yehuda, R., & Shalev, A. Y. (2017). Utilization of machine learning for prediction of post-traumatic stress: A re-examination of cortisol in the prediction and pathways to non-remitting PTSD. Translational Psychiatry, 7(3), e1070–e1070. 10.1038/tp.2017.38 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Garcia, M. A., & Delahanty, D. L. (2017). Oxytocin and other pharmacologic preventive interventions for posttraumatic stress disorder: Not a one-size-fits-all approach. Biological Psychiatry, 81(12), 977–978. 10.1016/j.biopsych.2017.04.001 [DOI] [PubMed] [Google Scholar]
- Geraerds, A. J., Haagsma, J. A., de Munter, L., Kruithof, N., de Jongh, M., & Polinder, S. (2019). Medical and productivity costs after trauma. PLoS One, 14(12), e0227131. 10.1371/journal.pone.0227131 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Haering, S., Meyer, C., Schulze, L., Conrad, E., Blecker, M. K., El-Haj-Mohamad, R., … Engel, S. (2024a). Sex and gender differences in risk factors for posttraumatic stress disorder: A systematic review and meta-analysis of prospective studies. Journal of Psychopathology and Clinical Science, 133(6), 429t. 10.1037/abn0000918 [DOI] [PubMed] [Google Scholar]
- Haering, S., Seligowski, A. V., Linnstaedt, S. D., Michopoulos, V., House, S. L., Beaudoin, F. L., … Powers, A. (2024b). Sex-dependent differences in vulnerability to early risk factors for posttraumatic stress disorder: Results from the AURORA study. Psychological Medicine, 54(11), 2876–2886. 10.1017/S0033291724000941 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Haering, S., Seligowski, A. V., Linnstaedt, S. D., Michopoulos, V., House, S. L., Beaudoin, F. L., … Stevens, J. S. (2024c). Disentangling sex differences in PTSD risk factors. Nature Mental Health, 2(5), 605–615. 10.1038/s44220-024-00236-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hinrichs, R., van Rooij, S. J., Michopoulos, V., Schultebraucks, K., Winters, S., Maples-Keller, J., … Jovanovic, T. (2019). Increased skin conductance response in the immediate aftermath of trauma predicts PTSD risk. Chronic Stress, 3, 2470547019844441. 10.1177/2470547019844441 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hoeboer, C. M., Karaban, I., Karchoud, J. F., Olff, M., & van Zuiden, M. (2024). Validation of the PCL-5 in Dutch trauma-exposed adults. BMC Psychology, 12(1), 456. 10.1186/s40359-024-01951-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hoeboer, C. M., Nava, F., Haagen, J. F., Broekman, B. F., van der Gaag, R. J., & Olff, M. (2025). Epidemiology of DSM-5 PTSD and ICD-11 PTSD and complex PTSD in The Netherlands. Journal of Anxiety Disorders, 110, 102963. 10.1016/j.janxdis.2024.102963 [DOI] [PubMed] [Google Scholar]
- Hovens, J. E., Van der Ploeg, H. M., Klaarenbeek, M. T. A., Bramsen, I., Schreuder, J. N., & Rivero, V. V. (1994). The assessment of posttraumatic stress disorder: With the Clinician Administered PTSD Scale: Dutch results. Journal of Clinical Psychology, 50(3), 325–340. 10.1002/1097-4679(199405)50:3<325::AID-JCLP2270500304>3.0.CO;2-M [DOI] [PubMed] [Google Scholar]
- Huygens, M. W., Vermeulen, J., Swinkels, I. C., Friele, R. D., Van Schayck, O. C., & De Witte, L. P. (2016). Expectations and needs of patients with a chronic disease toward self-management and eHealth for self-management purposes. BMC Health Services Research, 16(1), 1–11. 10.1186/s12913-016-1484-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karchoud, J. F., Haagsma, J., Karaban, I., Hoeboer, C., van de Schoot, R., Olff, M., & van Zuiden, M. (2024a). Long-term PTSD prevalence and associated adverse psychological, functional, and economic outcomes: A 12–15 year follow-up of adults with suspected serious injury. European Journal of Psychotraumatology, 15(1), 2401285. 10.1080/20008066.2024.2401285 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karchoud, J. F., Hoeboer, C. M., Piwanski, G., Haagsma, J. A., Olff, M., van de Schoot, R., & van Zuiden, M. (2024b). Towards accurate screening and prevention for PTSD (2-ASAP): Protocol of a longitudinal prospective cohort study. BMC Psychiatry, 24(1), 688. 10.1186/s12888-024-06110-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karstoft, K. I., Tsamardinos, I., Eskelund, K., Andersen, S. B., & Nissen, L. R. (2020). Applicability of an automated model and parameter selection in the prediction of screening-level PTSD in Danish soldiers following deployment: Development study of transferable predictive models using automated machine learning. JMIR Medical Informatics, 8(7), e17119. 10.2196/17119 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kessler, R. C. (2000). Posttraumatic stress disorder: The burden to the individual and to society. Journal of Clinical Psychiatry, 14(4), 4–14. [PubMed] [Google Scholar]
- Kim, R., Lin, T., Pang, G., Liu, Y., Tungate, A. S., Hendry, P. L., … Linnstaedt, S. D. (2023). Derivation and validation of risk prediction for posttraumatic stress symptoms following trauma exposure. Psychological Medicine, 53(11), 4952–4961. 10.1017/S003329172200191X [DOI] [PubMed] [Google Scholar]
- Kuhn, M. (2008). Caret package. Journal of Statistical Software, 28(5), 1–26.27774042 [Google Scholar]
- Langeland, W., & Olff, M. (2024). Sex and gender in psychotrauma research. European Journal of Psychotraumatology, 15(1), 2358702. 10.1080/20008066.2024.2358702 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Larracy, R., Phinyomark, A., & Scheme, E. (2021, November). Machine learning model validation for early stage studies with small sample sizes. In 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) (pp. 2314–2319). IEEE. [DOI] [PubMed]
- López, V., Fernández, A., García, S., Palade, V., & Herrera, F. (2013). An insight into classification with imbalanced data: Empirical results and current trends on using data intrinsic characteristics. Information Sciences, 250, 113–141.
- Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing System, 30. [Google Scholar]
- Maercker, A. (2025). The paradox of the biopsychological and sociocultural levels in post-traumatic stress disorder. World Psychiatry, 24(1), 87. 10.1002/wps.21275 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Marmar, C. R., Metzler, T. J., & Otte, C. (2004). The peritraumatic dissociative experiences questionnaire. The Guilford Press. [Google Scholar]
- Moons, K. G., Wolff, R. F., Riley, R. D., Whiting, P. F., Westwood, M., Collins, G. S., … Mallett, S. (2019). PROBAST: A tool to assess risk of bias and applicability of prediction model studies: Explanation and elaboration. Annals of Internal Medicine, 170(1), W1–W33. 10.7326/M18-1377 [DOI] [PubMed] [Google Scholar]
- Mouthaan, J., Sijbrandij, M., Reitsma, J. B., Gersons, B. P., & Olff, M. (2014). Comparing screening instruments to predict posttraumatic stress disorder. PLoS One, 9(5), e97183. 10.1371/journal.pone.0097183 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mrazek, P. J. & Haggerty, R. J. (Eds.). (1994). Reducing risks for mental disorders: Frontiers for preventive intervention research. Commit-tee on Prevention of Mental Disorders, Division of Behavioral Sciences and Mental Disorders, Institute of Medicine, National Academies Press. [PubMed] [Google Scholar]
- Olff, M., Bakker, A., Frewen, P., Aakvaag, H., Ajdukovic, D., … Brewer, D. (2020). Screening for consequences of trauma – an update on the global collaboration on traumatic stress. European Journal of Psychotraumatology, 11(1), 1. 10.1080/20008198.2020.1752504 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Papini, S., Pisner, D., Shumake, J., Powers, M. B., Beevers, C. G., Rainey, E. E., … & Warren, A. M. (2018). Ensemble machine learning prediction of posttraumatic stress disorder screening status after emergency room hospitalization. Journal of anxiety disorders, 60, 35–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Papini, S., Norman, S. B., Campbell-Sills, L., Sun, X., He, F., Kessler, R. C., … Stein, M. B. (2023). Development and validation of a machine learning prediction model of posttraumatic stress disorder after military deployment. JAMA Network Open, 6(6), e2321273–e2321273. 10.1001/jamanetworkopen.2023.21273 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Prins, A., Bovin, M. J., Smolenski, D. J., Marx, B. P., Kimerling, R., Jenkins-Guarnieri, M. A., … & Tiet, Q. Q. (2016). The primary care PTSD screen for DSM-5 (PC-PTSD-5): Development and evaluation within a veteran primary care sample. Journal of General Internal Medicine, 31(10), 1206–1211. 10.1007/s11606-016-3703-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Santos, J. L., Harnett, N. G., van Rooij, S. J., Ely, T. D., Jovanovic, T., Lebois, L. A., … Stevens, J. S. (2025). Social buffering of posttraumatic stress disorder: Longitudinal effects and neural mediators. Biological Psychiatry: Cognitive Neuroscience and Neuroimaging, 10(5), 531–541. 10.1016/j.bpsc.2024.11.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schultebraucks, K., Shalev, A. Y., Michopoulos, V., Grudzen, C. R., Shin, S. M., Stevens, J. S., … Galatzer-Levy, I. R. (2020). A validated predictive algorithm of post-traumatic stress course following emergency department admission after a traumatic stressor. Nature Medicine, 26(7), 1084–1088. 10.1038/s41591-020-0951-z [DOI] [PubMed] [Google Scholar]
- Schultebraucks, K., Sijbrandij, M., Galatzer-Levy, I., Mouthaan, J., Olff, M., & van Zuiden, M. (2021). Forecasting individual risk for long-term posttraumatic stress disorder in emergency medical settings using biomedical data: A machine learning multicenter cohort study. Neurobiology of Stress, 14, 100297. 10.1016/j.ynstr.2021.100297 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sharp, C., Fonagy, P., & Allen, J. G. (2012). Posttraumatic stress disorder: A social-cognitive perspective. Clinical Psychology: Science and Practice, 19(3), 229. 10.1111/cpsp.12002 [DOI] [Google Scholar]
- Sheehan, D. V., Lecrubier, Y., Sheehan, K. H., Amorim, P., Janavs, J., Weiller, E., … Dunbar, G. C. (1998). The Mini-International Neuropsychiatric Interview (MINI): The development and validation of a structured diagnostic psychiatric interview for DSM-IV and ICD-10. Journal of Clinical Psychiatry, 59(Suppl 20), 22–33. [PubMed] [Google Scholar]
- Siontis, G. C., Tzoulaki, I., Castaldi, P. J., & Ioannidis, J. P. (2015). External validation of new risk prediction models is infrequent and reveals worse prognostic discrimination. Journal of Clinical Epidemiology, 68(1), 25–34. 10.1016/j.jclinepi.2014.09.007 [DOI] [PubMed] [Google Scholar]
- Spinhoven, P. H., Ormel, J., Sloekers, P. P. A., Kempen, G. I. J. M., Speckens, A. E., & van Hemert, A. M. (1997). A validation study of the Hospital Anxiety and Depression Scale (HADS) in different groups of Dutch subjects. Psychological Medicine, 27(2), 363–370. 10.1017/S0033291796004382 [DOI] [PubMed] [Google Scholar]
- Stekhoven, D. J., & Bühlmann, P. (2012). MissForest—non-parametric missing value imputation for mixed-type data. Bioinformatics (Oxford, England), 28(1), 112–118. 10.1093/bioinformatics/btr597 [DOI] [PubMed] [Google Scholar]
- Tang, F., & Ishwaran, H. (2017). Random forest missing data algorithms. Statistical Analysis and Data Mining: The ASA Data Science Journal, 10(6), 363–377. 10.1002/sam.11348 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Therneau, T., Atkinson, B., Ripley, B., & Ripley, M. B. (2015). Package ‘rpart’. Available online: cran. ma. ic. ac. uk/web/packages/rpart/rpart. pdf (accessed on 20 April 2016).
- Tolin, D. F., & Foa, E. B. (2008). Sex differences in trauma and posttraumatic stress disorder: A quantitative review of 25 years of research. [DOI] [PubMed]
- Tortella-Feliu, M., Fullana, M. A., Pérez-Vigil, A., Torres, X., Chamorro, J., Littarelli, S. A., … de la Cruz, L. F. (2019). Risk factors for posttraumatic stress disorder: An umbrella review of systematic reviews and meta-analyses. Neuroscience & Biobehavioral Reviews, 107, 154–165. 10.1016/j.neubiorev.2019.09.013 [DOI] [PubMed] [Google Scholar]
- Vali, M., Nezhad, H. M., Kovacs, L., & Gandomi, A. H. (2025). Machine learning algorithms for predicting PTSD: A systematic review and meta-analysis. BMC Medical Informatics and Decision Making, 25(1), 34. 10.1186/s12911-024-02754-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- van der Mei, W. F., Barbano, A. C., Ratanatharathorn, A., Bryant, R. A., Delahanty, D. L., deRoon Cassini, T. A., … Shalev, A. Y. (2020). Evaluating a screener to quantify PTSD risk using emergency care information: A proof of concept study. BMC Emergency Medicine, 20(1), 16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- van Sonderen, E. (1997). Sociale Steun Lijst–Interacties (SSL-I) en Sociale Steun Lijst-Discrepanties (SSL-D). Centrum voor Gezondheidsvraagstukken Rijksuniversiteit Groningen.
- Van Vliet, I. M., & De Beurs, E. (2007). The MINI-International Neuropsychiatric Interview. A brief structured diagnostic psychiatric interview for DSM-IV en ICD-10 psychiatric disorders. Tijdschrift Voor Psychiatrie, 49(6), 393–397. [PubMed] [Google Scholar]
- Wang, Y., Chung, M. C., Wang, N., Yu, X., & Kenardy, J. (2021). Social support and posttraumatic stress disorder: A meta-analysis of longitudinal studies. Clinical Psychology Review, 85, 101998. 10.1016/j.cpr.2021.101998 [DOI] [PubMed] [Google Scholar]
- Weathers, F. W., Davis, M. T., Witte, T. K., & Domino, J. L. The posttraumatic stress disorder checklist for DSM-5 (PCL-5): Development. [DOI] [PubMed]
- Weathers, F. W., Ruscio, A. M., & Keane, T. M. (1999). Psychometric properties of nine scoring rules for the clinician-administered posttraumatic stress disorder scale. Psychological Assessment, 11(2), 124. 10.1037/1040-3590.11.2.124 [DOI] [Google Scholar]
- Weiss, D. S. (2007). The impact of event scale: Revised. In J. P. Wilson & C. S. Tang (Eds.), Cross-cultural assessment of psychological trauma and PTSD (pp. 219–238). Springer US. [Google Scholar]
- Wiseman, M., Hinks, M., Hallett, W., Hinks, M., Hallett, D., Blundell, J., Sweeney, E., Thorpe, C. M., … Swift-Gallant, A. (2023). Evidence that ovarian hormones, but not diet and exercise, contribute to the sex disparity in post-traumatic stress disorder. Journal of Psychiatric Research, 168, 213–220. 10.1016/j.jpsychires.2023.10.048 [DOI] [PubMed] [Google Scholar]
- Zhang, Z., Zhu, X., & Liu, D. (2022). Model of Gradient Boosting Random Forest Prediction. In: 2022 IEEE International Conference on Networking, Sensing and Control (ICNSC). IEEE; 2022. pp. 1–6.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data of the study and the code to produce the results described in this paper are available at Open Science Framework (OSF; https://osf.io/59eax/). The 2-ASAP and TraumaTIPS cohorts are registered in the FAIR Traumatic Stress Data Sets library of the Global Collaboration on Traumatic Stress (GCTS).

