Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Dec 1.
Published in final edited form as: Nat Med. 2025 Mar 5;31(6):1840–1846. doi: 10.1038/s41591-025-03560-7

Prediction of Mental Health Risk in Adolescents

Elliot D Hill 1,2, Pratik Kashyap 3, Elizabeth Raffanello 3, Yun Wang 4, Terrie E Moffitt 5,6,7, Avshalom Caspi 5,6,7, Matthew Engelhard 1,2,*, Jonathan Posner 3,*
PMCID: PMC12176513  NIHMSID: NIHMS2076603  PMID: 40044931

Abstract

Prospective prediction of mental health risk in adolescence can facilitate early preventive interventions. Using psychosocial questionnaires and neuroimaging measures from over 11,000 children in the Adolescent Brain and Cognitive Development (ABCD) study, we trained neural network models to stratify general psychopathology risk. The model trained on current symptoms accurately predicted which participants would convert into the highest psychiatric illness risk group in the following year, with an Area Under the Receiver Operating Characteristic curve (AUROC) of 0.84. The model trained solely on potential etiologies or disease mechanisms achieved an AUROC of 0.75 without relying on the child’s current symptom burden. Sleep disturbances emerged as the most influential predictor of high-risk status, surpassing adverse childhood experiences and family mental health history. Including neuroimaging measures did not significantly enhance predictive performance. These findings suggest that artificial intelligence (AI) models trained on readily available psychosocial questionnaires can effectively predict future psychiatric risk while highlighting potential targets for intervention. This is a promising step toward AI-based mental health screening for clinical decision support systems.

INTRODUCTION

Since the onset of the COVID-19 pandemic, mental illness rates among youth have risen significantly in the USA [1] and globally [2], adding strain to already overburdened mental health systems [3]. A key challenge in enhancing the effectiveness of mental health services is identifying youth most vulnerable or at the highest risk for psychiatric illness. Accurately predicting which youth in the general population will develop psychiatric problems would enable efficient allocation of preventive resources. To address this challenge, we trained neural network models [4] on longitudinal psychosocial and neurobiological data to predict future mental health risk, thereby providing an efficient approach for predicting psychiatric illness risk over time and identifying key contributing factors.

Understanding the diverse psychosocial and neurobiological factors contributing to youth mental health problems remains challenging [5]. Mental health issues seldom stem from a single cause; multiple factors typically influence them, each contributing a small yet significant added risk. Moreover, conventional studies assessing psychiatric risk often rely on categorical frameworks like the Diagnostic and Statistical Manual (DSM), which does not adequately address the high rates of comorbidity across psychiatric disorders. It may be more beneficial to characterize risk factors as contributing to psychopathology broadly rather than to specific psychiatric disorders. Addressing these challenges could lead to better predictors of psychiatric risk.

Based on longitudinal cohort research across the lifespan, investigators have recently proposed an approach to characterizing psychiatric illness based on one underlying dimension, the General Factor of Psychopathology or “p-factor.” The p-factor reflects shared variation across psychiatric disorders, much as the g-factor reflects shared variance across intelligence domains. Studies on the social and neurobiological mechanisms underlying youth mental health issues have used the p-factor to capture general psychopathology. For example, a recent study of the Adolescent Brain and Cognitive Development (ABCD) cohort associated higher p-factor scores with smaller global brain volume and surface area on MRI [6], and a second ABCD analysis associated the p-factor with connectivity in the default mode and dorsal attention networks [7].

While predictive models of psychiatric risk have been developed, and some have shown good accuracy, they have relied upon symptom burden as predictors [810]. Findings from these models typically can be summarized as pointing to current symptom burden as the key predictor of future burden. While useful for screening purposes, this focus on symptoms rather than underlying etiologies has not led to novel prevention strategies. An illustrative counterexample is the Framingham Risk Score [11], which indexes cardiovascular risk based in part on presumed etiologies (e.g., elevated lipids) rather than the disease or its symptoms, enabling lipid-lowering interventions as a prevention strategy. For psychiatric risk assessment to have a similar impact – not just by identifying risk but guiding prevention – it must also incorporate presumed etiologies or disease mechanisms. Our analyses include two approaches: a “symptom-driven” model using current symptom burden to predict future symptoms and a “mechanism-driven” model using predictors based on potential etiologies of mental illness. The symptom-driven approach provides a benchmark, or upper limit, of predictive accuracy, allowing us to evaluate the performance of the mechanism-driven approach in identifying etiological predictors of psychiatric risk.

Our study examined psychosocial and neurobiological factors associated with mental illness using data from the ongoing ABCD study, which includes over 11,000 youth and multiple assessments of their psychosocial environment and brain development collected over five years. These data were used to train neural network models to predict from current and previous assessment data which youth would have higher p-factor scores (indicative of a higher level of general psychopathology) one year later. It is possible to use traditional modeling tools in this setting, for example, by extracting summary or factor scores for each measurement instrument and applying a mixed modeling approach [12]. However, this requires simplifying assumptions about how to a) summarize measures, b) represent measurement history, and c) model relationships between outcomes and latent factors. In contrast, neural networks make fewer assumptions about data distribution, allowing relationships between predictors and outcomes to be entirely learned from data. While this flexibility can be a liability in smaller datasets, it is an advantage here due to the ABCD study’s size and uncertainty about the validity of the above assumptions.

We aimed to achieve five goals. First, we aimed to accurately predict mental health risk using current and past measurements. Second, we assessed whether accuracy could be maintained with accessible measurements (e.g., self- or parent-report questionnaires), avoiding expensive, difficult-to-access tests. Third, we examined if predictions remained accurate when using measurements of mechanisms, etiologies, or protective factors rather than symptoms. Fourth, we focused on predicting “conversion,” identifying children whose p-factor rose from below to above the 75th percentile (i.e., the most at-risk group) one year later. Fifth, we examined which psychosocial and neurobiological factors most influenced model predictions.

RESULTS

Cohort descriptive statistics

After applying our filtering criteria, the dataset contained 11,416 participants. The training, validation, and test set contained 9,132, 1,142, and 1,142 participants. The mean number of events per participant was 3.1. The baseline and 1–3-year follow-up events had 11,409, 7,954, 7,667, and 4,585 participants, respectively. The mean, minimum, and maximum ages were 11.4 ± 1.3, 8.9, and 15.4 years, respectively. There were 5,485 (47.7%) female and 6,027 (52.4%) male participants. There were 242 (2.1%) Asian, 1,682 (14.6%) Black, 2,310 (20.1%) Hispanic, 1,209 (10.5%) Other, and 6,069 (52.7%) White participants. The number of participants in all risk groups decreased with age, follow-up event, and event year due to loss-to-follow-up (Table 1).

Table 1. Distribution of Participant Measurements Across Mental Health Risk Categories Stratified by Demographic and Socioeconomic Subgroups.

Risk categories—No-risk, Low-risk, Moderate-risk, and High-risk—correspond to quartiles of the p-factor, with the High-risk group representing the highest quartile of general psychopathology. Subgroups are stratified by demographic variables like ADI quartile (Area Deprivation Index, divided into quartiles from least to most deprived) and Age (in years). Each cell provides the count of participants and the percentage of the total within the specified subgroup and risk category. The Total column shows the overall count and percentage of participants within each subgroup across all risk categories.

Variable Group No risk Low risk Moderate risk High risk Total
ADI quartile 1 2585 (27%) 2477 (27%) 2372 (26%) 2001 (22%) 9435 (26%)
ADI quartile 2 2408 (25%) 2400 (26%) 2385 (26%) 2116 (24%) 9309 (25%)
ADI quartile 3 2268 (24%) 2219 (24%) 2278 (25%) 2266 (25%) 9031 (25%)
ADI quartile 4 2227 (23%) 2039 (22%) 2199 (24%) 2621 (29%) 9086 (25%)
Age 9 767 (8%) 807 (9%) 946 (10%) 910 (10%) 3430 (9%)
Age 10 2003 (21%) 2145 (23%) 2186 (24%) 2084 (23%) 8418 (23%)
Age 11 2558 (27%) 2553 (28%) 2562 (28%) 2529 (28%) 10202 (28%)
Age 12 2364 (25%) 2056 (23%) 2146 (23%) 2111 (23%) 8677 (24%)
Age 13 1432 (15%) 1287 (14%) 1147 (12%) 1119 (12%) 4985 (14%)
Age 14 364 (4%) 287 (3%) 247 (3%) 251 (3%) 1149 (3%)
Event year 2016 140 (1%) 161 (2%) 178 (2%) 175 (2%) 654 (2%)
Event year 2017 1536 (16%) 1703 (19%) 1683 (18%) 1688 (19%) 6610 (18%)
Event year 2018 2692 (28%) 2752 (30%) 3005 (33%) 2759 (31%) 11208 (30%)
Event year 2019 2813 (30%) 2583 (28%) 2542 (28%) 2559 (28%) 10497 (28%)
Event year 2020 2084 (22%) 1788 (20%) 1674 (18%) 1687 (19%) 7233 (20%)
Event year 2021 223 (2%) 148 (2%) 152 (2%) 133 (1%) 656 (2%)
Follow-up event Baseline 2629 (28%) 2883 (32%) 3009 (33%) 2985 (33%) 11506 (31%)
Follow-up event 1-year 2809 (30%) 2664 (29%) 2772 (30%) 2567 (29%) 10812 (29%)
Follow-up event 2-year 2688 (28%) 2424 (27%) 2412 (26%) 2437 (27%) 9961 (27%)
Follow-up event 3-year 1362 (14%) 1164 (13%) 1041 (11%) 1015 (11%) 4582 (12%)
Race Asian 328 (3%) 195 (2%) 174 (2%) 93 (1%) 790 (2%)
Race Black 1594 (17%) 1171 (13%) 1034 (11%) 1174 (13%) 4973 (13%)
Race Hispanic 2005 (21%) 1758 (19%) 1794 (19%) 1763 (20%) 7320 (20%)
Race White 4766 (50%) 5133 (56%) 5219 (57%) 4823 (54%) 19941 (54%)
Race Other 795 (8%) 878 (10%) 1013 (11%) 1151 (13%) 3837 (10%)
Sex Female 4643 (49%) 4472 (49%) 4420 (48%) 3990 (44%) 17525 (48%)
Sex Male 4845 (51%) 4663 (51%) 4814 (52%) 5014 (56%) 19336 (52%)

Across the baseline and 1–3-year follow-up events, the conversion rate to high-risk was 11%—i.e., about 1 of 10 participants in the sample moved from a no-, low-, or medium-risk group into the high-risk group at the subsequent event. The rate of high-risk persistence was 66%—i.e., about 2 of 3 participants in the high-risk group at a given event remained in the high-risk group at the next event.

Hyperparameter tuning

Models incorporating temporal dependencies between measurements (RNN, LSTM, and Transformer) tended to have lower validation loss than simpler models (LM and MLP) that assume independent measurements. The LSTM achieved the lowest validation loss but was comparable to the RNN. For the questionnaires model, the hyperparameters that gave the lowest validation loss of 1.21 were: architecture = LSTM, hidden dimension = 191, number of layers = 1, dropout = 0.182, number of epochs = 56, learning rate = 0.003, momentum = 0.9, and weight decay = 1.61e-6 (Supplementary Table 1).

Model performance

For the mechanism-driven approaches, the models trained on the data-driven MRI predictor set did not converge, even when applying substantial dimensionality reduction and regularization (Supplementary Table 2). As a result, we focused on the smaller, theory-driven set of MRI features. However, even these did not improve model performance when added to the questionnaire predictor set (Supplementary Figure 1). Given these findings, we prioritized the questionnaire model (mechanism-driven approach), which provides strong performance and scalability.

Among the symptom-driven approaches, models using CBCL scales performed well, maintaining high predictive accuracy while requiring fewer questions, reducing the burden on participants. To streamline the analysis, we present results from two primary models:

  1. A mechanism-driven model using only questionnaire data.

  2. A symptom-driven model using CBCL scales.

Evaluation metrics for additional approaches are available in Supplementary Table 3.

The symptom-driven model CBCL scales model outperformed the mechanism-driven questionnaires model across all three high-risk scenarios (Figure 1). The mechanism-driven model trained on questionnaires had an AUROC for conversion, persistence, and agnostic (any prior risk group) of 0.75 ± 0.01, 0.69 ± 0.02, and 0.8 ± 0.01, respectively. The symptom-driven model trained on CBCL scales had an AUROC for conversion, persistence, and agnostic of 0.84 ± 0.01, 0.78 ± 0.01, and 0.89 ± 0.01, respectively. The models performed better when predicting the extreme compared to intermediate-risk groups; predicting which youth would be in the no-risk and high-risk groups was easier than in the low-risk and moderate-risk groups (Table 2).

Figure 1. Model Performance Stratified by Metric and High-Risk Groups.

Figure 1.

Receiver Operating Characteristic (ROC) and Precision-Recall (PR) curves comparing predictive model performance for mental health risk assessment for the test set (n = 1,142 participants). Each panel illustrates the performance of symptom-driven (CBCL scales) and mechanism-driven (Questionnaire) models:

(a) Conversion ROC curve: prediction of participants transitioning from no-, low-, or medium-risk groups to the high-risk group.

(b) Persistence ROC curve: prediction of participants that remain in the high-risk group.

(c) Agnostic ROC curve: prediction of participants entering the high-risk group from any prior group.

(d) Conversion PR curve: positive predictive value for participants converting to the high-risk group.

(e) Persistence PR curve: positive predictive value for participants persisting in the high-risk group.

(f) Agnostic PR curve: positive predictive value for predicting high-risk status regardless of prior group.

The curves highlight that while symptom-driven models (CBCL scales) outperform mechanism-driven models (Questionnaire), the sensitivity and precision of mechanism-driven models are strong.

Table 2. Predictive Performance Metrics for Mental Health Risk Stratified by Risk Scenarios and Groups.

The table compares model performance for high-risk prediction across three scenarios—Conversion, Persistence, and Agnostic—and stratifies results by risk groups (No-risk, Low-risk, Moderate-risk, and High-risk). Metrics reported include the Area Under the Receiver Operating Characteristic Curve (AUROC) and Average Precision (AP), denoted by mean ± 2 standard errors. For the AP metric, additional values in parentheses denote the prevalence of the high-risk scenario. The Predictor sets include the Questionnaires approach, reflecting the use of mechanism-driven predictors. Higher AUROC and AP values indicate better model performance in distinguishing participants likely to experience changes in their mental health risk status. (See Supplementary Glossary for additional explanation of terms).

Predictor set Metric Scenario No risk Low risk Moderate risk High risk
Questionnaires AUROC Conversion 0.73 ± 0.01 0.56 ± 0.01 0.63 ± 0.01 0.75 ± 0.01
Questionnaires AUROC Persistence 0.76 ± 0.04 0.69 ± 0.03 0.55 ± 0.02 0.69 ± 0.02
Questionnaires AUROC Agnostic 0.77 ± 0.01 0.63 ± 0.01 0.61 ± 0.01 0.8 ± 0.01
Questionnaires AP Conversion 0.56 ± 0.02 (0.34) 0.33 ± 0.01 (0.29) 0.35 ± 0.02 (0.27) 0.36 ± 0.02 (0.11)
Questionnaires AP Persistence 0.24 ± 0.06 (0.03) 0.15 ± 0.03 (0.05) 0.28 ± 0.02 (0.26) 0.76 ± 0.02 (0.66)
Questionnaires AP Agnostic 0.53 ± 0.02 (0.26) 0.3 ± 0.01 (0.22) 0.32 ± 0.01 (0.26) 0.59 ± 0.02 (0.26)
CBCL scales AUROC Conversion 0.82 ± 0.01 0.64 ± 0.01 0.72 ± 0.01 0.84 ± 0.01
CBCL scales AUROC Persistence 0.82 ± 0.03 0.79 ± 0.03 0.71 ± 0.02 0.79 ± 0.01
CBCL scales AUROC Agnostic 0.87 ± 0.01 0.72 ± 0.01 0.72 ± 0.01 0.88 ± 0.01
CBCL scales AP Conversion 0.68 ± 0.02 (0.34) 0.38 ± 0.02 (0.29) 0.42 ± 0.02 (0.27) 0.52 ± 0.03 (0.11)
CBCL scales AP Persistence 0.47 ± 0.07 (0.03) 0.33 ± 0.06 (0.05) 0.43 ± 0.03 (0.26) 0.84 ± 0.02 (0.66)
CBCL scales AP Agnostic 0.67 ± 0.02 (0.26) 0.37 ± 0.02 (0.22) 0.42 ± 0.01 (0.26) 0.75 ± 0.01 (0.26)

The mechanism-driven model’s predictive performance did not vary substantially across demographic subgroups (Table 3). Conversely, the symptom-driven performance increased with age, event year, and follow-up event but did not vary considerably across race, sex, or ADI quartile (Supplementary Table 3).

Table 3. Test Set Model Evaluation Metrics Stratified by Demographic Subgroups.

Predictive performance metrics stratified by demographic subgroups for high-risk mental health prediction. The table reports Area Under the Receiver Operating Characteristic Curve (AUROC) values for each subgroup across risk categories (No-risk, Low-risk, Moderate-risk, and High-risk). Subgroups include Sex (Female, Male) and Race/Ethnicity (Asian, Black, Hispanic, and others). AUROC values are presented as mean ± 2 standard errors, reflecting the model’s ability to distinguish participants within each demographic and risk category. Higher AUROC values indicate better model performance. The table demonstrates minimal variation in predictive performance across demographic subgroups, underscoring the model’s robustness and potential fairness across diverse populations. See Supplementary Glossary for additional explanation of terms.

Variable Group No risk Low risk Moderate risk High risk
Sex Female 0.77 ± 0.01 0.6 ± 0.02 0.62 ± 0.01 0.79 ± 0.01
Sex Male 0.78 ± 0.01 0.65 ± 0.01 0.6 ± 0.01 0.8 ± 0.01
Race Asian 0.7 ± 0.06 0.64 ± 0.07 0.65 ± 0.09 0.75 ± 0.09
Race Black 0.76 ± 0.02 0.61 ± 0.03 0.61 ± 0.03 0.77 ± 0.02
Race Hispanic 0.77 ± 0.02 0.59 ± 0.03 0.62 ± 0.02 0.79 ± 0.02
Race White 0.78 ± 0.01 0.64 ± 0.01 0.6 ± 0.01 0.8 ± 0.01
Race Other 0.77 ± 0.03 0.66 ± 0.03 0.58 ± 0.03 0.84 ± 0.02
Age 9 0.77 ± 0.03 0.57 ± 0.04 0.59 ± 0.03 0.79 ± 0.03
Age 10 0.77 ± 0.02 0.63 ± 0.02 0.6 ± 0.02 0.8 ± 0.02
Age 11 0.73 ± 0.02 0.61 ± 0.02 0.59 ± 0.02 0.79 ± 0.02
Age 12 0.79 ± 0.02 0.63 ± 0.02 0.62 ± 0.02 0.8 ± 0.02
Age 13 0.82 ± 0.02 0.68 ± 0.03 0.6 ± 0.03 0.82 ± 0.02
Age 14 0.85 ± 0.04 0.62 ± 0.06 0.71 ± 0.04 0.85 ± 0.04
Follow-up event Baseline 0.77 ± 0.02 0.63 ± 0.02 0.59 ± 0.02 0.8 ± 0.01
Follow-up event 1-year 0.76 ± 0.02 0.62 ± 0.02 0.59 ± 0.02 0.79 ± 0.01
Follow-up event 2-year 0.78 ± 0.02 0.62 ± 0.02 0.62 ± 0.02 0.8 ± 0.02
Follow-up event 3-year 0.8 ± 0.03 0.67 ± 0.03 0.65 ± 0.03 0.81 ± 0.02
ADI quartile 1 0.77 ± 0.02 0.62 ± 0.02 0.61 ± 0.02 0.81 ± 0.02
ADI quartile 2 0.78 ± 0.02 0.66 ± 0.02 0.61 ± 0.02 0.81 ± 0.02
ADI quartile 3 0.78 ± 0.02 0.61 ± 0.02 0.62 ± 0.02 0.77 ± 0.02
ADI quartile 4 0.76 ± 0.02 0.6 ± 0.02 0.58 ± 0.02 0.8 ± 0.02
Event year 2016 0.8 ± 0.08 0.64 ± 0.08 0.54 ± 0.07 0.77 ± 0.07
Event year 2017 0.77 ± 0.02 0.63 ± 0.03 0.62 ± 0.02 0.79 ± 0.02
Event year 2018 0.77 ± 0.02 0.61 ± 0.02 0.58 ± 0.02 0.81 ± 0.01
Event year 2019 0.77 ± 0.02 0.62 ± 0.02 0.61 ± 0.02 0.79 ± 0.02
Event year 2020 0.78 ± 0.02 0.66 ± 0.02 0.63 ± 0.02 0.82 ± 0.02
Event year 2021 0.76 ± 0.06 0.58 ± 0.11 0.61 ± 0.08 0.75 ± 0.07

SHAP values

The SHAP value analysis using the mechanism-driven (questionnaires-only) model indicated that sleep disturbances, prosocial behaviors, ACEs, family mental health history, and family conflict had the largest impact on high-risk predictions. Notably, sleep disturbances had a disproportionately large impact, and parent questionnaires influenced model predictions more strongly than youth questionnaires (Figure 2).

Figure 2. Predictor Category Shapley Additive Explanations (SHAP) Analysis.

Figure 2.

SHAP values represent the influence of each predictor category on the model’s predicted probability of a participant being classified as high-risk. Predictor categories are ranked by their absolute SHAP values sums, with higher values indicating a larger impact on model predictions. Absolute SHAP value sums are displayed as mean ± 95% confidence interval (n = 1,142 participants). The Sleep Disturbances category emerged as the most influential predictor, followed by Prosocial Behaviors, Adverse Childhood Experiences (ACEs), Family Mental Health History, and Family Conflict. Each predictor category is further stratified by respondent type (youth vs. parent), with parent-reported data generally showing more substantial influence on model predictions than youth-reported data. The questionnaire predictor responses can be found in Supplementary Table 4.

For most participants, more sleep disturbances increased predicted high-risk probability for the following year, although extreme levels of disturbance reduced predicted risk in a small subset. More adverse childhood experiences, family conflict, and family history of mental illness increased high-risk probability, whereas more prosocial behaviors and parental monitoring decreased it. Although older males had a slightly higher high-risk probability than younger females, the influence of sex and age was negligible compared to the other factors (Supplementary Figure 2). A list of the SHAP value sums for each predictor in the questionnaire model is presented in Supplementary Table 4.

Model generalization

The questionnaire model trained with the propensity score weighted loss function performed similarly to the unweighted loss function when predicting high-risk conversion, persistence, and agnostic groups (Supplementary Figure 3). The magnitude and relative ordering of the absolute SHAP value sums remained the same after reweighting.

The questionnaire model developed using site-based training, validation, and test splits maintained the same predictive performance as participant-based data splits when predicting high-risk conversion, persistence, and agnostic groups (Supplementary Figure 3).

Sensitivity analysis

The AUROC for the within-event and across-event p-factor models were comparable for the mechanism-driven questionnaires and symptom-driven CBCL scales models (Supplementary Figure 4).

DISCUSSION

Our study used artificial intelligence to prospectively predict mental health risk in youth. To our knowledge, this is the largest AI-based longitudinal study of generalized adolescent mental health risk prediction that assesses the relative importance of diverse predictors in making these prognostications.

Several findings stand out. First, our model predicted adolescent mental health (via the p-factor) one year after an assessment with relatively high accuracy. It effectively predicted high-risk conversion – predicting which youth would move from lower-risk groups to the highest-risk group. Given the low prevalence of high-risk conversion (11%), the model performed well on this subset of participants (Figures 1a and 1d). This is a promising step toward developing a clinical tool to preemptively identify at-risk patients who may require further evaluation to improve mental health outcomes.

Second, our model performed well with both the symptom-driven and mechanism-driven approaches, with the symptom-driven approach yielding slightly better results. This aligns with prior research showing that current symptom burden is often a strong predictor of future burden [10]. For example, the Chicago Adolescent Depression Risk Assessment achieves an AUROC of 0.8 for predicting future depressive episodes using current mood, anxiety, and affect regulation [8]. Relatedly, models predicting progression from at-risk to full-threshold schizophrenia [13] often rely on current symptoms and family history. While symptom-driven approaches may be useful for screening, they are susceptible to common methods bias, where predictors and outcomes come from the same source, potentially inflating associations.

In contrast, our mechanism-driven approach maintained strong predictive performance without relying on the child’s current symptom load. Its predictions were within a general, non-referred sample and included potential etiologies of illness, making it distinct from symptom-based approaches. By focusing on underlying mechanisms rather than symptom patterns, the mechanism-driven model offers the potential to identify actionable prevention targets.

Regarding potential prevention targets, sleep disturbances stood out as a robust predictor of high psychiatric illness risk. These were assessed with the 26-item Sleep Disturbance Scale for Children [14] (SDSC; see Supplementary Table 5 for specific items and scoring). Their predictive influence surpassed variables like adverse childhood experiences (ACEs), family mental health history, and socioeconomic status. Approximately 29% of the ABCD cohort had sleep disturbance scores exceeding clinical thresholds [14], indicating that pathological sleep disturbance was common but not universal. Notably, the relationship between sleep disturbances and future mental illness risk was non-linear; increased sleep disturbances raised risk up to a point, but extreme—and far less common—levels of sleep disturbances did not (Supplementary Figure 2). This could suggest that while moderate sleep disturbances portend mental illness, severe sleep problems may instead relate to non-psychiatric conditions. Importantly, our model’s performance remained stable during the pandemic, a period of potentially heightened sleep disruption (see Table 3). The stability of the model’s performance during the COVID-19 pandemic supports its generalizability under varying real-world conditions.

The importance of sleep disturbance in predicting mental illness is consistent with prior research. A meta-analysis [15] found sleep disturbances to be a precursor for several psychiatric disorders, including depression, anxiety, and bipolar disorder. Other studies have associated sleep disturbances with subsequent substance use disorders, suicide attempts, and suicide completion [16]. Sleep is vital for healthy neurodevelopment [17], and sleep disturbances, particularly in adolescence, are implicated as contributors to mental illness across diagnoses [1820]. Although further research is needed, our finding that sleep disturbances were an influential predictor of high psychiatric illness risk is potentially auspicious, as sleep disturbances are modifiable with evidence-based behavioral interventions [21]. Indeed, a recent clinical trial found that a sleep intervention effectively reduced symptoms across diagnoses [22].

Third, using the p-factor as an outcome supports the model’s broad clinical applicability as a transdiagnostic indicator of mental illness. The model’s predictions are not limited to a single diagnostic group. Additionally, the low variation in the model’s predictive performance across demographic groups (Table 3) suggests resilience to demographic biases. The symptom-driven model’s improved performance with participant age and event year suggests that the LSTM effectively utilized prior information, highlighting the benefit of longitudinal assessments for mental health prediction. In addition to its broad applicability, the p-factor also serves as a robust harbinger of future adverse psychiatric incidents. For example, a registry study of over 32,000 twins prospectively associated a general psychopathology factor (i.e., the p-factor) with increased suicidal behavior, substance overdoses, criminality, and new prescriptions for most psychiatric medications [23]. While our approach derived the p-factor from the multi-item CBCL, novel approaches, including computerized adaptive testing, aim to enable faster p-factor assessments [24,25].

Fourth, adding MRI-derived neurobiological measures did not significantly improve model performance beyond that of the questionnaires or CBCL models. This finding was consistent across both the theory-driven and data-driven approaches we employed and aligns with research showing marginal associations between brain measures and mental health [10,26]. This result underscores the model’s potential clinical utility as its predictors can rely on accessible, low-cost instruments scalable to most clinical settings. An alternative explanation for the lower performance of MRI predictors could be that they had half as many measurements as the psychosocial questionnaires because MRI scans were only taken every other year in the ABCD study. In contrast, the psychosocial predictors were measured every year. We used tabular MRI data for our experiments, primarily to compare them to the more scalable questionnaires; however, training computer vision models on ABCD’s raw imaging data may offer a promising avenue for future research.

Fifth, our findings indicate that the model’s performance was dominated by parent-reported data, with minimal contribution from youth-reported data (Figure 2). With their broader life experience, parents may better contextualize behaviors and detect deviations from developmental norms. In contrast, youth may lack external reference points required to recognize atypical experiences. Future work could explore ways to integrate youth perspectives into predictive models better.

Applying propensity weights to account for sampling biases in the ABCD study showed no change in the model’s performance metrics and SHAP values, supporting its reliability when applied to a population more representative of the U.S. Additionally, we conducted a site-specific analysis by excluding data from two ABCD sites during model training, using them exclusively for testing. This “spatial distribution shift” evaluation yielded consistent results, suggesting the model’s resilience to differences in assessment procedures, demographics, and other site-specific factors. Together, these analyses provide strong evidence for the model’s generalizability across diverse conditions.

While our findings are promising, additional steps are needed to enhance the model’s readiness for clinical deployment. Prospective evaluation in independent samples – such as real-world clinical settings where clinicians assess model predictions without influencing clinical decisions – will be critical [27]. Another priority is identifying a minimal set of questions that maintains high predictive performance while reducing patient burden. Sparsity-inducing loss functions (e.g., L1 penalization [28]) could aid in optimizing this balance. Further performance improvements may be achieved using accessible, low-cost data sources like medical records or behavioral assays. Lastly, while the ABCD study sampled a diverse range of demographics across the US, it is still possible that clinical populations are systematically different from the general US population. Thus, integrating methods that account for distribution shifts [29] into the modeling pipeline will be essential for improving clinical applicability. Future clinical applications could also include measures of adaptive functioning or functional impairment to augment mental health assessments.

This study demonstrates that AI models can accurately predict adolescent mental health prospectively using readily available questionnaires. Social environment and behavioral measures (particularly sleep disturbances) strongly influenced model-predicted psychiatric illness risk, whereas the impact of neurobiological measures on predicted risk was limited. Future research should validate these findings in clinical populations and assess whether AI-based early detection systems can help allocate mental health resources to patients who require preemptive intervention to improve their long-term outcomes. If validated, this approach could guide targeted preventive efforts, enhance early detection, and ultimately contribute to more efficient mental health care delivery for adolescents.

METHODS

Ethics Statement

Institutional review boards at each ABCD study site approved the study procedures. Written consent was obtained from all parents and verbal assent was given by all children. The Duke Health Institutional Review Board approved analysis of ABDC data related to this manuscript.

Glossary of Terms

For definitions of technical terms, refer to the Glossary provided in the Supplementary Information.

Description of Cohort

The ABCD study collected data from a community sample of 11,880 children from 21 research sites across the United States. Recruitment for the ABCD study used a multi-stage probability sampling method to capture the sociodemographic diversity of the U.S. population (for additional details, see [30]). Post-sampling propensity weights were provided to improve the sample’s representativeness, and we incorporated these weights in our generalization analyses [31]. All youth had in-person assessments once a year. Self-report and social assessments were taken yearly, whereas brain MRI scans were collected every two years. The baseline ABCD assessment was taken between 2017 and 2018 at ages 9–10, and four subsequent waves of assessment (1-year follow-up through 4-year follow-up; hereafter referred to as “events”) were included in our analysis. ABCD questionnaires were completed by either the youth or a parent. Our model predicted outcomes at each follow-up event based on predictors measured at all preceding events.

Predictors: Questionnaires

To assess participants’ social environment and behaviors, we included predictors from the following ABCD questionnaires in our analysis: Child Behavior Checklist (CBCL), family conflict, neighborhood safety, problem monitoring, prosocial behaviors, school risk and protective factors, screentime, sleep disturbances, parental rules, and family mental health history (see Supplementary Table 5 for predictors metadata including, the ABCD file, predictor names, questions, and responses). In addition, we also included measures of demographics (age and sex assigned at birth), adverse childhood experiences (ACEs), and socioeconomic status. We included age and sex even though our outcome (p-factor, see below) is adjusted for them because nonlinear interactions between these demographic variables and other predictors may not be accounted for in a linear adjustment. Because the ABCD study did not include a questionnaire specific to ACEs (e.g., CDC-Kaiser ACEs Questionnaire [32]), we derived an ACEs score for each participant based on other questionnaires (see Supplementary Table 6 for details). For socioeconomic status, we used the area deprivation index [33], the highest level of parental education, and annual household income. Lastly, we included the time elapsed since the baseline measurement (in years) as a predictor.

Predictors: MRI measures

We used two approaches for selecting MRI predictors: theory-driven and data-driven approaches.

In the theory-driven approach, we limited the MRI predictors to measures commonly associated with psychopathology [34], including transdiagnostic measures in youth [35]. These predictors included (i) functional connectivity estimates between the default mode, fronto-parietal, and cingulo-opercular networks (based on [36]); (ii). DTI-derived average fractional anisotropy within 21 white matter tracts (see Supplementary Table 5 for specific tracts); and (iii). mean activation (beta weights) from predefined regions of interest during the stop signal task [37], monetary incentive task [38], and emotional N-bask task [39]. This approach prioritized computational efficiency and reduced model complexity while preserving the study’s primary objectives.

To complement this targeted strategy, we conducted a data-driven analysis incorporating all available MRI measures from the ABCD study, encompassing 55,207 predictors. This broader analysis allowed us to explore the potential utility of less commonly studied MRI features, ensuring no potentially relevant information was excluded.

Outcome: P-factor

We performed factor analysis on the eight Childhood Behavior Checklist (CBCL) syndrome scales measured at all follow-up events to construct our outcome variable, the p-factor. The syndrome scales included anxious/depressed, withdrawn/depressed, somatic complaints, social problems, thought problems, attention problems, rule-breaking behavior, and aggressive behavior. A previous analysis of p-factor in ABCD demonstrated that factor loadings and scores are robust to the choice of factor analysis model and whether individual items versus CBCL subscales were used [40]. Therefore, we used a single-factor exploratory factor analysis model to reduce the dimensionality of the CBCL subscales to a single factor—the p-factor—after which we binned the p-factor scores into quartiles. The model’s task was to predict which quartile a participant’s p-factor would be in at the next follow-up event. For example, we tested whether predictors (e.g., questionnaires and MRI measures) at the baseline could predict a participant’s p-factor quartile measured at the 1-year follow-up event. We term quartiles 1–4 as “no-risk,” “low-risk,” “medium-risk,” and “high-risk,” respectively. We defined “conversion” as moving from the no-risk, low-risk, or medium-risk groups into the high-risk group at the next follow-up event, “persistence” as remaining in the high-risk group at the next follow-up event, and “agnostic” as being in the high-risk group at the next follow-up event irrespective of group in the prior event.

Data preprocessing

To make a p-factor prediction, a participant needed at least one p-factor measurement available from an event after the predictors were measured. Therefore, we did not include outcomes (i.e., p-factors) from the baseline event because no preceding data were available for this prediction. Similarly, we did not include predictors (questionnaires and MRI measures) from the 4-year follow-up event because there was no subsequent event in which the p-factor outcome could be assessed. Lastly, we excluded participants who had CBCL scales measured only at baseline.

From the questionnaires, we removed items that asked about the participant’s preferred language, the number of missing/answered questions, and summary scores (i.e., measures aggregated from multiple questions). In addition, we excluded predictors with more than 25% of their values missing. We standardized our predictors to z-scores (0 mean and unit variance). To impute missing values, we forward-filled predictors within participants; for predictors without values at preceding events, we imputed the mean value of the predictor across all participants for that event. Lastly, we split our data into independent training, validation, and test sets with proportions 8:1:1. The training set was used to learn the neural network parameters that best fit the data. The validation set was used to select optimal hyperparameters, and the test set was used to estimate out-of-sample model performance. To prevent data leakage (i.e., when the model erroneously learns information from the test set during training), all learnable preprocessing steps (i.e., z-score scaling and mean imputation) were fit only on the training set and then applied separately to the validation and test sets.

Model architecture

We trained neural network models to predict a participant’s p-factor quartile at a measurement event from the predictors measured at all prior events. We tested five neural network model architectures in increasing order of complexity: Linear Model (LM), Multilayer Perceptron (MLP) [41], Recurrent Neural Network (RNN) [42], Long Short-Term Memory (LSTM) [43], and Transformer [44]. The LM and MLP assume that future psychiatric risk (as measured by the p-factor) is conditionally independent of previous assessments, given information collected at the current assessment. In contrast, the RNN, LSTM, and Transformer learn temporal dependencies between p-factor predictions and the full history of past measurements. Importantly, we used a unidirectional RNN and LSTM to prevent the model from erroneously learning from measurements not yet available at the current prediction time. For the same reason, we used a causal attention mask for the Transformer.

We chose these sequence neural network models because they naturally handle repeated measurements and can efficiently learn high-dimensional, nonlinear relationships between predictors and outcomes. For the data-driven MRI predictor set, in addition to the architectures above, we used an additional autoencoder with linear encoder and decoder layers to reduce the dimensionality of the model input from 55,207 to 256 latent dimensions. Then, the learned latent encodings were used as the input for the five architectures listed above.

Training and hyperparameter tuning

The output dimension of all models was 4 (the number of predicted risk groups). We used multiclass cross-entropy (negative log-likelihood) as our loss function. We optimized the model’s parameters using stochastic gradient descent (SGD) [45] with Nesterov momentum [46]. We tuned the model architecture and optimization hyperparameters using 128 tuning trials [47]. The optimizer hyperparameter ranges for the number of epochs, initial learning rate, momentum, and weight decay [48] were [10, 100], [1e-4, 1e-2], [0.9, 0.99], and [1e-8, 1e-2], respectively. The model hyperparameter ranges for the dropout rate [49], hidden dimension size, and number of layers were [0.0, 0.75], [64, 256], and [1, 5]. For the Transformer, we set the number of attention heads to 4. The five model architectures were treated as hyperparameters during tuning. Finally, the model hyperparameters that achieved the lowest cross-entropy loss on the validation set were selected for evaluation on the test set. We used a learning rate scheduler to reduce the learning rate by a factor of 0.1 when the validation loss did not decrease after two epochs.

Model performance metrics

All model performance metrics were generated using the test set to estimate out-of-sample model performance (i.e., the model’s performance on data not used to train it). To evaluate the model’s predictive performance across all decision thresholds, we calculated the receiver operating characteristic (ROC) curve (true positive rate vs. false positive rate) and precision-recall (PR) curve (positive predictive value vs. true positive rate) for each quartile. To summarize the model’s performance, we calculated the area under the ROC curve (AUROC) and the average precision (AP), i.e., the area under the PR curve using the model’s predicted and observed outcomes for each quartile. All confidence intervals were generated by bootstrap resampling [50] with 1000 resamples of the test set predictions.

Scenarios for model performance

To assess the potential clinical utility of the model, we reasoned that the most clinically useful feature of this predictive model would be its ability to predict conversion (i.e., when a child’s mental health changes from low-risk to high-risk). A high likelihood of conversion should affect treatment planning. Conversely, if a child already has severe psychiatric illness and the model predicts that this child will remain in the high-risk group in the following year, such a prediction has less clinical utility because the treatment plan is unlikely to change.

Toward this goal, we evaluated the model under three scenarios. The first scenario was conversion from no-, low-, or moderate-risk to high-risk. Evaluating high-risk predictions for this scenario assesses the model’s ability to identify which participants will convert from a healthier to a more severe psychiatric state. The second scenario was persistence (remaining in the high-risk group from one year to the next). The third scenario was agnostic, or any prior risk group – that is, evaluating the model’s ability to predict being in the high-risk group, regardless of previous group status. Lastly, to measure how the model performance varied across demographic subgroups, we stratified the model performance metrics by sex, race/ethnicity, ADI quartile, age, follow-up event, and event year.

We also examined model performance across six different approaches based on the predictor subsets used in each approach. The subsets included (i) questionnaires, (ii) questionnaires and MRI measures, (iii) previous p-factors, (iv) CBCL scales, (v) questionnaires and CBCL scales, and (vi) questionnaires, MRI measures, and CBCL scales. For simplicity, we refer to approaches (i) and (ii) as “mechanism-driven” and approaches (iii–vi) as “symptom-driven.” Mechanism-driven approaches use distinct predictors and outcomes. In contrast, symptom-driven approaches involve predictors that overlap with the outcome measures, though they are measured at earlier time points (i.e., autoregressive predictors). For example, the p-factor (our outcome) is derived from CBCL scales; therefore, any approach incorporating CBCL scales as predictors is considered symptom-driven.

Symptom-driven models generally provide high predictive accuracy but are less likely to uncover actionable insights for interventions. For instance, while current moderate illness might predict later severe illness, such a prediction does not reveal underlying causes or suggest treatments. Conversely, mechanism-driven approaches focus on predictors from domains such as family conflict and neighborhood safety—factors thought to influence or protect against mental illness. These domains are distinct from those that define the p-factor. Analogous to the Framingham Risk Score in cardiovascular health, mechanism-driven approaches emphasize identifying modifiable mechanisms or etiologies contributing to mental illness, making them better suited for guiding targeted interventions. As a baseline for symptom-driven models, we evaluated the naive case of simply using the previous p-factor quartile to predict the following year’s p-factor quartile.

Predictor importance

We used Shapley additive explanations (SHAP) [51] to evaluate predictor importance. SHAP values estimate the influence of a participant’s predictor value on that participant’s predicted risk probability. To derive a single measure of importance for each predictor, we summed the SHAP values across each participant and event. To determine which questionnaire had the largest relative influence on model predictions, we aggregated the SHAP values by taking the absolute value of the sum of all predictors within a given questionnaire (e.g., SHAP values from each question in the family conflict questionnaire were summed). SHAP values must be estimated from a single model output; thus, we generated them using the predicted high-risk probability as we deemed it the most clinically relevant model output. We estimated SHAP values using expected gradients [52]. We generated 95% confidence intervals from 1000 bootstrap resamples.

Model generalization

To explore the generalizability of our model, we employed two complementary analyses: (i) Propensity Score Weighting and (ii) Site Analysis. First, to address sampling biases and enhance the representativeness of our sample, we weighted the cross-entropy loss function using propensity scores provided by the ABCD study [53]. These weights adjust for sampling biases, approximating a population more representative of the general U.S. population. By re-running our model with these adjusted weights, we tested the robustness of our findings under conditions that better reflect broader US population characteristics.

Second, we performed a site-level analysis to evaluate the model’s robustness to potential variability in assessment procedures and participant demographics across sites. Specifically, we excluded data from two ABCD sites during the model training and testing phases. These held-out sites were used solely for validation, enabling us to test model performance under “spatial” shifts—situations where data characteristics, such as questionnaire administration methods, order of administration, or site-specific demographic distributions, differ. This approach allowed us to assess how well the model performs when applied to previously unseen contexts.

Sensitivity analysis

We tested if estimating p-factors across versus within follow-up events affected model performance. When estimating factors across follow-up events, we assume the factor loadings do not change over time. In contrast, when we estimate factors within follow-up events, we assume that the corresponding factor loadings at different time points are statistically independent. These variations represent the two extremes of handling time in factor analysis, with more sophisticated modeling techniques likely falling somewhere in between.

Summary of methods

We trained neural network models on longitudinal psychosocial and MRI data from the ABCD cohort, which included 11,880 children aged 9–10 at baseline, to predict mental health risk one year later. Predictors comprised questionnaires capturing social environment, behaviors, and brain imaging measures. Data preprocessing involved standardizing predictors, imputing missing values, and splitting the data into stratified training, validation, and test sets. The outcome variable, the p-factor, was derived through factor analysis of CBCL syndrome scales and categorized into risk quartiles. Model performance was evaluated across various scenarios, including predicting conversion to high-risk status. Predictor importance was analyzed using SHAP values. Sensitivity analyses assessed the impact of different factor models on p-factor estimation and the model’s ability to generalize to different populations in the US (see Supplementary Figure 5 for an overview of the methods).

Supplementary Material

Supplementary tables 1-6
Supplementary material

ACKNOWLEDGEMENTS

M. E. is supported by grant K01-MH127309 from the National Institute of Mental Health (NIMH). T. M. and A. C. are supported by Grants R01-AG032282 and R01-AG069939 from the National Institute on Aging (NIA) and Grant MR/P005918/1 from the Medical Research Council. Health Data Science at Duke is supported by the National Center for Advancing Translational Sciences (NCATS), National Institutes of Health, through Grant Award Number UL1 TR002553. The Duke AI Health Data Science Fellowship Program is supported by the above grant, the Duke Department of Biostatistics & Bioinformatics, and Duke AI Health. The content of this publication is solely the responsibility of the authors and does not necessarily represent the official views of the NIH. The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.

Footnotes

COMPETING INTERESTS

All authors declare no competing interests.

CODE AVAILABILITY

The code for all data processing, modeling, and analysis can be found here.

DATA AVAILABILITY

The data used in this study were obtained from the Adolescent Brain Cognitive Development (ABCD) Study, a publicly available dataset managed by the National Institutes of Health (NIH). Researchers can request access to the ABCD Study data through the National Institute of Mental Health Data Archive (NDA) at: https://nda.nih.gov/abcd.

Access is restricted to qualified researchers affiliated with an institution, and approval is contingent on compliance with the NDA Data Use Certification, which is available at: https://nda.nih.gov/ndapublicweb/Documents/NDA+Data+Access+Request+DUC+FINAL.pdf. Data can be used for scientific research purposes only and cannot be used for commercial or non-research applications. Requests for access are typically reviewed within two to four weeks.

REFERENCES

  • 1.Xiao Y, Brown TT, Snowden LR, Chow JC-C, Mann JJ. COVID-19 Policies, Pandemic Disruptions, and Changes in Child Mental Health and Sleep in the United States. JAMA Netw Open. 2023;6:e232716. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Samji H, Wu J, Ladak A, Vossen C, Stewart E, Dove N, et al. Review: Mental health impacts of the COVID-19 pandemic on children and youth - a systematic review. Child Adolesc Ment Health. 2022;27:173–89. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.T K, R M, A H, E L, R A, C W, et al. Navigating inequities in the delivery of youth mental health care during the COVID-19 pandemic: perspectives of youth, families, and service providers. Can J Public Health Rev Can Sante Publique [Internet]. 2022. [cited 2024 Sep 3];113. Available from: https://pubmed.ncbi.nlm.nih.gov/35852728/ [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Schmidhuber J Deep learning in neural networks: An overview. Neural Netw. 2015;61:85–117. [DOI] [PubMed] [Google Scholar]
  • 5.Posner J The Role of Precision Medicine in Child Psychiatry: What Can We Expect and When? J Am Acad Child Adolesc Psychiatry. 2018;57:813. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Romer AL, Ren B, Pizzagalli DA. Brain Structure Relations With Psychopathology Trajectories in the ABCD Study. J Am Acad Child Adolesc Psychiatry. 2023;62:895–907. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Sripada C, Angstadt M, Taxali A, Kessler D, Greathouse T, Rutherford S, et al. Widespread attenuating changes in brain connectivity associated with the general factor of psychopathology in 9- and 10-year olds. Transl Psychiatry. 2021;11:575. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Voorhees BWV, Paunesku D, Gollan J, Kuwabara S, Reinecke M, Basu A. Predicting Future Risk of Depressive Episode in Adolescents: The Chicago Adolescent Depression Risk Assessment (CADRA). Ann Fam Med. 2008;6:503–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.King M, Walker C, Levy G, Bottomley C, Royston P, Weich S, et al. Development and Validation of an International Risk Prediction Algorithm for Episodes of Major Depression in General Practice Attendees: The PredictD Study. Arch Gen Psychiatry. 2008;65:1368–76. [DOI] [PubMed] [Google Scholar]
  • 10.Hou J, van Wingen G, van de Mortel L, Popma A, Smit D. Predicting the onset of mental health problems in adolescents [Internet]. OSF; 2023. [cited 2025 Mar 25]. Available from: https://osf.io/edfba_v1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Wilson PW, D’Agostino RB, Levy D, Belanger AM, Silbershatz H, Kannel WB. Prediction of coronary heart disease using risk factor categories. Circulation. 1998;97:1837–47. [DOI] [PubMed] [Google Scholar]
  • 12.Stroup WW. Generalized Linear Mixed Models: Modern Concepts, Methods and Applications. CRC Press; 2012. [Google Scholar]
  • 13.Shah J, Eack SM, Montrose DM, Tandon N, Miewald JM, Prasad KM, et al. Multivariate prediction of emerging psychosis in adolescents at high risk for schizophrenia. Schizophr Res. 2012;141:189–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Bruni O, Ottaviano S, Guidetti V, Romoli M, Innocenzi M, Cortesi F, et al. The Sleep Disturbance Scale for Children (SDSC). Construction and validation of an instrument to evaluate sleep disturbances in childhood and adolescence. J Sleep Res. 1996;5:251–61. [DOI] [PubMed] [Google Scholar]
  • 15.Pigeon WR, Bishop TM, Krueger KM. Insomnia as a Precipitating Factor in New Onset Mental Illness: a Systematic Review of Recent Findings. Curr Psychiatry Rep. 2017;19:44. [DOI] [PubMed] [Google Scholar]
  • 16.Pigeon WR, Pinquart M, Conner K. Meta-analysis of sleep disturbance and suicidal thoughts and behaviors. J Clin Psychiatry. 2012;73:e1160–1167. [DOI] [PubMed] [Google Scholar]
  • 17.Telzer EH, Goldenberg D, Fuligni AJ, Lieberman MD, Gálvan A. Sleep variability in adolescence is associated with altered brain development. Dev Cogn Neurosci. 2015;14:16–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Uccella S, Cordani R, Salfi F, Gorgoni M, Scarpelli S, Gemignani A, et al. Sleep Deprivation and Insomnia in Adolescence: Implications for Mental Health. Brain Sci. 2023;13:569. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Freeman D, Sheaves B, Waite F, Harvey AG, Harrison PJ. Sleep disturbance and psychiatric disorders. Lancet Psychiatry. 2020;7:628–37. [DOI] [PubMed] [Google Scholar]
  • 20.Harvey AG, Murray G, Chandler RA, Soehner A. Sleep Disturbance as Transdiagnostic: Consideration of Neurobiological Mechanisms. Clin Psychol Rev. 2011;31:225–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Lunsford-Avery JR, Bidopia T, Jackson L, Sloan JS. Behavioral Treatment of Insomnia and Sleep Disturbances in School-Aged Children and Adolescents. Child Adolesc Psychiatr Clin N Am. 2021;30:101–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Harvey AG, Dong L, Hein K, Yu SH, Martinez AJ, Gumport NB, et al. A randomized controlled trial of the Transdiagnostic Intervention for Sleep and Circadian Dysfunction (TranS-C) to improve serious mental illness outcomes in a community setting. J Consult Clin Psychol. 2021;89:537–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Pettersson E, Larsson H, D’Onofrio BM, Lichtenstein P. Associations Between General and Specific Psychopathology Factors and 10-Year Clinically Relevant Outcomes in Adult Swedish Twins and Siblings. JAMA Psychiatry. 2023;80:728–37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Moore TM, Calkins ME, Satterthwaite TD, Roalf DR, Rosen AFG, Gur RC, et al. Development of a computerized adaptive screening tool for overall psychopathology (“p”). J Psychiatr Res. 2019;116:26–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Jones JD, Boyd RC, Sandro AD, Calkins ME, Los Reyes AD, Barzilay R, et al. The General Psychopathology “p” Factor in Adolescence: Multi-Informant Assessment and Computerized Adaptive Testing. Res Child Adolesc Psychopathol. 2024; [DOI] [PubMed] [Google Scholar]
  • 26.Marek S, Tervo-Clemmens B, Calabro FJ, Montez DF, Kay BP, Hatoum AS, et al. Reproducible brain-wide association studies require thousands of individuals. Nature. 2022;603:654–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Antoniou T, Mamdani M. Evaluation of machine learning solutions in medicine. Cmaj. 2021;193:E1425–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Tibshirani R Regression Shrinkage and Selection Via the Lasso. J R Stat Soc Ser B Methodol. 1996;58:267–88. [Google Scholar]
  • 29.Quinonero-Candela J, Sugiyama M, Schwaighofer A, Lawrence ND. Dataset Shift in Machine Learning. MIT Press; 2022. [Google Scholar]

METHODS-ONLY REFERENCES

  • 30.Garavan H, Bartsch H, Conway K, Decastro A, Goldstein RZ, Heeringa S, et al. Recruiting the ABCD sample: Design considerations and procedures. Dev Cogn Neurosci. 2018;32:16–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Karcher NR, Barch DM. The ABCD study: understanding the development of risk for mental and physical health outcomes. Neuropsychopharmacol Off Publ Am Coll Neuropsychopharmacol. 2021;46:131–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Felitti VJ, Anda RF, Nordenberg D, Williamson DF, Spitz AM, Edwards V, et al. Relationship of childhood abuse and household dysfunction to many of the leading causes of death in adults: The Adverse Childhood Experiences (ACE) Study. Am J Prev Med. 1998;14:245–58. [DOI] [PubMed] [Google Scholar]
  • 33.Kind AJ, Jencks S, Brock J, Yu M, Bartels C, Ehlenbach W, et al. Neighborhood socioeconomic disadvantage and 30-day rehospitalization: a retrospective cohort study. Ann Intern Med. 2014;161:765–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Parlatini V, Itahashi T, Lee Y, Liu S, Nguyen TT, Aoki YY, et al. White matter alterations in Attention-Deficit/Hyperactivity Disorder (ADHD): a systematic review of 129 diffusion imaging studies with meta-analysis. Mol Psychiatry. 2023;28:4098–123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Xia J, Chen N, Qiu A. Unraveling Multimodal Brain Signatures: Deciphering Transdiagnostic Dimensions of Psychopathology in Adolescents. Adv Intell Syst. n/a:2300577. [Google Scholar]
  • 36.Gordon EM, Laumann TO, Adeyemo B, Huckins JF, Kelley WM, Petersen SE. Generation and Evaluation of a Cortical Area Parcellation from Resting-State Correlations. Cereb Cortex. 2016;26:288–303. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Verbruggen F, Logan GD. Response inhibition in the stop-signal paradigm. Trends Cogn Sci. 2008;12:418–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Knutson B, Westdorp A, Kaiser E, Hommer D. FMRI Visualization of Brain Activity during a Monetary Incentive Delay Task. NeuroImage. 2000;12:20–7. [DOI] [PubMed] [Google Scholar]
  • 39.Cohen A, Conley M, Dellarco D, Casey B. The impact of emotional cues on short-term and long-term memory during adolescence. Proc Soc Neurosci San Diego CA Novemb. 2016; [Google Scholar]
  • 40.Clark DA, Hicks BM, Angstadt M, Rutherford S, Taxali A, Hyde L, et al. The General Factor of Psychopathology in the Adolescent Brain Cognitive Development (ABCD) Study: A Comparison of Alternative Modeling Approaches. Clin Psychol Sci J Assoc Psychol Sci. 2021;9:169–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Murtagh F Multilayer perceptrons for classification and regression. Neurocomputing. 1991;2:183–97. [Google Scholar]
  • 42.Lipton ZC, Berkowitz J, Elkan C. A Critical Review of Recurrent Neural Networks for Sequence Learning [Internet]. arXiv; 2015. [cited 2024 Apr 5]. Available from: http://arxiv.org/abs/1506.00019 [Google Scholar]
  • 43.Graves A Long Short-Term Memory. In: Graves A, editor. Supervised Seq Label Recurr Neural Netw [Internet]. Berlin, Heidelberg: Springer; 2012. [cited 2024 Sep 13]. p. 37–45. Available from: 10.1007/978-3-642-24797-2_4 [DOI] [Google Scholar]
  • 44.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is All you Need. [Google Scholar]
  • 45.Bottou L Large-Scale Machine Learning with Stochastic Gradient Descent. In: Lechevallier Y, Saporta G, editors. Proc COMPSTAT2010. Heidelberg: Physica-Verlag HD; 2010. p. 177–86. [Google Scholar]
  • 46.Nesterov YE. A method of solving a convex programming problem with convergence rate O\bigl(k^2\bigr). Dokl Akad Nauk. Russian Academy of Sciences; 1983. p. 543–7. [Google Scholar]
  • 47.Bergstra J, Yamins D, Cox D. Making a Science of Model Search: Hyperparameter Optimization in Hundreds of Dimensions for Vision Architectures. Proc 30th Int Conf Mach Learn [Internet]. PMLR; 2013. [cited 2024 Sep 5]. p. 115–23. Available from: https://proceedings.mlr.press/v28/bergstra13.html [Google Scholar]
  • 48.Krogh A, Hertz J. A Simple Weight Decay Can Improve Generalization. Adv Neural Inf Process Syst [Internet]. Morgan-Kaufmann; 1991. [cited 2024 Sep 8]. Available from: https://proceedings.neurips.cc/paper_files/paper/1991/hash/8eefcfdf5990e441f0fb6f3fad709e21-Abstract.html [Google Scholar]
  • 49.Srivastava N, Hinton G, Krizhevsky A, Sutskever I, Salakhutdinov R. Dropout: a simple way to prevent neural networks from overfitting. J Mach Learn Res. 2014;15:1929–58. [Google Scholar]
  • 50.Efron B Bootstrap Methods: Another Look at the Jackknife. In: Kotz S, Johnson NL, editors. Breakthr Stat Methodol Distrib [Internet]. New York, NY: Springer; 1992. [cited 2024 Mar 15]. p. 569–93. Available from: 10.1007/978-1-4612-4380-9_41 [DOI] [Google Scholar]
  • 51.Lundberg SM, Lee S-I. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30. [Google Scholar]
  • 52.Sundararajan M, Taly A, Yan Q. Axiomatic attribution for deep networks. Int Conf Mach Learn. PMLR; 2017. p. 3319–28. [Google Scholar]
  • 53.Elliott MR, Valliant R. Inference for Nonprobability Samples. Stat Sci. 2017;32:249–64. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary tables 1-6
Supplementary material

Data Availability Statement

The data used in this study were obtained from the Adolescent Brain Cognitive Development (ABCD) Study, a publicly available dataset managed by the National Institutes of Health (NIH). Researchers can request access to the ABCD Study data through the National Institute of Mental Health Data Archive (NDA) at: https://nda.nih.gov/abcd.

Access is restricted to qualified researchers affiliated with an institution, and approval is contingent on compliance with the NDA Data Use Certification, which is available at: https://nda.nih.gov/ndapublicweb/Documents/NDA+Data+Access+Request+DUC+FINAL.pdf. Data can be used for scientific research purposes only and cannot be used for commercial or non-research applications. Requests for access are typically reviewed within two to four weeks.

RESOURCES