Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Sep 23.
Published before final editing as: Clin Psychol Sci. 2026 Sep 22:10.1177/21677026261473154. doi: 10.1177/21677026261473154

Leveraging machine learning to personalize depression treatment: A pre-registered study of 828 adults randomized to a digital single-session intervention or waitlist

EJ Jardas 1, Jacqueline Howard 2, Lorenzo Lorenzo-Luaces 1
PMCID: PMC13595997  NIHMSID: NIHMS2196400  PMID: 42775236

Abstract

Some digital single-session interventions (SSI) for depression appear effective, at least in youth, but not everyone benefits. The present study uses machine learning methods to develop a treatment matching algorithm for a digital SSI, the Common Elements Toolbox (COMET), versus a waitlist control. 828 adults with a current or past mental health problem were randomized to COMET or a waitlist control. Elastic net regularization models with 10-fold cross-validation were used to develop a Personalized Advantage Index (PAI) indicating the relative benefit of receiving COMET over the waitlist in 2-week post-treatment depressive symptoms. In the 20% held-out test data, PAI did not interact with treatment to predict depression severity post-treatment (β = 0.880, SE = 1.21, t = −0.72, p = .47), indicating that our treatment matching algorithm was not able to provide statistically significant recommendations. Even in a large sample, personalized treatment recommendations for digital SSIs are difficult to develop.

Keywords: digital SSIs, depression, treatment-matching, machine learning, randomized controlled trial, personalized advantage index


Depression is a leading cause of disability in the United States and worldwide (Murray et al., 2012; Patel et al., 2016; Whiteford et al., 2013). Yet, more than half of people with depression worldwide do not receive treatment (Kohn et al., 2004; National Institute of Mental Health, 2022). The gap between needing treatment and receiving it is often attributed to barriers like stigma, lack of available providers, and high cost of treatment (Andrade et al., 2014). Digital mental health interventions leverage the relatively high level of internet access across the United States to help connect resources to individuals who might not otherwise be able to access effective treatments (Ramos et al., 2024). Single-session interventions (SSIs) distill treatment skills into one brief, often self-guided session, requiring less time and money from patients and providers. Digital SSIs combine the scalability of brief, low-cost interventions with the accessibility of online treatments. Therefore, they have been touted as having the potential to reduce the burden of common mental health problems like depression (Kazdin & Rabbitt, 2013; Schleider & Beidas, 2022; Szkody et al., 2023).

The Common Elements Toolbox (COMET) is an SSI for youth and adults that incorporates elements from cognitive behavioral therapy (CBT) and positive psychology. COMET was seen as feasible by adolescents and school officials in a focus group (n = 21; Wasil et al., 2020) and was associated with improvements in perceived coping among US graduate students relative to pre-treatment (n = 189; Wasil et al., 2021). Moreover, two small randomized controlled trials (RCTs) of COMET found evidence supporting the efficacy of COMET. COMET led to reduced depressive symptoms relative to a study skills control among Kenyan adolescents (n = 103; Osborn et al., 2020) and relative to an attention control questionnaire among UK undergraduate students (n = 147; Lambert et al., 2025), although the latter study had high attrition and inadequate handling of missing data. Despite these early promising findings, in our recent RCT of 828 adult online workers, which is the most well-powered trial to date, COMET was not associated with significant improvements in depression compared to a waitlist control after 2, 4, or 8 weeks (Lorenzo-Luaces & Howard, 2023). The authors conducted various sensitivity analyses to handle missing data differently and subgroup analyses focusing on the most symptomatic patients but found no statistically or clinically significant effect of COMET. In another study in which adults could self-select into COMET, COMET was not associated with improvements in distress (n = 275; Lorenzo-Luaces & Howard, 2023). These findings can be interpreted as suggesting that COMET is either ineffective or only effective for specific subgroups (e.g., students).

Evidence supports the idea that subgroups of patients may benefit from certain interventions more than others. In an analysis of 51,853 patients from 306 trials, Kaiser et al. (2022) found that variability in depression scores was 9% higher in treatment groups than in control groups, suggesting that some patients do benefit more (or less) from psychotherapy than the group average. To our knowledge, no meta-analysis has use the methods Kaiser et al. (2022) employed to quantify heterogeneity of treatment effects for SSIs specifically. Terhorst et al. (2024) found that variability in the effects of internet-based interventions was highest among studies that had patients with more severe symptoms, suggesting that severity may need to play a role in treatment matching for SSIs and other digital interventions. Moreover, meta-analyses of SSIs and digital interventions for depression suggest there is substantial heterogeneity in treatment effects across studies. For example, Schleider & Weisz (2017)’s meta-analysis of youth SSIs found significant heterogeneity in effect size across studies (I2 = 89%) and Schleider et al. (2025)’s umbrella review of meta-analyses on SSIs across age groups found moderate heterogeneity (I2 = 43%, see also (Ball et al., 2024; Moshe et al., 2021). These results suggest that a moderate-to-high proportion of variability in effect size across studies is due to true heterogeneity rather than sampling error. That is, while some individuals may benefit from SSIs, some may not, and relatively little is known about how to identify those who will.

Indeed, the question of “what works for whom” has been longstanding in the field (Paul, 1967). Variability in treatment response can be influenced by a range of factors, including demographic variables like age, clinical characteristics like baseline severity or co-morbid symptoms, and psychosocial factors like life stressors or lack of social support (Kessler et al., 2017). However, it is unlikely that any one variable singularly predicts treatment response (Cohen & DeRubeis, 2018; DeRubeis, 2019). In the era of precision medicine, the field may be able to move beyond simple identification of moderating variables toward instead identifying specific individuals who are most likely to benefit from particular treatments (Kaiser et al., 2022; Simon & Perlis, 2010).

The goal of treatment matching appears increasingly feasible given the rise of machine learning (ML) approaches capable of detecting complex, high-dimensional interactions between person-level variables and treatment outcomes (Dwyer et al., 2018). These statistical predictions can be used to make treatment recommendations. For example, the Personalized Advantage Index (PAI) metric uses prediction models to estimate the expected benefit of one treatment over another for a given individual (DeRubeis et al., 2014). Treatment recommendations derived from PAI have been found to suggest meaningful differences in treatment outcomes for a variety of disorder symptoms, including depression, PTSD, somatic symptoms, and borderline personality disorder (Meinke et al., 2024). Nevertheless, a systematic review found a relatively small average effect size for the benefits of personalization (Cohen’s d ≈ 0.30) and cautioned that many PAI studies had small samples, poor protections against data leakage, and an overall high risk of bias, likely leading to an overestimation of the utility of PAI (Meinke et al., 2024). Indeed, the review found the most methodologically robust studies do not show evidence for a PAI-related outcome improvement. There is a need for more PAI development studies with robust methods and large samples.

Because SSIs are scalable and accessible, an argument could be made that they should be disseminated without attempts to “personalize” who receives the intervention (Lorenzo-Luaces et al., 2023a). However, research in mental health suggests that the fact that an intervention could be made accessible to everyone does not mean that everyone should access it. For example, L. Harvey et al. (2023) found that a school-based universal intervention based on dialectical behavior therapy for adolescents (n = 1,071) led to deterioration in depressive and anxiety symptoms relative to attending class as usual (see also Foulkes et al., 2026). It would be helpful to identify subgroups of individuals for whom an intervention does not work (Simon et al., 2022) in order to avoid potential negative effects and offer alternative treatments which may be more effective, especially because patients are less motivated to continue seeking treatment if they perceive prior experiences as unhelpful (Harris et al., 2020). This could be concerning, for example, if individuals perceive a digital SSI as unhelpful and then fail to seek a higher level of care which could benefit them. Indeed, a recent mega-study of twelve SSIs for 7,505 adults with depression found that some SSIs decreased readiness to change after 4 weeks (Kaveladze et al., 2026).

In the current study, we attempt to develop a treatment selection algorithm for a digital SSI (COMET) administered to 828 adult online workers endorsing a history of psychopathology (Lorenzo-Luaces & Howard, 2023). A major strength of the present study is its large sample (n = 828) and the fact that it was specifically designed to measure candidate predictors for treatment selection based on a review which identified replicated moderators of treatment outcomes among individuals with depression (Kessler et al., 2017). To assess the replicability of subgroup findings and reduce the risk of overfitting, we use a supervised machine learning approach with a prespecified training and test split. If successful, the resulting treatment selection model would be able to identify a subgroup of individuals predicted to improve more with COMET than with waiting. Importantly, the model may also identify a subsample of people whose depression is predicted to improve more with the waitlist than with COMET—that is, people for whom the treatment is possibly worse than no treatment at all.

Transparency and Openness

Pre-registration.

The study design of the parent RCT was pre-registered at clinicaltrials.gov (NCT: 05379881). The analyses for the parent RCT and the present analyses were pre-registered at https://osf.io/63yzh.

Data, materials, code, and online resources.

The data, analysis code, and materials for the present study are available at https://osf.io/n6egr/. Supplemental Materials are available on the journal website.

Reporting.

We report all data exclusions, all manipulations, and all measures in the study. The power analyses informing our sample size are available in the publication on the results of the parent RCT (Lorenzo-Luaces & Howard, 2023).

Ethical approval.

The study was approved by the human subject research ethics review board of Indiana University and was carried out in accordance with the provisions of the World Medical Association Declaration of Helsinki.

Methods

Procedure.

The present study used data from a parallel two-arm randomized controlled trial (1:1 allocation) comparing changes in depression and anxiety among online workers randomized to receive COMET or a waitlist control (“waitlist”). Online workers living in the United States were recruited from Prolific, an online research platform. The study was advertised as “testing the efficacy of an online wellness activity called the Common Elements Toolbox (COMET) [which] helps you learn skills that may help to improve your mood and well-being”. Workers were included in the trial if they replied “Yes” to the question, “Do you have—or have you had—a diagnosed, ongoing mental health illness/condition?” There were no other inclusion or exclusion criteria. Participants completed baseline measures of mental health characteristics, then were randomized to COMET or the waitlist. Participants completed follow-up surveys at 2-, 4-, and 8-weeks post-baseline. The current study examines outcomes at 2-weeks only.

Intervention.

COMET is a 4-module SSI containing two core principles of CBT, cognitive restructuring and behavioral activation, and two core principles of positive psychology, gratitude and self-compassion. COMET takes approximately 25–40 minutes to complete. It presents individuals with psychoeducation in the form of text and brief exercises. Participants were shown how to schedule pleasant activities (for behavioral activation), use the “ABCD” technique (for cognitive restructuring), reflect on “three good things” (for gratitude), and write self-compassionate statements (for self-compassion). After the four modules were presented, participants were asked to choose one module to practice in the following week. COMET was not found to be efficacious for depression over the waitlist in the trial from which the present data are drawn (SMD = −0.05, 95% CI [−0.20, 0.07]; Lorenzo-Luaces & Howard, 2023).

Outcome.

The outcome measure is depressive symptoms as measured by the PHQ-8, which includes all items on the PHQ-9 except the item assessing suicidal ideation. The PHQ-9 is a self-report questionnaire assessing the frequency of DSM-5 symptoms of a major depressive episode (Kroenke et al., 2001). It uses a 0 (“not at all”) to 3 (“nearly every day”) scale, with total scores ranging from 0 to 27. Higher scores indicate more frequent depressive symptoms. An individual participant data meta-analysis of 16,742 individuals demonstrated that the PHQ-8 and PHQ-9 are highly correlated (r = .996), have similar distributional properties, appear to measure the same construct, and have equivalent sensitivity and specificity when semi-structured interviews are used as a reference standard (Shin et al., 2019; Wu et al., 2020). Therefore, and in line with the parent trial the data are drawn from, we multiplied PHQ-8 scores to match the scale of the PHQ-9. This linear transformation preserves the original distribution of the PHQ-8 but allows scales to be more interpretable and comparable with trials which use the PHQ-9. Due to this transformation, we refer to the outcome as “PHQ-9 scores”.

Predictors.

The battery of predictors was designed based on a review of replicated predictors of treatment outcomes among depressed individuals (Kessler et al., 2017). See Table 1 for a description of all included predictors. They include age, age of onset of distress symptoms, race, marital status, educational attainment, yearly income, employment status, lifetime psychiatric medication use, self-rated physical health, weekly alcohol use, social support, anxiety (GAD-7), depression (PHQ-9), well-being (WHO-5), functioning (WSAS), insomnia (ISI-3), habitual use of emotion regulation strategies (cognitive reappraisal and expressive suppression subscales from the ERQ), childhood adversity (ACES), and personality traits (neuroticism, extraversion, openness, conscientiousness, and agreeableness subscales from the TIPI).

Table 1.

Sociodemographic and clinical characteristics for 828 participants with a history of mental health concerns randomized to the Common Elements Toolbox (COMET) vs. a waitlist control, split into 80–20% training and test sets.

 Variable Training Test
 Categorical Variables (Yes / No) Yes (%) Yes (%)
 Person of color 165 (24.9%) 46 (27.7%)
 LGBTQ+ 226 (34.1%) 57 (34.3%)
 Alone (not dating or married) 300 (45.3%) 60 (36.1%)
 Higher education degree (associates or more) 395 (59.7%) 84 (50.6%)
 Working (full or part time or as a student) 475 (71.8%) 114 (68.7%)
 Lifetime psychiatric medication use 552 (83.4%) 140 (84.3%)
 In poor health (or terrible health) 127 (19.2%) 28 (16.9%)
 Frequent drinking (multiple times per week) 122 (18.4%) 31 (18.7%)
 Has social support (more than one friend) 516 (77.9%) 123 (74.1%)
 Numeric Variables Mean (SD) Mean (SD)
 Age (years) 35.7 (11.9) 36.0 (12.2)
 Yearly income (range 1–9) 4.4 (2.1) 4.6 (2.2)
 Age of onset of internalizing symptoms (years) 13.5 (7.1) 14.1 (7.7)
 Insomnia (ISI, range 0–12) 6.0 (2.9) 5.7 (3.1)
 Adverse childhood experiences (ACEs, range 0–10) 3.2 (2.6) 3.3 (2.6)
 Cognitive reappraisal (ERQ, range 1–7) 4.5 (1.3) 4.4 (1.2)
 Expressive suppression (ERQ, range 1–7) 3.8 (1.5) 3.9 (1.5)
 Anxious distress (component) 0.0 (1.0) −0.0 (1.0)
 Dysphoric distress (component) −0.0 (1.0) 0.0 (1.0)
 Impairment (component) 0.0 (1.0) −0.1 (1.0)
 Prior use of COMET skills (component) −0.0 (1.0) 0.0 (1.0)
 System usability (SUS, range 20–100) 68.0 (14.9) 68.1 (13.3)
 Extraversion (TIPI, range 1–7) 2.9 (1.7) 2.9 (1.6)
 Agreeableness (TIPI, range 1–7) 5.2 (1.3) 5.2 (1.3)
 Conscientiousness (TIPI, range 1–7) 4.9 (1.5) 4.8 (1.5)
 Neuroticism (TIPI, range 1–7) 3.5 (1.5) 3.6 (1.5)
 Openness (TIPI, range 1–7) 5.2 (1.3) 5.1 (1.3)
 Desire to learn more about flexible thinking (range 1–9) 6.9 (2.0) 6.8 (2.0)
 Desire to learn more about gratitude (range 1–9) 6.6 (2.1) 6.4 (2.0)
 Desire to learn more about activity scheduling (range 1–9) 7.2 (1.7) 7.1 (1.8)
 Desire to learn more about self-compassion (range 1–9) 7.3 (1.9) 7.1 (2.0)
 Depression (PHQ-9; range 0–27) 10.4 (6.8) 9.9 (7.0)

SD: standard deviation; LGBTQ+: Lesbian, gay, bisexual, transgender, queer/questioning; ISI: Insomnia Severity Index; ERQ: Emotion Regulation Questionnaire; SUS: System Usability Scale; TIPI: Ten-Item Personality Inventory; PHQ-9: Patient Health Questionnaire.

Originally, gender identity (cisgender vs. transgender) and sexual orientation (LGB+ vs. not) were pre-registered to be included as predictors, in part because sexual and gender minority individuals are especially likely to seek out alternative treatment modalities like single-session interventions (Dobias et al., 2021). However, due to the relatively small number of transgender participants, these two variables were collapsed into one categorical predictor encompassing both gender identity and sexual orientation (e.g., LGBTQ+ or not). To investigate whether individuals’ preferences for treatment predict outcomes, questions about participants’ interest in learning each of the four COMET skills were included, which represents a deviation from the pre-registration. Similarly, to investigate whether individuals’ preexisting strengths influenced treatment effectiveness, questions about participants’ use of each COMET skill in the past two weeks were included.

Analytic Plan

Data preprocessing.

The data were randomly divided (80/20) into a training set (n = 662) and a held-out test set (n = 166), stratifying over treatment condition. Note that the decision to split the data into training and test sets was not pre-registered but follows machine learning best practice (Calster et al., 2026). The spread and intercorrelation of all baseline predictors were examined, and categorical variables were recoded as binary (0, 1; see Table 1). Using the training data, all data were normalized and missing values were imputed using a non-parametric missing value imputation technique based on random forests (Stekhoven & Bühlmann, 2012), using the “missForestPredict” package in R with 500 trees (Albu et al., 2025). To prevent leakage, the imputation model from the training data was saved and then applied to the held-out test data. That is, the imputation model used on the test data was the same as the training data. The imputation model converged after two iterations, yielding a mean normalized out-of-bag error (NMSE) of 0.63 across variables.

High dimensionality in the data required the use of principal component analysis (PCA) to reduce correlations between predictors. First, baseline depression (PHQ-9), anxiety (GAD-7), well-being (WHO-5), and functioning (WSAS) were highly correlated. To reduce dimensionality, we standardized the items of these four scales and applied a PCA to the items extract components related to baseline severity and impairment. Note that use of PCA was not pre-registered. To prevent leakage, the PCA was performed on the training data only (n = 662), and then loadings were projected into the held-out test data. The first three components were labelled anxious distress, dysphoric distress, and impairment, corresponding to the loadings of the different measures. They were retained for further analysis. Second, due to significant intercorrelations between each of the prior skill use items, PCA was applied to the four questions about prior skill use. Again, PCA was performed on the training data, then projected into the held-out test data to prevent leakage. Only one component was retained, indicating greater skill use. Results of the PCAs are detailed in the Supplemental Material.

To prevent leakage, a “recipe” was prepared in the training set and applied to the test set. Recipes are objects that specify the steps to clean and preprocess raw data before modeling, just as a chef’s recipe lists steps to transform raw ingredients into a prepared dish. Recipes are a commonly used technique for preventing data leakage and overfitting in machine learning. In our recipe, all variables except the outcome were normalized. Information about the mean and standard deviation of each variable in the training data were stored in the recipe so that it could be applied to normalize each variable in the test data without ‘leaking’ information from the test data into the training data (Kuhn et al., 2025). The recipe was derived from the training set. Then, the data were split into a waitlist training set (n = 335), COMET training set (n = 327), waitlist test set (n = 84), and COMET test set (n = 82).

Model development.

One model was built to predict participant response to COMET (n = 327), and a second model was built to predict participant response to the waitlist (n = 335), each using PHQ-9 scores after 2 weeks as the outcome. Each model used elastic net regularization (Zou & Hastie, 2005) with 10-fold cross-validation (Krstajic et al., 2014). All models included all candidate predictors. This is a deviation from our pre-registration, where we initially planned to pre-screen variables using model-based recursive partitioning. We deviated from this plan because pre-screening variables would require us to use the same data to identify variables to build the model as well as to test model performance, which may constitute a form ‘double-dipping’ that can lead to overfitting (Apicella et al., 2025; Calster et al., 2026; Kriegeskorte et al., 2009; Lorenzo-Luaces et al., 2021).

Once the two models (i.e., for COMET and for the waitlist) were developed, each was applied to the full training sample (n = 662) to estimate two predicted PHQ-9 scores for each participant: one for response to COMET and the other for response to the waitlist. These values were subtracted from one another (waitlistpred - COMETpred) to generate a numerical representation of the predicted benefit of receiving COMET over the waitlist for each participant. Note that the unit for PAI is PHQ-9 scores, so a PAI of 1.0 indicates that an individual is predicted to have one point lower depression (PHQ-9) in COMET, and a PAI of −1.0 indicates that an individual is predicted to have one point lower depression (PHQ-9) on the waitlist. Lastly, a final treatment model was developed (Salditt et al., 2024) using elastic net regularization and 10-fold cross-validation to estimate predicted differences in post-treatment PHQ-9 scores in the training set (n = 662). This prediction is termed the Personalized Advantage Index (PAI). Higher PAI indicate a higher predicted response to COMET relative to the waitlist. Gridsearch methods minimizing root mean squared error (RMSE) were used to select the optimal COMET, waitlist, and PAI models.

Model evaluation.

PAIs are developed in the training set by attempting to model treatment differences. Therefore, the training model may erroneously detect ‘differences’ that are not actually replicable (e.g., not found in new data). This is a common problem in machine learning typically referred to as ‘overfitting’ (Dwyer et al., 2018; Pargent et al., 2023). Overfitting occurs when models fit to random noise in the training data rather than (or in addition to) true underlying associations between variables. If a model’s ability to make accurate predictions is evaluated in the same training data the model (over)fit to, it is likely to yield an overly optimistic estimate of model performance. As a result, it is essential to test performance in data the model is naive to. Therefore, we use the held-out test data to evaluate model performance.

To characterize model performance, we must determine whether the model-estimated difference in PHQ-9 scores between the waitlist and COMET conditions (e.g., PAI) reflects actual differences in PHQ-9 scores between the waitlist and COMET conditions. In prior work, a common approach to treatment selection involves dichotomizing PAI at a specific threshold (e.g., zero) to generate treatment recommendations. Then, outcomes can be compared between participants who were randomly assigned to their PAI-indicated versus -contraindicated treatment. PAI are naturally continuous rather than dichotomous, so this approach requires researchers to choose a somewhat arbitrary cut point to derive treatment recommendations.

Therefore, in the present analysis, instead of using indicated versus contraindicated subgroups, we used a statistical approach more appropriate for a continuous variable. Specifically, in the test set (n = 166), we estimated the interaction between PAI and treatment condition followed by probing of the interaction using the Johnson-Neyman technique. The interaction term indicates whether the magnitude and direction of the treatment difference in 2-week PHQ-9 scores varies as a function of our PAI. Specifically, if there were a treatment-matching effect in our data, individuals with more positive PAI (i.e., lower predicted depressive symptoms with COMET) would be expected to exhibit correspondingly better outcomes if they were assigned to COMET. Similarly, individuals with more negative PAI (i.e., lower predicted depressive symptoms with the waitlist) would be expected to exhibit correspondingly better outcomes if they were assigned to the waitlist. The Johnson-Neyman technique (Hayes & Matthes, 2009; Johnson & Neyman, 1936) was applied to probe the interaction and identify the range of PAI at which the difference in depression outcomes between COMET and the waitlist condition crosses the p < .05 threshold. Importantly, in our data, the average difference between treatments was not significant at p < .05. The Johnson-Neyman technique can identify a subgroup of individuals for which there is a significant predicted difference between treatments at p < .05.

Results

Sample Characteristics

Sample characteristics are available in Table 1. A detailed description of the sample is available in Lorenzo-Luaces & Howard (2023). All participants from the parent RCT were included in the present study. In summary, the sample was approximately two-thirds female (65.7%) with a high proportion of LGBTQ+ individuals (34.8%) and representation of racial/ethnic minorities slightly below the U.S. population (73.2% White, 8.2% Latino, 6.2% Black, and 12% other; US Census Bureau, 2024). All participants were located in the United States. Regarding highest level of education completed, 43.8% had a high school diploma or less, 43.8% had an associates or bachelor’s degree, and 13.8% had a masters, doctoral, or professional degree.

Model Development

During model development in the training dataset, the predicted week 2 post-COMET PHQ-9 score for all individuals (i.e., whether they actually received COMET or not) was an average of 10.50 (SD = 4.55, min = −0.848, max = 21.50). Predicted week 2 post-waitlist PHQ-9 score was an average of 10.34 (SD = 5.06, min = −2.04, max = 22.90). PAI, which represent the predicted difference in PHQ-9 scores between receiving COMET and the waitlist, had an average of −0.17 (SD = 0.78; min = −2.71, max = 2.62) in the training data. Higher scores indicate a higher predicted benefit from COMET and lower negative scores indicate a higher predicted benefit from the waitlist. Scores close to zero indicate no predicted difference between the two. Participants’ PAI were similar in the held-out test data but with a somewhat more restricted range (M = −0.21, SD = 0.73, min = −2.15, max = 1.84). Within each treatment condition, actual week 2 PHQ-9 scores were highly correlated with predicted week 2 scores in both the training (COMET: r = .818, waitlist: r = .869) and held-out test data (COMET: r = .841, waitlist: r = .864). The results of hyperparameter tuning can be found in the Supplemental Material. Readers should be aware that estimates of model performance in the training set provides little if any evidence about whether PAI can be used for making personalized treatment recommendations. Rather, conclusions about model performance should be drawn from findings in the test set. Therefore, we do not report results about model performance in the training data. More information about performance in training is provided in the Supplemental Material only to characterize the extent of overfitting.

Treatment Matching

To evaluate whether the model we developed could be useful for treatment matching in new data, we used the held-out test dataset. In the test set (Figure 1), PAI did not interact with treatment to predict post-treatment depression severity (β = 0.880, SE = 1.21, t = −0.72, p = .47), suggesting that the effect of COMET did not vary systematically across levels of PAI. The Johnson-Neyman technique did not identify a range of PAI for which the predicted difference in PHQ-9 scores between COMET and the waitlist crosses the p < .05 threshold. Notably, echoing the findings of Ahuvia et al., (2023), the predicted responses to COMET vs. waitlist were very highly correlated in both the training (r = .993) as well as in the held-out test data (r = .993).

Figure 1.

Figure 1.

Predicted 2-week PHQ-9 scores by treatment condition (COMET versus the waitlist) across values of the Personalized Advantage Index (PAI) in a held-out test set of individuals with a history of psychopathology (n = 166).

Two week post-COMET and post-waitlist PHQ-9 scores as a function of the Personalized Advantage Index (PAI) in a held-out test set (n = 166). Shaded bands indicate 95% confidence intervals. There was no significant interaction between PAI and condition.

Discussion

We attempted to develop a treatment selection algorithm to guide allocation to a digital SSI relative to a waitlist. The broader goal of this work is to determine whether it is possible to identify, a priori, which individuals would benefit more from receiving a low-intensity intervention than from waiting, with the idea that this knowledge would be actionable in routine care. In the RCT from which these data were drawn, the digital SSI COMET was found to confer no benefit, on average, relative to the waitlist. Therefore, we explored whether we could predict differential treatment response to make treatment recommendations at the individual level. In the test dataset, no subgroup of individuals was identified as being more likely to benefit from COMET than from the waitlist.

Our findings join a growing number of large sample studies suggesting that the reality of precision treatment selection may be more complex and difficult to achieve than originally anticipated (Ahuvia et al., 2023; Van Bronswijk et al., 2024). Even in studies with encouraging findings, effects of using a treatment algorithm are often relatively small (Delgadillo et al., 2022). Small sample sizes have frequently been cited as a limitation in prior studies of treatment matching, with recommendations to include at least 300 participants per treatment arm (Luedtke et al., 2019). However, in our study, which had a relatively large sample, treatment matching was still unsuccessful. A recent comprehensive review of the literature by Cuijpers et al. (2026) similarly found that relatively little can be concluded about how to personalize treatment for depression, that machine learning studies attempting to do so are still in quite early stages of development, and that larger studies are needed. Indeed, most clinical prediction models are yield overly optimistic performance metrics due to methodological pitfalls (Calster et al., 2026), and PAI models have been found to be no exception (Meinke et al., 2024). In light of this literature, the inability to internally validate our findings in our held-out data suggests that caution around the current enthusiasm for personalized medicine approaches in mental health is warranted.

The present study had several limitations. First, the sample included online workers who indicated that they have or have had a diagnosed, ongoing mental health condition. Given the lack of specificity, participants may have included those for whom depression is not a key concern. Although the majority of participants would have met screening criteria for a depressive disorder based on the PHQ-9, some participants may have had primary concerns that would have been better addressed with other interventions. Second, our study may have lacked variables that would have predicted differential treatment outcomes, had they been measured and included in our models. For example, we only collected self-report variables. Finally, it is possible that personalized medicine effects are so small that even larger samples than the current one are required.

Despite these limitations, several factors are worth considering. First, our sample size was large relative to other PAI studies. For example, our sample size was larger than 17 of the 19 studies reviewed by Meinke et al. (2024), only slightly smaller than the samples by Ahuvia et al. (2023; n = 996), though much smaller than a non-randomized study by Schwartz et al. (2021; n = 1,379). While it is possible that even larger samples are required for this kind of research, it is worth asking to what extent personalized medicine approaches can be expected to improve overall treatment outcomes if matching effects are so small. Additionally, we undertook a methodologically rigorous prediction effort, including pre-registering our analytic plan and using strategies to minimize bias by attempting to prevent data leakage and using a held-out sample. Meinke et al. (2024) found that the more rigorous PAI studies produced smaller and non-significant treatment matching effects when compared to the least rigorous ones, which is consistent with our data. Another strength is that we combined a battery of predictors informed by prior research on moderators across a range of areas (e.g., demographics, features of depression, co-morbid symptoms, potential mechanisms, personality). Finally, the level of heterogeneity in our sample can be seen as a strength, because prediction efforts require the presence of variability in the sample. Prior efforts have been hampered by focusing on more heterogeneous samples (DeRubeis et al., 2014; Lorenzo-Luaces & DeRubeis, 2018).

One possible explanation for our null results is that the features used in the model may not contain sufficient information to enable treatment matching. Many predictors commonly used in psychotherapy research, such as age, gender, or diagnostic category, are only weakly related to differential treatment outcomes (Lorenzo-Luaces et al., 2021). Perhaps variables that are closer to mechanisms of treatment would be more relevant for predicting who will benefit from a particular intervention. It is possible that predictors that have been understudied in the literature and therefore not included in our study could provide additional predictive power. Moreover, in the context of digital SSIs, and especially with samples of online workers, it is possible that participants are inattentive or disengaged. Engagement appears to moderate digital SSI effectiveness, and removing noncompliant participants from trials can increase the apparent intervention effectiveness and reduce heterogeneity in treatment effects (Zainal et al., 2025). However, the original analysis of the current trial tested for inattention by triangulating self-report and meta-data and reviewing all text entry and found overall high levels of engagement.

The intervention itself may not have produced strong enough effects to support meaningful subgroup differentiation. In the parent RCT, COMET did not outperform a waitlist on average. In the absence of strong average treatment effects, the maximum size of interactions between treatment and individual differences (e.g., PAI) may be constrained, making it more difficult to detect meaningful subgroup differences (Lorenzo-Luaces, 2023a; McClelland & Judd, 1993; Pincus et al., 2011). Moreover, even in the held-out sample, the treatment- and waitlist-predicted outcomes were almost perfectly correlated, a finding that is similar to that of Ahuvia et al. (2023). Lack of differential treatment effects may be of particular concern for treatment selection of digital SSIs. Although SSIs are often lauded for their accessibility, their effects tend to be modest on average (Schleider et al., 2025; Schleider & Weisz, 2017). Nevertheless, even a small effect size in favor of personalized treatment has the potential to improve outcomes for some individuals (Nye et al., 2023) and may be cost-effective at a large scale (Delgadillo et al., 2022).

Future Directions

Although unguided digital SSIs appear promising in the treatment of youth distress, less work has focused on their efficacy for depressed adults. Future work may investigate whether brief digital interventions could be tailored to be more effective for depressed adults while still being more scalable than weekly one-on-one psychotherapy. For example, a review by Kim et al. (2023) identified several RCTs of single-session interventions for depression delivered by counselors or trained paraprofessionals that appeared to be effective relative to no treatment or a waitlist. Moreover, a recent RCT of a four-session, self-guided digital intervention focused on behavioral activation conferred a significant treatment effect for depression within a sample of online workers (Peipert & Lorenzo-Luaces, 2025). Digital interventions, which are brief but guided by a professional or paraprofessional (rather than self-guided), or interventions that are brief but longer than one session, may offer an additional avenue through which to provide more scalable treatments.

Moreover, a primary goal of treatment matching is to reduce the burden of depression by improving treatment outcomes. Given the difficulty of treatment matching thus far, more research could focus on reducing the burden of depression in other ways (Lorenzo-Luaces, 2023b). Indeed, a recent simulation by Cuijpers et al. (2025) found that no single innovation in psychotherapy is likely to be sufficient to raise patient treatment response to greater than 99% on its own. Rather, multiple incremental innovations are needed to address the depression treatment gap. For example, most individuals who need treatment do not receive it, and even fewer receive empirically supported treatments (A. G. Harvey & Gumport, 2015; Wang et al., 2005). According to a simulation, increasing the rate of individuals who receive treatment would reduce the burden of depression far more than improving outcomes among those who already do access it (Wilhelm et al., 2024). Major barriers to seeking treatment include structural barriers, like cost or low provider availability, as well as attitudinal barriers, such as lack of perceived need for therapy or beliefs that therapy will not work (Andrade et al., 2014).

Digital SSIs are not necessarily immune to these barriers. Even very small costs reduce willingness to try digital interventions (Lorenzo-Luaces et al., 2024), and attitudinal barriers, such as lack of perceived need for treatment or beliefs that treatment will not be effective, correlate to lower willingness to use digital interventions (Starvaggi & Lorenzo-Luaces, 2025). More research could focus on changing attitudes about and increasing awareness of digital mental health interventions. Finally, the shift toward personalized medicine is in part meant to avoid the process of “trial and error” which is common in routine care. Trial and error is associated with disengagement from the treatment process, even though individuals who remain in treatment appear to derive benefit from multiple attempts (Harris et al., 2020; Rush et al., 2006). Rather than attempting to forgo trial and error entirely through the use of personalized medicine, researchers could design studies that better simulate the process of undergoing multiple treatments to try to understand the most effective and cost-effective way to sequence treatment (Lorenzo-Luaces, 2023b; Lorenzo-Luaces & Fite, 2024).

Conclusion

Although the goal of identifying “what works for whom” remains compelling, our results suggest that it may be more difficult to achieve than initially assumed. The current study found no evidence of replicable subgroups who benefit more from a digital SSI than from waiting. Our findings highlight the importance of rigorous design in precision mental health research. Many published studies on treatment matching rely on small samples or models that are not validated in independent datasets (Meinke et al., 2024) or held-out test samples. Such practices may lead to unrealistic expectations about the promise of treatment-matching. Future work can focus on disseminating treatment, improving attitudinal barriers to care, and examining processes for sequential treatment.

Supplementary Material

1

Funding

This research was partially funded by grants KL2TR002530 and UL1TR002529 (Anantha Shekar, principal investigator) from the National Institutes of Health, National Center for Advancing Translational Sciences, Clinical and Translational Sciences Award, which provided support for LL-L. The work was also supported by the 2020 Global Mental Health Fellowship through the American Psychological Association (APA) and the International Union of Psychological Science (IUPsyS). EJ is supported by the National Science Foundation Graduate Research Fellowship Program under Grant No. 2240777.

Footnotes

Conflicts of Interest

Prof. Lorenzo-Luaces has received consulting fees from Syra Health, who had no involvement in the current work.

References

  1. Ahuvia IL, Mullarkey MC, Sung JY, Fox KR, & Schleider JL (2023). Evaluating a treatment selection approach for online single-session interventions for adolescent depression. Journal of Child Psychology and Psychiatry, 64(12), 1679–1688. 10.1111/jcpp.13822 [DOI] [PubMed] [Google Scholar]
  2. Albu E, Gao S, Wynants L, & Calster BV (2025). missForestPredict—Missing data imputation for prediction settings. PLOS ONE, 20(11), e0334125. 10.1371/journal.pone.0334125 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Andrade LH, Alonso J, Mneimneh Z, Wells JE, Al-Hamzawi A, Borges G, Bromet E, Bruffaerts R, De Girolamo G, De Graaf R, Florescu S, Gureje O, Hinkov HR, Hu C, Huang Y, Hwang I, Jin R, Karam EG, Kovess-Masfety V, … Kessler RC (2014). Barriers to mental health treatment: Results from the WHO World Mental Health surveys. Psychological Medicine, 44(6), 1303–1317. 10.1017/S0033291713001943 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Apicella A, Isgrò F, & Prevete R (2025). Don’t push the button! Exploring data leakage risks in machine learning and transfer learning. Artificial Intelligence Review, 58(11), 339. 10.1007/s10462-025-11326-3 [DOI] [Google Scholar]
  5. Ball J, Thompson Z, Meiser-Stedman R, & Chiu K (2024). Self-administered single session interventions for mental health in young people: A systematic review and meta-analysis. SSRN. 10.2139/ssrn.4856261 [DOI] [PubMed] [Google Scholar]
  6. Calster BV, van Smeden M, van Amsterdam W, Coemans M, Wynants L, & Steyerberg EW (2026). The enemies of reliable and useful clinical prediction models: A review of statistical and scientific challenges (Pt. 465–492). Annual Review of Statistics and Its Application, 13. https://doi.org/https://doi-org.proxyiub.uits.iu.edu/10.1146/annurev-statistics-042324-123749 [Google Scholar]
  7. Cohen ZD, & DeRubeis RJ (2018). Treatment selection in depression. Annual Review of Clinical Psychology, 14(1), 209–236. 10.1146/annurev-clinpsy-050817-084746 [DOI] [PubMed] [Google Scholar]
  8. Cuijpers P, Harrer M, & Furukawa T (2025). Assessing the strength of innovations in the treatment of depression. The British Journal of Psychiatry, 227(5), 810–813. 10.1192/bjp.2025.98 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Cuijpers P, Harrer M, & Furukawa TA (2026). Innovations to improve outcomes and uptake of psychotherapies for mental disorders: A state-of-the-art review. World Psychiatry, 25(1), 4–33. 10.1002/wps.70002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Delgadillo J, Ali S, Fleck K, Agnew C, Southgate A, Parkhouse L, Cohen ZD, DeRubeis RJ, & Barkham M (2022). Stratified care vs stepped care for depression: A cluster randomized clinical trial. JAMA Psychiatry, 79(2), 101. 10.1001/jamapsychiatry.2021.3539 [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. DeRubeis RJ (2019). The history, current status, and possible future of precision mental health. Behaviour Research and Therapy, 123, 103506. 10.1016/j.brat.2019.103506 [DOI] [PubMed] [Google Scholar]
  12. DeRubeis RJ, Cohen ZD, Forand NR, Fournier JC, Gelfand LA, & Lorenzo-Luaces L (2014). The personalized advantage index: Translating research on prediction into individualized treatment recommendations. A demonstration. PLoS ONE, 9(1), e83875. 10.1371/journal.pone.0083875 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Dobias ML, Schleider JL, Jans L, & Fox KR (2021). An online, single-session intervention for adolescent self-injurious thoughts and behaviors: Results from a randomized trial. Behaviour Research and Therapy, 147, 103983. 10.1016/j.brat.2021.103983 [DOI] [PubMed] [Google Scholar]
  14. Dwyer DB, Falkai P, & Koutsouleris N (2018). Machine learning approaches for clinical psychology and psychiatry. Annual Review of Clinical Psychology, 14(1), 91–118. 10.1146/annurev-clinpsy-032816-045037 [DOI] [PubMed] [Google Scholar]
  15. Foulkes L, Guzman Holst C, & Andrews JL (2026). Potential harm from universal school-based mental health interventions: Candidate mechanisms and future directions. Current Opinion in Psychology, 67, 102196. 10.1016/j.copsyc.2025.102196 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Harris MG, Kazdin AE, Chiu WT, Sampson NA, Aguilar-Gaxiola S, Al-Hamzawi A, Alonso J, Altwaijri Y, Andrade LH, Cardoso G, Cía A, Florescu S, Gureje O, Hu C, Karam EG, Karam G, Mneimneh Z, Navarro-Mateu F, Oladeji BD, … for the WHO World Mental Health Survey Collaborators. (2020). Findings From World Mental Health Surveys of the perceived helpfulness of treatment for patients With major depressive disorder. JAMA Psychiatry, 77(8), 830–841. 10.1001/jamapsychiatry.2020.1107 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Harvey AG, & Gumport NB (2015). Evidence-based psychological treatments for mental disorders: Modifiable barriers to access and possible solutions. Behaviour Research and Therapy, 68, 1–12. 10.1016/j.brat.2015.02.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Harvey L, White FA, Hunt C, & Abbott M (2023). Investigating the efficacy of a dialectical behaviour therapy-based universal intervention on adolescent social and emotional well-being outcomes. Behaviour Research and Therapy, 169, 104408. 10.1016/j.brat.2023.104408 [DOI] [PubMed] [Google Scholar]
  19. Hayes AF, & Matthes J (2009). Computational procedures for probing interactions in OLS and logistic regression: SPSS and SAS implementations. Behavior Research Methods, 41(3), 924–936. 10.3758/BRM.41.3.924 [DOI] [PubMed] [Google Scholar]
  20. Johnson PO, & Neyman J (1936). Tests of certain linear hypotheses and their application to some educational problems. Statistical Research Memoirs, 1, 57–93. [Google Scholar]
  21. Kaiser T, Volkmann C, Volkmann A, Karyotaki E, Cuijpers P, & Brakemeier E-L (2022). Heterogeneity of treatment effects in trials on psychotherapy of depression. Clinical Psychology: Science and Practice, 29(3), 294–303. 10.1037/cps0000079 [DOI] [Google Scholar]
  22. Kaveladze BT, Voelkel JG, Stagnaro MN, Huang M, Smock AE, Sullivan EK, Xu YM, McCall MP, Zapata JP, Ahmed SI, Bhattacharjee A, Georgieva I, Hernandez-Ramos R, Huber KS, Jennings JK, Kirk AC, Knowles RSM, Kornfield R, Lind MN, … Schleider JL (2026). A crowdsourced megastudy of 12 digital single-session interventions for depression in US adults. Nature Human Behaviour. 10.1038/s41562-026-02415-6 [DOI] [PubMed] [Google Scholar]
  23. Kazdin AE, & Rabbitt SM (2013). Novel models for delivering mental health services and reducing the burdens of mental illness. Clinical Psychological Science, 1(2), 170–191. 10.1177/2167702612463566 [DOI] [Google Scholar]
  24. Kessler RC, Van Loo HM, Wardenaar KJ, Bossarte RM, Brenner LA, Ebert DD, De Jonge P, Nierenberg AA, Rosellini AJ, Sampson NA, Schoevers RA, Wilcox MA, & Zaslavsky AM (2017). Using patient self-reports to study heterogeneity of treatment effects in major depressive disorder. Epidemiology and Psychiatric Sciences, 26(1), 22–36. 10.1017/S2045796016000020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Kim J, Ryu N, & Chibanda D (2023). Effectiveness of single-session therapy for adult common mental disorders: A systematic review. BMC Psychology, 11(1), 373. 10.1186/s40359-023-01410-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Kohn R, Saxena S, Levav I, & Saraceno B (2004). The treatment gap in mental health care. Bulletin of the World Health Organization. [PMC free article] [PubMed] [Google Scholar]
  27. Kriegeskorte N, Simmons WK, Bellgowan PSF, & Baker CI (2009). Circular analysis in systems neuroscience: The dangers of double dipping. Nature Neuroscience, 12(5), 535–540. 10.1038/nn.2303 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Kroenke K, Spitzer RL, & Williams JBW (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613. 10.1046/j.1525-1497.2001.016009606.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Krstajic D, Buturovic LJ, Leahy DE, & Thomas S (2014). Cross-validation pitfalls when selecting and assessing regression and classification models. Journal of Cheminformatics, 6(1), 1–15. 10.1186/1758-2946-6-10 [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Kuhn M, Wickham H, & Hvitfeldt E (2025). recipes: Preprocessing and Feature Engineering Steps for Modeling [R package version 1.3.1]. https://recipes.tidymodels.org/authors.html#citation [Google Scholar]
  31. Lambert J, Loades M, Marshall N, Higson-Sweeney N, Chan S, Mahmud A, Pile V, Maity A, Adam H, Sung B, Luximon M, MacLennan K, Berry C, & Chadwick P (2025). Investigating the efficacy of the web-based Common Elements Toolbox (COMET) single-session interventions in improving UK university student well-being: Randomized controlled trial. Journal of Medical Internet Research, 27, e58164. 10.2196/58164 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Lorenzo-Luaces L (2023a). Holy grails, personalized medicine, quantifying our promises, and the public health burden of psychopathology: A reflection on Ahuvia et al. (2023). Journal of Child Psychology and Psychiatry, 65(2), 248–250. 10.1111/jcpp.13914 [DOI] [PubMed] [Google Scholar]
  33. Lorenzo-Luaces L (2023b). Identifying active ingredients in cognitive-behavioral therapies: What if we didn’t? Behaviour Research and Therapy, 168, 104365. 10.1016/j.brat.2023.104365 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Lorenzo-Luaces L, & DeRubeis RJ (2018). Miles to go before we sleep: Advancing the understanding of psychotherapy by modeling complex processes. Cognitive Therapy and Research, 42(2), 212–217. 10.1007/s10608-018-9893-x [DOI] [Google Scholar]
  35. Lorenzo-Luaces L, & Fite RE (2024). Wider, faster, more: Reenvisioning depression treatment research in the United States. Practice Innovations, 9(1), 19–29. 10.1037/pri0000240 [DOI] [Google Scholar]
  36. Lorenzo-Luaces L, & Howard J (2023). Efficacy of an unguided, digital single-session intervention for internalizing symptoms in web-based workers: Randomized controlled trial. Journal of Medical Internet Research, 25, e45411. 10.2196/45411 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Lorenzo-Luaces L, Peipert A, De Jesús-Romero R, Lauren Rutter, & Natalie Rodriguez-Quitana. (2021). Personalized medicine and cognitive behavioral therapies for depression: Small effects, big problems, and bigger data. International Journal of Cognitive Therapy, 14, 59–85. https://doi.org/https://doi-org.proxyiub.uits.iu.edu/10.1007/s41811-020-00094-3 [Google Scholar]
  38. Lorenzo-Luaces L, Wasil A, Kacmarek CN, & DeRubeis R (2024). Race and socioeconomic status as predictors of willingness to use digital mental health interventions or one-on-one psychotherapy: National survey study. JMIR Formative Research, 8, e49780. 10.2196/49780 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Luedtke A, Sadikova E, & Kessler RC (2019). Sample size requirements for multivariate models to predict between-patient differences in best treatments of major depressive disorder. Clinical Psychological Science, 7(3), 445–461. 10.1177/2167702618815466 [DOI] [Google Scholar]
  40. McClelland GH, & Judd CM (1993). Statistical difficulties of detecting interactions and moderator effects. Psychological Bulletin, 114(2), 376–390. 10.1037/0033-2909.114.2.376 [DOI] [PubMed] [Google Scholar]
  41. Meinke C, Hornstein S, Schmidt J, Arolt V, Dannlowski U, Deckert J, Domschke K, Fehm L, Fydrich T, Gerlach AL, Hamm AO, Heinig I, Hoyer J, Kircher T, Koelkebeck K, Lang T, Margraf J, Neudeck P, Pauli P, … Hilbert K (2024). Advancing the personalized advantage index (PAI): A systematic review and application in two large multi-site samples in anxiety disorders. Psychological Medicine, 54(16), 4843–4855. 10.1017/S0033291724003118 [DOI] [PubMed] [Google Scholar]
  42. Moshe I, Terhorst Y, Philippi P, Domhardt M, Cuijpers P, Cristea I, Pulkki-Råback L, Baumeister H, & Sander LB (2021). Digital interventions for the treatment of depression: A meta-analytic review. Psychological Bulletin, 147(8), 749–786. 10.1037/bul0000334 [DOI] [PubMed] [Google Scholar]
  43. Murray CJL, Vos T, Lozano R, Naghavi M, Flaxman AD, Michaud C, Ezzati M, Shibuya K, Salomon JA, Abdalla S, Aboyans V, Abraham J, Ackerman I, Aggarwal R, Ahn SY, Ali MK, AlMazroa MA, Alvarado M, Anderson HR, … Lopez AD (2012). Disability-adjusted life years (DALYs) for 291 diseases and injuries in 21 regions, 1990–2010: A systematic analysis for the Global Burden of Disease Study 2010. The Lancet, 380(9859), 2197–2223. 10.1016/S0140-6736(12)61689-4 [DOI] [PubMed] [Google Scholar]
  44. National Institute of Mental Health. (2022). Mental Illness. https://www.nimh.nih.gov/health/statistics/mental-illness
  45. Nye A, Delgadillo J, & Barkham M (2023). Efficacy of personalized psychological interventions: A systematic review and meta-analysis. Journal of Consulting and Clinical Psychology, 91(7), 389–397. 10.1037/ccp0000820 [DOI] [PubMed] [Google Scholar]
  46. Osborn TL, Rodriguez M, Wasil AR, Venturo-Conerly KE, Gan J, Alemu RG, Roe E, Arango GS,Otieno BH, Wasanga CM, Shingleton R, & Weisz JR (2020). Single-session digital intervention for adolescent depression, anxiety, and well-being: Outcomes of a randomized controlled trial with Kenyan adolescents. Journal of Consulting and Clinical Psychology, 88(7), 657–668. 10.1037/ccp0000505 [DOI] [PubMed] [Google Scholar]
  47. Pargent F, Schoedel R, & Stachl C (2023). Best practices in supervised machine learning: A tutorial for psychologists. Advances in Methods and Practices in Psychological Science, 6(3), 1–35. https://doi.org/ 10.1177/25152459231162559 [DOI] [Google Scholar]
  48. Patel V, Chisholm D, Parikh R, Charlson FJ, Degenhardt L, Dua T, Ferrari AJ, Hyman S, Laxminarayan R, Levin C, Lund C, Mora MEM, Petersen I, Scott J, Shidhaye R, Vijayakumar L, Thornicroft G, & Whiteford H (2016). Addressing the burden of mental, neurological, and substance use disorders: Key messages from Disease Control Priorities, 3rd edition. The Lancet, 387(10028), 1672–1685. 10.1016/S0140-6736(15)00390-6 [DOI] [PubMed] [Google Scholar]
  49. Paul GL (1967). Strategy of outcome research in psychotherapy. Journal of Consulting Psychology, 31(2), 109–118. 10.1037/h0024436 [DOI] [PubMed] [Google Scholar]
  50. Peipert A, & Lorenzo-Luaces L (2025). Is there a treatment for the Turker blues? A fully remote nationwide randomized controlled trial of a digital intervention for depression in adult online workers. Journal of Consulting and Clinical Psychology, 93(11), 719–734. 10.1037/ccp0000982 [DOI] [PubMed] [Google Scholar]
  51. Pincus T, Miles C, Froud R, Underwood M, Carnes D, & Taylor SJ (2011). Methodological criteria for the assessment of moderators in systematic reviews of randomised controlled trials: A consensus study. BMC Medical Research Methodology, 11(1), 14. 10.1186/1471-2288-11-14 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Ramos G, Hernandez-Ramos R, Taylor M, & Schueller SM (2024). State of the science: Using digital mental health interventions to extend the impact of psychological services. Behavior Therapy, 55(6), 1364–1379. 10.1016/j.beth.2024.04.004 [DOI] [PubMed] [Google Scholar]
  53. Rush AJ, Trivedi MH, Wisniewski SR, Nierenberg AA, Stewart JW, Warden D, Niederehe G, Thase ME, Lavori PW, Lebowitz BD, McGrath PJ, Rosenbaum JF, Sackeim HA, Kupfer DJ, Luther J, & Fava M (2006). Acute and longer-term outcomes in depressed outpatients requiring one or several treatment steps: A STAR*D report. American Journal of Psychiatry, 163(11), 1905–1917. 10.1176/ajp.2006.163.11.1905 [DOI] [PubMed] [Google Scholar]
  54. Salditt M, Eckes T, & Nestler S (2024). A tutorial introduction to heterogeneous treatment effect estimation with meta-learners. Administration and Policy in Mental Health and Mental Health Services Research, 51(5), 650–673. 10.1007/s10488-023-01303-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Schleider JL, & Beidas RS (2022). Harnessing the single-session intervention approach to promote scalable implementation of evidence-based practices in healthcare. Frontiers in Health Services, 2, 997406. 10.3389/frhs.2022.997406 [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Schleider JL, & Weisz JR (2017). Little treatments, promising effects? Meta-analysis of single-session interventions for youth psychiatric problems. Journal of the American Academy of Child & Adolescent Psychiatry, 56(2), 107–115. 10.1016/j.jaac.2016.11.007 [DOI] [PubMed] [Google Scholar]
  57. Schleider JL, Zapata JP, Rapoport A, Wescott A, Ghosh A, Kaveladze B, Szkody E, & Ahuvia IL (2025). Single-session interventions for mental health problems and service engagement: Umbrella review of systematic reviews and meta-analyses. Annual Review of Clinical Psychology, 21(1), 279–303. 10.1146/annurev-clinpsy-081423-025033 [DOI] [PubMed] [Google Scholar]
  58. Schwartz B, Cohen ZD, Rubel JA, Zimmermann D, Wittmann WW, & Lutz W (2021). Personalized treatment selection in routine care: Integrating machine learning and statistical algorithms to recommend cognitive behavioral or psychodynamic therapy. Psychotherapy Research, 31(1), 33–51. 10.1080/10503307.2020.1769219 [DOI] [PubMed] [Google Scholar]
  59. Shin C, Lee S-H, Han K-M, Yoon H-K, & Han C (2019). Comparison of the usefulness of the PHQ-8 and PHQ-9 for screening for major depressive disorder: Analysis of psychiatric outpatient data. Psychiatry Investigation, 16(4), 300–305. 10.30773/pi.2019.02.01 [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Simon GE, & Perlis RH (2010). Personalized medicine for depression: Can we match patients with treatments? American Journal of Psychiatry, 167(December), 1445–1455. 10.1176/appi.ajp.2010.09111680 [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Simon GE, Shortreed SM, Rossom RC, Beck A, Clarke GN, Whiteside U, Richards JE, Penfold RB, Boggs JM, & Smith J (2022). Effect of Offering Care Management or Online Dialectical Behavior Therapy Skills Training vs Usual Care on Self-harm Among Adult Outpatients With Suicidal Ideation: A Randomized Clinical Trial. JAMA, 327(7), 630. 10.1001/jama.2022.0423 [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Starvaggi I, & Lorenzo-Luaces L (2025). Psychotherapy access barriers and interest in digital mental health interventions among adults with treatment needs: Survey study. JMIR Mental Health, 12, e65356. 10.2196/65356 [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Stekhoven DJ, & Bühlmann P (2012). MissForest—Non-parametric missing value imputation for mixed-type data. Bioinformatics, 28(1), 112–118. 10.1093/bioinformatics/btr597 [DOI] [PubMed] [Google Scholar]
  64. Szkody E, Chang Y-W, & Schleider JL (2023). Serving the underserved? Uptake, effectiveness, and acceptability of digital SSIs for rural American adolescents. Journal of Clinical Child & Adolescent Psychology, 1–14. 10.1080/15374416.2023.2272935 [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Terhorst Y, Kaiser T, Brakemeier E-L, Moshe I, Philippi P, Cuijpers P, Baumeister H, & Sander LB (2024). Heterogeneity of treatment effects in Internet- and mobile-based interventions for depression: A systematic review and meta-analysis. JAMA Network Open, 7(7), e2423241. 10.1001/jamanetworkopen.2024.23241 [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. US Census Bureau. (2024). U.S. Census Bureau QuickFacts: United States. Census.Gov. https://www.census.gov/quickfacts/fact/table/US/PST045224 [Google Scholar]
  67. Van Bronswijk SC, Howard J, & Lorenzo-Luaces L (2024). Data-driven personalized medicine approaches to cognitive-behavioral therapy allocation in a large sample: A reanalysis of the ENRICHED study. Journal of Affective Disorders, 356, 115–121. 10.1016/j.jad.2024.04.015 [DOI] [PubMed] [Google Scholar]
  68. Wang PS, Lane M, Olfson M, Pincus HA, Wells KB, & Kessler RC (2005). Twelve-month use of mental health services in the United States: Results from the National Comorbidity Survey replication. Archives of General Psychiatry, 62(6), 629–640. 10.1001/archpsyc.62.6.629 [DOI] [PubMed] [Google Scholar]
  69. Wasil AR, Park SJ, Gillespie S, Shingleton R, Shinde S, Natu S, Weisz JR, Hollon SD, & DeRubeis RJ (2020). Harnessing single-session interventions to improve adolescent mental health and well-being in India: Development, adaptation, and pilot testing of online single-session interventions in Indian secondary schools. Asian Journal of Psychiatry, 50, 101980. 10.1016/j.ajp.2020.101980 [DOI] [PubMed] [Google Scholar]
  70. Wasil AR, Taylor ME, Franzen RE, Steinberg JS, & DeRubeis RJ (2021). Promoting graduate student mental health during COVID-19: Acceptability, feasibility, and perceived utility of an online single-session intervention. Frontiers in Psychology, 12, 569785. 10.3389/fpsyg.2021.569785 [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Whiteford HA, Degenhardt L, Rehm J, Baxter AJ, Ferrari AJ, Erskine HE, Charlson FJ, Norman RE, Flaxman AD, Johns N, Burstein R, Murray CJ, & Vos T (2013). Global burden of disease attributable to mental and substance use disorders: Findings from the Global Burden of Disease Study 2010. The Lancet, 382(9904), 1575–1586. 10.1016/S0140-6736(13)61611-6 [DOI] [PubMed] [Google Scholar]
  72. Wilhelm M, Bauer S, Feldhege J, Wolf M, & Moessner M (2024). Alleviating the burden of depression: A simulation study on the impact of mental health services. Epidemiology and Psychiatric Sciences, 33, e19. 10.1017/S204579602400012X [DOI] [PMC free article] [PubMed] [Google Scholar]
  73. Wu Y, Levis B, Riehm KE, Saadat N, Levis AW, Azar M, Rice DB, Boruff J, Cuijpers P, Gilbody S, Ioannidis JPA, Kloda LA, McMillan D, Patten SB, Shrier I, Ziegelstein RC, Akena DH, Arroll B, Ayalon L, … Thombs BD (2020). Equivalency of the diagnostic accuracy of the PHQ-8 and PHQ-9: A systematic review and individual participant data meta-analysis. Psychological Medicine, 50(8), 1368–1380. 10.1017/S0033291719001314 [DOI] [PMC free article] [PubMed] [Google Scholar]
  74. Zainal NH, Benjet C, Albor Y, Nuñez-Delgado M, Zambrano-Cruz R, Contreras-Ibáñez CC, Cudris-Torres L, De La Peña FR, González N, Guerrero-López JB, Gutierrez-Garcia RA, Jiménez-Peréz AL, Medina-Mora ME, Patiño P, Cuijpers P, Gildea SM, Kazdin AE, Kennedy CJ, Luedtke A, … Kessler RC (2025). Statistical methods to adjust for the effects on intervention compliance in randomized clinical trials where precision treatment rules are being developed. International Journal of Methods in Psychiatric Research, 34(1), e70005. 10.1002/mpr.70005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Zou H, & Hastie T (2005). Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society Series B: Statistical Methodology, 67(2), 301–320. 10.1111/j.1467-9868.2005.00503.x [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

RESOURCES