ABSTRACT
Placebo effect represents a serious confounder for the assessment of treatment effect to the extent that it has become increasingly difficult to develop antidepressant medications appropriate for outperforming placebo. Treatment effect in randomized, placebo‐controlled trials, is usually estimated by the mean baseline adjusted difference of treatment response in active and placebo arms and is function of treatment‐specific and non‐specific effects. The non‐specific treatment effect varies subject by subject conditional to the individual propensity to respond to placebo. This effect is not estimable at an individual level using the conventional parallel‐group study design, since each subject enrolled in the trial is assigned to receive either active treatment or placebo, but not both. The objective of this study was to conduct a comparative analysis of the machine learning methodologies to estimate the individual probability of a non‐specific treatment effect. The estimated probability is expected to support novel methodological approaches for better controlling effect of excessively high placebo response. At this purpose, six machine learning methodologies (gradient boosting machine, lasso regression, logistic regression, support vector machines, k‐nearest neighbors, and random forests) were compared to the multilayer perceptrons artificial neural network (ANN) methodology for predicting the probability of individual non‐specific treatment response. ANN achieved the highest overall accuracy among all methods tested. A fivefold cross‐validation was used to assess performances and risks of overfitting of the ANN model. The analysis conducted without subjects with non‐specific effect indicated a significant increase of signal detection with significant increase in effect size.
Keywords: artificial intelligence, clinical trials, machine learning, non‐specific treatment response, placebo effect
Summary.
- What is the current knowledge on the topic?
-
○The non‐specific treatment effect (usually referred as placebo effect) represents a serious concern for the assessment of treatment effect in the development of antidepressant medications.
-
○
- What question did this study address?
-
○This study evaluate different machine learning methodologies for estimating the individual probability of shoving a non‐specific treatment response.
-
○
- What does this study add to our knowledge?
-
○This study identified the individual HAMD‐17 item changes from screening to baseline as predictors of non‐specific treatment response using an ANN model.
-
○
- How might this change clinical pharmacology or translational science?
-
○The estimated probability of non‐specific treatment response is expected to be instrumental for the development of novel methodological approaches for better controlling the effect of excessively high placebo response supporting a more effective development process of novel treatments for MDD.
-
○
1. Introduction
The non‐specific treatment effect, usually referred as placebo effect, represents a serious issue for the assessment of treatment effect in randomized, placebo‐controlled clinical trials (RCTs) conducted for the development of medications in major depressive disorders (MDD). This uncontrolled effect is so relevant that it has become increasingly difficult and expensive to develop antidepressant medications able to outperform placebo [1, 2].
The increasing level of placebo effect, preventing the assessment of potential benefits of new medications, has been also associated with a diminished ability to detect a signal of treatment effect leading to controversies about the efficacy of drug treatments [3, 4].
Randomized placebo‐controlled clinical trials are the gold standard methodology for testing safety and efficacy of new medications in MDD. In absence of any reference medications, the regulatory authorities currently require the proof that new medications perform better than placebo. In a meta‐analysis of monotherapy conducted on the FDA‐approved antidepressant trials in MDD, a higher placebo response rate has been shown to correlate either with a lower risk ratio of responding to antidepressant versus placebo (p < 0.001) or with higher antidepressant response rates (p < 0.001) [5]. This finding implies that the ability to detect a therapeutic effect depends on how well the placebo response is managed.
The unexpected high placebo response associated with modest antidepressant treatment response rates (about 50%–60%) is currently preventing the ability of detecting statistically significant drug effect in RCTs [6]. This finding was confirmed in a meta‐analysis, conducted on the US Food and Drug Administration database including 12 approved antidepressant drugs, 74 RCTs, and 12,564 patients. This analysis indicated that 49% of the clinical trials failed [7].
Non‐specific factors (such as expectation, natural course of depressive illness, regression to the mean, inappropriate subject selection, inadvertent supportive therapy) that are unrelated to the actual experimental medication have been advocated as potential factors associated with the improvements recorded in clinical trials [8].
The expectations of improvement is usually characterized by the adoption of an optimistic response criterion when deciding if the medication under evaluation is effective. This disposition is primarily generated by the interactions between the patients included in the trial and the investigator at the recruitment center. Confounding motivations, often implicit, may result in setting high expectations about the treatment outcome, with investigators desiring to heal the patient and patients wanting to please the investigators, leading to high improvement in the expectation of improvement.
Furthermore, the placebo response in antidepressant trials has been shown to grown by 7% per decade in the previous 30 years and that current estimates placed the average placebo response rate between 35% and 40% [9].
Various strategies have been attempted, with mixed degrees of success, to address a higher than anticipated placebo response. These strategies included increase of sample size as an accommodation to yield greater power, innovative study designs (e.g., staggered, blinded placebo run‐in phases or sequential parallel design methods—SPCD), enhanced inter‐rater reliability programs, surveillance of in‐study data to identify measurement error, site‐independent subject validation to minimize site‐biases, and enhanced patient education to minimize expectancy effects [8, 10]. More recently, novel analytic strategies based upon predictive statistical models have also been examined as a way to enhance signal detection [11]. This methodology, called band‐pass filter analysis, is based on signal detection theory, directly addresses non‐plausible placebo response rates at specific clinical trial sites by removing completely their data (of patients on either active treatment or placebo) from the analysis. To overcome the limited performances of these methodologies, new artificial intelligence (AI) based tools have been proposed for designing, conducting, and analyzing RCTs and for controlling and mitigating the increasing confounding effect of placebo response.
The observed treatment response in MDD trials is the resultant of a treatment‐specific and a non‐specific treatment response (NSTR) [12, 13, 14]. NSTR is considered as experimental noise and it is commonly referred as placebo effect. The NSTR is mainly due to expectations for improvement and to the effect of receiving personal care from clinicians. Otherwise, the “specific effect” is usually associated with the clinical benefit only driven by the properties of the active drug used in RCTs. The challenge of identifying specific and non‐specific responses is complex as these effects are not directly measured. When a subject responds to a treatment, it is difficult to determine whether the response is due to specific pharmacological action of the drug or to non‐specific effect. Furthermore, a subject may exhibit both specific and non‐specific effects and the degree of each can vary from subject to subject and from study to study. For this reason, there is an unmet need for developing novel methodologies to estimate, account and control the NSTR in the analysis of RCTs.
The assessment of treatment effect (TE) in RCTs in MDD is currently done by comparing the mean response of an active treatment to the mean response of the placebo arm in a parallel‐group study design. The assessment of the individual patient response to new treatments is not possible, since each subject enrolled in the trial is assigned to receive either new treatments or placebo, but not both. Therefore, the TE is not estimable at an individual level. The treatment effect is usually estimated as a single mean score (usually the baseline adjusted difference of the clinical score used to assess disease severity between active treatments and placebo estimate at study end) that is assumed to affect each treated patient; despite the fact that the contribution of each subject to the estimate of this mean score will vary subject by subject according to the unknown individual characteristics.
The baseline individual probability of experiencing a NSTR (also referred as the propensity to be non‐treatment‐specific responder) is considered as a major confounding factor for the assessment of TE. The larger is the baseline propensity to respond to non‐specific TE, the lower is the chance to detect any treatment‐specific effect. In the current clinical trial setting no methodologies are currently available for evaluating the comparability of the treatment arm with respect to the potential baseline unbalance on NSTR and to provide a NSTR‐adjusted estimate of the TE. On this basis, the key issues for a reliable estimate of TE is to find a measure of the individual probability of non‐specific response (prob‐NSTR) and then to integrate this estimation in appropriate analysis tools suitable for providing a prob‐NSTR‐adjusted estimate of TE.
AI and, in particular, machine learning (ML) methodologies have been shown to represent a powerful tool for improving the efficiency of drug development and, in particular, for describing and accounting for the effect of variable placebo response in the development of new CNS diseases [15, 16].
For example, it has been shown that ML models (support vector machine, k‐nearest neighbors, lasso regression, random forest, and gradient boosting machine) applied to early response viloxazine‐extended release after 2 weeks of treatment in pediatric patients with ADHD can predict efficacy outcome at Week 6 [17]. A random forest ML model has been also shown to represent a reliable tool for identifying pre‐treatment characteristics of patients most likely to respond to drug vs. placebo in MDD [18]. In addition, ML models (including support vector machine, gradient boosting, random forest and adaptive boosting) have been shown to provide an accurate prediction of the response to a panel of antidepressants and to provide a tool for optimizing treatment selection [19].
In this study, we present the comparison of different ML methodologies for estimating the individual probability of experiencing treatment‐specific‐ and non‐specific effects (prob‐NSTR) using information collected pre‐randomization (i.e., at screening and randomization time). Machine learning is the branch of AI that focuses on developing models and algorithms that let computers learn from data without being explicitly programmed [20]. Among the different algorithms, we are focusing on the classification algorithms, that is, the methods where the model tries to predict the correct classification of a subject by comparing the model predictions to observations. In this context, the following ML methodologies have been compared:
Logistic regression: models the probability of a binary outcome using a logistic function.
Support Vector Machine (SVM): finds the hyperplane that best separates different classes by maximizing the margin between them.
Random Forest: creates many individual decision trees using a random selection of data points and features.
k‐nearest neighbors: uses proximity to make classifications or predictions about the grouping of individual data points.
Gradient boosting machine: builds simpler prediction models sequentially where each model tries to predict the error left over by the previous model and then to combine the predictions of multiple weak learners to create a single, more accurate strong learner.
Lasso regression: enhances the linear regression methodology by making use of a regularization process. Regularization introduces a penalty for more complex models, effectively reducing their complexity and preventing overfitting.
ANN: multilayer perceptrons artificial neural network makes decisions in a manner similar to the human brain, by using processes that mimic the way neurons work. Every neural network consists of layers of nodes (artificial neurons) an input layer one or more hidden layers, and an output layer.
Prob‐NSTR is a complex and multifactorial score driven by a multitude of factors mediated by expectations, associative learning processes, and the quality of the patient‐physician interaction [21]. These data, not usually collected in a RCT, represent a study‐specific hidden layer of individual information affecting the individual prob‐NSTR. Our assumption is that the changes of the individual 17‐item score of the Hamilton depression rating scale (HAMD‐17), a clinician‐rated scale for the assessment of depression severity, represent a reliable description of different dimension of the depression suitable for estimating the individual prob‐NSRT. This probability is expected to be instrumental for the development of novel methodological approaches (such as the propensity weighting methodology) and for better controlling the effect of excessively high placebo response [22]. The scope of the proposed propensity weighting methodology is not to address the controversy issue of the potential additivity between placebo and active. The propensity weighted methodology assumes that the non‐specific treatment effect may not be necessary equal in the subjects enrolled in treatment arms, that the non‐specific treatment response varies individual to individual, and that the prob‐NSTR can be estimated using an AI model applied to the pre‐randomization data.
2. Methods
2.1. Data
The data from two clinical trials were used as a case study. The first one, study 810, was a randomized, double‐blind, parallel‐group, placebo‐controlled study evaluating efficacy and safety of paroxetine controlled release (CR) 12.5 mg/day (N = 156), paroxetine CR 25 mg/day (N = 154) versus placebo (N = 149) in MDD patients, mean age 39 years (± 0.5 StdErr). The primary efficacy endpoint was the change from baseline at Week 8 of the HAMD‐17 total score [23]. The second study, study 874 (NCT00067444), was a randomized, double‐blind, parallel‐group, placebo‐controlled fixed‐dose study evaluating the effect of paroxetine CR 12.5 mg/day (N = 168), paroxetine CR 25 mg/day (N = 177) versus placebo (N = 180) in MDD patients, mean age 67 years (±0.3 StdErr). The primary efficacy endpoint was the change from baseline at Week 8 in the HAMD‐17 scale [24].
The demographic data and total HAMD‐17 scores at screening and baseline for the two trials are presented in Table 1.
TABLE 1.
Demographic data and total HAMD‐17 scores at screening and baseline for the 810 and 874 trials.
| Study | Treatment | N | Variable | Mean | Std Error | CV% | L_95% | U_95% |
|---|---|---|---|---|---|---|---|---|
| 810 | 12.5 mg | 156 | Screening HAMD‐17 | 23.87 | 0.26 | 13.77 | 23.35 | 24.39 |
| Baseline HAMD‐17 | 23.13 | 0.23 | 12.5 | 22.67 | 23.59 | |||
| Age (y) | 38.37 | 0.98 | 31.78 | 36.44 | 40.3 | |||
| Weigh (kg) | 83.55 | 1.75 | 26.1 | 80.1 | 87 | |||
| 25 mg | 154 | Screening HAMD‐17 | 23.91 | 0.25 | 12.89 | 23.42 | 24.4 | |
| Baseline HAMD‐17 | 23.51 | 0.26 | 13.94 | 22.99 | 24.03 | |||
| Age (y) | 39.28 | 0.88 | 27.74 | 37.54 | 41.01 | |||
| Weigh (kg) | 83.98 | 1.69 | 24.98 | 80.64 | 87.32 | |||
| Placebo | 150 | Screening HAMD‐17 | 24.7 | 0.3 | 14.81 | 24.11 | 25.29 | |
| Baseline HAMD‐17 | 23.81 | 0.26 | 13.5 | 23.29 | 24.33 | |||
| Age (y) | 38.51 | 0.97 | 30.85 | 36.59 | 40.42 | |||
| Weigh (kg) | 86.1 | 2.09 | 29.66 | 81.98 | 90.22 | |||
| 874 | 12.5 mg | 168 | Screening HAMD‐17 | 23.26 | 0.32 | 17.86 | 22.62 | 23.89 |
| Baseline HAMD‐17 | 22.49 | 0.28 | 15.87 | 21.95 | 23.04 | |||
| Age (y) | 67.13 | 0.47 | 9.16 | 66.19 | 68.06 | |||
| Weigh (kg) | 79.64 | 1.28 | 20.7 | 77.12 | 82.16 | |||
| 25 mg | 177 | Screening HAMD‐17 | 23.22 | 0.31 | 17.67 | 22.61 | 23.83 | |
| Baseline HAMD‐17 | 23.08 | 0.3 | 17.17 | 22.5 | 23.67 | |||
| Age (y) | 67.03 | 0.49 | 9.76 | 66.06 | 68 | |||
| Weigh (kg) | 81.46 | 1.32 | 21.57 | 78.85 | 84.07 | |||
| Placebo | 180 | Screening HAMD‐17 | 23.08 | 0.28 | 16.12 | 22.54 | 23.63 | |
| Baseline HAMD‐17 | 22.71 | 0.3 | 17.66 | 22.12 | 23.3 | |||
| Age (y) | 67.97 | 0.5 | 9.91 | 66.98 | 68.96 | |||
| Weigh (kg) | 82.41 | 1.55 | 25.16 | 79.36 | 85.46 |
Abbreviations: CV% = coefficient of variation; L_95% & U_95% = lower and upper 95% confidence interval; Std Error = standard error.
Many potential predictors of the prob‐NSTR evaluated at screening and baseline can be considered such as demographic data, habits and quality of life, or disease‐related information, etc. For simplicity, we limited our exploration to the 17 individual items of the HAMD‐17 scale as these items are assumed to capture specific and independent symptoms of depression: (1) depressed mood, (2) feelings of guilt, (3) suicide, (4) insomnia initial, (5) insomnia middle, (6) insomnia delayed, (7) work and interests, (8) retardation, (9) Agitation, (10) anxiety—psychic, (11) anxiety—somatic, (12) somatic symptoms—gastrointestinal, (13) somatic symptoms—general, (14) genital symptoms, (15) hypochondriasis, (16) weight loss, (17) Insight [25].
2.2. Non‐specific Treatment Response
The non‐specific treatment response (i.e., the clinically relevant change from baseline at Week 8 of the placebo response) was defined using the clinician global impression‐improvement (CGI‐I) scale as this score represents a patient‐independent assessment done by the clinician on the patient's global disease severity prior to and after initiating the treatment. A CGI‐I between minimally improved (=3) and much improved (=2) was considered as an indicator of a clinically relevant response. This CGI‐I score was associated to a percent reduction from baseline at Week 8 in the HAMD‐17 total score of 41% using the equipercentile linking method [26, 27, 28]. A binary score (0 or 1) was then associated to each subject in the placebo arm as an indicator of the absence or presence of a clinically relevant signal of response when the change from baseline of HAMD‐17 total score at Week 8 was greater than 41%.
2.3. ML Models
The six previously described ML models were selected as they have been shown to represent an effective tool for dealing with multiple predictor variables in many different disease areas including CNS diseases [15, 29]. The performance of these models were compared to each other and to the ANN methodology [30]. Among the different methods used for implementing ANN models, the multilayer perceptron (MLP) artificial neural network was retained as this method has been shown to have superior and robust classification performance [30]. MLP consists of multiple layers of neurons that use nonlinear activation functions to learn complex data patterns. The implementation of a MLP ANN model required the definition of two hyperparameters that control the topology of the network: the number of hidden layers and the number of nodes in each layer. The definition of these parameters were done using a grid search analysis. The grid search procedure begins with the specification of possible value for each hyperparameter. In the present case, it was assumed to have three hidden layers and 20 possible values for the nodes in each layer. Then, the grid search evaluated the predictive performance associated with each combination of nodes in each layer. In case of the selection of only one node in a given layer, the procedure was restarted with a reduced number of layers. The optimality criteria used to assess the performance of each tested model was the AUC under the receiver operating characteristic (ROC) curve.
The ML analysis was conducted using the ‘caret’ library available in the statistical software R and the MLP ANN analysis was conducted using the ‘neuralnet’ library available in the statistical software R [31].
The ML models were expected to predict the probability of a clinically relevant signal of placebo response at Week 8 (prob‐NSTR) using the changes from screening to baseline of the 17 individual items of the HAMD‐17 scale. If the estimated prob‐NSTR was ≤ 0.5 then a binary score of 0 was associated to this subject, otherwise if prob‐NSTR > 0.5, a binary score of 1 was associated to the subject. The comparison between the model predicted and the observed binary score in the training dataset was used to qualify the predictive performances of the ML model.
The comparison of observed and model predicted binary responses was summarized using the following criteria:
TN or true negative: sum of observed negative responses with negative ML predictions.
FP or false positive: sum of observed negative responses with positive ML predictions.
TP or true positive: sum of observed positive responses with positive ML predictions.
FN or false negative: sum of observed positive responses with negative ML predictions.
These scores were used to evaluate the performances of the different ML models using the following criteria:
Sensitivity = , proportion of true positives that were correctly identified by the model.
Specificity = , proportion of true negatives as a fraction of total negatives
Accuracy = , proportion of correct predictions as a fraction of total predictions
Precision = , proportion of positive predictions that were correct.
ROC AUC = The ROC curve is the plot of the true positive rate (Sensitivity) against the false positive rate (1‐Specificity). The ROC AUC was computed with the 95% confidence intervals estimated with a 2000 replicates bootstrap analysis. ROC AUC is a measure for assessing the accuracy of a ML method which corresponds to the total proportion of correctly classified observations.
The schematic of the workflow used to estimate the individual prob‐NSTR in each study is presented in Figure 1. The data used in the analysis were organized in the following datasets:
Placebo dataset: including the changes from screening to baseline of the 17 individual items of the HAMD‐17 scale and the individual binary score indicating the absence or presence of a clinically relevant signal of placebo response at Week 8. The placebo dataset was then divided into a training set (accounting for data from 80% of patients) and a test set (data from the remaining 20% of patients). The training data, collected at screening and randomization, were used as predictors of a clinically relevant signal of placebo response at Week 8 using ML. The test data were used to assess the predictive performance of the ML model by comparing model predictions and observed data.
Analysis dataset: including the changes from screening to baseline of the 17 individual items of the HAMD‐17 scale of all subjects enrolled in the study (i.e., in the placebo, Dose 1, and Dose 2 arms). The ML model developed using the placebo dataset was applied to the changes from screening to baseline of the 17 individual items of the HAMD‐17 scale to predict the probability of NSTR of all subjects in the trial.
FIGURE 1.

ML modeling workflow.
The treatment effect (defined as the baseline corrected difference of the total HAMD‐17 score between active treatment and placebo at Week 8) and Cohen's d effect size were estimated in the total population and in the population without subjects with prob‐NSTR > 0.5 to explore the impact of the subjects with high prob‐NSTR on the study outcomes.
The final ANN models developed for study 810 and 874 were trained using a stratified grouped fivefold cross‐validation methodology to assess the performances and the final model validation. The main purpose of cross‐validation was to prevent overfitting, which occurs when a model is trained too well on the training data and performs poorly on new, unseen data [32]. The cross‐validation was also used to evaluate the robustness of the ANN model by computing the specificity, sensitivity, accuracy, precision/recall, ROC AUC, and the F1 score. The F1 score was selected as this score provides the mean of both accuracy and recall, offering a comprehensive evaluation of the model performances obtained through taking both false positives and false negatives into consideration, hence giving a balanced view [33]. The highest possible value of an F1 score is 1.0, indicating perfect precision and recall, and the lowest possible value is 0, if precision and recall are zero.
Two additional analyses were conducted to evaluate if and on which condition the ML model developed using data of one study can be generalized and prospectively used to predict the prob‐NSTR in other trials. In the first analysis, the final ANN model developed with the 810 study data was used to predict the prob‐NSTR of the 874 study data. In the second one the final ANN model developed with the 874 study data was used to predict the prob‐NSTR of the 810 study data. The comparison of the predictive performances of the two models was done using the 95% confidence intervals of the ROC AUCs.
3. Results
The observed longitudinal changes from baseline of the HAMD‐17 clinical scores in the placebo arms of the 810 and 874 studies are presented in Figure 2. The comparison of the time course of the placebo total HAMD‐17 score in the two studies showed similar profiles with a significantly higher response rate (~38%) in study 810.
FIGURE 2.

Longitudinal changes from baseline of the HAMD‐17 clinical scores in the placebo arms of the 810 and 874 studies.
The comparison of the specificity, sensitivity, accuracy, precision/recall and ROC AUC of the six machine learning methodologies (gradient boosting machine, lasso regression, logistic regression, support vector machines, k‐nearest neighbors, and random forests) used to predict the prob‐NSTR for the 810 and 874 studies together with the outcomes of the ANN methodology are presented in Figure 3. The performance of each of the ML methods was assessed using the same training and test datasets.
FIGURE 3.

Comparison of the estimated specificity, sensitivity, accuracy, precision/recall and ROC AUC (with the 95% confidence intervals) of the seven machine learning methodologies used to predict the prob‐NSTR for the 810 and 874 studies. The horizontal lines represent the 95% CI and the black dots represent the mean ROC AUC values.
The grid search in the ANN analysis indicated that the optimal number of layers was 3 and the optimal number of nodes per layer was 6, 7, and 13 for the 874 study, and 7, 11, and 5 for the 810 study, respectively. The random forest method performed better among the non‐ANN ML models either for the 810 or the 874 study. In any case, the ANN models showed better performance according to the value of the ROC AUC with the associated 95% confidence intervals and also according to the estimated scores for the sensitivity, specificity, accuracy, and precision as illustrated in Figure 3. Therefore, the ANN methodology was retained as the reference methodology for the assessment of the non‐specific treatment response.
The results of the fivefold cross‐validation of the final ANN models for the 810 and 874 studies are presented in Table 2. The comparison of the estimated mean F1 score together with the standard errors and the coefficient of variations of the six parameters estimated in the fivefold cross‐validation indicated good performances of the ANN models showing consistent results in the two studies evaluated.
TABLE 2.
Studies 810 & 874: Fivefold cross‐validation results of the final ANN models.
| Study | Variable | Mean | Std Error | CV% | L_95% | U_95% |
|---|---|---|---|---|---|---|
| 810 | Specificity | 0.90 | 0.04 | 10.59 | 0.78 | 1.02 |
| Sensitivity | 0.82 | 0.03 | 9.19 | 0.72 | 0.91 | |
| Accuracy | 0.84 | 0.02 | 4.86 | 0.79 | 0.89 | |
| Precision/recall | 0.91 | 0.04 | 9.24 | 0.81 | 1.02 | |
| ROC AUC | 0.89 | 0.01 | 2.48 | 0.86 | 0.91 | |
| F1 | 0.86 | 0.02 | 3.96 | 0.82 | 0.90 | |
| 874 | Specificity | 0.91 | 0.02 | 5.33 | 0.85 | 0.97 |
| Sensitivity | 0.84 | 0.05 | 12.38 | 0.71 | 0.97 | |
| Accuracy | 0.88 | 0.01 | 2.92 | 0.85 | 0.91 | |
| Precision/recall | 0.87 | 0.03 | 8.25 | 0.78 | 0.96 | |
| ROC AUC | 0.91 | 0.01 | 2.35 | 0.89 | 0.94 | |
| F1 | 0.85 | 0.02 | 5.71 | 0.79 | 0.91 |
Abbreviations: CV% = coefficient of variation; L_95% & U_95% = lower and upper 95% confidence interval; Std Error = standard error.
The ANN models predicted the probability of NSTR based on the definition of clinically relevant response as described in the methods. Subjects were classified according to the predicted probabilities of NSTR (> 0.5 = subject with non‐specific response, ≤ 0.5 = subject without non‐specific response). The comparison of the distribution of the subjects with specific and non‐specific responses in the two trials indicated that ~60% of the subjects in the 874 trial and only ~40% of the subjects in the 810 trial presented a non‐specific response (Figure 4).
FIGURE 4.

Distribution of the subjects with specific and non‐specific responses in the 810 and 874 trials.
The final neural network layouts for the ANN analyses of study 810 and 874 are presented in Figures S1 and S3. In these plots, the first column represents the list of the selected variables considered as predictors of the prob‐NSRT, the second column represents the combined items characterizing the first layer, the third column represents the combined items defining the second layer, and the fourth column represents the items defining the final layer. The lines connecting the nodes are color‐coded by sign (black increasing and gray decreasing effects). Except for the input nodes, each node is expected to emulate the function of a neuron using a nonlinear activation function. The size of the connecting lines (i.e., the weight of the connection) is analogous to the coefficient in a standard regression analysis, this size determines the relative importance of information associated with the connected variables processed in the network. The importance of each selected variable in the prediction of the individual prob‐NSRT was estimated as the sum of the product of raw input‐hidden and hidden‐output connection weights according to a methodology specifically developed for the neural networks with multiple hidden layers [34]. The relative importance of the selected variables for the prediction of prob‐NSRT in study 810 and 874 are presented in Figures S2 and S4. The distribution of the relative importance of the 17 predictive items in one study was different to the distribution of the 17 items in the other study. This finding indicates that a highly informative item in a study may be non‐informative for predicting the individual prob‐NSRT in another study.
Among of the individual HAMD‐17 items, the relative importance analysis identified the changes from screening to randomization of item 1 and 7 (depressed mood and work and interests) as the items with the larger importance in the assessment of prob‐NSTR in the two studies. This finding was in good agreement with the results of different individual item response analyses conducted on the HAMD‐17 items showing the relevance of items 1 and 7 in the prediction of the depressive severity [35, 36]. In the present analysis, the contribution of these items to the response was numerically different as this contribution was based on the ANN data‐driven models specifically developed using the data of each study.
The treatment effect and the effect size estimated in the population including all subjects and in the population without subject having a high probability of NSRT (i.e., prob‐NSRT > 0.5) are presented in Table 3. These results indicated that the control of the NSTR enhances signal detection, significantly increasing the effect size due to better control of the heterogeneity in the placebo response. These findings are consistent with the expected effect of a high placebo response on the estimable TE [37].
TABLE 3.
Treatment effect and effect size estimated in the total population and in the population without subject with non‐specific treatment response.
| Study | Treatment effect | Total population | Population with prob‐NSRT ≤ 0.5 | ||
|---|---|---|---|---|---|
| Mean HAMD‐17 change from baseline (95% CI) | Effect size | Mean HAMD‐17 change from baseline (95% CI) | Effect size | ||
| 810 | TE1 | −0.77 (−2.59, 1.05) | 0.11 | −3.94 (−6.70, −1.17) | 0.55 |
| TE2 | −2.16 (−4.04, −0.33) | 0.31 | −4.77 (−7.70, −1.84) | 0.67 | |
| 874 | TE1 | −0.47 (−2.22, 1.28) | 0.07 | −4.01 (−5.82, −2.19) | 0.72 |
| TE2 | −1.96 (−3.79, −0.13) | 0.26 | −4.99 (−6.93, −3.04) | 0.82 | |
Abbreviations: TE1 = Placebo‐12.5 mg arm at Week 8; TE2 = Placebo‐25 mg arm at Week 8.
The ANN model developed with the data of 810 (or 874) study was used to predict the prob‐NSTR in the study 874 (or 810) to evaluate the generatability of the ANN model for the prediction of prob‐NSTR in new trials. The generalizability of the ANN models was further assessed using the ROC AUC and the associated 95% CI estimated by comparing the NSTR predictions based on one trial with observations made in another trial. The results of the analysis indicated that the prediction of the outcomes of study 874 based on the model developed using data of the 810 trial were: ROC AUC = 0.73 (95% CI: 0.55–0.92). The outcomes of study 810 were predicted based on the model developed using data from the 874 trial: ROC AUC = 0.79 (95% CI: 0.56–1).
4. Discussion
The present study demonstrates how machine learning methods can help to identify pre‐randomization characteristics of patients most likely to show a non‐specific response to treatment. Given the large number of ML methodologies, an initial comparison of the alternative methods revealed an important difference in the performance of the different approaches and that the ANN model provided the most performing predictive performances.
The non‐specific response to treatment has been shown to represent a major issue in the assessment of clinical efficacy of novel antidepressant medications [37]. The problem of identifying specific and non‐ specific responses to a treatment in RCTs is complex as these effects are latent and not directly measured. When a subject responds to a drug treatment, we cannot assess whether the subject responded to a specific or to a non‐specific effect of the treatment. An additional complication is that a subject may respond to both specific and non‐ specific effects and the degree of each can vary from subject to subject. For this reason, there is an unmet need either for developing novel methodologies to differentiate the specific from the non‐specific response in subjects treated with active drug or to develop new methodologies for providing a NSTR‐adjusted TE. The present study was primarily aimed to identify the most performing ML model to estimate the prob‐NSTR in RCTs conducted in MDD. Among the different model evaluated, ANN achieved the highest overall accuracy in the two cases study evaluated. A fivefold cross‐validation confirmed the performances and indicated the limited risk of overfitting of the best performing ANN model.
To illustrate the clinical utility of the best performing ML method to estimate the probability of NSTR, we compared TE and effect size on the entire datasets for the 810 and 874 studies to the values estimated using only the subjects without NSTR. The results of the comparison indicated that the control of the NSTR enhance signal detection, with a significant increase of the effect size due to a better control of the heterogeneity in the response to placebo. The assessment of the generability of the predictive performances of the 17 individual items evaluated in one study showed a poorly performance in the estimation of the prob‐NSRT in another study. This lack of generalizability of the ANN model was also confirmed by the comparison of the relative importance of each selected item in the 810 and 874 studies in the prediction of the individual prob‐NSRT.
The proposed approach is very simple to be implemented since it is based on standard multivariate prognostic variables (the individual HAMD‐17 item scores estimated at the screening and randomization visits). The individual probability (prob‐NSTR) can be easily computed and used in at least three scenarios.
Scenario 1: Use the prob‐NSTR in a post hoc analysis conducted on data generated on historical RCT data to re‐assess the TE adjusted by the prob‐NSTR distribution.
Scenario 2: Prospectively use the prob‐NSTR in the analysis of new RCT using novel methodological approaches for controlling the effect of excessively high placebo response [22, 38]. The prospective use of prob‐NSTR would require to satisfy the following conditions: (i) the study was designed to collect screening and pre‐treatment baseline data, (ii) the criteria for assessing the clinical response to placebo were pre‐specified in the statistical analysis plan (SAP), (iii) the acceptable criteria for the predictive performance of the ANN model was prospectively defined in the SAP specifying that the acceptable ROC AUC cut‐offs should be statistically greater than the non‐informative threshold of 0.5.
Scenario 3: Use the prob‐NSTR score estimated using pre‐randomization data as an enrichment criteria in future randomized placebo‐controlled trials to select a study population in which detection of a drug effect is more likely than it would be in an unselected population [39]. The implementation of this scenario would require to dispose of a general ANN model. This should be initially trained on a large collection of patient‐level MDD trials linking pre‐randomization measurements to prob‐NSTR on large and pooled databases of RCTs (accounting for time‐varying placebo response and inter‐study variability). Then, this general model should be validated on external MDD datasets. In addition, we would also need to dispose of well‐grounded hypothesis on the relationship between pre‐randomization predictors and prob‐NSTR. This general model could then be used as ‘prior’ information in a Bayesian framework to estimate the individual probability of being a placebo responder in any new RCT. In the present study, the generalizability of the ANN models was preliminary assessed by comparing the acceptable level of the predictive performance of a model trained with the data of the 810 study on the test data set of the 874 study and by comparing the predictive performance of a model trained with the data of the 874 study on the test data set of the 810 study. This comparison indicated a poor predictive performance of the model developed in one study to predict the prob‐NSRT in another study.
Several limitations of the current investigation should be noted. First of all, the restricted number of RCTs evaluated, then the HAMD‐17 rating scale was the only clinical scale evaluated. Other relevant clinical scales such as MADRS, or PANSS have to be analyzed to replicate the results found. In addition, as the unpredictable high placebo response rate is one of the major factor associated with the failure of randomized clinical trials in a large majority of psychiatric disorders such as bipolar disorders, schizophrenia, anxiety, etc., the present approach would need to be also evaluated in trials conducted on these disorders.
In summary, the ML methodology provides a powerful data‐driven approach to empirically define the relationship between predictors and study outcomes. By definition, no mechanistical or semi‐mechanistical assumptions are used and/or required. This represents an advantage but also an important limitation that prevents, for example, to simulate the study outcomes when the predictor variables vary.
Finally, it is important to note that FDA has recently issued two discussion papers on the topic of using AI and ML in the development of drug/biologic products given the growing interest on the use of AI/ML methodologies [40, 41] and the increasing number of regulatory submissions to the FDA that included AI/ML [39].
The regulatory interest for using AI/ML methodologies emphasizes the interest of exploring novel AI/ML methodologies for improving the efficiency of clinical trials, and in particular for controlling the NSTR.
Author Contributions
R.G. and F.B.‐G. wrote the manuscript; R.G. designed the research; R.G. and F.B.‐G. performed the research and analyzed the data.
Ethics Statement
The authors have nothing to report.
Conflicts of Interest
R.G. and F.B.‐G. were paid consultants to Mapi pharma, Otsuka pharmaceutical, Auitifony Therapeutics; Tris Pharma; Orexia Therapeutics; Sunovion Pharmaceuticals; Chemopharma; Supernus Pharmaceuticals; Ironshore Pharmaceutical; Exeltis Pharma; UCB Pharma; Universal Pharma; Teva Pharmaceuticals; 4SC AG; Alfasigma; Recordati; CeNeRx BioPharma; GlaxoSmithKline; ViiV Healthcare; Hoffman‐LaRoche; Indivior; Johnson & Johnson Pharmaceutical Research & Development; Reckitt Benckiser; Relmada Therapeutics Inc.; KYE Pharmaceuticals; Orphan Europe; Singapore Agency for Science, Technology and Research (A*STAR); Amgen Inc.; Allecra Therapeutics; NDA Regulatory Service AB; Gilead Science Inc.; Theravance Biopharma; Sensorion SA; AstraZeneca.
Supporting information
Figures S1–S4.
Funding: The authors received no specific funding for this work.
Data Availability Statement
The analysis are based on previously published data; therefore, data are not available.
References
- 1. Khan A. and Brown W. A., “Antidepressants Versus Placebo in Major Depression: An Overview,” World Psychiatry 14 (2015): 294–300. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Salloum N. C., Fava M., Ball S., and Papakostas G. I., “Success and Efficiency of Phase 2/3 Adjunctive Trials for MDD Funded by Industry: A Systematic Review,” Molecular Psychiatry 25, no. 9 (2020): 1967–1974. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Kirsch I., Deacon B. J., Huedo‐Medina T. B., Scoboria A., Moore T. J., and Johnson B. T., “Initial Severity and Antidepressant Benefits: A Meta‐Analysis of Data Submitted to the Food and Drug Administration,” PLoS Medicine 5, no. 2 (2008): e45. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Rutherford B. R. and Roose S. P., “A Model of Placebo Response in Antidepressant Clinical Trials,” American Journal of Psychiatry 170, no. 7 (2013): 723–733. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Iovieno N. and Papakostas G. I., “Correlation Between Different Levels of Placebo Response Rate and Clinical Trial Outcome in Major Depressive Disorder: A Meta‐Analysis,” Journal of Clinical Psychiatry 73 (2012): 1300–1306. [DOI] [PubMed] [Google Scholar]
- 6. Papakostas G. I. and Fava M., “Does the Probability of Receiving Placebo Influence Clinical Trial Outcome? A Meta‐Regression of Double‐Blind, Randomized Clinical Trials in MDD,” European Neuropsychopharmacology 19, no. 1 (2009): 34–40. [DOI] [PubMed] [Google Scholar]
- 7. Turner E. H., Matthews A. M., Linardatos E., and Tell R. A., “Rosenthal R Selective Publication of Antidepressant Trials and Its Influence on Apparent Efficacy,” New England Journal of Medicine 358 (2008): 252–260. [DOI] [PubMed] [Google Scholar]
- 8. Fava M., Evins A., Dorer D., and Schoenfeld D., “The Problem of the Placebo Response in Clinical Trials for Psychiatric Disorders: Culprits, Possible Remedies, and a Novel Study Design Approach,” Psychotherapy and Psychosomatics 72, no. 3 (2003): 115–127. [DOI] [PubMed] [Google Scholar]
- 9. Walsh B. T., Seidman S. N., Sysko R., and Gould M., “Placebo Response in Studies of Major Depression: Variable, Substantial and Growing,” JAMA 287 (2002): 1840–1847. [DOI] [PubMed] [Google Scholar]
- 10. Kobak K. A., Kane J. M., Thase M. E., and Nierenberg A. A., “Why Do Clinical Trials Fail?: The Problem of Measurement Error in Clinical Trials: Time to Test New Paradigms?,” Journal of Clinical Psychopharmacology 27 (2007): 1–5. [DOI] [PubMed] [Google Scholar]
- 11. Merlo‐Pich E., Alexander R. C., Fava M., and Gomeni R., “A New Population‐Enrichment Strategy to Improve Efficiency of Placebo‐Controlled Clinical Trials of Antidepressant Drugs,” Clinical Pharmacology and Therapeutics 88, no. 5 (2010): 634–642. [DOI] [PubMed] [Google Scholar]
- 12. Palpacuer C., Gallet L., Drapier D., Reymann J. M., Falissard B., and Naudet F., “Specific and Non‐Specific Effects of Psychotherapeutic Interventions for Depression: Results From a Meta‐Analysis of 84 Studies,” Journal of Psychiatric Research 87 (2017): 95–104. [DOI] [PubMed] [Google Scholar]
- 13. Naudet F., Maria A. S., and Falissard B., “Antidepressant Response in Major Depressive Disorder: A Meta‐Regression Comparison of Randomized Controlled Trials and Observational Studies,” PLoS One 6, no. 6 (2011): e20811. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Petkova E., Tarpey T., and Govindarajulu U., “Predicting Potential Placebo Effect in Drug Treated Subjects,” International Journal of Biostatistics 5 (2009): 23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Sarker I. H., “Machine Learning: Algorithms, Real‐World Applications and Research Directions,” SN Computer Science 2, no. 3 (2021): 160. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Smith E. A., Horan W. P., Demolle D., et al., “Using Artificial Intelligence‐Based Methods to Address the Placebo Response in Clinical Trials,” Innovations in Clinical Neuroscience 19, no. 1–3 (2022): 60–70. [PMC free article] [PubMed] [Google Scholar]
- 17. Faraone S. V., Gomeni R., Hull J. T., et al., “Early Response to SPN‐812 (Viloxazine Extended‐Release) Can Predict Efficacy Outcome in Pediatric Subjects With ADHD: A Machine Learning Post‐Hoc Analysis of Four Randomized Clinical Trials,” Psychiatry Research 296 (2021): 113664. [DOI] [PubMed] [Google Scholar]
- 18. Zilcha‐Mano S., Roose S. P., Brown P. J., and Rutherford B. R., “A Machine Learning Approach to Identifying Placebo Responders in Late‐Life Depression Trials,” American Journal of Geriatric Psychiatry 26, no. 6 (2018): 669–677. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Taliaz D., Spinrad A., Barzilay R., et al., “Optimizing Prediction of Response to Antidepressant Medications Using Machine Learning and Integrated Genetic, Clinical, and Demographic Data,” Translational Psychiatry 11 (2021): 381. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Qiu Y., Zhu X., and Ma X., “Using Machine Learning Models to Identify the Risk of Depression in Middle‐Aged and Older Adults With Frequent and Infrequent Nicotine Use: A Cross‐Sectional Study,” Journal of Affective Disorders S0165‐0327, no. 24 (2024): e01417‐4. [DOI] [PubMed] [Google Scholar]
- 21. Fava G., Guidi J., Rafanelli C., and Rickels K., “The Clinical Inadequacy of the Placebo Model and the Development of an Alternative Conceptual Framework,” Psychotherapy and Psychosomatics 86 (2017): 332–340. [DOI] [PubMed] [Google Scholar]
- 22. Gomeni R., Bressolle‐Gomeni F., and Fava M., “A New Method for Analyzing Clinical Trials in Depression Based on Individual Propensity to Respond to Placebo Estimated Using Artificial Intelligence,” Psychiatry Research 327 (2023): 115367. [DOI] [PubMed] [Google Scholar]
- 23. Trivedi M. H., Pigotti T. A., Perera P., Dillingham K. E., Carfagno M. L., and Pitts C. D., “Effectiveness of Low Doses of Paroxetine Controlled Release in the Treatment of Major Depressive Disorder,” Journal of Clinical Psychiatry 65, no. 10 (2004): 1356–1364. [DOI] [PubMed] [Google Scholar]
- 24. Schaefer D., Pitts C., Lipschitz A., and Iyengar M., “Efficacy and Tolerability of Fixed, Low Dose Paroxetine CR in the Treatment of Major Depression in the Elderly Poster No NR701,” Presented at the American Psychiatric Association Annual Meeting, May 2005, accessed September 23, 2024, https://www.psychiatry.org/getattachment/f1320f18‐01ca‐43e7‐9e3f‐0826528702ee/am_newresearch_2005.pdf.
- 25. Hamilton M., “A Rating Scale for Depression,” Journal of Neurology, Neurosurgery, and Psychiatry 23 (1960): 56–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Leucht S., Fennema H., Engel R. R., Kaspers‐Janssen M., Lepping P., and Szegedi A., “What Does the MADRS Mean? Equipercentile Linking With the CGI Using a Company Database of Mirtazapine Studies,” Journal of Affective Disorders 210 (2017): 287–293. [DOI] [PubMed] [Google Scholar]
- 27. Leucht S., Fennema H., Engel R. R., Kaspers‐Janssen M., and Szegedi A., “Translating the HAM‐D Into the MADRS and Vice Versa With Equipercentile Linking,” Journal of Affective Disorders 226 (2018): 326–331. [DOI] [PubMed] [Google Scholar]
- 28. Kolen M. J. and Brennan R. L., “Observed Score Equating Using the Random Groups Design,” in Test Equating, Scaling, and Linking (New York, NY: Springer, 2017), 29–63. [Google Scholar]
- 29. Kuhn M., “Building Predictive Models in R Using the Caret Package,” Journal of Statistical Software 28, no. 5 (2008): 1–26.27774042 [Google Scholar]
- 30. Yu H., Samuels D. C., Zhao Y. Y., and Guo Y., “Architectures and Accuracy of Artificial Neural Network for Disease Classification From Omics Data,” BMC Genomics 20 (2019): 167–178. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. R Core Team , A Language and Environment for Statistical Computing (Vienna, Austria: R Foundation for Statistical Computing, 2022). [Google Scholar]
- 32. Lever J., Krzywinski M., and Altman N., “Model Selection and Overfitting,” Nature Methods 13 (2016): 703–704. [Google Scholar]
- 33. Aziz Taha A. and Hanbury A., “Metrics for Evaluating 3D Medical Image Segmentation: Analysis, Selection, and Tool,” BMC Medical Imaging 15 (2015): 1–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Olden J. D., Joy M. K., and Death R. G., “An Accurate Comparison of Methods for Quantifying Variable Importance in Artificial Neural Networks Using Simulated Data,” Ecological Modelling 178 (2004): 389–397. [Google Scholar]
- 35. Evans K. R., Sills T., DeBrota D. J., Gelwicks S., Engelhardt N., and Santor D., “An Item Response Analysis of the Hamilton Depression Rating Scale Using Shared Data From Two Pharmaceutical Companies,” Journal of Psychiatric Research 38, no. 3 (2004): 275–284. [DOI] [PubMed] [Google Scholar]
- 36. Santen G., Gomeni R., Danhof M., and Della P. O., “Sensitivity of the Individual Items of the Hamilton Depression Rating Scale to Response and Its Consequences for the Assessment of Efficacy,” Journal of Psychiatric Research 42, no. 12 (2008): 1000–1009. [DOI] [PubMed] [Google Scholar]
- 37. Enck P. and Klosterhalfen S., “The Placebo Response in Clinical Trials—The Current State of Play,” Complementary Therapies in Medicine 21, no. 2 (2013): 98–101. [DOI] [PubMed] [Google Scholar]
- 38. Gomeni R., Hopkins S., Bressolle‐Gomeni F., and Fava M., “Interpreting Clinical Trial Outcomes Complicated by Placebo Response With an Assessment of False‐Negative and True‐Negative Clinical Trials in Depression Using Propensity‐Weighting,” Translational Psychiatry 13 (2023): 388. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. FDA , “Guidance for Industry, Enrichment Strategies for Clinical Trials to Support Determination of Effectiveness of Human Drugs and Biological Products,” (2019), accessed September 23, 2024, https://www.fda.gov/media/121320/download.
- 40. Liu Q., Huang R., Hsieh J., et al., “Landscape Analysis of the Application of Artificial Intelligence and Machine Learning in Regulatory Submissions for Drug Development From 2016 to 2021,” Clinical Pharmacology and Therapeutics 113, no. 4 (2023): 771–774. [DOI] [PubMed] [Google Scholar]
- 41. FDA , “Discussion Papers on AI/ML,” accessed October 17, 2024, https://www.fda.gov/news‐events/fda‐voices/fda‐releases‐two‐discussion‐papers‐spur‐conversation‐about‐artificial‐intelligence‐and‐machine?utm_medium=email&utm_source=govdelivery.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Figures S1–S4.
Data Availability Statement
The analysis are based on previously published data; therefore, data are not available.
