Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2023 May 17.
Published in final edited form as: Clin Psychol Sci. 2022 Dec 22;11(3):458–475. doi: 10.1177/21677026221120236

Prediction of Attention-Deficit/Hyperactivity Disorder Diagnosis Using Brief, Low-Cost Clinical Measures: A Competitive Model Evaluation

Michael A Mooney 1,2, Christopher Neighbor 3, Sarah Karalunas 6, Nathan F Dieckmann 3,5, Molly Nikolas 4, Elizabeth Nousen 3, Jessica Tipsord 3, Xubo Song 1, Joel T Nigg 3,7
PMCID: PMC10191260  NIHMSID: NIHMS1827576  PMID: 37205171

Abstract

Proper diagnosis of ADHD is costly, requiring in-depth evaluation via interview, multi-informant and observational assessment, and scrutiny of possible other conditions. The increasing availability of data may allow the development of machine-learning algorithms capable of accurate diagnostic predictions using low-cost measures to supplement human decision-making. We report on the performance of multiple classification methods used to predict a clinician-consensus ADHD diagnosis. Methods ranged from fairly simple (e.g., logistic regression) to more complex (e.g., random forest), while emphasizing a multi-stage Bayesian approach. Classifiers were evaluated in two large (N>1000), independent cohorts. The multi-stage Bayesian classifier provides an intuitive approach consistent with clinical workflows, and was able to predict expert consensus ADHD diagnosis with high accuracy (>86%)—though not significantly better than other methods. Results suggest that parent and teacher surveys are sufficient for high-confidence classifications in the vast majority of cases, while an important minority require additional evaluation for accurate diagnosis.

Keywords: Attention deficit hyperactivity disorder, classification, machine learning

INTRODUCTION

The accurate diagnosis of attention-deficit/hyperactivity disorder (ADHD) is critically important given its prevalence, impairment, and cost. Yet, the literature reflects substantial and ongoing concerns about both over- and under-diagnosis and treatment (Costello et al., 2014; Fabiano & Haslam, 2020; Hamed et al., 2015; Kazda et al., 2021; Massuti et al., 2021; Simon et al., 2015) of this costly condition. A full evaluation of ADHD, requiring significant time and resources, involves a structured or semi-structured clinical interview, standardized ratings from parent and another informant from another setting (for children, typically the teacher), and evaluation of impairment, as well as comorbid conditions that might better explain the diagnosis (APA, 2013; Committee on Quality Improvement, Subcommittee on Attention-Deficit/Hyperactivity Disorder, 2000; National Guideline Centre (UK), 2018). Yet, surveys suggest that the majority of providers in primary care faced with evaluating ADHD report insufficient knowledge or resources to carry out a full evaluation in both children (French et al., 2019; Young et al., 2021) and adults (Adler et al., 2009; Faraone et al., 2004) contributing to substantial concerns about diagnostic accuracy, including both over- and under-diagnosis (Fabiano & Haslam, 2020; Kazda et al., 2021; Massuti et al., 2021) and leading to efforts to develop additional resources (Loskutova et al., 2021). Thus, while other factors may contribute, and the boundaries of all psychological diagnoses are fuzzy, there is consistent evidence that a lack of time and other resources and a lack of mental health expertise in front-line providers directly contributes to diagnostic inaccuracy and over-/under-treatment. Tools to aid clinician decision-making, particularly computational tools, may play a role in alleviating the significant demand on time and resources needed for ADHD assessment. Clinical decision aids are not meant to replace clinician judgement entirely, but may augment human decision-making, help reduce barriers and facilitate standards of care in settings where they are currently suboptimal.

Multiple approaches to assist diagnostic evaluation are under development in the field, using various contemporary computational tools. For example, computerized adaptive testing is a sophisticated method that uses item response theory and large item banks to develop rapid assessment tools (Gibbons et al., 2016, 2020). A similar approach was used to develop the ongoing NIMH PROMIS scales (Irwin et al., 2010). Such prior work has provided brief measures (Cella et al., 2007), and computer-assisted assessment is emerging (Hall et al., 2016; Nikolas et al., 2019; Slobodin et al., 2020). However, their integration into clinical practice requires adoption of new assessment procedures and tools that are not yet readily available to providers. A potentially simpler approach is to create a prediction algorithm that clinicians can apply to readily available standard clinical assessment tools using a machine aid to interpret the data.

Two questions are salient in relation to the possibility of a prediction algorithm. First, can brief, low-cost measures be used in prediction algorithms to accurately reproduce clinical diagnoses achieved by more costly multi-informant, multi-clinician formal evaluation? If so, they could greatly aid diagnosis and diagnostic confidence, particularly in settings where providers are broadly trained but may lack specific knowledge of psychological diagnostic standards. In addition, even in settings where diagnostic expertise for ADHD is present, such algorithms can provide potentially valuable cost savings. Second, what level of statistical complexity is needed to maximize diagnostic accuracy? Do complex contemporary machine learning algorithms outperform simpler and more familiar linear models at the sample sizes commonly available to local clinics and research studies? This is a methodological question that we also undertake.

The effort to use machine learning classifiers to identify clinical cases of psychiatric disorders on the basis of low-cost clinical data has been surprisingly sparse. In general, studies using machine learning to predict ADHD diagnosis have been conducted with small samples and have lacked external validation (Bledsoe et al., 2016; Christiansen et al., 2020; Duda et al., 2016; Dvorsky et al., 2016; Mueller et al., 2010; Wodka et al., 2008). To the relative neglect of simpler clinical tools, much of the recent literature on the prediction of ADHD or other psychiatric diagnosis has focused on brain imaging measures, with dozens of studies but with generally small samples, weak accuracy, or overfitting concerns (limited cross validation) (Rashid & Calhoun, 2020). High cost, along with small effect sizes for MRI associations (Marek et al., 2022), renders brain imaging currently sub-optimal for a brief, cost reducing ADHD diagnostic method. One study of ADHD used a computerized interview and teacher and parent rating scales to establish ADHD diagnosis and then related teacher ratings to establish prediction, but without external validation (Öztekin et al., 2021). They found no added benefit of MRI measures. Likewise, genetic risk factors (e.g., polygenic risk scores or individual variants) still have limited utility in terms of clinical diagnosis/prognosis (Martin, Daly, et al., 2019; Martin, Kanai, et al., 2019; Palk et al., 2019; Ronald et al., 2021), although they may be added to a diagnostic algorithm in some form in the relatively near future.

Here, we maximize clinical utility and potential translation by focusing on data types that are readily available in standard clinical practice. Our question is not whether we can improve prediction with biological or other measures that are currently outside the scope of DSM, but whether we can replicate a costly, gold-standard diagnosis with low-cost measures that are easily obtained in a wide range of settings where diagnosis may occur. Nonetheless, even focusing on these commonly used data types, it remains unclear what an optimal degree of density of assessment would be for ADHD diagnosis in a clinical setting. Parent and teacher ratings are strongly recommended for ADHD evaluation, yet these often do not agree. A parent interview may yield different responses than a rating scale. Rating scales utilize clinical “cut off scores” for ADHD that may or may not be optimal for individual prediction. Finally, cognitive tests remain controversial as ancillary information for ADHD diagnostic evaluation. They contribute to understanding of functional impairment, yet it is unknown whether they can also aid a diagnostic algorithm in cases of diagnostic uncertainty. For example, Nikolas et al. (2019) reported that in a simple regression model, certain cognitive measures enhanced ADHD prediction accuracy over and above rating scales (Nikolas et al., 2019). Here, we address questions about the incremental validity of these varying data types by examining the relative accuracy of algorithms based on (a) parent ratings alone, (b) parent + teacher ratings, and (c) parent + teacher + cognitive testing in predicting both a gold standard best-estimate diagnosis and a clinician structured interview diagnosis of ADHD.

A second major limitation of the clinical prediction literature is the emphasis on evaluating only complex machine learning approaches and neglecting simpler, lower cost algorithms that might enhance clinical efficiency. Such comparison of algorithms is critical for clinical implementation. For example, in a competitive modeling approach similar to the approach used here, Youngstrom et al. (Youngstrom et al., 2018) reported on a set of studies attempting to identify cases of bipolar disorder. They reported that complex machine learning algorithms, in that case Least Absolute Shrinkage and Selection Operation (LASSO) regression, did not markedly outperform the more familiar logistic regression classifier when validated on independent data.

Scrutiny of other mental health conditions along similar lines has been limited, and algorithm comparisons have not been conducted for ADHD. Here, we report on a competitive evaluation of multiple classification methods. In our selection of algorithms, we intentionally include those that mirror clinical decision-making. In particular, informal Bayesian logic—weighing new information in the context of existing evidence—is central to clinical decision-making (Gill et al., 2005). Because of this, we evaluate a multi-stage Bayesian prediction approach (a tree-augmented naïve Bayes classifier) to roughly parallel the sequential decision-making process that a time- and cost-conscious clinician uses. In the case of clinical evaluation for ADHD, the decision-making would typically start with a brief low-cost assessment (e.g., a single parent symptom checklist), then proceed to decisions about whether to add additional measures such as obtaining teacher ratings or cognitive testing, and so on. Thus, our sequence starts with the most easily-obtained measures (parent rating scales) and proceeds to teacher ratings, and then laboratory cognitive test measures.

In addition to the tree-augmented naïve Bayes classifier, we implemented a competitive modeling logic with the following classifiers, varying in computational complexity: a simple 2-level decision tree, standard logistic regression, regularized logistic regression, support vector machine, unrestricted-depth decision tree, random forest, and gradient boosted decision trees. All classification methods were evaluated in two large (N>1000), well-characterized, case-control cohorts to evaluate robustness and generalizability. The goal here was not to rediscover a high-cost algorithm, but rather to evaluate whether a prediction algorithm based on low-cost tools could reasonably match a high-cost multi-informant/multi-clinician assessment.

Because clinical prediction of this type has yet to be strongly evaluated for ADHD, the first and most basic challenge is to differentiate those with a clinical diagnosis (in this case, ADHD) from those without ADHD at sample sizes that may be typically available in local clinics. This paper illustrates an approach to tackling that challenge, by developing and evaluating multiple machine-learning classifiers, ranging in complexity, to differentiate ADHD from non-ADHD youth in a population with a moderate ADHD base rate. Once this is solved, the next steps are to address differentiation of ADHD from other disorders (i.e., the most typical clinical challenge) and prediction of clinical course to guide treatment decisions, all while considering different settings’ base rates.

TRANSPARENCY AND OPENNESS

Preregistration

None of the studies reported here were preregistered.

Data, Materials, Code, and Online Resources

Data for the Oregon-ADHD-1000 cohort is available on the NIMH Data Archive (NDA Collection 2857). All software packages used in the reported analyses are publicly available (see Methods), and specific code will be made available upon reasonable request to the authors. Supplemental materials referenced herein will be made available on the journal’s website.

Reporting

We report how we determined our sample size, all data exclusions, all manipulations, and all measures in the study.

Ethical Approval

The research reported here was approved by the institutional review board of Oregon Health & Science University and was conducted in accordance with the Declaration of Helsinki (6th revision, 2008).

MATERIALS AND METHODS

Sample Size and Participants

Methodologically, the relation between a machine learning model and sample size depends on many factors such as model nonlinearity and complexity, sample noise, dependency among the samples, and effectiveness of model optimization processes. It has been theoretically derived and empirically demonstrated (Abu-Mostafa et al., 2012; Bishop, 2006; Goodfellow et al., 2016) that complex machine learning models, such as multi-layer deep learning neural networks, are capable of capturing the complex, nonlinear relations between the model inputs and outputs. However, they require a large amount of data to train, due to the large number of model parameters to specify without overfitting. On the other hand, simple models, such as logistic regression models, can still be the appropriate model for a given data set, in cases of modest sample size, highly correlated samples (which reduces the effective sample size), noisy samples, and close-to-linear input-output relations. Thus, evaluation is necessary for the particular question at hand. Here, we opted for well-characterized samples that had sufficient sample size to test the primary question, that compared well to other similar efforts in the literature, and that might be similar to what most real settings would have available to train a local classifier.

The Oregon-ADHD-1000 is a community recruited case-control cohort (N=1423) of youth age 7-11 years (Karalunas et al., 2017; Mooney et al., 2020, 2021; Nigg et al., 2018, 2020). ADHD was deliberately oversampled to ensure an adequate range of clinical variation and of actual clinical cases in the data set. (We consider base rate issues later). To preserve the representativeness of the sample, we did not oversample for sex or other demographics, and these were not included in our algorithm to mimic clinician context. Human subjects and ethics approval was obtained from the local University Institutional Review Board. A parent/legal guardian provided written informed consent, and children provided written assent.

Recruitment was conducted by community outreach and a case-finding procedure to identify cases in the community, regardless of whether they had sought treatment or been previously diagnosed. After screening, a state-of-the-art, multi-informant, multi-method research diagnostic evaluation was conducted (see below). Children were excluded for disallowed medications (Supplemental Table S1), history of seizures or head injury, psychosis, mania, current major depressive episode, Tourette’s syndrome, autism and IQ<80. An ADHD assignment required all DSM-5 criteria to be met including parent-teacher convergence (both having either above-threshold rating scale scores or symptom counts). To increase the difficulty of the diagnostic case and make it more similar to real world decisions, the sample retained youth who were subthreshold for ADHD (5 symptoms + impairment) and others we label NOS for convenience—they had elevated symptoms, but by only one reporter, or were being medically treated for other conditions that made diagnosis highly ambiguous. As explained below, sensitivity analyses evaluated the effect of these cases.

Replication in a completely independent sample was evaluated with the Michigan-ADHD-1000 (Nikolas & Nigg, 2013, 2015). It is a cohort of 1064 youth ages 6-21 years, with the same recruitment and assessment procedures, but in a very different demographic population (central Michigan versus northwest Oregon). Its inclusion enabled a test of generalizability and performance in a completely independent sample to control any conclusions that might derive from single-sample over-fitting (despite our within-sample precautions, below).

ADHD Best-estimate Diagnosis

In both cohorts, diagnostic assignment followed a standard protocol. It included standardized, nationally-normed rating scales from parent and teacher, parent semi-structured clinical interview by a trained clinician with acceptable inter-clinician reliability and validity, child intellectual testing with carefully trained and supervised psychometrician administrators, and clinical observation and notes from two research assistants. Then, a best-estimate diagnostic team of two highly experienced clinicians (board certified child psychiatrist and a child clinical psychologist) reviewed all available data and arrived at a best-estimate diagnostic assignment, considering whether symptoms were better explained by an associated condition and whether impairment was severe enough to warrant diagnosis. The two clinicians agreed well (kappa>0.80) on ADHD assignment.

These clinician consensus diagnoses, which we refer to as best-estimate diagnoses, were used as the ground-truth classifications. The best-estimate team classified each subject as ADHD, control, subthreshold, or “other” (here called NOS for simplicity). The subthreshold category were children judged by the consensus best-estimate team to have impairing ADHD symptoms but insufficient to meet full diagnostic criteria (typically, 5 symptoms for DSM-IV under age 17). The NOS category included cases in which only one reporter reported elevated symptoms (90% of the NOS group), or cases with ADHD symptoms but also being treated for other conditions. Non-ADHD cases had 4 or fewer ADHD symptoms and had never been identified or treated for ADHD. Other psychopathology was free to vary in these community volunteers, although we excluded children with mania, a history of probable or definite psychosis, intellectual delay, or autism spectrum disorder or neurological injury.

For our primary analyses we collapse the children into two classes (ADHD, and non-ADHD); the subthreshold and NOS cases were all labeled as non-ADHD for that primary analysis. In sensitivity analyses we explored the effect of this type of decision versus a broader ADHD definition that included subthreshold cases and also how inclusion or exclusion of these cases from the training set affected classification learning.

The best-estimate team used some of the rating scales that were included in our low-cost classification models, our first analysis focused on whether those results could be approximated with a subset of their information. However, to control the potential “double-dip” of that approach, we also evaluated classifiers trained to predict diagnosis based on the Kiddie Schedule for Affective Disease and Schizophrenia-version E modified for DSM-IV (at the time of data collection) and checked for DSM-5 compliance (after DSM-5 was published). The KSADS-E was administered by a masters-degree or higher clinician (either masters in social work, in counseling, or in clinical psychology) trained to reliability and validity with a master coder who was in turn trained to validity with an outside expert. These interviewers all achieved adequate inter-rater reliability (kappa>0.80) with the master trainer. Recordings of their interviews were regularly reviewed by senior clinicians to prevent administration drift.

Predictive Measures

The initial modeling used the Oregon cohort. The following symptom measures were used as predictors in the classifiers: parent reported inattention and hyperactivity symptom counts from the ADHD Rating Scale (ADHDRS) (DuPaul et al., 1998) parent reported inattention, hyperactivity, executive functioning, learning problems, aggression, and peer relations scales from the Conners Rating Scale 3rd edition (Conners-3) (Conners, 2008), and teacher reported inattention and hyperactivity symptom counts from the ADHDRS. All parent and teacher reported measures were converted to T-scores to adjust for age and gender and again to simulate clinical decision-making. These symptom measures or others like them are often the first pieces of information used by a clinician when evaluating a child for ADHD.

In addition, cognitive measures from the following laboratory tests were used as predictors: Stop-Go task (Nigg, 1999; Schachar et al., 1995); Spatial Span forward and backward (De Luca et al., 2003); Digit Span forward and backward from the Wechsler Intelligence Scale for Children, Fourth Edition (WISC-IV) (Wechsler, 2003); Delis, Kaplan, and Kramer (DKEF) (Delis et al., 2001) version of the Stroop task (color, word, and color-word conditions); and DKEF Trail Making test (number sequencing and number-letter sequencing conditions).

For the Michigan cohort, parent reported symptoms were assessed with the ADHDRS (same as the Oregon cohort), and the cognitive problems and hyperactivity-impulsivity scales from the Revised Conners Parent Rating Scale (Conners-R) (Conners et al., 1998), due to the earlier era of data collection prior to publication of the Conners-3. Teacher reported inattention and hyperactivity symptoms were available from the ADHDRS. Again, these symptom measures were converted to T-scores to simulate clinician decision-making.

All of the cognitive measures available in the Oregon cohort were also available for the Michigan cohort, although a different version of the Trail Making test was used in Michigan (Reitan & Wolfson, 1985)). Supplemental Table S2 shows all predictive features available in each data set.

Data Cleaning / Pre-processing

Subjects without a best-estimate team diagnosis and those missing >80% of predictive features were excluded, leaving a total N=1423 and N=1057 subjects for analyses in the Oregon and Michigan cohorts, respectively. KSADS diagnoses were available for 1422 in the Oregon cohort and 1055 in the Michigan cohort.

For the primary analyses in the Oregon cohort, a stratified split of the data was done to maintain equal proportions of ADHD and non-ADHD cases in the sub-samples used for training (75%) and testing (25%) the classifiers. All variables were standardized (mean=0, standard deviation=1) using the StandardScaler method in the Scikit-learn Python package (Pedregosa et al., 2011). The test data was standardized relative to the mean and standard deviation of the training set, to bring the test data into the same range without data leakage from the test set.

Missing Data Imputation

For machine learning applications, addressing missing data must be done before the optimal predictive model is known (i.e., prior to classifier training). Because of this, model-based methods for handling missing data (e.g., maximum likelihood or regression), which depend on assumptions about the distributions of variables or the relationships between variables, are typically not appropriate. The best approach for machine learning tasks is unknown (and likely data set-specific), but almost certainly listwise deletion is suboptimal. Here, a k-nearest neighbor (KNN) imputation approach, implemented with the KNNImputer method (number of neighbors, k=7) in Scikit-learn, was used to impute missing values. The imputer model was fit with the training data only, to avoid data leakage from the test set, and was then applied to both the training and test sets. KNN imputation has been used in similar applications and has outperformed model-based methods and listwise deletion in terms of classification performance (Jerez et al., 2010). Alternative imputation methods were tested in simulated data (see Supplemental Materials), but provided little advantage at the cost of interpretive difficulty, hence we relied on the KNN approach here.

Classification Methods

We performed t-Distributed Stochastic Neighbor Embedding (t-SNE) (van der Maaten & Hinton, 2008) on the training data set for preliminary data exploration and visualization of how well the predictive features are able to separate the class labels. The TSNE method implemented in Scikit-learn was used with the following parameters: perplexity=30, distance=Euclidean.

Competitive Modeling.

We adopted a competitive modeling approach to evaluate the relative utility of different kinds of models, varying in complexity. We first implemented two relatively simple classification methods: unregularized logistic regression (LR-unregularized) and a decision tree limited to two levels (DT-simple), representing a simple algorithm that could easily be implemented in a clinic unaided by machine learning.

These methods were compared to the multi-stage TAN classifier (described below) and five other advanced methods: regularized logistic regression (LR), unrestricted-depth decision tree (DT), random forest (RF), support vector machine (SVM), and gradient boosted decision trees (GBDT). These classifiers were chosen because many have been used previously in similar applications (Bledsoe et al., 2016; Duda et al., 2016; Dvorsky et al., 2016), they span a reasonable range in terms of complexity and interpretability, and all are easily implemented using the popular Scikit-learn Python package (Pedregosa et al., 2011). Finally, we evaluated an ensemble classifier, which made classifications based on the average class probability across all of the six advanced classifiers (TAN, LR, DT, SVM, RF, and GBDT). This was done to test the hypothesis that the ensemble would outperform the individual classifiers, as is sometimes the case in similar applications (Tao et al., 2021).

Bayesian Classifier.

As noted earlier, to attempt to simulate the process of clinical decision-making so that the results would, in theory, be readily adapted to clinical use, we implemented a Bayesian classifier. To do so, we chose a multi-stage tree-augmented naïve Bayes (TAN) classifier. This model was chosen because it models dependencies between the predictive features, which we know exist. This is in contrast with standard naïve Bayes classifiers, which assume independence among predictors. The TAN classifiers were implemented using the bnclassify R package (Mihaljević et al., 2018). This multi-stage approach produces an initial prediction (class probability) based on a small subset of predictors. If a subject’s class probability is below a specified threshold (i.e., low-confidence), additional predictors are added to the classifier and the class probability is updated in subsequent stages. This is illustrated in Figure 1. We conducted a 4-stage approach as follows: Stage 1: parent ADHDRS; stage 2: parent Conners; stage 3: teacher ADHDRS; stage 4: cognitive measures (see Predictive Measures above).

Figure 1.

Figure 1.

Schematic of the multi-stage classification approach. Each step in the process is a tree-augmented naïve Bayes (TAN) classifier. At each stage, a subject’s prior class probability is updated based on a subset of predictors, and a posterior class probability is produced. In (A), each subject is carried through all stages and the final prediction is based on all data. In (B), if a subject’s posterior probability is above a specified threshold (>0.9; i.e., high-confidence) at any Stage, the final prediction is made at that point and subsequent Stages are skipped.

This multi-stage procedure was evaluated in two ways: (1) the final classification for all participants from stage 4 (i.e., all participants are classified using all predictive features; we refer to this as ‘TAN-stage4’; Figure 1A), or (2) treating an individual’s final classification as the earliest high-confidence classification (i.e., some children are classified using only a subset of predictive features; we refer to this as ‘TAN-earliest’; Figure 1B).

All predictive features were discretized before being input into the TAN classifier—that is, the continuous measures were summarized into a finite set of bins. This was necessary because the TAN classifier depends on the calculation of conditional probability tables from the training data. After standardizing and imputing the data, a z score of +/−1.0 was equivalent to a one-standard deviation position above or below the sample mean, roughly approximating what a clinician might use to gauge whether or not to further pursue a problem (Conners, 2003). In this way, the continuous measures were discretized into four bins: (x ≤ −1), (− 1 > x ≤ 0), (0 < x ≤ 1), (x > 1). However, alternative methods were evaluated in sensitivity analyses and results, which were generally supportive of the approach chosen, are included in the Supplemental Materials.

The Bayesian TAN classifier can naturally accommodate a multi-stage approach, by updating the class prior probability at each stage. While other methods could be adapted to a step-wise approach by building independent classifiers with different subsets of features, it is not clear the best way to incorporate information from one step to the next (i.e., the Bayesian logic is not automatic for these other methods). Therefore, we focus on the multi-stage approach for the TAN classifier only. Recognizing this fundamental difference between the methods, we later discuss the relative importance of specific predictive features, and prediction confidence.

General procedure for machine learning models.

A select set of classifier hyperparameters (Supplemental Table S3) were optimized using a grid search with 5-fold cross-validation as implemented in Scikit-learn (Olson et al., 2018). Default values were used for all other parameters. The hyperparameters that resulted in the highest mean cross-validation accuracy were then used for all subsequent classifier comparisons.

The Oregon cohort data set was split into training and test sets as described above (“Data Cleaning / Pre-processing” section) and the performance of each classifier was evaluated in three ways. First, we performed 5-fold cross validation within the Oregon cohort training set (i.e., mean performance across the 5 folds). Second, we trained the classifiers on the Oregon training set and then classified participants in the hold-out (Oregon) test set. Finally, to evaluate external cross-validation and generalizability, we training classifiers on the full Oregon cohort and then classified participants in the independent Michigan cohort. Accuracy, sensitivity (ADHD being the positive class), specificity, positive predictive value (PPV), and area under the receiver operating characteristic curve (AUC-ROC) are reported for all classifiers. PPV was calculated assuming both a 5% prevalence (PPV5) of ADHD in the general population (Song et al., 2019) (as might be seen in a general pediatrics practice) and a 50% prevalence (PPV50) (which may be more representative of the population seen in many outpatient psychiatric settings). Cohen’s kappa, which accounts for agreement by chance, is also reported for additional context.

Differences in the 5-fold cross-validation accuracy among classifiers was tested with a paired T-test. Given we were primarily interested in determining whether the more advanced methods outperformed unregularized logistic regression, we used a p-value threshold of 0.0071 (a Bonferroni correction for seven comparisons: TAN, LR, DT, DT-simple, SVM, RF, and GBDT) to denote significance. Significant differences observed for the 5-fold cross validation performance were confirmed with an additional 2-fold cross validation repeated 5 times (5x2-fold cross validation) (Dietterich, 1998).

In addition, we examined the improvement in classification performance in subsequent stages of the TAN classifier by testing for improved 5-fold cross-validation accuracy, as well as testing for overall improvement in prediction confidence across all 5 cross-validation folds and in the held-out test set (the hypothesis being that the difference between the predicted class probability and the true class probability—i.e., 1 for cases, 0 for controls—would decrease in subsequent stages of the classifier). Paired T-tests were used for both evaluations.

RESULTS

Description of Samples and cohorts

An overview of the datasets used for training and testing the ADHD classifiers in the Oregon-ADHD-1000 and for the generalizability analysis in the Michigan-ADHD-1000 is provided in Table 1. For primary analysis in the Oregon-ADHD-1000 data set, 75% of participants were randomly selected for training the classifiers (and to estimate performance using 5-fold cross validation), and the remaining 25% were used to validate classifier performance in a hold-out sample. Table 1 reveals no statistically significant differences between the training and test data sets in terms of demographic or clinical features. Differences between the Oregon and Michigan cohorts are notable in regard to generalizability and are discussed later.

Table 1.

Overview of samples used in the training and testing of the classifiers. The percentage of subjects with any current diagnosis of anxiety, mood, or disruptive (ODD/CD) disorder. *P-values <0.05 for statistical tests comparing the Michigan-1000 cohort to the entire Oregon-1000 cohort. P-ADHDRS=parent ADHD Rating Scale, T-ADHDRS=teacher ADHD Rating Scale. Family income measures are based on the following scale: 1=less than $25,000, 2=$25,000-$35,000, 3=$35,000-$50,000, 4=$50,000-$75,000, 5=$75,000-$100,000, 6=$100,000-$130,000, 7=$130,000-$150,000, 8=more than $150,000.

Oregon-1000 Michigan-1000
Training Set Test Set
Total N 1067 356 1057
ADHD cases 543 (50.9%) 181 (50.8%) 458 (43.3%)*
Controls 318 (29.8%) 104 (29.2%) 472 (44.7%)*
Other/NOS 206 (19.3%) 71 (19.9%) 127 (12.0%)*
% Males 60.1 62.4 56.5*
Mean Age (SD) 9.4 (1.5) 9.4 (1.7) 12.5 (3.1)*
% White / Caucasian 79.9% 77.0% 72.5%*
Family Income 4.4 (2.0) 4.5 (2.0) 3.8 (1.9)*
P-ADHDRS Int 62.3 (16.5) 61.3 (16.0) 55.0 (14.7)*
P-ADHDRS Hyp 59.2 (15.9) 58.6 (15.9) 51.6 (12.9)*
P-Conners Int 64.5 (16.4) 64.3 (15.8) 60.5 (14.8)*
P-Conners Hyp 63.0 (17.3) 62.2 (16.9) 59.0 (15.4)*
T-ADHDRS Int 51.8 (10.1) 52.1 (10.1) 48.8 (9.8)*
T-ADHDRS Hyp 51.0 (9.3) 51.3 (9.8) 48.3 (8.2)*

As an initial, qualitative examination of how well the predictive features were able to separate diagnostic classes, we visualized the Oregon training set by performing t-SNE (van der Maaten & Hinton, 2008), a non-linear dimensionality reduction technique, using the same subsets of predictive features as will be used in our multi-stage classifier (parent-reported symptom, teacher-reported symptoms, and cognitive measures). Since t-SNE maps the distribution of similarities between pairs of high-dimensional entities to a distribution in a low-dimensional space (here two dimensions), it can be useful for assessing the structure (e.g., presence of clusters) in high-dimensional data. Supplemental Figure S1 shows that parent and teacher ratings alone provided initial meaningful separation between ADHD and non-ADHD groups or clusters. As might be expected, those participants who the best-estimate procedure considered as either subthreshold or NOS appeared more difficult to classify.

Comparison of Classifier Performances on Predicting Best-estimate Diagnosis

To address our primary aims of determining (1) whether machine learning classification algorithms can accurately reproduce gold-standard clinical diagnoses of ADHD, and (2) whether more complex algorithms provide a meaningful benefit over simpler ones, we compared the ability of each classifier to predict the best-estimate team diagnosis in the Oregon cohort. The 5-fold cross-validation and test-set accuracies for all classifiers evaluated on the Oregon cohort are reported in Table 2. The table also includes accuracies on the subset of participants who were given a high-confidence prediction (operationalized as class probability >0.9). Additional performance measures for all classifiers are reported in Supplemental Tables S4 and S5. All classifiers were trained, using all discretized predictive features, to distinguish ADHD cases from non-ADHD controls (the non-ADHD controls included subthreshold and NOS cases to partially mimic the difficulty in clinical diagnosis). Receiver operating characteristic (ROC) curves for all classifiers are shown in Supplemental Figure S2.

Table 2.

Mean 5-fold cross-validation (CV) and test-set accuracies for the classifiers predicting best-estimate team diagnoses in the Oregon cohort. The classifiers included all discretized parent, teacher, and cognitive predictive features. The optimal hyperparameters used for each classifier are given in Supplemental Table S3. The test set is the Oregon 25% hold out sample (N=356). *Accuracies on the subset of the test set with high-confidence predictions (predicted class probability >0.9); the number of high-confidence predictions and the corresponding percentage of the test set are given in parentheses. SD=standard deviation, LR=logistic regression, DT=decision tree, TAN=tree-augmented naïve Bayes, RF=random forest, SVM=support vector machine, GBDT=gradient-boosted decision trees.

Classifier 5-fold CV (SD) Test Set Test Set, High-
confidence*
LR-unregularized 0.851 (0.011) 0.882 0.964 (N=223; 63%)
DT-simple 0.840 (0.011) 0.860 0.936 (N=251; 71%)
TAN-stage4 0.868 (0.014) 0.874 0.908 (N=326; 92%)
TAN-earliest 0.869 (0.015) 0.860 0.870 (N=339; 95%)
LR 0.859 (0.012) 0.876 0.971 (N=208; 58%)
DT 0.870 (0.015) 0.882 0.953 (N=191; 54%)
RF 0.884 (0.005) 0.910 0.969 (N=191; 54%)
SVM 0.877 (0.009) 0.924 0.975 (N=237; 67%)
GBDT 0.890 (0.004) 0.930 0.974 (N=229; 64%)
Ensemble 0.875 (0.013) 0.899 0.968 (N=219; 62%)

All classifiers were able to predict the best-estimate team diagnosis with 5-fold cross-validation accuracy >85% (base rate ~51%; kappa >0.70), except for the simple 2-level decision tree. Gradient boosted decision trees performed significantly better (89% mean 5-fold cross-validation accuracy; kappa = 0.78) than both “simpler” methods, unregularized logistic regression (accuracy = 85%; kappa = 0.70) and the 2-level decision tree (accuracy = 84%, kappa = 0.68). However, there were no significant differences among the more advanced classifiers (Supplemental Table S6). Cross-validation performance was representative of performance on the held-out Oregon test set, with test set accuracies ranging from 87.4% (kappa = 0.75) for TAN to 93% (kappa = 0.86) for gradient boosted decision trees.

For the TAN classifier, using the earliest high-confidence prediction (“TAN-earliest” in Table 2; Figure 1B) resulted in very similar overall performance compared to using the prediction after all 4 stages. However, the accuracy of high-confidence predictions in the test set was slightly lower for TAN-earliest.

Sensitivity analysis: Discrete vs. continuous predictors.

Because the TAN model was the only method that required discretized features, we examined whether discretization diminished the performance of the other methods. The performance of all classifiers, except for TAN, when using the continuous predictive measures, as well as the performance for different discretization methods, is reported in the Supplemental Materials (Supplemental Tables S7 and S8). We saw no significant differences in cross-validation accuracy when using discretized predictors vs. continuous predictors for any of the classification methods.

Classifier Performance Predicting KSADS Diagnosis

Because the best-estimate team used some of the same predictive features as used in the classifiers, there was a certain degree of non-independence between predictors and outcome. To address this, we also evaluated the ability of the classifiers to predict ADHD status defined using KSADS measures (i.e., the labels were determined independently from all predictive features used in the classifier). Agreement between best-estimate team and KSADS ADHD versus non-ADHD status was 85% (see Supplemental Table S9).

The performance of classifiers trained to predict KSADS labels in the Oregon cohort is reported in Supplemental Tables S10, S11, and S12. Overall, classification accuracy of KSADS diagnoses was only slightly lower than seen for the best-estimate team diagnoses. All classifiers achieved mean 5-fold cross-validation accuracy >82% (kappas > 0.65), with the TAN classifier showing the highest cross-validation accuracy (85%, kappa = 0.70) and regularized logistic regression the lowest (83%, kappa = 0.66). There were no significant differences in mean 5-fold cross-validation accuracies across the classification methods (Supplemental Table S13). And again, the 5-fold cross-validation accuracies were fairly representative of performance on the hold out test set, which ranged from 81% accuracy (kappa = 0.63) for the decision tree to 85% accuracy (kappa = 0.70) for gradient-boosted decision trees.

The Value of Confidence Thresholds to Determine Classifications

The clinical utility of a prediction algorithm is likely to be highly dependent on the perceived confidence (or certainty) of a given prediction. Furthermore, a measure of prediction confidence is useful for evaluating whether additional information is needed for adequate assessment. Hence, to again evaluate the value of different classification approaches, we compared the various classifiers based on how often they provided a high-confidence classification. In addition, we further examined the performance of the multi-stage TAN classifier with particular attention paid to the confidence of the classifications at each stage.

As expected, the accuracy of high-confidence predictions is significantly better than lower-confidence predictions. Therefore, we examined how often each classifier provided a high-confidence prediction (predicted class probability >0.9). As shown in the parenthetical numbers in the right-hand column of Table 2, the TAN-stage4 classifier strongly outperformed other methods in providing the most high-confidence predictions (92% of the test set; accuracy of 91% on this subset of high-confidence, test-set predictions), likely due to its ability to incorporate prior probabilities in subsequent stages of the classifier. The proportion of high-confidence classifications for other methods was far lower (Table 2).

In other words, the TAN classifier was able to provide high-confidence predictions for a higher percentage of the sample than the other methods. These results translate into the TAN classifier correctly classifying an additional 61 participants (17% of Oregon test set; 21 true positive ADHD cases) with high confidence, and incorrectly classifying an additional 14 participants (4% of Oregon test set; 7 false positives) with high-confidence, compared to the classifier with the next highest number of high-confidence classifications (DT-simple).

The multi-stage TAN approach allows for classifications to be assigned based on only a subset of the data, if that data produces a high-confidence classification (i.e., a class probability >0.9 in this case). Given the cost of gathering diagnostic data, particularly from laboratory tests, this approach could potentially save significant resources without sacrificing classification performance. Limited resources could be saved and used for the more difficult-to-classify cases.

We examined the proportion of subjects who could be classified with high confidence at each stage of the TAN classifier, as well as the accuracy of these high-confidence classifications. The 5-fold cross validation accuracy significantly improved from stage 2 (all parent ratings) to stage 3 (teacher ratings added) (paired T-test p-value = 0.0027). The improvement from stage1 to stage 2 and from stage 3 to stage 4 were minimal and not statistically significant (p-values of 0.104 and 0.128). However, the confidence of predictions, across all cross-validation folds, did improve—meaning that on average the predicted class probability moved closer to the true class—at all stages of the TAN classifier (all p-values <0.0002).

Supporting the cross-validation results, a large portion of the test-set subjects were classified with high-confidence (class probability >0.9) using only parent ratings (288/356 or 81%, stage 2) and that this portion increases when teacher ratings are added (312/356 or 88%, stage 3). The accuracy of these high-confidence classifications at stages 2 and 3 was 86% and 89% (kappas of 0.67 and 0.76), respectively, in the Oregon hold-out test set.

For the remaining 44 test-set subjects with low-confidence classifications (class probability <0.9) after stage 3 of the TAN classifier, the addition of data from cognitive tests resulted in: (1) the predicted class probability moving closer to the true class for N=35 (80%; 30 of these being correctly classified and 20 becoming high-confidence, correct classifications), (2) the cognitive measures reinforcing the stage-3 prediction for N=26 and contradicting the stage-3 prediction for N=18, and (3) a change in predicted class for N=10 (including 9 correctly classified; 4 with high confidence). Overall, when cognitive measures were included in the classifier (stage 4), the TAN classifier predicted 92% (326/356) of the Oregon test set with high confidence. The accuracy of these high-confidence predictions was 91% (kappa = 0.81).

At the same time, a small number of participants in the Oregon cohort test set were misclassified with high confidence by the TAN classifier (N=30; 8% of the test set). Of the 18 high-confidence false positives, 16 were either subthreshold (N=5) or NOS (N=11), highlighting the diagnostic difficulty of those cases. Of the 12 high-confidence false negatives, 10 had scores of 1=“minimal” on the KSADS for ADHD current impairment (scores range from 0=“none” to 3=“severe”; the other two had scores of 2=“moderate”), and all had a parent-reported SDQ impact score of 0 (0=“normal”, 1=“borderline”, 2=“abnormal”). In these cases, the best-estimate team likely made judgements about informant reliability or drew upon narrative notes from the clinical interviewer.

Class probabilities, at all stages of the TAN classifier, for all subjects in the Oregon test set are shown in Figure 2. It shows that (1) most participants were classified with high confidence (probability of ADHD ≥0.9) at Stage 3, prior to incorporating the cognitive measures; (2) a small subset of participants were never classified with high confidence, even when all predictors were used in the classifier; and (3) some participants have conflicting data at different stages—their class probabilities are highly variable, or “unstable”, across the stages (see further details in the sensitivity analysis reported below).

Figure 2.

Figure 2.

Predicted probability of an ADHD diagnosis (best-estimate team), at each stage of the TAN classifier for all participants in the Oregon test set. Most ADHD cases converge correctly towards a probability of ADHD of 1.0 across the 4 stages. Likewise, most controls converge correctly towards a probability of ADHD of 0. The majority of participants are predicted with high confidence (i.e., a predicted probability >0.9 or <0.1; dashed horizontal lines), by stage 3. Yet a small subset of participants had unstable predictions (large changes in predicted probability in subsequent stages)—these are significantly more likely to be other/NOS or subthreshold individuals (see Supplemental Materials). Stage 1 = parent ADHD Rating Scale, Stage 2 = parent Conners, Stage 3 = teacher ADHD Rating Scale, and Stage 4 = cognitive measures.

Sensitivity analysis: Impact of NOS or borderline classified participants.

Given the diagnostic uncertainty implied by the subthreshold and NOS categories, we investigated the effect of re-labeling or excluding these subjects from the data used to train the classifiers. The performance of classifiers trained on the full training data, with subthreshold and NOS participants all labeled as non-ADHD, is referred to as Training Scheme A. It was compared to the performance of two other training schemes. The results are displayed in Table 3. In Training Scheme B, the full training data set was used, but subthreshold cases were selected and re-labeled as ADHD cases. In Training Scheme C, all subthreshold and NOS participants were excluded from the training set. Scheme B led to reduced accuracy (Table 3). Scheme C resulted, not surprisingly, in slightly increased sensitivity (test-set sensitivity 0.939 vs. 0.917), but reduced specificity (test-set specificity 0.760 vs. 0.829).

Table 3.

Comparison of TAN classifier performance with three different training schemes. Training scheme A includes all training set subjects, with subthreshold and not-clean controls labeled as controls; Training scheme B re-labels subthreshold subjects as ADHD cases; and Training scheme C excludes subthreshold and not-clean control subjects from the training set (but they are retained in the test set). *CV performance measures are the average of measures across all 5-folds of the training set (average N=213). The full independent test set consists of N=356 subjects. Performance is also reported for only those subjects given a high-confidence (P ≥ 0.9) prediction (N=326, 323, and 336 for the three training schemes, respectively). Bolded values are the highest among the three training schemes for the test set, while italicized values are the highest among the three training schemes for CV. PPVx=positive predictive value assuming X% prevalence of ADHD.

Scheme A Scheme B Scheme C
Full CV/Test Set* 5-fold CV Test 5-fold CV Test 5-fold CV Test
Accuracy 0.868 0.874 0.866 0.854 0.857 0.851
Sensitivity 0.912 0.917 0.900 0.900 0.959 0.939
Specificity 0.823 0.829 0.818 0.805 0.750 0.760
PPV5 0.213 0.220 0.206 0.194 0.168 0.171
PPV50 0.837 0.843 0.832 0.820 0.793 0.796
AUC-ROC 0.947 0.942 0.937 0.921 0.910 0.915
High-confidence
Predictions Only
N=972 N=326 N=986 N=323 N=1017 N=336
Accuracy 0.903 0.908 0.886 0.885 0.870 0.872
Sensitivity 0.943 0.930 0.923 0.910 0.973 0.960
Specificity 0.862 0.884 0.835 0.851 0.761 0.778
PPV5 0.264 0.296 0.227 0.243 0.176 0.185
PPV50 0.872 0.889 0.848 0.859 0.803 0.812
AUC-ROC 0.959 0.950 0.949 0.934 0.913 0.918

In addition, we show that participants labeled as subthreshold or NOS are indeed more difficult to predict. For instance, participants who had unstable predictions across the four stages of the classifier (i.e., their predicted class switched from case to control or vice versa) were significantly more likely to be labeled as subthreshold or NOS (odds ratio = 4.29, Fisher’s Exact p-value = 0.000151) compared to cases and controls, whose class probabilities tended to converge to 1 or 0, respectively, across the stages. More details about the classification performance for participants with uncertain diagnoses is given in the Supplemental Materials.

Sensitivity analysis: The value of using impairment scores as an alternative to cognitive test measures.

As noted above, for high-confidence misclassifications, impairment scores tended to align with predictions in the majority of instances (e.g., 14 out of 18 high-confidence false positives had KSADS current impairment scores of “moderate” or “severe”, while 10 out of 12 high-confidence false negatives had KSADS current impairment scores of “minimal”). For participants with low-confidence predications after stage 3, we examined whether impairment scores would have been informative, given that these scores are easier to obtain than the cognitive tests. Of the 44 test-set participants with low-confidence predictions, 25 (57%) had impairment scores consistent with their true diagnosis (e.g., KSADS scores of “none” or “minimal” for controls, and scores of “moderate” or “severe” for ADHD cases). In contrast, 80% of these same participants benefited from the cognitive measures, in terms of prediction confidence, as described above.

Generalization to an Independent Cohort

Finally, we evaluated whether a classifier trained on the entire Oregon data set would perform well on the independent Michigan cohort. This test of generalizability is an essential part of algorithm evaluation before considering potential applications. As in the primary analyses done in the Oregon cohort, subjects categorized as subthreshold or NOS were labeled as controls for the purposes of training the classifiers in the full Oregon cohort. Because a different version of the Conners rating scale (with fewer subscales) was used in the Michigan cohort (due to its earlier period of collection), new classifiers were trained on the Oregon data excluding the subscales that were not available in the Michigan data.

Several differences between the two cohorts were expected to challenge the generalizability of the model, but also provided an important and beneficial opportunity to examine real-world generalizability/reproducibility questions (Table 1). Notably, the Michigan cohort included a significantly wider age range, and higher average age (12.5 vs. 9.4 years), a lower proportion of ADHD cases, a lower proportion of males, lower average family income, and higher proportion of non-white individuals. Finally, the overall distribution of symptom measures is shifted further into the clinical range in the Michigan cohort compared to the Oregon cohort—this difference remains when examining ADHD cases only (data not shown), suggesting the Michigan cohort is a somewhat more clinically severe ADHD sample.

Thus, it was encouraging that despite these differences, most classifiers showed impressive generalizability to the Michigan cohort, for classifying best-estimate team diagnosis. Table 4 shows these results. The TAN, RF, SVM, and GBDT methods were all able to classify ADHD and non-ADHD (including subthreshold and NOS) with >79% accuracy (kappas >0.56). Again, there appears to be a slight advantage to the more advanced classifiers, with GBDT performing best (82% accuracy, kappa = 0.63) and the 2-level decision tree worst (75% accuracy, kappa = 0.48) in the Michigan cohort. As seen in the Oregon cohort, the TAN provided considerably more high-confidence predictions (91%) than other methods (41%-69%).

Table 4.

Accuracy of the classifiers predicting best-estimate team diagnoses in the Michigan cohort. Also show (last column) are the accuracies for the subset of Michigan cohort subjects given high-confidence predictions (predicted class probability >0.9); the number of high-confidence predictions and the corresponding percentage of the Michigan cohort are given in parentheses. The classifiers were trained on data from the entire Oregon cohort, using all available (discretized) parent, teacher, and cognitive predictive features.

Classifier Training Set
(Oregon; N=1423)
Test Set
(Michigan;
N=1057)
Test Set (Michigan;
High-confidence)
LR-unregularized 0.878 0.779 0.895 (N=636; 60%)
DT-simple 0.850 0.761 0.879 (N=725; 69%)
TAN-stage4 0.877 0.792 0.820 (N=957; 91%)
LR 0.874 0.785 0.906 (N=585; 55%)
DT 0.878 0.770 0.917 (N=435; 41%)
RF 0.940 0.807 0.932 (N=474; 45%)
SVM 0.916 0.799 0.902 (N=673; 64%)
GBDT 0.909 0.822 0.909 (N=635; 60%)
Ensemble 0.919 0.807 0.916 (N=585; 55%)

Again, we observed slightly poorer performance across all methods when classifying KSADS diagnoses. The TAN performed best in that case, with an accuracy of 79% (kappa = 0.57) (Supplemental Tables S15 - S18). Overall, generalizability to a new cohort was supportive of the potential for these algorithms to achieve useful accuracy.

DISCUSSION

Our strategy was to use a competitive modeling approach to assess the ability of a variety of machine learning classification methods, ranging in complexity, to reproduce clinician best-estimate diagnosis of ADHD. A primary question was whether more complex algorithms provided a meaningful advantage over simpler ones. Among the methods evaluated was a multi-stage Bayesian approach (tree-augmented naïve Bayes classifier; TAN), which was chosen to reasonably simulate the sequential decision-making process of a clinician. All methods were evaluated in terms of both the accuracy and confidence of the given classifications.

Classification of gold-standard research diagnosis of ADHD was achieved with fairly high accuracy overall (>85%). In terms of classification accuracy, advanced methods appear to provided only incremental improvement over unregularized logistic regression and a simple (2-level) decision tree. However, this slight benefit diminished when predicting KSADS diagnosis, and when examining generalizability to a completely independent cohort. Likewise, overall performance was statistically similar across the more advanced classifiers, including the TAN.

At same time, however, because the TAN classifier was readily able to incorporate prior probabilities from early-stage predictions into the later stages, it provided high-confidence predictions for many more subjects than other classifiers. This may be a particularly advantageous feature of the TAN, given that prediction confidence will be an important factor for determining clinical utility. In the Oregon held-out test set (N=356), the high-confidence predictions from the TAN classifier resulted in an additional 40 high-confidence true negatives, 21 high-confidence true positives, 7 high-confidence false positives, and 7 high-confidence false negatives, compared to other methods. In other words, an additional 17% of the sample was correctly classified with high confidence, with only a 4% increase in high-confidence misclassifications.

Although overall accuracies were comparable across the more advanced classification methods, the Bayesian nature of the TAN classifier may align better with clinical decision-making processes, increasing interpretability and potential clinical utility. The multi-stage TAN classifier demonstrated performance comparable to other advanced methods, generalized well, and provided evidence that a large majority of subjects can be classified accurately with only low-cost parent and teacher symptom reports, even with very difficult borderline cases in the sample.

Our findings clarify the value of further cognitive assessment to aid diagnosis of ADHD in a small subset of cases. Thus, the data support the general best practice that cognitive testing is not diagnostic for ADHD. However, an important caveat is added here. The addition of data from laboratory cognitive tests provided a benefit, in terms of prediction confidence, for only a small subset of cases. Specifically, of the participants with low-confidence classifications based on parent and teacher measures alone, 68% (or 8% of the total test sample) benefited from further data from cognitive tests—meaning their classifications improved in confidence and were correct. Our findings are largely consistent with results from a smaller study of younger children (Oztekin et al., 2021). Though the benefit we observed is relatively small on the population level, it should not be dismissed outright, as incremental improvements may have meaningful impact for individual patients. Identifying those subjects who will benefit from further assessment (of all kinds), as well as identification of other low-cost/high-value predictors, should be a priority for future work.

The posterior class probability is an intuitive way to judge the confidence of a classification. Here we used a class probability cutoff of 0.9 to define a “high-confidence” prediction, but there is no firm guidance and this criterion was ultimately arbitrary. The optimal threshold to use for a particular classifier and outcome is still an open question, and should be studied further with the potential costs of misclassification in mind. Examining class probability as a measure of prediction confidence provided insight into those participants who were particularly difficult to classify. Not surprisingly, participants who were difficult to predict (those with low confidence predictions or unstable predictions across the stages of the multi-stage classifier) were also more likely to be given uncertain or “edge” diagnoses (subthreshold or other/NOS) by the best-estimate clinician team. Those with borderline symptoms or with only one informant reporting high symptoms are difficult to accurately evaluate. The unstable predictions suggests that expert judgement in those situations is not able to be matched by machine algorithm based only on informant reports, and further suggests that these instances can benefit from further assessments to obtain additional data to add accuracy and confidence to clinician efforts. A difficult point here is that many subthreshold cases are still accompanied by substantial impairment and may require treatment (Kirova et al., 2019), a point which mandates clinical judgement as well. Here, we observed some benefit from the addition of objective cognitive measures for a subset of those cases. However, there may be other types of data that would be informative in computational algorithms for this subset of difficult-to-classify individuals (e.g., child self-report data) that we did not evaluate.

Another conclusion suggested by our findings is that a training data set representing the full spectrum of ADHD presentations, including subthreshold cases and non-ADHD subjects with other psychiatric conditions, may be important for reducing false-positive classifications. A full spectrum provides better granularity which might help better defined decision boundaries for the ADHD class. This finding should be confirmed with additional data, given that the impact of uncertain or edge diagnoses may change as the size of training sets increase.

The current study has several important strengths, including: a larger sample size than most prior attempts at single case prediction in ADHD using machine learning, the inclusion of a completely independent validation cohort for independent generalizability (rare in this literature), the inclusion of difficult borderline diagnostic cases, a competitive modeling approach, and the examination of the relative importance of predictive measures in a multi-stage classification approach. Each of these provides a valuable contribution to the literature and in general, results are promising.

Certain limitations do bear mention. First, despite the relatively large samples used for training and evaluating the classifiers, and the availability of a generalizability cohort, substantial further work would be needed to evaluate generalizability and site-specific algorithms for clinical use. Second, the large amount of missing data among the cognitive measures, and the confound between diagnostic category and availability of these data, limited our ability to make strong conclusions about the utility of those measures in ambiguous cases. Best practices already specify that cognitive measures are not diagnostic for ADHD. That said, we still found that when those data were present, they did supplement diagnostic accuracy in a small subset of cases, suggesting that a minority of cases may benefit from these additional data before diagnosis is assigned. This is not to say that it should always be collected for diagnostic purposes, but rather that future work may identify those fairly rare situations when its collection would increase diagnostic confidence. (We note that these data may still be clinically useful for other purposes—characterizing frequently comorbid learning problems or understanding functional impairment.) Third, our reliance on a method (TAN) that required discretized measures could have resulted in loss of information, and therefore could have hurt performance—although our sensitivity analysis did not suggest this was the case. That symptom thresholds are typically used for diagnosis suggests that discretization of measures is appropriate (and possibly beneficial, due to removing redundant information) for the current application. The sensitivity analyses conducted, regarding missing data and discretization, lend confidence to our conclusions. And finally, though this applies to all diagnostic tools and not just the computational approaches we explored, it should be noted that, in practice, consideration should be given to the cost of different types of diagnostic errors (i.e., false positives vs. false negatives) and predictive performance should be tuned to maximize benefit and minimize potential harm.

To take full advantage of computational approaches like those reported here, future research will need to address the cost/benefit of implementing these approaches in a clinic, when the use of such tools is most appropriate, and the factors that influence clinician acceptance/resistance to their use.

In conclusion, the findings reported here point to the potential utility of machine learning algorithms to aid in the accuracy and confidence of identification of ADHD with data limited to rating scales supplemented, in a subset of cases, by additional testing. Future research should focus on (1) further validation of classifiers in larger and more diverse samples, including investigating predictive ability across a wider age range, (2) identification of difficult-to-classify subjects and those who will benefit from additional, more costly assessment (including, e.g., genetic risk scores), and (3) evaluation of the cost/benefit of prediction confidence thresholds that could guide clinical deployment of advanced classification algorithms.

Supplementary Material

1

ACKNOWLEDGEMENTS

This work was supported by funding from the National Institute of Mental Health (R37MH059105).

Footnotes

PRIOR VERSIONS

A prior version of this manuscript has been posted to the medRxiv preprint server: https://doi.org/10.1101/2021.12.23.21268330

REFERENCES

  1. Abu-Mostafa YS, Magdon-Ismail M, & Lin H-T (2012). Learning From Data. AMLBook. [Google Scholar]
  2. Adler L, Shaw D, Sitt D, Maya E, & Morrill MI (2009). Issues in the diagnosis and treatment of adults ADHD by primary care physicians. Primary Psychiatry, 16(5), 57–63. [Google Scholar]
  3. APA. (2013). Diagnostic and statistical manual of mental disorders: DSM-5™, 5th ed (pp. xliv, 947). American Psychiatric Publishing, Inc. 10.1176/appi.books.9780890425596 [DOI] [Google Scholar]
  4. Bishop CM (2006). Pattern Recognition and Machine Learning. Springer-Verlag. https://www.springer.com/gp/book/9780387310732 [Google Scholar]
  5. Bledsoe JC, Xiao D, Chaovalitwongse A, Mehta S, Grabowski TJ, Semrud-Clikeman M, Pliszka S, & Breiger D (2016). Diagnostic Classification of ADHD Versus Control: Support Vector Machine Classification Using Brief Neuropsychological Assessment. Journal of Attention Disorders, 1087054716649666. 10.1177/1087054716649666 [DOI] [PubMed] [Google Scholar]
  6. Cella D, Yount S, Rothrock N, Gershon R, Cook K, Reeve B, Ader D, Fries JF, Bruce B, Rose M, & PROMIS Cooperative Group. (2007). The Patient-Reported Outcomes Measurement Information System (PROMIS): Progress of an NIH Roadmap cooperative group during its first two years. Medical Care, 45(5 Suppl 1), S3–S11. 10.1097/01.mlr.0000258615.42478.55 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Christiansen H, Chavanon M-L, Hirsch O, Schmidt MH, Meyer C, Müller A, Rumpf H-J, Grigorev I, & Hoffmann A (2020). Use of machine learning to classify adult ADHD and other conditions based on the Conners’ Adult ADHD Rating Scales. Scientific Reports, 10(1), 18871. 10.1038/s41598-020-75868-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Committee on Quality Improvement, Subcommittee on Attention-Deficit/Hyperactivity Disorder. (2000). Clinical practice guideline: Diagnosis and evaluation of the child with attention-deficit/hyperactivity disorder. American Academy of Pediatrics. Pediatrics, 105(5), 1158–1170. 10.1542/peds.105.5.1158 [DOI] [PubMed] [Google Scholar]
  9. Conners CK (2003). Connors’ rating scales: Revised technical manual. Multi-Health Systems. [Google Scholar]
  10. Conners CK (2008). Conners 3rd edition manual. Multi-Health Systems. [Google Scholar]
  11. Conners CK, Sitarenios G, Parker JDA, & Epstein JN (1998). The Revised Conners’ Parent Rating Scale (CPRS-R): Factor Structure, Reliability, and Criterion Validity. Journal of Abnormal Child Psychology, 26(4), 257–268. 10.1023/A:1022602400621 [DOI] [PubMed] [Google Scholar]
  12. Costello EJ, He J, Sampson NA, Kessler RC, & Merikangas KR (2014). Services for adolescents with psychiatric disorders: 12-month data from the National Comorbidity Survey-Adolescent. Psychiatric Services (Washington, D.C.), 65(3), 359–366. 10.1176/appi.ps.201100518 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. De Luca CR, Wood SJ, Anderson V, Buchanan J-A, Proffitt TM, Mahony K, & Pantelis C (2003). Normative Data From the Cantab. I: Development of Executive Function Over the Lifespan. Journal of Clinical and Experimental Neuropsychology, 25(2), 242–254. 10.1076/jcen.25.2.242.13639 [DOI] [PubMed] [Google Scholar]
  14. Delis DC, Kaplan E, & Kramer JH (2001). Delis-Kaplan Executive Function System. Psychological Corporation. [Google Scholar]
  15. Dietterich TG (1998). Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms. Neural Computation, 10(7), 1895–1923. 10.1162/089976698300017197 [DOI] [PubMed] [Google Scholar]
  16. Duda M, Ma R, Haber N, & Wall DP (2016). Use of machine learning for behavioral distinction of autism and ADHD. Translational Psychiatry, 6(2), e732–e732. 10.1038/tp.2015.221 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. DuPaul G, Power T, Anastopoulos A, & Reid R (1998). ADHD Rating Scale—IV: Checklists, Norms, and Clinical Interpretation. Guilford Press. [Google Scholar]
  18. Dvorsky MR, Langberg JM, Molitor SJ, & Bourchtein E (2016). Clinical Utility and Predictive Validity of Parent and College Student Symptom Ratings in Predicting an ADHD Diagnosis. Journal of Clinical Psychology, 72(4), 401–418. 10.1002/jclp.22268 [DOI] [PubMed] [Google Scholar]
  19. Fabiano F, & Haslam N (2020). Diagnostic inflation in the DSM: A meta-analysis of changes in the stringency of psychiatric diagnosis from DSM-III to DSM-5. Clinical Psychology Review, 80, 101889. 10.1016/j.cpr.2020.101889 [DOI] [PubMed] [Google Scholar]
  20. Faraone SV, Spencer TJ, Montano CB, & Biederman J (2004). Attention-deficit/hyperactivity disorder in adults: A survey of current practice in psychiatry and primary care. Archives of Internal Medicine, 164(11), 1221–1226. 10.1001/archinte.164.11.1221 [DOI] [PubMed] [Google Scholar]
  21. French B, Sayal K, & Daley D (2019). Barriers and facilitators to understanding of ADHD in primary care: A mixed-method systematic review. European Child & Adolescent Psychiatry, 28(8), 1037–1064. 10.1007/s00787-018-1256-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Gibbons RD, Kupfer DJ, Frank E, Lahey BB, George-Milford BA, Biernesser CL, Porta G, Moore TL, Kim JB, & Brent DA (2020). Computerized Adaptive Tests for Rapid and Accurate Assessment of Psychopathology Dimensions in Youth. Journal of the American Academy of Child and Adolescent Psychiatry, 59(11), 1264–1273. 10.1016/j.jaac.2019.08.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Gibbons RD, Weiss DJ, Frank E, & Kupfer D (2016). Computerized Adaptive Diagnosis and Testing of Mental Health Disorders. Annual Review of Clinical Psychology, 12, 83–104. 10.1146/annurev-clinpsy-021815-093634 [DOI] [PubMed] [Google Scholar]
  24. Gill CJ, Sabin L, & Schmid CH (2005). Why clinicians are natural bayesians. BMJ : British Medical Journal, 330(7499), 1080–1083. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Goodfellow I, Bengio Y, & Courville A (2016). Deep Learning. MIT Press. [Google Scholar]
  26. Hall CL, Valentine AZ, Groom MJ, Walker GM, Sayal K, Daley D, & Hollis C (2016). The clinical utility of the continuous performance test and objective measures of activity for diagnosing and monitoring ADHD in children: A systematic review. European Child & Adolescent Psychiatry, 25(7), 677–699. 10.1007/s00787-015-0798-x [DOI] [PubMed] [Google Scholar]
  27. Hamed AM, Kauer AJ, & Stevens HE (2015). Why the Diagnosis of Attention Deficit Hyperactivity Disorder Matters. Frontiers in Psychiatry, 6, 168. 10.3389/fpsyt.2015.00168 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Irwin DE, Stucky B, Langer MM, Thissen D, Dewitt EM, Lai J-S, Varni JW, Yeatts K, & DeWalt DA (2010). An item response analysis of the pediatric PROMIS anxiety and depressive symptoms scales. Quality of Life Research: An International Journal of Quality of Life Aspects of Treatment, Care and Rehabilitation, 19(4), 595–607. 10.1007/s11136-010-9619-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Jerez JM, Molina I, García-Laencina PJ, Alba E, Ribelles N, Martín M, & Franco L (2010). Missing data imputation using statistical and machine learning methods in a real breast cancer problem. Artificial Intelligence in Medicine, 50(2), 105–115. 10.1016/j.artmed.2010.05.002 [DOI] [PubMed] [Google Scholar]
  30. Karalunas SL, Gustafsson HC, Dieckmann NF, Tipsord J, Mitchell SH, & Nigg JT (2017). Heterogeneity in development of aspects of working memory predicts longitudinal attention deficit hyperactivity disorder symptom change. Journal of Abnormal Psychology, 126(6), 774–792. 10.1037/abn0000292 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Kazda L, Bell K, Thomas R, McGeechan K, Sims R, & Barratt A (2021). Overdiagnosis of Attention-Deficit/Hyperactivity Disorder in Children and Adolescents: A Systematic Scoping Review. JAMA Network Open, 4(4), e215335. 10.1001/jamanetworkopen.2021.5335 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Kirova A-M, Kelberman C, Storch B, DiSalvo M, Woodworth KY, Faraone SV, & Biederman J (2019). Are subsyndromal manifestations of attention deficit hyperactivity disorder morbid in children? A systematic qualitative review of the literature with meta-analysis. Psychiatry Research, 274, 75–90. 10.1016/j.psychres.2019.02.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Loskutova NY, Lutgen CB, Callen EF, Filippi MK, & Robertson EA (2021). Evaluating a Web-Based Adult ADHD Toolkit for Primary Care Clinicians. Journal of the American Board of Family Medicine: JABFM, 34(4), 741–752. 10.3122/jabfm.2021.04.200606 [DOI] [PubMed] [Google Scholar]
  34. Marek S, Tervo-Clemmens B, Calabro FJ, Montez DF, Kay BP, Hatoum AS, Donohue MR, Foran W, Miller RL, Hendrickson TJ, Malone SM, Kandala S, Feczko E, Miranda-Dominguez O, Graham AM, Earl EA, Perrone AJ, Cordova M, Doyle O, … Dosenbach NUF (2022). Reproducible brain-wide association studies require thousands of individuals. Nature, 603(7902), 654–660. 10.1038/s41586-022-04492-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Martin AR, Daly MJ, Robinson EB, Hyman SE, & Neale BM (2019). Predicting polygenic risk of psychiatric disorders. Biological Psychiatry, 86(2), 97–109. 10.1016/j.biopsych.2018.12.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Martin AR, Kanai M, Kamatani Y, Okada Y, Neale BM, & Daly MJ (2019). Current clinical use of polygenic scores will risk exacerbating health disparities. Nature Genetics, 51(4), 584–591. 10.1038/s41588-019-0379-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Massuti R, Moreira-Maia CR, Campani F, Sônego M, Amaro J, Akutagava-Martins GC, Tessari L, Polanczyk GV, Cortese S, & Rohde LA (2021). Assessing undertreatment and overtreatment/misuse of ADHD medications in children and adolescents across continents: A systematic review and meta-analysis. Neuroscience and Biobehavioral Reviews, 128, 64–73. 10.1016/j.neubiorev.2021.06.001 [DOI] [PubMed] [Google Scholar]
  38. Mihaljević B, Bielza C, & Larrañaga P (2018). bnclassify: Learning Bayesian Network Classifiers. The R Journal, 10(2), 455–468. [Google Scholar]
  39. Mooney MA, Bhatt P, Hermosillo RJM, Ryabinin P, Nikolas M, Faraone SV, Fair DA, Wilmot B, & Nigg JT (2021). Smaller total brain volume but not subcortical structure volume related to common genetic risk for ADHD. Psychological Medicine, 51(8), 1279–1288. 10.1017/S0033291719004148 [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Mooney MA, Ryabinin P, Wilmot B, Bhatt P, Mill J, & Nigg JT (2020). Large epigenome-wide association study of childhood ADHD identifies peripheral DNA methylation associated with disease and polygenic risk burden. Translational Psychiatry, 10(1), 1–12. 10.1038/s41398-020-0710-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Mueller A, Candrian G, Kropotov JD, Ponomarev VA, & Baschera G-M (2010). Classification of ADHD patients on the basis of independent ERP components using a machine learning system. Nonlinear Biomedical Physics, 4(1), S1. 10.1186/1753-4631-4-S1-S1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. National Guideline Centre (UK). (2018). Attention deficit hyperactivity disorder: Diagnosis and management. National Institute for Health and Care Excellence (UK). http://www.ncbi.nlm.nih.gov/books/NBK493361/ [PubMed] [Google Scholar]
  43. Nigg JT (1999). The ADHD response-inhibition deficit as measured by the stop task: Replication with DSM-IV combined type, extension, and qualification. Journal of Abnormal Child Psychology, 27(5), 393–402. 10.1023/a:1021980002473 [DOI] [PubMed] [Google Scholar]
  44. Nigg JT, Gustafsson HC, Karalunas SL, Ryabinin P, McWeeney SK, Faraone SV, Mooney MA, Fair DA, & Wilmot B (2018). Working Memory and Vigilance as Multivariate Endophenotypes Related to Common Genetic Risk for Attention-Deficit/Hyperactivity Disorder. Journal of the American Academy of Child and Adolescent Psychiatry, 57(3), 175–182. 10.1016/j.jaac.2017.12.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Nigg JT, Karalunas SL, Gustafsson HC, Bhatt P, Ryabinin P, Mooney MA, Faraone SV, Fair DA, & Wilmot B (2020). Evaluating chronic emotional dysregulation and irritability in relation to ADHD and depression genetic risk in children with ADHD. Journal of Child Psychology and Psychiatry, 61(2), 205–214. 10.1111/jcpp.13132 [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Nikolas MA, Marshall P, & Hoelzle JB (2019). The role of neurocognitive tests in the assessment of adult attention-deficit/hyperactivity disorder. Psychological Assessment, 31(5), 685–698. 10.1037/pas0000688 [DOI] [PubMed] [Google Scholar]
  47. Nikolas MA, & Nigg JT (2013). Neuropsychological performance and attention-deficit hyperactivity disorder subtypes and symptom dimensions. Neuropsychology, 27(1), 107–120. 10.1037/a0030685 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Nikolas MA, & Nigg JT (2015). Moderators of Neuropsychological Mechanism in Attention- Deficit Hyperactivity Disorder. Journal of Abnormal Child Psychology, 43(2), 271–281. 10.1007/s10802-014-9904-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Olson RS, La Cava W, Mustahsan Z, Varik A, & Moore JH (2018). Data-driven advice for applying machine learning to bioinformatics problems. Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing, 23, 192–203. [PMC free article] [PubMed] [Google Scholar]
  50. Öztekin I, Finlayson MA, Graziano PA, & Dick AS (2021). Is there any incremental benefit to conducting neuroimaging and neurocognitive assessments in the diagnosis of ADHD in young children? A machine learning investigation. Developmental Cognitive Neuroscience, 49, 100966. 10.1016/j.dcn.2021.100966 [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Palk AC, Dalvie S, de Vries J, Martin AR, & Stein DJ (2019). Potential use of clinical polygenic risk scores in psychiatry – ethical implications and communicating high polygenic risk. Philosophy, Ethics, and Humanities in Medicine : PEHM, 14. 10.1186/s13010-019-0073-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, & Duchesnay É (2011). Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 12(85), 2825–2830. [Google Scholar]
  53. Rashid B, & Calhoun V (2020). Towards a brain-based predictome of mental illness. Human Brain Mapping, 41(12), 3468–3535. 10.1002/hbm.25013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Reitan RM, & Wolfson D (1985). The Halstead-Reitan neuropsychological test battery: Theory and clinical interpretation. Neuropsychology Press. [Google Scholar]
  55. Ronald A, Bode N. de, & Polderman TJC (2021). Systematic Review: How the Attention-Deficit/Hyperactivity Disorder Polygenic Risk Score Adds to Our Understanding of ADHD and Associated Traits. Journal of the American Academy of Child & Adolescent Psychiatry, 0(0). 10.1016/j.jaac.2021.01.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Schachar R, Tannock R, Marriott M, & Logan G (1995). Deficient inhibitory control in attention deficit hyperactivity disorder. Journal of Abnormal Child Psychology, 23(4), 411–437. 10.1007/BF01447206 [DOI] [PubMed] [Google Scholar]
  57. Simon AE, Pastor PN, Reuben CA, Huang LN, & Goldstrom ID (2015). Use of Mental Health Services by Children Ages Six to 11 With Emotional or Behavioral Difficulties. Psychiatric Services (Washington, D.C.), 66(9), 930–937. 10.1176/appi.ps.201400342 [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Slobodin O, Yahav I, & Berger I (2020). A Machine-Based Prediction Model of ADHD Using CPT Data. Frontiers in Human Neuroscience, 14, 383. 10.3389/fnhum.2020.560021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Song M, Dieckmann NF, & Nigg JT (2019). Addressing Discrepancies Between ADHD Prevalence and Case Identification Estimates Among U.S. Children Utilizing NSCH 2007-XSXS2012. Journal of Attention Disorders, 23(14), 1691–1702. 10.1177/1087054718799930 [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Tao X, Chi O, Delaney PJ, Li L, & Huang J (2021). Detecting depression using an ensemble classifier based on Quality of Life scales. Brain Informatics, 8(1), 2. 10.1186/s40708-021-00125-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. van der Maaten LJP, & Hinton GE (2008). Visualizing High-Dimensional Data Using t-SNE. Journal of Machine Learning Research, 9(nov), 2579–2605. [Google Scholar]
  62. Wechsler D (2003). Wechsler intelligence scale for children – Fourth edition (WISC-IV). The Psychological Corporation. [Google Scholar]
  63. Wodka EL, Loftis C, Mostofsky SH, Prahme C, Larson JCG, Denckla MB, & Mahone EM (2008). Prediction of ADHD in boys and girls using the D-KEFS. Archives of Clinical Neuropsychology, 23(3), 283–293. 10.1016/j.acn.2007.12.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Young S, Asherson P, Lloyd T, Absoud M, Arif M, Colley WA, Cortese S, Cubbin S, Doyle N, Morua SD, Ferreira-Lay P, Gudjonsson G, Ivens V, Jarvis C, Lewis A, Mason P, Newlove-Delgado T, Pitts M, Read H, … Skirrow C (2021). Failure of Healthcare Provision for Attention-Deficit/Hyperactivity Disorder in the United Kingdom: A Consensus Statement. Frontiers in Psychiatry, 12. https://www.frontiersin.org/article/10.3389/fpsyt.2021.649399 [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Youngstrom EA, Halverson TF, Youngstrom JK, Lindhiem O, & Findling RL (2018). Evidence-Based Assessment from Simple Clinical Judgments to Statistical Learning: Evaluating a Range of Options Using Pediatric Bipolar Disorder as a Diagnostic Challenge. Clinical Psychological Science: A Journal of the Association for Psychological Science, 6(2), 243–265. 10.1177/2167702617741845 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

RESOURCES