Abstract
This manuscript illustrates novel quantitative methods to estimate classification consistency in machine learning models used for screening measures. Screening measures are used in psychology and medicine to classify individuals into diagnostic classifications. In addition to achieving high accuracy, it is ideal for the screening process to have high classification consistency, which means that respondents would be classified into the same group every time if the assessment was repeated. Although machine learning models are increasingly being used to predict a screening classification based on individual item responses, methods to describe the classification consistency of machine learning models have not yet been developed. This paper addresses this gap by describing methods to estimate classification inconsistency in machine learning models arising from two different sources: sampling error during model fitting or measurement error in the item responses. These methods use data resampling techniques such as the bootstrap and Monte Carlo sampling. These methods are illustrated using three empirical examples predicting a health condition/diagnosis from item responses. R code is provided to facilitate the implementation of the methods. The paper highlights the importance of considering classification consistency alongside accuracy when studying screening measures and provides the tools and guidance necessary for applied researchers to obtain classification consistency indices in their machine learning research on diagnostic assessments.
Keywords: screening, machine learning, classification consistency, reliability
Estimating classification consistency of machine learning models with screening measures
Assessments are typically used in the areas of education, psychology, and medicine to classify individuals. Examples include the Psychopathology Checklist-Revised (PCL-R), which is a measure used to screen for psychopathy in a prison population (Hare, 2003), or the Geriatric Depression Scale (GDS), which is a measure used to identify older adults with depression (Sheikh & Yesavage, 1986). In this paper, we focus on screening assessments, although the themes and findings discussed generalize to other classification scenarios.
In the screening process, it is important that the decisions about the respondents be accurate and consistent. A decision is accurate when the screening procedure correctly identifies respondents who meet the criterion of interest (e.g., have a diagnosis). For example, if an individual has depression, the screening procedure should classify the individual into the depressed category. Indices for accuracy include classification accuracy, sensitivity, and specificity (Youngstrom, 2013). A decision is consistent when the screening procedure would classify a respondent into the same class if they were repeatedly screened over a short time period. For example, if the individual is assessed for depression both today and tomorrow, a consistent screening procedure would reach the same diagnostic conclusion on both occasions. Quantitative indices of accuracy and consistency provide validity evidence to support the use of the screening procedure to make decisions about respondents (AERA et al., 2014; SIOP, 2023).
In applied settings, the dominant approach to make decisions from screening measures is a psychometric approach, which entails scoring item responses and determining if the score is above or below a cut score (Youngstrom, 2013). However, machine learning techniques (Jordan & Mitchell, 2015; Mitchell, 1997; Yarkoni & Westfall, 2017) are increasingly being applied to the area of assessment (e.g., Fokkema et al., 2014; Gibbons et al., 2013; Gonzalez, 2021a, 2021b; Lu & Pektova, 2014; Wicks et al., 2016; Yan et al., 2004; Yarkoni, 2010), and specifically to the area of screening (e.g., Askland et al., 2015; Bone et al., 2016; Carpenter et al., 2016; Gibbons et al., 2013; Gibbons et al., 2016; Kim et al., 2021; Lötsch et al., 2018; Schulte-Rüther et al., 2022; Tennenhouse et al., 2020; Wshah et al., 2019; Zheng et al., 2020). In these screening applications, a machine learning model was trained to predict a known diagnosis directly from the item responses of screening measures, and respondents are then classified as having a diagnosis if their predicted probability of diagnosis is above a specified cutpoint (Gonzalez, 2021a). Because the diagnostic status was known (from a gold standard) for each individual, the researchers often estimate and report the prediction accuracy of their machine learning model. For example, Gibbons and colleagues (2013) developed a computerized screening tool by using decision trees to classify individuals with and without major depressive disorder (MDD; determined via the Structured Clinical Interview for DSM-IV Axis I Disorders; SCID) from responses on rating scale items measuring depressive symptomatology. The screening tool they developed had a sensitivity of .95 and specificity of .87. However, we are aware of no study that has examined how consistent the screening decisions based on machine learning models are, perhaps because researchers are not familiar with these methods.
One approach to estimate classification consistency is to administer the measure twice to the respondents, obtain predicted probabilities from the machine learning model for each measurement occasion, and calculate the number of classification agreements across the two administrations (e.g., Beiser et al., 2016; Teitelbaum & Carey, 2000). However, collecting multiple administrations would be impractical in many applied settings, or the machine learning analyses may be secondary analyses from data collected at a single administration. As such, the goal of this paper is to illustrate how to obtain estimates of classification consistency for screening decisions that are based on machine learning using only a single administration of the item set (e.g., Gonzalez et al., 2021; Gonzalez et al., 2023; Lathrop & Chen, 2014; Lee, 2010).
In the remainder of the paper, we provide some background on screening and then describe two sources of error that lead to classification inconsistencies in machine learning models, along with describing procedures for assessing classification consistency with respect to each source. The first source is sampling error during model training, whose influence can be estimated using procedures for model stability (e.g., Philip et al., 2018). The second source is measurement error, whose influence can be estimated using a novel method we developed that aligns more with how consistency is conceived in psychometric applications (Gonzalez et al., 2021). Both approaches rely on data resampling procedures (e.g., bootstrap and Monte Carlo resampling) to approximate a person’s expected responses across repeated administration of the measure. Then, we illustrate the methods using three empirical examples predicting a Major Depressive Disorder (MDD) diagnosis, a personality disorder (PD) diagnosis, and pseudo lifetime anxiety from items responses of measures that can be used for screening. Finally, we provide R code in the supplement to facilitate the implementation of these methods. Since consistency is conceptually similar to reliability (Haertel, 2006), an index that is often reported alongside any type of psychological instrument, we think that it would benefit machine learning and assessment research if it became standard practice to report consistency indices alongside accuracy indices.
Background on Methods to Classify Individuals Based on Screening Measures
Although there are many types of screening measures, this paper discusses screening measures that are comprised of items that measure a unidimensional construct that is related to a diagnosis of interest. The diagnosis is often determined by a gold standard that is accepted in the field (Youngstrom, 2013).1 In the screening process, individuals respond to the items and their responses are aggregated to determine a decision or classification. When researchers are developing a screening measure, they may conduct studies to determine the optimal approach to aggregate the item responses in order to make a recommendation to users on how to classify individuals based on the screener results (Gonzalez, 2021a). Here, we define “the optimal approach to aggregate the item responses” as the approach that yields the best classification performance (e.g., classification accuracy and consistency) in the population of interest (Gonzalez, 2021a).
For example, researchers can add the item responses to calculate a score, called a summed score, and then check if the summed score is above a pre-determined cutpoint. A second way to determine the diagnosis from item responses is to fit a latent variable model to the item responses, estimate a latent variable score (e.g., factor score or IRT score), and then check if the latent variable score is above a cutpoint. A third approach could be to use the item responses in a predetermined prediction model and obtain a probability of diagnosis. Note that a latent variable is not estimated when the items directly predict the diagnosis/condition in the prediction model, such as when a researcher uses logistic regression (a model that may be familiar to many psychologists; Lindhiem et al., 2020) to regress the diagnosis or condition on the item responses. Although logistic regression is often a viable option to try, it is not the only method to build a prediction model. In this paper, we discuss the area of machine learning, which offers a series of algorithms (e.g., random forest, decision trees, support vector machines, etc.) that can learn the relations between predictors and an outcome within a dataset (Yarkoni & Westfall, 2017). Note logistic regression can be treated as a machine learning algorithm too (James et al., 2013). The performance of machine learning algorithms is application-specific, so it is difficult to know a priori which algorithm might perform best in any given dataset (Wolpert & Macready, 1997). Similar to determining the optimal approach to aggregate item responses, researchers can go through a similar approach to determine which algorithm to use when predicting a diagnosis, often guided by maximizing the classification performance of the algorithm (Gonzalez, 2021a).
As we mention above, it is often not possible to determine a priori which way to aggregate the item responses. There are situations in which some of these methods might provide advantages over the others (Gonzalez, 2021a; 2021b). For example, summed scores are likely to work well when the items are unidimensional, reliable, and have a strong and similar linear relation with the diagnosis (e.g., no interaction effects among items). Latent variable scores are likely to work well in similar situations, but they have the advantage that they may be more reliable than the summed score. However, a latent variable score requires a well-fitting latent variable model to the item responses and specialized software to obtain the score. On the other hand, a prediction model (e.g., machine learning model) may perform better than summed scores or latent variable scores when the items have varying relations with the diagnosis, which might include nonlinear or interactive relations among the items. However, the summed scores or latent variable scores might outperform a prediction model when the relations are linear and only item main effects are present. Another disadvantage of machine learning models in the area of screening is that we know how to estimate classification accuracy of models (i.e., sensitivity, specificity, and classification accuracy), but we have limited knowledge about how to estimate classification consistency compared to the knowledge we have for psychometric methods (Livingston & Lewis, 1995; Rudner, 2005).
Overall, we recommend researchers to conduct studies to support the approach they chose to aggregate the item responses to determine a diagnosis. In no way do we intend to promote machine learning models for screening over other approaches that might be best for a specific application. Machine learning models might not always provide advantages for screening and may underperform compared to classifying based on summed scores or latent variable scores. It also may be the case that some machine learning models will have better theoretical justification than others. For instance, machine learning models that use theory to select the predictors may outperform models in which we have less knowledge of the problem. However, for users who have found that a machine learning model classifies better than other methods, the user might notice that there is a lack of development on methods to estimate classification consistency compared to classification consistency methods for summed scores (Livingston & Lewis, 1995) or latent variable scores (Rudner, 2005). As such, for the purposes of this paper, we assume that the assessment specialist using the screener has already determined that a machine learning model provides the best way to aggregate the item responses on the screener. In that situation, our paper discusses two sources of classification inconsistencies in machine learning models and proposes two approaches to obtain an estimate of how these sources affect the predictions of machine learning models used for screening.
Sources of classification inconsistency in machine learning models
Source of inconsistency #1: sampling error.
Because researchers rarely have access to population data, inferences about a population are made by collecting a sample of data from the population and fitting models to that sample. Sampling error occurs because the statistic estimated from a sample will differ from the true population parameter.
In a screening scenario involving machine learning, we can define classification consistency as the extent to which the same decisions are made when the algorithm has been trained in different samples drawn from the same population. In machine learning, work in this area is studied under the label of model stability, which is analogous to the concept of classification consistency given the definition mentioned above (see Philipp et al., 2018, for a review of procedures to estimate model stability). Generally, a model that is determined to be stable is a model that yields consistent classifications. For situations in which researchers do not have multiple samples, they could approximate the process by resampling the sample data many times. There are many approaches for data resampling, such as subsampling (splitting data into nonoverlapping samples, e.g., two independent parts), the jackknife, or the bootstrap. Philipp et al. (2018) indicate that the bootstrap with replacement is an appropriate resampling option for model stability estimation because the sampled datasets contain fewer repeated cases than the jackknife approach while achieving greater statistical power than the subsampling approach. Below, we describe a procedure discussed by Philipp et al. (2018) to estimate model stability (i.e., classification consistency) using only a single administration of the measure.
Procedure for Sampling Error:
Divide the dataset into a training and testing dataset, which is done to have independent parts where we build the model and where we evaluate the model.2
- Over many iterations (e.g., 1,000 iterations), do the following:
- Using the nonparametric bootstrap with replacement, draw two samples from the training dataset.
- Choose a machine learning algorithm.
- Train the machine learning model in step b to predict the diagnosis from item responses in each of the samples (model_s1 and model_s2).3
- For model_s1 and model_s2, obtain predicted class membership for each case in the testing dataset.
- Estimate classification agreement by comparing the predictions from model_s1 and model_s2.
- Save the results.
Report the mean and standard deviation of the classification agreement estimates across iterations.
High mean classification agreement across the pairs of models would suggest that the machine learning algorithm is stable because it yields screening decisions for the individuals in the testing dataset that are consistent each time that the model is trained using different samples. Low mean classification agreement would suggest that the machine learning algorithm is not stable because the screening decisions are not consistent – the model provided different classifications for the same individuals when the algorithm was trained in different versions of the dataset. In other words, model training is sensitive to sampling error. As such, model stability procedures can be used to infer how consistent are the screening decisions from a machine learning model.
Source of inconsistency #2: measurement error.
Screening measures intend to assess a latent construct that is associated with the diagnosis. Observed item responses reflect the level of the construct θ that respondents have and random measurement error. For example, Figure 1 (left) shows that for a particular θ, there are several possible observed summed scores associated with it solely due to measurement error. As such, measurement error could randomly place individuals above or below a cut score for screening (for more details, see Gonzalez et al., 2021, Gonzalez et al., 2023; Lee, 2010, Lathrop & Chen, 2014; Livingston & Lewis, 1995). We can take these possible observed scores and define classification consistency for machine learning models. The main idea of this procedure is to obtain those possible observed scores via Monte Carlo sampling by simulating from a data-generating model (Phillip et al., 2018), which cannot be achieved via data resampling like the bootstrap. In this case, we start by simulating data from a latent variable model, obtain many possible item responses that an individual with a specific θ could have given, estimate predicted probabilities using those possible item responses, and examine how their predicted probability of diagnosis varies (see Figure 1, right). In this case, the role of the latent variable model is to understand how much error is in the item responses and to simulate data. The latent variable is not used directly in the prediction model. To obtain the possible observed scores and estimate classification consistency, we can adapt a procedure described by Gonzalez et al. (2021) to machine learning models.
Figure 1.

Left. Distribution of summed scores at a fixed value of θ (e.g., θ = 1). Right. Distribution of predicted probabilities at a fixed value of θ (e.g., θ = 1)
Procedure for Measurement Error:
Divide the model into a training and a testing dataset, which is done to have independent parts where we build the model and where we evaluate the model.4
Train a machine learning model in the training dataset to predict the diagnosis from item responses.
In the testing dataset, fit a latent variable model (e.g., linear factor model or an item response theory model) to the item responses and save item parameters.5
Define a range of θ over which the user wants to estimate classification consistency. Note that the θ values are not estimated, but they are hand-picked.
- Over many iterations (e.g., 1,000), simulate possible item response patterns:
- Create vector of equally spaced values along the θ range (e.g., quadrature points). Because the θ values are hand-picked, they treated as fixed.
- Generate item responses using the item parameters in Step 3 and θ values in step 5a. Note that the item parameters are treated as fixed (i.e., the estimation variability is ignored) to simulate data.
Using the model fit in step 1, obtain predicted classes from the item responses simulated in Step 5b.
Estimate the proportion of predicted cases in class 0 (p0) and class 1 (p1) at each value of θ from Step 5a and estimate of classification consistency at that specific value of θ using p02 + p12 (Lee, 2010). The estimate at each level of θ is known as conditional classification consistency.
Estimate marginal classification consistency by averaging the conditional classification consistency estimates from Step 7 using either a raw average or a weighted average with quadrature weights.
In this case, a high classification consistency estimate would indicate that the machine learning model yields screening decisions that are consistent for individuals at a specific level of θ regardless of the measurement error in the item responses. On the other hand, if the classification consistency is low, then it would suggest that the measurement error could possibly lead to respondents being reclassified.
Below, we present three illustrations in which we show how the two procedures described above could work with different machine learning algorithms. Note that we are assuming that a researcher already determined that using a machine learning algorithm is the best way to combine the item responses to determine the condition of interest. The purpose of the illustrations is to demonstrate the methodology; we do not intend to make a substantive contribution to how depression, personality disorders, or anxiety are diagnosed or screened for.
Transparency and Openness
This paper aims to propose and illustrate quantitative methods, as such this study design and its analyses were not preregistered, and it did not seek research ethics committee approval (although all datasets presented received IRB approval during the original data collection). In each illustration, we discuss the availability of the data, materials, and code.
Illustration 1: Predicting Major Depressive Disorder from Affect Items
The dataset used was from a study by Gibbons et al. (2013) in which the goal was to build a screening tool for MDD. There were N = 629 cases in the dataset, in which 39% were treatment-seeking outpatients and 61% were nonpsychiatric controls, and all were between 18 and 80 years of age. 65% of participants identified as female and 80% as White. Further details of sample recruitment and composition can be found in Gibbons et al. (2013).
The dataset includes the participants’ item responses and a clinical diagnosis that was determined independently of the item responses. The clinical diagnosis was made by a trained interviewer using the Structured Clinical Interview for DSM-IV Axis I Disorders (SCID; First et al., 1995). 134 participants met the criteria for MDD and 495 participants did not meet the criteria for MDD. Therefore, the prevalence of MDD in this data was 21%. Participants responded to a bank of 88 items from existing scales to measure depression that aligned with the DSM-IV symptom criteria for a diagnosis of MDD. However, for simplicity in illustrating the methods, we used only the responses to 16 items measuring positive and negative affect. We chose these items because they seemed to measure one factor, so the latent variable model for the data would be relatively simple. Adding other symptoms might make the latent variable model more complex, but the machine learning model would be more accurate and consistent.
The dataset was randomly split in half into training and testing datasets. A parallel analysis (Horn, 1965) conducted using the psych R-package (Revelle, 2018) in the training dataset suggests that the 16 items are largely unidimensional (1st eigenvalue = 11.966, 2nd eigenvalue = .782) and that a one-factor model fits the data generally well (χ2(104) = 266.744, p < .001, RMSEA = .071, CFI = .999).
Random Forest Algorithm
The procedures discussed could be used to estimate the classification consistency of any machine learning algorithm that produces classifications. Here, we illustrate using a random forest. The random forest algorithm is an ensemble method that grows a series of decision trees in bootstrapped datasets and averages predictions across the decision trees (Breiman, 2001; Hastie et al., 2009). As the tree is grown, random forest randomly selects a subset of predictors to be used as candidates each time a tree branch is split. For example, if there are 20 predictors in the dataset, the random forest would define splits using only, say, four randomly chosen predictors in each region (note that the number of predictors can be user-specified or treated as a tuning parameter). The predicted probability of diagnosis is determined by the proportion of trees that classified the individual in the diagnosed class, and a predicted classification can be determined by comparing the predicted probability to a cutpoint. In this illustration, we used the randomForest R-package to train the model (Liaw & Wiener, 2002), and chose the cutpoint that maximized the Youden index (i.e., sensitivity + specificity – 1). For this example, random forest had the following classification performance in the testing dataset: accuracy = .872, sensitivity = .653, and specificity = .941. As such, the model does a better job at ruling out who has MDD than flagging those who have it.
Procedure
Data were obtained by contacting the original study authors (Gibbons et al., 2013), and code and research materials for our study are available by contacting the first author of this paper. The analyses were carried out in the R statistical environment (v.4.1.0). We conducted 2,000 iterations of the Procedure for Sampling Error and 2,000 iterations of the Procedure for Measurement Error. For the random forest specifications, we used 500 trees and fixed the number of randomly sampled predictors at each split to the value that is the square root of the number of items, which was 4 (this is often used as a rule-of-thumb for this tuning parameter; Hastie et al., 2009). For the Procedure for Measurement Error, we estimated in the testing dataset a unidimensional graded response model, which is an item response model appropriate for ordered-categorical data (Samejima, 1969), using the mirt R-package (Chalmers, 2012), and the item parameters were saved. Note that, for identification, the mean and variance of θ were set to 0 and 1, respectively. Also, we chose 41 equally spaced quadrature points in the range of θ = [−2, 2] to probe consistency. In the supplement, we provide simulated data based on the item parameters and R code to facilitate the implementation of these procedures and that researchers can adapt for their purposes.
Results
Figure 2 (top) shows the distribution of the classification agreement estimates from the Procedure for Sampling Error. Across samples, the mean and standard deviation was .941 (.019). The value is high, which suggest that the random forest model trained to screen individuals for MDD by the affect item responses yields consistent classifications. We can interpret this value as having a 94.1% chance that the model would give the same classification to an individual once the model is trained in slightly different versions of the training dataset. In other words, screening decisions from the model might not be adversely affected by sampling error.
Figure 2.

Top: Distribution of classification agreement estimates from the Procedure for Sampling Error, Step 2f for the affect example. Bottom: Conditional classification consistency estimates from the machine learning model as a function of theta from the Procedure for Measurement Error, Step 7 for the affect example.
Furthermore, Figure 2 (bottom) also shows the estimates of conditional classification consistency at a specific value on θ from the Procedure for Measurement Error. Here, the plot indicates that classification consistency depends on θ. For individuals who are below the mean of θ (i.e., θ < ~0) or when θ > ~1.2, the classification consistency is close to 1. As such, for individuals in those regions of θ, measurement error in the item responses is unlikely to change their MDD predicted classification. On the other hand, measurement error is likely to yield less consistent classifications for individuals whose latent variable score is 0 < θ < 1.2. In this case, classification consistency sharply drops, hits a minimum of .50 (the lowest it could go) at θ = ~.5 and then sharply increases again. As such, this approach identifies a critical region in which classifications are likely to be less consistent. Also, the unweighted marginal classification consistency was .929 and the weighted estimate was .919, which means that, across the range of θ, the classifications made by the machine learning model were highly consistent – measurement error is not likely to affect the predicted classification derived from machine learning models for most respondents, except those in a specific latent variable range.
Overall, the results suggest that the decisions made by a random forest that predicts MDD diagnosis from the affect item responses were highly consistent, and they were not severely affected by measurement error or sampling error during model fitting. As such, if random forest were to be used for screening using this set of items and population of respondents, it is expected that it would yield consistent screening decisions.
Illustration 2: Predicting a Personality Disorder Diagnosis from the IPDS Responses
The dataset used (Pilkonis, 2018a; 2018b) was from a study by Morse and Pilkonis (2007) in which the goal was to evaluate screening measures to predict personality disorders (PD). For this specific illustration, we will use data from the Iowa Personality Disorder Screen (IPDS; Langbehn et al., 1999), which consists of 19 self-report items that ask about 11 DSM-IV PD criteria. Item responses are scored on the 11 criteria by indicating if the participant endorsed or not any of the items that assessed the criteria. The diagnosis of any possible PD of the respondent was determined by consensus among a panel of clinical investigators, research team, and other judges during a diagnostic conference after a 6-month follow-up. During the 2-hour diagnostic conference, the panel received all data about the respondent (e.g., social and developmental history, life events, symptom, Axis I and Axis II interviews, etc.) prior to making a decision. Additional details of this study can be found in Morse and Pilkonis (2007).
There were N = 151 cases in the original dataset, but we analyzed a complete dataset in which the respondent answered all items and were judged by the panel on a PD diagnosis (N=146). The age range of the respondents was between 21 and 60 years of age (mean = 38.77), 69% identified themselves as female, and 87% were White. Also, in this sample, 60% (n = 87) of the individuals were identified as having a PD by the panel and 40% (n=59) of individuals did not receive a PD diagnosis.
The dataset was randomly split in half into training and testing datasets. A confirmatory factor analysis conducted in the training dataset suggests that the 11 PD criteria fit a one-factor model generally well (χ2(44) = 50.736 p = .225, RMSEA = .046, CFI = .971).
Boosted Trees
For this illustration, we used boosted trees to predict the PD diagnosis from the 11 symptom criteria. Similar to random forest, boosted trees are an ensemble method comprised of decision trees, but the trees are grown differently. Instead of bootstrapping data and growing the trees independently, boosted trees are grown sequentially in random subsamples of the data (i.e., stochastic gradient boosting) by up weighing datapoints that previous trees did not predict well. Each tree is typically grown to a predetermined size (e.g., one or two splits) instead of growing the tree until a stopping criterion is achieved (e.g., lack of improvement in a loss function). A predicted classification is obtained by a linear combination of the predicted values of all the trees rather than just a simple average. In this illustration, we used the gbm R-package to train the model (Greenwell et al., 2020), and chose the cutpoint that maximized the Youden index. For this example, boosted trees had the following classification performance in the testing dataset: accuracy = .603, sensitivity = .547, and specificity = .677. As expected, the PD diagnosis is difficult to predict (Morse & Pilkonis, 2007).
Procedure
All data used are publicly available in the institutional repository of the University of Pittsburgh, and can be accessed via the following URIs: http://d-scholarship.pitt.edu/35416/ and http://d-scholarship.pitt.edu/35426/ . Code and research materials for our study are available by contacting the first author of this paper. Similar to above, we conducted the analyses in the R statistical environment, with 2,000 iterations of the Procedure for Sampling Error, and 2,000 iterations of the Procedure for Measurement Error (fitting a unidimensional graded response model to the data). We relied on the defaults of the gbm package for fitting boosted trees – trees with a single split (i.e., tree stumps) were grown on random subsamples of 50% of the data. We used 5-fold cross-validation to determine the number of trees that should be used for prediction, with a maximum number of trees of 100. Similar to the example above, we chose 41 equally spaced quadrature points in the range of θ = [−2, 2] to probe consistency.
Results
Figure 3 (top) shows the distribution of the classification agreement estimates from the Procedure for Sampling Error. Across samples, the mean and standard deviation was .703 (.166). The value is substantially lower than in the random forest example above, which suggest that the boosted trees model trained to screen individuals for PD by the IPDS item responses does not yield consistent classifications. We can interpret this value as having a 70.3% chance that the model would give the same classification to an individual once the model is trained in slightly different versions of the training dataset.
Figure 3.

Top: Distribution of classification agreement estimates from the Procedure for Sampling Error, Step 2f for the personality disorder example. Bottom: Conditional classification consistency estimates from the machine learning model as a function of theta from the Procedure for Measurement Error, Step 7 for the personality disorder example.
Furthermore, Figure 3 (bottom) also shows the estimates of conditional classification consistency at a specific value on θ from the Procedure for Measurement Error. As we can see, the curve is jagged and nonsymmetrical even though we are approximating the location of the curve with 2,000 replications at each of the 41 quadrature points, which speaks to the prediction variability by the boosted trees. Perhaps more replications would make the curve smoother. Compared to random forest example, this curve is wider, which suggests that the θ region in which classifications are affected by measurement error is larger. Also, the curve hits a minimum around θ = 0 and does not reach 1 in −2 < θ < 2, so we can conclude that the screening procedure is only perfectly consistent in individuals who are at the extreme values of θ. Furthermore, we can see that classification consistency is lower than .8 in the range defined by −1 < θ < 1. The unweighted marginal classification consistency was .754 and the weighted estimate was .777, which means that, across the range of θ, the classifications made by the machine learning model were not consistent – an individual might have only a 75% chance of receiving a similar classification by the boosted trees simply due to measurement error. Recall that the boosted tree model had poor prediction accuracy, so for inaccurate prediction models classification consistency could be compromised.
Illustration 3: Predicting Pseudo Lifetime Anxiety Diagnosis from MASQ-GA Responses
The dataset used for this example came from the original calibration sample used to develop the PROMIS Anxiety measures, where legacy anxiety measures were also administered with the goal of linking the new measure with legacy measures (Pilkonis et al., 2011). One of the legacy measures administered was the Mood and Anxiety Symptom Questionnaire – General Distress-Anxious Symptom Scale (MASQ-GA). The MASQ-GA is an 11-item subscale from the 90-item measure MASQ scale that assesses the presence and severity of symptoms related to generalized anxiety. The items are rated on a Likert-type scale, ranging from 1 (“Not at all”) to 5 (“Extremely”). For illustrational purposes only, we will use the MASQ-GA item responses to predict a self-report item in which respondents indicated if they had ever been told by their doctor that they have anxiety (which we refer to as pseudo lifetime anxiety).
For this analysis, we analyzed a complete dataset from a sample recruited online via Polimetrix that is representative of the general population. There were N = 783 cases who attempted to complete the MASQ-GA, but we analyzed N=730 who had complete data and also responded to the item of pseudo lifetime anxiety. The age range of the respondents was between 18 and 88 years of age (mean = 50.48), 48.6% identified themselves as female, and 79.1% were White. Also, in this sample, 16.4% (n = 120) reported pseudo lifetime anxiety and 83.6% (n = 610) did not report pseudo lifetime anxiety.
The dataset was randomly split in half into training and testing datasets. A parallel analysis conducted in the training dataset suggests that the 11 MASQ-GA items are largely unidimensional (1st eigenvalue = 5.22, 2nd eigenvalue = 0.76) even though the one-factor model does not fit the data well (χ2(44) = 331.176, p < .001, RMSEA = .134, CFI = .984).
Relaxed Lasso Logistic Regression
For this example, we used relaxed lasso logistic regression to predict pseudo lifetime anxiety from the 11 MASQ-GA items. Logistic regression imposes a linear relation between the MASQ-GA item responses and the log of the odds of pseudo lifetime anxiety. The lasso component helps in the prediction and interpretability of the logistic model by adding a penalty to the loss function of logistic regression that controls the size of the regression coefficients and shrinks them toward zero. The lasso penalty involves the sum of the absolute size of the regression coefficient, and the impact of the penalty is controlled by a tuning parameter. If the regression coefficient of one of the items is reduced to zero, then the item does not have an influence on the prediction model. As such, the lasso penalty acts as a variable selection procedure where the variables selected are those with non-zero regression coefficients. A problem with the estimated regression coefficients with the lasso penalty is that they are downwardly biased toward zero, so one can “relax” the solution by refitting a logistic regression model without the lasso penalty only with the variables that remained in the model. In this illustration, we used the glmnet R-package to train the model (Friedman et al., 2002), and chose the cutpoint that maximized the Youden index. For this example, the lasso only selected two out of the 11 items to remain the model, and the classification performance in the testing dataset was an accuracy = .616, sensitivity = .667, and specificity = .607.
Procedure
All data used are publicly available in the HealthMeasures Dataverse and can be accessed via the following DOI: https://doi.org/10.7910/DVN/0NGAKG . Code and research materials for our study are available by contacting the first author of this paper. Similar to above, we conducted the analyses in the R statistical environment, with 2,000 iterations of the Procedure for Sampling Error, and 2,000 iterations of the Procedure for Measurement Error (fitting a unidimensional graded response model to the data). For the relaxed lasso logistic regression specification, we used 5-fold cross-validation to determine the size of the tuning parameter that determines the influence of the lasso penalty on the solution (Hastie et al., 2009). Also, we chose 41 equally spaced quadrature points in the range of θ = [−1, 3] to probe consistency.
Results
Figure 4 (top) shows the distribution of the classification agreement estimates from the Procedure for Sampling Error. Across samples, the mean and standard deviation was .789 (.106). The value is lower than in the random forest example above, but still generally high, which suggest that the relaxed lasso logistic regression model trained to screen individuals for pseudo lifetime anxiety by the MASQ-GA item responses yields consistent classifications. We can interpret this value as having an 80.3% chance that the model would give the same classification to an individual once the model is trained in slightly different versions of the training dataset.
Figure 4.

Top: Distribution of classification agreement estimates from the Procedure for Sampling Error, Step 2f for the MASQ-GA example. Bottom: Conditional classification consistency estimates from the machine learning model as a function of theta from the Procedure for Measurement Error, Step 7 for the MASQ-GA example.
Furthermore, Figure 4 (bottom) also shows the estimates of conditional classification consistency at a specific value on θ from the Procedure for Measurement Error. Similar to before, classification consistency depends on θ. For individuals who are below the mean of θ (i.e., θ < ~0.5), the classification consistency is close to 1. As such, for individuals in those regions of θ, measurement error in the item responses is unlikely to change their pseudo lifetime anxiety predicted classification. On the other hand, measurement error is likely to yield less consistent classifications for individuals whose latent variable score is 0.5 < θ < 3. In this case, classification consistency sharply drops, hits a minimum of .5 at θ = ~1.5 and then increases again. Also, the unweighted marginal classification consistency was .856 and the weighted estimate was .869, which means that, across the range of θ, the classifications made by the machine learning model were generally consistent. Note that the curve for the relaxed lasso logistic regression model is wider than in the random forest example, but narrower than the boosting example, which suggests that the region in which classifications are affected by measurement error is larger than in the random forest example, but better than in the boosting example.
Discussion
Estimates of classification consistency are important to assessment because they describe the reliability of decisions made from screening procedures. If the classifications from the screening procedure are not reliable, then the valid use of model predictions for decision-making would be compromised (AERA et al., 2014; SIOP, 2023). In this paper, we described two sources of classification inconsistencies, sampling error and measurement error, along with procedures to estimate the classification inconsistency arising from each source. These two procedures provide different information about classification inconsistencies and should be used together. Because the procedures are algorithm-independent, they can be conducted with any machine learning model. In the supplement, we included R code that enables assessment specialists to conduct the procedures presented in this paper. If we want to incorporate machine learning models to make valid decisions based on assessments, then it is important to evaluate the consistency of decisions in the screening process based on the machine learning algorithms.
Having illustrated the methods, we can now draw some comparisons about the steps that the two procedures take even though the procedures address different sources of inconsistency. The Procedure for Measurement Error can be thought of as a version of the Procedure for Sampling Error for situations in which we know the data-generating process (Philipp et al., 2018), but note that the model is only trained once in the Procedure for Measurement Error instead of training multiple models in the Procedure for Sampling Error. In the Procedure for Measurement Error, we are imposing a psychometric structure on the predictors used in the machine learning model by fitting a latent variable model, which facilitates the simulation of the possible observed scores. Even though determining a factor structure adds a layer of complication to the Procedure for Measurement Error, in many cases the factor structure of the screening measures has already been confirmed by other research studies or is well-known. In cases where the factor structure is not known, one could use the training dataset to explore potential factor structures and validate the structure in the testing dataset. On the other hand, the Procedure for Sampling Error does not have the complication of determining the data-generating model of the predictors. Also, the Procedure for Measurement Error is currently not able to handle situations in which the machine learning model has additional predictors that are not item responses, and this is not a problem with the Procedure for Sampling Error.6
Constraints on Generality
As mentioned above, the purpose of this paper is to demonstrate methodology used to determine classification consistency for machine learning models used for screening. For this purpose, we used existing datasets and had no control over the sample. We do not intend for the results of our illustrations to make clinical contributions to how depression, personality disorders, or anxiety are diagnosed or screened for in the general population.
Limitations
There were several limitations to the proposed procedures. For both procedures, the splitting of the sample into training and testing datasets could introduce prediction variability, and future work could address the best way to handle that source of variability. As noted above, the Procedure for Measurement Error assumes that the latent variable model fits the data well. In the case in which a latent variable model does not fit or is not theoretically feasible (e.g., the items assess a formative construct), then the Procedure for Measurement Error would not be useful. Also, another limitation of the procedure is that they are likely to work better in datasets with large sample sizes. In relation to our procedure, fitting a machine learning model or a latent variable model in dataset with a large sample size can provide more stable predictions and more precise parameter estimates than fitting them in datasets with smaller sample sizes. Sample size also becomes a factor when the researcher splits their data into a training and testing dataset. When researchers have a small sample, perhaps k-fold cross-validation would be a better approach to proceed. It is difficult to make suggestions of how large a sample size should be because strong relations among variables can offset the need for huge sample sizes. Furthermore, in situations in which the number of individuals diagnoses with the gold standard is small (i.e., a low base rate), there might be implications to the estimation of classification consistency. Beyond having trouble predicting the diagnosed class accurately, the estimate of classification consistency from the machine learning model could be artificially high simply because the algorithm would predict “no diagnosis” in a sample in which the vast majority of individuals do not have the diagnosis. As such, researchers should interpret the classification consistency estimate carefully or report another estimate such as Cohen’s kappa.
Future Directions
There are several areas of this works that would be explored to address the noted limitations. A future direction of this work is to examine if the Procedure for Measurement Error is robust to misspecifications in the psychometric model. Furthermore, one could investigate if there are ways to determine other possible item responses per person without relying on a latent variable model, such as examining responses by other individuals with similar response patterns. Regarding the Procedure for Sampling Error, future work could systematically investigate via simulations other resampling variants to estimate agreement across cases or investigate how to obtain estimates of agreement per respondent across bootstrap samples. In that work, it would be important to determine the best way to summarize that information to have an estimate of consistency. Also, it is important to examine plans for remediation when classifications are not consistent, such as adding more people to mitigate inconsistencies due to sampling error or more items to mitigate inconsistencies due to measurement error. Also, we expect that having high quality items (more reliable) would mitigate the effects of sampling error and measurement error – item responses that are more reliable are not likely to fluctuate as much as the responses to less reliable items. Furthermore, future work involves establishing new procedures to estimate consistency with respect to both measurement error and sampling error simultaneously and to disaggregate both sources of inconsistency. In addition, it would be important to find the best ways to estimate a confidence interval for the classification consistency estimates. Gonzalez (2023) showed how to estimate Bayesian credible intervals and bootstrap confidence intervals for classification consistency estimates of summed scores. However, Bayesian confidence intervals would not apply here because not all machine learning algorithms can be estimated using Bayesian inference, and it is unclear how to use bootstrap confidence intervals on a procedure that already relies on bootstrap datasets (i.e., the bootstrapped data would have to be bootstrapped again, so the same cases have even a higher probability of appearing in the bootstrapped sample).
Future directions also include studying the relation between psychometrics and machine learning algorithms in the estimation of classification consistency. As mentioned above, in situations in which the relation between the item responses and the diagnosis are linear and additive, the psychometric approaches commonly used for screening might be more consistent compared to the machine learning models. However, when the relation between item responses and the diagnoses are nonlinear or interactive, machine learning might provide some advantages over psychometric approaches (Gonzalez, 2021a). As such, continuing to understand the relative advantages and disadvantages of these methods would be an important area of future work. As far as we are aware of, there are no benefits or drawbacks to any of the machine learning methods in the estimation of classification consistency. As mentioned above, the performance of machine learning algorithms is specific to the application (i.e., there is no free lunch; Wolpert & Macready, 1997). We presume that researchers would be interested in estimating classification consistency on the machine learning algorithm that exhibited the best classification accuracy during the development of the screening measure. Furthermore, it is important to study the relationship between classification accuracy and consistency. Since there has to be some signal (i.e., reliability) in the item responses to detect relations with the outcome, one might infer a certain level of consistency once we determine that the model yields accurate predictions (i.e., high sensitivity and specificity; Gonzalez et al., 2023). Similarly, studying both classification accuracy and consistency could be useful for model selection (Philip et al., 2018). The comparison of screening approaches on accuracy alone may obscure meaningful differences between the methods in classification consistency. If accuracy is similar, we would select the framework that produces superior consistency; if accuracy differs, then we must consider both accuracy and consistency when selecting between procedures.
In conclusion, as researchers continue to incorporate machine learning methods for decision making, we need to evaluate if these models are fit for real-world applications. Therefore, continuing to develop tools to determine consistency of machine learning models is critical for the adoption of these methods for assessment. We encourage researchers to report the statistics presented in this paper to have a better understanding of the reliability of the decisions drawn from machine learning models.
Supplementary Material
Public Significance Statement.
Recently, methods for machine learning have been used to predict from a screening measure if individuals should be flagged for a condition (e.g., as depressed vs. not depressed), but it is unknown if the models provide consistent screening decisions if respondents were to repeatedly receive the screening measure. We propose statistical procedures to help researchers determine if a machine learning model is providing consistent screening decisions.
Acknowledgments
Dr. Georgeson was supported by the National Institute on Drug Abuse (DA053137). Dr. Pelham was supported by the National Institute on Drug Abuse (DA055935) and the National Institute on Alcohol Abuse and Alcoholism (AA030197). This article conducted secondary data analysis for method illustrations, so the datasets were used in the original publications. Data collection for one of the illustrations was supported by NIMH grant R01MH66302. The authors also thank Dr. Paul Pilkonis for making his data available for secondary analysis. Data by Dr. Pilkonis can be found in the Institutional Repository at the University of Pittsburgh. The authors thank the brown bag attendees at the University of North Carolina at Chapel Hill, University of Notre Dame, and University of California, Los Angeles for their valuable feedback.
Footnotes
Although not the central focus of the paper, an aspect that is critical to the development of the screener is the way that the diagnosis is established (e.g., structured clinical interview, consensus decision, etc.). If the diagnosis status is not well determined at the outset, then any screening decisions would be ineffective or misleading. That is a question of design that cannot be addressed with statistical methods.
If models are trained and evaluated on the same dataset, classification performance statistics might be overly inflated. Step 1 can also be replaced with k-fold cross-validation if needed.
By dividing the data, one trains the model using a dataset smaller than the full sample, which can impact classification performance. In machine learning models where a likelihood is available (e.g., logistic regression, lasso linear regression), one could obtain expected cross-validation estimates without needing to split the data. See Ding et al. (2018) for a review.
Similar to the Procedure for Sampling Error, one can replace this step with k-fold cross-validation, but one would have to make sure that the size of the test fold is large enough to estimate a latent variable model. As such, we expect that this procedure might work better with large samples, but more research is needed to verify this.
Note that, to evaluate the model, we generated data based on a latent variable model fit to the testing dataset to prevent overly optimistic classification results due to generating data based on the training dataset.
A potential solution to the Procedure for Measurement Error could be to estimate a latent variable score per person, , and then use to simulate possible observed scores per case and keep the non-item predictors associated with each constant. Future research can examine this extension.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association. [Google Scholar]
- Askland KD, Garnaat S, Sibrava NJ, Boisseau CL, Strong D, Mancebo M, ... & Eisen J. (2015). Prediction of remission in obsessive compulsive disorder using a novel machine learning strategy. International Journal of Methods in Psychiatric Research, 24(2), 156–169. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Beiser D, Vu M, & Gibbons R (2016). Test-retest reliability of a computerized adaptive depression screener. Psychiatric Services, 67(9), 1039–1041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bone D, Bishop SL, Black MP, Goodwin MS, Lord C, & Narayanan SS (2016). Use of machine learning to improve autism screening and diagnostic instruments: effectiveness, efficiency, and multi-instrument fusion. Journal of Child Psychology and Psychiatry, 57(8), 927–937. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Breiman L (2001). Random forests. Machine Learning, 45, 5–32. [Google Scholar]
- Carpenter KL, Sprechmann P, Calderbank R, Sapiro G, & Egger HL (2016). Quantifying risk for anxiety disorders in preschool children: a machine learning approach. PloS one, 11(11), e0165524. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chalmers RP, (2012). mirt: A Multidimensional Item Response Theory Package for the R Environment. Journal of Statistical Software, 48, 1–29. [Google Scholar]
- Ding J, Tarokh V, & Yang Y (2018). Model selection techniques: An overview. IEEE Signal Processing Magazine, 35(6), 16–34. [Google Scholar]
- First MB, Spitzer RL, Gibbon M, & Williams JB (1995). The structured clinical interview for DSM-III-R personality disorders (SCID-II). Part I: Description. Journal of Personality Disorders, 9, 83–91. [Google Scholar]
- Fokkema M, Smits N, Kelderman H, Carlier IV, & van Hemert AM (2014). Combining decision trees and stochastic curtailment for assessment length reduction of test batteries used for classification. Applied Psychological Measurement, 38, 3–17. [Google Scholar]
- Friedman J, Hastie T, Tibshirani R, (2010). Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33, 1–22. [PMC free article] [PubMed] [Google Scholar]
- Gibbons RD, Hooker G, Finkelman MD, Weiss DJ, Pilkonis PA, Frank E, Moore T, & Kupfer DJ (2013). The CAD-MDD: a computerized adaptive diagnostic screening tool for depression. The Journal of Clinical Psychiatry, 74, 669–674. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gibbons RD, Weiss DJ, Frank E, & Kupfer D (2016). Computerized adaptive diagnosis and testing of mental health disorders. Annual Review of Clinical Psychology, 12, 83–104. [DOI] [PubMed] [Google Scholar]
- Gonzalez O (2021a). Psychometric and machine learning approaches for diagnostic assessment and tests of individual classification. Psychological Methods, 26, 236–254. DOI: 10.1037/met0000317. [DOI] [PubMed] [Google Scholar]
- Gonzalez O (2021b). Psychometric and machine learning approaches to reduce the length of scales. Multivariate Behavioral Research, 56, 903–919. DOI: 10.1080/00273171.2020.1781585. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gonzalez O, Georgeson AR, & Pelham WE III. (2023). How accurate and consistent are score-based assessment decisions? A procedure using the linear factor model. Assessment, 30, 1640–1650. DOI: 10.1177/10731911221113568. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gonzalez O, Georgeson AR, Pelham WE III., & Fouladi RT. (2021). Estimating classification consistency of screening measures and quantifying the impact of measurement bias. Psychological Assessment, 37, 596–609. DOI: 10.1037/pas0000938. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Greenwell B, Boehmke B, Cunningham J, & GBM Developers (2020).gbm: Generalized Boosted Regression Models. R package version 2.1.8. [Google Scholar]
- Haertel EH (2006). Reliability. In Brennan RL (Ed.), Educational measurement, 4th ed. (pp. 65–110). Westport, CT: Praeger Publishers. [Google Scholar]
- Hastie T, Tibshirani R, & Friedman J (2009). The elements of statistical learning. Springer Series in Statistics. [Google Scholar]
- Horn JL (1965). A rationale and test for the number of factors in factor analysis. Psychometrika, 30, 179–185. [DOI] [PubMed] [Google Scholar]
- Jordan MI, & Mitchell TM (2015). Machine learning: Trends, perspectives, and prospects. Science, 349, 255–260. [DOI] [PubMed] [Google Scholar]
- Kim S, Lee HK, & Lee K (2021). Which PHQ-9 items can effectively screen for suicide? Machine learning approaches. International Journal of Environmental Research and Public Health, 18(7), 3339. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Langbehn DR, Pfohl BM, Reynolds S, Clark LA, Battaglia M, Bellodi L, Cadoret R, Grove W, Pilkonis P, & Links P (1999). The Iowa Personality Disorder Screen: Development and preliminary validation of a brief screening interview. Journal of Personality Disorders, 13, 75–89. [DOI] [PubMed] [Google Scholar]
- Lathrop QN, & Cheng Y (2014). A nonparametric approach to estimate classification accuracy and consistency. Journal of Educational Measurement, 51, 318–334. [Google Scholar]
- Lee WC (2010). Classification consistency and accuracy for complex assessments using item response theory. Journal of Educational Measurement, 47, 1–17. [Google Scholar]
- Liaw A, & Wiener M (2002). Classification and regression by randomForest. R news, 2, 18–22. [Google Scholar]
- Lindhiem O, Petersen IT, Mentch LK, & Youngstrom EA (2020). The importance of calibration in clinical psychology. Assessment, 27, 840–854. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Livingston SA, & Lewis C (1995). Estimating the consistency and accuracy of classifications based on test scores. Journal of Educational Measurement, 32, 179–197. [Google Scholar]
- Lötsch J, Sipilä R, Dimova V, & Kalso E (2018). Machine-learned selection of psychological questionnaire items relevant to the development of persistent pain after breast cancer surgery. British Journal of Anaesthesia, 121(5), 1123–1132. [DOI] [PubMed] [Google Scholar]
- Lu F, & Petkova E (2014). A comparative study of variable selection methods in the context of developing psychiatric screening instruments. Statistics in Medicine, 33, 401–421. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mitchell TM (1997). Machine learning. New York: McGraw-Hill. [Google Scholar]
- Morse JQ, & Pilkonis PA (2007). Screening for personality disorders. Journal of Personality Disorders, 21, 179–198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Philipp M, Rusch T, Hornik K, & Strobl C (2018). Measuring the stability of results from supervised statistical learning. Journal of Computational and Graphical Statistics, 27, 685–700. [Google Scholar]
- Pilkonis P (2018a) Sumsheet (Axis I&II Diagnoses, GAF, PD/Global Severity) - Personality Studies. [Dataset] (Unpublished). [Google Scholar]
- Pilkonis P (2018b) Iowa Personality Disorder Screen (IOWA) - Personality Studies. [Dataset] (Unpublished) [Google Scholar]
- Pilkonis PA, Choi SW, Reise SP, Stover AM, Riley WT, Cella D, & PROMIS Cooperative Group. (2011). Item banks for measuring emotional distress from the Patient-Reported Outcomes Measurement Information System (PROMIS®): depression, anxiety, and anger. Assessment, 18(3), 263–283. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Revelle W (2018) psych: Procedures for Personality and Psychological Research. R package version 1.8.12. [Google Scholar]
- Rudner LM (2005). Expected classification accuracy. Practical Assessment, Research & Evaluation, 10(13), 1–4 [Google Scholar]
- Samejima F (1969). Estimation of latent ability using a response pattern of graded scores. Psychometrika Monographs, 34(Suppl. 17). [Google Scholar]
- Sheikh JI, & Yesavage JA (1986). Geriatric Depression Scale (GDS): Recent evidence and development of a shorter version. Clinical Gerontologist: The Journal of Aging and Mental Health, 5, 165–173. [Google Scholar]
- Schulte-Rüther M, Kulvicius T, Stroth S, Wolff N, Roessner V, Marschik PB, ... & Poustka L. (2022). Using machine learning to improve diagnostic assessment of ASD in the light of specific differential and co-occurring diagnoses. Journal of Child Psychology and Psychiatry. [DOI] [PubMed] [Google Scholar]
- Society for Industrial and Organizational Psychology [SIOP] (2023). Considerations and recommendations for the validation and use of AI-based assessments for employee selection. https://www.siop.org/Research-Publications/Items-of-Interest/ArtMID/19366/ArticleID/7327/SIOP-Releases-Recommendations-for-AI-Based-Assessments
- Teitelbaum LM, & Carey KB (2000). Temporal stability of alcohol screening measures in a psychiatric setting. Psychology of Addictive Behaviors, 14, 401–404. [DOI] [PubMed] [Google Scholar]
- Tennenhouse LG, Marrie RA, Bernstein CN, & Lix LM (2020). Machine-learning models for depression and anxiety in individuals with immune-mediated inflammatory disease. Journal of Psychosomatic Research, 134, 110126. [DOI] [PubMed] [Google Scholar]
- Wshah S, Skalka C, & Price M (2019). Predicting posttraumatic stress disorder risk: a machine learning approach. JMIR Mental Health, 6(7), e13946. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wicks P, Hotopf M, Narayan VA, Basch E, Weatherall J, & Gray M (2016). It’s a long shot, but it just might work! Perspectives on the future of medicine. BMC medicine, 14(1), 1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wolpert DH, & Macready WG (1997). No free lunch theorems for optimization. IEEE transactions on evolutionary computation, 1, 67–82. [Google Scholar]
- Yan D, Lewis C, & Stocking M (2004). Adaptive testing with regression trees in the presence of multidimensionality. Journal of Educational and Behavioral Statistics, 29, 293–316. [Google Scholar]
- Yarkoni T, & Westfall J (2017). Choosing prediction over explanation in psychology: Lessons from machine learning. Perspectives on Psychological Science, 12, 1100–1122. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Youngstrom EA (2013). A primer on receiver operating characteristic analysis and diagnostic efficiency statistics for pediatric psychology: we are ready to ROC. Journal of Pediatric Psychology, 39, 204–221. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zheng Y, Cheon H, & Katz CM (2020). Using machine learning methods to develop a short tree-based adaptive classification test: Case study with a high-dimensional item pool and imbalanced data. Applied Psychological Measurement, 44, 499–514. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
