Skip to main content
Journal of Speech, Language, and Hearing Research : JSLHR logoLink to Journal of Speech, Language, and Hearing Research : JSLHR
. 2024 Feb 22;67(3):753–781. doi: 10.1044/2023_JSLHR-23-00273

Toward Generalizable Machine Learning Models in Speech, Language, and Hearing Sciences: Estimating Sample Size and Reducing Overfitting

Hamzeh Ghasemzadeh a,b,c,, Robert E Hillman a,b,d,e, Daryush D Mehta a,b,d,e
PMCID: PMC11005022  PMID: 38386017

Abstract

Purpose:

Many studies using machine learning (ML) in speech, language, and hearing sciences rely upon cross-validations with single data splitting. This study's first purpose is to provide quantitative evidence that would incentivize researchers to instead use the more robust data splitting method of nested k-fold cross-validation. The second purpose is to present methods and MATLAB code to perform power analysis for ML-based analysis during the design of a study.

Method:

First, the significant impact of different cross-validations on ML outcomes was demonstrated using real-world clinical data. Then, Monte Carlo simulations were used to quantify the interactions among the employed cross-validation method, the discriminative power of features, the dimensionality of the feature space, the dimensionality of the model, and the sample size. Four different cross-validation methods (single holdout, 10-fold, train–validation–test, and nested 10-fold) were compared based on the statistical power and confidence of the resulting ML models. Distributions of the null and alternative hypotheses were used to determine the minimum required sample size for obtaining a statistically significant outcome (5% significance) with 80% power. Statistical confidence of the model was defined as the probability of correct features being selected for inclusion in the final model.

Results:

ML models generated based on the single holdout method had very low statistical power and confidence, leading to overestimation of classification accuracy. Conversely, the nested 10-fold cross-validation method resulted in the highest statistical confidence and power while also providing an unbiased estimate of accuracy. The required sample size using the single holdout method could be 50% higher than what would be needed if nested k-fold cross-validation were used. Statistical confidence in the model based on nested k-fold cross-validation was as much as four times higher than the confidence obtained with the single holdout–based model. A computational model, MATLAB code, and lookup tables are provided to assist researchers with estimating the minimum sample size needed during study design.

Conclusion:

The adoption of nested k-fold cross-validation is critical for unbiased and robust ML studies in the speech, language, and hearing sciences.

Supplemental Material:

https://doi.org/10.23641/asha.25237045


Statistical analysis and methods are essential parts of the scientific discovery process, and they play important roles during the design of an experiment, its execution, and its final analysis. For example, statistical power analysis drives the decision about the required sample size during experimental design (Cohen, 2013; Jones et al., 2003). Additionally, inferential statistics allow for the study of a sample of a population whose findings can then generalize to the larger population (Field et al., 2012; Lowry, 2014). Statistical machine learning (ML) is a group of statistical tools and techniques that has gained a lot of attention in different branches of science, with particular application to fields that are relevant to this journal's audience, such as health care (Liu et al., 2019; Qayyum et al., 2020; Shen et al., 2021) and speech, language, and hearing sciences (Oleson et al., 2019). ML provides many advantages over conventional statistical analysis. For example, a single model can often incorporate features from different scales of measurement (nominal, ordinal, interval, and ratio). Also, the ability of ML to capture complex, high-dimensional, and nonlinear interactions among different features can provide a significant advantage over conventional statistical methods. ML can be categorized into two main groups: supervised and unsupervised ML. In supervised ML, true labels of the data are known a priori (e.g., knowing which speech samples are recorded from individuals with Parkinson's disease or from neurotypical individuals). Supervised ML encompasses a wide variety of methods ranging from simple logistic regression (Menard, 2002) to deep learning (Liu et al., 2019) and ensemble learning (Ghasemzadeh, 2019b; Sagi & Rokach, 2018). In unsupervised ML, unlabeled data are used to discover underlying structures and patterns of the data (e.g., doing the same study but not knowing which voice data came from individuals with Parkinson's disease; Theodoridis & Koutroumbas, 2009).

Studies using supervised ML in health care (including speech, language, and hearing sciences) can be categorized into two main groups depending on their aims: tool developing and knowledge developing. Tool-developing studies primarily use ML to “create a tool” for performing a task automatically. On the other hand, knowledge-developing studies (also referred to as ML-based science in the literature; Kapoor & Narayanan, 2022) primarily use ML as a robust statistical analysis and data mining tool to gain knowledge and/or answer research questions about certain phenomena. Tool-developing studies closely resemble an engineering approach, and their aims could vary from automated diagnosis and evaluation to segmentation of a signal of interest, to solving an inverse problem. Examples of using ML to automatically differentiate healthy individuals from those with disorders (for automated screening/diagnosis and assessment of treatments) include the identification of voice disorders using voice samples (Arjmandi & Pooyan, 2012; Arjmandi et al., 2011; Ghasemzadeh et al., 2015), Parkinson's disease using speech recordings (Ghasemzadeh et al., 2022; Tsanas et al., 2012), amyolateral sclerosis using speech recordings (Ghasemzadeh & Searl, 2018; Vieira et al., 2022), septic infants from the audio signal produced by their cry (Matikolaie & Tadj, 2022), abnormal vocal folds from laryngoscopic images (Cho & Choi, 2020), intelligibility assessment of aphasic speech (Le et al., 2016), swallowing disorder from high-resolution manometry recordings (Mielens et al., 2012), screening for autism spectrum disorder (Crippa et al., 2015; Thabtah & Peebles, 2020), predicting the hearing outcome in patients with sudden sensorineural hearing loss (Bing et al., 2018; Uhm et al., 2021), and predicting language and communication skills of infants based on their speech-evoked electroencephalography data (Wong et al., 2021). Studies that have utilized ML for segmentation applications include automated segmentation of the glottis from high-speed videoendoscopy recordings (Kist et al., 2021), detection of high-speed videoendoscopy frames with partial occlusion of the vocal folds (Yousef et al., 2022), tongue segmentation from ultrasound images (Hamed Mozaffari & Lee, 2019), segmentation of videofluoroscopic recordings for swallowing studies (Donohue et al., 2021), speech classification for improving the performance of hearing aids (Bhat et al., 2020), and identification of the type of stuttering-related events from speech (Alharbi et al., 2020; Bayerl et al., 2022). Estimating parameters of the phonatory system from acoustic signals (Gómez et al., 2018; Ibarra et al., 2021; Zhang, 2020), measuring the distance between a laser-projection endoscope and the vocal folds (Ghasemzadeh et al., 2020), and compensating for nonlinear image distortion in fiberoptic endoscopes and performing calibrated mm-measurements in laryngeal images (Ghasemzadeh et al., 2021) are some examples of solving an inverse problem in voice and speech science using ML.

On the other hand, in knowledge-developing studies, the main outcome is not the trained model itself but rather the knowledge that has been gained. One example from this category would be the application of ML for understanding differences in phonatory functions between vocally healthy individuals and patients with phonotraumatic vocal hyperfunction (PVH) using ambulatory recordings (Ghassemi et al., 2014; Mehta et al., 2015; Van Stan et al., 2020), or differences between mild and moderate severity of phonotrauma (Van Stan et al., 2023). Other examples would be the quantification of the information lost in the perceptual evaluation of voice due to inherent limitations of the human auditory system (Ghasemzadeh & Arjmandi, 2020), comparison between the information content of temporal and spectral features of connected speech for clinical evaluation of voice and speech (Ghasemzadeh et al., 2022), the optimum number of sensors for encoding articulatory information from electromagnetic articulography recordings (Wang et al., 2016), the effect of speech duration on classification uncertainty of persons with dementia (Ossewaarde et al., 2020), quantification of articulatory distinctiveness of different vowels and consonants (Wang et al., 2013), evaluating the relevance of different features in predicting hearing loss (Lenatti et al., 2022), and investigating the predictive power of different items of the Infant Monitoring Questionnaire on language outcomes (Armstrong et al., 2018).

The Problem of Overfitting in ML

Despite their growing popularity in the speech, language, and hearing sciences, ML models are susceptible to overfitting, which refers to models performing very well on one data set but poorly on similar, but independent, data sets. One contributing factor to overfitting is that the decision boundary of a model is determined during the training process, and the boundary could be complex, nonlinear, and in a high-dimensional space on the training data set. The second contributing factor could be due to improper implementation of model selection or hyperparameter optimization components, which may lead to data leakage (Kapoor & Narayanan, 2022) and overestimation of model performance. Model selection refers to the process of choosing the best-performing model out of many candidate models. Feature selection is a common example of model selection in speech, language, and hearing sciences, which determines the best subset of features or measures that yield the highest performance. Hyperparameter optimization refers to the process of tuning parameters of an ML algorithm (e.g., parameters of a support vector machine [SVM]) to get the maximum performance. Although the possibility of having a high-dimensional, complex, nonlinear decision boundary and model selection account for many of the primary advantages of ML over conventional statistical analysis, they raise the possibility of overfitting.

Proper utilization of an ML step termed cross-validation can both alleviate overfitting (Hawkins, 2004) and estimate the generalizability of the trained model (Theodoridis & Koutroumbas, 2009). Cross-validation is a method that uses some portion of the data called the training set for training the model and then evaluates the performance of the trained model on the remainder of the data called the testing set. Different cross-validation approaches are explained later in the Cross-Validation Methods section of this article. Reviewing the existing literature that used ML in the speech, language, and hearing sciences showed some important trends. For example, many of these investigations have employed simple cross-validation techniques only containing a single testing set (see, e.g., Ibarra et al., 2021; Van Stan et al., 2020; Vieira et al., 2022; Yousef et al., 2022; Zhang, 2020). Another significant observation was the sample size. While very large public data sets (several thousand to several hundred thousand samples) are available for machine vision (Huang et al., 2008) and speech processing (Panayotov et al., 2015), studies from the speech, language, and hearing fields are associated with small sample sizes (often a few tens of samples). Small sample sizes are especially susceptible to overfitting, with extra caution required for the proper application of cross-validation. This is especially true if the ML processing pipeline includes model selection components.

The problem of overestimation of ML performance in the presence of a model selection component has already been investigated in several studies. For example, two recent studies have reported very alarming findings in the application of ML in speech, language, and hearing literature (Berisha et al., 2021; Vabalas et al., 2019). One of these studies surveyed the literature on the application of ML in autism and found a strong negative correlation between sample size and reported accuracy (Vabalas et al., 2019). Similar findings have been reported on speech-based classifications of individuals with Alzheimer's disease (Berisha et al., 2021) and for individuals with different forms of cognitive impairments (Berisha et al., 2021). These observed trends are at odds with theoretical expectations and a strong indication of widespread flawed evaluation of ML in the literature, indicating the possibility of overfitting and overestimation of the performance. A recent study has described this situation as a “crisis in ML-based science” (Kapoor & Narayanan, 2022) that was mainly attributed to data leakage, which often refers to the existence of some overlap (sometimes very subtle) between training and testing sets and not having “truly” independent test set(s). Application of cross-validation without a “true” test set in the presence of a model selection component is a common culprit for data leakage. Prior studies have shown that cross-validations with independent test sets are robust to overfitting and can provide an unbiased estimate of ML performance (Krstajic et al., 2014; Parvandeh et al., 2020; Varoquaux et al., 2017). Simulated data based on Gaussian distributions have been used to study the effect of different cross-validations in the presence of feature selection and model parameter optimization (Vabalas et al., 2019).

Power Analysis in ML

When evaluating different cross-validation methods to reduce overfitting, prior literature has quantified their effect on ML performance. However, power analysis of ML and the effect of different cross-validations on the required sample size have not been addressed. For example, a recent meta-analysis of ML in medical imaging applications found that, out of 167 articles, only four articles included some procedures for determining sample size (Balki et al., 2019). In contrast, power analysis before conducting a study is a common practice in science and a ubiquitous component of grant applications to the National Institutes of Health. These highlight the importance of this missing piece in knowledge-developing ML studies. In conventional statistical tests, power analysis allows researchers to estimate the minimum sample size required for achieving a target statistical power. This analysis is often done based on effect sizes reported in the literature. However, with ML, power analysis can be defined in multiple ways. This study divides them into two broad categories: “the required sample size” and “the recommended sample size.” The required sample size follows the conventional definition, and it is the “minimum” sample size required to have enough statistical power to reject the null hypothesis. In contrast, the recommended sample size is based on reaching a target performance for the ML model. A common target performance in the ML community has been the ML accuracy reaching a plateau, after which only marginal improvement is observed for increasing the sample size beyond that point (Figueroa et al., 2012).

A comprehensive review of the “recommended sample size” approach was presented in a recent study (Viering & Loog, 2022). This definition of recommended sample size and the adopted methodology in Figueroa et al. (2012) and its referenced studies neglect the effect size reported in the literature, which is a significant source of information and a very powerful aspect of power analysis in conventional statistical analysis. Also, these methods are more appropriate for active learning paradigms (Figueroa et al., 2012), where ML performance gets evaluated continuously as batches of new samples are added to the data set (often through labeling the data and not study participant recruitment), and hence not applicable to estimating the required sample size during the design of a study. The initial recommended data set for these methods is also very large (100–200 samples; Figueroa et al., 2012), and the final sample size could be much larger than this number. Such sample sizes are practical in tool-developing studies, where often the data just need to be labeled. However, knowledge-developing studies involve experimental design and active recruitment of participants, and such high sample sizes would often be impractical.

In summary, the prior sample size estimation studies were focused on the recommended sample size and not the required sample size. Additionally, the prior works did not consider the effect of different cross-validations on the sample size. The current study incorporates the traditional definition of power analysis into the ML framework. This new approach would allow researchers to use the effect sizes reported in the literature for power analysis during the design of an ML study. Finally, an alternative definition of recommended sample size based on the effect sizes reported in the literature is presented in this study that relates to the optimality of feature selection and its ability to find the correct subset of features. We will argue why this alternative definition of recommended sample size is important and more relevant for knowledge-developing ML studies.

Study Aim and Research Questions

Optimization (determining the value of a parameter given a particular desired outcome) is an integral component of ML models. Two different types of optimizations can be differentiated in the context of ML: first, determining the optimum decision boundary between different classes, which is part of the training process and, hence, present in any supervised ML implementation; second, model selection, which may include feature selection, hyperparameter optimization, optimization of the architecture of a (deep or shallow) neural network, selecting between different classification algorithms, and so forth. The focus of this study is on ML implementations that also include a model selection component. Also, we paid specific attention to feature selection in this study. However, the general findings should apply to other model selection scenarios.

The main aim of this study was to show that the choice of cross-validation method has a profound impact on the statistical power of ML and to provide a strong motivation for researchers in speech, language, and hearing sciences and other disciplines to migrate from the commonly used single holdout, train–validation–test, and k-fold cross-validations to more robust and powerful methods. Specifically, this study quantified the statistical relationship among different cross-validation methods, the minimum required sample size, the number of extracted features, the number of selected features, and the discriminative power of selected features. The outcome of this study could be used to determine the required and recommended sample sizes of an ML study during its experimental design, to select the appropriate dimensionality of the feature space for an existing data set, and to help scientists with the proper implementation of their ML processing pipeline to achieve generalizable ML models. Last but not least, the availability of power analysis method for ML could lead to a wider application of ML in clinical science where researchers may have previously thought that the sample sizes for their studies could be too small to apply ML. To pursue this aim, four research questions were asked in this study.

  • Q1: How is the statistical power of an ML model affected by the employed cross-validation method?

  • Q2: How is the statistical confidence in an ML model affected by the employed cross-validation method?

  • Q3: What is the minimum required sample size for achieving a statistically significant outcome from an ML model with conventional power requirements (5% significance, 80% power [α = .05, 1 − β = .8])?

  • Q4: What is the recommended sample size for achieving a target statistical confidence from the final ML model?

This article is organized as follows. First, the cross-validation methods that were investigated in this study are presented. Then, the performance of ML on a clinical data set is reported to demonstrate the profound effect of different cross-validations on the statistical properties of the trained ML models. The implications of these differences and the rationale for using simulated data for the rest of the article are also presented in this section. The framework for generating simulated data and statistical evaluation of ML models are presented next. The comparison of cross-validation approaches using simulated data section is devoted to reporting the outcomes of three different experiments designed to answer the four research questions. Experiments 1 and 2 answer research questions Q1 and Q2, respectively. Experiment 3 was conducted to answer research questions Q3 and Q4. The results of the experiments are put into the larger context in the Discussion section, where more general implications of this work are presented, along with limitations of this study, generalization of its findings, and directions for future studies.

Method

Cross-Validation Methods

Feature selection is an optimization problem, and the cost function (the evaluation metric that is to be maximized) could be the accuracy of the classifier (Theodoridis & Koutroumbas, 2009). However, depending on which cross-validation is implemented, the cost function is computed differently. Therefore, our main hypothesis was that the choice of cross-validation would significantly affect the statistical characteristics (average accuracy, standard devition accuracy, and confidence) of the final optimized ML model. Although cross-validation comes in several different flavors, they can be categorized into two main types: two-split and three-split cross-validations. In the two-split method, the data set is divided into (a series of) two separate partitions of training and testing sets, whereas in the three-split method, the data set is divided into (a series of) three separate partitions of training, validation, and testing sets. This article compares two two-split methods and two three-split methods of cross-validation. MATLAB codes of our implementation are provided in the Supplemental Material S4.

Single Holdout Method

The single holdout approach is a two-split method and the simplest cross-validation approach. In this approach, based on a target splitting ratio, the data set is randomly divided into two disjoint sets: a training set and a testing set. The training set is used for creating the model. The accuracy of the classifier is computed by applying the trained model to the testing set. Therefore, the cost function of feature selection is the accuracy of the testing set. This means the feature set that maximizes the testing accuracy will be selected for the final optimized ML model. The main advantages of this approach are its implementation simplicity and low computational complexity. Figure 1A shows a schematic of this approach. In this study, we used a stratified version of the single holdout cross-validation to ensure that both training and testing sets had the same proportion of samples from each class, with a splitting ratio of 70% (training) to 30% (testing).

Figure 1.

The image illustrates data splitting for the 4 cross validation methods. A. For the single holdout method, the data is split into the testing and training sets. B. For the k fold method when k equals 3, the data split is as follows. 1. The data is split into testing and training sets. 2. The data is split into training, testing, and training sets. 3. The data is split into the training and testing sets. C. For the train validation test method, the data split is as follows. The data is split into the testing and training-validation sets. The training-validation set undergoes the following splits. 1. Validation and training. 2. Training validation and training. 3. Training and validation. D. For the nested k fold cross validation method, when k equals 3, the data is split between the outer k folds and the inner k folds for fold 1 to fold k. The outer k folds consists of the following data splits. 1. Testing and training-validation. 2. Training-validation, testing, followed by training- validation. 3. Training-Validation followed by testing. The data splits for the inner k folds consists of the following. 1. Validation and Training. 2. Training, validation, and training. 3. Training and validation.

Visualization of data splitting for the four cross-validation methods studied: (A) single holdout, (B) k-fold (k = 3), (C) train–validation–test, and (D) nested k-fold cross-validation (k = 3).

k-Fold Cross-Validation

The k-fold cross-validation approach is also a two-split method, where the data set is divided randomly and evenly into k disjoint sets. Each set contains 2n/k samples (n is the sample size per each class) and is considered the testing set for that fold. All the remaining samples constitute the training samples for that fold. Similar to the single holdout, the training set from each fold is used for training the model, and then that model is applied to the corresponding testing set to get an estimate of model performance. However, in contrast with the single holdout method, the training/testing process is performed multiple (k) times. The cost function of feature selection is the average accuracy over the k different testing sets. This means that the feature set that maximizes the average testing accuracy will be selected for the final optimized ML model.

The main advantages of this approach are that all samples contribute to the training and testing of the ML model, and the robustness of model performance can be obtained by computing the standard deviation of the k testing accuracies. Figure 1B shows a schematic of this approach with k = 3 folds for visualization purposes. In this study, we used stratified 10-fold cross-validation. The selection of k = 10 was motivated by its common use in the literature and a recent study supporting the suitability of this selection (Marcot & Hanea, 2021).

Train–Validation–Test Cross-Validation

The train–validation–test approach is a three-split method, where the data set is first split into two disjoint sets termed training–validation and testing sets. The training–validation set is then further split into two disjoint sets of training and validation sets using single holdout, k-fold cross-validation, or a different method. An ML model is trained using the training set and is then applied to the associated validation set to obtain an estimate of model performance. However, unlike the previous two-split methods, the cost function of feature selection is either validation accuracy (in the case of single holdout) or average validation accuracy over multiple splits (in the case of k-fold cross-validation). This means that the feature set that maximizes accuracy across validation set(s) will be selected for the final optimized model. The selected features are then used to train the final model using both training and validation samples. Critically, the final model is then applied to the testing samples to obtain a single value of model performance on a disjoint testing set.

The main advantage of this approach is that the testing set is completely independent of the model optimization process (i.e., feature selection in our case). Figure 1C shows a schematic of this approach with k = 3 folds within the training–validation set for visualization purposes. In this study, we used a stratified version and reserved 15% of the samples for the testing set. The remaining 85% of samples were divided into training and validation sets using 10-fold cross-validation.

Nested k-Fold Cross-Validation

The nested k-fold cross-validation approach is also a three-split method and is, basically, a k-fold (inner) cross-validation embedded inside another k-fold (outer) cross-validation. The nested k-fold cross-validation method thus consists of an outer k-fold cross-validation that divides the data into k different folds, with each outer fold consisting of training–validation and testing sets. Each training–validation set undergoes a secondary (inner) k-fold data splitting to create training and validation sets. In this sense, the testing sets of the outer folds represent multiple independent testing sets, the testing sets of the inner folds are intermediate validation sets, and the training sets of the inner folds are the training sets. This means that, unlike the previous three methods, feature selection will be executed k different times (one per each outer fold) and hence we will have k different estimates for the best feature set. In each outer fold, the feature set that maximizes the average accuracy over the k different validation sets (one for each inner fold) is selected as a candidate best feature set. Also, in contrast with the previous cross-validation methods, a consensus criterion is needed to select the best feature set among the k candidate feature sets (one from each outer fold; Parvandeh et al., 2020). Figure 1D shows a schematic of nested k-fold cross-validation with three outer folds and three inner folds per outer fold. In this study, we employed k = 10 and selected the feature set that had the highest probability of joint occurrence among the k different best feature sets (one per outer fold).

Comparing Cross-Validation Methods in a Real-World Clinical Example

The significant impact of different cross-validations is demonstrated here using ambulatory voice monitoring data. A primary focus of our research group has been determining the real-world vocal behavior/function associated with hyperfunctional voice disorders using ambulatory voice monitoring technology (Ghassemi et al., 2014; Mehta et al., 2015; Van Stan et al., 2020). To that end, six different voice measures capturing different aspects of phonation were extracted from the ambulatory data of patients with PVH and vocally healthy matched controls. The extracted measures included fundamental frequency (F0), cepstral peak prominence, sound pressure level, the difference between amplitudes of the first and second harmonics of the signal (H1–H2), smoothed duration of phonatory segments (voicing duration), and duration of nonphonatory segments (resting duration). The smoothed duration of phonatory segments was estimated from the outcome of the voice activity detector by connecting contiguous voiced segments that were separated by half a second or less of nonphonatory activity. No smoothing was applied to the durations of the nonphonatory segments. The methodology of processing the ambulatory data and the definition of these measures have been reported in an earlier study (Mehta et al., 2015). Similar to prior work (Ghassemi et al., 2014; Mehta et al., 2015; Van Stan et al., 2020), the day-long distribution of each measure was constructed and eight different distributional characteristics (mean, median, standard deviation, interquartile range, 5th percentile, 95th percentile, skewness, and kurtosis) were extracted. The final classification features were computed as distributional characteristics of each participant, averaged over all recording days, meaning each participant contributed only 1 data point to the data set. Therefore, the dimensionality of the feature space was 48 (six voice measures with eight distributional characteristics). Our data set included recordings from 153 females with PVH and 136 female controls. The data set was balanced by selecting 136 patients randomly. To make the findings robust to the random selection of patients and random splitting of the data, each analysis was repeated 500 times.

First, the average and standard deviation of testing accuracy for different numbers of selected features and different cross-validation methods are presented in Figure 2A. The results showed dissimilar trends for different cross-validation. Specifically, two-split cross-validation methods (single holdout, 10-fold) yielded higher average accuracy compared to the three-split cross-validation methods (train–validation–test, nested 10-fold), with single holdout tending to yield the highest average accuracy for each number of selected features. Another significant observation was the impact of different cross-validation methods on the variability of estimated testing accuracy. Specifically, testing accuracies of the methods with a single testing set (single holdout, train–validation–test) were quite sensitive to the random splitting of the data (larger error bars), whereas methods with multiple testing sets (10-fold, nested 10-fold) were quite robust (smaller error bars). In summary, the selection of cross-validation method has a significant impact on the distributional characteristics of the estimated performance of ML. The implications of these observations on the generalizability of the findings and the statistical power of the model will be discussed at the end of this section.

Figure 2.

2 graphs. The legend for the graphs is as follows. Dashed blue: Single holdout. Dashed red: 10 fold. Solid yellow: Train validation test. Dashed purple: Nested 10 fold. A. The first graph plots the accuracy with respect to the feature number. The feature number ranges from 1 to 10 in unit increments. The accuracy ranges from 0.6 to 0.9 in increments of 0.05. The dashed blue curve passes through (1, 0.72), (5, 0.82), and (10, 0.85). The dashed red curve passes through (1, 0.71), (5, 0.81), and (10, 0.83). The solid yellow curve passes through (1, 0.7), (5, 0.76), and (10, 0.78). The dashed purple curve overlaps the solid yellow curve. B. The second graph plots the accuracy with respect to the pair number. The pair number ranges from 25 to 125 in increments of 25. The dashed blue curve passes through (25, 0.9), (75, 0.83), and (125, 0.81). The dashed red curve passes through (25, 0.82), (75, 0.8), and (125, 0.8). The solid yellow curve passes through (25, 0.66), (75, 0.71), and (125, 0.75). The dashed purple curve almost overlaps the solid yellow curve. All values are estimates.

(A) The effect of different cross-validations on the average and standard deviation of testing accuracy for different numbers of selected features. (B) The effect of sample size on testing accuracy estimated based on different cross-validations.

Second, the effect of changing the sample size from 25 pairs to 125 pairs on the testing accuracy of the model with four selected features was evaluated. Similar to the previous analysis, the participants were selected randomly, and each analysis was repeated 500 times. The average and standard deviation of testing accuracy for different sample sizes are presented in Figure 2. Interestingly, the estimated performances of cross-validations with two splits (single holdout, 10-fold) declined as the sample size was increased, which does not follow the theoretical expectation. Conversely, testing accuracies of cross-validations with three splits (train–validation–test, nested 10-fold) increased as the sample size was increased. The implications of these observations on the generalizability of the findings and the statistical power of the model will be discussed at the end of this section.

Third, as we discussed earlier, the knowledge that is gained (vs. the trained model itself) is the primary objective of knowledge-developing ML studies. Looking at features that are selected and the orders in which they get selected is an important source of acquiring such knowledge. For example, in this clinical case study of ambulatory voice monitoring, each of the six voice measures relates to a different aspect of phonation. The finding that a distributional characteristic of H1–H2 is the best feature versus a distributional characteristic of F0 would have different implications regarding the etiology and/or compensatory behavior associated with PVH. To demonstrate the effect of different cross-validation methods on the outcome of feature selection, the probability of different features being selected at different steps of feature selection was computed from the 500 repetitions of the first analysis of this section. In this fashion, the first feature was the feature that provided the highest classification accuracy. The second feature was the feature (among the remaining ones) that provided the highest improvement in classification accuracy. The third and the fourth features were defined in a similar way.

Figure 3 shows the results. According to Figure 3A, the outcome of feature selection for the single holdout method was very unstable (low selection probabilities), meaning the outcome of feature selection was very sensitive to how the data were split into training and testing sets. That is, depending on the samples that were in the training set, different features (and even from different voice measures) were selected as the most informative (i.e., the highest classification accuracy) descriptors of PVH, which indicates low generalizability for the trained model. The situation for the train–validation–test method was better, where the first and second features were consistently selected (high selection probabilities). However, the third and fourth features had low selection probabilities. Lastly, 10-fold and nested 10-fold cross-validation methods produced models that were much more consistent, with the highest consistency of feature selection exhibited by nested 10-fold cross-validation.

Figure 3.

4 graphs compare the selection probabilities of the 4 cross validation methods by features and feature index. The legend for the features is as follows. Blue: first feature. Red: second feature. Yellow: third feature. Purple: fourth feature. The interpretation of the range of values of the feature index is as follows. 0 to 8: F 0. 9 to 16: CPP. 17 to 24: SPL. H1-H2: 25 to 32. Voicing duration: 33 to 40. Resting duration: 41 to 48. 1. A. Single holdout. The first feature has selection probabilities of 0.2 and 0.38 for SPL and H1-H2, respectively. B. 10 fold. The first feature has a selection probability of 1 for H1-H2. The second feature has a selection probability of 0.93 for H1-H2. The third feature has a selection probability of 0.67 for SPL. The fourth feature has a selection probability of 0.71 for CPP. C. Train validation test. The first feature has a selection probability of 0.95 for H1-H2. The second feature has a selection probability of 0.73 for H1-H2. The third feature has a selection probability of 0.33 for SPL. The fourth feature has a selection probability of 0.37 for CPP. D. Nested 10 fold. The first feature has a selection probability of 1 for H1-H2. The second feature has a selection probability of 0.95 for H1-H2. The third feature has a selection probability of 0.81 for SPL. The fourth feature has a selection probability of 0.86 for CPP.

A comparison of statistical confidence of four cross-validation methods: (A) single holdout, (B) 10-fold, (C) train–validation–test, and (D) nested 10-fold. The horizontal axes represent the 48 dimensions in the feature space (the six voice measures listed under the panels with eight distributional characteristics per voice measure). Within each voice measure category on the plots, the distributional characteristics, in order, are the mean, median, standard deviation, interquartile range, fifth percentile, 95th percentile, skewness, and kurtosis. CPP = cepstral peak prominence; SPL = sound pressure level.

These three analyses showed the impact of different cross-validation methods on the ML outcomes, highlighting the importance of employing appropriate cross-validation for ML in health care. Based on the result of Figure 2A, the average and standard deviation of the distribution of classification accuracy depended on how cross-validation was implemented. Figure 2B showed that increasing the sample size had dissimilar effects on the distribution of the accuracy of the classifier for different cross-validations. The statistical power of ML directly relates to the distribution of its accuracy; therefore, we can expect different statistical power for different implementations of cross-validations. Now the first question (research question Q1) is which method has the highest statistical power (and hence requires the smallest sample size). To answer this question, we need to systematically change the discriminative power of features, the sample size of the data set, the dimensionality of feature space, and so forth. Also, we need to have thousands of independent examples of each case so we can generate the distribution of the performance that is required for power analysis. Therefore, simulated data were used for the rest of this article.

Based on the results of Figure 3, the statistical confidence of the model varied significantly among the different cross-validation methods. The next question (research question Q2) is which method gives us the correct answer. With real data, we could not be sure which features are the most discriminative ones. Also, we cannot systematically change the discriminative power of features, the sample size of the data set, the dimensionality of feature space, and so forth and determine the recommended sample size for meeting a target model confidence. In contrast, with simulated data, we know the underlying distribution of each feature and, hence, which features should be selected. Also, we can change the parameters of the problem systematically. As a final note, more than 20 different voice measures have been presented in prior ambulatory voice monitoring studies (Cortés et al., 2018; Mehta et al., 2015; Van Stan et al., 2020). Additionally, up to 10 different distributional characteristics have been used in prior studies. This would add up to more than 200 possible features. A practical follow-up question would be, how many features can be extracted for a given sample size to have statistical confidence in the findings? The following simulation study will present methodologies to answer these important and practical questions.

Statistical Framework for ML Model Evaluation

This study assumed a balanced two-class classification problem (e.g., detecting the presence of a disorder in an equal number of patients and controls in a clinical application) with a sample size of 2n participants (n per class). A feature selection component was also included in the ML processing pipeline. Therefore, this study assumed a total of m independent computed features, of which only l were discriminative. All analyses were executed on a server running a Windows 2019 Server, with two Xenon Gold 6230 CPUs, and 768 GB of RAM. We used MATLAB R2021b for our analyses.

Given the aims of this study, features of positive (e.g., patient group) and negative (e.g., healthy control group) classes were generated from a series of multivariate Gaussian distributions. Additionally, to make the simulation tractable, we assumed all discriminative features to have the same discriminatory power of d (Cohen's d effect size metric; Cohen, 2013). Nm denotes an m-dimensional Gaussian distribution with μPμN being an m × 1 vector of means and ΣPΣN being an m × m matrix storing the covariance of the positive (negative) samples, where m is the number of features extracted. Equations 1 and 2 show the formula for data generation, where XP and XN are n × m matrices storing n samples from the positive class and n samples from the negative class, respectively.

XP~NmμPΣP (1)
XN~NmμNΣN (2)

For the negative class, we set μN=000T, and for the positive class, we used μP=dd00T, with the first l elements of μP equal to d and the remaining m − l elements equal to 0. Additionally, we used ΣN=ΣP=Im, where Im is an identity matrix with the size of m × m.

Comparing the definition of the XP and XN shows that only the first l features of the positive and negative classes had different distributions and hence were discriminative, and the remaining m − l features were irrelevant. The goal is then to find those l features and then train a classifier on them. This goal is an optimization problem and can be broken into two separate components: (a) selecting a proper cost function and (b) adopting a strategy for searching the solution space. This study adopted a wrapper-forward feature selection method due to its high performance and low computational cost (Ghasemzadeh, 2019a). The cost function of a wrapper method is the performance of a classifier. This means the subset of features that maximizes the classification performance gets selected. Forward feature selection refers to the search strategy. It is an iterative greedy search algorithm that starts with an empty selected set. At each stage, the remaining features are sequentially added to the selected set. The feature that leads to the maximum improvement of the cost function is added to the selected set. The process is repeated until the addition of new features does not improve the cost function. Finally, logistic regression was used as the classifier. Applying a logistic regression was primarily motivated by its widespread usage in speech, language, and hearing sciences, the simplicity of its algorithm that offers interpretability and produces a closed-form model, and the underlying distributions of the positive and negative classes. Specifically, the positive and negative samples were generated based on multivariate Gaussian distributions with the same covariance but different means; therefore, the true decision boundary should be a hyperplane. Logistic regression implements such a decision boundary and hence would be optimal for our problem. See Table 1 for a summary of symbols and parameters used in this study.

Table 1.

Definition of symbols and parameters used in this study.

Symbols Definition
H0 The null hypothesis
Ha The alternative hypothesis
α Probability of type I error
β Probability of type II error (1 − β is called power)
Nm m-dimensional Gaussian distribution
P Samples from the positive class
N Samples from the negative class
μP An m × 1 vector of means of Nm for positive samples
ΣP An m × m matrix storing the covariance of Nm for positive samples
Im Identity matrix with the size of m × m
µN An m × 1 vector of means of Nm for negative samples
ΣN An m × m matrix storing the covariance of Nm for negative samples
m The dimensionality of the feature space
n Sample size per each class in a balanced data set
l The number of discriminative (i.e., selected) features
d Cohen's d effect size of discriminative features
k The number of folds in k-fold and nested k-fold cross-validation
Cl,d Statistical confidence of the model (the probability of correct selection of d features out of l)
θi Parameters of fitted exponential curves
r Linear correlation coefficient
p p value
nr The required sample size to get a statistically significant outcome with conventional power requirements (α = .05, 1 − β = .8)
γdb The ratio of the number of positive to the number of negative samples in an unbalanced data set
γd The ratio of the larger Cohen's d to the smaller one if selected features had different discriminative powers.

Statistical Evaluation Criteria

To answer Q1 and Q3, the null (H0) and alternative (Ha) hypotheses need to be formed. There is an implicit assumption in the application of supervised ML that the extracted features have better discriminative power compared to random features. This assumption can be used to formulate H0 and Ha for a supervised ML problem.

Ha: Extracted features will perform significantly better than random features.

To formally test this directional hypothesis, we need to construct the distributions of Ha (i.e., the performance of the extracted features) and H0 (i.e., the performance of random features). For this study, the distribution of Ha can be estimated using Monte Carlo simulations by generating a series of XP (Equation 1) and XN (Equation 2) and passing them through our ML processing pipelines. The distribution of H0 can be estimated similarly by setting d = 0. Once distributions of H0 and Ha are estimated, they can be used for quantitative evaluations of different ML processing pipelines (e.g., different cross-validation approaches). This study has used the following evaluation criteria.

Distributional Characteristics of H0

The distributional characteristics (mean, standard deviation, confidence interval [CI]) of H0 are evaluation criteria that provide valuable insights into the statistical properties of different ML processing pipelines. For example, the mean accuracy of H0 for a balanced two-class problem should be 0.5 (on a scale of 0–1), and any gain over that number is often interpreted as the discriminative power of the extracted features. This interpretation is based on an implicit assumption that the implemented ML processing pipeline is not biased. The distribution of H0 would allow us to evaluate if such a bias exists and then quantify it. Additionally, based on a specific statistical significance level (α), we can find the CI of H0. Investigating Ha shows that the alternative hypothesis is directional. Therefore, the CI of H0 would be the range of values covering the minimum of the H0 distribution to its (1 − α) × 100 percentile. Assuming a normal distribution, the mean and variance of H0 contribute to its CI; therefore, the standard deviation of the H0 distribution was another evaluation criterion for comparing different cross-validation approaches.

Statistical Power and the Required Sample Size

Statistical power is the probability of a test producing true positive results, that is, the probability of detecting the presence of an effect if effect exists (Field et al., 2012). The distribution of Ha and the target probability of Type II error (β) are the factors determining the power of a test, where power is defined as 1 − β. Statistical power is closely related to the required sample size for a study to detect a significant effect with a reasonable probability. The required sample size depends on the H0 and Ha distributions and the selected values for α and β. To determine the required sample size for different cross-validation approaches, the distributions of H0 and Ha were estimated for different values of n (Equations 1 and 2). Then, for a given value of α, the upper bounds of the H0 CI, and, for a given value of β, the lower bounds of the Ha CI were estimated and plotted. The point of intersection between the two plots corresponds to the minimum required sample size. Figure 4 illustrates this for the commonly used values of α = .05 and β = .2 (power of .8), which were used throughout this study.

Figure 4.

A plot of the accuracy on the y axis and the pair number, n on the x axis. The y axis ranges from 0.48 to 0.62 in increments of 0.02. The x axis ranges from 50 to 500 in increments of 50. The legend is as follows. Solid blue curve: H subscript 0, alpha equals 0.05. Dashed red curve: H subscript a, beta equals 0.2. The blue curve passes through the points (50, 0.605), (210, 0.555), and (500, 0.535). The dashed red curve passes through the points (50, 0.48), (210, 0.555), and (500, 0.583). All values are estimates.

Empirical estimation of the minimum required sample size as the value of n at the intersection between the upper bound of the H0 confidence interval (given α = .05) and the lower bound of the Ha confidence interval (given β = .2).

Statistical Confidence

ML is a statistical tool, and hence, there is a chance that the outcome of an ML model provides incorrect results. Such inaccuracies are often quantified using performance metrics such as accuracy, sensitivity, and specificity. While it is well understood that the classifier outcome could be incorrect, the correctness of the model itself has received less attention. The correctness of the model can be defined in terms of the features that have been included (i.e., selected) in the model, as well as the coefficients (e.g., feature weights) of the model. The focus of this study is on the correctness of selected features.

Statistical confidence of an ML model is defined as the probability that correct features were selected and hence included in the ML model. In practice, statistical confidence is the probability of correct features being included in the final model if we repeat our experiment many times with different samples drawn from the population. This concept is especially important as it will have direct consequences on the interpretation that we make about models and the knowledge that we gain from them. As an example, let us consider the application of ML for understanding differences in vocal function and behavior between patients with a voice disorder and vocally healthy individuals. Depending on which features are being selected, we may attribute and rank different aspects of vocal function and/or behavior as the most likely contributors to causing and/or maintaining the disorder. Now, if the model has low confidence and other researchers try to replicate the study, there is a high probability that they will find very different results. This means, our conclusions about factors associated with the disorder could be unreliable, and hence, we may target incorrect (or at least suboptimal) aspects of phonation during efforts to prevent and treat the disorder. Statistical confidence is a metric that measures such an attribute. Referring to the definition of XP and XN (Equations 1 and 2), we know which features should be included in the final model; hence, the effect of different parameters and cross-validation approaches on the statistical confidence of the model can be quantified. The symbol Cl,d will be used to represent the statistical confidence of the model for different numbers of selected features. Specifically, Cl,d is the probability that, among the l selected features in an ML model, d of the features were correctly selected.

Comparison of Cross-Validation Approaches Using Simulated Data

Experiment 1: Distributional Characteristics of H0 and Statistical Power

Experiment 1 addressed research question Q1 (“How is the statistical power of an ML model affected by the employed cross-validation method?”) by investigating the effect of different parameters on obtaining statistically significant outcomes from a supervised ML model. First, the effects of the four cross-validation methods and different sample sizes on the distribution of H0 were investigated. Values of m = 20 (20 extracted features), l = 2 (two selected features), and d = 0 were used. Figure 5 shows the results of running each experiment 5,000 times, with each experiment defined by a given cross-validation method and sample size per class (50–500 samples, in increments of 50). The plots show the outcome of fitting a series of two-term exponential functions on the result of each cross-validation. The form of the exponential function was:

y=θ1eθ2x+θ3eθ4x (3)

where x and y are the predictor and outcome variables, respectively, and θ1, θ2, θ3, θ4 are parameters of the exponential function.

Figure 5.

2 graphs. In both graphs, the x axis represents the pair number and it ranges from 50 to 500 in increments of 50. The legend is as follows. Dotted blue: Single holdout. Dashed red: 10 fold. Solid yellow: Train validation test. Dashed purple: Nested 10 fold. Graph 1. In the first graph the y axis represents the accuracy and it ranges from 0.35 to 0.75 in increments of 0.05. The dotted blue curve runs between (50, 0.68) and (500, 0.56). The dashed red curve runs between (50, 0.64) and (500, 0.54). The solid yellow and the dashed purple curves overlap and run between (50, 0.5) and (500, 0.51). Graph 2. In the second graph, the y axis represents the one sided 95 percent cutoff and it ranges from 0.5 to 0.8 in increments of 0.05. The dotted blue curve runs between (50, 0.76) and (500, 0.59). The solid yellow and the dashed red curves are very close to each other and run near (50, 0.68), and (500, 0.54). The dashed purple curve runs between (50, 0.63) and (500, 0.54). All values are estimates.

H0 characteristics: (A) mean and standard deviation of H0 accuracy and (B) upper bound of H0 CI (one-sided, α = .05). CI = confidence interval.

In Figure 5A, the mean accuracies of the single holdout and 10-fold cross-validations significantly deviate from the theoretically expected value of 0.5 for a balanced binary classification with irrelevant (nondiscriminative) features. However, as the sample size increases, the mean accuracies of these two approaches start to converge toward the expected chance value of 0.5, indicating the accuracy that we saw at smaller sample sizes was due to overfitting and not learning the underlying patterns of the data (as there are no patterns for H0). On the other hand, the average accuracies of the other two methods are quite close to the expected value of 0.5, even for small sample sizes. Comparing the mean of H0 accuracy among different cross-validations clearly shows that the lack of an independent test set in the presence of an optimization task (e.g., feature selection, hyperparameter optimization, selection of the architecture of a neural network, etc.) leads to significant overestimation of the performance and should be avoided.

The whiskers of Figure 5A depict the standard deviation of the accuracy of different approaches across the 5,000 experimental trials. The most significant observations are that the train–validation–test and 10-fold cross-validation methods had the highest and the lowest standard deviations, respectively. Also, we see a significant decrease in the value of dispersion as we move from the simple train–validation–test to the more sophisticated method of nested 10-fold cross-validation. Finally, the upper bound of the H0 CI (one-sided, α = .05) depends on both the mean and the dispersion of the H0 distribution. Theoretically speaking, the upper bound increases with increases in the mean and dispersion of the H0 distribution. This is also reflected in Figure 5B. Based on this figure, for all approaches, the upper bound of H0 decreases monotonically with an increase in the sample size. However, we see significant differences among the approaches, especially in lower sample sizes. For example, with 50 pairs, the single holdout method has a significant chance (more than 5%) to get accuracies as high as 76.7% even if the features are irrelevant (i.e., random features). However, this value drops to 62% if nested 10-fold cross-validation is used instead.

Second, the interaction effect between the dimensionality of the feature space (m) and different cross-validation methods was investigated. Figure 6 shows the results. Based on Figure 6A, the bias (the deviation from the expected theoretical mean accuracy of 0.5) for single holdout and 10-fold cross-validation increases significantly as the number of extracted features increases, whereas methods with independent test sets (train–validation–test and nested k-fold) remain unbiased. Figure 6B represents the dependency of the upper bound of H0 CI. Based on this figure, the upper bound of H0 CI increases with the number of extracted features for the single holdout and 10-fold cross-validations, whereas this upper bound remains relatively constant for the methods with independent test sets.

Figure 6.

2 graphs. In both graphs, the x axis represents the number of extracted features, m and it ranges from 5 to 40 in increments of 5. The legend is as follows. Solid blue: Single holdout, 100 pairs. Dotted blue: Single holdout, 200 pairs. Solid red: 10 fold, 100 pairs. Dashed red: 10 fold, 200 pairs. Solid yellow: Train validation test, 100 pairs. Dashed yellow: Train validation test, 200 pairs. Solid purple: Nested 10 fold, 100 pairs. Dashed purple: Nested 10 fold, 200 pairs. A. Graph 1. In the first graph, the y axis represents the mean accuracy and it ranges from 0.5 to 0.75 in increments of 0.05. The solid blue curve passes through (5, 0.57) and (40, 0.65). The dotted blue and solid red curves are very close to each other and pass near (5, 0.55) and (40, 0.59). The dashed red curve passes through (5, 0.54) and (40, 0.57). The solid purple and the dashed purple curves remain constant at 0.5. B. Graph 2. In the second graph, the y axis represents the upper bound of H subscript 0 C I. The solid blue curve passes through (5, 0.65) and (40, 0.72). The solid yellow curve is constant at 0.64. The dashed yellow curve is constant at 0.6. The solid purple curve is constant at 0.57. The dashed purple curve is constant at 0.55. The dotted blue curve passes through (5, 0.61) and (40, 0.64). The solid red curve passes through (5, 0.6) and (40, 0.64). The dashed red curve passes through (5, 0.56) and (40, 0.6). All values are estimates.

Effect of the dimension of the feature space (m) on (A) mean H0 accuracy and (B) upper bound of H0 CI (one-sided, α = .05). CI = confidence interval.

Finally, the statistical powers of the different cross-validation methods were compared. To that end, the minimum required sample size for a study to detect a statistically significant effect (α = .05) with typical statistical power (1 − β = .8) was determined for each experimental condition. Figure 7 shows the results, where the markers are the outcome accuracies of the experiments, and the curves are polynomials of order 2 fitted on them. The portions of the figures that fall below the dashed black line correspond to underpowered cases. The results indicate that the train–validation–test and the single holdout methods have the lowest statistical power. For example, the train–validation–test method (see Figure 7C) with 50 pairs of samples would be underpowered (it's below the dashed black line, meaning 1 − β < .8) even with two highly discriminative features (d = 1) in our 20-dimensional feature space. The situation for the single holdout method (see Figure 7A) is slightly better. A different interpretation of the results would be, if we have two relatively good features with d = 0.6 in our 20-dimensional feature space, the train–validation–test needs at least 200 pairs of samples for having an 80% chance of getting a significant outcome, whereas the nested 10-fold method (see Figure 7D) only needs 100 pairs of samples. Finally, the 10-fold cross-validation (see Figure 7B) has slightly higher statistical power compared to the nested 10-fold method. The most likely reason for this observation would be the fact that nested k-fold uses fewer samples during the training and feature selection (in our case, 90%), whereas 10-fold cross-validation takes advantage of all the data.

Figure 7.

4 graphs plot the accuracy of the 4 cross validation methods as a function of the pair number. The legend is as follows. Dashed black: H subscript 0, d equals 0, alpha equals 0.05. Solid blue: H subscript a, d equals 0.4, beta equals 0.2. Solid yellow: H subscript a, d equals 0.6, beta equals 0.2. Solid pink: H subscript a, d equals 0.8, beta equals 0.2. Solid green: H subscript a, d equals 1, beta equals 0.2. A. Graph 1. Single holdout. The dashed black curve passes through (50, 0.75) and (500, 0.59). The solid blue curve passes through (50, 0.66) and (500, 0.61). The solid yellow curve passes through (50, 0.7) and (500, 0.65). The solid pink curve passes through (50, 0.74) and (500, 0.7). The solid green curve passes through (50, 0.76) and (500, 0.74). B. Graph 2. 10 fold. The dashed black curve passes through (50, 0.69) and (500, 0.55). The solid blue curve passes through (50, 0.62) and (500, 0.6). The solid yellow curve passes through (50, 0.65) and (500, 0.65). The solid pink curve passes through (50, 0.69) and (500, 0.7). The solid green curve passes through (50, 0.74) and (500, 0.75). C. Graph 3. Train validation test. The dashed black curve passes through (50, 0.69) and (500, 0.56). The solid blue curve passes through (50, 0.44) and (500, 0.56). The solid yellow curve passes through (50, 0.5) and (500, 0.64). The solid pink curve passes through (50, 0.56) and (500, 0.69). The solid green curve passes through (50, 0.63) and (500, 0.74). D. Graph 4. Nested 10 fold. The dashed black curve passes through (50, 0.63) and (500, 0.54). The solid blue curve passes through (50, 0.48) and (500, 0.59). The solid yellow curve passes through (50, 0.54) and (500, 0.65). The solid pink curve passes through (50, 0.61) and (500, 0.7). The solid green curve passes through (50, 0.66) and (500, 0.75). All values are estimates.

Comparison of the statistical power of the four cross-validation methods given varying simulated effect sizes and sample sizes. The portions of the figures that fall below the dashed black line correspond to underpowered cases: (A) single holdout, (B) 10-fold, (C) train–validation–test, and (D) nested 10-fold.

Experiment 2: Statistical Confidence of the Model

Experiment 2 addressed research question Q2 (“How is the statistical confidence in an ML model affected by the employed cross-validation method?”) by investigating the effect of different parameters on the statistical confidence of the final trained model. First, the effects of different cross-validation methods and sample sizes were investigated. Values of m = 20 (20 extracted features) and d = 0.8 were used. The generated feature set was then passed through feature selections with different cross-validations, and the two most discriminative features (l = 2) were determined. Each simulation was repeated 5,000 times, and then the probability of finding the correct features was computed.

Figure 8A shows the probability of at least one of the selected features being correct (C2,1), and Figure 8B represents the probability of both selected features being correct (C2,2). Based on these figures, models generated based on single holdout and nested 10-fold cross-validations had the lowest and highest statistical confidence, respectively. For example, the probability of both features being correct with 100 pairs is only about 20% with the single holdout (i.e., there is an 80% chance that the final model is based on at least one incorrect feature). However, this number increases to about 80% in nested 10-fold cross-validation. Additionally, comparing Figures 8A and 8B shows that the gap between the two methods increases as we aim for a correct multidimensional model (here the dimension of two). The nested 10-fold method has a slightly better performance compared to the 10-fold cross-validation method. This could be attributed to the voting mechanism described in the Cross-Validation Methods section for aggregating the outcomes of feature selection from different outer folds. Finally, the plots exhibit a saturation phenomenon, indicating there could be a “sweet spot” in terms of the recommended sample size required for getting a robust outcome.

Figure 8.

2 graphs. The legend for the graphs is as follows. Dotted blue: Single holdout. Dashed red: 10 fold. Solid yellow: Train validation test. Dashed purple: Nested 10 fold. In both the graphs, the x axis represents the pair number and it ranges from 50 to 500 in increments of 50. Graph 1. The y axis represents the probability of at least one being correct in percentage. The dotted blue curve passes through (50, 45), (250, 70), and (500, 82). The dashed red, solid yellow, and dashed purple curves almost overlap and pass near (50, 70), (250, 100), and (500, 100). Graph 2. The y axis represents the probability of both being correct in percentage. The dotted blue curve passes through (50, 10), (250, 45), and (500, 65). The dashed red, solid yellow, and dashed purple curves almost overlap and pass near (50, 41), (150, 90), (300, 100), and (500, 100). All values are estimates.

Effect of different cross-validation methods on the statistical confidence of the trained model: (A) (C2,1) and (B) (C2,2).

Next, the interaction effect between the dimensionality of the feature space (m) and cross-validation methods on the statistical confidence of the model was investigated. For this analysis, the value of d = 0.8 was used, and each simulation was repeated 5,000 times. Figure 9A shows the probability of at least one of the selected features being correct (C2,1), and Figure 9B represents the probability of both selected features being correct (C2,2). Based on these figures, the dimensionality of the feature space (m) and statistical confidence of the model are inversely related to each other. However, different cross-validation methods exhibit dissimilar behaviors. Specifically, whereas the single holdout is very sensitive to the dimensionality of the feature space, nested 10-fold cross-validation is very robust. Finally, there is an interaction effect between the sample size and the dimensionality of the feature space on the statistical confidence of the model. For example, if we have 100 pairs of samples and increase m from 5 to 40, C2,2 drops from 91% to 73%, respectively, in nested 10-fold. However, if we increase the sample size to 200 pairs, we will only see a decrease of 4 percentage point in statistical confidence. This result indicates that, for a feature with a target effect size (d), there could be a “sweet spot” in terms of the recommended sample size required for statistical confidence to become robust to an increase in the dimensionality of the feature space.

Figure 9.

2 graphs. The legend is as follows. Solid blue: Single holdout, 100 pairs. Dotted blue: Single holdout, 200 pairs. Solid red: 10 fold, 100 pairs. Dashed red: 10 fold, 200 pairs. Solid yellow: Train validation test, 100 pairs. Dashed yellow: Train validation test, 200 pairs. Solid purple: Nested 10 fold, 100 pairs. Dashed purple: Nested 10 fold, 200 pairs. In both graphs, the x axis represents the number of extracted features, m and it ranges from 5 to 40 in increments of 5. A. Graph 1. The y axis represents the probability that at least one is correct, in percentage and it ranges from 0 to 100 in increments of 20. The solid blue curve passes through (5, 78), and (40, 55). The dotted blue curve passes through (5, 82) and (40, 62). The solid yellow, red and purple curves are very close to each other and pass near (5, 92) and (40, 90). The dashed yellow, dashed red, and dashed purple curves almost overlap and pass near (5, 100) and (40, 95). B. Graph 2. The y axis represents the probability that both are correct in percentage and it ranges from 0 to 100 in increments of 20. The solid blue curve passes through (5, 52), (20, 25), and (40, 15). The dotted blue curve passes through (5, 65), (20, 40), and (40, 25). The solid yellow, solid red and solid purple curves are relatively close to each other and pass near (5, 90), (20, 80), and (40, 75). The dashed yellow, dashed red, and dashed purple curves are very close to each other and pass near (5, 99), (20, 97), and (40, 95). All values are estimates.

Effect of dimensionality of the feature space (m) on the statistical confidence of the trained model for nested 10-fold cross-validation: (A) C2,1 and (B) C2,2.

Experiment 3: Statistical Characteristics of Nested k-Fold Cross-Validation

Experiments 1 and 2 showed that nested k-fold cross-validation offers the best statistical power and statistical confidence of the model among the investigated methods. Experiment 3 addressed research questions Q3 and Q4 by quantifying the statistical properties of nested 10-fold cross-validation and providing models that would enable future studies to estimate the required and recommended sample sizes in the design phase of the study. Given the computational cost of running nested k-fold and the number of required analyses for this experiment, each set of parameters was simulated 2,000 times.

Based on Figure 6B, the 95th percentile of H0 for nested k-fold cross-validation was found to be relatively independent of the dimensionality of the feature space (m). We formally tested to see if the number of extracted features (m) and the number of selected features (l) had significant effects on the H0 CI. To that end, a two-way analysis of variance (ANOVA) in a 4 × 3 design was used with independent variables m [10, 20, 30, 40] and l [2, 3, 4]. The dependent variable was the upper bound of H0 CI (one-sided, α = .05). Table 2 presents the result, which did not show statistically significant effects of m and l on the upper bound of H0 CI, confirming the robustness of the upper bound of H0 CI in nested k-fold cross-validation to m and l.

Table 2.

Results of two-way 4 × 3 analysis of variance on the upper bound of the H0 confidence interval for nested 10-fold cross-validation.

Predictor Sum of squares df Mean square F p
Main effect: feature space dimensionality (m ∈ [10, 20, 30, 40]) 0.0017 3 0.0006 0.51 .67
Main effect: number of selected features (l ∈ [2, 3, 4]) 0.0042 2 0.0021 1.87 .15
Interaction effect 0.0093 6 0.0016 1.37 .22

Next, the minimum required number of pairs (nr) for a study using nested 10-fold cross-validation to detect a statistically significant effect (α = .05) with typical power (1 − β = .8) was computed for d ∈ [0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1], m ∈ [10, 20, 30, 40], l ∈ [2, 3, 4]. Figure 10 shows the results, which indicate that nr is inversely related to d and l and directly related to m, with d being the most influential factor. This means the function for estimating nr must integrate all three predictors.

Figure 10.

3 graphs. The legend for the graphs is as follows. Solid blue: m equals 10. Dashed red: m equals 20. Dashed yellow: m equals 30. Dotted purple: m equals 40. In all the 3 graphs, the y axis represents the required pair number and it ranges from 20 to 220 in increments of 20. The x axis represents the effect size in terms of Cohen\u2019s d and it ranges from 0.4 to 1 in increments of 0.1. A. Graph 1. The curves are very close to each other and pass near (0.4, 200), (0.7, 70), and (1, 45). Graph 2. The curves are very close to each other and pass near (0.4, 150), (0.7, 55), and (1, 30). Graph 3. The curves are very close to each other and pass near (0.4, 130), (0.7, 50), and (1, 30). All values are estimates.

Effect of the size of the feature space (m), the number of selected features (l), and effect size (d) on the required number of pairs (nr) for nested 10-fold cross-validation: (A) l = 2, (B) l = 3, and (C) l = 4.

To derive a closed-form equation for estimating nr, for each pair of m and l, a power curve with the general form of axb+c was used to estimate the value of nr from d. This resulted in 12 different curves. Then, the parameters of those curves (a, b, and c) were estimated from m and l based on three separate flat planes (one per parameter). Let d0, m0, and l0 denote the target parameters of a study for which we want to estimate the minimum required number of pairs (α = .05, 1 − β = .8). Equations 47 show the process, as follows:

a=39.376.718l0+0.263m0 (4)
b=1.9850.023l0+0.001m0 (5)
c=0.886+1.507l00.015m0 (6)
nr=ad0b+c (7)

The goodness of fit was evaluated using root-mean-square error (RMSE) and mean percent magnitude error (MPE). MPE was defined as the magnitude of the error divided by the true value expressed in percentage averaged over all testing instances. Using Equations 47 with the training samples (resubstitution error rate) resulted in an RMSE of 3.35 and MPE of 2.98%. That is, the expected magnitude error of using Equations 47 for estimating the number of required pairs (nr) was less than 3% on the training set. A MATLAB function (“Compute_RequiredSampleSize.m”) implementing Equations 47 has been provided in the Supplemental Material S3.

An independent test set was generated for the following sets of parameters: d ∈ [0.55, 0.75, 0.95, 1.2, 1.4], m ∈ [20, 30, 40], and l ∈ [2, 3, 4, 5, 6]. Our preliminary analysis showed the value of MPE over all test samples was 11.6%, which is substantially higher than the training MPE. Therefore, further investigation was carried out to find the source of the difference. Our investigation showed that correlation between MPE and the number of extracted features (i.e., m) was nonsignificant (p = .13). Similarly, correlation between MPE and discriminatory power of the features (i.e., d) was nonsignificant (p = .63). In contrast, MPE was highly correlated with the number of selected features (r = .82, p < .0001). Our further investigation showed that the value of MPE was quite different among the cases where l ∈ [2, 3, 4] and l ∈ [5, 6]. Going back to Experiment 3, Equations 47 were trained based on data points with l ∈ [2, 3, 4]. Therefore, the test samples with l ∈ [2, 3, 4] correspond to interpolation, whereas the test samples with l ∈ [5, 6] correspond to extrapolation cases. The interpolation cases (45 data points) had an MPE of 3.5%, which is quite close to the MPE over the training samples, and also had a nonsignificant correlation coefficient (p = .19). However, the MPE for extrapolation cases (30 data points) was 23.9%, with a high correlation (r = .73, p < .00001). In summary, Equations 47 are very reliable for estimating the required sample size for l ∈ [2, 3, 4], but the results become less reliable as we start to use them for extrapolating into models with higher dimensions (l5).

The statistical confidence of the final model (C2,2, i.e., the probability of both selected features being correct) using nested 10-fold cross-validation was computed for different numbers of pairs and d ∈ [0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1], m ∈ [10, 20, 30, 40], l = 2. The results are presented in Appendix Tables A1 and A2 with selected cases (d ∈ [0.4, 0.6, 0.8, 1], m = 10, l = 2) plotted in Figure 11.

Figure 11.

A graph of the model confidence, C subscript 2, comma, 2 on the y axis and the pair number on the x axis. The y axis ranges from 10 to 100 in increments of 10. The x axis ranges from 50 to 500 in increments of 50. The legend is as follows. Solid blue: d equals 0.4, dashed red: d equals 0.6, dashed yellow: d equals 0.8, and dotted purple: d equals 1. The solid blue curve passes through (50, 19), (250, 75), and (500, 92). The dashed red curve passes through (50, 40), (250, 95), and (500, 100). The dashed yellow curve passes through (50, 60), (150, 95), and (500, 100). The dotted purple curve passes through (50, 75), (100, 95), and (500, 100). All values are estimates.

Effect of d and n on the confidence of the model generated using nested 10-fold cross-validation when m = 10.

Finally, the recommended sample size in terms of the number of pairs (n) for achieving a target confidence for the final model (C2,2) can be obtained from Tables A1 and A2 in conjunction with linear interpolations. MATLAB functions (“Compute_NestedModelConfidence.m” and “Compute_RecommendedSampleSize.m”) implementing this procedure been provided in Supplemental Material S1 and Supplemental Material S2, respectively. For example, if we have two relatively good features with d = 0.6 in our 10-dimensional feature space, and we want to reach the confidence of 95%, we need to have about 215 pairs of samples. However, to achieve the same confidence with 40 extracted features, 342 pairs of samples need to be used.

Discussion

This study was primarily motivated by the observation that many ML-based studies in speech, language, and hearing sciences have adopted cross-validation methods with a single test set (i.e., single holdout or train–validation–test cross-validation approaches). Our initial investigation using real clinical data showed a profound impact of employed cross-validation on the estimated accuracy and its dispersion and the stability of the feature selection outcome. Therefore, our main hypothesis was that the statistical properties of the final trained model would be significantly affected by which cross-validation method was used, especially at smaller sample sizes (less than a few hundred) that are the most likely scenarios in these fields.

Three experiments were conducted to quantify the impact of different cross-validation methods on the statistical properties of the final trained models and to provide methods that would enable future studies to estimate the required and recommended sample sizes during the design phase based on the most robust cross-validation method that was evaluated in the current investigation (nested k-fold cross-validation). Table 3 presents a summary comparing the statistical properties of the four cross-validation approaches studied. We would like to emphasize that the three factors of bias (H0 mean accuracy), statistical power, and statistical confidence need to be considered simultaneously. For example, while 10-fold cross-validation has the highest power, its estimated accuracy (if the ML processing pipeline contains a model selection component) is biased and hence should not be used. Conversely, despite the fact that the train–validation–test method gives an unbiased estimate of the performance, the method has very low statistical power, meaning the distribution of H0 and Ha have a lot of overlap. Therefore, there is a significant probability for random features to produce high classification accuracies with the train–validation–test method (please refer to Figures 5A and 7). It is hoped that the results of this work will provide strong motivation for researchers to move away from single holdout, train–validation–test, and k-fold cross-validation approaches and instead consider using the unbiased and more robust nested k-fold cross-validation method in future studies.

Table 3.

A summary comparing the four cross-validation methods studied.

Method H0 mean accuracy H0 SD accuracy Robustness to large feature space Statistical power Statistical confidence Required sample size Recommended sample size Computational time
Single holdout Highest bias Low Lowest Low Lowest High Highest Lowest
k-fold High bias Lowest High Highest High Lowest Low Low
Train validation test Unbiased Highest Moderate Lowest Moderate Highest Moderate Low
Nested k-fold Unbiased Moderate Highest High Highest Low Lowest Highest

Note. Desirable model characteristics are marked with bold face letters.

Q1: How Is the Statistical Power of an ML Model Affected by the Employed Cross-Validation Method?

Experiment 1 answered this question by comparing the distribution of the null hypothesis H0 (i.e., when the features were irrelevant) and alternative hypothesis Ha (when some of the features were discriminative) for four different cross-validation methods. Results showed that the widely used single holdout method has significant bias (see Figure 5A) and therefore can significantly overestimate the true performance; at the same time, single holdout cross-validation leads to an H0 with a very heavy tail (see Figure 5B). This means that, even if the features are truly random (irrelevant features), there is a good chance for the single holdout (and k-fold) method to produce a high level of accuracy. On the other hand, not only does nested k-fold cross-validation give an unbiased estimate of the performance (see Figure 5A), but also its H0 distribution has a thinner higher tail (see Figure 5B). This means that the probability of getting a very high accuracy from irrelevant features from nested k-fold cross-validation is much lower than the single holdout cross-validation.

Another advantage of nested k-fold cross-validation is higher robustness to the dimensionality of the feature space. This means that the probability of getting high accuracy from irrelevant features increases with an increase in the number of extracted features if the single holdout method is used; in contrast, increasing the number of features does not artificially increase accuracy if nested k-fold cross-validation is used (see Figure 6B). Considering that most ML-based studies are exploratory in nature, they often start with a relatively high number of features, and therefore, using a method that is less affected by the number of extracted features is quite desirable.

The results of Experiment 1 also demonstrated that robust statistical results could be obtained with significantly smaller required sample sizes when using nested k-fold cross-validation as compared to the single holdout or train–validation–test methods (see Figure 7). Considering that data collection is usually the most costly part of a clinical study in terms of time and expense, a significant reduction in the required sample size could increase the likelihood of carrying out a study that may have previously been deemed impractical or too costly. This reduction of the required sample size of course comes at the expense of increased computational complexity of the nested k-fold cross-validation method. However, the availability of fast computers and high-performance computing resources reduces this concern, and the trade-off is quite favorable.

Q2: How Is the Statistical Confidence in an ML Model Affected by the Employed Cross-Validation Method?

Statistical confidence of the model was defined as the probability of correct features being selected and, hence, contributing to the classification outcome. Experiment 2 answered research question Q2 by quantifying the effect of the four cross-validation methods on the statistical confidence of the model. Based on Figure 8, the statistical confidence of the model based on nested k-fold cross-validation could be as much as 4 times higher than the confidence associated with the model based on single holdout cross-validation. The other important finding was the dissimilar behavior of statistical confidence of the model when the number of extracted features increases. Based on Figure 9B, not only does the nested k-fold cross-validation lead to significantly higher confidence in the model, but it is also more robust than the other methods to the dimensionality of the feature space. Therefore, whereas we need to be extra cautious and extract “as few” features as possible when employing single holdout cross-validation, we could relax this requirement and extract more features with nested k-fold cross-validation. In practice, this relaxation could significantly increase the chance of finding more powerful and discriminative features. These important characteristics should provide further motivation for shifting from the single holdout method to nested k-fold cross-validation.

Based on the results of Experiments 1 and 2, the nested k-fold cross-validation method exhibited the best performance, followed by the 10-fold cross-validation method. Train–validation–test cross-validation, which is a popular method in the ML community due to its unbiased estimate of accuracy (see Figure 5A) and the simplicity of its implementation, had the lowest statistical power (see Figure 7) and needed even more samples compared to the single holdout method. However, the statistical confidence of the model generated based on the train–validation–test was much higher than the single holdout method (see Figure 8). This higher confidence is attributed to our implementation of train–validation–test cross-validation. Specifically, the train–validation part was split using 10-fold cross-validation in our implementation, and based on Figure 8, employing 10-fold cross-validation leads to models with high statistical confidence. Then, the inferior statistical confidence of the model in the train–validation–test compared to the 10-fold cross-validation is attributed to the fact that the train–validation–test method only uses a portion of data for feature selection (85% in our implementation). However, if the train–validation part was split using a single holdout method (which in fact is very common in practice), the statistical confidence of the model based on the train–validation–test would have been lower than the statistical confidence of the single holdout method. In summary, the findings of this study advise against any application of single holdout and train–validation–test altogether.

Q3: What Is the Minimum Required Sample Size to Get a Statistically Significant Outcome From an ML Model With Conventional Power Requirements (5% significance, 80% power [α = .05, 1 − β = .8])?

Experiment 3 was devoted to answering research questions Q3 and Q4 for a model employing nested 10-fold cross-validation, which was the method with the best statistical properties. This model was described in Equations 47. The discriminative power of selected features (d) was the most influential factor in determining the minimum required sample size, followed by the dimensionality of the model (l) and then the number of extracted features (m). Additionally, the minimum required sample size in terms of the number of pairs (nr) decreases as the discriminative power of selected features (d) or dimensionality of the model (l) increases. In contrast, the minimum required sample size increases as the number of extracted features (m) increases.

This last observation may seem to contradict the robustness of the distribution of the null hypothesis regarding the number of extracted features (see Figure 6B and Table 2) and warrants further explanation. The minimum required sample size depends on the distributions of both null and alternative hypotheses. The distribution of the null hypothesis of nested k-fold cross-validation is relatively invariant to the number of extracted features (similar value of α for a fixed threshold). However, as more features get extracted (a larger m), the distribution of the alternative hypothesis changes and gets more overlapped with that of the null hypothesis (β increases). Therefore, the required sample size needs to be increased to reduce the value of β and attain similar statistical power.

Lastly, an example is provided to show how Equations 47 can be used to estimate the required sample size during the experimental design. Let us consider the example presented in Comparing Cross-Validation Methods in a Real-World Clinical Example section, where we know we want to extract 48 features prior to the execution of the study. Prior studies have reported the effect sizes in the range of 0.43–0.88 for such a problem (Van Stan et al., 2020). Plugging d0 = 0.66 (as an average expected effect size), m0 = 48, and l0 = 2 into Equations 47, we get nr = 89.3, meaning that we need to recruit a minimum of 90 pairs of participants to have the adequate power to run the analysis.

Q4: What Is the Recommended Sample Size for Achieving a Target Statistical Confidence of the Final ML Model?

The second part of Experiment 3 provided an answer to Q4 for nested 10-fold cross-validation. While we were not able to derive a closed-form formula for this case, different tables and MATLAB code (see Supplemental Material S2) are provided that can be used for linear interpolation and estimation of the recommended sample size for target values of d and m. It is noteworthy that using Tables A1 and A2 for estimating the recommended number of pairs (n) will result in a larger value compared to Equations 47. For example, we only need 89 pairs for d = 0.6, m = 20, and l = 2 to obtain a statistically significant effect (α = .05, 1 − β = .8). However, 89 pairs with the same parameters will provide 50% confidence in the model. To achieve model confidence of 80% and 95%, respectively, the sample size needs to be increased to 169 and 247 pairs. This wide gap is because Equations 47 estimate the required n for just having the statistical power to detect a significant effect. In comparison, having a model with high statistical confidence (a large C2,2) is a more restrictive requirement than getting a statistically significant outcome and, therefore, requires more samples. The number of pairs necessary for reaching a certain confidence in the model increases with a decrease in the discriminative power of features (d) and with an increase in the number of extracted features (m). As a final note, considering the negative effect of the high number of extracted features on the statistical power and statistical confidence of the model, keeping the number of extracted features relatively low is preferable.

ML solutions often have different optimization processes. The focus of this study was to quantify the statistical characteristics of different cross-validation methods in the presence of feature selection. However, other processes including hyperparameter optimization and optimization of the architecture (e.g., in deep learning) are also important and widely used in the ML community. While the specific findings (e.g., the method for estimating the required sample size) of this study are limited to cases with feature selection, the general findings can be applied to other optimization scenarios. For example, single holdout train–validation–test, and k-fold cross-validation methods should not be used for hyperparameter optimization and optimization of the architecture of deep networks. Instead, the nested k-fold cross-validation method should be employed. This argument is especially important for research with small sample sizes (less than a few hundred).

Computing Appropriate Dimensionality of the Feature Space

The presented power analysis method (Equations 47) was derived to assist researchers with determining the minimum sample size required during experimental design; however, the method can also be used for existing data sets. Specifically, during experimental design, researchers know the features that they want to use so the dimensionality of the feature space can be estimated, and their goal is to determine the required sample size. However, the situation reverses with existing data sets. That is, the sample size is already known, and the goal is to determine the appropriate dimensionality of the feature space. The presented method can also be used to estimate the maximum allowable dimensionality of the feature space. If the number of extracted features were higher than this estimated number, some features would need to be eliminated such that the dimensionality of the feature space gets below the maximum allowable dimension. We would like to emphasize that this elimination should be carried out independently from the labels of the data set; otherwise, the elimination procedure will most likely overestimate the performance of the model and will reduce the generalizability of the findings. Some examples of appropriate elimination would be using the literature (generated from different data sets) for removing inferior features, computing correlation between features of the data set and only keeping one among highly correlated features, or applying an unsupervised dimensionality reduction without any optimization based on the labels of the data set.

Lastly, an example is provided to show how Equations 47 can be used to estimate appropriate dimensionality of the feature space for an existing data set. Let's consider the example presented in Comparing Cross-Validation Methods in a Real-World Clinical Example section, where our existing data set has 136 pairs of participants. Plugging d0 = 0.66 (as an average expected effect size), m0 = 135, and l0 = 2 into Equations 47, we get nr = 135.2 ≈ 136, meaning that with 136 pairs of participants and d0 = 0.66, the maximum allowable dimensionality of the feature space would be 135.

Limitations

Despite the importance of our current findings, several limitations should be mentioned. Our findings were based on Monte Carlo simulations, and we had to make certain assumptions to keep the complexity of our simulations manageable. Those assumptions are responsible for most of the limitations of our findings. The first assumption was that ML models were being used for a binary classification problem. Although the findings would still be valid for a multiclass classification problem (e.g., that the single holdout method is the least reliable and nested k-fold cross-validation is the most reliable method), the results of Experiment 3 could not be applied directly to other classification problems and need to be simulated separately. The second assumption was that of a balanced data set (equal number of samples in both classes), which in practice may not be achievable or desirable. The third assumption was that all of the discriminative features each had equal discriminative power represented by the same value of d. The fourth assumption was that all discriminative features had Gaussian distributions, and hence, a linear classifier (logistic regression) was able to find the optimum decision boundary. The fifth assumption was that none of the extracted features were correlated, and the remaining (m − l) features were irrelevant to the classification problem. While it is desirable to extract only features that add novel information to the model and are uncorrelated, in practice, this is hardly ever achieved. Finally, it is important to note that the estimation of the minimum sample size of conventional statistical tests (e.g., t test, ANOVA, correlation, etc.) shares many of these assumptions and limitations, and in that regard, the methodology presented in Experiment 3 extends the application of widely used power analysis to include ML models.

Generalizability of Results to Unbalanced Data Sets and Other Classification Frameworks

As described in the Limitations section, the equations for the estimation of the required sample size were derived based on simulated balanced data sets with underlying Gaussian distributions and with equal discriminative power for selected features. This section includes modifications that would allow applications of power analysis to unbalanced data sets and features with unequal discriminative powers. We will also present preliminary results on a direction that could extend the application of the method to nonlinear feature spaces.

First, unbalanced data sets were investigated. Let γdb denote the ratio between the number of samples in the positive (npositive) and negative classes (nnegative). The following sets of parameters were used: m = 20, l = 2, d ∈ [0.6, 0.8, 1,] and γdb ∈ [1.2, 1.4, 1.6, 1.8, 2]. Two different scenarios were tested. In the first scenario, Equations 47 were used for estimating the required sample size of the smaller class. The error was always negative, with MPE equal to 44.2%. This indicates that using Equations 47 for estimating the required sample size of the smaller class leads to a significant overestimation. In the second scenario, Equations 47 were used for estimating an adjusted required sample size, where the adjusted required sample size was defined as the average of the sample sizes of the positive and negative classes (nnegative+npositive2). The error was reduced, with MPE equal to 12.4%. Further investigation indicated that, as the value of d increases, estimation of the adjusted required sample size becomes more robust to higher imbalance ratios (γdb). For example, if d = 0.6 and γdb=2, MPE for the adjusted required sample size was 70.6%; however, for d = 0.8 and γdb=2, MPE was reduced to 16%. Therefore, Equations 47 are relatively robust to moderately imbalanced data sets when employing the adjusted required sample size and if features were highly discriminative (d ≥ 0.8). The figures for this analysis are provided in Figure A1 in the Appendix. In summary, the “Compute_RequiredSampleSize.m” code provided along with this article can be used for unbalanced data sets; however, the estimated required sample size would be the average of the number of negative and positive samples. Estimation of the sample size of each class from the estimated average required sample size and γdb should be trivial.

Second, a limited number of cases with features with unequal effect sizes (d) were simulated. To that end, a series of 20-dimensional feature spaces (m=20) containing two discriminative features (l=2) were generated. The discriminatory power of the first feature was taken from the set d0.60.81, and the discriminatory power of the second feature was set equal to γd×d. The parameter γd captures the unequal discriminatory power between the two features, and γd∈[1.2,1.4,1.6,1.8,2]. The average discriminatory power of the two features (d0=dγd+12) was used in Equation 7. The value of MPE was 7.9%. Our further investigation indicated that, as the value of d increases, estimation of the required sample size becomes more robust to higher imbalance ratios (γd). For example, if d = 0.6 and γd=2, MPE was 23.1%. However, for d = 0.8 and γd=2, MPE was reduced to 16.7%. Therefore, Equations 47 are relatively robust to features with different discriminatory powers if both features were discriminative enough (d ≥ 0.8). The figures for this analysis are provided in Figure A2 in the Appendix. In summary, the “Compute_RequiredSampleSize.m” code provided along with this article can be used for features with different unequal discriminative powers; however, the average of their discriminative powers needs to be used.

Finally, the possibility of applying the method to nonlinear feature spaces and other classification algorithms was investigated. The findings of this study were based on simulated data sets with underlying Gaussian distributions. In that sense, quantification of their discriminative power using Cohen's d and the application of logistic regression were optimum selections. We would like to emphasize another rationale for parameterization of power analysis in terms of Cohen's d, and then we will present a possible direction for the generalization of power analysis to features with non-Gaussian distributions and nonlinear classifiers. Researchers often use effect sizes reported in the literature for estimating the minimum required sample size during the design of a study, and currently, the application of Cohen's d is ubiquitous in clinical sciences (Fritz et al., 2012). Therefore, the selection of Cohen's d was driven by this important practical consideration.

We will now present a possible modification that could extend the application of the presented power analysis to more general cases. However, an in-depth treatment of this topic is out of the scope of this article and needs to be pursued in an independent study. Our proposed solution is based on computing an “equivalent Cohen's d” for such general cases. The “equivalent Cohen's d” for a nonlinear feature space is the value of d that, if used to generate Gaussian features (please refer to Equations 1 and 2), produces similar classification accuracy with logistic regression as the problem under investigation. We will demonstrate this approach with bull's eye data sets with five different noise levels of [2, 2.5, 3, 3.5, 4]. Different noise levels model different amounts of overlap between the two classes. Specifically, for a noise level of 0, the data set will be two perfect co-centric circles with a between-class distance of 1 unit. The noise level of 2 corresponds to a uniform disturbance in the range of −1 to 1 (i.e., equal to the true between-class distance). The noise level of 4 corresponds to a uniform disturbance with a range that is 2 times larger than the true between-class distance. Figure A3 in the Appendix shows examples of this data set with low and high noise levels. An SVM with a polynomial kernel was adopted for learning the decision boundary between the two classes. For each noise level, the two dimensions of the bull's eye data set were augmented with 18 random features to create a 20-dimensional feature space ( m=20 ). Nested 10-fold cross-validation with polynomial SVM was used to find the two most discriminative features. Following the methodology presented in the Statistical Power and the Required Sample Size section (please refer to Figure 4), the true required sample size for rejecting the null hypothesis was determined for each noise level. The equivalent Cohen's d of each noise level was estimated from a look-up table that we generated from simulations of Gaussian features with a logistic regression classifier. The computed equivalent Cohen's d was then used to estimate the required sample size from Equations 47. The results showed an MPE of 8.6%, confirming the potential of this approach.

Directions for Future Studies

The study presented here can be extended in several directions. The first direction would be a rigorous and in-depth treatment of extending the power analysis to nonlinear feature spaces. Updating the power analysis method to better account for unbalanced data sets and unequal discriminative power for selected features are other directions. The current study showed that statistical properties of nested k-fold cross-validation are very desirable. However, only the value of k = 10 was implemented. While the general findings of the study should hold for other values of k, we expect an interaction effect among the power of the method, the sample size, and the value of k. A future study may suggest the appropriate value of k for a given problem such that the statistical power of the model is maximized. The current study used Cohen's d for Gaussian features and proposed equivalent Cohen's d for general feature spaces. However, investigation of the most appropriate way to quantify and report the effect size in ML would allow practitioners and researchers to use prior studies during experimental design and is a question for future research. A more theoretical treatment of power analysis and all of the aforementioned directions could significantly help scientists and practitioners with the appropriate and responsible usage of ML. Finally, evaluating the effect of different cross-validation methods on the performance of ML using Bayes Error is another direction for future studies.

Conclusions

Monte Carlo simulations were adopted in this study to quantify the statistical characteristics of different cross-validation methods. Our analyses showed significantly different statistical characteristics among different cross-validation methods. The model based on single holdout had very low statistical power (meaning it required many more samples to get a statistically significant outcome), very low statistical confidence (meaning it was very likely for incorrect features to be selected), and very high bias (meaning it was very likely to overestimate the performance of the model). The statistical characteristics of the popular method of train–validation–test were also very poor. Conversely, the model based on nested 10-fold cross-validation resulted in the highest statistical power and statistical confidence. Also, the estimated performance was unbiased. Our numerical analyses showed that the required sample size with a single holdout could be 50% higher than what would be needed if nested k-fold cross-validation were used. Also, statistical confidence in the model based on nested k-fold cross-validation was as much as four times higher than the statistical confidence in the single holdout-based model. Finally, computational models along with MATLAB code were provided to assist in power analysis for designing future ML studies to determine the required and recommended sample sizes when employing nested 10-fold cross-validation. Based on these findings, researchers are highly encouraged to avoid using the single holdout and train–validation–test methods and use the unbiased and more robust method of nested k-fold cross-validation. As a final note, if the ML processing pipeline does not include any model selection component (e.g., feature selection, hyperparameter optimization, optimization of the architecture of a deep or shallow neural network, selecting between different classification algorithms, etc.), k-fold cross-validation (no nesting) may be used due to its lower computational complexity relative to nested k-fold cross-validation, and should be avoided otherwise.

Data Availability Statement

Mass General Brigham and Mass General are not allowed to give access to data without the principal investigator for the human studies protocol first submitting a protocol amendment to request permission to share the data with a specific collaborator on a case-by-case basis. This policy is based on strict rules dealing with the protection of patient data and information. Anyone wishing to request access to the data used for Comparing Cross-Validation Methods in a Real-World clinical example section must contact Sarah DeRosa (sederosa@partners.org), Program Coordinator for Research and Clinical Speech-Language Pathology, Center for Laryngeal Surgery and Voice Rehabilitation, Massachusetts General Hospital.

The data used for Experiments 1–3 were generated using Monte Carlo simulations. Several MATLAB functions associated with this article are available; see Supplemental Materials S1S4. The latest version of the code can be retrieved from https://github.com/GhasemzadehHamzeh/ML_PowerAnalysis.

The “Compute_RequiredSampleSize.m” code takes the number of extracted features (m_0), the effect size of the best features (d_0), and the number of selected features (l_0) as the input parameters and computes the required number of pairs for having the statistical power for finding a statistically significant model (α = .05, 1 − β = .8). This code can also be used for unbalanced data sets; however, the estimated required sample size (nr) would be the average of the number of negative and positive samples. The sample size of each class can then be computed using the estimated average required sample size and γdb. Specifically, the required sample size of the smaller class will be nr×21+γdb and the required sample size of the larger class will be nr×2×γdb1+γdb. Similarly, this code can also be used for features with different effect sizes; the average of the effect sizes needs to be input as d_0. The “Compute_RecommendedSampleSize.m” code takes the number of extracted features (m_0), the effect size of the best features (d_0), and the target value of C2,2 (CI_0) as the input parameters and computes the recommended number of pairs for achieving the target value of C2,2. The “Compute_NestedModelConfidence.m” code takes the number of extracted features (m_0), the effect size of the best features (d_0), and the number of pairs in a data set (n_0) as the input parameters and computes the value of C2,2. The “Feature_Selection” toolbox provides an implementation of forward feature selection with four investigated cross-validation approaches. Please refer to the “Main_Test.m” file for examples of the usage of the code.

Supplementary Material

Supplemental Material S1. Compute_NestedModelConfidence.m.
JSLHR-67-753-s001.m (4.4KB, m)
Supplemental Material S2. Compute_RecommendedSampleSize.m.
JSLHR-67-753-s002.m (4.6KB, m)
Supplemental Material S3. Compute_RequiredSampleSize.m.
Supplemental Material S4. Feature_Selection.zip.
JSLHR-67-753-s004.zip (84.5KB, zip)

Acknowledgments

Research reported in this publication was supported by the National Institute on Deafness and Other Communication Disorders Grants T32 DC013017 (awarded to Christopher Moore and Cara Stepp), P50 DC015446 (awarded to Robert Hillman), and K99 DC021235 (awarded to Hamzeh Ghasemzadeh). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. The authors would like to acknowledge the Massachusetts General Hospital high-performance computing resource that made the execution of our intensive simulations possible.The authors would like to thank Jarrad Van Stan, Laura Toles, Katie Marks, AJ Ortiz, Rob Petit, and Dave Viggiano for study participant recruitment, data processing, and mobile application development of the ambulatory data that were used in the Comparing Cross-Validation Methods in a Real-World Clinical Example section.

Appendix

Table A1.

Effect of different parameters on the probability of both selected features to be correct (C2,2) when m = 10 or 20 features were extracted.

Discriminatory power (d), m = 10
Discriminatory power (d), m = 20
0.4 0.5 0.6 0.7 0.8 0.9 1 0.4 0.5 0.6 0.7 0.8 0.9 1
Number of pairs (n) 50 17.7 27.9 40.3 50.2 60.8 68.6 75.1 9.5 17.3 27 37.4 48.7 59.4 65.5
100 38.2 51.7 66.9 78.3 85.6 90.9 94.2 23.8 40.1 55.7 68 79 85.8 90.6
150 52.7 69.3 81.4 90.3 94.7 97.2 98.4 39.5 59.5 75 86.3 92.6 96.5 98
200 63.4 79.7 90.1 95.7 98.7 99.5 99.8 51.3 71.5 85.8 93.5 96 98.6 99.4
250 72.9 88.3 95.9 98.5 99.6 99.9 100 63.3 83.2 92.5 96.8 99.2 99.7 99.8
300 79.6 90.5 96.6 99.1 99.6 99.7 99.9 73.4 88.4 96.9 99.1 99.7 100 100
350 84.7 94.5 98.6 99.7 99.9 100 100 79 92 97.5 99.4 99.8 100 99.9
400 88.1 96.1 99.1 99.8 100 100 100 84.1 94.9 99 99.8 100 100 100
450 90.3 97.3 99.6 100 100 100 100 88.1 96.8 99.2 99.9 100 100 100
500 92.6 98.6 99.9 100 100 100 100 90.3 97.6 99.7 99.9 100 100 100

Table A2.

Effect of different parameters on the probability of both selected features to be correct (C2,2) when m = 30 or 40 features were extracted.

Discriminatory power (d), m = 30
Discriminatory power (d), m = 40
0.4 0.5 0.6 0.7 0.8 0.9 1 0.4 0.5 0.6 0.7 0.8 0.9 1
Number of pairs (n) 50 6.3 11.9 19.7 31.5 40.9 50.7 59.7 4.8 10.3 16.6 26.5 38.2 48.2 57.5
100 19.3 35.3 52.3 67.5 77.6 85.4 90.1 15.1 31.7 46.3 60.9 72.8 81.2 87.8
150 32.6 53.6 70.6 83.7 90.5 94.7 97.7 29.2 50.3 67.5 81 89.8 94.4 97.6
200 48.4 69.8 84.5 92.3 96.6 98.8 99.4 41.8 66.8 82.4 91.4 95.4 98.4 99.4
250 56.8 77.5 90.6 96.3 98.7 99.5 99.9 53.3 74.1 89.3 95.4 98.3 99.5 100
300 66.1 84 94.2 97.9 99.3 99.8 100 63.1 81.6 93 98 99.4 99.8 100
350 75.8 89.8 96.5 99.4 99.9 100 100 70.8 89.1 95.4 98.8 99.6 99.8 100
400 81.2 94.1 98.7 99.8 100 100 100 76.1 91.2 97.8 99.4 99.9 99.9 100
450 84.8 95.7 98.9 99.7 100 100 100 82.9 94.9 98.8 99.7 100 100 100
500 86.9 96.5 99.5 100 100 100 100 86.8 97.3 99.6 99.9 100 100 100

Figure A1.

2 graphs. In both graphs, the y axis represents the percent error and it ranges from negative 160 to 0 in increments of 20. The x axis represents the imbalance ratio, gamma subscript d b in increments of 1.2 to 2 in increments of 0.2. The legend is as follows. Solid blue: d equals 0.6. Dashed red: d equals 0.8. Dashed yellow: d equals 1. Graph 1. The solid blue line passes through (1.2, negative 10), (1.6, negative 55), and (2, negative 158). The dashed red line passes through (1.2, negative 10), and (2, negative 78). The dashed yellow line passes through (1.2, negative 5), and (2, negative 50). Graph 2. The solid blue line passes through (1.2, 5), (1.6, 2), and (2, 0). The dashed red line passes through (1.2, 2), (1.6, negative 2), and (2, negative 18). The solid blue line passes through (1.2, 0), (1.6, negative 30), and (2, negative 70). All values are estimates.

The effect of an unbalanced data set (e.g., unequal number of patients and controls) on the estimation error of the minimum required sample size when Equations 47 were used to estimate (A) the size of the smaller class and (B) an adjusted sample size, where the adjusted sample size was defined as the average of the sample sizes of the positive and negative classes. Please refer to the generalization of the results section for further details.

Figure A2.

2 graphs. In both graphs, the y axis represents the percent error and the x axis represents the discriminatory powers ratio, gamma subscript d. The y axis ranges from negative 180 to 0 in increments of 20. The x axis ranges from 1.2 to 2 in increments of 0.2. The legend is as follows. Solid blue line: d equals 0.6. Dashed red line: d equals 0.8. Dashed yellow line: d equals 1. Graph 1. The solid blue line passes through (1.2, negative 20), (1.6, negative 80), and (2, negative 170). The dashed red line passes through (1.2, negative 20), (1.6, negative 80), (1.8, negative 100), and (2, negative 155). The dashed yellow line passes through (1.2, negative 18), (1.6, negative 60), and (2, negative 140). Graph 2. The solid blue line passes through (1.2, negative 2) and (2, negative 22). The dashed red line passes through (1.2, negative 2), (1.6, negative 5), (1.8, negative 5), and (2, negative 20). The dashed yellow line passes through (1.2, 0), (1.6, negative 2), and (2, negative 18). All values are estimates.

The effect of having two features with unequal discriminative power on the estimation error of the minimum required sample size when Equations 47 were used with (A) the smaller effect size (d0=d) and (B) the average discriminatory power of the two features (d0=dγd+12). Please refer to the generalization of the results section for further details.

Figure A3.

2 scatterplots. In both the scatterplots, the x axis represents feature 1 and the y axis represents feature 2. Both axes range from negative 3 to 3 in unit increments. The legend is as follows. Asterisk: Class 1. Red dot: Class 2. Plot 1. Asterisks are marked between x values of negative 3 and 3 and y values of negative 3 and 3. Red dots are marked between x values of negative 2 and 2 and y values of negative 2 and 2. Plot 2. Asterisks are marked between x values of negative 4 and 4 and y values of negative 4 and 4. Red dots are marked between x values of negative 3 and 3 and y values of negative 3 and 3.

Bull's eye nonlinear feature space: (A) with low level of noise (noise level = 2) and (B) with high level of noise (noise level = 4). Please refer to the generalization of the results section for further details.

Funding Statement

Research reported in this publication was supported by the National Institute on Deafness and Other Communication Disorders Grants T32 DC013017 (awarded to Christopher Moore and Cara Stepp), P50 DC015446 (awarded to Robert Hillman), and K99 DC021235 (awarded to Hamzeh Ghasemzadeh). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.

References

  1. Alharbi, S., Hasan, M., Simons, A. J. H., Brumfitt, S., & Green, P. (2020). Sequence labeling to detect stuttering events in read speech. Computer Speech & Language, 62. 10.1016/j.csl.2019.101052 [DOI] [Google Scholar]
  2. Arjmandi, M. K., & Pooyan, M. (2012). An optimum algorithm in pathological voice quality assessment using wavelet-packet-based features, linear discriminant analysis and support vector machine. Biomedical Signal Processing and Control, 7(1), 3–19. 10.1016/j.bspc.2011.03.010 [DOI] [Google Scholar]
  3. Arjmandi, M. K., Pooyan, M., Mikaili, M., Vali, M., & Moqarehzadeh, A. (2011). Identification of voice disorders using long-time features and support vector machine with different feature reduction methods. Journal of Voice, 25(6), e275–e289. 10.1016/j.jvoice.2010.08.003 [DOI] [PubMed] [Google Scholar]
  4. Armstrong, R., Symons, M., Scott, J. G., Arnott, W. L., Copland, D. A., McMahon, K. L., & Whitehouse, A. J. O. (2018). Predicting language difficulties in middle childhood from early developmental milestones: A comparison of traditional regression and machine learning techniques. Journal of Speech, Language, and Hearing Research, 61(8), 1926–1944. 10.1044/2018_JSLHR-L-17-0210 [DOI] [PubMed] [Google Scholar]
  5. Balki, I., Amirabadi, A., Levman, J., Martel, A. L., Emersic, Z., Meden, B., Garcia-Pedrero, A., Ramirez, S. C., Kong, D., Moody, A. R., & Tyrrell, P. N. (2019). Sample-size determination methodologies for machine learning in medical imaging research: A systematic review. Canadian Association of Radiologists Journal, 70(4), 344–353. 10.1016/j.carj.2019.06.002 [DOI] [PubMed] [Google Scholar]
  6. Bayerl, S. P., Wagner, D., Nöth, E., & Riedhammer, K. (2022). Detecting dysfluencies in stuttering therapy using wav2vec 2.0. Interspeech, 2868–2872. https://doi.org/10.21437/Interspeech.2022-10908 [Google Scholar]
  7. Berisha, V., Krantsevich, C., Hahn, P. R., Hahn, S., Dasarathy, G., Turaga, P., & Liss, J. (2021). Digital medicine and the curse of dimensionality. NPJ Digital Medicine, 4(1), Article 153. 10.1038/s41746-021-00521-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Bhat, G. S., Shankar, N., & Panahi, I. M. S. (2020). Automated machine learning based speech classification for hearing aid applications and its real-time implementation on smartphone. 2020 42nd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), 956–959. [DOI] [PMC free article] [PubMed]
  9. Bing, D., Ying, J., Miao, J., Lan, L., Wang, D., Zhao, L., Yin, Z., Yu, L., Guan, J., & Wang, Q. (2018). Predicting the hearing outcome in sudden sensorineural hearing loss via machine learning models. Clinical Otolaryngology, 43(3), 868–874. 10.1111/coa.13068 [DOI] [PubMed] [Google Scholar]
  10. Cho, W. K., & Choi, S.-H. (2020). Comparison of convolutional neural network models for determination of vocal fold normality in laryngoscopic images. Journal of Voice, 36(5), 590–598. https://doi.org/10.1016/j.jvoice.2020.08.003 [DOI] [PubMed] [Google Scholar]
  11. Cohen, J. (2013). Statistical power analysis for the behavioral sciences. Academic Press. 10.4324/9780203771587 [DOI] [Google Scholar]
  12. Cortés, J. P., Espinoza, V. M., Ghassemi, M., Mehta, D. D., Van Stan, J. H., Hillman, R. E., Guttag, J. V, & Zanartu, M. (2018). Ambulatory assessment of phonotraumatic vocal hyperfunction using glottal airflow measures estimated from neck-surface acceleration. PLOS ONE, 13(12). 10.1371/journal.pone.0209017 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Crippa, A., Salvatore, C., Perego, P., Forti, S., Nobile, M., Molteni, M., & Castiglioni, I. (2015). Use of machine learning to identify children with autism and their motor abnormalities. Journal of Autism and Developmental Disorders, 45(7), 2146–2156. 10.1007/s10803-015-2379-8 [DOI] [PubMed] [Google Scholar]
  14. Donohue, C., Khalifa, Y., Mao, S., Perera, S., Sejdić, E., & Coyle, J. L. (2021). Characterizing swallows from people with neurodegenerative diseases using high-resolution cervical auscultation signals and temporal and spatial swallow kinematic measurements. Journal of Speech, Language, and Hearing Research, 64(9), 3416–3431. 10.1044/2021_JSLHR-21-00134 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Field, A., Miles, J., & Field, Z. (2012). Discovering statistics using R. Sage. [Google Scholar]
  16. Figueroa, R. L., Zeng-Treitler, Q., Kandula, S., & Ngo, L. H. (2012). Predicting sample size required for classification performance. BMC Medical Informatics and Decision Making, 12(1), 1–10. 10.1186/1472-6947-12-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Fritz, C. O., Morris, P. E., & Richler, J. J. (2012). Effect size estimates: Current use, calculations, and interpretation. Journal of Experimental Psychology: General, 141(1), 2–18. 10.1037/a0024338 [DOI] [PubMed] [Google Scholar]
  18. Ghasemzadeh, H. (2019a). Calibrated steganalysis of mp3stego in multi-encoder scenario. Information Sciences, 480, 438–453. 10.1016/j.ins.2018.12.035 [DOI] [Google Scholar]
  19. Ghasemzadeh, H. (2019b). Multi-layer architecture for efficient steganalysis of UnderMp3Cover in multiencoder scenario. IEEE Transactions on Information Forensics and Security, 14(1), 186–195. 10.1109/TIFS.2018.2847678 [DOI] [Google Scholar]
  20. Ghasemzadeh, H., & Arjmandi, M. K. (2020). Toward optimum quantification of pathology-induced noises: An investigation of information missed by human auditory system. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28, 519–528. 10.1109/TASLP.2019.2959222 [DOI] [Google Scholar]
  21. Ghasemzadeh, H., Deliyski, D., Ford, D., Kobler, J. B., Hillman, R. E., & Mehta, D. D. (2020). Method for vertical calibration of laser-projection transnasal fiberoptic high-speed videoendoscopy. Journal of Voice, 34(6), 847–861. 10.1016/j.jvoice.2019.04.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Ghasemzadeh, H., Deliyski, D. D., Hillman, R. E., & Mehta, D. D. (2021). Method for horizontal calibration of laser-projection transnasal fiberoptic high-speed videoendoscopy. Applied Sciences, 11(2), Article 822. 10.3390/app11020822 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Ghasemzadeh, H., Doyle, P. C., & Searl, J. (2022). Image representation of the acoustic signal: An effective tool for modeling spectral and temporal dynamics of connected speech. The Journal of the Acoustical Society of America, 152(1), 580–590. 10.1121/10.0012734 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Ghasemzadeh, H., & Searl, J. (2018, August 1–3). Modeling dynamics of connected speech in time and frequency domains with application to ALS [Paper presentation]. 11th International Conference on Voice Physiology and Biomechanics (ICVPB), East Lansing, MI, United States. [Google Scholar]
  25. Ghasemzadeh, H., Tajik Khass, M., Khalil Arjmandi, M., & Pooyan, M. (2015). Detection of vocal disorders based on phase space parameters and Lyapunov spectrum. Biomedical Signal Processing and Control, 22, 135–145. 10.1016/j.bspc.2015.07.002 [DOI] [Google Scholar]
  26. Ghassemi, M., Van Stan, J. H., Mehta, D. D., Zañartu, M., Cheyne, H. A., II, Hillman, R. E., & Guttag, J. V. (2014). Learning to detect vocal hyperfunction from ambulatory neck-surface acceleration features: Initial results for vocal fold nodules. IEEE Transactions on Biomedical Engineering, 61(6), 1668–1675. 10.1109/TBME.2013.2297372 [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Gómez, P., Schützenberger, A., Semmler, M., & Döllinger, M. (2018). Laryngeal pressure estimation with a recurrent neural network. IEEE Journal of Translational Engineering in Health and Medicine, 7, 1–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Hamed Mozaffari, M., & Lee, W.-S. (2019). Domain adaptation for ultrasound tongue contour extraction using transfer learning: A deep learning approach. The Journal of the Acoustical Society of America, 146(5), EL431–EL437. 10.1121/1.5133665 [DOI] [PubMed] [Google Scholar]
  29. Hawkins, D. M. (2004). The problem of overfitting. Journal of Chemical Information and Computer Sciences, 44(1), 1–12. 10.1021/ci0342472 [DOI] [PubMed] [Google Scholar]
  30. Huang, G. B., Mattar, M., Berg, T., & Learned-Miller, E. (2008). Labeled faces in the wild: A database for studying face recognition in unconstrained environments. Workshop on Faces in “Real-Life” Images: Detection, Alignment, and Recognition.
  31. Ibarra, E. J., Parra, J. A., Alzamendi, G. A., Cortés, J. P., Espinoza, V. M., Mehta, D. D., Hillman, R. E., & Zañartu, M. (2021). Estimation of subglottal pressure, vocal fold collision pressure, and intrinsic laryngeal muscle activation from neck-surface vibration using a neural network framework and a voice production model. Frontiers in Physiology, 12, Article 732244. 10.3389/fphys.2021.732244 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Jones, S., Carley, S., & Harrison, M. (2003). An introduction to power and sample size estimation. Emergency Medicine Journal, 20(5), 453–458. 10.1136/emj.20.5.453 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Kapoor, S., & Narayanan, A. (2022). Leakage and the reproducibility crisis in ML-based science. ArXiv. 10.48550/arXiv.2207.07048 [DOI] [PMC free article] [PubMed]
  34. Kist, A. M., Gómez, P., Dubrovskiy, D., Schlegel, P., Kunduk, M., Echternach, M., Patel, R., Semmler, M., Bohr, C., Dürr, S., Schutzenberger, A., & Dollinger, M. (2021). A deep learning enhanced novel software tool for laryngeal dynamics analysis. Journal of Speech, Language, and Hearing Research, 64(6), 1889–1903. 10.1044/2021_JSLHR-20-00498 [DOI] [PubMed] [Google Scholar]
  35. Krstajic, D., Buturovic, L. J., Leahy, D. E., & Thomas, S. (2014). Cross-validation pitfalls when selecting and assessing regression and classification models. Journal of Cheminformatics, 6(1), 1–15. 10.1186/1758-2946-6-10 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Le, D., Licata, K., Persad, C., & Provost, E. M. (2016). Automatic assessment of speech intelligibility for individuals with aphasia. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 24(11), 2187–2199. 10.1109/TASLP.2016.2598428 [DOI] [Google Scholar]
  37. Lenatti, M., Moreno-Sánchez, P. A., Polo, E. M., Mollura, M., Barbieri, R., & Paglialonga, A. (2022). Evaluation of machine learning algorithms and explainability techniques to detect hearing loss from a speech-in-noise screening test. American Journal of Audiology, 31(3S), 961–979. 10.1044/2022_AJA-21-00194 [DOI] [PubMed] [Google Scholar]
  38. Liu, X., Faes, L., Kale, A. U., Wagner, S. K., Fu, D. J., Bruynseels, A., Mahendiran, T., Moraes, G., Shamdas, M., Kern, C., Ledsam, J. R., Schmid, M. K., Balaskas, K, Topol, E. J., Bachmann, L. M.Keane, P. A., & Denniston, A. K. (2019). A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: A systematic review and meta-analysis. The Lancet Digital Health, 1(6), e271–e297. 10.1016/S2589-7500(19)30123-2 [DOI] [PubMed] [Google Scholar]
  39. Lowry, R. (2014). Concepts and applications of inferential statistics. Retrieved February 1, 2024, from http://vassarstats.net/textbook/
  40. Marcot, B. G., & Hanea, A. M. (2021). What is an optimal value of k in k-fold cross-validation in discrete Bayesian network analysis? Computational Statistics, 36(3), 2009–2031. 10.1007/s00180-020-00999-9 [DOI] [Google Scholar]
  41. Matikolaie, F. S., & Tadj, C. (2022). Machine learning-based cry diagnostic system for identifying septic newborns. Journal of Voice. 10.1016/j.jvoice.2021.12.021 [DOI] [PubMed] [Google Scholar]
  42. Mehta, D. D., Van Stan, J. H., Zañartu, M., Ghassemi, M., Guttag, J. V, Espinoza, V. M., Cortés, J. P., Cheyne, H. A., & Hillman, R. E. (2015). Using ambulatory voice monitoring to investigate common voice disorders: Research update. Frontiers in Bioengineering and Biotechnology, 3, Article 155. 10.3389/fbioe.2015.00155 [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Menard, S. (2002). Applied logistic regression analysis (Issue 106). Sage. 10.4135/9781412983433 [DOI] [Google Scholar]
  44. Mielens, J. D., Hoffman, M. R., Ciucci, M. R., McCulloch, T. M., & Jiang, J. J. (2012). Application of classification models to pharyngeal high-resolution manometry. Journal of Speech, Language, and Hearing Research, 55(3), 892–902. 10.1044/1092-4388(2011/11-0088) [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Oleson, J. J., Brown, G. D., & McCreery, R. (2019). The evolution of statistical methods in speech, language, and hearing sciences. Journal of Speech, Language, and Hearing Research, 62(3), 498–506. 10.1044/2018_JSLHR-H-ASTM-18-0378 [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Ossewaarde, R., Jonkers, R., Jalvingh, F., & Bastiaanse, R. (2020). Quantifying the uncertainty of parameters measured in spontaneous speech of speakers with dementia. Journal of Speech, Language, and Hearing Research, 63(7), 2255–2270. 10.1044/2020_JSLHR-19-00222 [DOI] [PubMed] [Google Scholar]
  47. Panayotov, V., Chen, G., Povey, D., & Khudanpur, S. (2015). LibriSpeech: An ASR corpus based on public domain audio books. 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5206–5210. [Google Scholar]
  48. Parvandeh, S., Yeh, H.-W., Paulus, M. P., & McKinney, B. A. (2020). Consensus features nested cross-validation. Bioinformatics, 36(10), 3093–3098. 10.1093/bioinformatics/btaa046 [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Qayyum, A., Qadir, J., Bilal, M., & Al-Fuqaha, A. (2020). Secure and robust machine learning for healthcare: A survey. IEEE Reviews in Biomedical Engineering, 14, 156–180. [DOI] [PubMed] [Google Scholar]
  50. Sagi, O., & Rokach, L. (2018). Ensemble learning: A survey. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 8(4). [Google Scholar]
  51. Shen, L., Kann, B. H., Taylor, R. A., & Shung, D. L. (2021). The clinician's guide to the machine learning galaxy. Frontiers in Physiology, 12, Article 658583. 10.3389/fphys.2021.658583 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Thabtah, F., & Peebles, D. (2020). A new machine learning model based on induction of rules for autism detection. Health Informatics Journal, 26(1), 264–286. 10.1177/1460458218824711 [DOI] [PubMed] [Google Scholar]
  53. Theodoridis, S., & Koutroumbas, K. (2009). Pattern recognition (4th ed.). Elsevier. [Google Scholar]
  54. Tsanas, A., Little, M. A., McSharry, P. E., Spielman, J., & Ramig, L. O. (2012). Novel speech signal processing algorithms for high-accuracy classification of Parkinson's disease. IEEE Transactions on Biomedical Engineering, 59(5), 1264–1271. 10.1109/TBME.2012.2183367 [DOI] [PubMed] [Google Scholar]
  55. Uhm, T., Lee, J. E., Yi, S., Choi, S. W., Oh, S. J., Kong, S. K., Lee, I. W., & Lee, H. M. (2021). Predicting hearing recovery following treatment of idiopathic sudden sensorineural hearing loss with machine learning models. American Journal of Otolaryngology, 42(2). 10.1016/j.amjoto.2020.102858 [DOI] [PubMed] [Google Scholar]
  56. Vabalas, A., Gowen, E., Poliakoff, E., & Casson, A. J. (2019). Machine learning algorithm validation with a limited sample size. PLOS ONE, 14(11). 10.1371/journal.pone.0224365 [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Van Stan, J. H., Burns, J., Hron, T., Zeitels, S., Panuganti, B. A., Purnell, P. R., Mehta, D. D., Hillman, R. E., & Ghasemzadeh, H. (2023). Detecting mild phonotrauma in daily life. The Laryngoscope, 133(11), 3094–3099. 10.1002/lary.30750 [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Van Stan, J. H., Mehta, D. D., Ortiz, A. J., Burns, J. A., Toles, L. E., Marks, K. L., Vangel, M., Hron, T., Zeitels, S., & Hillman, R. E. (2020). Differences in weeklong ambulatory vocal behavior between female patients with phonotraumatic lesions and matched controls. Journal of Speech, Language, and Hearing Research, 63(2), 372–384. 10.1044/2019_JSLHR-19-00065 [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Varoquaux, G., Raamana, P. R., Engemann, D. A., Hoyos-Idrobo, A., Schwartz, Y., & Thirion, B. (2017). Assessing and tuning brain decoders: Cross-validation, caveats, and guidelines. NeuroImage, 145(Pt. B), 166–179. 10.1016/j.neuroimage.2016.10.038 [DOI] [PubMed] [Google Scholar]
  60. Vieira, F. G., Venugopalan, S., Premasiri, A. S., McNally, M., Jansen, A., McCloskey, K., Brenner, M. P., & Perrin, S. (2022). A machine-learning based objective measure for ALS disease severity. NPJ Digital Medicine, 5(1), 1–9. 10.1038/s41746-022-00588-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Viering, T., & Loog, M. (2022). The shape of learning curves: A review. IEEE Transactions on Pattern Analysis and Machine Intelligence. [DOI] [PubMed] [Google Scholar]
  62. Wang, J., Green, J. R., Samal, A., & Yunusova, Y. (2013). Articulatory distinctiveness of vowels and consonants: A data-driven approach. Journal of Speech, Language, and Hearing Research, 56(5), 1539–1551. 10.1044/1092-4388(2013/12-0030) [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Wang, J., Samal, A., Rong, P., & Green, J. R. (2016). An optimal set of flesh points on tongue and lips for speech-movement classification. Journal of Speech, Language, and Hearing Research, 59(1), 15–26. 10.1044/2015_JSLHR-S-14-0112 [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Wong, P. C. M., Lai, C. M., Chan, P. H. Y., Leung, T. F., Lam, H. S., Feng, G., Maggu, A. R., & Novitskiy, N. (2021). Neural speech encoding in infancy predicts future language and communication difficulties. American Journal of Speech-Language Pathology, 30(5), 2241–2250. 10.1044/2021_AJSLP-21-00077 [DOI] [PubMed] [Google Scholar]
  65. Yousef, A. M., Deliyski, D. D., Zacharias, S. R. C., & Naghibolhosseini, M. (2022). Detection of vocal fold image obstructions in high-speed videoendoscopy during connected speech in adductor spasmodic dysphonia: A convolutional neural networks approach. Journal of Voice. Advance online publication. 10.1016/j.jvoice.2022.01.028 [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Zhang, Z. (2020). Estimation of vocal fold physiology from voice acoustics using machine learning. The Journal of the Acoustical Society of America, 147(3), EL264–EL270. 10.1121/10.0000927 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Material S1. Compute_NestedModelConfidence.m.
JSLHR-67-753-s001.m (4.4KB, m)
Supplemental Material S2. Compute_RecommendedSampleSize.m.
JSLHR-67-753-s002.m (4.6KB, m)
Supplemental Material S3. Compute_RequiredSampleSize.m.
Supplemental Material S4. Feature_Selection.zip.
JSLHR-67-753-s004.zip (84.5KB, zip)

Data Availability Statement

Mass General Brigham and Mass General are not allowed to give access to data without the principal investigator for the human studies protocol first submitting a protocol amendment to request permission to share the data with a specific collaborator on a case-by-case basis. This policy is based on strict rules dealing with the protection of patient data and information. Anyone wishing to request access to the data used for Comparing Cross-Validation Methods in a Real-World clinical example section must contact Sarah DeRosa (sederosa@partners.org), Program Coordinator for Research and Clinical Speech-Language Pathology, Center for Laryngeal Surgery and Voice Rehabilitation, Massachusetts General Hospital.

The data used for Experiments 1–3 were generated using Monte Carlo simulations. Several MATLAB functions associated with this article are available; see Supplemental Materials S1S4. The latest version of the code can be retrieved from https://github.com/GhasemzadehHamzeh/ML_PowerAnalysis.

The “Compute_RequiredSampleSize.m” code takes the number of extracted features (m_0), the effect size of the best features (d_0), and the number of selected features (l_0) as the input parameters and computes the required number of pairs for having the statistical power for finding a statistically significant model (α = .05, 1 − β = .8). This code can also be used for unbalanced data sets; however, the estimated required sample size (nr) would be the average of the number of negative and positive samples. The sample size of each class can then be computed using the estimated average required sample size and γdb. Specifically, the required sample size of the smaller class will be nr×21+γdb and the required sample size of the larger class will be nr×2×γdb1+γdb. Similarly, this code can also be used for features with different effect sizes; the average of the effect sizes needs to be input as d_0. The “Compute_RecommendedSampleSize.m” code takes the number of extracted features (m_0), the effect size of the best features (d_0), and the target value of C2,2 (CI_0) as the input parameters and computes the recommended number of pairs for achieving the target value of C2,2. The “Compute_NestedModelConfidence.m” code takes the number of extracted features (m_0), the effect size of the best features (d_0), and the number of pairs in a data set (n_0) as the input parameters and computes the value of C2,2. The “Feature_Selection” toolbox provides an implementation of forward feature selection with four investigated cross-validation approaches. Please refer to the “Main_Test.m” file for examples of the usage of the code.


Articles from Journal of Speech, Language, and Hearing Research : JSLHR are provided here courtesy of American Speech-Language-Hearing Association

RESOURCES