Abstract
In recent years, the cognitive neuroscience literature has come under criticism for containing many low-powered studies, limiting the ability to make reliable statistical inferences. Typically, the suggestion for increasing power is to collect more data with neural signals. However, many studies in cognitive neuroscience use parameters estimated from behavioral data in order to make inferences about neural signals (such as fMRI BOLD signal). In this paper, we explore how cognitive neuroscientists can learn more about their neuroimaging signal by collecting data on behavior alone and using alternative estimators designed to leverage this information. We demonstrate through simulation and mathematical derivations that knowing more about the marginal distribution of behavior can improve inferences about the mapping between cognitive processes and neural data. We analyze the magnitude of this benefit, finding that it depends on the desired estimand and several underlying study parameters. While in many cases the absolute gains in precision can be modest, our results demonstrate that, in realistic settings, additional behavioral data can lead to the same improvement in the precision of inferences more cheaply and easily than collecting additional data from subjects in a neuroimaging study. This means that when conducting a neuroimaging study, researchers now have another knob to turn in a design analysis: the number of subjects collected in the scanner and the number of behavioral subjects collected outside the scanner (in the laboratory or online).
Keywords: brain–behavior associations, hierarchical modeling, missing data
1. Introduction
One factor that limits progress in human cognitive neuroscience is statistical precision and power (Cremers et al., 2017; Munafò et al., 2019; Yarkoni, 2009). Many reported effects are small relative to the noise in neuroimaging signals, a problem compounded by the small sample sizes used in most studies. Within a hypothesis testing framework, a lack of precise estimates of some measurements can make it difficult to replicate reported effects. For example, even quite strong, known effects such as the relationship between motor movements and fMRI BOLD activation in sensorimotor regions can be unlikely to reach significance in current standard sample sizes of around 25–30 subjects (Poldrack et al., 2017). Having lower precision also means that any estimated effect that is significant will also have a high probability of the estimate being much greater in magnitude or even having the wrong sign compared with the true effect (Cremers et al., 2017; Gelman & Tuerlinckx, 2000; Yarkoni, 2009). However, due to a combination of factors (including lack of sufficient funding for large sample studies), research in human cognitive neuroscience tends to be under-powered in many cases (Poldrack et al., 2017).
Several remedies have been proposed for this situation, the most obvious being to simply collect more neuroimaging data (Munafò et al., 2019; Poldrack et al., 2017). Because the funding available to most laboratories is relatively small (compared with the cost of a large-scale neuroimaging study), this might require a move toward working in larger consortia around critical topics. Another approach might be to leverage the power of publicly available open datasets, which can allow for meta-analyses across several smaller experiments (Poldrack et al., 2017). However, increasing sample sizes is not the only option for dealing with a lack of precision. For example, increasing the signal-to-noise ratio of neuroimaging measures by creating more detailed statistical models of the neuroimaging signal (Lindquist et al., 2009), improving experimental design (Durnez et al., 2017), or even improving the measurement process itself (Feinberg & Yacoub, 2012; Lombardo et al., 2016) are all measures that are likely to help. In addition, it has been suggested that a focus on computational cognitive modeling and behavior can also allow for extracting more signal (Krakauer et al., 2017; Palmeri et al., 2017) by creating better models of the underlying cognitive and neural processes. In that theme, this paper explores a potentially under-studied option for improving the statistical precision of human cognitive neuroscience studies: collecting additional behavioral data without neural recordings.
We begin by noting that many common statistical analyses in cognitive neuroscience involve behavioral data, typically collected from the subjects who also provided neural data. In one common scenario, the interest is in the relationship between individual differences in behavioral traits and more static measures of neural activity, such as performance in a task and functional connectivity (e.g. Rosenberg et al., 2016). In another typical analysis, researchers are interested in how within-subject changes in a cognitive state variable over time relate to changes in the neural signal from a particular brain region, such as the prediction error in a reinforcement learning task (e.g., O’Doherty et al., 2003). One important, but often overlooked, point is that these behavioral measurements (e.g., survey questionnaires, behavioral responses to stimuli, summaries of task performance or estimates of cognitive model parameters) are themselves noisy estimates of an individual’s underlying cognitive processes and the distributions of these patterns that exist in the population. Thus our ability to infer relationships between brain and behavior is limited by sampling error and estimation error. Sampling error is a result of the fact that a small sample may not be representative of the population, so estimates of the relationship based on the sample may be far from the true relationship. A small sample, therefore, prevents us from being confident about the value of a parameter in a statistical model. Estimation error is a form of measurement error that arises because researchers cannot measure latent cognitive variables directly and instead use models to estimate their values. The relationship between neural recordings and cognitive variables measured with error will necessarily not be as strong as the relationship involving the true values, a notion sometimes known as regression dilution or attenuation. Ignoring these two sources of error can prevent us from learning as much as possible from neural data.
In the following, we demonstrate that one way to reduce both sources of error and improve statistical estimates of the relationship between the brain and behavior is to collect behavioral data from subjects without collecting corresponding neuroimaging data. This idea is somewhat counterintuitive in light of the fact that most common approaches in neuroimaging analyses assume that the only relevant data are from individuals whose neural activity was recorded. However, behavioral data (especially data collected over the Internet) are often significantly less resource intensive to collect than neuroimaging data, in terms of money and the time and effort of both researchers and human subjects. We show how using modern statistical methods that can leverage information from the additional behavioral data can help mitigate the two types of error. To anticipate one consequence of our analysis, we are able to provide a cost–benefit analysis (in terms of statistical power and precision) of collecting data from subjects with only behavior against collecting more expensive additional data from subjects with both behavior and neural recordings. In some cases, due to its relatively low cost, collecting more behavioral data instead of additional neuroimaging data can be a more effective use of resources. This result has implications for designing studies in cognitive neuroscience both prospectively and retrospectively (i.e., the re-analysis of archived open-science datasets.)
1.1. Notation
Throughout the article, we will discuss a situation where we have collected both neuroimaging data () and behavioral data () from Nxy participants and we will ask what the benefit is of collecting an additional Nx participants who only contribute behavioral data (from the same measurement as in the original experiment). We will always refer to constants, such as the number of participants, with capital letters (N). In addition, we refer to data vectors with lower case letters, and . Estimated latent variables from cognitive models (defined below) will be referred to as c when they vary within person (state variables) and θ when they only vary across people (trait variables).
1.2. Common paradigms for relating brain and behavior in neuroimaging studies
For the purposes of articulating our thesis, it will be useful to describe two common types of analyses used in cognitive neuroscience that use behavioral data to interpret neural signals and lay out some notation for the rest of the paper. While this is by no means an exhaustive list, we claim that many analyses can be understood as falling into two broad categories.
1.2.1. Regression (conditional distribution)
In many cognitive neuroscience studies, researchers are interested in how the conditional distribution of neural measures changes as a function of variation in cognitive states induced by a task. For instance, researchers might be interested in regions of the brain that track whether or not it is viewing a face (Kanwisher et al., 1997), the valence of the currently presented stimulus (Chikazoe et al., 2014), or the prediction error in a reinforcement learning model of decision making (O’Doherty et al., 2003). More formally, a subject i presented with a set of stimuli that occur on trials (or times) . Each stimulus presentation sit is assumed to induce some latent cognitive state cit that is associated with some neural activity yit. Researchers can identify neural signals related to the cognitive states by fitting a regression model
| (1) |
where ϵit is uncorrelated noise (for the purposes of this discussion, we ignore features of particular signals such as the fMRI BOLD hemodynamic response). However, this is not quite so straightforward because researchers need to derive estimates of the cognitive state for these analyses. The most common assumption is that the relevant states are closely tied with objective features of the stimulus, such as the color of an image or experimental conditions such as the reward offered for a choice, so cit can reasonably be replaced by the stimulus feature or experimental condition itself.
In other cases, the cognitive states of interest are tied to subjective features of the stimulus, such as the valence associated with sit or the typicality of sit for a particular category. In some cases, researchers collect behavioral ratings xit that depend on these latent states. Researchers then assume that the ratings come from a cognitive process C such that and obtain the cit from that model. In the simplest case, this can mean simply using the valence ratings given by subjects (Chikazoe et al., 2014) or using the average typicality from a group of subjects tested separately (Wilson-Mendenhall et al., 2015).
More recently, “model-based” neuroimaging analyses have investigated more complex relationships between cognitive states and behavior. In addition to the current stimulus sit, the assumption is that cit may also depend on θi latent cognitive “trait” variables that do not vary over the relevant time as well as past cognitive states . Researchers can then use the behavioral data to fit the model and obtain the relevant cit. These states may be, for instance, the reward prediction error or predicted value in a reinforcement learning model of sequential decision making (O’Doherty et al., 2003).
Because the latter two types of analyses involve behavioral data, those will be our focus in this paper. In particular, we will focus on the model-based case but nearly all of our analyses will also apply to the second case.
Modern multivariate pattern analysis techniques (Haxby, 2001; Polyn, 2005) invert this regression and try to predict the cognitive state from neural signals. Typically, this means predicting stimulus properties or experimental conditions, but in some cases the cognitive state will be inferred from behavioral data (e.g. Kragel & Polyn, 2016; Mack et al., 2013). While the example analyses in the rest of the paper do not directly apply to this setting, similar principles of collecting additional behavioral data may hold there as well.
1.2.2. Correlation (joint distribution)
In other cases, researchers are interested in the joint distribution of behavioral measures and neural measures (often across subjects). For instance, is there a relationship between individual differences in behavior and individual differences in static measures of the brain? Researchers often use a direct measure of a subject’s behavior xi such as answers to a questionnaire (Treadway et al., 2013) or performance on a behavioral task (such as a sustained attention task, Rosenberg et al., 2016). More recently (e.g., Homan et al., 2019), researchers have also used cognitive “trait” variables that can be related to behavioral data through cognitive models (i.e., θi above). Static measures can include measures such as functional connectivity or average resting-state activity that do not depend on time. Typically we only get one yi per subject and the measure is not always collected simultaneously to the task where xi is recorded. In this setting, researchers are typically interested in characterizing the joint distribution of neural measures yi and behavioral (xi) or latent cognitive (θi) trait variables. In particular, if the neural measures and behavioral variables can be approximated by a bivariate normal, researchers want to find neural measures where the correlation ρ is nonzero.
Another common analysis that involves correlation is Representational Similarity Analysis (Kriegeskorte, 2008). In these analyses, researchers correlate the similarities of the neural response to pairs of stimuli with the similarities of the two stimuli according to a model. Most often, these models are created based on stimulus properties but they are sometimes derived from behavioral ratings (Bruffaerts et al., 2013; Chikazoe et al., 2014). In these cases, our analyses about the benefits of additional behavioral data will apply as well.
We argue that these two categories of research design and inferential techniques cover a vast array of studies in cognitive neuroscience and neuroimaging in particular1. Each of these analyses can be thought of as attempting to learn about the conditional distribution of neural signals given behavioral data (regression) or the joint distribution of neural and behavioral data (correlation). The focus of this paper asks whether we can improve inferences about the relationship between brain and behavior in these two types of analyses by learning more about the marginal distribution of behavioral data, that is, collecting more behavioral data without neural data.
2. Two Sources of Error in Neuroimaging Analyses
As the Introduction laid out, there are two key sources of error in neuroimaging analyses. In this section, we describe each of the two sources of error in more detail as well as how they can affect the precision of inferences about the relationship between neural signals and behavior.
2.1. Sampling error
In a correlation analysis, we want to know what ρ is in the general population but estimate it by using the sample correlation r. With a small number of subjects, r can be quite far from the true ρ. An approximate standard error for ρ can be constructed as , where N is the number of subjects or, more generally, the number of data points used to estimate r (Fisher, 1915). As this makes clear, the precision of the estimate (and thus finding a ρ that we can say is nonzero with confidence) is constrained by the size of our sample. In many neuroimaging studies, this is often relatively small because neuroimaging data are expensive and time consuming to collect. In Figure 1, we show an example of where the smaller sample size can lead to noisy estimates and larger error bars.2 Sampling error can also be an issue in regression analyses as well. However, because in that case we are only interested in the conditional distribution of the neural signals, we cannot reduce sampling error by collecting more behavioral data so do not discuss it here.
Fig. 1.

Twenty-five samples (dark blue) from a multivariate Gaussian distribution are overlaid over an additional 100 samples (light blue) from the same distribution. Regression lines are plotted showing how estimates from smaller samples can add significant variance to the estimated correlation. Error bars show that we cannot rule out a zero correlation in the smaller sample.
2.2. Estimation error
In the analysis described above, we want to find a set of neural signals that have a linear relationship with a subject’s latent cognitive states .3 The true model we are interested in is
| (2) |
where in particular we are interested in which neural signals have a significant parameter (typically, at the group level). In practice, because we do not have direct access to cit, we typically replace it with estimates ĉit by fitting the model to the behavioral data xit. It is typically assumed that the model does not perfectly account for behavior and that there is some variability in responses beyond what can be predicted by the model. For instance, when fitting a reinforcement learning model, a decision noise parameter is often used to deal with variability beyond what can be accounted for by the latent representations in the model. Because we do not have infinite data, this typically means that several sets of parameters could account for the data even if the model is well identified. In other words, parameters are estimated with some uncertainty. This may seem obvious to many readers but standard techniques that are typically used in cognitive neuroscience do not easily allow this uncertainty to be taken into account in the above regression (c.f., Ly et al., 2017; B. M. Turner, 2015; B. M. Turner et al., 2013, 2017). Instead of the model above, researchers often simply plug in the estimated into the regression, that is,
| (3) |
In this model, the value of the slope estimate using the estimated latent variables, , will not necessarily be the same as when using the true parameter . In fact, it can be shown that on average, the least squares estimate of will be
| (4) |
where is the correlation between the true and the estimates (Wilson & Niv, 2015). This is perhaps not always relevant since the neuroimaging signal can often only be known up to a multiplicative constant. However, Wilson and Niv (2015) also showed that, as long as the estimates are uncorrelated with the regression error term ϵit, the expected t-statistic for the is
| (5) |
where Ti is the number of trials from subject i used to estimate the regression and is the variance of the residuals ϵit, due to noise in the neuroimaging signal. We can see that this t-statistic gets closer to 0 as gets further from 1. A smaller t-statistic means that the standard error (or any multiple of it such as a confidence interval) is more likely to overlap with 0. Thus having a lower correlation between the estimates and the true values, or, in other words, greater estimation error, can make it harder to find relevant signals in the typical regression setting. In Figure 2, we show an example of how estimation error can attenuate a regression fit.
Fig. 2.

Here we show a case where the latent behavioral measure is estimated with error. The correlation with the estimates is lower than it would be with the true latent behavioral measure.
Estimation error can also be relevant for correlation analyses looking at correlations between latent cognitive trait variables and neural signals. If we estimate the trait variable from behavioral data, Katahira (2016) showed that the correlation between estimates and structural neural signals can be decomposed as
| (6) |
Therefore, the correlation is upper bounded by the true correlation between the neural signals and latent trait variables and will decrease as the correlation between the estimates and the true values decreases (or as estimation error increases). Because the size of confidence intervals for a correlation coefficient only depends on the total number of subjects, the intervals for a correlation with a noisier estimate will be more likely to overlap with zero. This means that again, greater estimation error will make it harder to find relevant signals.
3. The Benefits of Additional Behavioral Data
Having laid out why these two errors can deteriorate our inferences, we now demonstrate how collecting additional behavioral data can help. In addition, for each case, we lay out a tradeoff between collecting cheaper behavioral data and more expensive imaging data in terms of how much each improves our inferences. These analyses show in which cases it may actually be more valuable to collect behavioral data than neuroimaging data, even if the main object of interest is the relationship between brain and behavior. To preface, Section 3.1 demonstrates how little known estimators are able to make use of additional behavioral data to reduce sampling error for correlation analyses. We then follow with simulations to understand where this estimator improves over standard ones. Section 3.2 then documents how the use of hierarchical modeling can allow for combining datasets and improving estimation for common regression analyses as well as correlation analyses. Further analytical work shows the benefits of decreasing estimation error for making group-level inferences about the relationship between cognitive processes and neural signals.
Critically, these are two distinct methods by which behavioral data can decrease sampling error and estimation error. For each, we show why each method works and then, using a combination of analytic derivations and simulations, investigate the tradeoff between collecting additional behavioral data vs. neuroimaging data when using these methods. As mentioned above, regression analyses commonly used in cognitive neuroimaging are not affected by the population variance of the behavioral variable. However, in the case of correlational analyses, or analyses of the full joint distribution of neural and behavioral measures more generally, both methods could be used, potentially in combination. The following Table 1 summarizes these situations more succinctly.
Table 1.
Proposed methods.
| Inference/error | Joint | Conditional |
|---|---|---|
| Sampling | Missing data estimators | N/A |
| Estimation | Hierarchical modeling | Hierarchical modeling |
3.1. How to decrease sampling error: Behavioral data as neuroimaging data with a missing variable
Typically we are interested in generalizing beyond our sample and characterizing the joint distribution of neural signals and behavioral variables in the population. However, as mentioned, most neuroimaging studies are limited by getting a relatively small sample. One somewhat counter-intuitive notion from the field of statistics focused on missing data is that we can often get better estimates of a parameter in the population (in terms of a lower mean squared error) by including data points that do not have every variable recorded in our analysis (Little & Rubin, 2019). This is because we are often assuming that variables have some joint distribution. If the correlation between two variables is not zero, then we can know something about the value of what a missing variable must be from knowing the value of the other variable. If we have collected neural and behavioral data from Nxy subjects and only behavioral data from an additional Nx subjects, Anderson (1957) showed that the maximum likelihood estimates of the mean and standard deviation of , as well as the correlation between and , ρ, are not the same when we include all of the data points as they would be if we only include the data with all variables recorded (the complete data). We can see this by rewriting the likelihood for a bivariate normal. If
| (7) |
where and then we can factor the likelihood as
| (8) |
The advantage of writing the joint distribution as we did above is that the first term is just the regression of on that is only defined for the data points that have both variables recorded. Therefore, we can get the maximum likelihood estimate for αxy, βxy, and σxy using standard least squares regression estimates with the Nxy data points. Maximizing the likelihood of the second term only depends on the data so we can use standard estimates of the mean and variance using all of the data. We can write the correlation as
| (9) |
Then plugging in the maximum likelihood estimates , and , we get Anderson’s factored likelihood estimator for ρ using all of the data:
| (10) |
Because the maximum likelihood estimate of σ x using all Nxy + Nx data points will not in general be the same as only using the Nxy data points, the maximum likelihood estimate of ρ using all the data will not be the same as using only the complete data.
In the Appendix, we show that there is a simple relationship between the standard Pearson r and Anderson (1957)’s . Specifically, if is the estimated standard deviation using only the Nxy subjects with neuroimaging data and is the estimated standard deviation using all Nxy + Nx subjects, then
| (11) |
To our knowledge, we are the first to show that only depends on the correlation estimate in the complete cases, r and the ratio of the variance of the behavioral variable in the complete cases and the full data set, . This form makes clear how the new information on the marginal distribution of the behavioral data contributes to the estimated measure of association. In Figure 3, we plot how changes as a function of these two variables, with the variance change plotted on the log scale. In general, if the variance increases when including the additional data, the correlation estimate increases and visa versa.
Fig. 3.

Anderson (1957) estimates ( ) of the correlation using all of the data as a function of Pearson’s r using only Nxy complete cases and the difference in the log of the variance of the behavioral data in the Nxy complete cases and all Nxy + Nx cases. In general, is modified up or down by the relationship between the variance of the behavioral data on the full dataset and that for the complete cases.
3.1.1. When does it work?
Garren (1998) showed analytically that asymptotically, the maximum likelihood estimate using all the data will have lower error than estimates for r using only the complete data. This result suggests that there can be cases where, if one variable is more expensive to collect than the other one (for instance neuroimaging data and behavioral data), the same quality of estimates of the correlation can be achieved for lower cost by collecting some data with the more expensive variable missing (Hocking & Smith, 1972). However, the result is an asymptotic result, meaning that it is only guaranteed to apply in the limit of infinitely large datasets. It is, therefore, still unclear when collecting more behavioral data will be better than collecting more neuroimaging data in practice. In the following, we conduct a simple simulation to explore this tradeoff.
Following the description above, suppose we have collected a behavioral variable () and a neuroimaging variable () from Nxy subjects. In addition, we collected just behavioral data from Nx subjects. For simplicity, we assume that we have a single behavioral measure and a single neuroimaging measure from each subject.
In order to specify the generative model with missing data, we can create a matrix of data from our neuroimaging subjects, an Nxy × 2 matrix with as the first column and as the second. The data matrix is generated for according to
| (12) |
where , and . The additional behavioral subjects then have distribution
| (13) |
For the purposes of our simulation, we set and .
Having generated data according to the above model, we now compare two ways of estimating the correlation, standard Pearson’s r using only data with both variables and the Anderson (1957) estimate of the correlation using all of the data.
As an example, we first compare estimates at for a study with 25 neuroimaging subjects. We then compare Pearson’s r and Anderson’s for seven values of the number of additional behavioral subjects (0, 25, 50, 100, 250, 500, and 1000). Finally, to compare the benefit of the additional behavioral data with collecting an additional neuroimaging subject, we compare all estimates with Pearson’s r with one additional neuroimaging subject. We simulate 1,000,000 datasets and estimate the root mean squared error of all of these estimates averaged across all simulations. We can see in Figure 4 that the Anderson estimator quickly improves over obtaining an additional neuroimaging subject.
Fig. 4.

Root mean squared estimation error, averaged over 1,000,000 simulated datasets, for the estimate of ρ from the two estimators as a function of the true correlation and number of additional behavioral subjects Nx for and Nxy = 25. We also compare with estimates from a neuroimaging dataset with Nxy + 1 subjects. The orange line is below the blue line showing the improved estimate that comes from one additional neuroimaging subject. The curved green line shows the Anderson (1957) changes in the estimate as the number of addition behavior-only subjects is added to the analysis.
Of note here is that collecting neuroimaging data can be quite expensive. For instance in 2022, collecting fMRI data at New York University’s Center for Brain Imaging costs researchers around $450 per hour of subject time. In addition to the price, collecting neuroimaging data can have many other nonmonetary costs such as requiring the subject to come into the laboratory, requiring the time of trained researchers or technicians to run the machine and only being able to run one subject at a time. In contrast, many behavioral studies can be easily run in parallel online without significant monitoring by researchers for a cost of around $11 per hour of subject time (Crump et al., 2013; Hara et al., 2018). Thus, if 40 behavioral subjects can improve inferences more than a single fMRI subject, this might be a worthwhile tradeoff in many experimental designs.
To make our results more general, we compare estimates using five values of Nxy (25, 50, 100, 500, and 1000) and five values of ρ from .1 to .9. For each set of , we generate 1,000,000 datasets and estimate the correlation using both methods. In Figure 5, we plot the benefit of collecting additional behavioral data in terms of the change in root mean squared error of collecting an additional neuroimaging subject and using the standard Pearson r. The gray dotted lines at 0 and -1 in Figure 5 reflect the blue and orange lines in Figure 4. The solid lines indicate the performance of the Anderson (1957) estimator (the green line in Fig. 4). In practice, the Anderson (1957) estimator is not beneficial in all conditions. This is indicated by the color of the solid line: in cases where it does not benefit, we color the solid line red.
Fig. 5.

Relative decrease in root mean squared estimation error, averaged over 1,000,000 simulated datasets, for the estimate of ρ from Anderson (1957) as a function of the true correlation, the number of neuroimaging subjects Nxy, and number of additional behavioral subjects Nx. Values are computed relative to the change in RMSE for a Pearson correlation with Nxy neuroimaging participants (0) compared with Nxy + 1 neuroimaging participants (-1). These values are indicated by the gray dashed lines. The solid line shows the Anderson (1957) changes in the estimate as the number of addition behavior-only subjects is added to the analysis. The solid lines are colored red in the conditions where the Anderson (1957) estimate is not beneficial and blue in the majority of cases where it is. Note that the scale of the y-axes is not fixed across plots but the gray dashed lines have the same meaning across plots. This is necessary in order to see the change with number of subjects.
Because we are emphasizing the tradeoff between collecting neuroimaging and behavioral data, the y-axis of Figure 5 is plotted relative to the precision benefit of one additional neuroimaging subject. However, in the Supplementary Material (Supplementary Fig. S3), we also show the absolute change in precision (in terms of root mean squared error) as a function of the parameters we discuss here.
Comparing the two methods of correlation estimation, we find that the effect of the additional behavioral data can be nonlinear in the size of the underlying neuroimaging dataset and the true value of ρ. Having too few neuroimaging data points to estimate a weak correlation (e.g., only 25 subjects for a correlation of .1) makes the behavioral data less useful and it can be better to collect more neuroimaging data. This finding resembles similar simulation results from Garren (1998). At low amounts of neuroimaging data and lower true values of ρ, the Anderson estimate of the correlation using all of the behavioral data can actually perform worse than the standard Pearson estimate. However, as the true ρ and number of neuroimaging subjects increase, the value of the additional data increases such that collecting additional behavioral data can be better than collecting a smaller amount of additional neuroimaging data. In some cases, collecting just 25 behavioral subjects can improve precision more than another neuroimaging subject. At the cost levels described above, this can mean that behavioral data can be a better investment than equivalently priced neuroimaging data. Therefore, a design analysis that takes costs into account has the potential to recommend collecting behavioral data rather than neuroimaging data.
However, we reiterate that this is not a general recommendation. As shown in Figure 5, it is not true across the whole parameter range and, perhaps in particular, it is not true for low sample sizes and low true correlation values, which is exactly where we might hope to gain in terms of precision. Why is it that the Anderson estimator performs worse here?
To answer this question, we can decompose the mean squared error into the bias and the variance, that is,
| (14) |
where Bias is defined as
| (15) |
As with mean squared error, we can also look at how various methods perform in terms of bias, that is, is the estimate of ρ equal to ρ on average? It is well known that Pearson’s r is biased downward such that, on average, correlation estimates will be lower than their true value (Olkin & Pratt, 1958). In Figure 6, we can see that the Anderson estimator of ρ always decreases the bias in estimating ρ, relative to an Nx of 0, where it is equivalent to the Pearson r. This means that the increase in mean squared error in the Anderson estimate for lower values of ρ and Nxy is due to the behavioral data adding variance to the estimate. It may be that the adjustment shown in Figure 3 is only likely to be in the right direction when there is enough data to estimate ρ and σ x precisely. We hope that future statistical work will address this question and potentially develop new estimators that can control this variance better in the low sample size, low correlation regime. For now, the choice of whether to use the Anderson (1957) estimator will depend on the research question. In the Supplementary Material, we investigate the performance of the Anderson estimator in real data from the Human Connectome Project (Van Essen et al., 2013) and show that it largely matches the performance expected from these simulations (Supplementary Fig. S1).
Fig. 6.

Bias in the Anderson (1957) estimate of ρ, over 1,000,000 simulated datasets, as a function of the true correlation, the number of neuroimaging subjects Nxy and number of additional behavioral subjects Nx. Note, as in the previous figure, that the axes here are not fixed across plots. This is necessary in order to see the change with number of subjects.
A side benefit of the above derivation in Eq. 11 is that it is easy to see how the estimate depends directly on variance estimates. At least since Gauss, statisticians have known to include an adjustment for the degrees of freedom when estimating an unbiased variance from a sample. In the traditional Pearson estimate of the correlation, these adjustments do not appear since they cancel out when the covariance is divided by the variance. However, this is not true of the Anderson ρ estimator, due to the term. In the original formulation Anderson (1957) and in all other treatments we have seen, researchers use the biased maximum likelihood estimate. We find that adjusting for the degree of freedom does improve estimates in the more likely scenarios of low number of joint data and low true correlations (Supplementary Fig. S2). We hope that future work can build on these results to discover improved estimators.
One somewhat counter-intuitive consequence of the above discussion is that Anderson’s estimator can help not only in studies of individual differences across subjects but also in studies using representational similarity analyses where the matrix of item–item similarities based on a model is correlated with the similarities based on the multivariate neural signal (Edelman et al., 1998; Kriegeskorte, 2008). If researchers simply increase the number of similarities computed for the model, the above analyses suggest that there may be cases where this increases the power to detect a correlation with neural measures. In practice, since representational similarity matrices are often created from behavioral ratings (e.g., Bruffaerts et al., 2013; Chikazoe et al., 2014), this is yet another way in which collecting additional behavioral data can increase power. While it is common in this literature to use Spearman’s rank correlation to assess similarities, we can put this within the linear correlation framework by first ranking the two variables and then computing the correlation. For Pearson’s r, this is exactly equivalent to computing the Spearman correlation.
3.2. How to decrease estimation error: Improving within-subject model estimates by collecting data from additional subjects
We now investigate how collecting additional behavioral data can also help reduce estimation error, making it applicable to the regression analyses described above. In our previous description of estimation error, we were agnostic as to how exactly parameters were estimated from behavioral data. The most common method of estimating cognitive model parameters for use in a model-based regression is to use maximum likelihood estimation (Myung, 2003). The maximum likelihood estimate of parameters θi for subject i is the set of parameters in a cognitive model C that maximizes the probability of that subject’s behavioral data that is,
| (16) |
One problem is that many neuroimaging experiments are relatively short so not many trials can be collected per subject. In complex cognitive models, when the number of data points is small, many sets of parameters can be consistent with the data and, therefore, two datasets generated from the cognitive model with the same θi could result in very different maximum likelihood estimates , that is, can be low on average. As demonstrated in the above estimation error section, this will result in lower power for estimating correlations and regression slopes when attempting to relate such variables to neural recordings.
Two methods have been used in the literature to increase and : sharing information across subjects or constraining the possible parameter values. One way to use information from other subjects’ data is to simply assume that all subjects have the same θ values, which means that they can be estimated using all of the behavioral data collected in the experiment (Daw, 2011). This increases the data used to fit the parameters by a factor of Nxy. The will still vary by subject because of the dependence on the stimuli in the experiment. If the variability in true parameters θi across subjects is small relative to the estimation error, this can be an effective, heuristic way to increase . However, this is not an option in a correlation analysis when we are often interested in the variation across subjects itself. In addition, many cognitive processes do have large individual differences, so this may actually decrease .
Another option is to constrain parameter values by adding a prior, P(θi|ϕ). If ϕ is chosen correctly, using a prior can decrease the likelihood of implausible parameter values such as a learning rate of 0 (Daw, 2011). We can then find the parameters that maximize the posterior, that is,
| (17) |
This is known as the maximum a posteriori (or MAP) estimation. By using the prior, we can force θi to be closer to more plausible values. However, it can be difficult to know in general whether your priors are constraining the parameters to be in the correct space. If your priors are wrong, this could also decrease .
One principled way to incorporate information from other subjects and use priors is hierarchical modeling (Gelman, 2006). This is an increasingly popular method in cognitive science and cognitive neuroscience (Ahn et al., 2011; Rouder & Lu, 2005) that assumes that there is a potentially nonuniform distribution P(θi|ϕ) of parameters in the population. Functionally, this constrains the possible values of θi as in the prior above, but because this is describing the population distribution, we can now estimate ϕ from all of the subjects’ data rather than setting ϕ by hand. The resulting estimate of θi will be a weighted average of the estimate from the population with the estimate that just uses an individual’s data. The weights are determined by the strength of the individual data. For a subject whose data provide less constraint on the parameters of the cognitive model (i.e., if they had fewer trials or made more inconsistent choices), their estimates will be closer to the prediction from the prior. If a subject made very consistent choices, we might already have a lot of information about the parameters of their cognitive processes and we do not need to use as much information from the population. Thus, hierarchical modeling provides a way to decide from the data how much we should pool information from the population in creating each individual estimate. Katahira (2016) demonstrates that estimates from a hierarchical model strictly dominate both individual maximum likelihood estimates and population maximum likelihood estimates in terms of bias and variance.
While there exist frequentist methods for estimating hierarchical models (such as restricted maximum likelihood as implemented in popular packages for fitting hierarchical linear models like lme4 (Bates et al., 2015)), for arbitrary cognitive models, fitting in this way often requires a complicated derivation or numerical integration, which can be challenging in high dimensions. With modern Bayesian inference tools such as Stan (Carpenter et al., 2017), it is usually much more straightforward to place a prior on ϕ and use Bayesian inference to compute a joint posterior over ϕ and θ. The posterior for the hierarchical model is then
| (18) |
In addition to Stan, many R or python packages have made hierarchical Bayesian versions of specific cognitive models particularly easy to fit (e.g., (hddm (Wiecki et al., 2013) for drift diffusion models or hBayesDM (Ahn et al., 2017)) for many types of decision-making models).
One feature of hierarchical modeling is that individual hierarchical estimates can often be improved simply by sampling more data from the population. That is, the cognitive model parameter estimates for subjects whose data you already have, that is, subjects who contributed neuroimaging data, can be improved by adding data from new subjects. This is because the error of the estimates of ϕ (or the standard deviation of the prior) will, for most models, converge toward the correct estimates with increasing the number of subjects.
To demonstrate this, we conduct a simulation with a very simple hierarchical model to show that additional behavioral data from new subjects can improve inference for latent cognitive parameters θi for subjects , Nxy of whom were collected in an original set neuroimaging experiment and Nx who were collected in a separate behavioral experiment. We assume each subject’s latent cognitive parameter θi is drawn from a Gaussian population distribution with mean μθ = 0 and variance = 1. Each subject then provides behavioral data bit, and to keep things simple, we assume that the cognitive state variables cit are essentially equal to bit such that
| (19) |
where is drawn from a uniform distribution between e−2 and e0 = 1 for each subject. Further simplifying, we summarize xit by its mean and assume that there are enough trials t that we can treat as known.
For the hierarchical model, we need to put a prior on . While in practice it is often better to use a weakly or strongly informative prior, for generality, in these simulations, we assume an improper uniform prior, that is, . Using the derivation in Gelman et al. (2013), we construct a grid approximation to the posterior for θ given . We now investigate how the correlation between the true subject parameters θi and their posterior mean estimate for the Nxy neuroimaging subjects changes as we add a number of new behavior-only subjects, Nx, to the experiment. Note that we are only computing the correlation within the original Nxy subjects. As we demonstrated in the section on estimation error, this correlation is what limits the power to detect true correlations or regression slopes between model estimates and neural signals.
Figure 7 shows that across a range of Nxy values, the addition of Nx new behavioral subjects can increase the correlation between model estimates for the Nxy neuroimaging subjects and true parameters .
Fig. 7.

Correlation between the estimated cognitive trait parameters and among the Nxy neuroimaging participants using the hierarchical model, averaged over 10,000 simulated datasets, as a function of the number of neuroimaging subjects Nxy and number of additional behavioral subjects Nx.
This implies that one way to increase (and, therefore, also increase ) for subjects in a neuroimaging experiment is to simply collect behavioral data from more subjects. However, the exact rate at which more subjects improve the model estimates will in general depend on the distribution of uncertainty in the individual θi as well as the particular form of the model you use. We also emphasize that while the increases in may be modest, it is possible that, for particular cognitive models, a small improvement in trait estimates could lead to a larger improvement in state estimates . It is these state estimates that are most commonly of interest in a model-based neuroimaging analysis (e.g., reward prediction error in a reinforcement learning model). In order to keep the current work at a high level of generality, we leave evaluating this possibility for particular cognitive models to future work. However, we do think there is reason to believe that this claim may be true for commonly used models. While Wilson and Niv (2015) show that, in a Pavlovian task with fixed rewards, the correlation between prediction error estimates and their true values in a simple Rescorla–Wagner learning model (Rescorla & Wagner, 1972) is not very sensitive to learning rate estimates, they also show that this sensitivity can be much higher in the context of drifting rewards with high variability and autocorrelation. This suggests a regime where modest increase in the accuracy of learning rate estimates could lead to significant improvement in prediction error estimates.
Beyond its use in improving hierarchical estimates of cognitive model parameters, collecting additional behavioral data can improve neuroimaging regression inferences in several other ways. While not computationally feasible yet for most neuroimaging analyses, joint modeling (B. M. Turner et al., 2013) is a recently proposed technique in which the cognitive model and the neuroimaging model are jointly fit, hierarchically, to both the neuroimaging and behavioral data. This can allow for even more efficient use of information by allowing both sources of data to inform the cognitive model parameters. In addition, it allows the uncertainty in the cognitive model parameters to be propagated into the estimates of the regression or correlation, resulting in more powerful estimates without the biases mentioned above. B. M. Turner et al. (2016) demonstrated how joint modeling can be used to improve multimodal inferences even when we do not have data for every mode for every row in the dataset (i.e., some subjects’ data were collected with EEG and others with fMRI). This can be generalized to the case we discuss where we only have rows in the dataset with only one type of data (e.g., behavior). Collecting additional behavioral data in experimental conditions or on novel stimuli not in the neuroimaging experiment could additionally assist in model selection. Choosing the correct model could potentially have a much larger impact on the correlation between model estimates and true parameters.
Finally, we can view reducing estimation error as relevant not only for model-based regression but also for regression approaches where the stimulus features are subjective or latent, such as the valence of a stimulus. The covariates in this case are often created from ratings based either on subjects in the experiment or outside behavioral subjects. Similar to their use with cognitive models, hierarchical models with additional behavioral data can reduce error in these stimulus ratings as well. Even just aggregating the average rating, however, will be improved through greater amounts of data per item. Indeed, going back to at least Kent and Rosanoff (1910), it is already quite common in cases like this to use large “norming” studies to understand stimuli, particularly for word stimuli.
3.2.1. When does it work?
Having demonstrated that it is possible to reduce estimation error in a neuroimaging dataset by collecting more behavioral data, we might now wonder when this is likely to help. As we mentioned above, the amount that hierarchical modeling can help will in general depend on the model and the experimental design. In addition, collecting additional behavioral data can improve estimates in other ways such as through model selection. In order to say something general about how improvements in estimation improve statistical power to find neuroimaging effects, we will simply assume that it is possible to increase the correlation between model estimates and true parameters using behavioral data by a certain amount and investigate how that changes estimates about the relationship with neural data.
In this paper, we have assumed that collecting data from additional subjects is the cheapest way to decrease estimation error (due to the availability of easy online data collection). In addition, we argued above that collecting data from more subjects reduces estimation error in hierarchical model estimates in a straightforward way through improved hyperparameter estimation. However, collecting additional behavioral data from the Nxy neuroimaging subjects will, in many cases, also increase . For instance, increasing the number of trials will necessarily decrease the standard error of the mean of some quantity computed over all of those trials. Which strategy is better will depend on many details but, as our analysis is in terms of the increased correlation between model estimates and true parameters rather than in terms of number of subjects, our results will apply equally to both strategies of reducing estimation error.
To begin, we will build on the derivation from Wilson and Niv (2015) where they showed how the least squares estimate of the regression slope is impacted by misestimation (i.e., equations 4 and 5). Wilson and Niv (2015) only solved this for the case where the analyst is trying to infer a single regression slope from the data (i.e., a nonhierarchical regression model). But in model-based neuroimaging analyses, it is rarely assumed that all subjects have the same effect size. Therefore, the most common analysis method in model-based analysis is to use a two-stage approach where a first-stage regression is fit to each subject (or each run) and then the subject regressions are aggregated at the group level by approximating a hierarchical model, testing whether the average effect is significantly different from zero (Beckmann et al., 2003; Woolrich et al., 2004). This means that the benefit of collecting an additional neuroimaging subject cannot be captured by the Wilson and Niv (2015) equations alone. Following Friston et al. (2002) and Beckmann et al. (2003), we assume a Gaussian hierarchical model of population effects:
| (20) |
| (21) |
where the noise distributions are both Gaussian:
| (22) |
| (23) |
Using the same notation as above, for subject i on trial t, cit is that subject’s cognitive state at that time and yit is the associated neural signal. represents the relationship between the cognitive state and the neural signal for a particular subject and βG represents that average relationship in the population. represents the variance of the magnitude of the relationship in the population and represents the additional variance in the neural signal that is not due to the cognitive state (either due to noise in the measurement or the signal being related to other unrelated cognitive processes as well).
In order to quantify the value of the improved model fit, we first must derive the power for estimating a reliable βG parameter given a particular relationship between the true model parameters and the model estimates. In the Appendix, we obtain an analytic expression under the assumption that the second-level model is estimated using the standard OLS method.
Rather than parameterizing this expression in terms of the raw effect size, βG, it turns out to be more illuminating standardize the effect size by the total standard deviation, . Following Wilson and Niv (2015), we call this standardized effect size the contrast-to-noise ratio (CNR). Additionally, in multilevel modeling, an important statistic is the intra-class correlation coefficient or ICC, which is the proportion of the total variance that is explained by the subject-level variance, that is, (Chen et al., 2018; Shrout & Fleiss, 1979). Given these terms, we get that the expected t-statistic under a particular study design is
| (24) |
We can then compute the power using standard formulas, and in Figure 8, we show how it depends on the model fit for several values of , T, and Nxy. As expected, improving the model fit increases power, although in a highly nonlinear way, meaning that the tradeoff for improving model fit depends on all of these parameters.
Fig. 8.

Power as a function of for a range of CNR, T, and Nxy values. ICC was set to .397, corresponding to a recent estimate in the literature (Elliott et al., 2020). In general, power increases as the number of subjects increases, as the number of trials increases, and as accuracy of the model estimates (the inverse of the estimation error) increases.
This leads us to ask the question “how much would additional behavioral data have to improve our latent variable estimates in order to improve power as much as another neuroimaging subject?” That is, what is the increase (δρ) in the correlation between estimated model parameters and true parameters () that is equivalent to increasing Nxy by 1? More formally, if P(t, N) is the power of the one-sample t-test with a t-stat of t and N subjects, we can write
| (25) |
and ask what value of δρ makes this equation true?
In order to keep our results general, we analyze this situation in terms of how much the correlation would have to increase, a dimensionless quantity, rather than a number of data points. This is because, as we noted above, this amount will depend on the specific cognitive model. In order to translate (δρ) into a specific number of subjects, one could conduct a simulation similar to the one displayed in Figure 7 to obtain the relationship between Nx and for the specific cognitive model of interest.
Since δρ must be between 0 and , this is a straightforward constrained root finding problem that we can solve exactly using numerical methods. We can solve this for several values of the other parameters (CNR, ICC, T, and Nxy) using Brent’s method (Brent, 1972).
Which settings of the parameters in equation 25 are reasonable for neuroimaging experiments? Wilson and Niv (2015) derive estimates for CNR from two different studies with a very wide range of .4 to 11. Their CNR definition was only normalized by σϵ and not the total standard deviation including the group variance as we have done here. Therefore, to translate those values to values in equation, we need to take into account the ICC. A recent meta-analysis of reliability across a wide-range of task-based neuroimaging studies found an average ICC of .397 (Elliott et al., 2020) corresponding to a between subjects variance that is twice as large as the neuroimaging noise, which seems to be quite large. With that value of ICC, Wilson and Niv (2015)’s CNR of .4 corresponds to a group CNR of which is approximately .31 and Wilson and Niv (2015)’s CNR of 11 corresponds to a group CNR of 8.5. We can see in Figure 8 that even a value of 1 for the CNR leads to somewhat implausible power estimates for most cognitive neuroscience task-based studies (Button et al., 2013; B. O. Turner et al., 2018). Therefore, we will assume that the CNR is closer to .31 in the following simulation. However, with the above derivation, it is easy to compute this for other values of the generative model and task parameters. Also of note in these plots is that for a particular value of Nxy, the power function will asymptote as a function of . This means that given a good enough model estimate, it is impossible to improve power more by improving the estimate than by collecting more subjects.
Figure 9 shows the solution to equation 25 in terms of δρ (the improvement to match the estimate with an additional neuroimaging subject) as a function of the correlation between estimate and true parameters with one more neuroimaging subject. Overall, as noted above, there is a range of parameter settings where improving model fit (potentially by collecting more behavioral subjects) cannot help improve the power more than collecting another neuroimaging subject. Wilson and Niv (2015) showed that, for a reinforcement learning model, it is fairly easy to obtain a high , even if the latent subject-level parameters are severely misestimated. This suggests that in that setting, model-fitting “isn’t necessary”, or at least improving the model fit may be unlikely to significantly affect power. However, this may not be the case for other cognitive models, although it seems unlikely that a model that is fit to the data will have a correlation as low as .2 or .3.
Fig. 9.

Solution to equation 25, showing the needed improvement in estimation quality to match an additional neuroimaging subject as a function of parameters for a range of the number of trials (T) and the number of neuroimaging subjects (Nxy). CNR is set to .31 and ICC is set to .397 (Elliott et al., 2020). Lines end when it is no longer possible to achieve better performance by increasing the correlation.
In future work, we hope to explore this in the context of other cognitive models. Assuming that correlations between model estimates and true parameters are reasonably high, and that a smaller δρ is possible to achieve, Figure 9 shows that improving the model fit with additional behavioral data is most useful when there is already a reasonably sized Nxy and T is smaller. While even 50 subjects are already significantly above average for a neuroimaging study, most use number of trials that are much higher than 50. One thing that this suggests is that using extra behavioral data may be particularly useful for neuroimaging studies with small number of trials. These are not common because of the low power they achieve but methods like those mentioned in the previous section may make them more feasible. This opens the door to more neuroimaging studies that are difficult to run with high number of trials such as long-term memory studies where subjects are unlikely to be able to remember a large number of items over a long delay. Finally, in this section, we focus on inferring the conditional relationship between cognitive model parameters and neural signals, assuming that the linear model is correct. However, there are many situations where it might be statistically reasonable or even preferable to use a linear model even if it is known that the neural signals are not exactly a linear function of the cognitive parameters (Buja et al., 2019; White, 1980). Recent work by Azriel et al. (2021) has shown that in this case, inference for regression parameters can be improved by additional unlabeled data as we show in the above work. Future work should also explore the value of behavioral data for model-based fMRI in this approximation regime.
4. Discussion
In this paper, we have presented two ways in which collecting additional behavioral data outside of a neuroimaging experiment has the potential to improve the precision of an estimate of the relationship between brain measures and cognition. We also attempted to quantify a tradeoff between collecting each type of data, showing how this could potentially affect the design of future studies or reanalyses of past work. In both cases, behavioral data had the potential to help but not in all cases.
Critically, our simulations reveal specific parameter regimes where these techniques are most effective, serving as a guide for researchers considering this approach. Regarding sampling error, we found that the benefit of the Anderson estimator is not uniform; it provides the substantial gains over collecting additional neuroimaging data when the original sample size is moderate (e.g., N > 25) and the true underlying correlation is moderate-to-strong ( ). Conversely, in very small samples looking for weak effects, the added variance from the behavioral data can outweigh the benefits. We hope that these somewhat mixed results can spur future work on how to most effectively leverage missing data approaches to combining neural and behavioral data for estimating correlations.
In contrast, the benefit of reducing estimation error via additional behavioral subjects (e.g., through hierarchical modeling) follows a different pattern. This approach is most advantageous when the precision of individual subject estimates is the limiting factor—specifically in studies with low trial counts (T) or high measurement noise. Thus, the hierarchical approach provides a clear benefit by allowing researchers to maximize power in resource-constrained neuroimaging designs, where collecting many trials per subject is infeasible.
The ideas proposed here do have additional limitations as well. For instance, while in this paper we have assumed that data collected with and without a neuroimaging recording device (such as in an fMRI scanner or with an EEG cap on) are exchangeable, that is unlikely to be exactly true in practice. This does not make the described methods impossible to use but does potentially require more complex models that allow for differences between the data collected under different modalities. That being said, there is relatively little work documenting the magnitude of these purported differences. If these differences were found to be large, this would likely have implications for learning about cognition through neuroimaging in general, not just with the statistical methods described here. We strongly encourage both future empirical work documenting these differences and future statistical work on how to aggregate inferences across data collection modalities.
Our analyses suggest important directions for future statistical work, in particular the potential to improve the Anderson (1957) correlation estimate by appropriately reducing the variance, especially in situations with low true correlations. In addition, our analyses suggest new regimes where behavioral data could be helpful; for instance, the potential for improving model fits with behavioral data could allow for feasibly running neuroimaging studies with small number of trials. In many cases, such as in long-term memory studies, it is not feasible to run large number of trials because participants cannot reasonably remember hundreds of items. Both of these directions will require more work to be put into practice, but we hope this paper will inspire collaborations between cognitive neuroscientists and statisticians to work more on ways of fusing data sources in order to learn about brain and behavior relationships.
We are, of course, not the first to have pointed out that collecting data with some modalities missing can still be used to improve inferences. Indeed, this was the perspective taken in B. M. Turner et al. (2016). We believe that we advance the literature in several important ways: (1) we connect this idea to the existing statistical literature with missing data; (2) we show that we can improve inferences not only for parameters of a cognitive model but also for the estimates of the relationship between brain and cognition, often a more important target of inference in cognitive neuroscience; (3) we provide new derivations that help understand (and improve, Supplementary Fig. S2) the Anderson ρ estimator, which we believe is a starting off point for combining behavioral data with neuroimaging data; (4) we provide a general treatment of the issue that is not tied to a particular cognitive model or dataset.
While we have mostly discussed collecting additional behavioral data in the context of designing a new study, it could also be a consideration in a study that reuses an old dataset. It could be that a model is not identifiable with a particular dataset, but it could become identifiable by collecting behavioral data in more conditions. With the advent of more open neuroimaging datasets, and particularly with the current difficulty in collecting new data, these methods point to a way to learn new things from old datasets simply by collecting more cheap behavioral data.
Supplementary Material
A. Appendix
A.1. Derivation of Relationship between Pearson R and Anderson
Here, we try to illuminate how the new estimate is different from the correlation estimate based on only the Nxy complete cases. Little and Rubin (2019) derive the missing data estimator in a different way, showing that
| (26) |
where r, , are the estimates of the correlation (i.e., Pearson’s r), the variance of the behavioral variable, and the variance of the neural variable among the Nxy complete cases. is the variance of y after adjusting for the estimate based on all of the data. Little and Rubin (2019) derive the formula for as
| (27) |
We can instead write this in terms of r and rearrange terms, that is,
| (28) |
Plugging this back in, we get
| (29) |
| (30) |
A.2. Power for Group-Level Parameters as a Function of Model Fit
From Wilson and Niv (2015), we know the distribution of estimates in the first level regression, that is,
| (31) |
Given the population distributions above and using standard formulas for linear combinations of Gaussian random variables, we can get
| (32) |
We can now write the distribution of first-level estimates as
| (33) |
The parameter we are interested in is βG, the average effect in the population. While there are more efficient but computationally expensive estimators (Beckmann et al., 2003; Woolrich et al., 2004), the most common way to estimate βG is the OLS method, that is,
| (34) |
is then just the mean of independent Gaussian random variables and will, therefore, have a Gaussian distribution with mean and variance
| (35) |
| (36) |
The first term is a constant. In the second term, because βc is a Gaussian random variable, the normalized squared sum has a noncentral chi-squared distribution,
| (37) |
and a noncentral chi-squared multiplied by a constant will have a generalized chi-squared distribution. Plugging in the mean of that distribution, the expectation of the variance is now:
| (38) |
| (39) |
Combining this with the expected estimate, we can get a t-statistic at the expected variance
| (40) |
Rather than parameterizing this in terms of the raw effect size, βG, it can be useful for understanding to standardize the variables by the total standard deviation, , similar to the derivation in Wilson and Niv (2015). For clarity, we define two new variables. Following Wilson and Niv (2015), we call the standardized effect size, , the contrast-to-noise ratio (CNR). In multilevel modeling, an important statistic is the intra-class correlation coefficient or ICC that is the proportion of the total variance that is explained by the subject-level variance, that is, (Chen et al., 2018; Shrout & Fleiss, 1979). If we multiply the right-hand side of the above equation by , we can now write the equation in terms of these two variables, that is,
| (41) |
Footnotes
Indeed, it may seem a bit strange to distinguish between these two types of analyses. While conceptually different, researchers frequently use them interchangeably. This is likely due to the fact that standard t-tests for regression slopes and correlation coefficients are mathematically identical. However, this is no longer true when we use alternative estimators like the ones described below and focus on precision rather than testing.
All figures and simulations in this paper are created using custom code in python 3.14.4 using the packages pandas (McKinney, 2010), numpy (van der Walt et al., 2011), matplotlib (Hunter, 2007), scipy (SciPy 1.0 Contributors et al., 2020) and seaborn (Waskom et al., 2020)
In the following we will focus on the case where we are interested in a single latent cognitive variable.
Data and Code Availability
Code for running simulations and generating the figures in this paper is available at https://github.com/dhalpern/getting_blood_from_a_stone. There are no data associated with this paper.
Author Contributions
David Halpern: Conceptualization, methodology, investigation, formal analysis, visualization, writing—original draft. Todd Gureckis: Conceptualization, writing—review and editing, funding acquisition.
Funding
This work was supported by the National Science Foundation under awards DRL-1631436 (T.G.).
Declaration of Competing Interest
The authors declare that they have no competing interests.
Supplementary Materials
Supplementary material for this article is available with the online version here: https://doi.org/10.1162/IMAG.a.1361#supplementary-data
References
- Ahn, W.-Y., Haines, N., & Zhang, L. (2017). Revealing neurocomputational mechanisms of reinforcement learning and decision-making with the hBayesDM package. Computational Psychiatry, 1, 24–57. 10.1162/CPSY_a_00002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ahn, W.-Y., Krawitz, A., Kim, W., Busemeyer, J. R., & Brown, J. W. (2011). A model-based fMRI analysis with hierarchical Bayesian parameter estimation. Journal of Neuroscience, Psychology, and Economics, 4(2), 95–110. 10.1037/a0020684 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Anderson, T. W. (1957). Maximum likelihood estimates for a multivariate normal distribution when some observations are missing. Journal of the American Statistical Association, 52(278), 200–203. 10.1080/01621459.1957.10501379 [DOI] [Google Scholar]
- Azriel, D., Brown, L. D., Sklar, M., Berk, R., Buja, A., & Zhao, L. (2021). Semi-supervised linear regression. Journal of the American Statistical Association, 117, 1–14. 10.1080/01621459.2021.1915320 [DOI] [Google Scholar]
- Bates, D., Mächler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48. 10.18637/jss.v067.i01 [DOI] [Google Scholar]
- Beckmann, C. F., Jenkinson, M., & Smith, S. M. (2003). General multilevel linear modeling for group analysis in FMRI. NeuroImage, 20(2), 1052–1063. 10.1016/S1053-8119(03)00435-X [DOI] [PubMed] [Google Scholar]
- Brent, R. P. (1972). Algorithms for minimization without derivatives. Prentice-Hall. 10.1090/s0025-5718-1975-0371062-9 [DOI] [Google Scholar]
- Bruffaerts, R., Dupont, P., Peeters, R., De Deyne, S., Storms, G., & Vandenberghe, R. (2013). Similarity of fMRI activity patterns in left perirhinal cortex reflects semantic similarity between words. Journal of Neuroscience, 33(47), 18597–18607. 10.1523/JNEUROSCI.1548-13.2013 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Buja, A., Brown, L., Berk, R., George, E., Pitkin, E., Traskin, M., Zhang, K., & Zhao, L. (2019). Models as approximations I: Consequences illustrated with linear regression. Statistical Science, 34(4), 523–544. 10.1214/18-STS693 [DOI] [Google Scholar]
- Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafò, M. R. (2013). Power failure: Why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365–376. 10.1038/nrn3475 [DOI] [PubMed] [Google Scholar]
- Carpenter, B., Gelman, A., Hoffman, M. D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M., Guo, J., Li, P., & Riddell, A. (2017). Stan: A probabilistic programming language. Journal of Statistical Software, 76, 1. 10.18637/jss.v076.i01 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen, G., Taylor, P. A., Haller, S. P., Kircanski, K., Stoddard, J., Pine, D. S., Leibenluft, E., Brotman, M. A., & Cox, R. W. (2018). Intraclass correlation: Improved modeling approaches and applications for neuroimaging. Human Brain Mapping, 39(3), 1187–1206. 10.1002/hbm.23909 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chikazoe, J., Lee, D. H., Kriegeskorte, N., & Anderson, A. K. (2014). Population coding of affect across stimuli, modalities and individuals. Nature Neuroscience, 17(8), 1114–1122. 10.1038/nn.3749 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cremers, H. R., Wager, T. D., & Yarkoni, T. (2017). The relation between statistical power and inference in fMRI. PLoS One, 12(11), e0184923. 10.1371/journal.pone.0184923 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Crump, M. J. C., McDonnell, J. V., & Gureckis, T. M. (2013). Evaluating Amazon’s mechanical turk as a tool for experimental behavioral research. PLoS One, 8(3), e57410. 10.1371/journal.pone.0057410 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Daw, N. D. (2011). Trial-by-trial data analysis using computational models: (Tutorial Review). In Delgado Mauricio R., Phelps Elizabeth A., and Robbins Trevor W. (Eds.), Decision making, affect, and learning. Oxford University Press. 10.1093/acprof:oso/9780199600434.003.0001 [DOI] [Google Scholar]
- Durnez, J., Blair, R., & Poldrack, R. A. (2017). Neurodesign: Optimal experimental designs for task fMRI (preprint). Neuroscience. 10.1101/119594 [DOI] [Google Scholar]
- Edelman, S., Grill-Spector, K., Kushnir, T., & Malach, R. (1998). Toward direct visualization of the internal shape representation space by fMRI. Psychobiology, 26(4), 309–321. 10.3758/BF03330618 [DOI] [Google Scholar]
- Elliott, M. L., Knodt, A. R., Ireland, D., Morris, M. L., Poulton, R., Ramrakha, S., Sison, M. L., Moffitt, T. E., Caspi, A., & Hariri, A. R. (2020). What is the test-retest reliability of common task-functional MRI measures? New empirical evidence and a meta-analysis. Psychological Science, 31(7), 792–806. 10.1177/0956797620916786 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Feinberg, D. A., & Yacoub, E. (2012). The rapid development of high speed, resolution and precision in fMRI. NeuroImage, 62(2), 720–725. 10.1016/j.neuroimage.2012.01.049 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fisher, R. A. (1915). Frequency distribution of the values of the correlation coefficient in samples from an indefinitely large population. Biometrika, 10(4), 507. 10.2307/2331838 [DOI] [Google Scholar]
- Friston, K., Penny, W., Phillips, C., Kiebel, S., Hinton, G., & Ashburner, J. (2002). Classical and Bayesian inference in neuroimaging: Theory. NeuroImage, 16(2), 465–483. 10.1006/nimg.2002.1090 [DOI] [PubMed] [Google Scholar]
- Garren, S. T. (1998). Maximum likelihood estimation of the correlation coefficient in a bivariate normal model with missing data. Statistics & Probability Letters, 38(3), 281–288. 10.1016/S0167-7152(98)00035-2 [DOI] [Google Scholar]
- Gelman, A. (2006). Multilevel (hierarchical) modeling: What it can and cannot do. Technometrics, 48(3), 432–435. 10.1198/004017005000000661 [DOI] [Google Scholar]
- Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., & Rubin, D. B. (2013). Bayesian data analysis (3rd ed.). Chapman and Hall/CRC. 10.1201/b16018 [DOI] [Google Scholar]
- Gelman, A., & Tuerlinckx, F. (2000). Type S error rates for classical and Bayesian single and multiple comparison procedures. Computational Statistics, 15(3), 373–390. 10.1007/s001800000040 [DOI] [Google Scholar]
- Hara, K., Adams, A., Milland, K., Savage, S., Callison-Burch, C., & Bigham, J. P. (2018). A Data-driven analysis of workers’ earnings on Amazon mechanical Turk. Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems - CHI ’18, 1–14. 10.1145/3173574.3174023 [DOI] [Google Scholar]
- Haxby, J. V. (2001). Distributed and overlapping representations of faces and objects in ventral temporal cortex. Science, 293(5539), 2425–2430. 10.1126/science.1063736 [DOI] [PubMed] [Google Scholar]
- Hocking, R. R., & Smith, W. B. (1972). Optimum incomplete multinormal samples. Technometrics, 14(2), 299–307. 10.1080/00401706.1972.10488916 [DOI] [Google Scholar]
- Homan, P., Levy, I., Feltham, E., Gordon, C., Hu, J., Li, J., Pietrzak, R. H., Southwick, S., Krystal, J. H., Harpaz-Rotem, I., & Schiller, D. (2019). Neural computations of threat in the aftermath of combat trauma. Nature Neuroscience, 22(3), 470–476. 10.1038/s41593-018-0315-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hunter, J. D. (2007). Matplotlib: A 2D graphics environment. Computing in Science & Engineering, 9(3), 90–95. 10.1109/MCSE.2007.55 [DOI] [Google Scholar]
- Kanwisher, N., McDermott, J., & Chun, M. M. (1997). The fusiform face area: A module in human extrastriate cortex specialized for face perception. The Journal of Neuroscience, 17(11), 4302–4311. 10.1523/JNEUROSCI.17-11-04302.1997 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Katahira, K. (2016). How hierarchical models improve point estimates of model parameters at the individual level. Journal of Mathematical Psychology, 73, 37–58. 10.1016/j.jmp.2016.03.007 [DOI] [Google Scholar]
- Kent, G. H., & Rosanoff, A. J. (1910). Part II. Association in insane subjects. In A study of association in insanity. (pp. 16–72). American Journal of Insanity. 10.1037/13767-002 [DOI] [Google Scholar]
- Kragel, J. E., & Polyn, S. M. (2016). Decoding episodic retrieval processes: Frontoparietal and medial temporal lobe contributions to free recall. Journal of Cognitive Neuroscience, 28(1), 125–139. 10.1162/jocn-a-00881 [DOI] [PubMed] [Google Scholar]
- Krakauer, J. W., Ghazanfar, A. A., Gomez-Marin, A., MacIver, M. A., & Poeppel, D. (2017). Neuroscience needs behavior: Correcting a reductionist bias. Neuron, 93(3), 480–490. 10.1016/j.neuron.2016.12.041 [DOI] [PubMed] [Google Scholar]
- Kriegeskorte, N. (2008). Representational similarity analysis – Connecting the branches of systems neuroscience. Frontiers in Systems Neuroscience, 2, 4. 10.3389/neuro.06.004.2008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lindquist, M. A., Meng Loh, J., Atlas, L. Y., & Wager, T. D. (2009). Modeling the hemodynamic response function in fMRI: Efficiency, bias and mis-modeling. NeuroImage, 45(1), S187–S198. 10.1016/j.neuroimage.2008.10.065 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Little, R. J. A., & Rubin, D. B. (2019). Statistical analysis with missing data (Third edition). Wiley. 10.1002/9781119482260 [DOI] [Google Scholar]
- Lombardo, M. V., Auyeung, B., Holt, R. J., Waldman, J., Ruigrok, A. N., Mooney, N., Bullmore, E. T., Baron-Cohen, S., & Kundu, P. (2016). Improving effect size estimation and statistical power with multi-echo fMRI and its impact on understanding the neural systems supporting mentalizing. NeuroImage, 142, 55–66. 10.1016/j.neuroimage.2016.07.022 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ly, A., Boehm, U., Heathcote, A., Turner, B. M., Forstmann, B., Marsman, M., & Matzke, D. (2017). A flexible and efficient hierarchical Bayesian approach to the exploration of individual differences in cognitive-model-based neuroscience. In Moustafa A.A. (Ed.), Computational models of brain and behavior (pp. 467–479). 10.1002/9781119159193.ch34 [DOI] [Google Scholar]
- Mack, M. L., Preston, A. R., & Love, B. C. (2013). Decoding the Brain’s algorithm for categorization from its neural implementation. Current Biology, 23(20), 2023–2027. 10.1016/j.cub.2013.08.035 [DOI] [PMC free article] [PubMed] [Google Scholar]
- McKinney, W. (2010). Data structures for statistical computing in python. Proceedings of the 9th Python in Science Conference, Austin, 28 June–3 July 2010, 56-61. 10.25080/Majora-92bf1922-00a [DOI] [Google Scholar]
- Munafò, M. R., Cremers, H. R., Wager, T. D., & Yarkoni, T. (2019). Power and design considerations in imaging research. In Casting light on the dark side of brain imaging (pp. 73–78). Elsevier. 10.1016/B978-0-12-816179-1.00011-6 [DOI] [Google Scholar]
- Myung, I. J. (2003). Tutorial on maximum likelihood estimation. Journal of Mathematical Psychology, 47(1), 90–100. 10.1016/S0022-2496(02)00028-7 [DOI] [Google Scholar]
- O’Doherty, J. P., Dayan, P., Friston, K., Critchley, H., & Dolan, R. J. (2003). Temporal difference models and reward-related learning in the human brain. Neuron, 38(2), 329–337. 10.1016/S0896-6273(03)00169-7 [DOI] [PubMed] [Google Scholar]
- Olkin, I., & Pratt, J. W. (1958). Unbiased estimation of certain correlation coefficients. The Annals of Mathematical Statistics, 29(1), 201–211. 10.1214/aoms/1177706717 [DOI] [Google Scholar]
- Palmeri, T. J., Love, B. C., & Turner, B. M. (2017). Model-based cognitive neuroscience. Journal of Mathematical Psychology, 76, 59–64. 10.1016/j.jmp.2016.10.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Poldrack, R. A., Baker, C. I., Durnez, J., Gorgolewski, K. J., Matthews, P. M., Munafò, M. R., Nichols, T. E., Poline, J.-B., Vul, E., & Yarkoni, T. (2017). Scanning the horizon: Towards transparent and reproducible neuroimaging research. Nature Reviews Neuroscience, 18(2), 115–126. 10.1038/nrn.2016.167 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Polyn, S. M. (2005). Category-specific cortical activity precedes retrieval during memory search. Science, 310(5756), 1963–1966. 10.1126/science.1117645 [DOI] [PubMed] [Google Scholar]
- Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and non-reinforcement. Classical Conditioning, Current Research and Theory, 2, 64–69. 10.1016/s0079-7421(08)60383-7 [DOI] [Google Scholar]
- Rosenberg, M. D., Finn, E. S., Scheinost, D., Papademetris, X., Shen, X., Constable, R. T., & Chun, M. M. (2016). A neuromarker of sustained attention from whole-brain functional connectivity. Nature Neuroscience, 19(1), 165–171. 10.1038/nn.4179 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rouder, J. N., & Lu, J. (2005). An introduction to Bayesian hierarchical models with an application in the theory of signal detection. Psychonomic Bulletin & Review, 12(4), 573–604. 10.3758/BF03196750 [DOI] [PubMed] [Google Scholar]
- SciPy 1.0 Contributors, Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R.,… van Mulbregt, P. (2020). SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods, 17(3), 261–272. 10.1038/s41592-019-0686-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. 10.1037/0033-2909.86.2.420 [DOI] [PubMed] [Google Scholar]
- Treadway, M. T., Buckholtz, J. W., & Zald, D. H. (2013). Perceived stress predicts altered reward and loss feedback processing in medial prefrontal cortex. Frontiers in Human Neuroscience, 7, 180. 10.3389/fnhum.2013.00180 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Turner, B. O., Paul, E. J., Miller, M. B., & Barbey, A. K. (2018). Small sample sizes reduce the replicability of task-based fMRI studies. Communications Biology, 1(1), 62. 10.1038/s42003-018-0073-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- Turner, B. M. (2015). Constraining cognitive abstractions through Bayesian modeling. In Forstmann B. U. & Wagenmakers E.-J. (Eds.), An introduction to model-based cognitive neuroscience (pp. 199–220). Springer New York. 10.1007/978-1-4939-2236-9ETM10 [DOI] [Google Scholar]
- Turner, B. M., Forstmann, B. U., Love, B. C., Palmeri, T. J., & Van Maanen, L. (2017). Approaches to analysis in model-based cognitive neuroscience. Journal of Mathematical Psychology, 76, 65–79. 10.1016/j.jmp.2016.01.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Turner, B. M., Forstmann, B. U., Wagenmakers, E.-J., Brown, S. D., Sederberg, P. B., & Steyvers, M. (2013). A Bayesian framework for simultaneously modeling neural and behavioral data. NeuroImage, 72, 193–206. 10.1016/j.neuroimage.2013.01.048 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Turner, B. M., Rodriguez, C. A., Norcia, T. M., McClure, S. M., & Steyvers, M. (2016). Why more is better: Simultaneous modeling of EEG, fMRI, and behavioral data. NeuroImage, 128, 96–115. 10.1016/j.neuroimage.2015.12.030 [DOI] [PubMed] [Google Scholar]
- van der Walt, S., Colbert, S. C., & Varoquaux, G. (2011). The NumPy array: A structure for efficient numerical computation. Computing in Science & Engineering, 13(2), 22–30. 10.1109/MCSE.2011.37 [DOI] [Google Scholar]
- Van Essen, D. C., Smith, S. M., Barch, D. M., Behrens, T. E., Yacoub, E., & Ugurbil, K. (2013). The WU-Minn Human Connectome Project: An overview. NeuroImage, 80, 62–79. 10.1016/j.neuroimage.2013.05.041 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Waskom, M., Botvinnik, O., Ostblom, J., Gelbart, M., Lukauskas, S., Hobson, P., Gemperline, D. C., Augspurger, T., Halchenko, Y., Cole, J. B., Warmenhoven, J., Ruiter, J. D., Pye, C., Hoyer, S., Vanderplas, J., Villalba, S., Kunter, G., Quintero, E., Bachant, P.,…Brian. (2020). Mwaskom/seaborn: V0.10.1 (April 2020). 10.5281/ZENODO.3767070 [DOI]
- White, H. (1980). Using least squares to approximate unknown regression functions. International Economic Review, 21(1), 149. 10.2307/2526245 [DOI] [Google Scholar]
- Wiecki, T. V., Sofer, I., & Frank, M. J. (2013). HDDM: Hierarchical Bayesian estimation of the drift-diffusion model in Python. Frontiers in Neuroinformatics, 7, 14. 10.3389/fninf.2013.00014 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wilson, R. C., & Niv, Y. (2015). Is model fitting necessary for model-based fMRI? PLoS Computational Biology, 11(6), e1004237. 10.1371/journal.pcbi.1004237 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wilson-Mendenhall, C. D., Barrett, L. F., & Barsalou, L. W. (2015). Variety in emotional life: Within-category typicality of emotional experiences is associated with neural activity in large-scale brain networks. Social Cognitive and Affective Neuroscience, 10(1), 62–71. 10.1093/scan/nsu037 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Woolrich, M. W., Behrens, T. E., Beckmann, C. F., Jenkinson, M., & Smith, S. M. (2004). Multilevel linear modelling for FMRI group analysis using Bayesian inference. NeuroImage, 21(4), 1732–1747. 10.1016/j.neuroimage.2003.12.023 [DOI] [PubMed] [Google Scholar]
- Yarkoni, T. (2009). Big correlations in little studies: Inflated fMRI correlations reflect low statistical power—Commentary on Vul et al. (2009). Perspectives on Psychological Science, 4(3), 294–298. 10.1111/j.1745-6924.2009.01127.x [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Code for running simulations and generating the figures in this paper is available at https://github.com/dhalpern/getting_blood_from_a_stone. There are no data associated with this paper.
