Abstract
Visualizations are vital for communicating scientific results. Historically, neuroimaging figures have only depicted regions that surpass a given statistical threshold. This practice substantially biases interpretation of the results and subsequent meta-analyses, particularly towards non-reproducibility. Here we advocate for a “transparent thresholding” approach that not only highlights statistically significant regions but also includes subthreshold locations, which provide key experimental context. This balances the dual needs of distilling modeling results and enabling informed interpretations for modern neuroimaging. We present four examples that demonstrate the many benefits of transparent thresholding, including: removing ambiguity, decreasing hypersensitivity to non-physiological features, catching potential artifacts, improving cross-study comparisons, reducing non-reproducibility biases, and clarifying interpretations. We also demonstrate the many software packages that implement transparent thresholding, several of which were added or streamlined recently as part of this work. A point-counterpoint discussion addresses issues with thresholding raised in real conversations with researchers in the field. We hope that by showing how transparent thresholding can drastically improve the interpretation (and reproducibility) of neuroimaging findings, more researchers will adopt this method.
INTRODUCTION
Beyond performing experiments and recording data, a scientist is responsible for synthesizing information, interpreting it within the context of prior knowledge, and communicating it effectively. This curation and distillation involves a careful balance of contextualization and concise communication, without losing meaningful information. Data visualization is itself an important analysis step with key processing choices to be made, though their impact is often overlooked and underappreciated. Here we examine these aspects for results reporting in neuroimaging, where typical datasets are complex and multi-dimensional. We start by describing how the field and its research questions have changed over time, and then discuss how the presentation of results in brain images should similarly evolve. The proposed improvements are straightforward and implementable in a large number of widely used software packages. Many of the presented examples use functional magnetic resonance imaging (FMRI) data, but the concepts and methods of improving results reporting are directly applicable to other imaging modalities.
Background: historical context
Localization was an early focus of neuroimaging researchers, who broadly conceptualized the brain as a set of discrete regions with well-defined functions to be identified and displayed, in line with empirical procedures from lesion based neuropsychology and early cognitive psychology (e.g., see Savoy, 2001). Unfortunately, neuroimaging signals—particularly FMRI recordings—are noisy, dynamic, and contain smooth patterns whose size and shape is hard to delineate. To combat these initial challenges, neuroscientists turned their attention to standard null hypothesis significance testing as a way to implement a clear filtering mechanism. Within this framework, strict thresholding was a key processing step for localizing a small number of regions of interest and reducing potential false positives (Forman et al., 1995; Nichols and Hayasaka, 2003; Smith and Nichols, 2009). An emphasis was placed on the results having a small number of suprathreshold clusters and rigorous boundaries of statistical significance. Subthreshold regions were interpreted as simply noise or non-neuronal features (negative-signed activations were even filtered in many cases, prior to interest in task-negative networks). Therefore, researchers presented results with strict or “opaque” thresholding: wherever a statistic value was subthreshold, all modeling results—including the effect estimate, p-value and statistic itself—are fully hidden from view and withheld from consideration.
However, this approach comes with inherent challenges. The adoption of univariate statistics on interdependent data like voxels and vertices requires statistical adjustment for multiple comparisons to reduce false positives across these thousands of tests. At the same time, this also causes meaningful effects to fall below the substantially decreased statistical power (Cremers et al., 2017; Lohmann et al., 2017; Bacchetti, 2013), causing a “tip of the iceberg effect” (Pang et al., 2023; Noble et al., 2024; Sundermann et al., 2024). Adopting multivariate analyses (e.g. searchlight-based multivariate pattern analysis, MVPA) does not solve this problem, since statistical peaks might not correspond to biologically meaningful locations, but instead to the center of maximally informative neighborhoods (Etzel et al., 2013). Consequently, replication studies using similarly stringent thresholds may yield results that appear strikingly different, potentially causing undue concerns about low reliability and highlighting the limitations of this traditional framework. Even within a single study, relevant results may fall below the strict corrective adjustments (i.e., fall “below the waterline”), biasing the evaluation and interpretation.
Over time, methodological and conceptual changes have also occurred in the field. Modern neuroimaging presents a more intricate and nuanced picture of brain activity. Network-based and connectomic studies have become much more common, shifting to a paradigm in which most functions involve the interaction of many parts of the brain to varying degrees (e.g., Calhoun et al, 2001; Calhoun et al., 2002; Greicius et al., 2003; Damoiseaux et al., 2006; Hagmann et al., 2008; Fair et al., 2009; Smith et al., 2009; Biswal et al., 2010; Van Essen et al., 2013; Pessoa, 2014). While regions with high statistical strength might still be particularly important, they do not simply indicate small regions turning on/off in isolation. Responses modulate, and importance also lies with other parts of their (or other) networks that might have weaker effects (e.g., see Noble et al. 2024). That is, even when focused on “significant clusters,” the subthreshold results across the rest of the brain will still provide necessary context for understanding their role and for having a more complete picture of how the brain is behaving. Gonzalez-Castillo et al. (2012)’s deep scanning study further showed brain activation at many scales and changing extents of regional activation with increased data. Recent approaches of modeling eigenmodes of brainwide function have reinforced the importance of visualizing nonlocal and subthreshold effects (Pang et al., 2023).
As a result, the idea of an exactly zero response in any gray matter seems unlikely. Treating it as such—which is what standard opaque thresholding effectively does—creates statistical issues (Cremers et al., 2017). Even if zero effects existed, concluding that subthreshold test statistics imply zero effect amounts to confirmation of the null hypothesis, which goes against the principles of hypothesis testing. In the view of modern neuroimaging, opaque thresholding also wastes meaningful information, since responses and effects are not simply localizable in an on/off manner, as responses occupy a continuous spectrum (Chen et al., 2022). This is one reason that clinical practitioners often include fully unthresholded images in their assessments, to see more context and reduce false negatives (Voets et al., 2025). In general, the results of a neuroimaging study should account for the non-dichotomous nature of data by including subthreshold information, both when the study authors are interpreting it and when readers are engaging with their work.
Reporting results: the past and the future
The issue of how figures are made might seem like merely a stylistic choice, but it is central to how scientists evaluate and interpret results, how readers assess them, and how meta-analyses compare them. Data visualization is an analysis step, and thresholding data is one of the final processing choices researchers make in a study. It has important consequences for evaluating and understanding results.
The current standard practice of opaque thresholding is rooted in the assumptions of the earliest neuroimaging studies. It has remained largely unchanged and, as a consequence, so have the basic figures and representations of results in neuroimaging studies, even though our understanding of brain function has grown in many ways. Previous work has noted problems with opaque thresholding (Allen et al., 2012; Chen et al., 2022; Taylor et al., 2023; Sundermann et al., 2024). Motivated by this, here we identify and focus on three ways that opaque thresholding negatively impacts the fundamental interpretations and comparisons of neuroimaging results:
Unrealistic biology: Opaque thresholding treats all subthreshold regions as if they had zero effect rather than simply statistically weaker observations. This creates an unrealistic ON/OFF picture of localized effects, and is not consistent with the current understanding of brain functioning.
Ambiguity: Brain regions are generally part of overlapping networks, rather than purely isolated and independent. Opaque thresholding removes the context of any results: how quickly effects drop off spatially, what network(s) a cluster is involved with, etc. Other parts of the brain might have higher uncertainty but they still contain useful context, such as evidence for the full range of effects, wider network interpretations, and comparisons.
Bias: The thresholding of a continuous brain effect or related metric mathematically introduces biases and hypersensitivity to small (often arbitrary) differences, such as between effects minimally below and above the threshold. These will negatively impact within-study evaluation and cross-study reproducibility.
As a consequence, opaque thresholding undermines the content of the results themselves. It compromises the ability to make holistic and accurate interpretations by study authors, readers, and meta-analyses at the most fundamental levels.
To improve neuroimaging visualization, Allen et al. (2012) proposed transparent thresholding as an effective way to include both suprathreshold and subthreshold information in reported results. This is a simple, meaningful, and straightforward solution whereby suprathreshold regions are highlighted in an image by being opaque and outlined, while subthreshold results are also included by using an opacity that decreases with their absolute statistical value. This balanced approach allows for having a concise summary of regions with the strongest effects (the same suprathreshold regions in existing figures), together with a graded assessment of activity across the rest of the brain (reporting information across the “missing majority” of the data that had been gathered and analyzed). By shifting to transparent thresholding, the data modeling and results are more accurately represented, no information is lost, and meaningful evidence is gained.
To date, this visualization approach has not been widely adopted in the field, but it has been applied effectively in several neuroimaging studies (a sample of these are provided in Table S1 of the Supplements). Its utility in providing more complete evaluations of results has also been shown in direct comparisons with standard, opaque thresholding. For example, Taylor et al. (2023) used transparent thresholding to reveal a previously underappreciated degree of consistency and reproducibility in the FMRI results of the Neuroimaging Analysis Replication and Prediction Study (NARPS; Botvinik-Nezer et al., 2020).
Here, we provide four examples that demonstrate the many benefits of retaining context in results. Each highlights an important aspect of reporting and interpreting results that is notably improved with transparent thresholding, namely: improving the understanding of a study, reducing hypersensitivity to arbitrary features (sample size, many processing choices and more), avoiding problematically selective reporting, and enhancing meta-analyses. In the later examples, we also argue that, relatedly, effect estimates from modeling should also be visualized whenever available. These provide an additional source of important information in both figures and comparisons. In the Discussion, we address several considerations and concerns that researchers in the field have raised about transparent thresholding. Changes to conventions often face hesitation, and these points are important to consider. But the scientific costs of using opaque thresholding are demonstrably high, and the benefits of showing context with transparent thresholding are both clear and abundant.
We note that one barrier to adopting a new approach is its availability to researchers. We highlight that transparent thresholding is now widely available across a large number of neuroimaging software packages. While a small number of toolboxes have previously included transparent thresholding, several others have been recently implemented or significantly streamlined this approach during this project. These tools are listed (with example images) in the Discussion, and further details are contained in the Supplements.
RESULTS
We present four examples using real data that illustrate the importance of showing context in brain images. These beneficial features include: removing ambiguity in interpreting the data (Ex. 1); accurately comparing data and assessing differences, both visually and in more formal meta-analyses (Ex. 2); avoiding hypersensitivity to non-physiological features, such as sample size (Ex. 3); and reinforcing robust and accurate interpretations (Ex. 4).
Example 1: context to reduce ambiguity
Fig. 1A shows group results from a set of Flanker task-based FMRI data (see Chen, Pine, et al. 2022) presented in the conventional style. In this figure, data processing includes opaque thresholding at a family-wise error (FWE) rate of 5% (via voxelwise p = 0.001 and cluster-correction N = 40 voxels). The displayed results only contain the lobes of a single cluster appearing in the right intraparietal sulcus. Since opaque thresholding is applied, no other information is conveyed for this slice than the fact that the absolute value of the statistic at every non-cluster location was subthreshold. Some potentially interesting clusters might be just one voxel below the estimated cut-off, but they remain entirely hidden. The reader can only interpret the biological implications of this study from this sparse statistical information alone: thus, only a single region appears to show significant response, and the activity seems fully lateralized in the right intraparietal sulcus.
Figure 1.
Results reporting examples, showing a single slice of task-based FMRI data (see Chen, Pine, et al., 2022). Each neuroimaging panel shows the same axial slice in MNI template space at z = 36S (image left = subject left), with thresholding is applied at voxelwise p = 0.001 and cluster size = 40 voxels (FWE = 5%). The data used for both overlay coloration and thresholding are the Z-score statistics. Panel A displays FMRI results using conventional strict (or opaque) thresholding, and shows one cluster in the right intraparietal sulcus. Panel B displays the same results with transparent thresholding (suprathreshold regions are opaque and outlined; subthreshold regions fade as the statistic decreases), revealing relevant context in the subthreshold regions that are hidden in A. Panel C shows a classic example from Anscombe (1973) of the risks of over-reducing data, here for a simple scatterplot. Panel D shows how the same considerations apply to neuroimaging: each dataset would have very different interpretations and biological implications, which can be appreciated with transparent thresholding (same colorbar as B), but when using opaque thresholding that context is lost and each slice reduces to the same image (that of panel A). Only by displaying the more full context with subthreshold visualization can the degeneracy be broken and results more accurately understood. Opaque thresholding removes context and can often lead to a misinterpretation of results.
The scale of information loss can be appreciated by viewing Fig. 1B, where the thresholding has been applied transparently to retain context in the same data. This reveals a brainwide set of nontrivial responses, many of which seem biologically relevant despite being statistically subthreshold. Note that the opacity fades quadratically with decreasing statistic value, so the most notable “extra” results are still near the threshold level (compare the colorbars in Figs. 1A-B). Seeing this context assists in interpreting the most significant regions, which are still highlighted with full opacity and outlines. It also reduces possible misinterpretation. For example, the richer context of the modeling results implies that the “lone lateralized cluster” actually has much stronger left-right symmetry than judged from Fig. 1A (though still with a larger response on the right). The arbitrary influence of opaque thresholding on laterality measures has been particularly noted in clinical contexts (Ruff et al., 2008; Seghier, 2008; Suarez et al., 2009), and though not yet widely adopted, transparent thresholding would greatly improve evaluations. We also see that opaque thresholding has over-reduced the results of the whole-brain modeling, while Fig. 1B contains a more accurate representation of the full study evidence and invites further research to explore the involvement of network components (several of which also appear focal and lateralized).
The interpretational consequences of over-reducing results are not just a quirk of FMRI, and cases from other fields provide useful lessons. For instance, in the classic example of Anscombe’s Quartet (Anscombe, 1973), four scatterplots show data that have the exact same summary statistics and correlation value, but with very different underlying patterns (see Fig. 1C). If the results of the correlation analysis are only reported as summary statistics, important information is lost and one is highly likely to make an entirely incorrect inference about the data itself. Only by keeping the contextual information of the plot can the ambiguity of the summary values be resolved. Acknowledgment of these issues has led researchers in many areas to take action in improving their statistical reporting: they show scatterplots of data (not just fit values); they use violin plots and raincloud plots to detail distributions; and more.
Identical reasoning further emphasizes the importance of using transparent thresholding to retain context in neuroimaging reports. Fig. 1D shows images of four sets of possible results, related to the Flanker task in Fig. 1A. Transparent thresholding allows us to see that there are major differences among them, but with opaque thresholding they each reduce exactly to Fig. 1A and are indistinguishable. That is, opaque thresholding inherently produces problematic ambiguity because the interpretation of the statistically significant cluster would be dramatically different across the four cases, with each image representing a very different biological implication: 1) fairly symmetric left-right activation (actual response); 2) anti-symmetric left-right activation; 3) strongly right-lateralized response; 4) likely a noise- or artifact-driven outcome rather than a task-related physiological one. From the loss of context with opaque thresholding, there is little choice but to apply Occam’s Razor and interpret the results as pointing toward strong or absolute laterality (something like Image #3 in Fig. 1D), which the fuller context does not support. By thresholding transparently, one observes a more complete representation of the results while still having the most significant regions highlighted.
Seeing results beyond the tip of the iceberg also improves localization, by revealing the spatial drop-off of the response (here, in terms of statistical value, but below we show how the effect evidence can be presented). This further provides useful information about the properties of the data itself such as its spatial smoothness, which reflects both acquisition and processing. Interestingly, spatial smoothness was highlighted as a major factor for varied outcomes in the NARPS project (Botvinik-Nezer et al., 2020), so having this information directly available in primary figures may be particularly useful for cross-study comparisons. Finally, subthreshold visualizations can help distinguish among potential underlying features, such as motion versus respiratory effects, additionally benefitting quality control and assurance efforts.
Example 2: context to improve cross-study comparison and meta-analysis
Retaining context is particularly important for comparing datasets and performing meta-analyses. The NARPS project gathered results from approximately 70 teams who had processed the same task-based FMRI data collection independently and reported on specific hypotheses. In their primary comparison of FMRI results, the NARPS authors assessed the similarity of binarized yes/no responses to nine region-specific hypotheses, based on opaquely thresholded statistical results. They reported observing a “substantial variability in reported binary results, with high levels of disagreement across teams on a majority of tested hypotheses.” They also noted that the similarity of teams’ opaquely thresholded maps, as measured by cluster overlap, was low. However, when they compared the unthresholded results, they found “a large cluster of teams had statistical maps that were strongly positively correlated with one another” and that “analyses of the underlying statistical parametric maps on which the hypothesis tests were based revealed greater consistency than expected from those inferences.” That is, comparisons on results processed with standard thresholding showed relative disagreement, while those without thresholding showed a large amount of agreement. The second message has tended to be greatly underemphasized, and this widely cited paper has overwhelmingly been referenced simply as evidence of high variability across processing pipelines. Instead, it should be viewed as a demonstration of the influence thresholding has as a processing step and an important warning of the biases of opaque thresholding, which preferentially tilt study comparisons towards non-reproducibility.
We can see how the choice of thresholding explains the opposing meta-analytic results—high variability when thresholding vs surprising similarity without it—by visualizing a representative set of teams’ results. Fig. 2A shows nine teams’ results from analyzing the same FMRI dataset for NARPS Hypotheses 1, and applying opaque thresholding at |Z| or |t| = 3 (equivalent quantities due to the degree of freedom count; see Supplements). While a few of the teams share somewhat similar suprathreshold clusters, several contain almost no results (3rd column) or sparse regions, yielding an impression of high variability across the teams, consistent with the NARPS authors’ primary meta-analysis findings. However, switching to transparent thresholding (Fig. 2B) immediately reveals that the statistical patterns are actually quite similar across the majority of teams, just with differing magnitudes. That is, the spatial patterns have high correlation, consistent with the secondary meta-analytic findings in NARPS. These observations are quantified and summarized in the similarity matrices shown in Fig. 2C and 2D, which are again consistent with NARPS’s primary and secondary meta-analyses, respectively. Taylor et al. (2023) showed how these patterns of predominantly high similarity (with a small number of outliers) persist across the full set of teams’ results and all hypotheses, when applying transparent thresholding.
Figure 2.
Visualizing results from 9 teams who participated in NARPS (Botvinik-Nezer et al., 2020); team IDs are shown in each panel. Panels A-B show Z and t statistic maps in the same sagittal slice of MNI space and thresholded at |Z| or |t| = 3. Panel A shows the results with opaque thresholding (in line with the study’s primary meta-analysis), which suggest high variability, inconsistency and disagreement across teams. Panel C shows the corresponding similarity matrix (using Dice coefficients for the binarized cluster maps), which quantifies the generally poor agreement. Panel B shows the same data with transparent thresholding (in line with the study’s second meta-analysis), where it becomes apparent that the results actually agree strongly for most subjects, but with varied strength. Panel D shows the corresponding similarity matrix (using Pearson correlation for the continuous statistic maps), showing the typically higher similarity. Transparent thresholding does not uniformly increase similarity, but allows for clearer interpretation of real differences (e.g., bottom right image). Opaque thresholding biases towards dissimilarity (e.g., 3rd column, top and middle). See Taylor et al. (2023) for similar comparisons across the full set of NARPS teams and hypotheses, where the same patterns hold.
This example demonstrates how having the full context in the images is key to understanding the full scope of results—particularly the kind of variability that is present.1 The dominant variability across the NARPS teams’ results is actually that of the magnitude of the statistics, rather than of the sign or spatial pattern of statistical maps. The opaquely thresholded maps cannot distinguish between these kinds of variability below an elevated cut-off point, and hence they bias the interpretation towards one of simply high variability and “lack of reproducibility.” Seeing the full context as part of the comparison helps to reduce and resolve this bias.
Note that while opaque thresholding biases comparisons towards non-reproducibility, using transparent thresholding to retain context does not necessarily increase similarity. It merely allows for better informed comparisons. Consider the third column in each of Fig. 2A and 2B: while transparent thresholding reveals meaningful statistical patterns in the upper two images, it also shows that the bottom image values are uniformly quite low and negligible by comparison. This is another case of resolving ambiguity, as described in Fig. 1: what ostensibly appeared to be three examples of the same thing under opaque thresholding were actually revealed to be distinct by using transparency. Seeing more context in the bottom-middle image reveals that its results have much lower similarity to the others above it, which is also reflected in its much lower correlation value (Fig. 2D). In all cases, these more informative observations can lead to a more impartial assessment of a particular team’s results, and encourage a helpful examination of the corresponding preprocessing and analytic choices.
Thus, while there is variability in the NARPS teams’ processing and results, two different meta-analysis approaches provide very different assessments of it both visually and quantitatively. The one associated with opaque thresholding (which hides useful information, leaves ambiguity, and biases the comparison towards dissimilarity) suggests high variability. In contrast, the one associated with transparent thresholding (which leaves context, provides more modeling evidence, and represents a more complete picture of the data) suggests widespread similarity with varied magnitude. That is, rather than being an inherent feature of the teams’ results, the outcome of high variability is primarily due to the processing choice of applying opaque thresholding prior to comparison, which introduces a strong bias. Instead, retaining context in the datasets improves both the mechanics and interpretation of the meta-analysis.2
Example 3: context to reduce hypersensitivity and instability of results
In this example, we contrast the hypersensitivity of opaquely thresholded results to non-physiological features with the relative stability of transparently thresholded ones. Aspects of this have been shown above in the discussion of the NARPS data, but here we demonstrate this point even more directly with a simple case of varying the number of subjects in a study.
Fig. 3 displays the results of a standard one-group analysis with the NARPS data (for Hyp. 2), using a two-sided t-test with cluster-based FWE = 5%. Clusters are outlined in white for visibility. The top row contains the results from analyzing the full set of 47 subjects used for group analysis after processing and quality control. Subsequent rows show results if the group size had been reduced by just one subject (arbitrarily chosen by order of subject ID). Fig. 3A shows the results with opaque thresholding applied—both the coverage and number of clusters change notably from row to row. In one case, removing just one subject decreased the number of clusters by 25%, and in another case it increased the count by 18%. The changes are not simply monotonic or convergent, and one cluster disappears and then reappears (in the left inferior parietal lobule; magenta arrow). A similarity matrix of the clusters (bottom row) shows the amount of variability across an extended set of group sizes (down to 17). Clearly, any interpretation of results with this opaque thresholding will be quite sensitive to group size.
Figure 3.
Each panel shows the same axial slices in MNI template space at z = 21S, 32S, 43S (image left = subject left) from NARPS data, Hyp. 2 and 4. The overlay values are effect estimates, in units of BOLD% signal change per dollar for this gambling task, and statistic values were used for thresholding (voxelwise p = 0.001; cluster-level FWE = 5%). All suprathreshold clusters are highlighted with white outlines, for visibility. The top row shows the full group number of subjects (Nsubj), and subsequent rows show results with 1 subject removed. Changes in cluster count (Nclust) are noted for each row. Changes in cluster results —in terms of both coverage and number—are more apparent in Panel A, where opaque thresholding is used. The changes are not simply convergent or monotonic. The magenta arrow highlights a cluster in the left inferior parietal lobule which disappears and reappears with varying Nsubj. The results with transparent thresholding in Panel B are less sensitive to Nsubj changes and also provide useful context. For example, the region highlighted with the magenta arrow appears to have left-right symmetry in negative BOLD response; this information is missed with opaque thresholding. The bottom of each column shows a similarity matrix for each thresholding style (as in Fig. 2), for an extended set of Nsubj. These reflect the striking sensitivity of opaque thresholding (Dice_all, left) with the more stable transparent thresholding (Corr_coef, right).
Fig. 3B displays the same results using transparent thresholding. The images are considerably more consistent across the small group size changes, reflecting less sensitivity to the arbitrary differences (e.g., quality control criteria, censoring thresholds, subject motion values, etc.). Note that the changes in cluster count are still known and still vary in the same way, but the contextualized maps allow for the reader’s evaluation to remain appropriately consistent and stable. The similarity matrix for these results (Fig. 3D) shows uniformly quite high values even down to the smallest group size. Beyond stability, additional benefits of transparency include having knowledge of the context itself. For example, the area highlighted with the magenta arrow appears to have notable left-right symmetry, which would be unknown in the opaque thresholding case.
Study results are always produced and examined in the context of prior information. The hypersensitivity of opaquely thresholded clusters makes them difficult to rely on for robust comparisons to other papers and even for meaningful evaluation of a study’s hypotheses. Opaque thresholding creates the dilemma of determining, for example, the appropriate final number of subjects from which to obtain the “correct” set of clusters. Moreover, it creates a potential incentive for p-hacking (Wichert et al., 2016): tweaking and selecting such parameters in a way that might reject the null hypothesis or match more closely with prior work. As shown here, transparent thresholding greatly reduces such temptations, as presented results are more stable and can still be discussed easily, even if just below threshold.
Example 4: using figures and context to avoid misinterpretation
As a final example of the importance of keeping context in figures to clearly communicate science results, we look back at one of the most widely known studies in FMRI, the “dead salmon study” (Bennett et al., 2009). This study had a single, simple message: when performing massively univariate voxelwise analyses in brain studies, one should adjust for multiple comparisons in some way, such as applying familywise error (FWE) or false discovery rate (FDR) adjustment. This message is clearly stated in the title of the paper, and it is repeatedly restated throughout the abstract and main text. However, the work has still been consistently misreferenced as showing that FMRI is unreliably susceptible to false results (predominantly outside the field) and even frequently misquoted within the field itself.
The part of the paper that unfortunately leaves room for misinterpretation is its lone figure. In practice, figures often leave stronger impressions of results with readers than text. The famous image is reproduced here in Fig. 4A, along with relevant descriptive information summarized from the original caption. The image shows opaquely thresholded results3 before adjusting for multiple comparisons, but not after doing so. The authors did perform multiple comparisons adjustment—indeed, that is the analysis step they are promoting—but they simply stated its outcome of “no clusters” in the text only. This latter part appears to be ignored by readers relatively frequently, leaving a false impression from the lone, pre-adjustment figure.
Figure 4.
Panel A shows the lone figure from the famous “dead salmon study” (Bennett et al., 2009; with permission of the authors). The figure is opaquely thresholded and only shows results before the recommended multiple comparisons adjustment; by not including the after image, many readers have misinterpreted the overall study message, even though it is clearly repeated throughout the text. Panel B shows a simple improvement to make the figure’s message clearer and reduce the likelihood of misinterpretation, by including both before- and after-adjustment images. Panel C shows the new validation salmon in the same manner as Panel B, replicating the original results with opaque thresholding. Panel D shows how more complete context can be added to further reduce risks of misinterpretation by thresholding transparently (suprathreshold regions outlined in green), displaying the effect estimate in units of BOLD % signal change as overlay colors, and even showing results outside the subject anatomy. This extra information would provide valuable evidence that any cluster that might survive here—which is possible even when including multiple comparisons adjustment—is likely noise due to the background pattern, high noise floor and likely low effect estimate value.
One clarifying step would be to include both the “before” and “after” cases in the image, such as in Fig. 4B. The figure’s primary message is now more clearly in line with the methodology and the authors’ intended purpose.
But the results reporting could still be improved further by using transparent thresholding. This is shown in Fig. 4D for a different fish,4 scanned during a standard flashing checkerboard stimulus paradigm (see Supplements for details). We again include both the “before” and “after” cases of statistical adjustment, which match those of the original dead salmon (Fig. 4C shows the new validation salmon results with opaque thresholding). By including subthreshold modeling results in the images in Fig. 4D, it is immediately apparent that the overall pattern is obviously quite noisy and unrelated to structure. This context is useful because it is quite possible that a noisy cluster could still survive all of the statistical adjustment and thresholding processes. Seeing the noisy subthreshold context would then provide useful (if not necessary) evidence that any such cluster was likely to be noise-related rather than a function-related cluster (similar to Image 4 in Fig. 1D).
In addition to transparent thresholding, Fig. 4D includes additional useful features. First, It includes data from the whole field of view, so one can judge the pattern of results inside the brain with respect to the “noise floor” background (Taylor et al., 2023). Second, the overlay is the effect estimate dataset (BOLD percent signal change), rather than just the statistics, which are still used for thresholding (Chen et al., 2017). In this way, the reader can appreciate that the effect estimate values are quite low within the subject, even though the statistic values there are relatively high in many places. These features apply even beyond applying transparent thresholds to “before” and “after” images, and should be used whenever possible. For example, seeing the very small effect size of a cluster that happened to survive would provide further useful evidence as to its true, noisy nature; this would apply to the cluster in the lower part of the image in Fig. 4C, if it had been just slightly larger. Thresholding transparently, displaying the effect estimate, and including the background FOV provide useful contextual information to solidify and clarify the interpretation.
DISCUSSION
Neuroimaging authors need to supply enough details for the analysis to be understood and replicable (e.g., Maumet et al., 2016; Nichols et al., 2017). They also need to balance presenting “digestible” results with retaining meaningful information content (Chen, Taylor, et al. 2022). The question of how much data to present and in what form is important, and we can turn again to Anscombe (1973), who commented on the value of graphs to provide useful contextual information beyond just summary statistics:
Graphs can have various purposes, such as: (i) to help us perceive and appreciate some broad features of the data, (ii) to let us look behind those broad features and see what else is there. Most kinds of statistical calculation rest on assumptions about the behavior of the data. Those assumptions may be false, and then the calculations may be misleading. We ought always to try to check whether the assumptions are reasonably correct; and if they are wrong we ought to be able to perceive in what ways they are wrong. Graphs are very valuable for these purposes.
In a near-exact parallel, we believe that transparent thresholding provides important value for understanding and evaluating results in neuroimaging. Many assumptions associated with opaque thresholding are inconsistent with the data, thereby increasing odds of misinterpretation and biases. In contrast, retaining context provides the benefits of both “appreciating broad features of the data” and letting readers “see what else is there.”
The four examples presented here illustrate these points and other benefits of retaining context in neuroimaging results. The primary benefit is to provide a clearer, deeper, and more accurate understanding of the data, for both authors and readers. This is the ultimate goal of a scientific experiment. Results reporting should prioritize having comprehensive information over artificial dichotomization. Removing context with opaque thresholding inserts a large amount of ambiguity and bias into results, tilting the scales towards misinterpretation.
The benefits of transparent thresholding—and costs of opaque thresholding—apply across modality and species. Whether analyzing FMRI, diffusion weighted imaging (DWI) or PET (Bishay et al., 2024) data in humans, macaques (Russ and Leopold, 2015) or rodents, showing subthreshold context improves interpretability. It can also be applied equivalently to voxelwise or ROI-based analyses (Chen et al., 2020; Taylor et al., 2023), as well as to both volumetric and surface-based analyses (Sava-Segal et al., 2024; Freund et al., 2025).
The issue of thresholding images touches at the core of the scientific endeavor, particularly for neuroimaging. An individual study rarely provides a definitive answer. Instead, empirical science is an iterative process that builds upon cumulative evidence from multiple studies. Using opaque thresholds treats an individual study as a standalone decision-making tool, which misrepresents the essence of scientific inquiry. In contrast, transparent thresholding assists that process by presenting results as widely informative evidence rather than as a narrowly defined “answer”. It is also directly in line with larger mathematical recommendations about improving the use and interpretation of statistics and p-values across broad scientific disciplines, which include having an “interpretation of results in context” and making “complete reporting” (Wasserstein and Lazar, 2016). By focusing on the broader context and appropriately including results that simply have higher uncertainty, researchers can better align with the collaborative and progressive nature of empirical investigation.
As demonstrated below in Figs 5–7, there are now a large number of software packages that make a similar form of transparent thresholding available to researchers. These encompass implementations for volumetric, surface- and region-based studies. Several of these were added or streamlined as part of this work. This methodological accessibility is an important practical step for the neuroimaging field (and it will likely continue to grow), allowing researchers to easily adopt the same form of transparent thresholding for their image visualization.
Figure 5.
Example images of transparent thresholding from various software implementations (and see Figs. 6 and 7 for more examples). Descriptions of the data and software usage are provided in the Supplements.
Figure 7.
Example images of transparent thresholding from various software implementations (and see Figs. 5 and 6 for more examples). Descriptions of the data and software usage are provided in the Supplements.
This approach is a useful complement to data sharing (such as via Neurovault (Gorgolewski et al., 2015), OSF, or another resource), which itself is beneficial to the field, but applying transparent thresholding is distinct and important on its own. Figures have a powerful and primary role in interpreting results, and therefore they should be as informative as possible within the presented publication in order to facilitate accurate evaluation. The examples presented above have demonstrated this. Showing opaque figures-as-usual and relying on readers to download and visualize the data again separately does not accomplish this effectively.
Reducing biases
Standard opaque thresholding, while done with well-meaning intentions, often inserts bias into subsequent analysis and meta-analysis. This practice both harms the ability to accurately assess reproducibility and acts to decrease reproducibility unnecessarily. The analyses of the NARPS data show both of these features, as directly demonstrated in Taylor et al. (2023) and the examples above, as well as indirectly noted in Botvinik-Nezer et al. (2020). As shown above, transparent thresholding showed increased reproducibility where appropriate (i.e., when results agreed but at different strengths) and showed low reproducibility where appropriate (when results disagreed or were essentially null). Thresholding is a processing choice, and these studies together strongly suggest that opaque thresholding is often detrimental to analyses and meta-analyses. Reproducibility has been a long-discussed topic in the field, and retaining context in results and figures is a clear step to help address it.
The adoption of transparent thresholding helps reduce several other biases in neuroimaging reporting, while still preserving the ability to highlight the most significant regions:
Type II errors, which may arise from overly strict multiple comparisons adjustments (Cremers et al., 2017), are reduced because subthreshold regions (particularly those just below the cutoff) are still visible. Such regions can still be assessed in the context of other studies or prior knowledge, without simply being treated (inaccurately) as “no effect”.
As discussed above, transparent thresholding greatly reduces incentives for p-hacking (Wicherts et al., 2016) to be able to report results in predetermined regions, or relatedly (and problematically) to “spin” results (Boutron and Ravaud, 2018).
Publication bias or the “file drawer problem” exists in neuroimaging, where findings with weak or limited statistical evidence are typically not reported. The result is that many meta-analyses misrepresent assessments (Jennings and Van Horn, 2012; Ioannidis et al., 2014). Transparent thresholding offers a way that findings with weak statistical evidence can be assessed and reported more informatively. That is, having multiple related studies with visible sub-threshold responses, represents critical information to include in meta-analyses. Additionally, displaying subthreshold regions meaningfully, even when some locations have suprathreshold ones, can help reduce publication bias.
Statistical thresholding itself is essentially a form of selective reporting, which was selected as the top factor contributing to reproducibility problems according to a recent Nature journal “reproducibility survey” of researchers (NPG, 2018). Consequently, overly stringent statistical thresholds and opaque thresholding may hinder rather than help the scientific process, as they can obscure meaningful patterns and impede the synthesis of findings across studies.
In addition to reducing those biases, the improved stability of transparent thresholding (see Ex. 3, above) greatly benefits studies of difficult to scan populations. While there has been a movement to increase group sizes, animal imaging studies (e.g., of nonhuman primates and rodents) remain small; in many cases, these have less than 10 subjects (Mandino et al., 2020), though these often have multiple sessions. In human clinical studies, group sizes also tend to be much smaller than standard research studies (Szucs and Ioannidis, 2020). The accrual of clinical participants is often limited by practical considerations of availability (e.g., rare diseases), cost (travel), complexity (health considerations and medical monitoring), and study length (drug trials). In these cases, opaque thresholding practices often result in few (or even no) regions above the statistical cut-off, potentially resulting in an inability to publish (i.e., publication bias) or an incentive to p-hack. Thus, the results that can be reported with typical thresholding are generally limited and tightly bound to the conditions of the sample itself, making them difficult to interpret and reproduce. (Data sharing is also typically more challenging for clinical studies.) While increased stability of results is not a substitute for having an adequately powered sample, transparent thresholding enables these smaller studies, acquired under difficult conditions, to contribute to the literature in a more meaningful way rather than adding to the “file drawer problem.” With transparent thresholding, the regions with larger uncertainty are clearly viewed as such, rather than being hidden or artificially pushed above thresholding by statistical maneuvering. At the same time, transparent thresholding helps reduce the likelihood of false negatives, which can be particularly important in clinical studies. It also parallels the trend in the clinical literature of moving away from the thresholding dichotomy, such as viewing results at multiple thresholds instead of a single arbitrary threshold (Voets, et al., 2025).
Forward looking goals
As another important focus on figures, we note that papers are increasingly extracted and processed by algorithms. NeuroSynth (Yarkoni et al., 2011) was an early example of a tool that automatically parsed publication text and aggregated information together. Today, large language models (LLMs) and multimodal LLMs (MLLMs) are increasingly applied to databases and libraries—some even have a specific focus on medical imaging—creating summary tools from both text and figures (Bhayana, 2024; Bzdok et al., 2024). Just as retaining accurate descriptions in text improves the results of these tools, so does (or surely will) having more informative figures.
The adoption of transparent thresholding facilitates and complements meta-analyses that include more data in neuroimaging. For example, auxiliary scatterplots of statistics and effects in MVPA results have revealed interesting spatial patterns (see Fig 4 of Visconti di Oleggio Castello et al., (2017)). If whole brain results are shared and used for cross-study comparisons, it is more consistent to have the results shown across the whole brain in the first place. It would be confusing and inconsistent to see meta-analyses point out differences in regions that were hidden in the initial papers. Moreover, proposed multiverse approaches (e.g., Dafflon et al., 2022; Lefort-Besnard et al., 2024) aim to combine results from multiple processing pipelines or statistical methods for a given dataset, similar to meta-analyses. Transparent thresholding facilitates reporting such “doubly probabilistic” results across the whole brain, where one would expect different sub-tests to have meaningful evidence to be reported.
Another way to improve both cross-study meta-analyses and within-study interpretations is to include effect estimates in the results, rather than only showing statistics (Halsey et al., 2015). Many neuroimaging modalities have physical units, such as DWI, and some, like FMRI, can be scaled to have meaningful units (Chen et al., 2017; Flournoy et al., 2020). These provide separate information about the data and its modeling, including the practical significance of results. In most areas of science, it would be inconceivable not to include effect estimates, as they form the basis of analysis and interpretation. Leaving these measures out wastes information (Chen et al., 2022) and increases the ambiguity of results. For example, showing effect estimates provides useful evidence about whether suprathreshold locations are more likely true or false positives (as noted in Ex. 4). They also enable more meaningful comparisons in meta-analyses than ones based on statistics alone (e.g., Maumet and Nichols, 2016).
Finally, transparent thresholding directly benefits quality control (QC) efforts during both data processing and results presentation. Several examples of this were provided in Reynolds et al. (2023), where artifacts in the acquired EPI time series would likely have gone unnoticed without using transparent thresholding in the creation of seed-based correlation QC images. In both Ex. 1 and 4 here, subthreshold patterns helped distinguish when any results might likely be due to noise or artifact, reducing the risk of false positives. In these and other cases, wider results reporting and retention of context provide greater confidence and clearer interpretability of figures.
Addressing comments, concerns and questions about transparent thresholding
While some in the neuroimaging field have been enthusiastic about using transparent thresholding to present more informative and less ambiguous results, others have been skeptical or raised critiques. Here we summarize some of the latter and address these points.
Thresholding at exactly p=0.001 and FWE = 5% is rigorous, and reporting any other results will harm reproducibility with false positives. Firstly, the proposed transparent thresholding still highlights the exact same regions above a given threshold (with full opacity and outlining). Secondly, threshold values themselves are typically set by convention, with various round numbers argued for at various points in history. Fisher initially wrote about p = 0.05 as useful, primarily because it corresponds to a round, two-tailed test value Z ≈ 2, but he also used other values such as p = 0.01 (Fisher, 1925). Many of his contemporary statisticians viewed such threshold choices as arbitrary and obfuscating (see Kennedy-Shaffer, 2019). More recently, one group of statisticians pushed to lower the canonical p-threshold to a still different value 0.005 (Benjamin et al., 2018), while another proposed rejecting p-value thresholds altogether (Amrhein & McShane, 2019). Even for clinical FMRI, there is no standard thresholding practice (Voets et al. 2025). This snapshot of a hundred-year-old debate alone shows there is no single, canonical threshold between “significance” and “insignificance,” and many modern statisticians view p-values as an unreliable focus (Halsey et al., 2015). Entirely hiding a cluster that has FWE= 5.01% is arbitrary and unscientific—such a practice itself actually harms reproducibility in the long term.
A given threshold may be arbitrary, but if everyone uses the same value, then results will still be on equal footing for comparisons. FMRI and many other kinds of neuroimaging data are noisy, with noise profiles changing across the field of view and with the scanner used. The underlying biological responses are continuous with varying magnitudes. In practice, no threshold will be consistent across studies, e.g. due to differing sample size and power, or noise characteristics and variability, or evolving acquisition strategies. Instead, applying a strict threshold for a continuous response variable also greatly increases sensitivity to non-physiological differences across studies, such as scanner type, field strength, number of subjects (see Ex. 3 above), voxel size, trial number/length, field inhomogeneities, modeling methods, etc. Opaque thresholding will strongly bias comparisons and meta-analyses towards irreproducibility and non-replicability (see the NARPS data discussion, above).
Science is about “storytelling” and transparent thresholding complicates the story of results by showing more things. Few stories of note have only main actors and no supporting cast and context—Romeo and Juliet (Shakespeare, 1597) would be a poor play if it contained no other characters. Perhaps fully unthresholded results are overly complicated to interpret for most, but transparently thresholded ones primarily add just near-significant regions and larger context. If there are not many near-threshold results, then the story stays terse. Alternatively, if there are many near-threshold results, then that is part of the story. In either scenario the reader learns from seeing the additional context. It will both facilitate storytelling and clarify a narrative, such as by suggesting which network a cluster belongs to. In Ex. 1, the additional context from transparent thresholding presents compelling evidence to challenge the narrative of strong laterality that opaque thresholding would have produced. As Einstein (probably) noted about science, “Everything should be made as simple as possible, but not simpler” (Shapiro, 2006). Storytelling should not be used as an excuse to sacrifice meaningful information.
We applied transparent thresholding, and now see too many regions to discuss—the paper will be too long. It does not seem necessary to have a detailed discussion about every single region with visible, subthreshold results. In many papers, researchers do not even write about each suprathreshold region. It would be logical to highlight any regions of particular interest, as determined from either prior research or background knowledge, and then the rest can remain as observable context and/or for potential relevance to future studies. Transparent thresholding also allows for further discussion of regions that show no visible response with transparent thresholding, especially if activity had been hypothesized there. In total, this should not make discussions more burdensome but instead more informative and clearer. It should also facilitate connecting the results to those of existing literature, functional networks and prior domain knowledge, by providing more globally informative results. (Additionally, if authors are concerned that a figure will not get published with transparent thresholding because there is a mess of blobs and artifacts near-threshold, this is a problem with their data and not something to be swept under the rug with opaque thresholding. Hiding issues in data to facilitate publication is generally considered poor scientific practice.)
We checked out the results at multiple thresholds within our group, so we feel confident about publishing with standard thresholding. Readers engage with scientific articles critically, a procedure that is facilitated by seeing more of the underlying evidence for themselves as they read. If transparent thresholding reveals no near-threshold results, then that simply reinforces the authors’ interpretation—nothing is lost. If transparency reveals some additional locations of interest, then readers will be aware even if it does not figure strongly into the authors’ interpretation. In fact, transparent thresholding should be viewed as a helpful tool to convince readers. For example, if the authors have used multiple thresholds and confirmed for themselves that their suprathreshold region is not just an extension of a near-threshold blob from the CSF, then they only strengthen their argument by including this information in their figures. In science, it benefits both the authors and readers to present the full picture. Moreover, as shown in Ex. 3, transparent thresholding will reduce the hypersensitivity of near-threshold results to arbitrary parameters (like number of subjects) and the incentives for p-hacking around varied thresholds.
- We will upload the unthresholded results to a public repository, so we will publish opaque images and people can explore the data themselves later. Making full results public is great5 and certainly enables meta-analyses, but separating the more complete results from the paper greatly diminishes the ability for critical understanding by the reader. It also places a burden of time and effort that not every reader will go through.6 A study should provide strong evidence for its interpretations, and transparent thresholding does a better job of this than opaque thresholding. Two well-known aphorisms apply:
- A picture is worth a thousand words. Figures are likely the most important messengers in a scientific study. We provided multiple examples here where there have been critical misinterpretations of written results, simply because figures were either lacking important information or were themselves missing. A summary figure can end up in talks and general discourse in ways that separate it from the complete story of the actual data in the repository and can even distort the intended message of the authors.
- First impressions are the most important. Reanalysis of public data might lead to a different interpretation from an initial study, but there will be a long lag before that update can enter the scientific conversation. Consider the widely cited NARPS study, which created a strong impression of poor FMRI reproducibility that is still echoed today. The follow-up analysis of the public data by Taylor et al. (2023) showing strong evidence for a different message was published three years later, in a separate journal, and with much less impact.
I’m a clinician. I need to know definite regions. You are the expert for your work, so you should be the decision maker from a reasonable set of evidence. Transparent thresholding presents meaningful results for you to interpret and from which to make informed judgments. In contrast, starting with opaque thresholding presents a predetermined decision based on an arbitrary cut-off and removes potentially useful information from your consideration—in short, it puts the scalpel in the hands of an academic. It is more scientific (and likely better for clinical outcomes) to share contextualized evidence for clinicians to interpret for their purposes. Some diagnostic specialties even use fully unthresholded maps regularly, but when greater digestibility is required, transparent thresholding helps reduce the risk of false negatives and misinterpretation. Consider the divergent laterality findings in Ex. 1. In special cases that a binarized image is needed (e.g., in surgery), then that can still be derived from a transparently thresholded one and likely with more confidence about the localization. Voets et al. (2025) discuss further issues for clinical applications.
I write for a non-technical audience, and I need to show a direct story for those outside the field. Even for non-neuroimaging readers, oversimplifying the results is problematic. Experience with the dead salmon (Ex. 4, above) and other studies has shown that. That manuscript made a clear and valid point, but sharing only the one opaquely thresholded image has resulted in repeated misinterpretations of their core message. With transparent thresholding and the other features discussed above, we posit that the dead salmon study would have been less easy to misinterpret. While its coverage in non-technical media might not have become as expansive, those it did reach would have gained a better understanding of science and the core issues. (And it would still have received wide attention within the field, because it is a great demonstration of a serious issue.) Scientific understanding should be as accurate as possible, for both those in the field and those outside of it.
Reviewers have criticized showing subthreshold results, complained about seeing more complicated spatial patterns, and do not like the lines around regions—this makes me hesitant to try to publish with this approach. While many researchers have successfully adopted transparent thresholding in the publications (see a partial list in Table S1), it does take time for new ideas to become accepted and commonplace. We hope that articles like this, which show transparent thresholding’s many benefits and demonstrate the significant problems with opaque thresholding, will help this way of displaying results become more commonplace. If spatial patterns are complicated because many results are slightly subthreshold, then it is likely even more important to use transparency, to reduce the hypersensitivity to non-physiological features and chance of misinterpretation, as shown above in Ex. 3. While transparent thresholding is not currently normative, it is plausible that manuscripts with opaque thresholding should or will eventually themselves be critiqued over what is not shown to readers. We hope that the examples and points raised in this paper, as well as those in Allen et al. (2012), Chen et al. (2022), Taylor et al. (2023), and Sundermann et al. (2024), can provide convincing rationales to the reviewers for this approach.
It is too difficult to implement this visualization. Transparent thresholding in data visualization is now available in a wide number of publicly available neuroimaging software packages (as well as in separately programmed implementations): in the original Trends-Matlab toolbox (https://trendscenter.org/x/datavis; Allen et al., 2012) and the related GIFT (http://trendscenter.org/software/gift); in AFNI (Cox, 1996) and preliminarily in the surface-based visualization of SUMA (Saad et al., 2004; Saad and Reynolds, 2012); in BrainVoyager (Goebel, 2012); in FSL’s FSLeyes (McCarthy, 2024; Smith et al., 2004); in NiiVue (Hanayik et al., 2023); in RMINC (Lerch et al., 2017) and the related MRIcrotome (https://github.com/Mouse-Imaging-Centre/MRIcrotome); in CIVET (Ad-Dab’bagh et al., 2006) and the related minc-toolkit-v2 (https://github.com/BIC-MNI/minc-toolkit-v2); in Nilearn (Nilearn contributors, 2025); and in bidspm (https://github.com/cpp-lln-lab/bidspm). Representative images are shown in Figs. 5–7, which encompass volumetric, surface-based and ROI-based cases. See Table S1 in the Supplements for further examples.7 We hope that increased use of transparent thresholding will see its implementations spread further.
CONCLUSION
Choosing how to threshold results is an important processing choice in FMRI and more widely across neuroimaging, including in many clinical applications. Transparent thresholding highlights the same strong regions as standard opaque thresholding but also retains brainwide context that is important for both authors and readers to see. This approach enhances understanding and helps to avoid misinterpretation in figures, which are key to presenting study results. It also provides a better framework for accurately comparing datasets and evaluating reproducibility. In contrast, opaque thresholding introduces strong biases and hypersensitivity to non-physiological features, harming within-study evaluation and cross-study reproducibility. Transparent thresholding is straightforward and easily implementable. In fact, it has already been included in a large number of packages and available scripts, providing an accessible and common way for neuroimagers display data. We hope that researchers in the field will move toward adopting this strategy when presenting results.
Supplementary Material
Figure 6.
Example images of transparent thresholding from various software implementations (and see Figs. 5 and 7 for more examples). Descriptions of the data and software usage are provided in the Supplements.
ACKNOWLEDGMENTS
The present work was completed as part of several authors’ official duties as Government employees. The views expressed do not necessarily reflect the views of the National Institutes of Health (NIH), the Department of Health and Human Services (HHS), or the United States Government. We appreciate the authors of the original “dead salmon” study (Bennett et al., 2009) giving their permission for reproduction of the study’s figure in this text. We also appreciate the authors of the original NARPS project (Botvinik-Nezer et al., 2020) for making public the results that had been submitted by all participating teams; this has been an important dataset for the neuroimaging community to consider and keep reanalyzing for new perspectives. PAT, DRG, PDL, JKR, RCR, and GC were supported by the NIMH Intramural Research Program (ZICMH002888) of the NIH/HHS, USA. AM was supported by a grant from NIH/NIA (R01AG083919). BER was supported by grants from NIMH (R01MH111439, RF1MH117040, R01MH124045, and P50MH109429) and NINDS (R01NS109498). BT et HA benefited from state aid managed by the Agence Nationale de la Recherche under the France 2030 program (reference ANR-22-PESN-0012), and have also received funding from the European Union’s Horizon 2020 Framework Programme for Research and Innovation under the Specific Grant Agreement HORIZON-INFRA-2022-SERV-B-01. CCG was supported by the Basque Government BERC 2022-2025 program, the Spanish State Research Agency through BCBL Severo Ochoa excellence accreditation CEX2020-001010/AEI/10.13039/501100011033, and the project PID2023-149410OB-100 funded by MCIU/AEI/10.13039/501100011033/FEDER, EU. CM has funding from consulting and lecturing for Siemens Healthineers on behalf of Evangelisches Krankenhaus Oldenburg. CP is funded by the Novo Nordisk Foundation (grant NNF20OC0063277). CR was supported by grants from the NIBIB/NIMH (RF1-MH133701) and NIDCD (P50-DC014664). DAH and JGC were supported by NIMH Intramural Research Program (ZIAMH002783); VR was supported by the NIMH Intramural Research Program (ZICMH002884); and PAB was supported by both of these grants. DAL was supported by NIMH Intramural Research Program (1ZICMH002899). EAGV was supported by UNAM PAPIIT projects IN213924 and IA201622, and is part of the CONAHCYT (3256252/629578) and DGAPA (781759) postdoctoral projects. GAD was supported by the Douglas Research Centre and funds from McGill University’s Healthy Brains for Healthy Lives initiative (a Canada First Research Excellence Fund Initiative). JWE was supported by NIMH Intramural Research Program (ZIAMH002857). JPL was supported by The Wellcome Centre for Integrative Neuroimaging, which is supported by core funding from the Wellcome Trust (203139/Z/16/Z and 203139/A/16/Z). LDR was supported by the NIH’s Office of Intramural Training and Education (OITE) through an Intramural Research Training Award (IRTA), and was also funded by the National Institute of Biomedical and Bioengineering Intramural program: Processing and Analysis of Quantitative Diffusion MRI Data, ZIA EB000088. LP was supported by a grant from NIMH (R01MH071589). MB was supported by a Flagship ERA-NET grant SoundSight (FRS-FNRS PINT-MULTI R.8008.19). MC was supported by the Canadian Institutes for Health Research, National Sciences Research Council of Canada, and HBHL, while also receiving salary support from Fonds de recherche du Québec - Santé. MGB was supported by grants from NIH NICHD (R03HD113915, R21HD108587). OFG and RG have financial interests tied to Brain Innovation company. SM has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No 101109770. ST was supported by UCSF RAP REAC Award 7504825. TEN was supported by NIH grants (R01DA048993, R01MH096906, R01EB026859, U19AG073585) of the NIH/HHS, USA. TH’s primary contributions were made while employed by FMRIB/Oxford University, and were supported by the Wellcome Centre for Integrative Neuroimaging (Oxford, UK) via funding from the Wellcome Trust (203139/Z/16/Z and 203139/A/16/Z) and the NIHR Oxford Health Biomedical Research Centre (NIHR203316); the views expressed are those of the author(s) and not necessarily those of the NIHR or the Department of Health and Social Care. VDC was supported by grants from NIH (R01MH123610) and NSF (2112455). YH was supported by a grant from NIH-NIBIB (P41EB019936). This work utilized the computational resources of the NIH HPC Biowulf cluster (https://hpc.nih.gov).
Footnotes
It also shows the importance of including at least some images of the underlying study datasets themselves in comparisons report. No figures of individual teams’ results were shown in the NARPS paper or its supplements, though the datasets were made publicly available.
We note that any formal meta-analysis here would be limited, since teams were only required to upload statistics data. This follows the unfortunately common practice that effect estimate information is rarely reported, even though it has many uses (Chen et al., 2017). With more detailed effect size data, a more accurate meta-analysis could have been conducted, such as with hierarchical modeling.
This study pre-dates the publication of the transparent thresholding idea by Allen et al. (2012).
The dataset from the original dead salmon study has been lost “upstream,” but this presents the opportunity to verify that the original findings replicate in another dead fish.
When possible—not all datasets can be shared, particularly in clinical studies.
An earlier version of this draft purposely omitted images of the original Anscombe Quartet for readers to look up online, as a simplified example of the burden caused by separating relevant figures from the main text. However, even this simple search (no downloads or software needed) was deemed too frustrating by coauthors, so the images were included here.
Interestingly, the majority of these studies that used transparent thresholding also presented effect estimates as overlays, providing further useful information as recommended here (see Ex. 4).
Contributor Information
Paul A. Taylor, Scientific and Statistical Computing Core, NIMH, NIH, Bethesda, MD, USA
Himanshu Aggarwal, Inria, CEA, Université Paris-Saclay, Palaiseau, 91120, France.
Peter A. Bandettini, Section on Functional Imaging Methods, NIMH, NIH, Bethesda, MD, USA
Marco Barilari, Crossmodal Perception and Plasticity Lab, Institute of Neuroscience (IoNS) and Institute of Research in Psychology (IPSY), Université Catholique de Louvain, 1348 Louvain-la-Neuve, Belgium.
Molly G. Bright, Department of Physical Therapy and Human Movement Sciences, Feinberg School of Medicine, Northwestern University, Chicago, IL, USA; Department of Biomedical Engineering, McCormick School of Engineering and Applied Sciences, Northwestern University, Evanston, IL, USA
César Caballero-Gaudes, Basque Center on Cognition, Brain and Language, San Sebastian-Donostia, Spain; Ikerbasque, Basque Foundation for Science, Bilbao, Spain.
Vince D. Calhoun, Tri-institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State, Georgia Tech, Emory, Atlanta, GA, USA
Mallar Chakravarty, Cerebral Imaging Centre, Douglas Mental Health University Institute, Montreal, QC, Canada; Department of Psychiatry, McGill University, Montreal, QC, Canada; Department of Biomedical Engineering, McGill University, Montreal, QC, Canada.
Gabriel A. Devenyi, Cerebral Imaging Centre, Douglas Mental Health University Institute, Montreal, QC, Canada; Department of Psychiatry, McGill University, Montreal, QC, Canada
Jennifer W. Evans, Experimental Therapeutics and Pathophysiology Branch, NIMH, NIH, Bethesda, MD, USA
Eduardo A. Garza-Villarreal, Department of Behavioral and Cognitive Neurobiology, Institute of Neurobiology, Universidad Nacional Autónoma de México campus Juriquilla, Querétaro, Mexico
Jalil Rasgado-Toledo, Department of Behavioral and Cognitive Neurobiology, Institute of Neurobiology, Universidad Nacional Autónoma de México campus Juriquilla, Querétaro, Mexico.
Rémi Gau, Inria, CEAUniversité Paris-Saclay, Palaiseau, 91120, France.
Daniel R. Glen, Scientific and Statistical Computing Core, NIMH, NIH, Bethesda, MD, USA
Rainer Goebel, Department of Cognitive Neuroscience, FPN, Maastricht University, Maastricht, NL; Brain Innovation, Maastricht, NL.
Javier Gonzalez-Castillo, Section on Functional Imaging Methods, NIMH, NIH, Bethesda, MD, USA; Basque Center on Cognition, Brain and Language, San Sebastian-Donostia, Spain.
Omer Faruk Gulban, Department of Cognitive Neuroscience, FPN, Maastricht University, Maastricht, NL; Brain Innovation, Maastricht, NL.
Yaroslav Halchenko, Department of Psychological and Brain Sciences, Dartmouth College, Hanover, NH, USA.
Daniel A. Handwerker, Section on Functional Imaging Methods, NIMH, NIH, Bethesda, MD, USA
Taylor Hanayik, Wellcome Centre for Integrative Neuroimaging, FMRIB, University of Oxford, Oxford, UK.
Peter D. Lauren, Scientific and Statistical Computing Core, NIMH, NIH, Bethesda, MD, USA
David A. Leopold, Systems Neurodevelopment Laboratory, NIMH, NIH, Bethesda, MD, USA
Jason P. Lerch, Wellcome Centre for Integrative Neuroimaging, Nuffield Department of Clinical Neurosciences, University of Oxford, Oxford, UK; Department of Medical Biophysics, University of Toronto, Toronto, ON, Canada
Christian Mathys, Institute of Radiology and Neuroradiology, Evangelisches Krankenhaus Oldenburg, Universitätsmedizin Oldenburg, Oldenburg, Germany; Research Center Neurosensory Science, Carl von Ossietzky Universität Oldenburg, Oldenburg, Germany.
Paul McCarthy, Wellcome Centre for Integrative Neuroimaging, FMRIB, Nuffield Department of Clinical Neurosciences, University of Oxford, UK.
Anke McLeod, Department of Radiology and Nuclear Medicine, University Hospital Magdeburg, Otto-Von-Guericke University, Magdeburg, Germany.
Amanda Mejia, Department of Statistics, Indiana University, Bloomington, USA.
Stefano Moia, Department of Cognitive Neuroscience, FPN, Maastricht University, Maastricht, NL.
Thomas E. Nichols, Big Data Institute, Li Ka Shing Centre for Health Information and Discovery, Nuffield Department of Population Health, University of Oxford, UK; Wellcome Centre for Integrative Neuroimaging, FMRIB, Nuffield Department of Clinical Neurosciences, University of Oxford, UK
Cyril Pernet, Neurobiology Research Unit, Rigshospitalet, Denmark.
Luiz Pessoa, Department of Psychology, University of Maryland, College Park, MD, USA.
Bettina Pfleiderer, Clinic of Radiology, Medical Faculty, University of Münster, Münster, Germany.
Justin K. Rajendra, Scientific and Statistical Computing Core, NIMH, NIH, Bethesda, MD, USA
Laura D. Reyes, Laboratory on Quantitative Medical Imaging, National Institute of Biomedical Imaging and Bioengineering, NIH, MD, USA
Richard C. Reynolds, Scientific and Statistical Computing Core, NIMH, NIH, Bethesda, MD, USA
Vinai Roopchansingh, Functional MRI Facility, NIMH, NIH, Bethesda, MD, USA.
Chris Rorden, McCausland Center for Brain Imaging, Department of Psychology, University of South Carolina, Columbia, SC 29208, USA.
Brian E. Russ, Center for Biomedical Imaging and Neuromodulation, Nathan Kline Institute, 140 Old Orangeburg Road, Orangeburg, NY 10962, USA; Nash Family Department of Neuroscience and Friedman Brain Institute, Icahn School of Medicine at Mount Sinai, One Gustave L. Levy Place, New York, NY 10029, USA; Department of Psychiatry, New York University at Langone, One, 8, Park Ave, New York, NY 10016, USA
Benedikt Sundermann, Institute of Radiology and Neuroradiology, Evangelisches Krankenhaus Oldenburg, Universitätsmedizin Oldenburg, Oldenburg, Germany; Research Center Neurosensory Science, Carl von Ossietzky Universität Oldenburg, Oldenburg, Germany; Clinic of Radiology, Medical Faculty, University of Münster, Münster, Germany.
Bertrand Thirion, Université Paris Saclay, France..
Salvatore Torrisi, University of California, San Francisco, Department of Radiology & Biomedical Imaging, SF CA; San Francisco Veteran Affairs Health Care System, SF CA.
Gang Chen, Scientific and Statistical Computing Core, NIMH, NIH, Bethesda, MD, USA.
REFERENCES
- Ad-Dab’bagh Y, Einarson D, Lyttelton O, Muehlboeck J-S, Mok K, Ivanov O, Vincent RD, Lepage C, Lerch J, Fombonne E, Evans AC (2006). The CIVET image-processing environment: A fully automated comprehensive pipeline for anatomical neuroimaging research. Proc. OHBM-2006. http://www.bic.mni.mcgill.ca/users/yaddab/Yasser-HBM2006-Poster.pdf [Google Scholar]
- Allen EA, Erhardt EB, Calhoun VD (2012). Data Visualization in the Neurosciences: overcoming the Curse of Dimensionality. Neuron 74:603–608. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Amrhein V, Greenland S, McShane B (2019). Scientists rise up against statistical significance. Nature 567:305–307 [DOI] [PubMed] [Google Scholar]
- Anscombe FJ (1973). Graphs in Statistical Analysis. The American Statistician 27(1):17–21 [Google Scholar]
- Bacchetti P (2013). Small sample size is not the real problem. Nat Rev Neuroscience 14, 585. [DOI] [PubMed] [Google Scholar]
- Benjamin DJ, Berger JO, Johannesson M, Nosek BA, Wagenmakers EJ, Berk R, Bollen KA, Brembs B, Brown L, Camerer C, Cesarini D, Chambers CD, Clyde M, Cook TD, De Boeck P, Dienes Z, Dreber A, Easwaran K, Efferson C, Fehr E, Fidler F, Field AP, Forster M, George EI, Gonzalez R, Goodman S, Green E, Green DP, Greenwald AG, Hadfield JD, Hedges LV, Held L, Hua Ho T, Hoijtink H, Hruschka DJ, Imai K, Imbens G, Ioannidis JPA, Jeon M, Jones JH, Kirchler M, Laibson D, List J, Little R, Lupia A, Machery E, Maxwell SE, McCarthy M, Moore DA, Morgan SL, Munafó M, Nakagawa S, Nyhan B, Parker TH, Pericchi L, Perugini M, Rouder J, Rousseau J, Savalei V, Schönbrodt FD, Sellke T, Sinclair B, Tingley D, Van Zandt T, Vazire S, Watts DJ, Winship C, Wolpert RL, Xie Y, Young C, Zinman J, Johnson VE (2018). Redefine statistical significance. Nat Hum Behav 2(1):6–10. [DOI] [PubMed] [Google Scholar]
- Bennett CM, Baird AA, Miller MB, Wolford GL (2009). Neural correlates of interspecies perspective taking in the post-mortem Atlantic salmon: an argument for proper multiple comparisons correction. J Serendipitous Unexpected Results 1:1–5. [Google Scholar]
- Bhayana R (2024). Chatbots and Large Language Models in Radiology: A Practical Primer for Clinical and Research Applications. Radiology 310(1):e232756. [DOI] [PubMed] [Google Scholar]
- Bishay S, Robb WH, Schwartz TM, Smith DS, Lee LH, Lynn CJ, Clark TL, Jefferson AL, Warner JL, Rosenthal EL, Murphy BA, Hohman TJ, Koran MEI (2024). Frontal and anterior temporal hypometabolism post chemoradiation in head and neck cancer: A real-world PET study. J Neuroimaging 34(2):211–216. [DOI] [PubMed] [Google Scholar]
- Biswal BB, Mennes M, Zuo XN, Gohel S, Kelly C, et al. (2010). Toward discovery science of human brain function. Proc Natl Acad Sci U S A 107(10):4734–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Botvinik-Nezer R, Holzmeister F, Camerer CF, Dreber A, Huber J, Johannesson M, et al. (2020). Variability in the analysis of a single neuroimaging dataset by many teams. Nature 582(7810):84–88. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boutron I, Ravaud P (2018). Misrepresentation and distortion of research in biomedical literature. Proc Natl Acad Sci U S A 115(11):2613–2619. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bowring A, Telschow F, Schwartzman A, Nichols TE (2019). Spatial confidence sets for raw effect size images. Neuroimage 203:116187. doi: 10.1016/j.neuroimage.2019.116187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bzdok D, Thieme A, Levkovskyy O, Wren P, Ray T, Reddy S (2024). Data science opportunities of large language models for neuroscience and biomedicine. Neuron 112(5):698–717. doi: 10.1016/j.neuron.2024.01.016. [DOI] [PubMed] [Google Scholar]
- Calhoun VD, Adali T, Pearlson GD, Pekar JJ (2001). A method for making group inferences from functional MRI data using independent component analysis. Hum Brain Mapp 14(3):140–51. doi: 10.1002/hbm.1048. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Calhoun VD, Pekar JJ, McGinty VB, Adali T, Watson TD, Pearlson GD (2002). Different activation dynamics in multiple neural systems during simulated driving. Hum Brain Mapp 16(3):158–67. doi: 10.1002/hbm.10032. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen G, Taylor PA, Cox RW (2017). Is the statistic value all we should care about in neuroimaging? Neuroimage. 147:952–959. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen G, Taylor PA, Qu X, Molfese PJ, Bandettini PA, Cox RW, Finn ES (2020). Untangling the relatedness among correlations, part III: Inter-subject correlation analysis through Bayesian multilevel modeling for naturalistic scanning. Neuroimage 216:116474. doi: 10.1016/j.neuroimage.2019.116474. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen G, Pine DS, Brotman MA, Smith AR, Cox RW, Haller SP (2021). Trial and error: A hierarchical modeling approach to test-retest reliability. Neuroimage 245:118647. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen G, Pine DS, Brotman MA, Smith AR, Cox RW, Taylor PA, Haller SP (2022). Hyperbolic trade-off: the importance of balancing trial and subject sample sizes in neuroimaging. NeuroImage 247:118786. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen G, Taylor PA, Stoddard J, Cox RW, Bandettini PA, Pessoa L (2022). Sources of information waste in neuroimaging: mishandling structures, thinking dichotomously, and over-reducing data. Aperture Neuro. 2: DOI: 10.52294/2e179dbf-5e37-4338-a639-9ceb92b055ea [DOI] [PMC free article] [PubMed] [Google Scholar]
- Coursey SE, Mandeville J, Reed MB, Hartung GA, Garimella A, Sari H, Lanzenberger R, Price JC, Polimeni JR, Greve DN, Hahn A, Chen JE (2024). On the analysis of functional PET (fPET)-FDG: baseline mischaracterization can introduce artifactual metabolic (de)activations. bioRxiv [Preprint] 2024.10.17.618550. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cox RW (1996). AFNI: software for analysis and visualization of functional magnetic resonance neuroimages. Comput Biomed Res 29(3):162–173. doi: 10.1006/cbmr.1996.0014 [DOI] [PubMed] [Google Scholar]
- Cremers HR, Wager TD, Yarkoni T (2017). The relation between statistical power and inference in fMRI. PLoS One 12(11):e0184923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dafflon J F Da Costa P, Váša F, Monti RP, Bzdok D, Hellyer PJ, Turkheimer F, Smallwood J, Jones E, Leech R (2022). A guided multiverse study of neuroimaging analyses. Nat Commun 13(1):3758. doi: 10.1038/s41467-022-31347-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Damoiseaux JS, Rombouts SA, Barkhof F, Scheltens P, Stam CJ, Smith SM, Beckmann CF (2006). Consistent resting-state networks across healthy subjects. Proc Natl Acad Sci U S A 103(37):13848–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Etzel JA, Zacks JM, Braver TS (2013). Searchlight analysis: promise, pitfalls, and potential. Neuroimage 78:261–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fair DA, Cohen AL, Power JD, Dosenbach NU, Church JA, Miezin FM, Schlaggar BL, Petersen SE (2009). Functional brain networks develop from a “local to distributed” organization. PLoS Comput Biol 5(5):e1000381. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fischl B, Dale AM (2000). Measuring the thickness of the human cerebral cortex from magnetic resonance images. Proc Natl Acad Sci U S A 97(20):11050–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fisher RA (1925). Statistical Methods for Research Workers, Oliver and Boyd, Edinburgh. [Google Scholar]
- Flournoy JC, Vijayakumar N, Cheng TW, Cosme D, Flannery JE, Pfeifer JH (2020). Improving practices and inferences in developmental cognitive neuroscience. Dev Cogn Neurosci 45:100807. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Forman SD, Cohen JD, Fitzgerald M, Eddy WF, Mintun MA, Noll DC (1995). Improved assessment of significant activation in functional magnetic resonance imaging (fMRI): use of a cluster-size threshold. Magn Reson Med 33(5):636–47. [DOI] [PubMed] [Google Scholar]
- Freund MC, Chen R, Chen G, Braver TS (2025). Complementary benefits of multivariate and hierarchical models for identifying individual differences in cognitive control. Imaging Neuroscience (2025) 3: imag_a_00447. 10.1162/imag_a_00447 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goebel R (2012). BrainVoyager--past, present, future. NeuroImage 62, 748–56. 10.1016/j.neuroimage.2012.01.083 [DOI] [PubMed] [Google Scholar]
- Gonzalez-Castillo J, Saad ZS, Handwerker DA, Inati SJ, Brenowitz N, Bandettini PA (2012). Whole-brain, time-locked activation with simple tasks revealed using massive averaging and model-free analysis. Proc Natl Acad Sci U S A 109(14):5487–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gorgolewski KJ, Storkey AJ, Bastin ME, Pernet CR (2012). Adaptive thresholding for reliable topological inference in single subject fMRI analysis. Front Hum Neurosci 6:245. doi: 10.3389/fnhum.2012.00245. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gorgolewski KJ, Varoquaux G, Rivera G, Schwarz Y, Ghosh SS, Maumet C, Sochat VV, Nichols TE, Poldrack RA, Poline JB, Yarkoni T, Margulies DS (2015). NeuroVault.org: a web-based repository for collecting and sharing unthresholded statistical maps of the human brain. Front Neuroinform. 9:8. doi: 10.3389/fninf.2015.00008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Greicius MD, Krasnow B, Reiss AL, Menon V (2003). Functional connectivity in the resting brain: a network analysis of the default mode hypothesis. Proc Natl Acad Sci U S A 100(1):253–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hagmann P, Cammoun L, Gigandet X, Meuli R, Honey CJ, Wedeen VJ, Sporns O (2008). Mapping the structural core of human cerebral cortex. PLoS Biol 6(7):e159. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Halsey LG, Curran-Everett D, Vowler SL, Drummond GB (2015). The fickle P value generates irreproducible results. Nat Methods 12(3):179–85. [DOI] [PubMed] [Google Scholar]
- Hanayik T, Rorden C, Drake C, Thual A, Taylor P, Hardcastle N, McCarthy P, Androulakis A, Markiewicz C, Nedelec P (2023). niivue/niivue: 0.37.0 (0.37.0). Zenodo. 10.5281/zenodo.8320937 [DOI] [Google Scholar]
- Ioannidis JP, Munafò MR, Fusar-Poli P, Nosek BA, David SP (2014). Publication and other reporting biases in cognitive sciences: detection, prevalence, and prevention. Trends Cogn Sci 18(5):235–41. doi: 10.1016/j.tics.2014.02.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jennings RG, Van Horn JD (2012). Publication bias in neuroimaging research: implications for meta-analyses. Neuroinformatics 10(1):67–80. doi: 10.1007/s12021-011-9125-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jernigan TL, Gamst AC, Fennema-Notestine C, Ostergaard AL (2003). More “mapping” in brain mapping: statistical comparison of effects. Hum Brain Mapp 19(2):90–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jung B, Taylor PA, Seidlitz PA, Sponheim C, Perkins P, Ungerleider LG, Glen DR, Messinger A (2021). A Comprehensive Macaque FMRI Pipeline and Hierarchical Atlas. NeuroImage 235:117997. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kennedy-Shaffer L (2019). Before p < 0.05 to Beyond p < 0.05: Using History to Contextualize p-Values and Significance Testing. Am Stat 73(Suppl 1):82–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kinsey S, Kazimierczak K, Camazón PA, Chen J, Adali T, Kochunov P, Adhikari B, Ford J, van Erp TGM, Dhamala M, Calhoun VD, Iraji A (2024). Networks extracted from nonlinear fMRI connectivity exhibit unique spatial variation and enhanced sensitivity to differences between individuals with schizophrenia and controls. Nat Ment Health. 2024;2(12):1464–1475. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lefort-Besnard J, Nichols TE, Maumet C (2024). Statistical Inference for Same Data Meta-Analysis in Neuroimaging Multiverse Analyzes. hal-04754078v2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lerch J, Hammill C, van Eede M, Cassel D (2017). RMINC: Statistical Tools for Medical Imaging NetCDF (MINC) Files. R package version 1.5.2.1, http://mouse-imaging-centre.github.io/RMINC. [Google Scholar]
- Lohmann G, Stelzer J, Müller K, Lacosse E, Buschmann T, Kumar VJ, Grodd W, Scheffler K (2017). Inflated false negative rates undermine reproducibility in task-based fMRI. bioRxiv 122788; doi: 10.1101/122788. [DOI] [Google Scholar]
- Luo WL, Nichols TE (2003). Diagnosis and exploration of massively univariate neuroimaging models. Neuroimage 19(3):1014–32. [DOI] [PubMed] [Google Scholar]
- Mandino F, Cerri DH, Garin CM, Straathof M, van Tilborg GAF, Chakravarty MM, Dhenain M, Dijkhuizen RM, Gozzi A, Hess A, Keilholz SD, Lerch JP, Shih YI, Grandjean J (2020). Animal Functional Magnetic Resonance Imaging: Trends and Path Toward Standardization. Front Neuroinform 13:78. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maumet C, Auer T, Bowring A, Chen G, Das S, Flandin G, Ghosh S, Glatard T, Gorgolewski KJ, Helmer KG, Jenkinson M, Keator DB, Nichols BN, Poline JB, Reynolds R, Sochat V, Turner J, Nichols TE (2016). Sharing brain mapping statistical results with the neuroimaging data model. Sci Data 3:160102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Maumet C, Nichols TE (2016). Minimal Data Needed for Valid and Accurate Image-Based fMRI Meta-Analysis. bioRxiv. doi: 10.1101/048249 [DOI] [Google Scholar]
- McCarthy P (2024). FSLeyes (1.11.0). Zenodo. 10.5281/zenodo.11047709 [DOI] [Google Scholar]
- Nichols T, Hayasaka S (2003). Controlling the familywise error rate in functional neuroimaging: a comparative review. Stat Methods Med Res 12(5):419–46. [DOI] [PubMed] [Google Scholar]
- Nichols TE, Das S, Eickhoff SB, Evans AC, Glatard T, Hanke M, Kriegeskorte N, Milham MP, Poldrack RA, Poline JB, Proal E, Thirion B, Van Essen DC, White T, Yeo BT (2017). Best practices in data analysis and sharing in neuroimaging using MRI. Nat Neurosci 20(3):299–303. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nilearn contributors (2025) ‘nilearn’. Zenodo. doi: 10.5281/zenodo.14697221. [DOI] [Google Scholar]
- Noble S, Curtiss J, Pessoa L, Scheinost D (2024). The tip of the iceberg: A call to embrace anti-localizationism in human neuroscience research. Imaging Neuroscience 2: 1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- NPG (2018). Checklists work to improve science. Nature 556:273–274. doi: 10.1038/d41586-018-04590-7. [DOI] [PubMed] [Google Scholar]
- Pang JC, Aquino KM, Oldehinkel M, Robinson PA, Fulcher BD, Breakspear M, Fornito A (2023). Geometric constraints on human brain function. Nature 618(7965):566–574. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pernet CR, Madan CR (2020). Data visualization for inference in tomographic brain imaging. Eur J Neurosci 51(3):695–705. [DOI] [PubMed] [Google Scholar]
- Pessoa L (2014). Understanding brain networks and brain organization. Phys Life Rev 11(3):400–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Posse S, Wiese S, Gembris D, Mathiak K, Kessler C, Grosse-Ruyken M, Elghahwagi B, Richards T, Dager S, Kiselev V (1999). Enhancement of BOLD-contrast sensitivity by single-shot multi-echo functional MR imaging. Magnetic Resonance in Medicine. 42:87–97. [DOI] [PubMed] [Google Scholar]
- Reynolds RC, Taylor PA, Glen DR (2023). Quality control practices in FMRI analysis: Philosophy, methods and examples using AFNI. Front. Neurosci. 16:1073800. doi: 10.3389/fnins.2022.1073800 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reynolds RC, Glen DR, Chen G, Saad ZS, Cox RW, Taylor PA (2024). Processing, evaluating and understanding FMRI data with afni_proc.py. Imaging Neuroscience 2:1–52. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ruff IM, Petrovich Brennan NM, Peck KK, Hou BL, Tabar V, Brennan CW, Holodny AI (20008). Assessment of the language laterality index in patients with brain tumor using functional MR imaging: effects of thresholding, task selection, and prior surgery. AJNR Am J Neuroradiol 29(3):528–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Russ BE, Leopold DA (2015). Functional MRI mapping of dynamic visual features during natural viewing in the macaque. Neuroimage 109:84–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Saad ZS, Reynolds RC, Argall B, Japee S, Cox RW (2004). SUMA: an interface for surface-based intra- and inter-subject analysis with AFNI. Presented at the 2nd IEEE International Symposium on Biomedical Imaging: Nano to Macro (IEEE Cat No. 04EX821), pp. 1510–1513 Vol. 2. [Google Scholar]
- Saad ZS, Reynolds RC (2012). SUMA. Neuroimage 62, 768–773. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sava-Segal C, Grall, C Finn ES(2024). Narrative ‘twist’ shifts within-individual neural representations of dissociable story features. bioRxiv preprint. https://www.biorxiv.org/content/10.1101/2025.01.13.632631v1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Savoy RL (2001). History and future directions of human brain mapping and functional neuroimaging. Acta Psychol (Amst) 107(1–3):9–42. [DOI] [PubMed] [Google Scholar]
- Seghier ML (2008). Laterality index in functional MRI: methodological issues. Magn Reson Imaging 26(5):594–601. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shakespeare W (1597). An Excellent Conceited Tragedie of Romeo and Juliet. John Danter; (London: ). [Google Scholar]
- Shapiro FR (2006). The New Yale Book of Quotations. Yale University Press; (New Haven: ), p231. [Google Scholar]
- Smith SM, Jenkinson M, Woolrich MW, Beckmann CF, Behrens TE, Johansen-Berg H, Bannister PR, De Luca M, Drobnjak I, Flitney DE, Niazy RK, Saunders J, Vickers J, Zhang Y, De Stefano N, Brady JM, Matthews PM (2004). Advances in functional and structural MR image analysis and implementation as FSL. Neuroimage. 23 Suppl 1:S208–19. [DOI] [PubMed] [Google Scholar]
- Smith SM, Fox PT, Miller KL, Glahn DC, Fox PM, Mackay CE, Filippini N, Watkins KE, Toro R, Laird AR, Beckmann CF (2009). Correspondence of the brain’s functional architecture during activation and rest. Proc Natl Acad Sci U S A 106(31):13040–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smith SM, Nichols TE (2009). Threshold-free cluster enhancement: addressing problems of smoothing, threshold dependence and localisation in cluster inference. Neuroimage 44(1):83–98. [DOI] [PubMed] [Google Scholar]
- Smith AR, White LK, Leibenluft E, McGlade AL, Heckelman AC, Haller SP, Buzzell GA, Fox NA, Pine DS (2020). The heterogeneity of anxious phenotypes: neural responses to errors in treatment-Seeking anxious and behaviorally inhibited youths. Journal of the American Academy of Child & Adolescent Psychiatry 59(6):759–769. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Suarez RO, Whalen S, Nelson AP, Tie Y, Meadows ME, Radmanesh A, Golby AJ (2009). Threshold-independent functional MRI determination of language dominance: a validation study against clinical gold standards. Epilepsy Behav 16(2):288–97. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sundermann B, Pfleiderer B, McLeod A, Mathys C (2024). Seeing more than the Tip of the Iceberg: Approaches to Subthreshold Effects in Functional Magnetic Resonance Imaging of the Brain. Clin Neuroradiol 34(3):531–539. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Szucs D, Ioannidis JP (2020). Sample size evolution in neuroimaging research: An evaluation of highly-cited studies (1990–2012) and of latest practices (2017–2018) in high-impact journals. Neuroimage 221:117164. [DOI] [PubMed] [Google Scholar]
- Taylor PA, Gotts SJ, Gilmore AW, Teves J, Reynolds RC (2022). A multi-echo FMRI processing demo including TEDANA in afni_proc.py pipelines. Proc. OHBM-2022. https://afni.nimh.nih.gov/pub/dist/OHBM2022/OHBM2022_tayloretal_apmulti.pdf [Google Scholar]
- Taylor PA, Reynolds RC, Calhoun V, Gonzalez-Castillo J, Handwerker DA, Bandettini PA, Mejia AF, Chen G (2023). Highlight Results, Don’t Hide Them: Enhance interpretation, reduce biases and improve reproducibility. Neuroimage 274:120138. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Taylor PA, Glen DR, Chen G, Cox RW, Hanayik T, Rorden C, Nielson DM, Rajendra JK, Reynolds RC (2024). A Set of FMRI Quality Control Tools in AFNI: Systematic, in-depth and interactive QC with afni_proc.py and more. Imaging Neuroscience 2: 1–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Turner BO, Paul EJ, Miller MB, Barbey AK (2018). Small sample sizes reduce the replicability of task-based fMRI studies. Communications biology 1(1):62. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vanduffel W, Fize D, Mandeville JB, Nelissen K, Van Hecke P, Rosen BR, Tootell RB, Orban GA (2001). Visual motion processing investigated using contrast agent-enhanced fMRI in awake behaving monkeys. Neuron 32(4):565–77. doi: 10.1016/s0896-6273(01)00502-5. [DOI] [PubMed] [Google Scholar]
- Van Essen DC, Smith SM, Barch DM, Behrens TE, Yacoub E, Ugurbil K; WU-Minn HCP Consortium (2013). The WU-Minn Human Connectome Project: an overview. Neuroimage 80:62–79. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Visconti di Oleggio Castello M, Halchenko YO, Guntupalli JS, Gors JD, Gobbini MI(2017). The neural representation of personally familiar and unfamiliar faces in the distributed system for face perception. Sci Rep 7(1):12237. doi: 10.1038/s41598-017-12559-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Voets N, Ashtari M, Beckmann C, Benjamin C, Benzinger T, Binder JR, ... Bookheimer S(2025). Consensus recommendations for clinical functional MRI applied to language mapping. Aperture Neuro 5. doi: 10.52294/001c.128149 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wasserstein RL Lazar NA (2016). The ASA Statement on p-Values: Context, Process, and Purpose. The American Statistician 70(2):129–133. doi: 10.1080/00031305.2016.1154108. [DOI] [Google Scholar]
- Wicherts JM, Veldkamp CLS, Augusteijn HEM, Bakker M, van Aert RCM, & van Assen MALM(2016). Degrees of freedom in planning, running, analyzing, and reporting psychological studies: A checklist to avoid p-hacking. Frontiers in Psychology, 7, Article 1832. doi: 10.3389/fpsyg.2016.01832 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yarkoni T, Poldrack RA, Nichols TE, Van Essen DC, Wager TD (2011). Large-scale automated synthesis of human functional neuroimaging data. Nat Methods. 8(8):665–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.







