Skip to main content
American Journal of Epidemiology logoLink to American Journal of Epidemiology
. 2017 Aug 21;186(6):636–638. doi: 10.1093/aje/kwx258

Invited Commentary: Can Issues With Reproducibility in Science Be Blamed on Hypothesis Testing?

Clarice R Weinberg *
PMCID: PMC5860396  PMID: 28938713

Abstract

In the accompanying article (Am J Epidemiol. 2017;186(6):646–647), Dr. Timothy Lash makes a forceful case that the problems with reproducibility in science stem from our “culture” of null hypothesis significance testing. He notes that when attention is selectively given to statistically significant findings, the estimated effects will be systematically biased away from the null. Here I revisit the recent history of genetic epidemiology and argue for retaining statistical testing as an important part of the tool kit. Particularly when many factors are considered in an agnostic way, in what Lash calls “innovative” research, investigators need a selection strategy to identify which findings are most likely to be genuine, and hence worthy of further study.

Keywords: epidemiologic methods, P values, reproducibility of results, significance testing


In the accompanying article in this issue of the Journal, Dr. Timothy Lash (1) provides a thought-provoking perspective on what has recently been discussed as a crisis of reproducibility in science. He asserts that the “culture” of null hypothesis significance testing is the primary culprit in this perceived failure of reproducibility. He shows that findings selected on the basis of statistical testing will tend to have exaggerated effect measures. Lash argues for methods that account for systematic errors and prior information, to help ensure reproducibility. He further asserts that efforts to improve reproducibility that do not begin by discarding the practice of significance testing are doomed to fail. While Lash finds much to criticize, his proposed alternatives are unclear, as is his vision for identifying findings that have been satisfactorily reproduced.

Concern about reproducibility goes well beyond epidemiology. The interested reader should see links provided on the website of the Reproducibility Project: Cancer Biology (2) for a project that is attempting to replicate preclinical cancer biology experiments reported in selected highly cited papers.

In his article (1), Lash begins by rightly pointing out that the probability that a “significant” result corresponds to a true finding depends on the research context. In endeavors that he terms “innovative,” the false-positive rate will be high, because most null hypotheses being tested are truly null. The analyst wants to know the positive predictive value for a finding: Given this low P value, what is the probability that this is a true association?

As a clinical analogy, if a dim-witted official were to serologically test nuns for human immunodeficiency virus (HIV), a large proportion of the positive test results would be false-positive. By contrast, in a population with a high a priori likelihood of HIV infection, a much lower fraction of positive test results would be false-positive. The investigator/analyst who wants to know the likelihood that a rejected null hypothesis corresponds to a true finding will need to have a way to account for the proportion of hypotheses tested that are truly null. The question is important because we would like to know which findings are worth devoting follow-up resources to—that is, which are likely to be true and ultimately replicable.

Another issue Lash raises is bias in estimation of measures of effect. Odds ratios that meet a threshold cutoff for statistical significance are likely to be exaggerated in comparison with the true parameter. His graphs illustrate that point nicely. The corresponding confidence intervals are also biased, and their coverage probabilities will fall below the nominal 95%. In the general setting, both biases are then exacerbated by publication bias: Journals favor papers with low P values. They understandably want the press attention and future citations that important findings attract. Then there is a “file drawer” problem, because negative findings are both boring and challenging to write about. As Lash acknowledges, those sources of selection bias will not be avoided by abstaining from explicit significance testing, even if journals and authors move toward estimation and confidence intervals (or Bayesian credible intervals).

Lash recommends mitigating estimation bias by using methods for shrinking the estimates toward the null, or methods of accounting (e.g., with resampling or Bayesian approaches) for bias and prior information. Such methods can be extremely helpful, and textbook Bayesianism is intuitively very appealing in that it builds on expanding knowledge. However, additional assumptions must be made, and Bayesians often ignore that prior knowledge and use uninformative priors.

One practice that has been underused is the construction of 1-sided P values or their siblings, 1-sided confidence intervals. In many settings, 1-sided inference could be defensible (with either frequentist or Bayesian approaches)—for example, for exposures that could cause cancer but would not plausibly protect people from cancer.

The same reproducibility-related issues that Lash highlights have played out in genetic studies, although Lash uses different terminology (3). Results from early work in genetics based on selected “candidate” genes were often disappointing. Following the completion of the Human Genome Project in 2003, genetic epidemiologists faced daunting reproducibility problems as they now carried out hundreds of thousands of hypothesis tests in “agnostic” locus-by-locus analyses across the genome. The bias seen in estimation that is discussed by Lash is what genetic epidemiologists call the “winner's curse” (4). It is now widely appreciated that genotype odds ratios for top “hits” are initially overestimated.

Genetic epidemiologists recognized that investigators using α = 0.05 would waste resources by following up thousands of false-positive findings. They initially attempted to control the “familywise” error rate by using a Bonferroni multiplier on the P values. This strategy seems misguided to me in settings where heritability is known for the disease under study, because the objective is not to address a “familywise” null hypothesis but to select genetic variants that are worth further investigation because they are associated with risk. In fact, the Bonferroni approach doubtless produced a huge number of type II errors because of its terrible stringency.

In a highly influential paper, Benjamini and Hochberg (5) proposed methods for controlling the false-discovery rate when many independent statistical tests are performed. Storey and Tibshirani (6) later introduced the “q value,” which measures the false-discovery rate and tolerates “weak dependency.” In recent big-data studies (including methylation, gene expression, and microbiome studies), P values have largely been replaced by false-discovery-rate corrected values.

In 2008, Wakefield (7) extended the methods in order to identify which findings are “noteworthy.” Wakefield pointed out that “the P value is by far the most commonly used measure, but requires careful calibration when the a priori probability of an association is small, and discards information by not considering the power associated with each test” (7, p. 641). He proposed a Bayes factor approach that accounts for information about the likely proportion of null hypotheses that are false and the distribution of the associated effect sizes. Kuo et al. (8) suggested a similarly straightforward conversion of P values to posterior probabilities.

Lash points out that hypothesis tests often rely on assumptions that are difficult to verify and that can distort the type I error rate when they fail. Investigators need to be aware that most hypothesis tests do require assumptions. Estimation procedures in general typically require additional assumptions, and bias correction procedures require even more. Even inclusion of a linear exposure term in a logistic model carries implicit assumptions. Such models can often be misspecified, leading us to target a poorly defined parameter.

I feel that Lash goes too far with his ultimate recommendation that we “discard the current culture of null hypothesis significance testing and its detrimental influences on study planning, data analysis, presentation of results, and inference” (1, p. 632). When there is a drowning, few would call for the lake to be drained, though many would want to improve water safety. We could do a much better job of teaching our students the valid uses, interpretations, and limitations of P values.

A good part of the perceived lack of replicability of research findings has to do with the common (and predictable) occurrence of an association that is statistically significant (and therefore thought to be real) in one study and then nonsignificant in another—“positive,” then “negative” (which often means nonsignificant). Authors and readers need to appreciate the role of sampling variability and think harder before characterizing such instances as “inconsistent findings.”

It is true that P values are imperfect measures of the extent of evidence against the null hypothesis, but confidence intervals have problems of their own. They are also frequentist by definition, and they are also easy to misinterpret. For example, it is often presumed that the probability is 95% that a given 95% confidence interval covers the true parameter. In fact, the 95% coverage is only guaranteed under a hypothetical scenario in which infinitely many similar and independent experiments are carried out. One can posit a probabilistic interpretation for “credible” intervals under Bayesian approaches, but the randomness that must be assumed for parameters will not sit well with many scientists.

There are also common situations where I find a P value more useful than an interval (9). For example, if the predictor of interest is a multicategorical variable, then the corresponding parameter is a vector in multidimensional space. Suppose, for example, that the predictor has 5 categories and these are unordered. In such a setting, one needs to consider a 4-dimensional confidence region (or credibility region) rather than separate confidence intervals, and such high-dimensional regions are devilishly hard to compute and depict in a research paper. In such a setting, the multidimensional region may exclude the null value, but the reader cannot discern this from the list of separate confidence intervals. Epidemiologists often fall into the trap of cherry-picking the one category whose interval excludes 1.0, in order to tell a story. I would rather see both the category-specific confidence intervals and the 4-degree-of-freedom χ2 value with its associated P value. That at least would tell me which high-dimensional confidence regions excluded the null vector.

If the multiple categories in the predictor are ordered—for example, for levels of education—then often the analyst would want to carry out a trend test, by coding the categories as 0, 1, 2, 3, and 4 and testing the coefficient against a null value of 0. But that trend-test coefficient has no particular meaning in itself, limiting the usefulness of its corresponding confidence interval. The P value provides a way to assess the degree to which the data are incompatible with equality across the categories.

A serious issue with confidence intervals is that they require us to know the correct model under the alternative hypothesis before we can identify the parameter to be estimated. Some years ago, my colleagues and I explored possible seasonal effects on early pregnancy loss, based on the North Carolina Early Pregnancy Study. The paper was published (with a P value) in Epidemiology (10). We did not have a particular alternative in mind but wanted to be able to detect a unimodal departure from the null hypothesis, where the null was simple constancy of risk across the days of the year of conception. We did this by fitting a sine wave using a harmonic logistic model for early pregnancy loss. This alternative model requires representing days of the year as points on the unit circle, trigonometrically transforming to the sine and cosine, and estimating the 2 coefficients. Although we found a significant departure from the null hypothesis, we do not believe the true alternative is likely to be in the form of a sine wave. Nevertheless, trigonometric regression offered a convenient and powerful analytical approach with which to assess departures from constancy of risk. Had we provided confidence intervals for the coefficients of the sine and cosine, most readers would not have known what to do with them.

Finally, Lash makes a compelling argument that the validity of significance testing requires that there be no selection bias and no measurement error. This would seem to be the final nail in the coffin, as few of us would feel comfortable claiming perfection. However, Lash's argument is not generally valid, because in many settings the null hypothesis stays null under selection-bias/measurement-error scenarios. For example, when we tested for seasonality of early pregnancy loss, random errors in estimating the date of conception would not have disturbed the flatness of the risk across seasons, so the null was preserved. Exceptions occur if errors are differential by case status, potentially invalidating both testing and estimation.

In my view, the concern about reproducibility/replicability in science has been overblown. Nevertheless, high-status journals—probably within each field—do have a troubling problem, because authors feel compelled to highlight the importance of their findings in order to get their papers accepted by leading journals, to get their next grant funded and to get promoted. The consequent distortion will be offset to some extent by the emergence of publications such as the Public Library of Science (PLOS) journals, which exclude importance from their criteria for publication and explicitly direct reviewers to care primarily about clarity and correctness. I am an optimist, however—I believe that science as an enterprise is inherently self-correcting. Provided that scientists can continue to be trustworthy and can continue to trust one another, the real associations will ultimately emerge and our work will trend toward the truth, even in epidemiology.

ACKNOWLEDGMENTS

Author affiliation: Biostatistics and Computational Biology Branch, National Institute of Environmental Health Sciences, Research Triangle Park, North Carolina (Clarice R. Weinberg).

This work was supported by the Intramural Research Program of the National Institute of Environmental Health Sciences under project Z01 ES040006.

I thank Drs. Marie Davidian, Dmitri Zaykin, Katie O'Brien, Richard Weinberg, and Allen Wilcox for comments on a draft of this article.

Conflict of interest: none declared.

REFERENCES

  • 1. Lash TL. The harm done to reproducibility by the culture of null hypothesis significance testing. Am J Epidemiol. 2017;186(6):627–635. [DOI] [PubMed] [Google Scholar]
  • 2. eLife Sciences Publications, Ltd Reproducibility Project: Cancer Biology. Investigating reproducibility in preclinical cancer research. 2017. https://elifesciences.org/collections/reproducibility-project-cancer-biology. Accessed March 13, 2017.
  • 3. Peng RD. Reproducible research and Biostatistics. Biostatistics. 2009;10(3):405–408. [DOI] [PubMed] [Google Scholar]
  • 4. Kraft P. Curses—winner's and otherwise—in genetic epidemiology. Epidemiology. 2008;19(5):649–651. [DOI] [PubMed] [Google Scholar]
  • 5. Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B Methodol. 1995;57(1):289–300. [Google Scholar]
  • 6. Storey JD, Tibshirani R. Statistical significance for genomewide studies. Proc Natl Acad Sci USA. 2003;100(16):9440–9445. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Wakefield J. Reporting and interpretation in genome-wide association studies. Int J Epidemiol. 2008;37(3):641–653. [DOI] [PubMed] [Google Scholar]
  • 8. Kuo CL, Vsevolozhskaya OA, Zaykin DV. Assessing the probability that a finding is genuine for large-scale genetic association studies. PLoS One. 2015;10(5):e0124107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Weinberg CR. It's time to rehabilitate the P-value. Epidemiology. 2001;12(3):288–290. [DOI] [PubMed] [Google Scholar]
  • 10. Weinberg CR, Moledor E, Baird DD, et al. Is there a seasonal pattern in risk of early pregnancy loss. Epidemiology. 1994;5(5):484–489. [PubMed] [Google Scholar]

Articles from American Journal of Epidemiology are provided here courtesy of Oxford University Press

RESOURCES