Skip to main content
. 2010 May 11;107(21):9546–9551. doi: 10.1073/pnas.0914005107

Fig. 2.

Fig. 2.

(A) The null distribution of the test statistic is affected by filtering on the maximum of within-class averages. In this example, all genes have a known common variance, the filter statistic is the maximum of within-class means, and the test statistic is a z-score. The unconditional distribution of the test statistic for nondifferentially expressed genes is a standard normal. Its conditional null distribution, given that the filter statistic (UI) exceeds a certain threshold (u), however, has much heavier tails. Using the unconditional null distribution to compute p-values after filtering would therefore be inappropriate. See SI Text for full details. (B and C) Overall variance filtering and the limma moderated t-statistic. Data for 5,000 nondifferentially expressed genes were generated according to the limma Bayesian model (n1 = n2 = 2, d0 = 3, Inline graphic). (B) Filtering on overall variance (θ = 0.5) preferentially eliminated genes with small si, causing gene-level standard deviation estimates for genes passing the filter (histogram) to be shifted relative to the unconditional distribution used to generate the data (dashed curve). The limma inverse χ2 model was unable to provide a good fit (solid curve) to the si passing the filter. (C) The fitting problems lead to a posterior degrees-of-freedom estimate of ∞. As a consequence, p-values were computed using an inappropriate null distribution, producing too many true-null p-values close to zero, i.e., loss of type I error rate control. An analogous analysis comparing biological replicates from the ALL study—so that real array data were used but no gene was expected to exhibit significant differential expression—yielded qualitatively similar results.