Global and Simultaneous Hypothesis Testing for High-Dimensional Logistic Regression Models

Rong Ma; T Tony Cai; Hongzhe Li

doi:10.1080/01621459.2019.1699421

. Author manuscript; available in PMC: 2022 Jan 1.

Published in final edited form as: J Am Stat Assoc. 2020 Jan 21;116(534):984–998. doi: 10.1080/01621459.2019.1699421

Global and Simultaneous Hypothesis Testing for High-Dimensional Logistic Regression Models

Rong Ma ^1,^2,³, T Tony Cai ^1,^2,³, Hongzhe Li ^1,^2,³

PMCID: PMC8375316 NIHMSID: NIHMS1550238 PMID: 34421157

Abstract

High-dimensional logistic regression is widely used in analyzing data with binary outcomes. In this paper, global testing and large-scale multiple testing for the regression coefficients are considered in both single- and two-regression settings. A test statistic for testing the global null hypothesis is constructed using a generalized low-dimensional projection for bias correction and its asymptotic null distribution is derived. A lower bound for the global testing is established, which shows that the proposed test is asymptotically minimax optimal over some sparsity range. For testing the individual coefficients simultaneously, multiple testing procedures are proposed and shown to control the false discovery rate (FDR) and falsely discovered variables (FDV) asymptotically. Simulation studies are carried out to examine the numerical performance of the proposed tests and their superiority over existing methods. The testing procedures are also illustrated by analyzing a data set of a metabolomics study that investigates the association between fecal metabolites and pediatric Crohn’s disease and the effects of treatment on such associations.

Keywords: False discovery rate, Global testing, Large-scale multiple testing, Minimax lower bound

1. INTRODUCTION

Logistic regression models have been applied widely in genetics, finance, and business analytics. In many modern applications, the number of covariates of interest usually grows with, and sometimes far exceeds, the number of observed samples. In such high-dimensional settings, statistical problems such as estimation, hypothesis testing, and construction of confidence intervals become much more challenging than those in the classical low-dimensional settings. The increasing technical difficulties usually emerge from the non-asymptotic analysis of both statistical models and the corresponding computational algorithms.

In this paper, we consider testing for high-dimensional logistic regression model:

log (\frac{π_{i}}{1 - π_{i}}) = X_{i}^{⊤} β, for i = 1, \dots, n .

(1)

where $β \in ℝ^{p}$ is the vector of regression coefficients. The observations are i.i.d. samples Z_i = (y_i,X_i) for i = 1,..,n, and we assume y_i | X_i ~ Bernoulli(π_i) independently for each i = 1, …, n.

1.1. Global and Simultaneous Hypothesis Testing

It is important in high-dimensional logistic regression to determine 1) whether there are any associations between the covariates and the outcome and, if yes, 2) which covariates are associated with the outcome. The first question can be formulated as testing the global null hypothesis H₀:β = 0; and the second question can be considered as simultaneously testing the null hypotheses H_0,i:β_i = 0 for i = 1, …, p. Besides such single logistic regression problems, hypothesis testing involving two logistic regression models with regression coefficients β⁽¹⁾ and β² in $ℝ^{p}$ is also important. Specifically, one is interested in testing the global null hypothesis H₀:β⁽¹⁾ = β⁽²⁾, or identifying the differentially associated covariates through simultaneously testing the null hypotheses $H_{0, i} : β_{i}^{(1)} = β_{i}^{(2)}$ for each i = 1, …, p.

Estimation for high-dimensional logistic regression has been studied extensively. van de Geer (2008) considered high-dimensional generalized linear models (GLMs) with Lipschitz loss functions, and proved a non-asymptotic oracle inequality for the empirical risk minimizer with the Lasso penalty. Meier, van de Geer, and Bühlmann (2008) studied the group Lasso for logistic regression and proposed an efficient algorithm that leads to statistically consistent estimates. Negahban et al. (2010) obtained the rate of convergence for the $l_{1}$ -regularized maximum likelihood estimator under GLMs using restricted strong convexity property. Bach (2010) extended tools from the convex optimization literature, namely self-concordant functions, to provide interesting extensions of theoretical results for the square loss to the logistic loss. Plan and Vershynin (2013) connected sparse logistic regression to one-bit compressed sensing and developed a unified theory for signal estimation with noisy observations.

In contrast, hypothesis testing and confidence intervals for high-dimensional logistic regression have only been recently addressed. van de Geer et al. (2014) considered constructing confidence intervals and statistical tests for single or low-dimensional components of the regression coefficients in high-dimensional GLMs. Mukherjee, Pillai, and Lin (2015) studied the detection boundary for minimax hypothesis testing in high-dimensional sparse binary regression models when the design matrix is sparse. Belloni, Chernozhukov, and Wei (2016) considered estimating and constructing the confidence regions for a regression coefficient of primary interest in GLMs. More recently, Sur, Chen, and Candès (2017) and Sur and Candès (2019) considered the likelihood ratio test for high-dimensional logistic regression under the setting that p / n → κ for some constant κ < 1 / 2, and showed that the asymptotic null distribution of the log-likelihood ratio statistic is a rescaled χ² distribution. Cai et al. (2017) proposed a global test and a multiple testing procedure for differential networks against sparse alternatives under the Markov random field model. Nevertheless, the problems of global testing and large-scale simultaneous testing for high-dimensional logistic regression models with p ≳ n remain unsolved.

In this paper, we first consider global and multiple testing for a single high-dimensional logistic regression model. The global test statistic is constructed as the maximum of squared standardized statistics for individual coefficients, which are based on a two-step standardization procedure. The first step is to correct the bias of the logistic Lasso estimator using a generalized low-dimensional projection (LDP) method, and the second step is to normalize the resulting nearly unbiased estimators by their estimated standard errors. We show that the asymptotic null distribution of the test statistic is a Gumbel distribution and that the resulting test is minimax optimal under the Gaussian design by establishing the minimax separation distance between the null space and alternative space. For large-scale multiple testing, data-driven testing procedures are proposed and shown to control the false discovery rate (FDR) and falsely discovered variables (FDV) asymptotically. The framework for testing for single logistic regression is then extended to the setting of testing two logistic regression models.

The main contributions of the present paper are threefold.

We propose novel procedures for both the global testing and large-scale simultaneous testing for high dimensional logistic regressions. The dimension p is allowed to be much larger than the sample size n. Specifically, we require $log p = O (n^{c_{1}})$ for the global test and $p = O (n^{c_{2}})$ for the multiple testing procedure, with some constant c₁, c₂ > 0. For the global alternatives characterized by the $l_{\infty}$ norm of the regression coefficients, the global test is shown to be minimax rate optimal with the optimal separation distance of order $\sqrt{log p / n}$ .
Following similar ideas in Ren, Zhang, and Zhou (2016) and Cai et al. (2017), our construction of the test statistics depends on a generalized version of the LDP method for bias correction. The original LDP method (Zhang and Zhang 2014) relies on the linearity between the covariates and outcome variable. For logistic regression, the generalized approach first finds a linearization of the regression function, and the weighted LDP is then applied. Besides its usefulness in logistic regression, the generalized LDP method is flexible and can be applied to other nonlinear regression problems (see Section 7 for a detailed discussion).
The minimax lower bound is obtained for the global hypothesis testing under the Gaussian design. The lower bound depends on the calculation of the χ²-divergence between two logistic regression models. To the best of our knowledge, this is the first lower bound result for high-dimensional logistic regression under the Gaussian design.

1.2. Other Related Work

We should note that a different but related problem, namely inference for high-dimensional linear regression, has been well studied in the literature. Zhang and Zhang (2014), van de Geer et al. (2014) and Javanmard and Montanari (2014a,b) considered confidence intervals and testing for low-dimensional parameters of the high-dimensional linear regression model and developed methods based on a two-stage debiased estimator that corrects the bias introduced at the first stage due to regularization. Cai and Guo (2017) studied minimaxity and adaptivity of confidence intervals for general linear functionals of the regression vector.

The problems of global testing and large-scale simultaneous testing for high-dimensional linear regression have been studied by Liu and Luo (2014), Ingster, Tsybakov, and Verzelen (2010) and more recently by Xia, Cai, and Cai (2018) and Javanmard and Javadi (2019). However, due to the nonlinearity and the binary outcome, the approaches used in these works cannot be directly applied to logistic regression problems. In the Markov random field setting, Ren, Zhang, and Zhou (2016) and Cai et al. (2017) constructed pivotal/test statistics based on the debiased LDP estimators for node-wise logistic regressions with binary covariates. However, the results for sparse high-dimensional logistic regression models with general continuous covariates remain unknown.

Other related problems include joint testing and false discovery rate control for high-dimensional multivariate regression (Xia, Cai, and Li 2018) and testing for high-dimensional precision matrices and Gaussian graphical models (Liu 2013; Xia, Cai, and Cai 2015), where the inverse regression approach and de-biasing were carried out in the construction of the test statistics. Such statistics were then used for testing the global null with extreme value type asymptotic null distributions or to perform multiple testing that controls the false discovery rate.

1.3. Organization of the Paper and Notations

The rest of the paper is organized as follows. In Section 2, we propose the global test and establish its optimality. Some comparisons with existing works are made in detail. In Section 3, we present the multiple testing procedures and show that they control the FDR/FDP or FDV/FWER asymptotically. The framework is extended to the two-sample setting in Section 4. In Section 5, the numerical performance of the proposed tests are evaluated through extensive simulations. In Section 6, the methods are illustrated by an analysis of a metabolomics study. Further extensions and related problems are discussed in Section 7. In Section 8, some of the main theorems are proved. The proofs of other theorems as well as technical lemmas, and some further discussions are collected in the online Supplementary Materials.

Throughout our paper, for a vector $a = {(a_{1}, \dots, a_{n})}^{⊤} \in ℝ^{n}$ , we define the $l_{p}$ norm $‖ a ‖_{p} = {(\sum_{i = 1}^{n} a_{i}^{p})}^{1 / p}$ , and the $l_{\infty}$ norm $‖ a ‖_{\infty} = {max}_{1 \leq j \leq n} | a_{i} |$ . $a_{- j} \in ℝ^{n - 1}$ stands for the subvector of a without the j the component. We denote diag(a₁, …, a_n) as the n × n diagonal matrix whose diagonal entries are a₁, …, a_n. For a matrix $A \in ℝ^{p \times q}$ , λi (A) stands for the i-th largest singular value of A and λ_max (A) = λ₁ (A), λ_min (A) = λ_p^q (A). For a smooth function f(x) defined on $ℝ$ , we denote $\dot{f} (x) = d f (x) / d x$ and $\ddot{f} (x) = d^{2} f (x) / d x^{2}$ . Furthermore, for sequences {a_n} and {b_n}, we write a_n = o(b_n) if $lim_{n} a_{n} / b_{n} = 0$ , and write a_n = O(b_n), a_n ≲ b_n or b_n ≳ a_n if there exists a constant C such that a_n ≤ Cb_n for all n. We write a_n ≍ b_n if a_n ≲ b_n and a_n ≳ b_n. For a set A, we denote |A| as its cardinality. Lastly, C, C₀, C₁, … are constants that may vary from place to place.

2. GLOBAL HYPOTHESIS TESTING

In this section, we consider testing the global null hypotheses

H_{0} : β = 0 vs . H_{1} : β \neq 0,

under the logistic regression model with random designs. The global testing problem corresponds to the detection of any associations between the covariates and the outcome.

Our construction of the global testing procedure begins with a bias-corrected estimator built upon a regularized estimator such as the $l_{1}$ -regularized M-estimator. For high-dimensional logistic regression, the $l_{1}$ -regularized M-estimator is defined as

\hat{β} = \underset{β}{argmin} {\frac{1}{n} \sum_{i = 1}^{n} [- y_{i} β^{⊤} X_{i} + log (1 + e^{β^{⊤} X_{i}})] + λ ‖ β ‖_{1}},

(2)

which is the minimizer of a penalized log-likelihood function. Negahban et al. (2010) showed that, when X_i are i.i.d. sub-gaussian, under some mild regularity conditions, standard high-dimensional estimation error bounds for $\hat{β}$ under the $l_{1}$ or $l_{2}$ norm can be obtained by choosing $λ ≍ \sqrt{log p / n}$ . Once we obtain the initial estimator $\hat{β}$ , our next step is to correct the bias of $\hat{β}$ .

For technical reasons, we split the samples so that the initial estimation step and the bias correction step are conducted on separate and independent datasets. Without loss of generality, we assume there are 2n samples, divided into two subsets $D_{1}$ and $D_{2}$ , each with n independent samples. The initial estimator $\hat{β}$ is obtained from $D_{1}$ . In the following, we construct a nearly unbiased estimator $\overset{ˇ}{β}$ based on $\hat{β}$ and the samples from $D_{2}$ , using the generalized LDP approach. Throughout the paper, the samples Z_i = (X_i, Y_i), i = 1, …, n, are from $D_{2}$ , which are independent of $\hat{β}$ . We would like to emphasize that the sample splitting procedure is only used to simplify our theoretical analysis, which does not make it a restriction for practical applications. Numerically, as our simulations in Section 5 show, sample splitting is in fact not needed in order for our methods perform well (see further discussions in Section 7).

2.1. Construction of the Test Statistic via Generalized Low-Dimensional Projection

Let X be the design matrix whose i-th row is X_i. We rewrite the logistic regression model defined by (1) as

y_{i} = f (β^{⊤} X_{i}) + ϵ_{i}

(3)

where f (u) = e^u / (1 + e^u) and ϵ_i is error term. To correct the bias of the initial estimator $\hat{β}$ , we consider the Taylor expansion of f (u_i) at ${\hat{u}}_{i}$ for u_i = β^⊤ X_i and ${\hat{u}}_{i} = {\hat{β}}^{⊤} X_{i}$

f (u_{i}) = f ({\hat{u}}_{i}) + \dot{f} ({\hat{u}}_{i}) (u_{i} - {\hat{u}}_{i}) + R e_{i}

where Re_i is the reminder term. Plug this into the regression model (3), we have

y_{i} - f ({\hat{u}}_{i}) + \dot{f} ({\hat{u}}_{i}) X_{i}^{⊤} \hat{β} = \dot{f} ({\hat{u}}_{i}) X_{i}^{⊤} β + (R e_{i} + ϵ_{i}) .

(4)

By rewriting the logistic regression model as (4), we can treat $y_{i} - f ({\hat{u}}_{i}) + \dot{f} ({\hat{u}}_{i}) X_{i}^{⊤} \hat{β}$ on the left hand side as the new response variable, whereas $\dot{f} ({\hat{u}}_{i}) X_{i}$ as the new covariates and Re_i + ϵ_i as the noise. Consequently, β can be considered as the regression coefficient of this approximate linear model.

The bias-corrected estimator, or, the generalized LDP estimator $\overset{ˇ}{β}$ is defined as

{\overset{ˇ}{β}}_{j} = {\hat{β}}_{j} + \frac{\sum_{i = 1}^{n} v_{i j} (y_{i} - f ({\hat{β}}^{⊤} X_{i}))}{\sum_{i = 1}^{n} v_{i j} \dot{f} ({\hat{β}}^{⊤} X_{i}) X_{i j}}, j = 1, \dots, p,

(5)

where X_ij is the j-th component of X_i and v_j = (v_1j, v_2j, …, v_nj) is the score vector that will be determined carefully (Ren, Zhang, and Zhou 2016; Cai et al. 2017). More specifically, we define the weighted inner product 〈·,·〉_n for any $a, b \in ℝ^{n}$ as ${〈 a, b 〉}_{n} = \sum_{i = 1}^{n} \dot{f} ({\hat{u}}_{i}) a_{i} b_{i}$ , and denote 〈·,·〉 as the ordinary inner product defined in Euclidean space. Combining (4) and (5), we can write

{\overset{ˇ}{β}}_{j} - β_{j} = \frac{〈 v_{j}, ϵ 〉}{{〈 v_{j}, x_{j} 〉}_{n}} + \frac{〈 v_{j}, R e 〉}{{〈 v_{j}, x_{j} 〉}_{n}} - \frac{{〈 v_{j}, h_{- j} 〉}_{n}}{{〈 v_{j}, x_{j} 〉}_{n}},

(6)

where $x_{j} \in ℝ^{n}$ denote the j-th column of X, $h_{- j} = X_{- j} ({\hat{β}}_{- j} - β_{- j})$ where $X_{- j} \in ℝ^{n} \times ℝ^{p - 1}$ is the submatrix of X without the j-th column, and Re = (Re₁, …, Re_n) with $R e_{i} = f (u_{i}) - f ({\hat{u}}_{i}) - \dot{f} ({\hat{u}}_{i}) (u_{i} - {\hat{u}}_{i})$ . We will construct score vector v_j so that the first term on the right hand side of (6) is asymptotically normal, while the second and third terms, which together contribute to the bias of the generalized LDP estimator ${\hat{β}}_{j}$ , are negligible.

To determine the score vector v_j efficiently, we consider the following node-wise regression among the covariates

x_{j} = X_{- j} γ_{j} + η_{j}, j = 1, \dots, p,

(7)

where $γ_{j} = arg {min}_{γ \in ℝ^{p - 1}} E [‖ x_{j} - X_{- j} γ ‖_{2}^{2}]$ and η_j is the error term. Intuitively, if we set $v_{j} = {\hat{W}}^{- 1} η_{j}$ for $\hat{W} = diag (\dot{f} ({\hat{u}}_{1}), \dots, \dot{f} ({\hat{u}}_{n}))$ , then it should follow that

{〈 v_{j}, h_{- j} 〉}_{n} \leq max_{k \neq j} | {〈 v_{j}, x_{k} 〉}_{n} | \cdot ‖ \hat{β} - β ‖_{1} = max_{k \neq j} | 〈 η_{j}, x_{k} 〉 | \cdot ‖ \hat{β} - β ‖_{1} \approx 0.

In practice, we use the node-wise Lasso to obtain an estimate of η_j. For X from $D_{2}$ and $\hat{β}$ obtained from $D_{1}$ , the score v_j is obtained by calibrating the Lasso-generated residue ${\hat{η}}_{j}$ , i.e.

v_{j} (λ) = {\hat{W}}^{- 1} {\hat{η}}_{j} (λ), {\hat{η}}_{j} (λ) = x_{j} - X_{- j} {\hat{γ}}_{j} (λ), {\hat{γ}}_{j} (λ) = \underset{b}{argmin} {\frac{‖ x_{j} - X_{- j} b ‖_{2}^{2}}{2 n} + λ ‖ b ‖_{1}} .

(8)

Clearly, v_j (λ) depends on the tuning parameter λ. Define the following quantities

ζ_{j} (λ) = max_{k \neq j} \frac{| {〈 v_{j} (λ), x_{k} 〉}_{n} |}{‖ v_{j} (λ) ‖_{n}}, τ_{j} (λ) = \frac{‖ v_{j} (λ) ‖_{n}}{| {〈 v_{j} (λ), x_{j} 〉}_{n} |} .

(9)

The tuning parameter λ can be determined through ζ_j (λ) and τ_j (λ) by the algorithm in Table 1, which is adapted from the algorithm in Zhang and Zhang (2014).

Table 1.

Computation of v_j from the Lasso (8)

Input:	An upper bound $ζ_{j}^{}$ for ζ_j with default value $ζ^{} = \sqrt{2 log p}$ ,
	tuning parameters κ₀ ∈[0,1] and κ₁ ∈(0,1];
Step 1:	If $ζ_{j} (λ) > ζ_{j}^{}$ for all λ = 0, set $ζ_{j}^{} = (1 + κ_{1}) inf_{λ > 0} ζ_{j} (λ)$ ;
	$λ \leftarrow max {λ : ζ_{j} (λ) \leq ζ_{j}^{}}, ζ_{j}^{} \leftarrow ζ_{j} (λ), τ_{j}^{*} \leftarrow τ_{j} (λ);$
Step 2:	$λ_{j} \leftarrow min {λ : τ_{j} (λ) \leq (1 + κ_{0}) τ_{j}^{*}};$
	v_i ← v_j(λ_j), τ_j ← τ_j(λ_j), ζ_j ← ζ_j(λ_j)
Input:	An upper bound $ζ_{j}^{}$ for ζ_j with default value $ζ^{} = \sqrt{2 log p}$ ,
Output:	λ_j, v_j, τ_j, ζ_j

Open in a new tab

Once ${\overset{ˇ}{β}}_{j}$ and τ_j are obtained, we define the standardized statistics

M_{j} = {\overset{ˇ}{β}}_{j} / τ_{j},

for j = 1, …, p. The global test statistic is then defined as

M_{n} = max_{1 \leq j \leq p} M_{j}^{2} .

(10)

2.2. Asymptotic Null Distribution

We now turn to the analysis of the properties of the global test statistic M_n defined in (10). For the random covariates, we consider both the Gaussian design and the bounded design. Under the Gaussian design, the covariates are generated from a multivariate Gaussian distribution with an unknown covariance matrix $Σ \in ℝ^{p \times p}$ . In this case, we assume

(A1). X_i ~ N(0, Σ) independently for each i = 1, …, n.

In the case of bounded design, we assume instead

(A2). X_i for i = 1, …, n are i.i.d. random vectors satisfying $E X_{i} = 0$ and max_1≤i≤n ‖ X_i‖_∞ ≤ T for some constant T > 0.

Define the $l_{1}$ ball

B_{1} (k) = {Ω = (ω_{i j}) \in ℝ^{p \times p} : max_{1 \leq i \leq p} \sum_{j = 1}^{p} min (| ω_{i j} | \sqrt{\frac{n}{log p}}, 1) \leq k} .

In general, $B_{1} (k)$ includes any matrix Ω whose rows ω_i are $l_{0}$ sparse with ‖ω_i‖₀≤ k or $l_{1}$ sparse with $‖ ω_{i} ‖_{1} \leq k \sqrt{log p / n}$ for all i = 1, …, p. The parameter space of the covariance matrix Σ and the regression vector β are defined as following.

(A3). The parameter space Θ(k) of $θ = (β, Σ) \in ℝ^{p} \times ℝ^{p \times p}$ satisfies

Θ (k) = {(β, Σ) : ‖ β ‖_{0} \leq k, M^{- 1} \leq λ_{min} (Σ) \leq λ_{max} (Σ) \leq M, Σ^{- 1} \in B_{1} (k)},

for some constant M ≥ 1. For convenience, we denote $Θ_{1} (k) = {β \in ℝ^{p} : ‖ β ‖_{0} \leq k}$ and $Θ_{2} (k) = {Σ \in ℝ^{p \times p} : M^{- 1} \leq λ_{min} (Σ) \leq λ_{max} (Σ) \leq M, Σ^{- 1} \in B_{1} (k)}$ , so that Θ(k) = Θ₁(k) × Θ₂(k).

The following theorem states that the asymptotic null distribution of M_n under either the Gaussian or bounded design is a Gumbel distribution.

Theorem 1. Let M_n be the test statistic defined in (10), D be the diagonal of Σ⁻¹ and (ζ_ij) = D^−1/2Σ⁻¹D^−1/2. Suppose max_1≤i<j≤p |ζ_ij|≤ c₀ for some constant 0 < c₀ < 1, log p = O(n^r) for some 0 < r < 1/5, and

1. under the Gaussian design, we assume (A1) (A3) and $k = o (\sqrt{n} / {log}^{3} p)$ ; or

2. under the bounded design, we assume (A2) (A3) and $k = o (\sqrt{n} / {log}^{5 / 2} p)$ .

Then under H₀, for any given $x \in ℝ$ ,

P_{θ} (M_{n} - 2 log p + loglog p \leq x) \to exp (\frac{1}{\sqrt{π}} exp (- x / 2)), as (n, p) \to \infty .

The condition that log p = o(n^r) for some 0 < r < 1/5 is consistent with those required for testing the global hypothesis in high-dimensional linear regression (Xia, Cai, and Cai 2018) and for testing two-sample covariance matrices (Cai, Liu, and Xia 2013). It allows the dimension p to be exponentially large comparing to the sample size n, which is much more flexible than the likelihood ratio test considered in Sur, Chen, and Candès (2017) and Sur and Candès (2019), where the dimension can only scale as p < n. Under the Gaussian design, it is required that the sparsity k is $o (\sqrt{n} / {log}^{3} p)$ whereas for the bounded design, it suffices that the sparsity k to be $o (\sqrt{n} / {log}^{5 / 2} p)$ .

Remark 1. The analysis can be extended to testing H₀:β_G = 0 versus H₁:β_G ≠ 0 for a given index set G. Specifically, we can construct the test statistic as $M_{G, n} = {max}_{i \in G} M_{j}^{2}$ and obtain a similar Gumbel limiting distribution by replacing p by | G |, as (n,|G|) → ∞. The sparsity condition thus should be forwarded to the set G.

Based on the limiting null distribution, the asymptotically α level test can be defined as

Φ_{α} (M_{n}) = I {M_{n} \geq 2 log p - loglog p + q_{α}},

where q_α is the 1 − α quantile of the Gumbel distribution with the cumulative distribution function $exp (- \frac{1}{\sqrt{π}} exp (- x / 2))$ , i.e.

q_{α} = - log (π) - 2 loglog {(1 - α)}^{- 1} .

The null hypothesis H₀ is rejected if and only if Φ_α(M_n) = 1.

2.3. Minimax Separation Distance and Optimality

In this subsection, we answer the question: “What is the essential difficulty for testing the global hypothesis in logistic regression.” To fix ideas, we begin with defining the minimax separation distance that measures such an essential difficulty for testing the global null hypothesis at a given level and type II error. In particular, we consider the alternative

H_{1} : β \in {β \in ℝ^{p} : ‖ β ‖_{\infty} \geq ρ, ‖ β ‖_{0} \leq k}

for some ρ > 0. This alternative concerns the detection of any discernible signals among the regression coefficients where the signals can be extremely sparse, which has interesting applications (see Xia, Cai, and Cai (2015)). Similar alternatives are also considered by Cai, Liu, and Xia (2013) and Cai, Liu, and Xia (2014).

By fixing a level α > 0 and a type II error probability δ > 0, we can define the δ-separation distance of a level α test procedure Φ_α for given design covariance Σ as

ρ (Φ_{α}, δ, Σ) = inf {ρ > 0 : inf_{β \in Θ_{1} (k) ‖ β ‖_{\infty} \geq ρ} P_{θ} (Φ_{α} = 1) \geq 1 - δ} = inf {ρ > 0 : sup_{β \in Θ_{1} (k) ‖ β ‖_{\infty} \geq ρ} P_{θ} (Φ_{α} = 0) \leq δ} .

(11)

The δ-separation distance ρ(Φ_α, δ,Θ(k)) over Θ(k) can thus be defined by taking the supremum over all the covariance matrices Σ ∈Θ₂(k), so that

ρ (Φ_{α}, δ, Θ (k)) = sup_{Σ \in Θ_{2} (k)} ρ (Φ_{α}, δ, Σ),

which corresponds to the minimal $l_{\infty}$ distance such that the null hypothesis H₀ is well separated from the alternative H₁ by the test Φ_α. In general, δ-separation distance is an analogue of the statistical risk in estimation problems. It characterizes the performance of a specific α-level test with a guaranteed type II error δ. Consequently, we can define the (α, δ)-minimax separation distance over Θ(k) and all the α-level tests as

ρ^{*} (α, δ, Θ (k)) = inf_{Φ_{α}} ρ (Φ_{α}, δ, Θ (k)) .

The definition of (α, δ)-minimax separation distance generalizes the ideas of Ingster (1993), Baraud (2002) and Verzelen (2012). The following theorem establishes the minimax lower bound of the (α, δ)-separation distance under the Gaussian design for testing the global null hypothesis over the parameter space Θ′(k) ⊂ Θ(k) defined as

Θ^{'} (k) = (Θ_{1} (k) \cap {β \in ℝ^{p} : ‖ β ‖_{2} ≲ {(n^{1 / 4} log p)}^{- 1}}) \times Θ_{2} (k) .

Theorem 2. Assume that α + δ ≤ 1. Under the Gaussian design, if (A1) and (A3) hold, (β, Σ) ∈ Θ′(k) and $k ≲ min {p^{γ}, \sqrt{n} / {log}^{3} p}$ for some 0 < γ <1 / 2 , then the (α, δ)-minimax separation distance over Θ′(k) has the lower bound

ρ^{*} (α, δ, Θ^{'} (k)) \geq c \sqrt{\frac{log p}{n}}

(12)

for some constant c > 0.

In order to show the above lower bound is asymptotically sharp, we prove that it is actually attainable under certain circumstances, by our proposed global test Φ_α. In particular, for the bounded design, we make the following additional assumption.

(A4). It holds that P_θ(max_1≤i≤n |β^⊤X_i|≥ C) = O(p^−c) for some constant C, c > 0.

Theorem 3. Suppose that log p = O(n^r) for some 0 < r < 1. Under the alternative $H_{1} : ‖ β ‖_{\infty} \geq c_{2} \sqrt{log p / n}$ for some c₂ > 0, and

(i) under the Gaussian design, assume that (A1) and (A3) hold, $‖ β ‖_{2} \leq C (loglog p) / \sqrt{log n}$ for $C \leq min {\sqrt{2 / λ_{max} (Σ)}, {(2 r \sqrt{2 λ_{max} (Σ)})}^{- 1}}$ , log p ≳ log^1+δ n for some δ > 0 and $k = o (\sqrt{n} / {log}^{3} p)$ ; or

(ii) under the bounded design, assume that (A2), (A3), and (A4) hold, and $k = o (\sqrt{n} / {log}^{5 / 2} p)$ .

Then we have P_θ(Φ_α(M_n) = 1) → 1 as (n, p) → ∞.

In Theorem 3, (A4) is assumed for the bounded case and $‖ β ‖_{2} = O (loglog p / \sqrt{log n})$ is required for the Gaussian case. In particular, since log p = O(n^r) for some 0 < r < 1, the upper bound $loglog p / \sqrt{log n}$ for ‖β‖₂ can be as large as $\sqrt{log n}$ . In Theorem 2, the minimax lower bound is established over (β, Σ) ∈Θ′(k), so that the same lower bound holds over a larger set

(β, Σ) \in (Θ_{1} (k) \cap {β \in ℝ^{p} : ‖ β ‖_{2} \leq loglog p / \sqrt{log n}}) \times Θ_{2} (k),

(13)

since $loglog p / \sqrt{log n} ≳ {(n^{1 / 4} log p)}^{- 1}$ . On the other hand, Theorem 3 (i) indicates an upper bound $ρ^{*} ≲ \sqrt{log p / n}$ attained by our proposed test under the Gaussian design over the set (13). These two results imply the minimax rate $ρ^{*} ≍ \sqrt{log p / n}$ and the minimax optimality of our proposed test over the set (13).

2.4. Comparison with Existing Works

In this section, we make detailed comparisons and connections with some existing works concerning global hypothesis testing in the high-dimensional regression literature.

Ingster, Tsybakov, and Verzelen (2010) addressed the detection boundary for high-dimensional sparse linear regression models, and more recently Mukherjee, Pillai, and Lin (2015) studied the detection boundary for hypothesis testing in high-dimensional sparse binary regression models. However, although both works obtained the sharp detection boundary for the global testing problem H₀:β = 0, their alternative hypotheses are different from ours. Specifically, Mukherjee, Pillai, and Lin (2015) considered the alternative hypothesis $H_{1} : β \in {β \in ℝ^{p} : ‖ β ‖_{0} \geq k, min {| β_{j} | : β_{k} \neq 0} \geq A}$ , which implies that β has at least k nonzero coefficients exceeding A in absolute values. Ingster, Tsybakov, and Verzelen (2010) considered the alternative hypothesis $H_{1} : β \in {β \in ℝ^{p} : ‖ β ‖_{0} \leq k, ‖ β ‖_{2} \geq ρ}$ , which concerns k sparse β with $l_{2}$ norm at least ρ. In fact, the proof of our Theorem 2 can be directly extended to such an alternative concerning the $l_{2}$ norm, which amounts to obtaining a lower bound of order $\sqrt{\frac{k log p}{n}}$ for high dimensional logistic regression. However, developing a minimax optimal test for such alternative is beyond the scope of the current paper.

Additionally, in contrast to the minimax separation distance considered in this paper, the papers by Ingster, Tsybakov, and Verzelen (2010) and Mukherjee, Pillai, and Lin (2015) considered the minimax risk (or the minimax total error probability) given by

inf_{Φ} sup_{Σ \in Θ_{2} (k)} Risk (Φ, Σ) = inf_{Φ} sup_{Σ \in Θ_{2} (k)} {max_{β \in H_{0}}} P_{θ} (Φ = 1) + max_{β \in Θ_{1} (k) {‖ β ‖}_{\infty} \geq ρ} P_{θ} (Φ = 0)},

(14)

where the infimum is taken over all tests Φ. This minimax risk can be also written as

inf_{Φ} sup_{Σ \in Θ_{2} (k)} Risk (Φ, Σ) = inf_{α \in (0, 1)} {α + inf_{Φ_{α}} sup_{Σ \in Θ_{2} (k)} sup_{β \in Θ_{1} (k) ‖ β ‖_{\infty} \geq ρ} P_{θ} (Φ_{α} = 0)} .

(15)

A comparison of (11) and (15) yields the slight difference between the two criteria, as one depends on a given Type I error α and the other doesn’t.

Moreover, these two papers considered different design scenarios from ours. In Ingster, Tsybakov, and Verzelen (2010), only the isotropic Gaussian design was considered. As a result, the optimal tests proposed therein rely highly on the independence assumption. In Mukherjee, Pillai, and Lin (2015), the general binary regression was studied under fixed sparse design matrices. In particular, the minimax lower and upper bounds were only derived in the special case of design matrices with binary entries and certain sparsity structures.

In comparison with the recent works of Sur, Chen, and Candès (2017), Candès and Sur (2018) and Sur and Candès (2019), besides the aforementioned difference in the asymptotics of (p, n), these two papers only considered the random Gaussian design, whereas our work also considered random bounded design as in van de Geer et al. (2014). In addition, Sur, Chen, and Candès (2017) and Sur and Candès (2019) developed the Likelihood Ratio (LLR) Test for testing the hypothesis $H_{0} : β_{j_{1}} = β_{j_{2}} = \dots = β_{j_{k}} = 0$ for any finite k. Intuitively, a valid test for the global null and p / n → κ ∈(0, 1 / 2) can be adapted from the individual LLR tests using the Bonferroni procedure. However, as our simulations show (Section 5), such a test is less powerful compared to our proposed test.

Lastly, our minimax results focus on the highly sparse regime k ≲ p^γ where γ ∈(0, 1 / 2). As shown by Ingster, Tsybakov, and Verzelen (2010) and Mukherjee, Pillai, and Lin (2015), the problem under the dense regime where γ ∈(1 / 2,1) can be very different from the sparse regime. Mostly likely, the fundamental difficulty of the testing problem changes in this situation so that different methods need to be carefully developed. We leave these interesting questions for future investigations.

3. LARGE-SCALE MULTIPLE TESTING

Denote by β the true coefficient vector in the model and denote $H_{0} = {j : β_{j} = 0, j = 1, \dots, p}, H_{1} = {j : β_{j} \neq 0, j = 1, \dots, p}$ . In order to identify the indices in $H_{1}$ , we consider simultaneous testing of the following null hypotheses

H_{0, j} : β_{j} = 0 vs . H_{1, j} : β_{j} \neq 0, 1 \leq j \leq p .

Apart from identifying as many nonzero β_j as possible, to obtain results of practical interest, we would like to control the false discovery rate (FDR) as well as the false discovery proportion (FDP), or the number of falsely discovered variables (FDV).

3.1. Construction of Multiple Testing Procedures

Recall that in Section 2, we define the standardized statistics $M_{j} = {\overset{ˇ}{β}}_{j} / τ_{j}$ , for j = 1, …, p. For a given threshold level t > 0, each individual hypothesis H_0,j:β_j = 0 is rejected if |M_j|≥ t. Therefore for each t, we can define

{FDP}_{θ} (t) = \frac{\sum_{j \in H_{0}} I {| M_{j} | \geq t}}{max {\sum_{j = 1}^{p} I {| M_{j} | \geq t}, 1}}, {FDR}_{θ} (t) = E_{θ} [FDP (t)],

and the expected number of falsely discovered variables ${FDV}_{θ} (t) = E_{θ} [\sum_{j \in H_{0}} I {| M_{j} | \geq t}]$

Procedure Controlling FDR/FDP.

In order to control the FDR/FDP at a pre-specified level 0 < α < 1, we can set the threshold level as

{\tilde{t}}_{1} = inf {0 \leq t \leq b_{p} : \frac{\sum_{j \in H_{0}} I {| M_{j} | \geq t}}{max {\sum_{j = 1}^{p} I {| M_{j} | \geq t}, 1}} \leq α},

(16)

for some b_p to be determined later.

In general, the ideal choice ${\tilde{t}}_{1}$ is unknown and needs to be estimated because it depends on the knowledge of the true null $H_{0}$ . Let G₀(t) be the proportion of the nulls falsely rejected by the procedure among all the true nulls at the threshold level t, namely, $G_{0} (t) = \frac{1}{p_{0}} \sum_{j \in H_{0}} I {| M_{j} 〉 \geq t}$ , where $p_{0} = | H_{0} |$ . In practice, it is reasonable to assume that the true alternatives are sparse. If the sample size is large, we can use the tails of normal distribution G(t) = 2 − 2Φ(t) to approximate G₀(t). In fact, it will be shown that, for $b_{p} = \sqrt{2 log p - 2 loglog p}$ in probability as (n, p) → ∞. To summarize, we have the following logistic multiple testing (LMT) procedure controlling the FDR and the FDP.

Procedure 1 (LMT). Let 0 < α < 1, $b_{p} = \sqrt{2 log p - 2 loglog p}$ and define

\hat{t} = inf {0 \leq t \leq b_{p} : \frac{p G (t)}{max {\sum_{j = 1}^{p} I {| M_{j} | \geq t}, 1}} \leq α} .

(17)

If $\hat{t}$ in (17) does not exist, then let $\hat{t} = \sqrt{2 log p}$ . We reject H_0,j Whenever $| M_{j} | \geq \hat{t}$ .

Procedure Controlling FDV.

For large-scale inference, it is sometimes of interest to directly control the number of falsely discovered variables (FDV) instead of the less stringent FDR/FDP, especially when the sample size is small (Liu and Luo 2014). By definition, the FDV control, or equivalently, the per-family error rate control, provides an intuitive description of the Type I error (false positives) in variable selection. Moreover, controlling FDV = r for some 0 < r < 1 is related to the family-wise error rate (FWER) control, which is the probability of at least one false positive. In fact, FDV control can be achieved by a suitable modification of the FDP controlling procedure introduced above. Specifically, we propose the following FDV (or FWER) controlling logistic multiple testing (LMT_V) procedure.

Procedure 2 (LMT_V). For a given tolerable number of falsely discovered variables r < p (or a desired level of FWER 0 < r < 1), let ${\hat{t}}_{F D V} = G^{- 1} (r / p)$ . H_0,j is rejected whenever $| M_{j} | \geq {\hat{t}}_{F D V}$ .

3.2. Theoretical Properties for Multiple Testing Procedures

In this section we show that our proposed multiple testing procedures control the theoretical FDR/FDP or FDV asymptotically. For simplicity, our theoretical results are obtained under the bounded design scenario. For FDR/FDP control, we need an additional assumption on the interplay between the dimension p and the parameter space Θ(k).

Recall that η_j = (η_j1, …, η_jn) for j = 1, …, p defined in (7). We define $F_{j k} = E_{θ} [η_{i j} η_{i k} / \dot{f} (u_{i})]$ for 1 ≤ j, k ≤ p, and $ρ_{j k} = F_{j k} / \sqrt{F_{i j} F_{k k}}$ . Denote $B (δ) = {(j, k) : | ρ_{j k} | \geq δ, i \neq j}$ and $A (ϵ) = B ({(log p)}^{- 2 - ϵ})$ .

(A5). Suppose that for some ϵ > 0 and q > 0, $\sum_{(j, k) \in A (ϵ) : j, k \in H_{0}} p^{\frac{2 | ρ_{j k} |}{1 + | ρ_{j k} |} + q} = O (p^{2} / {(log p)}^{2})$ .

The following proposition shows that M_j is asymptotically normal distributed and G₀(t) is well approximated by G(t).

Proposition 1. Under (A2) (A3) and (A4), suppose p = O(n^c) for some constant c > 0, $k = o (\sqrt{n} / {log}^{5 / 2} p)$ , then as (n, p) → ∞,

sup_{j \in H_{0}} sup_{0 \leq t \leq \sqrt{2 log p}} | \frac{P_{θ} (| M_{j} | \geq t)}{2 - 2 Φ (t)} - 1 | \to 0.

(18)

If in addition we assume (A5), then

sup_{0 \leq t \leq b_{p}} | \frac{G_{0} (t)}{G (t)} - 1 | \to 0

(19)

in probability, where Φ is the cumulative distribution function of the standard normal distribution and $b_{p} = \sqrt{2 log p - 2 loglog p}$ .

The following theorem provides the asymptotic FDR and FDP control of our procedure.

Theorem 4. Under the conditions of Proposition 1, for $\hat{t}$ defined in our LMT procedure, we have

lim_{(n, p) \to \infty} \frac{{FDR}_{θ} (\hat{t})}{α p_{0} / p} \leq 1, lim_{(n, p) \to \infty} P_{θ} (\frac{{FDP}_{θ} (\hat{t})}{α p_{0} / p} \leq 1 + ϵ) = 1

(20)

for any ϵ > 0.

For the FDV/FWER controlling procedure, we have the following theorem.

Theorem 5. Under (A2) (A3) and (A4), assume p = O(n^c) for some c > 0 and $k = o (\sqrt{n} / {log}^{5 / 2} p)$ .

Let r < p be the desired level of FDV. For ${\hat{t}}_{F D V}$ defined in our LMT_Vprocedure, we have $lim_{(n, p) \to \infty} \frac{{FDV}_{θ} ({\hat{t}}_{F D V})}{r p_{0} / p} \leq 1$ . In addition, if 0 < r < 1, we have $lim_{(n, p) \to \infty} \frac{{FWER}_{θ} ({\hat{t}}_{F D V})}{r p_{0} / p} \leq 1$ .

The above theoretical results are obtained under the dimensionality condition p = O(n^c), which is stronger than that of the global test. Essentially, the condition is needed to obtain the uniform convergence (18), whose form (as ratio) is stronger than the convergence in distribution in the ordinary sense (as direct difference).

4. TESTING FOR TWO LOGISTIC REGRESSION MODELS

In some applications, it is also interesting to consider hypothesis testing that involves two separate logistic regression models of the same dimension. Specifically, for $l = 1, 2$ and $i = 1, \dots, n_{l}$ , where $n_{1} ≍ n_{2}, y_{i}^{(l)} = f (β^{(l) ⊤} X_{i}^{(l)}) + ϵ_{i}^{(l)}$ , where f (u) = e^u / (1 + e^u), and $ϵ_{i}^{(l)}$ is a binary random variable such that $y_{i}^{(l)} | X_{i}^{(l)} ~ Bernoulli (f (β^{(l) ⊤} X_{i}^{(l)}))$ . The global null hypothesis H₀:β⁽¹⁾ = β⁽²⁾ implies that there is overall no difference in association between covariates and the response. If this null hypothesis is rejected, we are interested in simultaneously testing the hypotheses $H_{0, j} : β_{j}^{(1)} = β_{j}^{(2)}$ for each j = 1, …, p.

To test the global null H₀:β⁽¹⁾ = β⁽²⁾ against H₁:β⁽¹⁾ ≠ β⁽²⁾, we can first obtain ${\overset{ˇ}{β}}_{j}^{(l)}$ and $τ_{j}^{(l)}$ for each model, and then calculate the coordinate-wise standardized statistics $T_{j} = \frac{{\overset{ˇ}{β}}_{j}^{(1)}}{\sqrt{2} τ_{j}^{(1)}} - \frac{{\overset{ˇ}{β}}_{j}^{(2)}}{\sqrt{2} τ_{j}^{(2)}}$ , for j = 1, …, p. Define the global test statistic as $T_{n} = {max}_{1 \leq j \leq p} T_{j}^{2}$ , it can be shown that the limiting null distribution is also a Gumbel distribution. The α level global test is thus defined as Φ_α(T_n) = I{T_n ≥ 2log p−loglog p + q_α}, where q_α = −log(π) − 2loglog(1 − α)⁻¹. For multiple hypotheses testing of two regression vectors $H_{0, j} : β_{j}^{(1)} = β_{j}^{(2)}$ , we consider the test statistics T_j defined above. The two-sample multiple testing procedure controlling FDR/FDP is given as follows.

Procedure 3. Let 0 < α < 1 and define $\hat{t} = inf {0 \leq t \leq b_{p} : \frac{p G (t)}{max {\sum_{j = 1}^{p} I {| T_{j} | \geq t}, 1}} \leq α}$ . If the above $\hat{t}$ does not exist, let $\hat{t} = \sqrt{2 log p}$ . We reject H_0,j whenever $| T_{j} | \geq \hat{t}$ .

5. SIMULATION STUDIES

In this section we examine the numerical performance of the proposed tests. Due to the space limit, for both global and multiple testing problems, we only focus on the single regression setting, and report the results on two logistic regressions in the Supplementary Materials. Throughout our numerical studies, sample splitting was not used.

5.1. Global Hypothesis Testing

In the following simulations, we consider a variety of dimensions, sample sizes, and sparsity of the models. For alternative hypotheses, the dimension of the covariates p ranges from 100, 200, 300 to 400, and the sparsity k is set as 2 or 4. The sample sizes n are determined by the ratio r = p / n that takes values of 0.2, 0.4 and 1.2. To generate the design matrix X, we consider the Gaussian design with the blockwise-correlated covariates so that Σ = Σ_B, where Σ_B is a p × p blockwise diagonal matrix including 10 equal-sized blocks, whose diagonal elements are 1’s and off-diagonal elements are set as 0.7. Under the alternative, suppose $S$ is the support of the regression coefficients β and $| S | = k$ , we set $| β_{j} | = ρ 1 {j \in S}$ for j = 1, …, p and ρ = 0.75 with equal proportions of ρ and −ρ. We set κ₀ = 0 and κ₁ = 0.5.

To assess the empirical performance of our proposed test (“Proposed”), we compare our test with (i) a Bonferroni procedure applied to the p-values from univariate screening using MLE statistic (“U-S”), and (ii) to the method of Sur, Chen, and Candès (2017); Sur and Candès (2019) (“LLR”) in the setting where r = 0.2 and 0.4.

Table 2 shows the empirical type I errors of these tests at level α = 0.05 based on 1000 simulations. Figure 1 shows the corresponding empirical powers under various settings. As we expected, our proposed method outperforms the other two alternatives in all the cases (including the moderate dimensional cases where r = 0.2 and 0.4), and the power increases as n or p grows. In the rather lower dimensional setting where r = 0.2, the LLR performs almost as well as our proposed method.

Table 2.

Type I error with α = 0.05 for the proposed method (Proposed), the Bonferroni corrected univariate screening method (U-S) and the Bonferroni corrected likelihood ratio based method of Sur and Candès (2019) (LLR), for different n, p and k.

p / n	k = 2				k = 4
	p = 100	200	300	400	p = 400	600	800	1000
Proposed
0.2	0.052	0.066	0.042	0.054	0.058	0.050	0.046	0.070
0.4	0.038	0.054	0.062	0.054	0.046	0.050	0.060	0.074
1.2	0.026	0.044	0.042	0.045	0.014	0.044	0.054	0.054
U-S
0.2	0.040	0.032	0.024	0.018	0.018	0.022	0.028	0.034
0.4	0.050	0.032	0.024	0.020	0.028	0.028	0.032	0.046
1.2	0.028	0.038	0.024	0.020	0.032	0.018	0.034	0.014
LLR
0.2	0.050	0.050	0.068	0.040	0.058	0.044	0.046	0.034
0.4	0.084	0.070	0.048	0.056	0.062	0.042	0.058	0.064

Open in a new tab

Fig. 1 — Empirical power with α = 0.05 for the proposed method (Proposed), the Bonferroni corrected univariate screening method (U-S) and the Bonferroni corrected likelihood ratio based method of Sur and Candès (2019) (LLR). Top panel: k = 2; bottom panel: k = 4.

5.2. Multiple Hypotheses Testing

FDR Control.

In this case, we set p = 800 and let n vary from 600, 800, 1000, 1200 to 1400, so that all the cases are high-dimensional in the sense that p > n / 2. The sparsity level k varies from 40, 50 to 60. For the true positives, given the support $S$ such that $| S | = k$ , we set $| β_{j} | = ρ 1 {j \in S}$ for j = 1, …, p with equal proportions of ρ and −ρ. The design covariates X_i’s are generated from a $(| X_{i}^{⊤} β | < 3)$ -truncated multivariate Gaussian distribution with covariance matrix Σ = 0.01Σ_M, where Σ_M is a p × p blockwise diagonal matrix of 10 identical unit diagonal Toeplitz matrices whose off-diagonal entries descend from 0.1 to 0 (see Supplementary Material for the explicit form). The choice of κ₀ and κ₁ are the same as the global testing. Throughout, we set the desired FDR level as α = 0.2.

We compare our proposed procedure (denoted as “LMT”) with following methods: (i) the basic LMT procedure with b_p in (17) replaced by ∞ (“LMT0”), which is equivalent to applying the BH procedure (Benjamini and Hochberg 1995) to our debiased statistics M_j, (ii) the BY procedure (Benjamini and Yekutieli 2001) using our debiased statistics M_j (“BY”), implemented using the R function p.adjust(…,method=“BY”), (iii) a BH procedure applied to the p-values from univariate screening using the MLE statistics (“U-S”), and (iv) the knockoff method of Candès et al. (2018) (“Knockoff”). Figure 2 shows boxplots of the pooled empirical FDRs (see Supplementary Material for the case-by-case FDRs) and Figure 3 shows the empirical powers of these methods based on 1000 replications. Here the power is defined as the number of correctly discovered variables divided by the number of truly associated variables. As a result, we find that LMT and LMT0 correctly control FDRs and have the greatest power among all the cases. In particular, the power of LMT and LMT0 are almost the same, which increases as the sparsity decreases, the signal magnitude ρ increases, or the sample size n increases, while LMT0 has slightly inflated FDRs. The U-S method, although correctly controls the FDRs, has poor power, which is largely due to the dependence among the covariates.

Fig. 2 — Boxplots of the empirical FDRs across all the settings for α = 0.2.

Fig. 3 — Empirical power under FDR α = 0.2 for ρ = 3 (top) and ρ = 4 (bottom).

FDV Control.

For our proposed test that controls FDV (denoted as LMT_V), by setting desired FDV level r = 10, we apply our method to various settings. Specifically, we set ρ = 3, p ∈{800,1000,1200}, set k ∈{40,50,60}, and let n vary from 400, 600, 800 to 1000. The design covariates are generated similarly as the previous part. The resulting empirical FDV and powers are summarized in Table 3. Our proposed LMT_V has the correct control of FDV in all the settings and the power increases as n grows, k decreases, or p decreases.

Table 3.

Empirical performance of LMT_V with FDV level r = 10.

ρ	p	k	Empirical FDV				Empirical Power
			n = 400	600	800	1000	400	600	800	1000
		40	4.07	5.45	6.44	7.11	0.08	0.23	0.40	0.59
	800	50	4.30	6.29	7.27	8.26	0.06	0.16	0.32	0.49
		60	4.33	6.63	7.48	8.42	0.05	0.12	0.25	0.42
		40	3.30	4.59	5.79	6.82	0.06	0.18	0.35	0.52
3	1000	50	3.49	5.42	6.43	7.03	0.05	0.13	0.26	0.43
		60	3.68	5.47	7.29	7.97	0.03	0.09	0.20	0.34
		40	2.69	4.36	5.00	5.68	0.05	0.15	0.31	0.46
	1200	50	2.97	4.22	5.73	6.43	0.03	0.11	0.21	0.36
		60	2.78	4.91	5.91	7.25	0.02	0.07	0.16	0.27

Open in a new tab

6. REAL DATA ANALYSIS

We illustrate our proposed methods by analyzing a dataset from the Pediatric Longitudinal Study of Elemental Diet and Stool Microbiome Composition (PLEASE) study, a prospective cohort study to investigate the effects of inflammation, antibiotics, and diet as environmental stressors on the gut microbiome in pediatric Crohn’s disease (Lewis et al. 2015; Lee et al. 2015; Ni et al. 2017). The study considered the association between pediatric Crohn’s disease and fecal metabolomics by collecting fecal samples of 90 pediatric patients with Crohn’s disease at baseline, 1 week, and 8 weeks after initiation of either anti-tumor necrosis factor (TNF) or enteral diet therapy, as well as those from 25 healthy control children (Lewis et al. 2015). In details, an untargeted fecal metabolomic analysis was performed on these samples using liquid chromatography-mass spectrometry (LC-MS). Metabolites with more than 80% missing values across all samples were removed from the analysis. For each metabolite, samples with the missing values were imputed with its minimum abundance across samples. To avoid potential large outliers, for each sample, the metabolite abundances were further normalized by dividing 90% cumulative sum of the abundances of all metabolites. The normalized abundances were then log transformed and used in all analyses. The metabololomics annotation was obtained from Human Metabolome Database (Lee et al. 2015). In total, for each sample, abundances of 335 known metabolites were obtained and used in our analysis.

6.1. Association Between Metabolites and Crohn’s Disease Before and After Treatment

We first test the overall association between 335 characterized metabolites and Crohn’s disease by fitting a logistic regression using the data of 25 healthy controls and 90 Crohn’s disease patients at the baseline. We obtain a global test statistic of 433.88 with a p-value < 0.001, indicating a strong association between Crohn’s disease and fecal metabolites. At the FDR < 5%, our multiple testing procedure selects four metabolites, including C14:0.sphingomyelin, C24:1.Ceramide.(d18:1) and 3-methyladipate/pimelate (see Table 4). Recent studies have demonstrated that sphingolipid metabolites, particularly ceramide and sphingosine-1-phosphate, are signaling molecules that regulate a diverse range of cellular processes that are important in immunity, inflammation and inflammatory disorders (Maceyka and Spiegel 2014). In fact, ceramide acts to reduce tumor necrosis factor (TNF) release (Rozenova et al. 2010) and has important roles in the control of autophagy, a process strongly implicated in the pathogenesis of Crohn’s disease (Barrett et al. 2008; Sewell et al. 2012).

Table 4.

Significant metabolites associated with Crohn’s disease (coded as 1 in logistic regression) at the baseline, one week and 8 weeks after treatment with FDR < 5%. The refitted regression coefficients show the direction of the association.

Disease Stage	HMDB ID	Synonyms	Refitted Coefficient
Baseline	00885	C16:0.cholesteryl ester	4.45
	12097	C14:0.sphingomyelin	1.74
	04953	C24:1.Ceramide.(d18:1)	4.25
	00555	3-methyladipate/pimelate	−12.82
Week 1	06726	C20:4.cholesteryl ester	2.17
	12097	C14:0.sphingomyelin	2.06
	04949	C16:0.Ceramide.(d18:1)	0.87
	00555	3-methyladipate/pimelate	−6.10
	00056	beta-alanine	2.95
	00448	adipate	−4.50
Week 8	00883	valine	1.40
	00222	C16.carnitine	0.58
	00848	C18.carnitine	0.39
	00555	3-methyladipate/pimelate	−5.95
	00056	beta-alanine	0.63

Open in a new tab

We next investigate whether treatment of Crohn’s disease alters the association between metabolites and Crohn’s disease by fitting two separate logistic regressions using the metabolites measured one week or 8 weeks after the treatment. At each time point, a significant association is detected based on our global test (p-value < 0.001). One week after the treatment, we observe six metabolites associated with Crohn’s disease, including all four identified at the baseline and two additional metabolites, beta-alanine and adipate (see Table 4). The beta-alanine and adipate associations are likely due to that beta-alanine and adipate are important ingredients of the enteral nutrition treatment of Crohn’s disease. However, it is interesting that at 8 weeks after the treatment, valine, C16.carnitine and C18.carnitine are identified to be associated with Crohn’s disease together with 3-methyladipate/pimelate and beta-alanine. It is known that carnitine plays an important role in Crohn’s disease, which might be a consequence of the underlying functional association between Crohn’s disease and mutations in the carnitine transporter genes (Peltekova et al. 2004; Fortin 2011). Deficiency of carnitine can lead to severe gut atrophy, ulceration and inflammation in animal models of carnitine deficiency (Shekhawat et al. 2013). Our results may suggest that the treatment increases carnitine, leading to reduction of inflammation.

6.2. Comparison of Metabolite Associations Between Responders and Non-Responders

To compare the metabolic association with Crohn’s disease for responders (n = 47) and non-responders (n = 34) eight weeks after treatment, we fit two logistic regression models, responder versus normal control and non-responder versus normal control. Our global test shows that there is an overall difference in regression coefficients for responders and for non-responders when compared to the normal controls (p-value < 0.001). We next apply our proposed multiple testing procedure to identify the metabolites that have different regression coefficients in these two different logistic regression models. At the FDR < 0.05, our procedure identifies 9 metabolites with different regression coefficients (see Table 5). It is interesting that all these 9 metabolites have the same signs of the refitted coefficients, while the actual magnitudes of the associations between responders and non-responders when compared to the normal controls are different. Besides C24:4.cholesteryl ester, beta-alanine, valine, C18.carnitine and 3-methyladipate/pimelate that we observe in previous analyses, metabolites 5-hydroxytryptopha, nicotinate, and succinate also have differential associations between responders and non-responders when compared to the controls.

Table 5.

Significant metabolites identified via logistic regression of responder vs normal control and non-responder vs normal control for FDR ≤ 5%.

HMDB ID	Synonyms	Refitted Coefficients
		Responder vs.	Non-Responder vs.
		Normal	Normal
06726	C20:4.cholesteryl ester	0.139	1.854
01043	Linoleic.acid	−0.686	−0.388
00472	5-hydroxytryptophan	1.000	1.034
00056	beta-alanine	0.503	2.298
00883	valine	0.628	0.530
00848	C18.carnitine	1.100	0.457
01488	nicotinate	−1.936	−4.312
00254	succinate	0.750	1.508
00555	3-methyladipate/pimelate	−1.989	−4.209

Open in a new tab

7. DISCUSSION

In this paper, for both global and multiple testing, the precision matrix Ω = Σ⁻¹ of the covariates is assumed to be sparse and unknown. Node-wise regression among the covariates is used to learn the covariance structure in constructing the debiased estimator. However, if the prior knowledge of Ω = I is available, the algorithm can be simplified greatly. Specifically, instead of incorporating the Lasso estimators as in (8), we let $v_{j} = {\hat{W}}^{- 1} x_{j}$ and τ_j = ‖v_j‖_n/〈v_j, x_j〉 for each j = 1, …, p. The theoretical properties of the resulting global testing and multiple testing procedures still hold, while the computational efficiency is improved dramatically. However, from our theoretical analysis, even with the knowledge of Ω = I, the theoretical requirement for the model sparsity ( $k = o (\sqrt{n} / {log}^{3} p)$ in the Gaussian case and $k = o (\sqrt{n} / {log}^{5 / 2} p)$ in the bounded case) cannot be relaxed due to the nonlinearity of the problem.

Sample splitting was used in this paper for theoretical purpose. This is different from other works on inference in high-dimensional linear/logistic regression models, including Ingster, Tsybakov, and Verzelen (2010), van de Geer et al. (2014), Mukherjee, Pillai, and Lin (2015) and Javanmard and Javadi (2019), where sample splitting is not needed. However, as we discussed throughout the paper, the assumptions and the alternatives that we considered are different from those previous papers. In the case of high-dimensional logistic regression model, a sample splitting procedure seems unavoidable under the current framework of our technical analysis without making additional strong structural assumptions such as the sparse inverse Hessian matrices used in van de Geer et al. (2014) or the weakly correlated design matrices used in Mukherjee, Pillai, and Lin (2015). Our simulations showed that the sample splitting is actually not needed in order for our proposed methods to perform well. It is of interest to develop technical tools that can eliminate sample splitting in inference for high dimensional logistic regression models.

As mentioned in the introduction, the logistic regression model can be viewed as a special case of the single index model y = f (β^⊤x) + ϵ where f is a known transformation function (Yang et al. 2015). Based on our analysis, it is clear that the theoretical results are not limited to the sigmoid transfer function. In fact, the proposed methods can be applied to a wide range of transformation functions satisfying the following conditions: (C1) f is continuous and for any $u \in ℝ$ , 0 < f (u) < 1; (C2) for any $u_{1}, u_{2} \in ℝ$ , there exists a constant L > 0 such that $| \dot{f} (u_{1}) - \dot{f} (u_{2}) | \leq L | u_{1} - u_{2} |$ ; and (C3) for any constant C > 0, there exists δ > 0 such that for any $| u | \leq C, \dot{f} (u) \geq δ$ . Examples include but are not limited to the following function classes

Cumulative density functions: f (x) = P(X ≤ x) for some continuous random variable X supported on $ℝ$ . In particular, when X ~ N(0, 1), the resulting model becomes the probit regression.
Affine hyperbolic tangent functions: $f (x) = \frac{1}{2} tanh (a x + b) + 1$ for some parameter $a, b \in ℝ$ . In particular, (a, b) = (1, 0) corresponds to f (x) = e^x / (1 + e^x).
Generalized logistic functions: f (x) = (1 + e^−x)^−α for some α > 0.

Besides the problems we considered in this paper, it is also of interest to construct confidence intervals for functionals of the regression coefficients, such as ‖β‖₁,‖β‖₂, or θ^⊤β for some given loading vector θ. In modern statistical machine learning, logistic regression is considered as an efficient classification method (Abramovich and Grinshtein 2018). In practice, a predicted label with an uncertainty assessment is usually preferred. Therefore, another important problem is the construction of predictive intervals of the conditional probability π* associated with a given predictor X*. These problems are related to the current work and are left for future investigations.

8. PROOFS OF THE MAIN THEOREMS

In this section, we prove Theorems 1, Theorem 2 and Theorem 4 in the paper. The proofs of other results, including Theorems 3 and 5, Proposition 1 and the technical lemmas, are given in our Supplementary Materials.

Proof of Theorem 1

Define $F_{i j} = E [η_{i j}^{2} / \dot{f} (u_{i})]$ . Under H₀, $F_{i j} = 4 E [η_{i j}^{2}] = 4 / ω_{j j}$ , and by (A3), c < F_jj < C for j = 1, …, p and some constant C ≥ c > 0. Define statistics

{\tilde{M}}_{j} = \frac{〈 v_{j}, ϵ 〉}{‖ v_{j} ‖_{n}}, and {\overset{ˇ}{M}}_{j} = \frac{\sum_{i = 1}^{n} η_{i j} ϵ_{i} / \dot{f} (u_{i})}{\sqrt{n F_{j j}}}, j = 1, \dots, p .

and ${\tilde{M}}_{n} = {max}_{j} {\tilde{M}}_{j}^{2}$ , ${\overset{ˇ}{M}}_{n} = {max}_{j} {\overset{ˇ}{M}}_{j}^{2}$ . The following lemma shows that ${\tilde{M}}_{n}$ and therefore ${\overset{ˇ}{M}}_{n}$ are good approximations of M_n.

Lemma 1. Under the condition of Theorem 1, the following events

B_{1} = {| {\tilde{M}}_{n} - {\overset{ˇ}{M}}_{n} | = o (1)}, B_{2} = {| {\tilde{M}}_{n} - M_{n} | = o (\frac{1}{log p})},

hold with probability at least 1 − O(p^−c) for some constant c > 0.

It follows that under the event B₁ ∩ B₂, let y_p = 2log p − loglog p + x and ϵ_n = o(1), we have

P_{θ} ({\overset{ˇ}{M}}_{n} \leq y_{p} - ϵ_{n}) \leq P_{θ} (M_{n} \leq y_{p}) \leq P_{θ} ({\overset{ˇ}{M}}_{n} \leq y_{p} + ϵ_{n})

Therefore it suffices to prove that for any $t \in ℝ$ , as (n, p) → ∞,

P_{θ} ({\overset{ˇ}{M}}_{n} \leq y_{p}) \to exp (- \frac{1}{\sqrt{τ}} exp (- x / 2)) .

(21)

Now define ${\hat{M}}_{j} = \frac{\sum_{i = 1}^{n} {\hat{Z}}_{i j}}{\sqrt{n F_{i j}}}, j = 1, \dots, p .$ where ${\hat{Z}}_{i j} = v_{i j}^{0} ϵ_{i} 1 {| v_{i j}^{0} ϵ_{i} | \leq τ_{n}} - E [v_{i j}^{0} ϵ_{i} 1 {| v_{i j}^{0} ϵ_{i} | \leq τ_{n}}]$ for τ_n = log(p + n), $v_{i j}^{0} = η_{i j} / \dot{f} (u_{i})$ and ${\hat{M}}_{n} = {max}_{j} {\hat{M}}_{j}^{2}$ . The following lemma states that ${\hat{M}}_{n}$ is close to ${\overset{ˇ}{M}}_{n}$ .

Lemma 2. Under the condition of Theorem 1, $| {\overset{ˇ}{M}}_{n} - {\hat{M}}_{n} | = o (1)$ with probability at least 1 − O(p^−c) for some constant c > 0.

By Lemma 2, it suffices to prove that for any $t \in ℝ$ , as (n, p) → ∞,

P_{θ} ({\hat{M}}_{n} \leq y_{p}) \to exp (- \frac{1}{\sqrt{π}} exp (- x / 2)) .

(22)

To prove this, we need the classical Bonferroni inequality.

Lemma 3. (Bonferroni inequality) Let $B = \cup_{t = 1}^{p} B_{t}$ . For any integer k < p / 2, we have

\sum_{t = 1}^{2 k} {(- 1)}^{t - 1} A_{t} \leq P (B) \leq \sum_{t = 1}^{2 k - 1} {(- 1)}^{t - 1} A_{t},

(23)

where $A_{t} = \sum_{1 \leq i_{1} < \dots < i_{t} \leq p} P (B_{i_{1}} \cap \dots \cap B_{i_{t}})$ .

By Lemma 3, for any integer 0 < q < p / 2,

\sum_{d = 1}^{2 q} {(- 1)}^{d - 1} \sum_{1 \leq j_{1} \leq \dots \leq j_{d} \leq p} P_{θ} (\cap_{k = 1}^{d} A_{j_{k}}) \leq P_{θ} (max_{1 \leq j \leq p} {\hat{M}}_{j}^{2} \geq y_{p}) \leq \sum_{d = 1}^{2 p - 1} {(- 1)}^{d - 1} \sum_{1 \leq j_{1} < \dots < j_{d} \leq p} P_{θ} (\cap_{k = 1}^{d} A_{j_{k}}),

(24)

where $A_{j_{k}} = {{\hat{M}}_{j_{k}}^{2} \geq y_{p}}$ . Now let $w_{i_{j}} = {\hat{Z}}_{i j} / \sqrt{F_{j j}}$ for j = 1, …, p, and $W_{i} = {(w_{i, j_{1}}, \dots, w_{i, j_{d}})}^{⊤}$ for 1 ≤ i ≤ n. Define ‖a‖_min = min_1≤i≤d | a_i| for any vector $a \in ℝ^{d}$ . Then we have

P_{θ} (\cap_{k = 1}^{d} A_{j_{k}}) = P ((n^{- 1 / 2} \sum_{i = 1}^{n} W_{i} (_{min} \geq y_{p}^{1 / 2}) .

Then it follows from Theorem 1.1 in Zaitsev (1987) that

P_{θ} ((n^{- 1 / 2} \sum_{i = 1}^{n} W_{i} (_{min} \geq y_{p}^{1 / 2}) \leq P_{θ} (‖ N_{d} ‖_{min} \geq y_{p}^{1 / 2} - ϵ_{n} {(log p)}^{- 1 / 2}) + c_{1} d^{5 / 2} exp {- \frac{n^{1 / 2} ϵ_{n}}{c_{2} d^{3} τ_{n} {(log p)}^{1 / 2}}},

(25)

where c₁ > 0 and c₂ > 0 are constants, ϵ_n → 0 which will be specified later, and $N_{d} = (N_{m_{1}}, \dots, N_{m_{d}})$ is a normal random vector with $E (N_{d}) = 0$ and cov(N_d) = cov(W₁). Here d is a fixed integer that does not depend on n, p. Because log p = o(n^1/5), we can let ϵ_n → 0 sufficiently slow, say, $ϵ_{n} = \sqrt{{log}^{5} p / n}$ , so that for any large c > 0,

c_{1} d^{5 / 2} exp {- \frac{n^{1 / 2} ϵ_{n}}{c_{2} d^{3} τ_{n} {(log p)}^{1 / 2}}} = O (p^{- c}) .

(26)

Combining (24), (25) and (26), we have

P_{θ} (max_{1 \leq j \leq p} {\hat{M}}_{j}^{2} \geq y_{p}) \leq \sum_{d = 1}^{2 p - 1} {(- 1)}^{d - 1} \sum_{1 \leq j_{1} < \dots < j_{d} \leq p} P_{θ} (‖ N_{d} ‖_{min} \geq y_{p}^{1 / 2} - ϵ_{n} {(log p)}^{- 1 / 2}) + o (1) .

(27)

Similarly, one can derive

P_{θ} (max_{1 \leq j \leq p} {\hat{M}}_{j}^{2} \geq y_{p}) \geq \sum_{d = 1}^{2 p} {(- 1)}^{d - 1} \sum_{1 \leq j_{1} < \dots < j_{d} \leq p} P_{θ} (‖ N_{d} ‖_{min} \geq y_{p}^{1 / 2} + ϵ_{n} {(log p)}^{- 1 / 2}) + o (1) .

(28)

Now we use the following lemma from Xia, Cai, and Cai (2018).

Lemma 4. For any fixed integer d ≥ 1 and real number $t \in ℝ$ ,

\sum_{1 \leq j_{1} < \dots < j_{d} \leq p} P_{θ} (‖ N_{d} ‖_{min} \geq y_{p}^{1 / 2} \pm ϵ_{n} {(log p)}^{- 1 / 2}) = \frac{1}{d!} {(\frac{1}{\sqrt{π}} exp (- t / 2))}^{d} (1 + o (1)) .

It then follows from the above lemma, (27) and (28) that

\underset{n, p \to \infty}{limsup} P_{θ} (max_{1 \leq j \leq p} {\hat{M}}_{j}^{2} \geq y_{p}) \leq \sum_{d = 1}^{2 p} {(- 1)}^{d - 1} \frac{1}{d!} {(\frac{1}{\sqrt{π}} exp (- t / 2))}^{d},

\underset{n, p \to \infty}{liminf} P_{θ} (max_{1 \leq j \leq p} {\hat{M}}_{j}^{2} \geq y_{p}) \geq \sum_{d = 1}^{2 p - 1} {(- 1)}^{d - 1} \frac{1}{d!} {(\frac{1}{\sqrt{π}} exp (- t / 2))}^{d},

for any positive integer p. By letting p → ∞, we obtain (22) and the proof is complete. □

Proof of Theorem 2.

The proof essentially follows from the general Le Cam’s method described in Section 7.1 of Baraud (2002). The key elements can be summarized as the following lemma that reduces the lower bound problem to calculation of the total variation distance between two posterior distributions.

Lemma 5. Let $H_{1}$ be some subset in an $l_{2}$ bounded Hilbert space and ρ some positive number. Let μ_ρ be some probability measure on $H_{1} = {θ \in Θ, ‖ θ ‖ = ρ}$ . Set $P_{μ_{ρ}} = \int P_{θ} d μ_{ρ} (θ)$ , P₀ as the (posterior) distribution at the null, and denote by Φ_α the level-α tests, we have

inf_{Φ_{α}} sup_{θ \in H_{1}} P_{θ} (Φ_{α} = 0) \geq inf_{Φ_{α}} P_{μ_{ρ}} (Φ_{α} = 0) \geq 1 - α - T V (P_{μ_{ρ}}, P_{0}),

where $T V (P_{μ_{ρ}}, P_{0})$ denotes the total variation distance between $P_{μ_{ρ}}$ and P₀.

Now since by definition ρ*(Φ_α, δ, Θ(k)) ≥ ρ*(Φ_α, δ, Σ) for any Σ ∈Θ₂(k), by Lemma 5, it suffices to construct the corresponding $H_{1}$ for β ∈Θ_β(k) and find a lower bound ρ₁ = ρ(η) such that

\forall ρ \leq ρ_{1} inf_{Φ_{α}} P_{μ_{ρ}} (Φ_{α} = 0) \geq 1 - α - η = δ .

(29)

for fixed covariance Σ = I. In this case, an upper bound for the χ²-divergence between $P_{μ_{ρ}}$ and P₀, defined as $χ^{2} (P_{μ_{ρ}}, P_{0}) = \int \frac{{(d P_{μ_{ρ}})}^{2}}{d P_{0}} - 1$ , can be obtained by carefully constructing the alternative space $H_{1}$ . Since $T V (f, g) \leq \sqrt{χ^{2} (f, g)}$ (see p.90 of Tsybakov (2009)), it follows that $inf_{Φ_{α}} P_{μ_{ρ}} (Φ_{α} = 0) \geq 1 - α - \sqrt{χ^{2} (P_{μ_{ρ}}, P_{0})}$ . By choosing ρ₁ = ρ(η) such that for any ρ ≤ ρ₁, $χ^{2} (P_{μ_{ρ}}, P_{0}) \leq η^{2} = {(1 - α - δ)}^{2}$ , we have (29) holds. In the following, we will construct the alternative space $H_{1}$ and derive an upper bound of $χ^{2} (P_{μ_{ρ}}, P_{0})$ where P₀ corresponds to the null space $H_{0}$ defined at a single point β = 0. We divide the proofs into two parts. Throughout, the design covariance matrix is chosen as Σ = I.

Step 1: Construction of $H_{1}$ .

Firstly, for a set M, we define $l (M, n)$ as the set of all the n-element subsets of M. Let [1: p] ≡ {1, …, p}, so $l ([1 : p], k)$ contains all the k-element subsets of [1: p]. We define the alternative parameter space $H_{1} = {β \in ℝ^{p} : β_{j} = ρ 1 {j \in I} for I \in l ([1 : p], k)}$ . In other words, $H_{1}$ contains all the k-sparse vectors β(I) whose nonzero components ρ are indexed by I. Apparently, for any $β \in H_{1}$ , it follows ‖β‖_∞ = ρ and $H_{1} \subseteq Θ_{1} (k)$ .

Step 2: Control of $χ^{2} (P_{π_{H_{1}}}, P_{0})$ .

Let π denote the uniform prior of the random index set I over $l ([1 : p], k)$ . This prior induces a prior distribution $π_{H_{1}}$ over the parameter space $H_{1}$ . For ${0_{p}} = H_{0}$ , the corresponding joint distribution of the data ${(X_{i}, y_{i})}_{i = 1}^{n}$ is

f = \prod_{i = 1}^{n} p (X_{i}, y_{i}) = \frac{1}{{(2 π)}^{n p / 2}} \prod_{i = 1}^{n} \frac{1}{2} e^{- ‖ X_{i} ‖_{2}^{2} / 2} .

Similarly, the posterior distribution of the samples over the prior $π_{H_{1}}$ is denoted as

g = \prod_{i = 1}^{n} \int_{H_{i}} p (X_{i}, y_{j}; β) π_{H_{i}} = \frac{1}{(\begin{array}{l} p \\ k \end{array})} \sum_{β \in H_{1}} \prod_{i = 1}^{n} p (X_{i}, y_{i}; β) .

As a result, we have the following lemma controlling $χ^{2} (P_{π_{H_{1}}}, P_{0}) = χ^{2} (g, f)$ .

Lemma 6. Let $ρ^{2} = \frac{1}{n} log (1 + \frac{p}{h (η) k^{2}})$ where h(η) = [log(η² + 1)]⁻¹ and η = 1 − α − δ, then we have χ²(g, f) ≤ (1 − α − δ)².

Combining Lemma 5 and Lemma 6, we know that for α, δ > 0 and α + δ < 1, if $ρ = \sqrt{\frac{1}{n} log (1 + \frac{p}{h (η) k^{2}})}$ , then $\forall ρ^{'} \leq ρ, inf_{Φ_{α}} sup_{β \in Θ (k) : ‖ β ‖_{\infty} \geq ρ^{'}} P_{θ} (Φ_{α} = 0) \geq δ$ . Therefore, it follows that

ρ^{*} (α, δ, Θ (k)) \geq ρ^{*} (α, δ, I) ≳ \sqrt{\frac{1}{n} log (1 + \frac{p}{k^{2}})} .

(30)

Lastly, note that for the above chosen ρ, $H_{1} \subset Θ_{1} (k) \cap {β \in ℝ^{p} : ‖ β ‖_{2} ≲ {(n^{1 / 4} log p)}^{- 1}}$ when $k ≲ min {p^{γ}, \sqrt{n} / {log}^{3} p}$ for some 0 < γ < 1 / 2. This completes the proof. □

Proof of Theorem 4.

The proof follows similar arguments of the proof of Theorem 3.1 in Javanmard and Javadi (2019). We first consider the case when $\hat{t}$ , given by (17), does not exist. In this case, $\hat{t} = \sqrt{2 log p}$ and we consider the event $Ω_{0} = {\sum_{j \in H_{0}} I (| M_{j} | \geq \sqrt{2 log p}) \geq 1}$ that there are at least one false positive. In order to show the FDR/FDP can be controlled in this case, we show that

P_{θ} (Ω_{0}) \to 0, as (n, p) \to \infty .

(31)

Note that for $j \in H_{0}$ , we have $M_{j} = \frac{{\overset{ˇ}{β}}_{j}}{τ_{j}} = \frac{〈 v_{j}, ϵ 〉}{‖ v_{j} ‖_{n}} + \frac{〈 v_{j}, R e 〉}{‖ v_{j} ‖_{n}} - \frac{{〈 v_{j}, h_{- j} 〉}_{n}}{‖ v_{j} ‖_{n}}$ Then

P_{θ} (Ω_{0}) \leq P_{θ} (\sum_{j \in H_{0}} I (\frac{〈 v_{j}, ϵ 〉}{‖ v_{j} ‖_{n}} + \frac{〈 v_{j}, R e 〉}{‖ v_{j} ‖_{n}} - \frac{{〈 v_{j}, h_{- j}}}_{n}}{‖ v_{j} ‖_{n}} \geq \sqrt{2 log p}) \geq 1) + P_{θ} (\sum_{j \in H_{0}} I (\frac{〈 v_{j}, ϵ 〉}{‖ v_{j} ‖_{n}} + \frac{〈 v_{j}, R e 〉}{‖ v_{j} ‖_{n}} - \frac{{〈 v_{j}, h_{- j} 〉}_{n}}{‖ v_{j} ‖_{π}} \leq \sqrt{2 log p}) \geq 1) .

(32)

For any ϵ > 0, we can bound the first term by

P_{θ} (\sum_{j \in H_{0}} I (\frac{〈 v_{j}, ϵ 〉}{{‖ v_{j} ‖}_{k}} + \frac{〈 v_{f} R e 〉}{{‖ v_{j} ‖}_{v}} - \frac{{〈 v_{j}, h_{- j} 〉}_{n}}{{‖ v_{j} ‖}_{n}} \geq \sqrt{2 log p}) \geq 1) = P_{θ} (\sum_{j \in H_{0}} I ({\tilde{M}}_{j} \geq \sqrt{2 log p} + \frac{{〈 v_{j}, h_{- j} 〉}_{n}}{{‖ v_{j} ‖}_{n}} - \frac{〈 v_{j}, R e 〉}{{‖ v_{j} ‖}_{n}}) \geq 1) \leq P_{θ} (\sum_{j \in H_{0}} I ({\tilde{M}}_{j} \geq \sqrt{2 log p} - ϵ) \geq 1) + P_{θ} (max_{j \in H_{0}} | \frac{{〈 v_{j}, h_{- j} 〉}_{n}}{{‖ v_{j} ‖}_{n}} - \frac{〈 v_{j}, R e 〉}{{‖ v_{j} ‖}_{n}} | \geq ϵ) \leq p max_{j \in H_{0}} P_{θ} ({\tilde{M}}_{j} \geq \sqrt{2 log p} - ϵ) + P_{θ} (max_{j \in H_{0}} | \frac{{〈 v_{j}, h_{- j} 〉}_{n}}{{‖ v_{j} ‖}_{n}} - \frac{〈 v_{j}, R e 〉}{{‖ v_{j} ‖}_{n}} | \geq ϵ)

By the proof of Lemma 1, we know that $P_{θ} ({max}_{j \in H_{0}} | \frac{{〈 v_{j}, h_{- j} 〉}_{n}}{‖ v_{j} ‖_{n}} - \frac{〈 v_{j}, R e 〉}{‖ v_{j} ‖_{n}} | \geq ϵ) \to 0$ . In addition, for $j \in H_{0}, P_{θ} ({\tilde{M}}_{j} \geq \sqrt{2 log p} - ϵ) \leq P_{θ} ({\overset{ˇ}{M}}_{j} \geq \sqrt{2 log p} - 2 ϵ) + P_{θ} (| {\tilde{M}}_{j} - {\overset{ˇ}{M}}_{j} | \geq ϵ)$ , where ${max}_{j \in H_{0}} P_{θ} (| {\tilde{M}}_{j} - {\overset{ˇ}{M}}_{j} | \geq ϵ) = O (p^{- c})$ for some sufficiently large c > 0. Now since ${\overset{ˇ}{M}}_{j} = \frac{\sum_{i = 1}^{n} η_{i j} ϵ_{i} / \dot{f} (u_{i})}{\sqrt{n F_{j j}}}$ where $E \frac{η_{i j} ϵ_{i} / \dot{f} (u_{i})}{\sqrt{F_{j j}}} = 0$ and $Var (\frac{η_{i j} ϵ_{i} / \dot{f} (u_{i})}{\sqrt{F_{j j}}}) = 1$ , by Lemma 6.1 of Liu (2013), we have $sup_{0 \leq t \leq 4 \sqrt{log p}} | \frac{P_{θ} (| {\overset{ˇ}{M}}_{j} | \geq t)}{G (t)} - 1 | \leq C {(log p)}^{- 1}$ . Now let $t = \sqrt{2 log p} - 2 ϵ$ , we have $P_{θ} ({\overset{ˇ}{M}}_{j} \geq \sqrt{2 log p} - 2 ϵ) \leq G (\sqrt{2 log p} - 2 ϵ) + C \frac{G (\sqrt{2 log p} - 2 ϵ)}{log p}$ .

Hence $p {max}_{j \in H_{0}} P_{θ} ({\tilde{M}}_{j} \geq \sqrt{2 log p} - ϵ) \leq C p G (\sqrt{2 log p} - 2 ϵ) + O (p^{- c})$ , which goes to zero as (n, p) → ∞. By symmetry, we know that the second term in (32) also goes to 0. Therefore we have proved (31).

Now consider the case when $0 \leq \hat{t} \leq b_{p}$ holds. We have

{FDP}_{θ} (\hat{t}) = \frac{\sum_{j \in H_{0}} I {| M_{j} | \geq \hat{t}}}{max {\sum_{j = 1}^{p} I {| M_{j} | \geq \hat{t}}, 1}} \leq \frac{p_{0} G (\hat{t})}{max {\sum_{j = 1}^{p} I | M_{j} | \geq \hat{t}}, 1}} (1 + A_{p}),

where $A_{p} = sup_{0 \leq t \leq b_{p}} | \frac{\sum_{j \in H_{0}} I {| M_{j} | \geq t}}{p_{0} G (t)} - 1 |$ Note that by definition $\frac{p_{0} G (\hat{t})}{max {\sum_{j = 1}^{p} I {| M_{j} | \geq \hat{t}}, 1}} \leq \frac{p_{0} α}{p}$ . The proof is complete if A_p → 0 in probability, which has been shown by Proposition 1. □

Supplementary Material

Supp 1

NIHMS1550238-supplement-Supp_1.zip^{(1.1MB, zip)}

Footnotes

SUPPLEMENTARY MATERIALS

In the online Supplemental Materials, we prove Theorem 3, 5, Proposition 1, and the technical lemmas. The technical results and simulations concerning the two-sample tests discussed in Section 4 are also included.

References

Abramovich F, and Grinshtein V (2018), “High-dimensional classification by sparse logistic regression,” IEEE Transactions on Information Theory, 65, 3068–3079. [Google Scholar]
Bach F (2010), “Self-concordant analysis for logistic regression,” Electronic Journal of Statistics, 4, 384–414. [Google Scholar]
Baraud Y (2002), “Non-asymptotic minimax rates of testing in signal detection,” Bernoulli, 8, 577–606. [Google Scholar]
Barrett JC, Hansoul S, Nicolae DL, Cho JH, Duerr RH, Rioux JD, Brant SR, Silverberg MS, Taylor KD, Barmada MM, et al. (2008), “Genome-wide association defines more than 30 distinct susceptibility loci for Crohn’s disease,” Nature Genetics, 40, 955. [DOI] [PMC free article] [PubMed] [Google Scholar]
Belloni A, Chernozhukov V, and Wei Y (2016), “Post-selection inference for generalized linear models with many controls,” Journal of Business & Economic Statistics, 34, 606–619. [Google Scholar]
Benjamini Y, and Hochberg Y (1995), “Controlling the false discovery rate: a practical and powerful approach to multiple testing,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 289–300. [Google Scholar]
Benjamini Y, and Yekutieli D (2001), “The control of the false discovery rate in multiple testing under dependency,” The Annals of Statistics, 29, 1165–1188. [Google Scholar]
Cai TT, and Guo Z (2017), “Confidence intervals for high-dimensional linear regression: Minimax rates and adaptivity,” The Annals of Statistics, 45, 615–646. [Google Scholar]
Cai TT, Li H, Ma J, and Xia Y (2017), “Differential Markov random field analysis with applications to detecting differential microbial community structures,” Unpublished Manuscript. [DOI] [PMC free article] [PubMed] [Google Scholar]
Cai TT, Liu W, and Xia Y (2013), “Two-sample covariance matrix testing and support recovery in high-dimensional and sparse settings,” Journal of the American Statistical Association, 108, 265–277. [Google Scholar]
——— (2014), “Two-sample test of high dimensional means under dependence,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 76, 349–372. [Google Scholar]
Candès E, Fan Y, Janson L, and Lv J (2018), “Panning for gold: model-X knockoffs for high dimensional controlled variable selection,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 80, 551–577. [Google Scholar]
Candès EJ, and Sur P (2018), “The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression,” arXiv preprint arXiv:1804.09753. [DOI] [PMC free article] [PubMed]
Fortin G (2011), “L-Carnitine and intestinal inflammation,” in Vitamins & Hormones, Elsevier, vol. 86, pp. 353–366. [DOI] [PubMed] [Google Scholar]
Ingster YI (1993), “Asymptotically minimax hypothesis testing for nonparametric alternatives. I, II, III,” Mathematical Methods of Statiststics, 2, 85–114. [Google Scholar]
Ingster YI, Tsybakov AB, and Verzelen N (2010), “Detection boundary in sparse regression,” Electronic Journal of Statistics, 4, 1476–1526. [Google Scholar]
Javanmard A, and Javadi H (2019), “False discovery rate control via debiased lasso,” Electronic Journal of Statistics, 13, 1212–1253. [Google Scholar]
Javanmard A, and Montanari A (2014a), “Confidence intervals and hypothesis testing for high-dimensional regression.” Journal of Machine Learning Research, 15, 2869–2909. [Google Scholar]
——— (2014b), “Hypothesis testing in high-dimensional regression under the gaussian random design model: Asymptotic theory,” IEEE Transactions on Information Theory, 60, 6522–6554. [Google Scholar]
Lee D, Baldassano RN, Otley AR, Albenberg L, Griffiths AM, Compher C, Chen EZ, Li H, Gilroy E, Nessel L, et al. (2015), “Comparative effectiveness of nutritional and biological therapy in North American children with active Crohn’s disease,” Inflammatory Bowel Diseases, 21, 1786–1793. [DOI] [PubMed] [Google Scholar]
Lewis JD, Chen EZ, Baldassano RN, Otley AR, Griffiths AM, Lee D, Bittinger K, Bailey A, Friedman ES, Hoffmann C, et al. (2015), “Inflammation, antibiotics, and diet as environmental stressors of the gut microbiome in pediatric Crohn’s disease,” Cell Host & Microbe, 18, 489–500. [DOI] [PMC free article] [PubMed] [Google Scholar]
Liu W (2013), “Gaussian graphical model estimation with false discovery rate control,” The Annals of Statistics, 41, 2948–2978. [Google Scholar]
Liu W, and Luo S (2014), “Hypothesis testing for high-dimensional regression models,” Technical report. [Google Scholar]
Maceyka M, and Spiegel S (2014), “Sphingolipid metabolites in inflammatory disease,” Nature, 510, 58. [DOI] [PMC free article] [PubMed] [Google Scholar]
Meier L, van de Geer S, and Bühlmann P (2008), “The group lasso for logistic regression,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 70, 53–71. [Google Scholar]
Mukherjee R, Pillai NS, and Lin X (2015), “Hypothesis testing for high-dimensional sparse binary regression,” Annals of Statistics, 43, 352–381. [DOI] [PMC free article] [PubMed] [Google Scholar]
Negahban S, Ravikumar P, Wainwright MJ, and Yu B (2010), “A unified framework for high-dimensional analysis of M-estimators with decomposable regularizers,” Technical Report Number 979. [Google Scholar]
Ni J, Shen T-CD, Chen EZ, Bittinger K, Bailey A, Roggiani M, Sirota-Madi A, Friedman ES, Chau L, Lin A, et al. (2017), “A role for bacterial urease in gut dysbiosis and Crohn’s disease,” Science Translational Medicine, 9, eaah6888. [DOI] [PMC free article] [PubMed] [Google Scholar]
Peltekova VD, Wintle RF, Rubin LA, Amos CI, Huang Q, Gu X, Newman B, Van Oene M, Cescon D, Greenberg G, et al. (2004), “Functional variants of OCTN cation transporter genes are associated with Crohn’s disease,” Nature Genetics, 36, 471. [DOI] [PubMed] [Google Scholar]
Plan Y, and Vershynin R (2013), “Robust 1-bit compressed sensing and sparse logistic regression: A convex programming approach,” IEEE Transactions on Information Theory, 59, 482–494. [Google Scholar]
Ren Z, Zhang C-H, and Zhou HH (2016), “Asymptotic normality in estimation of large Ising graphical model,” Unpublished Manuscript. [Google Scholar]
Rozenova KA, Deevska GM, Karakashian AA, and Nikolova-Karakashian MN (2010), “Studies on the role of acid sphingomyelinase and ceramide in the regulation of tumor necrosis factor α (TNFα)-converting enzyme activity and TNFα secretion in macrophages,” Journal of Biological Chemistry, 285, 21103–21113. [DOI] [PMC free article] [PubMed] [Google Scholar]
Sewell GW, Hannun YA, Han X, Koster G, Bielawski J, Goss V, Smith PJ, Rahman FZ, Vega R, Bloom SL, et al. (2012), “Lipidomic profiling in Crohn’s disease: abnormalities in phosphatidylinositols, with preservation of ceramide, phosphatidylcholine and phosphatidylserine composition,” The International Journal of Biochemistry & Cell Biology, 44, 1839–1846. [DOI] [PMC free article] [PubMed] [Google Scholar]
Shekhawat PS, Sonne S, Carter AL, Matern D, and Ganapathy V (2013), “Enzymes involved in L-carnitine biosynthesis are expressed by small intestinal enterocytes in mice: Implications for gut health,” Journal of Crohn’s & Colitis, 7, e197–e205. [DOI] [PMC free article] [PubMed] [Google Scholar]
Sur P, and Candès EJ (2019), “A modern maximum-likelihood theory for high-dimensional logistic regression,” Proceedings of the National Academy of Sciences, 116, 14516–14525. [DOI] [PMC free article] [PubMed] [Google Scholar]
Sur P, Chen Y, and Candès EJ (2017), “The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled chi-square,” Probability Theory and Related Fields, 1–72. [Google Scholar]
Tsybakov AB (2009), Introduction to Nonparametric Estimation, Springer Series in Statistics. Springer, New York. [Google Scholar]
van de Geer S (2008), “High-dimensional generalized linear models and the lasso,” The Annals of Statistics, 36, 614–645. [Google Scholar]
van de Geer S, Bühlmann P, Ritov Y, and Dezeure R (2014), “On asymptotically optimal confidence regions and tests for high-dimensional models,” The Annals of Statistics, 42, 1166–1202. [Google Scholar]
Verzelen N (2012), “Minimax risks for sparse regressions: Ultra-high dimensional phenomenons,” Electronic Journal of Statistics, 6, 38–90. [Google Scholar]
Xia Y, Cai T, and Cai TT (2015), “Testing differential networks with applications to the detection of gene-gene interactions,” Biometrika, 102, 247–266. [DOI] [PMC free article] [PubMed] [Google Scholar]
——— (2018), “Two-sample tests for high-dimensional linear regression with an application to detecting interactions,” Statistica Sinica, 28, 63–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
Xia Y, Cai TT, and Li H (2018), “Joint testing and false discovery rate control in high-dimensional multivariate regression,” Biometrika, 105, 249–269. [DOI] [PMC free article] [PubMed] [Google Scholar]
Yang Z, Wang Z, Liu H, Eldar YC, and Zhang T (2015), “Sparse nonlinear regression: Parameter estimation and asymptotic inference,” arXiv preprint arXiv:1511.04514.
Zaitsev AY (1987), “On the Gaussian approximation of convolutions under multidimensional analogues of S.N. Bernstein’s inequality conditions,” Probability Theory and Related Fields, 74, 535–566. [Google Scholar]
Zhang C-H, and Zhang SS (2014), “Confidence intervals for low dimensional parameters in high dimensional linear models,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 76, 217–242. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supp 1

NIHMS1550238-supplement-Supp_1.zip^{(1.1MB, zip)}

[R1] Abramovich F, and Grinshtein V (2018), “High-dimensional classification by sparse logistic regression,” IEEE Transactions on Information Theory, 65, 3068–3079. [Google Scholar]

[R2] Bach F (2010), “Self-concordant analysis for logistic regression,” Electronic Journal of Statistics, 4, 384–414. [Google Scholar]

[R3] Baraud Y (2002), “Non-asymptotic minimax rates of testing in signal detection,” Bernoulli, 8, 577–606. [Google Scholar]

[R4] Barrett JC, Hansoul S, Nicolae DL, Cho JH, Duerr RH, Rioux JD, Brant SR, Silverberg MS, Taylor KD, Barmada MM, et al. (2008), “Genome-wide association defines more than 30 distinct susceptibility loci for Crohn’s disease,” Nature Genetics, 40, 955. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R5] Belloni A, Chernozhukov V, and Wei Y (2016), “Post-selection inference for generalized linear models with many controls,” Journal of Business & Economic Statistics, 34, 606–619. [Google Scholar]

[R6] Benjamini Y, and Hochberg Y (1995), “Controlling the false discovery rate: a practical and powerful approach to multiple testing,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 289–300. [Google Scholar]

[R7] Benjamini Y, and Yekutieli D (2001), “The control of the false discovery rate in multiple testing under dependency,” The Annals of Statistics, 29, 1165–1188. [Google Scholar]

[R8] Cai TT, and Guo Z (2017), “Confidence intervals for high-dimensional linear regression: Minimax rates and adaptivity,” The Annals of Statistics, 45, 615–646. [Google Scholar]

[R9] Cai TT, Li H, Ma J, and Xia Y (2017), “Differential Markov random field analysis with applications to detecting differential microbial community structures,” Unpublished Manuscript. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R10] Cai TT, Liu W, and Xia Y (2013), “Two-sample covariance matrix testing and support recovery in high-dimensional and sparse settings,” Journal of the American Statistical Association, 108, 265–277. [Google Scholar]

[R11] ——— (2014), “Two-sample test of high dimensional means under dependence,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 76, 349–372. [Google Scholar]

[R12] Candès E, Fan Y, Janson L, and Lv J (2018), “Panning for gold: model-X knockoffs for high dimensional controlled variable selection,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 80, 551–577. [Google Scholar]

[R13] Candès EJ, and Sur P (2018), “The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression,” arXiv preprint arXiv:1804.09753. [DOI] [PMC free article] [PubMed]

[R14] Fortin G (2011), “L-Carnitine and intestinal inflammation,” in Vitamins & Hormones, Elsevier, vol. 86, pp. 353–366. [DOI] [PubMed] [Google Scholar]

[R15] Ingster YI (1993), “Asymptotically minimax hypothesis testing for nonparametric alternatives. I, II, III,” Mathematical Methods of Statiststics, 2, 85–114. [Google Scholar]

[R16] Ingster YI, Tsybakov AB, and Verzelen N (2010), “Detection boundary in sparse regression,” Electronic Journal of Statistics, 4, 1476–1526. [Google Scholar]

[R17] Javanmard A, and Javadi H (2019), “False discovery rate control via debiased lasso,” Electronic Journal of Statistics, 13, 1212–1253. [Google Scholar]

[R18] Javanmard A, and Montanari A (2014a), “Confidence intervals and hypothesis testing for high-dimensional regression.” Journal of Machine Learning Research, 15, 2869–2909. [Google Scholar]

[R19] ——— (2014b), “Hypothesis testing in high-dimensional regression under the gaussian random design model: Asymptotic theory,” IEEE Transactions on Information Theory, 60, 6522–6554. [Google Scholar]

[R20] Lee D, Baldassano RN, Otley AR, Albenberg L, Griffiths AM, Compher C, Chen EZ, Li H, Gilroy E, Nessel L, et al. (2015), “Comparative effectiveness of nutritional and biological therapy in North American children with active Crohn’s disease,” Inflammatory Bowel Diseases, 21, 1786–1793. [DOI] [PubMed] [Google Scholar]

[R21] Lewis JD, Chen EZ, Baldassano RN, Otley AR, Griffiths AM, Lee D, Bittinger K, Bailey A, Friedman ES, Hoffmann C, et al. (2015), “Inflammation, antibiotics, and diet as environmental stressors of the gut microbiome in pediatric Crohn’s disease,” Cell Host & Microbe, 18, 489–500. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R22] Liu W (2013), “Gaussian graphical model estimation with false discovery rate control,” The Annals of Statistics, 41, 2948–2978. [Google Scholar]

[R23] Liu W, and Luo S (2014), “Hypothesis testing for high-dimensional regression models,” Technical report. [Google Scholar]

[R24] Maceyka M, and Spiegel S (2014), “Sphingolipid metabolites in inflammatory disease,” Nature, 510, 58. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R25] Meier L, van de Geer S, and Bühlmann P (2008), “The group lasso for logistic regression,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 70, 53–71. [Google Scholar]

[R26] Mukherjee R, Pillai NS, and Lin X (2015), “Hypothesis testing for high-dimensional sparse binary regression,” Annals of Statistics, 43, 352–381. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R27] Negahban S, Ravikumar P, Wainwright MJ, and Yu B (2010), “A unified framework for high-dimensional analysis of M-estimators with decomposable regularizers,” Technical Report Number 979. [Google Scholar]

[R28] Ni J, Shen T-CD, Chen EZ, Bittinger K, Bailey A, Roggiani M, Sirota-Madi A, Friedman ES, Chau L, Lin A, et al. (2017), “A role for bacterial urease in gut dysbiosis and Crohn’s disease,” Science Translational Medicine, 9, eaah6888. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R29] Peltekova VD, Wintle RF, Rubin LA, Amos CI, Huang Q, Gu X, Newman B, Van Oene M, Cescon D, Greenberg G, et al. (2004), “Functional variants of OCTN cation transporter genes are associated with Crohn’s disease,” Nature Genetics, 36, 471. [DOI] [PubMed] [Google Scholar]

[R30] Plan Y, and Vershynin R (2013), “Robust 1-bit compressed sensing and sparse logistic regression: A convex programming approach,” IEEE Transactions on Information Theory, 59, 482–494. [Google Scholar]

[R31] Ren Z, Zhang C-H, and Zhou HH (2016), “Asymptotic normality in estimation of large Ising graphical model,” Unpublished Manuscript. [Google Scholar]

[R32] Rozenova KA, Deevska GM, Karakashian AA, and Nikolova-Karakashian MN (2010), “Studies on the role of acid sphingomyelinase and ceramide in the regulation of tumor necrosis factor α (TNFα)-converting enzyme activity and TNFα secretion in macrophages,” Journal of Biological Chemistry, 285, 21103–21113. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R33] Sewell GW, Hannun YA, Han X, Koster G, Bielawski J, Goss V, Smith PJ, Rahman FZ, Vega R, Bloom SL, et al. (2012), “Lipidomic profiling in Crohn’s disease: abnormalities in phosphatidylinositols, with preservation of ceramide, phosphatidylcholine and phosphatidylserine composition,” The International Journal of Biochemistry & Cell Biology, 44, 1839–1846. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R34] Shekhawat PS, Sonne S, Carter AL, Matern D, and Ganapathy V (2013), “Enzymes involved in L-carnitine biosynthesis are expressed by small intestinal enterocytes in mice: Implications for gut health,” Journal of Crohn’s & Colitis, 7, e197–e205. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R35] Sur P, and Candès EJ (2019), “A modern maximum-likelihood theory for high-dimensional logistic regression,” Proceedings of the National Academy of Sciences, 116, 14516–14525. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R36] Sur P, Chen Y, and Candès EJ (2017), “The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled chi-square,” Probability Theory and Related Fields, 1–72. [Google Scholar]

[R37] Tsybakov AB (2009), Introduction to Nonparametric Estimation, Springer Series in Statistics. Springer, New York. [Google Scholar]

[R38] van de Geer S (2008), “High-dimensional generalized linear models and the lasso,” The Annals of Statistics, 36, 614–645. [Google Scholar]

[R39] van de Geer S, Bühlmann P, Ritov Y, and Dezeure R (2014), “On asymptotically optimal confidence regions and tests for high-dimensional models,” The Annals of Statistics, 42, 1166–1202. [Google Scholar]

[R40] Verzelen N (2012), “Minimax risks for sparse regressions: Ultra-high dimensional phenomenons,” Electronic Journal of Statistics, 6, 38–90. [Google Scholar]

[R41] Xia Y, Cai T, and Cai TT (2015), “Testing differential networks with applications to the detection of gene-gene interactions,” Biometrika, 102, 247–266. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R42] ——— (2018), “Two-sample tests for high-dimensional linear regression with an application to detecting interactions,” Statistica Sinica, 28, 63–92. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R43] Xia Y, Cai TT, and Li H (2018), “Joint testing and false discovery rate control in high-dimensional multivariate regression,” Biometrika, 105, 249–269. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R44] Yang Z, Wang Z, Liu H, Eldar YC, and Zhang T (2015), “Sparse nonlinear regression: Parameter estimation and asymptotic inference,” arXiv preprint arXiv:1511.04514.

[R45] Zaitsev AY (1987), “On the Gaussian approximation of convolutions under multidimensional analogues of S.N. Bernstein’s inequality conditions,” Probability Theory and Related Fields, 74, 535–566. [Google Scholar]

[R46] Zhang C-H, and Zhang SS (2014), “Confidence intervals for low dimensional parameters in high dimensional linear models,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), 76, 217–242. [Google Scholar]

PERMALINK

Global and Simultaneous Hypothesis Testing for High-Dimensional Logistic Regression Models

Rong Ma

T Tony Cai

Hongzhe Li

Abstract

1. INTRODUCTION

1.1. Global and Simultaneous Hypothesis Testing

1.2. Other Related Work

1.3. Organization of the Paper and Notations

2. GLOBAL HYPOTHESIS TESTING

2.1. Construction of the Test Statistic via Generalized Low-Dimensional Projection

Table 1.

2.2. Asymptotic Null Distribution

2.3. Minimax Separation Distance and Optimality

2.4. Comparison with Existing Works

3. LARGE-SCALE MULTIPLE TESTING

3.1. Construction of Multiple Testing Procedures

Procedure Controlling FDR/FDP.

Procedure Controlling FDV.

3.2. Theoretical Properties for Multiple Testing Procedures

4. TESTING FOR TWO LOGISTIC REGRESSION MODELS

5. SIMULATION STUDIES

5.1. Global Hypothesis Testing

Table 2.

Fig. 1.

5.2. Multiple Hypotheses Testing

FDR Control.

Fig. 2.

Fig. 3.

FDV Control.

Table 3.

6. REAL DATA ANALYSIS

6.1. Association Between Metabolites and Crohn’s Disease Before and After Treatment

Table 4.

6.2. Comparison of Metabolite Associations Between Responders and Non-Responders

Table 5.

7. DISCUSSION

8. PROOFS OF THE MAIN THEOREMS

Proof of Theorem 1

Proof of Theorem 2.

Step 1: Construction of H1.

Step 2: Control of χ2(PπH1,P0).

Proof of Theorem 4.

Supplementary Material

Footnotes

References

Associated Data

Supplementary Materials

ACTIONS

PERMALINK

RESOURCES

Similar articles

Cited by other articles

Links to NCBI Databases

Step 1: Construction of $H_{1}$ .

Step 2: Control of $χ^{2} (P_{π_{H_{1}}}, P_{0})$ .