Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2014 Feb 3.
Published in final edited form as: Stat Sin. 2012 Oct;22(4):1379–1401. doi: 10.5705/ss.2010.199

SEMIPARAMETRIC QUANTILE REGRESSION WITH HIGH-DIMENSIONAL COVARIATES

Liping Zhu 1, Mian Huang 2, Runze Li 3
PMCID: PMC3910001  NIHMSID: NIHMS383105  PMID: 24501536

Abstract

This paper is concerned with quantile regression for a semiparametric regression model, in which both the conditional mean and conditional variance function of the response given the covariates admit a single-index structure. This semiparametric regression model enables us to reduce the dimension of the covariates and simultaneously retains the flexibility of nonparametric regression. Under mild conditions, we show that the simple linear quantile regression offers a consistent estimate of the index parameter vector. This is a surprising and interesting result because the single-index model is possibly misspecified under the linear quantile regression. With a root-n consistent estimate of the index vector, one may employ a local polynomial regression technique to estimate the conditional quantile function. This procedure is computationally efficient, which is very appealing in high-dimensional data analysis. We show that the resulting estimator of the quantile function performs asymptotically as efficiently as if the true value of the index vector were known. The methodologies are demonstrated through comprehensive simulation studies and an application to a real dataset.

Key words and phrases: Dimension reduction, heteroscedasticity, linearity condition, local polynomial regression, quantile regression, single-index model

1. Introduction

Since the seminal work of Koenker and Bassett (1978), quantile regression has become an important statistical analytic tool. Koenker (2005) provided a comprehensive review of the topic of quantile regression. Nonparametric quantile regression has been extensively studied in situations where the covariate vector is univariate. For instance, Bhattacharya and Gangopadhyay (1990) studied kernel estimation and the nearest neighborhood estimation, and Fan, Hu, and Truong (1994) and Yu and Jones (1998) proposed local linear polynomial quantile regression. Koenker, Ng, and Portnoys (1994) proposed regression spline approaches for estimating the conditional quantile regression. When the covariate vector x is multivariate, Stone (1977) and Chaudhuri (1991) proposed fully nonparametric quantile regression. Due to the “curse of dimensionality”, however, the fully nonparametric quantile regression may not be very useful in practice. Thus, researchers considered alternatives by assuming structures on the regression function. De Gooijer and Zerom (2003), Yu and Lu (2004) and Horowitz and Lee (2005) proposed quantile regression for additive models. Wang, Zhu and Zhou (2009) considered quantile regression for partially linear varying-coefficient models. Wu, Yu and Yu (2010) suggested a back-fitting algorithm for single-index models. Based on their estimation equation, Kong and Xia (2010) investigated the Bahadur representation of single-index parameter estimators. Gannoun et al. (2004) considered quantile regression for multi-index models.

Let Y be a response variable, and x = (X1, …, Xp)T be a covariate vector. In this paper, we consider the following heteroscedastic single-index model:

Y=G(xTβ0)+σ(xTβ0)ε, (1.1)

where the error term ε is assumed to be independent of x, E(ε) = 0 and var(ε) = 1. The functions G(·) and σ(·) are unspecified, nonparametric smoothing functions. The parameter of interest is the direction of β0. To ensure the identifiability, it is assumed throughout this paper that ||β0|| = 1 with first nonzero element being positive. It is easy to see that model (1.1) includes linear quantile regression (Koenker and Bassett, 1978) as a special case thereof. Model (1.1) becomes a nonparametric regression model when p = 1, and the traditional single-index model (Ichimura, 1993) when σ(·) is a positive constant. Here we allow for the heteroscedasticity. The single-index structure in (1.1) retains the flexibility of nonparametric regression and allows the presence of high-dimensional covariates.

Chaudhuri, Doksum and Samarov (1997) proposed an estimation procedure for β0 using the average derivative approach, which takes partial derivatives of the conditional quantile with respect to the covariates x. The average derivative approach requires multivariate kernel regression, which is hence less useful in the presence of high-dimensional covariates. Wu, Yu and Yu (2010) proposed a back-fitting algorithm, which is computationally expensive when p is large. In this paper, we will propose a new estimation procedure for model (1.1). The newly proposed estimation procedure consists of two steps. We first estimate β0 using the linear quantile regression. We show that the resulting estimate is consistent. We then employ the local linear regression (Fan and Gijbels, 1996) to estimate the conditional quantile function of Y given xTβ0. This paper makes the following major contributions to the literature:

  1. Under mild conditions, we show that for any link function G(·) and τ ∈ (0, 1), the τ-th linear quantile regression coefficient for Y|x is proportional to β0 in the single-index model (1.1).

  2. We prove that the local linear regression estimate for the conditional quantile function based upon the root-n consistent estimate of β0 has the same asymptotic bias and variance as the local linear estimate with the true value of β0.

The theoretical result in (a) implies that the ordinary linear quantile regression actually results in root-n consistent estimates for β0. Thus, the quantile regression provides an effective tool to reduce the dimension of the covariate in the presence of high-dimensional covariates. This property enables us to reduce substantially the computational cost of the back-fitting algorithm (Wu, Yu and Yu, 2010). Thus our proposed procedure is appealing in high-dimensional data analysis. We demonstrate this issue in detail in our numerical analysis. In addition, variable selection procedures for quantile regression (Zou and Yuan, 2008; Wu and Liu, 2009; Belloni and Chernozhukov, 2011) may be applicable to model (1.1) with high-dimensional covariates when many covariates are not significant. In general, the regularization parameter in the penalized quantile regression should depend on the quantile index τ (Belloni and Chernozhukov, 2011). The result in (b) implies that the dimensionality of covariate does not affect the performance of the local linear estimate asymptotically. Thus, it may not be necessary to update the estimate of β0 iteratively once a root-n consistent estimator of β0 is available. It also ensures our proposed procedure to work properly in high-dimensional settings.

The rest of this article is organized as follows. In Section 2 we propose a two-step procedure to estimate the conditional quantile function. Asymptotic properties of the resulting estimates are also rigorously established in this section. We demonstrate the methodologies through comprehensive simulation studies and an application to a real data set in Section 3. We conclude this paper with a brief discussion in Section 4. All proofs are given in the Appendix.

2. A New Estimation Procedure

In this section we study the estimation of quantile regression for model (1.1). We prove that the linear quantile regression yields consistent estimate for β0. We further demonstrate that the quantile regression for model (1.1) can be reduced to the univariate quantile regression through replacing β0 with β̂0, the resulting estimate of the linear quantile regression.

2.1. Estimation of β0

Let ρτ(r) = τrr1(r < 0) be the check loss function, and Inline graphic(u, β) = E{ρτ(Y − u − βTx)}. Define

(uτ,βτ)=defargminu,β{Lτ(u,β)}. (2.1)

Theorem 1 below implies that the linear quantile regression offers a consistent estimate for β0 in (1.1) under mild conditions.

Theorem 1

If the covariate vector x in model (1.1) satisfies the following linearity condition

E(xβ0Tx)=var(x)β0{β0Tvar(x)β0}-1β0Tx, (2.2)

then βτ = κβ0 for some constant κ, where βτ is defined in (2.1).

The proof of Theorem 1 will be given in the Appendix. The linearity condition (2.2) plays an important role here. Theorem 1 implies that under the linearity condition, we are able to obtain a root n consistent estimator of the direction of β by the linear quantile regression. This will be further confirmed in Theorem 2. The linearity condition (2.2) is widely assumed in the context of sufficient dimension reduction. Li (1991) pointed out that (2.2) is satisfied when x follows an elliptically contour distribution. Hall and Li (1993) proved that this linearity condition always holds to a good approximation in single-index models of the form (1.1) when the dimension p of the covariates diverges. Thus, the linearity condition is typically regarded as mild, particularly when p is fairly large.

It is worth pointing out that similar results can be established for general convex loss functions under the linearity condition. When the quadratic loss function is used, however, the resulting estimate cannot recover any information beyond the conditional mean function. Similarly, the absolute loss function is designed to identify information in the conditional median function. More generally, the check loss function at any given τ is specifically designed to find the information relevant exclusively to the τ-th quantile function. Theorem 1 justifies partially the validity of the procedure suggested by Gannoun et al. (2004). However, their method can be excessive because the dimension-reduction step captures all information relevant to the whole conditional distribution function.

Theorem 1 also implies that the existing statistical procedures for the linear quantile regression may be directly applied to model (1.1). In particular, the existing variable selection procedure for the linear quantile regression with high-dimensional covariates (Wu and Liu, 2009; Zou and Yuan, 2008; Belloni and Chernozhukov, 2011) may be used to select significant variables for model (1.1). It is important to correctly set the range of the tuning parameter in variable selection procedure based on a penalized linear quantile regression. The range of the tuning parameter may explicitly or implicitly depend on the link function G, the quantile index τ and the density of error distribution. Data-driven method for tuning parameter selection is recommended.

Let (xi, Yi), i = 1, …, n, be a random sample from (x, Y), and

Lτn(u,β)=n-1i=1n{ρτ(Yi-u-xiTβ)}.

Let

(u^τ,β^τ)=defargminu,β{Lτn(u,β)}. (2.3)

The asymptotic normality of β̂τ is established in the following theorem, whose proof will be given in the Appendix.

Theorem 2

Assume that Λ = E{f*(ζτ|x)xxT} is invertible, where f*(z|x) denotes the conditional density of (YxTβτ) given x, and ζτ denotes the τ-th quantile of Y|x. Then n1/2(β̂τβτ) is asymptotically normal with mean zero and covariance matrix τ(1 − τ) Λ−1 E(xxT) Λ−1. If the “partial residual” term YxTβτ is independent of x, the asymptotic variance matrix of β̂τ reduces to τ(1 − τ) E(xxT)−1/{f*(ζτ)}2.

2.2. Estimation of conditional quantile function

By Theorem 1, βτ is proportional to β0, and hence the conditional distribution of Y|(xTβ0) is exactly the same as that of Y|(xTβτ). Consequently, Theorem 1 allows to estimate the quantile regression of Y|(xTβ0) through the conditional distribution of Y|(xTβτ). We denote the τ-quantile function of Y|(xTβτ) by Gτ;(·).

By using data points {( xiTβ^τ, Yi), i = 1, …, n}, we estimate Gτ(·) by using the local linear regression. For ease of presentation, let z=x0Tβτ,z^=x0Tβ^τ,Zi=xiTβτ,Z^i=xiTβ^τ. Local linear regression is used to approximate Gτ(Z) by

Gτ(Z)Gτ(z)+Gτ(z)(Z-z)

for Z in the neighborhood of z. Let

(a^,b^)=defargmina,b1ni=1nρτ{Yi-a-b(Z^i-z^)}K{(Z^i-z^)/h}.

Then, Ĝτ() = â and G^τ(z^)=b^.

In the sequel we derive the asymptotic property for the nonparametric estimate of Gτ(z). For any fixed x0 ∈ ℝp, the asymptotic normality of the local linear estimator Ĝτ(), is presented in the following theorem, whose proof will be given in the Appendix.

Theorem 3

Under the regularity conditions (i)–(iv) in Appendix A, if h → 0 and nh → ∞, then the asymptotic conditional bias and variance of the local linear quantile regression estimator Ĝτ() are given by

bias{G^τ(z^)}=Gτ(z)μ2h2/2+op(h2),var{G^τ(z^)}=(nh)-1Δ{1+op(1)},

where Δ=τ(1-τ)ν0/{f(z)fY2(ξτz)}=τ(1-τ)ν0σ2(z)/{f(z)fε2(ξτ)},ν0=-11K2(u)du, fε is the density function of ε and ξτ is its τ-th quantile. Furthermore, as n → ∞, h → 0 and nh → ∞,

(nh)1/2{G^τ(z^)-Gτ(z)-Gτ(z)μ2h2/2}

is asymptotically normal with mean zero and asymptotic variance Δ.

For univariate x, the asymptotic bias and variance of nonparametric quantile regression can be found in Fan, Hu and Truong (1994) and Yu and Jones (1998). We compared the asymptotic bias and variance of the local linear quantile regression estimator built upon {( xiTβ^τ, Yi), i = 1, …, n} with that of the estimator built upon {( xiTβτ, Yi), i = 1, …, n} in Fan, Hu and Truong (1994), and found they performed equally well asymptotically. This implies that the resulting estimate of the newly proposed procedure performs as well as an oracle estimate with replacing β̂τ with its true value.

Next we discuss how to discuss the issue of bandwidth selection. One usually chooses the bandwidth which minimizes the mean squared error (MSE) of nonparametric estimation. Theorem 3 indicates that the leading term of MSE of Ĝτ() is

MSE{G^τ(z^)}=14{μ2Gτ(z)}2h4+ν0nhf(z)τ(1-τ)fε2(ξτ)σ2(z).

Thus, the optimal variable bandwidth which minimizes the asymptotic MSE of Ĝτ() is

hopt(z^)=[4ν0τ(1-τ){μ2fε(ξτ)}2σ2(z)/f(z){Gτ(z)}2]1/5n-1/5. (2.4)

This implies that the local quantile regression achieves the optimal rate of convergence n2/5. To implement hopt(), one has to replace all unknowns in (2.4) with their consistent estimators. This is usually computationally inefficient although all involved unknown quantities are univariate nonparametric functions.

In the sequel we introduce a computationally efficient way to calculate the bandwidth for the quantile regression. We notice from Theorem 3 that the leading term of the asymptotic bias for the local linear quantile regression is the same as that for the local linear least squares estimator, whereas their asymptotic variances are different. The local least squares estimator (Fan and Gijbels, 1996) has the MSE of the form

MSE{G^LS(z^)}=14{μ2G(z)}2h4+ν0nhf(z)σ2(z),

which results in the optimal bandwidth

hmopt(z^)=[4ν0μ22σ2(z)/f(z){G(z)}2]1/5n-1/5. (2.5)

Comparing (2.5) with (2.4), we obtain that

hopt(z^)=hmopt(z^){τ(1-τ){fε(ξτ)}2{G(z)}2{Gτ(z)}2}1/5.

We next introduce a rule of thumb bandwidth selector. If the curvatures of the τ-th quantile function Gτ(z) and the conditional mean function G″(z) are similar, and the error ε is close to a standard normal random variable, then we can choose the approximately optimal bandwidth through

hopt(z^)=hmopt(z^){τ(1-τ){φ(ξτ)}2}1/5, (2.6)

where φ(·) denotes the probability density function of the standard normal distribution. There are many algorithms for calculating hmopt(z^) (Fan and Gijbels, 1996), thus hopt() can be rapidly calculated. One can estimate ξτ through the sample τ-th quantile of the residuals ε^i={Yi-G^(xiTβ^τ)}/σ^(xiTβ^τ),, for i = 1, …, n. An alternative way is to replace ξτ in (2.6) with the τ-th quantile of the standard normal population to further simplify the calculation.

The idea of this approximation is originated from Yu and Jones (1998) and Yu and Lu (2004). Although this approximation is built upon several assumptions, it provides a computationally efficient way to calculate the bandwidth for the quantile regression. We find it performs quite well in usual practice.

3. Simulations and Application

In this section we conduct simulations to compare the performance of the proposed methods with existing competitors and to illustrate the proposed methodology with a real data example.

3.1. Simulation

We generated 1000 datasets, each consisting of n = 500 observations, from

Y=sin{2(xTβ0)}+2exp{-16(xTβ0)2}+σ(xTβ0)ε, (3.1)

where the index parameter β0=(2,-2,-1,1,0,,0)T/10 is a p × 1 vector, and the covariate vector x = (X1, …, Xp)T is generated from a multivariate normal population with mean zero and variance-covariance matrix var(x) = (σij)p×p with σij = 0.5|ij|. The conditional mean function in model (3.1) was designed by Kai, Li and Zou (2010). In our simulations, we chose p = 10, 20 and 50, and considered five error distributions for ε: (i) the standard normal distribution N(0, 1); (ii) a mixture of two normal distributions 0.8N(0, 1) + 0.2N(0, 9); (iii) the Laplace distribution; (iv) the student-t distribution with 3 degrees of freedom t(3); and (v) the Cauchy distribution. The error term ε and the covariates x are mutually independent. We considered two cases for σ(xTβ0): σ(xTβ0) = 1 in homoscedasticity case (1), and σ(xTβ0) = exp (xTβ0) in heteroscedasticity case (2). We considered case (2) because it is often of interest to investigate the effect of heteroscedastic errors in quantile regressions. The response values contain outliers with larger probabilities in case (2) than in case (1). We compared the performance of our proposed procedures with the back-fitting algorithm proposed by Wu, Yu and Yu (2010). From our limited simulation experience, the back-fitting algorithm demands much computing time. We compared the simulation results of the back-fitting algorithm based on 100 and 1000 replications for one case, and found that the simulation results based on 100 replications are almost the same as those based on 1000 replications. Thus, we report the results of the back-fitting algorithm based on 100 replications in order to save computing time.

Performance in estimating β0

The index parameter β0 was estimated via a series of quantile regressions with τ = 0.025, 0.05, 0.5, 0.95 and 0.975, respectively. Table 3.1 depicts the averages of mean squared errors (MSE) of the estimate β̂τ, which is defined by MSE = ||β̂τ/||β̂τ||−β̂τ||2. We include three competitors into our comparison when p = 10: (i) the linear quantile regression formulated in (2.3); (ii) the back-fitting algorithm proposed by Wu, Yu and Yu (2010); and (iii) the ordinary least squares estimate (LSE). Li and Duan (1989) proved that this LSE is a consistent estimator of the direction of β0; thus, it serves naturally as a competitor here. It can be seen from Table 3.1 that the back-fitting algorithm performs the best in most scenarios. This is expected in that the back-fitting algorithm updates the nonparametric quantile function in a data-driven manner while estimating β0. By contrast, the linear quantile regression assumes a linear quantile function to reduce the computational complexity. Yet the linear quantile regression has also a satisfactory performance. In the particular homoscedastic case (1) with Cauchy errors, the linear quantile regression performs even better than the back-fitting algorithm in terms of the MSE values. In the homoscedasticity case (1) with standard normal errors, the LSE performs slightly better than the linear quantile regression. However, the linear quantile regression is superior to the LSE in all other cases. The improvement of linear quantile regression over the LSE is more significant in heteroscedasticity case (2) than in homoscedasticity case (1). This is expected because in that the performance of the LSE is sensitive to the presence of outliers. The quantile regression offers a more robust estimation in most scenarios.

Table 3.1.

The averages of MSEs of β̂τ for model (3.1) with p = 10.

case (1): σ (xTβ0) = 1 case (2): σ (xTβ0) = exp (xTβ0)

the τ-th quantile in % the τ-th quantile in %
2.5 5 50 LSE 95 97.5 2.5 5 50 LSE 95 97.5
Linear quantile regression

ε ~ N(0, 1)
0.200 0.191 0.140 0.113 0.292 0.305 0.173 0.164 0.107 0.265 0.230 0.245
ε ~ 0.8N(0, 1) + 0.2N(0, 9)
0.259 0.246 0.173 0.197 0.347 0.362 0.236 0.223 0.145 0.484 0.289 0.301
ε ~ Laplace distribution
0.144 0.136 0.097 0.110 0.219 0.230 0.127 0.121 0.082 0.255 0.186 0.197
ε ~ t(3)
0.249 0.239 0.171 0.217 0.344 0.359 0.222 0.210 0.139 0.498 0.290 0.303
ε ~ Cauchy distribution
0.376 0.359 0.237 1.441 0.452 0.474 0.371 0.354 0.215 1.430 0.420 0.442

Back-fitting algorithm

ε ~ N(0, 1)
0.137 0.118 0.069 0.164 0.195 0.105 0.094 0.048 0.117 0.135
ε ~ 0.8N(0, 1) + 0.2N(0, 9)
0.319 0.232 0.086 0.284 0.316 0.136 0.112 0.069 0.163 0.184
ε ~ Laplace distribution
0.119 0.104 0.045 0.128 0.167 0.098 0.087 0.035 0.115 0.125
ε ~ t(3)
0.269 0.218 0.089 0.260 0.314 0.162 0.146 0.083 0.142 0.168
ε ~ Cauchy distribution
0.598 0.612 0.164 0.645 0.655 0.336 0.257 0.158 0.305 0.394

Performance in coverage probability of prediction intervals

It is often of interest to evaluate the accuracy of quantile regression in offering the prediction interval of Y given xTβ0. To this end, we can estimate the τ-th and (1 − τ)-th quantile of the conditional distribution of Y given xTβ0, which can be used as the confidence interval at the level of (1 − 2τ) for τ < 0.5. That is, we use [G−1(τ|xTβ0), G−1(1 − τ|xTβ0)] to construct the prediction interval for Y|xTβ0 at the (1 − 2τ)×100% level, where G−1(τ|xTβ0) is the inverse function of Gτ(Y|xTβ0). That is, G−1(τ|xTβ0) is the τ-th quantile of Y|xTβ0. We choose τ = 0.025 and 0.05 to provide the confidence intervals at 95% and 90%. We expect the empirical coverage probability of such confidence interval to be close to the nominal level of (1 − 2τ).

We included three competitors into our comparison: (i) The back-fitting algorithm (Wu, Yu and Yu, 2010) which can simultaneously estimate the index parameter βτ and the quantile function; (ii) the quantile regression method introduced in Section 2.2 which is based on the estimated index parameter β̂τ; and (iii) the oracle estimator which provides the prediction interval by using the true index parameter β0. The oracle estimator serves as a benchmark for comparison purposes. The average and the standard deviation of the empirical coverage probabilities are summarized in Table 3.2.

Table 3.2.

The empirical coverage probabilities for model (3.1) with p = 10.

case (1): σ (xTβ0) = 1 case (2): σ (xTβ0) = exp (xTβ0)

quantile regression Back-fitting quantile regression Back-fitting

New Oracle algorithm New Oracle algorithm

Level 90% 95% 90% 95% 90% 95% 90% 95% 90% 95% 90% 95%
ε ~ N(0, 1)

aver. 89.4 94.5 89.8 94.8 90.5 95.5 90.1 95.1 90.9 95.5 91.1 95.9
stdev 0.5 0.4 0.5 0.4 0.5 0.4 0.5 0.4 0.5 0.4 0.5 0.4

ε ~ 0.8N(0, 1) + 0.2N(0, 9)

aver. 89.7 94.7 89.9 94.7 90.6 95.5 90.3 95.0 90.7 95.9 91.1 95.8
stdev 0.4 0.3 0.5 0.3 0.5 0.4 0.4 0.4 0.5 0.4 0.4 0.5

ε ~ Laplace distribution

aver. 89.7 94.7 89.8 94.8 90.6 95.6 90.3 95.1 90.9 95.5 91.1 95.9
stdev 0.4 0.3 0.5 0.4 0.6 0.4 0.4 0.4 0.5 0.4 0.5 0.4

ε ~ t(3)

aver. 89.7 94.7 89.8 94.8 90.6 95.5 90.3 95.1 90.8 95.4 91.2 95.9
stdev 0.4 0.3 0.5 0.4 0.6 0.4 0.4 0.4 0.5 0.4 0.5 0.4

ε ~ Cauchy distribution

aver. 90.4 95.4 90.4 95.7 90.9 95.7 90.8 95.6 91.0 95.7 91.2 95.9
stdev 0.5 0.4 0.5 0.5 0.7 0.4 0.5 0.4 0.5 0.4 0.5 0.5

Both the homoscedasticity and the heteroscedasticity cases reported in Table 3.2 show very similar messages. The coverage probabilities of three competitors are all close to the nominal levels for different τ values, although the back-fitting algorithm uses a more accurate estimate of β0. The standard deviations of the empirical probabilities are very small, indicating that the quantile regression offers a very reliable coverage probability, which is comparable with its oracle version. This phenomenon verifies the theoretical results in Theorem 3.

Computational efficacy

It is of interest to compare the computational efficacy of the newly proposed method with that of the back-fitting algorithm proposed by Wu, Yu and Yu (2010). Table 3.3 summarizes the average computing time in seconds used for estimating the index parameter and calculating the prediction intervals for one replication. It can be easily seen from Table 3.3 that the newly proposed method is much faster than the back-fitting algorithm. In addition, the computing time of the back-fitting algorithm varies over the error distributions.

Table 3.3.

The averages of computing times (in seconds) for model (3.1) with p = 10.

Case (1): σ (xTβ0) = 1 (2): σ (xTβ0) = exp (xTβ0)

Error New back-fitting New back-fitting
N(0, 1) 0.11 1018.57 0.10 1029.58
0.8N(0, 1) + 0.2N(0, 9) 0.11 1472.64 0.11 1490.92
Laplace distribution 0.11 1531.51 0.09 1551.81
t(3) 0.12 1386.17 0.10 1401.70
Cauchy distribution 0.13 1429.88 0.11 1426.16

Performance in high dimension

Next we investigate the performance of the proposed method when p is relative large. We still use model (3.1) for illustration purposes but increase the dimension p of the covariates from 10 to 20 and 50. Tables 3.4 and 3.5 report respectively the MSE values and the coverage probabilities for model (3.1) with p = 20 and 50. Table 3.4 together with Table 3.1 indicates that the MSE values of β̂τ increase significantly as the covariate dimension p increases due to the curse of dimensionality. However, the covariate dimension p has little effect on estimating the coverage probabilities of the prediction intervals, as can be seen from Tables 3.2 and 3.5. This benefits essentially from the single-index structure of the quantile function. These simulation results also partially support our theoretical investigations in Theorems 2 and 3, and confirm the advantage of our proposed methodology in high-dimensional setting.

Table 3.4.

The averages of MSEs of β̂τ for model (3.1) with p = 20 and 50.

case (1): σ (xTβ0) = 1 case (2): σ (xTβ0) = exp (xTβ0)

the τ-th quantile in % the τ-th quantile in %
2.5 5 50 LSE 95 97.5 2.5 5 50 LSE 95 97.5
p = 20

ε ~ N(0, 1)
0.380 0.365 0.271 0.230 0.516 0.535 0.371 0.353 0.231 0.439 0.450 0.468
ε ~ 0.8N(0, 1) + 0.2N(0, 9)
0.476 0.458 0.345 0.380 0.619 0.640 0.483 0.460 0.300 0.722 0.539 0.561
ε ~ Laplace distribution
0.294 0.280 0.213 0.232 0.438 0.458 0.290 0.275 0.185 0.446 0.400 0.419
ε ~ t(3)
0.444 0.427 0.331 0.387 0.606 0.628 0.470 0.445 0.290 0.736 0.535 0.559
ε ~ Cauchy distribution
0.651 0.624 0.447 1.633 0.754 0.778 0.701 0.671 0.431 1.597 0.708 0.733

p = 50

ε ~ N(0, 1)
0.757 0.733 0.586 0.504 0.941 0.967 0.840 0.804 0.539 0.780 0.863 0.888
ε ~ 0.8N(0, 1) + 0.2N(0, 9)
0.895 0.868 0.689 0.718 1.041 1.067 1.029 0.991 0.660 1.065 0.970 0.995
ε ~ Laplace distribution
0.670 0.643 0.496 0.506 0.878 0.906 0.704 0.671 0.465 0.756 0.817 0.844
ε ~ t(3)
0.890 0.859 0.667 0.734 1.008 1.034 1.009 0.972 0.656 1.083 0.972 0.997
ε ~ Cauchy distribution
1.175 1.145 0.889 1.777 1.253 1.283 1.313 1.275 0.865 1.749 1.165 1.189
Table 3.5.

The empirical coverage probabilities for model (3.1) with p = 20 and 50.

Case (1): σ (xTβ0) = 1 (2): σ (xTβ0) = exp (xTβ0)

New Oracle New Oracle

Error Level 90% 95% 90% 95% 90% 95% 90% 95%
p = 20

N(0,1) aver. 90.3 95.3 90.9 95.9 90.8 95.7 91.4 96.1
stdev 0.4 0.4 0.5 0.4 0.5 0.3 0.5 0.4

0.8N(0, 1) +0.2N(0, 9) aver. 90.4 95.3 90.9 95.7 90.8 95.6 91.3 95.9
stdev 0.4 0.3 0.5 0.4 0.4 0.3 0.5 0.5

Laplace aver. 90.4 95.4 91.0 95.8 90.8 95.7 91.4 96.0
stdev 0.4 0.3 0.5 0.5 0.4 0.3 0.5 0.5

t(3) aver. 90.3 95.3 90.8 95.7 90.8 95.6 91.2 96.0
stdev 0.4 0.4 0.5 0.4 0.5 0.4 0.5 0.4

Cauchy aver. 90.6 95.3 90.6 95.5 90.6 95.4 91.0 95.7
stdev 0.7 0.4 0.9 0.7 0.5 0.4 0.4 0.4

p = 50

N(0, 1) aver. 90.3 95.3 90.9 95.8 90.7 95.6 91.3 96.2
stdev 0.4 0.3 0.5 0.4 0.4 0.4 0.4 0.5

0.8N(0, 1) +0.2N(0, 9) aver. 90.4 95.3 90.9 95.7 90.6 95.5 91.2 95.9
stdev 0.4 0.3 0.5 0.4 0.5 0.3 0.5 0.4

Laplace aver. 90.3 95.3 91.0 95.8 90.6 95.5 91.3 96.1
stdev 0.4 0.3 0.6 0.5 0.4 0.3 0.5 0.4

t(3) aver. 90.4 95.3 90.9 95.7 90.7 95.5 91.3 96.0
stdev 0.4 0.3 0.5 0.4 0.5 0.4 0.5 0.4

Cauchy aver. 90.6 95.5 90.4 95.4 90.6 95.5 91.1 95.8
stdev 0.5 0.4 0.8 0.7 0.5 0.4 0.7 0.6

3.2. An application

We illustrate the proposed procedures by an empirical analysis of an automobile dataset (Johnson, 2003). It is of interest to know how the manufactures’ suggested retail price (MSRP) of vehicles depends upon different levels of attributes such as miles per gallon and horsepower. Thus the MSRP serves naturally as the response variable. This is the manufacturer’s assessment of a vehicle’s worth, and includes adequate profit for the manufacturer and the dealer. We use the logarithm transformation of the MSRP in U.S. dollars as the response (Y). In addition, there are seven features which possibly affect the resulting price of vehicles, such as the engine size (X1), number of cylinders (X2), horsepower (X3), average city miles per gallon (MPG, X4) and average highway MPG (X5), weight in pounds (X6) and wheel base in inches (X7). This dataset consists of 428 observations, sixteen of which have missing values. We remove those observations with missing values in our subsequent analysis. The remaining dataset has a total of eight variables and 412 observations. Both the response variable Y and the covariate vector x = (X1, …, X7)T are standardized marginally to have zero mean and unit variance in this empirical analysis. The fully nonparametric regression is not applicable to this particular case because the sample size is small compared with the dimensionality of the covariates.

We specify the heteroscedastic single-index model (1.1) for analyzing the new vehicle data. The index parameter is estimated at five different quantiles: τ = 0.025, 0.05, 0.50, 0.95 and 0.975. They yield similar estimates for the index parameter βτ at different quantiles, indicating that similar patterns between the MSRP and different features are revealed at different quantiles. Specifically, the estimated index parameter at τ = 0.5 is β̂τ = (0.3254, −0.1846, −0.6332, 0.3688, −0.3326, −0.4156, 0.1994)T. Table 3.6 reports that the standard deviation of β̂τ obtained from a nonparametric bootstrap procedure is (0.052, 0.046, 0.051, 0.069, 0.057, 0.043, 0.032)T. This implies that the horsepower (X3) is perhaps the most important factor that affects the suggested retail price (Y), followed by the weight in pounds (X6).

Table 3.6.

The average and standard deviation of the estimated index parameter based on 500 bootstrap samples of the new vehicle data.

εi sampled from X1 X2 X3 X4 X5 X6 X7
Fn aver. 0.325 −0.185 −0.633 0.369 −0.333 −0.416 0.199
stdev. 0.052 0.046 0.051 0.069 0.057 0.043 0.032

.5Fn+ .5N(0, .43322) aver. 0.305 −0.173 −0.630 0.351 −0.341 −0.428 0.191
stdev. 0.058 0.051 0.051 0.073 0.060 0.047 0.036

.5Fn + .5Cauchy aver. 0.303 −0.169 −0.628 0.314 −0.318 −0.421 0.184
stdev. 0.088 0.081 0.129 0.136 0.110 0.098 0.062

Using the data {( xiTβ^τ, Yi), i = 1, …, n}, we first estimate the regression function G(xTβ0) via the local least squares estimator. The estimated function and its 95% confidence interval are reported in Figure 3.1(C). We further conduct some exploratory data analysis on this dataset. The boxplot of the log-transformed MSRP, the response variable, is depicted in Figure 3.1(A), which shows that there are four vehicles whose prices are significantly higher than other vehicles. In addition, the histogram in Figure 3.1(B) reveals that the distribution of the log-transformed prices are highly skewed. This motivates us to further conduct empirical analysis of this dataset using the proposed procedure.

Figure 3.1.

Figure 3.1

(A) Boxplot of log-transformed retail price. (B) Histogram of log-transformed retail price. (C) Plot of Ĝ(·) and its 95% confidence interval using local linear least squares method. (D) Plot of Ĝτ(·) with τ = 2.5%, 50% and 97.5%. The 95% prediction interval of Y given x covers 96% points

Next we apply the quantile regression to this dataset. Figure 3.1(D) presents the 2.5% and 97.5% quantiles to give an approximate 95% prediction interval of Y. This prediction interval covers 95.87% of the data points, which is close to the nominal level. The line in the middle of Figure 3.1(D) presents the quantile regression at the level of 50%. It behaves similarly to the local least squares estimator in Figure 3.1(C) within the range of xTβ̂τ. When xTβ̂τ is near to the boundary, however, the local least squares estimate is clearly more sensitive to the outliers than the quantile estimate. This example demonstrates the effectiveness of the proposed procedure in terms of description of the conditional distribution of Y given x.

We further use a bootstrap procedure to demonstrate the performance of our method in the presence of outliers. Given the covariates we bootstrap new response variables by using the residuals. To calculate the residuals, we fix τ = 0.5 and estimate βτ and Gτ(xTβτ) to obtain ε^i=Yi-G^τ(xiTβ^τ) for i = 1, …, n. For ease of illustration, we denote by Fn the empirical distribution of the residuals ε̂i. We bootstrap 500 samples each of form {(xi, Yi), i = 1, …, n} where Yi=G^τ(xiTβ^τ)+ε^i, and εi is generated from three different distributions: (1) Fn; (2) 0.5Fn+0.5N(0, 0.43322); we consider this mixture because the standard deviation of the original residuals is exactly 0.4332; and (3) 0.5Fn + 0.5Cauchy. In case (3) the distribution of the response variable has a heavy tail.

We estimate the index parameter βτ at different quantiles, and summarize the average and standard deviation of β̂τ based on the bootstrap samples. Because the results for different quantiles show similar messages, we only report the results for τ = 0.5, which are depicted in Table 3.6. The averages of the estimated index parameter β̂τ are similar. This once again verifies that our method is robust to the presence of outliers. By contrast, the standard deviations of these three cases are slightly different. To be precise, the standard deviations of β̂τ in case (1) and (2) are very similar, both of which have smaller standard deviations than case (3). This can be interpreted to mean that the bootstrapped residuals in case (3) have substantially larger variance than in cases (1) and (2).

4. Discussion

To estimate the quantile regression with high-dimensional covariates, we impose a heteroscedastic single-index structure on the regression function. The heteroscedastic single-index structure allows us to reduce the dimension of covariates and simultaneously retains the flexibility of nonparametric regression. We propose a computationally efficient two-step estimation procedure to estimate the parameters involved in the quantile regression function. Asymptotic properties of the proposed procedures are studied.

It is remarkable here that, if (1.1) reduces to the homoscedastic single-index model (i.e., σ(·) remains constant), and the error ε is symmetric about zero, then a byproduct of this paper is that our estimation procedure offers a robust estimation for the conditional mean function. The simulation studies in this article verify this point implicitly. In addition, it is not necessary that the error term ε in model (1.1) is independent of x, in which case we can assume instead that E(x|xTβ0, ε) is a linear function of xTβ0. If the parameter of interest is the index parameter β0 only, the composite quantile regression (Zou and Yuan, 2008) can be adapted to improve the efficiency in estimating β0.

In this paper we suggest quantile regression to construct prediction intervals. This is in spirit a local pointwise prediction interval. How to construct global confidence band is an interesting question. Härdle and Song (2010) investigated the confidence band in univariate quantile regression. Their results are readily applicable to the single-index model (1.1) when β0 is known. Further research is needed to quantify the effect when β0 is replaced with its root-n consistent estimate.

Acknowledgments

Zhu was supported by a National Natural Science Foundation of China grant 11071077 and a National Institute on Drug Abuse (NIDA) grant R21-DA024260. Huang is the corresponding author and was supported by a funding through Pro-jet 211 Phase 3 of SHUFE, and Shanghai Leading Academic Discipline Project, B803. Li was supported by NSF grant DMS 0348869, National Natural Science Foundation of China grants 11028103 and and 10911120395, and National Institute on Drug Abuse (NIDA) grant P50-DA10075. The content is solely the responsibility of the authors and does not necessarily represent the official views of the NIDA or the NIH.

Appendix: Technical Conditions and Proofs

A.1. Technical conditions

The following technical conditions are imposed. They are not the weakest possible conditions, but they are imposed to facilitate the proofs.

  1. The quantile function Gτ(·) has a continuous and bounded second derivative.

  2. The density function f(·) of (xTβ) is positive and uniformly continuous for β in a neighborhood of β0. Furthermore, the density function of (xTβ0) is continuous and bounded away from zero and infinity on its support.

  3. The conditional density of Y given (xTβ0), denoted by fY(y|xTβ0), is continuous in (xTβ0) for each y ∈ ℝ. Moreover, there exist positive constants ε and δ and a positive function F(y|xTβ0) such that
    supxTβ-xTβ0εfY(yxTβ)F(yxTβ0).

    For any fixed value of (xTβ0), we assume that ∫ F(y|xTβ0) dy < ∞ and {ρτ(y-t)-ρτ(y)-ρτ(y)t}2F(yxTβ0)dy=0(t2), as t − 0.

  4. The kernel function K(·) is symmetric and has compact support [−1, 1]. It satisfies the first-order Lipschitz condition.

We briefly discuss these conditions. Condition (i) is commonly assumed for a link function in the literature. Condition (ii) guarantees the existence of any ratio terms with the density function f(xTβ0) appearing in the denominator. Condition (iii) is weaker than the Lipschitz condition of the function ρτ(·). The check loss function ρτ(·) is piecewise linear and non-differentiable at 0. We define that ρτ(u)=sign(u)+(2τ-1) for u ≠ 0 and ρτ(0)=0. With Taylor’s expansion, we obtain {ρτ(y-t)-ρτ(y)-ρτ(y)t}2F(yxTβ0)dy=t2{sign(u)-sign(y)}2F(yxTβ0)dy for some u* lies between yt and y. Without loss of generality we assume t > 0, then

-{sign(u)-sign(y)}2F(yxTβ0)dy=0t{sign(u)-sign(y)}2F(yxTβ0)dy40tF(yxTβ0)dy.

Thus, {ρτ(y-t)-ρτ(y)-ρτ(y)t}2F(yxTβ0)dy=o(t2) is implied by 0tF(yxTβ0)dy as t → 0. Thus condition (iii) is mild. Condition (iv) requires that the kernel function be a proper density function with compact support.

A.2. Proof of Theorem 1

Theorem 1 follows from the convexity of the check loss function and the linearity condition (2.2). It suffices to prove that, for any constant u ∈ ℝ, there exists a constant κ such that Inline graphic(u, β), ≥ Inline graphic(u, κβ0). Specifically,

Lτ(u,β)=E[E{ρτ(Y-u-βTx)β0Tx,ε}]E[ρτ{E(Yβ0Tx,ε)-u-E(βTxβ0Tx,ε)}]=E[ρτ{Y-u-E(βTxβ0Tx)}]=E{ρτ(Y-u-κβ0Tx)},

where κ=βTvar(x)β0/{β0Tvar(x)β0}. The first equality follows from the iterative law of conditional expectation; the first inequality follows from Jensen’s inequality and the convexity of ρτ; the second equality is true in view of model (1.1) and the last equality holds true by invoking the linearity condition (2.2) and the error ε is independent of x. This completes the proof.

A.3. Proof of Theorem 2

This proof follows from the quadratic approximation lemma (Hjort and Pollard, 1993). Let us present the quadratic approximation lemma here to facilitate the proof.

Quadratic Approximation Lemma

Suppose Inline graphic(β) is convex and can be represented as βTBβ/2+unTβ+an+Rn(β), where B is symmetric and positive definite, un is stochastically bounded, an is arbitrary, and Rn(β) goes to zero in probability for each β. Then β̂, the minimizer of Inline graphic(β), is only op(1) away from −B−1un. If un converges in distribution to u, then β̂ converges in distribution to −B−1u accordingly.

Now we use the quadratic approximation lemma to prove Theorem 2. Let αn = n−1/2 and a = n1/2(β̂τβτ). Further, let Lτn(u,β)=n-1i=1nρτ(Yi-u-xiTβ). By applying the identity (Knight, 1998) that

ρτ(x-y)-ρτ(x)=2[y{1(x0)-τ}+0y{1(xt)-1(x0)}dt], (A.1)

it follows that

Lτn(ξτ+αnb,βτ+αna)-Lτn(ξτ,βτ)=n-1i=1nρτ(Yi-ξτ-αnb-xiTβτ-αnxiTa)-n-1i=1nρτ(Yi-ξτ-xiTβτ)=n-1i=1nαn(xiTa+b){1(Yi-xiTβτξτ)-τ}+n-1i=1n0αn(xiTa+b){1(Yi-xiTβτξτ+t)-1(Yi-xiTβτξt)}dt.

The objective function Inline graphic(ζτ + αnb, βτ + αna) − Inline graphic(ζτ, βτ) is convex in (a, b). With a slight change of notation, we denote F(t|x) = E{1(YxTβτt)|x}. By Taylor’s expansion, it follows that

E{Lτn(ξτ+αnb,βτ+αna)-Lτn(ξτ,βτ)x1,,xn}=n-1i=1n[αn(xiTa+b){F(ξτxi)-τ}+ξτξτ+αn(xiTa+b)F(txi)-F(0xi)dt]=αnni=1n[(xiTa+b){F(ξτxi)-τ}+αn22f(ξτxi)(b+aTxi)2]+op(αn2). (A.2)

The first two terms are of order Op (1/n). The last term of (A.2) admits a quadratic function of (a, b). Following standard arguments, we obtain that

R2=Lτn(ξτ+αnb,βτ+αna)-Lτn(ξτ,βτ)-E{Lτn(ξτ+αnb,βτ+αna)-Lτn(ξτ,βτ)x1,,xn}=op(n-1).

This, together with the quadratic approximation lemma of Hjort and Pollard (1993) and (A.2), leads to the asymptotic normality of β̂.

A.4. Proof of Theorem 3

Recall the notations z=x0Tβτ,z^=x0Tβ^τ,Zi=xiTβτ,Z^i=xiTβ^τ. Let a0=Gτ(z)=Gτ(x0Tβτ),b0=Gτ(z),K^i=K{(Z^i-z^)/h}, i = K{(i)/h}, Ki = K {(Ziz)/h} and Ki=K{(Zi-z)/h} for i = 1, …, n.

We first show that, for β̂τ to be a root-n consistent estimate of βτ,

1ni=1n[ρτ{Yi-a-b(Z^i-z^)}K^i-ρτ{Yi-a-b(Zi-z)}Ki]=Op(n-1/2) (A.3)

Define

R11=n-1i=1nKi[ρτ{Yi-a-b(Z^i-z^)}-ρτ{Yi-a-b(Zi-z)}],R12=n-1i=1n(K^i-Ki)ρτ{Yi-a-b(Zi-z)},R13=E[ρτ{Yi-a-b(Z^i-z)}-ρτ{Yi-a-b(Zi-z)}](K^i-Ki).

Then, the left hand side of (A.3) can be decomposed into R11 + R12 + R13. Using the identity ρτ(x-y)-ρτ(x)=2[y{1(x0)-τ}+0y{1(xz)-1(x0)}dz], we have |ρτ(xy) − ρτ(x)| ≤ 4|y|. Invoking the root-n consistency of β̂τ, we obtain

n-1i=1n[Ki|ρτ{Yi-a-b(Z^i-z^)}-ρτ{Yi-a-b(Zi-z)}|]4bsupuK(u)n-1i=1n|Z^i-Zi|=4bn-1i=1nxiT(β^τ-βτ)=Op(n-1/2).

This indicates that |R11| = Op(n−1/2).

Next we deal with R12. R12 can be decomposed into

(β^τ-βτ)T(nh)-1i=1nK(Zi-zh)(xi-x0)ρτ{Yi-a-b(Zi-z)}{1+op(1)}.

Let zi = (xix0)/h. Then

(nh)-1i=1nK(Zi-zh)(xi-x0)ρτ{Yi-a-b(Zi-z)}=n-1i=1nK(ziTβτ)ziρτ(Yi-a-bhziTβτ)=Op(1),

which, together with the root-n consistency of β̂τ, proves that R12 = Op(n−1/2)

It remains to investigate the order of R13. Following similar arguments to deal with R11 and R12, we have R13 = Op(n−1/2).

Thus, (A.3) follows by combining the results of R1k’s, which implies that

n-1i=1nρτ{Yi-a-b(Z^i-z^)}K^i=n-1i=1nρτ{Yi-a-b(Zi-z)}Ki+op{(nh)-1/2}.

That is, the quantile regression estimate of Ĝ(·) based on {( xiTβ^τ, Yi), i = 1, …, n} is asymptotically as efficient as that based on {( xiTβ0, Yi), i = 1, …, n}.

The rest of the proof follows literally from Theorem 3 of Fan, Hu, and Truong (1994). Based on the auxiliary data points {( xiTβ0, Yi), i = 1, …, n}, the technique for establishing the asymptotic normality involves the convexity lemma (Pollard, 1991) and the quadratic approximation lemma (Hjort and Pollard, 1993). We only sketch the outline below.

For the sake of notational clarity, we write ξi = (aa0) + (bb0)(Ziz). By using (A.1) we obtain immediately that

i=1nρτ{Yi-a-b(Zi-z)}Ki-i=1nρτ{Yi-a0-b0(Zi-z)}Ki=2i=1n[ζi{1(Yi-a0-b0(Zi-z)0)-τ}+0ζi{1(Yi-a0-b0(Zi-z)t)-1(Yi-a0-b0(Zi-z)0)}dt]Ki. (A.4)

Let β=nh{a-a0,h(b-b0)}T. In the sequel, we show the first quantity in (A.4) approximately has the form 2unTβ, and the second quantity has the form βTBβ. Then we use the quadratic approximation lemma to complete the proof.

Denote by FY (y|Z) the conditional distribution of Y given Z. Then

2i=1nζi{1(Yi-a0-b0(Zi-z)0)-τ}Ki=2i=1nζi{FYi(a0+b0(Zi-z)Zi)-τ}Ki+2i=1nζi{1(Yi-a0-b0(Zi-z)0)-FYi(a0+b0(Zi-z)Zi)}Ki. (A.5)

The first term in (A.5) is a bias. To view this, we use Taylor’s expansion,

Gτ(Zi)=a0+b0(Zi-z)+Gτ(z)(Zi-z)2{1/2+op(1)},forZi-zh.

This together with another Taylor’s expansion entails that

FYi(a0+b0(Zi-z)Zi)=FYi(Gτ(Zi)-Gτ(z)(Zi-z)2{1/2+op(1)}Zi)=τ-fYi(Gτ(Zi)Zi)Gτ(z)(Zi-z)2{1/2+op(1)}.

With this result, this bias term can be simplified as follows.

2i=1nζi{FYi(a0+b0(Zi-z)Zi)-τ}Ki=-i=1nζifYi(Gτ(Zi)Zi)Gτ(z)(Zi-z)2{1+op(1)}Ki,

which is of the form

-1nhi=1nβT(1(Zi-z)/h)fYi(Gτ(Zi)Zi)Gτ(z)(Zi-z)2{1+op(1)}Ki. (A.6)

The second term in the right hand of (A.5) can be approximated as follows.

2i=1nζi{1(YiGτ(Zi))-τ}Ki{1+op(1)},

which is again of the form

2nhi=1nβT(1(Zi-z)/h){1(YiGτ(Zi))-τ}Ki{1+op(1)}. (A.7)

Let fY(y|Z) be the conditional density function of Y given Z. Then the second term in the right hand side of (A.4) can be approximated as follows.

2i=1n[0ζi{1(Yi-a0-b0(Zi-z)t)-1(Yi-a0-b0(Zi-z)0)}dt]Ki=i=1nfYi(a0+b0(Zi-z)Zi)ζi2Ki{1+op(1)}=i=1nfYi(Gτ(Zi)Zi)ζi2Ki{1+op(1)},

which is now of the form

1nhi=1nfYi(Gτ(Zi)Zi)βT(1{(Zi-z)/h}{(Zi-z)/h}{(Zi-z)/h}2)βKi{1+op(1)}. (A.8)

Combining the results of (A.5)–(A.8), we observe that (A.4) is now approximated by a quadratic function of β. By the quadratic approximation lemma, β is only op(1) away from −B−1un, where

B=1nhi=1nfYi(Gτ(Zi))(1{(Zi-Z)/h}{(Zi-z)/h}{(Zi-z)/h}2)Ki,

and

un=nh[1nhi=1n(1{(Zi-z)/h}){1(YiGτ(Zi))-τ}Ki-1nhi=1n(1{(Zi-z)/h})fYi(Gτ(Zi)Zi)Gτ(z)(Zi-z)2Ki/2].

It can be easily shown that, as n − ∞, B converges in probability to

fY(Gτ(z)z)fZ(z)(100μ2).

Because our target is the quantile function, it suffices to deal with the first element of un because the second element is concerned instead with its derivative. Without much difficulty, we can prove that the first quantity in the curly parenthesis converges in distribution to a normal population with zero mean and variance τ(1-τ)fZ(z)-11K2(u)du and the second quantity in the curly parenthesis converges in probability to h2Gτ(z)μ2{fY(Gτ(z)z)fZ(z)}.

Let fε denote the density function of the error term ε, and ξτ denote its τ-th quantile. We can easily verify that fY(Gτ(z)z)σ(z)=fε(ξτ) when model (1.1) is true. Then the asymptotic variance has a different appearance.

The proof of Theorem 3 is completed by combining the above results.

Contributor Information

Liping Zhu, Email: luz15@psu.edu, School of Statistics and Management, Shanghai University of Finance and Economics, Shanghai, 200433, P. R. China.

Mian Huang, Email: huang.mian@mail.shufe.edu.cn, School of Statistics and Management, Shanghai University of Finance and Economics, Shanghai, 200433, P. R. China.

Runze Li, Email: rli@stat.psu.edu, Department of Statistics and The Methodology Center, The Pennsylvania State University, University Park, Pennsylvania, 16802-2111, USA.

References

  1. Belloni A, Chernozhukov V. ℓ1-penalized quantile regression in high-dimensional sparse models. Annals of Statistics. 2011;39:82–130. [Google Scholar]
  2. Bhattacharya PK, Gangopadhyay AK. Kernel and nearest-neighbor estimation of a conditional quantile regression. Annals of Statistics. 1990;18:1400–1415. [Google Scholar]
  3. Chaudhuri P. Nonparametric estimates of regression quantiles and their local Bahahur representation. Annals of Statistics. 1991;19:760–777. [Google Scholar]
  4. Chaudhuri P, Doksum K, Samarov A. On average derivative quan-tile regression. Annals of Statistics. 1997;25:715–744. [Google Scholar]
  5. De Gooijer JG, Zerom D. On additive conditional quantiles with high dimensional covariates. Journal of the American Statistical Association. 2003;98:135–146. [Google Scholar]
  6. Fan J, Gijbels I. Local Polynomial Modelling and its Applications. London: Chapman and Hall; 1996. [Google Scholar]
  7. Fan J, Hu TC, Truong YK. Robust non-parametric function estimation. Scandinavian Journal of Statistics. 1994;21:433–446. [Google Scholar]
  8. Gannoun A, Girard S, Guinot C, Saracoo J. Sliced inverse regression in reference curves estimation. Computational Statistics and Data Analysis. 2004;46:103–122. [Google Scholar]
  9. Hall P, Li KC. On almost linearity of low dimensional projection from high dimensional data. Annals of Statistics. 1993;21:867–889. [Google Scholar]
  10. Härdle WK, Song S. Confidence bands in quantile regression. Econometric theory. 2010;26:1180–1200. [Google Scholar]
  11. Hjort NL, Pollard D. Asymptotics for minimisers of convex processes. 1993 Preprint available at www.stat.yale.edu/~pollard/Papers/convex.pdf.
  12. Horowitz JL, Lee S. Nonparametric estimation of an additive quantile regression model. Journal of the American Statistical Association. 2005;100:1238–1249. [Google Scholar]
  13. Ichimura H. Semiparametric least squares (SLS) and weighted SLS estimation of single-index models. Journal of Econometrics. 1993;58:71–120. [Google Scholar]
  14. Johnson RW. Kiplinger’s personal finance. Journal of Statistics Education. 2003;57:104–123. [Google Scholar]
  15. Kai B, Li R, Zou H. Local composite quantile regression smoothing: an efficient and safe alternative to local polynomial regression. Journal of the Royal Statistical Society Series B. 2010;72:49–69. doi: 10.1111/j.1467-9868.2009.00725.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Knight K. Limiting distribution for L1 regression estimators under general conditions. Annals of Statistics. 1998;26:755–770. [Google Scholar]
  17. Koenker R. Quantile Regression. Cambridge University Press; 2005. [Google Scholar]
  18. Koenker R, Bassett G. Regression quantiles. Econometrika. 1978;46:33–50. [Google Scholar]
  19. Koenker R, Ng P, Portnoys S. Quantiles smoothing splines. Biometrika. 1994;81:673–680. [Google Scholar]
  20. Kong E, Xia Y. Quantile estimation of a general single-index model. Working paper. 2010 Available at http://arxiv.org/abs/0803.2474.
  21. Li KC. Sliced inverse regression for dimension reduction. Journal of the American Statistical Association. 1991;86:316–327. doi: 10.1080/01621459.2018.1520115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Li KC, Duan N. Regression analysis under link violation. The Annals of Statistics. 1989;17:1009–1052. [Google Scholar]
  23. Pollard D. Asymptotics for least absolute deviation regression estimators. Econometric Theory. 1991;7:186–199. [Google Scholar]
  24. Stone CJ. Consistent nonparametric regression (with discussion) Annals of Statistics. 1977;5:595–645. [Google Scholar]
  25. Wang H, Zhu Z, Zhou J. Quantile regression in partially linear varying coefficient models. Annals of Statistics. 2009;37:3841–3866. doi: 10.1214/07-AOS561. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Wu Y, Liu Y. Variable selection in quantile regression. Statistica Sinica. 2009;19:801–817. [Google Scholar]
  27. Wu T, Yu K, Yu Y. Single-index quantile regression. Journal of Multivariate Statistics. 2010;101:1607–1621. [Google Scholar]
  28. Yu K, Lu Z. Local linear additive quantile regression. Scandinavian Journal of Statistics. 2004;31:333–346. [Google Scholar]
  29. Yu K, Jones MC. Local linear quantile regression. Journal of the American Statistical Association. 1998;93:228–237. [Google Scholar]
  30. Zou H, Yuan M. Composite quantile regression and the oracle model selection theory. Annals of Statistics. 2008;36:1108–1126. [Google Scholar]

RESOURCES