HIGH DIMENSIONAL COVARIANCE MATRIX ESTIMATION IN APPROXIMATE FACTOR MODELS

Jianqing Fan; Yuan Liao; Martina Mincheva

doi:10.1214/11-AOS944

. Author manuscript; available in PMC: 2012 May 30.

Published in final edited form as: Ann Stat. 2011 Jan 1;39(6):3320–3356. doi: 10.1214/11-AOS944

HIGH DIMENSIONAL COVARIANCE MATRIX ESTIMATION IN APPROXIMATE FACTOR MODELS

Jianqing Fan ¹, Yuan Liao ¹, Martina Mincheva ¹

PMCID: PMC3363011 NIHMSID: NIHMS367753 PMID: 22661790

Abstract

The variance covariance matrix plays a central role in the inferential theories of high dimensional factor models in finance and economics. Popular regularization methods of directly exploiting sparsity are not directly applicable to many financial problems. Classical methods of estimating the covariance matrices are based on the strict factor models, assuming independent idiosyncratic components. This assumption, however, is restrictive in practical applications. By assuming sparse error covariance matrix, we allow the presence of the cross-sectional correlation even after taking out common factors, and it enables us to combine the merits of both methods. We estimate the sparse covariance using the adaptive thresholding technique as in Cai and Liu (2011), taking into account the fact that direct observations of the idiosyncratic components are unavailable. The impact of high dimensionality on the covariance matrix estimation based on the factor structure is then studied.

Keywords and phrases: sparse estimation, thresholding, cross-sectional correlation, common factors, idiosyncratic, seemingly unrelated regression

1. Introduction

We consider a factor model defined as follows:

y_{it} = b_{i}^{'} f_{t} + u_{it},

(1.1)

where y_it is the observed datum for the ith (i = 1, …, p) asset at time t = 1, …, T; b_i is a K × 1 vector of factor loadings; f_t is a K × 1 vector of common factors, and u_it is the idiosyncratic error component of y_it. Classical factor analysis assumes that both p and K are fixed, while T is allowed to grow. However, in the recent decades, both economic and financial applications have encountered very large data sets which contain high dimensional variables. For example, the World Bank has data for about two-hundred countries over forty years; in portfolio allocation, the number of stocks can be in thousands and be larger or of the same order of the sample size. In modeling housing prices in each zip code, the number of regions can be of order thousands, yet the sample size can be 240 months or twenty years. The covariance matrix of order several thousands is critical for understanding the co-movement of housing prices indices over these zip codes.

Inferential theory of factor analysis relies on estimating Σ_u, the variance-covariance matrix of the error term, and Σ, the variance-covariance matrix of y_t = (y_1t, …, y_pt)′. In the literature, Σ = cov(y_t) was traditionally estimated by the sample covariance matrix of y_t:

Σ_{sam} = \frac{1}{T - 1} \sum_{t = 1}^{T} (y_{t} - \bar{y}) (y_{t} - \bar{y})',

which was always assumed to be pointwise root-T consistent. However, the sample covariance matrix is an inappropriate estimator in high dimensional settings. For example, when p is larger than T, Σ_sam becomes singular while Σ is always strictly positive definite. Even if p < T, Fan, Fan and Lv (2008) showed that this estimator has a very slow convergence rate under the Frobenius norm. Realizing the limitation of the sample covariance estimator in high dimensional factor models, Fan, Fan and Lv (2008) considered more refined estimation of Σ, by incorporating the common factor structure. One of the key assumptions they made was the cross-sectional independence among the idiosyncratic components, which results in a diagonal matrix $Σ_{u} = E u_{t} u_{t}^{'}$ . The cross-sectional independence, however, is restrictive in many applications, as it rules out the approximate factor structure as in Chamberlain and Rothschild (1983). In this paper, we relax this assumption, and investigate the impact of the cross-sectional correlations of idiosyncratic noises on the estimation of Σ and Σ_u, when both p and T are allowed to diverge. We show that the estimated covariance matrices are still invertible with probability approaching one, even if p > T. In particular, when estimating Σ⁻¹ and $Σ_{u}^{- 1}$ , we allow p to increase much faster than T, say, p = O (exp(T^α)), for some α ∈ (0, 1).

Sparsity is one of the commonly used assumptions in the estimation of high dimensional covariance matrices, which assumes that many entries of the off diagonal elements are zero, and the number of nonzero off-diagonal entries is restricted to grow slowly. Imposing the sparsity assumption directly on the covariance of y_t, however, is inappropriate for many applications of finance and economics. In this paper we use the factor model and assume that Σ_u is sparse, and estimate both Σ_u and $Σ_{u}^{- 1}$ using the thresholding method (Bickel and Levina (2008a), Cai and Liu (2011)) based on the estimated residuals in the factor model. It is assumed that the factors f_t are observable, as in Fama and French (1992), Fan, Fan and Lv (2008), and many other empirical applications. We derive the convergence rates of both estimated Σ and its inverse respectively under various norms which are to be defined later. In addition, we achieve better convergence rates than those in Fan, Fan and Lv (2008).

Various approaches have been proposed to estimate a large covariance matrix: Bickel and Levina (2008a, 2008b) constructed the estimators based on regularization and thresholding respectively. Rothman, Levina and Zhou (2009) considered thresholding the sample covariance matrix with more general thresholding functions. Lam and Fan (2009) proposed penalized quasi-likelihood method to achieve both the consistency and sparsistency of the estimation. More recently, Cai and Zhou (2010) derived the minimax rate for sparse matrix estimation, and showed that the thresholding estimator attains this optimal rate under the operator norm. Cai and Liu (2011) proposed a thresholding procedure which is adaptive to the variability of individual entries, and unveiled its improved rate of convergence.

The rest of the paper is organized as follows. Section 2 provides the asymptotic theory for estimating the error covariance matrix and its inverse. Section 3 considers estimating the covariance matrix of y_t. Section 4 extends the results to the seemingly unrelated regression model, a set of linear equations with correlated error terms in which the covariates are different across equations. Section 5 reports the simulation results. Finally, Section 6 concludes with discussions. All proofs are given in the appendix. Throughout the paper, we use λ_min(A) and λ_max(A) to denote the minimum and maximum eigenvalues of a matrix A. We also denote by ‖A‖_F, ‖A‖ and ‖A‖_∞ the Frobenius norm, operator norm and elementwise norm of a matrix A respectively, defined respectively as ‖A‖_F = tr^1/2(A′A), $‖ A ‖ = λ_{max}^{1 / 2} (A' A)$ , and ‖A‖_∞ = max_i,j |A_ij|. Note that, when A is a vector, both ‖A‖ and ‖A‖_F are equal to the Euclidean norm.

2. Estimation of Error Covariance Matrix

2.1. Adaptive thresholding

Consider the following approximate factor model, in which the cross-sectional correlation among the idiosyncratic error components is allowed:

y_{it} = b_{i}^{'} f_{t} + u_{it},

(2.1)

where i = 1, ‥, , p and t = 1, …, T; b_i is a K × 1 vector of factor loadings; f_t is a K × 1 vector of observable common factors, uncorrelated with u_it. Write

B = (b_{1}, \dots, b_{p})', y_{t} = (y_{1 t}, \dots, y_{pt})', u_{t} = (u_{1 t}, \dots, u_{pt})',

then model (2.1) can be written in a more compact form:

y_{t} = {Bf}_{t} + u_{t},

(2.2)

with E(u_t|f_t) = 0.

In practical applications, p can be thought of as the number of assets or stocks, or number of regions in spatial and temporal problems such as home price indices or sales of drugs, and in practice can be of the same order as, or even larger than T. For example, an asset pricing model may contain hundreds of assets while the sample size on daily returns is less than several hundreds. In the estimation of the optimal portfolio allocation, it was observed by Fan, Fan and Lv (2008) that the effect of large p on the convergence rate can be quite severe. In contrast, the number of common factors, K, can be much smaller. For example, the rank theory of consumer demand systems implies no more than three factors (e.g., Gorman (1981) and Lewbel (1991)).

The error covariance matrix

Σ_{u} = cov (u_{t}),

itself is of interest for the inferential theory of factor models. For example, the asymptotic covariance of the least square estimator of B depends on $Σ_{u}^{- 1}$ , and in simulating home price indices over a certain time horizon for mortgage based securities, a good estimate of Σ_u is needed. When p is close to or larger than T, estimating Σ_u is very challenging. Therefore, following the literature of high dimensional covariance matrix estimation, we assume it is sparse, i.e., many of its off-diagonal entries are zeros. Specifically, let Σ_u = (σ_ij)_p×p. Define

m_{T} = max_{i \leq p} \sum_{j \leq p} I (σ_{ij} \neq 0) .

(2.3)

The sparsity assumption puts an upper bound restriction on m_T. Specifically, we assume:

m_{T}^{2} = o (\frac{T}{K^{2} log p}) .

(2.4)

In this formulation, we even allow the number of factors K to be large, possibly growing with T.

A more general treatment (e.g., Bickel and Levina (2008a) and Cai and Liu (2011)) is to assume that the l_q norm of the row vectors of Σ_u are uniformly bounded across rows by a slowly growing sequence, for some q ∈ [0, 1). In contrast, the assumption we make in this paper, i.e., q = 0, has clearer economic interpretation. For example, the firm returns can be modeled by the factor model, where u_it represents a firm’s individual shock at time t. Driven by the industry-specific components, these shocks are correlated among the firms in the same industry, but can be assumed to be uncorrelated across industries, since the industry-specific components are not pervasive for the whole economy (Connor and Korajczyk (1993)).

We estimate Σ_u using the thresholding technique introduced and studied by Bickel and Levina (2008a), Rothman, Levina and Zhu (2009), Cai and Liu (2011), etc, which is summarized as follows: Suppose we observe data (X₁, …, X_T) of a p × 1 vector X, which follows a multivariate Gaussian distribution N (0, Σ_X). The sample covariance matrix of X is thus given by:

S_{X} = \frac{1}{T} \sum_{i = 1}^{T} (X_{i} - \bar{X}) (X_{i} - \bar{X})' = {(s_{ij})}_{p \times p} .

Define the thresholding operator by 𝒯_t(M) = (M_ijI (|M_ij| ≥ t)) for any symmetric matrix M. Then 𝒯_t preserves the symmetry of M. Let ${\hat{Σ}}_{X}^{𝒯} = 𝒯_{ω_{T}} (S_{X}), where ω_{T} = O (\sqrt{log p / T})$ . Bickel and Levina (2008a) then showed that:

‖ {\hat{Σ}}_{X}^{𝒯} - Σ_{X} ‖ = O_{p} (ω_{T} m_{T}) .

In the factor models, however, we do not observe the error term directly. Hence when estimating the error covariance matrix of a factor model, we need to construct a sample covariance matrix based on the residuals û_it before thresholding. The residuals are obtained using the plug-in method, by estimating the factor loadings first. Let b̂_i be the ordinary least square (OLS) estimator of b_i, and

û_{it} = y_{it} - {\hat{b}}_{i}^{'} f_{t} .

Denote by û_t = (û_1t, …, û_pt)′. We then construct the residual covariance matrix as:

{\hat{Σ}}_{u} = \frac{1}{T} \sum_{t = 1}^{T} û_{t} û_{t}^{'} = ({\hat{σ}}_{ij}) .

Note that the thresholding value $ω_{T} = O (\sqrt{log p / T})$ in Bickel and Levina (2008a) is in fact obtained from the rate of convergence of max_ij |s_ij − Σ_X,ij |. This rate changes when s_ij is replaced with the residual û_ij, which will be slower if the number of common factors K increases with T. Therefore, the thresholding value ω_T used in this paper is adjusted to account for the effect of the estimation of the residuals.

2.2. Asymptotic properties of the thresholding estimator

Bickel and Levina (2008a) used a universal constant as the thresholding value. As pointed out by Rothman, Levina and Zhu (2009) and Cai and Liu (2011), when the variances of the entries of the sample covariance matrix vary over a wide range, it is more desirable to use thresholds that capture the variability of individual estimation. For this purpose, in this paper, we apply the adaptive thresholding estimator (Cai and Liu (2011)) to estimate the error covariance matrix, which is given by

\begin{matrix} {\hat{Σ}}_{u}^{𝒯} = ({\hat{σ}}_{ij}^{𝒯}), {\hat{σ}}_{ij}^{𝒯} = {\hat{σ}}_{ij} I (| {\hat{σ}}_{ij} | \geq \sqrt{{\hat{θ}}_{ij}} ω_{T}) \\ {\hat{θ}}_{ij} = \frac{1}{T} \sum_{t = 1}^{T} {(û_{it} û_{jt} - {\hat{σ}}_{ij})}^{2}, \end{matrix}

(2.5)

for some ω_T to be specified later.

We impose the following assumptions:

Assumption 2.1. (i) {u_t}_t≥1 is stationary and ergodic such that each u_t has zero mean vector and covariance matrix Σ_u. In addition, the strong mixing condition in Assumption 3.2 holds.

(ii) There exist constants c₁, c₂ > 0 such that c₁ < λ_min(Σ_u) ≤ λ_max(Σ_u) < c₂, and c₁ < var(u_itu_jt) < c₂ for all i ≤ p, j ≤ p.

(iii) There exist r₁ > 0 and b₁ > 0, such that for any s > 0 and i ≤ p,

P (| u_{it} | > s) \leq exp (- {(s / b_{1})}^{r_{1}}) .

(2.6)

Condition (i) allows the idiosyncratic components to be weakly dependent. We will formally present the strong mixing condition in the next section. In order for the main results in this section to hold, it suffices to impose the strong mixing condition marginally on u_t only. Roughly speaking, we require the mixing coefficient

α (T) = sup_{A \in ℱ_{- \infty}^{0}, B \in ℱ_{T}^{\infty}} | P (A) P (B) - P (A \cap B) |

to decrease exponentially fast as T → ∞, where ( $ℱ_{- \infty}^{0}, ℱ_{T}^{\infty}$ ) are the σ-algebras generated by ${u_{t}}_{t = - \infty}^{0} and {u_{t}}_{t = T}^{\infty}$ respectively.

Condition (ii) requires the nonsingularity of Σ_u. Note that Cai and Liu (2011) allowed max_j σ_jj to diverse when direct observations are available. Condition (ii), however, requires that σ_jj should be uniformly bounded. In factor models, a uniform upper bound on the variance of u_it is needed when we estimate the covariance matrix of y_t later. This assumption is satisfied by most of the applications of factor models. Condition (iii) requires the distributions of (u_1t, …, u_pt) to have exponential-type tails, which allows us to apply the large deviation theory to $\frac{1}{T} \sum_{t = 1}^{T} u_{it} u_{jt} - σ_{ij}$ .

Assumption 2.2. There exist positive sequences κ₁(p, T) = o(1), κ₂(p, T) = o(1) and a_T = o(1), and a constant M > 0, such that for all C > M,

\begin{matrix} P (max_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} | u_{it} - û_{it} |^{2} > {Ca}_{T}^{2}) \leq O (κ_{1} (p, T)), \\ P (max_{i \leq p, t \leq T} | u_{it} - û_{it} | > C) \leq O (κ_{2} (p, T)) . \end{matrix}

This assumption allows us to apply thresholding to the estimated error covariance matrix when direct observations are not available, without introducing too much extra estimation error. Note that it permits a general case when the original “data” is contaminated, including any type of estimate of the data when direct observations are not available, as well as the case when data is subject to measurement of errors. We will show in the next section that in a linear factor model when {u_it}_i≤p,t≤T are estimated using the OLS estimator, the rate of convergence $a_{T}^{2} = (K^{2} log p) / T$ .

The following theorem establishes the asymptotic properties of the thresholding estimator ${\hat{Σ}}_{u}^{𝒯}$ , based on observations with estimation errors. Let $γ^{- 1} = 3 r_{1}^{- 1} + r_{2}^{- 1}$ , where r₁ and r₂ are defined in Assumptions 2.1, 3.2 respectively.

Theorem 2.1. Suppose γ < 1 and (log p)^6/γ−1 = o(T). Then under Assumptions 2.1 and 2.2, there exist C₁ > 0 and C₂ > 0 such that for ${\hat{Σ}}_{u}^{𝒯}$ defined in (2.5) with

ω_{T} = C_{1} (\sqrt{\frac{log p}{T}} + a_{T}),

we have,

P (‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ \leq C_{2} ω_{T} m_{T}) \geq 1 - O (\frac{1}{p^{2}} + κ_{1} (p, T) + κ_{2} (p, T)) .

(2.7)

In addition, if ω_Tm_T = o(1), then with probability at least $1 - O (\frac{1}{p^{2}} + κ_{1} (p, T) + κ_{2} (p, T))$ ,

λ_{min} ({\hat{Σ}}_{u}^{𝒯}) \geq 0.5 λ_{min} (Σ_{u}),

and

‖ {({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1} ‖ \leq C_{2} ω_{T} m_{T} .

Note that we derive the result (2.7) without assuming the sparsity on Σ_u, i.e., no restriction is imposed on m_T. When ω_Tm_T ≠ o(1), (2.7) still holds, but $‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖$ does not converge to zero in probability. On the other hand, the condition ω_Tm_T = o(1) is required to preserve the nonsingularity of ${\hat{Σ}}_{u}^{𝒯}$ asymptotically and to consistently estimate $Σ_{u}^{- 1}$ .

The rate of convergence also depends on the averaged estimation error of the residual terms. We will see in the next section that when the number of common factors K increases slowly, the convergence rate in Theorem 2.1 is close to the minimax optimal rate as in Cai and Zhou (2010).

3. Estimation of Covariance Matrix Using Factors

We now investigate the estimation of the covariance matrix Σ in the approximate factor model:

y_{t} = {Bf}_{t} + u_{t},

where Σ = cov(y_t). This covariance matrix is particularly of interest in many applications of factor models as well as corresponding inferential theories. When estimating a large dimensional covariance matrix, sparsity and banding are two commonly used assumptions for regularization (e.g., Bickel and Levina (2008a, 2008b)). In most of the applications in finance and economics, however, these two assumptions are inappropriate for Σ. For instance, the US housing prices in the county level are generally associated with a few national indices, and there is no natural ordering among the counties. Hence neither the sparsity nor banding is realistic for such a problem. On the other hand, it is natural to assume Σ_u sparse, after controling the common factors. Therefore, our approach combines the merits of both the sparsity and factor structures.

Note that

Σ = B cov (f_{t}) B' + Σ_{u} .

By the Sherman-Morrison-Woodbury formula,

Σ^{- 1} = Σ_{u}^{- 1} - Σ_{u}^{- 1} B {[cov {(f_{t})}^{- 1} + B' Σ_{u}^{- 1} B]}^{- 1} B' Σ_{u}^{- 1} .

When the factors are observable, one can estimate B by the least squares method: B̂ = (b̂₁, …, b̂_p)′, where,

{\hat{b}}_{i} = arg min_{b_{i}} \frac{1}{Tp} \sum_{t = 1}^{T} \sum_{i = 1}^{p} {(y_{it} - b_{i}^{'} f_{t})}^{2} .

The covariance matrix cov(f_t) can be estimated by the sample covariance matrix

\hat{cov} (f_{t}) = T^{- 1} XX' - T^{- 2} X 11' X',

where X = (f₁, …, f_T), and 1 is a T-dimensional column vector of ones. Therefore, by employing the thresholding estimator ${\hat{Σ}}_{u}^{𝒯}$ in (2.5), we obtain substitution estimators

{\hat{Σ}}^{𝒯} = \hat{B} \hat{cov} (f_{t}) \hat{B}' + {\hat{Σ}}_{u}^{𝒯},

(3.1)

and

{({\hat{Σ}}^{𝒯})}^{- 1} = {({\hat{Σ}}_{u}^{𝒯})}^{- 1} - {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B} {[\hat{cov} {(f_{t})}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}]}^{- 1} \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} .

(3.2)

In practice, one may apply a common thresholding λ to the correlation matrix of Σ̂_u, and then use the substitution estimator similar to (3.1). When λ = 0 (no thresholding), the resulting estimator is the sample covariance, whereas when λ = 1 (all off-diagonals are thresholded), the resulting estimator is an estimator based on the strict factor model (Fan, Fan and Lv (2008)). Thus we have created a path (indexed by λ) which connects the nonparametric estimate of covariance matrix to the parametric estimate.

The following assumptions are made.

Assumption 3.1. (i) {f_t}_t≥1 is stationary and ergodic.

(ii) {u_t}_t≥1 and {f_t}_t≥1 are independent.

In addition to the conditions above, we introduce the strong mixing conditions to conduct asymptotic analysis of the least square estimates. Let $ℱ_{- \infty}^{0}, and ℱ_{T}^{\infty}$ denote the σ-algebras generated by {(f_t, u_t) : −∞ ≤ t ≤ 0} and {(f_t, u_t) : T ≤ t ≤ ∞} respectively. In addition, define the mixing coefficient

α (T) = sup_{A \in ℱ_{- \infty}^{0}, B \in ℱ_{T}^{\infty}} | P (A) P (B) - P (AB) | .

The following strong mixing assumption enables us to apply the Bernstein’s inequality in the technical proofs.

Assumption 3.2. There exist positive constants r₂ and C such that for all t ∈ ℤ⁺,

α (t) \leq exp (- {Ct}^{r_{2}}) .

In addition, we impose the following regularity conditions.

Assumption 3.3. (i) There exists a constant M > 0 such that for all i, j and t, ${Ey}_{it}^{2} < M, {E f}_{it}^{2} < M$ , and |b_ij| < M.

(ii) There exists a constant r₃ > 0 with $3 r_{3}^{- 1} + r_{2}^{- 1} > 1$ , and b₂ > 0 such that for any s > 0 and i ≤ K,

P (| f_{it} | > s) \leq exp (- {(s / b_{2})}^{r_{3}}) .

(3.3)

Condition (ii) allows us to apply the Bernstein type inequality for the weakly dependent data.

Assumption 3.4. There exists a constant C > 0 such that λ_min(cov(f_t)) > C.

Assumptions 3.4 and 2.1 ensure that both λ_min(cov(f_t)) and λ_min(Σ) are bounded away from zero, which is needed to derive the convergence rate of ‖(Σ̂^𝒯)⁻¹ − Σ⁻¹‖ below.

The following lemma verifies Assumption 2.2, which derives the rate of convergence of the OLS estimator as well as the estimated residuals.

Let $γ_{2}^{- 1} = 1.5 r_{1}^{- 1} + 1.5 r_{3}^{- 1} + r_{2}^{- 1}$ .

Lemma 3.1. Suppose K = o(p), K⁴(log p)² = o(T) and (log p)^2/γ₂−1 = o(T). Then under the assumptions of Theorem 2.1 and Assumptions 3.1–3.4, there exists C > 0, such that

$P (max_{i \leq p} ‖ {\hat{b}}_{i} - b_{i} ‖ > C \sqrt{\frac{K log p}{T}}) = O (\frac{1}{p^{2}} + \frac{1}{T^{2}}),$
$P (max_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} | u_{it} - û_{it} |^{2} > \frac{{CK}^{2} log p}{T}) = O (\frac{1}{p^{2}} + \frac{1}{T^{2}}),$
$P (max_{i \leq p, t \leq T} | u_{it} - û_{it} | > CK {(log T)}^{1 / r_{3}} \sqrt{\frac{log p}{T}}) = O (\frac{1}{p^{2}} + \frac{1}{T^{2}}) .$

By Lemma 3.1 and Assumption 2.2, $a_{T} = K \sqrt{(log p) / T}$ , and κ₁(p, T) = κ₂(p, T) = p⁻² + T⁻². Therefore in the linear approximate factor model, the thresholding parameter ω_T defined in Theorem 2.1 is simplified to: for some positive constant $C_{1}^{'}$ ,

ω_{T} = C_{1}^{'} K \sqrt{\frac{log p}{T}} .

(3.4)

Now we can apply Theorem 2.1 to obtain the following theorem:

Theorem 3.1. Under the Assumptions of Lemma 3.1, there exist $C_{1}^{'} > 0 and C_{2}^{'} > 0$ such that the adaptive thresholding estimator defined in (2.5) with $ω_{T}^{2} = C_{1}^{'} \frac{K^{2} log p}{T}$ satisfies

$P (‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ \leq C_{2}^{'} m_{T} K \sqrt{\frac{log p}{T}}) = 1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}}) .$
If $m_{T} K \sqrt{\frac{log p}{T}} = o (1)$ , then with probability at least $1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}})$ ,
$λ_{min} ({\hat{Σ}}_{u}^{𝒯}) \geq 0.5 λ_{min} (Σ_{u}),$
and
$‖ {({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1} ‖ \leq C_{2}^{'} m_{T} K \sqrt{\frac{log p}{T}} .$

Remark 3.1. We briefly comment on the terms in the convergence rate above.

The term K appears as an effect of using the estimated residuals to construct the thresholding covariance estimator, which is typically small compared to p and T in many applications. For instance, the famous Fama-French three-factor model shows that K = 3 factors are adequate for the US equity market. In an empirical study on asset returns, Bai and Ng (2002) used the monthly data which contains the returns of 4883 stocks for sixty months. For their data set, T = 60, p = 4883. Bai and Ng (2002) determined K = 2 common factors.
As in Bickel and Levina (2008a) and Cai and Liu (2011), m_T, the maximum number of nonzero components across the rows of Σ_u, also plays a role in the convergence rate. Note that when K is bounded, the convergence rate reduces to $O_{p} (m_{T} \sqrt{(log p) / T})$ , the same as the minimax rate derived by Cai and Zhou (2010).

One of our main objectives is to estimate Σ, which is the p × p dimensional covarinace matrix of y_t, assumed to be time invariant. We can achieve a better accuracy in estimating both Σ and Σ⁻¹ by incorporating the factor structure than using the sample covariance matrix, as shown by Fan et al (2008) in the strict factor model case. When the cross sectional correlations among the idiosyncratic components (u_1t, …, u_pt) are in presence, we can still take advantage of the factor structure. This is particularly essential when direct sparsity assumption on Σ is inapproriate.

Assumption 3.5. ‖p⁻¹B′B−Ω‖ = o(1) for some K × K symmetric positive definite matrix Ω such that λ_min(Ω) is bounded away from zero.

Assumption 3.5 requires that the factors should be pervasive, i.e., impact every individual time series (Harding (2009)). It was imposed by Fan et al.(2008) only when they tried to establish the asymptotic normality of the covariance estimator. However, it turns out to be also helpful to obtain a good upper bound of ‖(Σ̂^𝒯)⁻¹ − Σ⁻¹‖, as it ensures that λ_max((B′Σ⁻¹B)⁻¹) = O(p⁻¹).

Fan et al. (2008) obtained an upper bound of ‖Σ̂^𝒯 − Σ‖_F under the Frobenius norm when Σ_u is diagonal, i.e., there was no cross-sectional correlation among the idiosyncratic errors. In order for their upper bound to decrease to zero, p² < T is required. Even with this restrictive assumption, they showed that the convergence rate is the same as the usual sample covariance matrix of y_t, though the latter does not take the factor structure into account. Alternatively, they considered the entropy loss norm, proposed by James and Stein (1961):

{‖ {\hat{Σ}}^{𝒯} - Σ ‖}_{Σ} = {(p^{- 1} tr [{({\hat{Σ}}^{𝒯} Σ^{- 1} - I)}^{2}])}^{1 / 2} = p^{- 1 / 2} {‖ Σ^{- 1 / 2} ({\hat{Σ}}^{𝒯} - Σ) Σ^{- 1 / 2} ‖}_{F} .

Here the factor p^−1/2 is used for normalization, such that ‖Σ‖_Σ = 1. Under this norm, Fan et al. (2008) showed that the substitution estimator has a better convergence rate than the usual sample covariance matrix. Note that the normalization factor p^−1/2 in the definition results in an averaged estimation error, which also cancels out the diverging dimensionality introduced by p. In addition, for any two p × p matrices A₁ and A₂:

\begin{matrix} {‖ A_{1} - A_{2} ‖}_{Σ} & = p^{- 1 / 2} {‖ Σ^{- 1 / 2} (A_{1} - A_{2}) Σ^{- 1 / 2} ‖}_{F} . \\ \leq ‖ Σ^{- 1 / 2} (A_{1} - A_{2}) Σ^{- 1 / 2} ‖ \\ \leq ‖ A_{1} - A_{2} ‖ \cdot λ_{max} (Σ^{- 1}) . \end{matrix}

Combining with the estimated low-rank matrix Bcov(f_t)B′, Theorem 3.1 implies the main theorem in this section:

Theorem 3.2. Suppose log T = o(p). Under the assumptions of Theorem 3.1 and Assumption 3.5, we have

$\begin{matrix} P ({‖ {\hat{Σ}}^{𝒯} - Σ ‖}_{Σ}^{2} \leq \frac{{CpK}^{2} {(log p)}^{2}}{T^{2}} + \frac{{Cm}_{T}^{2} K^{2} log p}{T}) = 1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}}), \\ P ({‖ {\hat{Σ}}^{𝒯} - Σ ‖}_{\infty}^{2} \leq \frac{{CK}^{2} log p + {CK}^{4} \log T}{T}) = 1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}}) . \end{matrix}$
If $m_{T} K \sqrt{\frac{log p}{T}} = o (1)$ , with probability at least $1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}})$ ,
$λ_{min} ({\hat{Σ}}^{𝒯}) \geq 0.5 λ_{min} (Σ_{u}),$
and
$‖ {({\hat{Σ}}^{𝒯})}^{- 1} - Σ^{- 1} ‖ \leq {Cm}_{T} K \sqrt{\frac{log p}{T}}$

Note that we have derived a better convergence rate of (Σ̂^𝒯)⁻¹ than that in Fan et al.(2008). When the operator norm is considered, p is allowed to grow exponentially fast in T in order for (Σ̂^𝒯)⁻¹ to be consistent.

We have also derived the maximum elementwise estimation ‖Σ̂^𝒯 − Σ‖_∞. This quantity appears in risk assessment as in Fan, Zhang and Yu (2008). For any portfolio with allocation vector w, the true portfolio variance and the estimated one are given by w′Σw and w′Σ̂^𝒯 w respectively. The estimation error is bounded by

| w' {\hat{Σ}}^{𝒯} w - w' Σ w | \leq {‖ {\hat{Σ}}^{𝒯} - Σ ‖}_{\infty} {‖ w ‖}_{1}^{2},

where ‖w‖₁, the l₁ norm of w, is the gross exposure of the portfolio.

4. Extension: Seemingly Unrelated Regression

A seemingly unrelated regression model (Kmenta and Gilbert (1970)) is a set of linear equations in which the disturbances are correlated across equations. Specifically, we have

y_{it} = b_{i}^{'} f_{it} + u_{it}, i \leq p, t \leq T,

(4.1)

where b_i and f_it are both K_i × 1 vectors. The p linear equations (4.1) are related because their error terms u_it are correlated, i.e., the covariance matrix

Σ_{u} = {({Eu}_{it} u_{jt})}_{p \times p}

is not diagonal.

Model (4.1) allows each variable y_it to have its own factors. This is important for many applications. In financial applications, the returns of individual stock depend on common market factors and sector-specific factors. In housing price index modeling, housing price appreciations depend on both national factors and local economy. When f_it = f_t for each i ≤ p, model (4.1) reduces to the approximate factor model (1.1) with common factors f_t.

Under mild conditions, running OLS on each equation produces unbiased and consistent estimator of b_i separately. However, since OLS does not take into account the cross sectional correlation among the noises, it is not efficient. Instead, statisticians obtain the best linear unbiased estimator (BLUE) via generalized least square (GLS). Write

\begin{matrix} y_{i} = (y_{i 1}, \dots, y_{iT})', T \times 1, X_{i} = (f_{i 1}, \dots, f_{iT})', T \times K_{i}, i \leq p, \\ y = (\begin{matrix} y_{1} \\ ⋮ \\ y_{p} \end{matrix}), X = (\begin{matrix} X_{1} & 0 & 0 \\ 0 & ⋱ & 0 \\ 0 & 0 & X_{p} \end{matrix}), B = (\begin{matrix} b_{1} \\ ⋮ \\ b_{p} \end{matrix}) . \end{matrix}

The GLS estimator of B is given by Zellner (1962):

{\hat{B}}_{GLS} = {[X' {({\hat{Σ}}_{u}^{- 1} \otimes I_{T})}^{- 1} X]}^{- 1} [X' {({\hat{Σ}}_{u}^{- 1} \otimes I_{T})}^{- 1} y],

(4.2)

where I_T denotes a T × T identity matrix, ⊗ represents the Kronecker product operation, and Σ̂_u is a consistent estimator of Σ_u.

In classical seemingly unrelated regression in which p does not grow with T, Σ_u is estimated by a two-stage procedure: (Kmenta and Gilbert (1970)): On the first stage, estimate B via OLS, and obtain residuals

û_{it} = y_{it} - {\hat{b}}_{i}^{'} f_{it} .

(4.3)

On the second stage, estimate Σ_u by

{\hat{Σ}}_{u} = ({\hat{σ}}_{ij}) = {(\frac{1}{T} \sum_{t = 1}^{T} û_{it} û_{jt})}_{p \times p} .

(4.4)

In high dimensional seemingly unrelated regression in which p > T, however, Σ̂_u is not invertible, and hence the GLS estimator (4.2) is infeasible.

By the sparsity assumption of Σ_u, we can deal with this singularity problem by using the adaptive thresholding estimator, and produce a consistent nonsingular estimator of Σ_u:

{\hat{Σ}}_{u}^{𝒯} = ({\hat{σ}}_{ij} I (| {\hat{σ}}_{ij} | > \sqrt{{\hat{θ}}_{ij}} ω_{T})), {\hat{θ}}_{ij} = \frac{1}{T} \sum_{t = 1}^{T} {(û_{it} û_{jt} - {\hat{σ}}_{ij})}^{2} .

(4.5)

To pursue this goal, we impose the following assumptions:

Assumption 4.1. For each i ≤ p,

{f_it}_t≥1 is stationary and ergodic.
{u_t}_t≥1 and {f_it}_t≥1 are independent.

Assumption 4.2. There exists positive constants C and r₂ such that for each i ≤ p, the strong mixing condition

α (t) \leq exp (- {Ct}^{r_{2}})

is satisfied by (f_it, u_t):

Assumption 4.3. There exist constants M and C > 0 such that for all i ≤ p, j ≤ K_i, t ≤ T

${Ey}_{it}^{2} < M$ , |b_ij| < M, and ${E f}_{it, j}^{2} < M$ .
min_i≤p λ_min(cov(f_it)) > C.

Assumption 4.4. There exists a constant r₄ > 0 with $3 r_{4}^{- 1} + r_{2}^{- 1} > 1$ , and b₃ > 0 such that for any s > 0 and i, j,

P (| f_{it, j} | > s) \leq exp (- {(s / b_{3})}^{r_{4}}) .

These assumptions are similar to those made in Section 3, except that here they are imposed on the sector-specific factors. The main theorem in this section is a direct application of Theorem 2.1, which shows that the adaptive thresholding produces a consistent nonsingular estimator of Σ̂_u.

Theorem 4.1. Let K = max_i≤p K_i and $γ_{3}^{- 1} = 1.5 r_{1}^{- 1} + 1.5 r_{4}^{- 1} + r_{2}^{- 1}$ ; suppose K = o(p), K⁴(log p)² = o(T) and (log p)^2/γ₃−1 = o(T). Under Assumptions 2.1, 4.1–4.4, there exist constants C₁ > 0 and C₂ > 0 such that the adaptive thresholding estimator defined in (4.5) with $ω_{T}^{2} = C_{1} \frac{K^{2} log p}{T}$ satisfies

$P (‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ \leq C_{2} m_{T} K \sqrt{\frac{log p}{T}}) = 1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}}) .$
If $m_{T} K \sqrt{\frac{log p}{T}} = o (1)$ , then with probability at least $1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}})$ ,
$λ_{min} ({\hat{Σ}}_{u}^{𝒯}) \geq 0.5 λ_{min} (Σ_{u}),$
and
$‖ {({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1} ‖ \leq C_{2} m_{T} K \sqrt{\frac{log p}{T}} .$

Therefore, in the case when p > T, Theorem 4.1 enables us to efficiently estimate B via feasible GLS:

{\hat{B}}_{GLS}^{𝒯} = {[X' {({({\hat{Σ}}_{u}^{𝒯})}^{- 1} \otimes I_{T})}^{- 1} X]}^{- 1} {[X' ({({\hat{Σ}}_{u}^{𝒯})}^{- 1} \otimes I_{T})}^{- 1} y] .

5. Monte Carlo Experiments

In this section, we use simulation to demonstrate the rates of convergence of the estimators Σ̂^𝒯 and (Σ̂^𝒯)⁻¹ that we have obtained so far. The simulation model is a modified version of the Fama-French three-factor model described in Fan, Fan, Lv (2008). We fix the number of factors, K = 3 and the length of time, T = 500, and let the dimensionality p gradually increase.

The Fama-French three-factor model (Fama and French (1992)) is given by

y_{it} = b_{i 1} f_{1 t} + b_{i 2} f_{2 t} + b_{i 3} f_{3 t} + u_{it},

which models the excess return (real rate of return minus risk-free rate) of the ith stock of a portfolio, y_it, with respect to 3 factors. The first factor is the excess return of the whole stock market, and the weighted excess return on all NASDAQ, AMEX and NYSE stocks is a commonly used proxy. It extends the capital assets pricing model (CAPM) by adding two new factors- SMB (“small minus big” cap) and HML (“high minus low” book/price). These two were added to the model after the observation that two types of stocks - small caps, and high book value to price ratio, tend to outperform the stock market as a whole.

We separate this section into three parts, calibration, simulation and results. Similar to Section 5 of Fan, Fan and Lv (2008), in the calibration part we want to calculate realistic multivariate distributions from which we can generate the factor loadings B, idiosyncratic noises ${u_{t}}_{t = 1}^{T}$ and the observable factors ${f_{t}}_{t = 1}^{T}$ . The data was obtained from the data library of Kenneth French’s website.

5.1. Calibration

To estimate the parameters in the Fama-French model, we will use the two-year daily data (ỹ_t, f̃_t) from Jan 1^st, 2009 to Dec 31^st, 2010 (T=500) of 30 industry portfolios.

Calculate the least squares estimator B̃ of ỹ_t = Bf̃_t + u_t, and take the rows of B̃, namely b̃₁ = (b₁₁, b₁₂, b₁₃), …, b̃₃₀ = (b_30,1, b_30,2, b_30,3), to calculate the sample mean vector μ_B and sample covariance matrix Σ_B. The results are depicted in Table 1. We then create a mutlivariate normal distribution N₃(μ_B, Σ_B), from which the factor loadings ${b_{i}}_{i = 1}^{p}$ are drawn.
For each fixed p, create the sparse matrix $Σ_{u} = D + ss' - diag {s_{1}^{2}, \dots, s_{p}^{2}}$ in the following way. Let û_t = ỹ_t − B̃f̃_t. For i = 1, …, 30, let σ̂_i denote the standard deviation of the residuals of the ith portfolio. We find min(σ̂_i) = 0.3533, max(σ̂_i) = 1.5222, and calculate the mean and the standard deviation of the σ̂_i’s, namely σ̄ = 0.6055 and σ_SD = 0.2621.

Let $D = diag {σ_{1}^{2}, \dots, σ_{p}^{2}}$ , where σ₁, …, σ_p are generated independently from the Gamma distribution G(α, β), with mean αβ and standard deviation α^1/2β. We match these values to σ̄ = 0.6055 and σ_SD = 0.2621, to get α = 5.6840 and β = 0.1503. Further, we create a loop that only accepts the value of σ_i if it is between min(σ̂_i) = 0.3533 and max(σ̂_i) = 1.5222.

Create s = (s₁, …, s_p)′ to be a sparse vector. We set each s_i ~ N(0, 1) with probability $\frac{0.2}{\sqrt{p} log p}$ , and s_i = 0 otherwise. This leads to an average of $\frac{0.2 \sqrt{p}}{log p}$ nonzero elements per each row of the error covariance matrix.

Create a loop that generates Σ_u multiple times until it is positive definite.
Assume the factors follow the vector autoregressive (VAR(1)) model f_t = μ + Φf_t−1 + ε_t for some 3 × 3 matrix Φ, where ε_t’s are i.i.d. N₃(0, Σ_ε). We estimate Φ, μ and Σ_ε from the data, and obtain cov(f_t). They are summarized in Table 2.

Table 1.

Mean and covariance matrix used to generate b

μ_B	Σ_B
1.0641	0.0475	0.0218	0.0488
0.1233	0.0218	0.0945	0.0215
−0.0119	0.0488	0.0215	0.1261

Open in a new tab

Table 2.

Parameters of f_t generating process

μ	cov(f_t)			Φ
0.1074	2.2540	0.2735	0.9197	−0.1149	0.0024	0.0776
0.0357	0.2735	0.3767	0.0430	0.0016	−0.0162	0.0387
0.0033	0.9197	0.0430	0.6822	−0.0399	0.0218	0.0351

Open in a new tab

5.2. Simulation

For each fixed p, we generate (b₁, …, b_p) independently from N₃(μ_B, Σ_B), and generate ${f_{t}}_{t = 1}^{T} and {u_{t}}_{t = 1}^{T}$ independently. We keep T = 500 fixed, and gradually increase p from 20 to 600 in multiples of 20 to illustrate the rates of convergence when the number of variables diverges with respect to the sample size.

Repeat the following steps N = 200 times for each fixed p:

Generate ${b_{i}}_{i = 1}^{p}$ independently from N₃(μ_B, Σ_B), and set B = (b₁, …, b_p)′.
Generate ${u_{t}}_{t = 1}^{T}$ independently from N_p(0, Σ_u).
Generate ${f_{t}}_{t = 1}^{T}$ independently from the VAR(1) model f_t = μ + Φf_t−1 + ε_t.
Calculate y_t = Bf_t + u_t for t = 1, …, T.
Set $ω_{T} = 0.10 K \sqrt{log p / T}$ to obtain the thresholding estimator (2.5) ${\hat{Σ}}_{u}^{𝒯}$ and the sample covariance matrices $\hat{cov} (f_{t}), {\hat{Σ}}_{y} = \frac{1}{T - 1} \sum_{t = 1}^{T} (y_{t} - \bar{y}) {(y_{t} - \bar{y})}^{T}$ .

We graph the convergence of Σ̂^𝒯 and Σ̂_y to Σ, the covariance matrix of y, under the entropy-loss norm ‖·‖_Σ and the elementwise norm ‖·‖_∞. We also graph the convergence of the inverses (Σ̂^𝒯)⁻¹ and ${\hat{Σ}}_{y}^{- 1}$ to Σ⁻¹ under the operator norm. Note that we graph that only for p from 20 to 300. Since T = 500, for p > 500 the sample covariance matrix is singular. Also, for p close to 500, Σ̂_y is nearly singular, which leads to abnormally large values of the operator norm. Lastly, we record the standard deviations of these norms.

5.3. Results

In Figures 1–3, the dashed curves correspond to Σ̂^𝒯 and the solid curves correspond to the sample covariance matrix Σ̂_y. Figure 1 and 2 present the averages and standard deviations of the estimation error of both of these matrices with respect to the Σ-norm and infinity norm, respectively. Figure 3 presents the averages and estimation errors of the inverses with respect to the operator norm. Based on the simulation results, we can make the following observations:

The standard deviations of the norms are negligible when compared to their corresponding averages.
Under the ‖·‖_Σ, our estimate of the covariance matrix of y, Σ̂^𝒯 performs much better than the sample covaraince matrix Σ̂_y. Note that, in the proof of Theorem 2 in Fan, Fan, Lv(2008), it was shown that:
${‖ {\hat{Σ}}_{y} - Σ ‖}_{Σ}^{2} = O_{p} (\frac{K^{3}}{Tp}) + O_{p} (\frac{p}{T}) + O_{p} (\frac{K^{3 / 2}}{T}) .$ (5.1)

For a small fixed value of K, such as K = 3, the dominating term in (5.1) is $O (\frac{p}{T})$ . From Theorem 4.1, and given that m_T = o(p^1/4), the dominating term in the convergence of ${‖ {\hat{Σ}}^{𝒯} - Σ ‖}_{Σ}^{2} is O_{p} (\frac{p}{T^{2}} + \frac{m_{T}^{2} log p}{T})$ . So, we would expect our estimator to perform better, and the simulation results are consistent with the theory.
Under the infinity norm, both estimators perform roughly the same. This is to be expected, given that the thresholding affects mainly the elements of the covariance matrix that are closest to 0, and the infinity norm depicts the magnitude of the largest elementwise absolute error.
Under the operator norm, the inverse of our estimator, (Σ̂^𝒯)⁻¹ also performs significantly better than the inverse of the sample covariance matrix.
Finally, when p > 500, the thresholding estimators ${\hat{Σ}}_{u}^{𝒯}$ and Σ̂^𝒯 are still nonsingular.

Fig 1 — Averages and standard deviations of ‖Σ̂^𝒯 − Σ‖_Σ (dashed curve) and ‖Σ̂_y − Σ‖_Σ (solid curve) over N = 200 iterations, as a function of the dimensionality p.

Fig 3 — Averages and standard deviations of ‖(Σ̂^𝒯)⁻¹ − Σ⁻¹‖ (dashed curve) and $‖ {\hat{Σ}}_{y}^{- 1} - Σ^{- 1} ‖$ (solid curve) over N = 200 iterations, as a function of the dimensionality p.

Fig 2 — Averages and standard deviations of ‖Σ̂^𝒯 − Σ‖_∞ (dashed curve) and ‖Σ̂_y − Σ‖_∞ (solid curve) over N = 200 iterations, as a function of the dimensionality p.

In conclusion, even after imposing less restrictive assumptions on the error covariance matrix, we still reach an estimator Σ̂^𝒯 that significantly outperforms the standard sample covariance matrix.

6. Conclusions and Discussions

We studied the rate of convergence of high dimensional covariance matrix of approximate factor models under various norms. By assuming sparse error covariance matrix, we allow for the presence of the cross-sectional correlation even after taking out common factors. Since direct observations of the noises are not available, we constructed the error sample covariance matrix first based on the estimation residuals, and then estimate the error covariance matrix using the adaptive thresholding method.We then constructed the covariance matrix of y_t using the factor model, assuming that the factors follow a stationary and ergodic process, but can be weakly-dependent. It was shown that after thresholding, the estimated covariance matrices are still invertible even if p > T, and the rate of convergence of (Σ̂^𝒯)⁻¹ and ${({\hat{Σ}}_{u}^{𝒯})}^{- 1}$ is of order $O_{p} ({Km}_{T} \sqrt{log p / T})$ , where K comes from the impact of estimating the unobservable noise terms. This demonstrates when estimating the inverse covariance matrix, p is allowed to be much larger than T.

In fact, the rate of convergence in Theorem 2.1 reflects the impact of unobservable idiosyncratic components on the thresholding method. Generally, whether it is the minimax rate when direct observations are not available but have to be estimated is an important question, which is left as a research direction in the future.

Moreover, this paper uses the hard-thresholding technique, which takes the form of σ̂_ij(σ_ij) = σ_ijI(|σ_ij| > θ_ij) for some pre-determined threshold θ_ij. Recently, Rothman et al (2009) and Cai and Liu (2011) studied a more general thresholding function of Antoniadis and Fan (2001), which admits the form σ̂_ij(θ_ij) = s(σ_ij), and also allows for soft-thresholding. It is easy to apply the more general thresholding here as well, and the rate of convergence of the resulting covariance matrix estimators should be straightforward to derive.

Finally, we considered the case when common factors are observable, as in Fama and French (1992). In some applications, the common factors are unobservable and need to be estimated (Bai (2003)). In that case, it is still possible to consistently estimate the covariance matrices using similar techniques as those in this paper. However, the impact of high dimensionality on the rate of convergence comes also from the estimation error of the unobservable factors. We plan to address this problem in a separate paper.

Acknowledgments

The research was partially supported by NIH Grant R01-GM072611, NSF Grant DMS-0704337, and NIH grant R01GM100474.

APPENDIX A: PROOFS FOR SECTION 2

A.1. Lemmas. The following lemmas are useful to be proved first, in which we consider the operator norm ‖A‖² = λ_max(A′A).

Lemma A.1. Let A be an m × m random matrix, B be an m × m deterministic matrix, and both A and B are semi-positive definite. If there exists a positive sequence ${c_{T}}_{T = 1}^{\infty}$ such that for all large enough T, λ_min(B) > c_T. Then

\begin{matrix} P (λ_{min} (A) \geq 0.5 c_{T}) \geq P (‖ A - B ‖ \leq 0.5 c_{T}), and \\ P (‖ A^{- 1} - B^{- 1} ‖ \leq \frac{2}{c_{T}^{2}} ‖ A - B ‖) \geq P (‖ A - B ‖ \leq 0.5 c_{T}) . \end{matrix}

Proof. For any v ∈ ℝ^m such that ‖v‖ = 1, under the event ‖A − B‖ ≤ 0.5c_T,

\begin{matrix} v' Av & = v' Bv - v' (B - A) v \geq λ_{min} (B) - ‖ A - B ‖ \\ \geq 0.5 c_{T} \end{matrix}

Hence λ_min(A) ≥ 0.5c_T.

In addition, still under the event ‖A − B‖ ≤ 0.5c_T,

\begin{matrix} ‖ A^{- 1} - B^{- 1} ‖ & = ‖ A^{- 1} (B - A) B^{- 1} ‖ \\ \leq λ_{min} {(A)}^{- 1} ‖ A - B ‖ λ_{min} {(B)}^{- 1} \\ = 2 c_{T}^{- 1} ‖ A - B ‖ . \end{matrix}

Q.E.D.

Lemma A.2. Suppose that the random variables Z₁, Z₂ both satisfy the exponential-type tail condition: There exist r₁, r₂ ∈ (0, 1) and b₁, b₂ > 0, such that ∀s > 0,

P (| Z_{i} | > s) \leq exp (1 - {(s / b_{i})}^{r_{i}}), i = 1, 2 .

Then for some r₃ and b₃ > 0, and any s > 0,

P (| Z_{1} Z_{2} | > s) \leq exp (1 - {(s / b_{3})}^{r_{3}}) .

(A.1)

Proof. We have, for any s > 0, $M = {({sb}_{2}^{r_{2} / r_{1}} / b_{1})}^{r_{1} / (r_{1} + r_{2})}$ , b = b₁b₂, and r = r₁r₂/(r₁ + r₂),

\begin{matrix} P (| Z_{1} Z_{2} | > s) & \leq P (M | Z_{1} | > s) + P (| Z_{2} | > M) \\ \leq exp (1 - {(s / b_{1} M)}^{r_{1}}) + exp (1 - {(M / b_{2})}^{r_{2}}) \\ = 2 exp (1 - {(s / b)}^{r}) . \end{matrix}

Pick up an r₃ ∈ (0, r), and b₃ > max{(r₃/r)^1/rb, (1 + log 2)^1/rb}, then it can be shown that F(s) = (s/b)^r − (s/b₃)^r₃ is increasing when s > b₃. Hence F(s) > F(b₃) > log 2 when s > b₃, which implies when s > b₃,

P (| Z_{1} Z_{2} | > s) \leq 2 exp (1 - {(s / b)}^{r}) \leq exp (1 - {(s / b_{3})}^{r_{3}}) .

When s ≤ b₃,

P (| Z_{1} Z_{2} | > s) \leq 1 \leq exp (1 - {(s / b_{3})}^{r_{3}}) .

Q.E.D.

Lemma A.3. Under the Assumptions of Theorem 2.1, there exists a constant C_r > 0 that does not depend on (p, T), such that when C > C_r,

$P (max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} u_{it} u_{jt} - σ_{ij} | > C \sqrt{\frac{log p}{T}}) = O (\frac{1}{p^{2}}),$
$P (max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} (û_{it} û_{jt} - u_{it} u_{jt} | > {Ca}_{T}) = O (\frac{1}{p^{2}} + κ_{1} (p, T)),$
$P (max_{i, j \leq p} | {\hat{σ}}_{ij} - σ_{ij} | > C (\sqrt{\frac{log p}{T}} + a_{T})) = O (\frac{1}{p^{2}} + κ_{1} (p, T)) .$

Proof. (i) By Assumption 2.1 and Lemma A.2, u_itu_jt satisfies the exponential tail condition, with parameter r₁/3 as shown in the proof of Lemma A.2. Therefore by the Bernstein’s inequality (Theorem 1 of Merlevède (2009)), there exist constants C₁, C₂, C₃, C₄ and C₅ > 0 that only depend on b₁, r₁ and r₂ such that for any i, j ≤ p, and $γ^{- 1} = 3 r_{1}^{- 1} + r_{2}^{- 1}$ ,

P (| \frac{1}{T} \sum_{t = 1}^{T} u_{it} u_{jt} - σ_{ij} | \geq s) \leq T exp (- \frac{{(T s)}^{γ}}{C_{1}}) + exp (- \frac{T^{2} s^{2}}{C_{2} (1 + {TC}_{3})}) + exp (- \frac{{(T s)}^{2}}{C_{4} T} exp (\frac{{(T s)}^{γ (1 - γ)}}{C_{5} {(log T s)}^{γ}})) .

Using Bonferroni’s method, we have

P (max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} u_{it} u_{jt} - σ_{ij} | > s) \leq p^{2} max_{i, j \leq p} P (| \frac{1}{T} \sum_{t = 1}^{T} u_{it} u_{jt} - σ_{ij} | > s) .

Let $s = C \sqrt{(log p) / T}$ for some C > 0. It is not hard to check that when (log p)^2/γ−1 = o(T) (by assumption), for large enough C,

p^{2} T exp (- \frac{{(T s)}^{γ}}{C_{1}}) + p^{2} exp (- \frac{{(T s)}^{2}}{C_{4} T} exp (\frac{{(T s)}^{γ (1 - γ)}}{C_{5} {(log T s)}^{γ}})) = o (\frac{1}{p^{2}}),

and

p^{2} exp (- \frac{T^{2} s^{2}}{C_{2} (1 + {TC}_{3})}) = O (\frac{1}{p^{2}}) .

This proves (i).

(ii) For some $C_{1}^{'} > 0$ such that

P (max_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} {(û_{it} - u_{it})}^{2} > C_{1}^{'} a_{T}^{2}) = O (κ_{1} (p, T)),

(A.2)

under the event ${{max}_{i \leq p} | \frac{1}{T} \sum_{t = 1}^{T} u_{it}^{2} - σ_{ii} | \leq {max}_{i \leq p} σ_{ii} / 4} \cap {{max}_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} {(û_{it} - u_{it})}^{2} \leq C_{1}^{'} a_{T}^{2}}$ , by Cauchy-Schwarz inequality,

\begin{matrix} Z & \equiv max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} (û_{it} û_{jt} - u_{it} u_{jt}) | \\ \leq max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} (û_{it} - u_{it}) (û_{jt} - u_{jt}) | + 2 max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} u_{it} (û_{jt} - u_{jt}) | \\ \leq max_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} {(û_{it} - u_{it})}^{2} + 2 \sqrt{max_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} u_{it}^{2}} \sqrt{max_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} {(û_{it} - u_{it})}^{2}} \\ \leq C_{1}^{'} a_{T}^{2} + 2 \sqrt{\frac{5}{4} max_{i \leq p} σ_{ii}} \sqrt{C_{1}^{'} a_{T}^{2}} . \end{matrix}

Since a_T = o(1), when $C > 3 \sqrt{C_{1}^{'} {max}_{i \leq p} σ_{ii}}$ , we have, for all large T,

{Ca}_{T} > C_{1}^{'} a_{T}^{2} + 2 \sqrt{\frac{5}{4} max_{i \leq p} σ_{ii}} \sqrt{C_{1}^{'} a_{T}^{2}},

and

P (Z \leq {Ca}_{T}) \geq 1 - P (max_{i \leq p} | \frac{1}{T} \sum_{t = 1}^{T} u_{it}^{2} - σ_{ii} | > max_{i \leq p} σ_{ii} / 4) - P (max_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} {(û_{it} - u_{it})}^{2} > C_{1}^{'} a_{T}^{2}) .

By part (i) and (A.2), P(Z ≤ Ca_T) ≥ 1 − O(p⁻² + κ₁(p, T)).

(iii) By (i) and (ii), there exists C_r > 0, when C > C_r, the displayed inequalities in (i) and (ii) hold. Under the event ${{max}_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} u_{it} u_{jt} - σ_{ij} | \leq C \sqrt{(log p) / T}} \cap {{max}_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} û_{it} û_{jt} - u_{it} u_{jt} | \leq {Ca}_{T}}$ , by the triangular inequality,

\begin{matrix} max_{i, j \leq p} | {\hat{σ}}_{ij} - σ_{ij} | & \leq max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} u_{it} u_{jt} - σ_{ij} | + max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} û_{it} û_{jt} - u_{it} u_{jt} | \\ \leq C (\sqrt{\frac{log p}{T}} + a_{T}) . \end{matrix}

Hence the desired result follows from part (i) and part (ii) of the lemma.

Q.E.D.

Lemma A.4. Under Assumptions 2.1, 2.2,

P (C_{L} \leq min_{i j} {\hat{θ}}_{ij} \leq max_{i j} {\hat{θ}}_{ij} \leq C_{U}) \geq 1 - O (\frac{1}{p^{2}} + κ_{1} (p, T) + κ_{2} (p, T)),

where

\begin{matrix} C_{L} = \frac{1}{45} min_{i j} var (u_{it} u_{jt}) \\ C_{U} = 3 max_{i \leq p} σ_{ii} + 4 max_{ij} var (u_{it} u_{jt}) . \end{matrix}

Proof. (i) Using Bernstein’s inequality and the same argument as in the proof of Lemma A.3(i), we have, there exists $C_{r}^{'} > 0, when C > C_{r}^{'}$ , and (log p)^6/γ−1 = o(T),

P (max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} {(u_{it} u_{jt} - σ_{ij})}^{2} - var (u_{it} u_{jt}) | > C \sqrt{\frac{log p}{T}}) = O (\frac{1}{p^{2}}) .

For some C > 0, under the event $\cap_{i = 1}^{4} A_{i}$ , where

\begin{matrix} A_{1} = {max_{i, j \leq p} | σ_{ij} - {\hat{σ}}_{ij} | \leq C (\sqrt{\frac{log p}{T}} + a_{T})} \\ A_{2} = {max_{i \leq p, t \leq T} | û_{it} - u_{it} | \leq min {\frac{1}{2}, \sqrt{{(20 max_{i} σ_{ii})}^{- 1} min_{i j} var (u_{it} u_{jt})}}} \\ A_{3} = {max_{i \leq p} | \frac{1}{T} \sum_{t = 1}^{T} u_{it}^{2} - σ_{ii} | \leq C \sqrt{\frac{log p}{T}}} \\ A_{4} = {max_{i, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} {(u_{it} u_{jt} - σ_{ij})}^{2} - var (u_{it} u_{jt}) | \leq C \sqrt{\frac{log p}{T}}}, \end{matrix}

we have, for any i, j, by adding and subtracting terms,

\begin{matrix} {\hat{θ}}_{i, j} & = \frac{1}{T} \sum_{t} {(û_{it} û_{jt} - {\hat{σ}}_{ij})}^{2} \\ \leq \frac{2}{T} \sum_{t} {(û_{it} û_{jt} - σ_{ij})}^{2} + 2 max_{i, j} {(σ_{ij} - {\hat{σ}}_{ij})}^{2} \\ \leq \frac{4}{T} \sum_{t} {(û_{it} - u_{it})}^{2} û_{jt}^{2} + \frac{4}{T} \sum_{t} {(û_{jt} - u_{jt})}^{2} u_{it}^{2} + \frac{4}{T} \sum_{t} {(u_{it} u_{jt} - σ_{ij})}^{2} + O (\frac{log p}{T} + a_{T}^{2}) \\ \leq 4 max_{it} | û_{it} - u_{it} |^{2} (max_{i} {\hat{σ}}_{ii} + max_{i} \frac{1}{T} \sum_{t} u_{it}^{2}) + 4 var (u_{it} u_{jt}) + O (\sqrt{\frac{log p}{T}} + \frac{log p}{T} + a_{T}^{2}) \\ \leq (2 C \sqrt{\frac{log p}{T}} + {Ca}_{T} + 2 max_{i} σ_{ii}) + 4 var (u_{it} u_{jt}) + o (1), \end{matrix}

where the O(.) and o(.) terms are uniformly in p an T. Hence under $\cap_{i = 1}^{4} A_{i}$ , for all large enough T, p, uniformly in i, j, we have

{\hat{θ}}_{i, j} \leq 3 max_{i \leq p} σ_{ii} + 4 max_{ij} var (u_{it} u_{jt}) .

Still by adding and subtracting terms, we obtain

\begin{matrix} \frac{1}{T} \sum_{t} {(u_{it} u_{jt} - σ_{ij})}^{2} \\ \leq \frac{4}{T} \sum_{t} {(u_{it} u_{jt} - û_{it} û_{jt})}^{2} + \frac{4}{T} \sum_{t} {(û_{it} û_{jt} - {\hat{σ}}_{ij})}^{2} + 4 {(σ_{ij} - {\hat{σ}}_{ij})}^{2} \\ \leq \frac{8}{T} \sum_{t} u_{it}^{2} {(u_{jt} - û_{jt})}^{2} + \frac{8}{T} \sum_{t} û_{jt}^{2} {(u_{it} - û_{it})}^{2} + 4 {\hat{θ}}_{ij} + O (\frac{log p}{T} + a_{T}^{2}) \\ \leq 8 max_{it} | û_{it} - u_{it} |^{2} (max_{i} {\hat{σ}}_{ii} + max_{j} \frac{1}{T} \sum_{t} u_{jt}^{2}) + 4 {\hat{θ}}_{ij} + o (1) . \end{matrix}

Under the event $\cap_{i = 1}^{4} A_{i}$ , we have

\begin{matrix} 4 {\hat{θ}}_{ij} + o (1) & \geq min_{ij} var (u_{it} u_{jt}) - C \sqrt{\frac{log p}{T}} - 8 max_{it} | û_{it} - u_{it} |^{2} [2 C \sqrt{\frac{log p}{T}} + {Ca}_{T} + 2 max_{i} σ_{ii}] \\ \geq \frac{1}{10} min_{ij} var (u_{it} u_{jt}) . \end{matrix}

Hence for all large T, p, uniformly in i, j, we have ${\hat{θ}}_{ij} \geq \frac{1}{45}$ min_ij var(u_itu_jt).

Finally, by Lemma A.3 and Assumption 2.2,

P (\cap_{i = 1}^{4} A_{i}) \geq 1 - O (\frac{1}{p^{2}} + κ_{1} (p, T) + κ_{2} (p, T)),

which completes the proof.

Q.E.D.

A.2. Proof of Theorem 2.1.

Proof. (i) For the operator norm, we have

‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ \leq max_{i \leq p} \sum_{j = 1}^{p} | {\hat{σ}}_{ij} I (| {\hat{σ}}_{ij} | \geq ω_{T} {\hat{θ}}_{ij}^{1 / 2}) - σ_{ij} |

By Lemma A.3 (iii), there exists C₁ > 0 such that the event

A_{1}^{'} = {max_{i, j \leq p} | {\hat{σ}}_{ij} - σ_{ij} | \leq C_{1} (\sqrt{\frac{log p}{T}} + a_{T})}

occurs with probability $P (A_{1}^{'}) \geq 1 - O (\frac{1}{p^{2}} + κ_{1} (p, T))$ . Let C > 0 be such that $C \sqrt{C_{L}} > 2 C_{1}$ , where C_L is defined in Lemma A.4. Let $ω_{T} = C (\sqrt{\frac{log p}{T}} + a_{T}), b_{T} = C_{1} (\sqrt{\frac{log p}{T}} + a_{T}), then \sqrt{C_{L}} ω_{T} > 2 b_{T}$ , and by Lemma A.4,

\begin{matrix} P (min_{ij} {\hat{θ}}_{ij}^{1 / 2} ω_{T} > 2 b_{T}) & \geq P (min_{ij} {\hat{θ}}_{ij}^{1 / 2} > \sqrt{C_{L}}) \\ \geq 1 - O (\frac{1}{p} + κ_{1} (p, T) + κ_{2} (p, T)) . \end{matrix}

Define the following events

\begin{matrix} A_{2}^{'} = {min_{ij} {\hat{θ}}_{ij}^{1 / 2} ω_{T} > 2 b_{T}} \\ A_{3}^{'} = {max_{ij} {\hat{θ}}_{ij}^{1 / 2} \leq C_{U}^{1 / 2}}, \end{matrix}

where C_U is defined in Lemma A.4. Under $\cap_{i = 1}^{3} A_{i}^{'}$ , the event $| {\hat{σ}}_{ij} | \geq ω_{T} {\hat{θ}}_{ij}^{1 / 2}$ implies |σ_ij| ≥ b_T, and the event $| {\hat{σ}}_{ij} | < ω_{T} {\hat{θ}}_{ij}^{1 / 2}$ implies $| σ_{ij} | < b_{T} + \sqrt{C_{U}} ω_{T}$ . We thus have, uniformly in i ≤ p, under $\cap_{i = 1}^{3} A_{i}^{'}$ ,

\begin{matrix} ‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ & \leq \sum_{j = 1}^{p} | {\hat{σ}}_{ij} I (| {\hat{σ}}_{ij} | \geq ω_{T} {\hat{θ}}_{ij}^{1 / 2}) - σ_{ij} | \\ \leq \sum_{j = 1}^{p} | {\hat{σ}}_{ij} - σ_{ij} | I (| {\hat{σ}}_{ij} | \geq ω_{T} {\hat{θ}}_{ij}^{1 / 2}) + \sum_{j = 1}^{p} | σ_{ij} | I (| {\hat{σ}}_{ij} | < ω_{T} {\hat{θ}}_{ij}^{1 / 2}) \\ \leq \sum_{j = 1}^{p} | {\hat{σ}}_{ij} - σ_{ij} | I (| σ_{ij} | \geq b_{T}) + \sum_{j = 1}^{p} | σ_{ij} | I (| σ_{ij} | < b_{T} + \sqrt{C_{U}} ω_{T}) \\ \leq b_{T} m_{T} + (b_{T} + \sqrt{C_{U}} ω_{T}) m_{T} \\ \leq (\sqrt{C_{L}} + \sqrt{C_{U}}) ω_{T} m_{T} . \end{matrix}

By Lemmas A.3(iii) and A.4, $P (\cap_{i = 1}^{3} A_{i}^{'}) \geq 1 - O (\frac{1}{p^{2}} + κ_{1} (p, T) + κ_{2} (p, T))$ , which proves the result. Q.E.D.

(ii) By part (i) of the theorem, there exists some C > 0,

P (‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ > C ω_{T} m_{T}) = O (\frac{1}{p^{2}} + κ_{1} (p, T) + κ_{2} (p, T)) .

By Lemma A.1,

\begin{matrix} P (λ_{min} ({\hat{Σ}}_{u}^{𝒯}) \geq 0.5 λ_{min} (Σ_{u})) & \geq P (‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ \leq 0.5 λ_{min} (Σ_{u})) \\ \geq 1 - O (\frac{1}{p^{2}} + κ_{1} (p, T) + κ_{2} (p, T)) . \end{matrix}

In addition, when ω_Tm_T = o(1),

\begin{matrix} P (‖ {({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1} ‖ \leq 2 ‖ Σ_{u}^{- 1} ‖ C ω_{T} m_{T}) \\ \geq P (‖ {({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1} ‖ \leq 2 ‖ Σ_{u}^{- 1} ‖ \cdot ‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖, ‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ \leq C ω_{T} m_{T}) \\ \geq P (‖ {({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1} ‖ \leq 2 ‖ Σ_{u}^{- 1} ‖ \cdot ‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖) - P (‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ > C ω_{T} m_{T}) \\ \geq P (‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ \leq 0.5 λ_{min} (Σ_{u})) - O (\frac{1}{p^{2}} + κ_{1} (p, T) + κ_{2} (p, T)) \\ \geq 1 - O (\frac{1}{p^{2}} + κ_{1} (p, T) + κ_{2} (p, T)), \end{matrix}

where the third inequality follows from Lemma A.1 as well.

Q.E.D.

APPENDIX B: PROOFS FOR SECTION 3

B.1. Proof of Theorem 3.1.

Lemma B.1. There exists C₁ > 0 such that,

$P (max_{i, j \leq K} | \frac{1}{T} \sum_{t = 1}^{T} f_{it} f_{jt} - {E f}_{it} f_{jt} | > C_{1} \sqrt{\frac{log T}{T}}) = O (\frac{1}{T^{2}}),$
$P (max_{k \leq K, i \leq p} | \frac{1}{T} \sum_{t = 1}^{T} f_{kt} u_{it} | > C_{1} \sqrt{\frac{log p}{T}}) = O (\frac{1}{p^{2}}) .$

Proof. (i) Let $Z_{ij} = \frac{1}{T} \sum_{t = 1}^{T} (f_{it} f_{jt} - E f_{it} f_{jt})$ . We bound max_ij |Z_ij| using Bernstein type inequality. Lemma A.2 implies that for any i and j ≤ K, f_itf_jt satisfies the exponential tail condition (3.3) with parameter r₃/3. Let $r_{4}^{- 1} = 3 r_{3}^{- 1} + r_{2}^{- 1}$ , where r₂ > 0 is the parameter in the strong mixing condition. By Assumption 3.3, r₄ < 1, and by the Bernstein inequality for weakly dependent data in Merlevède (2009, Theorem 1), there exist C_i > 0, i = 1, …, 5, for any s > 0

max_{i, j} P (| Z_{ij} | > s) \leq T exp (- \frac{{(T s)}^{r_{4}}}{C_{1}}) + exp (- \frac{T^{2} s^{2}}{C_{2} (1 + {TC}_{3})}) + exp (- \frac{{(T s)}^{2}}{C_{4} T} exp (\frac{{(T s)}^{r_{4} (1 - r_{4})}}{C_{5} {(log T s)}^{r_{4}}})) .

(B.1)

Using the Bonferroni inequality,

P (max_{i \leq K, j \leq K} | Z_{ij} | > s) \leq K^{2} max_{i, j} P (| Z_{ij} | > s) .

Let $s = C \sqrt{(log T) / T}$ . For all large enough C, since K² = o(T),

\begin{matrix} {TK}^{2} exp (- \frac{{(T s)}^{r_{4}}}{C_{1}}) + K^{2} exp (- \frac{{(T s)}^{2}}{C_{4} T} exp (\frac{{(T s)}^{r_{4} (1 - r_{4})}}{C_{5} {(log T s)}^{r_{4}}})) = o (\frac{1}{T^{2}}), \\ K^{2} exp (- \frac{T^{2} s^{2}}{C_{2} (1 + {TC}_{3})}) = O (\frac{1}{T^{2}}) . \end{matrix}

This proves part (i).

(ii) By Lemma A.2, and Assumptions 2.1(iii) and 3.3(ii), $Z_{ki, t}^{'} \equiv f_{kt} u_{it}$ satisfies the exponential tail condition (2.6) for the tail parameter 2r₁r₃/(3r₁ + 3r₃), as well as the strong mixing condition with parameter r₂. Hence again we can apply the Bernstein inequality for weakly dependent data in Merlevède (2009, Theorem 1) and the Bonferroni’s method on $Z_{ki, t}^{'}$ similar to (B.1) with the parameter $γ_{2}^{- 1} = 1.5 r_{1}^{- 1} + 1.5 r_{3}^{- 1} + r_{2}^{- 1}$ . It follows from $3 r_{1}^{- 1} + r_{2}^{- 1} > 1 and 3 r_{3}^{- 1} + r_{2}^{- 1} > 1$ that γ₂ < 1. Thus when $s = C \sqrt{(log p) / T}$ for large enough C, the term

pK exp (- \frac{T^{2} s^{2}}{C_{2} (1 + {TC}_{3})}) \leq p^{- 2},

and the rest terms on the right hand side of the inequality, multiplied by pK are of order o(p⁻²). Hence when (log p)^2/γ₂−1 = o(T) (which is implied by the Theorem’s assumption), and K = o(p), there exists C′ > 0,

P (max_{k \leq K, i \leq p} | \frac{1}{T} \sum_{t = 1}^{T} f_{kt} u_{it} | > C' \sqrt{\frac{log p}{T}}) = O (\frac{1}{p^{2}}) .

(B.2)

Q.E.D.

Proof of Lemma 3.1

Since $K \sqrt{log T} = o (\sqrt{T}), and λ_{min} (E f_{t} f_{t}^{'})$ is bounded away from zero, for large enough T, by Lemma B.1(i),
$\begin{matrix} P (‖ \frac{1}{T} XX' - E f_{t} f_{t}^{'} ‖ \leq 0.5 λ_{min} (E f_{t} f_{t}^{'})) \\ \geq & P (K max_{i \leq K, j \leq K} | \frac{1}{T} \sum_{t = 1}^{T} f_{it} f_{jt} - {E f}_{it} f_{jt} | \leq 0.5 λ_{min} (E f_{t} f_{t}^{'})) \\ \geq & 1 - O (\frac{1}{T^{2}}) . \end{matrix}$ (B.3)

Hence by Lemma A.1,
$P (λ_{min} (T^{- 1} XX') \geq 0.5 λ_{min} (E f_{t} f_{t}^{'})) \geq 1 - O (\frac{1}{T^{2}}) .$ (B.4)

As b̂_i − b_i = (XX′)⁻¹Xu_i, we have ${‖ {\hat{b}}_{i} - b_{i} ‖}^{2} = u_{i}^{'} X' {(XX')}^{- 2} {Xu}_{i}$ . For C′ > 0 such that (B.2) holds, under the event
$A \equiv {max_{k \leq K, i \leq p} | \frac{1}{T} \sum_{t = 1}^{T} f_{kt} u_{it} | \leq C' \sqrt{\frac{log p}{T}}} \cap {λ_{min} (T^{- 1} XX') \geq 0.5 λ_{min} (E f_{t} f_{t}^{'})},$
we have
$\begin{matrix} {‖ {\hat{b}}_{i} - b_{i} ‖}^{2} & \leq \frac{4}{λ_{min} {(E f_{t} f_{t}^{'})}^{2}} \sum_{k = 1}^{K} {(\frac{1}{T} \sum_{t = 1}^{T} f_{kt} u_{it})}^{2} \\ \leq \frac{4 K}{λ_{min} {(E f_{t} f_{t}^{'})}^{2}} max_{k \leq K, i \leq p} {(\frac{1}{T} \sum_{t = 1}^{T} f_{kt} u_{it})}^{2} \\ \leq \frac{4 KC'^{2} log p}{λ_{min} {(E f_{t} f_{t}^{'})}^{2} T} . \end{matrix}$

The desired result then follows from that $P (A) \geq 1 - O (\frac{1}{T^{2}} + \frac{1}{p^{2}})$ .
For C > max_i≤K ${E f}_{it}^{2}$ , we have, by Lemma B.1(i),
$P (\frac{1}{T} \sum_{t} {‖ f_{t} ‖}^{2} > CK) \leq P (K max_{k \leq K} | \frac{1}{T} \sum_{t = 1}^{T} f_{kt}^{2} - {E f}_{kt}^{2} | + K max_{k \leq K} {Ef}_{kt}^{2} > CK) = O (\frac{1}{T^{2}}) .$

The result then follows from
$max_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} | u_{it} - û_{it} |^{2} \leq max_{i \leq p} \frac{1}{T} \sum_{t} {‖ f_{t} ‖}^{2} {‖ {\hat{b}}_{i} - b_{i} ‖}^{2}$
and part(i).
By Assumption 3.3, for any s > 0,
$\begin{matrix} P (max_{t \leq T} ‖ f_{t} ‖ > s) & \leq TP (‖ f_{t} ‖ > s) \leq TK max_{k \leq K} P (f_{kt}^{2} > s^{2} / K) \\ \leq TK exp (- {(\frac{s}{b_{2} \sqrt{K}})}^{r_{3}}) . \end{matrix}$

When $s \geq C \sqrt{K} {(log T)}^{1 / r_{3}}$ for large enough C, i.e., $C^{r_{3}} > 4 b_{2}^{r_{3}}$ ,
$P (max_{t \leq T} ‖ f_{t} ‖ > C \sqrt{K} {(log T)}^{1 / r_{3}}) \leq T^{- 2} .$

The result then follows from
$max_{t \leq T, i \leq p} | u_{it} - û_{it} | = max_{t \leq T, i \leq p} | ({\hat{b}}_{i} - b_{i})' f_{t} | \leq max_{i} ‖ {\hat{b}}_{i} - b_{i} ‖ max_{t} ‖ f_{t} ‖,$
and Lemma 3.1(i). Q.E.D.

Proof of Theorem 3.1 Theorem 3.1 follows immediately from Theorem 2.1 and Lemma 3.1. Q.E.D.

B.2. Proof of Theorem 3.2 Part (i). Define

\begin{matrix} \begin{matrix} D_{T} = \hat{cov} (f_{t}) - cov (f_{t}), & C_{T} = \hat{B} - B, \end{matrix} \\ E = (u_{1}, \dots, u_{T}) . \end{matrix}

We have,

{‖ {\hat{Σ}}^{𝒯} - Σ ‖}_{Σ}^{2} \leq 4 {‖ {BD}_{T} B' ‖}_{Σ}^{2} + 24 {‖ B \hat{cov} (f) C_{T}' ‖}_{Σ}^{2} + 16 {‖ C_{T} \hat{cov} (f) C_{T}' ‖}_{Σ}^{2} + 2 {‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖}_{Σ}^{2} .

(B.5)

We bound the terms on the right hand side in the following lemmas.

Lemma B.2. There exists C > 0, such that

$P ({‖ D_{T} ‖}_{F}^{2} > \frac{{CK}^{2} log T}{T}) = O (T^{- 2});$
$P ({‖ C_{T} ‖}_{F}^{2} > \frac{CKp log p}{T}) = O (T^{- 2} + p^{- 2}) .$

Proof. (i) Similar to the proof of Lemma B.1(i), it can be shown that there exists C₁ > 0,

P (max_{i \leq K} | \frac{1}{T} \sum_{t = 1}^{T} f_{it} - {E f}_{it} | > C_{1} \sqrt{\frac{log T}{T}}) = O (T^{- 2}) .

Hence sup_K max_i≤K E|f_it| < ∞ implies that there exists C > 0 such that

P (max_{i, j \leq K} | \frac{1}{T} \sum_{t = 1}^{T} f_{it} \frac{1}{T} \sum_{t = 1}^{T} f_{jt} - {E f}_{it} {E f}_{jt} | > C \sqrt{\frac{log T}{T}}) = O (T^{- 2}) .

The result then follows from Lemma B.1(i) and that

{‖ D_{T} ‖}_{F}^{2} \leq K^{2} (max_{i, j \leq K} | \frac{1}{T} \sum_{t = 1}^{T} f_{it} f_{jt} - {E f}_{it} f_{jt} |^{2} + max_{i, j \leq K} | \frac{1}{T} \sum_{t = 1}^{T} f_{it} \frac{1}{T} \sum_{t = 1}^{T} f_{jt} - {E f}_{it} {E f}_{jt} |^{2}) .

(ii) We have C_T = EX′(XX′)⁻¹. By Lemma B.1 (ii), there exists C′ > 0 such that

P (max_{k, i} | \frac{1}{T} \sum_{t = 1}^{T} f_{kt} u_{it} | > C' \sqrt{\frac{log p}{T}}) = O (p^{- 2}) .

Under the event

A = {max_{k, i} | \frac{1}{T} \sum_{t = 1}^{T} f_{kt} u_{it} | \leq C' \sqrt{\frac{log p}{T}}} \cap {λ_{min} (T^{- 1} XX') \geq 0.5 λ_{min} (E f_{t} f_{t}^{'})},

${‖ C_{T} ‖}_{F}^{2} \leq 4 λ_{min}^{- 2} (E f_{t} f_{t}^{'}) {C'}^{2} p K (log p) / T$ , which proves the result since $λ_{min} (E f_{t} f_{t}^{'})$ is bounded away from zero and P(A) ≥ 1 − O (T⁻² + p⁻²) due to (B.4).

Q.E.D.

Lemma B.3. There exists C > 0 such that

$P ({‖ {BD}_{T} B' ‖}_{Σ}^{2} + {‖ B \hat{cov} (f_{t}) C_{T}^{'} ‖}_{Σ}^{2} > \frac{CK log p}{T} + \frac{{CK}^{2} log T}{Tp}) = O (T^{- 2} + p^{- 2});$
$P ({‖ C_{T} \hat{cov} (f) C_{T}' ‖}_{Σ}^{2} > \frac{{CpK}^{2} {(log p)}^{2}}{T^{2}}) = O (T^{- 2} + p^{- 2}) .$

Proof. (i) The same argument in Fan, Fan and Lv (2008), proof of Theorem 2 implies that

‖ B' Σ^{- 1} B ‖ \leq 2 ‖ cov {(f_{t})}^{- 1} ‖ = O (1) .

Hence

\begin{matrix} {‖ {BD}_{T} B' ‖}_{Σ}^{2} & = p^{- 1} tr (Σ^{- 1 / 2} {BD}_{T} B' Σ^{- 1} {BD}_{T} B' Σ^{- 1 / 2}) \\ = p^{- 1} tr (D_{T} B' Σ^{- 1} {BD}_{T} B' Σ^{- 1} B) \\ \leq p^{- 1} {‖ D_{T} B' Σ^{- 1} B ‖}_{F}^{2} \\ \leq O (p^{- 1}) {‖ D_{T} ‖}_{F}^{2} . \end{matrix}

(B.6)

On the other hand,

{‖ B \hat{cov} (f) C_{T}' ‖}_{Σ}^{2} \leq 8 T^{- 2} {‖ BXX' C_{T}' ‖}_{Σ}^{2} + 8 T^{- 4} {‖ BX 11' X' C_{T}' ‖}_{Σ}^{2} .

(B.7)

Respectively,

{\begin{matrix} {‖ BXX' C_{T}' ‖}_{Σ}^{2} & \leq p^{- 1} {‖ XX' C_{T}^{'} Σ^{- 1} ‖}_{F} {‖ C_{T} XX' B' Σ^{- 1} B ‖}_{F} \\ {‖ BX 11' X' C_{T}^{'} ‖}_{Σ}^{2} & \leq p^{- 1} {‖ X 11' X' C_{T}^{'} Σ^{- 1} ‖}_{F} ‖ C_{T} X 11' X' B' Σ^{- 1} B ‖ \end{matrix}}_{F} .

(B.8)

By Lemma B.1(i), and $E f_{t} f_{t}^{'} < \infty$ , P(‖XX′‖ > TC) = O(T⁻²) for some C > 0. Hence, Lemma B.2 (ii) implies

P ({‖ BXX' C_{T}' ‖}_{Σ}^{2} > C' TK log p) = O (T^{- 2} + p^{- 2})

(B.9)

for some C′ > 0. In addition, the eigenvalues of $\hat{cov} (f_{t}) = T^{- 1} XX' - T^{- 2} X 11' X'$ are all bounded away from both zero and infinity with probability at least 1 − O(T⁻²) (implied by Lemmas B.1(i), A.1, and Assumption 3.4). Hence for some C₁ > 0, with probability ast least 1 − O(T⁻²),

\begin{matrix} ‖ X 11' X' ‖ & \leq ‖ T XX' ‖ \leq T^{2} C_{1}, \\ {‖ BX 11' X' C_{T}^{'} ‖}_{Σ}^{2} & \leq O (p^{- 1}) {‖ X 11' X' ‖}^{2} {‖ C_{T} ‖}_{F}^{2} . \end{matrix}

(B.10)

The result then follows from the combination of (B.6)–(B.10), and Lemma B.2.

(ii) Straightforward calculation yields:

\begin{matrix} p {‖ C_{T} \hat{cov} (f) C_{T}' ‖}_{Σ}^{2} & = tr (C_{T} \hat{cov} (f) C_{T}' Σ^{- 1} C_{T} \hat{cov} (f) C_{T}' Σ^{- 1}) \\ \leq {‖ C_{T} \hat{cov} (f) C_{T}' Σ^{- 1} ‖}_{F}^{2} \\ \leq λ_{max}^{2} (Σ^{- 1}) λ_{max}^{2} (\hat{cov} (f_{t})) {‖ C_{T} ‖}_{F}^{4} . \end{matrix}

Since ‖cov(f_t)‖ is bounded, by Lemma B.1(i), $λ_{max}^{2} (\hat{cov} (f_{t}))$ is bounded with probability at least 1 − O(T⁻²). The result again follows from Lemma B.2(ii).

Proof of Theorem 3.2 Part (i)

We have
$\begin{matrix} {‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖}_{Σ} & = p^{- 1 / 2} {‖ Σ^{- 1 / 2} ({\hat{Σ}}_{u}^{𝒯} - Σ_{u}) Σ^{- 1 / 2} ‖}_{F} \\ \leq ‖ Σ^{- 1 / 2} ({\hat{Σ}}_{u}^{𝒯} - Σ_{u}) Σ^{- 1 / 2} ‖ \\ \leq ‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖ \cdot λ_{max} (Σ^{- 1}) . \end{matrix}$ (B.11)

Therefore, (B.5), (B.11), Theorem 3.1 and Lemmas B.2, B.3 yield the result, with the fact that (assuming log T = o(p))
$\frac{K log p}{T} + \frac{K^{2} log T}{Tp} + \frac{p K^{2} {(log p)}^{2}}{T^{2}} + \frac{m_{T}^{2} K^{2} log p}{T} = O (\frac{p K^{2} {(log p)}^{2}}{T^{2}} + \frac{m_{T}^{2} K^{2} log p}{T}) .$
For the infinity norm, it is straightforward to find that
${‖ {\hat{Σ}}^{𝒯} - Σ ‖}_{\infty} \leq {‖ 2 C_{T} cov (f_{t}) B' ‖}_{\infty} + {‖ {BD}_{T} B' ‖}_{\infty} + {‖ C_{T} cov (f_{t}) C_{T}^{'} ‖}_{\infty} + {‖ 2 {BD}_{T} C_{T}^{'} ‖}_{\infty} + {‖ C_{T} D_{T} C_{T}^{'} ‖}_{\infty} + {‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖}_{\infty} .$ (B.12)

By Assumption, both ‖B‖_∞ and ‖cov(f_t)‖_∞ are bounded uniformly in (p, K, T). In addition, let e_i be a p-dimensional column vector whose ith component is one with the remaining components being zeros. Then under the events ${‖ D_{T} ‖}_{\infty} \leq C \sqrt{(log T) / T}, {max}_{i \leq K, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} f_{it} u_{jt} | \leq C \sqrt{(log p) / T}, ‖ \frac{1}{T} XX' ‖ \leq C, {and max}_{i \leq p} ‖ {\hat{b}}_{i} - b_{i} ‖ \leq C \sqrt{K (log p) / T}$ , we have, for some C′ > 0,

\begin{matrix} {‖ 2 C_{T}^{'} cov (f_{t}) B ‖}_{\infty} & \leq 2 max_{i, j \leq p} ‖ e_{i}^{'} C_{T}^{'} cov (f_{t}) {Be}_{j} ‖ \\ \leq 2 max_{i \leq p} ‖ {\hat{b}}_{i} - b_{i} ‖ ‖ cov (f_{t}) ‖ max_{j \leq p} ‖ b_{j} ‖ \\ \leq C' K \sqrt{\frac{log p}{T}}, \end{matrix}

(B.13)

\begin{matrix} {‖ C_{T} ‖}_{\infty} & = max_{i, j \leq p} | e_{i}^{'} \frac{1}{T} EX' {(\frac{1}{T} XX')}^{- 1} e_{j} | \leq max_{i \leq p} ‖ e_{i}^{'} \frac{1}{T} EX' ‖ \cdot ‖ {(\frac{1}{T} XX')}^{- 1} ‖ \\ \leq \sqrt{K} max_{i \leq K, j \leq p} | \frac{1}{T} \sum_{t = 1}^{T} f_{it} u_{jt} | \cdot ‖ {(\frac{1}{T} XX')}^{- 1} ‖ \\ \leq C' \sqrt{K (log p) / T}, \end{matrix}

(B.14)

{‖ B' D_{T} B ‖}_{\infty} \leq K^{2} {‖ B ‖}_{\infty}^{2} {‖ D_{T} ‖}_{F} \leq C' K^{2} \sqrt{\frac{log T}{T}},

(B.15)

\begin{matrix} {‖ C_{T}^{'} cov (f_{t}) C_{T} ‖}_{\infty} & \leq max_{i, j \leq p} ‖ e_{i}^{'} C_{T}^{'} cov (f_{t}) C_{T} e_{j} ‖ \\ \leq max_{i \leq p} {‖ e_{i}^{'} C_{T}^{'} ‖}^{2} ‖ cov (f_{t}) ‖ \leq \frac{C' K^{2} log p}{T}, \end{matrix}

(B.16)

\begin{matrix} {‖ 2 B' D_{T} C_{T} ‖}_{\infty} & \leq 2 K^{2} {‖ B ‖}_{\infty} {‖ D_{T} ‖}_{\infty} {‖ C_{T} ‖}_{\infty} \\ = o (K^{2} \sqrt{\frac{log T}{T}}), \end{matrix}

(B.17)

and

{‖ C_{T}^{'} D_{T} C_{T} ‖}_{\infty} \leq K^{2} {‖ D_{T} ‖}_{\infty} {‖ C_{T} ‖}_{\infty}^{2} = o (K^{2} \sqrt{\frac{log T}{T}}) .

(B.18)

Moreover, the (i, j)th entry of ${\hat{Σ}}_{u}^{𝒯} - Σ_{u}$ is given by

{\hat{σ}}_{ij} I (| {\hat{σ}}_{ij} | \geq ω_{T} \sqrt{{\hat{θ}}_{ij}}) - σ_{ij} = {\begin{matrix} - σ_{ij}, & if | {\hat{σ}}_{ij} | < ω_{T} \sqrt{{\hat{θ}}_{ij}} \\ {\hat{σ}}_{ij} - σ_{ij}, & o . w . \end{matrix}

Hence ${‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖}_{\infty} \leq {max}_{i, j \leq p} | σ_{ij} - {\hat{σ}}_{ij} | + ω_{T} {max}_{i, j \leq p} \sqrt{{\hat{θ}}_{ij}}$ , which implies that with probability at least 1 − O(p⁻² + T⁻²),

{‖ {\hat{Σ}}_{u}^{𝒯} - Σ_{u} ‖}_{\infty} \leq C' K \sqrt{\frac{log p}{T}} .

(B.19)

The result then follows from the combination of (B.12)–(B.19), (B.4), and Lemmas 3.1,B.1.

Q.E.D.

B.3. Proof of Theorem 3.2 Part (ii). We first prove two technical lemmas to be used below.

Lemma B.4. (i) $λ_{min} (B' Σ_{u}^{- 1} B) \geq cp$ for some c > 0.

(ii) $‖ {[cov {(f)}^{- 1} + B' Σ_{u}^{- 1} B]}^{- 1} ‖ = O (p^{- 1})$ .

Proof. (i) We have

λ_{min} (B' Σ_{u}^{- 1} B) \geq λ_{min} (Σ_{u}^{- 1}) λ_{min} (B' B) .

It then follows from Assumption 3.5 that λ_min(B′B) > cp for some c > 0 and all large p. The result follows since ‖Σ_u‖ is bounded away from infinity.

(ii) It follows immediately from

λ_{min} (cov {(f_{t})}^{- 1} + B' Σ_{u}^{- 1} B) \geq λ_{min} (B' Σ_{u}^{- 1} B) .

Q.E.D.

Lemma B.5. There exists C > 0 such that,

$P (‖ \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B} - B' Σ_{u}^{- 1} B ‖ > {Cpm}_{T} K \sqrt{\frac{log p}{T}}) = O (\frac{1}{p^{2}} + \frac{1}{T^{2}});$
$P (‖ {[\hat{cov} {(f)}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}]}^{- 1} ‖ > \frac{C}{p}) = O (\frac{1}{p^{2}} + \frac{1}{T^{2}});$
for $G = {[\hat{cov} {(f)}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}]}^{- 1}$ ,
$P (‖ \hat{B} G \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} ‖ > C) = O (\frac{1}{p^{2}} + \frac{1}{T^{2}}) .$

Proof. (i) Let $H = ‖ \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B} - B' Σ_{u}^{- 1} B ‖$ .

H \leq 2 ‖ C_{T}^{'} Σ_{u}^{- 1} B ‖ + 2 ‖ C_{T}^{'} ({({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1}) B ‖ + ‖ B ″ ({({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1}) B ‖ + ‖ C_{T}^{'} Σ_{u}^{- 1} C_{T} ‖ + ‖ C_{T}^{'} ({({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1}) C_{T} ‖ .

The same argument of Fan, Fan and Lv (2008) (eq. 14) implies that ${‖ B ‖}_{F} = O (\sqrt{p})$ . Therefore, by Theorem 3.1 and Lemma B.2(ii), it is straightforward to verify the result.

(ii) Since ‖D_T ‖_F ≥ ‖D_T‖, according to Lemma B.2(i), there exists C′ > 0 such that with probability ast least 1 − O(T⁻²), $‖ D_{T} ‖ < C' K \sqrt{(log T) / T}$ . Thus by Lemma A.1, for some C″ > 0,

P (‖ \hat{cov} {(f_{t})}^{- 1} - cov {(f_{t})}^{- 1} ‖ < C ″ ‖ D_{T} ‖) \geq P (‖ D_{T} ‖ < C' K \sqrt{\frac{log T}{T}}) \geq 1 - O (T^{- 2}),

which implies

P (‖ \hat{cov} {(f_{t})}^{- 1} - cov {(f_{t})}^{- 1} ‖ < C ″ C' K \sqrt{\frac{log T}{T}}) \geq 1 - O (T^{- 2}) .

(B.20)

Now let $\hat{A} = \hat{cov} {(f_{t})}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}, and A = cov {(f_{t})}^{- 1} + B' Σ_{u}^{- 1} B$ . Then part (i) and (B.20) imply

P (‖ \hat{A} - A ‖ < C ″ C' K \sqrt{\frac{log T}{T}} + {Cpm}_{T} K \sqrt{\frac{log p}{T}}) \geq 1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}}) .

(B.21)

In addition, $m_{T} K \sqrt{(log p) / T} = o (1)$ . Hence by Lemmas A.1, B.4(ii), for some C > 0,

P (λ_{min} (\hat{A}) \geq C p) \geq P (‖ \hat{A} - A ‖ < C p) \geq 1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}}),

which implies the desired result.

(iii) By the triangular inequality, ${‖ \hat{B} ‖}_{F} \leq {‖ C_{T} ‖}_{F} + O (\sqrt{p})$ . Hence Lemma B.2(ii) implies, for some C > 0,

P ({‖ \hat{B} ‖}_{F} \leq C \sqrt{p}) \geq 1 - O (T^{- 2} + p^{- 2}) .

(B.22)

In addition, since $‖ Σ_{u}^{- 1} ‖$ is bounded, it then follows from Theorem 3.1 that $‖ {({\hat{Σ}}_{u}^{𝒯})}^{- 1} ‖$ is bounded with probability at least 1 − O(p⁻²+T⁻²). The result then follows from the fact that

P (‖ G ‖ > {Cp}^{- 1}) = O (\frac{1}{p^{2}} + \frac{1}{T^{2}}),

which is shown in part (ii).

Q.E.D.

To complete the proof of Theorem 3.2 Part (ii), we follow similar lines of proof as in Fan, Fan and Lv (2008). Using the Sherman-Morrison-Woodbury formula, we have

‖ {({\hat{Σ}}^{𝒯})}^{- 1} - Σ^{- 1} ‖ = ‖ {({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1} ‖ + ‖ ({({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1}) \hat{B} {[\hat{cov} {(f)}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}]}^{- 1} \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} ‖ + ‖ ({({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1}) \hat{B} {[\hat{cov} {(f)}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}]}^{- 1} \hat{B}' Σ_{u}^{- 1} ‖ + ‖ Σ_{u}^{- 1} (\hat{B} - B) {[\hat{cov} {(f)}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}]}^{- 1} \hat{B}' Σ_{u}^{- 1} ‖ + ‖ Σ_{u}^{- 1} (\hat{B} - B) {[\hat{cov} {(f)}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}]}^{- 1} B' Σ_{u}^{- 1} ‖ + ‖ Σ_{u}^{- 1} B ({[\hat{cov} {(f)}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}]}^{- 1} - {[cov {(f)}^{- 1} + B' Σ_{u}^{- 1} B]}^{- 1}) B' Σ_{u}^{- 1} ‖ = L_{1} + L_{2} + L_{3} + L_{4} + L_{5} + L_{6} .

(B.23)

The bound of L₁ is given in Theorem 3.1.

For $G = {[\hat{cov} {(f)}^{- 1} + \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} \hat{B}]}^{- 1}$ , then

L_{2} \leq ‖ {({\hat{Σ}}_{u}^{𝒯})}^{- 1} - Σ_{u}^{- 1} ‖ \cdot ‖ \hat{B} G \hat{B}' {({\hat{Σ}}_{u}^{𝒯})}^{- 1} ‖ .

(B.24)

It follows from Theorem 3.1 and Lemma B.5(iii) that

P (L_{2} \leq {Cm}_{T} K \sqrt{\frac{log p}{T}}) \geq 1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}}) .

The same bound can be achieved in a same way for L₃. For L₄, we have

L_{4} \leq {‖ Σ_{u}^{- 1} ‖}^{2} \cdot ‖ \hat{B} - B ‖ \cdot ‖ \hat{B} ‖ \cdot ‖ G ‖ .

It follows from Lemmas B.2, B.5(ii), and inequality (B.22) that

P (L_{4} \leq C \sqrt{\frac{K log p}{T}}) \geq 1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}}) .

The same bound also applies to L₅. Finally,

L_{6} \leq {‖ B ‖}^{2} {‖ Σ_{u}^{- 1} ‖}^{2} ‖ {\hat{A}}^{- 1} - A^{- 1} ‖ \leq {‖ B ‖}^{2} {‖ Σ_{u}^{- 1} ‖}^{2} ‖ \hat{A} - A ‖ \cdot ‖ {\hat{A}}^{- 1} ‖ \cdot ‖ A^{- 1} ‖,

where both Â and A are defined after inequality (B.20). By Lemma B.4(ii), ‖A⁻¹‖ = O(p⁻¹). Lemma B.5(ii) implies P(‖Â⁻¹ ‖ > Cp⁻¹) = O(p⁻² + T⁻²). Combining with (B.21), we obtain

P (L_{6} \leq {Cm}_{T} K \sqrt{\frac{log p}{T}}) \geq 1 - O (\frac{1}{p^{2}} + \frac{1}{T^{2}}) .

The proof is completed by combining L₁ ~ L₆. Q.E.D.

APPENDIX C: PROOFS FOR SECTION 4

The proof is similar to that of Lemma 3.1. Thus we sketch it very briefly. The OLS is given by

{\hat{b}}_{i} = {(X_{i}^{'} X_{i})}^{- 1} X_{i}^{'} y_{i}, i \leq p .

The same arguments in the proof of Lemma B.1 can yield, for large enough C > 0,

P (max_{i \leq p} ‖ {\hat{b}}_{i} - b_{i} ‖ > C \sqrt{\frac{K log p}{T}}) = O (\frac{1}{p^{2}} + \frac{1}{T^{2}}),

which then implies the rate of

max_{i \leq p} \frac{1}{T} \sum_{t = 1}^{T} {(u_{it} - û_{it})}^{2} \leq max_{i \leq p} {‖ {\hat{b}}_{i} - b_{i} ‖}^{2} \frac{1}{T} \sum_{t = 1}^{T} {‖ f_{it} ‖}^{2} .

The result then follows from a straightforward application of Theorem 2.1. Q.E.D.

Contributor Information

Jianqing Fan, Email: jqfan@princeton.edu.

Yuan Liao, Email: yuanliao@princeton.edu.

Martina Mincheva, Email: mincheva@princeton.edu.

REFERENCES

1.Antoniadis A, Fan J. Regularized wavelet approximations. J. Amer. Statist. Assoc. 2001;96:939–967. [Google Scholar]
2.Bai J. Inferential theory for factor models of large dimensions. Econometrica. 2003;71:135–171. [Google Scholar]
3.Bai J, Ng S. Determining the number of factors in approximate factor models. Econometrica. 2002;70:191–221. [Google Scholar]
4.Bickel P, Levina E. Covariance regularization by thresholding. Ann. Statist. 2008a;36:2577–2604. [Google Scholar]
5.Bickel P, Levina E. Regularized estimation of large covariance matrices. Ann. Statist. 2008b;36:199–227. [Google Scholar]
6.Cai T, Liu W. Adaptive thresholding for sparse covariance matrix estimation. J. Amer. Statist. Assoc. 2011;106:672–684. [Google Scholar]
7.Cai T, Zhou H. Manuscript. University of Pennsylvania; 2010. Optimal rates of convergence for sparse covariance matrix estimation. [Google Scholar]
8.Chamberlain G, Rothschild M. Arbitrage, factor structure and mean-variance analyssi in large asset markets. Econometrica. 1983;51:1305–1324. [Google Scholar]
9.Connor G, Korajczyk R. A Test for the number of factors in an approximate factor model. Journal of Finance. 1993;48:1263–1291. [Google Scholar]
10.Fama E, French K. The cross-section of expected stock returns. Journal of Finance. 1992;47:427–465. [Google Scholar]
11.Lam C, Fan J. Sparsistency and rates of convergence in large covariance matrix estimation. Ann. Statist. 2009;37:4254–4278. doi: 10.1214/09-AOS720. [DOI] [PMC free article] [PubMed] [Google Scholar]
12.Fan J, Fan Y, Lv J. High dimensional covariance matrix estimation using a factor model. J. Econometrics. 2008;147:186–197. [Google Scholar]
13.Fan J, Zhang J, Yu K. Manuscript. Princeton University; 2008. Asset location and risk assessment with gross exposure constraints for vast portfolios. [Google Scholar]
14.Gorman M. Some Engel curves. In: Deaton A, editor. Essays in the Theory and Measurement of Consumer Behavior in Honor of Sir Richard Stone. New York: Cambridge University Press; 1981. [Google Scholar]
15.Harding M. Manuscript. Stanford University; 2009. Structural estimation of high-dimensional factor models. [Google Scholar]
16.James W, Stein C. Estimation with quadratic loss. Proc. Fourth Berkeley Symp. Math. Statist. Probab; Univ. California Press; Berkeley. 1961. pp. 361–379. [Google Scholar]
17.Kmenta J, Gilbert R. Estimation of seemingly unrelated regressions with autoregressive disturbances. J. Amer. Statist. Assoc. 1970;65:186–196. [Google Scholar]
18.Lewbel A. The rank of demand systems: theory and nonparametric estimation. Econometrica. 1991;59:711–730. [Google Scholar]
19.Merlevède F, Peligrad M, Rio E. Manuscript. Université Paris Est.; 2009. A Bernstein type inequality and moderate deviations for weakly dependent sequences. [Google Scholar]
20.Rothman A, Levina E, Zhu J. Generalized thresholding of large covariance matrices. J. Amer. Statist. Assoc. 2009;104:177–186. [Google Scholar]
21.Zellner A. An efficient method of estimating seemingly unrelated regressions and tests for aggregation bias. J. Amer. Statist. Assoc. 1962;57:348–368. [Google Scholar]

[R1] 1.Antoniadis A, Fan J. Regularized wavelet approximations. J. Amer. Statist. Assoc. 2001;96:939–967. [Google Scholar]

[R2] 2.Bai J. Inferential theory for factor models of large dimensions. Econometrica. 2003;71:135–171. [Google Scholar]

[R3] 3.Bai J, Ng S. Determining the number of factors in approximate factor models. Econometrica. 2002;70:191–221. [Google Scholar]

[R4] 4.Bickel P, Levina E. Covariance regularization by thresholding. Ann. Statist. 2008a;36:2577–2604. [Google Scholar]

[R5] 5.Bickel P, Levina E. Regularized estimation of large covariance matrices. Ann. Statist. 2008b;36:199–227. [Google Scholar]

[R6] 6.Cai T, Liu W. Adaptive thresholding for sparse covariance matrix estimation. J. Amer. Statist. Assoc. 2011;106:672–684. [Google Scholar]

[R7] 7.Cai T, Zhou H. Manuscript. University of Pennsylvania; 2010. Optimal rates of convergence for sparse covariance matrix estimation. [Google Scholar]

[R8] 8.Chamberlain G, Rothschild M. Arbitrage, factor structure and mean-variance analyssi in large asset markets. Econometrica. 1983;51:1305–1324. [Google Scholar]

[R9] 9.Connor G, Korajczyk R. A Test for the number of factors in an approximate factor model. Journal of Finance. 1993;48:1263–1291. [Google Scholar]

[R10] 10.Fama E, French K. The cross-section of expected stock returns. Journal of Finance. 1992;47:427–465. [Google Scholar]

[R11] 11.Lam C, Fan J. Sparsistency and rates of convergence in large covariance matrix estimation. Ann. Statist. 2009;37:4254–4278. doi: 10.1214/09-AOS720. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R12] 12.Fan J, Fan Y, Lv J. High dimensional covariance matrix estimation using a factor model. J. Econometrics. 2008;147:186–197. [Google Scholar]

[R13] 13.Fan J, Zhang J, Yu K. Manuscript. Princeton University; 2008. Asset location and risk assessment with gross exposure constraints for vast portfolios. [Google Scholar]

[R14] 14.Gorman M. Some Engel curves. In: Deaton A, editor. Essays in the Theory and Measurement of Consumer Behavior in Honor of Sir Richard Stone. New York: Cambridge University Press; 1981. [Google Scholar]

[R15] 15.Harding M. Manuscript. Stanford University; 2009. Structural estimation of high-dimensional factor models. [Google Scholar]

[R16] 16.James W, Stein C. Estimation with quadratic loss. Proc. Fourth Berkeley Symp. Math. Statist. Probab; Univ. California Press; Berkeley. 1961. pp. 361–379. [Google Scholar]

[R17] 17.Kmenta J, Gilbert R. Estimation of seemingly unrelated regressions with autoregressive disturbances. J. Amer. Statist. Assoc. 1970;65:186–196. [Google Scholar]

[R18] 18.Lewbel A. The rank of demand systems: theory and nonparametric estimation. Econometrica. 1991;59:711–730. [Google Scholar]

[R19] 19.Merlevède F, Peligrad M, Rio E. Manuscript. Université Paris Est.; 2009. A Bernstein type inequality and moderate deviations for weakly dependent sequences. [Google Scholar]

[R20] 20.Rothman A, Levina E, Zhu J. Generalized thresholding of large covariance matrices. J. Amer. Statist. Assoc. 2009;104:177–186. [Google Scholar]

[R21] 21.Zellner A. An efficient method of estimating seemingly unrelated regressions and tests for aggregation bias. J. Amer. Statist. Assoc. 1962;57:348–368. [Google Scholar]

PERMALINK

HIGH DIMENSIONAL COVARIANCE MATRIX ESTIMATION IN APPROXIMATE FACTOR MODELS

Jianqing Fan

Yuan Liao

Martina Mincheva

Abstract

1. Introduction

2. Estimation of Error Covariance Matrix

2.1. Adaptive thresholding

2.2. Asymptotic properties of the thresholding estimator

3. Estimation of Covariance Matrix Using Factors

4. Extension: Seemingly Unrelated Regression

5. Monte Carlo Experiments

5.1. Calibration

Table 1.

Table 2.

5.2. Simulation

5.3. Results

Fig 1.

Fig 3.

Fig 2.

6. Conclusions and Discussions

Acknowledgments

APPENDIX A: PROOFS FOR SECTION 2

APPENDIX B: PROOFS FOR SECTION 3

APPENDIX C: PROOFS FOR SECTION 4

Contributor Information

REFERENCES

ACTIONS

PERMALINK

RESOURCES

Cite

Add to Collections

PERMALINK

HIGH DIMENSIONAL COVARIANCE MATRIX ESTIMATION IN APPROXIMATE FACTOR MODELS

Jianqing Fan

Yuan Liao

Martina Mincheva

Abstract

1. Introduction

2. Estimation of Error Covariance Matrix

2.1. Adaptive thresholding

2.2. Asymptotic properties of the thresholding estimator

3. Estimation of Covariance Matrix Using Factors

4. Extension: Seemingly Unrelated Regression

5. Monte Carlo Experiments

5.1. Calibration

Table 1.

Table 2.

5.2. Simulation

5.3. Results

Fig 1.

Fig 3.

Fig 2.

6. Conclusions and Discussions

Acknowledgments

APPENDIX A: PROOFS FOR SECTION 2

APPENDIX B: PROOFS FOR SECTION 3

APPENDIX C: PROOFS FOR SECTION 4

Contributor Information

REFERENCES

ACTIONS

PERMALINK

RESOURCES

Similar articles

Cited by other articles

Links to NCBI Databases