Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2015 Sep 1.
Published in final edited form as: J Econom. 2015 Mar 7;186(2):367–387. doi: 10.1016/j.jeconom.2015.02.015

Risks of Large Portfolios

Jianqing Fan *,, Yuan Liao , Xiaofeng Shi *
PMCID: PMC4504849  NIHMSID: NIHMS649195  PMID: 26195851

Abstract

The risk of a large portfolio is often estimated by substituting a good estimator of the volatility matrix. However, the accuracy of such a risk estimator is largely unknown. We study factor-based risk estimators under a large amount of assets, and introduce a high-confidence level upper bound (H-CLUB) to assess the estimation. The H-CLUB is constructed using the confidence interval of risk estimators with either known or unknown factors. We derive the limiting distribution of the estimated risks in high dimensionality. We find that when the dimension is large, the factor-based risk estimators have the same asymptotic variance no matter whether the factors are known or not, which is slightly smaller than that of the sample covariance-based estimator. Numerically, H-CLUB outperforms the traditional crude bounds, and provides an insightful risk assessment. In addition, our simulated results quantify the relative error in the risk estimation, which is usually negligible using 3-month daily data.

Keywords: High dimensionality, factor models, principal components, sparse matrix, volatility

1 Introduction

The potential of a portfolio’s loss is termed as the portfolio risk. There are two types of portfolio risks. The systematic risk is the risk inherent to the entire market, such as risk associated with interest rates, currencies, recession, war and political instability, etc. The systematic risk cannot be diversified away, even with a well-diversified portfolio. In contrast, specific risk (or idiosyncratic risk) refers to the risk that affects a very specific group of securities or even an individual security. For example, it can be the risk of price changes due to the unique circumstances of a specific stock. Unlike systematic risk, specific risk can be reduced through diversification.

Estimating and assessing the risk of a large portfolio is an important topic in financial econometrics and risk management. The risk of a given portfolio allocation vector wT is conveniently measured by (wTwT)1/2, in which Σ is a volatility (covariance) matrix of the assets’ returns. Often multiple portfolio risks are at interests and hence it is essential to estimate the volatility matrix Σ. The problem becomes challenging when the portfolio size is large. Suppose we have created a portfolio from two thousand assets and invested in a part of selected assets. The covariance matrix Σ involved then contains over two million unknown parameters. Yet, the sample size based on one year’s daily data is around 252. It is hard to assess the estimation accuracy when the estimation errors from more than two million parameters are aggregated. Hence some regularization method is recommended to estimate and assess risks.

We estimate and assess the risks of a given portfolio vector wT based on factor analysis. Two factor-based methods are compared, previously proposed by Fan et al. (2011, 2013). The first estimator assumes the factors to be known and observable. The second method deals with the case of unknown factors. In both cases, the factor model imposes a conditionally sparse structure, in that the idiosyncratic covariance is a large sparse matrix. This yields to an approximate factor model as in Chamberlain and Rothschild (1983), with a non-diagonal error covariance matrix.

We provide a new and practical method to assess the accuracy of risk estimation wT(^-)wT. In the literature (e.g. Fan, Zhang and Yu, 2012), this term has been bounded by

ξ^T=wT12^-max

where ||wT||1 is the gross exposure of the portfolio, and is bounded when there are no extreme positions in the portfolio. However, this upper bound depends on the unknown Σ, hence is not applicable in practice. In addition, numerical studies in this paper demonstrate that this upper bound is too crude: it is often of the same or even larger scale than the estimated risk. In contrast, we provide a high-confidence level upper bound (H-CLUB) for wT(^-)wT, which is of much smaller scale and easy to compute in practice. H-CLUB is constructed based on the confidence interval for the true risk. For the risk estimator wT^wT and a given ε ∈ (0, 1), we find an H-CLUB Û(ε) such that

P(wT(^-)wTU^(ε))1-ε.

In contrast, P(wT(^-)wTξ^T)=1. Hence H-CLUB is an upper bound for the risk estimation error with high confidence while the traditional bound ξ̂T is of full confidence.

For the inferential theory of the risk estimators with diversified portfolios, we prove that the effects of estimating the factor loadings and unknown factors are asymptotically negligible. Interestingly, it is found that when the dimensionality is larger than the sample size, the factor-based risk estimators have the same asymptotic variances no matter whether the factors are known or not. Hence the high dimensionality is in fact a bless for risk estimation instead of a curse from this point of view. In addition, the asymptotic variance of factor-based estimators is slightly smaller than that of the sample covariance-based estimator, but the difference is small. This demonstrates that the benefit of using a factor model is not in terms of a much smaller asymptotic variance, because the systematic risk cannot be diversified. Rather, factor analysis gives a strictly positive definite covariance estimator, which is essential to estimate the optimal portfolio allocation vector, and also interprets the structure of the portfolio risks.

Using our simulated results based on the model calibrated from the U.S. equity market data, we are able to quantify the relative error of the estimation error or coefficient of variation, defined as STD(wT^wT)/wT^wT, where STD(·) denotes the standard error of the estimated risk. Interestingly, this ratio is just a few percent and is approximately independent of the gross exposure ||wT||1 but sensitive to the length of the time series. On the other hand, we also quantify the relation between the crude bound and the practical H-CLUB. We find that ξ̂T is many times larger than Û(ε), and the ratio ξ̂T/Û(ε) increases as the gross exposure increases. A sampling technique that picks a random portfolio with a given gross exposure level is introduced, which can be useful for portfolio optimization and understanding the overall risks within a given level of gross exposure.

The interest on large portfolios surges recently. Pesaran and Zaffaroni (2008) examined the asymptotic behavior of the portfolio weights. Brodie et al. (2009) and Fan et al. (2012) addressed the problem of portfolio selection using a regularization penalty. Gomez and Gallon (2011) numerically compared several methods of covariance matrix estimation for portfolio management. In particular, the optimal portfolio selection involves inverting an estimated Σ, which is a challenging problem under a large number of assets. Gagliardini, Ossola and Scaillet (2010) considered a random coefficient model for an unbalanced panel, and focused on the observable factors, while we also study the inferential theory of the unobservable factor case. The recent works by Fan et al. (2011, 2013) are only concerned about covariance estimations and no inferential theories were studied. The literature is also found in Jacquier and Polson (2010), Antoine (2011), Chang and Tsay (2010), DeMiguel et al. (2009a, b), Ledoit and Wolf (2003), El Karoui (2010), Lai et al. (2011), Bannouh et al. (2012), Gandy and Veraart (2012), Bianchi and Carvalho (2011), among others.

The rest of the paper is organized as follows. Section 2 introduces risk estimators based on factor analysis under both known and unknown factors. Section 3 constructs the H-CLUB for each risk estimator based on the confidence interval for risks. Section 4 derives the limiting distributions of the risk estimators and compares their asymptotic variances. Section 5 presents simulation results. An empirical study is considered in Section 6. Finally, Section 7 concludes. All the proofs are given in the appendix.

Throughout the paper,

wT1=i=1Nwi

is used to denote the gross exposure of a given portfolio allocation vector. For a square matrix A, λmin(A) and λmax(A) represent its minimum and maximum eigenvalues; ||A||1 = maxiΣj |Aij|. Let ||A||max and ||A|| denote its element-wise sup-norm and operator norm, given by ||A||max = maxi,j |Aij| and A=λmax1/2(AA) respectively.

2 Estimation of Portfolio’s Risks

Let {Rt}t=1T be a strictly stationary time series of an N × 1 vector of observed excess returns and Σ = cov(Rt), often known as the volatility matrix. The portfolio risk of a given allocation vector wT is given by wTwT. With a covariance estimator Σ̂, a straightforward estimator of the portfolio risk is wT^wT. But how good such a substitution estimator is and how to assess its estimation accuracy when the dimension N is large relative to T are the questions addressed here.

The problem of estimating the risk of a given portfolio is challenging due to the high dimensionality of Σ. In most cases N can be much larger than T. We assume Σ to be time-invariant within a short period, which holds approximately for locally stationary time series. Recently, Chang and Tsay (2010) proposed a Cholesky decomposition approach to estimating the large covariance matrix, and used simulation to assess its performance. A natural alternative approach is through the factor analysis (e.g., Stock and Watson 2002, Bai 2003), because the assets’ returns are usually driven by a few market factors. Estimating Σ is possible when both the factor and the idiosyncratic components can be estimated well. We thus consider three estimators for estimating wTwT for a given wT, based on three different estimators Σ̂: sample covariance estimator, factor analysis with either observed or unobserved factors.

2.1 Sample-covariance-based estimator

The first estimator Σ̂ = S is the conventional sample covariance matrix based on {Rt}t=1T. Because we are mainly concerned about the variance, for simplicity and exposition, let us assume that the returns have mean zero and S=T-1t=1TRtRt. The asymptotic impact of using S on the risk management has been studied by Fan et al. (2008, 2012) when N is much larger than T. The sample covariance estimator does not require any structural assumption on the assets’ returns. It was shown by the aforementioned authors that for a given portfolio wT with a bounded gross exposure (that is, ||wT||1 is bounded),

wT(S-)wTwT12S-max=Op(logNT).

However, when N > T, it is well known that S is singular, and therefore may result in an estimated risk close to zero for certain portfolios.

2.2 Estimating risks based on factor analysis

We assume the true data generating process (DGP) of Rt to be an “approximate factor model” (Chamberlain and Rothschild 1983):

Rt=Bft+ut,tT, (2.1)

where B is an N × K matrix of factor loadings; ft is a K × 1 vector of common factors, and ut is an N × 1 vector of idiosyncratic error components. In contrast to N and T, here K is assumed to be fixed. The common factors may or may not be observable. For example, Fama and French (1992) identified three known factors that have successfully described the U.S. stock market. On the other hand, in an empirical study, Bai and Ng (2002) determined two unobservable factors for stocks traded on the New York Stock Exchange during 1994–1998.

Let cov(ft) and Σu = cov(ut) denote the covariance matrices of ft and ut, K × K and N × N respectively. Suppose ft and ut are uncorrelated. The factor model then implies the following decomposition of Σ:

=Bcov(ft)B+u. (2.2)

We shall assume that Σu is a sparse matrix in the sense that many off-diagonal elements of the covariance to be either zero or nearly so. The rationale of the sparsity here is, after the common factors are taken out, the remaining idiosyncratic components should be mostly weakly correlated with each other. The decomposition (2.2) also implies that the first K eigenvalues of Σ grow at the rate O(N), due to those of Bcov(ft)B′, while the remaining eigenvalues are bounded away from both zero and infinity. See, e.g., Onatski (2010) and Ahn and Horenstein (2013).

2.2.1 Known-factor-based estimator

We first consider the case of observable factors, and construct an estimator of Σ based on thresholding the covariance matrix of idiosyncratic errors. Suppose is the least squares estimator of B. The residual sample covariance matrix of ut is then given by

Su=T-1t=1T(u^t-u¯)(u^t-u¯)=(Su,ij)N×N,u^t=Rt-B^ft,u¯=1Tt=1Tu^t.

Let sij(·) : ℝ → ℝ be an entry-dependent adaptive thresholding function and for some thresholding parameter τijf>0,

sij(z)=0whenzτijf,andsij(z)-zτijf. (2.3)

The thresholding parameter is taken to be, for some positive constant C > 0,

τijf=CSu,iiSu,jjlogNT.

A simple example is the hard-thresholding sij(z)=zI(zτijf), namely, setting all correlation coefficients smaller than ClogN/T to zero. The soft-thresholding rule is given by sij(z)=(z-τijf)+. See Antoniadis and Fan (2001), Rothman et al. (2009) and Cai and Liu (2011) for detailed discussions of various thresholding functions.

Let

^u,ij={Su,iii=jsij(Su,ij),ij.

Let cov^(ft) denote the sample covariance of the common factors. Define the estimated covariance matrix as

^f=B^cov^(ft)B^+^u,^u=(^u,ij)N×N. (2.4)

2.2.2 Covariance matrix estimator with unknown factors

When the common factors are unobservable, we estimate Σ by “principal orthogonal complements thresholding” (POET), recently proposed by Fan et al. (2013). Because K, the number of factors, might also be unknown, this estimator uses a data-driven number of factors .

The POET works as follows: let λ̂1 ≥ ··· ≥ λ̂N be the ordered eigenvalues of the sample covariance S, whose corresponding eigenvectors are denoted by {ξ^j}j=1N. We then estimate Σ by

^P,K^=j=1K^λ^jξ^jξ^j+Ω^,Ω^=(Ωij)N×N,Ω^ij={k=K^+1Nλ^kξ^k,i2,i=jsij(k=K^+1Nλ^kξ^k,iξ^k,j),ij.

where sij(·) is the same adaptive thresholding function as before, based on an entry-dependent threshold τijP:

τijP=CΩ^iiΩ^jj(logNT+1N).

Here C is a user-specified constant to maintain the finite sample positive definiteness. Even when N/T → ∞, there is C* > 0 such that for any C > C*, both Σ̂f and Σ̂P,K̂ are strictly positive definite for a given finite sample. Previous simulated and empirical studies suggested that C = 0.5 is a good choice when sij is the soft thresholding (see Fan et al. 2013 and Fryzlewicz 2012).

The number of factors can be estimated by the information criterion as in Bai and Ng (2002):

K^=argmin0kM1Ntr(j=k+1Nλ^jξ^jξ^j)+k(N+T)NTlog(NTN+T), (2.5)

where M is a prescribed upper bound. The accuracy of the risk estimation wT(^P,K^-)wT is robust to over-estimating K, as the variance introduced by unnecessary factors are usually small. This is verified in our simulated studies.

Based on the factor analysis, our proposed risk estimator is either wT^fwT or wT^P,K^wT for a given portfolio allocation vector wT, depending on whether ft is observable. This paper focuses on the limiting distributions of these risk estimators and their assessment for a given diversified wT. We will see that under high dimensionality, the factor-based estimators have the same asymptotic variance, and is smaller than that of the sample covariance-based estimator.

3 Assessment of the Risk Estimation

This section proposes a new method to assess the estimated risks for a given portfolio allocation vector wT. We will assume ||wT||1c for some c ≥ 1, where ||wT||1 is the gross exposure of the portfolio. This prevents extreme positions. For simplicity, we shall also assume the number of factors to be fixed, which is easy to be relaxed to allow for slowly-growing K.

3.1 Data-dependent portfolio vector

One of the simplest examples of wT is the equally weighted (1/N) portfolios. DeMiguel et al. (2009b) empirically showed how such simple strategies can outperform more sophisticated strategies. There are also many methods proposed to choose data-dependent portfolios when the number of assets becomes large. Often one can estimate wT by solving constrained data-dependent optimization problems, or directly estimate the unknown parameters when the ideal portfolio wT depends on some unknown quantities (e.g., Brandt et al. (2009), El Karoui (2010), Yen (2013), Lai et al. (2011)).

While this paper focuses on assessing the risk of a given portfolio vector wT, we allow to use an estimated vector ŵT, which consistently estimates wT in the L1-norm:

w^T-wT1=op(1). (3.1)

The L1-convergence is the right notion for consistency in the current context, because the portfolio vector itself has finite gross exposure ||wT||1. This paper will study the statistical inference about w^Tw^T, which is the true risk of such types of data-dependent portfolios and strategies.

Example 3.1

The global minimum variance portfolio is the solution to the problem:

wTgmv=argminw(ww),suchthatwe=1

where e = (1,…,1), yielding wTgmv=-1e/(e-1e). Although this portfolio does not belong to the efficient frontier, Jagannathan and Ma (2003) showed that its performance is comparable with those of other tangency portfolios. Suppose u-11=O(1) and eΣ−1eCN for some C > 0 and all large N (in the factor model, a sufficient condition is eu-1B=o(N). Then ||Σ−1||1 = O(1) and wTgmv1(CN)-1-11N=O(1), so wTgmv has a finite gross exposure.

When N > T, we can estimate wTgmv using a factor model. This yields a data-dependent portfolio:

w^Tgmv=^-1ee^-1e,-1={^f-1knownfactors;^P,K^-1,unknownfactors.

To investigate the consistency of w^Tgmv, note that under the condition that the factors are pervasive (Assumption 4.3 (ii) below), the factor-based inverse covariance estimators of Fan et al. (2011, 2013) are L1-consistent, that is,

^f-1--11=op(1),^P,K^-1--11=op(1).

This implies, for ^-1=^f-1 or ^P,K^-1,

w^Tgmv-wTgmv11e-1e(^-1--1)e1+e(^-1--1e)(e^-1e)(e-1e)-1e1.

Note that there is C > 0 so that eΣ−1eCN and eΣ̂−1eCN with probability approaching one. The first term on the right-hand-side is bounded by Op(N−1)||Σ̂−1Σ−1||1 N = op(1); the second term is bounded by Op(N−2||Σ̂−1Σ−1||N)||Σ−1||1 N = op(1). Consequently, w^Tgmv is L1-consistent.

3.2 Measuring risks using full confidence bound

One of the main problems to be addressed is to assess the risk error Δ=w^T(^-)w^T when N > T. A commonly used upper bound for Δ is based on the following inequality:

Δw^T12^-maxξ^T, (3.2)

which is asymptotically tight for risk assessment in the sense that ξ̂T stochastically converges to zero. However, for the purpose of statistical inference, ξ̂T is infeasible as it depends on the true Σ. In addition, our simulation results have shown that the upper bound ξ̂T is actually too crude to be useful. Let us consider the following toy example.

Example 3.2

Consider three stocks with annualized returns that jointly follow a multi-variate Gaussian distribution Inline graphic (0, Σ) where Σ = 0.04 × I3. The equal-weight portfolio wT = (1/3, 1/3, 1/3)′ is constructed. Our task is to estimate the portfolio risk using the sample covariance matrix S based on the simulated 21-day (one month) returns.

The true value of portfolio variance is wTwT=0.0133, which corresponds to a true risk, by definition, (wTwT)1/2=11.55% per annum. Based on a simulated data set, the estimated portfolio variance wTSwT=0.0131, equivalent to a perceived risk (wTSwT)1/2=11.43% per annum. In addition, for this realization, we also get ξ^T=wT12-Smax=0.0248. Based on this upper bound, a simple calculation shows that wTwT[0,0.0379]. In other words, the true risk (wTwT)1/2 lies in [0, 19.46%], an interval that is too wide to be meaningful.

Note that the inequality (3.2) holds for every sampling sequence {Rt}t=1T. Hence ξ̂T is in fact an upper bound of full confidence, that is,

P(w^T(^-)w^Tξ^T)=1.

The toy example is typical in the sense that ξ̂T is already too crude for small portfolios. In statistical inference, often people use bounds of high confidence levels instead, e.g., quantities that bound Δ with a high probability. This paper pursues such a high-confidence-level upper bound (H-CLUB) based on the confidence interval.

3.3 High-Confidence-Level Upper Bound

We propose a confidence upper bound for Δ=w^T(^-)w^T to assess the estimation error of the portfolio risks. More specifically, for each proposed matrix estimator Σ̂ and any given ε > 0, we find a quantity Û(ε) such that for all large N and T,

P(w^T(^-)w^TU^(ε))1-ε.

Therefore, Û(ε) is an asymptotic (1−ε)100% confidence upper bound for Δ, which is obtained based on the limiting distribution of w^T^w^T. In addition, it is data-driven (up to user-specified tuning parameters), hence can be easily calculated in practice and used to construct confidence intervals for the true risks.

4 Limiting Distributions of Risk Estimators

4.1 Regularity conditions

We will treat N as an increasing function of T. Hence N grows via a fixed trajectory, e.g., N = NT = Tα for some α > 0, and can be faster than T, namely, α > 1.

We shall show that for a well behaved diversified ŵT,

Tw^T(^-)w^T=1Tt=1TZT,t+op(1),

where in the sample covariance based risk estimator, ZT,t=(wTRt)2-E(wTRt)2; in the factor based risk estimator, ZT,t=(wTBft)2-E(wTBft)2. Hence the limiting distribution of the estimated risk depends on that of t=1TZT,t. Note that {ZT,t}t=1T is a triangular array of weakly dependent random variables, and a triangular array central limit theorem for weakly dependent time series data (e.g., Peligrad 1996) can be applied. For this purpose, some technical assumptions are in order.

4.1.1 Data generating process

Assumption 4.1
  1. Assume that (2.1) holds. In addition, {Rt,ut,ft}t=1T is strictly stationary, {ut}t=1T and {ft}t=1T are independent, and Euit = Efjt = 0 for all i, j.

  2. There exist r1, r2 ∈ (0, 2] and b1, b2 > 0, such that for any s > 0,
    P(uit>s)exp(-(s/b1)r1),P(fjt>s)exp(-(s/b2)r2).
  3. There is C > 0 such that C−1 < λmin(Σu) ≤ λmax(Σu) < C, ||B||max < C and λmin(cov(ft)) > C−1.

For simplicity, we assume the true DGP to be a factor model for all the three risk estimators. In fact, when estimating Σ using the sample covariance only, assuming the factor structure on Rt is not necessary, as long as the exponential-tail condition is assumed directly on Rt. The exponential-tail condition (ii) enables us to apply the large deviation theory to control the uniform convergence of maxiN1Tt=1Tuitft and maxi,jN1Tt=1TRitRjt-ERitRjt, which are needed for the consistent estimation of high-dimensional covariance matrices. On the other hand, when financial noises or factors have heavier tails, the analysis will be much more technically involved for non-i.i.d. data. Research along that line will be left for the future. Note that in the factor model being considered: Rt = Bft + ut, the exponential-tail condition also ensures that both the error term and the factors have bounded finite moments, which also implies that wΣw indeed exists. Moreover, it is standard in the literature to assume that the eigenvalues of covariance matrices for ft and ut are bounded away from both zero and infinity.

The factor model is assumed to be conditionally sparse as follows:

Assumption 4.2

There is q ∈ [0, 1) such that

sNmaxiNj=1Nu,ijq=o(min{(T/logN)(1-q)/2,N(1-q)/2}). (4.1)

When q = 0 we define sN=maxiNj=1NI(u,ij0) as the maximum number of non-vanishing elements in each row. Assumption 4.2, though slightly stronger than those in Chamberlain and Rothschild (1983), is meaningful in practice. For example, when the idiosyncratic components represent firms’ individual shocks, they are either uncorrelated or weakly correlated among the firms across different industries, because the industry specific components are not pervasive for the whole economy (Connor and Korajczyk 1993).

Technically, this assumption is easy to satisfy by many sparse covariances as long as T/log N → ∞. For instance, when Σu is a diagonal matrix, sN = 1; when Σu is a block-diagonal matrix with uniformly bounded block sizes, sN is the maximum block size; when Σu is a banded matrix, sN is the width of the bands. In all these cases q = 0, and sN is bounded, while the right hand side of (4.1) diverges fast. Recently, Gagliardini et al. (2010) discussed a detailed example of block dependence and its relation to the sparsity assumption and the approximate factor structure.

The following assumption is standard for high-dimensional factor analysis (e.g., Bai 2003, Bai and Ng 2002, Stock and Watson 2002). Under these assumptions, the unknown factors and loadings can be consistently estimated.

Assumption 4.3
  1. There is M > 0 such that, E[N-1/2(usut-Eusut)]4<M and EN-1/2i=1Nbiuit4<M.

  2. As N → ∞, the eigenvalues of BB/N are bounded away from both zero and infinity.

4.1.2 Weakly dependent data

The autoregressive function of ZT,t plays a central role in the asymptotic variance of Δ=w^T(^-)w^T. For the sample covariance based risk estimator, the relevant object is ZT,t=(wTRt)2-E(wTRt)2. Hence we define the autoregressive function

γT(h)=cov((wTRt)2,(wTRt+h)2),h

which depends on T through dim(wT) = N = NT. For factor-based risk estimator, the relevant autoregressive function is defined for ZT,t=(wTBft)2-E(wTBft)2:

γf(h)=cov((wTBft)2,(wTBft+h)2)h,

Accordingly, the confidence interval of wTwT depends on

σT2=γT(0)+2h=1γT(h),σf2=γf(0)+2h=1γf(h). (4.2)

We make the following assumption on the strength of the serial dependence.

Assumption 4.4

As T, N → ∞,

  1. 1Th=1ThγT(h)=o(σT2),1Th=1Thγf(h)=o(σf2),

  2. When the factors are unknown, σf2N/T.

Condition (i) is standard for stationary weakly-dependent time series data. Let ρ(h) denotes the autocorrelation function of ZT,t, corresponding to either γT(h) or γf(h). Then by the dominated convergence theorem, a sufficient condition is that h=1ρ(h)< and 1 + 2 Σh≥1 ρ(h) is bounded away from zero, which is satisfied by many standard time series models such as ARMA(1,1). Condition (ii) is needed only when the common factors are not observable, so that a large dimension N helps to estimate the unknown factors accurately. This condition requires at least N/T → ∞ in order for the effect of estimating the common factors to be negligible in the asymptotic expansion of w^T(^-)w^T. See more explanations on this effect in Remark 4.4.

Next, we state the α-mixing condition. This condition enables us to apply the central limit theorem for the triangular array weakly dependent data. Let F-0 and FT denote the σ-algebras generated by {(Rt, ft, ut) : −∞ < t ≤ 0} and {(Rt, ft, ut) : Tt < ∞} respectively. Define the mixing coefficient

α(T)=supAF-0,BFTP(A)P(B)-P(AB).
Assumption 4.5

There exist r3 > 0, M1 > 0 and M2 ∈ (0, 1) such that: for all TInline graphic,

α(T)exp(-M1Tr3),andα(T)M2min{γT(0),γf(0)}.

The stated strong-mixing condition is also standard in the literature (e.g., Merlevède et al (2011)). Note that it also follows from the a-mixing condition that h=1γT(h)=O(1) and h=1γf(h)=O(1) (see Lemma B.6 in the appendix) and hence Condition 4.4 holds easily as noted before. Moreover, we also show in Lemma B.5 that γT(0) = O(1) and γf(0) = O(1). Hence by the definition of (4.2),

σf2γf(0)+2h=1γf(h)=O(1).

Similarly, σT2=O(1).

4.1.3 Diversified portfolio

We are particularly interested in the risks of diversified portfolios. These portfolios diversify away the idiosyncratic risk wTuwT, and hence the risks are mainly contributed by the factor risks wTBcov(ft)BwT. A diversified portfolio vector wT should not be dominated by a few of the cross-sectional units, and should satisfy the following technical condition:

Assumption 4.6

||wT||1 = O(1), wTuwT=o(σfsN-1N1/2-q/2T-1/2), and wTuwT=o(σf4+σfsN-1T-q/2(logN)-(1-q)/2).

Assuming ||wT||1 = O(1) ensures that the portfolio has a bounded exposure, which implies wTwT=O(1) and wT^wT=Op(1) because ||Σ̂||max = Op(1). This is formally proved by Lemma A.1 in the appendix.

The required upper bounds on wTuwT is a technical condition for our asymptotic analysis. Since the eigenvalues of Σu are bounded away from both zero and infinity, this condition can also be understood as requiring the same upper bounds on wT22. While the required upper bounds seem complicated and technical, the intuition is clear: ||wT||2 should decay as N → ∞, so that wT is diversified. Conditions of a similar spirit can be found, for instance, in Chudik et al. (2011, Assumption 2.2). Moreover, recall that q and sN ≥ 1 are defined in Assumption 4.2. In the usual case of q = 0 and sN = O(1), the required upper bounds in the above assumption are simplified to σf N1/2T−1/2 and σf4+σf(logN)-1/2 respectively.

The above condition is also related to the “risk ratio”:

wTuwTwTBcov(ft)BwT,

the ratio of the idiosyncratic risk relative to the factor risk. Assumption 4.6 then requires the idiosyncratic risk to be dominated by the factor risk. To illustrate the intuition, consider the following example.

Example 4.1

Consider a one-factor model on the asset returns with var(ft2)>0. Suppose {ft}t=1T are independent across t, and thus σf2=γf(0)=(wTB)4var(ft2). It is straightforward to show that Assumption 4.6 is satisfied as long as:

wTuwTwTBBwT=o(min{1/logN,N/T}). (4.3)

This condition is often satisfied by a diversified portfolio wT. For example, the equal-weight allocation wT = (1/N, ···, 1/N) gives wTuwT=O(N-1). For CN=N-1i=1Nbi=wTB, (4.3) holds as long as (logN)/N2=o(CN4) and T/N3=o(CN4). This is true since CN is often bounded away from zero and T/N3 → 0.

Finally, the following condition controls the error of using ŵT to estimate wT.

Assumption 4.7

For r1, r2 defined in Assumption 4.1,

w^T-wT1=op(min{σT2,σf2}×min{(logNT)-4/r1,(logT)-4/r2}).

This condition is only slightly stronger than the L1-consistency ||ŵTwT||1 = op(1), because often σT2 and σf2 are bounded away from zero, and log(NT) is growing slowly. As an example, for the global minimum variance portfolio in Example 3.1, suppose σT2>0 and σf2>0,

w^T-wT1=Op(^-1--11)=Op(sN(logNT)(1-q)/2).

Hence Assumption 4.7 is satisfied so long as sN does not grow too fast (in fact, for many examples of sparse matrices, sN is often bounded).

4.2 Sample covariance based risk estimator

Let us start with the risk estimator based on the sample covariance matrix S. The estimation error has an asymptotic expansion w^T(S-)w^T=1Tt=1TZT,t+R, where ZT,t=(wTRt)2-E(wTRt)2. The remaining term R arises from the use of estimated portfolio ŵT, which is asymptotically negligible under Assumption 4.7. To construct the confidence interval for the risk w^Tw^T, let us estimate the autocovariance function γT(h) by

γ^(h)=T-1t=1T-h((w^TRt)2-w^TSw^T)((w^TRt+h)2-w^TSw^T).

In particular, γ^(0)=T-1t=1T(w^TRt)4-(w^TSw^T)2. We employ the Newey-West estimator with Bartlett kernel to estimate the variance σT2=γT(0)+2h=1γT(h):

σ^2=γ^(0)+2h=1L(1-hL)γ^(h), (4.4)

where L = L(T) → ∞ is a truncation parameter (see Newey and West 1987). This estimator is always nonnegative in finite sample.

We are now ready to define the H-CLUB ÛS(ε) under the confidence level (1 − ε)100%, which is data-driven once a user-specified L is determined. Let zε/2 denote the upper ε/2 quantile of the standard normal distribution. Let

U^S(ε)=zε/2σ^/T.

Lemma 4.1

Under Assumptions 4.1, 4.4, 4.5, suppose L3=o(TσT4),LσT2, and h>Lγ(h)=o(σT2), then

σ^2-σT2=op(σT2)andU^S(ε)=o(logNT).
Remark 4.1

The conditions L3=o(TσT4),LσT2, and h>Lγ(h)=o(σT2) are called to control the aggregated errors of estimating γT(h) and the estimation bias when we use a truncated sum to approximate h=1γT(h). These estimation errors often arise in estimating the covariances with serially correlated data (see Newey and West 1987 and Andrews 1991), and the required conditions are satisfied with a slowly growing L, e.g., L3 = o(T) when σT2 is bounded away from zero.

The following theorem gives the limiting distribution of the estimated risk. It also demonstrates that ÛS(ε) is a valid H-CLUB for | w^T(S-)w^T|.

Theorem 4.1

Under the assumptions of Lemma 4.1, as T → ∞ and N = NT → ∞,

[var(t=1T(wTRt)2)]-1/2Tw^T(S-)w^TdN(0,1),

and for any ε > 0,

P(w^T(S-)w^TU^S(ε))1-ε.

By the delta-method, we have the following corollary for the risk estimation. Define

R^(w^T)=w^TSw^T,R(w^T)=w^Tw^T.
Corollary 4.1

Under the assumptions of Lemma 4.1, for any ε > 0, as T, N → ∞,

P(R^(w^T)-R(w^T)U^S(ε)/4w^TSw^T)1-ε.

4.3 Factor-based risk estimator

Let us now approach the problem via factor analysis. We assume Rt = Bft + ut, where in this section, {ft}t=1T are observed common factors. We show that the dominating term in the asymptotic expansion of w^T(^f-)w^T is

1Tt=1T(wTBft)2-E(wTBft)2

Hence the estimation error of the risk only comes from the systematic error brought by the common factors, and the risk component w^T(^u-)w^T introduced by the idiosyncratic error can be diversified away by a selected portfolio allocation vector.

To construct H-CLUB, we need to first estimate γf(h), the autocovariance function of (wTBft)2. For cov^(ft)=T-1t=1Tftft, define

γ^f(h)=T-1t=1T-h[(w^TB^ft+h)2-w^TB^cov^(ft)B^w^T][(w^TB^ft)2-w^TB^cov^(ft)B^w^T],

where is the least squares estimator of B. For some L = L(T) → ∞, apply the Newey-West estimator to estimate σf2=γf(0)+2h=1γf(h):

σ^f2=γ^f(0)+2l=1L(1-hL)γ^f(h), (4.5)

which is always nonnegative. Define

U^f(ε)=zε/2σ^f/T.

Let β=3(r1-1+r2-1+r3-1), where r1, r2, r3 are defined as in Assumptions 4.1, 4.5.

Lemma 4.2

Suppose (log N)2β+2 = o(T), and the truncation satisfies L(L+logN)/T+h>Lγf(h)=o(σf2), and Lσf2. Under Assumptions 4.1, 4.2, 4.4–4.6,

σ^f2-σf2=op(σf2),andU^f(ε)=o(logNT).
Remark 4.2

The condition L(L+logN)/T+h>Lγf(h)=o(σf2) and Lσf2 ensure that the effect of estimating the limiting distribution is asymptotically negligible. The first term L(L+logN)/T represents the error of estimating γf(0)+2l=1L(1-h/L)γf(h), while Σh>L γf(h) arises from approximating h=1γf(h) by a truncated sum. These conditions are easily satisfied with a slowly-growing L, for instance, when L3 = o(T) and σf2 is bounded away from zero.

The following theorem shows that Ûf(ε) is a valid H-CLUB for the risk estimation error. Technically, the estimation error for the factor loadings is asymptotically negligible even under high dimensionality.

Theorem 4.2

Suppose that the common factors are observable, and that the thresholded Σ̂f (2.4) is used as the covariance estimator. Under the assumptions of Lemma 4.2,

[var(t=1T(wTBft)2)]-1/2Tw^T(^f-)w^TdN(0,1),

and for any ε > 0,

P(w^T(^f-)w^TU^f(ε))1-ε.
Remark 4.3

Similar to Corollary 4.1, if we use R^f(w^T)=w^T^fw^T to estimate R(w^T)=w^Tw^T, then applying the delta-method yields

P(R^f(w^T)-R(w^T)U^f(ε)/4w^T^fw^T)1-ε.

Hence U^f(ε)/4w^T^fw^T is a valid H-CLUB for |f (ŵT) − R(ŵT)|.

It is also interesting to compare Ûf(ε) with ÛS(ε) and see if knowing the factor structure results in a reduced upper bound. This is equivalent to comparing the asymptotic variances of the estimated risks between a pure nonparametric estimator (sample covariance) and an estimator based on factor analysis. We will see in the following subsection that the factor-based risk estimator indeed gives a slightly smaller asymptotic variance.

Recently Gagliardini et al. (2010) considered a similar covariance thresholding problem for the risk premium in a random coefficient panel model, where the factors are assumed to be observable. They studied the inferential theory for estimating the risk premium. While we are based on similar frameworks of the conditionally sparse factor model, we focus on the inferential theory of the estimated risks for a given portfolio vector, and the impacts on risks from estimating large covariance matrices. In the following subsection, we also study the unobservable factor case.

4.4 Risk estimation with unknown factors

When the market assets’ returns are driven by a few unknown factors, one needs to handle the difficulty of not knowing the common factors in estimating the risk covariance matrix. In this case, the covariance estimator of Fan et al. (2013) is defined by

^P,K=j=1K^λ^jξ^jξ^j+Ω^. (4.6)

as described in Section 2.2. The risk estimator for a given portfolio ŵT is then w^T^P,K^w^T, which incorporates the case of unknown number of factors K. As we shall demonstrate below, using the consistent estimator of K does not affect the asymptotic behavior of the covariance estimator. In addition to the systematic risk t=1TZT,t as before, there are three other components in the asymptotic expansion of the risk: effects of estimating the unknown loadings, factors, as well as the idiosyncratic risk. All these three components are asymptotically negligible.

Under the conditional sparsity condition, Fan et al. (2011, 2013) showed that, when the common factors are observable,

^f-1--1=Op(sN(logNT)1/2-q/2). (4.7)

When the common factors are unobservable,

^P,K^-1--1=Op(sN(logNT+1N)1/2-q/2). (4.8)

where q and sN are defined in Assumption 4.2. The term 1/N in (4.8) is the price for not knowing ft. When T = o(N log N), the above convergence rates are the same. This also explains the condition (iii) in Assumption 4.4, which requires N/T → ∞ in the case when ft is unobservable. Intuitively, as the dimensionality increases, more information about the common factors is collected, and eventually the common factors can be treated as though they are known. Therefore, the effect of estimating the unknown factors on the estimated risk is negligible, and w^T^P,Kw^T and w^T^fw^T have the same limiting distribution. This is often true for the asset returns’ time series data. The number of assets can be in thousands while the sample size on the monthly returns over ten years is slightly larger than a hundred.

To define an H-CLUB for a factor model with unknown factors, we first apply the principal components method (e.g., Stock and Watson 2002) to estimate σf2. Let = (1, ···, T) be a × T matrix such that the rows of F^/T are the eigenvectors corresponding to the largest eigenvalues of the T × T matrix RR, where R = (R1, ···, RT). Let = RF̂′/T. Define

γ^P(h)=T-1t=1T-h[(w^TB^f^t+h)2-w^TB^B^w^T][(w^TB^f^t)2-w^TB^B^w^T].

For some L = L(T) → ∞, let

σ^P2=γ^P(0)+2h=1L(1-hL)γ^P(h),U^P(ε)=zε/2σ^P2/T. (4.9)

The following lemma characterizes the estimation error of σ^P2.

Lemma 4.3

Suppose conditions in Lemma 4.2 are satisfied and L=o(Nσf2). Under Assumptions 4.1–4.6,

σ^P2-σf2=op(σf2),U^P(ε)=o(logNT).
Remark 4.4

Compared to the conditions in Lemma 4.2, the extra requirement L=o(Nσf2) controls the effect of estimating the unknown factors. So the rates of convergence of estimating Σu and σf2 are the same when N(log N)/T → ∞. Intuitively, as the number of unknown factors is O(T), we require more cross sectional units (N) to accurately estimate them. This is also seen in the results in Bai (2003) and Fan et al. (2013).

The following theorem shows that ÛP(ε) is an H-CLUB for w^T(^P,K^-)w^T. Interestingly, w^T^P,K^w^T and w^T^fw^T have the same asymptotic limiting distribution. The price paid for not knowing the factors is asymptotically negligible.

Theorem 4.3

Suppose the common factors are unobservable, and Σ̂P, (4.6) is used as the covariance estimator. Under the assumptions of Lemma 4.3,

[var(t=1T(wTBft)2)]-1/2Tw^T(^P,K-)w^TdN(0,1),

and for any ε > 0,

P(w^T(^P,K^-)w^TU^P(ε))1-ε.
Remark 4.5

Similarly, if we define R^P(w^T)=w^T^P,K^w^T, then U^P(ε)/4w^T^P,K^w^T is a valid H-CLUB for |P(ŵT) − R(ŵT)|.

Knowing the factor-structure of the return Rt improves the estimation efficiency relative to the sample covariance estimator. This is demonstrated by the following theorem.

Theorem 4.4

Under the assumptions of Theorem 4.3,

var[t=1T(wTRt)2]>var[t=1T(wTBft)2].

In fact, the difference of the above two variances is small when wT is diversified enough, and this fact is further verified by our simulation results (see Tables 3 and 5 in Section 5). The reason is that the systematic risk cannot be diversified, and dominates the idiosyncratic risk. On the other hand, factor analysis gives a strictly positive definite covariance estimator, whereas the sample covariance may produce a risk estimator being zero for certain portfolio allocation vectors. The positive definiteness is particularly important to estimate the optimal portfolio allocation vector. Furthermore, factor analysis interprets the structure of portfolio’s risks. It is clearly seen in both Theorems 4.2, 4.3 that the idiosyncratic risks are diversified away by the portfolio allocation.

Table 3.

Averages and standard deviations of RE1 over 500 replications.

c = 1 c = 1.2 c = 1.4 c = 1.6 c = 1.8 c = 2
RE1
S
4.2552 (1.4595) 6.0757 (2.0517) 8.3303 (2.8590) 10.9470 (3.6818) 14.0553 (5.0410) 17.2724 (6.4417)
RE1
Σ̂f
4.2462 (1.4838) 6.0469 (2.0638) 8.3545 (2.9640) 10.9765 (3.7347) 14.0610 (5.1142) 17.3298 (6.5919)
RE1
Σ̂P,K̂
4.1813 (1.4888) 5.9698 (2.0866) 8.2153 (2.9936) 10.8263 (3.7560) 13.8465 (5.1758) 17.0570 (6.6062)
Table 5.

Averages and standard deviations of RE2 over 500 replications, with T = 200.

c = 1 c = 1.2 c = 1.4 c = 1.6 c = 1.8 c = 2
RE2
S
4.8968% (0.8440%) 4.8187% (0.8211%) 4.9171% (0.8694%) 4.8137% (0.8113%) 4.8808% (0.8586%) 4.8217% (0.8085%)
RE2
Σ̂f
4.8888% (0.8428%) 4.8139% (0.8152%) 4.8959% (0.8740%) 4.7921% (0.8043%) 4.8610% (0.8570%) 4.7995% (0.8021%)
RE2
Σ̂P,K̂
4.8918% (0.8443%) 4.8158% (0.8177%) 4.9015% (0.8746%) 4.7967% (0.8039%) 4.8636% (0.8583%) 4.8032% (0.8015%)

5 Monte Carlo Examples

In this section, we examine the finite-sample performance of both the full confidence upper bound ξ̂T defined in (3.2) and H-CLUB, based on three covariance estimators Σ̂, using portfolios wT with different gross exposure constraints.

Excess returns of the ith stock of a portfolio over the risk-free interest rate is assumed to follow the Fama-French three-factor model [Fama and French (1992)]:

Rit=λi1f1t+λi2f2t+λi3f3t+uit.

The first factor is the excess return of the whole equity market, while the second and third factors are SMB (“small minus big” cap) and HML (“high minus low” book/price) respectively. Using US equity market data, we calibrate a sub-model to generate the loadings bi = (λi1, λi2, λi3)′, the idiosyncratic noises ut and the factors ft = (f1t, f2t, f3t)′.

5.1 Calibration

To calibrate parameters in the model, we use the data on daily returns of S&P 500’s top 100 constituents ranked by market capitalization (on June 29th 2012), the data on 3-month Treasury bill rates, and daily return data of the Fama-French factors. They are obtained from COMPUSTAT database, the data library of Kenneth French’s website, and CRSP database respectively. The excess returns (t, t) are analyzed for the period from July 1st, 2008 to June 29th 2012, approximately 1000 trading days.

  1. Calculate the least square estimator of t = Bf̃t + ut, and compute the sample mean vector μB and sample covariance matrix ΣB of all the row vectors of . These parameters are reported in Table 1. The factor loadings {bi}i=1N of the simulated models are then generated from a trivariate Gaussian distribution Inline graphic(μB, ΣB).

  2. Assume that the factors follow the stationary vector autoregressive VAR(1) model ft = μ + Φft−1 + εt for some 3 × 3 matrix Φ, where εt follows i.i.d Inline graphic(0, Σε). The model parameters Φ, μ and Σε are calibrated using the daily excess returns of the Fama-French factors t. The covariance matrix cov(ft) is then obtained by solving the linear equation cov(ft) = Φcov(ft)Φ′ + Σε. Results are summarized in Table 2.

  3. The error covariance matrix is sparse in our setting. For each fixed N, it is created by Σu = 0D, where D = diag(σ1, ···, σp). To be more specific, σ1, ···, σp are generated independently from a Gamma distribution G(α, β), in which α and β are selected to match the sample mean and sample standard deviation of the 100 standard deviations of the errors ũt = tB̃f̃t (recall that each ũt is 100 dimensional; see also Fan, et al, 2008). An additional restriction is imposed on σi that only values in between the minimum and maximum of the standard deviation of ũt are accepted. We then generate the off-diagonal entries of the correlation matrix Σ0 independently from a Gaussian distribution, with mean and standard deviation equal to those of the sample correlations of the estimated residuals. Moreover, absolute values of the off-diagonal entries are set to no greater than 0.95. Finally the hard-thresholding is applied to make Σ0 sparse, where the threshold is set to be the smallest constant that makes Σ0 positive definite.

Table 1.

Mean and covariance used to generate bi

μB ΣB
0.9833 0.0921 −0.0178 0.0436
−0.1233 −0.0178 0.0862 −0.0211
0.0839 0.0436 −0.0211 0.7624

Table 2.

Parameters used to generate ft

μ Φ cov(ft)
0.0260 −0.1006 0.2803 −0.0365 3.2351 0.1783 0.7783
0.0211 −0.0191 −0.0944 0.0186 0.1783 0.5069 0.0102
−0.0043 0.0116 −0.0272 0.0272 0.7783 0.0102 0.6586

5.2 Representative portfolios

We examine the performance of H-CLUB based on wT with a couple of different gross exposures. For a given exposure c and given number of assets N, we randomly generate portfolios wT that satisfy i=1Nwi=1 and i=1Nwi=c. This task, which generates uniformly from the above set in ℝN, is of independent interest for portfolio optimization and research. We propose the following method.

Let w+ be the total long position and w be the total short position: w+ = (c+1)/2 and w = (c − 1)/2, where c = ||wT||1. For c = 1, there are no-short positions. For c > 1, there are both long and short positions (positive and negative components of wT). The identities (or indices) of long and short positions are hard to identify, but the following sampling scheme is a reasonable approximation: The positive positions are determined by a Bernoulli trial (N times) with probability of success w+/(w+ + w) = (c + 1)/(2c). Once the identities are determined, we can normalize them and the problem reduces to the case with c = 1. For the case with c = 1, the uniform distribution on the set { wi:i=1Nwi=1, wi ≥ 0} can be generated from a normalized exponential distribution:

wi=ζi/i=1Nζi,ζi~i.i.d.standardexponential.

Combining the above two steps, we can generate a randomly selected portfolio wT of size N with a gross exposure c as follows.

  1. Generate a positive integer k, the number of stocks with positive weights in wT, from a binomial distribution Bin(N,c+12c).

  2. Generate independently {ζi}i=1k from the standard exponential distribution and set each wi+=(c+1)ζi/(2j=1kζj), for i = 1, ···, k.

  3. Analogously compute for i = 1, ···, Nk, wi-=(1-c)ζi/(2j=1N-kζj), where {ζj}j=1N-k are obtained independently from the standard exponential distribution.

  4. Take the portfolio weights wT as a random permutation of the numbers {wi+}i=1k and {-wi-}i=1N-k.

5.3 Choosing the time lag L

The time lag can be chosen using the plug-in method, previously suggested by Newey and West (1994) and Andrews (1991). The plug-in method chooses L by minimizing the estimated mean squared error of the variance estimator σ̂2(L) of σT2 in (4.2). As suggested by Newey and West (1994), we find L* by

L=1.447T1/3(s^1/s^0)2/3

where

s^1=2h=1nhγ^(h),s^0=γ^(0)+2h=1nγ^(h),n=4(T/100)2/9.

We recommend initially setting L = L*, and then exercise some judgement about sensitivity of results to the choice of n and L. Some evidence from the literature suggests that the final result is less sensitive to n than to L (Silverman 1986). In the simulation and empirical studies below, L is chosen by this method.

5.4 Simulation

5.4.1 Data Generation

For each given gross exposure c, number of assets N and length of time series T, we generate 50 different models and 200 testing portfolios for each model, so a total of 10,000 portfolios are actually used. To be more specific, in any simulation with fixed c, N and T, we repeat the following steps for 50 times:

  1. Generate {bi}i=1N independently from Inline graphic(μB, ΣB). Set B = (b1, ···, bp)′.

  2. Generate {ut}t=1T independently from Inline graphic(0, Σu).

  3. Generate {ft}t=1T from a VAR(1) model ft = μ + Φft−1 + εt with parameters specified in the calibration part.

  4. Calculate Rt = Bft + ut for t = 1, ···, T.

  5. Calculate the sample covariance matrix S=T-1t=1T(yt-y¯)(yt-y¯); obtain the factor-based covariance estimators by soft-thresholding the sample correlation matrices.

  6. Generate 200 wT according to the method described in Section 5.2.

  7. (i) Compute the true risk R(wT)=wTwT; (ii) For Σ̂ = S, Σ̂f and Σ̂P,K̂, calculate Δ=wT(^-)wT,ξ^T=wT12^-max and U^(0.05)=1.96σ^/T; (iii) Calculate empirical coverage probabilities based on the 200 generated portfolios wT, i.e., the proportions of times that Δ < Û(0.05). Note that time lags L are determined by the plug-in method described in Section 5.3.

5.4.2 Output

We first produce the graph of risk domain by plotting averages of R(wT) as a function of c and N, followed by plots of Δ, ξ̂T and Û(ε) against N for sample-based, factor-based and POET-based estimators. Average absolute distance between the empirical coverage probability and the nominal level (95%) are also plotted against N, for all three kinds of estimators. Finally, fix the dimensionality N and increase the number of model generation to 500, we look at two ratio quantities, namely

RE1=ξ^TU^(ε)=wT12^-max1.96var^(wT^wT)andRE2=var^(wT^wT)2wTwT.

Means and standard deviations of RE1 and RE2 under a selection of c and T are reported, for all three kinds of estimators. Respective plots and tables are presented and analyzed in Section 5.5.

5.5 Results

In Figure 1, we gradually increase N from 20 to 600 in increments of 20. Averages of R(wT) over 10,000 portfolios are plotted against N. We produce the curves associated with four choices of c, all with T = 300. For the sake of comparison, we overlay four curves together on the same plot. The following observations can be made from Figure 1.

Figure 1.

Figure 1

Averages of annualized risks R(wT) over 10,000 portfolios.

  1. The average risk ranges from less than 30% to around 50% per annum.

  2. The average risk is higher for larger exposure c. This is consistent with the fact that portfolios with greater gross exposure are more volatile, and hence incur higher risks.

  3. Given a gross exposure c, as the portfolio size N increases, the average risk decreases. The rate of decline is very fast until N is around 150. This is consistent with the theory that as N increases, the portfolio becomes more diversified and the idiosyncratic risk is reduced through diversification.

Let N grow from 100 to 600 in grid of 20. In Figure 2, their respective averages of Δ=w^T(^-)w^T, ξ̂T and Û(ε) over 10,000 portfolios are plotted using estimators Σ̂ = S, Σ̂f and Σ̂P,K̂. In particular, c = 1.6 results in 130% long positions and 30% short positions (130/30 strategy). The 130/30 structure is popular in long-short funds. In each subfigure, the dashed curve corresponds to Δ, the solid curve corresponds to Û(ε), ε = 0.05, and the two-dash curve corresponds to ξ̂T. Based on these plots, we can observe the following features:

Figure 2.

Figure 2

Averages of Δ (dashed curve), Û(ε) with ε = 0.05 (solid curve) and ξ̂T (two-dash curve) over 10,000 portfolios for c = 1, 1.6, 2 and 3, based on three estimated covariance estimators.

  1. Solid curves lie entirely above dashed curves, ensuring Û(ε) as a valid H-CLUB. This feature maybe difficult to observe on Figure 2c and 2d with the presence of two-dash curves (ξ̂T). To make the comparison clearer, we also produce plots without the presence of ξ̂T, as shown in Figure 3 (a zoomed-in version of Figure 2).

  2. The full confidence upper bound ξ̂T is indeed a very crude bound and is much larger than Û(ε). The larger c is, the larger the difference is. This feature is further demonstrated in Table 3.

  3. H-CLUB slightly increases with larger N, but its degree of increase is much smaller than the crude bound ξ̂T.

Figure 3.

Figure 3

Averages of Δ (dashed curve), Û(ε) with ε = 0.05 (solid curve) over 10,000 portfolios. This is a zoom version of Figure 2

In order to further justify the validity of Û(ε) as an H-CLUB, we investigate the empirical coverage probabilities for all three estimators. We employ same settings as in Figure 2 and 3, plot average absolute distances between the empirical coverage probability and the nominal level 95% against dimensionality N in Figure 4. The average absolute distances vary from 1% to 5%, which is in an acceptable range. This is in line with our expectation which ensures Û(ε) as a valid H-CLUB of Δ.

Figure 4.

Figure 4

Average absolute distances between empirical coverage probabilities and the nominal level (95%).

Means and standard deviations (in parentheses) of RE1 for all three kinds of estimators are summarized in Table 3. Here we fix N = 600 and T = 300 with 500 replications. The ratio RE1 quantifies the relation between the full confidence bound and the H-CLUB. Numerical results justify our observations in Figure 2 in the sense that ξ̂T is in general many times greater than Û(ε). Moreover, RE1 increases dramatically as the exposure c increases.

Under the same setting, we also look at values of RE1 under multiple choices of ε. For illustrating purposes, only POET-based estimator with two gross exposures are considered. Means and standard deviations are summarized in Table 4. As ε decreases, it’s not difficult to identify simultaneous decline trends of RE1. This is due to the fact that Û(ε) grows as the confidence level increases.

Table 4.

Averages and standard deviations of RE1 over 500 replications using POET-based estimator.

ε = 0.05 ε = 0.01 ε = 0.005 ε = 0.001
RE1
c = 1
4.1813 (1.4888) 3.2413 (1.1541) 2.4030 (0.8556) 2.1498 (0.7654)
RE1
c = 1.6
10.8263 (2.9936) 8.3925 (2.3206) 6.2220 (1.7205) 5.5662 (1.5391)

Averages and standard deviations of relative error (see e.g. Corollary 4.1)

RE2=var^(wT^wT)/(2wTwT)

with two choices of T are summarized in Table 56, respectively. RE2 measures the accuracy of the perceived risk R^(wT)12 with respect to the true risk R(wT)12. Indeed by delta’s method, RE2ASD(R^(wT)12)/R(wT)12, where “ASD” stands for asymptotic standard deviation. From both tables, it is not difficult to observe that standard deviations are small when compared to their corresponding means. The results also show that the relative error are negligible, at around 3% ~ 5%, ensuring the estimate of R(wT) a high level of accuracy. More interestingly, we realize that this ratio is approximately independent of the gross exposure c but sensitive to the length of the time series. RE2 steadily decreases as T grows. We also observe from Tables 36 that the asymptotic variances (reflected by Û(ε)) of the estimators based on known and unknown factors are almost the same, and slightly smaller than that of the sample covariance estimator.

Table 6.

Averages and standard deviations of RE2 over 500 replications, with T = 400.

c = 1 c = 1.2 c = 1.4 c = 1.6 c = 1.8 c = 2
RE2
S
3.4783% (0.4438%) 3.4684% (0.4215%) 3.4831% (0.4611%) 3.4807% (0.4691%) 3.4979% (0.4392%) 3.4793% (0.4479%)
RE2
Σ̂f
3.4708% (0.4425%) 3.4607% (0.4183%) 3.4742% (0.4626%) 3.4659% (0.4649%) 3.4772% (0.4370%) 3.4553% (0.4438%)
RE2
Σ̂P,K̂
3.4745% (0.4431%) 3.4634% (0.4170%) 3.4775% (0.4617%) 3.4708% (0.4665%) 3.4811% (0.4380%) 3.4586% (0.4431%)

In addition, in order to examine the sensitivity to over-estimating K in the POET-based estimator, we compute RE2 with multiple choices of (the true K equals 3). Averages and standard deviations of RE2 are summarized in Tables 78.

Table 7.

Averages and standard deviations of RE2 over 500 replications, with T = 200.

c = 1 c = 1.2
RE2
Σ̂P,K̂( = 3)
4.8918% (0.8443%) 4.8158% (0.8177%)
RE2
Σ̂P,K̂( = 4)
4.8918% (0.8439%) 4.8154% (0.8178%)
RE2
Σ̂P,K̂( = 5)
4.8921% (0.8439%) 4.8155% (0.8178%)
RE2
Σ̂P,K̂( = 6)
4.8925% (0.8436%) 4.8144% (0.8185%)
RE2
Σ̂P,K̂( = 7)
4.8922% (0.8434%) 4.8145% (0.8192%)

Table 8.

Averages and standard deviations of RE2 over 500 replications, with T = 400.

c = 1 c = 1.2
RE2
Σ̂P,K̂( = 3)
3.4745% (0.4431%) 3.4634% (0.4170%)
RE2
Σ̂P,K̂( = 4)
3.4750% (0.4430%) 3.4635% (0.4170%)
RE2
Σ̂P,K̂( = 5)
3.4750% (0.4431%) 3.4639% (0.4171%)
RE2
Σ̂P,K̂( = 6)
3.4752% (0.4433%) 3.4636% (0.4172%)
RE2
Σ̂P,K̂( = 7)
3.4750% (0.4434%) 3.4635% (0.4175%)

In both tables, the average RE2 for = 3 … 7 are very close to each other, and hence the POET-based risk estimator is robust to over-estimating K.

6 Empirical Studies

We assess the performance of H-CLUB in a portfolio allocation. We use the daily excess returns of 100 industrial portfolios formed on the size and book to market ratio from the website of Kenneth French. The study period is from July 1st 2008 to June 29th 2012, which spans a total of 1000 trading days. At the end of each month the covariance matrix is estimated by three estimators, the sample covariance, the factor-based estimator, and the POET estimator, using daily returns of the preceding 12 months (T = 252). In particular, we employ the Fama-French three-factor model to construct the factor-based estimator. As in Section 5, we adopt plug-in method to estimate the time lags L, refer to Section 5.3 for more details. Two types of strategies are tested, namely the equally weighted portfolio, and the minimum variance portfolio. The optimal portfolios are constructed under two exposure constraints (c = 1 and c = 1.6). The equally weighted portfolio is given by ŵT = (1/N, ···, 1/N). The data-dependent minimum variance portfolio is given by

w^T=argminwT1=1,wT1=cwT^wT.

Portfolios are held for one month and rebalanced at the beginning of the next month.

Because the “true” volatility matrix Σ is unknown, to illustrate our method on the real data, at any time point, we look one month ahead and define the true Σ to be the sample covariance of daily excess returns over the future month. For example, at day 10, the corresponding Σ and R(ŵT) are defined to be

=121t=1131RtRt,R(w^T)=(w^Tw^T)1/2,

where ŵT is the portfolio defined above. On the other hand, the “sample covariance matrix” S in Table 9 below is the sample covariance of daily excess returns over the past twelve months. This is aggregated over the entire testing period. We choose ε = 0.01, i.e. Û(ε) is an empirical 99% upper bound. For each covariance matrix estimator and strategy, we study five quantities, whose respective averages over the whole study period are summarized in Table 9. Recall that U^(ε)/4w^T^w^T is the H-CLUB for the true risk error |(ŵT) − R(ŵT)|, see, for example, Corollary 4.1. All risks are annualized.

Table 9.

True risk errors and estimated risk errors based on the 100 Fama-French Industrial Portfolios.

Strategy Average of Δ(×10−4) Average of Û(0.01)(×10−4) Average of True Risk True Risk Error H-CLUB
Sample-based Covariance Estimator
Equal weighted 2.356 2.757 20.81% 11.18% 11.37%
Min variance (c = 1) 1.006 1.233 14.38% 7.00% 7.44%
Min variance (c = 1.6) 0.497 0.621 11.58% 4.69% 5.17%

Factor-based Covariance Estimator
Equal weighted 2.352 2.768 20.81% 11.16% 11.39%
Min variance (c = 1) 0.999 1.226 14.45% 6.95% 7.41%
Min variance (c = 1.6) 0.475 0.594 11.79% 4.52% 4.98%

POET-based Covariance Estimator
Equal weighted 2.353 2.768 20.81% 11.17% 11.38%
Min variance (c = 1) 1.005 1.231 14.38% 6.99% 7.43%
Min variance (c = 1.6) 0.490 0.626 11.59% 4.61% 5.22%

Here Δ=w^T(-^)wT^ and U^(0.01)=2.58(var^(w^T^w^T))1/2. True risk is R(ŵT). The True Risk Error is (w^Tw^T)1/2-(w^T^w^T)1/2 and H-CLUB is U^(0.01)/4w^T^w^T respectively.

By comparing the first two columns in Table 9, we observe that Û(ε) is uniformly greater that Δ, regardless of the strategies and the covariance matrix estimators. Moreover, the true risk errors are very large, around 40%–50% of the true risk. There are two possible causes: The sample covariance based on one month data (holding period) can be different from the true covariance matrix. It incurs substantial sampling variability. The second cause can be the non-stationarity of the financial returns. However, as shown in the two rightmost columns, results are still satisfactory in the sense that the U-CLUB’s are uniformly larger but quite close (< 1% per annum) to the true risk error.

7 Conclusions

In this paper we address the estimation and assessment for the risk of a large portfolio. The risk is estimated by a substitution of a good estimator of the volatility matrix. We study factor-based risk estimators, based on the approximate factor model with known factors and unknown factors. We derive the limiting distribution of the estimated risks under high dimensionality.

Given that the existing upper bound for the risk estimation error is too crude and not applicable in practice, we introduce a new method, H-CLUB, to assess the accuracy of the risk estimation based on the confidence intervals. Our numerical results demonstrate that the proposed upper bounds significantly outperform the traditional crude bounds, and provide insightful assessment of the estimation of the true portfolio risks.

The empirical study suggests that the financial excess returns may not be globally stationary. Our method also allows for locally stationary time series, as well as slowly time-varying covariance matrices through the localization in time (time-domain smoothing). In fact, the empirical study analyzes a kind of time-varying model, in which we estimate covariance matrices based on the data in rolling windows.

Acknowledgments

The research was partially supported by DMS-1206464 and NIH R01GM100474-01, NIH R01-GM072611.

A Proofs for the Sample Covariance

In this section, ZT,t=wTRtRtwT-EwTRtRtwT, and γT(h) = EZT,tZT,t+h. In particular, γT(0) = var(ZT,t).

We first prove that for σ̂ = S, σ̂f, σ̂P,, the risk estimator wT^wT=Op(1).

Lemma A.1 wT^wT=Op(1) and w^T^w^T=Op(1)

Proof

It is true that for σ̂ = S, σ̂f, or σ̂P,, we have ||σ̂σ||max = op(1) (see Theorem 3.2 of Fan et al. 2013). Hence ||σ̂||max ≤ ||σ||max + ||σ̂σ||max = Op(1). In addition, since ||wT||1 = O(1), we have wT^wTwT12^max=Op(1). Finally, because ||ŵTwT||1 = op(1), ||ŵT||1 = Op(1). The result then also follows from the inequality

w^T^w^Tw^T12^max=Op(1).

A.1 Proof of Lemma 4.1

Lemma A.2

maxtT ||ft|| = Op((log T)1/r2), maxiN,tT |uit| = Op((log NT)1/r1), maxiN,tT |Rit| = Op((log T)1/r2 + (log NT)1/r1).

Proof

Let s = b1(2 log NT)1/r1. Then

P(maxiN,tTuit>s)NTexp(-(s/b1)r1)exp(logNT-(s/b1)r1)0

Hence maxiN,tT |uit| = Op((log NT)1/r1). Similarly, maxtT||ft|| = Op((log T)1/r2).

maxi,tRitmaxibimaxtft+maxiN,tTuit=Op((logT)1/r2+(logNT)1/r1).

Lemma A.3

  1. (wTSwT)2-(wTwT)2=Op(T-1/2σT).

  2. maxhLT-1t=1T-h(wTRt)2(wTRt+h)2-E(wTRt)2(wTRt+h)2=Op(L/T).

  3. maxhLwTSwT-T-1t=1T-h(wTRt)2=Op(L2wTwT/T).

  4. maxhLwTSwT-T-1t=1T-h(wTRt+h)2=Op(L2wTwT/T).

Proof

Note that for any N × N matrix A = (aij), wTAwTAmaxwT12. Thus

(wTSwT)2-(wTwT)2S+maxwT12wT(S-)wT=Op(wT(S-)wT)=Op(T-1t=1TZT,t).

The Chebyshev inequality implies T-1t=1TZT,t=Op(T-1/2σT2).

  • (ii)
    Let Xt,h=(wTRt)2(wTRt+h)2. By the Chebyshev inequality, for any s > 0,
    P(maxhL1Tt=1TXt,h-EXt,h>s)LmaxhLP(1Tt=1TXt,h-EXt,h>s)LmaxhLvar(t=1TXt,h)T2s2.
    Note that maxhLvar(t=1TXt,h)=O(T) since maxhL var(Xt,h) = O(1) and maxhLt=1Tcov(X1,h,Xt+1,h)=O(1). Therefore, for arbitrarily small ε > 0, by choosing s>LM/(εT),P(maxhL1Tt=1TXt,h-EXt,h>s)<ε, which implies maxhL1Tt=1TXt,h-EXt,h=Op(L/T). In addition,
    P(maxhLXt,h-EXt,h>(LmaxhLvar(Xt,h)/ε)1/2)LmaxhLvar(Xt,h)LmaxhLvar(Xt,h)/ε=ε.
    Thus maxhLXt,h-EXt,h=Op(L), which implies
    maxhL1Tt=1T-hXt,h-EXt,hOp(L/T)+maxhLXt,h-EXt,hL/T=Op(L/T).
  • (iii)

    The left hand side is maxhLT-1t=T-h+1T(wTRt)2=maxT-L+1tT(wTRt)2L/T. For any s > 0, P(maxT-L+1tT(wTRt)2>s)LP((wTRt)2>s)LwTwT/s, which then implies maxT-L+1tT(wTRt)2=Op(LwTwT). The desired result then follows.

  • (iv)
    A similar argument as above shows maxT+1tT+L(wTRt)2=Op(LwTwT). Hence maxhLT-1t=T-h+1T(wTRt+h)2maxT+1tT+L(wTRt)2L/T=Op(L2wTwT/T). This implies that the desired quantity is bounded by a+Op(L2wTwT/T) where
    a=maxhL1Tt=1T[(wTRt)2-(wTRt+h)2]1Tt=1L(wTRt)2+1Tt=1L(wTRT+t)2.

    Note that 1Tt=1L(wTRt)2max1tL(wTRt)2L/T=Op(L2wTwT/T). Similarly we have 1Tt=1L(wTRT+t)2=Op(L2wTwT/T).

Lemma A.4 maxhLγ^(h)-γT(h)=Op(L/T+w^T-wT1((logT)4/r2+(logNT)4/r1))

Proof

The triangular inequality implies maxhLγ^(h)-γT(h)i=18ai, where

a1=maxhLT-1t=1T-h(wTRt)2(wTRt+h)2-E(wTRt)2(wTRt+h)2,a2=(wTSwT)2-(wTwT)2,a3=wTSwTmaxhLwTSwT-T-1t=1T-h(wTRt)2,a4=wTSwTmaxhLwTSwT-T-1t=1T-h(wTRt+h)2,a5=maxhL1Tt=1T-h(wTRt)2(wTRt+h)2-(w^TRt)2(w^TRt+h)2a6=maxhL1Tt=1T-h(wTRt)2wTSwT-(w^TRt)2w^TSw^Ta7=maxhL1Tt=1T-h(wTRt+h)2wTSwT-(w^TRt+h)2w^TSw^T,a8=(wTSwT)2-(w^TSw^T)2.

We have, wTSwTwT(S-)wT+wTwT=Op(wTwT+T-1/2σT2). It then follows from Lemma A.3 and σT2=O(1), L3 = O(T), wTwT=O(1) that ai=Op(L/T) for i = 1…4. In addition,

a5wT+w^T1(w^T12+wT12)w^T-wT1(maxiN,tTRit)4=Op(w^T-wT1((logT)4/r2+(logNT)4/r1)),a6wT-w^T1(maxitRit2Smax(wT13+w^T13+w12w^T+w^T12wT1)=Op(w^T-wT1((logT)2/r2+(logNT)2/r1)).

Term a7 is bounded in the same way as a6. Finally,

a8(wT12+w^T12)(wT1+w^T1)Smax2wT-w^T1=Op(wT-w^T1),

which implies maxhLγ^(h)-γT(h)=Op(L/T).

Proof of Lemma 4.1

By the triangular inequality, σ^2-σT2i=13bi, where

b1=γ^(0)-γT(0),b2=2h=1L(1-h/L)γ^(h)-γT(h),b3=2h>LγT(h),b4=2Lh=1LhγT(h)b22LmaxhLγ^(h)-γT(h)=Op(LL/T+Lw^T-wT1((logT)4/r2+(logNT)4/r1)).

Then

σ^2-σT2=Op(L3/2T-1/2+1Lh=1Lhγ(h)+h>LγT(h)+Lw^T-wT1((logT)4/r2+(logNT)4/r1)),

By the lemma’s conditions and Assumption 4.7, the first, third and the fourth terms are all stochastically dominated by σT2. As for the second term 1Lh=1Lhγ(h), note that h=1γT(h)/σT2<. Then by dominated convergence theorem, limT1Lh=1LhσT2γT(h)=h=1limThγT(h)1σT2LI(hL)=0 given that LσT2.

The second part U^S(ε)=o(logN/T) is due to σ^2=Op(σT2) as σT2-σ^2=op(σT2) and σT2=O(1)=o(logN), as N → ∞.

A.2 Proof of Theorem 4.1

Lemma A.5

  1. EZT,12=O(1) and maxlT |γT(l)| = O(1).

  2. For any K ∈ [m, T], var(t=1KZT,t)=KγT(0)+2Kh=1K(1-h/K)γT(h)=O(K).

Proof
  1. It suffices to show E(wTRt)4=O(1). In fact by maxiNERit4=O(1),E(wTRt)4=ijkl=1NwiwjwkwlERitRjtRktRltmaxiNERit4wT14=O(1). The second part follows immediately.

  2. It is well known that for a stationary process with zero mean, var(K-1t=1KZT,t)=K-1γT(0)+2K-1h=1K(1-h/K)γT(h), which implies the result.

Lemma A.6

Under the assumptions of Theorem 4.1,

[var(t=1T(wTRt)2)]-1/2TwT(S-)wTdN(0,1). (A.1)
Proof

The proof is based on Theorem 2.1 of Peligrad (1996). We have TwT(S-)wT=T-1/2t=1TZT,t. Define BT,K2=var(t=1KZT,t) and BT2=var(t=1TZT,t)=O(T). By Davydov’s inequality (Proposition 2.5 of Fan and Yao, 2003 with p = 1/2 and q = 1/4), there are constants M, M1, M2 > 0 such that for any integer h ≥ 0,

γT(h)8α(h)1/4(E(wTRt)2)1/2(E(wTRt)4)1/4=M2exp(-Mhr3/4)

where the last equality follows from the α-mixing condition and that E(wTRt)4=O(1). By the assumption that αR(T) = o(γT(0)), the correlation ρ(T) = |Corr(ZT,t, ZT,t+T)| ≤ |γT(T)|/γT(0) ≤ exp(− MTr3/4). To apply Theorem 2.1 of Peligrad (1996), we also need to check

limsupT1BT2t=1TEZT,t2<. (A.2)

and the Lindeberg condition: ∀ε > 0, 1BT2t=1TEZT,t2I(ZT,t>εBT)0.

Note that 1TBT2/γT(0)=[1+2h=1Tρ(h)](1+o(1)). Hence (A.2) holds because 1+2h=1TρT(h) is bounded away from zero. To verify the Lindeberg condition, note that there are c1, c2 > 0, for all large T, c1<1TBT2/γT(0)<c2. Then for any ε > 0, for some c > 0 and all large T,

1BT2t=1TEZT,t2I(ZT,t>εBT)cγ(0)EZT,t2I(ZT,t2>Tγ(0)ε2c1)

Because EZT,t2γ(0)<, hence by the dominated convergence theorem,

limTcγ(0)EZT,t2I(ZT,t2>Tγ(0)ε2c1)=cElimT1γ(0)ZT,t2I(ZT,t2>Tγ(0)ε2c1)=0.

Hence the conditions of Theorem 2.1 of Peligrad (1996) are satisfied, which implies BT-1t=1TZT,tdN(0,1), equivalent to (A.1).

Proofs of Theorem 4.1 and Corollary 4.1

Note that TσT2/BT21, we have

TBT[w^T(S-)w^T-wT(S-)wT]TBTS-maxw^T-w1(w^1+w1)Op(TσTlogNT)op(σT(logN)-4/r1)=op(1).

Hence by (A.1), TBT[w^T(S-)w^T]N(0,1). This also implies

TσT2w^T(S-)w^TdN(0,1). (A.3)

which also implies w^T(S-)w^T=Op(T-1/2σT2). Moreover, since σT2-σ^2=op(σT2),

Tw^T(S=)w^T|1σT2-1σ^2|=Tw^T(S-)w^Tσ^2-σT2σT2(1+op(1))=op(1).

It then follows from (A.3) that T/σ^2w^T(S-)w^TdN(0,1), which gives the H-CLUB. Corollary 4.1 follows straightforward from applying the delta-method.

B Proofs for the Factor-based Estimation

B.1 Proof of Lemma 4.2

Lemma B.1

maxhLγ^f(h)-γf(h)=Op((L+logN)/T+(logT)4/r2w^T-wT1).

Proof

The triangular inequality implies maxhLγ^f(h)-γf(h)i=14ai, where

a1=maxhLT-1t=1T-h(wTB^ft+h)2(wTB^ft)2-E(wTBft)2(wTBft+h)2,a2=(wTB^cov^(ft)B^wT)2-(wTBcov(ft)BwT)2a3=wTB^cov^(ft)B^wTmaxhLwTB^cov^(ft)B^wT-T-1t=1T-h(wTB^ft)2,a4=wTB^cov^(ft)B^wTmaxhLwTB^cov^(ft)B^wT-T-1t=1T-h(wTB^ft+h)2,a5=maxhLT-1t=1T-h(w^TB^ft+h)2(w^TB^ft)2-(wTB^ft+h)2(wTB^ft)2,a6=maxhLT-1t=1T-h(w^TB^ft+h)2(w^TB^cov^(ft)B^w^T)-(wTB^ft+h)2(wTB^cov^(ft)B^wT)a7=maxhLT-1t=1T-h(w^TB^ft)2(w^TB^cov^(ft)B^w^T)-(wTB^ft)2(wTB^cov^(ft)B^wT)a8=maxhLT-1t=1T-h(w^TB^cov^(ft)B^w^T)2-(wTB^cov^(ft)B^wT)2

a1 is bounded by a11 + a12, where

a11=maxhLT-1t=1T-h(wTBft+h)2(wTBft)2-E(wTBft)2(wTBft+h)2, and a12=maxhLT-1t=1T-h(wTBf^t+h)2(wTBft)2-(wTBft+h)2(wTBft)2.

Given the assumption that maxhLt=1Tcov[(wTBf1)2(wTBf1+h)2,(wTBf1+t)2(wTBf1+t+h)2]=O(1), the same argument of the proof of Lemma A.3(ii) implies a11=Op(L/T). On the other hand, by (B.14) of Fan et al. (2011), B^-Bmax=Op(logN/T), which implies wT(B-B)=Op(logN/T). It is then easy to show that a12=Op(logN/T). It follows that a1=Op((L+logN)/T). By the triangular inequality, a2=Op(logN/T). By the same argument of the proof of Lemma A.3, we have a3 = Op(L2/T) = a4. Moreover, because maxiN ||i|| = Op(1), maxiN,tTb^ift=Op((logT)1/r2),

a5w^T+w1w^T-wT1maxiN,tTb^ift4(wT12+w^T12)=Op((logT)4/r2w^T-wT1)a6w^T-wT1maxiN,tTb^ift2B^cov^(ft)B^max(wT12w^T1+wT13+w^T12w^T+wT1)=Op((logT)2/r2w^T-wT1).

Term a7 is bounded as the same as a6. Finally,

a8w^T-wT1B^cov^(ft)B^max2(w^T12+wT12)(w^T1+wT1)=Op(w^T-wT1).
Proof of Lemma 4.2

We have σ^f2-σf2i=13bi, where b1 = |γ̂f(0) − γf(0)|,

b2=2h=1L(1-h/L)γ^f(h)-γf(h),b3=2h>Lγf(h),b4=2Lh=1Lhγf(h)

By the lemma’s condition, b3=op(σf2) By Lemma B.1,

b1+b22LmaxhLγ^f(h)-γf(h)=Op(L(L+logN)/T+L(logT)4/r2w^T-wT1=op(σf2)

As for b4, note that h=1γf(h)/σf2<. By dominated convergence theorem,

limT1Lh=1Lhσf2γf(h)=h=1limThγf(h)1σf2LI(hL)=0

given that Lσf2. The second statement is due to σ^f2=op(logN).

B.2 Proof of Theorem 4.2

Write R = (R1, …, RT) be N × T; F = (f1, …, fT) be r × T, and cov^(ft)=FF/T. We have = RF′(FF′)−1. Define CT = B and DT=cov^(ft)-cov(ft). The we have the following decomposition: wT(^f-)wT=i=14di, where

d1=wTBDTBwT;d2=2wTCTcov^(ft)BwT;d3=wTCTcov^(ft)CTwT,d4=wT(^u-u)wT.

We now study each of the above four terms separately. Let E = (u1, …, uT) be N × T. Then CT = EF′(FF′)−1.

Lemma B.2
  1. FEwT=Op(T1/2(wTuwT)1/4(EwTut4)1/8).

  2. d2=Op(T-1/2(wTuwT)1/4(EwTut4)1/8).

Proof

We have,

EFEwT2=E[tr(wTEFFEwT)]=tr[E(FEwTwTEF)]=tr[E(FE(EwTwTEF)F)]=tr[E(FE(EwTwTE)F)].

Note that E(EwTwTE)=(E[utwTwTus])tt,sT=(cov(wTut,wTus))tt,sT. By Davydov’s inequality, (see, e.g., Proposition 2.5 of Fan and Yao, 2003 with p = 1/2 and q = 1/4), cov(wTut,wTus)8αf(t-s)1/4(wTuwT)1/2(EwTut4)1/4, where α(·) denotes the α-mixing coefficient. By t=1αf(t)1/4<, we have

EFEwT2=k=1rt=1Ts=1Tcov(wTut,wTus)E(fktfks)=O(1)(wTuwT)1/2(EwTut4)1/4t=1Ts=1Tαf(t-s)1/4=O(T(wTuwT)1/2(EwTut4)1/4),

which then implies (i). For part (ii), we have

d2=2wTBcov^(ft)(FF)-1FEwT=2TwTBFEwT2TwTBFEwT.

Now write B = (bij)iN,jK, then wTB2=j=1K(i=1Nwibij)2maxi,jbijKwT12=O(1).

Lemma B.3

For the factor-based thresholded error covariance matrix,

^u-u=Op(sN(logNT)1/2-q/2)
Proof

By Lemma 3.1 in Fan et al. (2011), we have, maxiNT-1t=1T(u^it-uit)2=Op(logN/T). The result then follows from Theorem A.1 in Fan et al. (2013).

Lemma B.4
  1. d3=Op(T-1(wTuwT)1/2(EwTut4)1/4).

  2. d4=Op(sN(logN/T)1/2-q/2wTuwT)

Proof
  1. Because ||(FF′)−1|| = Op(T−1),

    d3=T-1wTEF(FF)-1FEwT=Op(T-2FEwT2). It then follows from Lemma B.2.

  2. it follows from d4^u-uwT2λmin-1(u)^u-uwTuwT and Lemma B.3.

Lemma B.5
  1. E(wTBft)4=O(1) and E(wTRt)4=O(1);

  2. γT(0) = O(1) and γf (0) = O(1).

Proof
  1. For all s > 0, P(|fjt| > s) ≤ exp(−(s/b2)r2) implies
    maxjKEfjt4maxjKxP(fjt4>x)dxexp(-(s/b2)r2)ds<.
    We have, because the dimension of ft is fixed,
    E(wTBft)4wTB4Eft4wT14maxi,jbij4O(1)=O(1).
    In addition,
    E(wTRt)4=E(j=1NwT,jRjt)4=i,j,k,lNwT,iwT,jwT,kwT,lERitRjtRktRltmaxi,j,k,lERitRjtRktRltwT14maxjNERjt4wT14=O(1).
  2. The result follows from (i) and by observing that γT(0)=var((wTRt)2)E(wTRt)4 and that γf(0)=var((wTBft)2)E(wTBFt)4.

Lemma B.6

h=1γf(h)<, and h=1γT(h)<.

Proof

By Davydov’s inequality (Proposition 2.5 of Fan and Yao, 2003 with p = 1/2 and q = 1/4), there are constants M1, M2 > 0 such that for any integer h,

γf(h)8αf(h)1/4(E(wTBft)2)1/2(E(wTBft)4)1/4=M2exp(-Mhr3/4)

where the last equality follows from the α-mixing condition as well as the fact that E(wTBft)4=O(1) due to ||wT||1 = O(1) (Lemma B.5). The first result the follows from h=1exp(-Chr3)< for any C, r3 > 0. The proof of h=1γT(h)< follows from the same arguments.

Lemma B.7

T/σf2d1dN(0,1).

Proof

Let ZT,t=wTB(ftft-Eftft)BwT, which depends on T through dim(wT) = NT. Hence d1=T-1t=1TZT,t. Note that wTB2KBmax2wT12=O(1). Hence EZT,12=O(1). We define Bf,K2=var(t=1KZT,t) and Bf2=var(t=1TZT,t)=O(T). By the assumption that αf(T) = o(γf (0)), the correlation |Corr(ZT,t, ZT,t+T)| ≤ |γf(T)|/γf(0) = o(1). Moreover, the Lindeberg condition can be verified by the same lines as in the proof of Lemma A.6.

Hence the conditions of Theorem 2.1 of Peligrad (1996) are satisfied, which implies

Bf-1t=1TZT,tdN(0,1). (B.1)

Now Bf2=Tγf(0)+2Th=1Tγf(h)-2Th=1Thγf(h)/T. Because h=1Thγf(h)/T=o(γf(0)+2h=1γf(h)), we have T-1/2(σf2)-1/2t=1TZT,tdN(0,1).

Lemma B.8
Tσf2wT(^f-)wTdN(0,1). (B.2)
Proof

In fact, Tσf2wT(^f-)wT=Tσf2d1+Tσf2(d2+d3+d4). By Lemma B.7, it suffices to show that T/σf2(d2+d3+d4)=op(1). By Lemma B.2,

T/σf2d2=Op((wTuwT)1/4(EwTut4)1/8)/σf2)=Op((wTuwT)1/4/σf2)=op(1)

since EwTut4=O(1). Lemma B.4 implies T/σf2d3=Op((wTuwT)1/2(Tσf2)-1/2)=op(1) since wTuwT=o(σf4)=O(1). It also follows from Lemma B.4 that T/σf2d4=Op(wTuwTsN(σf2)-1/2(logN)1/2-q/2Tq/2)=op(1). This implies the desired result.

Proof of Theorem 4.2

By Theorem 3.2 of Fan et al. (2011), ^f-max=Op(logNTT). Note that Tσf2/Bf21, we have

TBf[w^T(^f-)w^T-wT(^f-)wT]TBf^f-maxw^T-w1(w^1+w1)Op(TσflogNTT)op(σf(logNT)-4/r1)=op(1). (B.3)

The first statement [var(t=1T(wTBft)2)]-1/2Tw^T(^f-)w^TdN(0,1) then follows from (B.2) and (B.3). They also yield

Tσf2w^T(^f-)w^TdN(0,1), (B.4)

and w^T(^f-)w^T=Op(T-1/2σf2). Moreover, since σf2-σ^f2=op(σf2),

Tw^T(^f-)w^Tσf-1-σ^f-1=Tw^T(^f-)w^Tσ^f2-σf2σf2(1+op(1))=op(1).

It then follows from (B.4) that T/σ^f2w^T(^f-)w^TdN(0,1), which gives the H-CLUB.

C Proofs for the POET-based Estimation

Let V denote the K × K diagonal matrix of the first K largest eigenvalues of S in decreasing order. Let = (1, …, T) be a K × T matrix such that the rows of F^/T are the eigenvectors corresponding to the r largest eigenvalues of the T × T matrix RR. Let = RF̂′/T. Define a K × K matrix

H=1TV-1F^FBB.

According to Stock and Watson (2002), and t can be treated as estimators of BH−1 and Hft respectively.

C.1 Proof of Lemma 4.3

Lemma C.1
  1. wTB^=Op(1), and wT(B^-BH-1)=Op(N-1/2+(logN/T)1/2)

  2. F^-HF2/T=T-1t=1Tf^t-Hft2=Op(N-1+T-2).

  3. wTE2=Op(T).

  4. T-1t=1T[f^tf^t-Hft(Hft)]=Op(N-1/2+T-1).

Proof
  1. By Lemma B.16 in an earlier version of Fan et al. (2013)1, ||||max ≤|| BH−1||max + ||B||max = Op(1). Thus wTB^2rB^max2wT12=Op(1). On the other hand, wT(B^-BH-1)2rB^-Bmax2wT12=Op(1/N+logN/T).

  2. By (A.1) in Bai (2003), the following identity holds:
    f^t-Hft=(V/N)-1(1Ts=1Tf^sE(usut)/N+1Ts=1Tf^sζst+1Ts=1Tf^sηst+1Ts=1Tf^sξst) (C.1)
    where ζst=usut/N-E(usut)/N,ηst=fsi=1Nbiuit/N, and ξst=fti=1pbiuis/N. It follows from Lemma C.7 in Fan et al. (2013) that
    1Tt=1T(1Ts=1Tf^isζst)2+1Tt=1T(1Ts=1Tf^isηst)2+1Tt=1T(1Ts=1Tf^isξst)2=Op(1N).
    Moreover, by Lemma C.9 of Fan et al. (2013), maxir1Tt=1T(f^t-Hft)i2=Op(1/T+1/N). Applying the inequality (a + b)2 ≤ 2a2 + 2b2 gives,
    1Tt=1T(1Ts=1Tf^isE(usut)/N)21Tt=1T(1Ts=1T[(f^s-Hfs)i+(Hfs)i]E(usut)/N)22Tt=1T(1Ts=1T(f^s-Hfs)iE(usut)/N)2+2Tt=1T(1Ts=1T(Hfs)iE(usut)/N)2.
    By the Cauchy-Schwarz inequality and that maxtTs=1TE(usut)/N2=O(1),
    2Tt=1T(1Ts=1T(f^s-Hfs)iE(usut)N)2maxir2Ts=1T(f^s-Hfs)i21Ts=1T(Eusut/N)2=Op(1T2+1NT).
    Also, 2Tt=1T(1Ts=1T(Hfs)iE(usut)/N)2Op(T-1)t=1T(1Ts=1TfsE(usut)/N)2. We have T-1t=1T(1Ts=1TfsE(usut)/N)2=Op(T-2) since
    E1Tt=1T(1Ts=1TfsE(usut)/N)2=1T2s=1Tl=1TEfsflEusutNEulutNmaxsTEfs2maxsT(1Ts=1TEusut/N)2=O(T-2).

    This implies maxirT-1t=1T(f^t-Hft)i2=Op(N-1+T-2), and thus T-1t=1Tf^t-Hft2=Op(N-1+T-2).

  3. EwTE2=Et=1T(i=1Nwiuit)2=Tmaxi,jEuitujtwT12=O(T). Thus wTE2=Op(T). Finally, (iv) follows from the Cauchy-Schwarz inequality and part (ii).

Lemma C.2

maxhLγ^P(h)-γf(h)=Op(L+logNT+1N+(logT)2/r2w^T-wT1).

Proof

The triangular inequality implies maxhLγ^P(h)-γf(h)i=18ai, where

a1=maxhLT-1t=1T-h(wTB^f^t+h)2(wTB^f^t)2-E(wTBft)2(wTBft+h)2,a2=(wTB^B^wT)2-(wTBBwT)2,a3=wTB^B^wTmaxhLwTB^B^wT-T-1t=1T-h(wTB^f^t)2,a4=wTB^B^wTmaxhLwTB^B^wT-T-1t=1T-h(wTB^f^t+h)2a5=maxhLT-1t=1T-h(w^TB^f^t+h)2(w^TB^f^t)2-(wTB^f^t+h)2(wTB^f^t)2,a6=maxhLT-1t=1T-h(w^TB^f^t+h)2(w^TB^B^w^T)-(wTB^f^t+h)2(wTB^B^wT)a7=maxhLT-1t=1T-h(w^TB^f^t)2(w^TB^B^w^T)-(wTB^f^t)2(wTB^B^wT)a8=maxhLT-1t=1T-h(w^TB^B^w^T)2-(wTB^B^wT)2

Here a1 is bounded by a11 + a12, where a11=maxhLT-1t=1T-h(wTBft+h)2(wTBft)2-E(wTBft)2(wTBft+h)2, and a12=maxhLT-1t=1T-h(wTB^f^t+h)2(wTB^f^t)2-(wTBft+h)2(wTBft)2.

As in the proof of Lemma B.1, a11=Op(L/T). It follows from the Cauchy-Schwarz inequality that

a12=Op(maxhLT-1t=1T-h(wTB^f^t+h)(wTB^f^t)-(wTBft+h)(wTBft))=Op(maxhL(T-1t=1T-h(wTB^f^t+h-wTBft+h)2)1/2+(T-1t=1T-h(wTB^ft-wTBft)2)1/2)=Op(wT(B^-BH-1)+wTBH-1(T-1t=1Tf^t-Hft2)1/2)

It follows from Lemma C.1 that a12=Op(logN/T+N-1/2), thus a1=Op((L+logN)/T+N-1/2). On the other hand, for g1, … , g5 defined in (C.2),

a2=Op(wT(B^B^-BB)wT)=Op(i=15gi)=Op(σf2/T)=op(L/T).a3=Op(1)maxhLwTB^(1Tt=T-h+1Tf^tf^t)B^wT=Op(1Tt=T-L+1Tf^t2).

We have 1Tt=T-L+1Tf^t22Tt=T-L+1Tf^t-Hft2+2Tt=T-L+1THft2. On one hand, E2Tt=T-L+1Tft2=O(L/T). On the other hand, by Theorem 3.3 in Fan et al. (2013), maxt||tHft|| = op(1), hence 2Tt=T-L+1Tf^t-Hft2=Op(L/T). Thus a3 = Op(L/T). Similarly, we have a4 = Op(L/T).

Morever, by Corollary 3.1 of Fan et al. (2013), maxi,tb^if^top(1)+maxi,tbift=op((logT)1/r2), and ||B̂B̂′||max = Op(1). Thus it follows from the same lines of proofs as those in Lemma B.1 that a5 + … + a8 = Op((log T)2/r2 ||ŵTwT||1).

Proof of Lemma 4.3

We have σ^P2-σf2i=13bi, where b1 = |γ̂P (0) − γf (0)|,

b2=2h=1L(1-h/L)γ^P(h)-γf(h),b3=2h>Lγf(h),b4=2Lh=1Lhγf(h)

By Lemma C.2, b1 + b2 ≤ 2L maxhL |γ̂P (h) − γf (h)|, which is Op(L(L+logN)/T+L(logT)4/r2w^T-wT1+L/N)=op(σf2) Hence the proof is the same as that of Lemma 4.2.

C.2 Proof of Theorem 4.3

If we write T = BH−1 and DT=T-1t=1Tftft-cov(ft), then wT(^P,K^-)wT=i=15gi, where

g1=wTBDTBwT,g2=wTCTBwT,g3=wTBH-1CTwT,g4=wT(Ω^-u)wT,g5=wTBH-11Tt=1T[f^tf^t-Hft(Hft)]H-1BwT. (C.2)

Recall the definition of d1 in Appendix B.2, g1 = d1. Thus it follows from Lemma B.7 that T/σf2g1dN(0,1). We first show that T/σf2gi are asymptotically negligible for i = 2, …, 5. These results are given in the following lemma.

Lemma C.3
  1. g2=Op(T-1/2(wTuwT)1/4+N-1/2+T-1),

  2. g3=Op(T-1/2(wTuwT)1/4+N-1/2+T-1).

  3. |g5| = Op(N−1/2 + T−1).

Proof

Using the facts that R = BF + E, = RF̂′/T and F̂F̂′/T = IK, we have

B-BH-1=BH-1(HF-F^)F^/T+E(F^-HF)/T+EFH/T.

Thus g2 = g21 + g22 + g23, where g21=wTBH-1(HF-F^)F^/TB^wT, g22=wTE(F^-HF)/TB^wT, and g23=wTEFH/TB^wT. It was shown by Fan et al. (2013) that ||H|| = Op(1) = ||H−1|. Thus by Lemma C.1, g21Op(1)HF-F^F^/T=Op(1/N+1/T2). In addition, g22Op(T)F^-HF/T=Op(1/N+1/T2). Finally, by Lemma B.2, g23wTEFOp(1/T)=Op(T-1/2(wTuwT)1/4). The proof for the convergence rate of |g3| is the same, so is omitted. Finally, the rate of convergence for |g5| follows from Lemma C.1.

Proof of Theorem 4.3

By Lemma C.3 and the assumption that Tσf2,σf2N/T and wTuwT=o(σf4), we have T/σf2gi=op(1) for i = 2, 3, 5. In addition, by Theorem 3.1 of Fan et al. (2013),

Ω^-u=Op(sN(logNT+1N)1/2-q/2),

which then implies that T/σf2g4=op(1). It then follows from Assumption 4.6 and Lemma B.7, T/σf2wT(^P,K^-)wTdN(0,1). By Theorem 3.2 of Fan et al. (2013), ^P,K^-max=Op(logNT+1N). Due to N/T → ∞ is assumed in the case of unknown factors,

TBf[w^T(^P,K^-)w^T-wT(^P,K^-)wT]TBf^P,K^-maxw^T-w1(w^1+w1)Op(TσflogNT)op(σf(logNT)-4/r1)+Op(TσfN)op(σf(logNT)-4/r1)=op(1). (C.3)

Therefore, [var(t=1T(wTBft)2)]-1/2Tw^T(^P,K^-)w^TdN(0,1), and

Tσf2w^T(^P,K-)w^TdN(0,1). (C.4)

Moreover, since σf2-σ^P2=op(σf2),

Tw^T(^P,K^-)w^Tσf-1-σ^P-1=Tw^T(^P,K^-)w^Tσ^P2-σf2σf2(1+op(1))=op(1).

It then follows from (C.4) that T/σ^P2w^T(^P,K-)w^TdN(0,1), which validates the H-CLUB.

C.3 Proof of Theorem 4.4

The theorem follows from the following lemma.

Lemma C.4

Suppose ft and ut are independent, and Eft = Eut = 0, then

var[t=1T(wTRt)2]=var[t=1T(wTBft)2]+var[t=1T(wTut)2]+var[2t=1TwTBftwTut].
Proof

Since Rt = Bft + ut, and ft and ut are independent,

var[t=1T(wTRt)2]=var[t=1T(wTBft)2+(wTut)2+2wTBftwTut]=var[t=1T(wTBft)2]+var[t=1T(wTut)2]+var[2t=1TwTBftwTut]+2cov[t=1T(wTBft)2+(wTut)2,2t=1TwTBftwTut].

It suffices to show the covariance term is zero. In fact, since EwTBftwTut=0,

cov[t=1T(wTBft)2,t=1TwTBftwTut]=sT,tTcov[(wTBfs)2,wTBftwTut]=s,tE(wTBfs)2wTBftwTut=s,tE(wTBfs)2wTBftEwTut=0.

Finally, as Eft = 0 implies EwTBft=0, we have

cov[t=1T(wTut)2,t=1TwTBftwTut]=sT,tTcov[(wTus)2,wTBftwTut]=s,tE(wTus)2wTBftwTut=s,tE(wTBft)E(wTus)2wTut=0.

Footnotes

References

  1. Ahn S, Horenstein A. Eigenvalue ratio test for the number of factors. Econometrica. 2013;81:1203–1227. [Google Scholar]
  2. Andrews D. Heteroskedasticity and autocorrelation consistent covariance matrix estimation. Econometrica. 1991;59:817–858. [Google Scholar]
  3. Antoine B. Portfolio Selection with estimation risk: a test-based approach. Journal of Financial Econometrics. 2011;10:164–197. [Google Scholar]
  4. Antoniadis A, Fan J. Regularized wavelet approximations. J Amer Statist Assoc. 2001;96:939–967. [Google Scholar]
  5. Bai J. Inferential theory for factor models of large dimensions. Econometrica. 2003;71:135–171. [Google Scholar]
  6. Bai J, Ng S. Determining the number of factors in approximate factor models. Econometrica. 2002;70:191–221. [Google Scholar]
  7. Bannouh K, Martens M, Oomen R, Dijk D. Realized mixed- frequency factor models for vast dimensional covariance estimation. 2012 Manuscript. [Google Scholar]
  8. Bianchi D, Carvalho C. Risk assessment in large portfolios: why imposing the wrong constraints hurts. 2011 manuscript. [Google Scholar]
  9. Bickel P, Levina E. Covariance regularization by thresholding. Ann Statist. 2008;36:2577–2604. [Google Scholar]
  10. Brandt M, Santa-Clara P, Valkanov R. Covariance regularization by parametric portfolio policie: exploiting characteristics in the cross-section of equity returns. The Review of Financial Studies. 2009;22:3411–3447. [Google Scholar]
  11. Brodie J, Daubechies I, Mol C, Giannone D, Loris I. Sparse and stable Markowitz portfolios. PNAS. 2009;106:12267–12272. doi: 10.1073/pnas.0904287106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Cai T, Liu W. Adaptive thresholding for sparse covariance matrix estimation. J Amer Statist Assoc. 2011;106:672–684. [Google Scholar]
  13. Chamberlain G, Rothschild M. Arbitrage, factor structure and mean-variance analyssi in large asset markets. Econometrica. 1983;51:1305–1324. [Google Scholar]
  14. Chang C, Tsay R. Estimation of covariance matrix via the sparse Cholesky factor with lasso. Journal of Statistical Planning and Inference. 2010;140:3858–3873. [Google Scholar]
  15. Chudik A, Pesaran H, Tosetti E. Weak and strong cross section dependence and estimation of large panels. The Econometrics Journal. 2011;14:C45–C90. [Google Scholar]
  16. Connor G, Korajczyk R. A Test for the number of factors in an approximate factor model. Journal of Finance. 1993;48:1263–1291. [Google Scholar]
  17. DeMiguel V, Garlappi L, Nogales F, Uppal R. A generalized approach to portfolio optimization: improving performance by constraining portfolio norms. Management Science. 2009a;55:798–812. [Google Scholar]
  18. DeMiguel V, Garlappi L, Nogales F, Uppal R. Optimal versus naive diversification: How inefficient is the 1/N portfolio strategy? Review of Financial Studies. 2009b;22:1915–1953. [Google Scholar]
  19. El Karoui N. High-dimensionality effects in the Markowitz problem and other quadratic programs with linear constraints: Risk underestimation. Annals of Statistics. 2010;38:3487–3566. [Google Scholar]
  20. Fama E, French K. The cross-section of expected stock returns. Journal of Finance. 1992;47:427–465. [Google Scholar]
  21. Fan J, Fan Y, Lv J. High dimensional covariance matrix estimation using a factor model. J Econometrics 2008 [Google Scholar]
  22. Fan J, Liao Y, Mincheva M. High dimensional covariance matrix estimation in approximate factor models. Ann Statist. 2011;39:3320–3356. doi: 10.1214/11-AOS944. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Fan J, Liao Y, Mincheva M. Large covariance estimation by thresholding principal orthogonal complements. (with discussion) Journal of the Royal Statistical Society Series B. 2013;75:603–680. doi: 10.1111/rssb.12016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Fan J, Zhang J, Yu K. Asset Allocation and Risk Assessment with Gross Exposure Constraints for Vast Portfolios. Journal of the American Statistical Association. 2012;107:592–606. doi: 10.1080/01621459.2012.682825. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Fan J, Yao Q. Nonlinear Time Series: Nonparametric and Parametric Methods. Springer-Verlag; New York: 2003. p. 576. [Google Scholar]
  26. Fryzlewicz P. High-dimensional volatility matrix estimation via wavelets and thresholding. Biometrika. 2012;100:921–938. [Google Scholar]
  27. Gagliardini P, Ossola E, Scaillet O. Time-varying risk premium in large cross-sectional equidity datasets. Swiss Finance Institute Occasional Paper Series. 2010:11–40. [Google Scholar]
  28. Gandy A, Veraart L. The effect of estimation in high-dimensional portfolios. Mathematical Finance. 2012;23:531–559. [Google Scholar]
  29. Gomez K, Gallon S. Comparison among high dimensional covariance matrix estimation methods. Revista Columbiana de Estadistica. 2011;34:567–588. [Google Scholar]
  30. Jacquier E, Polson N. Simulation-based-estimation in portfolio selection. 2010 Manuscript. [Google Scholar]
  31. Jagannathan R, Ma T. Risk Reduction in Large Portfolios: Why Imposing the Wrong Constraints Helps. Journal of Finance. 2003;54:1651–1683. [Google Scholar]
  32. Lai T, Xing H, Chen Z. Mean-variance portfolio optimization when means and covariances are unknown. Annals of Applied Statistics. 2011;5:798–823. [Google Scholar]
  33. Ledoit O, Wolf M. Improved estimation of the covariance matrix of stock returns with an application to portfolio selection. Journal of Empirical Finance. 2003;10:603–621. [Google Scholar]
  34. Merlevède F, Peligrad M, Rio E. A Bernstein type inequality and moderate deviations for weakly dependent sequences. Probab Theory Relat Fields. 2011;151:435–474. [Google Scholar]
  35. Newey W, West K. A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix. Econometrica. 1987;55:703–708. [Google Scholar]
  36. Newey W, West K. Automatic lag selection in covariance matrix estimation. Review of Economic Studies. 1994;61:631–654. [Google Scholar]
  37. Onatski A. Determining the number of factors from empirical distribution of eigenvalues. The Review of Economics and Statistics. 2010;92:1004–1016. [Google Scholar]
  38. Peligrad M. On the asymptotic normality of sequences of weak dependent random variables. Journal of Theoretical Probability. 1996;9:703–715. [Google Scholar]
  39. Pesaran MH, Zaffaroni P. Optimal asset allocation with factor models for large portfolios. Cesifo working paper. 2008:2326. [Google Scholar]
  40. Rothman A, Levina E, Zhu J. Generalized thresholding of large covariance matrices. J Amer Statist Assoc. 2009;104:177–186. [Google Scholar]
  41. Silverman B. Density estimation for statistics and data analysis. Chapman and Hall; London: 1986. [Google Scholar]
  42. Stock J, Watson M. Forecasting using principal components from a large number of predictors. J Amer Statist Assoc. 2002;97:1167–1179. [Google Scholar]
  43. Yen Y. Sparse weighted norm minimum variance portfolio. 2013 Manuscript. [Google Scholar]

RESOURCES