Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2018 Aug 3.
Published in final edited form as: J Econom Method. 2014 Mar 22;4(1):71–100. doi: 10.1515/jem-2012-0006

MULTIVARIATE FRACTIONAL REGRESSION ESTIMATION OF ECONOMETRIC SHARE MODELS

John Mullahy 1
PMCID: PMC6075740  NIHMSID: NIHMS978632  PMID: 30079291

Abstract

This paper describes and applies econometric strategies for estimating regression models of economic share data outcomes where the shares may take boundary values (zero and one) with nontrivial probability. The main focus of the paper is on the conditional mean structures of such data. The paper proposes an extension of the fractional regression methodology proposed by Papke and Wooldridge, 1996, 2008, in univariate cross-sectional and panel contexts. The paper discusses the stochastic aspects of share definition and measurement, and summarizes important features of the existing literature on econometric strategies for share model estimation. The paper then goes on to discuss the univariate fractional regression estimation strategies proposed by Papke and Wooldridge and to extend the fractional regression approach to estimation of and inference about regression models describing the multivariate share data. Some issues involving outcome aggregation/disaggregation are considered, as is a full likelihood estimation approach based on Dirichlet-multinomial models. The paper demonstrates the workings of these various empirical strategies by estimating models of financial asset portfolio shares using data from the 2001, 2004, and 2007 U.S. Surveys of Consumer Finances.

Prologue

Displayed in table 1a is an extract of nine observations on nine mutually exclusive and exhaustive categories of healthcare expenditures from a multiyear sample of the U.S. Medical Expenditure Panel Survey (MEPS); the specific quantities reported are the shares of total healthcare expenditures contributed by each of the nine expenditure categories. Table 1b exhibits an extract of nine observations on seventeen mutually exclusive and exhaustive two-digit categories of time use from a multiyear sample of the American Time Use Survey (ATUS); the detailed figures are the number of minutes reported being spent in each category of time use during the one-day, or 1,440-minute, time diary recall period. Finally, table 1c displays an extract of nine observations on ten mutually exclusive and exhaustive categories of financial assets from a multiyear sample of the Survey of Consumer Finances (SCF); the detailed data presented in this table are the shares of total financial assets accounted for by each of the ten financial asset categories.

Table 1.

a MEPS, Combined 1996–2007 Sample, Data Extract: Healthcare Expenditure Shares, Nine Consecutive Observations with Defined Shares (dupersid is MEPS Individual Case Identifier, with Data Presented in dupersid Sort Order; Boundary Solutions Shaded)
Observation Healthcare Expenditure Category Total
Office-
Based
Visits
Prescr.
Drugs
Inpatient
Stays
ER Visits Out-
patient
Visits
Dental
Visits
Home
Health
Care
Vision
Aids
Other
S & E
Year dupersid
1997 00014013 0 0 0 0 0 1 0 0 0 1.000
1997 00015011 1 0 0 0 0 0 0 0 0 1.000
1997 00015015 0 0 0 0 0 0 0 1 0 1.000
1997 00015022 0.045 0.126 0 0.828 0 0 0 0 0 1.000
1997 00018036 0 0 0.903 0.097 0 0 0 0 0 1.000
1997 00018059 0 0 0 1 0 0 0 0 0 1.000
1997 00018073 0 0.001 0.952 0.006 0.041 0 0 0 0 1.000
1997 00019014 0.108 0 0 0 0.481 0.411 0 0 0 1.000
1997 00020011 0 1 0 0 0 0 0 0 0 1.000
b American Time Use Survey, Combined 2003–2008 Sample, Data Extract: Two-Digit Time Use Categories, Nine Consecutive Observations (tucaseid is ATUS Individual Case Identifier, with Data Presented in tucaseid Sort Order; Boundary Solutions Shaded)
Observation Two-Digit Time Use Category Total
Personal
Care
Activities
Household
Activities
Caring
for & Helping
HH Members
Caring
for & Helping
Non-HH Members
Work &
Work-Related
Activities
Education Consumer
Purchases
Professional &
Personal
Care Svcs.
Household
Services
Govt. Svcs.
& Civic
Obligations
Eating &
Drinking
Socializing,
Relaxing,
& Leisure
Sports,
Exercise, &
Recreation
Religious & Spiritual
Activities
Volunteer
Activities
Telephone
Calls
Travelling
Year tucaseid
2003 20030807033472 540 240 0 0 0 0 0 0 0 0 45 540 0 0 0 75 0 1440
2003 20030807033485 615 0 10 0 540 0 0 0 0 0 60 140 0 0 0 0 75 1440
2003 20030807033487 480 60 0 0 45 0 0 0 0 0 120 320 290 0 0 0 125 1440
2003 20030807033489 1440 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1440
2003 20030807033490 705 90 0 0 0 0 0 0 0 0 0 450 35 0 0 150 10 1440
2003 20030807033491 600 0 0 477 0 0 0 0 0 0 0 60 0 201 0 0 102 1440
2003 20030807033494 720 0 0 470 0 0 0 0 0 0 10 240 0 0 0 0 0 1440
2003 20030807033495 570 270 0 0 0 0 0 0 0 0 90 380 0 0 0 70 60 1440
2003 20030807033498 705 30 0 0 180 0 0 0 0 0 20 315 0 0 0 0 190 1440
c Survey of Consumer Finances, Combined 2001, 2004, 2007 Sample, Data Extract: Financial Asset Shares, Nine Consecutive Observations with Defined Shares (yy1 is SCF Household Case Identifier, with Data Presented in yy1 Sort Order; Boundary Solutions Shaded)
Observation Financial Asset Category Total
Liquid
Assets
Quasi-
Liquid
Retir.
CDs Dir.
Held
Pooled
Funds
Savings
Bonds
Directly
Held
Stocks
Directly
Held
Bonds
Cash
Val.
Whole
Life
Other
Mgd.
Assets
Other
Fin.
Assets
Year yyi
2007 1531 1 0 0 0 0 0 0 0 0 0 1.000
2007 1532 .167 .131 0 0 0 0 0 .702 0 0 1.000
2007 1533 .146 .854 0 0 0 0 0 0 0 0 1.000
2007 1534 0 1 0 0 0 0 0 0 0 0 1.000
2007 1535 .255 .531 .204 0 .010 0 0 0 0 0 1.000
2007 1536 .049 .904 .014 0 0 .007 0 .026 0 0 1.000
2007 1537 0 0 0 0 0 0 0 1 0 0 1.000
2007 1538 .359 .513 0 0 0 .128 0 0 0 0 1.000
2007 1539 1 0 0 0 0 0 0 0 0 0 1.000

The samples from which these observations are extracted have structures that share two analytically important features. First, the multivariate outcomes are mutually exclusive and exhaustive shares of some total. Second, there is a nontrivial empirical frequency of shares that are realized at upper and lower boundary values.

1. Introduction

Multivariate outcomes measured as shares of some overall total arise in numerous contexts in applied microeconometrics. Whether the particular analysis focuses on time use (Mullahy and Robert, 2010), portfolio shares (e.g. Poterba and Samwick, 2001, 2002), consumer budgeting (see the references in section 3), market share analysis (Berry et al., 1995; Dubin, 2007), or some other topic, there will often arise -- as demonstrated by the above examples -- commonalities of the data structures under investigation with those exhibited in tables 1a–1c.

Letting yik represent the k-th outcome for the i-th individual, M denote the number of outcomes, and xi, i=1,…,N, be a p-vector of exogenous covariates, such data are characterized formally by the following:

yik[0,Bi], (1)
Pr(yik=0|xi)>0  and Pr (yik=Bi|xi)>0, (2)

and

m=1Myim=Bi  for all i, (3)

or in vector notation yi ∈ [0, Bi]M and 1'yi = Bi. Here Bi represents some finite total or upper bound: Bi=total annual healthcare spending in table 1a; Bi=B=1,440 minutes per day in the time use data in table 1b; and Bi=total financial assets in the data underlying the financial asset share data in table 1c. The Bi may or may not vary across i and may or may not be exogenous (e.g. consider healthcare spending or financial assets), although apart from the discussion in section 2 the discussion here proceeds as if the Bi are exogenous or can reasonably be conditioned on (more on this below).1 Considerations (1) and (3) are standard in the econometric share equation literature. Econometric strategies to handle of (2) in light of (1) and (3) are less studied and are the main focus of this paper.

This paper describes and applies econometric strategies for estimating regression models of various features of outcome data like those described above, with a main focus on conditional means. Specifically, for the analysis of the conditional mean structures of such data, this paper proposes an extension of the fractional regression methodology proposed by Papke and Wooldridge in univariate cross-sectional (Papke and Wooldridge, 1996) and panel (Papke and Wooldridge, 2008) contexts.2

The emphasis on conditional first-moment structures is central. The working premise is that the parameters of concern to the analyst are the set of conditional means E[yk | x], k=1,…,M, which are specified, with B treated as exogenous, to satisfy

E[yk|x](0,B),k=1,,M (4)

and

m=1ME[ym|x]=B. (5)

Note that in (4) the conditional means are assumed to span the open interval (0,B) rather than the closed interval [0,B] which the yk can occupy. While probably a reasonable assumption in general, this could be restrictive in some instances where for some values of x Pr (yk = 0|x) = 1 or Pr (yk = B|x) = 1 might be possible.3 The subsequent analysis accommodates either of these boundary probabilities being arbitrarily close to, but not identically, zero or one.

Henceforth the paper will work with the normalized outcomes or shares sk=yk/B instead of the yk themselves, with the vector of shares s satisfying s ∈ [0,1]M and 1's = 1. Moreover, the analysis will proceed under the assumption that the E[sk|x] have a parametric structure, i.e. E[sk|x] = ξk (x; α), where the generic common parameter vector α = [α1,…, αM] will generally be shared across the M conditional mean parameters ξk (x; α) to enforce the adding-up condition (5).

The plan for the remainder of the paper is as follows. Section 2 discusses the stochastic aspects of share definition and measurement. Section 3 summarizes salient features of the existing literature on econometric strategies for share model estimation. Section 4 highlights various features of the fractional regression estimation strategies proposed by Papke and Wooldridge (1996, 2008). Sections 5–8 are the methodological core of the paper: section 5 extends the fractional regression approach to the multivariate share model context; section 6 considers several issues involving inference and specification testing; section 7 offers some ideas on testing aggregation or disaggregation of outcome categories; and section 8 presents a Dirichlet-multinomial likelihood-based approach to estimating multivariate share models that accommodates important features of the observed outcomes. Section 9 presents an empirical example of the proposed methodologies using data on financial asset portfolio shares from the 2001, 2004, and 2007 U.S. Surveys of Consumer Finances. Section 10 concludes.

2. Share Definition and Stochastic Characteristics

The foundation of the empirical analysis is the joint distribution ϕ (y, x) of an M-vector of outcomes y0 and exogenous covariates x; extensions to situations involving endogenous x would be interesting to consider but are beyond this paper’s scope. From this, share measures may arise naturally in at least two ways that may imply different stochastic structures for the resulting econometric share models.

Suppose there are M quantities yk = gk (x, αk) + uk, k=1,…,M, that arise from some constrained optimization problem (utility maximization; cost minimization; portfolio composition optimization; etc.), where E[uk|x] = 0, k=1,…,M, so that the gk(․) are conditional means. The corresponding shares are given by

sk=ykm=1Mym=g(x,αk)+ukm=1Mg(x,αm)+um=g(x,αk)+ukY. (6)

Consider two alternative situations. In some cases (e.g. time use) there is a nonstochastic exogenous constraint B (e.g. 1,440 minutes per day) such that m=1Mym=Y=B. If the nonstochastic adding up restriction m=1Mg(x,αm)=B is required or enforced, then m=1Mum=0 and

sk=g(x,αk)+ukm=1Mg(x,αm)=g(x,αk)m=1Mg(x,αm)+vk=ξk(x,α)+vk, (7)

where the vk are conditionally mean-zero heteroskedastic errors so that E[sk|x] = ξk (x, α).

Alternatively, the total Y may simply be defined as the sum of the M stochastic quantities yk whose measurements share a common metric (e.g. currency units) without being subject to any analytically relevant exogenous constraint,4 in which case the share equations have the more general form sk = h(x, αk, uk)/H(x, α, u). In this instance, stochastic elements appear in both numerator and denominator of the share functions. As such, the derivation of E[sk|x] is no longer straightforward, requiring integration over the joint distribution of the entire vector of residuals u.

By making primitive first-moment assumptions along the lines of E[sk|x] = ξk (x, α), this paper proceeds for the most part under the assumption that the simpler structure (7) holds, although conceiving of E[sk|x] = ξk (x, α) as a first-order approximation via an expansion of h(x, αk, uk)/H(x, α, u) around u = 0 may also be reasonable.

3. Approaches to Econometric Share Model Estimation

This section provides a brief survey of approaches to econometric share model estimation that have been prominent in the literature.

Econometric Share Model Estimation

Much but not all of the econometric share equation literature focuses on the relationship between empirical share models and underlying constrained optimization behaviors yielding outcomes (e.g. commodity category expenditures or patterns of time use) that are shares of some particular total (e.g. money or time budgets). Early contributions to this literature include Christensen et al., 1975, and Wales and Woodland, 1977, who examine consumer demands and corresponding expenditure shares in utility maximization contexts. Subsequent studies have approached share equation estimation from theoretically motivated optimization models in which stochastic components are embedded to play particular roles (preference heterogeneity; technical or allocative inefficiency; etc.) in the optimization framework rather than being appended additively to nonstochastic share functions in what might be an ad hoc manner; such examples include Brown and Walker, 1989 and 1995, Chavas and Segerson, 1987, Kooreman and Kapteyn, 1987, and McElroy, 1987. Considine and Mount, 1984, derive a specification in which the set of share equations has a multinomial logit functional form; this is noteworthy because a multinomial logit form is at the core of the specifications proposed below. Dubin, 2007, uses nested multinomial logit market share models to estimate valuations of intangible assets. Fry et al., 1996, based in part on ideas developed in Aitchison, 1982, apply to the estimation of share models methods from compositional data analysis (CODA), which involve essentially modeling logs of ratios of shares.5 See also Theil, 1969, for an early related contribution.

Estimation of Share Models with Boundary or Corner Solutions

Fewer studies have tackled the thorny empirical problems that arise when observed shares take on corner or boundary solutions with nontrivial probabilities.6 Appealing to Kuhn-Tucker conditions and corresponding virtual or support prices, Lee and Pitt, 1986, propose a general empirical structure for multivariate share systems when corner solutions at zero are prominent in the data. Morey et al., 1995, explore a variety of statistical models to accommodate boundary outcomes, among these multinomial models for discretely measured (count) outcomes that share some features with the multivariate fractional regression models that are proposed below.

Since boundary solutions are a prominent feature of the share data of concern here, providing a general strategy for analyzing such outcomes is of some interest. Ad hoc fixes are not an appealing approach to the boundary solution phenomenon, particularly when the probability of such boundary outcomes is nontrivial.

Dirichlet Share Models

Woodland, 1979, proposed the Dirichlet distribution as a direct statistical model for shares without particular consideration of any underlying economic optimization framework. The Dirichlet density conditional on x is given by7

D(s|x;ψ)=(Γ(m=1Mzm(x,ψ))m=1MΓ(zm(x,ψ)))m=1Msm(zm(x,ψ)1), (8)

where zm(․) is a possibly nonlinear function of covariates and parameters. (Note that if any sk=0 then the density D(․)=0 and if any sk=1 then necessarily all the other sk=0. As such, the density is undefined in either event, thus precluding direct application of the Dirichlet model to the kinds of share data examined here where boundary values are prominent.

Yet it is useful for purposes of this paper to note that for the Dirichlet model, the conditional first moments are

E[sk|x]=zk(x,ψ)m=1Mzm(x,ψ),k=1,,M (9)

or

E[sk|x]=exp(xψk)m=1Mexp(xψm)=exp(x(ψkψM))1+m=1M1exp(x(ψmψM)),k=1,,M, (10)

using a natural exponential-with-linear-index parameterization for the z(․). Note that this corresponds to the standard functional form for multinomial logit probabilities even though for the Dirichlet model all M of the ψm are identified. Noteworthy for present purposes is that this conditional first-moment structure coincides with that of the multivariate fractional share model whose specification and estimation are discussed below.

4. Fractional Regression Estimation

Before exploring the multivariate estimation strategies that are the main focus of this paper, a brief overview of fractional regression (FREG) methods is worthwhile.8 The FREG model was proposed initially by Papke and Wooldridge (PW), 1996, in their study of voluntary individual contributions to retirement accounts in which the univariate dependent variable of interest is the fraction s ∈ [0,1] of allowable contributions made by individuals in their sample. The key result in the PW paper is that even when the outcomes take on values at the extremes of the bounded range they occupy -- i.e. s=0 or s=1 -- with nonzero probability, the FREG method provides consistent estimates of the parameters of a univariate conditional mean function so long as it is specified with the correct functional form and embedded in a suitable quasi-ML estimator or M-estimator. Specifically, PW, 1996, consider the case of univariate fractional outcome data (s) and conditional means E[s | x] with M=1, while PW, 2008, consider the panel data context with the added t-dimension, t=1,…,tmax, giving outcomes for the i-th individual at the t-th time period sitm=sit that are multivariate (tmax>1) in the t-dimension but univariate (M=1) in the m-dimension. In particular, in this case there are no implied adding-up restrictions of the form (5) to accommodate.

The basic idea underlying FREG estimation in the cross-sectional univariate outcome context is that if focus is exclusively on conditional first moments, then quasi-ML estimation when a correct parametric specification of the conditional first-moment structure is embedded in an exponential-family quasi-likelihood will yield consistent estimates of the first-moment parameters regardless of whether the nominal quasi-likelihood is true or not (Gourieroux et al., 1984a, b). In light of these arguments, PW suggest that the key property of a conditional first-moment model for a fractional outcome is that it obeys the boundary restrictions, i.e. E[s|x] = ξ (x) ∈ (0,1). PW then suggest that a general class of parametric functional forms that satisfy this restriction are distribution functions G(․) of continuous random variables, i.e. E[s|x] = ξ (x; ω) = G(x; ω) ∈ (0,1).

Thus direct specification of such a conditional mean structure embedded in an exponential-family quasi-likelihood whose maximization estimates such a distribution function should provide consistent estimates of ω so long as G(x; ω) is a correct specification of the conditional first moment. Bernoulli quasi-likelihoods G(x; ω)s × (1 − G(x; ω)1−s for the fractional (not binary) outcome measures s ∈ [0,1] are the obvious choice, with the particular functional form for G(x; ω) specified as a logit, probit, or other cumulative distribution function, typically with a linear index argument . Consistent inferences are straightforward, but will generally involve using robust sandwich or bootstrap covariance estimators since the share data will be underdispersed relative to the nominal Bernoulli model (see section 6).

The fractional logit ("FLOGIT") version of the FREG model with

E[s|x]=G(x;ω)=exp(xω)(1+exp(xω)) (11)

is the univariate foundation for the multivariate FREG estimator discussed now.

5. Multivariate Fractional Logit: Estimation

The central goal of this paper is to provide consistent estimation strategies to estimate properties of the conditional distribution of share data that enforce (12) and (13) and accommodate (14) and (15):

E[sk|x]=ξk(x;β)(0,1),k=1,,M, (12)
m=1ME[sm|x]=1, (13)
Pr(sk=0|x)0,k=1,,M, (14)
Pr(sk=1|x)0,k=1,,M, (15)

where β [β1, β2,…,βM]. The main concern in this section and the next is with estimation of the conditional first-moment structure of such data, i.e. ξ (x; β). Section 8 extends this inquiry to other features of the joint conditional probability models ϕ (s|x).

This and the following three sections offer a detailed exposition of the multivariate fractional logit ("MFLOGIT") estimator and its properties.9 PW and Gourieroux et al., 1984a,b provide the fundamental arguments to establish consistency for the multivariate/multinomial version of the univariate PW quasi-ML approach. Assume that the sample are independent draws from the (M+C)-variate distribution ϕ (s, x). Based on (10), specify the M conditional means to have a multinomial logit functional form in linear indexes as

E[sk|x]=ξk(x;β)=exp(xβk)m=1Mexp(xβm),k=1,,M. (16)

Note that this specification enforces both (12) and (13). Alternative specifications, e.g. ξk(x;φ)=xφk/m=1Mxφm, are estimable -- indeed, this is the conditional first-moment functional form implied in the Dirichlet simulations conducted by Woodland, 1979 -- but they admit the possibility of predicted shares falling outside the [0,1] interval at some values of x. This paper thus focuses on the specification (16) although the merits of competing first-moment specifications could be adjudicated empirically by conditional-moment or related tests.

As with the familiar multinomial logit estimator, some normalization is required since all M of the βk will not be separately identified in the multinomial quasi-likelihood; βM = 0 is used henceforth, giving

ξk(x;β)=exp(xβk)1+m=1M1exp(xβm),k=1,,M1 (17)

and

ξM(x;β)=11+m=1M1exp(xiβm). (18)

Owing in part to the normalization, interpretation of the signs and magnitudes of the βk is generally not straightforward.10 Typically much more interesting and useful in applications are the corresponding average partial effects (APEs) that are invariant with respect to the particular normalization selected. These are described in detail in Appendix 1.

Appealing to the quasi-ML estimation methods described by PW for the univariate case, one can define a multinomial logit quasi-likelihood function Q(․) that embeds the functional form (16) and that uses the observed shares sik ∈ [0,1] in place of the binary indicators that would typically populate a multinomial logit likelihood function, i.e.

Q(β)=i=1Nm=1Mξm(xi;β)sim. (19)

The log quasi-likelihood is

J(β)=log(Q(β))=i=1Nm=1Msim×log(ξm(xi;β)), (20)

with the corresponding p × (M − 1) estimating or score equations

J(β)βk=i=1Nxi'[sik(exp(xiβk)1+m=1M1exp(xiβm))],k=1,,M1, (21)

which are the same score equations as those corresponding to a standard multinomial logit estimator except that the sik are, in general, nonbinary.11 Consistency of the resulting β̂ follows from the arguments in PW and Gourieroux et al., 1984a, in particular that E[∂J(β)/∂βk] = ExE[∂J(β)/∂βk|x] = 0, k=1,…,M−1, given standard full-rank assumptions that ensure a unique maximizer.12

6. Multivariate Fractional Logit: Underdispersion, Inference, and Specification Testing

Underdispersion

Overdispersion (resp. underdispersion) is typically characterized in a univariate outcome context as a situation where the empirical variance of the distribution of some outcome y is greater than (resp. less than) the variance that would obtain if y followed a reference or nominal distribution ϕnom, possibly conditioned on covariates, i.e. Varemp (y|x) > Varnom (y|x) (resp. Varemp (y|x) < Varnom (y|x)), possibly enforcing the restriction Eemp [y|x] = Enom [y|x]. In the multivariate outcome context with outcome y an M×1 vector, a natural extension is to define overdispersion (resp. strict overdispersion) as the situation where the matrix difference Covemp (y|x) − Covnom (y|x) is positive semidefinite (resp. positive definite), and underdispersion (resp. strict underdispersion) as the situation where the Covnom (y|x) − Covemp (y|x) is positive semidefinite (resp. positive definite).

The nature of the multivariate share data analyzed here is such that the data necessarily manifest (conditional) underdispersion relative to their nominal multinomial counterparts in this matrix-definiteness sense. This can be seen as follows. Given the quasi-ML first-moment assumptions and multinomial nominal likelihood, it follows that Prnom (sk = 1|x) = Enom [sk|x] = Eemp [sk|x] = ξk (x). Thus for the share vector s, Covnom (s|x) is given by the multinomial covariance matrix

Covnom(s|x)=[ξ1(x)(1ξ1(x))ξ1(x)ξ2(x)ξ1(x)ξM(x)ξ1(x)ξ2(x)ξ2(x)(1ξ2(x))ξ1(x)ξM(x)ξM(x)(1ξM(x))], (22)

while Covemp (s|x) is given by

Covemp(s|x)=[E[(s1ξ1(x))2|x]E[(s1ξ1(x))(s2ξ2(x))|x]E[(s1ξ1(x))(sMξM(x))|x]E[(s1ξ1(x))(s2ξ2(x))|x]E[(s2ξ2(x))2|x]E[(s1ξ1(x))(sMξM(x))|x]E[(sMξM(x))2|x]]. (23)

Thus

Δ(x)=Covnom(s|x)Covemp(s|x)=[ξ1(x)E[s12|x]E[s1s2|x]E[s1sM|x]E[s1s2|x]ξ2(x)E[s22|x]E[s1sM|x]ξM(x)E[sM2|x]]. (24)

If Pr(sk ∈ (0,1)|x) > 0 for all k then each of the diagonal elements of Δ(x) is positive (since z > z2 for any z ∈ (0,1)). Note too that m=1mkME[sksm|x]=E[(skm=1mkMsm)|x]=E[sk(1sk)|x], so that the absolute value of the row sum of the off-diagonal elements of any row in Δ(x) equals the diagonal element in that row.

Definition

Asymmetric M×M matrix D is weakly diagonally dominant if Dkkm=1mkM|Dkm| for all k=1,…,M, and is diagonally dominant if the inequality holds strictly.

It follows that the matrix Δ(x) is weakly diagonally dominant. Furthermore, a diagonally dominant (resp. weakly diagonally dominant) matrix is positive definite (resp. positive semidefinite).13 As such, the empirical distribution of the share vector s conditional on x manifests underdispersion relative to the nominal multinomial quasi-likelihood. The implications of this for inference are considered below.

Inference

The asymptotic distribution of β̂ follows from the arguments in PW and in Gourieroux et al., 1984a,b. Specifically, given correct specification of the conditional first moments ξ(x) = E[s|x], N(β^β) is asymptotically N(0, VMFLOGIT) where

VMFLOGIT=AP1AQAp1 (25)

with

AP=Ex[(IM1x)'P(x)(IM1x)], (26)
AQ=Ex[(IM1x)'Q(x)(IM1x)], (27)
P(x)=[ξ1(x)(1ξ1(x))ξ1(x)ξ2(x)ξ1(x)ξ(M1)(x)ξ1(x)ξ2(x)ξ2(x)(1ξ2(x))ξ1(x)ξ(M1)(x)ξ(M1)(x)(1ξ(M1)(x))], (28)

and

Q(x)=[Var (s1|x)Cov (s1,s2|x)Cov (s1,s(M1)|x)Cov (s1,s2|x)Var (s2|x)Cov (s1,s(M1)|x)Var (s(M1)|x)]. (29)

Although estimating the model using quasi-ML methods will provide consistent estimates of the βk parameters, the corresponding inverse Hessian MNL covariance matrix will not be a consistent estimator of the true covariance matrix so long as Pr(sk ∈ (0,1)|x) > 0. Note that if the data were truly conditionally multinomially distributed, then N(β^β) would have asymptotic covariance matrix VMNL=AP1APAP1=AP1. The difference between VMNL and VMFLOGIT is

VMNLVMFLOGIT=AP1(APAQ)AP1. (30)

In turn, the matrix difference APAQ can be written as

APAQ=Ex[(IM1x)'(P(x)Q(x))(IM1x)]. (31)

Note that the matrix difference P(x) − Q(x) equals the matrix Δ(x) defined in eq. (24) with the M-th row and M-th column deleted. As such P(x) − Q(x) will in general be strictly diagonally dominant and, therefore, positive definite. Being quadratics in positive definite matrixes, it thus follows that APAQ and, therefore, AP1(APAQ)AP1, will themselves be positive definite. As such, VMNLVMFLOGIT will in general be positive definite. As will be seen below in section 9, inverse Hessian estimates of Cov(β̂) based on the standard multinomial logit quasi-likelihood yield estimated parameter t-statistics for the individual β̂mk that range in this application from about 1.1 to 2.4 times smaller than those obtained using a robust sandwich estimator.

Specification Testing

To assess the quality of the MFLOGIT first-moment specification fit, conditional moment tests can be conducted that are based on cross-products of a vector of functions of x and the estimated MFLOGIT residuals, Λ(x) × (s − Ê[s|x]).

The power of such tests to detect poorness of fit depends on the specification of Λ(x). The particular form of Λ(x) used here is suggested by the Hosmer-Lemeshow test strategy used commonly in the evaluation of binary logit models, i.e. Λ(x) is specified as a vector of indicator functions based on L sample quantiles of the Ê[sim|xi], i.e.

λmq=N1i=1N1(ξm(xi;β^)Jmq)×(simξm(xi;β^)),q=1,,L (32)

where the Jmq denote intervals on the real line defined by the sample quantiles of the Ê[sim|xi].

7. Multivariate Fractional Logit: Aggregation and Disaggregation of Outcome Categories

In empirical contexts where MFLOGIT-type estimation strategies might be applied it may sometimes be of interest to determine whether subsets of the outcome measures sk might sensibly be aggregated or pooled (to reduce dimensionality) or, if such data are available, disaggregated to (to refine detail). There are at least two fundamentally distinct circumstances in which aggregation considerations may arise in the analysis of share data. The first ("structured aggregation") occurs when the outcome data follow some natural and/or predefined tree structure. At each level of such a tree the share outcome categories at that level -- which are outcome subcategories by reference to the next-higher level -- are mutually exclusive and exhaustive and sum to one.14 The considerations raised above notwithstanding, testing for aggregation with such tree structures can more or less proceed using classical testing approaches. The second circumstance ("unstructured aggregation") arises when there are again mutually exclusive and exhaustive subcategory share outcomes that sum to one but in which there is either no natural and/or predefined tree structure or in which the analyst for some reason elects to ignore a given tree structure (e.g. to consider whether subcategories in different branches can be pooled with each other). In these circumstances, alternative testing strategies will generally be required. Prototypical examples of each type of aggregation are displayed in figure 1.15

Figure 1.

Figure 1

Structured and Unstructured Aggregation: Examples

Depending on the purpose of the analysis, aggregability of outcome categories could be characterized in a variety of ways that include: similarity of corresponding category-parameter vectors; similarity of corresponding category-partial effects; and others. Aggregation characterized by similarity of parameter vectors appears in this context most straightforward to assess; as such, given the MFLOGIT's first-moment structure, aggregation and disaggregation of outcome measures are tantamount to summation and proper subsetting, respectively. Likelihood-based strategies for testing for aggregation or disaggregation of categories in multinomial or nested logit models with discrete outcomes (e.g. likelihood-ratio tests) are well established in the multinomial logit literature (Cramer and Ridder, 1991; Hill, 1983). These approaches are not directly applicable in the MFLOGIT's first-moment/quasi-likelihood context, however.

Appendix 2 discusses several strategies based on robust Wald tests (a "bottom up" approach based on estimation of models for disaggregated outcomes) as opposed to Lagrange multiplier tests (a "top down" approach that would be more challenging to implement in the quasi-likelihood context) or tests based on criteria like mean squared error reduction. The merits of the approaches suggested here relative to such alternatives is for future research to assess.

8. Dirichlet-Multinomial Estimation

Given the discrepancy between the empirical and QMLE pseudo-multinomial second moment structures, it should in principle be possible to improve estimator efficiency if reasonable conditional second-moment assumptions can be made (Gourieroux et al., 1984a). While the share structure of the data provides some guidance on such specifications (e.g. that, like the first moments, the second moments must themselves be bounded) there appears to be little additional general guidance about second-moment specification offered by the data themselves.

A more-structured alternative approach is to postulate a probability model for the data that can describe the important features of the data recognizing, of course, that there is an inconsistency or robustness (relative to the MFLOGIT estimator) cost that may be incurred if such a probability model is incorrectly specified. Of course, circumstances may arise when the entire conditional probability structure of the multivariate outcomes is of interest, in which case the first-moment estimates offered by MFLOGIT will not be adequately informative.

One working probability model that exhibits underdispersion relative to a multinomial structure and that also accommodates positive probability mass16 for shares sk=0 and sk=1 is based on a Dirichlet mixture of multinomials ("DM") or multivariate negative hypergeometric (Johnson et al., 1997, pp. 80ff), which is the multivariate version of the beta-binomial distribution (Heckman and Willis, 1977).17 Imagine that some underlying multivariate T-trial counts n = [nk], k=1,…,M, follow a conditional DM probability model

DM(n|x;T)=Γ(T+1)Γ(m=1Mam(x))m=1MΓ(nm+am(x))Γ(T+m=1Mam(x))m=1M{Γ(nm)Γ(am(x))},n{0,1,,T}M,1'n=T, (33)

with ak (x) = exp(k) a natural parameterization. Letting A(x)=m=1Mam(x), the conditional marginal mean and marginal variance of the nk are given by E[nk|x;T] = Tak (x)/A(x) and

Var(nk|x;T)=T(T+A(x)1+A(x))(ak(x)A(x))(1A(x)), (34)

where the latter expression is seen to be (T + A(x))/(1 + A(x)) times the conditional marginal variance of the underlying T-trial multinomial distribution, taking ak (x)/A(x) as the multinomial probabilities.

Given these counts, the shares sk ∈ [0,1] -- or, more precisely, sk ∈ {0,1/T,2/T,…,1} -- are given by sk = nk/T, so that, in particular, E[sk|x] = ak (x)/A(x) and

Var(sk|x;T)=1T(T+A(x)1+A(x))(ak(x)A(x))(1A(x)). (35)

Since (T + A(x))/(T(1 + A(x))) is less than one for T>1, Var(sk|x;T) is smaller than the conditional variance of a one-trial multinomial distribution having a corresponding conditional first-moment or probability structure, which in turn is the quasi-likelihood model for the MFLOGIT. It is also easily shown that Var(sk|x;T) is decreasing in T.

Since T does not vanish from the probability model for or the conditional variance functions of the sj, the application of the DM(․) model in cases where T does not have a natural interpretation is clearly as an approximation to the true probability model. In some instances, a particular specification for T might be suggested naturally by the nature of the data's measures (e.g. 1,440 integer-measured minutes in a day or some integer-measured number of currency units in a budget). In other instances, however, specifying a value for T will be ad hoc. Moreover, when the measures of the observed share data do not follow a natural lattice structure (i.e. sk ∈ {0,1/r,2/r,…,r/r}) but instead are "continuously" measured, some coarsening of the data that maps sk into sjc will be required for nkc=Tskc to satisfy nkc{0,1,,T} as necessary for the DM distribution.

Two DM models are estimated here for comparison with the MFLOGIT results discussed above. For present purposes, two coarsenings of the data that imply different values for T are considered. For T=10 and T=100, these are given by nkTc=floor(Tsj) (i.e. ⌊Tsj⌋) for all shares except the share with the largest empirical marginal frequency while the coarsened measure corresponding to that share (k) is given by njkc=Tm=1mkMnmTc to ensure proper adding up. The resulting parameter and APE estimates are then compared with each other and with those obtained from the MFLOGIT estimator. Given the specification E[sk|x]=ak(x)/A(x)=exp(xζk)/m=1Mexp(xζm), and normalizing post-estimation on ζM (since all M ζM's are identified, unlike the MFLOGIT QMLE), some comparability of the point estimates of ζM and βk and, in particular, of the corresponding APE estimates might be expected if the DM likelihood is a reasonable approximation to the true probability model.

9. Modeling Financial Asset Portfolio Shares

Data and Estimation Sample

This section demonstrates some of the properties of the share model estimators and tests described above by estimating regression models of financial asset portfolio shares. Estimation of portfolio share models has been considered by Heaton and Lucas, 2000, and by Poterba and Samwick, 2001, 2002, among others. The data used are from the combined public use 2001, 2004, and 2007 U.S. Surveys of Consumer Finances (SCF). The SCF is a triennial sample whose collection is sponsored by the U.S. Federal Reserve to provide information on the financial circumstances of U.S. families (for details on the SCF, see Bucks et al, 2009). This combined sample comprises 13,379 household-level observations. Household-level sampling weights provided in the SCF are used in the computation of the APE estimates (note that the SCF weights are designed for within-year but not necessarily across-year weighting). This analysis focuses specifically on financial assets and the ten major subcategories of financial assets defined in the SCF listed in the top panel of table 2. Additional details on the data are provided in Appendix 3.

Table 2.

Survey of Consumer Finances, Combined 2001, 2004, 2007 Sample: Descriptive Statistics

Financial Asset Aggregate Category Shares,
Estimation Sample (N=12,723)
Mean Sample Percentages (Unweighted)
Unwtd. Weighted sim = 0 sim = 1 0 <sim< 1
Liquid Assets .332 .399 .018 .170 .812
Quasi-Liquid Retirement Accounts .292 .312 .355 .005 .640
CDs .039 .044 .823 .001 .176
Directly Held Pooled Investment Funds .074 .047 .745 .0002 .255
Savings Bonds .010 .013 .820 .001 .180
Directly Held Stocks .097 .050 .654 .001 .346
Directly Held Bonds .022 .004 .907 0 .093
Cash Value of Whole Life Insurance .065 .073 .669 .004 .326
Other Managed Assets .037 .026 .885 .0001 .115
Other Miscellaneous Financial Assets .032 .031 .258 .004 .738
Covariates,
Full and Estimation
Samples
Full Sample
(N=13,379)
Subsample with Defined Shares
(N=12,723)
Mean Min Max Mean Min Max
Unwtd. Wtd. Unwtd. Wtd.
Age 50.9 49.5 18 95 51.4 50.0 18 95
White .78 .73 0 1 .80 .75 0 1
Married .67 .59 0 1 .68 .60 1 1
Number of Kids .85 .82 0 10 .84 .80 0 10
High School Graduate .25 .32 0 1 .25 .32 0 1
Some College .16 .18 0 1 .16 .19 0 1
College Graduate .48 .35 0 1 .50 .37 0 1
Year 2004 .34 -- 0 1 .34 -- 0 1
Year 2007 .33 -- 0 1 .33 -- 0 1

These data are summarized in figure 2 and in the top panel of table 2. In the sample of 12,723 household observations with defined shares, it might be noted that only six households (.047%) have strictly positive shares for all ten financial asset categories, whereas for 2,348 (18.5%) of the sample's households some financial asset category share is 1.0. Covariates used in the analysis are: age in years of household head (Age); a dummy for race of survey respondent (White); a dummy for marital status of household head (Married); number of children in the household (Number of Kids); dummies for educational attainment (High School Graduate, Some College, and College Graduate); and survey year dummies (Year 2004, Year 2007).18 Descriptive statistics for these covariates are presented in the bottom panel of table 2.

Figure 2.

Figure 2

Survey of Consumer Finances, Combined 2001, 2004, 2007 Sample: Financial Assets (N=13,379) and Financial Asset Shares (N=12,723), by Age and Year (Weighted)

MFLOGIT Parameter Estimates and Inference

The MFLOGIT parameter estimates, normalized on the M-th (Other Financial Assets) category, are reported in table 3. Owing to the normalization, it is not straightforward to interpret the signs or magnitudes of these estimates. For purposes of hypothesis testing, however, such relative magnitudes may be informative, so asymptotic standard errors based on the robust sandwich estimator are presented in the table. Moreover, due to the large number of parameters (p × (M − 1)=90) estimated, a possible multiple comparisons situation arises for hypothesis testing. As such, and to accommodate the mutual dependence parameter estimates, the conservative false discovery rate (FDR) rejection criteria suggested by Benjamini and Yekutieli, 2001, are computed and point estimates with sufficiently small standard errors to meet these criteria are shaded in the table.

Table 3.

Financial Asset Shares, MFLOGIT Point Estimates and Robust Asymptotic Standard Errors (Normalization: βOther Financial Assets = 0; Shaded Entries Denote Conservative FDR Rejection Recommendation for H0: βkm = 0)

Financial Asset Category
Liquid
Assets
Quasi-
Liquid
Retir.
Accts.
CDs Dir.
Held
Pooled
Funds
Savings
Bonds
Directly
Held
Stocks
Directly
Held
Bonds
Cash
Val.
Whole
Life
Other
Managed
Funds
Age .001 .006 .049 .034 −.007 .039 .069 .020 .058
  s.e. .003 .003 .004 .003 .005 .003 .004 .003 .004
White −.264 .097 .203 .796 .508 .796 1.952 −.438 .946
  s.e. .102 .104 .139 .133 .192 .125 .288 .115 .181
Married .270 .873 .398 .890 .400 1.101 1.351 .546 .615
  s.e. .086 .088 .108 .100 .152 .097 .145 .101 .117
Number of Kids −.033 −.023 −.073 .001 .099 −.025 .021 .054 .012
  s.e. .039 .039 .052 .044 .054 .043 .059 .044 .055
High School Graduate −.024 .972 .672 1.405 .920 .945 1.287 .445 .888
  s.e. .154 .165 .193 .234 .340 .211 .421 .177 .256
Some College −.280 .898 .415 1.736 1.197 1.388 2.108 .271 1.113
  s.e. .161 .171 .204 .239 .341 .215 .413 .186 .262
College Graduate −.296 1.384 .678 2.681 .787 2.341 3.181 .161 1.923
  s.e. .145 .155 .183 .221 .330 .198 .388 .166 .239
Year 2004 −.041 −.069 −.208 −.164 −.014 −.332 −.034 −.486 −.331
  s.e. .100 .101 .125 .111 .161 .108 .137 .113 .128
Year 2007 −.249 −.118 −.296 −.216 −.336 −.508 −.394 −.714 −.656
  s.e. .100 .101 .122 .111 .166 .108 .140 .113 .130
Constant 2.611 .277 −3.184 −4.150 −2.285 −3.843 −9.356 −.199 −5.313
  s.e. .224 .231 .308 .305 .453 .298 .642 .259 .385

The extent of underdispersion in the data (relative to the nominal multinomial quasi-likelihood) can be appreciated in several ways. As suggested in section 6, the difference between the empirical multinomial logit (inverse Hessian) and robust sandwich parameter covariance matrixes is positive definite in the sample (the smallest eigenvalue of the matrix difference is positive). The ratios of non-robust to robust standard errors range from 1.12 to 2.37 over the 90 estimated parameters, suggesting a nontrivial degree of underdispersion.19

Average Partial Effects

Table 4 presents the weighted APE estimates and bootstrap 95% CIs (based on 500 bootstrap replications) across the M=10 outcomes. Recall that by construction the row sum of the APEs for each covariate will be zero. In this exercise, the schooling attainment and the year indicator variables are treated as groupwise dummies as discussed in Appendix 1.

Table 4.

Financial Asset Shares, MFLOGIT Estimates: Weighted APE Point Estimates and Bootstrap 95%-CI Lower and Upper Bounds (CIs based on 500 Bootstrap Replications and Hansen C2 Method)

Financial Asset Category
Liquid
Assets
Quasi-
Liquid
Retir.
Accts.
CDs Dir.
Held
Pooled
Funds
Savings
Bonds
Directly
Held
Stocks
Directly
Held
Bonds
Cash
Val.
Whole
Life
Other
Managed
Assets
Other
Financial
Assets
Age −.0037 −.0020 .0013 .0010 −.0002 .0016 .0007 .0005 .0012 −.0004
  CI-L −.0042 −.0024 .0011 .0008 −.0003 .0014 .0006 .0003 .0010 −.0006
  CI-U −.0033 −.0017 .0015 .0012 −.0001 .0018 .0009 .0007 .0014 −.0002
White −.0999 .0150 .0059 .0333 .0050 .0414 .0149 −.0355 .0196 .0003
  CI-L −.1174 −.0019 −.0009 .0270 .0022 .0348 .0127 −.0451 .0146 −.0069
  CI-U −.0824 .0304 .0126 .0398 .0079 .0490 .0169 −.0245 .0253 .0072
Married −.1030 .0783 −.0079 .0144 −.0019 .0325 .0094 −.0004 −.0008 −.0206
  CI-L −.1192 .0648 −.0141 .0088 −.0046 .0267 .0068 −.0081 −.0059 −.0272
  CI-U −.0879 .0909 −.0021 .0210 .0010 .0380 .0121 .0068 .0042 −.0144
Number of Kids −.0054 −.0014 −.0022 .0011 .0013 −.0006 .0006 .0051 .0009 .0007
  CI-L −.0115 −.0068 −.0052 −.0014 .0005 −.0033 −.0007 .0023 −.0018 −.0021
  CI-U .0004 .0037 .0006 .0035 .0022 .0021 .0021 .0081 .0031 .0033
High School Graduate −.1889 .1267 .0112 .0238 .0052 .0177 .0038 .0066 .0087 −.0146
  CI-L −.2153 .1056 .0022 .0176 .0009 .0083 .0007 −.0082 .0012 −.0255
  CI-U −.1607 .1486 .0216 .0309 .0096 .0257 .0065 .0214 .0156 −.0010
Some College −.2495 .1306 .0026 .0419 .0106 .0459 .0128 −.0002 .0160 −.0108
  CI-L −.2785 .1067 −.0081 .0343 .0056 .0351 .0086 −.0151 .0080 −.0238
  CI-U −.2186 .1534 .0130 .0509 .0158 .0559 .0171 .0134 .0240 .0039
College Graduate −.3540 .1714 −.0027 .0819 .0015 .0968 .0256 −.0288 .0300 −.0217
  CI-L −.3802 .1513 −.0110 .0748 −.0020 .0879 .0223 −.0424 .0237 −.0326
  CI-U −.3274 .1917 .0061 .0904 .0056 .1050 .0292 −.0158 .0361 −.0102
Year 2004 .0295 .0159 −.0027 −.0013 .0012 −.0149 .0020 −.0275 −.0062 .0040
  CI-L .0145 .0041 −.0087 −.0078 −.0015 −.0221 −.0014 −.0361 −.0111 −.0025
  CI-U .0443 .0289 .0035 .0048 .0043 −.0084 .0048 −.0196 −.0005 .0100
Year 2007 .0069 .0429 −.0003 .0040 −.0008 −.0174 −.0014 −.0324 −.0112 .0096
  CI-L −.0084 .0280 −.0066 −.0016 −.0036 −.0243 −.0038 −.0413 −.0163 .0036
  CI-U .0208 .0565 .0055 .0102 .0018 −.0106 .0014 −.0247 −.0063 .0159

Overall the estimated patterns of partial effects appear reasonable. The Age effects are consistent with the bulk of the estimation sample being of pre-retirement ages (70% are under age 60). Estimated patterns for Married and Number of Kids accord with the kinds of investment behaviors one would anticipate for such household structures for many of the shares (e.g. Liquid Assets, Quasi-Liquid Retirement Accounts, Directly Held Pooled Funds, Cash Value Whole Life). For many of the shares (e.g. Liquid Assets, Quasi-Liquid Retirement Accounts, Directly Held Pooled Funds, Directly Held Stocks), the estimated schooling attainment APEs are particularly large and precisely estimated, likely reflecting the schooling variables proxying for the most important human capital and wealth effects in these models. Finally the estimates for the year dummies show some interesting time trends for several of the share outcomes (e.g. Quasi-Liquid Retirement Accounts, Directly Held Stocks, Cash Value Whole Life, Other Managed Assets).

MFLOGIT Estimator Performance

For the conditional moment tests described in section 6, this application specifies the Λ(x) based on the vingtiles of each of the Ê[sik|xi] resulting in L=20 test indicators for each sk outcome. The sampling variation of these test indicators is estimated using 500 bootstrap replications, resulting in 95% CIs for each indicator-outcome combination as well as an overall χLM12 goodness-of-fit test statistic for the full multivariate model.

The results are summarized in figure 3, which depicts for each of the M=10 share outcomes and at each of the L=20 vingtiles the test statistic point estimate (N × λmq; dark line) and its bootstrap 95% CI (shaded area). Multiple comparisons concerns notwithstanding, the results depicted in figure 3 show only relatively few instances where the 95% CIs fail to cover zero, and these are typically in the tails with the exception of (Liquid Assets, Directly Held Stocks). In a few cases (Quasi-Liquid Retirement Accounts, CDs, Directly Held Stocks, Other Financial Assets) there is at least a suggestion of a U- or inverse U-shaped pattern across the vingtiles (underprediction in the tails and overprediction in the center, or vice-versa). Finally, the overall conditional moment χ1992 test statistic is 674.2 (p-value effectively zero).

Figure 3.

Figure 3

MFLOGIT Conditional Moment (CM) Tests Based on Percentiles of Estimated E[s|x]: Test Statistics N× λmq and Bootstrap 95%-CI Lower and Upper Bounds (CIs based on 500 Bootstrap Replications and Hansen C2 Method)

Several general model performance statistics are summarized in table 5. For this exercise, the MFLOGIT estimator was compared with a set of M=10 univariate linear regression (estimated by OLS) and univariate Tobit estimators that include the same covariate vector as used in the MFLOGIT models. The first performance criteria were out-of-sample MPE and MSE as assessed by an 80/20 cross-validation averaged over 100 replications. For this exercise, the linear model dominates MFLOGIT and Tobit on the MPE criterion while MFLOGIT generally dominates the linear model and Tobit on the MSE criterion. One obvious potential drawback of the linear model is that its predictions are not restricted to obey the [0,1] interval bounds. The rightmost columns of table 5 summarize the extent to which this is of concern in this empirical context. The two rightmost columns provide the cross-validation (averaged over replicates) and the in-sample frequencies with which the linear model predictions are less than zero (predictions greater than one were not observed). The cross-validation and in-sample frequencies are quite similar across the share categories, and suggest that the out-of-interval prediction problem is most severe for share categories with the smallest marginal empirical frequencies (e.g. Directly Held Bonds, Other Managed Assets).

Table 5.

Model Prediction Performance: 80/20 Cross-Validation for MPEs and MSEs, and Linear Model Predictions outside [0,1] Interval (Averages over 100 Replicates)

Out of Sample Predictions
(Best for Each Asset Category is Shaded)
Linear Model:
Fraction of
Predictions < 0
Mean Prediction Error Mean Squared Error
MFLOGIT Linear Tobit MFLOGIT Linear Tobit Out of
Sample
In-
Sample
Financial Asset Category Liquid Assets .0117 −.0002 −.0666 .1149 1141 .1187 0 0
Quasi-Liquid Retirement Accts. .0089 0003 .0082 1084 .1098 .1109 .0002 .0001
CDs .0014 0002 −.0031 0180 .0183 .0184 .0322 .0292
Directly Held Pooled Funds .0023 .0004 0001 0293 .0294 .0296 .0543 .0538
Savings Bonds .0004 0001 −.0048 00387 .00388 .0040 .0221 .0193
Directly Held Stocks .0018 −.0005 −.0034 03641 .0370 .03642 .0750 .0750
Directly Held Bonds .0005 00003 −.0007 0094 .0095 .0095 .1707 .1689
Cash Value of Whole Life .0019 −.0002 −.0110 03149 .03151 .0325 0 0
Other Managed Assets .0005 −.0004 −.0006 0196 .0197 .0197 .0987 .0984
Other Financial Assets .0014 0003 −.0043 01879 01879 .0190 .0007 .0006

Aggregation Testing

Figure 4 depicts the tree structure of the SCF categories. R=20 subcategories are specified for this analysis of structured aggregation as described in section 7 and Appendix 2, with the aggregation tests described in Appendix 2 applied to the three branches with #Ck>1 (Liquid Assets, Quasi-Liquid Retirement Funds, and Directly Held Bonds).20 The MFLOGIT point estimates of the θn parameters corresponding to those branches and the χ2 aggregation test statistics are presented in table 6.

Figure 4.

Figure 4

Surveys of Consumer Finances: Financial Asset Share Category Tree Structure (Subcategories Shaded in Light Gray Are the N=20 vn Used in the Empirical Analysis; Subcategories Shaded in Dark Gray Are Not Used in Empirical Analysis)

Table 6.

Disaggregated Share Model: MFLOGIT Point Estimates, Robust Asymptotic Standard Errors, and χ2 Aggregation Tests (Only Disaggregated Categories and Slope Parameters Shown; Normalization: θOther Financial Assets = 0)

Financial Asset Category
Liquid Assets Directly Held Pooled Funds Directly Held Bonds
Money
Market
Accounts
Checking
Accounts
Savings
Accounts
Call
Accounts
Category
Aggregate
(Reference)
Stock
Mutual
Funds
Tax-Free
Bond Mutual
Funds
Government
Bond Mutual
Funds
Other Bond
Mutual
Funds
Combination
& Other
Mutual
Funds
Category
Aggregate
(Reference)
Tax-Exempt
Bonds
Mortgage-
Backed
Bonds
U.S. Govt. &
Govt. Agcy.
Bonds &
Bills
Corporate &
Foreign
Bonds
Category
Aggregate
(Reference)
Subcategory % .167 .538 .281 .014 .668 .118 .031 .045 .138 .641 .059 .163 .137
Age .021 −.006 .001 .046 .001 .029 .054 .053 .051 .040 .034 .070 .093 .064 .076 .069
  s.e. .003 .003 .003 .006 .003 .003 .005 .008 .008 .005 .003 .005 .012 .010 .009 .004
White .087 −.337 −.313 1.411 .264 .751 .846 .766 1.405 1.017 .796 1.905 3.291 2.157 1.875 1.952
  s.e. .124 .105 .110 .309 .102 .138 .228 .301 .469 .228 .133 .348 .525 .540 .428 .288
Married .718 .101 .283 1.116 .270 .850 1.229 .738 .947 1.020 .890 1.543 1.177 1.009 1.254 1.351
  s.e. .103 .090 .094 .191 .086 .104 .164 .236 .249 .161 .100 .168 .361 .275 .265 .145
Number of Kids .017 −.041 −.049 .082 .033 −.009 −.037 .064 −.067 .106 .001 .006 −.028 .041 .095 .021
  s.e. .046 .041 .043 .072 .039 .045 .061 .089 .112 .067 .044 .065 .155 .099 .120 .059
H.S. Graduate .257 −.168 .147 .851 .024 1.506 .872 2.773 2.969 1.272 1.405 .895 3.925 1.908 15.35 1.287
  s.e. .196 .159 .166 .571 .154 .248 .396 .652 .652 .441 .234 .462 .905 .689 .401 .421
Some College .383 −.495 −.178 1.753 .280 1.793 1.377 3.450 3.220 1.710 1.736 1.797 4.672 2.809 15.91 2.108
  s.e. .203 .167 .175 .547 .161 .254 .387 .651 .650 .444 .239 .449 .879 .628 .380 .413
Coll. Graduate .917 −.761 −.240 2.684 .296 2.734 2.284 4.173 4.523 2.782 2.681 2.880 5.737 3.876 17.13 3.181
  s.e. .181 .150 .157 .511 .145 .235 .358 .609 .591 .410 .221 .414 .768 .572 .244 .388
Year 2004 −.085 .116 −.296 −.283 .041 −.266 −.340 −.081 .183 .408 −.164 −.047 .147 −.402 .279 .034
  s.e. .114 .105 .108 .177 .100 .114 .155 .245 .236 .176 .111 .151 .320 .228 .220 .137
Year 2007 −.221 −.137 −.439 −.650 .249 −.390 −.113 −.333 −.130 .582 .216 −.351 −.861 −.422 −.420 .394
  s.e. .114 .105 .109 .188 .100 .114 .157 .227 .248 .173 .111 .154 .406 .235 .237 .140
2 Test Stats.
  Within-Categ. 952.3 (d.f.=27, p<.0001) 153.5 (d.f.=36, p<.0001) 1,591.2 (d.f.=27, p<.0001)
  Overall 2,75.8 (d.f.=90, p<.0001)

The overall aggregation test for all three branches strongly rejects slope-parameter aggregation across all three branches. The individual category tests for aggregation also strongly suggest that aggregation is not reasonable for any of these three categories.21 Indeed, casual inspection of the individual slope-parameter point estimates indicates considerable variability within each main category.

Dirichlet-Multinomial Estimates

Two variants of the DM model were estimated here, these reflecting different degrees of data coarsening as discussed in section 8. Specifically, models for T=10 and T=100 were estimated. A useful, direct comparison between the DM and MFLOGIT estimators of concern here is in terms of their performance in estimating the conditional first-moment structure of the data, with these summarized most straightforwardly by comparing the point estimates of the estimated APEs. This comparison is presented in table 7. In most cases (the exceptions being the shaded cells) the MFLOGIT and DM APE estimates have the same signs. Quite broadly, the magnitudes of the point APE estimates roughly comparable, but typically larger for MFLOGIT than for either the T=10 or T=100 DM estimators. One consideration beyond the comparison of the APEs is the possible efficiency gain from using a full-likelihood estimation approach (DM) over a first-moment estimation approach (MFLOGIT). It turns out in this application that the efficiency gains are small at best and may be realized at the risk of nonrobustness relative to MFLOGIT. For the p × (M − 1) = 90 normalized parameters, the median of the ratio of MFLOGIT to DM robust standard errors is 1.05 for the T=100 model and 0.99 for the T=10 specification.

Table 7.

Financial Asset Shares, Weighted APE Point Estimates: MFLOGIT and DM (T=10 and T=100) Estimator Comparison (Shaded Cells Indicate Discordant Signs)

Financial Asset Category
Liquid
Assets
Quasi-
Liquid
Retir.
Accts.
CDs Dir.
Held
Pooled
Funds
Savings
Bonds
Directly
Held
Stocks
Directly
Held
Bonds
Cash
Val.
Whole
Life
Other
Managed
Assets
Other
Financial
Assets
Age MFLOGIT −.0037 −.0020 .0013 .0010 −.0002 .0016 .0007 .0005 .0012 −.0004
DM, T=100 −.0029 −.0016 .0010 .0007 −.0003 .0012 .0005 .0010 .0006 −.0001
DM, T=10 −.0018 −.0022 .0009 .0007 −.0001 .0012 .0005 .0003 .0007 −.0002
White MFLOGIT −.0999 .0150 .0059 .0333 .0050 .0414 .0149 −.0355 .0196 .0003
DM, T=100 −.0868 .0236 .0055 .0220 .0044 .0304 .0099 −.0162 .0101 −.0030
DM, T=10 −.0679 .0166 .0016 .0245 .0029 .0300 .0092 −.0259 .0113 −.0023
Married MFLOGIT −.1030 .0783 −.0079 .0144 −.0019 .0325 .0094 −.0004 −.0008 −.0206
DM, T=100 −.0978 .0702 −.0028 .0077 −.0032 .0217 .0050 .0112 .0010 −.0132
DM, T=10 −.0736 .0629 −.0049 .0084 −.0040 .0225 .0055 −.0010 −.0007 −.0150
N. of Kids MFLOGIT −.0054 −.0014 −.0022 .0011 .0013 −.0006 .0006 .0051 .0009 .0007
DM, T=100 −.0074 −.0005 −.0014 .0020 .0031 −.0003 .0007 .0049 .0003 −.0015
DM, T=10 −.0031 −.0014 −.0024 .0015 .0011 −.0002 .0007 .0039 .0008 −.0007
H.S. Grad. MFLOGIT −.1889 .1267 .0112 .0238 .0052 .0177 .0038 .0066 .0087 −.0146
DM, T=100 −.1410 .0844 .0071 .0156 .0116 .0141 .0018 .0127 .0020 −.0083
DM, T=10 −.1298 .0951 .0055 .0168 .0046 .0141 .0029 .0010 .0028 −.0130
Some Coll. MFLOGIT −.2495 .1306 .0026 .0419 .0106 .0459 .0128 −.0002 .0160 −.0108
DM, T=100 −.1911 .0904 .0001 .0272 .0147 .0375 .0093 .0070 .0088 −.0037
DM, T=10 −.1666 .0935 −.0017 .0283 .0075 .0364 .0082 −.0042 .0103 −.0120
Coll. Grad. MFLOGIT −.3540 .1714 −.0027 .0819 .0015 .0968 .0256 −.0288 .0300 −.0217
DM, T=100 −.2921 .1465 −.0057 .0568 .0026 .0776 .0173 −.0085 .0139 −.0085
DM, T=10 −.2380 .1302 −.0107 .0572 −.0009 .0737 .0180 −.0274 .0154 −.0176
Year 2004 MFLOGIT .0295 .0159 −.0027 −.0013 .0012 −.0149 .0020 −.0275 −.0062 .0040
DM, T=100 .0209 .0019 −.0028 −.0041 .0017 −.0053 −.0006 −.0117 −.0027 .0027
DM, T=10 .0129 .0125 −.0001 −.0013 .0012 −.0082 .0012 −.0159 −.0040 .0016
Year 2007 MFLOGIT .0069 .0429 −.0003 .0040 −.0008 −.0174 −.0014 −.0324 −.0112 .0096
DM, T=100 .0108 .0203 .0051 −.0038 −.0006 −.0106 −.0022 −.0167 −.0055 .0033
DM, T=10 .0013 .0304 .0032 .0024 −.0003 −.0123 −.0008 −.0197 −.0077 .0034

One test of overall goodness of fit for likelihood-based models like the DM is the information matrix (IM) test proposed by White, 1982 (see also Chesher, 1983, Lancaster, 1984, and Orme, 1990). With such a large model as that estimated here, it is not obvious whether the asymptotic properties of the IM test can be invoked given the available sample size, as the full model IM test has 4,095 degrees of freedom with a sample size of 12,723.22 Such considerations notwithstanding, the IM test statistic is computed for the T=10 and T=100 specifications using the method suggested by Lancaster, 1984. The DM model's χ40952 test statistics are 11,261 (T=10) and 10,575 (T=100) which, while suggesting a slightly better fit for T=100 than for T=10, still both have effective p-values of zero. For comparison, however, the MFLOGIT quasi-likelihood IM test statistic based on the coarsened T=10 data is 1.72E+07. More concretely perhaps, for all p × M=100 parameter estimates, the ratio of robust to inverse-Hessian standard error estimates ranges from .69 to 1.12 (median 1.02) for the T=10 specification and from .62 to 1.20 (median 1.03) for the T=100 specification.

It is also possible (though not undertaken here) to compute an overall χ2 goodness-of-fit test statistic (see Andrews, 1988) for the coarsened outcome cells or interesting aggregates thereof. In this spirit, the performance of the DM estimator in modeling the overall conditional probability structure of the data is summarized in figures 5 and 6 and in table 8. Figure 5 depicts for three of the share outcomes the coarsened marginal empirical frequency distribution juxtaposed with the estimated marginal empirical frequencies from the T=10 and T=100 specifications (the latter computed as the sample averages of the conditional frequencies). Figure 6 shows Lorenz curve summaries in which are plotted the T=100 marginal cumulative probability estimates against the corresponding cumulative 101 coarsened cell frequencies in the data. For Liquid Assets, Quasi-Liquid Retirement Accounts, Directly Held Pooled Funds, and Directly Held Stocks there are some noteworthy fit problems. Finally, table 8 summarizes the quality of fit of the DM estimators at the s=0 and s=1 endpoints. For most of the share outcomes, the T=100 estimator provides a much closer fit to the marginal empirical frequencies than does the T=10 estimator.

Figure 5.

Figure 5

Comparison of Actual (Light Blue) and Estimated (Dark Blue) Marginal Distributions, Dirichlet-Multinomial Model, Selected Outcomes, T=10 and T=100 (Estimated Distributions are Unweighted Sample Averages of Estimated Cell Probabilities)

Figure 6.

Figure 6

Lorenz Curve Plots of DM Estimates (T=100) against Corresponding Coarsened Data

Table 8.

DM Estimator Fit Performance: Pr (sk = 0) and Pr (sk = 1)

Financial Asset Category
Liquid
Assets
Quasi-
Liquid
Retir.
Accts.
CDs Dir.
Held
Pooled
Funds
Savings
Bonds
Directly
Held
Stocks
Directly
Held
Bonds
Cash
Val.
Whole
Life
Other
Managed
Assets
Other
Financial
Assets
Pr (sk = 0) Sample Frequency .018 .355 .823 .745 .820 .654 .907 .669 .885 .856
DM Estimate, T=10 .137 .450 .908 .811 .975 .760 .942 .857 .921 .935
DM Estimate, T=100 .105 .335 .840 .740 .910 .658 .910 .721 .889 .882
Pr (Sk = l) Sample Frequency .170 .005 .001 .0002 .001 .001 0 .004 .0001 .004
DM Estimate, T=10 .166 .028 .003 .003 .001 .004 .0004 .005 .001 .002
DM Estimate, T=100 .099 .012 .002 .001 .001 .002 .0002 .003 .001 .001

Overall, then, the merits of estimating and conducting inference using DM models in a full likelihood context are mixed. One presumably trades off some robustness relative to approaches like MFLOGIT in estimating first-moment models, with the differences in corresponding APE point estimates between the two approaches perhaps being nontrivial, at least in this empirical exercise. However, if one is interested in estimating the full conditional probability structure of such models (perhaps at the cost of using coarsened data), the DM approach may be a useful strategy to consider.

10. Summary and Discussion

With a central focus on estimation of conditional means, this paper has proposed econometric strategies for estimating regression models of economic share data in cases where shares assume values of zero and one with nontrivial probability. The main contribution has been to explore properties of an extension to share models of the fractional regression methodologies proposed by Papke and Wooldridge.

Several outstanding issues should be important items on the future research agenda. First involves further considerations of share category aggregation or disaggregation beyond those offered in section 7. A second consideration involves "covariate adjustment." For instance, in applications where understanding the determinants of shares net of the influence of some conditioning covariates is a prominent issue, the manner in which covariates are netted out is critical. How to effect this in a framework that involves bounded shares that obey adding-up restrictions is an open question.

Third, in some empirical contexts (e.g. like the portfolio share example presented here) there is necessarily selection on subsamples for which shares are defined (by nonzero denominators). This raises an important issue regarding for which populations inferences drawn from the estimated share models are relevant. While in some sense a garden variety selection problem, the issue of how to address this in the MFLOGIT or related estimation contexts remains unresolved.

Finally, a possible extension of this line of work would be to consider analogs to conditional logit (Hausman and McFadden, 1984) estimation that would fit into the fractional outcome data setting. While analogies to the discrete outcome random utility (RUM) framework are not obvious, the implied moment structures of such models might offer statistical tools for analyzing data where outcome-specific covariates or attributes are available. For instance, one might imagine a time use study in which time prices or wage rates for each outcome are available. Briefly, consider a situation where a vector of attributes for the k-th outcome is given by wk with the vector w = [wk, w−k] describing the entirety of such attributes over all outcomes. Then the first-moment share structure corresponding to a standard RUM model (with normalization wM=0) would be given by:

E[sk|wk,wk]=E[sk|w]=exp(wkδ)1+m=1M1exp(wmδ),k=1,,M, (36)

which could be extended to accommodate x in a mixed-logit structure,

E[sk|x,w]=exp(xβk+wkδ)1+m=1M1exp(xβm+wmδ),k=1,,M. (37)

Of course, in the absence of an underlying RUM structure, the influences of such outcome-specific attributes could also be captured in a standard MFLOGIT conditional mean model with

E[sk|x,w]=exp(xβk+wδk)1+m=1M1exp(xβm+wδm),k=1,,M, (38)

in which the δk would describe different patterns of own-w vs. cross-w effects across k=1,…,M.

Acknowledgments

For their thoughtful comments, suggestions, and discussions I am indebted to participants at presentations of various aspects of this work at the Rome Health Econometrics Workshop, Michigan State University, University of Arizona, University of Coimbra, University College Dublin, and UW-Madison, as well as to Badi Baltagi, Marguerite Burns, Ben Craig, Alberto Holly, Steve Koch, José Murteira, Tim Park, Stephanie Robert, Nilay Shah, João Santos Silva, Dave Vanness, Jeff Wooldridge, and two referees. All are absolved from any blame for the paper's shortcomings. Partial financial support from the Robert Wood Johnson Foundation Health & Society Scholars Program is acknowledged. Some of this work was completed as a visiting scholar at the UCD Geary Institute, which provided brilliant hospitality.

Appendix 1

Average Partial Effects for MFLOGIT Models

The general formula for the APE is

APE^mk=i=1N(win=1Nwn)PE^mki=i=1N(win=1Nwn)δE^[yim|xi]δxik,

where "δ" denotes either "Δ" or "∂" and where the wi are nonnegative weights that may be used to estimate, for instance, population average APEs (wi=1 for all i gives constant weight 1/N). Note that m=1MAPE^mk=0 due to the adding-up restriction. In the case where xik is a dummy variable, APEmk is computed as the (perhaps weighted) sample average, evaluated at β = β̂, of the difference

PEmki=ΔE[sim|xi]Δxik=exp(xk,iβm,k+βmk)1+j=1M1exp(xk,iβj,k+βjk)exp(xk,iβm,k)1+j=1M1exp(xk,iβj,k),

or the derivative

PEmki=E[sim|xi]xik=exp(xiβm)×(1+j=1M1exp(xiβj))×βmkj=1M1exp(xiβj)×βjk(1+j=1M1exp(xiβj))2,

where x−k,i is the vector xi for the i-th observation with the k-th element excluded.23

Hansen, 2010, gives the lower and upper limits of the C2 (1 − α) −CI for arbitrary statistic θ̂ as:

CIL(α,θ^,{θ^b})=θ^q{θ^b}(1.5α) and CIU(α,θ^,{θ^b})=θ^q{θ^b}(.5α),

where q{θ̂b} (τ) is the τ -th quantile of the bootstrap sampling distribution {θ̂b}. When θ^=APE^mk, there are at least three possibilities for bootstrapping to estimate the (1 − α) − CI:

  1. Accommodate variation in x and β̂ : Bootstrap the APEs via bootstrap draws (yb, xb) = [(yi(b), xi(b))] from (y, x) that estimate β̂b and, correspondingly, APEmkb(xb;β^b)=1Ni(b)=1NPEmki(b)(xi(b);β^b) accumulate the APEmkb (xb; β̂b) in the B-vector [APEmkb (xb; β̂b)], and base CIs on suitable percentiles of [APEmkb (xb; β̂b)].

  2. Accommodate variation in β̂ only: Bootstrap the APEs via bootstrap draws (yb, xb) from (y, x) that estimate β̂b and, correspondingly, APEmkb(x;β^b)=1Ni=1NPEmki(xi;β^b); accumulate the APEmkb (x; β̂b) in the B-vector [APEmkb (x; β̂b)]; and base CIs on suitable percentiles of [APEmkb (x; β̂b)].

  3. Accommodate variation in β̂ only, with weighting: Bootstrap the APEs via unweighted bootstrap replicates (yb, xb) from (y, x) that estimate β̂b and, correspondingly, APEmkbw(x;β^b)=i=1N(win=1Nwn)PEmki(xi;β^b); accumulate the APEmkbw(x;β^b) in the B-vector [APEmkbw(x;β^b)]; and base CIs on suitable percentiles of [APEmkbw(x;β^b)].

The estimates presented in tables 4 and 7 are based on approach (c), but in this application it turns out that the CIs estimated using (a), (b), or (c) are quite similar (tables showing alternatives are available on request).

Appendix 2

Approaches to Aggregation

Only structured aggregation will be considered here. Moreover, while the basic ideas generalize to multiple levels, this discussion considers only a simple two-level aggregation context in which there are M outcome categories or aggregates (e.g. the m=1,…,M sm measures) and R>M outcome subcategories or disaggregates (denoted vr, r=1,…,R) with m=1Msm=r=1Rvr=1. Define implicitly M index sets Cm via Σr∈Cm vr = sm, m=1,…,M, with m=1MCm={1,,R} and m=1MCm=. Thus, in the example depicted in the top panel of figure 1, C1 = {1, 2, 3}, C2 = {4}, …, and CM = {R − 2, R − 1, R}.

As such, and ignoring for now a necessary identifying parameter normalization,

E[sm|x]=E[nCmvn|x]=nCmE[vn|x]=nCm(exp(xθn)r=1Rexp(xθr))=nCm(exp(θn0+x1θn1)r=1Rexp(θr0+x1θr1))=nCm(exp(θn0+x1θn1)m=1MrCmexp(θr0+x1θr1)),m=1,,M.

Suppose for some k it holds that all elements of the set of slope parameter vectors {θn1 |n ∈ Ck} are identical and equal to (say) Θk. Then

E[sk|x]=exp(ln(qCkexp(θq0))+x1Θk)exp(ln(qCkexp(θq0))+x1Θk)+m=1mkMrCmexp(θr0+x1θr1)=exp(θ0k+x1Θk)exp(θ0k+x1Θk)+m=1mkMrCmexp(θr0+x1θr1),

where θ0k=ln(qCkexp(θq0)). That is, the subcategory outcomes vn, n ∈ Ck, aggregate in the sense that they share common slope coefficient vectors. While this is perhaps an obvious characterization of aggregation in the fractional share outcome setting, its deeper implications are less obvious. For instance, aggregation in this sense would imply that the aggregated subcategories all have the same conditional x1-elasticities but not the same conditional x1-partial effects.

Given considerations of structured aggregation in this slope-coefficient sense, at least two testing strategies are suggested.24 The first entails testing jointly the entirety of the equality restrictions implied if the subcategories under all categories having at least two subcategories simultaneously aggregate thusly. Note that if all outcome subcategories branching from the aggregated outcome categories aggregated in the slope-coefficient sense, it would follow that

E[nCkvn|x]=E[sk|x]=exp(θ0k+x1Θk)m=1Mexp(θ0m+x1Θm),k=1,,M.

Given estimates of the subcategory model, this test could be conducted as a straightforward Wald test of the implied parameter restrictions on the (p−1)-vectors θr1. Under a null hypothesis of slope-coefficient aggregation, such a test statistic would follow a large-sample χ(p1)(NM)2 distribution.

A second and likely more useful approach involves testing separately each candidate subcategory aggregation. For each k this entails testing jointly the equality of the elements in {θn1 |n ∈ Ck} via Wald tests. In isolation, such test statistics would follow null χ(#Ck1)×(p1)2 distributions. However, simultaneous testing of such restrictions across as many as M aggregate categories presents a multiple comparisons problem. Using the Benjamini-Yekutieli, 2001, conservative false discovery rate control approach (or related methods) to modify the set of reference rejection p-values may provide some protection against false positives.

Let θ = vec ([θ1, …, θN−1]) denote the p (N − 1) × 1 vector of estimable parameters and suppose N>M. Let θ be defined such that: (a) disaggregated subcategory outcomes from the same multi-subcategory branch have corresponding θk that are adjacent in θ; (b) that the first element of each θk is the "intercept" parameter, i.e. θk=[θk0,θk1']'; and (c) the K (0 ≤ K < M) aggregate outcome categories that branch to only single outcome subcategories correspond to the bottom rows of θ. It is assumed that the normalization θN = 0 has been imposed and, for simplicity of exposition, that subcategory vN is not being considered for aggregation with other subcategories. Then the Wald test statistics described in section 7 are given by the standard formula

wald(θ^)=(Rθ^)'(RV^(θ^)R')1(Rθ^),

where (θ̂) is the estimate of the robust MFLOGIT covariance estimator given in (25).

The test for the entirety of slope-parameter aggregations specifies R as follows. Let

Jp=[0(p1)×1,I(p1)],  dim(Jp)=(p1)×p,

and

Rk=[[blockdiag(#Ck1)(Jp)],[JpJpJp]],k=1,,(MK),  dim(Rk)=((p1)(#Ck1))×(p#Ck).

Then

R=[blockdiag(k=1,,(MK))([Rk]),0(p1)(NM)×p(K1)],  dim(R)=(p1)(NM)×p(N1).

In this case wald(θ̂) follows a χ(p1)(NM)2 distribution under the null.

For testing subcategory aggregation under a single outcome aggregate (say the m-th), the corresponding specification of R is

R=[0(p1)(#Cm1)×p(k<m#Ck),Rm,0(p1)(#Cm1)×p(N1km#Ck)],

with Rm specified as Rk above and dim(R) = (p − 1) (#Cm − 1) × p (N − 1). In this case wald(θ̂) follows a χ(p1)(#Cm1)2 distribution under the null.

Appendix 3

Survey of Consumer Finances Data

Downloaded in February and March 2010, the public use data files are contained in:

The financial asset data used in this analysis are derived at the household level by averaging over the five SCF "implicates" (imputed replicates) for each category, summing these household averages to obtain household total financial assets, and then obtaining the category shares as the ratios of the category implicate averages to this household total, so-defined. Specifically, for each survey year's sample, the Stata v.10 code used to compute the category shares is:

       sort yy1
       mac def nlfin "liq cds nmmf savb stoc bond cashli othma retq othf"
       mac def nlfina "aliq acds anmmf asavb astoc abond acashli aothma aretq aothf"
       foreach var of varlist $nlfin {
         by yy1: egen a`var'=mean(`var')
       }
       egen afin=rowtotal($nlfina)
       foreach var of varlist $nlfin {
         gen sh`var'=a`var'/afin
       }

This results in defined shares for 12,723 of the 13,379 total household-level observations.

Footnotes

1

The "i" subscript indexing observations will be suppressed henceforth unless useful for clarity.

2

Some applications may involve multivariate bounded outcome data that are not subject to adding-up restrictions like (5). The analysis of such "seemingly unrelated" outcomes might proceed generally by considering modifications of the framework proposed by Papke and Wooldridge, 2008, for panel data structures.

3

If Pr (yk = 0| x) = 1 or Pr (yk = B| x) = 1 for all x values subsequent analysis would likely be uninteresting and/or trivial.

4

This may be a reasonable characterization of the shares that are analyzed in the empirical analysis reported in section 9. Here the shares are the fractions of overall financial assets in each of ten financial asset categories. Even were the overall level of assets to be characterized as nonstochastic, the split between financial and nonfinancial assets and, therefore, total financial assets, would presumably be stochastic.

6

For instance, Kooreman and Kapteyn's elegant analysis of time use demands notes the potential for boundary solutions but then goes on to comment: "We will ignore the [boundary condition constraint equations in the theoretical model], which are binding for only a limited number of observations."

7

Woodland specifies a general functional form for the Dirichlet regression parameters but then uses a linear (not exponential) specification in his empirical analysis. Nonetheless, since the Dirichlet parameters must be positive, an exponential specification is appealing. (José Murteira, in a private communication, pointed out that the expression for the Dirichlet density in Woodland's equation (4) contains a typographical error: The product in the denominator of the correct expression is displayed as a summation in Woodland's paper.)

8

See Ramalho et al., 2010, for an excellent survey of fractional regression model estimation.

9

This estimation strategy has been applied in Mullahy, 2004, Mullahy and Robert, 2010, Koch, 2010, as well as in the transportation research literature (Sivakumar and Bhat, 2002; Ye and Pendyala, 2005).

11

Some canned multinomial logit estimation packages, e.g. Stata's mlogit, do not accommodate nonbinary sm. The estimates presented here and in section 8 are obtained using procedures written in Stata's Mata language, which are available on request.

12

In closing this section it might be noted that it would be feasible and straightforward to ignore the share system nature of the outcome data and estimate a first-moment structure with M independent binary FLOGIT models (11) yielding M corresponding ω̂k estimates and APEs. While straightforward, such an approach ultimately implies an unrealistic form of aggregation across the M categories (see section 7 for related discussion). Yet whether such an approach would have substantive implications for the estimates of parameters of interest like APEs is an open question; preliminary results using the SCF data (not reported here) show considerable similarity of the APE estimates from the MFLOGIT and the independent binary FLOGIT estimators.

13

See Graybill, 1983, Theorem 12.2.16, and Intriligator, 1971, equation B.7.10.

14

Well-known examples are two-to-six-digit NAICS/SIC industry definition codes, one-to-three-level expenditure hierarchies used in the Consumer Expenditure Survey, and the American Time Use Survey's two-, four-, and six-digit time-use categories.

15

While their purpose was different that this paper's, Cotterman and Peracchi, 1992 offer a useful conceptual discussion of structured vs. unstructured aggregation.

16

See Vanness and Hanmer, 2010, for a related discussion in a univariate and Bayesian context.

17

See Guimarães and Lindrooth, 2007, for an interesting econometric application of the DM distribution.

18

These specifications do not include Total Financial Assets as a covariate. Whether it is appropriate to do so entails issues akin to those involved in whether to include a measure of total expenditure as a covariate in a consumer expenditure share model.

19

Given the considerable size of the estimated parameter covariance matrix (4,095 unique elements), a comparison bootstrap covariance matrix was estimated using 1,000 bootstrap replicates to check the performance of the standard robust sandwich parameter covariance estimator. Across the 90 parameters estimated in this specification, the mean and median ratios of robust to bootstrap standard errors were .993 and .999, respectively, with a range of 0.86 to 1.06. To the extent that this result generalizes, use of the sandwich estimator in empirical applications of the MFLOGIT estimator may be reasonable.

20

The subcategories under the Quasi-Liquid Retirement Accounts and Other Managed Assets categories were ignored for computational reasons.

21

With only three subaggregate branches specified and in light of the large values of the realized aggregation test statistics, the multiple comparisons issues discussed above in section 7 are effectively irrelevant and are ignored here.

22

For comparison with the MFLOGIT quasi-likelihood, the parameter estimates are normalized against the Other Financial Assets baseline, thus reducing the dimensionality of the test.

23

When dummy variables are included in x as mutually exclusive and exhaustive (save an "omitted" category) members of sets of indicators -- e.g. race/ethnicity groups, educational attainment indicators -- setting up the discrete APE to capture the proper counterfactual is accomplished by zeroing out all of the group's dummy variables at baseline (i.e. setting all group dummies for all observations equal to the omitted category) and then setting the xik the variable in question equal to one for all observations.

24

These testing strategies follow the basic approach of Cramer and Ridder, 1991, except that Cramer and Ridder: are concerned with discrete outcomes; rely on likelihood-ratio tests; and do not address issues of multiple testing that are raised below.

References

  1. Aitchison J. The Statistical Analysis of Compositional Data. JRSS-B. 1982;44:139–177. [Google Scholar]
  2. Andrews DWK. Chi-Square Diagnostic Tests for Econometric Models: Theory. Econometrica. 1988;56:1419–1453. [Google Scholar]
  3. Benjamini Y, Yekutieli D. The Control of the False Discovery Rate in Multiple Testing under Dependency. Annals of Statistics. 2001;29:1165–1188. [Google Scholar]
  4. Berry S, Levinsohn J, Pakes A. Automobile Prices in Market Equilibrium. Econometrica. 1995;63:841–890. [Google Scholar]
  5. Billheimer D, Guttorp P, Fagan WF. Statistical Interpretation of Species Composition. JASA. 2001;96:1205–1214. [Google Scholar]
  6. Brown BW, Walker MB. The Random Utility Hypothesis and Inference in Demand Systems. Econometrica. 1989;57:815–829. [Google Scholar]
  7. Brown BW, Walker MB. Stochastic Specification in Random Production Models of Cost- Minimizing Firms. Journal of Econometrics. 1995;66:175–205. [Google Scholar]
  8. Bucks BK, Kennickell AB, Mach TL, Moore KB. Changes in U.S. Family Finances from 2004 to 2007: Evidence from the Survey of Consumer Finances. Federal Reserve Bulletin. 2009;95:A1–A55. [Google Scholar]
  9. Chavas J-P, Segerson K. Stochastic Specification and Estimation of Share Equation Systems. Journal of Econometrics. 1987;35:337–358. [Google Scholar]
  10. Chesher A. The Information Matrix Text: Simplified Calculation via a Score Test Implementation. Economics Letters. 1983;13:45–48. [Google Scholar]
  11. Christensen LR, Jorgenson DW, Lau LJ. Transcendental Logarithmic Utility Functions. American Economic Review. 1975;65:367–383. [Google Scholar]
  12. Considine TJ, Mount TD. The Use of Linear Logit Models for Dynamic Input Demand Systems. Review of Economics and Statistics. 1984;66:434–443. [Google Scholar]
  13. Cotterman R, Peracchi F. Classification and Aggregation: An Application to Industrial Classification in CPS Data. Journal of Applied Econometrics. 1992;7:31–51. [Google Scholar]
  14. Cramer JS, Ridder G. Pooling States in the Multinomial Logit Model. Journal of Econometrics. 1991;47:267–272. [Google Scholar]
  15. Crawford DL, Pollak RA, Vella F. Simple Inference in Multinomial and Ordered Logit. Econometric Reviews. 1998;17:289–299. [Google Scholar]
  16. Dubin JA. Valuing Intangible Assets with a Nested Logit Market Share Model. Journal of Econometrics. 2007;139:285–302. [Google Scholar]
  17. Fry JM, Fry TRL, McLaren KR. The Stochastic Specification of Demand Share Equations: Restricting Budget Shares to the Unit Simplex. Journal of Econometrics. 1996;73:377–385. [Google Scholar]
  18. Gourieroux C, Monfort A, Trognon A. Pseudo Maximum Likelihood Methods: Theory. Econometrica. 1984a;52:681–700. [Google Scholar]
  19. Gourieroux C, Monfort A, Trognon A. Pseudo Maximum Likelihood Methods: Applications to Poisson Models. Econometrica. 1984b;52:701–720. [Google Scholar]
  20. Graybill FA. Matrixes with Applications in Statistics. Belmont, CA: Wadsworth Publishing; 1983. [Google Scholar]
  21. Guimarães P, Lindrooth RC. Controlling for Overdispersion in Grouped Conditional Logit Models: A Computationally Simple Application of Dirichlet-Multinomial Regression. Econometrics Journal. 2007;10:439–452. [Google Scholar]
  22. Hansen B. Econometrics. Textbook manuscript, Dept. of Economics; UW-Madison: 2010. [Google Scholar]
  23. Hausman J, McFadden D. Specification Tests for the Multinomial Logit Model. Econometrica. 1984;52:1219–1240. [Google Scholar]
  24. Heaton J, Lucas D. Portfolio Choice and Asset Prices: The Importance of Entrepreneurial Risk. Journal of Finance. 2000;55:1163–1198. [Google Scholar]
  25. Heckman JJ, Willis RJ. A Beta-Logistic Model for the Analysis of Sequential Labor Force Participation by Married Women. Journal of Political Economy. 1977;85:27–58. [Google Scholar]
  26. Hill MA. Female Labor Force Participation in Developing and Developed Countries-Consideration of the Informal Sector. Review of Economics and Statistics. 1983;65:459–468. [Google Scholar]
  27. Intriligator MD. Mathematical Optimization and Economic Theory. Englewood Cliffs, NJ: Prentice-Hall; 1971. [Google Scholar]
  28. Johnson NL, Kotz S, Balakrishnan N. Discrete Multivariate Distributions. New York: Wiley; 1997. [Google Scholar]
  29. Koch SF. Fractional Multinomial Response Models with an Application to Expenditure Shares. Mimeo 2010 [Google Scholar]
  30. Kooreman P, Kapteyn A. A Disaggregated Analysis of the Allocation of Time within the Household. Journal of Political Economy. 1987;95:223–249. [Google Scholar]
  31. Lancaster T. The Covariance Matrix of the Information Matrix Test. Econometrica. 1984;52:1051–1053. [Google Scholar]
  32. Lee L-F, Pitt MM. Microeconometric Demand System with Binding Nonnegativity Constraints: The Dual Approach. Econometrica. 1986;54:1237–1242. [Google Scholar]
  33. McElroy MB. Additive General Error Models for Production, Cost, and Derived Demand or Share Systems. Journal of Political Economy. 1987;95:737–757. [Google Scholar]
  34. Morey ER, Waldman D, Assane D, Shaw D. Searching for a Model of Multiple-Site Recreation Demand That Admits Interior and Boundary Solutions. American Journal of Agricultural Economics. 1995;77:129–140. [Google Scholar]
  35. Mullahy J. Working Paper. University of Wisconsin; 2004. Squandering Time? Economic Aspects of Children's Time Use. [Google Scholar]
  36. Mullahy J, Robert SA. No Time to Lose: Time Constraints and Physical Activity in the Production of Health. Review of Economics of the Household. 2010;8:409–432. [Google Scholar]
  37. Orme C. The Small-Sample Performance of the Information-Matrix Test. Journal of Econometrics. 1990;46:309–331. [Google Scholar]
  38. Papke LE, Wooldridge JM. Econometric Methods for Fractional Response Variables with an Application to 401(k) Plan Participation Rates. Journal of Applied Econometrics. 1996;11:619–632. [Google Scholar]
  39. Papke LE, Wooldridge JM. Panel Data Methods for Fractional Response Variables with an Application to Test Pass Rates. Journal of Econometrics. 2008;145:121–133. [Google Scholar]
  40. Poterba JM, Samwick AA. Household Portfolio Allocation over the Life Cycle. In: Ogura S, Tachibanaki T, Wise DA, editors. Aging Issues in the United States and Japan. Vol. 2. Chicago: University of Chicago Press for NBER; 2001. [Google Scholar]
  41. Poterba JM, Samwick AA. Taxation and Household Portfolio Composition: U.S. Evidence from the 1980s and 1990s. Journal of Public Economics. 2002;87:5–38. [Google Scholar]
  42. Ramalho EA, Ramalho JJS, Murteira JMR. Alternative Estimating and Testing Empirical Strategies for Fractional Regression Models. Journal of Economic Surveys. 2010 (forthcoming) [Google Scholar]
  43. Sivakumar A, Bhat C. Fractional Split-Distribution Model for Statewide Commodity-Flow Analysis. Transportation Research Record. 2002;1790:80–88. [Google Scholar]
  44. Theil H. A Multinomial Extension of the Linear Logit Model. International Economic Review. 1969;10:251–259. [Google Scholar]
  45. Vanness DJ, Hanmer J. Presented at the 3rd Biennial Conference of the American Society of Health Economists. Cornell University; 2010. Health Utility Crosswalks: A Bayesian Beta Regression Approach. [Google Scholar]
  46. Wales TJ, Woodland AD. Estimation of the Allocation of Time for Work, Leisure, and Housework. Econometrica. 1977;45:115–132. [Google Scholar]
  47. White H. Maximum Likelihood Estimation of Misspecified Models. Econometrica. 1982;50:1–25. [Google Scholar]
  48. Woodland AD. Stochastic Specification and the Estimation of Share Equations. Journal of Econometrics. 1979;10:361–383. [Google Scholar]
  49. Ye X, Pendyala RM. A Model of Daily Time Use Allocation Using Fractional Logit Methodology. In: Mahmassani HS, editor. Transportation and Traffic Theory: Flow, Dynamics, and Human Interaction. Oxford: Pergamon, Elsevier Science Ltd; 2005. [Google Scholar]

RESOURCES