Skip to main content
Elsevier Sponsored Documents logoLink to Elsevier Sponsored Documents
. 2013 May;143(5):992–999. doi: 10.1016/j.jspi.2012.11.003

On conjugate families and Jeffreys priors for von Mises–Fisher distributions

Kurt Hornik a, Bettina Grün b,
PMCID: PMC3690539  PMID: 23805026

Abstract

This paper discusses characteristics of standard conjugate priors and their induced posteriors in Bayesian inference for von Mises–Fisher distributions, using either the canonical natural exponential family or the more commonly employed polar coordinate parameterizations. We analyze when standard conjugate priors as well as posteriors are proper, and investigate the Jeffreys prior for the von Mises–Fisher family. Finally, we characterize the proper distributions in the standard conjugate family of the (matrix-valued) von Mises–Fisher distributions on Stiefel manifolds.

Keywords: Bayesian inference, Conjugate prior, Jeffreys prior, von Mises–Fisher distribution

1. Introduction

A random unit length vector in Rd has a von Mises–Fisher (or Langevin, short: vMF) distribution with parameter θRd if its density with respect to the uniform distribution on the unit hypersphere Sd1={xRd:x=1} is given by

f(x|θ)=eθx/F10(;d/2;θ2/4),

where, using the rising factorial (ν)n=Γ(ν+n)/Γ(ν),

F10(;ν;z)=n=01(ν)nznn!=n=0Γ(ν)Γ(ν+n)znn!

is a generalized hypergeometric series and related to the modified Bessel function of the first kind Iν via

F10(;ν+1;κ2/4)=Iν(κ)Γ(ν+1)(κ/2)ν

(e.g., Mardia and Jupp, 1999, p. 168).

We note that the vMF distribution is commonly parameterized using polar coordinates, i.e., θ=κμ, where κ=θ and μSd1 are the concentration and mean direction parameters, respectively (if θ0, μ is uniquely determined as θ/θ). Using θ as the parameter, the family F of vMF distributions on Sd1 becomes a natural exponential family through the uniform distribution U on Sd1, commonly written as

f(x|θ)=eθxM(θ),

where in the vMF case, the cumulant transform M(θ) of U is given by

eM(θ)=Sd1eθxdU(x)=F10(;d/2;θ2/4).

Bayesian inference for the vMF distribution is first discussed in Mardia and El-Atoum (1976), who give conjugate priors for μ when κ is known, and derive the Jeffreys prior for the polar coordinates (μ, κ) parameterization. Guttorp and Lockhart (1988) introduce a Bayesian approach for finding the direction of a signal based on developing standard (e.g., Gutiérrez-Peña and Smith, 1997, Definition 3.1) conjugate priors for the von Mises (vM) distribution (i.e., for d=2) using the canonical (θ) parameterization. Damien and Walker (1999) present a full Bayesian analysis of circular data using the vM distribution by employing standard conjugate priors for the polar coordinates (μ, κ) parameterization, and developing a Gibbs sampler for this family of distributions. Nuñez-Antonio and Gutiérrez-Peña (2005) provide a full Bayesian analysis of directional (i.e., d2) data using the vMF distribution, again using standard (μ, κ) conjugate priors and obtaining samples from the posterior using a sampling-importance-resampling method found to outperform Gibbs sampling. Bangert et al. (2010) construct (possibly infinite) mixtures of vMF distributions using standard conjugate priors for the (μ, κ) parameterization and Dirichlet (process) priors for the mixing probabilities.

Interestingly, none of these references explicitly discuss when the employed priors (and respective posteriors) are actually proper, or whether the conjugate families obtained using the θ or (μ, κ) parameterizations are the same. In this paper, we settle these open issues, and also discuss Jeffreys priors for the general (d2) vMF family (Section 2). We also provide results for (matrix-valued) vMF distributions on Stiefel manifolds (Section 3).

2. Results

2.1. Propriety of priors from the standard conjugate family

In what follows, it will be convenient to write

Cd(κ)=1/F10(;d/2;κ2/4),

so that eM(θ)=Cd(θ). Let θ=θ(λ) be a suitable parameterization of θ. For a sample x1,,xn of independent, identically distributed (i.i.d.) observations from the vMF family F, the likelihood function for λ is given by

L(λ|s,n)=esθ(λ)nM(θ(λ))=Cd(θ(λ))nesθ(λ),

where s=x1++xn is the resultant of the sample. Following Gutiérrez-Peña and Smith (1997, Definition 3.1), the standard conjugate family for F relative to λ, denoted by Cλ(F), has densities

π(λ|s,ν)L(λ|s,ν).

Using such a prior with parameters s0 and ν0 will result in a posterior with parameters sn=s0+x1++xn and νn=ν0+n. As clearly

snνn=ν0ν0+ns0ν0+nν0+nx¯,

ν0 can be interpreted as the prior sample size, and sn/νn as a sample size weighted average of the “prior mean” s0/ν0 and the sample mean x¯.

The standard conjugate family Cθ(F) relative to the canonical parameter θ has several important properties, in particular the linear relationship between the posterior mean and the sample mean (Diaconis and Ylvisaker, 1979).

We note that the densities π(λ|s,ν) are usually taken relative to the Lebesgue measure, which does not quite fit the needs of the commonly used polar coordinates (μ, κ) parameterization of the vMF family F. Let us generally write η for the reference measure employed. Previous work using the (μ, κ) parameterization seem to take η as the product of the Lebesgue measure on [0,) (for κ) and the uniform distribution U on Sd1 (for μ), i.e., dηdκdU(μ). As for θ=κμ we have dθ=adκd1dκdU(μ) (where ad is the area of the unit hypersphere). The latter may be more natural as reference measure, turning the standard conjugate family relative to the polar coordinates (μ, κ) parameterization into the (obvious generalization) of what Gutiérrez-Peña and Smith (1997) call the DY-conjugate family for F relative to the parameterization.

Let Hλ;η denote the set of all hyperparameters s and ν for which π(λ|s,ν) is a proper distribution on the employed parameter space Λ (using η as reference measure), i.e.,

Hλ;η=(s,ν):ΛCd(θ(λ))νesθ(λ)dη(λ)<,

and let

J(α,β,ν)=0καCd(κ)νCd(βκ)dκ.

We have the following results.

Theorem 1

For the canonical parameterization θ of the vMF family and the Lebesgue measure as reference measure η,

Hθ;η={(s,ν):s<ν},

and the normalizing constant is the inverse of adJ(d1,s,ν).

In the following a parameter α is introduced which allows to cover both cases of reference measures when using the (μ, κ) parameterization: the Lebesgue measure (leading to α=d1) and the product of the Lebesgue measure on [0,) and the uniform distribution U on Sd1 which is employed in previous work (leading to α=0). Other choices for α lead to additional possible reference measures.

Theorem 2

For the polar coordinates (μ, κ) parameterization of the vMF family and the reference measure dη=καdκdU(μ) with α0,

Hκ,μ;η={(s,ν):s<νors=ν<12(α+1)/(d1)},

and the normalizing constant is the inverse of J(α,s,ν).

Note that if d2, 12(α+1)/(d1)>0 is equivalent to (d1)/2>α+1 or (d3)/2>α, which if α0 is only possible if d4. Thus, the set {(s,ν):s=ν<12(α+1)/(d1)} is non-empty only if d4 and α<(d3)/2, and clearly can only contain points for which s=ν<1.

For the proof, we use the following result.

Lemma 1

If α and β are nonnegative, J(α,β,ν)< if and only if β<ν or 0β=ν<12(α+1)/(d1).

Proof of Lemma 1

Using the asymptotic approximation Iν(κ)eκ/2πκ for κ and ν fixed (e.g., Abramowitz and Stegun, 1972, http://dlmf.nist.gov/10.40), we have

Cd(κ)κd/21/Id/21(κ)2πκ(d1)/2eκ.

Hence, for large κ, the integrand in J is “approximately proportional” to

κακν(d1)/2eνκ(βκ)(d1)/2eβκκα+(ν1)(d1)/2e(νβ)κ.

Thus, the integral diverges if ν<β, and converges if ν>β. If ν=β, convergence requires 1>α+(ν1)(d1)/2, or equivalently, ν1<2(α+1)/(d1) as asserted. □

Proof of Theorem 1

Transforming to polar coordinates θ=κμ, we obtain

RdCd(θ)νesθdθ=ad0Cd(κ)νSd1eκsμdU(μ)κd1dκ=ad0κd1Cd(κ)ν1Cd(κs)dκ=adJ(d1,s,ν),

interchanging the order of integration being justified by nonnegativity of the integrand. The assertion now follows from Lemma 1. □

Proof of Theorem 2

If dη=καdκdU(μ), we have

0Sd1Cd(κ)νeκsμκαdκdU(μ)=0καCd(κ)νCd(κs)dκ=J(α,s,ν),

whence the theorem follows by again using Lemma 1. □

We see that for the canonical parameterization, the hyperparameters giving proper distributions are the ones for which ν>0 and the “prior mean” s/ν lies in the interior of (the convex hull of) the unit hypersphere Sd1. This is not a coincidence: in fact, one can alternatively establish Theorem 1 (and equivalently, Theorem 2 for α=d1) without explicit convergence computations using the general results of Diaconis and Ylvisaker (1979), see also Gutiérrez-Peña and Smith (1997, Theorem 3.1). Let μ be a probability measure on (the Borel sets of) Rd with bounded support S and consider the natural exponential family through μ with density f(x|θ)=eθxM(θ), where eM(θ)=eθxdμ(x), and the standard conjugate family with densities π(θ|s,ν)=esθνM(θ) (with respect to the Lebesgue measure). As S is bounded and μ is finite, Θ={θ:M(θ)<}=Rd. Let X be the interior of the convex hull of S. Then by Theorem 1 of Diaconis and Ylvisaker (1979), if X is nonempty (and hence “the observation set is genuinely d-dimensional” Diaconis and Ylvisaker, 1979, p. 271) π(θ|s,ν) is proper if and only if ν>0 and s/νX. (The reference actually uses νs where we use s.) In the vMF case, S=Sd1, with convex hull the closed unit ball, and interior the open unit ball X={xRd:x<1}. Hence, π(θ|s,ν) is proper if and only if ν>0 and s/ν<1, or equivalently, s<ν, again establishing Theorem 1.

We also note that for αd1 and s<ν, Theorem 1 implies that

ad1Sd1Cd(κ)νeκsμκαdκdU(μ)ad1Sd1Cd(κ)νeκsμκd1dκdU(μ)=θ1Cd(θ)νesθdθ<,

so that s and ν give a proper conjugate distribution for the polar coordinates (μ, κ) parameterization with reference measure dη=καdκdU(μ). However, neither results for the case α>d1 nor necessity of the condition s<ν can be established using the general framework (and in fact, Theorem 2 shows that the condition is not necessary if d4 and 0α<(d3)/2).

2.2. Propriety of posteriors from improper standard conjugate priors

Quite interestingly, canonical priors employed in the literature are improper if a vague prior is intended (see for example Nuñez-Antonio and Gutiérrez-Peña, 2005, who use ν=0 and s=0). However, we note that if s=ν>0 and x1,,xnSd1, then s+x1++xns+x1++xnn+ν with equality if and only if x1==xn=s/ν which is a zero set for samples obtained from the vMF with fixed parameter θ. Hence intuitively, we expect that improper standard conjugate priors with s=ν>0 “almost always” yield proper posteriors. For the case s=ν=0 (as in the examples) we have x1++xnx1++xnn with equality if and only if x1==xn which is a zero set for samples obtained from the vMF with fixed parameter θ and n2.

This can be formalized as follows. Let π be the density (with respect to η) of a σ-finite measure on Λ and define μπ(n) on Ξ(n)=Sd1××Sd1 (n times), the space of all Sd1 valued samples of size n, via

μπ(n)(A)=ΛAf(x1|θ(λ))f(xn|θ(λ))dU(x1)dU(xn)π(λ)dη(λ)=AΛf(x1|θ(λ))f(xn|θ(λ))π(λ)dη(λ)dU(x1)dU(xn).

Writing gπ(x1,,xn)=Λf(x1|θ(λ))f(xn|θ(λ))π(λ)dη(λ),

μπ(n)(A)=Agπ(x1,,xn)dU(x1)dU(xn),

i.e., μπ(n) is a generalized “mixture” of the distribution of i.i.d. samples of size n from the vMF family. Let

An(s,ν)={(x1,,xn):s+x1++xn(n+ν)}.

Theorem 3

If s=ν>0 (s=ν=0), then μπ(n)(An(s,ν))=0 for all n1 (respectively, 2) and arbitrary π.

Proof of Theorem 3

Let ν>0. From the above, (x1,,xn)An(s,ν) if and only if x1==xn=s/ν. Clearly, for i.i.d. random variables (X1,,Xn) from the vMF distribution with parameter θ, P((X1,,Xn)An(s,ν)|θ)=An(s,ν)f(x1|θ)f(xn|θ)dU(x1)dU(xn)=0 and hence, μπ(n)(An(s,ν))=Λ0·π(λ)dη(λ)=0. As An(0,0) consists of all (x1,,xn) for which x1==xn, the assertion for the second case follows along the lines of the first case. □

If we use the canonical parameterization and π(θ)=eθsνM(θ), then by the above,

gπ(x1,,xn)=Θeθ(x1++xn)nM(θ)eθsνM(θ)dθ=adJ(d1,s+x1++xn,n+ν)

is infinite if and only if (x1,,xn)An(s,ν). If s=ν>0, this is a zero set under the product of uniforms on Ξ(n), and hence (again) μπ(n)(An(s,ν))=Agπ(x1,,xn)dU(x1)dU(xn)=0. On the other hand, if s>ν, then clearly μπ(n)(An(s,ν))=: in this sense, it is always possible to obtain improper posteriors when employing an improper standard conjugate prior with s>ν (we notice however that such priors are admittedly “strange”, as the corresponding prior sample means s/ν are outside the unit ball and hence “impossible”).

If X has a vMF distribution with parameter θ, the theory of regular exponential models (cf., e.g., Mardia and Jupp, 1999, pp. 32–33) implies that Eθ(X)=M(θ)/θ so that

Eθ(X)=dlog(Cd(θ))dθθθ=Cd(θ)Cd(θ)θθ=Ad(θ)θθ,

where one can show that the logarithmic derivative Ad of 1/Cd satisfies (Schou, 1978)

Ad(κ)=Id/2(κ)Id/21(κ),

that

Ad(κ)=1dκ1d2(d+2)κ3+O(κ5),

as κ0 (so that Ad(κ)/κ is in fact C provided we take its value at zero to be 1/d), and that

Ad(κ)=1(d1)21κ+(d1)(d3)81κ2+O(κ3)

as κ. Hence, for i.i.d. random variables X1,,Xn from the vMF distribution with parameter θ, X¯nEθ(X)=Ad(θ)θ/θ which is less than one in length, and hence

Snνn=s+X1++Xnn+νEθ(X)<1,

with probability one as n. Thus, if we write

pn(θ|s,ν)=Pθ((X1,,Xn)An(s,ν)),

pn(θ|s,ν)0 as n for all θ, and using continuity arguments one can easily see that this convergence is uniform on compact subsets of Θ=Rd. On the other hand, Eθ(X)1 for θ, so the convergence cannot be uniform over Θ. It would be very interesting to find the rate at which supθκpn(θ|s,ν) tends to zero, which would then allow one to characterize the improper prior densities π for which μπ(n)(An(s,ν))=Θpn(θ|s,ν)π(θ)dθ0 as n.

2.3. Jeffreys prior

A commonly suggested non-informative prior is the Jeffreys prior (Jeffreys, 1961), defined as the square root of the determinant of the Fisher information matrix relative to the parameterization employed. When using the canonical parameter, again by the theory of regular exponential models, I(θ)=varθ(X)=2M(θ)/θθ so that

I(θ)=Ad(θ)θθθθ+Ad(θ)1θIdθ1θ2θθ=Ad(θ)θθθθ+Ad(θ)θIdθθθθ=Ad(θ)θId+Ad(θ)Ad(θ)θθθθθ,

where Id denotes the d-dimensional identity matrix. By a well known result from linear algebra, det(Id+βvv)=1+βvv and thus det(γId+βvv)=γd(1+βvv/γ) and in particular if v=1, det(γId+βvv)=γd(1+β/γ)=γd1(β+γ), so that

det(I(θ))=Ad(θ)θd1Ad(θ),

generalizing the result obtained in Guttorp and Lockhart (1988) for the case d=2 (the sign in the reference is not correct).

Theorem 4

The Jeffreys prior π(θ)det(I(θ)) for the canonical parameter θ only depends on θ and behaves like 1/θ(d+1)/2 for θ.

Proof of Theorem 4

The first assertion is immediate by observing that det(I(θ)) depends on θ only via its length. To obtain the asymptotic behavior, we can use the asymptotics of Ad(κ) and the fact that Ad(κ)=1Ad(κ)(Ad(κ)+(d1)/κ) (Schou, 1978). Thus, as κ,

Ad(κ)11d12κ1d12κ+d1κ=11d12κ1+d12κ=(d1)24κ2,

such that for θ,

det(I(θ))1θd1(d1)24θ2=constθd+1.

We note that the Jeffreys prior “looks different” from the densities employed in the standard conjugate family Cθ(F) relative to the canonical parameter. Following Gutiérrez-Peña and Smith (1997), one can rigorously establish that it is not contained in this family by verifying that the skewness vector log(det(I(θ)))/θ of the vMF family is not linear in its mean parameter, which is straightforward from the above expression for det(I(θ)).

The Jeffreys prior with respect to the canonical parameterization is given by

π(θ)Ad(θ)θd1Ad(θ)=Ad(θ)θ(d1)/21Ad(θ)Ad(θ)+(d1)θ1/2.

If the polar coordinates (μ, κ) parameterization is employed two alternative parameterizations are possible for the mean direction parameter μ in order to obtain an unrestricted set of parameters. The (μ0,κ) parameterization with μ0=(μ1,,μd1) consists of the first d1 dimensions of the mean direction parameter μ and the (m,κ) parameterization uses the spherical polar coordinates for μ with m=(ϕ1,,ϕd1). If these parameterizations are used the Jeffreys prior derived for the canonical parameterization needs to be multiplied with the Jacobians which are given by

κd1/μd,

for the (μ0,κ) parameterization and

κd1j=1d2sin(ϕj)d1j,

for the (m,κ) parameterization. Note that the latter is also given in Mardia and El-Atoum (1976), which should have κ(d1)Ad(κ)κAd2(κ) instead of κAd(κ)κAd2(κ) and needs d1j instead of d2 in the exponents of the sinuses.

Clearly, the Jeffreys prior is not proper. The following shows that “almost all” posteriors obtained from it (and in fact, from arbitrary possibly improper priors which increase at most polynomially in θ) are proper for samples of size n2.

Theorem 5

Consider the canonical parameterization of the vMF family with the Lebesgue reference measure. Let π(θ) be O(θγ) for some finite γ0 as θ and Bn(π)={(x1,,xn):Θf(x1|θ)f(xn|θ)π(θ)dθ<}. Then for all n2, μπ(n)(Bn(π)c)=0.

Proof of Theorem 5

Writing s=x1++xn, we have

θ1f(x1|θ)f(xn|θ)π(θ)dθ=θ1eθ(x1++xn)nM(θ)π(θ)dθconst1Sd1eκsμCd(κ)nκγκd1dU(μ)dκ=const1κγ+d1Cd(κ)nCd(κs)dκ.

By Lemma 1, this is finite provided that s<n. Hence, Bn(π)cAn(0,0) which is a zero set under the product of uniforms on Ξ(n) provided that n2. □

2.4. Propriety of prior and posterior distributions in applications

Guttorp and Lockhart (1988) perform a full Bayesian analysis employing the canonical parameterization for 2-dimensional data. Rather than using the standard conjugate prior for θ, they use a flat prior on μ and the conjugate prior they derived for κ with μ known. Note that for the conjugate prior for κ to be proper, the same conditions on s and ν need to be satisfied as for the conjugate prior for θ. In order to parameterize the κ prior only the length of s and ν need to be specified. For their application, Guttorp and Lockhart (1988) use three different prior distributions for κ: a data-based, a high precision and a low precision prior. In all three cases the priors for κ are proper because s<ν. This is also clear by construction: the parameters of the priors are determined by specifying constraints for the moments of the prior distribution or quantities derived from the moments.

Damien and Walker (1999) employ the polar coordinates (μ, κ) parameterization with the Lebesque measure as reference measure, i.e., α=0. They use two different prior distributions in their two numerical examples. In the first example, they set all prior parameters to zero. This is the same prior Nuñez-Antonio and Gutiérrez-Peña (2005) use in their examples and refer to as vague prior. Given the updates for the posterior parameters as well as the interpretation of ν as the prior sample size, this seems to be an obvious choice. This prior is improper, but as shown in Section 2.2, posteriors will be proper almost surely for samples of size n2. In their second example, Damien and Walker (1999) use a flat prior for μ and a conjugate prior for κ with μ known. Referring to the low precision prior in Guttorp and Lockhart (1988), s=ν=5 are employed as parameters. The low precision prior in Guttorp and Lockhart (1988) actually is equal to s=9.8824 and ν=10, while the values for the high precision prior are s=4.99893 and ν=5. Interpreting ν as the prior sample size, larger values of ν imply more informative priors. However, the precision induced by the prior will depend on the average length given by s/ν. The closer this value is to 1 the higher is the precision induced by the prior. Using an approximation for large κ, Guttorp and Lockhart (1988) derived that the prior mean and variance of the precision parameter κ depend on ν as well as the difference νs. By rounding s to the same value as ν, Damien and Walker (1999) use an improper prior and the interpretation as a low precision prior, as induced by the prior mean and standard deviation, obviously is lost. As shown in Section 2.2, the posterior from this prior is almost surely proper for samples of size n1.

Bangert et al. (2010) use conditionally conjugate priors for μ and κ. The prior for μ is a von Mises–Fisher distribution with precision parameter equal to 0.1 and mean parameter equal to the mean direction of the data. For κ they use the conjugate prior for known μ. Setting the prior parameters equal to ν=5 and s=4.7, they employ a proper prior.

To sum up, these previous applications indicate that the non-informative but improper prior with s=ν=0 seems to be an obvious choice if no prior information is available. This seems to be unproblematic because the posteriors will be almost surely proper for sample sizes n2 by Theorem 3. If prior information on the precision parameter κ is to be included, a flat prior for μ is employed and the conjugate prior with known μ for κ. Another possibility, when using Gibbs sampling for estimation, is to employ conditionally conjugate priors. In general the parameters of the prior of κ are chosen to reflect the prior information available for the moments of κ, which leads to proper priors.

3. Extensions

The vMF family on Sd1 can straightforwardly be generalized to the vMF family on the Stiefel manifold Vk(Rd), the set of orthogonal k-frames in Rd, or equivalently,

Vk(Rd)={XRd×k:XX=Ik},

so that V1(Rd) corresponds to Sd1. The vMF family on Vk(Rd) has densities

f(X|A)=etr(AX)M(A),

with respect to the uniform distribution U on Vk(Rd), where

eM(A)=Vk(Rd)etr(AX)dU(X)=F10(;d/2;AA/4)

is a generalized hypergeometric function with matrix argument (e.g., Mardia and Jupp, 1999, p. 289). This family of distributions is useful as a probability distribution over orthonormal matrices and for example Hoff (2009) indicates that it arises as a posterior distribution for the orthonormal matrices in factor analysis when uniform priors are used. For a further discussion of this family of distributions in relation to orientation statistics see Downs (1972) and Khatri and Mardia (1977).

Clearly, Xvec(X) defines a one-to-one correspondence between Rd×k and Rdk with tr(XA)=vec(X)vec(A). The standard conjugate family for the vMF family on Vk(Rd) (relative to the canonical parameter) is thus given by the family of densities

p(A|S,ν)etr(SA)νM(A).

Let S2 denote the spectral norm (matrix 2-norm, the largest singular value) of S.

Theorem 6

The distributions in the standard conjugate family of the vMF family on the Stiefel manifold Vk(Rd) are proper if and only if S2<ν.

Proof of Theorem 6

The support of U is S=Vk(Rd), the convex hull of which is the closed unit ball in the spectral norm (e.g., Journée et al., 2010 or Gallivan and Absil, 2010), and hence has non-empty interior

X={XRd×k:X2<1}={XRd×k:XXIk}.

Using Theorem 1 of Diaconis and Ylvisaker (1979), the standard conjugate distributions are proper if and only if ν>0 and S/νX, or equivalently, if and only if S2<ν. □

The matrix vMF distributions are typically parameterized using the canonical parameter A. Alternatively, the analogue to the polar coordinates (μ, κ) parameterization in the vector case is using the (right) polar decomposition of A=MK, where the polar part (or orientation) M is in the Stiefel manifold Vk(Rd) and the elliptical part (or concentration) K is a symmetric, non-negative definite matrix (e.g., Mardia and Jupp, 1999, p. 286).

If A has full rank, K is the unique symmetric matrix root of AA, and (e.g., Cadet, 1996, adjusting for the different normalizations employed) dA=c(K)dKdU(M), where c(K)=ad,kdet(K)dki<j(λi+λj) with ad,k the (generalized) volume of Vk(Rd) and λ1,,λk>0 the eigenvalues of K. Hence,

Rd×ketr(SA)νM(A)dA=K0Vk(Rd)etr(SMK)(F10(;d/2;K2/4))νc(K)dKdU(M)=K0c(K)(F10(;d/2;K2/4))νVk(Rd)etr(KSM)dU(M)dK=K0c(K)F10(;d/2;KSSK/4)F10(;d/2;K2/4)νdK.

Thus, if we consider the standard conjugate family of the matrix vMF family on the Stiefel manifold Vk(Rd) relative to the polar coordinates parameterization A=MK with elements π(M,K|S,ν)etr(SMK)/F10(;d/2;K2/4)ν, and reference measures of the form c(K)αdKdU(M), then as discussed in Section 2 for the polar parameterization of the vector vMF distribution, if 0α1 Theorem 6 implies that distributions in this conjugate family are proper provided that S2<ν. Again, necessity of this condition for such values of α, or the characterization of the hyperparameters giving proper distributions if α>1 cannot be established.

For this, one needs to be able to characterize S and ν (and α) for which

Jk(α,S,ν)=K0c(K)αF10(;d/2;KSSK/4)F10(;d/2;K2/4)νdK<,

which seems quite challenging, requiring suitable “large K” asymptotics for F10(;d/2;K2/4) and F10(;d/2;KSSK/4). We note that Butler and Wood (2003) give Laplace approximations for F10 (and corresponding Bessel functions) of matrix arguments (but do not formally establish validity as an asymptotic approximation). Muirhead (1978, p. 22) gives an asymptotic approximation for F10(;d/2;AA/4) for the case where all singular values of A are large. For the above, a generalization to the case where some singular values are large is needed. We leave this for future research.

Acknowledgments

This research was funded by the Austrian Science Fund (FWF): V170-N18.

Contributor Information

Kurt Hornik, Email: Kurt.Hornik@wu.ac.at.

Bettina Grün, Email: Bettina.Gruen@jku.at.

References

  1. Abramowitz M., Stegun I.A. Dover; New York: 1972. Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables. [Google Scholar]
  2. Bangert, M., Hennig, P., Oelfke, U., 2010. Using an infinite von Mises–Fisher mixture model to cluster treatment beam directions in external radiation therapy. In: Proceedings of the Ninth International Conference on Machine Learning and Applications (ICMLA), pp. 746–751.
  3. Butler R.W., Wood A.T.A. Laplace approximation for Bessel functions of matrix argument. Journal of Computational and Applied Mathematics. 2003;155:359–382. [Google Scholar]
  4. Cadet A. Polar coordinates in Rnp; application to the computation of the Wishart and beta laws. Sankhya: The Indian Journal of Statistics Series A. 1996;58(1):101–114. [Google Scholar]
  5. Damien P., Walker S. A full Bayesian analysis of circular data using the von Mises distribution. The Canadian Journal of Statistics. 1999;27(2):291–298. [Google Scholar]
  6. Diaconis P., Ylvisaker D. Conjugate priors for exponential densities. The Annals of Statistics. 1979;7(2):269–281. [Google Scholar]
  7. Downs T.D. Orientation statistics. Biometrika. 1972;59(3):665–676. [Google Scholar]
  8. Gallivan, K.A., Absil, P.-A., 2010. Note on the Convex Hull of the Stiefel Manifold. Technical Report. FSU 10-06, Department of Mathematics, Florida State University. URL http://www.math.fsu.edu/∼aluffi/archive/paper386.pdf.
  9. Gutiérrez-Peña E, Smith A.F.M. Exponential and Bayesian conjugate families: review and extensions. Test. 1997;6:1–90. [Google Scholar]
  10. Guttorp P., Lockhart R.A. Finding the location of a signal: a Bayesian analysis. Journal of the American Statistical Association. 1988;83(402):322–330. [Google Scholar]
  11. Hoff P.D. Simulation of the matrix Bingham-von Mises–Fisher distribution, with applications to multivariate and relational data. Journal of Computational and Graphical Statistics. 2009;18:438–456. [Google Scholar]
  12. Jeffreys H. 3rd ed. Oxford University Press; Oxford: 1961. Theory of Probability. [Google Scholar]
  13. Journée M., Nesterov Y., Richtárik P., Sepulchre R. Generalized power method for sparse principal component analysis. Journal of Machine Learning Research. 2010;11:517–553. [Google Scholar]
  14. Khatri C.G., Mardia K.V. The von Mises–Fisher matrix distribution in orientation statistics. Journal of the Royal Statistical Society B. 1977;39(1):95–106. [Google Scholar]
  15. Mardia K.V., El-Atoum S.A.M. Inference for the von Mises–Fisher distribution. Biometrika. 1976;63(1):203–206. [Google Scholar]
  16. Mardia K.V., Jupp P.E. Wiley; 1999. Directional Statistics. Probability and Statistics. [Google Scholar]
  17. Muirhead R.J. Latent roots and matrix variates: a review of some asymptotic results. The Annals of Statistics. 1978;2(1):5–33. [Google Scholar]
  18. Nuñez-Antonio G., Gutiérrez-Peña E. A Bayesian analysis of directional data using the von Mises–Fisher distribution. Communications in Statistics—Simulation and Computation. 2005;34(4):989–999. [Google Scholar]
  19. Schou G. Estimation of the concentration parameter in von Mises–Fisher distributions. Biometrika. 1978;65(1):369–377. [Google Scholar]

RESOURCES