Nonparametric estimation of multivariate scale mixtures of uniform densities

Marios G Pavlides; Jon A Wellner

doi:10.1016/j.jmva.2012.01.001

. Author manuscript; available in PMC: 2012 Nov 1.

Published in final edited form as: J Multivar Anal. 2012 Jan 10;107:71–89. doi: 10.1016/j.jmva.2012.01.001

Nonparametric estimation of multivariate scale mixtures of uniform densities

Marios G Pavlides ^a,^*, Jon A Wellner ^b

PMCID: PMC3318987 NIHMSID: NIHMS359092 PMID: 22485055

Abstract

Suppose that U = (U₁, … , U_d) has a Uniform ([0, 1]^d) distribution, that Y = (Y₁, … , Y_d) has the distribution G on $R_{+}^{d}$ , and let X = (X₁, … , X_d) = (U₁Y₁, … , U_dY_d). The resulting class of distributions of X (as G varies over all distributions on $R_{+}^{d}$ ) is called the Scale Mixture of Uniforms class of distributions, and the corresponding class of densities on $R_{+}^{d}$ is denoted by $F_{SMU} (d)$ . We study maximum likelihood estimation in the family $F_{SMU} (d)$ . We prove existence of the MLE, establish Fenchel characterizations, and prove strong consistency of the almost surely unique maximum likelihood estimator (MLE) in $F_{SMU} (d)$ . We also provide an asymptotic minimax lower bound for estimating the functional f ↦ f(x) under reasonable differentiability assumptions on f ∈ $F_{SMU} (d)$ in a neighborhood of x. We conclude the paper with discussion, conjectures and open problems pertaining to global and local rates of convergence of the MLE.

Keywords: Nonparametric estimation, Monotonicity, Multivariate, Minimax, Consistency, Uniform, Mixture

1. Introduction and summary

Fix a non-negative integer k, and suppose that X₁, … , X_n are i.i.d. random variables distributed according to a density in the convex family of k-monotone densities (with respect to Lebesgue measure) on (0, ∞):

F_{k} ≔ {f_{k, G} (\cdot) \equiv \int_{0}^{\infty} k \frac{{(y - \cdot)}_{+}^{k - 1}}{y^{k}} d G (y) ∣ G \in g_{1}},

(1.1)

where $g_{1}$ will denote the set of all distribution functions on (0, ∞) grounded at 0. Here, we use the notation x₊ ≡ x · 1_[x≥0] for any $x \in R$ . It has been shown by Williamson [59] that the family $F_{k}$ is identifiably indexed by $g_{1}$ . In other words, if G₁, G₂ are distinct elements in $g_{1}$ , then f_k,G₁ (·) and f_k,G₂ (·) differ on a Lebesgue non-null set. Note that $F_{k}$ is exactly the collection of all scale mixtures of Beta (1, k) densities.

The Beta (1, 1) distribution is the standard uniform distribution, U(0, 1). Therefore, the class $F_{1}$ coincides with the class of all scale mixtures of uniform densities on (0, ∞). A well-known theorem by Khintchine (see, e.g., [16, p.158]) asserts that the class of densities on (0, ∞) with concave distribution functions is one and the same with our class $F_{1}$ . It can be seen that $F_{1}$ is also the class of all upper semi-continuous, non-increasing densities on (0, ∞). This class is induced by order restrictions, a term we use to explicitly mean that there exists a partial ordering (⪡) on the common support $X$ of the densities in $F_{1}$ such that f ∈ $F_{1}$ if and only if f is isotone with respect to this ordering: i.e., f ∈ $F_{1}$ if and only if f (x) ≤ f (y) whenever x , y ∈ $X$ such that x ⪡ y. In this case, (⪡) is the natural partial ordering, ≥, on (0, ∞).

Non-increasing, upper semi-continuous densities (in short, monotone densities) arise naturally via connections with renewal theory and uniform mixing (see, e.g., [60]). Maximum likelihood estimation of monotone densities on (0, ∞) was initiated by Grenander [18,19], with related work by Ayer et al. [3], Brunk [11], van Eeden [51-55]. Asymptotic theory of the MLE in $F_{1}$ (the Grenander estimator) was developed by Prakasa Rao [44] with later contributions by [20,21,8,9,30]. See [4] for descriptions of the behavior of the Grenander estimator at zero.

Nonparametric estimation in families of densities described by order restrictions goes back at least to the work of [18,19,11,12,45], with further development by Wegman [56-58], Sager [48,49]. Also see the books by Barlow et al. [5] and Robertson et al. [46]. [40-43] addressed estimation in various order restricted classes of multivariate densities from the perspective of the excess mass approach studied previously by e.g., [48,49,36]. Polonik shows that (under reasonable assumptions) the MLE in such classes exists and coincides with an estimator he constructs and calls the silhouette. Forcing the elements of the class to be upper semi-continuous, the MLE is seen to be unique. Brunk [11] also gives a graphical construction of the maximum likelihood estimator, and establishes L₁-consistency of the MLE.

In this paper, our goal is to extend the notion of “monotone densities” to higher dimensions; i.e., to densities on (0, ∞)^d with d > 1. Such an extension is not unique: for example, we may consider the family, $F_{BDD} (d)$ , of “block-decreasing densities” (a term coined by Biau and Devroye [6]) that contains all upper-semicontinuous densities on (0, ∞)^d that are non-increasing in each coordinate, while keeping all other coordinates fixed. This class was perhaps first introduced by Robertson [45]. The particular proper subclass of $F_{BDD} (d)$ studied here is the family $F_{SMU} (d)$ of all multivariate scale mixtures of uniform densities; i.e., the family of upper semi-continuous densities on (0, ∞)^d of the form

f_{G} (x) = \int_{{(0, \infty)}^{d}} (\frac{1}{∣ y ∣} 1_{(0, y]} (x)) d G (y), x \in {(0, \infty)}^{d}

(1.2)

for some $G \in g_{d}$ , the set of all distribution functions on (0, ∞)^d that are grounded (zero) at 0; here we use the notation $∣ y ∣ \equiv \prod_{i = 1}^{d} y_{i}$ for y = ( y₁, … , y_d)’ ∈ (0, ∞)^d. For any fixed $G \in g_{d}$ , it is clear that if Y = (Y₁, … , Y_d)’ is distributed according to G on (0, ∞)^d and if U₁, … , U_d are i.i.d. U(0, 1) (and independent of Y), then the vector X := (U₁Y₁, … , U_dY_d) is distributed according to f_G(·) on (0, ∞)^d.

Whereas the family $F_{BDD} (d)$ is characterized by order restrictions (and thus the results by Polonik apply), its subclass $F_{SMU}$ is not; as will be made more explicit in Section 2, densities in the class $F_{SMU}$ also satisfy non-negativity restrictions on their d-dimensional differences around all rectangles. Because of this additional shape restriction, estimation in this family requires separate treatment

A univariate parallelism to the latter point would be to consider the family $F_{2}$ in (1.1), induced by mixtures of triangular densities; this class can easily be seen to be exactly the class of all non-increasing, convex (and hence continuous) densities on (0, ∞). Thus $F_{2} \subset F_{1}$ is not an order-constrained class of densities, in contrast to its superclass $F_{1}$ . Convex densities arise in connection with Poisson process models for bird migration and scale mixtures of triangular densities (see, e.g., [26,2,32]). Estimation of non-increasing, convex densities on (0, ∞) was apparently initiated by Anevski [1] and was further pursued by Anevski [2] and Jongbloed [28]. The asymptotic distribution theory and further characterizations of the nonparametric MLE of such a density and its first derivative at a fixed point (both under reasonable assumptions) was obtained by Groeneboom et al. [24,25]. These authors show that the local rate of convergence of the MLE of the functional f ↦ f (x) is of the order n^⅖, whereas the Grenander estimator (the MLE in $F_{1}$ ) converges locally at the rate of only n^⅓.

The developments here have several motivations. One of these is to provide a multivariate family of shape-constrained densities with convergence rates for reasonable estimators which are (nearly) independent of the dimension d of the underlying space. As will be seen from the lower bound calculations in Section 4, it seems that the SMU class studied here may provide such a class. Another motivation comes from problems concerning multivariate analogues of interval censored data; see e.g. [27,61,62]. These apparently quite different models involve very similar mathematical considerations, and it might be helpful to develop methods for multivariate interval censored data problems by first studying the somewhat simpler SMU model.

Here is an outline of the remainder of the present paper. In Section 2, we provide characterizations of the family $F_{SMU} (d)$ that will prove useful in the sequel. Section 3 addresses existence, strong, pointwise consistency as well as L₁ and Hellinger consistency of a sequence of maximum likelihood estimators in $F_{SMU} (d)$ . In Section 4, we derive a local asymptotic minimax lower bound for estimation of f (x) at a fixed point x under for which f satisfies ∂^df (x)/(∂x₁ … ∂x_d) ≠ 0. The lower bound entails a rate of convergence of n^⅓ for all dimensions d and yields a constant depending on f which reduces to the known lower bound constant for d = 1. The paper concludes in Section 5 with a discussion of conjectures and open problems related with both the local (pointwise) and the global (L₁ and Hellinger) rates of convergence of the MLE in $F_{SMU} (d)$ .

2. Properties of the Scale Mixtures of Uniform family of densities

2.1. Properties of $F_{S M U} (d)$

A density function, f , on (0, ∞)^d will be called a (multivariate) Scale Mixture of Uniform densities if there exists a distribution function, G, on (0, ∞)^d such that

f (x) = f_{G} (x) = \int_{{(0, \infty)}^{d}} \frac{1}{∣ v ∣} 1_{(0, v]} (x) d G (v)

(2.1)

= \int_{v \geq x} \frac{1}{∣ v ∣} d G (v) for all x \in {(0, \infty)}^{d} .

(2.2)

It is clear from (2.2) that a SMU density is also a block-decreasing density: f_G(·) is non-increasing in each coordinate, while keeping all other coordinates fixed. Also, the map G ↦ f_G is identifiable in the following sense: if G₁ ≠ G₂, then f_G1 ≠ f_G2 on a set of positive Lebesgue measure; also see Theorem 2.3 below. The following lemma gives a formal statement of a slightly more general result. The proof is standard.

Lemma 2.1

Two upper semi-continuous and block-decreasing functions f and g on $R^{d}$ differ nowhere in the interior of their support or else on a Lebesgue non-negligible set.

The distribution function F_G corresponding to X ~ f_G is given by

F_{G} (x) = \int_{{(0, \infty)}^{d}} \frac{∣ x \land v ∣}{∣ v ∣} d G (v),

(2.3)

where ≤ denotes the natural partial ordering on $R^{d}$ , while

x \land v \equiv (x_{1}, \dots, x_{d}) \land (v_{1}, \dots, v_{d}) = (\min {x_{1}, v_{1}}, \dots, \min {x_{d}, v_{d}}),

and x∨ν ≡ (x1, … , x_d)∨(ν₁, … , ν_d) = (max{x₁, ν₁}, … , max{x_d, ν_d}). The distribution function F_G of X is generally not concave when d > 1, unlike the case when d = 1. An SMU density (and a block-decreasing density, in general) can possibly diverge at the origin, whereas the pointwise bound f (x) ≤ 1/|x| holds since, for x ∈ (0, ∞)^d we have

1 = \int_{{(0, \infty)}^{d}} f (y) dy \geq \int_{(0, x]} f (y) d y \geq ∣ x ∣ f (x) .

Further, a d-dimensional analogue of the proof of [13, Theorem 6.2, p. 173] can be used to show that

\lim_{∣ x ∣ \to \infty} {∣ x ∣ f (x)} = \lim_{x ↓ 0} {∣ x ∣ f (x)} = 0,

(2.4)

whenever f is a block-decreasing density on (0, ∞)^d.

For any two points x, y ∈ [0, ∞)^d, such that x ≤ y, we write [x, y] ≡ [x₁, y₁] × … × [x_d, y_d], [x, y) ≡ [x₁, y₁) × … × [x_d, y_d), (x, y] ≡ (x₁, y₁] × … × (x_d, y_d], (x, y) ≡ (x₁, y₁) × … × (x_d, y_d) for the natural closed, lower-closed upper open, lower open upper closed, and open rectangles respectively. Note that the closed rectangle [x, y] has (at most) 2^d vertices, the points u = (u₁, … , u_d) where each u_i is either x_i or y_i. Following [7], we write sgn_[x,y](u) ∈ {−1, 1}, the signum of the vertex u, according as the number of i, 1 ≤ i ≤ d, satisfying u_i = x_i is odd or even respectively.

Thus any two vertices defining an edge of the rectangle have alternating signs. Then, if u = (u₁, … , u_d) is some vertex of [x, y] and δ ∈ {−1,+1} is its signum, then (δ, u) is an element of the set

Δ_{d} [x, y] = {({(- 1)}^{\sum_{i = 1}^{d} {1_{[u_{i} = x_{i}]}}}, u) ∣ u \in {x_{1}, y_{1}} \times \dots \times {x_{d}, y_{d}}} .

Definition 2.1

For an upper semicontinuous and coordinatewise decreasing function g : (0, ∞)^d → [0, ∞) define the g-volume of a (possibly degenerate) rectangle [x, y) by:

V_{g} [x, y) = \sum_{(δ, u) \in Δ_{d} [x, y]} {δ g (u)},

(2.5)

provided that g is defined and is finite for all u in the summand. Correspondingly, for an upper semicontinuous and coordinatewise increasing function g : (0, ∞)^d → [0, ∞), we define the g-volume of a rectangle (x, y] by the sum on the right side of (2.5).

It is easily seen that for an SMU density, f_G, the f_G-volume of any rectangle [x, y) is always of the sign (−1)^d: indeed, consider (2.2) and observe that

{(- 1)}^{d} V_{f G} [x, y) = \int_{[x, y)} \frac{1}{∣ v ∣} d G (v) \geq 0 .

(2.6)

From (2.6), or, alternatively, from the fact that the class of sets [x, y) is a π-system which generates the Borel σ-field of subsets of [0, ∞)^d and then extending as in [7], it is clear that (−1)^dV_f extends uniquely to a (non-negative) measure on the Borel σ-field $B_{+}^{d} = B^{d} \cap {(0, \infty)}^{d}$ given by

{(- 1)}^{d} V_{f} (A) = \int_{A} \frac{1}{∣ v ∣} d G (v) for A \in B_{+}^{d};

in particular,

{(- 1)}^{d} V_{f} (x, y] = \int_{(x, y]} \frac{1}{∣ v ∣} d G (v) .

This argument extends easily to an arbitrary upper semicontinuous function g with the (−1)^dg-volumes of all rectangles [x, y) non-negative.

Lemma 2.2

Suppose that g is a non-negative, upper semi-continuous function satisfying (−1)^dV_g [x, y) ≥ 0 for all lower-closed upper open rectangles [x, y), and vanishing if any coordinate tends to ∞. Then (−1)^dV_g can be extended to a countably additive measure on $B_{+}^{d}$ .

Of course it is easy to exhibit a block-decreasing density that is not an SMU density: consider the uniform density on the closed triangle in $R_{+}^{2}$ with vertices (0, 0), (0, 1) and (1, 0). Then,

{(- 1)}^{2} V_{f} [(1 ∕ 8, 1 ∕ 8), (1 ∕ 2, 3 ∕ 4)) = - 2 < 0,

showing that this density is not an SMU density, even though it is block-decreasing.

The following theorem establishes identifiability of the mixing distribution G as well as providing a useful characterization of SMU densities.

Theorem 2.3

For the class of SMU densities $F_{S M U} (d) = {f_{G} : G \in g_{d}}$ with f_G as given in (2.1), f ∈ $F_{S M U} (d)$ if and only if f ≡ f_G, where G ∈ G_d is given by
$G (x) = \int_{{(0, \infty)}^{d}} {(- 1)}^{d} V_{f} (u, x] \cdot 1_{[u \leq x]} d u .$ (2.7)
Thus there is a one-to-one correspondence between $G \in g_{d}$ and $f_{G} \in F_{S M U} (d)$ .
Suppose that the Lebesgue density f on (0, ∞)^d is such that it converges to zero in each coordinate, while keeping all other coordinates fixed. Then, f is an SMU density if and only if (−1)^dV_f [x, y) ≥ 0 for all 0 ≤ x ≤ y.

Proof. (a) Suppose that f ≡ f_G, for $G \in g_{d}$ (recall that this implies that G(0) = 0), is an SMU density evaluated at an arbitrary x ∈ (0, ∞)^d as:

f (x) = \int_{{(0, \infty)}^{d}} \frac{1}{∣ y ∣} 1_{(0, x]} d G (y) = \int_{y_{1} \geq x_{1}} \dots \int_{y_{d} \geq x_{d}} \frac{1}{∣ y ∣} d G (y),

(2.8)

so that df (x) = (−1)^d|x|⁻¹ dG(x) and thus,

\begin{matrix} G (x) & = \int_{{(0, \infty)}^{d}} 1_{(0, x]} (y) ∣ y ∣ d {{(- 1)}^{d} f (y)} \\ = \int_{(0, x]} \int_{(0, x]} 1_{(0, y]} (u) d u d {{(- 1)}^{d} f (y)} \\ = \int_{(0, x]} {\int_{y \in (u, x]} d {{(- 1)}^{d} f (y)}} d u \\ = \int_{(0, x]} {(- 1)}^{d} V_{f} (u, x] d u, \end{matrix}

where the second to last equality follows by Fubini–Tonelli.

We will now show that G is unique: suppose that (2.8) above holds for $G = G_{i} \in g_{d}$ and i = 1, 2. Recall that this implies that G₁(0) = G₂(0) = 0 and, thus, G₀(·) ≡ G₁(·) – G₂(·) is such that G₀(0) = 0, ∫_{(0, ∞)^d} G₀(x) dx = 0 and

0 = \int_{{(0, \infty)}^{d}} \frac{1}{∣ y ∣} 1_{(0, x]} d G_{0} (y) = \int_{(0, x]} \frac{1}{∣ y ∣} d G_{0} (y)

(2.9)

holds for all x ∈ (0, ∞)^d and, thus, necessarily G₀(x) has to be independent of x and therefore everywhere equal to its value at 0: G₀(0) = 0. This completes the assertion of uniqueness, since G₁ ≡ G₂.

(b) If f is in $F_{SMU}$ , there exists $G \in g_{d}$ such that

f (x) = \int_{{(0, \infty)}^{d}} \frac{1}{∣ y ∣} 1_{(0, y]} (x) d G (y) = \int_{y \geq x} \frac{1}{∣ y ∣} d G (y),

so that it is easily seen that (−1)^dV_f [x, y) = ∫_[x,y) |y|⁻¹ dG(y) ≥ 0 holds true for all 0 ≤ x ≤ y.

On the other hand, assume that the Lebesgue density f is such that it converges to zero in each coordinate, while keeping all other coordinates fixed, and satisfies (−1)^dV_f [x, y] ≥ 0 for all 0 ≤ x ≤ y. By Lemma 2.2, this implies that for x₁ ≤ x₂ ≤ x, elements of (0, ∞)^d, we have (−1)^dV_f [x₁, x) ≥ (−1)^dV_f [x₂, x) and, letting x → ∞, this yields f (x₁) ≥ f (x₂) because we assumed that f vanishes as any one of its coordinates diverges to infinity, so that V_f [x_i, x) → (−1)^df (x_i) for i ∈ {1, 2}. Thus, f is block-decreasing.

Hence, by appealing to part (a), it thus suffices to show that G, as defined on (0, ∞)^d by (2.7) is a valid distribution function. Indeed, this is easily shown along the lines of the following sketch. In particular, (i) G is grounded at 0 trivially by inspection: G(0) = 0. (ii) By virtue of the fact that f is block-decreasing, 0 ≤ lim_|x|→∞ f (x) ≤ lim_|x|→∞{1/|x|} = 0 is is true and this can be used to show straightforwardly that lim_{x₁∧…∧x_d→∞} G(x₁, … , x_d) = 1. (iii) Similarly, it is an easy task to show that V_G(x, y] ≥ 0 for all 0 ≤ x ≤ y. Conditions (i)–(iii) are necessary and sufficient for G to be a bona-fide distribution function. This completes the proof.

2.2. Lebesgue measurability of block-decreasing functions

Now we note a technical fact concerning the (Lebesgue) measurability of block-decreasing functions which will be needed in our proofs in Section 3.2.

Proposition 2.4

Let f be a real-valued, non-negative function on (0, ∞)^d that is non-increasing and convergent to zero in each coordinate x_j, keeping all other coordinates fixed, as x_j coordinate tends to ∞. Then:

f is Lebesgue-measurable.
There exists such a function f that is not Borel-measurable. Such an f exists with f also satisfying sup{f (x) | x ∈ (0, ∞)^d} < ∞.

Proof. Proposition 2.4 (a) follows from Theorem 3 of [31]. Proposition 2.4 (b) is standard and follows from Proposition 1.2.2 in [50].

3. Existence and consistency of the MLE

Let X₁, … , X_n be i.i.d. random vectors distributed according to some density $f_{0} = f_{G_{0}} \in F_{SMU} (d)$ where f₀ is unknown. Our goal is to estimate the unknown SMU density, f₀, based on X₁, … , X_n. We will be interested in maximizing the likelihood function $f \mapsto \prod_{i = 1}^{n} f (X_{i})$ or, equivalently, the log-likelihood function $f \mapsto n P_{n} \log {f (X)}$ over $f \in F_{SMU} (d)$ where $P_{n} = n^{- 1} \sum_{i = 1}^{n} δ_{X_{i}}$ is the empirical measure of the data. Any such maximizer, ${\hat{f}}_{n} \in F_{SMU} (d)$ , should one exist, will be called a (nonparametric) maximum likelihood estimator of f₀, based on X₁, … , X_n. Since f₀ = f_G₀ is given by (2.1) it follows from Theorem 2.3 that estimation of $f_{0} \in F_{SMU}$ is equivalent to estimation of G₀.

3.1. On existence and uniqueness of an MLE

We begin with a definition followed by the main theorem of this subsection.

Definition 3.1 (Rectangular Grid Generated by Data)

Suppose that x₁, … , x_n are ( fixed or random) elements in (0, ∞)^d and suppose that xi = (x_i1, … , x_id)’ where i = 1, 2, … , n. Define the matrix A = [x_ij] ∈ M_n×d((0, ∞)) whose ith row is exactly $x_{i}^{'}$ , for i ∈ {1, 2, … , n}. Also let A^# = { (x(i₁),1, x(i₂),2, … , x(i_d),d) | i₁, … , i_d ∈ {1, 2, … , n}} denote the rectangular grid generated by A, where x_(i),j denotes the ith smallest element among x_1j, … , x_nj where i ∈ {1, 2, … , n} and j ∈ {1, 2, … , d}. In particular, x_* = (x_(1),1, x_(1),2, … , x_(1),d) and x_* = (x_(n),1, x_(n),2, … , x_(n),d) denote the element-wise minimum and maximum of x₁, … , x_n, respectively. For each fixed j ∈ {1, 2, … , d}, let n_j(A) := card({x_i,j | i = 1, 2, … , n}), and notice that we have: card $(A^{#}) = \prod_{j = 1}^{d} n_{j} (A) \equiv N \leq n^{d}$ .

Theorem 3.1 (Existence and Characterization of an MLE in $F_{S M U} (d)$ )

A maximum likelihood estimator (MLE), ${\hat{f}}_{n} \equiv f_{{\hat{G}}_{n}} \in F_{S M U} (d)$ of $f_{0} \equiv f_{G_{0}} \in F_{S M U} (d)$ almost surely exists, where ${\hat{G}}_{n} \in g_{d}$ is a purely-atomic probability measure, with at most n atoms, all of which are concentrated on A^#–the rectangular grid generated by the data X₁, … , X_n.
For almost all ω, the unique MLE, ${\hat{f}}_{n} \equiv f_{{\hat{G}}_{n}} \in F_{S M U} (d)$ , is completely characterized by the following Fenchel conditions:
$P_{n} {\frac{1_{[X \leq x]}}{\hat{f_{n}} (X)}} \leq ∣ x ∣; f o r a l l x \in {(0, \infty)}^{d},$ (3.1)

$a n d P_{n} {\frac{1_{[X \leq y]}}{\hat{f_{n}} (X)}} = ∣ y ∣; i f a n d o n l y i f$ (3.2)
y ∈ (0, ∞)^d satisfies ${\hat{G}}_{n} ({y}) > 0$ ; or, equivalently,
${(- 1)}^{d} \lim_{\in ↓ 0} {V_{\hat{f_{n}}} [y, y + \in 1)} > 0 .$

Maximum likelihood estimation in mixture models has been studied in general by Lindsay [34], and this material is nicely summarized in [35, Chapter 5]. To prove the present theorem, we will therefore appeal to the results in [35, Chapter 5] and [47]. We begin with three lemmas.

Lemma 3.2

The support set $y \equiv supp ({\hat{G}}_{n})$ of the mixing measure ${\hat{G}}_{n}$ of any MLE ${\hat{f}}_{n}$ is contained in the grid A^# ⊂ (0, ∞)^d generated by the observed data X₁, … , X_n; i.e., $y \subset A^{#}$ .

Proof. First we show that $y \subset (0, X^{*}]$ where X^* ≡ X₁ ∨ … ∨ X_n and the maximums are taken coordinatewise. If ${\hat{f}}_{n}$ maximizes $L_{n} (f) = n P_{n} \log f (x)$ over $f \in F_{SMU} (d)$ and there is some y ∈ (0, ∞)^d \(0, X^*] with $y \in y$ , then ${\hat{f}}_{n} (y) > 0$ . Since ${\hat{f}}_{n}$ is block decreasing, this implies that $0 < \int_{(0, X^{*}]} {\hat{f}}_{n} (x) d x \equiv β < 1$ . Then consider $\tilde{f} (x) \equiv ({\hat{f}}_{n} (x) ∕ β) 1_{(0, X^{*}]} (x)$ it is easily seen that $\tilde{f} \in F_{SMU} (d)$ and has greater likelihood than ${\hat{f}}_{n}$ , contradicting the assumption that ${\hat{f}}_{n}$ maximizes the likelihood. Thus $y \subset (0, X^{*}]$ , and we may restrict attention to the class of estimators with support contained in (0, X^*], say $K^{*} (d)$ . Suppose that ${\hat{f}}_{n} \in K^{*} (d)$ . Consider the mixing measure G̃_n defined by

{\tilde{G}}_{n} \equiv \sum_{j : W_{j} \in A^{#}} π_{j} δ_{W_{j}} ∕ \sum_{j : W_{j} \in A^{#}} π_{j} \equiv C \sum_{j : W_{j} \in A^{#}} π_{j} δ_{W_{j}}

where

π_{j} \equiv {(- 1)}^{d} V_{\hat{f_{n}}} [W_{j}, {W_{j}}^{+}) \cdot ∣ W_{j} ∣, for W_{j} \in A^{#}

where $W_{j}^{+} \in A^{#}$ defines the smallest rectangle above and right of W_j in the partition of [0, X^*] defined by the data. Then it is easy to see that

\tilde{f} (x) = \int_{{(0, \infty)}^{d}} \frac{1}{∣ u ∣} 1_{(0, u]} (x) d {\tilde{G}}_{n} (u)

satisfies

\begin{matrix} \tilde{f} (W_{j}) & = C \sum_{k : W_{k} \geq W_{j}} \frac{π_{j}}{∣ W_{j} ∣} \\ = C \sum_{k : W_{k} \geq W_{j}} {(- 1)}^{d} V_{\hat{f_{n}}} [W_{j}, W_{k}) \\ = C {(- 1)}^{d} V_{\hat{f_{n}}} [W_{j}, 2 X^{*}) = C \hat{f_{n}} (X_{j}), \end{matrix}

and this implies that

\tilde{f} (x) = C \sum_{j : W_{j} \in A^{#}} 1_{(W_{j}^{-}, W_{j}]} (x)

where $W_{i}^{-}$ defines the smallest rectangle below and to the left of W_j in the partition of [0, X^*] defined by the data. If ${\hat{f}}_{n} \neq \tilde{f}$ , then there exists $y \in (W_{j}^{-}, W_{j}]$ for some W_j ∈ A^# such that ${\hat{f}}_{n} (y) \neq \tilde{f} (y)$ , and then necessarily ${\hat{f}}_{n} (y) > \tilde{f} (y) = \tilde{f} (W_{j})$ .

This yields, since ${\tilde{f}}_{n} \in K^{*} (d)$ ,

\begin{matrix} 1 = & \int_{(0, X^{*})} \tilde{f} (x) d x = C \sum_{j : W_{j} \in A^{#}} {{\hat{f}}_{n} (W_{j}) \int_{(W_{j}^{-}, W_{j}]} d x} \\ < & C \sum_{j : W_{j} \in A^{#}} {\hat{f}}_{n} (W_{j}) \int_{(W_{j}^{-}, W_{j}]} {\hat{f}}_{n} (x) d x = C \int_{(0, X^{*})} {\hat{f}}_{n} (x) d x = C \end{matrix}

since $f \in K^{*} (d)$ . Thus f̃ has a greater log-likelihood than ${\hat{f}}_{n}$ , and it follows that $supp ({\hat{G}}_{n}) \subset A^{#}$ .

Now we can prove uniqueness of the MLEs ${\hat{f}}_{n}$ and ${\hat{G}}_{n}$ .

Lemma 3.3

There exists a set of points $y = {y_{1}, \dots, y_{m}} \subset {(0, \infty)}^{d}$ with m ≤ n such that a $F_{SMU} (d)$ density ${\hat{f}}_{n}$ with corresponding mixing measure ${\hat{G}}_{n}$ is the MLE only if $supp ({\hat{G}}_{n}) \subset y$ . Thus any MLE has the form

{\hat{f}}_{n} (x) = \sum_{j = 1}^{m} π_{j} \frac{1}{∣ y_{j} ∣} 1 (0, y_{j}] (x)

(3.3)

where π_j ≥ 0, $\sum_{j = 1}^{m} π_{j} = 1$ . Moreover, the vector ${({\hat{f}}_{n} (X_{i}))}_{i = 1}^{n}$ is unique.

Proof. As in [34,35], define Γ (u) ∈ (0, ∞)ⁿ by

Γ (u) ≔ (\frac{1}{∣ u ∣} 1_{(0, u]} (X_{1}), \dots, \frac{1}{∣ u ∣} 1_{(0, u)} (X_{n})),

and define the set Γ ≡ {Γ (u) | u ∈ (0, ∞)^d}. Then Γ is a closed and bounded, hence compact, subset of [0, ∞)ⁿ. Thus by Rockafellar [47, Theorem 17.2] $\bar{conv (Γ)} = conv (\bar{Γ}) = conv (Γ)$ is also a compact subset of [0, ∞)ⁿ. Thus the continuous function $\prod_{i = 1}^{n} z_{i}$ attains its supremum on conv(Γ ). Let $S = {argmax}_{z \in conv (Γ)} \sum_{i = 1}^{n} \log z_{i}$ . Since the intersection of Γ and the interior (0, ∞)ⁿ of [0, ∞)ⁿ is not empty, we have S ∈ (0, ∞)ⁿ. Since $\sum_{i = 1}^{n} \log z_{i}$ is strictly concave, S consists of a single point, $\hat{f} = {({\hat{f}}_{i})}_{i = 1}^{n} > 0$ . Therefore for any $MLE {\hat{f}}_{n}$ it follows that the vector ${({\hat{f}}_{n} (X_{i}))}_{i = 1}^{n}$ is unique. Note that the gradient of $\sum_{i = 1}^{n}$ log z_i at f̂ is proportional to $1 ∕ \hat{f} \equiv {(1 ∕ {\hat{f}}_{i})}_{i = 1}^{n}$

Now dim(conv(Γ )) = n; if we consider the n points u_i = X_i, then the n vectors Γ (u_i) = (1_{(0,X_i]}(X₁), … , 1_{(0,X_i]}(X_n))/|X_i|, i = 1, … , n, are almost surely linearly independent. (In fact, the matrix M with rows |X_i|Γ (X_i), i = 1, … , n has det(M) = 1 a.s. if the X_i’s are i.i.d. with any density f .) By Rockafellar [47, Theorem 27.4] the vector $1 ∕ \hat{f}$ belongs to the normal cone of conv(Γ) at f̂ . Since 1/f̂ > 0 we have f̂ ∈ ∂(conv(Γ )) and the plane τ defined by $\sum_{i = 1}^{n} z_{i} ∕ {\hat{f}}_{i} = n$ is a support plane of conv(Γ ) at f̂ . Thus for ν_i = 1/(nf̂_i), i = 1, … , n, it follows that

q (u) \equiv ∣ u ∣ - \sum_{i = 1}^{n} v_{i} 1_{(0, u]} (X_{i}) \geq 0

for all u ∈ [0, ∞)^d and q(u) = 0 if u = 0 or Γ (u) ∈ τ . We let $y$ denote the set of vectors u such that Γ (u) ∈ τ ; i.e., $Γ (y) = τ \cap Γ$ .

The intersection τ ∩ conv(Γ ) is an exposed face of conv(Γ ); see e.g. [47, p. 162]. By Rockafellar [47, Theorem 18.3], $τ \cap conv (Γ) = conv (Γ (y))$ , and by Theorem 18.1, $supp ({\hat{G}}_{n}) \subset y$ . This implies that for any MLE ${\hat{f}}_{n}$ , the support of the corresponding mixing measure ${\hat{G}}_{n}$ is a subset of $y$ , and thus any MLE has form (3.3) with $y_{j} \in y$ for j = 1, … ,m. To see that m ≤ n, note that $y_{j} \in y \subset A^{#}$ satisfy

∣ y_{j} ∣ = \sum_{i = 1}^{n} v_{i} 1_{(0, y_{j}]} (X_{i}) = 〈 v, ∣ y_{j} ∣ Γ (y_{j}) 〉, j = 1, \dots, m .

(3.4)

Suppose that the vectors ${∣ y_{j} ∣ Γ (y_{j})}_{j = 1}^{m}$ are linearly dependent; i.e.,

\sum_{j = 1}^{m} b_{j} ∣ y_{j} ∣ Γ (y_{j}) = 0

in $R^{n}$ for some b_j, j = 1, … ,m. Since all the coordinates of the |y_j|Γ (y_j) vectors take values in {0, 1}, this system of equations is algebraically equivalent to the same system in which all the b_j’s take only integer values, i.e., $b_{j} \in Z$ for j = 1, … ,m.

Then it follows on the one hand that

\begin{matrix} \sum_{j = 1}^{m} b_{j} 〈 v, ∣ y_{j} ∣ Γ (y_{j}) 〉 = & \sum_{j = 1}^{m} b_{j} \sum_{i = 1}^{n} v_{i} 1_{(0, y_{j}]} (X_{i}) \\ = & 〈 v, \sum_{j = 1}^{m} b_{j} ∣ y_{j} ∣ Γ (y_{j}) 〉 = 〈 v, 0 〉 = 0, \end{matrix}

and hence, by (3.4), $\sum_{j = 1}^{m} b_{j} ∣ y_{j} ∣ = 0$ , or, since y_j = W_{i_j} ∈ A^# for some i_j,

\sum_{j = 1}^{m} b_{j} ∣ W_{i_{j}} ∣ = 0

with all $b_{j} \in Z$ . But this equation has at most countably many solutions {|W_ij |, j = 1, … ,m}, and hence occurs with $P_{0}^{n}$ probability 0. That is, for any fixed vector $b = {(b_{j})}_{j = 1}^{k}$ with all $b_{j} \in Z$ , the function $f_{b} (X_{1}, \dots, X_{n}) = \sum_{j = 1}^{k} b_{j} ∣ W_{i_{j}} ∣$ has at most a finite number of zeros, so $P_{0}^{n} (f_{b} (X_{1}, \dots, X_{n}) = 0) = 0$ , and since $Z$ is countable $P_{0}^{n} (\cup_{b \in Z^{k}} {f_{b} (X_{1}, \dots, X_{n}) = 0}) = 0$ . Thus $P_{0}^{n} ({\cap_{b \in Z}}^{k} {f_{b} (X_{1}, \dots, X_{n}) \neq 0}) = 1$ . Hence it follows that the linear dependence condition only holds on an event with probability 0.

Thus the vectors |y_j|Γ (y_j), j = 1, … ,m are linearly independent almost surely $P_{0}^{n}$ , and hence m ≤ n ( $P_{0}^{n}$ -almost surely).

Lemma 3.4

The discrete mixing measure ${\hat{G}}_{n}$ which defines an MLE is $P_{0}^{n}$ -almost surely unique.

Proof. Suppose that there exist two different MLE’s ${\hat{f}}_{n}^{1}$ and ${\hat{f}}_{n}^{2}$ . then

{\hat{f}}_{n}^{l} (x) = \sum_{j = 1}^{m} π_{j}^{l} \frac{1}{∣ y_{j} ∣} 1_{(0, y_{j}]} (x), l = 1, 2,

where $π_{j}^{l} \geq 0$ and $\sum_{j = 1}^{m} π_{j}^{l} = 1$ for l = 1, 2. Therefore

δ_{n} (x) \equiv {\hat{f}}_{n}^{l} (x) - {\hat{f}}_{n}^{2} (x) = \sum_{j = 1}^{m} r_{j} \frac{1}{∣ y_{j} ∣} 1_{(0, y_{j}]} (x)

where $r_{j} \equiv π_{j}^{1} - π_{j}^{2}$ has at least n zeros (since we know that

{({\hat{f}}_{n}^{1} (X_{i}))}_{i = 1}^{n} = {({\hat{f}}_{n}^{2} (X_{i}))}_{i = 1}^{n} = {({\hat{f}}_{n} (X_{i}))}_{i = 1}^{n}

is unique). So, uniqueness holds if the vectors

{(1_{(0, y_{j}]} (X_{i}))}_{i = 1}^{n} \in {0, 1}^{n}, for j = 1, \dots, m \leq n

are (almost surely) linearly independent. But this follows from the proof of Lemma 3.3.

Theorem 3.1 does not assert that the MLE is always unique. An MLE is $P_{0}^{n}$ almost surely unique, but we now present an example in which there exist an infinite number of MLE’s.

Example 3.1 (A MLE in $F_{SMU}$ is Not Always Unique)

To be able to graphically illustrate the set Γ , in the proof of Theorem 3.1, we need to restrict consideration to n = 2 and in order that we be able to graphically illustrate the MLE(s) we need to restrict consideration to d = 2. Suppose that X₁ = (1, 3) and X₂ = (3, 2) are the observation points. The set

Γ \equiv {\frac{1}{u_{1} u_{2}} (1_{(0, u]} (X_{1}), 1_{(0, u]} (X_{2})) ∣ u = (u_{1}, u_{2}) \in {(0, \infty)}^{2}}

and its convex hull, Conv(Γ ), are illustrated in Fig. 1.

Using [35, Theorem 22, p. 118], it follows that any MLE, ${\hat{f}}_{2}$ , will have a unique value for $\hat{f} \equiv ({\hat{f}}_{2} (X_{1}), {\hat{f}}_{2} (X_{2}))$ that is given by $\hat{f} = ({\tilde{w}}_{1}^{- 1}, {\tilde{w}}_{2}^{- 1})$ where w̃ = (w̃₁,w̃₂) maximizes the function (w₁,w₂) ↦ log(w₁w₂) on the set

{(w_{1}, w_{2}) \in {(0, \infty)}^{2} ∣ \frac{w_{1}}{3} \leq 2 and \frac{w_{2}}{6} \leq 2} .

It is immediate that w̃ = (6, 12) from which we conclude that f̃ = (1/6, 1/12) has exactly two representations as convex combinations in terms of pairs of the points {A₁, A₂, A₃} (see Fig. 1(a) again):

(\frac{1}{6}, \frac{1}{12}) = \frac{1}{2} (0, \frac{1}{6}) + \frac{1}{2} (\frac{1}{3}, 0), and (\frac{1}{6}, \frac{1}{12}) = \frac{1}{4} (\frac{1}{3}, 0) + \frac{3}{4} (\frac{1}{9}, \frac{1}{9}) .

These two convex combinations yield two different maximum likelihood estimators, as shown in Fig. 2(a) and (b).

It should be noted, however, that infinitely many maximum likelihood estimators exist in this case since each convex combination of these two MLEs is again an MLE, by virtue of linearity of f_G (recall (2.1)) as a function of the mixing distribution, G.

3.2. Strong pointwise consistency of the MLE

Let X₁, X₂, … , X_n, … be the coordinate random elements on the (completed) infinite product space $(Ω^{\infty}, A^{\infty}, P^{\infty})$ such that these coordinates are i.i.d. according to f₀ ≡ f_G₀ on (0, ∞)^d. Let $A \in A^{\infty}$ be the event (with P^∞-probability one) that for each $n \in Z$ there exists a unique SMU density, ${\hat{f}}_{n} \equiv f_{{\hat{G}}_{n}}$ , maximizing the log-likelihood.

From Theorem 2.3 we have that for each $n \in N$ and a fixed ω ∈ A, there exists a unique Borel probability measure, ${\hat{G}}_{n}$ on ((0, ∞)^d, || · ||₂), such that

{\hat{f}}_{n} (x) = \int_{{(0, \infty)}^{d}} \frac{1}{∣ u ∣} 1_{(0, u)} (x) d {\hat{G}}_{n} (u) = \int_{u \geq x} \frac{1}{∣ u ∣} d {\hat{G}}_{n} (u)

(3.5)

holds true for all x ∈ (0, ∞)^d. We are ready to formulate and prove the following proposition.

Proposition 3.5 (Strong Consistency of the MLE in $F_{SMU}$ )

1. The sequence of maximum likelihood mixing distributions ${{\hat{G}}_{n}}_{n = 1}^{\infty}$ converges weakly to G₀ as n → ∞, P^∞-almost surely.
2. In addition, for Lebesgue almost all x ∈ (0, ∞)^d, ${\hat{f}}_{n} (x) \to {}_{a . s .}f_{0} (x)$ as n → ∞. In particular, if f₀ is continuous at x ∈ (0, ∞)^d, then
  $∣ {\hat{f}}_{n} (x) - f_{0} (x) ∣ \to_{a . s .} 0 a s n \to \infty .$
The sequence of maximum likelihood estimators, ${{\hat{f}}_{n}}_{n = 1}^{\infty}$ , is strongly consistent in the total variation (or L₁) and in the Hellinger metrics. That is,
$\int_{{(0, \infty)}^{d}} ∣ {\hat{f}}_{n} (x) - f_{0} (x) ∣ d x \to_{a . s .} 0 a s n \to \infty,$
and, with $h^{2} (p, q) = \int {\sqrt{p (x)} - \sqrt{q (x)}}^{2} d x$ ,
$h ({\hat{f}}_{n}, f_{0}) \to_{a . s .} 0 a s n \to \infty .$

Proof. (a) (i) To be able to apply Theorems 3.4, 3.5 and 3.7 of [39], with the refinement on page 143 of the same article, we need to provide the relevant setup as well as establish the assumptions of Pfanzagl’s theorems. We do this below.

Let $C_{0} ({(0, \infty)}^{d}, {∥ \cdot ∥}_{2})$ denote the set of all real-valued, continuous functions on (0, ∞)^d that vanish at ∞. Let Θ_* denote the set of all Borel sub-probability measures on (0, ∞)^d, equipped with the vague topology, τ , which makes the space a compact, metrizable, topological space, and thus with a countable base. It is also a convex subset of the linear space of all finite, signed, Borel measures on ((0, ∞)^d, || · ||₂). For clarity, the vague topology is the smallest topology that makes the functions

μ \mapsto \int_{{(0, \infty)}^{d}} g (x) d μ (x)

continuous, for each $g \in C_{0} ({(0, \infty)}^{d}, {∥ \cdot ∥}_{2})$ . By metrizability, the topology τ is completely characterized by convergent sequences, $θ_{n} \overset{v}{\Rightarrow} θ$ as n → ∞, on (Θ_*, τ ).

Let also Θ ⊆ Θ_* be the set of all Borel probability measures on (0, ∞)^d, and notice that μ ∈ Θ. Also, for each θ_* ∈ Θ_* there exists a unique c ∈ [0, 1] and a unique θ ∈ Θ, such that θ_* = cθ. Further, notice that letting m(ν, ·) ≡ fν (·), for each ν ∈ Θ_*, and $M_{n} (\cdot) \equiv P_{n} \log {m (\cdot, X)}$ , we have

M_{n} (θ_{*}) = \log {c} + M_{n} (θ) \leq M_{n} (θ), since c \in [0, 1],

whence, sup_{θ∈Θ_*} (M_n(θ)) = sup_θ∈Θ (M_n(θ)).

With reference measure the Lebesgue measure λ ≡ Q and for each ν ∈ Θ_*, let P_ν ∈ Θ_* be the sub-probability, Borel measure on ((0, ∞)^d, || · ||₂) with Radon–Nikodym derivative with respect to λ being f_ν , Lebesgue almost surely. Then by virtue of Fubini–Tonelli, P_ν ∈ Θ when and only when ν ∈ Θ. Also, notice that for each fixed x ∈ (0, ∞)^d, the functional ν ↦ fν (x) is not vaguely continuous at any ν ∈ Θ_* with a discontinuity point on the boundary of [x,∞). However, since for a fixed x ∈ (0, ∞)^d, the function $y \mapsto 1_{[x, \infty)} (y) ∕ ∣ y ∣$ is easily seen to be an upper semi-continuous function on (0, ∞)^d–vanishing at ∞, Doob [15, Theorem 10, p. 138], applies and asserts that the function ν ↦ f_ν (x) on (Θ_*, τ ) is itself (vaguely) upper semi-continuous. Since this holds for all x ∈ (0, ∞)^d, it holds almost-surely. Also, the mapping ν ↦ f_ν (x) is affine on Θ_* (and hence concave also).

It remains to establish that for each fixed τ -open subset U of Θ_*, the real-valued function T_U (·) on (0, ∞)^d defined by

T_{U} (x) = \sup_{ν \in U} {\int_{{(0, \infty)}^{d}} \frac{1}{∣ u ∣} 1_{(0, u]} (x) d ν (u)}

is a $A$ -measurable function. We can choose to take $A$ to be the Lebesgue σ-field, in which case measurability follows by observing that T_U (·) is a block-decreasing function and appeal to Proposition 2.4.

We now apply Theorem 3.4 of [39] to our setting and further appeal to the fact that a vaguely convergent sequence of probability measures with limit a probability measure, is, in fact, weakly convergent. This gives the desired conclusion: the random sequence of maximum likelihood mixing probability measures ${{\hat{G}}_{n}}_{n = 1}^{\infty}$ converges weakly to G₀ as n → ∞, P^∞-almost surely.

(ii) Combining the fact that, for each fixed x ∈ (0, ∞)^d, ν ↦ f_ν(x) is vaguely upper semi-continuous on Θ_* with the conclusion of part (a)(i), we get

\underset{n \to \infty}{\lim^{¯}} {f_{{\hat{G}}_{n}} (x)} \leq f_{0} (x); P^{\infty} - a . s . for all x \in {(0, \infty)}^{d} .

(3.6)

Let

F_{G_{0}} (\cdot) = \int_{{(0, \infty)}^{d}} \frac{∣ \cdot \land u ∣}{∣ u ∣} d G_{0} (u)

and

F_{{\hat{G}}_{n}} (\cdot) = \int_{{(0, \infty)}^{d}} \frac{∣ \cdot \land u ∣}{∣ u ∣} d {\hat{G}}_{n} (u)

be the distribution functions corresponding to the densities f₀(·) and ${\hat{f}}_{n} (\cdot)$ , respectively, $n \in N$ . These distribution functions are everywhere continuous on the Euclidean set (0, ∞)^d. In fact, since for each fixed x ∈ (0,∞)^d, the function u ↦ |x ∧ u| / |u| is bounded (by 1) and continuous on (0, ∞)^d, we then have that

F_{{\hat{G}}_{n}} (x) \to_{a . s .} F_{G_{0}} (x) for all x \in {(0, \infty)}^{d}

(3.7)

follows directly by the definition of almost sure weak convergence of the mixing random measures ${{\hat{G}}_{n}}_{n = 1}^{\infty}$ to G₀, established in part (a)(i).

Let B be the set of points on (0, ∞)^d at which f₀ is continuous. Then B^c has Lebesgue measure zero, λ(B_c ) = 0, exactly because f₀ is discontinuous on the boundary ∂[x₀, ∞) for a (possibly non-existent) x₀ ∈ (0, ∞)^d where P₀ is discontinuous (i.e., such that P₀({x₀}) > 0). Since P₀ can have at most countably many discontinuity points x₀ ∈ (0, ∞)^d and since λ(∂[x₀, ∞)) = 0, we get by countable subadditivity of λ that indeed λ(B_c ) = 0.

Fix arbitrary x ∈ B and ε > 0. Then, since f₀ is lower semi-continuous at x, there exists an open neighborhood U_x,ε of x such that for every y ∈ U_x,ε we have that f₀(y) > f₀(x) – ε. In particular, there exists an U_x,ε ∍ x_ε > x satisfying f₀(xε) > f₀(x) – ε. Since f₀ is block-decreasing, we have:

\frac{V_{F_{G_{0}}} (x, x_{\in})}{λ ((x, x_{\in}])} = \frac{\int_{(x, x_{\in}]} {f_{0} (y)} d y}{λ ((x, x_{\in}])} \geq f_{0} (x_{\in}) > f_{0} (x) - ∊ .

(3.8)

Further, for each fixed $n \in N$ , since ${\hat{f}}_{n} (\cdot)$ is block-decreasing (as a SMU density), we have

f_{{\hat{G}}_{n}} (x) \geq \frac{\int_{(x, x_{\in}]} {f_{{\hat{G}}_{n}} (y)} d y}{λ ((x, x_{\in}])}

(3.9)

= \frac{V_{F_{{\hat{G}}_{n}}} (x, x_{\in}]}{λ ((x, x_{\in}])} .

(3.10)

Eq. (3.7) further implies that

V_{F_{{\hat{G}}_{n}}} (x, x_{\in}] \to V_{F_{G_{0}}} (x, x_{\in}], as n \to \infty .

(3.11)

Combining Eqs. (3.8)–(3.11) and the fact that ε > 0 was arbitrary, we get

\underset{n \to \infty}{\lim_{¯}} {f_{{\hat{G}}_{n}} (x)} \geq f_{0} (x); P^{\infty} - a . s . for x \in B .

(3.12)

Eqs. (3.6) and (3.12) yield the assertion: for Lebesgue almost all x ∈ (0, ∞)^d (and, in particular, at the points of continuity of f ), $f_{{\hat{G}}_{n}} (x) \to {}_{a . s .}f_{0} (x)$ as n → ∞holds.

(b) Showing consistency in the L₁ (total-variation) norm is a direct consequence of part (a) (ii) and Glick’s Theorem, [17]; see also [14, p. 25].

Convergence in the Hellinger metric follows from the following well-known inequalities of [33, p.46]:

h^{2} (P, Q) \leq \frac{1}{2} {∥ P - Q ∥}_{L_{1}} \leq h (P, Q) {2 - h^{2} (P, Q)}^{\frac{1}{2}},

where $h^{2} (P, Q) = 2^{- 1} \int {(\sqrt{d P} - \sqrt{d Q})}^{2}$ is the squared Hellinger metric and || · ||_L₁ is the L₁-norm.

4. A local asymptotic minimax lower bound

Let X_i := (X_i,1, … , X_i,d)’ for i = 1, 2, … , n be i.i.d. random vectors from density $f \in F_{SMU} (d)$ . For a fixed x₀ ≡ (x_0,1, … , x_0,d)’ ∈ (0, ∞)^d, we want to estimate the functional T(f) := f(x₀) on the basis of X₁, … , X_n. We shall make the following assumption:

Assumption 4.1

Suppose that $f \in F_{SMU}$ is continuously differentiable at x₀, f (x₀) > 0, and, in particular, there exists an open ball A(x₀) around x₀ such that f is everywhere strictly positive on A(x₀) and where (∂/∂x_j)f(x₀) < 0 exist for all j ∈ {1, 2, … , d} and are continuous on A(x₀) ⊆ (0, ∞)^d. Further, we assume that the full mixed derivative of f exists, is continuous on A(x₀), and satisfies

{{(- 1)}^{d} \frac{\partial^{d} f}{\partial x_{1} \dots \partial x_{d}} (x) ∣}_{x = y} > 0 for all y \in A (x_{0}) .

Proposition 4.1

Suppose that $f \in F_{SMU}$ satisfies Assumption 4.1 at the fixed point x0 ∈ (0, ∞)^d. Then there is a sequence ${f_{n}} \subset F_{SMU}$ such that any estimator sequence {T_n} of f (x₀) satisfies

\begin{matrix} \underset{n \to \infty}{\lim_{¯}} \max {E_{f_{n}} {n^{\frac{1}{3}} ∣ T_{n} - f_{n} (x_{0}) ∣}, E_{f} {n^{\frac{1}{3}} ∣ T_{n} - f (x_{0}) ∣}} \\ \geq \frac{e^{- \frac{1}{3}}}{2^{d}} {3^{d - 1}}^{\frac{1}{3}} {{{(- 1)}^{d} \frac{\partial^{d} f (x)}{\partial x_{1} \dots \partial x_{d}} ∣}_{x = x_{0}} \cdot f (x_{0})}^{\frac{1}{3}} . \end{matrix}

(4.1)

Remark

The lower bound in Proposition 4.1 should be contrasted to a similar lower bound for estimation of f(x₀) for $f \in F_{BDD}$ which is derived by Pavlides [38]. In that case the natural hypothesis is ∂f (x₀)/∂x_i< 0 for i = 1, … , d, and the resulting rate of convergence is n^1/(d+2).

To prove Proposition 4.1 we will make use of the following lemma. It was established in the form presented here by [23]; see also Groeneboom and Jongbloed [22,29].

Lemma 4.2.

Let $F$ be a class of densities on a measurable space $(X, F)$ and f a fixed element of $F$ . Let $F_{f}$ denote any open Hellinger ball with center $f \in F$ . Assume that there exists a sequence ${f_{n}}_{n = 1}^{\infty} \subseteq F$ such that and

\lim_{n \to \infty} {\sqrt{n} h (f_{n}, f)} = α

(4.2)

and

\lim_{n \to \infty} ∣ T (f_{n}) - T (f) ∣ = β

(4.3)

both hold for some constants 0 < α, β < ∞, and where T is a functional on $F$ . Here, $h^{2} (f_{n}, f) \equiv 2^{- 1} \int {\sqrt{f_{n} (x)} - \sqrt{f (x)}}^{2} d μ (x)$ , is the Hellinger distance between the μ-densities f_n and f · Let l(·) be a convex function, symmetric about zero, which is non-decreasing on [0, ∞).

Then, it holds that

\underset{n \to \infty}{\lim_{¯}} {R_{n, 1} (F_{f})} \geq l (\frac{1}{4} β e^{- 2 α^{2}})

(4.4)

where $R_{n, 1} (F) \equiv \inf_{T_{n}} \sup_{g \in F} E_{g^{\otimes n}} {l (T_{n} - T (g))}$ is the minimax risk for estimating the functional T(f) based on n i.i.d observations from $F$ .

In particular, for the loss l(x) = |x| on we have

\underset{n \to \infty}{\lim_{¯}} {R_{n, ∣ \cdot ∣} (F_{f})} \geq l (\frac{1}{4} β e^{- 2 α^{2}})

(4.5)

Hereafter, fix an otherwise arbitrary vector h := (h₁, …, h_d) ∈ (0, ∞)^d, and define H := diag(h) ∈ M_d×d ((0, ∞)) . For each $k \in N$ , consider the perturbation rectangle

I_{n} (k) ≔ \otimes_{i = 1}^{d} [x_{0, i} - n^{- \frac{1}{k}} h_{i}, x_{0, i} + n^{- \frac{1}{k}} h_{i}],

only for those positive integers n ≥ n₀(k, x₀, h) for which I_n(k) ⊆ A(x₀) for all n ≥ n₀. The two-dimensional case, d = 2, is illustrated in Fig. 3.

Recall Assumption 4.1. Let b := (∂^d/∂x₁ … ∂x_d)f(x)|_x=x₀ and observe that (−1)^db > 0. Finally, define the functions h_n on I_n(3d) as follows:

h_{n} (y_{1}, \dots, y_{d}) ≔ {(- 1)}^{d} \prod_{i = 1}^{d} {1_{(x_{0, i}, x_{0, i} + n^{- \frac{1}{3 d}} h_{i}]} (y_{i}) - 1_{[{x_{0, i - n}}^{- \frac{1}{3 d}} h_{i}, x_{0, i}]} (y_{i})},

and

g_{n} (y) ≔ b \int_{u \underline{≻} y} {1_{I_{n} (3 d)} (u) \cdot h_{n} (u)} d u,

where we observe that g_n(y) ≥ 0 for all y ∈ I_n(3d), since x₀ is the center of the rectangle I_n(3d). In fact, consideration of the geometry of the definition of g_n(·) reveals that, for y ∈ I_n, g_n(y) is equal to (−1)^db > 0 times the volume of the rectangle [ν_n(y) ∧ y, ν_n(y) ∨ y], where ν_n(y) is defined as that vertex of I_n that is closest in L₂-distance from y ∈ I_n. Since I_n is a decreasing sequence of compact sets, it is then immediately clear that g_n(y) is (pointwise) non-increasing in $n \in N$ , for each fixed y ∈ (0, ∞)^d.

Assume that $f \in F_{SMU}$ , and for fixed vectors x₀, h ∈ (0, ∞)^d we further assume that f satisfies Assumption 4.1. For n ≥ n₀(3d, x₀, h), define the perturbed density, f_n of f at x₀, by

f_{n} (x) = {\begin{matrix} \frac{f (x) + θ g_{n} (x)}{d_{n}} : & if x \in I_{n} (3 d) \\ \frac{f (x)}{d_{n}} : & if x \in I_{n}^{c} (3 d) \end{matrix}

(4.6)

for some arbitrary but fixed θ ∈ (0, 1) and where d_n is the normalizing constant for f_n, uniquely determined by ∫_{(0, ∞)^d} f_n(x) dx = 1. We will see the importance of the value of b and the fact that 0 < θ < 1 in the following proposition that establishes that ${f_{n}}_{n \geq n_{1}} \subseteq F_{SMU} (d)$ for a sufficiently large $n_{1} \in N$ .

Proposition 4.3

There exists a positive integer n₁ := n₁(d, x₀, h) ≥ n₀(3d, x₀, h) such that $f_{n} \in F_{SMU}$ for all n ≥ n₁.

Proof. Since $f \in F_{SMU} (d)$ , we get from Theorem 2.3 that

V_{f} [x, y] \geq 0, for all d - boxes [x, y] .

(4.7)

From the definition of g_n(·), we see that its full, mixed partial derivative exists in a neighborhood of x₀. Hence, by definition and the fact that (−1)^db > 0 and θ ∈ (0, 1), we have that

\begin{matrix} {{(- 1)}^{d} \frac{\partial^{d} f_{n}}{\partial x_{1} \dots \partial x_{d}} (x) ∣}_{x = y} \geq & {{(- 1)}^{d} \frac{\partial^{d} f}{\partial x_{1} \dots \partial x_{d}} (x) ∣}_{x = y} - {(- 1)}^{d} b θ \\ = [{{(- 1)}^{d} \frac{\partial^{d} f}{\partial x_{1} \dots \partial x_{d}} (x) ∣}_{x = y} - {(- 1)}^{d} b] + (1 - θ) {(- 1)}^{d} b \\ \geq 2^{- 1} (1 - θ) {(- 1)}^{d} b > 0, \end{matrix}

(4.8)

where the second to last inequality follows from Assumption 4.1 that the full mixed partial derivative of f exists and is continuous at x₀ from which we get, by definition of continuity, that there exists a large enough positive integer n₁ := n₁(d, x₀, h) ≥ n₀(3d, x₀, h) such that

{{(- 1)}^{d} \frac{\partial^{d} f}{\partial x_{1} \dots \partial x_{d}} (x) ∣}_{x = y} - {(- 1)}^{d} b \geq - 2^{- 1} (1 - θ) {(- 1)}^{d} b

holds true for all y ∈ I_n(3d) and n ≥ n₁. The result in (4.8) suggests that

{(- 1)}^{d} V_{f n} [x, y] \equiv {(- 1)}^{d} \int_{(x, y]} {{\frac{\partial^{d} f_{n}}{\partial w_{1} \dots \partial w_{n}} (w) ∣}_{w = u}} d u \geq 0

holds true for all d-boxes (x, y] with x, y ∈ I_n(3d) and n ≥ n₁.

The last case not considered is the one that is exactly one between x and y, in the d-box [x, y], is an element of I_n(3d). See also Fig. 4. For this case, we can appeal to Lemma 2.2 by setting [x₀, y₀] := [x, y] ∩ I_n(3d)–the latter being well-defined as the intersection of two rectangles is itself an rectangle. Then, from Lemma 2.2 and (4.7), we have,

{(- 1)}^{d} V_{f n} [x, y] = {(- 1)}^{d} V_{f n} [x_{0}, y_{0}] + {(- 1)}^{d} \sum_{i = 1}^{m} {V_{f n} [x_{i}, y_{i}]} \geq 0 + 0 = 0,

exactly since $[x_{i}, y] \subseteq I_{n}^{c} (3 d)$ for all i ∈ {1, 2, … ,m} (where m is as defined in Lemma 2.2). For completeness, notice that we were not concerned above with end-point discontinuities of f (or f_n) on the entailed rectangle, subsets of I_n(3d), as, in fact, f (and f_n) is (are) continuous there for n ≥ n₁, by Assumption 4.1.

Fig. 4 — Perturbation rectangle I_n(k), for the case d = 2, with two rectangles intersecting I_n(k) but otherwise not subsets of it.

All these observations finally yield that (−1)^dV_fn [x, y] ≥ 0 holds true for all d-boxes [x, y] and thus Theorem 2.3 asserts that $f_{n} \in F_{SMU}$ for all n ≥ n₁.

We are ready to prove the main proposition of this section.

Proof. Recall Proposition 4.3. First, we establish that

\int_{l_{n}} g_{n} (x) d x = {(- 1)}^{d} b \prod_{i = 1}^{d} {h_{i}^{2}} \cdot n^{- \frac{2}{3}},

(4.9)

where, hereafter, I_n will be the short-hand form for I_n(3d). By definition, notice that,

\begin{matrix} \frac{1}{b} \int_{I_{n}} g_{n} (x) d x = & \int_{I_{n}} \int_{I_{n}} \prod_{i = 1}^{d} {1_{[x_{i} \leq u_{i}]}} h_{n} (u) d u d x \\ = & \int_{I_{n}} h_{n} (u) {\int_{I_{n}} 1_{(0, u]} (x) d x} d u \\ = & \int_{I_{n}} \prod_{i = 1}^{d} {u_{i} - (x_{0 i} - h_{i} n^{- \frac{1}{3 d}})} h_{n} (u) d u \\ = & \prod_{i = i}^{d} {\int_{x_{0 i} - h_{i} n^{- \frac{1}{3 d}}}^{x_{0 i} + h_{i} n^{- \frac{1}{3 d}}} ([u_{i} - (x_{0 i} - h_{i} n^{- \frac{1}{3 d}})] \\ \times [1_{[x_{0 i} - h_{i} n^{- \frac{1}{3 d}}, x_{0 i}]} (u_{i}) - 1_{(x_{0 i}, x_{0 i} + h_{i} n^{- \frac{1}{3 d}}]} (u_{i})]) d u_{i}} \\ = \prod_{i = 1}^{d} {\int_{x_{0 i} - h_{i} n^{- \frac{1}{3 d}}}^{x_{0 i}} [u_{i} - (x_{0 i} - h_{i} n^{- \frac{1}{3 d}})] d u_{i} \\ - \int_{x_{0 i}}^{x_{0 i} + h_{i} n^{- \frac{1}{3 d}}} [u_{i} - (x_{0 i} - h_{i} n^{- \frac{1}{3 d}})] d u_{i}} \\ = \prod_{i = 1}^{d} {\int_{0}^{h_{i} n^{- \frac{1}{3 d}}} [- y + h_{i} n^{- \frac{1}{3 d}}] d y - \int_{0}^{h_{i} n^{- \frac{1}{3 d}}} [w + h_{i} n^{- \frac{1}{3 d}}] d w} \\ = & \prod_{i = 1}^{d} {\int_{0}^{h_{i} n^{- \frac{1}{3 d}}} (- 2 y) d y} = {(- 1)}^{d} \prod_{i = 1}^{d} {h_{i}^{2} n^{- \frac{2}{3 d}}} = {(- 1)}^{d} \prod_{i = 1}^{d} {h_{i}^{2}} \cdot n^{- \frac{2}{3}}, \end{matrix}

thus yielding (4.9).

We next derive another equality, the most important fact about it being the factor n⁻¹ on the right hand side:

\int_{l_{n}} g_{n}^{2} (x) d x = {(\frac{8}{3})}^{d} b^{2} \prod_{i = 1}^{d} {h_{i}^{3}} \cdot n^{- 1} .

(4.10)

Before we start deriving (4.10), let us first define four rectangles $R_{j}^{i}$ with j = 1, 2, 3, 4 for each i ∈ {1, 2, …, d}:

$R_{1}^{i} = [x_{0 i} - h_{i} n^{- \frac{1}{3 d}}, x_{0 i}] \times [x_{0 i} - h_{i} n^{- \frac{1}{3 d}}, x_{0 i}],$
$R_{2}^{i} = [x_{0 i} - h_{i} n^{- \frac{1}{3 d}}, x_{0 i}] \times (x_{0 i}, x_{0 i} + h_{i} n^{- \frac{1}{3 d}}],$
$R_{3}^{i} = (x_{0 i}, x_{0 i} + h_{i} n^{- \frac{1}{3 d}}] \times [x_{0 i} - h_{i} n^{- \frac{1}{3 d}}, x_{0 i}],$
$R_{4}^{i} = (x_{0 i}, x_{0 i} + h_{i} n^{- \frac{1}{3 d}}] \times (x_{0 i}, x_{0 i} + h_{i} n^{- \frac{1}{3 d}}] .$

Then, by definition:

\begin{matrix} \frac{1}{b^{2}} \int_{l_{n}} g_{n}^{2} (x) d x = & \int_{l_{n}} {\int_{l_{n}} h_{n} (u) 1_{[x \leq u]} d u}^{2} d x \\ = & \int_{l_{n}} \int_{l_{n}} \int_{l_{n}} h_{n} (u) h_{n} (v) 1_{[x \leq u \land v]} d v d u d x \\ = & \int_{l_{n}} \int_{l_{n}} {\prod_{i = 1}^{d} [(u_{i} \land v_{i}) - (x_{0 i} - h_{i} n^{- \frac{1}{3 d}})] \times h_{n} (u) h_{n} (v)} d v d u \\ = & \prod_{i = 1}^{d} {\int_{R_{1}^{i} + R_{3}^{i}} [(u \land v) - (x_{0 i} - h_{i} n^{- \frac{1}{3 d}})] d v d u \\ - 2 \int_{R_{2}^{i}} ((u \land v) - (x_{0 i} - h_{i} n^{- \frac{1}{3 d}})] d v d u} \\ = & 2^{d} \prod_{i = 1}^{d} {S_{1 i} + S_{2 i} - S_{3 i}}, \end{matrix}

(4.11)

where the last equality follows by symmetry and Fubini–Tonelli and the integrals in the braces are to be evaluated below:

\begin{matrix} S_{1 i} \equiv & \int_{x_{0 i} - h_{i} n^{- \frac{1}{3 d}}}^{x_{0 i}} \int_{v}^{x_{0 i}} {v - (x_{0 i} - h_{i} n^{- \frac{1}{3 d}})} d u d v \\ = & \int_{x_{0 i} - h_{i} n^{- \frac{1}{3 d}}}^{x_{0 i}} {(x_{0 i} - v) (v - x_{0 i} + h_{i} n^{- \frac{1}{3 d}})} d v \\ = & \int_{x_{0 i} + h_{i} n^{- \frac{1}{3 d}}}^{h_{i} n^{- \frac{1}{3 d}}} {y (- y + h_{i} n^{- \frac{1}{3 d}})} d y [change of variable] \end{matrix}

while, again, by a change of variable argument:

\begin{matrix} S_{2 i} \equiv & \int_{x_{0 i}}^{x_{0 i} + h_{i} n^{- \frac{1}{3 d}}} \int_{v}^{x_{0 i} + h_{i} n^{- \frac{1}{3 d}}} {v - (x_{0 i} - h_{i} n^{- \frac{1}{3 d}})} d u d v \\ = & \int_{x_{0 i}}^{x_{0 i} + h_{i} n^{- \frac{1}{3 d}}} {[(x_{0 i} - v) + h_{i} n^{- \frac{1}{3 d}}] [(v - x_{0 i}) + h_{i} n^{- \frac{1}{3 d}}]} d v \\ = & \int_{0}^{h_{i} n^{- \frac{1}{3 d}}} {(- y + h_{i} n^{- \frac{1}{3 d}}) (y + h_{i} n^{- \frac{1}{3 d}})} d y, \end{matrix}

and similarly:

\begin{matrix} S_{3 i} \equiv & \int_{x_{0 i} - h_{i} n^{- \frac{1}{3 d}}}^{x_{0 i}} {h_{i} n^{- \frac{1}{3 d}} (v - x_{0 i} + h_{i} n^{- \frac{1}{3 d}})} d v \\ = & h_{i} n^{- \frac{1}{3 d}} \int_{0}^{h_{i} n^{- \frac{1}{3 d}}} {h_{i} n^{- \frac{1}{3 d}} - y} d y . \end{matrix}

Let now q_i := h_in^−1/3d, for i ∈ {1, 2, …, d}, and observe that

S_{1 i} + S_{2 i} - S_{3 i} = \int_{0}^{q_{i}} {y (q_{i} - y) + q_{i}^{2} - y^{2} + q_{i}^{2} - q_{i} y} d y = \dots = \frac{4}{3} h_{i}^{3} n^{- \frac{1}{d}},

so that plugging al these in (4.11) yields the desired (4, 10).

Now, recall from the definition of f_n that θ ∈ (0, 1) was arbitrary but fixed. Also, from ∫_{(0, ∞)^d}f_n(x) dx = 1 we can get an explicit expression for the normalizing constant d_n:

\begin{matrix} d_{n} = & \int_{l_{n}} f (x) d x + \int_{l_{n}^{c}} f (x) d x + θ \int_{l_{n}} g_{n} (x) d x \\ = & 1 + θ \int_{l_{n}} g_{n} (x) d x = 1 + {(- 1)}^{d} θ b \prod_{i = 1}^{d} {h_{i}^{2}} \cdot n^{- \frac{2}{3}}, \end{matrix}

(4.12)

where the second to last equality follows from ∫_{(0, ∞)^d}f(x) dx = 1, while the last equality follows from (4.9). Notice from (4.12) that d_n ↓ 1 as n ↑ ∞. Also, from the easily verifiable identity $g_{n} (x_{0}) = {(- 1)}^{d} b Π_{i = 1}^{d} {h_{i}} n^{- 1 ∕ 3}$ , we have

\begin{matrix} n^{\frac{1}{3}} ∣ f_{n} (x_{0}) - f (x_{0}) ∣ = & n^{\frac{1}{3}} ∣ \frac{f (x_{0}) + {(- 1)}^{d} b \prod_{i = 1}^{d} {h_{i}} n^{- \frac{1}{3}}}{d_{n}} - f (x_{0}) ∣ \\ = & ∣ n^{\frac{1}{3}} {\frac{1}{d_{n}} - 1} f (x_{0}) + \frac{{(- 1)}^{d} b θ \prod_{i = 1}^{d} {h_{i}}}{d_{n}} ∣ \\ \to & {(- 1)}^{d} b θ \prod_{i = 1}^{d} {h_{i}} (> 0), as n \to \infty . \end{matrix}

(4.13)

Also,

\begin{matrix} 2 n h^{2} (f_{n}, f) = & n \int_{l_{n}} {\sqrt{f_{n} (x)} - \sqrt{f (x)}}^{2} d x + n \int_{l_{n}^{c}} {\sqrt{f_{n} (x)} - \sqrt{f (x)}}^{2} d x \\ = & n \int_{l_{n}} {\frac{f_{n} (x) - f (x)}{\sqrt{f_{n} (x)} + \sqrt{f (x)}}}^{2} d x + δ_{n}^{2} \int_{l_{n}^{c}} f (x) d x, \end{matrix}

(4.14)

where,

\begin{matrix} δ_{n} \equiv & \sqrt{n} {1 - \frac{1}{\sqrt{d_{n}}}} = \sqrt{n} {\frac{\sqrt{d_{n}} - 1}{\sqrt{d_{n}}}} \\ = & \frac{\sqrt{n} {\sqrt{1 + O (n^{- \frac{2}{3}})} - 1}}{\sqrt{d_{n}}} \to 0, as n \to \infty, \end{matrix}

with the convergence on the last display following from (4.12). Applying this to (4.14), we have:

2 n h^{2} (f_{n}, f) = n \int_{l_{n}} {\frac{f_{n} (x) - f (x)}{\sqrt{f_{n} (x)} + \sqrt{f (x)}}}^{2} d x + o (1)

(4.15)

as n → ∞, because $0 \leq \int_{I_{n}^{c}} f (x) d x \leq 1$

For fixed $n \in N$ , such that f and g_n be continuous and strictly positive on I_n, let x_(n) and x⁽ⁿ⁾ denote, respectively, a minimizer and a maximizer of f on the compact set I_n. Let also y_(n) and y_(n) denote, respectively, a minimizer and a maximizer of g_n on the compact set I_n. Observe that, since I_n is a decreasing sequence of compact sets converging to {x₀}, all of x_(n), x⁽ⁿ⁾, y_(n) and y⁽ⁿ⁾ converage to x₀ as n → ∞. Also,

\begin{matrix} \sup_{x \in l_{n}} ∣ \frac{f_{n} (x) - f (x)}{f (x)} ∣ & = \sup_{x \in l_{n}} ∣ (\frac{1}{d_{n}} - 1) + \frac{θ g_{n} (x)}{d_{n} f (x)} ∣ \\ \leq (1 - \frac{1}{d_{n}}) + \frac{θ \sup_{x \in l_{n}} {g_{n} (x)}}{d_{n} \inf_{x \in l_{n}} {f (x)}} \\ \to 0, as n \to \infty, \end{matrix}

(4.16)

because g_n is pointwise non-increasing in $n \in N, g_{n} (x_{0}) = O (n^{- 1 ∕ 3})$ and f(x₀) > 0.

Also,

\begin{matrix} D_{1} (n) & \equiv \int_{l_{n}} {f_{n} (x) - f (x)}^{2} d x \\ = \frac{1}{d_{n}^{2}} \int_{l_{n}} {θ^{2} g_{n}^{2} (x) - O (n^{- \frac{2}{3}}) f (x) g_{n} (x) + O (n^{- \frac{4}{3}}) f^{2} (x)} d x \end{matrix}

and noticing that

0 \leq \int_{l_{n}} {g_{n} (x) f (x)} d x \leq f (x^{(n)}) \int_{l_{n}} ({g_{n} (x)} d x = O (n^{- \frac{2}{3}}),

so that,

\begin{matrix} n D_{1} (n) = & \frac{n}{d_{n}^{2}} {{(\frac{8}{3})}^{d} θ^{2} b^{2} \prod_{i = 1}^{d} {h_{i}^{3}} \cdot n^{- 1} + o (n^{- \frac{4}{3}})} \\ \to {(\frac{8}{3})}^{d} θ^{2} b^{2} \prod_{i = 1}^{d} {h_{i}^{3}}, as n \to \infty . \end{matrix}

(4.17)

Now, since f is block-decreasing, we have,

0 < f (x_{0} + n^{- \frac{1}{3 d}} I_{d} h) \leq f (x) \leq f (x_{0} - n^{- \frac{1}{3 d}} I_{d} h)

for all x ∈ I_n and n ≥ n₁. Hence,

\frac{n D_{1} (n)}{f (x_{0} - n^{- \frac{1}{3 d}} I_{d} h)} \leq n \int_{l_{n}} \frac{{f_{n} (x) - f (x)}^{2}}{f (x)} d x \leq \frac{n D_{1} (n)}{f (x_{0} + n^{- \frac{1}{3 d}} I_{d} h)}

which, ahead with Eq. (4.17) and sandwich, yields

n \int_{l_{n}} \frac{{f_{n} (x) - f (x)}^{2}}{f (x)} d x \to {(\frac{8}{3})}^{d} θ^{2} b^{2} \cdot \frac{\prod_{i = 1}^{d} {h_{i}^{3}}}{f (x_{0})}, as n \to \infty .

Applying all of the above to (4.15), and appealing to Lemma 2 of [29], we get

n h^{2} (f_{n}, f) = \frac{1}{8} \int_{l_{n}} \frac{{f_{n} (x) - f (x)}^{2}}{f (x)} d x + o (1)

(4.18)

\to \frac{8^{d - 1}}{3^{d} f (x_{0})} θ^{2} b^{2} \prod_{i = 1}^{d} {h_{i}^{3}}

(4.19)

as n → ∞, so that by applying (4.13) and (4.19) to Lemma 4.2, we get

\begin{matrix} \underset{n \to \infty}{\lim_{¯}} & \inf_{T_{n}} \max {E_{f_{n}} {n^{\frac{1}{3}} ∣ T_{n} - f_{n} (x_{0}) ∣}, E_{f} {n^{\frac{1}{3}} ∣ T_{n} - f (x_{0}) ∣}} \\ \geq \frac{1}{4} {{(- 1)}^{d} b} θ c \exp {- \frac{2^{3 d - 2}}{3^{d} f (x_{0})} θ^{2} b^{2} c^{3}} ≕ G_{f, x_{0}} (c, θ) \end{matrix}

where $c \equiv Π_{i = 1}^{d} {h_{i}}$ . For a fixed θ ∈ (0, 1) the maximum of G_f,x₀(c, θ) is attained at

c (θ) = {\frac{3^{d - 1} f (x_{0})}{2^{3 d - 2} θ^{2} b^{2}}}^{\frac{1}{3}}

and is equal to

G_{f} (c (θ), θ) = \frac{e^{- \frac{1}{3}}}{2^{d}} {3^{d - 1} θ}^{\frac{1}{3}} {{(- 1)}^{d} {\frac{\partial^{d} f (x)}{\partial x_{1} \dots \partial x_{d}} ∣}_{x = x_{0}} f (x_{0})}^{\frac{1}{3}},

the latter being an increasing function of θ ∈ (0, 1).

This implies that

\begin{matrix} \underset{n \to \infty}{\lim_{¯}} & \inf_{T_{n}} \max {E_{f_{n}} {n^{\frac{1}{3}} ∣ T_{n} - f_{n} (x_{0}) ∣}, E_{f} {n^{\frac{1}{3}} ∣ T_{n} - f (x_{0}) ∣}} \\ \geq \frac{e^{- \frac{1}{3}}}{2^{d}} {θ \cdot 3^{d - 1}}^{\frac{1}{3}} {{(- 1)}^{d} {\frac{\partial^{d} f (x)}{\partial x_{1} \dots \partial x_{d}} ∣}_{x = x_{0}} \cdot f (x_{0})}^{\frac{1}{3}} . \end{matrix}

Overall, we are allowed to take θ ↑ 1 in the above display, even if θ = 1 is not a valid configuration, yielding the lower bound in the wording of the proposition. The proof is thus complete.

5. Discussion and open problems

Once consistency has been established, interest focuses on rates of convergence of the MLE and other properties, including the behavior of ${\hat{f}}_{n}$ at zero and pointwise limiting distributions. We have the following conjectures concerning the MLE ${\hat{f}}_{n}$ for the class $F_{SMU} (d)$ . Work is currently underway on all of these further problems.

Conjecture 1

If f₀(0) < ∞, then we conjecture that $P_{0} ({\hat{f}}_{n} (0) \leq M {(\log n)}^{d - 1}) \to 1 f o r s o m e M > 0$ .

Conjecture 2

If f₀(0) < ∞ and f₀ is concentrated on [0, M1] for some 0 < M < ∞, then $h ({\hat{f}}_{n}, f_{0}) = O_{p} (n^{- 1 ∕ 3} {(\log n)}^{Y})$ for some γ depending only on d.

Concerning rates of convergence of the estimators at a fixed point, we do not yet have any upper bound results to accompany the lower bound results of Proposition 4.1. Thus there remain the following two possibilities: (a) the pointwise rate of convergence under Assumption 4.1 is n^1/3, and we expect convergence in distribution with the rate n^1/3, or, (b) the lower bound given in Proposition 4.1 is not yet sharp, and we should expect log terms in the rate (as might be expected from the covering number results of [10]). Our corresponding conjectures for these two possible scenarios are given below as Conjectures 3a and 3b respectively.

Conjecture 3a

Suppose that f₀ has ∂^df₀(x)/∂x₁…∂x_d continuous in a neighborhood of x₀ with

\partial^{d} f_{0} (x_{0}) \equiv {\frac{\partial^{d} f_{0} (x)}{\partial x_{1} \dots \partial x_{d}} ∣}_{x = x_{0}} \neq 0 .

Let ${W (t) : t \in R^{d}}$ be a 2^d-sided Brownian sheet process on $R^{d}$ and let

Y (t) \equiv \sqrt{f_{0} (x_{0})} W (t) + \frac{{(- 1)}^{d}}{2^{d}} {(- 1)}^{d} \partial^{d} f_{0} (x_{0}) {∣ t ∣}^{2} .

Then, in keeping with our lower bound results of Section 4, we conjecture that

n^{1 ∕ 3} ({\hat{f}}_{n} (x_{0}) - f_{0} (x_{0})) \to {}_{d}\partial^{d} H (t) ∣_{t = 0}

where the process $H$ is determined by

$H (t) \geq Y (t) for all t \in R^{d},$
$\int_{R^{d}} (H (t) - Y (t)) d (\partial^{d} H (t)) = 0, and$
$V_{\partial^{d} H} [u, v) \geq 0 for all u \leq v \in R^{d} .$

Partial results concerning Conjecture 3a were obtained in [37].

Conjecture 3b

As suggested in part by the covering number results of [10], the pointwise rate of convergence is (n/(log n)^d–1/2)^1/3. This would entail an improved version of Proposition 4.1. In this case, we do not yet have conjectures concerning the limiting distribution.

Acknowledgments

We owe thanks to Marina Meila, Fritz Scholz, and Arseni Seregin for helpful discussions concerning the proof of uniqueness, and especially Lemmas 3.3 and 3.4. We also thank the referees for several helpful suggestions and for catching a slip in a proof in the first version of the paper. The first author’s research was supported by NSF grant DMS-0503822. The second author’s research was supported by NSF grants DMS-0503822 and DMS-0804587 and NIH/NIAID grants 2R01 AI029168 and 4 R37 AI029168.

Footnotes

AMS 2000 subject classifications: 62G05, 62G07, 62G20, 62F20, 62H12

References

[1].Anevski D. Technical Report, dept. of Math. Statistics. Univ. of Lund; 1994. Estimating the derivative of a convex density. [Google Scholar]
[2].Anevski D. Estimating the derivative of a convex density. Stat. Neerl. 2003;57(2):245–257. [Google Scholar]
[3].Ayer M, Brunk HD, Ewing GM, Reid WT, Silverman E. An empirical distribution function for sampling with incomplete information. Ann. Math. Statist. 1955;26:641–647. [Google Scholar]
[4].Balabdaoui F, Jankowski H, Pavlides M, Seregin A, Wellner JA. On the Grenander estimator at zero. Statist. Sinica. 2011;21:873–879. doi: 10.5705/ss.2011.038a. [DOI] [PMC free article] [PubMed] [Google Scholar]
[5].Barlow RE, Bartholomew DJ, Bremner JM, Brunk HD. Wiley Series in Probability and Mathematical Statistics. John Wiley & Sons; London-New York-Sydney: 1972. Statistical inference under order restrictions, in: The Theory and Application of Isotonic Regression. [Google Scholar]
[6].Biau G, Devroye L. On the risk of estimates for block decreasing densities. J. Multivariate Anal. 2003;86(1):143–165. [Google Scholar]
[7].Billingsley P. Wiley Series in Probability and Mathematical Statistics. third ed. John Wiley & Sons Inc.; New York: 1995. Probability and measure. a Wiley-Interscience Publication. [Google Scholar]
[8].Birgé L. Estimating a density under order restrictions: nonasymptotic minimax risk. Ann. Statist. 1987;15(3):995–1012. [Google Scholar]
[9].Birgé L. The Grenander estimator: a nonasymptotic approach. Ann. Statist. 1989;17(4):1532–1549. [Google Scholar]
[10].Blei R, Gao F, Li WV. Metric entropy of high dimensional distributions. Proc. Amer. Math. Soc. 2007;135(12):4009–4018. electronic. [Google Scholar]
[11].Brunk HD. On the estimation of parameters restricted by inequalities. Ann. Math. Statist. 1958;29:437–454. [Google Scholar]
[12].Brunk HD. Nonparametric Techniques in Statistical Inference. Cambridge Univ. Press; London: 1970. Estimation of isotonic regression; pp. 177–197. Proc. Sympos., Indiana Univ., Bloomington, Ind., 1969. [Google Scholar]
[13].Devroye L. Nonuniform Random Variate Generation. Springer-Verlag; New York: 1986. [Google Scholar]
[14].Devroye L. A Course in Density Estimation, in: Progress in Probability and Statistics. Vol. 14. Birkhäuser Boston Inc; Boston, MA: 1987. [Google Scholar]
[15].Doob JL. Graduate Texts in Mathematics. Vol. 143. Springer-Verlag; New York: 1994. Measure Theory. [Google Scholar]
[16].Feller W. An Introduction to Probability Theory and Its Applications. second ed II. John Wiley & Sons Inc.; New York: 1971. [Google Scholar]
[17].Glick N. Consistency conditions for probability estimators and integrals of density estimators. Util. Math. 1974;6:61–74. [Google Scholar]
[18].Grenander U. On the theory of mortality measurement. I. Skand. Aktuarietidskr. 1956;39:70–96. [Google Scholar]
[19].Grenander U. On the theory of mortality measurement. II, Skand. Aktuarietidskr. 1957;39:125–153. [Google Scholar]
[20].Groeneboom P. Estimating a monotone density. Wadsworth Statist./Probab. Ser. Wadsworth; Proceedings of the Berkeley Conference in Honor of Jerzy Neyman and Jack Kiefer; Berkeley, Calif. 1983; pp. 539–555. Belmont, CA. [Google Scholar]
[21].Groeneboom P. Brownian motion with a parabolic drift and Airy functions. Probab. Theory Related Fields. 1989;81(1):79–109. [Google Scholar]
[22].Groeneboom P. Lecture Notes in Math. Vol. 1648. Springer; Berlin: 1996. Lectures on inverse problems, in: Lectures on Probability Theory and Statistics; pp. 67–164. Saint-Flour, 1994. [Google Scholar]
[23].Groeneboom P, Jongbloed G. Isotonic estimation and rates of convergence in Wicksell’s problem. Ann. Statist. 1995;23(5):1518–1542. [Google Scholar]
[24].Groeneboom P, Jongbloed G, Wellner JA. A canonical process for estimation of convex functions: the invelope of integrated Brownian motion +t4. Ann. Statist. 2001;29(6):1620–1652. [Google Scholar]
[25].Groeneboom P, Jongbloed G, Wellner JA. Estimation of a convex function: characterizations and asymptotic theory. Ann. Statist. 2001;29(6):1653–1698. [Google Scholar]
[26].Hampel FR. Design, Data & Analysis. John Wiley & Sons, Inc.; New York, NY, USA: 1987. Design, modelling, and analysis of some biological data sets; pp. 93–128. [Google Scholar]
[27].Jewell P. Nicholas, van der Laan Mark. Handbook of Statist. Vol. 23. Elsevier; Amsterdam: 2004. Current status data: review, recent developments and open problems; pp. 625–642. Advances in Survival Analysis. [Google Scholar]
[28].Jongbloed G. Ph.D. Thesis. Delft University; 1995. Three statistical inverse problems. [Google Scholar]
[29].Jongbloed G. Minimax lower bounds and moduli of continuity. Statist. Probab. Lett. 2000;50(3):279–284. [Google Scholar]
[30].Kim J, Pollard D. Cube root asymptotics. Ann. Statist. 1990;18(1):191–219. [Google Scholar]
[31].Lang R. A note on the measurability of convex sets. Arch. Math. (Basel) 1986;47(1):90–92. [Google Scholar]
[32].Lavee D, Safrie UN, Meilijson I. For how long do trans-saharan migrants stop over at an oasis? Ornis Scandinavica. 1991;22:33–44. [Google Scholar]
[33].Le Cam L. Springer Series in Statistics. Springer-Verlag; New York: 1986. Asymptotic Methods in Statistical Decision Theory. [Google Scholar]
[34].Lindsay BG. The geometry of mixture likelihoods: a general theory. Ann. Statist. 1983;11(1):86–94. [Google Scholar]
[35].Lindsay BG. Mixture Models: Theory, Geometry and Applications; NSF-CBMS Regional Conference Series in Probability and Statistics; IMS, Hayward CA. 1995. [Google Scholar]
[36].Müller DW, Sawitzki G. Excess mass estimates and tests for multimodality. J. Amer. Statist. Assoc. 1991;86(415):738–746. [Google Scholar]
[37].Pavlides M. Ph.D. Thesis. University of Washington; 2008. Nonparametric estimation of multivariate monotone densities. [Google Scholar]
[38].Pavlides M. Tech. rep. Frederick University; Nicosia, Cyprus: 2009. Local asymptotic minimax theory for block-decreasing densities. [Google Scholar]
[39].Pfanzagl J. Consistency of maximum likelihood estimators for certain nonparametric families, in particular: mixtures. J. Statist. Plann. Inference. 1988;19(2):137–158. [Google Scholar]
[40].Polonik W. Density estimation under qualitative assumptions in higher dimensions. J. Multivariate Anal. 1995;55(1):61–81. [Google Scholar]
[41].Polonik W. Measuring mass concentrations and estimating density contour clusters—an excess mass approach. Ann. Statist. 1995;23(3):855–881. [Google Scholar]
[42].Polonik W. Minimum volume sets and generalized quantile processes. Stochastic Process. Appl. 1997;69(1):1–24. [Google Scholar]
[43].Polonik W. The silhouette, concentration functions and ML-density estimation under order restrictions. Ann. Statist. 1998;26(5):1857–1877. [Google Scholar]
[44].Prakasa Rao BLS. Estimation of a unimodal density. Sankhyā Ser. A. 1969;31:23–36. [Google Scholar]
[45].Robertson T. On estimating a density which is measurable with respect to a σ-lattice. Ann. Math. Statist. 1967;38:482–493. [Google Scholar]
[46].Robertson T, Wright FT, Dykstra RL. Probability and Mathematical Statistics. John Wiley & Sons Ltd; Chichester: 1988. Order restricted statistical inference. Wiley Series in Probability and Mathematical Statistics. [Google Scholar]
[47].Rockafellar RT. Princeton Mathematical Series. Princeton University Press; Princeton, N.J: 1970. Convex Analysis. No. 28. [Google Scholar]
[48].Sager TW. An iterative method for estimating a multivariate mode and isopleth. J. Amer. Statist. Assoc. 1979;74:329–339. 366. part 1. [Google Scholar]
[49].Sager TW. Nonparametric maximum likelihood estimation of spatial patterns. Ann. Statist. 1982;10(4):1125–1136. [Google Scholar]
[50].Shorack GR. Springer Texts in Statistics. Springer-Verlag; New York: 2000. Probability for statisticians. [Google Scholar]
[51].van Eeden C. Statist. Afdeling S 188 (VP 5). Math. Centrum Amsterdam; 1956. Maximum likelihood estimation of ordered probabilities. [Google Scholar]
[52].van Eeden C. Maximum likelihood estimation of ordered probabilities. Nederl. Akad. Wetensch. Proc. Ser. A. 59 = Indag. Math. 1956;18:444–455. [Google Scholar]
[53].van Eeden C. II. Statist. Afdeling Rep. S 196 (VP7). Math. Centrum Amsterdam; 1956. Maximum likelihood estimation of ordered probabilities. [Google Scholar]
[54].van Eeden C. Maximum likelihood estimation of partially or completely ordered parameters. I, Nederl. Akad. Wetensch. Proc. Ser. A. 60 = Indag. Math. 1957;19:128–136. [Google Scholar]
[55].van Eeden C. Maximum likelihood estimation of partially or completely ordered parameters. II, Nederl. Akad. Wetensch. Proc. Ser. A. 60 = Indag. Math. 1957;19:201–211. [Google Scholar]
[56].Wegman EJ. A note on estimating a unimodal density. Ann. Math. Statist. 1969;40:1661–1667. [Google Scholar]
[57].Wegman EJ. Maximum likelihood estimation of a unimodal density function. Ann. Math. Statist. 1970;41:457–471. [Google Scholar]
[58].Wegman EJ. Maximum likelihood estimation of a unimodal density. II, Ann. Math. Statist. 1970;41:2169–2174. [Google Scholar]
[59].Williamson RE. Multiply monotone functions and their Laplace transforms. Duke Math. J. 1956;23:189–207. [Google Scholar]
[60].Woodroofe M, Sun J. A penalized maximum likelihood estimate of f(0+) when f is nonincreasing. Statist. Sinica. 1993;3(2):501–515. [Google Scholar]
[61].Wong GYC, Yu Q. Generalized MLE of a joint distribution function with multivariate interval-censored data. J. Multivariate Anal. 1999;69:155–166. [Google Scholar]
[62].Shaohua Yu, Qiqing Yu, Wong YC. George, Consistency of the generalized MLE of a joint distribution function with multivariate interval-censored data. J. Multivariate Anal. 2005;97:720–732. [Google Scholar]

[R1] [1].Anevski D. Technical Report, dept. of Math. Statistics. Univ. of Lund; 1994. Estimating the derivative of a convex density. [Google Scholar]

[R2] [2].Anevski D. Estimating the derivative of a convex density. Stat. Neerl. 2003;57(2):245–257. [Google Scholar]

[R3] [3].Ayer M, Brunk HD, Ewing GM, Reid WT, Silverman E. An empirical distribution function for sampling with incomplete information. Ann. Math. Statist. 1955;26:641–647. [Google Scholar]

[R4] [4].Balabdaoui F, Jankowski H, Pavlides M, Seregin A, Wellner JA. On the Grenander estimator at zero. Statist. Sinica. 2011;21:873–879. doi: 10.5705/ss.2011.038a. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R5] [5].Barlow RE, Bartholomew DJ, Bremner JM, Brunk HD. Wiley Series in Probability and Mathematical Statistics. John Wiley & Sons; London-New York-Sydney: 1972. Statistical inference under order restrictions, in: The Theory and Application of Isotonic Regression. [Google Scholar]

[R6] [6].Biau G, Devroye L. On the risk of estimates for block decreasing densities. J. Multivariate Anal. 2003;86(1):143–165. [Google Scholar]

[R7] [7].Billingsley P. Wiley Series in Probability and Mathematical Statistics. third ed. John Wiley & Sons Inc.; New York: 1995. Probability and measure. a Wiley-Interscience Publication. [Google Scholar]

[R8] [8].Birgé L. Estimating a density under order restrictions: nonasymptotic minimax risk. Ann. Statist. 1987;15(3):995–1012. [Google Scholar]

[R9] [9].Birgé L. The Grenander estimator: a nonasymptotic approach. Ann. Statist. 1989;17(4):1532–1549. [Google Scholar]

[R10] [10].Blei R, Gao F, Li WV. Metric entropy of high dimensional distributions. Proc. Amer. Math. Soc. 2007;135(12):4009–4018. electronic. [Google Scholar]

[R11] [11].Brunk HD. On the estimation of parameters restricted by inequalities. Ann. Math. Statist. 1958;29:437–454. [Google Scholar]

[R12] [12].Brunk HD. Nonparametric Techniques in Statistical Inference. Cambridge Univ. Press; London: 1970. Estimation of isotonic regression; pp. 177–197. Proc. Sympos., Indiana Univ., Bloomington, Ind., 1969. [Google Scholar]

[R13] [13].Devroye L. Nonuniform Random Variate Generation. Springer-Verlag; New York: 1986. [Google Scholar]

[R14] [14].Devroye L. A Course in Density Estimation, in: Progress in Probability and Statistics. Vol. 14. Birkhäuser Boston Inc; Boston, MA: 1987. [Google Scholar]

[R15] [15].Doob JL. Graduate Texts in Mathematics. Vol. 143. Springer-Verlag; New York: 1994. Measure Theory. [Google Scholar]

[R16] [16].Feller W. An Introduction to Probability Theory and Its Applications. second ed II. John Wiley & Sons Inc.; New York: 1971. [Google Scholar]

[R17] [17].Glick N. Consistency conditions for probability estimators and integrals of density estimators. Util. Math. 1974;6:61–74. [Google Scholar]

[R18] [18].Grenander U. On the theory of mortality measurement. I. Skand. Aktuarietidskr. 1956;39:70–96. [Google Scholar]

[R19] [19].Grenander U. On the theory of mortality measurement. II, Skand. Aktuarietidskr. 1957;39:125–153. [Google Scholar]

[R20] [20].Groeneboom P. Estimating a monotone density. Wadsworth Statist./Probab. Ser. Wadsworth; Proceedings of the Berkeley Conference in Honor of Jerzy Neyman and Jack Kiefer; Berkeley, Calif. 1983; pp. 539–555. Belmont, CA. [Google Scholar]

[R21] [21].Groeneboom P. Brownian motion with a parabolic drift and Airy functions. Probab. Theory Related Fields. 1989;81(1):79–109. [Google Scholar]

[R22] [22].Groeneboom P. Lecture Notes in Math. Vol. 1648. Springer; Berlin: 1996. Lectures on inverse problems, in: Lectures on Probability Theory and Statistics; pp. 67–164. Saint-Flour, 1994. [Google Scholar]

[R23] [23].Groeneboom P, Jongbloed G. Isotonic estimation and rates of convergence in Wicksell’s problem. Ann. Statist. 1995;23(5):1518–1542. [Google Scholar]

[R24] [24].Groeneboom P, Jongbloed G, Wellner JA. A canonical process for estimation of convex functions: the invelope of integrated Brownian motion +t4. Ann. Statist. 2001;29(6):1620–1652. [Google Scholar]

[R25] [25].Groeneboom P, Jongbloed G, Wellner JA. Estimation of a convex function: characterizations and asymptotic theory. Ann. Statist. 2001;29(6):1653–1698. [Google Scholar]

[R26] [26].Hampel FR. Design, Data & Analysis. John Wiley & Sons, Inc.; New York, NY, USA: 1987. Design, modelling, and analysis of some biological data sets; pp. 93–128. [Google Scholar]

[R27] [27].Jewell P. Nicholas, van der Laan Mark. Handbook of Statist. Vol. 23. Elsevier; Amsterdam: 2004. Current status data: review, recent developments and open problems; pp. 625–642. Advances in Survival Analysis. [Google Scholar]

[R28] [28].Jongbloed G. Ph.D. Thesis. Delft University; 1995. Three statistical inverse problems. [Google Scholar]

[R29] [29].Jongbloed G. Minimax lower bounds and moduli of continuity. Statist. Probab. Lett. 2000;50(3):279–284. [Google Scholar]

[R30] [30].Kim J, Pollard D. Cube root asymptotics. Ann. Statist. 1990;18(1):191–219. [Google Scholar]

[R31] [31].Lang R. A note on the measurability of convex sets. Arch. Math. (Basel) 1986;47(1):90–92. [Google Scholar]

[R32] [32].Lavee D, Safrie UN, Meilijson I. For how long do trans-saharan migrants stop over at an oasis? Ornis Scandinavica. 1991;22:33–44. [Google Scholar]

[R33] [33].Le Cam L. Springer Series in Statistics. Springer-Verlag; New York: 1986. Asymptotic Methods in Statistical Decision Theory. [Google Scholar]

[R34] [34].Lindsay BG. The geometry of mixture likelihoods: a general theory. Ann. Statist. 1983;11(1):86–94. [Google Scholar]

[R35] [35].Lindsay BG. Mixture Models: Theory, Geometry and Applications; NSF-CBMS Regional Conference Series in Probability and Statistics; IMS, Hayward CA. 1995. [Google Scholar]

[R36] [36].Müller DW, Sawitzki G. Excess mass estimates and tests for multimodality. J. Amer. Statist. Assoc. 1991;86(415):738–746. [Google Scholar]

[R37] [37].Pavlides M. Ph.D. Thesis. University of Washington; 2008. Nonparametric estimation of multivariate monotone densities. [Google Scholar]

[R38] [38].Pavlides M. Tech. rep. Frederick University; Nicosia, Cyprus: 2009. Local asymptotic minimax theory for block-decreasing densities. [Google Scholar]

[R39] [39].Pfanzagl J. Consistency of maximum likelihood estimators for certain nonparametric families, in particular: mixtures. J. Statist. Plann. Inference. 1988;19(2):137–158. [Google Scholar]

[R40] [40].Polonik W. Density estimation under qualitative assumptions in higher dimensions. J. Multivariate Anal. 1995;55(1):61–81. [Google Scholar]

[R41] [41].Polonik W. Measuring mass concentrations and estimating density contour clusters—an excess mass approach. Ann. Statist. 1995;23(3):855–881. [Google Scholar]

[R42] [42].Polonik W. Minimum volume sets and generalized quantile processes. Stochastic Process. Appl. 1997;69(1):1–24. [Google Scholar]

[R43] [43].Polonik W. The silhouette, concentration functions and ML-density estimation under order restrictions. Ann. Statist. 1998;26(5):1857–1877. [Google Scholar]

[R44] [44].Prakasa Rao BLS. Estimation of a unimodal density. Sankhyā Ser. A. 1969;31:23–36. [Google Scholar]

[R45] [45].Robertson T. On estimating a density which is measurable with respect to a σ-lattice. Ann. Math. Statist. 1967;38:482–493. [Google Scholar]

[R46] [46].Robertson T, Wright FT, Dykstra RL. Probability and Mathematical Statistics. John Wiley & Sons Ltd; Chichester: 1988. Order restricted statistical inference. Wiley Series in Probability and Mathematical Statistics. [Google Scholar]

[R47] [47].Rockafellar RT. Princeton Mathematical Series. Princeton University Press; Princeton, N.J: 1970. Convex Analysis. No. 28. [Google Scholar]

[R48] [48].Sager TW. An iterative method for estimating a multivariate mode and isopleth. J. Amer. Statist. Assoc. 1979;74:329–339. 366. part 1. [Google Scholar]

[R49] [49].Sager TW. Nonparametric maximum likelihood estimation of spatial patterns. Ann. Statist. 1982;10(4):1125–1136. [Google Scholar]

[R50] [50].Shorack GR. Springer Texts in Statistics. Springer-Verlag; New York: 2000. Probability for statisticians. [Google Scholar]

[R51] [51].van Eeden C. Statist. Afdeling S 188 (VP 5). Math. Centrum Amsterdam; 1956. Maximum likelihood estimation of ordered probabilities. [Google Scholar]

[R52] [52].van Eeden C. Maximum likelihood estimation of ordered probabilities. Nederl. Akad. Wetensch. Proc. Ser. A. 59 = Indag. Math. 1956;18:444–455. [Google Scholar]

[R53] [53].van Eeden C. II. Statist. Afdeling Rep. S 196 (VP7). Math. Centrum Amsterdam; 1956. Maximum likelihood estimation of ordered probabilities. [Google Scholar]

[R54] [54].van Eeden C. Maximum likelihood estimation of partially or completely ordered parameters. I, Nederl. Akad. Wetensch. Proc. Ser. A. 60 = Indag. Math. 1957;19:128–136. [Google Scholar]

[R55] [55].van Eeden C. Maximum likelihood estimation of partially or completely ordered parameters. II, Nederl. Akad. Wetensch. Proc. Ser. A. 60 = Indag. Math. 1957;19:201–211. [Google Scholar]

[R56] [56].Wegman EJ. A note on estimating a unimodal density. Ann. Math. Statist. 1969;40:1661–1667. [Google Scholar]

[R57] [57].Wegman EJ. Maximum likelihood estimation of a unimodal density function. Ann. Math. Statist. 1970;41:457–471. [Google Scholar]

[R58] [58].Wegman EJ. Maximum likelihood estimation of a unimodal density. II, Ann. Math. Statist. 1970;41:2169–2174. [Google Scholar]

[R59] [59].Williamson RE. Multiply monotone functions and their Laplace transforms. Duke Math. J. 1956;23:189–207. [Google Scholar]

[R60] [60].Woodroofe M, Sun J. A penalized maximum likelihood estimate of f(0+) when f is nonincreasing. Statist. Sinica. 1993;3(2):501–515. [Google Scholar]

[R61] [61].Wong GYC, Yu Q. Generalized MLE of a joint distribution function with multivariate interval-censored data. J. Multivariate Anal. 1999;69:155–166. [Google Scholar]

[R62] [62].Shaohua Yu, Qiqing Yu, Wong YC. George, Consistency of the generalized MLE of a joint distribution function with multivariate interval-censored data. J. Multivariate Anal. 2005;97:720–732. [Google Scholar]

PERMALINK

Nonparametric estimation of multivariate scale mixtures of uniform densities

Marios G Pavlides

Jon A Wellner

Abstract

1. Introduction and summary

2. Properties of the Scale Mixtures of Uniform family of densities

2.1. Properties of FSMU(d)

Lemma 2.1

Definition 2.1

Lemma 2.2

Theorem 2.3

2.2. Lebesgue measurability of block-decreasing functions

Proposition 2.4

3. Existence and consistency of the MLE

3.1. On existence and uniqueness of an MLE

Definition 3.1 (Rectangular Grid Generated by Data)

Theorem 3.1 (Existence and Characterization of an MLE in FSMU(d))

Lemma 3.2

Lemma 3.3

Lemma 3.4

Example 3.1 (A MLE in FSMU is Not Always Unique)

Fig. 1.

Fig. 2.

3.2. Strong pointwise consistency of the MLE

Proposition 3.5 (Strong Consistency of the MLE in FSMU)

4. A local asymptotic minimax lower bound

Assumption 4.1

Proposition 4.1

Remark

Lemma 4.2.

Fig. 3.

Proposition 4.3

Fig. 4.

5. Discussion and open problems

Conjecture 1

Conjecture 2

Conjecture 3a

Conjecture 3b

Acknowledgments

Footnotes

References

ACTIONS

PERMALINK

RESOURCES

Similar articles

Cited by other articles

Links to NCBI Databases

2.1. Properties of $F_{S M U} (d)$

Theorem 3.1 (Existence and Characterization of an MLE in $F_{S M U} (d)$ )

Example 3.1 (A MLE in $F_{SMU}$ is Not Always Unique)

Proposition 3.5 (Strong Consistency of the MLE in $F_{SMU}$ )