SCALING LIMITS OF A MODEL FOR SELECTION AT TWO SCALES

SHISHI LUO; JONATHAN C MATTINGLY

doi:10.1088/1361-6544/aa5499

. Author manuscript; available in PMC: 2017 Sep 1.

Published in final edited form as: Nonlinearity. 2017 Mar 15;30(4):1682–1707. doi: 10.1088/1361-6544/aa5499

SCALING LIMITS OF A MODEL FOR SELECTION AT TWO SCALES

SHISHI LUO ^1,³, JONATHAN C MATTINGLY ²

PMCID: PMC5580332 NIHMSID: NIHMS897715 PMID: 28867875

Abstract

The dynamics of a population undergoing selection is a central topic in evolutionary biology. This question is particularly intriguing in the case where selective forces act in opposing directions at two population scales. For example, a fast-replicating virus strain outcompetes slower-replicating strains at the within-host scale. However, if the fast-replicating strain causes host morbidity and is less frequently transmitted, it can be outcompeted by slower-replicating strains at the between-host scale. Here we consider a stochastic ball-and-urn process which models this type of phenomenon. We prove the weak convergence of this process under two natural scalings. The first scaling leads to a deterministic nonlinear integro-partial differential equation on the interval [0, 1] with dependence on a single parameter, λ. We show that the fixed points of this differential equation are Beta distributions and that their stability depends on λ and the behavior of the initial data around 1. The second scaling leads to a measure-valued Fleming-Viot process, an infinite dimensional stochastic process that is frequently associated with a population genetics.

Keywords: Markov chains, limiting behavior, evolutionary dynamics, Fleming–Viot process, scaling limits

1. Introduction

We study the model, introduced in [15], of a trait that is advantageous at a local or individual level but disadvantageous at a larger scale or group level. For example, an infectious virus strain that replicates rapidly within its host will outcompete other virus strains in the host. However, if infection with a heavy viral load is incapacitating and prevents the host from transmitting the virus, the rapidly replicating strain may not be as prevalent in the overall host population as a slow replicating strain.

A simple mathematical formulation of this phenomenon is as follows. Consider a population of m ∈ $ℕ$ groups. Each group contains n ∈ $ℕ$ individuals. There are two types of individuals: type I individuals are selectively advantageous at the individual (I) level and type G individuals are selectively advantageous at the group (G) level. Replication and selection occur concurrently at the individual and group level according to the Moran process [8] and are illustrated in Fig 1. Type I individuals replicate at rate 1 + s, s ≥ 0 and type G individuals at rate 1. When an individual gives birth, another individual in the same group is selected uniformly at random to die. To reflect the antagonism at the higher level of selection, groups replicate at a rate which increases with the number of type G indivduals they contain. As a simple case, we take this rate to be $w (1 + r \frac{k}{n})$ , where $\frac{k}{n}$ is the fraction of indivduals in the group that are type G, r ≥ 0 is the selection coefficient at the group level, and w > 0 is the ratio of the rate of group-level events to the rate of individual-level events. As with the individual level, the population of groups is maintained at m by selecting a group uniformly at random to die whenever a group replicates. The offspring of groups are assumed to be identical to their parent.

Schematic of the particle process. (a) *Left*: A population of m = 3 groups, each with n = 3 individuals of either type G (filled small circles) or type I (open small circles). *Middle*: A type I individual replicates in group 3 and a type G individual is chosen uniformly at random from group 3 to die. Right: Group 1 replicates and produces group 2′. Group 2 is chosen uniformly at random to die. (b) The states in (a) mapped to a particle process. *Left*: Group 2 has no type G individuals, represented by ball 2 in urn 0. Similarly, group 3 is represented by ball 3 in urn 2 and group 1 by ball 1 in urn 3. *Middle*: The number of type G individuals in group 3 decreases from two to one, therefore ball 3 moves to urn 1. *Right*: A group with zero type G individuals dies, while a group with three type G individuals is born. Therefore ball 2 leaves urn 0 and appears in urn 3 as ball 2′.

As illustrated in Fig 1, this two-level process is equivalent to a ball-and-urn or particle process, where each particle represents a group and its position corresponds to the number of type G individuals that are in it. We note that similar, though more general, particle models of evolutionary and ecological dynamics at multiple scales have been studied and we mention several particularly relevant works here. Dawson and Hochberg [6] also consider a population at multiple levels, albeit with one type per level, not two. In [4], Dawson and Greven consider a more general model, allowing for infinitely many hierarchical levels and for migration, selection, and mutation. Méléard and Roelly [16], in a study also inspired by host-pathogen interactions, investigate a model that allows for non-constant host and pathogen populations as well as mutation. In these more general contexts, determining long-term behavior is less straightforward than in our specific setting.

We now define the stochastic process that is the focus of this work. Let $X_{t}^{i}$ be the number of type G individuals in group i at time t. Then

μ_{t}^{m, n} : = \frac{1}{m} \sum_{i = 1}^{m} δ_{X_{t}^{i} / n}

is the empirical measure at time t for a given number of groups m and individuals per group n. δ_x(y) = 1 if x = y and zero otherwise. The $X_{t}^{i}$ are divided by n so that $μ_{t}^{m, n}$ is a probability measure on $E_{n} : = {0, \frac{1}{n}, \dots, 1}$ .

For fixed T > 0, $μ_{t}^{m, n} \in D ([0, T], P (E_{n}))$ , the set of càdlàg processes on [0, T] taking values in $P (E_{n})$ , where $P (S)$ is the set of probability measures on a set S. With the particle process described above, $μ_{t}^{m, n}$ has generator

(L^{m, n} ψ) (v) = \sum_{i, j} (R_{1} + w R_{2}) (v, v_{i, j}) [ψ (v_{i, j}) - ψ (v)]

(1)

where $v_{i j} : = v + \frac{1}{m} (δ_{\frac{j}{n}} - δ_{\frac{i}{n}})$ , $ψ \in C_{b} (P ([0, 1]))$ are bounded continuous functions, and $v \in P (E_{n}) \subset P ([0, 1])$ . The transition rates (R₁ + wR₂) are given by

R_{1} (v, v_{i j}) = {\begin{matrix} m v (\frac{i}{n}) i (1 - \frac{i}{n}) (1 + s) & if j = i - 1, i < n \\ m v (\frac{i}{n}) i (1 - \frac{i}{n}) & if j = i + 1, i > 0 \\ 0 & otherwise \end{matrix}}

and

R_{2} (v, v_{i j}) = m v (\frac{i}{n}) v (\frac{j}{n}) (1 + r \frac{j}{n}) .

R₁ represents individual-level events while R₂ represents group-level events.

2. Main results

We prove the weak convergence of this measure-valued process as m, n → ∞ under two natural scalings. The first scaling leads to a deterministic partial differential equation. We derive a closed-form expression for the solution of this equation and study its steady-state behavior. The second scaling leads to an infinite dimensional stochastic process, namely a Fleming-Viot process.

Let us briefly introduce some notation. By m, n → ∞ we mean a sequence {(m_k, n_k)}_k such that for any N, there is an n₀ such that if k ≥ n₀, m_k, n_k ≥ N. We define $〈 f, v 〉 = \int_{0}^{1} f (x) v (d x)$ where f is a test function and v a measure. Lastly, δ_x will denote the delta measure for both continuous and discrete state spaces.

To provide intuition for the two scalings and the corresponding limits, take ψ to be of the form ψ(v) = F (〈g, v〉), where g is some suitable function on [0, 1], and apply the generator in (1) to it:

L^{m, n} ψ (μ) = F^{'} \cdot {\sum_{i} [\frac{1}{n} g^{″} (\frac{i}{n}) - s g^{'} (\frac{i}{n})] \frac{i}{n} (1 - \frac{i}{n}) μ (\frac{i}{n}) + w r [\sum_{i} \frac{i}{n} g (\frac{i}{n}) μ (\frac{i}{n}) - \sum_{i} g (\frac{i}{n}) μ (\frac{i}{n}) \sum_{j} \frac{j}{n} μ (\frac{j}{n})]} + \frac{1}{m} w F^{″} \cdot {\sum_{i} g {(\frac{i}{n})}^{2} μ (\frac{i}{n}) - {(\sum_{i} g (\frac{i}{n}) μ (\frac{i}{n}))}^{2} + \frac{1}{2} r \sum_{i, j} {(g (\frac{j}{n}) - g (\frac{i}{n}))}^{2} \frac{j}{n} μ (\frac{j}{n}) μ (\frac{i}{n})} + o (\frac{1}{m}) + o (\frac{1}{n})

(2)

This suggests two natural scalings. The first is to take m, n → ∞ without rescaling any parameters. The g″ and F″ terms vanish and we have a deterministic process. The second is to let $s = \frac{σ}{n}$ , $r = \frac{ρ}{m}$ , and $\frac{n}{m} \to θ$ . The terms F″ and g″ no longer vanish and the process converges to a limit that is stochastic. The precise statement of the weak convergence of the finite state space system to the deterministic limit is in terms of a weak measure-valued solution to a partial differential equation:

Theorem 1

Suppose the particles in the system described by $μ_{t}^{m, n}$ are initially independently and identically distributed according to the measure $μ_{0}^{m, n}$ , where $μ_{0}^{m, n} \to μ_{0} \in P ([0, 1])$ as m, n → ∞. Then, as m, n → ∞, $μ_{t}^{m, n} \to μ_{t} \in D ([0, T], P ([0, 1]))$ , weakly, where μ_t solves the differential equation

\frac{d}{d t} 〈 f, μ_{t} 〉 = - 〈 x (1 - x) f^{'}, μ_{t} 〉 + λ [〈 x f, μ_{t} 〉 - 〈 f, μ_{t} 〉 〈 x, μ_{t} 〉]

(3)

for any positive-valued test function f ∈ C¹([0, 1]) and with initial condition 〈f, μ₀〉. Here, $λ : = \frac{w r}{s}$ and time has been sped up by a factor of s.

Throughout we will denote the measure-valued solutions to (3) by μ_t(dx). We note that strong, density-valued solutions, denoted by η_t(x), solve:

\frac{\partial}{\partial t} η_{t} = \frac{\partial}{\partial x} [x (1 - x) η_{t}] + λ η_{t} \cdot (x - \int_{0}^{1} y η_{t} (y) d y)

(4)

with initial density η₀(x). In this more transparent form one can see that the first term on the right is a flux term that transports density towards x = 0 whereas the second term is a forcing term that increases the density at values of x above the mean of the density. The flux corresponds to the individual-level moves: nearest neighbor moves in the particle system. The forcing term corresponds to group-level moves: moves to occupied sites in the particle system.

We will see that if we start with an initial measure μ₀ which is the sum of delta measures, then the solution μ_t retains the same form. More explicitly, if

μ_{0} (d x) = \sum_{i} a_{i} (0) δ_{x_{i} (0)} (d x)

where x_i(0) ∈ [0, 1], a_i(0) > 0, and $\sum a_{i} (0) = 1$ , then we will see (from Lemma 5) that the solution μ_t to (3) has the form

μ_{t} (d x) = \sum_{i} a_{i} (t) δ_{x_{i} (t)} (d x) .

Moreover, the parameters (a_i(t), x_i(t)) satisfy the following set of coupled equations

{\begin{cases} \frac{d x_{i}}{d t} = - x_{i} (1 - x_{i}) \\ \frac{d a_{i}}{d t} = λ a_{i} (x_{i} - 〈 y, μ_{t} 〉) = λ a_{i} (x_{i} - \sum_{j} a_{j} x_{j}) . \end{cases}

(5)

Notice that the positions of the delta masses change according to a negative logistic function, independently of the other masses and the density. The weight a_i increases at time t if the position of the particle x_i is above the mean, $\sum a_{j} x_{j}$ , and decreases if it is below the mean. To build intuition, it is instructive to consider some simple examples of this form.

Example 1

According to (5), if μ₀ = δ₁, then μ_t = μ₀. This can also be seen directly from (3). In the case of an initial condition containing some delta mass at 1, all of the rest of the mass will migrate towards zero. Eventually all of the mass will be below the mean as the mass at one will not move and will ever be increasing its mass as it is always above the mean. Once this happens it is clear that all of the mass will drain from all of the points not at one and hence μ_t → δ₁ as t → ∞. This reasoning holds in a more general setting and is included in Theorem 3.

Example 2

According to (5), if μ₀ = δ₀, then μ_t = μ₀. This too can be seen directly from (3). In the case of an initial condition containing no mass at one and only finite number of masses total, the mass will eventually all move towards zero and hence hence μ_t → δ₀ as t → ∞. If an infinite number of masses are allowed the situation is not as simple. Theorem 3 hints at the possible complications by giving an example of a density which is invariant.

These simple examples correspond to similarly straightforward biological scenarios: once the population is entirely composed of type I individuals or of type G individuals, it stays in that state. Though δ₀ is a fixed point of the system attracting many initial configurations, it is not Lyapunov stable. This means that even small perturbations of δ₀ can lead to an arbitrary large excursion away from δ₀ even though the system eventually returns to δ₀. Rather than making a precise statement which would require quantifying the size of a perturbation, consider the example of μ₀ = (1−ε)δ₀ +εδ₁_−α. As ε → 0, the distance between μ₀ and δ₀ goes to zero in any reasonable metric. If we write $μ_{t} = (1 - a_{t}) δ_{0} + a_{t} δ_{x_{t}}$ then as α → 0 one can ensure that the system spends arbitrarily long time with $x_{t} > \frac{1}{2}$ and hence a_t will grow to as close to one as one wants in this time. Thus the system could be described as making an an arbitrarily big excursion away from δ₀ even though μ_t → δ₀ as t → ∞.

It natural to ask if there are other fixed points beyond δ₀ and δ₁.

Lemma 2 (Fixed points)

The measures delta δ₀, δ₁, and densities in the Beta(λ − α, α) family of distributions:

\frac{1}{B (λ - α, α)} x^{λ - α - 1} {(1 - x)}^{α - 1}

with α ∈ (0, λ), are fixed points of (3). B(λ − α, α) is the normalizing constant that makes the density integrate to 1 over the interval [0, 1].

For measure-valued initial data, we show that the basins of attraction for the fixed points are determined by whether they charge the point x = 1 and their Hölder exponent around x = 1.

Theorem 3 (Steady state behavior)

Consider measure valued solution μ_t(dx) to (3) with initial probability measure μ₀(dx). If μ₀({1}) > 0 then

μ_{t} \to δ_{1} a s t \to \infty

and if μ₀([1 − ε, 1]) = 0 for some ε > 0 then

μ_{t} \to δ_{0} a s t \to \infty .

Alternatively, suppose that for some α > 0 and C > 0

x^{- α} μ_{0} ([1 - x, 1]) \to C a s x \to 0 .

If α < λ, then

μ_{t} (d x) \to Beta (λ - α, α) a s t \to \infty .

Otherwise, if α ≥ λ,

μ_{t} (d x) \to δ_{0} (d x) a s t \to \infty

The α < λ case in the above theorem is particularly relevant to theoretical evolutionary biology because it implies that coexistence of the two types is possible in this infinite population limit.

The results of Theorem 3 should be contrasted with the original Markov chain before taking the limit m, n → ∞. In the Markov chain, all individuals eventually become either entirely type G or type I. These two homogeneous states are absorbing states for the individual level dynamics. The population level state made of individuals that are all either homogeneous of type G or I is absorbing for the group level dynamics. Hence, the state of the system eventually becomes composed entirely of homogeneous groups of solely G or I and stays in that state for all future times. These two absorbing states of the Markov chain, with finite m and n, correspond to the states δ₀ and δ₁ in the scaling limit. Hence the natural discretization for the Beta distribution to the lattice ${\frac{k}{n} : 0 < k < n}$ , given by

\frac{1}{Z (m, n, λ, α)} {(\frac{k}{n})}^{λ - α - 1} {(1 - \frac{k}{n})}^{α - 1},

cannot be invariant. (Here Z is the normalization constant which ensures the probabilities sum to one.) However for large m and n, it is reasonable to expect it to be nearly invariant in the sense that if the initial states ${X_{i} (0) : 1 \leq i \leq m}$ are independent and distributed as the discrete Beta distribution then the Markov chain dynamics will keep the distribution close to the product of discretized Beta distributions for a long time. The expectation of this time will grow to infinity as m, n → ∞. In the context of evolutionary biology, this suggests that although a large finite population ultimately becomes fixed in one of two homogeneous states, it may be trapped for a long time in a state where both types coexist. Furthermore, this nearly invariant state should be similar to the discretization of the Beta distribution above.

We will not pursue a rigorous proof of this near or quasi invariance here. Nonetheless, we now briefly sketch the argument as we understand it, giving the central points. If the distribution of the Markov chain is close to a product of discretized Beta distributions, then the empirical mean will be highly concentrated around the mean of continuous Beta when m and n are large. Hence the generator projected on to any X_i is nearly decoupled from the other particles and close to being Markovian. More precisely, the dynamics of any fixed X_i is well approximated in this setting by the one-dimensional Markov chain obtained by replacing the mean of the empirical measure in the full generator with the mean of the Beta distribution. It is straightforward to see that for m and n large the discretized Beta distribution is an approximate left-eigenfunction of this one-dimensional generator with an eigenvalue which goes to zero as m, n → ∞.

All of these observations can be combined to show that if the systems starts in the product discretized Beta distribution then it will say close to the product discretized Beta distribution for a long time if m and n are large.

We now turn to the second scaling. Let $s = \frac{σ}{n}$ , $r = \frac{ρ}{m}$ , and $\frac{n}{m} \to θ$ , and let $v_{t}^{m, n}$ denote the empirical measure under this scaling. The terms F″ and g″ in the generator (1) no longer vanish and the process converges to a limit that is stochastic. Our weak convergence result is proved and stated in terms of a martingale problem.

Theorem 4

Suppose $\frac{n}{m} \to θ$ , w = O(1), $s = \frac{σ}{n}$ , $r = \frac{ρ}{m}$ , and we speed up time by a factor of n. Suppose the particles in the rescaled $v_{t}^{m, n}$ process are initially independently and identically distributed according to the measure $v_{0}^{m, n}$ where $v_{0}^{m, n} \to v_{0}$ as m, n → ∞. Then the rescaled process converges weakly to v_t as m, n → ∞, where ν_t satisfies the following martingale problem:

N_{t} (f) = 〈 f, v_{t} 〉 - 〈 f, v_{0} 〉 - \int_{0}^{t} 〈 A f, v_{z} 〉 d z - w θ ρ \int_{0}^{t} {\int_{0}^{1} \int_{0}^{1} f (x) V (z, v_{z}, y) Q (v_{z}; d x, d y)} d z

(6)

is a martingale with conditional quadratic variation

{〈 N (f) 〉}_{t} = 2 w θ \int_{0}^{t} [\int_{0}^{1} \int_{0}^{1} f (x) f (y) Q (v_{τ}; d x, d y)] d τ

(7)

where

A f (x) = x (1 - x) [\frac{d^{2}}{d x^{2}} f (x) - σ \frac{d}{d x} f (x)]

V (t, v, x) = x

Q (v; d x, d y) = v (d x) (δ_{x} (d y) - v (d y))

and f ∈ C²([0, 1]).

The drift part of the martingale (6) comprises a second order partial differential operator A and the centering term from the global jump dynamics (the expression in curly brackets). Note that A is in fact the generator of the Wright-Fisher diffusion with selection [8]. The entire process is a Fleming-Viot process [11]. Fleming-Viot processes frequently arise in models of population genetics (for example [3, 12]; see [10] for a review). In these contexts, the variable x can represent the geographical location of an individual, or as in the original paper of Fleming and Viot [11], the genotype of an individual (where genotype is a continuous instead of a discrete variable).

As an aside, it may be helpful to mention an alternative characterization of this Fleming-Viot process as an infinite system of ordinary stochastic differential equations. Instead of a martingale problem, where both m, n have been taken to ∞, we consider the n → ∞ limit first. In this case, we have a finite collection of delta masses (each of mass $\frac{1}{m}$ ) moving on the interval [0, 1]. The positions of these delta masses can be represented by a coupled system of stochastic differential equations (SDEs). From the generator equation (2), one can see that each SDE comprises a diffusion part (corresponding to the individual-level dynamics) and a jump process (corresponding to the group-level dynamics). Specifically, a delta mass jumps to the position of another delta mass according to a Poisson process with rates dependent on the positions of the delta masses. Donnelly & Kurtz [7] characterize a population process in terms of such a system of SDES and show that the infinite population limit corresponds to a martingale problem for the Fleming-Viot process.

We briefly discuss other scalings one might obtain from the particle system and their biological significance. The two scalings studied here correspond, respectively, to what is called ‘strong’ selection and ‘weak’ selection occurring at both levels. (In the field of theoretical evolutionary biology, strong selection is defined as the selection parameter being constant in the population size whereas weak selection has the selection parameter scaling with the inverse of the population size). One can also use the same techniques to characterize the limiting system when selection is strong at one level and weak at the other. The dynamics of these limiting systems are more straightforward. For example, if selection is weak at the individual level $(s = O (\frac{1}{n}))$ and strong at the group level (no rescaling of r), one can see from the generator equation (2) that the highest order term corresponds to selection at the group level. The limit is therefore deterministic and the steady-state is a population homogeneous in the group type with the largest proportion of type G individuals present in the initial state. Note that it is possible, by further rescaling w, to obtain a limit from a mixture of weak and strong selection. A biological interpretation of these observations is that for selection to manifest itself at two biological levels, the selective forces must be comparable in some sense: either both levels undergo the same type of selection (weak or strong), or if weak selection acts on one level but not the other, this weak selection must be compensated by a faster timescale.

The dynamical properties of the deterministic partial differential equation (3) are the focus of the next section. The proofs of weak convergence (Theorems 1 and 4) are deferred to section 4.

3. Properties of the deterministic limit

We begin with a closed-form expression for solutions to the deterministic partial differential equation (3).

Lemma 5

The solution to the deterministic partial differential equation (3) with initial measure μ₀ is given by

μ_{t} (d x) = (G_{t} μ_{0}) (d x) = (μ_{0} ϕ_{t}^{- 1}) (d x) \cdot w_{t} (x)

(8)

where

ϕ_{t}^{- 1} (x) = \frac{x}{e^{- t} + x (1 - e^{- t})}

w_{t} (x) = {[(e^{- t} + x (1 - e^{- t})) e^{t - \int_{0}^{t} h (z) d z}]}^{λ}

and h(t) satisfies h(t) = 〈x, μ_t〉

Remark 1

$(μ_{0} ϕ_{t}^{- 1}) (d x) : = μ_{0} (ϕ_{t}^{- 1} (d x))$ captures the changes in the initial data that are solely due to the flux term. This expression is also known as the push-forward measure of μ₀ under the dynamics of ϕ. As we will see in the proof, ϕ_t(x) is precisely the characteristic curve for the spatial variable x and includes a normalizing constant. The multiplication by w_t(x) captures the changes in the initial data that are due to the forcing term in (3) and includes a normalizing factor.

Remark 2

Density-valued solutions are given by

η_{t} (x) = η_{0} (ϕ_{t}^{- 1} (x)) \partial_{x} ϕ_{t}^{- 1} (x) \cdot w_{t} (x) = η_{0} (\frac{x}{e^{- t} + x (1 - e^{- t})}) {[e^{- t} + x (1 - e^{- t})]}^{λ - 2} e^{(λ - 1) t - λ \int_{0}^{t} h (z) d z}

(9)

To see this, suppose μ₀(dx) = η₀(x)dx. Then for any test function f,

\int_{0}^{1} f (x) (μ_{0} ϕ_{t}^{- 1}) (d x) = \int_{0}^{1} (f \circ ϕ_{t}) (x) μ_{0} (d x) = \int_{0}^{1} f (y) η_{0} (ϕ_{t}^{- 1} (y)) \partial_{y} ϕ_{t}^{- 1} (y) d y

The first equality follows from the change-of-variable property of push-forward measures and the second from a standard change of variables. The limits of integration do not change because 0 and 1 are fixed points of both ϕ_t and $ϕ_{t}^{- 1}$ .

Proof of Lemma 5

We apply the method of characteristics (see for example [17]) to obtain a formula for a density-valued solution. We then prove that the weak, measure-valued analog of this solution satisfies (3). Consider the following modification of (4):

\frac{\partial}{\partial t} ξ (t, x) = \frac{\partial}{\partial x} [x (1 - x) ξ (t, x)] + λ ξ (t, x) [x - h (t)]

(10)

where h(t) is a general function in time and ξ₀ ∈ C¹([0, 1]). Note that when $h (t) = \int_{0}^{1} y ξ (t, y) d y$ , this differential equation is equivalent to (4). To be clear about which equation we are solving, we use ξ(t, x) to denote solutions when h(t) is unspecified.

Rewriting (10):

(\frac{\partial}{\partial t} ξ, \frac{\partial}{\partial x} ξ, - 1) \cdot (1, - x (1 - x), [(1 - 2 x) + λ (x - h (t))] ξ) = 0

The second vector is therefore tangent to the solution surface and gives the rates of change for the t, x, and ξ coordinates. Let the initial condition be parameterized as (0, x, ξ₀(x)) = (0, p, ξ₀(p)). The t, x, and ξ coordinates change according to the characteristic equations

\begin{array}{l} \frac{d t}{d q} = 1 & t (0, p) = 0 \\ \frac{d x}{d q} = - x (1 - x) & x (0, p) = p \\ \frac{d ξ}{d q} = [(1 - 2 x (q, p)) + λ (x (q, p) - h (t (q, p)))] ξ & ξ (0, p) = ξ_{0} (p) \end{array}

where q is the parameter as we move through the solutions in time. The first two ordinary differential equations have solutions

\begin{array}{l} t (q, p) = q \\ x (q, p) = \frac{p}{p - (p - 1) e^{q}} = : ϕ_{q} (p) \end{array}

(11)

From this, the third differential equation can be solved exactly:

\frac{d ξ}{d q} = [1 + \frac{p}{p - (p - 1) e^{q}} (λ - 2) - λ h (q)] ξ

ξ (q, p) = ξ_{0} (p) \exp {q - λ \int_{0}^{q} h (z) d z + (λ - 2) \int_{0}^{q} \frac{p}{p - (p - 1) e^{z}} d z} = ξ_{0} (p) e^{q - λ \int_{0}^{q} h (z) d z} {[p (e^{- q} - 1) + 1]}^{(- λ + 2)}

Next, make the substitutions q = t and $p = ϕ_{t}^{- 1} (x)$ from (11) to obtain ξ in terms of t and x:

ξ (t, x) = (ξ_{0} \circ ϕ_{t}^{- 1}) (x) {[e^{- t} + x (1 - e^{- t})]}^{(λ - 2)} e^{(λ - 1) t - λ \int_{0}^{t} h (z) d z} = (ξ_{0} \circ ϕ_{t}^{- 1}) (x) \partial_{x} ϕ_{t}^{- 1} (x) \cdot w_{t} (x)

(12)

If h(t) satisfies $h (t) = \int_{0}^{1} y ξ (t, y) d y$ , then by definition, ξ(t, x) solves the partial differential equation (4). Conversely, if ξ(t, x) solves the partial differential equation (4), it also solves the differential equation (10) with $h (t) = \int_{0}^{1} y ξ (t, y) d y$ . Therefore this above expression, along with the condition $h (t) = \int_{0}^{1} y ξ (t, y) d y$ , are equivalent to solutions of (4).

To extend this result to measures, suppose we have a strong solution η_t(x) with initial condition η₀:

η_{t} (x) = (η_{0} \circ ϕ_{t}^{- 1}) (x) \partial_{x} ϕ_{t}^{- 1} (x) \cdot w_{t} (x)

Using a similar calculation as that in Remark 2, the measure μ_t(dx) corresponding to η_t(x) is given by

μ_{t} (d x) = (μ_{0} ϕ_{t}^{- 1}) (d x) \cdot w_{t} (x)

It remains to check that this satisfies the weak deterministic partial differential equation (3), with h(t) = 〈x, μ_t〉. The left hand side of the equation is

\frac{d}{d t} 〈 f, μ_{t} 〉 = \frac{d}{d t} \int_{0}^{1} f (x) w_{t} (x) (μ_{0} ϕ_{t}^{- 1}) (d x) = \frac{d}{d t} \int_{0}^{1} f (ϕ_{t} (x)) w_{t} (ϕ_{t} (x)) μ_{0} (d x)

Differentiating under the integral sign, expanding out the expressions for ∂_tϕ_t and ∂_t(w_t(ϕ_t(x)), and applying change of variables for push-forward measures again, we obtain

\frac{d}{d t} 〈 f, μ_{t} 〉 = - \int_{0}^{1} x (1 - x) f^{'} (x) w_{t} (x) (μ_{0} ϕ_{t}^{- 1}) (d x) + λ \int_{0}^{1} [x - h (t)] f (x) w_{t} (x) (μ_{0} ϕ_{t}^{- 1}) (d x)

This matches right hand side of the weak deterministic partial differential equation (3). □

In practice, the condition h(t) = 〈x, μ_t〉 is difficult to use. The following provides an equivalent and simpler condition.

Lemma 6

(Conservation of measure condition). Suppose ξ is a weak measure-valued solution to the deterministic partial differential equation (10) with initial condition $\int_{0}^{1} ξ_{0} (d x) = 1$ . Then

h (t) = \int_{0}^{1} y ξ (t, d y) if and only if \int_{0}^{1} ξ (t, d y) = 1 \forall t > 0

Proof

(⇒ direction) Suppose $h (t) = \int_{0}^{1} y ξ (t, d y)$ . Then ξ is a weak measure-valued solution to (3). Taking the test function f ≡ 1, we obtain

\frac{d}{d t} 〈 1, ξ 〉 = 0 + λ [〈 x, ξ 〉 - 〈 1, ξ 〉 〈 x, ξ 〉] = 0

Thus, if the initial data has total measure 1, 〈1, ξ〉 remains constant at 1 for all t ≥ 0. (⇐ direction) Suppose $\int_{0}^{1} ξ (t, d x) = 1 for all t > 0$ . Again take the test function f ≡ 1 but this time with unspecified h(t):

0 = \frac{d}{d t} 〈 1, ξ 〉 = 0 + λ [〈 x, ξ 〉 - 〈 1, ξ 〉 h (t)] = λ [〈 x, ξ 〉 - h (t)] .

For this to hold, we must have $h (t) = \int_{0}^{1} x ξ (t, d x)$ . □

The above lemmas imply that solutions μ_t(dx) to (3) can be obtained by using formula (8) from Lemma 5 and imposing the conservation of measure condition 〈1, μ_t〉 ≡ 1 from Lemma 6. We illustrate this with some examples of exactly solvable solutions for special choices of initial data. We will see that the long time behavior of the examples is consistent with results stated in Theorem 3.

Example 3

Initial measure concentrated at x₀ ∈ [0, 1], i.e. $μ_{0} = {δ_{x}}_{0}$ Using formula (8),

\int f (x) μ_{t} (d x) = \int f (x) w_{t} (x) δ_{x_{0}} (ϕ_{t}^{- 1} (d x)) = f (ϕ_{t} (x_{0})) w_{t} (ϕ_{t} (x_{0})) = \int f (x) w_{t} (x) δ_{ϕ_{t} (x_{0})} (d x)

Thus $μ_{t} (d x) = w_{t} (x) δ_{ϕ_{t} (x_{0})} (d x)$ . Imposing the conservation of measure condition gives $μ_{t} (d x) = δ_{ϕ_{t} (x_{0})} (d x)$ . In other words, an initial delta measure at x₀ moves as a delta measure along the x axis with position given by ϕ_t(x₀), the solution to the negative logistic equation with initial position x₀.

Example 4

Initial uniform density: η₀(x) = 1, i.e μ₀(dx) = dx Using formula (9),

η_{t} (x) = e^{(λ - 1) t - λ \int_{0}^{t} h (z) d z} {[e^{- t} + x (1 - e^{- t})]}^{(λ - 2)}

Imposing conservation of measure:

e^{(λ - 1) t - λ \int_{0}^{t} h (z) d z} = {[\int_{0}^{1} {[e^{- t} + x (1 - e^{- t})]}^{(λ - 2)} d x]}^{- 1} = {\begin{cases} \frac{(λ - 1) (1 - e^{- t})}{1 - e^{- (λ - 1) t}} & i f λ \neq 1 \\ \frac{1 - e^{- t}}{t} & i f λ = 1 \end{cases}

Thus,

η_{t} (x) = {\begin{cases} \frac{(λ - 1) (1 - e^{- t})}{1 - e^{- (λ - 1) t}} {[e^{- t} + x (1 - e^{- t})]}^{(λ - 2)} & i f λ \neq 1 \\ \frac{1 - e^{- t}}{t} {[e^{- t} + x (1 - e^{- t})]}^{(λ - 2)} & i f λ = 1 \end{cases}

Note that η₀ ≡ 1 corresponds to an initial condition satisfying the hypothesis of Theorem 3 with α = 1. As predicted when λ > 1, we obtain η(t, x) → (λ − 1)x^λ⁻² = Beta(λ – 1, 1) as t → ∞.

The following is an example with α > 1.

Example 5

If η₀(x) = 2(1 − x), i.e. μ₀ ([1 – x, 1]) = x², then the corresponding α from Theorem 3 is α = 2.

Using formula (9)

η_{t} (x) = 2 e^{(λ - 2) t - λ \int_{0}^{t} h (z) d z} (1 - x) {[e^{- t} + x (1 - e^{- t})]}^{(λ - 3)}

Imposing the condition in Lemma 6 to solve for the h(z) term

e^{(λ - 2) t - λ \int_{0}^{t} h (z) d z} = {[2 \int_{0}^{1} (1 - x) {[e^{- t} + x (1 - e^{- t})]}^{(λ - 3)} d x]}^{- 1} = {\begin{cases} \frac{(λ - 2) (1 - e^{- t})}{2} {[\frac{1}{(λ - 1) (1 - e^{- t})} - \frac{e^{- (λ - 1) t}}{(λ - 1) (1 - e^{- t})} - e^{- (λ - 2) t}]}^{- 1} & i f λ \neq 2 \\ \frac{{(1 - e^{- t})}^{2}}{2 t e^{- t}} & i f λ = 2 \end{cases}

As predicted by Theorem 3 for λ > 2 = α,

η_{t} (x) \to \frac{1}{2} (λ - 2) (λ - 1) (1 - x) x^{λ - 3} = Beta (λ - 2, 2)

as t → ∞.

Example 6

$η_{0} (x) = \frac{1}{c} \cdot 1_{[0, c]} (x)$ with c < 1.

Using formula (9)

η_{t} (x) = \frac{1}{c} 1 {x \leq ϕ_{t} (c)} w_{t} (x) \partial_{x} ϕ_{t}^{- 1} (x)

Since $ϕ_{t} (c) = \frac{c e^{- t}}{1 - c + c e^{- t}} \to 0$ as t → ∞, η_t(x) → 0 for any x > 0. Since η must have total mass 1, it follows that regardless of the value of λ, η_t(x)dx → δ₀(dx) for any c < 1. This can also be seen by applying Theorem 3 and noting that $μ_{0} ([1 - c, 1]) = \int_{1 - c}^{1} η_{0} (x) d x = 0 .$

We end these examples with solutions for μ₀ that are mixtures of delta measures and densities. First, note that it is straightforward to extend Example 3 to the case where $μ_{0} (d x) = \sum a_{i} δ_{x_{i}} (d x)$ is a linear combination of delta measures, a_i > 0 for all i. Applying (8), we obtain

μ (t, d x) = \sum_{i} a_{i} w_{t} (x) δ_{ϕ_{t} (x_{i})} (d x) = \sum_{i} a_{i} (t) δ_{x_{i} (t)} (d x)

where x_i(t) = ϕ_t(x_i) and $a_{i} (t) = a_{i} w_{t} (x) |_{x = x_{i} (t)}$ . Our earlier system of equations (5) is obtained from this and the definitions of ϕ_t(x) and w_t(x).

Second, we consider a combination of a delta measure and a density

μ_{0} (d x) = a {δ_{x}}_{0} (d x) + (1 - a) v_{0} (x) d x

Notice that the formula for the solution (8) at first seems linear in the initial condition:

\int f (x) μ_{t} (d x) = \int f (x) (G_{t} μ_{0}) (d x) = \int f (x) w_{t} (x) [a δ_{ϕ_{t} (x_{0})} (d x) + (1 - a) v_{0} (ϕ_{t}^{- 1} (x)) \partial_{x} ϕ_{t}^{- 1} (x) d x] = \int f (x) [a (G_{t} δ_{x_{0}}) (d x) + (1 - a) (G_{t} v_{0}) (d x)]

This gives $(G_{t} μ_{0}) (d x) = a (G_{t} {δ_{x}}_{0}) (d x) + (1 - a) (G_{t} v_{0}) (d x)$ . However, this notation is misleading because implicit in the G_t operator is the function h(t), the mean of the overall process over time. Here, h(t) involves both the delta measure and the density. The solution operator G_t is therefore not linear for this reason.

Nevertheless, we can still use this formula to obtain expressions for solutions. We illustrate this with a concrete example.

Example 7

Take x₀ = 0 and v₀(x) the density function for Beta(λ − α, α) with α ∈ (0, λ). Using the solution formula and direct calculation, we obtain

μ_{t} (d x) = a w_{t} (0) δ_{0} (d x) + (1 - a) w_{t} (x) (v_{0} \circ ϕ_{t}^{- 1}) (x) \partial_{x} ϕ_{t}^{- 1} (x) d x = e^{- λ \int_{0}^{t} h (z) d z)} {a δ_{0} (d x) + (1 - a) e^{(λ - α) t} v_{0} (x) d x}

Note in particular that μ_t remains a linear combination of δ₀ and the Beta distribution. The Beta distribution ultimately dominates because λ > α.

We now use Lemma 5 to show that Beta distributions, δ₀, and δ₁ are fixed points for the deterministic partial differential equation and thus provide a proof of Lemma 2 announced earlier in this note.

Proof of Lemma 2

Note that we could prove this lemma by substituting δ₀, δ₁, and the Beta distribution into the deterministic partial differential equation (3) and showing the right-hand side equals zero. Instead, we will show that these distribution are fixed points of the solution operator. Let v be the density of the Beta distribution,

v (x) = \frac{1}{B (λ - α, α)} x^{λ - α - 1} {(1 - x)}^{α - 1} .

The mean of v is $\frac{λ - α}{λ}$ . Using (9)

(G_{t} v) (x) = v (\frac{x}{e^{- t} + x (1 - e^{- t})}) {[e^{- t} + x (1 - e^{- t})]}^{λ - 2} e^{(λ - 1) t - (λ - α) t} = ν (x)

v is therefore a fixed point of the solution operator and hence is a fixed point of the deterministic partial differential equation.

For δ₀ and δ₁, we use Example 3 above to obtain $(G_{t} {δ_{x}}_{0}) (d x) = δ_{ϕ_{t} (x_{0})} (d x)$ . Since x₀ = 0 and x₀ = 1 are fixed points of ϕ_t, it follows that δ₀ and δ₁ are fixed points of G_t. □

We now prove when the fixed points are stable. We begin with a lemma which gives more general conditions than those given in Theorem 3 for the delta measure at zero to attract a given initial condition.

Lemma 7

If for some α ≥ λ > 0,

\lim_{x \to 0} x^{- α} μ_{0} ([1 - x, 1]) < \infty

then μ_t → δ₀ as t → ∞. In particular, this condition holds if μ₀([1 − ε, 1]) = 0 for some ε > 0.

To prove this and subsequent results, we will need the following technical lemma.

Lemma 8

Setting h(t) = 〈x, μ_t〉, the following two implications hold:

\int_{0}^{\infty} h (t) d t < \infty \Rightarrow h (t) \to 0 a s t \to \infty .

\int_{0}^{\infty} [1 - h (t)] d t < \infty \Rightarrow h (t) \to 1 a s t \to \infty .

Proof of Lemma 8

Since h(t) ≥ 0 and 1 − h(t) ≥ 0, the only obstruction to the implication is that h(t) (or 1 − h(t)) could have ever shorter and shorter intervals were they return to an order one value before returning to a value close to zero. This would require h(t) to have unbounded derivatives. However this is not possible since

\frac{d h}{d t} (t) = - (h - 〈 x^{2}, μ_{t} 〉) + λ (〈 x^{2}, μ_{t} 〉 - h^{2})

from which one easily see that $- 1 \leq \frac{d h}{d t} (t) \leq λ$ since 0 ≤ h − 〈x², μ_t〉 ≤ 1 and 0 ≤ 〈x², μ_t〉 − h ≤ 1. □

Proof of Lemma 7

As usual let h(t) = 〈x, μ_t〉. We begin by observing that if

\int_{0}^{\infty} h (t) d t < \infty

then h(t) → 0 as t → ∞ by Lemma 8 and μ_t → δ₀ as we wish to prove. Thus, we henceforth assume that $\int_{0}^{\infty} h (t) d t = \infty$ . Under this assumption, we will show that for any continuous function f

\int_{0}^{1} f (x) μ_{t} (d x) \to f (0) as t \to \infty .

Since f is continuous, given any ε > 0, there exists a δ > 0 so that |f(x) − f(0)| < ε whenever x ≤ δ. Hence

| \int_{0}^{1} f (x) μ_{t} (d x) - f (0) | \leq \int_{0}^{1} | f (x) - f (0) | μ_{t} (d x) \leq ε + \int_{δ}^{1} | f (x) - f (0) | μ_{t} (d x)

(13)

Now setting

\int_{δ}^{1} | f (x) - f (0) | μ_{t} (d x) = \int_{ϕ_{t}^{- 1} (δ)}^{1} | (f \circ ϕ_{t}) (x) - f (0) | (w_{t} \circ ϕ_{t}) (x) μ_{0} (d x) \leq 2 {‖ f ‖}_{\infty} \int_{ϕ_{t}^{- 1} (δ)}^{1} (w_{t} \circ ϕ_{t}) (y) μ_{0} (d y) .

Since for all $y \in [ϕ_{t}^{- 1} (δ), 1]$ and t > 0, we have

(w_{t} \circ ϕ_{t}) (y) \leq e^{λ t - λ \int_{0}^{t} h (s) d s}

we see that

\int_{δ}^{1} | f (x) - f (0) | μ_{t} (d x) \leq 2 {‖ f ‖}_{\infty} e^{λ t - λ \int_{0}^{t} h (s) d s} μ_{0} ([ϕ_{t}^{- 1} (δ), 1]) .

Now using the assumptions on μ₀ and that $ϕ_{t}^{- 1} (δ) \geq 1 - D e^{- 1}$ for some D > 0 and all t >0 one has that

e^{λ t - λ \int_{0}^{t} h (s) d s} μ_{0} ([ϕ_{t}^{- 1} (δ), 1]) \leq \hat{D} e^{- (α - λ) t - λ \int_{0}^{t} h (s) d s}

for some constant $\hat{D}$ and all t > 0. Since α ≥ λ and $\int_{0}^{\infty} h (s) d s = \infty$ this bound converges to zero as t → ∞ and the proof is complete as the ε in (13) was arbitrary. □

Proof of Theorem 3

We start with the setting when μ₀({1}) > 0 and begin by writing μ_t(dx) = a_tδ₁(dx) + (1 − a_t)ν_t(dx) for some time dependent process a_t ∈ [0, 1] with a₀ > 0 and some probability measure valued process ν_t(dx). As usual we define h(t) = 〈x, μ_t〉 and using the representation given in (8), one sees that a_t solves

\frac{d a_{t}}{d t} = λ a_{t} (1 - h (t)) \Rightarrow a_{t} = a_{0} \exp (λ \int_{0}^{t} [1 - h (s)] d s) .

Since 1 − h(t) ≥ 0, we know that $\int_{0}^{t} [1 - h (s)] d s$ converges as t → ∞. If it converges to ∞ then a_t also converges to ∞ since a₀ > 0. However this is impossible since a_t ∈ [0, 1] for all t ≥ 0. Thus, we conclude that $\int_{0}^{t} [1 - h (s)] d s < \infty$ . Then Lemma 8 implies that h(t) → 1 which in turn implies that μ_t → δ₁ as t → ∞.

We know turn to the setting when x⁻^αμ₀([1 − x, 1]) → C > 0 as x → 0. The case when λ ≤ α is already handled by Lemma 7 leaving only the case when λ > α > 0 to be proven. For x ∈ [0, 1], define U(x) = μ₀([0, x]). Since μ₀ is a probability measure we know that U has finite variation and is regular in the sense that both the right limit U(x⁺) and the left limit U(x⁻) exist, where U(x^±) = lim U(y) as y →^± x. At the extreme points, only the limit obtained by staying in [0, 1] is defined.

Now for any smooth function f of [0, 1], we have from (8) that

\int_{0}^{1} f (x) μ_{t} (d x) = Z_{t} \int_{0}^{1} f (x) g_{t} (x) (μ_{0} ϕ_{t}^{- 1}) (d x) = Z_{t} \int_{0}^{1} [(f g_{t}) \circ ϕ_{t}] (x) μ_{0} (d x)

where w_t(x) has been written as the product of g_t(x) = (e⁻^t + x(1 − e⁻^t))^λ and Z_t some positive, time dependent normalizing constant. It is enough to show that for some time positive, dependent constant K_t,

K_{t} \int_{0}^{1} [(f g_{t}) \circ ϕ_{t}] (x) μ_{0} (d x) \to \int_{0}^{1} f (x) x^{λ - α - 1} {(1 - x)}^{α - 1} d x as t \to \infty .

(14)

Since x ↦ f(x)g_t(x) is continuous on [0, 1], even if U(x) has discontinuities the integration by parts formula for Lebesgue-Stieltjes integrals produces

\int_{0}^{1} [(f g_{t}) \circ ϕ_{t}] (x) μ_{0} (d x) = (f g_{t} U) (1^{-}) - (f g_{t} U) (0^{+}) - \int_{0}^{1} \partial_{x} [(f g_{t}) \circ ϕ_{t}] (x) U (x) d x = [f g_{t} (U - 1)] (1^{-}) + [f g_{t} (1 - U)] (0^{+}) + \int_{0}^{1} \partial_{x} [(f g_{t}) \circ ϕ_{t}] (x) [1 - U] (x) d x .

Here we have used that ϕ_t is continuous with ϕ_t(1) = 1 and ϕ_t(0) = 0.

First observe that 1 − U(1⁻) = 0 since μ₀([1 – x, 1]) → 0 as x → 0 by assumption and that g_t(0⁺) = e⁻^λt. Hence

[f g_{t} (1 - U)] (1^{-}) + [f g_{t} (U - 1)] (0^{+}) = [U (0^{+}) - 1] f (0) e^{- λ t} .

(15)

Now turning to the integral term, applying the chain rule and changing variables to y = ϕ_t(x) produces

\int_{0}^{1} \partial_{x} [(f g_{t}) \circ ϕ_{t}] (x) [1 - U] (x) d x = \int_{0}^{1} [\partial_{x} (f g_{t}) \circ ϕ_{t}] (x) [1 - U] (x) (\partial_{x} ϕ_{t}) (x) d x = \int_{0}^{1} [\partial_{x} (f g_{t})] (y) [(1 - U) \circ ϕ_{t}^{- 1}] (y) d y

For any fixed x ∈ (0, 1) by direct calculation and use of the assumption on μ₀, one sees that

\begin{array}{r} \partial_{x} (f g_{t}) (x) \to \partial_{x} (x^{λ} f) (x) \\ e^{α t} (1 - U (ϕ_{t}^{- 1} (x))) = e^{α t} μ_{0} ([ϕ_{t}^{- 1} (x), 1]) \to C {(\frac{1 - x}{x})}^{α} \end{array}} as t \to \infty .

Combining these facts with (15) and the fact that e⁻⁽^λ⁻^α⁾^t → 0 as t → ∞ since λ > α produces

e^{α t} \int_{0}^{1} [(f g_{t}) \circ ϕ_{t}] (x) μ_{0} (d x) \to C \int_{0}^{1} \partial_{x} (x^{λ} f) (x) {(\frac{1 - x}{x})}^{α} d x as t \to \infty .

for some new positive constant C. Now since integration by parts implies that

\frac{1}{α} \int_{0}^{1} \partial_{x} (x^{λ} f) (x) {(\frac{1 - x}{x})}^{α} d x = \int_{0}^{1} f (x) x^{λ - α - 1} {(1 - x)}^{α - 1} d x

the last part of the proof is complete. □

4. Proofs of weak convergence

The proofs of Theorems 1 and 4 follow a standard procedure [13, 12, 3]. Both proofs require: (i) tightness of the sequence of stochastic processes – which implies a subsequential limit, and (ii) uniqueness of this limit. For the tightness of ${μ_{t}^{m, n}}_{m, n}$ on D([0, T], $P ([0, 1])$ ), it is sufficient, by Theorem 14.26 in Kallenberg [14] to show that ${〈 f, μ_{t}^{m, n} 〉}$ tight on D([0, T], ℝ) for any test function f from a countably dense subset of continuous, positive functions on [0, 1]. For the uniqueness of solutions to the partial differential equation in Theorem 1, we apply Gronwall’s inequality. For uniqueness of solutions to the martingale problem in Theorem 4, we apply a Girsanov theorem by Dawson [5].

4.1. Semimartingale property of multilevel selection process

It will be useful for what follows to treat $〈 f, μ_{t}^{m, n} 〉$ as a semimartingale. Below, $D_{x}^{+} f$ is the first order difference quotient of f taken from the right, $D_{x}^{-} f$ is the first order difference quotient of f taken from the left, and D_xxf is the second order difference quotient.

Lemma 9

For f ∈ C²([0, 1]) and $μ_{t}^{m, n}$ with generator L^m,n defined in (1),

〈 f, μ_{t}^{m, n} 〉 - 〈 f, μ_{0}^{m, n} 〉 = A_{t}^{m, n} (f) + M_{t}^{m, n} (f)

(16)

where $A_{t}^{m, n} (f)$ is a process of finite variation, $A_{t}^{m, n} (f) : = \int_{0}^{t} a_{z}^{m, n} (f) d z$ , with

a_{t}^{m, n} (f) = \sum_{i} μ_{t}^{m, n} (\frac{i}{n}) \frac{i}{n} (1 - \frac{i}{n}) [\frac{1}{n} D_{x x} f (\frac{i}{n}) - s D_{x}^{-} f (\frac{i}{n})] + w r {\sum_{j} μ_{t}^{m, n} (\frac{j}{n}) \frac{j}{n} f (\frac{j}{n}) - \sum_{i} μ_{t}^{m, n} (\frac{i}{n}) f (\frac{i}{n}) \sum_{j} μ_{t}^{m, n} (\frac{j}{n}) \frac{j}{n}}

(17)

and $M_{t}^{m, n} (f)$ is a càdlàg martingale with (conditional) quadratic variation

{〈 M^{m, n} (f) 〉}_{t} = \frac{1}{m} \int_{0}^{t} {\frac{1}{n} \sum_{i} μ_{z}^{m, n} (\frac{i}{n}) \frac{i}{n} (1 - \frac{i}{n}) [{(D_{x}^{+} f (\frac{i}{n}))}^{2} + (1 + s) {(D_{x}^{-} f (\frac{i}{n}))}^{2}] + w \sum_{i, j} μ_{z}^{m, n} (\frac{i}{n}) μ_{z}^{m, n} (\frac{j}{n}) (1 + r \frac{j}{n}) {(f (\frac{i}{n}) - f (\frac{j}{n}))}^{2}} d z

(18)

Proof

By Dynkin’s formula (see, for example, Lemma 17.21 in [14]),

ψ (μ_{t}^{m, n}) - ψ (μ_{0}^{m, n}) - \int_{0}^{t} (L^{m, n} ψ) (μ_{s}^{m, n}) d s

where ψ ∈ dom(L^m,n), is a càdlàg martingale. In particular, this is true for

ψ (μ_{t}^{m, n}) = F (〈 f, μ_{t}^{m, n} 〉)

where f ∈ C²([0, 1]) and F: ℝ → ℝ. Setting F (x) = x and plugging this f into (1):

(L^{m, n} 〈 f, \cdot 〉) (v) = \sum_{i} v (\frac{i}{n}) \frac{i}{n} (1 - \frac{i}{n}) [\frac{1}{n} D_{x x} f (\frac{i}{n}) - s D_{x}^{-} f (\frac{i}{n})] + w r {\sum_{j} v (\frac{j}{n}) \frac{j}{n} f (\frac{j}{n}) - \sum_{i} v (\frac{i}{n}) f (\frac{i}{n}) \sum_{j} v (\frac{j}{n}) \frac{j}{n}}

Thus,

〈 f, μ_{t}^{m, n} 〉 - 〈 f, μ_{0}^{m, n} 〉 - \int_{0}^{t} a_{z}^{m, n} (f) d z = M_{t}^{m, n} (f)

(19)

where $M_{t}^{m, n} (f)$ is some martingale and $a_{t}^{m, n} (f) = (L^{m, n} 〈 f, \cdot 〉) (μ_{t}^{m, n})$ . A_t(f) is a process of finite variation because for a given f, $a_{t}^{m, n} (f)$ is uniformly bounded in t.

Next, setting F (x) = x² and plugging this ψ into (1):

(L^{m, n} {〈 f, \cdot 〉}^{2}) (v) = 2 〈 f, v 〉 a_{t}^{m, n} (f) + \frac{1}{m n} \sum_{i} v (\frac{i}{n}) \frac{i}{n} (1 - \frac{i}{n}) [{(D_{x}^{+} f (\frac{i}{n}))}^{2} + (1 + s) {(D_{x}^{-} f (\frac{i}{n}))}^{2}] + \frac{w}{m} \sum_{i, j} v (\frac{i}{n}) v (\frac{j}{n}) (1 + r \frac{j}{n}) {(f (\frac{i}{n}) - f (\frac{j}{n}))}^{2}

Thus,

{〈 f, μ_{t}^{m, n} 〉}^{2} - {〈 f, μ_{0}^{m, n} 〉}^{2} - \int_{0}^{t} c_{z}^{m, n} (f) d z = martingale

(20)

where $c_{t}^{m, n} (f) = (L^{m, n} {〈 f, \cdot 〉}^{2}) (μ_{t}^{m, n})$ .

Alternatively, take $Y_{t} = 〈 f, μ_{t}^{m, n} 〉$ and apply Ito’s formula (for example, p78 in [18]) to $Y_{t}^{2}$ to obtain

{〈 f, μ_{t}^{m, n} 〉}^{2} - {〈 f, μ_{0}^{m, n} 〉}^{2} = 2 \int_{0}^{t} 〈 f, μ_{z} 〉 a_{z}^{m, n} (f) d z + {[M^{m, n} (f)]}_{t} + martingale

(21)

where [M^m,n(f)]_t is the quadratic variation process of $M_{t}^{m, n}$ . Since 〈M^m,n(f)〉_t is the compensator of [M^m,n(f)]_t,

{[M^{m, n} (f)]}_{t} - {〈 M^{m, n} (f) 〉}_{t}

is a martingale. Thus,

{〈 f, μ_{t}^{m, n} 〉}^{2} - {〈 f, μ_{0}^{m, n} 〉}^{2} - 2 \int_{0}^{t} 〈 f, μ_{z}^{m, n} 〉 a_{z}^{m, n} (f) d z - {〈 M^{m, n} (f) 〉}_{t} = martingale

(22)

The compensator 〈M^m,n(f)〉_t is a predictable process of finite variation (see p118 in[18]). By the Doob-Meyer inequality (p103 in [18]), the martingale in (22) is the same as the martingale in (20). Equating these martingale parts we obtain

2 \int_{0}^{t} 〈 f, μ_{z}^{m, n} 〉 a_{z}^{m, n} (f) d z + {〈 M^{m, n} (f) 〉}_{t} = \int_{0}^{t} c_{z}^{m, n} (f) d z .

(23)

Substituting in the expressions for $a_{z}^{m, n}$ and $c_{z}^{m, n}$ then gives the explicit expression for the conditional quadratic variation (18) in the statement of the lemma. □

4.2. Proof of deterministic limit

To prove Theorem 1, we need the two following lemmas. The first uses criteria in Billingsley [2] to show tightness of the sequence of processes $〈 f, μ_{t}^{m, n} 〉$ . The second uses Gronwall’s inequality to show uniqueness of solutions to the limiting system.

Lemma 10

The processes $〈 f, μ_{t}^{m, n} 〉$ , as a sequence in {(m, n)}, is tight for all positive-valued test functions f ∈ C¹([0, 1]).

Proof

By Theorem 13.2 in [2], a sequence of probability measures {P_n} on D([0, T], ℝ⁺) is tight if and only if (i) for all η > 0, there exists a such that

P_{n} (x : \sup_{t \in [0, T]} | x (t) | \geq a) \leq η for n \geq 1

and (ii) for all ε > 0 and η > 0, there exists δ ∈ (0, 1) and n₀ such that

P_{n} (x : w_{x}^{'} (δ) \geq ε) \leq η for all n > n_{0}

where w′ is the modulus of continuity for càdlàg processes and is defined

w_{x}^{'} (δ) : = \inf_{{t_{i}}} max_{1 \leq i \leq v} \sup_{s, t \in [t_{i - 1}, t_{i})} | x (s) - x (t) |

where {t_i} is a partition of [0, T] such that $max_{i} {t_{i} - t_{i - 1}} \leq δ$ and x ∈ D([0, T], ℝ⁺) is distributed according to P_n.

First, note that since $μ_{t}^{m, n}$ is a probability measure, we have

| 〈 f, μ_{t}^{m, n} 〉 | \leq {‖ f ‖}_{\infty}

for all t, m, and n. Thus, (i) holds.

For (ii), we have by Markov’s inequality:

P_{m, n} (w^{'} (δ) \geq ε) \leq \frac{1}{ε} E_{m, n} (w^{'} (δ))

(24)

where $w^{'} (δ) : = w_{〈 f, μ_{t}^{m, n} 〉}^{'} (δ)$ . We will use the fact that $〈 f, μ_{t}^{m, n} 〉$ is a pure jump process to bound the right-hand side. The process $〈 f, μ_{t}^{m, n} 〉$ has two types of jumps: nearest-neighbor, and occupied-site jumps. Nearest-neighbor jumps occur at rate

\sum_{i} m μ_{t}^{m, n} (\frac{i}{n}) i (1 - \frac{i}{n}) (2 + s) \leq \frac{m n}{4} (2 + s)

and have magnitude

| 〈 f, μ_{t}^{m, n} 〉 - 〈 f, μ_{t^{-}}^{m, n} 〉 | = | 〈 f, μ_{t^{-}}^{m, n} + \frac{1}{m} (δ_{\frac{i \pm 1}{n}} - δ_{\frac{i}{n}}) 〉 - 〈 f, μ_{t^{-}}^{m, n} 〉 | \leq \frac{1}{m n} max_{i} | D_{x}^{-} f (\frac{i}{n}) |

Occupied-site jumps occur at rate

\sum_{i, j} m μ_{t}^{m, n} (\frac{i}{n}) μ_{t}^{m, n} (\frac{j}{n}) (1 + r \frac{j}{n}) \leq m (1 + r)

and have magnitude

| 〈 f, μ_{t}^{m, n} 〉 - 〈 f, μ_{t^{-}}^{m, n} 〉 | = | 〈 f, μ_{t^{-}}^{m, n} + \frac{1}{m} (δ_{\frac{j}{n}} - δ_{\frac{i}{n}}) 〉 - 〈 f, μ_{t^{-}}^{m, n} 〉 | \leq \frac{2}{m} {‖ f ‖}_{\infty}

Putting this together,

E_{m, n} (w^{'} (δ)) \leq E_{m, n} [number of nearest-neighbor jumps in time δ] \cdot \frac{1}{m n} max_{i} | D_{x}^{-} f (\frac{i}{n}) | + E_{m, n} [number of occupied-site jumps in time δ] \cdot 2 \frac{1}{m} {‖ f ‖}_{\infty}

\leq \frac{m n}{4} (2 + s) δ \frac{1}{m n} max_{i} | D_{x}^{-} f (\frac{i}{n}) | + m (1 + r) δ \frac{2}{m} {‖ f ‖}_{\infty} = {\frac{2 + s}{4} max_{i} | D_{x}^{-} f (\frac{i}{n}) | + 2 (1 + r) {‖ f ‖}_{\infty}} δ

Because f ∈ C¹([0, 1]), the expression in curly brackets is uniformly bounded by C_f, a constant that depends on f but not on m nor n. Substituting the above into (24) we get that for $δ < \frac{ε η}{C_{f}}$ ,

P_{m, n} (w^{'} (δ) \geq ε) \leq η

for all m and n. Thus, both conditions for tightness are satisfied and $〈 f, μ_{t}^{m, n} 〉$ is tight. □

Lemma 11

The integro-partial differential equation (3) in Theorem 1 has a unique solution.

Proof

Suppose μ_t satisfies (3). Fix t ≥ 0 and let ψ_t(x) be a function of time t and space x. By the chain rule and the differential equation (3),

{\frac{d}{d t} 〈 ψ_{t}, μ_{t} 〉 = \frac{d}{d z} 〈 ψ_{z}, μ_{t} 〉 |}_{z = t} + {\frac{d}{d z} 〈 ψ_{t}, μ_{z} 〉 |}_{z = t} = 〈 \frac{\partial}{\partial t} ψ_{t}, μ_{t} 〉 - 〈 s x (1 - x) \frac{\partial ψ_{t}}{\partial x}, μ_{t} 〉 + w r [〈 x ψ_{t}, μ_{t} 〉 - 〈 ψ_{t}, μ_{t} 〉 〈 x, μ_{t} 〉]

〈 ψ_{t}, μ_{t} 〉 = 〈 ψ_{0}, μ_{0} 〉 + \int_{0}^{t} 〈 \frac{\partial}{\partial z} ψ_{z} (x) + G ψ_{z} (x), μ_{z} 〉 d z + w r \int_{0}^{t} 〈 x ψ_{z}, μ_{z} 〉 - 〈 ψ_{z}, μ_{z} 〉 〈 x, μ_{z} 〉 d z

(25)

where $G f = - s x (1 - x) \frac{\partial}{\partial x} f$ . Let P_t be the semigroup operator associated with G. In fact, using the method of characteristics (or Lemma 5 with λ = 0),

P_{t} f = f (\frac{x e^{- s t}}{1 - x + x e^{- s t}})

(26)

Now, set ψ_z(x) = P_t−zf(x) for 0 ≤ z ≤ t, where f ∈ C¹([0, 1]) is some test function. Substituting this into (25), we have

〈 P_{0} f, μ_{t} 〉 = 〈 P_{t} f, μ_{0} 〉 + \int_{0}^{t} 〈 \frac{\partial}{\partial z} P_{t - z} f (x) + G P_{t - z} f (x), μ_{z} 〉 d z + \int_{0}^{t} w r [〈 x P_{t - z} f, μ_{z} 〉 - 〈 P_{t - z} f, μ_{z} 〉 〈 x, μ_{z} 〉] d z

〈 f, μ_{t} 〉 = 〈 P_{t} f, μ_{0} 〉 + \int_{0}^{t} w r [〈 x P_{t - z} f, μ_{z} 〉 - 〈 P_{t - z} f, μ_{z} 〉 〈 x, μ_{z} 〉] d z

(27)

since $\frac{\partial}{\partial z} P_{t - z} f = - G P_{t - z} f$ . Thus, any μ_t that satisfies (3) also satisfies (27). We show that (27) has a unique solution, which in turn implies that (3) has a unique solution.

Suppose μ_t and ν_t both satisfy (27), with μ₀ = ν₀. Let t ≥ 0.

{‖ μ_{t} - ν_{t} ‖}_{T V} = \sup_{{‖ f ‖}_{\infty} \leq 1} 〈 f, μ_{t} 〉 - 〈 f, ν_{t} 〉 = \sup_{{‖ f ‖}_{\infty} \leq 1} {\int_{0}^{t} w r 〈 x P_{t - z} f, μ_{z} - ν_{z} 〉 + w r [〈 x, μ_{z} 〉 〈 P_{t - z} f, μ_{z} 〉 - 〈 x, ν_{z} 〉 〈 P_{t - z} f, ν_{z} 〉] d z}

(28)

We can bound the first term in the integrand by

w r | 〈 x P_{t - z} f, μ_{z} - ν_{z} 〉 | \leq w r {‖ μ_{z} - ν_{z} ‖}_{T V}

because ║xP_t−zf║_∞ ≤ ║P_t−zf║_∞≤║f║_∞≤1, where the first inequality follows from x ∈ [0, 1] and second from (26). For the second term in the integrand of (28), add and subtract 〈x, ν_z〉 〈P_t₋_zf, μ_z〉:

w r | 〈 x, μ_{z} 〉 〈 P_{t - z} f, μ_{z} 〉 - 〈 x, ν_{z} 〉 〈 P_{t - z} f, ν_{z} 〉 | = w r | 〈 x, μ_{z} - ν_{z} 〉 〈 P_{t - z} f, μ_{z} 〉 + 〈 x, ν_{z} 〉 〈 P_{t - z} f, μ_{z} - ν_{z} 〉 | \leq w r ({‖ P_{t - z} ‖}_{\infty} {‖ μ_{z} - ν_{z} ‖}_{T V} + {‖ μ_{z} - ν_{z} ‖}_{T V}) \leq w r ({‖ f ‖}_{\infty} + 1) {‖ μ_{z} - ν_{z} ‖}_{T V}

again, the inequalities follow from x ∈ [0, 1], ║P_tf║_∞ ≤ ║f║_∞ and also that μ_z and ν_z are probability measures. Substituting this back into (28),

{‖ μ_{t} - ν_{t} ‖}_{T V} \leq \int_{0}^{t} 3 w r {‖ μ_{z} - ν_{z} ‖}_{T V} d z

By Gronwall’s inequality, ║μ_t − ν_t║_TV = 0, so we have uniqueness. □

Proof of Theorem 1

The uniqueness of the limit is given by Lemma 11 and the tightness of the process by Lemma 10. It remains to show that ${〈 f, μ_{t}^{m, n} 〉}_{m, n}$ converges to the solution of (3). Recall from Lemma 9 that

〈 f, μ_{t}^{m, n} 〉 - 〈 f, μ_{0}^{m, n} 〉 = A_{t}^{m, n} (f) + M_{t}^{m, n} (f)

Since tightness implies relative compactness (Prohorov’s theorem), there exists a subsequence of $μ_{t}^{m, n}$ that converges to a limit, call it μ_t. Thus, $〈 f, μ_{t}^{m, n} 〉 \to 〈 f, μ_{t} 〉$ . We also have by $〈 f, μ_{0}^{m, n} 〉 \to 〈 f, μ_{0} 〉$ assumption. In addition,

A_{t}^{m, n} (f) = \int_{0}^{t} {\sum_{i} μ_{z}^{m, n} (\frac{i}{n}) \frac{i}{n} (1 - \frac{i}{n}) [\frac{1}{n} D_{x x} f (\frac{i}{n}) - s D_{x}^{-} f (\frac{i}{n})] + w r [\sum_{j} μ_{z}^{m, n} (\frac{j}{n}) \frac{j}{n} f (\frac{j}{n}) - \sum_{i} μ_{z}^{m, n} (\frac{i}{n}) f (\frac{i}{n}) \sum_{j} μ_{z}^{m, n} (\frac{j}{n}) \frac{j}{n}]} d z \to \int_{0}^{t} {〈 - x (1 - x) s \frac{d f}{d x}, μ_{z} 〉 + w r [〈 x f (x), μ_{z} 〉 - 〈 f (x), μ_{z} 〉 〈 x, μ_{z} 〉]} d z : = A_{t} (f)

The factor of $\frac{1}{m}$ in the quadratic variation (18) implies that $M_{t}^{m, n} \to 0$ as m, n → ∞. Therefore,

〈 f, μ_{t} 〉 - 〈 f, μ_{0} 〉 = A_{t} (f)

or,

\frac{d}{d t} 〈 f, μ_{t} 〉 = 〈 - x (1 - x) s \frac{d f}{d x}, μ_{t} 〉 + w r [〈 x f (x), μ_{t} 〉 - 〈 f (x), μ_{t} 〉 〈 x, μ_{t} 〉]

□

4.3. Proof of Fleming-Viot limit

The elementary proof for tightness in Theorem 1 does not easily carry over for the case of Theorem 4. We thus use a criterion by Aldous [1] to prove tightness for the martingale part of the stochastic process.

First, consider the semimartingale formulation of $〈 f, μ_{t}^{m, n} 〉$ (16) with the rescaled parameters $s = \frac{σ}{n}$ and $ρ = \frac{r}{m}$ . Let $E_{t}^{m, n} (f) : = \int_{0}^{t} e_{z}^{m, n} (f) d z$ and $N_{t}^{m, n} (f)$ denote the drift and martingale parts of $〈 f, ν_{t}^{m, n} 〉$ , the rescaled process. Then

E_{t}^{m, n} (f) = \int_{0}^{t} \sum_{i} ν_{z}^{m, n} (\frac{i}{n}) \frac{i}{n} (1 - \frac{i}{n}) [D_{x x} f (\frac{i}{n}) - σ D_{x}^{-} f (\frac{i}{n})] + w ρ \frac{n}{m} {\sum_{j} ν_{z}^{n} (\frac{j}{n}) \frac{j}{n} f (\frac{j}{n}) - \sum_{i} ν_{z}^{n} (\frac{i}{n}) f (\frac{i}{n}) \sum_{j} ν_{z}^{n} (\frac{j}{n}) \frac{j}{n}} d z

(29)

and

{〈 N^{m, n} (f) 〉}_{t} = \int_{0}^{t} {\frac{n}{m^{2}} \sum_{i} ν_{z}^{m, n} (\frac{i}{n}) \frac{i}{n} (1 - \frac{i}{n}) [{(D_{x}^{+} f (\frac{i}{n}))}^{2} + (1 + \frac{σ}{n}) {(D_{x}^{-} f (\frac{i}{n}))}^{2}] + w \frac{n}{m} \sum_{i, j} ν_{z}^{m, n} (\frac{i}{n}) ν_{t}^{m, n} (\frac{j}{n}) (1 + \frac{ρ}{m} \frac{j}{n}) {(f (\frac{i}{n}) - f (\frac{j}{n}))}^{2}} d z

(30)

Lemma 12

The processes $〈 f, ν_{t}^{m, n} 〉$ , as a sequence in {(m, n)}, is tight for all f ∈ C²([0, 1]).

Proof

Since $〈 f, ν_{t}^{m, n} 〉 = E_{t}^{m, n} (f) + N_{t}^{m, n} (f)$ , it suffices, by the triangle inequality applied to Billingsley’s tightness criterion (Theorem 13.2 in [2]), to show tightness of E^m,n(f) and N^m,n(f) separately.

For the tightness of the finite variation term $E_{t}^{m, n} (f)$ :

| e_{t}^{m, n} (f) | \leq \frac{1}{4} \sum_{i} ν_{z}^{m, n} (\frac{i}{n}) [| D_{x x} f (\frac{i}{n}) | + σ | D_{x}^{-} f (\frac{i}{n}) |] + w ρ \frac{n}{m} {\sum_{j} ν_{z}^{n} (\frac{j}{n}) \frac{j}{n} | f (\frac{j}{n}) | + \sum_{i} ν_{z}^{n} (\frac{i}{n}) | f (\frac{i}{n}) | \sum_{j} ν_{z}^{n} (\frac{j}{n}) \frac{j}{n}}

For a given γ > 0, we can choose n and m sufficiently large such that $\frac{n}{m} \in (θ - γ, θ + γ),$ $| D_{x x} f (\frac{i}{n}) | \leq {‖ f^{″} ‖}_{\infty} + γ$ ,and $| D_{x}^{-} f (\frac{i}{n}) | \leq {‖ f^{'} ‖}_{\infty} + γ$ . We thus obtain

| e_{t}^{m, n} (f) | \leq \frac{1}{4} [{‖ f^{″} ‖}_{\infty} + γ + σ ({‖ f^{'} ‖}_{\infty} + γ)] + 2 ω ρ (θ + γ) {‖ f ‖}_{\infty}

There are only a finite number of m and n for which this condition is not satisfied. Taking the maximum of the right-hand side of the above equation with the value of $| e_{t}^{m, n} (f) |$ for such m and n, we obtain that for all m and n,

| e_{t}^{m, n} (f) | \leq G_{f}

and therefore

\sup_{t \in [0, T]} | E_{t}^{m, n} (f) | \leq G_{f} T

where G_f is a constant that depends on f. Using the same conditions for tightness as in the proof of Theorem 1, condition (i) is satisfied because $E_{t}^{m, n} (f)$ is bounded uniformly in t, m, and n. Condition (ii) is satisfied because $| E_{t + δ}^{m, n} - E_{t}^{m, n} | \leq δ G_{f}$ for all t, m, and n and therefore we can always choose δ to be sufficiently small so that $| E_{t + δ}^{m, n} - E_{t}^{m, n} | \leq ε$ for some prescribed ε.

We will show tightness for the martingale part ${〈 N_{t}^{m, n} (f) 〉}_{t}$ using Aldous’ tightness condition (we use the result as stated in [9]). First, note that by equation (30),

{〈 N_{t}^{m, n} (f) 〉}_{t} \leq J_{f} t

for f ∈ C²([0, 1]), where J_f is a constant that depends on f. Thus for fixed t,

P_{m, n} (| N_{t}^{m, n} (f) | > a) \leq \frac{1}{a} E_{m, n} | N_{t}^{m, n} (f) | \leq \frac{1}{a} {(E_{m, n} {(N_{t}^{m, n} (f))}^{2})}^{1 / 2} = \frac{1}{a} {(E_{m, n} {〈 N_{t}^{m, n} (f) 〉}_{t})}^{1 / 2} \leq \frac{\sqrt{J_{f} t}}{a}

Given ε > 0, choose $a > \sqrt{\frac{J_{f} t}{ε}}$ and we have that $N_{τ}^{m, n} (f)$ is tight for each t. Next, let τ be a stopping time, bounded by T, and let ε > 0. For κ > 0,

P_{m, n} (| N_{τ + κ}^{m, n} (f) - N_{τ}^{m, n} (f) | \geq ε) \leq \frac{1}{ε} E_{m, n} | N_{τ + κ}^{m, n} (f) - N_{τ}^{m, n} (f) |

Now (suppressing subscripts on expected value for clarity),

E | N_{τ + κ}^{m, n} (f) - N_{τ}^{m, n} (f) | \leq {[E {(N_{τ + κ}^{m, n} (f) - N_{τ}^{m, n} (f))}^{2}]}^{1 / 2} = {[E (N_{τ + κ}^{m, n} {(f)}^{2} - N_{τ}^{m, n} {(f)}^{2} + 2 N_{τ}^{m, n} (f) (N_{τ}^{m, n} (f) - N_{τ + κ}^{m, n} (f))]}^{1 / 2} = {[E ({〈 N^{m, n} (f) 〉}_{τ + κ} - {〈 N^{m, n} (f) 〉}_{τ})]}^{1 / 2} \leq \sqrt{J_{f} κ}

Hence,

P_{m, n} (| N_{τ + κ}^{m, n} (f) - N_{τ}^{m, n} (f) | \geq ε) \leq \frac{1}{ε} \sqrt{J_{f} κ}

By taking $κ < \frac{ε^{4}}{\sqrt{J_{f}}}$ , we satisfy the conditions of Aldous’ stopping criterion.

□

Lemma 13

The martingale problem (6) and (7) has a unique solution.

Proof

The martingale problem with V (t, ν, x) = 0 corresponds to a neutral Fleming-Viot with linear mutation operator. Its uniqueness has previously been established (see for example [5]). To show uniqueness for nontrivial V, we use a Girsanov-type transform by Dawson [5]. It suffices to check that

\sup_{t, μ, x} | V (t, μ, x) | \leq V_{0} (a constant)

(31)

In our case, V (t, μ, x) = x and since x ∈ [0, 1], the condition is satisfied and the martingale problem has a unique solution. □

Proof of Theorem 4

The uniqueness of the limit is given by Lemma 13 and the tightness of the process by Lemma 12. To see that the limit is the martingale problem stated in Theorem 4, note that for a fixed t,

E_{t}^{m, n} (f) \to \int_{0}^{t} \int_{0}^{1} x (1 - x) [\frac{\partial^{2}}{\partial x^{2}} f (x) - σ \frac{\partial}{\partial s} f (x)] v_{z} (d x)

w ρ θ {\int_{0}^{1} x f (x) v_{z} (d x) - \int_{0}^{1} f (x) v_{z} (d x) \int_{0}^{1} x v_{z} (d x)} d z

as n;m →∞ and

〈 N^{m, n} (f) 〉 t \to \int_{0}^{t} w θ \int_{0}^{t} \int_{0}^{t} {(f (x) - f (y))}^{2} ν_{z} (d x) ν_{z} (d y) d z

Finally, notice that

\int_{0}^{1} \int_{0}^{1} {(f (x) - f (y))}^{2} v_{z} (d x) v_{z} (d y) = 2 \int_{0}^{1} \int_{0}^{1} f {(x)}^{2} v_{z} (d x) v_{z} (d y) - 2 \int_{0}^{1} \int_{0}^{1} f (x) f (y) v_{z} (d x) v_{z} (d y) = 2 \int_{0}^{1} \int_{0}^{1} f (x) f (y) v_{z} (d x) [δ_{x} (d y) - v_{z} (d y)]

and

\int_{0}^{1} x f (x) v_{z} (d x) - \int_{0}^{1} f (x) v_{z} (d x) \int_{0}^{1} x v_{z} (d x) = \int_{0}^{1} \int_{0}^{1} f (x) y v_{z} (d x) [δ_{x} (d y) - v_{z} (d x))]

satisfying the form of the martingale problem in the theorem. □

Acknowledgments

The authors would like to thank Mike Reed and Katia Koelle for their roles in the collaboration out of which this paper’s central model grew. We would also like to thank Rick Durrett of a number of useful discussions. SL would further like to thank Don Dawson, who sent her an early version of his notes on this topic [5], and Sylvie Méléard, who took time out of a conference to explain to her the elements of the proof for weak convergence. JCM would like to thank the NSF for its support though DMS-08-54879. SL would like to thank support from the NSF (grants NSF-EF-08-27416 and DMS-0942760), NIH (grant R01-GM094402), and the Simons Institute for the Theory of Computing.

Footnotes

Recommended by Professor Leonid Bunimovich

Mathematics Subject Classification numbers: 35F55, 35Q92, 37L40, 60J28, 60J68, 60J70

References

1.Aldous David. Stopping times and tightness. The Annals of Probability. 1978;6(2):335–340. [Google Scholar]
2.Billingsley Patrick. Convergence of Probability Measures. 2nd. Wiley-Interscience; 1999. [Google Scholar]
3.Champagnat Nicolas, Ferrière Régis, Méléard Sylvie. Unifying evolutionary dynamics: from individual stochastic processes to macroscopic models. Theoretical Population Biology. 2006;69(3):297–321. doi: 10.1016/j.tpb.2005.10.004. [DOI] [PubMed] [Google Scholar]
4.Dawson DA, Greven A. Hierarchically interacting Fleming-Viot processes with selection and mutation: Multiple space time scale analysis and quasi-equilibria. Electronic Journal of Probability. 1999;4(4):1–81. [Google Scholar]
5.Dawson Donald A. Introductory Lecture on Stochastic Population Systems, Technical Report Series of the Laboratory for Research Statistics and Probability. Carleton University -University of Ottawa; 2010. (Technical Report 451). [Google Scholar]
6.Dawson Donald A, Hochberg KJ. A multilevel branching model. Advances in Applied Probability. 1991;23:701–715. [Google Scholar]
7.Donnelly Peter, Kurtz Thomas G. Genealogical processes for Fleming-Viot models with selection and recombination. Annals of Applied Probability. 1999;9(4):1091–1148. [Google Scholar]
8.Durrett Richard. Probability Models for DNA Sequence Evolution. 2nd. Springer; 2008. [Google Scholar]
9.Etheridge Alison. An Introduction to Superprocesses. American Mathematical Society; 2000. [Google Scholar]
10.Ethier Stewart N, Kurtz Thomas G. Fleming-viot processes in population genetics. SIAM Journal on Control and Optimization. 1993;31(2):345–386. [Google Scholar]
11.Fleming Wendell H, Viot Michel. Some measure-valued markov processes in population genetics theory. Indiana University Mathematics Journal. 1979;28(5):817–843. [Google Scholar]
12.Fournier Nicolas, Méléard Sylvie. A microscopic probabilistic description of a locally regulated population and macroscopic approximations. Annals of Applied Probabability. 2004;14(4):1880–1919. [Google Scholar]
13.Joffe A, Metivier M. Weak convergence of sequences of semimartingales with applications to multi-type branching processes. Advances in Applied Probability. 1986;18:20–65. [Google Scholar]
14.Kallenberg Olav. Foundations of Modern Probability. 1st. Springer; 1997. [Google Scholar]
15.Luo Shishi. A unifying framework reveals key properties of multilevel selection. Journal of Theoretical Biology. 2014;341:41–52. doi: 10.1016/j.jtbi.2013.09.024. [DOI] [PubMed] [Google Scholar]
16.Méléard Sylvie, Roelly Sylvie. A host-parasite multilevel interacting process and continuous approximations. arXiv preprint arXiv:1101.4015. 2011 [Google Scholar]
17.Pinchover Yehuda, Rubinstein Jacob. An Introduction to Partial Differential Equations. Cambridge University Press; 2005. [Google Scholar]
18.Protter Philip E. Stochastic Integration and Differential Equations. 2nd. Springer; 2004. [Google Scholar]

[R1] 1.Aldous David. Stopping times and tightness. The Annals of Probability. 1978;6(2):335–340. [Google Scholar]

[R2] 2.Billingsley Patrick. Convergence of Probability Measures. 2nd. Wiley-Interscience; 1999. [Google Scholar]

[R3] 3.Champagnat Nicolas, Ferrière Régis, Méléard Sylvie. Unifying evolutionary dynamics: from individual stochastic processes to macroscopic models. Theoretical Population Biology. 2006;69(3):297–321. doi: 10.1016/j.tpb.2005.10.004. [DOI] [PubMed] [Google Scholar]

[R4] 4.Dawson DA, Greven A. Hierarchically interacting Fleming-Viot processes with selection and mutation: Multiple space time scale analysis and quasi-equilibria. Electronic Journal of Probability. 1999;4(4):1–81. [Google Scholar]

[R5] 5.Dawson Donald A. Introductory Lecture on Stochastic Population Systems, Technical Report Series of the Laboratory for Research Statistics and Probability. Carleton University -University of Ottawa; 2010. (Technical Report 451). [Google Scholar]

[R6] 6.Dawson Donald A, Hochberg KJ. A multilevel branching model. Advances in Applied Probability. 1991;23:701–715. [Google Scholar]

[R7] 7.Donnelly Peter, Kurtz Thomas G. Genealogical processes for Fleming-Viot models with selection and recombination. Annals of Applied Probability. 1999;9(4):1091–1148. [Google Scholar]

[R8] 8.Durrett Richard. Probability Models for DNA Sequence Evolution. 2nd. Springer; 2008. [Google Scholar]

[R9] 9.Etheridge Alison. An Introduction to Superprocesses. American Mathematical Society; 2000. [Google Scholar]

[R10] 10.Ethier Stewart N, Kurtz Thomas G. Fleming-viot processes in population genetics. SIAM Journal on Control and Optimization. 1993;31(2):345–386. [Google Scholar]

[R11] 11.Fleming Wendell H, Viot Michel. Some measure-valued markov processes in population genetics theory. Indiana University Mathematics Journal. 1979;28(5):817–843. [Google Scholar]

[R12] 12.Fournier Nicolas, Méléard Sylvie. A microscopic probabilistic description of a locally regulated population and macroscopic approximations. Annals of Applied Probabability. 2004;14(4):1880–1919. [Google Scholar]

[R13] 13.Joffe A, Metivier M. Weak convergence of sequences of semimartingales with applications to multi-type branching processes. Advances in Applied Probability. 1986;18:20–65. [Google Scholar]

[R14] 14.Kallenberg Olav. Foundations of Modern Probability. 1st. Springer; 1997. [Google Scholar]

[R15] 15.Luo Shishi. A unifying framework reveals key properties of multilevel selection. Journal of Theoretical Biology. 2014;341:41–52. doi: 10.1016/j.jtbi.2013.09.024. [DOI] [PubMed] [Google Scholar]

[R16] 16.Méléard Sylvie, Roelly Sylvie. A host-parasite multilevel interacting process and continuous approximations. arXiv preprint arXiv:1101.4015. 2011 [Google Scholar]

[R17] 17.Pinchover Yehuda, Rubinstein Jacob. An Introduction to Partial Differential Equations. Cambridge University Press; 2005. [Google Scholar]

[R18] 18.Protter Philip E. Stochastic Integration and Differential Equations. 2nd. Springer; 2004. [Google Scholar]

PERMALINK

SCALING LIMITS OF A MODEL FOR SELECTION AT TWO SCALES

SHISHI LUO

JONATHAN C MATTINGLY

Abstract

1. Introduction

Figure 1.

2. Main results

Theorem 1

Example 1

Example 2

Lemma 2 (Fixed points)

Theorem 3 (Steady state behavior)

Theorem 4

3. Properties of the deterministic limit

Lemma 5

Remark 1

Remark 2

Proof of Lemma 5

Lemma 6

Proof

Example 3

Example 4

Example 5

Example 6

Example 7

Proof of Lemma 2

Lemma 7

Lemma 8

Proof of Lemma 8

Proof of Lemma 7

Proof of Theorem 3

4. Proofs of weak convergence

4.1. Semimartingale property of multilevel selection process

Lemma 9

Proof

4.2. Proof of deterministic limit

Lemma 10

Proof

Lemma 11

Proof

Proof of Theorem 1

4.3. Proof of Fleming-Viot limit

Lemma 12

Proof

Lemma 13

Proof

Proof of Theorem 4

Acknowledgments

Footnotes

References

ACTIONS

PERMALINK

RESOURCES

Similar articles

Cited by other articles

Links to NCBI Databases