Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Feb 15;26(2):e70107. doi: 10.1111/1755-0998.70107

Estimating Effective Number of Breeders Under the Heterozygote‐Excess Equilibrium

Bing Yang 1, Dan Wang 1, Yujia Shen 1, Wenkai Li 1, Haiya Wang 2, Ruqi Wang 1, Ruixin Peng 1, Song Zhang 1, Zhiming Xia 2, Derek W Dunn 1, Kang Huang 1,, Baoguo Li 1,3
PMCID: PMC12907470  PMID: 41693055

ABSTRACT

Estimating the effective number of breeders per reproductive cycle or cohort (Nb) can contribute to understanding the viability of wild populations. However, previous Nb estimators based on heterozygote‐excess assume that parental genotypes are in accordance with Hardy–Weinberg equilibrium (HWE), and can only be applied to diploids. When these criteria are not fulfilled, estimates of expected heterozygosity may thus be biased. We extended a previous Nb estimator to autopolyploids and account for parental genotypes either in HWE or heterozygote‐excess equilibrium (HEE). Results include the following: the distributions of genotypes converge to HEE within nine generations (with a relative error of 103), the corrected estimators can asymptotically estimate Nb without bias, and estimations for autopolyploids are slightly better than those for diploids. The relationships between N^b and Nb and the influence of variable effective breeding numbers across generations on N^b were clarified. We find N^b mainly reflects Nb in the previous generation in diploids but is a weighted harmonic mean of Nb in previous generations in polyploids. Our methods are implemented in a software package, polygene v1.7, which is freely available at https://github.com/huangkang1987/polygene.

Keywords: effective number of breeders, Hardy–Weinberg equilibrium, heterozygote‐excess equilibrium, polyploids

1. Introduction

The effective population size, Ne, is an important parameter in population genetics in both theoretical and practical applications to animal and plant breeding, ecology, evolutionary biology and conservation biology (Falconer and Mackay 1996) and evolutionary theory (Charlesworth and Charlesworth 2010). Effective population size reveals the dynamics of genetic drift, selection, and gene flow and provides insights into the evolutionary history of species and populations, improving the understanding of complex patterns of speciation and adaptation. The effective number of breeders (Pudovkin et al. 1996) in one reproductive cycle (Nb), is a more informative metric to use than Ne in some cases, such as for populations with clearly defined age structure or with overlapping generations. Because Nb only requires data for a single breeding season rather than an entire life span, the term is easier to estimate than Ne (Schwartz et al. 2007).

Several methods have been developed to indirectly estimate Ne or Nb from genetic data, which can be broadly classified into two main categories: temporal and single‐sample methods. Temporal methods use genetic data from two or more temporally spaced samples obtained from the same population (Nei and Tajima 1981; Waples 1989), which is difficult for long‐lived species. In contrast, the single‐sample methods utilise genetic data from a single sample collected at a specific point in time to estimate the Ne or Nb of a population at or near to that time. Single‐sample methods require less time, cost and sampling effort than temporal methods, and have therefore recently increased in popularity (Waples 2006; Tallmon et al. 2008; Waples and Do 2008, 2010; Wang 2009, 2016; Huang et al. 2022). For discrete generation models, estimates from single‐sample and temporal methods differ in the timeframes to which they apply. The single‐sample methods estimate the inbreeding Ne, whereas the temporal methods operate based on the variance Ne (Waples 2005).

The heterozygote‐excess method is a single sample method and assumes the breeders are randomly sampled from a wider population. When there are few breeders, the binomial sampling error results in a heterozygote‐excess in the offspring (Pudovkin et al. 1996). There are several deficiencies in current heterozygote‐excess methods: (i) existing methods assume unrealistic conditions that parental genotypes are in accordance with HWE. However, HWE is not an equilibrium state in finite populations with a dioecious mating system and the distributions of genotypes will eventually converge to some new equilibrium states; (ii) the estimation of expected heterozygosity, and Selander's (1970) heterozygote‐excess index D, are biased. Although Pudovkin et al. (2009) applied a Gaussian correction to obtain an unbiased expected heterozygosity, the influence of a correlation between allele copies was not considered; (iii) current methods are developed for diploids and cannot be directly applied to autopolyploids. However, polyploidy has occurred in every ancestral plant lineage (Otto 2007), and is frequently observed in extant species (Burow et al. 2001; Comai 2005). Polyploidy plays a significant role in the formation and evolution of plant species, and has received much recent attention in evolutionary ecology research (Selmecki et al. 2015; Avni et al. 2017; Burns et al. 2021; Lovell et al. 2021).

Here, we aim to (i) extend the definitions of both Nb and Ne to account for autopolyploids, and quantify the relationship between these two metrics; (ii) develop a Nb estimator without the need to assume unrealistic conditions and correct any bias from expected heterozygosities; (iii) extend original and corrected estimators to the case of autopolyploids. We also use symbolic and numerical calculations to study the convergency of heterozygote‐excess index D and to compare the statistical performance of different estimators. By comparing the results of Pudovkin et al.'s (1996) estimator N^b,Pu and our corrected estimator N^b, (that is the estimator in which the genotypes of generation are assumed to follow HWE) by polygene v1.7 in two empirical datasets, we find that our estimator is more stable than N^b,Pu.

2. Methods

2.1. Ne and Nb for Autopolyploids

In the dioecious with random paring (DR) mating system (Weir and Hill 1980), in which, each offspring is reproduced by a new pairing, and there are F breeding females and M breeding males in the population for each generation, the effective number of breeders is defined as

Nb4FMF+M.

This definition is compatible with autopolyploids, and the value of Nb is close to the effective population size Ne in diploids.

In polyploids, the effective population size has several definitions. Some studies define Ne of an ideal polyploid population with N individuals as N (e.g., Pollak and Sabran 1996; Lynch et al. 2016); others define Ne as Nv/2, with v being the ploidy level (Hamilton 2009) because polyploids have a lower rate of genetic drift.

Here, we define Ne as the size of an ideal population with a haplotype sampling (HS) mating system (Huang et al. 2022) that has the same rate of decrease in heterozygosity λ at equilibrium as the real population. The HS mating system assumes that each individual is the product of randomly sampling v haplotypes with replacement from the previous generation, and is reduced to the DR mating system in haploid and diploid Wright–Fisher populations. Because the rate of decrease in heterozygosity may vary among generations, we use this at equilibrium state in order to define Ne:

Ne11λv.

We use this definition of Ne because: (i) the expression of λ of a polyploid ideal population with HS mating system is simple (Equation (2)); (ii) λ of a polyploid ideal population with N individuals is identical to that of a haploid Wright–Fisher population with Nv individuals, or a diploid Wright–Fisher populations with Nv/2 individuals; (iii) the value of Ne of a real polyploid population is still close to Nb.

2.2. The Rate of Decrease in Heterozygosity

According to the above definition of Ne, we need to derive the decay of HO and HE under two mating systems. The recurrence of heterozygosity can be expressed by matrix multiplication:

forHS:EHOEHE=v1NevNe1Nev1NevNe1NeHOHE=THSHOHE,
forDR:EHOEHE=v22v2v2v2v1NbvNb1NbHOHE=TDRHOHE.

Therefore, given initial heterozygosities, the observed and expected heterozygosities at generation t can be expressed by

EHO,t,HO,0|HE,0EHE,t,HO,0|HE,0=TtHO,0HE,0. (1)

The derivations are provided in Appendix A.

By performing eigen‐value decomposition for the coefficient matrix T, we obtain the values of λ under different mating systems:

λHS=11vNe,
λDR=c1+c24Nbv1, (2)

where

c1=8Nb+443Nb+2v+Nb+22v2,
c2=22v+3Nbv4Nb.

The corresponding eigen‐vectors for λHS and λDR are

UHS=11,UDR=Deq+11,

respectively, where Deq is the Selander's (1970) heterozygote‐excess index at heterozygote‐excess equilibrium, which is

Deq=4v4Nbv2+2v26v+c1v+4. (3)

The value of Ne for a polyploid population with a DR mating system can be expressed as a function of Nb by equating λHS=λDR:

Ne=2v2+Nbv+c12v.

By expanding the zeroth‐order Taylor series for the above equation at Nb, we find that Ne is approximately equal to Nb plus a constant:

NeNb+2v12v2.

The constant for v2,4,6,8,10 is 12,98,2518,4932,8150 and Ne is at most Nb+2 at v.

2.3. Parents in Accordance With HWE

Previous studies (Pudovkin et al. 1996, 2009; Huang et al. 2019) have derived and estimated Nb assuming that parental genotypes are in accordance with HWE (abbreviated hereafter to HWE0). The model assumes a dioecious ideal population, with each generation having an infinite number of individuals. In each generation, M males and F females are randomly sampled in order to produce the next generation, each individual generates an infinite number of gametes and these gametes randomly unite to generate the offspring.

According to Equation (1), we derive expected heterozygosities under HWE0:

EHO,1|HWE0=1 0TDR1 1THE,0,
EHE,1|HWE0=0 1TDR1 1THE,0. (4)

Therefore, Selander's heterozygote‐excess index is

ED1|HWE0=EHO,1|HWE0EHE,1|HWE01=1Nbv1.

By replacing ED1|HWE0 with D^ and Nb with N^b, we obtain an estimator of Nb assuming HWE0:

N^b,0=1D^v+1v.

The special cases for 2v10 are listed in Table 1.

TABLE 1.

The estimators of Nb.

v
N^b,Pu
a
N^b,0
a
N^b,
a
2
12D^+121+D^
12D^+12
12D^12D^
4
14D^18316D^+8732D^2
14D^+14
38D^3498D^
6
16D^718+5554D^+1085162D^2
16D^+16
518D^1092518D^
8
18D^1732+329128D^+5887512D^2
18D^+18
732D^21164932D^
10
110D^3150+1071250D^+208891250D^2
110D^+110
950D^36258150D^
a

N^b,Pu denotes the original Pudovkin et al. (1996) Nb estimator. N^b,0 and N^b, are corrected estimators assuming HWE0 and HWE, respectively. The expressions of N^b,Pu (with v>2), were expanded to second‐order Taylor series at D^=0.

2.4. Generation t in Accordance With HWE

Equation (4) can be extended to the scenario that the genotypes of generation t are in accordance with HWE (abbreviated hereafter to HWEt):

EHO,1|HWEt=1 0TDR1t1 1THE,t,
EHE,1|HWEt=0 1TDR1t1 1THE,t.

In this case, Selander's heterozygote‐excess index is

ED1|HWEt=4v42v26v+4+Nbv2+c1v+2c1vc2+c1c2c11t1. (5)

For t, ED1HWEt converges to Deq when Nb2, thus ED1HWE=Deq and we can estimate Nb under the assumption of HWE:

N^b,=2v1v2D^4v26v+22v24v+2D^. (6)

The special cases for 2v10 are listed in Table 1, with the computer derivations of N^b,0 and N^b, provided in File S1.

2.5. Pudovkin's (1996) Estimator

Pudovkin et al. (1996) original estimator can only be used for diploids and also assumes HWE0. Here, we use this method to expand the estimator to account for polyploids. The observed heterozygosity in generation 1 can be expressed by:

HO,1=v21v/22HF,0+HM,0+v22pF,0+pM,02pF,0pM,0
=v21v/22HF,0+HM,0+v22HE,1+12d02, (7)

here pF,0 and pM,0 denote the allele frequency of the breeding females and males in generation 0, respectively, and d0 is equal to pM,0pF,0. By taking the expectation of both sides of the above equation (Appendix B), and using the approximation

EHE,1λHE,0.

The expectation of HO,1 is approximately

2.5.

Therefore, the expectation of Selander's heterozygote‐excess index is

ED|HWE0=EHO,1|HWE0EHE,111Nbv2λ12Nbv1λ.

Substituting the expression of λ (Equation (2)) into the above equation, and using ED as known and Nb as the unknown, the expression of Nb as a function of ED can be solved. Then replacing Nb with N^b and ED with D^ (which can be calculated from the samples), we obtained the estimator of Nb.

For v=2, the result is identical to Pudovkin et al. (1996):

N^b,Pu=12D^+121+D^.

For v>2, the expression is complex, for example, in autotetraploids,

N^b,Pu=18D^15D^2+1+3D^21+2D^+25D^28D^12D^2.

For simplicity, we expand the second‐order Taylor series at D^=0, and list the estimators of Nb for 2v10 in Table 1. The computer derivations of N^b,Pu are provided in File S2.

2.6. Unbiased Estimation of HE

The ‘traditional’ estimator of HE is H^E=1kp^k2, which is the probability of sampling two non‐identical‐by‐state (non‐IBS) allele copies with replacement. However, the pairs for the same allele copy are always IBS, which results in an underestimation of HE. We use Bkia to derive the expectation of H^E, and Bkia is one if the a‐th allele copy in individual i is allele k, otherwise zero. Therefore,

H^E=1kp^k2=11n2v2ki,a,i,aBkiaBkia
=11n2v2k,i,aBkia+k,i,aaBkiaBkia+iik,a,aBkiaBkia.

Its expectation is

EH^E=11nvkEBkia+v1kEBkiaBkia+n1vkEBkiaBkia,

where n is the sample size, BkiaBkia and BkiaBkia are the products of indicators for allele pairs within the individual and between individuals, respectively, and it can be inferred that EBkia=pk, EBkiaBkia=CovBkia,Bkia+pk2, EBkiaBkia=CovBkia,Bkia+pk2.

The method of Pudovkin et al. (1996) ignores the correlation between alleles, and assumes kEBkiaBkia=kEBkiaBkia=1HE, then the above expectation can be simplified as

EH^E=11nv1+nv11HE,

Then, by substituting EH^E with H^E and HE with H~E, the corrected HE estimator is given by

H~E=nvnv1H^E.

We consider the influence of the correlation between alleles and only use the allele pairs between individuals to estimate HE. In this case, EBkia=pk, kEBkiaBkia=1HO, kEBkiaBkia=1HE. Then the expectation of HE is revised as

EH^E=11nv1+v11HO+n1v1HE.

Then substituting EH^E with H^E and HE with H~E, the corrected HE estimator is given by

H~E=H^EnvH^Ov1nvv.

We use D^ and D^ to distinguish the two heterozygote‐excess indices calculated from different HE estimator:

D^=H^OH~E1,D^=H^OH~E1. (8)

2.7. Simulation Method

Monte‐Carlo simulations were conducted to investigate the changes of D and Nb across generations and assess the statistical performance of various Nb estimators. A population with DR mating system was simulated, comprising an initial parental generation of M breeding males and F breeding females. Parental genotypes at L unlinked loci were randomly drawn according to either HWE or HEE (i.e., HWE0 or HEE0). To exclude the influence of locus polymorphism on the estimation, we generated offspring data with EHO=0.32 at diallelic markers (Appendix C). We used HO to evaluate locus polymorphism, because HE is sensitive to deviations from the HWE (e.g., Wahlund effect).

The parents then mate randomly to produce the offspring. We simulated R populations and calculated the harmonic mean, the root‐mean‐squared‐error (RMSE) of N^b and 1/N^b for each parameter combination. For each population, n individuals were sampled from the offspring generation and were genotyped at all L loci. The values of D^ and D^ were calculated by summing the heterozygosites across loci. For example,

D^=lH^O,llH~E,l1,

where H^O,l and H~E,l denotes the estimated observed and expected heterozygosities at locus l.

When D^<0, N^b is negative. The main reason for the negative value of the Nb estimate is that for the heterozygote excess method, when there is inbreeding or substructure in the population, there will be insufficient heterozygotes. Thus, negative D and Nb estimates will be obtained. Moreover, when D^ is close to zero, N^b can be very large (e.g., 10100). Although negative Nb estimates present conceptual challenges, these values still reflect significant information about genetic drift, thus the recommended approach involves calculating the harmonic mean across all estimates including negative ones as the measure of central tendency. Moreover, systematically ignoring negative or infinite estimates leads to estimation bias of N^ (Waples 2024, 2025). We carried out five simulations.

Simulation 1 compared the changes of theoretical and estimated D across generations to validate our derivations. The founders were under HWE and the population was reproduced for 12 generations, each generation had exactly Nb individuals with a sex ratio of 0.5. For each generation, all individuals were sampled to estimate D using Equation (8), while the theoretical D was calculated using Equation (5). Each allele copy in the founder generation is unique so as to offset genetic drift. The key parameters and their settings were as follows: the ploidy level v was set to 2,4,6; the effective population size Nb to 2,6,10; the number of alleles per locus K (here set as Nbv) to the corresponding products; the number of loci L to 2000; the sample size to n=Nb; and the number of simulation replicates R to 2000. D^ and D^ were calculated by summing the heterozygosities across 2000 populations and 2000 loci, for example, the value of D^ at generation t was calculated by

D^t=ptlH^O,ptlptlH~E,ptl1,

where H^O,ptl and H~E,ptl denote the estimated observed and expected heterozygosities of population p in generation t at locus l. The Jackknife method was not applied for this simulation. Further, to study the convergence speed of D, we also use symbolic calculation to obtain the number of generations Neq required to reach a predefined relative error.

Simulation 2 studied the behaviour of N^b, under population expansion and decline. The Nb before generation 1 was 2 (or 1024), hence we generated genotypes for generation 1 under HEE. After generation 1, the Nb value of each subsequent generation is twice or half that of the previous generation. Each generation had 1024 individuals and generation 10 is the final generation. The parameters were v2,4,6, Nb,t=2t+1 (or 210t), K=2, L=1000, n=1024 and R=10000. The Jackknife method was not applied for this simulation.

Simulation 3 studied the distribution of N^b to show asymptotically that N^b is unbiased, where v2,4,6, Nb10,30,100, K=2, L=2000, n=1000 and R=2000.

Simulation 4 evaluated the harmonic mean of N^b, the RMSE of N^b and 1/N^b as the number of loci L changes, where v2,4,6, Nb10,30,100, n=500, L200,400,,2000 and R=10,000.

Simulation 5 evaluated the harmonic mean of N^b, the RMSE of N^b and 1/N^b as the sample size n changes, where v2,4,6, Nb10,30,100, L=1000, n=100,200,,1000 and R=10,000.

Two additional estimators implemented in polygene v1.7, Nomura's (2008) estimator N^b,No and Huang et al.'s (2022) estimator N^b,Hu were also compared in Simulations 4 and 5. However, they operate under different assumptions and may perform suboptimally when their assumptions are violated. To ensure a fair comparison, we generated offspring data under their respective assumptions: N^b,Pu and N^b,0 assume HWE0, N^b, assumes HEE0, N^b,No does not rely on assumptions regarding the distribution of genotypes of parental generation and utilises HEE0, while N^b,Hu assumes equilibrium for the squared LD coefficient, rΔ2. This equilibrium was simulated by generating a founder generation under HEE and reproduced over approximately 6 generations using the equation 4.21/log1c6.

The C++ source code is provided in File S3, the simulation programme, parameter files and batch files are provided in File S4, and the results of processing scripts are provided in File S5.

2.8. The Number of Loci Required

We also derived the number of loci required to reach a predefined coefficient of variation (CV) under HEE0. We used CV instead of variance or RMSE, because these statistics are influenced by EN^b, and a smaller EN^b usually has a smaller VarN^b.

We first derived the moments of heterozygosities of given parental allele frequencies pf and pm (as detailed in Appendix D, Tables S1 and S2, File S6), then derived the variance of N^b, (as detailed in Appendix E, File S7). These procedures are briefly described as follows: (i) using the model described in Appendix C, we derived the moments of heterozygosities of given allele frequency p and α, where α is used to generate a heterozygote‐excess sample; (ii) by assuming a uniform a priori allele frequency distribution (Kimura 1954), p is eliminated from the conditions of these moments; (iii) according to Equation (6), the variance of N^b, at L loci VarN^b,L is approximately 4v12v4·Var1/D^L, where Var1/D^L is approximated by the variance of ratios (Stuart and Ord 1998).

By solving the equation CV=VarN^b,L/Nb (see File S8), we calculated the number of loci required to reach a predefined CV10%15%20% for a population with Nb10,30,100,300,1000,3000 and n50,100,,3000.

2.9. Empirical Data

Kamath et al. (2015) analysed 729 Yellowstone grizzly bears ( Ursus arctos ) from an isolated and well‐documented population in the Greater Yellowstone Ecosystem. These bears, born between 1962 and 2010, were sampled and genotyped at 20 microsatellite loci. Herein we analyse the genotype data of the sampled individuals born in the periods 1983–2008. We used the sampling method in Kamath et al. (2015), that is, to ensure sufficient sample sizes and account for intermittent breeding, bear samples were grouped into 3‐year overlapping cohorts based on birth year from 1983 to 2008. We calculated the estimates and 95% confidence intervals of N^b,Pu and N^b, by polygene v1.7 in this Yellowstone grizzly bear population. The 95% confidence intervals are estimated by the Jackknife method.

We also used another empirical genotypic dataset to evaluate the performance of N^b,Pu and N^b, in handling SNPs. The brook trout ( Salvelinus fontinalis ) dataset (Anne‐Laure et al. 2020) comprised genotype data from 1416 individuals across 50 populations, analysed at 14,779 high‐quality filtered SNPs.

3. Results

Simulations were performed using different combinations of the following population parameters: generation (t), ploidy level (v), number of loci (L), number of alleles per locus (K) and other parameters.

3.1. Simulation 1

The results of simulation 1 are shown in Figure 1. It can be found that D, D^ and D^ converge to Deq. Deq is smaller in higher Nb and ploidy, which is approximately 2v1Nbv2. D^ seems to be an unbiased estimator of D, while D^ is downwardly biased. The absolute bias of D^ decreases as Nb increases and can be reduced to a negligible level for Nb=10, but increases as v increases. This shows that the amplitude of the curves of D decreases as the number of generations increase (damped vibration) around Deq at v=2 or Nb=2.

FIGURE 1.

FIGURE 1

The values of D, D^ and D^ as a function of the number of generations reproduced. The subfigures in different columns are for different levels of ploidy. Empty and filled dots denote D^ and D^, respectively, and the solid curves denote D. Different colours denote different Nb, where red, blue and magenta objects denote Nb = 2, 6 and 10, respectively.

The numbers of generations Neq required to reach a predefined relative error are shown in Table 2. For polyploids, Neq increases as Nb increases, whilst in diploids, Neq decreases as Nb increases.

TABLE 2.

The number of generations to allow the relative error DDeqDeq<102,104 and 106.

Nb
Diploids Tetraploids Hexaploids Octoploids Decaploids
2 5/10/15 3/4/6 a 2/4/5 2/3/5 2/3/4
5 3/5/7 3/6/9 4/7/10 4/7/11 4/7/11
10 2/4/5 4/7/11 5/9/13 5/9/13 5/9/14
100 1/2/3 5/9/13 5/10/15 6/11/16 6/12/17
a

Cell values (e.g., 3/4/6) indicate the number of generations required for Nb=2 and tetraploids to achieve relative error thresholds of heterozygote excess index D below 102,104 and 106, respectively.

3.2. Simulation 2

The Nb for the preceding generation, the mean N^b,, the mean N^b, predicted by Equation (9) and by Equations (10) and (11) are shown in Figure S7. It can be found that: (i) N^b, does not estimate the Nb of the previous generation, but rather estimates some mean Nb across several preceding generations; (ii) in diploids, the Nb of the prior generation contributes more to the mean Nb, whereas its contribution is relatively less significant in polyploids; (iii) Equation (9) can precisely predict N^b, under any conditions, while Equations (10) and (11) exhibit decreased precision, particularly at lower Nb. This is because Equations (10) and (11) are first‐order Taylor series expansions at D0=0 and Deq,0=0.

3.3. Simulation 3

We used violin plots to show the distribution of N^b (Figure 2). We found that N^b,Pu is nearly unbiased when its core assumptions are met (diploids and HWE0), but is downwardly biased for polyploids (either HWE0 or HEE0) and upwardly biased for diploids with HEE0.

FIGURE 2.

FIGURE 2

Violin plots for N^b. The subfigures in different columns show cases for HWE0 or HEE0, and the subfigures in different rows show the results for different values of Nb. The blue, red and black symbols represent the probability density functions of N^b,Pu, N^b,0 and N^b,, respectively. The grey bars denote the mean estimate and the error bars denote the quartiles. Di, Tetra and Hexa denote results for diploids, autotetraploids and autohexaploids, respectively.

N^b,0 and N^b, are nearly unbiased under their own assumptions, but become biased when these assumptions are not met. Under HEE0, N^b,0 becomes downwardly biased in polyploids and upwardly biased in diploids. Under HWE0, N^b, becomes upwardly biased in polyploids and downwardly biased in diploids.

For larger values of Nb, although some estimators are still unbiased, variance tends to increase.

3.4. Simulation 4

The harmonic means of the estimates, along with the RMSE of the five Nb estimators and that of their reciprocals, are shown as functions of the number of loci L in Figure 3, Figures S1 and S2, respectively. The proportions of positive estimates of N^b,Pu, N^b,0 and N^b, under different numbers of loci are present in Figure S3.

FIGURE 3.

FIGURE 3

The harmonic means of the five N^b as functions of the number of loci L under their own assumptions. The different columns show the harmonic means for different levels of ploidy. The different rows show the harmonic means for different values of Nb. The blue empty dot, red filled dot, black line, magenta line and green line are the estimators N^b,Pu, N^b,0, N^b,, N^b,Hu and N^b,No, respectively. The black dotted line denotes the baseline which equals to the true Nb. Di, Tetra and Hexa denote results for diploids, autotetraploids and autohexaploids, respectively.

We found that N^b,Pu, N^b,0 and N^b, are asymptotically unbiased under their respective assumptions, but as Nb increases, for example when Nb=100, achieving asymptotic unbiasedness in the estimator becomes progressively more challenging. The harmonic mean of N^b,No quickly converges with more loci and closely matches the true value for larger Nb. The harmonic mean of N^b,Hu exhibits little to no change with the number of loci, and increases with the level of ploidy. For all five estimators, the reciprocals yield a dramatically reduced RMSE compared to the raw estimates.

The ploidy level slightly affects the harmonic means, the RMSE of Nb estimators and the RMSE of the reciprocals of these estimators: (i) for diploids, the harmonic means of these estimators are closer to the true value, in comparison, their RMSE values (and those of their reciprocals) are slightly lower in diploids than in polyploids; (ii) diploids require fewer loci than polyploids for these estimators to approach unbiasedness. In contrast, Nb has a more pronounced influence: (i) as Nb increases, more loci are required for these estimators to become nearly unbiased and achieving asymptotic unbiasedness becomes increasingly difficult; (ii) the RMSE of N^b, rises sharply from 0.3 to about 60 when the loci number L>1500 as Nb increases from 10 to 100, and the situation deteriorates further when fewer loci are used. In all cases, the RMSE is significantly lower for the reciprocals of these estimators compared to the estimators themselves.

3.5. Simulation 5

The harmonic means of the estimates, along with the RMSE of the five Nb estimators and that of their reciprocals, are shown as functions of sample size in Figure 4, Figures S4 and S5, respectively. The proportions of positive estimates of N^b,Pu, N^b,0 and N^b, across a range of sample sizes are present in Figure S6. In accordance with the current parameter settings, the harmonic means of the four estimators (N^b,Pu, N^b,0, N^b,Hu and N^b,) exhibit highly similar response patterns to variations in both sample size and number of loci, with largely consistent convergence behaviour and asymptotic values.

FIGURE 4.

FIGURE 4

The harmonic means of the five N^b as functions of the number of samples n under their own assumptions. The different columns show the harmonic means for different levels of ploidy. The different rows show the harmonic means for different values of Nb. The blue empty dot, red filled dot, black line, magenta line and green line are the estimators N^b,Pu, N^b,0, N^b,, N^b,Hu and N^b,No, respectively. The black dotted line denotes the baseline which equals to the true Nb. Di, Tetra and Hexa denote results for diploids, autotetraploids and autohexaploids, respectively.

However, for the estimator N^b,No, a pronounced discrepancy was observed in the response pattern of the harmonic mean to changes in sample size versus changes in the number of loci. The harmonic mean of N^b,No shows minimal variation with sample size and increases with the level of ploidy for Nb30. When Nb=100 with small sample size, the harmonic mean of N^b,No remains close to the true value. However, it decreases rapidly as the sample size increases.

3.6. The Number of Loci Required

The numbers of loci LCV required to reach a predefined CV as a function of n are shown in Figure 5. We found that LCV decreases as n, v and CV increase, while LCV increases as Nb increases. A polyploid sample with a large n and small Nb will result in the most precision in estimating Nb. When Nb100, the curves of LCV for different values of Nb are isometrically distributed.

FIGURE 5.

FIGURE 5

The number of loci required to reach a predefined coefficient of variation. The subfigures in different columns show the results from different predefined values of CV. The subfigures in different rows show the results for different levels of ploidy. Di, Tetra and Hexa denote results for diploids, autotetraploids and autohexaploids, respectively. In each subfigure, the different curves indicate various values of Nb.

3.7. Empirical Data

Tables S4 and S5 present the effective breeder number estimates N^b,Pu and N^b, (with 95% CIs) for grizzly bears ( Ursus arctos ; SSR data) and brook trout ( Salvelinus fontinalis ; SNP data), respectively. For grizzly bears, the harmonic means calculated across all 3‐years cohorts are 185.49 for N^b,Pu and 149.05 for N^b,. While in the case of brook trout data, the harmonic means across populations yield 87.92 for N^b,Pu and 70.65 for N^b,.

Our analysis reveals three key patterns: (i) in both datasets, for the vast majority of populations or years with positive estimates, N^b, produces consistently lower values than N^b,Pu, despite following a similar overall trend, and this systematic reduction is further supported by the final harmonic mean results; (ii) N^b, systematically exhibits lower variance yet wider 95% confidence intervals relative to N^b,Pu; and (iii) the rates of positive values vary across different marker datasets (SSR: 50% vs. SNP: 62%).

4. Discussion

4.1. Dynamics of D

We found from Equation (5) that the damped vibration of D at v=2 or Nb=2 in Figure 1 is due to the sign of c2+c1c2c1. If negative, the curve of D exhibits damped vibration. By solving the inequality c2+c1c2c1<0, we found that it holds in two scenarios: (i) v=2 and Nb>0, or (ii) v>2 and 0<Nb<2v3v2. This is consistent with our simulation (Figure 1).

According to numerical simulations (data not shown), we found D quickly converges to Deq, and reaches a relative error of 102,103,104,105,106,107 with at most 6,9,12,14,17,20 generations for 2v10. This suggests that as Nb changes, for example, due to population expansion or a bottleneck effect, D will quickly converge to a new Deq within several generations. Therefore, the values of D are mainly influenced by recent values of Nb, implying that our Nb estimator reflects recent values of Nb.

4.2. The Relationship Between N^b and Nb

For a given sample of n individuals, the pedigree has been determined. In this case, the expectation of H^O is still equal to v22v2HO,0+v2v2HE,0, because the ratio of the probabilities that two allele copies within an individual are from the same gamete or different gametes is not influenced by the pedigree, and is always equal to v2:v. The expectation of H~E for a given sample varies around v1NbvHO,0+Nb1NbHE,0, because the probability that two allele copies between individuals are from the same parents depends on the pedigree. We derived the expectation of H~E in Appendix F, and found that this is equal to v1NpvHO,0+Np1NpHE,0. Here, the term Np=4Psf+Psm represents the effective number of parents contributing to the sampled cohort, where Psf (or Psm) is the probability of two allele copies in eggs (or sperms) that form the sampled individuals (or breeding individuals) are from the same parent.

The expectation of Np was derived in Appendix F, and can be calculated as

ENp|rNb+NbNb4+8r8r22nn1,

where r is the proportion of males in the breeding individuals in the previous generation. Without a priori information, we assume r is uniformly distributed and take the integral of ENp over r0,1 to eliminate r:

ENpNb+NbNb832nn1.

The computer derivations are provided in File S9.

We therefore conclude that N^b actually estimates Np, N^b converges to Np as L, Np is greater than Nb if Nb>8/3, and Np converges to Nb as n.

4.3. N^b With Fluctuating Nb

For variation of Nb or Np across generations, we consider the transition from an initial state D0 to the new equilibrium state Deq for the non‐vibrated D (polyploids and Nb>2). According to Equation (1), the expectation of D in generation t is:

4.3. (9)

Expanding the first‐order Taylor series for ED1 at D0=0 and Deq,0=0, and eliminating the remainder of the Taylor series expansion, we obtained

ED1v22v2D0+v2v2Deq,0.

Therefore, ED1 can be approximated by a linear combination of D0 and Deq,0, then assuming the population reproduces for several more generations the expectation of D in generation t is

EDtv22v2tD0+v2v2·i=1tv22v2ti·Deq,i1.

Therefore, EDt is a linear combination of Deq in previous generations. For example, in autotetraploids and autohexaploids,

EDt,v=423·Deq,t1+29·Deq,t2+227·Deq,t3+,
EDt,v=6610·Deq,t1+24100·Deq,t2+961000·Deq,t3+.

When Nb is large, Deq2v1Nbv2 according to Equation (3). Assuming there are infinite numbers of both loci and samples, then D^ can be estimated as previously shown, that is, D^=EDt. Moreover, according to Equation (6), EDt2v1v2N^b. Replacing Deq,t with 2v1Nb,tv2 and EDt with 2v1v2N^b in the approximated expressions of EDt,v=4 and EDt,v=6, we obtained

1N^b,v=423·1Nb,t1+29·1Nb,t2+227·1Nb,t3+,
1N^b,v=6610·1Nb,t1+24100·1Nb,t2+961000·1Nb,t3+. (10)

This suggests that when N^b is previously estimated, 1/N^b would be the weighted sum of 1/Nb in previous generations with the weights forming a geometric progression.

For finite samples, Nb,t in Equation (10) should be replaced by Np. Clearly, the coefficients are geometric sequences and the first term Nb,t1 are more influential for autotetraploids than for higher levels of ploidy. This also explains why D converges more slowly at higher ploidy levels when Nb>2 (Figure 1, Table 2). It is noteworthy that because we expanded the Taylor series for D1 at D0=0 and Deq,0=0, the above equations are applicable for a high value of Nb but not when Nb=2.

We next discuss the scenario for diploids, according to Equations (1) and (3):

ED1,v=2=1D0Deq,01Deq,0+D0Deq,0Deq,02.

Similarly, by expanding the second‐order Taylor series at D0=0 and Deq,0=0 and eliminating the remainder, we obtained

ED1,v=2Deq,0D0Deq,0+Deq,02.

The above equation can be expanded to EDt,v=2Deq,t1EDt1Deq,t1+Deq,t12. Replacing EDt1 with Deq,t2EDt2Deq,t2+Deq,t22 and eliminating the third‐order terms, we obtained:

4.3.

Replacing Deq,t with 2v1Nb,tv2 and EDt,v=2 with 2v1N^bv2,

1N^b,v=21Nb,t11+1Nb,t11Nb,t2. (11)

Equation (11) implies that N^b mainly reflects the value of Nb in the previous generation in diploids. This also explains why diploids converged fastest when Nb>2 (Table 2, Figure S7).

It can also be inferred from Equations (10) and (11) that values of N^b at lower‐levels of ploidy reflect more recent values of Nb, especially when v=2. In addition, because D converges very fast, N^b is less likely to be influenced by the mutations.

4.4. Bias of Pudovkin's (1996) Estimator

There are two potential causes of bias in Pudovkin et al. (1996) estimator under its core assumption (HWE0): (i) the biased estimation of D, and (ii) the use of EHE,tλtHE,0 to approximate expected heterozygosities.

Biased estimation of D is a more serious problem for smaller values of Nb, which can cause an underestimation of D (Figure 1) and overestimation of Nb under HWE0 (Table 1). However, we used Nb10 in Figure 2 so any effects of a biased estimation of D are negligible. Hence, the use of EHE,tλtHE,0 to approximate expected heterozygosities is the main reason causing the bias under HWE0, and we use Equation (1) to correct this bias and obtain an asymptoticly unbiased estimator N^b,0. By comparing the expressions presented in Table 1, it can be shown that in diploids N^b,PuN^b,0, while N^b,0 is approximately N^b,Pu plus 38 and 1018 in autotetraploids and autohexaploids, respectively, if they use the same D estimator. This explains the downward bias of N^b,Pu under HWE0 (Figure 2). When values for Nb are high, this bias still exists but is concealed by the higher values of Nb and increased variance of N^b,Pu.

Under HEE0, both N^b,Pu and N^b,0 are upwardly biased for diploids but downwardly biased for polyploids (Figure 2). These biases are caused by the difference between ED1|HWE0 and Deq (Figure 1). N^b,Pu and N^b,0 assume HWE0, under which the expected heterozygote‐excess index is ED1|HWE0. However, it is decreased (or increased) to Deq in diploids (or polyploids) under HEE0. Given the reciprocal relationship between N^b and D^, N^b,Pu and N^b,0 are upwardly (or downwardly) biased in diploids (or polyploids) under HEE0. Moreover, the degrees of absolute bias of N^b,0 and N^b,Pu under HEE0 also differ; this is lower in diploids than polyploids and is caused by the difference between Deq and D1, which is slight in diploids but clear in polyploids (Nb=10, Figure 1).

4.5. Bias and Variance of Corrected Estimators

For a single locus l, H^O,l and H~E,l are correlated. However, D^ is calculated by summing the numerators and denominators across loci: D^=lH^O,llH~E,l1 according to Equation (8). Because the loci are assumed to be unlinked, the summation reduces the correlations between the numerators and the denominators . Thus, D^ is an asymptotically unbiased estimator of D, and the bias can be reduced to a negligible level when L is high (e.g., 100). Besides, the inverse 1/D^ can be expressed as 1D^=lH~E,llH^O,lH~E,l. Similarly, the summation reduces the correlations between the numerator and the denominator, and 1/D^ is also an asymptoticly unbiased estimator of 1/D. This suggests that our N^b,0 and N^b, are asymptoticly unbiased, which is consistent with our simulations (Figure 3).

In polyploids, Nb can still be unbiased, as estimated by N^b,0 or N^b, under their core assumptions, while variance may be either increased or decreased with increased level of ploidy. Two factors cause this phenomenon: (i) under the same conditions, polyploids have more within and between individual allele pairs than diploids. The observed and expected heterozygosities are therefore estimated with increased accuracy, as well as estimates for D and 1/D, which reduces the variance of N^b; (ii) because N^ba/D^ (Table 1), the coefficient a is higher in polyploids and increases the variance of N^b.

For variance, the increase in accuracy and precision by sampling one more individual is approximately equal to using two additional loci (Figures S2 and S5). Due to the cost of sequencing one extra individual usually being higher than using two additional loci, especially in polyploids, using more loci is probably in most cases more practical because this is feasible due to the current capabilities of modern genomic technologies, which enable the simultaneous detection of millions of genetic loci. But the number of individuals and loci should be balanced based on the study's goals, population characteristics, and available resources. In medical and agricultural research, larger sample sizes are common, while conservation focused ecological studies often use fewer individuals. For example, in our simulations for diploids with Nb=10, we achieve a 10% CV with 50 individuals and 1262 loci, and the same with 1000 individuals and 333 loci, which corresponds to more individuals and fewer loci. However, increasing the number of individuals further does not reduce the number of loci needed to maintain the same CV (Figure 5 and Table S3). From the isometric curves, logLCV increases by approximately 2 as logNb increases by 1, such that the relationship between Nb and LCV can be roughly described by LCVNb2 when Nb>100.

4.6. Application Scope

The extended definitions of Ne and Nb to polyploid species, along with the corrected Pudovkin et al. (1996) Nb estimator under HWE0 and HEE0, provide new opportunities for accurate estimation of population parameters both in polyploid and diploid organisms.

Our corrected estimator N^b, is appropriate for DR mating system, and performs well for small to medium sized populations with Nb100 (Figures 3 and 4, Figures S1 and S4), which are in need of monitoring to prevent extirpation. Moreover, it can be easily adapted for use in both diploid and polyploid species, and also performs well in polyploids. Regarding marker types, our estimator is applicable to both single nucleotide polymorphisms (SNPs) and simple sequence repeats (SSRs). Compared with N^b,Pu and N^b,0, we recommend use of N^b, in practice because it assumes more realistic conditions.

While the true Nb of the empirical dataset remains unknown, our assessments show that N^b, offers greater robustness than N^b,Pu, with wider 95% CIs (Tables S4 and S5). For high‐risk decisions, this reliability is paramount, conservative estimation presents a much lower risk than the severe consequences of overestimation.

It is crucial to emphasise that our method relies on the assumption of heterozygote excess. However, in many populations relevant to conservation biology, factors such as inbreeding and population subdivisionwhich often lead to heterozygote deficiencyare frequently observed. When applied to such populations, our method may yield negative estimates, thereby limiting its applicability. In such cases, simply discarding negative values may introduce bias. To enhance the applicability of the method in complex practical scenarios, the following adjustment strategies may be considered: (i) for inbred populations, if high‐resolution genotyping techniques such as resequencing are used, individuals whose Runs of Homozygosity (ROH) cover a specific locus can be excluded when calculating heterozygosity for that locus, thereby reducing the influence of inbreeding on heterozygosity estimation; (ii) for populations with substructure, subpopulations can first be identified using methods such as principal coordinate analysis (PCoA), K‐means clustering, or hierarchical clustering. Nb can then be estimated separately for each genetically homogeneous subpopulation, or the harmonic mean of Nb across subpopulations can be calculated to more accurately reflect the reproductive characteristics of the population.

It should be noted that the specific implementation of the above strategies is beyond the scope of this study. This research develops bias‐corrected, robust Nb estimator and extends them to autopolyploids, validated through simulations. We advise users to confirm heterozygosity excess before applying the method and to preprocess data or interpret results with the above strategies in mind. These limitations also highlight avenues for future methodological and applied research.

5. Conclusion

We extended the definition of Ne and Nb to polyploids, corrected Pudovkin et al. (1996) Nb estimator under HWE0 and HEE0, and extended and corrected Pudovkin et al. (1996) estimator to account for polyploids. From computer simulations, we found that genotypes can rapidly reach HEE, our corrected estimators can unbiasedly estimate Nb, and our estimates for polyploids are slightly better than those for diploids. We also discussed the relationship between N^b and Nb, and N^b in variable Nb across generations.

Author Contributions

B.Y., K.H., and B.L. conceptualised, designed and performed the study; K.H. and B.Y. developed methods; B.Y. and K.H. wrote the original draft paper, B.Y., K.H., D.W., Y.S., W.L, H.W., R.W., R.P., S.Z, Z.X. and D.W.D. reviewed and edited the paper; B.L. supervised the process of writing; K.H. and B.L. funded the study.

Funding

This work was supported by the National Natural Science Foundation of China (32570568, 32170515, 32371563, 32070453, 12171391) and the National Key R&D Program of China (2024YFF1307302).

Disclosure

Benefit‐sharing section: This study contains no new empirical data. The Yellowstone grizzly bears dataset and the brook trout dataset are retrieved from Dryad with https://doi.org/10.5061/dryad.s6764 (Kamath et al. 2015) and https://doi.org/10.5061/dryad.n02v6wwv6 (Anne‐Laure et al. 2020), respectively. Benefits from this research accrue from the sharing of our software and source code on public databases as described above.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Data S1: Supporting Information.

MEN-26-e70107-s010.zip (32.6MB, zip)

File S1: S1.nb, derivation of the corrected estimator (*.nb is the mathematica v9.0 notebook file).

MEN-26-e70107-s003.nb (22.4KB, nb)

File S2: S2.nb, derivation of Pudovkin et al. (1996) estimator.

MEN-26-e70107-s006.nb (10.7KB, nb)

File S3: S3.7z, C++ source code of the simulation programme.

MEN-26-e70107-s005.7z (42.9KB, 7z)

File S4: S4.7z, Simulation programme (64‐bit Windows PE executable, compiled by Visual Studio 2019), parameters files (Sim1.txt to Sim5.txt) and batch files to perform the Monte‐Carlo simulations (Sim1.bat to Sim5.bat).

File S5: S5.7z, matlab v8.2 scripts to process the simulation results (Sim1.m to Sim6.m) and resulting figures in PDF format.

MEN-26-e70107-s008.7z (29.1MB, 7z)

File S6: S6.nb, derivation of moments of homozygosities.

MEN-26-e70107-s002.nb (35.1KB, nb)

File S7: S7.nb, derivation of bias and variance of N^b,.

MEN-26-e70107-s009.nb (100KB, nb)

File S8: S8.nb, calculation of the number of loci required to reach a CV.

MEN-26-e70107-s004.nb (129.3KB, nb)

File S9: S9.nb, derivation of the expectation of Np.

MEN-26-e70107-s007.nb (19.2KB, nb)

Yang, B. , Wang D., Shen Y., et al. 2026. “Estimating Effective Number of Breeders Under the Heterozygote‐Excess Equilibrium.” Molecular Ecology Resources 26, no. 2: e70107. 10.1111/1755-0998.70107.

Data Availability Statement

The derivations of original and corrected estimators, and the source code and binary executables of the simulation programme are available in the Supporting Information of this article. Our methods have been implemented in a software package, polygene v1.7, which is freely available at https://github.com/huangkang1987/polygene.

References

  1. Anne‐Laure, F. , Maeva L., Martin L., et al. 2020. “Individual Genotypes of 1416 Brook Trout Genotyped at 14779 High‐Quality SNPs for Studying Local Adaptation and Maladaptation in Small Populations [Data Set].” Dryad. 10.5061/dryad.n02v6wwv6. [DOI]
  2. Avni, R. , Nave M., Barad O., et al. 2017. “Wild Emmer Genome Architecture and Diversity Elucidate Wheat Evolution and Domestication.” Science 357, no. 6346: 93–97. [DOI] [PubMed] [Google Scholar]
  3. Burns, R. , Mandáková T., Gunis J., et al. 2021. “Gradual Evolution of Allopolyploidy in Arabidopsis suecica .” Nature Ecology & Evolution 5, no. 10: 1367–1381. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Burow, M. D. , Simpson C. E., Starr J. L., and Paterson A. H.. 2001. “Transmission Genetics of Chromatin From a Synthetic Amphidiploid to Cultivated Peanut (Arachis hypogaea L.): Broadening the Gene Pool of a Monophyletic Polyploid Species.” Genetics 159, no. 2: 823–837. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Charlesworth, B. , and Charlesworth D.. 2010. Elements of Evolutionary Genetics. Roberts and Company. [Google Scholar]
  6. Comai, L. 2005. “The Advantages and Disadvantages of Being Polyploid.” Nature Reviews Genetics 6, no. 11: 836–846. [DOI] [PubMed] [Google Scholar]
  7. Falconer, D. S. , and Mackay T. F. C.. 1996. Introdutction to Quantitative Genetics. Pearson Education. [Google Scholar]
  8. Hamilton, M. B. 2009. Population Genetics. John Wiley & Sons. [Google Scholar]
  9. Huang, K. , Dunn D. W., Li W., Wang D., and Li B.. 2022. “Linkage Disequilibrium Under Polysomic Inheritance.” Heredity 128, no. 1: 11–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Huang, K. , Dunn D. W., Ritland K., et al. 2019. “Polygene: Population Genetics Analyses for Autopolyploids Based on Allelic Phenotypes.” Methods in Ecology and Evolution 11, no. 3: 448–456. [Google Scholar]
  11. Kamath, P. L. , Haroldson M. A., Luikart G., Paetkau D., Whitman C., and van Manen F.. 2015. “Multiple Estimates of Effective Population Size for Monitoring a Long‐Lived Vertebrate: An Application to Yellowstone Grizzly Bears.” Molecular Ecology 24, no. 22: 5507–5521. [DOI] [PubMed] [Google Scholar]
  12. Kimura, M. 1954. “Process Leading to Quasi‐Fixation of Genes in Natural Populations due to Random Fluctuation of Selection Intensities.” Genetics 39, no. 3: 280–295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Lovell, J. T. , MacQueen A. H., Mamidi S., et al. 2021. “Genomic Mechanisms of Climate Adaptation in Polyploid Bioenergy Switchgrass.” Nature 590, no. 7846: 438–444. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Lynch, M. , Ackerman M. S., Gout J. F., et al. 2016. “Genetic Drift, Selection and the Evolution of the Mutation Rate.” Nature Reviews Genetics 17, no. 11: 704–714. [DOI] [PubMed] [Google Scholar]
  15. Nei, M. , and Tajima F.. 1981. “Genetic Drift and Estimation of Effective Population Size.” Genetics 98, no. 3: 625–640. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Nomura, T. 2008. “Estimation of Effective Number of Breeders From Molecular Coancestry of Single Cohort Sample.” Evolutionary Applications 1, no. 3: 462–474. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Otto, S. P. 2007. “The Evolutionary Consequences of Polyploidy.” Cell 131, no. 3: 452–462. [DOI] [PubMed] [Google Scholar]
  18. Pollak, E. , and Sabran M.. 1996. “On the Theory of Partially Inbreeding Finite Populations. IV. The Effective Population Size for Polyploids Reproducing by Partial Selfing.” Mathematical Biosciences 135, no. 1: 69–84. [DOI] [PubMed] [Google Scholar]
  19. Pudovkin, A. I. , Zaykin D. V., and Hedgecock D.. 1996. “On the Potential for Estimating the Effective Number of Breeders From Heterozygote‐Excess in Progeny.” Genetics 144, no. 1: 383–387. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Pudovkin, A. I. , Zhdanova O. L., and Hedgecock D.. 2009. “Sampling Properties of the Heterozygote‐Excess Estimator of the Effective Number of Breeders.” Conservation Genetics 11, no. 3: 759–771. [Google Scholar]
  21. Schwartz, M. K. , Luikart G., and Waples R. S.. 2007. “Genetic Monitoring as a Promising Tool for Conservation and Management.” Trends in Ecology & Evolution 22, no. 1: 25–33. [DOI] [PubMed] [Google Scholar]
  22. Selander, R. K. 1970. “Behavior and Genetic Variation in Natural Populations.” American Zoologist 10, no. 1: 53–66. [DOI] [PubMed] [Google Scholar]
  23. Selmecki, A. M. , Maruvka Y. E., Richmond P. A., et al. 2015. “Polyploidy Can Drive Rapid Adaptation in Yeast.” Nature 519, no. 7543: 349–352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Stuart, A. , and Ord J. K.. 1998. Kendall's Advanced Theory of Statistics: Distribution Theory. 6th ed. Oxford University Press. [Google Scholar]
  25. Tallmon, D. A. , Koyuk A., Luikart G., et al. 2008. “ONESAMP: A Program to Estimate Effective Population Size Using Approximate Bayesian Computation.” Molecular Ecololgy Resources 8, no. 2: 299–301. [DOI] [PubMed] [Google Scholar]
  26. Wang, J. 2009. “A New Method for Estimating Effective Population Sizes From a Single Sample of Multilocus Genotypes.” Molecular Ecology 18, no. 10: 2148–2164. [DOI] [PubMed] [Google Scholar]
  27. Wang, J. 2016. “A Comparison of Single‐Sample Estimators of Effective Population Sizes From Genetic Marker Data.” Molecular Ecology 25, no. 19: 4692–4711. [DOI] [PubMed] [Google Scholar]
  28. Waples, R. S. 1989. “A Generalized Approach for Estimating Effective Population Size From Temporal Changes in Allele Frequency.” Genetics 121, no. 2: 379–391. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Waples, R. S. 2005. “Genetic Estimates of Contemporary Effective Population Size: To What Time Periods Do the Estimates Apply?” Molecular Ecology 14, no. 11: 3335–3352. [DOI] [PubMed] [Google Scholar]
  30. Waples, R. S. 2006. “A Bias Correction for Estimates of Effective Population Size Based on Linkage Disequilibrium at Unlinked Gene Loci.” Conservation Genetics 7, no. 2: 167–184. [Google Scholar]
  31. Waples, R. S. 2024. “Practical Application of the Linkage Disequilibrium Method for Estimating Contemporary Effective Population Size: A Review.” Molecular Ecology Resources 24: e13879. [DOI] [PubMed] [Google Scholar]
  32. Waples, R. S. 2025. “The Idiot's Guide to Effective Population Size.” Molecular Ecology 34: e17670. 10.1111/mec.17670. [DOI] [PubMed] [Google Scholar]
  33. Waples, R. S. , and Do C.. 2008. “LDNE: A Program for Estimating Effective Population Size From Data on Linkage Disequilibrium.” Molecular Ecology Resources 8, no. 4: 753–756. [DOI] [PubMed] [Google Scholar]
  34. Waples, R. S. , and Do C.. 2010. “Linkage Disequilibrium Estimates of Contemporary N e Using Highly Variable Genetic Markers: A Largely Untapped Resource for Applied Conservation and Evolution.” Evolutionary Applications 3, no. 3: 244–262. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Weir, B. S. , and Hill W. G.. 1980. “Effect of Mating Structure on Variation in Linkage Disequilibrium.” Genetics 95, no. 2: 477–488. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data S1: Supporting Information.

MEN-26-e70107-s010.zip (32.6MB, zip)

File S1: S1.nb, derivation of the corrected estimator (*.nb is the mathematica v9.0 notebook file).

MEN-26-e70107-s003.nb (22.4KB, nb)

File S2: S2.nb, derivation of Pudovkin et al. (1996) estimator.

MEN-26-e70107-s006.nb (10.7KB, nb)

File S3: S3.7z, C++ source code of the simulation programme.

MEN-26-e70107-s005.7z (42.9KB, 7z)

File S4: S4.7z, Simulation programme (64‐bit Windows PE executable, compiled by Visual Studio 2019), parameters files (Sim1.txt to Sim5.txt) and batch files to perform the Monte‐Carlo simulations (Sim1.bat to Sim5.bat).

File S5: S5.7z, matlab v8.2 scripts to process the simulation results (Sim1.m to Sim6.m) and resulting figures in PDF format.

MEN-26-e70107-s008.7z (29.1MB, 7z)

File S6: S6.nb, derivation of moments of homozygosities.

MEN-26-e70107-s002.nb (35.1KB, nb)

File S7: S7.nb, derivation of bias and variance of N^b,.

MEN-26-e70107-s009.nb (100KB, nb)

File S8: S8.nb, calculation of the number of loci required to reach a CV.

MEN-26-e70107-s004.nb (129.3KB, nb)

File S9: S9.nb, derivation of the expectation of Np.

MEN-26-e70107-s007.nb (19.2KB, nb)

Data Availability Statement

The derivations of original and corrected estimators, and the source code and binary executables of the simulation programme are available in the Supporting Information of this article. Our methods have been implemented in a software package, polygene v1.7, which is freely available at https://github.com/huangkang1987/polygene.


Articles from Molecular Ecology Resources are provided here courtesy of Wiley

RESOURCES