Skip to main content
PLOS One logoLink to PLOS One
. 2021 Mar 25;16(3):e0248808. doi: 10.1371/journal.pone.0248808

Superspreading of SARS-CoV-2 in the USA

Calvin Pozderac 1,*, Brian Skinner 1
Editor: Yury E Khudyakov2
PMCID: PMC7993775  PMID: 33765004

Abstract

A number of epidemics, including the SARS-CoV-1 epidemic of 2002-2004, have been known to exhibit superspreading, in which a small fraction of infected individuals is responsible for the majority of new infections. The existence of superspreading implies a fat-tailed distribution of infectiousness (new secondary infections caused per day) among different individuals. Here, we present a simple method to estimate the variation in infectiousness by examining the variation in early-time growth rates of new cases among different subpopulations. We use this method to estimate the mean and variance in the infectiousness, β, for SARS-CoV-2 transmission during the early stages of the pandemic within the United States. We find that σβ/μβ ≳ 3.2, where μβ is the mean infectiousness and σβ its standard deviation, which implies pervasive superspreading. This result allows us to estimate that in the early stages of the pandemic in the USA, over 81% of new cases were a result of the top 10% of most infectious individuals.

Introduction

The temporal growth of an epidemic is often characterized by either a time scale (such as the doubling time) [1, 2] or by the reproduction rate R0, which indicates the average number of new infections produced by each infected individual [3]. Estimates of R0 for the current pandemic of SARS-CoV-2 range from 1.4 to 3.8 [4–7]. Neither of these numbers, however, gives any information about the distribution of infectiousness among individuals—i.e., whether new infections arise relatively uniformly from all infected individuals, or whether new infections are driven primarily by a small number of highly infectious individuals. The latter case is commonly referred to as “superspreading”, and different epidemics exhibit superspreading to different degrees. For example, during the outbreak of SARS CoV-1 in 2002-2004, over 80% of cases were observed to result from the top 20% most infectious individuals [8, 9]. Understanding the degree of superspreading in the current pandemic of SARS-CoV-2 is crucial for developing strategies to mitigate continued spread and informing an educated reopening procedure [10–13].

Here we present a simple and direct method to understand how the infectiousness (also called the “reproduction rate” of the disease) varies among infected individuals. At late times after the onset of an epidemic, the number of infected individuals is large, and consequently any statistical fluctuations in the growth rate are relatively small, so that the growth rate is well characterized by the mean infectiousness, μβ. However, at early times, when there are relatively few cases, the growth rate is stochastic and the degree of randomness depends on the variance in infectiousness, σβ2, between individuals (Fig 1a). By examining the variance in growth rate across subpopulations at these early times (Fig 1b), we are able to infer the variation in the distribution of infectiousness. In our analysis we divide the US cases into counties and observe how the variance in growth rate across them evolves as the number of cases increases.

Fig 1.

Fig 1

(a) Illustration of the variance in early-time growth rate of new cases. At early times, there is noticeable variance in the growth rate between counties. As the number of cases grows, all counties stabilize towards the average growth rate I ∼ (1+ μβ)t, (dashed black line) where t is the number of days since the first case in a county. The counties shown are Boulder, CO (blue), St. Mary, LA (purple), Vanderburgh, IN (red), Mesa, CO (orange), and Jones, GA (green). (b) The number of daily infections per infected individual as a function of total infections. In the main figure, each point corresponds to a given county (across all US counties that never report ΔI < 0) at a given time point (within the first 14 days after the first infection reported in that county). As the number of cases increase, all counties converge to the mean infection rate. The mean (points) and variance (bars) of ΔI/I at a given I are shown in the inset. The variance decreases like (μβ+σβ2)/I (black lines).

Formalizing this idea, we present a derivation of the variance in the exponential growth rate, or number of new cases per infected individual per day, ΔI/I, using an SIR framework that incorporates a probability distribution for the infectiousness of a given individual. Our result implies a simple method for estimating the mean, μβ, and variance, σβ2, of the infectiousness β. We apply this method to data for COVID-19 cases in the USA, and find a mean infection rate of μβ = 0.18 cases/day and standard deviation of σβ ≳ 0.59 cases/day. Since the standard deviation is considerably larger than the mean, with σβ/μβ ≳ 3.2, we conclude that superspreading is prevalent. By our estimate, these results imply that at least 81% of new cases are caused by the top 10% of most infectious individuals. Our method, which uses only a direct measurement of variance in detected case data in the USA, is consistent with estimates of superspreading using surveillance data [14], secondary-case data [15], and more complicated estimates of cluster size distribution using Markov Chain Monte Carlo [16].

Results

Variance in growth rate in the SIR model

We derive a relation between the variance in the case growth rate and the variance in individual infectiousness between individuals in the population. We start with a standard discrete-time SIR model [17], which is governed by the following difference equations:

ΔS=-βISNΔI=βISN-rIΔR=rI (1)

Here, N is the total population and S, I, and R are the time-dependent numbers of susceptible, infected, and recovered individuals, respectively. The parameters β and r encode the infectiousness and recovery rate of a disease within a population. The time is effectively discretized into days by the available data, so we use ΔI rather than the usual time derivative, dI/dt. The SIR description typically assumes fixed values for β and r across the population. However, in superspreading contexts there is a substantial variance in the infectiousness within a population [8, 9, 18, 19]. We account for this variation by introducing a probability distribution of infectiousness, p(β), so that the probability for a randomly-selected individual to have infectiousness in the range (β, β + dβ) is given by is given by p(β)dβ.

For an individual with a given infectiousness, β, the probability of infecting exactly n others in a day follows the Poisson distribution, Pois(n;β). The probability that a randomly selected individual will infect n others is given by combining the Poisson distribution with the distribution p(β), giving

P(n)=∫0∞dβe-ββnn!p(β). (2)

The first two moments of P(n), μn and σn2, can be calculated independent of the form of p(β):

μn=∑n=0∞nP(n)=μβ (3)
σn2=∑n=0∞(n-μn)2P(n)=μβ+σβ2 (4)

Eq (4) represents the variance, among all infected individuals, of the number of new infections caused by a single person in a given day. When there are I active cases, the mean number of new cases per infected person, Δ(I + R)/I, is given by the average of I random variables drawn from the distribution P(n). By the central limit theorem, it follows that Var(Δ(I+R)/I)=σn2/I. Additionally, in the SIR model with a finite total population N, Δ(I+ R)/I = βS/N = β(1 − (I + R)/N) decreases as the susceptible population continually shrinks. Effectively, p(β) is scaled by the factor (1 − (I + R)/N), which represents the fraction of the population that remains susceptible. Consequently, μβ → μβ(1 − (I + R)/N) and σβ2→σβ2(1-(I+R)/N)2. Therefore the total variance in Δ(I+ R)/I follows:

Var(Δ(I+R)I)=μβ(1-I+RN)+σβ2(1-I+RN)2I (5)

This result becomes simpler in the limiting case where there is no significant change in the susceptible population (N → ∞) and no recovery (R → 0). In this limit, we retrieve the case of simple exponential growth, for which [20]

Var(ΔII)=μβ+σβ2I. (6)

In the limit σβ → 0, where every infected individual has the same infectiousness μβ, the variance in the average infection rate is simply μβ/I, which corresponds to the variance in a Poisson process with rate μβ.

In the case of SARS-CoV-2, it is well established that there are asymptomatic carriers [21–23] who transmit the virus without being detected, as well as other infections that are undetected or unreported. Current estimates typically predict that only 10 − 25% [24–26] of cases are detected. One can attempt to address this effect by assuming that there is a fixed detection probability, pdet, and that the entire infected population, regardless of symptoms, follows the same infectiousness distribution p(β). In this case, there are many more infected individuals, I ∼ Idet/pdet, than those detected, which reduces the statistical fluctuations in the growth rate and makes our calculation of σβ2 a lower bound. The effect of undetected cases is considered in more detail in the S3 Appendix. In order to be conservative (especially given the possibility that asymptomatic cases have a lower rate of infection than symptomatic ones [27, 28]), the results we present here use pdet = 1.

Data for COVID-19 in the USA

We now turn our attention to data for total detected cases of COVID-19 in the USA, taken from the publicly available data set at Ref. [29]. In the following analysis we limit our consideration to only a short timescale (∼14 days) after the first infection is detected in a given county. This limitation in time scale serves three main purposes; first, it is likely that through changes in policy, lockdown, social distancing, mask usage, etc., the average infectiousness within the population is time-dependent. By restricting ourselves to a relatively small window of early times, we may assume that there is a constant average infectiousness. Second, considering only beginning stages allows us to neglect the possible saturation of the susceptible population, effectively allowing us to take the N → ∞ limit. Finally, the recovery period for COVID-19 ranges from 7-14 days [30, 31] and so by considering this two week period, we can treat our system as if there is limited recovery and R → 0. These restrictions allow us to treat the USA data using the exponential case, Eq (6).

In our analysis, the population is divided into geographic regions and the variance is calculated across different trajectories I(t). The US cases are divided by county. For each county, we calculate the average number of new cases per current case per day, ΔI/I, for the first 14 days after the first infection is detected in that county. The variance in ΔI/I is then calculated among all counties that have a given fixed value of I (we present data only for values of I that have at least 250 corresponding counties). As shown in Fig 2, the US data generally follows the predicted ∼1/I trend. An unbiased fit of the data gives Var(ΔI/I) ∝ I−0.74. From Eq (6), we calculate μβ+σβ2 by averaging Var(ΔI/I) × I, weighted by the number of instances at each I value. One might worry that the main source of variation comes from differing average growth rates, μβ, in various counties (e.g. rural vs. urban). However, we show in the S2 Appendix that variance in μβ across counties is too small to explain the large observed variance in ΔI/I.

Fig 2. As the number of infections I in a given county increases, the variance in exponential growth rate, Var(ΔI/I), decreases as (μβ+σβ2)/I.

Fig 2

Each data point at a given I is calculated by taking the sample variance in ΔI/I across all counties when they have I cases. We observe that the USA data (blue) is inconsistent with a model of uniform infectiousness, or σβ = 0 (dashed red line). A fit to the data (solid black line) implies a large variance in infectiousness, such that σβ/μβ ≳ 3.2.

We calculate μβ from the entire USA population by averaging all values of ΔI/I weighted by the current number of infections. Equivalently, we sum the number of cases caused each day and then divide by the sum of the number of cases across those days. This procedure gives the mean infectiousness, μβ, and thus from Eq (6) and the fitted slope in Fig 2, we can infer σβ2.

This calculation yields μβ = 0.18 cases/day and σβ = 0.59 cases/day. The small value of μβ2/σβ2=0.096, equivalent to the dispersion parameter [16, 32, 33], provides clear evidence for superspreading during early stages of the COVID-19 pandemic in the United States. (See S7 Appendix for discussion about defining the dispersion parameter in terms of the daily infection rate).

These results for μβ and σβ can be used to further quantify the extent of superspreading under the assumption that p(β) follows a gamma distribution (as in Ref. [18]). In the Methods section we present a derivation of the cumulative share of infections, Y, caused by the top X portion of most infectious cases. The corresponding “Lorenz curve” Y(X) is plotted in Fig 3. This result implies (using our relatively conservative estimate of σβ) that 81% of new infections are produced by the top 10% of most infectious individuals, while only about 4.5% of cases arise from the 80% of infected individuals with the lowest infection rates.

Fig 3. An estimated Lorenz curve for SARS-CoV-2 infections in the USA, which displays the percentage of new cases that are caused by a given cumulative percentage of most infectious individuals (solid black).

Fig 3

A few points in the curve are highlighted (dashed grey lines): 61.7%, 81.4%, and 95.5% of new cases are caused by the top 5%, 10%, and 20% infectious cases, respectively. Accounting for undetected and asymptomatic cases would apparently make this curve steeper, corresponding to more severe superspreading.

Discussion

As we have shown, a wide distribution p(β) in infectiousness β leads to large statistical variation in the early-time growth rate of a disease. By calculating the variance in growth rate among different subpopulations one can infer the variance in p(β). Our result for COVID-19 cases in the USA suggests that σβ/μβ ≳ 3.2, implying a relatively severe superspreading. If we further assume that p(β) follows a gamma distribution (as in Ref. [18]), then we can produce a more direct estimate of the extent of superspreading (Fig 3). Our relatively simple and direct method, based on a calculation of variance in reported case data, can be contrasted with more complicated methods for inferring the dispersion parameter that are based on maximum likelihood estimation (e.g., Ref. [33] develops such a method using simulated data), cluster size distributions [16, 34], and surveillance or tracing data [14, 15]. These methods also tend to yield a lower-bound estimate for σβ/μβ. While studies based on testing and contact tracing (e.g., Refs. [18, 35–37]) remain the definitive method for assessing superspreading, the method we present here may provide a much simpler way of estimating its prevalence across a much larger population.

We emphasize that our analysis is unable to determine whether this large variance is a result of differing biological symptoms, social behavior, or other possible explanations. Additionally, this estimation is carried out for early times to minimize effects from a time varying p(β) and therefore predominantly speaks to the infectiousness prior to widespread lockdown measures.

We close by commenting on a number of complicating factors that we did not include in our analysis and which, one might suspect, could alter our primary finding of a large value of σβ/μβ. For example, we have assumed a uniform value of μβ across different geographic locations; we have neglected undetected cases; we have ignored the possible variation in detection rate pdet among different counties; we have effectively treated each county as an isolated population and have neglected cross-county interactions; and we have ignored the effects of the incubation period as well as the potential variation in incubation periods between individuals. In the Supplemental Information, we consider each of these mechanisms in turn and show that none of them can explain our result, so that our conclusion of prevalent superspreading of SARS-CoV-2 in the USA remains robust. In brief: the variation in μβ among different geographic locations is too small to explain the observed variance in growth rate [S2 Appendix]; neglecting undetected cases leads to an underestimate of the variance σβ2, so that our result is effectively a lower bound for the prevalence of superspreading [S3 Appendix]; variation in pdet between counties does not directly affect the variance in the growth rate (ΔIdet)/Idet, other than to provide an average of pdet < 1, which results in a lower-bound estimate of σβ2 [S4 Appendix]; cross county interactions tend to reduce the variance, so our result cannot be explained as a consequence of such interactions [S5 Appendix]; and variations in incubation period can only reduce the apparent variance in growth rate [S6 Appendix].

Methods

Data source

We use publicly available data taken from the data set provided by the Center for Systems Science and Engineering (CSSE) at Johns Hopkins University [29] to estimate μβ. Knowing μβ enables us to determine σβ by taking a best fit to Eq (6). Counties that recorded ΔI < 0 at any point are discarded from the analysis due to the potential for recording error; such counties comprise ∼20% of all counties.

Numerical simulation

We corroborate Eqs (5) and (6) using a numerical simulation of the trajectories of infection growth, I(t), for a given distribution p(β). Reference 18 has suggested that infectiousness follows a gamma distribution, and consequently, P(n)=NB(n;μβ2/σβ2,μβ/(μβ+σβ2)) where NB is the negative binomial distribution [10, 16]. Using this assumption, we simulate the growth of the epidemic by assuming that a given individual i, with infectiousness βi that is drawn randomly from p(β), generates a number ni of new cases each subsequent day that is drawn from Pois(ni;βi). The simulation results confirm Eqs (5) and (6), as shown in S1 Appendix. Numerical simulations were performed using Python; the primary analysis is publicly available [38] and the simulations are available upon request to the corresponding author.

Derivation of the curve Y(X)

Following Ref. [18], we assume that the distribution of infectiousness, p(β), follows a gamma distribution. This assumption also allows us to further quantify the degree of superspreading by deriving a mathematical relation for the curve Y(X), where Y represents the proportion of infections produced by the top X fraction of most infectious individuals. In particular, one can calculate the fraction of individuals Xβ0 with infectiousness larger than a given value β0, as well as the fraction of secondary infections Yβ0 that these individuals are expected to cause:

Xβ0=∫β0∞dβp(β)=Q(μβ2σβ2,β0μβσβ2) (7)
Yβ0=∫β0∞dβp(β)βμβ=Q(1+μβ2σβ2,β0μβσβ2), (8)

where Q is the Regularized Gamma function. By eliminating β0 we find

Y=Q(1+μβ2σβ2,Q-1(μβ2σβ2,X)). (9)

Fig 3 displays the cumulative share of infections, Y, caused by the top X portion of most infectious cases.

Supporting information

S1 Appendix. Simulations.

(PDF)

S2 Appendix. Variance in μβ.

(PDF)

S3 Appendix. Undetected cases.

(PDF)

S4 Appendix. Variance in testing.

(PDF)

S5 Appendix. Cross-county interactions.

(PDF)

S6 Appendix. Variance in incubation period.

(PDF)

S7 Appendix. Dispersion parameter comparison.

(PDF)

Acknowledgments

The authors are grateful to N. E. Skinner for helpful conversations.

Data Availability

We use publicly available data taken from the data set provided by the Center for Systems Science and Engineering (CSSE) at Johns Hopkins University. The paper describing the data can be found at: https://doi.org/10.1016/S1473-3099(20)30120-1. The data can be found at https://github.com/CSSEGISandData/COVID-19/tree/master/csse_covid_19_data/csse_covid_19_time_series under the filename time_series_covid19_confirmed_US.csv.

Funding Statement

The author(s) received no specific funding for this work.

References

  • 1. Muniz-Rodriguez K, Chowell G, Cheung CH, Jia D, Lai PY, Lee Y, et al. Doubling Time of the COVID-19 Epidemic by Province, China. Emerging Infectious Diseases. 2020;26(8):1912–1914. 10.3201/eid2608.200219 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Zhou L, Liu JM, Dong XP, McGoogan JM, Wu ZY. COVID-19 seeding time and doubling time model: an early epidemic risk assessment tool. Infectious Diseases of Poverty. 2020;9(76). 10.1186/s40249-020-00685-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Murray JD. Mathematical Biology. Springer-Verlag; 2003. [Google Scholar]
  • 4. Li Q, Guan X, Wu P, Wang X, et al. Early Transmission Dynamics in Wuhan, China, of Novel Coronavirus–Infected Pneumonia. New England Journal of Medicine. 2020;382(13):1199–1207. 10.1056/NEJMoa2001316 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Riou J, Althaus CL. Pattern of early human-to-human transmission of Wuhan 2019 novel coronavirus (2019-nCoV), December 2019 to January 2020. Eurosurveillance. 2020;25(4). 10.2807/1560-7917.ES.2020.25.4.2000058 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Sanche S, Lin YT, Xu C, Romero-Severson E, Hengartner N, Ke R. High Contagiousness and Rapid Spread of Severe Acute Respiratory Syndrome Coronavirus 2. 2020;26(7). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Liu Y, Gayle AA, Wilder-Smith A, Rocklöv J. The reproductive number of COVID-19 is higher compared to SARS coronavirus. Journal of Travel Medicine. 2020;27(2). 10.1093/jtm/taaa021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Galvani AP, May RM. Dimensions of superspreading. Nature. 2005;438:293–295. 10.1038/438293a [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Stein RA. Super-spreaders in infectious diseases. International Journal of Infectious Diseases. 2011;15(8):e510–e513. 10.1016/j.ijid.2010.06.020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Althouse BM, Wenger EA, Miller JC, Scarpino SV, Allard A, Hébert-Dufresne L, et al. Stochasticity and heterogeneity in the transmission dynamics of SARS-CoV-2; 2020. [DOI] [PMC free article] [PubMed]
  • 11. O’Donoghue AL, Dechen T, Pavlova W, Boals M, Moussa G, Madan M, et al. Super-Spreader Businesses and Risk of COVID-19 Transmission. medRxiv. 2020; 10.1101/2020.05.24.20112110 [DOI] [Google Scholar]
  • 12. Vespignani A, Tian H, Dye C, Lloyd-Smith JO, Eggo RM, Shrestha M, et al. Modelling COVID-19. Nature Reviews Physics. 2020; p. 1–3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Weiner Z, Wong G, Elbanna A, Tkachenko A, Maslov S, Goldenfeld N. Projections and early-warning signals of a second wave of the COVID-19 epidemic in Illinois. medRxiv. 2020. [Google Scholar]
  • 14. Lau MS, Grenfell B, Nelson K, Lopman B. Characterizing super-spreading events and age-specific infectivity of COVID-19 transmission in Georgia, USA. medRxiv. 2020; 10.1101/2020.06.20.20130476 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Hasan A, Susanto H, Kasim M, Nuraini N, Triany D, Lestari B. Superspreading in Early Transmissions of COVID-19 in Indonesia. medRxiv. 2020. 10.1101/2020.06.30.20143560 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Endo A, Abbott S, Kucharski AJ, Funk S. Estimating the overdispersion in COVID-19 transmission using outbreak sizes outside China. Wellcome Open Research. 2020;5:67. 10.12688/wellcomeopenres.15842.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Kermack WO, McKendrick AG, Walker GT. A contribution to the mathematical theory of epidemics. Proceedings of the Royal Society of London Series A, Containing Papers of a Mathematical and Physical Character. 1927;115(772):700–721. 10.1098/rspa.1927.0118 [DOI] [Google Scholar]
  • 18. Lloyd-Smith JO, Schreiber SJ, Kopp PE, Getz WM. Superspreading and the effect of individual variation on disease emergence. Nature. 2005;438:355–359. 10.1038/nature04153 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Sneppen K, Taylor RJ, Simonsen L. Impact of Superspreaders on dissemination and mitigation of COVID-19. medRxiv. 2020; 10.1101/2020.05.17.20104745 [DOI] [Google Scholar]
  • 20. Bliss CI, Fisher RA. Fitting the Negative Binomial Distribution to Biological Data. Biometrics. 1953;9(2):182. 10.2307/3001850 [DOI] [Google Scholar]
  • 21. Mahajan A, Solanki R, Sivadas N. Estimation of Undetected Symptomatic and Asymptomatic cases of COVID-19 Infection and prediction of its spread in USA. medRxiv. 2020; 10.1101/2020.06.21.20136580 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Mizumoto K, Kagaya K, Zarebski A, Chowell G. Estimating the asymptomatic proportion of coronavirus disease 2019 (COVID-19) cases on board the Diamond Princess cruise ship, Yokohama, Japan, 2020. Euro Surveillance. 2020;25(10). 10.2807/1560-7917.ES.2020.25.10.2000180 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Nishiura H, Kobayashi T, et al. Estimation of the asymptomatic ratio of novel coronavirus infections (COVID-19). Int J Infect Dis. 2020;94:154–155. 10.1016/j.ijid.2020.03.020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Pedersen M, Meneghini M. Quantifying undetected COVID-19 cases and effects of containment measures in Italy: Predicting phase 2 dynamics. 2020; 10.13140/RG.2.2.11753.85600 [DOI]
  • 25. Li R, Pei S, Chen B, Song Y, Zhang T, Yang W, et al. Substantial undocumented infection facilitates the rapid dissemination of novel coronavirus (SARS-CoV-2). Science. 2020;368(6490):489–493. 10.1126/science.abb3221 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Lu FS, Nguyen AT, Link NB, Lipsitch M, Santillana M. Estimating the Early Outbreak Cumulative Incidence of COVID-19 in the United States: Three Complementary Approaches. medRxiv. 2020; 10.1101/2020.04.18.20070821 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Wang Y, Tong J, Qin Y, Xie T, Li J, vi J, et al. Characterization of an asymptomatic cohort of SARS-COV-2 infected individuals outside of Wuhan, China. Clinical Infectious Diseases; 10.1093/cid/ciaa629 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Chu DK, Akl EA, Duda S, Solo K, Yaacoub S, Schünemann HJ, et al. Physical distancing, face masks, and eye protection to prevent person-to-person transmission of SARS-CoV-2 and COVID-19: a systematic review and meta-analysis. The Lancet. 2020;395:1950–1951. 10.1016/S0140-6736(20)31142-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Dong E, Du H, Gardner L. An interactive web-based dashboard to track COVID-19 in real time. The Lancet Infectious diseases. 2020;20(5):533–534. 10.1016/S1473-3099(20)30120-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Roman Wölfel VMC, Guggemos W, Seilmaier M, Zange S, Müller MA, Niemeyer D, et al. Virological assessment of hospitalized patients with COVID-2019. Nature. 2020;581:465–469. 10.1038/s41586-020-2196-x [DOI] [PubMed] [Google Scholar]
  • 31. Bar-On YM, Flamholz A, Phillips R, Milo R. Science Forum: SARS-CoV-2 (COVID-19) by the numbers. Elife. 2020;9:e57309. 10.7554/eLife.57309 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Hébert-Dufresne L, Althouse BM, Scarpino SV, Allard A. Beyond R0: Heterogeneity in secondary infections and probabilistic epidemic forecasting. medRxiv. 2020; 10.1101/2020.02.10.20021725 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Lloyd-Smith JO. Maximum Likelihood Estimation of the Negative Binomial Dispersion Parameter for Highly Overdispersed Data, with Applications to Infectious Diseases. PLOS ONE. 2007;2(2):1–8. 10.1371/journal.pone.0000180 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Kucharski AJ, Althaus CL. The role of superspreading in Middle East respiratory syndrome coronavirus (MERS-CoV) transmission. Eurosurveillance. 2015;20(25). 10.2807/1560-7917.ES2015.20.25.21167 [DOI] [PubMed] [Google Scholar]
  • 35. Althaus CL. Ebola superspreading. The Lancet Infectious Diseases. 2015;15(5):507–508. 10.1016/S1473-3099(15)70135-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Melsew YA, Gambhir M, Cheng AC, McBryde ES, Denholm JT, Tay EL, et al. The role of super-spreading events in Mycobacterium tuberculosis transmission: evidence from contact tracing. BMC Infectious Diseases. 2019;19(1):244. 10.1186/s12879-019-3870-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Adegboye OA, Elfaki F. Network analysis of mers coronavirus within households, communities, and hospitals to identify most centralized and super-spreading in the arabian peninsula, 2012 to 2016. Canadian Journal of Infectious Diseases and Medical Microbiology. 2018;2018. 10.1155/2018/6725284 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Pozderac C. Python code used for analysis and figures; https://github.com/calvinpozderac/COVID-19-Superspreading

Decision Letter 0

Yury E Khudyakov

27 Jan 2021

PONE-D-20-30570

Superspreading of SARS-CoV-2 in the USA

PLOS ONE

Dear Dr. POZDERAC,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

==============================

Your manuscript was reviewed by 2 experts in the field. Both reviewers identified many important issues in your submission which require your careful attention. Please review the attached comments and provide point-by-point responses.

==============================

Please submit your revised manuscript by Mar 13 2021 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols

We look forward to receiving your revised manuscript.

Kind regards,

Yury E Khudyakov, PhD

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Partly

Reviewer #2: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: No

Reviewer #2: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: In this study, Calvin and colleagues introduced a novel method to determine the variation of infectiousness among populations based on SIR model. And then authors used this method to illustrate there was significant variations of infectiousness among populations in USA known as superspreading events. The study is interesting and the result (superspreading events in USA) will also contribute to tailor the prevention and control policies for COVID-19 epidemic in USA. However, I have several concerns about this study.

The COVID-19 has highly varied incubation period (with mean of 5.2 days, and the distribution of incubation period was estimated as 12.5 days). The highly variable incubation time would result in surge for the number of confirmed cases in some specific times. As the method only used the number of confirmed case each day, it is important to know how the highly variable incubation time affect the result? As a novel method, it should be first tested in pervious data (such as SARS-CoV, Ebola, ZIKA etc) to illustrate the accuracy of the method. Then the result for COVID-19 would be more convincing.

Reviewer #2: In this article, the authors derive a relationship between the variance in epidemic growth rates and current size of the epidemic, formalizing the intuition that super spreading has the greatest impact on epidemic dynamics when the number of infected individuals is small. They then apply this relationship to COVID-19 case data from the U.S. to estimate superspreading of SARS-CoV-2.

Major comments

1. The authors compare their estimate of $\\mu_k^2/\\sigma_k^2$ base on mean and standard deviation of daily transmission rate to the dispersion parameters presented in the literature which are instead based on duration of infection transmission rate. These numerical estimates therefore seem incomparable without further transformation. For example, if the daily transmission rate was Gamma distributed with mean $\\mu_k$ and variance $\\sigma_k^2$ and each individual’s daily transmission rate was a independent draw from this distribution, I believe their duration of infection transmission rate (based on the manuscript’s assumption of a 14 day duration of infection) would have a distribution with the ratio of mean^2/variance being $14 \\mu_k^2/\\sigma_k^2$.

2. The authors point out that their estimates provide a lower bound on $sigma_k$ use this to identify prevalent superspreading throughout the USA. Given that superspreading of SARS-CoV-2 is generally accepted and the effort taken by the authors to quantify the effects of different assumptions, it might be more useful to use this methods of measuring superspreading to identify a range of possible values of $\\sigma_k$ based on plausible assumptions, that is, identify both lower and upper bounds on $\\sigma_k$, and compare this to the existing estimates in the literature.

3. The authors address many potential complications that could impact their findings, including varying testing rates in different counties (Appendix S4). A further complication would be differentially increasing testing rates, where some counties ramped up testing faster than others. Similarly, the estimates might be impacted by day of week patterns in reported cases, where weekly cases often show cyclical patterns indicating that testing and reporting of confirmed cases is variable based on the day of the week. While I doubt either of these would substantially affect the major findings of the paper, some discussion of the potential impact of this would be useful.

Minor comments

1. The choice of k to denote the transmission rate seem to be a likely source of confusion as k is frequently used as the dispersion parameter in literature on superspreading (e.g., see refs 10, 16, 18 from the manuscript).

2. Figure 1: Further details should be provided as to the set of counties and time periods included in part B of the figure.

3. Line 94: The authors state that the recovery period for COVID-19 is 14 days, but other literature suggest that viral shedding may decline and seroconversion may occur 7 days after symptom onset (e.g. Wölfel et al. 2020 Virological assessment of hospitalized patients with COVID-2019. Nature. https://doi.org/10.1038/s41586-020-2196-x). While this is negligible in the scheme of affecting the portion of the population susceptible, it nevertheless would be more accurate to provide a range here.

4. Figure 2: The caption for could be made clearer by indicating that the blue points are bucketed data where buckets are determined by confirmed cases. Additionally, clarification should be made as to how the number of confirmed cases was determined (at the end of the 14 day period?). Finally, the y-axis in Figure 2 needs additional labels to be informative.

5. I would strongly encourage the authors to make the code publicly available.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

Decision Letter 1

Yury E Khudyakov

8 Mar 2021

Superspreading of SARS-CoV-2 in the USA

PONE-D-20-30570R1

Dear Dr. POZDERAC,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Yury E Khudyakov, PhD

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Reviewers' comments:

Acceptance letter

Yury E Khudyakov

10 Mar 2021

PONE-D-20-30570R1

Superspreading of SARS-CoV-2 in the USA

Dear Dr. Pozderac:

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department.

If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org.

If we can help with anything else, please email us at plosone@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Yury E Khudyakov

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Appendix. Simulations.

    (PDF)

    S2 Appendix. Variance in μβ.

    (PDF)

    S3 Appendix. Undetected cases.

    (PDF)

    S4 Appendix. Variance in testing.

    (PDF)

    S5 Appendix. Cross-county interactions.

    (PDF)

    S6 Appendix. Variance in incubation period.

    (PDF)

    S7 Appendix. Dispersion parameter comparison.

    (PDF)

    Attachment

    Submitted filename: Response to Reviewers.pdf

    Data Availability Statement

    We use publicly available data taken from the data set provided by the Center for Systems Science and Engineering (CSSE) at Johns Hopkins University. The paper describing the data can be found at: https://doi.org/10.1016/S1473-3099(20)30120-1. The data can be found at https://github.com/CSSEGISandData/COVID-19/tree/master/csse_covid_19_data/csse_covid_19_time_series under the filename time_series_covid19_confirmed_US.csv.


    Articles from PLoS ONE are provided here courtesy of PLOS

    RESOURCES