Abstract
Background
As Nigeria has the sixth-highest population in the world and a significant amount of inbound and outbound travel, the characterization of SARS-CoV-2 genomic diversity across the country is critical for understanding novel pandemic dynamics. We describe the genomic diversity of SARS-CoV-2 in Nigeria throughout the COVID-19 pandemic and examine the coverage of Nigeria's genomic surveillance system.
Methods
Genome sequences and sample metadata were downloaded from the GISAID repository. A beta regression was used to test for a relationship between fully resolved nucleotide proportion over time, as a proxy for data quality. Sample and sequencing source were compared to assess geographic coverage.
Results
A total of 7759 COVID-19 sequences collected from February 2020 to March 2023 were included. The majority were collected in 2021 (76.6%) and South West (43%). Eleven states (30%) reported 10 or fewer SARS-CoV-2 genomes across the entire period. The genome sequences submitted to GISAID from Nigeria were of high quality with very few unresolved nucleotides. Waves 4 and 5, predominantly Omicron lineages, show higher diversity around position 23 kb than the other waves. Overall, the Nigeria Centre for Disease Control (NCDC) and state-run hospitals were the largest contributors to the sample collection efforts during this study period. However, the collection efforts shifted over time from NCDC in waves 1–3 to regional hospitals and other healthcare facilities in waves 4–5, although this pattern varied by geopolitical zone (GPZ). Sequencing efforts also shifted from research laboratories during the first waves to NCDC during waves 4 and 5.
Conclusions
The findings suggest the need for a coordinated sequencing strategy and standardized protocols to improve genomic surveillance during future outbreaks of existing and novel pathogens. A network of sequencing laboratories that includes at least one in each GPZ, linked to and coordinated by the national reference laboratory at NCDC might provide more balanced coverage for future pandemics and pathogen surveillance.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12864-025-12058-y.
Keywords: SARS-CoV-2, COVID-19, Genomic surveillance, Genomic diversity, Nigeria
Introduction
By January 28, 2024, more than 774 million Coronavirus disease 2019 (COVID-19) cases had been reported worldwide. However, only 9.6 million (1.2%) of these cases were reported in Africa [1]. Nigeria, the most populous country in Africa with an estimated 213 million people, contributed only 267 thousand cases (2.7%) to the cases reported by the continent by February 3, 2024 [2]. These numbers are considered an underestimation, especially when compared to published Nigerian COVID-19 seroprevalence survey results, which demonstrated that as much as 25% of the population had been infected with SARS-CoV-2 by October 2020 [3] and up to 79% had been infected by August 2021 in specific states across Nigeria that completed the surveys [4].
In response to the pandemic, the Government of Nigeria (GON) implemented a series of public health measures, including a lockdown on March 30, 2020, which was initially implemented only for high-burden and high-risk areas, including Lagos and Ogun states, and Federal Capital Territory (FCT). The lockdown expanded to other states in April 2020 and began to phase out in May 2020. The lockdown included school and workplace closures, bans on gatherings, curfews, and travel bans [5, 6]. In Nigeria, mask use became mandatory in early May 2020 [6, 7], and vaccine rollout among healthcare workers started in March 2021 and then expanded to the general eligible population in the last quarter of 2021 [8].
Like most countries in Africa [9, 10], the GON incorporated genomic surveillance as part of the country’s response, making Nigeria the first country in Africa to provide a complete SARS-CoV-2 sequence [11]. Genomic surveillance is essential to characterize transmission dynamics, circulation of the virus, and emergence of variants in the population, informing infection control measures [12]. In addition to being the most populous country in Africa, Nigeria has the sixth-highest population in the world and a significant amount of inbound and outbound travel [13–16]. Characterization of the genomic diversity of SARS-CoV-2 in areas of the world with a substantial population, population mixing, and mobility, such as Nigeria, is critical for understanding novel pandemic dynamics [12, 16, 17]. SARS-CoV-2 genomic literature from Nigeria is limited, and thus far, has primarily focused on the first waves of the pandemic [15, 18, 19] or specific states [16, 20]. This manuscript aims to describe the genomic diversity of SARS-CoV-2 in Nigeria throughout the COVID-19 pandemic and explore coverage of its genomic surveillance system during the pandemic response.
Methods
Genomic data
Data retrieval
Genome sequences and sample metadata were downloaded from the SARS-CoV-2 section of the GISAID (The Global Initiative on Sharing All Influenza Data) repository on March 29th, 2023 [21]. All SARS-CoV-2 sequences submitted to GISAID from Nigeria before the download date were downloaded for this analysis. Due to the size of the dataset [7, 993 viral genomes in total], the data had to be retrieved in two batches. The unique identifiers of each viral genome (“gisaid_epi_isl”) were downloaded for all viral genomes, then split into two smaller batches for full record retrieval, and accessible at 10.55876/gis8.240507dh. Data were stored on a private portion of the University of Minnesota computer cluster to comply with the GISAID data access agreement.
Metadata standardization and cleaning
Sample metadata from GISAID was manually edited to standardize the recorded names for the state of origin, sample collection facility, and sample sequencing facility. Samples with incomplete collection dates (year only, or month and year only) or invalid collection dates (before 1 October 2019 or after 29 March 2023) were excluded from the analysis. This date range was chosen to exclude samples that were labeled as being collected before the start of the viral pandemic [22] or after the date when the data were downloaded from GISAID. To improve the power of comparative analyses, samples were grouped into Nigeria’s national geopolitical zones (GPZs) based on the state of origin. Nigeria is divided into six GPZs: North Central (Benue, Federal Capital Territory, Kwara, Nasawara, and Plateau), North East (Adamawa, Bauchi, Borno, and Gombe), North West (Kaduna, Kano, Sokoto, and Zamfara), South East (Abia, Ebonyi, Enugu, and Imo), South South (Akwa Ibom, Delta, and Edo), and South West (Ekiti, Lagos, Ogun, Ondo, Osun, and Oyo).
Genome sequence processing
The genomes of the sequenced viral samples were retrieved in FASTA format. All sequences included in the analysis were aligned against the Wuhan-Hu-1 reference genome (Genbank accession NC_045512.2) with MAFFT version 7.475 [22]. The options used for alignment were “–keeplength –6merpair –addfragments –thread 16.” These parameters were chosen to make the alignment computationally tractable and preserve the coordinates of the sequences relative to the annotated reference genome.
The aligned genomes were separated by pandemic wave using the sample collection date. For this purpose, we estimated the date ranges of each wave using Nigeria’s weekly COVID-19 data reported by the World Health Organization [23]. Five specific waves were defined as followed: wave 1 (February 27, 2020 and October 25, 2020), wave 2 (October 26, 2020 and June 20, 2021), wave 3 (June 21, 2021 and November 21, 2021), wave 4 (November 22, 2021 and May 1, 2022), and wave 5 (May 2, 2022 to March 29, 2023).
Aligned genome sequences were masked to estimate pairwise diversity statistics and counts of mutations. The first 250nt and final 250nt of the genome were excluded due to high missing data fraction, likely from unreliable amplification of the ends of the linear genome. Only fully resolved nucleotides (A, C, G, or T) were used to calculate diversity and mutation frequencies; ambiguous nucleotides were excluded from these analyses. Ambiguous nucleotides were instead used to estimate the sequencing quality over time (see below). Masked regions were not used for any analyses. Average pairwise diversity was calculated on a per-nucleotide basis as the total number of pairwise differences in an alignment column divided by the number of pairwise comparisons considered. This is a measure of variability in a group of nucleotide sequences. We used this metric for identifying regions of the genome that were highly variable in temporal or geographic subsets of the genomes. Gaps were treated as missing data and not considered for the pairwise diversity calculations. Total mutation count was calculated as the number of fully resolved differences relative to the Wuhan-Hu-1 reference strain. This measure reports the total genetic divergence from the progenitor strain for a given genome sequence.
Beta regression was used to test for a relationship between fully resolved nucleotide proportion over time, as a proxy for data quality over time. The proportions of fully resolved nucleotides per-genome sequence were transformed to eliminate the limits of 0 and 1 using the following transformation:
![]() |
where y is the proportion of fully resolved nucleotides and n is the sample size. The beta regression was then fitted with Y as the response variable and the number of days since 2019–11-01 as the independent variable. The regression was implemented with the “betareg” package in the R statistical computing environment.
Metadata analyses
Frequencies and percentages were utilized to describe the demographic characteristics of the included sequences.
Additional data
COVID-19 reported new cases from Nigeria were extracted from the Our World in Data COVID-19 dataset [24]. Census 2016 Nigeria population distribution was provided by the National Bureau of Statistics [25].
Results
Demographic characteristics of sequenced samples
In this analysis, we included a total of 7,759 COVID-19 sequences collected from February 2020 to March 2023. Sixty percent of the sequences belonged to individuals aged 20–59 years old, and 48% were collected from male individuals. The majority of the sequenced samples were collected in 2021 (76.6%), and in South West GPZ (42.7%), followed by 27.2% in North-Central GPZ and 23.6% in South-South GPZ (Table 1, Fig. 1). Eleven states (30%) reported 10 or fewer SARS-CoV-2 genomes during the entire period. The main laboratories submitting sequences to GISAID during the pandemic in Nigeria were the African Centre of Excellence for Genomics of Infectious Diseases (ACEGID) (43.7%), the Nigeria Centre for Disease Control (NCDC) (31.6%), and the Nigerian Institute of Medical Research (NIMR) (6.3%).
Table 1.
Demographic characteristics of included sequenced SARS-CoV-2 samples, Nigeria from February 2020 to March 2023
| Characteristic | n (%) N = 7759 |
|---|---|
| Age | |
| 0–9 | 155 (2.0) |
| 10–19 | 352 (4.5) |
| 20–29 | 1220 (15.7) |
| 30–39 | 1460 (18.8) |
| 40–49 | 1144 (14.7) |
| 50–59 | 816 (10.5) |
| 60–69 | 455 (5.9) |
| 70–79 | 230 (3.0) |
| 80 and more | 115 (1.5) |
| Unknown | 1812 (23.4) |
| Sex | |
| Female | 2781 (35.8) |
| Male | 3711 (47.8) |
| Unknown | 1267 (16.3) |
| Geopolitical zones | |
| North-Central | 2108 (27.2) |
| North-East | 124 (1.6) |
| North-West | 278 (3.6) |
| South-East | 98 (1.2) |
| South-South | 1832 (23.6) |
| South-West | 3311 (42.7) |
| Unknown | 8 (0.1) |
| Submitting Sequencing Lab | |
| ACEGID | 3393 (43.7) |
| NCDC | 2449 (31.6) |
| NIMR | 489 (6.3) |
| Others | 1428 (18.4) |
| Year of sample collection | |
| 2020 | 455 (5.8) |
| 2021 | 5941 (76.6) |
| 2022 | 1346 (17.4) |
| 2023* | 17 (0.2) |
*March 29, 2023
African Centre of Excellence for Genomics of Infectious Diseases (ACEGID)
Nigeria Centre for Disease Control (NCDC)
Nigerian Institute of Medical Research (NIMR)
Fig. 1.
Distribution of population (%) (A) and SARS-CoV-2 genomes sequenced (B) in Nigeria
GISAID genomes are of high quality
The genome sequences submitted to GISAID from Nigeria were of high quality, and on average, the genomes analyzed in this study had very few unresolved nucleotides. After masking the first and last 250nt of the genome, the mean proportion of fully resolved nucleotides was 0.916 and the median proportion was 0.966. Only eight genome sequences had less than 50% of the genome resolved. This indicates that the isolation, amplification, and sequencing protocols used by laboratories in Nigeria were proficient at collecting viral genome data. The proportion of fully resolved nucleotides per genome sequence did change over time (Figs. 2 and 3). The Pearson correlation between the proportion of resolved nucleotides and days from November 1, 2019 (approximate start date of the COVID-19 pandemic) was significantly less than 0 (r = −0.232, p < 2.2E-16). A beta regression testing the proportion of resolved nucleotides against time further resulted in a significant negative relationship (coefficient = −1.88E-3, p < 2.2E-16), suggesting a decrease in resolved nucleotides over time. The genomic regions where nucleotide resolution decreases seem to be concentrated in the S locus (spike protein) and toward the termini of the genome sequence (Fig. 4).
Fig. 2.
Proportion of fully resolved nucleotides in SARS-CoV-2 genome sequences collected in Nigeria and deposited into GISAID. Most genome sequences have at least 90% of the genome resolved
Fig. 3.
Proportion of fully resolved nucleotides per genome sequence over time. 1 November, 2019 is used as an approximate start date of the COVID-19 pandemic. The first and last 250nt of the genome have been masked from this calculation
Fig. 4.
Changes in data quality by wave and position across the SARS-CoV-2 genome. Darker colors indicate higher values. Gapping and ambiguous nucleotide fraction increase over time and in specific regions of the genome, including the spike protein locus
There was significant variation in the proportion of resolved nucleotides per genome sequence across GPZs (one-way ANOVA F = 100.6, P < 2E-16). On average, the South-South zone had a significantly higher proportion of resolved nucleotides per genome (0.951) than any other GPZ. The North-Central and North-West GPZs had the lowest average proportion of resolved nucleotides (0.88 and 0.87, respectively). The power to detect differences in data quality among other GPZs is low because of the large imbalance in sequencing data representation among GPZs. However, there was substantial variation in the total number of viral genome sequences obtained from each GPZ; South-West region, which houses two of the three sequencing laboratories, had the highest number [3, 311 genomes], and South-East region had the lowest (98 genomes).
GISAID data recapitulate COVID spread in Nigeria
The collection date of the GISAID sequencing data and the number of COVID-19 new cases closely tracked each other (Figs. 5 and 6). The first wave, however, was not well represented in the sequencing data as the other waves. The assigned lineages of the viral genome sequences also recapitulated the known spread of major viral lineages throughout Nigeria (Fig. 6). The first wave, although poorly represented in the genome sequence data, is primarily composed of one of the original haplotypes of SARS-CoV-2. The second wave was largely made of a mixture of B.1.1.7 (“Alpha”) and B.1.525 (“Eta”) lineages. The third wave was largely made up of AY.36, a sub-variant of “Delta.” The fourth wave had a large representation of AY.36 at the start but shifted to BA.1 and BA.1.1 (“Omicron” and sub-variants). The fifth wave, like the fourth, is largely made up of sub-variants of “Omicron” (Table 2).
Fig. 5.
New cases of COVID-19 by date reported in Nigeria
Fig. 6.
Counts of SARS-CoV-2 genomes deposited in GISAID by date and viral lineage in Nigeria. Only PANGOLIN lineages with at least 50 represented samples are shown for clarity
Table 2.
Most abundant PANGOLIN lineages in each geopolitical area during each wave in Nigeria
| South- West | North-East | South-South | North- Central | South-East | North-West | |
|---|---|---|---|---|---|---|
| Wave 1 | B 1.1 | A | B.1 | B.1.1 | A | A |
| (73; 0.40) | (10; 0.33) | (18; 0.33) | (9; 0.19) | (2; 0.67) | (1; 0.25) | |
| Wave 2 | B.1.1.7 | B.1.525 | B.1.525 | B.1.525 | B.1.525 | L.3 |
| (218; 0.42) | (15; 0.39) | (98; 0.54) | (122; 0.41) | (26; 0.67) | (9; 0.75) | |
| Wave 3 | AY.36 | AY.36 | AY.36 | AY.36 | No samples | AY.36 |
| (920; 0.72) | (20; 0.8) | (456; 0.58) | (334; 0.65) | (52; 0.59) | ||
| Wave 4 | BA.1.1 | BA.1.1 | BA.1.1 | BA.1.1 | BA.1.1 | BA.1.1 |
| (449; 0.46) | (13; 0.59) | (340; 0.53) | (335; 0.35) | (17; 0.49) | (63; 0.40) | |
| Wave 5 | BE.1.1 | BE.1.1 | BE.1.1 | Unknown | BE.1.1 | BA.2.65 |
| (119; 0.33) | (3; 0.33) | (39; 0.24) | (9; 0.43) | (4; 0.25) |
Values are reported as “Lineage (Count; Proportion of sequenced genomes by wave and GPA)
Viral lineage succession in Nigeria
The spread and diversification of viruses throughout Nigeria are apparent beyond the PANGOLIN labels deposited into GISAID. Mutations relative to the reference Wuhan-Hu-1 strain increase with subsequent waves of viral spread (Fig. 7). Of the mutations represented here, only a subset is used to identify specific lineages and named variants. While the spike glycoprotein (S locus) seems to have the most mutations relative to the other open reading frames, other regions of the genome also accumulate many mutations, namely the nucleocapsid gene (N).
Fig. 7.
Per-wave mutation frequency relative to Wuhan-Hu-1. All mutated bases are counted as equal
Within waves, the virus samples have low sequence diversity through most of the genomes with “hotspots” of nucleotide diversity (Fig. 8). The second wave has a different pattern than the others in that this wave had a higher baseline diversity. This is likely because the second wave was predominantly a mixture of two different variants (B.1.17 and B.1.525). Across the genome, the S locus shows high diversity, especially in more advanced waves of viral spread. In the first wave, the sequences showed very little diversity in this region, and diversity rapidly increased with each wave. Waves 4 and 5, which were mostly made up of Omicron lineages, show higher diversity around position 23 kb than the other waves (Fig. 9).
Fig. 8.
Average pairwise diversity across the genome, separated by viral wave
Fig. 9.
Average pairwise diversity in spike, separated by viral wave
Value of public health agencies
Overall, NCDC and State-run hospitals were the largest contributors to the sample collection efforts during this study period. However, the collection efforts shifted over time. During waves 1, 2, and 3, NCDC was the primary source of collected samples. During waves 4 and 5, regional hospitals and other healthcare facilities were the primary sources of samples (Figure S1). This pattern varied by GPZ: in the North-East and South-South zones, regional universities and teaching hospitals were the primary collection facilities. In the North-West zone, privately run labs were the major contributors to sample collection. The North-Central region had a roughly even mix of hospitals, universities, private labs, and public health agencies contributing samples.
For sequencing facilities, the ACEGID and NCDC were the largest contributors to sequencing during this study period. Like with the sample collection, the efforts shifted over time: waves 1, 2, and 3 had a large contribution of sequencing data from ACEGID, and waves 4 and 5 had large contributions of sequences from NCDC (Figure S2). Hospitals affiliated with universities were involved throughout the study period. The contributions also varied by GPZ. NCDC, located centrally in the nation’s capital of Abuja in the north-central zone, contributed the most to sequences from the North-West, North-Central, and North-East zones. ACEGID stationed in South-West contributed the most to sequences from the South-South, and South-West zones. Genomes from the South-East zone were sequenced almost equally by both NCDC and ACEGID.
Discussion
Analyzing over 7,000 SARS-CoV-2 genomes collected from February 2020 to March 2023 during the COVID-19 pandemic, this manuscript aimed to describe the genomic diversity of the COVID-19 pandemic in Nigeria, while exploring the genomic surveillance coverage. Overall, the genomes submitted to GISAID were of high quality, with few unresolved nucleotides, although variability was observed between GPZ. The assigned lineages of the viral genome sequences matched the known spread of major viral lineages throughout Nigeria. As expected, viral diversity rapidly increased with each wave. However, SARS-CoV-2 had low sequence diversity within waves and hotspots [26–28].
The first wave, however, was not well represented in the sequencing data as the other waves. This is not surprising, as protocols for isolating and sequencing SARS-CoV-2 were likely not yet established during this time, and there were very few laboratories with the capacity for SARS-CoV-2 sequencing during the first phase of the pandemic. Furthermore, the genomic data were not fully representative of all Nigerian GPZs. South West (43%) and North Central (27%) contributed the majority of the sequenced samples, even though these GPZs only represent 20% and 15% of Nigeria’s population, respectively (Fig. 1). The latter was expected as the main sequencing institutions are based on these GPZs. ACEGID and NIMR are located in South West, while NCDC Reference Laboratory is based in North Central. Nevertheless, these findings suggest that there were large groups of people who may have been underrepresented during the COVID-19 genomic surveillance in Nigeria, a phenomenon also observed in many other African countries [29–32]. The discrepancies among GPZs regarding viral sample collection and genome sequencing may highlight discrepancies in access to genomics infrastructure for pathogen surveillance. For example, the South East had limited functional SARS-CoV-2 sequencing facilities during the outbreak and had to rely mostly on facilities in other GPZs for SARS-CoV-2 sequencing. The fact that few sequences came from that zone highlights the need to strengthen sample transport across all regions of the country to minimize the impact of differential sequencing facility capabilities across GPZs, to enable pathogen-genomic-based information to be sufficiently representative of the viral evolution across all regions of the country. In addition, considering the large size of the country, a network of sequencing laboratories that ensures at least one operating in each GPZ, all linked to and coordinated by the national reference laboratory at NCDC might provide more balanced coverage for future pandemics and for pathogen surveillance across the GPZs, especially with the likelihood of lockdown during such pandemics that restricts movement of people and samples.
Another limitation of the available genomic data includes the observation that the proportion of unresolved nucleotides and gapped sequences increases over time, suggesting a need to maintain viral sequencing assays to retain their utility. While we do not analyze the raw sequencing data in this study, ambiguous nucleotides and gapped sequences can be used as proxies for data quality. Ambiguous nucleotides represent uncertainty in the sequence, which can arise from low coverage or sequencing errors. In turn, these can be exacerbated by mutations in the priming regions of the amplification primers. Sequence similarity between the primers and divergent strains or species (e.g., other coronaviruses) could also contribute to ambiguity in the data by conflating species divergence with strain divergence. However, this would require that a patient sample be positive for multiple coronaviruses, or cross-contamination between multiple coronaviruses. Gapped sequences arise from lower coverage or true genetic deletions (mutation). While our analysis cannot distinguish the cause of a gapped sequence in the alignment, the natural mutational process can contribute to both causes of gapped sequences. Our results suggest that for genomic assays of pathogens to stay relevant, they must be updated to match the circulating strains. Long-read sequencing may address these limitations by reducing the need to amplify viral genomic segments as intensely as with typical short-read sequencing.
While the temporal patterns and identified lineages in the sequencing data recapitulate the patterns of viral spread through Nigeria, the sequencing data may not be fully representative of the circulating viruses in the community. Lack of standardized protocols, prioritization of samples with higher viral loads, and reagent supply chain challenges may have introduced bias into the lineage patterns identified in this analysis. This suggests the need for a coordinated sequencing strategy and standardized sequencing protocols across the nation to improve genomic surveillance during future outbreaks of both existing and novel pathogens. The first wave was also not sampled very well, likely because the protocols and infrastructure for sample isolation, amplification, and sequencing were still being established. The more recent samples may additionally suffer from greater uncertainty in lineage assignment due to the increased proportion of ambiguous nucleotides and gaps.
Overall, throughout the pandemic, Nigeria expanded and strengthened its genomic surveillance. Sample collection shifted from the national and state-level to regional healthcare facilities. Similarly, the sequencing effort shifted from private research laboratories to government institutes in waves 4–5. This decentralization strategy and transfer of knowledge and collaboration between African institutes and governments, as well as other global organizations and private sector partners, have been described across the continent [10].
Conclusions
Genomic surveillance played a key role in COVID-19 pandemic response and control in Nigeria, and the African Continent as a whole. Furthermore, the establishment of genomic surveillance has provided sequencing data to inform the global diversity of SARS-CoV-2 and served as a tracking tool for the virus's spread across the Continent and the World. It is essential to sustain the investment in genomic surveillance to establish a robust surveillance platform that addresses emerging, reemerging, and endemic infectious disease threats moving forward.
Supplementary Information
Acknowledgements
We gratefully acknowledge all data contributors, i.e., the Authors and their Originating laboratories responsible for obtaining the specimens and their submitting laboratories for generating the genetic sequence and metadata and sharing via the GISAID Initiative, on which this research is based.
Authorship for the INFORM Africa Research Study Group for National Institutes of Health (NIH) D-SI Africa Consortium.
Thomas J. Y. Kono1, Ezenwa J. Onyemata2, Natalia Blanco3, Chika K. Onwuamah4−5, Nnaemeka Ndodo6, Paul Oluniyi7, Olanrewaju Lawal8, Christina Riley9, Sophia Osawe2, Cheryl Baxter10, Anna Winters9, Chenfeng Xiong11, Christian T. Happi7, Babatunde L. Salako4, Ifedayo Adetifa6, Alash’le Abimiku2,3, Manhattan Charurat3, Kristen A. Stafford3, Timothy O’Connor3,Meagan Fitzpatrick3, Mohammad M. Sajadi3, Patrick Dakum2,3, Fati Murtala-Ibrahim2, Nifarta Andrew2, Aminu Musa2, Tolulope Adenekan2, Kenneth Ewerem2, Victoria Etuk2, Mmedorenyin Okon8, Weiyu Luo11, Xin Wu11, Tulio de Oliveira10, Eduan Wilkinson10, Houriiyah Tegally10, Jenicca Poongavanan10, Michelle Parker10, Danilo Silva10, Joicymara S Xavier10, Vivek Naranbhai12, Salim A Karim12, Kennedy Otwombe.13
1Minnesota Supercomputing Institute, University of Minnesota, Minneapolis, Minnesota, USA.
2International Research Center of Excellence, Institute of Human Virology Nigeria, Abuja, Nigeria.
3School of Medicine, University of Maryland, Baltimore, Maryland, United States.
8Department of Geography and Environmental Management, Faculty of Social Sciences University of Port Harcourt, Port Harcourt, Nigeria.
9Akros, Lusaka, Zambia.
10Centre for Epidemic Response and Innovation, Stellenbosch University, Stellenbosch, South Africa.
11Department of Civil and Environmental Engineering, College of Engineering, Villanova University, Villanova, Pennsylvania, United States.
12Centre for the AIDS Programme of Research in South Africa, Durban, South Africa.
13Consortium for Advanced Research Training in Africa (CARTA), Nairobi, Kenya.
Abbreviations
- ACEGID
The African Centre of Excellence for Genomics of Infectious Disease
- COVID-19
Coronavirus disease 2019
- FCT
Federal Capital Territory
- GISAID
The Global Initiative on Sharing All Influenza Data
- GON
Government of Nigeria
- GPZ
Geopolitical Zone
- N
Nucleocapsid gene
- NCDC
The Nigeria Centre for Disease and Control
- NIMR
The Nigerian Institute of Medical Research
- S locus
Spike glycoprotein
Authors’ contributions
Conception: TJYK, NB, MC, AA, KAS. Design: TJYK, NB, MC, AA, KAS. Data Acquisition: CKO, NN, PO, CTH, BLS, IA. Data Analysis: TJYK. Data Interpretation: TJYK, NB, MC, EJO, AA, KAS. Drafted Manuscript: TJYK, NB, KAS. Revise Draft: EJO, CKO, NN, PO, OL, CR, SO, CB, AW, CX, CTH, BLS, IA, AA, MC.
Funding
This research was funded by the National Institutes of Health, grant number U54TW012041.
Data availability
Genomic data is publicly available through GISAID and accessible at [https://doi.org/10.55876/gis8.240507dh].
Declarations
Ethics approval and consent to participate
The National Health Research Ethics Committee of Nigeria (NHREC/01/01/2007–19/01/2022), the University of Kwazulu-Natal Biomedical Research Ethics Committee (BREC/00003832/2022), Villanova University Institutional Review Board (IRB-FY2023-145) and the University of Maryland Baltimore Institutional Review Board (HP-00099829) approved the INFORM Africa protocol. Only publicly-available and non-identifiable data were used in this analysis.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Contributor Information
Natalia Blanco, Email: nblanco@ihv.umaryland.edu.
INFORM Africa Research Study Group:
Timothy O’Connor, Meagan Fitzpatrick, Mohammad M. Sajadi, Patrick Dakum, Fati Murtala-Ibrahim, Nifarta Andrew, Aminu Musa, Tolulope Adenekan, Kenneth Ewerem, Victoria Etuk, Mmedorenyin Okon, Weiyu Luo, Xin Wu, Tulio de Oliveira, Eduan Wilkinson, Houriiyah Tegally, Jenicca Poongavanan, Michelle Parker, Danilo Silva, Joicymara S. Xavier, Vivek Naranbhai, Salim A. Karim, and Kennedy Otwombe
References
- 1.World Health Organization (WHO) Health Emergencies Programmee. https://data.who.int/dashboards/covid19/cases?n=c. 2024. WHO COVID-19 Dashboard.
- 2.University of Oxford. Oxford Martin School. https://ourworldindata.org/coronavirus/country/nigeria. 2024. Our World Data:Nigeria: Coronavirus Pandemic Country Profile.
- 3.Audu RA, Stafford KA, Steinhardt L, Musa ZA, Iriemenam N, Ilori E, et al. Seroprevalence of SARS-CoV-2 in four states of Nigeria in October 2020: A population-based household survey. Nelson MI, editor. PLOS Global Public Health. 2022 Jun 17;2(6):e0000363. Available from: https://dx.plos.org/10.1371/journal.pgph.0000363. [DOI] [PMC free article] [PubMed]
- 4.Kolawole OM, Tomori O, Agbonlahor D, Ekanem E, Bakare R, Abdulsalam N, et al. SARS CoV-2 seroprevalence in selected states of high and low disease burden in Nigeria. JAMA Netw Open. 2022;E2236053. [DOI] [PMC free article] [PubMed]
- 5.Dan-Nwafor C, Ochu CL, Elimian K, Oladejo J, Ilori E, Umeokonkwo C, et al. Nigeria’s public health response to the COVID-19 pandemic: January to May 2020. J Glob Health. 2020;10(2). Available from: /pmc/articles/PMC7696244/. Cited 2023 Apr 12. [DOI] [PMC free article] [PubMed]
- 6.The Center for Policy Impact in Global Health. Nigeria’s Policy Response to COVID-19 . 2020 Jun.
- 7.Damilare Jacobs E, Ifeanyi Okeke M. A critical evaluation of Nigeria’s response to the first wave of COVID-19. Bull Natl Res Cent. 2022;46:44. 10.1186/s42269-022-00729-9. (Cited 2023 Apr 10). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.National Primary Health Care Development Plan. Nigeria: Accelerated COVID-19 Vaccines Introduction and Deployment Plan. 2021.
- 9.Tegally H, San JE, Cotten M, Moir M, Tegomoh B, Mboowa G, et al. The evolving SARS-CoV-2 epidemic in Africa: Insights from rapidly expanding genomic surveillance. Science (1979). 2022;378(6615). [DOI] [PMC free article] [PubMed]
- 10.Ochola R. The Case for Genomic Surveillance in Africa, vol. 10. Tropical Medicine and Infectious Disease: Multidisciplinary Digital Publishing Institute (MDPI); 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Ihekweazu C, Happi C, Omilabu S, Salako B, Abayomi A, Oluniyi P. 2020. https://virological.org/t/first-african-sars-cov-2-genome-sequence-from-nigerian-covid-19-case/421. First African SARS-CoV-2 genome sequence from Nigerian COVID-19 case.
- 12.Tosta S, Moreno K, Schuab G, Fonseca V, Segovia FMC, Kashima S, et al. Global SARS-CoV-2 genomic surveillance: What we have learned (so far). Infect Genet Evol. 2023;1:108. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Worldometer. Nigeria Population. 2025. Available from: https://www.worldometers.info/world-population/nigeria-population/. Cited 2025 Jun 8.
- 14.Worldometer. African Countries by Population. 2025. Available from: https://www.worldometers.info/world-population/nigeria-population/#:~:text=Nigeria%202025%20population%20is%20estimated,(and%20dependencies)%20by%20population. Cited 2025 Jun 8.
- 15.Olawoye IB, Oluniyi PE, Oguzie JU, Uwanibe JN, Kayode TA, Olumade TJ, et al. Emergence and spread of two SARS-CoV-2 variants of interest in Nigeria. Nat Commun. 2023 Dec 1;14(1). [DOI] [PMC free article] [PubMed]
- 16.Ozer EA, Simons LM, Adewumi OM, Fowotade AA, Omoruyi EC, Adeniji JA, et al. Multiple expansions of globally uncommon SARS-CoV-2 lineages in Nigeria. Nat Commun. 2022;13(1). [DOI] [PMC free article] [PubMed]
- 17.Abraham P, Cherian S, Potdar V. Genetic characterization of SARS-CoV-2 & implications for epidemiology, diagnostics & vaccines in India. Vol. 152, Indian Journal of Medical Research. Wolters Kluwer Medknow Publications; 2020. p. 12–5. [DOI] [PMC free article] [PubMed]
- 18.Ndodo N. Tracking SARS-CoV2 variants and strains: an overview of NCDC’s first 596 sequences, November 2021. J Public Health Africa. 2022;13:18–9. [Google Scholar]
- 19.Awoyelu EH, Oladipo EK, Adetuyi BO, Senbadejo TY, Oyawoye OM, Oloke JK. Phyloevolutionary analysis of SARS-CoV-2 in Nigeria. New Microbes New Infect. 2020;1:36. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Onwuamah C, Ahmed R, Toye E, Amoo O, Ndodo N, Momoh E, et al. The trend of SARS-CoV-2 variants from in metropolitan Lagos, the epicentre of the COVID-19 pandemic in Nigeria. 2021; 10.21203/rs.3.rs-1602118/v1
- 21.Khare S, Gurry C, Freitas L, Schultz MB, Bach G, Diallo A, et al. GISAID’s role in pandemic response. China CDC Wkly. 2021;3(49):1049–51. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Katoh K, Standley DM. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol. 2013;30(4):772–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.World Health Organization. https://covid19.who.int/region/afro/country/ng. 2023. Nigeria: COVID-19 situation.
- 24.Mathieu E, Ritchie H, Rodés-Guirao L, Appel C, Giattino C, Hasell J, et al. Published online at OurWorldinData.org. 2020 [cited 2025 Jun 5]. Coronavirus Pandemic (COVID-19). Available from: https://ourworldindata.org/coronavirus/country/nigeria. Cited 2025 Jun 5.
- 25.Dataphyte. Four things to consider as Nigeria’s census goes digital. 2025. Available from: https://archive.dataphyte.com/latest-reports/development/four-things-to-consider-as-nigerias-census-goes-digital/. Cited 2025 Jun 12.
- 26.Zeller M, Gangavarapu K, Anderson C, Smither AR, Vanchiere JA, Rose R, et al. Emergence of an early SARS-CoV-2 epidemic in the United States. Cell. 2021;184(19):4939-4952.e15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Lythgoe KA, Hall M, Ferretti L, de Cesare M, MacIntyre-Cockett G, Trebes A, et al. SARS-CoV-2 within-host diversity and transmission. Science. 2021. 10.1126/science.abg0821. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Wassenaar TM, Wanchai V, Buzard G, Ussery DW. The first three waves of the Covid-19 pandemic hint at a limited genetic repertoire for SARS-CoV-2. Vol. 46, FEMS Microbiology Reviews. Oxford University Press; 2022. [DOI] [PMC free article] [PubMed]
- 29.Sow MS, Togo J, Simons LM, Diallo ST, Magassouba ML, Keita MB, et al. Genomic characterization of SARS-CoV-2 in Guinea, West Africa. PLoS One. 2024 Mar 1;19(3 March). [DOI] [PMC free article] [PubMed]
- 30.Hamzaoui Z, Ferjani S, Medini I, Charaa L, Landolsi I, Ben Ali R, et al. Genomic surveillance of SARS-CoV-2 in North Africa: 4 years of GISAID data sharing. IJID Regions [Internet]. 2024 Jun 1 [cited 2025 Jun 10];11:100356. Available from: https://www.sciencedirect.com/science/article/pii/S2772707624000274 [DOI] [PMC free article] [PubMed]
- 31.Sisay A, Tshiabuila D, van Wyk S, Tesfaye A, Mboowa G, Oyola SO, et al. Molecular Epidemiology and Diversity of SARS-CoV-2 in Ethiopia, 2020–2022. Genes (Basel). 2023;14(3):705. Available from: https://www.mdpi.com/2073-4425/14/3/705/htm. Cited 2025 Jun 10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Morang’a CM, Ngoi JM, Gyamfi J, Amuzu DSY, Nuertey BD, Soglo PM, et al. Genetic diversity of SARS-CoV-2 infections in Ghana from 2020-2021. Nat Commun. 2022;13(1):1–11. Available from: https://www.nature.com/articles/s41467-022-30219-5. Cited 2025 Jun 10. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Genomic data is publicly available through GISAID and accessible at [https://doi.org/10.55876/gis8.240507dh].










