Skip to main content
BMC Genomics logoLink to BMC Genomics
. 2025 Oct 14;26:917. doi: 10.1186/s12864-025-12058-y

Genomic diversity and surveillance of SARS-CoV-2 in Nigeria

Thomas J Y Kono 1, Ezenwa J Onyemata 2, Natalia Blanco 3,, Chika K Onwuamah 4,5, Nnaemeka Ndodo 6, Paul Oluniyi 7, Olanrewaju Lawal 8, Christina Riley 9, Sophia Osawe 2, Cheryl Baxter 10, Anna Winters 9, Chenfeng Xiong 11, Christian T Happi 7, Babatunde L Salako 4, Ifedayo Adetifa 6, Alash’le Abimiku 2,3, Manhattan Charurat 3, Kristen A Stafford 3; INFORM Africa Research Study Group
PMCID: PMC12522711  PMID: 41087859

Abstract

Background

As Nigeria has the sixth-highest population in the world and a significant amount of inbound and outbound travel, the characterization of SARS-CoV-2 genomic diversity across the country is critical for understanding novel pandemic dynamics. We describe the genomic diversity of SARS-CoV-2 in Nigeria throughout the COVID-19 pandemic and examine the coverage of Nigeria's genomic surveillance system.

Methods

Genome sequences and sample metadata were downloaded from the GISAID repository. A beta regression was used to test for a relationship between fully resolved nucleotide proportion over time, as a proxy for data quality. Sample and sequencing source were compared to assess geographic coverage.

Results

A total of 7759 COVID-19 sequences collected from February 2020 to March 2023 were included. The majority were collected in 2021 (76.6%) and South West (43%). Eleven states (30%) reported 10 or fewer SARS-CoV-2 genomes across the entire period. The genome sequences submitted to GISAID from Nigeria were of high quality with very few unresolved nucleotides. Waves 4 and 5, predominantly Omicron lineages, show higher diversity around position 23 kb than the other waves. Overall, the Nigeria Centre for Disease Control (NCDC) and state-run hospitals were the largest contributors to the sample collection efforts during this study period. However, the collection efforts shifted over time from NCDC in waves 1–3 to regional hospitals and other healthcare facilities in waves 4–5, although this pattern varied by geopolitical zone (GPZ). Sequencing efforts also shifted from research laboratories during the first waves to NCDC during waves 4 and 5.

Conclusions

The findings suggest the need for a coordinated sequencing strategy and standardized protocols to improve genomic surveillance during future outbreaks of existing and novel pathogens. A network of sequencing laboratories that includes at least one in each GPZ, linked to and coordinated by the national reference laboratory at NCDC might provide more balanced coverage for future pandemics and pathogen surveillance.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12864-025-12058-y.

Keywords: SARS-CoV-2, COVID-19, Genomic surveillance, Genomic diversity, Nigeria

Introduction

By January 28, 2024, more than 774 million Coronavirus disease 2019 (COVID-19) cases had been reported worldwide. However, only 9.6 million (1.2%) of these cases were reported in Africa [1]. Nigeria, the most populous country in Africa with an estimated 213 million people, contributed only 267 thousand cases (2.7%) to the cases reported by the continent by February 3, 2024 [2]. These numbers are considered an underestimation, especially when compared to published Nigerian COVID-19 seroprevalence survey results, which demonstrated that as much as 25% of the population had been infected with SARS-CoV-2 by October 2020 [3] and up to 79% had been infected by August 2021 in specific states across Nigeria that completed the surveys [4].

In response to the pandemic, the Government of Nigeria (GON) implemented a series of public health measures, including a lockdown on March 30, 2020, which was initially implemented only for high-burden and high-risk areas, including Lagos and Ogun states, and Federal Capital Territory (FCT). The lockdown expanded to other states in April 2020 and began to phase out in May 2020. The lockdown included school and workplace closures, bans on gatherings, curfews, and travel bans [5, 6]. In Nigeria, mask use became mandatory in early May 2020 [6, 7], and vaccine rollout among healthcare workers started in March 2021 and then expanded to the general eligible population in the last quarter of 2021 [8].

Like most countries in Africa [9, 10], the GON incorporated genomic surveillance as part of the country’s response, making Nigeria the first country in Africa to provide a complete SARS-CoV-2 sequence [11]. Genomic surveillance is essential to characterize transmission dynamics, circulation of the virus, and emergence of variants in the population, informing infection control measures [12]. In addition to being the most populous country in Africa, Nigeria has the sixth-highest population in the world and a significant amount of inbound and outbound travel [1316]. Characterization of the genomic diversity of SARS-CoV-2 in areas of the world with a substantial population, population mixing, and mobility, such as Nigeria, is critical for understanding novel pandemic dynamics [12, 16, 17]. SARS-CoV-2 genomic literature from Nigeria is limited, and thus far, has primarily focused on the first waves of the pandemic [15, 18, 19] or specific states [16, 20]. This manuscript aims to describe the genomic diversity of SARS-CoV-2 in Nigeria throughout the COVID-19 pandemic and explore coverage of its genomic surveillance system during the pandemic response.

Methods

Genomic data

Data retrieval

Genome sequences and sample metadata were downloaded from the SARS-CoV-2 section of the GISAID (The Global Initiative on Sharing All Influenza Data) repository on March 29th, 2023 [21]. All SARS-CoV-2 sequences submitted to GISAID from Nigeria before the download date were downloaded for this analysis. Due to the size of the dataset [7, 993 viral genomes in total], the data had to be retrieved in two batches. The unique identifiers of each viral genome (“gisaid_epi_isl”) were downloaded for all viral genomes, then split into two smaller batches for full record retrieval, and accessible at 10.55876/gis8.240507dh. Data were stored on a private portion of the University of Minnesota computer cluster to comply with the GISAID data access agreement.

Metadata standardization and cleaning

Sample metadata from GISAID was manually edited to standardize the recorded names for the state of origin, sample collection facility, and sample sequencing facility. Samples with incomplete collection dates (year only, or month and year only) or invalid collection dates (before 1 October 2019 or after 29 March 2023) were excluded from the analysis. This date range was chosen to exclude samples that were labeled as being collected before the start of the viral pandemic [22] or after the date when the data were downloaded from GISAID. To improve the power of comparative analyses, samples were grouped into Nigeria’s national geopolitical zones (GPZs) based on the state of origin. Nigeria is divided into six GPZs: North Central (Benue, Federal Capital Territory, Kwara, Nasawara, and Plateau), North East (Adamawa, Bauchi, Borno, and Gombe), North West (Kaduna, Kano, Sokoto, and Zamfara), South East (Abia, Ebonyi, Enugu, and Imo), South South (Akwa Ibom, Delta, and Edo), and South West (Ekiti, Lagos, Ogun, Ondo, Osun, and Oyo).

Genome sequence processing

The genomes of the sequenced viral samples were retrieved in FASTA format. All sequences included in the analysis were aligned against the Wuhan-Hu-1 reference genome (Genbank accession NC_045512.2) with MAFFT version 7.475 [22]. The options used for alignment were “–keeplength –6merpair –addfragments –thread 16.” These parameters were chosen to make the alignment computationally tractable and preserve the coordinates of the sequences relative to the annotated reference genome.

The aligned genomes were separated by pandemic wave using the sample collection date. For this purpose, we estimated the date ranges of each wave using Nigeria’s weekly COVID-19 data reported by the World Health Organization [23]. Five specific waves were defined as followed: wave 1 (February 27, 2020 and October 25, 2020), wave 2 (October 26, 2020 and June 20, 2021), wave 3 (June 21, 2021 and November 21, 2021), wave 4 (November 22, 2021 and May 1, 2022), and wave 5 (May 2, 2022 to March 29, 2023).

Aligned genome sequences were masked to estimate pairwise diversity statistics and counts of mutations. The first 250nt and final 250nt of the genome were excluded due to high missing data fraction, likely from unreliable amplification of the ends of the linear genome. Only fully resolved nucleotides (A, C, G, or T) were used to calculate diversity and mutation frequencies; ambiguous nucleotides were excluded from these analyses. Ambiguous nucleotides were instead used to estimate the sequencing quality over time (see below). Masked regions were not used for any analyses. Average pairwise diversity was calculated on a per-nucleotide basis as the total number of pairwise differences in an alignment column divided by the number of pairwise comparisons considered. This is a measure of variability in a group of nucleotide sequences. We used this metric for identifying regions of the genome that were highly variable in temporal or geographic subsets of the genomes. Gaps were treated as missing data and not considered for the pairwise diversity calculations. Total mutation count was calculated as the number of fully resolved differences relative to the Wuhan-Hu-1 reference strain. This measure reports the total genetic divergence from the progenitor strain for a given genome sequence.

Beta regression was used to test for a relationship between fully resolved nucleotide proportion over time, as a proxy for data quality over time. The proportions of fully resolved nucleotides per-genome sequence were transformed to eliminate the limits of 0 and 1 using the following transformation:

graphic file with name d33e644.gif

where y is the proportion of fully resolved nucleotides and n is the sample size. The beta regression was then fitted with Y as the response variable and the number of days since 2019–11-01 as the independent variable. The regression was implemented with the “betareg” package in the R statistical computing environment.

Metadata analyses

Frequencies and percentages were utilized to describe the demographic characteristics of the included sequences.

Additional data

COVID-19 reported new cases from Nigeria were extracted from the Our World in Data COVID-19 dataset [24]. Census 2016 Nigeria population distribution was provided by the National Bureau of Statistics [25].

Results

Demographic characteristics of sequenced samples

In this analysis, we included a total of 7,759 COVID-19 sequences collected from February 2020 to March 2023. Sixty percent of the sequences belonged to individuals aged 20–59 years old, and 48% were collected from male individuals. The majority of the sequenced samples were collected in 2021 (76.6%), and in South West GPZ (42.7%), followed by 27.2% in North-Central GPZ and 23.6% in South-South GPZ (Table 1, Fig. 1). Eleven states (30%) reported 10 or fewer SARS-CoV-2 genomes during the entire period. The main laboratories submitting sequences to GISAID during the pandemic in Nigeria were the African Centre of Excellence for Genomics of Infectious Diseases (ACEGID) (43.7%), the Nigeria Centre for Disease Control (NCDC) (31.6%), and the Nigerian Institute of Medical Research (NIMR) (6.3%).

Table 1.

Demographic characteristics of included sequenced SARS-CoV-2 samples, Nigeria from February 2020 to March 2023

Characteristic n (%) N = 7759
Age
 0–9 155 (2.0)
 10–19 352 (4.5)
 20–29 1220 (15.7)
 30–39 1460 (18.8)
 40–49 1144 (14.7)
 50–59 816 (10.5)
 60–69 455 (5.9)
 70–79 230 (3.0)
 80 and more 115 (1.5)
 Unknown 1812 (23.4)
Sex
 Female 2781 (35.8)
 Male 3711 (47.8)
 Unknown 1267 (16.3)
Geopolitical zones
 North-Central 2108 (27.2)
 North-East 124 (1.6)
 North-West 278 (3.6)
 South-East 98 (1.2)
 South-South 1832 (23.6)
 South-West 3311 (42.7)
 Unknown 8 (0.1)
Submitting Sequencing Lab
 ACEGID 3393 (43.7)
 NCDC 2449 (31.6)
 NIMR 489 (6.3)
 Others 1428 (18.4)
Year of sample collection
 2020 455 (5.8)
 2021 5941 (76.6)
 2022 1346 (17.4)
 2023* 17 (0.2)

*March 29, 2023

African Centre of Excellence for Genomics of Infectious Diseases (ACEGID)

Nigeria Centre for Disease Control (NCDC)

Nigerian Institute of Medical Research (NIMR)

Fig. 1.

Fig. 1

Distribution of population (%) (A) and SARS-CoV-2 genomes sequenced (B) in Nigeria

GISAID genomes are of high quality

The genome sequences submitted to GISAID from Nigeria were of high quality, and on average, the genomes analyzed in this study had very few unresolved nucleotides. After masking the first and last 250nt of the genome, the mean proportion of fully resolved nucleotides was 0.916 and the median proportion was 0.966. Only eight genome sequences had less than 50% of the genome resolved. This indicates that the isolation, amplification, and sequencing protocols used by laboratories in Nigeria were proficient at collecting viral genome data. The proportion of fully resolved nucleotides per genome sequence did change over time (Figs. 2 and 3). The Pearson correlation between the proportion of resolved nucleotides and days from November 1, 2019 (approximate start date of the COVID-19 pandemic) was significantly less than 0 (r = −0.232, p < 2.2E-16). A beta regression testing the proportion of resolved nucleotides against time further resulted in a significant negative relationship (coefficient = −1.88E-3, p < 2.2E-16), suggesting a decrease in resolved nucleotides over time. The genomic regions where nucleotide resolution decreases seem to be concentrated in the S locus (spike protein) and toward the termini of the genome sequence (Fig. 4).

Fig. 2.

Fig. 2

Proportion of fully resolved nucleotides in SARS-CoV-2 genome sequences collected in Nigeria and deposited into GISAID. Most genome sequences have at least 90% of the genome resolved

Fig. 3.

Fig. 3

Proportion of fully resolved nucleotides per genome sequence over time. 1 November, 2019 is used as an approximate start date of the COVID-19 pandemic. The first and last 250nt of the genome have been masked from this calculation

Fig. 4.

Fig. 4

Changes in data quality by wave and position across the SARS-CoV-2 genome. Darker colors indicate higher values. Gapping and ambiguous nucleotide fraction increase over time and in specific regions of the genome, including the spike protein locus

There was significant variation in the proportion of resolved nucleotides per genome sequence across GPZs (one-way ANOVA F = 100.6, P < 2E-16). On average, the South-South zone had a significantly higher proportion of resolved nucleotides per genome (0.951) than any other GPZ. The North-Central and North-West GPZs had the lowest average proportion of resolved nucleotides (0.88 and 0.87, respectively). The power to detect differences in data quality among other GPZs is low because of the large imbalance in sequencing data representation among GPZs. However, there was substantial variation in the total number of viral genome sequences obtained from each GPZ; South-West region, which houses two of the three sequencing laboratories, had the highest number [3, 311 genomes], and South-East region had the lowest (98 genomes).

GISAID data recapitulate COVID spread in Nigeria

The collection date of the GISAID sequencing data and the number of COVID-19 new cases closely tracked each other (Figs. 5 and 6). The first wave, however, was not well represented in the sequencing data as the other waves. The assigned lineages of the viral genome sequences also recapitulated the known spread of major viral lineages throughout Nigeria (Fig. 6). The first wave, although poorly represented in the genome sequence data, is primarily composed of one of the original haplotypes of SARS-CoV-2. The second wave was largely made of a mixture of B.1.1.7 (“Alpha”) and B.1.525 (“Eta”) lineages. The third wave was largely made up of AY.36, a sub-variant of “Delta.” The fourth wave had a large representation of AY.36 at the start but shifted to BA.1 and BA.1.1 (“Omicron” and sub-variants). The fifth wave, like the fourth, is largely made up of sub-variants of “Omicron” (Table 2).

Fig. 5.

Fig. 5

New cases of COVID-19 by date reported in Nigeria

Fig. 6.

Fig. 6

Counts of SARS-CoV-2 genomes deposited in GISAID by date and viral lineage in Nigeria. Only PANGOLIN lineages with at least 50 represented samples are shown for clarity

Table 2.

Most abundant PANGOLIN lineages in each geopolitical area during each wave in Nigeria

South- West North-East South-South North- Central South-East North-West
Wave 1 B 1.1 A B.1 B.1.1 A A
(73; 0.40) (10; 0.33) (18; 0.33) (9; 0.19) (2; 0.67) (1; 0.25)
Wave 2 B.1.1.7 B.1.525 B.1.525 B.1.525 B.1.525 L.3
(218; 0.42) (15; 0.39) (98; 0.54) (122; 0.41) (26; 0.67) (9; 0.75)
Wave 3 AY.36 AY.36 AY.36 AY.36 No samples AY.36
(920; 0.72) (20; 0.8) (456; 0.58) (334; 0.65) (52; 0.59)
Wave 4 BA.1.1 BA.1.1 BA.1.1 BA.1.1 BA.1.1 BA.1.1
(449; 0.46) (13; 0.59) (340; 0.53) (335; 0.35) (17; 0.49) (63; 0.40)
Wave 5 BE.1.1 BE.1.1 BE.1.1 Unknown BE.1.1 BA.2.65
(119; 0.33) (3; 0.33) (39; 0.24) (9; 0.43) (4; 0.25)

Values are reported as “Lineage (Count; Proportion of sequenced genomes by wave and GPA)

Viral lineage succession in Nigeria

The spread and diversification of viruses throughout Nigeria are apparent beyond the PANGOLIN labels deposited into GISAID. Mutations relative to the reference Wuhan-Hu-1 strain increase with subsequent waves of viral spread (Fig. 7). Of the mutations represented here, only a subset is used to identify specific lineages and named variants. While the spike glycoprotein (S locus) seems to have the most mutations relative to the other open reading frames, other regions of the genome also accumulate many mutations, namely the nucleocapsid gene (N).

Fig. 7.

Fig. 7

Per-wave mutation frequency relative to Wuhan-Hu-1. All mutated bases are counted as equal

Within waves, the virus samples have low sequence diversity through most of the genomes with “hotspots” of nucleotide diversity (Fig. 8). The second wave has a different pattern than the others in that this wave had a higher baseline diversity. This is likely because the second wave was predominantly a mixture of two different variants (B.1.17 and B.1.525). Across the genome, the S locus shows high diversity, especially in more advanced waves of viral spread. In the first wave, the sequences showed very little diversity in this region, and diversity rapidly increased with each wave. Waves 4 and 5, which were mostly made up of Omicron lineages, show higher diversity around position 23 kb than the other waves (Fig. 9).

Fig. 8.

Fig. 8

Average pairwise diversity across the genome, separated by viral wave

Fig. 9.

Fig. 9

Average pairwise diversity in spike, separated by viral wave

Value of public health agencies

Overall, NCDC and State-run hospitals were the largest contributors to the sample collection efforts during this study period. However, the collection efforts shifted over time. During waves 1, 2, and 3, NCDC was the primary source of collected samples. During waves 4 and 5, regional hospitals and other healthcare facilities were the primary sources of samples (Figure S1). This pattern varied by GPZ: in the North-East and South-South zones, regional universities and teaching hospitals were the primary collection facilities. In the North-West zone, privately run labs were the major contributors to sample collection. The North-Central region had a roughly even mix of hospitals, universities, private labs, and public health agencies contributing samples.

For sequencing facilities, the ACEGID and NCDC were the largest contributors to sequencing during this study period. Like with the sample collection, the efforts shifted over time: waves 1, 2, and 3 had a large contribution of sequencing data from ACEGID, and waves 4 and 5 had large contributions of sequences from NCDC (Figure S2). Hospitals affiliated with universities were involved throughout the study period. The contributions also varied by GPZ. NCDC, located centrally in the nation’s capital of Abuja in the north-central zone, contributed the most to sequences from the North-West, North-Central, and North-East zones. ACEGID stationed in South-West contributed the most to sequences from the South-South, and South-West zones. Genomes from the South-East zone were sequenced almost equally by both NCDC and ACEGID.

Discussion

Analyzing over 7,000 SARS-CoV-2 genomes collected from February 2020 to March 2023 during the COVID-19 pandemic, this manuscript aimed to describe the genomic diversity of the COVID-19 pandemic in Nigeria, while exploring the genomic surveillance coverage. Overall, the genomes submitted to GISAID were of high quality, with few unresolved nucleotides, although variability was observed between GPZ. The assigned lineages of the viral genome sequences matched the known spread of major viral lineages throughout Nigeria. As expected, viral diversity rapidly increased with each wave. However, SARS-CoV-2 had low sequence diversity within waves and hotspots [2628].

The first wave, however, was not well represented in the sequencing data as the other waves. This is not surprising, as protocols for isolating and sequencing SARS-CoV-2 were likely not yet established during this time, and there were very few laboratories with the capacity for SARS-CoV-2 sequencing during the first phase of the pandemic. Furthermore, the genomic data were not fully representative of all Nigerian GPZs. South West (43%) and North Central (27%) contributed the majority of the sequenced samples, even though these GPZs only represent 20% and 15% of Nigeria’s population, respectively (Fig. 1). The latter was expected as the main sequencing institutions are based on these GPZs. ACEGID and NIMR are located in South West, while NCDC Reference Laboratory is based in North Central. Nevertheless, these findings suggest that there were large groups of people who may have been underrepresented during the COVID-19 genomic surveillance in Nigeria, a phenomenon also observed in many other African countries [2932]. The discrepancies among GPZs regarding viral sample collection and genome sequencing may highlight discrepancies in access to genomics infrastructure for pathogen surveillance. For example, the South East had limited functional SARS-CoV-2 sequencing facilities during the outbreak and had to rely mostly on facilities in other GPZs for SARS-CoV-2 sequencing. The fact that few sequences came from that zone highlights the need to strengthen sample transport across all regions of the country to minimize the impact of differential sequencing facility capabilities across GPZs, to enable pathogen-genomic-based information to be sufficiently representative of the viral evolution across all regions of the country. In addition, considering the large size of the country, a network of sequencing laboratories that ensures at least one operating in each GPZ, all linked to and coordinated by the national reference laboratory at NCDC might provide more balanced coverage for future pandemics and for pathogen surveillance across the GPZs, especially with the likelihood of lockdown during such pandemics that restricts movement of people and samples.

Another limitation of the available genomic data includes the observation that the proportion of unresolved nucleotides and gapped sequences increases over time, suggesting a need to maintain viral sequencing assays to retain their utility. While we do not analyze the raw sequencing data in this study, ambiguous nucleotides and gapped sequences can be used as proxies for data quality. Ambiguous nucleotides represent uncertainty in the sequence, which can arise from low coverage or sequencing errors. In turn, these can be exacerbated by mutations in the priming regions of the amplification primers. Sequence similarity between the primers and divergent strains or species (e.g., other coronaviruses) could also contribute to ambiguity in the data by conflating species divergence with strain divergence. However, this would require that a patient sample be positive for multiple coronaviruses, or cross-contamination between multiple coronaviruses. Gapped sequences arise from lower coverage or true genetic deletions (mutation). While our analysis cannot distinguish the cause of a gapped sequence in the alignment, the natural mutational process can contribute to both causes of gapped sequences. Our results suggest that for genomic assays of pathogens to stay relevant, they must be updated to match the circulating strains. Long-read sequencing may address these limitations by reducing the need to amplify viral genomic segments as intensely as with typical short-read sequencing.

While the temporal patterns and identified lineages in the sequencing data recapitulate the patterns of viral spread through Nigeria, the sequencing data may not be fully representative of the circulating viruses in the community. Lack of standardized protocols, prioritization of samples with higher viral loads, and reagent supply chain challenges may have introduced bias into the lineage patterns identified in this analysis. This suggests the need for a coordinated sequencing strategy and standardized sequencing protocols across the nation to improve genomic surveillance during future outbreaks of both existing and novel pathogens. The first wave was also not sampled very well, likely because the protocols and infrastructure for sample isolation, amplification, and sequencing were still being established. The more recent samples may additionally suffer from greater uncertainty in lineage assignment due to the increased proportion of ambiguous nucleotides and gaps.

Overall, throughout the pandemic, Nigeria expanded and strengthened its genomic surveillance. Sample collection shifted from the national and state-level to regional healthcare facilities. Similarly, the sequencing effort shifted from private research laboratories to government institutes in waves 4–5. This decentralization strategy and transfer of knowledge and collaboration between African institutes and governments, as well as other global organizations and private sector partners, have been described across the continent [10].

Conclusions

Genomic surveillance played a key role in COVID-19 pandemic response and control in Nigeria, and the African Continent as a whole. Furthermore, the establishment of genomic surveillance has provided sequencing data to inform the global diversity of SARS-CoV-2 and served as a tracking tool for the virus's spread across the Continent and the World. It is essential to sustain the investment in genomic surveillance to establish a robust surveillance platform that addresses emerging, reemerging, and endemic infectious disease threats moving forward.

Supplementary Information

Supplementary Material 1. (130.4KB, docx)

Acknowledgements

We gratefully acknowledge all data contributors, i.e., the Authors and their Originating laboratories responsible for obtaining the specimens and their submitting laboratories for generating the genetic sequence and metadata and sharing via the GISAID Initiative, on which this research is based.

Authorship for the INFORM Africa Research Study Group for National Institutes of Health (NIH) D-SI Africa Consortium.

Thomas J. Y. Kono1, Ezenwa J. Onyemata2, Natalia Blanco3, Chika K. Onwuamah4−5, Nnaemeka Ndodo6, Paul Oluniyi7, Olanrewaju Lawal8, Christina Riley9, Sophia Osawe2, Cheryl Baxter10, Anna Winters9, Chenfeng Xiong11, Christian T. Happi7, Babatunde L. Salako4, Ifedayo Adetifa6, Alash’le Abimiku2,3, Manhattan Charurat3, Kristen A. Stafford3, Timothy O’Connor3,Meagan Fitzpatrick3, Mohammad M. Sajadi3, Patrick Dakum2,3, Fati Murtala-Ibrahim2, Nifarta Andrew2, Aminu Musa2, Tolulope Adenekan2, Kenneth Ewerem2, Victoria Etuk2, Mmedorenyin Okon8, Weiyu Luo11, Xin Wu11, Tulio de Oliveira10, Eduan Wilkinson10, Houriiyah Tegally10, Jenicca Poongavanan10, Michelle Parker10, Danilo Silva10, Joicymara S Xavier10, Vivek Naranbhai12, Salim A Karim12, Kennedy Otwombe.13

1Minnesota Supercomputing Institute, University of Minnesota, Minneapolis, Minnesota, USA.

2International Research Center of Excellence, Institute of Human Virology Nigeria, Abuja, Nigeria.

3School of Medicine, University of Maryland, Baltimore, Maryland, United States.

8Department of Geography and Environmental Management, Faculty of Social Sciences University of Port Harcourt, Port Harcourt, Nigeria.

9Akros, Lusaka, Zambia.

10Centre for Epidemic Response and Innovation, Stellenbosch University, Stellenbosch, South Africa.

11Department of Civil and Environmental Engineering, College of Engineering, Villanova University, Villanova, Pennsylvania, United States.

12Centre for the AIDS Programme of Research in South Africa, Durban, South Africa.

13Consortium for Advanced Research Training in Africa (CARTA), Nairobi, Kenya.

Abbreviations

ACEGID

The African Centre of Excellence for Genomics of Infectious Disease

COVID-19

Coronavirus disease 2019

FCT

Federal Capital Territory

GISAID

The Global Initiative on Sharing All Influenza Data

GON

Government of Nigeria

GPZ

Geopolitical Zone

N

Nucleocapsid gene

NCDC

The Nigeria Centre for Disease and Control

NIMR

The Nigerian Institute of Medical Research

S locus

Spike glycoprotein

Authors’ contributions

Conception: TJYK, NB, MC, AA, KAS. Design: TJYK, NB, MC, AA, KAS. Data Acquisition: CKO, NN, PO, CTH, BLS, IA. Data Analysis: TJYK. Data Interpretation: TJYK, NB, MC, EJO, AA, KAS. Drafted Manuscript: TJYK, NB, KAS. Revise Draft: EJO, CKO, NN, PO, OL, CR, SO, CB, AW, CX, CTH, BLS, IA, AA, MC.

Funding

This research was funded by the National Institutes of Health, grant number U54TW012041.

Data availability

Genomic data is publicly available through GISAID and accessible at [https://doi.org/10.55876/gis8.240507dh].

Declarations

Ethics approval and consent to participate

The National Health Research Ethics Committee of Nigeria (NHREC/01/01/2007–19/01/2022), the University of Kwazulu-Natal Biomedical Research Ethics Committee (BREC/00003832/2022), Villanova University Institutional Review Board (IRB-FY2023-145) and the University of Maryland Baltimore Institutional Review Board (HP-00099829) approved the INFORM Africa protocol. Only publicly-available and non-identifiable data were used in this analysis.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Natalia Blanco, Email: nblanco@ihv.umaryland.edu.

INFORM Africa Research Study Group:

Timothy O’Connor, Meagan Fitzpatrick, Mohammad M. Sajadi, Patrick Dakum, Fati Murtala-Ibrahim, Nifarta Andrew, Aminu Musa, Tolulope Adenekan, Kenneth Ewerem, Victoria Etuk, Mmedorenyin Okon, Weiyu Luo, Xin Wu, Tulio de Oliveira, Eduan Wilkinson, Houriiyah Tegally, Jenicca Poongavanan, Michelle Parker, Danilo Silva, Joicymara S. Xavier, Vivek Naranbhai, Salim A. Karim, and Kennedy Otwombe

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (130.4KB, docx)

Data Availability Statement

Genomic data is publicly available through GISAID and accessible at [https://doi.org/10.55876/gis8.240507dh].


Articles from BMC Genomics are provided here courtesy of BMC

RESOURCES