Abstract
Purpose of review
Whole genome sequencing (WGS) has transformed bacterial strain typing, an essential tool for outbreak detection, antimicrobial resistance surveillance, and tracking clonal emergence across clinical, research, and public health settings. Herein, we will review recent advances in WGS-based bacterial strain typing methods for purposes of comparison and classification with a focus on improvements in variant identification, strain classification, and transmission assessment.
Recent findings
Advances in sequencing technologies as well as variant calling methodologies and parameter optimization have enhanced the precision and accuracy of single nucleotide variant identification. Hierarchical clustering of gene-by-gene strain typing, combined with novel data management and classification strategies, has improved standardized pathogen typing schemes in an effort to streamline inter-laboratory comparison. Additionally, novel approaches to defining transmission thresholds now better account for species-specific traits, while progress in metagenomic sequencing enables strain identification and tracking within mixed microbial communities.
Summary
Recent developments have enhanced the accuracy, portability, scalability, and standardization of bacterial typing methods, integrating variant calling and gene-by-gene approaches into unified genotyping systems. However, challenges still remain in nomenclature consistency, inter-laboratory variant calling compatibility, and capturing bacterial heterogeneity. Future work should focus on refining genotyping frameworks to enhance surveillance and optimize detection of pathogen transmission while accounting for microbial diversity across various environments.
Keywords: gene-by-gene strain typing, hierarchical clustering, metagenomics, transmission threshold, variant calling
INTRODUCTION
The assessment of microbial relatedness has been of paramount importance since the acceptance of germ theory in the late 19th century. Robert Koch's postulates are at the core of germ theory with the fourth postulate being particularly crucial to strain relatedness as it states, following experimental infection, the disease-causing microbe must be re-isolated in identical form from the newly infected organism. Nearly 150 years after Koch's work, the ability to accurately measure microbial relatedness, now primarily achieved through molecular techniques, has become increasingly important, as exemplified by the SARS-CoV-2 pandemic.
Herein, we update a 2021 review in this journal on bacterial strain typing [1], as we seek to cover its history, significance, and the impact of advancing whole-genome sequencing (WGS) technologies and analytical methodologies on defining genetic relatedness. We focus on how these developments have clarified some aspects of strain classification and comparison while raising new questions about what the best practices are to accurately and precisely standardize bacterial typing methods.
Box 1.
no caption available
A BRIEF HISTORY OF BACTERIAL RELATEDNESS ASSESSMENT IN THE PRE-GENOMIC ERA
Prior to the discovery of nucleic acid, bacteria were solely distinguished through phenotypic variation. Since the late 19th century, researchers have developed methods to phenotypically classify and differentiate bacteria that are still used to this day, such as Gram staining for morphology as well as catalase, oxidase, and acid production tests [2]. Advances in phenotypic-based bacterial typing continued through the mid-20th century, such as the development of serotyping and phage typing, which allowed researchers to cluster bacteria based on surface antigen composition (e.g., grouping of β-hemolytic streptococci [3]) and bacteriophage susceptibility (e.g., Salmonella Typhi [4]), respectively. Advances in molecular genetics and genomics have since linked specific genomic markers, such as surface antigen-encoding genes, to phenotypic traits, bridging traditional classification with the genomic era.
THE ARRIVAL OF MOLECULAR BASED BACTERIAL TYPING
The emergence of molecular typing techniques in the late 20th century leveraged genomic differences to achieve higher resolution than traditional phenotypic tests, enabling far more precise differentiation of bacterial strains. The first molecular techniques used to characterize genomic differences relied on restriction enzyme digestion followed by electrophoretic separation of DNA fragments, generating distinct banding patterns [5]. These methods, including pulsed-field gel electrophoresis (PFGE), provided a foundation for molecular epidemiology, with PFGE emerging as the gold standard for bacterial strain identification for nearly two decades [2,5,6]. Its widespread adoption led to the establishment of the U.S. Centers for Disease Control and Prevention PulseNet laboratory network in 1996, which early on encouraged investigators to develop best practices for typing schema definitions, database storage, and optimized standardized protocols across public health laboratories [6].
PCR-based multilocus sequence typing (MLST), introduced in 1998, served as a bridge between early molecular typing methods and modern WGS approaches. It enabled standardized, portable, and reproducible bacterial strain characterization by sequencing 7–8 conserved housekeeping genes [7]. The combination of alleles forms a sequence type (ST), which is cataloged in curated online databases for standardized strain comparison [8]. While MLST offers greater discriminatory power than PFGE, these techniques share limitations with earlier molecular approaches, including labor-intensive protocols and reliance on a small fraction of the genome [5]. These limitations have begun to be addressed with the advancement and greater accessibility of WGS-based bacterial typing.
WHOLE GENOME SEQUENCING: THE PRESENT AND FUTURE OF BACTERIAL TYPING
While WGS of bacteria serves numerous applications, one of its most critical uses is determining the degree of similarity between two or more strains. Despite clonal reproduction, bacterial genomes are highly dynamic, frequently altered by mobile genetic elements (MGEs), structural rearrangements, and homologous recombination, which complicates accurate assessment of genetic relatedness. At present, measures of WGS-based relatedness generally fall into two categories, namely single nucleotide variant (SNV) calling and ’gene-by-gene’ allelic approaches [5]. Methodologies of each strategy will be delineated before moving on to describing how the derived data can be used to assess strain relatedness.
GENETIC RELATEDNESS DEFINED BY NUCLEOTIDE VARIATION
Sequence alignment considerations
The alignment of two or more sequences to assess relatedness is fundamental to computational biology and serves as the cornerstone of how sequencing data is analyzed and interpreted (N.B.: for a comprehensive overview of whole genome alignment algorithms and methods, see [9,10]). Key alignment considerations, such as input data preprocessing (e.g., quality filtering and adapter trimming) as well as selecting appropriate alignment algorithms/software and tuning parameters (e.g., mismatch and gap penalties), enhance the accuracy of variant detection by reducing mapping artifacts and increasing the likelihood of proper alignment [9]. Given the small amount of error inherent in sequencing processes, sequencing depth, also known as coverage depth, is another critical factor in accurate alignments. When performing reference-based read mapping, selecting a reference strain that closely resembles the population of isolates being analyzed is crucial for accurate variant calling. A highly divergent reference can introduce alignment biases and miscalls, which can reduce variant calling sensitivity/specificity, as has been shown in Enterobacterales [11]. Another major consideration is genome masking, which involves excluding regions prone to high variability (e.g., MGEs or recombination hotspots) from read mapping and variant analysis to reduce false positives and improve alignment accuracy thereby avoiding artificial inflation of genetic differences; however, over masking can potentially omit informative genomic elements [12▪].
Variant calling parameter optimization
Once an alignment file has been generated, variant calling algorithms can be used to identify genetic differences, such as single nucleotide polymorphisms (SNPs) or insertions/deletions (INDELs) in the mapped reads relative to the reference genome. These programs typically generate variant call format (VCF) files, which can be used to compare strain-level differences or to infer phylogenetic relationships across larger cohorts.
The proficiency of aligner and variant calling tool combinations (i.e., variant calling pipelines/workflows) have been assessed, until recently, almost exclusively in the context of short-read (i.e., 50–500 bp sequence fragments) data [11,12▪,13,14▪]. Seah et al. found that with short-read sequencing data, as the ratio of INDELS to SNPs increases and as the INDEL length grows, variant calling precision and recall decline [11,14▪]. Furthermore, studies on benchmarking variant calling in Mycobacterium tuberculosis have highlighted challenges such as high false positive rates, particularly in hypervariable regions, and inconsistencies in masking and variant filtering, which complicate the implementation of standardized, reference-based short-read variant calling pipelines across diverse applications [12▪,13].
As previously alluded to when discussing masking, another important variant calling consideration is the choice of which genomic sites to include in a phylogenetic analysis, particularly when large numbers of strains are involved. A group of bacteria contain a core genome, consisting of genes/intergenic regions shared by all strains above a defined conservation threshold, and an accessory genome, which includes genes/intergenic regions variably present across strains. It is generally the core genome which is used to assess genetic relatedness. However, now that thousands of isolates can be studied with computational ease, the inverse relationship between group number and core genome size can result in suboptimal phylogenetic resolution. Taouk et al. recently demonstrated the value of considering the “soft-core” by analyzing sites that are present in most but not all strains in a given cohort. They found that soft-core SNP alignments with an inclusion cut-off of 95% presence retained more informative sites than “strict-core” (e.g., genomic sites present in >99%) approaches, improving phylogenetic accuracy in large, diverse microbial datasets of Neisseria gonorrhoeae and Salmonella enterica serovar Typhi and introduced an open-source tool for generating these alignments [15▪].
Long-read variant calling and synthesizing data
While most WGS-based bacterial typing has utilized short-read data, recent advances in long-read sequencing technologies are rapidly expanding their utility in this field [16,17▪▪]. Before these advancements, the long-read sequencing Oxford Nanopore Technologies (ONT) platform had high per-base error rates (∼1–3%), especially in homopolymer regions (e.g., AAAAAA), which limited the accuracy and feasibility of ONT-based variant calling [18]. Recent outbreak analyses using ONT-only approaches have indicated the transition from R9.4.1 to R10.4.1 flowcells (i.e., V10 to V14 chemistries) in addition to improved basecalling models (i.e., Guppy to Dorado basecalling) have made ONT-only variant calling a viable approach [17▪▪,19▪▪,20]. An important study that demonstrated the accuracy of long-read data for variant calling is Hall et al., where the authors benchmarked seven ONT long-read variant calling pipelines using fourteen fully assembled, error-free bacterial genomes [17▪▪]. Their study assessed the variant calling performance of various ONT sequencing models and coverage depths as well as included a comparison with Illumina short-read variant calling [17▪▪]. Their key finding was that ONT-based variant callers Clair3 [21] and DeepVariant [22] achieved significantly higher F1 scores (i.e., metric that balances precision and recall), using both high-accuracy and super-accurate (SUP) models, often outperforming Illumina short-read variant calling by an order of magnitude [17▪▪]. Given the relative ease and scalability of ONT sequencing, the long-read approach may become more widely utilized for WGS bacterial typing moving forward.
The capacity to use multiple types of WGS data inputs (e.g. short-read, long-read, etc.) for variant calling holds broad appeal given the flexibility afforded by the ability to analyze disparate data with distinct intrinsic error profiles. The authors of Bogaerts et al. demonstrated how Dorado SUP ‘duplex’ ONT data (i.e., data with both strands of DNA sequenced in contrast to a single strand) from the R10.4 flowcell could potentially be used interchangeably with Illumina short-read data [19▪▪]. However, there are still challenges in complete standardization and validation required to fully enable this interchangeability. Charron et al. have proposed a promising workflow that offers flexibility in utilizing short-read, long-read, or assembly data for detecting SNPs, INDELs, and SVs [23▪]. This approach allows researchers to obtain reproducible and high-confidence variant calls from publicly available data, which will be critical when trying to integrate disparately generated bacterial WGS datasets [23▪].
Alternative variant calling approaches
As previously noted, mapping of reads to a reference has long been fundamental to bacterial strain typing but also generates potential errors particularly when genetically dissimilar isolates are being studied. Thus, alternative variant calling methodologies that do not rely on canonical reference-based read mapping approaches are being increasingly utilized. The reference agnostic split k-mer alignment 2 (SKA2) tool uses a k-mer matching technique that can eliminate spurious SNPs resulting from misalignments often seen when mapping isolates against a highly divergent reference [24▪]. The updated, computationally efficient parsnp 2.0 performs a multisequence alignment across assemblies, in contrast to pairwise alignment, which also reduces reference-based biases and can accurately identify orthologous, conserved regions for core SNP calling [25▪].
Overall, reference-based variant calling remains foundational, but challenges such as alignment biases, errors induced by INDELs, and species-specific variability persist. Fortunately, long-read sequencing technologies, particularly ONT R10.4.1 with improved basecalling models and alignment algorithms, can enhance SNP/INDEL detection, potentially rivaling and surpassing Illumina short-read results. Emerging tools, including reference-free methods and hybrid workflows, offer greater flexibility and standardization, improving phylogenetic resolution and outbreak surveillance.
GENE-BY-GENE APPROACHES TO BACTERIAL STRAIN TYPING
Gene-by-gene allelic differences have been a gold standard for classifying and binning similar bacterial lineages ever since the introduction of MLST as aforementioned. With the advent of WGS, core genome MLST (cgMLST) built upon traditional 7–8 allele MLST by analyzing hundreds to thousands of conserved loci across the genome, offering far higher resolution [26]. Whole genome MLST (wgMLST) further extends this approach to include both core and accessory genes, enabling even finer strain differentiation while capturing broader genomic diversity within bacterial populations [27]. Importantly, publicly available cgMLST/wgMLST tools such as chewBBACA (i.e., BSR-Based Allele Calling Algorithm) have created an open, portable software suite that has significantly aided in cgMLST schema creation as well as facilitated these analyses across labs [28].
The preceding strain typing review in this journal provided a detailed characterization of gene-by-gene typing approaches as well as a conceptual framework of genetic relatedness [1]. Herein, we address recent efforts to improve the portability and scalability of bacterial typing systems, which aim to ‘future-proof’ evolving definitions of relatedness by enabling seamless integration of new data into increasingly complex classification frameworks.
Scalable and portable typing schema
As WGS data generation accelerates, a major challenge in bacterial classification is developing scalable, standardized schema that accurately reflect genetic relationships. Table 1 summarizes widely used strain typing databases that incorporate such standardized frameworks [8,29,30,31▪,32▪▪]. Traditional classification methods, such as MLST, rely on arbitrary integer-based clustering, which cannot capture the complexity of genomic variation of a grouping based solely on 7–8 housekeeping genes. Moreover, no information is provided in an MSLT schema regarding the genetic similarity of any two STs (e.g., ST1 vs. ST2). These limitations hinder the capacity to find sub-level clustering of groups for classification purposes as well as reduces the capacity to scale and interpret data across labs. Efforts are currently being made across well curated bacterial typing databases to alleviate these issues.
Table 1.
Examples of bacterial strain typing databases, software, and web servers
| Database | Genotyping method | Underlying software/algorithms | Strengths | Weaknesses | Citations |
| PubMLST | MLST, cgMLST, wgMLST | BIGSdb | Gold standard, comprehensive typing repository. Performs other in silico typing methods (e.g., lps typing). Houses a comparative genomics application, Genome Comparator, tool. |
Does not adhere to hierarchical level typing. BIGSdb for other species hosted on other platforms (i.e., BIGSdb-Pasteur web platform). |
[8,29] |
| EnteroBase | MLST, cgMLST, wgMLST | HierCC | Incorporates hierarchical clustering to cgMLST nomenclature. Provides phylogenetics, comparative genomics capacity. Improved user-interface (as of v1.2.0). |
Varying accuracy of particular typing schemes that are species dependent. Complex for beginners. |
[30,31▪] |
| chewie-ns | cgMLST, wgMLST | TypOn ontology | Open access compatibility with chewBBACA. Can submit novel cg/wgMLST schema. Does not rely on web server and has CLI functionality. |
High-confidence allelic schema of user-submitted data can be variable. | [32▪▪] |
| cgMLST.org | cgMLST | ridom SeqSphere+ | Standardized cgMLST schemas. Interoperable data formats. |
Primarily intended for use with proprietary ridom SeqSphere+ Software. Can’t incorporate new schema without SeqSphere+ software. |
None |
The authors of pHierCC developed a hierarchical multilevel clustering pipeline, integrated into EnteroBase in 2018, which efficiently scales with large datasets and improves sub-lineage resolution within cgMLST constructs compared to single-level clustering [33▪▪]. Hennart et al. expanded the cgMLST multilevel clustering approach by developing a ‘dual-barcoding’ schema [34▪▪]. This system integrates (1) Life Identification Numbers (LINs), a hierarchical and stable nomenclature assigning codes based on genomic similarity of closest assigned genome within a group and (2) Multilevel Single Linkage (MLSL) nomenclature, a SNP-based group identifier, to generate ‘cgLIN’ codes within BIGSdb versions ≥v1.34.0 [34▪▪]. Importantly, cgLIN codes not only offer a stable approach to future taxonomic classification but also afford backwards compatibility with existing MLST nomenclature [34▪▪]. Recently, cgLIN codes have been adopted more widely for surveillance of hypervirulent and multidrug-resistant K. pneumoniae[35–37].
Zhong et al. developed the distributed cgMLST (dcgMLST) schema, enabling decentralized, consistent strain typing by replacing complex cgMLST database structures with MD5 hashing values [38▪▪]. Using Neisseria strains, the authors used dcgMLST along with HierCC to show how this approach can identify sub-lineages within potential transmission networks as well as model epidemics over time [38▪▪]. Li et al. integrated dcgMLST into KleTy, a novel K. pneumoniae whole-genome typing pipeline that modularly characterizes core and accessory genome structures, enabling a portable, unified genotyping scheme for public health laboratories [39▪▪].
King et al. compared the aforementioned gene-by-gene approaches with a k-mer based clustering approach (i.e., Population Partitioning Using Nucleotide K-mers [PopPUNK] [40]) to determine if nucleotide diversity not captured in the former approaches affect capacity to resolve S. pneumoniae sub-lineages [41▪▪]. The cgMLST based methods using HierCC and LIN lineages, along with the k-mer-based PopPUNK, show high concordance, with clustering assignments validated by phylogenetic and pan-genome analyses, making all three effective for defining S. pneumoniae population structure [41▪▪].
These aforementioned advancements in gene-by-gene typing schema, including HierCC, cgLIN, dcgMLST, and PopPUNK, have improved resolution, scalability, and standardization. Future work should determine how these typing schemes can be optimally coalesced into a standardized uniform approach.
ESTABLISHING TRANSMISSION THRESHOLDS
Regardless of whether variant calling or a gene-by-gene approach is used to establish inter-strain genetic distance, identifying potential transmission networks relies on threshold cutoffs to distinguish closely related from unrelated strains [20,42,43]. However, species-specific genetic variability presents significant challenges in establishing thresholds which neither under- or overcall transmission [44]. A major challenge in this area is the paucity of epidemiologically confirmed transmission events in which sufficient numbers of strains have been whole genome sequenced to establish clear genetic thresholds as a gold standard. Numerous investigations have identified strains with negligible genetic differences causing infections in hospitalized persons with only minimal epidemiologic links [20,43,45▪]. Whether such observations represent cryptic healthcare transmission or community acquisition of closely related strains is not currently clear. Variation in mutation rates, recombination frequency, and genome plasticity within bacterial species can impact the accuracy of inferred relatedness, while within-host evolution and laboratory processing biases further complicate genetic comparisons (N.B.: for a seminal review on how these factors affect genetic relatedness interpretation, please see [46]).
Since publication of the previous review [1], there have been efforts to determine SNP cutoff values that define outbreak transmission events [45▪,47▪▪,48▪]. Mustapha et al. sequenced over 3000 clinical isolates, including isolates collected repeatedly from the same patient as well as epidemiologically confirmed outbreak isolates, to define species specific threshold values [47▪▪]. Talbot et al. analyzed differing SNP threshold values as well as different analytic techniques to measure genetic variation within and between MRSA populations to integrate with clinical data to refine transmission detection [45▪]. Miles-Jay et al. studied Clostridioides difficile dynamics in an intensive care unit and identified a SNP threshold of 2 to define probable transmission [48▪]. Such threshold values are increasingly being integrated with infection control initiatives to attempt to identify and mitigate healthcare transmission of bacterial pathogens [20,43,49]. Nevertheless, as Mustapha et al. suggest, SNP thresholds for determining relatedness are highly influenced by species-specific genetic diversity that can be largely affected by the population structure within a local epidemiological context, hence a rigid application of a pan species specific SNP threshold should be avoided [47▪▪].
To circumvent the limitations of arbitrary SNP thresholds, Hawken et al. built on previous phylogenetic methods [50] by introducing a genomic and epidemiological approach that defines threshold-free transmission clusters linking carbapenemase positive K. pneumoniae isolates to their most closely related imported strain [51▪▪]. This study indicated that based off surveillance of imported isolates at admission, their genetic distances were comparable to pairs of putative imported and acquired isolates respectively, suggesting that high false positives can occur in high endemic areas and threshold free approaches can mitigate these errors through identification of a putative common ancestor with relevant epidemiological data [51▪▪]. Another modelling study of foodborne outbreaks established a conceptual model to estimate and update genetic thresholds based on mutation rate (μ), duration of outbreak (D), and sampling dates to simulate and test capacity to distinguish outbreak from nonoutbreak cases [52▪▪]. The authors found that sensitivity remained high across simulations which is consistent with other studies in which the ability of WGS data to rule out transmission is quite robust. However, low μ and D values correlated with lower specificity, or the capacity of a low genetic distance between strains to accurately identify a transmission event [52▪▪]. Applying dynamic, locally informed SNP thresholds, guided by conceptual models that incorporate mutation rates and epidemiological data, could enhance hospital outbreak surveillance by improving the accuracy and precision of outbreak-related case identification.
METAGENOMICS
Advancements in metagenomic sequencing and metagenome-assembled genomes (MAGs) are enhancing the ability to compare and track strains within diverse microbial communities without the need for individual strain isolation, although many challenges remain [53,54]. There have been great strides in the past five years to create in silico tools that can track individual strains within complex communities [55,56▪▪,57▪,58]. In a seminal work, Valles-Colomer et al. used StrainPhlAn4 [55] to identify transmission of particular strains within the microbiome of household members via a metagenomic approach [59▪]. Zhou et al. introduced a reference-based microbial profiling pipeline for longitudinal metagenomic data that detects genome-wide SNVs and estimates strain proportions [56▪▪]. They validated their pipeline through simulations and real-world datasets and found that their approach outperformed existing metagenomic genotyping and deconvolution tools [56▪▪]. StrainScan is a novel strain-level composition analysis tool that leverages a tree-based k-mer indexing structure to improve strain identification accuracy and computational efficiency [57▪]. Advances in long-read metagenomics and improved deconvolution methods should make metagenomic strain typing a more widely adopted tool for surveillance in the future, particularly in longitudinal patient or environmental sampling.
CONCLUSION
With increasing affordability and accessibility, WGS has clearly emerged as the gold standard for bacterial strain classification and comparison. There have been significant recent strides made towards creating portable, scalable, and standardized bacterial typing methods using WGS data. Figure 1 provides an overview of how both variant calling and gene-by-gene approaches are moving towards more unified genotyping systems that are interchangeable across local, state, and national laboratories. Despite these gains, there remain challenges in agreed upon nomenclature as well as typing methodologies that can be utilized by health professionals other than skilled computational biologists/bioinformaticians. Comprehensive bioinformatics platforms, like the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), which integrates genotyping, annotation, phylogenetics, and AMR detection, is advancing user-friendly tools for standardized bacterial analysis for the wider public health community [60▪▪]. We anticipate in the next few years there will be increasing integration of bacterial strain typing with artificial intelligence as a means to both analyze and predict bacterial outbreaks [43,61]. There likely will always remain sampling biases that will preclude researchers from fully accounting for intrinsic heterogeneity of bacterial populations collected from both living and environmental samples. Like Heisenberg's uncertainty principle, bacterial genomes resist precise, static definitions – an important consideration for tracing transmission chains. Nevertheless, along with improved standardization of genotyping nomenclature, future work should strive towards improving the capture of this heterogeneity in order to improve robust pathogen surveillance networks.
FIGURE 1.
Advances and challenges in bacterial strain typing using whole-genome sequencing (WGS). Figure provides an overview of (1) WGS technologies that can be used for bacterial strain typing; (2) illustration of discriminatory power for each bacterial strain typing approach; (3) bullet-points on challenges, emerging solutions, and public health & research applications of bacterial strain typing with arrows indicating how each are interconnected. Figure was created with BioRender.
Acknowledgements
None.
Financial support and sponsorship
S.A.S. and B.M.H. receive partial support through the National Institute of Allergies and Infectious Diseases (NIAID) P01AI152999 grant.
Conflicts of interest
There are no conflicts of interest.
REFERENCES AND RECOMMENDED READING
Papers of particular interest, published within the annual period of review, have been highlighted as:
▪ of special interest
▪▪ of outstanding interest
REFERENCES
- 1.Simar SR, Hanson BM, Arias CA. Techniques in bacterial strain typing: past, present, and future. Curr Opin Infect Dis 2021; 34:339–345. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Franco-Duarte R, Černáková L, Kadam S, et al. Advances in chemical and biological methods to identify microorganisms—from past to present. Microorganisms 2019; 7:130. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Lancefield RC. A serological differentiation of human and other groups of hemolytic streptococci. J Exp Med 1933; 57:571–595. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Craigie J, Yen CH. The demonstration of types of B. Typhosus by means of preparations of type II Vi phage: I. principles and technique. Can Public Health J 1938; 29:448–463. [Google Scholar]
- 5.Uelze L, Grützke J, Borowiak M, et al. Typing methods based on whole genome sequencing data. One Health Outlook 2020; 2:3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Ribot EM, Freeman M, Hise KB, Gerner-Smidt P. PulseNet: entering the age of next-generation sequencing. Foodborne Pathog Dis 2019; 16:451–456. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Maiden MC, Bygraves JA, Feil E, et al. Multilocus sequence typing: a portable approach to the identification of clones within populations of pathogenic microorganisms. Proc Natl Acad Sci USA 1998; 95:3140–3145. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Jolley KA, Maiden MC. BIGSdb: scalable analysis of bacterial genome variation at the population level. BMC Bioinformatics 2010; 11:595. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Reinert K, Langmead B, Weese D, Evers DJ. Alignment of next-generation sequencing reads. Annu Rev Genomics Hum Genet 2015; 16:133–151. [DOI] [PubMed] [Google Scholar]
- 10.Saada B, Zhang T, Siga E, et al. Whole-genome alignment: methods, challenges, and future directions. Appl Sci 2024; 14:4837. [Google Scholar]
- 11.Bush SJ, Foster D, Eyre DW, et al. Genomic diversity affects the accuracy of bacterial single-nucleotide polymorphism – calling pipelines. Gigascience 2020; 9:giaa007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12▪.Marin M, Vargas R, Harris M, et al. Benchmarking the empirical accuracy of short-read sequencing across the M. tuberculosis genome. Bioinformatics 2022; 38:1781–1787. [DOI] [PMC free article] [PubMed] [Google Scholar]; Optimizing Illumina variant calling for M. tuberculosis enhances precision and recall while refining low-confidence and repeat regions. Future studies should assess the robustness of these optimization parameters across diverse M. tuberculosis sample sets.
- 13.Walter KS, Colijn C, Cohen T, et al. Genomic variant-identification methods may alter Mycobacterium tuberculosis transmission inferences. Microb Genom 2020; 6:mgen000418. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14▪.Seah YM, Stewart MK, Hoogestraat D, et al. In silico evaluation of variant calling methods for bacterial whole-genome sequencing assays. J Clin Microbiol 2023; 61:e0184222. [DOI] [PMC free article] [PubMed] [Google Scholar]; Authors created a short-read variant simulation pipeline to evaluate aligner/variant caller combinations. Pipelines that accounted for identifying base mismatches and high quality soft-clipped reads to identify insertions performed the best, highlighting optimal variant calling parameterization.
- 15▪.Taouk ML, Featherstone LA, Taiaroa G, et al. Exploring SNP filtering strategies: the influence of strict vs soft core. Microb Genom 2025; 11:001346. [DOI] [PMC free article] [PubMed] [Google Scholar]; An important consideration is which positions of a reference should be included as one balances sensitivity and specificity of variant calling. The authors found that more conservative, strict definitions of core sites limited phylogenetic signal, favoring a ‘soft-core’ alignment strategy, which may yield a more robust phylogenetic signal.
- 16.Lee H, Kim J, Lee J. Benchmarking datasets for assembly-based variant calling using high-fidelity long reads. BMC Genomics 2023; 24:148. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17▪▪.Hall MB, Wick RR, Judd LM, et al. Benchmarking reveals superiority of deep learning variant callers on bacterial nanopore sequence data. eLife 2024; 13: [DOI] [PMC free article] [PubMed] [Google Scholar]; This key study showed that ONT long-read variant calling, especially using convolutional neural network (CNN)-based methods, can outperform short-read approaches across diverse bacterial strains. High precision and recall were maintained even at 10× coverage, suggesting efficient data use and increased multiplexing potential, with further improvements expected from bacterial-specific CNN model training.
- 18.Quainoo S, Coolen JPM, Van Hijum SAFT, et al. Whole-genome sequencing of bacterial pathogens: the future of nosocomial outbreak analysis. Clin Microbiol Rev 2017; 30:1015–1063. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19▪▪.Bogaerts B, Van den Bossche A, Verhaegen B, et al. Closing the gap: Oxford Nanopore Technologies R10 sequencing allows comparable results to Illumina sequencing for SNP-based outbreak investigation of bacterial pathogens. J Clin Microbiol 2024; 62:e01576–e1623. [DOI] [PMC free article] [PubMed] [Google Scholar]; A long-read variant calling study of E. coli and L. monocytogenes outbreaks found that ONT R10 data produced phylogenies comparable to Illumina, with fewer filtering challenges than older R9 data. Mixed R10/Illumina datasets also yielded consistent phylogenetic signals, supporting the use of core genome alignment with combined data from newer ONT chemistries.
- 20.Wu C-T, Shropshire WC, Bhatti MM, et al. Rapid whole genome characterization of antimicrobial-resistant pathogens using long-read sequencing to identify potential healthcare transmission. Infect Control Hosp Epidemiol 2024; 1–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Zheng Z, Li S, Su J, et al. Symphonizing pileup and full-alignment for deep learning-based long-read variant calling. Nat Comput Sci 2022; 2:797–803. [DOI] [PubMed] [Google Scholar]
- 22.Poplin R, Chang P-C, Alexander D, et al. A universal SNP and small-indel variant caller using deep neural networks. Nat Biotechnol 2018; 36:983–987. [DOI] [PubMed] [Google Scholar]
- 23▪.Charron P, Kang M. VariantDetective: an accurate all-in-one pipeline for detecting consensus bacterial SNPs and SVs. Bioinformatics 2024; 40: [DOI] [PMC free article] [PubMed] [Google Scholar]; The inclusion of multiple variant callers improved accuracy of calls in this simulation study. This approach to leverage advantages of different variant callers may be useful approach in contrast to reliance on single, variant calling strategies.
- 24▪.Derelle R, Von Wachsmann J, Mäklin T, et al. Seamless, rapid, and accurate analyses of outbreak genomic data using split k-mer analysis. Genome Res 2024; 34:1661–1673. [DOI] [PMC free article] [PubMed] [Google Scholar]; This article provides an update on the reference free, split k-mer approach to variant calling (SKA2) which improves scalability and ability for future maintenance of the tool. Reference-free approaches such as SKA2 remove need for masking and appropriate species/sub-species reference, which improves usability.
- 25▪.Kille B, Nute MG, Huang V, et al. Parsnp 2.0: scalable core-genome alignment for massive microbial datasets. Bioinformatics 2024; 40: [DOI] [PMC free article] [PubMed] [Google Scholar]; The authors significantly improved the computational time for their core genome alignment tool, parsnp, which scales with large (i.e., >200 genomes) cohorts. Scalability and functionality with other ‘Harvest’ suite tools (i.e., alignment visualization tool, Gingr) creates a user-friendly interface.
- 26.Mellmann A, Harmsen D, Cummings CA, et al. Prospective genomic characterization of the German enterohemorrhagic Escherichia coli O104:H4 outbreak by rapid next generation sequencing technology. PLoS ONE 2011; 6:e22751. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Sheppard SK, Jolley KA, Maiden MCJ. A gene-by-gene approach to bacterial population genomics: whole genome MLST of campylobacter. Genes 2012; 3:261–277. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Silva M, Machado MP, Silva DN, et al. chewBBACA: a complete suite for gene-by-gene schema creation and strain identification. Microb Genom 2018; 4:e000166. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Jolley KA, Bray JE, Maiden MCJ. Open-access bacterial population genomics: BIGSdb software, the PubMLST.org website and their applications. Wellcome Open Res 2018; 3:124. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Zhou Z, Alikhan N-F, Mohamed K, et al. The EnteroBase user's guide, with case studies on Salmonella transmissions, Yersinia pestis phylogeny, and Escherichia core genomic diversity. Genome Res 2020; 30:138–152. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31▪.Dyer NP, Pauker B, Baxter L, et al. EnteroBase in 2025: exploring the genomic epidemiology of bacterial pathogens. Nucleic Acids Res 2025; 53:D757–D762. [DOI] [PMC free article] [PubMed] [Google Scholar]; The developers of Enterobase provided an update on functionality of their graphical user interface webserver which incorporates new HierCC nomenclature into their bacterial strain typing applications.
- 32▪▪.Mamede R, Vila-Cerqueira P, Silva M, et al. Chewie Nomenclature Server (chewie-NS): a deployable nomenclature server for easy sharing of core and whole genome MLST schemas. Nucleic Acids Res 2021; 49:D660–D666. [DOI] [PMC free article] [PubMed] [Google Scholar]; Chewie-NS is an important, open source, portable schema database that has allowed users to submit novel cgMLST allelic profiles. While more improved portable methods to share schema have been established, Chewie-NS schemes are still widely distributed and important for building novel species’ cgMLST databases.
- 33▪▪.Zhou Z, Charlesworth J, Achtman M. HierCC: a multilevel clustering scheme for population assignments based on core genome MLST. Bioinformatics 2021; 37:3645–3646. [DOI] [PMC free article] [PubMed] [Google Scholar]; This pivotal study introduced a scalable, hierarchical clustering framework for gene-by-gene classification, offering a robust alternative to traditional single-level approaches used in MLST and cgMLST. This study has importantly moved the field towards multilevel clustering that can better define closely related organisms and accounts for increasing amounts of data.
- 34▪▪.Hennart M, Guglielmini J, Bridel S, et al. A dual barcoding approach to bacterial strain nomenclature: genomic taxonomy of Klebsiella pneumoniae strains. Mol Biol Evol 2022; 39:msac135. [DOI] [PMC free article] [PubMed] [Google Scholar]; A novel, dual genetic identifier (barcoding) approach improved multilevel pathogen clustering, with multilevel single linkage clustering enabling backward compatibility with traditional MLST designations. This system incorporates stable Life Information (LIN) codes, allowing for more scalable and future-proof classification as genomic datasets expand.
- 35.Hu F, Pan Y, Li H, et al. Carbapenem-resistant Klebsiella pneumoniae capsular types, antibiotic resistance and virulence factors in China: a longitudinal, multicentre study. Nat Microbiol 2024; 9:814–829. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Selvaraj Anand S, Wu CT, Bremer J, et al. Identification of a novel CG307 sub-clade in third-generation-cephalosporin-resistant Klebsiella pneumoniae causing invasive infections in the USA. Microb Genom 2024; 10:001201. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Ikhimiukor OO, Zac Soligno NI, Akintayo IJ, et al. Clonal background and routes of plasmid transmission underlie antimicrobial resistance features of bloodstream Klebsiella pneumoniae. Nat Commun 2024; 15:6969. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38▪▪.Zhong L, Zhang M, Sun L, et al. Distributed genotyping and clustering of Neisseria strains reveal continual emergence of epidemic meningococcus over a century. Nat Commun 2023; 14: [DOI] [PMC free article] [PubMed] [Google Scholar]; This study presented a novel hash-based approach to enable classification of allelic profiles without requiring a centralized sequence database. Distributed cgMLST (dcgMLST) profiles are compact, easily transferable, and can be converted back into interpretable allelic sets across laboratories, supporting clustering systems like HierCC.
- 39▪▪.Li H, Liu X, Li S, et al. KleTy: integrated typing scheme for core genome and plasmids reveals repeated emergence of multidrug resistant epidemic lineages in Klebsiella worldwide. Genome Med 2024; 16: [DOI] [PMC free article] [PubMed] [Google Scholar]; A proof-of-principle study applying dcgMLST and HierCC to K. pneumoniae revealed plasmid sharing networks driving antimicrobial resistance across pathogenic clusters, highlighting the importance of chromosome–plasmid co-evolution in tracking highly transmissible pathogens.
- 40.Lees JA, Harris SR, Tonkin-Hill G, et al. Fast and flexible bacterial genomic epidemiology with PopPUNK. Genome Res 2019; 29:304–316. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41▪▪.King AC, Kumar N, Mellor KC, et al. Comparison of gene-by-gene and genome-wide short nucleotide sequence-based approaches to define the global population structure of Streptococcus pneumoniae. Microb Genom 2024; 10:001278. [DOI] [PMC free article] [PubMed] [Google Scholar]; A meta-analysis highlighted the need to move beyond traditional MLST typing, advocating for the integration of LIN codes, hierarchical cgMLST clustering, and reference-free methods like PopPUNK to improve bacterial classification systems.
- 42.Coll F, Raven KE, Knight GM, et al. Definition of a genetic relatedness cutoff to exclude recent transmission of meticillin-resistant Staphylococcus aureus: a genomic epidemiology analysis. Lancet Microbe 2020; 1:e328–e335. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Sundermann AJ, Chen J, Kumar P, et al. Whole-genome sequencing surveillance and machine learning of the electronic health record for enhanced healthcare outbreak detection. Clin Infect Dis 2022; 75:476–482. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Schürch AC, Arredondo-Alonso S, Willems RJL, Goering RV. Whole genome sequencing options for bacterial strain typing and epidemiologic analysis based on single nucleotide polymorphism versus gene-by-gene–based approaches. Clin Microbiol Infect 2018; 24:350–354. [DOI] [PubMed] [Google Scholar]
- 45▪.Talbot BM, Jacko NF, Petit RA, et al. Unsuspected clonal spread of methicillin-resistant staphylococcus aureus causing bloodstream infections in hospitalized adults detected using whole genome sequencing. Clin Infect Dis 2022; 75:2104–2112. [DOI] [PMC free article] [PubMed] [Google Scholar]; The authors tested three bioinformatic pipelines with two SNP thresholds and noted that clusters of MRSA collected over a year long period were detected with epidemiological links with patients that spanned months prior or without any epidemiological link detected suggesting missing links (e.g., transmission from healthcare workers) and/or community spread.
- 46.Didelot X, Walker AS, Peto TE, et al. Within-host evolution of bacterial pathogens. Nat Rev Microbiol 2016; 14:150–162. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47▪▪.Mustapha MM, Srinivasa VR, Griffith MP, et al. Genomic diversity of hospital-acquired infections revealed through prospective whole-genome sequencing-based surveillance. Msystems 2022; 7:e01384–e1421. [DOI] [PMC free article] [PubMed] [Google Scholar]; The authors were able to distinguish closely related species from a large cohort using either average nucleotide identity or PCA approaches. Importantly, when looking at between versus within species pairwise nucleotide variation, they noted large core genomic distance fluctuations across species, suggesting mutation rate should be accounted for when determining outbreak thresholds.
- 48▪.Miles-Jay A, Snitkin ES, Lin MY, et al. Longitudinal genomic surveillance of carriage and transmission of Clostridioides difficile in an intensive care unit. Nat Med 2023; 29:2526–2534. [DOI] [PMC free article] [PubMed] [Google Scholar]; In this outbreak analysis of Clostridioides difficile, using a stringent SNP threshold and epidemiological data suggested that cross-contamination within hospitals was unlikely, highlighting the importance of preventing progression from colonization to infection.
- 49.Blane B, Raven KE, Brown NM, et al. Evaluating the impact of genomic epidemiology of methicillin-resistant Staphylococcus aureus (MRSA) on hospital infection prevention and control decisions. Microb Genom 2024; 10:001235. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Snitkin ES, Won S, Pirani A, et al. Integrated genomic and interfacility patient-transfer data reveal the transmission pathways of multidrug-resistant Klebsiella pneumoniae in a regional outbreak. Sci Transl Med 2017; 9:eaan0093. [DOI] [PubMed] [Google Scholar]
- 51▪▪.Hawken SE, Yelin RD, Lolans K, et al. Threshold-free genomic cluster detection to track transmission pathways in health-care settings: a genomic epidemiology analysis. Lancet Microbe 2022; 3:e652–e662. [DOI] [PMC free article] [PubMed] [Google Scholar]; This threshold-free study on healthcare-associated infections integrated phylogenetic and epidemiological data to refine clustering of closely related isolates. Standardizing such approaches, which move beyond fixed SNV cutoffs, could enhance our ability to monitor strain transmission across diverse settings.
- 52▪▪.Duval A, Opatowski L, Brisse S. Defining genomic epidemiology thresholds for common-source bacterial outbreaks: a modelling study. Lancet Microbe 2023; 4:e349–e357. [DOI] [PMC free article] [PubMed] [Google Scholar]; A foodborne outbreak study used epidemiological and microbiological data to define genetic distance thresholds for identifying clusters. Similar modeling approaches that incorporate mutation rates and other relevant epidemiological data could improve interpretation of hospital-acquired infection transmission by accounting for species- or subspecies-specific variation, in contrast to utilizing single universal SNP thresholds, which may not account properly for cohort specific population structure.
- 53.Meziti A, XXX Rodriguez-R LM, Hatt JK, et al. The reliability of metagenome-assembled genomes (MAGs) in representing natural populations: insights from comparing MAGs against isolate genomes derived from the same fecal sample. Appl Environ Microbiol 2021; 87:e02593-20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Chen L-X, Anantharaman K, Shaiber A, et al. Accurate and complete genomes from metagenomes. Genome Res 2020; 30:315–333. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Blanco-Míguez A, Beghini F, Cumbo F, et al. Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. Nat Biotechnol 2023; 41:1633–1644. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56▪▪.Zhou B, Wang C, Putzel G, et al. An integrated strain-level analytic pipeline utilizing longitudinal metagenomic data. Microbiol Spectr 2024; 12:e01431–e1524. [DOI] [PMC free article] [PubMed] [Google Scholar]; LongStrain is a novel pipeline designed to improve genome-wide variant calling and strain proportion estimation in longitudinal metagenomic data, outperforming or matching existing tools in both accuracy and efficiency, particularly in host-associated microbiomes with dominant strains. However, its performance is limited in complex environments with multiple abundant strains or frequent strain switching, where accurate deconvolution remains a challenge, highlighting the need for further algorithmic refinement of metagenomic genotyping tools.
- 57▪.Liao H, Ji Y, Sun Y. High-resolution strain-level microbiome composition analysis from short reads. Microbiome 2023; 11:183. [DOI] [PMC free article] [PubMed] [Google Scholar]; The authors presented a k-mer-based method for clustering metagenomic samples which enabled more accurate pathogen identification, though it remains computationally intensive and requires high coverage to distinguish closely related organisms.
- 58.Olm MR, Crits-Christoph A, Bouma-Gregson K, et al. Banfield JF. inStrain profiles population microdiversity from metagenomic data and sensitively detects shared microbial strains. Nat Biotechnol 2021; 39:727–736. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59▪.Valles-Colomer M, Blanco-Míguez A, Manghi P, et al. The person-to-person transmission landscape of the gut and oral microbiomes. Nature 2023; 614:125–135. [DOI] [PMC free article] [PubMed] [Google Scholar]; This was an interesting microbiome study that examined person-to-person strain composition similarity as a means to determine transmission rates of microbial communities in settings with close personal contact.
- 60▪▪.Olson RD, Assaf R, Brettin T, et al. Introducing the Bacterial and Viral Bioinformatics Resource Center (BV-BRC): a resource combining PATRIC, IRD and ViPR. Nucleic Acids Res 2023; 51:D678–D689. [DOI] [PMC free article] [PubMed] [Google Scholar]; The Bacterial and Viral Bioinformatics Resource Center (BV-BRC) is a user-friendly web server that has multiple database, typing, phylogenetic, and other comparative genomics applications that can be readily used and shared across research communities. Adopting standardized bacterial strain typing nomenclature with these platforms holds much promise.
- 61.Liu CC, Hsiao WW. Machine learning reveals the dynamic importance of accessory sequences for Salmonella outbreak clustering. mBio 2025; e02650–e2724. [DOI] [PMC free article] [PubMed] [Google Scholar]


