Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Mar 9.
Published in final edited form as: Nat Microbiol. 2025 Oct 22;10(12):3258–3271. doi: 10.1038/s41564-025-02164-8

Long-read metagenomics for strain tracking after fecal microbiota transplant

Yu Fan 1, Mi Ni 1, Varun Aggarwala 1, Edward A Mead 1, Magdalena Ksiezarek 1, Lei Cao 1, Michael A Kamm 2, Thomas J Borody 3, Sudarshan Paramsothy 4, Nadeem O Kaakoush 5, Ari Grinspan 6, Jeremiah J Faith 1, Gang Fang 1,#
PMCID: PMC12967305  NIHMSID: NIHMS2142297  PMID: 41125958

Abstract

Accurate tracking of bacterial strains that stably engraft in fecal microbiota transplant (FMT) recipients is critical for understanding the determinants of strain engraftment, evaluating correlations with clinical outcomes, and guiding development of therapeutic consortia. While short-read sequencing has advanced FMT research, it faces challenges in strain-level de novo metagenomic assembly. Here, we described LongTrack, a method that uses long-read metagenomic assemblies for FMT strain tracking. LongTrack shows higher precision and specificity than short-read approaches especially when multiple strains co-exist in the same sample. We uncovered 648 of engrafted strains across six FMT cases of recurrent Clostridioides difficile infection and inflammatory bowel disease patients. Furthermore, long reads enabled assessment of the genomic and epigenomic stability of engrafted strains at the 5-year follow-up timepoint, revealing structural variations that may be associated with strain adaptation in a new host environment. Our findings support the use of long-read metagenomics to track microbial strains and their adaptations.


Fecal microbiota transplant (FMT) has emerged as a therapeutic strategy, transforming the treatment paradigms for patients with recurrent Clostridioides difficile infection (rCDI)1,2. By transferring gut microbiota from a healthy donor into a patient, FMT restores the patient’s gut microbial ecology3,4. Potential applications of FMT extend beyond rCDI onto treating a broader array of diseases5,6, such as inflammatory bowel disease (IBD)7,8 and cancer immunotherapy9,10 .

The FMT procedure operates on the principle of transferring gut microbiota, which constitutes diverse bacterial strains, the functional impact units of the gut microbiota11,12, from a healthy donor into a recipient3,13. Precise and comprehensive tracking of these strains within the recipient’s gut is essential for understanding factors that facilitate or hinder colonization4,13,14, providing vital information on how specific strains correlate with clinical outcomes14,15, which can guide the development of defined bacterial consortia, composed of strains with beneficial effects, for safer therapeutic applications1517.

High-throughput metagenomic shotgun sequencing studies have provided valuable insights into FMT strain engraftment in various diseases14,15,1720 primarily through short-read sequencing to detect single-nucleotide polymorphisms (SNPs) and/or gene clusters that track strains in the recipient after FMT18,19,21,22. A recent study using genome sequences from an extensive bacterial isolate collection15 for strain tracking showed that short-read approaches face challenges when multiple strains co-exist in a single microbiome sample2325. Once individual isolate genomes from donor and pre-FMT recipients are assembled, reliable tracking can be achieved using a strain’s unique k-mers in short-read metagenomic data15. However, standard culturing methods often miss certain bacterial taxa due to selective culture media and/or conditions26, and pose challenges for scaling up large cohorts. Long-read sequencing empowers de novo genome assembly of complete, strain-level resolution metagenome-assembled genomes (MAGs)2730, greatly complementing culture and short-read approaches31,32. With this rationale, we developed LongTrack, a rigorous method tailored to reliably use long-read MAGs for strain tracking (Fig. 1). Building upon systematically cultured bacterial strains from FMT donors and recipients for rCDI with 5-year follow-ups15 (Fig. 1A) as ground truth for evaluation (Fig. 1B) , we showed that LongTrack has higher precision and specificity than short-read approaches, especially in distinguishing co-existing strains within a microbiome sample (Fig. 1C). Additionally, taking the advantages of strain-level long read mapping31,32 and direct detection of bacterial DNA methylation3337, we identified structural variations (SVs) and epigenetic changes within strains as they transferred from donors to post-FMT recipients during follow-up (Fig. 1D).

Figure 1. Description of FMT samples, matched bacteria isolates and the core design of LongTrack for reliable FMT strain tracking and detecting intra-strain structural variations in engrafted strains.

Figure 1.

A, Stool samples were collected from three FMT donors and four FMT recipients at multiple time points over five years. Extensive bacterial culture, isolate genome assemblies, and short-read metagenomic sequencing have been performed in the previous study by Aggarwala et al and shown along with FMT samples. Long-read metagenomic sequencing data were generated for this study. B, Long-read data from donors’ and pre-FMT recipients’ samples were used to construct de novo metagenome-assembled genomes (long-read MAGs). The overlapping isolate genomes enable a systematic evaluation of the reliability of long-read MAGs, which serves as the foundation of LongTrack. C, Schematics of the core design of LongTrack for FMT strain tracking based on strain-level long-read (LR) MAGs. Strain-unique k-mers (k=31) were identified for each strain-level LR MAG from donor or pre-FMT recipient as the markers by eliminating common k-mers present in public database, or the other MAGs in the same FMT donor and recipient (Methods). To ensure specificity, strain-unique k-mers are only selected from shared regions across multiple strain-level LR MAGs within the same species. The presence of these final unique k-mers of a specific MAG (above a confidence score, Methods) in a post-FMT recipient’s short-read metagenomic data indicates that the corresponding MAG is present (green color) in the post-FMT recipient sample. D, Long-read metagenomic sequencing enables strain-specific detection of structure variations (SVs) that differ between donor and post-FMT samples. The references for putative genotypes were created and used to map and quantify specific SVs in each sample, allowing tracking of SV changes between hosts and/or over time. The long-read approach has a unique advantage in strain-level read mapping that allows reliable detection of intra-strain variations.

Results

High purity, strain-level long-read MAGs serve as a solid foundation for LongTrack.

The success of strain tracking in FMT depends on obtaining strain-resolved and high purity genomes from both the donor and the recipient before FMT, allowing differentiation among co-existing strains of the same species.

Bacterial isolation is an effective method for achieving this, as sequencing individual colonies yield discrete genomes with high purity and completeness. With this rationale, we built on the collection of 1000+ bacteria isolates with individually assembled genomes as ground truth15 to evaluate the long-read MAGs, and their accuracy in FMT strain tracking. We selected four FMT cases of rCDI (3 donors and 4 recipients, Supp Table S1) from which, 184 unique strains (105 species) were isolated with individually assembled genomes15 (Extended Data Fig. 1). We generated long-read sequencing data for 15 samples (Supp Table S2), and deeper short-read for two samples to enable a systematic evaluation along with LongTrack (Supp Table S3).

To construct strain-level long-read MAGs, we employed a pipeline tailored for long-read sequencing data, involving metagenomic assembly, haplotype phasing, and strain-level refinement (Extended Data Fig. 2, Methods). From four cases, we obtained a total of 460 long-read MAGs (460 strains across 312 species) (Supp Table S4), including 60 strains shared with cultured isolates (Fig. 2A). As Donor 1 (D1) had the most bacterial isolates (136 unique strains), we focused on D1 to demonstrate the advantage of long reads for de novo assembly of MAGs compared to short reads (Supp Table S3). We obtained 135 short-read MAGs from D1 (Fig. 2B) and 20 short-read MAGs from R1 using metaSPAdes38, which demonstrates better performance compared to MEGAHIT39 (Supp Fig. S1). Contamination analysis with CheckM40 for 460 long-read MAGs (from all 4 cases) and 155 short-read MAGs (from D1 and R1) revealed lower contamination and higher contig N50 values in long-read MAGs than short-read MAGs (Fig. 2C-D). Specifically, for D1, the assemblies yielded 124 long-read MAGs and 135 short-read MAGs with comparable completeness (Extended Data Fig. 3A). Although short-read data had the advantage in sequencing depth, long-read MAGs still consistently exhibit significantly lower contamination rates (Extended Data Fig. 3B) and larger N50 values (Extended Data Fig. 3C). These findings are consistent with previous studies that demonstrated the advantage of long-read sequencing for de novo metagenomic assembly across many microbiome samples24,2729,31.

Figure 2. High-purity, strain-level long-read MAGs serve as a solid foundation for LongTrack.

Figure 2.

A, The Venn diagram displays the number of strains that were recovered by long-read (LR) MAGs (yellow color) or isolate genomes from bacterial culture (red color) from 3 donors and 4 recipients, with the left bar plot illustrating examples of species with MAGs with modest-to-high relative abundance but were missed by culture, and with right bar plot illustrating examples of isolates that were cultured but not recovered by long-read sequencing because of their very low abundances. B, The Venn diagram displays the number of strains that can be recovered by LR MAGs (yellow color), short-read (SR) MAGs (blue) or isolate genomes (red) in Donor 1 (D1). LR MAGs include all the 23 strains recovered by SR MAGs, along with 8 additional strains not assembled by short reads. C, Contamination levels of 460 LR MAGs (from Donor 1–3 and Recipient 1–4) and 155 SR MAGs (from Donor 1 and Recipient 1) assessed using CheckM, showing significantly lower contamination in LR MAGs (p-value = 0.0048; two-sided Wilcoxon rank sum test; p-value < 0.01, **). D, Distributions of relative abundance (log scale) and contig N50 (kb) for LR MAGs (from Donor 1–3 and Recipient 1–4) and SR MAGs (from D1). E, Order-level taxonomic comparison of LR MAGs, SR MAGs, and isolate genomes from D1. F, Contamination level for 60 LR MAGs (from Donor 1–3 and Recipient 1–4) and 23 SR MAGs (from D1), evaluated across various species with a single strain or multiple strains. LR MAGs consistently show lower contamination (p-value = 0.0007 for single strain; p-value = 0.0086 for multiple strain; two-sided permutation test on the mean difference, n = 10,000 permutations; n=10000; p-value < 0.01, **). Contamination levels were assessed by mapping long-read and short-read MAGs to their corresponding isolate genomes and calculating the percentage of unaligned bases as a measure of contamination. G, LR sequencing has a unique advantage to resolve multiple co-existing strains from the same species over SR sequencing.

In our in-depth evaluation, we compared these MAGs to the isolate genomes. For D1, 31 and 23 isolate genomes were shared with long-read MAGs and short-read MAGs, respectively (Fig. 2B). Long-read MAGs include all 23 MAGs recovered by short-read MAGs along with 8 additional strains (Fig. 2B). Both long-read and short-read MAGs spanned the same 17 taxonomic orders, with 8 detected only by MAGs (Fig. 2E). Additionally, all MAGs overlapping with isolate genomes, including 60 long-read MAGs (across 4 cases) and 23 short-read MAGs (from D1), were used to evaluate the reliability of long-read MAGs in comparison with short-read MAGs (Supp Table S5). First, we assessed their contamination levels by mapping MAGs to their corresponding isolate genomes (Fig. 2F, Extended Data Fig. 4A), which provide more reliable insights than estimates by CheckM40 (Extended Data Fig. 4B). The contamination levels were lower for long-read MAGs (1.34%) compared to short-read MAGs (4.77%), for both species with a single and co-existing strains (Fig. 2F, Extended Data Fig. 4C-D). We further evaluated the k-mer similarity between MAGs and their corresponding isolate genomes, which consistently indicated long-read MAGs exhibit a higher similarity (Extended Data Fig. 4E). Importantly, long-read sequencing effectively recovered the same number of strains as bacterial culture from the same donor sample while these co-existing strains were not completely recovered in the short-read MAGs (Fig. 2G). We also identified 11 plasmids from the 60 shared isolate genomes of which 7 plasmids (Supp Table S6), including one circular, were recovered from long-read sequencing. However, none of the 7 plasmids were correctly binned with their matching chromosomes, highlighting that plasmid binning remains a challenge and need more specialized methods such as Hi-C41. While the shared strains demonstrate the reliability of long-read MAGs, the inherent methodological differences between metagenomics and bacterial isolation remain. Specifically, long-read MAGs, favor strains with modest-to-high relative abundance, such as Akkermansia muciniphila (0.63%) and Faecalibacterium prausnitzii (3.94%), while bacterial culture favors strains selected by culture media rather than abundance, such as Escherichia coli with abundance below 0.01% (Fig. 2A). In D1, 37.1% and 65.9% metagenomic reads can be mapped to the isolate genomes and long-read MAGs (Methods), respectively, with 76% mapping to the merged reference (isolate genomes and long-read MAGs) (Supp Fig. S2). Interestingly, the 31 shared strains accounted for 30.7% metagenomic reads, because of their relatively high abundance, highlighting the complementarity between long-read MAGs and bacterial culture.

LongTrack enables reliable use of Long-read MAGs for precise FMT strain tracking

Once individual isolate genomes are assembled, strain tracking can be reliably performed using unique k-mers (n=31) selected from isolate genomes and short-read sequencing of post-FMT samples, providing a ground truth for evaluating LongTrack15. This k-mer based strategy (Strainer) demonstrated15 superior performance compared to several SNP-based approaches15 including inStrain21, StrainFinder19, and ConStrains18, particularly for modest-to-low abundance strains and multiple co-existing strains. Our independent evaluation of LongTrack confirmed that the k-mer based approach is more sensitive than the SNP-based method (StrainFinder19), which tends to miss modest-to-low abundance strains (Extended Data Fig. 5).

However, Strainer is not compatible with long-read MAGs that often comprise incomplete MAGs (taxa with modest-to-low abundance) which can lead to false strain tracking. Therefore, to reliably use long-read MAGs for FMT strain tracking, we developed LongTrack, where for similar co-existing strains in one fecal sample, or two similar strains each in the donor sample and recipient pre-FMT samples, we selected unique k-mers only from their common genomic regions (Fig. 1C). This adaptation avoids unreliable k-mers that appear to be unique to a strain but are not, due to the incompleteness of long-read MAGs, highlighting a key difference between high-purity but incomplete long-read MAGs and the complete isolate genomes15 (Methods). We compared the number of unique k-mers used for strain tracking in multiple co-existing strains, using the original Strainer with high completeness isolate genomes versus LongTrack with high-purity long-read MAGs (Fig. 3A). Although LongTrack generally kept fewer unique k-mers compared to Strainer, it retained enough for reliable strain tracking in post-FMT samples (Fig 3B). In addition, we compared the sensitivity of LongTrack and culture-based strain tracking using the union of long-read MAGs and isolate genomes as the reference. Specifically, when assessing sensitivity across varying abundance ranges, LongTrack demonstrated 100% sensitivity for strains with relative abundance >0.01 and maintained 83% for 0.001 and 0.01, while culture-based strain tracking missed many taxa in these ranges. However, for ultra-low abundance strains (<1e–4), LongTrack’s sensitivity dropped sharply, while culture-based approach achieved >98% sensitivity (Extended Data Fig. 6). These results illustrate that LongTrack effectively captures strains with modest-to-high abundance, while culture-based methods better detect very low-abundance taxa. This evaluation highlights the complementarity of the two approaches for FMT strain tracking.

Figure 3. Evaluation of LongTrack for FMT strain tracking.

Figure 3.

A, Numbers of unique k-mers used for strain tracking in cases of multiple co-existing strains by the original Strainer method (with complete genomes assembled for isolates) and the LongTrack method (with long-read MAGs). Although some k-mers from shared regions are filtered out by LongTrack, enough unique k-mers remain for reliable strain tracking in post-FMT samples. B, Precision and specificity of strain tracking using LongTrack with long-read MAGs (Donor 1–3 and Recipient 1–4) versus Strainer with short-read MAGs (Donor 1). Across 3 donors and 4 recipients with multiple fecal sampling time points (326 unique strain presence/absence events), LongTrack with long-read MAGs achieved high precision (1.00) and specificity (1.00). In contrast, Strainer with short-read MAGs (n = 186) showed lower precision (0.96) and specificity (0.33). C, Presence (green) or absence (gray) of strains in post-FMT recipients determined by strain-specific unique k-mers selected from the ground truth (Strainer with isolate genomes), LongTrack with long-read MAGs, or Strainer with short-read MAGs. Discrepancies with the ground truth are indicated by an “X” marker. (C-E). The evaluation of strain tracking was performed for four FMT cases of rCDI: D1->R1 (C), D1->R2 (C), D2->R3 (D), and D3->R4 (E). F, Evaluation of LongTrack for FMT strain tracking using in silico spike-in data that represent the co-existence of multiple strains of the same species in the donor sample. Specifically, we spiked-in three strains of Phocaeicola massiliensis (Str 1–3) to D1, and a fourth strain of P. massiliensis (Str 4) to R1 pre-FMT. All four strains were added to post-FMT recipient R1. G, Evaluation of LongTrack for FMT strain tracking using in silico spike-in data that represent cases when there are high-similarity strains present in between the donor and pre-FMT samples. Specifically, we spiked-in ten pairs of strains, each with high similarity (ANI>99.5%), from ten different species, one strain to D1 and the other to pre-FMT recipient R1, similarly for all ten pairs. All twenty strains were added to post-FMT recipient R1.

We further compared precision and specificity of strain tracking based on LongTrack with long-read MAGs (across 4 cases) versus Strainer with short-read MAGs (D1), with multiple fecal sampling time points each. LongTrack achieved great precision (1.00) and specificity (1.00) (Fig. 3B) when using shared isolate genomes as the ground truth15, while Strainer with short-read MAGs had lower precision (0.96) and much lower specificity (0.33) (Fig. 3B), likely due to the false-positive tracking and missed the co-existing strains (Fig. 3C). Building upon the two well-characterized FMT cases, D1->R1 and D1-> R2, we found 100% consistency of strain tracking results between LongTrack with long-read MAGs and ground truth for all 31 shared strains (Fig. 3C). In contrast, Strainer with short-read MAGs correctly track 23 of the 31 strains (74%) but failed to track the other 8 strains in post-FMT samples as it could not resolve the multiple strains co-existing within the same donor sample. In addition, Strainer with short-read MAGs produced six false positives due to high contamination. For instance, the short-read MAGs Bifidobacterium adolescentis has 8.76% contamination level, with 28% of the contamination originating from Bacteroides vulgatus within the same sample (Supp Fig. S3), leading to unreliable k-mers and false positive tracking (Fig. 3C). Applying LongTrack to two additional rCDI FMT cases: D2->R3 and D3->R4 also showed high precision (Fig. 3D-E). We performed additional evaluations with two mock FMT scenarios with in silico spike-in (Methods). In the first evaluation, we spiked-in three strains of Phocaeicola massiliensis (Str 1–3) to D1, and a fourth strain of P. massiliensis (Str 4) to pre-FMT R1; all four were present in post-FMT R1 (Fig. 3F). LongTrack recovered all four strains and accurately tracked them in post-FMT R1, while short-read MAGs only recovered two of the four strains (Fig. 3F). In the second evaluation, we spiked-in ten pairs of strains with high similarity to D1 and pre-FMT R1. All twenty strains were added to post-FMT R1 (Fig. 3G). LongTrack correctly tracked all strains in post-FMT recipients, while Strainer with short-read MAGs missed two (Fig. 3G).

LongTrack helps better understand strain engraftment in rCDI and IBD patients

Building on the reliability of LongTrack-based FMT strain tracking as evaluated, we scale up strain tracking to more comprehensively understand strain engraftment in FMT. Across all the four cases of rCDI FMT, long-read sequencing recovered 396 MAGs with diverse taxonomy distribution (Fig. 4A, Supp Table S4) that were not captured by the previous culture-based study15.

Figure 4. LongTrack uncovered hundreds of engrafted strains across six FMT cases of rCDI and IBD patients.

Figure 4.

A, Diverse taxonomy distribution of long-read MAGs that are not cultured in the previous study by Aggarwala et al. B, Proportional engraftment of donor strains (PEDS) in individual post-FMT recipients at various time points of the 5-year follow-up. C, Proportional persistence of recipient strains (PPRS) in individual post-FMT recipients at various time points of the 5-year follow-up. D & E, PEDS and PPRS analysis for the two FMT cases for IBD patients R5 (D) and R6 (E), where three colors are used to represent the strains from the two Donors (D4 and D5) or the recipient. (F & G) The presence and absence of co-existing strains in the two IBD patients (R5, F, R6, G) after FMT that originated from the two Donors (D4 and D5) or the recipient.

We performed LongTrack to measure the engraftment of donor strains in the recipients, focusing on uncultured strains (Supp Table S7-S10). We used proportional engraftment of donor strains (PEDS) to summarize strain engraftment up to five years after FMT. We found consistent and stable engraftment of donor strains (Fig. 4B), with PEDS values comparable to cultured strains (Extended Data Fig. 7A). In these individuals, we report an average PEDS of 98% of uncultured strains in the first week, which stabilized ~96% at 8 weeks and remained 63% even 5 years later. This demonstrates that FMT can lead to a stable long-term engraftment of donor’s uncultured strains in rCDI patients. We also monitored the persistence of recipient strains post-FMT using proportional persistence of recipient strains (PPRS). We observed a major reduction (47%) in PPRS of recipient strains post-FMT, both for cultured and uncultured strains, indicating a major shift in the gut microbiota following FMT (Fig. 4C, Extended Data Fig. 7B). The absence of C. difficile in post-FMT R1 is noteworthy, suggesting that the transplantation led to a more balanced and health-promoting gut microbiota (Extended Data Fig. 8A).

We examined the individual trajectories of long-read MAGs recovered strains for donor-recipient pairs (Supp Table S7-S10, Extended Data Fig. 8) and found several strains that stably engrafted in the recipient after five years potentially linked to human health and diseases. One notable example, A. muciniphila , present in D1 and D2, is a mucin-degrading bacterium linked to various health benefits and improved treatment outcomes42. It helps prevent C. difficile overgrowth through the modulation of gut microbiota and metabolites, supports gut barrier integrity and immune response43. A. muciniphila strains from D1 and D2 demonstrated stable engraftment in post-FMT recipients, with persistence observed for up to five years (Extended Data Fig. 8A-B). F. prausnitzii is another species well resolved exclusively by long-read MAGs. and related to human health through short-chain fatty acids production44 and a potential role in maintaining gut health and preventing IBD45. F. prausnitzii strain from D1 was detected in 8-week post-FMT R1 but not detected 5-year post-FMT R1 (Extended Data Fig. 8A).

Next, we performed LongTrack to expand the tracking analysis to strains from post-FMT samples and reversely tracked them in earlier time points (Methods). This approach captures strains that were low abundance in donor or pre-FMT but increased in post-FMT samples, allowing their assembly from long reads. Indeed, this design helped us identify 35 additional strains (Extended Data Fig. 9A-C). All these cases have clear evidence that they were from either the donor sample or pre-FMT samples but were not recovered by long-read sequencing due to low abundance. Interestingly, we also observed previously undetected strains recovered from post-FMT recipients, but no evidence supported their presence in any of the earlier times, neither donor nor pre-FMT (Extended Data Fig. 9D-F). For example, Acidaminococcus intestini, present only in the 5-year post-FMT R1 (Extended Data Fig. 9D), was likely acquired from the environment sometime between week 8 and year 5, unless its abundance was below detections limit in all the earlier time points.

We further applied the LongTrack to two FMT cases for the treatment of IBD7,46 (Supp Table S11), which represents pre-FMT recipients with more complex microbiome than rCDI patients. The two IBD patients (R5 and R6) each received fecal samples from six donors (D4 - D9) (Extended Data Fig. 10), as multiple donors are often used in the FMT treatment of IBD 7,46. This greater complexity of FMT provided a great opportunity to demonstrate the power of LongTrack for strain tracking. We conducted long-read sequencing on two of the donors (D4 and D5) and the two pre-FMT recipients (Supp Table S12) and recovered 188 MAGs (Supp Table S13). Out of these strains, 64% and 35% strains from the two donors successfully engrafted in the two 16-week post-FMT recipients, respectively (Figs. 4D-E, Supp Table S14-S15). We also estimated PPRS and observed that 80% and 96% of strains from pre-FMT recipients were present in 16-week post-FMT recipients (Fig. 4D-E). Compared to the FMT cases for rCDI at a similar time point (8 weeks), we observed lower PEDS and higher PPRS in the FMT cases for IBD, which is consistent with the expectation that the gut microbiome of IBD patients is more diverse and resilient upon FMT introduction of donor strains7. We observed 19 and 20 coexisting strains (from donors and pre-FMT recipients) in the two IBD post-FMT recipients (Fig. 4F-G). Notably, while nearly all strains from the pre-FMT recipients remained stable for 16-week post-FMT, some donor strains were not detectable, suggesting endogenous strains in IBD patients are resilient in the same host even after FMT.

Long-read sequencing of post-FMT samples uncovered genomic structural variations within engrafted strains

It is crucial to understand how donor strains adapt to the recipient’s microbiome environment, which can drive genomic variations - that have not been extensively studied47. Long-read sequencing has great advantages in complex SV detection and direct detection of DNA methylation3337,48, enabling the assessment of the stability and dynamics of SVs and DNA methylation in individual strains across donor and post-FMT.

A systematic analysis of long-read sequencing data of the post-FMT recipients across various time points identified 38 SVs, including 6 deletions, 20 insertions, and 12 inversions, as determined by stringent filtering criteria (Methods). Comparison of read alignments of SV regions showed that both long-read sequencing methods provided clear and consistent flanking sequence alignments (Fig. 5A), illustrating they were from the same strain, whereas short-read sequencing plots displayed several mismatches, making it challenging to reliably attribute the sequences to the same strain (Fig. 5A). Out of the 38 SVs, 12 were identified in four strains that we were able to isolate from both donor and post-FMT recipient samples (Supp Table S16), allowing us to validate these SVs. Long-read sequencing of these cultured isolates confirmed paired donor and post-FMT recipient isolates are indeed the same strain, with high genome-wide average nucleotide identity (ANI) (>99.9%) (Supp Table S17). Among the 12 SVs, 11 alternative SV genotypes (91.67%) indeed exist in the same strain as validated by long-read sequencing of the isolates from donor and post-FMT samples, demonstrating the high reliability of SV detection within the same strain using long-read sequencing (Fig. 5B, Supp Fig. S4).

Figure 5. Long reads uncovered complex structural variations in engrafted strains during the 5-year follow-up after rCDI FMT.

Figure 5.

A, IGV plots illustrating the advantage of long reads in the detection of intra-strain structural variation (SV) detection, including a 43 bp deletion in Parabacteroides merdae and a 1250 bp insertion in Alistipes putredinis, over short-read sequencing. Flanking sequences of SVs show clear and consistent alignments in all long-read IGV plots, supporting that they are from the same strain. In contrast, short-read plots display several mismatches, making it difficult to reliably determine whether the sequences belong to the same strain. B, Heatmap summarizing SVs detected in donors and post-FMT recipients at early and late time points after FMT including inversions (INV), deletions (DEL), and insertions (INS). Samples with insufficient coverage over a specific SV region are marked by an asterisk (*). SVs confirmed in isolates from both donor and recipient are marked with a blue cross (+). SVs are also categorized by their location: intergenic regions (IR) and genic regions (GR). C, Examples of SVs with changing genotype frequencies across donor and post-FMT recipient over time. (1) inversion in Bacteroides stercoris; (2) deletion in Alistipes putredinis changes in genotype frequency both in the early (E) and late (L) stage post-FMT; (3) insertion in Gemmiger formicilis with frequency changes only at late stage post-FMT. Pie charts demonstrate the genotype frequencies for donor (D1 or D2) and post-FMT recipient (R1 or R3) at early and late time points after FMT. D, A complex inversion (shufflon) involving six genotypes is located on restriction endonuclease subunit S of the RMS system in Parabacteroides merdae. Genotype frequencies vary with major changes observed for Genotype 6 i.e. absence in donor and increases after FMT. E, DNA adenine methylation (6mA) motifs in the P. merdae strain were detected across donor and post-FMT recipient with different time points; CA6mAYNNNNNNRTC corresponds to Genotype 6 of this RMS phase variant, which is only detectable after FMT. F, DNA methylation motifs detected across several single strains demonstrating epigenomic stability in the donor (D1) and post-FMT recipient (R1) over time.

Notably, five SVs were identified in TonB-dependent receptor genes or their upstream regions (Fig. 5B). which are crucial for iron transport in Gram-negative49. The deletions in these genes might be associated with immune-mediated selective pressures50. Bacteroides stercoris D1 exhibited a high-frequency inversion exclusively in susC genes in late-stage post-FMT, potentially linked to immune adaptation within new host environment (Fig. 5C). A. putredinis D1, a deletion in IS110 family transposase gene (ISCaa7) was observed with a high prevalence in post-FMT (Fig. 5C), and may affect genetic diversity, evolution, and antibiotic resistance propagation, highlighting the impact on microbial adaptability51. Moreover, in Gemmiger formicilis D2, an insertion proximal to yesO gene (Fig. 5C), encoding an ABC transporter substrate-binding protein, likely plays a significant role in nutrient uptake and metabolic processes52.

We also observed some complex SVs that involve bacterial restriction-modification systems (RMS), which safeguard bacterial genomes by cleaving unmethylated foreign DNA while methylating self DNA35,5355. These SVs can alter the activity or specificity of DNA methyltransferases (MTases), driving genome-wide methylation changes that regulate gene expression and enhance phenotypic diversity56, promoting bacterial adaptation to antibiotics and host immune responses57. Long-read sequencing facilitates simultaneous detection of SVs and DNA methylation at single-molecule resolution, enabling linkage of RMS genotypes to methylation motifs, revealing the epigenomic impact of SVs. For example, we detected a complex SV spanning a Type I RMS S subunits of P. merdae. The nested inverted repeats create a complex genomic shuffling among six different genotypes (Fig. 5D). Notably, genotype 5, predominant in the donor, decreased sharply in post-FMT, while genotype 6 was absent in the donor but significantly increased in post-FMT. Because the S subunit determines the sequence recognition motif of the MTases, the genotype switches of this RMS locus can create genome-wide methylation changes41. To test this hypothesis, we used PacBio sequencing data, which offers advantages in detecting 6-methyladenine (6mA) at single-molecule resolution36,48,54, and indeed observed a DNA methylation motif CA6mAYNNNNNNRTC exclusively in post-FMT samples (Fig. 5E). This suggests some SVs can influence genome-wide changes in methylation and potentially help P. merdae adapt to the recipient’s gut environment. Additionally, we examined 11 unique Type II methylation motifs across ten single strains with high-coverage Nanopore long-read data (Fig. 5F), including one 5-methylcytosine (5mC) motif, two as 4-methylcytosine (4mC) motifs, and eight 6mA motifs. These motifs demonstrate consistency and stability in both donors and recipients over various time points.

Discussion

In this study, we present LongTrack, a method that uses long-read sequencing for precision FMT strain tracking, addressing a critical gap in the high-resolution tracking of microbial strains.

A key methodological innovation of LongTrack is the rigorous selection of strain-unique k-mers from long-read MAGs, compatible with the high-purity yet sometime incomplete long-read MAGs (for low-abundance strains), which are different from separately assembled bacterial isolate genomes15. LongTrack includes a critical step that retains only unique k-mers originating from assembled genomic regions shared across the conspecific strains both within and between FMT donor and recipient samples. Once reliable strain-unique k-mers are selected from long-read MAGs, subsequent strain tracking can be reliably performed using short-read microbiome sequencing data. This LongTrack framework was validated against ground truth, Strainer with isolate genomes from large-scale bacterial culture15.

The previous isolate-based method served as a valuable control for methods evaluation, although standard culture-based approaches often miss certain bacterial taxa due to the selective nature of culture media. For instance, A. muciniphila and F. prausnitzii were missed under standard culture conditions, despite being culturable from stool samples under more stringent conditions58,59. Nonetheless, the inclusion of isolate genomes in our study was primarily intended to evaluate LongTrack with long-read MAGs against Strainer with short-read MAGs, particularly in resolving co-existing strains. Importantly, pure and complete isolate genome assemblies provides a solid approach to evaluate strain tracking15.

A noteworthy observation is that several MAGs with low completeness (<50%) still exhibit high ANI (>99.6%) and k-mer similarity (>92%) to their corresponding cultured isolates (Supp Table S5). For example, Bacteroides caccae partial MAG (499 kb) contained sufficient unique k-mers to enable accurate strain tracking in post-FMT samples (Fig. 3C). with LongTrack. Sequencing depth for effective strain tracking using LongTrack is flexible and influenced by microbial complexity and strain abundance. For D1, we generated 23 Gb of long-read data to maximize recovery of strains shared with cultured isolates and enable a robust comparison. However, D2 and D3 yielded consistent accurate, and reliable tracking results with only 18 Gb and 15 Gb of data, respectively.

Applying LongTrack to rCDI and IBD FMT cases demonstrated applicability across diverse microbiome complexities. Due to the relatively small sample size, this study was mostly for methods evaluation and demonstration of broad applicability, instead of generalized findings on strain engraftment. Future studies on large cohorts could further explore the influence of strain richness on engraftment dynamics and its therapeutic implications46.

This study also attempts to assess the genomic dynamics of donor strains as they adapt to new microbial and host environments in recipients. Recent studies highlighted the importance of SVs, in addition to SNVs, in microbiome adaptation60,61. However, strain-level SVs detection across the longitudinal samples poses a challenge to short-read sequencing but can be better approached by long-read sequencing. Long reads have the advantages in detecting complex SVs like shufflons, dynamic DNA inversion systems known to influence bacterial adaptation by generating structural diversity62,63. Building on longitudinal samples collected over five years following FMT, we identified 38 SVs using stringent filtering criteria (requiring >5 coverage depth for high-confidence mapping) and note that the SV analysis is dependent on sequencing depth and additional SVs may be detected with deeper sequencing data.

While LongTrack offers insights, it faces limitations in recovering ultra-low-abundance strains. To mitigate this, we propose two complementary strategies. First, integrating long-read sequencing and culture-based methods to provide a more complete view of the FMT strain engraftment. Second, applying targeted enrichment methods, such as flow cytometry64 or bacterial methylome patterns65, to selectively sequence and analyze low-abundance taxa, thereby improving their recovery and tracking.

As the cost of long-read sequencing platforms continues to decrease, we expect LongTrack to be broadly used for high-resolution strain tracking in FMT studies, enabling better understanding of FMT and the development of more effective defined biotherapeutics. More broadly, LongTrack can support precise strain tracking in other biomedical studies to enhance our understanding of strain transmission, genomic and epigenomic variation, and their roles in health and diseases66.

Methods

Human participants

This study complied with relevant ethical regulations. Protocols were approved by the Mount Sinai Institutional Review Board (rCDI: HS#11–01669; IBD: HREC/13/SVH/69). Written informed consent was obtained before sample collection and analysis, and participants were compensated. We enrolled 11 adults (≥18 years) from previous studies7,15,46: three FMT donors and four FMT recipients with rCDI, and two FMT donors and two FMT recipients with IBD. Human participant health status is available in Supplementary Table 1 and 11.

Fecal sample collection and DNA extraction

Fecal DNA was extracted with a QIAamp PowerFecal Pro DNA Kit (Qiagen, 51804) according to the manufacturer’s instructions. Briefly, 64 – 436 mg of fecal sample, and 800 μl of Solution CD1 were added to the Bead Tube (provided). After gentle vortexing, the sample was homogenized with the TissueLyser II machine (Qiagen, 85300) for 10 min at 25 Hz speed (5 min each side). The homogenized samples were centrifuged to collect the supernatant at 15,000 x g for 1 min. The subsequent column-based extraction steps were performed as described in the QIAamp PowerFecal Pro DNA Kit protocol.

Library preparation and sequencing

Nanopore long-read sequencing

Native libraries were prepared using the Oxford Nanopore Technology Inc. (New York, NY, USA) Native barcoding kit (EXP-NBD104 and 114 for R9.4.1 flow cells, and SQK-NBD114.24 for R10.4.1 flow cells), Ligation Sequencing Kit (SQK-LSK109 for R9, and SQK-LSK114 for R10.4.1). Prior to library preparation, genomic DNA was sheared using Covaris g-tubes (Covaris, Woburn, MA, USA 520079) at 6000 RPM in an Eppendorf 5424R centrifuge (Eppendorf, Enfield, CT, USA, 5406000640) and cleaned up with AMPure XP beads (Beckman Coulter, Indianapolis, IN, USA, A63882) at 0.8 x concentration. Whole genome amplification (WGA) samples were prepared using REPLI-g Mini Kits (QIAGEN, Hilden, Germany, 150023) according to the protocol with 12.5 ng of input DNA per sample and 16 h incubation at 30 °C, followed by denaturation at 65 °C for 3 minutes. Next, WGA samples were treated with T7 endonuclease I (NEB M0302S) to maximize nanopore sequencing yield according to ONT documentation, then cleaned up with AMPure XP beads as above. Subsequently, samples were processed the same as above native libraries. Samples were sequenced on R9.4.1 and R10.4.1 MinION flow cells and post-run basecalled using Guppy (v5.0.7 for R9.4.1 data and 6.3.4 for R10.4.1 data) with the SUP model dna_r9.4.1_450bps_sup.cfg and dna_r10.4.1_e8.2_400bps_sup.cfg, respectively.

Illumina short-read sequencing

Library construction for Illumina sequencing was executed utilizing the seqWell Plexwell 384 kit. DNA was systematically barcoded, and ligation products underwent purification. Size selection and PCR amplification of the DNA fragments followed. The prepared library was then sequenced on an Illumina HiSeq platform, generating 150bp paired-end reads.

PacBio long-read sequencing

DNA SMRTbell libraries were prepared according to the manufacturer’s instructions for the SMRTbell Express Template Prep Kit 2.0, then each library was annealed to sequencing primer together with the Sequel 2.1 polymerase, followed by loading to the 8M Sequel II sequencing chip and 25M Revio sequencing chip.

Metagenomic assembly, strain-level separation and taxonomic annotation

Long-read metagenomic assembly

Initially, the raw long reads were subjected to a metagenomic assembly for long-read MAGs process using the metaFlye (v2.9)67 with “---nano-raw --meta --keep-haplotypes --iterations 2”. Species-level bins were generated with MaxBin v2.068 and taxonomically annotated using GTDB-Tk v1.6.069. For each species-level bin, Strainberry (v1.1)30 was used separate the contigs into the strain level based on haplotype phasing. Subsequently, these separated contigs were binned to strain level using the “binning” module of metaWRAP (v1.3.2)70. The quality of the reconstructed long-read MAGs was evaluated using the “bin_refinement” module in metaWRAP (v1.3.2)70. The strain-level long-read MAGs with completeness of more than 10% were selected for further analysis. Plasmids were identified usingviralVerify (v1.1)71 and PLASMe (v1.1)72. For viralVerify, only contigs shorter than 500 kbp with prediction scores greater than five were considered. Annotations labeled as ‘Plasmid’ or ‘Uncertain - plasmid or chromosomal’ were classified as plasmids. For PLASMe, only results with high precision were included. Overlapping plasmids identified by both methods were treated as the final results.

Short-read metagenomic assembly

We conducted a metagenomic assembly of the short reads using metaSPAdes (v3.15.2)38 and MEGAHIT (v1.1.3)39 with default settings to construct short-read MAGs. Subsequently, we utilized metaWRAP (v1.3.2)70 to bin and refine the reconstructed short-read MAGs. The completeness and contamination level for the short-read MAGs were estimated using CheckM (v.1.1.3)40 and taxonomically annotated using GTDB-tk (v1.6.0)69.

Evaluation of long-read MAGs

The previous study isolated, sequenced and assembled 184 unique strains from four rCDI cases15, which served as reference genomes. To detect shared strains between long-read MAGs/short-read MAGs and bacterial isolates, we utilized fastANI (1.32)73 with a 99.4% ANI cut-off74. We assessed the contamination level by mapping long-read and short-read MAG to its corresponding isolate genome, calculated as the percentage of unaligned bases. Mummer (v4.0.0b)75 was employed to identify the mapping regions between long-read MAG/short-read MAG and its corresponding isolate genome.. We also assessed the genomic similarity of shared strains using k-mer similarity as in the previous study15, with the following formula:

KmerSimilarity=NumberofsharedkmerbetweenGenomeAandGenomeBLessertotalkmercountofeitherGenomeAandGenomeB

The k-mer similarity between two genomes A and B was determined by generating a hash for genomes with k-mer size 31 and quantifying the proportion of k-mers shared in both genomes A and B divided by the lesser total k-mer count of either

The proportion of reads associated with each strain/MAG

We analyzed long-read and short-read metagenomic sequencing data from D1 by mapping to a merged set of genomes, which included genomes from bacterial isolates and uncultured long-read/short-read MAGs. For the long-read data, we used 136 isolate genomes and 105 uncultured long-read MAGs. The proportion of reads associated with each strain/MAG was determined by calculating the percentage of metagenomic reads that were successfully mapped to each genome/MAG in this merged set. Similarly, the short-read data were aligned to the same 136 isolate genomes and 120 uncultured short-read MAGs, to quantify the proportion of reads associated with each strain/MAG.

LongTrack framework designation for FMT strain tracking based on strain-unique k-mers

The LongTrack framework for FMT strain tracking follows a three-step algorithmic process (Fig. 1C).

Step 1: Identification of strain-unique k-mers from each long-read MAG

a. Extract strain-unique k-mers from each long-read MAG

Following established methods15, LongTrack extracts k-mers (k=31) for each strain-level long-read MAG from the donor or pre-FMT recipient using KMC (version 3.0.0)76. Strain-unique k-mers are then identified by eliminating common k-mers shared in 100,000 bacterial genomes in the NCBI database (as of 2018–11-20), 116 unrelated metagenomic samples from a previous study15, and other long-read MAGs within the same FMT case.

b. Obtain strain-unique k-mers in co-existing strains

Relying solely on strain-unique k-mers identified in Step 1a can lead to false strain-tracking results due to the varying completeness of long-read MAGs, especially for similar co-existing strains. For instance, as shown in Fig. 1C, Strains A1 and A2 both belong to Species A and are co-existing strains from the same donor. If the k-mers identified for Strain A1 in Step 1a correspond to genomic regions missing in Strain A2 due to incomplete assembly, these k-mers may also appear in Strain A2 simply because of the high genomic similarity between the two strains. In such cases, the k-mers are not truly unique to Strain A1 but are instead masked in Strain A2 by assembly gaps, leading to unreliable results.

To address this issue, LongTrack introduces methodological adaptations to ensure the reliable selection of strain-unique k-mers and to minimize false positives. Specifically, for similar co-existing strains within a single fecal sample or closely related strains found separately in the donor and pre-FMT recipient samples, LongTrack applies a filter to retain only those unique k-mers originating from genomic regions shared among multiple co-existing strains in same FMT case. Mummer (v4.0.0b)75 is used to identify these common genomic regions.

This approach prevents the erroneous selection of k-mers that might seem strain-specific but are unreliable due to MAG incompleteness—a critical distinction between high-purity long-read MAGs and complete genomes from bacterial isolates. The resulting retained strain-unique k-mers serve as markers for each long-read MAG, enabling reliable strain tracking in subsequent steps.

Step 2: Assign short reads with final strain-unique k-mers to metagenomic samples

In the second step, the final strain-unique k-mers identified in Step 1b were utilized to assign short reads from post-FMT recipient microbiome data to their corresponding long-read MAGs. Short reads containing strain-unique k-mers were mapped back to their respective long-read MAGs using Bowtie 2.4.177, with parameters optimized for accuracy (‘very-sensitive,’ ‘no-mixed,’ and ‘no-discordant’). The successful alignment of short reads to genomic regions containing strain-unique k-mers confirmed the presence of the corresponding MAG in the post-FMT recipient, facilitating accurate strain tracking.

Step 3: Discard results with low confidence score

To ensure the reliability of strain tracking, we calculated a confidence score by comparing read enrichment in five unrelated metagenomic samples (negative controls) as done in the previous study15. Results with low confidence scores (< 95% confidence) were discarded to maintain the accuracy of strain tracking.

Evaluation of LongTrack for FMT strain tracking

Sensitivity evaluation of LongTrack and culture-based approach (Strainer)

To evaluate the sensitivity of LongTrack and the culture-based approach (Strainer), we used the union of all strains detected either through isolate genomes or MAGs as the reference set. The analysis was stratified by relative abundance thresholds to assess detection performance across varying microbial abundance levels, yielding valuable insights into the sensitivity of LongTrack. Sensitivity was defined as True Positives / (True Positives + False Negatives).

Evaluation of LongTrack for FMT strain tracking through comparison with ground truth

For a systematic assessment of LongTrack for FMT strain tracking, we utilized the shared strains between long-read MAGs/short-read MAGs and genomes of bacterial isolates in four FMT cases of rCDI. This evaluation was conducted by comparing tracking results in the post-FMT recipients at different time points based on Strainer with isolate genomes (ground truth) and LongTrack/Strainer with short-read MAGs. We applied key metrics:

  1. Precision: True Positives / (True Positives + False Positives), the proportion of correct tracking results relative to all tracking results predicted by the LongTrack/Strainer with short-read MAGs.

  2. Specificity: True Negative / (True Negative + False Positives), evaluates the effectiveness of LongTrack/Strainer with short-read MAGs in accurate strain tracking

  3. Accuracy, the number of consistent tracking results between ground truth and LongTrack/Strainer with short-read MAGs divided by the total number of tracking results.

Evaluation of FMT strain tracking between LongTrack and StrainFinder

To evaluate FMT strain tracking between LongTrack and StrainFinder19, we utilized in silico spike-in datasets (Extended Data Fig. 5). For donor D1, we spiked in 3% Clostridium tyrobutyricum (60x) and 2% Turicibacter sanguinis (40x) into long-read microbiome dataset. The simulated Nanopore long reads were generated using nanosim (v3.0.0)78 for metagenomic assembly. Similarly, we produced simulated short-read data with identical abundances of these strains for D1. For the R1 post-FMT, we spiked in short reads to the short-read metagenomic sequencing data, which was used for strain tracking. We designed two distinct spike-in datasets: A) The first dataset A featured a composition of 0.25% (5x) Clostridium tyrobutyricum, along with paired 150bp short reads. Furthermore, this dataset included 0.1% (2x) Turicibacter sanguinis, with paired 150bp short reads. B) The second dataset B presented a composition of 2.5% (50x) Clostridium tyrobutyricum, accompanied by paired 150bp short reads. Additionally, this dataset contained 1% (20x) Turicibacter sanguinis, once again with paired 150bp short reads. We simulated paired short reads using randomreads.sh from the bbmap package (v38.86)79. Subsequently, LongTrack and StrainFinder were performed for strain tracking analysis.

Evaluation of LongTrack for FMT strain tracking when donor and pre-FMT samples contain high-similarity strains

To evaluate LongTrack/Strainer with short-read MAGs for FMT strain tracking when donor and pre-FMT samples contain high-similarity strains, we spiked-in ten pairs of strains (ANI>99.5%) from ten different species, one strain to D1 and the other to R1 pre-FMT, similarly for all ten pairs (Fig. 3F). All twenty strains were added to R1 post-FMT. To facilitate metagenomic assembly, we generated simulated long-read data/short-read data for D1 and R1 pre-FMT into their long-read/short-read microbiome datasets. Furthermore, we simulated short-read paired reads to R1 post-FMT short-read microbiome datasets. The simulated Nanopore long reads and short reads were generated using the nanosim (v3.0.0)78 and the bbmap randomreads.sh (v38.86)79. LongTrack/Strainer with short-read MAGs was performed for strain tracking analysis.

Evaluation of LongTrack for FMT strain tracking when one donor sample contain co-existing strains of the same species

To evaluate LongTrack/Strainer with short-read MAGs for FMT strain tracking when one donor sample contain co-existing strains of the same species, we spiked-in three strains of Phocaeicola massiliensis (Strains 1–3) to D1 and a fourth strain of P. massiliensis (Strain 4) to R1 pre-FMT (Fig. 3H). For metagenomic assembly, long-read data/short-read data were generated for these strains in D1 and R1 pre-FMT microbiome datasets. Furthermore, we simulated short-read paired reads of 4 strains (Strains 1–4) to R1 post-FMT short-read microbiome datasets. The simulated Nanopore long reads and short reads were generated using the nanosim (v3.0.0)78 and bbmap randomreads.sh (v38.86)79. LongTrack/Strainer with short-read MAGs was performed for strain tracking analysis.

Application of LongTrack

Application of LongTrack in strain tracking of rCDI cases

We initially retrieved long-read MAGs from 3 donors and 4 pre-FMT recipients. Using LongTrack, we examined whether these MAGs were present in the metagenomic sequencing data of the corresponding post-FMT recipients across multiple time points.

Application of LongTrack in strain tracking of IBD cases

In the IBD cases, the recipients (R5 and R6) had received fecal microbiota transplantation from six donors (D4-D9), and strain tracking was performed for long-read MAGs from two donors (D4 and D5). To avoid the influence of common k-mers from the remaining four donors, we excluded the k-mers present in their short-read metagenomic sequencing data. Subsequently, we assessed whether these MAGs were present in the metagenomic sequencing data of the corresponding post-FMT recipients across multiple time points using LongTrack.

LongTrack-based reverse tracking and environmental strain identification

We initially retrieved 402 uncultured long-read MAGs from post-FMT recipients (R1-R4). We then identified long-read MAGs shared between the post-FMT recipient and the respective donor or pre-FMT recipient using an ANI exceeding 99.4% in the same FMT cases. We confirmed that 39 long-read MAGs are unique to the post-FMT recipients and investigated whether they were present in the metagenomic sequencing data of the corresponding donor and pre-FMT recipient using LongTrack. If the long-read MAGs were not tracked in microbiome data of donor, pre-FMT or any earlier time points of post-FMT, we considered them to be environmental strains.

Detection of structural variation in engrafted strains using long-read sequencing

We utilized NGMLR (v 0.2.7)80 to perform the alignment of long-read microbiome data of post-FMT recipients (R1-R4) across various time points to the uncultured long-read MAGs and isolated genomes from the respective rCDI donors (D1-D3). Following the alignment process, we employed Sniffles (v1.0.12)81 for the detection of SVs. To ensure the reliability of SV detection, we applied several filtering criteria: 1) We excluded SVs involving multiple strains within the same species, as these are more prone to mismatches and may not accurately represent SVs. 2) For mapping quality control, we retained aligned reads with an identity greater than 98%, a length more than 1000 bp, and a MAPQ (Mapping Quality) score 60. 3) We kept SVs that had a coverage depth greater than 5 in the donor and at least one post-FMT, ensuring that the detected structural variations were well-supported by sequencing data.

Validation of structural variations using isolate long-read sequencing

We collected 18 isolates for SV validation (detailed list in Supp Table S16). Native libraries were prepared using the ONT Native Barcoding Kit (SQK-NBD114.24) for R10.4.1 flow cells and the Ligation Sequencing Kit (SQK-LSK114). Samples were sequenced on R10.4.1 MinION flow cells, and basecalling was performed post-run using Guppy (v6.3.4) with the SUP model dna_r10.4.1_e8.2_400bps_sup.cfg. Single-isolate long reads were aligned to corresponding long-read MAGs with NGMLR (v0.2.7)80; SVs were called with Sniffles (v1.0.12)81.

Methylation analysis using long read sequencing

De novo methylation motif discovery using Nanopore long-read sequencing

We utilized Nanodisco (v1.0.2)33 to identify de novo methylation motifs in long-read MAGs from the D1. First, raw fast5 reads from both native and WGA metagenomic sequencing data of D1 were aligned to the long-read MAGs with “nanodisco preprocess”. The alignments for each long-read MAG were then segregated. Only long-read MAGs exhibiting coverage above 20x were retained. We then computed current differences (PA) between the native (retaining methylation) and WGA (methylation-free) data for each genomic position using “nanodisco difference”. To conclude, “nanodisco motif and refine” was employed for de novo methylation motif discovery based on these current differences. For D1 5-year, R1 post-FMT 8-week, and R1 post-FMT 5-year, the raw fast5 reads from native microbiome data were aligned to the retained long-read MAGs from D1 with “nanodisco preprocess” and separated the alignments into each MAG. Subsequently, we computed the current differences by comparing the native data with the WGA data from D1 for each MAG using “nanodisco difference”. “Nanodisco refine” was employed to verify the presence of D1 detected methylation motifs in the samples of D1 5-year, R1 post-FMT 8-week, and R1 post-FMT 5-year based on the current differences.

De novo methylation motif discovery using PacBio long-read sequencing

For PacBio long-read data of D1 (generated by the Sequel II platform), the raw reads pre-processing was performed with SMRTlink (v8.0). Raw multiplexed files were first de-multiplexed with lima. We mapped the subreads to long-read MAGs using pbalign (v0.4.1) with default parameters and separated the alignments into each MAG. The IPD ratio was calculated with ipdSummary (v2.4.1) from the alignments for each MAG. The de novo methylation motifs were detected by using motifMaker (SMRTlink v8.0) with default parameters. For PacBio long-read data of D1 5-year, R1 8-week and R1 5-year (generated by the Revio platform), the raw reads pre-processing was performed with SMRTlink (v13.0), and methylation motif detection was performed by our in-house scripts.

Statistics

No statistical method was used to predetermine sample size; all available samples were analyzed. No data was excluded. The experiments were not randomized. Blinding was not relevant in our study. Data distribution was assumed to be normal but it was not formally tested. For unpaired long-read versus short-read MAGs comparisons of CheckM-based contamination, completeness, and N50, we used a two-sided Wilcoxon rank-sum test. Contamination percentages relative to isolate genomes were compared using two-sided permutation tests on the mean difference (10,000 permutations). For single comparisons, no multiple-comparison adjustment was applied. Analyses were performed in R 4.3.1.

Extended Data

Extended Data Fig. 1. Description of sample collection and FMT study design for CDI patients.

Extended Data Fig. 1

Stool samples were collected from three donors and four recipients at multiple time points. For the illustration of Long Track, strain tracking was performed for long-read MAGs assembled from donors and recipients using short-read metagenomic data generated across all the samples shown in the figure (Methods). The Venn diagram displays the number of strains that were recovered by long-read (LR) MAGs (yellow color) or isolate genomes from bacterial culture (red color) from 3 donors and 2 recipients

Extended Data Fig. 2. Illustration of Long Track workflow’s implementation in this study for generating strain-level long-read MAGS.

Extended Data Fig. 2

Initially, the raw long reads from donor and pre-FMT recipient underwent metagenomic assembly using the metaFlye (v2.9). The resulting Metagenome-assembled genomes (MAGS) were then partitioned into individual bins corresponding to the species level using MaxBin (v2.0). Strainberry (v1.1) was used to separate the contigs into strain level based on haplotype phasing for each species-level bin. Subsequently, metaWRAP (v1.3.2) was employed to bin the separated contigs into strain level. Finally, the strain-level long-read MAGs with completeness of more than 10% were chosen for FMT strain tracking in post-FMT recipients (Methods).

Extended Data Fig. 3. Comparative evaluation of long-read MAGs and short-read MAGS from Donor 1.

Extended Data Fig. 3

A. Completeness levels of 124 long-read MAGS (LR MAGs) and 135 short-read MAGS (SR MAGs) from Donor 1 reveal no significant differences. B. Contamination levels of 124 LR MAGS and 135 SR MAGs from Donor 1 assessed using CheckM, showing significantly lower contamination in LR MAGs (p=0.043 ; two-sided Wilcoxon rank sum test; p<0.05, *). C. Contig N50 (kb) comparison for 124 LR MAGs and 135 SR MAGs Donor 1, showing significantly higher N50 in LR MAGs (p = 2.2e-16; two-sided Wilcoxon rank sum test; p < 0.01 **).

Extended Data Fig. 4. Assessment of contamination level in long-read/short-read MAGS.

Extended Data Fig. 4

A. Schematic of contamination level by mapping long-read and short-read MAG to its corresponding isolate genome, calculated as the percentage of unaligned bases contamination; B. Comparative analysis of contamination levels determined by comparing to the isolate genome and by using the CheckM tool. A total of 60 long-read (LR) MAGs from Donors 1–3 and Recipients 1–4, and 23 short-read (SR) MAGs from Donor 1 (D1) were analyzed. (C-D) Comparison of contamination levels in corresponding strains between LR MAGs and SR MAGS from D1. (C) Bar plots illustrate contamination levels in corresponding strains between LR MAGs and SR MAGs for species with multiple co-existing strains and species with a single strain, emphasizing the consistently lower contamination observed in LR MAGS. D. Contamination levels for 31 LR MAGs and 23 SR MAGs from D1, categorized by species with either a single strain or multiple co-existing strains. LR MAGs consistently demonstrated lower contamination levels (p-value = 0.0048 for single strain; p-value = 0.0086 for multiple strain; two-sided permutation test on the mean difference, n = 10,000 permutations; p-value < 0.01, **) E. Genomic similarity of shared strains between the long-read (LR) MAGs/short-read (SR) MAGs and their corresponding isolate genomes determined by k-mer similarity (Methods).

Extended Data Fig. 5. Evaluation of FMT strain tracking between Long Track and StrainFinder as a representative of SNP-based inference approach.

Extended Data Fig. 5

A. For the donor sample D1, in silico spike-in dataset was designed with 3% Clostridium tyrobutyricum and 2% Turicibacter sanguinis added to the microbiome dataset. For the post-FMT recipient R1, the spike-in data set A (low abundance) had a composition of 0.25% Clostridium tyrobutyricum and 0.1% Turicibacter sanguinis, while the spike-in data set B (high abundance) had a composition of 2.5% Clostridium tyrobutyricum and 1% Turicibacter sanguinis (Methods). B. The strain tracking results indicate that the SNP-based approach, StrainFinder, works well for strains with relatively modest-to-high abundance (data set B), but tends to miss strains with modest-to-low abundance (data set A). The chart represents the presence (green) or absence (gray) of strains in post-FMT recipients based on spike-in design (considering ground truth), Long Track with long-read MAGs, or StrainFinder. Discrepancies with the ground truth are indicated by an “X” marker.

Extended Data Fig. 6. Sensitivity evaluation of Long Track and culture-based approach (Strainer).

Extended Data Fig. 6

The sensitivity evaluation of Long Track and the culture-based approach (Strainer) was performed using the union of all strains detected either through isolate genomes or MAGs as the reference set. To evaluate sensitivity across different microbial abundance levels, the reference set was stratified by relative abundance thresholds, which provides a comprehensive benchmark that encompasses the diversity of detectable strains, including those uniquely identified by each method. Numbers in brackets on the x-axis indicate the number of reference strains (n) in each bin.

Extended Data Fig. 7. LongTrack uncovered engraftment of cultured strains across FMT cases in rCDI patients.

Extended Data Fig. 7

A, Proportional engraftment of donor strains (PEDS) for shared cultured strains is shown across individual post-FMT recipients at various time points of the 5-year follow up. B, Proportional persistence of recipient strains (PPRS) for shared cultured strains is shown across individual post-FMT recipients at various time points of the 5-year follow up.

Extended Data Fig. 8. LongTrack uncovered uncultured of engrafted strains of rCDI.

Extended Data Fig. 8

The long-read MAGs uncultured by the previous study were utilized for strain tracking to determine if they were present in the post-FMT recipients for FMT cases of CDI: The results of tracking across three FMT cases D1->R1 (A), D2->R3 (B), and D3->R4 (C), which describe the presence (green) or absence (gray) of strains based on LongTrack.

Extended Data Fig 9. Reverse strain tracking of earlier samples based on long-read MAGs assembled from post-FMT recipients using LongTrack.

Extended Data Fig 9

The long-read MAGs assembled from post-FMT recipient samples were utilized for strain tracking to determine if they were also present in the donor and pre-FMT recipient metagenomic data using LongTrack (Reverse tracking). The results of reverse tracking across three FMT cases: D1 -> R1, D2 -> R3 and D3 -> R4 are shown in A-C, which describe the presence (green) or absence (gray) of strains based on LongTrack. (D-F) Illustration of bacterial strains that may have been acquired from environmental sources, as determined by long-read MAGs assembled from the five-year post-FMT recipients’ samples but absent in all the earlier samples. The results across three FMT cases: D1 -> R1, D2 -> R3 and D3 -> R4 are shown in D-F, which describe the presence (green) or absence (gray) of strains based on LongTrack.

Extended Data Fig 10. Description of samples collection and FMT study design for two IBD patients.

Extended Data Fig 10

Stool samples were collected from six donors and two recipients at multiple time points. The two recipients (IBD R5 and IBD R6; orange) had received FMT from six donors (IBD D4 - D9; purple and gray). For the illustration of LongTrack, strain tracking was performed for long-read MAGs assembled from two donors (IBD D4 and D5; purple) using short-read metagenomic data generated across all the samples shown in the figure (Methods)

Supplementary Material

Supplementary Figures
Supplementary Tables

Acknowledgments

This work was supported in part by the staff and resources of the Microbiome Translational Center and Department of Scientific Computing at the Icahn School of Medicine at Mount Sinai. This work was supported by grant no. R35 GM139655 (G.F.) from the National Institutes of Health. G.F. is a Hirschl Research Scholar by Irma T. Hirschl/Monique Weill-Caulier Trust and a Nash Family Research Scholar.

Footnotes

Competing interests

S.P. has served as a consultant / steering committee member for Vedanta Biosciences and has received speaker / advisory board fees from AbbVie, Dr Falk Pharma, Ferring, Janssen and Takeda. J.F. is a consultant for Vedanta Biosciences and Genfit. The remaining authors declare no competing interests.

Data availability

All sequencing data generated for this study have been uploaded to the Sequence Read Archive under the BioProject PRJNA1101882.

Code Availability

Customized scripts of LongTrack and a tutorial with supporting data are available at https://github.com/fanglab/LongTrack

Reference

  • 1.Feuerstadt P et al. SER-109, an Oral Microbiome Therapy for Recurrent Clostridioides difficile Infection. N Engl J Med 386, 220–229 (2022). [DOI] [PubMed] [Google Scholar]
  • 2.Khanna S et al. Efficacy and Safety of RBX2660 in PUNCH CD3, a Phase III, Randomized, Double-Blind, Placebo-Controlled Trial with a Bayesian Primary Analysis for the Prevention of Recurrent Clostridioides difficile Infection. Drugs 82, 1527–1538 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Ooijevaar RE, Terveer EM, Verspaget HW, Kuijper EJ & Keller JJ Clinical Application and Potential of Fecal Microbiota Transplantation. Annu Rev Med 70, 335–351 (2019). [DOI] [PubMed] [Google Scholar]
  • 4.Li SS et al. Durable coexistence of donor and recipient strains after fecal microbiota transplantation. Science 352, 586–589 (2016). [DOI] [PubMed] [Google Scholar]
  • 5.Pamer EG Fecal microbiota transplantation: effectiveness, complexities, and lingering concerns. Mucosal Immunol 7, 210–214 (2014). [DOI] [PubMed] [Google Scholar]
  • 6.Bethlehem L et al. Microbiota therapeutics for inflammatory bowel disease: the way forward. The Lancet Gastroenterology & Hepatology 9, 476–486 (2024). [DOI] [PubMed] [Google Scholar]
  • 7.Paramsothy S et al. Multidonor intensive faecal microbiota transplantation for active ulcerative colitis: a randomised placebo-controlled trial. The Lancet 389, 1218–1228 (2017). [DOI] [PubMed] [Google Scholar]
  • 8.Rossen NG et al. Findings From a Randomized Controlled Trial of Fecal Transplantation for Patients With Ulcerative Colitis. Gastroenterology 149, 110–118.e4 (2015). [DOI] [PubMed] [Google Scholar]
  • 9.Sevcikova A et al. The Impact of the Microbiome on Resistance to Cancer Treatment with Chemotherapeutic Agents and Immunotherapy. International Journal of Molecular Sciences 23, 488 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Davar D et al. Fecal microbiota transplant overcomes resistance to anti-PD-1 therapy in melanoma patients. Science 371, 595–602 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhao S et al. Adaptive evolution within the gut microbiomes of healthy people. Cell Host Microbe 25, 656–667.e8 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Lloyd-Price J et al. Strains, functions and dynamics in the expanded Human Microbiome Project. Nature 550, 61–66 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Porcari S et al. Key determinants of success in fecal microbiota transplantation: From microbiome to clinic. Cell Host & Microbe 31, 712–733 (2023). [DOI] [PubMed] [Google Scholar]
  • 14.Ianiro G et al. Variability of strain engraftment and predictability of microbiome composition after fecal microbiota transplantation across different diseases. Nat Med 28, 1913–1923 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Aggarwala V et al. Precise quantification of bacterial strains after fecal microbiota transplantation delineates long-term engraftment and explains outcomes. Nat Microbiol 6, 1309–1318 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Duvallet C et al. Framework for rational donor selection in fecal microbiota transplant clinical trials. PLOS ONE 14, e0222881 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Podlesny D et al. Identification of clinical and ecological determinants of strain engraftment after fecal microbiota transplantation using metagenomics. CR Med 3, (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Luo C et al. ConStrains identifies microbial strains in metagenomic datasets. Nat Biotechnol 33, 1045–1052 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Smillie CS et al. Strain-tracking reveals the determinants of bacterial engraftment in the human gut following fecal microbiota transplantation. Cell Host Microbe 23, 229–240.e5 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Schmidt TSB et al. Drivers and determinants of strain dynamics following fecal microbiota transplantation. Nat Med 28, 1902–1912 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Olm MR et al. InStrain enables population genomic analysis from metagenomic data and sensitive detection of shared microbial strains. Nat Biotechnol 39, 727–736 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Blanco-Míguez A et al. Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. Nat Biotechnol 41, 1633–1644 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Van Rossum T, Ferretti P, Maistrenko OM & Bork P Diversity within species: interpreting strains in microbiomes. Nat Rev Microbiol 18, 491–506 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Gounot J-S et al. Genome-centric analysis of short and long read metagenomes reveals uncharacterized microbiome diversity in Southeast Asians. Nat Commun 13, 6044 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Sczyrba A et al. Critical Assessment of Metagenome Interpretation – a benchmark of computational metagenomics software. Nat Methods 14, 1063–1071 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Tidjani Alou M et al. State of the Art in the Culture of the Human Microbiota: New Interests and Strategies. Clinical Microbiology Reviews 34, 10.1128/cmr.00129-19 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Moss EL, Maghini DG & Bhatt AS Complete, closed bacterial genomes from microbiomes using nanopore sequencing. Nat Biotechnol 38, 701–707 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Gehrig JL et al. Finding the right fit: evaluation of short-read and long-read sequencing approaches to maximize the utility of clinical microbiome data. Microbial Genomics 8, 000794 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Kim CY, Ma J & Lee I HiFi metagenomic sequencing enables assembly of accurate and complete genomes from human gut microbiota. Nat Commun 13, 6367 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Vicedomini R, Quince C, Darling AE & Chikhi R Strainberry: automated strain separation in low-complexity metagenomes using long reads. Nat Commun 12, 4485 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Agustinho DP et al. Unveiling microbial diversity: harnessing long-read sequencing technology. Nat Methods 21, 954–966 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Chen L et al. Short- and long-read metagenomics expand individualized structural variations in gut microbiomes. Nat Commun 13, 3175 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Tourancheau A, Mead EA, Zhang X-S & Fang G Discovering multiple types of DNA methylation from bacteria and microbiome using nanopore sequencing. Nat Methods 18, 491–498 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Beaulaurier J et al. Metagenomic binning and association of plasmids with bacterial host genomes using DNA methylation. Nat Biotechnol 36, 61–69 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Beaulaurier J, Schadt EE & Fang G Deciphering bacterial epigenomes using modern sequencing technologies. Nat Rev Genet 20, 157–172 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Kong Y et al. Critical assessment of DNA adenine methylation in eukaryotes using quantitative deconvolution. Science 375, 515–522 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Fang G et al. Genome-wide mapping of methylated adenine residues in pathogenic Escherichia coli using single-molecule real-time sequencing. Nat Biotechnol 30, 1232–1239 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Nurk S, Meleshko D, Korobeynikov A & Pevzner PA metaSPAdes: a new versatile metagenomic assembler. Genome Res 27, 824–834 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Li D, Liu C-M, Luo R, Sadakane K & Lam T-W MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics 31, 1674–1676 (2015). [DOI] [PubMed] [Google Scholar]
  • 40.Parks DH, Imelfort M, Skennerton CT, Hugenholtz P & Tyson GW CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res 25, 1043–1055 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Kent AG, Vill AC, Shi Q, Satlin MJ & Brito IL Widespread transfer of mobile antibiotic resistance genes within individual gut microbiomes revealed through bacterial Hi-C. Nat Commun 11, 4379 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Bae M et al. Akkermansia muciniphila phospholipid induces homeostatic immune responses. Nature 608, 168–173 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Wu Z et al. Akkermansia muciniphila Ameliorates Clostridioides difficile Infection in Mice by Modulating the Intestinal Microbiome and Metabolites. Front Microbiol 13, 841920 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Leylabadlo HE et al. The critical role of Faecalibacterium prausnitzii in human health: An overview. Microbial Pathogenesis 149, 104344 (2020). [DOI] [PubMed] [Google Scholar]
  • 45.Quévrain E et al. Identification of an anti-inflammatory protein from Faecalibacterium prausnitzii, a commensal bacterium deficient in Crohn’s disease. Gut 65, 415–425 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Chen-Liaw A et al. Gut microbiota strain richness is species specific and affects engraftment. Nature 637, 422–429 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Durrant MG & Bhatt AS Microbiome genome structure drives function. Nat Microbiol 4, 912–913 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Kong Y, Mead EA & Fang G Navigating the pitfalls of mapping DNA and RNA modifications. Nat Rev Genet 24, 363–381 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Noinaj N, Guillier M, Travis J Barnard & Buchanan, S. K. TonB-Dependent Transporters: Regulation, Structure, and Function. Annual Review of Microbiology 64, 43–60 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Dingemans J et al. The deletion of TonB-dependent receptor genes is part of the genome reduction process that occurs during adaptation of Pseudomonas aeruginosa to the cystic fibrosis lung. Pathog Dis 71, 26–38 (2014). [DOI] [PubMed] [Google Scholar]
  • 51.Partridge SR, Kwong SM, Firth N & Jensen SO Mobile Genetic Elements Associated with Antimicrobial Resistance. Clinical Microbiology Reviews 31, 10.1128/cmr.00088-17 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Maqbool A et al. The substrate-binding protein in bacterial ABC transporters: dissecting roles in the evolution of substrate specificity. Biochemical Society Transactions 43, 1011–1017 (2015). [DOI] [PubMed] [Google Scholar]
  • 53.Sánchez-Romero MA & Casadesús J The bacterial epigenome. Nat Rev Microbiol 18, 7–20 (2020). [DOI] [PubMed] [Google Scholar]
  • 54.Beaulaurier J et al. Single molecule-level detection and long read-based phasing of epigenetic variations in bacterial methylomes. Nat Commun 6, 7438 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Oliveira PH & Fang G Conserved DNA Methyltransferases: A Window into Fundamental Mechanisms of Epigenetic Regulation in Bacteria. Trends in Microbiology 29, 28–40 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Seib KL, Srikhanta YN, Atack JM & Jennings MP Epigenetic Regulation of Virulence and Immunoevasion by Phase-Variable Restriction-Modification Systems in Bacterial Pathogens. Annu Rev Microbiol 74, 655–671 (2020). [DOI] [PubMed] [Google Scholar]
  • 57.Oliveira PH et al. Epigenomic characterization of Clostridioides difficile finds a conserved DNA methyltransferase that mediates sporulation and pathogenesis. Nat Microbiol 5, 166–180 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Lopez-Siles M, Duncan SH, Garcia-Gil LJ & Martinez-Medina M Faecalibacterium prausnitzii: from microbiology to diagnostics and prognostics. ISME J 11, 841–852 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Wu H et al. Strategies for high cell density cultivation of Akkermansia muciniphila and its potential metabolism. Microbiol Spectr 12, e0238623 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Zhernakova DV et al. Host genetic regulation of human gut microbial structural variation. Nature 625, 813–821 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.West PT, Chanin RB & Bhatt AS From genome structure to function: insights into structural variation in microbiology. Current Opinion in Microbiology 69, 102192 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Furuta Y et al. Methylome Diversification through Changes in DNA Methyltransferase Sequence Specificity. PLOS Genetics 10, e1004272 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Srikhanta YN, Fox KL & Jennings MP The phasevarion: phase variation of type III DNA methyltransferases controls coordinated switching in multiple genes. Nat Rev Microbiol 8, 196–206 (2010). [DOI] [PubMed] [Google Scholar]
  • 64.Ji P, Zhang Y, Wang J & Zhao F MetaSort untangles metagenome assembly by reducing microbial community complexity. Nat Commun 8, 14306 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Cao L et al. mEnrich-seq: methylation-guided enrichment sequencing of bacterial taxa of interest from microbiome. Nat Methods 21, 236–246 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Sarkar A et al. Microbial transmission in the social microbiome and host health and disease. Cell 187, 17–43 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Kolmogorov M et al. metaFlye: scalable long-read metagenome assembly using repeat graphs. Nat Methods 17, 1103–1110 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Wu Y-W, Simmons BA & Singer SW MaxBin 2.0: an automated binning algorithm to recover genomes from multiple metagenomic datasets. Bioinformatics 32, 605–607 (2016). [DOI] [PubMed] [Google Scholar]
  • 69.Chaumeil P-A, Mussig AJ, Hugenholtz P & Parks DH GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics 36, 1925–1927 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Uritskiy GV, DiRuggiero J & Taylor J MetaWRAP—a flexible pipeline for genome-resolved metagenomic data analysis. Microbiome 6, 158 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Antipov D, Raiko M, Lapidus A & Pevzner PA Metaviral SPAdes: assembly of viruses from metagenomic data. Bioinformatics 36, 4126–4129 (2020). [DOI] [PubMed] [Google Scholar]
  • 72.Tang X, Shang J, Ji Y & Sun Y PLASMe: a tool to identify PLASMid contigs from short-read assemblies using transformer. Nucleic Acids Res 51, e83 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Jain C, Rodriguez-R LM, Phillippy AM, Konstantinidis KT & Aluru S High throughput ANI analysis of 90K prokaryotic genomes reveals clear species boundaries. Nat Commun 9, 5114 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Rodriguez-R LM et al. An ANI gap within bacterial species that advances the definitions of intra-species units. mBio 15, e02696–23 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Marçais G et al. MUMmer4: A fast and versatile genome alignment system. PLOS Computational Biology 14, e1005944 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Kokot M, Dlugosz M & Deorowicz S KMC 3: counting and manipulating k-mer statistics. Bioinformatics 33, 2759–2761 (2017). [DOI] [PubMed] [Google Scholar]
  • 77.Langmead B & Salzberg SL Fast gapped-read alignment with Bowtie 2. Nat Methods 9, 357–359 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Yang C, Chu J, Warren RL & Birol I NanoSim: nanopore sequence read simulator based on statistical characterization. GigaScience 6, gix010 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Bushnell B BBMap: A Fast, Accurate, Splice-Aware Aligner. (2014). [Google Scholar]
  • 80.Sedlazeck FJ et al. Accurate detection of complex structural variations using single molecule sequencing. Nat Methods 15, 461–468 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Smolka M et al. Detection of mosaic and population-level structural variants with Sniffles2. Nat Biotechnol 1–10 (2024) doi: 10.1038/s41587-023-02024-y. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Figures
Supplementary Tables

Data Availability Statement

All sequencing data generated for this study have been uploaded to the Sequence Read Archive under the BioProject PRJNA1101882.

RESOURCES