Abstract
Summary
Bacterial phase variation enables reversible, locus-specific phenotypic switching, often driven by DNA inversion (invertons). To identify these events, researchers commonly rely on sequencing reads that provide orientation-specific support. Metagenomic sequencing, which captures total genetic material independent of cultivation, offers a powerful platform for the comprehensive study of invertons. However, computational inverton calling from metagenomic data is difficult at low sequencing depth: hard read-support cutoffs can miss true events, while sequence-only predictors lack read-backed interpretability and uncertainty quantification. To address this, we present TPMM, a three-component posterior mixture model for inverton calling in metagenomic data. TPMM explicitly incorporates sequencing depth to formulate inverton detection as a probabilistic mixture problem. Starting from candidates flanked by inverted repeats, the model classifies the candidates into noise, low-probability, or high-probability inversion signals using read evidence. Finally, TPMM assigns posterior probabilities as soft labels and applies cumulative Bayesian False Discovery Rate control to robustly identify true invertons. On two real gut metagenomic datasets, TPMM agrees well with PhaseFinder at high depth but recovers substantially more invertons under systematic downsampling, demonstrating superior performance in sparse-data regimes. We further examine potential reversible inversion elements in viral genomes and provide supporting analyses, suggesting a broader scope for inversion-mediated regulation.
Availability
The source code of TPMM is available via: https://github.com/KennyxxD/TPMM
1 Introduction
Bacterial phase variation is a widespread strategy that enables frequent, reversible switching of genetic states at specific loci (Henderson et al. 1999, van der Woude and Bäumler 2004, Trzilova and Tamayo 2021), thereby generating phenotypic heterogeneity even within clonal populations. In contrast to adaptation driven by point mutations, phase variation typically restricts variability to a small set of hypermutable loci. This localization allows populations to rapidly sample alternative phenotypes while limiting the fitness costs associated with an elevated mutation rate across the rest of the genome. Phase variation is therefore thought to play a central role in pathogen infection and immune evasion, and in the long-term persistence of commensals under fluctuating selective pressures, including host immunity, antibiotic exposure and spatial heterogeneity (Coyne et al. 2008, Porter et al. 2020). Among several molecular mechanisms of phase variation, reversible DNA inversion is particularly prominent (Ben-Assa et al. 2020, Trzilova and Tamayo 2021). This process is typically mediated by site-specific recombinases (invertases) recognizing flanking inverted repeats (IRs) to catalyze segment inversion (Roche-Hakansson et al. 2007). Compared with nucleotide substitutions or indels, inversion tends to occur at higher rates and is reversible, enabling stable coexistence of subpopulations with distinct phenotypic states. Functionally, a common regulatory architecture is the invertible promoter: in the ON orientation, the promoter drives transcription of downstream genes or operons, whereas in the OFF orientation transcription is inhibited (Yan et al. 2022). For example, several capsular polysaccharide loci in Bacteroides fragilis are regulated by invertible promoters that flip between ON and OFF orientations, thereby turning expression of the corresponding biosynthesis operons on or off (Troy et al. 2010). Invertible segments can also include alternative regulatory elements such as terminators, altering transcriptional output via additional mechanisms. Beyond intergenic invertons, gene-intersecting inversions may generate alternative protein isoforms or modify molecular specificity (Grundy and Howe 1984, Sandulache et al. 1984). Collectively, inverton-mediated regulation can act both as an expression switch and as a mechanism shaping microbial structural and functional interactions with their environment and hosts.
In metagenomic data, active inversion is evidenced by sequencing reads supporting both orientations of a locus (Troy et al. 2010). Accordingly, the numbers of reads supporting the forward versus reverse orientation, and their ratio, are commonly used features for inverton detection and quantification. PhaseFinder, the first tool to exploit this, identifies invertons by quantifying read support for in silico generated orientations (Jiang et al. 2019). However, PhaseFinder relies on fixed thresholds, which often excludes true invertons in low-coverage regions. To address this limitation, a second class of methods primarily relies on sequence signals. DeepInverton (Wen et al. 2024) is designed based on the inverton sequence rather than read evidence. Moreover, by operating independently of reads, DeepInverton lacks direct read-backed interpretability and principled treatment of sampling uncertainty.
Here, we develop TPMM, a three-component posterior mixture model for probabilistic inverton calling. TPMM operates on the full set of candidate loci generated by PhaseFinder prior to its default filtering, specifically retaining loci with at least one supporting reverse read (via paired-end or spanning methods) and a total read count exceeding a fixed threshold. We propose an explicit generative framework for read support evidence, decomposing the candidate loci into a three-component posterior mixture representing noise, low-probability, and high-probability signals. The noise component captures spurious support for reverse orientation reads, while the low-probability and high-probability components represent true inversions occurring at low and high probabilities in the environment. The model outputs posterior probabilities as soft labels, allowing for the quantification of evidence strength and uncertainty without relying on an arbitrary binary cutoff. To identify the final set of true invertons for downstream analysis, we further control cumulative Bayesian False Discovery Rate on the posterior outputs, converting the soft labels into final hard calls (Newton et al. 2004).
We validated TPMM on two real metagenomic datasets, including a human gut dataset and a rat gut dataset. At high read depth, TPMM shows strong agreement with PhaseFinder, indicating that it preserves conclusions when evidence is sufficient. Based on this, we performed downsampling on both datasets at three different ratios to emulate progressively lower read depth and prove that TPMM recovers substantially more invertons than PhaseFinder, demonstrating improved recovery under low-depth conditions. Moreover, the invertons identified by TPMM are significantly enriched in downstream genes, underscoring their biological significance. Finally, building on previous reports of reversible DNA inversions in bacteriophages (Giphart-Gassler et al. 1982, Van de Putte and Goosen 1992), we systematically searched for putative invertons in metagenomically recovered viral genomes.
2 Methodology
2.1 Overview and data representation
As illustrated in Fig. 1A, PhaseFinder initiates by using EMBOSS einverted to locate inverted repeats. It then generates an augmented reference containing both forward and reverse sequences in silico for each candidate for metagenomic read alignment. Alignment evidence is derived from Paired-end (Pe) and Spanning (Span) methods (Supplementary Section 1, available as supplementary data at Bioinformatics online), which are classified into Forward (F) or Reverse (R) orientations. A significant accumulation of R reads relative to total depth indicates the presence of an active DNA inversion. However, the default PhaseFinder algorithm relies on hard-filtering thresholds to identify positive calls, specifically requiring at least 5 paired-end reverse reads (), 3 spanning reverse reads (), and a reverse orientation frequency of at least 1%. Although effective for high-coverage isolated genomes, this approach lacks the flexibility required for metagenomic data. Given the variable sequencing depth inherent in complex microbial communities, the reliance on hard thresholds inevitably leads to the omission of invertons under low-depth metagenomic data.
Figure 1.

Overview of the pipeline. (A) Acquisition of read support evidence. PhaseFinder is utilized to quantify read counts supporting both the forward and reverse orientations. (B) Statistical inference and false discovery control. After initial thresholding of read evidence, the data is analyzed using a Three-component Posterior Mixture Model (TPMM) to estimate posterior probabilities for each locus via the Expectation-Maximization (EM) algorithm. Candidate invertons are ranked in ascending order of their posterior probability of being noise () to calculate the cumulative Bayesian False Discovery Rate (BFDR). The final set of invertons is identified by selecting the top K candidates such that , where is a user-defined threshold.
Hence, we propose TPMM, the Three-component Posterior Mixture Model. Instead of applying a rigid cutoff, TPMM evaluates read support to generate a probabilistic soft label. This strategy allows the model to adapt dynamically to varying read counts, ensuring the recovery of valid biological signals even when read support is scarce. The workflow of TPMM is shown in Figure 1B. It defines the input data for a candidate locus i as the raw results of read support:
| (1) |
Where F and R denote the number of reads supporting the forward and reverse orientations, respectively. We define the total sequencing depth for candidate i within a specific channel as . We initially process each metagenomic sample independently using its corresponding sequencing reads. For each sample, we apply a relaxed pre-filtering step to retain potential low-probability signals, keeping candidates with a total number of reads () and at least one supporting reverse read in either channel (). The filtered candidates from all samples in the same environment are then collected together and processed using the framework below.
2.2 Three-component posterior mixture model (TPMM)
We assume that the observed read counts arise from a mixture of three distinct latent biological states (Supplementary Section 2, available as supplementary data at Bioinformatics online), denoted by :
Noise (): Background sequencing errors where no inversion exists.
Low-probability (): Rare inversion events present in a small fraction of the population.
High-probability (): Common inversion events in this environment.
The marginal likelihood for the observed data is a weighted sum of the conditional likelihoods for each component:
| (2) |
Where represents the mixing proportion of state k, subject to . The set of model parameters is defined as:
| (3) |
Which includes the mixing proportions, the signal inversion rates (), and the noise regression coefficients (). These parameters are described in detail below.
2.2.1 Component-wise likelihood
We model the number of reverse reads in channel c as a Binomial random variable, conditioned on the total depth and an inversion probability . We assume that, conditional on the latent state k and the observed total depths , the reverse reads from the two channels are approximately independent:
| (4) |
Here, represents the expected proportion of reverse reads for candidate i in state k. We define these probabilities differently for signal and noise components. For true invertons, we estimate a constant inversion rate for each channel. Thus, and . Here, we address the label-switching aspect of TPMM identifiability. In the unconstrained model, the labels of the two signal components are interchangeable. To assign consistent biological interpretations to these components, we impose an ordering constraint on their average reverse-read probabilities:
| (5) |
Other identifiability considerations are provided in Supplementary Section 3, available as supplementary data at Bioinformatics online. In addition, sequencing error rates often fluctuate with sequencing depth. To capture this, we model the noise parameter using a logistic regression function dependent on the total read depth . To reduce nuisance parameters, we share this noise rate across both channels:
| (6) |
Where is the sigmoid function, and are regression coefficients estimated from the data.
2.2.2 Estimation and inference
We estimate the parameter set via the Expectation-Maximization (EM) algorithm by maximizing the total log-likelihood . The detailed procedure is provided in Supplementary Section 4, available as supplementary data at Bioinformatics online. After convergence, we compute the posterior probability that candidate i belongs to state k, denoted as . We define the probability that a candidate is a true inverton () as the sum of the posterior probabilities for the low- and high-probability signal components:
| (7) |
To determine the final set of positive calls, we employ cumulative Bayesian False Discovery Rate (BFDR) approach. We define the local false discovery rate (lfdr) for candidate i as its noise probability (). Candidates are ranked in increasing order of . For a cutoff at rank K, the cumulative BFDR is calculated as:
| (8) |
We select the largest K such that (where is the desired error rate), providing a principled method to control false positives without relying on arbitrary read count thresholds. Parameter estimation stability increases with the number of candidate loci. Based on the datasets used in this study, reliable convergence was achieved with approximately 1,500 or more loci (Figure S2).
3 Results
We analyzed two metagenomic datasets: a rat gut study (Wen et al. 2025) and a human gut study (Lewis et al. 2015). We then compared our results with PhaseFinder and DeepInverton. The rat dataset, collected from diverse habitats in Hong Kong, represents complex environmental stressors that may drive inverton turnover. The human dataset comprises 26 samples from healthy controls and 343 from individuals with Crohn’s disease, spanning multiple therapeutic regimens and longitudinal timepoints. This design captures dynamic perturbations in the gut environment, thereby increasing the likelihood of detecting inverton events.
Raw sequencing reads underwent quality control using Trimmomatic (v0.39) to remove adapters and low-quality bases. The parameters are provided in Table S3. To minimize host contamination, filtered reads were aligned to the corresponding host reference genomes using Bowtie2 (v2.3.5.1), and aligned reads were discarded. The remaining high-quality, non-host reads were assembled de novo into contigs using MEGAHIT (v1.2.9) with default settings. Following preprocessing, 88 rat and 365 human gut metagenomes were retained for downstream analysis with approximately and reads, respectively.
First, we characterized the identified invertons by examining their supporting read distributions, demonstrating how the probabilistic model stratifies candidates based on evidence strength rather than arbitrary cutoffs. Second, we benchmarked our results against the standard output of the default PhaseFinder algorithm to evaluate concordance, specifically analyzing the read evidence profiles underlying the unique detections identified by each method. Next, to evaluate the impact of sequencing depth on detection reliability, we conducted a downsampling experiment, quantifying the recovery of high-confidence invertons across systematically reduced coverage levels. Subsequently, to assess functional significance, we investigated the genomic context of the identified invertons, specifically testing for enrichment near genes using a contig-preserving permutation framework that controls for local gene density. Finally, analysis of the human gut metagenomic data revealed a small number of putative invertons on high confidence viral contigs, including a complete Microviridae contig with a putative inverton overlapping the major capsid protein gene. We further evaluated these candidates using viral contig classification, genome completeness assessment, evidence from read mapping and coding sequence annotation.
3.1 TPMM helps define a high-confidence inverton set
3.1.1 Output results and comparison with PhaseFinder
We applied TPMM to each candidate to infer its latent component membership and obtained posterior probabilities for the three mixture components. In total, we identified 3,517 invertons in the human gut dataset and 764 in the rat gut dataset. To examine whether the estimated results behave consistently with read evidence, we summarized the empirical distributions of and and constructed bins to ensure adequate coverage across the grid. We then visualized the binned relationship between and read evidence in Fig. 2. As shown in Fig. 2, for a fixed bin, increases with , reflecting stronger support for the inversion. Conversely, for a fixed bin, decreases as increases (equivalently, as decreases). For downstream analyses, we controlled the cumulative BFDR at and obtained positive calls. The detailed calculation procedure, as well as an additional consistency check between and , is provided in Figure S3. Then we compared the TPMM positive calls with PhaseFinder results. Here, we mainly focused on the calls uniquely identified by these two methods. The blue box in Fig. 2 highlights TPMM-unique positives, which concentrate in bins with smaller but relatively higher . To investigate the underlying cause, we examined the detailed read distribution and found that these discrepancies are primarily driven by insufficient absolute read support in one of PhaseFinder’s evidence channels (e.g., paired-end or spanning) due to its minimum supporting-read thresholds. In contrast, the red box marks PhaseFinder-unique calls, which are enriched at larger but lower . To illustrate the most pronounced disagreements between two methods, we calculated the net enrichment of TPMM-unique calls over PhaseFinder-unique calls. Moreover, we provided Venn diagrams for both datasets in Figure S4. For downstream analysis, we defined a conservative reference set as the intersection of TPMM and PhaseFinder calls at full depth, ensuring that subsequent sensitivity comparisons are anchored on events supported by both approaches.
Figure 2.

Heatmap of read evidence and posterior for selected samples from the human gut dataset. : total informative reads for candidate i. : inversion-supporting reads for candidate i. Color: median within each bin. Blank bins: no observations. Blue box: TPMM-unique positives. Red box: PhaseFinder-unique positives. Yellow star: largest enrichment of TPMM-unique positives. Purple star: largest enrichment of PhaseFinder-unique calls.
3.1.2 Comparison with DeepInverton
In addition, we applied DeepInverton to both datasets with default parameters the authors provide. In the human gut dataset, the tool identified 1,310 candidate invertons, yet only 12 possess reverse read support. Similarly, in the rat gut dataset, 4,103 invertons were predicted, none of which were supported by reverse reads. This disconnection between prediction and physical evidence suggests that this tool, which relies on DNA sequences rather than alignment data, lacks the reliability and biological interpretability necessary to accurately identify active DNA inversions.
3.2 TPMM retains higher inverton recovery under downsampling
As mentioned above, most of invertons detected by DeepInverton in both datasets lack reverse read support. Therefore, we only compared TPMM and PhaseFinder in this experiment. To assess robustness under low-depth sequencing conditions, we performed a downsampling analysis on both datasets. For each sample, we randomly subsampled 20%, 40%, and 80% of the original reads (denoted as DS20, DS40, and DS80), respectively. Subsequently, TPMM and PhaseFinder were run independently on each of the six downsampled datasets.
Then, we defined the high-confidence set described in previous section as , representing the intersection of invertons identified by both methods on the full-coverage datasets: . For each down-sampled condition, we mapped the identified invertons to based on genomic coordinates. We then quantified: (i) the fraction of recovered by both methods, and (ii) the proportion of recovered exclusively by either TPMM or PhaseFinder. As illustrated in Fig. 3, the subset of invertons recovered uniquely by TPMM (red bar) was consistently larger than that recovered uniquely by PhaseFinder (blue bar) across both datasets, indicating that TPMM maintains higher sensitivity under reduced sequencing depth.
Figure 3.

Composition of calls recovered under downsampling. DS20, DS40, and DS80 denote downsampled datasets retaining 20%, 40%, and 80% of the original reads. Left: Human dataset; Right: Rat dataset. Stacked bars represent the proportion of identified by both methods (grey), detected only by TPMM (red), and detected only by PhaseFinder (blue).
Specifically, in the human gut dataset, TPMM uniquely recovered 25.54%, 22.15%, and 5.80% of the reference set under DS20, DS40, and DS80, respectively, whereas PhaseFinder uniquely recovered only 1.07%, 1.96%, and 1.68%. Similarly, in the rat gut dataset, TPMM uniquely recovered 49.15%, 44.14%, and 9.21%, compared to 2.54%, 3.90%, and 5.26% for PhaseFinder. Notably, this performance advantage become more pronounced as sequencing depth decreased, demonstrating that TPMM is more robust in preserving the detection of reference invertons when read depth is severely limited. We further evaluated whether the improved recovery of TPMM was accompanied by increased false positive calls. TPMM achieved higher recall without an obvious increase in proxy FDP, and this pattern was consistent across five repeated DS20 downsampling experiments (Supplementary Tables S4 and S5, available as supplementary data at Bioinformatics online).
3.3 Enrichment of downstream genes near invertons
Previous studies have reported that invertons can modulate downstream gene expression (Chatzidaki-Livanis et al. 2010, Taketani et al. 2015). Consequently, we investigated whether the invertons identified by TPMM in two datasets are located upstream of genes more frequently than expected by chance. Detailed definitions of downstream regions are provided in Supplementary Section 5, available as supplementary data at Bioinformatics online. We first performed de novo gene prediction on all contigs using Prodigal (Hyatt et al. 2010) to define coding sequence (CDS) start coordinates. We then quantified the proximity of invertons to downstream CDS start sites and assessed statistical significance using a contig-preserving permutation framework that accounts for contig structure and local gene density.
Briefly, for a given window size w, we determined whether at least one annotated CDS start site falls within w bp downstream of the inverton end coordinate, taking strand orientation into account. Let denote the observed hit rate, defined as the fraction of invertons with at least one downstream CDS start within the specified distance. To generate an appropriate null model, we preserved both the host contig and the length of each inverton. Specifically, for each inverton i with length on contig , we randomly sampled a replacement interval of length uniformly along , ensuring the interval remains within contig boundaries. This procedure was repeated for B replicates to generate a null distribution of hit rates for each window size evaluated.
Figure 4A (human gut) and B (rat gut) display the observed hit rates compared against the null expectation across a continuous range of window sizes. Crucially, the observed hit rate consistently exceeds the null expectation in both datasets, a trend that is particularly pronounced at small window sizes. The gap between the observed data and the null model is widest at the shortest distances, indicating that invertons are preferentially located in the immediate vicinity of downstream coding sequences. While both the observed and null rates naturally increase as the window expands—reflecting the cumulative probability of encountering a gene over larger distances—the distinct enrichment observed at short distances suggests a tight physical linkage and potential regulatory relationship between the identified invertons and their adjacent genes.
Figure 4.

Enrichment of downstream coding sequences near identified invertons. The observed hit rate (, black line) is compared to a null distribution generated by contig-preserving permutation. The red line represents the mean of the null distribution, and the light blue shaded, diagonally hatched area indicates the 95% confidence interval. (A) Analysis of the human dataset. (B) Analysis of the rat dataset.
3.4 Putative invertons identified in viral genomes
Although invertons are well characterized in bacterial genomes, their occurrence in viral genomes recovered from metagenomic data remains poorly understood. Because our initial analyses did not distinguish between bacterial and viral origins, we next performed viral sequence classification to determine whether TPMM detected invertons were on metagenomically assembled viral genomes. Given the high viral diversity and complex environmental pressures of the human gut, we selected the human dataset for this investigation. We classified contigs using four independent viral identification tools: DeepVirFinder (Ren et al. 2020), VirSorter2 (Roux et al. 2015), ViraLM (Peng et al. 2024), and geNomad (Camargo et al. 2024). To ensure high specificity, we retained only those contigs supported by at least three tools, resulting in 11,415 high-confidence viral contigs. Genome quality was further assessed using CheckV (Nayfach et al. 2021). By restricting our analysis to contigs with at least “Medium” completeness, we obtained a final set of 1403 quality-filtered viral contigs. Within this set, we identified putative invertons on 16 contigs, comprising 4 complete, 2 high-quality, and 10 medium-quality genomes.
To examine the genomic context and potential functional relevance of these candidate invertons, we predicted CDSs on the 16 contigs using Prodigal and annotated protein functions via BLASTP. The results reveal that 8 of the 16 candidate invertons are located within 1,000bp of confirmed viral CDS start or fall directly within CDS boundaries (not host-derived) (Fig. 5A). These patterns suggest that, analogous to bacterial invertons which modulate gene expression or protein structure, putative viral invertons may contribute to regulatory or structural variation. Due to the absence of metatranscriptomic data, we could not directly assess expression changes of flanking genes. Therefore, to illustrate potential functional impacts, we examined a representative candidate inverton on contig “SRR2145362_contig9143,” which falls within a CDS boundary. This contig, classified as complete, contains a predicted invertible segment defined by inverted repeats at positions 2653–2666 and 2775–2788. The invertible interval lies entirely within a single CDS (2083–3993). BLASTN analysis revealed extensive alignment to Microviridae reference sequences, and functional annotation identified the CDS as a major capsid protein (MCP). To evaluate the potential consequences of inversion, we computationally generated the reverse-complement of the invertible segment and re-annotated the modified contig. Notably, the original single MCP CDS was predicted as two distinct CDSs after inversion (Fig. 5B). This observation aligns with reports that Microviridae capsid proteins contain conserved modular segments (Kirchberger and Ochman 2023), suggesting that inversion at this locus could reconfigure the MCP coding structure. Given that Microviridae are single-stranded DNA viruses, we hypothesize that this recombination event likely occurs on the double-stranded replicative form during viral replication. Furthermore, we extended our analysis to another human gut dataset and observed similar results, which are detailed in Supplementary Section 10, available as supplementary data at Bioinformatics online.
Figure 5.

Genomic context and impact of identified invertons. (A) Schematic representation of coding sequences (CDSs) located within 1000 bp downstream of or directly spanning the inverton region. Genes are color-coded by predicted function, including hypothetical proteins, major capsid protein (MCP), DNA polymerase, peptidase, and TerD domain-containing proteins. White rectangles with diagonal stripes indicate inverton sequences. (B) Representative example of a Microvirus inverton impacting the major capsid protein (MCP) gene.
4 Conclusion and discussion
In this study, we introduce TPMM, a three-component posterior mixture model designed to improve inverton detection under low-depth metagenomes. Building on the read evidence, TPMM returns posterior probabilities as soft labels, and further enables hard calling under BFDR control, reducing reliance on fixed thresholds when evidence is sparse. Validation on two gut metagenomic datasets demonstrates that TPMM recovers more invertons as read counts decrease, demonstrating practical utility in low-depth settings. Furthermore, invertons detected by TPMM exhibit clear enrichment in downstream genes, supporting their biological relevance. Together, these results suggest that a well-specified probabilistic model provides a reliable and interpretable alternative to deep learning-based approaches. Notably, in the human gut dataset, 3152 of the 3517 TPMM-detected invertons were detected in Crohn’s disease samples, whereas 365 were detected in control samples. While this distribution may be influenced by sample imbalance, it suggests that an elevated number of invertons can potentially serve as a disease biomarker in future studies with balanced sample sizes. In addition, we identified putative invertons on metagenomically assembled viral genomes and provide supporting analyses and evidence. To our knowledge, this is the first report of putative invertons identified from metagenomically assembled viral genomes in the human gut, expanding the potential scope of inversion mediated regulation.
TPMM also has limitations that motivate future improvements. First, small candidate sets may limit empirical information for mixture-model fitting, reducing paramter stability. Second, TPMM remains dependent on read evidence and cannot address candidates with no reverse read support. For such cases, TPMM cannot produce a reliable detection output. Third, TPMM detects phase variation but does not estimate forward and reverse orientation proportions, which would require a calibrated model accounting for coverage, mapping bias, and related factors. Future work may incorporate stronger regularization or hierarchical modeling to improve robustness with small candidate sets, and explore integration with sequence-level signals to enable more comprehensive inference.
Supplementary Material
Acknowledgement
We thank Prof. Fuyong Li and his student Yi Yang from Zhejiang University for the discussion about the problem formulation.
Contributor Information
Yi Lu, Department of Electrical Engineering, City University of Hong Kong, Hong Kong SAR, China.
Jiaojiao Guan, Department of Electrical Engineering, City University of Hong Kong, Hong Kong SAR, China.
Yang Shen, Department of Electrical Engineering, City University of Hong Kong, Hong Kong SAR, China.
Jiayu Shang, Department of Information Engineering, Chinese University of Hong Kong, Hong Kong SAR, China.
Yanni Sun, Department of Electrical Engineering, City University of Hong Kong, Hong Kong SAR, China.
Supplementary material
Supplementary material is available at Bioinformatics online.
Conflicts of interest
The authors have no conflicts of interest to declare.
Funding
This work is supported by the Hong Kong Research Grants Council (RGC) General Research Fund (GRF) [11209823], the City University of Hong Kong projects [9667256, 9678241, 7020092] and Institute of Digital Medicine.
Data availability
All data and codes used for this study are available online via: https://github.com/KennyxxD/TPMM
Author contributions
Yi Lu (Conceptualization, Data curation, Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing), Jiaojiao Guan (Data curation, Formal analysis, Writing – review & editing), Yang Shen (Formal analysis, Methodology, Writing – review & editing), Jiayu Shang (Methodology, Writing – review & editing), and Yanni Sun (Conceptualization, Formal analysis, Funding acquisition, Methodology, Supervision, Validation, Writing – review & editing)
References
- Ben-Assa N, Coyne MJ, Fomenkov A et al. Analysis of a phase-variable restriction modification system of the human gut symbiont Bacteroides fragilis. Nucleic Acids Res 2020;48:11040–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Camargo AP, Roux S, Schulz F et al. Identification of mobile genetic elements with genomad. Nat Biotechnol 2024;42:1303–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chatzidaki-Livanis M, Weinacht KG, Comstock LE. Trans locus inhibitors limit concomitant polysaccharide synthesis in the human gut symbiont Bacteroides fragilis. Proc Natl Acad Sci USA 2010;107:11976–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Coyne MJ, Chatzidaki-Livanis M, Paoletti LC et al. Role of glycan synthesis in colonization of the mammalian gut by the bacterial symbiont Bacteroides fragilis. Proc Natl Acad Sci USA 2008;105:13099–104. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Giphart-Gassler M, Plasterk RHA, van de Putte P. G inversion in bacteriophage Mu: a novel way of gene splicing. Nature 1982;297:339–42. [DOI] [PubMed] [Google Scholar]
- Grundy FJ, Howe MM. Involvement of the invertible g segment in bacteriophage mu tail fiber biosynthesis. Virology 1984;134:296–317. [DOI] [PubMed] [Google Scholar]
- Henderson IR, Owen P, Nataro JP. Molecular switches – the on and off of bacterial phase variation. Mol Microbiol 1999;33:919–32. [DOI] [PubMed] [Google Scholar]
- Hyatt D, Chen G-L, LoCascio PF et al. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics 2010;11:119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jiang X, Hall AB, Arthur TD et al. Invertible promoters mediate bacterial phase variation, antibiotic resistance, and host adaptation in the gut. Science 2019;363:181–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kirchberger PC, Ochman H. Microviruses: a world beyond phix174. Annu Rev Virol 2023;10:99–118. [DOI] [PubMed] [Google Scholar]
- Lewis JD, Chen EZ, Baldassano RN et al. Inflammation, antibiotics, and diet as environmental stressors of the gut microbiome in pediatric Crohn’s disease. Cell Host Microbe 2015;18:489–500. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nayfach S, Camargo AP, Schulz F et al. Checkv assesses the quality and completeness of metagenome-assembled viral genomes. Nat Biotechnol 2021;39:578–85. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Newton MA, Noueiry A, Sarkar D et al. Detecting differential gene expression with a semiparametric hierarchical mixture method. Biostatistics 2004;5:155–76. [DOI] [PubMed] [Google Scholar]
- Peng C, Shang J, Guan J et al. ViraLM: empowering virus discovery through the genome foundation model. Bioinformatics 2024;40:btae704. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Porter NT, Hryckowian AJ, Merrill BD et al. Phase-variable capsular polysaccharides and lipoproteins modify bacteriophage susceptibility in bacteroides thetaiotaomicron. Nat Microbiol 2020;5:1170–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ren J, Song K, Deng C et al. Identifying viruses from metagenomic data using deep learning. Quant Biol 2020;8:64–77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Roche-Hakansson H, Chatzidaki-Livanis M, Coyne MJ et al. Bacteroides fragilis synthesizes a DNA invertase affecting both a local and a distant region. J Bacteriol 2007;189:2119–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Roux S, Enault F, Hurwitz BL et al. VirSorter: mining viral signal from microbial genomic data. PeerJ 2015;3:e985. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sandulache R, Prehm P, Kamp D. Cell wall receptor for bacteriophage mu g(+). J Bacteriol 1984;160:299–303. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Taketani M, Donia MS, Jacobson AN et al. A phase-variable surface layer from the gut symbiont Bacteroides thetaiotaomicron. mBio 2015;6:e01339–15. 10.1128/mbio.01339-15 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Troy EB, Carey VJ, Kasper DL et al. Orientations of the Bacteroides fragilis capsular polysaccharide biosynthesis locus promoters during symbiosis and infection. J Bacteriol 2010;192:5832–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Trzilova D, Tamayo R. Site-specific recombination – how simple dna inversions produce complex phenotypic heterogeneity in bacterial populations. Trends Genet 2021;37:59–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Van de Putte P, Goosen N. DNA inversions in phages and bacteria. Trends Genet 1992;8:457–62. [DOI] [PubMed] [Google Scholar]
- van der Woude MW, Bäumler AJ. Phase and antigenic variation in bacteria. Clin Microbiol Rev 2004;17:581–611, table of contents. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wen C, Guan J, Uea-Anuwong T et al. Dissecting the gut microbial communities and resistomes of wild rats from different ecological areas in Hong Kong. Environ Res 2025;283:122108. [DOI] [PubMed] [Google Scholar]
- Wen J, Zhang H, Chu D et al. Deep learning revealed the distribution and evolution patterns for invertible promoters across bacterial lineages. Nucleic Acids Res 2024;52:12817–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yan W, Hall AB, Jiang X. Bacteroidales species in the human gut are a reservoir of antibiotic resistance genes regulated by invertible promoters. NPJ Biofilms Microbiomes 2022;8:1. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data and codes used for this study are available online via: https://github.com/KennyxxD/TPMM
