Abstract
Investigating the lethal effect of multi-gene knockout is essential for discovering novel antibiotic targets and metabolic engineering. Unlike single genes or gene pairs, three-gene combinations involve more intricate interactions, making experimental screening time-consuming. Computational methods, particularly Genome-scale metabolic model (GEM)-based Flux balance analysis (FBA), requires constructing new GEMs from experimental data, limiting its use for new species. Moreover, using FBA for three-gene knockout screening could take several years. Therefore, a faster and GEM-independent approach is needed to facilitate genome-wide three-gene knockout screening. Here, we introduce Tripleknock for predicting the lethal effects of three-gene knockouts. Tripleknock was trained using genome-derived protein sequence features from Escherichia coli K-12 MG1655, and three-gene knockout simulations using FBA. The model uses a threshold of 90% reduction in cell growth to define lethal effect as the prediction output. Compared to FBA, Tripleknock yields predictions ~ 20 × faster and achieves an average cross-species F1 score of 0.77 on six Enterobacteriaceae pathogens (FBA-defined labels). To mitigate information leakage, we additionally evaluated Tripleknock under a strict gene-level disjoint Train/Val/Test split (genes left out). We further benchmarked against an essentiality-rule baseline and literature-curated (n = 37) E. coli triple perturbations for external validation; no false positives were observed among predicted lethal cases (FP = 0) in this small curated set. To the best of our knowledge, this provides a reproducible baseline and evaluation framework for bacterial triple-knockout lethality prediction.
Supplementary Information
The online version contains supplementary material available at 10.1038/s41598-026-46272-9.
Keywords: Triple-gene knockout, Lethal effect, Bacteria, Antibiotic design, Deep learning
Subject terms: Biotechnology, Computational biology and bioinformatics, Microbiology
Introduction
Gene knockout is a widely-used method to investigate the relationship between genotype and phenotype. In metabolic engineering, gene knockouts can generate bacterial strains with increased production of specific metabolites1, while in antibiotic design, gene-knockout screening can facilitate the discovery of novel drug targets2. Single-gene knockout studies have paved the way for identifying genes important for survival and growth. These genes are critical for maintaining basic cellular functions, such as metabolism, DNA replication, and cellular integrity. Absence of essential genes will disrupt cellular processes, typically halting growth or leading to cell death3–5. The identification of essential genes has been primarily achieved through experimental methods such as gene editing, RNA interference, and CRISPR/Cas96,7. In addition to experimental methods, computational approaches have also been developed to predict essential genes across species based on genomic data8,9.
Compared with single-gene knockout, two-gene knockout can reveal novel genetic interactions that are not evident in single-gene knockout studies. Among these interactions, Synthetic Lethal (SL)10 and Synthetic Rescue (SR)11 are two important types of gene interaction. SL is defined as when the knockout of both genes results in a severe loss of cell growth or even cell death, while the knockout of either gene does not lead to this inactivation, indicating that both genes are functionally interdependent. On the other hand, SR refers to a situation where the inhibition of one gene (either gene A or gene B) alone reduces cell growth, but the simultaneous inhibition of both genes rescues the phenotype, thus allowing the cell to survive. SR is also a potential new mechanism in cancer drug resistance12–14. These interactions provide valuable insights into gene dependencies and redundancies in cellular networks. Algorithms and experimental methods for understanding SL and SR effects are also reported15–17.
Expanding to three-gene knockouts, the complexity of genetic interactions increases even further compared to two-gene interactions. Exploring three-gene knockout combination is important and has shown promise in metabolic engineering and therapeutic interventions. For example, a three-gene knockout has also been reported to enhance polyhydroxyalkanoate (PHA) production in E. coli18. These interactions could result in unexpected phenotypic effects, and understanding them may provide new avenues for therapeutic interventions. For pathogenic species, investigating the lethal effect of multi-gene knockout could facilitate the finding of new antibiotic targets. However, multi-gene knockout remains a significant challenge, for the difficulties of accessing sufficient experimental data. And the lack of experimental data limits the development of computational methods. Fortunately, GEMs coupled with FBA19,20 provide reliable results for simulating cell growth when genes are knocked out. But, although it’s the most popular simulation methods in the past over 20 years, it faces limitations when applied to multi-gene knockout. First, the construction of GEMs for each species is labor-intensive and time-consuming21–24. Even if the new species is closely related to a species that already has GEM, a new GEM must be built de novo, which makes it difficult to generalize the results of gene knockout across related species. Second, GEMs only include approximately one-third of genes in the genome, which means the searching space is largely limited. Recent advancements in deep learning have shown promise in predicting gene knockout effects without relying on GEM construction25–27. Given these constraints, an accurate, GEM-independent and user-friendly prediction model for multi-gene’s lethal effects would be a powerful tool for biologists who are working on metabolic engineering or multi-target antibiotics finding.
In response to these challenges, we developed Tripleknock to predict FBA-defined lethal outcomes of three-gene knockouts from genome-derived protein sequence features. We evaluate the model across related Enterobacteriaceae species and, critically, under leakage-aware gene-disjoint splits where entire genes are left out. We further compare against a simple essentiality-rule baseline and perform literature-based experimental validation to assess real-world consistency.
Results
Overview of Tripleknock’s algorithmic architecture
The Tripleknock workflow consists of three main components (Fig. 1). First, a two-stage pre-trained autoencoder is used for genome compression (Fig. 1a). This tandem autoencoder model is designed to process genomes with triple gene knockouts and sequentially reduces both the dimensionality of gene features and the number of genes, resulting in an approximately 1,000-fold reduction in genome matrix size. Second, phenotype labels are generated using FBA, where all feasible three-gene knockout combinations in the GEM of E. coli K-12 MG1655 are simulated and categorized based on their predicted growth outcomes (Fig. 1b). Third, a classification model based on a multilayer perceptron (MLP) is used to predict the viability of triple knockout genotypes (Fig. 1c). This classifier receives a combined feature consisting of the compressed genome and the triple knockout genes, and processes it through fully connected layers to produce a binary prediction. Both the autoencoder and the MLP models are available for direct download and application for users.
Fig. 1.
Workflow of Tripleknock. (a) Construction of 2-mer encoded protein feature matrices and genome-level compression via a two-stage autoencoder pipeline. (b) Feature engineering and three-gene knockout simulation by FBA. (c) Predictions using concatenated features of knocked-out genes and genomes, done by MLP-based binary classification of cell growth. FC, fully connected layer. Yellow boxes refer to feature engineering, green boxes indicates the layers in three sub architectures, blue boxes indicate the labeling method, and the purple boxes are the final binary output of Tripleknock.
We plotted the loss function of the three sub-models during training (Supplementary Fig. 2). Autoencoder 1 (AE1) demonstrates rapid convergence. A horizontal dashed green line indicates the theoretical background noise level (1/400 = 0.0025), which reflects the average frequency of a 2-mer assuming uniform distribution (Tables 1, 2). The model’s reconstruction error is significantly below this threshold, indicating that the autoencoder captures non-random structural patterns in the input data, which may reflect biologically relevant features (Supplementary Fig. 2a). AE2 also achieves low reconstruction error, indicating that it captures non-random structure in the genome-level input representation (Supplementary Fig. 2b).
Table 1.
Gene-level disjoint (genes left out): train/val/test three-way disjoint (stricter).
| Training size (N) | Test AUC (mean ± sd) | Val AUC (mean ± sd) |
|---|---|---|
| 1 k | 0.589 ± 0.053 | 0.603 ± 0.052 |
| 5 k | 0.616 ± 0.047 | 0.599 ± 0.045 |
| 10 k | 0.607 ± 0.021 | 0.607 ± 0.024 |
| 20 k | 0.597 ± 0.025 | 0.582 ± 0.030 |
| 40 k | 0.610 ± 0.036 | 0.584 ± 0.026 |
| 100 k | 0.582 ± 0.023 | 0.586 ± 0.028 |
| 1000 k | 0.583 ± 0.026 | 0.583 ± 0.033 |
| 2000 k | 0.586 ± 0.030 | 0.582 ± 0.034 |
| 3000 k | 0.583 ± 0.034 | 0.579 ± 0.033 |
Table 2.
Stratified performance on the 3000 k gene-level disjoint test set by confidence threshold (t) and essential-gene count (n_essential).
| t | n_essential | N_total | Coverage (mean ± sd) | ACCURACY (mean ± sd) |
|---|---|---|---|---|
| 0.1 | 0 | 1,921,859 | 0.641 ± 0.026 | 0.559 ± 0.027 |
| 0.1 | 1 | 855,784 | 0.285 ± 0.017 | 0.546 ± 0.051 |
| 0.1 | 2 | 124,739 | 0.042 ± 0.007 | 0.643 ± 0.078 |
| 0.1 | 3 | 5980 | 0.002 ± 0.001 | 0.724 ± 0.094 |
| 0.2 | 0 | 1,857,364 | 0.619 ± 0.027 | 0.561 ± 0.028 |
| 0.2 | 1 | 825,239 | 0.275 ± 0.017 | 0.548 ± 0.053 |
| 0.2 | 2 | 120,309 | 0.040 ± 0.007 | 0.648 ± 0.079 |
| 0.2 | 3 | 5787 | 0.002 ± 0.001 | 0.730 ± 0.094 |
| 0.3 | 0 | 1,777,888 | 0.593 ± 0.027 | 0.563 ± 0.029 |
| 0.3 | 1 | 787,967 | 0.263 ± 0.016 | 0.550 ± 0.055 |
| 0.3 | 2 | 114,972 | 0.038 ± 0.006 | 0.654 ± 0.081 |
| 0.3 | 3 | 5542 | 0.002 ± 0.001 | 0.741 ± 0.094 |
| 0.4 | 0 | 1,656,588 | 0.552 ± 0.028 | 0.566 ± 0.031 |
| 0.4 | 1 | 730,799 | 0.244 ± 0.014 | 0.554 ± 0.058 |
| 0.4 | 2 | 106,666 | 0.036 ± 0.006 | 0.663 ± 0.085 |
| 0.4 | 3 | 5149 | 0.002 ± 0.001 | 0.755 ± 0.091 |
Bold indicates the highest value within each subgroup.
For MLP training, we estimated that the time required for generating and storing all intermediate feature files during training will be approximately 70 days; considering cross-validation, this will increase to near 1 year. Therefore, only random 5% and 10% subsets were used to train the model (Supplementary Fig. 2c-d) at combination level. In both cases, the training on 5% and 10% subsets yielded similar training curves, justified the use of subsampling method. The time of data preparation (FBA simulation) and training of sub-models within Tripleknock framework is collected (Table 3). The combination-level model is released for users, whereas the strict gene-disjoint setting is used to provide a leakage-aware estimate of generalization to unseen genes.
Table 3.
FBA simulation and training time of models in Tripleknock.
| Step | Time/day |
|---|---|
| FBA simulation | ~ 100 (100-threads) |
| Model training | |
| Autoencoder 1 | 8.15 |
| Autoencoder 2 | 19.72 |
| MLP 5% | ~ 5 |
| MLP 10% | ~ 10 |
| Total training time | ~ 30–40 |
We further considered the risk that the combination-level cross-validation can allow the same genes (and gene pairs) to appear in both training and evaluation splits. This can substantially inflate performance estimates. We adopt a leakage-aware evaluation where entire genes are left out and the gene sets are mutually exclusive across Train/Val/Test. Under this strict protocol (Table 1), validation and test performance remain comparable and stays consistently above random despite the strict unseen-gene setting across training scales; at 3000 k samples, Val AUC is 0.579 ± 0.033 and Test AUC is 0.583 ± 0.034. Details of the more gene pair level disjoint and combination-level setting and their inflated estimates are provided in the Supplementary for completeness.
We further stratified the strict gene-disjoint test set by prediction confidence (|y_score − 0.5|≥ t) and essential-gene count (Table 2). Higher accuracy is observed in small high-confidence strata (e.g., n_essential = 3 at t = 0.4), and we interpret this as a diagnostic of structured errors rather than an overall performance gain.
Baseline comparison on “new” synthetic-lethal triples
We compared Tripleknock with an essentiality-rule baseline that predicts a triple as lethal if any member is single-gene lethal. In exhaustive screening, 380,719,699 triples (~ 67%) contain no single-gene lethal member, yet this subset includes 123,630 lethal triples. By construction, the baseline has near-zero recall on these ‘new’ lethal triples. Because it reduces prediction to single-gene essentiality, this rule cannot capture synthetic-lethal triples composed entirely of non-essential genes, whereas Tripleknock achieves substantially higher AUC (Table 4). Note that Table 4 uses random fivefold CV within E. coli (i.e., a non–gene-disjoint, in-distribution setting) and is intended as a diagnostic comparison to isolate the limitation of the essentiality-rule baseline. Leakage-aware generalization to unseen genes is evaluated separately under the strict gene-disjoint protocol in Table 1.
Table 4.
Comparison with the essential-gene rule baseline (Random fivefold CV within E. coli, combination-level, same splits for both methods). *Dataset 1: Mixed no-essential and with-essential data. *Dataset 2: No-essential gene, ‘new triples’.
| Dataset setting | N_total | Tripleknock AUC (mean ± sd) | Baseline AUC (mean ± sd) |
|---|---|---|---|
| Dataset 1 | 494,520 | 0.998 ± 0.000 | 0.750 ± 0.001 |
| Dataset 2 | 247,260 | 0.999 ± 0.000 | 0.500 ± 0.000 |
Literature-based experimental validation
Because no public database systematically reports bacterial triple-lethal knockouts, we curated 37 E. coli triple-gene perturbations from published studies. We assigned labels based on the source text (viable vs. not obtainable/nonviable) to determine whether each perturbation was lethal. These source texts, along with the triple-gene data and references, have been uploaded to GitHub. On this literature-based set, Tripleknock achieves TN = 23, FP = 0, FN = 5, TP = 9 (Accuracy = 0.865; ROC-AUC = 0.821). Several false negatives may reflect label uncertainty and protocol mismatch: ‘unable to obtain/construct’ can arise from technical constraints rather than measured lethality, and experimental growth criteria may not align with our FBA-based 90% growth cutoff. This indicates a conservative behavior in this small set, with reliable lethal calls (FP = 0). Differences in strain background, media, and phenotype definitions mean this validation is necessarily approximate, and we treat it as proof-of-concept external consistency (Table 5; confusion matrix Fig. 2).
Table 5.
Performance of Tripleknock on published experimental data.
| Gene A | b_number A | Gene B | b_number B | Gene C | b_number C | Real label | Triple knock | Species | References |
|---|---|---|---|---|---|---|---|---|---|
| argD | b3359 | astC | b1748 | gabT | b2662 | 0 | 0 | E. coli K-12 | 28 |
| recA | b2699 | deoR | b0840 | nupG | b2964 | 0 | 0 | E. coli VH33 | 29 |
| lpxL | b1054 | lpxM | b1855 | lpxP | b2378 | 0 | 0 | E. coli | 30 |
| rpoS | b2741 | rmf | b0953 | ompC | b2215 | 1 | 1 | E. coli K-12 | 31 |
| dnaA | b3702 | priC | b0467 | rnhA | b0214 | 0 | 0 | E. coli K-12 | 32 |
| trk | b3849 | kup | b3747 | kdp | b0698 | 0 | 0 | E. coli | 33 |
| trkA | b3290 | argE | b3957 | holD | b4372 | 0 | 0 | E. coli | 33 |
| holD | b4372 | dinB | b0231 | polB | b0060 | 0 | 0 | E. coli | 33 |
| pncA | b1768 | xapA | b2407 | nadC | b0109 | 0 | 0 | E. coli BW25113 | 34 |
| ybiS | b0819 | erfK | b1990 | ycfS | b1113 | 0 | 0 | E. coli | 35 |
| hemF | b2436 | ybdK | b0581 | gdhA | b1761 | 0 | 0 | E. coli BW25113 | 36 |
| hemF | b2436 | ybdK | b0581 | gadB | b1493 | 0 | 0 | E. coli BW25113 | 36 |
| hemF | b2436 | gdhA | b1761 | gadB | b1493 | 0 | 0 | E. coli BW25113 | 36 |
| pepA | b4260 | pepD | b0237 | pepN | b0932 | 0 | 0 | E. coli BL21 | 37 |
| uvrA | b4058 | xth | b1749 | nfo | b2159 | 0 | 0 | E. coli | 38 |
| xth | b1749 | nfi | b3998 | nfo | b2159 | 0 | 0 | E. coli | 38 |
| nth | b1633 | nfo | b2159 | xth | b1749 | 1 | 1 | E. coli BW535 | 39 |
| dacA | b0632 | dacB | b3182 | dacC | b0839 | 0 | 0 | E. coli | 40 |
| secB | b3609 | tig | b0436 | dnaJ | b0015 | 0 | 0 | E. coli | 41 |
| ykfI | b0245 | ypjF | b2646 | cbtA | b2005 | 0 | 0 | E. coli K-12 | 42 |
| yafW | b0246 | cbeA | b2004 | yfjZ | b2645 | 0 | 0 | E. coli K-12 | 42 |
| clcB | b1592 | nhaA | b0019 | nhaB | b1186 | 0 | 0 | E. coli | 43 |
| polB | b0060 | dinB | b0231 | umuD | b1183 | 0 | 0 | E. coli | 44 |
| polB | b0060 | dinB | b0231 | umuC | b1184 | 0 | 0 | E. coli | 44 |
| dnaA | b3702 | recB | b2820 | rep | b3778 | 1 | 1 | E. coli | 45 |
| rep | b3778 | dnaA | b3702 | recC | b2822 | 1 | 0 | E. coli | 45 |
| amiA | b2435 | yciM | b1280 | amiB | b4169 | 1 | 1 | E. coli | 46 |
| dut | b3640 | degP | b0161 | pstC | b3727 | 1 | 0 | E. coli | 47 |
| pitA | b3493 | ugp | b3452 | pst | b3726 | 1 | 0 | E. coli | 48 |
| uvrD | b3813 | ruvA | b1861 | recQ | b3822 | 1 | 1 | E. coli | 49 |
| uvrD | b3813 | ruvB | b1860 | recQ | b3822 | 1 | 1 | E. coli | 49 |
| bamE | b2617 | skp | b0178 | degP | b0161 | 1 | 0 | E. coli | 50 |
| bamC | b2477 | skp | b0178 | degP | b0161 | 0 | 0 | E. coli | 50 |
| rarA | b0892 | recG | b3652 | ruvB | b1860 | 1 | 1 | E. coli | 51 |
| rarA | b0892 | ruvB | b1860 | recQ | b3822 | 1 | 1 | E. coli | 51 |
| rnhA | b0214 | rnhB | b0183 | recA | b2699 | 1 | 1 | E. coli | 52 |
| yhdP | b4472 | tamB | b4221 | ydbH | b1381 | 1 | 0 | E. coli | 53 |
Fig. 2.

Confusion matrix of Tripleknock on experimental data.
Cross-species prediction
We evaluated the cross-species generalizability of the Tripleknock framework by applying it to six pathogenic bacterial species within the Enterobacteriaceae family. For each species, we generated three independent batches of 20,000 unique triple-gene knockouts (60,000 total) and simulated growth for each batch using FBA. The 10% training version of Tripleknock was then used to predict lethal effect based on the amino acid sequences of six pathogenic bacterial species. The resulting average F1 scores for each species were mapped onto a phylogenetic tree (Fig. 3a), demonstrating that Tripleknock achieved an overall cross-species average F1 score of 0.77, with particularly strong performance observed in E. coli and Shigella, where F1 scores exceeded 0.83. To investigate the impact of training data size on predictive performance, we compared classification metrics (Accuracy, Precision, Recall, and F1-score) across models trained with 5% and 10% of the data (Fig. 3b). The differences were not statistically significant according to the Wilcoxon signed-rank test, suggesting that, for this cross-species FBA-simulation benchmark, using 10% of the E. coli training triples provides a practical trade-off between training cost and performance. Because cross-species evaluation uses genomes from other species, all test genes are unseen relative to the E. coli training genome; therefore, this benchmark complements (but does not replace) the leakage-aware unseen-gene evaluation in Table 1. In practice, many genes in these related Enterobacteriaceae species have close homologs to E. coli genes, so this cross-species benchmark reflects near-homology transfer and is typically less adversarial than the strict gene-disjoint split within E. coli. Accordingly, all subsequent interpretability analyses were conducted using models trained on 10% of the data, and the version of Tripleknock released for public use is likewise based on this setting. In terms of computational efficiency, Tripleknock delivers predictions consistent with FBA outputs approximately 20 to 30 times faster (depending on hardware) when executed on an L4 GPU within the Google Colab environment. Testing on Colab also suggested the usage and capability of Tripleknock for users.
Fig. 3.
Cross-species prediction. (a) Phylogenetic tree of the seven selected species within the Enterobacteriaceae family. Branch lengths are indicated above the grey lines, and confidence values are shown below. Genome IDs are displayed to the right of each node. Average F1 scores for each species are annotated in red. The training data usage is 10%. (b) Performance of Tripleknock trained on 5% and 10% of E. coli K-12 data and evaluated on six pathogenic species. Error bars represent the standard error of the mean. The average values are labeled above each bar. Statistical significance between the two training ratios for each species was assessed using the Wilcoxon signed-rank test.
Model interpretation
We evaluated interpretability on three representative test species based on their prediction performance: the species with the highest F1 score, the one with the lowest F1 score, and a randomly selected Shigella species. We present volcano plots (Fig. 4a–c) as descriptive class separability in the normalized input space (mean difference and Welch’s t-test), emphasizing that these statistics do not quantify feature importance for a nonlinear MLP. To obtain model-consistent attributions, we computed Integrated Gradients with respect to the model logit and summarized attributions at the predefined block level (KO1/KO2/KO3/AE), and we corroborated these findings using grouped permutation importance measured by AUC drop (Fig. 4d–i).
Fig. 4.
Data separability and model-based attribution across three test species. (a-c) Volcano plots showing feature separability in the normalized input space (descriptive statistics, not MLP feature importance). The x-axis represents the mean difference between the two groups, the y-axis shows the -log10-transformed p-value from a Welch’s t-test. The horizontal blue dashlines indicate the p-value 0.01. (d-f) Integrated Gradients (IG) attribution aggregated at the feature-block level, computed with respect to the model logit and summarized as mean absolute attribution (normalized to proportions) for KO1 (1–400), KO2 (401–800), KO3 (801–1200), and AE (1201–2400). (g-i) Grouped permutation importance measured as AUC drop after permuting each feature block across samples; larger drops indicate stronger reliance on that block.
Discussion
Three-gene knockout screens for lethal phenotypes are valuable for metabolic engineering and are increasingly important in synthetic-lethal–based disease therapy. However, there remain limited end-to-end approaches for bacterial three-gene knockout outcome prediction. Two major limitations of GEM-based FBA are the time required to screen combinatorial knockouts and the strong dependence on a well-curated, organism-specific GEM. To address these constraints and the growing interest in multi-gene phenotype prediction, we used lethality as a representative phenotype and performed a comprehensive GEM-based in silico screen of all possible three-gene knockouts in E. coli K-12, collecting post-knockout growth outcomes as training data. We developed Tripleknock to predict FBA-defined lethal outcomes of three-gene knockouts, and we further assessed external consistency on 37 literature-curated experimental perturbations and across six pathogenic species.
Our interpretability analyses provide complementary views of the data structure and the model’s decision process. The volcano plots (descriptive separability) show that, as cross-species performance decreases, feature-wise differences between positive and negative classes shrink and more points collapse toward the central axis (Δμ ≈ 0), consistent with weakened class separability and increased distribution shift in the target species. In contrast, model-based attribution (Integrated Gradients on the logit and grouped permutation importance) indicates that, in the current training setting, predictions are driven primarily by features derived from the three knocked-out genes, while the genome-context (AE-derived) block contributes marginally on these test species. This does not imply that genome-wide context is biologically irrelevant; rather, it suggests that the present embedding and training objective may not yet extract transferable contextual signals beyond the KO genes. These findings suggest a future direction that a deeper understanding of genome-wide features is essential to improve the model’s transferability across species54.
We aim to produce a user-friendly tool for experimental biologists who may not familiar with coding, therefore, during the exploration of features integration, we weighed two choices: (1) Whether to incorporate the gene–reaction–metabolite graph information from GEMs. We didn’t because manage to make predictions without depending on GEM. The use of GEM presents a challenge, as constructing a high-quality GEM model for a new species is not straightforward. (2) Whether to use both DNA and protein sequences for feature extraction. In small-scale tests, adding DNA information did not provide a clear improvement. Also, because whole-genome sequencing mainly uses CDS, DNA and protein sequences contain substantial overlapping and redundant information. Therefore, we chose to use protein-sequence 2-mer features as the raw input, and after integration and normalization, we fed them into the neural network. The above considerations make sure Tripleknock to function end-to-end and is friendly for biologists who are not familiar with coding, solely based on an annotated genome file, without relying on external or supplementary materials. Some other questions will also be considered. For instance, this study employs a binary classification approach using solely sequence data, which may be inadequate for more complex predictive tasks. Moreover, our prediction does not incorporate experimental metabolic data, such as multi-omics datasets. It is estimated that integrating these diverse data sources could enhance the accuracy of cross-species predictions. Under the strict gene-disjoint Train/Val/Test protocol, generalization to entirely unseen genes remains modest (AUC ~ 0.58), reflecting the inherent challenges of OOD prediction from sequence-only features and motivating future incorporation of richer functional/contextual information. Sequence similarity represents just one of many conserved features across different species; other factors, such as metabolic network-based information55, also contribute significantly to cross-species prediction. Therefore, in future work, we will also incorporate metabolic network structure—beyond sequence conservation—to improve prediction accuracy and cross-species generalization.
In summary, to the best of our knowledge, we developed an end-to-end deep learning framework Tripleknock, which provides a reproducible framework for triple-knockout lethality prediction and benchmarking. While we acknowledge the performance boundaries under strict unseen-gene evaluation, Tripleknock remains an effective and user-friendly tool for biologists who are working on three-target antibiotics finding, triple-gene lethal screening and related metabolic engineering areas.
Methods
Computational environment
The FBA simulation and model training were performed on a server running an Ubuntu 18.04.2 LTS system. The hardware included 4 Intel(R) Xeon(R) Gold 6252 CPUs operating at 2.10 GHz, and 2 NVIDIA A100-PCIE-40 GB GPUs. The software environment was based on Python 3.8.19, CUDA Version 11.7. Cross-species evaluations were conducted on Google Colab using an L-4 GPU. Model interpretability was conducted on Google Colab using a CPU. No software version-dependent errors were encountered, demonstrating the robustness and portability of Tripleknock.
Datasets
E. coli K-12 MG1655, a non-pathogenic strain, was used for training. For testing, we selected one E. coli strain, one Salmonella enterica, three Shigella species, and one Klebsiella, all of which are pathogenic. These seven species belong to the Enterobacteriaceae family (Table 6). The GEMs were downloaded from the Biochemical, Genetic and Genomic knowledgebase (BiGG), a database of high-quality GEMs56. The protein Coding DNA Sequences (CDS) for each species were downloaded from the National Center for Biotechnology Information (NCBI) in FASTA protein format. A horizontal bar plot shows the number of genes annotated in the genome (CDS genes) versus those included in the GEMs across seven species (Supplementary Fig. 1). All of the distributions of protein length of the seven species are right-skewed, indicating that most proteins are relatively short, typically under 600 amino acids. A long tail exists, representing a small number of much longer proteins, some exceeding 2000 amino acids.
Table 6.
Selected bacteria in training and test set.
| Organism | BiGG | NCBI | BV-BRC | Pathogenic |
|---|---|---|---|---|
| E. coli str. K-12 substr. MG1655 | iML1515 | NC_000913.3 | 511,145.12 | No |
| E. coli O157:H7 str. EDL933 | iZ_1308 | AE005174.2 | 155,864.43 | Yes |
| Salmonella enterica subsp. enterica serovar Typhimurium str. LT2 | STM_v1_0 | AE006468.2 | 99,287.12 | Yes |
| Shigella boydii CDC 3083–94 | iSbBS512_1146 | CP001063.1 | 344,609.11 | Yes |
| Shigella dysenteriae Sd197 | iSDY_1059 | CP000034.1 | 300,267.13 | Yes |
| Shigella flexneri 2a str. 301 | iSF_1195 | AE005674.2 | 198,214.7 | Yes |
| K. pneumoniae subsp. pneumoniae MGH 78,578 | iYL1228 | CP000647.1 | 272,620.9 | Yes |
The ‘Bacterial Genome Tree’ tool in the Bacterial and Viral Bioinformatics Resource Center (BV-BRC)57 was used for building the reference phylogenetic tree for cross-species prediction (Supplementary Table 1). This was used for cross-species prediction, during which we used Tripleknock to receive an annotated genome sequence to make direct three-gene knockout prediction on these related species.
Pretrained tandem autoencoders
The first process generated all two-gene pairs from a total of 4305 genes (
), in the E. coli K-12 MG1655 genome. The total number of combinations is 9,264,360. Removing each two-gene combination from
will lead to an n-2 gene set, which is represented as a matrix whose shape is 4303 × 400. The AE1 was trained to compress gene feature vectors from 400 dimensions down to 3, and the AE2 reduces the dimensionality of gene number to 400. The training process used all data over 6 epochs, with a batch size of 2000. The Eq. (1) is the loss function used for training is Mean Squared Error (MSE) to minimize the reconstruction error between the input and the output.
![]() |
1 |
where:
indexes scalar entries of the inputN is the number of input gene features in AE1, the number of genes in AE2.
is the original input.
is the reconstructed output from the autoencoder (the predicted feature vector).
To accommodate Tripleknock for potential single-gene knockout on E. coli MG1655, we explicitly add one zero vector row to the input matrix (making it 4304 × 400). This ensures that, when simulating single-gene knockouts for a genome of 4305 genes, every gene in E. coli K-12 MG1655 is represented and the matrix multiplication aligns with the model’s input dimensions. Without adding zero rows, predicting a single-gene knockout would require randomly omitting a gene to satisfy the multiplication rules.
Given a two-gene knockout genome matrix
, we create an extended matrix
by appending a zero row as Eq. (2):
![]() |
2 |
Before training, each feature vector is normalized using min–max normalization to ensure that the values across different genes are on a comparable scale.
For each feature vector
, i indexes scalar entries of the input, we apply the min–max normalization as Eq. (3):
![]() |
3 |
If max(
) = min(
), then:
![]() |
4 |
The genome matrix
is compressed to
by the encoder of AE1, which then is compressed to
by the encoder of AE2.
Dropout regularization was applied to both autoencoders (0.35 in AE1; 0.2 and 0.3 in AE2). This autoencoder pipeline had a total of ~ 32.9 million parameters.
Gene knockout simulation by FBA
We first generated all three-gene combinations from the single gene set of the GEM of E. coli K-12 MG1655 (BiGG: iML1515), which has 1515 genes (manually removed s0001). For each three gene combination, the knockout simulation was performed using CobraPy58 with the default target function by maximizing its biomass. We calculated whether the cell growth is less than 10% of the initial growth by Eq. (5):
![]() |
5 |
After performing the knockout for all combinations, we proceed with data cleaning by removing infeasible knockout combinations and excluding those that involve the gene b2092 and gene b4104, which are pseudogenes based on the NCBI record (accessed by Mar 09, 2022). The theoretical number of possible three-gene combinations is:
![]() |
6 |
After data cleaning and merging, we obtain 576,107,619 valid combinations, with 417 missing due to infeasible FBA results. Among all valid combinations, 380,596,069 were labeled as 0, and 195,511,550 were labeled as 1. The ratio of positive to negative labels is 1:1.95. The merged data (12.1 GB) was randomly shuffled and then split into 10 parts for gene combination-level training and output as the final released model for users. The merged triple-knockout dataset was used to construct leakage-aware evaluation splits. For gene-level disjoint evaluation, we partitioned genes into folds and formed Train/Val/Test gene sets that are mutually exclusive (Gtrain ∩ Gval = ∅, Gtrain ∩ Gtest = ∅, Gval ∩ Gtest = ∅). Triples were included in a split only if all three genes belonged to the corresponding gene set. Due to the resulting reduction in usable triples, we sampled up to 3000 k triples for the primary experiments.
The complete simulation of all three-gene knockout combinations was conducted over a period of three months, utilizing 100 parallel computational threads.
MLP classifier
For each three-gene combination {geneX, geneY, geneZ}, the two parts of the features are constructed and concatenated as input for the classifier.
The first part is three-gene 2-mer frequency features. Each gene in
has a 2-mer frequency vector of size
, which are concatenated into a single feature vector of size
:
![]() |
7 |
Then, z-score normalization is applied:
![]() |
8 |
where
and
denote the mean and standard deviation of the vector, and
= 10–8 prevents division by zero.
The second part is remaining genome features via dual autoencoders. The remaining genome is defined as:
![]() |
9 |
Each gene in
is represented by a 2-mer frequency vector in
, forming a matrix
. To maintain the consistent shape, two zero vectors are added:
![]() |
10 |
This matrix is passed through a tandem autoencoder pipeline, reducing the dimensionality to:
![]() |
11 |
The compressed matrix Z is flattened and normalized:
![]() |
12 |
The final input to the MLP is:
![]() |
13 |
The input
is passed to an MLP classifier with sigmoid output, predicting the probability of a lethal 3-gene knockout. The model is trained with Binary Cross Entropy Loss (BCELoss), defined as:
![]() |
14 |
where:
y is the true label.
is the predicted probability.
The MLP was trained with Adam (lr = 0.001) using early stopping based on the designated validation split in the leakage-aware protocol. We report results under the strict gene-disjoint Train/Val/Test setting as the primary evaluation (Table 1), and provide additional split variants and per-fold details in the Supplementary.
Baseline definition (essentiality rule)
We implemented a simple essential-gene rule baseline: a triple is predicted lethal if any member is single-gene lethal/essential (based on single-gene FBA in E. coli). To quantify performance differences, we constructed two evaluation sets (Table 4). Dataset 1 (Mixed data) includes (i) lethal triples without any single-gene lethal member (“new lethal”), (ii) lethal triples containing at least one single-gene lethal gene, and (iii) matched non-lethal triples, yielding a balanced dataset. Dataset 2 (No-essential data) contains only triples with no single-gene lethal genes and balanced lethal/non-lethal labels. For both datasets, we ran random fivefold cross-validation and applied the baseline rule and Tripleknock to the same test folds, reporting mean ± SD test AUC. This random-CV setup is used only for an in-distribution baseline diagnosis; leakage-aware evaluation is reported under the gene-disjoint protocol (Table 1).
Literature-based experimental validation
We searched published multi-gene knockout studies and extracted a small set of relevant experimental phenotypes. Specifically, we curated 37 E. coli triple-gene perturbations from the literature and manually assigned labels based on the source text: triples reported as successfully constructed/growing were labeled 0 (viable), whereas statements such as “unable to obtain/construct” or equivalent nonviability descriptions were labeled 1 (lethal). Because some of these phrases are not direct biochemical growth measurements, we treat this curated labeling as a pragmatic but noisy proxy for non-viability. We then evaluated Tripleknock on this experimentally grounded set. We released it on our GitHub repository with the original source-text criteria used for labeling to ensure transparency and reuse.
Model interpretability
1) Descriptive statistics of class separability
To characterize distributional differences in the input feature space between lethal and non-lethal KOs, we performed a univariate separability analysis on the normalized input features. These statistics describe class separability in the feature space and do not quantify feature importance for the nonlinear MLP.
Let
be the feature matrix of the test set and
the labels. We denote the subsets
and
as the amples with
and
, respectively. For each feature
, we computed the mean difference:
![]() |
15 |
where
and
are the feature-wise means within the positive and negative groups. Statistical significance of the distributional difference was assessed using the two-sided Welch’s t-test:
![]() |
16 |
For visualization, we used the transformed score:
![]() |
17 |
Volcano plots were generated with
and
.
2) Model-based attribution analysis
To interpret the trained MLP in a manner consistent with nonlinear feature interactions, we applied model-based attribution methods. Specifically, we used Integrated Gradients (IG) to quantify feature contributions to the model output, and we validated the conclusions using grouped permutation importance (AUC drop).
First, for Integrated Gradients part, we let
denote the model output probability, and let
denote the corresponding logit:
![]() |
18 |
Because sigmoid outputs can saturate near 0 or 1, IG was computed with respect to the logit
to obtain more stable gradients.
Given an input
and a baseline
(we use the feature-wise mean baseline), the IG attribution for feature
is:
![]() |
19 |
which we approximate with m steps:
![]() |
20 |
To obtain a global importance score per feature, we aggregated absolute attributions over a subset of test samples:
![]() |
21 |
Because the last 2400-dimensional input vector consists of four predefined blocks—KO1 (1–400), KO2 (401–800), KO3 (801–1200), and AE (1201–2400)—we further summarized attribution at the block level by summing
within each block and normalizing to proportions.
Second, to provide a model-agnostic validation of the attribution results, we computed grouped permutation importance at the same block granularity. For each block B, we permuted the corresponding features across samples (breaking the association between that block and the label while preserving marginal distributions), recomputed predictions, and evaluated the AUC. The importance of block B is defined as:
![]() |
22 |
where
denotes the feature matrix after permuting only the columns in block B. Larger
indicates stronger reliance of the model on that block.
All interpretability analyses were performed on the same normalized input representation used during model training.
Supplementary Information
Acknowledgements
Part of the analysis was performed on the High-Performance Computing Platform of Peking University, and Biomedical Computing Platform of National Biomedical Imaging Center of Peking University.
Author contributions
G.P.X. and Z.H. led the research. G.P.X. collected the data, developed the network architecture and training, and conducted the interpretability analysis. H.J., G.J., and J.X. contributed numerous ideas throughout the project.
Funding
This work was supported by the National Science and Technology Major Project (2025ZD01901200) and the National Natural Science Foundation of China (32570752, 32070667).
Data availability
The original code for data collection, which includes three gene knockout FBA simulations, GEM models, and model training using an autoencoder and MLP classifier, is stored on GitHub at the following link: https://github.com/Peneapple/Tripleknock. On the main page of GitHub, we also offer a tutorial for new users unfamiliar with coding to run Tripleknock on Google Colab.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Alper, H., Jin, Y.-S., Moxley, J. F. & Stephanopoulos, G. Identifying gene targets for the metabolic engineering of lycopene biosynthesis in Escherichia coli. Metab. Eng.7, 155–164 (2005). [DOI] [PubMed] [Google Scholar]
- 2.Tamae, C. et al. Determination of antibiotic hypersensitivity among 4,000 single-gene-knockout mutants of Escherichia coli. J. Bacteriol.190, 5981–5988 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Gerdes, S. Y. et al. Experimental determination and system level analysis of essential genes in Escherichia coli MG1655. J. Bacteriol.185, 5673–5684 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Jordan, I. K., Rogozin, I. B., Wolf, Y. I. & Koonin, E. V. Essential genes are more evolutionarily conserved than are nonessential genes in bacteria. Genome Res.12, 962–968 (2002). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Li, X., Li, W., Zeng, M., Zheng, R. & Li, M. Network-based methods for predicting essential genes or proteins: a survey. Brief. Bioinform.21, 566–583 (2020). [DOI] [PubMed] [Google Scholar]
- 6.Baba, T. et al. Construction of Escherichia coli K‐12 in‐frame, single‐gene knockout mutants: the Keio collection. Mol. Syst. Biol.2, 2006.0008 (2006). [DOI] [PMC free article] [PubMed]
- 7.Peters, J. M. et al. A comprehensive, CRISPR-based functional analysis of essential genes in bacteria. Cell165, 1493–1506 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Campos, T. L., Korhonen, P. K. & Young, N. D. Cross-predicting essential genes between two model eukaryotic species using machine learning. Int. J. Mol. Sci.10.3390/ijms22105056 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Marques de Castro, G., Hastenreiter, Z., Silva Monteiro, T. A., Martins da Silva, T. T. & Pereira Lobo, F. Cross-species prediction of essential genes in insects. Bioinformatics38, 1504–1513 (2022). [DOI] [PubMed] [Google Scholar]
- 10.Shen, J. P. & Ideker, T. Synthetic lethal networks for precision oncology: promises and pitfalls. J. Mol. Biol.430, 2900–2912 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Sahu, A. D. et al. Genome-wide prediction of synthetic rescue mediators of resistance to targeted and immunotherapy. Mol. Syst. Biol.15, e8323 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Bunting, S. F. et al. 53BP1 inhibits homologous recombination in Brca1-deficient cells by blocking resection of DNA breaks. Cell141, 243–254 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Jaspers, J. E. et al. Loss of 53BP1 causes PARP inhibitor resistance in Brca1-mutated mouse mammary tumors. Cancer Discov.3, 68–81 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Dev, H. et al. Shieldin complex promotes DNA end-joining and counters homologous recombination in BRCA1-null cells. Nat. Cell Biol.20, 954–965 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Motter, A. E., Gulbahce, N., Almaas, E. & Barabási, A.-L. Predicting synthetic rescues in metabolic networks. Mol. Syst. Biol.4, 168 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Szamecz, B. et al. The genomic landscape of compensatory evolution. PLoS Biol.12, e1001935 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Wang, S. et al. KG4SL: Knowledge graph neural network for synthetic lethality prediction in human cancers. Bioinformatics37, i418–i425 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Biran, N. et al. Triple knockout of frdC gltA and pta genes enhanced PHA production in Escherichia coli. Asia-Pac. J. Mol. Biol. Biotechnol.26, 11–18 (2018). [Google Scholar]
- 19.Orth, J. D., Thiele, I. & Palsson, B. Ø. What is flux balance analysis?. Nat. Biotechnol.28, 245–248 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Gu, C., Kim, G. B., Kim, W. J., Kim, H. U. & Lee, S. Y. Current status and applications of genome-scale metabolic models. Genome Biol.20, 121 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Edwards, J. S. & Palsson, B. Ø. The Escherichia coli MG1655 in silico metabolic genotype: its definition, characteristics, and capabilities. Proc. Natl. Acad. Sci. U. S. A.97, 5528–5533 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Feist, A. M. et al. A genome-scale metabolic reconstruction for Escherichia coli K-12 MG1655 that accounts for 1260 ORFs and thermodynamic information. Mol. Syst. Biol.3, 121 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Monk, J. M. et al. iML1515, a knowledgebase that computes Escherichia coli traits. Nat. Biotechnol.35, 904–908 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Bordbar, A., Monk, J. M., King, Z. A. & Palsson, B. O. Constraint-based models predict metabolic and associated cellular functions. Nat. Rev. Genet.15, 107–120 (2014). [DOI] [PubMed] [Google Scholar]
- 25.Sahu, A., Blätke, M.-A., Szymański, J. J. & Töpfer, N. Advances in flux balance analysis by integrating machine learning and mechanism-based models. Comput. Struct. Biotechnol. J.19, 4626–4640 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Passi, A. et al. Genome-scale metabolic modeling enables in-depth understanding of big data. Metabolites12, 14 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Zhou, J. & Troyanskaya, O. G. Predicting effects of noncoding variants with deep learning-based sequence model. Nat. Methods12, 931–934 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Guzmán, G. I. et al. Model-driven discovery of underground metabolic functions in Escherichia coli. Proc. Natl. Acad. Sci. U. S. A.112, 929–934 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Borja, G. M. et al. Engineering Escherichia coli to increase plasmid DNA production in high cell-density cultivations in batch mode. Microb. Cell Factories11, 132 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Vorachek-Warren, M. K., Ramirez, S., Cotter, R. J. & Raetz, C. R. H. A triple mutant of Escherichia coli lacking secondary acyl chains on lipid A. J. Biol. Chem.277, 14194–14205 (2002). [DOI] [PubMed] [Google Scholar]
- 31.Samuel Raj, V. et al. Decrease in cell viability in an RMF, sigma38, and OmpC triple mutant of Escherichia coli. Biochem. Biophys. Res. Commun.299, 252–257 (2002). [DOI] [PubMed]
- 32.Hinds, T. & Sandler, S. J. Allele specific synthetic lethality between priC and dnaAts alleles at the permissive temperature of 30°C in E. coli K-12. BMC Microbiol.4, 47 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Durand, A., Sinha, A. K., Dard-Dascot, C. & Michel, B. Mutations affecting potassium import restore the viability of the Escherichia coli DNA polymerase III holD mutant. PLoS Genet.12, e1006114 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Dong, W.-R. et al. New function for Escherichia coli xanthosine phophorylase (xapA): Genetic and biochemical evidences on its participation in NAD+ salvage from nicotinamide. BMC Microbiol.14, 29 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Sanders, A. N. & Pavelka, M. S. Phenotypic analysis of Eschericia coli mutants lacking L,D-transpeptidases. Microbiol. Read. Engl.159, 1842–1852 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Ye, C. et al. Metabolic engineering of Escherichia coli BW25113 for the production of 5-aminolevulinic acid based on CRISPR/Cas9 mediated gene knockout and metabolic pathway modification. J. Biol. Eng.16, 26 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Jing, Z. et al. Multiplex gene knockout raises Ala-Gln production by Escherichia coli expressing amino acid ester acyltransferase. Appl. Microbiol. Biotechnol.107, 3523–3533 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Harrison, L., Brame, K. L., Geltz, L. E. & Landry, A. M. Closely opposed apurinic/apyrimidinic sites are converted to double strand breaks in Escherichia coli even in the absence of exonuclease III, endonuclease IV, nucleotide excision repair and AP lyase cleavage. DNA Repair5, 324–335 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Janion, C., Sikora, A., Nowosielska, A. & Grzesiuk, E. E. coli BW535, a triple mutant for the DNA repair genes xth, nth, and nfo, chronically induces the SOS response. Environ. Mol. Mutagen.41, 237–242 (2003). [DOI] [PubMed] [Google Scholar]
- 40.Edwards, D. H. & Donachie, W. D. Construction of a triple deletion of penicillin-binding proteins 4, 5, and 6 in Escherichia coli. In Bacterial Growth and Lysis: Metabolism and Structure of the Bacterial Sacculus (eds de Pedro, M. A. et al.) 369–374 (Springer US, 1993). 10.1007/978-1-4757-9359-8_44. [Google Scholar]
- 41.Ullers, R. S., Ang, D., Schwager, F., Georgopoulos, C. & Genevaux, P. Trigger factor can antagonize both SecB and DnaK/DnaJ chaperone functions in Escherichia coli. Proc. Natl. Acad. Sci. U. S. A.104, 3101–3106 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Wen, Z., Wang, P., Sun, C., Guo, Y. & Wang, X. Interaction of type IV toxin/antitoxin systems in cryptic prophages of Escherichia coli K-12. Toxins9, 77 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Chamberlain, M. et al. The differential effects of anesthetics on bacterial behaviors. PLoS ONE12, e0170089 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Csörgő, B., Fehér, T., Tímár, E., Blattner, F. R. & Pósfai, G. Low-mutation-rate, reduced-genome Escherichia coli: An improved host for faithful maintenance of engineered genetic constructs. Microb. Cell Fact.11, 11 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Rao, T. V. P. & Kuzminov, A. Robust linear DNA degradation supports replication–initiation-defective mutants in Escherichia coli. G3 Genes Genomes Genetics12, jkac228 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Nicolaes, V. et al. Insights into the function of YciM, a heat shock membrane protein required to maintain envelope integrity in Escherichia coli. J. Bacteriol.196, 300–309 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Ting, H., Kouzminova, E. A. & Kuzminov, A. Synthetic lethality with the dut defect in Escherichia coli reveals layers of DNA damage of increasing complexity due to uracil incorporation. J. Bacteriol.190, 5841–5854 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Stasi, R., Neves, H. I. & Spira, B. Phosphate uptake by the phosphonate transport system PhnCDE. BMC Microbiol.19, 79 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Lestini, R. & Michel, B. UvrD controls the access of recombination proteins to blocked replication forks. EMBO J.26, 3804–3814 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Rigel, N. W., Schwalm, J., Ricci, D. P. & Silhavy, T. J. BamE modulates the Escherichia coli beta-barrel assembly machine component BamA. J. Bacteriol.194, 1002–1008 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Jain, K., Wood, E. A. & Cox, M. M. The rarA gene as part of an expanded RecFOR recombination pathway: negative epistasis and synthetic lethality with ruvB, recG, and recQ. PLoS Genet.17, e1009972 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Kouzminova, E. A., Kadyrov, F. F. & Kuzminov, A. RNase HII saves rnhA mutant Escherichia coli from R-loop-associated chromosomal fragmentation. J. Mol. Biol.429, 2873–2894 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Douglass, M. V., McLean, A. B. & Trent, M. S. Absence of YhdP, TamB, and YdbH leads to defects in glycerophospholipid transport and cell morphology in gram-negative bacteria. PLOS Genet.18, e1010096 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Avsec, Ž et al. Effective gene expression prediction from sequence by integrating long-range interactions. Nat. Methods18, 1196–1203 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Chen, C., Liao, C. & Liu, Y.-Y. Teasing out missing reactions in genome-scale metabolic networks through hypergraph learning. Nat. Commun.14, 2375 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.King, Z. A. et al. BiGG Models: A platform for integrating, standardizing and sharing genome-scale models. Nucleic Acids Res.44, D515–D522 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Olson, R. D. et al. Introducing the bacterial and viral bioinformatics resource center (BV-BRC): a resource combining PATRIC IRD and ViPR. Nucleic Acids Res.51, D678–D689 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Ebrahim, A., Lerman, J. A., Palsson, B. O. & Hyduke, D. R. COBRApy: Constraints-based reconstruction and analysis for Python. BMC Syst. Biol.7, 1–6 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The original code for data collection, which includes three gene knockout FBA simulations, GEM models, and model training using an autoencoder and MLP classifier, is stored on GitHub at the following link: https://github.com/Peneapple/Tripleknock. On the main page of GitHub, we also offer a tutorial for new users unfamiliar with coding to run Tripleknock on Google Colab.

























