Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Apr 10;16:16161. doi: 10.1038/s41598-026-46272-9

Tripleknock: predicting lethal effect of three-gene knockout in bacteria by deep learning

Peter X Geng 1, Jiaheng Hou 1, Jinyuan Guo 1,2, Xiaoqing Jiang 1, Huaiqiu Zhu 1,3,✉
PMCID: PMC13201573  PMID: 41963398

Abstract

Investigating the lethal effect of multi-gene knockout is essential for discovering novel antibiotic targets and metabolic engineering. Unlike single genes or gene pairs, three-gene combinations involve more intricate interactions, making experimental screening time-consuming. Computational methods, particularly Genome-scale metabolic model (GEM)-based Flux balance analysis (FBA), requires constructing new GEMs from experimental data, limiting its use for new species. Moreover, using FBA for three-gene knockout screening could take several years. Therefore, a faster and GEM-independent approach is needed to facilitate genome-wide three-gene knockout screening. Here, we introduce Tripleknock for predicting the lethal effects of three-gene knockouts. Tripleknock was trained using genome-derived protein sequence features from Escherichia coli K-12 MG1655, and three-gene knockout simulations using FBA. The model uses a threshold of 90% reduction in cell growth to define lethal effect as the prediction output. Compared to FBA, Tripleknock yields predictions ~ 20 × faster and achieves an average cross-species F1 score of 0.77 on six Enterobacteriaceae pathogens (FBA-defined labels). To mitigate information leakage, we additionally evaluated Tripleknock under a strict gene-level disjoint Train/Val/Test split (genes left out). We further benchmarked against an essentiality-rule baseline and literature-curated (n = 37) E. coli triple perturbations for external validation; no false positives were observed among predicted lethal cases (FP = 0) in this small curated set. To the best of our knowledge, this provides a reproducible baseline and evaluation framework for bacterial triple-knockout lethality prediction.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-46272-9.

Keywords: Triple-gene knockout, Lethal effect, Bacteria, Antibiotic design, Deep learning

Subject terms: Biotechnology, Computational biology and bioinformatics, Microbiology

Introduction

Gene knockout is a widely-used method to investigate the relationship between genotype and phenotype. In metabolic engineering, gene knockouts can generate bacterial strains with increased production of specific metabolites1, while in antibiotic design, gene-knockout screening can facilitate the discovery of novel drug targets2. Single-gene knockout studies have paved the way for identifying genes important for survival and growth. These genes are critical for maintaining basic cellular functions, such as metabolism, DNA replication, and cellular integrity. Absence of essential genes will disrupt cellular processes, typically halting growth or leading to cell death3–5. The identification of essential genes has been primarily achieved through experimental methods such as gene editing, RNA interference, and CRISPR/Cas96,7. In addition to experimental methods, computational approaches have also been developed to predict essential genes across species based on genomic data8,9.

Compared with single-gene knockout, two-gene knockout can reveal novel genetic interactions that are not evident in single-gene knockout studies. Among these interactions, Synthetic Lethal (SL)10 and Synthetic Rescue (SR)11 are two important types of gene interaction. SL is defined as when the knockout of both genes results in a severe loss of cell growth or even cell death, while the knockout of either gene does not lead to this inactivation, indicating that both genes are functionally interdependent. On the other hand, SR refers to a situation where the inhibition of one gene (either gene A or gene B) alone reduces cell growth, but the simultaneous inhibition of both genes rescues the phenotype, thus allowing the cell to survive. SR is also a potential new mechanism in cancer drug resistance12–14. These interactions provide valuable insights into gene dependencies and redundancies in cellular networks. Algorithms and experimental methods for understanding SL and SR effects are also reported15–17.

Expanding to three-gene knockouts, the complexity of genetic interactions increases even further compared to two-gene interactions. Exploring three-gene knockout combination is important and has shown promise in metabolic engineering and therapeutic interventions. For example, a three-gene knockout has also been reported to enhance polyhydroxyalkanoate (PHA) production in E. coli18. These interactions could result in unexpected phenotypic effects, and understanding them may provide new avenues for therapeutic interventions. For pathogenic species, investigating the lethal effect of multi-gene knockout could facilitate the finding of new antibiotic targets. However, multi-gene knockout remains a significant challenge, for the difficulties of accessing sufficient experimental data. And the lack of experimental data limits the development of computational methods. Fortunately, GEMs coupled with FBA19,20 provide reliable results for simulating cell growth when genes are knocked out. But, although it’s the most popular simulation methods in the past over 20 years, it faces limitations when applied to multi-gene knockout. First, the construction of GEMs for each species is labor-intensive and time-consuming21–24. Even if the new species is closely related to a species that already has GEM, a new GEM must be built de novo, which makes it difficult to generalize the results of gene knockout across related species. Second, GEMs only include approximately one-third of genes in the genome, which means the searching space is largely limited. Recent advancements in deep learning have shown promise in predicting gene knockout effects without relying on GEM construction25–27. Given these constraints, an accurate, GEM-independent and user-friendly prediction model for multi-gene’s lethal effects would be a powerful tool for biologists who are working on metabolic engineering or multi-target antibiotics finding.

In response to these challenges, we developed Tripleknock to predict FBA-defined lethal outcomes of three-gene knockouts from genome-derived protein sequence features. We evaluate the model across related Enterobacteriaceae species and, critically, under leakage-aware gene-disjoint splits where entire genes are left out. We further compare against a simple essentiality-rule baseline and perform literature-based experimental validation to assess real-world consistency.

Results

Overview of Tripleknock’s algorithmic architecture

The Tripleknock workflow consists of three main components (Fig. 1). First, a two-stage pre-trained autoencoder is used for genome compression (Fig. 1a). This tandem autoencoder model is designed to process genomes with triple gene knockouts and sequentially reduces both the dimensionality of gene features and the number of genes, resulting in an approximately 1,000-fold reduction in genome matrix size. Second, phenotype labels are generated using FBA, where all feasible three-gene knockout combinations in the GEM of E. coli K-12 MG1655 are simulated and categorized based on their predicted growth outcomes (Fig. 1b). Third, a classification model based on a multilayer perceptron (MLP) is used to predict the viability of triple knockout genotypes (Fig. 1c). This classifier receives a combined feature consisting of the compressed genome and the triple knockout genes, and processes it through fully connected layers to produce a binary prediction. Both the autoencoder and the MLP models are available for direct download and application for users.

Fig. 1.

Fig. 1

Workflow of Tripleknock. (a) Construction of 2-mer encoded protein feature matrices and genome-level compression via a two-stage autoencoder pipeline. (b) Feature engineering and three-gene knockout simulation by FBA. (c) Predictions using concatenated features of knocked-out genes and genomes, done by MLP-based binary classification of cell growth. FC, fully connected layer. Yellow boxes refer to feature engineering, green boxes indicates the layers in three sub architectures, blue boxes indicate the labeling method, and the purple boxes are the final binary output of Tripleknock.

We plotted the loss function of the three sub-models during training (Supplementary Fig. 2). Autoencoder 1 (AE1) demonstrates rapid convergence. A horizontal dashed green line indicates the theoretical background noise level (1/400 = 0.0025), which reflects the average frequency of a 2-mer assuming uniform distribution (Tables 1, 2). The model’s reconstruction error is significantly below this threshold, indicating that the autoencoder captures non-random structural patterns in the input data, which may reflect biologically relevant features (Supplementary Fig. 2a). AE2 also achieves low reconstruction error, indicating that it captures non-random structure in the genome-level input representation (Supplementary Fig. 2b).

Table 1.

Gene-level disjoint (genes left out): train/val/test three-way disjoint (stricter).

Training size (N) Test AUC (mean ± sd) Val AUC (mean ± sd)
1 k 0.589 ± 0.053 0.603 ± 0.052
5 k 0.616 ± 0.047 0.599 ± 0.045
10 k 0.607 ± 0.021 0.607 ± 0.024
20 k 0.597 ± 0.025 0.582 ± 0.030
40 k 0.610 ± 0.036 0.584 ± 0.026
100 k 0.582 ± 0.023 0.586 ± 0.028
1000 k 0.583 ± 0.026 0.583 ± 0.033
2000 k 0.586 ± 0.030 0.582 ± 0.034
3000 k 0.583 ± 0.034 0.579 ± 0.033

Table 2.

Stratified performance on the 3000 k gene-level disjoint test set by confidence threshold (t) and essential-gene count (n_essential).

t n_essential N_total Coverage (mean ± sd) ACCURACY (mean ± sd)
0.1 0 1,921,859 0.641 ± 0.026 0.559 ± 0.027
0.1 1 855,784 0.285 ± 0.017 0.546 ± 0.051
0.1 2 124,739 0.042 ± 0.007 0.643 ± 0.078
0.1 3 5980 0.002 ± 0.001 0.724 ± 0.094
0.2 0 1,857,364 0.619 ± 0.027 0.561 ± 0.028
0.2 1 825,239 0.275 ± 0.017 0.548 ± 0.053
0.2 2 120,309 0.040 ± 0.007 0.648 ± 0.079
0.2 3 5787 0.002 ± 0.001 0.730 ± 0.094
0.3 0 1,777,888 0.593 ± 0.027 0.563 ± 0.029
0.3 1 787,967 0.263 ± 0.016 0.550 ± 0.055
0.3 2 114,972 0.038 ± 0.006 0.654 ± 0.081
0.3 3 5542 0.002 ± 0.001 0.741 ± 0.094
0.4 0 1,656,588 0.552 ± 0.028 0.566 ± 0.031
0.4 1 730,799 0.244 ± 0.014 0.554 ± 0.058
0.4 2 106,666 0.036 ± 0.006 0.663 ± 0.085
0.4 3 5149 0.002 ± 0.001 0.755 ± 0.091

Bold indicates the highest value within each subgroup.

For MLP training, we estimated that the time required for generating and storing all intermediate feature files during training will be approximately 70 days; considering cross-validation, this will increase to near 1 year. Therefore, only random 5% and 10% subsets were used to train the model (Supplementary Fig. 2c-d) at combination level. In both cases, the training on 5% and 10% subsets yielded similar training curves, justified the use of subsampling method. The time of data preparation (FBA simulation) and training of sub-models within Tripleknock framework is collected (Table 3). The combination-level model is released for users, whereas the strict gene-disjoint setting is used to provide a leakage-aware estimate of generalization to unseen genes.

Table 3.

FBA simulation and training time of models in Tripleknock.

Step Time/day
FBA simulation  ~ 100 (100-threads)
Model training
Autoencoder 1 8.15
Autoencoder 2 19.72
MLP 5%  ~ 5
MLP 10%  ~ 10
Total training time  ~ 30–40

We further considered the risk that the combination-level cross-validation can allow the same genes (and gene pairs) to appear in both training and evaluation splits. This can substantially inflate performance estimates. We adopt a leakage-aware evaluation where entire genes are left out and the gene sets are mutually exclusive across Train/Val/Test. Under this strict protocol (Table 1), validation and test performance remain comparable and stays consistently above random despite the strict unseen-gene setting across training scales; at 3000 k samples, Val AUC is 0.579 ± 0.033 and Test AUC is 0.583 ± 0.034. Details of the more gene pair level disjoint and combination-level setting and their inflated estimates are provided in the Supplementary for completeness.

We further stratified the strict gene-disjoint test set by prediction confidence (|y_score − 0.5|≥ t) and essential-gene count (Table 2). Higher accuracy is observed in small high-confidence strata (e.g., n_essential = 3 at t = 0.4), and we interpret this as a diagnostic of structured errors rather than an overall performance gain.

Baseline comparison on “new” synthetic-lethal triples

We compared Tripleknock with an essentiality-rule baseline that predicts a triple as lethal if any member is single-gene lethal. In exhaustive screening, 380,719,699 triples (~ 67%) contain no single-gene lethal member, yet this subset includes 123,630 lethal triples. By construction, the baseline has near-zero recall on these ‘new’ lethal triples. Because it reduces prediction to single-gene essentiality, this rule cannot capture synthetic-lethal triples composed entirely of non-essential genes, whereas Tripleknock achieves substantially higher AUC (Table 4). Note that Table 4 uses random fivefold CV within E. coli (i.e., a non–gene-disjoint, in-distribution setting) and is intended as a diagnostic comparison to isolate the limitation of the essentiality-rule baseline. Leakage-aware generalization to unseen genes is evaluated separately under the strict gene-disjoint protocol in Table 1.

Table 4.

Comparison with the essential-gene rule baseline (Random fivefold CV within E. coli, combination-level, same splits for both methods). *Dataset 1: Mixed no-essential and with-essential data. *Dataset 2: No-essential gene, ‘new triples’.

Dataset setting N_total Tripleknock AUC (mean ± sd) Baseline AUC (mean ± sd)
Dataset 1 494,520 0.998 ± 0.000 0.750 ± 0.001
Dataset 2 247,260 0.999 ± 0.000 0.500 ± 0.000

Literature-based experimental validation

Because no public database systematically reports bacterial triple-lethal knockouts, we curated 37 E. coli triple-gene perturbations from published studies. We assigned labels based on the source text (viable vs. not obtainable/nonviable) to determine whether each perturbation was lethal. These source texts, along with the triple-gene data and references, have been uploaded to GitHub. On this literature-based set, Tripleknock achieves TN = 23, FP = 0, FN = 5, TP = 9 (Accuracy = 0.865; ROC-AUC = 0.821). Several false negatives may reflect label uncertainty and protocol mismatch: ‘unable to obtain/construct’ can arise from technical constraints rather than measured lethality, and experimental growth criteria may not align with our FBA-based 90% growth cutoff. This indicates a conservative behavior in this small set, with reliable lethal calls (FP = 0). Differences in strain background, media, and phenotype definitions mean this validation is necessarily approximate, and we treat it as proof-of-concept external consistency (Table 5; confusion matrix Fig. 2).

Table 5.

Performance of Tripleknock on published experimental data.

Gene A b_number A Gene B b_number B Gene C b_number C Real label Triple knock Species References
argD b3359 astC b1748 gabT b2662 0 0 E. coli K-12 28
recA b2699 deoR b0840 nupG b2964 0 0 E. coli VH33 29
lpxL b1054 lpxM b1855 lpxP b2378 0 0 E. coli 30
rpoS b2741 rmf b0953 ompC b2215 1 1 E. coli K-12 31
dnaA b3702 priC b0467 rnhA b0214 0 0 E. coli K-12 32
trk b3849 kup b3747 kdp b0698 0 0 E. coli 33
trkA b3290 argE b3957 holD b4372 0 0 E. coli 33
holD b4372 dinB b0231 polB b0060 0 0 E. coli 33
pncA b1768 xapA b2407 nadC b0109 0 0 E. coli BW25113 34
ybiS b0819 erfK b1990 ycfS b1113 0 0 E. coli 35
hemF b2436 ybdK b0581 gdhA b1761 0 0 E. coli BW25113 36
hemF b2436 ybdK b0581 gadB b1493 0 0 E. coli BW25113 36
hemF b2436 gdhA b1761 gadB b1493 0 0 E. coli BW25113 36
pepA b4260 pepD b0237 pepN b0932 0 0 E. coli BL21 37
uvrA b4058 xth b1749 nfo b2159 0 0 E. coli 38
xth b1749 nfi b3998 nfo b2159 0 0 E. coli 38
nth b1633 nfo b2159 xth b1749 1 1 E. coli BW535 39
dacA b0632 dacB b3182 dacC b0839 0 0 E. coli 40
secB b3609 tig b0436 dnaJ b0015 0 0 E. coli 41
ykfI b0245 ypjF b2646 cbtA b2005 0 0 E. coli K-12 42
yafW b0246 cbeA b2004 yfjZ b2645 0 0 E. coli K-12 42
clcB b1592 nhaA b0019 nhaB b1186 0 0 E. coli 43
polB b0060 dinB b0231 umuD b1183 0 0 E. coli 44
polB b0060 dinB b0231 umuC b1184 0 0 E. coli 44
dnaA b3702 recB b2820 rep b3778 1 1 E. coli 45
rep b3778 dnaA b3702 recC b2822 1 0 E. coli 45
amiA b2435 yciM b1280 amiB b4169 1 1 E. coli 46
dut b3640 degP b0161 pstC b3727 1 0 E. coli 47
pitA b3493 ugp b3452 pst b3726 1 0 E. coli 48
uvrD b3813 ruvA b1861 recQ b3822 1 1 E. coli 49
uvrD b3813 ruvB b1860 recQ b3822 1 1 E. coli 49
bamE b2617 skp b0178 degP b0161 1 0 E. coli 50
bamC b2477 skp b0178 degP b0161 0 0 E. coli 50
rarA b0892 recG b3652 ruvB b1860 1 1 E. coli 51
rarA b0892 ruvB b1860 recQ b3822 1 1 E. coli 51
rnhA b0214 rnhB b0183 recA b2699 1 1 E. coli 52
yhdP b4472 tamB b4221 ydbH b1381 1 0 E. coli 53

Fig. 2.

Fig. 2

Confusion matrix of Tripleknock on experimental data.

Cross-species prediction

We evaluated the cross-species generalizability of the Tripleknock framework by applying it to six pathogenic bacterial species within the Enterobacteriaceae family. For each species, we generated three independent batches of 20,000 unique triple-gene knockouts (60,000 total) and simulated growth for each batch using FBA. The 10% training version of Tripleknock was then used to predict lethal effect based on the amino acid sequences of six pathogenic bacterial species. The resulting average F1 scores for each species were mapped onto a phylogenetic tree (Fig. 3a), demonstrating that Tripleknock achieved an overall cross-species average F1 score of 0.77, with particularly strong performance observed in E. coli and Shigella, where F1 scores exceeded 0.83. To investigate the impact of training data size on predictive performance, we compared classification metrics (Accuracy, Precision, Recall, and F1-score) across models trained with 5% and 10% of the data (Fig. 3b). The differences were not statistically significant according to the Wilcoxon signed-rank test, suggesting that, for this cross-species FBA-simulation benchmark, using 10% of the E. coli training triples provides a practical trade-off between training cost and performance. Because cross-species evaluation uses genomes from other species, all test genes are unseen relative to the E. coli training genome; therefore, this benchmark complements (but does not replace) the leakage-aware unseen-gene evaluation in Table 1. In practice, many genes in these related Enterobacteriaceae species have close homologs to E. coli genes, so this cross-species benchmark reflects near-homology transfer and is typically less adversarial than the strict gene-disjoint split within E. coli. Accordingly, all subsequent interpretability analyses were conducted using models trained on 10% of the data, and the version of Tripleknock released for public use is likewise based on this setting. In terms of computational efficiency, Tripleknock delivers predictions consistent with FBA outputs approximately 20 to 30 times faster (depending on hardware) when executed on an L4 GPU within the Google Colab environment. Testing on Colab also suggested the usage and capability of Tripleknock for users.

Fig. 3.

Fig. 3

Cross-species prediction. (a) Phylogenetic tree of the seven selected species within the Enterobacteriaceae family. Branch lengths are indicated above the grey lines, and confidence values are shown below. Genome IDs are displayed to the right of each node. Average F1 scores for each species are annotated in red. The training data usage is 10%. (b) Performance of Tripleknock trained on 5% and 10% of E. coli K-12 data and evaluated on six pathogenic species. Error bars represent the standard error of the mean. The average values are labeled above each bar. Statistical significance between the two training ratios for each species was assessed using the Wilcoxon signed-rank test.

Model interpretation

We evaluated interpretability on three representative test species based on their prediction performance: the species with the highest F1 score, the one with the lowest F1 score, and a randomly selected Shigella species. We present volcano plots (Fig. 4a–c) as descriptive class separability in the normalized input space (mean difference and Welch’s t-test), emphasizing that these statistics do not quantify feature importance for a nonlinear MLP. To obtain model-consistent attributions, we computed Integrated Gradients with respect to the model logit and summarized attributions at the predefined block level (KO1/KO2/KO3/AE), and we corroborated these findings using grouped permutation importance measured by AUC drop (Fig. 4d–i).

Fig. 4.

Fig. 4

Data separability and model-based attribution across three test species. (a-c) Volcano plots showing feature separability in the normalized input space (descriptive statistics, not MLP feature importance). The x-axis represents the mean difference between the two groups, the y-axis shows the -log10-transformed p-value from a Welch’s t-test. The horizontal blue dashlines indicate the p-value 0.01. (d-f) Integrated Gradients (IG) attribution aggregated at the feature-block level, computed with respect to the model logit and summarized as mean absolute attribution (normalized to proportions) for KO1 (1–400), KO2 (401–800), KO3 (801–1200), and AE (1201–2400). (g-i) Grouped permutation importance measured as AUC drop after permuting each feature block across samples; larger drops indicate stronger reliance on that block.

Discussion

Three-gene knockout screens for lethal phenotypes are valuable for metabolic engineering and are increasingly important in synthetic-lethal–based disease therapy. However, there remain limited end-to-end approaches for bacterial three-gene knockout outcome prediction. Two major limitations of GEM-based FBA are the time required to screen combinatorial knockouts and the strong dependence on a well-curated, organism-specific GEM. To address these constraints and the growing interest in multi-gene phenotype prediction, we used lethality as a representative phenotype and performed a comprehensive GEM-based in silico screen of all possible three-gene knockouts in E. coli K-12, collecting post-knockout growth outcomes as training data. We developed Tripleknock to predict FBA-defined lethal outcomes of three-gene knockouts, and we further assessed external consistency on 37 literature-curated experimental perturbations and across six pathogenic species.

Our interpretability analyses provide complementary views of the data structure and the model’s decision process. The volcano plots (descriptive separability) show that, as cross-species performance decreases, feature-wise differences between positive and negative classes shrink and more points collapse toward the central axis (Δμ ≈ 0), consistent with weakened class separability and increased distribution shift in the target species. In contrast, model-based attribution (Integrated Gradients on the logit and grouped permutation importance) indicates that, in the current training setting, predictions are driven primarily by features derived from the three knocked-out genes, while the genome-context (AE-derived) block contributes marginally on these test species. This does not imply that genome-wide context is biologically irrelevant; rather, it suggests that the present embedding and training objective may not yet extract transferable contextual signals beyond the KO genes. These findings suggest a future direction that a deeper understanding of genome-wide features is essential to improve the model’s transferability across species54.

We aim to produce a user-friendly tool for experimental biologists who may not familiar with coding, therefore, during the exploration of features integration, we weighed two choices: (1) Whether to incorporate the gene–reaction–metabolite graph information from GEMs. We didn’t because manage to make predictions without depending on GEM. The use of GEM presents a challenge, as constructing a high-quality GEM model for a new species is not straightforward. (2) Whether to use both DNA and protein sequences for feature extraction. In small-scale tests, adding DNA information did not provide a clear improvement. Also, because whole-genome sequencing mainly uses CDS, DNA and protein sequences contain substantial overlapping and redundant information. Therefore, we chose to use protein-sequence 2-mer features as the raw input, and after integration and normalization, we fed them into the neural network. The above considerations make sure Tripleknock to function end-to-end and is friendly for biologists who are not familiar with coding, solely based on an annotated genome file, without relying on external or supplementary materials. Some other questions will also be considered. For instance, this study employs a binary classification approach using solely sequence data, which may be inadequate for more complex predictive tasks. Moreover, our prediction does not incorporate experimental metabolic data, such as multi-omics datasets. It is estimated that integrating these diverse data sources could enhance the accuracy of cross-species predictions. Under the strict gene-disjoint Train/Val/Test protocol, generalization to entirely unseen genes remains modest (AUC ~ 0.58), reflecting the inherent challenges of OOD prediction from sequence-only features and motivating future incorporation of richer functional/contextual information. Sequence similarity represents just one of many conserved features across different species; other factors, such as metabolic network-based information55, also contribute significantly to cross-species prediction. Therefore, in future work, we will also incorporate metabolic network structure—beyond sequence conservation—to improve prediction accuracy and cross-species generalization.

In summary, to the best of our knowledge, we developed an end-to-end deep learning framework Tripleknock, which provides a reproducible framework for triple-knockout lethality prediction and benchmarking. While we acknowledge the performance boundaries under strict unseen-gene evaluation, Tripleknock remains an effective and user-friendly tool for biologists who are working on three-target antibiotics finding, triple-gene lethal screening and related metabolic engineering areas.

Methods

Computational environment

The FBA simulation and model training were performed on a server running an Ubuntu 18.04.2 LTS system. The hardware included 4 Intel(R) Xeon(R) Gold 6252 CPUs operating at 2.10 GHz, and 2 NVIDIA A100-PCIE-40 GB GPUs. The software environment was based on Python 3.8.19, CUDA Version 11.7. Cross-species evaluations were conducted on Google Colab using an L-4 GPU. Model interpretability was conducted on Google Colab using a CPU. No software version-dependent errors were encountered, demonstrating the robustness and portability of Tripleknock.

Datasets

E. coli K-12 MG1655, a non-pathogenic strain, was used for training. For testing, we selected one E. coli strain, one Salmonella enterica, three Shigella species, and one Klebsiella, all of which are pathogenic. These seven species belong to the Enterobacteriaceae family (Table 6). The GEMs were downloaded from the Biochemical, Genetic and Genomic knowledgebase (BiGG), a database of high-quality GEMs56. The protein Coding DNA Sequences (CDS) for each species were downloaded from the National Center for Biotechnology Information (NCBI) in FASTA protein format. A horizontal bar plot shows the number of genes annotated in the genome (CDS genes) versus those included in the GEMs across seven species (Supplementary Fig. 1). All of the distributions of protein length of the seven species are right-skewed, indicating that most proteins are relatively short, typically under 600 amino acids. A long tail exists, representing a small number of much longer proteins, some exceeding 2000 amino acids.

Table 6.

Selected bacteria in training and test set.

Organism BiGG NCBI BV-BRC Pathogenic
E. coli str. K-12 substr. MG1655 iML1515 NC_000913.3 511,145.12 No
E. coli O157:H7 str. EDL933 iZ_1308 AE005174.2 155,864.43 Yes
Salmonella enterica subsp. enterica serovar Typhimurium str. LT2 STM_v1_0 AE006468.2 99,287.12 Yes
Shigella boydii CDC 3083–94 iSbBS512_1146 CP001063.1 344,609.11 Yes
Shigella dysenteriae Sd197 iSDY_1059 CP000034.1 300,267.13 Yes
Shigella flexneri 2a str. 301 iSF_1195 AE005674.2 198,214.7 Yes
K. pneumoniae subsp. pneumoniae MGH 78,578 iYL1228 CP000647.1 272,620.9 Yes

The ‘Bacterial Genome Tree’ tool in the Bacterial and Viral Bioinformatics Resource Center (BV-BRC)57 was used for building the reference phylogenetic tree for cross-species prediction (Supplementary Table 1). This was used for cross-species prediction, during which we used Tripleknock to receive an annotated genome sequence to make direct three-gene knockout prediction on these related species.

Pretrained tandem autoencoders

The first process generated all two-gene pairs from a total of 4305 genes (Inline graphic), in the E. coli K-12 MG1655 genome. The total number of combinations is 9,264,360. Removing each two-gene combination from Inline graphic will lead to an n-2 gene set, which is represented as a matrix whose shape is 4303 × 400. The AE1 was trained to compress gene feature vectors from 400 dimensions down to 3, and the AE2 reduces the dimensionality of gene number to 400. The training process used all data over 6 epochs, with a batch size of 2000. The Eq. (1) is the loss function used for training is Mean Squared Error (MSE) to minimize the reconstruction error between the input and the output.

graphic file with name d33e2362.gif 1

where:

  • Inline graphic indexes scalar entries of the input

  • N is the number of input gene features in AE1, the number of genes in AE2.

  • Inline graphic is the original input.

  • Inline graphic is the reconstructed output from the autoencoder (the predicted feature vector).

To accommodate Tripleknock for potential single-gene knockout on E. coli MG1655, we explicitly add one zero vector row to the input matrix (making it 4304 × 400). This ensures that, when simulating single-gene knockouts for a genome of 4305 genes, every gene in E. coli K-12 MG1655 is represented and the matrix multiplication aligns with the model’s input dimensions. Without adding zero rows, predicting a single-gene knockout would require randomly omitting a gene to satisfy the multiplication rules.

Given a two-gene knockout genome matrix Inline graphic, we create an extended matrix Inline graphic by appending a zero row as Eq. (2):

graphic file with name d33e2413.gif 2

Before training, each feature vector is normalized using min–max normalization to ensure that the values across different genes are on a comparable scale.

For each feature vector Inline graphic, i indexes scalar entries of the input, we apply the min–max normalization as Eq. (3):

graphic file with name d33e2431.gif 3

If max(Inline graphic) = min(Inline graphic), then:

graphic file with name d33e2445.gif 4

The genome matrix Inline graphic is compressed to Inline graphic by the encoder of AE1, which then is compressed to Inline graphic by the encoder of AE2.

Dropout regularization was applied to both autoencoders (0.35 in AE1; 0.2 and 0.3 in AE2). This autoencoder pipeline had a total of ~ 32.9 million parameters.

Gene knockout simulation by FBA

We first generated all three-gene combinations from the single gene set of the GEM of E. coli K-12 MG1655 (BiGG: iML1515), which has 1515 genes (manually removed s0001). For each three gene combination, the knockout simulation was performed using CobraPy58 with the default target function by maximizing its biomass. We calculated whether the cell growth is less than 10% of the initial growth by Eq. (5):

graphic file with name d33e2479.gif 5

After performing the knockout for all combinations, we proceed with data cleaning by removing infeasible knockout combinations and excluding those that involve the gene b2092 and gene b4104, which are pseudogenes based on the NCBI record (accessed by Mar 09, 2022). The theoretical number of possible three-gene combinations is:

graphic file with name d33e2491.gif 6

After data cleaning and merging, we obtain 576,107,619 valid combinations, with 417 missing due to infeasible FBA results. Among all valid combinations, 380,596,069 were labeled as 0, and 195,511,550 were labeled as 1. The ratio of positive to negative labels is 1:1.95. The merged data (12.1 GB) was randomly shuffled and then split into 10 parts for gene combination-level training and output as the final released model for users. The merged triple-knockout dataset was used to construct leakage-aware evaluation splits. For gene-level disjoint evaluation, we partitioned genes into folds and formed Train/Val/Test gene sets that are mutually exclusive (Gtrain ∩ Gval = ∅, Gtrain ∩ Gtest = ∅, Gval ∩ Gtest = ∅). Triples were included in a split only if all three genes belonged to the corresponding gene set. Due to the resulting reduction in usable triples, we sampled up to 3000 k triples for the primary experiments.

The complete simulation of all three-gene knockout combinations was conducted over a period of three months, utilizing 100 parallel computational threads.

MLP classifier

For each three-gene combination {geneX, geneY, geneZ}, the two parts of the features are constructed and concatenated as input for the classifier.

The first part is three-gene 2-mer frequency features. Each gene in Inline graphic has a 2-mer frequency vector of size Inline graphic, which are concatenated into a single feature vector of size Inline graphic:

graphic file with name d33e2539.gif 7

Then, z-score normalization is applied:

graphic file with name d33e2545.gif 8

where Inline graphic and Inline graphic denote the mean and standard deviation of the vector, and Inline graphic = 10–8 prevents division by zero.

The second part is remaining genome features via dual autoencoders. The remaining genome is defined as:

graphic file with name d33e2566.gif 9

Each gene in Inline graphic is represented by a 2-mer frequency vector in Inline graphic, forming a matrix Inline graphic. To maintain the consistent shape, two zero vectors are added:

graphic file with name d33e2584.gif 10

This matrix is passed through a tandem autoencoder pipeline, reducing the dimensionality to:

graphic file with name d33e2590.gif 11

The compressed matrix Z is flattened and normalized:

graphic file with name d33e2596.gif 12

The final input to the MLP is:

graphic file with name d33e2602.gif 13

The input Inline graphic is passed to an MLP classifier with sigmoid output, predicting the probability of a lethal 3-gene knockout. The model is trained with Binary Cross Entropy Loss (BCELoss), defined as:

graphic file with name d33e2612.gif 14

where:

y is the true label.

Inline graphic is the predicted probability.

The MLP was trained with Adam (lr = 0.001) using early stopping based on the designated validation split in the leakage-aware protocol. We report results under the strict gene-disjoint Train/Val/Test setting as the primary evaluation (Table 1), and provide additional split variants and per-fold details in the Supplementary.

Baseline definition (essentiality rule)

We implemented a simple essential-gene rule baseline: a triple is predicted lethal if any member is single-gene lethal/essential (based on single-gene FBA in E. coli). To quantify performance differences, we constructed two evaluation sets (Table 4). Dataset 1 (Mixed data) includes (i) lethal triples without any single-gene lethal member (“new lethal”), (ii) lethal triples containing at least one single-gene lethal gene, and (iii) matched non-lethal triples, yielding a balanced dataset. Dataset 2 (No-essential data) contains only triples with no single-gene lethal genes and balanced lethal/non-lethal labels. For both datasets, we ran random fivefold cross-validation and applied the baseline rule and Tripleknock to the same test folds, reporting mean ± SD test AUC. This random-CV setup is used only for an in-distribution baseline diagnosis; leakage-aware evaluation is reported under the gene-disjoint protocol (Table 1).

Literature-based experimental validation

We searched published multi-gene knockout studies and extracted a small set of relevant experimental phenotypes. Specifically, we curated 37 E. coli triple-gene perturbations from the literature and manually assigned labels based on the source text: triples reported as successfully constructed/growing were labeled 0 (viable), whereas statements such as “unable to obtain/construct” or equivalent nonviability descriptions were labeled 1 (lethal). Because some of these phrases are not direct biochemical growth measurements, we treat this curated labeling as a pragmatic but noisy proxy for non-viability. We then evaluated Tripleknock on this experimentally grounded set. We released it on our GitHub repository with the original source-text criteria used for labeling to ensure transparency and reuse.

Model interpretability

1) Descriptive statistics of class separability

To characterize distributional differences in the input feature space between lethal and non-lethal KOs, we performed a univariate separability analysis on the normalized input features. These statistics describe class separability in the feature space and do not quantify feature importance for the nonlinear MLP.

Let Inline graphic be the feature matrix of the test set and Inline graphic the labels. We denote the subsets Inline graphic and Inline graphic as the amples with Inline graphic and Inline graphic, respectively. For each feature Inline graphic, we computed the mean difference:

graphic file with name d33e2696.gif 15

where Inline graphic and Inline graphic are the feature-wise means within the positive and negative groups. Statistical significance of the distributional difference was assessed using the two-sided Welch’s t-test:

graphic file with name d33e2709.gif 16

For visualization, we used the transformed score:

graphic file with name d33e2715.gif 17

Volcano plots were generated with Inline graphic and Inline graphic.

2) Model-based attribution analysis

To interpret the trained MLP in a manner consistent with nonlinear feature interactions, we applied model-based attribution methods. Specifically, we used Integrated Gradients (IG) to quantify feature contributions to the model output, and we validated the conclusions using grouped permutation importance (AUC drop).

First, for Integrated Gradients part, we let Inline graphic denote the model output probability, and let Inline graphic denote the corresponding logit:

graphic file with name d33e2745.gif 18

Because sigmoid outputs can saturate near 0 or 1, IG was computed with respect to the logit Inline graphic to obtain more stable gradients.

Given an input Inline graphic and a baseline Inline graphic(we use the feature-wise mean baseline), the IG attribution for feature Inline graphic is:

graphic file with name d33e2769.gif 19

which we approximate with m steps:

graphic file with name d33e2778.gif 20

To obtain a global importance score per feature, we aggregated absolute attributions over a subset of test samples:

graphic file with name d33e2785.gif 21

Because the last 2400-dimensional input vector consists of four predefined blocks—KO1 (1–400), KO2 (401–800), KO3 (801–1200), and AE (1201–2400)—we further summarized attribution at the block level by summing Inline graphic within each block and normalizing to proportions.

Second, to provide a model-agnostic validation of the attribution results, we computed grouped permutation importance at the same block granularity. For each block B, we permuted the corresponding features across samples (breaking the association between that block and the label while preserving marginal distributions), recomputed predictions, and evaluated the AUC. The importance of block B is defined as:

graphic file with name d33e2803.gif 22

where Inline graphic denotes the feature matrix after permuting only the columns in block B. Larger Inline graphic indicates stronger reliance of the model on that block.

All interpretability analyses were performed on the same normalized input representation used during model training.

Supplementary Information

Supplementary Information. (775.6KB, docx)

Acknowledgements

Part of the analysis was performed on the High-Performance Computing Platform of Peking University, and Biomedical Computing Platform of National Biomedical Imaging Center of Peking University.

Author contributions

G.P.X. and Z.H. led the research. G.P.X. collected the data, developed the network architecture and training, and conducted the interpretability analysis. H.J., G.J., and J.X. contributed numerous ideas throughout the project.

Funding

This work was supported by the National Science and Technology Major Project (2025ZD01901200) and the National Natural Science Foundation of China (32570752, 32070667).

Data availability

The original code for data collection, which includes three gene knockout FBA simulations, GEM models, and model training using an autoencoder and MLP classifier, is stored on GitHub at the following link: https://github.com/Peneapple/Tripleknock. On the main page of GitHub, we also offer a tutorial for new users unfamiliar with coding to run Tripleknock on Google Colab.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Alper, H., Jin, Y.-S., Moxley, J. F. & Stephanopoulos, G. Identifying gene targets for the metabolic engineering of lycopene biosynthesis in Escherichia coli. Metab. Eng.7, 155–164 (2005). [DOI] [PubMed] [Google Scholar]
  • 2.Tamae, C. et al. Determination of antibiotic hypersensitivity among 4,000 single-gene-knockout mutants of Escherichia coli. J. Bacteriol.190, 5981–5988 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Gerdes, S. Y. et al. Experimental determination and system level analysis of essential genes in Escherichia coli MG1655. J. Bacteriol.185, 5673–5684 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Jordan, I. K., Rogozin, I. B., Wolf, Y. I. & Koonin, E. V. Essential genes are more evolutionarily conserved than are nonessential genes in bacteria. Genome Res.12, 962–968 (2002). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Li, X., Li, W., Zeng, M., Zheng, R. & Li, M. Network-based methods for predicting essential genes or proteins: a survey. Brief. Bioinform.21, 566–583 (2020). [DOI] [PubMed] [Google Scholar]
  • 6.Baba, T. et al. Construction of Escherichia coli K‐12 in‐frame, single‐gene knockout mutants: the Keio collection. Mol. Syst. Biol.2, 2006.0008 (2006). [DOI] [PMC free article] [PubMed]
  • 7.Peters, J. M. et al. A comprehensive, CRISPR-based functional analysis of essential genes in bacteria. Cell165, 1493–1506 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Campos, T. L., Korhonen, P. K. & Young, N. D. Cross-predicting essential genes between two model eukaryotic species using machine learning. Int. J. Mol. Sci.10.3390/ijms22105056 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Marques de Castro, G., Hastenreiter, Z., Silva Monteiro, T. A., Martins da Silva, T. T. & Pereira Lobo, F. Cross-species prediction of essential genes in insects. Bioinformatics38, 1504–1513 (2022). [DOI] [PubMed] [Google Scholar]
  • 10.Shen, J. P. & Ideker, T. Synthetic lethal networks for precision oncology: promises and pitfalls. J. Mol. Biol.430, 2900–2912 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Sahu, A. D. et al. Genome-wide prediction of synthetic rescue mediators of resistance to targeted and immunotherapy. Mol. Syst. Biol.15, e8323 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Bunting, S. F. et al. 53BP1 inhibits homologous recombination in Brca1-deficient cells by blocking resection of DNA breaks. Cell141, 243–254 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Jaspers, J. E. et al. Loss of 53BP1 causes PARP inhibitor resistance in Brca1-mutated mouse mammary tumors. Cancer Discov.3, 68–81 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Dev, H. et al. Shieldin complex promotes DNA end-joining and counters homologous recombination in BRCA1-null cells. Nat. Cell Biol.20, 954–965 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Motter, A. E., Gulbahce, N., Almaas, E. & Barabási, A.-L. Predicting synthetic rescues in metabolic networks. Mol. Syst. Biol.4, 168 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Szamecz, B. et al. The genomic landscape of compensatory evolution. PLoS Biol.12, e1001935 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wang, S. et al. KG4SL: Knowledge graph neural network for synthetic lethality prediction in human cancers. Bioinformatics37, i418–i425 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Biran, N. et al. Triple knockout of frdC gltA and pta genes enhanced PHA production in Escherichia coli. Asia-Pac. J. Mol. Biol. Biotechnol.26, 11–18 (2018). [Google Scholar]
  • 19.Orth, J. D., Thiele, I. & Palsson, B. Ø. What is flux balance analysis?. Nat. Biotechnol.28, 245–248 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Gu, C., Kim, G. B., Kim, W. J., Kim, H. U. & Lee, S. Y. Current status and applications of genome-scale metabolic models. Genome Biol.20, 121 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Edwards, J. S. & Palsson, B. Ø. The Escherichia coli MG1655 in silico metabolic genotype: its definition, characteristics, and capabilities. Proc. Natl. Acad. Sci. U. S. A.97, 5528–5533 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Feist, A. M. et al. A genome-scale metabolic reconstruction for Escherichia coli K-12 MG1655 that accounts for 1260 ORFs and thermodynamic information. Mol. Syst. Biol.3, 121 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Monk, J. M. et al. iML1515, a knowledgebase that computes Escherichia coli traits. Nat. Biotechnol.35, 904–908 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Bordbar, A., Monk, J. M., King, Z. A. & Palsson, B. O. Constraint-based models predict metabolic and associated cellular functions. Nat. Rev. Genet.15, 107–120 (2014). [DOI] [PubMed] [Google Scholar]
  • 25.Sahu, A., Blätke, M.-A., Szymański, J. J. & Töpfer, N. Advances in flux balance analysis by integrating machine learning and mechanism-based models. Comput. Struct. Biotechnol. J.19, 4626–4640 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Passi, A. et al. Genome-scale metabolic modeling enables in-depth understanding of big data. Metabolites12, 14 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Zhou, J. & Troyanskaya, O. G. Predicting effects of noncoding variants with deep learning-based sequence model. Nat. Methods12, 931–934 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Guzmán, G. I. et al. Model-driven discovery of underground metabolic functions in Escherichia coli. Proc. Natl. Acad. Sci. U. S. A.112, 929–934 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Borja, G. M. et al. Engineering Escherichia coli to increase plasmid DNA production in high cell-density cultivations in batch mode. Microb. Cell Factories11, 132 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Vorachek-Warren, M. K., Ramirez, S., Cotter, R. J. & Raetz, C. R. H. A triple mutant of Escherichia coli lacking secondary acyl chains on lipid A. J. Biol. Chem.277, 14194–14205 (2002). [DOI] [PubMed] [Google Scholar]
  • 31.Samuel Raj, V. et al. Decrease in cell viability in an RMF, sigma38, and OmpC triple mutant of Escherichia coli. Biochem. Biophys. Res. Commun.299, 252–257 (2002). [DOI] [PubMed]
  • 32.Hinds, T. & Sandler, S. J. Allele specific synthetic lethality between priC and dnaAts alleles at the permissive temperature of 30°C in E. coli K-12. BMC Microbiol.4, 47 (2004). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Durand, A., Sinha, A. K., Dard-Dascot, C. & Michel, B. Mutations affecting potassium import restore the viability of the Escherichia coli DNA polymerase III holD mutant. PLoS Genet.12, e1006114 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Dong, W.-R. et al. New function for Escherichia coli xanthosine phophorylase (xapA): Genetic and biochemical evidences on its participation in NAD+ salvage from nicotinamide. BMC Microbiol.14, 29 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Sanders, A. N. & Pavelka, M. S. Phenotypic analysis of Eschericia coli mutants lacking L,D-transpeptidases. Microbiol. Read. Engl.159, 1842–1852 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Ye, C. et al. Metabolic engineering of Escherichia coli BW25113 for the production of 5-aminolevulinic acid based on CRISPR/Cas9 mediated gene knockout and metabolic pathway modification. J. Biol. Eng.16, 26 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Jing, Z. et al. Multiplex gene knockout raises Ala-Gln production by Escherichia coli expressing amino acid ester acyltransferase. Appl. Microbiol. Biotechnol.107, 3523–3533 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Harrison, L., Brame, K. L., Geltz, L. E. & Landry, A. M. Closely opposed apurinic/apyrimidinic sites are converted to double strand breaks in Escherichia coli even in the absence of exonuclease III, endonuclease IV, nucleotide excision repair and AP lyase cleavage. DNA Repair5, 324–335 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Janion, C., Sikora, A., Nowosielska, A. & Grzesiuk, E. E. coli BW535, a triple mutant for the DNA repair genes xth, nth, and nfo, chronically induces the SOS response. Environ. Mol. Mutagen.41, 237–242 (2003). [DOI] [PubMed] [Google Scholar]
  • 40.Edwards, D. H. & Donachie, W. D. Construction of a triple deletion of penicillin-binding proteins 4, 5, and 6 in Escherichia coli. In Bacterial Growth and Lysis: Metabolism and Structure of the Bacterial Sacculus (eds de Pedro, M. A. et al.) 369–374 (Springer US, 1993). 10.1007/978-1-4757-9359-8_44. [Google Scholar]
  • 41.Ullers, R. S., Ang, D., Schwager, F., Georgopoulos, C. & Genevaux, P. Trigger factor can antagonize both SecB and DnaK/DnaJ chaperone functions in Escherichia coli. Proc. Natl. Acad. Sci. U. S. A.104, 3101–3106 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Wen, Z., Wang, P., Sun, C., Guo, Y. & Wang, X. Interaction of type IV toxin/antitoxin systems in cryptic prophages of Escherichia coli K-12. Toxins9, 77 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Chamberlain, M. et al. The differential effects of anesthetics on bacterial behaviors. PLoS ONE12, e0170089 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Csörgő, B., Fehér, T., Tímár, E., Blattner, F. R. & Pósfai, G. Low-mutation-rate, reduced-genome Escherichia coli: An improved host for faithful maintenance of engineered genetic constructs. Microb. Cell Fact.11, 11 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Rao, T. V. P. & Kuzminov, A. Robust linear DNA degradation supports replication–initiation-defective mutants in Escherichia coli. G3 Genes Genomes Genetics12, jkac228 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Nicolaes, V. et al. Insights into the function of YciM, a heat shock membrane protein required to maintain envelope integrity in Escherichia coli. J. Bacteriol.196, 300–309 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Ting, H., Kouzminova, E. A. & Kuzminov, A. Synthetic lethality with the dut defect in Escherichia coli reveals layers of DNA damage of increasing complexity due to uracil incorporation. J. Bacteriol.190, 5841–5854 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Stasi, R., Neves, H. I. & Spira, B. Phosphate uptake by the phosphonate transport system PhnCDE. BMC Microbiol.19, 79 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Lestini, R. & Michel, B. UvrD controls the access of recombination proteins to blocked replication forks. EMBO J.26, 3804–3814 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Rigel, N. W., Schwalm, J., Ricci, D. P. & Silhavy, T. J. BamE modulates the Escherichia coli beta-barrel assembly machine component BamA. J. Bacteriol.194, 1002–1008 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Jain, K., Wood, E. A. & Cox, M. M. The rarA gene as part of an expanded RecFOR recombination pathway: negative epistasis and synthetic lethality with ruvB, recG, and recQ. PLoS Genet.17, e1009972 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Kouzminova, E. A., Kadyrov, F. F. & Kuzminov, A. RNase HII saves rnhA mutant Escherichia coli from R-loop-associated chromosomal fragmentation. J. Mol. Biol.429, 2873–2894 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Douglass, M. V., McLean, A. B. & Trent, M. S. Absence of YhdP, TamB, and YdbH leads to defects in glycerophospholipid transport and cell morphology in gram-negative bacteria. PLOS Genet.18, e1010096 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Avsec, Ž et al. Effective gene expression prediction from sequence by integrating long-range interactions. Nat. Methods18, 1196–1203 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Chen, C., Liao, C. & Liu, Y.-Y. Teasing out missing reactions in genome-scale metabolic networks through hypergraph learning. Nat. Commun.14, 2375 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.King, Z. A. et al. BiGG Models: A platform for integrating, standardizing and sharing genome-scale models. Nucleic Acids Res.44, D515–D522 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Olson, R. D. et al. Introducing the bacterial and viral bioinformatics resource center (BV-BRC): a resource combining PATRIC IRD and ViPR. Nucleic Acids Res.51, D678–D689 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Ebrahim, A., Lerman, J. A., Palsson, B. O. & Hyduke, D. R. COBRApy: Constraints-based reconstruction and analysis for Python. BMC Syst. Biol.7, 1–6 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Information. (775.6KB, docx)

Data Availability Statement

The original code for data collection, which includes three gene knockout FBA simulations, GEM models, and model training using an autoencoder and MLP classifier, is stored on GitHub at the following link: https://github.com/Peneapple/Tripleknock. On the main page of GitHub, we also offer a tutorial for new users unfamiliar with coding to run Tripleknock on Google Colab.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES