Skip to main content
BMC Genomics logoLink to BMC Genomics
. 2026 May 22;27:505. doi: 10.1186/s12864-026-12884-8

The interplay of specialized metabolites, antibiotic resistance, and virulence in Enterococcus faecium: in silico analysis of bacterial genomes

Sarah Sonbol 1, Laila Abdel Samad 2, Farah Al-Marzooq 3, Laila Ziko 4,
PMCID: PMC13198060  PMID: 42174466

Abstract

Background

Enterococcus faecium is a prominent Gram-positive nosocomial pathogen associated with high morbidity and mortality worldwide. Identification of its antimicrobial resistance (AMR) genes, virulence factors, and biosynthetic gene clusters (BGCs) encoding specialized metabolites is essential as a preliminary step toward developing novel therapeutic interventions. This study aimed to explore the relationship between BGCs, AMR and virulence attributes in 81 E. faecium genomes—retrieved from the Genome Taxonomy Database (GTDB)—using different computational tools, including antiSMASH, GECCO and Clinker. We assessed the prevalence, co-occurrence and associations between all identified traits.

Results

Nearly 50% of the 334 total detected BGCs were identified as RiPP-like clusters, with the majority predicted to encode bacteriocin-related products. Certain clusters of cyclic lactone autoinducers (CAL-BGCs) showed genomic rearrangements or unique sequences, suggesting diversity in quorum-sensing potential or signaling systems. Several primary metabolic gene clusters (MGCs) were also identified, including the arginine deiminase system, PFOR-II, and gallic acid metabolism, suggesting roles in energy adaptation and gut colonization. Both BGCs and MGCs show relatively high conserved profiles. The virulence genes: acm, scm, fss3, sgrA and ecbA were present at variable levels, implying differences in host-cell binding abilities and biofilm formation. AMR profiling showed an extremely high resistance profile in most genomes, where 43 AMR genes were detected, such as vancomycin resistance clusters (vanA-type), aac(6')-Ii, and msrC. Concurrent resistance for 5 or more classes of antibiotics was detected in approximately 70% of the studied genomes. Epidemiological integration showed strong sequence type (ST) structuring across geography and clinical source, with lineage-dependent variation in AMR burden but relatively conserved virulence and BGC profiles.

Conclusion

This study highlights the metabolic and genomic plasticity of E. faecium, especially its dynamic AMR profile, which reflects its high pathogenicity and environmental resilience. Our results underscore the importance of the concurrent exploration of specialized metabolites, resistance and virulence traits to gain insights into the pathogen's survival mechanisms and identify potential targets for novel therapeutics. Nevertheless, further experimental validation of these computational results will deepen our understanding of this pathogen and will aid in the battle against AMR bacteria.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12864-026-12884-8.

Keywords: Enterococcus faecium, BGCs, Specialized metabolites, Virulence, AMR genes

Background

Globally, as a result of the escalating increase of multidrug-resistant (MDR) bacteria, the number of healthcare-associated infections has risen tremendously. According to the WHO, ESKAPE pathogens (Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, and Enterobacter sp.) are spearheading the microbial pathogens; hence, they are considered of top priority, and developing new antibiotics is needed urgently [1, 2]. Enterococci are Gram-positive, non-motile, facultative anaerobes commonly found as commensals in the gastrointestinal tract of humans and other animals [24]. Enterococcal infections were mainly reported in immunocompromised patients; however, their frequency experienced a surge in the late 1970 s [4]. This rise was mostly from Enterococcus faecalis due to its intrinsic virulence, the excessive use of broad-spectrum antibiotics and the introduction of more invasive medical devices [4]. By the 1990 s, an escalation of the number of nosocomial infections due to MDR Enterococcus faecium was observed [4]. Currently, vancomycin-resistant Enterococcus faecium (VREfm) has become one of the most significant threats in healthcare settings [1, 4].

While antibiotic resistance is the primary driver of its clinical persistence, the pathogen’s capacity to synthesize specialized metabolites (SMs) likely provides a competitive advantage in host colonization and inter-species interactions [5]. Unlike primary metabolites, SMs are non-essential for the basic survival and energy metabolism of the microbe [6]. Microbial SMs, also known as natural products (NPs), are low molecular weight compounds of high chemical diversity that enable bacteria to adapt to their surroundings, among other functions [5, 7]. Usually, the genes encoding SMs are clustered together as Biosynthetic Gene Clusters (BGCs) [5, 7, 8]. These BGCS are involved in several biochemical pathways, and they code for several SM-producing enzymes, such as polyketide synthases (PKSs), non-ribosomal peptide synthetases (NRPSs), terpene synthases, cyclases, and ribosomally synthesized and post-translationally modified peptides (RiPPs) [9]. Advances in sequencing and the availability of a myriad of microbial genomes across various databases facilitated computational mining for BGCs.

While different studies have documented the antimicrobial resistance (AMR) abilities of E. faecium [1, 4], the co-occurrence and interplay between its biosynthetic potential, virulence factors and primary metabolic adaptation remain poorly understood. Consequently, in this study, we aimed to integrate genomic mining of BGCs with comprehensive AMR and virulence profiling to elucidate the survival mechanisms that underpin this pathogen’s resilience in healthcare settings.

Materials and methods

Genome sequences

The Genome Taxonomy Database (GTDB; https://gtdb.ecogenomic.org/) [10] was used to retrieve the genome sequence of E. faecium in February 2024. We selected assembled complete genomes (CheckM completeness ≥ 99.6), with no contamination (CheckM contamination = 0) and no ambiguous bases. Assembled genomes with an N50 contig length of less than 2,500,000—which covers more than 75% of the expected gene size—were excluded to ensure a high quality of the assembled genomes. A total of 81 genomes were retrieved for further analysis. The accession numbers of the genome assemblies included in this study are provided in Supplementary Table S1 in the first column, and the individual biosample accession numbers are denoted in the seventh column. All further analyses were performed on the retrieved 81 genomes.

BGC detection, alignment and visualization

Biosynthetic gene clusters (BGCs) were then predicted using two different tools: a homology-based tool, antiSMASH 7.0 [11], and a tool that detects de novo BGCs, GECCO (GEne Cluster prediction with COnditional random fields) v0.9.10, which is based on a machine-learning approach rather than a rule-based one [12]. Relaxed conditions were applied when using antiSMASH to determine well-defined clusters, in addition to partial clusters, which are missing one or more functional parts, as the standard parameters are optimized for highly conserved BGCs [11], which may not be applicable with noncanonical BGCs, as shown by different studies [13, 14]. In GECCO, the default parameters were used: a predicted BGC must contain a minimum of three consecutive genes to be considered as such, and the used minimum probability for a gene to be considered a part of a BGC was 0.8, as lower numbers were found to reduce the accuracy of the findings [12]. All of the predicted proto-clusters were then compiled together. Clinker v0.0.31 [15] was used to align the selected BGCs for RiPP-like and cyclic lactone auto-inducer BGCs. Only a subset of the BGCs were aligned using clinker, so only representative BGCs were aligned for each group. PRISM 4 was also used to predict the structures of the SMs [16].

Detection of primary metabolism gene clusters (GCs) and structure prediction

In order to predict primary metabolic pathways in selected E. faecium genomes, gutSMASH version 1.0 web-based prediction tool was used [17].

Detection of AMR and virulence genes

Antimicrobial resistance (AMR) genes were identified using the Comprehensive Antibiotic Resistance Database (CARD 4.0.1, https://card.mcmaster.ca/) through the Resistance Gene Identifier (RGI v6.0.5; accessed on 30/11/2025). Genome assemblies were analyzed using the “perfect and strict hits only” criteria with default parameters. To complement this analysis and ensure robustness, AMR gene screening was also performed using ABRicate (v1.2.0; accessed on 30/11/2025; https://github.com/tseemann/ABRICATE). Hits were filtered based on a minimum sequence identity of ≥ 90% and coverage of ≥ 90%.

Furthermore, virulence genes were identified using the Virulence Factor Database (VFDB), available at http://www.mgc.ac.cn/VFs/ (accessed on 30/11/2025) [18]. The results were verified via ABRicate (v1.2.0; accessed on 30/11/2025; https://github.com/tseemann/ABRICATE). Gene detection was performed using default parameters based on > 90% identity and coverage to the reference database [19, 20].

Molecular typing

The sequence type (ST) was determined using the public databases for molecular typing and microbial genome diversity (PubMLST; https://pubmlst.org/organisms/enterococcus-faecium). Allele detection was based on 100% identity and 100% coverage. Major sequence types (ST17, ST18, ST80, ST736) were compared for AMR burden, virulence gene prevalence, and BGC distribution using Kruskal–Wallis tests followed by pairwise Dunn's post-hoc tests when significant. Fisher's exact tests assessed specific gene associations with individual STs. Gene prevalence heatmaps stratified by ST were generated to visualize ST-specific genomic features.

Gene co-occurrence network analysis

Presence/absence data for 54 genes across all E. faecium isolates were extracted from the above-mentioned datasets. Genes were categorized into three functional groups: specialized metabolites (n = 6), virulence factors (n = 5), and antimicrobial resistance genes (n = 43). Pairwise gene associations were assessed using the phi coefficient (φ), which is equivalent to the Pearson correlation coefficient for binary (presence/absence) data. The phi coefficient (φ) was used to quantify association strength and direction (range: −1 to + 1). Statistical significance was determined using Fisher's exact test with Benjamini–Hochberg correction for multiple testing (adjusted p < 0.05 considered significant). Fisher’s exact test was used as the default method to assess the significance of associations. The chi-square test was applied only when all expected cell counts were greater than five, in accordance with standard statistical guidelines. The Jaccard index quantified co-occurrence similarity, and odds ratios with 95% confidence intervals were calculated to measure the magnitude of association.

Co-occurrence networks were constructed using correlation-based adjacency matrices [21]. Multiple threshold combinations were tested to determine optimal network structure. The final network employed conservative criteria: |φ|≥ 0.30 and FDR-adjusted p < 0.05. This threshold was selected based on established effect-size guidelines by Cohen et al. [20], in which φ ≥ 0.30 indicates a medium effect size, ensuring biologically meaningful associations while minimizing false positives.

Networks were visualized as undirected, weighted graphs, with nodes representing genes and edges representing significant associations [22] with edge weights proportional to the strength of the correlation. Positive correlations indicate that genes co-occur, while negative correlations indicate mutual exclusion. Identified associations represent statistical co-occurrence, not causal relationships [20].

Integration of genomic and epidemiological data

Genomic profiles, including AMR genes, virulence determinants and BGCs, were integrated with epidemiological metadata (ST, isolation source, geographic location and host disease category) using isolate identifiers as unique keys. Geographic locations were grouped into broader regions (USA, India, China, and Other) and isolation sources were grouped into clinically relevant categories (Blood, Rectal/Fecal, Urine, and Other). Disease categories were similarly consolidated to ensure adequate group sizes for statistical testing.

Sequence type and metadata associations

Associations between sequence type and epidemiological variables (geographic region, isolation source and disease category) were evaluated using Pearson’s chi-square tests. Only sequence types with sufficient representation were included in visualizations to avoid overplotting.

AMR, virulence and BGC burden

AMR burden was calculated as the total number of detected resistance genes per isolate. Virulence and BGC burden were defined analogously as the number of detected genes within each category. Differences in AMR burden across metadata categories were assessed using the Kruskal–Wallis test.

Gene-level association testing

Individual AMR, virulence and BGC genes were tested for association with geographic region, isolation source and sequence type using Fisher’s exact tests. To account for multiple comparisons, p-values were adjusted using the Benjamini–Hochberg false discovery rate (FDR) method.

Data visualization

Co-occurrence network and heatmaps were visualized using R packages (igraph, corrplot and pheatmap) [22]. Network layouts employed the Fruchterman-Reingold force-directed algorithm. Node size was scaled proportionally to gene prevalence, and node color indicates functional category [23]. Statistical analyses were performed in R version 4.4.0.

To visualize epidemiological structuring of lineages, a chord diagram was generated using the circlize package in R. The diagram linked sequence types to grouped geographic regions, isolation sources, and disease categories, with chord width proportional to the number of isolates in each ST–metadata combination. Only sequence types with ≥ 5 isolates were included in the chord diagram to enhance interpretability. Heatmaps were generated using the pheatmap package to display the prevalence of selected AMR and virulence genes across sequence types.

All statistical analyses and visualizations were performed using R version 4.4.0, with two-sided p-values < 0.05 considered statistically significant.

Results

BGCs detected in the selected E. faecium genomes

Upon utilizing antiSMASH, a total of 334 BGCs were detected that pertain to four BGC types: RiPP-like (166), RRE-containing (2), type III polyketide synthase (81), and cyclic lactone auto-inducer (85) BGCs (Fig. 1A). Both RiPP-like and RRE BGCs are associated with the detection of ribosomally synthesized and post-translationally modified peptides (RiPPs); however, they differ in their levels of specificity and function. While RRE identification relies on the detection of the RRE (RiPP recognition element) domain, RiPP-like BGC detection relies on broader profiles that may include the RRE domain and also other characteristic features [11].

Fig. 1.

Fig. 1

Hierarchical clustering of E. faecium genomes based on the detected BGCs, virulence and AMR genes. A Heatmap for the BGCs detected by both rule-based and machine learning approaches. B Virulence genes heatmap. C AMR genes heatmap

Upon using GECCO, 83 BGCs were identified in the genomes included in the study. Out of the total detected BGCs, 42 were identified as RiPP BGCs (most of which had already been detected by antiSMASH), 38 as saccharide BGCs and 3 were unknown BGCs (Fig. 1A). The following protein domains were identified in the unknown BGCs: PF09221 and two PF10551 domains in the first BGC; PF00665, PF09221 and PF10551 in the second; and PF01721, two PF00665, and PF03412 in the third. Based on the Pfam database, PF09221 and PF01721 are class II bacteriocin-related peptides, while the other identified domains encode for putative transposases or integrases [24].

Using PRISM, 119 BGCs were detected, with 81 polyketide BGCs and 38 class II/III bacteriocin BGCs. All detected BGCs were also detected by antiSMASH. Only two structures were predicted using PRISM (Table S2), both pertaining to the class II/III bacteriocin type.

Clinker results for aligned CAL-BGCs showed that almost all clusters were closely related, either having the exact arrangement, an inverted arrangement, or a few inserted or deleted genes, typically transposases or insertion sequences (Figure S1A). However, only three clusters were totally different from the rest (Figure S1A). Concerning the RiPP-like BGCs (Figure S1B), similarly aligned RiPP-like BGCs showed a shared conserved core structure amongst representative genomes, with repeated blocks of high sequence similarity across clusters. These correspond to genes encoding the precursor peptide and core modification enzymes typical of RiPP biosynthesis. As illustrated in Figure S1B, the central RiPP cassette exhibited a similar orientation and is essentially identical in the upper and middle clusters. However, slight variability is observed in the flanking regions where differences in gene content and arrangement were evident across genomes. On the other hand, other regions deviate from the main pattern by demonstrating notably truncated versions without flanking regions, which may represent more divergent RiPP-like loci.

Primary metabolism GCs identified in the selected E. faecium strains

Four main pathways were detected, namely, the arginine to hydrogen-carbonate pathway, the PFOR-II pathway, the pyruvate to acetate-formate pathway, and the gallic acid metabolism pathway.

Virulence factor genes detected in the selected E. faecium strains, and dominant clones

Five different genes encoding virulence factors were detected in the included genomes, namely scm (collagen adhesin protein), fss3 (Enterococcus faecalis surface protein Fss3 fibrinogen-binding protein), acm (collagen adhesin precursor), sgrA (cell wall-anchored protein SgrA), and ecbA (encoding a microbial surface component recognizing adhesive matrix molecules called EcbA, involved in bacterial adhesion and biofilm formation) (Fig. 1B). Major clonal lineages are ST736, ST17, ST18, and ST80, each was detected in 10–20% of the strains.

AMR genes detected in the selected E. faecium strains

Forty-three distinct AMR genes were detected within the genomes (Fig. 1C), in the range of 2–20 AMR genes per genome. Based on genotypes, multidrug resistance was observed in a high proportion of strains: 96.3% carried > 3 AMR genes across multiple classes and almost 70% carried ≥ 15 AMR genes conferring resistance to at least 5 classes of antibiotics, suggesting an extensive drug resistance phenotype. Alarmingly, 81.5% of the strains were vanA-positive, which causes resistance to vancomycin, with multiple variants of the gene (vanA, vanB, vanHA, vanHB, vanHM, vanM, vanRA, vanRB, vanRM, vanSA, vanSB, vanSM, vanWB, vanXA, vanXB, vanXM, vanYA, vanYB, vanYM and vanZA). Near-universal aminoglycoside resistance was also noted based on the genotype, with multiple AMR genes such as aad(6), aac(6′)-Ie-aph(2″)-Ia, aac(6')-Ii and ant(6)-Ia, aph(3')-IIIa. Macrolide resistance genes were also found, namely, efmA, msrC, ermA, ermB, and ermT. Tetracycline resistance genes were detected as well (tet(L), tetM, tetS, and tetU), in addition to other classes of antibiotics (SAT-4, catA8, dfrF, dfrG, fexA, lnuG, lsaE, optrA, and poxtA) (Fig. 1C).

The counts of all identified BGCs, virulence and AMR genes in the studied genomes are depicted in Figure S2.

Correlation and network analysis

Among the 58 genes investigated, 51 were included in the network, as the rest were consistent in all samples. As shown in Fig. 2, 113 significant associations were identified (p < 0.05, |r|> 0.2), of which 102 (90.3%) were positive correlations indicating co-occurrence, and 11 (9.7%) were negative correlations indicating mutual exclusion.

Fig. 2.

Fig. 2

Network visualization of significant gene–gene associations in E. faecium (n = 51 nodes, 113 edges). Node color indicates functional category, while edge color and style represent association type: solid red lines indicate positive correlations (co-occurrence), dashed blue lines indicate negative correlations (mutual exclusion). Colored oval regions delineate communities detected by Louvain clustering

Eight distinct communities were identified (Fig. 3A and Figure S3), with the largest cluster containing the vanA operon (vanA, vanHA, vanRA, vanSA, vanXA, vanYA and vanZA) plus diverse AMR genes (aac(6')-Ie-aph(2'')-Ia, aph(3')-IIIa, ermB, sat-4, aad(6), efmA and tetU), one virulence factor (fss3), and two BGCs (cyclic-lactone-autoinducer, RRE-containing). The complete vanB operon showed perfect co-occurrence for vanB, vanHB, vanRB, vanSB, vanWB, vanXB, and vanYB. The same was noted for the vanM operon (vanHM, vanM, vanRM, vanSM, vanXM and vanYM). Additionally, we identified a mixed cluster containing tetracycline resistance genes (tet(L), tetM, tetS, tetU), folate pathway inhibitors (dfrF, dfrG), ermT, and metabolite biosynthesis genes. Other clusters were enriched in newer resistance mechanisms, including lsaE, optrA, and poxtA, with the virulence factor sgrA. Singleton communities containing single genes (ermA and msrC) were also observed (Figure S3).

Fig. 3.

Fig. 3

Hierarchical clustering of the detected AMR, virulence genes & BGCs in the E. faecium genomic dataset. A Gene correlation heatmap for the significant values only. Heatmap showing pairwise Pearson correlation coefficients between genes with significant associations (p < 0.05). Red indicates positive correlations (co-occurrence), blue indicates negative correlations (mutual exclusion), and color intensity represents correlation strength (scale: −1.0 to 1.0). Dendrogram branches indicate similarity in co-occurrence patterns. B Top gene associations heatmap

Perfect co-occurrence was observed between all vanB operon gene pairs (21 associations), core vanA operon genes: vanA-vanHA, vanA-vanXA, vanHA-vanXA, and BGCs of cyclic-lactone-autoinducer and RRE-containing (Fig. 3A). On the other hand, perfect mutual exclusion was detected between the cyclic lactone autoinducer and efmA, as well as between RRE-containing and efmA. Strong negative correlations were identified between lsaE (linezolid resistance) and vanA operon components: vanA, vanHA, vanXA, vanRA, vanSA, vanYA and vanZA, and between scm and tetU (Fig. 3A).

Cross-category associations revealed an AMR dominant pattern, representing vancomycin resistance operons and linked resistance determinants. AMR-virulence showed limited associations, suggesting independent acquisition patterns, whereas AMR-specialized metabolites exhibited negative associations, particularly the exclusion of efmA from metabolite-producing strains. The heatmap (Fig. 3B) clearly depicts the top association among the genes investigated in this study.

Integration of genomic features with clinical and epidemiological metadata

AMR gene burden among the 81 isolates varied significantly by isolation source (Kruskal–Wallis χ2 = 15.14, p = 0.0098, FDR-adjusted p = 0.027) and geographic region (χ2 = 18.41, p = 0.018, FDR-adjusted p = 0.027), but not by host disease status (χ2 = 6.93, p = 0.33). Rectal/fecal isolates carried the highest AMR burden (median = 15 genes, IQR: 14–15), compared to blood/bloodstream isolates (median = 13 genes, IQR: 12–15). Geographically, USA isolates exhibited higher resistance (median = 15 genes) than Chinese (median = 7–8 genes) and Indian (median = 13–15 genes) isolates.

Gene-specific associations with metadata

To determine whether differences in AMR burden reflected quantitative expansion of shared genes or qualitative shifts in gene composition, individual gene–metadata associations were assessed using Fisher’s exact tests. One resistance determinant, tet(L), showed a significant source-specific distribution (FDR-adjusted p = 0.0198), with a higher prevalence in blood isolates (75.0%) than in rectal samples (32.5%; OR = 6.05).

Sequence type distribution across epidemiological variables

A total of 21 distinct sequence types were identified, with ST(736) (n = 16), ST(17) (n = 12), ST(18) (n = 11), and ST(80) (n = 11) representing the most frequent lineages. Sequence type distribution differed significantly across isolation source (χ2 = 150.92, df = 60, p = 8.73 × 10⁻1⁰), geographic region (χ2 = 202.22, df = 60, p < 2.2 × 10⁻1⁶) and disease category (χ2 = 130.82, df = 60, p = 3.5 × 10⁻⁷), demonstrating strong epidemiological structuring of lineages (Fig. 4A, B).

Fig. 4.

Fig. 4

The association of STs with AMR, virulence genes and epidemiological variables. A Heatmap of key AMR and virulence genes across STs. Prevalence (%) of selected AMR and virulence genes across STs. Color intensity reflects the proportion of isolates within each ST carrying the corresponding gene. Clustering was performed based on Euclidean distance to highlight lineage-specific genomic patterns. B Chord diagram of the associations between STs and epidemiological variables. The diagram links sequence types (STs) with grouped geographic regions, isolation sources, and disease categories. Each sector represents either an ST or a metadata category, and chord width is proportional to the number of isolates in each ST–metadata combination. Only STs represented by ≥ 5 isolates were included to enhance interpretability.

Geographically, ST(736), ST(17), and ST(18) were predominantly observed among USA isolates, whereas ST(80) was enriched in India. Regarding the isolation source, ST(736) was strongly associated with bloodstream isolates, whereas ST(17) and ST(18) were primarily recovered from rectal samples. Disease distribution also differed across STs, with ST(736) and ST(80) frequently linked to bacteremia/sepsis (Fig. 4A, B).

AMR burden across sequence types

AMR gene burden differed among sequence types. ST(17) exhibited the highest mean of resistance gene count (mean = 16.2 genes, median = 17), followed by ST(18) (mean = 15.9 genes, median = 16). ST(736) and ST(80) displayed moderately lower AMR burdens (means = 14.1 and 14.7 genes, respectively), whereas ST(78) showed comparatively lower resistance (mean = 10.4 genes). These findings indicate lineage-associated differences in resistance accumulation.

Virulence and BGC profiles across STs

Virulence gene burden showed limited variation across STs (mean range 3.5–4.5 genes), suggesting a relatively conserved accessory virulome within the population. No variation in BGC content was detected across sequence types in this dataset.

Individual gene associations with sequence type and metadata

Beyond overall AMR burden, several individual resistance determinants demonstrated significant associations with epidemiological variables after FDR correction. Resistance genes dfrF and tetM were strongly associated with geographic region (FDR < 5 × 10⁻⁸), with markedly higher prevalence among USA-associated sequence types. Similarly, vanA, vanHA, and related vancomycin resistance genes exhibited significant geographic enrichment.

At the lineage level, tet(L) and dfrF showed highly significant associations with sequence type (FDR < 10⁻12), indicating strong ST-dependent distribution of specific resistance determinants.

Virulence genes also displayed lineage-associated patterns. The adhesin genes ecbA and scm were significantly associated with both isolation source and sequence type (FDR < 0.001), while fss3 and sgrA demonstrated moderate ST-associated enrichment. In contrast, no significant associations were detected for BGC elements across metadata categories.

Figures 4A and B show the association of ST with the source of isolates and key AMR and virulence genes. This highlights the strong epidemiological structuring of lineages across geographic and clinical contexts.

Discussion

In this study, we inspected the potential relationships and co-occurrence profiles of different BGCs, AMR and virulence genes identified in 81 assembled E. faecium genomes. Our search detected several predicted classes of BGCs in E. faecium pertaining to RiPP-like, RRE-containing (RRE-element-containing cluster), type III polyketide synthase, and cyclic lactone auto-inducer BGCs. E. faecium genomes seem to have a common BGC signature (Fig. 1A), with those four BGCs constituting possible pathogenic gene clusters. GECCO identified fewer, but similar, RiPP-like BGCs to those detected by antiSMASH, and it did not detect type III polyketide synthase and cyclic lactone auto-inducer BGCs, indicating a higher sensitivity of the used relaxed conditions in antiSMASH. However, GECCO was able to identify de novo saccharide BGCs. Further analysis of BGCs identified by GECCO was limited, as the tool relies on detecting protein domains rather than chemical structure similarities [12]. The few predicted SM structures from E. faecium warrant further investigation, possibly by other tools [25].

Studying detected BGCs is a key factor in identifying potential antimicrobials and antimicrobial targets. For instance, in terms of survival and self-defense against other microbes, bacteria produce bacteriocins [26]. It was reported earlier that a subset of E. faecium strains are resistant to class II bacteriocins [27], which is expected as E. faecium strains are reported to produce them [28]. Our results corroborate these reports; as in the included genomes, some BGC hits shared known BGCs, such as enterocin A, which is a bacteriocin [26]. These were classified by antiSMASH as RiPP-like BGCs. Enterocin A was the first enterocin to be purified and genetically characterized [29]. BGCs similar to carnobacteriocin XY (Cbn XY) were also detected in the comprised E. faecium genomes, hinting that these strains can possibly produce CbnXY-related bacteriocins, which can kill other Gram-positive strains [30, 31]. BGCs similar to those coding for enterocin NKR-5-3B were found in the included E. faecium genomes; that is consistent with previous studies reporting that the same species produces it [32]. BGCs with similarity to the lantibiotic lactocin S BGC were detected as well, although experimentally, it was proven that E. faecium produces rather a cytolysin, a two-peptide bacteriocin [33]. In fact, both lactocin S and cytolysin have antimicrobial properties, but their mechanisms of action, target cells and regulation mechanisms differ [34]. Moreover, after refining the analysis of the detected “unknown” clusters by GECCO, we believe they could be classified as potential bacteriocin-like clusters as well, based on the co-localization of structural precursor class II bacteriocin peptides (PF01721/PF09221); a dedicated maturation machinery, specifically the C39 Peptidase (PF03412) [35]; and other putative transposases and integrases (PF10551 & PF00665) normally detected in BGCs [36, 37]. These clusters were labeled by GECCO as “unknown” because its algorithm detected a high probability of core biosynthetic Pfams; however, after detecting different BGCs, GECCO tries to assign them to different classes based on their domain arrangement and sequence homology to experimentally characterized BGCs in the Minimum Information about a Biosynthetic Gene cluster (MIBiG) database [12]. For these “unknown” BGCs, the specific peptide sequence shared low homology with previously characterized bacteriocins in the MIBiG 2.0 database; thus, they were assigned as unknown BGCs. Hence, it is essential to validate the computational findings experimentally and verify whether the detected BGCs encode for novel bacteriocins or not. The detected output from antiSMASH, GECCO and PRISM intersect in terms of the classes of the bacteriocins detected, with PRISM being able to classify the bacteriocins better functionally.

Using GECCO, identified BGCs fell within the groups of RiPPs, saccharides, and unknown BGCs. Identified RiPPs and saccharides could serve as potential targets for novel antimicrobials. For instance, Secreted antigen A (SagA), a RiPP essential for cell wall metabolism and bacterial growth in E. faecium, was also found to be an active player in inducing immune response in the host. As SagA deletion renders the bacterial cells more susceptible to cell-wall-acting antibacterials, they may act as potential targets for new antibacterials [38]. In addition, it was reported before that sugar metabolism genes and sugar transport genes were abundant amongst E. faecium clinical isolates [39]. An earlier study proved that targeting phosphoenolpyruvate: sugar phosphotransferase system (PTS) enzymes kills the E. faecium pathogen, and hence, further exploring the saccharide BGCs can possibly help in developing antimicrobials targeting these systems [39]. Although identifying novel E. faecium BGCs is useful, experimental validation is necessary to determine the exact function of the detected de novo BGCs and their role in the survival and pathogenicity of E. faecium, as well as their potential pharmaceutical applications.

Additionally, BGCs can encode SMs involved in a variety of quorum sensing-controlled regulatory processes, such as pathogenicity and virulence. Quorum sensing is a population-density-based regulation process. Many bacterial species can produce small molecules known as autoinducers that can be secreted into the outer environment. When these autoinducers reach a certain concentration threshold sensed by bacterial surface receptors, a cascade of signal transduction is initiated [40]. Quorum sensing systems are well-documented in both Gram-negative and Gram-positive bacteria. Gram-positive organisms, including Staphylococci and Enterococci, predominantly employ ribosomally synthesized and post-translationally modified autoinducing peptides (AIPs) [40, 41]. For instance, in species belonging to Staphylococcus, mature AIPs, which contain either a cyclic thiolactone or lactone ring [42], are detected by the AgrC transmembrane histidine kinase when they reach a sufficiently high concentration. This initiates a signal transduction cascade and the expression of various genes [43]. In parallel, the LuxS/Autoinducer-2 (AI-2) system produces a furanosyl borate diester signal that can be found in both Gram-positive and Gram-negative bacteria and can facilitate inter-species communication [43]. In E. faecalis, many virulence factors are regulated by a quorum sensing circuit that shows similarity to the one found in S. aureus. Quorum-sensing (QS) systems in other enterococcal species have not been fully characterized yet [43]. However, E. faecium is known for its biofilm formation potential, which is an important virulence trait that quorum-sensing circuits could control [3]. Understanding quorum sensing in E. faecium is crucial for developing inhibitors that disrupt intercellular bacterial communication. Targeting the QS circuit in Enterococcus species has great potential in inhibiting biofilm formation and attenuating the pathogenicity and virulence of this pathogen [44].

In our work, we detected BGCs in which some are predicted to produce cyclic lactone autoinducers (CAL). These candidate gene clusters contain genes encoding AI-2E family transporters, histidine kinases, AgrB-like accessory regulators, other transporters and DNA-binding proteins. Almost all genomes showed similar CAL-BGCs, with a few differences in arrangements or with the insertion and deletion of different transposases and insertion sequences; however, two of the studied genomes possessed different extra putative CAL-BGCs (Figure S1A). Although antiSMASH annotated the identified loci as CAL-BGCs, functional interpretation of these regions in E. faecium necessitates careful studying. AntiSMASH predictions are primarily based on rule-driven pattern recognition and homology to previously characterized biosynthetic architectures rather than direct evidence of metabolite production or chemical structure [11]. In Enterococci, experimentally validated quorum sensing systems are predominantly peptide-dependent and regulatory in nature rather than driven by canonical small-molecule lactone synthases such as in the fsr system in E. faecalis [45]. The gene composition of the CAL-like loci detected here, which is enriched with histidine kinases, transport-associated proteins and regulatory elements, is therefore consistent with potential roles in signal perception, environmental sensing, or regulatory modulation. Notably, transporters such as AI-2E family proteins may participate in intercellular signaling or metabolite trafficking without implying autoinducer synthesis [46]. Moreover, phylogenomic analyses across diverse bacterial taxa indicate that the presence of histidine kinases and regulatory domains within predicted biosynthetic loci does not necessarily imply a biosynthetic function, as these components may instead mediate regulatory or signaling processes associated with cluster activity [47]. Consequently, in the absence of biochemical characterization or gene expression evidence linking these clusters to lactone synthesis, we should regard them as candidate signaling or regulatory modules pending experimental validation.

It is also useful to study the modes of the bacterium’s primary metabolism, as it may offer insights into both commensal and pathogenic behaviors, which may allow specific targeting of these pathways. Primary metabolites constitute the largest share of the chemical output of the gut microbiome [17]. As they are essential for bacterial survival and energy utilization, it was interesting to investigate whether we could detect any differences in the primary metabolic gene clusters (MGCs) among the studied genomes. Nevertheless, the same MGCs were detected in all examined genomes, indicating their essential role in survival and energy metabolism in E. faecium. The arginine-to-hydrogen-carbonate pathway (the same as the arginine deiminase (ADI) pathway) [48] is employed by similar microbes to convert L-arginine to L-ornithine, producing ATP. Therefore, it is a unique pathway that generates energy for the E. faecium. Another pathway with its GCs detected in the E. faecium genomes was the pyruvate:ferredoxin oxidoreductase (PFOR) II pathway [49]. This pathway explains the mode of anaerobic respiration that E. faecium employs to survive in the absence of oxygen. The PFOR II pathway enables the bacterium to synthesize pyruvate as well as being employed in the reverse direction to form acetyl-CoA from pyruvate [50]. Another metabolic pathway detected in the comprised genomes was the pyruvate to acetate-formate pathway, which employs pyruvate formate-lyase (PFL) [51, 52]. In anaerobic conditions, this pathway enables the bacterium to metabolize pyruvate into acetyl-CoA and formate. Genes pertaining to the gallic acid metabolism pathway were also identified in the selected E. faecium genomes [53]. It is interesting to further explore this pathway in E. faecium, as a part of the gut microbiome, as gallic acid is usually a product of edible food and a component of many fruits, and it needs to be metabolized to have advantageous effects on the human gut, such as antioxidant and anticancer effects [54]. This observation suggests that E. faecium can be a beneficial member of the human gut microbiome and could have a possible probiotic effect, albeit it needs further experimental validation [54]. Furthermore, as these energy-acquisition pathways are conserved among potentially highly resistant strains, as shown from the detected AMR profile in this study, they may serve as stable targets for metabolic inhibitors.

Different virulence factors belonging to microbial surface component-recognizing adhesive matrix molecules (MSCRAMM) were detected in the studied genomes, such as acm, scm, fss3, sgrA and ecbA genes (Fig. 1B). It is well-documented that these genes have a role in virulence and biofilm formation [55]. The acm gene in particular was ubiquitous in all the genomes in the study dataset. Acm, EcbA and Scm are virulence factors of E. faecium that enable the pathogenic bacterium to bind to collagen [55]. Another detected MSCRAMM that alternatively binds to fibrinogen was Fss3 [56]. In addition, we detected the gene encoding the surface adhesin SgrA in most genomes. This adhesin facilitates the binding of the pathogen to the host cell [57].

We also noticed that a few genomes were devoid of fss3, sgrA and ecbA genes. This may point to the possibility that biofilm formation is less important in environments from which these strains were isolated. One of these strains was isolated from pet food, and another from a probiotic product. Although the remaining samples were fecal, they may be commensals with lower biofilm-forming abilities. Although biofilm formation is an important virulence factor, it is not clear whether these strains have lower pathogenic abilities, as other virulent factors may contribute to their potential pathogenicity.

The diversity and abundance of the AMR genes detected within the E. faecium genomes point towards the high antibiotic resistance characteristic of this species (Fig. 1C). One particular AMR gene detected in the entirety of the genomes was aac(6')-Ii, which encodes an aminoglycoside acetyltransferase that is chromosomally encoded [58]. The same aforementioned resistance gene was detected in isolates obtained from in-patient subjects in hospitals [58]. Another AMR gene commonly detected amongst most of the comprised genomes was the msrC gene, which is chromosomally encoded in the bacterial genome and confers resistance to streptogramin B, erythromycin and several other macrolide antibiotics [59]. The mscrC gene pertains to the ABC-F efflux family and is involved in the efflux and transport of antibiotics outside of the bacterium [60]. Several AMR genes were detected in most of the genomes simultaneously, namely vanYA, vanZA, vanRA, vanSA, vanXA, vanHA, and vanA. The first six genes are the assortment of the vanA-type cluster, for which the bacterium’s resistance to vancomycin is attributed to [61]. Additionally, the detected vanA gene encodes an enzyme that synthesizes D-Ala-D-Lac, an alternative substrate that would hinder vancomycin from binding to its target [62]. The presence of those particular genes together in the same genomes strongly suggests that they would all confer vancomycin resistance to the genomes harboring them. The increasing number of genomes containing vancomycin-resistant genes raises concerns about the potential threat this bacterium poses to humans. Furthermore, some of the genes uniquely detected in one genome were catA8, vanXM and vanRM. Indeed, catA8 was reported to be an enzyme that inactivates chloramphenicol antibiotic, as it encodes a chloramphenicol acetyltransferase (CAT), and it is less common than catA9 [63]. Two genes pertaining to the vanM cluster were also uniquely detected in the included genomes, which points to the potential presence of the vanM cluster and further augments the problem of glycopeptide and vancomycin resistance in those strains, as it functions in a similar way to the vanA cluster [64]. Antimicrobial resistance assays such as the Minimum Inhibitory Concentration (MIC) assays could be performed to correlate the detected genotype with the expected phenotype.

The observed gene co-occurrence network reveals a highly structured genetic organization in E. faecium, dominated by three distinct vancomycin resistance operons that show perfect internal co-occurrence (Figs. 2 and 3A, B). The vanB operon exhibits complete linkage, indicating these genes are physically linked on a single mobile genetic element, most likely a conjugative transposon of the Tn1549 family [64, 65]. Similarly, the vanA and vanM operons maintain high internal cohesion, consistent with their organization on transferable plasmids or integrative conjugative elements [65, 66]. The perfect co-occurrence of genes within resistance operons reflects their functional interdependence, as vancomycin resistance requires coordinated expression of multiple genes for modification of cell wall precursors. Physical linkage on mobile genetic elements ensures co-inheritance and prevents separation of functionally coupled genes. This has important implications, as targeting individual genes within these operons will not eliminate resistance because the entire genetic cassette is transmitted as an indivisible unit during horizontal gene transfer [67]. In contrast, chromosomally encoded resistance genes may exhibit stable lineage-specific inheritance patterns. Therefore, observed gene co-occurrence patterns may reflect a combination of clonal expansion and mobile genetic element–mediated transfer. Future studies incorporating plasmid reconstruction and long-read sequencing would help resolve the genomic context of these determinants.

The perfect mutual exclusion between metabolite BGCs (CAL-BGCs and RRE-containing) and the efmA multidrug efflux pump strongly suggests the presence of distinct E. faecium lineages or clades in this population. This pattern indicates two non-overlapping genetic backgrounds. One lineage that is characterized by specialized metabolite production, including quorum-sensing or other signaling systems (CAL-BGCs) and RiPP-like compounds. These isolates lack efmA, while another lineage is characterized by efmA-mediated multidrug efflux. These isolates lack these particular SM biosynthesis pathways. This genetic segregation may reflect ecological niche divergence, as well as clonal expansion for hospital-adapted clones versus community-associated strains. It may also reflect functional incompatibility, with competitive exclusion or fitness costs preventing coexistence of these genetic modules. The high prevalence of metabolite genes suggests these are core features of certain E. faecium lineages, possibly providing competitive advantages in the gut microbiome through SMs encoded by CAL-BGCs and antimicrobial peptide production. Moreover, the strong negative correlation between lsaE and the vanA operon reveals potential fitness trade-offs between different resistance mechanisms. lsaE confers resistance to linezolid and other oxazolidinones, representing a newer resistance threat, while vanA provides vancomycin resistance, a long-established mechanism. Several explanations for this mutual exclusion include metabolic burden mediating simultaneous maintenance of both resistance systems, which may impose excessive fitness costs, particularly in the absence of antibiotic pressure [68]. It may also indicate plasmid incompatibility, as vanA (typically plasmid-borne) and lsaE may reside on incompatible replicons. Lineage association is also possible, as these resistance genes may be preferentially associated with the distinct lineages. This pattern has clinical significance, as sequential antibiotic pressure may select for different resistance profiles. If vancomycin use is reduced in favor of linezolid, the population structure may shift from vanA-dominant to lsaE-dominant strains rather than accumulating both resistance mechanisms. This finding may also imply that rotation between vancomycin and linezolid may effectively change the population structure and prevent the emergence of a pan-resistant super-lineage. It is important to note that the detected associations reflect statistical co-occurrence patterns and may not reflect causal relationships. Co-occurrence may result from physical gene linkage on plasmids or transposons, co-selection under antimicrobial pressure, or clonal population structure. Further experimental validation is required to establish mechanistic relationships. This could be done, for instance, by RNA-Seq or other transcriptomics studies to validate the co-transcription of co-occurring genes under certain stresses.

In terms of molecular epidemiology, the strains investigated belonged to dominant sequence types from high-risk clones, such as ST736, ST17, ST18, and ST80, most of which are epidemic clones associated with high virulence and resistance (Fig. 4A, B) [69, 70]. This fact also underscores the importance of our study, which explores the genomic properties of these globally dominant clones. Importantly, although AMR genes exhibited strong lineage- and geography-associated structuring, BGC content remained highly conserved across STs. This suggests that resistance acquisition and clonal expansion are primarily driven by selective pressures on AMR determinants rather than by diversification of SM repertoires. The relatively stable BGC architecture may reflect conserved ecological functions, whereas AMR appears more dynamically shaped by horizontal gene transfer and clinical antibiotic pressure. Furthermore, it is noteworthy that the integration of both epidemiological metadata along with the genomic data may provide a roadmap for tracking the spread of “high-risk clones”. For instance, in our limited dataset, ST(736) was associated with bloodstream infections, and ST(17) and ST(18) were associated with the highest AMR burden. Perhaps such findings may help public health laboratories to use sequence typing as a predictor of clinical outcome after experimental validation that these genes are actually being transcribed.

A limitation of this study is that the definitive genomic localization of AMR and BGC sequences could not be determined due to the use of draft genome assemblies and utilizing in silico analyses. Consequently, it was not possible to distinguish between plasmid- and chromosome-associated resistance genes. While the assembly quality was sufficient for gene identification and functional annotation, the presence of repetitive flanking regions often precludes the assembly of complete circularized replicons. Future studies incorporating long-read sequencing or experimental validation (e.g., Southern blotting) would be required to resolve gene localization.

Lastly, it is important to study the resistance of AMR pathogens, including E. faecium, to identify targets for novel treatments or for other antibiotics to which they have not yet developed resistance. Studying the primary MGCs, specialized metabolism BGCs, AMR genes, and virulence genes is useful in the overall evaluation of the pathogen considering the different discussed attributes in order to eradicate the pathogen (Fig. 5, Figure S2). We herein suggest looking at metabolism, virulence factor genes, and antimicrobial resistance genes, collectively, to combat these harmful pathogens.

Fig. 5.

Fig. 5

A pathogenomic signature approach to understand and combat pathogens, like E. faecium

Conclusion

This study provides an integrative genomic analysis of Enterococcus faecium, revealing the diversity of its biosynthetic gene clusters (BGCs), antimicrobial resistance (AMR) genes, virulence factors, and primary metabolic pathways. Our findings highlight that E. faecium possesses a conserved core of identified BGCs, such as RiPP-like and cyclic lactone autoinducer BGCs, with several genomes harboring novel or divergent clusters, possibly contributing to strain-specific traits such as quorum-sensing coordination or other regulatory pathways (Fig. 5). The relatively conserved BGCs and MGCs may reflect stable ecological functions and act as potential targets for new therapeutics. In addition, the predominance of vanA-type VRE, combined with extensive multidrug resistance, highlights the critical public health threat posed by these isolates in clinical and community settings, with limited therapeutic options due to extensive resistance and the risk of treatment failure. This also underscores the urgent need for the development of novel therapeutics. A high-risk virulence profile was noted in the strains, reflecting their potential high pathogenicity, invasiveness, enhanced adhesion, colonization, persistence, and ability to survive under environmental stress. Gene co-occurrence network analysis reveals non-random patterns of genetic association in E. faecium, with distinct communities representing functionally related gene modules. Both positive and negative correlations provide insights into the environmental forces shaping genome composition and adaptation. The combination of universal adherence factors, high prevalence of invasive enzymes, and extensive antibiotic resistance creates a potentially highly virulent pathogen profile. These isolates are high-risk healthcare-associated pathogens that require strict infection control and aggressive treatment protocols. Hence, this work underscores the necessity of integrating metabolic, resistance, and virulence data to be used as a foundational dataset for future functional studies and targeted therapeutic strategies to combat E. faecium infections.

Supplementary Information

12864_2026_12884_MOESM1_ESM.pptx (7.2MB, pptx)

Supplementary Material 1: Figure S1. Gene cluster architecture and alignment of the selected BGCs in E. faecium genomes for: (A) cyclic lactone auto-inducer BGCs and (B) RiPP-like BGCs. Figure S2. Hierarchical clustering of E. faecium genomes based on the detected BGCs, virulence, and AMR genes. Figure S3. Gene co-occurrence network with communities of genes for BGCs, AMR and virulence from E. faecium genomes.

12864_2026_12884_MOESM2_ESM.pptx (498.8KB, pptx)

Supplementary Material 2: Table S1. Assembly data for the samples. Table S2. PRISM results for each isolate.

Authors’ contributions

L.Z. & F. A. conceived the idea of the study. All authors contributed in the bioinformatics tools running and visualization of the figures and tables. All authors wrote and revised the manuscript.

Funding

The authors received no funding for this study.

Data availability

All the data that is related to the study are in the figures and tables, both in the main and supplementary files. The genomes included in the study were retrieved from Genome Taxonomy Database, GTDB (https://gtdb.ecogenomic.org). All the genomes used in this study have their accession numbers listed in Supplementary Table S1, in which the first column has the Genbank accession numbers of the assemblies, and the seventh column has the biosample accession number.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.De Oliveira DMP, Forde BM, Kidd TJ, Harris PNA, Schembri MA, Beatson SA, et al. Antimicrobial resistance in ESKAPE pathogens. Clin Microbiol Rev. 2020. 10.1128/CMR.00181-19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Jain A, Oli AK, Kulkarni S, Kulkarni RD, Chandrakanth K. A review on drug resistance patho-mechanisms in ESKAPE bacterial pathogens. Novel Res Microbiol J. 2024. 10.21608/nrmj.2024.287047.1549. The National Information and Documentation Centre (NIDOC) affiliated to Academy of Scientific Research and Technology(ASRT). [Google Scholar]
  • 3.Aruwa CE, Chellan T, S’thebe NW, Dweba Y, Sabiu S. ESKAPE pathogens and associated quorum sensing systems: new targets for novel antimicrobials development. Health Sci Rev. 2024. 10.1016/j.hsr.2024.100155. [Google Scholar]
  • 4.Zhong Z, Kwok LY, Hou Q, Sun Y, Li W, Zhang H, et al. Comparative genomic analysis revealed great plasticity and environmental adaptation of the genomes of Enterococcus faecium. BMC Neurosci. 2019;20. 10.1186/s12864-019-5975-8. [DOI] [PMC free article] [PubMed]
  • 5.Abdelsalam NA, Elhadidy M, Saif NA, Elsayed SW, Mouftah SF, Sayed AA, et al. Biosynthetic gene cluster signature profiles of pathogenic Gram-negative bacteria isolated from Egyptian clinical settings. Microbiol Spectr. 2023;11. 10.1128/spectrum.01344-23. [DOI] [PMC free article] [PubMed]
  • 6.Kayser O, Averesch NJH. Technical Biochemistry: The Biochemistry and Industrial Use of Natural Products. Tech Biochem Biochem Indust Use Nat Prod. 2025. 10.1007/978-3-658-47121-7. [Google Scholar]
  • 7.Medema MH. Computational Genomics of Specialized Metabolism: from Natural Product Discovery to Microbiome Ecology. mSystems. 2018;3. 10.1128/msystems.00182-17. [DOI] [PMC free article] [PubMed]
  • 8.Kenshole E, Herisse M, Michael M, Pidot SJ. Natural product discovery through microbial genome mining. Curr Opin Chem Biol. 2021. 10.1016/j.cbpa.2020.07.010. [DOI] [PubMed] [Google Scholar]
  • 9.Tobias NJ, Bode HB. Heterogeneity in bacterial specialized metabolism. J Mol Biol. 2019. 10.1016/j.jmb.2019.04.042. [DOI] [PubMed] [Google Scholar]
  • 10.Parks DH, Chuvochina M, Rinke C, Mussig AJ, Chaumeil PA, Hugenholtz P. GTDB: An ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Res. 2022;50. 10.1093/nar/gkab776. [DOI] [PMC free article] [PubMed]
  • 11.Blin K, Shaw S, Augustijn HE, Reitz ZL, Biermann F, Alanjary M, et al. AntiSMASH 7.0: new and improved predictions for detection, regulation, chemical structures and visualisation. Nucleic Acids Res. 2023;51:W46–50. 10.1093/nar/gkad344. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Carroll LM, Larralde M, Fleck JS, Ponnudurai R, Milanese A, Cappio E, et al. Accurate de novo identification of biosynthetic gene clusters with GECCO. bioRxiv. 2021.
  • 13.Santos-Aberturas J, Chandra G, Frattaruolo L, Lacret R, Pham TH, Vior NM, et al. Uncovering the unexplored diversity of thioamidated ribosomal peptides in Actinobacteria using the RiPPER genome mining tool. Nucleic Acids Res. 2019;47. 10.1093/nar/gkz192. [DOI] [PMC free article] [PubMed]
  • 14.Fernandez-Cantos MV, Garcia-Morena D, Yi Y, Liang L, Gómez-Vázquez E, Kuipers OP. Bioinformatic mining for RiPP biosynthetic gene clusters in Bacteroidales reveals possible new subfamily architectures and novel natural products. Front Microbiol. 2023;14. 10.3389/fmicb.2023.1219272. [DOI] [PMC free article] [PubMed]
  • 15.van den Belt M, Gilchrist C, Booth TJ, Chooi YH, Medema MH, Alanjary M. CAGECAT: The CompArative GEne Cluster Analysis Toolbox for rapid search and visualisation of homologous gene clusters. BMC Bioinformatics. 2023;24. 10.1186/s12859-023-05311-2. [DOI] [PMC free article] [PubMed]
  • 16.Skinnider MA, Johnston CW, Gunabalasingam M, Merwin NJ, Kieliszek AM, MacLellan RJ, et al. Comprehensive prediction of secondary metabolite structure and biological activity from microbial genome sequences. Nat Commun. 2020;11. 10.1038/s41467-020-19986-1. [DOI] [PMC free article] [PubMed]
  • 17.Pascal Andreu V, Roel-Touris J, Dodd D, Fischbach MA, Medema MH. The gutSMASH web server: automated identification of primary metabolic gene clusters from the gut microbiota. Nucleic Acids Res. 2021;49:W263–70. 10.1093/nar/gkab353. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Al-Marzooq F, Ghazawi A, Allam M, Collyns T, Saleem A. Novel Variant of New Delhi Metallo-Beta-Lactamase (blaNDM-60) Discovered in a Clinical Strain of Escherichia coli from the United Arab Emirates: An Emerging Challenge in Antimicrobial Resistance. Antibiotics. 2024;13. 10.3390/antibiotics13121158. [DOI] [PMC free article] [PubMed]
  • 19.Lipworth S, Crook D, Walker AS, Peto T, Stoesser N. Exploring uncatalogued genetic variation in antimicrobial resistance gene families in Escherichia coli: an observational analysis. Lancet Microbe. 2024;5. 10.1016/S2666-5247(24)00152-6. [DOI] [PMC free article] [PubMed]
  • 20.Cohen J. Statistical power analysis for the behavioural sciences. Hillside. NJ: Lawrence Earlbaum Associates; 1988. [Google Scholar]
  • 21.Al-Aamri A, Taha K, Al-Hammadi Y, Maalouf M, Homouz Di. Analyzing a co-occurrence gene-interaction network to identify disease-gene association. BMC Bioinform. 2019;20. 10.1186/s12859-019-2634-7. [DOI] [PMC free article] [PubMed]
  • 22.Muller H, Acquati F. Topological Properties of Co-Occurrence Networks in Published Gene Expression Signatures. Bioinform Biol Insights. 2008;2. 10.4137/bbi.s518. [DOI] [PMC free article] [PubMed]
  • 23.CsardiG, Nepusz T. The igraph software package for complex network research. Complex syst. 2006;1695. https://igraph.org
  • 24.Mistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, Sonnhammer ELL, et al. Pfam: The protein families database in 2021. Nucleic Acids Res. 2021;49. 10.1093/nar/gkaa913. [DOI] [PMC free article] [PubMed]
  • 25.Terlouw BR, Biermann F, Vromans SPJM, Zamani E, Helfrich EJN, Medema MH. RAIChU: automating the visualisation of natural product biosynthesis. J Cheminform. 2024;16. 10.1186/s13321-024-00898-x. [DOI] [PMC free article] [PubMed]
  • 26.Wu Y, Pang X, Wu Y, Liu X, Zhang X. Enterocins: classification, synthesis, antibacterial mechanisms and food applications. Molecules. 2022. 10.3390/molecules27072258. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Geldart K, Kaznessis YN. Characterization of class IIa bacteriocin resistance in Enterococcus faecium. Antimicrob Agents Chemother. 2017;61. 10.1128/AAC.02033-16. [DOI] [PMC free article] [PubMed]
  • 28.Vimont A, Fernandez B, Hammami R, Ababsa A, Daba H, Fliss I. Bacteriocin-producing Enterococcus faecium LCW 44: A high potential probiotic candidate from raw camel milk. Front Microbiol. 2017;8. 10.3389/fmicb.2017.00865. [DOI] [PMC free article] [PubMed]
  • 29.Javed A, Masud T, Ul Ain Q, Imran M, Maqsood S. Enterocins of Enterococcus faecium, emerging natural food preservatives. Ann Microbiol. 2011. 10.1007/s13213-011-0223-8. [Google Scholar]
  • 30.Almeida-Santos AC, Novais C, Peixe L, Freitas AR. Enterococcus spp. as a producer and target of bacteriocins: a double-edged sword in the antimicrobial resistance crisis context. Antibiotics. 2021. 10.3390/antibiotics10101215. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Tedim AP, Almeida-Santos AC, Lanza VF, Novais C, Coque TM, Freitas AR, et al. Bacteriocin distribution patterns in Enterococcus faecium and Enterococcus lactis: bioinformatic analysis using a tailored genomics framework. Appl Environ Microbiol. 2024;90:e0137624. 10.1128/aem.01376-24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Himeno K, Rosengren KJ, Inoue T, Perez RH, Colgrave ML, Lee HS, et al. Identification, Characterization, and Three-Dimensional Structure of the Novel Circular Bacteriocin, Enterocin NKR-5-3B, from Enterococcus faecium. Biochemistry. 2015;54. 10.1021/acs.biochem.5b00196S. [DOI] [PubMed]
  • 33.Gilmore MS. Enterococcus faecalis cytolysin and lactocin S of Lactobacillus sake. Antonie van Leeuwenhoek, Int J Gen Mol Microbiol. 1996. 10.1007/BF00399418. [DOI] [PubMed]
  • 34.Biernaum G. Genetics of Lantibiotic Biosynthesis. Comprehensive Natural Products Chemistry. 1999; 275–304. 10.1016/B978-0-08-091283-7.00096-5.
  • 35.Havarstein LS, Diep DB, Nes IF. A family of bacteriocin ABC transporters carry out proteolytic processing of their substrates concomitant with export. Mol Microbiol. 1995;16. 10.1111/j.1365-2958.1995.tb02295.x. [DOI] [PubMed]
  • 36.Donia MS, Cimermancic P, Schulze CJ, Wieland Brown LC, Martin J, Mitreva M, et al. A systematic analysis of biosynthetic gene clusters in the human microbiome reveals a common family of antibiotics. Cell. 2014;158. 10.1016/j.cell.2014.08.032. [DOI] [PMC free article] [PubMed]
  • 37.Zotchev SB. Inter-species horizontal transfer of biosynthetic gene clusters: an evolutionary driver for chemical diversity in bacterial communities. Essays Biochem. 2026. 10.1042/EBC20250014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Klupt S, Fam KT, Zhang X, Chodisetti PK, Mehmood A, Boyd T, et al. Secreted antigen a peptidoglycan hydrolase is essential for enterococcus faecium cell separation and priming of immune checkpoint inhibitor therapy. Elife. 2024;13. 10.7554/eLife.95297. [DOI] [PMC free article] [PubMed]
  • 39.Hallenbeck M, Chua M, Collins J. The role of the universal sugar transport system components PtsI (EI) and PtsH (HPr) in Enterococcus faecium. FEMS Microbes. 2024;5. 10.1093/femsmc/xtae018. [DOI] [PMC free article] [PubMed]
  • 40.Moreno-Gámez S, Hochberg ME, van Doorn GS. Quorum sensing as a mechanism to harness the wisdom of the crowds. Nat Commun. 2023;14. 10.1038/s41467-023-37950-7. [DOI] [PMC free article] [PubMed]
  • 41.Sharma A, Singh P, Sarmah BK, Nandi SP. Quorum sensing: its role in microbial social networking. Res Microbiol. 2020. 10.1016/j.resmic.2020.06.003. [DOI] [PubMed] [Google Scholar]
  • 42.Ji G, Pei W, Zhang L, Qiu R, Lin J, Benito Y, et al. Staphylococcus intermedius produces a functional agr autoinducing peptide containing a cyclic lactone. J Bacteriol. 2005;187. 10.1128/JB.187.9.3139-3150.2005. [DOI] [PMC free article] [PubMed]
  • 43.Mull RW, Harrington A, Sanchez LA, Tal-Gan Y. Cyclic Peptides that Govern Signal Transduction Pathways: From Prokaryotes to Multi-Cellular Organisms. Curr Top Med Chem. 2018;18. 10.2174/1568026618666180518090705. [DOI] [PMC free article] [PubMed]
  • 44.Lanka S, Katta A, Kovvali M, Pandrangi S. Enterococcus faecium Virulence Factors and Biofilm Components: Synthesis, Structure, Function, and Inhibitors. ESKAPE Pathogens. Singapore: Springer Nature Singapore; 2024. pp. 209–226. 10.1007/978-981-99-8799-3_7.
  • 45.Ali IAA, Lévesque CM, Neelakantan P. Fsr quorum sensing system modulates the temporal development of Enterococcus faecalis biofilm matrix. Mol Oral Microbiol. 2022;37. 10.1111/omi.12357. [DOI] [PubMed]
  • 46.Pereira CS, Thompson JA, Xavier KB. AI-2-mediated signalling in bacteria. FEMS Microbiol Rev. 2013. 10.1111/j.1574-6976.2012.00345.x. [DOI] [PubMed] [Google Scholar]
  • 47.Rodriguez-Sanchez AC, Gónzalez-Salazar LA, Rodriguez-Orduña L, Cumsille Á, Undabarrena A, Camara B, et al. Phylogenetic classification of natural product biosynthetic gene clusters based on regulatory mechanisms. Front Microbiol. 2023;14. 10.3389/fmicb.2023.1290473. [DOI] [PMC free article] [PubMed]
  • 48.Noens EEE, Lolkema JS. Convergent evolution of the arginine deiminase pathway: the ArcD and ArcE arginine/ornithine exchangers. Microbiologyopen. 2017;6. 10.1002/mbo3.412. [DOI] [PMC free article] [PubMed]
  • 49.Khademian M, Imlay JA. Do reactive oxygen species or does oxygen itself confer obligate anaerobiosis? The case of Bacteroides thetaiotaomicron. Mol Microbiol. 2020;114. 10.1111/mmi.14516. [DOI] [PMC free article] [PubMed]
  • 50.Furdui C, Ragsdale SW. The role of pyruvate ferredoxin oxidoreductase in pyruvate synthesis during autotrophic growth by the Wood-Ljungdahl pathway. J Biol Chem. 2000;275. 10.1074/jbc.M003291200. [DOI] [PubMed]
  • 51.Amador-Noguez D, Feng XJ, Fan J, Roquet N, Rabitz H, Rabinowitz JD. Systems-level metabolic flux profiling elucidates a complete, bifurcated tricarboxylic acid cycle in Clostridium acetobutylicum. J Bacteriol. 2010;192. 10.1128/JB.00490-10. [DOI] [PMC free article] [PubMed]
  • 52.Doi Y, Ikegami Y. Pyruvate formate-lyase is essential for fumarate-independent anaerobic glycerol utilization in the Enterococcus faecalis strain W11. J Bacteriol. 2014;196. 10.1128/JB.01512-14. [DOI] [PMC free article] [PubMed]
  • 53.Esteban-Torres M, Santamaría L, Cabrera-Rubio R, Plaza-Vinuesa L, Crispie F, de las Rivas B, et al. A diverse range of human gut bacteria have the potential to metabolize the dietary component gallic acid. Appl Environ Microbiol. 2018;84. 10.1128/AEM.01558-18. [DOI] [PMC free article] [PubMed]
  • 54.Rahman MM, Siddique N, Hasnat S, Rahman MT, Rahman M, Alam M, et al. Genomic insights into the probiotic potential and genes linked to gallic acid metabolism in Pediococcus pentosaceus MBBL6 isolated from healthy cow milk. PLoS One. 2024;19. 10.1371/journal.pone.0316270. [DOI] [PMC free article] [PubMed]
  • 55.Soltani S, Arshadi M, Getso MI, Aminharati F, Mahmoudi M, Pourmand MR. Prevalence of virulence genes and their association with biofilm formation in vre faecium isolates from Ahvaz, Iran. J Infect Dev Ctries. 2018;12. 10.3855/jidc.10078. [DOI] [PubMed]
  • 56.Sillanpää J, Nallapareddy SR, Houston J, Ganesh VK, Bourgkogne A, Singh KV, et al. A family of fibrinogen-binding MSCRAMMs from Enterococcus faecalis. Microbiology (N Y). 2009;155. 10.1099/mic.0.027821-0. [DOI] [PMC free article] [PubMed]
  • 57.Hendrickx APA, Van Luit-Asbroek M, Schapendonk CME, Van Wamel WJB, Braat JC, Wijnands LM, et al. SgrA, a nidogen-binding LPXTG surface adhesin implicated in biofilm formation, and EcbA, a collagen binding MSCRAMM, are two novel adhesins of hospital-acquired Enterococcus faecium. Infect Immun. 2009;77. 10.1128/IAI.00275-09. [DOI] [PMC free article] [PubMed]
  • 58.Costa Y, Galimand M, Leclercq R, Duval J, Courvalin P. Characterization of the chromosomal aac(6’)-Ii gene specific for Enterococcus faecium. Antimicrob Agents Chemother. 1993;37. 10.1128/AAC.37.9.1896. [DOI] [PMC free article] [PubMed]
  • 59.Werner G, Hildebrandt B, Witte W, Murray BE, Singh KV. The newly described msrC gene is not equally distributed among all isolates of Enterococcus faecium [1] (multiple letters). Antimicrob Agents Chemother. 2001. 10.1128/AAC.45.12.3672-3673.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Singh KV, Malathum K, Murray BE. Disruption of an Enterococcus faecium species-specific gene, a homologue of acquired macrolide resistance genes of staphylococci, is associated with an increase in macrolide susceptibility. Antimicrob Agents Chemother. 2001;45. 10.1128/AAC.45.1.263-266.2001. [DOI] [PMC free article] [PubMed]
  • 61.He Q, Hou Q, Wang Y, Li J, Li W, Kwok LY, et al. Comparative genomic analysis of Enterococcus faecalis: Insights into their environmental adaptations. BMC Genomics. 2018;19. 10.1186/s12864-018-4887-3. [DOI] [PMC free article] [PubMed]
  • 62.Marshall CG, Broadhead G, Leskiw BK, Wright GD. D-Ala-D-Ala ligases from glycopeptide antibiotic-producing organisms are highly homologous to the enterococcal vancomycin-resistance ligases VanA and VanB. Proc Natl Acad Sci U S A. 1997;94. 10.1073/pnas.94.12.6480. [DOI] [PMC free article] [PubMed]
  • 63.Šeputiene V, Bogdaite A, Ružauskas M, Sužiedeliene E. Antibiotic resistance genes and virulence factors in Enterococcus faecium and Enterococcus faecalis from diseased farm animals: Pigs, cattle and poultry. Pol J Vet Sci. 2012;15. 10.2478/v10181-012-0067-6. [PubMed]
  • 64.Xu X, Lin D, Yan G, Ye X, Wu S, Guo Y, et al. vanM, a new glycopeptide resistance gene cluster found in Enterococcus faecium. Antimicrob Agents Chemother. 2010;54. 10.1128/AAC.01710-09. [DOI] [PMC free article] [PubMed]
  • 65.Zheng B, Tomita H, Inoue T, Ike Y. Isolation of VanB-type Enterococcus faecalis strains from nosocomial infections: First report of the isolation and identification of the pheromone-responsive plasmids pMG2200, encoding VanB-type vancomycin resistance and a Bac41-type bacteriocin, and pMG2201, encoding erythromycin resistance and cytolysin (Hly/Bac). Antimicrob Agents Chemother. 2009;53. 10.1128/AAC.00754-08. [DOI] [PMC free article] [PubMed]
  • 66.Huang YC, Chen FJ, Huang IW, Wu HC, Kuo SC, Huang TW, et al. Clonal expansion of Tn1546-like transposon-carrying vancomycin-resistant Enterococcus faecium, a nationwide study in Taiwan, 2004-2018. J Glob Antimicrob Resist. 2024;39. 10.1016/j.jgar.2024.06.005. [DOI] [PubMed]
  • 67.Partridge SR, Kwong SM, Firth N, Jensen SO. Mobile genetic elements associated with antimicrobial resistance. Clin Microbiol Rev. 2018. 10.1128/CMR.00088-17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Starikova I, Al-Haroni M, Werner G, Roberts AP, Sørum V, Nielsen KM, et al. Fitness costs of various mobile genetic elements in enterococcus faecium and enterococcus faecalis. J Antimicrob Chemother. 2013;68. 10.1093/jac/dkt270. [DOI] [PMC free article] [PubMed]
  • 69.El Zowalaty ME, Lamichhane B, Falgenhauer L, Mowlaboccus S, Zishiri OT, Forsythe S, et al. Antimicrobial resistance and whole genome sequencing of novel sequence types of Enterococcus faecalis, Enterococcus faecium, and Enterococcus durans isolated from livestock. Sci Rep. 2023;13. 10.1038/s41598-023-42838-z. [DOI] [PMC free article] [PubMed]
  • 70.Wang G, Kamalakaran S, Dhand A, Huang W, Ojaimi C, Zhuge J, et al. Identification of a novel clone, ST736, among Enterococcus faecium clinical isolates and its association with daptomycin nonsusceptibility. Antimicrob Agents Chemother. 2014;58. 10.1128/AAC.02683-14. [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

12864_2026_12884_MOESM1_ESM.pptx (7.2MB, pptx)

Supplementary Material 1: Figure S1. Gene cluster architecture and alignment of the selected BGCs in E. faecium genomes for: (A) cyclic lactone auto-inducer BGCs and (B) RiPP-like BGCs. Figure S2. Hierarchical clustering of E. faecium genomes based on the detected BGCs, virulence, and AMR genes. Figure S3. Gene co-occurrence network with communities of genes for BGCs, AMR and virulence from E. faecium genomes.

12864_2026_12884_MOESM2_ESM.pptx (498.8KB, pptx)

Supplementary Material 2: Table S1. Assembly data for the samples. Table S2. PRISM results for each isolate.

Data Availability Statement

All the data that is related to the study are in the figures and tables, both in the main and supplementary files. The genomes included in the study were retrieved from Genome Taxonomy Database, GTDB (https://gtdb.ecogenomic.org). All the genomes used in this study have their accession numbers listed in Supplementary Table S1, in which the first column has the Genbank accession numbers of the assemblies, and the seventh column has the biosample accession number.


Articles from BMC Genomics are provided here courtesy of BMC

RESOURCES