Abstract
Borosins are ribosomally synthesized and post-translationally modified peptides (RiPPs) containing backbone α-N-methylations. These modifications confer favorable pharmacokinetic properties including increased membrane permeability and resistance to proteolytic degradation. Previous studies have biochemically and bioinformatically explored several borosins revealing (1) numerous domain architectures and (2) diverse core regions lacking conserved sequence elements. Due to these characteristics, large-scale computational identification of borosin biosynthetic genes remains challenging and often requires additional, time-intensive manual inspection. This work builds upon previous findings and updates the genome-mining tool RODEO to automatically evaluate borosin biosynthetic gene clusters (BGCs) and identify putative precursor peptides. Using the new RODEO module, we provide an updated analysis of borosin BGCs identified in the NCBI database. From our dataset, we bioinformatically predict and experimentally characterize a new fused borosin domain architecture, in which the modified natural product core is encoded N-terminal to the methyltransferase domain. Additionally, we demonstrate that a borosin precursor peptide is a native substrate of shewasin A, a reported aspartyl peptidase with no previously identified substrates. Shewasin A requires post-translational modification of the leader peptide for proteolytic maturation, a feature not previously observed in RiPPs. Overall, this work provides a user-friendly and open-access tool for the analysis of borosin BGCs and we demonstrate its utility to uncover additional biosynthetic strategies within the borosin class of RiPPs.
Introduction
Borosins are a class of ribosomally synthesized and post-translationally modified peptides (RiPPs) characterized by the presence of amide backbone α-N-methylations.1 In RiPP biosynthesis, the precursor peptide typically contains an N-terminal leader region, which is recognized by enzymes that modify the C-terminal core region into the mature RiPP. These modifying enzymes are often encoded in proximity to the precursor peptide, together forming a biosynthetic gene cluster (BGC). The maturation process includes removal of the leader region by a peptidase that may or may not be locally encoded.2 The first discovered borosin, omphalotin A, was isolated from the basidiomycete fungus Omphalotus olearius.1,3 The omphalotin BGC contained a novel RiPP domain architecture in which the core region was fused to the enzyme that modifies it: an α-N-methyltransferase (NMT) domain.1 Since the initial characterization of the omphalotin BGC, bioinformatic and biochemical studies have uncovered additional borosin BGCs in fungi, archaea, and bacteria4–6 with a diverse set of BGC and NMT domain architectures. Among the domain architectures are three distinct types of “fused borosins,” such as the type found in O. olearius, and seven types of “split borosins” that encode the precursor peptide and NMT as separate genes. To date, >20 borosin NMTs have been biochemically characterized.1,4–7 However, other steps of borosin maturation remain understudied, including leader peptide cleavage. The only characterized protease involved in borosin biosynthesis is OphP, the macrocyclase from the omphalotin BGC.8,9
Computational identification of borosin BGCs is challenging for numerous reasons. Borosin NMTs are homologs of the large tetrapyrrole (TP) methylase family (PF00590, >175,000 members) and are difficult to differentiate from homologs not involved in borosin biosynthesis. Furthermore, identification of borosin precursor peptides presents additional nontrivial obstacles. Whereas RiPP precursor peptides are typically <100 amino acids (AAs), borosin precursor peptides range from ~60 to ~600 AAs, and are not always discretely encoded genes.5,6 Additionally, the α-N-methylation of the peptide backbone appears indiscriminate, as borosin core sequences often lack defining characteristics. Borosins may be enriched in aliphatic, polar, or acidic residues, with methylation occurring on any of these residue types. Borosins also vary in the number of methylations that occur; some sequences contain just two methylation sites while others contain >40, as in the case of SliA from Spirosoma liguale DSM 74.5 When identifying precursor peptides, previous studies have relied on manual inspection of open reading frames (ORFs)4,5 or constructing sequence similarity networks (SSNs) of short ORFs found near the putative borosin NMTs.6 These methods are time-intensive and infeasible for large-scale bioinformatic analysis. Thus, we sought to develop a streamlined approach to analyzing and assessing potential borosin BGCs.
RODEO (Rapid ORF Description & Evaluation Online) is a computational tool used to analyze genes and short ORFs neighboring an input protein query.10 Several specialized modules have been developed to evaluate and identify putative precursor peptides for various RiPP classes.10–15 In this work, we have updated RODEO with the capability to automatically identify putative borosin precursor peptides using heuristic scoring, hidden Markov models (HMMs), motif analysis, and support vector machine classification. Using the new borosin module, we have collected and analyzed all detectable borosin BGCs found in the NCBI non-redundant protein database. During the analysis, we assigned precursor peptides to borosin NMTs that previously were lacking a known substrate (i.e., orphan NMTs). Additionally, we identified and biochemically verified a new type of fused borosin wherein the core region is found N-terminal to the NMT domain, which we have designated as type 0 borosins. Finally, we have demonstrated that a borosin precursor peptide is the native substrate for shewasin A, a previously characterized pepsin-like aspartyl peptidase from Shewanella amazonensis. Shewasin A requires an α-N-methylated leader peptide to release a core peptide devoid of α-N-methylations, a paradigm not previously observed in RiPP biosynthesis.
Results & Discussion
Development of a borosin RODEO scoring module.
We first sought a method to robustly differentiate borosin NMTs from non-borosin methylases within the tetrapyrrole (TP) methylase family. Previous work relied on a combination of custom HMMs to identify putative borosin NMTs, but many of the sequences retrieved using the individual custom HMMs were manually identified as false positives.5 BLASTP searches were performed against the NCBI non-redundant (nr) protein database using each characterized borosin NMT as a query (Table S1). Retrieved sequences were analyzed using hmmscan16 and a subset of the results was plotted based on the expectation value (e-value) for the TP methylases (PF00590) and custom BorosinMT HMMs (Fig. S1). We then inspected and assigned sequences as either borosin or non-borosin methyltransferases (SI Methods). The graph comparing BorosinMT and PF00590 e-values showed two distinct trends that clearly separated borosin from non-borosin sequences, except at very poor e-values.
Next, we developed a set of heuristics to identify borosin precursor peptides based on accumulated knowledge of borosin BGCs (Table S2). Most identified borosin BGCs contain a conserved five-helical bundle, termed the Borosin Binding Domain (BBD), that is crucial for precursor peptide recognition by the NMT.7 Fused borosins (types I-III) contain a BBD and core sequence fused to an NMT domain. Some split borosin BGCs contain a BBD-NMT fusion (types V-VIII), while others contain a BBD-precursor peptide fusion (types IV, IX, and X) (Fig. 1). Several groups of split borosins contain other conserved sequences within the BGC. These include several distinct sequence motifs found in a subset of borosin BGCs from Legionellales, which lack a BBD6 and a previously unreported conserved region often found in split borosin BGCs with a BBD-methyltransferase fusion (types VI, VII, and VIII). We constructed custom HMMs using these conserved features to aid precursor peptide identification (Supplemental Dataset 1). The presence of these sequences, and in some cases the e-value similarity to HMMs, were included in the heuristic scoring. Additionally, borosin core regions often contain repeating motifs or are enriched in certain amino acids at their C-terminus. Repeating motifs from characterized borosins were identified using the MEME suite17 and were incorporated into the heuristic scoring alongside the enrichment of certain amino acids in the precursor peptide.
Figure 1: Workflow for bioinformatically mining borosin BGCs.

Workflow used to create a comprehensive dataset of all detected borosin BGCs. α-N-Methyltransferase (NMT) homologs are identified by BLASTP and filtered by e-value to the BorosinMT HMM to reduce computation time. Custom HMM matches in UniProtKB are identified using hmmsearch. After duplicate removal, the list of unique NCBI protein accession identifiers were analyzed by RODEO with borosin scoring enabled. The RODEO-predicted precursor peptides were then analyzed and assigned to different borosin types. BBD, borosin binding domain. TPR, tetratricopeptide repeat. GGDEF, GGDEF-sequence-containing domain. In the domain architecture diagrams: blue indicates the BBD, yellow indicates borosin core regions, and green indicates a transmembrane domain.
After defining the heuristic-, motif-, and HMM-based scoring metrics, we tested the borosin module on a set of high-confidence predicted borosin BGCs, but quickly learned this alone was insufficient to accurately identify borosin precursor peptides (Fig. S2–S3, SI methods). Therefore, we trained a support vector machine (SVM) classifier to identify ORFs as borosin or non-borosin using local attributes extracted from the genomic sequence (Table S3). Features used include gene co-occurrence, amino acid properties, and many of the heuristic scoring features. The addition of SVM classification significantly increased module accuracy and effectively eliminated BorosinMT false positives. To decrease computation time, methyltransferases with an e-value higher than E-5 to the BorosinMT HMM were filtered prior to RODEO analysis (Fig. 1).
Bioinformatics analysis of the borosin BGC dataset
A dataset of borosin precursor peptides was curated using the newly devised RODEO module (SI Methods). This dataset contains 2699 predicted borosin NMTs with 3064 predicted precursor peptides across 2596 BGCs (Supplemental Dataset 2). While the RODEO module correctly matched 1565 precursor peptides for previously reported (putative) borosin NMTs, the module found an additional 1127 precursor peptides that had not been associated with borosin biosynthesis. Precursor peptides were not found for 733 orphan NMTs identified in previous bioinformatic studies,5,6 even after analysis using the new RODEO module. Over 50% of these NMTs i) could not be retrieved, ii) were part of fragmented sequencing reads, or iii) were assigned as non-borosin TP methylases based on manual inspection and the previously described e-value cutoffs. The remaining NMTs had a potential borosin precursor peptide in the genomic neighborhood, but further efforts are needed to validate if these are borosin BGCs. These BGCs may represent new borosin architectures, those with distally encoded precursor peptides, evolutionary relics that no longer have a precursor peptide, or the TP methylase homolog carries out some other biological role. Following identification, we assigned predicted borosin precursor peptides to their respective borosin types using similarity to HMMs, gene co-occurrence, and sequence information.5 An SSN was generated from NMTs in the RODEO-mined borosin dataset (Fig. 2, S4). Many of the largest SSN groups contain biochemically characterized representatives, with groups 7 and 8 being the exceptions. SSN group 3 is comprised of NMT-BBD fusions from borosin types I-III and VI-VIII. Group 3 is taxonomically more widely disseminated, appearing in bacteria and fungi, relative to groups that discretely encode the NMT and BBD. Several unusual borosin BGC architectures were also identified, including BBD-GNAT acetyltransferase fusions and a BGC with 11 BBD-fused precursor peptides (Fig. S5).
Figure 2: SSN of borosin methyltransferases.

Borosin NMTs with a RODEO-identified precursor peptide were used to generate an SSN (n = 2124). Sequences for OphMA, PgiMA and AboMA are absent in the NCBI protein database and were thus manually added. SSN is viewed at an alignment score of 90.
We next investigated the sequence relationships of borosin NMTs and BBDs. A phylogenetic tree was created using 469 diversity-maximized NMTs from the RODEO-enabled borosin BGC dataset, including members from each borosin type (SI Methods, Fig. S6–S7). A custom script14 excised the region scoring highest for a BBD HMM from either the borosin NMT or split precursor peptide in the BGCs (Fig. S8–S9). A second tree was generated from these sequences, and the sum of the branch distances for the NMT and BBD trees were analyzed (Fig. S10). This analysis found that characterized borosin NMTs are more similar to each other than to the PF00590 TP methylase outgroup, indicating that borosins likely evolved once. Excised BBDs do not follow this trend, with characterized BBDs showing similar relatedness between each other and the outgroup. This may result from conservation of secondary structure within the BBD rather than specific sequence motifs for precursor peptide recognition. Similar examples of this exist in RiPP biosynthesis, most notably with RiPP precursor recognition elements found in many prokaryotic RiPP BGCs.18
The borosin NMT from Streptomyces sp. NRRL S-118 (SspM; WP_158827804.1) and the associated BBD are the characterized split borosin pair most closely related to the fused borosins and outgroups. To get a more accurate root and branching order of borosin NMT evolution, SspM homologs in the NCBI nr database were identified using BLASTP, and a larger phylogenetic tree was generated (Fig. S11–S12). Analysis of the resulting branch distances revealed several groups of RODEO-predicted type IV borosin NMTs positioned between the outgroups and characterized borosins (Fig. S13). As all characterized borosins are more similar to the predicted type IV borosins than the outgroups, we speculate that borosins originated from pathways encoding a discretely encoded NMT that acted upon a BBD-fused precursor peptide. The borosin NMTs from Legionellales lacking a BBD are closely related to other borosin NMTs, with distances comparable to other split borosin NMTs. This presumably indicates that the Legionellales BGCs originated from BBD-containing borosins that later lost the BBD.
While rare examples exist of fused borosins in bacteria, split borosins have not been identified in fungi. To search for evidence of fungi acquiring borosin BGCs through horizontal gene transfer (HGT), we compared the G+C content of the borosin NMT gene to the G+C content of the entire contig that encodes the NMT. No evidence for HGT was observed for fungal borosins from this analysis. However, several HGT candidates were found in bacteria that would involve transfer from a lower G+C organism to one with a higher G+C content (Fig. S14). Additionally, SSN groups 2 and 4, which are dominated by Shewanella and Legionella, respectively, contain many split borosin BGCs with nearby phage integrases. These mobile genetic elements may have enabled distribution of borosins within these genera (Fig. S15). Overall, borosins appear to have originated with a split BBD-fused precursor peptide and diverged into two groups: a Pseudomonadota-dominated group with BBD-precursor peptides fusions and a taxonomically diverse group with NMT-BBD fusions.
Discovery and characterization of type 0 borosins
Upon reviewing the excised BBDs from the RODEO-enabled analysis above, we discovered a new group of fused borosins designated as type 0. Unlike other fused borosins (types I-III), where the BBD and core region are encoded C-terminal to the NMT, the BBD and putative core of type 0 borosins are located at the N-terminus (Fig. 3). RODEO identified 32 putative type 0 borosins based on the relative HMM alignment positions in BBD-fused NMTs (Fig. S4). Two examples, FpeAM from Flavihumibacter petaseus NBRC 10605419 and SkoAM from Segetibacter koreensis DSM 18137,20 were selected for experimental validation. To confirm NMT activity and core sequences, we performed heterologous expression, purification, and analysis by high-resolution tandem mass spectrometry (HR-MS/MS). Up to 14 methylations in FpeAM and 10 in SkoAM were detected in the predicted core region between the BBD and methyltransferase domains (Fig. 3, S16).
Figure 3: Type 0 borosin architecture and LC-MS/MS verification.

A. Comparison of the overall domain architectures of the type 0 and type I borosin precursors including the NMT (orange), BBD (blue), and core peptide (yellow) domains. B. Methylation pattern of the SkoAM core region determined by LC-MS/MS HCD fragmentation. Detected b+ and y+ ions are labeled above and below the peptide sequence, respectively. Methylated residues are indicated by orange circles. C. MS2 spectrum corresponding to the annotated peptide in panel B, peaks corresponding to detected b+ and y+ ions are labeled. A mass cutoff of 10.0 ppm and minimum relative peak intensity of 0.1% was used for mass annotation.
All currently published crystal structures of borosin NMTs show a dimeric conformation that facilitates an intermolecular mechanism in which one subunit modifies the core peptide of the other.7,21,22 In contrast, SkoAM was determined to be monomeric by size-exclusion chromatography (Fig. S17). This observation raises the possibility that type 0 borosins may catalyze intramolecular α-N-methylation, where each NMT modifies its own core peptide in cis rather than in trans. AlphaFold 2.023 structural predictions of both monomeric and homodimeric SkoAM support the proposal of intramolecular catalysis based on the relative orientation of the BBD and NMT domains (Fig. S17). Unfortunately, efforts to confirm the predicted monomeric stoichiometry of SkoAM suggested by size-exclusion were inconclusive using trans-complementation assays similar to those previously used to establish intermolecular activity for OphMA.1 Residual activity observed from a SkoAM triple active site variant (R214A Y218A Y240A) led to ambiguous results and intramolecular activity of type 0 borosins remains unconfirmed.
Shewasin A is a borosin peptidase
RODEO streamlines the analysis of ORFs neighboring the gene-of-interest. Since proteolytic cleavage is a pivotal step in RiPP biosynthesis,2 we scanned for peptidases that commonly occur in borosin BGCs. Some borosin-encoding organisms provide a local copy a PF00026 pepsin-like aspartyl peptidase within the borosin BGC, while others distally encode such enzymes (Fig. S18). Of the borosin NMTs in the RODEO-curated dataset (n = 2699), an aspartyl peptidase was found locally in 194 cases (7%). PSI-BLAST searches of the NCBI nr database yielded 441 aspartyl peptidases after removing incomplete fragments. Upon RODEO analysis, an additional 73 putative borosin BGCs were identified. Of the remaining aspartyl peptidases, 56 were found in organisms with a predicted borosin BGC, and 118 (including all 84 examples from Eukarya) were not associated with borosins. A phylogenetic tree of the peptidases shows a clear distinction between eukaryotic and bacterial sequences, with those peptidases found in borosin BGCs being less related to the eukaryotic peptidases (Fig. S19). Inspection of the genomic neighborhood around the non-borosin peptidases did not provide evidence of involvement in RiPP biosynthesis.
Pepsin-like aspartyl peptidases are primarily found in eukaryotes, with few identified in bacteria.24,25 One such peptidase from bacteria that has been partially characterized is shewasin A from Shewanella amazonensis. Two studies demonstrated protease activity for shewasin A and shewasin D (an ortholog) against a panel of non-native substrates. However, the native substrates for both remained unknown. Our RODEO-mined borosin dataset revealed that shewasin A and D reside in putative borosin BGCs, indicating a potential role in borosin maturation.
To determine if shewasin A is a borosin peptidase, we first evaluated if the nearby borosin NMT (SamM1) acted upon the predicted BBD-fused precursor peptide (SamA1). Genes samM1 and samA1 were heterologously expressed in E. coli and purified by nickel affinity chromatography. LC-MS/MS analysis of in-gel-trypsinized SamA1 revealed a single α-N-methylation on Phe52 (Fig. S20). Expressions over 24 and 48 h had >99% conversion to the methylated product. Expression with a less active SamM1 R65A Y69A variant resulted in <10% conversion, indicating α-N-methylation is SamM1 dependent (Fig. S21–S23). These results confirm that SamM1 is a borosin NMT and SamA1 is the cognate substrate.
Next, shewasin A was heterologously expressed, purified, and assayed using methylated and unmethylated SamA1 with monitoring by SDS-PAGE (Fig. 4). Efficient cleavage of α-N-methylated SamA1 was observed, while only basal proteolytic activity was detected with unmethylated SamA1. Activity was observed up to pH 6.0, with significant precipitation of the protease at pH 4.2 and 5.0 (Fig. S24). Proteolytic activity was not observed using an inactive D37A variant.26 These results indicate that SamA1 is a native substrate of shewasin A (hereafter SamP), and activity is α-N-methylation and pH dependent. Previous studies reported modest SamP proteolytic activity only at acidic pH between 2.5-5.5.26 In addition, shewasin D was localized to the pH-neutral cytosol in the native organism.27 Our results reconcile these seemingly contradictory observations by demonstrating that, while SamP is still active under acidic conditions, it is more stable and active on its native substrate at near-neutral pH. While most A1 family peptidases are only active at acidic pHs, some are active in near-neutral conditions.28,29 This shift is proposed to be caused by a Thr to Ala substitution near the active site, which is found in SamP (shewasin A) and shewasin D (Fig. S25).28
Figure 4: SamP (shewasin A) is a borosin protease.

A. Putative S. amazonensis borosin BGC. Genes expressed in this study are labeled. The sequence of the precursor peptide SamA1 is shown (bottom) with the putative core sequence bolded and underlined. LC-MS/MS localized α-N-methylated Phe52 is circled in orange, with an arrow denoting the SamP cut site. B. SDS-PAGE of in vitro reactions containing methylated (SamA1/M1) and unmethylated (SamA1/M1 R65A Y69A; SamA1) SamA1 and variants (SamA1Ala-Ala /M1; SamA1Gly-Gly /M1) with and without SamP (+/−).
Finally, LC-MS/MS analysis of in vitro reactions localized SamP cleavage at the amide bond between Phe53 and Pro54 (Fig. S26). Previous work probing SamP substrate recognition found a strong preference for Leu, Phe, Tyr, and Asp at the P1 position as well as bulky, hydrophobic residues at P2.27 Moderate preference for neutral residues at positions P3/P1’/P2’ and Asp at P4’ were also observed. The SamA1 cleavage site is mostly consistent with these results (Fig. 4). Notable exceptions are Pro at the P1’ position and Glu at P3’. Given the strong preference for bulky residues at P2 and P1, substrate variants replacing the Phe-Phe motif in the core peptide with Ala-Ala (SamA1Ala-Ala) or Gly-Gly (SamA1Gly-Gy) were assayed with SamP. A small amount of proteolytic activity was observed for SamA1Ala-Ala, but not for SamA1Gly-Gly (Fig. S24, S26). LC-MS/MS analysis of trypsinized SamA1Ala-Ala and SamA1Gly-Gly revealed >97% and >72% methylation, respectively (Fig. S21–S23). This indicates that α-N-methylation is not solely sufficient for SamP recognition. Further work detailing SamP substrate scope and specificity is ongoing and will be reported elsewhere.
In attempts to identify SamP activity in the native host, LC-MS/MS analysis was performed on extracts from S. amazonensis cultures overexpressing a plasmid-encoding copy of the sam BGC. The peptide PAEDEPEKEQQENEEEDEAPKA was detected in cell extracts, corresponding to the C-terminus of SamA1 sans three C-terminal residues, presumably removed by a host protease (Fig. S27). These results support our SamP in vitro results and confirm α-N-methylation is not found in the proteolytically released C-terminal core peptide. Of note, a putative acetyltransferase (PF13302) is also encoded in the sam BGC, suggesting additional modifications may be present in the final natural product (Fig. 4). Acetylated products were not detected in extracts (Fig. S27). Future work will delve further into the biological activity and final structure of the S. amazonensis borosin metabolite, since studies have shown that anionic peptides can have specific and potent antimicrobial activities30,31 Nonetheless, our data redefines the activity profile and substrate scope of rare bacterially encoded pepsins, now revealed as borosin pathway maturases. Furthermore, our work reveals a previously unseen role for post-translational modifications in RiPP biosynthesis, where α-N-methylation on the leader peptide acts as a recognition element for proteolytic release rather than as a modification of the natural product.
Conclusions
Borosins are an expansive RiPP class with diverse core sequences and biosynthetic protein domain architectures. Several factors, including the extent and variability of the class-defining backbone α-N-methylation, have previously necessitated time-consuming manual inspection to identify putative borosin precursor peptides. Our newly developed borosin RODEO module has enabled a streamlined approach for borosin BGC analysis, including precursor peptide identification. This large-scale identification allowed for evolutionary analysis of borosin BGCs that suggest borosin methyltransferases evolved once, then diverged into two main groups. The divide between the Pseudomonadota-dominated group containing BBD-fused precursor peptides and the taxonomically diverse group containing BBD-fused methyltransferase may correlate with differences in biological function; further studies on the bioactivities of the peptide products may shed light on the benefit and pattern of propagation of borosin BGCs. Additionally, analysis of >2500 RODEO-mined borosin BGCs revealed a new type of fused borosin with a core sequence N-terminal to the methyltransferase domain. Finally, we demonstrated that the previously studied pepsin-like aspartic protease SamP (shewasin A) is a borosin maturase, marking SamP as only the second α-N-methylation–dependent protease currently known. In vitro and mutagenesis experiments revealed that SamP preferentially cleaves the α-N-methylated precursor peptide SamA1 to yield a highly anionic, unmethylated peptide. While ongoing work will be necessary to confirm the final natural product and details of SamP specificity, these data report the first instance of a post-translational modification requirement on a RiPP leader peptide for natural product maturation. All in all, our newly developed RODEO module provides a new large-scale computational tool for studying borosin BGCs and demonstrates its utility to facilitate targeted enzyme and natural product discoveries.
Methods
Bioinformatic mining and analysis of borosin biosynthetic gene clusters (BGCs)
Borosin module.
To generate a list of potential borosin BGCs, we used characterized borosin N-methyltransferase (NMT) sequences as queries for protein BLAST searches against the NCBI non-redundant (nr) database (Table S1). High confidence borosin BGCs were assigned by manually inspecting BGCs for precursor peptides and protein architectures similar to those already known. The methyltransferase and precursor peptide sequences were analyzed, and HMMER332 was used to create custom profile Hidden Markov Models (HMMs) (Supplementary Dataset 1) to identify more distant homologs. High confidence precursor peptides were analyzed for motifs using XSTREME from the MEME suite.17 These features were used alongside amino acid content for heuristic scoring and support vector machine (SVM) classification.
The SVM was trained and tested with 414 positives and 7838 negatives using a 5-fold cross validation. The positive dataset includes characterized and high confidence predicted borosin precursor peptides. The negative dataset includes hypothetical non-borosin open reading frames (ORFs) from within borosin and non-borosin RiPP BGCs. The borosin scoring module accuracy was assessed using the following equations (Fig. S2, Fig. S3):
| Equation S1: |
| Equation S2: |
| Equation S3: |
Bioinformatic analysis of borosin BGCs.
To generate a list of potential borosin BGCs for RODEO to analyze, we used characterized borosin NMT sequences as queries for BLASTP searches against the NCBI non-redundant (nr) database (Table S3) keeping up to 10,000 hits for each query. Additionally, hmmsearch32 was used to query the UniProtKB database for each of the borosin NMT, borosin binding domain (BBD), and precursor HMMs; these sequences were mapped to their corresponding NCBI accession identifiers. After removing duplicates, the compiled sequences were analyzed by RODEO using the borosin scoring module. BGCs with multiple queried accessions were dereplicated and precursor peptides with a cutoff score of 16 or higher were considered valid and included in the final borosin BGC dataset.
An analysis of the phylogenetic distribution of borosin methyltransferases with a RODEO predicted precursor peptide was performed. All sequence similarity networks (SSNs) were generated using the Enzyme Function Initiative Enzyme Similarity Tool (EFI-EST) (https://efi.igb.illinois.edu/).33 SSN visualization was performed using organic layout within Cytoscape.34 Multiple sequence alignments were performed using MUSCLE,35 and phylogenetic trees were built using the FastTree36 with the default Jones-Taylor-Thornton model. Trees were visualized using the interactive Tree of Life (iTOL) website (http://itol.embl.de/).37 The sum of branch lengths were compared using MEGA 11.38 The protein sequence of OphMA, PgiMA1, and AboMA are not available in UniProtKB or NCBI and thus were added manually.
In vitro assays of SamA1 and SamP
Purified 6His-SamA1 (with or without SamM1 in complex) was added to a tube with SamP at ratios ranging from 1:10 – 1:1000 (enzyme:substrate). Reactions were performed in 50-100 mM sodium citrate pH 5.0, 2-(N-morpholino)ethanesulfonic acid (MES) pH 6.0, HEPES pH 7.0 or HEPES pH 8.0 with 100 mM KCl, and with or without 8% dimethyl sulfoxide (DMSO). Subsequent assays used 50 mM MES pH 6.0 once the optimal conditions were identified. For SDS-PAGE analysis, reactions were quenched by adding 1/4 volume 5X SDS-PAGE loading dye (250 mM TrisCl pH 6.8; 5% (v/v) β-mercaptoethanol; 10% (w/v) SDS; 0.25% (w/v) bromophenol blue; 50% (v/v) glycerol) and boiling for 5 minutes before running on SDS-PAGE. For LC-MS/MS analysis, reactions were quenched by adding 1 M ammonium bicarbonate solution, pH 8.0 to a final concentration of 100 mM. Samples were treated with iodoacetamide (described in Supporting Information), desalted using C18 ZipTips, then run on LC-MS/MS (described below).
Peptide mass spectrometric analysis
LC-MS/MS analysis of digested peptides was performed as described previously.5 Briefly, data were recorded on a Thermo Scientific Fusion or Fusion Lumos mass spectrometer equipped with a Dionex Ultimate 3000 UHPLC system using a nLC column (200 mm × 75 μm) packed using Vydac 5-μm particles with a 300 Å pore size (Hichrom Limited). Elution was performed with a linear gradient using water with 0.1% FA (solvent A) and ACN with 0.1% FA (solvent B) at a flow rate of 0.3 μL min−1. The column was equilibrated with 20% solvent B for 5 min, followed by a linear increase of solvent B to 95% over 32 min and a final elution step with 95% solvent B for 2 min. Mass spectra were acquired in positive ion mode. Full MS was done at a resolution of 60,000 [automatic gain control (AGC) target, 4 × 105; maximum injection time (IT), 50–100 ms; range, 300–1800 m/z], and data-dependent as well as targeted MS/MS was performed at a resolution of 15,000 (AGC target, 5 × 105; maximum IT, 100–500 ms; isolation window, 2.2 m/z) using higher-energy collisional dissociation (HCD) or electron-transfer dissociation (ETD). Data were processed using Thermo Fisher Xcalibur software and MaxQuant v1.6.10.145. Methylation state abundance for trypsinized SamA1 variants were calculated by summing the peak areas (as ion counts) from extracted ion chromatograms (EICs) of the theoretical monoisotopic 2+ and 3+ parent masses +/− 10 millimass unit (mmu) for each methylation state, dividing by the total peak areas of all methylation states for that variant, and multiplied by 100. Percentages were averaged and standard deviation was calculated from three technical replicates for each variant.
Sam BGC overexpression and metabolite extraction from S. amazonensis SB2B
The sam BGC was cloned into the pBBR1MCS5 expression vector using methods described above then transformed into the diaminopimelic acid (DAP) auxotrophic E. coli WM3064 strain. The plasmid was then conjugated into the S. amazonensis SB2B host strain by mating the two strains on DAP-supplemented LB plates followed by selection on DAP-deficient LB supplemented with gentamicin. For expression, 1:100 volume of an overnight culture was inoculated into TB + gentamicin and grown at 30 °C shaking at 170 rpm until OD600 reached 0.4–0.6. Expression was induced with 0.2 mM IPTG and incubated at 30 °C shaking at 170 rpm for 24 h. Cells were harvested by centrifugation at 4000 × g for 20 min at 4 °C then flash-frozen in liquid nitrogen and stored at −80 °C until extraction. For extraction, the biomass was resuspended in 80% methanol and sonicated on ice for 5 min. Debris was removed by centrifugation at 4000 × g for 30 min followed by sequential filtration through Whatman paper and then a 0.22 μm filter. The sample was then dried down in a rotary evaporator then resuspended in water. Sample was then processed through a Pierce™ Graphite Spin Column (Thermo Scientific, Waltham, MA) following manufacturer’s instructions then analyzed by LC-MS/MS using methods described above.
Supplementary Material
Supplemental Dataset 1: Custom HMMs (.txt)
Supplemental Dataset 2: Precursor Dataset (.xlsx)
Supplemental Dataset 3: Phylogenetic trees in Newick format (.nwk)
Supplemental Dataset 4: Primers, genes, plasmids, and protein sequences used in this study (.xlsx)
Acknowledgements
We thank L. Daigh for the preliminary work on the borosin module and thank J. Gralnick for the strain S. amazonensis SB2B.
Funding
This work was supported by the National Institutes of Health (R35 GM133475 for M.F.F., R01 GM123998 for D.A.M., and T32-GM136629 for R.S.C.). S.R.D. is supported by the Diffenbaugh Fellowship.
Footnotes
Additional methods can be found in Supporting Information
Accession Codes:
FpeAM: WP_083990299.1
SkoAM: WP_018611879.1
LepM1: WP_061716904.1
ParM: WP_007623905.1
SspM: WP_031073184.1
SamM: WP_011758902.1
SonM: WP_011071665.1
RceM: WP_012568691.1
SliM: ADB39711.1
AinM: WP_083827882.1
PmoM: WP_051555776.1
SurM1: WP_165337934.1
LulM: WP_149193073.1
The authors declare the following competing financial interest(s): M.F.F. is an inventor on patents WO2017EP58327, US20190112583A1, WO2021168399, and US20230090771A1.
References
- (1).van der Velden NS; Kälin N; Helf MJ; Piel J; Freeman MF; Künzler M Autocatalytic Backbone N-Methylation in a Family of Ribosomal Peptide Natural Products. Nat. Chem. Biol 2017, 13 (8), 833–835. 10.1038/nchembio.2393. [DOI] [PubMed] [Google Scholar]
- (2).Montalbán-López M; Scott TA; Ramesh S; Rahman IR; Heel A. J.van; Viel JH; Bandarian V; Dittmann E; Genilloud O; Goto Y; Burgos MJG; Hill C; Kim S; Koehnke J; Latham JA; Link AJ; Martínez B; Nair SK; Nicolet Y; Rebuffat S; Sahl H-G; Sareen D; Schmidt EW; Schmitt L; Severinov K; Süssmuth RD; Truman AW; Wang H; Weng J-K; Wezel G. P. van; Zhang Q; Zhong J;Piel J; Mitchell DA; Kuipers OP; Donk W. A. van der. New Developments in RiPP Discovery, Enzymology and Engineering. Nat. Prod. Rep 2021, 38 (1), 130–239. 10.1039/D0NP00027B. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (3).Sterner O; Etzel W; Mayer A; Anke H Omphalotin, a New Cyclic Peptide with Potent Nematicidal Activity from Omphalotus Olearius. Nat. Prod. Lett 1997, 10 (1), 33–38. 10.1080/10575639708043692. [DOI] [Google Scholar]
- (4).Quijano MR; Zach C; Miller FS; Lee AR; Imani AS; Künzler M; Freeman MF Distinct Autocatalytic α-N-Methylating Precursors Expand the Borosin RiPP Family of Peptide Natural Products. J. Am. Chem. Soc 2019, 141 (24), 9637–9644. 10.1021/jacs.9b03690. [DOI] [PubMed] [Google Scholar]
- (5).Imani AS; Lee AR; Vishwanathan N; De Waal F; Freeman MF Diverse Protein Architectures and α-N-Methylation Patterns Define Split Borosin RiPP Biosynthetic Gene Clusters. ACS Chem. Biol 2022, 17 (4), 908–917. 10.1021/ACSCHEMBIO.1C01002/SUPPL_FILE/CB1C01002_SI_005.XLSX. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (6).Cho H; Lee H; Hong K; Chung H; Song I; Lee J-S; Kim S Bioinformatic Expansion of Borosins Uncovers Trans-Acting Peptide Backbone N-Methyltransferases in Bacteria. Biochemistry 2022, 61 (3), 183–194. 10.1021/acs.biochem.1c00764. [DOI] [PubMed] [Google Scholar]
- (7).Miller FS; Crone KK; Jensen MR; Shaw S; Harcombe WR; Elias MH; Freeman MF Conformational Rearrangements Enable Iterative Backbone N-Methylation in RiPP Biosynthesis. Nat. Commun 2021, 12 (1), 5355. 10.1038/s41467-021-25575-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (8).Matabaro E; Kaspar H; Dahlin P; Bader DLV; Murar CE; Staubli F; Field CM; Bode JW; Künzler M Identification, Heterologous Production and Bioactivity of Lentinulin A and Dendrothelin A, Two Natural Variants of Backbone N-Methylated Peptide Macrocycle Omphalotin A. Sci. Rep 2021, 11 (1), 3541. 10.1038/s41598-021-83106-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (9).Matabaro E; Witte L; Gherlone F; Vogt E; Kaspar H; Künzler M Promiscuity of Omphalotin A Biosynthetic Enzymes Allows de Novo Production of Non-Natural Multiply Backbone N-Methylated Peptide Macrocycles in Yeast. ChemBioChem 2024, 25 (3), e202300626. 10.1002/cbic.202300626. [DOI] [PubMed] [Google Scholar]
- (10).Tietz JI; Schwalen CJ; Patel PS; Maxson T; Blair PM; Tai H-C; Zakai UI; Mitchell DA A New Genome-Mining Tool Redefines the Lasso Peptide Biosynthetic Landscape. Nat. Chem. Biol 2017, 13 (5), 470–478. 10.1038/nchembio.2319. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (11).Schwalen CJ; Mitchell D Discovery of Antibiotic Peptides from Novelty-Prioritized Natural Product Genome Mining. FASEB J. 2017, 31 (S1), 939.8–939.8. 10.1096/fasebj.31.1_supplement.939.8. [DOI] [Google Scholar]
- (12).Hudson GA; Mitchell DA RiPP Antibiotics: Biosynthesis and Engineering Potential. Curr. Opin. Microbiol 2018, 45, 61–69. 10.1016/j.mib.2018.02.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (13).Walker MC; Eslami SM; Hetrick KJ; Ackenhusen SE; Mitchell DA; van der Donk WA Precursor Peptide-Targeted Mining of More than One Hundred Thousand Genomes Expands the Lanthipeptide Natural Product Family. BMC Genomics 2020, 21 (1), 387. 10.1186/s12864-020-06785-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (14).Georgiou MA; Dommaraju SR; Guo X; Mast DH; Mitchell DA Bioinformatic and Reactivity-Based Discovery of Linaridins. ACS Chem. Biol 2020, 15 (11), 2976–2985. 10.1021/acschembio.0c00620. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (15).Ramesh S; Guo X; DiCaprio AJ; De Lio AM; Harris LA; Kille BL; Pogorelov TV; Mitchell DA Bioinformatics-Guided Expansion and Discovery of Graspetides. ACS Chem. Biol 2021, 16 (12), 2787–2797. 10.1021/acschembio.1c00672. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (16).Eddy SR Accelerated Profile HMM Searches. PLOS Comput. Biol 2011, 7 (10), e1002195. 10.1371/JOURNAL.PCBI.1002195. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (17).Bailey TL; Johnson J; Grant CE; Noble WS The MEME Suite. Nucleic Acids Res 2015, 43, W39–W49. 10.1093/nar/gkv416. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (18).Burkhart BJ; Hudson GA; Dunbar KL; Mitchell DA A Prevalent Peptide-Binding Domain Guides Ribosomal Natural Product Biosynthesis. Nat. Chem. Biol 2015, 11 (8), 564–570. 10.1038/nchembio.1856. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (19).Zhang NN; Qu JH; Yuan HL; Sun YM; Yang JS Flavihumibacter Petaseus Gen. Nov., Sp. Nov., Isolated from Soil of a Subtropical Rainforest. Int. J. Syst. Evol. Microbiol 2010, 60 (Pt 7), 1609–1612. 10.1099/IJS.0.011957-0. [DOI] [PubMed] [Google Scholar]
- (20).An DS; Lee HG; Im WT; Liu QM; Lee ST Segetibacter Koreensis Gen. Nov., Sp. Nov., a Novel Member of the Phylum Bacteroidetes, Isolated from the Soil of a Ginseng Field in South Korea. Int. J. Syst. Evol. Microbiol 2007, 57 (Pt 8), 1828–1833. 10.1099/IJS.0.64803-0. [DOI] [PubMed] [Google Scholar]
- (21).Song H; van der Velden NS; Shiran SL; Bleiziffer P; Zach C; Sieber R; Imani AS; Krausbeck F; Aebi M; Freeman MF; Riniker S; Künzler M; Naismith JH A Molecular Mechanism for the Enzymatic Methylation of Nitrogen Atoms within Peptide Bonds. Sci. Adv 2018, 4 (8). 10.1126/sciadv.aat2720. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (22).Ongpipattanakul C; Nair SK Molecular Basis for Autocatalytic Backbone N-Methylation in RiPP Natural Product Biosynthesis. ACS Chem. Biol 2018, 13 (10), 2989–2999. 10.1021/acschembio.8b00668. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (23).Jumper J; Evans R; Pritzel A; Green T; Figurnov M; Ronneberger O; Tunyasuvunakool K; Bates R; Žídek A; Potapenko A; Bridgland A; Meyer C; Kohl SAA; Ballard AJ; Cowie A; Romera-Paredes B; Nikolov S; Jain R; Adler J; Back T; Petersen S; Reiman D; Clancy E; Zielinski M; Steinegger M; Pacholska M; Berghammer T; Bodenstein S; Silver D; Vinyals O; Senior AW; Kavukcuoglu K; Kohli P; Hassabis D Highly Accurate Protein Structure Prediction with AlphaFold. Nature 2021, 596, 583. 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (24).Rawlings ND; Bateman A Pepsin Homologues in Bacteria. BMC Genomics 2009, 10, 437. 10.1186/1471-2164-10-437. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (25).Dunn BM Structure and Mechanism of the Pepsin-Like Family of Aspartic Peptidases. Chem. Rev 2002, 102 (12), 4431–4458. 10.1021/cr010167q. [DOI] [PubMed] [Google Scholar]
- (26).Simões I; Faro R; Bur D; Kay J; Faro C Shewasin A, an Active Pepsin Homolog from the Bacterium Shewanella Amazonensis. FEBS J. 2011, 278 (17), 3177–3186. 10.1111/J.1742-4658.2011.08243.X. [DOI] [PubMed] [Google Scholar]
- (27).Leal AR; Cruz R; Bur D; Huesgen PF; Faro R; Manadas B; Wlodawer A; Faro C; Simões I Enzymatic Properties, Evidence for in Vivo Expression, and Intracellular Localization of Shewasin D, the Pepsin Homolog from Shewanella Denitrificans. 2016. 10.1038/srep23869. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (28).Ido E; Han HP; Kezdy FJ; Tang J Kinetic Studies of Human Immunodeficiency Virus Type 1 Protease and Its Active-Site Hydrogen Bond Mutant A28S. J. Biol. Chem 1991, 266 (36), 24359–24366. 10.1016/S0021-9258(18)54237-X. [DOI] [PubMed] [Google Scholar]
- (29).Yamauchi T; Nagahama M; Hori H; Murakami K Functional Characterization of Asp-317 Mutant of Human Renin Expressed in COS Cells. FEBS Lett. 1988, 230 (1–2), 205–208. 10.1016/0014-5793(88)80672-0. [DOI] [PubMed] [Google Scholar]
- (30).Miller A; Matera-Witkiewicz A; Mikołajczyk A; Wątły J; Wilcox D; Witkowska D; Rowińska-Żyrek M Zn-Enhanced Asp-Rich Antimicrobial Peptides: N-Terminal Coordination by Zn(II) and Cu(II), Which Distinguishes Cu(II) Binding to Different Peptides. Int. J. Mol. Sci 2021, 22 (13), 6971. 10.3390/ijms22136971. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (31).Harris F; Dennison SR; Phoenix DA Anionic Antimicrobial Peptides from Eukaryotic Organisms. Curr. Protein Pept. Sci 2009, 10 (6), 585–606. 10.2174/138920309789630589. [DOI] [PubMed] [Google Scholar]
- (32).Mistry J; Finn RD; Eddy SR; Bateman A; Punta M Challenges in Homology Search: HMMER3 and Convergent Evolution of Coiled-Coil Regions. Nucleic Acids Res. 2013, 41 (12), e121. 10.1093/nar/gkt263. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (33).Zallot R; Oberg N; Gerlt JA The EFI Web Resource for Genomic Enzymology Tools: Leveraging Protein, Genome, and Metagenome Databases to Discover Novel Enzymes and Metabolic Pathways. Biochemistry 2019, 58 (41), 4169–4182. 10.1021/acs.biochem.9b00735. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (34).Shannon P; Markiel A; Ozier O; Baliga NS; Wang JT; Ramage D; Amin N; Schwikowski B; Ideker T Cytoscape: A Software Environment for Integrated Models of Biomolecular Interaction Networks. Genome Res. 2003, 13 (11), 2498–2504. 10.1101/gr.1239303. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (35).Edgar RC MUSCLE: A Multiple Sequence Alignment Method with Reduced Time and Space Complexity. BMC Bioinformatics 2004, 5 (1), 113. 10.1186/1471-2105-5-113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (36).Price MN; Dehal PS; Arkin AP FastTree 2 – Approximately Maximum-Likelihood Trees for Large Alignments. PLOS ONE 2010, 5 (3), e9490. 10.1371/journal.pone.0009490. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (37).Letunic I; Bork P Interactive Tree Of Life (iTOL) v4: Recent Updates and New Developments. Nucleic Acids Res. 2019, 47 (W1), W256–W259. 10.1093/nar/gkz239. [DOI] [PMC free article] [PubMed] [Google Scholar]
- (38).Tamura K; Stecher G; Kumar S MEGA11: Molecular Evolutionary Genetics Analysis Version 11. Mol. Biol. Evol 2021, 38 (7), 3022–3027. 10.1093/molbev/msab120. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplemental Dataset 1: Custom HMMs (.txt)
Supplemental Dataset 2: Precursor Dataset (.xlsx)
Supplemental Dataset 3: Phylogenetic trees in Newick format (.nwk)
Supplemental Dataset 4: Primers, genes, plasmids, and protein sequences used in this study (.xlsx)
