Skip to main content
Applications in Plant Sciences logoLink to Applications in Plant Sciences
. 2025 Jun 3;13(4):e70009. doi: 10.1002/aps3.70009

The Computer‐Assisted Sequence Annotation (CASA) workflow for enzyme discovery

Gemma R Takahashi 1, Franchesca M Cumpio 1, Carter T Butts 2, Rachel W Martin 1,3,
PMCID: PMC12319702  PMID: 40766899

Abstract

Premise

With the advent of inexpensive nucleic acid sequencing and automated annotation at the level of basic functionality, the central problem of enzyme discovery is no longer finding active sequences, it is determining which ones are suitable for further study. This requires annotation that goes beyond sequence similarity to known enzymes and provides information at the sequence and structural levels.

Methods

Here we introduce a workflow for generating highly informative, richly annotated sequence alignments from protein sequence data. Computer‐Assisted Sequence Annotation (CASA) is a freely available Python‐based workflow designed to automate portions of novel protein characterization, while producing a human‐interpretable final output.

Results

We demonstrate CASA using one enzyme from the Drosera capensis genome. The workflow generates detailed annotations providing comparisons to known reference sequences. In addition to sequence similarity and predicted function, user‐specified features such as active site residues, disulfide bonds, and substrate‐binding residues can be displayed, and these can then be combined with downstream analyses to gain new insights into enzyme structure and function.

Discussion

This work demonstrates the utility of detailed annotations and protein structure prediction for choosing protein targets for biochemistry or structural biology from nucleic acid sequence data. The toolchain is freely available along with instructions and representative examples.

Keywords: Drosera capensis, enzyme discovery, genome annotation, protease, protein sequence analysis, protein sequence annotation


The rapid improvement of high‐throughput nucleic acid sequencing has resulted in the discovery of vast numbers of sequences coding for functional proteins; however, these are not fully characterized in practice. It is now relatively straightforward to sequence a genome or transcriptome and translate putative genes to obtain a “parts list” with thousands of protein sequences. Sequence‐based annotation methods can divide these sequences into families and predict some aspects of function. However, realizing the full utility of these enzymes still requires experimental characterization, which is expensive and time‐consuming. Strategic target selection is required to ensure that resources are allocated to the most promising targets.

Sequence‐based functional annotation is a fundamental part of enzyme discovery. Automated sequence‐based annotation can make predictions about enzyme class and some sequence features. In terms of time and resources, these methods are much less expensive than their molecular modeling and experimental counterparts, and they can be adapted to diverse datasets. This makes them ideal for the high‐throughput annotation of proteins as a first step in parsing large sets of sequence data. Bioinformatics tools are often featured in sequence‐based workflows for consistent comparison and alignment with well‐annotated reference proteins. These programs are generally optimized for examining large numbers of sequences simultaneously and providing alignments and sequence conservation in a standardized format. To maximize their utility for researchers interested in choosing specific targets for experimental characterization, the results obtained from these programs must be supplemented with additional sequence features like active site residues, propeptides, disulfide bonds, metal ion binding sites, or others specified by the users. Here we introduce the Computer‐Assisted Sequence Annotation (CASA) workflow, which is a set of Python scripts designed to facilitate sequence analysis and novel enzymes that have well‐characterized homologs. CASA builds on the output of useful but very general bioinformatics tools by adding user‐defined sequence features in a human‐readable format with a convenient graphical output. Importantly, the goal of CASA is to support the investigation of target sequences by the analyst, rather than to replace scientific judgment with an automated procedure. It is thus intended to be complementary to other elements of the researcher's toolkit and to facilitate integration of biological knowledge into the interpretation of sequence data.

Examples of freely available resources for protein sequence analysis include the Basic Local Alignment Search Tool (BLAST) (Altschul et al., 1990; Camacho et al., 2009), which searches for similar sequences and ranks them in terms of identity, and Clustal Omega (Sievers et al., 2011; Sievers and Higgins, 2018), which aligns input sequences and displays sequence conservation among selected proteins. Other tools, such as InterProScan (Jones et al., 2014), can be used to automate repetitive parts of protein characterization. InterProScan can, among other features, add functional annotations to proteins in the form of Gene Ontology (GO) (Ashburner et al., 2000) and MetaCyc (Caspi et al., 2014) codes. It is high‐throughput; can be run on DNA, RNA, or protein sequences; and has a well‐documented command‐line interface, but it does not readily display the data in graphical form.

UniProt has features that are useful at both the machine‐readable and human‐readable levels. Once reference sequences have been identified for a particular new sequence, it is necessary to manually validate predicted features in novel proteins. For this purpose, UniProt Tools includes a protein alignment viewer, complete with feature annotations of UniProt entries. Its user‐friendly design and online interface make it useful for manually curated comparisons of a few proteins at a time; however, there is no command‐line version, limiting the number of sequences that can be examined at once. A major goal of CASA is to bridge this gap between mass functional annotation and manual curation of novel proteins.

The development of CASA was influenced by several useful bioinformatics tools that aid in the retrieval, processing, and visualization of large biological datasets. Packages like UniProtR (Soudy et al., 2020), uniprot (The UniProt Consortium, 2024), UniProt.ws (Carlson and Ortutay, 2017), queryup (Voisinne, 2019), and BioServices (Cokelaer et al., 2013) facilitate tasks including general programmatic access to UniProt web services, retrieval of protein sequence information, and certain types of visualization. Additionally, open source, community‐driven projects like BioPython (Cock et al., 2009) and Bioconductor (Huber et al., 2015) have made processing biological data and running bioinformatics tools more accessible. These kinds of packages encourage the development of custom scripts for specific workflows and provide a foundation for tools like CASA. For more targeted curation and visualization of sequence data and annotations, there are tools like MView (Brown et al., 1998), JalView (Waterhouse et al., 2009), UniProt Tools (Zaru et al. 2023), the Protein Analysis Toolbox (Gracy and Chiche, 2005), and the MPI Bioinformatics Toolkit (Gabler et al., 2020). CASA takes inspiration from both high‐throughput, general tools and targeted analysis toolkits to deaggregate and make sense of large protein sequence data in a systematic manner.

CASA is most similar to other protein characterization tools in the early steps that rely on database searches and sequence comparison. Two such tools are OrthoMCL (Li et al., 2003) and OrthoFinder (Emms and Kelly, 2015), both of which identify orthologous groups of proteins between eukaryotic genomes. In addition to contributing to phylogenetic inference, this method of clustering can annotate function in novel proteins, a goal that is shared by tools like InterProScan (Jones et al., 2014). Still, broad functional annotation alone does not tell a full story of protein function; the manual annotation of key residues in individual proteins remains an important part of novel protein characterization, especially if those proteins are unique in sequence or structure.

Manual annotation can make use of features that are not always annotated in reference databases. For example, nepenthesins, digestive enzymes from Nepenthes pitcher plants, are distinguished from other plant aspartic proteases (APs) by a region with (often) four cysteine residues (Takahashi et al., 2005) known as a nepenthesin‐type AP‐specific insert (NAP‐specific insert) (Athauda et al., 2004). Disulfide bonds between these cysteines may contribute to the unusual stability of nepenthesins at a range of pH values and temperatures (Takahashi et al., 2005). Nepenthesin I (Athauda et al., 2004) from Nepenthes gracilis Korth., a typical representative of the nepenthesin subfamily, is included in UniProt's manually curated Swiss‐Prot database under the accession code NEP1_NEPGR (The UniProt Consortium, 2024). Although nepenthesin I is known to contain a NAP‐specific insert (Athauda et al., 2004), this feature is not specifically annotated in the NEP1_NEPGR UniProt entry. Notably, the two NAP‐specific insert disulfide bonds proposed in Takahashi et al. (2005) are annotated in NEP1_NEPGR and were added via manual assertion (The UniProt Consortium, 2024). CASA facilitates the addition of these types of features to alignments by combining steps commonly used in mass protein functional annotation (automated database searches, alignment) with those necessary for more detailed annotation (individual feature comparison, refinement of inference using outside references or tools). In doing so, it streamlines and organizes the often tedious but valuable process of manual protein sequence annotation.

CASA can be used to perform a BLAST search (Altschul et al., 1990; Camacho et al., 2009) of a FASTA‐formatted list of proteins against the most recent, manually curated Swiss‐Prot database (Boutet et al., 2016; Poux et al., 2017), generate a Clustal Omega (Sievers et al., 2011; Sievers and Higgins, 2018) alignment from a FASTA‐formatted list of proteins, download a table of UniProt annotations for valid UniProt entries in either a FASTA‐ or valid Clustal‐formatted list of proteins, and generate a publication‐quality scalable vector graphics (SVG) file from a valid Clustal‐formatted file. The four base scripts that encode the above functions can be run independently or in sequence and are freely available from a GitHub repository (https://github.com/grtakaha/CASA; see Data Availability Statement). Here, we describe the operation of these scripts and provide example workflows, illustrated using novel enzyme sequences from the carnivorous plant Drosera capensis L. (Slack, 1979).

METHODS

Workflow overview

At its core, CASA is a manager of four Python scripts (Figure 1): search_proteins.py, retrieve_annotations.py, alignment.py, and clustal_to_svg.py. Each relies on one or all of the included support modules (helpers.py, api_funcs.py, protein_classes.py), as well as the required command‐line tools (BLAST+ and Clustal Omega). These scripts are independent of operating system; there are Windows, MAC, and Linux options for all required installations. A brief description of each script is included in Appendix S1 (see Supporting Information with this article) and use recommendations can be found in the README file on GitHub (https://github.com/grtakaha/CASA/; see Data Availability Statement). The version of CASA used here was built, tested, and run on Ubuntu 24.04.1 LTS with Python 3.12.3 (pandas 2.2.3 and requests 2.32.3), Clustal Omega 1.2.4 (Sievers et al., 2011; Sievers and Higgins, 2018), and BLAST 2.16.0+ (Camacho et al., 2009). Although CASA includes options for the user to adjust BLAST and Clustal Omega settings, it primarily uses the default settings for the BLASTP and ClustalO algorithms. One exception is the ‐num_alignments option for BLASTP, which can be specified by CASA's ‐nr option; here, this number is set to five in order to limit the size of the final SVG alignments. The BLAST command used in default CASA runs is: blastp ‐query [QUERY] ‐outfmt 6 ‐out [OUT] ‐num_alignments 5 ‐db [SWISS‐PROT]. The Clustal Omega command used in default CASA runs is: clustalo ‐‐infile [INFILE] ‐‐outfile [OUTFILE] ‐‐seqtype Protein ‐‐force ‐‐outfmt clustal. Defaults for BLASTP 2.16.0+ include: ‐num_descriptions 500, ‐matrix BLOSUM62, ‐evalue 10, and ‐soft‐masking false. Clustal Omega 1.2.4 also uses a fast clustering algorithm, mBed (Blackshields et al., 2010), instead of calculating full distance matrices; because the number of sequences recommended here does not exceed ClustalO's default cluster size of 100, alignments should be generated with full distance matrices regardless. Detailed option descriptions can be found in the specified versions' README files and help command outputs.

Figure 1.

Figure 1

CASA overview. The four scripts managed by CASA.py are shown in different quadrants. Inputs and outputs for each script are indicated by arrows; file formats with arrows pointing toward CASA.py (outputs) are passed to the tool manager before being used as inputs for the next script. The suggested step order is indicated above the name of each script. Alignment SVG files output by clustal_to_svg.py are not passed back through CASA.py, and therefore do not have an arrow. All output file formats are saved to the designated output directory (option ‐o) when CASA.py is initially run. search_proteins.py: input = FASTA, output = FASTA. retrieve_annotations.py: input = Clustal or FASTA, output = annotations. alignment.py: input = FASTA, output = Clustal. clustal_to_svg.py: input = Clustal and annotations, output = SVG.

For the initial organization and visualization of target proteins, we recommend executing all four CASA scripts in succession; the order is designated by the ‐ord option: ‐ord blast annotate align svg. Six optional steps are further described in Appendix S1, and valid script orders can be found in the README file on GitHub (see Data Availability Statement). This basic workflow takes one FASTA file with at least one target protein as input and runs all four scripts in CASA on each target protein.

The workflow begins with search_proteins.py, which creates separate directories for each of the input target protein sequences. By default, search_proteins.py also downloads and creates a BLAST database for the latest Swiss‐Prot release, and then runs BLASTP (Camacho et al., 2009) using each target protein as a query against this Swiss‐Prot subject database. Results, which include target protein query–Swiss‐Prot subject pairs, are saved in the tabular BLAST outfmt 6 ([QUERY].tsv). All BLAST hit sequences are then downloaded through UniProt's REST API (The UniProt Consortium, 2024) and combined into a single FASTA file along with the target protein (all.fasta).

Next, retrieve_annotations.py retrieves residue features for each BLAST hit via UniProt's REST API (The UniProt Consortium, 2024). By default, these include active sites (Active site), disulfide bonds (Disulfide bond), propeptides (Propeptide), and signal sequences (Signal). All features are saved to a table (all.ann).

The all.fasta file from search_proteins.py, which contains one target protein query and all BLAST hits for that query against a subject database, is then aligned by a call to ClustalO (Sievers et al., 2011; Sievers and Higgins, 2018) through alignment.py. By default, alignment.py outputs a Clustal‐formatted file (alignment.clustal).

In the final step, annotations (all.ann) and the alignment (alignment.clustal) are submitted to clustal_to_svg.py, which generates an annotated set of SVG files (0.svg, 1.svg, etc.). Multiple SVG files are generated when an alignment exceeds the set page size (i.e., 7.5 × 9 in).

This use of CASA consists of just one command, which can be run from any command‐line interface. As it runs, it writes a log of commands and key information to the command interface; for large runs, we recommend redirecting this output to a file for future reference. Notably, annotated feature residues and alignment positions are printed at the same time that SVG files are generated, making it possible to tell which alignment SVGs contain target annotations. For more example commands, use cases, downstream applications, and CASA options, see Appendix S1 and the README file on GitHub (see Data Availability Statement).

RESULTS

The visualization of functional sequence regions within proteins is often needed prior to experimental characterization. This is usually done by comparing target proteins to well‐annotated references, and can be a tedious process. The basic protocol, which is summarized in Figure 1, is meant to simplify and organize the generation of files used in detailed protein annotation workflows, and is recommended as a first look at a new sequence. The example given here, a putative digestive protease from D. capensis, was chosen to demonstrate key features of CASA. Here we show an annotated sequence alignment that enables the visualization of starts and stops, labels important features of the novel protein, and indicates sequence similarities with well‐annotated reference sequences as well as features that are useful in downstream structural or biochemical analyses.

Case study: Cysteine protease from Drosera capensis

Here we demonstrate the application of CASA using DCAP_6781, a putative cysteine protease identified from the D. capensis genome (Butts et al., 2016ab). CASA was initially run using a FASTA‐formatted file that included the DCAP_6781 protein sequence as input; options were selected to execute all four scripts (‐ord blast annotate align svg), retrieve five BLAST hits (‐nr 5), and output an annotated alignment SVG with no numbers (‐n FALSE), no Clustal Omega conservation codes (‐c FALSE), and reformatted UniProt IDs (‐u TRUE) (Figure 2). Active site residues and disulfide bonds were identified manually using this initial alignment SVG. CASA has built‐in options and features for the purpose of manual residue number identification; more information is available on the GitHub repository (see Data Availability Statement).

Figure 2.

Figure 2

The DCAP_6781 alignment, prepared using default CASA settings and subsequent custom runs. (A) The putative cysteine protease DCAP_6781 is shown aligned to five Swiss‐Prot entries: SAG39_ORYSI, SAG39_ORYSJ, SAG12_ARATH, CEP1_ARATH, and CYSEP_PHAVU. A legend has been manually added to the bottom right corner. (B) DCAP_6781 aligned to five Swiss‐Prot entries and dionain 1 (A0A0E3GLN3_DIOMU). Dionain 1's propeptide was added manually to the annotation file used to generate this figure. Residues are colored according to their biochemical property: blue = positive, red = negative, green = hydrophobic, yellow = special, and black = other. Colored circles indicate conserved residues. Residues that are part of annotated features are highlighted by colored rectangles: green = signal sequence, blue = active site residue, orange = disulfide bond cysteine, purple = propeptide.

Two of the original five BLAST hits contained annotated propeptides, but neither cleavage site aligned well with DCAP_6781's sequence. CASA was run two additional times with different settings in order to further identify a propeptide cleavage site for DCAP_6781. Dionain 1 (UniProt ID: A0A0E3GLN3_DIOMU), a homologous enzyme from Dionaea muscipula J. Ellis, was chosen as an additional reference. The sequence for dionain 1, which was identical to that presented in Risør et al. (2016), was manually downloaded from the UniProt website (https://www.uniprot.org/) and added to a FASTA file with DCAP_6781 and its five original BLAST hits. CASA was run on this new FASTA file, starting from the alignment step, to retrieve UniProt annotations for dionain 1. From these annotations, it was apparent that dionain 1's propeptide was not explicitly annotated in its UniProt entry. We added this propeptide annotation manually and ran CASA again from the SVG‐generating step, this time using the modified annotations and the last alignment as input. Editing of these output SVG files for figures was performed in Inkscape v1.4. Files and commands used can be found in an Open Science Foundation repository (https://osf.io/xnmha/; see Data Availability Statement).

By default, CASA retrieves and visualizes active site residues (Active site), propeptides (Propeptide), signal sequences (Signal), and disulfide bond cysteines (Disulfide bond) listed in the online entries for valid UniProt proteins. These annotations, denoted by colored rectangles in each alignment SVG (Figure 2), can be used to identify the same features in the unannotated target protein upon examination of the alignment. There are several conserved features among the six proteins (DCAP_6781 and five Swiss‐Prot entries) (Figure 2A); disulfide bond cysteines and active site residues are conserved among all six sequences. DCAP_6781's active site can therefore be putatively identified as a catalytic triad made of cysteine (C152), histidine (H289), and asparagine (N310). Similarly, the cysteines likely to be involved in disulfide bonds are C149, C183, C192, C225, C283, and C335, although some of the reference features shown in Figure 2A (e.g., propeptides) share less similarity with DCAP_6781. It is important to consider external information when annotating low‐similarity features.

Two of the Swiss‐Prot reference proteins for DCAP_6781 contain propeptides, but neither cleavage site shares much sequence similarity with DCAP_6781 itself. Realigning with outside sequences can provide more information. DCAP_6781 does align well with a homologous protease from Dionaea muscipula called dionain 1 (UniProt ID: A0A0E3GLN3_DIOMU), which has an experimentally solved structure that is available for comparison. Its proposed propeptide cleavage site (Risør et al., 2016) is conserved among several D. capensis cysteine proteases (Butts et al., 2016b), including DCAP_6781. The DCAP_6781 alignment after adding dionain 1 and manually annotating its propeptide, prior to final SVG generation, is shown in Figure 2B. Removing DCAP_6781's propeptide at this cleavage site (D127‐V128) can improve downstream molecular modeling and experimental analyses. Such insights complement the information obtained from sequence data and provide novel hypotheses for subsequent experimental investigation.

DISCUSSION

CASA: Features and user‐specified options

CASA enables detailed functional annotation of novel proteins using multiple sources and an iterative scientific process. These semi‐automated workflows can inform experimental design by revealing key residues and improve the biological relevance of molecular models with knowledge of active structures. Examples of downstream applications for this information, like protein structure prediction and molecular dynamics simulations, can be found in Appendix S1.

Because users have different goals and preferences for sequence alignments, CASA also allows the default BLASTP, ClustalO, or CASA settings to be modified. For example, users and developers have sometimes interpreted BLAST options and outputs differently (González‐Pech et al., 2018; Shah et al., 2019); users who would like more control over BLAST and Clustal Omega settings in CASA can make use of the ‐bopts and ‐copts options, respectively. CASA itself avoids making any functional or structural predictions, and instead provides a way to organize and visualize protein sequence data in preparation for downstream, user‐guided analysis. For this reason, CASA does not include its own BLAST filtering algorithms, although it does limit BLAST hits using the ‐num_alignments option from BLASTP, specified through CASA's ‐nr option. There is a limit to the number of sequences that can physically fit on a 7.5 × 9 in page in a CASA SVG file; limiting the number of initial BLAST hits ensures that clustal_to_svg.py is able to generate readable outputs.

The CASA scripts are inherently modular in order to provide maximum control over filtering, annotation, or other steps involved in protein characterization pipelines. Users can customize their workflows by executing the CASA scripts either one at a time or in batches, redirecting inputs and outputs where desired. For example, if users choose to perform their own BLAST filtering, they may run search_proteins.py from CASA (‐ord blast), use their own scripts to filter BLAST hits, and then pass those hits to the other three scripts in CASA (‐ord annotate align svg). Alternatively, a user may choose to include annotations from different sources; for this purpose, they may run search_proteins.py, alignment.py, and retrieve_annotations.py (‐ord blast align annotate) before adding custom annotations to all.ann files from retrieve_annotations.py. These annotations can then be passed to clustal_to_svg.py (‐ord svg) with individual commands or with scripts.

CASA also includes options not demonstrated here (see Appendix S1 and the GitHub repository for more details): the ‐‐features option (‐f OR ‐‐features) allows users to retrieve alternative features from UniProt and specify their SVG colors, and the ‐db option is used to specify an alternative BLAST subject database. For ease of residue identification, CASA can add Clustal information, like conservation codes and residue numbers, back into SVG alignments using its ‐c and ‐n options, respectively. Additionally, SVG files are initially generated with sequence‐accurate titles for each line and residue; these can be seen in the Inkscape XML editor and are meant to aid the user in identifying specific residue numbers. Users may also have need for general functional annotation in downstream analyses. For this reason, retrieve_annotations.py (‐ord annotate) outputs a metadata file (all.metadata) that includes descriptions, Enzyme Commission (EC) numbers, and Protein Data Bank (PDB) structure IDs for all valid UniProt protein inputs. In this way, CASA encourages users to customize and iterate over their workflows.

Case study: Cysteine protease from Drosera capensis

Cysteine proteases are common in plant defenses against herbivory and pathogen infection. Well‐known examples include papain from papaya (Carica papaya L.) and bromelain from pineapple (Ananas comosus (L.) Merr.). In D. capensis, our initial analysis revealed 44 cysteine protease paralogs, and many plants have more. These enzymes provide a potentially rich source of proteases with novel functionality; however, many of them are likely to be functionally equivalent to known enzymes. Predicting which of many closely related sequences may have different substrate specificity or other variations in chemistry is a major goal. In our case study, we present a detailed characterization of key features of the novel protease DCAP_6781 from the genome of D. capensis, which is a typical example of a cysteine protease. This specific protein was chosen after downstream examination of its predicted structure revealed an unusually active site architecture (Appendix S1).

In carnivorous plants like D. capensis, free amino acids from protease‐catalyzed protein degradation can be used as building blocks for other proteins or as alternative sources of nitrogen (Kováčik et al., 2012b). For example, phenylalanine ammonia lyase (PAL) liberates nitrogen from phenylalanine, generating cinnamic acid, a precursor to phenolic metabolites; freed nitrogen can in turn be used in amino acid metabolism. In Drosera spatulata Labill., its activity, and thus the level of total soluble phenols, was shown to increase under low nitrogen conditions (Kováčik et al., 2012a). Relatedly, phenol levels did not increase in prey‐fed D. capensis (Kováčik et al., 2012b); it can be hypothesized that free amino acids derived from protease activity may mitigate the need for PAL‐liberated nitrogen in amino acid metabolism (Kováčik et al., 2012a).

When examining a new protein sequence, it is often useful to locate functional sequence features. In digestive proteases, these include active site residues, signal peptides, and propeptides. Many enzymes are expressed as pro‐enzymes or zymogens, which have sequence regions that are part of the protein during and just after translation but are cleaved later to yield the mature enzyme. This is common with proteases, where the immediate activation of the enzyme in the cytoplasm would result in uncontrolled cleavage. Therefore, many proteases have a propeptide that blocks the active site until the enzyme reaches its destination, often the vacuole or the extracellular space. Proteins that are exported from the cell after expression generally also have a signal sequence, which often takes the form of an N‐terminal helix that serves as a binding site for the secretion machinery to recognize. Signal sequences often have low sequence identity even among related proteins, but can be identified by amino acid composition using automated tools such as SignalP (Teufel et al., 2021). In contrast, propeptides generally lack common properties and must be predicted by comparison to similar sequences for which experimental data are available. CASA allows users to integrate hand‐curated knowledge from external sources with their automated protein characterization workflows.

Conclusions

CASA was designed to automate repetitive portions of novel protein characterization, while leaving room for an iterative scientific process upon further analysis. It complements high‐throughput, aggregate data workflows by allowing users to manually curate and regenerate custom results, but is still capable of automating and organizing outputs from large datasets. CASA is intended to provide a first look at novel protein sequences, and it is most useful when coupled with other targeted bioinformatics tools or molecular modeling techniques. The information inferred from CASA's annotated alignments can both improve and be strengthened by downstream applications. As raw sequence data continues to accumulate, the efficient, manual curation of rich protein annotations will provide a foundation for future high‐throughput methodologies.

AUTHOR CONTRIBUTIONS

G.R.T. and R.W.M. designed the study and performed sequence analysis and molecular model annotation. G.R.T. wrote and tested the scripts, F.M.C. performed sequence analysis and molecular model annotation, and C.T.B. performed molecular dynamics simulations. G.R.T., C.T.B., and R.W.M. made primary contributions and edits to the final manuscript. All authors approved the final version of the manuscript.

Supporting information

Appendix S1. Detailed description of the CASA scripts, examples highlighting features of CASA, and potential downstream analysis using protein structure prediction and molecular modeling.

ACKNOWLEDGMENTS

The authors acknowledge Jose Uribe and Karen S. Campos (University of California, Irvine) for translating the abstract into Spanish. This work was supported by the National Institutes of Health (grant R01GM144964 to R.W.M. and C.T.B.). G.R.T. was supported by GAANN award P200A210024 to the Department of Molecular Biology and Biochemistry at the University of California, Irvine.

Takahashi, G. R. , Cumpio F. M., Butts C. T., and Martin R. W.. 2025. The Computer‐Assisted Sequence Annotation (CASA) workflow for enzyme discovery. Applications in Plant Sciences 13(4): e70009. 10.1002/aps3.70009

This article is part of the special issue “Advances in analyzing and engineering plant metabolic diversity.”

DATA AVAILABILITY STATEMENT

All commands, CASA outputs, and predicted structures are available in an Open Science Foundation data repository (https://osf.io/xnmha/; DOI 10.17605/OSF.IO/XNMHA). Scripts can be found at https://github.com/grtakaha/CASA/.

REFERENCES

  1. Altschul, S. F. , Gish W., Miller W., Myers E. W., and Lipman D. J.. 1990. Basic local alignment search tool. Journal of Molecular Biology 215: 403–410. [DOI] [PubMed] [Google Scholar]
  2. Ashburner, M. , Ball C. A., Blake J. A., Botstein D., Butler H., Cherry J. M., Davis A. P., et al. 2000. Gene Ontology: Tool for the unification of biology. Nature Genetics 25: 25–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Athauda, S. B. P. , Matsumoto K., Rajapakshe S., Kuribayashi M., Kojima M., Kubomura‐Yoshida N., Iwamatsu A., et al. 2004. Enzymic and structural characterization of nepenthesin, a unique member of a novel subfamily of aspartic proteinases. Biochemical Journal 381: 295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Blackshields, G. , Sievers F., Shi W., Wilm A., and Higgins D. G.. 2010. Sequence embedding for fast construction of guide trees for multiple sequence alignment. Algorithms for Molecular Biology 5: 21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Boutet, E. , Lieberherr D., Tognolli M., Schneider M., Bansal P., Bridge A. J., Poux S., et al. 2016. UniProtKB/Swiss‐Prot, the manually annotated section of the UniProt KnowledgeBase: How to use the entry view. Methods in Molecular Biology 1374: 23–54. [DOI] [PubMed] [Google Scholar]
  6. Brown, N. P. , Leroy C., and Sander C.. 1998. MView: A web‐compatible database search or multiple alignment viewer. Bioinformatics 14: 380–381. [DOI] [PubMed] [Google Scholar]
  7. Butts, C. T. , Bierma J. C., and Martin R. W.. 2016a. Novel proteases from the genome of the carnivorous plant Drosera capensis: Structural prediction and comparative analysis. Proteins: Structure, Function, and Bioinformatics 84: 1517–1533. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Butts, C. T. , Zhang X., Kelly J. E., Roskamp K. W., Unhelkar M. H., Freites J. A., Tahir S., and Martin R. W.. 2016b. Sequence comparison, molecular modeling, and network analysis predict structural diversity in cysteine proteases from the Cape sundew, Drosera capensis . Computational and Structural Biotechnology Journal 14: 271–282. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Camacho, C. , Coulouris G., Avagyan V., Ma N., Papadopoulos J., Bealer K., and Madden T. L.. 2009. BLAST+: Architecture and applications. BMC Bioinformatics 10: 421. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Carlson, M. , and Ortutay C.. 2017. UniProt.ws: R interface to UniProt web services. R package version 2. Website https://bioconductor.org/packages/UniProt.ws [accessed 24 April 2025].
  11. Caspi, R. , Altman T., Billington R., Dreher K., Foerster H., Fulcher C. A., Holland T. A., et al. 2014. The MetaCyc database of metabolic pathways and enzymes and the BioCyc collection of Pathway/Genome Databases. Nucleic Acids Research 42: D459–D471. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Cock, P. J. A. , Antao T., Chang J. T., Chapman B. A., Cox C. J., Dalke A., Friedberg I., et al. 2009. Biopython: Freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics 25: 1422–1423. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Cokelaer, T. , Pultz D., Harder L. M., Serra‐Musach J., and Saez‐Rodriguez J.. 2013. BioServices: A common Python package to access biological Web Services programmatically. Bioinformatics 29: 3241–3242. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Emms, D. M. , and Kelly S.. 2015. OrthoFinder: Solving fundamental biases in whole genome comparisons dramatically improves orthogroup inference accuracy. Genome Biology 16: 157. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Gabler, F. , Nam S.‐Z., Till S., Mirdita M., Steinegger M., Söding J., Lupas A. N., and Alva V.. 2020. Protein sequence analysis using the MPI bioinformatics toolkit. Current Protocols in Bioinformatics 72: e108. [DOI] [PubMed] [Google Scholar]
  16. González‐Pech, R. A. , Stephens T. G., and Chan C. X.. 2018. Commonly misunderstood parameters of NCBI BLAST and important considerations for users. Bioinformatics 35: 2697–2698. [DOI] [PubMed] [Google Scholar]
  17. Gracy, J. , and Chiche L.. 2005. PAT: A protein analysis toolkit for integrated biocomputing on the web. Nucleic Acids Research 33: W65–W71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Huber, W. , Carey V. J., Gentleman R., Anders S., Carlson M., Carvalho B. S., Bravo H. C., et al. 2015. Orchestrating high‐throughput genomic analysis with Bioconductor. Nature Methods 12: 115–121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Jones, P. , Binns D., Chang H.‐Y., Fraser M., Li W., McAnulla C., McWilliam H., et al. 2014. InterProScan 5: Genome‐scale protein function classification. Bioinformatics 30: 1236–1240. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Kováčik, J. , Klejdus B., and Repčáková K.. 2012a. Phenolic metabolites in carnivorous plants: Inter‐specific comparison and physiological studies. Plant Physiology and Biochemistry 52: 21–27. [DOI] [PubMed] [Google Scholar]
  21. Kováčik, J. , Klejdus B., Štork F., and Hedbavny J.. 2012b. Prey‐induced changes in the accumulation of amino acids and phenolic metabolites in the leaves of Drosera capensis L. Amino Acids 42: 1277–1285. [DOI] [PubMed] [Google Scholar]
  22. Li, L. , Stoeckert C. J., and Roos D. S.. 2003. OrthoMCL: Identification of ortholog groups for eukaryotic genomes. Genome Research 13: 2178–2189. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Poux, S. , Arighi C. N., Magrane M., Bateman A., Wei C.‐H., Lu Z., Boutet E., et al. 2017. On expert curation and scalability: UniProtKB/Swiss‐Prot as a case study. Bioinformatics 33: 3454–3460. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Risør, M. W. , Thomsen L. R., Sanggaard K. W., Nielsen T. A., Thøgersen I. B., Lukassen M. V., Rossen L., et al. 2016. Enzymatic and structural characterization of the major endopeptidase in the Venus flytrap digestion fluid. Journal of Biological Chemistry 291: 2271–2287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Shah, N. , Nute M. G., Warnow T., and Pop M.. 2019. Misunderstood parameter of NCBI BLAST impacts the correctness of bioinformatics workflows. Bioinformatics 35: 1613–1614. [DOI] [PubMed] [Google Scholar]
  26. Sievers, F. , and Higgins D. G.. 2018. Clustal Omega for making accurate alignments of many protein sequences. Protein Science 27: 135–145. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Sievers, F. , Wilm A., Dineen D., Gibson T. J., Karplus K., Li W., Lopez R., et al. 2011. Fast, scalable generation of high‐quality protein multiple sequence alignments using Clustal Omega. Molecular Systems Biology 7: 539. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Slack, A. 1979. Carnivorous plants. MIT Press, Cambridge, Massachusetts, USA. [Google Scholar]
  29. Soudy, M. , Anwar A. M., Ahmed E. A., Osama A., Ezzeldin S., Mahgoub S., and Magdeldin S.. 2020. UniprotR: Retrieving and visualizing protein sequence and functional information from Universal Protein Resource (UniProt knowledgebase). Journal of Proteomics 213: 103613. [DOI] [PubMed] [Google Scholar]
  30. Takahashi, K. , Athauda S. B. P., Matsumoto K., Rajapakshe S., Kuribayashi M., Kojima M., Kubomura‐Yoshida N., et al. 2005. Nepenthesin, a unique member of a novel subfamily of aspartic proteinases: Enzymatic and structural characteristics. Current Protein & Peptide Science 6: 513–525. [DOI] [PubMed] [Google Scholar]
  31. Teufel, F. , Armenteros J. J. A., Johansen A. R., Gíslason M. H., Pihl S. I., Tsirigos K. D., Winther O., et al. 2021. SignalP 6.0 achieves signal peptide prediction across all types using protein language models. biorXiv 447770 [preprint]. Available at: 10.1101/2021.06.09.447770 [posted 10 June 2021; accessed 24 April 2025]. [DOI]
  32. The UniProt Consortium . 2024. UniProt: The Universal Protein Knowledgebase in 2025. Nucleic Acids Research 53: D609–D617. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Voisinne, G. 2019. queryup: Query the UniProt REST API using R. Centre d'Immunologie de Marseille‐Luminy, Aix Marseille Université, INSERM, CNRS, Marseille, France. Website: https://github.com/VoisinneG/queryup/ [accessed 24 April 2025].
  34. Waterhouse, A. M. , Procter J. B., Martin D. M. A., Clamp M., and Barton G. J.. 2009. Jalview Version 2—A multiple sequence alignment editor and analysis workbench. Bioinformatics 25: 1189–1191. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Zaru, R. , Orchard S., and The UniProt Consortium . 2023. UniProt Tools: BLAST, Align, Peptide Search, and ID Mapping. Current Protocols 3: e697. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix S1. Detailed description of the CASA scripts, examples highlighting features of CASA, and potential downstream analysis using protein structure prediction and molecular modeling.

Data Availability Statement

All commands, CASA outputs, and predicted structures are available in an Open Science Foundation data repository (https://osf.io/xnmha/; DOI 10.17605/OSF.IO/XNMHA). Scripts can be found at https://github.com/grtakaha/CASA/.


Articles from Applications in Plant Sciences are provided here courtesy of Wiley

RESOURCES