Abstract
Endogenous protein tagging in Caenorhabditis elegans enables direct visualization and manipulation of proteins in vivo, providing native readouts of expression, localization, and dynamics. No coordinated effort currently exists to comprehensively tag proteins on a large scale, resulting in patchy coverage that limits comprehensive proteome analyses. We systematically reviewed 2,500 primary research articles, identifying 778 that report novel endogenous tags, and integrated these with the Caenorhabditis Genetics Center strain records to catalog >90% of all existing tagged alleles. In total, we found that 1,554 unique genes (~8% of the proteome) have been endogenously tagged. Gene Ontology enrichment analysis revealed that cytoskeletal proteins, transcription factors, and RNA-binding proteins dominate the tagged proteome, while membrane proteins, metabolic enzymes, and mitochondrial components remain largely untagged, reflecting both technical barriers and research priorities that have shaped the last decade of tagging efforts. We created WormTagDB (https://wormtagdb.rc.duke.edu), an interactive, community-updatable resource that consolidates all known endogenously tagged alleles and provides precomputed CRISPR guide and homology-arm primer designs for N- and C-terminal knock-ins across all protein-coding genes. This will enable researchers to easily identify existing alleles to prevent redundant strain generation and rapidly initiate new knock-in experiments. A systematic effort to tag every C. elegans gene would deliver the first complete metazoan visual proteome, providing comprehensive insights into protein localization, dynamics, and regulation, revealing new protein associations and molecular processes.
Keywords: CRISPR-Cas9, genome editing, knock-in allele, fluorescent proteins, epitopes, visual proteome, primer design
Summary
Endogenous protein tagging allows researchers to visualize protein localization and function in living C. elegans, but no comprehensive catalog of tagged alleles exists. We manually surveyed 2,500 research papers and CGC records, identifying 1,554 tagged genes (~8% of the proteome). We created WormTagDB, a centralized community-updatable database that consolidates all known tagged alleles and provides precomputed CRISPR guide RNAs and primer designs for every protein-coding gene. This prevents redundant strain generation and eliminates time-consuming reagent design. Our analysis reveals biases in current tagging efforts and establishes a roadmap for completing the first metazoan visual proteome.
Introduction
The insertion of fluorescent, epitope, and affinity tags directly into native gene loci, known as endogenous tagging, has become a foundational tool in modern cell and developmental biology (Leonetti et al. 2016; Kanca et al. 2017). In this approach, the tag sequence is fused in-frame with the coding region of the target gene, allowing the resulting fusion protein to be expressed under native regulatory control and translated as a contiguous polypeptide. By preserving endogenous expression levels and regulatory contexts, this approach allows researchers to conduct biochemical and microscopy studies of protein dynamics, localization, and interactions in vivo without the confounding effects of overexpression or ectopic promoters (Frøkjær-Jensen et al. 2008; Nance and Frøkjær-Jensen 2019). Endogenously tagged alleles now underpin a wide array of imaging, biochemical, and functional studies across model organisms (Keeley et al. 2020; Cho et al. 2022; Xu et al. 2022).
Over the last three decades, techniques have evolved to achieve efficient endogenous tagging in Caenorhabditis elegans. Early attempts at gene editing via homologous recombination proved technically challenging with a low frequency of edits (e.g. Broverman et al. 1993), limiting practical utility. A major advance came in the early 2000s with the development of transposon-based recombination methods. Using excision of a Tc1 transposon inserted in the target gene to stimulate homologous recombination, Barrett et al. (2004) published the first fluorescently tagged protein expressed from its native locus in the worm by inserting GFP into the C-terminus of the frm-3 gene. While still technically challenging and inefficient, this work demonstrated that precise tagging at endogenous loci was feasible. Building on the principle of transposon-stimulated editing, the MosTIC system introduced a more generalizable method using engineered Mos1 elements (Robert and Bessereau 2007). This approach enabled targeted excision of a Mos1 insertion followed by homology-directed repair to insert GFP into the native unc-5 locus. The development of CRISPR-Cas9 genome editing in 2013 marked a turning point in endogenous fluorophore tagging in the worm (e.g. Dickinson et al. 2013). Cas9-directed double-strand breaks, coupled with homology-directed repair, allowed precise tag insertion at endogenous loci. The development of user-friendly tools, including the self-excising cassette (SEC) system (Dickinson et al. 2015), co-CRISPR screening (Arribere et al. 2014; El Mouridi et al. 2017), and cloning-free protocols (Paix et al. 2015), increased efficiency and led to the generation of tagged alleles becoming a routine practice.
Despite the widespread adoption of endogenous tagging in C. elegans, there is not yet a centralized, searchable inventory of labelled genes. WormBase and the Caenorhabditis Genetics Center (CGC) serve as vital resources for cataloguing alleles and strains, but the coverage of endogenously tagged alleles is incomplete, with many only documented in the text or supplementary information of individual research articles. This fragmentation hinders both experimental planning and community-wide coordination. Researchers can also duplicate existing alleles unknowingly, missing already-validated alleles that could accelerate studies. For example, (Riga et al. 2021) C-terminally tagged hmr-1 with GFP years after (Marston et al. 2016) published an identical allele, while Nguyen and Phillips (2021) and Charlesworth et al. (2021) independently and near-simultaneously generated functionally identical FLAG::GFP::alg-3 alleles.
Here, we performed a manual survey of the C. elegans literature, screening 2,500 papers to identify ~90% of published alleles in which a fluorescent or other protein tag was inserted into a gene’s endogenous locus. We integrated this dataset with CGC strain records to generate a comprehensive inventory of endogenously tagged genes in C. elegans. We developed WormTagDB (https://wormtagdb.rc.duke.edu), an interactive RShiny platform, to make this resource publicly accessible, searchable, and community-updatable. We also analyzed Gene Ontology (GO) term enrichment to uncover systematic trends in the current tagged gene set. Based on these findings, we propose a roadmap for a coordinated, community-driven effort to more rapidly advance the tagging of the C. elegans proteome.
1. Literature survey of endogenously tagged genes
To systematically identify C. elegans endogenously tagged genes, we performed a keyword-based search using WormBase’s Textpresso platform. Search terms were chosen to capture studies employing genome editing for native locus tagging, including “knock-in”, “knock in”, “tagged”, “CRISPR”, “gRNA”, “homology arm”, and “repair template.” This search returned 11,015 publications, each scored and ranked by keyword relevance. We manually reviewed the top 2,500 papers. This number was chosen as discovery curve analysis with an exponential saturation model, which indicated ~90% gene coverage would be achieved by examining ~2,424 papers (Supplementary Figure 1). Within this group of papers, 778 reported the generation of one or more new endogenously tagged alleles (Supplementary Table 1). Each allele was annotated for gene identity, tag type(s), tag location, and additional metadata when available. This dataset was then combined with the list of endogenously tagged strains cataloged by the CGC to generate the final endogenously tagged database (Supplementary Table 2).
Figure 1 provides an overview of the compiled dataset, including the number of newly tagged genes reported since 2013, the distribution of tag types, and the relative proportions of tag placements. In total, we identified 2,812 different alleles, which tagged 1,554 unique genes. By examining the genes endogenously tagged over time, we found that tagging activity expanded dramatically following the introduction of CRISPR-Cas9 genome editing in 2013 (Figure 1A). However, since 2019, the number of unique genes tagged per year has been a steady ~200 new genes per year rather than increasing, suggesting a plateau in technique innovation and expertise among laboratories in the field.
Figure 1 |. Landscape of endogenously tagged genes in C. elegans.

A) Cumulative number of unique genes tagged endogenously over time. Tagging has increased linearly since 2019, with an average rate of ~194 new genes per year (linear fit: R2 = 0.9968, p < 2e-16). B) Overlap between genes tagged in the published literature and those available from the CGC. C) Number of distinct alleles per gene. Most genes (63.1%) have only one reported tagged allele, while a minority have multiple independently generated alleles. D) Distribution of tag types across all genes. Many genes are tagged with multiple tag classes, including fluorescent proteins, epitopes, degrons, affinity tags, and others. E) Tag insertion positions across all alleles. C-terminal tagging is most common (62.9%), followed by N-terminal (33.7%) and internal insertions (2.9%). A small number of alleles (0.5%) carry multiple tags at different positions.
Given that researchers usually contribute their most significant published strains to the CGC, we expected, and found, a considerable overlap of tagged genes described in the literature and those found in the CGC database (Figure 1B). However, 47% of all tagged genes are not available from the CGC. In many cases, genes have been tagged in multiple ways by different labs, so there are more generated alleles than tagged genes. Of all the published alleles, 92% had allele designations, which were used to cross-reference with the CGC database. This revealed that most alleles (70%) have not been deposited with the CGC and that 42% of CGC alleles are not yet described in the literature, although this number may decrease with a more exhaustive literature survey. We next examined the proportions of genes that had multiple tagged alleles. Most genes (63.1%) were represented by a single allele, 19.7% by two, 8.3% by three, while the remaining 8.9% have been tagged in four or more alleles (Figure 1C). In some cases, these reflect different tags or insertion sites designed for specific applications (e.g., dual-color imaging, biochemical assays). In others, they represent redundant alleles generated independently by different labs (e.g. 177 genes have two or more alleles tagged with similar fluorophores by different labs).
The most frequently used tags were fluorophores and epitopes (Figure 1D). Among these, GFP (44.0%), mNeonGreen (17.0%), and FLAG tags (30.1%) were most common, with the majority placed at the C-terminal end of the protein (62.9%), followed by the N-terminus (33.7%), while only a minority (2.9%) were internally inserted (Figure 1E). This bias at the C-terminus likely reflects the biological and practical advantages of placing tags at the protein’s C-terminus, such as reduced chance of disruption of regulatory elements near the transcriptional start site, and that N-terminal SEC insertions can be challenging to create as they behave as loss-of-function alleles prior to excision (Dickinson et al. 2015).
Taken together, these observations reveal that to date, without a centralized effort, ~8% of the C. elegans proteome has been tagged. In addition, more than 10% of these proteins have been redundantly labelled. These results indicate the feasibility and progress towards a complete proteome, but also underscore the need for a centralized tracking system and communication of available strains between researchers for greater efficiency.
2. The Current Tagging Landscape
To better understand how existing tagging efforts have been distributed across the C. elegans genome, we compared the set of 1,554 tagged genes to the genome-wide background of 14,956 protein-coding genes annotated with at least one GO term. Overall, 10.0% of genes with GO terms have been tagged (1,495 genes), compared with 7.8% of all genes. To identify representation within broader GO categories, GO terms were clustered by semantic similarity, reducing 6,800 individual terms into 728 broader parent categories (414 Biological Processes, 197 Molecular Functions, and 117 Cellular Compartments) that consolidate related functions and gene sets. We then performed GO term enrichment analysis to determine if there was overrepresentation or underrepresentation of tagged genes among functional categories (Figure 2A). Relatively few functional categories were underrepresented for endogenously tagged alleles, whereas many categories showed strong enrichment.
Figure 2 |. Functional and disease enrichment patterns reveal selective biases in endogenous tagging efforts.

A) Volcano plot showing Gene Ontology (GO) term enrichment analysis. Each point represents a GO parent category following clustering by semantic similarity (728 total categories). B) Human disease-associated gene enrichment analysis showing log2(odds ratio) for disease categories significantly enriched or depleted among tagged genes. Numbers in parentheses indicate (tagged genes/total genes) in each disease category.
The depleted set of parent terms skews toward membrane-associated and enzyme-heavy categories. Broad membrane terms are underrepresented (membrane 390/6508, 6.0%), with low coverage in organelle interiors (mitochondrial matrix 7/177, 4.0%) and G Protein-Coupled Receptor (GPCR) functions (GPCR activity 6/484, 1.2%; peptide GPCR 1/208, 0.48%). Glycosylation and redox pathways are similarly underrepresented (glycosyltransferase activity 2/297, 0.7%; hexosyltransferase activity 0/75, 0%; oxidoreductase activity 27/497, 5.4%). This pattern is consistent with the difficulty of tagging multi-pass transmembrane and organelle-targeted proteins without disrupting folding, trafficking, import, or activity (Harner et al. 2011; Soave et al. 2021).
The enriched GO categories among endogenously tagged genes are likely due to a convergence of scientific interest and where fluorescent protein tagging has been technically feasible. Cytoskeletal and chromosome-linked components are highly represented (microtubule cytoskeleton 69/133, 51.9%; chromosome 154/314, 49.0%), in line with longstanding interest in spindle assembly, division, and morphogenesis (e.g. Honda et al. 2017; Nishida et al. 2021; Xu et al. 2024). Collagen-containing extracellular matrix also stands out (25/29, 86.2%), consistent with extensive studies of cuticular collagens, basement-membrane dynamics, and tissue remodeling (Keeley et al. 2020; Jayadev et al. 2022; Adams et al. 2023; Ragle et al. 2025). Broader enriched categories include RNA regulation (ribonucleoprotein granule 92/164, 56.0%; regulatory ncRNA-mediated gene silencing 62/217, 28.6%), protein-protein and protein-DNA interactions (protein binding 641/2856, 22.4%; nucleic acid binding 372/1308, 28.4%, DNA-binding transcription factor activity 177/624, 28.4%) and developmental biology terms including cell differentiation (177/453, 39.1%), multicellular development (248/620, 40.0%), larval development (141/345, 40.9%), and neurogenesis (57/107, 54.3%).
Human disease-associated genes are also enriched, with 18.0% tagged (813/4517) compared to 7.8% of all genes. Enrichment analysis of specific disease categories revealed strong overrepresentation of tagged genes in several rare disease groups, including tubulinopathy, polymicrogyria, ptosis, facial paralysis, and multiple lissencephaly subtypes (Figure 2B). This is a direct result of these diseases being driven by defects in microtubule biology, which is a well-tagged protein class. Atrial heart septal defect is likewise enriched as myosins (unc-54, myo-2) and homologs of transcriptional regulators involved in heart development (e.g. ceh-28, elt-1) have been tagged. Finally, multiple cancers, some of which encompassed larger gene sets (e.g., colorectal and prostate cancer), appear due to their overlap with broad enriched classes, including transcriptional and cell-cycle regulators. On the other hand, while many diseases have very few or zero tagged genes, only four categories - heroin dependence, opiate dependence, liver disease, and P. falciparum malaria - were significantly depleted for tagged genes. The genes associated with these diseases are strongly enriched for metabolic and detoxification GO terms, including xenobiotic metabolism, monooxygenase and oxidoreductase activity, and heme/iron binding, reflecting that components in these pathways remain almost entirely untagged in C. elegans.
3. Best Practices and Tagging Considerations
An important element for accelerating tagging of the C. elegans proteome is establishing a pipeline to guide researchers in designing, generating, validating, and sharing endogenously tagged alleles (Dickinson et al. 2015; Nance and Frøkjær-Jensen 2019). Below, we concisely review the best practices for each step of the process from choosing protein tags and their insertion locations to allele validation.
A critical first step in all protein tagging efforts is a detailed analysis of the protein’s structure, functional domains, and post-translational modifications. Tags inserted near enzymatic active sites, binding motifs, regulatory domains, or transmembrane segments can impair protein activity, stability, or localization. Likewise, motifs such as nuclear localization or export signals, signal peptides, residues subject to N-terminal processing (e.g., initiator methionine removal or myristoylation, signal peptide cleavage), and internal proteolytic cleavage sites must be carefully considered, as tagging in these regions can disrupt essential functions or result in the tag being removed during processing (Tian et al. 2004; Huang et al. 2010). Importantly, “N-terminal tagging” need not mean placing the tag at the initiating methionine. For proteins with a cleaved N-terminal signal peptide, the tag should be positioned immediately downstream of the cleavage site, effectively generating an N-terminally tagged mature protein while preserving signal peptide function (Keeley et al. 2020). Internal tagging is also a viable strategy when the region of interest is structurally permissive and shared across isoforms. This can involve inserting the tag into an internal exon, or be inserted as part of a synthetic exon that includes minimal splice sites (e.g. 5’ CAG, 3’ GTA; Lim and Burge 2001) engineered into an intron of the target gene (Keeley et al. 2020). Bioinformatic tools such as AlphaFold, UniProt, and Pfam provide valuable structural and domain annotations, including catalytic cores, domain boundaries, signal peptides, and intrinsically disordered regions. Disordered regions and surface-exposed internal loops often provide flexible and permissive sites for tag insertion. Combining these annotations with quantitative features such as how surface exposed each amino acid position is (relative solvent accessibility), and the evolutionary sequence conservation (indicating functional regions) can reveal permissive tagging locations (Xu et al. 2024; Zinski et al. 2024).
The diversity of transcript isoforms presents an additional layer of complexity. Alternative transcription start sites or splicing events may alter N- or C-terminal coding sequences, subcellular targeting signals, or expression profiles. In many cases, isoforms arise from alternative transcription start sites or exon skipping near the 5′ end of the gene, meaning that internal or C-terminal tagging is more likely to label all isoforms simultaneously (Ramani et al. 2011; Weinreb et al. 2024). However, this generalization does not hold universally, and the isoform structure of each gene should be carefully assessed using available gene models or transcriptomic data. When possible, tags should be inserted into exons shared by all biologically relevant isoforms or the canonical or most highly expressed isoform unless a specific variant is the intended target.
Practical considerations during CRISPR-mediated knock-in can also influence tag placement. For example, in strategies that rely on the self-excising cassette (SEC) system (Dickinson et al. 2015), N-terminal tagging involves an initial insertion that separates the promoter and target gene coding sequence, producing a knockout allele prior to selection cassette excision. This step can be problematic for essential genes, where loss of function is lethal or causes selection bottlenecks. In such cases, C-terminal tagging is often more feasible, as animals will carry the desired insertion without passing through a strong loss-of-function intermediate generation. However, a recently reported SEC variant (NSEC) avoids this N-terminal knockout intermediate by embedding the selection cassette in a reverse orientation within an artificial intron that is spliced out of the target gene’s mRNA, preserving native gene function throughout strain construction (Gibney and Pani 2025). This innovation removes this common barrier to N-terminal endogenous tagging.
Once the optimal location for tag insertion has been established, tag selection is another important feature to be considered. Bright fluorescent proteins such as GFP, mNeonGreen, and mScarlet are widely used for live-cell imaging of protein localization and dynamics, but their relatively large size (~27 kDa) can hinder proper folding, trafficking, or function, especially when tagging small, highly structured, or sterically restricted proteins such as ribosomal subunits or transcription factors (Noma et al. 2017; Costa et al. 2023). In such cases, split fluorophore systems provide a valuable alternative (Cabantous and Waldo 2006). For example, tagging with the small 1.8kDa sfGFP(11) peptide, combined with tissue-specific expression of the complementary 24kDa sfGFP(1–10) fragment, minimizes the size of the inserted tag while enabling spatially restricted reconstitution of fluorescence (Hefel and Smolikove 2019; Costa et al. 2023). Animal viability can also be affected by the fluorophore used in tagging, likely due to differences in fluorophore protein folding rates (Srinivasan et al. 2025).
For high-resolution or multiplexed imaging, self-labeling enzyme tags such as HaloTag (Los et al. 2008) and SNAP-tag (Keppler et al. 2004) enable covalent attachment of synthetic dyes with increased brightness, photostability, spectral range, and labeling speed compared to genetically encoded fluorescent proteins (D. Zhang et al. 2023; De Luis et al. 2025). These tags offer exceptional flexibility by allowing the user to choose from a wide array of small-molecule ligands conjugated to custom fluorophores, making them especially valuable for applications requiring high signal-to-noise and optical precision, including super-resolution microscopy and single-molecule tracking (Kompa et al. 2023; Catapano et al. 2025), time-lapse imaging, and pulse-chase experiments (Borchers et al. 2025). These self-labeling tags can also be paired with specialized dyes to enable FRET-based detection of protein interactions or conformational changes, sensing of local pH, redox state, or ion concentrations (Cook et al. 2023). These capabilities make HaloTag and SNAP-tag highly versatile tools for precisely describing dynamic protein behaviors.
Smaller epitope tags, such as FLAG, HA, and V5, are widely used for biochemical applications, including Western blotting, immunoprecipitation (IP), chromatin immunoprecipitation (ChIP), and affinity purification, and also fluorescence immunostaining (Terpe 2003; Dickinson et al. 2015; De Luis et al. 2025). These tags consist of short amino acid sequences (typically 5–30 residues) that are recognized by specific antibodies. Their compact size minimizes (but does not eliminate, e.g., Barker et al. 2025) the likelihood of interfering with protein folding, localization, or function, even when multiple copies are inserted in tandem to increase detection sensitivity (Stowers 2025). Having multiple distinct epitopes available, each recognized by highly specific antibodies, is advantageous for multiplexed experiments in which different proteins are tagged and detected simultaneously, or for tandem immunoaffinity purification (DeCaprio and Kohl 2019; De Luis et al. 2025).
In addition to tagging endogenous proteins for visualization or biochemical studies, many strategies combine endogenous tagging of proteins with conditional control elements to enable temporal or spatial manipulation of the native protein. One widely used approach is the auxin-inducible degron (AID) system, which allows tissue-specific and temporally controlled protein depletion by tagging the protein of interest with a minimal degron sequence and co-expressing the plant-derived TIR1 F-box protein (Zhang et al. 2015; Ashley et al. 2021). Upon auxin treatment, TIR1 recruits the SCF E3 ligase complex, leading to ubiquitination and proteasomal degradation of the degron-tagged target. A conceptually similar strategy that requires no exogenous small molecule treatment is the ZIF-1/ZF1 system, in which a short ZF1 degron tag is recognized by the endogenous ZIF-1 adaptor protein, which likewise recruits an E3 ligase complex (Armenti et al. 2014). Nanobodies can also be used to degrade proteins tagged with epitopes, such as a GFP-nanobody fused to ZIF-1 that degrades GFP-tagged proteins (Wang et al. 2017) and NbALFA fused to E3 ubiquitin ligases that degrades ALFA-tagged proteins (Yang et al. 2025). Although not formally a protein tag, knock-in protease-cleavable sites in which an internal site for a sequence-specific protease (e.g., 7aa TEV protease site) is engineered into the protein allow for inducible cleavage and inactivation upon protease expression (Harder et al. 2008; Das et al. 2024). Conditional expression of tagged proteins from the native locus can also be achieved using FRT-based recombinase systems, in which FLP is used to fuse a protein tag in a tissue-specific or inducible manner (Muñoz-Jiménez et al. 2017).
To preserve native protein function, tags should be connected to the protein of interest via a flexible linker, typically composed of small, uncharged residues such as glycine and serine. These glycine-serine (Gly-Ser) repeats provide rotational freedom and reduce structural interference between the tag and adjacent protein domains. Common linkers include (Gly-Gly-Gly-Gly-Ser)1–3 motifs (i.e., GGGGS, repeated 1–3 times), which have been widely adopted due to their flexibility, hydrophilicity, and low immunogenicity (Chen et al. 2013). The optimal linker length depends on several factors, including the size and rigidity of the tag, the structural constraints of the target protein, and the location of the insertion site (e.g., near a folded domain vs. in a disordered region). In C. elegans, the 9aa flexlink (Gly-Ala-Ser)3 motif was designed in the original SEC constructs (Dickinson et al. 2015), but recent studies have found that increasing this length to 18aa (Gly-Ala-Ser)6 (Keeley et al. 2020) or 30aa (Gibney et al. 2025) improves viability when tagging particular genes.
Following strain generation, initial functional checks should confirm that the knock-in does not appear to interfere with the endogenous protein’s function. Animals should be monitored for developmental, behavioral, and fertility defects by measuring growth rates of individual worms or time to plate starvation after plating a set number of worms on a defined amount of food (Srinivasan et al. 2025). However, this approach might not identify subtle phenotypes. An absence of signal or unexpected protein localization may also indicate problems such as disrupted tagged protein function, cleavage of the tag from the protein, disrupted splicing, high turnover of the protein, or low tagged protein levels. Western blotting can help determine cleavage of the tag, but an absence of signal can be difficult to interpret (Keeley et al. 2020; Srinivasan et al. 2025).
Equally important is comprehensive documentation of the tagging design, including the tag and linker sequences, precise genomic insertion site, tag orientation, and validation results, both positive and negative. Depositing this information in public repositories (e.g., WormBase, CGC, WormTagDB) maximizes reproducibility and enables other researchers to build on prior work while helping to identify cases in which specific tags or insertion sites subtly interfere with protein function.
4. The Endogenous Tagging Database – WormTagDB
To support community-wide access to the curated dataset of endogenously tagged C. elegans genes, we developed WormTagDB, an interactive web platform available at https://wormtagdb.rc.duke.edu. Built using the RShiny framework (Chang et al. 2025), the database enables researchers to browse, filter, and download tagged allele data through a simple interface. Each entry includes gene name, allele designation, tag type, tag position, source (literature or CGC), and publication reference (when available), thus facilitating direct communication and reagent sharing between labs.
While the CGC plays an essential role in strain preservation and distribution, it is not feasible for the CGC to archive every published strain. Thus, many strains remain housed in individual labs and may not be formally deposited or easily discoverable. WormTagDB addresses this gap by consolidating tagged allele metadata from both the literature and CGC records, thus offering a centralized, up-to-date resource to facilitate direct lab-to-lab sharing of strains.
To ensure researchers can easily access all experimentally relevant details of the tagged alleles, WormTagDB allows users to search for specific genes, alleles, or tags, and to filter entries by tag type (e.g., GFP, mCherry, FLAG), tag position (N-terminal, C-terminal, internal), and source (literature vs CGC). In addition, users can search for GO terms of interest (e.g. “basement membrane” or “autophagy”) to retrieve all tagged alleles of genes annotated with those terms. Results are displayed in an interactive table, and users can download the full or filtered databases as CSV files. To identify systematic functional trends in the current tagging landscape, WormTagDB also includes a GO term enrichment analysis module. This feature compares the functional annotations of tagged genes against a genome-wide background, highlighting biological processes, molecular functions, and cellular components that are either over- or under-represented among tagged loci. This analysis can help researchers identify saturated categories as well as areas where tagged alleles are still lacking.
To accelerate CRISPR-based endogenous tagging in C. elegans, WormTagDB provides ready-to-use reagent designs for 99.6% of protein-coding genes: predicted N- and C-terminal insertion coordinates (WBcel235), sgRNA sequences, and primer sequences to amplify homology arms (Supplementary Figure 2, Supplementary Table 3). Using a chain-aware pipeline that accounts for post-translational processing (initiator Met removal, signal/transit/pro-peptides, lipidation sites, and curated chain boundaries in Uniprot), we delineated the mature polypeptide. N-terminal insertion sites were positioned immediately downstream of all predicted N-terminal processing and lipidation sites, while C-terminal sites were positioned immediately upstream of C-terminal processing and lipidation sites. Lower-confidence sequence motif predictions (PTS1/PTS2, KDEL, di-Lys, CAAX) were reported as warning flags but do not influence site selection. We selected optimal guide sequences with the minimal number of off-target sequences and with cut sites closest to the target insertion location (Supplementary Figure 2). We designed primers to amplify the flanking homology arms required for homology-directed repair, introducing silent mutations as needed to prevent Cas9 from cutting the repair template. An earlier version of this pipeline was used to successfully generate 167 tagged strains (Supplementary Table 4). This comprehensive resource of the sequence reagents to tag 19,886 genes is accessible in the “Guide Predictions” page of WormTagDB with an interactive IGV viewer to visualize the sequence locations and orientations, substantially lowering the barrier for entry to high-throughput tagging experiments.
WormTagDB also includes a community submission form to encourage direct contributions from researchers. Users can report newly generated endogenously tagged alleles along with associated metadata such as tag type, insertion site, linker sequence, and validation information. Using the same form, attempted knock-ins that were non-viable can also be reported. Submitted entries will be incorporated into the database to expand and update the resource, ensuring it reflects the most current tagging efforts. To help coordinate ongoing work, the site also features a “reserve genes” form. This allows researchers to indicate alleles they plan to generate, optionally including their lab and contact information to be displayed in the database. By making planned projects more visible, this feature helps prevent unintentional duplication of effort, enables potential collaborators to connect early, and fosters a more coordinated, efficient approach toward systematically tagging the C. elegans proteome.
5. A Roadmap for Whole-Proteome Tagging
Since the advent of CRISPR-Cas9-mediated genome editing, the C. elegans community has made steady progress toward tagging the complete proteome. To date, 1,554 unique genes (~8% of the proteome) have been tagged, with new additions appearing at a rate of nearly 200 genes per year (Figure 1A). At this rate, the complete proteome will not be tagged for approximately 100 years. A coordinated, large-scale effort is thus needed to accelerate gene tagging. This should include laboratories sharing protocols, developing communication channels, and tracking projects to increase efficiency and train new research groups to increase participation. Individual labs can then align efforts toward the common goal of a complete tagged proteome, while maintaining flexibility to pursue their own scientific questions.
To facilitate this large-scale tagging effort, we have provided predictions of appropriate N- and C-terminal knock-in sites along with proximal guide RNA sequences and homology arm primers for all C. elegans protein-coding genes on WormTagDB (Supplementary Table 3). This resource greatly streamlines the first step in knock-in generation - the manual design of these sequence reagents. This makes endogenous tagging more accessible to the broader C. elegans research community and will enable high-throughput tagging campaigns that would otherwise be more time-consuming.
The time required for microinjection can be another significant bottleneck in worm genome engineering. Manual injection typically requires 1–2 hours per construct, while a robotic microinjector described by Pan et al. (2024) reduces this time to ~15 minutes. Further, improved reagent mixes (Paix et al. 2015) can decrease the number of required injections, pushing efficiency even higher. With moderate uptake of robotic injection at institutions, the pace and accessibility of proteome-wide tagging could be accelerated.
For systematic proteome tagging, we propose that the self-excising cassette (SEC) strategy should be the community standard. In this approach, animals carrying the insertion are selected by drug resistance and roller phenotypes encoded on the cassette, which is subsequently removed by heat-shock induction (Dickinson et al. 2015; Gibney and Pani 2025). Although SEC requires ~2–3 weeks and multiple generations to obtain homozygous, scar-free alleles, it has several decisive advantages over Co-CRISPR for large-scale projects. These include a simple workflow that does not rely on extensive PCR screening and fewer person-hours, allowing greater overall efficiency. For groups that have not previously generated knock-ins, SEC also provides a straightforward entry point into genome editing with clear phenotypic markers and a robust, well-documented workflow.
Another important element is to standardize shared donor backbones and protein tags. As many SEC donor backbones are readily available from Addgene (e.g., Dickinson et al. 2015; Gibney and Pani 2025), backbone design is already largely standardized across the community. A remaining source of variability is fluorophore tag choice. To enable quantitative, cross-laboratory comparisons with shared strains, the community should converge on a core set (1–2 per color) of standardized fluorophores (blue, green, red, and far-red) and phase out legacy tags such as GFP and mCherry. For example, mNeonGreen and mStayGold are brighter than GFP and are also fast-folding and photostable (Ko and Mizumoto 2025). For red fluorescence, mScarlet-I3 provides excellent performance and folding time (Cao et al. 2024), while new blue proteins such as mTagBFP2 or Electra (Papadaki et al. 2022) and far-red options such as miRFP680 or miRFP713 (H. Zhang et al. 2023) allow straightforward multiplexing of up to four colors with commonly-used filter sets. Beyond fluorophores, attaching an additional small epitope, such as 3xFLAG, would enable biochemical isolation of all tagged proteins (e.g. Dickinson et al. 2015). We suggest the ALFA tag should be tested further in C. elegans, as it displays higher affinity than FLAG (Götzke et al. 2019). The higher affinity also allows 1xALFA or 2xALFA configurations (Igreja et al. 2022), thus avoiding repetitive sequences that can cause problems during construct assembly. While covering routine biochemistry applications, ALFA also allows nanobody-based relocalization (Xu et al. 2022), degradation (Liu et al. 2024; Yang et al. 2025), and fluorescence or super-resolution imaging (Westlund et al. 2023; Quintin et al. 2025).
We propose WormTagDB as a coordination hub to register planned and newly generated alleles. This will allow researchers to avoid duplicate tagging and facilitate strain sharing and tagging designs. Each entry captures the essential metadata (gene, tag position, fluorophore, etc). Crucially, negative results, such as failed tag positions and problematic linkers, can also be recorded. This registry will improve project planning and make strain sharing routine. WormTagDB can also be used by Undergraduate Research Experiences courses (CUREs) for tagging projects (Hastie et al. 2019; McDonald et al. 2023; Herrera Sandoval et al. 2024).
With standardized tags, a shared set of donor backbones, and a common registry, WormTagDB initiates a practical, scalable path to tag the C. elegans proteome. As tagged allele inventories continue to grow, large-scale imaging pipelines will be able to process whole-worm datasets using automated segmentation, registration, and representation learning techniques to extract organ- and cell-level localization and dynamics. These maps can then be integrated with single-cell expression and perturbation data into a functional protein atlas. Because the worm is transparent and genetically tractable, C. elegans was the first multicellular animal to have its entire developmental cell lineage mapped and its genome fully sequenced (Corsi et al. 2015). The worm is now poised to become the first metazoan to possess a comprehensive resource of endogenously tagged alleles to reveal the localization and functions of the entire proteome.
Methods
To identify endogenously tagged C. elegans alleles described in the literature, a keyword-based search was performed using WormBase's Textpresso Central portal (https://wb-textpresso.alliancegenome.org/tpc/search), updated as of September 1st, 2025. The search included the following terms commonly associated with endogenous genome editing: “knock-in”, “knock in”, “tagged”, “CRISPR”, “gRNA”, “homology arm”, and “repair template”. This search returned 11,015 papers, each assigned a keyword relevance score by the Textpresso engine. The top 2,500 papers were manually reviewed in descending order of score, including both main text and supplementary materials. Of these, 778 were found to describe the generation of one or more novel endogenously tagged alleles. For each paper, all newly generated alleles representing the insertion of protein tags at native genomic loci were recorded. Endogenously tagged strains obtained from other sources, such as previously published studies or external collaborators, were excluded unless the paper described their de novo generation. In cases where multiple functionally identical alleles were reported for the same tagged gene (e.g., independent insertions of gene-1::GFP from separate microinjections or founders), only one representative allele was recorded to avoid redundancy. In rare cases, specific alleles were described in more than one publication, and in these cases, only the first publication was kept. Constructs introduced via MosSCI, multicopy transgenes, or overexpression arrays were excluded from the analysis. Also excluded were self-cleaving transcriptional reporters (e.g. T2A, SL2) and endogenously tagged mutant loci. This literature-curated allele list was merged with the set of endogenously tagged alleles available through the Caenorhabditis Genetics Center (CGC), based on CGC strain records as of September 1st, 2025. The final dataset represents a manually curated, up-to-date inventory of endogenously tagged C. elegans alleles as of this date.
GO annotations for C. elegans protein-coding genes were obtained from Ensembl BioMart and evidence codes from org.Ce.eg.db (v3.20.0) (Carlson 2017). GO term similarity and hierarchical clustering were performed using the rrvgo (v1.18.0) Bioconductor package (Sayols 2023), with semantic similarity calculated via the Wang method and a similarity threshold of 0.7 for clustering related BP and MF terms and 0.6 for CC terms. The GO terms in each cluster with the largest count of Entrez Gene identifiers annotated to that GO term or to its child nodes in the ontology were selected as the parent term for the cluster. Enrichment was assessed by calculating the proportion of tagged genes per GO term and comparing it to the proportion expected if genes were tagged at random across the genome. GO terms were grouped by parent category to reveal higher-order functional trends in tagging coverage. Disease ontology (DO) annotations were obtained from Alliance of Genome Resources v8.1.0 (https://www.alliancegenome.org/downloads). The number of tagged genes in each GO term or disease category was compared to the total number of C. elegans genes associated with that GO term/disease to calculate an odds ratio, which was log2-transformed for visualization. Statistical significance was determined using Fisher’s exact test with a false discovery rate (FDR) threshold of 0.05. For visualization, a log2(odds ratio) cutoff of 7 and −7 was used in place of infinite values to allow GO terms/diseases with completely tagged or untagged gene sets to be plotted on the same scale.
To make the curated dataset accessible to the broader research community, we developed the interactive web application WormTagDB using the RShiny framework. The app allows users to browse, search, and filter all identified endogenously tagged alleles by gene, tag type, tag position, GO terms, and human disease associations. Each entry includes the gene name, tag, allele designation, source, and strain information where available. Users can also download the complete dataset and submit new alleles through an integrated community submission form. The app is publicly available at https://wormtagdb.rc.duke.edu and will be updated regularly to reflect newly published alleles. The code is available at https://github.com/jakeleyhr/WormTagDB.
To design genome-wide N- and C-terminal tagging sites and corresponding CRISPR guides, we created custom Python scripts (https://github.com/jakeleyhr/CRISPR-Guide-and-Primer-Design-Pipeline-for-C.-elegans). We began with a list of all protein-coding C. elegans genes and selected the Ensembl canonical transcript and corresponding UniProt ID, and inferred mature protein boundaries in a “chain-aware” manner by parsing curated features (initiator methionine removal, signal peptides, propeptides, transit peptides, lipidation sites, and annotated chain regions). The N-terminal insertion site was positioned immediately downstream of all predicted N-terminal processing events (initiator-Met removal and any signal, transit, or pro-peptide segments), and the C-terminal site was positioned immediately upstream of residues expected to be removed. Insertion sites were positioned 5aa proximal to any terminal lipidation sites. Simple motif searches for PTS1, PTS2, KDEL, di-Lysine, and CAAX motifs were also performed, but as these were of lower confidence, they did not influence insertion site choice and instead were listed as potential warning flags. N- and C-terminal insertion coordinates were defined at the nucleotide immediately 5′ of the relevant codon boundary in transcript orientation by converting amino-acid indices to CDS positions (WBcel235). For each site, ±50 bp of genomic sequence was retrieved, and both strands were scanned for SpCas9 NGG targets, evaluating candidates as 20bp protospacer + NGG PAM with cut positions at the expected offset 3bp upstream of the PAM site. Candidates were filtered for composition: GC between 25–80%, no 7+ bp long homopolymers, and divided into guides suitable for “in vivo transcription” or “in vitro transcription only” based on the absence/presence of 4+ consecutive Ts. Candidates with GGG PAM sites or NGG PAM sites where the next non-guide base was G (NGG+G) were also excluded. Potential off-target cut sites were identified using FlashFry (McKenna and Shendure 2018). Guides were ranked by the minimal number of off-target sequences and then by the proximity of the cut site to the intended insertion coordinate. The best single “in vivo” guide was selected for each target site, with an additional “in vitro” guide reported if it had a cut site closer to the insertion site.
For each N- or C-terminal target locus, we designed primers to amplify the left (Primer 1 and Primer 2) and right (Primer 3 and Primer 4) homology arms used to flank the knock-in cassette. We first extracted ±100 bp of genomic sequence surrounding the predicted insertion site, and Primers 2 and 3 were directly anchored to the genomic sequence immediately adjacent to the insertion site: Primer 2 comprised the 30–35 bp immediately upstream, and Primer 3 the 30–35 bp immediately downstream. 4-to-5 silent mutations were introduced at codons in the primers corresponding to the 3′ end of the guide and PAM site to prevent cutting of the vector or repaired allele. When a guide cut site was located far enough from the insertion site that not all required silent mutations could be placed within Primers 2 or 3 while still retaining at least 15 unmodified bases at the 3′ end, an overlapping two-primer design was used. In these cases, auxiliary primers (Primers 2A or 3A) containing the necessary silent mutations were generated with ~20 bp overlaps to the modified Primers 2 or 3 (renamed Primers 2B or 3B).
Primers 1 and 4 were then identified using the Primer3 program (Untergasser et al. 2012) to pair appropriately with Primers 2 and 3, respectively, with the goal of generating homology arm amplicons between ~500–1000 bp and with closely matched annealing temperatures. If no valid pairing was found within the default search window, the maximum allowable arm length was incrementally expanded up to 10kb. All candidate Primer 1 and 4 sequences were subjected to BLAT searches against the C. elegans genome, and when possible, only primers without predicted off-target matches were selected. Finally, genotyping primers flanking the entire edited region were designed using Primer3, again using BLAT searches on candidates to ensure high specificity. These primers can also be used in a two-step workflow in which the locus is first amplified by genotyping PCR, followed by nested or overlapping PCR to obtain each homology arm.
Supplementary Material
Acknowledgements
We thank Ariel Pani and Daniel Dickinson for their comments on the manuscript and encouragement from Geraldine Seydoux.
Funding
D.R.S is funded by the National Institutes of Health (R35GM118049), the Wellcome Trust (226804/Z/22/Z), by FourPoints Innovation, a portfolio company of certain funds managed by Deerfield Management Company, L.P., and Duke University discretionary funds. W.Z. is funded by the National Natural Science Foundation of China grants 32370822 and 31970919.
Funding Statement
D.R.S is funded by the National Institutes of Health (R35GM118049), the Wellcome Trust (226804/Z/22/Z), by FourPoints Innovation, a portfolio company of certain funds managed by Deerfield Management Company, L.P., and Duke University discretionary funds. W.Z. is funded by the National Natural Science Foundation of China grants 32370822 and 31970919.
Data Availability
All data necessary for confirming the conclusions presented in the article are represented fully within the article figures and supplementary information. The data is viewable at the WormTagDB website (https://wormtagdb.rc.duke.edu), and all relevant code is available on GitHub (https://github.com/jakeleyhr/WormTagDB and https://github.com/jakeleyhr/CRISPR-Guide-and-Primer-Design-Pipeline-for-C.-elegans).
References
- Adams JRG et al. 2023. Nanoscale patterning of collagens in C. elegans apical extracellular matrix. Nat Commun. 14(1):7506. 10.1038/s41467-023-43058-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Armenti ST, Lohmer LL, Sherwood DR, Nance J. 2014. Repurposing an endogenous degradation system for rapid and targeted depletion of C. elegans proteins. Development. 141(23):4640–4647. 10.1242/dev.115048 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Arribere JA et al. 2014. Efficient Marker-Free Recovery of Custom Genetic Modifications with CRISPR/Cas9 in Caenorhabditis elegans. Genetics. 198(3):837–846. 10.1534/genetics.114.169730 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ashley GE et al. 2021. An expanded auxin-inducible degron toolkit for Caenorhabditis elegans Greenstein D, editor. Genetics. 217(3):iyab006. 10.1093/genetics/iyab006 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Barker B et al. 2025. Epitope tags are not created equal: Disruption of cellular function of a translation factor by a short viral tag. MicroPublication Biol. [published online ahead of print]. 10.17912/micropub.biology.001452 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Barrett PL, Fleming JT, Göbel V. 2004. Targeted gene alteration in Caenorhabditis elegans by gene conversion. Nat Genet. 36(11):1231–1237. 10.1038/ng1459 [DOI] [PubMed] [Google Scholar]
- Borchers C, Osburn K, Roh HC, Aoki ST. 2025. In vivo pulse-chase in Caenorhabditis elegans reveals intestinal histone turnover changes upon starvation. J Biol Chem. 301(7):110299. 10.1016/j.jbc.2025.110299 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Broverman S, MacMorris M, Blumenthal T. 1993. Alteration of Caenorhabditis elegans gene expression by targeted transformation. Proc Natl Acad Sci. 90(10):4359–4363. 10.1073/pnas.90.10.4359 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cabantous S, Waldo GS. 2006. In vivo and in vitro protein solubility assays using split GFP. Nat Methods. 3(10):845–854. 10.1038/nmeth932 [DOI] [PubMed] [Google Scholar]
- Cao WX et al. 2024. Comparative analysis of new mScarlet-based red fluorescent tags in Caenorhabditis elegans. Genetics. 228(2):iyae126. 10.1093/genetics/iyae126 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Carlson M. 2017. org.Ce.eg.db: Genome wide annotation for Worm.
- Catapano C et al. 2025. Long-Term Single-Molecule Tracking in Living Cells using Weak-Affinity Protein Labeling. Angew Chem Int Ed. 64(1):e202413117. 10.1002/anie.202413117 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chang W et al. 2025. shiny: Web Application Framework for R. https://shiny.posit.co/
- Charlesworth AG et al. 2021. Two isoforms of the essential C. elegans Argonaute CSR-1 differentially regulate sperm and oocyte fertility. Nucleic Acids Res. 49(15):8836–8865. 10.1093/nar/gkab619 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen X, Zaro JL, Shen W-C. 2013. Fusion protein linkers: Property, design and functionality. Adv Drug Deliv Rev. 65(10):1357–1369. 10.1016/j.addr.2012.09.039 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cho NH et al. 2022. OpenCell: Endogenous tagging for the cartography of human cellular organization. Science. 375(6585):eabi6983. 10.1126/science.abi6983 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cook A, Walterspiel F, Deo C. 2023. HaloTag-Based Reporters for Fluorescence Imaging and Biosensing. ChemBioChem. 24(12):e202300022. 10.1002/cbic.202300022 [DOI] [PubMed] [Google Scholar]
- Corsi AK, Wightman B, Chalfie M. 2015. A Transparent Window into Biology: A Primer on Caenorhabditis elegans. Genetics. 200(2):387–407. 10.1534/genetics.115.176099 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Costa DS et al. 2023. The Caenorhabditis elegans anchor cell transcriptome: ribosome biogenesis drives cell invasion through basement membrane. Development. 150(9):dev201570. 10.1242/dev.201570 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Das M et al. 2024. Condensin I folds the Caenorhabditis elegans genome. Nat Genet. 56(8):1737–1749. 10.1038/s41588-024-01832-5 [DOI] [PubMed] [Google Scholar]
- De Luis M, Xu S, Zinn K. 2025. Fluorescent labeling of proteins in vitro and in vivo using encoded peptide tags. J Biol Chem. 301(6):110229. 10.1016/j.jbc.2025.110229 [DOI] [PMC free article] [PubMed] [Google Scholar]
- DeCaprio J, Kohl TO. 2019. Tandem Immunoaffinity Purification Using Anti-FLAG and Anti-HA Antibodies. Cold Spring Harb Protoc. 2019(2):pdb.prot098657. 10.1101/pdb.prot098657 [DOI] [PubMed] [Google Scholar]
- Dickinson DJ et al. 2015. Streamlined Genome Engineering with a Self-Excising Drug Selection Cassette. Genetics. 200(4):1035–1049. 10.1534/genetics.115.178335 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dickinson DJ, Ward JD, Reiner DJ, Goldstein B. 2013. Engineering the Caenorhabditis elegans genome using Cas9-triggered homologous recombination. Nat Methods. 10(10):1028–1034. 10.1038/nmeth.2641 [DOI] [PMC free article] [PubMed] [Google Scholar]
- El Mouridi S et al. 2017. Reliable CRISPR/Cas9 Genome Engineering in Caenorhabditis elegans Using a Single Efficient sgRNA and an Easily Recognizable Phenotype. G3 GenesGenomesGenetics. 7(5):1429–1437. 10.1534/g3.117.040824 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Frøkjær-Jensen C et al. 2008. Single-copy insertion of transgenes in Caenorhabditis elegans. Nat Genet. 40(11):1375–1383. 10.1038/ng.248 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gibney T et al. 2025. FGF-dependent, polarized SOS activity orchestrates directed migration of C. elegans muscle progenitors independently of canonical effectors in vivo. bioRxiv. [published online ahead of print]. 10.1101/2025.04.11.648432 [DOI] [Google Scholar]
- Gibney TV, Pani AM. 2025. Engineering the C. elegans genome with a nested, self-excising selection cassette. bioRxiv. [published online ahead of print]. 10.1101/2025.05.01.651742 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Götzke H et al. 2019. The ALFA-tag is a highly versatile tool for nanobody-based bioscience applications. Nat Commun. 10(1):4403. 10.1038/s41467-019-12301-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Harder B et al. 2008. TEV Protease-Mediated Cleavage in Drosophila as a Tool to Analyze Protein Functions in Living Organisms. BioTechniques. 44(6):765–772. 10.2144/000112884 [DOI] [PubMed] [Google Scholar]
- Harner M, Neupert W, Deponte M. 2011. Lateral release of proteins from the TOM complex into the outer membrane of mitochondria: Lateral release from the TOM complex. EMBO J. 30(16):3232–3241. 10.1038/emboj.2011.235 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hastie E, Sellers R, Valan B, Sherwood DR. 2019. A Scalable CURE Using a CRISPR/Cas9 Fluorescent Protein Knock-In Strategy in Caenorhabditis elegans. J Microbiol Biol Educ. 20(3):70. 10.1128/jmbe.v20i3.1847 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hefel A, Smolikove S. 2019. Tissue-Specific Split sfGFP System for Streamlined Expression of GFP Tagged Proteins in the Caenorhabditis elegans Germline. G3 GenesGenomesGenetics. 9(6):1933–1943. 10.1534/g3.119.400162 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Herrera Sandoval C, Borchers C, Aoki ST. 2024. An effective Caenorhabditis elegans CRISPR training module for high school and undergraduate summer research experiences in molecular biology. Biochem Mol Biol Educ. 52(6):656–665. 10.1002/bmb.21856 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Honda Y et al. 2017. Tubulin isotype substitution revealed that isotype combination modulates microtubule dynamics in C. elegans embryos. J Cell Sci. 130(9):1652–1661. 10.1242/jcs.200923 [DOI] [PubMed] [Google Scholar]
- Huang Y, Wilkinson G, Willars G. 2010. Role of the signal peptide in the synthesis and processing of the glucagon-like peptide-1 receptor. Br J Pharmacol. 159(1):237–251. 10.1111/j.1476-5381.2009.00517.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Igreja C et al. 2022. Application of ALFA-Tagging in the Nematode Model Organisms Caenorhabditis elegans and Pristionchus pacificus. Cells. 11(23):3875. 10.3390/cells11233875 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jayadev R et al. 2022. A basement membrane discovery pipeline uncovers network complexity, regulators, and human disease associations. Sci Adv. 8(20):eabn2265. 10.1126/sciadv.abn2265 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kanca O, Bellen HJ, Schnorrer F. 2017. Gene Tagging Strategies To Assess Protein Expression, Localization, and Function in Drosophila. Genetics. 207:389–412 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Keeley DP et al. 2020. Comprehensive Endogenous Tagging of Basement Membrane Components Reveals Dynamic Movement within the Matrix Scaffolding. Dev Cell. 54(1):60–74.e7. 10.1016/j.devcel.2020.05.022 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Keppler A et al. 2004. Labeling of fusion proteins with synthetic fluorophores in live cells. Proc Natl Acad Sci. 101(27):9955–9959. 10.1073/pnas.0401923101 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ko S, Mizumoto K. 2025. Comparison among bright green fluorescent proteins in C. elegans. MicroPublication Biol. [published online ahead of print]. 10.17912/micropub.biology.001447 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kompa J et al. 2023. Exchangeable HaloTag Ligands for Super-Resolution Fluorescence Microscopy. J Am Chem Soc. 145(5):3075–3083. 10.1021/jacs.2c11969 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leonetti MD et al. 2016. A scalable strategy for high-throughput GFP tagging of endogenous human proteins. Proc Natl Acad Sci. 113(25). 10.1073/pnas.1606731113 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lim LP, Burge CB. 2001. A computational analysis of sequence features involved in recognition of short introns. Proc Natl Acad Sci. 98(20):11193–11198. 10.1073/pnas.201407298 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu Z et al. 2024. Protein degradation by small tag artificial bacterial E3 ligase. bioRxiv. [published online ahead of print] [accessed 2025 Sep 5]. http://biorxiv.org/lookup/doi/10.1101/2024.09.07.611797. 10.1101/2024.09.07.611797 [DOI] [Google Scholar]
- Los GV et al. 2008. HaloTag: A Novel Protein Labeling Technology for Cell Imaging and Protein Analysis. ACS Chem Biol. 3(6):373–382. 10.1021/cb800025k [DOI] [PubMed] [Google Scholar]
- Marston DJ et al. 2016. MRCK-1 Drives Apical Constriction in C. elegans by Linking Developmental Patterning to Force Generation. Curr Biol. 26(16):2079–2089. 10.1016/j.cub.2016.06.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- McDonald K et al. 2023. Using CRISPR knock-in of fluorescent tags to examine isoform-specific expression of EGL-19 in C. elegans. MicroPublication Biol. [published online ahead of print]. 10.22002/MFRV8-KWW12 [DOI] [PMC free article] [PubMed] [Google Scholar]
- McKenna A, Shendure J. 2018. FlashFry: a fast and flexible tool for large-scale CRISPR target design. BMC Biol. 16(1):74. 10.1186/s12915-018-0545-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Muñoz-Jiménez C et al. 2017. An Efficient FLP-Based Toolkit for Spatiotemporal Control of Gene Expression in Caenorhabditis elegans. Genetics. 206(4):1763–1778. 10.1534/genetics.117.201012 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nance J, Frøkjær-Jensen C. 2019. The Caenorhabditis elegans Transgenic Toolbox. Genetics. 212(4):959–990. 10.1534/genetics.119.301506 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nguyen DAH, Phillips CM. 2021. Arginine methylation promotes siRNA-binding specificity for a spermatogenesis-specific isoform of the Argonaute protein CSR-1. Nat Commun. 12(1):4212. 10.1038/s41467-021-24526-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nishida K et al. 2021. Expression Patterns and Levels of All Tubulin Isotypes Analyzed in GFP Knock-In C. elegans Strains. Cell Struct Funct. 46(1):51–64. 10.1247/csf.21022 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Noma K, Goncharov A, Ellisman MH, Jin Y. 2017. Microtubule-dependent ribosome localization in C. elegans neurons. eLife. 6:e26376. 10.7554/eLife.26376 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Paix A, Folkmann A, Rasoloson D, Seydoux G. 2015. High Efficiency, Homology-Directed Genome Editing in Caenorhabditis elegans Using CRISPR-Cas9 Ribonucleoprotein Complexes. Genetics. 201(1):47–54. 10.1534/genetics.115.179382 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pan P et al. 2024. Robotic microinjection enables large-scale transgenic studies of Caenorhabditis elegans. Nat Commun. 15(1):8848. 10.1038/s41467-024-53108-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Papadaki S et al. 2022. Dual-expression system for blue fluorescent protein optimization. Sci Rep. 12(1):10190. 10.1038/s41598-022-13214-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Quintin S, Saad MI, Amann G, Reymann A-C. 2025. In vivo detection of ALFA-tagged proteins in C. elegans with a transgenic fluorescent nanobody. MicroPublication Biol. [published online ahead of print]. 10.17912/micropub.biology.001542 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ragle JM et al. 2025. Multiscale patterning of a model apical extracellular matrix revealed by systematic endogenous protein tagging. bioRxiv. [published online ahead of print]. 10.1101/2025.05.14.653803 [DOI] [Google Scholar]
- Ramani AK et al. 2011. Genome-wide analysis of alternative splicing in Caenorhabditis elegans. Genome Res. 21(2):342–348. 10.1101/gr.114645.110 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Riga A et al. 2021. Caenorhabditis elegans LET-413 Scribble is essential in the epidermis for growth, viability, and directional outgrowth of epithelial seam cells Nance J, editor. PLOS Genet. 17(10):e1009856. 10.1371/journal.pgen.1009856 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Robert V, Bessereau J-L. 2007. Targeted engineering of the Caenorhabditis elegans genome following Mos1-triggered chromosomal breaks. EMBO J. 26(1):170–183. 10.1038/sj.emboj.7601463 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sayols S. 2023. rrvgo: a Bioconductor package for interpreting lists of Gene Ontology terms. MicroPublication Biol. [published online ahead of print]. 10.17912/micropub.biology.000811 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Soave M et al. 2021. Detection of genome-edited and endogenously expressed G protein-coupled receptors. FEBS J. 288(8):2585–2601. 10.1111/febs.15729 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Srinivasan S et al. 2025. A collagen IV fluorophore knock-in toolkit reveals trimer diversity in C. elegans basement membranes. J Cell Biol. 224(6):e202412118. 10.1083/jcb.202412118 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stowers RS. 2025. Multimerized epitope tags for high-sensitivity protein detection. G3 GenesGenomesGenetics. 15(6):jkaf070 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Terpe K. 2003. Overview of tag protein fusions: from molecular and biochemical fundamentals to commercial systems. Appl Microbiol Biotechnol. 60(5):523–533. 10.1007/s00253-002-1158-6 [DOI] [PubMed] [Google Scholar]
- Tian G-W et al. 2004. High-Throughput Fluorescent Tagging of Full-Length Arabidopsis Gene Products in Planta. Plant Physiol. 135(1):25–38. 10.1104/pp.104.040139 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Untergasser A et al. 2012. Primer3—new capabilities and interfaces. Nucleic Acids Res. 40(15):e115–e115. 10.1093/nar/gks596 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang S et al. 2017. A toolkit for GFP-mediated tissue-specific protein degradation in C. elegans. Development. 144(14):2694–2701. 10.1242/dev.150094 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Weinreb A et al. 2024. Alternative splicing across the C. elegans nervous system. [accessed 2025 Aug 7]. http://biorxiv.org/lookup/doi/10.1101/2024.05.16.594567. 10.1101/2024.05.16.594567 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Westlund E et al. 2023. Application of nanotags and nanobodies for live cell single-molecule imaging of the Z-ring in Escherichia coli. Curr Genet. 69(2–3):153–163. 10.1007/s00294-023-01266-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu J et al. 2022. Protein visualization and manipulation in Drosophila through the use of epitope tags recognized by nanobodies. eLife. 11:e74326. 10.7554/eLife.74326 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu K et al. 2024. AlphaFold2-guided engineering of split-GFP technology enables labeling of endogenous tubulins across species while preserving function Bullock S, editor. PLOS Biol. 22(8):e3002615. 10.1371/journal.pbio.3002615 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang H et al. 2025. Redirecting E3 ubiquitin ligases for targeted protein degradation with heterologous recognition domains. J Biol Chem. 301(1):108077. 10.1016/j.jbc.2024.108077 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang D et al. 2023. Design of a palette of SNAP-tag mimics of fluorescent proteins and their use as cell reporters. Cell Discov. 9(1):56. 10.1038/s41421-023-00546-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang H et al. 2023. Quantitative assessment of near-infrared fluorescent proteins. Nat Methods. 20(10):1605–1616. 10.1038/s41592-023-01975-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang L, Ward JD, Cheng Z, Dernburg AF. 2015. The auxin-inducible degradation (AID) system enables versatile conditional protein depletion in C. elegans. Development. (142):4374–4384. 10.1242/dev.129635 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zinski J et al. 2024. EpicTope: narrating protein sequence features to identify non-disruptive epitope tagging sites. bioRxiv. [published online ahead of print] [accessed 2025 Aug 4]. 10.1101/2024.03.03.583232 [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data necessary for confirming the conclusions presented in the article are represented fully within the article figures and supplementary information. The data is viewable at the WormTagDB website (https://wormtagdb.rc.duke.edu), and all relevant code is available on GitHub (https://github.com/jakeleyhr/WormTagDB and https://github.com/jakeleyhr/CRISPR-Guide-and-Primer-Design-Pipeline-for-C.-elegans).
