Skip to main content
ACS AuthorChoice logoLink to ACS AuthorChoice
. 2026 Jun 11;17(13):2494–2503. doi: 10.1021/acschemneuro.6c00123

Neuropeptide Diversity Encoded in Newly Sequenced Crustacean Genomes Reveals Signaling Roles during Feeding

Lauren Fields 1, Vu Ngoc Huong Tran 2, Thao Duong 1, Tina C Dang 2, Kendra G Selby 1, Lingjun Li 1,2,3,4,*
PMCID: PMC13329902  PMID: 42276778

Abstract

Neuropeptides are chemically diverse signaling molecules that regulate physiology and behavior, yet many species lack neuropeptidomic characterization due to sparse genomic annotation. Here, we integrate genomic sequence information with neuropeptidomics to define the neuropeptidomes of two widely used crustacean model organisms, Callinectes sapidus and Cancer borealis. Using a curated multispecies precursor database, tBLASTn alignment, signal peptide detection, and in silico processing, we predicted more than 23,000 putative peptides across both genomes, including numerous sequences bearing hallmarks of mature neuropeptides. Mass spectrometry-based profiling provided experimental support for many predicted peptides and revealed substantial chemical diversity, including novel allatostatin B and C isoforms, an insulin-like peptide B-chain-like isoform in C. sapidus, and the first report of natalisin peptides in C. borealis. Notably, we observed an atypical precursor architecture in which a single prohormone encoded two distinct neuropeptide families, suggesting previously unrecognized modes of neuropeptide copackaging and signaling. Finally, leveraging a feeding perturbation model, we observed tissue-specific differences in the abundance of newly identified peptides in the thoracic ganglion, commissural ganglion, and pericardial organs, consistent with functional neuroendocrine roles. Together, this work expands the known repertoire of crustacean neuropeptides, provides a resource for comparative peptide biology, and establishes a genome-enabled framework for discovery of endogenous peptide signaling molecules in newly sequenced species.

Keywords: neuropeptide, genomics, mass spectrometry, Cancer borealis, Callinectes sapidus, endogenous


graphic file with name cn6c00123_0008.jpg


graphic file with name cn6c00123_0006.jpg

Introduction

Neuropeptides are critical signaling molecules secreted by neurons that play essential roles in regulating physiology and have garnered wide attention in healthcare as both biomarkers and therapeutics. Insulin and glucagon-like peptide (GLP-1) peptides are renowned for their contributions to modern medicine and notably originate as prohormones that encode bioactive peptides. − Despite their tremendous importance, much remains unknown about neuropeptides. This knowledge gap stems in part from the complexities of comodulation, in which two or more neuropeptides act together to achieve a shared function. One established approach for unraveling the intricate players behind these complex outcomes is to leverage invertebrate model organisms, which provide a simpler testbed for parsing regulatory functions and applying environmental or biological perturbations to delineate their impact on neuropeptide responses. While rats, mice, Drosophila, nematodes, and other organisms have been used in this arena, a model organism of particular prominence for neurochemical investigations is the crab. Its popularity can be attributed in part to the relative simplicity of the crustacean stomatogastric nervous system (STNS), which governs many homeostatic processes. For example, other simple model organisms, such as the nematode (e.g., C. elegans) have approximately 300 neurons in their nervous system, whereas the stomatogastric ganglion of crustaceans such as C. borealis, have approximately 30 neurons, making them ideal models for neuropeptide research.

A variety of crustacean species have been used to elucidate neuropeptide-driven processes, and as a result, some of the most well-characterized neural circuits originate from these organisms. For example, the gastric mill and pyloric rhythm, central to feeding behavior, are most robustly characterized in the Jonah crab, Cancer borealis. , Additionally, Callinectes sapidus, or the blue crab, has been extensively used to study a variety of environmental-based stressors, including hypoxia, − temperature sensitivity, pH perturbations, and more. Despite this wealth of knowledge, progress in this area has historically been limited due to the lack of genome assemblies. Two of the most widely used crustacean model systems, C. sapidus and C. borealis, were only recently assembled and published in 2021 and 2024, respectively. , In the interim, neuropeptidomic insights have been gained through in silico prediction from transcriptomics, experimental methods such as de novo sequencing by mass spectrometry (MS), and related approaches. However, genome-level information provides unprecedented access to the neuropeptidome. Neuropeptides are highly specialized signaling molecules, with high sequence similarity among peptides that perform distinct functional roles. Thus, although Edman degradation and later de novo sequencing via MS have helped characterize neuropeptide sequences, these methods are restricted to peptides in their processed forms. Transcriptomics has also been heavily applied; however, genome-level analysis provides complementary and independent sequence information, enabling detection of variants that may not be represented in available transcriptomic data sets, particularly for species with limited transcriptomic coverage. − For example, the most comprehensive C. borealis transcriptome-based neuropeptidome to date was derived from pooled nervous system tissues, and would overlook peptides expressed primarily in the peripheral neuroendocrine organs such as the pericardial organs or the sinus glands. Genome-based prediction is tissue- and condition-agnostic, providing access to the full peptide encoding potential of the organism independent of sampling decisions.

With the emergence of more sophisticated prediction models, we sought to use a genome-driven in silico workflow to predict the neuropeptidomes of Callinectes sapidus (blue crab) and Cancer borealis (Jonah crab) from their recently published genomes. Our findings revealed several previously unreported neuropeptides, including novel isoforms and an atypical precursor architecture encoding two distinct neuropeptide families. Moreover, this work demonstrates that by integrating genome-based in silico predictions with existing transcriptome-derived predictions and data from de novo sequencing and other experimental methods, we can generate a more comprehensive representation of the neuropeptidome. This expanded landscape included notable discoveries, such as insulin-like peptide (ILP) isoforms and a single neuropeptide precursor encoding two distinct neuropeptide families.

Finally, we applied our findings to one of the most established biological contexts for Jonah crabs: feeding. We evaluated neuroendocrine tissues postfeeding using our newly curated in silico peptide database and used these predictions to assess patterns in neuropeptide abundance across tissues. In particular, we observed several neuropeptides displaying differential abundance between the thoracic ganglion (TG) and pericardial organs (POs), consistent with tissue-specific neuroendocrine roles.

Results and Discussion

To elucidate neuropeptides from the genome, we first compiled a database spanning a wide variety of neuropeptide families from diverse crustacean species previously reported in UniProt (Supplemental File 1). To ensure broad coverage, we included searches explicitly targeting key neuropeptide families (Figure A). Approximately 40% of the precursor sequences used as queries were classified as level 4 according to UniProt, representing predicted sequences (Figure B). Another 50% of the precursors came from levels 2 or 3, indicating moderate evidence, representing transcript-level evidence and putative identification based on homology, respectively. Just 10% of precursors reflected the highest level of confidence with protein-level evidence (level 1). Finally, upon further inspection, approximately half of the queries corresponded to true peptides, including fully processed mature peptides (‘peptide’), sequences with peptide-like features but uncertain processing (‘peptide-like’), partial peptide sequences (‘peptide fragment’), or peptides mapping to the flanking regions of a precursor (‘precursor-related peptide’). The remaining sequences were classified as precursors, including confirmed precursors, isoforms, or truncated precursor sequences (Figure C). After obtaining precursor sequences, we aligned them to the C. sapidus and C. borealis genomes using tBLASTn, identifying genomic regions with high similarity to known precursors (Figure D). As expected, most database entries originated from shrimp and crab species, reflecting the substantial body of crustacean neuropeptide research in the literature (Figure E). A detailed list of organisms included in the search is provided in Figure S1. The resulting genomic hits were then translated and evaluated for the presence of a signal peptide region. Because signal peptides are a defining feature of neuropeptide precursors, only sequences containing a putative signal peptide were retained for downstream analysis. Once signal peptides were identified, putative neuropeptides were predicted by inducing cleavage at basic residues or at sites directly following the signal peptide via NeuroPred.

1.

1

Distribution of precursors used for in silico prediction of neuropeptides. A) Number of entries per neuropeptide family. B) Level of UniProt documentation for the entry. Level 1 represents the most confident identifications, with support at the protein level. Subsequent levels 2, 3, and 4 represent decreasing confidence, representing support at the transcript level, homology-based identification, and predicted protein. Level 5 (protein uncertain) were omitted from this analysis. C) True query classifications as peptide or precursor, with additional delineation. D) General workflow for prediction. E) Organisms used for prediction. Abbreviations: vitellogenesis inhibiting hormone (VIH), gonad inhibiting hormone (GIH), mandibular organ-inhibiting hormone (MOIH), crustacean cardioactive peptide (CCAP), red pigment concentrating hormone (RPCH), molt-inhibiting hormone (MIH), crustacean hyperglycemic hormone (CHH), precursor-related peptide (PRP).

Prediction of Putative Neuropeptides

Applying this workflow, we identified 14,569 potential peptides for C. borealis and 8,892 peptides for C. sapidus (Figure S2). Predicted peptides for C. borealis and C. sapidus are detailed in Supplemental File 2 and Supplemental File 3, respectively. While this number is substantial, these peptides were derived from only 66 putative precursors in C. sapidus, for example. This relatively small number of precursor genes likely reflects the sparse annotation of both genomes rather than a true biological ceiling, as both assemblies were recently published and remain incompletely annotated. To evaluate the validity of these predictions, we first examined peptides for features indicative of maturity. In general, mature neuropeptides that undergo C-terminal amidation exhibit several hallmark characteristics, including an amidated C-terminus preceded by a glycine residue and a dibasic cleavage site. Additional maturity criteria, including motif conservation, were also considered to account for mature peptides that do not undergo C-terminal amidation. The mature peptides identified in C. borealis and C. sapidus are listed in Tables S1 and S2, respectively. Interestingly, we also identified several mature natalisin peptides, a notable finding given their conserved roles in arthropod neurobiology.

To further contextualize these findings, we evaluated the topology of the identified natalisin peptides. All mature natalisin peptides were preceded by a signal peptide and mapped to the same precursor, which encoded a variety of amidated and nonamidated peptides (Figure A). The amidated peptides exhibited the expected pattern of glycine prior to the cleavage site (Table S3). Regarding signal peptides, the lengths predicted by SignalP for C. sapidus and C. borealis displayed a similar distribution, ranging from as few as 9 to as many as 49 residues (Figure B). Because basic residues are essential for proteolytic processing, we examined the identities of cleavage sites at the N-terminus (Figure C) and C-terminus (Figure D) of mature C. borealis peptides relative to their precursor sequences. While dibasic cleavage sites were most common, we also observed a handful of monobasic cleavages, which have been reported elsewhere. This was further supported by WebLogo analysis, where we evaluated the homology of peptides extending beyond the flanking basic residues illustrated in Figure C,D. Interestingly, when plotting mature peptides, distinct conserved residues were most prominently visible in C. borealis (Figure S3A), where a glycine clearly dominates immediately upstream of the N-terminus cleavage site. Given the significance of glycine for C-terminal amidation of neuropeptides, it is possible that glycine is mechanistically important at the N-terminus as well. While more subtle, conservation was also observed in C. sapidus (Figure S3B), with potential motifs preceding the cleavage site.

2.

2

Mature neuropeptide attributes from findings. A) Precursor topography of a C. borealis natalisin precursor, annotated by color to denote the signal peptide, neuropeptide regions, cleavage sites, amidated neuropeptides, and prepropeptides. Classifications are based on annotation and inference from homologous precursor proteins reported in UniProt and the literature. In this context, ‘neuropeptide’ refers to the predicted mature peptide region, whereas ‘prepropeptide’ refers to the broader precursor-derived sequence context and does not imply lack of biological function. B) Distribution of signal peptide lengths for C. borealis and C. sapidus putative precursors. C) N-termini and D) C-termini sequence distribution for identified, confirmed peptides in C. borealis. “N/A” refers to terminal sequences lacking a basic residue. Chromosomal mapping of the mature precursors found in E) Cancer borealis and F) Callinectes sapidus. Abbreviations: red pigment concentrating hormone (RPCH), pigment dispersing hormone (PDH).

We next mapped these mature precursors back onto the genomes. In C. borealis, 36 neuropeptides mapped to four precursor genes, each located on a distinct chromosome (Figure E). Notably, the natalisin precursor encoded 16 peptides, while the allatostatin B precursor produced 18 peptides (Figure E). For C. sapidus, we mapped 16 mature neuropeptides to six genes across five chromosomes (Figure F). The most prolific family in the C. sapidus analysis was ecdysis-triggering hormone, supported by five mature neuropeptides (Table S4).

Isoform Detection via Mass Spectrometry

To experimentally validate our in silico predictions, we performed MS analysis of C. sapidus and C. borealis to assess the biological context of these predicted peptides within their respective genomes. We then mapped experimentally identified peptides, obtained from database searching of MS results, back onto their corresponding precursors. Notably, these identifications provided additional evidence supporting the accuracy and relevance of our predicted neuropeptidome. For example, several novel peptides (e.g., GAWGKR, FQGSWGKR) appeared repeatedly within the C. borealis precursor derived from transcript g30501. As is well-known, such repetition often hints at a conserved neuropeptide sequence. While these peptides are not deeply characterized as neuropeptides, we used their recurrence as a guide to annotate the precursor more comprehensively.

Using this approach, we also identified novel isoforms of known neuropeptides derived from the genome. For example, we detected several neuropeptides corresponding to the allatostatin B (AST-B) family. Although the identified fragments were relatively short, they highlighted conserved repeats that could be extrapolated to mature neuropeptides. For instance, the peptide QGSWGKR appeared four times within the precursor (Figure A). The QGSW sequence clearly represented the C-terminal end of the peptide, where the -GKR was consistent with peptide amidation followed by a dibasic cleavage site. By extending this sequence upstream toward the preceding dibasic cleavage site (i.e., KR), several known mature peptides including NNWSKFQGSWamide, − TSWGKFQGSWamide, ,− , NNNWSKFQGSWamide ,, and GGWNKFQGSWamide, , were readily apparent, all of which have previously been reported.

3.

3

Peptides observed via mass spectrometry were used to map and extract mature neuropeptides from A) a natalisin precursor in C. borealis, B) an insulin-like peptide (ILP) in C. sapidus, and C) an allatostatin-C peptide in C. sapidus. Signal peptides are shown in red, predicted cleavage sites are highlighted in purple, and putative peptides are highlighted in pink. Green underlines indicate peptides detected by MS/MS database searching, exhibiting a 1 Da mass shift consistent with C-terminal amidation. The adjacent GKR motif denotes the predicted amidation and dibasic cleavage site. Green highlighting between known and predicted sequences indicates homology between experimentally discovered (‘known’) and in silico predicted (‘predicted’) peptides.

Similarly, the conserved GAW motif appeared within the same precursor, accompanied by a KR cleavage site approximately 6–7 residues upstream (Figure A). This organization produced peptides AGWSSMRGAWamide, , AWSNLQGAWamide, and AGWSSLWGAWamide, each flanked by KR and GKR at the N- and C-termini, respectively. Peptides AWSNLQGAWamide and AGWSSMRGAWamide have previously been identified, , as have the closely related isoforms AGWSSLQGAWamide, , AGWSSLKGAWamide, , and AGWSSTSGAWamide. However, this is the first report of AGWSSLWGAWamide, a plausible isoform within the AST-B family. Also included within this precursor were peptides TPDDTPEHGLQGSER, GEEIQAAED, and STNWSSLRGAA (Figure A), which are isoforms of known neuropeptides TPDDTPEHGLQVSED, GEEIQDAEE, and SGDWSSLRGAW, ,, respectively. Curiously, TPDDTPEHGLQGSER, GEEIQAAED, and STNWSSLRGAA belong to the AST-B family, while the remaining peptides encoded on this precursor, including those bearing the QGSW and GAW motifs, correspond to natalisin neuropeptides. A similar event was also found in an RYamide putative precursor within C. sapidus, where two RFamide peptides and one RYamide peptide were encoded on a single precursor (Table S4).

We identified a similar trend in isoforms in a corazonin precursor within C. sapidus (Figure B). The precursor, obtained via homology with Daphnia galeata, documented as a pro-corazonin peptide (UniProt: A0A8J2RDF5), yielded a single peptide from the thoracic ganglion (TG), LANELNRVCK. While not immediately apparent, when mapped in context of its signal peptide on the precursor, two consecutive arginine residues were found upstream, producing peptides SPRTLQEGGLVKQGE and LCGWRLANELNRVCKGVYNMPTVSTNALFYLKGRA, each an isoform of peptides previously predicted via the C. borealis transcriptome. The latter peptide resembles the B-chain of an ILP, while the former is an ILP precursor-related peptide isoform. It is anticipated that the B-chain directly engages with an A-chain ILP, with the documented A-chain appearing as GLSAECCRKACSVSELAGYCY in the literature. Critically, the cysteines within the observed isoform are retained, suggesting preserved functionality. While the A chain was not observed, it can be hypothesized that further mechanistic studies may shed light on a comodulating peptide counterpart. Finally, another set of isoforms was established through an AST-C precursor, where several conserved, yet novel neuropeptides were observed in C. sapidus (Figure C). Altogether, this work provides a new space for neuropeptide discovery and for drawing functional connections.

Motif Evaluation of Putative Peptides

Motifs are a central component of all neuropeptide research efforts. They have been used to identify peptides, hypothesize peptide function, and even bridge findings between vertebrates and invertebrates. Thus, we employed MotifQuest to search for prominent motifs within our predicted neuropeptides from C. sapidus and C. borealis.

In examining the MotifQuest results, we observed several notable motifs that are consistent with those previously reported. For example, the motif CYFNPISCF was found multiple times in the C. sapidus results, consistent with AST-C neuropeptides. Additionally, in C. sapidus, we identified known motifs such as FGXRL (e.g., pyrokinin), YEXD and YDDD (e.g., CHH amide), GASR, PSRA, and QGLG (e.g., CPRP). In C. borealis, known motifs included those for tachykinin peptides: FLGMR, FYGXR, LGXR. CHHamide motifs were also identified, including WPPS, YEED, and GSLP. RYamide motifs were also identified, such as FYSQRY and GGXR. This level of similarity between predicted motifs and experimentally validated motifs was reflected in results obtained from a multiple sequence alignment between predicted and known motifs, showing a wide distribution of similarity scores centered around 50% for both species (Figure A). Upon closer examination, we observed the motif GSDES, which appeared 108 times within the C. sapidus predictions. While we focused mainly on full motifs or consecutive conserved regions of residues, partial motifs or a variable amino acid flanked by two conserved regions have been shown to convey peptide conservation. This same trend was evident in WebLogo analysis of all sequences, where the GSDES motif extended to a variable motif GXXGSDES, along with several other conserved regions (Figure B). In addition to GSDES, several other motifs were found with high frequency in C. borealis (Figure C) and C. sapidus (Figure D).

4.

4

Analysis of conserved motifs in putative predicted neuropeptides. A) Fuzzy multiple sequence alignment between known neuropeptide motif database and predicted neuropeptides for C. sapidus (upper) and C. borealis (lower). B) WebLogo analysis of all matching peptides for motif GSDES, identified 108 times in C. sapidus predictions. Frequency of top ten motifs found for C) C. borealis and D) C. sapidus..

Global Perspective on Neuropeptide Identifications and Potential Functional Roles in Feeding

We observed a substantial increase in the number of identifications in the TG compared to the brain and SGs (Figure A), contrary to typical observations. Additionally, it was compelling that many peptides were detected across all tissues in both C. borealis (Table S5) and C. sapidus (Table S6). Given that feeding is a well-established context for crustacean neuropeptide signaling, we next evaluated our novel peptides in the context of feeding. As a pilot experiment to demonstrate the utility of our predicted peptide database in a biologically relevant context, we compared crabs dissected 30 min postfeeding with unfed controls and observed peptides with differential abundance across the TG, PO, and CoG tissues (Figure B). Three peptides HSNSRGSERamide (adipokinetic), QRLQWLR (AST-B), and RSSRQamide (RFamide-related) exhibited reduced abundances in the PO and elevated levels in the TG, while Q­(Gln→pyro-Glu)­VFEDRamide exhibited the opposite trend (i.e., increased abundance in the PO and reduced levels in the TG), as determined by mean TIC-normalized intensity across technical replicates.

5.

5

Biological analysis of predicted neuropeptides. A) Number of tissue-specific peptide identifications from the brain, sinus glands (SG), and thoracic ganglion (TG) in C. borealis and C. sapidus. B) Ratio of fed to unfed peptide abundance in C. borealis in commissural ganglion (CoG), pericardial organs (PO), and TG. Peptide families are noted below the corresponding peptides. Abbreviation: allatostatin-B (AST-B), crustacean hyperglycemic hormone (CHH). Error bars represent the standard deviation across technical replicates (n = 3).

Given the early stage of genome-enabled neuropeptide research in these species, we anticipated identifying neuropeptide isoforms, putative novel neuropeptides, and other biologically relevant variants made accessible by the recent release of both genomes. However, it should be noted that both genomes remain sparsely annotated, which limits the overall yield of predicted neuropeptides. Nevertheless, our study uncovered several novel candidate neuropeptides.

It is well-established that neuropeptide precursors can encode multiple neuropeptides; however, the extent and organization of this phenomenon are not well-understood. To address this, we mapped our predicted peptides onto their corresponding precursors. Natalisin peptides contain the conserved FXXXRamide motif at their C-termini, often occurring as multiple repeats within the same gene. Their amidation arises from a glycine followed by arginine and/or lysine at dibasic cleavage sites. Initially identified in insects, natalisin peptides have since been reported in crayfish through next-generation sequencing, and more recently through MS in the American lobster. To our knowledge, this is the first instance of natalisin peptide detection in C. borealis. This finding also provided new insight into precursor architecture, as we observed natalisin and AST-B peptides simultaneously encoded on the same neuropeptide precursor. This discovery is particularly intriguing, as genes are typically considered to encode a single neuropeptide family. , To our knowledge, this represents one of the first documented cases in which two distinct peptide families are encoded on a single precursor, suggesting that coordination of distinct neuropeptide families within a single precursor may represent an underappreciated mechanism for neuropeptide signaling.

From a global perspective, we examined identifications obtained through database searches for putative neuropeptides in C. borealis and C. sapidus using their respective predicted databases. Database searching confirmed the presence of numerous predicted peptides in the empirical MS data sets from both species. Although there are relatively few reports on neuropeptides in the TG and the tissue is lipid rich, this framework suggests that this tissue may simply be unexplored, and that an unbiased genome-based search can provide deeper insight into its peptide composition.

In the context of feeding, four peptides showed differential abundance patterns between fed and unfed states, all observed in the PO and TG. It is well established that the PO is connected to the STNS, while the TG is connected to the brain. , Thus, the observation that these two tissues, both peripheral to the heart, display opposing fold changes for the same peptides, is consistent with tissue-specific neuroendocrine roles for these peptides in the context of feeding. Mechanistic resolution of these patterns, for example, distinguishing increased secretion from reduced biosynthesis, will require transcript-level measurements such as RT-qPCR on prohormone mRNAs, which we identify as a priority for future studies.

In this work, we present the first, genome-derived prediction and experimental validation of the neuropeptidomes of Callinectes sapidus and Cancer borealis. By integrating a curated, multispecies neuropeptide precursor database with genomic alignment, signal peptide detection, and subsequent cleavage prediction, we generated a database of putative neuropeptides from two newly assembled and sparsely annotated crustacean genomes. This approach yielded more than 23,000 predicted peptides across both species, including numerous mature neuropeptide sequences and several previously unreported peptide families and isoforms. Collectively, these findings expand the known neuropeptide landscape in crustaceans, illuminating the rich peptide diversity encoded within newly sequenced genomes, and demonstrating the power of combining in silico genomic prediction with experimental MS. This work provides a valuable resource for future studies on neuropeptide evolution, signaling, and physiology, opening new avenues for mechanistic exploration of peptide function in the STNS, and related neuroendocrine organs, establishing a framework readily applicable to other newly sequenced invertebrate species.

Methods

Ethical Use of Animals in Research

No institutional approval is required for working with invertebrates, and all experiments were performed under national and local guidelines and regulations.

Putative precursor identification

For alignment of the genomes with known neuropeptides, tBLASTn (National Center for Biotechnology Information, Bethesda, MD; http://blast.ncbi.nlm.nih.gov/Blast.cgi) was used, restricting to data from either Callinectes sapidus (GCA_020233015.1) or Cancer borealis (GCA_041682235.1) species. Query neuropeptide precursor sequences corresponding to crustacean sequences were extracted via UniProt. Query precursors were obtained via three search strategies: searching for the term “neuropeptide” in UniProt, filtering for crustacean species; searching for particular neuropeptide families in UniProt by name as a keyword, filtering for crustacean species; and referencing the recent publication of the lobster neuropeptidome. The neuropeptide families queried were the following: adipokinetic, allatostatin, bursicon, crustacean cardioactive peptide (CCAP), crustacean hyperglycemic hormone (CHH), corazonin, CHH precursor-related peptide (CPRP), diuretic hormone-31 (DH-31), ecdysis, elevin, FLRFamide, FMRFamide, gonadoliberin, GSEFLamide, HIGSLYRamide, insulin, molt-inhibiting hormone (MIH), mandibular organ-inhibiting hormone (MOIH), myosuppressin, natalisin, neuroparsin, orcokinin, orcomyotropin, proctolin, pyrokinin, RFamide, red pigment concentrating hormone (RPCH), RYamide, SIFamide, short neuropeptide F (sNPF), sulfakinin, tachykinin, and trissin.

Aligned sequences were filtered to an E-value less than 6 for loose searches, and less than 0.001 for restricted searches. Alignment was conducted using both PAM30 and BLOSUM62 algorithms, accommodating neuropeptide precursors of a variety of lengths. Sequence identity was also filtered to exclude hits with sequence identity less than 50%. Subsequently, aligned portions were translated and assessed for signal peptides using SignalP 6.0. Proteins containing a signal peptide with a confidence score of 0.70 or greater were retained. Remaining candidate protein sequences were cleaved using NeuroPred and evaluated for sequence motifs.

Novel Peptide Motif Evaluation

MotifQuest was used to evaluate potential novel motifs that may be present in these identifications. In brief, peptide sequences were parsed from predicted neuropeptide databases using the SeqIO.parse function in BioPython, after which all sequences were aligned with Clustal Omega. Following alignment, pairwise evolutionary divergence between sequences was quantified using the ClustalW distance matrix function, which reports amino acid positional differences, where lower values indicate greater similarity. ClustalW was selected for integration into the MotifQuest workflow because it measures sequence similarity within an evolutionary context, consistent with the premise that conserved motifs arise from evolutionary pressures. Hierarchical clustering was then performed using the SciPy fcluster function applied to the ClustalW-generated distance matrix, with the clustering threshold (T-value) determining how the dendrogram was segmented into distinct motif groups. This approach allowed identification of related sequence families, paralleling the method used to classify neuropeptides. MotifQuest scored the resulting motifs using a framework that extends the weighted coverage method previously reported by incorporating a normalization step based on motif frequency within the full database. Motifs were ranked according to frequency of occurrence within the predicted peptide database.

Similarity between predicted and known motifs was assessed using a fuzzy sequence similarity approach, obtained via the rapidfuzz Python package (v. 3.14.1), which assigned a similarity score of 0–100 based on weighted positional identity, accommodating partial matches at variable positions. For WebLogo analysis, sequences extending 20 residues upstream and downstream of the cleavage site were manually extracted and aligned, setting the cleavage position to zero.

Identification of Putative Mature Neuropeptides

Putative mature neuropeptides were identified based on several criteria: (i) cleavage at mono- or dibasic residues located at the N- and C-termini; (ii) for peptide families known to undergo C-terminal amidation (e.g., RFamide, RYamide, AST-B), the presence of a C-terminal +1 glycine followed by dibasic residues, consistent with established neuropeptide processing mechanisms, was required; (iii) aligned with sequence motifs reported for several crustacean neuropeptide families; , and (iv) cross-referenced with published mature neuropeptides from other invertebrate species. For the purposes of classification, “peptide” refers to a fully processed, mature sequence. Sequences bearing canonical neuropeptide features but lacking complete processing evidence were classified as “peptide-like”. Partial MS/MS-identified sequences were labeled “peptide fragment”, and sequences derived from precursor flanking regions were categorized as “precursor-related peptide”.

Sample Preparation

C. sapidus and C. borealis crabs were obtained from Global Market (Madison, WI) and housed in an artificial seawater tank at a salinity concentration of 30 ppt. Crabs were equilibrated for at least 2 weeks prior to sacrifice. Crabs were dissected following anesthetization on ice for 30 min. From both species, the brain, commissural ganglion (CoG), paired pericardial organs (PO), thoracic ganglion (TG), and paired sinus glands (SG) were obtained and heat stabilized via Denator.

Neuropeptides were extracted from tissue by probe sonication in acidified methanol followed by centrifugation for 1 h at 16k ×g, 4 °C, and the supernatant was retained and dried. For each condition, tissues from three individual crabs were pooled to generate a single biological replicate. Following pooling, samples were desalted using OMIX C18 tips (Agilent) according to manufacturer instructions, eluting sequentially in 25%, 50%, and 75% acetonitrile in 0.1% formic acid in water. Prior to MS analysis, samples were reconstituted in 20 μL of 0.1% FA. To ensure consistency in quantity of peptides injected, samples were evaluated via nanodrop (NanoDrop One, Thermo Fisher Scientific) at 205 nm, reflecting peptide bond absorbance.

Mass Spectrometry Data Acquisition

Untargeted neuropeptide profiling was performed using LC-MS/MS on a Thermo Q-Exactive HF mass spectrometer interfaced with a Dionex Ultimate 3000 liquid chromatography system. Peptide separation employed mobile phase A consisting of 0.1% FA in water and mobile phase B containing 0.1% FA in acetonitrile. The gradient program increased from 10% to 20% B over 70 min, followed by 20% to 95% B over an additional 20 min, at a constant flow rate of 300 nL/min. Full MS data were collected in profile mode from m/z 200 to m/z 2000 at a resolving power of 60,000. The automatic gain control (AGC) target for MS1 scans was set at 1 × 106 with a maximum injection time of 250 ms. MS/MS spectra were acquired in centroid mode, selecting the ten most intense precursor ions for HCD fragmentation with a 30-s dynamic exclusion. For DDA, instrument parameters included a resolution of 15,000, a 2.0 Th isolation window, an NCE of 30, a maximum injection time of 120 ms, an AGC target of 2 × 105, and a fixed first mass of m/z 100. All samples were analyzed in technical triplicate.

Feeding Study

Prior to feeding experiments, Jonah crabs (C. borealis) were equilibrated and fed as described previously. An unfed control crab was also placed on ice at the same time. C. borealis crabs were fed 4 g thawed tilapia, allowed to rest for 30 min following completion of feeding, and subsequently sacrificed with the aforementioned tissue collection strategy. Tissues were stored at −80 °C until MS analysis. Crabs were dissected and tissues were prepared for MS analysis, as described herein. MS analysis was performed on a Thermo Q-Exactive mass spectrometer coupled to a Waters Acquity liquid chromatography system using the same parameters as those described within.

Database Searching from Mass Spectrometry Data

EndoGenius was utilized to evaluate the presence of the predicted peptides within empirical data sets. Analysis parameters were described elsewhere. In brief, precursor and fragment ion tolerances were set to 20 ppm and 0.02 Da, respectively. Variable PTMs included C-terminal amidation, oxidation of M, pyro-Glu from E, and pyro-Glu from Q. All identifications were obtained at an EndoGenius score threshold of 1000.

Quantification and Statistics

For peptide quantification, abundances were normalized against total ion current (TIC) prior to evaluation. Peptide abundances were quantified using the precursor intensity extracted from the.MS2 format spectral files by EndoGenius. Precursor intensities were subsequently normalized to the TIC of each technical replicate to account for run-to-run variation. Prior to MS analysis, sample concentrations were normalized by nanodrop to ensure equivalent injection amounts across technical replicates.

To quantify, the peptides were required to be identified in at least two of the three technical replicates. Abundance was determined as the mean TIC normalized precursor intensity ± the standard deviation. For feeding calculations, the ratio of fed/unfed abundance was calculated as the log2 of the mean TIC-normalized intensity in the fed condition divided by that in the unfed condition. To avoid undefined values for zero-intensity observations, a small pseudo count offset (10–8) was applied. Error bars were calculated as the propagated standard error of the mean across technical replicates.

Supplementary Material

cn6c00123_si_001.pdf (479.6KB, pdf)
cn6c00123_si_002.xlsx (324.5KB, xlsx)
cn6c00123_si_003.xlsx (60KB, xlsx)
cn6c00123_si_004.xlsx (1.4MB, xlsx)

Acknowledgments

The TOC Figure and Figure 1D were produced via BioRender. Portions of this work are derived from L.F.’s doctoral dissertation, “Elucidating the crustacean neuropeptidome through innovative multiplexed data-independent acquisition mass spectrometry and bioinformatics approaches” (University of Wisconsin–Madison, 2025), available through the UW–Madison Libraries (catalog record on 1584705081).

All mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the MassIVE partner repository with the data set identifier: MSV000100298. All code for data analysis and figure generation is available via GitHub at https://github.com/lingjunli-research/genome-guided-crustacean-neuropeptide-prediction.

The Supporting Information is available free of charge at acs.org. The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acschemneuro.6c00123.

  • Multispecies database generation (Figure S1), predicted neuropeptide comparison (Figure S2), precursor sequence evaluation (Figure S3), mature neuropeptides (Tables S1 and S2), annotated neuropeptide precursors (Tables S3 and S4), and commonly found precursors (Tables S5 and S6) (PDF)

  • Multispecies database searched (Supplemental File 1) (XLSX)

  • C. borealis neuropeptide predictions (Supplemental File 2) (XLSX)

  • C. sapidus neuropeptide predictions (Supplemental File 3) (XLSX)

Conceptualization: L.F. and L.L.; implementation and data acquisition: L.F.; data interpretation: L.F., V.N.H.T., T.D., T.C.D., K.G.S., L.L.; funding acquisition: L.L.; drafting manuscript: L.F. and L.L.; manuscript revision: all coauthors.

This work was supported in part by National Institutes of Health (NIH) through grants R01DK071801 and R01NS029346 (L.L.) and the National Science Foundation (NSF) through the grant CHE-2108223 (L.L.). L.F. was supported in part by the National Institute of General Medical Sciences of the National Institutes of Health under Award Number T32GM008505 (Chemistry–Biology Interface Training Program), the 2024 Eli Lilly and Company/ACS Analytical Graduate Fellowship, and a predoctoral fellowship supported by the NIH, under Ruth L. Kirschstein National Research Service Award (NRSA) from the National Institutes of Health-General Medical Sciences F31GM156104. T.C.D. was supported in part by the National Institute of General Medical Sciences of the National Institute of Health under Award 5T32GM141013 (Molecular and Cellular Pharmacology Training Program) and a SciMed Graduate Research Scholars Fellowship through the University of Wisconsin-Madison. L.L. would like to acknowledge NIH Grants R01AG052324, R01AG078794, S10OD028473, S10OD025084, and S10RR029531, as well as funding support from a Vilas Distinguished Achievement Professorship and a Charles Melbourne Johnson Professorship with funding provided by the Wisconsin Alumni Research Foundation and University of Wisconsin-Madison School of Pharmacy.

The authors declare no competing financial interest.

References

  1. Holst J. J.. Glucagon-like peptide-1: Are its roles as endogenous hormone and therapeutic wizard congruent? J. Intern Med. 2022;291(5):557–573. doi: 10.1111/joim.13433. [DOI] [PubMed] [Google Scholar]
  2. Zheng Z., Zong Y., Ma Y., Tian Y., Pang Y., Zhang C., Gao J.. Glucagon-like peptide-1 receptor: mechanisms and advances in therapy. Signal Transduct Target Ther. 2024;9(1):234. doi: 10.1038/s41392-024-01931-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Sato I., Arima H., Ozaki N., Watanabe M., Goto M., Hayashi M., Banno R., Nagasaki H., Oiso Y.. Insulin inhibits neuropeptide Y gene expression in the arcuate nucleus through GABAergic systems. J. Neurosci. 2005;25(38):8657–8664. doi: 10.1523/JNEUROSCI.2739-05.2005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. DeLaney K., Li L.. Neuropeptidomic Profiling and Localization in the Crustacean Cardiac Ganglion Using Mass Spectrometry Imaging with Multiple Platforms. J. Am. Soc. Mass Spectrom. 2020;31(12):2469–2478. doi: 10.1021/jasms.0c00191. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Christie A. E., Stemmler E. A., Dickinson P. S.. Crustacean neuropeptides. Cell. Mol. Life Sci. 2010;67(24):4135–4169. doi: 10.1007/s00018-010-0482-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Fields L., Dang T. C., Tran V. N. H., Ibarra A. E., Li L.. Decoding Neuropeptide Complexity: Advancing Neurobiological Insights from Invertebrates to Vertebrates through Evolutionary Perspectives. ACS Chem. Neurosci. 2025;16(9):1662–1679. doi: 10.1021/acschemneuro.5c00053. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Skiebe P.. Neuropeptides are ubiquitous chemical mediators: Using the stomatogastric nervous system as a model system. J. Exp Biol. 2001;204(Pt 12):2035–2048. doi: 10.1242/jeb.204.12.2035. [DOI] [PubMed] [Google Scholar]
  8. Cook A. P., Nusbaum M. P.. Feeding state-dependent modulation of feeding-related motor patterns. J. Neurophysiol. 2021;126(6):1903–1924. doi: 10.1152/jn.00387.2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Christie A. E., Skiebe P., Marder E.. Matrix of neuromodulators in neurosecretory structures of the crab Cancer borealis. J. Exp Biol. 1995;198(Pt 12):2431–2439. doi: 10.1242/jeb.198.12.2431. [DOI] [PubMed] [Google Scholar]
  10. Chung J. S., Zmora N.. Functional studies of crustacean hyperglycemic hormones (CHHs) of the blue crab, Callinectes sapidus - the expression and release of CHH in eyestalk and pericardial organ in response to environmental stress. FEBS J. 2008;275(4):693–704. doi: 10.1111/j.1742-4658.2007.06231.x. [DOI] [PubMed] [Google Scholar]
  11. Bell G. W., Eggleston D. B., Noga E. J.. Molecular keys unlock the mysteries of variable survival responses of blue crabs to hypoxia. Oecologia. 2010;163(1):57–68. doi: 10.1007/s00442-009-1539-y. [DOI] [PubMed] [Google Scholar]
  12. Bell G. W., Eggleston D. B., Noga E. J.. Environmental and physiological controls of blue crab avoidance behavior during exposure to hypoxia. Biol. Bull. 2009;217(2):161–172. doi: 10.1086/BBLv217n2p161. [DOI] [PubMed] [Google Scholar]
  13. Brouwer M., Larkin P., Brown-Peterson N., King C., Manning S., Denslow N.. Effects of hypoxia on gene and protein expression in the blue crab, Callinectes sapidus. Marine Environmental Research. 2004;58:787–792. doi: 10.1016/j.marenvres.2004.03.094. [DOI] [PubMed] [Google Scholar]
  14. Lehtonen M. P., Burnett L. E.. Effects of Hypoxia and Hypercapnic Hypoxia on Oxygen Transport and Acid-Base Status in the Atlantic Blue Crab, Callinectes sapidus, During Exercise. J. Exp Zool A Ecol Genet Physiol. 2016;325(9):598–609. doi: 10.1002/jez.2054. [DOI] [PubMed] [Google Scholar]
  15. Sauer C. S., Li L.. Mass Spectrometric Profiling of Neuropeptides in Response to Copper Toxicity via Isobaric Tagging. Chem. Res. Toxicol. 2021;34(5):1329–1336. doi: 10.1021/acs.chemrestox.0c00521. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Polinski J. M., O’Donnell T. P., Bodnar A. G.. Chromosome-level reference genome for the Jonah crab, Cancer borealis. G3 (Bethesda) 2025;15:jkae254. doi: 10.1093/g3journal/jkae254. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Bachvaroff T. R., McDonald R. C., Plough L. V., Chung J. S.. Chromosome-level genome assembly of the blue crab, Callinectes sapidus. G3 (Bethesda) 2021;11(9):1–11. doi: 10.1093/g3journal/jkab212. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Christie A. E., Lundquist C. T., Nassel D. R., Nusbaum M. P.. Two novel tachykinin-related peptides from the nervous system of the crab Cancer borealis. J. Exp Biol. 1997;200(Pt 17):2279–2294. doi: 10.1242/jeb.200.17.2279. [DOI] [PubMed] [Google Scholar]
  19. Christie A. E., Cieslak M. C., Roncalli V., Lenz P. H., Major K. M., Poynton H. C.. Prediction of a peptidome for the ecotoxicological model Hyalella azteca (Crustacea; Amphipoda) using a de novo assembled transcriptome. Mar Genomics. 2018;38:67–88. doi: 10.1016/j.margen.2017.12.003. [DOI] [PubMed] [Google Scholar]
  20. Christie A. E., Hull J. J., Richer J. A., Geib S. M., Tassone E. E.. Prediction of a peptidome for the western tarnished plant bug Lygus hesperus. Gen. Comp. Endrocrinol. 2017;243:22–38. doi: 10.1016/j.ygcen.2016.10.008. [DOI] [PubMed] [Google Scholar]
  21. Christie A. E., Roncalli V., Cieslak M. C., Pascual M. G., Yu A., Lameyer T. J., Stanhope M. E., Dickinson P. S.. Prediction of a neuropeptidome for the eyestalk ganglia of the lobster Homarus americanus using a tissue-specific de novo assembled transcriptome. Gen. Comp. Endrocrinol. 2017;243:96–119. doi: 10.1016/j.ygcen.2016.11.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Christie A. E.. Neuropeptide discovery in Proasellus cavaticus: Prediction of the first large-scale peptidome for a member of the Isopoda using a publicly accessible transcriptome. Peptides. 2017;97:29–45. doi: 10.1016/j.peptides.2017.09.003. [DOI] [PubMed] [Google Scholar]
  23. Christie A. E.. Expansion of the neuropeptidome of the globally invasive marine crab Carcinus maenas. Gen. Comp. Endrocrinol. 2016;235:150–169. doi: 10.1016/j.ygcen.2016.05.013. [DOI] [PubMed] [Google Scholar]
  24. Christie A. E., Chi M.. Prediction of the neuropeptidomes of members of the Astacidea (Crustacea, Decapoda) using publicly accessible transcriptome shotgun assembly (TSA) sequence data. Gen. Comp. Endrocrinol. 2015;224:38–60. doi: 10.1016/j.ygcen.2015.06.001. [DOI] [PubMed] [Google Scholar]
  25. Christie A. E.. In silico prediction of a neuropeptidome for the eusocial insect Mastotermes darwiniensis. Gen. Comp. Endrocrinol. 2015;224:69–83. doi: 10.1016/j.ygcen.2015.06.006. [DOI] [PubMed] [Google Scholar]
  26. Wu, W. ; Fields, L. ; DeLaney, K. ; Buchberger, A. R. ; Li, L. . An Updated Guide to the Identification, Quantitation, and Imaging of the Crustacean Neuropeptidome. In Methods Mol Biol; Schrader, M. , Fricker, L. , Eds.; Peptidomics, Vol. 2758; Humana, 2024; pp 255–289. 10.1007/978-1-0716-3646-6_ [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Southey B. R., Amare A., Zimmerman T. A., Rodriguez-Zas S. L., Sweedler J. V.. NeuroPred: a tool to predict cleavage sites in neuropeptide precursors and provide the masses of the resulting peptides. Nucleic Acids Res. 2006;34(Web Server issue):W267–272. doi: 10.1093/nar/gkl161. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Lu G., Tran V. N. H., Wu W., Ma M., Li L.. Neuropeptidomics of the American Lobster Homarus americanus. J. Proteome Res. 2024;23(5):1757–1767. doi: 10.1021/acs.jproteome.3c00925. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Jiang H., Lkhagva A., Daubnerova I., Chae H. S., Simo L., Jung S. H., Yoon Y. K., Lee N. R., Seong J. Y., Zitnan D.. et al. Natalisin, a tachykinin-like signaling system, regulates sexual activity and fecundity in insects. Proc. Natl. Acad. Sci. U. S. A. 2013;110(37):E3526–3534. doi: 10.1073/pnas.1310676110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Veenstra J. A.. Mono- and dibasic proteolytic cleavage sites in insect neuroendocrine peptide precursors. Arch Insect Biochem Physiol. 2000;43(2):49–63. doi: 10.1002/(SICI)1520-6327(200002)43:2<49::AID-ARCH1>3.0.CO;2-M. [DOI] [PubMed] [Google Scholar]
  31. Hui L., Xiang F., Zhang Y., Li L.. Mass spectrometric elucidation of the neuropeptidome of a crustacean neuroendocrine organ. Peptides. 2012;36(2):230–239. doi: 10.1016/j.peptides.2012.05.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Hui L., D’Andrea B. T., Jia C., Liang Z., Christie A. E., Li L.. Mass spectrometric characterization of the neuropeptidome of the ghost crab Ocypode ceratophthalma (Brachyura, Ocypodidae) Gen. Comp. Endrocrinol. 2013;184:22–34. doi: 10.1016/j.ygcen.2012.12.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Ma M., Bors E. K., Dickinson E. S., Kwiatkowski M. A., Sousa G. L., Henry R. P., Smith C. M., Towle D. W., Christie A. E., Li L.. Characterization of the Carcinus maenas neuropeptidome by mass spectrometry and functional genomics. Gen. Comp. Endrocrinol. 2009;161(3):320–334. doi: 10.1016/j.ygcen.2009.01.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. DeKeyser S. S., Kutz-Naber K. K., Schmidt J. J., Barrett-Wilt G. A., Li L.. Imaging mass spectrometry of neuropeptides in decapod crustacean neuronal tissues. J. Proteome Res. 2007;6(5):1782–1791. doi: 10.1021/pr060603v. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. DeLaney K., Hu M., Hellenbrand T., Dickinson P. S., Nusbaum M. P., Li L.. Mass Spectrometry Quantification, Localization, and Discovery of Feeding-Related Neuropeptides in Cancer borealis. ACS Chem. Neurosci. 2021;12(4):782–798. doi: 10.1021/acschemneuro.1c00007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Behrens H. L., Chen R. B., Li L. J.. Combining microdialysis, nanoLC-MS, and MALDI-TOF/TOF to detect neuropeptides secreted in the crab, Anal. Chem. 2008;80(18):6949–6958. doi: 10.1021/ac800798h. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Christie A. E., Pascual M. G.. Peptidergic signaling in the crab Cancer borealis: Tapping the power of transcriptomics for neuropeptidome expansion. Gen. Comp. Endrocrinol. 2016;237:53–67. doi: 10.1016/j.ygcen.2016.08.002. [DOI] [PubMed] [Google Scholar]
  38. Fields L., Vu N. Q., Dang T. C., Yen H. C., Ma M., Wu W., Gray M., Li L.. EndoGenius: Optimized Neuropeptide Identification from Mass Spectrometry Datasets. J. Proteome Res. 2024;23(8):3041–3051. doi: 10.1021/acs.jproteome.3c00758. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Duve H., Johnsen A. H., Scott A. G., Thorpe A.. Allatostatins of the tiger prawn, Penaeus monodon (Crustacea: Penaeidea) Peptides. 2002;23(6):1039–1051. doi: 10.1016/S0196-9781(02)00035-9. [DOI] [PubMed] [Google Scholar]
  40. Kalimullina L. B., Kalkamanov Kh A., Akhmadeev A. V., Zakharov V. P., Sharafullin I. F.. Structural bases for neurophysiological investigations of amygdaloid complex of the brain. Sci. Rep. 2015;5(1):17052. doi: 10.1038/srep17052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Dang T. C., Fields L., Li L.. MotifQuest: An Automated Pipeline for Motif Database Creation to Improve Peptidomics Database Searching Programs. J. Am. Soc. Mass Spectrom. 2024;35(8):1902–1912. doi: 10.1021/jasms.4c00192. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Stemmler E. A., Bruns E. A., Cashman C. R., Dickinson P. S., Christie A. E.. Molecular and mass spectral identification of the broadly conserved decapod crustacean neuropeptide pQIRYHQCYFNPISCF: the first PISCF-allatostatin (Manduca sexta- or C-type allatostatin) from a non-insect. Gen. Comp. Endrocrinol. 2010;165(1):1–10. doi: 10.1016/j.ygcen.2009.05.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Saideman S. R., Ma M., Kutz-Naber K. K., Cook A., Torfs P., Schoofs L., Li L., Nusbaum M. P.. Modulation of rhythmic motor activity by pyrokinin peptides. J. Neurophysiol. 2007;97(1):579–595. doi: 10.1152/jn.00772.2006. [DOI] [PubMed] [Google Scholar]
  44. Elphick M. R., Mirabeau O., Larhammar D.. Evolution of neuropeptide signalling systems. J. Exp Biol. 2018;221(Pt 3):jeb151092. doi: 10.1242/jeb.151092. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Fields L., Wu W., Dang T. C., Ibarra A. E., Gray M., Li L.. Unlocking the Neuropeptidome using a Novel Endogenous Peptidomics Framework. bioRxiv. 2025 doi: 10.1101/2025.06.12.659356. [DOI] [Google Scholar]
  46. Veenstra J. A.. The power of next-generation sequencing as illustrated by the neuropeptidome of the crayfish Procambarus clarkii. Gen. Comp. Endrocrinol. 2015;224:84–95. doi: 10.1016/j.ygcen.2015.06.013. [DOI] [PubMed] [Google Scholar]
  47. Fisher J. M., Sossin W., Newcomb R., Scheller R. H.. Multiple neuropeptides derived from a common precursor are differentially packaged and transported. Cell. 1988;54(6):813–822. doi: 10.1016/S0092-8674(88)91131-2. [DOI] [PubMed] [Google Scholar]
  48. Hook V., Lietz C. B., Podvin S., Cajka T., Fiehn O.. Diversity of Neuropeptide Cell-Cell Signaling Molecules Generated by Proteolytic Processing Revealed by Neuropeptidomics Mass Spectrometry. J. Am. Soc. Mass Spectrom. 2018;29(5):807–816. doi: 10.1007/s13361-018-1914-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Li L., Kelley W. P., Billimoria C. P., Christie A. E., Pulver S. R., Sweedler J. V., Marder E.. Mass spectrometric investigation of the neuropeptide complement and release in the pericardial organs of the crab, Cancer borealis. J. Neurochem. 2003;87(3):642–656. doi: 10.1046/j.1471-4159.2003.02031.x. [DOI] [PubMed] [Google Scholar]
  50. Chen R., Xiao M., Buchberger A., Li L.. Quantitative neuropeptidomics study of the effects of temperature change in the crab Cancer borealis. J. Proteome Res. 2014;13(12):5767–5776. doi: 10.1021/pr500742q. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Pearson W. R.. Selecting the Right Similarity-Scoring Matrix. Curr. Protoc Bioinformatics. 2013;43(1):3 5 1–3 5 9. doi: 10.1002/0471250953.bi0305s43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Teufel F., Almagro Armenteros J. J., Johansen A. R., Gislason M. H., Pihl S. I., Tsirigos K. D., Winther O., Brunak S., von Heijne G., Nielsen H.. SignalP 6.0 predicts all five types of signal peptides using protein language models. Nat. Biotechnol. 2022;40(7):1023–1025. doi: 10.1038/s41587-021-01156-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Cock P. J., Antao T., Chang J. T., Chapman B. A., Cox C. J., Dalke A., Friedberg I., Hamelryck T., Kauff F., Wilczynski B.. et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics. 2009;25(11):1422–1423. doi: 10.1093/bioinformatics/btp163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Sievers F., Higgins D. G.. Clustal Omega for making accurate alignments of many protein sequences. Protein Sci. 2018;27(1):135–145. doi: 10.1002/pro.3290. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Thompson J. D., Higgins D. G., Gibson T. J.. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice. Nucleic Acids Res. 1994;22(22):4673–4680. doi: 10.1093/nar/22.22.4673. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Virtanen P., Gommers R., Oliphant T. E., Haberland M., Reddy T., Cournapeau D., Burovski E., Peterson P., Weckesser W., Bright J.. et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat. Methods. 2020;17(3):261–272. doi: 10.1038/s41592-019-0686-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Hoyle C. H.. Neuropeptide families and their receptors: evolutionary perspectives. Brain Res. 1999;848(1–2):1–25. doi: 10.1016/S0006-8993(99)01975-7. [DOI] [PubMed] [Google Scholar]
  58. Eipper B. A., Stoffers D. A., Mains R. E.. The biosynthesis of neuropeptides: peptide alpha-amidation. Annu. Rev. Neurosci. 1992;15:57–85. doi: 10.1146/annurev.ne.15.030192.000421. [DOI] [PubMed] [Google Scholar]
  59. Tran V. N. H., Duong T. U., Fields L., Tourlouskis K., Beaver M., Li L.. cNPDB: A comprehensive empirical crustacean neuropeptide database. bioRxiv. 2025 doi: 10.1101/2025.07.29.667494. [DOI] [Google Scholar]
  60. DeLaney K., Hu M., Wu W., Nusbaum M. P., Li L.. Mass spectrometry profiling and quantitation of changes in circulating hormones secreted over time in Cancer borealis hemolymph due to feeding behavior. Anal Bioanal Chem. 2022;414(1):533–543. doi: 10.1007/s00216-021-03479-1. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

cn6c00123_si_001.pdf (479.6KB, pdf)
cn6c00123_si_002.xlsx (324.5KB, xlsx)
cn6c00123_si_003.xlsx (60KB, xlsx)
cn6c00123_si_004.xlsx (1.4MB, xlsx)

Data Availability Statement

All mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the MassIVE partner repository with the data set identifier: MSV000100298. All code for data analysis and figure generation is available via GitHub at https://github.com/lingjunli-research/genome-guided-crustacean-neuropeptide-prediction.


Articles from ACS Chemical Neuroscience are provided here courtesy of American Chemical Society

RESOURCES