Abstract
Red swamp crayfish, Procambarus clarkii, is the most cultured freshwater crayfish species. It attracts significant research attention due to its considerable economic importance. However, the limited availability of genome information has impeded further genetic studies and breeding programs. By utilizing Illumina, PacBio, and Hi-C sequencing technologies, we present a more comprehensive and continuous chromosome-level assembly for P. clarkii than the published one. The final genome size is 4.03 Gb, consisting of 2,358 scaffolds with a N50 of 42.87 Mb. Notably, 3.68 Gb, corresponding to 91.42% of the genome, was anchored to 94 chromosomes. The assembly comprises 70.64% repetitive sequences, including 5.21% tandem repeats and 65.40% transposable elements. Additionally, a total of 4,456 non-coding RNAs and 28,852 protein-coding genes were predicted in the P. clarkii genome, with 96.26% of the genes were annotated. This high-quality genome assembly not only represents a significant improvement for the genome of P. clarkii and provides insights into the unique genome evolution, but also offers valuable information for developing freshwater aquaculture and accelerating genetic breeding.
Subject terms: Genome, Evolutionary genetics
Background & Summary
Freshwater crayfish, comprising over 640 species, are crucial components of aquatic ecosystems, with Procambarus clarkii (Girard 1852) standing out as a representative species of significant economic value1. In 2019, the global production of P. clarkii reached 2.17 million tons, valued at over US$14 billion2. China dominates production with a staggering 95%, while the United States follows as the second-largest producer, contributing only 4% to the global production.
Native to the southern United States and northern Mexico, P. clarkii was introduced into China in 1929 and successfully established populations in more than 40 countries across the five continents, making it the most widely distributed freshwater crayfish globally3,4. Moreover, P. clarkii demonstrates remarkable adaptability in various freshwater ecosystems, including rivers, lakes, ponds, and, notably, rice fields5,6. In China, the cultivation area of P. clarkii has rapidly increased to 29.50 million mu in 2023, with the rice-crayfish integrated farming model accounting for 85.76%.
Despite being the most cultured crayfish species, the limited genetic information of P. clarkii has hindered further genetic studies and breeding programs. To address this, a chromosome-scale genome of P. clarkii was conducted using Illumina short-reads, PacBio long-reads, and Hi-C sequencing data (Fig. 1). A contig-level genome, consisting of 15,724 contigs with an N50 of 488.66 kb, represents a substantial improvement over the published genome7, which had an N50 of 216.75 kb. After Hi-C scaffolding, 2,358 scaffolds were obtained with an N50 of 42.87 Mb and a total length of 4.03 Gb. The largest 94 scaffolds were anchored to 94 chromosomes, spanning 3.68 Gb and corresponding to 91.42% of the assembled sequences. This genome includes 70.64% repeat sequences, 4,456 noncoding RNAs, and 28,852 genes, of which 96.26% were functionally annotated. The high-quality genome promises valuable insights into the molecular basis of P. clarkii, enriching crustacean genomic resources and providing foundational data for genome-wide selective breeding initiatives.
Fig. 1.
The pipelines overview of P. clarkii chromosome-level genome assembly and annotation.
Methods
Sample collection
Specimens of P. clarkii were obtained from a research base situated in Wuhan, Hubei, China (30°24′21″ N, 113°46′33″ E), where the rice-crayfish farming model was implemented. A total of twelve tissues were collected: ovary, hepatopancreas, gill, and muscle from a female crayfish, as well as brain, eyestalk, epidermis, stomach, heart, intestine, antennal gland, and testis from a male crayfish. All samples were immediately stored at −80 °C. Genomic DNA of muscle was extracted and subsequently used for Illumina, PacBio, and Hi-C sequencing, while total RNA of all twelve tissues was extracted for subsequent Iso-Seq and RNA-Seq. Both DNA and RNA were stored at −80 °C prior to sequencing.
Illumina sequencing and genome survey
The Illumina HiSeq X-ten instrument was used to sequence two libraries (270 bp and 500 bp), obtaining 576.87 Gb of raw data (Table 1). After filtering using SOAPnuke8 (v1.5.3) with parameter settings “-l 10 -q 0.4 -n 0.01”, resulting in 467.83 Gb of clean data. The clean data corresponded to approximately 117-fold haploid genome coverage. Of these, 120 Gb data were performed a genome survey analysis using a 21-mer frequency distribution with Jellyfish9 (v1.1.12) and GenomeScope10 (v2.0). This result estimated that the genome of P. clarkii is approximately 3.91 Gb in size, the heterozygosity rate is 1.01%, and the repeat content is 75.7% (Fig. 2a).
Table 1.
Statistics of sequencing data.
| Data type | Instrument/Platform | Raw Data (Gb) | Clean Data (Gb) | N50 (bp) | GC (%) |
|---|---|---|---|---|---|
| WGS short reads | Illumina Hiseq X-ten | 576.87 | 467.83 | — | 45.23 |
| PacBio long reads | PacBio Sequel | 431.89 | — | 19,710 | 42.22 |
| Hi-C reads | MGISEQ-2000 | 847.06 | 835.58 | — | 43.65 |
| Iso-seq subreads | PacBio Sequel | 31.09 | — | 3,424 | 43.16 |
| Transcriptome short reads | MGISEQ-2000 | 136.46 | 109.23 | — | 38.57 |
Fig. 2.
Genome assembly of P. clarkii. (a) The 21-mer analysis of the genome. len, estimated genome size in bp; uniq, unique percent; het, heterozygosity rate; kcov, kmer coverage; err, error rate; dup, duplication rate; k, kmer. (b) Hi-C assembly of chromosome interactive heatmap. Numbers on the axes show the chromosome length in Gb. The numbers at the bottom of the figure represents the chromosomes number. The ends of each link map to assembled scaffolds, with red dots represented the intersections of their corresponding loci in contigs. (c) Genomic features in a sliding window of 1 Mb. From the outermost to the innermost circle: i, chromosome ideograms; ii, protein-coding gene density; iii, TE content density; iv, LINE/CR1 content density; v, LTR/Gypsy content density; vi, GC content density.
PacBio sequencing and genome assembly
Twenty-four SMRTbell libraries were constructed with 20-kb DNA inserts, and subsequently sequenced using the PacBio Sequel instrument. PacBio sequencing generated 431.89 Gb of raw data with an N50 of 19.71 kb (Table 1). Subsequently, Falcon11 (v0.2.2) and SMARTdenovo12 (v1.0.0) were used to assemble the genome of P. clarkii, obtaining three assemblies: FALCON_1, FALCON_2, and SMARTdenovo versions (Table 2). The FALCON_1 and SMARTdenovo versions were assembled with default parameters, while the FALCON_2 version was assembled with following parameters “pa_DBsplit_option: -s600 -x2000; pa_HPCdaligner_option: -v -M24 -D200 -h600 -e.75 -s1000 -l4500 -k18; overlap_filtering_setting:–max_diff 90–max_cov 180–min_cov 10–bestn 10–n_core 16–min_len 10000”.
Table 2.
Different genome assemblies of P. clarkia.
| FALCON_1 | FALCON_2 | SMARTdenovo | |
|---|---|---|---|
| N50 (Number) | 238,067 bp (6,978) | 342,212 bp (4,480) | 319,513 bp (2,580) |
| N60 (Number) | 169,238 bp (10,281) | 256,456 bp (6,458) | 231,795 bp (3,964) |
| N70 (Number) | 108,440 bp (15,155) | 184,151 bp (9,147) | 164,297 bp (5,898) |
| N80 (Number) | 59,535 bp (23,429) | 120,262 bp (13,057) | 111,317 bp (8,677) |
| N90 (Number) | 36,170 bp (38,034) | 65,344 bp (19,721) | 66,461 bp (13,018) |
| Contig Number | 72,012 | 34,869 | 22,275 |
| Max Contig | 2,482,788 bp | 5,857,453 bp | 7,134,225 bp |
| Total | 6,634,750,326 bp | 5,839,445,558 bp | 3,746,471,481 bp |
Comparing the total length of the three assemblies with the estimated genome size, we selected the FALCON_2 assembly for further genome polishing with FinisherSC13 (v1.22) and Pilon14 (v1.22). Then, the redundant contigs which shared > = 60% non-repetitive-region 17-mers (depth <100) with other sequences were filtered using the TrimDup15 tool (https://github.com/gigascience/rabbit-genome-assembler). After removing redundancies, we obtained a 4.02 Gb genome assembly, which is consist of 15,724 contigs with a N50 of 488.66 kb (Table 3).
Table 3.
Optimization of the P. clarkii genome assembly.
| FALCON_2 | Polish | De-redundancy | Hi-C | |
|---|---|---|---|---|
| N50 (Number) | 342,212 bp (4,480) | 343,380 bp (4,480) | 488,656 bp (2,377) | 42,870,550 bp (40) |
| N60 (Number) | 256,456 bp (6,458) | 257,297 bp (6,458) | 386,822 bp (3,305) | 36,398,550 bp (51) |
| N70 (Number) | 184,151 bp (9,147) | 184,825 bp (9,145) | 294,505 bp (4,495) | 34,570,824 bp (62) |
| N80 (Number) | 120,262 bp (13,057) | 120,720 bp (13,053) | 212,348 bp (6,097) | 29,397,883 bp (75) |
| N90 (Number) | 65,344 bp (19,721) | 65,466 bp (19,715) | 127,446 bp (8,509) | 19,119,945 bp (91) |
| Contig/Scaffold Number | 34,869 | 34,869 | 15,724 | 2,358 |
| Max Contig | 5,857,453 bp | 5,872,046 bp | 5,872,046 bp | 65,481,133 bp |
| Total | 5,839,445,558 bp | 5,858,056,529 bp | 4,022,935,872 bp | 4,029,667,872 bp |
Hi-C sequencing and scaffolding
After sequencing on the BGI-Shenzhen MGISEQ-2000 platform, 847.06 Gb of raw Hi-C reads were obtained and then filtered, resulting in 835.58 Gb filtered data (Table 1). Bowtie216 (v2.2.5) was used to align the Hi-C data to the FALCON_2 assembly with the parameters “--very-sensitive -L 30 --score-min L, -0.6, -0.2 --end-to-end –reorder”, obtaining 766,705,231 unique mapped read pairs. Of these, the valid interaction pairs identified using HiC-Pro17 (v2.5.0) achieved an impressive success rate of 83.61%. Subsequently, we used Juicer18 (v1.6.2) and 3D-dna19 (v180922) to anchor the contigs into chromosome-length scaffolds. The chromosome-level assembly contains 4.03 Gb in 2,358 scaffolds, with a N50 of 42.87 Mb (Table 3). Total length of 94 chromosomes was 3.68 Gb, with lengths range from 15.93 Mb to 65.48 Mb (Fig. 2b).
Repeat and noncoding RNAs annotation
Tandem repeats were annotated by Tandem Repeats Finder20 (TRF, v4.09), while transposable elements (TEs) were annotated using de novo and homology-based predictions. We used RepeatMasker to annotate TEs in a de novo library constructed by RepeatModeler21 (v2.0.2). Meanwhile, we used RepeatMasker (v4.0.7) and RepeatProteinMask (v3.30) to identify the TEs with the Repbase transposable element library22. Using TEclass23 (v2.1.3), the unknown TEs predicted above were further classified to four categories: DNA transposons, long terminal repeats (LTRs), long interspersed nuclear elements (LINEs), and short interspersed nuclear elements (SINEs). To sum, 70.64% repeats were identified in this genome, including 5.24% tandem repeats and 65.40% TEs (Table 4). This result is compared to the estimated repeat content of 75.7%.
Table 4.
Classification of repeat elements in the P. clarkii genome.
| Types | Software | Repeat Size (bp) | % of genome |
|---|---|---|---|
| Tandem repeats | TRF | 210,947,850 | 5.24 |
| Transposable elements | RepeatMasker | 741,299,344 | 18.43 |
| RepeatProteinMask | 781,654,467 | 19.43 | |
| RepeatModeler | 2,481,286,644 | 61.68 | |
| Combined TEs | 2,630,818,761 | 65.40 | |
| Total | — | 2,841,766,611 | 70.64 |
The tRNA was identified using tRNAscan-SE24 (v1.3.138), while the miRNA was identified using the miRbase25 (v22) database. Based on the Rfam26 (v12.0) database, the rRNA was predicted by RNAmmer27 (v1.2), while the snRNA was predicted with Infernal28 (v1.1). A total of 837 tRNAs, 28 miRNAs, 2,819 rRNAs, and 772 snRNA were predicted (Table 5).
Table 5.
ncRNA statistics of P. clarkii genome.
| Type | Number | Total length (bp) | Average length (bp) | % of genome |
|---|---|---|---|---|
| tRNA | 837 | 62,940 | 75.19713 | 0.16 |
| miRNA | 28 | 2,145 | 76.60714 | 0.01 |
| rRNA | 2,819 | 304,302 | 107.9468 | 0.76 |
| snRNA | 772 | 132,946 | 172.2098 | 0.33 |
Gene prediction and function annotation
In this study, we performed gene prediction in a repeat-masked genome through transcriptome-based, homology-based, and de novo prediction methods. For transcriptome-based prediction, one PacBio Iso-seq library from muscle and eleven Illumina RNA-seq libraries from other tissues (including brain, eyestalk, gill, epidermis, stomach, heart, hepatopancreas, intestine, antennal gland, testis, and ovary) were constructed. Sequencing yielded 31.09 Gb of Iso-Seq data and 68.23 Gb of RNA-seq filtered data (Table 1). These data were aligned to the P. clarkii genome using Hisat229 (v2.1.0) with the parameters “–sensitive–no-mixed–no-discordant -X 1000 -I 1”, followed by assembly using StringTie30 (v1.0.4). For homology-based prediction, BLASTN31 (v2.2.26) was used to align protein sequences from D. melanogaster, H. azteca, and L. vannamei against the P. clarkii genome, as well as GeneWise32 (v2.4.1) was then performed to compare the aligned proteins with the genome sequences to extract exon-intron information. Meanwhile, we performed de novo assembled for Iso-Seq and RNA-Seq data using Trinity33 (v2.5.1) and predicted the gene structures with TransDecoder34 (v5.5.0). De novo gene prediction was carried out using AUGUSTUS35 (v3.4.0). The results from the homology-based, transcriptome-based, and de novo predictions were integrated using EvidenceModeler36 (v5.0.2). A total of 28,852 genes were successfully predicted in the P. clarkii genome, with an average exon number of 4.5. The average lengths of gene, coding sequence (CDS), exon, and intron were 17,176 bp, 1,217 bp, 270 bp, and 3,833 bp, respectively (Fig. 2c, Table 6).
Table 6.
Summary statistics of predicted protein-coding genes of P. clarkii.
| Species | Total number of genes | Average gene length (bp) | Average CDS length (bp) | Average exon number | Average exon length | Average intron length |
|---|---|---|---|---|---|---|
| P. clarkii | 28,852 | 17,176.21 | 1,217.33 | 4.50 | 270.43 | 3,833.82 |
For functional annotation, six databases were used: NCBI non-redundant proteins (NR), SwissProt, TrEMBL, InterPro, Gene Ontology (GO), and KEGG. A total of 27,774 genes (96.26%) were successfully annotated with at least one functional protein database (Table 7). Among these, 11,223 genes were annotated in six databases, 6,734 in five databases, 4,141 in four databases, 3,121 in three databases, 1,758 in two databases, and 797 in only one database (Fig. 3).
Table 7.
Statistical results of gene function annotation.
| Database | NR | SwissProt | TrEMBL | InterPro | GO | KEGG | Total |
|---|---|---|---|---|---|---|---|
| Number | 26,774 | 19,662 | 27,568 | 20,858 | 13,121 | 23,265 | 27,774 |
| Percentage | 92.80% | 68.15% | 95.55% | 72.29% | 45.48% | 68.15% | 96.26% |
Fig. 3.
Statistical results of gene annotation in the P. clarkii genome.
Data Records
All the raw sequencing data, including Illumina WGS, PacBio WGS, Hi-C WGS, Illumina RNA-Seq, and PacBio Iso-seq, have been submitted to the Sequence Read Archive (SRA) of National Center for Biotechnology Information (NCBI) under BioProject accession number PRJNA1123252, with accession number SRP51624437. The genome assembly has been deposited at GenBank under the accession JBFOCG00000000038. The annotation file was deposited at the Figshare repository39.
Technical Validation
We assessed the quality of P. clarkii genome assembly in the four aspects. Firstly, contig N50 has achieved more than double the improvement to 488.66 Kb, which is much higher than published assembly (GCA_020424385.2) of P. clarkii or other closely related species7,40–50. Secondly, we aligned the Illumina sequencing data back to the genome using BWA51 (v0.7.15), revealing a high mapping rate of 98.92%. Thirdly, we obtained 94.6% completed and 2.3% fragmented BUSCOs52 (v4.0) using the arthropoda_odb9 database. Fourthly, we analyzed the reasons for the differences in genome size of the two assemblies. Compared to the assessed genome size (3.07 Gb) of the published genome, our estimated size (3.91 Gb) is more consistent with the flow cytometry results (2 C = 8.4 ± 0.45 pg)53,54. We also assessed collinearity with the previous published genome using MUMMER55 (v3.0) with the parameters “–mum–mincluster 500”. As a result, 2.63 Gb of sequences of our assembly and 2.29 Gb of published genomes showed 1-to-1 Mummer-mapped alignments, while the other sequences with a high level of repeat content (77.76% and 75.78%) were classified as Mummer-unmapped, including many-to-many alignments and mismatches (Table 8). In addition, 135 Gb of our Illumina reads were aligned to the two genomes using BWA. The ratio of mapped reads to the two genomes was similar, but the mapped ratio of paired reads to our assembly (88.29%) is significantly higher than that to published genomes (79.38%) (Table 9). A total of 62.88 Mb and 148.29 Mb regions of the tow genomes have the coverage depth >100, and the accumulated mapped bases reach 17.40 Gb and 44.58 Gb, respectively. The theory genome size of these high-depth regions was estimated by accumulated mapped bases/Coverage depth (33X), resulting in 527.27 Mb and 1.32 Gb, respectively (Table 9). These results imply that our assembly has the higher integrity in high repeat regions.
Table 8.
The statistics of collinearity between the two P. clarkii genomes.
| This genome | Published genome (GCA_020424385.2) | ||
|---|---|---|---|
| Assembly Size | 4,029,667,872 bp | 2,735,413,669 bp | |
| Mummer-mapped (1-to-1) | Size | 2,629,236,616 bp | 2,293,548,872 bp |
| Repeat | 1,590,050,739 bp | 1,336,647,316 bp | |
| Repeat ratio | 60.48% | 58.28% | |
| Gene Number | 23,823 | ||
| Mummer-unmapped (many-to-many matches, or mismatch) | Size | 1,400,431,256 bp | 441,864,797 bp |
| Repeat | 1,089,010,611 bp | 334,852,967 bp | |
| Repeat ratio | 77.76% | 75.78% | |
| Gene Number | 5,029 |
Table 9.
The statistics of collinearity between the two P. clarkii genomes.
| This genome | Published genome | |
|---|---|---|
| Total reads number | 906,993,191 | |
| Total bases | 135,158,456,739 bp | |
| BWA mapped reads number | 896,506,599 | 892,692,378 |
| BWA mapped reads ratio | 98.84% | 97.97% |
| Paired reads number | 800,760,657 | 723,269,574 |
| Paired reads ratio | 88.29% | 79.38% |
| Unpaired reads number | 95,745,942 | 169,422,804 |
| Unpaired reads ratio | 10.55% | 18.59% |
| Total bases of high-depth region (Coverage Depth > 100) | 62,884,502 bp | 148,290,844 bp |
| Accumulated mapped bases | 17,399,429,682 bp | 44,581,751,584 bp |
| Theory genome size | 527,255,445 bp | 1,350,962,169 bp |
Acknowledgements
The authors also kindly acknowledge Dr. Kosuke Nakanishi (The University of Shiga Prefecture, Japan), Dr. Yew-Hu Chien (National Taiwan Ocean University, Taiwan, China) and Dr. Jinshui Zheng (Huazhong Agricultural University, Wuhan, China) for their insightful discussions and experimental assistance. This work was supported by National Key Research and Development Program of China (2022YFD2400700), Major Project of Hubei Hongshan Laboratory (2021HSZD002), Hubei Key Research and Development Program (2021BBA231), and Hubei Agricultural Sciences and Technology Innovation Center (2021-620-000-001-33).
Author contributions
Z.G., initiated, managed, and drove the genome project. R.H. and M.L. conceived the study. M.L. and Z.X. collected the animal material. M.X. and L.Q. managed data and provided bioinformatic support. M.L. and M.X. analyzed the data, with input from Z.G., R.H., J.W. and S.Z., M.L. carried out the experiments with Y.L. and X.L., M.L. and R.H. wrote the manuscript with contributions from Z.X., C.B., J.W., S.Z. and Z.G. All authors read and approved the final manuscript.
Code availability
All software and tools were used following the parameters mentioned in the Methods section or with default parameters.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Mingcong Liao, Meng Xu, Ruixue Hu.
References
- 1.Kozák, P. et al. Crayfish biology and culture, Vodňany, Czech Republic: University of South Bohemia in České Budějovice, Faculty of Fisheries and Protection of Waters, (2015).
- 2.FAO. FAO yearbook: fishery and aquaculture statistics 2019, (2021).
- 3.Oficialdegui, F. J., Sánchez, M. I. & Clavero, M. One century away from home: how the red swamp crayfish took over the world. Rev Fish Biol Fish30, 121–135, 10.1007/s11160-020-09594-z (2020). 10.1007/s11160-020-09594-z [DOI] [Google Scholar]
- 4.Oficialdegui, F. J. et al. Unravelling the global invasion routes of a worldwide invader, the red swamp crayfish (Procambarus clarkii). Freshwater Biol64, 1382–1400, 10.1111/fwb.13312 (2019). 10.1111/fwb.13312 [DOI] [Google Scholar]
- 5.Loureiro, T. G., Anastácio, P. M. S. G., Araujo, P. B., Souty-Grosset, C. & Almerão, M. P. Red swamp crayfish: biology, ecology and invasion-an overview. Nauplius23, 1–19, 10.1590/S0104-64972014002214 (2015). 10.1590/S0104-64972014002214 [DOI] [Google Scholar]
- 6.Chen, L. et al. The Microbiome Structure of a Rice-Crayfish Integrated Breeding Model and Its Association with Crayfish Growth and Water Quality. Microbiol Spectr10, e02204-21, 10.1128/spectrum.02204-21 (2022). 10.1128/spectrum.02204-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Xu, Z. et al. A chromosome-level reference genome of red swamp crayfish Procambarus clarkii provides insights into the gene families regarding growth or development in crustaceans. Genomics113, 3274–3284, 10.1016/j.ygeno.2021.07.017 (2021). 10.1016/j.ygeno.2021.07.017 [DOI] [PubMed] [Google Scholar]
- 8.Chen, Y. et al. SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience7, gix120, 10.1093/gigascience/gix120 (2018). 10.1093/gigascience/gix120 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Marcais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics27, 764–770, 10.1093/bioinformatics/btr011 (2011). 10.1093/bioinformatics/btr011 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Vurture, G. W. et al. GenomeScope: fast reference-free genome profiling from short reads. Bioinformatics33, 2202–2204, 10.1093/bioinformatics/btx153 (2017). 10.1093/bioinformatics/btx153 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Chin, C.-S. et al. Phased diploid genome assembly with single-molecule real-time sequencing. Nat Methods13, 1050–1054, 10.1038/nmeth.4035 (2016). 10.1038/nmeth.4035 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Liu, H. L., Wu, S. G., Li, A. & Ruan, J. SMARTdenovo: A de novo assembler using long noisy reads. Gigabyte2021, 1–9, 10.20944/preprints202009.0207.v1 (2021). 10.20944/preprints202009.0207.v1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Lam, K.-K., LaButti, K., Khalak, A. & Tse, D. FinisherSC: a repeat-aware tool for upgrading de novo assembly using long reads. Bioinformatics31, 3207–3209, 10.1093/bioinformatics/btv280 (2015). 10.1093/bioinformatics/btv280 [DOI] [PubMed] [Google Scholar]
- 14.Walker, B. J. et al. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS One9, e112963, 10.1371/journal.pone.0112963 (2014). 10.1371/journal.pone.0112963 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.You, M. et al. A heterozygous moth genome provides insights into herbivory and detoxification. Nat Genet45, 220–225, 10.1038/ng.2524 (2013). 10.1038/ng.2524 [DOI] [PubMed] [Google Scholar]
- 16.Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat Methods9, 357–359, 10.1038/nmeth.1923 (2012). 10.1038/nmeth.1923 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Servant, N. et al. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol16, 1–11, 10.1186/s13059-015-0831-x (2015). 10.1186/s13059-015-0831-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Durand, N. C. et al. Juicer Provides a One-Click System for Analyzing Loop-Resolution Hi-C Experiments. Cell Syst3, 95–98, 10.1016/j.cels.2016.07.002 (2016). 10.1016/j.cels.2016.07.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Dudchenko, O. et al. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science356, 92–95, 10.1126/science.aal3327 (2017). 10.1126/science.aal3327 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Benson, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res27, 573–580, 10.1093/nar/27.2.573 (1999). 10.1093/nar/27.2.573 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc Natl Acad Sci USA117, 9451–9457, 10.1073/pnas.1921046117 (2020). 10.1073/pnas.1921046117 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Bao, W., Kojima, K. K. & Kohany, O. Repbase Update, a database of repetitive elements in eukaryotic genomes. Mob DNA6, 1–6, 10.1186/s13100-015-0041-9 (2015). 10.1186/s13100-015-0041-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Abrusán, G., Grundmann, N., DeMester, L. & Makalowski, W. TEclass—a tool for automated classification of unknown eukaryotic transposable elements. Bioinformatics25, 1329–1330, 10.1093/bioinformatics/btp084 (2009). 10.1093/bioinformatics/btp084 [DOI] [PubMed] [Google Scholar]
- 24.Chan, P. P. & Lowe, T. M. tRNAscan-SE: searching for tRNA genes in genomic sequences. Methods Mol Biol1962, 1–14, 10.1007/978-1-4939-9173-0_1 (2019). 10.1007/978-1-4939-9173-0_1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Kozomara, A., Birgaoanu, M. & Griffiths-Jones, S. miRBase: from microRNA sequences to function. Nucleic Acids Res47, D155–D162, https://doi.org/MiRBase (2019). 10.1093/nar/gky1141 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Nawrocki, E. P. et al. Rfam 12.0: updates to the RNA families database. Nucleic Acids Res43, D130–D137, 10.1093/nar/gkx1038 (2015). 10.1093/nar/gkx1038 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Lagesen, K. et al. RNAmmer: consistent and rapid annotation of ribosomal RNA genes. Nucleic Acids Res35, 3100–3108, 10.1093/nar/gkm160 (2007). 10.1093/nar/gkm160 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Nawrocki, E. P. & Eddy, S. R. Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics29, 2933–2935, 10.1093/bioinformatics/btt509 (2013). 10.1093/bioinformatics/btt509 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Kim, D., Paggi, J. M., Park, C., Bennett, C. & Salzberg, S. L. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol37, 907–915, 10.1038/s41587-019-0201-4 (2019). 10.1038/s41587-019-0201-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Pertea, M. et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol33, 290–295, 10.1038/nbt.3122 (2015). 10.1038/nbt.3122 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Camacho, C. et al. BLAST+: architecture and applications. BMC Bioinform10, 1–9, 10.1186/1471-2105-10-421 (2009). 10.1186/1471-2105-10-421 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Birney, E., Clamp, M. & Durbin, R. GeneWise and Genomewise. Genome Res14, 988–995, 10.1101/gr.1865504 (2004). 10.1101/gr.1865504 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Grabherr, M. G. et al. Trinity: reconstructing a full-length transcriptome without a genome from RNA-Seq data. Nat Biotechnol29, 644, 10.1038/nbt.1883 (2011). 10.1038/nbt.1883 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Grabherr, M. G. et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol29, 644–652, 10.1038/nbt.1883 (2011). 10.1038/nbt.1883 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Stanke, M. & Morgenstern, B. AUGUSTUS: a web server for gene prediction in eukaryotes that allows user-defined constraints. Nucleic Acids Res33, W465–W467, 10.1093/nar/gki458 (2005). 10.1093/nar/gki458 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol9, 1–22, 10.1186/gb-2008-9-1-r7 (2008). 10.1186/gb-2008-9-1-r7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.NCBI Sequence Read Archivehttp://identifiers.org/ncbi/insdc.sra:SRP516244 (2024).
- 38.NCBI GenBankhttps://identifiers.org/ncbi/insdc:JBFOCG000000000 (2024).
- 39.Liao, M. C. Genome annotation for the red swamp crayfish Procambarus clarkii. figshare10.6084/m9.figshare.24589233 (2023). 10.6084/m9.figshare.24589233 [DOI]
- 40.Shao, C. et al. The enormous repetitive Antarctic krill genome reveals environmental adaptations and population insights. Cell186, 1279–1294, 10.1016/j.cell.2023.02.005 (2023). 10.1016/j.cell.2023.02.005 [DOI] [PubMed] [Google Scholar]
- 41.Chen, H. et al. The chromosome-level genome of Cherax quadricarinatus. Sci Data10, 215, 10.1038/s41597-023-02124-z (2023). 10.1038/s41597-023-02124-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Zhao, M. et al. A chromosome-level genome of the mud crab (Scylla paramamosain estampador) provides insights into the evolution of chemical and light perception in this crustacean. Mol Ecol Resour21, 1299–1317, 10.1111/1755-0998.13332 (2021). 10.1111/1755-0998.13332 [DOI] [PubMed] [Google Scholar]
- 43.Uengwetwanit, T. et al. A chromosome-level assembly of the black tiger shrimp (Penaeus monodon) genome facilitates the identification of growth-associated genes. Mol Ecol Resour21, 1620–1640, 10.1111/1755-0998.13357 (2021). 10.1111/1755-0998.13357 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Polinski, J. M. et al. The American lobster genome reveals insights on longevity, neural, and immune adaptations. Sci Adv7, eabe8290, 10.1126/sciadv.abe829 (2021). 10.1126/sciadv.abe829 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Li, B. Y. et al. Chromosome-level genome assembly of the aphid parasitoid Aphidius gifuensis using Oxford Nanopore sequencing and Hi-C technology. Mol Ecol Resour21, 941–954, 10.1111/1755-0998.13308 (2021). 10.1111/1755-0998.13308 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Jin, S. et al. A chromosome-level genome assembly of the oriental river prawn, Macrobrachium nipponense. Gigascience10, giaa160, 10.1093/gigascience/giaa160 (2021). 10.1093/gigascience/giaa160 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Cui, Z. X. et al. The Chinese mitten crab genome provides insights into adaptive plasticity and developmental regulation. Nat Commun12, 1–13, 10.1038/s41467-021-22604-3 (2021). 10.1038/s41467-021-22604-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Jeong, C. B. et al. The genome of the harpacticoid copepod Tigriopus japonicus: Potential for its use in marine molecular ecotoxicology. Aquat Toxicol222, 105462, 10.1016/j.aquatox.2020.105462 (2020). 10.1016/j.aquatox.2020.105462 [DOI] [PubMed] [Google Scholar]
- 49.Zhang, X. J. et al. Penaeid shrimp genome provides insights into benthic adaptation and frequent molting. Nat Commun10, 356, 10.1038/s41467-018-08197-4 (2019). 10.1038/s41467-018-08197-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Gutekunst, J. et al. Clonal genome evolution and rapid invasive spread of the marbled crayfish. Nat Ecol Evol2, 567–573, 10.1038/s41559-018-0467-9 (2018). 10.1038/s41559-018-0467-9 [DOI] [PubMed] [Google Scholar]
- 51.Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics25, 1754–1760, 10.1093/bioinformatics/btp324 (2009). 10.1093/bioinformatics/btp324 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Felipe, A. S., Robert, M. W., Panagiotis, I., Evgenia, V. K. & Evgeny, M. Z. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics31, 3210–3212, 10.1093/bioinformatics/btv351 (2015). 10.1093/bioinformatics/btv351 [DOI] [PubMed] [Google Scholar]
- 53.Shi, L., Yi, S. & Li, Y. Genome survey sequencing of red swamp crayfish Procambarus clarkii. Mol Biol Rep45, 799–806, 10.1007/s11033-018-4219-3 (2018). 10.1007/s11033-018-4219-3 [DOI] [PubMed] [Google Scholar]
- 54.Jimenez, A. G., Kinsey, S. T., Dillaman, R. M. & Kapraun, D. F. Nuclear DNA content variation associated with muscle fiber hypertrophic growth in decapod crustaceans. Genome53, 161–171, (2010). 10.1139/G09-095 [DOI] [PubMed] [Google Scholar]
- 55.Kurtz, S. et al. Versatile and open software for comparing large genomes. Genome Biol5, 1–9, 10.1186/gb-2004-5-2-r12 (2004). 10.1186/gb-2004-5-2-r12 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Citations
- NCBI Sequence Read Archivehttp://identifiers.org/ncbi/insdc.sra:SRP516244 (2024).
- NCBI GenBankhttps://identifiers.org/ncbi/insdc:JBFOCG000000000 (2024).
- Liao, M. C. Genome annotation for the red swamp crayfish Procambarus clarkii. figshare10.6084/m9.figshare.24589233 (2023). 10.6084/m9.figshare.24589233 [DOI]
Data Availability Statement
All software and tools were used following the parameters mentioned in the Methods section or with default parameters.



