Abstract
Coconut, Cocos nucifera L., is the top agricultural export of the Philippines with about 1.2 billion pesos in value. The Philippine coconut industry, however, has been under major threat from an outbreak of the armored coconut scale insects. There are two species believed to cause an outbreak: Aspidiotus destructor Signoret and Aspidiotus rigidus Reyne. These are sibling species and are hard to differentiate using morphological identification. In fact, A. rigidus was once thought to be a subspecies of A. destructor and was misidentified as A. destructor during the early phases of the coconut scale insect outbreak because it is the only known species native to the Philippines. Aspidiotus rigidus has recently been identified as an invasive species and found to be the cause of the outbreak. The native species A. destructor should not be overlooked as a subject of research since it is still present alongside the A. rigidus, and it continues to infest several perennial crops. The need for a research approach of these insects at the molecular level has required the use of transcriptomics. Transcriptome datasets offer a way at investigating how genes are expressed, how species differ from each other, and how phenotypes came to be, among others. Transcriptomics offers such deeper understanding and can be used to develop methods for pest management. Because A. rigidus transcriptome has recently become available, it is imperative to have the dataset for its sibling species, A. destructor. This serves as a foundational resource; the first publicly available transcriptome assembly for the species. This will provide additional knowledge on how the two sibling species differ and assess their capacity to cause outbreak.
The data here represents the first transcriptome profile of the A. destructor using Illumina HiSeq 4000 paired- end sequencing. Pair-end reads were assembled de novo with Trinity. Raw fastq reads have been deposited in NCBI-SRA (SRR17085744 and SRR17085743). The Trinity-based transcriptome assembly have also been deposited in the NCBI-SRA (SUB10341747). Two additional assemblies were also generated and have been deposited in NCBI-SRA: an assembly clustered using CD-HIT-EST (SUB10341752) and an assembly sorted according to its longest assembly via a custom script (SUB10341753).
Keywords: Coconut scale insect, Transcriptomics, Aspidiotus, Cocos nucifera
Specifications Table
| Subject | Biology |
| Specific subject area | Transcriptome assembly and evaluation |
| Type of data | Raw RNA-seq data, Transcriptome assemblies |
| Data collection | Total RNA was extracted coconut scale insects 100 adult female and treated with DNase I. Oligo(dT) was used to isolate mRNA which was then fragmented and used in the cDNA synthesis. Short fragments were purified and resolved with EB buffer for end reparation and single nucleotide A (adenine) addition. Adapters were then ligated to the short fragments The suitable fragments were selected for PCR amplification. During the QC steps, Agilent 2100 Bioanalyzer and ABI StepOnePlus Real-Time PCR System were used in quantification and qualification of the sample libraries. The libraries produced using TruSeq RNA Library Kit as described above were sequenced using Illumina HiSeq 4000. |
| Data source location | Functional Genomics Laboratory National Institute of Molecular Biology and Biotechnology University of the Philippines Diliman, Quezon City 1101, Philippines |
| Data accessibility |
All data referred to in your data article must be made publicly available prior to publication. The raw Illumina sequencing data in FASTQ were deposited at NCBI SRA database under accession numbers SRR17085744 and SRR17085743. The transcriptome assemblies were deposited at NCBI TSA (Accession No. GGKE00000000). Accession numbers SUB10341747, SUB10341752, and SUB10341753 are, for the Trinity assembly, CD-HIT-EST-clustered assembly, and longest isoform-sorted assembly, respectively |
| Related research article |
1. Value of the Data
-
•
The data set is the first deposited source of RNA-seq raw reads and assemblies for the transcriptome for the coconut scale insect, Aspidiotus destructor Signoret.
-
•
The datasets provide other researchers the opportunity to discover genes related to several physiological processes involved in A. destructor or genes associated with its adaptive traits. Molecular markers can be developed from the available sequences and be used for its identification and distinction from its sibling species, A. rigidus. Molecular markers can also be developed for population genetic analysis.
-
•
The raw sequence reads will be accessible for further processing and analysis by researchers.
2. Background
The Philippines is the world’s second-largest producer of coconut products [1]. About 3.6 million hectares are planted with coconut trees, supporting 2.5 million farmers and contributing significantly to the nation’s GDP with 14.7 million tons of annual production [2]. From 2014 to 2018, coconut exports earned an average of 91.4 billion pesos yearly [2]. Although growth slowed to 0.9 % in late 2019 [3], coconut remains one of the country’s key crops and exports [4]. A major factor behind the slow growth is pest infestation, particularly from Aspidiotus species, armored scale insects (Hemiptera: Diaspididae) first found infecting coconut palms in Batangas in 2009 [1]. Two species are involved: Aspidiotus rigidus Reyne and Aspidiotus destructor Signoret. A. rigidus mainly infests coconut (Cocos nucifera), mangosteen, and banana, and is the main outbreak species in the Philippines [5]. A. destructor is polyphagous, attacking over 75 genera in 44 plant families, including coconut, mango, guava, and papaya [6]. It feeds on the undersides of leaves but rarely causes severe damage in older trees or those on well-drained soils [7].
3. Data Description
The datasets presented here were generated following the workflow described in Fig. 1. These datasets include the raw Illumina sequencing data obtained through sequencing of adult female Aspidiotus destructor Signoret (Hemiptera: Diaspididae) RNA with RNA Integrity Number (RIN) of 6.2 (Supplementary Figure 1) The raw Illumina sequencing data in FASTQ were deposited to NCBI SRA database under accession no. SRR17085744 and SRR17085743. De novo assembly of reads was performed using Trinity [8]. Two additional assemblies were also generated: an assembly clustered using CD-HIT-EST [9] and an assembly sorted according to its longest assembly via a custom script (Supplementary Table 1). Three Trinity-based assemblies were generated because each represents a different level of biological detail and redundancy, serving distinct downstream purposes. The raw Trinity de novo assembly retains all isoforms and transcript variants, maximizing biological completeness and suitability for annotation and isoform-level analyses but at the cost of high redundancy. The CD-HIT-EST–clustered assembly reduces sequence redundancy by collapsing highly similar transcripts, making it more appropriate for expression quantification, assembly quality assessment, and comparative analyses. The longest-isoform-per-gene assembly further simplifies the dataset by selecting a single representative transcript per Trinity gene, which is particularly useful for BUSCO [10] completeness assessment, orthology inference, and gene-level comparisons. Providing all three assemblies ensures broad reusability of the dataset and follows best practices for de novo transcriptome resources in non-model organisms. An overview of the RNA-seq data is presented in Table 1 and Fig. 2. The transcriptome assemblies were deposited to NCBI TSA database with accession numbers SUB10341747 for the Trinity assembly, SUB10341752 for the CD-HIT-EST-clustered assembly, and SUB10341753 for the longest isoform-sorted assembly. Assembly statistics were generated using Transrate [11] (Table 2 and are further summarized in Fig. 3 Assembly analysis for completeness using BUSCO is presented in Fig. 4, respectively. The contigs from the Trinity assembly were annotated against the NCBI nr database using Diamond-BLASTX [12]. Out of 107,407 assembled transcripts, the functional annotation file returned 41,085 unique sequence IDs and the taxonomic annotation file returned 18,593 unique sequence IDs. The functional annotation file is generated (Supplementary Table 2) and the taxonomic annotation file is generated separately (Supplementary Table 3). Abundance analysis was also performed by aligning the raw RNA-Seq reads against the three assemblies and were tabulated (Supplementary Table 4)
Fig. 1.
Overview of the RNA-seq data generation, transcriptome assembly, and data deposition workflow for Aspidiotus destructor. Adult female insects reared under laboratory conditions were used for total RNA extraction and Illumina paired-end RNA sequencing. Raw RNA-seq reads were deposited in the NCBI Sequence Read Archive (SRA). De novo transcriptome assembly was performed using Trinity, followed by non-redundant transcript clustering with CD-HIT-EST and representative transcript selection based on the longest isoform. The resulting transcriptome assemblies were deposited in the NCBI Transcriptome Shotgun Assembly (TSA) database.
Table 1.
Data production of RNA-seq. Generated using FASTQC version 0.11.5 [13].
| KLADe-1_1.fastq | KLADe_1_2.fastq | KLADe_2_1.fastq | KLADe_2_2.fastq | |
|---|---|---|---|---|
| File Type | Conventional base calls | |||
| Encoding | Illumina 1.5 | |||
| Sequence length | 100 | |||
| Total Sequences | 54,447,557 | 54,447,557 | 59,721,012 | 59,721,012 |
| Sequences flagged as poor quality | 0 | 0 | 0 | 0 |
| %GC | 39 | 39 | 40 | 40 |
| Percent of sequences remaining if deduplicated | 33.86 | 39.48 | 34.24 | 39.79 |
Fig. 2.
FastQC analysis demonstrating of the paired-end reads of Aspidiotus destructor Signoret RNA bioreplicates sequenced using Illumina HiSeq4000. The y-axis represents the Phred scores, and the x-axis represents the base positions in the reads (bp). All of the reads are above Q30 or these reads have 100 % base call accuracy (graphs at the top represent the paired end reads 1 and 2 of bioreplicate 1, graphs at the bottom represent the pair-end reads 1 and 2 of bioreplicate 2).
Table 2.
Assembly statistics. Generated using TransRate version 1.0.3 [11].
| Assembly | Trinity | CD-HIT-EST Clustered | Longest-Isoform Sorted |
|---|---|---|---|
| n_seqs | 107,407 | 93,360 | 69,179 |
| smallest | 200 | 200 | 201 |
| largest | 21,129 | 21,129 | 21,129 |
| n_bases | 147,369,086 | 108,260,737 | 59,701,807 |
| mean_len | 1371.37657 | 1159.13961 | 862.71794 |
| n_under_200 | 0 | 0 | 0 |
| n_over_1k | 38,990 | 28,762 | 13,836 |
| n_over_10k | 576 | 317 | 157 |
| n_with_orf | 34,685 | 25,598 | 12,656 |
| mean_orf_percent | 47.04794 | 46.47073 | 47.93761 |
| n90 | 468 | 378 | 291 |
| n70 | 1803 | 1430 | 753 |
| n50 | 3055 | 2624 | 2058 |
| n30 | 4670 | 4105 | 3636 |
| n10 | 7730 | 7047 | 6675 |
| gc | 0.38572 | 0.38455 | 0.38412 |
| bases_n | 0 | 0 | 0 |
| proportion_n | 0 | 0 | 0 |
| fragments | 114,168,569 | 114,168,569 | 114,168,569 |
| fragments_mapped | 42,720,402 | 55,093,671 | 104,085,511 |
| p_fragments_mapped | 0.37419 | 0.48256 | 0.91168 |
| good_mappings | 41,588,412 | 53,662,073 | 99,301,163 |
| p_good_mapping | 0.36427 | 0.47002 | 0.86978 |
| bad_mappings | 1131,990 | 1431,598 | 4784,348 |
| potential_bridges | 0 | 0 | 0 |
| bases_uncovered | 100,585,784 | 54,818,424 | 1022,713 |
| p_bases_uncovered | 0.68254 | 0.50636 | 0.01713 |
| contigs_uncovbase | 67,019 | 51,129 | 19,504 |
| p_contigs_uncovbase | 0.62366 | 0.54744 | 0.28184 |
| contigs_uncovered | 107,460 | 93,397 | 69,202 |
| p_contigs_uncovered | 1 | 1 | 1 |
| contigs_lowcovered | 107,460 | 93,397 | 69,202 |
| p_contigs_lowcovered | 1 | 1 | 1 |
| contigs_segmented | 12,166 | 11,915 | 8289 |
| p_contigs_segmented | 0.11321 | 0.12757 | 0.11978 |
| score | 0.07234 | 0.14817 | 0.47486 |
| optimal_score | 0.17523 | 0.2276 | 0.50458 |
| cutoff | 0.26934 | 0.17225 | 0.03944 |
| weighted | 0 | 0 | 0 |
Fig. 3.
Summary of transcript numbers across transcriptome assembly refinement and annotation steps for Aspidiotus destructor. The bar plot shows the proportion of transcripts retained relative to the initial Trinity-assembled transcriptome (n = 107,407) after non-redundant clustering with CD-HIT-EST, representative transcript selection based on the longest isoform per cluster, and functional annotation using DIAMOND against the NCBI nr database. Absolute transcript counts at each stage are indicated below each bar.
Fig. 4.
BUSCO assessment for the three transcriptome assemblies of lab-reared Aspidiotus destructor Signoret. The three assemblies (called trinity, cdhit, and longest) were assessed using datasets Eukaryota, Metazoa, Arthropoda, Insecta, and Hemiptera.
4. Experimental Design, Materials and Methods
4.1. Insect material and RNA extraction
Total RNA isolation was accomplished using TRIzol as the homogenizing reagent. One hundred (100) adult female coconut scale insects, A. destructor Signoret (Hemiptera: Diaspididae), reared on squash (Cucurbita moschata Duch) in the laboratory were collected and homogenized for 2 min while incubated on ice. It was followed by incubation at room temperature for 5 min. After which, 20 µL of chloroform was added and the sample was vortexed. It was incubated at room temperature for 3 min followed by centrifugation at 12,000 xg at 4 °C for 15 min. The aqueous phase was obtained and one volume of isopropanol was added. This was incubated at room temperature for 10 min and centrifuged at 12,000 xg at 4 °C for 10 min. Supernatant was discarded and 100 µL of 75 % ethanol was used to wash the pellet. It was centrifuged at 12,000 xg at 4 °C for 5 min. The supernatant was discarded and the pellet was air-dried. Isolated RNA was suspended in nuclease free water and stored in −80 °C before sending for sequencing. Two bioreplicates were prepared for sequencing.
4.2. Transcriptome assembly and analyses
The raw data of the two bioreplicates which were sequenced using Illumina HiSeq 4000. Scripts used the pipeline are in Supplementary Table 3. Reads were analyzed using FASTQC version 0.11.5 [13]. From the two RNA-Seq bioreplicates, three assemblies were generated: First, a de novo transcriptome assembly was generated with Trinity version 2.8.6 [8] by pooling the two RNA-Seq FASTQ files as input. Second, a CD-HIT-EST clustered assembly was generated using CD-HIT version 4.8.1 [9] using the de novo Trinity assembly as input. Third, an assembly sorted according to their longest isoform using the get_longest_isoform_seq_per_trinity_gene.pl custom script from github.com/trinityrnaseq [14] was also generated using the CD-HIT-EST assembly as input. The quality of the assemblies was assessed by Transrate version 1.0.3 [11]. All the assemblies were also analyzed for completeness using BUSCO version 3.1.0 [10].
4.3. Transcript abundance analysis
Functional annotation was performed using Diamond version 0.9.30 [12] with the Trinity assembly as input against the NCBI nr database. Abundance analysis was performed by aligning the raw RNA-Seq reads against the three assemblies. The alignment was performed using HISAT2 version 2.1.0 [15], and the transcript abundance in FPKM and TPM were calculated using StringTie version 2.1.0 [16]. The top 25 abundant transcripts were retrieved.
Limitations
The NCBI nr database used in this study was obtained in 2020, and all software employed for data processing in this article was also from 2020. Annotation coverage of the overall assembly for A. destructor appears low using the NCBI NR database in 2020. An insect-specific database closest to the coconut scale insect may improve assembly as data curation continues and analysis of newly collected and established samples progresses.
Ethics Statement
The authors have read and follow the ethical requirements for publication in Data in Brief and confirming that the current work does not involve human subjects, animal experiments, or any data collected from social media platforms.
CRediT Author Statement
Ma. Anita M. Bautista: Conceptualization, Methodology, Formal Analysis, Data curation, Writing- Original draft preparation, Writing- Reviewing and Editing. John Michael C. Egana: Methodology, Data Curation. Maria Almira S. Cleofe: Methodology, Data Curation. Jherico R. Geronca: Methodology, Data curation, Formal Analysis, Writing- Original draft preparation, Writing- Reviewing and Editing
Acknowledgments
Acknowledgements
This work was supported by the Balik-PhD grant of the Office of the Vice President for Academic Affairs, University of the Philippines to MAM Bautista. The authors wish to thank Dr. Barbara L. Caoili of the Institute of Weed Science, Entomology and Plant Pathology, College of Agriculture and Food Science, University of the Philippines Los Baños, College, Laguna Philippines for providing the Aspidiotus destructor samples.
Declaration of Competing Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Footnotes
Supplementary material associated with this article can be found, in the online version, at doi:10.1016/j.dib.2026.112619.
Appendix. Supplementary materials
Data Availability
NCBI TSAAspidiotus destructor Lab-reared CD-HIT-EST clustered assembly (Original data)
NCBI TSAAspidiotus destructor Lab-reared Longest isoform sorted assembly (Original data)
NCBI SRARNA-seq of Aspidiotus destructor Signoret Biological Replicate 1 (Original data)
NCBI SRARNA-seq of Aspidiotus destructor Signoret Biological Replicate 2 (Original data)
NCBI TSAAspidiotus destructor Lab-reared Trinity assembly (Original data)
References
- 1.Watson G.W., Adalla C.B., Shepard B.M., Carner G.R. Aspidiotus rigidus Reyne (Hemiptera: diaspididae): a devastating pest of coconut in the Philippines. Agric. For. Entomol. 2015;17(1):1–8. doi: 10.1111/afe.12078. [DOI] [Google Scholar]
- 2.Philippine Coconut Authority Coconut: major export crop of Filipino farmers. Philippine Coconut Authority. 2019 https://pca.gov.ph/index.php [Google Scholar]
- 3.Philippine Statistics Authority Major non-food and industrial crops quarterly bulletin, October–December 2019. Philippine Statistics Authority. 2019 https://psa.gov.ph [Google Scholar]
- 4.Castillo M., Ani P.A. The Philippine coconut industry: status, policies and strategic directions for development. FFTC Agricultural Policy Platform. 2019 https://ap.fftc.agnet.org [Google Scholar]
- 5.Caoili B.L., Lit I.L., Jr., Yap S.A., Guerrero M.S., Sapin G.D., Sandoval R.F.C., Belen J.M., Latina R.A., Barbecho N.M. Identification of the Philippine outbreak populations of the coconut scale insects using molecular markers with morphological confirmation. Philip. Entomol. 2014;28(2):202–225. [Google Scholar]
- 6.Davidson J.A., Miller D.R. 4 (B) Elsevier; Amsterdam, the Netherlands: 1990. Ornamental plants; pp. 603–632. [Google Scholar]
- 7.CABI Aspidiotus destructor (coconut scale) CABI. 2019 https://www.cabi.org [Google Scholar]
- 8.Grabherr M.G., Haas B.J., Yassour M., Levin J.Z., Thompson D.A., Amit I.…Regev A. Full-length transcriptome assembly from RNA-seq data without a reference genome. Nat. Biotechnol. 2011;29(7):644–652. doi: 10.1038/nbt.1883. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Fu L., Niu B., Zhu Z., Wu S., Li W. CD-HIT: accelerated for clustering the next-generation sequencing data. Bioinformatics. 2012;28(23):3150–3152. doi: 10.1093/bioinformatics/bts565. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Simão F., Waterhouse R., Ioannidis P., Krivetseva E., Zdobnov E. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics. 2015;31(19) doi: 10.1093/bioinformatics/btv351. [DOI] [PubMed] [Google Scholar]
- 11.Smith-Unna R., Boursnell C., Patro R., Hibberd J.M., Kelly S. TransRate: reference-free quality assessment of de novo transcriptome assemblies. Genome Res. 2016;26(8):1134–1144. doi: 10.1101/gr.196469.115. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Buchfink B., Xie C., Huson D.H. Fast and sensitive protein alignment using DIAMOND. Nat. Methods. 2015;12(1):59–60. doi: 10.1038/nmeth.3176. [DOI] [PubMed] [Google Scholar]
- 13.S. Andrews, (2010). Babraham bioinformatics-FastQC a quality control tool for high throughput sequence data. Retrieved from www.bioinformatics.babraham.ac.uk/.
- 14.Haas B.J., Papanicolaou A., Yassour M., Grabherr M., Blood P.D., Bowden J.…Regev A. De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nat. Protoc. 2013;8(8):1494–1512. doi: 10.1038/nprot.2013.084. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Kim D., Paggi J.M., Park C., Bennett C., Salzberg S.L. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat. Biotechnol. 2019;37(8):907–915. doi: 10.1038/s41587-019-0201-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Pertea M., Kim D., Pertea G.M., Leek J.T., Salzberg S.L. Transcript-level expression analysis of RNA-seq experiments with HISAT, StringTie and Ballgown. Nat. Protoc. 2016;11(9):1650–1667. doi: 10.1038/nprot.2016.095. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
NCBI TSAAspidiotus destructor Lab-reared CD-HIT-EST clustered assembly (Original data)
NCBI TSAAspidiotus destructor Lab-reared Longest isoform sorted assembly (Original data)
NCBI SRARNA-seq of Aspidiotus destructor Signoret Biological Replicate 1 (Original data)
NCBI SRARNA-seq of Aspidiotus destructor Signoret Biological Replicate 2 (Original data)
NCBI TSAAspidiotus destructor Lab-reared Trinity assembly (Original data)




