Skip to main content
Wellcome Open Research logoLink to Wellcome Open Research
. 2026 Mar 19;8:107. Originally published 2023 Mar 1. [Version 2] doi: 10.12688/wellcomeopenres.19066.2

The genome sequence of the Common Yellow Sally, Isoperla grammatica (Poda, 1761) (Plecoptera: Perlodidae)

Emma McSwan 1, Caleala Clifford 2; Natural History Museum Genome Acquisition Lab; Darwin Tree of Life Barcoding collective; Wellcome Sanger Institute Tree of Life programme; Wellcome Sanger Institute Scientific Operations: DNA Pipelines collective; Tree of Life Core Informatics collective, Craig R Macadam 3, Benjamin W Price 4; Darwin Tree of Life Consortiuma
PMCID: PMC13033136  PMID: 41913785

Version Changes

Revised. Amendments from Version 1

In Version  2 of this data note we have added information to the Background section to contextualise our sequencing project with respect to other molecular studies in Plecoptera, and as part of the Darwin Tree of Life project. We have expanded on the results of the sequencing runs (new Table 1), and added more detail to the assembly methods section.  A new section on "Assembly quality assessment" has been added to the Methods. We have replaced Figure 5 with a new labelled version in PretextView for readability.  In the Data availability section we added links to the Tree of Life production pipeline suite and to sequence and metadata for this species.

Abstract

We present a genome assembly from an individual male Isoperla grammatica (the Common Yellow Sally; Arthropoda; Insecta; Plecoptera; Perlodidae). The genome sequence is 874.6 megabases in span. Most of the assembly is scaffolded into 14 chromosomal pseudomolecules, including the assembled X 1 and X 2 chromosomes. The mitochondrial genome has also been assembled and is 16.2 kilobases in length. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Keywords: Isoperla grammatica, Common Yellow Sally, genome sequence, chromosomal, Plecoptera

Species taxonomy

Eukaryota; Metazoa; Ecdysozoa; Arthropoda; Hexapoda; Insecta; Pterygota; Neoptera; Polyneoptera; Plecoptera; Perloidea; Perlodidae; Isoperlinae; Isoperla; Isoperla grammatica (Poda, 1761) (NCBI:txid552050).

Background

Isoperla grammatica ( Figure 1) is a western Palaearctic species found across Europe from France to Romania, south to Sicily and north to the Baltic and Fennoscandia ( Dewalt et al., 2023). It is found throughout Britain and Ireland and can be very common in some watercourses.

Figure 1. Isoperla grammatica © Jon Mortin (CC BY) Source: https://www.inaturalist.org/photos/146228992.

Figure 1.

It is considered a eurytherm ( Graf et al., 2009) and is typically occurs in high densities in all lotic water types with stable and unstable substrata, amongst moss, leaf packets and gravel, and is often present in rivers with slight organic enrichment ( Baars & Kelly-Quinn, 2006; Costello, 1988; Frost, 1942). The widespread distribution of this species indicates that it has no preference for particular pH conditions and has been found in both neutral and episodically acidic waters ( Feeley et al., 2011; Feeley, 2012; Feeley & Kelly-Quinn, 2014; Murphy et al., 2013).

Isoperla grammatica has a univoltine life cycle ( Frost, 1942; Smith et al., 2000) with larvae present for part of two summers and the intervening winter ( Elliott, 1967; Langford, 1971; Malmqvist & Sjöström, 1989; Ulfstrand, 1968). Research in Norway and Britain indicated that eggs need warm temperatures of 7 to 12°C to initiate development, but the optimum incubation temperature is 16°C ( Elliott, 1991; Elliott, 1995; Lillehammer et al., 1989; Saltveit & Lllehammer, 1984). Larvae occur all year round in small numbers and across various sizes, indicating variability in larval growth ( Langford, 1971; Malmqvist & Sjöström, 1989; Smith et al., 2000). However, larvae typically grow rapidly in autumn and spring, although winter growth has also been noted where water temperatures were suitable ( Malmqvist & Sjöström, 1989).

Although diatom and other algal matter are also ingested, the larvae of I. grammatica are carnivorous from very early instars ( Graf et al., 2009; Malmqvist & Sjöström, 1989; Malmqvist et al., 1991). Larvae are also highly selective in their prey items ( Williams, 1987). Of the prey items found, Chironomidae and Simuliidae seem to dominate ( Elliott, 2004; Malmqvist & Sjöström, 1989; Malmqvist et al., 1991), with Williams highlighting a preference for Baetidae in Wales ( Williams, 1987). Elliott ( 2000, 2004) indicated that the feeding behaviour was by active search and was limited to the hours of dusk and dawn, with little activity during the day or at night. The adults feed on a range of pollens, fungi and fine particulate organic matter ( Tierno de Figueroa & Sánchez-Ortega, 1999).

Another contig-level assembly for this species is also available (GCA_001676475.1; submitted by Hannah C Macdonald) (data obtained via NCBI datasets, O’Leary et al., 2024). In addition, work using transcriptome and whole genome sequencing to examine the evolutionary history and taxonomy of Plecoptera has been published ( Letsch et al., 2021; South et al., 2021).

We report a chromosome-level complete genome sequence for Isoperla grammatica. This assembly was generated as part of the Darwin Tree of Life Project, which aims to generate high-quality reference genomes for all named eukaryotic species in Britain and Ireland to support research, conservation, and the sustainable use of biodiversity ( Blaxter et al., 2022).

Genome sequence report

Sequencing data

PacBio sequencing of the Isoperla grammatica specimen generated 25.55 Gb (gigabases) from 2.94 million reads, which were used to assemble the genome. GenomeScope2.0 analysis estimated the haploid genome size at 672.52 Mb, with a heterozygosity of 3.14% and repeat content of 47.49%. These estimates guided expectations for the assembly. Based on the estimated genome size, the sequencing data provided approximately 34× coverage. Hi-C sequencing produced 109.98 Gb from 728.32 million reads, which were used to scaffold the assembly. RNA sequencing data were also generated and are available in public sequence repositories, but not used in the assembly. Table 1 summarises the specimen and sequencing details.

Table 1. Specimen and sequencing data for BioProject PRJEB53729.

Platform PacBio HiFi Hi-C RNA-seq
ToLID ipIsoGram3 ipIsoGram4 ipIsoGram7
Specimen ID NHMUK014360609 NHMUK014360651 NHMUK014361594
BioSample (source individual) SAMEA7520999 SAMEA7521000 SAMEA7521375
BioSample (tissue) SAMEA7521099 SAMEA7521100 SAMEA7521453
Tissue whole organism whole organism whole organism
Instrument Sequel II Illumina NovaSeq 6000 Illumina HiSeq 4000
Run accessions ERR9878388; ERR9878387 ERR9881687 ERR9881692
Read count total 2.94 million 728.32 million 29.52 million
Base count total 25.55 Gb 109.98 Gb 4.46 Gb

Assembly statistics

Manual assembly curation of the assembly was done to to confirm chromosome boundaries. We also corrected 400 missing or mis-joins and removed 47 haplotypic duplications, reducing the assembly length by 1.18% and the scaffold number by 24.31%, and increasing the scaffold N50 by 27.94%.

The final assembly has a total length of 874.6 Mb in 682 sequence scaffolds with a scaffold N50 of 56.8 Mb ( Table 2). Most (95.16%) of the assembly sequence was assigned to 14 chromosomal-level scaffolds, representing 12 autosomes, and the X 1 and X 2 sex chromosomes. Chromosome-scale scaffolds confirmed by the Hi-C data have been named in order of size. ( Figure 2Figure 5; Table 3). The scaffold order and orientation are uncertain in the following regions: chromosome 8 (29.54–40.94 Mb), chromosome 9 (2.58–20.91 Mb), and chromosome 11 (24.65–31.18 Mb). While not fully phased, the assembly deposited is of one haplotype. Contigs corresponding to the second haplotype have also been deposited.

Figure 2. Genome assembly of Isoperla grammatica, ipIsoGram3.1: metrics.

Figure 2.

The BlobToolKit Snailplot shows N50 metrics and BUSCO gene completeness. The main plot is divided into 1,000 size-ordered bins around the circumference with each bin representing 0.1% of the 874,600,353 bp assembly. The distribution of scaffold lengths is shown in dark grey with the plot radius scaled to the longest scaffold present in the assembly (137,496,831 bp, shown in red). Orange and pale-orange arcs show the N50 and N90 scaffold lengths (56,757,646 and 40,559,246 bp), respectively. The pale grey spiral shows the cumulative scaffold count on a log scale with white scale lines showing successive orders of magnitude. The blue and pale-blue area around the outside of the plot shows the distribution of GC, AT and N percentages in the same bins as the inner plot. A summary of complete, fragmented, duplicated and missing BUSCO genes in the insecta_odb10 set is shown in the top right. An interactive version of this figure is available at https://blobtoolkit.genomehubs.org/view/ipIsoGram3.1/dataset/CAMDTW01/snail.

Figure 5. Genome assembly of Isoperla grammatica, ipIsoGram3.1: Hi-C contact map.

Figure 5.

Hi-C contact map of the ipIsoGram3.1 assembly, visualised using PretextView and PretextSnapshot. Chromosomes are shown in order of size from left to right and top to bottom. An interactive version of this figure may be viewed in HiGlass at https://genome-note-higlass.tol.sanger.ac.uk/l/?d=FHx1cM8_RE6963zmzuaDVw.

Table 2. Genome assembly statistics.

Assembly name ipIsoGram3.1
Assembly accession GCA_945910005.1
Alternate haplotype accession GCA_945909985.1
Assembly level chromosome
Span (Mb) 874.58
Number of chromosomes 14
Number of contigs 2,826
Contig N50 0.65 Mb
Number of scaffolds 682
Scaffold N50 56.76 Mb
Sex chromosomes X 1 and X 2
Organelles Mitochondrion: 16.16 kb
Metric (benchmark) Values achieved
Consensus quality QV (≥ 40) Primary: 55.3; alternate: 56.2; combined: 55.8
k-mer completeness (≥ 95%) Primary: 68.59%; alternate: 56.02%; combined: 98.25%
BUSCO * (S > 90%; D < 5%) C:99.3%[S:96.6%,D:2.7%],
F:0.3%,M:0.4%,n:1,367
Percentage of assembly
assigned to chromosomes
(≥ 90%)
95.16%

* BUSCO scores based on the insecta_odb10 BUSCO set using 5.3.2. C = complete [S = single copy, D = duplicated], F = fragmented, M = missing, n = number of orthologues in comparison. A full set of BUSCO scores is available at https://blobtoolkit.genomehubs.org/view/ipIsoGram3.1/dataset/CAMDTW01/busco.

Figure 3. Genome assembly of Isoperla grammatica, ipIsoGram3.1: BlobToolKit Blob plot.

Figure 3.

Scaffolds are coloured by phylum. Circles are sized in proportion to scaffold length. Histograms show the distribution of scaffold length sum along each axis. An interactive version of this figure is available at https://blobtoolkit.genomehubs.org/view/ipIsoGram3.1/dataset/CAMDTW01/blob.

Figure 4. Genome assembly of Isoperla grammatica, ipIsoGram3.1: cumulative sequence.

Figure 4.

BlobToolKit cumulative sequence plot. The grey line shows cumulative length for all scaffolds. Coloured lines show cumulative lengths of scaffolds assigned to each phylum using the buscogenes taxrule. An interactive version of this figure is available at https://blobtoolkit.genomehubs.org/view/ipIsoGram3.1/dataset/CAMDTW01/cumulative.

Table 3. Chromosomal pseudomolecules in the genome assembly of Isoperla grammatica, ipIsoGram3.

INSDC accession Chromosome Size (Mb) GC%
OX246737.1 1 137.5 39.3
OX246745.1 X1 45.93 39.7
OX246738.1 2 115.38 39
OX246739.1 3 75.84 39.6
OX246746.1 X2 45.27 38.6
OX246740.1 4 60.58 39.5
OX246741.1 5 56.76 39.9
OX246742.1 6 48.65 41
OX246743.1 7 48.25 41
OX246744.1 8 46.81 41.3
OX246747.1 9 41.61 41.5
OX246748.1 10 41.44 40.8
OX246749.1 11 40.56 42.4
OX246750.1 12 26.74 39.9
OX246751.1 MT 0.02 31.3
- unplaced 43.26 40

The mitochondrial genome was also assembled (length 16.16 kb, OX246751.1). This sequence is included as a contig in the multifasta file of the genome submission and as a standalone record.

The primary assembly has a BUSCO v5.3.2 ( Manni et al., 2021) completeness of 99.3% (single 96.6%, duplicated 2.7%), using the insecta_odb10 reference set. The combined primary and alternate assemblies achieve an estimated QV of 55.8. The k-mer completeness is 68.59% for the primary assembly, 56.02% for the alternate haplotype, and 98.25% for the combined assemblies.

Methods

Sample acquisition and nucleic acid extraction

Two Isoperla grammatica specimens (specimen ID NHMUK014360609, ToLID ipIsoGram3 and specimen ID NHMUK014360651, ToLID ipIsoGram4) were collected from River Test, Great Bridge, Hampshire (latitude 51.00, longitude –1.50) on 19 March 2019. The specimens were taken from freshwater by Emma McSwan (Environment Agency) using a kick-net. The specimen was also identified by Emma McSwan and snap-frozen in a dry shipper at the Natural History Museum, London. A specimen used for RNA sequencing (specimen ID NHMUK014361594, ToLID ipIsoGram7) was collected by Caleala Clifford (Natural Resources Wales) from River Taff Fawr, Garwnant, UK (latitude 51.81, longitude –-3.44) on 19 March 2019, and snap-frozen in a dry shipper at the Natural History Museum, London.

DNA was extracted at the Tree of Life laboratory, Wellcome Sanger Institute (WSI). The ipIsoGram3 specimen was weighed and dissected on dry ice. The tissue was cryogenically disrupted to a fine powder using a Covaris cryoPREP Automated Dry Pulveriser, receiving multiple impacts. High molecular weight (HMW) DNA was extracted using the Qiagen MagAttract HMW DNA extraction kit. HMW DNA was sheared into an average fragment size of 12–20 kb in a Megaruptor 3 system with speed setting 30. Sheared DNA was purified by solid-phase reversible immobilisation using AMPure PB beads with a 1.8X ratio of beads to sample to remove the shorter fragments and concentrate the DNA sample. The concentration of the sheared and purified DNA was assessed using a Nanodrop spectrophotometer and Qubit Fluorometer and Qubit dsDNA High Sensitivity Assay kit. Fragment size distribution was evaluated by running the sample on the FemtoPulse system.

RNA was extracted from tissue of ipIsoGram7 in the Tree of Life Laboratory at the WSI using TRIzol, according to the manufacturer’s instructions. RNA was then eluted in 50 μl RNAse-free water and its concentration assessed using a Nanodrop spectrophotometer and Qubit Fluorometer using the Qubit RNA Broad-Range (BR) Assay kit. Analysis of the integrity of the RNA was done using Agilent RNA 6000 Pico Kit and Eukaryotic Total RNA assay.

Sequencing

Pacific Biosciences HiFi circular consensus DNA sequencing libraries were constructed according to the manufacturers’ instructions. Poly(A) RNA-Seq libraries were constructed using the NEB Ultra II RNA Library Prep kit. DNA and RNA sequencing was performed by the Scientific Operations core at the WSI on Pacific Biosciences SEQUEL II (HiFi) and Illumina HiSeq 4000 (RNA-Seq). Hi-C data were also generated from ipIsoGram4 using the Arima v2 kit and sequenced on the NovaSeq 6000 instrument.

Genome assembly and curation

Prior to assembly of the PacBio HiFi reads, a database of k-mer counts ( k = 31) was generated from the filtered reads using FastK. GenomeScope2 ( Ranallo-Benavidez et al., 2020) was used to analyse the k-mer frequency distributions, providing estimates of genome size, heterozygosity, and repeat content.

Assembly of PacBio HiFi reads was carried out with Hifiasm ( Cheng et al., 2021) and haplotypic duplication was identified and removed with purge_dups ( Guan et al., 2020). The Hi-C reads ( Rao et al., 2014) were mapped to the primary contigs using bwa-mem2 ( Vasimuddin et al., 2019) and then scaffolded using YaHS ( Zhou et al., 2023). The assembly was checked for contamination and corrected as described previously ( Howe et al., 2021). Manual curation was performed using HiGlass ( Kerpedjiev et al., 2018) and PretextView ( Harry, 2022). The curation process is documented at https://gitlab.com/wtsi-grit/rapid-curation. PretextSnapshot was used to generate a Hi-C contact map of the final assembly.

The mitochondrial genome was assembled using MitoHiFi ( Uliano-Silva et al., 2023), which performed annotation using MitoFinder ( Allio et al., 2020).

Assembly quality assessment

The Merqury.FK tool (Rhie et al., 2020) was run in a Singularity container ( Kurtzer et al., 2017) to evaluate k-mer completeness and assembly quality for the primary and alternate haplotypes using the k-mer databases ( k = 31) computed prior to genome assembly. The analysis outputs included assembly QV scores and completeness statistics.

The genome was analysed within the BlobToolKit environment ( Challis et al., 2020) and BUSCO scores ( Manni et al., 2021) were calculated. A Hi-C map for the final assembly was produced using bwa-mem2 ( Vasimuddin et al., 2019) in the Cooler file format ( Abdennur & Mirny, 2020).

Table 4 contains a list of relevant software tool versions and sources.

Table 4. Software tools and versions used.

Ethics and compliance issues

The materials that have contributed to this genome note have been supplied by a Darwin Tree of Life Partner. The submission of materials by a Darwin Tree of Life Partner is subject to the Darwin Tree of Life Project Sampling Code of Practice. By agreeing with and signing up to the Sampling Code of Practice, the Darwin Tree of Life Partner agrees they will meet the legal and ethical requirements and standards set out within this document in respect of all samples acquired for, and supplied to, the Darwin Tree of Life Project. All efforts are undertaken to minimise the suffering of animals used for sequencing. Each transfer of samples is further undertaken according to a Research Collaboration Agreement or Material Transfer Agreement entered into by the Darwin Tree of Life Partner, Genome Research Limited (operating as the Wellcome Sanger Institute), and in some circumstances other Darwin Tree of Life collaborators.

Funding Statement

This work was supported by Wellcome through core funding to the Wellcome Sanger Institute (206194, <a href=https://doi.org/10.35802/206194>https://doi.org/10.35802/206194</a>) and the Darwin Tree of Life Discretionary Award (218328, <a href=https://doi.org/10.35802/218328>https://doi.org/10.35802/218328</a>).

The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

[version 2; peer review: 1 approved, 3 approved with reservations]

Data availability

European Nucleotide Archive: Isoperla grammatica (common yellow sally). Accession number PRJEB53729; https://identifiers.org/ena.embl/PRJEB53729.

The genome sequence is released openly for reuse by the Wellcome Sanger Institute. The Isoperla grammatica genome sequencing initiative is part of the Darwin Tree of Life (DToL) project. All raw sequence data and the assembly have been deposited in INSDC databases. The genome will be annotated using available RNA-Seq data and presented through the Ensembl pipeline at the European Bioinformatics Institute. Raw data and assembly accession identifiers are reported in Table 1 and Table 2.

Production code used in genome assembly at the WSI Tree of Life is available at https://github.com/sanger-tol. And 4 are

Metadata for specimens, BOLD barcode results, spectra estimates, sequencing runs, contaminants and pre-curation assembly statistics are given at https://links.tol.sanger.ac.uk/species/552050.

Author information

Members of the Natural History Museum Genome Acquisition Lab are listed here: https://doi.org/10.5281/zenodo.4790043.

Members of the Darwin Tree of Life Barcoding collective are listed here: https://doi.org/10.5281/zenodo.4893703.

Members of the Wellcome Sanger Institute Tree of Life programme are listed here: https://doi.org/10.5281/zenodo.4783585.

Members of Wellcome Sanger Institute Scientific Operations: DNA Pipelines collective are listed here: https://doi.org/10.5281/zenodo.4790455.

Members of the Tree of Life Core Informatics collective are listed here: https://doi.org/10.5281/zenodo.5013541.

Members of the Darwin Tree of Life Consortium are listed here: https://doi.org/10.5281/zenodo.4783558.

References

  1. Abdennur N, Mirny LA: Cooler: scalable storage for Hi-C data and other genomically labeled arrays. Bioinformatics. 2020;36(1):311–316. 10.1093/bioinformatics/btz540 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Allio R, Schomaker-Bastos A, Romiguier J, et al. : MitoFinder: Efficient automated large‐scale extraction of mitogenomic data in target enrichment phylogenomics. Mol Ecol Resour. 2020;20(4):892–905. 10.1111/1755-0998.13160 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Baars JR, Kelly-Quinn M: The Plecoptera of Irish freshwaters - species distribution, status and association with environmental parameters. Report to the Heritage Council. Reference no. 14525. [Preprint], 2006. [Google Scholar]
  4. Blaxter M, Mieszkowska N, Di Palma F, et al. : Sequence locally, think globally: the Darwin Tree of Life project. Proc Natl Acad Sci U S A. 2022;119(4): e2115642118. 10.1073/pnas.2115642118 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Challis R, Richards E, Rajan J, et al. : BlobToolKit - interactive quality assessment of genome assemblies. G3 (Bethesda). 2020;10(4):1361–1374. 10.1534/g3.119.400908 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Cheng H, Concepcion GT, Feng X, et al. : Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021;18(2):170–175. 10.1038/s41592-020-01056-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Costello MJ: A review of the distribution of stoneflies (Insecta, Plecoptera) in Ireland. Proc R Ir Acad B. 1988;88B(4):1–22. Reference Source [Google Scholar]
  8. DeWalt RE, Maehr MD, Hopkins H, et al. : Plecoptera Species File Online. Version 5.0/5.0. [2024/03/23].2023. Reference Source
  9. Elliott JM: The life histories and drifting of the plecoptera and ephemeroptera in a dartmoor stream. J Anim Ecol. 1967;36(2):343–362. 10.2307/2918 [DOI] [Google Scholar]
  10. Elliott JM: The effect of temperature on egg hatching for three populations of Isoperla grammatica and one population of Isogenus nubecula (Plecoptera: Perlodidae). Entomol Gaz. 1991;42:61–65. [Google Scholar]
  11. Elliott JM: Egg hatching and ecological partitioning in carnivorous stoneflies (Plecoptera). C R Acad Sci III. 1995;318(2):237–243. [Google Scholar]
  12. Elliott JM: Prey switching in four species of carnivorous stoneflies. Freshw Biol. 2004;49(6):709–720. 10.1111/j.1365-2427.2004.01222.x [DOI] [Google Scholar]
  13. Elliott JM: Contrasting diel activity and feeding patterns of four species of carnivorous stoneflies. Ecol Entomol. 2000;25(1):26–34. 10.1046/j.1365-2311.2000.00229.x [DOI] [Google Scholar]
  14. Feeley HB: The impact of mature conifer forest plantations on the hydrochemical and ecological quality of headwater streams in Ireland, with particular reference to episodic acidification. National University of Ireland (UCD), 2012. [Google Scholar]
  15. Feeley HB, Kelly-Quinn M: Re-examining the effects of episodic acidity on macroinvertebrates in small conifer-forested streams in Ireland and empirical evidence for biological recovery. Biology and Environment: Proceedings of the Royal Irish Academy. 2014;114B(3):205–218. 10.3318/bioe.2014.18 [DOI] [Google Scholar]
  16. Feeley HB, Kerrigan C, Fanning P, et al. : Longitudinal extent of acidification effects of plantation forest on benthic macroinvertebrate communities in soft water streams: evidence for localised impact and temporal ecological recovery. Hydrobiologia. 2011;671:217–226. 10.1007/s10750-011-0719-z [DOI] [Google Scholar]
  17. Frost WE: River Liffey survey: IV. The fauna of the submerged “mosses” in an acid and an alkaline water. Proc R Ir Acad B. 1942;47:293–369. Reference Source [Google Scholar]
  18. Graf W, et al. : Plecoptera: Volume 2.In: A. and D.H. (eds) Schmidt-Kloiber (ed.) Distribution and Ecological Preferences of European Freshwater Organisms. 2009;1–262. [Google Scholar]
  19. Guan D, McCarthy SA, Wood J, et al. : Identifying and removing haplotypic duplication in primary genome assemblies. Bioinformatics. 2020;36(9):2896–2898. 10.1093/bioinformatics/btaa025 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Harry E: PretextView (Paired REad TEXTure Viewer): A desktop application for viewing pretext contact maps. 2022; (Accessed: 19 October 2022). Reference Source [Google Scholar]
  21. Howe K, Chow W, Collins J, et al. : Significantly improving the quality of genome assemblies through curation. GigaScience. Oxford University Press, 2021;10(1): giaa153. 10.1093/gigascience/giaa153 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Kerpedjiev P, Abdennur N, Lekschas F, et al. : HiGlass: web-based visual exploration and analysis of genome interaction maps. Genome Biol. 2018;19(1): 125. 10.1186/s13059-018-1486-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Kurtzer GM, Sochat V, Bauer MW: Singularity: scientific containers for mobility of compute. PLoS ONE. 2017;12(5): e0177459. 10.1371/journal.pone.0177459 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Langford TE: The distribution, abundance and life-histories of stoneflies (Plecoptera) and mayflies (Ephemeroptera) in a British river, warmed by cooling-water from a power station. Hydrobiologia. 1971;38:339–377. 10.1007/BF00036844 [DOI] [Google Scholar]
  25. Letsch H, Simon S, Frandsen PB, et al. : Combining molecular datasets with strongly heterogeneous taxon coverage enlightens the peculiar biogeographic history of stoneflies (Insecta: Plecoptera). Syst Entomol. 2021;46(4):952–967. 10.1111/syen.12505 [DOI] [Google Scholar]
  26. Lillehammer A, Brittain JE, Saltveit SJ, et al. : Egg development, nymphal growth and life cycle strategies in Plecoptera. Holarctic Ecology. 1989;12(2):173–186. 10.1111/j.1600-0587.1989.tb00836.x [DOI] [Google Scholar]
  27. Malmqvist B, Sjöström P: The life cycle and growth of Isoperla grammatica and I. difformis (Plecoptera) in southernmost Sweden: intra and interspecific considerations. Hydrobiologia. 1989;175:97–108. 10.1007/BF00765120 [DOI] [Google Scholar]
  28. Malmqvist B, Sjöström P, Frick K: The diet of two species of Isoperla (Plecoptera: Perlodidae) in relation to season, site, and sympatry. Hydrobiologia. 1991;213:191–203. 10.1007/BF00016422 [DOI] [Google Scholar]
  29. Manni M, Berkeley MR, Seppey M, et al. : BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol Biol Evol. 2021;38(10): 4647–4654. 10.1093/molbev/msab199 [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Murphy JF, Davy-Bowker J, McFarland B, et al. : A diagnostic biotic index for assessing acidity in sensitive streams in Britain. Ecol Indic. 2013;24:562–572. 10.1016/j.ecolind.2012.08.014 [DOI] [Google Scholar]
  31. O’Leary NA, Cox E, Holmes JB, et al. : Exploring and retrieving sequence and metadata for species across the Tree of Life with NCBI Datasets. Sci Data. 2024;11(1): 732. 10.1038/s41597-024-03571-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Ranallo-Benavidez TR, Jaron KS, Schatz MC: GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes. Nat Commun. 2020;11(1):1432. 10.1038/s41467-020-14998-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Rao SSP, Huntley MH, Durand NC, et al. : A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell. 2014;159(7):1665–80. 10.1016/j.cell.2014.11.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Rhie A, McCarthy SA, Fedrigo O, et al. : Towards complete and error-free genome assemblies of all vertebrate species. Nature. 2021;592(7856):737–746. 10.1038/s41586-021-03451-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Saltveit SJ, Lllehammer A: Studies on egg development in the Fennoscandian Isoperla species (Plecoptera). Annales de Limnologie. 1984;20(1–2):91–94. 10.1051/limn/1984027 [DOI] [Google Scholar]
  36. Smith C, Good M, Murphy JF, et al. : Life-history patterns and spatial and temporal overlap in an assemblage of lotic Plecoptera in the Araglin Catchment Study Area, Ireland. Archive für Hydrobiologie. 2000;150(1):117–132. 10.1127/archiv-hydrobiol/150/2000/117 [DOI] [Google Scholar]
  37. South EJ, Skinner RK, Edward DeWalt R, et al. : A New Family of Stoneflies (Insecta: Plecoptera), Kathroperlidae, fam. n., with a Phylogenomic Analysis of the Paraperlinae (Plecoptera: Chloroperlidae). Insect Syst Divers. 2021;5(4). 10.1093/isd/ixab014 [DOI] [Google Scholar]
  38. Tierno de Figueroa JM, Sánchez-Ortega A: Imaginal feeding of certain systellognathan stonefly species (Insecta: Plecoptera). Ann Entomol Soc Am. 1999;92(2):218–221. 10.1093/aesa/92.2.218 [DOI] [Google Scholar]
  39. Ulfstrand S: Life cycles of benthic insects in Lapland streams (Ephemeroptera, Plecoptera, Trichoptera, Diptera Simuliidae). Oikos. 1968;19(2):167–190. 10.2307/3565005 [DOI] [Google Scholar]
  40. Uliano-Silva M, Ferreira JGRN, Krasheninnikova K, et al. : MitoHiFi: a python pipeline for mitochondrial genome assembly from PacBio high fidelity reads. BMC Bioinformatics. 2023;24(1): 288. 10.1186/s12859-023-05385-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Vasimuddin M, Misra S, Li H, et al. : Efficient architecture-aware acceleration of BWA-MEM for multicore systems. In: 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS).IEEE,2019;314–324. 10.1109/IPDPS.2019.00041 [DOI] [Google Scholar]
  42. Williams DD: A laboratory study of predator prey interactions of stoneflies and mayflies. Freshw Biol. 1987;17(3): 471–490. 10.1111/j.1365-2427.1987.tb01068.x [DOI] [Google Scholar]
  43. Zhou C, McCarthy SA, Durbin R: YaHS: Yet another Hi-C Scaffolding tool. Bioinformatics. Edited by C. Alkan, 2023;39(1):btac808. 10.1093/bioinformatics/btac808 [DOI] [PMC free article] [PubMed] [Google Scholar]
Wellcome Open Res. 2026 Apr 8. doi: 10.21956/wellcomeopenres.23612.r151496

Reviewer response for version 2

Jinjun Cao 1

This data note presents a high-quality, chromosome-level genome assembly for  Isoperla grammatica (Common Yellow Sally), a stonefly species found throughout Britain and Ireland. The paper includes detailed methods for sample collection, DNA/RNA extraction, sequencing, assembly, and quality assessment. Three specimens were collected from different locations in the UK in March 2019, with one used for the primary genome assembly.

The description of certain key steps in the Methods section is overly general, which compromises the reproducibility of the study. The rigor of some result descriptions needs to be improved.

1. For Mitochondrial genome assembly: The description of MitoHiFi in “Methods -> Genome assembly and curation” is too brief. Clarification should be added regarding the input data (e.g., whether they were extracted from preliminarily assembled contigs by Hifiasm or assembled de novo from raw HiFi reads), along with a brief explanation of the reference mitochondrial genome or parameter settings used for annotation.

2. For Hi-C data mapping: The manuscript states that “Hi-C reads were mapped to the primary contigs using bwa-mem2” but provides no mapping parameters. Basic alignment parameters (such as default settings or specific parameters employed) should be supplemented, as this is standard practice.

Are sufficient details of methods and materials provided to allow replication by others?

Partly

Is the rationale for creating the dataset(s) clearly described?

Yes

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Partly

Reviewer Expertise:

Phylogeny and Evolution of Plecoptera

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard, however I have significant reservations, as outlined above.

Wellcome Open Res. 2026 Mar 28. doi: 10.21956/wellcomeopenres.23612.r151233

Reviewer response for version 2

R Edward DeWalt 1

This version I believe is ready to be accepted.

A few improvements...

1. A citation with my name should be DeWalt not Dewalt. It is correct in the reference.

The sentence "It is considered a eurytherm (Graf et al., 2009) and is typically occurs..." should read "It is considered a eurytherm (Graf et al., 2009) and typically occurs...."

Otherwise, I am happy with the revision.

Are sufficient details of methods and materials provided to allow replication by others?

No

Is the rationale for creating the dataset(s) clearly described?

No

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Yes

Reviewer Expertise:

Plecoptera taxonomy and systematics.

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard.

Wellcome Open Res. 2023 Jul 3. doi: 10.21956/wellcomeopenres.21139.r59815

Reviewer response for version 1

R Edward DeWalt 1

The authors present a genome sequence for the stonefly Isoperla grammatica. There is little to no rationale provided for this work. What is the Darwin Tree of Life project? How is the dataset important? How will it used?

I assume that the identity of the specimens used is I. grammatica. Craig Macadam can certain identify this species, but he is not listed as the identifier nor is it said he confirmed the identification by McSwan. It would be good to state that this basic effort has been done to ensure the identity of the specimens use. I am disturbed by the total lack of voucher material. The presence of vouchers is a basic tenant of taxonomic science and without it repeatability is compromised! Not even photo vouchers are presented. The image of the adult provided is only an exemplar and matches no specimens sequenced in this work.

While this may be the only instance of high quality genome sequences for I. grammatic, other work using transcriptome and whole genome sequencing have recently been published (South et al. 2021a 1 & b 2 , Letch et al. 2021 3 ). These should be cited. One of the other reviewers asked for a reference for the distribution of the this species. DeWalt et al. 2023 could be used.

One other reviewer thought the biology section was too long. I don't agree.

Are sufficient details of methods and materials provided to allow replication by others?

No

Is the rationale for creating the dataset(s) clearly described?

No

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Yes

Reviewer Expertise:

Plecoptera taxonomy and systematics.

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard, however I have significant reservations, as outlined above.

References

  • 1. : A New Family of Stoneflies (Insecta: Plecoptera), Kathroperlidae, fam. n., with a Phylogenomic Analysis of the Paraperlinae (Plecoptera: Chloroperlidae). Insect Systematics and Diversity .2021;5(4) : 10.1093/isd/ixab014 10.1093/isd/ixab014 [DOI] [Google Scholar]
  • 2. : Combining molecular datasets with strongly heterogeneous taxon coverage enlightens the peculiar biogeographic history of stoneflies (Insecta: Plecoptera). Systematic Entomology .2021;46(4) : 10.1111/syen.12505 952-967 10.1111/syen.12505 [DOI] [Google Scholar]
  • 3. : Phylogenomics of the North American Plecoptera. Systematic Entomology .2021;46(1) : 10.1111/syen.12462 287-305 10.1111/syen.12462 [DOI] [Google Scholar]
Wellcome Open Res. 2025 Sep 24.
Tree of Life Team Sanger 1

Thank you for reviewing this data note. We have resubmitted a new version and have found your comments helpful. Responses to some of your points are below: 

1. Specimen identification. 

We are unable to supply a voucher photograph, although specimen identification was done using reliable keys. We did a BOLD barcode check on the DNA during assembly, and the results are available on the TOLQC page (also now linked in the article): 

https://tolqc.cog.sanger.ac.uk/darwin/insects/Isoperla_grammatica/

2. Other work using molecular data. 

We have expanded on this in the introduction. 

The Tree of Life team

Wellcome Open Res. 2023 Jun 22. doi: 10.21956/wellcomeopenres.21139.r59818

Reviewer response for version 1

Tilman Schell 1

Comments to the authors:

McSwan et al. present the de novo genome assembly of Isoperla grammatica. PacBio HiFi reads and Arima Hi-C data were utilized to achieve chromosome level scaffolds. The presented analyses show a high quality of the assembly.

The manuscript is well written and methods are described in a way to ensure reproducibility. Nevertheless, I do have some remarks that need to be clarified before indexing.

Major remarks:

Not sure if this applies more to the journal formatting but without line numbers pointing to certain locations in the text.

In the section “Genome sequence report” you state that the scaffolds represent chromosomes. Is there any previous work showing the karyotype of this species? If so, you would need to state this here. Otherwise you should be more careful by stating that the chromosome-level scaffolds are probably representing mostly complete chromosomes based on Hi-C data.

Additionally, using “(chromosome-scale) scaffolds” and “chromosomes” synonymous, is not applicable as long as you could e.g. proof correspondence of assembled sequence and actual molecules by staining experiments.

Is there genome size estimate of this species? It further underlines the quality of an assembly, if total assembly length and estimated genome size are similar.

It would be helpful to state the total HiFi Gbp and HiFi N50s of the different SMRT cells. I assume that there are multiple SMRT cells sequenced as you state “Pacific Biosciences HiFi circular consensus DNA sequencing libraries [...]” in the “Sequencing” section and there are two accession numbers in Table 1. Please clarify this a bit more in the main text.

The methods on mt genome assembly are very brief, please describe this in more detail. Did you used contigs or reads as input? Did you filter out contigs representing (parts of the) mt genome?

Mapping of Hi-C reads is not described at all. Please add this.

Minor remarks:

General:

Please label y-axis to be readable from the right.

Just to be sure: You created RNAseq data but you did not use it in this study, or?

What parameters were used in the applied tools? All were running with default settings? Please add this to the methods.

Particular:

Table 1: The QV value is presented here but it is not described how this value was computed. Please add this to the methods.

“Percentage of assembly mapped to chromosomes”: Probably you mean “anchored in scaffolds”.

How were the sex chromosomes (better: scaffolds representing sex chromosomes) identified? This is not described.

Figure 3: “GC-coverage plot”: As far as I am aware this is called “blob plot”.

Was there any filtering involved here? If yes, you should state this in the methods. If not it looks surprisingly clean, as you state a whole (?) specimen was used for DNA extraction.

Did you check the blob plot before scaffolding? In the worst case there are parts of contaminants scaffolded into the larger scaffolds and finally not visible anymore.

Figure 5: As the other reviewer already stated, it is not very constructive to show a contact map without showing the scaffold and contig boundaries.

Are the two scaffolds representing the sex chromosomes in the area with the reduced signal? Maybe this would be worth a note in the figure caption.

In the second paragraph of “Genome sequencing report” you stated uncertain regions to be in “chromosome 8, 9 and 11”. Please rephrase to “scaffold 8, 9 and 11”.

In section “Genome assembly” you state “Pretext” but this should be either “rapid curation pipeline” (more precise) or “PretextView” (as in Table 3).

Are sufficient details of methods and materials provided to allow replication by others?

No

Is the rationale for creating the dataset(s) clearly described?

Yes

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Yes

Reviewer Expertise:

Bioinformatics, de novo genome assembly and annotation

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard, however I have significant reservations, as outlined above.

Wellcome Open Res. 2023 Jun 13. doi: 10.21956/wellcomeopenres.21139.r59819

Reviewer response for version 1

Chen Zhiteng 1

The work presents a genome assembly from Isoperla grammatica, which is an important resource for the genomics of Plecoptera. However, the following comments should be addressed before acceptance for indexing.

  1. In first paragraph of Background, please cite a reference for the distribution.

  2. In Background, there is too much information on the biology. Please add information on genomics and genetics of I. grammatica, such as how many genomes of the family Perlodidae have been sequenced.

  3. In the results, there is no information on the mitochondrial genome. Please clarify it, and compare with the published mitogenomes of I. grammatica in GenBank.

  4. Figure 5, the boundary of the chromosomes is not clear, please check.

Are sufficient details of methods and materials provided to allow replication by others?

Yes

Is the rationale for creating the dataset(s) clearly described?

Yes

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Yes

Reviewer Expertise:

Taxonomy, genomics, phylogeny

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard, however I have significant reservations, as outlined above.

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Data Availability Statement

    European Nucleotide Archive: Isoperla grammatica (common yellow sally). Accession number PRJEB53729; https://identifiers.org/ena.embl/PRJEB53729.

    The genome sequence is released openly for reuse by the Wellcome Sanger Institute. The Isoperla grammatica genome sequencing initiative is part of the Darwin Tree of Life (DToL) project. All raw sequence data and the assembly have been deposited in INSDC databases. The genome will be annotated using available RNA-Seq data and presented through the Ensembl pipeline at the European Bioinformatics Institute. Raw data and assembly accession identifiers are reported in Table 1 and Table 2.

    Production code used in genome assembly at the WSI Tree of Life is available at https://github.com/sanger-tol. And 4 are

    Metadata for specimens, BOLD barcode results, spectra estimates, sequencing runs, contaminants and pre-curation assembly statistics are given at https://links.tol.sanger.ac.uk/species/552050.


    Articles from Wellcome Open Research are provided here courtesy of The Wellcome Trust

    RESOURCES