Skip to main content
Wellcome Open Research logoLink to Wellcome Open Research
. 2025 Sep 26;7:35. Originally published 2022 Feb 1. [Version 2] doi: 10.12688/wellcomeopenres.17577.2

The genome sequence of the Small Skipper, Thymelicus sylvestris (Poda, 1761)

Alex Hayward 1, Ryan Biscocho 1; Darwin Tree of Life Barcoding collective; Wellcome Sanger Institute Tree of Life programme; Wellcome Sanger Institute Scientific Operations: DNA Pipelines collective; Tree of Life Core Informatics collective; Darwin Tree of Life Consortiuma
PMCID: PMC12595296  PMID: 41209393

Version Changes

Revised. Amendments from Version 1

In version 2 of this data note we have added information about karyotyping of the Hesperiidae family to the Background section.  Information about genome annotation by Ensembl has been added.  We have used an additional table to clarify which specimens were used for different types of sequence data. We have replaced the chromosome map in Figure 5 with a labelled version.

Abstract

We present a genome assembly from an individual male Thymelicus sylvestris (the Small Skipper; Arthropoda; Insecta; Lepidoptera; Hesperiidae). The genome sequence is 471 megabases in span. Most of the assembly (99.97%) is scaffolded into 28 chromosomal pseudomolecules, including the Z sex chromosome. Gene annotation of this assembly on Ensembl identified 13,941 protein-coding genes. The mitochondrial genome was also assembled and is 17.1 kilobases in length.

Keywords: Thymelicus sylvestris, small skipper, genome sequence, chromosomal, Lepidoptera

Species taxonomy

Eukaryota; Metazoa; Ecdysozoa; Arthropoda; Hexapoda; Insecta; Pterygota; Neoptera; Endopterygota; Lepidoptera; Glossata; Ditrysia; Hesperioidea; Hesperiidae; Hesperiinae; Hesperiini; Thymelicus; Thymelicus sylvestris (Poda, 1761) (NCBI:txid272628).

Background

The Small Skipper ( Thymelicus sylvestris) is a butterfly within the skipper family Hesperiidae. The skippers are named for their characteristically quick, darting flight. The common name of T. sylvestris is a clear reference to its small size: the adult wingspan ranges from 27–34 mm ( Tomlinson & Still, 2002). However, it is not the smallest of the skippers, with four other British species being an equivalent size or smaller. Similar to other skippers, T. sylvestris has golden-orange wings with clear sex brands on males, but it can be distinguished by a lack of coloured patches on its wings and a dull brown or orange colouration to its antennae ( Tomlinson & Still, 2002).

Thymelicus sylvestris is widespread across the European continent with a habitat range paralleling that of other skipper species. This range encompasses the northernmost reaches of Morocco and Algeria all the way to the bordering regions between the Baltic states and Russia ( Tolman & Lewington, 2008). However, it is noticeably absent from northern Scandinavia, Corsica and Sardinia ( Tolman & Lewington, 2008). In the British Isles the small skipper is found across most of Wales and England with recent trends showing a northward expansion in range, beyond the England-Scotland border. Recently, T. sylvestris individuals have also been observed in Ireland, where they had not been reported previously ( Harding & Jacob, 2013). Thymelicus sylvestris populations appear stable and it is listed as a species of least concern by the IUCN ( van Swaay et al., 2010).

The small skipper is a habitat generalist ( Louy et al., 2007) and can be found in open areas with long grass, such as rough grasslands and roadside verges ( Tomlinson & Still, 2002). It is most associated with Yorkshire fog ( Holcus lanatus), its main food plant, on which it often basks and lays eggs from June to July. Females are known to be meticulous with their choice of oviposition sites, spending up to 15 minutes inspecting potential host plants prior to laying eggs ( Tolman & Lewington, 2008). After approximately a month, eggs hatch into caterpillars which develop through 5 instar stages. Come winter, caterpillars spin cocoons within which they undergo diapause. The caterpillars re-emerge in spring, constructing a ‘leaf tube’ by joining together the ends of a leaf, where they live and feed, moving to new leaves as necessary. Small skipper caterpillars usually pupate by June, with adult butterflies emerging in July, to spend their remaining days in tall grassland until the summer’s end in September.

Most skippers retain the ancestral lepidopteran haploid number n = 31, which is also inferred to be ancestral for Hesperiidae overall ( Lukhtanov, 2014). Within Hesperiinae, the tribe Thymelicini most commonly shows n = 29, with a report of n = 27 for Thymelicus sylvestris based on cytology ( Lukhtanov, 2014). We present the first chromosome-level genome sequence for Thymelicus sylvestris, the Small Skipper. The karyotype reported here is n = 28.

The assembly was generated as part of the Darwin Tree of Life Project, which aims to generate high-quality reference genomes for all named eukaryotic species in Britain and Ireland to support research, conservation, and the sustainable use of biodiversity.

Genome sequence report

The genome was sequenced from a single male T. sylvestris collected from Ruan Minor, Cornwall, UK (latitude 49.9942295, longitude –5.1974720) ( Figure 1). Two other specimens were used for Hi-C scaffolding and RNA sequencing. Table 1 summarises the specimen and sequencing details. PacBio sequencing generated 20.16 Gb (gigabases) from 1.99 million reads, which were used to assemble the genome. Based on the estimated genome size, the sequencing data provided approximately 40× coverage. Hi-C sequencing produced 122.17 Gb from 809.08 million reads, which were used to scaffold the assembly. RNA sequencing data were also generated and are available in public sequence repositories.

Figure 1. Fore and hind wings of the Thymelicus sylvestris specimen from which the genome was sequenced.

Figure 1.

Dorsal (left) and ventral (right) surface view of wings from specimen EN_RM_1378 (ilThySylv1) from Ruan Minor, Cornwall, UK, used to generate Pacific Biosciences and 10X genomics data.

Table 1. Specimen and sequencing data for BioProject PRJEB45673.

Platform PacBio HiFi Hi-C RNA-seq
ToLID ilThySylv1 ilThySylv3 ilThySylv2
Specimen ID SAN0000837 SAN0000839 SAN0000838
BioSample (source individual) SAMEA7523279 SAMEA7523281 SAMEA7523280
BioSample (tissue) SAMEA7523375 SAMEA7523377 SAMEA7523376
Tissue whole organism whole organism whole organism
Instrument Sequel II Illumina NovaSeq 6000 Illumina HiSeq 4000
Run accessions ERR6608659 ERR6363315 ERR6363314
Read count total 1.99 million 809.08 million 38.45 million
Base count total 20.16 Gb 122.17 Gb 5.81 Gb

Manual assembly curation corrected 9 missing/misjoins and removed 3 haplotypic duplications, reducing the assembly size by 0.06% and the scaffold number by 20.00%.

The final assembly has a total length of 471 Mb in 32 sequence scaffolds with a scaffold N50 of 16.61 Mb ( Table 2). Of the assembly sequence, 99.97% was assigned to 28 chromosomal-level scaffolds, representing 27 autosomes (numbered by sequence length), and the Z sex chromosome ( Figure 2Figure 5; Table 3). The assembly curator noted that this differs from the published karyotype for this species ( Lukhtanov, 2014), but the available data did not support making additional joins. One chromosome is very small scaffold, at about 8 Mb, which could plausibly have been overlooked in earlier cytological work; alternatively, differences between cytological and genome-assembly interpretations could explain the discrepancy.

Figure 2. Genome assembly of Thymelicus sylvestris, ilThySylv1.1: metrics.

Figure 2.

BlobToolKit Snailplot shows N50 metrics and BUSCO gene completeness. The main plot is divided into 1,000 size-ordered bins around the circumference with each bin representing 0.1% of the 470,727,450 bp assembly. The distribution of chromosome lengths is shown in dark grey with the plot radius scaled to the longest chromosome present in the assembly (37,236,842 bp, shown in red). Orange and pale-orange arcs show the N50 and N90 chromosome lengths (17,253,319 and 11,732,099 bp), respectively. The pale grey spiral shows the cumulative chromosome count on a log scale with white scale lines showing successive orders of magnitude. The blue and pale-blue area around the outside of the plot shows the distribution of GC, AT and N percentages in the same bins as the inner plot. A summary of complete, fragmented, duplicated and missing BUSCO genes in the lepidoptera_odb10 set is shown in the top right. An interactive version of this figure is available at https://blobtoolkit.genomehubs.org/view/ilThySylv1.1/dataset/CAJVQR01/snail.

Figure 5. Genome assembly of Thymelicus sylvestris, ilThySylv1.1: Hi-C contact map.

Figure 5.

Hi-C contact map of the ilThySylv1.1 assembly, visualised in PretextView and PretextSnapshot. Chromosomes are arranged in size order from left to right and top to bottom.

Table 2. Genome data for Thymelicus sylvestris, ilThySylv1.1.

Assembly name ilThySylv1.1
Assembly accession GCA_911387775.1
Alternate haplotype accession GCA_911387695.1
Assembly level chromosome
Span (Mb) 470.71
Number of chromosomes 28
Number of contigs 46
Contig N50 16.61 Mb
Number of scaffolds 32
Scaffold N50 17.25 Mb
Sex chromosomes Z
Organelles Mitochondrion: 17.11 kb

Figure 3. Genome assembly of Thymelicus sylvestris, ilThySylv1.1: GC coverage.

Figure 3.

BlobToolKit GC-coverage plot. Scaffolds are coloured by phylum. Circles are sized in proportion to scaffold length. Histograms show the distribution of scaffold length sum along each axis. An interactive version of this figure is available at https://blobtoolkit.genomehubs.org/view/ilThySylv1.1/dataset/CAJVQR01/blob.

Figure 4. Genome assembly of Thymelicus sylvestris, ilThySylv1.1: cumulative sequence.

Figure 4.

BlobToolKit cumulative sequence plot. The grey line shows cumulative length for all scaffolds. Coloured lines show cumulative lengths of scaffolds assigned to each phylum using the buscogenes taxrule. An interactive version of this figure is available at https://blobtoolkit.genomehubs.org/view/ilThySylv1.1/dataset/CAJVQR01/cumulative.

Table 3. Chromosomal pseudomolecules in the genome assembly of Thymelicus sylvestris, ilThySylv1.1.

INSDC accession Chromosome Size (Mb) GC%
OU426886.1 1 20.99 36.0
OU426887.1 2 20.74 36.9
OU426888.1 3 20.31 36.3
OU426889.1 4 19.76 36.3
OU426890.1 5 19.60 36.5
OU426891.1 6 19.09 35.7
OU426892.1 7 18.50 35.9
OU426893.1 8 18.16 35.9
OU426894.1 9 17.98 35.8
OU426895.1 10 17.90 36.0
OU426896.1 11 17.25 35.6
OU426897.1 12 17.04 35.8
OU426898.1 13 16.99 36.2
OU426899.1 14 16.61 36.0
OU426900.1 15 16.11 36.2
OU426901.1 16 16.10 35.9
OU426902.1 17 15.35 35.9
OU426903.1 18 15.05 36.0
OU426904.1 19 14.87 36.2
OU426905.1 20 14.87 36.6
OU426906.1 21 14.38 36.2
OU426907.1 22 13.89 36.3
OU426908.1 23 11.73 36.0
OU426909.1 24 11.38 36.7
OU426910.1 25 10.65 36.1
OU426911.1 26 9.84 36.9
OU426912.1 27 8.15 36.4
OU426885.1 Z 37.24 35.7
OU426913.1 MT 0.02 16.7

The assembly has a BUSCO ( Simão et al., 2015) v5.1.2 completeness of 98.5% (single 98.1%, duplicated 0.5%) using the lepidoptera_odb10 reference set. While not fully phased, the assembly deposited is of one haplotype. Contigs corresponding to the second haplotype have also been deposited.

Genome annotation report

The Thymelicus sylvestris genome assembly (GCA_911387775.1) was annotated by Ensembl at the European Bioinformatics Institute (EBI). This annotation includes 38,357 transcribed mRNAs from 13,941 protein-coding and 6 462 non-coding genes. The average transcript length is 18,220.25 bp, with an average of 1.88 coding transcripts per gene and 7.23 exons per transcript. For further information about the annotation, please refer to the annotation page on Ensembl.

Methods

Specimen acquisition and nucleic acid extraction

Three male T. sylvestris (ilThySylv1, ilThySylv2 and ilThySylv3) specimens were collected from Ruan Minor, Cornwall, UK (latitude 49.9942295, longitude –5.1974720) using a net by Alex Hayward in May 2019. The samples were identified by the same person, and then snap-frozen on dry ice.

DNA was extracted from the whole organism of ilThySylv1 (specimen ID: EN_RM_1378) at the Wellcome Sanger Institute (WSI) Scientific Operations core from the whole organism using the Qiagen MagAttract HMW DNA kit, according to the manufacturer’s instructions. RNA from whole organism tissue of ilThySylv2 was extracted in the Tree of Life Laboratory at the WSI using TRIzol, according to the manufacturer’s instructions. RNA was then eluted in 50 μl RNAse-free water and its concentration assessed using a Nanodrop spectrophotometer and Qubit Fluorometer using the Qubit RNA Broad-Range (BR) Assay kit. Analysis of the integrity of the RNA was done using the Agilent RNA 6000 Pico Kit and Eukaryotic Total RNA assay.

Sequencing

Pacific Biosciences HiFi circular consensus and 10X Genomics Chromium read cloud sequencing libraries were constructed according to the manufacturers’ instructions. Poly(A) RNA-Seq libraries were constructed using the NEB Ultra II RNA Library Prep kit. Sequencing was performed by the Scientific Operations core at the Wellcome Sanger Institute on Pacific Biosciences SEQUEL II (HiFi), Illumina HiSeq X (10X) and Illumina HiSeq 4000 (RNA-Seq) instruments. Hi-C data were generated from head tissue of ilThySylv3 in the Tree of Life Laboratory using the Arima Hi-C+ kit and sequenced on an Illumina NovaSeq 6000 instrument.

Genome assembly

Assembly was carried out with Hifiasm ( Cheng et al., 2021). Haplotypic duplication was identified and removed with purge_dups ( Guan et al., 2020). One round of polishing was performed by aligning 10X Genomics read data to the assembly with longranger align, calling variants with freebayes ( Garrison & Marth, 2012). The assembly was then scaffolded with Hi-C data ( Rao et al., 2014) using SALSA2 ( Ghurye et al., 2019). The assembly was checked for contamination and corrected using the gEVAL system ( Chow et al., 2016) as described previously ( Howe et al., 2021). Manual curation was performed using gEVAL, HiGlass ( Kerpedjiev et al., 2018) and Pretext. The mitochondrial genome was assembled using MitoHiFi ( Uliano-Silva et al., 2021). The genome was analysed and BUSCO scores generated within the BlobToolKit environment ( Challis et al., 2020). Table 3 contains a list of all software tool versions used, where appropriate.

Table 3. Software tools used.

Data availability

European Nucleotide Archive: Thymelicus sylvestris (small skipper). Accession number PRJEB45673; https://identifiers.org/ena.embl/PRJEB45673. The genome sequence is released openly for reuse. The Thymelicus sylvestris genome sequencing initiative is part of the Darwin Tree of Life Project (PRJEB40665), the Sanger Institute Tree of Life Programme (PRJEB43745) and Project Psyche (PRJEB71705).

All raw sequence data and the assembly have been deposited in INSDC databases. Raw data and assembly accession identifiers are reported in Table 1 and Table 2.

Production code used in genome assembly at the WSI Tree of Life is available at https://github.com/sanger-tol.

Funding Statement

This work was supported by Wellcome through core funding to the Wellcome Sanger Institute (206194) and the Darwin Tree of Life Discretionary Award (218328). AH is supported by a Biotechnology and Biological Sciences Research Council (BBSRC) David Phillips Fellowship (BB/N020146/1).

The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

[version 2; peer review: 4 approved]

Author information

Members of the Darwin Tree of Life Barcoding collective are listed here: https://doi.org/10.5281/zenodo.5744972.

Members of the Wellcome Sanger Institute Tree of Life programme are listed here: https://doi.org/10.5281/zenodo.5744840.

Members of Wellcome Sanger Institute Scientific Operations: DNA Pipelines collective are listed here: https://doi.org/10.5281/zenodo.5746904.

Members of the Tree of Life Core Informatics collective are listed here: https://doi.org10.5281/zenodo.5743293.

Members of the Darwin Tree of Life Consortium are listed here: https://doi.org/10.5281/zenodo.5638618.

References

  1. Challis R, Richards E, Rajan J, et al. : BlobToolKit - interactive quality assessment of genome assemblies. G3 (Bethesda). 2020;10(4):1361–74. 10.1534/g3.119.400908 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Cheng H, Concepcion GT, Feng X, et al. : Haplotype-Resolved de Novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021;18(2):170–75. 10.1038/s41592-020-01056-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Chow W, Brugger K, Caccamo M, et al. : gEVAL - a web-based browser for evaluating genome assemblies. Bioinformatics. 2016;32(16):2508–10. 10.1093/bioinformatics/btw159 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Garrison E, Marth G: Haplotype-Based variant detection from short-read sequencing.arXiv: 1207.3907.2012. Reference Source [Google Scholar]
  5. Ghurye J, Rhie A, Walenz BP, et al. : Integrating Hi-C links with assembly graphs for chromosome-scale assembly. PLoS Comput Biol. 2019;15(8): e1007273. 10.1371/journal.pcbi.1007273 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Guan D, McCarthy SA, Wood J, et al. : Identifying and removing haplotypic duplication in primary genome assemblies. Bioinformatics. 2020;36(9):2896–98. 10.1093/bioinformatics/btaa025 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Harding J, Jacob M: Addition of Small Skipper butterfly (Thymelicus Sylvestris) to the Irish List and notes on the essex skipper (Thymelicus Lineola) (Lepidoptera: Hesperiidae). The Irish Naturalists’ Journal. 2013;32(2):142–44. Reference Source [Google Scholar]
  8. Howe K, Chow W, Collins J, et al. : Significantly improving the quality of genome assemblies through curation. GigaScience. 2021;10(1): giaa153. 10.1093/gigascience/giaa153 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Kerpedjiev P, Abdennur N, Lekschas F, et al. : HiGlass: web-based visual exploration and analysis of genome interaction maps. Genome Biol. 2018;19(1):125. 10.1186/s13059-018-1486-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Louy D, Habel JC, Schmitt T, et al. : Strongly diverging population genetic patterns of three skipper species: the role of habitat fragmentation and dispersal ability. Conserv Genet. 2007;8(3):671–81. 10.1007/s10592-006-9213-y [DOI] [Google Scholar]
  11. Lukhtanov VA: Chromosome number evolution in skippers (Lepidoptera, Hesperiidae). Comp Cytogenet. 2014;8(4):275–291. 10.3897/CompCytogen.v8i4.8789 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Rao SS, Huntley MH, Durand NC, et al. : A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell. 2014;159(7):1665–80. 10.1016/j.cell.2014.11.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Simão FA, Waterhouse RM, Ioannidis P, et al. : BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics. 2015;31(19):3210–12. 10.1093/bioinformatics/btv351 [DOI] [PubMed] [Google Scholar]
  14. Tolman T, Lewington R: The most complete guide to the butterflies of Britain and Europe.Harper Collins, London,2008. [Google Scholar]
  15. Tomlinson D, Still R: Britain’s butterflies.WILDGuides,2002. Reference Source [Google Scholar]
  16. Uliano-Silva M, Nunes JGF, Krasheninnikova K, et al. : marcelauliano/MitoHiFi: mitohifi_v2.0.2021. 10.5281/zenodo.5205678 [DOI] [Google Scholar]
  17. van Swaay C, Wynhoff I, Verovnik R, et al. : IUCN red list of threatened species: Thymelicus Sylvestris. IUCN Red List of Threatened Species. 2010. Reference Source [Google Scholar]
Wellcome Open Res. 2026 Jan 6. doi: 10.21956/wellcomeopenres.27568.r134447

Reviewer response for version 2

Xueyan Li 1

This paper “The genome sequence of the small skipper, Thymelicus sylvestris (Poda, 1761)” have now been improved well. I think it is ok now.

Are sufficient details of methods and materials provided to allow replication by others?

Partly

Is the rationale for creating the dataset(s) clearly described?

Yes

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Yes

Reviewer Expertise:

Insect genome and evolution

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard.

Wellcome Open Res. 2025 Nov 27. doi: 10.21956/wellcomeopenres.27568.r138439

Reviewer response for version 2

Vlad Dinca 1

In this data note, the authors have successfully generated a chromosome-level genome assembly of a male Thymelicus sylvestris (Lepidoptera: Hesperiidae). The mitochondrial genome has been assembled as well.

Although I am not an expert in several of the methodological approaches used, as far as I can tell, the manuscript is technically sound and uses methodologies and pipelines that are fairly well-established through use by the Wellcome Sanger Institute.

The genome appears to be of high quality and scaffolded into 28 chromosomal pseudomolecules (including the Z sex chromosome).

A few additional comments are included below.

I would say it is absent from most of Scandinavia, as it is basically present in Denmark.

Holcus lanatus does indeed seem to be one of the preferred host plants, but several others have been reported across Europe. Thus, it may be good to list a few other species on which larvae have been reported (see for example Clarke 2024 -

https://doi.org/10.1002/ece3.10834).

The species can be on wing earlier, at least during May as well. Please update this. The specimen analysed in this note is later mentioned to have been collected in May 2019.

Please also update the mention that the adults emerge in July. Across Europe, depending on location, the flight time can start even at the end of April and will usually end in August. The “core” of the flight time is usually between May-July.

It may also be useful adding 1-2 sentences about what is currently known about the genetic structure of this species across Europe. See for example Hinojosa et al. (2019 -

https://doi.org/10.1111/mec.15153).

Are sufficient details of methods and materials provided to allow replication by others?

Yes

Is the rationale for creating the dataset(s) clearly described?

Yes

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Yes

Reviewer Expertise:

Lepidoptera taxonomy, phylogeography, conservation.

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard.

Wellcome Open Res. 2025 Nov 8. doi: 10.21956/wellcomeopenres.27568.r136729

Reviewer response for version 2

Xiaoling Fan 1

This manuscript presents a chromosomal-level genome assembly for the Small Skipper,  Thymelicus sylvestris, achieved through the integration of PacBio HiFi, 10X, RNA-seq, and Hi-C data from fresh specimens. The assembly's exceptional continuity (scaffold N50 of 17.25 Mb) and completeness (98.5% BUSCO score) establish it as a foundational and highly reliable resource for future research in skipper butterfly genomics.

Are sufficient details of methods and materials provided to allow replication by others?

Yes

Is the rationale for creating the dataset(s) clearly described?

Yes

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Yes

Reviewer Expertise:

systematics

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard.

Wellcome Open Res. 2022 Jun 9. doi: 10.21956/wellcomeopenres.19436.r50767

Reviewer response for version 1

Xueyan Li 1

This paper “The genome sequence of the small skipper, Thymelicus sylvestris (Poda, 1761)“ assembled the chromosomal-level genome of Thymelicus sylvestris. Due to the HiFi, 10X Genomics, Hi-C Illumina and RNA-seq, the assembly quality is relatively high. The high-quality reference genomes of more butterfly species will be helpful to investigate the evolution of butterfly diversity.

In general, this paper is well written with a concise description. Nevertheless, the authors didn’t provide genome annotation (repeat and genes), which, if unavailable, will limit the use of the data, especially for non-bioinformatic specialists. So, please provide genome annotation information.

Several revisions are needed for the data.

  • First, how many chromosomes for this species, please describe it and cite the related literature.

  • Second, how the authors identified the Z sex chromosome?

Minor revisions

  • Background: “The skippers are named for their characteristic quick, darting flight.“ to “The skippers are named for their characteristically quick, darting flight.”.

Are sufficient details of methods and materials provided to allow replication by others?

Partly

Is the rationale for creating the dataset(s) clearly described?

Yes

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Yes

Reviewer Expertise:

Insect genome and evolution

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard, however I have significant reservations, as outlined above.

Wellcome Open Res. 2025 Sep 23.
Tree of Life Team Sanger 1

Thank you for reviewing this data note. We have made changes to the article in Version 2 and our responses to your comments are given below.

  • “Nevertheless, the authors didn’t provide genome annotation (repeat and genes), which, if unavailable, will limit the use of the data, especially for non-bioinformatic specialists. So, please provide genome annotation information.”

Response: Genome annotation is not part of our assembly pipeline but is done externally by Ensembl at the European Bioinformatics Institute. We have added information and a link to the Ensembl page for the genome annotation, which has been supplied since version 1 of the data note was published.

  • First, how many chromosomes for this species, please describe it and cite the related literature.

Response: Thank you for this suggestion. We have added information to the background section about the related literature. Previously cytology done in 1960 found a chromosome count of n = 27. Although this was the expected chromosome count for our assembly, we found an additional chromosome, and the data did not support further joins being made.

  • Second, how the authors identified the Z sex chromosome?

Response: The Z chromosome is identified based on synteny to related species and the presence of conserved BUSCO genes.

Wellcome Open Res. 2022 Feb 21. doi: 10.21956/wellcomeopenres.19436.r48488

Reviewer response for version 1

Francois Olivier Hebert 1,2

General comments:

This short and well written manuscript presents the genome sequence of a male Thymelicus sylvestris (small skipper) sample in the context of the Darwin Tree of Life project. The genome assembly in this work was generated with a combination of PacBio HiFi reads and additional 10X sequencing from a single male individual, and RNAseq and HiC sequencing for genome improvement. They report a final genome sequence of 471 Mb, which is 98.5% complete according to its BUSCO score. Although I am not an expert per say in entomology or specifically in skipper biology, I found this piece of work interesting, written in a straightforward manner and in a simple form. This work is totally worth publishing because it provides a very high-quality reference genome sequence for a species that was still lacking from public databases. It allows the research community to expand its general knowledge on butterfly genomics and it also provides additional data that can be used in the future to perform comparative analyses and better understand insect genome evolution on a broader sense. The authors did a good job at generating a complete and highly contiguous/complete genome sequence.

Although this manuscript is very clear and the assembly of high quality (props to the team for the cleaness of this assembly), I would have liked to learn a little bit more about the reason(s) behind this genome sequence, i.e., the “what we know”, “what we don’t know”, and what it is that this genome sequence potentially brings to the table to better understand what we don’t know about this species or its related taxa. Just to put the work into context. Maybe on the form of just a sentence on what the ultimate goal of the Darwin Tree of Life project is? And maybe a little bit more details in the methods in terms of command lines and scripts used to run the programs (which would augment the repeatability of this nice work). But apart from that, great work! Down there are some comments that could increase the quality of the paper if addressed by the authors:

Specific comments:

  • Suggestion: just a tiny minor detail, but maybe specify in the abstract that it’s 27 chromosomes + Z chromosome, for a total of 28 chromosomes? It is well mentioned in the main text, but it could be phrased differently in the abstract so that it is does not create any ambiguity when just reading the abstract by itself.

  • Impressive quality in terms of BUSCO score: little duplicated BUSCO genes and 98% single is very good, especially in animal species with (typically) large quantities of repeated elements.

  • On that note, is there a report somewhere of the percentage of repeats in this assembly? Or maybe this information is not relevant in the context of the DToL project, but I would have liked to have a sense of how repeated that genome is. I presume that there is no gene annotation report so far? If yes, this would be valuable to provide details on the numbers of genes, gene densities on the chromosomes, and other relevant functional annotation metrics.

  • Written in the “general comments” section, but again here: in the introduction, I think it would be interesting for the readership to learn a little bit more about the reasons behind this work. Why produce a genome sequence of this species in particular? How is it part of a larger initiative with a particular goal, for instance here, sequencing numerous species with a super large comparative genomics perspective in mind (i.e., understanding the evolutionary genomics / mapping the genome of all animals living in a common geographic location)? In other words, a brief rationale justifying the dataset.

  • In the section Genome sequence report, first paragraph, it says that the genome was sequenced from a single individual male, but later in the Methods section, the authors report the capture of three samples. In Table 1, it looks like these three samples were used to generate different libraries of different sequencing technologies, all involved in the genome assembly method. I was thinking that the authors could specify this detail in the first paragraph of the “Genome sequence report” section when they mention the different technologies used in their study, i.e., which technology was used to sequence which sample (and maybe the reason why three samples?)? Perhaps I’m missing something here in terms of assembly strategy, but in the end, if 3 different samples were used, the final genome assembly is not entirely from a single individual… or is it? Why not use these 3 technologies on the same sample? I know that extracting enough DNA material in a single individual insect can sometimes be quite difficult (low yields), so maybe this is the reason why the authors chose to use three samples? Just a short clarification on that note would elevate the quality of the paper… to my humble opinion.

  • I am by no means an expert in Lepidoptera, although I’ve assembled a lepidopteran genome in the past, I am wondering, out of curiosity, how many chromosomes are expected in this species specifically? Is it 28 or 31? I know chromosome numbers are quite variable in Hesperidae, but isn’t it supposed to be 31 usually (or ancestrally?)? Maybe the authors could briefly comment on that topic in the manuscript, i.e., if this number of chromosomes is expected or not? If there are 98% complete single BUSCO group, we can presume that no chromosomes are missing though, but maybe there are 28 chromosomes assembled here (out of 32 scaffolds?), but there should be 30 or 31 and the rest is in the “unanchored” contigs and remaining scaffolds? Could the authors, please, clarify this?

  • Data accessibility: I respect so much the fact that everything from this manuscript seems to be transparent and open source. Maybe I missed something in the data provided with this work, but is there a mention somewhere of the scripts or command lines used by the authors to produce this assembly? Or maybe all the programs were used with default values (which is sometimes recommended anyways when reading these tools’ manual)? The protocols are appropriate, and the work is technically sound, but it lacks a little bit of method details to be repeatable.

Are sufficient details of methods and materials provided to allow replication by others?

Partly

Is the rationale for creating the dataset(s) clearly described?

Partly

Are the datasets clearly presented in a useable and accessible format?

Yes

Are the protocols appropriate and is the work technically sound?

Yes

Reviewer Expertise:

Functional genomics, genome assembly, RNAseq, physiology

I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard.

Wellcome Open Res. 2025 Sep 23.
Tree of Life Team Sanger 1

Thank you for reviewing this data note. We have made the following changes in Version 2, which we hope will provide more clarity.

  • We have clarified the chromosome count in the abstract as you recommend. The chromosome map in Figure 3 is now labelled, which we hope will improve the data note.

  • As the genome has now been annotated, we have provided a summary of the annotated and a link to the information on Ensembl.

  • We have added a table showing the specimen and sequencing information. Three different individuals were used for PacBio HiFi, Illumina Hi-C and RNA-seq. The genome is that of only one individual, the one used for PacBio sequencing.

  • We have added a paragraph on the expected chromosome count for the family of skippers, and we compare the chromosome number recovered here to the expected count. There are no additional chromosomes in the unassembled scaffolds, as the sequence is almost completely assembled.

  • The main rationale for sequencing this species is that it is on the Darwin Tree of Life list, and we have

  • We have added a link to the software repository for Tree of Life assemblies in the Data availability section.

The Tree of Life team

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Data Availability Statement

    European Nucleotide Archive: Thymelicus sylvestris (small skipper). Accession number PRJEB45673; https://identifiers.org/ena.embl/PRJEB45673. The genome sequence is released openly for reuse. The Thymelicus sylvestris genome sequencing initiative is part of the Darwin Tree of Life Project (PRJEB40665), the Sanger Institute Tree of Life Programme (PRJEB43745) and Project Psyche (PRJEB71705).

    All raw sequence data and the assembly have been deposited in INSDC databases. Raw data and assembly accession identifiers are reported in Table 1 and Table 2.

    Production code used in genome assembly at the WSI Tree of Life is available at https://github.com/sanger-tol.


    Articles from Wellcome Open Research are provided here courtesy of The Wellcome Trust

    RESOURCES