Skip to main content
Scientific Data logoLink to Scientific Data
. 2025 Oct 20;12:1656. doi: 10.1038/s41597-025-05945-2

A high-quality chromosome-level genome assembly and annotation of Anqing Six-end-white pig (Sus scrofa)

Wei Zhang 1,2,3,4, Mei Zhou 1,2,3,4, Linqing Liu 1,2,3,4, Miao Lu 1,2,3, Shiguang Su 1,2,3,4,, Chonglong Wang 1,2,3,4,
PMCID: PMC12537974  PMID: 41115900

Abstract

The Anqing Six-end-white pig, a precious Chinese autochthonous breed, is mainly distributed in the Anhui Province, China. It is well documented for its excellent meat quality and crude-feed tolerance. However, the current lack of high-quality Anqing Six-end-white pig genome assembly poses a significant limitation for elucidating its germplasm characteristics and implementing genomic-level conservation. In this study, we assembled a high-quality chromosome-level genome; the genome size was 2.66 Gb with contig N50 = 90.48 Mb and scaffold N50 = 143.10 Mb. There were 23 gaps in the final assembled genome. A total of 1.16 Gb repeat sequences were identified and accounting for 43.52% of the genome. A total of 20,809 protein-coding genes were identified, and 99.18% of these genes were annotated. The predicted non-coding genes included 848 miRNAs, 4544 tRNA, 253 rRNA, and 2156 snRNAs in the genome. Genome completeness was 98.67% using BUSCO and 99.81% using Compleasm. This chromosome-level assembly provides a robust scientific foundation for the conservation, breeding, and exploration of genetic traits in the Anqing Six-end-white pig.

Subject terms: Genome, Evolution

Background & Summary

Pig (Sus scrofa), initially domesticated from wild boars independently in the Near East and China approximately 10,000 years ago14, have evolved into indispensable contributors to anthropogenic systems through their diverse roles. Nutritionally, they supply essential proteins, vitamins, and trace minerals critical for human dietary requirements5. Economically, industrialized swine husbandry drives agricultural economies, thus generating substantial revenue streams and employment opportunities5. Ecologically, swine has an exceptional capacity for sustainable nutrient cycling5. In translational medicine, the anatomical and physiological homology of pigs and humans has established them as pivotal preclinical models6. Culturally, pig hold profound symbolic significance in Eastern traditions, especially in Chinese cultural practices5. European and Asian pig exhibit different phenotypic and genomic characteristics7. The current pig genome, Sscrofa 11.1, was assembled using the Duroc Jersey pig, which was a most continuous assemblies and provide a pivotal role in mining the germplasm-related genes, thereby serving as an indispensable foundation for strategic genetic improvement initiatives in swine populations8. However, reliance on the pig genome Sscrofa 11.1 for evolutionary and population genetic analyses of Asian pig populations may result in the incomplete detection of Asian variations, hence, compromising the comprehensive characterization of their germplasm characteristics.

Anqing Six-end-white pig (Fig. 1), a representative native pig breed with a dual-purpose meat-lard type, is primarily distributed in surrounding mountainous areas including Taihu, Wangjiang, and Susong counties in China. In 2007, there were nearly 600 Anqing Six-white pigs. Recent data from 2019 show that the population has reached 4,617. Two national-level conservation farms were established to protect Anqing Six-white pigs. This breed is deeply loved by the local people, not only for providing meat protein and increasing farmers’ income but also for their position in local culture and history. In our prior investigations of Anqing Six-end-white pig, several candidate genes were identified that play important roles in regulating meat quality traits, lipid metabolism, and fat deposition912. However, the genetic basis of the characteristics in Anqing Six-end-white pig, has not yet been fully elucidated. The high-quality reference genome of Chinese indigenous pig breeds is a powerful tool to elucidate the characteristics of native breeds. Several high-quality pig genomes have been assembled, including Huai pig13, Chenghua pig14, Bamei pig15, and others16,17. However, a high-quality genome assembly of Anqing Six-end-white pig is still lacking, which greatly limits the elucidate of their germplasm characteristics and protection at the genomic level.

Fig. 1.

Fig. 1

The Anqing Six-end-white pig.

In this study, we assembled the first chromosome-level Anqing Six-end-white pig genome by combining short reads, PacBio HiFi (high fidelity) reads, and Hi-C (High-throughput chromosome conformation capture) sequencing data. The genome size and heterozygosity of Anqing Six-end-white pig were estimated to be 2.4 Gb and 0.61% according to 17-mer analysis of 129.62 Gb (57x) short reads. After de novo assembly with 229.34 Gb (95x) HiFi reads, the genome size was 2.69 Gb, composed of 62 contigs, and the contig N50 was 90.48 Mb with a 42.64% GC content. The final assembly genome (2.66 Gb) was anchored to 20 chromosomes (18 autosomes plus one X and Y), with scaffold N50 = 143.10 Mb through Hi-C assisted scaffolding. There were 23 gaps in the final assembled genome. The Anqing Six-end-white pig assembly captured 38 telomeres and 20 centromeres. Repeat sequences in the Anqing Six-end-white pig genome were annotated, and a 1.16 Gb repeat sequences (approximately 43.52% of the assembled genome) were identified. A total of 20,809 protein-coding genes were identified in our assembly, which harbored 36,142 transcripts with an average of 9.48 exons per gene. Non-coding RNAs in the assembly were annotated, and 848 miRNA, 4544 tRNA, 253 rRNA, and 2156 snRNA were identified. This study established a robust scientific foundation for the conservation, selective breeding, and exploration of superior genetic traits in the Anqing Six-end-white pig, offering critical genomic insights to safeguard germplasm resources and enhance agricultural sustainability.

Methods

Sample collection

A one-year-old male Anqing Six-end-white pig from the national Anqing Six-end-white pig conservation farm in Anqing city, Anhui province, China (30°19′N, 116°33′E) was used for genome sequencing (short reads, PacBio HiFi, and Hi-C) and assembly. Forty-two tissues of three one-year-old male Anqing Six-end-white pig were collected and rapidly frozen in liquid nitrogen and stored at −80°C for RNA sequencing, including the heart, liver, spleen, lung, kidney, stomach, longissimus dorsi muscle, adipose, duodenum, jejunum, ileum, cecum, colon and rectum. The pigs used in this study were healthy, and no genetic defects were observed in it or its parents. This study was conducted in accordance with and was approved by the Animal Care Committee of the Anhui Academy of Agricultural Sciences (Hefei, People’s Republic of China; Approval No. AAAS2023-4).

Sequencing

Genomic DNA was extracted from the blood sample using a standard phenol–chloroform method and assessed using a 0.5% agarose gel and Nanodrop spectrophotometer (Thermo Fisher Scientific, Waltham, MA, USA).

For short read sequencing, the DNA were processed with fragmented, size selected, end-repaired, A-tailed, ligated to paired-end adaptors, PCR amplificated, and sequenced on the BGISEQ DNBseq-T7 sequencing platform at OneMore Technology Co., Ltd. (Wuhan, China). In total, 141.65 Gb 150 bp paired-end reads were generated. Given that the original sequencing data may contain adapter sequences, low-quality bases, and undetected bases, Fastp18 (v0.23.2) software was used to filter the above information and 129.62 Gb data (Table 1) were retained to estimate the genome size of Anqing Six-end-white pig.

Table 1.

Sequencing data used for the Anqing Six-end-white pig genome assembly, which include the sequencing types, sequencing platform, sequencing tissue, average read length, total bases and sequencing coverage.

Data Sequencer Tissue Average read length(bp) Total bases (Gb) Sequencing depth
HiFi reads Pacbio_Revio Blood 18,509.21 229,34 94.89
short reads (paired-end) BGISEQ DNBseq-T7 Blood 150 129.62 51.85
Hi-C reads BGISEQ DNBseq-T7 Blood 150 249.05 103.04
RNA-Seq reads Illumina NovaSeq 6000 heart, liver, spleen, lung, kidney, stomach, longissimus dorsi muscle, adipose, duodenum, jejunum, ileum, cecum, colon and rectum 150 An average of 10.97 Gb for each tissue

For SMRT sequencing library construction, genomic DNA was sheared, end-repaired, and ligated to SMRTbell adapters after exonuclease-mediated overhang removal. DNA damage was repaired prior to size selection (BluePippin system). Post-ligation exonuclease digestion eliminated the unligated fragments. Library quality was verified suing Qubit quantification and Fragment Analyzer sizing. Sequencing generated 12.39 million CCS reads (229.34 Gb, Table 1).

Hi-C libraries were constructed using 10 mL blood from a male Anqing Six-white pig. Chromatin was crosslinked (paraformaldehyde), digested (restriction enzyme), and ligated to preserve spatial interactions. After blunt-end repair with biotinylated dNTPs, crosslinks were reversed and the DNA was sheared (300–500 bp). Biotin-tagged fragments were enriched using streptavidin beads, followed by Illumina adapter ligation and PCR optimization. Library quality was verified by fluorometry (Qubit), fragment analysis (Bioanalyzer), and qPCR. Sequencing generated 257 Gb raw data, which were filtered to 249 Gb clean data (Table 1) with Fastp v0.23.2.

Total RNA was extracted from each tissue using Trizol reagent (Invitrogen, Carlsbad, CA, USA) according to the manufacturer’s instructions. Quantity and purity were analysed using a Bioanalyzer 2100 and RNA 6000 Nano Labchip Kit (Agilent, Palo Alto, CA, USA). Only RNA samples with suitable RNA electrophoresis results (28 S/18 S ≥ 1.0) and RNA integrity number (RIN) ≥ 7.5 could be analysed further. The RNA-seq libraries were sequenced on an Illumina NovaSeq 6000, which generated 150-bp paired-end reads. An average of 10.97 Gb for each tissue were obtained (Table 1). These RNA-seq data were used for whole genome protein-coding gene prediction.

Genome size estimation and De novo genome assembly

Prior to genome assembly, key genomic features such as genome size and heterozygosity rate composition could be computationally estimated using the K-mer method. The 129.62 Gb clean data were subjected to a 17-mer frequency distribution analysis using GCE19 (v1.0.0) software (Fig. 2). The following equation was used to estimate the genome size of Anqing Six-end-white pig: G = K-mer number/K-mer Depth (where K-num is the total number of 17-mers, and K-depth denotes the K-mer depth, and G represents the genome size). To account for the confounding effects of genome heterozygosity on genome size estimation, K-mer exhibiting a depth of 1 were identified as sequencing errors. The derived error rate was used to refine the genome size calculation. The formula used was Revised Gsize = G x (1-Error Rate). The genome size and heterozygosity of Anqing Six-end-white pig were estimated to be 2.4 Gb and 0.61%.

Fig. 2.

Fig. 2

The frequency distribution of k-mer for Anqing Six-end-white pig genome (k = 17).

De novo assembly of Anqing Six-end-white pig was performed using PacBio HiFi data and HiFiasm20 (v0.19.6) software. The HiFiasm assembly pipeline was constructed in three critical phases: (1) error correction and haplotype phasing of raw reads; (2) iterative construction of assembly graphs; and (3) refinement to generate chromosome-level contigs. To eliminate heterozygosity-induced redundancy, the purge_haplotigs21 (v1.0.4) software was applied to filter misassembled contigs by integrating the read coverage patterns and pairwise sequence alignment metrics. Contaminant sequences were systematically identified and removed by alignment against the NT database (v2023.07.04, https://ftp.ncbi.nlm.nih.gov/blast/db/) using BLASTN22 (v2.11.0+). The genome size was 2.69 Gb, composed of 62 contigs, and the contig N50 was 90.48 Mb with a GC content of 42.64% (Table 2).

Table 2.

Assembly statistics for the Anqing Six-end-white pig.

Stat Type Polished Genome Hi-C-based Chromosome-level Genome
Contig Length(bp) Contig Number Contig Length(bp) Contig Number Scaffold Length (bp) Scaffold Number
N50 90,477,219 10 90,477,219 10 143,096,131 8
N60 68,892,161 13 68,892,161 13 141,975,503 9
N70 61,806,158 17 61,806,158 17 135,795,207 11
N80 52,468,784 22 53,390,153 21 91,738,786 14
N90 30,994,865 28 31,031,647 27 69,194,913 17
Longest 291,651,123 1 291,651,123 1 291,651,123 1
Total length 2,695,021,661 62 2,660,439,557 53 2,660,441,857 30

Hi-C assisted scaffolding

Following quality control of Hi-C data, the validated reads were processed through a hierarchical assembly pipeline: (1) data integrity was assessed using HiCUP23 (v0.7.2); (2) alignment to the draft genome was performed with Juicer24 (v1.5.6); and (3) chromosome-scale scaffolding was achieved through iterative integration of 3d-DNA (v180.922, https://github.com/theaidenlab/3d-dna) and HapHiC25 (v1.0.2), with manual curation guided by JuiceBox26 (v1.11.08) visualization of contact matrices. The final genome assembly was 2.66 Gb with contig N50 = 90.48 Mb and scaffold N50 = 143.10 Mb (Table 2). There were 23 gaps in the final assembled genome. The chromosomal anchoring efficiency was 98.80%. A heatmap was drawn for the interactions between the chromosomes (Fig. 3). The comparison between Anqing Six-end-white pig assembled genome and Sscrofa 11.1 was shown in Table 3.

Fig. 3.

Fig. 3

The Hi-C heatmap of chromosome interaction in Anqing Six-end-white pig at a 1 Mb resolution. The color from light to dark indicates an increase in interaction strength.

Table 3.

The genome comparison of Anqing Six-end-white pig and Sscrofa 11.1.

Statistic Anqing Six-end-white pig Sscrofa 11.1
Total length of assembly (Gb) 2.66 2.50
Max length (Mb) 291.65 274.33
Contig N50 (Mb) 90.48 48.2
scaffold N50 (Mb) 143.10 88.2
Number of chromosomes 20 20
Gaps 23 103
Protein-coding genes 20,809 20,661
Average length of CDS (bp) 1,649 1,668

Repetitive element identification

Repetitive sequence annotation was performed through an integrated pipeline combining three complementary strategies: (1) de novo prediction of tandem repeats using TRF27 (v4.09) software; (2) homology-based identification via RepeatMasker28 (v open-4.0.9) and RepeatProteinMask against the RepBase database (v20181026, http://www.girinst.org/repbase); and (3) a custom repeat library was constructed by integrating RepeatModeler29 (v open-1.0.11) and LTR_FINDER_parallel30 (v1.0.7), and was subsequently applied RepeatMasker (v open-4.0.9) to systematically annotate repetitive elements across the genome by de novo predictions. Consensus annotations were generated by merging non-redundant predictions from all approaches. A total of 1.16 Gb repeat sequences (approximately 43.52% of the assembled genome) were identified, of which 99.8% were classified as known repeat sequences (Table 4) Fig. 4.

Table 4.

The detail information of repetitive sequences in the Anqing Six-end-white pig genome.

RepBase TEs TE Proteins De novo Combined TEs
Length (bp) %in Genome Length (bp) %in Genome Length (bp) %in Genome Length (bp) %in Genome
DNA 75,137,988 2.82 3,918,537 0.15 27,754,199 1.04 76,552,602 2.88
LINE 504,231,281 18.95 234,087,295 8.8 623,860,264 23.45 761,043,543 28.61
SINE 22,069,496 0.83 0 0 19,348,005 0.73 31,208,933 1.17
LTR 142,420,859 5.35 6,495,402 0.24 72,484,253 2.72 145,581,048 5.47
Satellite 156,660,535 5.89 0 0 3,766,729 0.14 156,954,886 5.9
Other 2,398 0 459 0 0 0 2,857 0
Unknown 1,202,506 0.05 8,610 0 1,400,927 0.05 2,612,043 0.1
Total 891,443,376 33.51 244,476,387 9.19 747,749,479 28.11 1,157,828,521 43.52

Fig. 4.

Fig. 4

Schematic illustration of the genetic features of the Anqing Six-end-white pig genome in the 500 kb nonoverlapping windows. (a) GC content distribution. (b) Gene density distribution. (c) Total repeat density distribution. (d) LTR density distribution. (e) LINE density distribution. (f) DNA-TE density distribution.

Protein-coding genes prediction

Gene annotation was performed through an integrative multi-evidence framework: (1) ab initio predictions were generated using Augustus31 (v3.3.2) and Genscan32 under default parameters; (2) homology-based inference integrated cross-species protein alignments from Homo sapiens (GCF_000001405.40_GRCh38.p14), Mus musculus (GCF_000001635.27_GRCm39), Phacochoerus africanus (GCF_016906955.1_ROS_Pafr_v1), and Sus scrofa (GCF_000003025.6_Sscrofa11.1) through miniprot33 (v0.11-r234); and (3) transcriptomic evidence was incorporated through HISAT234 (v2.1.0)-guided RNA-seq alignments, followed by transcript assembly with StringTie35 (v1.3.5) and ORF prediction via TransDecoder (v5.5.0, https://github.com/TransDecoder/TransDecoder). BUSCO36 (Benchmarking Universal Single-Copy Orthologs, v5.7.1) was used for predictions based on Augustus (v3.3.2) and lineage-specific ortholog libraries (laurasiatheria_odb10). MAKER237 (v2.31.10) was used to integrate preliminary gene models into a consensus set. Final refinement was performed using HiFAP (developed in-house). In total, 20,809 protein-coding genes were identified, which harbored 36,142 transcripts with an average of 9.48 exons per gene (Table 5).

Table 5.

The detail information of protein-coding gene of Anqing Six-end-white pig.

Method Gene set Protein coding gene number Average gene length (bp) Average CDS length (bp) Average exon per gene Average exon length (bp) Average intron length (bp)
De novo Genscan 55,108 27,992 1,061 6.59 160.96 4,816
AUGUSTUS 29,226 25,201 1,011 6.03 167.53 4,807
Homolog Phacochoerus africanus 50,599 48,483 1,967 11.58 169.8 4,395
Sus Scrofa 66,249 58,367 2,034 12.05 168.8 5,097
Mus musculus 96,754 57,868 2,037 11.84 172.03 5,149
Homo sapiens 145,739 56,242 2,021 12.07 167.36 4,896
Trans ORF RNAseq 13,190 62,793 1,855 11.62 490.18 5,378
BUSCO 12,233 40,934 1,862 11.05 168.56 3,889
MAKER 22,200 50,136 1,436 9.34 447.28 5,508
HiFAP 20,809 46,025 1,649 9.48 325.69 5,063

Gene function annotation

Protein-coding genes were functionally analyzed using ten datasets: NR_Annotation, SwissProt_Annotation, TrEMBL_Annotation, KOG_Annotation, TF_Annotation, InterPro_Annotation, GO_Annotation, KEGG_ALL_Annotation, KEGG_KO_Annotation, and Pfam_Annotation. Functional annotation was performed through dual complementary strategies: (1) sequence similarity analysis using diamond38 blastp against TrEMBL39, SwissProt39, NR40, KOG (https://ftp.ncbi.nih.gov/pub/COG/KOG/), COG41, GO42, and KEGG43 databases, with KOBAS44 (v3.0) mapping KEGG orthologs to metabolic pathways; and (2) domain/motif profiling through InterProScan45 (v5.61–93.0) integrating InterPro46 databases to obtain protein conserved sequences, motifs, domains and other information. Hmmscan from the HMMER3 (v3.3.1) software was used to annotate conserved sequence information, such as transcription factors, Pfam, and motif based on multiple sequence alignment and the hidden Markov model47. A total of 20,638 genes (99.18% of predicted genes) were annotated (Table 6).

Table 6.

The gene function annotation of Anqing Six-end-white pig assembly.

Type Number Percent (%)
Annotation NR 20,563 98.82
SwissProt 19,867 95.47
TrEMBL 20,241 97.27
KOG 17,896 86
TF 5,432 26.1
InterPro 19,767 94.99
GO 15,436 74.18
KEGG_ALL 20,469 98.37
KEGG_KO 14,716 70.72
Pfam 18,850 90.59
Total Annotated 20,638 99.18
Total 20,809

Annotation of non-coding RNAs (ncRNA)

RNA annotation employed species-specific approaches based on molecular features: (1) tRNA were identified using tRNAscan-SE48; (2) rRNA were detected via BLASTN against conserved ribosomal sequences from phylogenetically proximal species (Sus scrofa and Phacochoerus africanus); and (3) miRNAs and snRNAs were annotated with INFERNAL49 (v1.1.4) by scanning against the Rfam50 (v14.8) database. The predicted non-coding genes included 848 miRNAs, 4544 tRNA, 253 rRNA, and 2156 snRNAs in the Anqing Six-end-white pig genome (Table 7).

Table 7.

The annotation of non-coding RNAs in the Anqing Six-end-white pig assembly.

Type Copy Average length (bp) Total length (bp) Percentage of genome (%)
miRNA miRNA 848 79 66,852 0.002513
tRNA tRNA 4,544 76 343,418 0.012908
rRNA 18S 15 1,830 27,443 0.001032
28S 13 2,029 26,383 0.000992
5.8S 9 153 1,380 0.000052
5S 216 114 24,592 0.000924
snRNA snRNA 2,156 108 233,781 0.008787

Detecting telomeric and centromeric regions

Telomeric regions were systematically identified across the Anqing Six-end-white pig genome using quarTeT51 software (version 1.1.4), which detects the consensus telomeric repeat motif AACCCT. Among all chromosomes, 18 exhibited telomeric repeats at both termini, whereas two chromosomes showed terminal repeats at a single end (Table 8). According to the results of repeat annotation in the genome annotation, regions with high repeat distribution were selected as candidate centromere inputs, and srf52 software ((https://github.com/lh3/srf)) was used for centromere prediction. The srf software identifies multiple centromere candidate regions and multiple repeat types, and then selects the candidate regions to obtain the predicted centromere region (Table 8). Visualization of porcine telomeres and centromeres is shown in Fig. 5.

Table 8.

The telomeric and centromeric regions of Anqing Six-end-white pig genome.

Chromosome ID Chromosome length Centromere start Centromere end Telomeric start Telomeric end
chr1 291651123 94304684 102432636 94304684 102432636
chr2 167358806 53390491 60379944 53390491 60379944
chr3 145918656 21760542 24597310 21760542 24597310
chr4 135795207 46873747 48303035 46873747 48303035
chr5 114037171 45184459 47337867 45184459 47337867
chr6 175657206 41414533 43141051 41414533 43141051
chr7 135030497 28660374 29642917 28660374 29642917
chr8 141975503 54701714 55595496 54701714 55595496
chr9 143096131 65226772 66100890 65226772 66100890
chr10 89390217 37063066 37775566 37063066 37775566
chr11 84968970 37859363 39289098 37859363 39289098
chr12 68089547 30285789 35047072 30285789 35047072
chr13 214496550 2725761 4424935 2725761 4424935
chr14 159369480 1400131 3864969 1400131 3864969
chr15 163383680 9907411 16350331 9907411 16350331
chr16 91738786 1998232 3157604 1998232 3157604
chr17 68113027 2655250 4232892 2655250 4232892
chr18 69194913 1790 13167136 1790 13167136
chrX 136280807 55540690 59138797 55540690 59138797
chrY 33018292 1530593 18385119 1530593 18385119

Fig. 5.

Fig. 5

Overview of the telomeres and centromeres. The upper part of the chromosome represents the overall length of the chromosome and the location of the assembled telomeres and centromeres. The lower part of the chromosome represents the distribution of the assembled conting on the corresponding position of the chromosome.

Data Records

The assembled genome has been deposited at NCBI GenBank under the accession number GCA_050231125.153. The PacBio HiFi, Hi-C, short reads, and RNA-seq data were submitted to the Genome Sequence Archive of the China National Center for Bioinformation (https://ngdc.cncb.ac.cn/gsa/) under the project PRJCA03772254,55. The accession number of CRR1702913 and CRR1702914 for Hi-C. The accession number of CRR1702915 for short reads. The accession number of CRR1702916 and CRR1702917 for PacBio HiFi. The accession number of from CRR1702918 to CRR1702959 for RNA-seq. Additionally, files containing the protein-coding gene annotation, non-coding RNA prediction, and repeat annotation of Anqing Six-end-white pig have been deposited in the Figshare database56 (10.6084/m9.figshare.28943891).

Technical Validation

Multiple evaluation metrics were utilized to assess the quality and robustness of the Anqing Six-end-white pig assembly genome. Firstly, the short reads, PacBio HiFi reads, and RNA-seq data were mapped to the genome with bwa57 (v0.7.12-r1039), minimap258 (v2.24-r1122), and Hisat2, respectively. Alignment rates were 99.14% for short-read data, 99.99% for PacBio HiFi reads, and 89.43% for RNA-seq data. Secondly, following short-read alignment, variants (SNPs and InDels) were identified using SAMtools59 (v1.9), Picard (v1.124), and GATK60 (v4.4.0.0). The homozygous SNP rate (0.001%), homozygous InDel rate (0.001%), heterozygous SNP rate (0.346%), and heterozygous InDel rate (0.069%). The exceptionally low homozygous rates indicate high assembly accuracy. Thirdly, The Merqury61 software (v1.3) was employed to assess the quality values (QV) of Anqing Six-end-white pig genome using a combination of short-read and PacBio HiFi reads. The results revealed that the QV based on short-read and PacBio HiFi reads were 49.3113 and 72.6645 respectively. Fourthly, BUSCO36 (v5.7.1) and Compleasm62 (v0.2.6) analyses were conducted to evaluate the completeness of our assembly by laurasiatheria_odb10. BUSCO analyses revealed that 12,071(98.67%) of the 12,234 conserved single-copy genes in our assembly, of which 12034 were single, 37 were duplicated, and 78 were fragmented matches (Table 9). Compleasm analyses revealed that 12,211(99.81%) of the 12,234 conserved single-copy genes in our assembly, of which 12193 were single, 18 were duplicated, and 9 were fragmented matches (Table 10). Overall, these findings indicate the high quality of the genome assembly.

Table 9.

The BUSCO results of the Anqing Six-end-white pig assembly.

Type Assembly Annotation
Proteins Percentage (%) Proteins Percentage (%)
Complete BUSCOs 12,071 98.67 11,948 97.66
Complete Single-Copy BUSCOs 12,034 98.37 11,899 97.26
Complete Duplicated BUSCOs 37 0.3 49 0.4
Fragmented BUSCOs 78 0.64 44 0.36
Missing BUSCOs 85 0.69 242 1.98
Total BUSCO groups searched 12,234 100 12,234 100

Table 10.

The Compleasm results of the Anqing Six-end-white pig assembly.

Type Assembly Annotation
Proteins Percentage (%) Proteins Percentage (%)
Complete BUSCOs 12,211 99.81 11,951 97.69
Complete Single-Copy BUSCOs 12,193 99.66 11,458 93.66
Complete Duplicated BUSCOs 18 0.15 493 4.03
Fragmented BUSCOs 9 0.07 47 0.38
Missing BUSCOs 14 0.11 236 1.93
Total BUSCO groups searched 12,234 100 12,234 100

Acknowledgements

This work was financially supported by the grants from Anhui Provincial Financial Agricultural Germplasm Resources Protection and Utilization Fund Project, The Special Fund for Anhui Agricultural Research System (AHCYJSTX-05-12, AHCYJSTX-05-23), Anhui Provincial Science and Technology Mission Project, Anhui Province Academic and Technical Leader Candidate Project (No.2022H300), 2025 Institutional Research Program of Anhui Academy of Agricultural Sciences (2025YL043). We thank Wuhan Onemore-tech Co. Ltd. for their kindly helps in data analysis.

Author contributions

C.-L.W. and W.Z. conceived and designed the experiments; W.Z. designed the analytical strategy and performed analysis processes. M.Z., S.-G.S. and L.-Q.L. performed the sampling; M.L. extracted the genomic DNA and RNA; The initial draft of the manuscript was authored by W.Z., who incorporated detailed feedback from all the authors. All authors read and approved the final manuscript.

Code availability

The software versions, settings and parameters used are described below. No custom code was used during this study for the curation and/or validation.

Fastp v0.23.2: -l 50

GCE v1.0.0: -k 17

HiFiasm v0.19.6: default

purge_haplotigs v1.0.4: default

BLASTN v2.11.0+: -evalue 0.00001 -max_hsps 1

HiCUP v0.7.2: default

Juicer v1.5.6: default

3d-DNA v180.922: -r 0

HapHiC v1.0.2: --NM 3

JuiceBox v1.11.08: default

TRF v4.09: 2 7 7 80 10 50 2000 -d -h

RepeatMasker v open-4.0.9: -nolow -no_is -norna

RepeatModeler v open-1.0.11: default

LTR_FINDER_parallel v1.0.7: default

Miniprot v0.11-r234: --gff-only -O 11 -E 1 -F 23 -C 1 -B 5 -G 200000 -j 1

HISAT2 v2.1.0: --dta

StringTie v1.3.5: default

TransDecoder v5.5.0: default

BUSCO v5.7.1: --miniprot -m genome -f --offline

MAKER2 v2.31.10: max_dna_len = 3000000 min_contig = 10000 pred_flank = 500 min_protein = 30

HMMER3 v3.3.1: hmmsearch -E 1e-05 --domE 1e-05

Diamond v2.0.14: --evalue 1e-05

tRNAscan-Se v1.3.1: -q

Rfam v14.8: cmscan --rfam –nohmmonly

quarTeT v1.1.4: -c animal

bwa v0.7.12-r1039: default

minimap2 v2.24-r1122: default

SAMtools v1.9: default

Picard v1.124: default

GATK v4.4.0.0: default

Competing interests

We declare that we do not have any commercial or associative interests that represent conflicts of interest in connection with the submitted work.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Shiguang Su, Email: ssg92@163.com.

Chonglong Wang, Email: ahwchl@163.com.

References

  • 1.Ai, H. et al. Adaptation and possible ancient interspecies introgression in pigs identified by whole-genome sequencing. Nature genetics47, 217–225, 10.1038/ng.3199 (2015). [DOI] [PubMed] [Google Scholar]
  • 2.Frantz, L. A. F. et al. Evidence of long-term gene flow and selection during domestication from analyses of Eurasian wild and domestic pig genomes. Nature genetics47, 1141–1148, 10.1038/ng.3394 (2015). [DOI] [PubMed] [Google Scholar]
  • 3.Groenen, M. A. M. et al. Analyses of pig genomes provide insight into porcine demography and evolution. Nature491, 393–398, 10.1038/nature11622 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Larson, G. et al. Worldwide phylogeography of wild boar reveals multiple centers of pig domestication. Science307, 1618–1621, 10.1126/science.1106927 (2005). [DOI] [PubMed] [Google Scholar]
  • 5.China National Commission of Animal Genetic Resource. Animal Genetic Resource in China. Pigs; China Agriculture Press: Beijing, China, (2011).
  • 6.Swindle, M. M., Makin, A., Herron, A. J., Clubb, F. J. Jr & Frazier, K. S. Swine as models in biomedical research and toxicology testing. Veterinary pathology49(2), 344–356, 10.1177/0300985811402846 (2012). [DOI] [PubMed] [Google Scholar]
  • 7.Li, M. et al. Comprehensive variation discovery and recovery of missing sequence in the pig genome using multiple de novo assemblies. Genome research27, 865–874, 10.1101/gr.207456.116 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Warr, A. et al. An improved pig reference genome sequence to enable pig genetics and genomics research. Gigascience 9, 10.1093/gigascience/giaa051 (2020). [DOI] [PMC free article] [PubMed]
  • 9.Wang, Y. L. et al. RNA sequencing analysis of the longissimus dorsi to identify candidate genes underlying the intramuscular fat content in Anqing Six-end-white pigs. Animal genetics54(3), 315–327, 10.1111/age.13308 (2023). [DOI] [PubMed] [Google Scholar]
  • 10.Zhang, W. et al. Identification of Signatures of Selection by Whole-Genome Resequencing of a Chinese Native Pig. Frontiers in genetics11, 566255, 10.3389/fgene.2020.566255 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhang, X. D. et al. Association between ADSL, GARS-AIRS-GART, DGAT1, and DECR1 expression levels and pork meat quality traits. Genetics and molecular research14(4), 14823–30, 10.4238/2015.November.18.47 (2015). [DOI] [PubMed] [Google Scholar]
  • 12.Hu, H. et al. Comparative analysis of meat sensory quality, antioxidant status, growth hormone and orexin between Anqingliubai and Yorkshire pigs. J. Appl. Anim. Res47, 357–361, 10.1080/09712119.2019.1643729 (2019). [Google Scholar]
  • 13.Du, H. et al. Chromosome-level genome assembly of Huai pig (Sus scrofa). Scientific data11(1), 1072, 10.1038/s41597-024-03921-w (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Wang, Y. et al. A chromosome-level genome of Chenghua pig provides new insights into the domestication and local adaptation of pigs. International journal of biological macromolecules270, 131796, 10.1016/j.ijbiomac.2024.131796 (2024). [DOI] [PubMed] [Google Scholar]
  • 15.Li, D. et al. Pangenome and genome variation analyses of pigs unveil genomic facets for their adaptation and agronomic characteristics. iMeta3(6), e257, 10.1002/imt2.257 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Ma, H. et al. Long-read assembly of the Chinese indigenous Ningxiang pig genome and identification of genetic variations in fat metabolism among different breeds. Molecular ecology resources22(4), 1508–1520, 10.1111/1755-0998.13550 (2022). [DOI] [PubMed] [Google Scholar]
  • 17.Zhou, R. et al. The Meishan pig genome reveals structural variation-mediated gene expression and phenotypic divergence underlying Asian pig domestication. Molecular ecology resources21(6), 2077–2092, 10.1111/1755-0998.13396 (2021). [DOI] [PubMed] [Google Scholar]
  • 18.Chen, S. et al. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics34(17), i884–i890, 10.1093/bioinformatics/bty560 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Liu, B. et al. Estimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects. Quantitative Biology35, 62–67, 10.1016/S0925-4005(96)02015-1 (2013). [Google Scholar]
  • 20.Cheng, H., Concepcion, G. T., Feng, X., Zhang, H. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nature methods18(2), 170–175, 10.1038/s41592-020-01056-5 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Roach, M. J., Schmidt, S. A. & Borneman, A. R. Purge Haplotigs: allelic contig reassignment for third-gen diploid genome assemblies. BMC bioinformatics19(1), 460, 10.1186/s12859-018-2485-7 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Altschul, S. F., Gish, W., Miller, W., Myers, E. W. & Lipman, D. J. Basic local alignment search tool. Journal of molecular biology215(3), 403–410, 10.1016/S0022-2836(05)80360-2 (1990). [DOI] [PubMed] [Google Scholar]
  • 23.Wingett, S. et al. HiCUP: pipeline for mapping and processing Hi-C data. F1000Research4, 1310, 10.12688/f1000research.7334.1 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Durand, N. C. et al. Juicer Provides a One-Click System for Analyzing Loop-Resolution Hi-C Experiments. Cell systems3(1), 95–8, 10.1016/j.cels.2016.07.002 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Zeng, X. et al. Chromosome-level scaffolding of haplotype-resolved assemblies using Hi-C data without reference genomes. Nature plants10(8), 1184–1200, 10.1038/s41477-024-01755-3 (2024). [DOI] [PubMed] [Google Scholar]
  • 26.Durand, N. C. et al. Juicebox Provides a Visualization System for Hi-C Contact Maps with Unlimited Zoom. Cell systems3(1), 99–101, 10.1016/j.cels.2015.07.012 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Benson, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic acids research27(2), 573–580, 10.1093/nar/27.2.573 (1999). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Smit AFA, Hubley R, Green P. RepeatMasker Open-3.0. [http://repeatmasker.org] (1996)
  • 29.Price, A. L., Jones, N. C. & Pevzner, P. A. De novo identification of repeat families in large genomes. Bioinformatics21, i351–i358, 10.1093/bioinformatics/bti1018 (2005). [DOI] [PubMed] [Google Scholar]
  • 30.Ou, S. & Jiang, N. LTR_FINDER_parallel: parallelization of LTR_FINDER enabling rapid identification of long terminal repeat retrotransposons. Mobile DNA10, 48, 10.1186/s13100-019-0193-0 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Stanke, M. et al. AUGUSTUS: ab initio prediction of alternative transcripts. Nucleic acids research34, W435–W439, 10.1093/nar/gkl200 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Burge, C. & Karlin, S. Prediction of complete gene structures in human genomic DNA. Journal of molecular biology268(1), 78–94, 10.1006/jmbi.1997.0951 (1997). [DOI] [PubMed] [Google Scholar]
  • 33.Slater, G. S. & Birney, E. Automated generation of heuristics for biological sequence comparison. BMC bioinformatics6, 31, 10.1186/1471-2105-6-31 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kim, D., Paggi, J. M., Park, C., Bennett, C. & Salzberg, S. L. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nature biotechnology37(8), 907–915, 10.1038/s41587-019-0201-4 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Pertea, M. et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nature biotechnology33(3), 290–295, 10.1038/nbt.3122 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Manni, M., Berkeley, M. R., Seppey, M., Simão, F. A. & Zdobnov, E. M. BUSCO Update: Novel and Streamlined Workflows along with Broader and Deeper Phylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes. Molecular biology and evolution38(10), 4647–4654, 10.1093/molbev/msab199 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Holt, C. & Yandell, M. MAKER2: an annotation pipeline and genome-database management tool for second-generation genome projects. BMC bioinformatics12, 491, 10.1186/1471-2105-12-491 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Buchfink, B., Reuter, K. & Drost, H. G. Sensitive protein alignments at tree-of-life scale using DIAMOND. Nature methods18(4), 366–368, 10.1038/s41592-021-01101-x (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Coudert, E. et al. Annotation of biologically relevant ligands in UniProtKB using ChEBI. Bioinformatics39(1), btac793, 10.1093/bioinformatics/btac793 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Sayers, E. W. et al. Database resources of the national center for biotechnology information. Nucleic acids research49(D1), D10–D17, 10.1093/nar/gkaa892 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Galperin, M. Y. et al. COG database update: focus on microbial diversity, model organisms, and widespread pathogens. Nucleic acids research49(D1), D274–D281, 10.1093/nar/gkaa1018 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Ashburner, M. et al. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium. Nature genetics25(1), 25–29, 10.1038/75556 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Ogata, H. et al. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic acids research27(1), 29–34, 10.1093/nar/27.1.29 (1999). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Bu, D. et al. KOBAS-i: intelligent prioritization and exploratory visualization of biological functions for gene enrichment analysis. Nucleic acids research49(W1), W317–W325, 10.1093/nar/gkab447 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Jones, P. et al. InterProScan 5: genome-scale protein function classification. Bioinformatics30(9), 1236–1240, 10.1093/bioinformatics/btu031 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Paysan-Lafosse, T. et al. InterPro in 2022. Nucleic acids research51(D1), D418–D427, 10.1093/nar/gkac993 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Eddy, S. R. Profile hidden Markov models. Bioinformatics.14(9), 755–763, 10.1093/bioinformatics/14.9.755 (1998). [DOI] [PubMed] [Google Scholar]
  • 48.Chan, P. P., Lin, B. Y., Mak, A. J. & Lowe, T. M. tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic acids research49(16), 9077–9096, 10.1093/nar/gkab688 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Nawrocki, E. P. & Eddy, S. R. Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics29, 2933–2935, 10.1093/bioinformatics/btt509 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Griffiths-Jones, S., Moxon, S., Marshall, M., Khanna, A. & Eddy, S. R. & Bateman, A. Rfam: annotating non-coding RNAs in complete genomes. Nucleic acids research33, D121–D124, 10.1093/nar/gki081 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Lin, Y. et al. quarTeT: a telomere-to-telomere toolkit for gap-free genome assembly and centromeric repeat identification. Horticulture research10(8), uhad127, 10.1093/hr/uhad127 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Zhang, Y., Chu, J., Cheng, H. & Li, H. De novo reconstruction of satellite repeat units from sequence data. Genome research33(11), 1994–2001, 10.1101/gr.278005.123 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.NCBI GenBankhttps://identifiers.org/ncbi/insdc.gca:GCA_050231125.1 (2025).
  • 54.NGDC Genome Sequence Archive.https://ngdc.cncb.ac.cn/gsa/browse/CRA024101 (2025).
  • 55.NGDC Genome Warehouse.https://ngdc.cncb.ac.cn/gwh/Assembly/93577/show (2025).
  • 56.Zhang, W. et al. A high-quality Chromosome-level genome assembly and annotation of Anqing Six-end-white pig (Sus scrofa). Figshare10.6084/m9.figshare.28943891 (2025). [DOI] [PMC free article] [PubMed]
  • 57.Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics25(14), 1754–1760, 10.1093/bioinformatics/btp324 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics34(18), 3094–3100, 10.1093/bioinformatics/bty191 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Li, H. et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics25(16), 2078–9, 10.1093/bioinformatics/btp352 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.McKenna, A. et al. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome research20(9), 1297–1303, 10.1101/gr.107524.110 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Rhie, A. et al. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biology21, 245, 10.1186/s13059-020-02134-9 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Huang, N. & Li, H. Compleasm: a faster and more accurate reimplementation of BUSCO. Bioinformatics39(10), btad595, 10.1093/bioinformatics/btad595 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Citations

  1. NCBI GenBankhttps://identifiers.org/ncbi/insdc.gca:GCA_050231125.1 (2025).
  2. NGDC Genome Sequence Archive.https://ngdc.cncb.ac.cn/gsa/browse/CRA024101 (2025).
  3. NGDC Genome Warehouse.https://ngdc.cncb.ac.cn/gwh/Assembly/93577/show (2025).
  4. Zhang, W. et al. A high-quality Chromosome-level genome assembly and annotation of Anqing Six-end-white pig (Sus scrofa). Figshare10.6084/m9.figshare.28943891 (2025). [DOI] [PMC free article] [PubMed]

Data Availability Statement

The software versions, settings and parameters used are described below. No custom code was used during this study for the curation and/or validation.

Fastp v0.23.2: -l 50

GCE v1.0.0: -k 17

HiFiasm v0.19.6: default

purge_haplotigs v1.0.4: default

BLASTN v2.11.0+: -evalue 0.00001 -max_hsps 1

HiCUP v0.7.2: default

Juicer v1.5.6: default

3d-DNA v180.922: -r 0

HapHiC v1.0.2: --NM 3

JuiceBox v1.11.08: default

TRF v4.09: 2 7 7 80 10 50 2000 -d -h

RepeatMasker v open-4.0.9: -nolow -no_is -norna

RepeatModeler v open-1.0.11: default

LTR_FINDER_parallel v1.0.7: default

Miniprot v0.11-r234: --gff-only -O 11 -E 1 -F 23 -C 1 -B 5 -G 200000 -j 1

HISAT2 v2.1.0: --dta

StringTie v1.3.5: default

TransDecoder v5.5.0: default

BUSCO v5.7.1: --miniprot -m genome -f --offline

MAKER2 v2.31.10: max_dna_len = 3000000 min_contig = 10000 pred_flank = 500 min_protein = 30

HMMER3 v3.3.1: hmmsearch -E 1e-05 --domE 1e-05

Diamond v2.0.14: --evalue 1e-05

tRNAscan-Se v1.3.1: -q

Rfam v14.8: cmscan --rfam –nohmmonly

quarTeT v1.1.4: -c animal

bwa v0.7.12-r1039: default

minimap2 v2.24-r1122: default

SAMtools v1.9: default

Picard v1.124: default

GATK v4.4.0.0: default


Articles from Scientific Data are provided here courtesy of Nature Publishing Group

RESOURCES