Skip to main content
BMC Genomics logoLink to BMC Genomics
. 2026 May 11;27:580. doi: 10.1186/s12864-026-12936-z

Genome-wide association analysis identifies the genetic basis of body size traits in Tan sheep (Ovis aries)

Wei Ding 1, Yuwei Yang 1, Xiaoming Ma 1, Caijuan Yue 1, Li Jia 1, Jinyang Tian 2, Xin Chen 2, ZiQing Niu 2, Yan Ma 2, Minshan Sun 3, Hehua EEr 1,✉
PMCID: PMC13330306  PMID: 42115864

Abstract

Background

Tan sheep is a Chinese indigenous breed prized for dual-purpose production of premium pelts and flavorful mutton, featuring a distinctive fat-tailed phenotype. The genetic basis of its body conformation remains largely unexplored.

Results

We conducted the first whole-genome sequencing-based genome-wide association study in 249 seven-month-old Tan sheep. Comprehensive population genetics analyses confirmed five distinct genetic clusters corresponding to the five sampling sources, supporting the cohort’s suitability for the study. Using a mixed linear model, we identified 690 significant SNPs associated with eight traits: 10 for body weight, 2 for body length, 353 for body height, 105 for chest girth, 203 for tail length, 11 for tail circumference, and 6 for cannon circumference. Functional annotation revealed compelling candidate genes, including FBLN1 (weight/height), FGFBP1 (cannon circumference), and MAP3K20/FMNL3 (tail circumference).

Conclusions

These findings illuminate the genetic architecture of growth and tail development in Tan sheep, providing a valuable genomic resource for molecular breeding strategies aimed at enhancing meat production while preserving superior pelt traits.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12864-026-12936-z.

Keywords: Tan sheep, Genome-wide association study (GWAS), Whole-genome sequencing, Body size traits, Tail circumference

Background

Sheep (Ovis aries) are one of the most economically important livestock species worldwide, providing meat, wool, and leather [1, 2]. In China, the sheep population exceeds 160 million, with indigenous breeds constituting about 60% of the genetic resources [3]. Body size traits, such as body weight, body length, chest girth, and height, are critical determinants of meat production efficiency and overall productivity in sheep breeding programs [2, 4–6]. Understanding the genetic architecture of these polygenic traits is therefore essential for enhancing breeding efficiency through marker-assisted selection.

In recent years, genome-wide association studies (GWAS) have become a powerful tool for dissecting the genetic basis of complex traits in sheep. For instance, Yang et al. (2025) identified three candidate genes (SLC9C1, VSTM2A, FRG1) associated with chest girth and cannon circumference in Hulunbuir sheep using a custom 40 K SNP chip [2]. Similarly, studies on Alpine Merino sheep have integrated GWAS prior information into GS models, improving the prediction accuracy for 14-month live weight and highlighting genes such as NCAPG and LCORL as key growth regulators [7]. However, the majority of existing GWAS have relied on medium-density SNP arrays (e.g., 50 K), which are prone to ascertainment bias and may miss rare or breed-specific variants [8, 9]. Moreover, most studies have focused on commercial meat or wool breeds, often overlooking indigenous breeds with distinct phenotypic adaptations [9].

Tan sheep, a renowned indigenous carpet-wool breed from the Ningxia region, is prized for its high-quality “two-fur” pelts and flavorful meat [10]. Beyond pelt quality, Tan sheep produce tender, flavorful mutton with even fat distribution, holds geographical indication status, positioning the breed as a dual-purpose genetic resource [10, 11]. Notably, despite its economic importance, the genetic management of Tan sheep breeding populations faces challenges such as incomplete or incorrect pedigree records due to long-term enclosed breeding, which may elevate the risk of inbreeding depression and hinder genetic progress [10]. However, generations of selection focused primarily on wool and pelt characteristics have engendered a modest body frame and slower growth rate compared to specialized mutton breeds (e.g., Dorper, Suffolk), limiting the realization of its meat production potential [11–13]. Critically, Tan sheep possess a distinctive fat-tailed phenotype (quantified via tail width, length, and circumference), a key metabolic adaptation that also directly influences carcass weight, fat distribution, and local market value, yet it has been systematically omitted from all previous ovine GWAS frameworks [14–17]. However, the genetic basis of body size and tail morphology in Tan sheep remains entirely unexplored.

To decipher the unique genetic architecture of body conformation in this dual-purpose breed, the present study employed whole-genome sequencing (WGS) in a cohort of 249 Tan sheep sampled from five distinct management systems and geographic locations across Ningxia. This high-resolution approach aims to overcome the inherent ascertainment bias of chip-based studies and to capture a comprehensive spectrum of genetic variation. A GWAS was subsequently conducted for eight key conformation traits, including the economically relevant yet genetically unexplored tail measurements. The objectives were to identify significant trait-associated loci, to annotate candidate genes and pathways, and to provide a foundational genomic resource for the molecular breeding of Tan sheep. This work is expected to deliver novel insights into the genetic regulation of growth and tail development, thereby facilitating the precise genetic improvement and sustainable conservation of this valuable indigenous breed.

Materials and methods

Breeds and animals

The Tan sheep used in this study were supplied and bred by different farms or companies. Specifically, sheep from the Yanchi Tan Sheep Breeding Farm in Ningxia Hui Autonomous Region were labeled as ‘X’; free-range Tan sheep from farmers in Huianbao Town, Yanchi County, were labeled as ‘Y’; sheep from Ningxia Tongxin Yiyang Modern Animal Husbandry Co., Ltd. were labeled as ‘H’ (Black Tan sheep, with uniformly black fleece) and ‘B’; and sheep from Ningxia Wuzhong Hongsibao Tianyuan Agriculture & Animal Husbandry Sci-Tech Development Co., Ltd. were labeled as ‘N’. All Tan sheep involved in this study were managed under standard farm conditions with free access to water and feed. The sampling procedures (e.g., blood collection) were performed by trained personnel under strict aseptic conditions. No animals were sacrificed specifically for this study; samples were collected opportunistically during routine farm management and commercial processing.

Genotypic and phenotypic data

Phenotypic traits, including body weight, body length (BL), body height (BH), chest girth (CG), tail width (TW), tail length (TL), tail circumference (TC), and cannon circumference (CC), were investigated in 7-month-old Tan sheep with reference to the Chinese National Standard (GB/T 2033 − 2008). Measurements were recorded using specialized tools with animals standing naturally on flat surfaces. Body weight was measured in the morning before feeding. Body length was measured as the straight-line distance from the anterior edge of the shoulder joint to the posterior edge of the ischial tuberosity. Body height was measured as the vertical distance from the highest point of the shoulder blades (withers) to the ground. Chest girth was measured around the chest immediately behind the shoulder blades. Tail width was measured at the widest part of the tail. Tail length was measured from the anterior edge of the first caudal vertebra to the end of the tail. TC was measured as the horizontal circumference around the widest part of the tail. Cannon circumference was measured as the circumference around the upper one-third of the left forelimb cannon bone, at its narrowest point.

Sample collection and DNA extraction

3 ml blood was collected in 10-ml BD Vacutainer blood collection tubes containing EDTA, the tubes were placed on dry ice, transferred to the laboratory, and stored at − 20 °C until DNA extraction. The sample was lysed in 1 mL of lysis buffer (containing 100 µL 20 mg/mL Proteinase K and 100 µL 20% SDS) at 65 °C for 30 min. After cooling to room temperature, the lysate was centrifuged at maximum speed for 10 min. The supernatant was transferred to a new tube and mixed with an equal volume of phenol: chloroform: isoamyl alcohol (25:24:1), followed by centrifugation at maximum speed for 20 min. The aqueous phase was transferred to a fresh tube, and DNA was precipitated by adding 2/3 volume of isopropyl alcohol (with optional 1/10 volume of 3 M sodium acetate), inverting several times, and incubating at -20 °C for 1 h. The DNA pellet was collected by centrifugation at 18,213 × g for 10 min, washed with 1 mL of 75% ethanol, and air-dried. Finally, the pellet was resuspended in 25–100 µL of TE Buffer. DNA concentration was measured using a Qubit Fluorometer, and integrity and purity were assessed by agarose gel electrophoresis.

Construction of library and sequencing

1–1.5 µg genomic DNA was randomly fragmented by Covaris, the fragmented DNA was selected by Agencourt AMPure XP-Medium kit to an average size of 200–400 bp. The selected fragments were through end-repair, 3’adenylated, adapters-ligation, PCR Amplifying and the products were recovered by the AxyPrep Mag PCR clean up Kit. The double stranded PCR products were heat denatured and circularized by the splint oligo sequence. The single strand circle DNA (ssCir DNA) were formatted as the final library and qualified by QC. The qualified libraries were sequenced on MGISEQ2000 platform (performed by Bioyi Biotechnology Co., Ltd. Wuhan, China).

Bioinformatic analysis

Raw reads were first filtered using the fastp (v.0.20.0) (set to default parameters) to remove low quality reads, adapters, and reads containing poly-N [18]. Clean data were mapped onto the reference genome sequence using bwa software (0.7.17) [19]. The Ovis aries (v1.0) reference genome data was downloaded from NCBI (https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_040805955.1/) [20]. GATK (v4.2.2.0) was used for mark duplicates, filtered SNP and indel, and then GATK Haplotype Caller (-ERC GVCF) was used for calling SNP analysis [21]. The variant sites were annotated by ANNOVAR (v2018-04-16) software with default parameters after filtered [22].

EIGENSOFT (v7.2.1) (https://www.hsph.harvard.edu/alkes-price/software/) was applied in PCA analysis [23, 24]. Population structure analysis was used admixture (v1.3.0) [25, 26]. Treebest (v1.9.2) (http://treesoft.sourceforge.net/treebest.shtml) was used for phylogenetic tree construction employing neighbor-joining method and p-distance model. A final consensus tree was validated by bootstrap with 1,000 replicated data sets. Stacks package (v2.61, --min-mac 5 -r 0.8 --min-maf 0.05 --max-obs-het 0.7) was used for Genetic Diversity analysis [27]. PopLDdecay (-MaxDist 10000) was applied to calculate LD between two SNPs with in certain distance (10,000 kb) on a same chromosome [28]. The population genetic FST index was calculated by using vcftools (v0.1.16) (--FST-window-size 200000 -- FST-window-step 20000) software [29]. A GWAS for all traits based on Mixed Linear Model (MLM) was conducted using the GEMMA (https://github.com/genetics-statistics/GEMMA/) software [30]. Sampling source and first three principal components were adjusted for population stratification. A genomic relationship matrix calculated using the -gk option was included as a random effect via the -k option. The -lmm 4 option was used to perform the Wald test. Given the LD structure typical of the ovine genome, where adjacent SNPs are non-independent, we applied a significance threshold of -log10(p) ≥ 6. This threshold is commonly used in whole-genome sequencing-based GWAS to balance discovery power with false-positive control.

Results

Phenotypic analysis

Investigation of eight phenotypic traits across 249 samples revealed that Weight, Body Length (BL), Body Height (BH), Chest Girth (CG), Tail Length (TL), Tail Circumference (TC), and Cannon Circumference (CC) all exhibited typical normal distribution characteristics. In contrast, the distribution of Tail Width (TW) showed a positively skewed pattern (Fig. 1). Furthermore, body weight showed significant positive correlations with all other traits. The correlation with CG was the highest, reaching 0.74, while correlations with BL, BH, TC, and CC ranged between 0.55 and 0.66. Additionally, significant correlations exceeding 0.5 were also observed between BL and both BH and CG (Fig. 1). Summary statistics (mean, standard deviation, and range) for each trait are provided in Supplementary Table 1.

Fig. 1.

Fig. 1

Distribution and correlation of eight traits. Weight, Body length (BL), Body height (BH), Chest girth (CG), Tail width (TW), Tail length (TL), Tail circumference (TC), Cannon circumference (CC). ***: p < 0.001, **: p < 0.01, *: p < 0.05

Identification of SNP markers

Whole-genome sequencing of 249 samples yielded a total of 10.19 Tb of data, with a reference genome alignment rate exceeding 98% and an average sequencing depth of ≥ 9.34X. The quality of the data was sufficient for subsequent analyses (Supplementary Tables 2 and 3). After initial filtering, 63,014,900 SNPs and 8,163,337 indels were identified (Fig. 2A, Supplementary Tables 4 and 5). Further filtering using vcftools resulted in 23,024,781 high-quality SNPs retained for downstream analysis. These indels and SNPs were primarily distributed in intergenic, intronic, and exonic regions, as well as in the upstream and downstream regions of genes (Fig. 2B). Among these SNPs, there were 1.4 billion homozygous and 841 million heterozygous loci for the Transition type, and 592 million homozygous and 411 million heterozygous loci for the Transversion type (Fig. 2C).

Fig. 2.

Fig. 2

The distribution of SNPs and the analysis of SNPs and indels. A Chromosome distribution map. B The location information of SNPs and indels. C Transition and transversion statistics of SNPs

Evolutionary analysis

A neighbor-joining tree constructed based on genetic distance revealed that the individuals were primarily divided into five major clades. Most individuals from the same collection site clustered into a single branch, indicating potentially similar genetic backgrounds. However, a small number of individuals from the B, N, and Y groups, which originated from different collection sites, showed some intermixing within the tree (Fig. 3A). Principal Component Analysis (PCA) yielded clustering results highly consistent with the phylogenetic relationships. The sample points for the B, N, and Y groups were located close to each other, while the H and J groups were positioned distantly, suggesting greater genetic divergence between these two groups and the other subpopulations (Fig. 3B). All populations exhibited a sharp LD decay, with the R² value decreasing from 0.1 to 0.05 within the first 100 kb. This rapid decay indicates low background linkage noise and supports the precision of the marker-trait correlation analyses performed (Fig. 3C).

Fig. 3.

Fig. 3

The phylogenetic evolution tree and principal component of the all accessions. A Evolutionary tree. Blue, Orange, yellow, green, skyblue, B Principal component analysis (PCA) of 249 sheep. C LD decay analysis. Yanchi Tan Sheep Breeding Farm in Ningxia Hui Autonomous Region were labeled as ‘X’; free-range Tan sheep from farmers in Huianbao Town, Yanchi County, were labeled as ‘Y’; sheep from Ningxia Tongxin Yiyang Modern Animal Husbandry Co., Ltd. were labeled as ‘H’ (Black Tan sheep, with uniformly black fleece) and ‘B’; and sheep from Ningxia Wuzhong Hongsibao Tianyuan Agriculture & Animal Husbandry Sci-Tech Development Co., Ltd. were labeled as ‘N’

Population structure analysis.

The cross-validation (CV) error was calculated for each K value ranging from 2 to 9. The CV error decreased sharply from K = 2 to K = 5, after which the rate of decline slowed considerably. Given the plateau in CV error beyond K = 5 and consistent with the PCA and phylogenetic results (Fig. 3), K = 5 was selected as the optimal number of ancestral populations, corresponding to the five sampling sources (Fig. 4A–C). Furthermore, the kinship estimation and PCA confirmed the effectiveness of the sampling strategy.

Fig. 4.

Fig. 4

Population structure of sheep. A Cross-validation of K values. B Heat map of relationship between samples. C Stacked diagram of the Tan sheep community structure

Screening of related SNPs and candidate genes

Based on the obtained high-quality SNPs, a significance threshold was set at -log10(p-value) ≥ 6. SNPs with p-values below this threshold were considered significantly associated with the phenotypes. Through GWAS analysis, structural genes located within a defined flanking region (100 kb upstream and downstream) of these significantly associated SNPs were annotated. Subsequent analyses excluded SNPs that could not be annotated to any gene within these 100 kb regions.

The results revealed 10 SNPs significantly associated with body weight, corresponding to candidate genes R6Z07_004364 (Atlastin-2), R6Z07_004365 (PGAP2), R6Z07_005574 (RIBC2), R6Z07_005575 (Fibulin-1), R6Z07_005653, R6Z07_005654 (EEF1A lysine methyltransferase 2), R6Z07_021178 (60 S ribosomal protein L17), and R6Z07_015136 (Eukaryotic translation initiation factor 2) (Fig. 5A, Supplementary Table 6). For body length, 2 significantly associated SNPs were identified, corresponding to two genes: R6Z07_015136 and R6Z07_015423 (Fig. 5B, Supplementary Table 6).

Fig. 5.

Fig. 5

Manhattan and QQ plots of seven phenotypic traits (tail width excluded from GWAS). A Weight; B Body length; C Body height; D Chest girth; E Tail length; F Tail circumference; G Cannon circumference. The dashed line represents the threshold line for a -log10(p-value) = 6. In the Q_Q plots, the red line represents the expected value and the observed value

A total of 353 SNPs were significantly associated with body height, which were mapped to 33 candidate genes, including R6Z07_000371, R6Z07_001745, R6Z07_005574, R6Z07_005575 (Fibulin-1), R6Z07_008398 (PSME1), R6Z07_008399-R6Z07_008403, R6Z07_009773, R6Z07_009774 (Retinol dehydrogenase 10), R6Z07_016996-R6Z07_016997 (ENPP4/5), R6Z07_017005 (Adhesion G protein-coupled receptor F5), R6Z07_017021-3 (Beta-defensin 114), R6Z07_017130 (Proline-rich protein 3), R6Z07_017131 (Guanine nucleotide-binding protein-like 1), R6Z07_017132 (Ribonuclease P protein subunit p21), R6Z07_017133 (E3 ubiquitin-protein ligase TRIM39), R6Z07_017134, R6Z07_017135, R6Z07_017136 (Tonsoku-like protein), R6Z07_017138-9 (Tripartite motif-containing protein 26), R6Z07_017140 (E3 ubiquitin-protein ligase TRIM15), R6Z07_017141 (Tripartite motif-containing protein 10), R6Z07_017170-3 (Olfactory receptor 2H2) (Fig. 5C, Supplementary Table 6).

For chest girth, 105 significant SNPs were identified, corresponding to 39 genes including R6Z07_005683 (Collagen alpha-2(I)), R6Z07_005683 (N-acetylneuraminate 9-O-acetyltransferase), R6Z07_005685 (Epsilon-sarcoglycan), R6Z07_007995 (Serine protease HTRA3), R6Z07_015084 (Iroquois-class homeodomain protein), and so on (Fig. 5D, Supplementary Table 6).

Regarding tail length, 203 associated SNPs were found, mapped to 78 candidate genes, including R6Z07_002004 (Collagen alpha-4), R6Z07_002005 (Calpain-7), R6Z07_002006 (SH3 domain-binding protein 5), and so on (Fig. 5E, Supplementary Table 6).

For TC, 11 significant SNPs were identified, and eleven genes were successfully annotated within their flanking regions, including R6Z07_000096 (Sialidase-2), R6Z07_002893 (Mitogen-activated protein kinase kinase kinase 20), R6Z07_002894 (formin-like protein 3), R6Z07_005653, R6Z07_005654 (EEF1A lysine methyltransferase 2), R6Z07_012020 (phosducin), R6Z07_012021 (olfactory receptor-like protein OLF2), R6Z07_021330 (inactive 1-aminocyclopropane-1-carboxylate synthase-like protein 2), R6Z07_021331 (1-aminocyclopropane-1-carboxylate synthase-like protein 1), R6Z07_014647, R6Z07_014648(Low-density lipoprotein receptor-related protein 4) (Fig. 5F, Supplementary Table 6).

Finally, 6 SNPs were significantly associated with cannon circumference, and these were annotated to three genes: R6Z07_007303 (neuromedin-U receptor 2), R6Z07_007972 (fibroblast growth factor-binding protein 1), and R6Z07_015860 (Stabilizer of axonemal microtubules 2) (Fig. 5G, Supplementary Table 6).

Discussion

The present study represents the first whole-genome sequencing (WGS)-based dissection of body conformation genetics in Tan sheep, revealing breed-specific loci that govern growth and the unique fat-tailed phenotype. By sequencing 249 individuals, we identified 10 weight-associated SNPs, 354 height-associated SNPs, and trait-specific loci for tail morphology provides a foundational genomic resource for precision breeding of this indigenous Chinese breed.

The annotated genes within our association peaks reveal compelling biological connections to growth regulation. Notably, Fibulin-1 (FBLN1), associated with both body weight and height, encodes an extracellular matrix glycoprotein critical for elastic fiber assembly and skeletal muscle integrity. Beyond its role in skeletal development and balance, FBLN1 is closely linked to the maintenance of muscle stem cell states and stem cell homeostasis. Furthermore, it serves as a core component of microfibrils in the adipose tissue extracellular matrix, directly regulating adipocyte differentiation and adipose tissue remodeling. It participates in systemic metabolic regulation and is closely associated with fat accumulation [31, 32]. Consistent with our findings, FBLN1 has been implicated in growth regulation in other mammals, though its role in sheep GWAS has not been previously highlighted. Whether this gene is involved in regulating body weight and height in Tan sheep warrants further investigation. Intriguingly, PGAP2 (post-GPI attachment to proteins 2) is associated with body weight. This gene regulates lipid raft composition and cell surface signaling—processes fundamental to adipocyte differentiation and energy metabolism [33]. This association is particularly relevant given the Tan sheep’s propensity for intramuscular fat deposition, and aligns with recent GWAS in Hulunbuir sheep linking lipid metabolism genes to body size traits.

Moreover, we identified Fibroblast growth factor-binding protein 1 (FGFBP1) as being significantly associated with cannon circumference. This protein enhances FGF signaling activity, which is a critical pathway regulating limb skeletal growth and muscle development [34–37], suggesting a direct link between this pathway and cannon circumference. This finding aligns with research in Chinese Merino sheep, which also identified FGF11 as a candidate gene associated with body size [38] Further supporting the relevance of developmental signaling pathways to growth traits, VSTM2A—identified as one of three candidate genes associated with body size traits in Hulunbuir sheep—is a key secretory protein during early adipocyte differentiation. It positively regulates adipogenesis by activating the BMP4 signaling pathway [2, 39]. In mice, knockout of VSTM2A disrupts glucose and lipid metabolism homeostasis, leading to adipocyte hypertrophy [39]. Notably, the BMP and FGF signaling pathways are known to extensively coordinate during cell differentiation processes, such as in adipocytes and osteoblasts [40], highlighting a potential interconnected regulatory network underlying body size variation. Across these comparisons, we note that direct SNP-level alignment between studies is challenging due to differences in genotyping platforms and reference genome versions; nevertheless, the consistency at the gene and pathway levels supports the biological relevance of our findings.

As a breed characterized by substantial caudal fat deposition, Tan sheep also exhibited multiple genes closely linked to fat accumulation within the genomic regions associated with TC, such as R6Z07_002893 (Mitogen-activated protein kinase kinase kinase 20, MAP3K20) and R6Z07_002894 (formin-like protein 3, FMNL3). As an upstream activator of the MAPK signaling pathway, MAP3K20 can activate the extracellular signal-regulated kinase 1/2 (ERK1/2) cascade. During adipose development, sustained activation of the ERK1/2 pathway promotes the transition of adipose precursor cells from the proliferation phase to the differentiation phase, ultimately driving adipocyte maturation and lipid accumulation [41, 42]. The MAPK pathway also regulates adipocyte cytoskeletal dynamics and lipolysis. FMNL3, as a cytoskeleton regulatory factor, functions in high coordination with downstream effectors of the MAPK pathway, suggesting a potential collaborative role in maintaining adipose tissue homeostasis [43, 44]. To our knowledge, these two genes have not been previously reported in ovine GWAS for tail traits, reflecting the breed-specific nature of fat-tail development in Tan sheep.

Our population structure analyses (Figs. 3 and 4) revealed that the five Tan sheep populations formed distinct genetic clusters, with limited admixture observed only in a few individuals from the B, N, and Y groups. Populations from the Yanchi Tan Sheep Breeding Farm (X) and the black-coated Tongxin population (H) exhibited relatively high genetic integrity, whereas free-range (Y) and some commercial populations (N) showed slightly greater heterogeneity. Similar patterns have been reported in other indigenous breeds: Antonopoulou et al. (2025) demonstrated that Greek Florina and Chios sheep from national stationary stocks maintained lower heterozygosity and clearer genetic differentiation than Karagouniko sheep from a commercial farm, underscoring the value of structured breeding programs for conserving indigenous genetic resources [45]. Our findings reinforce this notion and suggest that the conservation of Tan sheep should prioritize the maintenance of genetically distinct lineages to preserve the breed’s unique genetic heritage while enabling selective breeding for improved meat and pelt production.

Certainly, in this study, unlike the research design employed for Hulunbuir sheep, we did not establish age-stratified groups. Instead, our analysis primarily relied on individuals at 7 months of age. Phenotypic measurements taken at this early stage may be relatively susceptible to environmental influences. Furthermore, non-genetic factors such as rearing conditions and nutritional composition of feed were not accounted for in assessing body measurement traits. The functional implications of the candidate genes identified herein are presently inferred solely from bioinformatics analyses, lacking molecular experimental validation. Therefore, these findings necessitate further investigation and verification.

In addition to these limitations, the relatively large number of significant SNPs, especially for body height, likely reflects the polygenic architecture of growth traits and the influence of linkage disequilibrium. Importantly, the mixed linear model used in our analysis effectively controlled for population structure, as confirmed by well-calibrated Q-Q plots. Future identity by descent based mapping or fine-mapping studies will help refine these candidate signals.

Conclusion

This study conducted the first WGS-based GWAS on 249 Tan sheep, identifying significant SNPs and candidate genes associated with eight body size and tail traits. Genes like FBLN1, FGFBP1, MAP3K20, and FMNL3 are functionally linked to growth, fat accumulation, and tail development, revealing breed-specific genetic regulation networks. These findings fill the gap in genomic research on Tan sheep’s body size phenotype and provide high-quality genetic markers for marker assisted selection. The results support precise breeding to enhance meat yield while preserving superior wool and pelt traits, laying a foundation for sustainable conservation and genetic improvement of this indigenous breed.

Supplementary Information

12864_2026_12936_MOESM1_ESM.xlsx (25.2KB, xlsx)

Supplementary Material 1: Supplementary Table 1. Summary statistics of phenotypic traits.

12864_2026_12936_MOESM2_ESM.xlsx (60.3KB, xlsx)

Supplementary Material 2: Supplementary Table 2. Sequence data statistical.

12864_2026_12936_MOESM3_ESM.xlsx (11.7KB, xlsx)

Supplementary Material 3: Supplementary Table 3. Mapping data statistical.

12864_2026_12936_MOESM4_ESM.xls (17.1KB, xls)

Supplementary Material 4: Supplementary Table 4. SNPs statistical after initial filtering.

12864_2026_12936_MOESM5_ESM.xls (12.5KB, xls)

Supplementary Material 5: Supplementary Table 5. indels statistical after initial filtering.

12864_2026_12936_MOESM6_ESM.xlsx (32.3KB, xlsx)

Supplementary Material 6: Supplementary Table 6. SNP and Genes significantly associated with the phenotypes.

Acknowledgements

We are grateful to Henan Assist Research Biotechnology Co., Ltd. (Zhengzhou, China) for assisting in the bioinformatics analysis.

Abbreviations

WGS

Whole-genome sequencing

GWAS

Genome-wide association study

BL

Body length

BH

Body height

CG

Chest girth

TL

Tail length

TC

Tail circumference

CC

Cannon circumference

TW

Tail width

PCA

Principal component analysis

CV

Cross-validation

Authors’ contributions

HE, WD, YY, and XM designed the study. WD, YY, CY, LJ, and JT performed experiments, conducted statistical analysis, and drafted original manuscript. XC, ZN, YM, and MS contributed to the analysis of metabolomics data. All authors have given approval to the final version of the manuscript.

Funding

The research project is supported by the Joint Fund of the National Natural Science Foundation of China, with the number U21A20246.

Data availability

All the raw data used in this study have been deposited at **National Genomics Data Center** BioProject ID: [PRJCA054441] (https:/ngdc.cncb.ac.cn/gsub/submit/bioproject/PRJCA054441) (https://ngdc.cncb.ac.cn/).

Declarations

Ethics approval and consent to participate

The authors declare that all methods were carried out in accordance with relevant guidelines and regulations. All experiments were performed in strict compliance according to the guidelines of the Science and Technology Ethics Committee of the Ningxia Academy of Agriculture and Forestry Sciences (approval No.NKZRDF2022006).

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Jia YL. Report on domestic animal genetic resources in China. In: Beijing: Chinese Agriculture Publishers: 2004; 2004.
  • 2.Li NN, He WZ, Ye YF, He MI, Di TI, Hao XY, Ding CQ, Yang YJ, Wang L, Wang XC. Integrated metabolomics and proteomics analyses reveal the molecular mechanism underlying the yellow leaf phenotype of Camellia sinensis. Hortic Plant J. 2025;11(1):417–30. [Google Scholar]
  • 3.Chen SY, Duan ZY, Sha T, Xiangyu J, Wu SF, Zhang YP. Origin, genetic diversity, and population structure of Chinese domestic sheep. Gene. 2006;376(2):216–23. [DOI] [PubMed] [Google Scholar]
  • 4.Ahmad SF, Khan NN, Ganai NA, Shanaz S, Rather MA, Alam S. Multivariate quantitative genetic analysis of body weight traits in Corriedale sheep. Trop Anim Health Prod. 2021;53(2):197. [DOI] [PubMed] [Google Scholar]
  • 5.Kemper KE, Visscher PM, Goddard ME. Genetic architecture of body size in mammals. Genome Biol. 2012;13(4):244. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Erdenee S, Akhatayeva Z, Pan CY, Cai Y, Xu HW, Chen H, Lan XY. An insertion/deletion within the CREB1 gene identified using the RNA-sequencing is associated with sheep body morphometric traits. Gene. 2021;775:145444. [DOI] [PubMed] [Google Scholar]
  • 7.Li CL, Li JY, Wang HF, Zhang R, An XJ, Yuan C, Guo TT, Yue YJ. Genomic Selection for Live Weight in the 14th Month in Alpine Merino Sheep Combining GWAS Information. Animals-Basel. 2023;13(22):3516. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Xu SS, Li MH. Recent advances in understanding genetic variants associated with economically important traits in sheep (Ovis aries) revealed by high-throughput screening technologies. Front Agricultural Sci Eng. 2017;4(3):279–88. [Google Scholar]
  • 9.Tuersuntuoheti M, Zhang J, Zhou W, Zhang CL, Liu C, Chang Q, Liu S. Exploring the growth trait molecular markers in two sheep breeds based on Genome-wide association analysis. PLoS ONE. 2023;18(3):e0283383. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Shen N, Wang T, Gan Q, Liu S, Wang L, Jin B. Plant flavonoids: Classification, distribution, biosynthesis, and antioxidant activity. Food Chem 2022, 383. [DOI] [PubMed]
  • 11.Cheng J, Zhang X, Xu D, Zhang D, Zhang Y, Song Q, Li X, Zhao Y, Zhao L, Li W, et al. Relationship between rumen microbial differences and traits among Hu sheep, Tan sheep, and Dorper sheep. J Anim Sci. 2022;100(9):1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Yang P, Wang G, Jiang S, Chen M, Zeng J, Pang Q, Du D, Zhou M. Comparative analysis of genome-wide copy number variations between Tibetan sheep and White Suffolk sheep. Anim Biotechnol. 2023;34(4):986–93. [DOI] [PubMed] [Google Scholar]
  • 13.McGowan E, Coffey M, Simm G, Mrode R. Modelling growth in Suffolk and Charollais sheep populations using random regression models and validation of constrained polynomial correlation values. Animal: Int J Anim bioscience. 2023;17(5):100792. [DOI] [PubMed] [Google Scholar]
  • 14.Li J, Tang C, Yang Y, Hu Y, Zhao Q, Ma Q, Yue X, Li F, Zhang J. Characterization of meat quality traits, fatty acids and volatile compounds in Hu and Tan sheep. Front Nutr. 2023;10:1072159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Zhang X, Liu C, Kong Y, Li F, Yue X. Effects of intramuscular fat on meat quality and its regulation mechanism in Tan sheep. Front Nutr. 2022;9:908355. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Yue C, Wang J, Shen Y, Zhang J, Liu J, Xiao A, Liu Y, Eer H, Zhang QE. Whole-genome DNA methylation profiling reveals epigenetic signatures in developing muscle in Tan and Hu sheep and their offspring. Front veterinary Sci. 2023;10:1186040. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Ma HEE, Xie L, Ma X, Ma J, Yue X, Ma C, Liang Q, Ding X, Li W. Genetic polymorphism association analysis of SNPs on the species conservation genes of Tan sheep and Hu sheep. Trop Anim Health Prod. 2020;52(3):915–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Chen S, Zhou Y, Chen Y, Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018;34(17):i884–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25(14):1754–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Luo LY, Wu H, Zhao LM, Zhang YH, Huang JH, Liu QY, Wang HT, Mo DX, Eer HH, Zhang LQ, et al. Telomere-to-telomere sheep genome assembly identifies variants associated with wool fineness. Nat Genet. 2025;57(1):218–30. [DOI] [PubMed] [Google Scholar]
  • 21.Van der Auwera GA, Carneiro MO, Hartl C, Poplin R, del Angel G, Levy-Moonshine A, Jordan T, Shakir K, Roazen D, Thibault J et al. From FastQ data to high confidence variant calls: the genome analysis toolkit best practices pipeline. Current protocols in bioinformatics : 2018, 43(1):11.10.11–11.10.33. [DOI] [PMC free article] [PubMed]
  • 22.Yang H, Wang K. Genomic variant annotation and prioritization with ANNOVAR and wANNOVAR. Nat Protoc. 2015;10(10):1556–66. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Price AL, Patterson NJ, Plenge RM, Weinblatt ME, Shadick NA, Reich D. Principal components analysis corrects for stratification in genome-wide association studies. Nat Genet. 2006;38(8):904–9. [DOI] [PubMed] [Google Scholar]
  • 24.Yang JA, Lee SH, Goddard ME, Visscher PM. GCTA: A Tool for Genome-wide Complex Trait Analysis. Am J Hum Genet. 2011;88(1):76–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Francis RM. pophelper: an R package and web app to analyse and visualize population structure. Mol Ecol Resour. 2017;17(1):27–32. [DOI] [PubMed] [Google Scholar]
  • 26.Alexander DH, Novembre J, Lange K. Fast model-based estimation of ancestry in unrelated individuals. Genome Res. 2009;19(9):1655–64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Catchen J, Hohenlohe PA, Bassham S, Amores A, Cresko WA. Stacks: an analysis tool set for population genomics. Mol Ecol. 2013;22(11):3124–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Zhang C, Dong SS, Xu JY, He WM, Yang TL. PopLDdecay: a fast and effective tool for linkage disequilibrium decay analysis based on variant call format files. Bioinformatics. 2019;35(10):1786–8. [DOI] [PubMed] [Google Scholar]
  • 29.Danecek P, Auton A, Abecasis G, Albers CA, Banks E, DePristo MA, Handsaker RE, Lunter G, Marth GT, Sherry ST, et al. The variant call format and VCFtools. Bioinformatics. 2011;27(15):2156–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Zhou X, Stephens M. Genome-wide efficient mixed-model analysis for association studies. Nat Genet. 2012;44(7):821–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Raman R, Antony M, Nivelle R, Lavergne A, Zappia J, Guerrero-Limón G, da Silva CC, Kumari P, Sojan JM, Degueldre C et al. The Osteoblast Transcriptome in Developing Zebrafish Reveals Key Roles for Extracellular Matrix Proteins Col10a1a and Fbln1 in Skeletal Development and Homeostasis. Biomolecules 2024, 14(2). [DOI] [PMC free article] [PubMed]
  • 32.Muthu ML, Reinhardt DP. Fibrillin-1 and fibrillin-1-derived asprosin in adipose tissue function and metabolic disorders. J Cell Commun Signal. 2020;14(2):159–73. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Kinoshita T. Biosynthesis and biology of mammalian GPI-anchored proteins. Open biology. 2020;10(3):190290. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Abud HE, Skinner JA, Cohn MJ, Heath JK. Multiple functions of fibroblast growth factors in vertebrate development. Biochem Soc Symp. 1996;62:39–50. [PubMed] [Google Scholar]
  • 35.Yamaguchi TP, Rossant J. Fibroblast growth factors in mammalian development. Curr Opin Genet Dev. 1995;5(4):485–91. [DOI] [PubMed] [Google Scholar]
  • 36.Ornitz DM, Marie PJ. Fibroblast growth factors in skeletal development. Curr Top Dev Biol. 2019;133:195–234. [DOI] [PubMed] [Google Scholar]
  • 37.Goldfarb M. Functions of fibroblast growth factors in vertebrate development. Cytokine Growth Factor Rev. 1996;7(4):311–25. [DOI] [PubMed] [Google Scholar]
  • 38.He S, Di J, Han B, Chen L, Liu M, Li W. Genome-Wide Scan for Runs of Homozygosity Identifies Candidate Genes Related to Economically Important Traits in Chinese Merino. Animals: open access J MDPI 2020, 10(3). [DOI] [PMC free article] [PubMed]
  • 39.Secco B, Camiré É, Brière MA, Caron A, Billong A, Gélinas Y, Lemay AM, Tharp KM, Lee PL, Gobeil S, et al. Amplification of Adipogenic Commitment by VSTM2A. Cell Rep. 2017;18(1):93–106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Han BY, Tian DH, Li X, Liu SJ, Tian F, Liu DH, Wang S, Zhao K. Multiomics Analyses Provide New Insight into Genetic Variation of Reproductive Adaptability in Tibetan Sheep. Mol Biol Evol 2024, 41(3). [DOI] [PMC free article] [PubMed]
  • 41.Merrett JE, Xie J, Psaltis PJ, Proud CG. MAPK-interacting kinase 2 (MNK2) regulates adipocyte metabolism independently of its catalytic activity. Biochem J. 2020;477(14):2735–54. [DOI] [PubMed] [Google Scholar]
  • 42.Yang YZLM, Peng W, Su XL, Yue BL, Shu S, Wang JK, Fu CQ, Zhong JC, Wang H. Epigenomics Analysis of the Suppression Role of SIRT1 via H3K9 Deacetylation in Preadipocyte Differentiation. Int J Mol Sci 2023, 24(14). [DOI] [PMC free article] [PubMed]
  • 43.Pan MH, Wan X, Wang HH, Pan ZN, Zhang Y, Sun SC. FMNL3 regulates FASCIN for actin-mediated spindle migration and cytokinesis in mouse oocytes†. Biol Reprod. 2020;102(6):1203–12. [DOI] [PubMed] [Google Scholar]
  • 44.Hetheridge C, Scott AN, Swain RK, Copeland JW, Higgs HN, Bicknell R, Mellor H. The formin FMNL3 is a cytoskeletal regulator of angiogenesis. J Cell Sci. 2012;125(Pt 6):1420–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Antonopoulou D, Symeon G, Zaralis K, Avdi M, Frydas IS, Giantsis IA. Genome-Wide Association Study (GWAS) on Reproductive Seasonality in Indigenous Greek Sheep Breeds: Insights into Genetic Integrity. Curr Issues Mol Biol. 2025;47(4):279. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

12864_2026_12936_MOESM1_ESM.xlsx (25.2KB, xlsx)

Supplementary Material 1: Supplementary Table 1. Summary statistics of phenotypic traits.

12864_2026_12936_MOESM2_ESM.xlsx (60.3KB, xlsx)

Supplementary Material 2: Supplementary Table 2. Sequence data statistical.

12864_2026_12936_MOESM3_ESM.xlsx (11.7KB, xlsx)

Supplementary Material 3: Supplementary Table 3. Mapping data statistical.

12864_2026_12936_MOESM4_ESM.xls (17.1KB, xls)

Supplementary Material 4: Supplementary Table 4. SNPs statistical after initial filtering.

12864_2026_12936_MOESM5_ESM.xls (12.5KB, xls)

Supplementary Material 5: Supplementary Table 5. indels statistical after initial filtering.

12864_2026_12936_MOESM6_ESM.xlsx (32.3KB, xlsx)

Supplementary Material 6: Supplementary Table 6. SNP and Genes significantly associated with the phenotypes.

Data Availability Statement

All the raw data used in this study have been deposited at **National Genomics Data Center** BioProject ID: [PRJCA054441] (https:/ngdc.cncb.ac.cn/gsub/submit/bioproject/PRJCA054441) (https://ngdc.cncb.ac.cn/).


Articles from BMC Genomics are provided here courtesy of BMC

RESOURCES