Skip to main content
Genome Biology logoLink to Genome Biology
. 2026 Jun 5;27:252. doi: 10.1186/s13059-026-04121-y

DNA methylation-based immune cell profiling in mouse blood using MouseRS-CMD

Samuel R Reynolds 1,7,#, Hannah G Stolrow 1,2,#, Min Kyung Lee 1, Steven C Pike 1, Fred W Kolling IV 2, Jacqueline Y Channon 3, Daniel W Mielcarz 4, Gary A Ward 4, Steven N Fiering 3, Jennifer Fields 3, Lucas A Salas 1,2,5, Brock C Christensen 1,2,5,6,✉
PMCID: PMC13459285  PMID: 42249416

Abstract

We present MouseRS-CMD, a murine DNA methylation-based immune cell deconvolution tool. Neutrophils, monocytes, natural killer cells, B cells, CD4 + and CD8 + T cells were purified. DNA methylation was measured with the Illumina Mouse Methylation BeadChip at > 285,000 CpG loci for reference library development and in eye and terminal bleed whole blood with flow cytometry measurements for validation. The IDOL algorithm identified an optimal reference library of 300 CpGs with RMSE < 2.7 for four cell types and < 3.5 for all cell types. MouseRS-CMD enables immune profiling in archival samples and reduces confounding from cell type heterogeneity in molecular studies.

Supplementary Information

The online version contains supplementary material available at 10.1186/s13059-026-04121-y.

Keywords: DNA methylation, Epigenetics, Immune profiling, Mus musculus, Neutrophils, Natural killer cells, B cells, Monocytes, CD8 + T cells, CD4 + T cells

Background

Murine models have been widely used in experimental research because they are phylogenetically and physiologically similar to humans, can be genetically manipulated to mimic human disease, reproduce quickly, have an accelerated lifespan, and are cost-effective. Comparative genomic studies have shown that ninety percent of the mouse, Mus musculus, and human genomes can be divided into conserved syntenic regions [1]. This similarity has advanced the study of cancer biology, immunology, drug development, and numerous other medical fields. Although Mus musculus is an effective model of human biology, key biological differences exist between the two species. First, mouse models lack the high genetic variability present in humans as mouse strains are inbred. Not only do lab mice lack genetic variability, but they also lack variability in immunological experience. Compared with feral mice and humans, who are exposed to multitudes of pathogens throughout their lives, lab mice, having been raised in a laboratory setting, will be exposed to fewer and less diverse pathogens [2]. Additionally, there are differences in the development and response of the innate and adaptive immune system in mice and humans. Mice often differ in many of the immunoglobulin receptors used when compared with humans, alongside possessing a more prominent innate compartment [3]. Moreover, leukocyte proportions in the periphery differ vastly, where human blood is neutrophil rich (50–70% neutrophils, 30–50% lymphocytes) and mouse blood is composed mostly of lymphocytes (75–90% lymphocytes, 10–25% neutrophils) [3]. Therefore, when modeling human disease in mice, monitoring and understanding granular changes in the immune system is crucial.

Currently, mouse blood is acquired through the submandibular vein, tail vein, retro-orbital bleeding, or cardiac puncture for a terminal bleed to determine blood immune profiles. Traditional methods, such as flow cytometry, are then used to determine the cell type composition of the sample. While flow cytometry is accurate and offers a high resolution immune profile, samples must be processed or specially preserved immediately as intact cells are required. Alternatives such as RNA-seq-based cell-type deconvolution [4] permit some level of sample preservation for future analysis. However, like cell membranes, RNA is labile and requires special consideration for optimal preservation. Notably, the accuracy of RNA-based deconvolution approaches is inherently limited by intrinsic variation of cell-specific transcript number within cell fraction, and the relatively low feature space of genes to define cell-specific reference gene profiles.

DNA methylation, the covalent addition of a methyl group to the 5-carbon of cytosine at CpG dinucleotides, plays a critical role in cell differentiation and gene expression, resulting in cell-type specific epigenetic patterns. Cell-specific DNA methylation patterns have been used to develop highly accurate and widely used cell-type deconvolution tools in human blood [5–9], brain tissue [10, 11], tumors [12], and other tissue types [13, 14]. Methylation cytometry uses cell-specific reference DNA methylation profiles to deconvolve cell-type proportions from typical biospecimens with high accuracy and standardized measurement tools. Advancements in array technology, from Illumina’s HumanMethylation27k microarray to their Infinium MethylationEPIC, have made these deconvolution tools possible. With the advent of the Infinium Mouse Methylation BeadChip, Zhou et al. demonstrated that murine immune cell types and tissues show distinct methylation patterns [15]. Schönung et al. further demonstrated this by constructing a reference library using murine bone marrow samples, enabling prediction of five cell types, including four immune cell types: T cells, B cells, neutrophils, and monocytes [16]. Beyond these studies, there is currently no publicly available murine blood immune cell-type specific reference deconvolution library. Our work builds upon prior research to develop a more granular and immune focused deconvolution library, enabling detailed profiling of the mouse immune system. In this study, we establish a reference library for the deconvolution of six blood immune cell proportions in Mus musculus using the Illumina Mouse Methylation BeadChip. Our library deconvolves bulk whole blood samples into six immune cell types (B cells, CD4 and CD8 T cells, monocytes, natural killer cells (NK), and polymorphonuclear neutrophils (PMN)). This reference library, hereafter named MouseRS-CMD (cell mixture deconvolution) can be applied to data from the Infinium Mouse Methylation BeadChip to study the immune profiles of mice in response to exposures, disease conditions, and treatment.

Results

An overview of study design, sample collection, flow cytometry, and DNA methylation measures for development of mouse blood DNA methylation deconvolution is shown in Fig. 1. We measured DNA methylation and performed rigorous quality assessment and control for all sample data. The final reference dataset included the following cell types: polymorphonuclear neutrophils (PMN), monocytes (Mono), B cells (B), T-helper CD4 + cells (CD4), T-cytotoxic CD8 + cells (CD8), and natural killer cells (NK). The estimated purity of reference samples is based on commonly accepted CD marker definitions (Additional file 1: Table S1). The purified immune cell types were flow-sorted using whole blood collected from two male and two female BALB/c and two male and two female C57BL/6 mice at an average age of 14.1 months (n = 8, Additional file 1: Table S2). DNA was extracted from the flow-sorted, purified cell type samples as well as whole blood samples taken via cardiac puncture terminal bleeds and retro-orbital eye bleeds (n = 30). DNA methylation data was obtained using the Illumina Infinium Mouse Methylation BeadChip which measures > 285,000 CpG loci genome-wide (Fig. 1).

Fig. 1.

Fig. 1

Study design overview for development of mouse whole blood DNA methylation cytometry deconvolution. (created with Biorender.com). A Whole blood was collected from a total of eight mice via cardiac puncture terminal bleed for the isolation and purification of immune cells (B cell, CD4 T cell, CD8 T cell, NK cell, monocyte, and PMN). Additional whole blood aliquots were taken via both cardiac puncture terminal bleed and retro-orbital eye bleeds. Cell proportions were quantified by flow cytometry for each whole blood sample. B DNA methylation profiles were generated using the Infinium Mouse Methylation BeadChip for both the purified cell types and whole blood samples (from terminal and eye bleeds). C DNA methylation data from purified cell types were used to identify a candidate CpG library as input for the IDOL algorithm. Whole blood samples (terminal and eye bleeds) were split into training and testing sets to evaluate the reference library

DNA methylation data from the purified cell types (B Cell (n = 8), CD4 T cell (n = 8), CD8 T cell (n = 8), monocyte (n = 6), NK (n = 7), and PMN (n = 7)) was compared using a series of pairwise t-tests to identify a candidate CpG library of 1,798 cell-type specific CpGs that would be used to select the optimal cell reference deconvolution library of CpG sites (Fig. 1C). The candidate library was used as an input to the IDOL algorithm. Multiple models were iteratively trained based on library size from size 200 to 700 CpGs with a step size of 50 CpG sites. The R2 and root mean square error (RMSE) values across library sizes are shown in Additional file 2: Fig. S1.

Figure 2 shows average methylation levels of the 300 CpG IDOL-identified library across the purified cell-type samples. To evaluate the 300 CpG IDOL-optimized library against the limma-based deconvolution library, we split whole blood eye and terminal bleed samples into training and testing sets using an 80/20 split. The training set included 22 whole blood samples (eye bleed n = 3 (13.6%), and terminal bleed n = 19 (86.4%)), and the testing included 8 samples (eye bleed, n = 5 (62.5%); terminal bleed, n = 3 (37.5%), Fig. 1C). The estimated cell-type proportions using IDOL-optimized library are shown in Fig. 3. Overall, the performance of our IDOL library for DNA methylation-based cell-type deconvolution was high. Test set RMSE values (representative of the percent error in cell-type estimation) were < 2.7 for all cell types except B cells and PMN, which had RMSEs of 3.47 and 3.41, respectively. Using independent testing samples, we then compared a deconvolution reference library derived using the limma-based method in the Meffil package to our IDOL cell-type reference deconvolution library and observed consistently improved R2 and RMSE values for our IDOL library. More specifically, RMSE values for the limma-based deconvolution library were higher for all cell types compared with the IDOL results for training and testing samples. In addition, the R2 values for the IDOL-optimized library were higher for all cell types in both training and testing set performance evaluations except for an equivalent R2 with the limma library for B lymphocytes (R2 = 0.92), and slightly higher R2 for CD8 T cells with the limma library (R2 = 0.74) versus IDOL CD8 T-cell results for the validation samples (R2 = 0.69, Fig. 3).

Fig. 2.

Fig. 2

Heatmap of the IDOL-optimized 300 CpG library showing the average beta value per cell type. Each column represents the average beta value across the MouseRS-CMD 300 CpG library for a given cell type. Cell types are abbreviated as follows: B, B cell; CD4, CD4 T cell; CD8, CD8 T cell; Mono, Monocyte; NK, Natural Killer Cell; PMN, Polymorphonuclear Neutrophils

Fig. 3.

Fig. 3

The IDOL-optimized library improves deconvolution when compared to the limma-identified library. A Estimated cell-type proportions per sample for the training set, R2 and root mean square error (RMSE) were estimated for the IDOL-based method projections per cell type. B Estimated cell-type proportions per sample for the testing set, R2 and RMSE were estimated for the IDOL-based method projections per cell type. C Estimated cell-type proportions per sample for the testing set, R2 and RMSE were estimated for the limma-based method projections per cell type. The cell types are abbreviated as follows: B, B cell; CD4, CD4 T cell; CD8, CD8 T cell; Mono, Monocyte; NK, Natural Killer Cell; PMN, Polymorphonuclear Neutrophils

The mean absolute error (MAE) per cell type, error per cell type, global error per method, and total absolute error per sample are shown in Fig. 4 for both limma-based and IDOL-based deconvolution. The IDOL-based library has a lower average total error across cell types and shows improved classification of the natural killer and monocyte fractions compared to the limma-based library. Genomic and functional contexts of the 300 CpG IDOL-identified library are shown in Table 1 with the majority of the CpGs mapping to TSSBody (78%) and open sea regions (82%). Additionally, Table 1 includes odds ratios (OR) and 95% confidence intervals (CI) comparing the MouseRS IDOL-optimized 300 CpG library to background probes (after filtering for probes mapping to sex and mitochondrial chromosomes, SNPs, and non-CpG probes). Additional file 2: Fig. S2 illustrates the distribution of functional regions across our 300 CpG library within each genomic context. Examples of the differentially methylated CpGs that map to leukocyte-specific genes are illustrated in Fig. 5. Of the lymphocyte-specific loci, CpG sites located on Cd19, Cd4, and Cd8a were found to be differentially methylated in B cells, CD4 T cells, and CD8 T cells, respectively. In total, 173 GO pathways (146 biological processes, 23 cellular components, and 4 molecular functions) were statistically significant (false discovery rate < 0.05) after array bias correction. Among others, several GO pathways were tracked to the parent GO terms: regulation of gamma-delta T cell differentiation, regulation of mast cell degranulation, B cell receptor signaling pathway, T cell differentiation, cytokine-mediated signaling pathway, and positive regulation of inflammatory response.

Fig. 4.

Fig. 4

Comparison of limma-based (Meffil) and IDOL-based deconvolution algorithms. A Mean absolute error per cell type, (B) error per cell type (C) global error per method, (D) and total absolute error per sample

Table 1.

Genomic and functional contexts of CpGs identified in the IDOL-optimized deconvolution library. The number and percentage of CpGs within each context across the IDOL-optimized 300 CpG library. Odds ratios (OR) and 95% confidence intervals (CI) were calculated using Fisher’s Exact Test comparing the 300 CpG library vs. the background probes (after filtering for probes mapping to sex and mitochondrial chromosomes, SNPs, and non-CpG probes).

n = 300 n (%) OR (95% CI)
Genomic context
 CpG island 8 (3) 0.23 (0.1, 0.45)*
 North shelf 9 (3) 1.01 (0.45, 1.94)
 North shore 13 (4) 0.69 (0.36, 1.20)
 South shelf 9 (3) 1.20 (0.54, 2.31)
 South shore 15(5) 0.92 (0.51, 1.54)
 Open sea 246 (82) 1.76 (1.31, 2.41)*
Functional context
 TSSBody 233 (78) 1.67 (1.27, 2.23)*
 TSS1500 14 (5) 0.67 (0.36, 1.15)
 TSS200 6 (2) 0.57 (0.21, 1.25)
 Intergenic 47 (16) 0.65 (0.47, 0.89)*

*P-value < 0.05

Fig. 5.

Fig. 5

Examples of critical CpGs for cell-type deconvolution selected by IDOL. Each plot shows the distribution of beta values across six cell types (B cells, CD4 T cells, CD8 T cells, monocytes, NK cells, and PMNs) at specific CpG loci mapping to leukocyte-specific genes (e.g., Cd19, Cd4, Cd8a)

We assessed the MouseRS-CMD library in 50 isolated cell samples (B, CD4, CD8, NK, and PMN) from Zhou et al. (GSE184410), excluding monocyte samples, which were included in our reference (Fig. 6A) [15]. The sorted B-cell and CD4 T-cell samples were largely composed of their expected cell type, with sorted B cell samples having a mean B cell proportion of 91.2% (SD: 8.8) and CD4 T-cell samples had a mean CD4 T-cell proportion of 80.4% (SD: 15.4, Fig. 6B). CD8 T-cell samples had a lower mean purity of 67.0% but with high variance (SD: 27.8), where a subset of six samples appeared more pure (Fig. 6B). NK and PMN samples had heterogeneity in deconvoluted cell type composition, with samples on average comprising less than 50% of the expected cell type (Fig. 6B).

Fig. 6.

Fig. 6

Evaluation of MouseRS-CMD library in publicly available sorted immune cell samples from GSE184410 dataset. A Immune cell type proportion in sorted cell samples from GSE184410 quantified using deconvolution with the MouseRS-CMD library. Each stacked bar chart corresponds to sorted cell type samples, for which the proportion of target cell type is approximately 100%. B Summary statistics of the estimated proportion of the expected (sorted) cell type across samples from panel A. C Unsupervised clustering heatmap of DNA methylation levels among CpGs in the MouseRS-CMD library for purified cell type samples from the GSE184410 dataset (left panel) and our MouseRS reference samples (right panel). The cell types are abbreviated as follows: B, B cell; CD4, CD4 T cell; CD8, CD8 T cell; Mono, Monocyte; NK, Natural Killer Cell; PMN, Polymorphonuclear Neutrophils

Unsupervised clustering of the DNA methylation data among the 300 CpGs in the MouseRS-CMD library was performed across sorted cell type samples from the GSE184410 dataset and our MouseRS dataset (Fig. 6C). Clustering by cell type was less consistent in the GSE184410 data (Fig. 6C left panel) than purified cell type samples from our data (Fig. 6C right panel). Compared with samples from GSE184410, more distinct DNA methylation patterns were observed across our MouseRS purified cell type samples, with all cell types clustering by lineage, including clear grouping of myeloid and lymphoid populations.

Next, using the GSE184410 isolated cell data, we constructed an IDOL-optimized reference library, herein known as GSE184410-CMD, and compared its performance to our MouseRS-CMD. Whole blood samples with known cell type proportions quantified by flow cytometry are critical in IDOL library development. To enable a direct 1:1 comparison with the MouseRS-CMD library, we utilized the whole blood samples with known cell type proportions from our dataset. Library sizes ranging from 200 to 450 (500 iterations per library size), increasing by 50 CpG per library, were evaluated. Across all library sizes, MouseRS demonstrated consistently better model fit (higher R2) and accuracy (lower RMSE) (Additional file 2: Fig. S3).

The performance of cell type deconvolution libraries from MouseRS isolated cell type data was relatively stable across library size. In contrast, libraries generated using GSE184410 data, showed greater variability in IDOL performance, with noticeable reductions in R2 at 350 and 450 CpGs (Additional file 2: Fig. S3). RMSE values for libraries from MouseRS data ranged from 7.6 to 8.4, compared to substantially higher errors (20.2–26.3) observed for libraries derived from GSE184410 data. For both datasets, the 300 CpG library performed best overall (MouseRS: R2 = 0.895, RMSE = 7.6; GSE184410: R2 = 0.86, RMSE = 20.17) (Additional file 2: Fig. S3).

Next, we performed unsupervised clustering of the DNA methylation data among the 300 CpGs in the GSE184410 IDOL-optimized library across sorted cell type samples from the GSE184410 dataset and the MouseRS dataset (Fig. 7). In the GSE184410 samples, B cells cluster together, while CD4, CD8, and monocyte samples largely clustered by cell type. However, NK cells and PMN samples were less consistently separated (Fig. 7A), whereas samples from our data set cluster perfectly by cell type (Fig. 7B).

Fig. 7.

Fig. 7

DNA methylation patterns of the GSE184410 IDOL-optimized 300 CpG library in the GSE184410 and MouseRS dataset. Heatmap of DNA methylation beta values for CpGs in the IDOL-optimized 300 CpG library derived from the GSE184410 dataset, shown across purified cell types in two datasets (A) the GSE184410 sorted cell dataset used to generate the library and (B) MouseRS dataset. Cell types are abbreviated as follows: B, B cell; CD4, CD4 T cell; CD8, CD8 T cell; Mono, Monocyte; NK, Natural Killer cell; PMN, Polymorphonuclear neutrophils

Deconvolution performance was then compared between MouseRS-CMD and GSE184410-CMD using predictions from our whole blood testing samples with known cell type proportions. MouseRS-CMD exhibited lower cell type-specific error and global error compared to GSE184410-CMD IDOL-optimized library (Fig. 8A and B). Across all cell types, the MouseRS-CMD library had significantly lower RMSE in both training (p = 0.031) and testing sets (p = 0.031). In the testing data, the RMSE values for MouseRS-CMD ranged from 1.19 to 3.47, whereas in GSE184410-CMD RMSE values ranged from 2.43 to 6.89, with cell type specific errors nearly twice as large for GSE184410-CMD compared to MouseRS-CMD. In the testing dataset, all R2 values were higher for MouseRS-CMD, except for CD8 T cells (MouseRS-CMD: 0.69, GSE184410-CMD: 0.79).

Fig. 8.

Fig. 8

Comparison of IDOL-optimized MouseRS-CMD and GSE184410-CMD algorithms in whole blood samples using flow cytometry quantification as ground truth. A Error per cell type observed in the testing whole blood samples. B Global error per CMD algorithm in testing whole blood samples. C Comparison of RMSE and R2 between MouseRS-CMD and GSE184410-CMD across training and testing whole blood samples for each cell type. RMSE was significantly lower for the MouseRS-CMD library in both the training and testing sets

Lastly, we evaluated MouseRS-CMD in the publicly available data from Barrett et al. (GSE236260) (Additional file 1: Table S3). In that study, BALB/c mice were exposed to medroxyprogesterone acetate and 7,12-dimethylbenzanthracene (P/D) to induce mammary gland tumors. Some mice were also treated with mifepristone (MIF), which is a progesterone antagonist that reduces the risk of mammary tumor development [17]. Whole blood was collected by cardiac puncture at 23 weeks or when tumor reached 200mm3 at the time of euthanasia [17]. We assessed whether immune composition differed among the four groups of P/D-exposed mice that all developed mammary gland tumors: P/D+ MIF- and P/D+ MIF+ mice on a normal diet and P/D+ MIF- and P/D+ MIF+ mice on a ketogenic diet (Fig. 9). We observed no significant differences in B cells or monocytes across the four groups. However, among the MIF- mice, there were significant diet-associated differences in CD4 T cells (p = 0.029) and CD8 T cells (p = 0.018), with higher proportions observed in the normal diet group. NK cells also differed significantly by diet in the MIF- groups (p = 0.00032), with higher proportions in the normal diet group. Within the ketogenic diet group, NK cell proportions were higher in MIF+ compared to MIF- mice (p = 0.017). Lastly, we observed a significant difference in PMNs between ketogenic and normal diets among MIF- mice (p = 0.0071), with higher proportions in the ketogenic diet group. After FDR adjustment (Benjamini-Hochberg), PMN and NK results remained significant (p < 0.05).

Fig. 9.

Fig. 9

MouseRS-CMD derived blood immune cell proportions in P/D-exposed mice with mammary gland tumors stratified by MIF status and diet (normal vs. ketogenic). Diet associated differences were observed in MIF- mice, including CD4 T cells (p = 0.029), CD8 T cells (p = 0.018), NK cells (p = 0.00032). Within diet differences showed a significant difference in NK cells between MIF- and MIF+ mice (p = 0.017). Cell types are abbreviated as follows: B, B cell; CD4, CD4 T cell; CD8, CD8 T cell; Mono, Monocyte; NK, Natural Killer cell; PMN, Polymorphonuclear neutrophils

Discussion

We offer a new method for DNA methylation-based immune cell typing in mice that leverages cell-type-specific DNA methylation profiles. We identify an optimized cell-specific reference DNA methylation library using flow-sorting purified polymorphonuclear neutrophil (PMN), monocyte, B cells, natural killer, and CD4 + and CD8 + T cells and measure genome-scale DNA methylation at > 285,000 CpG sites with the Illumina Infinium Mouse Methylation BeadChip. We make MouseRS-CMD available as an R package. This library offers improved accuracy over the Meffil-identified library, with a decrease in total error across the six immune cell types. The Identifying Optimal Libraries (IDOL) algorithm, which ranks cell-specific hyper- and hypomethylated CpGs based on t-statistics, improves accuracy and minimizes noise associated with larger libraries. The MouseRS-CMD library contains probes associated with critical biological pathways involved in immune development and differentiation, as well as some that are less well-characterized. By iteratively selecting CpG loci that optimize prediction performance across the six leukocyte subtypes, the IDOL algorithm identifies sensitive and specific differentially methylated sites that comprise the final deconvolution library.

This library includes loci that map to genes well established in immunology, especially regarding the leukocyte-specific CpGs, as well as other novel CpG sites. Of the lymphocyte-specific loci, CpG sites located on Cd19, Cd4, and Cd8a were found to be differentially methylated in B cells, CD4 T cells, and CD8 T cells, respectively. Cd19 is a transmembrane glycoprotein that is a part of the immunoglobulin superfamily. Cd19 is a known biomarker of B cells and is crucial in B cell maturation and proliferation [18, 19]. The Cd4 gene encodes a class II major histocompatibility (MHC) cell surface glycoprotein and is specific to helper T cells [20]. The Cd8a gene encodes a class I MHC cell surface glycoprotein and is specific to cytotoxic T cells [20]. Cell subpopulations that are not accounted for in the MouseRS reference library will most likely be classified by the closest related cell type. This can potentially impact the accuracy of the model as these cell types are not currently included in the cell reference library. Further studies are needed to increase the granularity of the deconvolution model and deconvolve these subpopulations. Additionally, flow-sorted immune cell proportions were used as the ground truth for the training of our model. To improve upon this, artificial mixtures could be used to provide an additional and arguably more accurate ground truth. Future work could further refine reference library selection using emerging optimization methods [21].

The MouseRS cell-specific reference library we define has numerous potential applications in preclinical settings, including DNA methylation-based immune profiling and reduction of cell-type specific confounding in epigenome-wide association studies (EWAS). Only a small amount of blood (e.g. retro-orbital eye bleed) is needed to extract the 250 ng of DNA needed for the Illumina Infinium Mouse Methylation BeadChip. Because our blood immune profiling method is DNA based, it provides study design advantages, including flexibility for longitudinal sampling and use of archival samples. In addition to evaluating the relation of preclinical therapeutics and potential toxicants on murine immune status, this method will enable EWAS to adjust for potential confounding due to cell-type-specific DNA methylation effects. Performing EWAS helps elucidate epigenetic changes related to disease states and exposures. However, without controlling for cell-type heterogeneity, the identified CpGs may not represent cell-type-independent effects. In humans, well-established methods for DNA-based cell typing have provided robust tools for peripheral blood immune cell deconvolution [5–9].

DNA methylation-based mouse immune profiling using methylation cytometry presents opportunities to advance our understanding of biological aging processes in mice. Epigenetic clocks for mice have been developed in tissue and blood using whole genome bisulfite sequencing (WGBS), reduced-representation bisulfite sequencing (RRBS), and more recently, microarrays, using the Infinium Mouse Methylation BeadChip [22–25]. Peripheral blood immune profiles are known to change with age in both humans [26, 27] and mice [28]. In humans, studies have demonstrated that residual confounding from immune cell composition can influence DNA methylation-based clock estimates, potentially limiting the rigor and reproducibility of aging research [29, 30]. Leveraging the development of DNA methylation-based immune profiling in mice enables disentangling the roles of cell type-dependent and cell type-independent processes in epigenetic age estimates, improving their accuracy and reliability.

Correcting for immune cell composition using the MouseRS reference will help reduce confounding in mouse DNA methylation studies. Murine models are widely used to study human diseases and are essential for preclinical drug efficacy trials [2, 31]. As a result, our murine immune reference library could enable the monitoring of immune responses to investigational drugs by providing detailed immune profiles throughout the duration of these trials. MouseRS-CMD improves the study of immune responses in murine models by elucidating cell-type heterogeneity and reducing cell-type-dependent confounding in molecular studies using whole blood samples. Additionally, due to the stability of DNA methylation, archived frozen samples can be used [32]. In contrast, flow cytometry requires live or intact cells and must be performed within 48 hours of sample collection or special cryopreservation used for reliable results. This makes MouseRS-CMD a valuable alternative, as it allows researchers to estimate cell-type proportions from standard frozen archival samples, expanding the flexibility and utility of archived specimens.

Conclusions

MouseRS-CMD, a novel DNA methylation-based murine immune cell deconvolution tool, accurately determines proportions of polymorphonuclear neutrophils, monocytes, B cells, natural killer cells, and CD4 + and CD8 + T cells. Unlike flow cytometry, this reference library can analyze immune profiles from frozen samples, making it a valuable resource for murine omics studies of whole blood. By reducing cell-type-dependent confounding, it can enhance the accuracy and interpretation of other molecular studies, as well. The principles of cell-specific reference library identification for MouseRS-CMD extend across cell types, offering the potential for future development of DNA methylation reference libraries to enable cell-type deconvolution in other mouse tissues.

Methods

Sample collection

A total of eight different mice were used in the isolation and purification of immune cells from whole blood samples. Cardiac puncture for a terminal bleed was used to collect blood from two male and two female mice each of two different strains purchased from the Jackson Laboratory (JAX, BALB/c, JAX #000651 and C57BL/6, JAX #000664) (Fig. 1A). 1 mL of each whole blood sample was incubated with 20 mL of 1 × BD FACS™ Lysing Solution (BD Biosciences, Catalog # 349,202) on a shaker for 10 min at room temperature to lyse red blood cells. Cell pellets were washed twice with PBS/EDTA. 20 µg of mouse Fc block was added to the cell pellet. Antibodies specific for lymphocyte or myeloid panels were added to the cell pellets. Cell pellets were washed with PBS/EDTA. The single-cell suspensions from whole blood samples were then sorted with a Fluorescence-Activated Cell Sorter (FACS; FACSAria IIIu (BD Biosciences)) to isolate B cells, CD4 and CD8 T cells, monocytes, and PMN cell types with subset-defining markers listed in Additional file 1: Table S1. Antibody staining panels were adapted from OMIP-32. The lymphocyte panel included NKp46 (CD335), CD62L, CD44, CD8, CD49b, CD45.2, TCRbeta, B220, and CD4 markers. The myeloid panel included Ly6C, Siglec-F, MHCII, CD11b, Ly6G, CD11c, and CD45.2 markers [33]. Additionally, a total of 32 mice from four different strains (BALB/c, JAX #000651, C57BL/6 JAX #000664, FVB, ICR/CD-1) were used for whole blood sample collection via cardiac puncture for terminal bleeds and retro-orbital bleeds for eye bleed samples. Splenocytes were also collected from spleen samples in these mice. Mouse whole blood and splenocytes were analyzed for cell subset frequencies via flow cytometry for eventual comparison to DNA methylation data. Single-cell suspensions were prepared as above, then stained with the following 7-color flow cytometry panel: CD4, CD3, CD45, CD11b, Ly6G, B220, and CD8a. Stained blood leukocytes and splenocytes were then analyzed on a MacsQuant10 flow cytometer to determine the frequencies of the subsets sorted above.

DNA was extracted from the whole blood samples and FACS-purified B cells, CD4, CD8, NK, monocytes, and PMNs using the QIAmp DNA Micro Kit (Qiagen, Catalog #: 56,304) according to the manufacturer's protocol. Purified gDNA samples were quantified using Invitrogen Qubit 3.0 Fluorometer broad range assay (Thermo Scientific, Catalog #Q32853). All samples were bisulfite converted using the Zymo EZ DNA methylation (Zymo Research, Catalog #D5001), and DNA methylation profiles were assayed with Infinium Mouse Methylation BeadChip containing 285,000 CpG sites (Illumina, Catalog #20,041,558). Additionally, DNA methylation data from six purified monocyte samples from Zhou et al.’s (2022) dataset [15], were also used. These samples were isolated using magnetic-activated cell sorting (MACS).

Data preprocessing, normalization, and quality control

The raw IDAT files from the Infinium Mouse Methylation BeadChip arrays were preprocessed in R 4.3.1 using minfi [34]. The minfi pipeline was adapted to the processing of the Infinium Mouse Methylation BeadChip array by using compatible annotation files [35, 36]. The data was background corrected using Noob (normal-out-of-band), and dye bias was corrected using a linear dye bias correction. Cross-reactive, polymorphic and probes on the X and Y chromosomes and mitochondrial DNA were filtered. Beta values, a measure of methylation, and defined as the ratio of methylated intensity over total intensity (M/(M + U)) constrained between 0 (unmethylated) and 1 (methylated), were calculated. Ten whole blood samples were removed due to poor quality, and eight purified monocyte samples were removed due to poor sample quality. The ten samples were removed following the first iteration of training due to high total error between predicted and observed cell type proportions. These were considered potential biological outliers. Additionally, the 8 monocyte samples were removed. These samples showed poor monocyte isolation during flow cytometry, likely due to overlapping surface markers between monocytes and neutrophils, which may have led to cross contamination and reduced purity. These monocytes were replaced by six MACS isolated monocyte samples taken from Zhou et al.'s (2022) dataset [15].

Deconvolution methods

Two deconvolution methods were used in this study: limma-based deconvolution (empirical Bayes moderated t-test using limma) and IDentification of Optimal Libraries (IDOL) for deconvolution (a machine learning protocol for CpG site selection optimization) [37, 38]. The limma-based deconvolution was used as a benchmark to which we compared the IDOL-based model.

Limma-based deconvolution

The limma-based reference library was identified using Meffil’s meffil.cell.type.specific.methylation function and resulted in a 350 CpG benchmark library. Meffil.cell.type.specific.methylation employs limma’s lmFit and eBayes functions [39]. lmFit fits a linear model for each probe, while eBayes uses these models to compute moderated t-statistics given a contrast matrix comparing the six cell types.

Candidate CpG selection

After preprocessing, the CandidateDMRFinder.v2 function from the IDOL library was used to select a library of candidate cell-type specific CpGs to be used in the IDOL algorithm. CandidateDMRFinder.v2 identifies differentially methylated loci (DML) based on the procedure described in Koestler et al. (2016). A series of two-sample t-tests compare the mean methylation beta-values for J CpGs between K cell types against the mean methylation beta values of the remaining K—1 cell types. Their t-statistics rank the J CpGs, and the top M hyper- and hypomethylated DMLs are selected based on the largest absolute value of the t-statistics for each of the K comparisons. For this study, M was set to 150. Of the 281,681 CpGs available after preprocessing, 1,798 were identified as candidate DMLs. These candidate CpGs were then used as input into the IDOL algorithm.

IDOL algorithm and model training

To identify an optimal deconvolution library, the IDOL algorithm was applied where iteratively designed libraries were evaluated against ground truth cell quantities determined using flow cytometry on whole blood specimens. Full details of the IDOL algorithm are presented in Koestler et al. (2016). Briefly, the IDOL algorithm employs a training dataset composed of samples with DNA methylation data whose underlying cell-type proportions are known to identify a set of probes that establish an optimal reference library for the deconvolution of cell mixtures. A series of t-tests comparing the mean CpG-specific methylation between each leukocyte cell type with the mean methylation across the remaining cell types identified a candidate list of probes containing leukocyte differentially methylated loci (L-DML) for each of the six cell types: PMNs, monocytes, B lymphocytes, natural killer cells, and CD4 + and CD8 + T cells. This candidate list consisted of 2*M*K total L-DMLs, where M represents the number of hyper- and hypomethylated L-DMLs selected based on the largest absolute value of the t-statistics for each of the K cell types and forms the IDOL algorithm search space. L-DML subsets of size < 2*M*K are iteratively selected and examined for their prediction accuracy in deconvolving the training dataset samples. In this study, we iterated from a library size of 200 to 700 with a step size of 50 to train eleven different models. In the first iteration of the IDOL algorithm, all 2*M*K CpGs in the candidate library have an equal probability of being selected to be included in the IDOL-selected library. We applied the constrained projection/quadratic programming approach to obtain the cell-composition estimates for each sample in the training dataset. The R2 and root mean square error (RMSE) values were then calculated for these predictions for each cell type using the estimated and known cell proportions in each sample. Next, the contribution of each individual CpG in the IDOL-selected library to the accuracy of cell-composition estimates is calculated by removing CpGs one by one from the randomly IDOL-selected library, followed by computation of the R2 and RMSE values based on cell-composition estimates obtained using the new library. The algorithm then modifies the probability of each CpG being selected in subsequent IDOL iterations based on its contributions to the accuracy of the model, upweighting CpGs that improve the accuracy and downweighting those that do not. The process is repeated at each library size for 500 iterations, with the algorithm eventually converging on an “optimal” library for deconvolution or a library with the lowest error and highest accuracy. The model was trained and tested on 30 whole blood samples with matching flow-sorted immune cell-type proportions. The data was split into a training (n = 22) and testing set (n = 8), stratified on the mouse strain using the initial_split function from the package tidymodels [40].

Genomic and functional context

Illumina’s Infinium Mouse Methylation Gene Probe Annotation File [35] and Infinium Mouse Methylation Manifest File [36] were used to provide the genomic and functional context.

The probes present in the new MouseRS library were tested for enrichment using the Gene Ontology (GO) pathways and PANTHER18.0 [41, 42].

Supplementary Information

13059_2026_4121_MOESM1_ESM.xlsx (12.9KB, xlsx)

Additional file 1: Tables S1-S3. Table S1. Flow cytometry markers in whole blood samples. Table S2. Purified cell type and whole blood sample characteristics. Table S3. Characteristics of the GSE184410 dataset used to construct the IDOL-optimized reference library.

13059_2026_4121_MOESM2_ESM.docx (379.9KB, docx)

Additional file 2: Figs. S1-S3. Performance of cell-type deconvolution as a function of CpG reference library size. Fig. S2. Distribution of gene features across CpG genomic contexts. Fig. S3. Comparison of performance across reference library sizes when trained on MouseRS and GSE184410 datasets.

13059_2026_4121_MOESM3_ESM.docx (5.7MB, docx)

Additional file 3: Point-by-point response to reviewers.

Acknowledgements

Illumina provided material support in the form of Infinium Mouse Methylation Beadchips and reagents. Methylation array measurements were carried out at Dartmouth Cancer Center in the Genomics and Molecular Biology Shared Resource, supported by P30CA023108 from the National Cancer Institute (NCI). We would like to acknowledge the Dartmouth DartLab Immunoassays & Flow Cytometry Core for their invaluable assistance with flow cytometry of the purified cell types. Flow cytometry was carried out in DartLab, the Immune Monitoring and Flow Cytometry Shared Resource (RRID:SCR_019165) at Dartmouth which is supported by NCI Cancer Center Support Grant 5P30CA023108. We are grateful for their expertise and support, which greatly contributed to the success of this research.

Peer review information

Nicolae Radu Zabet, Andrew Cosgrove and Veronique van den Berghe were the primary editors of this article and managed its editorial process and peer review in collaboration with the rest of the editorial team. The peer-review history is available as Additional file 3 and in the online version of this article.

Authors’ contributions

BCC and LAS conceptualized and designed the study. SRR and HGS drafted the manuscript. SRR, HGS, LAS, SCP contributed to data processing and analysis. MKL processed samples and performed DNA extractions. JYC, GAW, and DWM conducted flow cytometry experiments. JF performed mouse work, including the collection of terminal and eye bleeds and harvesting of mouse spleens. FK led the DNA methylation microarray data collection. HGS, BCC, LAS, SRR, DWM, MKL, SCP, JYC, and SNF critically revised the manuscript. All authors reviewed and approved the final manuscript.

Funding

This work was supported by the U.S. National Institutes of Health grants R01CA253976 and P30GM149408 (Brock C. Christensen), and the National Cancer Institute: Dartmouth Cancer Center Support Grant P30CA023108.

Data availability

The data generated and used in this study are publicly available on Gene Expression Omnibus (GEO) with the accession number GSE284254 [43]. Additional datasets used in this work are available on GEO under accession numbers GSE184410 [44] and GSE236260 [45]. All relevant code for this manuscript has been deposited in Zenodo [46] (https://doi.org/10.5281/zenodo.20126537) and is also available on GitHub [47] (https://github.com/SalasLab/MouseRS), both released under the GNU General Public License version 3.

Declarations

Ethics approval and consent to participate

All mouse work performed for this study was conducted in accordance with protocols approved by the Institutional Animal Care and Use Committee (IACUC) at Dartmouth.

Consent for publication

Not applicable.

Competing interests

LAS and BCC are co-founders of Cellintec, which had no role in this work. BCC serves as an advisor for Guardant Health, which had no role in this work. 

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Samuel R. Reynolds and Hannah G. Stolrow contributed equally to this work.

References

  • 1.Mouse Genome Sequencing Consortium, Waterston RH, Lindblad-Toh K, Birney E, Rogers J, Abril JF, et al. Initial sequencing and comparative analysis of the mouse genome. Nature. 2002;420(6915):520–62. 10.1038/nature01262 PubMed PMID: 12466850. [DOI] [PubMed]
  • 2.Masopust D, Sivula CP, Jameson SC. Of mice, dirty mice, and men: using mice to understand human immunology. J Immunol. 2017;199(2):383–8. 10.4049/jimmunol.1700453. (PubMed PMID: 28696328; PubMed Central PMCID: PMC5512602). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Mestas J, Hughes CCW. Of mice and not men: Differences between mouse and human immunology. J Immunol. 2004;172(5):2731–8. 10.4049/jimmunol.172.5.2731. [DOI] [PubMed] [Google Scholar]
  • 4.Chen Z, Huang A, Sun J, Jiang T, Qin FXF, Wu A. Inference of immune cell composition on the expression profiles of mouse tissue. Sci Rep. 2017;7(1):40508. 10.1038/srep40508. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Arneson D, Yang X, Wang K. MethylResolver—a method for deconvoluting bulk DNA methylation profiles into known and unknown cell contents. Commun Biol. 2020;3(1):1–13. 10.1038/s42003-020-01146-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Houseman EA, Accomando WP, Koestler DC, Christensen BC, Marsit CJ, Nelson HH, et al. DNA methylation arrays as surrogate measures of cell mixture distribution. BMC Bioinformatics. 2012;13(1):86. 10.1186/1471-2105-13-86. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Salas LA, Koestler DC, Butler RA, Hansen HM, Wiencke JK, Kelsey KT, et al. An optimized library for reference-based deconvolution of whole-blood biospecimens assayed using the Illumina HumanMethylationEPIC BeadArray. Genome Biol. 2018;19(1):64. 10.1186/s13059-018-1448-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Salas LA, Zhang Z, Koestler DC, Butler RA, Hansen HM, Molinaro AM, et al. Enhanced cell deconvolution of peripheral blood using DNA methylation for high-resolution immune profiling. Nat Commun. 2022;13(1):761. 10.1038/s41467-021-27864-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Teschendorff AE, Breeze CE, Zheng SC, Beck S. A comparison of reference-based algorithms for correcting cell-type heterogeneity in epigenome-wide association studies. BMC Bioinformatics. 2017;18(1):105. 10.1186/s12859-017-1511-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Pike SC, Wiencke JK, Zhang Z, Molinaro AM, Hansen HM, Koestler DC, et al. Glioma immune microenvironment composition calculator (GIMiCC): a method of estimating the proportions of eighteen cell types from DNA methylation microarray data. Acta Neuropathol Commun. 2024;12(1):170. 10.1186/s40478-024-01874-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhang Z, Wiencke JK, Kelsey KT, Koestler DC, Molinaro AM, Pike SC, et al. Hierarchical deconvolution for extensive cell type resolution in the human brain using DNA methylation. Front Neurosci. 2023;17:1198243. 10.3389/fnins.2023.1198243. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Zhang Z, Wiencke JK, Kelsey KT, Koestler DC, Christensen BC, Salas LA. HiTIMED: hierarchical tumor immune microenvironment epigenetic deconvolution for accurate cell type resolution in the tumor microenvironment using tumor-type-specific DNA methylation data. J Transl Med. 2022;20:516. 10.1186/s12967-022-03736-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Muse ME, Bergman DT, Salas LA, Tom LN, Tan JM, Laino A, et al. Genome-scale DNA methylation analysis identifies repeat element alterations that modulate the genomic stability of melanocytic nevi. J Invest Dermatol. 2022;142(7):1893-1902.e7. 10.1016/j.jid.2021.11.025PubMedPMID:34871578;PubMedCentralPMCID:PMC9163203. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Muse ME, Carroll CD, Salas LA, Karagas MR, Christensen BC. Application of Novel Breast Biospecimen Cell-Type Adjustment Identifies Shared DNA Methylation Alterations in Breast Tissue and Milk with Breast Cancer-Risk Factors. Cancer Epidemiol Biomarkers Prev. 2023;32(4):550–60. 10.1158/1055-9965.EPI-22-0405PubMedPMID:36780234;PubMedCentralPMCID:PMC10068446. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Zhou W, Hinoue T, Barnes B, Mitchell O, Iqbal W, Lee SM, et al. DNA methylation dynamics and dysregulation delineated by high-throughput profiling in the mouse. Cell Genomics. 2022;2(7):100144. 10.1016/j.xgen.2022.100144PubMedPMID:35873672;PubMedCentralPMCID:PMC9306256. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Schönung M, Hartmann M, Krämer S, Stäble S, Hakobyan M, Kleinert E, et al. Dynamic DNA methylation reveals novel cis-regulatory elements in mouse hematopoiesis. Exp Hematol. 2023;117:24-42.e7. 10.1016/j.exphem.2022.11.001. (PubMed PMID: 36368558). [DOI] [PubMed] [Google Scholar]
  • 17.Barrett JE, Herzog CM, Aminzadeh-Gohari S, Redl E, Ishaq Parveen I, Rothärmel J, et al. Epigenetic signatures in surrogate tissues are able to assess cancer risk and indicate the efficacy of preventive measures. Commun Med. 2025;5(1):1–10. 10.1038/s43856-025-00779-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Wang K, Wei G, Liu D. CD19: a biomarker for B cell development, lymphoma diagnosis and therapy. Exp Hematol Oncol. 2012;1(1):36. 10.1186/2162-3619-1-36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Wang Y, Carter RH. CD19 Regulates B Cell Maturation, Proliferation, and Positive Selection in the FDC Zone of Murine Splenic Germinal Centers. Immunity. 2005;22(6):749–61. 10.1016/j.immuni.2005.04.012. [DOI] [PubMed] [Google Scholar]
  • 20.Miceli MC, Parnes JR. The roles of CD4 and CD8 in T cell activation. Semin Immunol. 1991;3(3):133–41 (PubMed PMID: 1909592). [PubMed] [Google Scholar]
  • 21.Guo X, Teschendorff AE. Guidelines on optimizing DNA methylation reference panels for cell-type deconvolution. Commun Biol. 2026;9(1):454. 10.1038/s42003-026-09745-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Petkovich DA, Podolskiy DI, Lobanov AV, Lee SG, Miller RA, Gladyshev VN. Using DNA methylation profiling to evaluate biological age and longevity interventions. Cell Metab. 2017;25(4):954-960.e6. 10.1016/j.cmet.2017.03.016PubMedPMID:28380383;PubMedCentralPMCID:PMC5578459. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Stubbs TM, Bonder MJ, Stark AK, Krueger F, Bolland D, Butcher G, et al. Multi-tissue DNA methylation age predictor in mouse. Genome Biol. 2017;18(1):68. 10.1186/s13059-017-1203-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Wang T, Tsui B, Kreisberg JF, Robertson NA, Gross AM, Yu MK, et al. Epigenetic aging signatures in mice livers are slowed by dwarfism, calorie restriction and rapamycin treatment. Genome Biol. 2017;18(1):57. 10.1186/s13059-017-1186-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Perez-Correa JF, Tharmapalan V, Geiger H, Wagner W. Epigenetic clocks for mice based on age-associated regions that are conserved between mouse strains and human. Front Cell Dev Biol. 2022;10:902857. 10.3389/fcell.2022.902857. (PubMed PMID: 35721486; PubMed Central PMCID: PMC9204067). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Thin KA, Cross A, Angsuwatcharakon P, Mutirangura A, Puttipanyalears C, Edwards SW. Changes in immune cell subtypes during ageing. Arch Gerontol Geriatr. 2024;122:105376. 10.1016/j.archger.2024.105376. [DOI] [PubMed] [Google Scholar]
  • 27.Nissen E, Reiner A, Liu S, Wallace RB, Molinaro AM, Salas LA, et al. Assessment of immune cell profiles among post-menopausal women in the Women’s Health Initiative using DNA methylation-based methods. Clin Epigenetics. 2023;15(1):69. 10.1186/s13148-023-01488-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Teo YV, Hinthorn SJ, Webb AE, Neretti N. Single-cell transcriptomics of peripheral blood in the aging mouse. Aging. 2023;15(1):6–20. 10.18632/aging.204471. (PubMed PMID: 36622281; PubMed Central PMCID: PMC9876630). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Zhang Z, Reynolds SR, Stolrow HG, Chen J, Christensen BC, Salas LA. Deciphering the role of immune cell composition in epigenetic age acceleration: insights from cell‐type deconvolution applied to human blood epigenetic clocks. Aging Cell. 2023;23(3):e14071. 10.1111/acel.14071. (PubMed PMID: 38146185; PubMed Central PMCID: PMC10928575). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Tong H, Guo X, Jacques M, Luo Q, Eynon N, Teschendorff AE. Cell-type specific epigenetic clocks to quantify biological age at cell-type resolution. Aging. 2024;16(22):13452–504. 10.18632/aging.206184. (PubMed PMID: 39760516; PubMed Central PMCID: PMC11723652). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Kim WY, Sharpless NE. Drug efficacy testing in mice. Curr Top Microbiol Immunol. 2012;355:19–38. 10.1007/82_2011_160. (PubMed PMID: 21823029; PubMed Central PMCID: PMC3732649.). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Li Y, Pan X, Roberts ML, Liu P, Kotchen TA, Cowley AW Jr, et al. Stability of global methylation profiles of whole blood and extracted DNA under different storage durations and conditions. Epigenomics. 2018;10(6):797–811. 10.2217/epi-2018-0025. (PubMed PMID: 29683333). [DOI] [PubMed] [Google Scholar]
  • 33.Unsworth A, Anderson R, Haynes N, Britt K. OMIP-032: Two multi-color immunophenotyping panels for assessing the innate and adaptive immune cells in the mouse mammary gland. Cytometry A. 2016;89(6):527–30. 10.1002/cyto.a.22867. [DOI] [PubMed] [Google Scholar]
  • 34.Aryee MJ, Jaffe AE, Corrada-Bravo H, Ladd-Acosta C, Feinberg AP, Hansen KD, et al. Minfi: A flexible and comprehensive Bioconductor package for the analysis of Infinium DNA methylation microarrays. Bioinformatics. 2014;30(10):1363–9. 10.1093/bioinformatics/btu049. (PubMed PMID: 24478339; PubMed Central PMCID: PMC4016708). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Herzog C. IlluminaMouseMethylationanno.12.v1.mm10. GitHub. [Internet]. 2022. Available from: https://github.com/chiaraherzog/IlluminaMouseMethylationanno.12.v1.mm10.
  • 36.Herzog C. IlluminaMouseMethylationmanifest. GitHub [Internet]. 2022. Available from: https://github.com/chiaraherzog/IlluminaMouseMethylationmanifest?tab=readme-ov-file.
  • 37.Ritchie ME, Phipson B, Wu D, Hu Y, Law CW, Shi W, et al. Limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43(7):e47. 10.1093/nar/gkv007. (PubMed PMID: 25605792; PubMed Central PMCID: PMC4402510). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Koestler DC, Jones MJ, Usset J, Christensen BC, Butler RA, Kobor MS, et al. Improving cell mixture deconvolution by identifying optimal DNA methylation libraries (IDOL). BMC Bioinformatics. 2016;17(1):120. 10.1186/s12859-016-0943-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Min JL, Hemani G, Davey Smith G, Relton C, Suderman M. Meffil: Efficient normalization and analysis of very large DNA methylation datasets. Bioinformatics. 2018;34(23):3983–9. 10.1093/bioinformatics/bty476. (PubMed PMID: 29931280; PubMed Central PMCID: PMC6247925). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Kuhn M, Wickham H. Tidymodels: A collection of packages for modeling and machine learning using tidyverse principles [Internet]. 2020. Available from: https://www.tidymodels.org.
  • 41.Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, et al. Gene Ontology: Tool for the unification of biology. Nat Genet. 2000;25(1):25–9. 10.1038/75556. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Thomas PD, Ebert D, Muruganujan A, Mushayahama T, Albou L, Mi H. PANTHER: Making genome‐scale phylogenetics accessible to all. Protein Sci. 2022;31(1):8–22. 10.1002/pro.4218. (PubMed PMID: 34717010; PubMed Central PMCID: PMC8740835). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Reynolds SR, Stolrow HG, Lee M, Pike SC, Kolling F, Channon JY, Mielcarz DW, Ward GA, Fields J, Salas LA, Christensen BC. DNA methylation-based immune cell profiling in mouse blood with methylation cytometry. Datasets. Gene Expression Omnibus. 2026. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE284254. [DOI] [PMC free article] [PubMed]
  • 44.Zhou W, Shen H, Laird PW. Mouse DNA methylation atlas using Infinium Mouse Methylation Beadchips. Datasets. Gene Expression Omnibus. 2022. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE184410.
  • 45.Widschwendter M, Barrett J. DNA methylation in mouse cancer model with ketogenic diet and mifepristone treatment. Datasets. Gene Expression Omnibus. 2025. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE236260.
  • 46.Reynolds SR, Stolrow HG, Salas LA, Christensen BC. 2026. Zenodo. 10.5281/zenodo.20126537.
  • 47.Reynolds SR, Stolrow HG, Salas LA, Christensen, BC. GitHub. 2026. https://github.com/SalasLab/MouseRS.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

13059_2026_4121_MOESM1_ESM.xlsx (12.9KB, xlsx)

Additional file 1: Tables S1-S3. Table S1. Flow cytometry markers in whole blood samples. Table S2. Purified cell type and whole blood sample characteristics. Table S3. Characteristics of the GSE184410 dataset used to construct the IDOL-optimized reference library.

13059_2026_4121_MOESM2_ESM.docx (379.9KB, docx)

Additional file 2: Figs. S1-S3. Performance of cell-type deconvolution as a function of CpG reference library size. Fig. S2. Distribution of gene features across CpG genomic contexts. Fig. S3. Comparison of performance across reference library sizes when trained on MouseRS and GSE184410 datasets.

13059_2026_4121_MOESM3_ESM.docx (5.7MB, docx)

Additional file 3: Point-by-point response to reviewers.

Data Availability Statement

The data generated and used in this study are publicly available on Gene Expression Omnibus (GEO) with the accession number GSE284254 [43]. Additional datasets used in this work are available on GEO under accession numbers GSE184410 [44] and GSE236260 [45]. All relevant code for this manuscript has been deposited in Zenodo [46] (https://doi.org/10.5281/zenodo.20126537) and is also available on GitHub [47] (https://github.com/SalasLab/MouseRS), both released under the GNU General Public License version 3.


Articles from Genome Biology are provided here courtesy of BMC

RESOURCES