Abstract
Sampling lipids from the skin surface via noninvasive sebum collection is painless, efficient, and has already demonstrated clinical potential in preliminary biomarker research. Replacing blood-based sampling with sebum sampling will allow these benefits to be leveraged as well as enable researchers to better include protected populations, such as children, who tend to be excluded from biomarker research. However, for researchers to obtain meaningful lipidomic data sets in routine studies, more information is needed regarding sebum sampling feasibility and variability. Herein, we assess the feasibility of forehead sebum sampling from child participants. Furthermore, using high resolution mass spectrometry in combination with both supervised and unsupervised classification, we determine the variability of the lipid profile according to the age, sex, and biological relatedness of 98 participants, ages four to 73. Outcomes of this study indicate that age plays a notable role in the sebum lipid profile, while both biological sex and biological relatedness have little to no impact on these lipids. Importantly, sebum sampling from the forehead proved to be a successful sampling approach for children, as children appeared enthusiastic about participation and samples were collected rapidly and painlessly. Together, these results equip researchers with information necessary for designing effective, sebum-based biomarker studies, and they demonstrate that sebum collection is an ideal sampling approach for biomarker studies in children.
Graphical Abstract

Introduction
Sebum, the oily residue that coats the exterior of the human body, may be a useful sample source for biomarker studies. Sebum consists of many different lipids (wax esters, glycerides, squalene, fatty acids, cholesterol, and cholesterol esters) which are secreted directly onto the skin surface by the sebaceous gland.1–3 Sebum sampling is a noninvasive sampling approach that allows for the efficient, painless collection of lipids directly from the skin surface1–3 and thus may be an ideal alternative to traditional blood-based sampling methods. Sebum sampling would be particularly ideal for biomarker exploration studies, as preliminary research has already demonstrated that skin surface lipids correlate with various disease states.4–11 Moreover, sebum sampling could enable researchers to better include populations that tend to be excluded from biomarker studies. Current biomarker research, which relies predominantly on blood-based sampling protocols, often excludes protected populations, such as children, and evidence suggests that this may be due to difficulties in obtaining enough child patient samples.12,13 Alternatively, sebum sampling allows for many samples to be collected quickly and noninvasively, without any discomfort to the patient, making this approach particularly optimal for pediatrics.
Before the benefits of sebum sampling can be leveraged for large-scale, routine studies in children, more information is needed regarding sample variability and collection feasibility from this population for ESI-MS analyses. Since variability between like samples (i.e. healthy controls) can make it difficult for machine learning algorithms to identify subtle distinctions between unlike samples (i.e. disease vs healthy controls), it is important that sebum variability be well understood and accounted for in the biomarker study design. Specifically, it is not yet clear if demographics, such as participant age,14–29 sex,11,15,24,31–34 or other biological factors, impact the sebum lipid profile.
Prior comparative studies of sebum between adults and children have leveraged techniques like FT-IR18,21 and GC22,23,25 or TLC16. These methods can provide general information about lipid classes that change, but they are not the go-to method for modern lipidomics biomarker studies, which heavily leverage ESI-MS, a method that can detect thousands of compounds per sample and can be used to identify the exact compounds changing in abundance.. Herein, we apply ESI-MS to lipid extracts from adults and children, affording insight into the molecular changes between these groups for the first time.
The studies conducted herein explore how other types of variability, for example sex and biological relatedness, may impact the complex lipid profiles detected by ESI-MS in datasets that include children. While some researchers have identified sex-based correlations between sebum samples,11,15,24,31–34 others have reported a lack thereof.25–27,35–38 (We note that the large diversity in techniques and sample collection methods used for these studies likely contributes to their inconsistent findings.) The study described herein affords the opportunity to further explore sex-based differences in a large and diverse data set of sebum samples analyzed by ESI-MS. Additionally, another biologically relevant variable that could affect the sebum lipid profile is the biological relatedness of participants (ex. children and their parents). This variable has not been assessed at all. If genetics and/or diet plays a notable role in an individual’s sebum lipid profile, one would expect that samples from related individuals would appear more similar to each other. Thus, sebum sample differences related to participant age, sex, and biological relatedness could provide valuable insight into biological differences that are measurable between these groups.
Previous skin surface studies have sampled sebum from small numbers of child participants,14–16,18–23,25,26,39 but most of these sebum-based protocols are not ideal for children, as they expose participants to chemicals,16,20,22,23,25,26 require participants to dedicate one hour or more to sampling,19,20,26 or employ collection methods applicable to less informative analyses than the gold-standard approach to lipidomics, mass spectrometry.14–16,18–21,23,39 As such, it is important that the feasibility of sebum sampling (particularly methods that cause no discomfort and are ideal for ESI-MS analysis downstream), be adequately developed and leveraged.
Herein, using an adaptation of a previously described sebum collection and lipid analysis workflow,40 we advance sebum sampling lipidomics by assessing sebum profile variability according to participant age, participant sex, and the biological relatedness of participants. Additionally, when lipidomic differences are evident between sample groups, we identify the specific m/z values (and the lipid identities) that contribute most to group discrimination. Moreover, we evaluate the feasibility of this sebum sampling method for pediatrics and comment on its capacity for large-scale biomarker studies that include children. These findings will allow researchers to make informed decisions when choosing the best practices for biomarker exploration workflows, thus enabling the acquisition of meaningful lipidomic data sets via sampling that is painless, efficient, and ideal for all ages.
Experimental
Sebum Sampling
A total of 98 participants (ages 4–73) donated sebum samples for this study, the majority of which consisted of parents and their biological children. Specific participant demographics are shown in Table S1 of the Supporting Information. Forehead sebum samples were collected in compliance with The University of Kansas’ Human Research Protection Program for Human Subjects using a noninvasive sampling approach adapted from previous studies.35,40 All samples were collected on a single day over the span of three hours at The University of Kansas’ Carnival of Chemistry, an annual event taking place in November. Parents, guardians, and children gave informed consent and answered a series of questions regarding age, biological sex, and biological relatedness to one another. Then, a researcher placed two pieces of clean aluminum foil (approximately 1 cm × 2 cm each) onto the forehead of each participant. The forehead was chosen as the sample site to maximize lipid collection and allow for a consistent sampling surface area. Foil pieces were placed one at a time, near the center of the forehead, in adjacent locations, and moderate pressure was applied for five seconds each. This was accomplished using a kim wipe wrapped around a gloved finger to avoid excessive contamination of the foil substrate by glove-originating compounds. Foil pieces were then removed and placed on a kim wipe (collection side up) inside a covered petri dish, and the dish was labeled according to an anonymous code that corresponded to that of the participant’s consent form. Samples were then immediately transported to a nearby lab. For each new participant, this procedure was repeated with a new glove and kim wipe. Each child who participated in the study, as well as those who chose to decline participation, received a small prize. Two forehead sebum samples were collected from each of the 98 participants, resulting in a total of 196 samples.
Sample Preparation
Once the foil samples arrived at the nearby lab, each foil sample was folded in half (collection side out) using acetone-rinsed tweezers, placed in a glass vial, labeled accordingly, and immediately stored at approximately −20 °C. After all samples had been collected (approximately three hours), samples were promptly transported to a nearby lab building and remained under cold storage (−20 °C) until further sample preparation and MS analysis (approximately three weeks).
Prior to MS analysis, the lipid content from a single foil sample collected from each of the 98 participants was extracted in organic solvent, desalted, and diluted following a previously de-scribed protocol.35,40 This extraction procedure removes the salts that are naturally present in sweat that would be collected in the sample. Briefly, 400 μL of dichloromethane was added to each vial, samples were sonicated for 10 minutes, and foil pieces were then discarded. Distilled water (200 μL) was added to each sample vial for subsequent liquid-liquid extraction. Samples were sonicated for 10 minutes and then left to stand for 20 minutes to ensure sufficient aqueous/organic layer separation. The aqueous component was discarded, and the remaining organic component was stored at −20 °C until MS analysis. Before MS analysis, samples were brought to room temperature, vortexed, and diluted by a factor of 20 with 1:4 dichloromethane/methanol.
Flow Injection ESI-MS
The injection order of the sebum samples (98 in total) was randomized with respect to age, sex, and family relatedness. High resolution flow injection electrospray ionization mass spectrometry (ESI-MS) was accomplished using a Waters Acquity UPLC (Milford, MA) coupled to an Orbitrap Fusion Tribrid mass spectrometer (Thermo Fisher Scientific, San Jose, CA). Flow injection and ESI-MS parameters followed that of prior experiments.35,40 Briefly, 5 μL of sample was injected into the mass spectrometer at 25 μL/min. The elution was isocratic for a two-minute duration and consisted of 50% Mobile Phase A (50:50 methanol/water; 5 mM ammonium acetate) and 50% Mobile Phase B (20:79:1 acetone/2-propanol/water; 5 mM ammonium acetate). All MS spectra were acquired in positive ion mode, with each scan covering the range of m/z 250 – 1300, at a resolution of 120,000 for m/z 200 (XCalibur V. 4.6.67.17). AGC and maximum injection time were 4 ×105 and 246 ms, respectively. Spray voltage, capillary temperature, sheath gas, and auxiliary gas were as follows: 2.8 kV, 275 °C, 10 units, and 6 units.
Data Processing
Raw MS data collected from 98 sebum samples were input into a data matrix, processed, and then assessed via supervised (XGBoost)41 and unsupervised (PCA) classification, as explained previously.35 To construct the data matrix, the raw MS spectra were first converted to MS1 format via Raw Convertor (version 1.2.0.1).42 Then, using an in-house R script (R version 4.1.0),43 data from the MS1 files (across retention time 0.9 – 1.1 min) were extracted into a data matrix where each row represented an m/z “bin” and each column represented a sample. The m/z “bins,” or features, corresponded to 0.01 Da intervals; each mass spectrum (ranging 250 – 1300 m/z) was divided into 0.01 Da increments, and the peak intensities within each 0.01 Da increment were summed and input into the matrix. The resulting feature matrix of m/z bins (rows) and samples (columns) enabled the MS data to be compatible with subsequent machine learning applications.
Prior to machine learning, the feature matrix was filtered and normalized. To eliminate unused bins, rows (m/z bins) were only kept in the matrix if at least 1% of the samples contained a nonzero number. Then, sample normalization was performed by dividing each value in the matrix by the sum of intensities from the corresponding column, thereby allowing each matrix value to represent a fractional portion of the corresponding sample’s total ion counts. This step accounted for the fact that different participants may produce different overall abundances of lipids and thus enabled the assessment of different types of lipids between individuals. Batch normalization was then performed to account for MS batch effects, and this was accomplished by the addition of 1 × 10−10 to each matrix value (to eliminate zero values), log2 transformation, and then the removeBatchEffect function via R package “limma”.44 The vector generated for removeBatchEffect corresponded to sample injection order and consisted of a four-variable vector, in which the first 24 samples injected were considered “batch 1,” the next 25 samples were “batch 2,” the next 24 samples were “batch 3,” and the last 25 samples were “batch 4.” After normalization, noise was reduced by keeping only the m/z bins with approximately the top 10% of median intensity values, resulting in a final feature matrix of 3800 rows (m/z bins) and 98 columns (samples). The Supporting Information shows the raw data matrix (Table S2) and the normalized feature matrix used in each analysis (Table S3). Prior to machine learning applications, all feature matrices were transposed.
Data Analysis
Sebum samples were assessed according to participant age, biological sex, and biological relatives. These analyses were accomplished via supervised classification using XGBoost (extreme gradient boosting, R package “xgboost”) and unsupervised classification using PCA (principal component analysis, R packages “factoextra” and “ggplot2”).45,46 The XGBoost hyperparameters were not optimized prior to classification, and these included: the booster (gbtree), objective (binary:logistic), eta (0.3), gamma (0), max_depth (6), min_child_weight (1), subsample (1), colsample_bytree (1), and nrounds (50). To prevent bias, leave-one-out cross-validation (LOOCV) was employed so that the sample being tested was removed from the training set. Since XGBoost hyperparameters were left unoptimized, results from classification with LOOCV represent the accuracy that would be otherwise obtainable from a larger study utilizing a separate test data set.47 To assess XGBoost classification outcomes, the classification accuracy (the percentage of correctly classified samples) and AUC value resulting from the corresponding ROC curve (package “proc”),48 were calculated for each analysis. When classification outcomes indicated notable differences between sample groups, the XGBoost importance matrix was extracted, and the m/z bins most useful in discriminating between the groups were identified. Then, the m/z values that corresponded to the m/z bins with the highest gain values were identified in the raw MS data. Using these m/z values, lipid class identification was accomplished using the Lipid Maps Structure Database (LMSD)49 with a mass accuracy threshold of < 4 ppm.
Results and Discussion
A total of 98 forehead sebum samples were collected from child and adult participants and analyzed via high resolution mass spectrometry in combination with supervised (XGBoost) and unsupervised (PCA) classification, as depicted in Figure 1. Sebum lipid differences corresponding to participant age, participant sex, and the biological relatedness of participants was assessed, and the feasibility of forehead sebum sampling for pediatric participants was evaluated.
Figure 1.

Workflow used for forehead sebum sampling, sample preparation, and ESI-MS data analysis.
1. Sebum sampling feasibility for pediatric studies
Before sebum sampling can be implemented into large-scale biomarker studies, the feasibility of this sampling approach for protected populations should be assessed. As such, the feasibility of forehead sebum sampling from children was evaluated herein. Samples were collected as part of the Carnival of Chemistry, a free, annual event hosted by the University of Kansas’ Department of Chemistry in which faculty members and college students demonstrate and explain scientific concepts to families in the community. Sample collection for this project was conducted at a designated table alongside other scientific demonstrations. Children and their guardians gave informed consent, and the study was explained in age-appropriate language so that all parties could comprehend.
This method of sebum sampling proved to be an efficient, painless sampling approach ideal for child patients. Children of all ages appeared excited to participate and ask questions to the researchers, and most guardians were eager to allow their children to participate. The sample collection, which consisted of placing two small pieces of clean aluminum foil onto the forehead for five seconds each, was painless and rapid. There was a consistent line of families eager to participate in this project throughout the entire duration of the event. The event lasted three hours, and in total 196 samples were collected from 98 participants across 41 different families. A histogram of the participants, grouped by age, is shown in Supplemental Figure S1.
Though no significant issues were encountered throughout this sample collection, it was concluded that more samples could have been collected during this three-hour time span had the researchers anticipated the overwhelming enthusiasm of community members, especially children, and thus allocated more researchers to tasks such as sample labeling, collection, and storage. In the future, more samples could be collected with the same sampling protocol but an increased number of researchers, as participants were not the limiting reagent in this study. Nonetheless, this approach to collecting forehead sebum samples from children and their families proved to be efficient for researchers (five seconds per sample collection) and painless for patients. Outcomes suggest that, when applicable, sebum sampling would be an efficient, non-invasive alternative to traditional blood draw sampling methods, and this approach would be particularly ideal for biomarker studies inclusive of child patients.
2. Differences between children and adults
The purpose of the first experiment was to assess whether children and adult sebum samples are measurably different. If children and adults produce different sebum lipid profiles, it is possible they may also produce different sebum biomarkers of disease, information that will be useful to biomarker researchers. Furthermore, variability among healthy participants of different ages could make it difficult for machine learning algorithms to identify small differences between healthy and disease patient groups, so age-based sebum variability will be an important factor for researchers to be aware of and thus account for in the biomarker study design. A total of 54 children (ages 4–15) and 44 adults (ages 29–73) donated forehead sebum samples for this study. Using PCA and XGBoost classification, age-based differences were assessed across this data set of 98 sebum samples and then re-assessed after removing samples belonging to older children (ages 11–15).
The PCA plot in Figure 2A shows 98 data points, each representative of a different participant’s sebum sample. The PCA data clusters of children and adult participants are readily discernable, suggesting child and adult sebum lipid profiles are measurably distinct. XGBoost classification outcomes also indicate age-based differences are present (82% accuracy, AUC = 0.84, Fig. 2B). In previous studies50 using a similar sampling protocol, sebum was collected from the same individuals at several time points throughout the day over the course of a year. Those analyses showed that the timing of collection contributed negligibly to the overall lipidomic variability compared with the biological differences between individuals. As a result, the PCA and XGBoost outcomes observed here are unlikely to reflect sampling time effects and instead represent age-related variation in sebum composition. Since the literature suggests child sebum production changes around the age of puberty,14,18,23,25 it was hypothesized that children of puberty age may produce lipids that are different from younger children, and that this may impact the child vs. adult classification results.
Figure 2.

Unsupervised and supervised classification results for forehead sebum samples collected from 54 children (ages 4–15) and 44 adults (ages 29–73). (A) Principal component analysis (PCA) of 98 forehead sebum samples; concentration ellipses are produced using RStudio, packages “factoextra” and “ggplot2.” (B) ROC curve generated from the XGBoost classification results of the same 98 samples from panel A (accuracy = 82%, AUC = 0.836).
To determine whether the samples collected from older children have a notable impact on the classification outcomes of children vs adults, the five samples collected from child participants ages 11–15 were removed from the data set, and the remaining 93 samples were reclassified. Figure 3A shows the PCA plot from this reclassification, and child and adult sample groups remain readily discernable from one another. Upon reclassification, the XGBoost child vs. adult classification accuracy increased from 82% to 88%, as well as the AUC value from 0.84 to 0.92 (Fig. 3B). Moreover, none of the misclassified samples belong to the youngest children, those ages 4–7. Since removing the samples collected from older children improves the classification results, and none of the youngest participant samples are misclassified, it is likely that the sebum samples of young children are more similar to one another than they are to older children, and thus age plays a notable role in the sebum lipid profile.
Figure 3.

Unsupervised and supervised classification results for forehead sebum samples collected from 49 children (ages 4–10) and 44 adults (ages 29–73). (A) Principal component analysis (PCA) of 93 forehead sebum samples; concentration ellipses are produced using RStudio, packages “factoextra” and “ggplot2.” (B) ROC curve generated from the XGBoost classification results of the same 93 samples from panel A (accuracy = 88%, AUC = 0.916).
To further assess the lipidomic differences between children and adults, the XGBoost importance matrix was generated from the classification of 49 child (ages 4–10) and 44 adult (ages 29–73) samples. MS peaks corresponding to m/z 549.489 and m/z 551.505 represented the most important features in XGBoost age-based discrimination (i.e. corresponding to the m/z bins yielding the highest gain values). Lipid Maps Structure Database search results indicate both MS peaks are likely diacylglycerols. Both forms were substantially less abundant in children compared to adults.
While several different isomeric diacylglycerols match this mass, the two most abundant in humans are dehydrated forms of DG 16.0/16.0 (m/z 551.504) and DG 16.0/16.1 (m/z 549.489). The former comprises about a third of the fatty acids in lipid membranes, and the latter is also an abundant species in membranes and throughout the body. The masses observed in the spectra, and reported above, correspond to the dehydrated forms of these lipids; that form is generated in high abundance when these lipids are ionized under the conditions used herein, as shown in Supplemental Figure S2, which includes ESI-MS data of this compound, which is commercially available. The composition of DG 16.0/16.0 in the fingerprint sample was further verified by matching its MS/MS data to that of the authentic standard. The MS/MS data for DG 16.0/16.1 was also acquired, and the data are also consistent with the assignment. See Supplemental Figures S2 and S3.
We are unaware of any previous reports showing that children have lower concentrations of these lipids, however the biological observation could perhaps be explained by the fact that children between the ages of 4 and 10 are growing rapidly, and these lipids, are the primary building blocks of both PC and PE, which are key components of lipid membranes throughout the body and in the brain. In retrospect, it may be logical for the body to shuttle these lipids into pathways that use them for growth and development, leaving relatively fewer of them detectable in sebum. This example shows how new hypothesis about biological processes could be built from these data and tested in future experiments.
To explore the relative impact of biological and technical sources of variability, pooled data were compared (Figure S4). A pooled sample comprising of the first half of the experimental run showed no major spectral differences compared to a pool from the second half, confirming stable instrument performance throughout the analysis (Figure S4A). This observation is consistent with previous evaluations of this analytical method, which demonstrated < 5% relative standard deviation across replicate injections.35 In contrast, a comparison of the pooled spectra from child and adult samples revealed clear differences (Figure S4B). When looking at the m/z 520–570 mass range, these differences become more apparent, with the features in child samples exhibiting higher signal intensities than those in adult samples (Figure S4C). Together, these findings indicate that age-based biological variability is the dominant contributor to the observed differences, whereas technical variability is minimal.
3. Differences between children of different age groups
Since the results presented in the previous section suggested that sebum lipids differ between children of different age groups, the goal of this experiment was to further assess this hypothesis— an endeavor that could aid in designing biomarker studies focused on childhood diseases. Therefore, in this analysis, the 54 child samples were classified according to two age groups, and lipidomic differences were assessed. The “younger” children were considered those ages 4–9 (42 samples) and the “older” children were considered those ages 10–15 (12 samples).
The PCA plot in Figure 4A shows discernable sample clusters corresponding to child age groups, and the supervised classification results also suggest these groups are measurably different (81% accuracy, AUC = 0.758, Fig. 3B). Together, these results imply that younger children produce sebum lipids that are different from those of older children, and thus participant age is a variable that researchers should be aware of when designing sebum-based biomarker studies for pediatric applications. Furthermore, the XGBoost importance matrix reveals that the lipids contributing most to discrimination between these two age groups are those corresponding to m/z 425.274 and m/z 403.289. Consultation with the LMSD suggests both these MS peaks represent glycerophospholipids; m/z 425.274 corresponds to PI(35:2), and m/z 403.289 corresponds to PA(44:6). Among the participants tested here, glycerophospholipids differ notably between children ages 4–9 and 10–15, and thus researchers should take this information into account when assessing sebum lipid biomarkers of childhood diseases.
Figure 4.

Unsupervised and supervised classification results for forehead sebum samples collected from 42 young children (ages 4–9) and 12 older children (ages 10–15). (A) Principal component analysis (PCA) of 54 forehead sebum samples; concentration ellipses are produced using RStudio, packages “factoextra” and “ggplot2.” (B) ROC curve generated from the XGBoost classification results of the same 54 samples from panel A (accuracy = 81%, AUC = 0.758).
4. A lack in differences between males and females
Since intragroup variability can make it more difficult for machine learning algorithms to identify intergroup differences, potential sex-based sebum differences will be important for biomarker researchers to be aware of. To assess sex-based sebum differences, supervised and unsupervised male-female classification was performed on adult samples (41 total) and child samples (54 total) independently.
Figure 5A shows the PCA plot of 14 male and 27 female sebum samples collected from 41 adult participants (ages 29–73). Note: only 41 of the 44 total adult samples were used in this analysis as not every adult participant indicated their biological sex on the consent form. The PCA plot shows overlap between male and female sample groups, and 61% of the samples are correctly classified using XGBoost (AUC = 0.48, Fig. 5B), suggesting adult male and female sebum lipids are not measurably different from one another.
Figure 5.

Unsupervised and supervised classification results for forehead sebum samples collected from 14 adult males and 27 adult females. (A) Principal component analysis (PCA) of 41 forehead sebum samples; concentration ellipses are produced using RStudio, packages “factoextra” and “ggplot2.” (B) ROC curve generated from the XGBoost classification results of the same 41 samples from panel A (accuracy = 61%, AUC = 0.476).
A similar outcome was obtained from the sebum samples collected from child participants. Figure 6A shows the PCA plot of 28 male and 26 female sebum samples collected from 54 child participants (ages 4–15). As observed in the adult data set, the PCA plot shows an overlap of male and female sample groups, and sex-based differences are not effectively classified by XGBoost (46% accuracy, AUC = 0.64, Fig. 6B), suggesting the male and female sebum samples of child participants are not measurably different from one another. Together, these results reveal that biological sex does not notably impact the types of lipids produced in forehead sebum, regardless of participant age, and thus is not a variable researchers need to account for when designing sebum-based biomarker studies. Other studies have arrived at similar conclusions.25–27,35,37
Figure 6.

Unsupervised and supervised classification results for forehead sebum samples collected from 28 male children and 26 female children. (A) Principal component analysis (PCA) of 54 forehead sebum samples; concentration ellipses are produced using RStudio, packages “factoextra” and “ggplot2.” (B) ROC curve generated from the XGBoost classification results of the same 54 samples from panel A (accuracy = 46%, AUC = 0.637).
5. A lack in similarities among biological relatives
The goal of this experiment was to assess the sebum lipid profiles of biologically related participants, a variable which, to our knowledge, has yet to be investigated. In this experiment, forehead sebum samples collected from biologically related participants (i.e. children and their biologically related siblings, parents, and grandparents) were assessed using unsupervised classification. PCA results (shown in Fig. 7) revealed a lack in discernable family-related clusters, suggesting that the sebum lipids produced by biologically related participants are not any more similar to one another than they are to other, non-related participants. As such, clinical researchers designing sebum-based biomarker studies can expect the genetic relatedness of participants to have minimal to no impact on the resulting sebum lipid profile. This outcome also suggests that personal care products, for example, shampoo, that would be presumably shared within a household, do not dominate the clustering of the samples. We note, however, that others have determined that these compounds are detectable on the skin surface.51
Figure 7.

Principal component analysis (PCA) of 98 forehead sebum samples collected from children and their biological relatives. “Families” refer to biologically related siblings, parents, and/or grandparents. (A) PCA plot of 98 samples, with three arbitrarily chosen families, denoted as families 1–3, color-coded according to family unit. (B) PCA plot of 98 samples, with three arbitrarily chosen families, denoted as families 4–6, color-coded according to family unit. (C) PCA plot of 98 samples, with three arbitrarily chosen families, denoted as families 7–9, color-coded according to family unit.
Conclusions
The outcomes of this study reveal important information that will enable researchers to choose best practices when designing sebum-based biomarker studies. In turn, this will allow researchers to leverage the benefits of sebum sampling: efficient, non-invasive sample collection that is ideal for protected populations. Herein, we have demonstrated that this sebum sampling method is ideal for biomarker studies inclusive of child participants. Children were enthusiastic about participation; sample collection took only five seconds per sample, and participants experienced no pain or discomfort. In addition, we have assessed the extent to which participant demographics (age, sex, and biological relatedness) impact the sebum lipid profile. The results suggest that age has a notable influence on the sebum lipid profile. Sebum samples collected from children appear to be measurably different from those collected from adults, with analytes corresponding to m/z 549.489 and m/z 551.505 contributing most to age differentiation among these participants. Furthermore, sebum samples of young children (ages 4–9) appear to be measurably different from those of older children (ages 10–15), with analytes corresponding to m/z 425.274 and m/z 403.289 contributing most to age differentiation among these participants. However, the biological sex and biological relatedness of participants does not appear to measurably impact the sebum lipid profile. The results reported here equip the metabolomics community with information regarding sebum sampling variability and feasibility that will allow researchers to obtain meaningful lipidomic data sets from sebum-based biomarker studies that are inclusive of all ages.
Supplementary Material
Figure S1: Histograms of participants by age.
Figure S2. Mass spectrometry data of lipid standard.
Figure S3. Mass spectrometry data of the key lipids differentiating children and adults.
Figure S4. Butterfly plots comparing samples pooled by injection order or age.
Table S1: Participant demographics (xlsx)
Table S3: Normalized feature matrices (xlsx)
Table S2: Raw data matrix (xlsx)
The Supporting Information is available free of charge on the ACS Publications website.
ACKNOWLEDGMENT
The authors would like to thank several volunteers who assisted with the Carnival of Chemistry: Braysen Miller, Ian Isom, Jane March, Jason Applegate, Hannah Chern, Mike Kim, and Miyuru De Silva.
MI would like to thank the Madison and Lila Self Graduate Fellowship for funding. HD acknowledges support from NIH grant R35GM156211.
Craiyon was used to assist in generating portions of Fig. 1. ChatGPT was used to assist in generating some of the code. The authors take full responsibility for the accuracy of the code.
Footnotes
CONFLICT OF INTEREST
The authors declare no competing financial interest.
DATA AVAILABILITY
The raw MS data is included in the Supporting Information.
REFERENCES
- 1.Isom M; Desaire H Skin Surface Sebum Analysis by ESI-MS. Biomolecules. 2024, 14 (7), 790. DOI: 10.3390/biom14070790. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Géhin C; Tokarska J; Fowler SJ; Barran PE; Trivedi DK No Skin off Your Back: The Sampling and Extraction of Sebum for Metabolomics. Metabolomics. 2023, 19, 21. DOI: 10.1007/s11306-023-01982-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Elpa DP; Chiu HY; Wu SP; Urban PL Skin Metabolomics. Trends Endocrinol. Metab 2021, 32, 66–75. DOI: 10.1016/j.tem.2020.11.009. [DOI] [PubMed] [Google Scholar]
- 4.Sinclair E; Trivedi DK; Sarkar D; Walton-Doyle C; Milne J; Kunath T; Rijs AM; de Bie RMA; Goodacre R; Silverdale M; Barran P Metabolomics of sebum reveals lipid dysregulation in Parkinson’s disease. Nat Commun. 2021, 12 (1), 1592. DOI: 10.1038/s41467-021-21669-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Briganti S; Truglio M; Angiolillo A; Lombardo S; Leccese D; Camera E; Picardo M; Di Costanzo A Application of Sebum Lipidomics to Biomarkers Discovery in Neurodegenerative Diseases. Metabolites. 2021, 11 (12), 819. DOI: 10.3390/metabo11120819. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Sarkar D; Sinclair E; Lim SH; Walton-Doyle C; Jafri K; Milne J; Vissers JPC; Richardson K; Trivedi DK; Silverdale M; Barran P Paper Spray Ionization Ion Mobility Mass Spectrometry of Sebum Classifies Biomarker Classes for the Diagnosis of Parkinson’s Disease. JACS Au. 2022, 2 (9), 2013–2022. DOI: 10.1021/jacsau.2c00300. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Spick M; Lewis HM; Frampas CF; Longman K; Costa C; Stewart A; Dunn-Walters D; Greener D; Evetts G; Wilde MJ; Sinclair E; Barran PE Skene DJ; Bailey MJ An integrated analysis and comparison of serum, saliva and sebum for COVID-19 metabolomics. Sci Rep. 2022, 12 (1), 11867. DOI: 10.1038/s41598-022-16123-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Spick M; Longman K; Frampas C; Lewis H; Costa C; Walters DD; Stewart A; Wilde M; Greener D; Evetts G; Trivedi D; Barran P; Pitt A; Bailey M Changes to the sebum lipidome upon COVID-19 infection observed via rapid sampling from the skin. EClinicalMedicine. 2021, 33, 100786. DOI: 10.1016/j.eclinm.2021.100786. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Yin H; Qiu Z; Zhu R; Wang S; Gu C; Yao X; Li W Dysregulated lipidome of sebum in patients with atopic dermatitis. Allergy. 2023, 78 (6), 1524–1537. DOI: 10.1111/all.15569. [DOI] [PubMed] [Google Scholar]
- 10.Shetage SS; Traynor MJ; Brown MB; Galliford TM; Chilcott RP Application of sebomics for the analysis of residual skin surface components to detect potential biomarkers of type-1 diabetes mellitus. Sci Rep. 2017, 7 (1), 8999. DOI: 10.1038/s41598-017-09014-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Cavallo A; Camera E; Bottillo G; Maiellaro M; Truglio M; Marini F; Chavagnac-Bonneville M; Fauger A; Perrier E; Pigliacelli F; Picardo M; Cristaudo A; Mariano M Biosignatures of defective sebaceous gland activity in sebum-rich and sebum-poor skin areas in adult atopic dermatitis. Exp Dermatol. 2024, 33, e15066. DOI: 10.1111/exd.15066. [DOI] [PubMed] [Google Scholar]
- 12.Shores DR; Everett AD Children as Biomarker Orphans: Progress in the Field of Pediatric Biomarkers. Journal of Pediatrics. 2018, 193, 14–20.e31. DOI: 10.1016/j.jpeds.2017.08.077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Savage WJ; Everett AD Biomarkers in pediatrics: children as biomarker orphans. Proteomics Clin Appl. 2010, 4 (12), 915–921. DOI: 10.1002/prca.201000062. [DOI] [PubMed] [Google Scholar]
- 14.Cotterill JA; Cunliffe WJ; Williamson B; Bulusu L Age and sex variation in skin surface lipid composition and sebum excretion rate. Br J Dermatol. 1972, 87 (4), 333–340. DOI: 10.1111/j.1365-2133.1972.tb07419.x. [DOI] [PubMed] [Google Scholar]
- 15.Man MQ; Xin SJ; Song SP; Cho SY; Zhang XJ; Tu CX; Feingold KR; Elias PM Variation of skin surface pH, sebum content and stratum corneum hydration with age and gender in a large Chinese population. Skin Pharmacol Physiol. 2009, 22 (4), 190–199. DOI: 10.1159/000231524. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Ramasastry P; Downing DT; Pochi PE; Strauss JS Chemical composition of human skin surface lipids from birth to puberty. J Invest Dermatol. 1970, 54 (2), 139–144. DOI: 10.1111/1523-1747.ep12257164. [DOI] [PubMed] [Google Scholar]
- 17.Starr NJ; Johnson DJ; Wibawa J; Marlow I; Bell M; Barrett DA; Scurr DJ Age-Related Changes to Human Stratum Corneum Lipids Detected Using Time-of-Flight Secondary Ion Mass Spectrometry Following in Vivo Sampling. Anal Chem. 2016, 88 (8), 4400–4408. DOI: 10.1021/acs.analchem.5b04872. [DOI] [PubMed] [Google Scholar]
- 18.Hemmila A; McGill J; Ritter D Fourier transform infrared reflectance spectra of latent fingerprints: a biometric gauge for the age of an individual. J Forensic Sci. 2008, 53 (2), 369–376. DOI: 10.1111/j.1556-4029.2007.00649.x. [DOI] [PubMed] [Google Scholar]
- 19.Pochi PE; Strauss JS; Downing DT Age-related changes in sebaceous gland activity. J Invest Dermatol. 1979, 73 (1), 108–111. DOI: 10.1111/1523-1747.ep12532792. [DOI] [PubMed] [Google Scholar]
- 20.Jacobsen E; Billings JK; Frantz RA; Kinney CK; Stewart ME; Downing DT Age-Related Changes in Sebaceous Wax Ester Secretion Rates in Men and Women. J Invest Dermatol. 1985, 85 (5), 483–485. DOI: 10.1111/1523-1747.ep12277224. [DOI] [PubMed] [Google Scholar]
- 21.Antoine KM; Mortazavi S; Miller AD; Miller LM Chemical differences are observed in children’s versus adults’ latent fingerprints as a function of time. J Forensic Sci. 2010, 55 (2), 513–518. DOI: 10.1111/j.1556-4029.2009.01262.x. [DOI] [PubMed] [Google Scholar]
- 22.Buchanan MV; Asano K; Bohanon A Chemical characterization of fingerprints from adults and children. Proceedings of the SPIE. 1997, 2941, 89–95. [Google Scholar]
- 23.Sansone-Bazzano G; Cummings B; Seeler AK; Reisner RM Differences in the lipid constituents of sebum from pre-pubertal and pubertal subjects. Br J Dermatol. 1980, 103 (2), 131–137. DOI: 10.1111/j.1365-2133.1980.tb06581.x. [DOI] [PubMed] [Google Scholar]
- 24.Zhou Z; Zare RN Personal Information from Latent Fingerprints Using Desorption Electrospray Ionization Mass Spectrometry and Machine Learning. Anal. Chem 2017, 89, 2, 1369–1372. DOI: 10.1021/acs.analchem.6b04498. [DOI] [PubMed] [Google Scholar]
- 25.Akutsu N; Ooguri M; Onodera T; Kobayashi Y; Katsuyama M; Kunizawa N; Hirao T; Hosoi J; Masuda Y; Yoshida S; Takahashi M; Tsuchiya T; Tagami H Functional characteristics of the skin surface of children approaching puberty: age and seasonal influences. Acta Derm Venereol. 2009, 89 (1), 21–27. DOI: 10.2340/00015555-0548. [DOI] [PubMed] [Google Scholar]
- 26.Shetage SS; Traynor MJ; Brown MB; Raji M; Graham-Kalio D; Chilcott RP Effect of ethnicity, gender and age on the amount and composition of residual skin surface components derived from sebum, sweat and epidermal lipids. Skin Res Technol. 2014, 20 (1), 97–107. DOI: 10.1111/srt.12091. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Wilhelm KP; Cua AB; Maibach HI Skin aging. Effect on transepidermal water loss, stratum corneum hydration, skin surface pH, and casual sebum content. Arch Dermatol. 1991, 127 (12), 1806–1809. DOI: 10.1001/archderm.127.12.1806. [DOI] [PubMed] [Google Scholar]
- 28.Croxton RS; Baron MG; Butler D; Kent T; Sears VG Variation in amino acid and lipid composition of latent fingerprints. Forensic Sci Int. 2010, 199 (1–3), 93–102. DOI: 10.1016/j.forsciint.2010.03.019. [DOI] [PubMed] [Google Scholar]
- 29.Mukherjee S; Mitra R; Maitra A; Gupta S; Kumaran S; Chakrabortty A; Majumder PP Sebum and Hydration Levels in Specific Regions of Human Face Significantly Predict the Nature and Diversity of Facial Skin Microbiome. Sci Rep. 2016, 6, 36062. DOI: 10.1038/srep36062. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Girod A; Ramotowski R; Weyermann C Composition of fingermark residue: A qualitative and quantitative review. Forensic Sci. Int 2012, 223, 10–24. DOI: 10.1016/j.forsciint.2012.05.018. [DOI] [PubMed] [Google Scholar]
- 31.Roh M; Han M; Kim D; Chung K Sebum output as a factor contributing to the size of facial pores. Br J Dermatol. 2006, 155 (5), 890–894. DOI: 10.1111/j.1365-2133.2006.07465.x. [DOI] [PubMed] [Google Scholar]
- 32.Shetage SS; Traynor MJ; Brown MB; Chilcott RP Sebomic identification of sex- and ethnicity-specific variations in residual skin surface components (RSSC) for bio-monitoring or forensic applications. Lipids Health Dis. 2018, 17 (1), 194. DOI: 10.1186/s12944-018-0844-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Ní Raghallaigh S; Bender K; Lacey N; Brennan L; Powell FC The fatty acid profile of the skin surface lipid layer in papulopustular rosacea. Br J Dermatol. 2012, 166 (2), 279–287. DOI: 10.1111/j.1365-2133.2011.10662.x. [DOI] [PubMed] [Google Scholar]
- 34.Agrawal K; Hassoun LA; Foolad N; Borkowski K; Pedersen TL; Sivamani RK; Newman JW Effects of atopic dermatitis and gender on sebum lipid mediator and fatty acid profiles. Prostaglandins Leukot Essent Fatty Acids. 2018, 134, 7–16. DOI: 10.1016/j.plefa.2018.05.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Isom M; Go EP; Desaire H Groomed Fingerprint Sebum Sampling: Reproducibility and Variability According to Anatomical Collection Region and Biological Sex. Molecules. 2025, 30(3), 726. DOI: 10.3390/molecules30030726. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Jacobi U; Gautier J; Sterry W; Lademann J Gender-related differences in the physiology of the stratum corneum. Dermatology. 2005, 211 (4), 312–317, DOI: 10.1159/000088499. [DOI] [PubMed] [Google Scholar]
- 37.O’Neill KC; Hinners P; Jin Lee Y Potential of triacylglycerol profiles in latent fingerprints to reveal individual diet, exercise, or health information for forensic evidence. Anal. Methods 2020, 12 (6), 792–798. DOI: 10.1039/C9AY02652E. [DOI] [Google Scholar]
- 38.Sadowski T; Klose C; Gerl MJ; Wójcik-Maciejewicz A; Herzog R; Simons K; Reich A; Surma MA Large-scale human skin lipidomics by quantitative, high-throughput shotgun mass spectrometry. Sci Rep. 2017, 7, 43761. DOI: 10.1038/srep43761. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Williams DK; Brown CJ; Bruker J Characterization of children’s latent fingerprint residues by infrared microspectroscopy: forensic implications. Forensic Sci Int. 2011, 206 (1–3), 161–165. DOI: 10.1016/j.forsciint.2010.07.033. [DOI] [PubMed] [Google Scholar]
- 40.Isom M; Go EP; Desaire H Enabling Lipidomic Biomarker Studies for Protected Populations by Combining Noninvasive Fingerprint Sampling with MS Analysis and Machine Learning. J. Proteome Res 2024, 23(8), 2805–2814. DOI: 10.1021/acs.jproteome.3c00368. [DOI] [PubMed] [Google Scholar]
- 41.Chen T; Guestrin C XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; ACM: New York, NY. 2016, 785– 794. [Google Scholar]
- 42.He L; Diedrich J; Chu Y-Y; Yates JR Extracting accurate precursor information for tandem mass spectra by RawConverter. Anal. Chem 2015, 87(22), 11361– 11367. DOI: 10.1021/acs.analchem.5b02721. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.R Core Team. R: A language and environment for statistical computing. R Foundation for Statistical Computing: Vienna, Austria. 2020. [Google Scholar]
- 44.Ritchie ME; Phipson B; Wu D; Hu Y; Law CW; Shi W; Smyth GK limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015, 43 (7), e47. DOI: 10.1093/nar/gkv007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Kassambara A; Mundt F Factoextra: Extract and Visualize the Results of Multivariate Data Analyses. R Package Version 1.0.7, 2020. [Google Scholar]
- 46.Wickham H ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag: New York, NY, 2016. [Google Scholar]
- 47.Desaire H How (Not) to Generate a Highly Predictive Biomarker Panel Using Machine Learning. J Proteome Res. 2022, 21 (9), 2071–2074. DOI: 10.1021/acs.jproteome.2c00117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Robin X; Turck N; Hainard A; Tiberti N; Lisacek F; Sanchez J-C; Müller M pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinform. 2011, 12, 77. DOI: 10.1186/1471-2105-12-77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Conroy MJ; Andrews RM; Andrews S; Cockayne L; Dennis EA; Fahy E; Gaud C; Griffiths WJ; Jukes G; Kolchin M; et al. LIPID MAPS: Update to Databases and Tools for the Lipidomics Community. Nucleic Acids Res. 2024, 52, D1677–D1682 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Chua AE; Go EP; Desaire H Exploring Sample Storage Conditions for the Mass Spectrometric Analysis of Extracted Lipids from Latent Fingerprints. Biomolecules. 2025, 15(4), 477. DOI: 10.3390/biom15040477. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Bouslimani A; Porto C; Rath CM; Wang M; Guo Y; Gonzalez A; Berg-Lyon D; Ackermann G; Moeller Christensen GJ; et al. Molecular cartography of the human skin surface in 3D. Proc Natl Acad Sci U S A. 2015, 112(17), E2120–2129. DOI: 10.1073/pnas.1424409112. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Figure S1: Histograms of participants by age.
Figure S2. Mass spectrometry data of lipid standard.
Figure S3. Mass spectrometry data of the key lipids differentiating children and adults.
Figure S4. Butterfly plots comparing samples pooled by injection order or age.
Table S1: Participant demographics (xlsx)
Table S3: Normalized feature matrices (xlsx)
Table S2: Raw data matrix (xlsx)
Data Availability Statement
The raw MS data is included in the Supporting Information.
