Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Apr 13.
Published in final edited form as: Environ Res. 2026 Apr 1;300:124393. doi: 10.1016/j.envres.2026.124393

A two-step Bayesian clustering approach to relate PFAS mixture profiles and dietary patterns in early pregnancy

Xuzhi Wang 1, Tamarra James-Todd 2,3, Ami R Zota 4, Luke Shawler 1, Pi-I D Lin 5, Emily Oken 5,6, Jorge Chavarro 6, Sharon Sagiv 7, Briana JK Stephenson 1
PMCID: PMC13070428  NIHMSID: NIHMS2163476  PMID: 41933863

Abstract

Diet is a major source of exposure to per- and polyfluoroalkyl substances (PFAS). However, few studies have investigated dietary patterns in relation to differences observed in the PFAS exposure patterns. We conducted a cross-sectional analysis within the Project Viva cohort of 1,383 pregnant women enrolled at their first prenatal visit (1999–2002), who completed a validated food frequency questionnaire (158 items) and had plasma concentrations of six PFAS. We implemented a Bayesian repulsive Gaussian mixture model (BRGM) to identify PFAS exposure patterns. These patterns defined our PFAS subpopulations for a subsequent analysis, where we derived dietary patterns that accounted for differences amongst these subpopulations. Using a robust profile clustering (RPC) model, we detected dietary patterns both across the entire cohort and within specific PFAS subpopulations. The BRGM model identified six PFAS subpopulations with distinct PFAS exposure levels. Across the entire cohort, we identified six dietary patterns, plus one additional local dietary pattern per PFAS subpopulation identified by the RPC model. Notably, we observed clear consumption differences between high- and low-PFAS exposure subpopulations. Subpopulations with higher PFAS concentrations showed greater consumption of sugar-sweetened beverages, packaged and processed condiments, decaf coffee (during pregnancy), skim milk, and poultry, while those with lower PFAS concentrations tended to show greater consumption of vegetables, fruits, regular beer (before pregnancy), hot cereal, and sweet baked goods. Our approaches provided novel insights into the relationship between prenatal exposure to PFAS and dietary intake by identifying dietary patterns unique to different PFAS subpopulations.

Keywords: dietary patterns, PFAS, robust profile clustering, Bayesian repulsive Gaussian mixture model, exposure analysis

1. Introduction

Per- and polyfluoroalkyl substances (PFAS) are fluorinated synthetic compounds widely used in industrial and consumer products. Their applications include non-stick and stain-resistant coatings, firefighting foams, food packaging materials, upholstery fabrics, and carpets (Lindstrom et al., 2011). PFAS are highly persistent in the environment, and many have long elimination half-lives in the human body, ranging from 2 to 7 years (Olsen et al., 2007). U.S. adults are commonly exposed to multiple PFAS, primarily through the consumption of contaminated food and water (Egeghy and Lorber, 2011). Pregnancy is a sensitive window of exposure to harmful chemicals, like PFAS, with long-lasting health implications on both maternal and child health. Previous studies have identified associations between prenatal exposure to PFAS and various adverse maternal and neonatal outcomes, including increased diastolic blood pressure, a higher risk of gestational hypertension, altered maternal and neonatal thyroid function, and increased risks of preterm birth, miscarriage, and preeclampsia (Gao et al., 2021; Preston et al., 2022; Preston et al., 2020; Preston et al., 2018).

Diet is considered a major source of PFAS exposure in the general population without occupational exposure or high intake of contaminated water (Domingo, 2012; Fraser et al., 2013). PFAS contamination in food can result from bioaccumulation in food chains or contact with PFAS-containing cookware and food packaging (Lin et al., 2020). In addition, PFAS-contaminated sludge has been applied as fertilizer on crops (Zhou et al., 2024). Previous studies assessing the relationships between diet and PFAS have reported that higher PFAS concentrations may be associated with higher intake of coffee and tea, meat, fish and other seafood, low-fiber and high-fat foods, and various packaged foods. In contrast, lower PFAS concentrations have been associated with better overall dietary quality, vegetables, fruits, dietary fiber, as well as food prepared at home (Hampson et al., 2024; Lin et al., 2020; Zhou et al., 2019; Christensen et al., 2017; Sultan et al., 2023; Dzierlengaa et al., 2021). In cohorts of pregnant women specifically, higher PFAS concentrations have been associated with greater intake of milk and cheese, fish and seafood, red meat, poultry, animal offal, and high-fat and fried foods. Lower PFAS concentrations, by contrast, have been linked with higher intake of plant-based foods, wheat, cereals, and soy products (Eick et al., 2023; Papadopoulou et al., 2019; Huo et al., 2023; DeLuca et al., 2025; Fabelova et al., 2022; Shu et al., 2018; Tian et al., 2018; Wang et al., 2016). In Project Viva, a cohort study of pregnant women and their children, Seshasayee et al. (2021) found that children whose dietary patterns included fish, ice cream, soda, and other frequently packaged foods (e.g., candy and salad dressing) had higher PFAS concentrations (Seshasayee et al., 2021).

Previous studies have focused on examining multiple PFAS concentrations separately using models like multivariate linear regressions (Eick et al., 2023; Lin et al., 2020; Preston et al., 2018). However, since individuals are ubiquitously exposed to multiple PFAS, an increasing number of studies are investigating the combined effects of PFAS mixtures using weighted quantile sum (WQS) regression and Bayesian kernel machine regression (BKMR) (Bobb et al., 2018; Carrico et al., 2015; Pan et al., 2024; Preston et al., 2022; Preston et al., 2020; Rosato et al., 2022). While all of these methods have found great utility in the field, none of them are able to explicitly account for population heterogeneity or identify latent exposure patterns from mixtures of PFAS that are correlated and coexisting. Algorithm-based clustering methods, such as k-means, have been used in a few studies to group individuals with similar PFAS exposure profiles, but they require a predefined number of clusters, fail to account for cluster membership uncertainty, and may produce redundant or small clusters (Kalloo et al., 2020; Schildroth et al., 2021). The Bayesian repulsive Gaussian mixture (BRGM) model, which was recently introduced by Xie and Xu, could effectively address the limitations in these current approaches (Xie and Xu, 2019). BRGM is an unsupervised model-based clustering approach that accounts for population heterogeneity and identifies subgroups in the population that share similar exposure patterns of PFAS mixtures. These patterns are both identifiable and easily interpretable. These population subgroups define subpopulations of PFAS mixture patterns that can be further analyzed to understand how food consumption habits share similarities or differences across the different subpopulations.

Current PFAS-diet research (Eick et al., 2023; Hampson et al., 2024; Lin et al., 2020; Odediran and Obeng-Gyasi, 2024; Seshasayee et al., 2021; Sultan et al., 2023) faces parallel limitations, relying on methods that either (1) group food by PFAS explanatory power while ignoring population heterogeneity (such as using reduced rank regression), (2) obscure the contribution of individual food through composite score (such as adherence scores or indices), or (3) examine food in isolation (Eick et al., 2023; Switkowski et al., 2019; Woo Baidal et al., 2018). Latent class models are another unsupervised model-based clustering approach that has been frequently used to identify the population heterogeneity of food consumption habits in a population (Hernandez et al., 2024; Stephenson et al., 2020a; Stephenson et al., 2020b; Stephenson and Willett, 2023), but they are restricted to a global clustering assumption, preventing the ability to see how these shared behaviors may differ by exposure to different PFAS mixture subgroups. In contrast, robust profile clustering (RPC) provides a more flexible unsupervised approach that simultaneously identifies global dietary patterns shared across the entire population, and local patterns specific to particular PFAS exposure subpopulations. When combined with BRGM model for PFAS exposure profiling, this two-step clustering approach offers a layered insight on how two exposures interrelate with one another. To our knowledge no existing studies have investigated the relationships between dietary patterns and PFAS mixtures using such a flexible model-based two-step clustering approach.

Our study aims to address this gap with two specific objectives: (1) to derive and identify PFAS exposure profiles in pregnant women, and (2) to characterize dietary consumption patterns while accounting for these PFAS exposure profiles.

2. Methods

2.1. Study population

The study population consisted of pregnant women enrolled in Project Viva, a prospective prebirth cohort designed to investigate the effects of pre- and perinatal factors on long-term health outcomes (Oken et al., 2015). Project Viva enrolled pregnant women during their first prenatal visit (median 9.6 weeks of gestation) between 1999 and 2002 at a multispecialty group practice in eastern Massachusetts. At the first study visit immediately after the women’s initial prenatal visit during the first trimester, trained research staff obtained written informed consent, conducted a brief interview, provided take-home self-administered questionnaires, and collected blood samples. The following demographic variables from the questionnaires were included in our analysis: participants’ age (continuous, in years), educational level (high school degree or less, some college or an associate’s degree, 4 years of college, or graduate degree), smoking history (never smoked, former smoker, or smoked during pregnancy), race/ethnicity (Hispanic/Latina, White, Black, or others), household income (≤$40,000, $40,001-$70,000, >$70,000, or “don’t know”), and employment status (yes or no). Our analysis also included a validated food frequency questionnaire (FFQ) that assessed the participants’ diet during early pregnancy (Fawzi et al., 2004). Eligible participants were fluent in English, had singleton gestations, were less than 22 weeks’ gestation, and planned to deliver in eastern Massachusetts. Project Viva included 2,670 pregnancies, of which 2,128 resulted in live births with follow-up. After excluding multiple enrollments (n=28), the analysis was restricted to 2,100 unique women with their first enrollment records. Of the 2,100 mothers who had first live births, 1,627 provided maternal PFAS measurements, and of those, 1,383 completed the first trimester FFQ and had no missing values in the 158 food items that we selected (Figure 1).

Figure 1. Study sample selection.

Figure 1.

PFAS: per- and polyfluoroalkyl substance; BRGM: Bayesian repulsive Gaussian mixture model; FFQ: food frequency questionnaire.

All study protocols were approved by the institutional review boards of all participating institutions, and all participating women provided written informed consent.

2.2. PFAS measurements in plasma

The collection, storage, and quantification of PFAS were detailed previously (Sagiv et al., 2015). Briefly, maternal plasma samples were collected from participants during their initial prenatal visit and stored until analysis. At the Division of Laboratory Sciences at the Centers for Disease Control and Prevention (Atlanta, GA), the samples were analyzed for concentrations of eight PFAS [perfluorohexane sulfonate, PFHxS; perfluorooctane sulfonate, PFOS; perfluorooctanoate, PFOA; perfluorononanoate, PFNA; perfluorodecanoate, PFDA; 2-(N-ethyl-perfluorooctane sulfonamido) acetate, EtFOSAA; 2-(N-methyl-perfluorooctane sulfonamide) acetate, MeFOSAA; perfluorooctane sulfonamide, PFOSA] using solid-phase extraction coupled with isotope dilution high-performance liquid chromatography-tandem mass spectrometry, as previously described (Kato et al., 2011). PFDA and PFOSA were detected in less than 50% of the samples and were therefore excluded from further analyses, while the other six PFAS were detected in 99–100% of the samples.

The limits of detection (LOD) were 0.2 ng/mL for PFOS and 0.1 ng/mL for all other PFAS. PFAS concentration values below the LOD were initially replaced by LOD/2 (Preston et al., 2022). For PFOS, replacing values <LOD with LOD/2 led to a small artificial cluster consisting of only three participants when applying BRGM. Therefore, to avoid spurious clustering driven by these few observations, the three PFOS values <LOD were replaced with a constant value of 2.6, corresponding to the next lowest detected value. Due to skewness in their distributions, all PFAS concentrations were log-transformed on the original scale, and these log-transformed PFAS concentrations were used in the analysis.

2.3. Dietary intake assessments

Dietary intake was assessed using a semi-quantitative FFQ, which was completed by pregnant women during their first trimester. The FFQ was developed based on the extensively validated Willet FFQ and was modified for use among pregnant women (Hu et al., 1997; Rifas-Shiman et al., 2009; Rimm et al., 1992). The FFQ included 167 questions on the average consumption frequency of specific foods from the last menstrual period up to the date of FFQ completion within the first trimester of pregnancy. Of the 167 FFQ food items, two were removed for redundancy (fruit, vegetables). Seven were removed for sparse consumption, with <6% of the participants reporting intake [regular beer during pregnancy, light beer during pregnancy, liquor during pregnancy, white wine during pregnancy, red wine during pregnancy, oat bran (added to food), and other bran (added to food)]. These food items were excluded because their low levels of consumption limit their informativeness for stable clustering using BRGM and RPC. We focused on the remaining 158 FFQ food items for our subsequent analysis.

Each food item in the FFQ included a specified serving size, and the participants were asked to report how often they consumed that serving size on average, ranging from no consumption to at least once a day. Therefore, each food item had 2–5 levels on the frequency of consumption, and the number of levels may vary depending on the specific food item. To avoid sparse consumption levels, for each food item we collapsed consumption levels with fewer than 5% of participants, so that at least 5% of participants were included in each level. This resulted in foods being collapsed to 2-level, 3-level, 4-level, and 5-level variables. A detailed description of the food items in each of the food groups was provided in Figures S1S5 in the Supplementary Materials.

2.4. Statistical Methods

Our statistical analysis consisted of two steps. First, we generated PFAS exposure profiles based on the joint distribution of the six PFAS variables included for analysis. This was performed using the BRGM model among the 1,627 participants who provided PFAS measurements. In the second step, we restricted the participants to those who had both PFAS measurements and completed first trimester FFQ (n=1,383). We defined the generated PFAS exposure profiles as subpopulations. We then identified dietary patterns shared across the overall population and those localized to each subpopulation. We performed the analysis in the second step using the RPC model.

2.4.1. Bayesian repulsive Gaussian mixture model

Gaussian mixture models are a popularly used clustering approach to identify shared behaviors within a heterogeneous population amongst a set of normally distributed continuous exposures. These shared behaviors are grouped together into latent subgroups, known as cluster profiles. Each cluster profile can be described using a multivariate gaussian distribution for each of the exposures included in the model. When the number of variables and sample size is large, Bayesian nonparametric Gaussian mixture models, such as the Dirichlet Process Mixture Model (DPMM) are typically implemented to allow flexibility to handle the large set of exposure variables and the unknown number of cluster profiles (Miller and Harrison, 2013; Miller and Harrison, 2017). However, DPMM is prone to generate redundant and extraneous clusters, compromising identifiability. This is due to the type of prior set on the parameters that describe each individual cluster. To address this limitation, the BRGM model was developed to introduce a repulsive prior on cluster parameters, encouraging substantial separation between clusters. This reduces potential redundancies in the clusters and facilitates interpretation of the clusters as biologically meaningful subpopulations (Xie and Xu, 2019).

Mathematically, the BRGM model uses the following notations. Let xi=xi1,xi2,,xip denote the vector of p PFAS exposure variables for subject i(1,2,,n), where n is the total number of subjects. We assume

f(xiK,wk,μk,Σkk=1K)=k=1Kwkfkxiμk,Σk

where K is the number of clusters, which is random and determined by the data. Let wkk=1K represent the mixture proportions for the K clusters, where each wk0 and k=1Kwk=1. Let fkxiμk,Σk denote the multivariate Gaussian density for cluster k with mean vector, μk=(μ1k,,μpk) and covariance matrix Σk of dimensions p×p.

A unique feature of the BRGM model is the introduction of the repulsive prior to ensure that the clusters are well-separated. The degree of repulsion is controlled by a hyperparameter g0 in the repulsive prior. We set g0=100 to impose strong shrinkage and encourage well-separated clusters. This choice was motivated by prior work that recommends relatively large values of g0 in multivariate clustering, with a larger value used here to reflect the higher dimensionality of our data (Xie and Xu, 2019). Further details are provided in the section Repulsive priors for cluster centers in BRGM in Supplementary Materials. In practical terms, the BRGM model groups women into a small number of distinct PFAS exposure profiles, avoiding redundant clusters.

2.4.2. Robust profile clustering

Latent class models have been applied to categorical dietary data to identify shared dietary patterns across a large set of food items. However, the traditional latent class model operates under a global clustering assumption, where each latent class, or clustered profile, shares the same food consumption behaviors for all food items. Given the complex heterogeneity of diet within a diverse population, this assumption is often unrealistic. For example, populations that exhibit different PFAS exposure profiles may also vary in the consumption of different foods in their diet. The RPC model relaxes this assumption by introducing local clusters that allow foods to deviate from their ‘global’ dietary patterns based on a subpopulation they belong to (e.g., PFAS exposure profiles derived from the first step) (Stephenson et al., 2020a). The probabilistic framework of the RPC model comprises three components: global clustering, where dietary patterns are shared amongst the overall population; local clustering, where dietary patterns are shared within a defined subpopulation (e.g. PFAS exposure profile); and deviation indicator, which indicates whether a food item will assume the ‘global’ dietary pattern or the ‘local’ dietary pattern.

Global clustering

The global dietary clusters capture the latent dietary behaviors shared across the overall population. The assignment of global clusters is dependent on the following probability vector, π=π1,π2,,πK0, where K0 is the total number of global clusters, and π is the vector of probabilities describing the individual membership probabilities for each of the K0 global clusters. Given the assignment to a global cluster h1,2,,K0, different consumption levels for food item j(1,2,,q) are characterized by the probability vector θ0jh=(θ0j1h,,θ0jdjh), where dj is the total number of consumption levels for food item j.

Local clustering

The local dietary clusters identify shared latent dietary behaviors that are specific to a defined subpopulation. These clusters are driven by the foods that differ from their global dietary patterns within each subpopulation. The assignment of the local clusters for subjects within each subpopulation is determined similarly to the global cluster assignments, where λsi=(λ1si,,λKssi), and λsi indicates the probabilities that subject i in subpopulation si belongs to each of the KS local clusters, si(1,,S) represents the subpopulation index for subject i(1,n), and Ks denotes the number of local clusters. Given the assignment to a local cluster l1,,Ks within subpopulation si, the probability vector of dj consumption levels for food j is denoted as θ1jlsi=(θ1j1lsi,,θ1jdjlsi).

Deviation indicator

Each food item j is assumed to describe the dietary patterns at either the global level or local level for each subject i, based on a variable deviation indicator Gij, where Gij=1 for the global level and Gij=0 for the local level. This variable deviation indicator is determined by the probability vj(s)=PGij=1si=s, which represents the probability of food item j being allocated to the global level versus the local level for subject i within subpopulation s. A higher vj(s) value indicates that individuals in subpopulation s consume food j similarly to the overall population, meaning the food follows the global clustering pattern. Conversely, a lower vj(s) value suggests their consumption of food j differs from the population norm, instead reflecting a unique local pattern specific to that PFAS subpopulation. This parameter essentially measures how strongly each food item’s consumption aligns with either the broader population trends or distinctive subgroup behaviors. To determine the allocation of each food item, we impose a conservative threshold of 0.45 to the deviation indicator probabilities. In other words, foods with a probability vj(s)>45% are classified as global, while foods with a probability vj(s)45% are classified as local. We chose 0.45 because this threshold ensures that only foods with a probability vj(s)45% are classified as local, such that it introduces more informative local clustering than the threshold of 50%.

For more details on likelihood construction, please refer to the section Likelihood of RPC in Supplementary Materials. In practical terms, the RPC model identifies dietary patterns that are shared across the entire cohort and those that are specific to each derived PFAS exposure profile.

2.4.3. Posterior computation

Both models were estimated using Bayesian approaches. We assumed no previous knowledge of the model parameters to allow the observed data to drive posterior inference. Details on prior selection, posterior estimation, and code implementation were described previously in (Xie and Xu, 2019) for the BRGM model and in (Stephenson et al., 2020a; Stephenson et al., 2020b; Stephenson and Willett, 2023) for the RPC model. The BRGM model was analyzed using MATLAB 2023a, and the RPC model was implemented in R 4.1.2 using the code on GitHub available at https://github.com/smwu/WRPC (Wu et al., 2024).

3. Results

Demographic characteristics of the 1,627 participants who provided PFAS measurements were summarized in Table 1. Pregnant participants had an average age of 32 years. The majority of the participants were employed (72%), White (68%), completed four years of college or had a graduate degree (64%), had a total household income of >$70000 (53%), and never smoked (67%).

Table 1.

Demographics by PFAS subpopulations in 1627 participants.

Variable Total
(n=1627)
Subpopulation 1
(n=115)
Subpopulation2
(n=591)
Subpopulation 3
(n=639)
Subpopulation 4
(n=29)
Subpopulation 5
(n=196)
Subpopulation 6
(n=57)
Age (years), mean (SD) 31.8 (5.17) 31.3 (5.57) 31.6 (5.18) 31.9 (5.04) 32.5 (3.57) 32.4 (5.33) 32.1 (5.77)
Employment, n (%)
 Yes 1174 (72.2%) 88 (76.5%) 425 (71.9%) 463 (72.5%) 22 (75.9%) 136 (69.4%) 40 (70.2%)
 No 203 (12.5%) 10 (8.7%) 63 (10.7%) 83 (13.0%) 3 (10.3%) 31 (15.8%) 13 (22.8%)
 Missing 250 (15.4%) 17 (14.8%) 103 (17.4%) 93 (14.6%) 4 (13.8%) 29 (14.8%) 4 (7.0%)
Race and Ethnicity, n (%)
 Hispanic/Latina 115 (7.1%) 6 (5.2%) 32 (5.4%) 53 (8.3%) 2 (6.9%) 17 (8.7%) 5 (8.8%)
 White 1102 (67.7%) 85 (73.9%) 411 (69.5%) 435 (68.1%) 15 (51.7%) 120 (61.2%) 36 (63.2%)
 Black 251 (15.4%) 17 (14.8%) 88 (14.9%) 96 (15.0%) 3 (10.3%) 37 (18.9%) 10 (17.5%)
 Other 141 (8.7%) 7 (6.1%) 49 (8.3%) 51 (8.0%) 9 (31.0%) 19 (9.7%) 6 (10.5%)
 Missing 18 (1.1%) 0 (0%) 11 (1.9%) 4 (0.6%) 0 (0%) 3 (1.5%) 0 (0%)
Education, n (%)
 High school degree or less 190 (11.7%) 13 (11.3%) 72 (12.2%) 79 (12.4%) 1 (3.4%) 18 (9.2%) 7 (12.3%)
 Some college or an associate’s degree 381 (23.4%) 38 (33.0%) 133 (22.5%) 158 (24.7%) 5 (17.2%) 35 (17.9%) 12 (21.1%)
 4 years of college 587 (36.1%) 33 (28.7%) 232 (39.3%) 225 (35.2%) 8 (27.6%) 72 (36.7%) 17 (29.8%)
 Graduate degree 451 (27.7%) 31 (27.0%) 143 (24.2%) 173 (27.1%) 15 (51.7%) 68 (34.7%) 21 (36.8%)
 Missing 18 (1.1%) 0 (0%) 11 (1.9%) 4 (0.6%) 0 (0%) 3 (1.5%) 0 (0%)
Total Household Income, n (%)
 <=$40,000 221 (13.6%) 14 (12.2%) 64 (10.8%) 92 (14.4%) 6 (20.7%) 34 (17.3%) 11 (19.3%)
 40,001–70,000 354 (21.8%) 32 (27.8%) 124 (21.0%) 139 (21.8%) 4 (13.8%) 43 (21.9%) 12 (21.1%)
 >$70,000 869 (53.4%) 56 (48.7%) 322 (54.5%) 351 (54.9%) 16 (55.2%) 103 (52.6%) 21 (36.8%)
 Don’t know 59 (3.6%) 4 (3.5%) 25 (4.2%) 21 (3.3%) 0 (0%) 4 (2.0%) 5 (8.8%)
 Missing 124 (7.6%) 9 (7.8%) 56 (9.5%) 36 (5.6%) 3 (10.3%) 12 (6.1%) 8 (14.0%)
Smoking Status, n (%)
 Former Smoker 299 (18.4%) 25 (21.7%) 121 (20.5%) 112 (17.5%) 6 (20.7%) 31 (15.8%) 4 (7.0%)
 Smoked During Pregnancy 215 (13.2%) 15 (13.0%) 94 (15.9%) 81 (12.7%) 2 (6.9%) 17 (8.7%) 6 (10.5%)
 Never Smoked 1097 (67.4%) 74 (64.3%) 369 (62.4%) 440 (68.9%) 21 (72.4%) 146 (74.5%) 47 (82.5%)
 Missing 16 (1.0%) 1 (0.9%) 7 (1.2%) 6 (0.9%) 0 (0%) 2 (1.0%) 0 (0%)

The BRGM model identified six PFAS exposure profiles from the 1,627 participants. These PFAS exposure profiles defined six respective subpopulations which were utilized in the subsequent RPC model. The PFAS exposure levels across six subpopulations were displayed in Figure 2. PFAS subpopulation 1 (highest PFAS profile) demonstrated the highest average levels across all PFAS concentration variables, followed by subpopulation 2 (second highest PFAS profile). PFAS subpopulations 3 (intermediate PFAS profile) and 4 (highest PFNA profile) showed intermediate average PFAS exposure levels. PFAS subpopulations 5 (second lowest PFAS profile) and 6 (lowest PFAS profile) showed an overall lower level of PFAS exposures, with subpopulation 6 showing the lowest levels. Notably, in subpopulation 4 (highest PFNA profile), five out of six PFAS concentrations were at intermediate levels, except for PFNA, which had the highest average level. Additionally, subpopulation 4 (highest PFNA profile) included some extremely low PFHxS values compared to the other subpopulations. PFAS subpopulation 1 (highest PFAS profile) had median log-transformed PFOS and PFOA concentrations approximately two-fold and three-fold higher, respectively, than those observed in subpopulation 6 (lowest PFAS profile). The two largest PFAS subpopulations were subpopulations 2 (36.3%) and 3 (39.2%). Subpopulations 4 (1.8%) and 6 (3.5%) had the smallest proportions of participants assigned. Subpopulations 1 (7.1%) and 5 (12%) had intermediate proportions of participants.

Figure 2. Distribution of individual plasma PFAS concentration across each PFAS subpopulation.

Figure 2.

Each panel corresponds to a PFAS subpopulation (cluster). Within each panel, the y-axis represents the values of the log-transformed plasma PFAS [log(ng/mL)], while the boxes denote individual PFAS concentration in the following order (left to right): EtFOSAA, MeFOSAA, PFHxS, PFNA, PFOA, and PFOS. Subpopulation 1: highest average PFAS exposure levels; Subpopulation 2: second highest average PFAS exposure levels; Subpopulation 3: intermediate average PFAS exposure levels; Subpopulation 4: intermediate average PFAS exposure levels but highest PFNA levels and lower PFHxS levels; Subpopulation 5: second lowest average PFAS exposure levels; Subpopulation 6: lowest average PFAS exposure levels.

Subpopulation-specific demographic and PFAS characteristics were summarized in Table 2. PFAS subpopulation 1 (highest PFAS profile) had the highest proportion of White participants (74%) and the highest proportion of former smokers before pregnancy (22%). PFAS subpopulation 2 (second highest PFAS profile) had the highest proportion of participants who smoked during pregnancy (16%). PFAS subpopulation 3 (intermediate PFAS profile) showed intermediate levels for most demographic variables. PFAS subpopulations 4 (highest PFNA profile) (78%) and 5 (second lowest PFAS profile) (70%) had a higher proportion of participants who completed four years of college or had a graduate degree. PFAS subpopulation 6 (lowest PFAS profile) had the highest proportion of participants who never smoked (83%).

Table 2.

Subpopulation-specific characteristics.

Subpopulation Demographic characteristics PFAS characteristics
Subpopulation 1 (7.1%) Highest proportion of White participants and former smokers before pregnancy Highest average levels across all PFAS exposure variables
Subpopulation 2 (36.3%) Highest proportion of participants who smoked during pregnancy Second highest average PFAS exposure levels
Subpopulation 3 (39.2%) Intermediate levels for most demographic variables Intermediate average PFAS exposure levels
Subpopulation 4 (1.8%) Highest proportion of participants who completed four years of college or had a graduate degree Highest PFNA exposure level and some extremely low PFHxS exposure values; other PFAS variables show intermediate average exposure levels.
Subpopulation 5 (12%) Relatively high proportion of participants who completed four years of college or had a graduate degree Second lowest average PFAS exposure levels
Subpopulation 6 (3.5%) Highest proportion of participants who never smoked Lowest average PFAS exposure levels

Of the 1,627 participants with maternal plasma PFAS concentrations, 1,383 with complete first trimester FFQ data were included in the subsequent RPC analysis. Compared to the full sample (n=1,627), the subsample had a higher proportion of participants who were employed (83% vs. 72%) and White (73% vs. 68%). Other demographic and PFAS characteristics were comparable between the two groups. Demographic and PFAS characteristics of these 1,383 participants were summarized in Table S1 and Figure S6 in the Supplementary Materials.

Global dietary patterns

The RPC model identified a total of six global dietary patterns (DP), with the proportion of participants in each pattern as follows: DP A (29.2%), DP B (9.2%), DP C (10.6%), DP D (18.5%), DP E (24.3%), and DP F (8.2%). The heatmap in Figure 3, where food items were first grouped by the number of consumption levels and then by food categories, illustrated the consumption level with the highest posterior probability (“consumption mode”), for a given food within each global dietary pattern. Figure S7 showed the same global dietary patterns but with food items grouped only by food categories. Global DP A (“High Sugar & High Fat”) favored a diet high in added sugar, fat and alcohol, with higher consumption of soda, coffee, sweets (cookies, chocolate, brownie), French fries, butter, cream, margarine, beer, and wine. Global DP B (“Low Diversity & High Protein”) had the least diet diversity, favoring no consumption of the majority of the foods (i.e., 149 out of 158) queried. However, this pattern favored the consumption of foods with high protein and refined grain, including liver, eggbeaters or egg whites, chicken/turkey, and white rice. Global DP C (“Low Diversity & Processed”) also had low diet diversity but favored the consumption of processed foods (soda, French fries, chocolate, corn/corn chips, white bread, bagels, and added margarine) and animal-based foods (milk, chicken/turkey, bacon, and cheese). Global DP D (“Sandwich and Carbs”) favored higher consumption of diet sodas, common salad and sandwich ingredients (romaine lettuce, chicken or turkey sandwich, crackers, salad dressing, and cheese), as well as simple carbs (pasta, bagels, and dark bread). Global DP E (“Moderate and Mixed”) had a higher probability of consuming alcohol, coffee, snacks (crackers, cookies, pretzels, and ice cream), vegetables, fruits, and fish. Global DP F (“Meat & Sweets-Driven”) favored the consumption of meat, poultry, baked desserts (cake and pie), and soda.

Figure 3. Consumption modes for six global dietary patterns.

Figure 3.

All food items are first categorized into five groups by how consumption frequency was reported, separated by bold black horizontal lines: 2-level, 3-level, 4-level, 5-level, and other foods. Within each group, the food items are ordered by broader food categories. More details of these five groups of foods can be found in Supplementary Figures S1S5. An alternative ordering approach can be found in Figure S7 and Figure S9. The consumption modes for the first 4 food groups are color-coded as follows: 1. dark blue (no consumption); 2. light blue [(at least) 1–3 times/month]; 3. yellow: [(at least) once a week]; 4. orange [(at least) 2–4 times/week]; 5. red [at least) 5–6 times/week]. Interpretations for the consumption modes in the ‘other foods’ group can be found in Supplementary Figure S5. Global dietary pattern A: “High Sugar and High Fat”; Global dietary pattern B: “Low Diversity and High Protein”; Global dietary pattern C: “Low Diversity and Processed”; Global dietary pattern D: “Sandwich and Carbs”; Global dietary pattern E: “Moderate and Mixed”; Global dietary pattern F: “Meat and Sweets-Driven”.

The distribution of global dietary patterns across PFAS subpopulations was illustrated in Figure 4. All six global dietary patterns were identified in each of the six PFAS subpopulations. Global DP A (“High Sugar and High Fat”) was the largest and most prominent dietary pattern across PFAS subpopulations 1–4 (29%−33%), which had higher or intermediate average PFAS exposure levels. PFAS subpopulation 5 with the second lowest average PFAS exposure levels showed higher proportions of both Global DP A (“High Sugar and High Fat”, 21%) and Global DP C (“Low Diversity and Processed”, 21%). Global DP B (“Low Diversity and High Protein”, 27%) and Global DP E (“Moderate and Mixed”, 27%) were the most common dietary patterns in PFAS subpopulation 6 with the lowest average PFAS exposure levels.

Figure 4. Distribution of global dietary patterns across PFAS subpopulations.

Figure 4.

The y-axis represents six PFAS subpopulations, and the x-axis shows the proportion of each global dietary pattern within each PFAS subpopulation. Each bar is divided into segments representing different global dietary patterns, with the segment sizes and values corresponding to their relative proportions within the respective PFAS subpopulation. Subpopulation 1: highest average PFAS exposure levels; Subpopulation 2: second highest average PFAS exposure levels; Subpopulation 3: intermediate average PFAS exposure levels; Subpopulation 4: intermediate average PFAS exposure levels but highest PFNA levels and lower PFHxS levels; Subpopulation 5: second lowest average PFAS exposure levels; Subpopulation 6: lowest average PFAS exposure levels; Global dietary pattern A: “High Sugar and High Fat”; Global dietary pattern B: “Low Diversity and High Protein”; Global dietary pattern C: “Low Diversity and Processed”; Global dietary pattern D: “Sandwich and Carbs”; Global dietary pattern E: “Moderate and Mixed”; Global dietary pattern F: “Meat and Sweets-Driven”.

Local dietary patterns

While all participants across the six PFAS subpopulations were assigned to global dietary patterns, they each exhibited unique dietary behaviors specific to their subpopulation. Each PFAS subpopulation identified one local dietary pattern. Supplementary Figure S8 illustrated the posterior probability of each food item being assigned to the global vs the local level (i.e., vj(s)). Based on the 45% vj(s) threshold, of the 158 food items, 108 showed a tendency to deviate from their global dietary patterns and were assigned to the local level in at least one PFAS subpopulation (Figure 5 and Figure S9). Figures 5 and S9 displayed the same local dietary patterns but used different ordering approaches. Among these foods, 11 foods were classified as local across all six PFAS subpopulations (Figure 6), while the remaining foods were considered as local in some PFAS subpopulations and global in others.

Figure 5. Consumption modes within each PFAS subpopulation for 108 foods with local deviations.

Figure 5.

Grey spaces indicate that those food items were allocated to the global level within those PFAS subpopulations. Food items in red were more likely to be consumed in high PFAS concentration subpopulations, while food items in green were more likely to be consumed in low PFAS concentration subpopulations. All food items are first categorized into five groups by how consumption frequency was reported, separated by bold black horizontal lines: 2-level, 3-level, 4-level, 5-level, and other foods. Within each group, the food items are ordered by broader food categories. More details of these five groups of foods can be found in Supplementary Figures S1S5. An alternative ordering approach can be found in Figure S7 and Figure S9. The consumption modes for the first 4 food groups are color-coded as follows: 1. dark blue (no consumption); 2. light blue [(at least) 1–3 times/month]; 3. yellow: [(at least) once a week]; 4. orange [(at least) 2–4 times/week]; 5. red [(at least) 5–6 times/week]. Interpretations for the consumption modes in the ‘other foods’ group can be found in Supplementary Figure S5. Subpopulation 1: highest average PFAS exposure levels; Subpopulation 2: second highest average PFAS exposure levels; Subpopulation 3: intermediate average PFAS exposure levels; Subpopulation 4: intermediate average PFAS exposure levels but highest PFNA levels and lower PFHxS levels; Subpopulation 5: second lowest average PFAS exposure levels; Subpopulation 6: lowest average PFAS exposure levels.

Figure 6. Probability distributions of consumption for 11 foods classified as local across all six PFAS subpopulations.

Figure 6.

The y-axis denotes consumption level probability, while the x-axis represents subpopulation index. Each bar is divided into segments representing different consumption levels, with the segment sizes corresponding to their probabilities within each PFAS subpopulation. The consumption levels are color-coded as follows: 1. dark blue (no consumption); 2. light blue [(at least) 1–3 times/month]; 3. yellow: [(at least) once a week]. Subpopulation 1: highest average PFAS exposure levels; Subpopulation 2: second highest average PFAS exposure levels; Subpopulation 3: intermediate average PFAS exposure levels; Subpopulation 4: intermediate average PFAS exposure levels but highest PFNA levels and lower PFHxS levels; Subpopulation 5: second lowest average PFAS exposure levels; Subpopulation 6: lowest average PFAS exposure levels.

To determine the relationship between consumption of individual food items and PFAS exposure levels, we grouped foods according to whether they were more commonly consumed in higher PFAS versus lower PFAS subpopulations, as shown in Figure 5 and Table 3. We identified two groups of food items: the high PFAS food group and the low PFAS food group. The high PFAS food group consisted of foods more likely to be consumed by PFAS subpopulations with higher PFAS exposure levels (e.g., subpopulations 1–4). The low PFAS food group included foods more likely to be consumed by PFAS subpopulations with lower PFAS exposure levels (e.g., subpopulations 3–6). For the high PFAS food group, regular orange juice was either not consumed or consumed at least 1–3 times per month at the global level for participants in PFAS subpopulation 1(highest PFAS profile). In contrast, regular orange juice favored no consumption at the local level across all other PFAS subpopulations. In other words, no matter what global dietary patterns the participants in subpopulations 2–6 were assigned to, they were likely to not consume regular orange juice. Similarly, in PFAS subpopulation 1 (highest PFAS profile), added margarine was classified as global, with participants either consuming it at least five to six times a week (Global DP A: “High Sugar and Fat” and Global DP C: “Low Diversity and Processed”), at least 1–3 times a month (Global DP B: “Low Diversity and High Protein” and Global DP F: “Meat and Sweets-Driven”), or not consuming it at all (Global DP D: “Sandwich and Carbs” and Global DP E: “Moderate and Mixed”). For all remaining subpopulations, those participants favored no consumption of added margarine.

Table 3.

Food groups associated with different PFAS consumption levels.

Food group Food items included Characteristics
The high PFAS food group Regular orange juice, other sugary-sweetened beverages (before pregnancy), added margarine, decaf coffee (during pregnancy), skim milk, chicken or turkey with skin More likely to be consumed by PFAS subpopulations with higher PFAS exposure levels (e.g., subpopulations 1–4)
The low PFAS food group Kale, mustard, or chard greens, dark orange squash, yams or sweet potatoes, peaches, apple sauce, regular beer (before pregnancy), hot cereal, sweet rolls, ready-made sweet pastries, donuts More likely to be consumed by PFAS subpopulations with lower PFAS exposure levels (e.g., subpopulations 3–6)

Decaf coffee (during pregnancy), other sugary-sweetened beverages (before pregnancy), chicken or turkey with skin, and skim milk had a mode of no consumption at the local level in the PFAS subpopulations with overall lower exposure levels (i.e., subpopulations 4, 5 or 6). In contrast, these foods remained in accordance with the global dietary patterns among the subpopulations with medium to high PFAS exposure levels (i.e. subpopulations 1–3), where each food favored some level of consumption in at least one global dietary pattern.

In the low PFAS food group, there were 12 food items that were more likely to be consumed in the subpopulations with lower PFAS exposure levels. Regular beer (before pregnancy) favored no consumption at the local level in the greater exposed PFAS subpopulations, while they were either not consumed or consumed once per week among all other PFAS subpopulations with lower PFAS exposure levels. Kale, mustard, or chard greens, hot cereal, sweet rolls, apple sauce, dark orange squash, yams or sweet potatoes, ready-made sweet pastries, donuts, and peaches were allocated to the global level in subpopulations with medium to low PFAS exposure levels (subpopulations 4–6). At the global level, the consumption modes for these foods varied across different global dietary patterns. In contrast, subpopulations with medium to high PFAS exposure levels (subpopulations 1–3) were more likely to deviate in favor of no consumption for these foods at the local level.

Locally, 11 food items shared a mode of no consumption across all six PFAS subpopulations. Despite the similarity in their consumption mode, these 11 food items differed in the overall distribution considering all possible consumption levels (Figure 6). For example, subpopulation 6 (lowest PFAS profile) had the highest consumption probability of caffeinated tea and eggbeaters of at least once a week. Similarly, subpopulation 3 (intermediate PFAS profile) had the highest consumption probability of some level of consumption of tomato juice compared to other subpopulations.

Lastly, some foods favored higher consumption in some subpopulations. For example, PFAS subpopulation 4, which had the highest PFNA level, favored higher consumption of shellfish, other fish, whole eggs, soup, oranges, and added butter. PFAS subpopulation 1 (highest PFAS profile) favored higher consumption for popcorn and pasta. PFAS subpopulation 2 (second highest PFAS profile) favored more consumption of pancakes, whole eggs, and other fruit juices. Similarly, PFAS subpopulation 3 (intermediate PFAS profile) favored higher levels of consumption of pancakes and whole eggs.

4. Discussion

In our cohort of pregnant women, we applied two novel Bayesian clustering methods in a two-step process to identify dietary patterns accounting for the differences identified in the PFAS exposure profiles. In the first step we identified six PFAS exposure profiles which exhibited distinct overall PFAS exposure levels (e.g., high, intermediate, or low), across different participant demographics. For example, the patterns with the highest overall PFAS exposure levels across all PFAS had the highest proportion of White participants. In the second step, the derived PFAS exposure profiles were utilized to define subpopulations upon which dietary intake could vary. We identified six global dietary patterns shared across the overall cohort, as well as local dietary patterns unique to each PFAS subpopulation.

Our findings suggest that dietary patterns rich in processed and packaged foods are associated with higher PFAS exposure profiles, whereas more plant-forward dietary patterns tend to align with lower PFAS exposure profiles during pregnancy. Specifically, subpopulations with higher overall PFAS exposure levels favored consumption of sugar-sweetened beverages, packaged and processed condiments, decaf coffee (during pregnancy), skim milk, and poultry. Notably, while fish and other seafood have been commonly linked to higher PFAS concentrations in previous studies (Eick et al., 2023; Huo et al., 2023; Shu et al., 2018; Tian et al., 2018; Wang et al., 2016), we found that only the subpopulation with the highest PFNA exposure favored higher consumption of seafood (e.g., shellfish and other fish), eggs, soup, oranges, and added butter. In contrast, subpopulations with lower overall PFAS exposure levels favored the consumption of vegetables (kale, mustard, or chard greens, dark orange squash, yams or sweet potatoes), fruits (peaches and apple sauce), regular beer (before pregnancy), and hot cereal. Despite the consumption of some sweet baked goods, including sweet rolls, pastries, and donuts, the overall dietary patterns in the subpopulations with lower overall PFAS exposure levels were generally more nutrient-dense and plant-based.

Most of our findings were consistent with prior studies. Specifically, previous studies and our study found that fish and other seafood, packaged foods or condiments, coffee, milk, and poultry were associated with higher overall PFAS mixture profiles or multiple PFAS exposure. In contrast, vegetables, fruits, and cereal were associated with lower overall PFAS mixture profiles or multiple PFAS exposure in prior studies and our analysis. Notably, we observed higher intake of shellfish and other fish only in the subpopulation with the highest PFNA exposure levels. This may be attributed to the more granular categorization of food items in our study, which included 158 FFQ items and separated seafood into shellfish, dark meat fish, other fish, and canned tuna. One previous cross-sectional study among adults (65% female) with pre-diabetes reported that the strongest PFAS-diet associations were between PFNA and the consumption of fried fish and other fish/shellfish, which further supports our findings (Lin et al., 2020).

Notable differences were also observed with prior studies. Several foods that showed higher consumption in subpopulations with lower PFAS levels were either not reported or were reported to be associated with higher PFAS exposure in previous studies. These food items include regular beer (before pregnancy), which was not reported previously, and sweet baked goods, which have previously been reported to be positively associated with PFAS exposure. Reasons behind these differences may be attributed to: 1) our analysis included more granularity of food items allowing a more comprehensive picture of participant’s diet; 2) our two-step approach, which combined two novel clustering methods (i.e. BRGM, RPC) revealed greater insight into the diet-PFAS relationship for heterogeneous populations, not previously captured in conventional dietary analyses (Eick et al., 2023); 3) Our findings reflect differences in population characteristics, packaging-related exposure sources, and/or cooking materials, which were not fully captured through the FFQ. For example, previous studies have suggested that PFAS can enter the food chain through migration from food contact materials, such as food packaging and coated frying pans (Carnero et al, 2021). Additional dietary assessments are needed to further investigate the use of food packaging and cooking materials, which were not available for this study.

Our study had several strengths. First, we analyzed PFAS mixtures by identifying latent PFAS exposure profiles using the novel BRGM model, which allowed us to characterize the varying exposures of PFAS within the study population and better understand how each individual PFAS chemical contributed to the overall mixture of PFAS exposure. Implementation of the BRGM model allowed the data to derive the number of appropriate clusters via a penalization component to reduce complexity. The model also encouraged separation between clusters and generated meaningful PFAS subpopulations with distinct exposure profiles for subsequent analysis (Xie and Xu, 2019). Second, the RPC model offered several advantages to improve our understanding of how dietary patterns varied by different PFAS exposure profiles (Stephenson et al., 2020a; Stephenson et al., 2020b; Stephenson and Willett, 2023). As in BRGM, the RPC model was able to determine the number of dietary clusters directly from the data without the pre-specification of cluster number. The joint-stratified latent class structure allowed us to identify which dietary patterns were shared amongst the overall study population and which patterns were localized based on their subpopulation, defined by their PFAS exposure profile. This dual layer of the model avoided the need for subpopulation-specific analyses, which may reduce statistical power, yield misleading characterization of dietary patterns, and compromise generalizability.

Our analysis was met with some limitations. First, our study sample consisted of pregnant women who were predominantly non-Hispanic White, well-educated, had middle or upper income, and resided in eastern Massachusetts. This limited our generalization to all pregnant women in the US or other populations in Massachusetts that may share different demographic backgrounds. Second, our analysis results may be sensitive to certain model parameters. For example, the repulsion parameter in the BRGM model that controls the separation and identifiability of the different PFAS profiles is defined at the discretion of the researcher. More conservative approaches could yield a different number of profiles and different characterization. We chose our parameter, based on prior simulation studies and clinical interpretability of the respective results. Another model parameter of note is the threshold value in the RPC model to determine whether a food item should assume a global or subpopulation-specific pattern. We selected a more conservative threshold of 0.45. Different threshold values would yield different food items being defined and characterized at the two levels. Third, dietary assessments used in Project Viva were self-reported. Consequently, our results were prone to reporting bias and measurement errors. The overreporting of fruits and vegetables on FFQ as well as the underreporting of foods with high caloric intake was shown to be possible in previous studies, which may be a result of social desirability bias (Amanatidis et al., 2001; Haraldsdottir, 1993; Shu et al., 2004). Nevertheless, the use of FFQ, which is calibrated for use in pregnancy, is still a reliable epidemiologic tool to assess diet intake (Fawzi et al., 2004; McCullough et al., 2021). Fourth, our analysis did not include information about food packaging, which is another source of PFAS we were unable to account for in the data, due to limitations in the FFQ data (Hampson et al., 2024; Seshasayee et al., 2021). Assessments which include data on food packaging as part of the dietary assessment could lead to better understanding of diet-PFAS relationships in future studies. Fifth, as with all PFAS-related studies, some PFAS concentration values were below LOD and required imputation. However, such values below LOD were rare, and the imputation method used in our study was shown to produce accurate estimation of the mean and standard deviation in skewed distributions (Hornung and Reed, 1989). Finally, dietary intake was assessed during the first trimester and may not fully represent dietary patterns across the entire pregnancy, as eating habits can change during later gestation. Therefore, our findings may not be generalizable to the entire course of pregnancy. However, previous analyses from the Project Viva cohort have suggested that average intakes of foods and energy-adjusted nutrients changed little between the first and second trimesters (Rifas-Shiman et al., 2006; Rifas-Shiman et al., 2009), although these studies did not evaluate changes through the third trimester.

In conclusion, our study effectively identified heterogeneity in PFAS mixtures and captured dietary consumption differences across subpopulations with distinct PFAS exposure profiles. Our approaches, including the BRGM model and the RPC model, provided novel insights into the relationship between prenatal exposure to PFAS and dietary intake. Future directions may include examining whether dietary determinants of PFAS exposure vary by trimester, exploring health outcomes associated with the RPC-derived dietary patterns, and evaluating dietary changes from pregnancy to mid-life with the inclusion of additional dietary assessments for this cohort.

Supplementary Material

1

Highlights.

  • We observed six PFAS exposure subpopulations in 1,383 Project Viva pregnant women.

  • We identified six global dietary patterns shared across the overall cohort.

  • We observed one additional local dietary pattern within each PFAS subpopulation.

  • Dietary patterns differed between high- and low-PFAS exposure subpopulations.

Acknowledgements

The authors thank Ethan Powell for data cleaning, wrangling, and preliminary analysis of dietary intake data to motivate the appropriate analysis for this study. We also thank Project Viva participants and staff for their dedication to the study.

Funding sources

Project Viva is supported by grants from the National Institutes of Health and NIEHS (R01HD034568, R24ES030894, R01HD096032, R01ES031065, and U54 AG062322).

Abbreviations

PFAS

Per- and polyfluoroalkyl substances

BRGM

Bayesian repulsive Gaussian mixture model

RPC

Robust profile clustering

FFQ

Food frequency questionnaire

LOD

Limits of detection

DPMM

Dirichlet Process Mixture Model

DP

Dietary pattern

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Declaration of interests

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data availability

Data used in the analysis and reporting of this study are not publicly available, but can be made available upon request from the Project Viva study website: https://www.projectviva.org/for-investigators

References

  1. Amanatidis S, et al. , 2001. Comparison of two frequency questionnaires for quantifying fruit and vegetable intake. Public Health Nutr. 4, 233–9. [DOI] [PubMed] [Google Scholar]
  2. Bobb JF, et al. , 2018. Statistical software for analyzing the health effects of multiple concurrent exposures via Bayesian kernel machine regression. Environ Health. 17, 67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Carrico C, et al. , 2015. Characterization of Weighted Quantile Sum Regression for Highly Correlated Data in a Risk Analysis Setting. J Agric Biol Environ Stat. 20, 100–120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. DeLuca NM, et al. , 2025. Associations between self-reported consumption of foods and serum PFAS concentrations in a sample of pregnant women in the United States. Environ Res. 276, 121461. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Domingo JL, 2012. Health risks of dietary exposure to perfluorinated compounds. Environ Int. 40, 187–195. [DOI] [PubMed] [Google Scholar]
  6. Dzierlenga MW, et al. , 2021. The concentration of several perfluoroalkyl acids in serum appears to be reduced by dietary fiber. Environ Int. 146, 106292. [DOI] [PubMed] [Google Scholar]
  7. Egeghy PP, Lorber M, 2011. An assessment of the exposure of Americans to perfluorooctane sulfonate: a comparison of estimated intake with values inferred from NHANES data. J Expo Sci Environ Epidemiol. 21, 150–68. [DOI] [PubMed] [Google Scholar]
  8. Eick SM, et al. , 2023. Dietary predictors of prenatal per- and poly-fluoroalkyl substances exposure. J Expo Sci Environ Epidemiol. 33, 32–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Fabelova L, et al. , 2023. PFAS levels and exposure determinants in sensitive population groups. Chemosphere. 313, 137530. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Fawzi WW, et al. , 2004. Calibration of a semi-quantitative food frequency questionnaire in early pregnancy. Ann Epidemiol. 14, 754–62. [DOI] [PubMed] [Google Scholar]
  11. Fraser AJ, et al. , 2013. Polyfluorinated compounds in dust from homes, offices, and vehicles as predictors of concentrations in office workers’ serum. Environ Int. 60, 128–36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Fulay AP, et al. , 2018. Associations of the dietary approaches to stop hypertension (DASH) diet with pregnancy complications in Project Viva. Eur J Clin Nutr. 72, 1385–1395. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Gao X, et al. , 2021. Per- and polyfluoroalkyl substances exposure during pregnancy and adverse pregnancy and birth outcomes: A systematic review and meta-analysis. Environ Res. 201, 111632. [DOI] [PubMed] [Google Scholar]
  14. Hampson HE, et al. , 2024. Associations of dietary intake and longitudinal measures of per- and polyfluoroalkyl substances (PFAS) in predominantly Hispanic young Adults: A multicohort study. Environ Int. 185, 108454. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Haraldsdottir J, 1993. Minimizing error in the field: quality control in dietary surveys. Eur J Clin Nutr. 47 Suppl 2, S19–24. [PubMed] [Google Scholar]
  16. Hernandez E, et al. , 2024. Toddler dietary patterns from the INSIGHT randomized clinical trial comparing responsive parenting versus control: A latent class analysis. Obesity (Silver Spring). 32, 141–149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Hornung RW, Reed LD, 1989. Estimation of Average Concentration in the Presence of Nondetectable Values. Applied Occupational and Environmental Hygiene. 5(1), 46–51. [Google Scholar]
  18. Hu FB, et al. , 1997. Dietary fat intake and the risk of coronary heart disease in women. N Engl J Med. 337, 1491–9. [DOI] [PubMed] [Google Scholar]
  19. Huo X, et al. , 2023. Dietary and maternal sociodemographic determinants of perfluoroalkyl and polyfluoroalkyl substance levels in pregnant women. Chemosphere. 332, 138863. [DOI] [PubMed] [Google Scholar]
  20. Kalloo G, et al. , 2020. Exposures to chemical mixtures during pregnancy and neonatal outcomes: The HOME study. Environ Int. 134, 105219. [DOI] [PubMed] [Google Scholar]
  21. Kato K, et al. , 2011. Improved selectivity for the analysis of maternal serum and cord serum for polyfluoroalkyl chemicals. J Chromatogr A. 1218, 2133–7. [DOI] [PubMed] [Google Scholar]
  22. Lin PD, et al. , 2020. Dietary characteristics associated with plasma concentrations of per- and polyfluoroalkyl substances among adults with pre-diabetes: Cross-sectional results from the Diabetes Prevention Program Trial. Environ Int. 137, 105217. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Lindstrom AB, et al. , 2011. Polyfluorinated compounds: past, present, and future. Environ Sci Technol. 45, 7954–61. [DOI] [PubMed] [Google Scholar]
  24. McCullough ML, et al. , 2021. The Cancer Prevention Study-3 FFQ Is a Reliable and Valid Measure of Nutrient Intakes among Racial/Ethnic Subgroups, Compared with 24-Hour Recalls and Biomarkers. J Nutr. 151, 636–648. [DOI] [PubMed] [Google Scholar]
  25. Miller JW, Harrison MT, A simple example of Dirichlet process mixture inconsistency for the number of components. Advances in Neural Information Processing Systems, 2013, pp. 199–206. [Google Scholar]
  26. Miller JW, Harrison MT, 2017. Mixture Models With a Prior on the Number of Components. Journal of the American Statistical Association. 113, 340–256. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Odediran A, Obeng-Gyasi E, 2024. Association between Combined Metals and PFAS Exposure with Dietary Patterns: A Preliminary Study. Environments. 11. [Google Scholar]
  28. Oken E, et al. , 2015. Cohort profile: project viva. Int J Epidemiol. 44, 37–48. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Olsen GW, et al. , 2007. Half-life of serum elimination of perfluorooctanesulfonate, perfluorohexanesulfonate, and perfluorooctanoate in retired fluorochemical production workers. Environ Health Perspect. 115, 1298–305. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Pan S, et al. , 2024. Applications of mixture methods in epidemiological studies investigating the health impact of persistent organic pollutants exposures: a scoping review. J Expo Sci Environ Epidemiol. [Google Scholar]
  31. Papadopoulou E, et al. , 2019. Diet as a Source of Exposure to Environmental Contaminants for Pregnant Women and Children from Six European Countries. Environ Health Perspect. 127, 107005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Preston EV, et al. , 2022. Early-pregnancy plasma per- and polyfluoroalkyl substance (PFAS) concentrations and hypertensive disorders of pregnancy in the Project Viva cohort. Environ Int. 165, 107335. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Preston EV, et al. , 2020. Prenatal exposure to per- and polyfluoroalkyl substances and maternal and neonatal thyroid function in the Project Viva Cohort: A mixtures approach. Environ Int. 139, 105728. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Preston EV, et al. , 2018. Maternal Plasma per- and Polyfluoroalkyl Substance Concentrations in Early Pregnancy and Maternal and Neonatal Thyroid Function in a Prospective Birth Cohort: Project Viva (USA). Environ Health Perspect. 126, 027013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Radesky JS, et al. , 2008. Diet during early pregnancy and development of gestational diabetes. Paediatr Perinat Epidemiol. 22, 47–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Ramirez Carnero A, et al. , 2021. Presence of Perfluoroalkyl and Polyfluoroalkyl Substances (PFAS) in Food Contact Materials (FCM) and Its Migration to Food. Foods. 10. [Google Scholar]
  37. Rifas-Shiman SL, et al. , 2009. Dietary quality during pregnancy varies by maternal characteristics in Project Viva: a US cohort. J Am Diet Assoc. 109, 1004–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Rifas-Shiman SL, et al. , 2006. Changes in dietary intake from the first to the second trimester of pregnancy. Paediatr Perinat Epidemiol. 20, 35–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Rimm EB, et al. , 1992. Reproducibility and validity of an expanded self-administered semiquantitative food frequency questionnaire among male health professionals. Am J Epidemiol. 135, 1114–26; discussion 1127–36. [DOI] [PubMed] [Google Scholar]
  40. Rosato I, et al. , 2022. How to investigate human health effects related to exposure to mixtures of per- and polyfluoroalkyl substances: A systematic review of statistical methods. Environ Res. 205, 112565. [DOI] [PubMed] [Google Scholar]
  41. Sagiv SK, et al. , 2015. Sociodemographic and Perinatal Predictors of Early Pregnancy Per- and Polyfluoroalkyl Substance (PFAS) Concentrations. Environ Sci Technol. 49, 11849–58. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Schildroth S, et al. , 2021. Correlates of Persistent Endocrine-Disrupting Chemical Mixtures among Reproductive-Aged Black Women. Environ Sci Technol. 55, 14000–14014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Seshasayee SM, et al. , 2021. Dietary patterns and PFAS plasma concentrations in childhood: Project Viva, USA. Environ Int. 151, 106415. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Shu H, et al. , 2018. Temporal trends and predictors of perfluoroalkyl substances serum levels in Swedish pregnant women in the SELMA study. PLoS One. 13, e0209255. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Shu XO, et al. , 2004. Validity and reproducibility of the food frequency questionnaire used in the Shanghai Women’s Health Study. Eur J Clin Nutr. 58, 17–23. [DOI] [PubMed] [Google Scholar]
  46. Stephenson BJK, et al. , 2020a. Robust Clustering with Subpopulation-specific Deviations. J Am Stat Assoc. 115, 521–537. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Stephenson BJK, et al. , 2020b. Empirically Derived Dietary Patterns Using Robust Profile Clustering in the Hispanic Community Health Study/Study of Latinos. J Nutr. 150, 2825–2834. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Stephenson BJK, Willett WC, 2023. Racial and ethnic heterogeneity in diets of low-income adult females in the United States: results from National Health and Nutrition Examination Surveys from 2011 to 2018. Am J Clin Nutr. 117, 625–634. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Sultan H, et al. , 2023. Dietary per- and polyfluoroalkyl substance (PFAS) exposure in adolescents: The HOME study. Environ Res. 231, 115953. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Switkowski KM, et al. , 2019. Associations of protein intake in early childhood with body composition, height, and insulin-like growth factor I in mid-childhood and early adolescence. Am J Clin Nutr. 109, 1154–1163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Tian Y, et al. , 2018. Determinants of plasma concentrations of perfluoroalkyl and polyfluoroalkyl substances in pregnant women from a birth cohort in Shanghai, China. Environ Int. 119, 165–173. [DOI] [PubMed] [Google Scholar]
  52. Wang B, et al. , 2016. Perfluoroalkyl and polyfluoroalkyl substances in cord blood of newborns in Shanghai, China: Implications for risk assessment. Environ Int. 97, 7–14. [DOI] [PubMed] [Google Scholar]
  53. Woo Baidal JA, et al. , 2018. Association of vitamin E intake at early childhood with alanine aminotransferase levels at mid-childhood. Hepatology. 67, 1339–1347. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Wu SM, et al. , 2024. Derivation of outcome-dependent dietary patterns for low-income women obtained from survey data using a supervised weighted overfitted latent class analysis. Biometrics. 80, ujae122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Xie F, Xu Y, 2019. Bayesian Repulsive Gaussian Mixture Model. Journal of the American Statistical Association. 115, 187–203. [Google Scholar]
  56. Zhang Y, et al. , 2023. Association of Early Pregnancy Perfluoroalkyl and Polyfluoroalkyl Substance Exposure With Birth Outcomes. JAMA Netw Open. 6, e2314934. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Zhou T, et al. , 2024. Occurrence, fate, and remediation for per-and polyfluoroalkyl substances (PFAS) in sewage sludge: A comprehensive review. J Hazard Mater. 466, 133637. [DOI] [PubMed] [Google Scholar]
  58. Zhou W, et al. , 2019. Dietary intake, drinking water ingestion and plasma perfluoroalkyl substances concentration in reproductive aged Chinese women. Environ Int. 127, 487–494. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

Data Availability Statement

Data used in the analysis and reporting of this study are not publicly available, but can be made available upon request from the Project Viva study website: https://www.projectviva.org/for-investigators

RESOURCES