Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2024 Nov 21.
Published in final edited form as: Environ Sci Technol. 2023 Aug 24;57(46):18104–18115. doi: 10.1021/acs.est.3c00343

Towards advancing precision environmental health: Developing a customized exposure burden score to PFAS mixtures to enable equitable comparisons across population subgroups, using mixture item response theory

Shelley H Liu 1,*, Leah Feuerstahler 2, Yitong Chen 1, Joseph M Braun 3, Jessie P Buckley 4
PMCID: PMC11106720  NIHMSID: NIHMS1991101  PMID: 37615359

Abstract

Quantifying a person’s cumulative exposure burden to PFAS mixtures is important for risk assessment, biomonitoring, and reporting results to participants. However, different people may be exposed to different sets of PFAS due to heterogeneity in exposure sources and patterns. Applying a single measurement model for the entire population (e.g. by summing concentrations of all PFAS analytes), assumes that each PFAS analyte is equally informative to PFAS exposure burden for all individuals. This assumption may not hold if PFAS exposure sources systematically differ within the population. However, the socio-demographic, dietary and behavioral characteristics that underlie systematic exposure differences may not be known, or may be due to a combination of these factors. Therefore, we used mixture item response theory, an unsupervised psychometrics method, to develop a customized exposure burden scoring algorithm. This scoring algorithm ensures that PFAS burden scores can be equitably compared across population subgroups. We applied our methods to PFAS biomonitoring data from the United States National Health and Nutrition Examination Survey (2013-2018). Using mixture item response theory, we found that higher household income was associated with higher PFAS burden. Low-income and middle-income Asian Americans had significantly higher PFAS burden compared with non-Hispanic Whites and other race/ethnicity groups with similar incomes. However, some disparities were hidden when using summed PFAS concentrations. This work demonstrates that our summary PFAS burden metric accounting for sources of exposure variation may be a more fair and informative estimate of PFAS exposure.

Keywords: Per- and polyfluoroalkyl substances, chemical mixtures, psychometrics, item response theory, latent variables, algorithmic bias, fairness, precision environmental health

Introduction:

Per- and polyfluoroalkyl substances (PFAS) are used in products ranging from oil/water repellant textiles to flame retardants to food production.[13] Many PFAS are persistent chemicals, and exposure to PFAS mixtures is ubiquitous.[4] Exposure routes are primarily through diet and drinking water, and to a lesser extent through household items and dust. Recent literature calls for PFAS to be regulated as a single class,[5, 6] as PFAS may exert similar health effects. Further, a ban on all PFAS, a policy which is being considered in the European Union, would help avoid “regrettable substitution”, in which bans of one PFAS may lead to substitution of another PFAS that exerts similar detrimental health effects.[7] Accordingly, researchers and policymakers may be interested in quantifying an individual’s cumulative exposure burden to PFAS mixtures. However, quantifying our exposure burden to PFAS is challenging, because PFAS are a large class of thousands of chemicals, but current human biomonitoring studies routinely measure only ~10-20 PFAS analytes. We previously developed a PFAS exposure burden scoring algorithm[8] to quantify an individual’s latent (underlying) exposure burden to similar PFAS analytes, that may or may not have been assayed. Using item response theory (IRT), we developed a data-driven measurement model in which PFAS analytes were differentially informative to the overall PFAS exposure burden.

However, this approach uses the same scoring algorithm for all persons, ignoring potential sub-population differences. Moreover, the socio-demographic, dietary and behavioral characteristics that underlie systematic exposure differences may not be known, or may be due to a complex combination of these factors. In other words, there may be latent (hidden) subpopulations that differ systematically in their exposure patterns. Thus, customized scoring algorithms for PFAS exposure burden ensure that the quantification of PFAS exposure burden is equitable and informative for all people. To improve estimates of a person’s underlying exposure burden to the PFAS chemical class, we aimed to additionally account for between-person differences in diets, behaviors, and drinking water sources that may expose individuals to different sets of PFAS. We hypothesized that using a one-size-fits-all approach to quantify individual’s exposure burden can mask disparities in exposure burden across population subgroups, whereas accounting for differences in exposure sources could provide a more equitable approach to comparing PFAS burden across population subgroups.

Different population subgroups may be exposed to different sets of PFAS, perhaps due to differing combinations of dietary and behavioral habits and socio-demographic factors. If we assume a single measurement model for the entire population, we are assuming that a specific PFAS analyte is equally informative to PFAS exposure burden for all individuals (as is the assumption when using summed PFAS concentrations as the summary metric). This assumption may not hold if PFAS exposure sources systematically differ within the population, and corresponding, the PFAS exposure profiles differ systematically within a population. Taken together, we need to ensure that individual differences in estimated PFAS burden are truly reflective of the underlying differences in their exposure burden, and not artefacts of the scoring algorithm we used (e.g., that the algorithm is more appropriate for some individuals than others). This topic has been understudied in the environmental epidemiology literature, which primarily focuses on identifying key drivers of health outcomes within a mixture and their interactions, not on quantifying cumulative exposure burden. In order to advance precision environmental health, we need to optimally and equitably quantify exposure burden to mixtures. Future work to intervene on PFAS first requires an understanding of total exposure burden to PFAS, and ensuring that the summary metric used is fair and informative for all people.

Mixture item response theory (MixIRT) [914] combines item response theory with latent class analysis (LCA) to identify both a latent continuum of individual differences (e.g., a latent burden score) and latent classes. Mixture IRT methods have been applied to a wide range of topics in the psychology literature, including personality assessments,[12, 15] tobacco dependence criteria,[16] and assessment of risky behaviors.[10] In our setting, mixture IRT allows us to simultaneously identify latent subpopulations that differ systematically in their exposure patterns, identify unique measurement models of PFAS burden corresponding to the latent subpopulations, and provide data-driven algorithms to predict customized exposure burden that are based on a participant’s weighted likelihood of belonging to each latent subpopulation. To our knowledge, mixture IRT has not been applied in the environmental epidemiology literature.

We used national biomonitoring data from the US National Health and Nutrition Examination Survey (NHANES) to simultaneously identify latent subpopulations with distinct exposure profiles, and identify corresponding measurement models for PFAS exposure burden within each latent subpopulation. We then estimated customized (MixIRT) PFAS exposure burden scores, which are weighted by an individual’s likelihood of belonging to each subpopulation. We examine whether this customized exposure burden scoring algorithm, using mixture IRT, provides additional sensitivity to detect health effects and disparities across socio-demographic groups. We validate our models built on 2013-2016 NHANES data with the 2017-2018 NHANES sample. Lastly, we describe the unique sociodemographic and dietary characteristics of the latent subpopulations.

Methods:

Data:

We analyzed publicly available data from NHANES, a recurring cross-sectional survey of the noninstitutionalized civilian US population with details provided elsewhere.[17] To provide nationally representative estimates of biomonitoring in the US population, the NHANES oversamples Hispanic, non-Hispanic Black, and Asian individuals to ensure diversity in race/ethnicity. During each two-year cycle, a one-third subsample of participants aged 12 years and older are selected for measurement of PFAS concentrations in serum.[18, 19] The CDC laboratory quantified nine PFAS analytes in NHANES. We used eight PFAS analytes for our analysis; Sb-PFOA was excluded, because it has low detection frequency (10% in 2017-2018 NHANES), and our previous analysis suggested that Sb-PFOA yielded poor item fit in the item response theory models and was not informative to the burden score.[8] The remaining eight PFAS analytes analyzed studied here were: perfluorodecanoic acid (PFDeA), perfluorohexane sulfonic acid (PFHxS), 2-(N-methylperfluoroctanesulfonamido) acetic acid (Me-PFOSA-AcOH), perfluorononanoic acid (PFNA), perfluoroundecanoic acid (PFUA), n-perfluorooctanoic acid (n-PFOA), n-perfluorooctane sulfonic acid (n-PFOS), and perfluoromethylheptane sulfonic acid isomers (Sm-PFOS).[18, 20] NHANES provides survey-weights for each individual to enable generalization of findings to the overall US population. We used survey-weighted (using PFAS subsample weights) quantile cutoffs for categorizing each PFAS analyte.

We combined the 2013-2014 and 2015-2016 NHANES cycles. Prior to combining the two cycles, we first tested whether any PFAS analytes exhibited differential item functioning (DIF) by cycle. DIF detection investigates whether an item (e.g. PFAS analyte) functions the same way across manifest (known) groups. The presence of DIF suggests that for individuals with the same level of the latent trait, participants from different subgroups have a different probability of having a certain response to an item. Conceptually, DIF may exist across time because some PFAS analytes are phased out of production while others are phased into production. If no DIF was exhibited by NHANES cycle for any PFAS analyte, this would suggest that it was reasonable to combine the NHANES cycle years into a larger dataset. Using the R package “Lordif”,[21] we evaluated DIF for NHANES cycle, using ordinal logistic regression models with a McFadden’s pseudo R^2 change of 2% as the critical value.[22] We tested the 2013-2014 and 2015-2016 PFAS analytes by discretizing continuous PFAS analyte concentrations into quantiles using survey-weighted quantile cutoffs from the 2017-2018 (see our previous work;[8] Supplementary Table 1). As no PFAS analytes exhibited DIF across the 2013-2014 and 2015-2016 cycles, we chose to combine NHANES 2013-2016 cycles into a larger dataset, which we used for subsequent analyses. The NHANES 2017-2018 data was used as a validation sample.

MixIRT:

Mixture IRT allows for subpopulations to differ in regards to the processes that underlie their item responses. We conceptualize complex, multi-factorial exposure sources to be the processes that underlie a participant’s PFAS exposure profile. We hypothesized that different segments of the population may have different types of exposure sources, but how they are differentiated is largely unknown. We further hypothesized that we can identify different latent classes that have different systematic exposure to PFAS, and thus may need more than one measurement model to represent PFAS exposure burden in the population. Therefore, we tested whether a 2-class, 3-class, 4-class, 5-class or 6-class model would provide a better fit to the data, compared with assuming a single measurement model fitting the entire population. In this approach, rather than using known factors (e.g. age, race/ethnicity, sex), we assume that there are both known and unknown factors which systematically affect exposure to sets of PFAS, and this will be captured by the latent classes. In other words, we hypothesized that there may be more than one scoring algorithm that is appropriate to quantify PFAS burden in the population, which is due to an unknown array of demographic, dietary and behavior variables that necessitate having different scoring algorithms. Rather than using a different scoring algorithm for each age by sex by race/ethnicity by diet by drinking water source by etc. strata, we instead identified if there were latent subpopulations characterized by different scoring algorithms of PFAS burden.

We implemented mixture IRT, using mixture graded response models, which are used for ordinal data. Suppose that we have measured m = 1, …, M PFAS analytes. Each ordinal PFAS analyte can take on Km possible values (categories), allowing for differing numbers of quantiles for different PFAS analytes (mixed item types). For example, we can have a Km=2 category ordinal PFAS analyte, with k = 1 representing non-detect and k = 2 representing detects. Or we can have a Km=4 category analyte, with k = 1 representing PFAS analyte concentration <25th percentile, k = 2 representing 25th – 50th percentile, k = 3 representing 50th – 75th percentile, and k = 4 representing > 75th percentile. Further, suppose the population contains c = 1, …, C latent classes, which are assumed to be mutually exclusive and jointly exhaustive. Each latent class c is characterized by a latent class proportion, pi_c.

For a given subject i, xi,m represents their polytomous level on the mth PFAS analyte, such that xi,m can take on Km possible values. θic denotes the latent PFAS burden for individual i in latent class c. The mixture graded response IRT model relates an individual’s observed level of a PFAS analyte xi,m, to an individual’s latent PFAS burden θi via the following formulae.

The probability of scoring at or above the given response option k, given the level of θ and latent class c is:

Pr(xi,mk|c,θic)=exp{αmc(θicβm,kc)}1+exp{αmc(θicβm,kc)}whereβ2<β3<<βKm.

The probability of scoring at the current response option, k, is:

Pr(xi,m=k|c,θi)=Pr(xi,mk|c,θic)Pr(xi,mk+1|c,θic),

where Prxi,m1c,θic)=1 and Prxi,m=Kmc,θic)=Prxi,mKmc,θic) because Km is the highest response category. θic represents the latent PFAS burden for individual i in class c, and is often assumed to follow a standard normal distribution, with mean 0 and standard deviation 1. mc denotes the discrimination parameter for PFAS analyte m, in class c, which indicates how well the analyte can distinguish between participants with very similar latent PFAS burdens. Analytes with high discrimination parameters provide more information about latent PFAS burden differences across participants, while analytes with low discrimination parameters do not provide much information and may not need to be included in the scale. Item discrimination estimates can theoretically range from negative infinity to infinity, but will usually be positive as negative discriminations may indicate that the item should be reverse coded, because it is negatively correlated with the latent variable. Functionally, increasing item discrimination values above approximately 3.50 has a negligible effect on the model predictions.

βm,k denotes the threshold parameter corresponding to PFAS analyte m for response k. The thresholds are on the same scale as θi, that is, a standard normal. For an analyte m, there will be (Km-1) thresholds, which represent the cut points between the Km analyte quantiles. In the case of a binary (detect/non-detect) analyte, the threshold refers to how high an individual’s latent PFAS burden needs to be to have a 0.5 probability of having a particular PFAS analyte detected. In the case of an ordinal analyte, the threshold refers to how high an individual’s latent PFAS burden needs to be to have a 0.5 probability of having that particular observed quantile of exposure, or a higher observed quantile of exposure, on that specific analyte. Specifically, if an analyte is coded into quartiles, there will be 3 threshold parameters. The first threshold is the latent trait score at which an individual has a 50/50 chance of having exposure 1 vs. 2 or 3 or 4 for that analyte. The second threshold is the latent trait score at which an individual has a 50/50 chance of having exposure 1 or 2 vs. 3 or 4 for that analyte; the third threshold is the latent trait score at which an individual has a 50/50 chance of having exposure 1 or 2 or 3 vs. 4 for that analyte.

Because IRT models have no inherent scale, practitioners need to set the scale in order to interpret the exposure burden scores. In a standard IRT model (without the mixture component), the scale is typically set by restricting the mean of the latent variable to be zero, and variance to be 1 (e.g. latent variable has a standard normal distribution). In mixture IRT models, a scale can be set by restricting the distribution within each latent class to be standard normal.

In order to identify the number of latent classes, we first set the PFAS exposure burden distribution within each class to be standard normal. We henceforth refer to this as the unconstrained model, which assumes that the latent mean and variance are fixed across the latent classes, for identifiability purposes. The item thresholds are allowed to vary across latent classes. We identified the optimal number of latent classes using established metrics of information criterion based on the log-likelihood: Akaike information criterion, Bayesian information criterion, sample-size adjusted BIC (SABIC). These criteria differ based on the penalty that is imposed for model complexity. Lower values indicated better fit.

Mixture IRT was implemented using “mirt: A Multidimensional Item Response Theory Package for the R Environment” [23] using the standard EM algorithm with fixed quadrature. We used fixed starting values that were internally determined by the “mirt” package. We evaluated the model performance of using fixed starting values versus using random starting values generated from different seeds. Comparing the log-likelihoods, we found that the models using fixed starting values had the lowest −2 log-likelihoods. When the log-likelihoods were nearly identical between models with fixed versus random starting values, the group membership was also nearly identical. We also tried imposing priors on the item discrimination and threshold parameters, using lognormal and normal priors, respectively. However, when we imposed priors and used fixed starting values, the “mirt” package generated errors. Hence, we did not use priors on the final models.

Identifying anchor item(s) for MixIRT:

The unconstrained model imposes a hindrance on interpretability, as in practice, we would not expect each latent class to have the same mean and variance for the PFAS exposure burden. Thus, instead of using this restriction to set the scale for the latent trait, we instead sought to identify anchor item(s) that would enable us to set the same scale across latent classes, in order to have scale comparability. By using one or more anchor item(s), in which we set the parameters of that item to be the same across latent classes, we can set the scale of the PFAS exposure burden without needing to impose restrictions on the mean and variance of each latent class.

We tested each PFAS analyte as a potential anchor item. We used a previously reported approach, in which for each PFAS analyte, we conducted a likelihood ratio test comparing the −2logl for the model that constrains parameters for one PFAS analyte to be the same across the four latent classes, to the unconstrained model. If the likelihood ratio test was not significant for that PFAS analyte (alpha = 0.05), then it suggested that the fit of the model did not substantially differ from that of the unconstrained model, and hence that PFAS analyte could be used as the anchor item. After identifying the anchor item(s), we ran another MixIRT model with four latent classes, in which item parameters for the anchor item(s) were constrained across the latent classes and were allowed to vary for all other items.

We then calculated exposure burden scores that are weighted by probability of class membership in the different latent classes, using weighted expected a posteriori (EAP) scores.

For an IRT model with a single latent class, the EAP score for person i may be calculated from the following equation:

θ^i,EAP=Σq=1QθqL(θq|xi)p(θq)Σq=1QL(θq|xi)p(θq)

where q indexes Q quadrature points such that θq values represent a series of latent burden values (e.g., a uniform sequence from −4 to 4), Lθxi is the log-likelihood of a given θ value given person i’s response pattern xi, and pθq is the density of the prior distribution of ability (e.g., standard normal) at point θq. For a mixture IRT model, an EAP score may be calculated as a weighted average of class-specific EAP scores, where individual-level weights equal each participants’ posterior probability of belonging to each latent class.

Associations with cardio-metabolic outcomes:

We sought to determine whether PFAS exposure burden was significantly associated with cardio-metabolic outcomes. These particular cardiometabolic outcomes were selected because they were previously used in our manuscript demonstrating the application of IRT to calculate PFAS burden scores.[8] Here, C-reactive protein was excluded because it was not available in for the 2013-2014 cycle of NHANES. The cardiometabolic outcomes were: total cholesterol (mg/dL), high-density lipoprotein (mg/dL), non-high-density lipoprotein (mg/dL), triglycerides (mg/dL), systolic blood pressure (mmHg) and diastolic blood pressure (mmHg). Non-high-density lipoprotein (non-HDL) was calculated by subtracting high-density lipoprotein (HDL) cholesterol from total cholesterol. Regression models were adjusted for sex, age (in years), race/ethnicity (Mexican American, other Hispanic, non-Hispanic white, non-Hispanic black, non-Hispanic Asian, other race including multi-racial) as a proxy for racism/structural inequity,[24] ratio of family income to poverty, and body-mass-index (kilogram per meter squared). The family income to poverty ratio was calculated by dividing family income by poverty guidelines specific to the NHANES survey year, family size and geographic location, and is henceforth referred to as the socioeconomic status (SES) index.[19, 24, 25]

Characterizing differences in socio-demographic, dietary and behavioral variables:

We used the non-parametric Kruskal Wallis test to identify significant differences in predictor variables between the latent classes. Responses of “don’t know”, or “refused” were treated as missing data, as our goal was to characterize the classes, we did not multiply impute missing data. We also calculated the percent difference between median levels of the PFAS exposure across race/ethnicity group (using non-Hispanic White as reference) and across socioeconomic status (SES are classified into SES<1, SES between 1-2, SES between 2-3, SES between 3-4 and SES>=4 using the family income to poverty ratio; SES<1 was used as the reference, indicating family income below the poverty line). We used the equation (median_group – median_ref)/|(median_ref)|*100 to calculate the percent difference.

Results:

Table 1 presents the unweighted summary statistics for the combined 2013-2014 and 2015-2016 NHANES cycles (N=3915). 51.9% of the participants were female; 36.0% were Non-Hispanic White, 20.9% were Non-Hispanic Black, 16.9% were Mexican-American, 11.4% were Other Hispanic, 10.8% were Non-Hispanic Asian, and the remainder were other race, including multi-racial. The median age of the sample was 43 [interquartile range (IQR) 25, 61] years.

Table 1: Summary statistics of the sample (2013-2014 and 2015-2016 NHANES).

In the 2013-2014 NHANES sample, there were N=1954 participants with complete data on all PFAS analytes. After excluding pregnant women (N=15), N=1939 participants remained. In the 2015-2016 NHANES sample, there were N=1993 with complete data on all PFAS analytes. After excluding pregnant women (N=17), N=1976 remained. Summary statistics below are reported for N=3915 participants with complete data on all PFAS analytes.

Percentage (%) % Missing
n 3915
Race/ethnicity 0
Mexican-American 16.9
Other Hispanic 11.4
Non-Hispanic White 36.0
Non-Hispanic Black 20.9
Non-Hispanic Asian 10.8
Other Race, including Multi-Racial 3.9
Sex 0
Male 48.1
Female 51.9
Median (interquartile range)
Age 43.00 [25.00, 61.00] 0
Family income to poverty ratio 1.93 [1.02, 3.82] 8.94
PFDeA (ng/mL) 0.20 [0.07, 0.30] 0
PFHxS (ng/mL) 1.20 [0.70, 2.20] 0
Me-PFOSA-AcOH (ng/mL) 0.07 [0.07, 0.20] 0
PFNA (ng/mL) 0.60 [0.40, 1.00] 0
PFUA (ng/mL) 0.07 [0.07, 0.20] 0
n-PFOA (ng/mL) 1.60 [1.00, 2.50] 0
n-PFOS (ng/mL) 3.30 [1.90, 5.90] 0
Sm-PFOS (ng/mL) 1.40 [0.70, 2.50] 0
Body mass index 27.6 [23.5, 32.3] 0.97
Total cholesterol (mg/dL) 179.0 [154.00, 208.0] 0.03
HDL cholesterol (mg/dL) 51.0 [42.0, 62.0] 0.03
non-HDL (mg/dL) 125.0 [100.0, 155.0] 0.03
Triglycerides (reference) (mg/dL) 113.0 [73.0, 179.0] 0.13
Systolic BP (mmHg) 118.7 [109.3, 131.3] 2.76
Diastolic BP (mmHg) 68.0 [60.0, 75.3] 2.76

Calibrating the mixture IRT scoring algorithm for the customized MixIRT PFAS burden score:

We used mixture graded response models to develop a customized exposure burden score for each participant. Supplemental Table 1 presents the quantile cutoffs for the PFAS analytes, which were used to discretize the PFAS analyte concentrations into ordinal data. We then determined the optimal number of latent classes was 4 (BIC = 54709.58) (Supplemental Table 2), indicating that there were four measurement models. Using likelihood ratio tests to compare the fits of the unconstrained model to the model that constrains parameters for one PFAS to be the same across the four latent classes (Supplemental Table 3), only PFNA had a non-significant p-value; thus, we used PFNA as the anchor item. We then ran the final model, which used mixture graded response model with 4 latent classes, constrained the item parameters for PFNA to be the same for each latent class, and allowed item parameters for all other PFAS analytes to vary. Item parameters (Supplemental Table 4) were presented for each latent class. Because PFNA was identified as the anchor item, the item parameters for PFNA remained the same for each class, enabling estimation of exposure burden to be on the same scale for all participants. The final PFAS exposure burden scores were calculated using weighted EAP, so that the burden score for each participant is weighted by their probability of membership in each latent class. Thus, this approach assumed that there were four measurement models to quantify PFAS burden for participants in the sample, and a participants’ burden score was a weighted combination of their burden scores estimated from each measurement model, weighted by how likely each measurement model fit them. For a detailed explanation of the different measurement models identified, see Supplementary Appendix 1. This exposure burden score, derived from the mixture IRT approach, is henceforth referred to as: “MixIRT PFAS burden score”.

Calibrating the overall IRT scoring algorithm for the PFAS burden score:

To enable comparisons with the mixture IRT scoring algorithm for estimating PFAS burden, we also calculated an overall IRT model using the same quantile cutoffs for each PFAS analyte (Supplemental Table 1). This approach assumes that there is a single measurement model to quantify PFAS exposure burden for the entire sample of participants. The item parameters for the overall IRT model are presented in Supplementary Table 5. This exposure burden score is henceforth referred to as “PFAS burden score”.

Scoring algorithm using summed PFAS concentrations:

We also calculated summed PFAS concentrations of the eight PFAS analytes to enable comparisons with the two previously defined PFAS burden scores. This is henceforth called “summed PFAS concentrations”.

The MixIRT PFAS burden score was highly correlated with the PFAS burden score (Pearson’s ρ = 0.95). The two burden scores were moderately correlated with summed PFAS concentrations (Pearson’s ρ were 0.69 for the MixIRT PFAS burden score, and 0.74 for the PFAS burden score).

Comparing scoring algorithms and sensitivity to detect associations between PFAS and cardiometabolic outcomes:

We assessed adjusted associations of PFAS with cardiometabolic outcomes in the 2013-2016 NHANES, comparing effect sizes between the three scoring algorithms (MixIRT PFAS burden score, PFAS burden score, and summed PFAS concentrations). Both burden scores were significantly associated with total cholesterol, non-HDL and HDL (all p<0.001), but not with triglycerides, systolic or diastolic blood pressure (Figure 1A; Supplementary Table 6). An IQR increase in the MixIRT PFAS burden score was associated with 4.2 [2.5, 5.8] mg/dL, 2.8 [1.2, 4.5] mg/dL and 1.3 [0.7, 1.9] mg/dL, higher total cholesterol, non-HDL cholesterol, and HDL cholesterol, respectively. Meanwhile, an IQR increase in the PFAS burden score was associated with 5.7 [3.6, 7.8] mg/dL, 4.4 [2.3, 6.5] mg/dL and 1.3 [0.5, 2.0] mg/dL higher total cholesterol, non-HDL cholesterol, and HDL cholesterol, respectively. Summed PFAS concentrations were not associated with any cardiometabolic outcomes.

Figure 1:

Figure 1:

Figure 1:

Adjusted effect size and 95% CI corresponding to a 1 IQR increase in mixture IRT (MixIRT) PFAS burden, PFAS burden, and summed concentrations across cardiometabolic outcomes. MixIRT PFAS burden scores were calculated from the mixture IRT model with PFNA as an anchor item. PFAS burden scores were calculated from the overall IRT model. Associations were adjusted for gender, age, race/ethnicity, ratio of family income to poverty and BMI.

A: Adjusted effect size and 95% CI corresponding to a 1 IQR increase in PFAS burden and in summed concentrations across cardiometabolic outcomes for NHANES 2013-2014 and 2015-2016 cycles.

B: Adjusted effect size and 95% CI corresponding to a 1 IQR increase in PFAS burden and in summed concentrations across cardiometabolic outcomes for NHANES 2017-2018 cycle.

We used the 2017-2018 cycle of NHANES as a validation dataset, using the mixIRT and overall IRT models that were developed using the 2013-2016 data. The MixIRT PFAS burden score was highly correlated with the PFAS burden score (Pearson’s ρ = 0.94). There was low-moderate correlation between the PFAS burden scores and the summed PFAS concentrations (Pearson’s ρ = 0.33, 0.32, respectively). Results were similar to the 2013-2016 findings, showing strong replication (Figure 1B; Supplementary Table 7). Notably, we observed larger magnitudes of association in the 2017-2018 cycle for total cholesterol and HDL compared to estimates in the 2013-2016 data across for the MixIRT PFAS burden score, PFAS burden score, and summed PFAS concentrations. For total cholesterol, effect sizes corresponding to a 1 IQR increase in PFAS burden were 7.2 [4.6, 9.8] mg/dL, 10.0 [6.5, 13.5] mg/dL, and 2.4 [0.5, 4.2] mg/dL, respectively. For non-HDL, effect sizes corresponding to a 1 IQR increase in PFAS burden were 6.0 [3.4, 8.6] mg/dL, 9.0 [5.5, 12.5] mg/dL, and 1.8 [0.0, 3.7] mg/dL, respectively. We also detected a significant association between MixIRT PFAS burden score and HDL and DBP: Effect sizes of 1.2 [0.3, 2.0] mg/dL for HDL and 0.8 [0.0, 1.7] mmHg for DBP corresponded to a 1 IQR increase in the MixIRT PFAS burden score. These associations were not significant using PFAS burden score or summed PFAS concentrations. No associations were detected with triglycerides or SBP for either PFAS burden score or summed PFAS concentrations.

Comparing scoring algorithms and sensitivity to detect disparities in PFAS exposure across socio-demographic groups:

Using 2013-2016 NHANES, we assessed disparities in PFAS exposure across race/ethnicity groups. We calculated the percent difference between median levels of the MixIRT PFAS burden score in each race/ethnicity group compared to the median score in non-Hispanic Whites (Figure 2A and Supplementary Table 8A). Mexican American, Other Hispanic and Other Race had significantly lower MixIRT PFAS burden scores compared to Non-Hispanic White (−88.5% for Mexican American, −30.8% for Other Hispanic, and −34.6% for Other Race), while Non-Hispanic Asian had significantly higher MixIRT PFAS burden (88.5% higher compared to Non-Hispanic White). We detected same direction but smaller effect sizes using summed PFAS (−33.0% for Mexican American, −19.7% for Other Hispanic, −16.9% for Other Race and 8.1% for Non-Hispanic Asian compared to Non-Hispanic White). Raw values of the median and IQR of the scores are presented in Supplementary Table 9A.

Figure 2:

Figure 2:

Figure 2:

Comparision of MixIRT PFAS burden scores and summed PFAS concentrations across race/ethnicity groups compared to those metrics in non-Hispanic Whites, in the NHANES 2013-2016 cycles and 2017-2018 cycle (validation data). The non-parametric Wilcoxon Mann Whitney test was used to compare distributions of MixIRT PFAS burden scores and summed PFAS concentrations for each race/ethnicity group against those of non-Hispanic Whites. Significant differences were denoted by asterisks (* p<0.05; ** p<0.01; *** p<0.001). We present the percent difference between median MixIRT PFAS burden scores in each race/ethnicity group compared to the median scores in non-Hispanic Whites, using the equation (median_group − median_nhw)/|(median_nhw)|*100. As an example, median MixIRT PFAS burden was −0.49 for Mexican Americans and −0.26 for non-Hispanic Whites. Thus, the percent difference was −0.49 − (−0.26)/|−0.26| * 100 = −88.5%. This was similarly done for the summed PFAS concentrations. As an example, median summed PFAS concentrations was 7.17 ng/mL for Mexican Americans and 10.70 ng/mL for non-Hispanic Whites; thus, the percent difference was (7.17 − 10.70)/|10.70|*100 = −33%. The percent differences and raw values of MixIRT PFAS burden score and summed PFAS concentrations are presented in Supplementary Table 8 and Supplementary Table 9.

A: Disparities in MixIRT PFAS burden scores and summed PFAS concentrations across race/ethnicity groups compared to those metrics in non-Hispanic Whites, in the NHANES 2013-2016 cycles

B: Disparities in MixIRT PFAS burden scores and summed PFAS concentrations across race/ethnicity groups compared to those metrics in non-Hispanic Whites, in the 2017-2018 cycle (validation data).

We then validated these findings using the 2017-2018 cycle of NHANES (Figure 2B, Supplementary Table 8B and Supplementary Table 9B). Mexican American and Other Race were found to have significantly lower MixIRT PFAS burden compared to Non-Hispanic White (−96.3% and −59.3% respectively), and Non-Hispanic Asian had significantly higher MixIRT PFAS burden (25.9%), but we did not detect significant difference for Other Hispanic in the 2017-2018 NHANES cycle. We detected similar differences using summed PFAS concentrations (−36.4% for Mexican American Vs, Non-Hispanic White, −17.6% for Other Hispanic Vs. Non-Hispanic White, and −25.4% for Other Race Vs. Non-Hispanic White) except that no significant difference was seen when comparing Non-Hispanic Asian to Non-Hispanic White. Raw values of the median and IQR of the scores are presented in Supplementary Table 9B.

We then assessed disparities in PFAS exposure across socio-economic groups in all three scoring algorithms, in both 2013-2016 NHANES and 2017-2018 NHANES as validation. Across both datasets, increasing household income was monotonically associated with significantly higher PFAS exposure for all scoring algorithms. Figure 3 shows the percent difference across the SES groups. Supplementary Table 10 and Supplementary Table 11 present the percent differences with p values and the raw values of median and IQR of the scores.

Figure 3:

Figure 3:

Figure 3:

Disparities in MixIRT PFAS burden scores and summed PFAS concentrations across socioeconomic status (SES) categories, in the NHANES 2013-2016 cycles and 2017-2018 cycle (validation data). (SES are classified into SES<1, SES between 1-2, SES between 2-3, SES between 3-4 and SES>=4 using the family income to poverty ratio. We used SES<1, indicating the family income is below the poverty line, as the reference.) The non-parametric Wilcoxon Mann Whitney test was used to compare distributions of MixIRT PFAS burden scores and summed PFAS concentrations for SES 1-2, 2-3, 4 and above against those of SES less than 1. Significant differences were denoted by asterisks (* p<0.05; ** p<0.01; *** p<0.001). We present the percent difference between median MixIRT PFAS burden scores and summed PFAS concentrations in each SES category compared to those of SES<1.The percent differences and raw values of MIxIRT PFAS burden score and summed PFAS concentrations are presented in Supplementary Tables 10 and Supplementary Tables 11.

A: Disparities in MixIRT PFAS burden scores and summed PFAS concentrations across SES categories (using SES<1 as the reference) in the NHANES 2013-2016 cycles.

B: Disparities in MixIRT PFAS burden scores and summed PFAS concentrations across SES categories (using SES<1 as the reference) in the NHANES 2017-2018 cycle.

To assess whether there is confounding effect between race/ethnicity groups and SES groups, we also calculated the percent difference between median levels of the MixIRT PFAS burden score of Non-Hispanic Asian compared to Non-Hispanic White stratified by SES categories (low income, SES less than 1; middle income, SES between 1 and 3; higher income, SES equal to and greater than 3). (Figure 4 and Supplementary Table 12). Non-Hispanic Asians had consistent higher MixIRT PFAS burden compared to Non-Hispanic White for those with low income (SES<1) and middle income (SES within 1-3) both in the 2013-2016 NHANES and the 2017-2018 NHANES cycles. We did not detect significant percent difference for those SES>3. Using summed PFAS concentrations, we detected the same direction of disparities except that we found Non-Hispanic Asian had lower summed PFAS concentration compared to Non-Hispanic White in the SES>3 group. Supplementary Table 13 shows the raw values of median and IQR of the scores. Supplementary Table 14 presents the median age (IQR) of Non-Hispanic Asian and Non-Hispanic White stratified by SES categories. Non-Hispanic Asian did not have significantly older age compared to Non-Hispanic White, which suggested our finding that non-Hispanic Asian had higher PFAS burden in the low income (SES<1) and middle income (SES in 1-3) groups were not confounded by age.

Figure 4.

Figure 4

Figure 4

A: Percent difference between median MixIRT PFAS burden scores and median summed PFAS concentrations of Non-Hispanic Asians versus those of Non-Hispanic Whites, stratified by SES categories in the 2013-2016 NHANES cycles. Low income indicates family income to poverty ratio less than 1; middle income indicates family income to poverty ratio greater or equal to 1 and less than 3, and higher income indicates family income to poverty ratio greater or equal to 3. SES categories were collapsed in this way to ensure reasonable sample size of non-Hispanic Asians in each category. The percent differences and raw values of MIxIRT PFAS burden scores and summed PFAS concentrations are presented in Supplementary Tables 12 and Supplementary Tables 13.

B: Percent difference between median MixIRT PFAS burden scores and median summed PFAS concentrations of Non-Hispanic Asians versus those of Non-Hispanic Whites, stratified by SES categories in the 2017-2018 NHANES cycle. SES categories were collapsed in this way to ensure reasonable sample size of non-Hispanic Asians in each category. Low income indicates family income to poverty ratio (SES) less than 1; middle income indicates family income to poverty ratio greater or equal to 1 and less than 3, and higher income indicates family income to poverty ratio greater or equal to 3. Significant differences were denoted by asterisks (* p<0.05; ** p<0.01; *** p<0.001).

Secondary analysis of the mixture IRT scoring algorithm: Comparisons using most likely class for each participant

Although the MixIRT PFAS burden scores are weighted by a participant’s likelihood of belonging to the four latent classes, we also sought to characterize differences in the subpopulations based on the most likely class for each participant. The proportion of individuals in each class, based on their most likely class, were: Class 1, 16.9%; Class 2, 33.1%; Class 3, 23.5%; and Class 4, 26.5%. Concentrations of individual PFAS differed across the four classes (Supplemental Table 15). Class 1 had the highest median concentrations of PFUA, PFDeA, n-PFOS, and PFAS burden scores and the highest summed PFAS concentrations (median 12.54 [IQR: 5.80, 20.60] ng/mL). Class 2 had the highest median concentrations of n-PFOA. Class 3 had lower concentrations of all PFASs. Class 4 had higher median concentrations of PFHxS and Sm-PFOS.

Supplemental Table 16 presents the differences in socio-demographic variables across the four latent subpopulations based on the most likely class for each participant. Class 1 was predominantly female (64.8%) with a higher proportion of non-Hispanic Asian (32.4%) participants. Class 2 had the highest median SES (2.26) whereas Class 3 had the lowest SES (1.58), highest proportion with very low food security (11.9%), and highest proportion of households ever receiving emergency food (15.7%) and FS benefit (50.8%). Class 4 was predominantly male (63.5%) with the highest proportion of non-Hispanic White (47.1%) participants and oldest median age (48 years).

We also found notable differences between classes with respect to diet (Supplemental Table 17). The classes differed significantly on frequency of milk consumption in the past 30 days (p<0.001), with the highest proportion (40.8%) of participants in Class 4 drinking milk once a day or more. Class 4 also had the highest proportion of participants who were a regular milk drinker for most of all or their life, including childhood (40.7% in Class 4, versus 29.7 – 35.4% in Class 1-3; p<0.001). There was a significant difference in prevalence of drinking tap water across the 4 classes (p=0.002), with Class 4 having the highest proportion of drinking tap water (82.9% in Class 4, versus 75.4%-79.6% in Class 1-3). Lastly, the classes differed on their fish consumption in the past 30 days, with Class 1 having the highest proportion of those who consumed fish (79.7% vs. 56.9-68.8%).

Discussion:

We used unsupervised item response theory methods to address how to optimally and equitably quantify an exposure burden score to PFAS mixtures that is fair and informative for all people. Rather than assuming a single algorithm, our approach estimated a person’s total exposure burden to PFAS mixtures, accounting for the fact that different people have different diets and behaviors that expose them to different sets of PFAS chemicals in complex ways. We are the first to use novel data science approaches (mixture item response theory) to address this. Our findings show that by quantifying PFAS burden using mixture item response theory, we could detect more precise associations of PFAS burden with health outcomes. Further, we could uncover disparities in PFAS burden across race/ethnicity groups, which are hidden if we assume a single scoring algorithm is sufficient for the entire population. For example, we found that Asian Americans have significantly higher PFAS burden compared with non-Hispanic Whites and all other race/ethnicity groups, even when controlling for socio-economic status. However, some disparities were hidden if we used summed PFAS concentrations. Our work suggests that biomonitoring and risk assessment may consider using a summary PFAS exposure metric that accounts for exposure heterogeneity, so that the summary metric is fair and informative for all people. Indeed, our method may be applied when implementing clinical recommendations from a National Academies of Science, Engineering and Medicine panel that recommended additional health screening for people with elevated serum/plasma PFAS concentrations [26].

Using the 2013-2016 NHANES data and the mixture IRT approach, we identified four latent subpopulations characterized by distinct PFAS exposure profiles and measurement models of exposure burden. We linked these scores by identifying an anchor item, PFNA, for which parameters were held constant across the four latent subpopulations, so the estimated EAP scores could be compared. Note that the factor scores were weighted by the probability of subgroup membership, and that the subpopulations were not identified with any variables other than the PFAS concentrations. However, these subgroups differed based on demographic and dietary characteristics. Further, different PFAS were informative to PFAS burden for different subgroups.

We found that PFAS burden scores calculated using mixture IRT had increased sensitivity to detect associations with health outcomes and disparities across demographic groups, compared with PFAS burden scores calculated using an overall IRT model for the entire sample. In the 2013-2016 NHANES, we found that for both the overall and subpopulation-specific scoring algorithms, PFAS burden was significantly associated with total cholesterol, non-HDL and HDL. However, only using the subpopulation-specific scoring algorithm was PFAS burden associated with diastolic blood pressure. Summed PFAS concentrations were not associated with any cardiometabolic biomarker. Using the 2017-2018 NHANES as the validation dataset, we found the same pattern of associations for the subpopulation-specific PFAS burden (associations with total cholesterol, non-HDL, HDL and diastolic blood pressure). Meanwhile, burden scores for the overall IRT approach and the summed PFAS concentration approach were only associated with total cholesterol and non-HDL. In both the model building (2013-2016 NHANES) and validation (2017-2018 NHANES) datasets, PFAS burden calculated using any of the three scoring algorithms was not associated with systolic blood pressure or triglycerides. Taken together, this suggests that using a mixture IRT scoring approach for PFAS burden may have increased sensitivity to detect health effects. Previous studies have reported associations between single PFAS analytes with serum lipids.[27, 28] Some previous studies have found null associations between individual PFAS and blood pressure,[29, 30] while others have found significant positive associations of PFAS with systolic and diastolic blood pressure.[31]

Reducing environmental health disparities and promoting environmental justice is a priority area of the US National Institute of Environmental Health Sciences.[32] Here, we found that using a mixture IRT approach, we uncovered disparities in PFAS burden among race/ethnicity groups as a proxy for racism/structural inequity,[24] that were not seen when using an overall IRT approach or the summed PFAS concentrations approach. Specifically, we found that non-Hispanic Asians had significantly higher PFAS burden than non-Hispanic Whites. One way to understand how the mixture IRT approach can reveal patterns beyond an overall IRT approach is to assume that there do exist subpopulations who are exposed to different combinations of PFAS. In such a case, fitting an overall IRT model would conceptually result in a type of average of the models that would be fit to these subpopulations (if the subpopulations were known exactly). Although the mixture IRT approach may not perfectly identify all relevant subpopulations, it may provide a more realistic representation of the major groups that are affected, and thus provide a higher validity model of the various ways in which PFAS burden manifests. In the context of this study, we found that Non-Hispanic Asians had higher PFAS burden compared with non-Hispanic Whites. We found that Mexican Americans had lower PFAS burden compared with non-Hispanic Whites, and this was consistent across all scoring algorithms for PFAS burden. Previous studies have found that PFAS analyte concentrations vary across race/ethnicity[33] and socio-economic status.[34]

Our findings quantifying PFAS burden using mixture IRT provided stronger effect estimates to detect associations with health outcomes and disparities, echoes similar findings in the psychology literature, which found that using mixture IRT increased the predictive validity of personality test scores compared with using the same IRT model for all participants.[12, 15] Specifically, when comparing self-rated personality test scores of participants, the mixture IRT test scores had higher correlation with how supervisors, colleagues and subordinates rated those same participants.[15] The authors suggested that this was because mixture IRT could uncover different approaches to how subgroups fill out a questionnaire, and thus different measurement models were needed for the different subgroups. In our setting, this implies that different population subgroups have higher PFAS burden through different exposure patterns.

Our approach has limitations. Findings may be sensitive to the chosen anchor item(s). In lieu of expert scientific knowledge to select an anchor item, we followed established psychometric practices for identifying an anchor item. We conducted a sensitivity analysis that found that very similar differences between latent classes for the model with the anchor item, and for the unconstrained model. The parameter estimates, particularly for the discrimination parameter, should not be over-interpreted, as values above approximately 3.50 are functionally very similar. We tried placing priors on the discrimination and threshold parameters, using lognormal and normal distributions, respectively, but had poorer model fit and errors in the algorithm, so we did not use priors on any parameters. Also, mixture IRT models can be hampered by identifying global vs. local solutions. To guard against this, we fit the mixture models several times with various seeds, but found very similar classification results across these solutions. Another limitation is that we relied on dietary questionnaire data from NHANES; there is evidence of systematic misreporting errors on diet questionnaires, such as systematic underreporting of energy intake, but not all foods are misreported equally.[35, 36] Therefore, how the latent classes differ with regards to dietary intake should be interpreted with limitations of dietary questionnaire data in mind. Disparities in PFAS burden across SES levels has been previously reported in the literature.[37] Reasons for this disparity may be complex: People with higher SES may have more fish consumption, more takeout food consumption, among others. Tracing exposure sources is complex but a necessary next step.[38]

Strengths of our mixture IRT approach include that it accounts for the fact that there is more than one algorithm to quantify latent PFAS burden within a diverse population. The estimated PFAS burden scores are weighted EAP scores, like a customized exposure burden scoring algorithm. For each individual, the PFAS burden scores are weighted by their probability of belonging in each latent subgroup. While we characterize the subgroup profiles for each person based on their most likely group, people can have for instance 60% probability of belonging to class 1, 20% probability of belonging to class 2, 10% probability of belonging to class 3 and 4. The weighted EAP score takes this into account. Notably, we showed that using this approach, we can identify significant associations with health outcomes as found using other methods of quantifying PFAS burden. However, for some outcomes, like HDL and diastolic blood pressure, in which the literature has found mixed findings, we show that the mixture IRT approach for calculating PFAS burden is consistent in its findings for the model building and validation datasets. This approach can be useful for researchers who want to assess associations between PFAS mixture burden and health outcomes, or investigate socio-demographic disparities in PFAS burden, in addition to the de facto approach of using summed PFAS concentrations. Another benefit is that we can build the model, and use that model to predict exposure burden for new participants’ data. Mixture IRT enabled us to identify different measurement models for different subgroups, accounting for the fact that sources of exposure to PFAS are complex and somewhat unknown. Our approach differs from other latent class analyses of chemical exposure profiles,[39] in that we are not only characterizing distinct profiles of exposure mixtures, but we are also identifying whether the subpopulations are characterized by different measurement models. In practice, this is important because different groups may have very different exposure sources. Hence, a analyte may be differentially informative to a participants’ overall PFAS burden for different groups.

If we only used latent class analysis (LCA), we could have identified latent subpopulations with different exposure profiles to PFAS mixtures. However, LCA does not allow us to quantify PFAS exposure burden within each class. Importantly, using MixIRT, we could place the PFAS exposure burdens within each latent class onto the same scale, by identifying and using appropriate anchor item(s), which are item(s) whose parameters are invariant across classes.[25] We could also determine whether the most informative PFAS analytes differ across latent classes, if some items are salient (e.g. individuals tend to have higher exposure levels) for some latent classes compared to others, and if a analyte provides the most information at a different range of exposure burden for one latent class compared with another. MixIRT is closely related to differential item functioning (DIF), which refers to when items do not function the same way for manifest (known) subgroups, such as age and sex.[13] However, the subgroups may not be known in advance, hence the need for MixIRT. Previous applications of this approach in the psychology literature suggests that use of mixture IRT applied to personality tests increases the predictive validity of the test scores, compared with fitting an overall IRT model to the entire sample.[12, 15]

In the environmental epidemiology literature, approaches to quantify exposure burden to chemical mixtures, independent of pre-specific health outcomes, are lacking. Most statistical approaches for chemical mixtures are supervised, focusing on identifying key drivers, interactions, and overall mixture effects specific to a pre-specified health outcome. Those methods analyze relationships between the specific chemical analytes measured in the sample and the health outcome. They treat the specific chemical analytes that were chosen to be measured in the sample as if they are all the possible chemical exposures to that chemical class. However, for biomonitoring and report-back, we may not be as interested in the specific set of PFASs that a specific laboratory assays, but rather we are more interested in the underlying exposure burden to similar PFAS analytes, both measured and unmeasured, as they can act in aggregate to affect health. For example, currently the EU CONTAM Panel is regulating PFAS based on the sum of four PFASs (PFOA, PFNA, PFHxS and PFOS).[40] However, the summed approach is a special case of a latent variable model known as a parallel factor model [41], but may not be optimal to quantify PFAS exposure burden for the entire population. Our findings suggest that this summed approach could mask associations with some health outcomes and demographic disparities in exposure burden. More customized methods of quantifying PFAS exposure burden, such as the mixture IRT approach, provides another tool towards regulations of PFAS. Further, quantifying latent exposure burden will also enable researchers and policymakers to identify disparities in exposure burden, towards ensuring environmental justice and identifying solutions to protect those at increased risk.[42, 43]

To our knowledge, this is the first application of mixture IRT in the environmental epidemiology literature. To better inform PFAS exposure assessment towards policy, regulation, remediation, it is important to have informative and fair estimates of our total exposure burden to PFAS chemicals, that accounts for complex heterogeneity in exposure sources and patterns. An important aspect of risk assessment is exposure assessment – in particular, understanding how exposure differs in vulnerable groups; we have shown that the mixture IRT approach can help with exposure characterization, revealing differences in PFAS burden across population groups that are hidden when we just use summed PFAS concentrations. In order to advance precision environmental health, we need to optimally and equitably quantify exposure burden to mixtures. Future work to intervene on PFAS first requires a good understanding of total exposure burden to PFAS, and ensuring that the summary metric used is fair and informative for all people. Mixture IRT is a promising approach to quantify latent exposure burden to chemical mixtures that accounts for complex heterogeneity in exposure sources and patterns, towards quantifying a customized exposure burden index for an individual.

Supplementary Material

Supplementary Material

Supplementary Table 1: Survey weighted quantile cutoffs for frequently detected PFAS analytes (calculated from NHANES 2017-2018).

Supplementary Table 2: Comparison of model fit to identify optimal number of latent classes.

Supplementary Table 3: Comparison of model fits to identify anchor item(s), for 4 latent classes (the optimal number of classes determined in Supplementary Table 2.

Supplementary Table 4: Parameter estimates for the mixture model using PFNA as an anchor item.

Supplementary Appendix 1: Discrimination parameters and item thresholds.

Supplementary Table 5: Parameter estimates by PFAS analyte for the graded response model (overall IRT model), using data from NHANES 2013-2014 and 2015-2016 cycles.

Supplementary Table 6: Adjusted effect size and 95% CI corresponding to a 1 IQR increase in MixIRT PFAS burden score, PFAS burden score, and summed PFAS concentration across cardiometabolic outcomes for NHANES 2013-2014 and 2015-2016 cycles.

Supplementary Table 7: Adjusted effect size and 95% CI corresponding to a 1 IQR increase MixIRT PFAS burden score, PFAS burden score, and summed PFAS concentration across cardiometabolic outcomes for NHANES 2017-2018 cycle.

Supplementary Table 8: The percent differences of MixIRT PFAS burden score and summed PFAS concentrations across race/ethnicity groups compared to those metrics in non-Hispanic Whites, in the NHANES 2013-2016 cycles and 2017-2018 cycle (validation data).

Supplementary Table 9: MixIRT PFAS burden score, PFAS burden score and summed PFAS concentration stratified by race/ethnicity group.

Supplementary Table 10: The percent differences of MixIRT PFAS burden score and summed PFAS concentrations across SES categories (using SES<1 as the reference), in the NHANES 2013-2016 cycles and 2017-2018 cycle (validation data).

Supplementary Table 11: MixIRT PFAS burden score, PFAS burden score and summed PFAS concentrations stratified by SES category.

Supplementary Table 12: Percent difference between median levels of the MixIRT PFAS burden score of Non-Hispanic Asian compared to Non-Hispanic White stratified by SES categories.

Supplementary Table 13: MixIRT PFAS burden score, PFAS burden score and summed PFAS concentrations across race/ethnicity group stratified by SES category.

Supplementary Table 14: Median age (IQR) of Non-Hispanic Asian compared to Non-Hispanic White stratified by SES categories.

Supplementary Table 15: Comparison of PFAS concentrations across latent classes.

Supplementary Table 16: Comparison of socio-demographic factors across latent classes.

Supplementary Table 17: Comparison of diet variables across latent classes.

Synopsis:

We used mixture item response theory methods to address how to optimally and equitably quantify an exposure burden score to PFAS mixtures that is fair and informative for all people.

Funding sources:

S.H.L. was supported by National Institute for Environmental Health Sciences (NIEHS) R03ES033374 and National Institute of Child Health and Human Development (NICHD) K25HD104918. J.P.B. was supported by NIEHS R03ES033374, R01ES030078, and R01ES033252. J.M.B was supported by NIEHS R01 ES032386. Joseph Braun has been compensated for serving as an expert witness for plaintiffs in litigation over PFAS contaminated drinking water.

References

  • 1.Buck RC, Franklin J, Berger U, Conder JM, Cousins IT, de Voogt P, Jensen AA, Kannan K, Mabury SA, van Leeuwen SP: Perfluoroalkyl and polyfluoroalkyl substances in the environment: terminology, classification, and origins. Integr Environ Assess Manag 2011, 7(4):513–541. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Cousins IT, DeWitt JC, Gluge J, Goldenman G, Herzke D, Lohmann R, Miller M, Ng CA, Scheringer M, Vierke L et al. : Strategies for grouping per- and polyfluoroalkyl substances (PFAS) to protect human and environmental health. Environ Sci Process Impacts 2020, 22(7):1444–1460. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Sunderland EM, Hu XC, Dassuncao C, Tokranov AK, Wagner CC, Allen JG: A review of the pathways of human exposure to poly- and perfluoroalkyl substances (PFASs) and present understanding of health effects. J Expo Sci Environ Epidemiol 2019, 29(2):131–147. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Fenton SE, Ducatman A, Boobis A, DeWitt JC, Lau C, Ng C, Smith JS, Roberts SM: Per- and Polyfluoroalkyl Substance Toxicity and Human Health Review: Current State of Knowledge and Strategies for Informing Future Research. Environ Toxicol Chem 2021, 40(3):606–630. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Kwiatkowski CF, Andrews DQ, Birnbaum LS, Bruton TA, DeWitt JC, Knappe DRU, Maffini MV, Miller MF, Pelch KE, Reade A et al. : Scientific Basis for Managing PFAS as a Chemical Class. Environmental Science & Technology Letters 2020, 7(8):532–543. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Braun JM: Enhancing Regulations to Reduce Exposure to PFAS-Federal Action on “Forever Chemicals”. N Engl J Med 2023, 388(21):1924–1926. [DOI] [PubMed] [Google Scholar]
  • 7.Trager R: Efforts underway in Europe to ban PFAS compounds. In: Chemistry World. Royal Society of Chemistry; 2021. [Google Scholar]
  • 8.Liu SH, Kuiper JR, Chen Y, Feuerstahler L, Teresi J, Buckley JP: Developing an Exposure Burden Score for Chemical Mixtures Using Item Response Theory, with Applications to PFAS Mixtures. Environ Health Perspect 2022, 130(11):117001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.De Ayala RJ, Santiago SY: An introduction to mixture item response theory models. J Sch Psychol 2017, 60:25–40. [DOI] [PubMed] [Google Scholar]
  • 10.Finch WH, Pierson EE: A Mixture IRT Analysis of Risky Youth Behavior. Front Psychol 2011, 2:98. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Jasper F, Nater UM, Hiller W, Ehlert U, Fischer S, Witthoft M: Rasch scalability of the somatosensory amplification scale: a mixture distribution approach. J Psychosom Res 2013, 74(6):469–478. [DOI] [PubMed] [Google Scholar]
  • 12.Maij-de Meij AM, Kelderman H, van der Flier H: Fitting a Mixture Item Response Theory Model to Personality Questionnaire Data: Characterizing Latent Classes and Investigating Possibilities for Improving Prediction. Applied Psychological Measurement 2008, 32(8):611–631. [Google Scholar]
  • 13.Maij-de Meij AM, Kelderman H, van der Flier H: Improvement in Detection of Differential Item Functioning Using a Mixture Item Response Theory Model. Multivariate Behav Res 2010, 45(6):975–999. [DOI] [PubMed] [Google Scholar]
  • 14.Sen S, Cohen AS: Applications of Mixture IRT Models: A Literature Review. Measurement: Interdisciplinary Research and Perspectives 2019, 17(4):177–191. [Google Scholar]
  • 15.Egberink IJL, Meijer RR, Veldkamp BP: Conscientiousness in the workplace: Applying mixture IRT to investigate scalability and predictive validity. Journal of Research in Personality 2010, 44:232–244. [Google Scholar]
  • 16.Muthen B, Asparouhov T: Item response mixture modeling: Application to tobacco dependence criteria. Addictive Behaviors 2006, 31:1050–1066. [DOI] [PubMed] [Google Scholar]
  • 17.Zipf G, Chiappa M, Porter KS, Ostchega Y, Lewis BG, Dostal J: National health and nutrition examination survey: plan and operations, 1999–2010. Vital Health Stat 2013, 56:1–37. [PubMed] [Google Scholar]
  • 18.CDC: Laboratory Procedure Manual: Perfluoroalkyl and polyfluoroalkyl substances. In. Organic Analytical Toxicology Branch, Division of Laboratory Sciences. [Google Scholar]
  • 19.CDC: National Health and Nutrition Examination Survey Data, 2017-2018. In. Edited by Centers for Disease Control and Prevention NCfHS. Hyattsville, MD: U.S. Department of Health and Human Services. [Google Scholar]
  • 20.Chalmers RP: mirt: A Multidimensional Item Response Theory Package for the R Environment. Journal of Statistical Software 2012, 48:1–29. [Google Scholar]
  • 21.Choi SW, Gibbons LE, Crane PK: lordif: An R Package for Detecting Differential Item Functioning Using Iterative Hybrid Ordinal Logistic Regression/Item Response Theory and Monte Carlo Simulations. J Stat Softw 2011, 39(8):1–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Lameijer CM, van Bruggen SGJ, Haan EJA, Van Deurzen DFP, Van der Elst K, Stouten V, Kaat AJ, Roorda LD, Terwee CB: Graded response model fit, measurement invariance and (comparative) precision of the Dutch-Flemish PROMIS Upper Extremity V2.0 item bank in patients with upper extremity disorders. BMC Musculoskeletal Disorders 2020, 21:170. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Chalmers RP: mirt: A multidimensional item response theory package for the R environment. Journal of Statistical Software 2012, 48(6):1–29. [Google Scholar]
  • 24.Kaufman JD, Hajat A: Confronting Environmental Racism. Environmental health perspectives 2021, 129(5):51001–51001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Cho S-J, Cohen AS, Bottge B: Detecting Intervention Effects Using a Multilevel Latent Transition Analysis with a Mixture IRT Model. Psychometrika 2013, 78(3):576–600. [DOI] [PubMed] [Google Scholar]
  • 26.National Academies of Sciences E, and Medicine: Guidance on PFAS Exposure, Testing, and Clinical Follow-Up. 2022(Washington, DC: The National Academies Press.). [PubMed] [Google Scholar]
  • 27.Fan Y, Li X, Xu Q, Zhang Y, Yang X, Han X, Du G, Xia Y, Wang X, Lu C: Serum albumin mediates the effect of multiple per-and polyfluoroalkyl substances on serum lipid levels. Environ Pollut 2020, 266(Pt 2):115138. [DOI] [PubMed] [Google Scholar]
  • 28.Jain RB, Ducatman A: Associations between lipid/lipoprotein levels and perfluoroalkyl substances among US children aged 6-11 years. Environ Pollut 2018, 243(Pt A):1–8. [DOI] [PubMed] [Google Scholar]
  • 29.Canova C, Zeddi MJ, Barbieri G, Gion M, Dapra F, Russo F, Fletcher T, Pitter G: Perfluoroalkyl substances and blood pressure in exposed young population in the Veneto Region, Italy. European Journal of Public Health 2020, 30(5). [Google Scholar]
  • 30.Lin PD, Cardenas A, Hauser R, Gold DR, Kleinman KP, Hivert MF, Calafat AM, Webster TF, Horton ES, Oken E: Per- and polyfluoroalkyl substances and blood pressure in pre-diabetic adults-cross-sectional and longitudinal analyses of the diabetes prevention program outcomes study. Environ Int 2020, 137:105573. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Pitter G, Zare Jeddi M, Barbieri G, Gion M, Fabricio ASC, Dapra F, Russo F, Fletcher T, Canova C: Perfluoroalkyl substances are associated with elevated blood pressure and hypertension in highly exposed young adults. Environ Health 2020, 19(1):102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Environmental health disparties and environmental justice [https://www.niehs.nih.gov/research/supported/translational/justice/index.cfm] [Google Scholar]
  • 33.Park SK, Peng Q, Ding N, Mukherjee B, Harlow SD: Determinants of per- and polyfluoroalkyl substances (PFAS) in midlife women: Evidence of racial/ethnic and geographic differences in PFAS exposure. Environmental Research 2019, 175:186–199. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Buekers J, Colles A, Cornelis C, Morrens B, Govarts E, Schoeters G: Socio-economic status and health: Evaluation of human biomonitored chemical exposure to per- and polyfluorinated substances across status. Int J Environ Res Public Health 2018, 15:2818. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Ravelli MN, Schoeller DA: Traditional Self-Reported Dietary Instruments Are Prone to Inaccuracies and New Approaches Are Needed. Front Nutr 2020, 7:90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Gibney M, Allison D, Bier D, Dwyer J: Uncertainty in human nutrition research. Nature Food 2020, 1(5):247–249. [Google Scholar]
  • 37.Tyrrell J, Melzer D, Henley W, Galloway TS, Osborne NJ: Associations between socioeconomic status and environmental toxicant concentrations in adults in the USA: NHANES 2001-2010. Environment International 2013, 59:328–335. [DOI] [PubMed] [Google Scholar]
  • 38.Menichetti G, Ravandi B, Mozaffarian D, Barabási A-L: Machine learning prediction of the degree of food processing. Nature Communications 2023, 14(1):2312. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Carroll R, White AJ, Keil AP, Meeker JD, McElrath TF, Zhao S, Ferguson KK: Latent classes for chemical mixtures analyses in epidemiology: an example using phthalate and phenol exposure biomarkers in pregnant women. J Expo Sci Environ Epidemiol 2020, 30(1):149–159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.EFSA, Schrenk D, Bignami M, Bodin L, Chipman JK, del Maxo J: Risk to human health related to the presence of perfluoroalkyl substances in food. EFSA Journal 2020, 18(9):e06223. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.McNeish D, Wolf MG: Thinking twice about sum scores. Behav Res Methods 2020, 52(6):2287–2305. [DOI] [PubMed] [Google Scholar]
  • 42.Varshavsky JR, Zota AR, Woodruff TJ: A novel method of calculating potency-weighted cumulative phthalates exposure with implications for identifying racial/ethnic disparities among U.S. reproductive-aged women in NHANES 2001-2012. Environ Sci Technol 2016, 50(19):10616–10624. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Morello-Frosch R, Zuk M, Jerrett M, Shamasunder B, Kyle AD: Understanding the cumulative impacts of inequalities in environmental health: implications for policy. Health Aff 2011, 30(5):879–887. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material

Supplementary Table 1: Survey weighted quantile cutoffs for frequently detected PFAS analytes (calculated from NHANES 2017-2018).

Supplementary Table 2: Comparison of model fit to identify optimal number of latent classes.

Supplementary Table 3: Comparison of model fits to identify anchor item(s), for 4 latent classes (the optimal number of classes determined in Supplementary Table 2.

Supplementary Table 4: Parameter estimates for the mixture model using PFNA as an anchor item.

Supplementary Appendix 1: Discrimination parameters and item thresholds.

Supplementary Table 5: Parameter estimates by PFAS analyte for the graded response model (overall IRT model), using data from NHANES 2013-2014 and 2015-2016 cycles.

Supplementary Table 6: Adjusted effect size and 95% CI corresponding to a 1 IQR increase in MixIRT PFAS burden score, PFAS burden score, and summed PFAS concentration across cardiometabolic outcomes for NHANES 2013-2014 and 2015-2016 cycles.

Supplementary Table 7: Adjusted effect size and 95% CI corresponding to a 1 IQR increase MixIRT PFAS burden score, PFAS burden score, and summed PFAS concentration across cardiometabolic outcomes for NHANES 2017-2018 cycle.

Supplementary Table 8: The percent differences of MixIRT PFAS burden score and summed PFAS concentrations across race/ethnicity groups compared to those metrics in non-Hispanic Whites, in the NHANES 2013-2016 cycles and 2017-2018 cycle (validation data).

Supplementary Table 9: MixIRT PFAS burden score, PFAS burden score and summed PFAS concentration stratified by race/ethnicity group.

Supplementary Table 10: The percent differences of MixIRT PFAS burden score and summed PFAS concentrations across SES categories (using SES<1 as the reference), in the NHANES 2013-2016 cycles and 2017-2018 cycle (validation data).

Supplementary Table 11: MixIRT PFAS burden score, PFAS burden score and summed PFAS concentrations stratified by SES category.

Supplementary Table 12: Percent difference between median levels of the MixIRT PFAS burden score of Non-Hispanic Asian compared to Non-Hispanic White stratified by SES categories.

Supplementary Table 13: MixIRT PFAS burden score, PFAS burden score and summed PFAS concentrations across race/ethnicity group stratified by SES category.

Supplementary Table 14: Median age (IQR) of Non-Hispanic Asian compared to Non-Hispanic White stratified by SES categories.

Supplementary Table 15: Comparison of PFAS concentrations across latent classes.

Supplementary Table 16: Comparison of socio-demographic factors across latent classes.

Supplementary Table 17: Comparison of diet variables across latent classes.

RESOURCES