Abstract
Objective
To assess the feasibility of applying machine learning (ML) methods to imputation in the Medical Expenditure Panel Survey (MEPS).
Data Sources
All data come from the 2016–2017 MEPS.
Study Design
Currently, expenditures for medical encounters in the MEPS are imputed with a predictive mean matching (PMM) algorithm in which a linear regression model is used to predict expenditures for events with (donors) and without (recipients) data. Recipient events and donor events are then matched based on the smallest distance between predicted expenditures, and the donor event's expenditures are used as the recipient event's imputation. We replace linear regression algorithm in the PMM framework with ML methods to predict expenditures. We examine five alternatives to linear regression: Gradient Boosting, Random Forests, Extreme Random Forests, Deep Neural Networks, and a Stacked Ensemble approach. Additionally, we introduce an alternative matching scheme, which matches on a vector of predicted expenditures by sources of payment instead of a single total expenditure prediction to generate potentially superior matches.
Data Collection
Study data is derived from a large federal survey.
Principal Findings
ML algorithms perform better at both prediction and matching imputation than Ordinary Least Squares (OLS), the most common prediction algorithm used in PMM. On average, the Stacked Ensemble approach that combines all the ML algorithms performs best, improving expenditure prediction R2 by 108% (0.156 points) and final imputation R2 by 227% (0.397 points). Matching on a prediction vector also improves alignment of sources of payments between donor and recipient events.
Conclusions
ML algorithms and an alternative matching scheme improve the overall quality of expenditure PMM imputation in the MEPS. These methods may have additional value in other national surveys that currently rely on PMM or similar methods for imputation.
Keywords: imputation, machine learning, medical expenditures, MEPS, predictive mean matching
What is known on this topic
Accurate imputations are important to ensure data quality.
Predictive Mean Matching imputation is widely used for imputation by a number of important surveys.
Machine learning potentially offers improvements to predictive accuracy over current estimation methods.
What this study adds
We apply emerging machine learning techniques to imputation problems.
We demonstrate that using machine learning algorithms in a Predictive Mean Matching framework leads to better imputations.
We show that the flexibility of machine learning techniques can provide options for further matching improvements.
1. INTRODUCTION
Item nonresponse is a perennial issue for survey data collection, stemming from respondents failing to completely answer survey questions due to refusal or inability to answer. Item nonresponse has become more prevalent over time across a number of large public surveys, 1 increasing the importance of effectively addressing the issue. Statistical agencies providing public data from such surveys face important choices in how to release data suffering from item nonresponse and must balance statistical accuracy concerns with ease of use for external users, who have varying levels of technical experience. One option is to provide directly imputed data for the end user, either through multiple or single imputations. For example, the National Health Interview Survey provides multiple imputation files for income, which end users can use to construct estimates using specialized software routines. 2 Other large surveys, such as the Current Population Survey (CPS), produce a single set of data in which missing values have been imputed with single values using a statistical method such as hot‐deck imputation. 3
Absent the provision of imputed data, external researchers may try to avoid the issue by relying on complete case analysis, wherein any observations or items with missing data are dropped from the analysis, an approach that can lead to significantly biased results. 4 , 5 Additionally, high‐quality imputations are important as the data are used to create predictions and estimates that have bearing on real‐world decisions. Therefore, developing better methods to compensate for missing data is an active area of research, and a variety of different approaches have been proposed.
For federal surveys that disseminate public use files, hot‐deck imputation to replace the missing values with plausible values is a common practice. 6 Many different methods fall into the hot‐deck imputation family, all of which share the common property of replacing a responding unit's missing value with a value already existing in the data from a similar donor unit. These methods mainly differ on how the donors are matched to the recipient data, and better matches between recipients and donors lead to better overall imputations.
Predictive mean matching (PMM) is a commonly used hot deck approach due to its ability to generate close matches. 4 , 7 PMM produces a matching score by weighting donor and recipient characteristics based on how strongly they predict the variable with missing values. Variables in PMM are usually weighted by ordinary least squares (OLS). OLS has the advantage of being a well‐known estimator with well‐studied properties. However, machine learning (ML) methods have demonstrated better results than OLS in prediction. 8
ML approaches have several attributes that result in better predictions than OLS. In contrast with OLS, which can only combine predictors linearly, ML algorithms can search more complex function spaces, including interactions and nonlinearities, to find the best predictions. ML methods are also more flexible with input data and can estimate models with very sparse data or when there are more predictors than observations, unlike OLS which assumes the predictor matrix is full rank.
The predictive advantage of ML over OLS has been demonstrated in a number of contexts, including medical expenditures. 9 Previous application of ML methods to healthcare spending has mainly involved risk adjustment using claims data, and only focuses on prediction of expenditures. 10 , 11 , 12 , 13 , 14 , 15 , 16 This literature has often found improvements in predictive accuracy of ML algorithms over OLS, however, in some instances, OLS performs better. In considering algorithms it is important to note that simpler models, such as OLS, may generalize better to out‐of‐sample data and thus be preferable to more complex ML models. 17 Additionally, while direct prediction of expenditures is important, better predictions do not necessarily lead to better imputations in a PMM context. We extend this existing literature to show that better predictions can also lead to higher‐quality matches for imputation.
Additionally, we extend the PMM literature by examining an alternative matching scheme. Standard PMM methods predict and match on one value. In instances where multiple related values are to be imputed, potentially better imputations can be made by jointly predicting and matching on all values. For example, when matching on a variable with multiple sub‐components, single‐dimension matching could result in matches that have misaligned sub‐components. Matching on the main component and the sub‐components makes it more likely that donors and recipients are similar across all dimensions.
We study the potential for ML to improve the quality of PMM used in processing event‐level expenditure data collected in the Medical Expenditure Panel Survey (MEPS), the primary source of official United States statistics on individual and family‐level medical spending. Considering the wide use of the MEPS by both researchers and policymakers, any improvements to expenditure data directly translate into better research and more accurately informed public policies, with impacts on a variety of real‐world outcomes. The event‐level expenditures, from which national estimates of healthcare spending are derived, are the central purpose of the MEPS, making nonresponse a particular concern. Large proportions of missing expenditures would result in inadequate data to produce the estimates for which the MEPS was established. Therefore, it is necessary to impute expenditures to ensure as complete and accurate expenditure data as possible.
In the context of PMM, assuming better generalization and predictions lead to better matches, the replacement of OLS with more sophisticated ML algorithms should lead to better imputation performance. Our goal is to improve the quality of MEPS imputation through two innovations. First, we examine if replacing OLS predictions with better ML predictions creates higher‐quality imputations in the PMM framework. Second, we assess if matching on a prediction vector of total expenditure and its components results in better matches than simply matching on a single prediction of total expenditure alone.
In the remainder of this paper, we first describe the MEPS data and the current imputation process to provide the background for our proposed changes. Next, we describe the ML algorithms we test as replacement candidates for OLS and how we assess the algorithms. We then outline our proposed change to use multiple sources of payment predictions in the matching process. Next, we present the results of these changes. Finally, we discuss the limitations and implications of our results.
2. METHODS
2.1. Data
We train our models using MEPS data from 2016 to 2017. 18 The MEPS is a nationally representative sample of the U.S. civilian noninstitutionalized population. Households are interviewed five times over a two‐year period with detailed questions on health care utilization and expenditures, health insurance, and health status, as well as a wide variety of social, demographic, and economic characteristics for everyone in the household. In all interviews, respondents are asked to enumerate every visit to a medical care provider since the previous interview and provide charge and expenditure data for those visits.
For events that occur at an emergency department (ED), a hospital (HS), a hospital outpatient department (OP), or medical office (MV) interviewers conduct a follow‐back survey, the Medical Provider Component (MPC), with a sample of providers to obtain billing and other administrative data relating to those events. The MPC is the primary source of expenditure data in the MEPS because household‐reported data are generally incomplete. 19 Where MPC data are available, they replace household‐reported data. The MPC data also form the donor pool for the MEPS imputation process (see Appendix S1) for these types of health care events.
The MPC collects total expenditure and expenditures for 10 potential sources of payment: Out‐of‐pocket, Medicare, Medicaid, Private Insurance, Veteran's Administration, Tricare/CHAMPUS, other federal, other state and local sources, worker's compensation, and other sources. Imputing values for the total and each potential source is the primary goal of the MEPS imputation process. For this work, payments from other Federal Sources are too rare to be estimated and are thus combined with other sources.
The extent of expenditure imputation varies across event types. Figure 1 shows trends in the proportion of events that are imputed each year in the MEPS from 1999 to 2017. Imputation rates in 2017 for ED, HS, and OP events range between 32% and 37%, while expenditures are imputed for between 52% and 63% of MV events each year. This disparity is due to the much larger number of household‐reported MV events and lower sampling rate for the MPC due to costs and other data collection limitations, however within MV collection more complex (and thus more expensive) MV events are prioritized, reducing the imputation rate for expensive visits. Overall, combining these four event types, just over 55% of events were imputed in 2017. Given these relatively high rates of missingness, even small improvements in imputation methods can have a substantial impact on the accuracy of the MEPS expenditure data.
FIGURE 1.

Trends in the proportion of medical events with imputed total expenditure values in the MEPS. Medical events are divided by type: emergency department visits (ED), inpatient hospitalizations (HS), office‐based visits (MV), and outpatient visits in hospital settings (OP).
To assess the performance of the current MEPS imputation and its alternatives, we select only events with complete MPC data. Though there are potential errors in this data arising from various causes, such as transcription errors during data collection, these data are similar to administrative data and represent the most accurate measures available. We split these event data into two groups, with 80% of the events forming the training data for the ML algorithms (analogous to the donor pool in the PMM imputation process) and 20% forming an out‐of‐sample test set (akin to the recipient pool).
Appendix Table 1 describes the donor and recipient data for each event type. HS events have the smallest training and validation datasets with 2,592 and 667 events respectively, while MV events have the largest with 74,794 and 18,610 events respectively. As is expected for a random 80–20 split, total expenditures between all the donor and recipient data are comparable. For predictors, we select the variables that are currently used in the MEPS imputation process for events of these types in which all sources of payments are known. A full list of predictors used can be found in Appendix Table 2.
2.2. Current MEPS event expenditure imputation process
The MEPS uses a PMM process to impute missing expenditure values for each encounter with the medical system. This process first estimates a predictive model of mean total expenditures using an OLS regression trained on donor events. Predicted expenditures for both donor and recipient events are then generated from this model. Next, recipients are matched to donors with the most similar predicted total expenditure and the recipients are assigned the donor's total expenditure value as their imputed total expenditure. Additionally, any missing source of payment expenditures on the recipient record is filled with the corresponding expenditures from the donor event. Therefore, it is important to align donors and recipients by sources of payment in addition to total expenditure to ensure realistic payment patterns. Since standard PMM only offers one dimension upon which to align (total expenditure), matching is conducted within groups with similarly reported sources of payment availability (termed XGROUPs). For example, recipients with private insurance and no out‐of‐pocket expenditures are only matched with donors of that group, and not with a donor with private insurance and out‐of‐pocket expenditures. Further details can be found in the Appendix S1.
We seek to improve on the existing imputation process in two ways. First, we attempt to generate better predictions (and subsequently better matches) for the expenditure targets by replacing the OLS algorithm in PMM with more predictive ML algorithms. Second, we try to improve alignment of sources of payment between donors and recipients by predicting and matching the full vector of payment sources and total expenditures instead of only total expenditures, thereby eliminating XGROUPs as explained below.
Broadly, we first assess ML and OLS predictions by comparing predictions of total expenditure in the hold‐out test events from models estimated with the training events. Next, we evaluate if those predictions result in better imputations for the hold‐out test events by applying PMM matching using each algorithms' prediction of total expenditure. Finally, we examine if PMM matching can be improved by using a vector of predictions from the algorithm that shows the best performance in predicting and imputing just total expenditure.
2.3. Machine learning algorithm candidates
Prediction error consists of two parts: the bias, or how closely a model approximates the underlying data generating function, and the variance, or how well a model generalizes to out‐of‐sample data. The bias of OLS can be reduced to an extent by including additional predictors. However, this added complexity tends to increase the variance of the prediction error. ML generates better predictions than OLS by allowing more complex models (thereby reducing bias) and/or increasing the generalizability of the model (thereby reducing variance). To replace the OLS predictions, we assess the predictive ability of four individual algorithms and a stacked ensemble algorithm.
The first algorithm we examine is the random forest. Random forests (RF) are sets of regression trees, a prediction method in which values are predicted based on successive splits in the values of the predictors which optimize the similarity of the resulting groups' outcomes, until a stopping criterion is reached. Each regression tree is trained on a sub‐sample of the training data and then is aggregated into a final prediction by the RF. The non‐linear partitioning creates complex individual trees which result in lower bias, while the final aggregation reduces the variance to which each individual tree is prone.
Next, we consider variations of the RF: Extreme Random Forest (XRF) and Gradient Boosting Machines (GBM). With XRF the partitioning is done at random instead of optimal predictor values. This tends to reduce the variance of the error without substantially increasing the bias. 20 GBM builds its regression trees sequentially, with each simple tree trained on the error of the prior tree, in contrast with RF which aggregates the trees in the final step. This typically results in more bias than RF models, but lower variance.
We also examine Deep Neural Networks (DNN) and a Stacked Ensemble (SE) of all the algorithms. DNN consist of sets of nodes, which weight and aggregate their inputs, apply an activation function, and pass the resulting output to be inputs for subsequent nodes. At its simplest, a neural network with a single node and no activation function is equivalent to a linear regression. By layering multiple sets of nodes, the DNN creates vastly more complex models, and can become universal function approximators. 21 This not only generates low bias, but certain architectures have also been shown to generalize well to out‐of‐sample data, generating low variance as well. 22 Finally, the SE is an approach which seeks to further reduce variance error by combining the results of all the ML algorithms. More complete details of the algorithm families we examine can be found in the Appendix S1.
Each family of ML algorithms has a series of hyperparameters that can potentially affect the accuracy of the results. Since the best set of hyperparameter values is unique to each prediction problem, there is no way to ensure any given set of values is the optimal set. To aid in model selection, we use the AutoML feature of the H2O.ai ML platform. The AutoML tool conducts a grid search over common hyperparameter values for each ML algorithm to identify models with strong predictive performance. To limit training time, we examine only the first 75 models in AutoML's search tree (which includes one RF, one XFR, 12 GBMs, and 61 DNNs).
Models are then compared based on prediction accuracy using a five‐fold cross‐validation approach. The best model from each family is then selected for final comparison on a hold‐out test data set. Cross‐validation reduces the potential of selecting overfitted models as accuracy is based and models are selected on “out‐of‐sample” validation data.
2.4. Comparing predictions and imputations with current matching approach
We first train all algorithms (OLS, RF, XRF, GB, DNN, SE) on the 80% cut of the events with complete MPC data, henceforth called the “donors”, on a single target, the square root of total expenditure, the transformation currently used by MEPS to convert the highly skewed expenditure distribution.
Next, we obtain predictions of total expenditure for both the donor events and the 20% testing set representing the recipients, match recipients to donors based on the minimum distance function within existing XGROUPS and assign the matched donor's actual expenditures as the recipient's imputed expenditures. Finally, we assess the accuracy of the imputations by comparing the imputed values to the recipient's actual MPC values.
Our preferred metric is the out‐of‐sample R2, with higher values indicating better performance. It should be noted that the R2 measure can take negative values when calculated on out‐of‐sample data not used to estimate the model, as the prediction error can exceed the total variation in the data. Although the R2 is a commonly used metric, relative model performance can vary with the choice of evaluation metric. We assess performance with a variety of other metrics in the Appendix S1, all of which show equivalent results to the R2. While a random 80/20 missingness pattern does not reflect reality, because algorithms all face the same data, their expected performance can be assessed.
2.5. Assessing imputations based on matching with vector of predictions
After assessing the prediction and imputation performance for the single target expenditure, we expand the ML training and evaluation to include the multiple targets from each potential source of payment. In contrast with OLS, the ML algorithms can jointly estimate multiple targets, however, we elect to predict each of the sources of payment separately and then use those predictions as inputs for the prediction of the total expenditure. To limit computational effort, we only examine one algorithm, the best‐performing model from the single target exercise.
The final prediction vector then consists of all ten predicted values (the total expenditure value, as well as each of the nine payment sources, examined). As with the previous step, we match the recipient group with donors based on the vector of predictions instead of a single prediction of total expenditure. Critically, we do not impose a predefined XGROUP structure on this matching, instead allowing the prediction vector to guide the matches. We match based on the minimum Euclidean distance between the recipient and donor prediction vectors (details in the Appendix S1) and assign all expenditure values from the closest donor to the recipient.
To assess the performance of this matching approach, separately for total expenditures and each source of payment, we compare the vector of imputed values with the vector of actual values, again using R2 as the comparison metric. We also compare the results obtained from the vector matching with the baseline set of results from matching with OLS from our first analysis which represents the closest approximation to the current imputation process.
3. RESULTS
3.1. Prediction with machine learning algorithms
Table 1, Panel A compares the out‐of‐sample prediction R2 statistics, or the fit of the recipient predictions to their true values, for the algorithms examined. The currently used algorithm, OLS, does the worst job at predicting the out‐of‐sample values for every event type, with R2 statistics ranging from 0.086 to 0.365. Appendix Table 3 reports alternative evaluation metrics for these measures which show similar results. The individual ML algorithms all offer similar improvements in performance, with the individual algorithm offering the largest performance gain varying by event type.
TABLE 1.
Prediction accuracy of machine learning algorithms
| Algorithm | Event type | |||
|---|---|---|---|---|
| Emergency department (ED) | Hospitalization (HS) | Outpatient (OP) | Medical visit (MV) | |
| Panel A: Prediction accuracy | ||||
| Ordinary least squares (OLS) | 0.27 | 0.365 | 0.264 | 0.086 |
| Gradient boosting (GB) | 0.318 | 0.413 | 0.428 | 0.333 |
| Random forest (RF) | 0.318 | 0.411 | 0.396 | 0.341 |
| Extreme random forest (XRF) | 0.32 | 0.415 | 0.408 | 0.333 |
| Deep neural net (DNN) | 0.283 | 0.424 | 0.375 | 0.219 |
| Stacked ensemble (SE) | 0.358 | 0.46 | 0.44 | 0.352 |
| Panel B: Imputation accuracy | ||||
| Ordinary least squares (OLS) | −0.438 | 0.071 | −0.127 | −0.298 |
| Gradient boosting (GB) | 0.06 | 0.345 | 0.038 | 0.019 |
| Random forest (RF) | 0.263 | 0.396 | 0.014 | 0.052 |
| Extreme random forest (XRF) | 0.172 | 0.393 | 0.015 | 0.047 |
| Deep neural net (DNN) | −0.062 | 0.37 | −0.173 | −0.195 |
| Stacked ensemble (SE) | 0.25 | 0.408 | 0.069 | 0.069 |
Note: Values represent the imputation R2 for the square root of total expenditure on out‐of‐sample respondents' predictions (Panel A) and imputations (Panel B) created by matching with in‐sample values based on proximity of the predicted expenditures of both the in‐ and out‐of‐sample observations. Observation matching is conducted within pre‐defined XGROUPS representing source of payment patterns to ensure similar expenditure profiles for matches.
However, the SE approach, which not only combines the predictions of the four best of family ML algorithms reported but also the predictions of the other 71 models estimated during the automatic grid search, gives the best prediction performance overall. Additionally, the SE model offers significant improvements over OLS. The amount of improvement in prediction accuracy over OLS varies by event type, ranging from a 0.088‐point improvement for ED events to a 0.266‐point improvement for MV events.
The variance in prediction accuracy and improvements across event types is also notable. The OLS algorithm does a reasonable job predicting HS event total expenditures, but there is a steep drop off in accuracy when predicting other event types, particularly MV events. This can probably be attributed to the relatively homogenous expense of hospital stays since hospital expenditures are closely related to the length of hospitalization in prediction models. Conversely, MV events have significant heterogeneity in expenditures because they range from simple well visits to more complex consultations or treatments, such as oncological services, which can be very expensive. 23 While it is difficult to fit these heterogeneities into a linear model using OLS, the ML algorithms are more flexible, allowing for better predictions on this type of data.
We continue to follow the PMM imputation process by matching these predictions with donor data based on the single value prediction of total expenditure. The imputation R2 statistics, comparing the matched donor values to actual values, can be found in Table 1, Panel B. Alternative evaluation metrics can be found in Appendix Table 4 which shows similar results. The matching process reduces the accuracy of the imputations across algorithms. For OLS, the matching process results in a negative R2 statistic for three of the four event types, suggesting that a simple mean imputation would give closer values than the matching process. As with the prediction step, the ML algorithms consistently outperform the OLS results and, except for the DNN algorithm, generate matches with positive R2 statistics. The SE approach continues to give the best overall performance, though the RF algorithm results in slightly better matches for ED events.
Considering the higher accuracy of direct prediction when compared to the imputation matches raises the question of dispensing with matching altogether in favor of imputing missing values directly with the first‐step predictions. Figure 2, Panel A overlays the distributions of the predicted square root of total expenditures of the SE algorithm on the distributions of the true values for the recipient data. The actual distribution of total expenditures is right skewed for all event types. However, the distribution of the predicted values tends to concentrate more values near the modes of the actual distribution, with fewer predictions in the long tails. Figure 2, Panel B overlays the imputed expenditure distribution on the actual. In contrast with the prediction's distribution, the imputed distribution is much more in line with the underlying true distribution, preserving observations in the long tails and reducing the concentration of values at the modes. Similar figures for the current OLS matching can be found in Appendix Figure 1.
FIGURE 2.

Distributions of predicted and imputed total expenditures for recipients using a stacked ensemble prediction algorithm. Medical events are divided by type: emergency department visits (ED), inpatient hospitalizations (HS), office‐based visits (MV), and outpatient visits in hospital settings (OP).
3.2. Imputation matching by vector of predictions
The replacement of OLS with ML in the PMM process results in significantly better imputations for a single target, the square root of total expenditures in this case. Next, we assess matching based on a vector of predictions, the square roots of all available expenditure variables, instead of based on a single target. This assessment will determine whether a more flexible vector‐matching approach outperforms the current “XGROUP” matching process. Table 2 reports the R2 statistics for out‐of‐sample predictions obtained from the Stacked Ensemble approach for all sources of payment. As expected, the prediction accuracy for total expenditure compares favorably to the results reported in Table 1 as the only difference between these two sets of models is the addition of the predicted source of payment sub‐components as additional predictors.
TABLE 2.
Prediction accuracy of a machine learning stacked ensemble for a vector of medical expenditures
| Expenditure type | Event type | |||
|---|---|---|---|---|
| Emergency department (ED) | Hospitalization (HS) | Outpatient (OP) | Medical visit (MV) | |
| Out of pocket | 0.714 | 0.503 | 0.643 | 0.717 |
| Medicare | 0.754 | 0.758 | 0.597 | 0.64 |
| Medicaid | 0.658 | 0.711 | 0.659 | 0.716 |
| Private | 0.649 | 0.698 | 0.656 | 0.63 |
| VA | 0.991 | 0.634 | 0.915 | 0.802 |
| Tricare/CHAMPUS | 0.675 | 0.739 | 0.486 | 0.848 |
| Other state and Local | −0.616 | 0.997 | 0.484 | 0.813 |
| Worker's comp | 0.709 | ‐ | 0.528 | 0.716 |
| Other | 0.630 | −0.044 | 0.655 | 0.739 |
| Total expenditure | 0.254 | 0.452 | 0.410 | 0.332 |
Note: Values represent the prediction R2 for the square root of expenditure on out‐of‐sample respondents. Expenditure types are the official MEPS sources of payments with the exception of Other Federal Sources which has been combined with Other payments due to the relative rarity of this source of payment.
Prediction accuracies for the individual sources of payment are significantly higher than those for overall payment. This is likely due to the inclusion of source of payment indicators as predictors when training the model. It becomes trivially easy for the ML algorithms to predict a value of $0 for respondents that report a particular source of payment is not available to them, which would inherently increase the accuracy of the predictions. Despite the triviality of these predictions, they play an important role in the vector‐matching process as they contain information on the pattern of sources of payment available to both the recipient and the donor, which we wish to align as closely as possible.
The SE approach has trouble in predicting Other State and Local expenditures for ED events and Other expenditures for HS events. For these sources of payments, the prediction R2 statistic is negative. Given the sparsity of these types of expenditure in the respective event types, it is likely the algorithms did not have enough examples from which to learn for robust prediction. Training with more data would provide additional examples and possibly lead to better prediction accuracy.
Table 3 compares the imputation accuracy for matching based on a single target prediction within XGROUPS for both the OLS and SE algorithms and for matching based on the vector of expenditure predictions from the SE algorithm. For reference, the imputation accuracy for Total Expenditure for the single target matching is the same imputation accuracy reported in Table 1 Panel B for the respective algorithms. Table 3 extends this metric to all sources of payment.
TABLE 3.
Imputation accuracy of matching on a single target versus a vector of targets
| Source of payment | Event type | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Emergency department (ED) | Hospitalization (HS) | Outpatient (OP) | Medical visit (MV) | |||||||||
| OLS ‐ Single | SE ‐ Single | SE ‐ Vector | OLS ‐ Single | SE ‐ Single | SE ‐ Vector | OLS ‐ Single | SE ‐ Single | SE ‐ Vector | OLS ‐ Single | SE ‐ Single | SE ‐ Vector | |
| Out of pocket | −0.913 | −1.39 | 0.554 | −17.9 | −17.8 | 0.224 | −0.427 | −0.312 | 0.522 | −0.149 | 0.014 | 0.66 |
| Medicare | −1.36 | −0.001 | 0.665 | 0.103 | 0.283 | 0.759 | −1.16 | −1.61 | 0.31 | −0.699 | −0.2 | 0.505 |
| Medicaid | −0.129 | −0.076 | 0.594 | −0.151 | −0.251 | 0.708 | −0.086 | −0.357 | 0.483 | −0.233 | −0.12 | 0.615 |
| Private | −1.71 | 0.151 | 0.556 | −0.134 | −0.14 | 0.655 | −0.568 | 0.07 | 0.49 | −0.615 | −0.322 | 0.437 |
| Veteran's Admin. | −0.082 | 0.325 | 0.81 | −0.006 | −0.006 | 0.145 | 0.114 | 0.127 | 0.166 | 0.316 | 0.74 | 0.792 |
| Tricare/CHAMPUS | −4.38 | −10.6 | 0.716 | 0.084 | −0.03 | 0.727 | 0.02 | 0.016 | 0.446 | −0.672 | −1.43 | 0.734 |
| Other state and Local | −28.6 | −9.46 | 0.432 | −0.014 | −0.003 | −0.003 | 0.019 | 0.019 | 0.133 | 0.771 | 0.515 | 0.931 |
| Worker's Comp | −0.066 | −0.007 | 0.881 | ‐ | ‐ | ‐ | −0.021 | −0.001 | 0.523 | 0.18 | 0.186 | 0.684 |
| Other | −0.04 | −0.293 | 0.71 | −0.01 | −0.723 | 0.205 | −0.096 | −0.105 | 0.626 | −0.797 | −0.088 | 0.723 |
| Total Expenditure | −0.438 | 0.25 | 0.208 | 0.071 | 0.408 | 0.438 | −0.127 | 0.069 | 0.162 | −0.298 | 0.069 | 0.072 |
Note: Values represent the imputation R2 for the square root of expenditure on out‐of‐sample respondents matched with in‐sample true values based on proximity of the predicted expenditures of both the in‐ and out‐of‐sample observations. Predictions are made using Ordinary Least Squares (OLS) and a Stacked Ensemble (SE) of machine learning algorithms. Observation matching for the single target columns is conducted within pre‐defined XGROUPS representing source of payment patterns to ensure similar expenditure profiles for matches. Matching for the vector columns does not impose a group structure on source of payment patterns. Expenditure types are the official MEPS sources of payments with the exception of Other Federal Sources which has been combined with Other payments due to the relative rarity of this source of payment.
Matching based on the single target of total expenditure results in significant discrepancies in the expenditure sub‐components, despite the attempts to align sources of payment using the hard classification of the XGROUPS. Across event types, matching based on OLS predictions routinely results in negative R2 statistics that are large in magnitude. Switching to matching based on SE predictions, while improving the accuracy of imputed total expenditures, only does slightly better when considering the imputation accuracy for individual sources of payment.
In contrast, matching based on the vector of predictions produces a much better alignment of sources of payments between the recipient and donor. The imputation R2 statistics for the individual sources of payment routinely exceeds that of the total expenditure imputation, with a majority exceeding 0.5. Expanding the number of matching targets not only improves source of payment alignment, but it does not adversely affect the imputation accuracy of overall expenditures. In fact, vector matching improves (or only slightly reduces, as in the case of ED events) the accuracy of imputed total expenditures.
4. DISCUSSION
Replacing OLS with a more sophisticated ML approach in the PMM process yields both better predictions and better imputations in the MEPS. While all ML algorithms showed improvement over OLS, the Stacked Ensemble performed best, improving prediction performance by 108% and final imputation performance by 227% on average. These results are comparable to prior research on the accuracy of medical expenditure prediction by ML. 11 , 16 Additionally, we demonstrate that matching on a vector of predictions instead of a single prediction creates closer alignment between donor and recipient payment patterns. These improvements suggest an expanded role for ML, not only in the MEPS imputation process but in broader survey data preparation applications.
It is not evident a priori that ML would improve PMM performance. OLS can perform as well as or better than ML algorithms in predicting medical expenditures, particularly in the tails of the distribution. 14 , 24 However, in predicting average MEPS expenditures, OLS does not seem to consistently generalize to the out‐of‐sample data (as measured by prediction R2). ML algorithms, on the other hand, more consistently predict similar values for in‐ and out‐of‐sample data, leading to better correspondence between donors and recipients in the matching process. However, despite the superior generalization by ML algorithms to out‐of‐sample data, there is still significant drop‐off in accuracy between direct predictions and matching imputations.
Direct prediction, while providing greater accuracy than PMM, shrinks towards the mean and thus fails to adequately replicate the underlying distribution. This is of concern in imputing survey data as the distribution of data matters more than any single individual value, as it is the distribution from which inferences about the population parameters are made. Although matching reduces accuracy, it restores the shape of the underlying distribution. Consequently, matching remains a preferred method of imputation as maintaining the underlying distribution is an important goal of MEPS expenditure imputation.
In the MEPS expenditure data, which can be divided into sub‐component expenditures (sources of payments), we demonstrated that ML algorithms offer improvements to the matching process as well. Using OLS limits matching to a single prediction (total expenditure), which can create significant mismatches in sources of payments even with a hard classification scheme, which attempts to match events with similar expenditure components. This rigidity in matching pools limits potential matches, thereby reducing match quality.
ML algorithms provide flexibility in the number of target predictions each model can produce. Jointly predicting both total expenditures and its components creates an opportunity to match recipients with donors across a vector of predictions instead of a single prediction. This approach eliminates the hard classifications scheme, but still matches recipients with donors that have similar expenditure patterns. Moreover, this vector‐matching approach provides flexibility for potentially better matches that may have otherwise been precluded by XGROUPS.
Our study found significant improvements to the PMM process resulting from these two changes: shifting from OLS to a ML approach and matching based on a vector of expenditures. Not only are imputations for total expenditures shown to be markedly better, but alignment of the patterns of sources of payment is substantially improved as well. The success of these approaches for the MEPS data suggests potential gains for other survey imputation processes currently using PMM.
However, there are limitations to ML algorithms. First, unlike OLS, which under certain assumptions is guaranteed to be the best linear estimator, there is no way to know if any particular ML model is the optimal one (or even close to it). In this respect, choosing the hyperparameters for ML is much more of an art than a science, and results can be sensitive to ad hoc choices made by the modeler. Next, these approaches require significant amounts of both data and computing power, and thus may be limited in application to smaller imputation problems with fewer resources. Finally, if the data to be imputed are not overly complex, these ML algorithms may be excessive and less accessible in terms of transparency for the end users of the imputed data.
This work examined the feasibility of adapting a PMM approach to imputing missing medical expenditure values in the MEPS to exploit recent advances in ML. A Stacked Ensemble ML approach with matching on a vector of predicted values performed substantially better than OLS with single target matching, the current method for imputation in MEPS. This better predictive performance also generated better matches and imputations in the PMM process. As the MEPS expenditure data are widely used in both policy and research, and imputation accuracy substantially impacts these data, improving PMM imputation with ML is actively being considered for adoption in the MEPS data production process, possibly as early as the 2021 or 2022 data years. Further, these results suggest that ML is a promising approach for improving imputation in a variety of other settings as well.
Supporting information
Appendix S1. Supporting Information.
ACKNOWLEDGEMENT
We express our gratitude to Steve Hill, Ed Miller, Joel Cohen, the participants of the AHRQ CFACT seminar series, and our anonymous reviewers for helpful feeback resulting in tremendous improvement to the manuscript. The authors have no external funding sources or other conflicts of interest to report.
McClellan C, Mitchell E, Anderson J, Zuvekas S. Using machine‐learning algorithms to improve imputation in the medical expenditure panel survey. Health Serv Res. 2023;58(2):423‐432. doi: 10.1111/1475-6773.14115
The views expressed herein are solely those of the authors. No official endorsement by the Agency for Healthcare Research and Quality is intended or should be inferred.
DATA AVAILABILITY STATEMENT
The data that support the findings of this study are available from the Agency for Healthcare Research and Quality. Restrictions apply to the availability of these data.
REFERENCES
- 1. Meyer BD, Mok WK, Sullivan JX. Household surveys in crisis. J Econ Perspect. 2015;29(4):199‐226. [Google Scholar]
- 2. NHIS . Multiple imputation of family income in 2019 national health interview survey: Methods. NCHS Methods Report. National Center for Health Statistics; 2020. [Google Scholar]
- 3. CPS . Imputation of Unreported Data Items. Census Bureau Technical Report. U.S. Census Bureau; 2021. [Google Scholar]
- 4. Little RJ, Rubin DB. Statistical Analysis with Missing Data. Vol 793. John Wiley & Sons; 2019. [Google Scholar]
- 5. Ibrahim JG, Chen M‐H, Lipsitz SR, Herring AH. Missing‐data methods for generalized linear models: a comparative review. J Am Stat Assoc. 2005;100(469):332‐346. [Google Scholar]
- 6. Andridge RR, Little RJ. A review of hot deck imputation for survey non‐response. Int Stat Rev. 2010;78(1):40‐64. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Little RJ. Missing‐data adjustments in large surveys. J Bus Econ Stat. 1988;6(3):287‐296. [Google Scholar]
- 8. Mullainathan S, Spiess J. Machine learning: an applied econometric approach. J Econ Perspect. 2017;31(2):87‐106. [Google Scholar]
- 9. Morid MA, Kawamoto K, Ault T, Dorius J, Abdelrahman S. Supervised learning methods for predicting healthcare costs: systematic literature review and empirical evaluation. AMIA Annual Symposium Proceedings: 2017. American Medical Informatics Association; 2017:1312. [PMC free article] [PubMed] [Google Scholar]
- 10. Duncan I, Loginov M, Ludkovski M. Testing alternative regression frameworks for predictive modeling of health care costs. North Am Actuarial J. 2016;20(1):65‐87. [Google Scholar]
- 11. Irvin JA, Kondrich AA, Ko M, et al. Incorporating machine learning and social determinants of health indicators into prospective risk adjustment for health plan payments. BMC Public Health. 2020;20(1):1‐10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Kaushik S, Choudhury A, Dasgupta N, Natarajan S, Pickett LA, Dutt V. Using LSTMs for predicting patient's expenditure on medications. 2017 International Conference on Machine Learning and Data Science (MLDS): 2017. IEEE; 2017:120‐127. [Google Scholar]
- 13. Mazumdar M, Lin J‐YJ, Zhang W, et al. Comparison of statistical and machine learning models for healthcare cost data: a simulation study motivated by oncology care model (OCM) data. BMC Health Serv Res. 2020;20(1):1‐12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Rose S. A machine learning framework for plan payment risk adjustment. Health Serv Res. 2016;51(6):2358‐2374. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Shrestha A, Bergquist S, Montz E, Rose S. Mental health risk adjustment with clinical categories and machine learning. Health Serv Res. 2018;53:3189‐3206. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Vimont A, Leleu H, Durand‐Zaleski I. Machine learning versus regression modelling in predicting individual healthcare costs from a representative sample of the nationwide claims database in France. Eur J Health Econ. 2022;23(2):211‐223. [DOI] [PubMed] [Google Scholar]
- 17. Kolassa S. Sometimes It's better to Be simple than correct. Foresight Int J Appl Forecasting. 2016;40:20‐26. [Google Scholar]
- 18. MEPS Data Overview. https://meps.ahrq.gov/mepsweb/data_stats/data_overview.jsp
- 19. Mitchell E, Anderson J, Biener A. Undercounting of healthcare utilization in the medical expenditure panel survey. AHRQ Workieng Paper 20001 . 2020.
- 20. Geurts P, Ernst D, Wehenkel L. Extremely randomized trees. Mach Learn. 2006;63(1):3‐42. [Google Scholar]
- 21. Hornik K, Stinchcombe M, White H. Multilayer feedforward networks are universal approximators. Neural Netw. 1989;2(5):359‐366. [Google Scholar]
- 22. Neal B, Mittal S, Baratin A, Tantia V, Scicluna M, Lacoste‐Julien S, Mitliagkas I. A modern take on the bias‐variance tradeoff in neural networks. arXiv Preprint arXiv:181008591. 2018.
- 23. Park J, Look KA. Health care expenditure burden of cancer Care in the United States. INQUIRY J Health Care Organ Provision Financing. 2019;56:0046958019880696. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Park S, Basu A. Alternative evaluation metrics for risk adjustment methods. Health Econ. 2018;27(6):984‐1010. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix S1. Supporting Information.
Data Availability Statement
The data that support the findings of this study are available from the Agency for Healthcare Research and Quality. Restrictions apply to the availability of these data.
