Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Apr 22.
Published in final edited form as: J Racial Ethn Health Disparities. 2025 Apr 22;13(4):2513–2522. doi: 10.1007/s40615-025-02435-4

Race and Ethnicity Data in the Electronic Health Records: New Insights Through Comparison with American Community Survey Microdata

Rocio Rosa-Lebron 1, Aubrey Limburg 2, Timothy S Carey 3, Victoria M Udalova 2, Barbara Entwisle 1,4,*
PMCID: PMC12771273  NIHMSID: NIHMS2114087  PMID: 40261485

Abstract

The collection of race and ethnicity information varies across data sources which impacts our ability to conduct high quality research focused on population health, generally, and racial and ethnic disparities in health, specifically. This research examines concordance in racial/ethnic identification between two sources by linking individual-level electronic health record (EHR) data (2017–2019) from a public integrated health delivery system in North Carolina to American Community Survey (ACS) microdata (2001–2017). We find that concordance is high for individuals who identify as non-Hispanic Black, non-Hispanic White, and Hispanic, but considerably lower for other non-White, non-Hispanic individuals, particularly for American Indian and Alaska Native patients. Given their detailed health information, EHR data have the potential to support research focused on population health and racial and ethnic disparities. Results from this study provide information regarding data quality and future applications of this work to expand population research.

Keywords: race, ethnicity, electronic health records, missing data, data quality

Introduction

There is no universally accepted method for measuring race and ethnicity given that categorization and identification can change over time and vary by place [1]. To standardize race and ethnicity for federal data collection and reporting, the Office of Management and Budget (OMB) has promulgated recommendations – first in 1977, with revisions in 1997 and 2024 – about approach (self-reports are preferred), format (which question(s) to ask and in what order), and response categories [2]. States, businesses, and hospitals have adopted but also adapted these recommendations, resulting in inconsistencies across organizations. Organizations also differ in the priority given to the collection of race and ethnicity data and, thus, its quality and completeness. Variation within and between sources abounds with consequences for the assessment of racial and ethnic health disparities, especially when disparate data sources are pooled, integrated, or otherwise compared [3–5]. This paper links race and ethnicity data derived from electronic health records (referred to as EHRs) to American Community Survey (ACS) microdata to assess differences in the classification of individuals who appear in both datasets. The overall goal of this work is to inform research focused on racial and ethnic health disparities by considering the data quality of EHRs and their implication for population health research.

Electronic Health Records: Possibilities and Challenges

Electronic health record (EHR) data, created through patient contact with healthcare delivery systems, are a valuable data source for studying health disparities [6]. Their use expanded dramatically after the passage of the 2009 American Recovery and Reinvestment Act, which required healthcare systems to adopt and demonstrate “meaningful use” of EHRs by 2014 [7]. Recent examples include studies of racial and ethnic disparities in lung cancer diagnosis [8], provider-patient communication [9], blood pressure control among patients with hypertension [10], and inflammatory bowel disease [11], to name a few. Each capitalizes on the strengths of EHRs, which include up-to-date and detailed information on patient visits, lab reports, diagnoses, medications, and the like for very large patient populations [12].

Despite their strengths, the use of EHRs to describe racial and ethnic health disparities faces several challenges. One significant challenge stems from missing data on race and ethnicity, which varies across health systems and can be substantial [4,13,14]. In a recent report based on the 56 health care institutions in the National COVID Collaborative Cohort (NC3), 11.3% of patients were missing data on race and an additional 8% were recorded as refusals [3]. Nearly 60% of patients were missing data on race in one studied EHR system [15]. Researchers have taken a number of approaches to managing these missing data, such as incorporating them into a residual category along with refusals [10,16,17], imputing missing values [4], or simply omitting them from analysis altogether [8,11]. A systematic review identified 15 studies that examined the completeness of race and ethnicity data [18]. Two-thirds of those reported that missing data were not missing at random, much less missing completely at random, which could introduce bias into statistical estimates. In addition to possible challenges related to missing information, there are also challenges related to consistent classification of race and ethnicity information across sources.

Racial and Ethnic Classification

Race is a contextual identity [19], and the race that people identify as (or are identified as, in the case of getting race data from clinical notes) may vary across settings both inside and outside of the healthcare context [20]. For example, Harris and Sim [20] examined how students identified their race when at home versus when at school. Though the same question with the same categories was asked in both instances, they found that while 6.8% of students identified as multiracial when at school, only 3.6% of those same students identified as multiracial when at home. Brown and colleagues [21] found similar variations in school versus home context among Hispanics: of the students in their sample who identified as Hispanic at school, about 80% also identified as Hispanic when interviewed at home. Healthcare settings also provide a distinct context that could affect the reporting of racial identity, as non-white patients report being uncomfortable giving this information to healthcare providers for fear of discrimination [22].

Additionally, within the healthcare context, how the data are collected also varies, as is the degree to which healthcare professional prioritize collecting race and ethnicity data [4,23,24]. Despite federal recommendations, non-federal agencies frequently adapt racial/ethnic classification to their unique context, resulting in inconsistencies across organizations. For example, the North Carolina-based integrated health system from which we draw the EHR data used in this study follows 1997 OMB guidance by collecting ethnicity separately from race and utilizing the minimum categories set forth in the 1997 revision. However, contrary to those guidelines, it neither requests nor accommodates multiple responses to the race question, which is not unusual [4]. These responses are coded as Other.

In EHRs, the Other category can be a bit of a catchall. Patients who respond in ways that do not correspond to prescribed categories (see Table 1) are classified as “other.” This includes patients who identify with a racial subgroup rather than the panethnic category (e.g., Chinese rather than Asian) and patients such as those from the Middle East or North Africa who, until 2024, were classified as White under the 1997 OMB standards [4,25] as well as those who identify with two racial groups. Other can also include Hispanic patients who do not identify with any of the standard racial groups. Since the 1997 revision of the OMB standards, there has been tension with the treatment of Hispanic/Latino as an ethnicity and not a race [26]. For example, in both the 2010 and 2020 Census, those who identified as Hispanic/Latino made up a majority of those who identified as Some Other Race only [27,28], suggesting that how Hispanic/Latino individuals classify themselves differs from the existing OMB categories.

Table 1.

Race and Ethnicity Classification in EHR and ACS Data

EHR Data ACS Data
Ethnicity
Hispanic or Latino Hispanic, Latino, or Spanish origin [Mexican, Mexican American, Chicano; Puerto Rican; Cuban; another (print)]
Not Hispanic or Latino Not of Hispanic, Latino, or Spanish origin
Patient Refused Missing (this is not present in ACS data due to allocation methods)
Unknown
Race
White or Caucasian White
Black or African American Black or African American
American Indian or Alaska Native American Indian or Alaska Native [print tribe]
Asian Asian [Asian Indian; Japanese; Chinese; Korean; Filipino; Vietnamese; Other Asian (print)]
Native Hawaiian or Other Pacific Islander Native Hawaiian or Pacific Islander [Native Hawaiian; Guamanian or Chamorro; Samoan; Other Pacific Islander (print)]
Other Some Other Race (print)
Two or more races
Missing / Patient Refused Missing (this is not present in ACS data due to allocation methods)

Note: The population of individuals identifying as Native Hawaiian or Other Pacific Islander in North Carolina is very small (<0.5%), they were grouped with the Other race group for proceeding analyses.

In the health system of interest during the 2016–2019 period that is our focus, patients or a proxy (for example, the parents or guardian of a child) provided the information directly to a clerk or as part of an intake questionnaire during the first visit. Sometimes, patients could also enter in the information directly during electronic registration. Clerks were instructed to accept whatever answer they were given and specifically not to push if there was resistance and not to fill the information in based on their own observations nor impute the information. Whether they adhered to this protocol consistently is unknown. Additionally, lack of training can create inconsistencies. It may not be clear to clerks or to patients why race and ethnicity information is needed [4]. Finally, new immigrants to the U.S. who are not familiar with the U.S. racial system or categories may have different understandings of race and ethnicity than what exists in the U.S., which may also impact how they answer questions on race/ethnicity [29].

Concordance Studies

In order to evaluate and address the challenges of EHR data for studying racial and ethnic disparities, previous research has linked EHRs to other health records such as billing systems and clinical observations [30,31], clinical trials [32], registries [24], and other health-focused research conducted outside hospital settings such as databases from healthcare related surveys [15,33,34]. Despite the innovative linkage strategies, many of these studies are not generalizable to larger patient populations given their specialized samples such as cancer survivors [35], low-income patients recruited into a tobacco cessation trial [32], or adolescents admitted to a psychiatric inpatient unit [36]. Sample sizes tend to be modest, typically less than 1,000 [e.g., 34], although there are exceptions [15]. In order to evaluate and improve the quality of EHR data, this study links EHR data to individual-level responses to a nationally representative sample of individuals in the American Community Survey. The linked data provides new information on the race and ethnicity of patients for whom that information is missing in the EHR as well as new insights about consistency in the classification of race and ethnicity information across sources.

Data and Methods

Data

This study compares race and ethnicity data from two sources: EHRs from 2016–2019 and American Community Survey (ACS) restricted microdata, 2001–2017. The EHR data1 come from a large public integrated healthcare delivery system located in North Carolina, whose mission is to serve the health and medical needs of the citizens of the state [37]. The state has been growing rapidly in recent years and is now the 9th largest in the country. The Hispanic population has been growing particularly rapidly, accounting for almost 11% of the overall population in 2020. The Asian population has also been growing rapidly, although from a much smaller base, accounting for 3% of the population in 2020 [38]. Immigration played a major role in these changes: about 40% of the Hispanic population were foreign-born in 2019 [39], and about 60% of the Asian population were foreign-born in 2016 [40]. Almost 40% of all immigrants in North Carolina were undocumented in 2016 (39%), most of whom (56%) were from Mexico [41].

This research builds on previous research in which we drew a disproportionate stratified random sample of about 200,000 patients aged 25–74 with at least two visits (e.g., in-patient visit or out-patient visit) between 2016 and 2019 [37]. As described earlier, race and ethnicity information were entered at registration into the EHR. Review of contemporaneous training materials used by the health system indicated that staff were to accept answers from patients, including allowing refusals. Staff were not to impute information or rely on appearance. For this research, we oversampled patients who identified as Black, Asian, or Other Race and Hispanic or Other Ethnicity. We also oversampled those for whom race or ethnicity was missing, allowing us to assess classification and completeness separately. For example, while 10% of patients were missing race information in the source data, 28% are missing race in the selected sample.

Limited patient demographic information (including race and ethnicity) for the sample were transferred securely to the Census Bureau IT environment, where anonymized person-level identifiers called Protected Identification Keys (PIKs) were assigned [42]. PIKs were assigned by matching personally identifiable information (PII), including Social Security number, name, date of birth, sex, and address, in the EHRs to references files maintained at the Census Bureau [42]. A more thorough explanation of how PIKs were assigned for these data can be found in a previous article written by the authors [37]. From there, patients with PIKs were linked to restricted ACS microdata with PIKs. For the purposes of this study, we focus exclusively on the subset of sampled EHRs for which PIKs and ACS matches were available. The success of PIK assignments and linkage to the ACS, discussed in detail elsewhere, varies by race and ethnicity [5,43]. We will return to this point in the discussion below.

The ACS is an annual cross-sectional survey of about 1–1.5% of the U.S. population designed as a replacement for the decennial census long-form in 2001. Participation in the ACS is mandatory and response rates have historically topped 90% [44], although coverage rates vary by race and ethnicity [45]. Critically, the ACS is based on a probability sample designed to be nationally representative. Our analysis is based on 29,000 patients linked from the EHRs to the ACS. In the ACS, a reference person reports both their own race and ethnicity and the race and ethnicity of all other household members. Importantly, only a small percent of data on race and ethnicity is imputed (allocated) in the ACS—3.3% for race and 2.4% for Hispanic origin in our sample of matched records. Because the ACS imputes missing data, it enables us to distinguish completeness from racial classification in the EHRs and to identify the race and ethnicity of patients with missing data.

Measures

Race and ethnicity information in EHR and ACS data are similar but have important differences. Notably, the ACS provides more information about racial and ethnic categories. For example, rather than “Hispanic or Latino” as in the EHRs, the ACS response category is “Hispanic, Latino, or Spanish origin (Mexican, Mexican American, Chicano, Puerto Rican, Cuban, another (print)).” As another example, rather than “Asian” as in the EHRs, the ACS response categories are “Asian Indian, Japanese, Chinese, Korean, Filipino, Vietnamese, and Other Asian (print).” Such detail may be particularly helpful, especially clarity regarding the inclusion of South Asians as well as East and Southeast Asians. Table 1 shows how we harmonized the categories for the two sources. Although the ACS allows for the selection of multiple race categories, persons identified as multiracial in the EHRs were included in the Other category. For the purposes of this concordance study, those who identified as Some Other Race or two or more races were combined into the Other race category when reporting on ACS data. Finally, as noted earlier, there is no category in the ACS comparable to “missing,” “refused,” or “unknown.”

Although race and ethnicity information were collected separately in each data source, they are commonly looked at together, and as noted earlier, OMB guidelines have moved in that direction as well. Following standard practice, our analysis focuses on a combined measure of race and ethnicity: non-Hispanic White, non-Hispanic Black, non-Hispanic American Indian and Alaskan Native (AIAN), non-Hispanic Asian, non-Hispanic Other race, and Hispanic2. However, we did examine the concordance of race and ethnicity separately as well; tables for these analyses are in the Supplementary Materials. When looking at the degree to which classification in the EHR agrees with the ACS, there is very little difference between race alone and the combined race-ethnicity measure. Thus, we use the combined race-ethnicity measures for this study.

Analyses

We assess concordance between race/ethnicity as recorded in the EHRs and the ACS from two perspectives. First, taking the EHRs as our starting point, we examine the extent to which patient race/ethnicity agrees with the ACS. Missing data are included in this assessment (N=29,000). We refer to this as the EHR→ACS perspective. Second, taking the ACS as our starting point, we examine the extent to which racial/ethnic classification agrees with the EHRs. Given the absence of a category for “missing” in the ACS, cases with missing data in the EHRs are excluded (N=20,500). We refer to this as the ACS→EHR perspective. Using both perspectives allows us to better capture the different ways people identify across the two data sources, as well as how the different categories used in the EHRs and ACS data can impact identification. For each perspective, we present heat maps showing levels of concordance and Sankey diagrams3 tracing the nature of the differences.

Results

The heat map in Figure 1a shows concordance from the EHR→ACS perspective. Concordance is highest for non-Hispanic White (97.1%) and non-Hispanic Black (95.9%) patients, closely followed by Hispanic (87.6%) and non-Hispanic Asian (87.2%) patients. Concordance for non-Hispanic AIAN patients is lower (68.0%) and strikingly low for non-Hispanic Other patients (7.4%). The latter likely reflects differences in what constitutes Other in the two datasets. In the EHRs, Other includes any response outside of the specified race categories, which may include such responses as “human race.” In the ACS, Other typically refers to the Some Other Race category, though in this instance we have also included the Native Hawaiian and Pacific Islander into the Other category. In addition, although the numbers are likely to be very small, the race of patients identifying as Middle Eastern/North African would be coded as Other in the EHRs but as White in the ACS.

Fig 1a and 1b. Heatmaps of Concordance between EHR and ACS data (left) and ACS to EHR (right).

Fig 1a and 1b.

Source: Electronic Health Records (2017–2019), 1-year American Community Survey (ACS) microdata (2001–2017).

Notes: AIAN = American Indian or Alaska Native and nH = non-Hispanic; The Census Bureau has reviewed this data product to ensure appropriate access, use, and disclosure avoidance protection of the confidential source data used to produce this product (Data Management System (DMS) number: 7519212, Disclosure Review Board (DRB) approval number: CBDRB-FY23-POP001–0160). All numeric values were rounded according to U.S. Census Bureau disclosure protocols to preserve data privacy.

A similar analysis is performed from the ACS→EHR, focusing on patients matched to ACS without missing race/ethnicity data in the EHRs4. As shown in Figure 1b, levels of concordance are highest for non-Hispanic Black (92.2%), non-Hispanic Whites (88.4%), and Hispanic (86.7%) patients. Although still high, concordance is slightly lower for non-Hispanic Asian (78.3%) and non-Hispanic AIAN (75%) patients. Of those discordant AIAN and Asian patients, 11% and 17.3% were identified as Other in the EHR, respectively. Patients identified as Other, non-Hispanic had the lowest concordance rate of all, at only 33%.

The Sankey diagrams in Figure 2 provide a visualization of the percentages shown in the heat maps. Of key interest in Figure 2a is how race/ethnicity is recorded in the ACS for patients who refused to answer either question in the EHRs or for whom EHR data were missing for some other reason. These patients were oversampled in the EHR data to provide sufficient cases for analysis. As shown, a disproportionate number of them identified as non-Hispanic White in the ACS, 73.7% compared to 56% for the overall EHR sample. The contrast is even more dramatic if the representation of non-Hispanic Whites among patients with missing data (73.7%) is compared with the representation of non-Hispanic Whites among patients with observed data (44.5%). All other racial/ethnic groups are underrepresented among patients with missing data.

Fig 2a and 2b. Sankey Diagrams of Concordance between EHR and ACS data (left) and ACS and EHR data (right).

Fig 2a and 2b.

Source: Electronic Health Records (2017–2019), 1-year American Community Survey (ACS) microdata (2001–2017).

Notes: AIAN = American Indian or Alaska Native and nH = non-Hispanic; The Census Bureau has reviewed this data product to ensure appropriate access, use, and disclosure avoidance protection of the confidential source data used to produce this product (Data Management System (DMS) number: 7519212, Disclosure Review Board (DRB) approval number: CBDRB-FY23-POP001–0160). All numeric values were rounded according to U.S. Census Bureau disclosure protocols to preserve data privacy.

Additionally, patients with discordant race/ethnicity (other than non-Hispanic White patients) are more often identified as non-Hispanic White in the ACS than any other racial/ethnic group, sometimes substantially. For example, in Figure 2a, 8.7% of patients who identified as Hispanic in the EHR are identified as non-Hispanic White in the ACS, as are 12.0% of non-Hispanic AIAN, 5.6% of non-Hispanic Asian, and 1.4% of non-Hispanic Black. A substantial portion, 42.2%, of those identified as non-Hispanic Other in the EHRs are identified as non-Hispanic White in the ACS.

Discussion

Patterns and Comparisons

This paper linked race and ethnicity data derived from EHRs from a large, public, integrated health delivery system in North Carolina to ACS microdata to assess levels of concordance between the two sources. Overall, non-Hispanic White and non-Hispanic Black individuals had the highest rates of race/ethnicity concordance across the two datasets. High levels of concordance for non-Hispanic White individuals have been found in previous work when comparing EHRs with other types of data, e.g., clinical trials, registries, and adjacent data collections. In a systematic review that identified 16 studies reporting such comparisons, data accuracy measured ranged from 81% to 99% for this group, with most studies exceeding 90% [13]. Our results fell within that range. Results for non-Hispanic Black patients are more variable. Johnson and colleagues [13] report a range of 70% to 99% for this group, with most above 93%. We found a similar pattern in our study, but others have found lower rates of concordance for non-Hispanic Black patients relative to non-Hispanic White patients [35].

Concordance was also high for Hispanic patients, though lower than that for non-Hispanic White and non-Hispanic Black patients. The status of Hispanic/Latino as a race or an ethnicity has long been in contention in the United States, with many Hispanic and non-Hispanic individuals treating the group as a race [48]. However, OMB guidance has previously treated the group as an ethnicity, which is reflected in the two-question format used in both the ACS and EHR data. Our results suggest that treating Hispanic/Latino as its own category can help increase concordance for that group because it better aligns with how Hispanics view themselves. In their systematic review, Johnson and colleagues [13] report concordance for Hispanics ranging between 41% and 91%, with most above 80%, which our results also echoed.

Concordance was also lower for non-Hispanic Asian patients. In their review, Johnson colleagues [13] report concordance ranging from 35% to 97% for Asian individuals, with most above 75%. For non-Hispanic Asian individuals, these differences could reflect differences in response categories between the two sources. In the ACS, the Asian category is separated out into specific country groups, which makes it clear who should identify as Asian. Meanwhile, the EHRs only give a singular, pan-Asian category without any other guidance [49]. As the dominant image of Asians in the U.S. is based on East Asia [50], those who find themselves excluded from the Asian category, particularly South Asians [51], may be likely to choose “Other.” Indeed, while 87.2% of patients identified as Asian in the EHR were also identified as Asian in the ACS, only 78.3% of those identified as Asian in the ACS were identified as Asian in the EHRs, with most (80%) of the remainder classified as “other.” And yet, other studies have found low concordance. The fact that our study focused on North Carolina may impact the concordance rates for Hispanic and Asian patients as well. Around 60% of Asian individuals in North Carolina are immigrants [40] as are about 40% of Hispanic individuals [52]. Immigrants may be unfamiliar with U.S. racial categories and understandings of race, which may explain the lower concordance rates.

Concordance was notably low for AIAN patients, mirroring results found in other studies [53]. These low rates fit with existing literature that discusses the difficulties with classifying native identity [54,55]. Additionally, the context of North Carolina may complicate AIAN identification. Eight tribes are recognized by the state, but only one is federally recognized, the Eastern Band of Cherokee. The largest tribe in the state, the Lumbee, is partially recognized [56]. Lack of federal recognition may play a role in the likelihood that individuals from this tribe report AIAN racial identity, though more research is needed to know for sure.

There was low concordance for individuals classified as Other. A large percentage (42%) of patients identifying as Other in the EHRs were classified as non-Hispanic White in the ACS. To some extent, this reflects the greater representation of non-Hispanic White individuals overall. Other factors may also be at work. At the time this study was conducted, individuals with origins in the Middle East and North America (MENA) were classified as White in the ACS (back coded as White if they selected Some Other Race and filled in a MENA identity [57]). However, MENA people have a complicated relationship to whiteness [25]. Because there is no MENA category in the EHR data and no instructions for how individuals who identify as MENA should respond, MENA individuals could be classified as other in the EHR, but White in the ACS. The next largest group (24%) of patients identifying as Other in the EHRs are classified as Asian in the ACS. We commented before on how differently the EHRs and ACS capture this ethnoracial identity. Finally, a very small portion (7%) of individuals identified as Other in the EHR were also identified as Other in the ACS, compared to if you look in the other direction (ACS to EHR is 33%). This suggests that Other as a residual category appears to have quite different meanings in the two sources.

Stepping back, our findings also show the importance of modality, question wording, and response categories for racial/ethnic classification. During the period of study, there was no single way to ask race and ethnicity in the EHRs, even though a modified version of the OMB guidelines were followed. Response categories were broad, not well defined, and listed alphabetically (e.g., for race: AIAN first, “White Caucasian” last). The ACS utilized standard questions, included clear instructions for answering them, provided more detail in the listing of response categories, and listed them in order of relative size. Groups with the highest concordance in our study, Hispanic and non-Hispanic White and Black individuals, are also those for whom questions and categories are most similar across the two data sources. Lower concordance was found for non-Hispanic Asian and AIAN individuals, where categories were less similar.

Missing Data

There was another clear difference between EHRs and the ACS: completeness of race and ethnicity reporting. EHRs are notably incomplete in reporting race; sometimes, more than half of the data are missing [15]. In the integrated health system that is the focus of our study, 10% were missing5. In contrast, race and ethnicity reporting in the ACS is complete, with only a small percent (<5%) imputed. This enabled us to identify the race and ethnicity for patients missing this information in the EHRs. Our findings show that in the EHR, the majority (73.7%) of the missing category is identified as non-Hispanic White in the ACS, a greater percentage than expected based on their representation in the overall sample. Indeed, research has shown that White individuals do not necessarily realize that they have a race and consider race to only be something non-White groups possess [58]. It is worth noting that this pattern of results does not align with that reported in other studies. Previous studies report that Hispanic, Asian, and AIAN patients are more likely than White patients to be missing [18]. It is possible that the state sociopolitical context contributes to this pattern. Whatever the explanation, the cases do not appear to be missing at random. Either way, restricting analysis to complete cases is likely to yield biased results [3,23,24].

Limitations

Our study, which compares race and ethnicity data between EHRs and ACS microdata, has many strengths, including being able to simultaneously evaluate missing data and identify differences in racial and ethnic classification between datasets. Of course, there are also limitations. One is that we draw EHRs from a single public integrated healthcare system in a single state, and as a consequence, the generalizability of the results is unclear. As discussed earlier, health systems vary in how they collect race and ethnicity data, the response categories included, and levels of missing data. The health system that is at the heart of our study prioritizes the collection of race and ethnicity data, and implements a standard protocol, but uses broad categories and does not allow for multiple responses. The state of North Carolina itself is distinctive in many ways, including the rapid growth of Hispanic and Asian populations. The immigrants who account for much of this growth may still be assimilating to racial categories in the U.S. [29]. Finally, previous research has shown considerable variation among states in the concordance of race and ethnicity as recorded in Medicaid records and Census Bureau sources [5,59]. While Medicaid records differ from EHRs, these studies raise questions about the generalizability of our results that should be addressed in future research.

A second concern about the generalizability of our results stems from responses coming from different years. EHRs from 2016–2019 are matched to 2001–2017 ACS. Racial and ethnic identities are fluid, specific to place, time, and context which may lead people to change the racial and ethnic group with which they identify [33,60]. Cook and colleagues [18] suggest that these changes in definition may be the reason that concordance for Hispanic/Latino individuals is lower than that of non-Hispanic Whites. If so, our results may underestimate concordance between race and ethnicity recorded in the EHRs and in the ACS. This is a topic worthy of further study.

Finally, PIK assignment and ACS matches are selective: of our original sample of around 200,000 EHRs, only about 14% meet both criteria [37]. It is possible that patients with the most differences EHR data have already been removed from the analytic sample. For example, PIK assignment relied partially (along with four other components) on Social Security numbers (SSNs), potentially excluding populations less likely to have SSNs such as undocumented immigrants [5]. Additionally, ACS matches may be lower among groups with lower coverage rates. Coverages rates6 may be affected by a variety of factors such as renters, people living in multi-family units, and those who distrust the government [43]. These factors may disproportionately impact racial and ethnic minorities, such as Black individuals, and lead to lower coverage rate for that group [45].

Conclusion

Given detailed health information, EHR data have the potential to support research focused on population health, generally, and racial and ethnic disparities, specifically. However, given the limitations of the race and ethnicity information currently available in EHR data, it is important to evaluate the quality of these data. In this research, we relied on an integrated dataset containing individual-level EHR data linked with American Community Survey (ACS) microdata – which is the nation’s largest nationally representative survey – as a means of investigating race and ethnicity identification in EHR data. Comparing these two sources made it possible to evaluate the quality of EHR data and investigate race and ethnicity information for patients who are missing this information. Our results suggest that ACS data may be used to enhance missing data from the EHRs, particularly for groups with high concordance such as non-Hispanic Black, non-Hispanic White, and Hispanic patients. Additionally, our results show that EHRs can be studied in relation to a representative source population. Factors related to who appears in EHRs can be understood, modeled, and incorporated in studies of population health and health disparities.

Supplementary Material

Supplementary Material

Acknowledgements:

This research is a collaborative effort involving the University of North Carolina at Chapel Hill (UNC) and the Enhancing Health Data (EHealth) Program at the Census Bureau. We are grateful for support from all parties: the Carolina Population Center (NICHD P2C HD050924); UNC Translational and Clinical Sciences (TraCS) Institute (CTSA UL1TR002489); Enhancing Health Data (EHealth) program (census.gov/ehealth). The goal of this partnership is to perform innovative data linkages between EHRs and restricted Census Bureau microdata in order to produce high-quality statistics and advance health research. Safeguarding the disclosure of data derived from EHRs to the Census Bureau fulfills the responsibility of the EHR owner to protect this information under the Health Insurance Portability and Accountability Act (HIPAA). Under Title 13 of the U.S. Code, the Census Bureau is authorized to collect information from various entities and is required to maintain its confidentiality, using it solely for statistical purposes. Aggregate statistics derived from EHRs and Census Bureau microdata are generated using rigorous disclosure avoidance techniques, ensuring full compliance with the confidentiality requirements outlined in Title 13 U.S.C., Section 9. Furthermore, all data products undergo review and approval by the Census Bureau’s Disclosure Review Board before they are publicly released (list relevant authorization numbers here). All numeric values were rounded according to U.S. Census Bureau disclosure protocols to preserve data privacy. Any opinions and conclusions expressed herein are those of the authors and do not reflect the view of the U.S. Census Bureau.

Funding:

We are grateful for support from the Carolina Population Center (NICHD P2C HD050924), the UNC Translational and Clinical Sciences (TraCS) Institute (CTSA UL1TR002489), and the Enhancing Health Data (EHealth) program at the US Census Bureau (census.gov/ehealth).

Footnotes

Statements and Declarations

Competing Interests: The authors have no relevant financial or non-financial interests to disclose.

Ethical Approval: This study is a partnership between the University of North Carolina at Chapel Hill and the US Census Bureau. The goal of the partnership is to perform innovative data linkages between electronic health records (EHRs) and restricted Census Bureau microdata in order to produce high-quality statistics and advance health research. Both parties provided oversight. The study was reviewed and approved by the UNC Institutional Review Board (#19–0549). Safeguarding the disclosure of data derived from EHRs to the Census Bureau fulfills the responsibility of the EHR owner to protect this information under the Health Insurance Portability and Accountability Act (HIPAA). Under Title 13 of the U.S. Code, the Census Bureau is authorized to collect information from various entities and is required to maintain its confidentiality, using it solely for statistical purposes. Aggregate statistics derived from EHRs and Census Bureau microdata are generated using rigorous disclosure avoidance techniques, ensuring full compliance with the confidentiality requirements outlined in Title 13 U.S.C., Section 9. Furthermore, all data products undergo review and approval by the Census Bureau’s Disclosure Review Board before they are publicly released (Data Management System (DMS) number: 7519212: Disclosure Review Board (DRB) approval number: CBDRB-FY23-POP001–0160). All numeric values were rounded according to U.S. Census Bureau disclosure protocols to preserve data privacy.

Data Statement: Under Title 13 of the U.S. Code, the Census Bureau is authorized to collect information from various entities and is required to maintain its confidentiality, thus data for this study cannot be shared.

1

Data are derived from EHR records, but do not contain the records themselves.

2

Given the very small numbers of Native Hawaiian and Other Pacific Islander individuals in our North Carolina patient population, we combined them with Other.

3

For more information on Sankey diagrams and how they have been used for data visualization, see Schmidt [46,47].

4

There is no category for missing data in the ACS, so it is not possible to assess concordance.

5

We oversampled so that we could better study this group.

6

The Census Bureau defines the coverage rate to be the ratio of the ACS population or housing estimate of an area or group to the independent estimate for that area or group, times 100 [61].

References

  • 1.Davenport L The Fluidity of Racial Classifications. Annu Rev Polit Sci. 2020;23:221–40. [Google Scholar]
  • 2.Office of Management and Budget. Revisions to OMB’s Statistical Policy Directive No. 15: Standards for Maintaining, Collecting, and Presenting Federal Data on Race and Ethnicity. 2024. [cited 2024 Oct 1];Federal Register Vol 89, No. 62. Available from: https://www.federalregister.gov/documents/2024/03/29/2024-06469/revisions-to-ombs-statistical-policy-directive-no-15-standards-for-maintaining-collecting-and
  • 3.Cook L, Espinoza J, Weiskopf NG, Mathews N, Dorr DA, Gonzales KL, et al. Issues With Variability in Electronic Health Record Data About Race and Ethnicity: Descriptive Analysis of the National COVID Cohort Collaborative Data Enclave. JMIR Med Inform. 2022;10:e39235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Owosela BO, Steinberg RS, Leslie SL, Celi LA, Purkayastha S, Shiradkar R, et al. Identifying and improving the “ground truth” of race in disparities research through improved EMR data reporting. A systematic review. International Journal of Medical Informatics. 2024;182:105303. [DOI] [PubMed] [Google Scholar]
  • 5.Limburg A, Young J, Carey TS, Chelminski PR, Udalova VM, Entwisle B. Assessing Electronic Health Records for Describing Racial and Ethnic Health Disparities: A Research Note. Demography. 2024;61:1325–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Rumball-Smith J, Bates DW. The Electronic Health Record and Health IT to Decrease Racial/Ethnic Disparities in Care. Journal of Health Care for the Poor and Underserved. 2018;29:58–62. [DOI] [PubMed] [Google Scholar]
  • 7.Blumenthal D, Tavenner M. The “Meaningful Use” Regulation for Electronic Health Records. N Engl J Med. 2010;363:501–4. [DOI] [PubMed] [Google Scholar]
  • 8.Kim SJ, Kery C, An J, Rineer J, Bobashev G, Matthews AK. Racial/Ethnic disparities in exposure to neighborhood violence and lung cancer risk in Chicago. Social Science & Medicine. 2024;340:116448. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Schut RA. Racial disparities in provider-patient communication of incidental medical findings. Social Science & Medicine. 2021;277:113901. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Cooper‐DeHoff RM, Fontil V, Carton T, Chamberlain AM, Todd J, O’Brien EC, et al. Tracking Blood Pressure Control Performance and Process Metrics in 25 US Health Systems: The PCORnet Blood Pressure Control Laboratory. Journal of the American Heart Association. 2021;10:e022224. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Barnes EL, Nowell WB, Venkatachalam S, Dobes A, Kappelman MD. Racial and Ethnic Distribution of Inflammatory Bowel Disease in the United States. Inflammatory Bowel Diseases. 2022;28:983–7. [DOI] [PubMed] [Google Scholar]
  • 12.Casey JA, Schwartz BS, Stewart WF, Adler NE. Using Electronic Health Records for Population Health Research: A Review of Methods and Applications. Annu Rev Public Health. 2016;37:61–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Johnson JA, Moore B, Hwang EK, Hickner A, Yeo H. The accuracy of race & ethnicity data in US based healthcare databases: A systematic review. The American Journal of Surgery. 2023;226:463–70. [DOI] [PubMed] [Google Scholar]
  • 14.Smith MA, Gigot M, Harburn A, Bednarz L, Curtis K, Mathew J, et al. Insights into measuring health disparities using electronic health records from a statewide network of health systems: A case study. J Clin Trans Sci. 2023;7:1–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Polubriaginof FCG, Ryan P, Salmasian H, Shapiro AW, Perotte A, Safford MM, et al. Challenges with quality of race and ethnicity data in observational databases. Journal of the American Medical Informatics Association. 2019;26:730–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Boehmer TK, Koumans EH, Skillen EL, Kappelman MD, Carton TW, Patel A, et al. Racial and Ethnic Disparities in Outpatient Treatment of COVID-19 ― United States, January–July 2022. MMWR Morb Mortal Wkly Rep. 2022;71:1359–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wiltz JL, Feehan AK, Molinari NM, Ladva CN, Truman BI, Hall J, et al. Racial and Ethnic Disparities in Receipt of Medications for Treatment of COVID-19 — United States, March 2020–August 2021. MMWR Morb Mortal Wkly Rep. 2022;71:96–102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Cook LA, Sachs J, Weiskopf NG. The quality of social determinants data in the electronic health record: a systematic review. Journal of the American Medical Informatics Association. 2021;29:187–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Cornell S, Hartmann D. Mapping the Terrain: Definitions. Ethnicity and Race: Making Identities in a Changing World. Thousand Oaks, California: Pine Forge Hills; 2006. p. 15–38. [Google Scholar]
  • 20.Harris DR, Sim JJ. Who Is Multiracial? Assessing the Complexity of Lived Race. American Sociological Review. 2002;67:614–27. [Google Scholar]
  • 21.Brown JS, Hitlin S, Elder Jr GH. The Greater Complexity of Lived Race: An Extension of Harris and Sim. Social Science Quarterly. 2006;87:411–31. [Google Scholar]
  • 22.Baker DW, Cameron KA, Feinglass J, Georgas P, Foster S, Pierce D, et al. Patients’ attitudes toward health care providers collecting information about their race and ethnicity. J Gen Intern Med. 2005;20:895–900. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Al-Sahab B, Leviton A, Loddenkemper T, Paneth N, Zhang B. Biases in Electronic Health Records Data for Generating Real-World Evidence: An Overview. J Healthc Inform Res. 2024;8:121–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Lee SJC, Grobe JE, Tiro JA. Assessing race and ethnicity data quality across cancer registries and EMRs in two hospitals. Journal of the American Medical Informatics Association. 2016;23:627–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Maghbouleh N, Schachter A, Flores RD. Middle Eastern and North African Americans may not be perceived, nor perceive themselves, to be White. Proc Natl Acad Sci USA. 2022;119:e2117940119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Porter SR, Snipp CM. Measuring Hispanic Origin: Reflections on Hispanic Race Reporting. The ANNALS of the American Academy of Political and Social Science. 2018;677:140–52. [Google Scholar]
  • 27.Humes K, Jones NA, Ramirez R. Overview of Race and Hispanic Origin: 2010. Washington, DC: US Census Bureau; 2011. Mar p. 24. Report No.: C2010BR-02. [Google Scholar]
  • 28.Jones NA, Marks R, Ramirez R, Rios-Vargas M. 2020 Census Illuminates Racial and Ethnic Composition of the Country [Internet]. Census.gov. 2021. [cited 2022 Dec 9]. Available from: https://www.census.gov/library/stories/2021/08/improved-race-ethnicity-measures-reveal-united-states-population-much-more-multiracial.html [Google Scholar]
  • 29.Roth W Race Migrations: Latinos and the Cultural Transformation of Race. Palo Alto: Stanford University Press; 2012. [Google Scholar]
  • 30.Cusick MM, Sholle ET, Davila MA, Kabariti J, Cole CL, Campion TR. A Method to Improve Availability and Quality of Patient Race Data in an Electronic Health Record System. Appl Clin Inform. 2020;11:785–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Sholle ET, Pinheiro LC, Adekkanattu P, Davila MA, Johnson SB, Pathak J, et al. Underserved populations with missing race ethnicity data differ significantly from those with structured race/ethnicity documentation. Journal of the American Medical Informatics Association. 2019;26:722–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Klinger EV, Carlini SV, Gonzalez I, Hubert SSt, Linder JA, Rigotti NA, et al. Accuracy of Race, Ethnicity, and Language Preference in an Electronic Health Record. J GEN INTERN MED. 2015;30:719–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Magaña López M, Bevans M, Wehrlen L, Yang L, Wallen GR. Discrepancies in Race and Ethnicity Documentation: a Potential Barrier in Identifying Racial and Ethnic Disparities. J Racial and Ethnic Health Disparities. 2017;4:812–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Samalik JM, Goldberg CS, Modi ZJ, Fredericks EM, Gadepalli SK, Eder SJ, et al. Discrepancies in Race and Ethnicity in the Electronic Health Record Compared to Self-report. J Racial and Ethnic Health Disparities [Internet]. 2022. [cited 2023 Jan 18]; Available from: https://link.springer.com/10.1007/s40615-022-01445-w [DOI] [PubMed] [Google Scholar]
  • 35.Codden RR, Sweeney C, Ofori-Atta BS, Herget KA, Wigren K, Edwards S, et al. Accuracy of patient race and ethnicity data in a central cancer registry. Cancer Causes Control. 2024;35:685–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Sojka PC, Maron MM, Dunsiger SI, Belgrave C, Hunt JI, Brannan EH, et al. Evaluation of Reliability Between Race and Ethnicity Data Obtained from Self-report Versus Electronic Health Record. J Racial and Ethnic Health Disparities [Internet]. 2024. [cited 2024 Jul 25]; Available from: 10.1007/s40615-024-02041-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Udalova V, Carey TS, Chelminski PR, Dalzell L, Knoepp P, Motro J, et al. Linking Electronic Health Records to the American Community Survey: Feasibility and Process. Am J Public Health. 2022;112:923–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.America Counts Staff. North Carolina: 2020 [Internet]. America Counts Story, U.S. Census Bureau. 2021. [cited 2024 Sep 4]. Available from: https://www.census.gov/library/stories/state-by-state/north-carolina-population-change-between-census-decade.html [Google Scholar]
  • 39.Cline M Hispanic Population is Fastest Growing Population in North Carolina | NC OSBM [Internet]. North Carolina Office of State Budget and Management. 2023. [cited 2023 Nov 6]. Available from: https://www.osbm.nc.gov/blog/2023/05/01/hispanic-population-fastest-growing-population-north-carolina [Google Scholar]
  • 40.Dreier A Asian Americans in North Carolina [Internet]. North Carolina Justice Center. 2016. [cited 2024 Jun 16]. Available from: https://www.ncjustice.org/publications/asian-americans-in-north-carolina/ [Google Scholar]
  • 41.Pew Research Center. U.S. unauthorized immigrant population estimates by state, 2016 [Internet]. Pew Research Center. 2019. [cited 2024 Aug 22]. Available from: https://www.pewresearch.org/race-and-ethnicity/feature/u-s-unauthorized-immigrants-by-state/ [Google Scholar]
  • 42.Wagner D, Layne M. The Person Identification Validation System (PVS): Applying the Center for Administrative Records Research and Applications’ (CARRA) Record Linkage Software. U.S. Census Bureau; 2014. Jul. Report No.: Working Paper #2014–01. [Google Scholar]
  • 43.Bond B, Brown JD, Luque A, O’Hara A. The Nature of the Bias When Studying Only Linkable Person Records: Evidence from the American Community Survey. U.S. Census Bureau; 2014. Apr. Report No.: Working Paper #2014–08. [Google Scholar]
  • 44.U.S. Census Bureau. American Community Survey (ACS) response rates [Internet]. 2022. [cited 2024 Sep 4]. Available from: https://www.census.gov/acs/www/methodology/sample-size-and-data-quality/response-rates/
  • 45.U.S. Census Bureau. American Community Survey (ACS) coverage rates [Internet]. 2022. [cited 2024 Sep 4]. Available from: https://www.census.gov/acs/www/methodology/sample-size-and-data-quality/coverage-rates/
  • 46.Schmidt M The Sankey Diagram in Energy and Material Flow Management: Part I History. Journal of Industrial Ecology. 2008;12:82–94. [Google Scholar]
  • 47.Schmidt M The Sankey Diagram in Energy and Material Flow Management: Part II: Methodology and Current Applications. Journal of Industrial Ecology. 2008;12:173–85. [Google Scholar]
  • 48.Hitlin S, Brown JS, Elder GH. Measuring Latinos: Racial vs. Ethnic Classification and Self-Understandings. Social Forces. 2007;86:587–611. [Google Scholar]
  • 49.Adia AC, Nguyen KH, Ponce NA. EHR Data and Inclusion of Multiracial Asian American, Native Hawaiian, and Pacific Islander People—Opportunities for Advancing Data-Centered Equity in Health Research. JAMA Netw Open. 2024;7:e240719. [DOI] [PubMed] [Google Scholar]
  • 50.Lee J, Ramakrishnan K. Who counts as Asian. Ethnic and Racial Studies. 2020;43:1733–56. [Google Scholar]
  • 51.Kurien P, Purkayastha B. Why Don’t South Asians in the U.S. Count As “Asian”?: Global and Local Factors Shaping Anti-South Asian Racism in the United States. Sociological Inquiry. 2024;94:351–68. [Google Scholar]
  • 52.Tippett R The Hispanic/Latino Community in North Carolina (2017) [Internet]. The University of North Carolina at Chapel Hill, Carolina Demography. 2017. [cited 2024 Sep 4]. Available from: https://carolinademography.cpc.unc.edu/2017/10/10/the-hispaniclatino-community-in-north-carolina/ [Google Scholar]
  • 53.Conrick KM, Mills B, Schreuder AB, Wardak W, Vil CSt, Dotolo D, et al. Disparities in Misclassification of Race and Ethnicity in Electronic Medical Records Among Patients with Traumatic Injury. J Racial and Ethnic Health Disparities. 2023;1–5. [DOI] [PubMed] [Google Scholar]
  • 54.Liebler CA. Counting America’s First Peoples. The ANNALS of the American Academy of Political and Social Science. 2018;677:180–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Liebler CA, Bhaskar R, Porter (née Rastogi) SR. Joining, Leaving, and Staying in the American Indian/Alaska Native Race Category Between 2000 and 2010. Demography. 2016;53:507–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.American Indian Center. FAQs About American Indians [Internet]. UNC American Indian Center. 2024. [cited 2024 Jul 31]. Available from: https://americanindiancenter.unc.edu/resources/faqs-about-american-indians/ [Google Scholar]
  • 57.Ennis S, Tiv M, Fernandez LE, Bhaskar R, Porter S. Examining Racial Identity Responses Among People with Middle Eastern and North African Ancestry in the American Community Survey. U.S. Census Bureau; 2024. Mar. Report No.: CES 24–14. [Google Scholar]
  • 58.Lewis AE. “What Group?” Studying Whites and Whiteness in the Era of “Color-Blindness.” Sociological Theory. 2004;22:623–46. [Google Scholar]
  • 59.Fernandez LE, Rastogi S, Ennis SR, Noon JM. Evaluating Race and Hispanic Origin Responses of Medicaid Participants Using Census Data. U.S. Census Bureau; 2015. Apr. Report No.: Working Paper #2015–01. [Google Scholar]
  • 60.Saperstein A, Penner AM. Racial Fluidity and Inequality in the United States. American Journal of Sociology. 2012;118:676–727. [Google Scholar]
  • 61.U.S. Census Bureau. Coverage Rates Definitions [Internet]. Census.gov. 2022. Available from: https://www.census.gov/programs-surveys/acs/methodology/sample-size-and-data-quality/coverage-rates-definitions.html [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material

RESOURCES