Since the first US Census in 1790, members of the public have expressed concerns about census intrusion into personal privacy. Public anxiety about census privacy appears to have peaked in the middle decades of the twentieth century and has generally declined since, except for a small uptick around the 2000 Census (Ruggles and Magnuson 2023). For almost two centuries, Census officials have responded to the public’s privacy concerns with promises of confidentiality. This strategy has been ineffective; promises to keep personal information within the government do not address the core concern about government prying.
The Census Bureau maintains that strong confidentiality guarantees are essential for maximizing response rates to censuses and surveys, but little evidence supports this view. Indeed, experimental studies have consistently found that assurances of confidentiality actually increase concerns about confidentiality and reduce response rates to surveys (Berman, McCombs, and Boruch 1977; Frey 1986; Reamer 1979; Singer, Hippier, and Schwarz 1992). A Census Bureau analysis in the 1990s found that promises of confidentiality had no significant impact on response rates (Dillman et al. 1996).
The US Census Bureau recently implemented a new disclosure control strategy that marks a “sea change for the way that official statistics are produced and published” (Garfinkel, Abowd, and Powazek 2018, p. 136). The new disclosure control system adds deliberate error to every population statistic for every geographic unit smaller than a state, including metropolitan areas, cities, and counties.
Population data describing small geographic areas are essential for core political functions like drawing boundaries for state legislative districts and for the US House of Representatives. Towns, cities, counties, states, and the federal government use small area statistics for planning and policy purposes, ranging from decisions about delivery of public services to infrastructure needs. Moreover, economists and other social scientists rely on these demographic data to understand social changes and to evaluate policy outcomes. There is no historical precedent and no demonstrated need for introducing deliberate error into every population statistic for geographic units below the state level. These steps do nothing to allay public concern about the invasion of privacy by the government or about government misuse of data.
This article provides a critical history of the Census Bureau’s evolving rationale for implementing this new approach to disclosure control. The Census Bureau justified the change by repeatedly claiming that the disclosure control system used for the 2010 Census revealed the confidential responses of millions of respondents. That claim, I contend, is entirely unsupported. Moreover, this new approach has undermined the utility of the 2020 Census data for scientific and policy research (Ruggles et al. 2019; Hotz et al. 2022; Muralidhar and Domingo-Ferrer 2023a). Ironically, the new system probably provides less confidentiality protection than the targeted disclosure control system it replaced (Kenny et al. 2023; Muralidhar et al. 2024).
Census Bureau Disclosure Controls, 1970–2010
The Census Bureau has used several techniques to protect the confidentiality of respondents over the past half century (McKenna 2018). These traditional statistical disclosure control techniques introduced uncertainty into published data. The most important disclosure control method was “swapping,” which means exchanging a small percentage of households with similar households from a nearby area. In the 2000 and 2010 data, the swapped households had to match on their size and on the number of adults, but they could differ on any other characteristic. Every household had a chance of being swapped, but the algorithm especially targeted records with the most disclosure risk. Swapping focused on households containing persons with a unique combination of characteristics within their census “block.” Census blocks are the smallest geographic unit described by the census. There were about 6.4 million inhabited blocks in 2010, and the median person resided on a block with 109 people. Small census blocks pose the greatest disclosure risk, so the smaller the block, the higher the rate of swapping (Zayatz et al. 2009).
A concrete example of swapping, reported in the New York Times (Hansen 2018), involved the residents of Liberty Island, the site of the Statue of Liberty, which was a census block in 2010. The 2010 Census data show that the block included only one household with two people, a man in his 60s and a woman in her 40s, both identified as Asian. As it turns out, the Times had interviewed the actual couple in 2011 for an article about Liberty Island (Vadukul 2011). The real residents in 2010 were a 59-year-old husband and a 49-year-old wife, both identified as White. Without disclosure control, the published census would have revealed the confidential census responses of the only actual two residents on Liberty Island; in practice, their responses were protected by swapping their household with a household from a nearby block that also consisted of two adults.
The process of data cleaning and editing introduces additional uncertainty into census data. When census information is missing or inconsistent—for a particular attribute of an individual, for an entire individual, or for a whole household—the Census Bureau substitutes (or imputes) information from a similar nearby person or household (Cantwell 2021). This imputation approach is conceptually similar to swapping, but the objective is to improve the data by reducing the number of missing cases, not to protect confidentiality.
Swapping introduces error into some counts, because the swapped households typically do not match on all characteristics. Prior to the 2020 Census, however, swapping and other disclosure control methods did not alter the counts of total population, voting age adults, housing units, or housing occupancy status at any geographic level. Swapping led to some error on other characteristics, but the Census Bureau concluded that “the impact in terms of introducing error into the estimates was much smaller than errors from sampling, non-response, editing, and imputation” (McKenna 2018, p. 24).
The traditional Census Bureau disclosure control strategy focused on ensuring that the identity of respondents—such as name, address, or Social Security number—cannot be inferred from Census publications. The Census Bureau implemented targeted disclosure controls to prevent reidentification attacks, so that an outside adversary cannot positively identify which person provided any particular response. The protections in place from 1970 through 2010 worked extremely well to meet that standard (Lauger, Wisniewski, and McKenna 2014). Indeed, there is not a single documented case of anyone outside the Census Bureau revealing the responses of a particular identified person using data from the decennial census.
The Database Reconstruction Theorem
Despite the unblemished record of swapping-based disclosure control, in 2017 the Census Bureau decided that such methods were inadequate for the 2020 Census and that an entirely new approach was needed. The Bureau justified this decision by citing a “database reconstruction theorem” developed by computer scientists. The Census Bureau argued repeatedly: “The database reconstruction theorem is the death knell for traditional data publication systems from confidential sources” (for example, Abowd 2017, 2018a; Abowd et al. 2019).
The database reconstruction theorem, developed by Dinur and Nissim (2003), examined confidentiality protection in a hypothetical database. Dinur and Nissim envisioned a set of hospital records that include a secret binary medical condition (coded 0 or 1) along with a number of attributes, such as age and sex. In this exercise, researchers using the database are allowed to query the database and obtain the number of cases with a positive condition among patients with any given set of attributes; for example, they might ask for the number of positive cases among women aged 30–39. To protect privacy, the hospital adds random noise to the result of each query. Dinur and Nissim proved that if the added noise is relatively small compared with the size of the database, it may be possible to reconstruct the original data given a sufficient number of random queries.
The hypothetical database query system envisioned by Dinur and Nissim (2003) bears little resemblance to published census tabulations (Muralidhar and Domingo-Ferrer 2023a, 2023b). The census does not allow random queries; instead, the Census Bureau decides in advance on a set of queries and presents the results in published tables. The overwhelming majority of possible queries cannot be answered by reference to these published tables. The Census Bureau’s swapping algorithm used prior to 2020 added some uncertainty to the published tables, but it was very different from the random noise envisioned by Dinur and Nissim. The database reconstruction theorem relies on random noise being injected at the time of each query; the underlying data can be revealed because the random noise varies with each query. By contrast, the uncertainty injected through swapping by the pre-2020 censuses was done only once. Moreover, the uncertainty added by the Census Bureau was not random; rather, it was targeted towards those cases likely to pose the greatest disclosure risk, such as residents of small blocks or persons with unique characteristics on their block.
Dinur and Nissim themselves point out that noise-infused data in a static database not subject to repeated queries can effectively protect confidentiality; they call this “the CD Model, where users get a ‘private’ version of the database (written on a CD)” (Dinur and Nissim 2003, p. 206). The CD model is essentially the same disclosure control strategy employed by the Census Bureau prior to 2020 (Muralidhar and Domingo-Ferrer 2023b). Thus, the database reconstruction theorem does not demonstrate a disclosure threat in published census tables.
The Census Bureau’s Database Reconstruction Experiment
To demonstrate the threat of database reconstruction for census confidentiality, in 2016 the Census Bureau embarked on an experiment to reconstruct individual-level census responses using only the published tables from the 2010 census. The first public acknowledgment of the experiment came in a July 2018 presentation entitled “Staring Down the Database Reconstruction Theorem.” The presentation provided no actual results of the experiment, but it did claim that “the confidential micro-data from the hundred percent detail file can be reconstructed quite accurately” using only published tables (Abowd 2018a).
The first information on the results of the database reconstruction experiment appeared in a December 2018 article in the New York Times, the same article with the example of swapping on Liberty Island described above (Hansen 2018). Based on an interview with Census Bureau Chief Scientist John Abowd, the article explained that the goal of the reconstruction was to identify the age, sex, race, and Hispanic ethnicity for each individual in every census block in the country. Within each block, the Bureau generated a set of individual-level records consistent with the published tables. The article reported: “By this summer, Mr. Abowd and his team had completed their reconstruction for nearly every part of the country. When they matched their reconstructed data to the actual, confidential records—again comparing just block, sex, age, race and ethnicity—they found about 50 percent of people matched exactly” (Hansen 2018, p. 7). Over the next eight months, Census Bureau staff made many additional public presentations and posted a 26-part “tweetorial” describing the threats exposed by the database reconstruction experiment (for example, Abowd 2018a, 2018b, 2018c, 2019; Abowd et al. 2019). These presentations provided no details about how the reconstructions worked, citing a need to keep the specifics confidential until a paper describing the reconstruction could be peer-reviewed for publication (Abowd et al. 2019).
The most detailed description of the Census database reconstruction experiment was the by-product of a lawsuit. On March 10, 2021, the State of Alabama filed a lawsuit against the Department of Commerce and the Census Bureau in US District Court, arguing that the proposed introduction of deliberate error to stymie database reconstruction—which would involve skewing every statistic below the state level—violated the right of the state to receive accurate tabulations of population (Alabama v. Commerce, 3:21-cv-211-RAH [2021]). (I served as an expert witness for the plaintiffs.) The case was dismissed on June 29, 2021, without prejudice to the underlying issues, on technical grounds related to questions of timing and standing. But the litigation resulted in a “declaration” from the US Census Bureau that included a twelve-page appendix providing a substantially more detailed description of the database reconstruction than had previously been available (Abowd 2021). My description below of the Census Bureau experiment draws mainly on this appendix.
The goal of the Census Bureau database reconstruction experiment was to convert tabular data describing the characteristics of census blocks in 2010 into microdata describing the characteristics of each individual on the block. Rearranging tabular data describing population characteristics into an individual-level format need not be complicated. Consider Table 1, a two-by-two table with information on eight people, broken down by race and sex. Table 2 shows exactly the same information as Table 1, but now expressed as individual-level microdata.
Table 1.
Hypothetical Population Data in Tabular Format
| White | Black | |
|---|---|---|
| Male | 2 | 1 |
| Female | 3 | 2 |
Table 2.
Hypothetical Population Data in Microdata Format
| Case number | Race | Sex |
|---|---|---|
| 1 | White | Male |
| 2 | White | Male |
| 3 | White | Female |
| 4 | White | Female |
| 5 | White | Female |
| 6 | Black | Male |
| 7 | Black | Female |
| 8 | Black | Female |
The individual-level characteristics that the Census Bureau attempted to reconstruct, based on published tables, were age, sex, race, and Hispanic ethnicity for the residents of each census block. The published 2010 Census includes a table (number P12A-I) that provides a breakdown of age by sex by race and Hispanic origin for every census block. As with Tables 1 and 2, it is simple to rearrange this table to obtain an individual-level dataset with block, age, sex, and race/ethnicity. Those microdata will be perfect replicas of the tabular data, although like the tabular data they include errors due to swapping, misreporting, nonresponse, and imputation. The detail of such constructed microdata, however, is limited. Most importantly, most ages are given in five-year bins (for example, 0–4, 5–9, and so on). Only five race groups are included, and other races are grouped into “some other race” or “two or more races.” Moreover, combinations of Hispanic origin and race are limited to Hispanic or Non-Hispanic and “White Alone.”
Thus, to reconstruct the database at the block level with greater detail, the Census Bureau turned to additional tables. One of these, Table P14, gives sex by single year of age for persons less than 20 years old on each block. If, for example, there is only one female aged 5–9 on a given block, then Table P14 can provide a specific age for that child. Another important table is PCT012A-N, which provides single years of age by race, ethnicity, and sex, but only at the census tract level. There are multiple blocks in every tract. If a tract contains just one person who is in a given age group, and who is of a specified race and sex, PCT012A-N can be used to assign a specific age to that individual.
Altogether, the Census Bureau used data from nine tables. The reconstruction experiment constructed a system of simultaneous equations consistent with the published tables. The investigators then solved the equations to create a set of hypothetical individual-level records consistent with all nine tables. The Census Bureau then compared those reconstructed microdata with the original census data by searching each block in the reconstructed data for cases that exactly match a case in the census on age, sex, race, and Hispanic ethnicity.
Database Reconstruction: Results and Critique
The declaration from the Alabama court case provided precise match rates between the reconstructed data and the “true” census data. The Census analysts were able to find matches in the reconstructed data for 46.48 percent of the original unswapped Census microdata. How should this finding be interpreted?
The Census Bureau asserted that the experiment showed that “the micro-data from the confidential 2010 Hundred-percent Detail File (HDF) can be accurately reconstructed” using the published census tables (Abowd et al. 2019). Contrary to Bureau claims, an error rate of over 50 percent, however, is not an “accurate” reconstruction. Moreover, an outside attacker would have no means of knowing which of the reconstructed records matched someone in the real population, so the reconstructed data would not allow an outsider to positively identify any particular census respondent (Ruggles et al. 2019).
Perhaps unsurprisingly, given the task the Census Bureau had set for itself, most of the errors were because of age. As noted, the key table needed for reconstruction (P12A-I) provides only binned ages, mostly in five-year age groups, at the census block level. Individual years of age are only available at the census tract level. Inferring single years of age from five-year age bins was the biggest challenge facing the Census Bureau’s database reconstruction experiment.
As an illustration of the problem, consider Census Tract 5.01 in Laramie, Wyoming. In 2010, Tract 5.01 had 13 Black, non-Hispanic males in the 25–29 year age bin, spread over nine blocks. Muralidhar (2022) showed that just for these 13 individuals, there are 308,000 different ways to assign the exact ages from the tract-level data. The Census Bureau’s reconstruction procedure treats each of these 308,000 possibilities as equally likely and assigns the first one encountered. In the entire country, there are trillions of ways to “reconstruct” the population that are consistent with the published tabulations. The overwhelming majority of these reconstructions do not match the real population.1
How should we evaluate a match rate of 46 percent between the database reconstruction and the original census data? Any analysis of efficacy needs a control group. If a clinical trial found that 46 percent of patients who received a particular medication recovered from a disease, that would not prove that the medication is effective. One would also need to know what percent recovered among those receiving a placebo. Similarly, to evaluate the database reconstruction experiment, it is not sufficient to count the matches between the reconstructed population and the real population. Rather, we need a baseline to assess how much the reconstruction experiment outperforms a null model of random guessing.
To investigate the role of chance in the Census Bureau’s database reconstruction, Ruggles and Van Riper (2022) constructed a simple Monte Carlo simulation. The analysis estimated that a randomly selected person drawn from the 2010 population would exactly match the age and sex of somebody on a randomly selected block 52.6 percent of the time. One would therefore expect the Census Bureau to be “correct” on age, sex, and block most of the time even if they had never looked at the tabular data from 2010 and had instead just assigned ages and sexes to a hypothetical population at random.
The simulation by Ruggles and Van Riper (2022) does not factor in race or ethnicity, but because of high residential segregation, most blocks are highly homogenous with respect to race and ethnicity. If we assign everyone on each block the most frequent race and ethnicity of the block, then race and ethnicity assignment is correct in 77.8 percent of cases (Manson et al. 2022). Using that method to adjust the random age-sex combinations described above, 40.9 percent of cases would be expected to match on all five characteristics to a respondent on the same block. That does not differ greatly from the Census Bureau’s reported 46.5 percent match rate for its reconstructed data (Abowd 2021, App B, p. 3). The analysis suggested that among the minority of cases where the Census Bureau did find a match between their hypothetical population and a real person, most of those matches would be expected to occur purely by chance.
Recently, a group of Census Bureau analysts argued that the Ruggles and Van Riper analysis “was severely flawed” (Jarmin et al. 2023). To estimate a baseline match rate, Ruggles and Van Riper (2022) estimated the likelihood that a randomly-selected individual would be present on a randomly-selected block. Jarmin et al. argue that Ruggles and Van Riper should instead have closely modeled the matching procedure used by the Census Bureau, and they provided previously undocumented details of that algorithm. It is straightforward to adjust the Monte Carlo simulation to conform to the Census Bureau’s specifications (Ruggles and Van Riper 2023). The modification reduces the match rate obtained by random assignment, but the overall conclusion is unchanged. The revised simulation still shows that the great majority (68.6 percent) of the matches between the Census Bureau’s reconstructed data and the original census data would be expected with random assignment of age and sex and systematic assignment of race and ethnicity. Thus, relatively little additional information about census respondents is gained by drawing on the published census tables in the reconstruction exercise.
The main point is that a match rate is not a valid indicator of disclosure risk. Dick et al. (2023)—prominent advocates of the database reconstruction theorem—endorse this core argument. They write that Ruggles and Van Riper (2022) “does raise the important question of choosing appropriate baselines and what we should measure, beyond top-level reconstruction rates, to indicate when we should view reconstruction attacks as worrisome” (Dick et al. 2023, p. 2).
The Census Bureau Reidentification Experiment
Reconstruction and reidentification are very different, and only reidentification poses a meaningful risk to census respondents (Muralidhar and Domingo-Ferrer 2023a). The reconstructed data produced by the Census Bureau contain no personal identifiers. A disclosure attack requires reidentification, meaning that the attacker must accurately link the reconstructed data to an external database that reveals people’s names.
The standard procedure in reidentification studies is to look for unique combinations of variables in individual-level research data (such as survey data) that match the same variables in an external source. Those unique matches are considered vulnerable to reidentification. A reidentification is confirmed only if the identity of the person—ordinarily verified by name—is the same in the research data and in the external data (McKenna 2019; McKenna and Haubach 2019). Thus, for example, a medical record in an anonymized research dataset containing exact date of birth, sex, and zip code might be uniquely matched to public voter records that include the same three variables, thus revealing the name and address of the patient (Barth-Jones 2012). That reidentification would be “confirmed” only if the name on the internal medical records matched the name in the voter file.
Initially, Census Bureau presentations on database reconstruction consistently maintained that although the reconstruction was accurate, the risk of reidentification was small (for example, Abowd 2018a, 2018b; Abowd et al. 2019). In February 2019, the Census Bureau abruptly reversed course, asserting that they had established “confirmed reidentifications” for 17 percent of the 2010 population, exposing the confidential responses of 53 million identified respondents (Abowd 2019).
For the reidentification component of the Census experiment, the Census Bureau used financial and marketing data from five commercial vendors: Experian, Targus, Veteran Service Group of Illinois, InfoUSA, and Melissa. The information in the commercial files was drawn from credit agency reports, magazine subscription records, utility bills, property tax records, voter registration data, and other sources. These files consistently provided information on name, address, age, and sex, and sometimes other variables such as race and Hispanic origin. Shortly after the 2010 Census, the Census Bureau merged all the files, removed the duplicates, and matched them to the 2010 Census returns using the information on name, address, age, and sex. At that time, the goal of the study was to assess the feasibility of using commercial data to improve or substitute for enumerated census data (Rastogi and O’Hara 2012).
Figure 1 diagrams the reidentification process. Each case in matched commercial data should have the same name, address, age, and sex as the corresponding record in the original census file. The Census Bureau assigned a unique numeric code called a “Protected Identification Key” (PIK) to each individual; for each matched individual, an identical PIK appeared in both the original census file and in the matched commercial file.
Figure 1. The Census Bureau Reidentification Experiment.

Source: Author’s creation.
The first step in the reidentification experiment was to search the matched commercial file for a record that matched each record in the reconstructed microdata on block, sex, and age. This was done by “looping through all the records in the reconstructed microdata file produced from the reconstruction, find the first record in the source file [the commercial data] that matches exactly on block, sex, and age” (Abowd 2021: App. B, p. 7). This match yields the “putative” reidentifications, which represent 45 percent of the population.2 The Census Bureau then attached the PIK code from the commercial data to the putative reidentifications.
As the Census declaration in the Alabama v. Commerce lawsuit noted: “Putative reidentifications are not necessarily correct” (Abowd 2021: App. B, p. 8). The putative reidentifications are just cases where a row of the reconstructed microdata has the same block, age, and sex as an individual in the commercial database. If there are multiple people in the reconstructed microdata of the same age and sex—as often occurred—the first one encountered is considered the putative reidentification.
The Census Bureau considers putative reidentifications “confirmed” if they match someone in the original census on the PIK number, race, and Hispanic origin in addition to block, age, and sex. A core problem, however, is that the reconstructed data do not have a PIK; the PIK in the putative file is copied from the commercial data. The Census Bureau assigned PIKs to both the commercial data and the original census data based on name, address, age, and sex (Wagner and Lane 2014). Because the PIK code in the putative data comes from the commercial data, it ordinarily should match a PIK in the original census data. Those cases should also match the census on block, age, and sex, because address, age, and sex were used to assign the PIKs in the first place. Using PIK, block, age, and sex to confirm putative matches is therefore circular: all a match really means is that someone in the commercial data also appears in the census.
To be designated as confirmed, putative cases must also match the original census data on race and Hispanic origin (in addition to block, age, sex, and PIK). As noted earlier, because of the high residential segregation in the United States, most blocks have little diversity with respect to race and ethnicity. Indeed, 78 percent of the population from the 2010 Census identified with the most frequent race and ethnicity found on their block. One would therefore expect the great majority of putative reidentifications to be confirmed just because most people’s race and ethnicity matches the race and ethnicity of most other people on their block.
Given the circular reasoning, one would anticipate a high rate of confirmation of putative reidentifications. The Census Bureau, however, reports strikingly low confirmation rates. The confirmed reidentifications comprise just 38 percent of the putative reidentifications, or 17 percent of the whole population (that is, 38 percent of the 45 percent of putative reidentifications were confirmed). The most plausible explanation for such a low confirmation rate is that the commercial data are highly incomplete and inaccurate.3
The reidentification procedure used by the Census Bureau may sound superficially similar to a standard reidentification study, but it is actually very different. Reidentification studies use names to confirm identity. In the Census Bureau experiment, the external source—the matched commercial file—had already been matched to the original census data using name, address, age, and sex prior to the experiment. The Census Bureau could not use independently acquired information on name, as would be done in a standard reidentification study, because the reconstructed data do not include independent information on name. By the standards and procedures used in prior Census Bureau reidentification studies, none of the reidentifications of the 2010 reconstructed data would be considered confirmed (McKenna 2019; McKenna and Haubach 2019).
Implications of the Reidentification Experiment
Does the Census Bureau’s reidentification experiment demonstrate a realistic threat to confidentiality? The Census Bureau argues that their reidentification could reveal a respondent’s confidential responses to the race and ethnicity questions (Abowd 2021; Garfinkel 2023). The idea is that an external attacker could infer an individual’s race and Hispanic origin by matching reconstructed data on age, sex, and block to an external source that revealed identity, thus making a putative reidentification.
Based on the Census Bureau’s own analysis, this approach would be highly inaccurate; 62 percent of the putative reidentifications were definitely incorrect. The exercise would also be pointless: as Francis (2022) pointed out, the Census Bureau’s elaborate attack strategy is far less reliable than simply inferring race and ethnicity based on the characteristics of the block. By assigning the modal race and ethnicity of the block, one can accurately guess race and ethnicity in about 75 percent of cases. For the 11 percent of the population residing on perfectly homogeneous blocks, one can infer race and ethnicity with perfect accuracy (except for uncertainty introduced by swapping and imputation).4
In a recent working paper, Census Bureau analysts acknowledge that Francis (2022) was correct in the great majority of cases: for the bulk of the population, the database reconstruction and reidentification exercise could not help an outsider guess race and ethnicity any better than guessing the modal race and ethnicity of the block (Abowd et al. 2023). The Census Bureau analysts argue, however, that reconstruction and reidentification are effective for a particular subset of the reconstructed data: reconstructed rows that do not have the modal race and ethnicity of their block and that are unique on their block with respect to binned age and sex. They call these cases “nonmodal uniques.” Because these cases are unique on their block with respect to binned age and sex, their race and ethnicity can usually be read directly from Table P12A-I; no reconstruction is needed. Among this subset, the paper reports that about one in six cases (representing 0.19 percent of the total population) matched someone in the commercial data on binned age and sex. These are the “putative reidentifications.” The authors conclude that the putative reidentifications of nonmodal uniques “definitively show” that the published tables “result in confidentiality breaches” (Abowd et al. 2023, p. 47).
Without access to confidential internal census data, an outside attacker on the 2010 Census would have no means to gauge the reliability of attempted reidentifications of the nonmodal uniques. Because swapping targets cases with unique characteristics, a potential attacker would likely assume an exceptionally high error rate for this group. Abowd et al. (2023) have now revealed a somewhat lower error rate for the putative nonmodal uniques than one might expect.5 Nevertheless, many of the race and ethnic inferences are still incorrect, and an outside attacker has no means of confirming whether any particular inference is true.
The database reconstruction and reidentification as implemented by the Census Bureau posed no realistic threat to the confidentiality of the 2010 Census. As explained above, the database reconstruction theorem does not apply to published census tables. The Census Bureau’s reconstructed data were decidedly inaccurate, performing little better than a random number generator. Even if the Census Bureau could somehow improve the quality of the reconstruction, the only thing to reconstruct would be the swapped version of the data, so there would always be uncertainty about whether any given row of the reconstruction appeared on the block in real life. Reidentification studies ordinarily require a match on name or another reliable identifier to confirm any putative reidentification (McKenna 2019). The reconstructed data have no such identifiers; therefore, the only confirmed links possible are between the commercial data and the census, not between the reconstructed data and the census.
In a 2019 blog post, the Acting Census Bureau Director acknowledged: “The accuracy of the data our researchers obtained from this study is limited, and confirmation of re-identified responses requires access to confidential internal Census Bureau information … more than half of these matches are incorrect, and an external attacker has no means of confirming them” (Jarmin 2019). Five years later, that assessment has proven valid: the kind of reconstruction and reidentification attack used in the Census Bureau’s experiment does not allow positive identification of any census respondents.
Differential Privacy and Noise-Infused Tabular Census Data
In September 2017, the Census Bureau announced to the Census Scientific Advisory Committee that the 2020 Census would abandon traditional disclosure controls such as swapping and instead use a new “differentially private” disclosure avoidance system to ensure confidentiality, in response to the threat allegedly posed by the database reconstruction and reidentification (Garfinkel 2017).6 The application of differential privacy to census data represents a radical departure from established Census Bureau precedents.
Instead of guaranteeing that census responses cannot be tied to particular individuals, differential privacy guarantees that the presence or absence of any individual case from a database should not significantly affect any database query. The requirement that database outputs do not significantly change when any individual’s data is added or removed has profound implications. In effect, under differential privacy, it is prohibited to reveal characteristics of an individual even if the identity of that individual is effectively concealed. In other words, differential privacy is more concerned about database reconstruction than about reidentification. This redefinition of privacy makes disclosure control significantly more challenging compared with the traditional Census Bureau focus on preventing positive identification of respondents.7
The big advantage of the new definition of privacy is that it is relatively simple to formalize, and that formalization yields a metric summarizing a database’s level of “privacy” in a single number. The core metric used in the differential privacy literature is epsilon (ε), which is often referred to as the “privacy-loss budget.” When ε is large, noise infusion is limited and confidentiality protection (under the new definition) is low; when ε is small and near-zero, noise infusion is large, and disclosure control is high. Of course, adding a high degree of randomness to many data parameters can also make census data less useful, or not useful at all, for purposes of public policy and research (Dwork et al. 2006; Bambauer, Muralidhar, and Sarathy 2014).
Based on the theory of differential privacy, the Census Bureau implemented an elaborate procedure to inject random noise into tabular census data. The noise-infusion algorithm is clearly documented in reports of the National Academies of Science, Engineering, and Medicine (Committee on National Statistics 2020, 2023a; Sullivan and Cork 2022). There are five main steps. The process begins by constructing tabular data from the individual-level census returns. Second, a controlled amount of random statistical noise is added to each cell of the tables; the amount of noise used varies according to geographic level and variable. The third step is post-processing, to make sure that the tables do not include logical impossibilities (such as negative population counts) and are internally consistent. Fourth, the Census Bureau converts the processed tables into microdata using the same database reconstruction method they had used to support a need for differentially private census data. Finally, the reconstructed microdata are tabulated to prepare standard census tables for public release.8
To test this data production algorithm, the Census Bureau released, between October 2019 and April 2021, a series of six “demonstration files” based on the 2010 Census (Van Riper, Kugler, and Schroeder 2020–2023). These demonstration files allowed external investigators to assess the usability of the noise-infused data for research and public planning purposes. Most external analysts concluded that all these demonstration files were unfit for critical research and policy applications. Because of the post-processing step, the disclosure avoidance system not only introduced unacceptable levels of random error for many applications of the census but also introduced systematic biases. For example, the demonstration files systematically reduced the size of urban and suburban populations and increased rural populations. The demonstration files also reduced minority populations where minorities are highly concentrated and increased them in areas with lower minority concentration (Santos-Lozada, Howard, and Verdery 2020). The test data undercounted mixed-race and mixed-partisan precincts, posing serious concerns for redistricting (Kenny et al. 2021). They substantially distorted net migration estimates, making migration rate calculations unusable in about half of counties (Winkler et al. 2022). They distorted COVID-19 mortality rates and measures, sometimes causing mortality rates to exceed 100 percent (Hauer and Santos-Lozada 2021). The demonstration data also introduced systematic error into measures of residential segregation (Asquith et al. 2022).
Newspaper reports and social media piled on, highlighting anomalies and internal inconsistencies in the demonstration files (for example, Capps 2021; Menger 2021; Schneider 2021; Wines 2022). In some cases, the noise infusion algorithm made occupied neighborhoods vanish; in other cases, the algorithm populated uninhabited blocks. The data included hundreds of thousands of “Lord of the Flies Blocks” consisting entirely of children with no adults present; “Mermaid Blocks” consisting of people residing in vacant housing units located in lakes and rivers; and “Ghost Blocks” with occupied homes but zero population.
When the Census Bureau announced in June 2021 the final specifications for the electoral redistricting data, they shocked the data user community by specifying far less noise than they had used for the demonstration files. As a result, the corresponding privacy-loss budget ε—the statistic summarizing the amount of error introduced—was also many times higher than is ordinarily contemplated by privacy researchers. Recall that when the summary metric ε is large, noise infusion is limited, and confidentiality protection is low. The range of ε in the differential privacy literature generally runs from 0.01 to 5.0, but many analysts argue that, to guarantee privacy, ε should not greatly exceed 1.0 (Lee and Clifton 2011; Dwork 2011). Several years ago, Apple announced that it would use a differential privacy approach to protecting personal data, but with a value of ε = 14. Frank McSherry, one of the co-inventors of differential privacy, remarked at the time that “anything much bigger than one is not a very reassuring guarantee.” He argued, “Apple has put some kind of handcuffs on in how they interact with your data. It just turns out those handcuffs are made out of tissue paper.” McSherry went on to describe Apple’s disclosure controls as “relatively pointless” (as reported in Greenberg 2017; Domingo-Ferrer, Sánchez, and Blanco-Justicia 2021).
In response to the scholarly and public criticisms of the demonstration files, the Census Bureau used an even higher ε = 19.61 for the final specifications of the redistricting data file. Because the scale of the privacy loss budget is exponential, this privacy-loss budget for the redistricting file was “exponentially higher” than the highest budget used in any of the demonstration files that had been released earlier (US Census Bureau 2021). Accordingly, the Census Bureau’s implementation of differential privacy provides minimal data security. Indeed, this new approach probably provides less confidentiality protection than the traditional disclosure controls (mainly swapping) used by the Census Bureau before 2020.
Since the release of the 2020 data, the Census Bureau has produced new demonstration files for the 2010 Census designed using the differentially private disclosure controls actually used for the 2020 Census. The demonstration files allowed external analysts to assess the impact of the new disclosure controls on the usability of the 2020 data. The studies conducted to date suggest that the reduction in the level of noise has improved data quality, but also that the noise-infused data remain substantially inferior to the data based on disclosure avoidance techniques originally used in 2010, and the noise-infused data remain unusable for some applications. For example, one analysis found that the disclosure control introduced errors exceeding 5 percent in the number of young children for 27 percent of school districts (Committee on National Statistics 2023b, p. 34). Another analysis revealed dramatic differences in household composition among the elderly in the differentially private data (Committee on National Statistics 2023b, p. 44). Mueller and Santos-Lozada (2022) demonstrated that even with ε = 19.61, the noise infusion introduced an unacceptable level of error for small populations, especially non-Whites and inhabitants of rural areas, raising serious questions about the validity of the approach. The final redistricting data still include many internal inconsistencies, including 91,000 blocks with occupied housing units but no people, 101,000 blocks with occupied households but only children present, and 309,000 blocks with people in vacant housing units.
The Census Bureau committed to using differentially private disclosure control in 2017, before the algorithms for adding noise were developed and tested. The design of the system and software was still at an early stage, and implementation ultimately proved far more difficult than anticipated. Indeed, at the time of this writing in early 2024, the Census Bureau is still developing algorithms needed to produce critical tables for the 2020 Census. The implementation of differentially private disclosure controls has led to a years-long delay in the delivery of census tabulations.
When it became clear in June 2021 that the initial algorithm was producing data unfit for use, at the last minute the Census Bureau reduced the noise infusion to levels far below the normal standard required for disclosure protection. Despite the reduced noise, many academic users remain dissatisfied with data quality, and planners and policymakers have lost faith in the reliability of the 2020 statistics for small areas. As more tables were released, the Census Bureau was forced to reduce the amount of noise still further and to raise the privacy-loss budget to the previously unheard-of level of ε = 39.9 (US Census Bureau 2022). To maintain some degree of disclosure control, the Bureau has also substantially reduced the number of published statistics from the 2020 Census relative to the 2010 Census.
There is no evidence that the Census Bureau’s current implementation of differential privacy, using extremely small levels of random noise, protects confidentiality as well as traditional statistical disclosure control. An evaluation commissioned by the Census Bureau concluded that the 2020 disclosure control as implemented “does not provide any comforting guarantees” of confidentiality (JASON 2022, p. 116) and “not enough is known about whether the privacy mechanisms as implemented are sufficient to mitigate the disclosure risks that motivated adoption of formal privacy” (p. 9). Kenny et al. (2023) found that noise-infused data substantially increase the odds of correctly guessing the racial identification of particular respondents. Accordingly, the disclosure control system implemented by the Census Bureau not only reduces the utility of the data, but also provides minimal confidentiality protection (Muralidhar et al. 2024).
Synthetic Microdata
The Census Bureau is on the brink of another disclosure control blunder. The American Community Survey (ACS) Public Use Microdata Sample (PUMS) provides annual samples describing the demographic and economic characteristics of about 3.4 million individuals and 1.3 million households. The ACS PUMS continues a series of large microdata samples produced by the Census Bureau since 1962, providing detailed characteristics of individual respondents. To protect confidentiality, the Census Bureau does not identify places with less than 100,000 population, and the Bureau uses the traditional tools of swapping, top-coding of high values for continuous variables such as income and age, and perturbation of some ages. The ACS PUMS is also protected because it is a sample of just 1 percent of the population. Even if one finds a unique match of certain characteristics between a respondent in the survey and an individual in an external source, one can never be sure that the match is unique in the population as a whole.
The risk of reidentification is very low; the most recent reidentification study of the American Community Survey microdata showed that 0.017 percent of respondents were vulnerable to possible reidentification, but 78 percent of those putative reidentifications were false, and an outsider would have no means of determining which ones were correct (Ramachandran et al. 2012). There is no documented case in which a response to the American Community Survey has ever been linked to an identified person by someone without access to internal Census Bureau data.
Despite this strong record of confidentially protection, the Census Bureau is planning to replace the American Community Survey microdata with “fully synthetic” data to bolster confidentiality (Rodríguez 2021; Daily 2022). The idea of fully synthetic microdata is to develop models describing the interrelationships of all the variables in the data and then use random draws to construct a simulated population consistent with those models (Abowd et al. 2020).
Synthetic data captures relationships between variables only if those relationships have been anticipated in advance and intentionally included in the models. Accordingly, synthetic microdata are poorly suited to studying unanticipated relationships, which impedes new discovery. The large size of the American Community Survey means that it is possible to study small population subgroups, but synthetic data cannot capture all the ways in which interrelationships among variables can vary across subgroups. For example, fully synthetic ACS data would certainly incorporate a general relationship between education and income, but would not assess that relationship separately for every population subgroup. The relationship between education and income might be very different for American Indians in South Dakota compared with Asian Indians in the Queens borough of New York City. Nor would synthetic ACS data capture the myriad possible interrelationships among the characteristics of different family members, such as the relationship between a person’s education and their spouse’s education.
These limitations are important, because the American Community Survey microdata are among the most intensively used scientific data sources in the world and are bedrock resources for demographic and economic research. Tens of thousands of academic researchers, planners, and policy makers rely on the ACS, and according to Google Scholar they generate over 10,000 publications per year. Common topics of analysis include poverty, inequality, immigration, internal migration, ethnicity, residential segregation, disability, transportation, fertility, marriage, occupational structure, education, and family composition. If public use data become unusable or inaccessible, the quantity and quality of research about US policies, the economy, and social structure will decline dramatically.
The Census Bureau acknowledges that the synthetic data will not be reliable enough to support research applications. Consequently, a central element of the plan is to provide “validation” services to researchers. In theory, the researchers will conduct their analyses on the synthetic data and then submit their code to the Census Bureau, which will run the code against the internal “true” data. The Census Bureau will then put the output through disclosure control procedures and give the results back to the researchers (Abowd et al. 2020).
This strategy has multiple flaws. Investigators need access to real data for exploratory analyses to discover relevant variables to incorporate in their analyses, not just for validation. Moreover, as Muralidhar (2023)9 points out, the validation procedure would potentially disclose more information than the current system does, because the validation server would use the original unswapped data; given a sufficient number of validation queries, the underlying data could be subject to exact reconstruction as in the Dinur and Nissim (2003) scenario. The plan is impractical; the Census Bureau lacks the resources to provide this validation service at scale. Executing the plan would be expensive, partly because all the output from the validations must undergo full disclosure review before being released to the researcher. To provide service comparable to the current usage of the American Community Survey microdata, the Census Bureau would have to validate hundreds of thousands of analyses per year.10 If the ACS microdata are replaced with fully synthetic data, that would represent the most damaging loss in access to data describing the US population in the history of our statistical system. The Census Bureau has not provided evidence of disclosure risk from existing practices that would justify this radical change.
Balancing Privacy and Usefulness of Data
Differentially private noise injection was an inappropriate disclosure control choice for the census. The Census Bureau never attempted to weigh realistic measures of disclosure risk under alternative disclosure control methods against the harm of producing an unreliable census (Hotz et al. 2022). The noise-injection algorithm is a blunt instrument that adds deliberate error to every statistic below the state level. Differentially private noise injection is indiscriminate; unlike swapping, it does not target the most vulnerable respondents. Even the total population of New York City is perturbed. This is pointless: tabular data for large populations do not need disclosure control, because there are no cases with unique combinations of characteristics.
For small areas, disclosure control for statistical publications is essential. Without disclosure control, the census responses of the only two people residing on Liberty Island in 2010 would have been compromised, as would the responses of millions of others. Anyone with a unique set of characteristics on a census block is potentially identifiable. Swapping is an attractive tool for disclosure control because it does minimal damage to accuracy, preserves the counts of the number of people and the number of adults at every level of geography, and can be effectively targeted to focus on people at high risk of disclosure. It is not necessary to swap every case with a unique set of characteristics; one need only swap a sufficient proportion to create uncertainty, so that an outside attacker can never be confident of a particular respondent’s identity.
Census law prohibits the positive identification of respondents. The law does not prohibit publication of statistics that can help an attacker guess someone’s characteristics based on place of residence. Indeed, all census statistics improve the chances of guessing respondent characteristics. If I know only that someone lived in Vermont on census day in 2010, I can guess with 94.4 percent confidence that the person identified as White and non-Hispanic (Manson et al. 2022). That is far more precise than estimates of race and ethnicity responses based on the Census Bureau reconstruction and reidentification experiment. In fact, even the tiny subgroup with the highest reidentification accuracy—described by the Census Bureau as the “nonmodal uniques”—has about the same error rate (Abowd et al. 2023).
An external report on the new disclosure controls commissioned by the Census Bureau argues that “the risk that matters is if the released data allows an adversary to make inferences about an individual’s characteristics with more accuracy and confidence than could be done without the data released by the Census Bureau” (JASON 2022, p. 114). This interpretation—if taken literally—would have profound consequences. If the Census Bureau could not publish anything that improves the chances of guessing someone’s characteristics, then all publication of population characteristics—even at the national level—would be prohibited.
To reduce the risk of disclosure of census responses, the Census Bureau could take steps to reduce the number of tabulations with unique combinations of characteristics. Many census blocks are extremely small: in 2010 there were 195,339 blocks with a single person, and 76 million people resided in blocks with fewer than 50 people (Manson et al. 2022). There are not many use cases for such tiny blocks; they could be consolidated into neighboring blocks to reduce disclosure threats. If the Census Bureau consolidated the blocks with fewer than 50 people by merging them with other blocks, that would eliminate 68 percent of the cases that have unique combinations of age, sex, and block. If blocks smaller than 100 were eliminated, that would take care of over 90 percent of the unique age-sex-block combinations.
A similar strategy for minimizing disclosure risk would be reducing the detail of Census tables provided at the block level. Although the Census Bureau has greatly reduced the number of tables to be produced for the 2020 Census, the Bureau oddly did not eliminate the key table that underlies their database reconstruction, the P12A-I table that shows age by sex by race and Hispanic origin for every census block. Surprisingly, the Census Bureau is substantially expanding the detail provided in this table by tabulating a full set of the interactions between Hispanic origin and race, as well as many race combinations (US Census Bureau 2023b). The additional information makes database reconstruction easier and more accurate. The Census Bureau has not explained the reasons for providing this extra detail.
When uncertainty is added to protect confidentiality—either through swapping or through targeted noise injection—that should not affect the total counts of population. Population size is by far the most important census statistic for small areas and should be reported accurately; revealing true population counts need not compromise anyone’s privacy. Accordingly, swapping or noise injection should be designed to add uncertainty to census responses on detailed age, sex, race, ethnicity, or family relationship, without altering the total population counts.
There is a compelling case for the social benefit of broadening access to reliable data from federal agencies (Commission on Evidence-Based Policymaking 2018). Access to high-quality data makes social research and policy formation more reliable, less expensive, and more reproducible. The Foundations for Evidence-Based Policymaking Act of 2018 (2019) requires each federal agency to develop open data plans and make federal data publicly available by default (Commission on Evidence-Based Policymaking 2018). In the ensuing years, however, federal agencies have reduced public access to reliable data because of unproven worries about confidentiality. To ensure that government agencies do not curtail broad access to rich and reliable data without clearly demonstrated need, we need further legislation in two areas. First, we need to clarify that disclosure control laws protect against positive identification of particular respondents, and do not prohibit publishing population characteristics. Second, we should mandate that agencies do not withdraw access to data without balancing the social cost of losing data against realistic measures of the risk of harm to individuals (Hotz et al. 2022).
Supplementary Material
Acknowledgments
■ Work was supported by Alfred P. Sloan Foundation grant G-2019-12589, “Implications of Differential Privacy on Decennial Census Data Accuracy and Utility,” and Eunice Kennedy Shriver National Institute of Child Health and Human Development grant P2C HD041023, “Minnesota Population Center.” I am grateful for the extensive feedback I received from Margo Anderson, Andrew Beveridge, Josep Domingo-Ferrer, Paul Francis, V. Joseph Hotz, Miriam L. King, Diana Magnuson, Robert Moffitt, Krish Muralidhar, Theresa Sullivan, and David Van Riper, as well as JEP editors Heidi Williams, Nina Pavcnik and Erik Hurst. I especially appreciate the exceptional efforts of Managing Editor Timothy Taylor, who reorganized and largely rewrote the entire article.
Footnotes
For supplementary materials such as appendices, datasets, and author disclosure statements, see the article page at https://doi.org/10.1257/jep.38.2.201.
Recently the Census Bureau has improved its success metric for reconstruction by dropping their attempt to infer single years of age, and instead using the binned age groups (Hawes 2022; Abowd and Hawes 2023; Abowd et al. 2023). Using binned ages increased the match rate between the original census data and the reconstructed data from 46.5 to 91.9 percent. Because this reconstruction is effectively just a rearrangement of Table P012A-I into microdata format—much like the rearrangement shown above in Table 2—one might expect the match rate to be perfect. It is imperfect solely because of errors in race and ethnicity (these can arise for non-White Hispanics and for multiple-race respondents) and because of error introduced by swapping. Binned ages are generally not unique on their block and therefore pose much lower disclosure risk than do exact ages.
This is potentially confusing because the 45 percent putative match rate is so close to the 46 percent match rate between the reconstructed data and the original census on block, age, sex, race and Hispanic origin. This is purely a coincidence; the two matched sets are overlapping but distinct.
In the Census Bureau’s newest reidentification estimates, they use binned ages as described in note 1. Then, instead of using the relatively low-quality commercial data, they use the original census data to define putative reidentifications, on the grounds that an external attacker might somehow have access to extremely high quality data (Abowd and Hawes 2023; Abowd et al. 2023). The confirmed match rate (in which there is a match between the putative reidentifications and the original data on Protected Identification Key, block, sex, age group, race, and ethnicity) is 75.5 percent. Given that this exercise is effectively matching the original individual-level census data back to itself, that confirmation rate seems surprisingly low; by definition, 100 percent of the putative rows will match the original census data on PIK, block, sex, and age group. The mismatches can result only from incorrect race or ethnicity or from swapping.
Jarmin et al. (2023) complain that Francis (2022) ignores the effects of swapping and imputation in 2010 (even though the Census Bureau’s reconstruction also ignores swapping and imputation). The results in Alabama v. Commerce (3:21-cv-211-RAH [2021]) suggest that the impact of swapping is small compared with the extremely high error rate in the database reconstruction.
In the past, the Census Bureau kept such statistics secret, because they might enable an attacker to gauge the reliability of an attempted reidentification. Abowd et al. (2023) for the first time reveal the error rate on race and ethnicity for the subset of nonmodal uniques who match someone in the commercial data on age and sex.
In describing the plan, Dajani et al. (2017) describes an “agreement with the Department of Justice” under which the “Census Bureau will provide exact counts at the Census block level” for the total population and the adult (age 18+) population. In August 2020, the Census Bureau posted a revised version of that paper that “supersedes the 2017 version” (Abowd et al. 2020). The paper specifies that exact population counts would not be provided at the block level, nor at the tract, county, or city levels. The only exact population counts in the 2020 Census would be for entire states; every statistic below the state level would have random errors added. In the new iteration, the agreement between the Census Bureau and the Justice Department was not mentioned.
Abowd (2018c: 15) argues that there is a legal basis for the new definition of privacy: “Re-identification risk is only one part of the Census Bureau’s statutory obligation to protect confidentiality. The statute also requires protection against exact attribute disclosure.” This suggests that confidentiality law prohibits revealing respondent characteristics even if the respondent’s identity is protected. The statute in question prohibits “any publication whereby the data furnished by any particular establishment or individual under this title can be identified” (Title 13 USC. § 9, Public Law 87-813). Previously, the Census Bureau interpreted “particular individual” to be an individual whose identity is known.
The procedure was used for the initial redistricting file (P.L. 94-171) and for the “demographic and housing characteristics” file (DHC). The detailed demographic and housing characteristics file (DDHC-A) and the yet-to-be released DDHC-B file and supplemental DHC file use an entirely different noise-infusion algorithm developed by an outside contractor (US Census Bureau 2023a).
This was pointed out by Muralidhar in a personal communication in 2023.
The Census Bureau has suggested that users who need reliable American Community Survey microdata could gain access through the network of Federal Statistical Research Data Centers (FSRDCs). This would be challenging: there are currently only a little over 200 FSRDC projects using census data across the entire network, and disclosure review is already a significant bottleneck. Providing broad access through the data centers would require expansion of the capacity of these centers by multiple orders of magnitude.
References
- Abowd John M. 2017. “Research Data Centers, Reproducible Science, and Confidentiality Protection: The Role of the 21st Century Statistical Agency.” Presentation, Summer DemSem, Wisconsin Federal Statistical RDC, June 5, 2017. https://www2.census.gov/cac/sac/meetings/2017-09/role-statistical-agency.pdf. [Google Scholar]
- Abowd John M. 2018a. “Staring-Down the Database Reconstruction Theorem.” Presentation, Joint Statistical Meetings, Vancouver, BC, July 30, 2018. https://www.census.gov/content/dam/Census/newsroom/press-kits/2018/jsm/jsm-presentation-database-reconstruction.pdf. [Google Scholar]
- Abowd John M. 2018b. “The U.S. Census Bureau Adopts Differential Privacy.” Speech, 24th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, London, August 23, 2018. https://dl.acm.org/doi/10.1145/3219819.3226070. [Google Scholar]
- Abowd John M. 2018c. “Tweetorial: formal privacy for social scientists.” Twitter, November 30, 2018. https://twitter.com/john_abowd/status/1068645579872497664. [Google Scholar]
- Abowd John M. 2019. “Staring Down the Database Reconstruction Theorem.” Presentation, American Association for the Advancement of Science Annual Meeting, Washington, DC, February 16, 2019. https://bpb-us-e1.wpmucdn.com/blogs.cornell.edu/dist/4/7616/files/2019/04/2019-02-16-Abowd-AAAS-Slides-Saturday-330-500-session-FINAL-as-delivered-1iqsdg2.pdf. [Google Scholar]
- Abowd John M. 2021. Declaration of John M. Abowd. Case no. 3:21-CV-211-RAH-ECM-KCN, US District Court for the Middle District of Alabama. April 13. [Google Scholar]
- Abowd John M., Adams Tamara, Ashmead Robert, Darais David, Dey Sourya, Garfinkel Simson L., Goldschlag Nathan et al. 2023. “The 2010 Census Confidentiality Protections Failed, Here’s How and Why.” NBER Working Paper 31995. [Google Scholar]
- Abowd John M., Benedetto Gary L., Garfinkel Simson L., Dahl Scot A., Dajani Aref N., Graham Matthew, Hawes Michael B. et al. 2020. “The Modernization of Statistical Disclosure Limitation at the U.S. Census Bureau.” https://www2.census.gov/adrm/CED/Papers/CY20/2020-009-AbowdBenedettoGarfinkelDahletal-The%20modernization%20of.pdf. [Google Scholar]
- Abowd John M., and Hawes Michael B.. 2023. “Confidentiality Protection in the 2020 US Census of Population and Housing.” Annual Review of Statistics and Its Application 10: 119–44. [Google Scholar]
- Abowd John M., Schmutte Ian M., Sexton William N., and Vilhuber Lars. 2019. “Why the Economics Profession Cannot Cede the Discussion of Privacy Protection to Computer Scientists.” Speech, Allied Social Science Associations Annual Meeting, Atlanta, GA, January 5. https://ecommons.cornell.edu/handle/1813/60836. [Google Scholar]
- Asquith Brian, Hershbein Brad, Kugler Tracy, Reed Shane, Ruggles Steven, Schroeder Jonathan, Yesiltepe Steve, and Van Riper David. 2022. “Assessing the Impact of Differential Privacy on Measures of Population and Racial Residential Segregation.” Harvard Data Science Review (S2). [Google Scholar]
- Bambauer Jane, Muralidhar Krishnamurty and Sarathy Rathindra. 2014. “Fool’s Gold: An Illustrated Critique of Differential Privacy.” Vanderbilt Journal of Entertainment and Technology Law 16 (4): 701–55. [Google Scholar]
- Barth-Jones Daniel. 2012. “The Debate over ‘Re-identification’ of Health Information: What Do We Risk?” Health Affairs Forefront. 10.1377/forefront.20120810.021952. [DOI] [Google Scholar]
- Berman John, McCombs Harriet, and Boruch Robert. 1977. “Notes on the Contamination Method: Two Small Experiments in Assuring Confidentiality of Response.” Sociological Methods and Research 6 (1): 45–63. [Google Scholar]
- Cantwell Pat. 2021. “How We Complete the Census When Households or Group Quarters Don’t Respond.” US Census Bureau, April 16 2021. https://www.census.gov/newsroom/blogs/random-samplings/2021/04/imputation-when-households-or-group-quarters-dont-respond.html. [Google Scholar]
- Capps Kriston. 2021. “Data Scientists Square Off Over Trust and Privacy in 2020 Census.” Bloomberg News, August 12. https://www.bloomberg.com/news/articles/2021-08-12/data-scientists-ask-can-we-trust-the-2020-census. [Google Scholar]
- Commission on Evidence-Based Policymaking. 2018. The Promise of Evidence-Based Policymaking. Washington, DC: Commission on Evidence-Based Policymaking. [Google Scholar]
- Committee on National Statistics. 2020. 2020 Census Data Products: Data Needs and Privacy Considerations: Proceedings of a Workshop. Washington, DC: National Academies Press. [PubMed] [Google Scholar]
- Committee on National Statistics. 2023a. Assessing the 2020 Census: Final Report. Washington, DC: National Academies Press. [Google Scholar]
- Committee on National Statistics. 2023b. 2020 Census Data Products: Demographic and Housing Characteristics File: Proceedings of a Workshop. Washington, DC: National Academies Press. [PubMed] [Google Scholar]
- Daily Donna. 2022. “Disclosure Avoidance Protections for the American Community Survey.” US Census Bureau, December 14. https://www.census.gov/newsroom/blogs/random-samplings/2022/12/disclosure-avoidance-protections-acs.html. [Google Scholar]
- Dajani Aref N., Lauger Amy D., Singer Phyllis E., Kifer Daniel, Reiter Jerome P., Machanavajjhala Ashwin, Garfinkel Simson L. et al. 2017. “The Modernization of Statistical Disclosure Limitation at the U.S. Census Bureau.” United Nations Economic Commission for Europe, Joint UNECE/Eurostat work session on statistical data confidentiality, Skopje, September 20–22. https://unece.org/fileadmin/DAM/stats/documents/ece/ces/ge.46/2017/3_census_bureau.pdf. [Google Scholar]
- Dick Travis, Dwork Cynthia, Kearns Michael, Liu Terrance, Roth Aaron, Vietri Giuseppe, and Wu Zhiwei Steven. 2023. “Confidence-Ranked Reconstruction of Census Microdata from Published Statistics.” Proceedings of the National Academy of Sciences 120 (8): e2218605120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dillman Don A., Singer Eleanor, Clark Jon R., and Treat James B.. 1996. “Effects of Benefits Appeals, Mandatory Appeals, and Variations in Statements of Confidentiality on Completion Rates for Census Questionnaires.” Public Opinion Quarterly 60: 376–89. [Google Scholar]
- Dinur Irit, and Nissim Kobbi. 2003. “Revealing Information While Preserving Privacy.” In Proceedings of the Twenty-Second ACM Sigmod-Sigact-Sigart Symposium on Principles of Database Systems, 202–10. New York: Association for Computing Machinery. [Google Scholar]
- Domingo-Ferrer Josep, Sánchez David, and Blanco-Justicia Alberto. 2021. “The Limits of Differential Privacy (And Its Misuse in Data Release and Machine Learning).” Communications of the ACM 64 (7): 33–35. [Google Scholar]
- Dwork Cynthia. 2011. “A Firm Foundation for Private Data Analysis.” Communications of the ACM 54 (1): 86–95. [Google Scholar]
- Dwork Cynthia, McSherry Frank, Nissim Kobby, and Smith Adam. 2006. “Calibrating Noise to Sensitivity in Private Data Analysis.” In Theory of Cryptography, edited by Halevi Shai, Rabin Tal, 265–84. Heidelberg: Springer Berlin. [Google Scholar]
- Foundations for Evidence-Based Policymaking Act of 2018. P. L. No. 115–435. 132 Stat. 5529. (2019). [Google Scholar]
- Francis Paul. 2022. “A Note on the Misinterpretation of the US Census Re-identification Attack.” In Privacy in Statistical Databases, edited by Domingo-Ferrer Josep and Laurent Maryline, 299–311. Cham, Switzerland: Springer. [Google Scholar]
- Frey James H. 1986. “An Experiment with a Confidentiality Reminder in a Telephone Survey.” Public Opinion Quarterly 50 (2): 267–69. [Google Scholar]
- Garfinkel Simson. 2017. “Modernizing Disclosure Avoidance: Report on the 2020 Disclosure Avoidance Subsystem as Implemented for the 2018 End-to-End Test.” Presentation, Census Scientific Advisory Committee, Suitland, MD, September 15, 2017. https://www2.census.gov/cac/sac/meetings/2017-09/garfinkel-modernizing-disclosure-avoidance.pdf. [Google Scholar]
- Garfinkel Simson. 2023. “Comment to Muralidhar and Domingo-Ferrer (2023)—Legacy Statistical Disclosure Limitation Techniques Were Not an Option for the 2020 US Census of Population and Housing.” Journal of Official Statistics 39 (3): 399–410. [Google Scholar]
- Garfinkel Simson L., Abowd John M., and Powazek Sarah. 2018. “Issues Encountered Deploying Differential Privacy.” In WPES’18 Proceedings of the 2018 Workshop on Privacy in the Electronic Society, 133–37. New York: Association for Computing Machinery. [Google Scholar]
- Greenberg Andy. 2017. “How One of Apple’s Key Privacy Safeguards Falls Short.” Wired Magazine, September 15. https://www.wired.com/story/apple-differential-privacy-shortcomings/. [Google Scholar]
- Hansen Mark. 2018. “To Reduce Privacy Risks, the Census Plans to Report Less Accurate Data.” New York Times, December 5. https://www.nytimes.com/2018/12/05/upshot/to-reduce-privacy-risks-the-census-plans-to-report-less-accurate-data.html. [Google Scholar]
- Hauer Mathew E., and Santos-Lozada Alexis R.. 2021. “Differential Privacy in the 2020 Census Will Distort COVID-19 Rates.” Socius 7: 2378023121994014. [Google Scholar]
- Hawes Michael. 2022. “Reconstruction and Re-identification of the Demographic and Housing Characteristics File (DHC).” Presentation, Census Scientific Advisory Committee, March 17–18, 2022. https://www2.census.gov/about/partners/cac/sac/meetings/2022-03/presentation-reconstruction-and-reidentification-of-the-dhc.pdf. [Google Scholar]
- Hotz V. Joseph, Bollinger Christopher R., Komarova Tatiana, Manski Charles F., Moffitt Robert A., Nekipelov Denis, Sojourner Aaron, and Spencer Bruce D.. 2022. “Balancing Data Privacy and Usability in the Federal Statistical System. Proceedings of the National Academy of Sciences 119 (31): e2104906119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jarmin Ron. 2019. “Census Bureau Adopts Cutting Edge Privacy Protections for 2020 Census.” US Census Bureau, February 15. https://www.census.gov/newsroom/blogs/random-samplings/2019/02/census_bureau_adopts.html. [Google Scholar]
- Jarmin Ron S., Abowd John M., Ashmead Robert, Cumings-Menon Ryan, Goldschlag Nathan, Hawes Michael B., Keller Sallie Ann et al. 2023. “An In-Depth Examination of Requirements for Disclosure Risk Assessment.” Proceedings of the National Academy of Sciences 120 (43): e2220558120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- JASON. 2022. Consistency of Data Products and Formal Privacy Methods for the 2020 Census. (JSR-21-02, January 11, 2022). The MITRE Corporation. https://perma.cc/XJS8-ADX6. [Google Scholar]
- Kenny Christopher T., Kuriwaki Shiro, McCartan Cory, Rosenman Evan T. R., Simko Tyler, and Imai Kosuke. 2021. “The Use of Differential Privacy for Census Data and its Impact on Redistricting: The Case of the 2020 US Census.” Science Advances 7 (41): eabk3283. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kenny Christopher T., Kuriwaki Shiro, McCartan Cory, Rosenman Evan T. R., Simko Tyler, and Imai Kosuke. 2023. “Comment: The Essential Role of Policy Evaluation for the 2020 Census Disclosure Avoidance System.” Harvard Data Science Review (S2). [Google Scholar]
- Lauger Amy, Wisniewski Billy, and McKenna Laura. 2014. “Disclosure Avoidance Techniques at the U.S. Census Bureau: Current Practices and Research.” Research Report Series, Disclosure Avoidance #2014-02. Washington, DC: US Census Bureau. [Google Scholar]
- Lee Jaewoo, and Clifton Chris. 2011. “How Much Is Enough? Choosing ε for Differential Privacy.” In International Conference on Information Security, edited by Lai Xuejia, Zhou Jianying, and Li Hui, 325–40. Heidelberg: Springer Berlin. [Google Scholar]
- Manson Steven, Schroeder Jonathan, Van Riper David, Kugler Tracy, and Ruggles Steven. 2022. “IPUMS [Google Scholar]
- National Historical Geographic Information System: Version 17.0 [dataset].” Minneapolis, MN: IPUMS. 10.18128/D050.V17.0 (accessed June 3, 2023). [DOI] [Google Scholar]
- McKenna Laura. 2018. “Disclosure Avoidance Techniques Used for the 1970 through 2010 Decennial Censuses of Population and Housing.” US Census Bureau Working Paper CES-18-47. [Google Scholar]
- McKenna Laura. 2019. “U.S. Census Bureau Reidentification Studies.” US Census Bureau Working Paper CED-WP-2019-008. [Google Scholar]
- McKenna Laura, and Haubach Matthew. 2019. “Legacy Techniques and Current Research in Disclosure Avoidance at the US Census Bureau.” US Census Bureau Working Paper CED-WP-2019-005. [Google Scholar]
- Menger Elyse. 2021. “Mermaids and Census Privacy Concerns.” Applied Geographic Solutions September 2. https://appliedgeographic.com/2021/09/mermaids-and-census-privacy-concerns/. [Google Scholar]
- Mueller J. Tom, and Santos-Lozada Alexis R.. 2022. “The 2020 US Census Differential Privacy Method Introduces Disproportionate Discrepancies for Rural and Non-White Populations.” Population Research and Policy Review 41 (4): 1417–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Muralidhar Krishnamurthy. 2022. “A Re-examination of the Census Bureau Reconstruction and Reidentification Attack.” In Privacy in Statistical Databases, edited by Domingo-Ferrer Josep and Laurent Maryline, 312–23. Cham, Switzerland: Springer. [Google Scholar]
- Muralidhar Krishnamurthy, and Domingo-Ferrer Josep. 2023a. “Database Reconstruction Is Not So Easy and Is Different from Reidentification.” Journal of Official Statistics 39 (3): 381–98. [Google Scholar]
- Muralidhar Krishnamurthy, and Domingo-Ferrer Josep. 2023b. “A Rejoinder to Garfinkel (2023)—Legacy Statistical Disclosure Limitation Techniques for Protecting 2020 Decennial US Census: Still a Viable Option.” Journal of Official Statistics 39 (3): 411–20. [Google Scholar]
- Muralidhar Krishnamurthy, Domingo-Ferrer Josep, Sanchez David, and Ruggles Steven. 2024. “Protecting Vulnerable Respondents: A Critical Analysis of the Privacy-Preserving Methods of the 2010 and 2020 Decennial Census.” Unpublished. [Google Scholar]
- Ramachandran Aditi, Singh Lisa, Porter Edward, and Nagle Frank. 2012. “Exploring Re-Identification Risks in Public Domains.” In Tenth Annual International Conference on Privacy, Security and Trust, 35–42. Piscataway, NJ: Institute of Electrical and Electronics Engineers. [Google Scholar]
- Rastogi Sonya and O’Hara Amy. 2012. 2010 Census Match Study: Final Report. 2010 Census Planning Memoranda Series 247. Washington, DC: US Census Bureau. [Google Scholar]
- Reamer Frederic G. 1979. “Protecting Research Subjects and Unintended Consequences: The Effect of Guarantees of Confidentiality.” Public Opinion Quarterly 43 (4): 497–506. [Google Scholar]
- Rodríguez Rolando A. 2021. “Disclosure Avoidance and the American Community Survey.” Presentation, 2021 ACS Dara Users Conference, May 20. https://acsdatacommunity.prb.org/discussion-forum/m/2021-acs-conference-files/147/download. [Google Scholar]
- Ruggles Steven, Fitch Catherine, Magnuson Diana, and Schroeder Jonathan. 2019. “Differential Privacy and Census Data: Implications for Social and Economic Research.” AEA Papers and Proceedings 109: 403–08. [Google Scholar]
- Ruggles Steven, and Magnuson Diana L.. 2023. “‘It’s None of Their Damn Business’: Privacy and Disclosure Control in the U.S. Census, 1790–2020.” Population and Development Review 49 (3): 651–79. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ruggles Steven, and Van Riper David. 2022. “The Role of Chance in the Census Bureau Database Reconstruction Experiment.” Population Research and Policy Review 41 (3): 781–88. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ruggles Steven, and Van Riper David. 2023. “Response to the Census Bureau Critique of ‘The Role of Chance in the Census Bureau Database Reconstruction Experiment.” Open Science Foundation Working Paper. https://osf.io/my64r. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sánchez David, Domingo-Ferrer Josep, Muralidhar Krishnamurty. 2023. “Confidence-Ranked Reconstruction of Census Records from Aggregate Statistics Fails to Capture Privacy Risks and Reidentifiability” Proceedings of the National Academy of Sciences 120 (18): e2303890120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Santos-Lozada Alexis R., Howard Jeffrey T., and Verdery Ashton M.. 2020. “How Differential Privacy Will Affect our Understanding of Health Disparities in the United States.” Proceedings of the National Academy of Sciences 117 (24): 13405–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schneider Mike. 2021. “People, Homes Vanish Due to 2020 Census’ New Privacy Method.” Associated Press, October 31. https://apnews.com/article/religion-wisconsin-new-york-tampa-florida-68c96e7eb701da74ae7c8df3c3476705. [Google Scholar]
- Singer Eleanor, Hippier Hans-Jürgen, and Schwarz Norbert. 1992. “Confidentiality Assurances in Surveys: Reassurance or Threat?” International Journal of Public Opinion Research 4 (3): 256–68. [Google Scholar]
- Sullivan Teresa A., and Cork Daniel L., eds. 2022. Understanding the Quality of the 2020 Census: Interim Report. Washington, DC: National Academies Press. [Google Scholar]
- State of Alabama et al. v. US Department of Commerce et al. 3:21-CV-211-RAH-ECM-KCN (2021). [Google Scholar]
- US Census Bureau. 2021. “Census Bureau Sets Key Parameters to Protect Privacy in 2020 Census Results.” Press Release CB21-CN.42, June 9. https://www.census.gov/newsroom/press-releases/2021/2020-census-key-parameters.html.
- US Census Bureau. 2022. “Privacy Loss Budget Allocation.” https://www2.census.gov/programs-surveys/decennial/2020/program-management/data-product-planning/2010-demonstration-data-products/02-Demographic_and_Housing_Characteristics/2022-03-16_Summary_File/2022-03-16_Privacy-Loss_Budget_Allocations.pdf.
- US Census Bureau. 2023a. “Detailed Demographic and Housing Characteristics File A (Detailed DHC-A) Proof of Concept.” https://www2.census.gov/programs-surveys/decennial/2020/program-management/data-product-planning/2010-demonstration-data-products/03-Detailed_ DHC-A/2023-01-31/Proof_of_Concept.pdf.
- US Census Bureau. 2023b. “2020 Census Data Table Guide.” https://www2.census.gov/programs-surveys/decennial/2020/program-management/data-table-guide-dhc-dp.xlsx.
- Vadukul Alex. 2011. “Relaxing under Liberty’s Shadow.” New York Times, August 12. https://www.nytimes.com/2011/08/14/nyregion/on-sundays-david-luchsinger-relaxes-under-lady-libertys-shadow.html. [Google Scholar]
- Van Riper David, Kugler Tracy, and Schroeder Jonathan. 2020–2023. “IPUMS NHGIS Privacy-Protected 2010 Census Demonstration Data [database].” Minneapolis, MN: IPUMS. https://www.nhgis.org/privacy-protected-2010-census-demonstration-data (accessed June 3, 2023). [Google Scholar]
- Wagner Deborah, and Layne Mary. 2014. “The Person Identification Validation System (PVS): Applying the Center for Administrative Records Research and Applications’ (CARRA) Record Linkage Software.” US Census Bureau CARRA Working Paper 2014–01. [Google Scholar]
- Wezerek Gus, and Van Riper Daniel. 2020. “Changes to the Census Could Make Small Towns Disappear.” New York Times, February 6. https://www.nytimes.com/interactive/2020/02/06/opinion/census-algorithm-privacy.html. [Google Scholar]
- Wines Michael. 2022. “The 2020 Census Suggests That People Live Underwater. There’s a Reason.” New York Times, April 21. https://www.nytimes.com/2022/04/21/us/census-data-privacy-concerns.html. [Google Scholar]
- Winkler Richelle L., Butler Jaclyn L., Curtis Katherine J., and Egan-Robertson David. 2022. “Differential Privacy and the Accuracy of County-Level Net Migration Estimates” Population Research and Policy Review 41 (2): 417–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zayatz Laura, Lucero Jason, Massell Paul, and Ramanayake Asoka. 2009. “Disclosure avoidance for Census 2010 and American Community Survey Five-Year Tabular Data Products.” Research Report Series 2009-10. Washington, DC: Statistical Research Division, US Census Bureau. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
