Skip to main content
American Journal of Epidemiology logoLink to American Journal of Epidemiology
. 2025 Jan 3;194(5):1179–1181. doi: 10.1093/aje/kwae483

Insufficient sample size or insufficient attention to marginalized populations? A practical guide to moving observational research forward

Naomi Harada Thyden 1,
PMCID: PMC13016713  PMID: 39756380

Editor’s note: The opinions expressed in this article are those of the authors and do not necessarily reflect the views of the American Journal of Epidemiology.

Introduction

In epidemiology, public health, and population health, large sample sizes and precise estimates (small confidence intervals) are ideal. But if there are important questions that have not been answered and those two things are not possible, how should we proceed?

In 2020, COVID-19 mortality inequities by race/ethnicity, Derek Chauvin’s murder of George Floyd, and increased discrimination against Asian Americans converged to put a spotlight on persistent structural inequities in the United States. Across fields, including public health and epidemiology, there is now a broader understanding that structural discrimination causes health inequities.1-4

As part of that understanding, calls have emerged for more data on marginalized groups. We need data on Native Hawaiian/Pacific Islanders,5 and disaggregated ethnicity data on Asian Americans6-8 and Black immigrants.9 We need data on transgender men and women, nonbinary people, and other sexual and gender minority populations.10,11 We need data on American Indian and Alaska Native populations12-14 that respect differences between sovereign nations. Yet it is still common practice to exclude and obscure these groups in population-level analyses because of insufficient sample size”.

As social epidemiologists, we use epidemiologic methods to study racism within structures such as housing, education, and the criminal legal system. We also need to examine racism within structures in our own field. Many normative statistical practices originated from people who were successful during a time when only a small subgroup of people was allowed to become successful. At the most extreme end, two founders of statistics—Pearson of the “Pearson Correlation Coefficient” and Fisher of “Fisher’s exact test”—were Professors of Eugenics.15

Therefore, it is worth revisiting whether common practices around sample size requirements in observational data are actually best practices or not, and whether they strike the right balance between privacy and our ability to document the experiences of marginalized groups. The following are methodological recommendations for analyzing small samples from marginalized populations in observational data. Implementing these practices will improve the rigor and relevance of health inequities research.

Recommendations for researchers and public health practitioners

Recognize that descriptive analyses have smaller sample size requirements

Rules of thumb for adequate sample sizes often assume that the end goal is multivariable models, which typically require larger sample sizes than descriptive analyses. However, descriptive epidemiology is an essential tool to address health inequities16,17 despite its reputation as being less impressive than causal epidemiology.16 Exposures and outcomes that are more common among marginalized groups than dominant groups are especially likely to be inadequately described. For example, the potent childhood exposure of a parent or child death is concentrated among American Indians but that fact had not been quantified in the literature until recently.18,19 As we integrate structural racism theory into our research practices, we need to acknowledge that important exposures and outcomes have likely been systematically understudied and are therefore in earlier stages of research.

Report overall N's of each racial/ethnic subgroup in your data, even if you do not analyze them further

This recommendation encourages transparency. In situations where marginalized groups actually do have sufficient sample size to be analyzed further, this practice will allow authors to be explicit about their analytic decisions, or to reconsider them. The way in which many papers report race/ethnicity data make it impossible for a reader or reviewer to assess whether it was justifiable to exclude or combine racial/ethnic groups. Peer reviewers encounter assertions of insufficient sample size without supporting data, and further examination of public documentation sometimes reveals hundreds or thousands of discarded participants. It is standard practice to detail exactly how many participants are excluded for a variety of other reasons, and marginalized identities should be treated with similar attentiveness. An example of how to implement this recommendation is: even if analytic models use a four-category race/ethnicity variable, Table 1 or an appendix should display univariates of whatever granular race/ethnicity categories are available. Because univariates do not pose as much of a threat to disclosure as bivariates or model results do, this is one approach to balance transparency and privacy.

Explore ways to thoughtfully increase the analytic sample of the group you are interested in

In situation when sample sizes from marginalized groups are too small to analyze with bivariates or regression models, thoughtfully consider who else is similar enough to be categorized together. While the tendency can be to rely on a “standard” approach such as a four-category race/ethnicity variable, many datasets are able to support a more nuanced approach. For instance, participants who fall into more than one racial/ethnic category could intentionally be classified into the category with fewer people, if it makes theoretical sense to do so. If country of origin is too sparse, perhaps region of origin9 is reasonable. Researchers can perform sensitivity analyses to help determine if groups they are considering combining are similar enough across other relevant variables.

Publish estimates produced from small samples of marginalized groups even if they are imprecise

Some health inequities between groups are so pronounced that they are statistically detectable even with small sample sizes. Other times, analyses from small samples will produce imprecise estimates (ie, large confidence intervals), which are not ideal but are not nothing. In isolation, imprecise estimates might not be convincing, but if publications across datasets produce similar effect size estimates, we have essentially increased the sample size and made progress toward precision.

Avoid analyzing or reporting “other” or “multi-racial” as racial/ethnic categories

A scenario where it is not appropriate to combine categories is the haphazard “other” label. Grouping people into an “other” category erases them from the data and from potential public health action.20 For instance, 14 states did not specify the American Indian/Alaska Native population on their COVID-19 public data dashboard.20 Presented next to racial/ethnic group averages, an “other” estimate might be understandably but erroneously interpreted by the reader as a simple average of the estimates from the remaining racial/ethnic groups. However, an “other” estimate is actually a weighted average, weighted toward whichever remaining group happens to be the largest in that particular dataset—a group which is rarely identified.

Similarly, “multi-racial” is usually too vague to be useful because it encompasses such a wide variety of identities. A person with a white parent and a Black parent is likely perceived as Black in American society, and a person with a white parent and an Asian parent is likely perceived as Asian in American society, yet in research they are often grouped together as “multi-racial.” One option is to create categories for specific combinations of reported racial and ethnic groups (eg, “white and Asian”), which can exhibit distinct health patterns.21 Another option is to create racial/ethnic categories that are not mutually exclusive—in other words, a participant who is multi-racial could be included in each separate race category, with the total across races summing to greater than one hundred percent.

Recommendations for data owners

Facilitate the implementation of the above recommendations

Data owners can aid or hinder researchers’ ability to carry out their aims. For example, there are publicly funded datasets well suited (and even well powered) to study race and racism that are underutilized in part because researchers are unjustifiably denied access to them. In lieu of expertise in race and racism with which to judge an application’s merits, data owners sometimes rely too heavily on sample size considerations. This results in unfairly denied proposals and untapped public resources. In other words, structural racism researchers encounter structural racism in their efforts to obtain data. To combat this, data owners should allow researchers to make a case for analyzing smaller samples than the data owners are accustomed to with their dataset.

Reconsider sample size requirements with the size of marginalized populations in mind

When members of a dominant group make decisions, they often do so with their own community in mind. Perhaps in the past, researchers assumed they could simply recruit more participants to reach sample size goals. However, some marginalized groups simply do not exist in numbers large enough to satisfy sample size requirements. In these cases, are we willing to exclude entire communities from the scope of statistics? Data owners can consider decreasing minimum sample size requirements for populations whose data would otherwise be discarded.

Become familiar with data ownership frameworks outside of the mainstream

Mainstream frameworks for data ownership and ethics emphasize well-founded concerns about privacy and statistical rigor (which therefore prioritize sample size considerations), while complementary frameworks emphasize other value systems. For example, the Collective benefit, Authority to control, Responsibility and Ethics principles guide the use of indigenous data.22 These principles place at the forefront “people” and “purpose”—for instance, indigenous peoples should be involved in how data are collected and interpreted, and the purpose of the resulting analyses should be to their benefit.

Conclusions

It is time to re-examine whether norms about sample sizes in observational data are best practices or not. Knee jerk “insufficient sample size” justifications for failing to collect and analyze data on marginalized communities is dismissive and harmful. This commentary provides concrete steps that researchers and data owners can take to analyze data from smaller marginalized groups with the same care and attention shown to dominant groups. Still, there are no shortcuts for cultural humility23 and respect for people’s lived experiences.

As Abigail Echo-Hawk, Executive Vice President of the Seattle Indian Health Board, and Director of the Urban Indian Health Institute, said, “We are a small population of people because of genocide. No other reason. If you eliminate us in the data, we do not exist.”24 If our standard practices systematically disadvantage minoritized groups, then we need to find ways to correct them.

Acknowledgments

An early version of this commentary was presented as invited remarks at the Interdisciplinary Association for Population Health Science (IAPHS) Conference in 2023 in Baltimore, MD, for the 2023 Student Award sponsored by Mary Amelia Center for Women’s Health Equity Research at Tulane University School of Public and Tropical Medicine.

Funding

The author gratefully acknowledges support from the Minnesota Population Center (P2C HD041023) funded through a grant from the Eunice Kennedy Shriver National Institute for Child Health and Human Development (NICHD).

Conflict of interest

None declared.

References

  • 1. Bailey  ZD, Krieger  N, Agenor  M, et al.  Structural racism and health inequities in the USA: evidence and interventions. Lancet.  2017;389(10077):1453-1463. 10.1016/s0140-6736(17)30569-x [DOI] [PubMed] [Google Scholar]
  • 2. Bailey  ZD, Feldman  JM, Bassett  MT. How structural racism works - racist policies as a root cause of U.S. racial health inequities. N Engl J Med.  2021;384(8):768-773. 10.1056/NEJMms2025396 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Adkins-Jackson  PB, Chantarat  T, Bailey  ZD, et al.  Measuring structural racism: a guide for epidemiologists and other health researchers. Am J Epidemiol.  2022;191(4):539-547. 10.1093/aje/kwab239 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Gee  GC, Ford  CL. Structural racism and health inequities: old issues, new Directions1. Du Bois Rev.  2011;8(1):115-132. 10.1017/S1742058X11000130 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Morey  BN, Chang  RC, Thomas  KB, et al.  No equity without data equity: data reporting gaps for native Hawaiians and Pacific islanders as structural racism. J Health Polit Policy Law.  2022;47(2):159-200. 10.1215/03616878-9517177 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Yom  S, Lor  M. Advancing health disparities research: the need to include Asian American subgroup populations. J Racial Ethn Health Disparities.  2021;9(6):2248-2282. 10.1007/s40615-021-01164-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Muramatsu  N, Chin  MH. Battling structural racism against Asians in the United States: call for public health to make the “invisible” visible. J Public Health Manag Pract.  2022;28(Supplement 1):S3-S8. 10.1097/PHH.0000000000001411 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Yi  SS. Taking action to improve Asian American health. Am J Public Health.  2020;110(4):435-437. 10.2105/AJPH.2020.305596 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Larimore  S, Ifatunji  M, Lee  H, et al.  Geographic variation in reproductive health among the black population in the US: an analysis of nativity, region of origin, and division of residence. Popul Res Policy Rev.  2021;40(1):33-59. 10.1007/s11113-020-09629-0 [DOI] [Google Scholar]
  • 10. Kress  AC, Asberry  A, Taillepierre  JD, et al.  Collection of data on sex, sexual orientation, and gender identity by US public health data and monitoring systems, 2015–2018. Int J Environ Res Public Health.  2021;18(22):12189. 10.3390/ijerph182212189 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Poteat  TC, Logie  CH, van der  Merwe  LLA. Advancing LGBTQI health research. The Lancet.  2021;397(10289):2031-2033. 10.1016/S0140-6736(21)01057-6 [DOI] [PubMed] [Google Scholar]
  • 12. Rhodes  KL, Echo-Hawk  A, Lewis  JP, et al.  Centering data sovereignty, tribal values, and practices for equity in American Indian and Alaska native public health systems. Public Health Rep.  2023;139(1_suppl):10S-15S. 10.1177/00333549231199477 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Small-Rodriguez  D, Akee  R. Identifying disparities in health outcomes and mortality for American Indian and Alaska native populations using tribally disaggregated vital statistics and health survey data. Am J Public Health.  2021;111(S2):S126-S132. 10.2105/AJPH.2021.306427 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Mays  VM, Echo-Hawk  A, Cochran  SD, et al.  Data equity in American Indian/Alaska native populations: respecting sovereign nations’ right to meaningful and usable COVID-19 data. Am J Public Health.  2022;112(10):1416-1420. 10.2105/AJPH.2022.307043 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Krieger  N. Structural racism, health inequities, and the two-edged sword of data: structural problems require structural solutions. Front. Public Health.  2021;9:301. 10.3389/fpubh.2021.655447 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Fox  MP, Murray  EJ, Lesko  CR, et al.  On the need to revitalize descriptive epidemiology. Am J Epidemiol.  2022;191(7):1174-1179. 10.1093/aje/kwac056 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Ward  JB, Gartner  DR, Keyes  KM, et al.  How do we assess a racial disparity in health? Distribution, interaction, and interpretation in epidemiological studies. Ann Epidemiol.  2019;29(January):1-7. 10.1016/j.annepidem.2018.09.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Thyden  NH, Schmidt  NM, Osypuk  TL. The unequal distribution of sibling and parent deaths by race and its effect on attaining a college degree. Ann Epidemiol.  2020;45:76-82.e1. 10.1016/j.annepidem.2020.03.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Thyden  NH, Slaughter-Acey  J, Widome  R, et al.  Family deaths in the early life course and their association with later educational attainment in a longitudinal cohort study. Soc Sci Med.  2023;333:116161. 10.1016/j.socscimed.2023.116161 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Urban Indian Health Institute . Data genocide of American Indians and Alaska Natives in COVID-19 data. Seattle, WA: Urban Indian Health Institute; 2021. [Google Scholar]
  • 21. Tabb  KM, Gavin  AR, Smith  DC, et al.  Self-rated health among multiracial young adults in the United States: findings from the add health study. Ethn Health.  2019;24(5):495-511. 10.1080/13557858.2017.1346175 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Carroll  S, Garba  I, Figueroa-Rodríguez  O, et al.  The CARE principles for indigenous data governance. Data Sci J.  2020;19:43. 10.5334/dsj-2020-043 [DOI] [Google Scholar]
  • 23. Yeager  KA, Bauer-Wu  S. Cultural humility: Essential foundation for clinical researchers. Appl Nurs Res.  2013;26(4):251-256. 10.1016/j.apnr.2013.06.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Nagle  R.  Native Americans being left out of the US coronavirus data and labelled as ‘other’. The Guardian, London, United Kingdom. April 24, 2020. https://www.theguardian.com/us-news/2020/apr/24/us-native-americans-left-out-coronavirus-data [Google Scholar]

Articles from American Journal of Epidemiology are provided here courtesy of Oxford University Press

RESOURCES