Skip to main content
Journal of the American Medical Informatics Association: JAMIA logoLink to Journal of the American Medical Informatics Association: JAMIA
. 2020 Sep 20;27(12):2016–2019. doi: 10.1093/jamia/ocaa133

Addressing health disparities in the Food and Drug Administration’s artificial intelligence and machine learning regulatory framework

Kadija Ferryman 1,
PMCID: PMC7727393  PMID: 32951036

Abstract

The exponential growth of health data from devices, health applications, and electronic health records coupled with the development of data analysis tools such as machine learning offer opportunities to leverage these data to mitigate health disparities. However, these tools have also been shown to exacerbate inequities faced by marginalized groups. Focusing on health disparities should be part of good machine learning practice and regulatory oversight of software as medical devices. Using the Food and Drug Administration (FDA)'s proposed framework for regulating machine learning tools in medicine, I show that addressing health disparities during the premarket and postmarket stages of review can help anticipate and mitigate group harms.

Keywords: health disparities, artificial intelligence, health policy, machine learning

INTRODUCTION

The United States has stark health disparities. Racial and ethnic minority groups have lower overall life expectancy and worse outcomes for major diseases, including heart disease, and higher prevalence of others such as diabetes.1 Though some population-specific biological differences explain some of these disparate outcomes, gaps between groups can largely be attributed to social factors, such as differences in exposure to environmental hazards, access to high-quality food, and effects of structural racism and discrimination in medical facilities and society at large.2,3 The coronavirus disease 2019 (COVID-19) pandemic brings these disparities into sharp relief.4 Addressing these social determinants of health is key to mitigating health disparities and achieving health equity in the United States.

Though addressing social issues are important, technological developments offer promising avenues to pursue. The exponential growth of health data from devices, health applications, and electronic health records coupled with the development of data analysis tools such as machine learning (ML) offer opportunities to leverage these data to mitigate health disparities. However, these tools have also been shown to exacerbate inequities faced by marginalized groups, such as recidivism risk scoring algorithms that have higher error rates for Black defendants and higher error rates for women in mortality prediction algorithms.5,6 Despite the growing evidence of these kinds of negative impacts of algorithmic technologies, the recently proposed regulatory framework governing ML applications in health care from the Food and Drug Administration (FDA) makes explicit mention of neither the potential for group harms nor how these tools could either exacerbate or mitigate health disparities. As ML gains ground in health, considerations of health disparities and the potential for group harms must be central to comprehensive regulatory policy. In this commentary, I argue that in order for the FDA to ensure patient safety with the use ML in medicine, regulation must (1) require attention to health disparities and the potential for group harms as part of the clinical evaluation process in premarket review and (2) continue this focus on disparities and group harms through a “health equity review” that is part of postmarket, real-world performance monitoring. This health equity review can be an inclusive process that includes multiple stakeholders, including device manufacturers, patients, clinicians, and informatics professionals.

THE FDA’S PROPOSED REGULATORY FRAMEWORK FOR ARTIFICIAL INTELLIGENCE AND ML IN MEDICINE

The FDA has issued guidance regarding the use of SaMD (software as a medical device), recognizing the growing importance of data analytics tools in health.7 In April of 2019, the agency proposed a regulatory framework for ML tools in medicine, noting that these tools demand distinctive regulatory consideration because they may change and adapt after their initial development and deployment, and use new data to improve their performance and effectiveness.8 The document features the agency’s effort “to reimagine an approach to premarket review for artificial intelligence (AI)/ML-driven software modifications” that is a “new, total product lifecycle (TPLC) regulatory approach.”8 This approach has several components: a demonstration of good ML practices by the device maker; an assessment of the risk of the ML tool as well as plans for updating it with new data or new algorithm architecture; an ongoing review by device makers to check whether any changes to the ML tool as it is deployed deviates significantly from the premarket proposed changes; and ongoing, postmarket transparency about the tool’s performance through updates to the FDA, collaborators, clinicians, and the public.

As part of premarket review, manufacturers must have good ML practices in place, which the document describes as demonstrating the ML tool’s valid clinical associations, analytical validity, and clinical validity. However, these having these “good practices” in place may not be sufficient to prevent group harms. A recent article by Obermeyer et al9 provides an instructive example of the missteps that can befall an algorithm that does not consider health disparities and the potential for group harm as part of good ML practices. In this case, the amount of money spent on patient care was selected as a proxy measurement of health status, as generally more money is spent on sick people than healthy people.9 The purpose of the model was to direct more healthcare resources to sicker people. However, this tool ended up favoring healthier white patients over sicker Black patients. If the model developers had been required to demonstrate knowledge of health disparities that exist in their domain of interest, they may have discovered that in the United States, less money is spent on Black patients, even when they are sicker.10 This discovery could have led to adjustments to their input data, and scrutiny of the validity of the association between healthcare expenditure and illness before the model was developed. This example shows that even though the model worked, in the sense that it used accurate data and “correctly processed” these data as the FDA outlines in the proposed framework, it still favored white patients over Black patients.

Rajkomar et al11 have argued that even attention to how a model performs between groups may not be enough to ensure that health disparities are not exacerbated. For example, they note that it is impossible for a model to have both equalized odds and equal positive and negative predictive value across groups.11 What this means is that model developers must make these kinds of decisions about performance, and these decisions could have significant impacts, depending on the groups in question and the existing health disparities. The FDA’s current regulatory framework does not draw attention to this in its description of good ML practices. Under the current framework, there is little substantive difference between a decision for the model to have equal performance across groups or to have equal positive and negative predictive values across groups, as either choice would fulfill the requirements for the model to be analytically valid. This is a blind spot in the proposed regulation that must be addressed.

In addition to demonstrating good ML practices, and submitting an algorithm change protocol detailing how the ML tool might change as it learns from ingesting data when it is deployed, the FDA also proposes that device manufacturers conduct “real-world performance monitoring” that documents how the tool performs under actual clinical conditions after the device has been approved. Currently, the proposed regulation asks device makers to update the FDA on the tool’s performance and suggests updating others, such as clinicians and patients. The document also lists potential changes that the agency should be made aware of like “change in inputs,” and updates to the “specifications or compatibility of any impacted supporting devices.”8 The real-world performance of the tool is defined in terms of purely technical aspects, and does not include any mention of how nontechnical factors might impact the tool’s performance, effectiveness, and intended purpose, especially on groups already experiencing health disparities.

Consideration of factors that are known to influence health disparities such as differential healthcare access, structural discrimination, and clinician bias should be part of a health equity review that would be part of the larger process of real-world performance monitoring of AI/ML devices. IDx-DR, the first AI-powered diagnostic SaMD approved by the FDA, can serve as a thought example to illustrate the potential benefits of including a health equity review as part of AI/ML performance monitoring. IDx-DR is an AI-powered tool that detects diabetic retinopathy (an eye disease caused by diabetes); this disease is marked by significant racial disparities. African Americans have a 4 times higher risk of developing this condition than do non-Hispanic whites, and this group is also less likely than whites to have eye examinations, which can detect diseases such as diabetic retinopathy (DR).12

The IDx-DR tool allows for the screening of DR in primary care settings, as the system analyzes the images and provides the diagnosis, rather than the clinician. Because nonspecialist clinicians can use the tool as part of primary care, the IDx-DR tool could lower some of the barriers to receiving eye exams. As an article on the tool notes, “[f]or people with diabetes, autonomous AI systems have the potential to improve earlier detection of DR, and thereby lessen the suffering caused by blindness and visual loss.”13

There is some evidence that cost of testing and lack of or limited health insurance can explain why African Americans are less likely to receive eye examinations than whites are. However, this gap may also be due to clinicians being less likely to offer eye examinations to Black patients.14 Though the data on this particular factor are emerging, it dovetails with other research that shows that clinicians are less likely to offer Black patients multiple forms of medical treatment and screening.15–17

A health equity review that would be part of periodic real-world performance monitoring of the IDx-DR tool would include information on its use across different racial groups. Ostensibly, because this tool is being implemented in primary care, its use across groups should mirror population demographics. If the use rates deviate from what is expected, then over- or underuse of the tool in specific groups should be investigated. For example, if the tool is being used less often with Black patients, this could signal a form of clinician bias, in which the screening tool might not be used as often due to a belief that Black patients would not proceed with follow-up care. The tool might continue to improve its performance in disease detection overall, and even could demonstrate better disease detection within particular groups, but these improvements might be overshadowed if certain groups are less likely to be offered and receive the test over time. If this were the case, instead of lowering barriers, this AI tool might be exacerbating health disparities. Alternately, overuse of the tool among Black patients might suggest that clinicians see this group as at a higher risk of DR, and might be using the tool as a way to respond to higher vulnerability in this patient population. In this case, using the IDx-DR tool in the “real-world” clinical setting might be both accurately detecting this disease overall and helping to ameliorate health disparities.

RECOMMENDATIONS

The FDA regulates products to ensure that they are safe for consumers. With ML tools in health, health disparities are a safety issue, as inattention to potential negative impacts can increase the risk and danger to groups already marginalized and discriminated against in health care. Consideration of health disparities must not be out of scope, or an optional dimension to consider when developing ML tools for medicine. To that end, considerations of health disparities can be integrated into the FDA’s AI/ML regulation in 4 ways, in both the premarket and postmarket stages (Figure 1):

Figure 1.

Figure 1.

This figure is adapted from the FDA’s Overlay of the TPLC approach on AI/ML workflow, in the Proposed Regulatory Framework for Modifications to Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD). Author’s recommendations are in bold italics.

Premarket review and good ML practices

1. Knowledge of Health Disparities. Device manufacturers should include evidence of knowledge of health disparities in their clinical domain of interest. Having this background knowledge and context can aid in the design of the ML tool as well as in the monitoring of its performance. This could be included as part of their demonstration of culture of quality and organizational excellence. This kind of background knowledge will help manufacturers appropriately identify which groups in their clinical area of interest should be classified as protected or marginalized groups.

2. Data Bias Review. Device manufacturers should document decisions that were made regarding the data that were used to train the model, including the representativeness of protected groups in the sample(s). This should include whether there is a representative sample, or if these groups are under- or overrepresented. The data bias review should also document an examination of the potential for latent biases in data, such as data that reflect histories of unequal access to health care, data that may be present but not as informative for all groups, or data that reflect racial, gender, or other corrections that may be clinically questionable or disputed. The Algorithm Change Protocol would include checks to ensure that any new data added to the model would be examined for data bias.

3. Group Impacts and Performance Decisions: Device manufacturers should document decisions that were made regarding the performance of the model, including discussions of how different choices can impact groups already experiencing health disparities in the clinical condition of interest. The rationale for how these decisions were made would ideally include input from multiple stakeholders, including clinicians and patients from groups experiencing health disparities in the clinical domain of interest.

Postmarket real-world performance monitoring

4. Health Equity Review: Device manufacturers’ updates to the FDA should include specific information on how the ML tool is impacting protected groups, and whether any new group differences have emerged because of the deployment of the tool. This update would also include details about any activities undertaken to address unfairness or group harms, even if these activities do not involve changes to the model’s inputs or architecture. Device makers may also identify how the use of their tool has mitigated health disparities. Ideally, this update would reflect an inclusive process that includes engagement, involvement, and feedback from multiple stakeholders including patients, clinicians, informatics professionals, and others.

The FDA’s Office of Minority Health and Health Equity could be involved in implementing these recommendations. This office currently works on issues such as diversifying clinical trials and research on population differences in disease biomarkers, treatments, and health communication and messaging. The office could help guide device manufacturers in expanding their good ML practices to include attention to health disparities, as well as provide assistance in building diverse partnerships for postmarket evaluation and monitoring.

CONCLUSION

These recommendations could be part of a first step in a journey toward comprehensive ML policy for advancing health equity. The FDA has argued that regulatory processes for software used in medicine should be “agile” in order to “accommodate the faster rate of development and potential for innovation in software-based products.”18 This is a laudable goal. However, an emphasis on “agility” should not come at the expense of those already experiencing unfair health access, treatment, and health outcomes. ML offers promises of efficiency, cost savings, and better health outcomes, but regulation of this burgeoning field must explicitly address and call out health disparities, rather than leave unmentioned how these tools could have significant impacts on marginalized groups—both positive and negative. The current COVID-19 pandemic is yet another illustration of just how entrenched health disparities are in the United States. Health care and society at large are once again in the position of reacting to drastic differences in disease incidence, treatment, and outcomes in this pandemic. Though ML tools alone will not address the multiple historical and structural forces that influence health disparities, the FDA could take a more proactive stance in regulation so that these tools anticipate that there could be increased risks for harm for those already experiencing health disparities. The guidance for ML tools in health is still in development, and not yet final, so there is an opportunity to think carefully and deliberately about good governance of AI/ML in medicine.

AUTHOR CONTRIBUTIONS

KF is the sole author of this work.

CONFLICT OF INTEREST STATEMENT

The author has no competing interests to declare.

REFERENCES

  • 1. Centers for Disease Control and Prevention. CDC health disparities and inequalities report—United States, 2013. MMWR Suppl 2013; 62 (3): 1–184. [PubMed] [Google Scholar]
  • 2. Smedley BD, Stith AY, Nelson AR, eds. Institute of Medicine, Committee on Understanding and Eliminating Racial and Ethnic Disparities in Health Care Unequal Treatment: Confronting Racial and Ethnic disparities in Healthcare. Washington, DC: National Academies Press; 2003. [PubMed] [Google Scholar]
  • 3. Butkus R, Rapp K, Cooney TG, Engel LS; Health and Public Policy Committee of the American College of Physicians. Envisioning a better U.S. Health Care System for all: reducing barriers to care and addressing social determinants of health. Ann Intern Med 2020; 172 (2_Supplement): S50. [DOI] [PubMed] [Google Scholar]
  • 4. Eligon J, Burch A. Questions of bias in Covid-19 treatment add to the mourning for Black families. The New York Times. May 10, 2020. https://www.nytimes.com/2020/05/10/us/coronavirus-african-americans-bias.html Accessed May 12, 2020.
  • 5. Angwin J, Larson J, Mattu S, Kirchner L. (2019). Machine bias: There’s software used across the country to predict future criminals. and it’s biased against blacks. 2016. https://www. propublica. org/article/machine-bias-risk-assessments-in-criminal-sentencing Accessed May 12, 2020.
  • 6. Chen IY, Szolovits P, Ghassemi M.. Can AI help reduce disparities in general medical and mental health care? AMA J Ethics 2019; 21 (2): 167–79. [DOI] [PubMed] [Google Scholar]
  • 7.U.S. Food and Drug Administration. Software as a Medical Device (SaMD). http://www.fda.gov/medical-devices/digital-health/software-medical-device-samd Accessed July 2, 2019.
  • 8.U.S. Food and Drug Administration. Proposed Regulatory Framework for Modifications to Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD). https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-and-machine-learning-software-medical-device Accessed April 15, 2020.
  • 9. Obermeyer Z, Powers B, Vogeli C, Mullainathan S.. Dissecting racial bias in an algorithm used to manage the health of populations. Science 2019; 366 (6464): 447–53. [DOI] [PubMed] [Google Scholar]
  • 10. Benjamin R. Assessing risk, automating racism. Science 2019; 366 (6464): 421–2. [DOI] [PubMed] [Google Scholar]
  • 11. Rajkomar A, Hardt M, Howell MD, Corrado G, Chin MH.. Ensuring fairness in machine learning to advance health equity. Ann Intern Med 2018; 169 (12): 866–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Fathy C, Patel S, Sternberg P Jr, Kohanim S.. Disparities in adherence to screening guidelines for diabetic retinopathy in the United States: a comprehensive review and guide for future directions. Semin Ophthalmol 2016; 31 (4): 364–77. [DOI] [PubMed] [Google Scholar]
  • 13. Abràmoff MD, Lavin PT, Birch M, Shah, et al. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. NPJ Digit Med 2018; 1 (1): 39. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Ellish NJ, Royak-Schaler R, Passmore SR, Higginbotham EJ.. Knowledge, attitudes, and beliefs about dilated eye examinations among African-Americans. Invest Ophthalmol Vis Sci 2007; 48 (5): 1989–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Bonham VL. Race, ethnicity, and pain treatment: striving to understand the causes and solutions to the disparities in pain treatment. J Law Med Ethics 2001; 28 (4_suppl): 52–68. [DOI] [PubMed] [Google Scholar]
  • 16. Bach PB, Cramer LD, Warren JL, et al. Racial differences in the treatment of early-stage lung cancer. N Engl J Med 1999; 341 (16): 1198–205. [DOI] [PubMed] [Google Scholar]
  • 17. Wenneker JB. Racial inequalities in the use of procedures for patients with ischemic heart disease in Massachusetts. JAMA 1989; 261 (2): 253–7. [PubMed] [Google Scholar]
  • 18.U.S. Food and Drug Administration. Developing a software precertification program: A working model. https://www.fda.gov/media/119722/download Accessed May 27, 2020.

Articles from Journal of the American Medical Informatics Association : JAMIA are provided here courtesy of Oxford University Press

RESOURCES