McCradden et al. (2022) propose to close the “AI chasm” between algorithms and clinically meaningful application using the norms of evidence-based medicine (EBM) and clinical research, with the rationale that prospective and randomized methods control for biases. While prospective trials can control for confounding bias, and randomization can control for selection bias within the study population, these methods do not address larger problems of representation embedded in the models themselves, especially because of social inequalities inherent in the underlying data. The norms of clinical research have not successfully addressed the broader issues of fairness, and thus will not solve those problems for algorithmic evaluation.
McCradden et al. argue that a clash between the epistemic and ethical cultures of computer science and clinical research accounts for the AI chasm. They describe the epistemic divide between computer science and clinical research as coming down to a difference in methods between these disciplines. They characterize a central tension between the data-driven culture of AI/Machine Learning (ML) and the purpose of research ethics to protect participants from exploitation, particularly in terms of consent and data protection. According to McCradden et al., the main ethical challenge arising from implementation of ML in a clinical context is the potential deviation from standard of care, with bias presenting a source of risk to patients. In order to mitigate these risks, they therefore advocate for rigorous evaluation of medical ML using clinical research norms and methods. Their approach focuses on randomization and prospective study designs to control for biases that threaten the effective translation of AI/ML to clinical applications. However, in flattening the epistemic and ethical clash to a matter of conflicting methods, this approach ignores key issues regarding how bias is defined and addressed.
Bias can be broadly defined as a type of systematic error that can affect scientific investigations at each stage, from initial research question to interpretation and applications of the findings (Sica 2006). While AI/ML and clinical research generally focus on a technical sense of bias as systematic error, ethics is intent on addressing bias that is unfair, leading to unjust outcomes and disparate impacts in subpopulations (Cho 2021). The well-known case of the widely-used healthcare algorithm that exhibited racial bias by reducing the number of Black patients assigned for extra care is a strong example of the kind of bias that is of ethical concern (Obermeyer et al. 2019). In proposing a process that relies on clinical research methods to address bias, McCradden et al. overlook that ethics is primarily concerned with addressing unfair bias, and that the methods of clinical research that they propose to use to evaluate ML are not designed to address issues of fairness (Jadad and Enkin 2007).
There are many different types of bias that have been identified in algorithmic and clinical contexts. The design of randomized or prospective clinical research has generally focused on minimizing selection bias and confounding bias (Lambert 2011). Selection bias refers to systematic error produced by selecting a study population that is not representative of the population sample. Randomization is generally used in order to control for selection bias.
Confounding bias refers to the “spurious association made between the outcome and a factor that is not itself causally related to the outcome, and occurs if the factor is associated with a range of other characteristics that do increase the outcome risk.” (Lambert 2011) Confounding factors may hide a true association between variables or falsely indicate an association between variables, such as an intervention and outcome. Prospective trials, as well as randomization, can be used to address confounding bias, Bias cannot be completely eliminated from scientific investigation. However, in order to effectively mitigate bias, in the ethical sense, processes must be employed that recognize and address the broader social and organizational forces that shape medical data and lead to unfair outcomes.
Clinical research methods are insufficient to address the context-based nature of data and the problems of representation embedded in the models themselves (Miceli, Posada, and Yang 2021). There are a number of ways that social inequalities can be embedded in the underlying data used for developing medical ML. One obvious source of bias is when training data underrepresents people from relevant sub-populations. While this potential source of algorithmic bias has received repeated attention, a study of medical AI/ML approved by the FDA found that many AI/ML datasets had inadequate representation of minoritized populations or reporting of representativeness was not provided (Wu et al. 2021).
Social inequality can be embedded in data in ways that vary according to the type of data and require attention to sources of bias that are context-dependent and at the systems level. For example, data taken from Electronic Health Records (EHRs) may reflect social inequalities in how minoritized populations are perceived and treated by health care providers– such as Black patients being undertreated for pain or under-diagnosed for certain health conditions. EHRs are also more likely to be missing data for low socioeconomic status patients, an issue which can lead to further unfair outcomes of resulting tools (Gianfrancesco et al. 2018). The classification of groups according to social or governmentally-defined categories, such as “Latino” or “Asian” can also serve to obscure underlying factors affecting health disparities in the data (Kauh, Read, and Scheitler 2021) A recent study of racial bias in an algorithm that provides personalized prognostication of outcomes for patients receiving a left ventricular-assist device (LVAD) noted that there were a number of systemic issues, that arise upstream from data collection, that impacted if and how data from Black patients appeared in the datasets used to develop the algorithm, such as whether Black patients had access to local cardiologists or hospital or the influence of socioeconomic factors on a person’s candidacy to receive an LVAD (Kostick-Quenet et al. 2022).
Traditional clinical research methods can be used to mitigate some forms of technical bias, but are not formulated to address forms of bias that arise from societal and institutional inequities. An approach that seeks to use such methods to evaluate the use of a specific tool in a specific context is insufficient for mitigating the multitude of ways that social inequities at the systems- and structural levels (Kostick-Quenet et al. 2022; Miceli, Posada, and Yang 2021). Mitigating bias, in the ethical sense, is a necessary part of bridging the AI chasm in order to reduce health disparities and produce medical AI/ML that provide social benefits for diverse populations. Doing so requires acknowledgement that data are shaped by historical and social inequities and commitment to a broader realignment of the norms and values of AI/ML and clinical research. Developers of healthcare AI/ML will need to take on the responsibility of a sustained commitment to identifying the social and structural factors that lead to inequities in data and AI/ML development, as well as to the ethical norms and values that support fairness and justice in AI/ML applications in healthcare.
FUNDING
MKC was funded by NIH grant 1R01HG010476) and the Greenwall grant, NM was funded by NIMH grant K01MH118375-01A1.
Footnotes
DISCLOSURE STATEMENT
No potential conflict of interest was reported by the authors.
REFERENCES
- Cho MK 2021. Rising to the challenge of bias in health care AI. Nature Medicine 27 (12): 2079–81. doi: 10.1038/s41591-021-01577-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gianfrancesco MA, Tamang S, Yazdany J, and Schmajuk G. 2018. Potential biases in machine learning algorithms using electronic health record data. JAMA Internal Medicine 178 (11): 1544–7. doi: 10.1001/jamainternmed.2018.3763. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jadad A, and Enkin M. 2007. Bias in randomized controlled trials. In Randomized controlled trials: Questions, answers and musings, 29–47. New York: John Wiley & Sons, Ltd. doi: 10.1002/9780470691922.ch3. [DOI] [Google Scholar]
- Kauh TJ, Read JG, and Scheitler AJ. 2021. The critical role of racial/ethnic data disaggregation for health equity. Population Research and Policy Review 40 (1):1–7. doi: 10.1007/s11113-020-09631-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kostick-Quenet KM, Cohen IG, Gerke S, Lo B, Antaki J, Movahedi F, Njah H, Schoen L, Estep JE, and Blumenthal-Barby JS. 2022. Mitigating racial bias in machine learning. The Journal of Law, Medicine & Ethics 50 (1):92–100. doi: 10.1017/jme.2022.13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lambert J 2011. Statistics in brief: How to assess bias in clinical studies? Clinical Orthopaedics and Related Research 469 (6):1794–6. doi: 10.1007/s11999-010-1538-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McCradden MD, Anderson JA, Stephenson EA, Drysdale E, Erdman L, Goldenberg A, and Zlotnik Shaul R. 2022. A research ethics framework for the clinical translation of healthcare machine learning. The American Journal of Bioethics 22(5): 8–12. doi: 10.1080/15265161.2021.2013977. [DOI] [PubMed] [Google Scholar]
- Miceli M,. Posada J, and Yang T. 2021. “Studying up machine learning data: Why talk about bias when we mean power?” ArXiv:2109.08131 [Cs], September. http://arxiv.org/abs/2109.08131.
- Obermeyer Z, Powers B, Vogeli C, and Mullainathan S. 2019. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366 (6464):447–53. doi: 10.1126/science.aax2342. [DOI] [PubMed] [Google Scholar]
- Sica GT 2006. Bias in research studies. Radiology 238 (3): 780–9. doi: 10.1148/radiol.2383041109. [DOI] [PubMed] [Google Scholar]
- Wu E, Wu K, Daneshjou R, Ouyang D, Ho DE, and Zou J. 2021. How medical AI devices are evaluated: Limitations and recommendations from an analysis of FDA approvals. Nature Medicine 27 (4):582–4. doi: 10.1038/s41591-021-01312-x. [DOI] [PubMed] [Google Scholar]
