While artificial intelligence (AI) holds great promise in improving various domains of clinical practice – from administrative tasks to diagnostic tools and clinical decision support, the potential for bias raises significant concerns related to patient harm. This article provides a brief overview of bias in artificial intelligence models intended for use in health care settings, its clinical impact, and offers a simple framework for mitigating bias in AI models.
Defining Bias in Artificial Intelligence
Bias in AI models is often described as algorithmic bias and can be broadly classified into 2 forms – (1) Inherent bias in underlying datasets1 and (2) Labeling bias i.e. use of a incorrect/error-prone training endpoint or model input.
Inherent bias
Given that disparities remain pervasive in healthcare, existing clinical research datasets and electronic health records (EHRs) unfortunately harbor biases related to societal, sex, gender, racial, and economic differences. For example, the underrepresentation of women in clinical trials and datasets curated from these studies contributes to the limited understanding of sex differences in the pathobiology of many cardiovascular diseases and therapeutic options. Also, because racial and ethnic minority individuals often have limited access to care, EHRs often provide an incomplete picture of overall health and outcomes in these groups. As such, training AI models with these datasets can inadvertently lead to inaccurate predictions of health indices and outcomes in specific patient groups and reinforcement of health disparities. Essentially, inherent bias in this context refers to biases present in underlying datasets either due to underrepresentation or misrepresentation.
Labeling bias
This may occur when models use proxy variables believed to represent other unmeasured variables. Using proxy variables in AI models can lead to a series of unpredictable domino effects. For example, a widely used commercial prediction algorithm by health systems demonstrated racial bias leading to Black patients being less likely to be referred for specialized care programs. An in-depth analysis of the algorithm identified that it was trained to predict healthcare costs, and this was being used as a proxy for illness. This resulted in racially biased predictions because at a given level of health (based on the number of comorbid conditions), Black patients generated lower health care costs compared with White patients (likely due to differential access to care) and with this algorithm, lower healthcare costs were assumed to be equivalent to being less sick. As such, health care costs did not accurately reflect health status in Black patients2. Importantly, algorithmic bias may also be missed if inadequate metrics of effectiveness are used such as – model validation in small or non-diverse cohorts, or provision of aggregate model performance without evaluating effects within subgroups. Hence, there is an urgent need for thorough validation and bias assessments prior to deployment of AI tools for use in clinical care.
Impact of Algorithmic Bias in Clinical Contexts
The impact of algorithmic bias in clinical care settings can range from minor/low risk to significant or life-threatening consequences. The potential harm that algorithmic bias poses to humans has been demonstrated with algorithms deployed in non-health care settings such as commercial facial recognition software – showing poorer performance among dark skinned females (gender shades study) and recidivism prediction tools used by law enforcement agencies – demonstrated to overestimate risk among African Americans. Recently, a deep learning model was developed to predict incident heart failure within 5 years using EHR data from patient encounters at a single institution. Data input was 12-lead electrocardiograms, and the training target (outcome of interest) was incident heart failure – determined using SNOMED clinical codes. The AI model was shown to perform poorly in young Black patients, particularly women. To address model bias, the following steps were taken: (1) retraining the model using equal samples of racial and ethnic groups, (2) training separate models for each racial group, and (3) incorporating demographic variables into the model. However, these strategies did not successfully mitigate the biased model predictions3. With this example, multiple attempts were made to address inherent bias (to minimize the impact of underrepresentation of Black patients) using recommended methods; however, labeling bias may have contributed to model’s disparate performance. The training target requires accurate and objective identification of incident heart failure. The use of clinical and or diagnosis codes for health outcome ascertainment, as was done in this study, is known to be error prone, and as such, it is at risk for labeling bias. In addition, the choice of a 5-year future risk prediction comes with the potential risk for labeling errors as the model assumes that EHRs accurately and completely account for longitudinal health outcomes, although clinical encounters are often episodic and fragmented (care at other facilities). Some have suggested the use of race-specific thresholds may potentially address algorithmic disparities. However, the use of broad race categories as a mitigation strategy in this context is problematic given significant heterogeneity within racial groups and concerns that have been highlighted more recently about the use of race in clinical risk score/prediction models4. Furthermore, this strategy may wrongly reinforce biological differences when in fact, race is a social construct. Alternative approaches to consider would be to incorporate social determinants of health (if the data are available) in the model or reconsider the use of AI for tasks such as these.
Strategies to Address Algorithmic Bias in AI Models for Clinical Use
To effectively address algorithmic bias, it is important to adopt strategies that target the different forms of bias. These strategies are broadly referred to as responsible AI. Pending the development of national and or institution specific guidelines/code of conduct for AI in healthcare, clinicians, researchers, and model developers intending to either develop or use AI models should consider the proposed framework5 below:
(1) Inclusivity – intentional inclusion of diverse datasets (for model training) and team members. In the fields of cardiology and cardiac electrophysiology, it is important to ensure adequate representation of women as well as racial and ethnic minority groups. Large, multi-site collaborations for development of AI algorithms, efforts to encourage safe data sharing, and public-private partnerships should be considered.
(2) Specificity - ensuring the intended task, training endpoint, or research question is sufficiently objective and specific, in addition to availability of accurately labeled data. If this is not possible, one might need to reconsider whether the use of an AI algorithm is appropriate.
(3) Transparency – providing standard reporting metrics regarding the training data, AI related study protocols and trials; model description, limitations, and interpretability as appropriate.
(4) Validation – conducting a thorough evaluation of model performance in various patient populations and diverse clinical settings. Ideally, this would include retrospective and prospective studies, clinical trials – to provide information on the impact of AI tools on clinical outcomes, and implementation studies – to address context-specific integration of AI tools into clinical practice.
It should be noted that despite all mitigation strategies highlighted above, AI models are unlikely to always provide perfect discrimination. As such, it will be important to balance trade-offs in the implementation of AI in clinical practice. To that end, evaluating the following is recommended: explainability vs. accuracy and or reproducibility, relative risks vs. potential impact, benefits vs. cost, and targeted vs. broad population-based interventions. For example, fully autonomous AI systems in healthcare, without human oversight, are likely to have the highest risk of harm when a wrong decision is made if the intervention directly impacts patient care. However, their use for administrative tasks in healthcare may not necessarily carry similar risks. Lastly, ethical and legal implications of AI use must be weighed in relation to who bears responsibility for clinical decisions highlighting the importance of AI regulation and governance.
Conclusion
In conclusion, the use of AI in clinical care will transform clinical practice as it is currently known. As much as clinicians anticipate a positive impact from it, they must be aware of the potential for negative or harmful consequences if these tools are not deployed responsibly or do not employ strategies that mitigate the risk of bias.
Funding Sources:
Dr. Adedinsewo is supported by the Mayo Building Interdisciplinary Research Careers in Women’s Health (BIRCWH) Program funded by the National Institutes of Health [grant number K12 HD065987]. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Abbreviations
- AI
Artificial intelligence
- EHR
Electronic health record
- SNOMED
Systematized Nomenclature of Medicine
Footnotes
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
Disclosures: The authors declare no conflicts of interest.
References
- 1.Obermeyer Z, Nissan R, Stern M, Eaneff S, Bembeneck EJ, Mullainathan S. Algorithmic bias playbook. Center for Applied AI at Chicago Booth 2021:7–8. [Google Scholar]
- 2.Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 2019;366(6464):447–453. (https://science.sciencemag.org/content/sci/366/6464/447.full.pdf). [DOI] [PubMed] [Google Scholar]
- 3.Kaur D, Hughes JW, Rogers AJ, et al. Race, Sex, and Age Disparities in the Performance of ECG Deep Learning Models Predicting Heart Failure. Circulation: Heart Failure 2024;17(1):e010879. DOI: doi: 10.1161/CIRCHEARTFAILURE.123.010879. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Basu A Use of race in clinical algorithms. Sci Adv 2023;9(21):eadd2704. (In eng). DOI: 10.1126/sciadv.add2704. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Perez-Downes JC, Tseng AS, McConn KA, et al. Mitigating Bias in Clinical Machine Learning Models. Current Treatment Options in Cardiovascular Medicine 2024:1 [Google Scholar]
