Abstract.
Purpose
Breast ultrasound suffers from low positive predictive value and specificity. Artificial intelligence (AI) proposes to improve accuracy, reduce false negatives, reduce inter- and intra-observer variability and decrease the rate of benign biopsies. Perpetuating racial/ethnic disparities in healthcare and patient outcome is a potential risk when incorporating AI-based models into clinical practice; therefore, it is necessary to validate its non-bias before clinical use.
Approach
Our retrospective review assesses whether our AI decision support (DS) system demonstrates racial/ethnic bias by evaluating its performance on 1810 biopsy proven cases from nine breast imaging facilities within our health system from January 1, 2018 to October 28, 2021. Patient age, gender, race/ethnicity, AI DS output, and pathology results were obtained.
Results
Significant differences in breast pathology incidence were seen across different racial and ethnic groups. Stratified analysis showed that the difference in output by our AI DS system was due to underlying differences in pathology incidence for our specific cohort and did not demonstrate statistically significant bias in output among race/ethnic groups, suggesting similar effectiveness of our AI DS system among different races ( for all).
Conclusions
Our study shows promise that an AI DS system may serve as a valuable second opinion in the detection of breast cancer on diagnostic ultrasound without significant racial or ethnic bias. AI tools are not meant to replace the radiologist, but rather to aid in screening and diagnosis without perpetuating racial/ethnic disparities.
Keywords: artificial intelligence, deep learning, breast ultrasound, breast cancer, racial and ethnicity bias
1. Introduction
Breast ultrasound (US) detects small, invasive, node-negative cancers. It is used both as a diagnostic tool in those with signs or symptoms of breast pathology and as a supplemental tool in screening mammography for dense breasts.1 Assessment of a finding includes shape, orientation, margin, echogenicity, and posterior features as well as pre-test probability and individual risk profile. In the United States, the primary goal for breast cancer screening is to maximize cancer detection and decrease mortality. U.S. Preventive Services Task Force found that early detection of breast cancer has reduced the number of deaths from breast cancer for women 40 years and older.2 Breast ultrasound, particularly in women with dense breasts, detects mammographically occult malignancies and high-risk lesions.3 One prospective American College of Radiology Imaging Network protocol 6666 study showed that the addition of ultrasound to screening mammography increased cancer detection by 2.7% to 4.6% per 1000 women screened.1,4 However, breast ultrasound alone has low positive predictive value and low specificity with a high recall rate.4,5 False positives are more common on ultrasound, especially for younger women, who are more likely to have dense breasts as well as rapidly growing cancers that present clinically between screening exams.4 The harm of false positives creates a barrier to implementing ultrasound as a screening tool. However, a potential delay in diagnosis of breast cancer has significant consequences for treatment, management, and patient prognosis.6,7 Therefore, breast radiologists strive towards both early cancer detection and high positive predictive value.
The cost to Medicare of breast cancer screening exceeds $1 billion annually.8 Regional variation of screening-related cost is significant and driven by the use of new and often more expensive technology.8 For women 75 years and older, the annual screening-related expenditure exceeds $410 million. In regions with higher screening costs, women were more likely to be diagnosed with early-stage cancer.8 Therefore, adoption of new technology is important to validate the high cost for integration.
Artificial intelligence (AI) has been shown to reduce false negative mammography interpretations and increase detection of early-stage malignancies with potential simultaneous reduction in benign breast biopsies.9,10 An AI-based decision support (DS) system for breast US interpretation can distinguish a clinically significant threshold when it reports a lesion as suspicious versus probably benign, as it determines biopsy versus imaging follow-up for a patient. A breast imaging reporting and database system (BI-RADS) 3 lesion is proven to have a risk of malignancy greater than typically benign findings, but .11 The malignancy rate increases as the patient’s age increases, and thus BI-RADS 3 lesions are a clinically important threshold.
Perpetuating disparities is a potential risk when incorporating AI based models into clinical practice. AI-based systems can affect decision-making by optimizing clinical workflows, enhancing diagnoses, and supporting clinical management, but there is equal concern of these algorithms creating or amplifying certain biases.12,13 Although bias within the health system has been well established, AI-driven underdiagnosis or overdiagnosis based on patient demographics has yet to be explored.13 Underdiagnosis, specifically, is important for radiologists, as missing a diagnosis can change the course of clinical management, especially in the setting of breast cancer.
Koios DS for breast ultrasound (Koios Medical, Chicago, Illinois, United States) is an FDA-approved AI DS system that incorporates machine learning to generate a probability of malignancy for user-selected region of interest (ROI) based on two orthogonal views of a breast lesion and provides a BI-RADS aligned assessment. Our retrospective review assesses whether our AI DS system demonstrates bias based on race or ethnicity.
2. Materials and Methods
This study was designed as a quality study to determine if an FDA-approved software purchased by our institution had any racial or ethnic bias. This study was approved by the Institutional Review Board and was HIPAA-compliant. Using our Mammography Information System (MIS) (Ikonopedia, Richardson, Texas, United States), 4020 ultrasound-guided biopsy cases were identified from January 1, 2018 to October 28, 2021 from the nine breast imaging facilities within our health system that perform breast ultrasound. Pathology results were reviewed and classified as either benign or malignant. We compared the assessment of our AI DS system with the biopsy-proven pathology across each race/ethnicity. Non-primary breast malignancies, axillary lymph node biopsies, cyst fine needle aspirations, as well as patients with multiple biopsies on the same day were excluded, resulting in a total of 1810 biopsy proven cases.
We utilized images from the day of biopsy or diagnostic ultrasound prior to biopsy. We chose an ROI over one radial and anti-radial or transverse and sagittal view and analyzed each biopsy proven breast lesion with our AI DS system (Fig. 1). Patient demographics, including age, self-reported race/ethnicity, and gender were obtained from MIS (Ikonopedia, Richardson, Texas, United States).
Fig. 1.
Koios DS for breast ultrasound evaluating two orthogonal views of a breast lesion, producing a BI-RADS aligned assessment. Pathology: (benign, BI-RADS 2) hyalinized fibroadenoma; (probably benign, BI-RADS 3) fibroadenoma; (suspicious, BI-RADS 4A-4B) invasive poorly differentiated duct carcinoma; and (probably malignant, BI-RADS 4C+) invasive ductal carcinoma moderately differentiated.
Koios DS utilizes a collection of algorithms from over 2 million ultrasound images to aid in characterizing lesions suspicious for cancer. It is an FDA-approved computer-assisted diagnostic (CADx) software for lesions seen on ultrasound. Output from this AI DS system included the following categories: benign, probably benign, suspicious, and highly suspicious. Benign pathology encompassed AI DS categories of benign and probably benign and are considered equivalent to BI-RADS 2 and 3, respectively (Fig. 1). Similarly, malignant pathology encompassed AI DS categories of suspicious and highly suspicious and are considered equivalent to BI-RADS 4A/B and 4C+, respectively (Fig. 1). Cases were stratified by race/ethnicity, AI DS output and pathology. Chi-square test and Kruskal–Wallis tests were used for statistical analysis. Statistical significance was set at 0.05 or less.
3. Results
Self-reported race distribution showed 44.1% white, 19.5% Black/African American, 10.4% Asian, and 26% other (Table 1). Self-reported ethnicity distribution showed 33.3% Hispanic, 66.7% non-Hispanic (Table 1). Significant difference in distribution of race and ethnicity were observed in relation to machine type and facility () (Table 1). Our patient demographics showed a mean age of 51.4 years old with a range of 11 to 103 years old (Table 1). There was a significant difference in distribution of age among patients of different race/ethnicity () (Table 1). Average age for white patients was 53.7 years old, for Black patients 50.8 years old, for Asian patients 47.7 years old, and other patients 49.3 years old. Average age for Hispanic patients was 48.7 years old and non-Hispanic patients 52.7 years old. 98.6% of our patient population was female and 1.4% male (Table 1). There was no significant difference in distribution of gender amongst racial/ethnic groups (Table 1). There was a significant difference in distribution of different racial and ethnic groups by machine type and facility (Table 1). For example, most of our cases were from the Dubin Breast Center (46.7%), which demonstrated a distribution of 40.6% white patients, 55.2% Black patients, 38.1% Asian patients and 54% other. At Dubin, 52.7% of patients were Hispanic and 43.7% were non-Hispanic. We had four different machine types used across all nine facilities, GE (19.2%), Phillips (48%), Toshiba (1.3%), and unknown (31.5%) (Table 1).
Table 1.
Demographics by race/ethnicity. This composite table illustrates the distribution of patients according to their respective race and ethnicity in conjunction with various patient, ultrasound machine, and facility attributes. Notably, significant disparities in patient race/ethnicity distribution have been observed in relation to age, machine type, and facility ().
| All | White | Black/African American | Asian | Other | p | Hispanic | Non-Hispanic | p | ||
|---|---|---|---|---|---|---|---|---|---|---|
| N (%) | 1810 (100.0%) | 798 (44.1%) | 353 (19.5%) | 189 (10.4%) | 470 (26.0%) | 602 (33.3%) | 1208 (66.7%) | |||
| Age (mean (STD)) | 51.4 (15.8) | 53.7 (15.7) | 50.8 (16.2) | 47.7 (12.7) | 49.3 (16.3) | <0.001 | 48.7 (16.3) | 52.7 (15.4) | <0.001 | |
| Gender | Female | 1785 (98.6) | 782 (98.0) | 350 (99.2) | 188 (99.5) | 465 (98.9) | 0.224 | 593 (98.5) | 1192 (98.7) | 0.886 |
| Male | 25 (1.4) | 16 (2.0) | 3 (0.8) | 1 (0.5) | 5 (1.1) | 9 (1.5) | 16(1.3) | |||
| Machine | GE | 348 (19.2) | 176 (22.1) | 69 (19.5) | 43 (22.8) | 60 (12.8) | <0.001 | 93 (15.4) | 255 (21.1) | <0.001 |
| Philips | 868 (48.0) | 340 (42.6) | 197 (55.8) | 75 (39.7) | 256 (54.5) | 319 (53.0) | 549 (45.4) | |||
| Toshiba | 24 (1.3) | 19(2.4) | 4 (1.1) | 0 (0.0) | 1 (0.2) | 1 (0.2) | 23 (1.9) | |||
| Unknown | 570 (31.5) | 263 (33.0) | 83 (23.5) | 71 (37.6) | 153 (32.6) | 189 (31.4) | 381 (31.5) | |||
| Facility | BC | 191 (10.6) | 88 (11.0) | 34 (9.6) | 26 (13.8) | 43 (9.1) | <0.001 | 53 (8.8) | 138 (11.4) | <0.001 |
| BH | 42 (2.3) | 17 (2.1) | 18 (5.1) | 2 (1.1) | 5 (1.1) | 9 (1.5) | 33 (2.7) | |||
| BR | 250 (13.8) | 134 (16.8) | 43 (12.2) | 33 (17.5) | 40 (8.5) | 64 (10.6) | 186 (15.4) | |||
| DU | 845 (46.7) | 324 (40.6) | 195 (55.2) | 72 (38.1) | 254 (54.0) | 317 (52.7) | 528 (43.7) | |||
| HU | 111 (6.1) | 107 (13.4) | 1 (0.3) | 2 (1.1) | 1 (0.2) | 4 (0.7) | 107 (8.9) | |||
| KH | 24 (1.3) | 19(2.4) | 4 (1.1) | 0 (0.0) | 1 (0.2) | 1 (0.2) | 23 (1.9) | |||
| PH | 235 (13.0) | 79 (9.9) | 45(12.7) | 39 (20.6) | 72 (15.3) | 93 (15.4) | 142 (11.8) | |||
| QN | 28 (1.5) | 12(1.5) | 3 (0.8) | 6 (3.2) | 7 (1.5) | 11 (1.8) | 17(1.4) | |||
| SL | 84 (4.6) | 18 (2.3) | 10 (2.8) | 9 (4.8) | 47 (10.0) | 50 (8.3) | 34 (2.8) |
Out of 1810 cases, 78.1% of our biopsy proven pathology was benign and 21.9% was malignant (Table 2). There was statistically significant difference in distribution of pathology incidence across different racial groups () and ethnicities () (Table 2). The output of our AI DS system among different racial and ethnic patient groups showed statistically significant difference in distribution (, , respectively) (Table 2). For example, white patients have relative increased output from AI DS of suspicious and highly suspicious (49.6% and 20.7%, respectively), which correlates to having a higher percentage of malignant pathology compared to other racial groups in our study (25.6% malignant pathology) (Table 2). Relative to other groups, self-reported other race patients had overall increased output of benign from AI DS system (22.8%), which correlated to having a higher percentage of benign pathology (81.7%) (Table 2). Hispanic patients showed relative increased output of benign (22.8%) compared to non-Hispanic patients, which correlated to more benign pathology proven cases (81.7%). Additionally, non-Hispanic patients had more output of suspicious and highly suspicious (48.8% and 19.9%), which correlated to more malignant pathology cases (23.7%) compared to Hispanic patients (18.3%) (Table 2).
Table 2.
AI diagnostic system (DS) and pathology results by race/ethnicity. This composite table presents patient data categorized by race and ethnicity, juxtaposed with the outcomes generated by the AI DS system and actual pathology results. Notably, significant disparities in pathology incidence have been observed across different racial groups () and ethnicities (). Furthermore, distinctions have emerged in AI DS output based on race () and ethnicity (). As a result, a stratified analysis into benign and malignant pathology, as detailed in Table 3, is imperative to ascertain whether these statistical differences in AI DS output are attributable to racial disparities in incidence of pathology or other variables.
| All | White | Black/African American | Asian | Other | p | Hispanic | Non-Hispanic | p | ||
|---|---|---|---|---|---|---|---|---|---|---|
| AIDS | Benign | 347 (19.2) | 130 (16.3) | 78(22.1) | 32 (16.9) | 107 (22.8) | 0.012 | 137 (22.8) | 210 (17.4) | 0.006 |
| Probably benign | 250 (13.8) | 107 (13.4) | 45 (12.7) | 32 (16.9) | 66 (14.0) | 82(13.6) | 168 (13.9) | |||
| Suspicious | 875 (48.3) | 396 (49.6) | 164 (46.5) | 91 (48.1) | 224 (47.7) | 285 (47.3) | 590 (48.8) | |||
| Highly suspicious | 338 (18.7) | 165 (20.7) | 66 (18.7) | 34 (18.0) | 73 (15.5) | 98 (16.3) | 240 (19.9) | |||
| Pathology | 1414 (78.1) | 594 (74.4) | 280 (79.3) | 156 (82.5) | 384 (81.7) | 0.006 | 492 (81.7) | 922 (76.3) | 0.032 | |
| Malignant | 396 (21.9) | 204 (25.6) | 73 (20.7) | 33 (17.5) | 86 (18.3) | 110 (18.3) | 286 (23.7) |
Differences in output among racial and ethnic groups did not persist when stratified by malignant versus benign pathology (Table 3). There was no significant difference in distribution of output among different racial groups for benign and malignant biopsy proven cases (, , respectively) and for different ethnic groups for benign and malignant biopsy proven cases (, , respectively). AI DS system incorrectly diagnosed 4 pathology proven malignant cases as benign and 15 pathology proven malignant cases as probably benign (Table 3).
Table 3.
AI DS by race/ethnicity stratified by benign and malignant pathology. The previously demonstrated differences in Table 2 have not persisted, suggesting that the differences in AI DS output were due to underlying differences in pathology incidences. Furthermore, these disparities are eliminated within in the benign and malignant groups, suggesting similar effectiveness of the AI DS system among different races ( for all).
| Pathology | AIDS | All | White | Black/African American | Asian | Other | p | Hispanic | Non-Hispanic | p |
|---|---|---|---|---|---|---|---|---|---|---|
| Begin | Benign | 343 (24.3) | 129 (21.7) | 75 (26.8) | 32 (20.5) | 107 (27.9) | 0.18 | 135 (27.4) | 208 (22.6) | 0.051 |
| Probably benign | 235 (16.6) | 99 (16.7) | 41 (14.6) | 31 (19.9) | 64 (16.7) | 80 (16.3) | 155 (16.8) | |||
| Suspicious | 722 (51.1) | 321 (54.0) | 135 (48.2) | 78 (50.0) | 188 (49.0) | 242 (49.2) | 480 (52.1) | |||
| Highly suspicious | 114 (8.1) | 45 (7.6) | 29 (10.4) | 15 (9.6) | 25 (6.5) | 35 (7.1) | 79 (8.6) | |||
| Malignant | Benign | 4 (1.0) | 1 (0.5) | 3 (4.1) | 0 (0.0) | 0 (0.0) | 0.528 | 2 (1.8) | 2 (0.7) | 0.783 |
| Probably benign | 15 (3.8) | 8 (3.9) | 4 (5.5) | 1 (3.0) | 2 (2.3) | 2 (1.8) | 13 (4.5) | |||
| Suspicious | 153 (38.6) | 75 (36.8) | 29 (39.7) | 13 (39.4) | 36 (41.9) | 43 (39.1) | 110 (38.5) | |||
| Highly suspicious | 224 (56.6) | 120 (58.8) | 37 (50.7) | 19 (57.6) | 48 (55.8) | 63 (57.3) | 161(56.3) |
4. Discussion
Some of the most important risk factors for breast cancer are age over 40 years old, history of cancer in first degree relatives, early menarche and late childbearing, and Caucasian race.14,15 80% of breast cancers are diagnosed in women 50 years or older with 50% occurring in women between 50 and 69 years old.14 Our results reflect some of these known risk factors. In our specific cohort, we find that white patients are more likely to have breast cancer, demonstrating a higher percentage of malignant pathology relative to other racial groups (Table 2). This may be related to a relatively older average age of white patients in our cohort, average 53.7 years old (Table 1). The second oldest age group was Black patients, with average age 50.8 years old (Table 1) and they also displayed the second highest malignant pathology results (20.7%) (Table 2).
Our race/ethnicity distribution differed from the national rates of race/ethnicity distribution, which may reflect the unique diversity across Manhattan, Queens, Brooklyn, and Long Island. Our distribution showed 44.1% white, 19.5% Black/African American, 10.4% Asian, and 26% other (Table 1). Self-reported ethnicity distribution showed 33.3% Hispanic and 66.7% non-Hispanic (Table 1). For comparison, Breast Cancer Surveillance Consortium (1994 to 2009) national rates of racial/ethnicity distribution in breast cancer surveillance show 73.4% white, Black/African American 6.91%, Asian or pacific islander 6.44%, American Indian or Alaska native 1.71%, mixed 1.29%, and Hispanic 10.1%.16 One of our limitations was that we did not have any self-reported Native American or Alaska native population.
Research has shown that there is more genetic variation within a race than between races, and that race is more of a social construct than a biological one.17 When evaluating disparities and discrimination related to race and ethnicity, we are looking at how social and cultural construct rather than genetic ancestry affects patient outcomes.17 By dividing analysis into benign and malignant pathology proven cases (Table 3), we removed possible confounding factors, such as statistically significant differences in distribution of pathology or age. Our AI DS system did not show significant difference in distribution of output among racial and ethnic groups when stratified by benign and malignant cases, suggesting a similar effectiveness among different racial and ethnic groups despite underlying differences in pathology incidence (Tables 2 and 3). In other words, although our results showed a difference in distribution of benign and malignant pathology per race and ethnicity, this was not due to AI DS system bias, but rather an accurate reflection of the pathologic characteristics of our patient population, possibly due to other confounding factors. A limitation in our study is that we have taken biopsy proven cases and retrospectively tested our AI DS system. Therefore, we do not know cases that may have been missed by AI DS system or deemed to not require a biopsy by the radiologist in real practice and therefore, we do not know true false negatives. AI DS may also provide false positive reports as well due to this limiting factor. Therefore, we cannot accurately assess accuracy, sensitivity, and specificity. Prospective studies may be able to truly assess the accuracy of an AI DS system and whether this affects its bias among racial and ethnic groups.
Biases are defined as differences in performance against or in favor of a subpopulation for a predictive task, or when an algorithm systematically favors one outcome over another.12,13 Biases can be introduced throughout any of the steps in the development process of AI, such as data collection, data selection, model training, and deployment.12 For example, Seyyed-Kalantari et al showed that AI models can demonstrate different accuracy in chest X-ray diagnosis across racial and demographic groups based only on the provided image.13 In this particular study, compared to white and male patients, patients who are Black and female were more often incorrectly identified as healthy (false negative). AI DS systems are based on a prediction of outcomes from shown examples, therefore, when training data reflects real-world disparities, unintended bias can be a potential consequence during machine learning.12 We must be cautious that although our AI DS system performed without bias in our data set, there is a possibility that a model can appear unbiased in one setting but may display bias in another, as AI-based algorithms tend to overfit the data on which they are trained.12 We do not have information on the patient demographics or data set that our AI DS system was trained on. In addition, we only tested one AI DS system, which does not generalize to all AI-based models.
AI DS system incorrectly diagnosed 4 pathology proven malignant cases as benign and 15 pathology proven malignant cases as probably benign (Table 3). Probably benign (BI-RADS 3) cases have a risk of malignancy greater than typical benign findings, but ; therefore, this result would be considered as discordant with a malignant pathology proven case. Although the benign cases were missed in white and Black patients, these disparities were not statistically significant. Upon review of missed malignant cases by our AI DS system, many of the lesions demonstrated typically benign features on ultrasound, such as well-circumscribed borders, oval shape, parallel, and anechoic with posterior acoustic shadowing. A comparison between breast radiologists’ assessment and AI DS output and the difference in accuracy may help determine if AI DS system may be missing nuances within a case, such as patient age, patient positioning, or artifacts. Given that we had different users of different levels of experience (first-year radiology residents to breast radiology attendings with years of experience) use AI DS system interface on images in a blinded fashion, we cannot account for differences in experience that may translate to false negatives. For example, if a radiologist decides to only highlight a cystic portion of the mass and miss the solid component, this may result in a false negative. Ultrasound machine type, technique and transducer frequency may also affect image quality. Subtle architectural distortions can be missed when an AI DS system is assessing a lesion based on still images taken by different technologists with different levels of experience. Future studies can evaluate whether these factors introduce variability in AI DS output accuracy.
We want to emphasize that an AI DS system was not designed to be a stand-alone interpretation, but a tool that can be used by the radiologist. As opposed to CADx for mammography or magnetic resonance imaging (MRI), AI DS for breast ultrasound is operator dependent in multiple ways. Ultrasound, innately, is an examination that depends on real-time scanning to distinguish artifact shadowing from true posterior shadowing.1 Its use is predicated on the operator detecting a lesion, on the image quality, the radiologist’s decision to interrogate the lesion with the AI DS system, as well as the ROI drawn by the AI DS or reader. We did not assess radiologists and the use of this system concurrently in real time to assess their agreement.
Our study included nine breast imaging facilities, which include academic faculty practice and hospital-based imaging breast imaging centers dispersed within Manhattan, Queens, Brooklyn, and Long Island. Ultrasound, stereotactic, and MRI guided biopsies are offered at our institutions. Method of biopsy is dependent on which modality the lesion has been seen or assessed. For example, lesions seen on ultrasound would undergo an ultrasound guided biopsy versus a lesion only seen on mammography would undergo stereotactic biopsy. Although the locations of our facilities differ, they are served by one faculty practice group and exams can be cross-read throughout the system, thus there is relatively good standardization in our study. However, significant disparities in patient race and ethnicity distribution have been observed in relation to age, machine type, and facility with clustering seen in specific machine types and more patient volume in certain facilities, such as Dubin Breast Center (DU) (Table 1). Some facilities may utilize one machine type more than another. The significant difference in distribution of racial and ethnic groups may reflect the inherent differences in distribution per neighborhood within New York, United States. Studies have shown that a variety of performance across racial and ethnic groups can be related to the imaging facility and access rather than personal characteristics.7,18 Mammographic accuracy, for example, has been shown to vary greatly based on institution and radiologist, with poorer performance at facilities that treat predominantly minority and low-income women.18 Facilities that serve mostly racial/ethnic minority groups may have a history of chronic underinvestment, as NYC neighborhoods differ drastically in insurance coverage rates, income level, breast cancer rates.19 Structural racism in and outside of healthcare centers negatively impact racial and ethnic minorities and are associated with worse healthcare access and outcomes.19 It would be valuable to look at additional factors such as healthcare literacy, insurance coverage, reimbursement policies per patient subpopulation utilizing specific sites.7 We also did not have access to socioeconomic status of our patients, which would be a valuable additional factor for future analysis. Recognizing these disparities can help to allocate more resources to decrease disproportionate negative impact in the diagnostic work up for breast cancer on racial and ethnic minorities and underserved populations, such as patient navigators, translators, and same-day biopsy services.7 It would be interesting to evaluate and compare the accuracy of an AI DS system and breast radiologists and whether they are impacted by facility, neighborhood, or machine type.
Non-primary breast malignancies, axillary lymph node biopsies, cyst fine needle aspirations, as well as patients with multiple biopsies on the same day were excluded. Exclusion of these observations can lead to underestimated or overestimated effects. Another limitation is that we looked only at biopsied cases, such that true screening ultrasound cases with BI-RADS 2 or 3 could have changed distribution of our AI DS system’s output. Lack of prior images is associated with higher false positive rates.3,18 This issue may be more common among populations with less access to care. Availability of prior comparison studies have been shown to reduce false positive recalls for all breast imaging modalities, decreasing from 20.9% in first screening ultrasound to 10.7% in the second and third year follow-up.4 Future analysis comparing AI DS accuracy between first and second/third look ultrasounds may be helpful to assess if AI DS can reduce false positive recalls.
5. Conclusion
Our results show promise that an AI DS system and newer technologies may serve as a valuable second opinion in the detection of breast cancer on diagnostic ultrasound without significant racial or ethnic bias.
Biographies
Clara Koo, MD, is a diagnostic radiology resident at Mount Sinai Hospital in New York. She earned her MD degree from Icahn School of Medicine at Mount Sinai.
Biographies of the other authors are not available.
Contributor Information
Clara Koo, Email: clara.koo2@mountsinai.org.
Anthony Yang, Email: anthony.yang@mountsinai.org.
Colton Welch, Email: colton.welch@mountsinai.org.
Vipashyana Jadav, Email: Vipashyana.Jadav@mountsinai.org.
Liana Posch, Email: liana.posch@mountsinai.org.
Nicholas Thoreson, Email: Nicholas.Thoreson@mountsinai.org.
Darrell Morris, Email: darrell.morris@mountsinai.org.
Fatima Chouhdry, Email: Fatima.Chouhdry@mountsinai.org.
Janet Szabo, Email: janet.szabo@mountsinai.org.
David Mendelson, Email: david.mendelson@mountsinai.org.
Laurie R. Margolies, Email: laurie.margolies@mountsinai.org.
Disclosures
A donation by the Glasgold Family Foundation supported the acquisition of the Koios DS for breast ultrasound but has not supported this investigation or any of the investigators. This study is not sponsored by Koios and is not a shared project. Results will be published and therefore available for review by anyone, including the company, but there is no direct involvement.
Code and Data Availability
Data cannot be publicly made available as it was pulled from our database of patients throughout the hospital system.
References
- 1.Berg W. A., et al. , “Training the ACRIN 6666 Investigators and effects of feedback on breast ultrasound interpretive performance and agreement in BI-RADS ultrasound feature analysis,” AJR Am. J. Roentgenol. 199(1), 224–235 (2012). 10.2214/AJR.11.7324 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Siu A. L. and U.S. Preventive Services Task Force, “Screening for Breast Cancer: U.S. preventive services task force recommendation statement,” Ann. Intern. Med. 164(4), 279–296 (2016). 10.7326/M15-2886 [DOI] [PubMed] [Google Scholar]
- 3.Weigert J., Steenbergen S., “The Connecticut experiments second year: ultrasound in the screening of women with dense breasts,” Breast J. 21(2), 175–180 (2015). 10.1111/tbj.12386 [DOI] [PubMed] [Google Scholar]
- 4.Berg W. A., et al. , “Ultrasound as the primary screening test for breast cancer: analysis from ACRIN 6666,” J. Natl. Cancer Inst. 108(4), djv367 (2015). 10.1093/jnci/djv367 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Bahl M., “Artificial intelligence for breast ultrasound: will it impact radiologists’ accuracy?” J. Breast Imaging 3(3), 312–314 (2021). 10.1093/jbi/wbab022 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Jacobson J. S., et al. , “Breast biopsy and race/ethnicity among women without breast cancer,” Cancer Detect. Prev. 30(2), 129–133 (2006). 10.1016/j.cdp.2006.02.002 [DOI] [PubMed] [Google Scholar]
- 7.Lawson M. B., et al. , “Multilevel factors associated with time to biopsy after abnormal screening mammography results by race and ethnicity,” JAMA Oncol. 8(8), 1115–1126 (2022). 10.1001/jamaoncol.2022.1990 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Gross C. P., et al. , “The cost of breast cancer screening in the Medicare population,” JAMA Intern. Med. 173(3), 220–226 (2013). 10.1001/jamainternmed.2013.1397 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Lehman C. D., et al. , “Diagnostic accuracy of digital screening mammography with and without computer-aided detection,” JAMA Intern. Med. 175(11), 1828–1837 (2015). 10.1001/jamainternmed.2015.5231 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Mango V. L., et al. , “Should we ignore, follow, or biopsy? Impact of artificial intelligence decision support on breast ultrasound lesion assessment,” Am. J. Roentgenol. 214(6), 1445–1452 (2020). 10.2214/AJR.19.21872 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Berg W. A., “BI-RADS 3 on screening breast ultrasound: what is it and what is the appropriate management?” J. Breast Imaging 3(5), 527–538 (2021). 10.1093/jbi/wbab060 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Huang J., et al. , “Evaluation and mitigation of racial bias in clinical machine learning models: scoping review,” JMIR Med. Inf. 10(5), e36388 (2022). 10.2196/36388 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Seyyed-Kalantari L., et al. , “Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations,” Nat. Med. 27, 2176–2182 (2021). 10.1038/s41591-021-01595-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Kamińska M., et al. , “Breast cancer risk factors,” Prz. Menopauzalny 14(3), 196–202 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Jatoi I., Sung H., Jemal A., “The emergence of the racial disparity in U.S. breast-cancer mortality,” New Engl. J. Med. 386(25), 2349–2352 (2022). 10.1056/NEJMp2200244 [DOI] [PubMed] [Google Scholar]
- 16.“Breast Cancer Surveillance Consortium website,” Dataset (HHSN261201100031C). tools.bcsc-scc.org/dataexplorer/.
- 17.Gichoya J. W., et al. , “AI recognition of patient race in medical imaging: a modelling study,” Lancet Digit. Health 4(6), e406–e414 (2022). 10.1016/S2589-7500(22)00063-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.McCarthy A. M., et al. , “Racial differences in false-positive mammogram rates: results from the ACRIN Digital Mammographic Imaging Screening Trial (DMIST),” Med. Care 53(8), 673–678 (2015). 10.1097/MLR.0000000000000393 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Kang J. X., Levanon Seligson A., Dragan K. L., “Identifying New York City neighborhoods at risk of being overlooked for interventions,” Prev. Chronic Dis. 17, E32 (2020). 10.5888/pcd17.190325 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Data cannot be publicly made available as it was pulled from our database of patients throughout the hospital system.

