Skip to main content
Cell Reports Medicine logoLink to Cell Reports Medicine
. 2025 Dec 16;6(12):102527. doi: 10.1016/j.xcrm.2025.102527

Contrastive learning enhances fairness in pathology artificial intelligence systems

Shih-Yen Lin 1,23, Pei-Chen Tsai 1,2,23, Fang-Yi Su 1,2,23, Chun-Yen Chen 2, Fuchen Li 1, Junhan Zhao 1,3, Yuk Yeung Ho 1,4, Tsung-Lu Michael Lee 5, Elizabeth Healey 1,6, Po-Jen Lin 1,7, Ting-Wan Kao 1, Dmytro Vremenko 1, Thomas Roetzer-Pejrimovsky 8,9, Lynette Sholl 10,11, Deborah Dillon 10, Nancy U Lin 12, David Meredith 10, Keith L Ligon 10,11, Ying-Chun Lo 13, Nipon Chaisuriya 13,14, David J Cook 13, Adelheid Woehrer 8,15, Jeffrey Meyerhardt 11, Shuji Ogino 10,16,17,18, MacLean P Nasrallah 19, Jeffrey A Golden 20, Sabina Signoretti 10,11, Jung-Hsien Chiang 2,, Kun-Hsing Yu 1,10,21,22,24,∗∗
PMCID: PMC12765949  PMID: 41406939

Summary

AI-enhanced pathology evaluation systems hold significant potential to improve cancer diagnosis but frequently exhibit biases against underrepresented populations due to limited diversity in training data. Here, we present the Fairness-aware Artificial Intelligence Review for Pathology (FAIR-Path), a framework that leverages contrastive learning and weakly supervised machine learning to mitigate bias in AI-based pathology evaluation. In a pan-cancer AI fairness analysis spanning 20 cancer types, we identify significant performance disparities in 29.3% of diagnostic tasks across demographic groups defined by self-reported race, gender, and age. FAIR-Path effectively mitigates 88.5% of these disparities, with external validation showing a 91.1% reduction in performance gaps across 15 independent cohorts. We find that variations in somatic mutation prevalence among populations contribute to these performance disparities. FAIR-Path represents a promising step toward addressing fairness challenges in AI-powered pathology diagnoses and provides a robust framework for mitigating bias in medical AI applications.

Keywords: artificial intelligence, AI, deep learning, pathology, cancer diagnosis, bias mitigation, fairness, algorithmic bias, contrastive learning, weakly supervised learning

Graphical abstract

graphic file with name fx1.jpg

Highlights

  • Standard pathology AI systems show performance disparities across demographic groups

  • Our FAIR-Path framework mitigates 88.5% of these disparities

  • External validation demonstrates a 91.1% reduction in diagnostic performance gaps

  • Population-level somatic mutation variations contribute to AI model bias


Standard AI systems for cancer pathology interpretation exhibit differential performance across demographic groups. Lin et al. address this critical challenge by establishing the FAIR-Path framework, which substantially reduces diagnostic disparities and enhances model fairness. This work represents a crucial step toward ensuring effective AI-assisted cancer evaluation for all populations.

Introduction

Cancer is a leading cause of death in developed countries.1 Pathology evaluation plays an indispensable role in diagnosing cancers. Recent advances in computational pathology, driven by progress in artificial intelligence (AI), have shown great potential to enhance cancer pathology diagnoses. Several studies have demonstrated AI’s capabilities in detecting tumors,2,3,4 classifying cancer subtypes,5,6,7 predicting cancer-related genomic profiles,8,9,10 and estimating patient survival outcomes.2,3,5,8,11,12 These findings point to a promising future where data-driven algorithms could significantly enhance diagnostic accuracy and patient care. In addition, AI has been shown to alleviate clinicians’ workload by automating repetitive tasks, enabling them to focus more on patient-specific care.13,14 For example, AI applications in pathology aid in timely detection and treatment by predicting key prognostic factors, allowing for earlier interventions. Computational pathology shows the potential for intraoperative diagnostics and decisions using frozen sections, such as determining the extent of malignant tissue removal to prevent overtreatment.7,15

However, algorithmic bias remains a significant challenge in deploying AI models for pathology diagnosis.16 Large-scale pathology datasets, such as The Cancer Genome Atlas Program17 (TCGA) and the Clinical Proteomic Tumor Analysis Consortium18 (CPTAC), predominantly consist of Caucasian patients, with African Americans, Asians, and other minority racial groups accounting for only 17.4% of patients.19 This imbalance introduces biases into AI-driven pathology diagnosis models, as these systems often reflect the bias present in their training data.20 When coupled with racial disparities in cancer risk factors21,22 and social determinants of health,23 using biased datasets without caution could result in underdiagnosis and exacerbate health inequities in underrepresented groups. In addition, previous studies have shown that pathology images contain features that correlate with race24 and hospital site,25 which may lead to shortcut learning26 during model training. These issues can result in unequal diagnostic performance and perpetuate disparities in healthcare outcomes.27 Therefore, auditing AI models for bias has become a critical prerequisite of clinical deployment.28

Various bias mitigation methods have been proposed in the deep learning literature. For example, some studies address data imbalances between demographic groups through reweighting, resampling,29 data augmentation,30 or generating synthesized minority group data.31 Others have sought to learn fair representations by suppressing information related to demographic attributes or other confounding factors (e.g., tissue source site) using contrastive learning,32 feature disentanglement,33 domain generalization,34 generative models,31 network pruning,35 or adversarial learning.36,37 Additional studies attempted to validate the efficacy of bias mitigation methods in various disease diagnosis models.38

In cancer pathology, several studies have attempted to address fairness in AI-driven diagnostic models. For example, Ktena et al. demonstrated that generative models can reduce performance disparities between tissue source sites in pathology classification models while enhancing overall performance.31 Hosseini et al. proposed proportionally fair federated learning39 to improve fairness among hospitals. Vaidya et al. investigated how data preprocessing methods, model architectures, and bias mitigation methods affect racial disparities in subtyping breast and lung carcinomas and predicting IDH1 mutations in gliomas. They also revealed that state-of-the-art models for lung cancer subtyping exhibited significant biases related to attributes such as insurance type, income, and age.24 These studies provide valuable insights into understanding bias in computational pathology while underscoring the need for continued efforts in bias mitigation. Despite these efforts, fairness analysis and bias mitigation in computational pathology remain limited to small datasets and selected cancer types. A systematic analysis of demographic biases and bias mitigation methods in general cancer diagnosis tasks remains lacking.

In this study, we performed a systematic fairness analysis in computational pathology, evaluating critical cancer detection and classification tasks across 20 cancer types and eight datasets. Our analysis revealed that standard deep learning models8 exhibited biases and performance disparities against underrepresented demographic subgroups defined by race, gender, and age. To address these issues, we introduce Fairness-aware Artificial Intelligence Review for Pathology (FAIR-Path), a machine learning framework that mitigates bias through a fairness-aware contrastive learning process. Our pan-cancer evaluation, conducted across diverse cancer types and diagnostic tasks using multiple independent cohorts from seven medical centers or study populations, shows that FAIR-Path reduces up to 88.5% of the biases found in baseline diagnosis systems. Our approach lays the foundation for developing unbiased AI systems for cancer pathology, advancing equity in cancer diagnoses and management.

Results

Study cohort summary

We collected 28,732 whole-slide cancer pathology images (WSIs) from 14,456 cancer patients, encompassing 20 cancer types across five medical centers and three nationwide study cohorts. These digital pathology images include 5,834 formalin-fixed paraffin-embedded (FFPE) slides and 15,330 frozen section slides from the TCGA dataset17 (from 8,694 patients), 1,248 frozen section slides from CPTAC-3 studies18 (from 1,367 patients), 1,573 FFPE slides from the Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trials40 (from 475 patients), 759 FFPE slides from the Medical University of Vienna, Austria41 (Vienna; from 701 patients), 3,267 FFPE slides from the Dana-Farber Cancer Institute (DFCI; from 3,160 patients), 166 FFPE slides from Mayo Clinic (Mayo; from 166 patients), 157 FFPE and 165 frozen section slides from the Hospital of the University of Pennsylvania (UPenn; from 164 patients), and 121 FFPE and 112 frozen section slides from Mass General Brigham (MGB; from 154 patients). Detailed patient demographics are summarized in Table S1.

Overview of FAIR-Path

Figure 1 illustrates the overall workflow of the FAIR-Path system. First, we tessellated the WSIs into patches and removed those containing mostly blank backgrounds based on their color profiles. Our model employed a multiple instance learning (MIL) architecture8 with a two-stage training pipeline to ensure fairness in the learned image representations. In the first stage, we pretrained the MIL feature extractor with our proposed fairness-aware supervised contrastive learning objective, enhancing the similarity between features from images with the same ground-truth labels. A non-discrimination loss was incorporated to mitigate demographic-related but diagnostic-irrelevant attributes by penalizing similarities between samples sharing demographic attributes but differing in diagnostic categories. After this stage, we froze the model parameters of the MIL feature extractor and fine-tuned the classifier in a weakly supervised manner. Detailed descriptions of the model architecture and training procedure are provided in the methods section.

Figure 1.

Figure 1

The Fairness-Aware Artificial Intelligence Review for Pathology (FAIR-Path) framework

(A) The workflow of FAIR-Path. We first tessellated whole-slide pathology images (WSIs) into patches and removed those with predominantly blank backgrounds. FAIR-Path employs a two-stage training pipeline with multiple-instance learning to extract diagnostic signals irrespective of demographic attributes. In the first stage, supervised contrastive learning identifies diagnosis-related features independent of demographic-related imaging signals. In the second stage, we froze the pretrained feature extractor and fine-tuned the multilayer perceptron (MLP) classifiers.

(B) Patient cohorts. We collected 28,732 WSIs from 14,456 cancer patients and evaluated FAIR-Path across 82 classification scenarios, encompassing 27 diagnostic tasks, 20 cancer types, and eight cohorts. To assess internal validity, we quantified performance across patient subgroups defined by self-reported race, gender, and age using cohorts from The Cancer Genome Atlas (TCGA). To evaluate the generalizability of FAIR-Path, we employed 15 independent cohorts from the National Cancer Institute’s (NCI) Clinical Proteomic Tumor Analysis Consortium (CPTAC), the Prostate, Lung, Colorectal and Ovarian (PLCO) screening trials, the Medical University of Vienna, Austria (MUV), and institutional cohorts from Dana-Farber Cancer Institute (DFCI), Mass General Brigham (MGB), Mayo Clinic (Mayo), and the Hospital of the University of Pennsylvania (UPenn).

(C) Performance comparison. Standard deep learning models showed significant performance disparities (one-sided bootstrap test, p value < 0.05) across race, gender, and age groups, as measured by the equal opportunity metric. These disparities affected 8.75% of cancer detection, 12.1% of histological type classification, and 10.5% of subtype classification tasks (top left). FAIR-Path mitigated these biases in 88.5% of affected tasks, achieving 100% resolution in cancer detection and subtype classification, and 80.0% in histological type classification (top right). Bottom row: percentage of biased tasks in standard models and resolution rates by FAIR-Path, stratified by self-reported race, gender, and age groups.

See also Table S1.

Pan-cancer AI fairness analysis revealed demographic biases in AI-driven cancer pathology diagnosis

We conducted a pan-cancer fairness analysis to quantify demographic bias in baseline AI models across 82 classification scenarios, including 17 tumor detection tasks, 6 cancer type classification tasks, and 4 cancer subtype classification tasks, using TCGA datasets (see Table S2 for the list of classification scenarios). For each task, we trained a standard AI model8 with the same architecture as FAIR-Path but without incorporating demographic fairness considerations. We developed separate models for FFPE and frozen section slides due to their intrinsic differences and distinct clinical utility. We evaluated fairness with respect to self-reported race, gender, and age. We measured performance disparities using equal opportunity (EOpp; disparity in recall) and equal balanced accuracy (EBAcc). We compared FAIR-Path against existing fairness correction methods, including demographic-balanced reweighting29 (DBR), adaptive sensitive reweighting (ASR),42 learning fair representations (LFR),43 AdaFair,44 and synthetic minority over-sampling technique (SMOTE).45

Standard AI models show significant race, gender, or age bias in 29.3% (11 out of 37) of cancer diagnostic tasks (Figure 2). In FFPE-based models for assisting final cancer diagnosis, baseline models showed significant demographic disparities in three cancer classification tasks for at least one fairness measure (Figures 2A and 2B). For instance, in classifying lung adenocarcinoma (LUAD) vs. lung squamous cell carcinoma (LUSC), we observed significant racial disparity in equal opportunity for LUAD (EOpp (LUAD); bootstrap test p value = 0.048), along with gender disparities in EOpp (LUAD; p < 0.001) and EBAcc (p = 0.009). In differentiating invasive ductal carcinoma (IDC) from invasive lobular carcinoma (ILC) of the breast, the baseline model showed significant age disparities in EOpp(ILC) (p = 0.039) and EBAcc (p = 0.044). Additionally, the IDC/mixed IDC (IDC+) vs. ILC task showed significant racial disparities in EOpp(IDC+) (p = 0.045) and EOpp(ILC) (p = 0.040).

Figure 2.

Figure 2

Standard AI models showed significant performance disparities in cancer diagnosis tasks across population groups

Each fairness metric is represented by a different color, with higher values indicating greater disparity. Data are represented as micro-averaged metrics. Asterisks denote statistically significant bias.

(A) Cancer-type classification (FFPE samples). We observed significant racial and gender disparities for LUAD vs. LUSC classification.

(B) Subtype classification (FFPE). We identified significant racial disparities in classifying IDC and related ductal carcinomas vs. ILC and age disparities in the IDC vs. ILC classification.

(C) Cancer-type classification (frozen sections). We found significant racial disparities in the classification of GBM vs. LGG, KIRC vs. KICH, KIRC vs. KIRP, and KIRP vs. KICH.

(D) Subtype classification (frozen sections). No significant disparities were observed.

(E) Tumor detection (frozen sections). We found significant racial disparities in BRCA detection; significant gender disparities in KIRP and THCA detection; and significant age disparities in KIRP and STAD detection.

Abbreviations: BAC vs. LUAD: bronchoalveolar carcinoma vs. other lung adenocarcinoma subtypes; sUCEC vs. nsUCEC: serous vs. non-serous uterine corpus endometrial carcinoma; IDC vs. ILC: invasive ductal carcinoma vs. invasive lobular carcinoma; IDC+ vs. ILC: invasive ductal carcinoma or invasive ductal carcinoma mixed with lobular carcinoma vs. invasive lobular carcinoma. ∗bootstrap test p value < 0.05. See also Figures S1 and S2; Tables S1 and S2.

Frozen section-based models for intraoperative decision support revealed significant disparities in four classification tasks (Figures 2C and 2D). We observed significant racial disparities in the classification of glioblastoma multiforme (GBM) vs. low-grade glioma (LGG; EOpp(GBM), p = 0.022), renal clear cell carcinoma (KIRC) vs. kidney chromophobe (KICH; EOpp(KIRC), p < 0.001), KIRC vs. renal papillary cell carcinoma (KIRP; EOpp(KIRC), p = 0.040), and KIRP vs. KICH (EOpp(KIRP), p = 0.004). Across the 17 tumor detection tasks using frozen section slides from the TCGA cohorts, baseline models showed significant demographic disparity in four tasks (Figure 2E). Specifically, we identified racial disparities in breast cancer (BRCA) detection (EOpp(normal), p = 0.021; EBAcc, p = 0.025). The KIRP detection task showed gender- (EOpp(tumor), p = 0.005) and age-related disparities (EOpp(normal), p = 0.039; EBAcc, p = 0.025). In addition, the thyroid carcinoma (THCA) detection task exhibited gender-related disparity (EOpp(tumor), p = 0.040), while the stomach adenocarcinoma (STAD) detection task showed age-related disparity (EOpp(tumor), p = 0.002).

We found performance disparities mostly in underrepresented groups defined in medical studies,46,47,48 including racial minorities, females, or older age groups. Interestingly, 3 tasks showed lower performance in groups not underrepresented in the training cohorts. For example, we observed lower performance for male patients in LUAD vs. LUSC classification using FFPE samples. Younger age groups also showed lower performance in IDC vs. ILC classification using FFPE samples and in tumor detection of KIRP and STAD (Figure S1A). Pairwise analysis between each self-reported race group revealed that most of the significant racial biases were against African American patients. Racial disparities against Asian patients were observed in only one task (IDC/mixed IDC vs. ILC using FFPE), likely due to insufficient statistical power from smaller sample sizes (Figure S1B).

Additionally, these biases could not be fully explained by tissue procurement site, scanner type, magnification, patient stage, or patch selection strategies (Figure S2). These results suggested a more complex mechanism of model bias beyond technical factors and differences in healthcare access.

Data imbalances, genomic variations, and morphological differences contributed to the demographic biases of standard pathology AI models

We further investigated sampling and molecular factors contributing to demographic biases in baseline models. Many TCGA cohorts exhibit demographic imbalances (Table S1), and several diagnostic tasks show different demographic compositions between diagnostic classes (Table S2). We conducted subsample analyses by varying group imbalance and class distribution shifts and discovered that imbalanced sample sizes between demographic subgroups significantly correlate with larger biases (Figure 3A; Figure S3A). Furthermore, discrepancies in diagnosis prevalence between demographic groups, even with balanced sample sizes, contributed to performance disparities (Figure 3B; Figure S3B).

Figure 3.

Figure 3

Data imbalances and morphological differences contributed to performance disparities among demographic groups in cancer diagnosis tasks

(A) A subsample analysis shows that performance disparities between majority and minority populations increased as the sample size imbalance across demographic groups enlarged. The Spearman correlation between fairness metrics and sample size balance is shown.

(B) A subsample analysis shows greater performance disparities as differences in diagnostic category distribution between majority and minority populations in the training set widened. The Spearman correlation between fairness metrics and class balance in the minority groups is shown.

(C) Variations in somatic mutation rates partially explained performance discrepancies in baseline models across patient groups. We summarized the tasks where baseline models exhibited significant performance disparities between patient groups, alongside covariates that simultaneously showed significant differences between demographic groups and significant associations with predictive errors in baseline models. Whiskers in box plots indicate 1.5 × the interquartile range. ∗Wald test p value < 0.05; ∗∗Wald test p value < 0.01.

(D) Demographically distinct imaging features correspond to differences in cellular composition, including neoplastic, inflammatory, connective, and epithelial cell densities. Data are represented as mean ± SEM. Statistical comparisons between demographic subgroups were performed using the Mann-Whitney U test.

BRCA: breast invasive carcinoma; LUSC: lung squamous cell carcinoma; LUAD: lung adenocarcinoma. ∗Mann-Whitney U test p < 0.05; ∗∗p < 10−3; ∗∗∗p < 10−5. See also Figures S3 and S4; Tables S3–S5.

We further examined the relationships between the observed performance biases and the incidence of genomic mutations (see Methods for details). We found that differential incidence rates of somatic mutations correlate with performance gaps between groups. For example, variations in TP53 mutation rates are associated with differential error rates in IDC/mixed IDC vs. ILC, LUAD vs. LUSC, and GBM vs. LGG classification. Similarly, differences in CDH1 mutation rates were linked to significant racial disparity in IDC/mixed IDC vs. ILC and age discrepancy in IDC vs. ILC classification. These genetic mutations have significantly different prevalences in patient groups defined by age, gender, and self-reported race (Wald test p value < 0.05) and are significantly associated with predictive errors in standard AI baseline models (Wald test p value < 0.05; Figure 3C). Moreover, these somatic mutations are detectable from WSIs by AI models with significant accuracy (AUROCs = 0.556–0.769, bootstrap test p value < 0.05; see Table S3), a finding consistent with previous literature.3

In addition, many cancer types associated with demographic bias showed demographically distinct imaging characteristics. Supervised machine learning analysis using pathology imaging features in lung cancer patients revealed significant distinctions by race (AUROC = 0.628, Delong test p value< 0.001) and gender (AUROC = 0.629, p value< 0.001). For breast cancer patients, imaging characteristics are significantly predictive of race (AUROC = 0.641, p value < 0.001) and age (AUROC = 0.571, p -value < 0.001). We found significant racial distinctions across renal cancer subtypes, including KIRC vs. KIRP (AUROC = 0.683, p value < 0.001), and KIRC vs. KICH (AUROC = 0.745, p value < 0.001). Stratification by cancer subtype reduced the significance of racial distinctions in lung and renal cancers. In breast cancer, however, morphological patterns associated with race and age persisted within ductal and lobular subtypes (AUROC = 0.543–0.631, p < 0.05; Table S4). Stratified analyses showed that such distinctions in imaging characteristics remain significant after controlling for other demographic attributes and prevalent somatic mutations (Figure S4).

We next investigated the cellular basis for these demographic distinctions and found significant differences in tumor composition (Figure 3D). We observed consistent racial differences across breast cancer subtypes, where tumors from African American patients had lower inflammatory and connective cell densities and higher overall and neoplastic cellularity than tumors from Caucasian patients (Mann–Whitney U test, p value < 0.001). Epithelial cell density was lower for African American patients in mixed IDC tumors but higher in ILC tumors (p value < 0.001). Age correlated with distinct morphological profiles across breast cancer subtypes: tumors from elderly patients consistently showed lower overall, inflammatory, connective, and epithelial cell densities, but a higher neoplastic cell density (p value < 0.001). Our findings are consistent with existing evidence that a subset of AI-extracted pathology features is strongly associated with demographic variables.24

FAIR-Path framework mitigates demographic biases in cancer pathology diagnosis

We employed FAIR-Path to mitigate the fairness issues discovered by our systematic fairness evaluation. For each task where baseline models showed significant bias, we used the same model architecture and trained it with FAIR-Path’s fairness-aware contrastive learning approach (Figure 1). We also reported group-stratified and overall performance using a diverse set of performance metrics (see Table S2).

For FFPE-based tasks, FAIR-Path successfully mitigated all significant biases in standard models (Figure 4A). In contrast, existing fairness correction methods demonstrated resolution rates of up to 60% for EOpp, while one competing method (SMOTE) also achieved a 100% resolution rate for EBAcc. For frozen section diagnostic tasks, FAIR-Path mitigated 77.8% of the significant EOpp biases and 75% of the EBAcc biases (Figure 4B). This substantially outperformed the leading competing methods, ASR and DBR, which achieved mitigation rates of 55.6% for EOpp and 50% for EBAcc. Notably, bias correction by FAIR-Path exhibited no significant decrease in overall performance in nearly all tasks (bootstrap test p value > 0.05). The only exception was the race correction for LUAD vs. LUSC classification (p value < 0.001). In this case, the leading competing methods (ASR and DBR) failed to mitigate the bias and exacerbated the disparities (Figures 4C and 4D).

Figure 4.

Figure 4

FAIR-Path more effectively mitigated biases in critical cancer diagnostic tasks compared to existing methods

Data are represented as micro-averaged metrics.

(A) Bias mitigation in diagnostic tasks using FFPE samples. FAIR-Path (red bars) eliminated significant biases in both equal opportunity and equal balanced accuracy across all tasks. In contrast, success rates for other methods ranged from 33.4% to 66.7%. Red asterisks indicate significant disparities.

(B) Bias mitigation in diagnostic tasks using frozen section samples. FAIR-Path mitigated biases in 75% of tasks with equal balanced accuracy bias and 77.8% of those with equal opportunity bias, respectively. This surpassed other fairness correction methods, which corrected only 33.3%–55.6% of such tasks.

(C) FAIR-Path improves both fairness and performance in FFPE-based pathology evaluation. For tasks with significant bias in baseline models (black dots), FAIR-Path (red crosses) consistently reduced bias while increasing the overall AUROC.

(D) We observed similar gains by FAIR-Path in frozen section samples. Among tasks showing performance disparities, FAIR-Path (red crosses) consistently reduced bias in baseline models (black dots) and enhanced overall AUROC.

ASR: adaptive sensitive reweighting; LFR: learning fair representation; SMOTE: synthetic minority over-sampling technique. ∗bootstrap test p value < 0.05. See also Figures S5 and S6; Tables S1, S2, S6, and S7.

Furthermore, bias mitigation by FAIR-Path preserved pathological signals for clinically relevant biomarkers with demographic variations in prevalence (Table S5). In addition, FAIR-Path model ensembles can achieve enhanced fairness across multiple demographic attributes (Figure S5A). In scenarios where FAIR-Path successfully mitigated baseline biases, intersectional fairness analyses revealed no significant bias across subgroups (Figure S5B). We project that fairness improvements by FAIR-Path could prevent thousands of misdiagnoses among minority patients in the U.S. annually if implemented nationwide. For example, in classifying IDC/mixed IDC vs. ILC among African American female patients, the fairness improvement of FAIR-Path will prevent the misdiagnosis of approximately 5,902 African American female patients each year (Table S6). However, some biases persisted after mitigation through FAIR-Path, occurring exclusively in frozen section evaluation tasks (GBM vs. LGG classification and KIRC vs. KIRP classification).

We further conducted ablation studies to validate the efficacy of our framework (see Methods). First, an otherwise identical supervised contrastive learning framework without the proposed fairness objective exhibited significant demographic bias in 4 of 8 tasks, demonstrating that contrastive learning alone is insufficient to ensure fairness (Figure S6A). Second, we tested whether demographic bias is a proxy for site-related factors by adapting FAIR-Path to correct procurement sites. Correcting site-related bias did not eliminate demographic disparities, as 9 out of 12 tasks retained persistent demographic bias after site correction (Figure S6B). This result indicates that demographic bias arises from mechanisms more complex than site-specific variations, reinforcing the need for our explicit fairness-aware approach. Lastly, we assessed the reliability of FAIR-Path in low-data regimes by training it on progressively smaller subsets of the training data. We found that FAIR-Path’s ability to mitigate bias was preserved under reduced training sample sizes, demonstrating its robustness (Figure S6C).

Independent validation demonstrated the generalizability of fairness enhancement by FAIR-Path

We tested the generalizability of FAIR-Path on FFPE and frozen section slides from 15 independent validation cohorts not involved in model development. After excluding tasks with insufficient minority group samples, we evaluated 39 classification settings across all diagnostic tasks, tissue sources, and demographic attributes (Table S2).

We found that standard pathology AI models exhibited significant biases in 60% of tumor detection tasks, 40% of cancer type classification tasks, and half of cancer subtype classification tasks across independent validation cohorts (Figure 5). In particular, we observed significant gender disparities in GBM vs. LGG classification in the DFCI FFPE cohort and gender disparities in both the MGB and DFCI FFPE cohorts. For KIRC vs. KIRP classification, we found significant gender bias in the DFCI FFPE cohort. The LUAD vs. LUSC classification models showed significant age bias in the CPTAC cohorts, along with both gender and age biases in the PLCO cohorts. In the COAD vs. READ classification, the standard model demonstrated significant age and gender biases in the DFCI cohort. For bronchoalveolar carcinoma detection from other LUAD in the PLCO cohort, the baseline model showed significant age bias. In tumor detection tasks using the CPTAC cohorts, standard models exhibited significant age disparities for detecting both KIRC and LUAD. In addition, only 3 out of 9 tasks that showed significant bias in independent cohorts also demonstrated significant bias in the TCGA cohorts. These differences highlight distributional shifts between cohorts and the robustness of FAIR-Path.

Figure 5.

Figure 5

Baseline models showed significant performance disparities in cancer diagnosis tasks across population groups in external validation cohorts

Higher values reflect greater disparity. Significant biases are marked with asterisks.

(A) Cancer-type classification (FFPE samples). Baseline models showed significant performance disparities between demographic groups in 41.7% of these tasks, including COAD vs. READ in the DFCI cohort, GBM vs. LGG in the DFCI and UPenn cohorts, KIRC vs. KIRP in the DFCI cohort, and LUAD vs. LUSC in the PLCO cohort.

(B) Subtype classification (FFPE). BAC vs. LUAD classification showed significant age disparity in the PLCO cohort.

(C) Cancer-type classification (frozen sections). LUAD vs. LUSC classification showed significant age disparity in the CPTAC cohort.

(D) Tumor detection (frozen sections). 50% of these tasks showed significant age disparity, including the detection of KIRC and LUAD in the CPTAC cohort.

∗bootstrap test p value <0.05. See also Tables S1 and S2.

Among the biased diagnostic tasks in external FFPE cohorts, FAIR-Path successfully mitigated 88.9% of the observed EOpp bias and 100% of the EBAcc bias (Figure 6A). In comparison, other methods were less effective: LFR achieved the next highest resolution rates, resolving 55.6% of EOpp bias and 100% of EBAcc bias, while other methods ranged from 44.4 to 55.6% for EOpp and up to 66.7% for EBAcc. The only instance where FAIR-Path did not fully mitigate bias was in the COAD vs. READ classification in the DFCI dataset. Notably, none of the other methods succeeded in this task. All FAIR-Path-corrected tasks in the independent FFPE cohorts showed no significant drop in overall AUROC performance (one-sided bootstrap test p > 0.05; Figure 6C). In contrast, although ASR, LFR, and DBR also showed no significant AUROC drop, AdaFair and SMOTE caused significant AUROC declines in 30% and 10% of tasks, respectively.

Figure 6.

Figure 6

FAIR-Path mitigated biases in cancer diagnostic tasks across 15 external validation cohorts, outperforming existing methods

(A) FFPE-based diagnostic tasks. FAIR-Path (red bars) is compared to standard deep learning models (black bars) and other correction methods. Red asterisks indicate significant bias. FAIR-Path resolved 88.9% of significant baseline biases in equal opportunity and 100% in equal balanced accuracy. LFR achieved the second-highest resolution rates, resolving 55.6% and 100% for equal opportunity and equal balanced accuracy, respectively.

(B) Diagnostic tasks using frozen sections. FAIR-Path corrected all significant baseline biases, whereas other methods achieved resolution rates of 80%–100% for equal opportunity and 75%–100% for equal balanced accuracy, respectively.

(C) Fairness vs. performance in FFPE-based tasks with significant baseline bias. FAIR-Path (red markers) reduced bias without compromising AUROC, outperforming baseline models (black dots) and other methods. Stars indicate no significant bias or AUROC drops. Arrows indicate significant change from baseline: an upward arrow indicates significant bias, a leftward arrow indicates a drop in AUROC, and an upper-left arrow indicates both.

(D) Fairness vs. performance in frozen section diagnostic tasks with significant baseline bias. FAIR-Path resolved biases without significantly reducing overall AUROC in tasks including LUAD vs. LUSC classification and KIRC detection. In the LUAD detection task, AdaFair, SMOTE, and FAIR-Path all exhibited significant AUROC drops compared to the baseline model.

ASR: adaptive sensitive reweighting; LFR: learning fair representation; SMOTE: synthetic minority over-sampling technique. ∗bootstrap test p value < 0.05. See also Tables S1 and S2.

In the frozen section cohorts, FAIR-Path eliminated all significant biases (Figure 6B). AdaFair and LFR also achieved complete bias mitigation, while other fairness correction methods showed slightly lower resolution rates: ASR and DBR resolved 66.7% of EOpp biases and 100% of EBAcc biases, and SMOTE achieved 100% resolution for EOpp and 66.7% for EBAcc. Notably, FAIR-Path showed a significant AUROC drop only in the LUAD detection task, where we observed more substantial performance-fairness trade-offs with AdaFair and LFR (Figure 6D).

FAIR-Path enhanced fairness in pathology foundation models

Pathology foundation models3,5,49—large models pretrained on diverse datasets using self-supervised learning—have drawn substantial interest for their strong performance and generalizability. However, prior studies have shown that these models exhibit demographic disparities when fairness is not explicitly addressed during downstream training.24 To demonstrate the utility of FAIR-Path in this context, we evaluated fairness metrics for three state-of-the-art pathology foundation models (CHIEF,3 UNI,5 and GigaPath49) in cancer classification tasks across TCGA cohorts. Our result showed that, without a fairness constraint, foundation model-based diagnostic classifiers displayed significant performance disparities by age, gender, and self-reported race, with CHIEF exhibiting the highest bias rate (77.8%), followed by UNI (66.7%) and GigaPath (66.7%; see Figure 7A). Applying FAIR-Path successfully mitigated these disparities in 73.7%, 66.7%, and 57.1% of tasks for CHIEF, GigaPath, and UNI, respectively (Figure 7B). These findings demonstrate the effectiveness of FAIR-Path in enhancing fairness in both standard end-to-end training pipelines and state-of-the-art foundation models. Notably, the racial bias in KIRC vs. KIRP classification, which we observed consistently across all foundation models, remained significant after applying FAIR-Path. Our analysis revealed a notable difference between this task and those where bias was consistently mitigated (IDC vs. ILC and sUCEC vs. nsUCEC). In the latter cases, we found significant racial disparities in the prevalence of common mutations, whereas no such disparities were observed in KIRC or KIRP (Table S7). These findings suggest that FAIR-Path is more effective when demographic bias is linked to differences in genomic profiles but may be less so when disparities originate from non-mutational factors.

Figure 7.

Figure 7

FAIR-Path mitigates bias in pathology foundation models

Data are represented as micro-averaged metrics.

(A) Pathology foundation models exhibited race, gender, and age bias in classifying cancer types and subtypes. When trained without demographic bias correction, all evaluated foundation models (CHIEF, GigaPath, and UNI) exhibited significant performance disparity in 66.7%–77.8% of the classification tasks. Asterisks indicate statistically significant disparities between demographic subgroups.

(B) In tasks where standard pathology foundation models (CHIEF, GigaPath, and UNI) exhibited significant performance disparities (black bars), FAIR-Path (red bars) eliminated these biases in 73.7%, 66.7%, and 57.1% of tasks for each model, respectively.

∗bootstrap test p value < 0.05. See also Tables S1 and S2.

Discussion

Despite rapid advancements in AI, fairness issues in real-world medical diagnostic tasks remain a pressing concern. Our pan-cancer fairness analysis demonstrated that standard AI diagnostic models and advanced pathology foundation models can inadvertently discriminate against patients from underrepresented subpopulations, even when overall performance appears strong. AI models can encode representation biases from demographically imbalanced pre-training datasets, and their downstream training is vulnerable to shortcut learning—a risk that persists regardless of the feature encoder’s representational power unless fairness constraints are explicitly applied. These findings highlight the importance of fairness evaluation and the urgent need for equitable frameworks in AI pathology to prevent perpetuating health disparities. To address this, we introduce FAIR-Path, a fairness-aware contrastive learning framework designed to enhance fairness in AI-empowered pathology image evaluation. Through extensive experiments across multiple cancer types and external validation involving five medical centers and three nationwide study cohorts, we demonstrated that FAIR-Path, coupled with fairness auditing, successfully reduced performance disparities across demographic subgroups in both standard AI diagnostic models and pathology foundation models. This framework represents a promising step forward in addressing the ethical and fairness challenges in pathology AI diagnosis.

We found that demographic attributes are associated with tumor histology differences detectable by AI models. In older patients, tumors consistently exhibit increased stromal volume and diminished inflammatory infiltration compared to younger patients. This pattern aligns with established features of aging tissues, including fibroblast senescence, extracellular matrix accumulation, and immune exclusion.50,51 Similarly, tumors from African American patients show a higher density of neoplastic cells and a lower presence of immune and stromal components compared to those from Caucasian patients. This histologic profile is consistent with prior reports of elevated levels of CD163+ tumor-associated macrophages52 and exhausted CD8+ T cells53 in tumors from individuals of African ancestry, both of which are linked to impaired immune surveillance. Higher microvessel density has also been reported in tumors from African American patients, suggesting enhanced angiogenesis and stromal remodeling.54 Collectively, these observations indicate that morphological variations associated with demographic factors are embedded in tissue architecture and can be inadvertently internalized by AI models.

Our analyses revealed that differences in sample sizes, class distribution, and morphological variations between demographic groups contributed to disparate performance in pathology diagnosis. The significant demographic biases we identified align with well-documented demographic differences in disease prevalence for several cancer subtypes, including renal,55,56 lung,57 brain,58 breast,59 and stomach cancer.60 In addition, we showed that variations in mutation prevalence, particularly for TP53 and CDH1, partially explain performance disparities in baseline models across several diagnostic tasks. For example, TP53, the most commonly mutated gene in advanced-stage cancers,61 has shown significant differences in mutation rates across racial62,63,64,65,66,67 and gender-based groups.68 Conversely, CDH1 mutations are important markers for invasive lobular carcinoma (ILC),69 are more prevalent in late-onset breast cancer,70 and are more common in Caucasian patients with breast, gastric, and colorectal cancers.71 Pathology AI can identify cell morphology associated with genetic aberations,3,72 which is also further confirmed by our experimental results. Disparities in gene mutation incidence across populations may contribute to subtle morphological differences, leading to variations in AI model performance across demographic groups that are not fully attributable to technical factors.25 This is supported by our analyses, where demographic bias persisted after accounting for technical variables. Thus, AI-based cancer diagnostic tools should account for these morphological differences to minimize bias. This underscores the importance of implementing fairness-aware data collection processes and training strategies.

We observed that FAIR-Path’s efficacy is consistently lower for frozen section samples. This reduced efficacy may arise from the greater variability in image quality inherent to frozen section slides,73 attributable to variable staining profiles, ice-crystal artifacts, and thicker samples. To address these limitations, potential solutions include rigorous quality-control pipelines, advanced data preprocessing, and domain adaptation techniques.74,75,76,77 Ultimately, building a large, multi-institutional repository of standardized frozen section slides will support these technical solutions and further reduce biases. These strategies chart a path toward enhanced fairness and reliability for intra-operative diagnostic assessment.

In summary, our systematic pan-cancer analysis revealed that standard computational pathology models frequently exhibit significant demographic biases, disproportionately affecting underrepresented racial groups, female patients, and older individuals. In addition, we identified the molecular and sampling factors contributing to these disparities. By incorporating fairness-aware contrastive learning, our FAIR-Path framework substantially mitigated these performance gaps across a wide range of pathology evaluation tasks in diverse external cohorts. This enhanced fairness in AI-driven pathology models will lead to more equitable diagnostic assessments, reducing the risk of misdiagnosis and enhancing patient outcomes.

Limitations of the study

Our study has a few limitations. First, our racial group investigation was limited to Caucasian, African American, and Asian patients. Biases against other racial minorities, such as Hispanics, Native Americans, Alaska Natives, or Pacific Islanders, could not be reliably evaluated due to insufficient sample sizes. In addition, our pan-cancer fairness analyses were limited to individual demographic attributes. Understanding how multiple demographic factors interact could provide further insights into the complex mechanisms of algorithmic bias, but intersectional analysis78,79 is particularly challenging due to the limited sample sizes of intersectional subgroups defined by multiple demographic factors. Furthermore, although we explained some performance disparities through differences in sample sizes, disease category distribution, and genomic profiles, other potential explainers of bias—such as disparities in socioeconomic status,23 population differences between geographical areas, or differences in comorbidities46,80—remain to be explored. Finally, systematic fairness evaluation of emerging foundation models and vision-language models3,81,82 is needed to further examine their performance in diverse populations.

Resource availability

Lead contact

Data access inquiries should be directed to the lead contact, Kun-Hsing Yu (kun-hsing_yu@hms.harvard.edu), and will be forwarded to the managers of these institutional datasets. These data are strictly only for non-commercial academic use.

Materials availability

This study did not generate new unique reagents.

Data and code availability

  • The digital pathology slides, summarized clinical data, and biospecimen information for The Cancer Genome Atlas (TCGA) and Clinical Proteomic Tumor Analysis Consortium (CPTAC) cohorts are available through the National Cancer Institute Genomic Data Commons (https://portal.gdc.cancer.gov/). Summarized molecular data for these cohorts are available from cBioPortal (https://www.cbioportal.org/). Data from the Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trials can be obtained from the NCI Cancer Data Access System (https://cdas.cancer.gov/plco/). Data from the Medical University of Vienna are available through the EBRAINS portal (https://www.ebrains.eu/). Datasets from Mayo Clinic, Dana-Farber Cancer Institute (DFCI), University of Pennsylvania (UPenn), and Mass General Brigham (MGB) are not publicly available due to patient privacy concerns and data use agreement requirements. To request access to these data, please contact the lead contact, Kun-Hsing Yu (kun-hsing_yu@hms.harvard.edu). The requestor must describe the objectives of the research project for which the data will be used. We aim to forward all requests to the managers of these institutional datasets within 2 weeks, and these requests will be evaluated according to institutional policies. Data access will be considered for research purposes and non-commercial use only. In order to ensure patient privacy, access to personally identifiable information or sensitive clinical information will not be provided, and requests for data access must rigorously adhere to the consent agreements established with study participants.

  • All original code has been deposited at GitHub and is publicly available as of the date of publication. The code for our model and data analyses can be found at https://github.com/hms-dbmi/fairpath.

  • Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

Acknowledgments

We thank Alessia Hughes, Faith McDonald, Catherine Burroughs, and Mariam Kapanadze for their administrative support. We thank Sunni Lin for her technical assistance and Dr. Mu-Hung Tsai for his help in grouping pathology subtypes. K.-H.Y. is partly supported by the National Institute of General Medical Sciences grant R35GM142879, the National Heart, Lung, and Blood Institute grant R01HL174679, the Department of Defense Peer Reviewed Cancer Research Program Career Development Award HT9425-23-1-0523, the Research Scholar Grant RSG-24-1253761-01-ESED (grant DOI: https://doi.org/10.53354/ACS.RSG-24-1253761-01-ESED.pc.gr.193749) from the American Cancer Society, Google Research Scholar Award, and the Harvard Medical School Dean’s Innovation Award. J.-H.C. is supported in part by the National Science and Technology Council (NSTC), Taiwan, under research grant numbers NSTC 113-2917-I-006-009, NSTC 112-2634-F-006-003, and NSTC 113-2321-B-006-023. P.-C.T. is supported by the Doctoral Students Scholarship from Xin Miao Education Foundation. F.-Y.S. is partly supported by the National Science and Technology Council (NSTC), Taiwan, under research grant number NSTC 114-2917-I-006-016. We thank the AWS Cloud Credits for Research, Microsoft Azure for Research Award, the NVIDIA GPU Grant Program, and the Extreme Science and Engineering Discovery Environment (XSEDE) at the Pittsburgh Supercomputing Center (allocation TG-BCS180016) for their computational support.

Author contributions

Conceptualization, S.-Y.L., P.-C.T., F.-Y.S., T.-L.M.L., J.-H.C., and K.-H.Y.; methodology, S.-Y.L., F.-Y.S., C.-Y.C., and K.-H.Y.; validation, S.-Y.L., P.-C.T., F.L., and E.H.; investigation, S.-Y.L., P.-C.T., F.-Y.S., T.-W.K., D.V., J.Z., C.-Y.C., F.L., Y.Y.H., E.H., P.-J.L., J.-H.C., and K.-H.Y.; resource, T.R.-P., L.S., D.D., N.U.L., D.M., K.L.L., Y.-C.L., N.C., D.J.C., A.W., J.M., S.O., M.P.N., J.A.G., S.S., and K.-H.Y.; data curation, T.R.-P., L.S., D.D., N.U.L., D.M., K.L.L., Y.-C.L., N.C., D.J.C., A.W., J.M., S.O., M.P.N., and K.-H.Y.; writing – original draft, S.-Y.L., P.-C.T., F.-Y.S., F.L., P.-J.L., and K.-H.Y.; writing – review & editing, all authors; visualization, S.-Y.L., P.-C.T., J.Z., T.-W.K., and Y.Y.H.; supervision, J.-H.C. and K.-H.Y.; project administration, J.-H.C. and K.-H.Y.; funding acquisition, J.-H.C. and K.-H.Y.

Declaration of interests

K.-H.Y. is an inventor of U.S. Patent 10,832,406. This patent is assigned to Harvard University and is not directly related to this manuscript. K.L.L. was a consultant of Travera, BMS, Servier, Integragen, LEK, and Blaze Bioscience, received equity from Travera, and has research funding from BMS and Lilly (not related to this work). D.V. is a co-founder and shareholder of Vectorly, Inc. (not related to this work).

Declaration of generative AI and AI-assisted technologies in the writing process

After preparing the initial manuscript, the authors used ChatGPT to edit selected sections to improve readability. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.

STAR★Methods

Key resources table

REAGENT or RESOURCE SOURCE IDENTIFIER
Deposited data

Digital pathology slides and summarized clinical data (TCGA) National Cancer Institute – Genomic Data Commons (GDC) https://portal.gdc.cancer.gov
Digital pathology slides (CPTAC-3) National Cancer Institute/GDC https://portal.gdc.cancer.gov
Summarized molecular profiles (TCGA/CPTAC) cBioPortal https://cbioportal.org
PLCO Cancer Screening Trial data NCI Cancer Data Access System (CDAS) https://cdas.cancer.gov/plco/
Medical University of Vienna cohort portal EBRAINS https://ebrains.eu

Software and algorithms

FAIR-Path source code (model and analysis) GitHub https://github.com/hms-dbmi/fairpath
NVIDIA NGC PyTorch base docker image (development environment) NVIDIA NGC 22.03
Python (main program language) Python 3.8
PyTorch (external validation) PyTorch 2.0.0
NumPy (numerical operations) NumPy 1.22.4
Scikit-learn (metric calculation) Scikit-learn 1.1.2
SciPy (Mann–Whitney U test, Spearman rank correlation test, bootstrapping) SciPy 1.16.1
WORC (DeLong test) WORC 3.7.0
statsmodels (GLM) statsmodels 0.14.2

Experimental model and study participant details

Human participants

All patients provided written informed consent at the time of enrollment. Our multi-center study was approved by the Harvard Medical School Institutional Review Board (IRB25-0188 for the SmartPath Research Network and IRB 20–1509). We obtained digital hematoxylin and eosin (H&E)-stained pathology slides from the National Cancer Institute-sponsored TCGA, CPTAC-3, and PLCO Cancer Screening Trials. In addition, we digitized 4,747 slides from the Mayo Clinic, MGB, UPenn, DFCI, and the Medical University of Vienna, Austria. Digital slides from the TCGA, CPTAC-3, and PLCO cohorts were scanned at 40X or 20X magnification. We scanned pathology slides from the Mayo Clinic at 40X with Aperio GT450 scanners, with pixel widths of 0.263 microns per pixel. We digitized slides from MGB and UPenn by Hamamatsu S210 scanners at 40X magnification, with pixel widths ranging from 0.221 to 0.253 microns per pixel. All WSIs used in this study were derived from pre-treatment tumor samples, and thus the potential differential access to treatments will not affect our analyses. We digitized slides from the Medical University of Vienna with a Hamamatsu NanoZoomer 2.0 HT slide scanner at 40X magnification with a pixel width of 0.228 microns per pixel. We scanned slides from DFCI using Aperio ScanScope scanners at 40X (0.248 microns per pixel) or 20X (0.496 microns per pixel). Participant age, gender, and race distributions for TCGA cohorts are summarized in Table S1 of the main text; demographics for external validation cohorts are summarized in the “Patient demographics for external cohorts” table. Analyses in this study explicitly evaluate performance across demographic attributes (age, gender, self-reported race) and include site-, scanner-, and magnification-stratified assessments reported in the Results and Table S3.

Method details

Whole-slide preprocessing and patch sampling

We preprocessed the whole-slide images (WSIs) by tiling each image into non-overlapping patches of 512 × 512 pixels at 20X magnification. To separate tissue regions from the background, we converted the WSIs to grayscale and excluded pixels with intensity values greater than 225 via foreground filtering. We removed patches with an insufficient size (<512 × 512 pixels) or low tissue content (a foreground ratio <0.1). This filtering strategy effectively eliminates regions with minimal biological signal without selectively enriching for any specific microenvironment. We then divided the list of image patches into equal-length segments and evenly sampled 200 patches per patient from these segments. This strategy ensures a balanced representation, preventing the model from overemphasizing high-tumor regions while retaining potential morphological patterns relevant to demographic factors.

Fairness-aware contrastive pretraining

FAIR-Path employs a multiple instance learning (MIL) architecture with ResNet83 as the tile-level feature extractor, which aligns with standard pathology AI frameworks such as CLAM2 and MOMA8 that employed ResNet as the backbone for image feature extraction. We selected ResNet-18 to balance computational efficiency and performance during the contrastive learning phase of our fairness-enhancement framework. Following tile-level encoding, a MIL module aggregates features across all tiles, which are then processed by a two-layer multilayer perceptron (MLP) projector.

During the pretraining process, we trained the slide-level encoder, along with a two-layer multilayer perceptron (MLP) projector, using supervised contrastive loss and non-discrimination loss. The pretraining was performed in an end-to-end manner, where all model parameters were trained during the process. After training the FAIR-Path slide-level encoder, we froze its parameters and trained a three-layer MLP for cancer diagnostic tasks using a group-balanced reweighting approach.29

For randomly sampled data triplets, {bi,yi,si}i=,,N , in a batch, where bi,yi, and si denotes the bag of the input samples, target labels, and demographic attribute labels, respectively, with an equal number of samples drawn from each demographic group and random sampling performed within each group. After augmentation, we obtained 2N data triplets in a batch, {bˆi,yˆi,sˆi}i=1,,2N. The feature extractor encodes patches into tile-level features, which are grouped into bags corresponding to the original slides and input into the attention-based MIL module. The MIL module computes attention weights, which are multiplied by their respective embeddings and then summed to form the bag representations, ziZ for i = 1, …, 2N. We then projected these representations into another feature space, where we compute the fairness-aware supervised contrastive loss using supervised contrastive loss and non-discrimination loss.

The non-discrimination loss (LNDL) extends supervised contrastive learning84 by addressing the fairness constraint. Mathematically,

LNDL=ziϵZlog{1|Ztp(i)|ztpϵZtp(i)exp(zi·ztp/τ)zjϵ(Za(i)Znn(i))exp(zi·ztp/τ)}

where the set of representations that share the same diagnostic category but have different demographic attributes from zi is denoted as Ztp{ztpZa(i):yˆp=yˆi,sˆpsˆi}. Similarly, the set of representations with both different diagnostic categories and demographic attributes from zi is denoted as Znn{znnZa(i):yˆpyˆi,sˆpsˆi}.

We incorporated the non-discrimination loss LNDL into the training process to mitigate the model’s tendency to encode imaging features associated with demographic attributes rather than diagnostic categories, which could lead to performance differences across population groups. By pushing apart representations with the same demographic attribute but different diagnostic labels, our model learns to focus on diagnosis-related features instead of demographic characteristics. This enhances the model’s fairness by reducing the risk of bias related to these demographic attributes. As a result, the model becomes less likely to unfairly favor or penalize certain population groups, thereby achieving more equitable outcomes across different demographic groups.

Our model training process is guided by a loss function, denoted as Lcontrast, which combines both the supervised contrastive loss and the non-discrimination loss. I.e.,

Lcontrast=Lsup+LNDL

We trained our contrastive learning models with SGD using a learning rate of 5 × 10−3, a batch size of 8, and 100 epochs with a cosine annealing learning rate scheduler.

Classifier training

We conducted pretraining in an end-to-end manner, with all model parameters optimized during the process. In the second stage, we froze the parameters of the pretrained encoders and trained a three-layer MLP for cancer diagnosis using a simple group-balanced reweighting approach.29 We trained the classifier with an SGD using a learning rate of 10−4, a batch size of 8, and 50 epochs with a cosine annealing learning rate scheduler. In our experiments, the hyperparameters were set as above across all tasks, and no hyperparameter tuning was required.

Sample size and allocation to analysis groups

We included all pathology diagnostic tasks with sufficient sample sizes (≥8 patients) in each diagnostic category and demographic subgroups for performance and fairness evaluation. This resulted in 27 classification scenarios, including 17 tumor detection tasks, 6 cancer type diagnosis tasks, and 4 cancer subtype diagnostic tasks across 20 cancer types. We repeated the process independently for FFPE and frozen section samples. For each task, we employed a 4-fold cross-validation (CV) to train the model, with a training-validation-test split ratio of 2:1:1 for each fold. There was no patient overlap among training, validation, and testing sets. We selected the number of epochs with the highest AUROC in the validation set. After finalizing the model, test set metrics for each fold were aggregated to provide the final evaluation results for our developmental cohorts.

To evaluate the generalizability of our framework, we performed external validation on patient cohorts from CPTAC-3, PLCO, DFCI, Mayo Clinic, UPenn, MGB, and the Medical University of Vienna. For the external validation experiments, we only included tasks with baseline results from the TCGA cohort and sufficient sample sizes for each cancer type in each demographic subgroup (N > 8). We listed all tasks included in our validation study in Table S2.

Baseline training details

In the baseline model, we adopted ResNet-18 as the feature extractor and replaced the last fully connected layer with a linear layer and a rectified linear unit (ReLU).85 We employed gated attention-based multiple instance learning to aggregate patches from the same slides. For more efficient training, we used a ResNet-18 backbone pretrained on ImageNet and subsequently fine-tuned it using pathology images during model training. We trained the baseline models using stochastic gradient descent (SGD) with a learning rate of 5 × 10−3 and a batch size of 8. Because the supervised contrastive learning framework mitigated bias for one demographic attribute at a time, we trained separate FAIR-Path models for each demographic characteristic.

External validation of FAIR-Path

To evaluate the generalizability of our framework, we performed external validation on patient cohorts from CPTAC-3, PLCO, DFCI, Mayo Clinic, UPenn, MGB, and the Medical University of Vienna. For the external validation experiments, we only included tasks with baseline results from the TCGA cohort and sufficient sample sizes for each cancer type in each demographic subgroup (N > 8). We listed all tasks included in our validation study in Table S2.

Based on the slide preservation method (frozen section or FFPE), demographic attributes (race, age, or gender), and the type of diagnostic tasks (tumor detection, cancer type classification, or subtype classification), we selected the corresponding baseline and FAIR-Path models trained on the TCGA dataset. For each model, we fixed the embeddings extracted by the slide-level encoder and retrained only the subsequent classification layers. To account for domain differences, we re-estimated the normalization statistics (i.e., mean and variance) of all batch normalization layers in the model using data from the external cohorts. We implemented our methods using PyTorch 2.0.0.

Patch selection

To assess the potential impact of patch selection on the demographic bias in standard AI models, we compared the resulting fairness from our patch selection method with that of two alternative patch selection strategies. The first strategy is a random selection approach, where 200 patches were randomly selected from each WSI during each iteration of the training process. The second patch selection strategy involves random sampling of 200 patches in a demographic-balanced manner, where the samples were drawn from each patient group with equal sampling probabilities. As a representative example, we compared the racial biases from these patch selection strategies on cancer type and subtype classification tasks using FFPE cohorts.

Evaluation metrics of model fairness

We evaluated the fairness of our models’ performance metrics, including recall, balanced accuracy, accuracy, and AUROC. We focus on the two standard fairness measurements, including equal opportunity (EOpp) and equal balanced accuracy (EBAcc), given their clinical significance. We also provide additional fairness measurements in Table S2, including equalized odds (EOdd), difference of accuracy (AccDiff), and difference of AUROC (AUCDiff).

For a class in a binary classification task, we defined its EOpp (denoted as EOppc) as the difference between the highest and lowest recall values for class across subgroups defined by a demographic attribute (S):

EOppc=maxsS(Recallsc)minsS(Recallsc)

EBAcc quantifies the maximum between-group discrepancy in balanced accuracy (BA), where BA is the macro-average of recalls for all ground truth classes:

EBAcc=maxsϵS(BAs)minsϵS(BAs)

where:

BA=1|C|cϵCRecallc

Equalized Odds quantifies the bias through the maximum absolute between-group disparity across true positive rate (TPR) and false positive rate (FPR), thereby promoting fairness across both positive and negative categories:

EOdd=max(|FPRunprivilegedFPRprivileged|,|TPRunprivilegedTPRprivileged|)
EBAcc=maxsϵS(BAs)minsϵS(BAs),BA=1|C|cϵCRecallc

Similarly, we defined AccDiff and AUCDiff as the maximum between-group discrepancies in overall accuracy and AUROC, respectively:

AccDiff=maxsϵS(Accs)minsϵS(Accs)

and

AUCDiff=maxsϵSAUROCsminsϵSAUROCs

The performance and fairness metrics listed above were estimated based on model predictions for each individual. To assess performance disparities measured by each fairness metric, we employed non-parametric, one-sided bootstrap hypothesis testing86 to evaluate the significance of the observed bias. For each task, we first aggregated the labels and predictive scores from the test cohorts across all 4-folds. Subsequently, for each demographic group, we resampled the label-score pairs with replacement from the combined testing dataset. The bootstrapped fairness metrics were estimated from these resampled pairs. In total, we generated 10,000 bootstrap samples for each metric. For a given fairness metric, the P-value was defined as the percentage of bootstrap samples that exhibited greater bias than the observed fairness metric. A model was considered significantly biased for a given fairness metric if the P-value was less than 0.05.

Subsampling analyses for imbalance and shifts

We quantified how sample size imbalance and shifts in class distribution between demographic groups affect performance disparities in a deep learning model for common pathology diagnostic tasks through a series of subsample analyses. We focused on 13 cancer diagnostic tasks from the TCGA cohorts that contained more than four samples for each cancer type across all demographic groups. To ensure sufficient sample sizes across all demographic strata, we employed a stratified random split where the testing partition in each stratum was set to be at least half of the size of the smallest stratum defined by the targeted demographic attribute (e.g., age, gender, or self-reported race). The training dataset was then sampled from the remaining data in each stratum. There was no patient overlap between the training and testing sets.

First, we investigated the impact of sample size balance on model fairness. To achieve this goal, we varied the number of samples for the majority and minority groups while keeping the total dataset size at 200 samples. The balance between the minority and majority groups ranged from being completely balanced (100 majority and 100 minority) to only including data from the majority group (200 majority and 0 minority).

Next, we examined the impact of distribution shifts in diagnostic categories within each demographic group while maintaining overall sample size balance across groups (e.g., 100 majority and 100 minority patients). In the majority group, sample sizes for each cancer type were perfectly balanced (e.g., 50% for each cancer type), while we introduced varying degrees of diagnostic category imbalance in the minority group (from 100:0 to 50:50).

In the sample size balance evaluation experiment, we excluded tasks in which the minority group’s training sample size was less than 100. In the distribution shift experiment, we excluded tasks where the minority group sample size was fewer than 50 across both classes or fewer than 25 within a single class. We excluded the COAD vs. READ classification task from both experiments due to a consistently low AUC. For each experiment, we repeated the training and evaluation process 20 times with different balance ratios, and we used two-sided Spearman correlation tests to examine the relationships between sample size imbalance, category imbalance, and performance disparities between majority and minority groups. The Spearman correlation test assumes ordinal or continuous data, monotonic relationships, paired observations, and independence of observations.

Demographic separability

We evaluated the extent to which pathology image representation contained information related to demographic attributes by conducting supervised machine learning analyses. For cancer types where baseline AI models for cancer diagnosis showed significant demographic bias, we extracted slide-level WSI features using the trained baseline AI models. We then built logistic regression models to classify patients’ race, gender, or age using the extracted features and 4-fold cross-validation. We reported the AUROCs of demographic attribute classification and employed the one-sided DeLong test to determine if the AUROC is significantly greater than 0.5.

For cancer types showing significant distinction between demographic groups in their image features (Delong test P-value <0.05), we further investigated their cellular attributes contributing to demographic separability. For each patient, we selected up to 500 of the most accurately predicted patches by the demographic attribute classification model. Subsequently, we applied CellViT87 to segment and classify the cells into four primary cell types: neoplastic, inflammatory, epithelial, and connective or soft tissue cells. Necrosis was excluded from this analysis, as it was only detected in 0.22% of the selected patches. We reported the density of each cell type and assessed the differences in cell-type distributions across demographic groups using the Mann–Whitney U test. These analyses collectively assess how demographic information is encoded in tissue morphology and identify cellular correlates underlying such signals.

GLM analyses for tissue/genomic factors

We further examined the relationships between performance disparities and tissue or genomic characteristics in our pathology samples. For each WSI, we obtained the percentages of eight tissue types (tissues with lymphocyte infiltration, monocyte infiltration, necrosis, neutrophil infiltration, normal cells, stromal cells, tumor cells, and tumor nuclei) curated by the TCGA study consortium from GDC. In addition, we retrieved mutational profiles for the five genes with the highest mutation prevalence and those related to FDA-approved targeted therapies from cBioPortal. The genomic mutations investigated in our analyses are detailed in Table S8. Our analysis included genomic mutations both related and unrelated to the cancer type in question, allowing us to investigate the influence of potential disease-independent factors. We investigated germline MET mutation in KIRP, which is characteristic of the familial KIRP variant,56 but found no significant performance disparity due to limited sample size88,89,90 (N = 3 in TCGA cohort; Mann-Whitney U-test P-value = 0.96).

We focused on cancer type or subtype classification tasks where the baseline models exhibited significant performance disparity. We excluded tissue percentage analyses for FFPE samples because fewer than 1% of them contain such information. Among the classification scenarios showing significant performance bias in the baseline model, we employed generalized linear models to quantify tasks that showed both 1) significant associations between each tissue or genomic factor and the demographic attribute of interest and 2) significant associations between each tissue or genomic factor and the model’s predictive errors. For this purpose, we employed generalized linear models (GLM), with the assumptions of independent observations, linearity in parameters, binomial distribution for demographic attributes and genomic mutations, Gaussian distributions for predictive error and tissue percentages, and no perfect multicollinearity. For variables modeled with a Gaussian distribution, we employed an identity link function, whereas for those modeled with a binomial distribution, a logit link function was used. We used two-sided P-values from Wald tests to determine statistical significance.

Fairness analysis accounting for technical factors

We conducted stratified bootstrap tests to evaluate whether patient stage or technical heterogeneity contributes to demographic disparities in model performance. Factors examined included patient stage as well as technical factors such as tissue procurement sites, scanner types, and scanning magnification. For each factor, the dataset was stratified to preserve the internal structure of that variable. Within each stratum, 1000 random samples via bootstrapped resampling were used to generate a null distribution of fairness metrics under the constraint that the stratification factor remained fixed. The observed fairness metrics were then compared against their respective null distributions to determine whether the observed disparities could be explained by patient stage or technical variability alone.

Ablation studies

We conducted three ablation analyses to demonstrate the efficacy of our model design.

  • 1)

    FAIR-Path versus Supervised Contrastive Learning: While supervised contrastive learning has shown promise in improving generalization, its ability to mitigate demographic bias remains unquantified. To investigate this, we compared the racial fairness of FAIR-Path to an otherwise identical model trained without the non-discrimination objective, using only the supervised contrastive loss. We evaluated and compared these models across all cancer type and subtype classification tasks using FFPE cohorts.

  • 2)

    Correcting Hospital-Related Bias: To investigate whether the observed demographic bias can be mitigated by addressing the representation disparity between hospital sites, we trained a variant of FAIR-Path that aims to correct the hospital-related bias. For this purpose, the hospital sites are treated as demographic attributes during the fairness-aware contrastive learning and the downstream training process. Hospital sites with two or fewer samples were excluded from the analysis. We selected 12 tasks in which the baseline models exhibited demographic biases. After correction for hospital sites, we then evaluated the bias related to the demographic attributes (race, gender, and age) on these site-corrected classification models.

  • 3)

    Impact of Training Sample Size: To investigate the robustness of our fairness-aware contrastive learning framework under varying amounts of training data, we conducted an experiment to assess how the amount of training data affects its fairness performance. Using the LUAD vs. LUSC classification task as a case study, we trained the FAIR-Path model on five different proportions of the training data (20%, 40%, 60%, 80%, and 100%) in the contrastive learning stage. Subsequently, all samples were used to fine-tune the downstream classification task.

Quantification and statistical analysis

All statistical analyses were performed in Python 3.8 (NGC PyTorch 22.03 base docker image uses Python 3.8) using NumPy 1.22.4 and scikit-learn 1.1.2. Mann–Whitney U tests and Spearman rank correlation tests were conducted with SciPy 1.16.1. Generalized linear model analyses were conducted using statsmodels 0.14.2. The DeLong tests were conducted using the statistics.delong function from the WORC 3.7.0 package. Bootstrap significance tests were performed via our Python implementation provided in the repository (see key resources table) with results computed on pooled per-patient predictions. Unless otherwise specified, n denotes the number of patients (per-patient predictions aggregated across slides). Statistical significance was assessed at α = 0.05.

Published: December 16, 2025

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.xcrm.2025.102527.

Contributor Information

Jung-Hsien Chiang, Email: jchiang@mail.ncku.edu.tw.

Kun-Hsing Yu, Email: kun-hsing_yu@hms.harvard.edu.

Supplemental information

Document S1. Figures S1–S6 and Tables S1 and S3–S8
mmc1.pdf (876.1KB, pdf)
Table S2. Detailed experimental results for standard AI models and FAIR-Path, related to Figures 2 and 4–7
mmc2.xlsx (213.1KB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (10.1MB, pdf)

References

  • 1.Siegel R.L., Miller K.D., Jemal A. Cancer statistics, 2020. CA Cancer J. Clin. 2020;70:7–30. doi: 10.3322/caac.21590. [DOI] [PubMed] [Google Scholar]
  • 2.Lu M.Y., Williamson D.F.K., Chen T.Y., Chen R.J., Barbieri M., Mahmood F. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat. Biomed. Eng. 2021;5:555–570. doi: 10.1038/s41551-020-00682-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Wang X., Zhao J., Marostica E., Yuan W., Jin J., Zhang J., Li R., Tang H., Wang K., Li Y., et al. A pathology foundation model for cancer diagnosis and prognosis prediction. Nature. 2024;634:970–978. doi: 10.1038/s41586-024-07894-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Marostica E., Barber R., Denize T., Kohane I.S., Signoretti S., Golden J.A., Yu K.H. Development of a Histopathology Informatics Pipeline for Classification and Prediction of Clinical Outcomes in Subtypes of Renal Cell Carcinoma. Clin. Cancer Res. 2021;27:2868–2878. doi: 10.1158/1078-0432.Ccr-20-4119. [DOI] [PubMed] [Google Scholar]
  • 5.Chen R.J., Ding T., Lu M.Y., Williamson D.F.K., Jaume G., Song A.H., Chen B., Zhang A., Shao D., Shaban M., et al. Towards a general-purpose foundation model for computational pathology. Nat. Med. 2024;30:850–862. doi: 10.1038/s41591-024-02857-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Ektefaie Y., Yuan W., Dillon D.A., Lin N.U., Golden J.A., Kohane I.S., Yu K.H. Integrative multiomics-histopathology analysis for breast cancer classification. npj Breast Cancer. 2021;7:147. doi: 10.1038/s41523-021-00357-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Nasrallah M.P., Zhao J., Tsai C.C., Meredith D., Marostica E., Ligon K.L., Golden J.A., Yu K.H. Machine learning for cryosection pathology predicts the 2021 WHO classification of glioma. Med. 2023;4:526–540.e4. doi: 10.1016/j.medj.2023.06.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Tsai P.C., Lee T.H., Kuo K.C., Su F.Y., Lee T.L.M., Marostica E., Ugai T., Zhao M., Lau M.C., Väyrynen J.P., et al. Histopathology images predict multi-omics aberrations and prognoses in colorectal cancer patients. Nat. Commun. 2023;14:2102. doi: 10.1038/s41467-023-37179-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Yu K.H., Berry G.J., Rubin D.L., Ré C., Altman R.B., Snyder M. Association of Omics Features with Histopathology Patterns in Lung Adenocarcinoma. Cell Syst. 2017;5:620–627.e3. doi: 10.1016/j.cels.2017.10.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Yu K.H., Wang F., Berry G.J., Ré C., Altman R.B., Snyder M., Kohane I.S. Classifying non-small cell lung cancer types and transcriptomic subtypes using convolutional neural networks. J. Am. Med. Inform. Assoc. 2020;27:757–769. doi: 10.1093/jamia/ocz230. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Yu K.H., Hu V., Wang F., Matulonis U.A., Mutter G.L., Golden J.A., Kohane I.S. Deciphering serous ovarian carcinoma histopathology and platinum response by convolutional neural networks. BMC Med. 2020;18:236. doi: 10.1186/s12916-020-01684-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Yu K.H., Zhang C., Berry G.J., Altman R.B., Ré C., Rubin D.L., Snyder M. Predicting non-small cell lung cancer prognosis by fully automated microscopic pathology image features. Nat. Commun. 2016;7 doi: 10.1038/ncomms12474. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Fogel A.L., Kvedar J.C. Artificial intelligence powers digital medicine. npj Digit. Med. 2018;1 doi: 10.1038/s41746-017-0012-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Alowais S.A., Alghamdi S.S., Alsuhebany N., Alqahtani T., Alshaya A.I., Almohareb S.N., Aldairem A., Alrashed M., bin Saleh K., Badreldin H.A., et al. Revolutionizing healthcare: the role of artificial intelligence in clinical practice. BMC Med. Educ. 2023;23:689. doi: 10.1186/s12909-023-04698-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Ozyoruk K.B., Can S., Darbaz B., Başak K., Demir D., Gokceler G.I., Serin G., Hacisalihoglu U.P., Kurtuluş E., Lu M.Y., et al. A deep-learning model for transforming the style of tissue images from cryosectioned to formalin-fixed and paraffin-embedded. Nat. Biomed. Eng. 2022;6:1407–1419. doi: 10.1038/s41551-022-00952-9. [DOI] [PubMed] [Google Scholar]
  • 16.Yu K.H., Healey E., Leong T.Y., Kohane I.S., Manrai A.K. Medical Artificial Intelligence and Human Values. N. Engl. J. Med. 2024;390:1895–1904. doi: 10.1056/NEJMra2214183. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Tomczak K., Czerwińska P., Wiznerowicz M. The Cancer Genome Atlas (TCGA): an immeasurable source of knowledge. Contemp. Oncol. 2015;19:A68–A77. doi: 10.5114/wo.2014.47136. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Edwards N.J., Oberti M., Thangudu R.R., Cai S., McGarvey P.B., Jacob S., Madhavan S., Ketchum K.A. The CPTAC Data Portal: A Resource for Cancer Proteomics Research. J. Proteome Res. 2015;14:2707–2713. doi: 10.1021/pr501254j. [DOI] [PubMed] [Google Scholar]
  • 19.Spratt D.E. Are we inadvertently widening the disparity gap in pursuit of precision oncology? Br. J. Cancer. 2018;119:783–784. doi: 10.1038/s41416-018-0223-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Chen R.J., Wang J.J., Williamson D.F.K., Chen T.Y., Lipkova J., Lu M.Y., Sahai S., Mahmood F. Algorithmic fairness in artificial intelligence for medicine and healthcare. Nat. Biomed. Eng. 2023;7:719–742. doi: 10.1038/s41551-023-01056-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Ashktorab H., Kupfer S.S., Brim H., Carethers J.M. Racial Disparity in Gastrointestinal Cancer Risk. Gastroenterology. 2017;153:910–923. doi: 10.1053/j.gastro.2017.08.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Özdemir B.C., Dotto G.P. Racial Differences in Cancer Susceptibility and Survival: More Than the Color of the Skin? Trends Cancer. 2017;3:181–197. doi: 10.1016/j.trecan.2017.02.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Du X.L., Lin C.C., Johnson N.J., Altekruse S. Effects of individual-level socioeconomic factors on racial disparities in cancer treatment and survival. Cancer. 2011;117:3242–3251. doi: 10.1002/cncr.25854. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Vaidya A., Chen R.J., Williamson D.F.K., Song A.H., Jaume G., Yang Y., Hartvigsen T., Dyer E.C., Lu M.Y., Lipkova J., et al. Demographic bias in misdiagnosis by computational pathology models. Nat. Med. 2024;30:1174–1190. doi: 10.1038/s41591-024-02885-z. [DOI] [PubMed] [Google Scholar]
  • 25.Howard F.M., Dolezal J., Kochanny S., Schulte J., Chen H., Heij L., Huo D., Nanda R., Olopade O.I., Kather J.N., et al. The impact of site-specific digital histology signatures on deep learning model accuracy and bias. Nat. Commun. 2021;12:4423. doi: 10.1038/s41467-021-24698-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Geirhos R., Jacobsen J.-H., Michaelis C., Zemel R., Brendel W., Bethge M., Wichmann F.A. Shortcut learning in deep neural networks. Nat. Mach. Intell. 2020;2:665–673. doi: 10.1038/s42256-020-00257-z. [DOI] [Google Scholar]
  • 27.Yu K.H., Beam A.L., Kohane I.S. Artificial intelligence in healthcare. Nat. Biomed. Eng. 2018;2:719–731. doi: 10.1038/s41551-018-0305-z. [DOI] [PubMed] [Google Scholar]
  • 28.Administration, F.a.D. (2019). Proposed regulatory framework for modifications to artificial intelligence/machine learning (AI/ML)-based software as a medical device (SaMD).
  • 29.Kamiran F., Calders T. Data preprocessing techniques for classification without discrimination. Knowl. Inf. Syst. 2012;33:1–33. doi: 10.1007/s10115-011-0463-8. [DOI] [Google Scholar]
  • 30.Franchet C., Schwob R., Bataillon G., Syrykh C., Péricart S., Frenois F.-X., Penault-Llorca F., Lacroix-Triki M., Arnould L., Lemonnier J., et al. Bias reduction using combined stain normalization and augmentation for AI-based classification of histological images. Comput. Biol. Med. 2024;171 doi: 10.1016/j.compbiomed.2024.108130. [DOI] [PubMed] [Google Scholar]
  • 31.Ktena I., Wiles O., Albuquerque I., Rebuffi S.-A., Tanno R., Roy A.G., Azizi S., Belgrave D., Kohli P., Cemgil T., et al. Generative models improve fairness of medical classifiers under distribution shifts. Nat. Med. 2024;30:1166–1173. doi: 10.1038/s41591-024-02838-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Shen A., Han X., Cohn T., Baldwin T., Frermann L. Contrastive Learning for Fair Representations. arXiv. 2021 doi: 10.48550/arXiv.2109.10645. Preprint at. [DOI] [Google Scholar]
  • 33.Tartaglione E., Barbano C.A., Grangetto M. End: Entangling and disentangling deep representations for bias correction. arXiv. 2021 doi: 10.48550/arXiv.2103.02023. Preprint at. [DOI] [Google Scholar]
  • 34.Pham T.-H., Zhang X., Zhang P. Fairness and Accuracy Under Domain Generalization. arXiv. 2023 doi: 10.48550/arXiv.2301.13323. Preprint at. [DOI] [Google Scholar]
  • 35.Wu Y., Zeng D., Xu X., Shi Y., Hu J. FairPrune: Achieving Fairness Through Pruning for Dermatological Disease Diagnosis. arXiv. 2022 doi: 10.48550/arXiv.2203.02110. Preprint at. [DOI] [Google Scholar]
  • 36.Zhang B.H., Lemoine B., Mitchell M. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society. Association for Computing Machinery; 2018. Mitigating Unwanted Biases with Adversarial Learning. [Google Scholar]
  • 37.Madras D., Creager E., Pitassi T., Zemel R. In: Proceedings of Machine Learning Research. Jennifer D., Andreas K., editors. PMLR; 2018. Learning Adversarially Fair and Transferable Representations; pp. 3384-–3393. [Google Scholar]
  • 38.Zong Y., Yang Y., Hospedales T. MEDFAIR: Benchmarking Fairness for Medical Imaging. arXiv. 2022 doi: 10.48550/arXiv.2210.01725. Preprint at. [DOI] [Google Scholar]
  • 39.Hosseini S.M., Sikaroudi M., Babaie M., Tizhoosh H.R. Proportionally fair hospital collaborations in federated learning of histopathology images. IEEE Trans. Med. Imaging. 2023;42:1982–1995. doi: 10.1109/TMI.2023.3234450. [DOI] [PubMed] [Google Scholar]
  • 40.Zhu C.S., Pinsky P.F., Kramer B.S., Prorok P.C., Purdue M.P., Berg C.D., Gohagan J.K. The prostate, lung, colorectal, and ovarian cancer screening trial and its associated research resource. J. Natl. Cancer Inst. 2013;105:1684–1693. doi: 10.1093/jnci/djt281. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Roetzer-Pejrimovsky T., Moser A.-C., Atli B., Vogel C.C., Mercea P.A., Prihoda R., Gelpi E., Haberler C., Höftberger R., Hainfellner J.A., et al. the Digital Brain tumour atlas, an open histopathology resource. Sci. Data. 2022;9:55. doi: 10.1038/s41597-022-01157-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Zemel R., Wu Y., Swersky K., Pitassi T., Dwork C. In: Proceedings of the 30th International Conference on Machine Learning. Sanjoy D., David M., editors. PMLR; 2013. Learning Fair Representations. [Google Scholar]
  • 43.Krasanakis E., Spyromitros-Xioufis E., Papadopoulos S., Kompatsiaris Y. Proceedings of the 2018 world wide web conference. 2018. Adaptive sensitive reweighting to mitigate bias in fairness-aware classification; pp. 853–862. [Google Scholar]
  • 44.Iosifidis V., Ntoutsi E. Proceedings of the 28th ACM international conference on information and knowledge management. 2019. Adafair: Cumulative fairness adaptive boosting; pp. 781–790. [Google Scholar]
  • 45.Chawla N.V., Bowyer K.W., Hall L.O., Kegelmeyer W.P. SMOTE: synthetic minority over-sampling technique. J. Artif. Intell. Res. 2002;16:321–357. doi: 10.1613/jair.953. [DOI] [Google Scholar]
  • 46.Bickell N.A., Wang J.J., Oluwole S., Schrag D., Godfrey H., Hiotis K., Méndez J.E., Guth A.A. Missed opportunities: racial disparities in adjuvant breast cancer treatment. J. Clin. Oncol. 2006;24:1357–1362. doi: 10.1200/JCO.2005.04.5799. [DOI] [PubMed] [Google Scholar]
  • 47.Bouchardy C., Rapiti E., Blagojevic S., Vlastos A.T., Vlastos G. Older female cancer patients: importance, causes, and consequences of undertreatment. J. Clin. Oncol. 2007;25:1858–1869. doi: 10.1200/JCO.2006.10.4208. [DOI] [PubMed] [Google Scholar]
  • 48.Duma N., Vera Aguilera J., Paludo J., Haddox C.L., Gonzalez Velez M., Wang Y., Leventakos K., Hubbard J.M., Mansfield A.S., Go R.S., Adjei A.A. Representation of Minorities and Women in Oncology Clinical Trials: Review of the Past 14 Years. J. Oncol. Pract. 2018;14:e1–e10. doi: 10.1200/JOP.2017.025288. [DOI] [PubMed] [Google Scholar]
  • 49.Xu H., Usuyama N., Bagga J., Zhang S., Rao R., Naumann T., Wong C., Gero Z., González J., Gu Y., et al. A whole-slide foundation model for digital pathology from real-world data. Nature. 2024;630:181–188. doi: 10.1038/s41586-024-07441-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Fane M., Weeraratna A.T. How the ageing microenvironment influences tumour progression. Nat. Rev. Cancer. 2020;20:89–106. doi: 10.1038/s41568-019-0222-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Sprenger C.C., Plymate S.R., Reed M.J. Aging-related alterations in the extracellular matrix modulate the microenvironment and influence tumor progression. Int. J. Cancer. 2010;127:2739–2748. doi: 10.1002/ijc.25615. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Omilian A.R., Cannioto R., Mendicino L., Stein L., Bshara W., Qin B., Bandera E.V., Zeinomar N., Abrams S.I., Hong C.C., et al. CD163(+) macrophages in the triple-negative breast tumor microenvironment are associated with improved survival in the Women's Circle of Health Study and the Women's Circle of Health Follow-Up Study. Breast Cancer Res. 2024;26:75. doi: 10.1186/s13058-024-01831-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Yao S., Cheng T.Y.D., Elkhanany A., Yan L., Omilian A., Abrams S.I., Evans S., Hong C.C., Qi Q., Davis W., et al. Breast Tumor Microenvironment in Black Women: A Distinct Signature of CD8+ T-Cell Exhaustion. J. Natl. Cancer Inst. 2021;113:1036–1043. doi: 10.1093/jnci/djaa215. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Martin D.N., Boersma B.J., Yi M., Reimers M., Howe T.M., Yfantis H.G., Tsai Y.C., Williams E.H., Lee D.H., Stephens R.M., et al. Differences in the Tumor Microenvironment between African-American and European-American Breast Cancer Patients. PLoS One. 2009;4 doi: 10.1371/journal.pone.0004531. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Ullah A., Yasinzai A.Q.K., Daino N., Tareen B., Jogezai Z.H., Sadia H., Jamil N., Baloch G., Karim A., Badini K., et al. Papillary Renal Cell Carcinoma: Demographics, Survival Analysis, Racial Disparities, and Genomic Landscape. J. Kidney Cancer VHL. 2023;10:33–42. doi: 10.15586/jkcvhl.v10i4.294. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Cancer Genome Atlas Research Network. Linehan W.M., Spellman P.T., Ricketts C.J., Creighton C.J., Fei S.S., Davis C., Wheeler D.A., Murray B.A., Schmidt L., et al. Comprehensive Molecular Characterization of Papillary Renal-Cell Carcinoma. N. Engl. J. Med. 2016;374:135–145. doi: 10.1056/NEJMoa1505917. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Fu J.B., Kau T.Y., Severson R.K., Kalemkerian G.P. Lung Cancer in Women: Analysis of the National Surveillance, Epidemiology, and End Results Database. Chest. 2005;127:768–777. doi: 10.1378/chest.127.3.768. [DOI] [PubMed] [Google Scholar]
  • 58.Ostrom Q.T., Cote D.J., Ascha M., Kruchko C., Barnholtz-Sloan J.S. Adult glioma incidence and survival by race or ethnicity in the United States from 2000 to 2014. JAMA Oncol. 2018;4:1254–1262. doi: 10.1001/jamaoncol.2018.1789. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Cristofanilli M., Gonzalez-Angulo A., Sneige N., Kau S.W., Broglio K., Theriault R.L., Valero V., Buzdar A.U., Kuerer H., Buchholz T.A., Hortobagyi G.N. Invasive lobular carcinoma classic type: response to primary chemotherapy and survival outcomes. J. Clin. Oncol. 2005;23:41–48. doi: 10.1200/jco.2005.03.111. [DOI] [PubMed] [Google Scholar]
  • 60.Rawla P., Barsouk A. Epidemiology of gastric cancer: global trends, risk factors and prevention. Prz. Gastroenterol. 2019;14:26–38. doi: 10.5114/pg.2018.80001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Mendiratta G., Ke E., Aziz M., Liarakos D., Tong M., Stites E.C. Cancer gene mutation frequencies for the U.S. population. Nat. Commun. 2021;12:5961. doi: 10.1038/s41467-021-26213-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Cancer Genome Atlas Research Network Comprehensive genomic characterization of squamous cell lung cancers. Nature. 2012;489:519–525. doi: 10.1038/nature11404. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Cancer Genome Atlas Research Network. Brat D.J., Verhaak R.G.W., Aldape K.D., Yung W.K.A., Salama S.R., Cooper L.A.D., Rheinbay E., Miller C.R., Vitucci M., et al. Comprehensive, integrative genomic analysis of diffuse lower-grade gliomas. N. Engl. J. Med. 2015;372:2481–2498. doi: 10.1056/NEJMoa1402121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Whelan K., Dillon M., Strickland K.C., Pothuri B., Bae-Jump V., Borden L.E., Thaker P.H., Haight P., Arend R.C., Ko E., et al. TP53 mutation and abnormal p53 expression in endometrial cancer: Associations with race and outcomes. Gynecol. Oncol. 2023;178:44–53. doi: 10.1016/j.ygyno.2023.09.009. [DOI] [PubMed] [Google Scholar]
  • 65.van Beek E.J.A.H., Hernandez J.M., Goldman D.A., Davis J.L., McLaughlin K., Ripley R.T., Kim T.S., Tang L.H., Hechtman J.F., Zheng J., et al. Rates of TP53 Mutation are Significantly Elevated in African American Patients with Gastric Cancer. Ann. Surg Oncol. 2018;25:2027–2033. doi: 10.1245/s10434-018-6502-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Ashing K.T., Jones V., Bedell F., Phillips T., Erhunmwunsee L. Wolters Kluwer Health; 2022. Calling Attention to the Role of Race-Driven Societal Determinants of Health on Aggressive Tumor Biology: A Focus on Black Americans. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Sanjeevaiah A., Cheedella N., Hester C., Porembka M.R. Gastric cancer: recent molecular classification advances, racial disparity, and management implications. J. Oncol. Pract. 2018;14:217–224. doi: 10.1200/JOP.17.00025. [DOI] [PubMed] [Google Scholar]
  • 68.Haupt S., Haupt Y. Cancer and tumour suppressor p53 encounters at the juncture of sex disparity. Front. Genet. 2021;12 doi: 10.3389/fgene.2021.632719. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Ciriello G., Gatza M.L., Beck A.H., Wilkerson M.D., Rhie S.K., Pastore A., Zhang H., McLellan M., Yau C., Kandoth C., et al. Comprehensive Molecular Portraits of Invasive Lobular Breast Cancer. Cell. 2015;163:506–519. doi: 10.1016/j.cell.2015.09.033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Cho S.Y., Park J.W., Liu Y., Park Y.S., Kim J.H., Yang H., Um H., Ko W.R., Lee B.I., Kwon S.Y., et al. Sporadic Early-Onset Diffuse Gastric Cancers Have High Frequency of Somatic CDH1 Alterations, but Low Frequency of Somatic RHOA Mutations Compared With Late-Onset Cancers. Gastroenterology. 2017;153:536–549.e26. doi: 10.1053/j.gastro.2017.05.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Adib E., El Zarif T., Nassar A.H., Akl E.W., Abou Alaiwi S., Mouhieddine T.H., Esplin E.D., Hatchell K., Nielsen S.M., Rana H.Q., et al. CDH1 germline variants are enriched in patients with colorectal cancer, gastric cancer, and breast cancer. Br. J. Cancer. 2022;126:797–803. doi: 10.1038/s41416-021-01673-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Kather J.N., Heij L.R., Grabsch H.I., Loeffler C., Echle A., Muti H.S., Krause J., Niehues J.M., Sommer K.A.J., Bankhead P., et al. Pan-cancer image-based detection of clinically actionable genetic alterations. Nat. Cancer. 2020;1:789–799. doi: 10.1038/s43018-020-0087-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Shabihkhani M., Lucey G.M., Wei B., Mareninov S., Lou J.J., Vinters H.V., Singer E.J., Cloughesy T.F., Yong W.H. The procurement, storage, and quality assurance of frozen blood and tissue biospecimens in pathology, biorepository, and biobank settings. Clin. Biochem. 2014;47:258–266. doi: 10.1016/j.clinbiochem.2014.01.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Wang D., Shelhamer E., Liu S., Olshausen B., Darrell T. Tent: Fully test-time adaptation by entropy minimization. arXiv. 2020 doi: 10.48550/arXiv.2006.10726. Preprint at. [DOI] [Google Scholar]
  • 75.Ho M.M., Dubey S., Chong Y., Knudsen B., Tasdizen T. IEEE; 2025. F2FLDM: Latent Diffusion Models with Histopathology Pre-trained Embeddings for Unpaired Frozen Section to FFPE Translation; pp. 4382–4391. [Google Scholar]
  • 76.Falahkheirkhah K., Guo T., Hwang M., Tamboli P., Wood C.G., Karam J.A., Sircar K., Bhargava R. A generative adversarial approach to facilitate archival-quality histopathologic diagnoses from frozen tissue sections. Lab. Invest. 2022;102:554–559. doi: 10.1038/s41374-021-00718-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Nakhli R., Rich K., Zhang A., Darbandsari A., Shenasa E., Hadjifaradji A., Thiessen S., Milne K., Jones S.J.M., McAlpine J.N., et al. VOLTA: an enVironment-aware cOntrastive ceLl represenTation leArning for histopathology. Nat. Commun. 2024;15:3942. doi: 10.1038/s41467-024-48062-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Hankivsky O., Doyal L., Einstein G., Kelly U., Shim J., Weber L., Repta R. The odd couple: using biomedical and intersectional approaches to address health inequities. Glob. Health Action. 2017;10 doi: 10.1080/16549716.2017.1326686. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Turan J.M., Elafros M.A., Logie C.H., Banik S., Turan B., Crockett K.B., Pescosolido B., Murray S.M. Challenges and opportunities in examining and addressing intersectional stigma and health. BMC Med. 2019;17:7. doi: 10.1186/s12916-018-1246-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Duffy G., Clarke S.L., Christensen M., He B., Yuan N., Cheng S., Ouyang D. Confounders mediate AI prediction of demographics in medical imaging. npj Digit. Med. 2022;5:188. doi: 10.1038/s41746-022-00720-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Wang X., Yang S., Zhang J., Wang M., Zhang J., Yang W., Huang J., Han X. Transformer-based unsupervised contrastive learning for histopathological image classification. Med. Image Anal. 2022;81 doi: 10.1016/j.media.2022.102559. [DOI] [PubMed] [Google Scholar]
  • 82.Kang M., Song H., Park S., Yoo D., Pereira S. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023. Benchmarking self-supervised learning on diverse pathology datasets; pp. 3344–3354. [Google Scholar]
  • 83.He K., Zhang X., Ren S., Sun J. Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. Deep residual learning for image recognition; pp. 770–778. [Google Scholar]
  • 84.Khosla P., Teterwak P., Wang C., Sarna A., Tian Y., Isola P., Maschinot A., Liu C., Krishnan D. In: Supervised contrastive learning. Larochelle H., Ranzato M., Hadsell R., Balcan M.F., Lin H., editors. Curran Associates, Inc.; 2020. pp. 18661–18673. [Google Scholar]
  • 85.Nair V., Hinton G.E. Proceedings of the 27th International Conference on Machine Learning (ICML-10) 2010. Rectified linear units improve restricted boltzmann machines; pp. 807–814. [Google Scholar]
  • 86.MacKinnon J.G. Handbook of Computational Econometrics. Wiley; 2009. Bootstrap hypothesis testing; pp. 183–213. [DOI] [Google Scholar]
  • 87.Hörst F., Rempe M., Heine L., Seibold C., Keyl J., Baldini G., Ugurel S., Siveke J., Grünwald B., Egger J., Kleesiek J. Cellvit: Vision transformers for precise cell segmentation and classification. Med. Image Anal. 2024;94 doi: 10.1016/j.media.2024.103143. [DOI] [PubMed] [Google Scholar]
  • 88.Verine J., Pluvinage A., Bousquet G., Lehmann-Che J., de Bazelaire C., Soufir N., Mongiat-Artus P. Hereditary Renal Cancer Syndromes: An Update of a Systematic Review. Eur. Urol. 2010;58:701–710. doi: 10.1016/j.eururo.2010.08.031. [DOI] [PubMed] [Google Scholar]
  • 89.Orphanet Hereditary papillary renal cell carcinoma. 2022. https://www.orpha.net/en/disease/detail/47044
  • 90.Jacoba I.M., Lu Z. Hereditary papillary renal cell carcinoma. Semin. Diagn. Pathol. 2024;41:28–31. doi: 10.1053/j.semdp.2023.12.002. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1–S6 and Tables S1 and S3–S8
mmc1.pdf (876.1KB, pdf)
Table S2. Detailed experimental results for standard AI models and FAIR-Path, related to Figures 2 and 4–7
mmc2.xlsx (213.1KB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (10.1MB, pdf)

Data Availability Statement

  • The digital pathology slides, summarized clinical data, and biospecimen information for The Cancer Genome Atlas (TCGA) and Clinical Proteomic Tumor Analysis Consortium (CPTAC) cohorts are available through the National Cancer Institute Genomic Data Commons (https://portal.gdc.cancer.gov/). Summarized molecular data for these cohorts are available from cBioPortal (https://www.cbioportal.org/). Data from the Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trials can be obtained from the NCI Cancer Data Access System (https://cdas.cancer.gov/plco/). Data from the Medical University of Vienna are available through the EBRAINS portal (https://www.ebrains.eu/). Datasets from Mayo Clinic, Dana-Farber Cancer Institute (DFCI), University of Pennsylvania (UPenn), and Mass General Brigham (MGB) are not publicly available due to patient privacy concerns and data use agreement requirements. To request access to these data, please contact the lead contact, Kun-Hsing Yu (kun-hsing_yu@hms.harvard.edu). The requestor must describe the objectives of the research project for which the data will be used. We aim to forward all requests to the managers of these institutional datasets within 2 weeks, and these requests will be evaluated according to institutional policies. Data access will be considered for research purposes and non-commercial use only. In order to ensure patient privacy, access to personally identifiable information or sensitive clinical information will not be provided, and requests for data access must rigorously adhere to the consent agreements established with study participants.

  • All original code has been deposited at GitHub and is publicly available as of the date of publication. The code for our model and data analyses can be found at https://github.com/hms-dbmi/fairpath.

  • Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.


Articles from Cell Reports Medicine are provided here courtesy of Elsevier

RESOURCES