Abstract
The batch effect is a nonbiological variation that arises from technical differences across different batches of data during the data generation process for acquisition-related reasons, such as collection of images at different sites or using different scanners. This phenomenon can affect the robustness and generalizability of computational pathology- or radiology-based cancer diagnostic models, especially in multi-center studies. To address this issue, we developed an open-source platform, Batch Effect Explorer (BEEx), that is designed to qualitatively and quantitatively determine whether batch effects exist among medical image datasets from different sites. A suite of tools was incorporated into BEEx that provide visualization and quantitative metrics based on intensity, gradient, and texture features to allow users to determine whether there are any image variables or combinations of variables that can distinguish datasets from different sites in an unsupervised manner. BEEx was designed to support various medical imaging techniques, including microscopy and radiology. Four use cases clearly demonstrated the ability of BEEx to identify batch effects and validated the effectiveness of rectification methods for batch effect reduction. Overall, BEEx is a scalable and versatile framework designed to read, process, and analyze a wide range of medical images to facilitate the identification and mitigation of batch effects, which can enhance the reliability and validity of image-based studies.
Introduction
Batch effects are nonbiological variations arising from technical differences across different batches of data during the data generation process, such as in images curated from different sites or obtained using different scanners. These acquisition-related factors impact the imaging data in a manner that impedes the ability of machine learning algorithms to generalize to data from new sites. In the literature, researchers showed that they can use machine learning classifiers to predict the originating sites of data (1), which indicates that there is abundant site-related batch effect information embedded in the image data. The causes of batch effects are varied. In digital pathology, batch effects can arise from the use of different scanners, the method of image preprocessing, differences in operator handling, and the manner in which the slides are prepared. In radiology, batch effects can arise from scanner and protocol variability, acquisition parameters, and image preprocessing methods.
Batch effects pose substantial obstacles in the field of medical image analysis, particularly with regard to machine learning applications. When algorithms are trained on one batch or cohort, they may not perform well when applied to another and may lack robustness or generalizability. Previously, researchers have demonstrated the extent and impact of batch effects in various contexts. Kothari et al. demonstrated that batch effects can change morphological image features and decrease the prediction performance in multi-batch studies (2). Stopsack et al. demonstrated that in half of the 20 protein biomarkers obtained from 1448 men in two nationwide cohort studies, more than 10% of the biomarker variance was attributable to batch differences (3). Their study proved that the differences between different batches could be attributed more to batch effects than to different patient or tumor characteristics. An increasing number of researchers have focused on multi-center studies (4,5) to validate the robustness and generalizability of artificial intelligence (AI) models. Therefore, there is an urgent need for efficient tools that can assess batch effects in multi-center settings.
In the context of genomic analysis, Zhu et al. presented BatchServer, an open-source R/Shiny-based platform for batch effect analysis on proteomic and transcriptomic data (6). However, the platform is incompatible with medical imaging. Stopsack et al. (3) demonstrated the influence of batch effects on tissue microarrays (TMAs) and employed diverse algorithms to alleviate them. A quantitative quality control tool, MRQy (7), was specifically designed for magnetic resonance imaging (MRI) data quality control, and can quantify site- or scanner-specific variations in image resolution or contrast as well as imaging artifacts such as noise or inhomogeneity. Similarly, for histological images, HistoQC (8), a tool for rapid quality control, has been proposed to delineate artifacts and identify cohort-level outliers. Chen et al. utilized HistoQC-generated image variables to train a random forest to classify the originating sites of whole slide images (WSIs), and found batch effects in at least three laboratories (9). However, to the best of our knowledge, no open-source, pre-analytic platform has been specifically designed to explore image-based batch effects on medical images obtained from multiple centers. Therefore, to enable researchers to perform batch effect analysis prior to the validation of AI models in a multi-center setting, we present the open-source Batch Effect Explorer (BEEx), which is implemented via Python and is compatible with a wide range of types of medical images.
Materials and Methods
Image Data for Multi-center Studies
To prove that modern AI diagnostic and prognostic models are robust and generalizable, many studies are currently testing these models on multi-center datasets. In addition, researchers are not only focusing on a single modality, but also combining information from different modalities. To demonstrate the effectiveness of BEEx for analyzing the batch effect among different datasets and modalities, we prepared four multi-center case studies as shown in Figure 1(A), each of which focuses on a specific type of medical image: WSI, TMA, MRI, and computed tomography (CT). These image types are widely used in pathological and radiological studies and play a crucial role in clinical diagnostics, treatment planning, and biomedical research.
Figure 1:

Overview of the Batch Effect Explorer (BEEx) Platform. (A) Four case studies that include TMA, WSI, MR, and CT Images. (B) Illustration of the interpretative flowchart of BEEx platform including different analyzers with different purpose: (a) Image Overviewer directly examines batch effects by visual appearance. (b) Distribution Analysis with UMAP investigates batch effect based site-specific clusters, (c) Hierarchical Clustering visualizes group of features driving the clusters using a heatmap, (d) Distribution Analysis with Violin Plots helps reveal and trace features contributing to batch effect, (e) PVCA numerically partitions the variance in data represented by batch and other clinical related factors, and (f) BES measures the degree of batch effects cross cohorts.
Case Study-WSI:
This case study examined four distinct hepatocellular carcinoma (HCC) WSI cohorts from different medical institutions. These cohorts included 162 images of 156 patients from Guangdong Provincial People’s Hospital (GDPH), 290 images of 274 patients from The Cancer Genome Atlas (TCGA), 190 images of 182 patients from The First Affiliated Hospital of Kunming Medical University (FAHKU), and 167 images of 166 patients from The First Affiliated Hospital Zhejiang University School of Medicine (FAHZU).
Case Study-VTMA:
This case study involved three virtual TMA cohorts consisting of 293 patients with breast cancer and 293 images: VTMA-1 (102 patients and 102 images), VTMA-2 (94 patients and 94 images), and VTMA-3 (97 patients and 97 images). These virtual TMA spots were digitally extracted from the tumor regions of three resected tissue WSI cohorts: Shandong University Qilu Hospital, Qingdao University Hospital, and the Second Hospital of Shandong University. Note that all histological sections were removed from preserved tissues from different medical institutes. The recut tissues were processed, stained in the different laboratories while using the same reagents and antibodies, and subsequently scanned using the same scanner.
Case Study-MRI-A and -B:
This case study focused on prostate cancer (PCa) MRIs and drew data from three distinct institutes: Shanghai Sixth People’s Hospital (SSPH), which contributed 100 images from 100 patients; Renming Hospital of Wuhan University (RHWU), which provided 100 images from 100 patients; and Yichang Central People’s Hospital (YCPH), which supplied 77 images from 77 patients. To determine whether BEEx can detect the reduction in batch effects among datasets processed using batch effect rectification methods and evaluate the downstream task performance after batch effect rectification (10), we investigated Case Study-MRI-B based on Case Study-MRI-A with batch effect rectification applied.
Case Study-CT:
This was a control case study with three pseudo-cohorts: GDPH-1, GDPH-2, and GDPH-3. These pseudo-cohorts were artificially created by randomly selecting cases from the same abdominal CT urography cohort from Guangdong Provincial People’s Hospital. In this case study, each pseudo-cohort comprised 100 images from 100 patients.
Ethics Statement
This retrospective study was conducted in accordance with the ethical guidelines of the Declaration of Helsinki. Ethical approvals were obtained from the following institutional review boards: GDPH (KY2023-357-02), FAHKU (KYLX2024-171), and FAHZU (KY2024-0376) for the WSI datasets; the School of Basic Medical Sciences of Shandong University (ECSBMSSDU2023-1-52) for the VTMA datasets; SSPH (2022-KY-073(K)) for the MRI datasets; and GDPH (KY2023-146-01) for the CT dataset. All institutional review boards waived the requirement for informed consent from patients due to the retrospective nature of the study.
Clinical Trial Number
This is a retrospective study and does not require a clinical trial number.
Batch Effect Explorer
BEEx platform includes three interconnected modules: I) Preprocessor: This module is responsible for loading and preprocessing images, segmentation masks, and clinical data, and preparing them for subsequent analysis. II) Feature Extractor: This module focuses on extracting image intensity-, gradient-, and texture-based features from images for downstream analyses. III) Analyzer: This module provides batch effect assessment and visualization. It also includes a suite of tools for various analyses as shown in Figure 1(B), including Image Overviewer, Distribution Analysis with Violin Plots (11), Uniform Manifold Approximation and Projection (UMAP, RRID:SCR_018217) (12), Hierarchical Clustering (RRID:SCR_014673) (13), Principal Variation Component Analysis (PVCA, RRID:SCR_001356)) (14), and the proposed Batch Effect Score (BES).
Preprocessor
The preprocessor helps prepare and process data for the feature extractor and analyzer. BEEx supports a wide range of image types and formats for medical images, e.g. NIfTI (.nii.gz), DICOM (.dcm), portable network graphics (.png), and OpenSlide (.svs). In addition, BEEx provides helper scripts to assist users in converting DICOM files into the NIfTI format. Pre-segmented region of interest (ROI) masks and clinical data are optional during preprocessing. If ROI masks are specified, BEEx analyzes the batch effects only within the areas of interest indicated by the ROI masks. Clinical data, if available, are used to perform PVCA.
Feature Extractor
In this module, BEEx extracts various intensity-, gradient-, and texture-based features from the processed images. The extraction of these image features was inspired by two quality check toolkits MRQy (RRID:SCR_025779) (7) and HistoQC (RRID:SCR_025780) (8). For digital pathology images, BEEx extracts (n=36) features, including color histograms, brightness, and contrast. For radiology images, BEEx extracts at most (n=27) features, including the signal-to-noise ratio, mean of the foreground intensity, and entropy focus criterion. A detailed description of each feature is provided in Supplementary Tables S1 and S2.
BEEx Analyzers
BEEx employs six analyzers to visualize and evaluate the batch effects from different perspectives: Image Overviewer, Distribution Analysis with UMAP, Hierarchical Clustering, Distribution Analysis with Violin Plots, PVCA, and the proposed Batch Effect Score. Specifically, Image Overviewer and Distribution Analysis with UMAP evaluate the existence of batch effect through direct visual examination. Hierachical Clustering and Distribution Analysis with Violin Plots explore batch effects on feature level. PVCA and BES quantitively measure the degree of batch effect.
Image Overviewer directly examines batch effects based on visual appearance. It resizes raw images and arranges them into one single image plot such that, the visual differences and characteristics of the image (e.g., color, style, and shape) from different cohorts become more distinct and easier to examine. A flowchart of Image Overviewer is shown in Figure S1.
Distribution Analysis with UMAP projects image features from different cohorts into a two-dimensional space for better batch effect examination (12). By reducing the dimensionality of the data, we can visualize complex patterns and structures more easily and check for site-specific clusters. If a batch effect is present, distinguishable clusters are present in the plot. The pseudo code of Distribution Analysis with UMAP is provided in Algorithm S1.
Hierarchical Clustering also known as clustergram further interprets the clustering results of Distribution Analysis with UMAP, it creates a hierarchy of clusters and visualizes the features driving the clusters using a heatmap(13). The resulting clusters are visualized using a heatmap with color gradients to represent the numerical values of the image features, which allows us to visualize the relationships, similarities, and differences between different image samples. If a batch effect exists, the samples in the analyzer plot are primarily grouped based on their cohort. The pseudo code for hierarchical clustering is given in Algorithm S2.
Distribution Analysis with Violin Plots illustrates the distribution of image features calculated by the Feature Extractor for different cohorts, which helps to reveal and trace the potential batch effect of different cohorts (11). That is, to determine which or what kinds of features contribute to batch effects, this analyzer performs the Wilcoxon rank-sum test in a pairwise manner (15) on the extracted image features from different cohorts. The Wilcoxon rank-sum test is a nonparametric statistical hypothesis test used to compare independent samples (i.e., image features from different cohorts in our case) and assess whether their distributions differ significantly. If the feature values of one cohort are significantly different from those of the other cohorts (p < 0.05), this cohort is highlighted in red in the subviolin plots. Thus, users can assess the degree of the batch effect not only at the overall level but also at the cohort level. Moreover, tracing the features related to batch effects may help users better understand the extent of variability in their data and enlighten them on how batch effects can be alleviated.
Principal Variance Component Analysis (PVCA) combines principal component analysis (PCA) and variance component analysis (VCA) to numerically partition the variance in data represented by batch and other clinical related factors into different major components (14). This can be useful for identifying and quantifying major sources of variability in the data. If batch effects exist, the batch variance in the plot is one of the main contributors to the overall variation. The implementation of PVCA is shown in the Algorithm S3.
Batch Effect Score (BES) serves as another numeric index of the degree of batch effects. It is a measure of the batch effect through the assessment of the stability of sample clustering in cohorts based on consensus clustering (16). We assumed that cohorts with a high degree of batch effects contained samples that were stably clustered in multiple clustering runs. Thus, given resampling iterations , the sampled datasets can be defined as with a pre-defined proportion, e.g., 80%. In each resampling iteration , let denote the connectivity matrix corresponding to dataset , where is the number of all samples. The results of clustering of pairs of samples can be defined as entries of the following matrix as:
| (1) |
Let be the indicator matrix. Here, equals 1 when items and are both present in the sampled dataset . Otherwise, it equals 0. The consensus matrix can be defined as the normalized sum of the connectivity matrices of all the resampled datasets:
| (2) |
In other words, the consensus matrix records the number of times items and are assigned to the same cluster, divided by the total number of times both items are selected in the same resampled datasets. Therefore, the consensus matrix measures how stably and robustly the sample pairs are clustered together.
Let indicate whether items and belong to the same cohort; we then, define it as follows:
| (3) |
Here, we define the BES as the area under the curve (AUC) (17) of the consensus matrix and cohort matrix where item :
| (4) |
| (5) |
| (6) |
| (7) |
When the value of the BES approaches 1, the samples are easily and stably clustered according to their cohorts, indicating a high degree of batch effects among their cohorts. In contrast, if the value of the BES is close to 0.5, the samples are randomly clustered, indicating a low degree of batch effects among the cohorts. If the BES value is below 0.5, it may indicate data issues within the dataset, such as mistakenly treating data from a single center as if they were from different centers, or significant disparities in the number of samples from different centers. In such cases, users should review their data and consider balancing the sample sizes to address these issues. Note that the BES can be used not only for all cohorts in a dataset but also for cohort-level within a multi-center dataset.
Getting started with BEEx
Easily runnable GUI versions of BEEx for Linux and Windows can be found at https://github.com/wuusn/beex, along with the sample data used in this study for quick reproduction. Figure 2 shows the BEEx GUI.
Figure 2:

Screenshot of the BEEx GUI. Users need to specify the location of the image files of each cohort with a few clicks. After clicking the “Run” button with the selected analyzers, the results pop up in separate windows.
Technical requirements: 64-bit operating system and at least 16 GB of RAM. However, a larger RAM size is recommended for WSI processing and for large numbers of images. Figure S2 illustrates how the time cost of BEEx varies with different numbers of centers and images.
BEEx supports a large number of medical image formats, including basic formats (.png, .jpg, and .tiff), OpenSlide (.svs), NIfTI (.nii.gz), and DICOM (.dcm).
All BEEx codes are open-source and can be found in our GitHub repository (https://github.com/wuusn/beex), along with current issues and proposed code changes. Figure 3 shows the use of BEEx from the command line. Figures S3 to S8 provide a step-by-step tutorial on how to use the BEEx GUI.
Figure 3:

Example of running code from the command line. The user needs to create a configuration .yaml file indicating the basic information about the dataset and then run the Python script with the .yaml file as the input.
If you need help, feel free to open an issue in the GitHub repository.
Data Availability
The WSI, VTMA, MRI, and CT data analyzed in this study are publicly available at https://doi.org/10.6084/m9.figshare.25271215.v1. The TCGA cohort of WSI data analyzed in this study was obtained from The Cancer Genome Atlas at https://portal.gdc.cancer.gov/projects/TCGA-LIHC.
Code Availability
The Python code for the BEEx framework is available on GitHub at https://github.com/wuusn/beex. A reproducible capsule (https://codeocean.com/capsule/2528710/tree/v1) is hosted on the CodeOcean platform.
Results
In this study, we demonstrated the use of BEEx through four case studies, each of which focused on a specific type of medical image. In Case Study-WSI, the average image sizes at 40× magnification for GDPH, TCGA, FAHKU, and FAHZU were 77492 × 79768, 96965 × 68054, 100825 × 80537, and 84246 × 64355 pixels, respectively. The average ages within these cohorts were 55 ± 11, 60 ± 12, 52 ± 11, and 55 ± 10, respectively. The male-to-female ratios in these cohorts were (85.3% and 14.7%), (68.6% and 31.4%), (91.2% and 8.8%), and (81.9% and 18.1%), respectively. In Case Study-VTMA, each TMA image had a size of 2800 × 2800 pixels at 40× magnification. VTMA-1, VTMA-2, and VTMA-3 contained 19 non-invasive and 83 invasive tumor images, one non-invasive and 93 invasive tumor images, and 45 non-invasive and 52 invasive tumor images, respectively. Additionally, they contained 7 low, 51 medium, and 44 high histological grade tumor images; 8 low, 48 medium, and 38 high histological grade tumor images; and 13 low, 42 medium, and 42 high histological grade tumor images, respectively. In Case Study-MRI, the average ages within SSPH, RHWU, and YCPH were 71 ± 9, 71 ± 8, and 69 ± 9, respectively. Additionally, the average values of the prostate-specific antigen (PSA) test were 20.65 ± 20.77, 75.64 ± 146.14, and 15.40 ± 24.57, respectively. The imaging matrices for these cohorts were 178 × 132, 256 × 256, and 224 × 224 pixels, respectively. In Case Study-CT, the average ages of GDPH-1, GDPH-2, and GDPH-3 were 62 ± 12, 61 ± 11, and 64 ± 12, respectively. The male-to-female ratios in these cohorts were (68.37% and 31.63%), (58.16% and 41.84%), and (60% and 40%), respectively. All the imaging matrices for these cohorts were 512 × 512. In Case Study-VTMA, the results of BEEx demonstrate that there was a low degree of batch effects. Conversely, in Case Study-WSI, Case Study-MRI-A and -B, BEEx detected a high degree of batch effects among these cohorts. In the Case Study-MRI, BEEx revealed that Case Study-MRI-B had a lower degree of batch effects than did Case Study-MRI-A. In Case Study-CT, BEEx did not detect batch effects in the dataset.
BEEx reveals a high degree of batch effects in a WSI-based multi-center HCC dataset.
In this case study, we examined a dataset that included four HCC whole-slide image cohorts: GDPH, TCGA, FAHKU, and FAHZU. Clear batch effects were observed among the four cohorts in the dataset. However, BEEx did not identify any significant batch effects between GDPH and FAHKU.
As shown in Figure 4(a), differences in hematoxylin and eosin (H&E) staining styles and intensities among the cohorts are noticeable. More specifically, TCGA and FAHZU exhibit significantly different staining styles from those of GDPH and FAHKU, whereas the staining styles of GDPH and FAHKU are not significantly from each other. Similarly, FAHZU and TCGA each have distinct clusters, whereas GDPH and FAHKU overlap, as shown in Figure 4(b), indicating a high degree of batch effects among FAHZU, TCGA, and the combination of GDPH and FAHKU, as well as a low degree of batch effects between GDPH and FAHKU. Moreover, in the Distribution Analysis with Violin Plots, as shown in Figure 4(c), most tissue-, texture-, and color-based features (e.g., Dark Tissue Ratio, Stain H std, Contrast RMS) distinguish all cohorts significantly (p < 0.05), whereas some features fail to distinguish cohorts (e.g., RGB 1 Brightness mean) or can only separate one or two cohorts significantly (e.g., YUV 2 Brightness mean and HSV 1 Brightness std). Based on the results of hierarchical clustering (Figure 4 (d)), the samples from TCGA and FAHZU were primarily grouped based on their cohorts; however, the samples from GDPH and FAHKU were mixed. In the PVCA, as shown in Figure 4(e), the batch was the major contributor (63.48%) to the overall variation. Furthermore, the overall BES of Case Study-WSI is 0.8417. The cohort-to-cohort BES values of TCGA to GDPH, TCGA to FAHKU, TCGA to FAHZU, GDPH to FAHKU, GDPH to FAHZU, and FAHKU to FAHZU are 0.9956, 0.9673, 0.9055, 0.5523, 0.9939, and 0.8839, respectively.
Figure 4:

Visualization of BEEx analyzers in Case Study-WSI. In this case study, BEEx detected a high degree of batch effects. TCGA and FAHZU show a significantly different staining style from GDPH and FAHKU in the Image Overviewer (a). Moreover, there is no apparent difference between the staining styles of GDPH and FAHKU. Similarly, in the Distribution Analysis with UMAP (b), there are three noticeable distinct clusters, according to the cohorts. FAHZU and TCGA each have a separate cluster, whereas GDPH and FAHKU overlap, sharing a same separated cluster. In the Distribution Analysis with Violin Plots (c), most features can separate these four cohorts significantly. In the Hierarchical Clustering (d), almost all samples of TCGA and FAHZU are clustered to their cohort in the right, whereas the samples of GDPH and FAHKU were mixed in the left. In the PVCA (e), Batch explains 63.48%, making it the largest contributor to the overall data variance.
BEEx finds a low degree of batch effects in a TMA-based multi-center BCa dataset.
In this dataset, all the TMAs were stained using the same reagents and antibodies and processed in three different laboratories. As shown in Figure 5(a), the variations in H&E staining styles across these cohorts are barely discernible to the human eye, and the UMAP plot shown in Figure 5(b) does not exhibit clear clusters. Nevertheless, a few subvisual features were detected by our Distribution Analysis Violin Plots (Figure 5(c)), demonstrating that some color- and texture-based features can significantly differentiate these cohorts (p<0.05, highlighted in red). This is because the staining of slides from different cohorts was conducted in different laboratories. In addition, the clustergram shown in Figure 5(d) demonstrates that the samples are not grouped mainly based on the cohort to which they belong. However, most samples of VTMA-3 are clustered on the left and VTMA-2 on the right, reflecting batch effects between VTMA-3 and VTMA-2. As shown in Figure 5(b), VTMA-2 shares a different cluster below VTMA-1 and VTMA-3. The PVCA shown in Figure 5(e) depicts the overall batch effects, where batch contributes only 2.96% to the overall data variance, which is lower than overgrade (a tumor feature), which contributes 11.48%. Instead, the unknown random effect residuals explain most of the variance in the data. In addition, the overall BES of Case Study-VTMA is 0.5072. The cohort-to-cohort BES values for VTMA-1 to VTMA-2, VTMA-1 to VTMA-3, and VTMA-2 to VTMA-3 are 0.5008, 0.5073, and 0.6339, respectively.
Figure 5:

Visualization of BEEx analyzers in Case Study-VTMA. In this case study, BEEx demonstrates a low degree of batch effects. There are no clear color or shape differences in the Image Overviewer (a). Moreover, the Distribution Analysis with UMAP (b) shows no distinct clusters, although VTMA-2 appears to share a different cluster below from VTMA-1 and VTMA-3. In the Distribution Analysis with Violin Plots (c), several color- and texture-based features significantly separate these cohorts. In Hierarchical Clustering (d), most samples of VTMA-3 are clustered in the left, whereas those of VTMA-2 are clustered in the right. Finally, in the PVCA (e), Batch explains only 2.96% of the overall variance of the data.
BEEx validates batch effect rectification with downstream task performance improvement in an MRI-based multi-center PCa dataset.
In this case study, we examined a dataset that consisted of three PCa MRI cohorts: SSPH, RHWU, and YCPH. BEEx successfully revealed the existence of batch effects in Case Study-MRI-A and -B and verified the reduction of batch effects in Case Study-MRI-B, where a batch effect rectification method was applied to Case Study-MRI-A (10). In addition, the case with batch effect rectification (Case Study-MRI-B) showed better downstream task performance than the case without batch effect rectification (Case Study-MRI-A).
Here, we examine the batch effects of Case Study-MRI-A, which did not apply a batch effect rectification method (10). Figure 6(a) shows that the image contrast intensities and resolutions for these cohorts differ. Noticeably, SSPH appears to have significantly smoother boundaries than the other two cohorts, possibly because the post-processing algorithm was applied after image capture. In the UMAP visualization of the distribution analysis shown in Figure 6(b), three distinct clusters are evident according to the cohorts. In the Distribution Analysis Violin Plots shown in Figure 6(c), most metadata features (e.g., VRX and ROWS) distinguish these cohorts significantly (p<0.05), and some measurement features can also separate cohorts (e.g., CPP and SNR1) significantly to some extent. In the clustergram shown in Figure 6(d), most samples are grouped based on their cohorts. In the PVCA shown in Figure 6(e), the batch variance (63.12%) is the primary contributor to the overall variation. The overall BES value of Case Study-MRI-A is 0.8760. The cohort-to-cohort BES values of the RHWU to SSPH, RHWU to YCPH, and SSPH to YCPH are 0.9703, 0.9876, and 0.9554, respectively. In conclusion, BEEx found that Case Study-MRI-A had a high degree of batch effects and that most features could significantly separate multi-center cohorts.
Figure 6:

Visualization of BEEx analyzers in Case Study-MRI-A. Image Overviewer (a) shows the visual difference among the cohorts. At the same time, Distribution Analysis with UMAP (b) shows three clearly distinct clusters, according to their cohorts. In the Distribution Analysis with Violin Plots (c) most metadata and measurement features can significantly separate cohorts to some extent (with p < 0.05). In the Hierarchical Clustering (d), almost all samples are grouped based on their cohorts. PVCA (e) shows that Batch explains 63.18%, making it the largest contributor to the overall data variance.
In Case Study-MRI-B, we employed a batch effect rectification method (10) that combined generative models (18) with an adversarial training strategy (19) to synthesize high-b-value MRI without rectal artifacts. After deploying batch effect rectification, the AUC performance for the downstream PCa diagnosis task improved by 6% at the patient level (10). In parallel, BEEx successfully validated the decrease in batch effects in Case Study-MRI-B compared to Case Study-MRI-A. As shown in Figure 7(a), the Image Overviewer demonstrates less visual variation than that shown in Figure 6(a). Although the UMAP plot shown in Figure 7(b) still has three clusters, some samples from different cohorts overlap, further indicating that feature-level differences are decreased via the batch effect rectification method. In the Distribution Analysis with Violin Plots, as shown in Figure 7(c), only two features (i.e., RNG and PSNR) can significantly (p < 0.05) separate three cohorts, whereas most features become unable to separate cohorts or can only separate one cohort. Similarly, in the clustergram shown in Figure 7(d), the result of sample grouping by cohorts is not as clear as before, as shown in Figure 6(d). In the PVCA in Figure 7(e), the data variance contributed by the batch decreases noticeably to 5.08%. In addition, the overall BES of Case Study-MRI-B decreases to 0.7939. The cohort-to-cohort BES values of the RHWU to SSPH, RHWU to YCPH, and SSPH to YCPH were 0.7688, 0.7922, and 0.9662, respectively.
Figure 7:

Visualization of BEEx analyzers in Case Study-MRI-B following the deployment of a batch effect rectification method showing a decreased degree of batch effects compared with the results of Case Study-MRI-A. Image Overviewer (a) shows the visual differences among cohorts. Distribution Analysis with UMAP (b) reveals three clusters based on the cohort, although some samples from RHWU overlapped with the other cohorts. Distribution Analysis with Violin Plots (c) demonstrates that most features cannot separate cohorts significantly or can only separate one cohort significantly. Hierarchical Clustering (d) shows that some samples are grouped according to cohorts, whereas other subgroups are mixed. The results of PVCA (e) demonstrate that Batch explains 5.08% of the overall data variance.
BEEx finds no batch effects in a CT-based control case study.
Case Study-CT is a control case study in which three pseudo-cohorts, GDPH-1, GDPH-2, and GDPH-3, were artificially created using randomly selected cases from the same dataset. As expected, BEEx found no batch effect among the cohorts. Specifically, there are no visual differences between the three cohorts (Figure 8(a)). Moreover, the UMAP shown in Figure 8(b) does not show clearly separated clusters, according to the cohorts. Similar to the violin plots shown in Figure 8(c), each subplot shares a similar distribution, indicating that cohorts GDPH-1, GDPH-2, and GDPH-3 have the same distribution. In addition, the clustergram shown in Figure 8(d) demonstrates that the samples are not grouped based on their cohorts. Finally, the batch contributes to 0.06% of the overall data variance as shown in Figure 8(e). The overall BES value is 0.4987, and the cohort-to-cohort BES values of GDPH-1 to GDPH-2, GDPH-1 to GDPH-3, and GDPH-2 to GDPH-3 are 0.4976, 0.5009, and 0.5009, respectively, which are close to the results of random guessing.
Figure 8:

BEEx results of Case Study-CT. This is a control case study that demonstrates no batch effects among all analyzers. There is no visual difference among these cohorts, as shown in the Image Overviewer (a). In the Distribution Analysis with UMAP (b), BEEx does not observe clearly separated clusters, according to the cohorts. In addition, the Distribution Analysis Violin Plots in (c) shows that each cohort shares a similar feature distribution without significant differences. Similarly in Hierarchical Clustering (d), the samples are not grouped according to their cohort. Batch contributes only 0.06% to the overall variance in PVCA (e).
Discussion
Recently, AI models have gained significant attention in addressing clinical challenges. However, owing to the heterogeneity of the tumor microenvironment, population variations, and interobserver variability among experts, researchers need to validate the reliability and applicability of extant AI models through multi-center studies (5,20,21). Nevertheless, batch effects, ubiquitous non-biological variations in multi-center datasets, often affect model performance. A trained model affected by batch effects typically lacks generalizability and is prone to misinterpretations and wrong conclusions. Consequently, there is an urgent need for effective tools that can evaluate batch effects in multi-center settings.
To the best of our knowledge, existing platforms specifically designed to investigate batch effects neither support medical images (6) nor are they open-source (2,3), as shown in Table S3. Therefore, to empower researchers to conduct batch effects analysis prior to validating AI models in a multi-center study, we introduced BEEx, an open-source platform, for batch effect analysis. BEEx is a scalable and versatile framework designed to read, process, and analyze a wide range of medical images for batch-effect evaluation. It is fully automated and features both a GUI and a command-line interface. Additionally, it includes configuration files that enable analysis across different cases and are freely available to the public. With the help of BEEx, users can quantify the severity of batch effects among cohorts, trace possible sources of batch effects, and gain insights to address and validate the methods used for batch effect rectification.
Focusing on pathological images, we conducted two case studies in which BEEx effectively revealed batch effects via multiple visualization methods. In a case study involving radiological image data, we utilized BEEx for batch effect examination before and after batch effect rectification and observed a significant reduction in the batch effect. Hence, we can use BEEx not only to observe batch effects in data prior to multi-center modeling but also as a validation tool for batch effect rectification methods.
The effectiveness of BEEx has been consistently demonstrated across multiple case studies, thereby reaffirming its utility in detecting and analyzing batch effects. These batch effects, which are often overlooked, can significantly affect the results of experiments by introducing artificial variation, generating spurious associations, reducing reproducibility, and resulting in misleading statistical analyses. Therefore, their identification and mitigation are of utmost importance in medical image analysis. As the features provided by BEEx are basic features used for image quality control, they may not fully meet the needs of all users. Nevertheless, because BEEx is an open-source framework, users are free to integrate additional features.
The ability of BEEx to detect these batch effects, as evidenced in case studies involving TMAs, WSIs, and MR and CT images, shows its useful application in diverse research contexts. Researchers using BEEx as a prescreening tool prior to the development of a multi-center AI-based cancer diagnostic model can confidently assess batch effects in their data, identify their sources, and quantify their impact. This will allow them to mitigate batch effects by using appropriate normalization techniques to enhance the reliability and validity of their findings.
Supplementary Material
Significance:
BEEx is a pre-screening tool for image-based analyses that allows researchers to evaluate batch effects in multi-center studies and determine their origin and magnitude to facilitate development of accurate AI-based cancer models.
Acknowledgements
This work was funded by The National Key R&D Program of China, 2023YFC3402800; National Natural Science Foundation of China (No. 82272084, No.82302130); The National Natural Science Foundation of China Excellent Young Scientists Fund (Overseas) (No. 22HAA01598); Regional Innovation and Development Joint Fund of National Natural Science Foundation of China (No.U22A20345); National Science Fund for Distinguished Young Scholars of China (No. 81925023); Guangdong Provincial Key Laboratory of Artificial Intelligence in Medical Image Analysis and Application (No. 2022B1212010011); High-level Hospital Construction Project (No.DFJHBF202105); Zhejiang Province Health Major Science and Technology Program of National Health Commission Scientific Research Fund (No. WKJ-ZJ-2426).
Footnotes
Conflict of Interest Declaration
The authors declare that they have NO affiliations with or involvement in any organization or entity with any financial interest in the subject matter or materials discussed in this manuscript.
Software Environment
BEEx requires a 64-bit Windows, Linux, or Mac operating system with at least 16 GB of RAM. It is implemented using a Python 3.11.5 programming environment with the openslide-python 1.30, Pillow 10.0.1, MedPy 0.4.0, and Pydicom 2.4.3 (RRID:SCR_002573) for medical image reading; Numpy 1.25.2 (RRID:SCR_008633), pandas 2.1.1, scikit-image 0.21.0, and scikit-learn 1.3.1 for data processing and analysis; and matplotlib 3.8.0 and seaborn 0.13.0 for visualization. The BEEx GUI was based on PyQt and PySide6 6.7.1.
Blinding
All the patients remain anonymous in this study.
Power Analysis
Power analysis is not applicable to our study.
References
- 1.Schmitt M, Maron RC, Hekler A, Stenzinger A, Hauschild A, Weichenthal M, et al. Hidden variables in deep learning digital pathology and their potential to cause batch effects: prediction model study. J Med Internet Res. 2021;23:e23436. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Kothari S, Phan JH, Stokes TH, Osunkoya AO, Young AN, Wang MD. Removing batch effects from histopathological images for enhanced cancer diagnosis. IEEE J Biomed Health Inform. 2014;18:765–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Stopsack KH, Tyekucheva S, Wang M, Gerke TA, Vaselkiv JB, Penney KL, et al. Extent, impact, and mitigation of batch effects in tumor biomarker studies using tissue microarrays. eLife. 2021;10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Wu Y, Koyuncu CF, Toro P, Corredor G, Feng Q, Buzzy C, et al. A machine learning model for separating epithelial and stromal regions in oral cavity squamous cell carcinomas using H&E-stained histology images: A multi-center, retrospective study. Oral Oncol. 2022;131:105942. [DOI] [PubMed] [Google Scholar]
- 5.Lu C, Bera K, Wang X, Prasanna P, Xu J, Janowczyk A, et al. A prognostic model for overall survival of patients with early-stage non-small cell lung cancer: a multicentre, retrospective study. Lancet Digit Health. 2020;2:e594–606. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Zhu T, Sun R, Zhang F, Chen G-B, Yi X, Ruan G, et al. Batchserver: A web server for batch effect evaluation, visualization, and correction. J Proteome Res. 2021;20:1079–86. [DOI] [PubMed] [Google Scholar]
- 7.Sadri AR, Janowczyk A, Zhou R, Verma R, Beig N, Antunes J, et al. Technical Note: MRQy - An open-source tool for quality control of MR imaging data. Med Phys. 2020;47:6029–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Janowczyk A, Zuo R, Gilmore H, Feldman M, Madabhushi A. HistoQC: An Open-Source Quality Control Tool for Digital Pathology Slides. JCO Clin Cancer Inform. 2019;3:1–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Chen Y, Zee J, Smith A, Jayapandian C, Hodgin J, Howell D, et al. Assessment of a computerized quantitative quality control tool for whole slide images of kidney biopsies. J Pathol. 2021;253:268–78. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Hu L, Guo X, Zhou D, Wang Z, Dai L, Li L, et al. Development and Validation of a Deep Learning Model to Reduce the Interference of Rectal Artifacts in MRI-based Prostate Cancer Diagnosis. Radiol Artif Intell. 2024;6:e230362. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Hintze JL, Nelson RD. Violin Plots: A Box Plot-Density Trace Synergism. The American Statistician. 1998;52:181–4. [Google Scholar]
- 12.McInnes L, Healy J, Saul N, Großberger L. UMAP: uniform manifold approximation and projection. JOSS. 2018;3:861. [Google Scholar]
- 13.Murtagh F, Contreras P. Algorithms for hierarchical clustering: an overview. WIREs Data Mining Knowl Discov. 2012;2:86–97. [Google Scholar]
- 14.Li J, Bushel PR, Chu T-M, Wolfinger RD. Principal variance components analysis: estimating batch effects in microarray gene expression data. In: Scherer A, editor. Batch effects and noise in microarray experiments. Chichester, UK: John Wiley & Sons, Ltd; 2009. page 141–54. [Google Scholar]
- 15.Cuzick J. A Wilcoxon-type test for trend. Stat Med. 1985;4:87–90. [DOI] [PubMed] [Google Scholar]
- 16.Monti S, Tamayo P, Mesirov J, Golub T. Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data. Mach Learn. 2003;52:91–118. [Google Scholar]
- 17.Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143:29–36. [DOI] [PubMed] [Google Scholar]
- 18.Hu L, Zhou D, Zha Y, Li L, He H, Xu W, et al. Synthesizing High-b-Value Diffusion-weighted Imaging of the Prostate Using Generative Adversarial Networks. Radiology: Artificial Intelligence. 2021;e200237. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Hu L, Zhou D, Xu J, Lu C, Han C, Shi Z, et al. Protecting Prostate Cancer Classification from Rectal Artifacts via Targeted Adversarial Training. IEEE J Biomed Health Inform. 2024;1–14. [DOI] [PubMed] [Google Scholar]
- 20.Zhao S, Chen D-P, Fu T, Yang J-C, Ma D, Zhu X-Z, et al. Single-cell morphological and topological atlas reveals the ecosystem diversity of human breast cancer. Nat Commun. 2023;14:6796. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Hölscher DL, Bouteldja N, Joodaki M, Russo ML, Lan Y-C, Sadr AV, et al. Next-Generation Morphometry for pathomics-data mining in histopathology. Nat Commun. 2023;14:470. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The WSI, VTMA, MRI, and CT data analyzed in this study are publicly available at https://doi.org/10.6084/m9.figshare.25271215.v1. The TCGA cohort of WSI data analyzed in this study was obtained from The Cancer Genome Atlas at https://portal.gdc.cancer.gov/projects/TCGA-LIHC.
