Skip to main content
PLOS Digital Health logoLink to PLOS Digital Health
. 2026 Sep 28;5(9):e0001744. doi: 10.1371/journal.pdig.0001744

Foundation model-powered deep learning of endometrial histology for predicting the cumulative live birth of an in vitro fertilization cycle

Nianbo Xu 1, Andy Chun Hang Chen 1,2,3, Hanzhang Ruan 1,2, Donglin Yang 4, Renjie Liao 5,6,7, Xiaojuan Qi 4, Sze Wan Fong 1, Dandan Cao 3, Lu Yu 8, Yuanhua Huang 2,8, William Shu Biu Yeung 1,2,3, Ernest Hung Yu Ng 1,3,*, Yin Lau Lee 1,2,3,*
Editor: Feng Liu9
PMCID: PMC13618903  PMID: 42804460

Abstract

Objective and reproducible assessment of endometrial receptivity is essential for optimizing in vitro fertilization (IVF) success, yet traditional histological dating suffers from observer variability. This study investigated whether deep learning of hematoxylin and eosin histology images could support cumulative live birth prediction in IVF. An end-to-end ResNet-18 was compared with a UNI2-h-based pipeline, using UNI2-h as a frozen feature extractor. Ten-fold cross-validation ensembles were developed from natural-cycle endometrial biopsies and evaluated in an internal held-out cohort with live birth outcomes. Additional phase-based evaluation was performed, which tested the performance in distinguishing LH + 7 versus non-LH + 7 phases. A luminal epithelium (LE)-focused design was also assessed to examine whether concentrating on maternal-embryo interface could improve fertility-oriented learning. The whole-slide ResNet-18 model performed well in internal outcome-based testing but near random in external phase-based testing. In contrast, the UNI2-h model with mean pooling and a multilayer perceptron classifier showed less divergent performance, with ensembled AUROCs of 0.74 ± 0.03 in internal outcome-based evaluation, and 0.90 ± 0.05 in external phase-based testing. Despite a smaller training set, LE-focused models retained comparable performance. In internal outcome-based testing, LE-focused ResNet-18 and UNI2-h models achieved AUROCs of 0.71 ± 0.03 and 0.74 ± 0.09 respectively. External phase-based testing yielded AUROCs of 0.80 ± 0.16 for ResNet-18, and 0.84 ± 0.06 for UNI2-h. Grad-CAM review of LE-focused ResNet-18 models showed attention commonly on LE-alone or mixed with adjacent stroma. In multimodal analyses, integrated models incorporating histology significantly outperformed the clinical metadata-only model. Integrated model’s feature weighting showed dominant histology outputs, smaller contributions from estradiol and maternal age, and negligible contributions from progesterone, endometrial thickness, and BMI. These findings support histology-based AI for fertility-oriented endometrial assessment and highlight biologically informed design and repurposed foundation models as promising bases for clinically meaningful prediction.

Author summary

Clinicians still lack an objective way to assess whether the endometrium is in an optimal state to support a successful pregnancy. Traditional histological assessment relies on pathologists’ interpretations and vary among observers, while newer molecular tests remain costly and uncertain reliability. In this study, we explored whether artificial intelligence (AI) could extract useful fertility-related information from routine endometrial histology slides. We compared a conventional training approach with a more modern strategy that repurposed an existing AI model with pre-existing histology knowledge. Two testing cohorts, including an internal cohort based on live-birth outcomes and an external cohort evaluated on luteal-phase classification were included. The repurposed approach showed more consistent performance between the two testing cohorts, while the conventional approach fell to near-random in the external phase-based cohort. We also tested whether focusing on endometrial luminal epithelium, where the embryo first makes contact, could improve prediction. Surprisingly, luminal epithelium-focused models performed comparably to the whole-slide approach despite using far fewer training images. These findings support AI-based histology analysis as a promising approach that warrants further refinement for more objective endometrial assessments. Combining foundation models with biologically informed design could provide a useful direction for future fertility-related diagnostics.

Introduction

The global birth rate remains low as many countries report their birth rate lower than the replacement level of 2.1 births per woman. The current cumulative live birth rate per cycle remained at ~30–40% only [1] despite advancements in in vitro fertilization (IVF). Poor embryo quality is often cited as the primary factor in infertility, but the probability of a live birth per a single euploid embryo transfer stays at approximately 60% [2,3], suggesting that maternal factors beyond embryo euploidy contribute to the outcomes. Consequently, a large proportion of patients require multiple transfer cycles, which can be particularly challenging for patients with a limited number of good quality embryos. Repeated failures are also significantly correlated with a higher risk of anxiety and depression even after accounting for socioeconomic confounding factors [4,5].

Optimal endometrial preparation is essential for the establishment of a successful pregnancy [6]. Yet, clinicians still lack an accurate and low cost method to assess endometrial receptivity and predict live birth outcomes [7]. Conventional non-invasive markers, such as ultrasound-based endometrial thickness and vascularization fail to serve as reliable predictive tools because of conflicting results reported [8–10]. The Noyes criteria, which provided descriptive morphological standards to date endometrial samples, has long been the “gold standard” of endometrial histological dating since its introduction in 1950 [11]. However, the technique is labor intensive and exhibits significant observer variability, compromising its reliability [12,13]. Advancements in understanding the molecular signatures of the window of implantation (WOI) have provided hopes of novel assays for endometrial dating. Up-to-date, the true clinical value of currently available molecular assays, like the Endometrial Receptivity Analysis (ERA), remains controversial due to unproven benefits in predicting IVF success [14–16]. Thus, there is a critical need for an objective and low-cost method to evaluate endometrial receptivity.

Since the introduction of the initial whole slide image (WSI) scanner in the 1990s, computational pathology has evolved from simple feature extraction to advanced deep learning algorithms capable of interpreting complex high-dimensional features [17]. Deep learning models are commonly developed either by full training from scratch or by adapting pretrained foundation models. The full training approach, typically utilizing Convolutional Neural Networks (CNNs), enables task-specific learning but is inherently data-demanding and sensitive to institutional biases such as imaging and staining variations [18,19]. Consequently, the field is increasingly shifting toward foundation models as a more robust alternative [20–22]. Unlike conventional approaches requiring large task-specific datasets, foundation models such as UNI [21] and Virchow [22] are pretrained on massive histology archives, making the utilization of complex Vision Transformer (ViT) backbones practically feasible. The increasing availability of these foundation models represents a critical methodological advance to make the development of task-specific pathological deep learning models easier, especially when sample sizes are restricted [23].

The application of deep learning based artificial intelligence (AI) model in healthcare is a fast-growing field with immense potential [24]. While AI has been widely tested for its application in optimizing stimulation protocols and embryo selection, its use in predicting endometrial receptivity remains limited. By training a ResNet-18 model from scratch, a widely adopted efficient architecture for natural and medical image analysis [25,26], to differentiate endometrial samples based on fertility outcomes, we reported the earliest application of deep learning in endometrial receptivity [27]. While the reported models achieved satisfactory live birth prediction with an accuracy of 75%, the wider applicability of the model which was derived exclusively from patients undergoing Hormone Replacement Therapy (HRT) was constrained. Endometrium in HRT cycles exhibits distinct histological profiles compared to those in natural cycles [28,29]. Therefore, the generalizability of such a model to natural cycles remains unverified. Furthermore, a significant knowledge gap remains regarding the transferability of modern foundation models to the endometrial function-oriented tasks. Many widely used histopathology foundation models have been developed and extensively benchmarked in cancer-heavy settings [20–22]. It is currently unknown whether these cancer-specific extractors can effectively characterize the subtle, non-malignant morphological changes associated with endometrial receptivity in healthy tissues.

In this study, we aimed to first evaluate the potential transferability of the HRT-trained ResNet-18 model for assessing the performance using WSIs obtained in natural cycles. We also evaluated whether the histopathology pretrained UNI2-h foundation model could provide a simpler, more robust starting point for digital pathology-based fertility outcome prediction. As luminal epithelium (LE) of endometrium is the first point of contact between maternal tissue and the implanting embryo, we also explored an LE-focused modelling approach specific to the fertility-oriented predictive goal.

Materials and methods

Ethical approval

The study protocol was approved by the Institutional Review Boards of the University of Hong Kong/Hospital Authority Hong Kong West Cluster (IRB: UW 25–573) and the Ethics Committee of the University of Hong Kong-Shenzhen Hospital (IRB reference number [2026]087).

Study cohorts and sample collection

Archived anonymized paraffin-embedded tissues with known pregnancy outcomes from a previous study [30] were used in this study. IVF patients who had regular cycles and underwent their first IVF cycles at the Centre of Assisted Reproduction and Embryology, The University of Hong Kong - Queen Mary Hospital and the Assisted Reproduction Centre of Kwong Wah Hospital from 2017 to 2021 were included in this study. They were asked to donate endometrial biopsy samples in their natural cycles. Individuals with [1] an abnormal uterine cavity as determined by saline infusion sonogram or hysteroscopy; [2] untreated uterine pathologies such as endometrial polyps or fibroids; [3] untreated hydrosalpinx; and [4] use of donor oocytes were excluded. Pipelle samplers (CCD Laboratories, Paris, France) were used to obtain the endometrial biopsy at 7 days after a luteinizing hormone (LH) surge (LH + 7). Biopsied samples were fixed and embedded in paraffin blocks. From each block, 4 μm tissue sections were prepared for hematoxylin and eosin (H&E) staining, and the stained slides were digitized as high resolution WSIs at 40 × magnification (approximately 0.220 µm per pixel, or mpp) using a Hamamatsu NanoZoomer S210 slide scanner (Hamamatsu Photonics K.K., Shizuoka Prefecture, Japan). The digitalized WSI files were stored natively in NDPI files first before transformed for further formatting change and downstream processing. In total, 183 internal samples were included in this study, of which 152 samples were assigned for training and validation purposes, and another 31 samples were assigned as an internal held-out test set. Essential metadata and clinical characteristics were summarized in S1 Table.

We also included an external phase-based cohort consisting of 21 fertile patients WSIs collected at natural cycle from the University of Hong Kong-Shenzhen Hospital. The patient sample condition and tissue collection details were described previously [31]. All donors had achieved a healthy live birth within two years of collection either via spontaneous pregnancy or a preimplantation genetic testing (PGT)-validated euploid embryo transfer. The collected samples were biopsied at five different endometrial menstrual phases, including LH + 3 (n = 3), LH + 5 (n = 3), LH + 7 (n = 9), LH + 9 (n = 3), and LH + 11 (n = 3). Paraffin block embedding and H&E staining were conducted at the University of Hong Kong-Shenzhen Hospital. The slides were scanned at 20 × scanning resolution (approximately equivalent to 0.440 mpp) using a KF-FL-020 WSI scanner (KFBIO, Zhejiang, China). The resulting digitalized WSI files were stored natively in KFB file format before transformed for further processing.

Ovarian stimulation and embryo transfer

Clinical management and IVF procedures were conducted according to the centers’ standard operating procedures as described earlier [30]. Briefly, ovarian stimulation using a gonadotropin-releasing hormone antagonist protocol was used and the starting dosage of FSH depended on the antral follicle count and body mass index of the patients. Final oocyte maturation was triggered with recombinant human chorionic gonadotropin (hCG), followed by oocyte retrieval 34–36 hours later. Embryos were cultured to reach cleavage or blastocyst stage and eventually selected for the subsequent fresh or frozen embryo transfer (FET) cycles. The blastocysts were graded based on the Gardner scoring system [32].

Outcome data collection

The primary outcome was the cumulative live birth, defined as a live birth occurring after 22 completed weeks of gestation [33] from a stimulated IVF cycle and its subsequent FET cycles within 6 months after start of ovarian stimulation. The outcomes were assigned as either live birth or unsuccessful outcomes. This binary classification was utilized as the ground truths for training and assessing the performance of the developed models in the internal dataset. Ultimately, among the 152 samples utilized for training purposes, 79 were labeled as unsuccessful and 73 were labeled as live birth. The 31 internal held-out testing cohort consisted of 20 unsuccessful and 11 live birth samples.

As the external phase-based cohort comprised fertile donors who did not undergo embryo transfer, direct cumulative live birth outcome labels were unavailable. Therefore, the labels of external phase-based testing cohort were inferred by the timing of the biopsy relative to the LH surge. LH + 7 was considered as receptive, as it is widely accepted to be the optimal timepoint for embryo transfer [31,34]. Our previous time-series single-cell data further supported marked transcriptomic transitions before and after LH + 7, with stromal and glandular epithelial cells showing major changes from LH + 5 to LH + 7, and epithelial cells exhibiting a further sharp transition at late LH + 9 [31]. Thus, a non-identical but similar phase-based binary label was assigned, with LH + 7 samples (n = 9) representing the live birth outcome samples, and pre-receptive (LH + 3, + 5) and post-receptive (LH + 9, + 11) samples collectively representing unsuccessful outcome samples (n = 12). This external testing design therefore evaluated a biologically relevant but non-equivalent target, focusing on distinguishing receptive-phase LH + 7 endometrium from non-LH + 7 samples, rather than directly validating cumulative live birth prediction.

Histology-based model selection and training parameter settings

This study utilized two separate deep learning architectures, the ImageNet 1K-pretrained ResNet-18 [26] and the histopathology pretrained UNI2-h [21], to achieve the same goal of developing a binary classification model to differentiate endometrial H&E WSIs based on their live birth outcome potentials. All the models generated were trained on a remote Linux server implemented with three available NVIDIA A100 Tensor Core GPUs, except for the HRT-trained ResNet-18 model (HRT-model), which was trained with details as described previously [27].

For all models, an optimal threshold was determined independently for each cross-validation fold by maximizing Youden’s J statistic on the validation cohort (i.e., predicted output higher than the calculated threshold interpreted as live birth; values lower than the threshold interpreted as unsuccessful). Adapting the cross-validation framework from previous studies, we implemented a 10-fold cross-validation strategy to construct an ensemble model, an approach demonstrated to enhance robustness and generalizability compared to individual models [27,35]. In brief, an ensemble was formed by selecting the best-performing model from each fold. The predicted outcomes were averaged with equal weighting to produce a final predicted probability (soft voting), and optimal thresholds were similarly averaged to establish a global threshold.

The adopted CNN-based ImageNet 1K-pretrained ResNet-18 was maintained largely consistent with previous studies by utilizing a linear classifier head for binary classification [27]. Since this model is designed to operate as a patch-level classifier, slide-level predictions were generated by calculating the mean of the predicted probability scores across all patches extracted from a single WSI (Fig 1). To better encounter overfitting issues with the limited sample size, the architecture was modified to include a dropout layer after global average pooling (GAP). Furthermore, conventional data augmentation techniques were applied during training, including random horizontal and vertical flips, rotation, and color jitter. The models were trained using the AdamW optimizer with a cosine decay learning-rate scheduler and a warmup period to prevent destabilizing model weights. Detailed hyperparameter settings, including learning rates, batch sizes, and weight decay values, are summarized in S2 Table.

Fig 1. Overview of deep learning pipelines for endometrial fertility outcome prediction.

Fig 1

The workflow utilizes input patches derived either from whole slide images (WSIs) or manually annotated LE/peri-LE regions. The samples included two outcomes, either a cumulative live birth (CLB) or unsuccessful outcome. Two distinct modeling strategies were implemented. One includes an ImageNet 1K-pretrained ResNet-18 model trained end-to-end as a patch-level classifier, with patch-level probability scores processed independently and averaged in the end to generate a final slide-level binary prediction (CLB vs. Unsuccessful). The second strategy utilizes a histopathology pre-trained UNI2-h foundation model as a feature extractor. Extracted patch features were aggregated via either a global averaged pooling (GAP) or an attention-based pooling strategy (ACMIL). The aggregated slide-level representation was classified using either a linear classifier head or a non-linear MLP head to predict the slide-level binary outcome. Figure partially created in BioRender [https://biorender.com/f125abi].

The ViT-H/14-based histopathology pretrained UNI2-h foundation model was utilized as a feature extractor to extract high dimensional features from image patches [21]. Considering the size of the UNI2-h model at 681M parameters and the relatively limited abundance of our available sample set, the model’s weights were kept completely frozen during the study. To map these patch-level features to slide-level predictions, two distinctive feature aggregation strategies were implemented (Fig 1). One configuration utilized a GAP mechanism, in which the patch-level features extracted from a single slide were averaged to form a slide-level bag. In parallel, an Attention-Challenging Multiple Instance Learning (ACMIL) mechanism [36] was adapted. ACMIL replaced the equal averaging with a learnable attention network which differentially weight patches based on their discriminative relevance. Averaged or weighted feature outputs from the aggregation mechanisms were eventually passed on to either a conventional linear or a non-linear Multi-Layer Perceptron (MLP) classifier head for the final binary prediction. The MLP variant comprised two hidden layers with 768 and 384 dimensions, and provided a non-linear projection to adapt the UNI2-h features to the current specific endometrial task. For all the UNI2-h extracted feature-based training, one NVIDIA A100 GPU was used. The hyperparameter settings were kept largely consistent across the WSI models, except for necessary adaptations to fit specific model’s input expectations. Detailed hyperparameter settings are shown in S2 Table.

For the LE-focused models specifically trained with patches containing LE and peri-LE stroma regions, the architectures and major parameter settings were kept consistent with the WSI approaches (S2 Table). Specifically, the ResNet-18 LE models were trained using the ImageNet 1K-pretrained backbone combined with a linear classifier head, while the UNI2-h LE models utilized the frozen feature extractor combined with an MLP head.

Patch image extraction and preprocessing

Direct processing of WSIs in full resolution during training remains impractical due to their typical gigapixel-scale dimensions. Therefore, all WSIs were cropped into smaller patch images in order to facilitate efficient training. Although the maximum native scanning resolution differed between the two cohorts, with internal samples scanned at 40× and external samples scanned at 20 × , all subsequent training and analyses were standardized to patches generated at 20 × resolution (0.440 mpp) for minimizing technical variations. An exception was the internal dataset used for testing the pre-developed HRT model, which utilized an unchanged 40× (0.220 mpp) resolution at 4832 × 6864 pixels to match its specific input settings as described previously [27].

For all WSI models, automated segmentation was performed using the Trident toolkit [37], and only patches with at least 35% of tissue were included. The size of the resulting patch datasets varied slightly depending on the expected input pixel dimension of specific model. For the pipeline involving the ResNet-18 WSI models, all involved patches were extracted at 20 × resolution with a dimension of 224 × 224 pixels. This dataset comprised a training set of 152 patients (92,738 patches), an internal testing set of 31 patients (22,361 patches), and an external phase-based testing set of 21 patients (33,011 patches). In parallel, for pipeline involving the UNI2-h WSI models, all involved patches were extracted at 20 × resolution with dimensions of 256 × 256 pixels. This dataset comprised a training set of 152 patients (85,827 patches), an internal testing set of 31 patients (17,713 patches), and an external phase-based testing set of 21 patients (31,873 patches).

In contrast, the LE-focused workflow relied on manual annotation. Approximate LE-containing region annotations were drawn using QuPath, a WSI viewing and processing tool, followed by automatic generation of fixed-size squares to cover the annotated Regions of Interest (ROIs). These squares were manually adjusted to maximize LE coverage while minimizing the inclusion of glandular structures and were subsequently extracted using a custom Groovy script. Unlike the WSI approach which covers the whole tissue, conducting separate rounds of manual annotation for each pipeline could introduce bias, leading to inconsistencies in areas such as ROI centering and stromal proportion between the two model inputs. Hence, a single set of ‘base’ patches was generated first to ensure consistent anatomical coverage across the two pipelines. Specifically, a series of base patches was first extracted at 260 × 260 pixels (20 × resolution) from the internal and external cohort, and subsequently center-cropped to match the input dimensions of each architecture (224 × 224 pixels for ResNet-18-based and 256 × 256 pixels for UNI2-h feature-based). Consequently, the patch counts were identical for both pipelines: a training set of 124 patients (3,552 patches), an internal testing set of 24 patients (657 patches), and an external phase-based testing set of 18 patients (783 patches).

Visualization and quantification of model attention regions

The gradient-weighted class activation mapping (Grad-CAM) approach [38] was used to visualize the attention regions on the original patch images of LE models. All patch images from the internal testing cohort (657 images) were utilized to generate the attention heatmaps. In these heatmaps, warmer (red) areas highlighted regions the model deemed more important for its prediction, while cooler (blue) areas had weaker contributions. Semi-quantitative analysis of the Grad-CAM-generated heatmaps was performed by manually evaluating the most prominent high-attention area within the LE-containing patches using the internal testing cohort. Attention categories were assigned according to the predominant localization of the highlighted region. Considering that the patches primarily consisted of LE, peri-LE stroma, and lumen, which were generally easily distinguishable on H&E images, Grad-CAM attention was prespecified into four broad categories comprising LE-predominant, mixed LE and peri-LE stroma, peri-LE stroma-predominant, and patterns on unclassified regions such as lumen and potential artifacts. Other cell types potentially associated with fertility outcomes, such as endothelial and immune cells, were not classified separately because they could not be identified consistently across patches and Grad-CAM provided only approximate spatial localization. Formal categorization was performed independently using anonymized patch identifiers and images, without access to model predictions, prediction probabilities, prediction correctness, patient outcomes, or other clinical information. Based on the average AUROCs across the internal and external sets, three best-performing models from the 10-fold ensembles were selected to visualize for their attention tendencies. Semi-quantitative summaries of attention category distributions were calculated separately for each model and reported as mean proportions with standard deviations. Native ACMIL attention weights from the ACMIL-based WSI pipelines were also directly visualized on WSIs for qualitative inspection of any potentially unique attention patterns.

Multimodal analysis for clinical metadata integration

To evaluate the potential of multimodal analysis, logistic regression models were constructed using predicted probabilities from the WSI and LE histology models, with or without maternal age (Age), body mass index (BMI), endometrial thickness (EnTh), and biopsy-day serum estradiol (E2) and progesterone (P4) levels within the internal test set. EnTh was determined via transvaginal ultrasound prior to the biopsy for sample collection. Five logistic regression settings were evaluated including: a histology-only model integrating WSI and LE models’ probabilities; an integrative seven-factor histology-clinical model additionally incorporating Age, BMI, EnTh, E2 and P4; and three clinical metadata-only models incorporating Age, BMI and EnTh; E2 and P4; or all five clinical variables. To maximize training efficiency while preventing data leakage, we leveraged the 10-fold cross-validation framework of the histology models to generate out-of-fold predictions for the entire training cohort. These out-of-fold probabilities served as the input features for training the logistic regression models. Multimodal analyses containing histology-based probabilities as inputs were restricted to samples with both WSI and LE model outputs. To examine feature contribution patterns in the most integrative seven-factor model, standardized coefficients weights were calculated by multiplying each raw coefficient by the feature’s standard deviation, allowing the presented coefficients weights to be directly comparable. SHapley Additive exPlanations (SHAP) summary plots were additionally generated for the internal cohorts to summarize the direction and distribution of each feature’s contribution to the predictions.

Computational software packages

The ImageNet 1K-pretrained ResNet-18 (https://download.pytorch.org/models/resnet18-5c106cde.pth) was trained and the histopathology pretrained UNI2-h foundation model (https://huggingface.co/MahmoodLab/UNI2-h) was utilized directly in this study. The UNI2-h model weights were accessed via application through an institutional email-linked HuggingFace account as requested. Training of models within the study and Grad-CAM visualization largely relied on functions and tools available from the PyTorch machine learning library on Python (PyTorch version = 2.4.1; Python version = 3.10.0) [39]. Logistic regression models for multimodal analysis were implemented using scikit-learn (version = 1.7.1), and SHAP-based model interpretation was performed using the SHAP package (version = 0.49.1). QuPath (version = 0.5.1) was the primary WSI inspection software. Trident (version 0.2.0) is a toolkit developed by the same group that created the UNI2-h [37]. The toolkit was used in all our reported ResNet-18 and UNI2-h-related patch preparation tasks, with Tiff patch images generated for the ResNet-18 workflow and patch-level feature vectors extracted for the UNI2-h-related workflow. Patch image generation for the LE-focused model was performed using QuPath and a Groovy script. DeLong tests were implemented following the fast DeLong algorithm described by Sun and Xu [40].

Model performance evaluation and reported metrics

The performances of the developed 10-fold ensemble models were evaluated based on three common machine learning performance metrics, including accuracy, balanced accuracy, and area under the receiver operating characteristic curve (AUROC). To monitor training stability and convergence among different architectures and classifier heads, performance trajectories were tracked across all 10 folds and reported as the mean of available folds per epoch to account for varying early-stopping points. To address the potential misleading interpretation of accuracy considering the study’s test sample sets were not in 1:1 balanced ratio, balanced accuracy was utilized to define best epochs during training and validation sessions.

Reported models were compared in two major settings using DeLong test, including with a chance-level random classifier with AUROC = 0.5, and paired model-to-model comparisons. Comparison between the WSI and the LE model was restricted to samples with predictions from both pipelines to ensure paired testing. Reported P values were adjusted using the Benjamini-Hochberg procedure within each cohort and comparison group, with statistical significance defined as an adjusted P < 0.05. To supplement AUROC-based discrimination and threshold-dependent classification metrics, model calibration was assessed using calibration curves and Brier scores. Each model was analyzed separately in the internal cumulative-live-birth testing cohort and the external phase-based testing cohort.

Results

Limited transferability of the HRT model to natural cycle WSIs

To evaluate the potential transferability of the existing deep learning model, we first assessed the performance of the reported HRT model [27] using the natural cycle LH + 7 WSIs archived in our centers. The input settings of all the patches, including patch resolution, dimensions, and cropping procedures, were prepared to match its published HRT model. Despite the currently applied input-level standardization, the model yielded an AUROC of 0.55 ± 0.08 and a balanced accuracy of 55.85 ± 1.37%, indicating a limited discriminative capability in this natural cycle cohort (S1 Fig; S3 Table).

Natural cycle endometrial WSI-trained new models for live birth outcome prediction

We next explored new models using a new WSI workflow for the natural cycle dataset (Fig 2A). We first retrained a new model based on the same ResNet-18 architecture by incorporating adjusted normalization parameters and augmentation strategies tailored to the new dataset. To investigate whether feature representations learned from large pan-cancer datasets could be transferable to a non-cancerous endometrium, live birth-oriented task, we also implemented the UNI2-h foundation model as a frozen feature extractor, evaluating its performance across different aggregation strategies (GAP vs. ACMIL) and classifier heads (linear vs. MLP) with detailed performance metrics summarized in Table 1.

Fig 2. WSI endometrial live birth outcome prediction models.

Fig 2

(A) Schematic diagram showing the training workflow of the WSI models developed in this study. (B-D) ROC curves illustrating the performances of internal and external phase-based test sets of ResNet-18 model (B), UNI2-h GAP-Linear and GAP-MLP models (C) and UNI2-h ACMIL-Linear and UNI2-h ACMIL-MLP models (D). Figure partially created in BioRender [https://biorender.com/f125abi].

Table 1. Performance metrics of WSI deep learning models for endometrial live birth outcome prediction.

Performance Metric ± SD ResNet-18 WSI UNI2-h

GAP-Linear WSI
UNI2-h

GAP-MLP WSI
UNI2-h

ACMIL-Linear WSI
UNI2-h

ACMIL-MLP WSI
Internal Testing Set
Sensitivity 0.727 ± 0.105 0.636 ± 0.315 0.727 ± 0.141 0.636 ± 0.200 0.636 ± 0.148
Specificity 0.800 ± 0.136 0.600 ± 0.235 0.700 ± 0.166 0.650 ± 0.197 0.750 ± 0.137
PPV 0.667 ± 0.165 0.467 ± 0.145 0.571 ± 0.073 0.500 ± 0.169 0.583 ± 0.160
NPV 0.842 ± 0.041 0.750 ± 0.106 0.824 ± 0.075 0.765 ± 0.086 0.789 ± 0.052
F1 0.696 ± 0.055 0.538 ± 0.184 0.640 ± 0.050 0.560 ± 0.073 0.608 ± 0.075
Accuracy 0.774 ± 0.063 0.613 ± 0.056 0.710 ± 0.071 0.645 ± 0.071 0.710 ± 0.068
Balanced Accuracy 0.764 ± 0.042 0.618 ± 0.057 0.714 ± 0.046 0.643 ± 0.045 0.693 ± 0.060
AUC 0.836 ± 0.025 0.627 ± 0.066 0.741 ± 0.030 0.727 ± 0.031 0.750 ± 0.066
External Test Set
Sensitivity 1.000 ± 0.365 0.778 ± 0.292 0.556 ± 0.216 0.444 ± 0.333 0.556 ± 0.235
Specificity 0.083 ± 0.309 0.583 ± 0.377 1.000 ± 0.121 0.917 ± 0.382 1.000 ± 0.085
PPV 0.450 ± 0.177 0.583 ± 0.197 1.000 ± 0.168 0.800 ± 0.320 1.000 ± 0.316
NPV 1.000 ± 0.409 0.778 ± 0.352 0.750 ± 0.106 0.688 ± 0.252 0.750 ± 0.086
F1 0.621 ± 0.209 0.667 ± 0.167 0.714 ± 0.149 0.571 ± 0.198 0.714 ± 0.259
Accuracy 0.476 ± 0.090 0.667 ± 0.142 0.810 ± 0.094 0.714 ± 0.109 0.810 ± 0.105
Balanced Accuracy 0.542 ± 0.093 0.681 ± 0.119 0.778 ± 0.103 0.681 ± 0.083 0.778 ± 0.119
AUC 0.444 ± 0.140 0.815 ± 0.164 0.898 ± 0.051 0.778 ± 0.075 0.870 ± 0.095

PPV = positive predictive value; NPV = negative predictive value; AUC = area under the receiver operating characteristic curve

1. ResNet-18 model with end-to-end training.

The ResNet-18 ensemble model trained on the internal natural cycle WSI dataset (ResNet-18 WSI model) demonstrated strong predictive capability when evaluated on the internal testing set (AUROC = 0.84 ± 0.03; balanced accuracy = 76.36 ± 4.23%) (Fig 2B). However, its performance on the external phase-based testing cohort collapsed rapidly during training, resulting in a final ensemble model achieving an AUROC of only 0.44 ± 0.14 and a balanced accuracy of 54.17 ± 9.31%, representing a poorer performance towards the external phase-based testing setting (Fig 2B; S2A Fig). DeLong test further indicated that the internal testing AUROC was significantly above that of a chance-level random classifier (P < 0.001), while the external phase-based testing AUROC remained not significantly different from chance level (P = 0.738).

2. Frozen UNI2-h feature extractor with global average pooling.

The ensemble model utilizing the frozen UNI2-h feature extractor with a standard linear classifier head (UNI2-h GAP-Linear WSI model) indicated a limited overall predictive capability (Fig 2C). For the internal outcome-based testing cohort, the model achieved a non-significant AUROC of 0.63 ± 0.07 (P = 0.258) with a balanced accuracy of 61.82 ± 5.74%. A higher performance was observed in the external phase-based testing cohort, where the model has yielded a significant AUROC of 0.82 ± 0.16 (P = 0.002) with a balanced accuracy of 68.06 ± 11.86%.

Subsequently, the impact of coupling the feature extractor with a non-linear MLP classifier head (UNI2-h GAP-MLP WSI model) was evaluated (Fig 2C). While training trajectories across the internal cohorts remained generally consistent between both classifiers, the MLP variant demonstrated more rapid convergence on the external cohort during the training session (S2B Fig). For the internal testing cohort, the model achieved a significant AUROC of 0.74 ± 0.03 (P = 0.034) with a balanced accuracy of 71.36 ± 4.61%. Upon external phase-based testing, the model also yielded a significant AUROC of 0.90 ± 0.05 (P < 0.001) with a balanced accuracy of 77.78 ± 10.33%.

3. Frozen UNI2-h feature extractor with weighted attention aggregation.

The application of the ACMIL aggregation strategy was further evaluated with either a standard linear classifier head or an MLP head (Fig 2D). On the internal testing cohort, the model coupled with a linear head (UNI2-h ACMIL-Linear WSI model) achieved a significant AUROC of 0.73 ± 0.03 (P = 0.034) with a balanced accuracy of 64.32 ± 4.52%. The model exhibited a comparable generalizability on the external phase-based test set, which yielded an AUROC of 0.78 ± 0.08 (P = 0.012) with a balanced accuracy of 68.06 ± 8.31%.

The model with ACMIL aggregation strategy utilizing the non-linear MLP classifier head (UNI2-h ACMIL-MLP WSI model) demonstrated significant and elevated AUROC values compared to random (Fig 2D). For the internal outcome-based testing cohort, this approach achieved an AUROC of 0.75 ± 0.07 (P = 0.031) with a balanced accuracy of 69.32 ± 6.03%. Upon external phase-based testing, the model maintained generally comparable to its GAP-based counter version, with an AUROC of 0.87 ± 0.09 (P < 0.001) and a balanced accuracy of 77.78 ± 11.87%. Model’s training trajectories did not indicate a clear advantage in convergence or stability for the ACMIL implementation compared with GAP (S2C Fig).

To further investigate the attention-weighting patterns of the ACMIL framework, we visualized the attention weights directly on the WSIs to identify whether particular histological regions were consistently prioritized by the model. Qualitative visual inspection revealed relatively sparse attention distribution without a strong and clear concentration towards a specific, recognizable morphological structure (S3 Fig).

Performance of deep learning models trained on LE-containing regions

To address the concern that WSI-based approaches may dilute the existence of LE features, we developed biologically-targeted models trained exclusively on manually annotated LE regions (Fig 3A). Predictive performance was assessed across two distinct models trained with the LE-containing patches (namely ResNet-18 LE and UNI2-h LE models), with detailed performance metrics summarized in Table 2. Compared with the ResNet-18 WSI model, the ResNet-18 LE model has demonstrated a less divergent AUROC pattern across the internal outcome-based and external phase-based testing cohorts (Fig 3B; S2D Fig). For the internal testing set, the ResNet-18 LE model achieved a non-significant AUROC of 0.71 ± 0.03 (P = 0.090), with a balanced accuracy of 55.71 ± 5.75%. In the external phase-based testing cohort, the model achieved a significant AUROC of 0.80 ± 0.16 (P = 0.033) and a balanced accuracy of 70.00% ± 12.30%.

Fig 3. LE-targeted endometrial live birth outcome prediction models.

Fig 3

(A) Schematic diagram showing the establishment of deep learning models trained on LE-containing regions. (B, C) ROC curves comparing model performance on the internal outcome-based (left) and external phase-based (right) testing cohorts for the LE-targeted ResNet-18 model (B) and LE-targeted UNI2-h model with an MLP classifier head (C). (D) Representative Grad-CAM visualizations from correctly predicted CLB and unsuccessful outcome cases. Higher attention areas are illustrated in warmer colours and lower attention areas are illustrated in cooler colours. (E) Semi-quantitative analysis of the average attention tendency of the models. Data are presented as the averaged percentage composition of attention categories from top-3 performing models within the 10-fold ensemble [(LE alone (blue), LE+stroma (purple), stroma alone (red), and unclassified (grey)], presented in Mean ± SD. Figure partially created in BioRender [https://biorender.com/f125abi].

Table 2. Performance metrics of LE-focused deep learning models for endometrial live birth outcome prediction.

Performance Metric ± SD ResNet-18 LE UNI2-h LE
Internal Testing Set
Sensitivity 0.400 ± 0.181 0.900 ± 0.175
Specificity 0.714 ± 0.136 0.643 ± 0.158
PPV 0.500 ± 0.098 0.643 ± 0.090
NPV 0.625 ± 0.081 0.900 ± 0.094
F1 0.444 ± 0.086 0.750 ± 0.097
Accuracy 0.583 ± 0.052 0.750 ± 0.081
Balanced Accuracy 0.557 ± 0.057 0.771 ± 0.080
AUC 0.707 ± 0.027 0.743 ± 0.086
External Test Set
Sensitivity 1.000 ± 0.159 0.750 ± 0.281
Specificity 0.400 ± 0.320 0.800 ± 0.151
PPV 0.571 ± 0.133 0.750 ± 0.076
NPV 1.000 ± 0.411 0.800 ± 0.147
F1 0.727 ± 0.083 0.750 ± 0.172
Accuracy 0.667 ± 0.142 0.778 ± 0.079
Balanced Accuracy 0.700 ± 0.123 0.775 ± 0.094
AUC 0.800 ± 0.164 0.837 ± 0.063

PPV = positive predictive value; NPV = negative predictive value; AUC = area under the receiver operating characteristic curve

The UNI2-h LE model also yielded performance comparable to its WSI variant (Fig 3C; S2D Fig). On the internal outcome-based testing cohort, the model achieved a significant AUROC of 0.74 ± 0.09 (P = 0.034) with a balanced accuracy of 77.14 ± 7.98%. Upon external phase-based testing, the model’s AUROC maintained at 0.84 ± 0.06 (P = 0.002) with a balanced accuracy of 77.50 ± 9.39%.

To further examine the spatial localization of model attention, Grad-CAM was utilized to visualize the high-attention regions driving the predictions of the three top-performing ResNet-18 LE models within the 10-fold ensemble, defined by their averaged AUROCs (0.79, 0.78, and 0.77). The resulting heatmaps were manually categorized according to the predominant localization of the highlighted compartments: (1) LE alone (attention primarily on LE), (2) mixed (attention spanning both LE and peri-LE stroma without clear preference), and (3) peri-LE stroma alone (attention primarily on peri-LE stroma), and (4) unclassified regions, including luminal spaces and potential artifacts (Fig 3D). Semi-quantitative analysis of the top-3 models showed that regions of higher attention were predominantly associated with LE-related compartments rather than peri-LE stroma-alone or unclassified regions. When stratified by predicted pregnancy outcome, patches showing LE alone and mixed attentions remained the two major categories in both CLB and unsuccessful cases, accounting for 44.33 ± 9.98% and 39.03 ± 5.20% of patches in CLB cases, and 38.83 ± 8.02% and 38.25 ± 7.34% in unsuccessful cases, respectively (Fig 3E). In contrast, patches showing high attention on peri-LE stroma alone were less frequent, accounting for 13.75 ± 4.92% and 16.37 ± 1.31% of patches in CLB and unsuccessful cases, respectively. Unclassified patches with attention restricted to empty luminal spaces lacking clear biological relevance, accounted for only a minor proportion in both groups (2.89 ± 2.68% and 6.55 ± 1.89% for CLB and unsuccessful cases, respectively). These results highlighted that Grad-CAM’s localization-based visualization was more focused on LE region.

Comparisons of histology model performances

Paired DeLong tests were used for pairwise comparisons among the individual histology model configurations (S4 Fig). In the internal outcome-based testing cohort, none of the comparisons reached statistical significance after Benjamini-Hochberg adjustment. The two largest internal AUROC differences were between specific model configurations. ResNet-18 WSI and UNI2-h GAP-MLP WSI differed by ΔAUROC = 0.096, comparing end-to-end training with a foundation model-based pipeline. UNI2-h GAP-Linear WSI and UNI2-h GAP-MLP WSI differed by ΔAUROC = 0.114, comparing linear and MLP classifier heads. The remaining internal comparisons showed smaller differences, with the smallest difference observed between UNI2-h GAP-MLP WSI and UNI2-h ACMIL-MLP WSI (ΔAUROC = 0.009), comparing GAP and ACMIL aggregation. In the external phase-based testing cohort, UNI2-h GAP-MLP WSI achieved a significantly higher AUROC than ResNet-18 WSI, comparing the foundation model-based pipeline with end-to-end training (ΔAUROC = 0.454, P = 0.004). The remaining external comparisons were not statistically significant, with ΔAUROC ranging from 0.028 to 0.093. The smallest external difference was again observed between UNI2-h GAP-MLP WSI and UNI2-h ACMIL-MLP WSI (ΔAUROC = 0.028).

Model calibration for probability prediction

Calibration curves and Brier scores were presented for each histology model separately in the internal outcome-based testing cohort and the external phase-based testing cohort (S5 Fig). The ResNet-18 WSI model showed the largest internal-external discrepancy, with the lowest Brier score in the internal testing cohort (0.161) but the highest Brier score in the external phase-based testing cohort (0.331), consistent with its divergent discriminative performance across the two testing settings. Among the UNI2-h WSI feature-based pipelines, Brier scores were generally within a narrower range. The UNI2-h ACMIL-MLP WSI model showed the lowest internal Brier score among the UNI2-h WSI models (0.205), while the UNI2-h GAP-MLP WSI model showed the lowest external Brier score (0.143). For the LE-focused models, the ResNet-18 LE model yielded Brier scores of 0.224 and 0.239 in the internal and external testing cohorts, respectively, while the UNI2-h LE model yielded Brier scores of 0.213 and 0.193, respectively.

Multimodal integration of histology-based model outputs and clinical variables

To assess the potential benefit of multimodal integration, logistic regression models were constructed using predicted probabilities from the UNI2-h GAP-MLP WSI and UNI2-h LE pipelines on the internal outcome-based cohort, with or without maternal age (Age), BMI, endometrial thickness (EnTh), and E2 and P4 levels (detailed metrics in S4 Table). A model integrating the WSI and LE outputs achieved an AUROC of 0.80 (P = 0.006) with a balanced accuracy of 63.57% (Fig 4A). When Age, BMI, EnTh, E2, and P4 were further incorporated, the integrative seven-factor model achieved an AUROC of 0.77 (P = 0.034) with a balanced accuracy of 70.71% (Fig 4B). The clinical metadata-only model incorporating Age, BMI, EnTh, E2, and P4 yielded a non-significant AUROC of 0.52 (P = 0.903) with a balanced accuracy of 52.73% (Fig 4C). Two simpler Age + BMI + EnTh and E2 + P4 clinical metadata-only models were also evaluated, with non-significant AUROCs of 0.541 and 0.532, respectively (S6A, S6B Fig). In paired comparisons, both the WSI + LE model and the integrative seven-factor model achieved significantly improved discriminatory ability than the five-factor clinical metadata-only model (P = 0.026 for both comparisons). No significant difference was reached between the seven-factor and the WSI + LE models (P = 0.607).

Fig 4. Multimodal integration of histology-based models with clinical and hormonal variables.

Fig 4

ROC curves for the histology-only WSI + LE model (A), the integrative seven-factor LE + WSI + Age + BMI + EnTh + E2 + P4 model (B), and the Age + BMI + EnTh + E2 + P4 five-factor clinical metadata-only model (C), evaluated in the internal outcome-based testing cohort. (D) Standardized coefficients of the integrative seven-factor logistic regression model, showing the relative direction and magnitude of each feature’s contribution to the model output.

To further examine feature weighting patterns within the multimodal framework, standardized logistic regression coefficients and SHAP-based feature impacts were visualized. The standardized coefficients showed that the WSI model probability (0.529) and LE model probability (0.432) had the largest positive contributions to the model output, followed by E2 (0.253) and the maternal age with a negative weight (-0.213). Compared to the other features, P4 (0.052), EnTh (0.037), and BMI (0.028) contributed minimally with relatively lower impact (Fig 4D). Consistent with this pattern, SHAP analysis from either internal testing (S6C Fig) or training cohort (S6D Fig) demonstrated a broader distribution of feature impact for the WSI- and LE-based probabilities, smaller effects from E2 and maternal age, and minimal effects from P4, EnTh, and BMI. Together, these results summarized the relative contribution patterns of histology-based and clinical inputs within the multimodal framework.

Discussion

This study explored the application of deep learning to predict cumulative live birth outcomes using human endometrial H&E WSIs. The reported ImageNet 1K-pretrained HRT model [27] was initially tested with our natural cycle endometrial WSIs. Even though patch preparation was matched to its published settings, the model showed poor predictive power. This limited transferability likely reflected the residual domain shift, biological differences between HRT and natural-cycle endometrium, and sample-size uncertainty, rather than implementation settings alone. The broader medical-imaging literature similarly emphasizes that model reliability is affected by pipeline factors, including sample preparation, patch generation, preprocessing, and validation designs [41,42]. Such domain shift might have been exacerbated by the small development cohort [27], further compromising stability and cross-institutional practicality. HRT samples form distinct sub-clusters from natural-cycle counterparts, with significant transcriptional shifts in the endometrial epithelial populations [43]. Histological studies similarly demonstrates HRT affects the thickness and the sizes of glandular and vascular space [29,44], further widening the discrepancy between the reported model and the test set.

In this study, two model development strategies were evaluated for cumulative live birth prediction using samples from natural cycles. The CNN-based ResNet-18 model was utilized as an end-to-end training pipeline due to its efficient and well-established architecture for medical image recognition tasks [45]. ResNet-18’s simplicity also made it practical for limited-data settings, where overfitting is a major concern [26,46]. In parallel, UNI2-h [21,47] was selected as one of the newest publicly accessible histopathology-pretrained foundation models at the time of our model development. Its ViT-Huge architecture and large-scale histopathology pretraining enable more complex histological inference. Nevertheless, its pretraining focused predominantly on cancer data, but not endometrial analysis or fertility outcomes. Thus, we evaluated whether frozen UNI2-h could be repurposed for non-malignant endometrial live-birth prediction.

The two approaches exhibited different performance tendencies. While the ResNet-18 model trained on internal natural cycle samples achieved relatively high predictive performance within the internal outcome-based testing cohort, its performance collapsed to near-random levels in the external phase-based testing cohort. This discrepancy may arise from two non-mutually exclusive factors. First, despite standardizing patches by pixel dimensions and physical coverage based on scanner-specific mpp metadata, residual upstream differences like tissue processing, sectioning, staining, scanner platforms may still have introduced institution-specific image characteristics [41,42]. Such effects may particularly impact end-to-end training, which can learn center-specific patterns from limited single-center data. Second, the external cohort lacked IVF outcome-based validation, instead evaluating receptive LH + 7 versus non-LH + 7 phase discrimination. The performance gap may reflect cross-institutional variation and limited target transferability, potentially similar to the previously assessed HRT model. Further outcome-matched validation with stronger statistical power is required.

In comparison, the UNI2-h feature-based pipeline completely lacks the ability to fine-tune its feature extractor’s logic. Despite, the pipeline eventually demonstrated models with less divergent performances across the internal outcome-based and external phase-based testing cohorts. With the standard linear-probe design, the UNI2-h GAP-Linear underperformed on the internal outcome-based test cohort, contrasting with the effective linear-probe performance reported in the original UNI benchmark [21]. Notably, healthy superficial endometrium and the subtle fertility-associated morphological features are likely underrepresented in UNI2-h’s knowledgebase of training. Li et al. [48] proposed that linear probes may be insufficient when the target differs substantially from the pretraining patterns. We therefore implemented a non-linear MLP head to better adapt UNI2-h-derived features to the endometrial fertility-oriented task. As expected, MLP-head models achieved higher AUROCs than their linear-head counterparts. Specifically, the UNI2-h GAP-MLP achieved significance in both internal and external cohorts, unlike the linear variant (external-only). This confirmed that frozen UNI2-h representations with MLP head are feasible for non-malignant endometrial live birth prediction without feature-extractor modification. Nevertheless, DeLong comparisons showed no significant difference between classifier heads, suggesting potential MLP utility without establishing statistical superiority given the current testing sample size. External generalizability for live birth prediction would require larger, outcome-matched cohorts. Beyond classifier-head adaptation, pretrained model selection may also influence downstream performance. Recent representational-similarity analysis showed that pathology foundation models (e.g., UNI2-h and Virchow2) extract divergent representations from the same histological images despite similar ViT-Huge architectures [49]. Whole-slide pretrained models, such as TITAN and Prov-GigaPath, further broaden the available strategies by incorporating broader slide-level context [50]. Thus, future comparisons across different feature-extraction strategies are essential to identify the optimal starting point for endometrial live birth-oriented analysis.

The study further implemented the ACMIL attention strategy [36] to evaluate whether prioritizing informative WSI regions prevents the dilution of critical features, a method known to improve performance in cancer-oriented studies [21,23,51]. However, comparable AUROC performance was achieved in GAP-MLP and ACMIL-MLP-based models, indicating ACMIL did not clearly outperform GAP. Calibration analysis provided a probability-level perspective of model performance beyond binary discrimination [52]. The UNI2-h ACMIL-MLP model maintained relatively low Brier scores across both testing cohorts, whereas the UNI2-h GAP-MLP model showed the highest internal Brier score. The differential weighting mechanism of ACMIL likely down-weighted the less-informative patches, making it potentially more reliable for probabilistic estimation than global pooling. Therefore, the ACMIL-MLP design can be preferable in clinical settings requiring patient-level probability estimates. For cumulative live birth prediction, larger-cohort comparisons should evaluate the most suitable aggregation approach. Post hoc recalibration may further be applied to improve the reliability of probability estimates used for individualized risk counselling regarding the likelihood of live birth [53].

ACMIL attention was sparse without clustered or consistent concentration within a specific histological compartment. This may be compatible with live birth prediction, where decision-relevant morphology may be more spatially distributed than in cancer-detection tasks [23]. Therefore, sparse weighting may reflect endometrial biopsy heterogeneity and not biologically uninformative. For example, asynchronous glands with mixed menstrual-stage epithelia have been observed in otherwise normal endometrial biopsies [54–56], which complicate phase assessment. Such within-slide heterogeneity may likewise affect attention interpretation in this fertility-oriented task. Further improvement requires deeper biological understanding of endometrial morphology and quantification of high-attention regions. One practical extension would be to quantify the interpretable histomics-centred features, such as nuclear morphology, cellular density, glandular architecture, and spatial cell-type relationships [57] in ACMIL-weighted patches, assessing whether high-attention regions are statistically associated with measurable morphological characteristics beyond visual inspection. Recent studies have demonstrated the use of AI-assisted endometrial cell annotation and targeted morphological analysis with explainable metrics such as epithelial-to-stromal ratio [58] and peri-gland stroma nuclear size [59]. However, no broadly available and cross-institutionally validated model yet reliably annotates LE, glandular epithelium, and stroma, while manual endometrial WSIs annotation remains impractical. Developing compartment-aware endometrial annotation models alongside outcome-oriented deep learning prediction models would enable objective attention-weighted histomics, interpretable handcrafted-feature baselines, and scalable ROI-based analysis. Complementing morphology and histomics-based approaches, recent work has aligned AI prediction heatmaps with Visium spatial transcriptomic profiles to validate molecular features associated with model-predicted regions [60]. Applying a similar strategy in endometrium could help determine whether attention patterns correspond to distinct cell states and potentially gene-expression profiles.

The LE-focused pipeline of analysis represents a manually defined, biologically informed ROI-based modelling strategy. The LE is the primary maternal-embryo interface during early implantation. Mouse in vivo studies, have highlighted that the LE undergoes morphological reshaping for implantation, affecting subcellular features including apical microvilli protrusions, apical tight junctions, and basolateral membranes [61–63]. However, capturing LE within a standard automated WSI patch generation pipeline is challenging because LE is sparse as compared to other cell types in endometrium. While only a minority of patients lacked identifiable LE, our dataset of LE-containing patches was substantially smaller than that of the standard WSI-based patches, with training set to be approximately 25 times smaller. The actual existence of LE within the WSI pipeline could be further compromised since the LE naturally has higher chance of falling into edge-case tiles dominated by empty spaces, which were discarded during the upstream patch generation procedures.

Beyond the LE, the adjacent peri-LE stroma may also represent a potentially underrepresented region worth specific emphasis. Histologically, the dense superficial stroma directly beneath the LE is the stratum compactum, whereas the deeper looser region is the stratum spongiosum [64]. Peri-LE stroma remained understudied until recent spatial transcriptomics suggests stromal WNT and NOTCH expressions may vary with distance from the luminal surface [65]. Therefore, our LE-containing ROI design was intended to enrich both LE and the potentially unique peri-LE stroma for separate evaluation. Surprisingly, despite the massive disparity in data volume of learning, models trained exclusively on LE-containing regions have achieved performance generally comparable to the WSI models. Together with the established biological roles of the LE and adjacent peri-LE stroma, these findings align with the expectation that this compartment represents a biologically enriched, morphologically informative reservoir for fertility-oriented endometrial modelling.

To visualize the morphological regions that were likely associated with model predictions, we employed Grad-CAM on the three best-performing LE-focused ResNet-18 models. This patch-level attention visualization was not extended to UNI2-h feature-based pipeline because reliable pixel-level saliency was difficult under slide-level feature aggregation. Semi-quantitative review showed that attention more often targeted LE-containing regions, either mainly the LE compartment or jointly with adjacent peri-LE stroma, while peri-LE stroma alone was less frequent. Unclassified artifacts or lumen-dominant regions were relatively rare. Thus, this localization tendency was consistent with the intended LE-focused design. Nevertheless, Grad-CAM provides only approximate spatial localization and does not specify the underlying mechanistic drivers. Reliable automated endometrial compartment annotation, together with more standardized attention quantification methods, would be essential for broader implementation of this interpretability framework and for more informative assessment of attention tendencies. Potentially, a multiscale design allowing more variable peri-LE context may also improve model robustness [57,66], as the physical boundary of the peri-LE stroma compartment has not yet been clearly defined by a fixed depth or anatomical cutoff [67,68].

Finally, multimodal integration was performed to explore the potential benefit of combining two histology-based models covering different regions of focus, with maternal clinical (age, BMI, and EnTh) and hormonal variables (E2 and P4). Such multimodal designs are increasingly explored in studies aiming to develop clinically useful predictive models [69]. Both WSI + LE and seven-factor models significantly outperformed chance and the clinical metadata-only model (Age + BMI + EnTh + E2 + P4). Standardized coefficients and SHAP analyses consistently ranked the two histology-model probabilities as the most impactful features. E2 and maternal age had small positive and negative associations, respectively; P4, EnTh, and BMI contributed minimally.

The inverse maternal age-live-birth association aligned with the known age-related decline [70]. The positive E2 and limited P4 associations were biologically plausible, consistent with a study that midluteal E2 levels were significantly associated with live birth rates once a minimum P4 level is reached [71]. The association between BMI and live birth remains cohort-dependent and controversial [72,73]. Limited P4 and EnTh weights do not imply biological irrelevance. Both variables are known to have minimal thresholds pregnancy supports [9,71,74]; thereby, a simpler threshold-based implementation may be more appropriate. Overall, these findings highlight the dominant contribution of histology-based predictions while demonstrating feasibility of incorporating complementary clinical variables. Future studies may explore earlier-stage multimodal integration to capture non-linear morphological–clinical relationships [57].

This study has several limitations. First, the patient-level sample size remained small although our cohort was larger than that of our previous HRT model [27]. Freezing the UNI2-h feature extractor reduced training complexity, but the low testing cohorts still weaken performance stability. Larger prospective cohorts are required to confirm the most suitable modelling strategy before wider implementation. Second, the external cohort was outcome-mismatched. Therefore, external performance should not be interpreted as direct validation of live birth prediction. Outcome-matched multicenter cohorts incorporating harmonization approaches (e.g., StainGAN) [75] are required to evaluate cross-center generalizability and mitigate technical variations. Third, the interpretability analyses remain limited to approximate spatial localization and could not support precise cellular or mechanistic interpretation from Grad-CAM saliency maps alone [76,77]. The manually defined LE workflow and lack of a validated endometrial compartment annotator further constrained scalable histomic analysis and the inclusion of a handcrafted morphological baseline for comparison.

From a clinical perspective, this predictive framework is best positioned as a supportive tool for pathologists and IVF clinicians when considering whether and when to proceed with embryo transfer. Unlike conventional dating, which is limited by biological and observer variability, this outcome-oriented pipeline offers a standardized, extendable, and quantifiable use of routine H&E images. On the other hand, the high cost of existing tests, such as transcriptomic array-based ERA, which costs US$800 or more per test, renders large prospective studies financially challenging [78]. The relatively lower cost of the proposed H&E image-based framework could practically facilitate larger validation and subsequent broader clinical implementation. Integrating biologically informed ROI modelling with automated histomic quantification could further enhance its comprehensiveness and interpretability. Of note, false positives could provide undue reassurance before embryo transfer, while false negatives could cause unnecessary delay or extra assessment. Thus, model thresholds should align with clinical priorities and undergo prospective validation. The clinical value of this framework relative to existing receptivity assessments also requires separate evaluation. With prospective validation and improved biological interpretation, it may enable more objective and personalized embryo-transfer planning.

In conclusion, this study showed that the histopathology pretrained UNI2-h foundation model could be effectively repurposed for endometrial fertility-related analysis without task-specific end-to-end training of the feature extractor. Although an endometrium-specific foundation model can be theoretically more ideal, such datasets are difficult to obtain in practice. Further outcome-matched external validation will be needed to formally assess generalizability for live-birth prediction. The LE-focused analyses further highlighted that a biologically informed design could retain performance comparable to WSI-based approaches despite using substantially fewer patches during development. Together, these findings support the LE and peri-LE regions and repurposed pathology models with ROI designs as promising starting points for endometrial histology analysis.

Supporting information

S1 Fig. HRT model performance on natural cycle samples.

An ROC curve showing the performance of the reported 3-fold ensemble HRT-model on the natural cycle LH + 7 WSIs.

(TIF)

pdig.0001744.s001.TIF (352.5KB, TIF)
S2 Fig. Summary of model predictive performance trajectories during training sessions.

Summary figures illustrate progression of model predictive performance in AUROC (y-axis) over training epochs (x-axis). Performance is displayed for three cohorts including validation, internal testing, and external testing cohorts. All discussed models‘ performance trajectories were presented, including (A) ResNet-18 WSI model trajectory represented in blue solid line; (B) UNI2-h GAP-Linear and UNI2-h GAP-MLP model trajectories presented in blue and red solid lines; (C) UNI2-h ACMIL-Linear and UNI2-h ACMIL-MLP model trajectories presented in blue and red solid lines; (D) ResNet-18 LE and UNI2-h LE model trajectories presented in blue and red solid lines, with their WSI counter variants (ResNet-18 WSI and UNI2-h GAP-MLP model trajectories) plotted in purple and yellow dotted lines for easier comparison.

(TIF)

S3 Fig. Visualization of attention distributions in the ACMIL-based UNI2-h model.

Representative heatmaps display patch-level attention tendency on WSIs for a correctly predicted CLB case (A) and Unsuccessful (B) case. Warmer colors indicate higher attention weights.

(TIF)

pdig.0001744.s003.TIF (3.6MB, TIF)
S4 Fig. Histology-based model pairwise performance comparison summary.

Summary of model performance comparisons between specified models, specifically in the internal outcome-based (A) and the external phase-based (B) testing cohorts. Δ value presents difference of AUROC between model A and B within each specified pair. Formal statistics performed with DeLong test corrected for multiple comparison (p < 0.05 indicated with *).

(TIF)

pdig.0001744.s004.TIF (737.9KB, TIF)
S5 Fig. Summary of per model calibration metrics.

Calibration curves presented for each single reported models, including (A) ResNet-18 WSI, (B) UNI2-h GAP-Linear WSI, (C) UNI2-h GAP-MLP WSI, (D) UNI2-h ACMIL-Linear WSI, (E) UNI2-h ACMIL-MLP WSI, (F) ResNet-18 LE, and (G) UNI2-h LE model. Solid line presents model’s calibration curve on the internal test set; dotted line presents model’s calibration curve on the external set.

(TIF)

S6 Fig. Additional multimodal model evaluations and feature-impact analysis.

ROC curves for the Age + BMI + EnTh clinical metadata-only model (A) and the hormone-only E2 + P4 model (B). SHAP summary plot showing feature-contribution patterns across the internal test (C) and out-of-fold development samples (D) for the integrative seven-factor model. Positive SHAP values indicate a shift toward live birth prediction, negative SHAP values indicate a shift toward unsuccessful outcome prediction. Each dot represents one sample, with colour representing each samples’ original feature value, with red indicating a higher and blue representing a lower original value.

(TIF)

pdig.0001744.s006.TIF (867.8KB, TIF)
S1 Table. Summary of clinical characteristics of the internal cohort by outcome group.

(XLSX)

pdig.0001744.s007.xlsx (8.3KB, xlsx)
S2 Table. Model training hyperparameter setting summary.

(XLSX)

pdig.0001744.s008.xlsx (9.4KB, xlsx)
S3 Table. Evaluation of the HRT-trained ResNet-18 model on the internal natural cycle cohort.

(XLSX)

pdig.0001744.s009.xlsx (8.3KB, xlsx)
S4 Table. Performance metrics of logistic regression models with multimodal integration.

(XLSX)

pdig.0001744.s010.xlsx (8.6KB, xlsx)

Data Availability

Sensitive raw clinical data and images are prohibited to be shared publicly due to legal and ethical restrictions. De-sensitized extracted features and checkpoints can be made available upon reasonable request to obsgyn@hku.hk. All codes essential to fully reproduce the training, testing, and heatmap visualizations are freely available on GitHub (https://github.com/AlexXuNB/Human_endometrial_fertility_histology_DL_study) and through an unrestricted Zenodo repository (DOI: 10.5281/zenodo.19730652). The repository includes a clear README instruction on how each analysis pipeline can be executed.

Funding Statement

This study was partly supported by the Health and Medical Research Fund (HMRF 04151546 to Yin Lau Lee; HMRF 10212996 to Yin Lau Lee) from the Food and Health Bureau, Government of the Hong Kong Special Administrative Region. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Saket Z, Källén K, Lundin K, Magnusson Å, Bergh C. Cumulative live birth rate after IVF: trend over time and the impact of blastocyst culture and vitrification. Hum Reprod Open. 2021;2021(3):hoab021. doi: 10.1093/hropen/hoab021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Wang YA, Farquhar C, Sullivan EA. Donor age is a major determinant of success of oocyte donation/recipient programme. Hum Reprod. 2012;27(1):118–25. doi: 10.1093/humrep/der359 [DOI] [PubMed] [Google Scholar]
  • 3.Pirtea P, De Ziegler D, Tao X, Sun L, Zhan Y, Ayoubi JM, et al. Rate of true recurrent implantation failure is low: results of three successive frozen euploid single embryo transfers. Fertil Steril. 2021;115(1):45–53. doi: 10.1016/j.fertnstert.2020.07.002 [DOI] [PubMed] [Google Scholar]
  • 4.Ni Y, Shen H, Yao H, Zhang E, Tong C, Qian W, et al. Differences in fertility-related quality of life and emotional status among women undergoing different IVF treatment cycles. Psychol Res Behav Manag. 2023;16:1873–82. doi: 10.2147/PRBM.S411740 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Pasch LA, Gregorich SE, Katz PK, Millstein SG, Nachtigall RD, Bleil ME, et al. Psychological distress and in vitro fertilization outcome. Fertil Steril. 2012;98(2):459–64. doi: 10.1016/j.fertnstert.2012.05.023 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Lutjen P, Trounson A, Leeton J, Findlay J, Wood C, Renou P. The establishment and maintenance of pregnancy using in vitro fertilization and embryo donation in a patient with primary ovarian failure. Nature. 1984;307(5947):174–5. doi: 10.1038/307174a0 [DOI] [PubMed] [Google Scholar]
  • 7.Vaegter KK, Lakic TG, Olovsson M, Berglund L, Brodin T, Holte J. Which factors are most predictive for live birth after in vitro fertilization and intracytoplasmic sperm injection (IVF/ICSI) treatments? Analysis of 100 prospectively recorded variables in 8,400 IVF/ICSI single-embryo transfers. Fertility and sterility. 2017;107(3):641–8. [DOI] [PubMed] [Google Scholar]
  • 8.Gallos ID, Khairy M, Chu J, Rajkhowa M, Tobias A, Campbell A, et al. Optimal endometrial thickness to maximize live births and minimize pregnancy losses: analysis of 25,767 fresh embryo transfers. Reprod Biomed Online. 2018;37(5):542–8. doi: 10.1016/j.rbmo.2018.08.025 [DOI] [PubMed] [Google Scholar]
  • 9.Mahutte N, Hartman M, Meng L, Lanes A, Luo Z-C, Liu KE. Optimal endometrial thickness in fresh and frozen-thaw in vitro fertilization cycles: an analysis of live birth rates from 96,000 autologous embryo transfers. Fertility and Sterility. 2022;117(4):792–800. [DOI] [PubMed] [Google Scholar]
  • 10.Ng EHY, Chan CCW, Tang OS, Yeung WSB, Ho PC. The role of endometrial blood flow measured by three-dimensional power Doppler ultrasound in the prediction of pregnancy during in vitro fertilization treatment. Eur J Obstet Gynecol Reprod Biol. 2007;135(1):8–16. doi: 10.1016/j.ejogrb.2007.06.006 [DOI] [PubMed] [Google Scholar]
  • 11.Noyes RW, Hertig AT, Rock J. Dating the endometrial biopsy. Obstet Gynecol Surv. 1950;5(4):561–4. doi: 10.1097/00006254-195008000-00044 [DOI] [PubMed] [Google Scholar]
  • 12.Murray MJ, Meyer WR, Zaino RJ, Lessey BA, Novotny DB, Ireland K, et al. A critical analysis of the accuracy, reproducibility, and clinical utility of histologic endometrial dating in fertile women. Fertil Steril. 2004;81(5):1333–43. doi: 10.1016/j.fertnstert.2003.11.030 [DOI] [PubMed] [Google Scholar]
  • 13.Coutifaris C, Myers ER, Guzick DS, Diamond MP, Carson SA, Legro RS, et al. Histological dating of timed endometrial biopsy tissue is not related to fertility status. Fertil Steril. 2004;82(5):1264–72. doi: 10.1016/j.fertnstert.2004.03.069 [DOI] [PubMed] [Google Scholar]
  • 14.Arian SE, Hessami K, Khatibi A, To AK, Shamshirsaz AA, Gibbons W. Endometrial receptivity array before frozen embryo transfer cycles: a systematic review and meta-analysis. Fertil Steril. 2023;119(2):229–38. doi: 10.1016/j.fertnstert.2022.11.012 [DOI] [PubMed] [Google Scholar]
  • 15.Díaz-Gimeno P, Ruiz-Alonso M, Blesa D, Bosch N, Martínez-Conejero JA, Alamá P, et al. The accuracy and reproducibility of the endometrial receptivity array is superior to histology as a diagnostic method for endometrial receptivity. Fertil Steril. 2013;99(2):508–17. doi: 10.1016/j.fertnstert.2012.09.046 [DOI] [PubMed] [Google Scholar]
  • 16.Cohen AM, Ye XY, Colgan TJ, Greenblatt EM, Chan C. Comparing endometrial receptivity array to histologic dating of the endometrium in women with a history of implantation failure. Syst Biol Reprod Med. 2020;66(6):347–54. doi: 10.1080/19396368.2020.1824032 [DOI] [PubMed] [Google Scholar]
  • 17.Pantanowitz L, Sharma A, Carter AB, Kurc T, Sussman A, Saltz J. Twenty years of digital pathology: an overview of the road travelled, what is on the horizon, and the emergence of vendor-neutral archives. J Pathol Inform. 2018;9:40. doi: 10.4103/jpi.jpi_69_18 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Srinidhi CL, Ciga O, Martel AL. Deep neural network models for computational histopathology: A survey. Med Image Anal. 2021;67:101813. doi: 10.1016/j.media.2020.101813 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Tellez D, Litjens G, Bándi P, Bulten W, Bokhorst J-M, Ciompi F, et al. Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology. Med Image Anal. 2019;58:101544. doi: 10.1016/j.media.2019.101544 [DOI] [PubMed] [Google Scholar]
  • 20.Lu MY, Chen B, Williamson DFK, Chen RJ, Liang I, Ding T, et al. A visual-language foundation model for computational pathology. Nat Med. 2024;30(3):863–74. doi: 10.1038/s41591-024-02856-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Chen RJ, Ding T, Lu MY, Williamson DFK, Jaume G, Song AH, et al. Towards a general-purpose foundation model for computational pathology. Nat Med. 2024;30(3):850–62. doi: 10.1038/s41591-024-02857-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Vorontsov E, Bozkurt A, Casson A, Shaikovski G, Zelechowski M, Liu S. Virchow: A million-slide digital pathology foundation model. 2023. https://arxiv.org/abs/230907778
  • 23.Lu MY, Williamson DFK, Chen TY, Chen RJ, Barbieri M, Mahmood F. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat Biomed Eng. 2021;5(6):555–70. doi: 10.1038/s41551-020-00682-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Dimitriadis I, Zaninovic N, Badiola AC, Bormann CL. Artificial intelligence in the embryology laboratory: a review. Reprod Biomed Online. 2022;44(3):435–48. doi: 10.1016/j.rbmo.2021.11.003 [DOI] [PubMed] [Google Scholar]
  • 25.Ehteshami Bejnordi B, Veta M, Johannes van Diest P, van Ginneken B, Karssemeijer N, Litjens G, et al. Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer. JAMA. 2017;318(22):2199–210. doi: 10.1001/jama.2017.14585 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 770–8. 10.1109/cvpr.2016.90 [DOI]
  • 27.Li T, Liao R, Chan C, Greenblatt EM. Deep learning analysis of endometrial histology as a promising tool to predict the chance of pregnancy after frozen embryo transfers. J Assist Reprod Genet. 2023;40(4):901–10. doi: 10.1007/s10815-023-02745-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Adams SM, Terry V, Hosie MJ, Gayer N, Murphy CR. Endometrial response to IVF hormonal manipulation: comparative analysis of menopausal, down regulated and natural cycles. Reprod Biol Endocrinol. 2004;2:21. doi: 10.1186/1477-7827-2-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Pan Y, Li B, Wang Z, Wang Y, Gong X, Zhou W, et al. Hormone replacement versus natural cycle protocols of endometrial preparation for frozen embryo transfer. Front Endocrinol (Lausanne). 2020;11:546532. doi: 10.3389/fendo.2020.546532 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Lee YL, Ruan H, Lee KC, Fong SW, Yue C, Chen ACH, et al. Attachment of a trophoblastic spheroid onto endometrial epithelial cells predicts cumulative live birth in women aged 35 and older. Fertil Steril. 2023;120(2):268–76. doi: 10.1016/j.fertnstert.2023.03.013 [DOI] [PubMed] [Google Scholar]
  • 31.Cao D, Liu Y, Cheng Y, Wang J, Zhang B, Zhai Y, et al. Time-series single-cell transcriptomic profiling of luteal-phase endometrium uncovers dynamic characteristics and its dysregulation in recurrent implantation failures. Nat Commun. 2025;16(1):137. doi: 10.1038/s41467-024-55419-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Gardner DK, Schoolcraft WB. Culture and transfer of human blastocysts. Curr Opin Obstet Gynecol. 1999;11(3):307–11. doi: 10.1097/00001703-199906000-00013 [DOI] [PubMed] [Google Scholar]
  • 33.Zegers-Hochschild F, Adamson GD, Dyer S, Racowsky C, de Mouzon J, Sokol R. The International glossary on infertility and fertility care, 2017. Human Reproduction. 2017;32(9):1786–801. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.An BGL, Chapman M, Tilia L, Venetis C. Is there an optimal window of time for transferring single frozen-thawed euploid blastocysts? A cohort study of 1170 embryo transfers. Hum Reprod. 2022;37(12):2797–807. doi: 10.1093/humrep/deac227 [DOI] [PubMed] [Google Scholar]
  • 35.Muller D, Soto-Rey I, Kramer F. An analysis on ensemble learning optimized medical image classification with deep convolutional neural networks. IEEE Access. 2022;10:66467–80. doi: 10.1109/access.2022.3182399 [DOI] [Google Scholar]
  • 36.Zhang Y, Li H, Sun Y, Zheng S, Zhu C, Yang L. Attention-challenging multiple instance learning for whole slide image classification. In: European conference on computer vision, 2024.
  • 37.Zhang A, Jaume G, Vaidya A, Ding T, Mahmood F. Accelerating data processing and benchmarking of ai models for pathology. 2025.
  • 38.Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In: 2017 IEEE International Conference on Computer Vision (ICCV). 2017. 618–26. 10.1109/iccv.2017.74 [DOI]
  • 39.Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, et al. Pytorch: An imperative style, high-performance deep learning library. Adv Neural Inform Process Syst. 2019;32. [Google Scholar]
  • 40.Sun X, Xu W. Fast implementation of delong’s algorithm for comparing the areas under correlated receiver operating characteristic curves. IEEE Signal Process Lett. 2014;21(11):1389–93. doi: 10.1109/lsp.2014.2337313 [DOI] [Google Scholar]
  • 41.Neha NF. Radiomics in medical imaging: methods, applications, and challenges. 2026. https://doi.org/arXiv:260200102 [DOI] [PMC free article] [PubMed]
  • 42.Howard FM, Dolezal J, Kochanny S, Schulte J, Chen H, Heij L, et al. The impact of site-specific digital histology signatures on deep learning model accuracy and bias. Nat Commun. 2021;12(1):4423. doi: 10.1038/s41467-021-24698-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Marečková M, Garcia-Alonso L, Moullet M, Lorenzi V, Petryszak R, Sancho-Serra C, et al. An integrated single-cell reference atlas of the human endometrium. Nat Genet. 2024;56(9):1925–37. doi: 10.1038/s41588-024-01873-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Wahab M, Thompson J, Hamid B, Deen S, Al-Azzawi F. Endometrial histomorphometry of trimegestone-based sequential hormone replacement therapy: a weighted comparison with the endometrium of the natural cycle. Hum Reprod. 1999;14(10):2609–18. doi: 10.1093/humrep/14.10.2609 [DOI] [PubMed] [Google Scholar]
  • 45.Takahashi S, Sakaguchi Y, Kouno N, Takasawa K, Ishizu K, Akagi Y, et al. Comparison of vision transformers and convolutional neural networks in medical image analysis: a systematic review. J Med Syst. 2024;48(1):84. doi: 10.1007/s10916-024-02105-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Chen J, Pan T, Zhu Z, Liu L, Zhao N, Feng X, et al. A deep learning-based multimodal medical imaging model for breast cancer screening. Sci Rep. 2025;15(1):14696. doi: 10.1038/s41598-025-99535-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Bareja R, Carrillo-Perez F, Zheng Y, Pizurica M, Nandi TN, Shen J. Evaluating vision and pathology foundation models for computational pathology: A comprehensive benchmark study. medRxiv. 2025. doi: 10.1101/2025.05.08.25327250 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Li J, Hu J, Sun Q, Yan R, Ouyang M, Guan T, et al. Can we simplify slide-level fine-tuning of pathology foundation models?. 2025. https://arxiv.org/abs/250220823
  • 49.Mishra V, Lotter W. Comparing computational pathology foundation models using representational similarity analysis. 2025. https://doi.org/arXiv:250915482 [PMC free article] [PubMed]
  • 50.Ding T, Wagner SJ, Song AH, Chen RJ, Lu MY, Zhang A, et al. A multimodal whole-slide foundation model for pathology. Nat Med. 2025;31(11):3749–61. doi: 10.1038/s41591-025-03982-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Ilse M, Tomczak J, Welling M. Attention-based deep multiple instance learning. In: International conference on machine learning. 2018.
  • 52.Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW, Topic Group ‘Evaluating diagnostic tests and prediction models’ of the STRATOS initiative. Calibration: The Achilles heel of predictive analytics. BMC Med. 2019;17(1):230. doi: 10.1186/s12916-019-1466-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Binuya MAE, Engelhardt EG, Schats W, Schmidt MK, Steyerberg EW. Methodological guidance for the evaluation and updating of clinical prediction models: A systematic review. BMC Med Res Methodol. 2022;22(1):316. doi: 10.1186/s12874-022-01801-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Thornburgh I, Anderson MC. The endometrial deficient secretory phase. Histopathology. 1997;30(1):11–5. doi: 10.1046/j.1365-2559.1997.d01-554.x [DOI] [PubMed] [Google Scholar]
  • 55.Stewart CJR, Bharat C, Leake R. Asynchronous glands in secretory pattern endometrium: Clinical associations and immunohistological changes. Histopathology. 2015;67(1):39–47. doi: 10.1111/his.12620 [DOI] [PubMed] [Google Scholar]
  • 56.Russell P, Hey-Cunningham A, Berbic M, Tremellen K, Sacks G, Gee A, et al. Asynchronous glands in the endometrium of women with recurrent reproductive failure. Pathology. 2014;46(4):325–32. doi: 10.1097/PAT.0000000000000111 [DOI] [PubMed] [Google Scholar]
  • 57.Neha F, Bhati D, Shukla DK. Phenotyping of histology imaging data with histomics. AI. 2026;7(6):228. doi: 10.3390/ai7060228 [DOI] [Google Scholar]
  • 58.Lee S, Arffman RK, Komsi EK, Lindgren O, Kemppainen J, Kask K, et al. Dynamic changes in AI-based analysis of endometrial cellular composition: Analysis of PCOS and RIF endometrium. J Pathol Inform. 2024;15:100364. doi: 10.1016/j.jpi.2024.100364 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Raudonis V, Bartasiene R, Minajeva A, Saare M, Drejeriene E, Kozlovskaja-Gumbriene A, et al. Towards metric-driven difference detection between receptive and nonreceptive endometrial samples using automatic histology image analysis. Appl Sci. 2024;14(13):5715. doi: 10.3390/app14135715 [DOI] [Google Scholar]
  • 60.Zeng Q, Klein C, Caruso S, Maille P, Allende DS, Mínguez B, et al. Artificial intelligence-based pathology as a biomarker of sensitivity to atezolizumab-bevacizumab in patients with hepatocellular carcinoma: a multicentre retrospective study. Lancet Oncol. 2023;24(12):1411–22. doi: 10.1016/S1470-2045(23)00468-0 [DOI] [PubMed] [Google Scholar]
  • 61.Murphy CR. Uterine receptivity and the plasma membrane transformation. Cell Res. 2004;14(4):259–67. doi: 10.1038/sj.cr.7290227 [DOI] [PubMed] [Google Scholar]
  • 62.Ye X. Uterine luminal epithelium as the transient gateway for embryo implantation. Trends Endocrinol Metab. 2020;31(2):165–80. doi: 10.1016/j.tem.2019.11.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Jones-Paris CR, Paria S, Berg T, Saus J, Bhave G, Paria BC, et al. Embryo implantation triggers dynamic spatiotemporal expression of the basement membrane toolkit during uterine reprogramming. Matrix Biol. 2017;57–58:347–65. doi: 10.1016/j.matbio.2016.09.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Tandulwadkar S, Pal B. Hysteroscopy simplified by masters. Springer; 2020. [Google Scholar]
  • 65.Mathyk B, Schwartz A, DeCherney A, Ata B. A critical appraisal of studies on endometrial thickness and embryo transfer outcome. Reprod Biomed Online. 2023;47(4):103259. doi: 10.1016/j.rbmo.2023.103259 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Deng R, Cui C, Remedios LW, Bao S, Womick RM, Chiron S. Cross-scale multi-instance learning for pathological image diagnosis. Medical Image Analysis. 2024;94:103124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Tempest N, Soul J, Hill CJ, Caamaño Gutierrez E, Hapangama DK. Cell type and region-specific transcriptional changes in the endometrium of women with RIF identify potential treatment targets. Proc Natl Acad Sci U S A. 2025;122(11):e2421254122. doi: 10.1073/pnas.2421254122 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Garcia-Alonso L, Handfield L-F, Roberts K, Nikolakopoulou K, Fernando RC, Gardner L, et al. Mapping the temporal and spatial dynamics of the human endometrium in vivo and in vitro. Nat Genet. 2021;53(12):1698–711. doi: 10.1038/s41588-021-00972-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Boehm KM, Khosravi P, Vanguri R, Gao J, Shah SP. Harnessing multimodal data integration to advance precision oncology. Nat Rev Cancer. 2022;22(2):114–26. doi: 10.1038/s41568-021-00408-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.McLernon DJ, Steyerberg EW, Te Velde ER, Lee AJ, Bhattacharya S. Predicting the chances of a live birth after one or more complete cycles of in vitro fertilisation: population based study of linked cycle data from 113 873 women. BMJ. 2016;355. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Alsbjerg B, Jensen MB, Elbaek HO, Laursen R, Povlsen BB, Anderson R, et al. Midluteal serum estradiol levels are associated with live birth rates in hormone replacement therapy frozen embryo transfer cycles: a cohort study. Fertil Steril. 2024;121(6):1000–9. doi: 10.1016/j.fertnstert.2024.04.006 [DOI] [PubMed] [Google Scholar]
  • 72.Zheng Z, Zhang X, Wu F, Liao H, Zhao H, Zhang M, et al. Effect of BMI on cumulative live birth rates in patients that completed IVF treatment: a retrospective cohort study of 16,126 patients. Endocr Connect. 2024;13(3):e230105. doi: 10.1530/EC-23-0105 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Beshar I, Milki AA, Gardner RM, Zhang WY, Johal JK, Bavan B. Elevated body mass index in modified natural cycle frozen euploid embryo transfers is not associated with live birth rate. J Assisted Reprod Genetics. 2023;40(5):1055–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Takaya Y, Matsubayashi H, Kitaya K, Nishiyama R, Yamaguchi K, Takeuchi T, et al. Minimum values for midluteal plasma progesterone and estradiol concentrations in patients who achieved pregnancy with timed intercourse or intrauterine insemination without a human menopausal gonadotropin. BMC Res Notes. 2018;11(1):61. doi: 10.1186/s13104-018-3188-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Shaban MT, Baur C, Navab N, Albarqouni S. Staingan: stain style transfer for digital histological images. In: 2019 IEEE 16th International Symposium on Biomedical Imaging (ISBI 2019), 2019. 953–6. 10.1109/isbi.2019.8759152 [DOI]
  • 76.Adebayo J, Gilmer J, Muelly M, Goodfellow I, Hardt M, Kim B. Sanity checks for saliency maps. Adv Neural Inform Process Syst. 2018;31. [Google Scholar]
  • 77.Draelos RL, Carin L. Use HiResCAM instead of Grad-CAM for faithful explanations of convolutional neural networks. arXiv preprint. 2020. 10.48550/arXiv.2011.08891 [DOI]
  • 78.Lensen S, Shreeve N, Barnhart KT, Gibreel A, Ng EHY, Moffett A. In vitro fertilization add-ons for the endometrium: it doesn’t add-up. Fertility and Sterility. 2019;112(6):987–93. [DOI] [PubMed] [Google Scholar]
PLOS Digit Health. doi: 10.1371/journal.pdig.0001744.r001

Decision Letter 0

Feng Liu, Fnu Neha

14 Jun 2026

Response to Reviewers Revised Manuscript with Track Changes Manuscript Journal Requirements:

1. Please amend your detailed Financial Disclosure statement. This is published with the article. It must therefore be completed in full sentences and contain the exact wording you wish to be published.

a. State the initials, alongside each funding source, of each author to receive each grant.

b. State what role the funders took in the study. If the funders had no role in your study, please state: “The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.”

c. If any authors received a salary from any of your funders, please state which authors and which funders.

2.  We note that you have indicated that there are restrictions to data sharing for this study. For studies involving human research participant data or other sensitive data, we encourage authors to share de-identified or anonymized data. However, when data cannot be publicly shared for ethical reasons, we allow authors to make their data sets available upon request. For information on unacceptable data access restrictions, please see http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions.

Before we proceed with your manuscript, please address the following prompts:

a) If there are ethical or legal restrictions on sharing a de-identified data set, please explain them in detail (e.g., data contain potentially identifying or sensitive patient information, data are owned by a third-party organization, etc.) and who has imposed them (e.g., a Research Ethics Committee or Institutional Review Board, etc.). Please also provide contact information for a data access committee, ethics committee, or other institutional body to which data requests may be sent.

b) If there are no restrictions, please upload the minimal anonymized data set necessary to replicate your study findings to a stable, public repository and provide us with the relevant URLs, DOIs, or accession numbers. Please see http://www.bmj.com/content/340/bmj.c181.long for guidelines on how to de-identify and prepare clinical data for publication. For a list of recommended repositories, please see https://journals.plos.org/plosone/s/recommended-repositories. You also have the option of uploading the data as Supporting Information files, but we would recommend depositing data directly to a data repository if possible.

Please update your Data Availability statement in the submission form accordingly.

3. Some material included in your submission may be copyrighted. According to PLOS’s copyright policy, authors who use figures or other material (e.g., graphics, clipart, maps) from another author or copyright holder must demonstrate or obtain permission to publish this material under the Creative Commons Attribution 4.0 International (CC BY 4.0) License used by PLOS journals. Please closely review the details of PLOS’s copyright requirements here: PLOS Licenses and Copyright. If you need to request permissions from a copyright holder, you may use PLOS's Copyright Content Permission form.

Please respond directly to this email or email the journal office and provide any known details concerning your material's license terms and permissions required for reuse, even if you have not yet obtained copyright permissions or are unsure of your material's copyright compatibility.

Potential Copyright Issues:

Fig 1,2,3: Please confirm whether you drew the images / clip-art within the figure panels by hand. If you did not draw the images, please provide (a) a link to the source of the images or icons and their license / terms of use; or (b) written permission from the copyright holder to publish the images or icons under our CC-BY 4.0 license. Alternatively, you may replace the images with open source alternatives. See these open source resources you may use to replace images / clip-art:

- https://commons.wikimedia.org

- https://openclipart.org/

Additional Editor Comments (if provided)

Comments from the Editorial Office: To ensure transparency, we are informing you that the previous Academic Editor acted as Reviewer 3 on your manuscript.

Reviewers' Comments:

Comments to the Author

1. Does this manuscript meet PLOS Digital Health’s publication criteria?>

Reviewer #1: Yes

Reviewer #2: Partly

Reviewer #3: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?-->?>

Reviewer #1: N/A

Reviewer #2: Yes

Reviewer #3: N/A

**********

3. Have the authors made all data underlying the findings in their manuscript fully available (please refer to the Data Availability Statement at the start of the manuscript PDF file)??>

The PLOS Data policy

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English??>

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

Reviewer #1: The paper presents an interesting application of deep learning and histopathology foundation models for predicting cumulative live birth outcomes from endometrial histology WSIs in IVF settings. The comparison between conventional ResNet-18 training and the UNI2-h foundation model is technically relevant, and the inclusion of luminal epithelium-focused analysis adds biological interpretability to the framework. The use of external validation, Grad-CAM analysis, multimodal integration, and discussion on domain transferability between cancer-oriented foundation models and non-malignant reproductive histology strengthens the manuscript’s contribution. The study also demonstrates good methodological transparency by reporting cross-validation strategies, preprocessing pipelines, and detailed evaluation metrics.

However, several aspects require further clarification and improvement before publication. The sample size remains relatively limited for developing robust histopathology AI models, especially for external validation and LE-focused experiments, which may affect statistical reliability and generalizability. Although the authors discuss overfitting and domain shift, the manuscript would benefit from stronger analysis of stain normalization, scanner variability, and batch-effect mitigation strategies across institutions. In addition, the biological interpretation of the learned representations remains somewhat limited, as Grad-CAM only provides coarse localization rather than mechanistic insight into fertility-associated morphological phenotypes. The multimodal analysis is also relatively simple, relying only on logistic regression with age and endometrial thickness, whereas additional reproductive and hormonal variables could further improve prediction.

The literature review should also be expanded to better position the work within computational pathology and histomics-based representation learning. The authors are encouraged to discuss and cite "Phenotyping of Histology Imaging Data with Histomics", which provides important insights into phenotype-driven histology representation learning, quantitative tissue characterization, and interpretable histopathological feature extraction. This reference would strengthen the manuscript’s discussion regarding biologically meaningful tissue phenotyping, representation-level learning, and morphology-driven computational pathology pipelines, particularly in the context of fertility-oriented endometrial analysis.

Reviewer #2: This manuscript presents an interesting and clinically relevant application of deep learning and foundation models for predicting IVF cumulative live birth outcomes using endometrial histology whole-slide images. The comparison between conventional CNN-based approaches and the UNI2-h foundation model is valuable, and the biologically informed luminal epithelium (LE)-focused analysis adds originality to the study. The inclusion of external validation and explainability analyses are important strengths.

The manuscript is generally well written and technically sound overall. However, several methodological and interpretability concerns should be addressed before publication.

Major comments:

The external cohort labeling strategy requires further clarification and discussion. The external dataset did not use actual IVF/live birth outcomes but instead inferred labels based on LH timing. While biologically reasonable, this is not equivalent to true clinical live birth outcomes and may affect external validity and model performance interpretation.

The relatively limited patient-level sample size raises concerns regarding overfitting and generalizability, particularly given the large number of extracted patches and multiple model configurations evaluated. Additional discussion regarding statistical uncertainty and robustness across institutions would strengthen the manuscript.

The manuscript reports AUROC and balanced accuracy values across models, but no formal statistical comparison between model performances is presented. Statistical testing comparing AUROCs or confidence interval overlap analysis would strengthen claims regarding superiority or robustness of specific architectures.

Although the Grad-CAM analysis is informative, the semi-quantitative categorization appears partially subjective. The manuscript would benefit from additional clarification regarding annotation criteria, reviewer blinding, and reproducibility of the interpretability assessment.

The clinical applicability of this approach should be discussed in greater depth. It would be valuable to better contextualize how this workflow could realistically integrate into IVF practice and how it compares with currently available receptivity assessment tools.

Minor comments:

Terminology regarding “live birth prediction,” “fertility prediction,” and “receptivity prediction” should be standardized throughout the manuscript for clarity.

Additional discussion comparing UNI2-h with other recent pathology foundation models would help contextualize model selection.

Minor language editing and sentence shortening in several sections would improve readability.

Overall, this is a promising and innovative study with meaningful translational potential. The concerns above appear addressable and would substantially strengthen the manuscript.

Reviewer #3: This manuscript presents an interesting and clinically relevant application of computational pathology by leveraging foundation models and deep learning for predicting cumulative live birth outcomes from endometrial histology. The comparison between conventional CNN-based approaches and the histopathology foundation model (UNI2-h), together with the biologically motivated luminal epithelium (LE)-focused analysis, provides novelty and translational potential. The inclusion of external validation and explainability analyses using Grad-CAM further strengthens the study. However, several methodological, statistical, biological, and presentation-related limitations need to be addressed before publication.

The manuscript would benefit from deeper methodological rigor, stronger biological interpretation, clearer discussion of limitations, and broader integration with existing quantitative pathology literature. In particular, the study positions itself within computational pathology but overlooks important literature on histomics and quantitative image phenotyping, which are highly relevant to the interpretation and biological validation of histology-based AI models.

Suggested references that should be discussed: "Phenotyping of Histology Imaging Data with Histomics",

"Radiomics in Medical Imaging: Methods, Applications, and Challenges"

These works provide important perspectives on quantitative tissue phenotyping, interpretable feature extraction, reproducibility, feature stability, and biological representation learning, all of which are directly relevant to the proposed endometrial histology analysis framework.

1. Limited Dataset Size Relative to Model Complexity

A major concern is the relatively small patient cohort used to train and evaluate extremely large deep learning models.

The training cohort consists of only 152 patients, while UNI2-h contains hundreds of millions of parameters. Although the feature extractor is frozen, the study still attempts to learn fertility-related representations from a limited number of cases.

The manuscript repeatedly attributes superior performance to foundation-model knowledge transfer, but it remains unclear whether observed improvements are statistically significant or merely reflect sampling variability.

The authors should discuss:

Dataset-size limitations more explicitly.

Risks of model instability and optimistic performance estimates.

Statistical power limitations.

Whether performance differences between models are statistically significant.

Confidence intervals alone are insufficient to demonstrate superiority.

2. External Validation Design is Problematic

The external cohort does not contain actual IVF outcome labels.

Instead, fertility status is inferred from menstrual phase timing:

LH+7 → labeled as successful

LH+3, LH+5, LH+9, LH+11 → labeled as unsuccessful

This introduces a substantial conceptual issue.

The internal model is trained to predict cumulative live birth outcomes, whereas external validation evaluates the ability to distinguish endometrial phases.

These are fundamentally different prediction targets.

Consequently, the reported external AUROC values may not truly represent generalization for live birth prediction.

The manuscript should explicitly acknowledge this limitation and avoid presenting external performance as a direct validation of IVF outcome prediction.

3. Insufficient Biological Interpretation

Although Grad-CAM visualizations are provided, biological interpretation remains superficial.

The study concludes that LE-associated regions contribute to prediction, but does not explain:

Which morphological characteristics drive predictions.

Whether attention corresponds to known implantation biology.

Whether glandular, stromal, vascular, inflammatory, or epithelial features are implicated.

The discussion remains largely model-centric rather than biology-centric.

More detailed pathological interpretation is necessary to support clinical relevance.

4. Missing Quantitative Histomics Perspective

The manuscript treats the problem entirely as a deep-learning classification task.

However, computational pathology increasingly emphasizes interpretable quantitative tissue phenotyping through histomics-based approaches.

The authors should discuss how their deep learning framework relates to established histomic descriptors such as:

Nuclear morphology

Cellular density

Glandular architecture

Spatial organization

Tissue texture patterns

The manuscript would be significantly strengthened by discussing findings in the context of:

Phenotyping of Histology Imaging Data with Histomics

This reference provides an important framework for linking computational image features with biologically meaningful tissue phenotypes.

5. Lack of Comparison with Handcrafted Feature Approaches

The study compares multiple deep learning architectures but does not include any traditional quantitative pathology baseline.

A comparison against:

Histomics features,

Morphological descriptors,

Classical machine learning models,

would provide valuable insight into whether foundation models truly outperform interpretable quantitative approaches.

Without such comparison, it is difficult to determine the added value of the proposed deep learning framework.

6. Interpretability Claims Are Overstated

The manuscript repeatedly emphasizes biological interpretability.

However, Grad-CAM alone does not provide mechanistic interpretability.

Grad-CAM merely highlights image regions associated with model decisions and does not explain:

Which features were learned,

Why predictions were made,

Whether highlighted regions are causally relevant.

The authors should moderate interpretability claims and acknowledge limitations of attention-based explanations.

7. No Assessment of Model Calibration

Clinical deployment requires more than discrimination metrics.

The study reports:

AUROC

Accuracy

Balanced accuracy

but does not evaluate calibration.

Important metrics such as:

Calibration curves

Brier score

Expected calibration error (ECE)

should be considered.

A model with good AUROC may still generate poorly calibrated probabilities unsuitable for clinical decision-making.

8. Limited Discussion of Reproducibility and Domain Shift

The study demonstrates substantial performance degradation of the HRT-trained model when applied to natural-cycle data.

This observation highlights a broader issue of domain shift.

However, the discussion does not sufficiently address:

Scanner variability

Staining variability

Institution-specific biases

Preprocessing variability

Batch effects

These challenges are central topics in quantitative imaging research.

The authors may benefit from discussing reproducibility principles described in:

Radiomics in Medical Imaging: Methods, Applications, and Challenges

which provides a comprehensive discussion of feature robustness, reproducibility, and external validation.

9. Statistical Analysis Needs Strengthening

The manuscript reports mean ± SD values across folds.

However, no statistical comparisons are presented between:

ResNet-18 vs UNI2-h

GAP vs ACMIL

Linear vs MLP

WSI vs LE models

Without statistical testing, claims regarding superiority remain largely descriptive.

Appropriate significance testing or confidence interval comparisons should be considered.

10. Clinical Utility Remains Unclear

Even the best-performing models achieve AUROCs around 0.84–0.90.

The manuscript does not explain:

How predictions would alter IVF decision-making.

Whether performance exceeds existing clinical assessment approaches.

Whether cost-benefit advantages exist.

How false-positive and false-negative predictions would affect patients.

The clinical translation pathway remains insufficiently developed.

**********

what does this mean?). If published, this will include your full peer review and any attached files.

Do you want your identity to be public for this peer review?  If you choose “no”, your identity will remain anonymous but your review may still be made public.

For information about this choice, including consent withdrawal, please see our Privacy Policy

Reviewer #1: No

Reviewer #2: No

Reviewer #3: No

**********

Figure resubmission:

Reproducibility: --> -->-->To enhance the reproducibility of your results, we recommend that authors of applicable studies deposit laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols-->?>

PLOS Digit Health. doi: 10.1371/journal.pdig.0001744.r003

Decision Letter 1

Feng Liu

27 Aug 2026

Response to ReviewersRevised Manuscript with Track ChangesManuscriptJournal Requirements: Additional Editor Comments (if provided): Reviewers' Comments:

Comments to the Author

Reviewer #1: All comments have been addressed

Reviewer #2: (No Response)

**********

publication criteria?>

Reviewer #1: Yes

Reviewer #2: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?-->?>

Reviewer #1: N/A

Reviewer #2: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available (please refer to the Data Availability Statement at the start of the manuscript PDF file)??>

The PLOS Data policy

Reviewer #1: Yes

Reviewer #2: No

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English??>

Reviewer #1: Yes

Reviewer #2: Yes

**********

Review Comments to the Author

Reviewer #1: Good work.

Reviewer #2: Thank you for the substantial revisions. The manuscript has improved considerably, and the authors have addressed the major methodological and statistical concerns raised in the previous round. In particular, the addition of formal DeLong testing with multiple-comparison adjustment, calibration assessment, clearer discussion of limited sample size and model uncertainty, improved treatment of technical domain shift, and more cautious interpretation of model comparisons have strengthened the study.

One point would benefit from further clarification before publication.

The external cohort is now appropriately described as an “external phase-based testing” cohort rather than a true external validation cohort for cumulative live birth. However, the Abstract still presents the internal and external AUROC values in parallel, which may give readers the impression that the external cohort directly validates live-birth prediction. Because the external cohort distinguishes LH+7 from non-LH+7 samples rather than using actual IVF/live-birth outcomes, I suggest making this distinction explicit in the Abstract at the point where the external performance results are reported. The conclusion should also continue to avoid implying that these external results constitute direct clinical validation of live-birth prediction.

Aside from this point, the authors have responded comprehensively to the previous reviewer concerns, and the revised manuscript is substantially stronger.

**********

what does this mean?). If published, this will include your full peer review and any attached files.

Do you want your identity to be public for this peer review?  If you choose “no”, your identity will remain anonymous but your review may still be made public.

For information about this choice, including consent withdrawal, please see our Privacy Policy

Reviewer #1: No

Reviewer #2: No

**********

Figure resubmission:

Reproducibility: --> -->-->To enhance the reproducibility of your results, we recommend that authors of applicable studies deposit laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols-->?>

PLOS Digit Health. doi: 10.1371/journal.pdig.0001744.r005

Decision Letter 2

Feng Liu

9 Sep 2026

Foundation model-powered deep learning of endometrial histology for predicting the cumulative live birth of an in vitro fertilization cycle

PDIG-D-26-00729R2

Dear Dr Lee,

We are pleased to inform you that your manuscript 'Foundation model-powered deep learning of endometrial histology for predicting the cumulative live birth of an in vitro fertilization cycle' has been provisionally accepted for publication in PLOS Digital Health.

Before your manuscript can be formally accepted you will need to complete some formatting changes, which you will receive in a follow-up email from a member of our team.

Please note that your manuscript will not be scheduled for publication until you have made the required changes, so a swift response is appreciated.

IMPORTANT: The editorial review process is now complete. PLOS will only permit corrections to spelling, formatting or significant scientific errors from this point onwards. Requests for major changes, or any which affect the scientific understanding of your work, will cause delays to the publication date of your manuscript.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they'll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact digitalhealth@plos.org.

Thank you again for supporting Open Access publishing; we are looking forward to publishing your work in PLOS Digital Health.

Best regards,

Feng Liu, Ph.D.

Section Editor

PLOS Digital Health

***********************************************************

Additional Editor Comments (if provided):

Reviewer Comments (if any, and for reference):

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Fig. HRT model performance on natural cycle samples.

    An ROC curve showing the performance of the reported 3-fold ensemble HRT-model on the natural cycle LH + 7 WSIs.

    (TIF)

    pdig.0001744.s001.TIF (352.5KB, TIF)
    S2 Fig. Summary of model predictive performance trajectories during training sessions.

    Summary figures illustrate progression of model predictive performance in AUROC (y-axis) over training epochs (x-axis). Performance is displayed for three cohorts including validation, internal testing, and external testing cohorts. All discussed models‘ performance trajectories were presented, including (A) ResNet-18 WSI model trajectory represented in blue solid line; (B) UNI2-h GAP-Linear and UNI2-h GAP-MLP model trajectories presented in blue and red solid lines; (C) UNI2-h ACMIL-Linear and UNI2-h ACMIL-MLP model trajectories presented in blue and red solid lines; (D) ResNet-18 LE and UNI2-h LE model trajectories presented in blue and red solid lines, with their WSI counter variants (ResNet-18 WSI and UNI2-h GAP-MLP model trajectories) plotted in purple and yellow dotted lines for easier comparison.

    (TIF)

    S3 Fig. Visualization of attention distributions in the ACMIL-based UNI2-h model.

    Representative heatmaps display patch-level attention tendency on WSIs for a correctly predicted CLB case (A) and Unsuccessful (B) case. Warmer colors indicate higher attention weights.

    (TIF)

    pdig.0001744.s003.TIF (3.6MB, TIF)
    S4 Fig. Histology-based model pairwise performance comparison summary.

    Summary of model performance comparisons between specified models, specifically in the internal outcome-based (A) and the external phase-based (B) testing cohorts. Δ value presents difference of AUROC between model A and B within each specified pair. Formal statistics performed with DeLong test corrected for multiple comparison (p < 0.05 indicated with *).

    (TIF)

    pdig.0001744.s004.TIF (737.9KB, TIF)
    S5 Fig. Summary of per model calibration metrics.

    Calibration curves presented for each single reported models, including (A) ResNet-18 WSI, (B) UNI2-h GAP-Linear WSI, (C) UNI2-h GAP-MLP WSI, (D) UNI2-h ACMIL-Linear WSI, (E) UNI2-h ACMIL-MLP WSI, (F) ResNet-18 LE, and (G) UNI2-h LE model. Solid line presents model’s calibration curve on the internal test set; dotted line presents model’s calibration curve on the external set.

    (TIF)

    S6 Fig. Additional multimodal model evaluations and feature-impact analysis.

    ROC curves for the Age + BMI + EnTh clinical metadata-only model (A) and the hormone-only E2 + P4 model (B). SHAP summary plot showing feature-contribution patterns across the internal test (C) and out-of-fold development samples (D) for the integrative seven-factor model. Positive SHAP values indicate a shift toward live birth prediction, negative SHAP values indicate a shift toward unsuccessful outcome prediction. Each dot represents one sample, with colour representing each samples’ original feature value, with red indicating a higher and blue representing a lower original value.

    (TIF)

    pdig.0001744.s006.TIF (867.8KB, TIF)
    S1 Table. Summary of clinical characteristics of the internal cohort by outcome group.

    (XLSX)

    pdig.0001744.s007.xlsx (8.3KB, xlsx)
    S2 Table. Model training hyperparameter setting summary.

    (XLSX)

    pdig.0001744.s008.xlsx (9.4KB, xlsx)
    S3 Table. Evaluation of the HRT-trained ResNet-18 model on the internal natural cycle cohort.

    (XLSX)

    pdig.0001744.s009.xlsx (8.3KB, xlsx)
    S4 Table. Performance metrics of logistic regression models with multimodal integration.

    (XLSX)

    pdig.0001744.s010.xlsx (8.6KB, xlsx)
    Attachment

    Submitted filename: Response_to_Reviewers_Endometrial_Histology_final.docx

    pdig.0001744.s011.docx (49.9KB, docx)
    Attachment

    Submitted filename: Response_to_Reviewers_PDIG-D-26-00729R2_minor_revision.docx

    pdig.0001744.s012.docx (31.8KB, docx)

    Data Availability Statement

    Sensitive raw clinical data and images are prohibited to be shared publicly due to legal and ethical restrictions. De-sensitized extracted features and checkpoints can be made available upon reasonable request to obsgyn@hku.hk. All codes essential to fully reproduce the training, testing, and heatmap visualizations are freely available on GitHub (https://github.com/AlexXuNB/Human_endometrial_fertility_histology_DL_study) and through an unrestricted Zenodo repository (DOI: 10.5281/zenodo.19730652). The repository includes a clear README instruction on how each analysis pipeline can be executed.


    Articles from PLOS Digital Health are provided here courtesy of PLOS

    RESOURCES