Abstract
Objective and reproducible assessment of endometrial receptivity is essential for optimizing in vitro fertilization (IVF) success, yet traditional histological dating suffers from observer variability. This study investigated whether deep learning of hematoxylin and eosin histology images could support cumulative live birth prediction in IVF. An end-to-end ResNet-18 was compared with a UNI2-h-based pipeline, using UNI2-h as a frozen feature extractor. Ten-fold cross-validation ensembles were developed from natural-cycle endometrial biopsies and evaluated in an internal held-out cohort with live birth outcomes. Additional phase-based evaluation was performed, which tested the performance in distinguishing LH + 7 versus non-LH + 7 phases. A luminal epithelium (LE)-focused design was also assessed to examine whether concentrating on maternal-embryo interface could improve fertility-oriented learning. The whole-slide ResNet-18 model performed well in internal outcome-based testing but near random in external phase-based testing. In contrast, the UNI2-h model with mean pooling and a multilayer perceptron classifier showed less divergent performance, with ensembled AUROCs of 0.74 ± 0.03 in internal outcome-based evaluation, and 0.90 ± 0.05 in external phase-based testing. Despite a smaller training set, LE-focused models retained comparable performance. In internal outcome-based testing, LE-focused ResNet-18 and UNI2-h models achieved AUROCs of 0.71 ± 0.03 and 0.74 ± 0.09 respectively. External phase-based testing yielded AUROCs of 0.80 ± 0.16 for ResNet-18, and 0.84 ± 0.06 for UNI2-h. Grad-CAM review of LE-focused ResNet-18 models showed attention commonly on LE-alone or mixed with adjacent stroma. In multimodal analyses, integrated models incorporating histology significantly outperformed the clinical metadata-only model. Integrated model’s feature weighting showed dominant histology outputs, smaller contributions from estradiol and maternal age, and negligible contributions from progesterone, endometrial thickness, and BMI. These findings support histology-based AI for fertility-oriented endometrial assessment and highlight biologically informed design and repurposed foundation models as promising bases for clinically meaningful prediction.
Author summary
Clinicians still lack an objective way to assess whether the endometrium is in an optimal state to support a successful pregnancy. Traditional histological assessment relies on pathologists’ interpretations and vary among observers, while newer molecular tests remain costly and uncertain reliability. In this study, we explored whether artificial intelligence (AI) could extract useful fertility-related information from routine endometrial histology slides. We compared a conventional training approach with a more modern strategy that repurposed an existing AI model with pre-existing histology knowledge. Two testing cohorts, including an internal cohort based on live-birth outcomes and an external cohort evaluated on luteal-phase classification were included. The repurposed approach showed more consistent performance between the two testing cohorts, while the conventional approach fell to near-random in the external phase-based cohort. We also tested whether focusing on endometrial luminal epithelium, where the embryo first makes contact, could improve prediction. Surprisingly, luminal epithelium-focused models performed comparably to the whole-slide approach despite using far fewer training images. These findings support AI-based histology analysis as a promising approach that warrants further refinement for more objective endometrial assessments. Combining foundation models with biologically informed design could provide a useful direction for future fertility-related diagnostics.
Introduction
The global birth rate remains low as many countries report their birth rate lower than the replacement level of 2.1 births per woman. The current cumulative live birth rate per cycle remained at ~30–40% only [1] despite advancements in in vitro fertilization (IVF). Poor embryo quality is often cited as the primary factor in infertility, but the probability of a live birth per a single euploid embryo transfer stays at approximately 60% [2,3], suggesting that maternal factors beyond embryo euploidy contribute to the outcomes. Consequently, a large proportion of patients require multiple transfer cycles, which can be particularly challenging for patients with a limited number of good quality embryos. Repeated failures are also significantly correlated with a higher risk of anxiety and depression even after accounting for socioeconomic confounding factors [4,5].
Optimal endometrial preparation is essential for the establishment of a successful pregnancy [6]. Yet, clinicians still lack an accurate and low cost method to assess endometrial receptivity and predict live birth outcomes [7]. Conventional non-invasive markers, such as ultrasound-based endometrial thickness and vascularization fail to serve as reliable predictive tools because of conflicting results reported [8–10]. The Noyes criteria, which provided descriptive morphological standards to date endometrial samples, has long been the “gold standard” of endometrial histological dating since its introduction in 1950 [11]. However, the technique is labor intensive and exhibits significant observer variability, compromising its reliability [12,13]. Advancements in understanding the molecular signatures of the window of implantation (WOI) have provided hopes of novel assays for endometrial dating. Up-to-date, the true clinical value of currently available molecular assays, like the Endometrial Receptivity Analysis (ERA), remains controversial due to unproven benefits in predicting IVF success [14–16]. Thus, there is a critical need for an objective and low-cost method to evaluate endometrial receptivity.
Since the introduction of the initial whole slide image (WSI) scanner in the 1990s, computational pathology has evolved from simple feature extraction to advanced deep learning algorithms capable of interpreting complex high-dimensional features [17]. Deep learning models are commonly developed either by full training from scratch or by adapting pretrained foundation models. The full training approach, typically utilizing Convolutional Neural Networks (CNNs), enables task-specific learning but is inherently data-demanding and sensitive to institutional biases such as imaging and staining variations [18,19]. Consequently, the field is increasingly shifting toward foundation models as a more robust alternative [20–22]. Unlike conventional approaches requiring large task-specific datasets, foundation models such as UNI [21] and Virchow [22] are pretrained on massive histology archives, making the utilization of complex Vision Transformer (ViT) backbones practically feasible. The increasing availability of these foundation models represents a critical methodological advance to make the development of task-specific pathological deep learning models easier, especially when sample sizes are restricted [23].
The application of deep learning based artificial intelligence (AI) model in healthcare is a fast-growing field with immense potential [24]. While AI has been widely tested for its application in optimizing stimulation protocols and embryo selection, its use in predicting endometrial receptivity remains limited. By training a ResNet-18 model from scratch, a widely adopted efficient architecture for natural and medical image analysis [25,26], to differentiate endometrial samples based on fertility outcomes, we reported the earliest application of deep learning in endometrial receptivity [27]. While the reported models achieved satisfactory live birth prediction with an accuracy of 75%, the wider applicability of the model which was derived exclusively from patients undergoing Hormone Replacement Therapy (HRT) was constrained. Endometrium in HRT cycles exhibits distinct histological profiles compared to those in natural cycles [28,29]. Therefore, the generalizability of such a model to natural cycles remains unverified. Furthermore, a significant knowledge gap remains regarding the transferability of modern foundation models to the endometrial function-oriented tasks. Many widely used histopathology foundation models have been developed and extensively benchmarked in cancer-heavy settings [20–22]. It is currently unknown whether these cancer-specific extractors can effectively characterize the subtle, non-malignant morphological changes associated with endometrial receptivity in healthy tissues.
In this study, we aimed to first evaluate the potential transferability of the HRT-trained ResNet-18 model for assessing the performance using WSIs obtained in natural cycles. We also evaluated whether the histopathology pretrained UNI2-h foundation model could provide a simpler, more robust starting point for digital pathology-based fertility outcome prediction. As luminal epithelium (LE) of endometrium is the first point of contact between maternal tissue and the implanting embryo, we also explored an LE-focused modelling approach specific to the fertility-oriented predictive goal.
Materials and methods
Ethical approval
The study protocol was approved by the Institutional Review Boards of the University of Hong Kong/Hospital Authority Hong Kong West Cluster (IRB: UW 25–573) and the Ethics Committee of the University of Hong Kong-Shenzhen Hospital (IRB reference number [2026]087).
Study cohorts and sample collection
Archived anonymized paraffin-embedded tissues with known pregnancy outcomes from a previous study [30] were used in this study. IVF patients who had regular cycles and underwent their first IVF cycles at the Centre of Assisted Reproduction and Embryology, The University of Hong Kong - Queen Mary Hospital and the Assisted Reproduction Centre of Kwong Wah Hospital from 2017 to 2021 were included in this study. They were asked to donate endometrial biopsy samples in their natural cycles. Individuals with [1] an abnormal uterine cavity as determined by saline infusion sonogram or hysteroscopy; [2] untreated uterine pathologies such as endometrial polyps or fibroids; [3] untreated hydrosalpinx; and [4] use of donor oocytes were excluded. Pipelle samplers (CCD Laboratories, Paris, France) were used to obtain the endometrial biopsy at 7 days after a luteinizing hormone (LH) surge (LH + 7). Biopsied samples were fixed and embedded in paraffin blocks. From each block, 4 μm tissue sections were prepared for hematoxylin and eosin (H&E) staining, and the stained slides were digitized as high resolution WSIs at 40 × magnification (approximately 0.220 µm per pixel, or mpp) using a Hamamatsu NanoZoomer S210 slide scanner (Hamamatsu Photonics K.K., Shizuoka Prefecture, Japan). The digitalized WSI files were stored natively in NDPI files first before transformed for further formatting change and downstream processing. In total, 183 internal samples were included in this study, of which 152 samples were assigned for training and validation purposes, and another 31 samples were assigned as an internal held-out test set. Essential metadata and clinical characteristics were summarized in S1 Table.
We also included an external phase-based cohort consisting of 21 fertile patients WSIs collected at natural cycle from the University of Hong Kong-Shenzhen Hospital. The patient sample condition and tissue collection details were described previously [31]. All donors had achieved a healthy live birth within two years of collection either via spontaneous pregnancy or a preimplantation genetic testing (PGT)-validated euploid embryo transfer. The collected samples were biopsied at five different endometrial menstrual phases, including LH + 3 (n = 3), LH + 5 (n = 3), LH + 7 (n = 9), LH + 9 (n = 3), and LH + 11 (n = 3). Paraffin block embedding and H&E staining were conducted at the University of Hong Kong-Shenzhen Hospital. The slides were scanned at 20 × scanning resolution (approximately equivalent to 0.440 mpp) using a KF-FL-020 WSI scanner (KFBIO, Zhejiang, China). The resulting digitalized WSI files were stored natively in KFB file format before transformed for further processing.
Ovarian stimulation and embryo transfer
Clinical management and IVF procedures were conducted according to the centers’ standard operating procedures as described earlier [30]. Briefly, ovarian stimulation using a gonadotropin-releasing hormone antagonist protocol was used and the starting dosage of FSH depended on the antral follicle count and body mass index of the patients. Final oocyte maturation was triggered with recombinant human chorionic gonadotropin (hCG), followed by oocyte retrieval 34–36 hours later. Embryos were cultured to reach cleavage or blastocyst stage and eventually selected for the subsequent fresh or frozen embryo transfer (FET) cycles. The blastocysts were graded based on the Gardner scoring system [32].
Outcome data collection
The primary outcome was the cumulative live birth, defined as a live birth occurring after 22 completed weeks of gestation [33] from a stimulated IVF cycle and its subsequent FET cycles within 6 months after start of ovarian stimulation. The outcomes were assigned as either live birth or unsuccessful outcomes. This binary classification was utilized as the ground truths for training and assessing the performance of the developed models in the internal dataset. Ultimately, among the 152 samples utilized for training purposes, 79 were labeled as unsuccessful and 73 were labeled as live birth. The 31 internal held-out testing cohort consisted of 20 unsuccessful and 11 live birth samples.
As the external phase-based cohort comprised fertile donors who did not undergo embryo transfer, direct cumulative live birth outcome labels were unavailable. Therefore, the labels of external phase-based testing cohort were inferred by the timing of the biopsy relative to the LH surge. LH + 7 was considered as receptive, as it is widely accepted to be the optimal timepoint for embryo transfer [31,34]. Our previous time-series single-cell data further supported marked transcriptomic transitions before and after LH + 7, with stromal and glandular epithelial cells showing major changes from LH + 5 to LH + 7, and epithelial cells exhibiting a further sharp transition at late LH + 9 [31]. Thus, a non-identical but similar phase-based binary label was assigned, with LH + 7 samples (n = 9) representing the live birth outcome samples, and pre-receptive (LH + 3, + 5) and post-receptive (LH + 9, + 11) samples collectively representing unsuccessful outcome samples (n = 12). This external testing design therefore evaluated a biologically relevant but non-equivalent target, focusing on distinguishing receptive-phase LH + 7 endometrium from non-LH + 7 samples, rather than directly validating cumulative live birth prediction.
Histology-based model selection and training parameter settings
This study utilized two separate deep learning architectures, the ImageNet 1K-pretrained ResNet-18 [26] and the histopathology pretrained UNI2-h [21], to achieve the same goal of developing a binary classification model to differentiate endometrial H&E WSIs based on their live birth outcome potentials. All the models generated were trained on a remote Linux server implemented with three available NVIDIA A100 Tensor Core GPUs, except for the HRT-trained ResNet-18 model (HRT-model), which was trained with details as described previously [27].
For all models, an optimal threshold was determined independently for each cross-validation fold by maximizing Youden’s J statistic on the validation cohort (i.e., predicted output higher than the calculated threshold interpreted as live birth; values lower than the threshold interpreted as unsuccessful). Adapting the cross-validation framework from previous studies, we implemented a 10-fold cross-validation strategy to construct an ensemble model, an approach demonstrated to enhance robustness and generalizability compared to individual models [27,35]. In brief, an ensemble was formed by selecting the best-performing model from each fold. The predicted outcomes were averaged with equal weighting to produce a final predicted probability (soft voting), and optimal thresholds were similarly averaged to establish a global threshold.
The adopted CNN-based ImageNet 1K-pretrained ResNet-18 was maintained largely consistent with previous studies by utilizing a linear classifier head for binary classification [27]. Since this model is designed to operate as a patch-level classifier, slide-level predictions were generated by calculating the mean of the predicted probability scores across all patches extracted from a single WSI (Fig 1). To better encounter overfitting issues with the limited sample size, the architecture was modified to include a dropout layer after global average pooling (GAP). Furthermore, conventional data augmentation techniques were applied during training, including random horizontal and vertical flips, rotation, and color jitter. The models were trained using the AdamW optimizer with a cosine decay learning-rate scheduler and a warmup period to prevent destabilizing model weights. Detailed hyperparameter settings, including learning rates, batch sizes, and weight decay values, are summarized in S2 Table.
Fig 1. Overview of deep learning pipelines for endometrial fertility outcome prediction.

The workflow utilizes input patches derived either from whole slide images (WSIs) or manually annotated LE/peri-LE regions. The samples included two outcomes, either a cumulative live birth (CLB) or unsuccessful outcome. Two distinct modeling strategies were implemented. One includes an ImageNet 1K-pretrained ResNet-18 model trained end-to-end as a patch-level classifier, with patch-level probability scores processed independently and averaged in the end to generate a final slide-level binary prediction (CLB vs. Unsuccessful). The second strategy utilizes a histopathology pre-trained UNI2-h foundation model as a feature extractor. Extracted patch features were aggregated via either a global averaged pooling (GAP) or an attention-based pooling strategy (ACMIL). The aggregated slide-level representation was classified using either a linear classifier head or a non-linear MLP head to predict the slide-level binary outcome. Figure partially created in BioRender [https://biorender.com/f125abi].
The ViT-H/14-based histopathology pretrained UNI2-h foundation model was utilized as a feature extractor to extract high dimensional features from image patches [21]. Considering the size of the UNI2-h model at 681M parameters and the relatively limited abundance of our available sample set, the model’s weights were kept completely frozen during the study. To map these patch-level features to slide-level predictions, two distinctive feature aggregation strategies were implemented (Fig 1). One configuration utilized a GAP mechanism, in which the patch-level features extracted from a single slide were averaged to form a slide-level bag. In parallel, an Attention-Challenging Multiple Instance Learning (ACMIL) mechanism [36] was adapted. ACMIL replaced the equal averaging with a learnable attention network which differentially weight patches based on their discriminative relevance. Averaged or weighted feature outputs from the aggregation mechanisms were eventually passed on to either a conventional linear or a non-linear Multi-Layer Perceptron (MLP) classifier head for the final binary prediction. The MLP variant comprised two hidden layers with 768 and 384 dimensions, and provided a non-linear projection to adapt the UNI2-h features to the current specific endometrial task. For all the UNI2-h extracted feature-based training, one NVIDIA A100 GPU was used. The hyperparameter settings were kept largely consistent across the WSI models, except for necessary adaptations to fit specific model’s input expectations. Detailed hyperparameter settings are shown in S2 Table.
For the LE-focused models specifically trained with patches containing LE and peri-LE stroma regions, the architectures and major parameter settings were kept consistent with the WSI approaches (S2 Table). Specifically, the ResNet-18 LE models were trained using the ImageNet 1K-pretrained backbone combined with a linear classifier head, while the UNI2-h LE models utilized the frozen feature extractor combined with an MLP head.
Patch image extraction and preprocessing
Direct processing of WSIs in full resolution during training remains impractical due to their typical gigapixel-scale dimensions. Therefore, all WSIs were cropped into smaller patch images in order to facilitate efficient training. Although the maximum native scanning resolution differed between the two cohorts, with internal samples scanned at 40× and external samples scanned at 20 × , all subsequent training and analyses were standardized to patches generated at 20 × resolution (0.440 mpp) for minimizing technical variations. An exception was the internal dataset used for testing the pre-developed HRT model, which utilized an unchanged 40× (0.220 mpp) resolution at 4832 × 6864 pixels to match its specific input settings as described previously [27].
For all WSI models, automated segmentation was performed using the Trident toolkit [37], and only patches with at least 35% of tissue were included. The size of the resulting patch datasets varied slightly depending on the expected input pixel dimension of specific model. For the pipeline involving the ResNet-18 WSI models, all involved patches were extracted at 20 × resolution with a dimension of 224 × 224 pixels. This dataset comprised a training set of 152 patients (92,738 patches), an internal testing set of 31 patients (22,361 patches), and an external phase-based testing set of 21 patients (33,011 patches). In parallel, for pipeline involving the UNI2-h WSI models, all involved patches were extracted at 20 × resolution with dimensions of 256 × 256 pixels. This dataset comprised a training set of 152 patients (85,827 patches), an internal testing set of 31 patients (17,713 patches), and an external phase-based testing set of 21 patients (31,873 patches).
In contrast, the LE-focused workflow relied on manual annotation. Approximate LE-containing region annotations were drawn using QuPath, a WSI viewing and processing tool, followed by automatic generation of fixed-size squares to cover the annotated Regions of Interest (ROIs). These squares were manually adjusted to maximize LE coverage while minimizing the inclusion of glandular structures and were subsequently extracted using a custom Groovy script. Unlike the WSI approach which covers the whole tissue, conducting separate rounds of manual annotation for each pipeline could introduce bias, leading to inconsistencies in areas such as ROI centering and stromal proportion between the two model inputs. Hence, a single set of ‘base’ patches was generated first to ensure consistent anatomical coverage across the two pipelines. Specifically, a series of base patches was first extracted at 260 × 260 pixels (20 × resolution) from the internal and external cohort, and subsequently center-cropped to match the input dimensions of each architecture (224 × 224 pixels for ResNet-18-based and 256 × 256 pixels for UNI2-h feature-based). Consequently, the patch counts were identical for both pipelines: a training set of 124 patients (3,552 patches), an internal testing set of 24 patients (657 patches), and an external phase-based testing set of 18 patients (783 patches).
Visualization and quantification of model attention regions
The gradient-weighted class activation mapping (Grad-CAM) approach [38] was used to visualize the attention regions on the original patch images of LE models. All patch images from the internal testing cohort (657 images) were utilized to generate the attention heatmaps. In these heatmaps, warmer (red) areas highlighted regions the model deemed more important for its prediction, while cooler (blue) areas had weaker contributions. Semi-quantitative analysis of the Grad-CAM-generated heatmaps was performed by manually evaluating the most prominent high-attention area within the LE-containing patches using the internal testing cohort. Attention categories were assigned according to the predominant localization of the highlighted region. Considering that the patches primarily consisted of LE, peri-LE stroma, and lumen, which were generally easily distinguishable on H&E images, Grad-CAM attention was prespecified into four broad categories comprising LE-predominant, mixed LE and peri-LE stroma, peri-LE stroma-predominant, and patterns on unclassified regions such as lumen and potential artifacts. Other cell types potentially associated with fertility outcomes, such as endothelial and immune cells, were not classified separately because they could not be identified consistently across patches and Grad-CAM provided only approximate spatial localization. Formal categorization was performed independently using anonymized patch identifiers and images, without access to model predictions, prediction probabilities, prediction correctness, patient outcomes, or other clinical information. Based on the average AUROCs across the internal and external sets, three best-performing models from the 10-fold ensembles were selected to visualize for their attention tendencies. Semi-quantitative summaries of attention category distributions were calculated separately for each model and reported as mean proportions with standard deviations. Native ACMIL attention weights from the ACMIL-based WSI pipelines were also directly visualized on WSIs for qualitative inspection of any potentially unique attention patterns.
Multimodal analysis for clinical metadata integration
To evaluate the potential of multimodal analysis, logistic regression models were constructed using predicted probabilities from the WSI and LE histology models, with or without maternal age (Age), body mass index (BMI), endometrial thickness (EnTh), and biopsy-day serum estradiol (E2) and progesterone (P4) levels within the internal test set. EnTh was determined via transvaginal ultrasound prior to the biopsy for sample collection. Five logistic regression settings were evaluated including: a histology-only model integrating WSI and LE models’ probabilities; an integrative seven-factor histology-clinical model additionally incorporating Age, BMI, EnTh, E2 and P4; and three clinical metadata-only models incorporating Age, BMI and EnTh; E2 and P4; or all five clinical variables. To maximize training efficiency while preventing data leakage, we leveraged the 10-fold cross-validation framework of the histology models to generate out-of-fold predictions for the entire training cohort. These out-of-fold probabilities served as the input features for training the logistic regression models. Multimodal analyses containing histology-based probabilities as inputs were restricted to samples with both WSI and LE model outputs. To examine feature contribution patterns in the most integrative seven-factor model, standardized coefficients weights were calculated by multiplying each raw coefficient by the feature’s standard deviation, allowing the presented coefficients weights to be directly comparable. SHapley Additive exPlanations (SHAP) summary plots were additionally generated for the internal cohorts to summarize the direction and distribution of each feature’s contribution to the predictions.
Computational software packages
The ImageNet 1K-pretrained ResNet-18 (https://download.pytorch.org/models/resnet18-5c106cde.pth) was trained and the histopathology pretrained UNI2-h foundation model (https://huggingface.co/MahmoodLab/UNI2-h) was utilized directly in this study. The UNI2-h model weights were accessed via application through an institutional email-linked HuggingFace account as requested. Training of models within the study and Grad-CAM visualization largely relied on functions and tools available from the PyTorch machine learning library on Python (PyTorch version = 2.4.1; Python version = 3.10.0) [39]. Logistic regression models for multimodal analysis were implemented using scikit-learn (version = 1.7.1), and SHAP-based model interpretation was performed using the SHAP package (version = 0.49.1). QuPath (version = 0.5.1) was the primary WSI inspection software. Trident (version 0.2.0) is a toolkit developed by the same group that created the UNI2-h [37]. The toolkit was used in all our reported ResNet-18 and UNI2-h-related patch preparation tasks, with Tiff patch images generated for the ResNet-18 workflow and patch-level feature vectors extracted for the UNI2-h-related workflow. Patch image generation for the LE-focused model was performed using QuPath and a Groovy script. DeLong tests were implemented following the fast DeLong algorithm described by Sun and Xu [40].
Model performance evaluation and reported metrics
The performances of the developed 10-fold ensemble models were evaluated based on three common machine learning performance metrics, including accuracy, balanced accuracy, and area under the receiver operating characteristic curve (AUROC). To monitor training stability and convergence among different architectures and classifier heads, performance trajectories were tracked across all 10 folds and reported as the mean of available folds per epoch to account for varying early-stopping points. To address the potential misleading interpretation of accuracy considering the study’s test sample sets were not in 1:1 balanced ratio, balanced accuracy was utilized to define best epochs during training and validation sessions.
Reported models were compared in two major settings using DeLong test, including with a chance-level random classifier with AUROC = 0.5, and paired model-to-model comparisons. Comparison between the WSI and the LE model was restricted to samples with predictions from both pipelines to ensure paired testing. Reported P values were adjusted using the Benjamini-Hochberg procedure within each cohort and comparison group, with statistical significance defined as an adjusted P < 0.05. To supplement AUROC-based discrimination and threshold-dependent classification metrics, model calibration was assessed using calibration curves and Brier scores. Each model was analyzed separately in the internal cumulative-live-birth testing cohort and the external phase-based testing cohort.
Results
Limited transferability of the HRT model to natural cycle WSIs
To evaluate the potential transferability of the existing deep learning model, we first assessed the performance of the reported HRT model [27] using the natural cycle LH + 7 WSIs archived in our centers. The input settings of all the patches, including patch resolution, dimensions, and cropping procedures, were prepared to match its published HRT model. Despite the currently applied input-level standardization, the model yielded an AUROC of 0.55 ± 0.08 and a balanced accuracy of 55.85 ± 1.37%, indicating a limited discriminative capability in this natural cycle cohort (S1 Fig; S3 Table).
Natural cycle endometrial WSI-trained new models for live birth outcome prediction
We next explored new models using a new WSI workflow for the natural cycle dataset (Fig 2A). We first retrained a new model based on the same ResNet-18 architecture by incorporating adjusted normalization parameters and augmentation strategies tailored to the new dataset. To investigate whether feature representations learned from large pan-cancer datasets could be transferable to a non-cancerous endometrium, live birth-oriented task, we also implemented the UNI2-h foundation model as a frozen feature extractor, evaluating its performance across different aggregation strategies (GAP vs. ACMIL) and classifier heads (linear vs. MLP) with detailed performance metrics summarized in Table 1.
Fig 2. WSI endometrial live birth outcome prediction models.

(A) Schematic diagram showing the training workflow of the WSI models developed in this study. (B-D) ROC curves illustrating the performances of internal and external phase-based test sets of ResNet-18 model (B), UNI2-h GAP-Linear and GAP-MLP models (C) and UNI2-h ACMIL-Linear and UNI2-h ACMIL-MLP models (D). Figure partially created in BioRender [https://biorender.com/f125abi].
Table 1. Performance metrics of WSI deep learning models for endometrial live birth outcome prediction.
| Performance Metric ± SD | ResNet-18 WSI | UNI2-h GAP-Linear WSI |
UNI2-h GAP-MLP WSI |
UNI2-h ACMIL-Linear WSI |
UNI2-h ACMIL-MLP WSI |
|---|---|---|---|---|---|
| Internal Testing Set | |||||
| Sensitivity | 0.727 ± 0.105 | 0.636 ± 0.315 | 0.727 ± 0.141 | 0.636 ± 0.200 | 0.636 ± 0.148 |
| Specificity | 0.800 ± 0.136 | 0.600 ± 0.235 | 0.700 ± 0.166 | 0.650 ± 0.197 | 0.750 ± 0.137 |
| PPV | 0.667 ± 0.165 | 0.467 ± 0.145 | 0.571 ± 0.073 | 0.500 ± 0.169 | 0.583 ± 0.160 |
| NPV | 0.842 ± 0.041 | 0.750 ± 0.106 | 0.824 ± 0.075 | 0.765 ± 0.086 | 0.789 ± 0.052 |
| F1 | 0.696 ± 0.055 | 0.538 ± 0.184 | 0.640 ± 0.050 | 0.560 ± 0.073 | 0.608 ± 0.075 |
| Accuracy | 0.774 ± 0.063 | 0.613 ± 0.056 | 0.710 ± 0.071 | 0.645 ± 0.071 | 0.710 ± 0.068 |
| Balanced Accuracy | 0.764 ± 0.042 | 0.618 ± 0.057 | 0.714 ± 0.046 | 0.643 ± 0.045 | 0.693 ± 0.060 |
| AUC | 0.836 ± 0.025 | 0.627 ± 0.066 | 0.741 ± 0.030 | 0.727 ± 0.031 | 0.750 ± 0.066 |
| External Test Set | |||||
| Sensitivity | 1.000 ± 0.365 | 0.778 ± 0.292 | 0.556 ± 0.216 | 0.444 ± 0.333 | 0.556 ± 0.235 |
| Specificity | 0.083 ± 0.309 | 0.583 ± 0.377 | 1.000 ± 0.121 | 0.917 ± 0.382 | 1.000 ± 0.085 |
| PPV | 0.450 ± 0.177 | 0.583 ± 0.197 | 1.000 ± 0.168 | 0.800 ± 0.320 | 1.000 ± 0.316 |
| NPV | 1.000 ± 0.409 | 0.778 ± 0.352 | 0.750 ± 0.106 | 0.688 ± 0.252 | 0.750 ± 0.086 |
| F1 | 0.621 ± 0.209 | 0.667 ± 0.167 | 0.714 ± 0.149 | 0.571 ± 0.198 | 0.714 ± 0.259 |
| Accuracy | 0.476 ± 0.090 | 0.667 ± 0.142 | 0.810 ± 0.094 | 0.714 ± 0.109 | 0.810 ± 0.105 |
| Balanced Accuracy | 0.542 ± 0.093 | 0.681 ± 0.119 | 0.778 ± 0.103 | 0.681 ± 0.083 | 0.778 ± 0.119 |
| AUC | 0.444 ± 0.140 | 0.815 ± 0.164 | 0.898 ± 0.051 | 0.778 ± 0.075 | 0.870 ± 0.095 |
PPV = positive predictive value; NPV = negative predictive value; AUC = area under the receiver operating characteristic curve
1. ResNet-18 model with end-to-end training.
The ResNet-18 ensemble model trained on the internal natural cycle WSI dataset (ResNet-18 WSI model) demonstrated strong predictive capability when evaluated on the internal testing set (AUROC = 0.84 ± 0.03; balanced accuracy = 76.36 ± 4.23%) (Fig 2B). However, its performance on the external phase-based testing cohort collapsed rapidly during training, resulting in a final ensemble model achieving an AUROC of only 0.44 ± 0.14 and a balanced accuracy of 54.17 ± 9.31%, representing a poorer performance towards the external phase-based testing setting (Fig 2B; S2A Fig). DeLong test further indicated that the internal testing AUROC was significantly above that of a chance-level random classifier (P < 0.001), while the external phase-based testing AUROC remained not significantly different from chance level (P = 0.738).
2. Frozen UNI2-h feature extractor with global average pooling.
The ensemble model utilizing the frozen UNI2-h feature extractor with a standard linear classifier head (UNI2-h GAP-Linear WSI model) indicated a limited overall predictive capability (Fig 2C). For the internal outcome-based testing cohort, the model achieved a non-significant AUROC of 0.63 ± 0.07 (P = 0.258) with a balanced accuracy of 61.82 ± 5.74%. A higher performance was observed in the external phase-based testing cohort, where the model has yielded a significant AUROC of 0.82 ± 0.16 (P = 0.002) with a balanced accuracy of 68.06 ± 11.86%.
Subsequently, the impact of coupling the feature extractor with a non-linear MLP classifier head (UNI2-h GAP-MLP WSI model) was evaluated (Fig 2C). While training trajectories across the internal cohorts remained generally consistent between both classifiers, the MLP variant demonstrated more rapid convergence on the external cohort during the training session (S2B Fig). For the internal testing cohort, the model achieved a significant AUROC of 0.74 ± 0.03 (P = 0.034) with a balanced accuracy of 71.36 ± 4.61%. Upon external phase-based testing, the model also yielded a significant AUROC of 0.90 ± 0.05 (P < 0.001) with a balanced accuracy of 77.78 ± 10.33%.
3. Frozen UNI2-h feature extractor with weighted attention aggregation.
The application of the ACMIL aggregation strategy was further evaluated with either a standard linear classifier head or an MLP head (Fig 2D). On the internal testing cohort, the model coupled with a linear head (UNI2-h ACMIL-Linear WSI model) achieved a significant AUROC of 0.73 ± 0.03 (P = 0.034) with a balanced accuracy of 64.32 ± 4.52%. The model exhibited a comparable generalizability on the external phase-based test set, which yielded an AUROC of 0.78 ± 0.08 (P = 0.012) with a balanced accuracy of 68.06 ± 8.31%.
The model with ACMIL aggregation strategy utilizing the non-linear MLP classifier head (UNI2-h ACMIL-MLP WSI model) demonstrated significant and elevated AUROC values compared to random (Fig 2D). For the internal outcome-based testing cohort, this approach achieved an AUROC of 0.75 ± 0.07 (P = 0.031) with a balanced accuracy of 69.32 ± 6.03%. Upon external phase-based testing, the model maintained generally comparable to its GAP-based counter version, with an AUROC of 0.87 ± 0.09 (P < 0.001) and a balanced accuracy of 77.78 ± 11.87%. Model’s training trajectories did not indicate a clear advantage in convergence or stability for the ACMIL implementation compared with GAP (S2C Fig).
To further investigate the attention-weighting patterns of the ACMIL framework, we visualized the attention weights directly on the WSIs to identify whether particular histological regions were consistently prioritized by the model. Qualitative visual inspection revealed relatively sparse attention distribution without a strong and clear concentration towards a specific, recognizable morphological structure (S3 Fig).
Performance of deep learning models trained on LE-containing regions
To address the concern that WSI-based approaches may dilute the existence of LE features, we developed biologically-targeted models trained exclusively on manually annotated LE regions (Fig 3A). Predictive performance was assessed across two distinct models trained with the LE-containing patches (namely ResNet-18 LE and UNI2-h LE models), with detailed performance metrics summarized in Table 2. Compared with the ResNet-18 WSI model, the ResNet-18 LE model has demonstrated a less divergent AUROC pattern across the internal outcome-based and external phase-based testing cohorts (Fig 3B; S2D Fig). For the internal testing set, the ResNet-18 LE model achieved a non-significant AUROC of 0.71 ± 0.03 (P = 0.090), with a balanced accuracy of 55.71 ± 5.75%. In the external phase-based testing cohort, the model achieved a significant AUROC of 0.80 ± 0.16 (P = 0.033) and a balanced accuracy of 70.00% ± 12.30%.
Fig 3. LE-targeted endometrial live birth outcome prediction models.

(A) Schematic diagram showing the establishment of deep learning models trained on LE-containing regions. (B, C) ROC curves comparing model performance on the internal outcome-based (left) and external phase-based (right) testing cohorts for the LE-targeted ResNet-18 model (B) and LE-targeted UNI2-h model with an MLP classifier head (C). (D) Representative Grad-CAM visualizations from correctly predicted CLB and unsuccessful outcome cases. Higher attention areas are illustrated in warmer colours and lower attention areas are illustrated in cooler colours. (E) Semi-quantitative analysis of the average attention tendency of the models. Data are presented as the averaged percentage composition of attention categories from top-3 performing models within the 10-fold ensemble [(LE alone (blue), LE+stroma (purple), stroma alone (red), and unclassified (grey)], presented in Mean ± SD. Figure partially created in BioRender [https://biorender.com/f125abi].
Table 2. Performance metrics of LE-focused deep learning models for endometrial live birth outcome prediction.
| Performance Metric ± SD | ResNet-18 LE | UNI2-h LE |
|---|---|---|
| Internal Testing Set | ||
| Sensitivity | 0.400 ± 0.181 | 0.900 ± 0.175 |
| Specificity | 0.714 ± 0.136 | 0.643 ± 0.158 |
| PPV | 0.500 ± 0.098 | 0.643 ± 0.090 |
| NPV | 0.625 ± 0.081 | 0.900 ± 0.094 |
| F1 | 0.444 ± 0.086 | 0.750 ± 0.097 |
| Accuracy | 0.583 ± 0.052 | 0.750 ± 0.081 |
| Balanced Accuracy | 0.557 ± 0.057 | 0.771 ± 0.080 |
| AUC | 0.707 ± 0.027 | 0.743 ± 0.086 |
| External Test Set | ||
| Sensitivity | 1.000 ± 0.159 | 0.750 ± 0.281 |
| Specificity | 0.400 ± 0.320 | 0.800 ± 0.151 |
| PPV | 0.571 ± 0.133 | 0.750 ± 0.076 |
| NPV | 1.000 ± 0.411 | 0.800 ± 0.147 |
| F1 | 0.727 ± 0.083 | 0.750 ± 0.172 |
| Accuracy | 0.667 ± 0.142 | 0.778 ± 0.079 |
| Balanced Accuracy | 0.700 ± 0.123 | 0.775 ± 0.094 |
| AUC | 0.800 ± 0.164 | 0.837 ± 0.063 |
PPV = positive predictive value; NPV = negative predictive value; AUC = area under the receiver operating characteristic curve
The UNI2-h LE model also yielded performance comparable to its WSI variant (Fig 3C; S2D Fig). On the internal outcome-based testing cohort, the model achieved a significant AUROC of 0.74 ± 0.09 (P = 0.034) with a balanced accuracy of 77.14 ± 7.98%. Upon external phase-based testing, the model’s AUROC maintained at 0.84 ± 0.06 (P = 0.002) with a balanced accuracy of 77.50 ± 9.39%.
To further examine the spatial localization of model attention, Grad-CAM was utilized to visualize the high-attention regions driving the predictions of the three top-performing ResNet-18 LE models within the 10-fold ensemble, defined by their averaged AUROCs (0.79, 0.78, and 0.77). The resulting heatmaps were manually categorized according to the predominant localization of the highlighted compartments: (1) LE alone (attention primarily on LE), (2) mixed (attention spanning both LE and peri-LE stroma without clear preference), and (3) peri-LE stroma alone (attention primarily on peri-LE stroma), and (4) unclassified regions, including luminal spaces and potential artifacts (Fig 3D). Semi-quantitative analysis of the top-3 models showed that regions of higher attention were predominantly associated with LE-related compartments rather than peri-LE stroma-alone or unclassified regions. When stratified by predicted pregnancy outcome, patches showing LE alone and mixed attentions remained the two major categories in both CLB and unsuccessful cases, accounting for 44.33 ± 9.98% and 39.03 ± 5.20% of patches in CLB cases, and 38.83 ± 8.02% and 38.25 ± 7.34% in unsuccessful cases, respectively (Fig 3E). In contrast, patches showing high attention on peri-LE stroma alone were less frequent, accounting for 13.75 ± 4.92% and 16.37 ± 1.31% of patches in CLB and unsuccessful cases, respectively. Unclassified patches with attention restricted to empty luminal spaces lacking clear biological relevance, accounted for only a minor proportion in both groups (2.89 ± 2.68% and 6.55 ± 1.89% for CLB and unsuccessful cases, respectively). These results highlighted that Grad-CAM’s localization-based visualization was more focused on LE region.
Comparisons of histology model performances
Paired DeLong tests were used for pairwise comparisons among the individual histology model configurations (S4 Fig). In the internal outcome-based testing cohort, none of the comparisons reached statistical significance after Benjamini-Hochberg adjustment. The two largest internal AUROC differences were between specific model configurations. ResNet-18 WSI and UNI2-h GAP-MLP WSI differed by ΔAUROC = 0.096, comparing end-to-end training with a foundation model-based pipeline. UNI2-h GAP-Linear WSI and UNI2-h GAP-MLP WSI differed by ΔAUROC = 0.114, comparing linear and MLP classifier heads. The remaining internal comparisons showed smaller differences, with the smallest difference observed between UNI2-h GAP-MLP WSI and UNI2-h ACMIL-MLP WSI (ΔAUROC = 0.009), comparing GAP and ACMIL aggregation. In the external phase-based testing cohort, UNI2-h GAP-MLP WSI achieved a significantly higher AUROC than ResNet-18 WSI, comparing the foundation model-based pipeline with end-to-end training (ΔAUROC = 0.454, P = 0.004). The remaining external comparisons were not statistically significant, with ΔAUROC ranging from 0.028 to 0.093. The smallest external difference was again observed between UNI2-h GAP-MLP WSI and UNI2-h ACMIL-MLP WSI (ΔAUROC = 0.028).
Model calibration for probability prediction
Calibration curves and Brier scores were presented for each histology model separately in the internal outcome-based testing cohort and the external phase-based testing cohort (S5 Fig). The ResNet-18 WSI model showed the largest internal-external discrepancy, with the lowest Brier score in the internal testing cohort (0.161) but the highest Brier score in the external phase-based testing cohort (0.331), consistent with its divergent discriminative performance across the two testing settings. Among the UNI2-h WSI feature-based pipelines, Brier scores were generally within a narrower range. The UNI2-h ACMIL-MLP WSI model showed the lowest internal Brier score among the UNI2-h WSI models (0.205), while the UNI2-h GAP-MLP WSI model showed the lowest external Brier score (0.143). For the LE-focused models, the ResNet-18 LE model yielded Brier scores of 0.224 and 0.239 in the internal and external testing cohorts, respectively, while the UNI2-h LE model yielded Brier scores of 0.213 and 0.193, respectively.
Multimodal integration of histology-based model outputs and clinical variables
To assess the potential benefit of multimodal integration, logistic regression models were constructed using predicted probabilities from the UNI2-h GAP-MLP WSI and UNI2-h LE pipelines on the internal outcome-based cohort, with or without maternal age (Age), BMI, endometrial thickness (EnTh), and E2 and P4 levels (detailed metrics in S4 Table). A model integrating the WSI and LE outputs achieved an AUROC of 0.80 (P = 0.006) with a balanced accuracy of 63.57% (Fig 4A). When Age, BMI, EnTh, E2, and P4 were further incorporated, the integrative seven-factor model achieved an AUROC of 0.77 (P = 0.034) with a balanced accuracy of 70.71% (Fig 4B). The clinical metadata-only model incorporating Age, BMI, EnTh, E2, and P4 yielded a non-significant AUROC of 0.52 (P = 0.903) with a balanced accuracy of 52.73% (Fig 4C). Two simpler Age + BMI + EnTh and E2 + P4 clinical metadata-only models were also evaluated, with non-significant AUROCs of 0.541 and 0.532, respectively (S6A, S6B Fig). In paired comparisons, both the WSI + LE model and the integrative seven-factor model achieved significantly improved discriminatory ability than the five-factor clinical metadata-only model (P = 0.026 for both comparisons). No significant difference was reached between the seven-factor and the WSI + LE models (P = 0.607).
Fig 4. Multimodal integration of histology-based models with clinical and hormonal variables.

ROC curves for the histology-only WSI + LE model (A), the integrative seven-factor LE + WSI + Age + BMI + EnTh + E2 + P4 model (B), and the Age + BMI + EnTh + E2 + P4 five-factor clinical metadata-only model (C), evaluated in the internal outcome-based testing cohort. (D) Standardized coefficients of the integrative seven-factor logistic regression model, showing the relative direction and magnitude of each feature’s contribution to the model output.
To further examine feature weighting patterns within the multimodal framework, standardized logistic regression coefficients and SHAP-based feature impacts were visualized. The standardized coefficients showed that the WSI model probability (0.529) and LE model probability (0.432) had the largest positive contributions to the model output, followed by E2 (0.253) and the maternal age with a negative weight (-0.213). Compared to the other features, P4 (0.052), EnTh (0.037), and BMI (0.028) contributed minimally with relatively lower impact (Fig 4D). Consistent with this pattern, SHAP analysis from either internal testing (S6C Fig) or training cohort (S6D Fig) demonstrated a broader distribution of feature impact for the WSI- and LE-based probabilities, smaller effects from E2 and maternal age, and minimal effects from P4, EnTh, and BMI. Together, these results summarized the relative contribution patterns of histology-based and clinical inputs within the multimodal framework.
Discussion
This study explored the application of deep learning to predict cumulative live birth outcomes using human endometrial H&E WSIs. The reported ImageNet 1K-pretrained HRT model [27] was initially tested with our natural cycle endometrial WSIs. Even though patch preparation was matched to its published settings, the model showed poor predictive power. This limited transferability likely reflected the residual domain shift, biological differences between HRT and natural-cycle endometrium, and sample-size uncertainty, rather than implementation settings alone. The broader medical-imaging literature similarly emphasizes that model reliability is affected by pipeline factors, including sample preparation, patch generation, preprocessing, and validation designs [41,42]. Such domain shift might have been exacerbated by the small development cohort [27], further compromising stability and cross-institutional practicality. HRT samples form distinct sub-clusters from natural-cycle counterparts, with significant transcriptional shifts in the endometrial epithelial populations [43]. Histological studies similarly demonstrates HRT affects the thickness and the sizes of glandular and vascular space [29,44], further widening the discrepancy between the reported model and the test set.
In this study, two model development strategies were evaluated for cumulative live birth prediction using samples from natural cycles. The CNN-based ResNet-18 model was utilized as an end-to-end training pipeline due to its efficient and well-established architecture for medical image recognition tasks [45]. ResNet-18’s simplicity also made it practical for limited-data settings, where overfitting is a major concern [26,46]. In parallel, UNI2-h [21,47] was selected as one of the newest publicly accessible histopathology-pretrained foundation models at the time of our model development. Its ViT-Huge architecture and large-scale histopathology pretraining enable more complex histological inference. Nevertheless, its pretraining focused predominantly on cancer data, but not endometrial analysis or fertility outcomes. Thus, we evaluated whether frozen UNI2-h could be repurposed for non-malignant endometrial live-birth prediction.
The two approaches exhibited different performance tendencies. While the ResNet-18 model trained on internal natural cycle samples achieved relatively high predictive performance within the internal outcome-based testing cohort, its performance collapsed to near-random levels in the external phase-based testing cohort. This discrepancy may arise from two non-mutually exclusive factors. First, despite standardizing patches by pixel dimensions and physical coverage based on scanner-specific mpp metadata, residual upstream differences like tissue processing, sectioning, staining, scanner platforms may still have introduced institution-specific image characteristics [41,42]. Such effects may particularly impact end-to-end training, which can learn center-specific patterns from limited single-center data. Second, the external cohort lacked IVF outcome-based validation, instead evaluating receptive LH + 7 versus non-LH + 7 phase discrimination. The performance gap may reflect cross-institutional variation and limited target transferability, potentially similar to the previously assessed HRT model. Further outcome-matched validation with stronger statistical power is required.
In comparison, the UNI2-h feature-based pipeline completely lacks the ability to fine-tune its feature extractor’s logic. Despite, the pipeline eventually demonstrated models with less divergent performances across the internal outcome-based and external phase-based testing cohorts. With the standard linear-probe design, the UNI2-h GAP-Linear underperformed on the internal outcome-based test cohort, contrasting with the effective linear-probe performance reported in the original UNI benchmark [21]. Notably, healthy superficial endometrium and the subtle fertility-associated morphological features are likely underrepresented in UNI2-h’s knowledgebase of training. Li et al. [48] proposed that linear probes may be insufficient when the target differs substantially from the pretraining patterns. We therefore implemented a non-linear MLP head to better adapt UNI2-h-derived features to the endometrial fertility-oriented task. As expected, MLP-head models achieved higher AUROCs than their linear-head counterparts. Specifically, the UNI2-h GAP-MLP achieved significance in both internal and external cohorts, unlike the linear variant (external-only). This confirmed that frozen UNI2-h representations with MLP head are feasible for non-malignant endometrial live birth prediction without feature-extractor modification. Nevertheless, DeLong comparisons showed no significant difference between classifier heads, suggesting potential MLP utility without establishing statistical superiority given the current testing sample size. External generalizability for live birth prediction would require larger, outcome-matched cohorts. Beyond classifier-head adaptation, pretrained model selection may also influence downstream performance. Recent representational-similarity analysis showed that pathology foundation models (e.g., UNI2-h and Virchow2) extract divergent representations from the same histological images despite similar ViT-Huge architectures [49]. Whole-slide pretrained models, such as TITAN and Prov-GigaPath, further broaden the available strategies by incorporating broader slide-level context [50]. Thus, future comparisons across different feature-extraction strategies are essential to identify the optimal starting point for endometrial live birth-oriented analysis.
The study further implemented the ACMIL attention strategy [36] to evaluate whether prioritizing informative WSI regions prevents the dilution of critical features, a method known to improve performance in cancer-oriented studies [21,23,51]. However, comparable AUROC performance was achieved in GAP-MLP and ACMIL-MLP-based models, indicating ACMIL did not clearly outperform GAP. Calibration analysis provided a probability-level perspective of model performance beyond binary discrimination [52]. The UNI2-h ACMIL-MLP model maintained relatively low Brier scores across both testing cohorts, whereas the UNI2-h GAP-MLP model showed the highest internal Brier score. The differential weighting mechanism of ACMIL likely down-weighted the less-informative patches, making it potentially more reliable for probabilistic estimation than global pooling. Therefore, the ACMIL-MLP design can be preferable in clinical settings requiring patient-level probability estimates. For cumulative live birth prediction, larger-cohort comparisons should evaluate the most suitable aggregation approach. Post hoc recalibration may further be applied to improve the reliability of probability estimates used for individualized risk counselling regarding the likelihood of live birth [53].
ACMIL attention was sparse without clustered or consistent concentration within a specific histological compartment. This may be compatible with live birth prediction, where decision-relevant morphology may be more spatially distributed than in cancer-detection tasks [23]. Therefore, sparse weighting may reflect endometrial biopsy heterogeneity and not biologically uninformative. For example, asynchronous glands with mixed menstrual-stage epithelia have been observed in otherwise normal endometrial biopsies [54–56], which complicate phase assessment. Such within-slide heterogeneity may likewise affect attention interpretation in this fertility-oriented task. Further improvement requires deeper biological understanding of endometrial morphology and quantification of high-attention regions. One practical extension would be to quantify the interpretable histomics-centred features, such as nuclear morphology, cellular density, glandular architecture, and spatial cell-type relationships [57] in ACMIL-weighted patches, assessing whether high-attention regions are statistically associated with measurable morphological characteristics beyond visual inspection. Recent studies have demonstrated the use of AI-assisted endometrial cell annotation and targeted morphological analysis with explainable metrics such as epithelial-to-stromal ratio [58] and peri-gland stroma nuclear size [59]. However, no broadly available and cross-institutionally validated model yet reliably annotates LE, glandular epithelium, and stroma, while manual endometrial WSIs annotation remains impractical. Developing compartment-aware endometrial annotation models alongside outcome-oriented deep learning prediction models would enable objective attention-weighted histomics, interpretable handcrafted-feature baselines, and scalable ROI-based analysis. Complementing morphology and histomics-based approaches, recent work has aligned AI prediction heatmaps with Visium spatial transcriptomic profiles to validate molecular features associated with model-predicted regions [60]. Applying a similar strategy in endometrium could help determine whether attention patterns correspond to distinct cell states and potentially gene-expression profiles.
The LE-focused pipeline of analysis represents a manually defined, biologically informed ROI-based modelling strategy. The LE is the primary maternal-embryo interface during early implantation. Mouse in vivo studies, have highlighted that the LE undergoes morphological reshaping for implantation, affecting subcellular features including apical microvilli protrusions, apical tight junctions, and basolateral membranes [61–63]. However, capturing LE within a standard automated WSI patch generation pipeline is challenging because LE is sparse as compared to other cell types in endometrium. While only a minority of patients lacked identifiable LE, our dataset of LE-containing patches was substantially smaller than that of the standard WSI-based patches, with training set to be approximately 25 times smaller. The actual existence of LE within the WSI pipeline could be further compromised since the LE naturally has higher chance of falling into edge-case tiles dominated by empty spaces, which were discarded during the upstream patch generation procedures.
Beyond the LE, the adjacent peri-LE stroma may also represent a potentially underrepresented region worth specific emphasis. Histologically, the dense superficial stroma directly beneath the LE is the stratum compactum, whereas the deeper looser region is the stratum spongiosum [64]. Peri-LE stroma remained understudied until recent spatial transcriptomics suggests stromal WNT and NOTCH expressions may vary with distance from the luminal surface [65]. Therefore, our LE-containing ROI design was intended to enrich both LE and the potentially unique peri-LE stroma for separate evaluation. Surprisingly, despite the massive disparity in data volume of learning, models trained exclusively on LE-containing regions have achieved performance generally comparable to the WSI models. Together with the established biological roles of the LE and adjacent peri-LE stroma, these findings align with the expectation that this compartment represents a biologically enriched, morphologically informative reservoir for fertility-oriented endometrial modelling.
To visualize the morphological regions that were likely associated with model predictions, we employed Grad-CAM on the three best-performing LE-focused ResNet-18 models. This patch-level attention visualization was not extended to UNI2-h feature-based pipeline because reliable pixel-level saliency was difficult under slide-level feature aggregation. Semi-quantitative review showed that attention more often targeted LE-containing regions, either mainly the LE compartment or jointly with adjacent peri-LE stroma, while peri-LE stroma alone was less frequent. Unclassified artifacts or lumen-dominant regions were relatively rare. Thus, this localization tendency was consistent with the intended LE-focused design. Nevertheless, Grad-CAM provides only approximate spatial localization and does not specify the underlying mechanistic drivers. Reliable automated endometrial compartment annotation, together with more standardized attention quantification methods, would be essential for broader implementation of this interpretability framework and for more informative assessment of attention tendencies. Potentially, a multiscale design allowing more variable peri-LE context may also improve model robustness [57,66], as the physical boundary of the peri-LE stroma compartment has not yet been clearly defined by a fixed depth or anatomical cutoff [67,68].
Finally, multimodal integration was performed to explore the potential benefit of combining two histology-based models covering different regions of focus, with maternal clinical (age, BMI, and EnTh) and hormonal variables (E2 and P4). Such multimodal designs are increasingly explored in studies aiming to develop clinically useful predictive models [69]. Both WSI + LE and seven-factor models significantly outperformed chance and the clinical metadata-only model (Age + BMI + EnTh + E2 + P4). Standardized coefficients and SHAP analyses consistently ranked the two histology-model probabilities as the most impactful features. E2 and maternal age had small positive and negative associations, respectively; P4, EnTh, and BMI contributed minimally.
The inverse maternal age-live-birth association aligned with the known age-related decline [70]. The positive E2 and limited P4 associations were biologically plausible, consistent with a study that midluteal E2 levels were significantly associated with live birth rates once a minimum P4 level is reached [71]. The association between BMI and live birth remains cohort-dependent and controversial [72,73]. Limited P4 and EnTh weights do not imply biological irrelevance. Both variables are known to have minimal thresholds pregnancy supports [9,71,74]; thereby, a simpler threshold-based implementation may be more appropriate. Overall, these findings highlight the dominant contribution of histology-based predictions while demonstrating feasibility of incorporating complementary clinical variables. Future studies may explore earlier-stage multimodal integration to capture non-linear morphological–clinical relationships [57].
This study has several limitations. First, the patient-level sample size remained small although our cohort was larger than that of our previous HRT model [27]. Freezing the UNI2-h feature extractor reduced training complexity, but the low testing cohorts still weaken performance stability. Larger prospective cohorts are required to confirm the most suitable modelling strategy before wider implementation. Second, the external cohort was outcome-mismatched. Therefore, external performance should not be interpreted as direct validation of live birth prediction. Outcome-matched multicenter cohorts incorporating harmonization approaches (e.g., StainGAN) [75] are required to evaluate cross-center generalizability and mitigate technical variations. Third, the interpretability analyses remain limited to approximate spatial localization and could not support precise cellular or mechanistic interpretation from Grad-CAM saliency maps alone [76,77]. The manually defined LE workflow and lack of a validated endometrial compartment annotator further constrained scalable histomic analysis and the inclusion of a handcrafted morphological baseline for comparison.
From a clinical perspective, this predictive framework is best positioned as a supportive tool for pathologists and IVF clinicians when considering whether and when to proceed with embryo transfer. Unlike conventional dating, which is limited by biological and observer variability, this outcome-oriented pipeline offers a standardized, extendable, and quantifiable use of routine H&E images. On the other hand, the high cost of existing tests, such as transcriptomic array-based ERA, which costs US$800 or more per test, renders large prospective studies financially challenging [78]. The relatively lower cost of the proposed H&E image-based framework could practically facilitate larger validation and subsequent broader clinical implementation. Integrating biologically informed ROI modelling with automated histomic quantification could further enhance its comprehensiveness and interpretability. Of note, false positives could provide undue reassurance before embryo transfer, while false negatives could cause unnecessary delay or extra assessment. Thus, model thresholds should align with clinical priorities and undergo prospective validation. The clinical value of this framework relative to existing receptivity assessments also requires separate evaluation. With prospective validation and improved biological interpretation, it may enable more objective and personalized embryo-transfer planning.
In conclusion, this study showed that the histopathology pretrained UNI2-h foundation model could be effectively repurposed for endometrial fertility-related analysis without task-specific end-to-end training of the feature extractor. Although an endometrium-specific foundation model can be theoretically more ideal, such datasets are difficult to obtain in practice. Further outcome-matched external validation will be needed to formally assess generalizability for live-birth prediction. The LE-focused analyses further highlighted that a biologically informed design could retain performance comparable to WSI-based approaches despite using substantially fewer patches during development. Together, these findings support the LE and peri-LE regions and repurposed pathology models with ROI designs as promising starting points for endometrial histology analysis.
Supporting information
An ROC curve showing the performance of the reported 3-fold ensemble HRT-model on the natural cycle LH + 7 WSIs.
(TIF)
Summary figures illustrate progression of model predictive performance in AUROC (y-axis) over training epochs (x-axis). Performance is displayed for three cohorts including validation, internal testing, and external testing cohorts. All discussed models‘ performance trajectories were presented, including (A) ResNet-18 WSI model trajectory represented in blue solid line; (B) UNI2-h GAP-Linear and UNI2-h GAP-MLP model trajectories presented in blue and red solid lines; (C) UNI2-h ACMIL-Linear and UNI2-h ACMIL-MLP model trajectories presented in blue and red solid lines; (D) ResNet-18 LE and UNI2-h LE model trajectories presented in blue and red solid lines, with their WSI counter variants (ResNet-18 WSI and UNI2-h GAP-MLP model trajectories) plotted in purple and yellow dotted lines for easier comparison.
(TIF)
Representative heatmaps display patch-level attention tendency on WSIs for a correctly predicted CLB case (A) and Unsuccessful (B) case. Warmer colors indicate higher attention weights.
(TIF)
Summary of model performance comparisons between specified models, specifically in the internal outcome-based (A) and the external phase-based (B) testing cohorts. Δ value presents difference of AUROC between model A and B within each specified pair. Formal statistics performed with DeLong test corrected for multiple comparison (p < 0.05 indicated with *).
(TIF)
Calibration curves presented for each single reported models, including (A) ResNet-18 WSI, (B) UNI2-h GAP-Linear WSI, (C) UNI2-h GAP-MLP WSI, (D) UNI2-h ACMIL-Linear WSI, (E) UNI2-h ACMIL-MLP WSI, (F) ResNet-18 LE, and (G) UNI2-h LE model. Solid line presents model’s calibration curve on the internal test set; dotted line presents model’s calibration curve on the external set.
(TIF)
ROC curves for the Age + BMI + EnTh clinical metadata-only model (A) and the hormone-only E2 + P4 model (B). SHAP summary plot showing feature-contribution patterns across the internal test (C) and out-of-fold development samples (D) for the integrative seven-factor model. Positive SHAP values indicate a shift toward live birth prediction, negative SHAP values indicate a shift toward unsuccessful outcome prediction. Each dot represents one sample, with colour representing each samples’ original feature value, with red indicating a higher and blue representing a lower original value.
(TIF)
(XLSX)
(XLSX)
(XLSX)
(XLSX)
Data Availability
Sensitive raw clinical data and images are prohibited to be shared publicly due to legal and ethical restrictions. De-sensitized extracted features and checkpoints can be made available upon reasonable request to obsgyn@hku.hk. All codes essential to fully reproduce the training, testing, and heatmap visualizations are freely available on GitHub (https://github.com/AlexXuNB/Human_endometrial_fertility_histology_DL_study) and through an unrestricted Zenodo repository (DOI: 10.5281/zenodo.19730652). The repository includes a clear README instruction on how each analysis pipeline can be executed.
Funding Statement
This study was partly supported by the Health and Medical Research Fund (HMRF 04151546 to Yin Lau Lee; HMRF 10212996 to Yin Lau Lee) from the Food and Health Bureau, Government of the Hong Kong Special Administrative Region. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
References
- 1.Saket Z, Källén K, Lundin K, Magnusson Å, Bergh C. Cumulative live birth rate after IVF: trend over time and the impact of blastocyst culture and vitrification. Hum Reprod Open. 2021;2021(3):hoab021. doi: 10.1093/hropen/hoab021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Wang YA, Farquhar C, Sullivan EA. Donor age is a major determinant of success of oocyte donation/recipient programme. Hum Reprod. 2012;27(1):118–25. doi: 10.1093/humrep/der359 [DOI] [PubMed] [Google Scholar]
- 3.Pirtea P, De Ziegler D, Tao X, Sun L, Zhan Y, Ayoubi JM, et al. Rate of true recurrent implantation failure is low: results of three successive frozen euploid single embryo transfers. Fertil Steril. 2021;115(1):45–53. doi: 10.1016/j.fertnstert.2020.07.002 [DOI] [PubMed] [Google Scholar]
- 4.Ni Y, Shen H, Yao H, Zhang E, Tong C, Qian W, et al. Differences in fertility-related quality of life and emotional status among women undergoing different IVF treatment cycles. Psychol Res Behav Manag. 2023;16:1873–82. doi: 10.2147/PRBM.S411740 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Pasch LA, Gregorich SE, Katz PK, Millstein SG, Nachtigall RD, Bleil ME, et al. Psychological distress and in vitro fertilization outcome. Fertil Steril. 2012;98(2):459–64. doi: 10.1016/j.fertnstert.2012.05.023 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Lutjen P, Trounson A, Leeton J, Findlay J, Wood C, Renou P. The establishment and maintenance of pregnancy using in vitro fertilization and embryo donation in a patient with primary ovarian failure. Nature. 1984;307(5947):174–5. doi: 10.1038/307174a0 [DOI] [PubMed] [Google Scholar]
- 7.Vaegter KK, Lakic TG, Olovsson M, Berglund L, Brodin T, Holte J. Which factors are most predictive for live birth after in vitro fertilization and intracytoplasmic sperm injection (IVF/ICSI) treatments? Analysis of 100 prospectively recorded variables in 8,400 IVF/ICSI single-embryo transfers. Fertility and sterility. 2017;107(3):641–8. [DOI] [PubMed] [Google Scholar]
- 8.Gallos ID, Khairy M, Chu J, Rajkhowa M, Tobias A, Campbell A, et al. Optimal endometrial thickness to maximize live births and minimize pregnancy losses: analysis of 25,767 fresh embryo transfers. Reprod Biomed Online. 2018;37(5):542–8. doi: 10.1016/j.rbmo.2018.08.025 [DOI] [PubMed] [Google Scholar]
- 9.Mahutte N, Hartman M, Meng L, Lanes A, Luo Z-C, Liu KE. Optimal endometrial thickness in fresh and frozen-thaw in vitro fertilization cycles: an analysis of live birth rates from 96,000 autologous embryo transfers. Fertility and Sterility. 2022;117(4):792–800. [DOI] [PubMed] [Google Scholar]
- 10.Ng EHY, Chan CCW, Tang OS, Yeung WSB, Ho PC. The role of endometrial blood flow measured by three-dimensional power Doppler ultrasound in the prediction of pregnancy during in vitro fertilization treatment. Eur J Obstet Gynecol Reprod Biol. 2007;135(1):8–16. doi: 10.1016/j.ejogrb.2007.06.006 [DOI] [PubMed] [Google Scholar]
- 11.Noyes RW, Hertig AT, Rock J. Dating the endometrial biopsy. Obstet Gynecol Surv. 1950;5(4):561–4. doi: 10.1097/00006254-195008000-00044 [DOI] [PubMed] [Google Scholar]
- 12.Murray MJ, Meyer WR, Zaino RJ, Lessey BA, Novotny DB, Ireland K, et al. A critical analysis of the accuracy, reproducibility, and clinical utility of histologic endometrial dating in fertile women. Fertil Steril. 2004;81(5):1333–43. doi: 10.1016/j.fertnstert.2003.11.030 [DOI] [PubMed] [Google Scholar]
- 13.Coutifaris C, Myers ER, Guzick DS, Diamond MP, Carson SA, Legro RS, et al. Histological dating of timed endometrial biopsy tissue is not related to fertility status. Fertil Steril. 2004;82(5):1264–72. doi: 10.1016/j.fertnstert.2004.03.069 [DOI] [PubMed] [Google Scholar]
- 14.Arian SE, Hessami K, Khatibi A, To AK, Shamshirsaz AA, Gibbons W. Endometrial receptivity array before frozen embryo transfer cycles: a systematic review and meta-analysis. Fertil Steril. 2023;119(2):229–38. doi: 10.1016/j.fertnstert.2022.11.012 [DOI] [PubMed] [Google Scholar]
- 15.Díaz-Gimeno P, Ruiz-Alonso M, Blesa D, Bosch N, Martínez-Conejero JA, Alamá P, et al. The accuracy and reproducibility of the endometrial receptivity array is superior to histology as a diagnostic method for endometrial receptivity. Fertil Steril. 2013;99(2):508–17. doi: 10.1016/j.fertnstert.2012.09.046 [DOI] [PubMed] [Google Scholar]
- 16.Cohen AM, Ye XY, Colgan TJ, Greenblatt EM, Chan C. Comparing endometrial receptivity array to histologic dating of the endometrium in women with a history of implantation failure. Syst Biol Reprod Med. 2020;66(6):347–54. doi: 10.1080/19396368.2020.1824032 [DOI] [PubMed] [Google Scholar]
- 17.Pantanowitz L, Sharma A, Carter AB, Kurc T, Sussman A, Saltz J. Twenty years of digital pathology: an overview of the road travelled, what is on the horizon, and the emergence of vendor-neutral archives. J Pathol Inform. 2018;9:40. doi: 10.4103/jpi.jpi_69_18 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Srinidhi CL, Ciga O, Martel AL. Deep neural network models for computational histopathology: A survey. Med Image Anal. 2021;67:101813. doi: 10.1016/j.media.2020.101813 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Tellez D, Litjens G, Bándi P, Bulten W, Bokhorst J-M, Ciompi F, et al. Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology. Med Image Anal. 2019;58:101544. doi: 10.1016/j.media.2019.101544 [DOI] [PubMed] [Google Scholar]
- 20.Lu MY, Chen B, Williamson DFK, Chen RJ, Liang I, Ding T, et al. A visual-language foundation model for computational pathology. Nat Med. 2024;30(3):863–74. doi: 10.1038/s41591-024-02856-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Chen RJ, Ding T, Lu MY, Williamson DFK, Jaume G, Song AH, et al. Towards a general-purpose foundation model for computational pathology. Nat Med. 2024;30(3):850–62. doi: 10.1038/s41591-024-02857-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Vorontsov E, Bozkurt A, Casson A, Shaikovski G, Zelechowski M, Liu S. Virchow: A million-slide digital pathology foundation model. 2023. https://arxiv.org/abs/230907778
- 23.Lu MY, Williamson DFK, Chen TY, Chen RJ, Barbieri M, Mahmood F. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat Biomed Eng. 2021;5(6):555–70. doi: 10.1038/s41551-020-00682-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Dimitriadis I, Zaninovic N, Badiola AC, Bormann CL. Artificial intelligence in the embryology laboratory: a review. Reprod Biomed Online. 2022;44(3):435–48. doi: 10.1016/j.rbmo.2021.11.003 [DOI] [PubMed] [Google Scholar]
- 25.Ehteshami Bejnordi B, Veta M, Johannes van Diest P, van Ginneken B, Karssemeijer N, Litjens G, et al. Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer. JAMA. 2017;318(22):2199–210. doi: 10.1001/jama.2017.14585 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 770–8. 10.1109/cvpr.2016.90 [DOI]
- 27.Li T, Liao R, Chan C, Greenblatt EM. Deep learning analysis of endometrial histology as a promising tool to predict the chance of pregnancy after frozen embryo transfers. J Assist Reprod Genet. 2023;40(4):901–10. doi: 10.1007/s10815-023-02745-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Adams SM, Terry V, Hosie MJ, Gayer N, Murphy CR. Endometrial response to IVF hormonal manipulation: comparative analysis of menopausal, down regulated and natural cycles. Reprod Biol Endocrinol. 2004;2:21. doi: 10.1186/1477-7827-2-21 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Pan Y, Li B, Wang Z, Wang Y, Gong X, Zhou W, et al. Hormone replacement versus natural cycle protocols of endometrial preparation for frozen embryo transfer. Front Endocrinol (Lausanne). 2020;11:546532. doi: 10.3389/fendo.2020.546532 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Lee YL, Ruan H, Lee KC, Fong SW, Yue C, Chen ACH, et al. Attachment of a trophoblastic spheroid onto endometrial epithelial cells predicts cumulative live birth in women aged 35 and older. Fertil Steril. 2023;120(2):268–76. doi: 10.1016/j.fertnstert.2023.03.013 [DOI] [PubMed] [Google Scholar]
- 31.Cao D, Liu Y, Cheng Y, Wang J, Zhang B, Zhai Y, et al. Time-series single-cell transcriptomic profiling of luteal-phase endometrium uncovers dynamic characteristics and its dysregulation in recurrent implantation failures. Nat Commun. 2025;16(1):137. doi: 10.1038/s41467-024-55419-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Gardner DK, Schoolcraft WB. Culture and transfer of human blastocysts. Curr Opin Obstet Gynecol. 1999;11(3):307–11. doi: 10.1097/00001703-199906000-00013 [DOI] [PubMed] [Google Scholar]
- 33.Zegers-Hochschild F, Adamson GD, Dyer S, Racowsky C, de Mouzon J, Sokol R. The International glossary on infertility and fertility care, 2017. Human Reproduction. 2017;32(9):1786–801. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.An BGL, Chapman M, Tilia L, Venetis C. Is there an optimal window of time for transferring single frozen-thawed euploid blastocysts? A cohort study of 1170 embryo transfers. Hum Reprod. 2022;37(12):2797–807. doi: 10.1093/humrep/deac227 [DOI] [PubMed] [Google Scholar]
- 35.Muller D, Soto-Rey I, Kramer F. An analysis on ensemble learning optimized medical image classification with deep convolutional neural networks. IEEE Access. 2022;10:66467–80. doi: 10.1109/access.2022.3182399 [DOI] [Google Scholar]
- 36.Zhang Y, Li H, Sun Y, Zheng S, Zhu C, Yang L. Attention-challenging multiple instance learning for whole slide image classification. In: European conference on computer vision, 2024.
- 37.Zhang A, Jaume G, Vaidya A, Ding T, Mahmood F. Accelerating data processing and benchmarking of ai models for pathology. 2025.
- 38.Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In: 2017 IEEE International Conference on Computer Vision (ICCV). 2017. 618–26. 10.1109/iccv.2017.74 [DOI]
- 39.Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, et al. Pytorch: An imperative style, high-performance deep learning library. Adv Neural Inform Process Syst. 2019;32. [Google Scholar]
- 40.Sun X, Xu W. Fast implementation of delong’s algorithm for comparing the areas under correlated receiver operating characteristic curves. IEEE Signal Process Lett. 2014;21(11):1389–93. doi: 10.1109/lsp.2014.2337313 [DOI] [Google Scholar]
- 41.Neha NF. Radiomics in medical imaging: methods, applications, and challenges. 2026. https://doi.org/arXiv:260200102 [DOI] [PMC free article] [PubMed]
- 42.Howard FM, Dolezal J, Kochanny S, Schulte J, Chen H, Heij L, et al. The impact of site-specific digital histology signatures on deep learning model accuracy and bias. Nat Commun. 2021;12(1):4423. doi: 10.1038/s41467-021-24698-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Marečková M, Garcia-Alonso L, Moullet M, Lorenzi V, Petryszak R, Sancho-Serra C, et al. An integrated single-cell reference atlas of the human endometrium. Nat Genet. 2024;56(9):1925–37. doi: 10.1038/s41588-024-01873-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Wahab M, Thompson J, Hamid B, Deen S, Al-Azzawi F. Endometrial histomorphometry of trimegestone-based sequential hormone replacement therapy: a weighted comparison with the endometrium of the natural cycle. Hum Reprod. 1999;14(10):2609–18. doi: 10.1093/humrep/14.10.2609 [DOI] [PubMed] [Google Scholar]
- 45.Takahashi S, Sakaguchi Y, Kouno N, Takasawa K, Ishizu K, Akagi Y, et al. Comparison of vision transformers and convolutional neural networks in medical image analysis: a systematic review. J Med Syst. 2024;48(1):84. doi: 10.1007/s10916-024-02105-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Chen J, Pan T, Zhu Z, Liu L, Zhao N, Feng X, et al. A deep learning-based multimodal medical imaging model for breast cancer screening. Sci Rep. 2025;15(1):14696. doi: 10.1038/s41598-025-99535-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Bareja R, Carrillo-Perez F, Zheng Y, Pizurica M, Nandi TN, Shen J. Evaluating vision and pathology foundation models for computational pathology: A comprehensive benchmark study. medRxiv. 2025. doi: 10.1101/2025.05.08.25327250 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Li J, Hu J, Sun Q, Yan R, Ouyang M, Guan T, et al. Can we simplify slide-level fine-tuning of pathology foundation models?. 2025. https://arxiv.org/abs/250220823
- 49.Mishra V, Lotter W. Comparing computational pathology foundation models using representational similarity analysis. 2025. https://doi.org/arXiv:250915482 [PMC free article] [PubMed]
- 50.Ding T, Wagner SJ, Song AH, Chen RJ, Lu MY, Zhang A, et al. A multimodal whole-slide foundation model for pathology. Nat Med. 2025;31(11):3749–61. doi: 10.1038/s41591-025-03982-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Ilse M, Tomczak J, Welling M. Attention-based deep multiple instance learning. In: International conference on machine learning. 2018.
- 52.Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW, Topic Group ‘Evaluating diagnostic tests and prediction models’ of the STRATOS initiative. Calibration: The Achilles heel of predictive analytics. BMC Med. 2019;17(1):230. doi: 10.1186/s12916-019-1466-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Binuya MAE, Engelhardt EG, Schats W, Schmidt MK, Steyerberg EW. Methodological guidance for the evaluation and updating of clinical prediction models: A systematic review. BMC Med Res Methodol. 2022;22(1):316. doi: 10.1186/s12874-022-01801-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Thornburgh I, Anderson MC. The endometrial deficient secretory phase. Histopathology. 1997;30(1):11–5. doi: 10.1046/j.1365-2559.1997.d01-554.x [DOI] [PubMed] [Google Scholar]
- 55.Stewart CJR, Bharat C, Leake R. Asynchronous glands in secretory pattern endometrium: Clinical associations and immunohistological changes. Histopathology. 2015;67(1):39–47. doi: 10.1111/his.12620 [DOI] [PubMed] [Google Scholar]
- 56.Russell P, Hey-Cunningham A, Berbic M, Tremellen K, Sacks G, Gee A, et al. Asynchronous glands in the endometrium of women with recurrent reproductive failure. Pathology. 2014;46(4):325–32. doi: 10.1097/PAT.0000000000000111 [DOI] [PubMed] [Google Scholar]
- 57.Neha F, Bhati D, Shukla DK. Phenotyping of histology imaging data with histomics. AI. 2026;7(6):228. doi: 10.3390/ai7060228 [DOI] [Google Scholar]
- 58.Lee S, Arffman RK, Komsi EK, Lindgren O, Kemppainen J, Kask K, et al. Dynamic changes in AI-based analysis of endometrial cellular composition: Analysis of PCOS and RIF endometrium. J Pathol Inform. 2024;15:100364. doi: 10.1016/j.jpi.2024.100364 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Raudonis V, Bartasiene R, Minajeva A, Saare M, Drejeriene E, Kozlovskaja-Gumbriene A, et al. Towards metric-driven difference detection between receptive and nonreceptive endometrial samples using automatic histology image analysis. Appl Sci. 2024;14(13):5715. doi: 10.3390/app14135715 [DOI] [Google Scholar]
- 60.Zeng Q, Klein C, Caruso S, Maille P, Allende DS, Mínguez B, et al. Artificial intelligence-based pathology as a biomarker of sensitivity to atezolizumab-bevacizumab in patients with hepatocellular carcinoma: a multicentre retrospective study. Lancet Oncol. 2023;24(12):1411–22. doi: 10.1016/S1470-2045(23)00468-0 [DOI] [PubMed] [Google Scholar]
- 61.Murphy CR. Uterine receptivity and the plasma membrane transformation. Cell Res. 2004;14(4):259–67. doi: 10.1038/sj.cr.7290227 [DOI] [PubMed] [Google Scholar]
- 62.Ye X. Uterine luminal epithelium as the transient gateway for embryo implantation. Trends Endocrinol Metab. 2020;31(2):165–80. doi: 10.1016/j.tem.2019.11.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Jones-Paris CR, Paria S, Berg T, Saus J, Bhave G, Paria BC, et al. Embryo implantation triggers dynamic spatiotemporal expression of the basement membrane toolkit during uterine reprogramming. Matrix Biol. 2017;57–58:347–65. doi: 10.1016/j.matbio.2016.09.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Tandulwadkar S, Pal B. Hysteroscopy simplified by masters. Springer; 2020. [Google Scholar]
- 65.Mathyk B, Schwartz A, DeCherney A, Ata B. A critical appraisal of studies on endometrial thickness and embryo transfer outcome. Reprod Biomed Online. 2023;47(4):103259. doi: 10.1016/j.rbmo.2023.103259 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Deng R, Cui C, Remedios LW, Bao S, Womick RM, Chiron S. Cross-scale multi-instance learning for pathological image diagnosis. Medical Image Analysis. 2024;94:103124. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Tempest N, Soul J, Hill CJ, Caamaño Gutierrez E, Hapangama DK. Cell type and region-specific transcriptional changes in the endometrium of women with RIF identify potential treatment targets. Proc Natl Acad Sci U S A. 2025;122(11):e2421254122. doi: 10.1073/pnas.2421254122 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Garcia-Alonso L, Handfield L-F, Roberts K, Nikolakopoulou K, Fernando RC, Gardner L, et al. Mapping the temporal and spatial dynamics of the human endometrium in vivo and in vitro. Nat Genet. 2021;53(12):1698–711. doi: 10.1038/s41588-021-00972-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Boehm KM, Khosravi P, Vanguri R, Gao J, Shah SP. Harnessing multimodal data integration to advance precision oncology. Nat Rev Cancer. 2022;22(2):114–26. doi: 10.1038/s41568-021-00408-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.McLernon DJ, Steyerberg EW, Te Velde ER, Lee AJ, Bhattacharya S. Predicting the chances of a live birth after one or more complete cycles of in vitro fertilisation: population based study of linked cycle data from 113 873 women. BMJ. 2016;355. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Alsbjerg B, Jensen MB, Elbaek HO, Laursen R, Povlsen BB, Anderson R, et al. Midluteal serum estradiol levels are associated with live birth rates in hormone replacement therapy frozen embryo transfer cycles: a cohort study. Fertil Steril. 2024;121(6):1000–9. doi: 10.1016/j.fertnstert.2024.04.006 [DOI] [PubMed] [Google Scholar]
- 72.Zheng Z, Zhang X, Wu F, Liao H, Zhao H, Zhang M, et al. Effect of BMI on cumulative live birth rates in patients that completed IVF treatment: a retrospective cohort study of 16,126 patients. Endocr Connect. 2024;13(3):e230105. doi: 10.1530/EC-23-0105 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Beshar I, Milki AA, Gardner RM, Zhang WY, Johal JK, Bavan B. Elevated body mass index in modified natural cycle frozen euploid embryo transfers is not associated with live birth rate. J Assisted Reprod Genetics. 2023;40(5):1055–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Takaya Y, Matsubayashi H, Kitaya K, Nishiyama R, Yamaguchi K, Takeuchi T, et al. Minimum values for midluteal plasma progesterone and estradiol concentrations in patients who achieved pregnancy with timed intercourse or intrauterine insemination without a human menopausal gonadotropin. BMC Res Notes. 2018;11(1):61. doi: 10.1186/s13104-018-3188-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Shaban MT, Baur C, Navab N, Albarqouni S. Staingan: stain style transfer for digital histological images. In: 2019 IEEE 16th International Symposium on Biomedical Imaging (ISBI 2019), 2019. 953–6. 10.1109/isbi.2019.8759152 [DOI]
- 76.Adebayo J, Gilmer J, Muelly M, Goodfellow I, Hardt M, Kim B. Sanity checks for saliency maps. Adv Neural Inform Process Syst. 2018;31. [Google Scholar]
- 77.Draelos RL, Carin L. Use HiResCAM instead of Grad-CAM for faithful explanations of convolutional neural networks. arXiv preprint. 2020. 10.48550/arXiv.2011.08891 [DOI]
- 78.Lensen S, Shreeve N, Barnhart KT, Gibreel A, Ng EHY, Moffett A. In vitro fertilization add-ons for the endometrium: it doesn’t add-up. Fertility and Sterility. 2019;112(6):987–93. [DOI] [PubMed] [Google Scholar]
