Abstract
Hypoxic Ischemic Encephalopathy (HIE) represents a brain dysfunction, affecting approximately 1 to 5 per 1000 full-term neonates. The precise delineation and segmentation of HIE-related lesions in neonatal brain Magnetic Resonance Images (MRI) are pivotal in advancing outcome predictions, identifying patients at high risk, elucidating neurological manifestations, and assessing treatment efficacies. Despite its importance, the development of algorithms for segmenting HIE lesions from MRI volumes has been impeded by data scarcity. Addressing this critical gap, we organized the first BONBID-HIE challenge with diffusion MRI data (Apparent Diffusion Coefficient (ADC) maps) for HIE lesion segmentation, in conjunction with the MICCAI 2023. Totally 14 algorithms were submitted, employing a gamut of cutting-edge automatic machine-learning-based segmentation algorithms. Our comprehensive analysis of HIE lesion segmentation and submitted algorithms facilitates an in-depth evaluation of the current technological zenith, outlines directions for future advancements, and highlights persistent hurdles. To foster ongoing research and benchmarking, the annotated HIE dataset, developed algorithm dockers, and unified evaluation codes are accessible through a dedicated online platform (https://bonbid-hie2023.grand-challenge.org).
Index Terms—: Brain injury, Lesion segmentation, Machine Learning, MRI, Challenge, Benchmark, Hypoxic Ischemic Encephalopathy, Algorithm Comparison, Algorithm development
I. Introduction
Neonatal hypoxic-ischemic encephalopathy (HIE) remains a significant public health concern, characterized by brain injury due to insufficient blood and oxygen supply to the brain. HIE affects approximately 1 to 5 per 1000 termborn neonates worldwide each year, with an estimated annual cost exceeding $2 billion in the United States alone, not accounting for the substantial burden on affected families [1]–[3]. Despite the adoption of Therapeutic Hypothermia (TH) as the standard of care, there is a substantial proportion of HIE patients (35%−50%) experiencing adverse neurocognitive outcomes [4]–[6]. Reducing mortality and morbidity associated with HIE is, therefore, a critical public health objective. Globally, there are currently 180 ongoing clinical trials related to HIE, spanning 33 countries and five continents (Figure 1, data retrieved on clinicaltrials.gov on 11/2024) [7]–[13]. Early identification of patients at high risk for adverse outcomes sooner after initiation of therapy remains a significant challenge, as clinical outcomes are often not reliably measurable until the age of two years [14]. This highlights the urgent need for accurate and reliable biomarkers to enable an early prognosis and outcome prediction.
Fig. 1.

There are 180 ongoing trials related to HIE spreading over 33 countries and 5 continents (data retrieved on clinicaltrials.gov on 11/2024).
Accurate identification and segmentation of HIE-related lesions in neonatal brain magnetic resonance images (MRIs) is a critical step toward this goal. Clinical trials often rely on the NRN scoring system [15], [16], which is based on expert assessments of lesion extent and location, to predict neurocognitive outcomes. Lesion patterns in regions such as the basal ganglia, thalamus, and watershed areas are associated with distinct neurocognitive impairments, including motor, language, and executive dysfunction [16]–[19].
Machine learning, especially deep learning, in the detection and segmentation of lesions remains largely unexplored in HIE. Developing robust segmentation algorithms is challenging due to two interrelated obstacles:(i) Data scarcity and annotation limitations. Compared to neurological disorders such as brain tumors [23], [24], Alzheimer’s disease [25], [26], and ischemic stroke [21], [27], [28], neonatal HIE lacks sufficient publicly available imaging benchmarks. High-quality datasets containing annotated MRIs alongside clinical and outcome information remain rare. One reason is that integrating imaging with longitudinal clinical and neurodevelopmental outcomes requires long-term follow-up, which is often challenging in both clinical and research settings. Additionally, expert annotation of HIE lesions demands specialized knowledge of neonatal neuroimaging, further limiting the scalability of dataset curation. (ii) Algorithmic challenges posed by lesion characteristics. Unlike tumors [22], HIE lesions are typically small (< 1%) and diffuse (multi-focal) [20], [29], with over half of our patient cohort exhibiting such characteristics. This makes segmentation of HIE MRI data more challenging compared to tasks involving adult brain tumors, where lesions are generally larger and more focal, as illustrated in Figure 2. These characteristics lead to extreme class imbalance and an increased risk of false negatives in conventional segmentation pipelines.
Fig. 2.

Comparison of small, diffuse lesions in HIE [20] with large, focal lesions in ISLES-SPES (acute stroke outcome/penumbra estimation) [21] and BraTS [22]. The boxplot quantifies the percentage of brain volume affected by lesions across all patients in each dataset.
To tackle these challenges, we collected a cohort of 133 cases with expert-annotated lesions over a decade of dedicated research [1], [20], forming the the Boston Neonatal Brain Injury Dataset for Hypoxic-Ischemic Encephalopathy (BONBID-HIE) [20]. To accelerate advancements in this domain, we organized the BONBID-HIE Lesion Segmentation Challenge. This challenge provided a direct, fair and independently controlled comparison of automated methods on this rigorously curated public dataset and platform. The challenge was conducted as a satellite event at the 26th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2023) in Vancouver, Canada. The event garnered global interest, with more than 140 registrations and 14 successful submissions in the test phase, including docker containers for algorithms and comprehensive method descriptions, which were published in our workshop proceedings [30].
In summary, this paper introduces the BONBID-HIE Lesion Segmentation Challenge. Despite the limited availability of data, the ultimate goal was to alleviate the burden on medical experts in diagnosing and prognosing HIE. The challenge includes the publicly accessible BONBID-HIE dataset, the submitted algorithm containers and their corresponding results, and the accompanying online validation tools as ongoing benchmarking resources. Our primary goal was to provide brain MRIs for HIE to inspire new research directions in HIE lesion segmentation. The event provided a unique platform for participants to explore machine learning methods and their practical applications in segmenting small and diffuse lesions in HIE. Concurrently, the challenge aimed to foster interdisciplinary research and collaboration, further advancing efforts in HIE outcome prediction. Additionally, we hope that this dataset will also be valuable for other small and diffuse lesion segmentation tasks.
The paper is organized as follows: Section 2 describes the BONBID-HIE challenge, including data characteristics, manual lesion annotations, and the evaluation process. Section 3 presents the challenge results, along with a statistical analysis of the performance of different algorithms. Section 4 discusses failure cases, method limitations, and future directions for HIE challenges. Section 5 concludes the paper.
II. 1st BONBID-HIE Lesion Segmentation Challenge
The 1st BONBID-HIE Lesion Segmentation Challenge was held as an online challenge and workshop in conjunction with MICCAI 2023. It was designed to promote continuous engagement by allowing new groups to access the training and test data, submit their segmentations, and automatically compare and rank their results against all previous submissions.
A. Data Settings and Annotations
Setting.
133 cases were retrospectively collected from a cohort of neonates diagnosed with HIE at Massachusetts General Hospital with Institutional Review Board approval from Massachusetts General Hospital and Boston Children’s Hospital to anonymize, curate, annotate, and release them [20]. The image data was divided into three sets: 85 cases for training, 4 cases for docker sanity validation, and 44 cases for testing. For each case, the provided inputs included Apparent Diffusion Coefficient (ADC) maps and ZADC, which were used as inputs for the algorithm containers. The output of the algorithm was the corresponding binary lesion segmentation map for each case.
Image Preprocessing.
MRIs were acquired on either GE 1.5T Signa scanner (N=52, scanned during 2001–2012) or SIEMENS 3T Trio scanner (N=81, scanned during 2012–2018). Diffusion tensor imaging has the protocol as follows: and 30 diffusion directions (SIEMENS scanner) or and 6 diffusion directions (GE scanner). The MRI preprocessing pipeline applied in this study was previously established and validated for the BONBID-HIE dataset [20]. Specifically, preprocessing included N4 bias field correction to address intensity nonuniformities [31], field-of-view normalization to ensure consistent spatial coverage across images [32], and multi-atlas skull stripping tailored explicitly for ADC maps [33]. For additional details, please refer to [20]. All images retained their original voxel sizes without resampling. Consequently, segmentation masks were evaluated in each case’s voxel space. This preserved resolution, avoided interpolation artifacts, and maintained clinically relevant spatial accuracy.
ADC Maps.
Diffusion-weighted imaging (DWI) is widely used in brain MRI to probe the microstructural properties of tissue by sensitizing the MRI signal to water molecule motion [34]. The ADC is a quantitative metric derived from DWI that reflects the magnitude of water diffusion within a voxel. ADC is computed from DWI acquired with different b-values by modeling the signal attenuation according to a mono-exponential relationship: , where and denote the signal intensities at b-value and at , respectively. Taking the natural logarithm of this equation and fitting a linear regression across the acquired b-values yields the ADC estimate: . DWI and the resulting ADC maps are an appropriate modality to detect hypoxic ischemic brain injury in neonates diagnosed with HIE in the first few days after therapeutic hypothermia. Hypoxic ischemic injury leads to restricted water diffusion, which appears bright on DWI and dark on ADC maps, with these signal changes occurring earlier and being more clearly visible than signal changes on conventional T1- or T2-weighted sequences [35]–[38]. Therefore, DWI and ADC maps offer higher early sensitivity in detecting HIE related brain injury [39]–[42].
ZADC Maps.
We developed ZADC maps to normalize and make ADC values comparable across brain voxel locations [20]. ZADC maps quantify location-specific deviations from normal, which is important for abnormal region segmentation. To generate ZADC maps, the following steps are performed: (1) A normative ADC atlas is constructed from the scans of 13 neonates, capturing the mean and standard deviation of ADC values at each voxel [43]. (2) A deformation field is computed, which maps each voxel in the patient’s ADC map to its anatomically corresponding location in the atlas space. The normal range of ADC variation is defined by the mean and standard deviation for each voxel. (3) The patient’s ADC value at voxel is converted to a -score:
| (1) |
Therefore, the ZADC value at voxel location quantifies how many standard deviations the patient’s ADC value deviates from the mean normal ADC value at that anatomical location. For details on this process, please refer to [20].
Annotations.
HIE lesions were initially identified in radiology reports using ADC maps and structural MRIs as primary imaging sequences [44], [45]. To translate these descriptions into voxel-level annotations, a two-step expert consensus approach was employed. First, lesions were manually annotated on ADC maps based on neuroradiology reports, with adjustments for integrity across axial, coronal, and sagittal planes. In cases of uncertainty or disagreement, consensus was reached through discussions among three experienced pediatric neuro-radiologists. This multi-expert consensus process provided a robust and unbiased set of lesion annotations, marking the first ADC-based HIE lesion annotations [20].
B. Dataset Characteristics
Lesion Statistics.
Figure 3 illustrates the spatial distribution of HIE lesions in neonates, presented as a frequency map overlaid on a normal brain atlas. The overlay on the normal atlas provides a clear visual comparison, illustrating the common regions impacted by HIE and their deviation from typical brain ADC values. As shown in the figure, HIE lesions are predominantly concentrated in the thalamus, perirolandic cortex, expected location of the corticospinal tract, midbrain and superior vermis.
Fig. 3.

Statistical lesion atlas quantifying the voxel-wise lesion frequency in our cohort of N=133 patients in the normal 0–14 days ADC atlas space from BONBID-HIE dataset [20].
Lesion Distribution.
All Lesion percentages in our paper refer to the proportion of predicted lesion volume relative to the whole brain volume. As shown in Table I, more than 50% of HIE lesions are < 1% of the brain. This imbalance challenges machine learning algorithms, limiting performance optimization. In Figure 4, we illustrate the varying percentages of brain injury in HIE, displaying corresponding ADC, ZADC, and lesion segmentation maps.
TABLE I.
Distribution of brain injury extent in neonates with HIE
| Lesion percentage | # of cases | Percentage |
|---|---|---|
| <1% | 74 | 55.64% |
| [1%, 5%) | 26 | 19.55% |
| [5%, 100%) | 33 | 5.26% |
Fig. 4.

Representative cases from the BONBID-HIE dataset visualized across a spectrum of brain injury percentages. Each row shows axial slices from two patients with similar lesion percentage categories. For each patient, three images are shown: the original ADC map, the corresponding ZADC map, and the expert-annotated ground truth (GT) lesion mask. Lesions are categorized based on their volume relative to the whole brain: < 1%, [1%, 5%), [5%, 50%), and [50%, 100%). The percent brain injury is indicated below each case. This figure illustrates the wide variability in lesion extent and appearance, highlighting the challenge of segmenting small and multifocal HIE lesions.
C. Training and Testing
The postmenstrual age (PMA) at the time of MRI scan was 4.2 ± 2.6 days for the training set and 3.3 ± 1.3 days for the testing set. Due to the limited number of cases, we prioritized balancing lesion percentage and scanner types across the training and testing subsets. Lesion distribution was stratified as follows: in the training set, 46 small lesions (< 1%), 16 medium lesions (1–5%), and 25 large lesions (> 5%); and in the testing set, 22 small, 12 medium, and 10 large lesions (Table II). The training and testing data were also split while taking into account two scanner types (SIEMENS 3T and GE 1.5T) to ensure variability and robustness under different acquisition conditions.
TABLE II.
Number of patients scanned by different scanners in patient subgroups with different lesion volumes.
| Split | Scanner | Total | |||
|---|---|---|---|---|---|
| < 1% | [1%, 5%] | > 5% | |||
| Train&Val | GE 1.5T | 15 | 2 | 16 | 33 |
| SIEMENS 3T | 37 | 12 | 7 | 56 | |
| Test | GE 1.5T | 8 | 5 | 6 | 19 |
| SIEMENS 3T | 14 | 7 | 4 | 25 |
D. Data and Evaluation Availability
One of the primary objectives of the BONBID-HIE is to provide an open-source repository to support the continuous development of algorithms. The BONBID-HIE lesion segmentation dataset is publicly available through Zenodo1. A unified portal for accessing benchmarks, submitting algorithm Docker containers, and performing automatic evaluations is available at challenge website2. All submitted algorithm dockers can also be accessed here. The evaluation codes are available on the challenge’s GitHub repository3.
E. Thresholding ZADC as baseline method
We provide a baseline method for the BONBID-HIE 2023 challenge. This method was originally described in our dataset paper [20]. ZADC quantifies how many standard deviations the patient’s ADC value deviates from the mean normal ADC value at the corresponding anatomical location. As reported in [20], by simply thresholding ZADC at −2, the DICE for lesion prediction using this method achieves a value of 0.54 on the entire dataset (N=133). When evaluated on this test set (N=44), the DICE score of ZADC (−2) is 0.58 (Table IV).
TABLE IV.
(A) Evaluation of submitted methods on HIE lesion segmentation using Dice, MASD, and NSD metrics. The table ranks 14 participating algorithms and the baseline. (B) Label fusion using MV, STAPLE, and SBA algorithms are presented for the top 2, 5, 10, and all methods. Greyed-out rows indicate the top-performing methods for each corresponding metric or subgroup.
| Ranking | Participants | DICE (↑) | MASD (↓) | NSD (↑) |
|---|---|---|---|---|
| 1 | civalab | 62.2 ± 24.4% | 2.2 ± 2.5 mm | 75.6 ± 24.1% |
| 2 | xleratorxlerator9 | 62.3 ± 23.9% | 2.4 ± 2.6 mm | 74.9 ± 23.6% |
| 3 | frimpz | 57.4 ± 23.9% | 2.7 ± 3.3 mm | 73.4 ± 24.9% |
| 4 | schlauglab | 58.0 ± 25.6% | 2.6 ± 3.0 mm | 72.7 ± 24.8% |
| 5 | rajroy | 57.1 ± 24.1% | 2.7 ± 3.4 mm | 73.2 ± 25.1% |
| 6 | IWM | 57.7 ± 25.1% | 2.9 ± 3.6 mm | 72.2 ± 25.2% |
| 7 | imad.toubal | 53.4 ± 24.3% | 2.5 ± 2.6 mm | 71.6 ± 24.3% |
| 8 | UNeImage | 56.3 ± 25.4% | 3.0 ± 3.3 mm | 71.5 ± 24.8% |
| 9 | ngzvh | 53.8 ± 24.6% | 3.0 ± 3.3 mm | 69.9 ± 25.2% |
| 10 | ashwin_dhakal | 49.1 ± 25.1% | 3.4 ± 3.7 mm | 67.7 ± 24.5% |
| 11 | civa | 50.0 ± 26.3% | 3.5 ± 3.4 mm | 67.8 ± 23.7% |
| 12 | arda.aydn | 48.3 ± 22.8% | 3.5 ± 3.1 mm | 61.8 ± 24.0% |
| 13 | punithakumar | 50.0 ± 28.1% | 5.2 ± 9.4 mm | 62.7 ± 28.9% |
| 14 | tiansong_philips | 40.7 ± 25.7% | 4.7 ± 7.2 mm | 57.8 ± 26.6% |
| Baseline | ZADC (−2) | 58.3 ± 26.4% | 3.2 ± 3.3 mm | 67.1 ± 25.0% |
| (b) Accuracies by the consensus of algorithms. | ||||
| MV fusion [59] | 61.5 ± 23.5% | 2.4 ± 2.5 mm | 74.5 ± 23.3% | |
| Top 2 | STAPLE fusion [60] | 61.8 ± 24.2% | 2.3 ± 2.4 mm | 75.3 ± 23.9% |
| SBA fusion [61] | 63.0 ± 23.7% | 2.3 ± 2.5 mm | 76.2 ± 24.2% | |
| Top 5 | MV fusion | 62.9 ± 23.8% | 2.3 ± 2.9 mm | 76.5 ± 24.7% |
| STAPLE fusion | 60.7 ± 23.4% | 2.4 ± 2.7 mm | 75.0 ± 23.8% | |
| SBA fusion | 61.7 ± 24.7% | 2.6 ± 3.3 mm | 74.1 ± 26.3% | |
| Top 10 | MV fusion | 61.8 ± 23.8% | 2.4 ± 3.2 mm | 76.1 ± 24.6% |
| STAPLE fusion | 56.8 ± 24.2% | 2.6 ± 2.8 mm | 71.8 ± 24.2% | |
| SBA fusion | 59.7 ± 26.1% | 3.0 ± 3.8 mm | 72.3 ± 27.4% | |
| all | MV fusion | 61.9 ± 24.3% | 2.5 ± 3.3 mm | 75.9 ± 25.1% |
| STAPLE fusion | 54.5 ± 24.8% | 2.7 ± 3.0 mm | 70.2 ± 24.5% | |
| SBA fusion | 54.3 ± 30.8% | 4.1 ± 6.8 mm | 64.2 ± 32.6% | |
F. Evaluation metrics and ranking
Evaluation metrics.
Evaluation metrics are DICE coefficient (DICE), and Mean Average Surface Distance (MASD) [53] and Normalized Surface Distance (NSD) [53] in different percentages of lesion (< 1%, 1% ~ 5%, > 5%).
DICE.
The DICE coefficient evaluates the overlap between two segmented regions, measuring the similarity between the predicted and ground truth segments, with a higher value indicating better agreement. DICE is computed as,
| (2) |
where and are the predicted segmentation lesion volume and ground truth lesion volume, respectively.
MASD.
MASD measures the average distance between the surfaces of the predicted and ground truth segmentations, assessing how closely the boundaries align, with a lower value indicating better performance. The MASD between two sets of boundary points of and of is calculated as:
| (3) |
where: and . is the Euclidean distance between points and . and denote the number of points in sets and , respectively. is the average of the minimum distances from each point in to the set . is the average of the minimum distances from each point in to the set . Therefore, the MASD is the mean of these two average surface distances, providing a symmetric measure of the average surface distance between the two boundaries.
NSD.
NSD evaluates the proportion of surface points on the predicted segmentation that lie within a specified distance from the ground truth surface, providing a normalized measure of boundary alignment, with a higher value indicating better alignment. The NSD between two surfaces and within a maximum tolerated distance is given by:
| (4) |
where: is the border region of within the maximum tolerated surface distance, which defines the set of points that lie within a distance from the surface . A point is considered to lie within this region if: . Similarly, is the border region of within the maximum tolerated distance . is the number of boundary points of that lie within the border region of . is the number of boundary points of that lie within the border region of . and denote the total number of boundary points in and , respectively. The NSD measures the proportion of boundary points of each surface that are within the maximum tolerated distance of the other surface’s border region.
For the NSD metric, we used , which aligns with the typical in-plane voxel spacing of our MRI data [20]. Specifically, the GE 1.5T scanner provided images with a resolution of 1.5 × 1.5 × (2.0–4.0) mm3, while the SIEMENS 3T scanner offered a uniform resolution of 2 × 2 × 2 mm3. Given this range, we chose the 2.0 mm threshold to reflect the most common voxel dimension across datasets, consistent with standard practice and our imaging protocol in [20]. Moreover, the segmented lesions in our neonatal brain MRI are relatively small. It is a clinically acceptable boundary tolerance for this specific task. Any deviation beyond this distance would likely be considered a significant segmentation error. A value of 2 mm for boundary tolerance in neonatal brain MRI lesion segmentation is clinically justified because this matches the minimum voxel size used in clinical diffusion MRI for HIE [54]. Other non-BONBID-related independent neonatal lesion segmentation studies also reported an average surface distance of at 2 mm or slightly above as the central indicator of credible segmentation quality by atlas-based approach [55], conventional machine learning [56], [57] or deep learning [58], directly linking deviations beyond this range with significant segmentation error and reduced clinical utility.
Handling of empty segmentation masks.
When the predicted mask is empty while the ground truth mask is non-empty, the MASD becomes undefined. Prior to evaluation, we reviewed all algorithms for instances of empty predicted masks: among the 14 algorithms, one produced a single empty mask, another produced three, and the remaining 12 each produced two. This indicates that all submitted algorithms occasionally generated empty predictions in this small and diffuse lesion segmentation task, even top performing methods. For these specific cases, a DICE score of zero and an NSD of zero were assigned. For MASD, we applied two computation strategies: (1) Mean value substitution: the missing value was replaced with the mean MASD computed from the respective algorithm’s valid predictions across the test set (Reported in Table IV); and (2) Max value substitution: to penalize the empty prediction masks, the missing value was replaced with the maximum MASD observed from that algorithm’s predictions across the test set (Reported in Table V). We chose this conservative strategy instead of imposing stronger penalties because the lesions are often small, and some patients may not present with visible lesions [20]; in such cases, empty outputs may represent clinically plausible scenarios.
TABLE V.
Performance of all teams and ZADC on MASD, where empty predictions were replaced by each algorithm’s maximum predicted MASD value (Mean ± SD).
| ID | Team | MASD |
|---|---|---|
| 1 | civalab | 2.7 ± 3.1 mm |
| 2 | xleratorxlerator9 | 2.9 ± 3.4 mm |
| 3 | frimpz | 3.4 ± 4.8 mm |
| 4 | schlauglab | 3.3 ± 4.1 mm |
| 5 | rajroy | 3.5 ± 4.9 mm |
| 6 | IWM | 3.5 ± 4.3 mm |
| 7 | imad.toubal | 2.9 ± 3.4 mm |
| 8 | UNeImage | 3.5 ± 4.2 mm |
| 9 | ngzvh | 3.6 ± 4.2 mm |
| 10 | ashwin_dhakal | 4.0 ± 4.7 mm |
| 11 | civa | 3.8 ± 4.0 mm |
| 12 | arda.aydn | 4.1 ± 4.1 mm |
| 13 | punithakumar | 7.2 ± 13.0 mm |
| 14 | tiansong_philips | 7.6 ± 12.9 mm |
| 15 | ZADC (−2) | 3.9 ± 4.6 mm |
Participation.
Different algorithms (typically in articles by different first authors [30]) were allowed, even if the authors were from the same research group. Multiple submissions of the same algorithms, e.g., differences only in the parameter settings, were not allowed.
Rankings.
The BONBID-HIE challenge employed a case-wise ranking method, which accounts for the significant variability in the complexity of patient cases. This ranking scheme has been successfully used in other challenges involving small and diffuse lesions, such as the subacute ischemic stroke lesion segmentation [21]. The evaluation process involved (1) computing the DICE, MASD, and NSD values for each case, (2) establishing each team’s rank based on these metrics across all cases, and then (3) calculating the mean rank over all three evaluation measures to determine the final team rankings.
III. Results
A. Submitted algorithms and leaderboard
During the challenge testing phase, 14 algorithms, software dockers, and method descriptions were submitted. All submitted dockers were evaluated on the hidden test set (N=44) on our challenge platform. Table III contains an overview of the methods submitted by the participating groups in the challenge (details are in the challenge proceedings [30] and challenge method supplementary file on our website github4). The submitted algorithms span a broad array of approaches leveraging anatomical information about HIE, data augmentation, training strategies, model architecture, and integration with traditional machine learning methods.
TABLE III.
Overview of all submitted methods on BONBID-HIE 2023 Lesion Segmentation Challenge. Details of each algorithm’s description can be found in supplementary file.
| Team color | Participants | Brief description of algorithms |
|---|---|---|
|
|
civalab [46] | Integrated Swin-UNETR with random forest |
|
|
xleratorxlerator9 [47] | Decoder denoising self-pretraining and finetuning |
|
|
frimpz [48] | Swin-UNETR |
|
|
schlauglab [49] | Heavy augmentation with 3-D ResUNet |
|
|
rajroy | Swin-UNETR |
|
|
IWM [50] | 3D ResUNet with heavy augmentation |
|
|
imad.toubal | Swin-UNETR |
|
|
UNetImage | Voxel specific logistic regression |
|
|
ngzvh | Swin-UNETR |
|
|
ashwin dhakal | 3D-UNet |
|
|
civa | 3D-UNet |
|
|
arda.aydn [51] | SegResNet with Reciprocal Transformation |
|
|
SVCC [52] | nnUNet |
|
|
tiansong philips | Swin-UNETR |
In summary, the top two methods exemplify different strategies for addressing the challenges of HIE lesion segmentation. Civalab’s approach integrates Swin-UNETR [62] with traditional machine learning method Random Forest, focusing on refining segmentation accuracy. xleratorxlerator9’s method emphasizes denoising and pretraining to enhance the quality of initial feature representations prior to fine-tuning for specific tasks. As indicated in the Table III, the majority of teams preferred advanced deep learning models, particularly Swin-UNETR and 3D-UNet variants, often paired with extensive data augmentation. Only the UNetImage team employed a traditional logistic regression model for lesion segmentation. The diversity of approaches reflects the ongoing experimentation and adaptation within the field, with some teams opting for traditional machine learning methods while others explored the potential of state-of-the-art deep learning methods.
Table IV summarizes the leaderboard rankings of all submitted algorithms. The ranking of the participating teams demonstrates a gradual improvement in the performance of the ranked approaches. For further evaluation of the impact of empty predictions, Table V reports the MASD results obtained when the maximum MASD from each algorithm’s valid predictions across the test set is used to represent the MASD for empty predictions. It is noteworthy that the variability in evaluation metrics across the teams does not differ significantly between any two sequentially ranked teams.
B. Winning Method
The Top 1 method civalab combined a Swin-UNETR with a random forest classifier [63]. In the first stage, the Swin-UNETR processed the 3D HIE images using two channels (ADC and ZADC) to generate a lesion probability map. The second stage involved a local refinement strategy: this probability map, alongside the input channels, was segmented into 5 × 5 2D windows and fed into a random forest classifier for a more localized lesion probability prediction. This two-stage approach, along with the use of the random forest for local refinement, played a role in mitigating the overfitting tendencies often associated with large parameter networks like Swin-UNETR, leading to more accurate and robust segmentations. Additionally, their proposed log Hausdorff distance loss aimed to regularize the 3D anatomical shape of HIE regions and implicitly optimize surface distance metrics.
The second-ranked method xleratorxlerator9 also employed a two-stage process, beginning with Label Aware Denoising Pretraining (LADP). In the first stage, a deep learning model was pretrained using a denoising method LADP, which strategically applied increasing levels of noise to regions surrounding lesion contours. This pretraining aimed to enable the model to learn more robust and relevant features specifically for distinguishing lesion from surrounding tissue. The second stage leveraged the learned representations from the pretraining phase for the downstream segmentation task. For complete method description, please refer to [47].
In summary, both top methods used a two-stage strategy, benefiting from the separation of feature learning and refinement. Civalab used a global-to-local prediction approach with a random forest, while xleratorxlerator9 applied lesion-focused denoising pretraining.
C. Impact of different lesion volumes
Table VI evaluates the performance of submitted methods on HIE lesion cases, categorized by lesion size: < 1%, [1%, 5%], and > 5% of the brain volume. Lesions smaller than 1% are particularly important as they account for over 50% of the cases in the HIE dataset, underscoring the need for algorithms capable of accurately detecting and segmenting these tiny, often diffuse, lesions. As shown in the table, current methods are more effective at segmenting larger lesions, which Lesion < 1% of the whole brain volume are generally easier to delineate, as evidenced by higher DICE scores, reaching up to 0.83 for lesions greater than 5%. In contrast, performance drops notably for smaller lesions, where even the top methods achieve DICE scores only in the 0.60–0.69 range. This trend is further illstrated that the best-performing methods for lesions smaller than 1% managed to achieve a DICE score of just around 0.50. The MASD and NSD metrics also reflect the increased complexity and variability associated with segmenting these small, diffuse lesions, indicating lower precision and accuracy. These results show the significant challenge of segmenting HIE lesions and highlight the pressing need for more advanced and refined algorithms capable of handling these difficult cases.
TABLE VI.
Evaluation of submitted methods on various HIE lesion cases.
| Participants | DICE (↑) | MASD(↓) | NSD(↑) |
|---|---|---|---|
| civalab | 50.0 ± 24.9% | 3.4 ± 3.0 mm | 66.9 ± 28.6% |
| xleratorxlerator9 | 49.0 ± 23.6% | 3.9 ± 3.0 mm | 64.4 ± 27.3% |
| frimpz | 44.5 ± 22.8% | 4.2 ± 4.2 mm | 62.5 ± 29.2% |
| schlauglab | 43.7 ± 24.1% | 4.0 ± 3.6 mm | 62.1 ± 28.6% |
| rajroy | 42.8 ± 22.1% | 4.3 ± 4.3 mm | 61.3 ± 29.3% |
| IWM | 41.8 ± 22.4% | 4.9 ± 4.3 mm | 59.1 ± 28.8% |
| imad.toubal | 38.3 ± 21.0% | 3.8 ± 3.1 mm | 59.3 ± 27.6% |
| UNeImage | 40.8 ± 22.4% | 4.9 ± 3.8 mm | 59.0 ± 27.4% |
| ngzvh | 38.4 ± 20.9% | 4.8 ± 3.9 mm | 56.9 ± 28.1% |
| ashwin_dhakal | 31.7 ± 17.9% | 5.5 ± 4.3 mm | 53.6 ± 25.5% |
| civa | 35.2 ± 22.6% | 5.5 ± 3.9 mm | 55.7 ± 24.8% |
| arda.aydn | 37.8 ± 21.2% | 4.8 ± 3.9 mm | 56.4 ± 27.2% |
| punithakumar | 35.8 ± 26.9% | 8.7 ± 12.3 mm | 51.0 ± 32.7% |
| tiansong_philips | 24.0 ± 17.6% | 7.3 ± 9.3 mm | 45.4 ± 28.4% |
| ZADC (−2) | 42.7 ± 24.3% | 5.2 ± 3.7 mm | 52.8 ± 25.6% |
| Lesion [1%, 5%] of the whole brain volume | |||
| civalab | 66.8 ± 15.8% | 1.6 ± 0.9 mm | 80.1 ± 15.2% |
| xleratorxlerator9 | 69.3 ± 14.2% | 1.4 ± 0.7 mm | 82.5 ± 11.8% |
| frimpz | 60.3 ± 13.9% | 1.7 ± 0.9 mm | 79.0 ± 13.0% |
| schlauglab | 63.5 ± 16.7% | 2.0 ± 1.3 mm | 77.6 ± 15.8% |
| rajroy | 62.2 ± 13.4% | 1.6 ± 0.8 mm | 80.5 ± 11.0% |
| IWM | 66.0 ± 12.2% | 1.5 ± 0.6 mm | 81.2 ± 9.5% |
| imad.toubal | 58.7 ± 12.7% | 1.6 ± 0.7 mm | 79.1 ± 11.4% |
| UNeImage | 64.5 ± 16.3% | 1.7 ± 0.7 mm | 80.0 ± 13.3% |
| ngzvh | 59.3 ± 14.4% | 1.8 ± 0.7 mm | 77.6 ± 12.9% |
| ashwin_dhakal | 55.1 ± 12.6% | 1.9 ± 0.8 mm | 76.1 ± 12.5% |
| civa | 52.1 ± 17.3% | 2.2 ± 1.0 mm | 72.6 ± 15.3% |
| arda.aydn | 55.4 ± 17.1% | 2.5 ± 1.1 mm | 69.8 ± 16.3% |
| punithakumar | 56.3 ± 18.7% | 2.7 ± 2.3 mm | 70.6 ± 17.7% |
| tiansong_philips | 47.4 ± 16.8% | 3.2 ± 2.3 mm | 64.7 ± 18.1% |
| ZADC (−2) | 67.4 ± 17.9% | 2.1 ± 0.8 mm | 76.9 ± 12.7% |
| Lesion > 5% of the whole brain volume | |||
| civalab | 83.4 ± 12.0% | 0.7 ± 0.4 mm | 89.6 ± 9.0% |
| xleratorxlerator9 | 83.3 ± 12.7% | 0.7 ± 0.5 mm | 88.8 ± 12.0% |
| frimpz | 82.3 ± 11.8% | 0.7 ± 0.4 mm | 90.5 ± 7.9% |
| schlauglab | 83.0 ± 12.2% | 0.7 ± 0.3 mm | 90.0 ± 6.7% |
| rajroy | 82.4 ± 11.8% | 0.7 ± 0.4 mm | 90.5 ± 7.9% |
| IWM | 82.7 ± 14.9% | 0.6 ± 0.4 mm | 90.3 ± 8.9% |
| imad.toubal | 80.5 ± 12.9% | 0.7 ± 0.3 mm | 89.8 ± 6.7% |
| UNeImage | 80.7 ± 14.4% | 0.7 ± 0.5 mm | 88.7 ± 11.0% |
| ngzvh | 81.0 ± 12.8% | 0.8 ± 0.3 mm | 89.2 ± 7.2% |
| ashwin_dhakal | 80.3 ± 13.8% | 0.8 ± 0.4 mm | 88.8 ± 8.1% |
| civa | 79.8 ± 13.5% | 0.8 ± 0.4 mm | 88.5 ± 7.6% |
| arda.aydn | 62.8 ± 21.1% | 2.1 ± 1.1 mm | 64.0 ± 21.2% |
| punithakumar | 73.7 ± 20.2% | 1.3 ± 0.8 mm | 78.8 ± 17.6% |
| tiansong_philips | 69.1 ± 20.5% | 1.4 ± 0.6 mm | 76.8 ± 14.1% |
| ZADC (−2) | 81.6 ± 14.5% | 0.8 ± 0.5 mm | 87.0 ± 12.6% |
D. Statistical analysis
In order to assess potential statistically significant performance differences across teams, we also performed a pairwise comparison. Each pair of methods was compared using the one-sided Wilcoxon signed-rank test [64], a robust nonparametric test designed to determine whether one method consistently outperforms another. As shown in Figure 6, in the DICE metric evaluation, “civalab” and “xleratorxlerator9” emerged as the top-performing methods. Statistical analysis revealed no significant difference between these two methods , indicating comparable performance. Furthermore, both methods demonstrated superiority over all 11 remaining methods. “Schlauglab” and “frimpz” were identified as the next highest-ranking methods, with “schlauglab” statistically outperforming seven other methods and “frimpz” showing superiority over the same number of competitors. However, both were statistically inferior to “civalab” and “xleratorxlerator9.” In the NSD metric analysis, “civalab” and “xleratorxlerator9” continued to dominate, while “schlauglab” was found to statistically outperform six other teams. Similarly, “frimpz” outperformed seven methods, with only one method showing better performance. This analysis highlights the dominance of “civalab” and “xleratorxlerator9” in DICE and the strength of “schlauglab” and “frimpz” in NSD, confirming their high rankings and consistent performance in the initial leaderboard.
Fig. 6.

Comparative analysis of 14 participating methods in terms of DICE and NSD metrics, evaluated using a one-sided Wilcoxon signed-rank test. Each node represents a method, with edges indicating statistically significant superiority of the originating method over the destination method. The number of outgoing and incoming edges, denoted as (#out / #in), reflects the relative strength of each method, with higher outdegrees and lower in-degrees signaling stronger performance. Node color saturation corresponds to this performance ratio, with more saturated colors highlighting more robust methods. Methods with identical edge counts exhibit comparable performance.
E. Impact of different field strength
Table II presents the distribution of patients in different lesion sizes across 1.5T GE and 3T SIEMENS scanners. Although efforts were made to split the test set with similar distributions, the SIEMENS scanner data contains nearly twice as many small and diffuse cases compared to the GE scanner data due to inherent data distribution. Cases acquired from different scanners can exhibit significant variations in appearance. A robust automatic HIE lesion segmentation method should effectively handle variations across scanners. To assess this, we evaluated each method’s performance separately for each scanner type and conducted a Mann–Whitney U test [65] to compare the results between scanners for each algorithm.
Figure 7 illustrates the performance variability of different teams’ algorithms, as measured by DICE scores. These methods show slightly better performance on GE data than on SIEMENS, as indicated by higher median DICE scores for several approaches, though this difference is not consistent across all methods. The boxplots demonstrate that the top two teams, “civalab” and “xleratorxlerator9”, achieve stable performance across both scanners, with similar median scores and interquartile ranges. Other algorithms display greater variability, possibly due to challenges in detecting small, diffuse lesions that are more common in 3T SIEMENS scans. However, Mann–Whitney U test results indicate that these differences are not statistically significant . Overall, the top two methods exhibit strong robustness across both 1.5T GE and 3T SIEMENS data, while the outliers in several methods suggest difficulties with specific lesion types or imaging conditions.
Fig. 7.

Boxplot of DICE scores across different teams’ algorithms on two types of scanners: GE 1.5T (red) and SIEMENS 3T (blue). The x-axis represents different teams, while the y-axis shows the DICE score values.
F. Combining the participants’ results by label fusion
To assess whether label fusion can enhance lesion segmentation results, we applied three typical algorithms: Majority Voting (MV) [59], which assigns a lesion label when the majority of algorithms concur; STAPLE [60], which calculates global weights for each algorithm to reduce the influence of less accurate algorithms; and Shape-Based Averaging (SBA) [61], which iteratively improves performance by excluding the least accurate algorithms compared to a weighted majority voting.
We conducted label fusion using the top 2, top 5, top 10, and all methods. Table IV (b) and Figure 5 illustrated that the SBA of the top 2 algorithms yielded the better DICE, MASD and NSD, offered marginal improvement over the top 2 methods in DICE. Additionally, when using MV, STAPLE, or SBA, the negative influence of multiple less accurate algorithms correlated segmentations resulted in lower accuracy compared to at least the two top-ranked algorithms. This observation, that ensemble results may not outperform the best single algorithm, aligns with findings in the segmentation of subacute stroke lesions [21], which are also small and diffuse.
Fig. 5.

Boxplot comparison of Dice scores across different label fusion strategies (MV, SBA, STAPLE) and individual team methods. The boxplots show the distribution of Dice scores for the top 2, top 5, top 10, and all algorithms, followed by each individual method.
G. Failure Cases
We categorize the failure cases of the top two algorithms and the majority voting results into five main categories, as illustrated in Figure 8: (1) Small and Diffuse Lesions: HIE cases often involve lesions affecting less than 1% of brain tissue, posing significant challenges due to extreme data imbalance; (2) Atypical Lesion Locations: lesions in less common regions are underrepresented in the training data, leading to suboptimal performance in these atypical cases. (3) Discrepancies of ZADC: variations in lesions can result in conflicting information between ZADC and ADC maps, leading to inaccurate segmentation inputs. (4) Uncertainty in Boundary Regions: ambiguous areas with inconsistent annotations, such as boundaries, making training more difficult and sometimes degrading segmentation performance. As shown in the figure, all methods struggle at boundary regions, where distinctions are challenging. (5) Composition of Multiple Factors: most algorithm failures stem from multiple concurrent issues.
Fig. 8.

Representative failure cases in HIE lesion segmentation. Each row illustrates a different challenge contributing to segmentation errors. Columns show the ADC map, ZADC map, ground truth (GT), and error maps from the top-1, top-2, and majority-vote (MV) methods. Error maps highlight true positives (yellow), false negatives (red), and false positives (blue), with expert annotations in green. These examples illustrate the complexity and ambiguity of neonatal HIE lesion segmentation and provide insights into common failure cases.
IV. Discussion
Performance of segmentation algorithms on HIE.
Despite advances in automated segmentating HIE lesions in our challenge, their performance and robustness still fall short of expectations (Overall DICE 0.62 and lesion < 1% DICE 0.5). Continuous improvement is expected through the expansion of training datasets to include larger and more diverse patient populations, along with innovations in training strategies and machine learning architectures. Future research should prioritize improving the robustness of automatic segmentation systems and addressing failure cases as discussed above, to achieve more reliable performance across varying lesion sizes in clinical HIE MRI settings.
Comparisons with ZADC maps.
We provided a baseline method based on clinical knowledge rather than deep learning algorithms. This approach applies a straightforward thresholding of ZADC maps at a value of −2. Despite the lack of complex modeling or machine learning algorithms, this method ranked 3rd when compared with other machine learning-based methods submitted to the challenge (Table IV). This result highlights three important insights: (1) Clinical Knowledge: The baseline method, which is purely based on clinical comprehension of ADC maps, demonstrates that conventional, knowledge-driven approaches can still be highly competitive in medical imaging tasks. Although machine learning approaches frequently produce innovation and enhanced accuracy, the efficacy of this more straightforward strategy emphasizes the importance of established clinical concepts in achieving successful outcomes. (2) Interpretability versus Complexity: A primary advantage of this method is its interpretability. In contrast to deep learning models, which frequently function as “black boxes”, such threshold-based method provides transparent and explicable criteria for its predictions. This transparency is a crucial benefit in medical contexts, where explainability is essential for clinical decisionmaking. (3) Potential for Hybrid Approaches: The success of this baseline method suggests that forthcoming research may gain from the integration of clinical expertise with machine learning methodologies. Hybrid models that integrate datadriven methodologies with domain-specific expertise may surpass solely machine learning-based methods while preserving a level of clinical interpretability.
Consensus of algorithms.
Consensus among algorithms generally yields higher performance, as demonstrated in segmentation tasks for large focal lesion regions such as the Brain Tumor Segmentation (BraTS) challenge [22]. In these cases, consensus methods often outperform the single best algorithm, whether in the segmentation of the whole tumor region, tumor core, or active tumor area. However, this improvement does not always hold for small and diffuse lesions, such as in the ISLES-SISS (Subacute Ischemic Stroke Segmentation) task [21]. In such cases, consensus among algorithms may not outperform the top-performing single method, which we have also observed in the HIE lesion segmentation task.
Evaluation metrics.
While we reported separate results for DICE, MASD, and NSD, and based our overall final ranking on these scores, these metrics may not adequately capture the performance of algorithms on small and diffuse lesions [53], [66]. However, there is no predefined metric that effectively evaluates small and diffuse lesions. DICE coefficient, although widely used, can disproportionately penalize slight mismatches in small regions, leading to lower scores even when the segmentation is clinically acceptable. Similarly, MASD and NSD, which rely on surface distance calculations, may struggle to accurately evaluate tiny or diffuse lesions where the boundary is not well-defined. Recent work on mesh-based metrics [67] compute distances directly in the mesh domain, which uses marching cubes or flying edges to generate surface meshes and accounts for boundary element sizes when computing distances. This approach may help reduce discretization artifacts and evaluate distance calculations more faithfully to reflect the underlying geometry of the segmentation. Incorporating mesh-based metrics in future evaluations may provide more reliable assessments for lesions with poorly defined boundaries.
Besides, the NSD metric is influenced by the choice of value [66]. Currently, we are using , although it matches the most common in-plane voxel size in our dataset, it also reflects a clinically acceptable tolerance for small neonatal brain lesions. However, when the tolerance parameter approaches the the same order of magnitude as the image resolution, the distances between predicted and reference boundaries become discretized [66]. In the current study, only a single set of expert annotations was available, so we could not determine based on inter-observer variability. Our ongoing work involves collecting multiple expert annotations from a larger multi-center dataset. In future work, we will determine by considering both the clinical tolerance for our lesion segmentation task and expert-to-expert surface distances, following the approach of Nikolov et al. [66]. We will also investigate how alternative values influence metric sensitivity and ranking stability.
Inter-observer variance.
One limitation of this study is the absence of multiple independent annotations for each subject. BONBID-HIE is the first public dataset for HIE lesion segmentation contains only a single annotation per subject, which was derived from a multi-expert consensus [20]. The lack of independent, multi-expert annotations hinders the ability to quantify intra- and inter-reader variability. Our ongoing multi-center work, spanning 21 sites and over 500 cases, incorporate multi-rater segmentations or consensus-driven labeling pipelines to address this limitation.
Clinical knowledge guided algorithms.
Most participating teams approached HIE lesion segmentation from an artificial intelligence perspective, utilizing state-of-the-art methods tailored for this task. Looking ahead, we encourage the formation of interdisciplinary teams that combine expertise from both the clinical and AI fields. By doing so, we aim to foster the development of algorithms that are guided by clinical knowledge, ensuring that the solutions are more aligned with clinical explanations, needs and practices. Besides, it remains to be elucidated how sex, age at MRI scan (in days), and other clinical factors will influence or contribute to lesion segmentation accuracy for HIE cases.
Beyond segmentation and toward outcome prediction.
In the second challenge, we further enhanced the clinical relevance of HIE lesion segmentation by integrating it with HIE outcome prediction. The BONBID-HIE 2023 challenge has concluded, but the benchmarking platform remains available for future research. A follow-up challenge, held in association with MICCAI 2024, introduced an outcome prediction track and a modified segmentation task, along with additional unannotated data to support unsupervised and semi-supervised approaches. The results of this challenge will be presented in a separate article. We have introduced a new track focused on using MRIs to predict 2-year neurocognitive outcomes. Moreover, MRI data alone may not be sufficient for accurate outcome prediction. Therefore, in our upcoming challenges over the next few years, we plan to release clinical information and encourage the research community to incorporate this alongside MRI data for outcome prediction tasks.
V. Summary and Conclusion
Fourteen state-of-the-art approaches provided valuable insights and highlighted the persistent challenges in accurately segmenting HIE lesions. Our findings demonstrate that while segmentation of HIE lesions in MRI is feasible, the results are still suboptimal, particularly for lesions occupying less than 1% of brain volume, where DICE scores averaged only 0.50. Overall, the average DICE score across all cases was 0.62, underscoring the inherent difficulty of this task. Looking ahead to the next iteration of BONBID-HIE, we plan to expand the dataset by including data from additional sites and to introduce a neurocognitive outcome prediction task, linking MRI findings with practical clinical outcomes. Ultimately, we hope to facilitate the integration of AI methods with MRI and clinical data into clinical practice and to promote more interdisciplinary research in this field.
Acknowledgements
This work was funded, in part, by the Harvard Medical School and Boston Children’s Hospital through Early Career Development Fellowship, Thrasher Research Fund Early Career Awards, NIH R21NS121735, R61NS126792, and R03HD104891.
We thank all the participants of the challenge and workshop for their valuable contributions and engagement.
Footnotes
References
- [1].Weiss Rebecca J, Bates Sara V, Song Ya’nan, Zhang Yue, Herzberg Emily M, Chen Yih-Chieh, Gong Maryann, Chien Isabel, Zhang Lily, Murphy Shawn N, et al. Mining multi-site clinical data to develop machine learning MRI biomarkers: application to neonatal hypoxic ischemic encephalopathy. Journal of translational medicine, 17(1):1–16, 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Graham Ernest M, Ruis Kristy A, Hartman Adam L, Northington Frances J, and Fox Harold E. A systematic review of the role of intrapartum hypoxia-ischemia in the causation of neonatal encephalopathy. American Journal of Obstetrics and Gynecology, 199(6):587–595, 2008. [DOI] [PubMed] [Google Scholar]
- [3].Lee Anne CC, Kozuki Naoko, Blencowe Hannah, Vos Theo, Bahalim Adil, Darmstadt Gary L, Niermeyer Susan, Ellis Matthew, Robertson Nicola J, Cousens Simon, et al. Intrapartum-related neonatal encephalopathy incidence and impairment at regional and global levels for 2010 with trends from 1990. Pediatric Research, 74(1):50–72, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Shankaran Seetha, Laptook Abbot R, Ehrenkranz Richard A, Tyson Jon E, McDonald Scott A, Donovan Edward F, Fanaroff Avroy A, Poole W Kenneth, Wright Linda L, Higgins Rosemary D, et al. Whole-body hypothermia for neonates with hypoxic-ischemic encephalopathy. New England Journal of Medicine, 353(15):1574–1584, 2005. [DOI] [PubMed] [Google Scholar]
- [5].Edwards A David, Brocklehurst Peter, Gunn Alistair J, Halliday Henry, Juszczak Edmund, Levene Malcolm, Strohm Brenda, Thoresen Marianne, Whitelaw Andrew, and Azzopardi Denis. Neurological outcomes at 18 months of age after moderate hypothermia for perinatal hypoxic ischaemic encephalopathy: synthesis and meta-analysis of trial data. Bmj, 340, 2010. [Google Scholar]
- [6].Azzopardi Denis V, Strohm Brenda, Edwards A David, Dyet Leigh, Halliday Henry L, Juszczak Edmund, Kapellou Olga, Levene Malcolm, Marlow Neil, Porter Emma, et al. Moderate hypothermia to treat perinatal asphyxial encephalopathy. New England Journal of Medicine, 361(14):1349–1358, 2009. [DOI] [PubMed] [Google Scholar]
- [7].Laptook Abbot R, Shankaran Seetha, Tyson Jon E, Munoz Breda, Bell Edward F, Goldberg Ronald N, Parikh Nehal A, Ambalavanan Namasivayam, Pedroza Claudia, Pappas Athina, et al. Effect of therapeutic hypothermia initiated after 6 hours of age on death or disability among newborns with hypoxic-ischemic encephalopathy: a randomized clinical trial. Jama, 318(16):1550–1560, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Shankaran Seetha, Laptook Abbot R, Pappas Athina, McDonald Scott A, Das Abhik, Tyson Jon E, Poindexter Brenda B, Schibler Kurt, Bell Edward F, Heyne Roy J, et al. Effect of depth and duration of cooling on death or disability at age 18 months among neonates with hypoxic-ischemic encephalopathy: a randomized clinical trial. Jama, 318(1):57–67, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [9].Liu Zulian, Xiong Tengbin, and Meads Catherine. Clinical effectiveness of treatment with hyperbaric oxygen for neonatal hypoxic-ischaemic encephalopathy: systematic review of chinese literature. Bmj, 333(7564):374, 2006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].Potter Molly, Rosenkrantz Ted, and Fitch R Holly. Behavioral and neuroanatomical outcomes in a rat model of preterm hypoxic-ischemic brain injury: effects of caffeine and hypothermia. International Journal of Developmental Neuroscience, 70:46–55, 2018. [DOI] [PubMed] [Google Scholar]
- [11].Nuñez-Ramiro Antonio, Benavente-Fernández Isabel, Valverde Eva, Cordeiro Malaika, Blanco Dorotea, Boix Hector, Cabañas Fernando, Chaffanel Mercedes, Fernández-Colomer Belén, Fernández-Lorenzo Jose Ramón, et al. Topiramate plus cooling for hypoxic-ischemic encephalopathy: a randomized, controlled, multicenter, double-blinded trial. Neonatology, 116(1):76–84, 2019. [DOI] [PubMed] [Google Scholar]
- [12].Liang Shi-Peng, Chen Qian, Cheng Yi-Bing, Xue Ying-Ying, and Wang Hai-Jun. Comparative effects of monosialoganglioside versus citicoline on apoptotic factor, neurological function and oxidative stress in newborns with hypoxic-ischemic encephalopathy. Coll Physicians Surg Pak, 29(4):324–327, 2019. [Google Scholar]
- [13].Cotten C Michael, Murtha Amy P, Goldberg Ronald N, Grotegut Chad A, Smith P Brian, Goldstein Ricki F, Fisher Kimberley A, Gustafson Kathryn E, Waters-Pick Barbara, Swamy Geeta K, et al. Feasibility of autologous cord blood cells for infants with hypoxic-ischemic encephalopathy. The Journal of pediatrics, 164(5):973–979, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [14].Laptook Abbot R, Shankaran Seetha, Barnes Patrick, Rollins Nancy, Do Barbara T, Parikh Nehal A, Hamrick Shannon, Hintz Susan R, Tyson Jon E, Bell Edward F, et al. Limitations of conventional magnetic resonance imaging as a predictor of death or disability following neonatal hypoxic-ischemic encephalopathy in the late hypothermia trial. The Journal of pediatrics, 230:106–111, 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Shankaran Seetha, McDonald Scott A, Laptook Abbot R, Hintz Susan R, Barnes Patrick D, Das Abhik, Pappas Athina, Higgins Rosemary D, Ehrenkranz Richard A, Goldberg Ronald N, et al. Neonatal magnetic resonance imaging pattern of brain injury as a biomarker of childhood outcomes following a trial of hypothermia for neonatal hypoxic-ischemic encephalopathy. The Journal of pediatrics, 167(5):987–993, 2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [16].Shankaran Seetha, Barnes Patrick D, Hintz Susan R, Laptook Abbott R, Zaterka-Baxter Kristin M, McDonald Scott A, Ehrenkranz Richard A, Walsh Michele C, Tyson Jon E, Donovan Edward F, et al. Brain injury following trial of hypothermia for neonatal hypoxic–ischaemic encephalopathy. Archives of Disease in Childhood-Fetal and Neonatal Edition, 97(6):F398–F404, 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17].Li Yi, Wisnowski Jessica L, Chalak Lina, Mathur Amit M, McKinstry Robert C, Licona Genesis, Mayock Dennis E, Chang Taeun, Van Meurs Krisa P, Wu Tai-Wei, et al. Mild hypoxic-ischemic encephalopathy (hie): timing and pattern of mri brain injury. Pediatric research, 92(6):1731–1736, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [18].Lee Bo Lyun, Gano Dawn, Rogers Elizabeth E, Xu Duan, Cox Stephany, Barkovich A James, Li Yi, Ferriero Donna M, and Glass Hannah C. Long-term cognitive outcomes in term newborns with watershed injury caused by neonatal encephalopathy. Pediatric research, 92(2):505–512, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].Weeke Lauren C, Groenendaal Floris, Mudigonda Kalyani, Blennow Mats, Lequin Maarten H, Meiners Linda C, van Haastert Ingrid C, Benders Manon J, Hallberg Boubou, and de Vries Linda S. A novel magnetic resonance imaging score predicts neurodevelopmental outcome after perinatal asphyxia and therapeutic hypothermia. The Journal of pediatrics, 192:33–40, 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [20].Bao Rina, Song Ya’nan, Bates Sara V, Weiss Rebecca J, Foster Anna N, Jaimes Camilo, Sotardi Susan, Zhang Yue, Hirschtick Randy L, Grant P Ellen, et al. Boston neonatal brain injury data for hypoxic ischemic encephalopathy (bonbid-hie): I. mri and lesion labeling. Scientific Data, 12(1):53, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Maier Oskar, Menze Bjoern H, Gablentz Janina Von der, Häni Levin, Heinrich Mattias P, Liebrand Matthias, Winzeck Stefan, Basit Abdul, Bentley Paul, Chen Liang, et al. Isles 2015-a public evaluation benchmark for ischemic stroke lesion segmentation from multispectral mri. Medical image analysis, 35:250–269, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Menze Bjoern H, Jakab Andras, Bauer Stefan, Kalpathy-Cramer Jayashree, Farahani Keyvan, Kirby Justin, Burren Yuliya, Porz Nicole, Slotboom Johannes, Wiest Roland, et al. The multimodal brain tumor image segmentation benchmark (brats). IEEE Transactions on Medical Imaging, 34(10):1993–2024, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [23].Menze Bjoern H, Jakab Andras, Bauer Stefan, Kalpathy-Cramer Jayashree, Farahani Keyvan, Kirby Justin, Burren Yuliya, Porz Nicole, Slotboom Johannes, Wiest Roland, et al. The multimodal brain tumor image segmentation benchmark (brats). IEEE transactions on medical imaging, 34(10):1993–2024, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Bakas Spyridon, Akbari Hamed, Sotiras Aristeidis, Bilello Michel, Rozycki Martin, Kirby Justin S, Freymann John B, Farahani Keyvan, and Davatzikos Christos. Advancing the cancer genome atlas glioma mri collections with expert segmentation labels and radiomic features. Scientific data, 4(1):1–13, 2017. [Google Scholar]
- [25].Jack Clifford R Jr, Bernstein Matt A, Fox Nick C, Thompson Paul, Alexander Gene, Harvey Danielle, Borowski Bret, Britson Paula J, Whitwell Jennifer L., Ward Chadwick, et al. The alzheimer’s disease neuroimaging initiative (adni): Mri methods. Journal of Magnetic Resonance Imaging: An Official Journal of the International Society for Magnetic Resonance in Medicine, 27(4):685–691, 2008. [Google Scholar]
- [26].Mueller Susanne G, Weiner Michael W, Thal Leon J, Petersen Ronald C, Jack Clifford R, Jagust William, Trojanowski John Q, Toga Arthur W, and Beckett Laurel. Ways toward an early diagnosis in alzheimer’s disease: the alzheimer’s disease neuroimaging initiative (adni). Alzheimer’s & Dementia, 1(1):55–66, 2005. [Google Scholar]
- [27].Winzeck Stefan, Hakim Arsany, McKinley Richard, Pinto José AADSR, Alves Victor, Silva Carlos, Pisov Maxim, Krivov Egor, Belyaev Mikhail, Monteiro Miguel, et al. Isles 2016 and 2017-benchmarking ischemic stroke lesion outcome prediction based on multispectral mri. Frontiers in neurology, 9:679, 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [28].Petzsche Moritz R Hernandez, de la Rosa Ezequiel, Hanning Uta, Wiest Roland, Valenzuela Waldo, Reyes Mauricio, Meyer Maria, Liew Sook-Lei, Kofler Florian, Ezhov Ivan, et al. Isles 2022: A multicenter magnetic resonance imaging stroke lesion segmentation dataset. Scientific data, 9(1):762, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [29].Bao Rina, Weiss Rebecca J, Bates Sara V, Song Ya’nan, He Sheng, Li Jingpeng, Bjornerud Alte, Hirschtick Randy L, Grant P Ellen, and Ou Yangming. Paradise: Personalized and regional adaptation for hie disease identification and segmentation. Medical Image Analysis, 102:103419, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [30].Bao Rina, Grant Ellen, Kirkpatrick Andrew, Wachs Juan, and Ou Yangming, editors. AI for Brain Lesion Detection and Trauma Video Action Recognition, volume 14567 of Lecture Notes in Computer Science, Cham, 2024. Springer. doi: 10.1007/978-3-031-71626-3. [DOI] [Google Scholar]
- [31].Tustison Nicholas J, Avants Brian B, Cook Philip A, Zheng Yuanjie, Egan Alexander, Yushkevich Paul A, and Gee James C. N4itk: improved n 3 bias correction. IEEE transactions on medical imaging, 29(6):1310–1320, 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [32].Ou Yangming, Lilla Zöllei Xiao Da, Retzepi Kallirroi, Murphy Shawn N, Gerstner Elizabeth R, Rosen Bruce R, Grant P Ellen, Kalpathy-Cramer Jayashree, and Gollub Randy L. Field of view normalization in multi-site brain mri. Neuroinformatics, 16:431–444, 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [33].Ou Yangming, Gollub Randy L, Retzepi Kallirroi, Reynolds Nathaniel, Pienaar Rudolph, Pieper Steve, Murphy Shawn N, Grant P Ellen, and Zöllei Lilla. Brain extraction in pediatric adc maps, toward characterizing neuro-development in multi-platform and multi-institution clinical images. NeuroImage, 122:246–261, 2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [34].Gaddamanugu Siddhartha, Shafaat Omid, Sotoudeh Houman, Sarrami Amir Hossein, Rezaei Ali, Saadatpour Zahra, and Singhal Aparna. Clinical applications of diffusion-weighted sequence in brain imaging: beyond stroke. Neuroradiology, 64(1):15–30, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [35].Lawrence Russell K and Inder Terrie E. Anatomic changes and imaging in assessing brain injury in the term infant. Clinics in perinatology, 35(4):679–693, 2008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [36].Vermeulen RJ, Fetter WPF, Hendrikx L, Van Schie PEM, Van Der Knaap MS, and Barkhof F. Diffusion-weighted mri in severe neonatal hypoxic ischaemia: the white cerebrum. Neuropediatrics, 34(02):72–76, 2003. [DOI] [PubMed] [Google Scholar]
- [37].Liauw Lishya, van Wezel-Meijler Gerda, Veen Sylvia, Van Buchem MA, and van der Grond Jeroen. Do apparent diffusion coefficient measurements predict outcome in children with neonatal hypoxic-ischemic encephalopathy? American Journal of Neuroradiology, 30(2):264–270, 2009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [38].Forbes Kirsten PN, Pipe James G, and Bird Roger. Neonatal hypoxic-ischemic encephalopathy: detection with diffusion-weighted mr imaging. American journal of neuroradiology, 21(8):1490–1496, 2000. [PMC free article] [PubMed] [Google Scholar]
- [39].Wolf Ronald L, Zimmerman Robert A, Clancy Robert, and Haselgrove John H. Quantitative apparent diffusion coefficient measurements in term neonates for early detection of hypoxic-ischemic brain injury: initial experience. Radiology, 218(3):825–833, 2001. [DOI] [PubMed] [Google Scholar]
- [40].Sayed Alaa A, Omar Nagham NM, Refaat Nafisa H, and Mahmoud Mohammed K. Role of diffusion-weighted magnetic resonance imaging in detection of neonatal hypoxic-ischemic encephalopathy. Journal of Current Medical Research and Practice, 5(1):115–120, 2020. [Google Scholar]
- [41].Huang Benjamin Y and Castillo Mauricio. Hypoxic-ischemic brain injury: imaging findings from birth to adulthood. Radiographics, 28(2):417–439, 2008. [DOI] [PubMed] [Google Scholar]
- [42].Cimpersek M, Meglic NP, Panjan DP, Skofljanec A, and Popovic KS. The role of diffusion weighted imaging and magnetic resonance imaging scoring system in assessing the effectiveness of treatment with hypothermia in neonates with hypoxic-ischemic encephalopathy. Neonat Pediatr Med, 3(135):2, 2017. [Google Scholar]
- [43].Ou Yangming, Zöllei Lilla, Retzepi Kallirroi, Castro Victor, Bates Sara V, Pieper Steve, Andriole Katherine P, Murphy Shawn N, Gollub Randy L, and Grant Patricia Ellen. Using clinically acquired mri to construct age-specific adc atlases: Quantifying spatiotemporal adc changes from birth to 6-year old. Human Brain Mapping, 38(6):3052–3068, 2017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [44].Douglas-Escobar Martha and Weiss Michael D. Hypoxic-ischemic encephalopathy: a review for the clinician. JAMA pediatrics, 169(4):397–403, 2015. [DOI] [PubMed] [Google Scholar]
- [45].Wei Ruili, Wang Chaonan, He Fangping, Hong Lirong, Zhang Jie, Bao Wangxiao, Meng Fangxia, and Luo Benyan. Prediction of poor outcome after hypoxic-ischemic brain injury by diffusion-weighted imaging: A systematic review and meta-analysis. Plos One, 14(12):e0226295, 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [46].Toubal Imad Eddine, Kazemi Elham Soltani, Rahmon Gani, Kucukpinar Taci, Almansour Mohamed, Ho Mai-Lan, and Palaniappan Kannappan. Fusion of deep and local features using random forests for neonatal hie segmentation. AI FOR BRAIN LESION DETECTION AND TRAUMA VIDEO ACTION RECOGNITION: First Bonbid, 14567:3, 2024. [Google Scholar]
- [47].Ninalga Dean. Label aware denoising pretraining. CMBES Proceedings, 46, 2024. [Google Scholar]
- [48].Kazemi Elham Soltani, Toubal Imad Eddine, Rahmon Gani, Kucukpinar Taci, Almansour Mohamed, Ho Mai-Lan, and Palaniappan Kannappan. Enhancing lesion segmentation in the bonbid-hie challenge: An ensemble strategy. AI FOR BRAIN LESION DETECTION AND TRAUMA VIDEO ACTION RECOGNITION, 14567:14, 2024. [Google Scholar]
- [49].Koirala Chiranjeewee Prasad, Mohapatra Sovesh, and Schlaug Gottfried. An ensemble approach for segmentation of neonatal hie lesions. In AI FOR BRAIN LESION DETECTION AND TRAUMA VIDEO ACTION RECOGNITION, pages 23–27. Springer, 2023. [Google Scholar]
- [50].Wodzinski Marek and Müller Henning. Improving segmentation of hypoxic ischemic encephalopathy lesions by heavy data augmentation: Contribution to the bonbid challenge. In AI FOR BRAIN LESION DETECTION AND TRAUMA VIDEO ACTION RECOGNITION, pages 28–33. Springer, 2023. [Google Scholar]
- [51].Aydın M Arda, Abdinli Elvin, and Unal Gozde. Segresnet based reciprocal transformation for bonbid-hie lesion segmentation. In booktitle=AI FOR BRAIN LESION DETECTION AND TRAUMA VIDEO ACTION RECOGNITION,, pages 39–44. Springer, 2023. [Google Scholar]
- [52].Tahmasebi Nazanin and Punithakumar Kumaradevan. A deep neural network approach for the lesion segmentation from neonatal brain magnetic resonance imaging. In AI FOR BRAIN LESION DETECTION AND TRAUMA VIDEO ACTION RECOGNITION, pages 34–38. Springer, 2023. [Google Scholar]
- [53].Reinke Annika, Tizabi Minu D, Sudre Carole H, Eisenmann Matthias, Rädsch Tim, Baumgartner Michael, Acion Laura, Antonelli Michela, Arbel Tal, Bakas Spyridon, et al. Common limitations of image processing metrics: A picture story. arXiv preprint arXiv:2104.05642, 2021. [Google Scholar]
- [54].Onda Kengo, Catenaccio Eva, Chotiyanonta Jill, Chavez-Valdez Raul, Meoded Avner, Soares Bruno P, Tekes Aylin, Spahic Harisa, Miller Sarah C, Parker Sarah-Jane, et al. Development of a composite diffusion tensor imaging score correlating with short-term neurological status in neonatal hypoxic–ischemic encephalopathy. Frontiers in Neuroscience, 16:931360, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [55].Beare Richard J, Chen Jian, Kelly Claire E, Alexopoulos Dimitrios, Smyser Christopher D, Rogers Cynthia E, Loh Wai Y, Matthews Lillian G, Cheong Jeanie LY, Spittle Alicia J, et al. Neonatal brain tissue classification with morphological adaptation and unified segmentation. Frontiers in neuroinformatics, 10:12, 2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [56].Weisenfeld Neil I and Warfield Simon K. Automatic segmentation of newborn brain MRI. Neuroimage, 47(2):564–572, 2009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [57].Yu Xintian, Zhang Yanjie, Lasky Robert E, Datta Sushmita, Parikh Nehal A, and Narayana Ponnada A. Comprehensive brain MRI segmentation in high risk preterm newborns. PloS one, 5(11):e13874, 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [58].Richter Leonie and Fetit Ahmed E. Accurate segmentation of neonatal brain MRI with deep learning. Frontiers in Neuroinformatics, 16:1006532, 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [59].Xu Lei, Krzyzak Adam, and Suen Ching Y. Methods of combining multiple classifiers and their applications to handwriting recognition. IEEE transactions on systems, man, and cybernetics, 22(3):418–435, 1992. [Google Scholar]
- [60].Warfield Simon K, Zou Kelly H, and Wells William M. Simultaneous truth and performance level estimation (STAPLE): an algorithm for the validation of image segmentation. IEEE transactions on medical imaging, 23(7):903–921, 2004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [61].Rohlfing Torsten and Maurer Calvin R Jr. Shape-based averaging for combination of multiple segmentations. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 838–845. Springer, 2005. [Google Scholar]
- [62].Hatamizadeh Ali, Nath Vishwesh, Tang Yucheng, Yang Dong, Roth Holger R, and Xu Daguang. Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images. In International MICCAI Brainlesion Workshop, pages 272–284. Springer, 2021. [Google Scholar]
- [63].Toubal Eddine, Kazemi Elham Soltani, Rahmon Gani, and Kucukpinar Taci. Fusion of deep and local features using random forests for neonatal hie segmentation. AI for Brain Lesion Detection and Trauma Video Action Recognition, pages 1–12, 2023. [Google Scholar]
- [64].Wilcoxon Frank. Individual comparisons by ranking methods. In Breakthroughs in statistics: Methodology and distribution, pages 196–202. Springer, 1992. [Google Scholar]
- [65].Fay Michael P and Proschan Michael A. Wilcoxon-mann-whitney or t-test? on assumptions for hypothesis tests and multiple interpretations of decision rules. Statistics surveys, 4:1, 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [66].Nikolov Stanislav, Blackwell Sam, Zverovitch Alexei, Mendes Ruheena, Livne Michelle, De Fauw Jeffrey, Patel Yojan, Meyer Clemens, Askham Harry, Romera-Paredes Bernardino, et al. Deep learning to achieve clinically applicable segmentation of head and neck anatomy for radio-therapy. arXiv preprint arXiv:1809.04430, 2018. [Google Scholar]
- [67].Podobnik Gašper and Vrtovec Tomaž. Metrics revolutions: Ground-breaking insights into the implementation of metrics for biomedical image segmentation. arXiv preprint arXiv:2410.02630, 2024. [Google Scholar]
