Skip to main content
RSNA Journals logoLink to RSNA Journals
. 2025 Jun 24;315(3):e242416. doi: 10.1148/radiol.242416

A Data-Centric Approach to Deep Learning for Brain Metastasis Analysis at MRI

Laurens Topff 1,2,✉, Liliana Petrychenko 1,2, Neeraj Jain 1,2,3, Sara Lingier 4, Jeroen Bertels 4, Patricio Astudillo 4, Milan Prosec 1, Pablo Menéndez Fernández-Miranda 5,6, Olivier Gevaert 7,8, Marion Smits 9,10, Sophie Derks 9,11,12, Eline Verhaak 13,14, Patrick E J Hanssens 13,14, Enrique Marco de Lucas 15,16, Rodrigo Sutil 15,16, Pablo D Dominguez 17, Adina Negoita 18, Ernst Visser 1, David Corral Fontecha 19, Loes M M Braun 1, Dieta Brandsma 20, Jacob J Visser 9, Erik R Ranschaert 21,22, Kevin B W Groot Lipman 1, Regina G H Beets-Tan 2,23
Editor: Sven Haller
PMCID: PMC12207652  PMID: 40552999

Abstract

Background

With the increasing incidence of brain metastases (BMs), artificial intelligence models have shown promise in assisting with the detection and volumetric analysis of lesions at MRI. However, current models are limited in identifying small lesions and lack generalizability.

Purpose

To develop a generalizable deep learning system for detecting, segmenting, and longitudinally tracking BMs of any size at MRI.

Materials and Methods

In this retrospective study, a data-centric approach to deep learning model development was used. A multicenter dataset was collected, comprising pre- and/or posttreatment MRI scans from patients with BMs and MRI scans from patients with cancer without BMs (December 2015 to August 2023). Iterative data annotation by radiologists with systematic quality control increased the consistency of reference segmentations. A modified nnU-Net framework, with robust data preprocessing and augmentation, was used. Lesion-wise detection metrics and segmentation performance, Dice similarity coefficient, and normalized surface distance were evaluated.

Results

In total, 1985 scans from 1623 patients (mean age, 62.0 years ± 12.2 [SD]; 743 female patients, 157 patients of unknown sex), with 5552 BMs, were included. BMs were present in 64.8% of the scans (1286 of 1985), 36.0% (463 of 1286) of which were posttreatment scans. The model was trained on 1451 scans acquired on 30 different scanners. In internal testing (n = 223), sensitivity was 98.0% (95% CI: 96.3, 99.0; 449 of 458 lesions). In external testing (n = 311), sensitivity was 97.4% (95% CI: 96.2, 98.2; 935 of 960; P = .58), with a mean of 0.6 false positives per patient. The sensitivity remained high for all lesion sizes, including those less than 3 mm in diameter (93.3% [95% CI: 89.1, 96.0]; 196 of 210). Median Dice similarity coefficient was 0.89 and 0.90 for the internal and external test datasets, respectively (P = .13). Median normalized surface distance was 0.99 for both datasets.

Conclusion

The deep learning system demonstrated high performance and generalizability in detecting and segmenting BMs of all sizes on pre- and posttreatment MRI scans.

© RSNA, 2025

Supplemental material is available for this article.


graphic file with name radiol.242416.VA.jpg


Summary

A data-centric approach prioritizing data quality enabled the development of a generalizable deep learning system that accurately detected and segmented brain metastases of all sizes on MRI scans.

Key Results

  • ■ In this retrospective, multicenter study (1623 patients, 1985 MRI scans), a deep learning system achieved high sensitivity (97.4%) for brain metastasis detection in external testing, including for the smallest lesions (<3 mm; 93.3%).

  • ■ False-positive detections were more frequent on scans with brain metastases than on negative scans (mean, 0.9 vs 0.1), yet most positive scans showed no false positives (median, 0).

  • ■ The model demonstrated strong segmentation performance in external testing (median Dice similarity coefficient, 0.90; median normalized surface distance, 0.99).

Introduction

Brain metastases (BMs) are the most common malignant brain tumors in adults and significantly impact morbidity and mortality in patients with advanced solid malignancies (1,2). Improvements in systemic therapies have led to more effective control of extracranial disease and have prolonged patient survival, resulting in an increased incidence of BMs over time (1). MRI plays an essential role in the diagnosis and management of BMs, with treatment decisions driven by the size, number, and location of lesions identified at imaging (2–4).

While radiologists are generally proficient in reading MRI scans for BMs, detection rates for small lesions tend to be lower (5,6), potentially leading to suboptimal staging and treatment planning. Stereotactic radiosurgery is increasingly performed in patients with multiple BMs, making the planning process substantially more time-consuming (3,7). Posttreatment scan interpretation is associated with additional challenges, particularly when multiple lesions are treated at different time points and exhibit heterogeneous responses.

Artificial intelligence (AI)–based image analysis has emerged as a promising tool to support radiologists in BM assessment. Multiple deep learning models have been developed for the detection and segmentation of BMs, achieving good results for large lesions (8,9). However, detecting lesions less than 6 mm in size remains difficult, with models demonstrating low sensitivity or high false-positive (FP) rates (8,9). Furthermore, studies have often been limited to single-center data or pretreatment scans (6,10–13). The substantial variability in MRI scans from different scanners and institutions necessitates the development of robust models that can perform reliably across various clinical settings.

Traditionally, the field of AI has prioritized improvements in neural network architectures and training techniques to improve model performance, an approach known as model-centric AI (14). However, there is growing recognition that large, diverse, and well-annotated datasets are crucial for successfully translating AI models from development to clinical impact (15,16). In recent years, data-centric AI has emerged as a new paradigm, shifting the focus from model refinement to data quality enhancement. The goal of this data-centric approach is to improve the generalizability and trustworthiness of AI models (14,15).

The aim of this study, which was based on data-centric principles, was to develop a generalizable deep learning system for the detection, segmentation, and longitudinal tracking of BMs of any size on MRI scans. The purpose of this system is to assist radiologists and other physicians in the diagnosis, treatment planning, and monitoring of BMs.

Materials and Methods

Study Design

A retrospective study was performed to develop and evaluate an imaging-based deep learning model for the detection and segmentation of BMs on MRI scans. In addition, the performance of existing registration software was evaluated for longitudinal analysis by coregistering consecutive scans and tracking lesion segmentations produced by the deep learning model. The study was approved by the institutional review boards of the participating institutions and was performed in compliance with the Health Insurance Portability and Accountability Act. The requirement for informed consent was waived. This study was a collaboration between academic institutions and an industry partner, Robovision, which provided technical support and access to the data annotation platform. Robovision did not provide financial support for the study. All study data and information were controlled by authors who were not employees of or consultants for Robovision.

Patients and Imaging

Adult patients (≥18 years old) with primary extracranial cancer who underwent brain MRI for screening, diagnosis, treatment planning, or posttreatment follow-up of BMs were eligible for inclusion. MRI scans of the brain consisted of three-dimensional contrast-enhanced T1-weighted series. The AI model was developed using data from a multicenter convenience sample of patients from the Netherlands Cancer Institute (March 2018 to February 2022), Stanford Hospital (April 2017 to November 2022), and Clínica Universidad de Navarra (February 2022 to January 2023) and public Mathematical Oncology Laboratory data from five institutions (17). Patients were exclusively assigned to either the training set or the internal test set via a stratified random split (85:15) (Table S1). Additionally, an external test set was collected from three centers: Hospital Universitario Marqués de Valdecilla (February 2018 to August 2023), Erasmus University Medical Center (March 2016 to August 2021), and Elisabeth-TweeSteden Hospital (December 2015 to January 2018). The sample from Erasmus University Medical Center was a convenience sample of patients with melanoma, and the sample from Elisabeth-TweeSteden Hospital was a previously reported sample of patients who underwent focal radiation therapy (18).

Medical exclusion criteria were meningeal metastases, more than 50 lesions, suspected radiation necrosis, prior brain surgery, or primary intracranial tumors. Technical exclusion criteria were a section thickness greater than 2 mm or a magnetic field strength less than 1.5 T. Nondiagnostic scans with severe artifacts or missing imaging data were also excluded.

Reference Standard

The total tumor volume of each BM was manually segmented. The annotation instructions underwent iterative refinement to clarify ambiguities (Appendix S1). A multistage process was followed for data annotation, with initial segmentation by a general radiologist (L.P., P.M.F.M., or A.N., with 1–4 years of experience) or a closely supervised medical student (M.P., E. Visser, or D.C.F.). The readers consulted radiology reports and follow-up scans for guidance and iteratively refined their segmentation. All cases were subsequently reviewed by a neuroradiologist (L.T., with 7 years of experience) to ensure the consistency of the segmentations. The test datasets were further reviewed by another neuroradiologist (N.J., with 9 years of experience). The reference segmentations of the test sets were created entirely through manual annotation, without any automated assistance. A collaborative platform (RVAI version 3.11; Robovision) was used for data curation and quality control. Small lesions were segmented using magnified views to allow precise delineation. Segmentation errors, such as stray voxels, were cleaned using scripts.

To assess interrater variability, two neuroradiologists (L.T. and N.J.) independently segmented 100 randomly selected scans with one to 10 lesions each from the external test set. Discrepancies were adjudicated by a third neuroradiologist (L.M.M.B., with 7 years of experience), leading to the final reference standard.

Data Preprocessing

The brain extraction algorithm HD-BET was applied (19), and the brain masks were dilated to ensure that small peripheral lesions were not removed. Furthermore, default preprocessing steps from the nnU-Net framework were employed, including resampling to a median voxel spacing of 0.5 × 0.5 × 1.0 mm (20). Further details are provided in Appendix S2.

Model

The BrainMets model was developed using a modified version of the three-dimensional full-resolution setup within the nnU-Net framework (20). The model was adapted to incorporate transformer blocks via the architecture of Vaswani et al (21) (Fig S1). The patch size was 128 × 192 × 80 voxels. Orthogonal initialization was used for the model weight matrices (22). Additional details are available in Appendix S3.

Training

Data augmentation was performed using the extensive augmentation techniques provided in the nnU-Net framework (20). Modifications to the training hyperparameters included reducing the initial learning rate from 0.01 to 0.005. Training was extended from 1000 to 2000 epochs. The batch size during training was two. An adapted loss function combining soft Dice loss with cross-entropy loss was chosen to improve lesion detection sensitivity. A voxelwise weighting scheme inversely correlated with lesion size was incorporated. Fivefold cross-validation was applied to the training set, with the weights of the final epoch used for testing. Additional details are available in Appendix S3.

The code is available in a public repository (https://github.com/nki-radiology/BrainMets). The repository includes a page where researchers can upload MRI scans for analysis using the BrainMets model.

Postprocessing

Predictions from each fold’s models were averaged, and a probability threshold of 0.5 was used to binarize the output. Lesions were defined via connected component analysis and assigned a confidence score by averaging the probabilities of each voxel. For the longitudinal analysis, consecutive scans were coregistered with Elastix version 5.1.0 (23), and overlapping lesions were matched (Appendix S4).

Evaluation

True-positive detections were defined as predicted lesions having any overlap with the reference mask. Lesion-wise sensitivity, positive predictive value, and F1 score are reported, stratified by lesion size groups based on maximum Feret diameter. To account for variability in the number of lesions per scan and the number of scans per patient, metrics were averaged per scan and subsequently per patient, as appropriate. Free-response receiver operating characteristic curves were generated. Additionally, the ability of the model to discriminate between positive and negative scans was assessed. Segmentation performance was evaluated for true-positive lesions using the Dice similarity coefficient (DSC) and normalized surface distance (NSD), with a tolerance distance of 1 mm. Bland-Altman plots were constructed for volumes.

The performance of the BrainMets model was compared with that of the baseline nnU-Net (20) and nnFormer (24) models. Additionally, the model was tested on the public UCSF-BMSR (University of California San Francisco Brain Metastases Stereotactic Radiosurgery) dataset via its independently created annotations (Appendix S5) (25).

Interrater variability was assessed by comparing reader segmentations and evaluating readers’ detection performance against the final reference standard.

Statistical Analysis

CIs were calculated using the Wilson-Brown method or bootstrapping. Dataset characteristics and results were compared using the χ2 test or the Mann-Whitney U test. Spearman correlation analysis was performed to assess relationships between predicted and reference lesion counts and volumes. Model performance was compared between BrainMets and the baseline models using the McNemar and Wilcoxon signed-rank tests, with P values corrected via the Holm-Bonferroni method. P < .05 was considered to indicate a statistically significant difference. Statistical analyses were performed by two authors (L.T. and K.B.W.G.L.) using Python (version 3.12; Python Software Foundation) and GraphPad Prism (version 10.2; GraphPad Software).

Results

Patient and Imaging Characteristics

A total of 5162 patients were evaluated for inclusion (Fig 1). Patient exclusions (1329 of 5162; 25.7%) were due to meningeal metastases (n = 444), technical image criteria (n = 264), prior brain surgery (n = 247), radiation necrosis (n = 154), having more than 50 lesions (n = 91), severe artifacts (n = 42), primary intracranial tumor (n = 32), age less than 18 years (n = 30), or missing imaging data (n = 25). After random selection of eligible patients, 1623 patients (mean age, 62.0 years ± 12.2 [SD]; 743 female patients, 723 male patients, 157 patients of unknown sex) with 1985 brain MRI scans were included. The most common primary tumor types were lung cancer, melanoma, and breast cancer (Table 1). BMs were present in 64.8% of the scans (1286 of 1985), 36.0% (463 of 1286) of which were posttreatment scans. Prior treatment of BMs included systemic therapy (45.4%, 210 of 463), radiation therapy (26.1%, 121 of 463), combined systemic and radiation therapy (24.2%, 112 of 463), or unspecified therapy (4.3%, 20 of 463).

Figure 1:

Flowchart of patient inclusion and exclusion. BM = brain metastasis, CUN = Clínica Universidad de Navarra, EMC = Erasmus University Medical Center, ETZ = Elisabeth-TweeSteden Hospital, HUMV = Hospital Universitario Marqués de Valdecilla, MOLAB = Mathematical Oncology Laboratory, NKI = Netherlands Cancer Institute.

Flowchart of patient inclusion and exclusion. BM = brain metastasis, CUN = Clínica Universidad de Navarra, EMC = Erasmus University Medical Center, ETZ = Elisabeth-TweeSteden Hospital, HUMV = Hospital Universitario Marqués de Valdecilla, MOLAB = Mathematical Oncology Laboratory, NKI = Netherlands Cancer Institute.

Table 1:

Patient and Imaging Characteristics by Dataset

Characteristic Training Set Internal Test Set External Test Set P Value*
Sites NKI, Stanford, CUN, MOLAB NKI, Stanford, CUN, MOLAB EMC, HUMV, ETZ
No. of patients 1159 204 260
Age (y) 63 (55–71) [19–94] 62 (54–72) [20–91] 62 (54–70) [19–85] .18
Sex .65
 Female 522 (45.0) 95 (46.6) 126 (48.5)
 Male 498 (43.0) 91 (44.6) 134 (51.5)
 Unknown 139 (12.0) 18 (8.8) 0 (0)
Primary cancer <.001
 Lung 426 (36.8) 67 (32.8) 66 (25.4)
 Melanoma 304 (26.2) 61 (29.9) 132 (50.8)
 Breast 120 (10.4) 23 (11.3) 16 (6.2)
 Genitourinary 55 (4.7) 10 (4.9) 19 (7.3)
 Gastrointestinal 52 (4.5) 13 (6.4) 9 (3.5)
 Other 53 (4.6) 9 (4.4) 17 (6.5)
 Unknown 149 (12.9) 21 (10.3) 1 (0.4)
No. of MRI scans 1451 223 311
Scans per patient 1 (1–1) [1–6] 1 (1–1) [1–4] 1 (1–1) [1–2] <.001
Presence of BMs <.001
 Positive scan 962 (66.3) 114 (51.1) 210 (67.5)
 Negative scan 489 (33.7) 109 (48.9) 101 (32.5)
Treatment status† .17
 Pretreatment scan 503 (52.3) 77 (67.5) 138 (65.7)
 Posttreatment scan 364 (37.8) 27 (23.7) 72 (34.3)
 Unknown 95 (9.9) 10 (8.8) 0 (0)
MRI field strength <.001
 3 T 772 (53.2) 131 (58.7) 132 (42.4)
 1.5 T 679 (46.8) 92 (41.3) 179 (57.6)
MRI scanner manufacturer <.001
 Philips 788 (54.3) 139 (62.3) 214 (68.8)
 Siemens 395 (27.2) 65 (29.1) 10 (3.2)
 General Electric 268 (18.5) 19 (8.5) 87 (28.0)
No. of MRI scanner models 30 17 15

Note.—Categorical data are presented as numbers of patients or scans, with percentages in parentheses, and continuous data are presented as medians, with IQRs in parentheses and ranges in brackets. BM = brain metastasis, CUN = Clínica Universidad de Navarra, EMC = Erasmus University Medical Center, ETZ = Elisabeth-TweeSteden Hospital, HUMV = Hospital Universitario Marqués de Valdecilla, MOLAB = Mathematical Oncology Laboratory, NKI = Netherlands Cancer Institute, Stanford = Stanford Hospital.

*

Comparison of the internal test set versus the external test set using the χ2 test for categorical variables and the Mann-Whitney U test for continuous variables. Missing values were omitted from the analyses.

†

Among patients with BMs.

The model was trained on 1451 MRI scans, acquired on 30 different types of scanners from three manufacturers (details in Table S2). The internal test set included 223 scans (114 with BMs, 109 without) from 204 patients, and the external test set included 311 scans (210 with BMs, 101 without) from 260 patients (Tables 1, S3).

Reference BM Characteristics

The radiologists identified 5552 BMs on 1286 scans, with an average of four lesions per scan (Tables 2, S4). There was high variability in lesion size, with diameters ranging from 1 to 95 mm, and the number of voxels per lesion ranged from two to 1.4 million. Lesions were grouped by diameter: less than 3 mm (21.2%; 1179 of 5552), 3 mm to less than 6 mm (28.6%; 1587 of 5552), 6 mm to less than 12 mm (23.4%; 1297 of 5552), and 12 mm or greater (26.8%; 1489 of 5552).

Table 2.

Brain Metastasis Characteristics by Dataset

Characteristic Training Set Internal Test Set External Test Set P Value*
No. of lesions 4134 458 960
Lesions per scan .54
 Mean 4 ± 6 4 ± 5 5 ± 6
 Median 2 (1–4) 2 (1–4) 2 (1–6)
 Range 1–50 1–27 1–47
Volume (mL) .009
 Mean 1.8 ± 6.7 1.9 ± 5.1 1.5 ± 5.1
 Median 0.1 (0.01–0.5) 0.1 (0.02–0.8) 0.1 (0.01–0.5)
 Range 0.001–110 0.001–40.5 0.002–79.8
Diameter (mm) <.001
 Mean 10.2 ± 10.8 11.8 ± 11.7 9.8 ± 9.7
 Median 5.9 (3.3–12.7) 7.1 (4–14.9) 6.1 (3.3–12.3)
 Range 1–84.6 1–94.9 1–66.9
Lesion size group .002
 Diameter < 3 mm 904 (21.9) 65 (14.2) 210 (21.9)
 Diameter 3 to <6 mm 1199 (29.0) 128 (27.9) 260 (27.1)
 Diameter 6 to <12 mm 936 (22.6) 117 (25.5) 244 (25.4)
 Diameter ≥ 12 mm 1095 (26.5) 148 (32.3) 246 (25.6)

Note.—Categorical data are presented as numbers of lesions, with percentages in parentheses. Continuous data are presented as means ± SDs, medians with IQRs in parentheses, and ranges.

*

Comparison of the internal test set versus the external test set using the χ2 test for categorical variables and the Mann-Whitney U test for continuous variables.

Model Detection of BMs

The BrainMets model showed a sensitivity of 98.0% (95% CI: 96.3, 99.0; 449 of 458 lesions) in the internal test set and 97.4% (95% CI: 96.2, 98.2; 935 of 960) in the external set (P = .58). The sensitivity was high for metastases of all sizes (Table 3), including the smallest lesions (<3 mm) (external test set, 93.3% [95% CI: 89.1, 96.0]; 196 of 210). Examples of small BMs detected by the model are shown in Figures 2, S2, and S3. The mean sensitivity per patient was 98.2% (95% CI: 96.9, 98.9) in the external test set, with a similar sensitivity for posttreatment versus pretreatment scans (98.1% vs 98.2%; P = .97) (Table S5). There was a strong correlation between the predicted and reference number of lesions per scan (ρ = 0.929–0.967; Fig S4). The lesion-wise positive predictive value and F1 score were 83.2% (95% CI: 80.9, 85.3; 935 of 1124) and 0.90, respectively, in external testing. The free-response receiver operating characteristic curve is shown in Figure 3A. A comparison of the results with those of previously published models is provided in Table S6.

Table 3.

Detection Performance of the BrainMets Model versus Baseline Models in Internal and External Testing

Metric Internal Test Set External Test Set
BrainMets nnU-Net nnFormer BrainMets nnU-Net nnFormer
No. of true positives by lesion size
 All sizes 449 427 403 935 872 811
 Diameter < 3 mm 60 50 36 196 157 105
 Diameter 3 to <6 mm 126 117 109 254 236 227
 Diameter 6 to <12 mm 115 113 112 239 234 233
 Diameter ≥ 12 mm 148 147 146 246 245 246
No. of false negatives by lesion size
 All sizes 9 31 55 25 88 149
 Diameter < 3 mm 5 15 29 14 53 105
Diameter 3 to <6 mm 2 11 19 6 24 33
 Diameter 6 to <12 mm 2 4 5 5 10 11
 Diameter ≥ 12 mm 0 1 2 0 1 0
Sensitivity (%) by lesion size
 All sizes 98.0 (96.3, 99.0) 93.2 (90.6, 95.2) 88.0 (84.7, 90.7) 97.4 (96.2, 98.2) 90.8 (88.8, 92.5) 84.5 (82.1, 86.6)
 Diameter < 3 mm 92.3 (83.2, 96.7) 76.9 (65.4, 85.5) 55.4 (43.3, 66.8) 93.3 (89.1, 96.0) 74.8 (68.5, 80.2) 50.0 (43.3, 56.7)
 Diameter 3 to <6 mm 98.4 (94.5, 99.7) 91.4 (85.3, 95.1) 85.2 (78.0, 90.3) 97.7 (95.1, 98.9) 90.8 (86.6, 93.7) 87.3 (82.7, 90.8)
 Diameter 6 to <12 mm 98.3 (94.0, 99.7) 96.6 (91.5, 98.7) 95.7 (90.4, 98.2) 98.0 (95.3, 99.1) 95.9 (92.6, 97.8) 95.5 (92.1, 97.5)
 Diameter ≥ 12 mm 100 (97.5, 100) 99.3 (96.3, 100) 98.6 (95.2, 99.8) 100 (98.5, 100) 99.6 (97.7, 100) 100 (98.5, 100)
Per-patient metrics
 Mean sensitivity per patient (%) 98.1 (95.6, 99.2) 96.1 (92.9, 97.7) 93.7 (90.8, 95.9) 98.2 (96.9, 98.9) 94.8 (92.1, 96.4) 91.6 (88.5, 93.8)
Mean no. of false positives per patient 0.7 (0.5, 1.0) 0.3 (0.3, 0.5) 0.2 (0.2, 0.3) 0.6 (0.5, 0.8) 0.3 (0.2, 0.3) 0.3 (0.2, 0.3)
 With BMs 1.3 (1.0, 1.8) 0.4 (0.3, 0.6) 0.2 (0.1, 0.3) 0.9 (0.7, 1.3) 0.3 (0.3, 0.4) 0.3 (0.2, 0.4)
 Without BMs 0.2 (0.1, 0.4) 0.3 (0.2, 0.4) 0.2 (0.2, 0.4) 0.1 (0.03, 0.1) 0.2 (0.1, 0.3) 0.3 (0.2, 0.4)
 Mean PPV per patient 79.9 (74.9, 84.2) 91.1 (87.9, 93.7) 94.1 (90.6, 96.4) 85.6 (82.5, 88.4) 91.5 (88.9, 93.7) 91.1 (88.1, 93.5)
 Mean F1 score per patient 0.86 (0.82, 0.89) 0.92 (0.90, 0.94) 0.93 (0.90, 0.95) 0.90 (0.88, 0.92) 0.91 (0.89, 0.93) 0.89 (0.86, 0.91)

Note.—True positives and false negatives are lesion counts; sensitivity data are presented as percentages; and the remaining data are presented as means of the mean per patient. Data in parentheses are 95% CIs. BM = brain metastasis, PPV = positive predictive value.

Figure 2:

Small brain metastases detected by the BrainMets model on axial contrast-enhanced T1-weighted MRI scans. Below each image are zoomed-in images of the area in the white box, showing the reference segmentation by radiologists (green outline) and the predicted segmentation by the BrainMets model (red outline). (A) Image in a 58-year-old male patient with melanoma reveals a small enhancing lesion, 1.7 mm in diameter, in the left parietal lobe. This newly detected brain metastasis comprises only six voxels. The patient had two additional small brain metastases (not shown). (B) Image in a 53-year-old male patient with melanoma demonstrates a single enhancing lesion of 2.7 mm in the left frontal lobe, consistent with a new solitary brain metastasis. (C) Image in an 80-year-old male patient with non–small cell lung cancer shows a new small enhancing lesion of 1.7 mm in the left occipital lobe. The patient had two other larger brain metastases (not shown).

Small brain metastases detected by the BrainMets model on axial contrast-enhanced T1-weighted MRI scans. Below each image are zoomed-in images of the area in the white box, showing the reference segmentation by radiologists (green outline) and the predicted segmentation by the BrainMets model (red outline). (A) Image in a 58-year-old male patient with melanoma reveals a small enhancing lesion, 1.7 mm in diameter, in the left parietal lobe. This newly detected brain metastasis comprises only six voxels. The patient had two additional small brain metastases (not shown). (B) Image in a 53-year-old male patient with melanoma demonstrates a single enhancing lesion of 2.7 mm in the left frontal lobe, consistent with a new solitary brain metastasis. (C) Image in an 80-year-old male patient with non–small cell lung cancer shows a new small enhancing lesion of 1.7 mm in the left occipital lobe. The patient had two other larger brain metastases (not shown).

Figure 3:

Brain metastasis detection performance of the BrainMets model. (A) Graph shows free-response receiver operating characteristic curves for the external test dataset overall and for each reference lesion diameter subgroup. (B) Graph shows the distribution of predicted lesion diameters for false positives in the internal and external test datasets. (C) Chart shows the classification of false positives in the internal and external test datasets. (D) Graph shows the effects of thresholding lesion confidence scores on the mean sensitivity and number of false positives per scan for the external test dataset. The data in A and D are from scans positive for brain metastases.

Brain metastasis detection performance of the BrainMets model. (A) Graph shows free-response receiver operating characteristic curves for the external test dataset overall and for each reference lesion diameter subgroup. (B) Graph shows the distribution of predicted lesion diameters for false positives in the internal and external test datasets. (C) Chart shows the classification of false positives in the internal and external test datasets. (D) Graph shows the effects of thresholding lesion confidence scores on the mean sensitivity and number of false positives per scan for the external test dataset. The data in A and D are from scans positive for brain metastases.

Failure Analysis

The false-negative metastases (n = 34) in both test sets were predominantly small (<3 mm) or faintly enhancing (detailed overview in Figs S5–S7). The mean number of FPs per patient was 0.7 and 0.6 in the internal and external test sets, respectively. The mean number of FPs was higher for scans with BMs than for those without BMs (0.9 vs 0.1, respectively). The median number of FPs for positive scans was zero (IQR, 0–1). The FPs were mostly small blood vessels (Figs 3B, 3C, S8, and S9).

FPs could be reduced with a limited effect on sensitivity by thresholding the confidence score (Fig 3D) or voxel count (Fig S10) of the predicted lesion. For example, when the confidence score threshold was increased from 0.5 to 0.8, the mean number of FPs per positive scan in the external test set was reduced from 0.9 to 0.6, with the mean sensitivity per scan decreasing slightly, from 98.2% to 98.1%.

Binary Scan-Level Classification

The model showed 100% sensitivity for binary classification of scans as positive or negative for the presence of BMs in both the internal (114 of 114 scans) and external (210 of 210) test sets (Fig 4). The binary scan-level specificities were 89% (97 of 109) and 94% (95 of 101) for the internal and external test sets, respectively.

Figure 4:

Scan-level prediction performance of the BrainMets model. The binary confusion matrices compare the predicted scan-level presence of brain metastases to the reference scan-level outcomes in the internal (A) and external (B) test datasets. A true-positive scan-level prediction was defined as detection of at least one true-positive lesion.

Scan-level prediction performance of the BrainMets model. The binary confusion matrices compare the predicted scan-level presence of brain metastases to the reference scan-level outcomes in the internal (A) and external (B) test datasets. A true-positive scan-level prediction was defined as detection of at least one true-positive lesion.

Model Segmentation of BMs

For the comparison of predicted segmentations with reference segmentations, the median Dice similarity coefficient was 0.89 (IQR, 0.77–0.94) for lesions in the internal test set and 0.90 (IQR, 0.80–0.94; P = .13) for lesions in the external test set; the median normalized surface distance was 0.99 (IQR, 0.94–1) for lesions in the internal test set and 0.99 (IQR, 0.95–1; P = .03) for lesions in the external test set (Table 4). Predicted and reference lesion volumes showed a strong correlation (ρ = 0.958–0.965) and were compared using a Bland-Altman plot (Fig 5). Longitudinal analysis of lesions enabled the calculation of volume growth, as illustrated in Figure 6 (detailed results in Appendix S4).

Table 4:

Segmentation Performance of the BrainMets Model versus Baseline Models in Internal and External Testing

Metric Internal Test Set External Test Set
BrainMets nnU-Net nnFormer BrainMets nnU-Net nnFormer
Dice similarity coefficient by lesion size
 All sizes 0.89 (0.77–0.94) 0.89 (0.80–0.94) 0.89 (0.82–0.94) 0.90 (0.80–0.94) 0.91 (0.83–0.95) 0.91 (0.84–0.94)
 Diameter < 3 mm 0.77 (0.63–0.84) 0.80 (0.55–0.86) 0.78 (0.66–0.84) 0.77 (0.60–0.87) 0.77 (0.62–0.89) 0.81 (0.73–0.88)
 Diameter 3 to <6 mm 0.80 (0.69–0.88) 0.82 (0.74–0.87) 0.84 (0.77–0.88) 0.86 (0.78–0.91) 0.88 (0.78–0.91) 0.87 (0.79–0.91)
 Diameter 6 to <12 mm 0.90 (0.85–0.93) 0.90 (0.85–0.94) 0.90 (0.85–0.93) 0.91 (0.87–0.94) 0.92 (0.88–0.94) 0.91 (0.88–0.93)
 Diameter ≥ 12 mm 0.94 (0.91–0.96) 0.95 (0.91–0.97) 0.94 (0.91–0.96) 0.95 (0.93–0.96) 0.96 (0.93–0.97) 0.95 (0.93–0.97)
Normalized surface distance by lesion size
 All sizes 0.99 (0.94–1) 0.99 (0.95–1) 0.99 (0.95–1) 0.99 (0.95–1) 0.99 (0.96–1) 0.99 (0.96–1)
 Diameter < 3 mm 1 (0.92–1) 1 (0.97–1) 1 (0.97–1) 1 (0.96–1) 1 (0.94–1) 1 (1–1)
 Diameter 3 to <6 mm 1 (0.96–1) 1 (0.96–1) 1 (0.97–1) 1 (0.98–1) 1 (0.98–1) 1 (0.98–1)
 Diameter 6 to <12 mm 1 (0.97–1) 1 (0.96–1) 1 (0.97–1) 0.99 (0.97–1) 1 (0.97–1) 0.99 (0.96–1)
 Diameter ≥ 12 mm 0.98 (0.92–1) 0.98 (0.92–1) 0.98 (0.91–0.99) 0.97 (0.92–0.99) 0.98 (0.94–0.99) 0.97 (0.94–0.99)

Note.—Data are presented as medians, with IQRs in parentheses. Segmentation metrics are relative to reference segmentations by radiologists and were calculated only for true-positive lesion detections, which varied in number across models (as shown in Table 3). The normalized surface distance had a tolerance distance of 1 mm.

Figure 5:

Segmentation performance and volume analysis of the BrainMets model. Data represent detected brain metastases in the internal (n = 449; green) and external (n = 935; red) test sets. (A) Scatterplot of Dice similarity coefficient versus reference lesion volume on a logarithmic scale. (B) Scatterplot of predicted versus reference lesion volume. The dashed line represents the identity line (y = x). The Spearman correlation coefficient (ρ) is shown for each dataset. (C) Bland-Altman plot comparing predicted and reference lesion volume. The mean difference was 0.10 mL (95% limits of agreement: −2.49, 2.70 [green lines]) for the internal test set and −0.02 mL (95% limits of agreement: −3.13, 3.09 [red lines]) for the external test set. The black line represents zero mean difference. The x-axis is on a logarithmic scale. An expanded view (right) shows the data for mean lesion volumes less than 0.1 mL.

Segmentation performance and volume analysis of the BrainMets model. Data represent detected brain metastases in the internal (n = 449; green) and external (n = 935; red) test sets. (A) Scatterplot of Dice similarity coefficient versus reference lesion volume on a logarithmic scale. (B) Scatterplot of predicted versus reference lesion volume. The dashed line represents the identity line (y = x). The Spearman correlation coefficient (ρ) is shown for each dataset. (C) Bland-Altman plot comparing predicted and reference lesion volume. The mean difference was 0.10 mL (95% limits of agreement: −2.49, 2.70 [green lines]) for the internal test set and −0.02 mL (95% limits of agreement: −3.13, 3.09 [red lines]) for the external test set. The black line represents zero mean difference. The x-axis is on a logarithmic scale. An expanded view (right) shows the data for mean lesion volumes less than 0.1 mL.

Figure 6:

Automated volumetric assessment and longitudinal analysis of brain metastases in a 72-year-old male patient with non–small cell lung cancer. The BrainMets system processed both the pretreatment MRI scan for stereotactic radiosurgery planning and the 3-month posttreatment scan. Lesions were automatically matched between scans. The top images are three-dimensional renderings of the brain with model-predicted lesion segmentations at both time points. A mixed response is evident: Metastases with decreased volume after treatment are shown in blue, metastases with increased volume are shown in red, and a new lesion is shown in yellow. The bottom images are example predicted segmentations (colored outlines) for three lesions on matched pre- and posttreatment axial contrast-enhanced T1-weighted MRI sections. The BrainMets system correctly detected all brain metastases on both the pretreatment (n = 8) and posttreatment (n = 9) scans. Reference segmentations are not shown. DSC = Dice similarity coefficient, NSD = normalized surface distance, Vol. = predicted volume.

Automated volumetric assessment and longitudinal analysis of brain metastases in a 72-year-old male patient with non–small cell lung cancer. The BrainMets system processed both the pretreatment MRI scan for stereotactic radiosurgery planning and the 3-month posttreatment scan. Lesions were automatically matched between scans. The top images are three-dimensional renderings of the brain with model-predicted lesion segmentations at both time points. A mixed response is evident: Metastases with decreased volume after treatment are shown in blue, metastases with increased volume are shown in red, and a new lesion is shown in yellow. The bottom images are example predicted segmentations (colored outlines) for three lesions on matched pre- and posttreatment axial contrast-enhanced T1-weighted MRI sections. The BrainMets system correctly detected all brain metastases on both the pretreatment (n = 8) and posttreatment (n = 9) scans. Reference segmentations are not shown. DSC = Dice similarity coefficient, NSD = normalized surface distance, Vol. = predicted volume.

Comparison with Baseline Models

The sensitivity of the BrainMets model (97.4%; 935 of 960 lesions) was greater than that of the baseline nnU-Net model (90.8%; 872 of 960 lesions; corrected P < .001) and the nnFormer model (84.5%; 811 of 960 lesions; corrected P < .001) in external testing (Tables 3, S7). However, the BrainMets model resulted in more FPs per patient than both the baseline nnU-Net model and the nnFormer model (mean, 0.6 vs 0.3 for both; corrected P < .001 and P = .01, respectively). The segmentation performance was mostly similar (Table 4).

Model Evaluation Using Public Data

Further evaluation was conducted using the public UCSF-BMSR dataset, containing 324 pre- and posttreatment scans from 223 patients (25). This sample had a higher mean lesion count per scan, at 10 (range, 1–127), and included meningeal metastases. The BrainMets model showed a sensitivity of 89.0% (95% CI: 87.8, 90.0; 2979 of 3349 lesions), with a mean sensitivity per patient of 90.6% (95% CI: 88.3, 92.3) (Table S8). The mean and median number of FPs per patient were 2.6 and one, respectively. The median Dice similarity coefficient and normalized surface distance were 0.78 (IQR, 0.67–0.85) and 0.91 (IQR, 0.75–0.98), respectively (Table S9). Example lesion results, including apparent FPs, which may represent lesions overlooked in the reference segmentations, are shown in Figures S11–S13.

Interrater Variability

Neuroradiologist A demonstrated 98.0% sensitivity (95% CI: 96.1, 99.0; 388 of 396 lesions) and neuroradiologist B demonstrated 87.9% sensitivity (95% CI: 84.3, 90.7; 348 of 396; P < .001) compared to the final (adjudicated) reference standard. The BrainMets model achieved a sensitivity of 98.0% (95% CI: 96.1, 99.0; 388 of 396) on the same sample, that is, no evidence of a difference in sensitivity from neuroradiologist A (P > .99) and higher sensitivity than neuroradiologist B (P < .001). The median interrater Dice similarity coefficient for the two neuroradiologists was 0.91 (IQR, 0.86–0.94); the model showed a Dice similarity coefficient of 0.91 (IQR, 0.84–0.95) compared with the final reference standard on the same sample.

Discussion

The BrainMets model demonstrated strong generalizability across a multicenter dataset, with high detection and segmentation performance for brain metastases (BMs) of all sizes. At external testing, the model showed a sensitivity of 97.4%, maintaining a high sensitivity of 93.3% for the smallest lesions (<3 mm in diameter). These results were achieved by prioritizing data quality throughout model development. This data-centric approach contrasts with a traditional model-centric approach, which focuses on optimizing model architecture and training techniques. While the aim of the model-centric strategy is to find the best possible model for a given dataset, the data-centric approach of our study was focused on creating the highest-quality dataset for model training. A heterogeneous dataset comprising 1985 scans from over 30 different MRI scanners was curated to capture the variability of BMs and scanning protocols encountered in clinical practice. Rigorous data annotation by radiologists and systematic cleaning procedures were implemented to improve the consistency of segmentation, with particular attention given to the smallest metastases. Radiologists’ domain expertise was important to this data-centric approach, contributing to data representativeness, the design of the comprehensive annotation scheme, and the error analysis approach. In addition, the performance of the model on small metastases was improved by using the extensive data preprocessing and augmentation techniques of the nnU-Net framework (20), enhancing the model architecture with transformer blocks, and enhancing the training objective with a size-dependent weighted loss function.

The performance of the BrainMets model appears to surpass that of previous models, particularly in detecting small lesions, which are clinically relevant for treatment planning but are often challenging to identify. To our knowledge, only Yin et al (26) achieved comparably high sensitivity, reporting values of 88.9%–95.5% for their model on external data, with 89% sensitivity for lesions 3 mm or smaller. However, their model was limited to detection without segmentation, was trained on single-center data, and included only newly diagnosed lesions. The high sensitivity of the BrainMets model was accompanied by false detections, averaging 0.6 FPs per patient, with higher rates for positive scans than negative scans (0.9 vs 0.1). This FP rate is lower than or comparable to those reported in most previous studies (8). Notably, the median number of FPs for positive scans was zero, indicating that most scans had no FPs. These results show that while the model performs well, radiologist review of AI output remains essential. Most FPs were small blood vessels that can be easily dismissed. However, there is a risk that some AI findings may be indeterminate, and follow-up imaging may be required. In a study by Luo et al (27), AI assistance increased readers’ detection sensitivity for BMs but also led to more FPs, particularly for less-experienced readers.

Comparing BrainMets model results with published model results presents challenges due to dataset variability; benchmarking on public datasets facilitates objective comparison. However, certain public datasets exhibit missing or incomplete lesion segmentation (28–30), while others show reduced image quality due to preprocessing (31), limiting their suitability for small metastasis detection. The BrainMets model was evaluated on the public UCSF-BMSR dataset (25) and achieved high performance, with a mean sensitivity per patient of 90.6%, which was lower than that on our collected data. The FP rate of the BrainMets model in the UCSF-BMSR dataset was also higher than that in our collected data. These differences are primarily attributable to the public dataset having a higher lesion count per scan and some errors in the reference segmentations, rather than representing limited generalizability. Nevertheless, this evaluation provides valuable insight into the performance of the model on independently annotated data and emphasizes the importance of high-quality public datasets for fair model comparisons.

While previous studies have explored the incorporation of multiple imaging sequences as model inputs to improve detection performance, such as combining gradient-echo and spin-echo sequences (32) or gradient-echo and black-blood imaging (10), these approaches may limit generalizability because of their reliance on specialized sequences not routinely used in many hospitals. To increase applicability, we used only a single three-dimensional contrast-enhanced T1-weighted sequence as input in our study. The use of three-dimensional imaging facilitated the precise registration of scans from different time points, which is crucial for longitudinal tracking of small lesion volumes. Volumetric assessment of BMs has demonstrated greater consistency than diameter measurements (33) and may enable earlier detection of disease progression during treatment response assessment (34). The BrainMets model achieved high segmentation performance (Dice similarity coefficient in the external test set, 0.90) and demonstrated accurate lesion tracking, supporting its potential for facilitating volumetric analysis and treatment monitoring.

The BrainMets model was developed as a supportive tool to aid radiologists during MRI scan interpretation, requiring seamless integration of its AI output into clinical workflows. Radiologists must be able to efficiently interact with AI-generated findings, accept true lesions, and reject FP detections. The BrainMets model outputs lesion measurements and segmentations in machine-readable, standard-based formats (including Digital Imaging and Communications in Medicine Structured Report, or DICOM SR) for integration in clinical viewers (35). Accepted findings can then automatically populate radiology reports in compatible reporting systems.

Our study has several limitations. First, the retrospective nature of the study and the nonconsecutive patient selection may have introduced selection bias. Second, the exclusion of patients with meningeal metastases, radiation necrosis, and resection cavities may limit the model’s applicability and lead to an underestimation of FP rates in real-world scenarios. The exclusion criteria were chosen to optimize model training for its primary task, namely, detecting parenchymal BMs. Future development could incorporate the excluded conditions through additional training data and adapted brain extraction algorithms to retain extra-axial regions. However, expanding the scope of the model would require careful consideration of the increased complexity of the data annotation process and potential trade-offs in detection performance. Finally, the study did not assess the impact of the model on radiologist performance. Prospective studies are needed to evaluate the utility of the model in clinical workflows and its effects on patient outcomes.

In conclusion, a generalizable deep learning system for detecting and segmenting brain metastases on both pre- and posttreatment MRI scans was developed via a data-centric approach. The ability of the model to accurately detect small metastases and longitudinally track lesion volumes makes it a promising tool to assist physicians in the management of brain metastases.

Acknowledgments

Acknowledgments

We would like to express our sincere gratitude to Robovision for its collaboration on this project, including providing the data annotation platform that was instrumental in this research. Special thanks go to Stephane Willaert, MSc, MBA, of Robovision for his invaluable support throughout the project. We also extend our appreciation to ImFusion for supplying the medical image viewer integrated into the data annotation platform, which greatly facilitated our work. Additionally, we thank neuroradiologist Joppe Schneiders, MD, PhD, for his valuable input on the annotation scheme.

Funding: Authors declared no funding for this work.

Disclosures of conflicts of interest: L.T. Valorization of the software developed in this study is being considered, for which institution may receive future royalties. L.P. No relevant relationships. N.J. No relevant relationships. S.L. No relevant relationships. J.B. Employee of Robovision. P.A. Support for the present manuscript from Robovision. M.P. No relevant relationships. P.M.F.M. No relevant relationships. O.G. Grants to institution from the National Cancer Institute, AstraZeneca, Food and Drug Administration, Saudi Company for Artificial Intelligence, Owkin, Onc. AI, University of California, Berkeley, and Roche Molecular Systems; honorarium for lecture from Princess Margaret Cancer Center, University of Toronto; support for attending meetings and/or travel from AZ Delta, Indonesia Ministry of Health, and Global AI Summit; and patent S21-177, provisional patent S22-425, and provisional patent S24-079. M.S. Consulting fees paid to institution from Bracco and board member of the Dutch Society of Radiology (unpaid) and European Society of Radiology (unpaid). S.D. No relevant relationships. E. Verhaak No relevant relationships. P.E.J.H. No relevant relationships. E.M.d.L. No relevant relationships. R.S. No relevant relationships. P.D.D. No relevant relationships. A.N. No relevant relationships. E. Visser No relevant relationships. D.C.F. No relevant relationships. L.M.M.B. No relevant relationships. D.B. No relevant relationships. J.J.V. Grants or contracts to institution from Qure.ai, Enlitic, Promedius, AstraZeneca, and Philips; payment or honoraria to institution for lectures, presentations, speakers bureaus, manuscript writing, or educational events from Roche; support for attending meetings or travel from Qure.ai; participation on a data and safety monitoring board or advisory board for Quibim, Contextflow, and Noaber Foundation; chair of the scientific committee of the European Society of Medical Imaging Informatics (unpaid); chair of the European Society of Radiology Value-based Radiology Subcommittee (unpaid); junior editor for European Journal of Radiology (unpaid); member of the RSNA Common Data Elements Steering Committee (unpaid); and stock or stock options from Quibim and Contextflow. E.R.R. No relevant relationships. K.B.W.G.L. Public private partnership grant (TKI-LSH) from Health-Holland with ScreenPoint Medical, and chairman of the Young Club of the European Society of Medical Imaging Informatics. R.G.H.B.T. No relevant relationships.

Abbreviations:

AI
artificial intelligence
BM
brain metastasis
FP
false positive

References

  • 1. Lamba N , Wen PY , Aizer AA . Epidemiology of brain metastases and leptomeningeal disease . Neuro Oncol 2021. ; 23 ( 9 ): 1447 – 1456 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Aizer AA , Lamba N , Ahluwalia MS , et al . Brain metastases: a Society for Neuro-Oncology (SNO) consensus review on current management and future directions . Neuro Oncol 2022. ; 24 ( 10 ): 1613 – 1646 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Le Rhun E , Guckenberger M , Smits M , et al. ; EANO Executive Board and ESMO Guidelines Committee . EANO–ESMO clinical practice guidelines for diagnosis, treatment and follow-up of patients with brain metastasis from solid tumours . Ann Oncol 2021. ; 32 ( 11 ): 1332 – 1347 . [DOI] [PubMed] [Google Scholar]
  • 4. Derks SHAE , van der Veldt AAM , Smits M . Brain metastases: the role of clinical imaging . Br J Radiol 2022. ; 95 ( 1130 ): 20210944 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Qu J , Zhang W , Shu X , et al . Construction and evaluation of a gated high-resolution neural network for automatic brain metastasis detection and segmentation . Eur Radiol 2023. ; 33 ( 10 ): 6648 – 6658 . [DOI] [PubMed] [Google Scholar]
  • 6. Rudie JD , Weiss DA , Colby JB , et al . Three-dimensional U-Net convolutional neural network for detection and segmentation of intracranial metastases . Radiol Artif Intell 2021. ; 3 ( 3 ): e200204 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Rozati H , Chen J , Williams M . Overall survival following stereotactic radiosurgery for ten or more brain metastases: a systematic review and meta-analysis . BMC Cancer 2023. ; 23 ( 1 ): 1004 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Wang TW , Hsu MS , Lee WK , et al . Brain metastasis tumor segmentation and detection using deep learning algorithms: a systematic review and meta-analysis . Radiother Oncol 2024. ; 190 : 110007 . [DOI] [PubMed] [Google Scholar]
  • 9. Ozkara BB , Chen MM , Federau C , et al . Deep learning for detecting brain metastases on MRI: a systematic review and meta-analysis . Cancers (Basel) 2023. ; 15 ( 2 ): 334 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Park YW , Jun Y , Lee Y , et al . Robust performance of deep learning for automatic detection and segmentation of brain metastases using three-dimensional black-blood and three-dimensional gradient echo imaging . Eur Radiol 2021. ; 31 ( 9 ): 6686 – 6695 . [DOI] [PubMed] [Google Scholar]
  • 11. Ziyaee H , Cardenas CE , Yeboa DN , et al . Automated brain metastases segmentation with a deep dive into false-positive detection . Adv Radiat Oncol 2023. ; 8 ( 1 ): 101085 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Li R , Guo Y , Zhao Z , et al . MRI-based two-stage deep learning model for automatic detection and segmentation of brain metastases . Eur Radiol 2023. ; 33 ( 5 ): 3521 – 3531 . [DOI] [PubMed] [Google Scholar]
  • 13. Liang Y , Lee K , Bovi JA , et al . Deep learning-based automatic detection of brain metastases in heterogenous multi-institutional magnetic resonance imaging sets: an exploratory analysis of NRG-CC001 . Int J Radiat Oncol Biol Phys 2022. ; 114 ( 3 ): 529 – 536 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Liang W , Tadesse GA , Ho D , et al . Advances, challenges and opportunities in creating data for trustworthy AI . Nat Mach Intell 2022. ; 4 ( 8 ): 669 – 677 . [Google Scholar]
  • 15. Zhang A , Xing L , Zou J , Wu JC . Shifting machine learning for healthcare from development to deployment and from models to data . Nat Biomed Eng 2022. ; 6 ( 12 ): 1330 – 1345 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Kann BH , Hosny A , Aerts HJWL . Artificial intelligence for clinical oncology . Cancer Cell 2021. ; 39 ( 7 ): 916 – 927 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Ocaña-Tienda B , Pérez-Beteta J , Villanueva-García JD , et al . A comprehensive dataset of annotated brain metastasis MR images with clinical and radiomic data . Sci Data 2023. ; 10 ( 1 ): 208 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Verhaak E , Schimmel WCM , Gehring K , Emons WHM , Hanssens PEJ , Sitskoorn MM . Health-related quality of life after Gamma Knife radiosurgery in patients with 1–10 brain metastases . J Cancer Res Clin Oncol 2021. ; 147 ( 4 ): 1157 – 1167 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Isensee F , Schell M , Pflueger I , et al . Automated brain extraction of multisequence MRI using artificial neural networks . Hum Brain Mapp 2019. ; 40 ( 17 ): 4952 – 4964 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Isensee F , Jaeger PF , Kohl SAA , Petersen J , Maier-Hein KH . nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation . Nat Methods 2021. ; 18 ( 2 ): 203 – 211 . [DOI] [PubMed] [Google Scholar]
  • 21. Vaswani A , Shazeer N , Parmar N , et al . Attention is all you need . arXiv 1706.03762 [preprint] https://arxiv.org/abs/1706.03762. Posted June 12, 2017. Updated August 2, 2023. Accessed July 3, 2023 . [Google Scholar]
  • 22. Saxe AM , McClelland JL , Ganguli S . Exact solutions to the nonlinear dynamics of learning in deep linear neural networks . arXiv 1312.6120 [preprint] https://arxiv.org/abs/1312.6120. Posted December 20, 2013. Updated February 19, 2014. Accessed July 3, 2023 .
  • 23. Klein S , Staring M , Murphy K , Viergever MA , Pluim JPW . elastix: a toolbox for intensity-based medical image registration . IEEE Trans Med Imaging 2010. ; 29 ( 1 ): 196 – 205 . [DOI] [PubMed] [Google Scholar]
  • 24. Zhou HY , Guo J , Zhang Y , Yu L , Wang L , Yu Y . nnFormer: interleaved transformer for volumetric segmentation . arXiv 2109.03201 [preprint] https://arxiv.org/abs/2109.03201. Posted September 7, 2021. Updated February 4, 2022. Accessed July 3, 2023 .
  • 25. Rudie JD , Saluja R , Weiss DA , et al . The University of California San Francisco Brain Metastases Stereotactic Radiosurgery (UCSF-BMSR) MRI Dataset . Radiol Artif Intell 2024. ; 6 ( 2 ): e230126 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Yin S , Luo X , Yang Y , et al . Development and validation of a deep-learning model for detecting brain metastases on 3D post-contrast MRI: a multi-center multi-reader evaluation study . Neuro Oncol 2022. ; 24 ( 9 ): 1559 – 1570 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Luo X , Yang Y , Yin S , et al . False-negative and false-positive outcomes of computer-aided detection on brain metastasis: secondary analysis of a multicenter, multireader study . Neuro Oncol 2023. ; 25 ( 3 ): 544 – 556 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Grøvik E , Yi D , Iv M , Tong E , Rubin D , Zaharchuk G . Deep learning enables automatic detection and segmentation of brain metastases on multisequence MRI . J Magn Reson Imaging 2020. ; 51 ( 1 ): 175 – 182 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Wang Y , Duggar WN , Caballero DM , et al . A brain MRI dataset and baseline evaluations for tumor recurrence prediction after Gamma Knife radiotherapy . Sci Data 2023. ; 10 ( 1 ): 785 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Link KE , Schnurman Z , Liu C , et al . Longitudinal deep neural networks for assessing metastatic brain cancer on a large open benchmark . Nat Commun 2024. ; 15 ( 1 ): 8170 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Moawad AW , Janas A , Baid U , et al . The Brain Tumor Segmentation (BraTS-METS) Challenge 2023: brain metastasis segmentation on pre-treatment MRI . arXiv 2306.00838 [preprint] https://arxiv.org/abs/2306.00838. Posted June 1, 2023. Updated December 9, 2024. Accessed August 10, 2023 .
  • 32. Yun S , Park JE , Kim N , Park SY , Kim HS . Reducing false positives in deep learning–based brain metastasis detection by using both gradient-echo and spin-echo contrast-enhanced MRI: validation in a multi-center diagnostic cohort . Eur Radiol 2024. ; 34 ( 5 ): 2873 – 2884 . [DOI] [PubMed] [Google Scholar]
  • 33. Bauknecht HC , Romano VC , Rogalla P , et al . Intra- and interobserver variability of linear and volumetric measurements of brain metastases using contrast-enhanced magnetic resonance imaging . Invest Radiol 2010. ; 45 ( 1 ): 49 – 56 . [DOI] [PubMed] [Google Scholar]
  • 34. Ocaña-Tienda B , Pérez-Beteta J , Romero-Rosales JA , et al . Volumetric analysis: rethinking brain metastases response assessment . Neurooncol Adv 2024. ; 6 ( 1 ): vdad161 . [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Tejani AS , Cook TS , Hussain M , Sippel Schmidt T , O’Donnell KP . Integrating and adopting AI in the radiology workflow: a primer for standards and Integrating the Healthcare Enterprise (IHE) profiles . Radiology 2024. ; 311 ( 3 ): e232653 . [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Radiology are provided here courtesy of Radiological Society of North America

RESOURCES